# CLOSED-ITEMS — finished work, compressed > **What this is.** Every register row that reached a terminal state, compressed to its title, the > version it shipped in, its evidence paths, and any sentence that states a RULE rather than a > narrative. **Nothing was deleted:** each entry names the commit that holds its full original text, > and `git show :documentation/backlog/OPEN-ITEMS.md` returns it verbatim. > > **Why it exists (operator ruling, 2026-08-22).** `OPEN-ITEMS.md` had grown to 672 KB across 286 > entries, over half of it finished work, with one single entry at 16 KB. A file that cannot be read > is a file that cannot be checked — and this project has already paid for that twice: a record > nobody could find because it sat inside an entry about something else, and a finding rediscovered > because nobody could see it. The register now holds **open work only**, so its size tracks the work > rather than the project's age. > > **A sibling rather than the bottom of the register**, deliberately: appending to the same file keeps > the byte count and the scroll, which is the thing being fixed. > > **Load-bearing reasoning was NOT compressed away.** Where a closed row states a rule, a fence or a > deliberate refusal, that sentence is carried here verbatim under **Reasoning kept**. Rules that > outlive their work item also live in their proper homes — `workspace-CLAUDE.md` standing rules, > `felhom.eu/CLAUDE.md`, `CONTEXT.md`, and the architecture folder — and this file is not their > primary record. > > **This file is not the register.** Nothing here is open. `OPEN-ITEMS.md` remains the single source > of truth for open work; `scripts/one_register_gate.py` enforces that against `ROADMAP.md`. --- | **R-361** | **The pre-restore safety dump overwrote the app's own DB dump, and the comment beside it said it could not.** Shipped in controller v0.221.0 (+v0.221.1). Evidence: `audits/DRILL-r361-2026-08-22/evidence/`. **Reasoning kept:** *`DumpOne` writes `-.sql` — the app's canonical dump, the name the replay loop matches EXACTLY — so nothing else may ever be written to it.* The fix is a DESTINATION, not a rename: `DumpOneTo` takes the final path and derives its own `.tmp` from it, so neither the destination nor the scratch file can collide with a nightly dump running beside it. **`DumpOne`'s signature did not move** — it has callers outside this concern. **The manifest no longer lists the undo copies:** every consumer of `Manifest.DBDumps` was grepped and named — three, all inside `recovery_unit.go`, none reading it for recovery. **AND THAT CHANGE MADE ANOTHER UNREACHABLE:** a stable `db_dumps` let `CaptureRecoveryUnit`'s already-current early return fire, and the undo-copy prune sat after it — four copies on disk against a cap of three, counted live. The prune now runs ABOVE the check; it is housekeeping on the dump directory and is independent of whether the manifest needs rewriting. **PROVEN LIVE the only way it can be:** the canonical dump's sha256, unchanged across a restore — `docmost` `5d35678349bb…`, `bookstack` `7837aa5de295…`, both byte-identical before and after. A test asserting merely that the undo copy exists passes just as well when the app's backup was destroyed. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.221.1, 2026-08-23) | full text: `git show a8caa0fdde7c:documentation/backlog/OPEN-ITEMS.md` | | **R-379** | **The pre-restore undo copy was valid, was named to the customer, and no product action could apply it.** Shipped in controller v0.220.0 (+v0.220.1, v0.220.2). Evidence: `audits/DRILL-r379-rollback-2026-08-22/evidence/`. **Reasoning kept:** *R-379 and R-380 were ONE failure with ONE fix — both ended with a half-restored database and the only difference was whether it looked broken.* **The undo set is matched on THE RUN'S OWN STAMP, never on the `pre-restore-` prefix** (four copies coexisted on one app in one afternoon; a prefix match replays an arbitrary older state) **and never just the first file** (a two-database app would have had one restored and the other left half-written). **The rollback RE-DISCOVERS the container** — the undo file is stable, the container is not: the DB-only start re-creates it, and v0.220.0's own first live run held an app for 30 s of `waitDBReady` against a dead id while its data was recoverable. **No unit test saw that: they all inject the import seam and never look at container identity.** | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.220.1, 2026-08-22; docmost and bookstack both rolled back to byte-identical prior state) | full text: `git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md` | | **R-380** | **A failed MariaDB replay left a partially-applied database behind an app reporting `health=healthy`.** Shipped in controller v0.220.0. Evidence: `audits/DRILL-r379-rollback-2026-08-22/evidence/13-step2-verify.txt`. **Reasoning kept:** **no engine flag closes this** — `--single-transaction` was added to the Postgres import and does make it all-or-nothing, but **MariaDB's DDL is not transactional**, so a partial apply there is unavoidable at the engine. The flag is a belt; the rollback is the fix, and this row must not be read as saying otherwise. Proven live: `bookstack`'s `migrations` table back at **102 rows**, the exact cell the defect was measured in. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.220.0, 2026-08-22) | full text: `git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md` | | **R-381** | **The restore-failure message pasted raw engine stderr — including rows out of the customer's own database — into the Hungarian customer surface.** Shipped in controller v0.220.0. **Reasoning kept:** the full engine text now goes to the operator log, **which never had it before — the diagnostic was ADDED, not removed**. Measured: 407 bytes (Postgres) and 615 (MariaDB, whose middle was an `INSERT INTO migrations VALUES (…)` listing); now 257 bytes with no engine tokens. **A red-proof for this PASSED and the test was hollow**: it injected below `ImportDump`, so a leak reintroduced inside `ImportDump` could not fail it. The guard now sits at that layer. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.220.0, 2026-08-22) | full text: `git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md` | | **R-382** | **The reconstitution's summary log line omitted the volume count it already held.** Shipped in controller v0.220.0. Proven live: `0 file(s) placed, 3 volume(s) replayed, 1 DB dump(s) replayed`. | **CLOSED — SHIPPED** (controller v0.220.0, 2026-08-22) | full text: `git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md` | | **R-356** | **The off-site restore refused every app that has no data drive — it asked "does this app have an HDD path?" to answer "is this app installed?", and for 40 of 53 catalogue apps the honest answer to the first is permanently no.** Shipped in controller v0.219.0. Evidence: `audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/`. **Reasoning kept:** *the restore destination is resolved by the SAME rule as the capture destination — the drive if the app has one, the system data path otherwise (`Manager.GetAppDrivePath`, one expression). The refusal that protects a drive app from being restored onto the wrong disk applies to apps that HAVE a drive to get wrong.* **An app with no drive is not misconfigured** — `01-topology-and-trust.md` §8 carries the `[DESIGN]` marker; between 19 and 22 August that design was called a defect four times. **Deployment is asked of `ListDeployedStacks()` and FAILS CLOSED on a nil provider:** "cannot tell" must not become "go ahead" when the caller's next act is a write. **Two different failures get two different sentences** — installed-but-no-resolvable-data-root has its own refusal and its own route; widening `nincs telepítve` to cover it would send a customer to reinstall a running app and hide the real fault. **Measured, and load-bearing: 53 templates, 13 `needs_hdd: true`, 40 `false`** (catalogue @ `459766cb1639`). **The capture side's raw `GetStackHDDPath` is FENCED and was not changed** — capture resolves an app's declared `userdata`/`import` file legs against that value, and a system-data fallback there would write a snapshot claiming to hold files it does not. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.219.0, 2026-08-22; `privatebin` on `demo-hp`: planted, backed up, deleted, restored, 15/15 files byte-identical including two Hungarian accented names) | full text: `git show e18668f9e19f:documentation/backlog/OPEN-ITEMS.md` | | **R-216** | **A correct recovery code was reported to the customer as wrong.** Shipped in 0.120.0, v0.125.0. | **SHIPPED** (controller v0.201.0 + hub v0.97.0/0.97.1) — **but see R-223**: the feature does not work on a NEW box until the manifest vouches agent 0.125.0. Until then such a box is correctly HELD, not lied to | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-218** | **Succeeding at recovery stopped the box asking for what it still needed.** Shipped in v0.203.0. Evidence: `documentation/tests/part4-rewalk-2026-08-06/journal.md`. | **CLOSED 2026-08-06 — controller v0.203.0, proven live.** *(State corrected 2026-08-06: this field read REOPENED while the body below already recorded the fix shipped and proven. The history of the over-claim is kept deliberately — it is why the row is worded as it is.)* **The over-claim, as it stood: the fix covered the DECLARATION half only.** Measured on the R-201 re-walk: the box declared, and **`offsiteheal` re-staged the secret at 11:44:57** saying *"the box re-consumes on its next cycle"* — **the next cycle came and went** (`host-report` 11:55:46, `Received report` 11:55:54, a full cycle **with a positive control that it ran**) **and the credential was still not consumed.** 23 minutes after the re-stage the box's last off-site-apply attempt was still the pre-re-stage one. A census of the customer-reachable actions on `/backups/remote` (`config`, `reset`, `run`, `toggle`) found **none that fetches a staged credential**, and the only lever is `systemctl restart felhom-controller-bootstrap.service` **inside the guest** — which worked in **18 s** (Campaign 11 measured 17), confirming nothing was wrong with the credential, the target or the key: **the only thing missing is anything at all to trigger a retry.** **This is the FIRST of the two dead ends that keep the recovery journey failing** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-219** | **The listing the screen promises could never render on the shape it exists for.** | **SHIPPED** (controller v0.201.0) — the unlock now places the key, brings the tier up, then lists | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-217** | **An unreadable store reported as "opened, with unattributable content".** | **SHIPPED** (controller v0.201.0) — opened / empty / unreadable are three distinguishable states | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-222** | **Reaching for a RETAINED earlier package read as a wrong code.** | **SHIPPED** (controller v0.201.0 + hub v0.97.0) — the ACK carries `superseded_present`/`superseded_at` and the screen names the situation. **It states what the hub knows and promises nothing** — the read path is still unbuilt (R-199's inventory) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-215** | **`GET /recovery` rendered the recovery story on a box that never had off-site backups.** | **SHIPPED** (controller v0.201.0) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-220** | **After a rebuild the customer's drives cannot be re-enrolled, and the refusal names an impossible action.** | **CLOSED 2026-08-06 — shipped in agent v0.127.0 and PROVEN LIVE on a genuinely rebuilt box.** The fix is **corroborated, not a widened prefix**: a mountpoint outside `/mnt/felhom-drives` is forgiven only when the SAME device is also mounted under the managed path — a pairing only Felhom's own enrolment produces, so a disk another system is using at `/srv/data` or even `/mnt/someone-elses-disk` is still refused (own test + red-proof). Read from `/proc/mounts` deliberately: the lsblk invocation is pinned verbatim in the sudoers file, so switching to plural `MOUNTPOINTS` would have shipped a sudoers change with the binary. Fail-safe: an unreadable mount table corroborates nothing. **Measured on the Part 4 venue after a real guest purge, with both raw mounts still present on the surviving host:** `/disks/candidates` returned both drives in `attach` and `initialize` (before the fix: two empty lists), and both **re-attached through the customer endpoint** (`registered: true`). The customer-facing refusal was corrected in controller v0.203.0. | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-221** | **A rebuilt box cannot run the escrow ceremony at all.** | **CLOSED 2026-08-08 — agent v0.128.0.** `Apply` re-asserts the seed BEFORE the idempotent early return; the return itself is kept and pinned by a zero-Proxmox-calls assertion. **The writer was established at `file:line` rather than assumed** — see the follow-through section below | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-223** | **The Day-0 manifest vouched agent 0.120.0 while the recovery feature needs 0.125.0 — and a reinstall DOWNGRADES a box that was fixed by hand.** Shipped in 0.120.0, 0.125.0, 0.192.0. | **CLOSED 2026-08-05.** Golden **0.201.0** baked in the drill VM (658 165 766 B, sha `e730d7cab343eb35…f007654`, **round-trip verified from Gitea**), then manifest set in one save: `agent=0.125.0 golden=0.201.0 min_agent=0.125.0`. A fresh install now lands on current agent AND current controller | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-224** | **Every non-code failure on the unlock path is reported to the customer as a statement about their code.** **Reasoning kept:** **And R-216's gate cannot catch it**: the box's own ring reads `recovery capability gate: offsite_key_recovery=yes (source=version)` — the gate discriminates the agent's **age**, not its **reachability**, so a dead agent of the right version sails through the guard whose own comment says *"An attemp | **CLOSED 2026-08-06 — controller v0.202.0 + agent v0.126.0.** The discriminator is now a VALUE: `escrow.ErrBundleFetch` → **HTTP 502** at the agent, `agentapi.RecoveryRefusal` carrying the status at the controller, and `ClassifyRecoveryFailure` mapping it to one of five classes **from the value, never the text**. **PROVEN LIVE on the venue**, same wrong code, only the hub's reachability changed: `hub up → 400 "…did not open the sealed bundle"` · `hub REJECTed → 502 "…could not be fetched — the recovery code was NOT used"` · `hub restored → 400`. Red-proof: deleting the agent case reproduces `got 400, want 502` with the wrong-code sentence. **Coupled `MinAgent 0.126.0`** — an older agent answers 400 for both causes, so the reading is withheld and the 400 degrades to NEUTRAL; the gate blocks nothing. **The customer-facing messages were NOT re-driven end-to-end**: `/recovery` correctly redirects since F7 set the old data aside, and restoring that state is the reconfiguration §11 forbids — they are covered by handler tests + red-proofs | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-225** | **The remote store reports `0 pillanatkép · 0 / 50 GB` when the box cannot read it — directly above a card stating the store holds backups.** | **CLOSED 2026-08-06 — controller v0.202.0.** `StatsKnown` is a **named** state (the `OffsiteInventory.Empty` pattern), because zero is what an unread store and an empty one both look like and `omitempty` makes "absent" and "0" the same bytes. The fill bar renders only when the fill is known — a 0 %-wide bar is a picture of emptiness. **PROVEN LIVE both ways**: before a run the venue read „a pillanatképek száma még ismeretlen"; after one, „2 pillanatkép … / 50 GB". A measured zero still says zero | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-226** | **M1 — the only message that tells a customer to check their typing — is unreachable on any box that has re-escrowed.** | **CLOSED 2026-08-06 — controller v0.202.0.** The retained-package message now names **both** possibilities and restores the ten-words prompt, because the two are indistinguishable at the engine and saying so is the honest thing. It still does not promise the earlier package can be opened. Red-proof: removing the clause makes the prompt unreachable again | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-228** | **After „I do not want the old data", the set-aside history becomes invisible — the box records where it is and shows it to nobody.** | **CLOSED 2026-08-06 — controller v0.202.0.** `OrphanedRenamedTo` is surfaced as two facts and stops. **It does not promise the history can be reopened** — it cannot be, by anyone, today (R-199's inventory is unbuilt) — and the set-aside **confirmation copy was corrected** for the same reason: *"a helyreállítási kód nélkül többé nem lesznek megnyithatók"* implied that WITH the code they could be. The field's own comment said "recovery-code-recoverable", the same over-promise in the code. **PROVEN LIVE**: the notice renders on the venue | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-227** | **A controller restart mid-unlock returns a raw English `Bad Gateway`.** | **CLOSED 2026-08-06 — controller v0.202.0, partially and stated as such.** **The layer that answers is traefik**, whose config this repo generates — but traefik v3 serves no static files, so a branded proxy page needs a **new always-up container** for every 502 on the box: **scoped, not built**. Shipped: the unlock posts via `fetch` and answers a gateway failure in Hungarian in-page. **Progressive enhancement — with no JS the plain POST still shows the proxy's error** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-234** | **An off-site run reports success while silently omitting an app the customer just switched on.** Shipped in v0.205.0. | **CLOSED 2026-08-06** — controller v0.205.0 | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-236** | **After a guest rebuild the hub never re-stages the off-site credential.** | **CLOSED 2026-08-06 — not a defect** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-237** | **After a successful recovery the customer is shown no backups at all, because the restore surface is keyed on apps that are currently installed and currently marked for future remote backup.** | **CLOSED 2026-08-06** — controller v0.204.0: the list is now built from `OffsiteInventoryList` (the repository's own snapshot tags). Installed-ness became a property OF a row, never a filter; an unreadable store renders as UNKNOWN **and keeps the action offered**; `felhom-offbox` and `_shares` are excluded. 7 new tests incl. a rendered-page test for the rebuilt shape, and a red-proof that keys the list back on installed-and-toggled apps. | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-238** | **„Teljes visszaállítás előkészítése" accepts the click and does nothing.** Shipped in v0.204.0. | **CLOSED 2026-08-06** — controller v0.204.0 | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-239** | **The fixes are written, tested, pushed — and a machine installed tonight gets none of them.** Shipped in 0.127.0, 0.203.0, 0.204.0. Evidence: `tests/finalwalk-r201-2026-08-07/journal.md`, `tests/golden-0.205.0-2026-08-07/`. | **CLOSED 2026-08-07** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-241** | **The credential self-heal, succeeding, locks the customer out of their own recovery.** Shipped in v0.206.0, v0.98.0. Evidence: `audits/SPIKE-r241-recovery-offer-2026-08-07.md`, `tests/finalwalk-r201-2026-08-07/journal.md`. | **FIXED 2026-08-07 — v0.206.0 / hub v0.98.0** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-247** | **The box is being told something false, in its own words, and it recommends the destructive act.** Shipped in v0.206.0. | **CLOSED 2026-08-08** — controller v0.209.0. The field is received and the box tells the two conditions apart; see the Campaign-12 follow-through section below. The WRONG FLAG itself is R-246 (operator act, hub-side) and the customer-facing card copy is unchanged — both stated rather than folded in | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-249** | **The retrieval passphrase ships in the customer page's HTML, so any headless read puts it in a transcript.** Shipped in v0.206.0, v0.207.0. **Reasoning kept:** **Severity MEDIUM:** it is a live per-customer secret that fetches the whole config (`GET /api/v1/config/` with `X-Retrieval-Password`), but the exposure is to someone who can already read the operator page — a defence-in-depth failure, not a boundary crossed. | **CLOSED 2026-08-08** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-252** | **After a rebuild the restore refuses because the data drives are not registered, and nothing on the recovery path says so.** Shipped in v0.207.0. | **CLOSED 2026-08-08** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-253** | **The restore page promises it will reinstall the app, and the restore then refuses because the app is not installed — in the customer's own language, three lines apart.** Shipped in v0.207.0. | **CLOSED 2026-08-08** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-212** | **The orphaned-ciphertext deletion HALTED: the stores on the storage box do not match this register's record.** | **CLOSED 2026-08-05 — all three deleted after the operator confirmed the corrected list** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-94** | **~~A hand-synced version constant drifts, and the gate that would catch it is never run~~** Shipped in 1.22.0, 9.9.9. | **CLOSED — SHIPPED** (hub v0.87.0, 2026-08-02) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-110** | **`main` is the installer's publish channel — there is no staging.** Shipped in 1.22.0, v1.23.0, v4.4.0. **Reasoning kept:** E-2a's `felhom-backup-target-apply` (`:2116`) is installed **0755 to `/usr/local/sbin` and root-fenced in sudoers**, validated only by `bash -n` — a root-executed artifact taken from `main` with no pinned integrity, which is this row's class exactly. Fixing only (i) leaves a tagged installer pulling nine untagged files from `main` at run time — a staging story that is false in the place it matters most, since one of those nine (`felhom-backup-target-apply`) is installed **0755 into `/usr/local/sbin` and root-fenced in sudoers**, validated only b | **CLOSED — SHIPPED** (installer v1.23.0, 2026-08-03) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-111** | **The Day-0 artifact channel is 17 agent releases stale — a box installed today gets agent `0.96.0`, not `0.113.0`.** Shipped in 0.113.0, 0.114.0, 0.161.0. Evidence: `audits/E2D-fresh-vm-2026-07-29.md`. **Reasoning kept:** **The global controller floor was deliberately NOT raised**: the golden now bakes 0.185.1, so a fresh box needs no self-update, and raising it would have been an unnecessary fleet-wide write. | **SHIPPED 2026-07-29 — the channel now serves agent 0.113.0 + golden 0.185.1** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-115** | **Publishing is a remembered step, and it was forgotten within eight hours of being documented as forgettable.** Shipped in 0.113.0, 0.114.0, 0.119.0. **Reasoning kept:** **Class: → R-29, one layer up** — a control that exists and is never walked; deliberately NOT given its own ID. It calls the existing `publish-agent.sh` rather than reimplementing it, refuses a dirty or unpushed tree, refuses to re-release an existing version (one version name must never mean two binaries), and **deliberately does not vouch** — vouching points machines at a version and stays the operator's ac | **CLOSED — SHIPPED** (`release-agent.sh` + `check-published-versions.py`, 2026-08-03) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-116** | **The drive-absent alarm and its recovery were a MISMATCHED PAIR — absent fired the GENERIC `storage_disconnected`, return the SPECIFIC `backup_target_restored`; `backup_target_absent` never fired at all** Shipped in 0.185.1, v0.115.0, v1.25.0. Evidence: `audits/R116-v0116-2026-07-30.md`, `audits/SPIKE-r117-bind-liveness-2026-07-30.md`. | **SHIPPED + PROVEN-LIVE** (agent v0.116.0, 2026-07-30) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-120** | **The golden baked a controller that predated R-114 + R-112, so a FRESH box showed the customer the WRONG absent-target message** Shipped in 0.113.0, 0.116.0, 0.156.0. Evidence: `audits/R120-golden-rebake-2026-07-30.md`. | **CLOSED — golden rebaked + PROVEN-LIVE, and the class now has an ENFORCED gate** (golden 0.186.0 + hub v0.82.0, 2026-07-30) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-117** | **A drive's guest bind becomes a DEAD MOUNT while every signal reads healthy — and it happens in TWO ways, only one of which the original framing covered.** Shipped in 0.113.0, 0.117.0. Evidence: `audits/R117-v0117-2026-07-30.md`. **Reasoning kept:** **No block I/O proven by strace** (only `/proc/self/mountinfo`, **0** statfs) — the Part 1 `CLAUDE.md` fence applied to its own first consumer. | **SHIPPED + PROVEN-LIVE** (agent **v0.117.0**, 2026-07-30) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-113** | **The drive-absent gate CANNOT FIRE on device loss — E-2b's alarm is wired to an unreachable condition.** Shipped in 0.113.0, 0.114.0, v0.185.0. Evidence: `audits/E2D-fresh-vm-2026-07-29.md`, `audits/SESSION-C-2026-07-29.md`. | **SHIPPED + PROVEN-LIVE** (agent v0.114.0, 2026-07-29) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-112** | **E-2's degraded banner and offer have NO UI CONSUMER — the endpoint is correct and the customer never sees it.** Shipped in v0.185.1. Evidence: `audits/E2D-fresh-vm-2026-07-29.md`, `audits/SESSION-C-2026-07-29.md`. | **SHIPPED + PROVEN-LIVE** (controller v0.186.0, 2026-07-29) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-114** | **On target-drive loss the customer is told the wrong story and offered the drive that just vanished.** Shipped in 0.113.0. Evidence: `audits/E2D-fresh-vm-2026-07-29.md`, `audits/SESSION-C-2026-07-29.md`. | **SHIPPED + PROVEN-LIVE** (controller v0.186.0, 2026-07-29) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-29** | **The green gates are not enforced anywhere — one was RED for 16 releases before anyone ran it.** Shipped in 1.19.0, 1.22.0, v0.129.0. | **CLOSED — both halves shipped** (2026-08-02) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-86** | **Restore-tests are interval-scheduled, not backup-aligned** | **CLOSED — SHIPPED + PROVEN-LIVE 2026-08-03** (agent **v0.121.0**, hub **v0.91.0**) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-87** | **The restic tier is never restore-tested** **Reasoning kept:** **And R-95 still applies:** that credential can delete, so a restic restore-test must never be able to write to the repo CC | **READY — RE-RANKED UP 2026-08-03 (R-86 closed)** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-185** | **The agent cannot see the host backup tier's archives on demo-felhom — the PVE token has no ACL on `/storage/felhom-backup`, so the content listing returns EMPTY where root sees three archives.** Shipped in v0.123.0. | **CLOSED — SHIPPED + PROVEN-LIVE 2026-08-03** (agent **v0.123.0**, installer **1.24.0**) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-186** | **A released agent binary's sha256 cannot be reproduced from its tag.** Shipped in v0.120.1, v0.121.0, v0.121.2. | **CLOSED — SHIPPED + MEASURED 2026-08-03** (agent **v0.122.0**) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-187** | **R-115's one-command release had never actually run its publish leg — the first real use died there.** Shipped in v0.121.0. | **CLOSED — SHIPPED 2026-08-03** (`felhom-agent`) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-188** | **Every agent release has a ~50 % chance of emailing the operator a CI failure for a release that is correct.** Shipped in 0.121.2, v0.121.0, v0.121.1. | **CLOSED — SHIPPED 2026-08-03** (agent **v0.122.0**) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-189** | **A passing restore-test can be invisible to the hub forever — and R-86 made that window a week instead of a day.** Shipped in v0.121.1. | **CLOSED — SHIPPED + PROVEN-LIVE 2026-08-03** (agent **v0.122.0**) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-195** | **A customer with no machine ever bound e-mailed an `expected_dbdump_missed` ERROR every morning.** Shipped in v0.73.0. **Reasoning kept:** Fail-**open** on a read error (an unreadable binding must never SUPPRESS a real alarm), and the deferral is LOGGED with its own counter (the v0.73.0 Part-7 precedent: a quiet check must not look like a check that did not run). | **SHIPPED** (hub **v0.92.0**, 2026-08-04) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-196** | **`escrow_stale` is wired to the ONE path that does not change the repo password, and absent from the path that does.** Shipped in v0.95.0. Evidence: `audits/DRILL-r201-night-run-2026-08-04.md`, `audits/SPIKE-offsite-credential-recovery-2026-08-04.md`. | **CLOSED 2026-08-05 — hub v0.95.0** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-197** | **The hub holds both halves of the evidence that a box's offsite DATA key changed, and reads neither.** Shipped in v0.78.0, v0.93.0. Evidence: `audits/SPIKE-offsite-credential-recovery-2026-08-04.md`. **Reasoning kept:** The in-between shapes (a first-ever hash, a hash-less supersession) are LOGGED rather than dropped, so *"we chose not to alarm"* and *"the check did not run"* never look identical. | **SHIPPED** (hub **v0.93.0**, 2026-08-04) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-198** | **The hub's superseded-escrow retention does NOT retain the offsite repository password — and the ceremony the system tells the customer to run is what destroys the last copy.** Shipped in v0.92.0, v0.93.0. Evidence: `audits/RECON-offsite-dr-chain-2026-08-04.md`. | **SHIPPED** (hub **v0.93.0**, 2026-08-04) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-199** | **The hub serves recovery blobs on two endpoints that have no client anywhere in the system.** Evidence: `audits/RECON-offsite-dr-chain-2026-08-04.md`. | **SHIPPED + PROVEN-LIVE 2026-08-04** (hub **v0.94.0**, agent **v0.125.0**) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-203** | **A customer-declared MANDATORY data directory was silently absent from the off-site snapshot while the run reported `ok`.** Evidence: `audits/DRILL-r201-offsite-recovery-2026-08-04.md`. | **SHIPPED + PROVEN-LIVE 2026-08-04** (controller **v0.197.0**) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-204** | **A rebuilt box can recover its off-site key and still cannot use it: the remedy that reconfigures the tier is the thing that blocks the recovery.** Shipped in v0.198.0, v0.199.0, v0.95.0. Evidence: `audits/DRILL-r201-night-run-2026-08-04.md`. **Reasoning kept:** **Live on demo-felhom 9201, nothing restarted (`restarts=0`, container older than both mints): the superseded code returned „Hibás vagy lejárt kód" and the current one was accepted first time.** **Item 2 (a re-issue marks a healthy escrow stale) — CLOSED, → R-196.** Test-proven; deliberately NOT fir **What is deliberately NOT automated: the escrow ceremony.** A credential is replaceable; the recovery code is not. | **ALL FOUR ITEMS CLOSED 2026-08-05** (items 1–3 controller v0.198.0 + hub v0.95.0; item 4 controller v0.199.0 + hub v0.96.0) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-192** | **`offsite_delivery_stuck` tells the operator the opposite of what the detector measured, and the self-heal silently refuses for exactly the reason the message denies.** Shipped in 0.187.0, 0.192.0, v0.199.0. Evidence: `audits/RECON-offsite-dr-chain-2026-08-04.md`, `audits/SPIKE-offsite-credential-recovery-2026-08-04.md`. | **CLOSED 2026-08-05 — the guard's scoping half closed BY REPLACEMENT** (hub v0.96.0) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-193** | **A guest rebuild silently drops the off-site app-data tier, and nothing restages the credential.** Shipped in 0.156.0, 0.187.0, 0.192.0. Evidence: `audits/RECON-offsite-dr-chain-2026-08-04.md`, `audits/SPIKE-offsite-credential-recovery-2026-08-04.md`. | **CLOSED 2026-08-05 — controller v0.200.0** (credential half v0.199.0/v0.96.0; the recovery SCREEN v0.200.0) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-90** | **~~ep0 RAM headroom — 4 GiB swap survived its first reboot 2026-07-27; 3.8 GB RAM unchanged~~** | **CLOSED — the operator rescaled ep0 to a CX33 on 2026-08-03** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-97** | **~~Whole-guest backup tier had no hub signal; quiesce blamed the apps~~** Shipped in v0.79.0. | **SHIPPED** (controller v0.177.0 + hub v0.78.0/v0.79.0, 2026-07-27) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-100** | **~~A restic offsite tier that fails every night never goes stale on the hub — `isStale` counted from `LastRun`** **Reasoning kept:** The real defect is **defeated defence in depth**: the hub-side *pull* net was anchored on a field the failing controller keeps refreshing, so it could not compensate for a lost *push* (cf. | **SHIPPED + PROVEN-LIVE** (controller v0.181.0 + hub v0.80.0, 2026-07-28) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-101** | **~~Tier-2 `LastRun` is written on failure and rendered to the customer as „Legutóbbi másolat" — including in t** | **SHIPPED + PROVEN-LIVE** (controller v0.182.0, 2026-07-28) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-108** | **~~Network storage can host an app's namespace, and FileBrowser binds a network share at its ROOT~~** Evidence: `audits/R108-network-app-namespace-2026-07-30.md`. | **SHIPPED + PROVEN-LIVE** (controller v0.187.0, 2026-07-30) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-109** | **~~The DR recipe records no backup target~~** Evidence: `audits/R106-R109-recipe-completeness-2026-07-30.md`. | **SHIPPED + PROVEN-LIVE** (agent v0.118.1 + hub v0.83.0, 2026-07-30) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-106** | **~~The DR recipe records the PBS namespace as `"root"` on every box~~** | **SHIPPED + PROVEN-LIVE** (agent v0.118.1, 2026-07-30) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-122** | **~~`AssembleDRRecipe` silently DROPPED `offsite_restic` — the offsite recovery location never reached any reci** | **SHIPPED** (hub v0.83.0, 2026-07-30) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-125** | **A "test through the production path" is only true up to the seam it injects at.** Shipped in v0.118.0. Evidence: `audits/R106-R109-recipe-completeness-2026-07-30.md`. | **FIXED** (agent v0.118.1) — filed for the DOCTRINE point | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-128** | **~~`build-felhom-iso.sh:44` comments that `ISO_VERSION` "aligns with felhom-host-install SCRIPT_VERSION" — a c** | **CLOSED** (iso v1.26.0, 2026-07-31) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-154** | **~~`[first-boot]` is automated-install-only and nothing in the Felhom tree said so~~** Evidence: `audits/SPIKE-universal-iso-3-2026-07-31.md`. | **CLOSED** (iso v1.26.0, 2026-07-31) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-155** | **~~`iso-repack.sh` refuses any ISO without `auto-installer-mode.toml`, blocking the no-`answer.toml` posture~~** | **CLOSED** (iso v1.26.0, 2026-07-31) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-156** | **An app's data is neither persisted nor backed up, and it reports healthy.** Shipped in 26.6.1. **Reasoning kept:** **Provenance, stated because it decides the row:** the observation is `docker ps -a` on demo-hp's **guest 9201** returning empty, supplied with the 2026-08-02 task; **this session did not re-measure** (documentation-only, every box fenced). | **CLOSED — all three apps fixed** (papra template, 2026-08-03) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-157** | **`bootrecon`'s start-ONCE sweep misses the boot orphan it exists to recover — TWO mechanisms.** | **CLOSED — SHIPPED + PROVEN-LIVE** (B: controller v0.189.0; A: v0.190.0, 2026-08-02) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-170** | **The drive-backed boot gate infers a customer's Stop from a container count.** Shipped in v0.190.0. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.190.0, 2026-08-02) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-171** | **The boot sweep started apps whose data drive was ABSENT — a regression introduced by v0.189.0, now FIXED.** Shipped in v0.189.0. Evidence: `audits/DIAG-bootrecon-drive-absent-2026-08-02.md`. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.190.0, 2026-08-02) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-172** | **A false `host_stale` alarm fires when the hub's SQLite refuses two consecutive host reports.** **Reasoning kept:** **Retry options (b) and (c) were deliberately NOT taken** — with readers no longer blocking writers a surviving `SQLITE_BUSY` would be a real signal, and a retry would hide it; revisit only on evidence. | **CLOSED — SHIPPED + PROVEN-LIVE** (hub v0.88.0, 2026-08-02) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-174** | **The app-stop guard's crash recovery started apps onto MISSING drives — a regression in v0.189.0 code.** Shipped in v0.189.0. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.191.0, 2026-08-02) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-175** | **`07-backup-architecture.md` §7.5 states ONE box's size bound as if it were the fleet's.** Shipped in 0.192.0, 7.5.1. Evidence: `audits/SPIKE-r165-mp1-merge-2026-08-02.md`. | **CLOSED — FIXED 2026-08-03** (same pass as R-165) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-183** | **A fresh install fetched the vouched agent BINARY and its sixteen CONFIG files from two different refs, and nothing compared them.** Shipped in v0.120.0. **Reasoning kept:** **Why it is a defect and not only untidiness:** these files are the agent's own operating surface — its systemd unit, its sudoers, its guarded wrappers — and `configs/felhom-backup-target-apply` is installed **0755 into `/usr/local/sbin` and root-fenced in sudoers**, validated only by `bash -n`. | **CLOSED — SHIPPED** (installer v1.23.0, 2026-08-03) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-182** | **A full disk tells the operator about ONE app and silently swallows every other app's refusal for an hour.** Shipped in v0.194.0, v0.90.0, v0.90.1. | **CLOSED — SHIPPED** (controller v0.194.0 + hub v0.90.0/.1, 2026-08-03) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-181** | **The capture floor guards the cheap leg and not the leg that fills the volume — and its refusal message asserts an invariant the code does not provide.** Shipped in v0.192.0, v0.193.0, v0.193.1. | **CLOSED — SHIPPED** (controller v0.193.0 + v0.193.1, 2026-08-03) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-178** | **The merged golden (0.192.0) is built and published but NO BOX HAS BEEN REINSTALLED FROM IT, and it is deliberately UNVOUCHED.** Shipped in 0.119.0, 0.120.0, 0.192.0. | **CLOSED — BOTH BOXES REINSTALLED AND PROVEN (2026-08-03)** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-158** | **A local Tier-1 app-data backup failure reaches no hub channel.** Shipped in v0.78.0. | **CLOSED BY R-167 — SHIPPED + PROVEN-LIVE** (controller v0.191.0 + hub v0.89.0, 2026-08-02) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-159** | **wishlist's data landed in an ANONYMOUS volume — never backed up, orphaned by a redeploy.** | **SHIPPED** (`templates/wishlist/docker-compose.yml`, 2026-08-02) — filed to record the CLASS | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-160** | **gramps-web persisted three paths and wrote to none of them.** | **SHIPPED** (`templates/gramps-web/docker-compose.yml`, 2026-08-02) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-163** | **`mp1` is RETENTION, not staging — and it is sized as if it were neither.** Shipped in v0.192.0. | **CLOSED by R-165 — the ceiling it describes no longer exists** (golden v3.0.0, 2026-08-03) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-165** | **Merge `mp1` into `mp0` — the dedicated 20 G backup partition stops existing.** Shipped in 0.192.0, v0.192.0. Evidence: `audits/SPIKE-r165-phase0-2026-08-03.md`. | **SHIPPED — golden `build-golden.sh` v3.0.0 + agent v0.120.0 + controller v0.192.0 (B2), 2026-08-03. IMPLEMENTED — the LAYOUT is proven live on both boxes (R-178, 2026-08-03); the BULKHEAD'S REPLACEMENT IS NOT (→ R-181)** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-166** | **App state gets a desired/observed model with its own store.** Shipped in v0.189.0. | **SHIPPED + PROVEN-LIVE** (controller v0.189.0, 2026-08-02) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-167** | **Storage monitoring and backup alerts.** Shipped in v0.191.1, v0.191.2. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.191.0/.1/.2 + hub v0.89.0, 2026-08-02) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-168** | **~~CI: no runner exists, and with trunk-based pushes CI can DETECT but not BLOCK~~** Shipped in 0.1.0. Evidence: `audits/SPIKE-ci-runner-2026-08-02.md`. | **SHIPPED — and the alarm is DEMONSTRATED** (2026-08-02) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-205** | **`RootFsPressureDespiteHousekeeping` can never fire** Evidence: `audits/SPIKE-dooplex-buildcache-2026-08-05.md`. | **CLOSED — SHIPPED + RED-PROVEN LIVE** (`homelab-manifests` `6808a4b`, 2026-08-05) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-258** | **C3 — the customer's per-app backup tick is green on the PRESENCE of a restore point, and its only red condition is a GLOBAL one.** | **CLOSED 2026-08-08 — controller v0.210.0.** `appDumpVerdict` reads THIS app's own dump result; three states, no icon when nothing is known. **Recency deliberately not added** — see the observation in the follow-through section | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-259** | **C4 — a disk read that FAILS renders as „0.0 GB / 0.0 GB (0%)" in the nominal colour, on the dashboard's most-looked-at meter.** | **CLOSED 2026-08-08 — controller v0.210.0.** `readDiskUsage` reports success; `SystemInfo.DiskKnown`/`HDDKnown`; the template draws no figure, no percentage and no meter fill when unknown. **The hub leg is deliberately NOT fixed and is now R-266** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-260** | **C5 — the agent reports at least eight decision-bearing facts the hub models NOWHERE, and the sharpest one blinds the check that answers „can the operator get into this box".** | **CLOSED 2026-08-08** — the class is GATED (G-1, `scripts/wire_contract_gate.py`) and the sharpest instance is fixed (hub v0.99.0). The remaining unconsumed facts are **R-264, OPEN** — allowlisted with reasons, which is not the same as decided. See the follow-through section below | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-265** | **A CI run can fail with NO LOG PERSISTED, and the alarm mail then points the operator at a log that does not exist.** | **CLOSED 2026-08-08 — `timeout-minutes: 5` on the gates job, and the alarm mail now states elapsed seconds and qualifies its own "names itself in the run log" sentence.** ⚠ **The unknown is NOT closed and must not be read as closed:** whether the `if: failure()` alarm fires for a REAPED job is still unverified. The timeout makes the reap unreachable in practice; it does not answer what happens inside one | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-267** | **The Configuration page is 2.6× faster and is still ~10 s, and the remaining cost is ONE Gitea call whose latency swings 20× with load.** Shipped in 0.100.2, v0.100.0. | **CLOSED 2026-08-08 — hub v0.101.0 + a registry prune.** **Final: cold 5.4 s, warm 0.14 s** (was 26.2 s). Three serialisation legs took it to 9.85 s mean, the 60 s in-memory memo took the warm path to a quarter-second, and the prune halved what remains of the cold path. **⚠ TWO CORRECTIONS TO THIS ROW'S OWN EARLIER TEXT, because both were wrong and both mattered.** **(1) "Only 50 generic versions exist" WAS NOT A COUNT, IT WAS A PAGE LIMIT.** `?type=generic&limit=1000` returns at most 50; the 50 I measured was exactly the cap, and three older agent versions (0.81.0, 0.80.0, 0.79.0) only became visible after the first 30 deletions moved them onto page one. **An unpaginated listing is not evidence of a total** — this repo's own "an empty listing is not evidence of emptiness" rule, walked into while measuring. **(2) THE OPERATOR'S "REDUCE THE NUMBER OF ARTIFACTS" WAS THE BETTER CALL AND MY MEASUREMENT SAID OTHERWISE.** I reported it helps "sub-linearly" and "is not the lever". Measured after: trimming to 10+10 took the COLD load from 13.4 s to 5.4 s — a 2.5× improvement on the path the memo cannot help, because the fan-out is per-version. Recorded rather than quietly dropped (the R-96 standing rule). **Pruned to the newest 10 per package on the operator's rule**, with the live-vouched golden/agent/floor asserted into the KEEP set before a single DELETE was issued; 33 deletions, all HTTP 204, and golden 0.210.0 / agent 0.128.0 / agent 0.127.0 verified still fetchable afterwards. `drill-r50` runs agent 0.113.0, now deleted — flagged to the operator first; it is a disposable nested drill VM and only its re-download path is gone | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-268** | **A live per-guest local-API token was printed into a session transcript.** Shipped in 169.254.253. | **CLOSED — ROTATED + PROVEN LIVE 2026-08-09** (rehearsal pre-phase, `audits/REHEARSAL-byo-reinstall-2026-08-09.md` §3) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-273** | **RANK 1 — the hub vouched an agent version that was never git-tagged, and every install fleet-wide now fails at step 5/8.** Shipped in 0.127.0, 0.128.0, v0.127.0. | **CLOSED 2026-08-09 — tag pushed, install PROVEN** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-278** | **demo-felhom's off-site tier has never completed a run and has been stuck for six days.** Shipped in 0.200.0, v0.93.0. | **CLOSED 2026-08-10 — protection RESTORED, and the recovery it waited for could never have worked** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-280** | **RANK 1 — after a reinstall the data drive cannot be re-attached through ANY dashboard route, and the restore page promises it is "two clicks".** | **CLOSED — controller v0.211.0, delivered via golden 0.211.0 (vouched 2026-08-10)** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-281** | **The hub said NOTHING through an entire reinstall — and the tripwire for a sealed-backup unseal did not fire on a real unseal.** | **WITHDRAWN 2026-08-09** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-293** | **CENSUS, 2026-08-10 — no machine that is not ours can be in the state that cost demo-felhom its history, and here is the whole population.** | **CLOSED-INFORMATIONAL 2026-08-10** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-294** | **The orphan card promises restorability that the box rendering it cannot evaluate — specified, not implemented.** Evidence: `documentation/design/SPEC-orphan-card-copy-2026-08-10.md`. | **CLOSED — controller v0.211.0; see R-299 for the sentence it missed** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-296** | **The orphan card's OTHER sentence makes the same promise, and the spec says it is fine.** | **CLOSED — shipped in controller v0.212.0 (R-299); verified: the sentence at backups_remote.html:98 was replaced and the stem guard covers it** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-297** | **An install took whatever golden was lying around.** Shipped in 0.153.0, 0.210.0, 0.213.0. | **CLOSED — observed live + PUBLISHED as `installer-v1.27.0` (both refs bumped)** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-299** | **The orphan card's OTHER sentence made the same unevaluable promise, and the spec called it accurate.** Shipped in v0.211.0. | **CLOSED — controller v0.212.0** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-300** | **Our own uninstall left the thing that makes our own reinstall refuse.** Shipped in 0.0.0, 10.0.2, 127.0.0. | **CLOSED — observed live + PUBLISHED as `installer-v1.27.0` (both refs bumped)** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-301** | **The abandon countdown banner makes the retired promise a third time, and as a flat statement.** **Reasoning kept:** the customer chose to abandon a recovery offer that exists — which is why it was NOT changed (this session was fenced to the orphan card). | **CLOSED — premise CONFIRMED and fixed in controller v0.213.0 (R-302)** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-302** | **The abandon banner promised retrieval it could not see was still true — fixed by PINNING a fingerprint at the decision.** Shipped in v0.213.0. **Reasoning kept:** Empty is not a match on either side; a countdown started before v0.213.0 carries no pin and takes the cautious branch (deliberately NOT backfilled). | **CLOSED — controller v0.213.0** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-305** | **The R-300 cleanup fires exactly once per machine, and the second reinstall hits the original wall.** Shipped in 0.0.0, v1.27.0. | **CLOSED — superseded by R-316** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-307** | **`demo-felhom` carries a LIVE abandon countdown that this drill did not start — and the end state says there should be none.** **Reasoning kept:** The drill's fence forbade starting, shortening or triggering a countdown, and none was; but its required end state was *"no abandon countdown anywhere"*, and one exists. | **CLOSED — countdown cancelled 2026-08-12 on the operator's ruling** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-308** | **~~The stored controller password no longer opens `demo-felhom`~~ — WITHDRAWN 2026-08-12, this was MY BUG, not a defect.** | **WITHDRAWN — not a defect (my error)** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-309** | **The day-0 runbook says pushing the installer publishes it. It has not since R-110.** Shipped in 1.25.0, 1.27.0. Evidence: `documentation/runbooks/day0-install.md`. | grep -m1 '^SCRIPT_VERSION'`. **Measured while writing it: served `1.28.0`, `main` `1.28.0`, both pins `installer-v1.28.0` — the three agreeing is the observation; any one alone is not.** **The claim was copied elsewhere and the copy was hunted:** `audits/SPIKE-universal-iso-3-2026-07-31.md:184` said the same thing and **cited `day0-install.md` as its source**, which is how it spread. It was **true on the day it was written** (R-110 shipped 2026-08-03), so the dated finding is kept verbatim and carries a SUPERSEDED note rather than being rewritten — falsifying a dated record to tidy it is its own defect. Two other hits are correct in context: `hostinstall_gates.py:198` states the consequence of the manifest LOSING its tag, and the 2026-08-12 drill record already names the sentence as false | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-311** | **A correct recovery code for a retained package stopped being reported as wrong.** Shipped in 0.126.0, 0.128.0, 0.129.0. | **CLOSED — shipped + delivered: hub v0.103.0 + agent v0.129.0 + controller v0.214.0** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-316** | **The removal now genuinely reverses the installation — R-305's once-per-machine defect closed.** Shipped in 0.0.0, v1.27.0, v1.28.0. | **CLOSED — shipped + published, observed on the cycle that actually fails** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-318** | **No honest marker exists that says Felhom installed dnsmasq on a machine already in the field, and none can be invented.** Shipped in v1.27.0. **Reasoning kept:** `/var/log/dpkg.log` does record the install — and is a **timestamp**, which the standing rule refuses as a heuristic dressed as a fact. | **CLOSED — established, no action possible for existing boxes** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-319** | **The guest-network watchdog finally has a reader — the first of R-264's twenty-one, and it is the repair COUNT that matters, not the state.** Shipped in 0.92.0, v0.92.0. | **CLOSED — shipped hub-side 2026-08-13** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-320** | **Evidence has been destroyed twice in three days, in the same place, by the same act.** Shipped in v1.28.0. Evidence: `audits/DRILL-retained-key-2026-08-12.md`, `audits/REPORT-r316-installer-v1.28.0-2026-08-13.md`. **Reasoning kept:** **The rule, now standing rule 5 in `workspace-CLAUDE.md` (so it loads in every session) and repeated where a session actually meets it — `runbooks/target-selection.md`, `RUNBOOK-rehearsal-v3.md`, and the `PROMPT-TEMPLATE.md` report section: evidence is copied off the machine at the end of the phase | **CLOSED — rule written, four homes** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-321** | **A box on which reporting is deliberately switched off still alarms as stale, then down.** Shipped in v0.105.0. | **CLOSED — shipped hub v0.105.0, both doors** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-322** | **The claim guard has never scanned the hub, and the hub sends the customer's first sentence.** **Reasoning kept:** **Recommended shape, and the reason it is not one line:** the gate is invoked by `controller_gates.py`, so pointing it at a sibling repo makes a controller gate fail on a felhom.eu edit — the cross-repo lesson from G-1 (a gate needing a sibling passes locally and exits INCONCLUSIVE in CI, and must n | **CLOSED 2026-08-13 by R-324** — `scripts/hub_copy_gate.py`, registered in `repo_gates.py`, scanning 95 hub files for retired names and four declared customer surfaces for retrieval stems, with a plant→convict→remove→pass selftest that caught a defect in its own instrument on the first run. The stem list IS shared (`scripts/customer_copy_vocab.py`) and no controller gate was made to depend on a felhom.eu clone; the controller gate's adoption of the shared list is R-325, and until it happens the two are drift-checked rather than left to diverge | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-323** | **The third near-homograph — the five-word phrase is „Tulajdonosi jelmondat” now.** | **CLOSED — shipped hub v0.105.0** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-324** | **The hub's customer copy is under a guard for the first time — and the guard has been watched catching, ignoring and releasing.** | **CLOSED — shipped, selftest green** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-326** | **"Which claims are unproven?" is a question a machine can answer now — and the number everyone was repeating answered a different question.** | **CLOSED — shipped** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-328** | **The disk alert was emailed to nobody, and one word is the whole reason.** Evidence: `audits/DIAG-smart-passed-trap-2026-08-14.md`. | **CLOSED — controller v0.215.0, PROVEN LIVE 2026-08-14.** Now `"warning"`, and `DiskAlertKind.Severity()` is exported so the contract is assertable from any package rather than duplicated as a literal. **The proof is a side-by-side pair pushed through the REAL hub event endpoint** from demo-hp's controller: severity `"warning"` → stored `warning`, `notification_log` **id 689, channel `operator`, status `sent`**; the identical push at `"warn"` → stored **`info`**, and **no `notification_log` row exists at all**. Pinned by `TestNotifyDiskHealthDegraded_SeverityRoutes`, which asserts membership of the hub's accepted set (not just the literal) and names both hub locations; its red-proof — restoring `"warn"` — fails all three assertions | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-334** | **CLOSED 2026-08-18 — golden 0.216.0 baked, published and VOUCHED; CI green by run id.** Shipped in 0.214.0, 0.215.0, 0.216.0. Evidence: `documentation/tests/golden-*`, `documentation/tests/golden-0.214.0-2026-08-12`. | **CLOSED 2026-08-18.** Baked from `RUNBOOK-manual-build.md` §4.0+§4.1 in the DooPlex drill VM and published: **`GOLDEN_VERSION=0.216.0`**, **`GOLDEN_SHA256=ac004dc90d8cefccc5448377892f9cff3a4c3e1e27d0e11129120e38ac31c34b`**, 656,970,239 bytes at `…/generic/felhom-golden/0.216.0/golden.tar.zst`. **The published bytes were verified, not just the script's print** — the artifact was downloaded back out of Gitea and hashed, and it matches. **Vouched by the operator, all THREE fields together**, confirmed by reading the hub's own store rather than the save: `artifact_golden_version=0.216.0`, `artifact_agent_version=0.129.0`, `artifact_min_agent=0.129.0` (2026-08-18 11:00:59–11:01:00), and the hub's recorded sha256 matches the downloaded artifact. The R-216 shape was checked on the machine: `MinAgent` 0.129.0 is **equal to**, not above, the newest **published** agent. **`golden_currency_gate.py` rc=0 and `repo_gates.py --fast` rc=0 — all nine gates — and CI is GREEN BY RUN ID: run **353**, `head_sha 7d81681d6`, conclusion `success`** (the two prior runs 351/352 on this same afternoon were red on exactly this row, which is the contrast). That push needed **no `--no-verify`** — the first of the day that did not. Evidence: `documentation/tests/golden-0.216.0-2026-08-18/`, report `REPORT-golden-0.216.0.md`. **Closed with the run id quoted deliberately**: this row was re-confirmed once and widened once, and closing it on a local green a third time would have left the same ambiguity | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-335** | **One physical disk was walked TWICE per run, and the second walk sustained it against itself.** Shipped in v0.215.0. **Reasoning kept:** **This is the shape standing rule 3 warns about: an absent alarm was not evidence — the two artefacts had to be read AGAINST each other** — — CC | **CLOSED — controller v0.216.0, 2026-08-14.** Each `diskKey` is evaluated once per run; both entries stay marked `seen` so neither looks like a disappeared disk, and the card still renders both storage rows (the dedup is about state and alerts, not display). Pinned by `TestDiskCheck_SameDiskTwiceIsEvaluatedOnce`; companion red-proof run and reverted — deleting the guard makes the first sighting emit `Kind:2` (Hiba-from-sectors) at 8 sectors | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-344** | **`felhom-agent` leaks one TCP connection to PBS per poll cycle, forever, on both sides — and it is the whole of the ep0 descriptor leak.** Shipped in 0.129.0, 0.130.0. Evidence: `audits/SPIKE-ep0-established-connections-2026-08-20.md`. **Reasoning kept:** **The proof obligation is the fd count, not the diff:** per standing rule 3 the positive observable is ep0's ESTAB count going FLAT between proxy restarts, measured over a window long enough to matter — a green test suite proves nothing here, and a 30-minute window proves nothing here either (that e **control 4 cycles -> 4 leaks; fixed 4 cycles -> 0 leaks.** **Positive observable per standing rule 3** (a zero leak is equally consistent with "the agent stopped working"): the fixed box's four poll cycles are in ep0's log, and the boxes' other traffic is near-identical (libwww-perl 924 vs 926, pro | **CLOSED 2026-08-20 — fixed, proven live on both boxes, published and vouched** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-347** | **The R-344 fix exists on two demo boxes by hand and NOWHERE ELSE — a box installed from the current image still ships the leaking agent.** Shipped in 0.129.0, 0.130.0, 0.216.0. Evidence: `documentation/runbooks/publish-train-rules.md`. | **CLOSED 2026-08-20 — published, vouched, and the fleet reconciled onto the published bytes** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-351** | **The restore never read back where the backup said the data lived, and a second press started a second restore.** | \.NamespaceRoot\b' --include=*.go` found **no non-test reader anywhere** — the reconstitution opened the manifest (`offbox_reconstitute.go:235`) purely for the coherence stamp and resolved its destination from the LIVE app instead. **A restore into a destination different from the recorded one therefore succeeded silently, under a green message.** **(b) The second press.** All seven restore handlers gated on `backupMgr.IsRunning()` — the CONCURRENCY flag, acquired *inside* the goroutine (`offbox_reconstitute.go:180`) **after** the handler returned. Established with a test before any change: both the reconstitute and place handlers answered „…elindult" and **overwrote the first restore's op/stack**. The wizard had read the correct flag since v0.154.0 and said so in a comment; the handlers were never moved over. **(c)** The banner gated its terminal result on a page-local `sawRunning`, so a restore that finished before the page opened — the 8.666 s OpenGist restore — was shown to nobody. | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-354** | **The off-site full restore has NO named-volume leg — the tar is in the unit, in the snapshot and in the checking folder, and is never replayed.** Shipped in 0.217.0, 0.218.0. | **CLOSED — SHIPPED + PROVEN-LIVE 2026-08-22** (controller **v0.218.0**) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-355** | **`paperless-ngx`'s PostgreSQL is dumped into a directory for a stack that does not exist, so its unit has never contained a database dump — and the destructive restore therefore takes no safety dump and tells the customer the app has no database.** Shipped in 0.217.0, 0.218.0. | **CLOSED — SHIPPED + PROVEN-LIVE 2026-08-22** (controller **v0.218.0**) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-339** | **The hub was SILENT when it lost sight of the off-site stores — and a 9 h 37 m outage proved it.** **Reasoning kept:** That is **correct for a fill signal** — a missing reading must never be mistaken for 0%, which is why degraded data drives no band transition — but the consequence was that a completely dead off-site endpoint and a healthy one were **indistinguishable on the operator channel**. | **SHIPPED — hub v0.106.0, 2026-08-18.** Reachability is now a second, independent signal: consecutive failed fetch windows counted per checker, `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default **3 windows (≈30–45 min)**, with paired `*_recovered` all-clears wired into `recoveredPairedDownTypes` — necessary because both recoveries are severity `info` and `severityNotifies` drops `info`. Threshold tunable via `alerting.box_unreachable_windows`. **The fill logic is untouched**: no threshold, throttle, band or escalate-once behaviour changed. Evidence: `internal/monitor/box_reachability_test.go` (Scenarios A–F) + `internal/notify/dispatcher_box_reachability_test.go` (the cross-package wiring, asserting an actual operator mail), plus three companion red-proofs each seen failing with a message naming the right cause | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-370** | **PROCESS: between 2026-08-19 and 2026-08-22 the reviewing side called a documented architectural decision a defect, in four places, because it read the register and live source and never `documentation/architecture/`.** Evidence: `documentation/architecture/`. **Reasoning kept:** R-352 (re-framed), R-369 The record is corrected in place with the framing marked rather than deleted, per the standing rule that a document which quietly changes its mind teaches nobody. | **CLOSED — corrected 2026-08-22** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-96** | **Two standing rules were agreed in chat and never committed** Evidence: `documentation/runbooks/workspace-CLAUDE.md:48-70`. **Reasoning kept:** **Two standing rules were agreed in chat and never committed** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-27, size XS, roadmap state `idea — found 2026-07-27`.** Moved verbatim; nothing added or reinterpreted. | **CLOSED — migrated from ROADMAP 2026-08-22** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-107** | **No offsite action unpacks the named-volume tars Tier-3 captures on every run.** Shipped in v0.218.0. | **CLOSED — migrated from ROADMAP 2026-08-22** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-383** | **The double-failure message told the customer their previous state was saved, and named a file that was not there.** Shipped in controller v0.222.0. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/`. **Reasoning kept:** *One of the two ways a rollback fails is that the undo copy is missing — so the sentence was most likely to be false in exactly the case it was printed.* **Do NOT simply drop the filename:** an operator needs it, and R-351's lesson is that a refusal naming nothing forces someone to remember what the product already knows — so the absent case still names WHERE the file should have been. **A zero-length dump counts as MISSING**, because a 0-byte file restores nothing and calling it present is the same false reassurance one step smaller. The check is `os.Stat` and deliberately not an integrity test: this runs at the end of a failed restore on a machine that may be unwell, and presence is the honest claim available there. | **CLOSED — SHIPPED** (controller v0.222.0, 2026-08-23; `undoCopyPhrase`, four cases, plus an AST seam test that the message is still wired to the builder) | full text: `git show 1eb64bec5183:documentation/backlog/OPEN-ITEMS.md` | | **R-384** | **An app whose DATABASE had died raised no dead-app alarm — the wrong question answered first.** Shipped in controller v0.222.0. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/`. **Reasoning kept:** *The defect was the ORDER of two questions, not the `unhealthy` exclusion.* "Is a SUPERVISED member dead?" and "is a RUNNING member failing its healthcheck?" are different questions, and the second was answering the first — a dying database drags its own front end `unhealthy`, so the symptom the fault causes was what suppressed the alarm for it. **`IsDownState` is byte-identical and `unhealthy` stays excluded** — an unhealthy container is RUNNING, and folding it in reintroduces the flapping that exclusion exists to stop; **no new state was minted**, `StateDegraded` already means this. **Two things had to move and either alone leaves the defect standing:** the hoist, AND widening "some members are up" from `running > 0` to *any member not in the down bucket* — the old guard made the R-51 block unreachable in precisely the case it was written for. **The register's own suggested fix was WRONG and is recorded as such:** it proposed a sustained-`unhealthy` threshold on the `crashLoopAfter` model; the actual defect needed no threshold at all. **PROVEN LIVE the only way it can be** — the same fixture that printed `0 currently down` on 2026-08-22 printed **`1 currently down`** on 2026-08-23, with `app_start_failed` 7 s after the stop and the banner reading *„…nem fut: BookStack (degraded)"*. Scenario D measured **0 alarms across 9 scans** through a full stop→start cycle. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.222.0, 2026-08-23) | full text: `git show 1eb64bec5183:documentation/backlog/OPEN-ITEMS.md` | | **R-329** | **`app_start_failed` was emitted with severity `"warn"`, so every one of them was delivered to nobody.** Shipped in controller v0.223.0 (+ hub v0.107.0). Evidence: `audits/DRILL-r329-r386-2026-08-23/`. **Reasoning kept:** *The vocabulary is EXACT and it is the HUB's, not ours* — `{info, warning, error, critical}`; anything else is coerced to `info` at ingest and dropped by `severityNotifies` before BOTH legs. **This was the SECOND occurrence** (`DiskAlertKind.Severity` until v0.215.0), and its comment had recorded the lesson — **a comment is not a guard**, so the guard is now an AST walk over the whole controller, with the six variable-passing call sites registered by name because a walk cannot follow a variable and *an unlisted limit is not a limit, it is a hole*. **The register's own framing was that the DECISION was the work** — should a stopped app mail the customer at all? Answered: **operator always, customer OFF by default**, because `processOperator` never consults customer preferences, so one word fixed the operator leg and left the customer leg exactly where the ruling wanted it. **Deliberately NOT added to `operatorOnlyEvents`** — that would make the new toggle visible, flickable and structurally incapable of delivering. **Measured on the live hub DB: 91 events stored all-time, ZERO notification rows before the fix; one operator row, `warning`/`sent`, after it.** | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.223.0 + hub v0.107.0, 2026-08-23) | full text: `git show 68a9f5475cd2:documentation/backlog/OPEN-ITEMS.md` | | **R-386** | **A single-container app stopped out of band raised no alarm, and a comment stated the opposite as settled fact.** Shipped in controller v0.223.0. Evidence: `audits/DRILL-r329-r386-2026-08-23/`. **Reasoning kept:** *the state test was guessing at something the product already knows.* `DesiredState` records the customer's intent, has **exactly one writer**, and is tri-state; `StateExited` never survives aggregation, so no state test can separate an out-of-band stop from a customer stop. The ruling: `Stopped` → no alarm, `Running` → **alarm**, **absent → UNKNOWN, keep today's behaviour AND announce it**. *Reading unknown as "nobody asked" would, on the first cycle after upgrade, e-mail about every app any owner ever deliberately stopped — fleet-wide, from a field that predates the intent it is being asked about.* **A rule without a mechanism is a wish:** every such suppression sets `IntentUnknown` and the names are logged at INFO, so an operator can answer *"how many apps am I blind to?"*. **`failedRestart` must still lift a `Stopped` intent or F-CRIT-1 re-opens.** **Fenced act: adding a `DesiredState` WRITER** — twelve of `StopStack`'s fourteen callers are machines. Proven live: alarm 24 s after an out-of-band `docker compose stop`, heartbeat `1 currently down` against the previous day's `0`; and with intent removed, suppressed *plus* the log line naming the app. **0 of 8 deployed apps on `demo-hp` carry an absent intent.** | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.223.0, 2026-08-23) | full text: `git show 68a9f5475cd2:documentation/backlog/OPEN-ITEMS.md` | | **R-389** | **Only the FIRST broken app per hour reached the operator — the cooldown key named the event type, not the app.** Shipped in hub v0.108.0. Evidence: `audits/DRILL-cooldown-grain-2026-08-23/`. **Reasoning kept:** the fix is a THIRD SIBLING of `cooldownTierSuffix`/`cooldownRunSuffix`, separate for the reason the second one's docstring already gives — *the existing two keep byte-identical semantics for every type that uses them.* **`cooldownStackSuffix` takes the EVENT TYPE as well as the details, unlike its siblings, and that asymmetry is the whole safety property:** `tier` and `run_id` appear only on types that want that grain, `stack_name` does not. **`perAppCooldownEvents` is a named allow-list with `app_start_failed` and nothing else** — *the backup family's cooldown is coarse ON PURPOSE (R-97a, R-182) so one full disk sends one digest rather than one mail per app*, and this is not hypothetical: **`crossdrive_failed` is severity `error`, reaches the operator leg, and carries `stack_name` through a different struct**, so a payload-shape rule would have split it silently. **The fenced act is adding an entry for a type whose family has a digest or a coarse-by-design cooldown.** `app_start_failed` qualifies precisely because it has NO digest — there is no `apps_down_run` the way `backup_run_failures` summarises a run. **The hour is unchanged; the grain was the complaint.** Fail-soft: absent or malformed details degrade to the old key and the mail still goes. **PROVEN LIVE 2026-08-23:** two apps four minutes apart gave **2 sent / 0 suppressed** where the same shape gave 1 and 1 the day before, each repeat suppressed under its OWN key (`…:opengist`, `…:calibre-web`) against the previous day's shared `key=demo-hp:app_start_failed`; and `crossdrive_failed` for two different apps stayed **coarse** under `key=demo-hp:crossdrive_failed`, byte-identical to the derived v0.107.0 value. **AND IT WAS NEVER FILED UNTIL THE DAY IT WAS FIXED** — it lived in a REPORT.md observations paragraph, which is why gate 11 now exists. | **CLOSED — SHIPPED + PROVEN-LIVE** (hub v0.108.0, 2026-08-23) | full text: `git show 45659bdc5a2f:documentation/backlog/OPEN-ITEMS.md` |