Files
felhom.eu/documentation/backlog/CLOSED-ITEMS.md
T
admin f41a1a0ad8
gates / gates (push) Failing after 18s
R-411/408/407, R-414, R-412a leg 1, R-410, R-406 CLOSED; determination + live evidence
Part 2.1's determination is the first artifact: the scratch resolver was consciously OUT OF
SCOPE for R-356, not excluded on state-only grounds - established from R-356's own commit
08eb1a6, whose tests say 'the prepared scratch still resolves ... only the DESTINATION
moves'. So 07 section 6.3's rule applies and now has a FOURTH consumer, and the section
says so.

Live evidence: the collision rerun on demo-hp with the sampler positively controlled first
(12 locks=1 across a real check, 4 locks=0 quiet), showing unlock --remove-all 0 times where
the drill saw it twice; and the proof reaching verdict pass on demo-felhom - the box that
could not run it at all - recorded where last_proof_result had been ABSENT every night.

Capability map: the off-site proof row now records that the nightly firing IS proven (it ran
unattended at 05:30 on demo-hp) and that a driveless box can now be proved.

Register: six rows closed and compressed. OPEN 176 -> 170, CLOSED 152 -> 158.
2026-09-01 10:36:54 +02:00

93 KiB
Raw Blame History

CLOSED-ITEMS — finished work, compressed

What this is. Every register row that reached a terminal state, compressed to its title, the version it shipped in, its evidence paths, and any sentence that states a RULE rather than a narrative. Nothing was deleted: each entry names the commit that holds its full original text, and git show <commit>:documentation/backlog/OPEN-ITEMS.md returns it verbatim.

Why it exists (operator ruling, 2026-08-22). OPEN-ITEMS.md had grown to 672 KB across 286 entries, over half of it finished work, with one single entry at 16 KB. A file that cannot be read is a file that cannot be checked — and this project has already paid for that twice: a record nobody could find because it sat inside an entry about something else, and a finding rediscovered because nobody could see it. The register now holds open work only, so its size tracks the work rather than the project's age.

A sibling rather than the bottom of the register, deliberately: appending to the same file keeps the byte count and the scroll, which is the thing being fixed.

Load-bearing reasoning was NOT compressed away. Where a closed row states a rule, a fence or a deliberate refusal, that sentence is carried here verbatim under Reasoning kept. Rules that outlive their work item also live in their proper homes — workspace-CLAUDE.md standing rules, felhom.eu/CLAUDE.md, CONTEXT.md, and the architecture folder — and this file is not their primary record.

This file is not the register. Nothing here is open. OPEN-ITEMS.md remains the single source of truth for open work; scripts/one_register_gate.py enforces that against ROADMAP.md.


| R-361 | The pre-restore safety dump overwrote the app's own DB dump, and the comment beside it said it could not. Shipped in controller v0.221.0 (+v0.221.1). Evidence: audits/DRILL-r361-2026-08-22/evidence/. Reasoning kept: DumpOne writes <stack>-<dbtype>.sql — the app's canonical dump, the name the replay loop matches EXACTLY — so nothing else may ever be written to it. The fix is a DESTINATION, not a rename: DumpOneTo takes the final path and derives its own .tmp from it, so neither the destination nor the scratch file can collide with a nightly dump running beside it. DumpOne's signature did not move — it has callers outside this concern. The manifest no longer lists the undo copies: every consumer of Manifest.DBDumps was grepped and named — three, all inside recovery_unit.go, none reading it for recovery. AND THAT CHANGE MADE ANOTHER UNREACHABLE: a stable db_dumps let CaptureRecoveryUnit's already-current early return fire, and the undo-copy prune sat after it — four copies on disk against a cap of three, counted live. The prune now runs ABOVE the check; it is housekeeping on the dump directory and is independent of whether the manifest needs rewriting. PROVEN LIVE the only way it can be: the canonical dump's sha256, unchanged across a restore — docmost 5d35678349bb…, bookstack 7837aa5de295…, both byte-identical before and after. A test asserting merely that the undo copy exists passes just as well when the app's backup was destroyed. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.221.1, 2026-08-23) | full text: git show a8caa0fdde7c:documentation/backlog/OPEN-ITEMS.md | | R-379 | The pre-restore undo copy was valid, was named to the customer, and no product action could apply it. Shipped in controller v0.220.0 (+v0.220.1, v0.220.2). Evidence: audits/DRILL-r379-rollback-2026-08-22/evidence/. Reasoning kept: R-379 and R-380 were ONE failure with ONE fix — both ended with a half-restored database and the only difference was whether it looked broken. The undo set is matched on THE RUN'S OWN STAMP, never on the pre-restore- prefix (four copies coexisted on one app in one afternoon; a prefix match replays an arbitrary older state) and never just the first file (a two-database app would have had one restored and the other left half-written). The rollback RE-DISCOVERS the container — the undo file is stable, the container is not: the DB-only start re-creates it, and v0.220.0's own first live run held an app for 30 s of waitDBReady against a dead id while its data was recoverable. No unit test saw that: they all inject the import seam and never look at container identity. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.220.1, 2026-08-22; docmost and bookstack both rolled back to byte-identical prior state) | full text: git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md | | R-380 | A failed MariaDB replay left a partially-applied database behind an app reporting health=healthy. Shipped in controller v0.220.0. Evidence: audits/DRILL-r379-rollback-2026-08-22/evidence/13-step2-verify.txt. Reasoning kept: no engine flag closes this — --single-transaction was added to the Postgres import and does make it all-or-nothing, but MariaDB's DDL is not transactional, so a partial apply there is unavoidable at the engine. The flag is a belt; the rollback is the fix, and this row must not be read as saying otherwise. Proven live: bookstack's migrations table back at 102 rows, the exact cell the defect was measured in. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.220.0, 2026-08-22) | full text: git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md | | R-381 | The restore-failure message pasted raw engine stderr — including rows out of the customer's own database — into the Hungarian customer surface. Shipped in controller v0.220.0. Reasoning kept: the full engine text now goes to the operator log, which never had it before — the diagnostic was ADDED, not removed. Measured: 407 bytes (Postgres) and 615 (MariaDB, whose middle was an INSERT INTO migrations VALUES (…) listing); now 257 bytes with no engine tokens. A red-proof for this PASSED and the test was hollow: it injected below ImportDump, so a leak reintroduced inside ImportDump could not fail it. The guard now sits at that layer. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.220.0, 2026-08-22) | full text: git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md | | R-382 | The reconstitution's summary log line omitted the volume count it already held. Shipped in controller v0.220.0. Proven live: 0 file(s) placed, 3 volume(s) replayed, 1 DB dump(s) replayed. | CLOSED — SHIPPED (controller v0.220.0, 2026-08-22) | full text: git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md | | R-356 | The off-site restore refused every app that has no data drive — it asked "does this app have an HDD path?" to answer "is this app installed?", and for 40 of 53 catalogue apps the honest answer to the first is permanently no. Shipped in controller v0.219.0. Evidence: audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/. Reasoning kept: the restore destination is resolved by the SAME rule as the capture destination — the drive if the app has one, the system data path otherwise (Manager.GetAppDrivePath, one expression). The refusal that protects a drive app from being restored onto the wrong disk applies to apps that HAVE a drive to get wrong. An app with no drive is not misconfigured — 01-topology-and-trust.md §8 carries the [DESIGN] marker; between 19 and 22 August that design was called a defect four times. Deployment is asked of ListDeployedStacks() and FAILS CLOSED on a nil provider: "cannot tell" must not become "go ahead" when the caller's next act is a write. Two different failures get two different sentences — installed-but-no-resolvable-data-root has its own refusal and its own route; widening nincs telepítve to cover it would send a customer to reinstall a running app and hide the real fault. Measured, and load-bearing: 53 templates, 13 needs_hdd: true, 40 false (catalogue @ 459766cb1639). The capture side's raw GetStackHDDPath is FENCED and was not changed — capture resolves an app's declared userdata/import file legs against that value, and a system-data fallback there would write a snapshot claiming to hold files it does not. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.219.0, 2026-08-22; privatebin on demo-hp: planted, backed up, deleted, restored, 15/15 files byte-identical including two Hungarian accented names) | full text: git show e18668f9e19f:documentation/backlog/OPEN-ITEMS.md | | R-216 | A correct recovery code was reported to the customer as wrong. Shipped in 0.120.0, v0.125.0. | SHIPPED (controller v0.201.0 + hub v0.97.0/0.97.1) — but see R-223: the feature does not work on a NEW box until the manifest vouches agent 0.125.0. Until then such a box is correctly HELD, not lied to | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-218 | Succeeding at recovery stopped the box asking for what it still needed. Shipped in v0.203.0. Evidence: documentation/tests/part4-rewalk-2026-08-06/journal.md. | CLOSED 2026-08-06 — controller v0.203.0, proven live. (State corrected 2026-08-06: this field read REOPENED while the body below already recorded the fix shipped and proven. The history of the over-claim is kept deliberately — it is why the row is worded as it is.) The over-claim, as it stood: the fix covered the DECLARATION half only. Measured on the R-201 re-walk: the box declared, and offsiteheal re-staged the secret at 11:44:57 saying "the box re-consumes on its next cycle" — the next cycle came and went (host-report 11:55:46, Received report 11:55:54, a full cycle with a positive control that it ran) and the credential was still not consumed. 23 minutes after the re-stage the box's last off-site-apply attempt was still the pre-re-stage one. A census of the customer-reachable actions on /backups/remote (config, reset, run, toggle) found none that fetches a staged credential, and the only lever is systemctl restart felhom-controller-bootstrap.service inside the guest — which worked in 18 s (Campaign 11 measured 17), confirming nothing was wrong with the credential, the target or the key: the only thing missing is anything at all to trigger a retry. This is the FIRST of the two dead ends that keep the recovery journey failing | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-219 | The listing the screen promises could never render on the shape it exists for. | SHIPPED (controller v0.201.0) — the unlock now places the key, brings the tier up, then lists | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-217 | An unreadable store reported as "opened, with unattributable content". | SHIPPED (controller v0.201.0) — opened / empty / unreadable are three distinguishable states | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-222 | Reaching for a RETAINED earlier package read as a wrong code. | SHIPPED (controller v0.201.0 + hub v0.97.0) — the ACK carries superseded_present/superseded_at and the screen names the situation. It states what the hub knows and promises nothing — the read path is still unbuilt (R-199's inventory) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-215 | GET /recovery rendered the recovery story on a box that never had off-site backups. | SHIPPED (controller v0.201.0) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-220 | After a rebuild the customer's drives cannot be re-enrolled, and the refusal names an impossible action. | CLOSED 2026-08-06 — shipped in agent v0.127.0 and PROVEN LIVE on a genuinely rebuilt box. The fix is corroborated, not a widened prefix: a mountpoint outside /mnt/felhom-drives is forgiven only when the SAME device is also mounted under the managed path — a pairing only Felhom's own enrolment produces, so a disk another system is using at /srv/data or even /mnt/someone-elses-disk is still refused (own test + red-proof). Read from /proc/mounts deliberately: the lsblk invocation is pinned verbatim in the sudoers file, so switching to plural MOUNTPOINTS would have shipped a sudoers change with the binary. Fail-safe: an unreadable mount table corroborates nothing. Measured on the Part 4 venue after a real guest purge, with both raw mounts still present on the surviving host: /disks/candidates returned both drives in attach and initialize (before the fix: two empty lists), and both re-attached through the customer endpoint (registered: true). The customer-facing refusal was corrected in controller v0.203.0. | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-221 | A rebuilt box cannot run the escrow ceremony at all. | CLOSED 2026-08-08 — agent v0.128.0. Apply re-asserts the seed BEFORE the idempotent early return; the return itself is kept and pinned by a zero-Proxmox-calls assertion. The writer was established at file:line rather than assumed — see the follow-through section below | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-223 | The Day-0 manifest vouched agent 0.120.0 while the recovery feature needs 0.125.0 — and a reinstall DOWNGRADES a box that was fixed by hand. Shipped in 0.120.0, 0.125.0, 0.192.0. | CLOSED 2026-08-05. Golden 0.201.0 baked in the drill VM (658 165 766 B, sha e730d7cab343eb35…f007654, round-trip verified from Gitea), then manifest set in one save: agent=0.125.0 golden=0.201.0 min_agent=0.125.0. A fresh install now lands on current agent AND current controller | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-224 | Every non-code failure on the unlock path is reported to the customer as a statement about their code. Reasoning kept: And R-216's gate cannot catch it: the box's own ring reads recovery capability gate: offsite_key_recovery=yes (source=version) — the gate discriminates the agent's age, not its reachability, so a dead agent of the right version sails through the guard whose own comment says "An attemp | CLOSED 2026-08-06 — controller v0.202.0 + agent v0.126.0. The discriminator is now a VALUE: escrow.ErrBundleFetch → HTTP 502 at the agent, agentapi.RecoveryRefusal carrying the status at the controller, and ClassifyRecoveryFailure mapping it to one of five classes from the value, never the text. PROVEN LIVE on the venue, same wrong code, only the hub's reachability changed: hub up → 400 "…did not open the sealed bundle" · hub REJECTed → 502 "…could not be fetched — the recovery code was NOT used" · hub restored → 400. Red-proof: deleting the agent case reproduces got 400, want 502 with the wrong-code sentence. Coupled MinAgent 0.126.0 — an older agent answers 400 for both causes, so the reading is withheld and the 400 degrades to NEUTRAL; the gate blocks nothing. The customer-facing messages were NOT re-driven end-to-end: /recovery correctly redirects since F7 set the old data aside, and restoring that state is the reconfiguration §11 forbids — they are covered by handler tests + red-proofs | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-225 | The remote store reports 0 pillanatkép · 0 / 50 GB when the box cannot read it — directly above a card stating the store holds backups. | CLOSED 2026-08-06 — controller v0.202.0. StatsKnown is a named state (the OffsiteInventory.Empty pattern), because zero is what an unread store and an empty one both look like and omitempty makes "absent" and "0" the same bytes. The fill bar renders only when the fill is known — a 0 %-wide bar is a picture of emptiness. PROVEN LIVE both ways: before a run the venue read „a pillanatképek száma még ismeretlen"; after one, „2 pillanatkép … / 50 GB". A measured zero still says zero | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-226 | M1 — the only message that tells a customer to check their typing — is unreachable on any box that has re-escrowed. | CLOSED 2026-08-06 — controller v0.202.0. The retained-package message now names both possibilities and restores the ten-words prompt, because the two are indistinguishable at the engine and saying so is the honest thing. It still does not promise the earlier package can be opened. Red-proof: removing the clause makes the prompt unreachable again | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-228 | After „I do not want the old data", the set-aside history becomes invisible — the box records where it is and shows it to nobody. | CLOSED 2026-08-06 — controller v0.202.0. OrphanedRenamedTo is surfaced as two facts and stops. It does not promise the history can be reopened — it cannot be, by anyone, today (R-199's inventory is unbuilt) — and the set-aside confirmation copy was corrected for the same reason: "a helyreállítási kód nélkül többé nem lesznek megnyithatók" implied that WITH the code they could be. The field's own comment said "recovery-code-recoverable", the same over-promise in the code. PROVEN LIVE: the notice renders on the venue | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-227 | A controller restart mid-unlock returns a raw English Bad Gateway. | CLOSED 2026-08-06 — controller v0.202.0, partially and stated as such. The layer that answers is traefik, whose config this repo generates — but traefik v3 serves no static files, so a branded proxy page needs a new always-up container for every 502 on the box: scoped, not built. Shipped: the unlock posts via fetch and answers a gateway failure in Hungarian in-page. Progressive enhancement — with no JS the plain POST still shows the proxy's error | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-234 | An off-site run reports success while silently omitting an app the customer just switched on. Shipped in v0.205.0. | CLOSED 2026-08-06 — controller v0.205.0 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-236 | After a guest rebuild the hub never re-stages the off-site credential. | CLOSED 2026-08-06 — not a defect | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-237 | After a successful recovery the customer is shown no backups at all, because the restore surface is keyed on apps that are currently installed and currently marked for future remote backup. | CLOSED 2026-08-06 — controller v0.204.0: the list is now built from OffsiteInventoryList (the repository's own snapshot tags). Installed-ness became a property OF a row, never a filter; an unreadable store renders as UNKNOWN and keeps the action offered; felhom-offbox and _shares are excluded. 7 new tests incl. a rendered-page test for the rebuilt shape, and a red-proof that keys the list back on installed-and-toggled apps. | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-238 | „Teljes visszaállítás előkészítése" accepts the click and does nothing. Shipped in v0.204.0. | CLOSED 2026-08-06 — controller v0.204.0 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-239 | The fixes are written, tested, pushed — and a machine installed tonight gets none of them. Shipped in 0.127.0, 0.203.0, 0.204.0. Evidence: tests/finalwalk-r201-2026-08-07/journal.md, tests/golden-0.205.0-2026-08-07/. | CLOSED 2026-08-07 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-241 | The credential self-heal, succeeding, locks the customer out of their own recovery. Shipped in v0.206.0, v0.98.0. Evidence: audits/SPIKE-r241-recovery-offer-2026-08-07.md, tests/finalwalk-r201-2026-08-07/journal.md. | FIXED 2026-08-07 — v0.206.0 / hub v0.98.0 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-247 | The box is being told something false, in its own words, and it recommends the destructive act. Shipped in v0.206.0. | CLOSED 2026-08-08 — controller v0.209.0. The field is received and the box tells the two conditions apart; see the Campaign-12 follow-through section below. The WRONG FLAG itself is R-246 (operator act, hub-side) and the customer-facing card copy is unchanged — both stated rather than folded in | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-249 | The retrieval passphrase ships in the customer page's HTML, so any headless read puts it in a transcript. Shipped in v0.206.0, v0.207.0. Reasoning kept: Severity MEDIUM: it is a live per-customer secret that fetches the whole config (GET /api/v1/config/<id> with X-Retrieval-Password), but the exposure is to someone who can already read the operator page — a defence-in-depth failure, not a boundary crossed. | CLOSED 2026-08-08 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-252 | After a rebuild the restore refuses because the data drives are not registered, and nothing on the recovery path says so. Shipped in v0.207.0. | CLOSED 2026-08-08 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-253 | The restore page promises it will reinstall the app, and the restore then refuses because the app is not installed — in the customer's own language, three lines apart. Shipped in v0.207.0. | CLOSED 2026-08-08 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-212 | The orphaned-ciphertext deletion HALTED: the stores on the storage box do not match this register's record. | CLOSED 2026-08-05 — all three deleted after the operator confirmed the corrected list | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-94 | A hand-synced version constant drifts, and the gate that would catch it is never run Shipped in 1.22.0, 9.9.9. | CLOSED — SHIPPED (hub v0.87.0, 2026-08-02) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-110 | main is the installer's publish channel — there is no staging. Shipped in 1.22.0, v1.23.0, v4.4.0. Reasoning kept: E-2a's felhom-backup-target-apply (:2116) is installed 0755 to /usr/local/sbin and root-fenced in sudoers, validated only by bash -n — a root-executed artifact taken from main with no pinned integrity, which is this row's class exactly. Fixing only (i) leaves a tagged installer pulling nine untagged files from main at run time — a staging story that is false in the place it matters most, since one of those nine (felhom-backup-target-apply) is installed 0755 into /usr/local/sbin and root-fenced in sudoers, validated only b | CLOSED — SHIPPED (installer v1.23.0, 2026-08-03) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-111 | The Day-0 artifact channel is 17 agent releases stale — a box installed today gets agent 0.96.0, not 0.113.0. Shipped in 0.113.0, 0.114.0, 0.161.0. Evidence: audits/E2D-fresh-vm-2026-07-29.md. Reasoning kept: The global controller floor was deliberately NOT raised: the golden now bakes 0.185.1, so a fresh box needs no self-update, and raising it would have been an unnecessary fleet-wide write. | SHIPPED 2026-07-29 — the channel now serves agent 0.113.0 + golden 0.185.1 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-115 | Publishing is a remembered step, and it was forgotten within eight hours of being documented as forgettable. Shipped in 0.113.0, 0.114.0, 0.119.0. Reasoning kept: Class: → R-29, one layer up — a control that exists and is never walked; deliberately NOT given its own ID. It calls the existing publish-agent.sh rather than reimplementing it, refuses a dirty or unpushed tree, refuses to re-release an existing version (one version name must never mean two binaries), and deliberately does not vouch — vouching points machines at a version and stays the operator's ac | CLOSED — SHIPPED (release-agent.sh + check-published-versions.py, 2026-08-03) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-116 | The drive-absent alarm and its recovery were a MISMATCHED PAIR — absent fired the GENERIC storage_disconnected, return the SPECIFIC backup_target_restored; backup_target_absent never fired at all Shipped in 0.185.1, v0.115.0, v1.25.0. Evidence: audits/R116-v0116-2026-07-30.md, audits/SPIKE-r117-bind-liveness-2026-07-30.md. | SHIPPED + PROVEN-LIVE (agent v0.116.0, 2026-07-30) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-120 | The golden baked a controller that predated R-114 + R-112, so a FRESH box showed the customer the WRONG absent-target message Shipped in 0.113.0, 0.116.0, 0.156.0. Evidence: audits/R120-golden-rebake-2026-07-30.md. | CLOSED — golden rebaked + PROVEN-LIVE, and the class now has an ENFORCED gate (golden 0.186.0 + hub v0.82.0, 2026-07-30) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-117 | A drive's guest bind becomes a DEAD MOUNT while every signal reads healthy — and it happens in TWO ways, only one of which the original framing covered. Shipped in 0.113.0, 0.117.0. Evidence: audits/R117-v0117-2026-07-30.md. Reasoning kept: No block I/O proven by strace (only /proc/self/mountinfo, 0 statfs) — the Part 1 CLAUDE.md fence applied to its own first consumer. | SHIPPED + PROVEN-LIVE (agent v0.117.0, 2026-07-30) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-113 | The drive-absent gate CANNOT FIRE on device loss — E-2b's alarm is wired to an unreachable condition. Shipped in 0.113.0, 0.114.0, v0.185.0. Evidence: audits/E2D-fresh-vm-2026-07-29.md, audits/SESSION-C-2026-07-29.md. | SHIPPED + PROVEN-LIVE (agent v0.114.0, 2026-07-29) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-112 | E-2's degraded banner and offer have NO UI CONSUMER — the endpoint is correct and the customer never sees it. Shipped in v0.185.1. Evidence: audits/E2D-fresh-vm-2026-07-29.md, audits/SESSION-C-2026-07-29.md. | SHIPPED + PROVEN-LIVE (controller v0.186.0, 2026-07-29) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-114 | On target-drive loss the customer is told the wrong story and offered the drive that just vanished. Shipped in 0.113.0. Evidence: audits/E2D-fresh-vm-2026-07-29.md, audits/SESSION-C-2026-07-29.md. | SHIPPED + PROVEN-LIVE (controller v0.186.0, 2026-07-29) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-29 | The green gates are not enforced anywhere — one was RED for 16 releases before anyone ran it. Shipped in 1.19.0, 1.22.0, v0.129.0. | CLOSED — both halves shipped (2026-08-02) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-86 | Restore-tests are interval-scheduled, not backup-aligned | CLOSED — SHIPPED + PROVEN-LIVE 2026-08-03 (agent v0.121.0, hub v0.91.0) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-185 | The agent cannot see the host backup tier's archives on demo-felhom — the PVE token has no ACL on /storage/felhom-backup, so the content listing returns EMPTY where root sees three archives. Shipped in v0.123.0. | CLOSED — SHIPPED + PROVEN-LIVE 2026-08-03 (agent v0.123.0, installer 1.24.0) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-186 | A released agent binary's sha256 cannot be reproduced from its tag. Shipped in v0.120.1, v0.121.0, v0.121.2. | CLOSED — SHIPPED + MEASURED 2026-08-03 (agent v0.122.0) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-187 | R-115's one-command release had never actually run its publish leg — the first real use died there. Shipped in v0.121.0. | CLOSED — SHIPPED 2026-08-03 (felhom-agent) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-188 | Every agent release has a ~50 % chance of emailing the operator a CI failure for a release that is correct. Shipped in 0.121.2, v0.121.0, v0.121.1. | CLOSED — SHIPPED 2026-08-03 (agent v0.122.0) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-189 | A passing restore-test can be invisible to the hub forever — and R-86 made that window a week instead of a day. Shipped in v0.121.1. | CLOSED — SHIPPED + PROVEN-LIVE 2026-08-03 (agent v0.122.0) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-195 | A customer with no machine ever bound e-mailed an expected_dbdump_missed ERROR every morning. Shipped in v0.73.0. Reasoning kept: Fail-open on a read error (an unreadable binding must never SUPPRESS a real alarm), and the deferral is LOGGED with its own counter (the v0.73.0 Part-7 precedent: a quiet check must not look like a check that did not run). | SHIPPED (hub v0.92.0, 2026-08-04) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-196 | escrow_stale is wired to the ONE path that does not change the repo password, and absent from the path that does. Shipped in v0.95.0. Evidence: audits/DRILL-r201-night-run-2026-08-04.md, audits/SPIKE-offsite-credential-recovery-2026-08-04.md. | CLOSED 2026-08-05 — hub v0.95.0 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-197 | The hub holds both halves of the evidence that a box's offsite DATA key changed, and reads neither. Shipped in v0.78.0, v0.93.0. Evidence: audits/SPIKE-offsite-credential-recovery-2026-08-04.md. Reasoning kept: The in-between shapes (a first-ever hash, a hash-less supersession) are LOGGED rather than dropped, so "we chose not to alarm" and "the check did not run" never look identical. | SHIPPED (hub v0.93.0, 2026-08-04) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-198 | The hub's superseded-escrow retention does NOT retain the offsite repository password — and the ceremony the system tells the customer to run is what destroys the last copy. Shipped in v0.92.0, v0.93.0. Evidence: audits/RECON-offsite-dr-chain-2026-08-04.md. | SHIPPED (hub v0.93.0, 2026-08-04) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-199 | The hub serves recovery blobs on two endpoints that have no client anywhere in the system. Evidence: audits/RECON-offsite-dr-chain-2026-08-04.md. | SHIPPED + PROVEN-LIVE 2026-08-04 (hub v0.94.0, agent v0.125.0) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-203 | A customer-declared MANDATORY data directory was silently absent from the off-site snapshot while the run reported ok. Evidence: audits/DRILL-r201-offsite-recovery-2026-08-04.md. | SHIPPED + PROVEN-LIVE 2026-08-04 (controller v0.197.0) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-204 | A rebuilt box can recover its off-site key and still cannot use it: the remedy that reconfigures the tier is the thing that blocks the recovery. Shipped in v0.198.0, v0.199.0, v0.95.0. Evidence: audits/DRILL-r201-night-run-2026-08-04.md. Reasoning kept: Live on demo-felhom 9201, nothing restarted (restarts=0, container older than both mints): the superseded code returned „Hibás vagy lejárt kód" and the current one was accepted first time. Item 2 (a re-issue marks a healthy escrow stale) — CLOSED, → R-196. Test-proven; deliberately NOT fir What is deliberately NOT automated: the escrow ceremony. A credential is replaceable; the recovery code is not. | ALL FOUR ITEMS CLOSED 2026-08-05 (items 1–3 controller v0.198.0 + hub v0.95.0; item 4 controller v0.199.0 + hub v0.96.0) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-192 | offsite_delivery_stuck tells the operator the opposite of what the detector measured, and the self-heal silently refuses for exactly the reason the message denies. Shipped in 0.187.0, 0.192.0, v0.199.0. Evidence: audits/RECON-offsite-dr-chain-2026-08-04.md, audits/SPIKE-offsite-credential-recovery-2026-08-04.md. | CLOSED 2026-08-05 — the guard's scoping half closed BY REPLACEMENT (hub v0.96.0) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-193 | A guest rebuild silently drops the off-site app-data tier, and nothing restages the credential. Shipped in 0.156.0, 0.187.0, 0.192.0. Evidence: audits/RECON-offsite-dr-chain-2026-08-04.md, audits/SPIKE-offsite-credential-recovery-2026-08-04.md. | CLOSED 2026-08-05 — controller v0.200.0 (credential half v0.199.0/v0.96.0; the recovery SCREEN v0.200.0) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-90 | ep0 RAM headroom — 4 GiB swap survived its first reboot 2026-07-27; 3.8 GB RAM unchanged | CLOSED — the operator rescaled ep0 to a CX33 on 2026-08-03 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-97 | Whole-guest backup tier had no hub signal; quiesce blamed the apps Shipped in v0.79.0. | SHIPPED (controller v0.177.0 + hub v0.78.0/v0.79.0, 2026-07-27) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-100 | ~~A restic offsite tier that fails every night never goes stale on the hub — isStale counted from LastRun Reasoning kept: The real defect is defeated defence in depth: the hub-side pull net was anchored on a field the failing controller keeps refreshing, so it could not compensate for a lost push (cf. | SHIPPED + PROVEN-LIVE (controller v0.181.0 + hub v0.80.0, 2026-07-28) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-101 | ~~Tier-2 LastRun is written on failure and rendered to the customer as „Legutóbbi másolat" — including in t | SHIPPED + PROVEN-LIVE (controller v0.182.0, 2026-07-28) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-108 | Network storage can host an app's namespace, and FileBrowser binds a network share at its ROOT Evidence: audits/R108-network-app-namespace-2026-07-30.md. | SHIPPED + PROVEN-LIVE (controller v0.187.0, 2026-07-30) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-109 | The DR recipe records no backup target Evidence: audits/R106-R109-recipe-completeness-2026-07-30.md. | SHIPPED + PROVEN-LIVE (agent v0.118.1 + hub v0.83.0, 2026-07-30) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-106 | The DR recipe records the PBS namespace as "root" on every box | SHIPPED + PROVEN-LIVE (agent v0.118.1, 2026-07-30) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-122 | ~~AssembleDRRecipe silently DROPPED offsite_restic — the offsite recovery location never reached any reci | SHIPPED (hub v0.83.0, 2026-07-30) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-125 | A "test through the production path" is only true up to the seam it injects at. Shipped in v0.118.0. Evidence: audits/R106-R109-recipe-completeness-2026-07-30.md. | FIXED (agent v0.118.1) — filed for the DOCTRINE point | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-128 | ~~build-felhom-iso.sh:44 comments that ISO_VERSION "aligns with felhom-host-install SCRIPT_VERSION" — a c | CLOSED (iso v1.26.0, 2026-07-31) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-154 | [first-boot] is automated-install-only and nothing in the Felhom tree said so Evidence: audits/SPIKE-universal-iso-3-2026-07-31.md. | CLOSED (iso v1.26.0, 2026-07-31) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-155 | iso-repack.sh refuses any ISO without auto-installer-mode.toml, blocking the no-answer.toml posture | CLOSED (iso v1.26.0, 2026-07-31) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-156 | An app's data is neither persisted nor backed up, and it reports healthy. Shipped in 26.6.1. Reasoning kept: Provenance, stated because it decides the row: the observation is docker ps -a on demo-hp's guest 9201 returning empty, supplied with the 2026-08-02 task; this session did not re-measure (documentation-only, every box fenced). | CLOSED — all three apps fixed (papra template, 2026-08-03) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-157 | bootrecon's start-ONCE sweep misses the boot orphan it exists to recover — TWO mechanisms. | CLOSED — SHIPPED + PROVEN-LIVE (B: controller v0.189.0; A: v0.190.0, 2026-08-02) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-170 | The drive-backed boot gate infers a customer's Stop from a container count. Shipped in v0.190.0. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.190.0, 2026-08-02) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-171 | The boot sweep started apps whose data drive was ABSENT — a regression introduced by v0.189.0, now FIXED. Shipped in v0.189.0. Evidence: audits/DIAG-bootrecon-drive-absent-2026-08-02.md. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.190.0, 2026-08-02) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-172 | A false host_stale alarm fires when the hub's SQLite refuses two consecutive host reports. Reasoning kept: Retry options (b) and (c) were deliberately NOT taken — with readers no longer blocking writers a surviving SQLITE_BUSY would be a real signal, and a retry would hide it; revisit only on evidence. | CLOSED — SHIPPED + PROVEN-LIVE (hub v0.88.0, 2026-08-02) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-174 | The app-stop guard's crash recovery started apps onto MISSING drives — a regression in v0.189.0 code. Shipped in v0.189.0. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.191.0, 2026-08-02) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-175 | 07-backup-architecture.md §7.5 states ONE box's size bound as if it were the fleet's. Shipped in 0.192.0, 7.5.1. Evidence: audits/SPIKE-r165-mp1-merge-2026-08-02.md. | CLOSED — FIXED 2026-08-03 (same pass as R-165) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-183 | A fresh install fetched the vouched agent BINARY and its sixteen CONFIG files from two different refs, and nothing compared them. Shipped in v0.120.0. Reasoning kept: Why it is a defect and not only untidiness: these files are the agent's own operating surface — its systemd unit, its sudoers, its guarded wrappers — and configs/felhom-backup-target-apply is installed 0755 into /usr/local/sbin and root-fenced in sudoers, validated only by bash -n. | CLOSED — SHIPPED (installer v1.23.0, 2026-08-03) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-182 | A full disk tells the operator about ONE app and silently swallows every other app's refusal for an hour. Shipped in v0.194.0, v0.90.0, v0.90.1. | CLOSED — SHIPPED (controller v0.194.0 + hub v0.90.0/.1, 2026-08-03) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-181 | The capture floor guards the cheap leg and not the leg that fills the volume — and its refusal message asserts an invariant the code does not provide. Shipped in v0.192.0, v0.193.0, v0.193.1. | CLOSED — SHIPPED (controller v0.193.0 + v0.193.1, 2026-08-03) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-178 | The merged golden (0.192.0) is built and published but NO BOX HAS BEEN REINSTALLED FROM IT, and it is deliberately UNVOUCHED. Shipped in 0.119.0, 0.120.0, 0.192.0. | CLOSED — BOTH BOXES REINSTALLED AND PROVEN (2026-08-03) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-158 | A local Tier-1 app-data backup failure reaches no hub channel. Shipped in v0.78.0. | CLOSED BY R-167 — SHIPPED + PROVEN-LIVE (controller v0.191.0 + hub v0.89.0, 2026-08-02) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-159 | wishlist's data landed in an ANONYMOUS volume — never backed up, orphaned by a redeploy. | SHIPPED (templates/wishlist/docker-compose.yml, 2026-08-02) — filed to record the CLASS | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-160 | gramps-web persisted three paths and wrote to none of them. | SHIPPED (templates/gramps-web/docker-compose.yml, 2026-08-02) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-163 | mp1 is RETENTION, not staging — and it is sized as if it were neither. Shipped in v0.192.0. | CLOSED by R-165 — the ceiling it describes no longer exists (golden v3.0.0, 2026-08-03) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-165 | Merge mp1 into mp0 — the dedicated 20 G backup partition stops existing. Shipped in 0.192.0, v0.192.0. Evidence: audits/SPIKE-r165-phase0-2026-08-03.md. | SHIPPED — golden build-golden.sh v3.0.0 + agent v0.120.0 + controller v0.192.0 (B2), 2026-08-03. IMPLEMENTED — the LAYOUT is proven live on both boxes (R-178, 2026-08-03); the BULKHEAD'S REPLACEMENT IS NOT (→ R-181) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-166 | App state gets a desired/observed model with its own store. Shipped in v0.189.0. | SHIPPED + PROVEN-LIVE (controller v0.189.0, 2026-08-02) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-167 | Storage monitoring and backup alerts. Shipped in v0.191.1, v0.191.2. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.191.0/.1/.2 + hub v0.89.0, 2026-08-02) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-168 | CI: no runner exists, and with trunk-based pushes CI can DETECT but not BLOCK Shipped in 0.1.0. Evidence: audits/SPIKE-ci-runner-2026-08-02.md. | SHIPPED — and the alarm is DEMONSTRATED (2026-08-02) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-205 | RootFsPressureDespiteHousekeeping can never fire Evidence: audits/SPIKE-dooplex-buildcache-2026-08-05.md. | CLOSED — SHIPPED + RED-PROVEN LIVE (homelab-manifests 6808a4b, 2026-08-05) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-258 | C3 — the customer's per-app backup tick is green on the PRESENCE of a restore point, and its only red condition is a GLOBAL one. | CLOSED 2026-08-08 — controller v0.210.0. appDumpVerdict reads THIS app's own dump result; three states, no icon when nothing is known. Recency deliberately not added — see the observation in the follow-through section | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-259 | C4 — a disk read that FAILS renders as „0.0 GB / 0.0 GB (0%)" in the nominal colour, on the dashboard's most-looked-at meter. | CLOSED 2026-08-08 — controller v0.210.0. readDiskUsage reports success; SystemInfo.DiskKnown/HDDKnown; the template draws no figure, no percentage and no meter fill when unknown. The hub leg is deliberately NOT fixed and is now R-266 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-260 | C5 — the agent reports at least eight decision-bearing facts the hub models NOWHERE, and the sharpest one blinds the check that answers „can the operator get into this box". | CLOSED 2026-08-08 — the class is GATED (G-1, scripts/wire_contract_gate.py) and the sharpest instance is fixed (hub v0.99.0). The remaining unconsumed facts are R-264, OPEN — allowlisted with reasons, which is not the same as decided. See the follow-through section below | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-265 | A CI run can fail with NO LOG PERSISTED, and the alarm mail then points the operator at a log that does not exist. | CLOSED 2026-08-08 — timeout-minutes: 5 on the gates job, and the alarm mail now states elapsed seconds and qualifies its own "names itself in the run log" sentence. ⚠ The unknown is NOT closed and must not be read as closed: whether the if: failure() alarm fires for a REAPED job is still unverified. The timeout makes the reap unreachable in practice; it does not answer what happens inside one | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-267 | The Configuration page is 2.6× faster and is still ~10 s, and the remaining cost is ONE Gitea call whose latency swings 20× with load. Shipped in 0.100.2, v0.100.0. | CLOSED 2026-08-08 — hub v0.101.0 + a registry prune. Final: cold 5.4 s, warm 0.14 s (was 26.2 s). Three serialisation legs took it to 9.85 s mean, the 60 s in-memory memo took the warm path to a quarter-second, and the prune halved what remains of the cold path. ⚠ TWO CORRECTIONS TO THIS ROW'S OWN EARLIER TEXT, because both were wrong and both mattered. (1) "Only 50 generic versions exist" WAS NOT A COUNT, IT WAS A PAGE LIMIT. ?type=generic&limit=1000 returns at most 50; the 50 I measured was exactly the cap, and three older agent versions (0.81.0, 0.80.0, 0.79.0) only became visible after the first 30 deletions moved them onto page one. An unpaginated listing is not evidence of a total — this repo's own "an empty listing is not evidence of emptiness" rule, walked into while measuring. (2) THE OPERATOR'S "REDUCE THE NUMBER OF ARTIFACTS" WAS THE BETTER CALL AND MY MEASUREMENT SAID OTHERWISE. I reported it helps "sub-linearly" and "is not the lever". Measured after: trimming to 10+10 took the COLD load from 13.4 s to 5.4 s — a 2.5× improvement on the path the memo cannot help, because the fan-out is per-version. Recorded rather than quietly dropped (the R-96 standing rule). Pruned to the newest 10 per package on the operator's rule, with the live-vouched golden/agent/floor asserted into the KEEP set before a single DELETE was issued; 33 deletions, all HTTP 204, and golden 0.210.0 / agent 0.128.0 / agent 0.127.0 verified still fetchable afterwards. drill-r50 runs agent 0.113.0, now deleted — flagged to the operator first; it is a disposable nested drill VM and only its re-download path is gone | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-268 | A live per-guest local-API token was printed into a session transcript. Shipped in 169.254.253. | CLOSED — ROTATED + PROVEN LIVE 2026-08-09 (rehearsal pre-phase, audits/REHEARSAL-byo-reinstall-2026-08-09.md §3) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-273 | RANK 1 — the hub vouched an agent version that was never git-tagged, and every install fleet-wide now fails at step 5/8. Shipped in 0.127.0, 0.128.0, v0.127.0. | CLOSED 2026-08-09 — tag pushed, install PROVEN | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-278 | demo-felhom's off-site tier has never completed a run and has been stuck for six days. Shipped in 0.200.0, v0.93.0. | CLOSED 2026-08-10 — protection RESTORED, and the recovery it waited for could never have worked | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-280 | RANK 1 — after a reinstall the data drive cannot be re-attached through ANY dashboard route, and the restore page promises it is "two clicks". | CLOSED — controller v0.211.0, delivered via golden 0.211.0 (vouched 2026-08-10) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-281 | The hub said NOTHING through an entire reinstall — and the tripwire for a sealed-backup unseal did not fire on a real unseal. | WITHDRAWN 2026-08-09 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-293 | CENSUS, 2026-08-10 — no machine that is not ours can be in the state that cost demo-felhom its history, and here is the whole population. | CLOSED-INFORMATIONAL 2026-08-10 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-294 | The orphan card promises restorability that the box rendering it cannot evaluate — specified, not implemented. Evidence: documentation/design/SPEC-orphan-card-copy-2026-08-10.md. | CLOSED — controller v0.211.0; see R-299 for the sentence it missed | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-296 | The orphan card's OTHER sentence makes the same promise, and the spec says it is fine. | CLOSED — shipped in controller v0.212.0 (R-299); verified: the sentence at backups_remote.html:98 was replaced and the stem guard covers it | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-297 | An install took whatever golden was lying around. Shipped in 0.153.0, 0.210.0, 0.213.0. | CLOSED — observed live + PUBLISHED as installer-v1.27.0 (both refs bumped) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-299 | The orphan card's OTHER sentence made the same unevaluable promise, and the spec called it accurate. Shipped in v0.211.0. | CLOSED — controller v0.212.0 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-300 | Our own uninstall left the thing that makes our own reinstall refuse. Shipped in 0.0.0, 10.0.2, 127.0.0. | CLOSED — observed live + PUBLISHED as installer-v1.27.0 (both refs bumped) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-301 | The abandon countdown banner makes the retired promise a third time, and as a flat statement. Reasoning kept: the customer chose to abandon a recovery offer that exists — which is why it was NOT changed (this session was fenced to the orphan card). | CLOSED — premise CONFIRMED and fixed in controller v0.213.0 (R-302) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-302 | The abandon banner promised retrieval it could not see was still true — fixed by PINNING a fingerprint at the decision. Shipped in v0.213.0. Reasoning kept: Empty is not a match on either side; a countdown started before v0.213.0 carries no pin and takes the cautious branch (deliberately NOT backfilled). | CLOSED — controller v0.213.0 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-305 | The R-300 cleanup fires exactly once per machine, and the second reinstall hits the original wall. Shipped in 0.0.0, v1.27.0. | CLOSED — superseded by R-316 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-307 | demo-felhom carries a LIVE abandon countdown that this drill did not start — and the end state says there should be none. Reasoning kept: The drill's fence forbade starting, shortening or triggering a countdown, and none was; but its required end state was "no abandon countdown anywhere", and one exists. | CLOSED — countdown cancelled 2026-08-12 on the operator's ruling | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-308 | The stored controller password no longer opens demo-felhom — WITHDRAWN 2026-08-12, this was MY BUG, not a defect. | WITHDRAWN — not a defect (my error) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-309 | The day-0 runbook says pushing the installer publishes it. It has not since R-110. Shipped in 1.25.0, 1.27.0. Evidence: documentation/runbooks/day0-install.md. | grep -m1 '^SCRIPT_VERSION'. **Measured while writing it: served 1.28.0, main 1.28.0, both pins installer-v1.28.0— the three agreeing is the observation; any one alone is not.** **The claim was copied elsewhere and the copy was hunted:**audits/SPIKE-universal-iso-3-2026-07-31.md:184said the same thing and **citedday0-install.mdas its source**, which is how it spread. It was **true on the day it was written** (R-110 shipped 2026-08-03), so the dated finding is kept verbatim and carries a SUPERSEDED note rather than being rewritten — falsifying a dated record to tidy it is its own defect. Two other hits are correct in context:hostinstall_gates.py:198states the consequence of the manifest LOSING its tag, and the 2026-08-12 drill record already names the sentence as false | full text:git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-311** | **A correct recovery code for a retained package stopped being reported as wrong.** Shipped in 0.126.0, 0.128.0, 0.129.0. | **CLOSED — shipped + delivered: hub v0.103.0 + agent v0.129.0 + controller v0.214.0** | full text:git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-316** | **The removal now genuinely reverses the installation — R-305's once-per-machine defect closed.** Shipped in 0.0.0, v1.27.0, v1.28.0. | **CLOSED — shipped + published, observed on the cycle that actually fails** | full text:git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-318** | **No honest marker exists that says Felhom installed dnsmasq on a machine already in the field, and none can be invented.** Shipped in v1.27.0. **Reasoning kept:**/var/log/dpkg.logdoes record the install — and is a **timestamp**, which the standing rule refuses as a heuristic dressed as a fact. | **CLOSED — established, no action possible for existing boxes** | full text:git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-319** | **The guest-network watchdog finally has a reader — the first of R-264's twenty-one, and it is the repair COUNT that matters, not the state.** Shipped in 0.92.0, v0.92.0. | **CLOSED — shipped hub-side 2026-08-13** | full text:git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-320** | **Evidence has been destroyed twice in three days, in the same place, by the same act.** Shipped in v1.28.0. Evidence:audits/DRILL-retained-key-2026-08-12.md, audits/REPORT-r316-installer-v1.28.0-2026-08-13.md. **Reasoning kept:** **The rule, now standing rule 5 in workspace-CLAUDE.md(so it loads in every session) and repeated where a session actually meets it —runbooks/target-selection.md, RUNBOOK-rehearsal-v3.md, and the PROMPT-TEMPLATE.mdreport section: evidence is copied off the machine at the end of the phase | **CLOSED — rule written, four homes** | full text:git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-321** | **A box on which reporting is deliberately switched off still alarms as stale, then down.** Shipped in v0.105.0. | **CLOSED — shipped hub v0.105.0, both doors** | full text:git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-322** | **The claim guard has never scanned the hub, and the hub sends the customer's first sentence.** **Reasoning kept:** **Recommended shape, and the reason it is not one line:** the gate is invoked bycontroller_gates.py, so pointing it at a sibling repo makes a controller gate fail on a felhom.eu edit — the cross-repo lesson from G-1 (a gate needing a sibling passes locally and exits INCONCLUSIVE in CI, and must n | **CLOSED 2026-08-13 by R-324** — scripts/hub_copy_gate.py, registered in repo_gates.py, scanning 95 hub files for retired names and four declared customer surfaces for retrieval stems, with a plant→convict→remove→pass selftest that caught a defect in its own instrument on the first run. The stem list IS shared (scripts/customer_copy_vocab.py) and no controller gate was made to depend on a felhom.eu clone; the controller gate's adoption of the shared list is R-325, and until it happens the two are drift-checked rather than left to diverge | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-323** | **The third near-homograph — the five-word phrase is „Tulajdonosi jelmondat” now.** | **CLOSED — shipped hub v0.105.0** | full text:git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-324** | **The hub's customer copy is under a guard for the first time — and the guard has been watched catching, ignoring and releasing.** | **CLOSED — shipped, selftest green** | full text:git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-326** | **"Which claims are unproven?" is a question a machine can answer now — and the number everyone was repeating answered a different question.** | **CLOSED — shipped** | full text:git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-328** | **The disk alert was emailed to nobody, and one word is the whole reason.** Evidence:audits/DIAG-smart-passed-trap-2026-08-14.md. | **CLOSED — controller v0.215.0, PROVEN LIVE 2026-08-14.** Now "warning", and DiskAlertKind.Severity()is exported so the contract is assertable from any package rather than duplicated as a literal. **The proof is a side-by-side pair pushed through the REAL hub event endpoint** from demo-hp's controller: severity"warning"→ storedwarning, notification_log**id 689, channeloperator, status sent**; the identical push at "warn" → stored **info**, and **no notification_logrow exists at all**. Pinned byTestNotifyDiskHealthDegraded_SeverityRoutes, which asserts membership of the hub's accepted set (not just the literal) and names both hub locations; its red-proof — restoring "warn"— fails all three assertions | full text:git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-334** | **CLOSED 2026-08-18 — golden 0.216.0 baked, published and VOUCHED; CI green by run id.** Shipped in 0.214.0, 0.215.0, 0.216.0. Evidence:documentation/tests/golden-, documentation/tests/golden-0.214.0-2026-08-12. | **CLOSED 2026-08-18.** Baked from RUNBOOK-manual-build.md §4.0+§4.1 in the DooPlex drill VM and published: **GOLDEN_VERSION=0.216.0**, **GOLDEN_SHA256=ac004dc90d8cefccc5448377892f9cff3a4c3e1e27d0e11129120e38ac31c34b**, 656,970,239 bytes at …/generic/felhom-golden/0.216.0/golden.tar.zst. **The published bytes were verified, not just the script's print** — the artifact was downloaded back out of Gitea and hashed, and it matches. **Vouched by the operator, all THREE fields together**, confirmed by reading the hub's own store rather than the save: artifact_golden_version=0.216.0, artifact_agent_version=0.129.0, artifact_min_agent=0.129.0(2026-08-18 11:00:59–11:01:00), and the hub's recorded sha256 matches the downloaded artifact. The R-216 shape was checked on the machine:MinAgent 0.129.0 is **equal to**, not above, the newest **published** agent. **golden_currency_gate.pyrc=0 andrepo_gates.py --fastrc=0 — all nine gates — and CI is GREEN BY RUN ID: run **353**,head_sha 7d81681d6, conclusion success** (the two prior runs 351/352 on this same afternoon were red on exactly this row, which is the contrast). That push needed **no --no-verify** — the first of the day that did not. Evidence: documentation/tests/golden-0.216.0-2026-08-18/, report REPORT-golden-0.216.0.md. **Closed with the run id quoted deliberately**: this row was re-confirmed once and widened once, and closing it on a local green a third time would have left the same ambiguity | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-335** | **One physical disk was walked TWICE per run, and the second walk sustained it against itself.** Shipped in v0.215.0. **Reasoning kept:** **This is the shape standing rule 3 warns about: an absent alarm was not evidence — the two artefacts had to be read AGAINST each other** — — CC | **CLOSED — controller v0.216.0, 2026-08-14.** EachdiskKeyis evaluated once per run; both entries stay markedseenso neither looks like a disappeared disk, and the card still renders both storage rows (the dedup is about state and alerts, not display). Pinned byTestDiskCheck_SameDiskTwiceIsEvaluatedOnce; companion red-proof run and reverted — deleting the guard makes the first sighting emit Kind:2(Hiba-from-sectors) at 8 sectors | full text:git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | **R-344** | **felhom-agentleaks one TCP connection to PBS per poll cycle, forever, on both sides — and it is the whole of the ep0 descriptor leak.** Shipped in 0.129.0, 0.130.0. Evidence:audits/SPIKE-ep0-established-connections-2026-08-20.md. **Reasoning kept:** **The proof obligation is the fd count, not the diff:** per standing rule 3 the positive observable is ep0's ESTAB count going FLAT between proxy restarts, measured over a window long enough to matter — a green test suite proves nothing here, and a 30-minute window proves nothing here either (that e **control 4 cycles -> 4 leaks; fixed 4 cycles -> 0 leaks.** **Positive observable per standing rule 3** (a zero leak is equally consistent with "the agent stopped working"): the fixed box's four poll cycles are in ep0's log, and the boxes' other traffic is near-identical (libwww-perl 924 vs 926, pro | **CLOSED 2026-08-20 — fixed, proven live on both boxes, published and vouched** | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-347** | **The R-344 fix exists on two demo boxes by hand and NOWHERE ELSE — a box installed from the current image still ships the leaking agent.** Shipped in 0.129.0, 0.130.0, 0.216.0. Evidence:documentation/runbooks/publish-train-rules.md. | **CLOSED 2026-08-20 — published, vouched, and the fleet reconciled onto the published bytes** | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | **R-351** | **The restore never read back where the backup said the data lived, and a second press started a second restore.** | \.NamespaceRoot\b' --include=*.go found no non-test reader anywhere — the reconstitution opened the manifest (offbox_reconstitute.go:235) purely for the coherence stamp and resolved its destination from the LIVE app instead. A restore into a destination different from the recorded one therefore succeeded silently, under a green message. (b) The second press. All seven restore handlers gated on backupMgr.IsRunning() — the CONCURRENCY flag, acquired inside the goroutine (offbox_reconstitute.go:180) after the handler returned. Established with a test before any change: both the reconstitute and place handlers answered „…elindult" and overwrote the first restore's op/stack. The wizard had read the correct flag since v0.154.0 and said so in a comment; the handlers were never moved over. (c) The banner gated its terminal result on a page-local sawRunning, so a restore that finished before the page opened — the 8.666 s OpenGist restore — was shown to nobody. | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-354 | The off-site full restore has NO named-volume leg — the tar is in the unit, in the snapshot and in the checking folder, and is never replayed. Shipped in 0.217.0, 0.218.0. | CLOSED — SHIPPED + PROVEN-LIVE 2026-08-22 (controller v0.218.0) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-355 | paperless-ngx's PostgreSQL is dumped into a directory for a stack that does not exist, so its unit has never contained a database dump — and the destructive restore therefore takes no safety dump and tells the customer the app has no database. Shipped in 0.217.0, 0.218.0. | CLOSED — SHIPPED + PROVEN-LIVE 2026-08-22 (controller v0.218.0) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-339 | The hub was SILENT when it lost sight of the off-site stores — and a 9 h 37 m outage proved it. Reasoning kept: That is correct for a fill signal — a missing reading must never be mistaken for 0%, which is why degraded data drives no band transition — but the consequence was that a completely dead off-site endpoint and a healthy one were indistinguishable on the operator channel. | SHIPPED — hub v0.106.0, 2026-08-18. Reachability is now a second, independent signal: consecutive failed fetch windows counted per checker, pbsdr_box_unreachable / offsite_box_unreachable (severity warning) past a default 3 windows (≈30–45 min), with paired *_recovered all-clears wired into recoveredPairedDownTypes — necessary because both recoveries are severity info and severityNotifies drops info. Threshold tunable via alerting.box_unreachable_windows. The fill logic is untouched: no threshold, throttle, band or escalate-once behaviour changed. Evidence: internal/monitor/box_reachability_test.go (Scenarios A–F) + internal/notify/dispatcher_box_reachability_test.go (the cross-package wiring, asserting an actual operator mail), plus three companion red-proofs each seen failing with a message naming the right cause | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-370 | PROCESS: between 2026-08-19 and 2026-08-22 the reviewing side called a documented architectural decision a defect, in four places, because it read the register and live source and never documentation/architecture/. Evidence: documentation/architecture/. Reasoning kept: R-352 (re-framed), R-369 The record is corrected in place with the framing marked rather than deleted, per the standing rule that a document which quietly changes its mind teaches nobody. | CLOSED — corrected 2026-08-22 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-96 | Two standing rules were agreed in chat and never committed Evidence: documentation/runbooks/workspace-CLAUDE.md:48-70. Reasoning kept: Two standing rules were agreed in chat and never committed MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-27, size XS, roadmap state idea — found 2026-07-27. Moved verbatim; nothing added or reinterpreted. | CLOSED — migrated from ROADMAP 2026-08-22 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-107 | No offsite action unpacks the named-volume tars Tier-3 captures on every run. Shipped in v0.218.0. | CLOSED — migrated from ROADMAP 2026-08-22 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-383 | The double-failure message told the customer their previous state was saved, and named a file that was not there. Shipped in controller v0.222.0. Evidence: audits/DRILL-r384-dead-db-alarm-2026-08-23/. Reasoning kept: One of the two ways a rollback fails is that the undo copy is missing — so the sentence was most likely to be false in exactly the case it was printed. Do NOT simply drop the filename: an operator needs it, and R-351's lesson is that a refusal naming nothing forces someone to remember what the product already knows — so the absent case still names WHERE the file should have been. A zero-length dump counts as MISSING, because a 0-byte file restores nothing and calling it present is the same false reassurance one step smaller. The check is os.Stat and deliberately not an integrity test: this runs at the end of a failed restore on a machine that may be unwell, and presence is the honest claim available there. | CLOSED — SHIPPED (controller v0.222.0, 2026-08-23; undoCopyPhrase, four cases, plus an AST seam test that the message is still wired to the builder) | full text: git show 1eb64bec5183:documentation/backlog/OPEN-ITEMS.md | | R-384 | An app whose DATABASE had died raised no dead-app alarm — the wrong question answered first. Shipped in controller v0.222.0. Evidence: audits/DRILL-r384-dead-db-alarm-2026-08-23/. Reasoning kept: The defect was the ORDER of two questions, not the unhealthy exclusion. "Is a SUPERVISED member dead?" and "is a RUNNING member failing its healthcheck?" are different questions, and the second was answering the first — a dying database drags its own front end unhealthy, so the symptom the fault causes was what suppressed the alarm for it. IsDownState is byte-identical and unhealthy stays excluded — an unhealthy container is RUNNING, and folding it in reintroduces the flapping that exclusion exists to stop; no new state was minted, StateDegraded already means this. Two things had to move and either alone leaves the defect standing: the hoist, AND widening "some members are up" from running > 0 to any member not in the down bucket — the old guard made the R-51 block unreachable in precisely the case it was written for. The register's own suggested fix was WRONG and is recorded as such: it proposed a sustained-unhealthy threshold on the crashLoopAfter model; the actual defect needed no threshold at all. PROVEN LIVE the only way it can be — the same fixture that printed 0 currently down on 2026-08-22 printed 1 currently down on 2026-08-23, with app_start_failed 7 s after the stop and the banner reading „…nem fut: BookStack (degraded)". Scenario D measured 0 alarms across 9 scans through a full stop→start cycle. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.222.0, 2026-08-23) | full text: git show 1eb64bec5183:documentation/backlog/OPEN-ITEMS.md | | R-329 | app_start_failed was emitted with severity "warn", so every one of them was delivered to nobody. Shipped in controller v0.223.0 (+ hub v0.107.0). Evidence: audits/DRILL-r329-r386-2026-08-23/. Reasoning kept: The vocabulary is EXACT and it is the HUB's, not ours — {info, warning, error, critical}; anything else is coerced to info at ingest and dropped by severityNotifies before BOTH legs. This was the SECOND occurrence (DiskAlertKind.Severity until v0.215.0), and its comment had recorded the lesson — a comment is not a guard, so the guard is now an AST walk over the whole controller, with the six variable-passing call sites registered by name because a walk cannot follow a variable and an unlisted limit is not a limit, it is a hole. The register's own framing was that the DECISION was the work — should a stopped app mail the customer at all? Answered: operator always, customer OFF by default, because processOperator never consults customer preferences, so one word fixed the operator leg and left the customer leg exactly where the ruling wanted it. Deliberately NOT added to operatorOnlyEvents — that would make the new toggle visible, flickable and structurally incapable of delivering. Measured on the live hub DB: 91 events stored all-time, ZERO notification rows before the fix; one operator row, warning/sent, after it. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.223.0 + hub v0.107.0, 2026-08-23) | full text: git show 68a9f5475cd2:documentation/backlog/OPEN-ITEMS.md | | R-386 | A single-container app stopped out of band raised no alarm, and a comment stated the opposite as settled fact. Shipped in controller v0.223.0. Evidence: audits/DRILL-r329-r386-2026-08-23/. Reasoning kept: the state test was guessing at something the product already knows. DesiredState records the customer's intent, has exactly one writer, and is tri-state; StateExited never survives aggregation, so no state test can separate an out-of-band stop from a customer stop. The ruling: Stopped → no alarm, Running → alarm, absent → UNKNOWN, keep today's behaviour AND announce it. Reading unknown as "nobody asked" would, on the first cycle after upgrade, e-mail about every app any owner ever deliberately stopped — fleet-wide, from a field that predates the intent it is being asked about. A rule without a mechanism is a wish: every such suppression sets IntentUnknown and the names are logged at INFO, so an operator can answer "how many apps am I blind to?". failedRestart must still lift a Stopped intent or F-CRIT-1 re-opens. Fenced act: adding a DesiredState WRITER — twelve of StopStack's fourteen callers are machines. Proven live: alarm 24 s after an out-of-band docker compose stop, heartbeat 1 currently down against the previous day's 0; and with intent removed, suppressed plus the log line naming the app. 0 of 8 deployed apps on demo-hp carry an absent intent. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.223.0, 2026-08-23) | full text: git show 68a9f5475cd2:documentation/backlog/OPEN-ITEMS.md | | R-389 | Only the FIRST broken app per hour reached the operator — the cooldown key named the event type, not the app. Shipped in hub v0.108.0. Evidence: audits/DRILL-cooldown-grain-2026-08-23/. Reasoning kept: the fix is a THIRD SIBLING of cooldownTierSuffix/cooldownRunSuffix, separate for the reason the second one's docstring already gives — the existing two keep byte-identical semantics for every type that uses them. cooldownStackSuffix takes the EVENT TYPE as well as the details, unlike its siblings, and that asymmetry is the whole safety property: tier and run_id appear only on types that want that grain, stack_name does not. perAppCooldownEvents is a named allow-list with app_start_failed and nothing else — the backup family's cooldown is coarse ON PURPOSE (R-97a, R-182) so one full disk sends one digest rather than one mail per app, and this is not hypothetical: crossdrive_failed is severity error, reaches the operator leg, and carries stack_name through a different struct, so a payload-shape rule would have split it silently. The fenced act is adding an entry for a type whose family has a digest or a coarse-by-design cooldown. app_start_failed qualifies precisely because it has NO digest — there is no apps_down_run the way backup_run_failures summarises a run. The hour is unchanged; the grain was the complaint. Fail-soft: absent or malformed details degrade to the old key and the mail still goes. PROVEN LIVE 2026-08-23: two apps four minutes apart gave 2 sent / 0 suppressed where the same shape gave 1 and 1 the day before, each repeat suppressed under its OWN key (…:opengist, …:calibre-web) against the previous day's shared key=demo-hp:app_start_failed; and crossdrive_failed for two different apps stayed coarse under key=demo-hp:crossdrive_failed, byte-identical to the derived v0.107.0 value. AND IT WAS NEVER FILED UNTIL THE DAY IT WAS FIXED — it lived in a REPORT.md observations paragraph, which is why gate 11 now exists. | CLOSED — SHIPPED + PROVEN-LIVE (hub v0.108.0, 2026-08-23) | full text: git show 45659bdc5a2f:documentation/backlog/OPEN-ITEMS.md |

2026-08-30 — the off-site store gets checked (controller v0.227.0/v0.227.1)

Two rows closed, one CORRECTED and deliberately left open. Full original text: git show <this commit> -- documentation/backlog/OPEN-ITEMS.md.

ID Title Shipped Evidence
R-359 The off-site restic store was never verified by anything, ever controller v0.227.0/v0.227.1 documentation/tests/r359-integrity-2026-08-30/
R-397 NotifyIntegrityOK/NotifyIntegrityFailed had no caller, and the product advertised a weekly check that did not exist controller v0.227.0 documentation/tests/r359-integrity-2026-08-30/

R-398 is deliberately NOT a row here. It was proposed for closure in the same pass and was CORRECTED instead — the premise was wrong, the seam already existed — so it stays in OPEN-ITEMS.md as the record. It is written as prose rather than a table row because a register row in both files is exactly what closed_register_gate.py convicts on.

The rules these leave behind:

  • The integrity check TAKES the single-writer flag and SKIPS rather than waits. resticStep escalates to unlock --remove-all on a lock error and is only safe while every caller holds that mutex; a check without it can strip a LIVE prune's lock. Never remove that guard.
  • Due-ness, not a weekday. R-341 is the other shape: a dated check quietly missed and never caught up. No Weekly primitive was added; a daily job that asks "is it due?" catches up after downtime.
  • "I could not look" is not "I looked and it is broken". Skipped / unreachable / failed are three facts. A timeout is unreachable, never damage. A failure advances due-ness; a skip does not.
  • A success that mails nobody is a design choice, not a gap. backup_integrity_ok is severity info and is dropped before both delivery legs. A weekly success e-mail is how alerts stop being read.
  • ⚠ THE STRUCTURE CHECK DOES NOT CATCH SILENT CORRUPTION. Measured: a pack corrupted without a size change returned no errors were found, exit 0. Only --read-data* caught it. An ok at the shipped depth means the index and the snapshot graph are sound — narrower than the word suggests. That is R-399, open.
  • A damage classifier must match PHRASES, not words. "pack ", "tree " and "snapshot " all appear in restic's ORDINARY progress output; the first draft would have called a healthy run corrupt. The negative control caught it — which is why a control that has only seen the failing case is worth nothing.
  • The restic exec seam has always existed (SetOffboxRunner). R-398 said otherwise and was wrong.

2026-08-30 — the restore tells the truth (controller v0.226.0) + R-395

Six rows closed. Full original text: git show e027b5d9 -- documentation/backlog/OPEN-ITEMS.md. Compressed here to title, shipping version, evidence, and the sentences that state a RULE.

ID Title Shipped Evidence
R-353 A local unit restore reported a bare completion whether it returned an entire dataset or nothing controller v0.226.0 documentation/audits/evidence-r353-r360-live-2026-08-30/live-validation.txt — live sentence A(z) opengist: 1 adatkötet visszaállítva — az alkalmazás újraindult.
R-357 The destructive reconstitute had no free-space gate; all three that existed guarded non-destructive paths controller v0.226.0 seam tests only — NOT live-validated, by design
R-358 OffboxFullScratchReady asked "non-empty directory", which is what a failed restic run leaves controller v0.226.0 documentation/audits/evidence-r353-r360-live-2026-08-30/live-validation.txt — marker {"schema":1,…,"full":false}, gate logged place-to-live closed
R-360 The verification-copy delete gated on the concurrency flag, which a verification restore never holds controller v0.226.0 documentation/audits/evidence-r353-r360-live-2026-08-30/live-validation.txt — refused in the live flag state; planted canary survived
R-396 A unit-only verification restore unlocked the DESTRUCTIVE full restore controller v0.226.0 same evidence; found while answering R-358's open question
R-395 STATUS.md contradicted itself about the controller version doc fix, same session STATUS.md at e027b5d9

The rules these leave behind — the reason the rows are kept rather than deleted:

  • A restore outcome is a claim about THE BACKUP, never about the app. R-355 extended to the Tier-1 path. The off-site twin has SafetyDump as an honest discriminator; the local path has none, so no claim about the app is available to it at all. Not merely unproven — unprovable from a manifest: §6.3 records that an absent dump has causes that say nothing about the app.
  • Zero-replayed has two causes and they are opposite news. "The backup held no data" and "the backup listed data that did not come back" must never share a sentence.
  • The destructive reconstitute uses NO headroom margin, matching PlaceOffsiteRestore. The ×1.1 elsewhere exists because that gate PREDICTS a download; this one measures a tree that already exists.
  • Fail closed when a probe reads ≤ 0. free < need with need == 0 is FALSE, so an unmeasurable input sails through — a gate present and inert, which is worse than no gate because it reads as protection.
  • A hidden button is not a guard. Template enable-flags control a button; the handler must refuse.
  • One boolean must not drive three intents (R-396): ScratchReady answered "is there a scratch" while being consumed as "may we place" and "may we destructively restore".
  • Never restate a version in a second place on the same page (R-395). Live versions belong in the hub, never in a doc.
  • A doc comment claiming a guard exists is why nobody looks for the missing guard (R-360). Correct such a sentence in place; do not delete it. | R-399 | How deep should the off-site integrity check go — Viktor ruled full depth. Shipped in controller v0.228.0, 2026-08-31. Evidence: felhom-controller/REPORT.md (v0.228.0) — restic argv observed from the guest at both depths on demo-hp. Reasoning kept: the structure check does not detect a size-preserving pack corruption — measured 2026-08-30, plain restic check reported no errors were found and exited 0 over a damaged pack that every read-data form caught. That is the reason for the default and it is what should stop anyone turning it back down to save four seconds. An empty value means "not configured", therefore the default; off is the off token, because a setting with no off switch is not a setting. A malformed value falls back to the DEFAULT, never to structure — falling back to structure would silently remove the protection on a typo, which is R-357's shape. Superseded by R-401 for anything about a large store. Original text: git show 300d7e8:documentation/backlog/OPEN-ITEMS.md | | R-400 | A third of the debug page posted to endpoints that did not exist — and three of the seven fetched on page LOAD. Shipped in controller v0.228.0, 2026-08-31. 24 referenced / 17 dispatched became 18 / 18. backup/crossdrive implemented (proven live: real Tier-2 copies for three apps); backup/infra, hub/infra-push, dr/infra-status, storage/watchdog-status and both storage/simulate-* deleted with their panels and JavaScript. Reasoning kept: implement or delete FIRST, register the gate SECOND — a registered-but-failing gate refuses every push. Keep handleDebugAPI's exact-match switch with its NotFound default; a prefix match would have made the defect invisible instead of merely silent. A panel left behind renders nothing forever, which is how this class hides. A debug control that simulates or mutates storage state is deleted unless a live need can be shown — that is where drives get unenrolled and data gets stranded. Enforced by controller/scripts/debug_route_gate.py, both directions, red-proofed. Original text: git show 300d7e8:documentation/backlog/OPEN-ITEMS.md | | R-102 (was C9-F4) | Tier-2 wrote a full recovery-unit/ mirror on every run and no code path read it - RecoveryUnitPath joined a hard-coded backups/primary/, so in the one failure Tier-2 exists for the surviving copy was unopenable. Shipped in controller v0.229.0: four unit-directory-relative path primitives in appbackup, RestoreFromRecoveryUnitAt(stack, unitDir), RestoreTier2Unit. Evidence: audits/DRILL-r102-tier2-unit-2026-08-31/. Reasoning kept: THE SOURCE MOVES; THE DESTINATION DOES NOT - unitDir changes only where a unit is READ from; data still lands in the live volumes and the live database container, resolved by GetAppDrivePath exactly as the capture is, because a restore that also relocated an app's data would be a migration wearing a restore's label. And: a directory that exists is not a package - the Tier-2 route refuses fail-closed unless the mirror carries a parseable manifest. | CLOSED 2026-08-31 - controller v0.229.0, PROVEN-LIVE with the primary unit moved aside (07 §8 row 3b -> PROVEN, 28.65 s; row 4 stays PARTIAL - the drive-loss JOURNEY is still unexercised) | full text: git show 1623a4d5b5d5:documentation/backlog/OPEN-ITEMS.md | | R-103 (was C9-F1b) | The Tier-2 no-coverage refusal named the working action but did not route to it - it sent the customer to a button on another page for data that R-102 made restorable on the page they were already looking at. Shipped in controller v0.229.0: POST /backup/tier2/unit-restore and „Teljes visszaállítás a másolatból” on the Tier-2 row. Evidence: audits/DRILL-r102-tier2-unit-2026-08-31/. Reasoning kept: a destructive operation reached from a non-destructive surface must carry the difference in the CONFIRM, not in the label - the two actions stay two buttons because they are two promises, and the confirm names the copy's date, differently when that date is only an attempt clock (R-101). And: two questions, two predicates - CanRestore() was NOT widened to cover the unit; one predicate answering two questions is R-356, which refused 40 running apps for months. And: tier2UnitNotCoveredMsg was NOT deleted, because it is appended where the FILE restore ran and is still exactly true of it. | CLOSED 2026-08-31 - controller v0.229.0, PROVEN-LIVE (the refusal now carries tier2UnitAvailableMsg, verified at the endpoint) | full text: git show 1623a4d5b5d5:documentation/backlog/OPEN-ITEMS.md | | R-87 | The restic tier was never restore-tested — RE-SCOPED by its own spike to "prove the off-site snapshot still CONTAINS a recoverable unit". Shipped in controller v0.231.0 + hub v0.110.0. Evidence: tests/r87-offsite-proof-2026-08-31/; reasoning: audits/SPIKE-restic-restore-test-2026-08-31.md. Reasoning kept: The weekly check proves the stored bytes are the bytes we stored; it cannot tell us we stored the WRONG thing. The acceptance rule has TWO parts and the obvious one is a trap — "everything declared is present" passes a hollow unit, which is the shape it exists to catch. The expectation comes from INSIDE the unit, never the live box: the snapshot may predate the app's shape. The volume half is an EXISTENCE check and not a name match — the naming held on all eight real units, but "held on eight" is not "derivable" (R-355), and half a rule that is true beats a whole rule that is invented. THREE outcomes: pass, fail, and cannot-judge — collapsing the third hides a gap in one direction and alarms on our own blind spot in the other. It proves the snapshot CONTAINS a recoverable unit; it does NOT prove a restore puts data back into a running app — §8 matrix row 4 was deliberately NOT moved. The proof's scratch is a SEPARATE root because the job deletes on every path, and sharing the customer's root would mean a nightly job deleting a copy the customer is looking at. | CLOSED 2026-08-31 — SHIPPED + PROVEN-LIVE (controller v0.231.0, hub v0.110.0) | full text: git show 303129e:documentation/backlog/OPEN-ITEMS.md | | R-406 | Two unrelated findings shared the identifier R-133. Resolved 2026-09-01 by renumbering the hub-uniqueness finding to R-415. Reasoning kept: citations were MEASURED before choosing — 3 for hub-uniqueness, 5 for the plaintext break-glass credential — and the FEWER-cited one moved; this is the opposite of the task's literal instruction, whose stated ground ("the older number has the longer reference trail") the measurement contradicts; the principle was followed and the letter was not; the within-register duplicate rule was deliberately NOT added in the commit that removed its only subject — a guard whose red-proof can only be a planted fixture is not this project's standard (R-416). | CLOSED — RENUMBERED (2026-09-01) | full text: git show 22e1c95:documentation/backlog/OPEN-ITEMS.md | | R-410 | golden_currency_gate.py was satisfied by a DIRECTORY NAME — mkdir turned it green with no bake behind it. Shipped in felhom.eu, 2026-09-01. Reasoning kept: a directory name is a label; GOLDEN_SHA256=<64 hex> is a fact only a completed publish produces; the self-test ships a POSITIVE CONTROL, without which "it fails on an empty directory" would be satisfied by a gate that fails on everything; directories that look right and hold nothing are printed by name rather than silently ignored, so a half-finished bake is visible. | CLOSED — SHIPPED (felhom.eu, 2026-09-01) | full text: git show 22e1c95:documentation/backlog/OPEN-ITEMS.md | | R-414 | The nightly off-site proof was INERT on a box with no registered data drive, every night, with only a WARN. Shipped in controller v0.232.0. Evidence: audits/R411-R414-2026-09-01/, determination in 00-part2.1-determination.md. Reasoning kept: the scratch resolver was consciously OUT OF SCOPE for R-356, not excluded — its own tests say "the scratch still resolves … only the DESTINATION moves"; the fallback is SCOPED because the two callers ask different questions, and one predicate answering both is the R-356 defect itself — unit-only may fall back (§7: a driveless app's unit already lives on the system data path, "intended, not a defect"), a full restore may not (§2.2: state-only tier); absence on last_proof_result already means "controller too old", so a second meaning on one field is the StatsKnown trap one level up; a cannot_run is recorded but does NOT advance per-snapshot due-ness, or the app would never be retried once a drive is registered. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.232.0, 2026-09-01) | full text: git show 22e1c95:documentation/backlog/OPEN-ITEMS.md | | R-407 | restic check DOES write a lock file, and the comment above it said it never writes. Corrected in controller v0.232.0. Reasoning kept: corrected in place, not deleted — R-360's rule is that a comment claiming a guard is why nobody looks for the missing one; the same paragraph now carries the fact that restic stats also takes a lock. | CLOSED — CORRECTED IN PLACE (controller v0.232.0, 2026-09-01) | full text: git show 22e1c95:documentation/backlog/OPEN-ITEMS.md | | R-408 | RestoreOffboxScratch took no single-writer flag while a comment asserted every off-site operation did. Shipped in controller v0.232.0. Reasoning kept: the real deliverable is the WALK, not the acquire — the sentence was false for months and nothing checked it, the ninth instance of this project's most-repeated class; it is an AST pass and not strings.Contains, because a commented-out call still contains the string; adding a line to offsiteExempt is a deliberate act and belongs in the commit that adds it; the R-87 proof's exemption is kept HONEST by a second test that fails if that path ever gains unlockStale, routes through resticStep, or loses --no-lock. | CLOSED — SHIPPED (controller v0.232.0, 2026-09-01) | full text: git show 22e1c95:documentation/backlog/OPEN-ITEMS.md | | R-411 | A background job deleted the lock of a live customer restore and logged it as a crash that did not happen. Shipped in controller v0.232.0. Evidence: audits/DRILL-soak-2026-08-31/phase1-lock-collision/, audits/R411-R414-2026-09-01/. Reasoning kept: restic stats TAKES a repository lock — the fact nobody had, and the one that made the chain reachable; restic check takes one too, restic snapshots and restic list do not; the fix was wider than the row — FOUR entry points were unflagged, three of them found by R-408's walk rather than by the report; the escalation in resticStep was NOT removed — real stale locks exist and it clears them; the defect was that a sibling could be live. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.232.0, 2026-09-01) | full text: git show 22e1c95:documentation/backlog/OPEN-ITEMS.md | | R-403 | A poorer copy deleted a richer one: an EMPTY recovery unit on the primary drive was mirrored over a COMPLETE copy on the second drive, with --delete. Shipped in controller v0.230.0. MEASURED before it was fixed — on the shipped v0.229.0, on demo-hp: 120 082 104 B (4 database dumps + 3 volume tars) -> 7 036 B (none of either) in one nightly run, recorded as a success. Evidence: audits/DRILL-r403-tier2-delete-2026-08-31/. Reasoning kept: hollowness is a MANIFEST question, never a size question - a unit with a fat compose capture and no dumps is the dangerous shape and a 360-byte unit for a tiny app is healthy; absent or unparseable manifest counts as hollow, fail closed. The guard fences ONE shape and not shrinking - 07 §8 row 5's derived-copy rebuild is a DESIGN DECISION, --delete stays, the data legs are untouched, and only source-hollow-over-destination-complete is refused (§8.2 records the exception beside the rule so nobody 'fixes' it back). The rehydrate happens INSIDE the restore - the hollow manifest was written two seconds later by the 5-minute capture job, so any follow-up job races it; and the capture is deliberately NOT guarded, because a capture describing an empty drive as empty is correct and guarding it would make the manifest lie. A warning that fires on everything costs the same as the comforting lie it replaces - the first draft flagged 'package older than the run', which is true of every healthy app, and four healthy apps on the box would have been warned. | CLOSED 2026-08-31 - controller v0.230.0, PROVEN-LIVE both ways (the loss reproduced on v0.229.0, then the same state preserved on v0.230.0 with all 7 files sha256-identical) | full text: git show 66156c619fd2:documentation/backlog/OPEN-ITEMS.md |