From 71b8c8c62bfa0b4f0960b8ec52528116f8fbf735 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Sat, 3 Oct 2026 09:17:22 +0200 Subject: [PATCH] =?UTF-8?q?Backlog=20triage=20Part=20B:=20125=20finished?= =?UTF-8?q?=20rows=20+=2020=20id-less=20rows=20moved=20to=20CLOSED-ITEMS?= =?UTF-8?q?=20(full=20text=20at=209e2786c);=20open=20rows=20normalised=20t?= =?UTF-8?q?o=20one=206-column=20shape;=20narratives=20archived=20verbatim;?= =?UTF-8?q?=20closed=5Fregister=5Fgate=20RULE=203=20refuses=20a=20finished?= =?UTF-8?q?=20row=20in=20OPEN-ITEMS=20(decoys,=20seen=20red);=20rules=20re?= =?UTF-8?q?homed=20to=20CONTEXT=20+=2007=20=C2=A711;=20loose=20notes=20tri?= =?UTF-8?q?aged;=20R-814..R-819=20filed;=20register=20444=20->=20325?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- .claude/rules/docs.md | 2 + CLAUDE.md | 3 +- CONTEXT.md | 25 + documentation/PROMPT-TEMPLATE.md | 15 +- .../architecture/07-backup-architecture.md | 14 + ...-f9-storage-registration-gap-2026-06-14.md | 0 .../{backlog => archive}/FIX-M18-NOTES.md | 0 .../{backlog => archive}/FIX-M19-NOTES.md | 0 .../FOLLOWUP-golden-default-controller-tag.md | 0 .../OPEN-ITEMS-narratives-2026-10-03.md | 361 ++++++ .../audits/DRILL-golden-098-2026-07-03.md | 2 +- documentation/backlog/CLOSED-ITEMS.md | 160 +++ ...WUP-nas-automount-guest-reboot-reassert.md | 3 + documentation/backlog/OPEN-ITEMS.md | 1047 +++++------------ documentation/backlog/README.md | 42 +- documentation/backlog/ROADMAP.md | 8 +- .../backlog/SPEC-r85-phase4-5-2026-07-26.md | 3 + scripts/CHANGELOG.md | 14 + scripts/closed_register_gate.py | 33 +- scripts/instructions_gate.py | 27 + scripts/register_table.py | 106 ++ scripts/test_gate_decoys.py | 30 +- 22 files changed, 1080 insertions(+), 815 deletions(-) rename documentation/{backlog => archive}/DIAGNOSIS-f9-storage-registration-gap-2026-06-14.md (100%) rename documentation/{backlog => archive}/FIX-M18-NOTES.md (100%) rename documentation/{backlog => archive}/FIX-M19-NOTES.md (100%) rename documentation/{backlog => archive}/FOLLOWUP-golden-default-controller-tag.md (100%) create mode 100644 documentation/archive/OPEN-ITEMS-narratives-2026-10-03.md create mode 100644 scripts/register_table.py diff --git a/.claude/rules/docs.md b/.claude/rules/docs.md index 55107bdd..a3bdb7cc 100644 --- a/.claude/rules/docs.md +++ b/.claude/rules/docs.md @@ -20,6 +20,8 @@ repo. Sibling repos point here; this is where the pointed-at thing must actually | logging levels and phrasing | `runbooks/logging-conventions.md` | | a spike or campaign result | `audits/` | | every open finding | `backlog/OPEN-ITEMS.md` | +| every finished finding, compressed | `backlog/CLOSED-ITEMS.md` | +| finished notes and history moved out of a register | `archive/` | | the operator's one-screen view | root `STATUS.md` | **Do not restate a fact that has a home** — point at it. Re-check an address rather than trusting one diff --git a/CLAUDE.md b/CLAUDE.md index 135dd880..dfdada6b 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -131,7 +131,8 @@ something, not only sessions that touch `documentation/` — which is why it is - **`CHANGELOG.md` + `REPORT.md`** in every repo touched (see the workspace root for the rule, and the parallel-session caveat above). - **`REUSE.md`**, if a shared helper or pattern moved (same commit). -- **`OPEN-ITEMS.md`** — every finding, with a number. +- **`OPEN-ITEMS.md`** — every finding, with a number. **A row you close moves to `CLOSED-ITEMS.md` in the + same commit** — `closed_register_gate.py` RULE 3 refuses a finished row left in the open register. - **Root `STATUS.md`** — at the end of every session in which something shipped, broke or was decided. It is a **view** of `OPEN-ITEMS.md`; nothing may exist only there. One screen, written for the operator in plain language, and deliberately **not** `CONTEXT.md`. diff --git a/CONTEXT.md b/CONTEXT.md index 2b26b445..ed0ffbcb 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -25,6 +25,31 @@ > reverse:** the severity scale (P1 now · P2 before the first paying customer · P3 during the first customers · P4 later) > and the eleven categories, both written at the top of `OPEN-ITEMS.md`. +> **Rules carried out of rows closed 2026-10-03 (`PROMPT-TEMPLATE.md` §N.7 item 2).** The triage moved 125 finished +> rows to `CLOSED-ITEMS.md`; these sentences stated a rule and had no other durable home. Verbatim, with their row. +> Decisions: **[DESIGN]** R-245 — *"An automatic ending should trigger on the harm, with a dated warning, never on a +> date alone."* R-303 — *"the wrong fix (suppressing the orphan card during a countdown) would hide a real second +> fault."* R-312 — *"retention is an operator-only capability, and nothing anywhere may promise the customer can +> perform it themselves"* (all three: `07` §11, "Decided, and deliberately not built"). Method and traps: R-201 — +> the drill's pass condition is *"a byte-identical sentinel sha256, not 'the store opened'"*, and *"a good snapshot +> is not durable against a later bad run on the same day."* R-209 — *"`RequiresMountsFor` on a path with NO mount +> unit is a SILENT NO-OP"*; *"`-X` is load-bearing — overlayfs stacking rides `trusted.overlay.*`."* R-520 — +> *"`pct mount` maps the guest's ROOTFS ONLY and does not apply the guest's own internal mounts"*; *"An empty +> directory is not evidence of an absent file."* R-582 — *"a guard ported between languages must be re-derived from +> what the claim IS in the new language, not translated word for word"*. R-583 — *"the surface you would use to +> CHECK a feature is the one most worth checking first — a broken instrument that reports success is worse than a +> broken feature."* R-601 — *"a 'no access' claim must list what was tried; it does not say the list makes the claim +> true"*; *"Six failed routes to a stale address are six failures of one assumption, not six pieces of evidence."* +> R-611 — *"the report's FIRST section is 'not done', even when empty"*, and every part of a brief skipped, +> shortened or changed is listed there with the reason. R-623 — *"an instrument that can report a success as a +> timeout is not a measurement"*. R-633 — *"A refusal from the wrong rule is not evidence for the new one."* R-715 — +> a catalog gate refuses a probe without a before/after measurement in its comment. R-736 — *"the Remove button runs +> `RemoveStack`, not `DeleteStack`; `docker image ls` without `-a` hides the untagged digest-pulled images"*. R-745 — +> *"the agent rolls back to the RUNNING image (what `/etc/felhom-controller-image` named), never to a previous one +> the controller hands it"* (`09` decision 56's wording may say otherwise — R-817). R-752 — a settings-only change +> reaches the stack file at the next sync (≤15 min) and the running app only at the next `compose up -d` +> (Restart/Start, an Update, or a backup's restart). + > **2026-10-02 (afternoon) — the persistence re-sweep (R-801/R-788 CLOSED), controller v0.288.0 (R-800), golden 0.288.0.** > Rulings `09` §3 66 (licences: Tandoor like SparkyFitness; Emby/Plex/n8n kept; lawyer review before the first paying > customer, R-802) and 67 (userdata stays on every remove; dialog + result name it — `07` §6.5). The gate now sends diff --git a/documentation/PROMPT-TEMPLATE.md b/documentation/PROMPT-TEMPLATE.md index 2235aeb0..b3e2dc8d 100644 --- a/documentation/PROMPT-TEMPLATE.md +++ b/documentation/PROMPT-TEMPLATE.md @@ -394,8 +394,9 @@ or **what is open changed**), update **all four** in the SAME session: > later by an overnight drill that planted files and watched them not come back**, and shipped as > R-354. The work was right the first time; only the filing was missing. -**Report which `OPEN-ITEMS.md` rows the task opened, closed or re-ranked** (§15). Every row carries -an owner — a row nobody owns is how items got lost in the first place. +**Report which `OPEN-ITEMS.md` rows the task opened, closed (and moved to `CLOSED-ITEMS.md`), narrowed +or re-ranked** (§15). Every row carries an owner — a row nobody owns is how items got lost in the first +place. ### N.6 Website version bump (if controller/hub version is shown on the site). @@ -406,10 +407,14 @@ over half of it finished work, one entry at 16 KB. **A file that cannot be read be checked**, and this project has paid for that twice — a record nobody could find because it sat inside an entry about something else, and a finding rediscovered because nobody could see it. -1. **Compress what this session closed.** A closed row keeps its title, the version it shipped in, its - evidence paths, and any sentence stating a rule. Everything else goes, and it moves to - `backlog/CLOSED-ITEMS.md`. **Nothing is deleted:** the compressed entry names the commit whose +1. **Move what this session closed — IN THE SAME COMMIT THAT CLOSES IT.** A closed row keeps its title, the + version it shipped in, its evidence paths, and any sentence stating a rule. Everything else goes, and it + moves to `backlog/CLOSED-ITEMS.md`. **Nothing is deleted:** the compressed entry names the commit whose `git show` returns the full original text. **Open rows are not touched — their detail is doing a job.** + **This is a gate, not a habit (2026-10-03):** `scripts/closed_register_gate.py` RULE 3 refuses a push + whose `OPEN-ITEMS.md` holds a row LEADING with a finished word (CLOSED, SHIPPED, FIXED, DECIDED, …). + The rule existed from 2026-08-22 without a gate and was followed only sometimes: on 2026-10-03 the open + register held 113 finished rows, a quarter of the file. 2. **Rehome live reasoning before compressing it away.** If a closed entry carries the reason a rule exists or a fence sits where it does, that reasoning moves — to `CONTEXT.md` if it is a decision, to the owning `architecture/*.md` if it is a shape. **Where it is a decision, mark the resulting diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index e666451f..ad4315f7 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -1381,6 +1381,20 @@ to now *implement* D5 remains an open scheduling decision, not a blocked one. --- +**Decided, and deliberately not built — carried here 2026-10-03 when their register rows moved to +`backlog/CLOSED-ITEMS.md`** (the reasoning is in `CONTEXT.md`, "Rules carried out of rows closed 2026-10-03"): + +- **[DESIGN] No automatic abandon on a date (R-245, decided 2026-08-07).** A customer who never decides is not + auto-abandoned after 30 days. An automatic ending, if ever built, triggers on the HARM (quota — old history + blocking new backups), with a dated warning, never on a date alone. Reopens on quota. +- **[DESIGN] `markOrphaned` keeps no guard against an active abandon countdown (R-303, decided 2026-08-13).** The + co-render is made harmless, not impossible; suppressing the orphan card during a countdown would hide a real + second fault. Reopens on a real-world sighting. +- **[DESIGN] No in-product route from the recovery screen to a set-aside store (R-312, decided 2026-08-13).** + Retention is an operator-only capability, and nothing anywhere may promise the customer can perform it + themselves. Re-evaluate on a real customer request. Its fixture, `demo-felhom`'s unrecoverable set-aside store + (R-313), is kept until R-312 ships or is abandoned. + ## 12. Evidence index | Claim | Grade | Source | diff --git a/documentation/backlog/DIAGNOSIS-f9-storage-registration-gap-2026-06-14.md b/documentation/archive/DIAGNOSIS-f9-storage-registration-gap-2026-06-14.md similarity index 100% rename from documentation/backlog/DIAGNOSIS-f9-storage-registration-gap-2026-06-14.md rename to documentation/archive/DIAGNOSIS-f9-storage-registration-gap-2026-06-14.md diff --git a/documentation/backlog/FIX-M18-NOTES.md b/documentation/archive/FIX-M18-NOTES.md similarity index 100% rename from documentation/backlog/FIX-M18-NOTES.md rename to documentation/archive/FIX-M18-NOTES.md diff --git a/documentation/backlog/FIX-M19-NOTES.md b/documentation/archive/FIX-M19-NOTES.md similarity index 100% rename from documentation/backlog/FIX-M19-NOTES.md rename to documentation/archive/FIX-M19-NOTES.md diff --git a/documentation/backlog/FOLLOWUP-golden-default-controller-tag.md b/documentation/archive/FOLLOWUP-golden-default-controller-tag.md similarity index 100% rename from documentation/backlog/FOLLOWUP-golden-default-controller-tag.md rename to documentation/archive/FOLLOWUP-golden-default-controller-tag.md diff --git a/documentation/archive/OPEN-ITEMS-narratives-2026-10-03.md b/documentation/archive/OPEN-ITEMS-narratives-2026-10-03.md new file mode 100644 index 00000000..c6eefcc4 --- /dev/null +++ b/documentation/archive/OPEN-ITEMS-narratives-2026-10-03.md @@ -0,0 +1,361 @@ +# OPEN-ITEMS — the narrative sections, as they stood on 2026-10-03 + +> **History, not a register.** These sections sat between the register tables of +> `documentation/backlog/OPEN-ITEMS.md` until the 2026-10-03 triage. They are moved here word for word +> (`git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` holds them in place). Every row they discuss is in +> `OPEN-ITEMS.md` (open) or `CLOSED-ITEMS.md` (finished). A status word below is the status ON THE DATE OF +> ITS SECTION and may be stale — the registers are the authority. The ranking paragraphs ("Why the TOP +> READY rows rank this way", "The 2026-08-02 intake, ranked") are superseded by +> `documentation/audits/backlog-triage-2026-10-03/RECOMMENDATION.md`. + +## Operator rulings — 2026-08-04 + +Recorded here because a ruling that lives only in a conversation binds nobody (the R-96 standing rule). + +1. **Run the recovery drill, after R-198.** R-198 shipped in hub v0.93.0; **the drill is the next + session.** Design: `audits/RECON-offsite-dr-chain-2026-08-04.md` §10 — demo-hp, ~3–4 h, a recovery + code created and KEPT, a sentinel file, wipe, reinstall, recover, and **pass = a byte-identical + sha256, not "the repository opened"**. → R-201 +2. **Delete the orphaned ciphertext** (~1.2 GB across the two demo boxes, in set-aside restic stores + nothing prunes). **STILL OWED — not done in v0.93.0.** It is a destructive act on a protected + endpoint and belongs to a session that is scoped for it, not to a release that ships a schema + change. → R-193 +3. **Accept the risk on R-193 candidate (c)** — no repository password retained on the Proxmox host. + **This is what makes R-198 load-bearing rather than tidy:** with no host-retained copy, the + customer-present recovery path is the ONLY way back from a rebuild, and that path runs entirely + through the retained identity blob. → R-193, R-199, R-200, R-201 + +~~**Still open and untouched by v0.93.0:** R-199, R-200, R-201.~~ **UPDATED 2026-08-04 (evening), after +hub v0.94.0 + agent v0.125.0 + controller v0.195.0:** + +- **R-199 — CLOSED, proven on hardware.** Chain links 6–8 are assembled and walked. The offsite + repository password came back out of the sealed bundle **byte-identical** to the one on disk + (`c60c8bc737a6…` from three independent sources: the box's file, the recovered bundle, and the hash + the hub already stored). +- **R-200 — the plumbing half shipped**, the customer-facing form did not, deliberately. +- **R-201 — PASSED 2026-08-04 (night run).** A customer's file survived a machine rebuild and came + back byte-identical, through the customer's own restore flow. **It passed only because a person was + there:** four manual interventions stood between the recovered key and the restored file, none of + them in any design document → R-204. +- **R-202 — untouched.** The orphan card still promises recoverability unconditionally. +- **The orphaned ciphertext deletion (~1.2 GB) is STILL OWED** — ruling 2, above. + +**UPDATED 2026-08-05, after controller v0.198.0 + hub v0.95.0 (R-204):** + +- **R-204 items 1–3 — CLOSED.** The reset code works without a restart; a re-issue no longer marks a + healthy escrow stale (→ **R-196 CLOSED**); a unit restore states what it did NOT restore. +- **R-204 item 4 — OPEN and unstarted:** a rebuilt box cannot obtain an off-site credential unaided. + It needs an operator ruling on the one-shot credential design → **R-193**. +- **Still open and untouched by this session, stated so nothing is presumed closed by association:** + **R-202** (the orphan card's unconditional promise), **the ~1.2 GB orphaned-ciphertext deletion** + (ruling 2 — still owed, still needs its own scoped session), and **R-198's retention, which remains + UNIT-PROVEN ONLY** — nothing has superseded a key in production, and proving it needs a SECOND + deliberate wipe. That retention drill is the next item, and it is not this session's. + +v0.93.0 made the key survive. v0.94.0/v0.125.0/v0.195.0 make it come back. v0.197.0 got the file into +the snapshot and the drill got it out again. **v0.198.0/v0.95.0 remove three of the four crutches the +drill needed — the fourth is R-193, and until it goes the recovery is still operator-assisted.** + +## R-201 — THE RE-WALK, 2026-08-06 (attended) + +**The question was asked a second time, on the fixed build, on a brand-new appliance built from the +published ISO. The answer is still no — but it is a nearer no.** + +| half | verdict | +|---|---| +| **the data** | **PASS** — all three sentinels byte-identical, including a 12 MB binary and an accented Hungarian filename **whose NAME BYTES are also byte-identical** (verified as hex, not as rendered text). Restored in **23 s** out of the pre-destruction snapshot `a7bc23bd`, through the customer's own two-step full-restore flow | +| **the journey** | **FAIL — two dead ends**, against Phase 1's four. One needed a command line **inside the guest**, one a **Proxmox-host** action | + +**RTO: the unaided figure is STILL UNDEFINED**, because the unaided journey still does not complete. +Attended: login 11:42:22 → key placed **+45 s** → tier up **+24 m 12 s** (after intervention 1) → +all three sentinels restored and verified **+30 m 13 s** (after intervention 2). **The 30 m figure must +not be quoted as the customer number.** The only segment that reflects the product working alone is +**23 seconds** to pull 12.8 MB back once everything was in place. + +**The two dead ends:** **R-218's consume half** (row corrected above) and **R-220** (drives +unenrollable after a rebuild — reproduced and red-proved again; without it no app can be redeployed, +and without a redeployed app the restore page is empty, which is R-213's territory and follows from +R-220 rather than being separate). + +**What PASSED and is worth keeping:** the recovery screen **appeared without being sought** +(`/` → `/launcher` → `/recovery`); it answered all three of its questions and its seal date matched the +hub's `created_at` exactly; the emailed reset code worked **first try**; the unlock took **1.528 s** — +a real unseal — and placed the key; and **R-225's fix was seen working in the wild** (the store read +„a pillanatképek száma még ismeretlen" rather than a false zero). + +**R-216 part 4 reproduced live:** the reinstall **downgraded the agent 0.126.0 → 0.125.0**, back to the +vouched version — an operator's hand-fix undone by the very event that makes recovery necessary. + +⚠ **THE DELIVERY GAP, and it is owed.** A fresh install landed on controller **0.201.0** / agent +**0.125.0** — the vouched versions, **neither carrying the fixes**. They were installed **by hand**. +Fleet delivery needs a golden carrying 0.202.0 **and** a vouched agent 0.126.0. **Nothing was vouched; +that is the operator's act.** **This re-walk proves the JOURNEY on the fixed build; it does NOT prove a +customer would receive that build.** + +Evidence: `tests/rewalk-r201-2026-08-06/journal.md`. + +## CAMPAIGN 11 — the recovery journey, 2026-08-05 + +**The whole journey was walked end to end for the first time, on a throwaway appliance built from the +published ISO. The data came back byte-identical; the journey did not exist.** **Ten findings from +Phases 1 and 3, R-214 … R-223** (seven fixed in controller **v0.201.0** + hub **v0.97.0/0.97.1**; +**three deliberately still open**, each blocking a real flow), **plus five from Phase 2's injected +faults, R-224 … R-228.** Evidence: `tests/campaign11-evidence-2026-08-05/journal.md` (Phases 0/1/3) +and `journal-phase24.md` (Phases 2/4). Campaign document: +`audits/CAMPAIGN-11-recovery-journey-2026-08-05.md`. + +> **Phase 2's verdict in one line.** The cryptography, the retention and the transport all work and +> are now proven live. **What fails is being told the truth:** a mistyped code, a hub outage, a +> stopped agent and a correct code for a retained earlier package all produce **one** message, and +> three of the four are wrong. + +### Phase 2 — the injected faults, 2026-08-05/06 (unattended) + +> **ALL FIVE CLOSED 2026-08-06** in controller **v0.202.0** + agent **v0.126.0**. The rule they now +> enforce, stated so it outlives them: **on the unlock path the customer is blamed only after a real +> attempt REFUSED their code; every other outcome, including an unclassifiable one, says something +> else.** +> +> **STILL OPEN AND DELIBERATELY UNTOUCHED BY THAT WORK — said explicitly rather than left to +> inference:** **R-214** (the console never stops showing a stale pairing code), **R-220** (a rebuilt +> box's drives cannot be re-enrolled — *currently worked around BY HAND on the campaign venue*, which +> is the only reason an app could be deployed there at all), **R-221** (a rebuilt box cannot run the +> escrow ceremony), **R-213** (putting files back), **R-202** (the orphan card's unconditional +> promise — now the **last** place on that surface still promising recoverability, two doors from +> where R-228 removed the same promise). + +Eleven faults, each judged on **the message, not the outcome**, each with a positive control proving +the fault was real. Full observables: `tests/campaign11-evidence-2026-08-05/journal-phase24.md`. + +## Instruction files — deferred half, 2026-08-06 + +**Recorded against existing rows by Phase 2:** + +- **R-216 — §4.1 is now MEASURED, not deduced.** The previous session could only offer two absences. + The box's own `/settings` renders „Minimális verzió (üzemeltető) **0.200.0**" (`GetFloor()`, whose + only writer is the report-ACK handler; **both hold branches serve `Floor=""`**, pinned by + `managed_floor_test.go:94`), and a **cold-started** controller logs + `settle-gate: GO — at/above floor 0.200.0 (we are 0.201.0)` against the same line reading + `floor still unknown after 1m30s` while the hold was in force. The hub's HELD lines ran every + 15 min to 22:12:06 and stopped, with a liveness control proving the hub kept logging. **The floor + is served.** +- **⚠ A correction to how that positive was to be taken.** `SetFloor`'s line is `u.dbg(...)`, gated on + `cfg.Logging.Level == "debug"` and written to the **logger** — it can **never** reach the logx debug + ring, so it cannot appear in `/api/debug/logs` at any level. A controller restart alone would not + have produced it. Confirmed with a level census on the ring first (1196 DEBUG / 2802 INFO / 2 WARN), + so the absence was known to be structural rather than evidential. +- **R-218 — the live half is STILL NOT MEASURED, deliberately.** The fix is present and correct + (`needsOffsiteCredential` now retires on the **target**, not the key), but the venue **has** a + target, so the box correctly does not declare; declaring here would be the bug. The state that + exercises it is shape (a), which the venue no longer holds. **Recorded as not measured rather than + inferred from the unit test.** +- **R-217 — its fix HELD under exactly its fault** (F5): with the store blocked after a successful + unlock, the page rendered M3 and **no listing block at all**; the three false-claim strings are + absent, verified in UTF-8 with accented positive controls present. +- **R-215 — its fix is present** and the `GET /recovery` gate consults the same predicate as the POST + sibling. +- **R-199's back-pointer in `architecture/00-capability-map.md` was already added** — the brief lists + it as owed; it is present on the escrow-recovery row, explicitly labelled as the omitted + back-pointer. **No action taken; the brief's assumption was stale.** + +**Untouched by this session, stated so nothing is presumed closed by association:** **R-213** (putting +files back — the half the recovery screen deliberately does not do) and **R-202** (the orphan card's +unconditional promise, which CAMPAIGN-11 §7 step 7 measured the customer-facing cost of). + +## CAMPAIGN 12 — the class sweep, 2026-08-08 (unattended) + +Eight rows, **grouped by class so the classes are visible as classes**. Full method, controls, blind +spots and the Part-4 gating ranking: `audits/CAMPAIGN-12-class-sweep-2026-08-08.md`. Gating candidates +are in `ROADMAP.md`, not here. **C1 produced no new instance** and has no row, deliberately. + +**⚠ A correction the campaign owed to its own brief:** the task described C5's `escrow_stale` instance +as "closed individually". **It is not — it is R-247, `READY`.** The live repo is the source. + +## CAMPAIGN 12 follow-through — G-1 built, R-260 and R-247 closed, 2026-08-08 + +**The gate was built BEFORE the fixes and was seen failing on 40 fields**, captured verbatim in +`documentation/tests/wire-contract-gate-2026-08-08/BEFORE.md`. That order was the method, not +bureaucracy: the night before, `deadcode` was rejected for class C6 precisely because it was made to +prove itself first and found neither of the two defects it was meant for. + +**⚠ A COUNT THIS SESSION'S OWN PROMPT GOT WRONG, corrected against the repo rather than quoted.** The +prompt said *"465 emitted tags, eight unreachable"*. R-260's wording was "at least eight +DECISION-BEARING facts", never eight tags in total, and its own census already listed more. Measured +on the three declared wires: **210 tags checked, 51 skipped, 40 convicted.** Two prompt claims were +wrong this week and both were caught the same way. + +**Two things the gate's CONTROL caught before it was trusted**, each a defect in the instrument: + +1. **A substring false negative.** `grep -F healed_at` also matches `privsep_healed_at`, so a + genuinely dropped field read as received. R-260 named `healed_at`, so its absence from the output + was the tell. Now a whole-token regex. +2. **`dr_recipe` is not wholly opaque.** The hub stores each half as `json.RawMessage` and re-emits + nested shapes verbatim — but the TOP-LEVEL section keys are decoded by `hostHalfShape` / + `appHalfShape`, and **those are allow-lists**: a section an emitter adds is silently dropped until + named in both. That already cost `offsite_restic` (R-122). The gate is therefore opaque BELOW + depth 1, so the sections are checked; treating the whole subtree as opaque would have put R-122's + shape back outside its reach. + +**THE FORTY, BY DISPOSITION.** Full per-field reasons live in the gate's own `ALLOWLIST`, where each +entry is a claim someone can re-check. + +| # | field(s) | direction | decision | what changed | +|---|---|---|---|---| +| 1 | `oob.operator_key_configured` | agent → hub | **RECEIVE AND ACT ON IT** | decoded as a pointer; `oobDegraded` now fails when the key is absent, and the alert NAMES it. Hub v0.99.0 | +| 2 | `oob.wg_handshake_age_s`, `oob.healed_at` | agent → hub | **RECEIVE, for the message only** | decoded into `HostOOBRow` and put in the event payload; deliberately NOT in the predicate — widening the check beyond the fact that is now arriving is how a check stops being read | +| 3 | `escrow_stale` | hub → controller | **RECEIVE AND ACT ON IT** | `report.EscrowStatus.Stale`; the box tells a withheld hash from a hash-less one. Controller v0.209.0. **This is R-247** | +| 4 | `host.{cpu_temp_c,loadavg,memory_total_bytes,memory_used_bytes,uptime_seconds}`, `system.{load_avg_1,5,15, memory_total_mb, memory_used_mb, temperature_celsius, uptime_seconds}` | agent/controller → hub | **NO CONSUMER WANTED — redundant** | allowlisted: the hub decodes `cpu_percent` / `memory_percent` / `disk_percent` from the same stanzas and every threshold is expressed on those | +| 5 | `guests.spec.{disk_bytes,memory_bytes}` | agent → hub | **NO CONSUMER WANTED — redundant** | allowlisted: guest sizing is hub-owned INTENT (the manifest), not mirrored reality | +| 6 | `storage_targets.smart.model_name` | agent → hub | **NO CONSUMER WANTED — redundant** | allowlisted: a display label with no threshold on it; `smart.health` and every counter the hub bands on ARE decoded | +| 7 | `wireguard.last_handshake_age_s` | agent → hub | **NO CONSUMER WANTED — redundant** | allowlisted: wgsync reconciles from its own state, and the OOB path's own handshake age is now decoded | +| 8 | `guest_net` + its 7 children, `selfupdate_pending(_version)`, `mgmt_plane.healed_recently`, `pbs_dr.applied_at`, `restore_tests.mount_{parity,inventory}`, `config_hash`, `reporting_disabled`, `stacks`, `storage.migrated_to`, `backup.last_db_dump`, `backup.last_integrity_check` | agent/controller → hub | **NO CONSUMER TODAY, AND ONE IS ARGUABLY OWED** | allowlisted **against R-264, which stays OPEN**. Allowlisting is not deciding, and the entries say so | + +**A C7 instance found while doing it, and corrected.** `HostReport.SelfUpdatePending`'s own comment +claimed *"The hub reads an absent field as pending=false, the correct default."* The hub has no field +for it and reads nothing either way. Comment corrected in the agent (no version bump — comment only). + +**The hub's OOB test fixture was part of why this survived.** `oobReport()` omitted +`operator_key_configured` entirely, so every pre-existing scenario ran against a report shape **no +released agent produces**. Fixed, and the new tests drive the raw JSON decode boundary — a test that +builds the receiving struct by hand cannot see a field that never decodes, which is the whole class. + +**Explicitly still open, untouched by this session:** R-246 (the wrong stale flag on `demo-hp` — +clearing it is an operator act hub-side), R-255, R-256, R-257, R-258, R-259, R-261, R-262, R-263, and +**C7's test-comment half**, which Campaign 12 recorded as *owed, not done* (60 of 2652 production +invariant comments sampled; none of the 1440 test comments). + +## The seed that never ran twice, and three pictures that were not true — 2026-08-08 + +Four defects of one family: something the box already knows, either thrown away or drawn as its +opposite. Agent **v0.128.0**, controller **v0.210.0**, `gates.yml` (no hub change, no hub bump). + +**R-221's writer, ESTABLISHED at `file:line` rather than assumed** — the prompt asked for this and it +was owed. `step_agent_config` (`felhom.eu/scripts/felhom-host-install.sh:2396`) renders `agent.json` +from `base = {}` unless an explicit `--preserve-from` is passed (flag `:1246`, defaulting empty at +`:256`), and writes it with `O_TRUNC` (`:2579`). **The render never writes an `escrow` section at +all** — grep over the whole heredoc returns zero hits. The pbsdr marker lives host-side +(`/pbsdr/marker.json`) and survives. So a rebuild keeps the marker and takes the key: +same descriptor, same hash, early return, seed never re-runs. **The attribution in R-221 was +correct.** A rebuild is nonetheless only the case that was measured — the same hole opens for a +hand-edited or restored config, which is the honest reason the fix is at the seam and not in the +installer. + +**The idempotent early return was KEPT**, and that is load-bearing: it stops a converged box +re-running Proxmox operations every 60 s. +`TestSeedReasserted_OnConvergedTick_WithZeroProxmoxCalls` asserts **zero** recorded runner calls on +that tick, so a "fix" that simply deleted the return fails the test. Verified by mutation. + +**The §7.3 truth table as implemented** (R-258): + +| this app's own most recent dump result | restore point | verdict | +| any of its databases failed | yes | `error` — cross | +| all clean | yes | `ok` — tick | +| none recorded (no database / no run yet) | yes | **no icon**, time only, title „Erről a mentésről nincs eredményünk." | +| any | no | no tier-1 row at all, unchanged | + +**An existing test was asserting the defect and was corrected, not deleted.** +`TestBuildAppBackupRows_Tier1FromRestorePoints` expected `"ok"` for a `FullBackupStatus` with **no +`LastDBDump` at all** — a green tick derived from nothing but a file's existence, i.e. Scenario G. +Its real subject, the `Tier1LastRun` time, is unchanged. + +**The convention is now ruled** (§7.2, `CONTEXT.md` S-39): a `…Known bool` companion beside the +figures. `ROADMAP.md` G-3 was blocked on that decision and is unblocked. + +**Six red-proofs, every one demonstrated failing and restored, each with the mutation asserted +applied.** The one that matters: Scenario A **fails against today's tree** with the intended message +— so the test tests the defect. + +**Explicitly still open, untouched by this session:** R-246, R-255, R-256, R-257, R-261, R-262, +R-263, **R-264** (the twenty-one undecided facts — a design session of its own), R-240, R-243, +R-202, R-213, R-244, R-214/R-235, and **C7's test-comment half**, which Campaign 12 recorded as +*owed, not done*. **G-8's other half** (a hub-side check that notices a *vouch* has been forgotten) +was deliberately not built: it is hub work whose payoff is a daily email, and this session already +ends with a bake-and-vouch cycle in front of the operator. + +## Why the TOP READY rows rank this way + +This covers the next few only — it is deliberately **not** a full ordering of the table above, so that +there is one ranking to maintain rather than two. + +1. **R-95** — the largest *data* exposure: the tier holding the customer's documents and photos is the + one whose credential can delete. **THIS PARAGRAPH WAS STALE UNTIL 2026-09-01 AND ITS OLD TEXT IS + NAMED SO THE CORRECTION IS NOT MISTAKEN FOR A RE-RANK.** It said the snapshot mitigation was + *"armed (daily 00:00, keep 7), but it has taken zero snapshots so far"*. **Both halves were wrong:** + seven daily snapshots do exist (R-429), and the word "armed" was withdrawn by the spike that same + day. **What is true now:** the snapshots exist and the box cannot write into them, but **no account + we hold can read anything out of one** (R-433, measured — 777,600 names, zero hits, controlled), so + they do not yet bound this exposure. The root cause is untouched either way — the box can still + `forget --prune` its own repo, from two call sites. **The ORDER of this list is unchanged and is + Viktor's**; only the facts under item 1 were corrected. +2. **R-94** — **de-ranked 2026-07-29.** The prior rationale ("until it moves every hub-driven + install gets the pre-R-82 default") was false: the constant selects no script and every install + already fetches 1.22.0. What remains is a wrong label plus two pieces of dead safety equipment — + a drift gate nobody runs and a test that compares a constant to itself. Cheap and worth doing; + not high-consequence, and it blocks nothing. +3. ~~**R-86**~~ — **CLOSED 2026-08-03**, agent v0.121.0 + hub v0.91.0, proven live on demo-felhom. +4. **R-87** — **SPIKED 2026-08-31 and now a DECISION, not work.** It also spent 2026-08-22..31 in + `CLOSED-ITEMS.md` by mistake while this paragraph ranked it fourth and pointed at nothing + (R-405). The spike measured it rather than designing it: a scratch restore of all 8 apps costs + 25 s against the 40.3 s weekly check, but it would have caught ONE of the five drill-found + restore defects. **Recommendation: build the NARROW version — prove the snapshot still + CONTAINS a recoverable unit — or close the row.** Viktor's call; see + `audits/SPIKE-restic-restore-test-2026-08-31.md`. *The 2026-08-03 rationale, kept:* + R-86 built most of what it was waiting for (per-archive due-ness, a + proof that names its archive, and a staleness window that learns a tier's rhythm). What is left is + restic-specific — there is no scratch-guest analogue — so it still needs its own design, but it is + no longer waiting on a scheduling model that did not exist. +5. ~~**R-185**~~ — **CLOSED 2026-08-03**, agent v0.123.0 + installer 1.24.0, proven live on both demo + boxes. The silence was fixed as well as the grant: the box now asks whether it may READ each tier + it depends on, because an empty listing cannot distinguish forbidden from newborn. +6. ~~**R-189**~~ — **CLOSED 2026-08-03** with **R-188** and **R-186**, agent v0.122.0. The three + reporting/release signals that misreported their own work are fixed; **R-185 is the one that + remains open from that group** and is untouched by this — it is a missing storage ACL on + demo-felhom, not a reporting defect. +7. **R-110** — last **because it is not a READY row**: the ruling is the operator's, not CC's, and + there is nothing for CC to build until it lands. Ranked here rather than omitted because it is the + only item on this page about the *publish channel* of the most privileged artifact Felhom ships, + and today's exposure is zero — which makes now the cheapest moment it will ever be to decide. + +### The 2026-08-02 intake (R-156 … R-164), ranked + +Filed in one pass from Campaign 10, its two spikes, and the 53-template catalog persistence sweep. +**R-156 and R-157 had lived only in audit documents** — the identical *"minted in a spike doc and +never carried across"* failure the register already records for R-153/R-154/R-155, caught by the +sweep's own §8.0 while it was happening. R-158 was minted by a second session on the same day for an +unrelated finding, which is why the sweep's proposals were renumbered to R-159…R-162 at filing time. + +1. **R-157** — highest: a `deployed: true` app can stay **down indefinitely** after a power cut or + hard reset, and in mechanism **B** nothing reports it on any channel (`0 currently down`). It is the + only row here where the customer loses service and has no signal at all. +2. **R-156** — the class is now detectable and two of three apps are fixed; what remains is **papra's + referral**, one app, well understood. *(Promoted 2026-08-02: R-161 was ranked here because nothing + ran the gate; it now has a mandated entry point, so R-156's residue is the larger remaining item.)* +4. **R-163** — a real ceiling that silently caps local backup once an app outgrows `mp1`, and it gates + Tier-2 and Tier-3 as well. Ranked below the above only because **overflow itself is safe today** — + it refuses per app and preserves the last good unit byte-identical. **RE-FRAMED 2026-08-02:** no + longer waiting on a ratio — decision **D-a** merges `mp1` away, so the row is now the record of the + constraint and the work moves to **R-165** (with **R-167** shipping in the same step). R-165 inherits + this rank; it is the highest-ranked item that must land **before any external install**. +5. **R-158** — the gap that makes R-163 dangerous: cross the size line and **one page** tells you. On + its own it is a notification gap, not a silent failure, which is why it sits here and not higher. +6. **R-164** — blocked on a predicate, no customer impact today; it only becomes urgent if the unit + size in R-163 is judged unacceptable, since the tar-drop is the cheapest way to halve it. +7. **R-161** — **de-ranked 2026-08-02, ruled and shipped at reduced scope.** The gate now has one + mandated entry point (`catalog_gates.py`), which is the shape that actually gets run here. What is + left is the automatic half, and that is sufficient while **one** person touches templates — so it + ranks low by design, not by neglect. Revisit when a second does. +8. **R-162** — `WATCHING` only. A limitation that fails closed; revisit if a non-overlay driver ships. + +**R-159 and R-160 are SHIPPED** and are not ranked; they are filed to record the class, and R-159's +class (an image `VOLUME` at an unmounted path) is still live — `immich-server` has one today. + + + + diff --git a/documentation/audits/DRILL-golden-098-2026-07-03.md b/documentation/audits/DRILL-golden-098-2026-07-03.md index 0a24666d..4a08fd44 100644 --- a/documentation/audits/DRILL-golden-098-2026-07-03.md +++ b/documentation/audits/DRILL-golden-098-2026-07-03.md @@ -5,7 +5,7 @@ Source findings: `DRILL-day0-cleanroom-2026-07-03.md` §9 **B5** (golden bakes a pre-floor controller → mandatory manual D.1b on every fresh install) and **B1** (controller-bootstrap only fires at boot; the hot-plugged config mount needed the installer's reboot crutch). Also closes the stale backlog note -`documentation/backlog/FOLLOWUP-golden-default-controller-tag.md`. +`documentation/archive/FOLLOWUP-golden-default-controller-tag.md` (moved from `backlog/` 2026-10-03). **Verdict (short):** both findings are FIXED in the product. A fresh Day-0 install now lands controller **0.98.3** on first boot and self-manages from there (no D.1b), and the baked diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 6fcae516..c8531400 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -26,6 +26,166 @@ --- +## 2026-10-03 — the triage: finished rows moved out of the open register + +> Every row below sat in `OPEN-ITEMS.md` with a finished LEADING verdict (or was verified finished against +> live source on this day, or folded into an older duplicate). Full original text of each: +> `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`. Since this day `closed_register_gate.py` RULE 3 +> refuses a finished row left in the open register. + +| Row | What | Closed | Full text | +|---|---|---|---| +| **R-88a** | **~~Failing backup re-quiesces every 5 min, no backoff~~** **Reasoning kept:** «breaker 15m→4h, per-tier, never permanent» | SHIPPED (controller v0.176.0, 2026-07-27). Live on both boxes; breaker 15m→4h, per-tier, never permanent. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-88b** | **~~`/backup/due` cannot say *unknown*~~** | SHIPPED + PROVEN-LIVE (agent v0.105.0 + controller v0.178.0, 2026-07-27). `age_state=unknown` captured on real hardware during a deliberate ep0 outage; controller deferred, zero app stacks stopped. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-123** | **R-105 and R-106 were `READY` in `ROADMAP.md` with no row on THIS page** | CLOSED 2026-10-03 (triage, verified) — its last residue — "R-105 still needs a row" — is done: R-105 has its own row; the process gap is R-369's gate | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-131** | **`sess-f` is a fourth orphaned scratch customer** | CLOSED 2026-10-03 (triage, verified) — `sess-f` was already gone by 2026-08-05 (`documentation/tests/campaign11-evidence-2026-08-05/journal.md:24-27`) | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-176** | **Two prerequisites for the R-165 merge are UNMEASURED, and both are cheap.** | (a) ANSWERED 2026-08-03 (P1: PASS). (b) NOT REQUIRED — operator ruling: every node is reinstalled, none migrated (withdrawn, not deferred; CONTEXT.md:2076). — residue NOT-A-FINDING (2026-10-03): its residue ("re-run (a) once against agent v0.120.0") is moot: every node was reinstalled, none migrated (`CONTEXT.md:2076`) | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/SPIKE-r165-phase0-2026-08-03.md | +| **R-191** | **Every weekly offsite backup UPLOADS successfully and then FAILS the job on a prune the box is deliberately not allowed to do — on both demo boxes.** **Reasoning kept:** «R-89 moved PBS pruning SERVER-SIDE — *"boxes set `keep_last: 0`, ep0 runs prune jobs; box tokens stay write-only, never widen the grant"*.» | CLOSED — SHIPPED 2026-08-04 (installer 1.25.0; both live boxes corrected to offsite `keep_last: 0`, `prune_pbs_allowed=false`). ep0 prune jobs verified first (18 tasks OK). hostinstall_gates.py asserts it (red-proved). The 'next weekly run OK' observation: offsite snapshots dated 2026-08-11 recorded in audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md:118. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md | +| **R-201** | **Nothing in the offsite DR chain has ever been exercised past the ceremony — and the one live proof that exists predates the field it is cited for.** **Reasoning kept:** «**The pass condition is unchanged** — a byte-identical sentinel sha256, not "the store opened"» «**a good snapshot is not durable against a later bad run on the same day.**» | CLOSED 2026-08-07 — BOTH HALVES PASS on the fifth walk (DATA 2026-08-04 night run; JOURNEY 2026-08-07, zero guest command lines). Does NOT claim a smooth journey (R-252/R-253, both since CLOSED 2026-08-08) nor shape (c)'s positive half. Walls along the way: R-203, R-204, R-241. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/DRILL-r201-night-run-2026-08-04.md, audits/RECON-offsite-dr-chain-2026-08-04.md, tests/finalwalk-r201-2026-08-07/journal.md, tests/walk5-r201-2026-08-07/journal.md | +| **R-202** | **NOW EVIDENCED, NOT ARGUED (2026-08-10).** | CLOSED 2026-10-03 (triage, verified) — the false recoverability promise was removed (R-294/R-299; `controller/internal/web/templates/backups_remote.html:93-108`) | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-209** | **~~Should the containerd store move to SSD2 at all?~~** **Reasoning kept:** «**A TRAP was found while proving the guard, and it is the reusable part: `RequiresMountsFor` on a path with NO mount unit is a SILENT NO-OP**» «**`-X` is load-bearing** — overlayfs stacking rides `trusted.overlay.*`.» | EXECUTED 2026-08-05 on operator ruling — reboot validation DEFERRED → R-209a (still open, WATCHING, OPEN-ITEMS.md:339). Zero-loss move verified on four observables; storageReserved on SSD2 0 → 80 GB; pre-move tree moved aside. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/SPIKE-dooplex-buildcache-2026-08-05.md | +| **R-214** | **The physical console never stops asking to be paired.** | CLOSED 2026-09-20 - ISO 1.28.0+/R-535, proven on a fresh install (localisation slice 6's walk on ISO 1.29.0: last console paint after self-bind is the bilingual bound banner). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/i18n-slice6-2026-09-20/drill/screens/22-console-after-claim.png | +| **R-229** | **The instruction-file rightsizing landed for `felhom-controller` and the workspace root; three pieces were deliberately deferred.** | CLOSED 2026-10-03 (triage, verified) — legs (a), (b), (c) CLOSED 2026-08-06 in the row's own text; leg (d) moved to R-230 | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-233** | **The golden bake's acceptance checks were a list of strings the script does not print.** **Reasoning kept:** «**The general lesson:** a runbook's pass markers must be **copied from a captured log, never written from memory**» | CLOSED 2026-08-06 — fixed in `RUNBOOK-manual-build.md` (§4.1): markers re-captured from the real log, pre-gate URL corrected to `golden.tar.zst`, token moved off the command line into an in-VM runner script, positive control on the token-leak grep, vouch step rewritten. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: runbooks/RUNBOOK-manual-build.md | +| **R-245** | **Should a customer who never decides be auto-abandoned after 30 days? RECORDED, NOT BUILT.** **Reasoning kept:** «**The real harm, if it comes, is QUOTA** — old history blocking new backups — and **that is a condition, not a calendar**. An automatic ending should trigger on the harm, with a dated warning, never on a date alone. **If this is ever built, build it that way.**» | DECIDED 2026-08-07 — not built; reopens on quota. Built instead: escalating reminders (1/3/7/14 days) and operator levers `--abandon-extend` / `--abandon-stop`. Re-filed 2026-08-08 as a decision taken. Reopens verbatim: "THE CONDITION THAT REOPENS IT, which the reasoning already names: QUOTA — old set-aside history blocking new backups." | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-248** | **A flag that changes behaviour is visible to nobody who would look for it.** | FOLDED into R-246 2026-10-03 — the same ruling: give `stale_at` a visible, evidence-bearing setter, or retire it | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-254** | **The same render-then-hide pattern R-249 fixed is live in two more places, and one of them carries a real per-install secret.** **Reasoning kept:** «`escrow_handlers.go`'s rule that a secret is revealed by an XHR and never templated server-side into HTML.» «**A form must carry what it submits.**» | CLOSED 2026-08-08 — both sites closed in controller v0.208.0: `POST /apps//initial-credentials/reveal` (re-reads container, no-store, CSRF, logged) and `POST /stacks//auto-field/reveal` for deployed pages; pre-deploy hidden input kept deliberately. Gate `scripts/secret_in_markup_gate.py`; its blind spot → R-255. Measured exposure: none; rotation not indicated. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: scripts/secret_in_markup_gate.py | +| **R-272** | **RANK 1 — Felhom's `--uninstall` leaves the exact condition that makes Felhom's own reinstall REFUSE.** | CLOSED 2026-10-03 (triage, verified) — the uninstall purges the dnsmasq Felhom installed (R-316; `scripts/felhom-host-install.sh:858-882`, `_dnsmasq_purge_owned`) | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-295** | **One name per secret — CONTROLLER HALF SHIPPED.** | CLOSED 2026-10-03 (triage, verified) — hub half shipped in hub v0.104.0 (`4d6ec7c`, hub `CHANGELOG.md:844,905`); the controller half shipped v0.211.0 | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-303** | **`markOrphaned` has no guard against an active abandon countdown — the co-render is made HARMLESS, not IMPOSSIBLE.** **Reasoning kept:** «the wrong fix (suppressing the orphan card during a countdown) would hide a real second fault.» | DECIDED 2026-08-13 — left as it is; reopens on a real-world sighting. Verbatim trigger: "an observation of the combined state occurring OUTSIDE a constructed test." Related: R-302. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-312** | **There is no in-product route from the recovery screen to a set-aside store, and building one is not wiring — it is new surface.** **Reasoning kept:** «retention is an **operator-only** capability, and nothing anywhere may promise the customer can perform it themselves» | DECIDED 2026-08-13 — deliberately not built; re-evaluate on a real customer request. Verbatim trigger: "a real need appearing — one request from a customer who is not us." Related: R-304, R-311. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-313** | **`demo-felhom`'s set-aside store is UNRECOVERABLE — 36 snapshots whose key we destroyed ourselves.** | DECIDED 2026-08-13 — kept as a test fixture; delete when R-312 ships or is abandoned. Verbatim: "when R-312 is built, or when R-312's ruling above is made permanent. On either event, delete it deliberately and record why". Related: R-198, R-307, R-312. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-341** | **Does the fd slope change after the PBS 4.2.5 upgrade? — two dated checks, and the answer is expected to be NO.** **Reasoning kept:** «the anchor must be `ps -o lstart= -p $MainPID`, NOT `systemctl show -p ActiveEnterTimestamp`, which reads 03:54:54Z for this generation (the upgrade re-exec'd the proxy; systemd never saw a stop, `NRestarts` is still 0) and would put the rate ~15% low.» | CLOSED 2026-08-30 — second check taken (fd back to baseline 17, ESTAB 0) but it cannot answer the question: the leak was REMOVED mid-interval by our own fix (R-344, agent 0.130.0). Question moot; fix confirmed holding at 12 days. First check 2026-08-20: unchanged, 201.6 fd/day. R-336 keeps the request-rate scaling item; ActiveEnterTimestamp trap → R-346. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/SPIKE-ep0-established-connections-2026-08-20.md, audits/evidence-r341-plus7d-2026-08-30/step1-fd-and-sockets.txt, evidence-ep0-pbs-upgrade-2026-08-18/stop1-ruling.txt | +| **R-343** | **The managed controller floor was raised 0.214.0 → 0.216.0 — and it was NOT the no-op it was expected to be: it moved a live box nine seconds later.** | CLOSED 2026-10-03 (triage, verified) — obsolete — its only step was "confirm 0.216.0 is healthy in normal operation"; both demo boxes have run every release since and are on 0.288.0 (`STATUS.md`, 2026-10-02) | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-369** | **There are TWO registers, only one calls itself the source of truth, and work filed in the other is invisible to every standing rule that says "grep the register".** | CLOSED 2026-10-03 (triage, verified) — RULED 2026-08-22 (one register) and enforced by `scripts/one_register_gate.py` (`ef6ac6f`) | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-378** | **A status word inside a longer verdict fooled the session's own compressor, moving six still-open rows into the closed file.** **Reasoning kept:** «Match the LEADING verdict.» | CLOSED 2026-08-22 — corrected in the same session; the six rows (R-123, R-190, R-214, R-264, R-295, R-352) restored verbatim from commit `fddfe00ce268`. Successor: R-369. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-385** | **A controller was built, baked AND vouched with no CHANGELOG entry of its own, and every gate stayed green.** **Reasoning kept:** «**Membership, not `baked > released`, deliberately:** a comparison against the newest heading alone goes green the moment any later entry is written, leaving the unrecorded version permanently unrecorded and the gate permanently silent about it.» | CLOSED — 2026-08-23. v0.221.1 given its own CHANGELOG heading (commit da75603); scripts/golden_currency_gate.py now requires the baked version's own `## vX.Y.Z` heading anywhere in the CHANGELOG (membership); INCONCLUSIVE exit 2 preserved; both directions red-proofed. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/DRILL-r384-dead-db-alarm-2026-08-23/evidence/gate-0*.txt | +| **R-387** | **The hub REWRITES an unknown severity and says nothing, and the guard built to catch that sits downstream of the rewrite.** **Reasoning kept:** «**The coercion STAYS; only the silence is fixed** — a rejected event is a LOST event, and losing an alarm is worse than mis-routing one.» | CLOSED — hub v0.107.0, 2026-08-23. Coercion kept; a WARN now names customer, event type, rejected value and consequence. Dispatcher severity branch kept (it is the only guard for monitor-originated events). Proven live with an `error` control silent. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/DRILL-r329-r386-2026-08-23/evidence/live-19-scenarioH-after.txt | +| **R-398** | **`resticStep` is not a seam, so no test can drive any restic-backed path.** | CLOSED 2026-10-03 (triage, verified) — premise corrected 2026-08-30; the execution test exists (`controller/internal/backup/r358_scratch_marker_test.go:200-205`, `TestR358_MarkerOrderingIsExecuted`) and the `resticStepFn` seam was deliberately NOT built — it would hide the `unlock --remove-all` escalation the test asserts | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-405** | **R-87 sat in CLOSED-ITEMS.md for nine days while still open; the ranking paragraph ranked it fourth pointing at nothing.** | CLOSED 2026-08-31 — corrected + gated in the same session: R-87 restored verbatim from `ef6ac6f^`; `scripts/closed_register_gate.py` is the 12th gate; R-398 stub turned to prose. Gate's residual holes in its docstring; hole 4 is R-406. Predecessor R-378. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-413** | **R-87's proof caught a naturally-produced hollow snapshot, end to end, unattended.** | CLOSED 2026-08-31 — the claim it upgrades is recorded (opengist `volumes_expected_none_captured`, one `offsite_proof_empty` at error; four apps ahead passed). Related R-87, R-412. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-428** | **The decoy-coverage gate identified a repository by its DIRECTORY NAME.** | CLOSED 2026-09-01 — fixed, and kept as the class's best example (felhom.eu CI job 490 found it; the gate now identifies a repo by which registered runner FILE exists under the root; verified under a renamed directory). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-432** | **A customer's own sub-account can REACH the snapshot door and is REFUSED writes — but sees it EMPTY, so per-file recovery is not product-reachable.** | ANSWERED 2026-09-01 — negatively; the panel-read next step is WITHDRAWN as unnecessary (777,600 exact snapshot names, zero hits; /.zfs/snapshot is a different filesystem from /home). Per-file recovery unreachable from the box entirely → R-433. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/evidence-drill-r95-recovery-2026-09-01/ | +| **R-434** | **The snapshot-drop alarm promises a recovery that cannot be performed.** **Reasoning kept:** «when a verdict changes which fact it counts from, the alarm text has to change with it, or the operator acts on a promise nobody can keep.» «**That sentence is true under EVERY possible answer to the provider questions, so it never needs a second rewrite** — which is the whole reason it was not blocked.» | CLOSED 2026-09-01 — hub v0.111.1; the promise was DELETED (not replaced) with a sentence true under every provider answer; three tests in `hub/internal/monitor/offsite_r434_test.go`, red-proofed. The "blocked on R-433" verdict above was mine and it was wrong. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-463** | **The day the catalog moves postgres:16 to 17, ELEVEN apps are affected and the container image will NOT perform the conversion. (P2)** **Reasoning kept:** «A later move of any of the three needs its own two-venue proof: the engine gate enforces it per app.» «PostgreSQL majors are converted BY THE BOX as a guarded-update step — save everything from the old engine, start the new one empty, load it back, check.» | CLOSED 2026-09-30 — 8 of 11 moved by the box's own conversion (docmost, paperless-ngx, tandoor, claper, calcom, rallly, outline, sparkyfitness; controller v0.273.0+, `09` decisions 16/42/43); 3 stay by decision 42's rule (zipline, adventurelog, immich). A later move of any of the three needs its own two-venue proof: the engine gate enforces it per app. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md, audits/update-night-2026-09-21/24-Q5-postgres-conversion-costed.md, audits/pg-calcom-claper-2026-09-28/, audits/pg-last-six-2026-09-30/ | +| **R-497** | **No product channel ever gives the customer the „Tulajdonosi jelmondat”, yet the self-bind mail says they received it at setup. (P2)** | CLOSED — hub v0.113.0 (2026-09-14, ArgoCD Synced): customer page renders the hand-over sentence; tests red first (`passphrase_handover_test.go`); mail change pinned by test, not read live. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/DOORSTEP-walk-1270-2026-09-14.md | +| **R-500** | **[P3-LOW] The dashboard shows the last backup in UTC while every backup page shows it in local time — two different clock times for one backup.** | CLOSED 2026-10-03 (triage, verified) — controller v0.283.0 — the dashboard's last-backup time uses `fmtTime`, local time (`dashboard.html:154`; controller `CHANGELOG.md:222`) | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-505** | **A fresh box on customer tester-1 connects its Cloudflare tunnel but receives NO routes, so the dashboard answers 503. (P1)** | CLOSED 2026-09-29 — proven end to end on a fresh box. Cause: the tester-1 tunnel had no public hostnames; operator added a route 2026-09-14; box-side hop proven by the 2026-09-29 new-household drill. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-506** | **[P3-LOW] `day0-install.md` A.1 says "the controller manages per-app hostnames itself via the tunnel" — it does not.** | CLOSED 2026-10-03 (triage, verified) — the day-0 runbook names the published-route step and says the controller creates no routes or DNS (`documentation/runbooks/day0-install.md:47-53`) | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-511** | **A rebuilt customer's box keeps its ep0 PBS token, and the DR tier can be neither provisioned nor re-issued. (P2)** | CLOSED 2026-09-16 — proven live on a fresh box: hub v0.114.0 re-issue ADOPTS the endpoint token (shipped 2026-09-15); made live by R-534's ep0 grant. Token-only release on host delete not built → R-526. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/evidence-p1fixes-2026-09-15/ | +| **R-520** | **A power cut during a guarded Update leaves no record that an update was running. (P3)** **Reasoning kept:** «**`pct mount` maps the guest's ROOTFS ONLY and does not apply the guest's own internal mounts**, and 9202 keeps `/var/lib/docker` on its own ext4 mount under the `mp0` volume, so from the host that path is an EMPTY STUB.» «**An empty directory is not evidence of an absent file.**» | CLOSED 2026-09-21 — measured on a real version change (scratch 9202, controller v0.260.0, uptime-kuma 2.4.0→2.5.0 cut in `pulling`; box put the pin back itself). The `starting`-phase cut → R-610. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/update-arc-2026-09-21/ | +| **R-524** | **When the catalog moves an app back to an older version, a box that already updated shows „Frissítés elérhető” — and the offered Update is a downgrade. (P2)** **Reasoning kept:** «**Ahead is NARROW on purpose:** every differing service must be orderable AND newer, or the verdict falls back to Behind — this gate can BLOCK an update, so it errs towards letting one run.» | CLOSED 2026-09-21 — controller v0.260.0: `stacks.CatalogOrder` four verdicts incl. Ahead; `UpdatePreflight` refuses `downgrade` (409); three red-proofs; `09` §3 decision 10 (decided by CC unattended — operator may reverse). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-534** | **The off-site tier cannot be provisioned or adopted for a rebuilt box: the hub's endpoint token lacks Datastore.Modify. (P1)** | CLOSED 2026-09-16 — grant given and proven end to end: DatastoreAdmin for the hub's `felhom@pbs` on `/datastore/felhom-offsite` only (narrowest role carrying Modify; PBS has no custom roles); re-issue then ADOPTED on a fresh box. Closes R-511. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/evidence-drill-0243-2026-09-16/phase1-pbsdr.txt, audits/evidence-backup-promise-2026-09-16/phaseC-ep0-grant.txt, audits/evidence-backup-promise-2026-09-16/phaseC-reissue.txt | +| **R-535** | **The box's console still says it waits for pairing long after the box is bound — and promises the screen refreshes itself. (P2)** | CLOSED 2026-09-16 — shipped in ISO 1.28.0 (published); on-screen effect not photographed (`print_bound_banner`). Later SEEN on screen by R-214's closure 2026-09-20. Deliberately does not name the dashboard URL nor reflect the later claim. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/evidence-drill-0243-2026-09-16/screens/33-console-after-claim.png | +| **R-536** | **The hub is told „Alkalmazás telepítve” the moment a deploy is ACCEPTED, so an install that never finishes is recorded as completed. (P2)** **Reasoning kept:** «The accept-time `app.yaml` is deliberately NOT deleted on failure — it is the crash-safe record with `Deployed:false` and it holds the settings the customer typed; the state every surface reads is `not_deployed`.» | CLOSED 2026-09-16 — controller v0.244.0 + hub v0.116.0: `app_deploy_started` at accept, `app_deployed` from the async end, `app_deploy_failed` (warning); both new types in `allowedEventTypes` and `customerMessages`; two red-proofs. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-537** | **The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and shows the data-drive size, but the tier-1 unit contains NO drive-side app data (P1-HIGH)** **Reasoning kept:** «The contents label is computed PER TIER from what that tier captures: Tier 1 says „Adatok" only when the app's data really is in the volumes the unit captured» | CLOSED 2026-09-16 — controller v0.244.0, proven live on demo-hp; label computed per tier; re-proven on a fresh box (ISO 1.28.0). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/evidence-drill-0243-2026-09-16/phase2-f10.txt | +| **R-538** | **A tier-1 app restore reports plain success, leaves Nextcloud listing files whose bytes were never backed up, and destroys the app's own trash (P1-HIGH)** **Reasoning kept:** «never present a DB-only restore of a class-A app as a complete one.» «A unit restore refuses before anything is touched when the unit cannot return the app's drive-side files, and names the route that can.» | CLOSED 2026-09-16 — controller v0.244.0, proven live on demo-hp; unit restore refuses when the unit cannot return the drive-side files; re-proven on a fresh box, off-site route returned 5/5 photos sha256-identical. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/evidence-drill-0243-2026-09-16/phase2-f10.txt, audits/evidence-backup-promise-2026-09-16/phaseE-photos.txt | +| **R-543** | **Off-site ON by default is not off-site WORKING: a fresh box's tier 3 waits on escrow and nothing asks the household. (P1)** **Reasoning kept:** «The pause is untouched: it is the zero-knowledge escrow design, and this row was never about the mechanism.» | CLOSED 2026-09-16 — shipped in controller v0.245.0 and proven live: escrow-pending bar on every authenticated page (via `executeTemplate`) + tier-1 sentence rendered by `tier3State`; both red-proofed; live on 9202 (paused) and 9201 (escrowed). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/evidence-recovery-code-2026-09-16/ | +| **R-557** | **Localisation slice 2 — Go-side customer strings follow the language. (P3)** **Reasoning kept:** «Plurals are a BUNDLE rule (a key with `.one`/`.other` takes its count first), not a call-site flag.» «EXCEPT one producer: `"Sikeres — nincs mentésre jelölt alkalmazás"` (controller/internal/backup/offbox.go) must stay Hungarian until **R-570** closes, because the page's legacy fallback still reads it on boxes that have not run off-site since 0.251.0.» | CLOSED 2026-09-18 - controller v0.252.0 + v0.253.0 + v0.254.0 (release A: 226 literals + `i18n_go_parity.py`; B: all 179 error literals keyed via `util.MsgError`; C: saved notes in box language, globe switch). Gaps filed: R-570, R-572..R-578; R-566 closed with it. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-558** | **Localisation slice 3 — the hub's customer e-mails follow the household's language. (P3)** **Reasoning kept:** «The bind page is per-language with `expired` pinned to Hungarian (the language would otherwise be the oracle the text refuses to be).» | CLOSED 2026-09-18 - hub v0.118.1 + controller v0.256.1 (hub v0.118.0/v0.118.1 + controller v0.256.0/v0.256.1; 56 Hungarian mail goldens unchanged; proven live twice). Gaps filed: R-581..R-585; R-555 closed with it. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-559** | **Localisation slice 4 — the console banner and the download page in English. (P3)** | CLOSED 2026-09-18 — ISO 1.29.0 published (sha256 `dceacae5…e94829`), proof installs on both menu entries + reboot; `felhom.eu/en/download` live; release gate G16 rewritten per ruling 1b. Gaps filed: R-586, R-587, R-588. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-560** | **Localisation slice 5 — catalog cards, settings and first steps in English. (P3)** **Reasoning kept:** «lists replace WHOLE and every other list is matched by its own key (`env_var`, option `value`, `match_group`, `target`, `path`), never by position;» «**The blocks are GENERATED from a flat `{path: english}` map**, not hand-written: a mistyped `env_var` is INERT on the box rather than an error, and the generator can only write paths that exist on the Hungarian side.» | CLOSED 2026-09-20 — controller v0.257.0, catalog fully translated (catalog `e81d41e` + three batches; 1 031/1 032 strings, ceiling 1), fleet floor 0.257.0 (min_agent 0.131.0). Gaps filed: R-589..R-594. — residue NOT-A-FINDING (2026-10-03): its residue (Peti's box and tester-1 taking floor 0.257.0 unattended) is moot: Peti's box was retired 2026-09-25 and tester-1 takes every floor like any box | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-561** | **Localisation slice 6 — the volunteer guide in English, then a stranger's first hour in English; closes R-516. (P3)** | CLOSED 2026-09-20 - controller v0.258.0, the English guide, and the walk (`runbooks/VOLUNTEER-first-hour.en.md`; R-589/R-590/R-573 fixed; R-214 closed as a side effect). Verdict: not yet ready for an English tester (R-596). Gaps filed: R-596..R-599. R-516 does NOT close. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/DRILL-first-hour-en-0258-2026-09-20.md, audits/i18n-slice6-2026-09-20/R-516-item-by-item.md, runbooks/VOLUNTEER-first-hour.en.md | +| **R-566** | **Three page titles built in Go around an app name stay Hungarian in the English browser tab. (P3)** | CLOSED 2026-09-18 - controller v0.252.0 (four `%s` keys via `data["TitleArgs"]`; pinned by `TestParameterisedPageTitles` and `i18n_go_parity.py`). Part of R-557. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-572** | **[P3-LOW] Two copy-producing template helpers have no English form, so an English page renders „vasárnap" and „%d órája" in Hungarian.** | CLOSED 2026-10-03 (triage, verified) — controller v0.258.0 — `pruneLabel`/`nextPruneLabel` deleted as dead code (`controller/internal/web/funcmap.go:368-377`) | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-580** | **[P3-LOW] `curl -w '%{redirect_url}'` prints Basic-auth credentials back into the session transcript.** | FOLDED into R-132 2026-10-03 — the same `curl -w %{redirect_url}` credential echo, seen again 2026-09-18 | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-582** | **An English copy guard written from the Hungarian one matches ordinary words instead of the claim. (P3)** **Reasoning kept:** «**The general form worth keeping: a guard ported between languages must be re-derived from what the claim IS in the new language, not translated word for word — and a guard that convicts 141 true sentences is worse than no guard, because it earns an allowlist entry per sentence and then nobody reads it.**» | CLOSED 2026-09-18 - hub v0.118.0 (recorded for the lesson; three decoys incl. an innocent control): patterns carry the modal and admit an adverb; gate scans bundle VALUES only. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-583** | **The test-notification mail was the one customer mail that did not follow the language. (P3)** **Reasoning kept:** «**The general form worth keeping: the surface you would use to CHECK a feature is the one most worth checking first — a broken instrument that reports success is worse than a broken feature.**» | CLOSED 2026-09-18 - hub v0.118.1 (`mail.test.subject`/`mail.test.body`, red-proofed, proven live in both languages 74 s apart on demo-hp). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-589** | **The update badge is Hungarian on an English app page — and it is the badge, not a corner case. (P3)** **Reasoning kept:** «**The lesson, because it cost a reviewer pass and half a task brief: a reviewer who reads ONE producer cannot see a SECOND producer that overrides it** — reading `updatebadge.go` alone gives exactly the wrong answer.» | CLOSED 2026-09-21 — shipped in v0.258.0; the row was stale (English built in `web.localeFuncs` "updateBadge", pinned by `TestUpdateBadgeFollowsTheLanguage`; verified at controller `19ef0329ab66`). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/DRILL-first-hour-en-0258-2026-09-20.md | +| **R-590** | **[P3-LOW] The data-folder card tells an English household, in Hungarian, whether its files are backed up.** | CLOSED 2026-10-03 (triage, verified) — controller v0.258.0 — `consequenceFor` renders through message keys in the household's language (`controller/internal/web/datapath_card.go:32-50`) | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-592** | **Two defects in the new catalog copy gate, each found by its own decoy rather than by reading it. (P3)** | CLOSED 2026-09-20 — fixed in the same session, decoys added (`scripts/test_gate_decoys.py`, 33 cases): coverage scoped to named apps, credential and ASCII-stem bare-substring matches. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-595** | **The catalog's new copy gate could not RUN in CI at all — six pushes red, six alarm mails, while the local hook was green. (P2)** **Reasoning kept:** «**The general form, and the reason this is P2 rather than P3:** a new gate is written and tested on the machine that has every library, and the runner deliberately has none.» | CLOSED 2026-09-20 — degraded mode; CI job 791 green (catalog `18a6d2d`: without PyYAML the gate runs the freeze check and prints what it did not check; five decoys with PyYAML shadowed). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-596** | **The claim page — the FIRST screen an English household touches — is English chrome with HUNGARIAN messages. (P1)** **Reasoning kept:** «It was **deleted, not translated**: a translated dead field would have read for ever after as evidence that this page's title is decided in the handler.» | CLOSED 2026-09-21 — controller v0.259.0, proven live (fourteen sites via `s.msg(r, "claim.msg.*")`; dead `data["Title"]` deleted; language chain pinned by a test; live on guest 9201 both languages; red-proofed). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-597** | **The setup code is three Hungarian words inside an otherwise fully English e-mail to an English household. (P2)** **Reasoning kept:** «**The task's proposed "read it over the phone" filter was MEASURED and NOT adopted** — it removes 5270 of 7772 words (68%, 12.92 → 11.29 bits/word) and would make this list stricter than the one the product already uses for the code a household writes on paper during a disaster;» «The hub does not own that secret and no row was opened for it: a second definition here is the drift `backupTargetAbsentText` already demonstrates across two repos.» | CLOSED 2026-09-21 — hub v0.119.0: `RandomPassphraseFor(lang, use)`; setup code 4 en words (51.7 bits), owner passphrase 6 en (77.5); entropy floor computed from embedded lists, red-proofed. Recovery code was never Hungarian (agent, EFF list). Decision in source, operator may reverse. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-598** | **The Backup page's two protection warnings are Hungarian on an English dashboard. (P2)** **Reasoning kept:** «**The English is asserted to carry the same NEGATION the Hungarian does** ("protects against corrupted files, **but not** against a disk failure"); an English sentence that promised disk-failure protection would be worse than leaving it Hungarian.» | CLOSED 2026-09-21 — controller v0.259.0; the two warnings proven by render test, not live (tier names proven live on 9201; `degradedMessageFor` returns a KEY). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-601** | **demo-hp is unreachable — WRONG, WITHDRAWN THE SAME DAY. The box was never down; MY ROUTES WERE. (P2)** **Reasoning kept:** «**The standing rule says a "no access" claim must list what was tried; it does not say the list makes the claim true.**» «Six failed routes to a stale address are six failures of one assumption, not six pieces of evidence.» | CLOSED 2026-09-21 — withdrawn, the claim was false; the routes are fixed (both `~/.ssh/config` entries repointed to 192.168.0.104; `nodes.md` corrected; tailscale not installed on demo-hp). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-608** | **The controller swaps ITSELF in the middle of a guarded app update, and 04:30 sits inside the proposed auto-update window. (P2)** **Reasoning kept:** «**THE PROPERTY THAT MATTERS MOST IS THAT THE LOCK DOES NOT LATCH:** `Stack.Updating` is cleared on done, failed AND held, so a HELD app does not block the controller's own updates — including the release that might fix whatever held it.» «**`stacks` never imports `selfupdate`.**» | CLOSED 2026-09-21 — controller v0.261.0: `Manager.AnyUpdating()` → `Updater.SetAppUpdatingCheck`; `UpdatePreflight` refuses `self_updating`; lock does not latch (`TestR608_LockReleasesAfterHold`); four red-proofs. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-609** | **An update refusal has a machine-readable reason inside the process and none on the wire, so an unattended caller cannot tell "wait" from "never" (P3-LOW)** **Reasoning kept:** «`busy`/`updating`/`deploying`/`migrating`/`self_updating` are TRANSIENT and `held`/`downgrade` are TERMINAL until a person acts.» | CLOSED 2026-09-21 — controller v0.261.0: 409 body gains data.reason additively; the held refusal in actionStack (before UpdatePreflight) carries it too; red-proofed. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-611** | **A session reported "everything is done" over a phase it had silently skipped — the process failure, not the missing measurement (P3-LOW)** **Reasoning kept:** «the report's FIRST section is "not done", even when empty.» «Every part and scenario of a brief is listed there if it was skipped, shortened or changed, with the reason.» | CLOSED 2026-09-21 — run by the successor session (scenarios F and G); "not done" is now the report's first section. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/update-arc-gaps-2026-09-21/ | +| **R-614** | **A stale `update_phase` survives a remove and redeploy, so a freshly installed app can read "Frissitve" before it has ever been updated (P3-LOW)** **Reasoning kept:** «The name is the only thing a new install shares with the old one, so the record has to go when the app does.» | CLOSED 2026-09-22 — controller v0.262.0: RemoveStack calls ClearUpdateState; red-proofed. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-623** | **The unattended-update caller turned every SUCCESS into a `timeout`, then refused to press that app again — the instrument, not the box (P3-LOW)** **Reasoning kept:** «an instrument that can report a success as a timeout is not a measurement» | CLOSED 2026-09-21 — fixed in audits/update-arc-gaps-2026-09-21/unattended-caller.py (unwraps data). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/update-arc-gaps-2026-09-21/unattended-caller.py | +| **R-627** | **Nothing checked that the register is a well-formed table; an append ate two rows' state cells and R-254 was broken for 45 days (P2-MEDIUM)** | CLOSED 2026-09-22 — gate 14 (scripts/register_shape_gate.py in repo_gates.py), four decoys, register repaired 317→315. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-628** | **An empty search of a mailbox I do not control was turned into a claim about what a THIRD PARTY had done, and entered the register as fact (P2-MEDIUM)** **Reasoning kept:** «*an empty search may be reported as "absent from the place I looked", never as "it did not happen" — and only after a control query that MUST hit has been seen to hit.*» | CLOSED 2026-09-22 — rule recorded in the Gmail-access memory; R-433 corrected with the real answers. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-629** | **The drill catalog sent the operator 47 CI-failure alarms in one night; the drill method did not mention CI (P2-MEDIUM)** | CLOSED 2026-09-22 — has_actions false on admin/app-catalog-drill, `09` §6.5 method updated; the 47 mails are the operator's to clear. — residue NOT-A-FINDING (2026-10-03): its residue (47 unread CI mails) is the operator's inbox housekeeping, not product work | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-630** | **`paperless-ngx`'s health probe has never run on any box; a successful update stops the app (P2-MEDIUM, raised to P1)** **Reasoning kept:** «A stack with no probe is not healthy and not failing — it is SETTLED ON CONTAINER STATE (`09` §3), and never a reason to stop a running app.» | CLOSED 2026-09-22 — controller v0.262.0 + catalog: the no-probe wait settles on container state, the target is decidable (explicit container field), the gate refuses an ambiguity. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/the-28-2026-09-22/sidejobs/r630-controller-words.txt, audits/the-28-2026-09-22/sidejobs/r630.json | +| **R-631** | **Five templates cannot be judged by the probe gate at all, and one is correct only by accident (P3-LOW)** **Reasoning kept:** «Add `expect: {status: 200}` to that template — a change that looks like a tightening — and home-assistant goes permanently unhealthy and every successful update of it starts stopping it.» | CLOSED 2026-09-22 — all five read live on guest 9202, all five probes correct; home-assistant's fragility is now a number (401 with no expect). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/the-28-2026-09-22/sidejobs/r631.json | +| **R-632** | **Twenty-eight of the 53 templates have never been deployed by any update drill (P3-LOW)** | CLOSED 2026-09-22 — all 28 walked in one night (26 deployed, 6 proven); findings → R-630, R-633, R-634. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/DRILL-the-28-2026-09-22.md, audits/probe-fix-2026-09-22/not-judged.json | +| **R-633** | **A remove sent while a restore runs reports success, deletes the record, and leaves a container restarting forever with a live public route (P2-MEDIUM)** **Reasoning kept:** «`down` returning 0 is a request, not a result» «A refusal from the wrong rule is not evidence for the new one.» | CLOSED 2026-09-22 — controller v0.262.0 (busy guard + verified teardown) + v0.262.1 (typed RemoveBusyError → HTTP 409); both halves proven live on 9202. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/the-28-2026-09-22/apps/gokapi/came-back-evidence.txt, audits/the-28-2026-09-22/apps/gokapi/log.txt | +| **R-701** | **demo-hp's whole-guest restore test can never run: every 6 h it picks the right archive and the space preflight refuses it (P3-LOW)** | CLOSED — option (b), 2026-09-28 (operator, `09` §3 decision 44): restore_storage → nvme-scratch, FelhomAgentStore granted on /storage/nvme-scratch; manual and first scheduled cycle passed. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/version-travel-2026-09-26/D1/D1-cycle-demo-hp.txt, audits/evidence-golden-0276-2026-09-28/phaseD1-reclaim.txt, audits/logins-nvme-2026-09-28/C/ | +| **R-702** | **Every claper install creates an admin `admin@claper.co` with the public password `claper`, published on the household's domain (P1-HIGH)** | CLOSED — claper fixed 2026-09-28 (catalog 9dc8a05, controller v0.279.0 after_install); the class → R-707. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/pg-calcom-claper-2026-09-28/box/C0-claper-default-admin.txt | +| **R-703** | **calcom v6.2.0 cannot start at its catalog memory limit — a fresh install crash-loops and the box stops it (P2)** | CLOSED — catalog 9555e73, 2026-09-28: memory 768M → 1536M (anon peak 817 MiB, 53%). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/pg-calcom-claper-2026-09-28/box/C0-calcom-crash.txt, audits/pg-calcom-claper-2026-09-28/box/C0-calcom-memory.txt | +| **R-708** | **grafana falls back to password `admin` when its admin field is empty (P3-LOW)** **Reasoning kept:** «no default in the compose (`${GF_SECURITY_ADMIN_PASSWORD:?}` refuses to start instead).» | CLOSED — 2026-09-29 (catalog d0e7e2e): `${GF_SECURITY_ADMIN_PASSWORD:?…}` refuses to start empty. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/login-gate-2026-09-29/D/D4-grafana-r708.txt | +| **R-709** | **The deploy page writes the generated admin passwords of installed apps into its HTML (P3-LOW)** | CLOSED — 2026-09-29, controller v0.280.0: password field renders empty with a reveal eye; red-proofs RP14/RP15; live on 9202. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/login-gate-2026-09-29/D/D6-r709-live.txt | +| **R-710** | **An app installed before its template gained an `after_install:` is never warned about its default login (P2-MEDIUM)** | CLOSED — 2026-09-29, controller v0.280.0 (RP13): absent record is "not run yet" only 30 min after install; „Megváltoztattam" press; live on demo-hp. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/login-gate-2026-09-29/A/A2-page-warning-after.txt, audits/login-gate-2026-09-29/A/A3-demo-hp-changed-it.txt | +| **R-711** | **About a dozen class-4 apps keep open sign-up after their first admin exists — the setup gate does not close that (P2-MEDIUM)** | CLOSED — 2026-09-29 (decision 47, controller v0.281.0, catalog 6faf432): per-app signup_block written when the gate opens; wanderer → R-714. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/login-gate-2026-09-29/B/B-VERDICT.md, audits/gate-rollout-2026-09-29/ | +| **R-712** | **wger refused every browser sign-in behind traefik: "CSRF verification failed" (P2-MEDIUM)** | CLOSED — 2026-09-29 (catalog d0e7e2e): CSRF_TRUSTED_ORIGINS + X_FORWARDED_PROTO_HEADER_SET; proven on a fresh install. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/login-gate-2026-09-29/D/D2-live.txt | +| **R-713** | **claper's `after_install` pastes the household's password into Elixir code, and the controller does not refuse a value that would break such code (P3-LOW)** **Reasoning kept:** «a code-bound value holding a quote, backslash, `$`, `{`, `}`, backtick or line break is refused» | CLOSED — 2026-09-29, controller v0.281.0 (RP24) + catalog 6faf432: code-bound unsafe values refused, `${NAME¦base64}` added; proven live. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/gate-rollout-2026-09-29/ | +| **R-714** | **wanderer cannot be gated: its web part calls its own database host through the public name (P2-MEDIUM)** | CLOSED — 2026-09-29 (decision 48, controller 0.282.0): resolved without a gate; sign-up closed by "Close sign-up now" block + PUBLIC_DISABLE_SIGNUP; proven on 9202. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/signup-lock-2026-09-29/ | +| **R-715** | **The setup gate's probe reads only an HTTP-200 JSON object, so three apps with a real status get the button (P3-LOW)** **Reasoning kept:** «a catalog gate that refuses a probe without a before/after measurement in its comment.» | CLOSED — 2026-09-29, controller v0.282.0 (RP29, RP30): list indexes in field, done_status:; new catalog gate check-probe-measured.py (5 decoys). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/signup-lock-2026-09-29/ | +| **R-716** | **Apps installed before controller 0.281.0 keep their open sign-up — decision 47 closes it only where the box opened the gate (P3-LOW)** **Reasoning kept:** «a catalog change never touches an installed app» | CLOSED — 2026-09-29: operator ruled A (decision 49); built in controller v0.282.0 and pressed on demo-hp adventurelog/opengist and demo-felhom opengist. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/gate-rollout-2026-09-29/0/P0-3-demo-boxes-after-push.txt, audits/signup-lock-2026-09-29/ | +| **R-720** | **A new household's apps are not in the off-site copy: every app starts „3. mentés Kikapcsolva" (P2-MEDIUM)** | CLOSED 2026-09-30 — controller v0.283.0 (decision 50): fresh install on an off-site box switches the app ON; one-press offer for older apps. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/evidence-fixes-first-tester-2026-09-30/partA/ | +| **R-721** | **The household presses Stop during a whole-guest backup, and the backup starts the app again (P2-MEDIUM)** **Reasoning kept:** «unquiesce restarts only stacks whose desired state is still running» | CLOSED 2026-09-30 — controller v0.283.1, proven live (v0.283.0 was wrong in production; v0.283.1 wires the adapter and pins the production types). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/evidence-drill-new-household-2026-09-30/phase1/step10-stop-undone-by-quiesce.log, audits/evidence-fixes-first-tester-2026-09-30/partE/ | +| **R-722** | **The volunteer guide is stale in five places a volunteer reads literally (P2-MEDIUM)** | CLOSED 2026-09-30 — both guides rewritten as measured (operator: CC rewrites); every changed line dated. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: runbooks/VOLUNTEER-first-hour.md | +| **R-727** | **The whole-guest restore test picks a PREVIOUS box's archive, fails on its key, and the ✗ is labelled with the wrong tier (P2-MEDIUM)** **Reasoning kept:** «The restore test skips an archive written with another key (logged by name).» | CLOSED 2026-09-30 — agent v0.138.0 + decision 51 (restore test skips archives written with another key); ✗ card names the tier in controller v0.283.0; ep0 drill archives in tester-1 removed; delivered by signed jobs to both demo boxes. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/evidence-fixes-first-tester-2026-09-30/partC/ | +| **R-730** | **The published ISO 1.29.0 was built from a tree that git cannot name, so it cannot be tagged (P3-LOW)** **Reasoning kept:** «`scripts/iso/build-felhom-iso.sh` now runs a clean-tree gate right after argument parsing: any uncommitted or untracked change, or HEAD ≠ `origin/main`, refuses the build (no bypass flag); the manifest records `repo-commit` from the gate and `iso-version-tag : iso-v`.» «the `installer-v*` tags are the host-install SCRIPT's versions (`SCRIPT_VERSION` 1.28.0 on main; the website serves `/scripts/` from `installer-v1.28.0`), a different number line from the ISO — an `installer-v1.29.0` tag would block the next script release.» | CLOSED 2026-09-30 — the gate + test; 1.29.0 untagged by decision (its build tree is not a commit). build-felhom-iso.sh clean-tree gate (no bypass), manifest records iso-v; test scripts/iso/test/clean-tree.sh red-proofed. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/pg-last-six-2026-09-30/E/E4-installer-tag.txt, scripts/iso/test/clean-tree.sh | +| **R-732** | **immich's FIRST start at the catalog pin could not finish on the bench: its database container was OOM-killed at 512 MiB (P2-MEDIUM)** | CLOSED 2026-09-30 — catalog `56c4888` (768M + v3.2.4); older step 0b8272068aab36bf re-proven at 768M (catalog `48440ce`, `63a96b0`); an installed immich gets it with its next guarded Update. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/pg-last-six-2026-09-30/F/immich-first-start-oom.txt, audits/immich-first-start-2026-09-30/A-cause.md, audits/more-night-apps-2026-09-30/ | +| **R-735** | **The test bench generates a `password:N:special` deploy field WITHOUT a special character, so calibre-web refuses it (P3-LOW)** | CLOSED 2026-09-30 — catalog `5b1972b` (`_gen` mirrors `randomWithSpecial`; test seen failing first). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/more-night-apps-2026-09-30/bench/apps/calibre-web/, A/R735-red.txt, A/R735-green.txt | +| **R-736** | **Removing or updating an app never deletes the image it pulled, so a box's Docker disk fills until an install or update is refused (P2-MEDIUM)** **Reasoning kept:** «Keep set box-wide at delete time; exact id; no pass while an update runs; one summary line per pass; a one-time sweep.» «the Remove button runs `RemoveStack`, not `DeleteStack`; `docker image ls` without `-a` hides the untagged digest-pulled images» | CLOSED 2026-09-30 — controller v0.284.2, floor 0.284.2 (decision 53, option A; 0.284.0/0.284.1 never floored). Old controller images → R-745. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/more-night-apps-2026-09-30/box/T2-images-removed-by-name.txt, audits/night-rulings-2026-09-30/ | +| **R-737** | **wger's app login API answers 500 on a CORRECT password: the template sets no `JWT_PRIVATE_KEY` (P3-LOW)** | CLOSED 2026-09-30 — catalog `45d8482` (RSA pair generated once by wger's `manage.py generate-jwt-keys`, kept 0600 on the data volume; bench-proven, not proven on a box). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/more-night-apps-2026-09-30/B/wger-probe.txt, audits/night-rulings-2026-09-30/ | +| **R-740** | **A same-tag upstream security fix (postgres:18-alpine, redis:7-alpine, mariadb:12.3 …) reaches NO box, because the catalog never records a re-test of a tag at a new digest (P2-MEDIUM)** **Reasoning kept:** «Not a cron job (runbook `runbooks/monthly-floating-retest.md` says why); who presses it monthly is the operator's word (STATUS).» | CLOSED 2026-09-30 — built (catalog `6a3ead9`); the monthly run is a standing step. Decision 52 option A: re-test entries + `scripts/retest-floating.py`; proven end to end on 9202. Exact-tag re-pushes → R-743. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/more-night-apps-2026-09-30/C/C3-floating-repush.txt, audits/more-night-apps-2026-09-30/C/C4-same-tag-retest.txt, runbooks/monthly-floating-retest.md, audits/night-rulings-2026-09-30/ | +| **R-741** | **For a few seconds after a fresh install, an `after_install` app answers its PUBLIC default login through the front door (P3-LOW)** **Reasoning kept:** «an `after_install` app is installed HELD behind the setup gate's door until the login is replaced (or the household says it changed it).» | CLOSED 2026-09-30 — controller v0.284.2 (install hold: an after_install app is installed held behind the setup gate's door until the login is replaced; red-proofs RP-IH1..4). mealie lock → R-747. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/more-night-apps-2026-09-30/A/A3-calibre-default-login-window.txt, audits/night-rulings-2026-09-30/ | +| **R-742** | **zipline 4.8.0 cannot be reached from 4.6.1 in one step: it refuses to start until the database ran the release before it (P3-LOW)** | CLOSED 2026-09-30 — catalog `a9700e2` + `fb87030` (two steps 4.6.1 → 4.7.0 → 4.8.0, each proven on both venues). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/more-night-apps-2026-09-30/bench/apps/zipline/, audits/more-night-apps-2026-09-30/box/zipline/ | +| **R-743** | **Exact version tags are re-pushed under the same name too — not only floating lines (P3-LOW)** **Reasoning kept:** «decisions 54 (a CC session the operator starts with the standing brief runs it monthly) and 55 (every app with a proven ladder; `--engines-only` a switch).» | CLOSED 2026-10-01 — decisions 54, 55; both tags re-tested (nextcloud catalog `3b59dfb`, sonarr `1a37032`). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/night-rulings-2026-09-30/A/, audits/retest-2026-10/ | +| **R-744** | **outline's fixture cannot seed outline 1.10.1: no `csrfToken` cookie after `installation.create` (P3-LOW)** | CLOSED 2026-10-01 — fixture fixed, proven on 9202 (accepts `__Host-csrfToken`; catalog `9fc7052`). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/rulings-2026-10-01/D/D2-outline-fixture-9202.txt | +| **R-745** | **Old CONTROLLER images are never deleted: ~50 versions on each demo box (P3-LOW)** **Reasoning kept:** «the agent rolls back to the RUNNING image (what `/etc/felhom-controller-image` named), never to a previous one the controller hands it.» | CLOSED 2026-10-01 — controller v0.285.0, floor 0.285.0 (decision 56: keep running + previous, delete older/untagged). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/night-rulings-2026-09-30/E/, audits/rulings-2026-10-01/B/ | +| **R-746** | **`image_digest.resolve` ignores a `@digest` in its argument — it answers the TAG's current digest (P3-LOW)** | CLOSED 2026-10-01 — catalog 804884a (manifest asked by digest; malformed digest refused; scripts/test_image_digest.py red-proofed). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/rulings-2026-10-01/D/D1-r746-digest.txt | +| **R-748** | **The register-shape gate skipped every row whose id has a letter suffix (R-88a, R-88b, R-209a) (P3-LOW)** | CLOSED 2026-09-30 — `scripts/register_shape_gate.py` (`R-\d+[a-z]?`; decoy `suffix-row-eaten-state`). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-749** | **`retest-floating.py` could never start on a fresh bench: it checked for `/opt/upg/upgrade-test.py` before the step that copies it (P3-LOW)** | CLOSED 2026-10-01 — catalog 9e53205. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/rulings-2026-10-01/A/ | +| **R-750** | **The registry no longer holds controller releases older than 0.213.0 — something removed them, and nothing records what (P3-LOW)** **Reasoning kept:** «`gitea-image-prune.sh` keeps the newest 20 + every version in use (floor, vouched golden, vouched agent, `min_agent`, the golden's baked images, the running hub), refuses (exit 3) when that list is unreadable» | CLOSED 2026-10-01 — decision 62, misc-scripts c9d5ed5 (cause: manual `gitea-image-prune.sh --keep 7` run 2026-08-22/23, HM-024; prune now keeps newest 20 + every in-use version, exit 3 when unreadable). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/rulings-2026-10-01/B/, audits/lockouts-2026-10-01/C/C1-registry-read.txt, audits/calibre-name-and-prune-2026-10-01/B/ | +| **R-751** | **The image clean-up after an app update could crash the whole controller: nil stack dereference when the app was gone (P2)** | CLOSED 2026-10-01 — controller v0.285.0, floor 0.285.0 (TestRetainImagesAfterUpdate_AppGoneDoesNotPanic seen panicking first). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/rulings-2026-10-01/B/B1-red-proofs.txt | +| **R-752** | **Four more catalog apps let a stranger lock the household out with wrong passwords for a known login name — like mealie (P3-LOW)** **Reasoning kept:** «Installed apps: a settings-only change reaches the stack file at the next sync (images equal, ≤15 min) and the running app at the next `compose up -d` — Restart/Start (measured: the env changed only at Restart), an Update, or a backup's restart (`backup.go:972`, read).» | CLOSED 2026-10-01 — decisions 58–61 (wger catalog `82fff32`; BookStack, Grafana kept; calibre-web generated ADMIN_USER catalog `e9f50b5`). Installed calibre-web's invented name → R-757. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/lockouts-2026-10-01/B/ | +| **R-753** | **Behind the tunnel every visitor reaches an app with the SAME address, so every per-address guard is an 'everyone' guard (P3-LOW)** **Reasoning kept:** «cloudflared alone on `felhom-tunnel` at the fixed `172.16.253.2`, traefik trusts forwarded headers from it only, an entrypoint middleware removes client-writable host/path/address headers» | CLOSED 2026-10-01 — controller v0.286.1 (decision 63; v0.286.0 never floored; catalog `04e9516`..`50e4fb4`) (rows R-776..R-779 carry what is left). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/lockouts-2026-10-01/A/A1-client-address.txt, audits/visitors-2026-10-01/A/ | +| **R-754** | **`01-topology-and-trust.md` §7 says cloudflared runs on the Proxmox HOST; on every box it runs INSIDE the guest (P3-LOW)** **Reasoning kept:** «the data path is not independent of the guest.» | CLOSED 2026-10-01 — document corrected (operator: the build is right). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-765** | **A new app that builds its login from a `type: password` value at EVERY start loses the login on a restore after a remove — Radicale (P2-MEDIUM)** **Reasoning kept:** «a restore carries only `type: secret` values (`stacks.PortableSecretEnvVars`, the D5 ruling — a `type: password` is an internet-reachable login and stays out of the drive's backup)» «the login file is written on the FIRST start only and lives on the data volume.» | CLOSED 2026-10-01 — Radicale's template, before publishing (login file written on first start only, on the data volume). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/new-apps-2026-10-01/box/radicale-attempt1/restore.txt, audits/new-apps-2026-10-01/box/radicale/restore.txt | +| **R-767** | **MeTube has no login at all, by design, and the box has no permanent household-only door — so it is not built (P3-LOW)** | CLOSED 2026-10-02 — controller v0.287.0 + catalog (MeTube published) behind the family gate, decision 64, no exception. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/new-apps-2026-10-01/FIT.md, audits/family-gate-2026-10-02/A/items.txt, audits/family-gate-2026-10-02/B/box/metube-fresh.txt | +| **R-772** | **A health probe that finds NO container to probe records `healthy: true` (P3-LOW)** **Reasoning kept:** «a not-run probe records `healthy: false, not_checked: true`, is looked at on the next tick» | CLOSED 2026-10-01 — controller v0.286.1 (not-run probe records healthy false, not_checked true; red-proofs RP-D1, RP-D1b). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/new-apps-2026-10-01/box/karakeep-768M/neg-and-crawl.txt, audits/visitors-2026-10-01/D/ | +| **R-773** | **After a remove + restore, an app's sign-up route block (decision 47) is gone (P3-LOW)** **Reasoning kept:** «a REMOVED app restored from its backup gets the lock record (`opened_by: restore`) and the block written before anything starts» | CLOSED 2026-10-01 — controller v0.286.0 (restored app gets lock record `opened_by: restore` + block before start; red-proof RP-D2). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/new-apps-2026-10-01/box/karakeep/restore.txt, audits/visitors-2026-10-01/D/r773-live.txt | +| **R-780** | **A permanent household gate with family accounts — the spike PASSED; the build waits for the operator's go (P2-MEDIUM)** **Reasoning kept:** «**Build requirement F1:** exceptions must be anchored (`PathPrefix(/api/v1/opds)` also matched `/api/v1/opdsx`).» | CLOSED 2026-10-02 — controller v0.287.0 (decision 64; internal/family, Család card, family_gate fields, catalog gate check-family-gate.py; Grimmory and MeTube published). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/permanent-gate-2026-10-01/VERDICT.md, audits/family-gate-2026-10-02/A/items.txt, audits/family-gate-2026-10-02/E/floor.txt | +| **R-787** | **No catalog app's LICENCE has been checked against Felhom being a paid service (P2-MEDIUM)** **Reasoning kept:** «New apps carry their licence in record row 0.1.» | CLOSED 2026-10-02 — the read (decisions are R-784, R-789..R-795). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/licences-2026-10-02/TABLE.md | +| **R-788** | **The volume-persistence gate reads an EMPTY declared volume as CLEAN (P3-LOW)** **Reasoning kept:** «an app that wrote nothing is UNDETERMINED (exit 2), never a pass.» | CLOSED 2026-10-02 — catalog (empty declared volume after the exercise is UNDETERMINED, red-proofed; APP_EXERCISE added). Bind mounts → R-805. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/family-gate-2026-10-02/C-metube-bench/C2-volume-persistence-metube.txt, audits/persistence-sweep-2026-10-02/A/RP-R788-empty-volume.txt | +| **R-790** | **Emby is proprietary: a personal, non-commercial, non-transferable licence (P2-MEDIUM)** **Reasoning kept:** «KEPT — the household is the licensee; Felhom installs the official, unchanged image.» | CLOSED 2026-10-02 — operator ruling (kept), decision 66; part of the lawyer's review R-802. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/licences-2026-10-02/TABLE.md | +| **R-791** | **n8n's Sustainable Use License allows own internal or personal use; distribution only free of charge for non-commercial purposes (P3-LOW)** **Reasoning kept:** «KEPT — the household is the user; the official, unchanged image.» | CLOSED 2026-10-02 — operator ruling (kept), decision 66; part of the lawyer's review R-802. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/licences-2026-10-02/TABLE.md | +| **R-792** | **Plex is proprietary: the HOUSEHOLD is the licensee under its own Plex account (P3-LOW)** **Reasoning kept:** «KEPT — the household is the licensee under its own Plex account.» | CLOSED 2026-10-02 — operator ruling (kept), decision 66; part of the lawyer's review R-802. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/licences-2026-10-02/TABLE.md | +| **R-795** | **recipe-importer, Felhom's own image, has no licence file (P3-LOW)** **Reasoning kept:** «Felhom's own code — no action unless it is shared with anyone.» | CLOSED 2026-10-02 — operator ruling (no action unless shared), decision 66. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/licences-2026-10-02/TABLE.md | +| **R-800** | **[P2-MEDIUM] "Delete your data from the hard drive" keeps the app's files in the household's userdata folder — and neither the dialog nor the result says so.** | CLOSED 2026-10-03 (triage, verified) — controller v0.288.0 (`580b656`, decision 67) — the remove dialog and result name the kept userdata folder; live on 9202 (`0945332`) | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| **R-801** | **The volume-persistence gate never sent a request to ANY app: it read the routed port from label VALUES, not the label NAME (P2-MEDIUM)** | CLOSED 2026-10-02 — catalog (the fixed gate `routed_ports()` + the re-sweep of all 58: CLEAN 39 · UNDETERMINED 19 · BROKEN 0). Follow-ups R-805..R-807. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/family-gate-2026-10-02/B/volume-persistence-metube-exerciser-diag.txt, audits/family-gate-2026-10-02/B/RP-R801-routed-ports.txt, audits/persistence-sweep-2026-10-02/A/sweep/TABLE.md | +| **R-803** | **papra cannot start under its own memory limit: a fresh install is OOM-killed in its migration and crash-loops (P1-HIGH)** | CLOSED 2026-10-02 — catalog (papra 768M, mem_request 256M; measured from birth, 0 kills, persistence CLEAN; no box ran papra). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/persistence-sweep-2026-10-02/A/sweep/papra-oom-diagnosis.txt, audits/persistence-sweep-2026-10-02/A/papra/SUMMARY.md | + +**Rows that never had an `R-` id** (the 2026-07-28 table of campaign findings and watches): + +| Row | What | Closed | Full text | +|---|---|---|---| +| E-2d | **Prove E-2 on a fresh VM — a real host-install 1.22.0 run, Case B, a claimable customer, add a drive and unplug it.** | CLOSED — PARTIALLY PROVEN (2026-07-29). C1, C2 proven; C3, C4 proven live; C5 FAILED → R-116 (the single named open leg; R-116 since CLOSED in CLOSED-ITEMS). Per Session-C runbook §9 a failed claim closes as partially proven, no re-run. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/E2D-fresh-vm-2026-07-29.md, audits/SESSION-C-2026-07-29.md | +| — (watch, line 285) | **Storage Box **snapshots** on `storage-box-pool-1` — plan SET (daily 00:00, keep 7) but **0 taken yet**** | CLOSED 2026-10-03 (triage) — snapshots exist: R-429 recorded seven daily Storage Box snapshots (2026-09-01) | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| — (watch, line 289) | **demo-felhom's next weekly PBS backup (newest is 2026-07-26)** | CLOSED 2026-10-03 (triage) — folded into R-91, whose trigger is exactly this backup | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| — (watch, line 290) | **demo-felhom's next restore-test (84 h cadence, last 2026-07-27 06:38 UTC)** | CLOSED 2026-10-03 (triage) — restore tests have run on demo-felhom on their cadence ever since (R-672, R-689 in `CLOSED-ITEMS.md`) | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| F-CRIT-2 | **~~A failed offsite backup left a phantom snapshot that RESET the tier's freshness clock — 7 days silent~~** **Reasoning kept:** «`NewestArchiveTime` now counts only plausibly-complete entries (measured 1 MiB floor; undecidable ⇒ not counted).» | SHIPPED + PROVEN-LIVE (agent v0.106.0, 2026-07-28). Campaign fault 2 replayed on demo-hp: phantom rejected + logged once, tier DUE and backed up, no thrash on the inverse. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| F-CRIT-1 | **~~An app that fails to restart after a quiesce never alarms on any channel~~** | SHIPPED + PROVEN-LIVE (controller v0.179.0, 2026-07-28). Live on demo-hp: alarmed 9s after grace expiry; a deliberate user stop stayed silent through 9 dead-app scans. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| F-A1 | **~~A restore-test in flight made a healthy backup report as FAILED (HTTP 409 read as a tier failure)~~** **Reasoning kept:** «409 → contention: tier stays DUE, dropped before anything stops (15m), and BLOCKED alarm if contention outlives the agent's 120m ceiling (3h).» | SHIPPED + PROVEN-LIVE (controller v0.179.0, 2026-07-28). Hub DB: 409 → 0 operator emails, real failure → 1. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| C9-F1 | **~~Tier-2 „Fájlok visszaállítása" is offered for apps whose copy has no restorable file leg; restores 0 files and reports all files present~~** **Reasoning kept:** «`Tier2RestoreCoverage` refuses UP FRONT without stopping the app and NAMES the working action; a run that proceeds claims only what it **examined** and discloses that the database and volumes are not covered.» | SHIPPED + PROVEN-LIVE (controller v0.183.0, 2026-07-28). Live on demo-felhom: bookstack refused without restart; paperless re-run byte-identical, 16/16 docs. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| C9-F2 | **~~An app in a Docker crash loop never alarms on any channel; `StateRestarting` is in no down-set~~** **Reasoning kept:** «`StateRestarting` deliberately NOT added to `IsDownState` (that alarms on every deploy fleet-wide)» | SHIPPED + PROVEN-LIVE (controller v0.183.0, 2026-07-28). Sustained restarting becomes down after `crashLoopAfter`=5m; dashboard counter uses the same predicate. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| C9-F3 → R-104 | **An **interrupted offsite run leaves an exclusive restic lock the self-heal cannot reach**: `resticStep` (`offbox.go:634-648`) has `unlock --remove-all`, but `ensureOffboxRepo`'s probe fails first, `classifyResticProbe` (`offbox.go:77-93`) has no lock case → …** | FOLDED into R-104 — the same finding; R-104 has its own open row | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| D5 | **~~Move app secrets into the LOCAL recovery unit so Tier-1/Tier-2 restore stop needing the guest~~** **Reasoning kept:** «**Operator ruling 2026-07-30: `type: secret` travels (45 fields), `type: password` NEVER (7) plus a code register (`vaultwarden/ADMIN_TOKEN`); plaintext.** The exclusion is what LICENSES the plaintext — coupled, not independent.» «the register is **code, not a catalog flag** (a boundary a catalog push can move is not a boundary — R-97a).» «**Precedence: the UNIT WINS** over the guest, because the unit's secrets were captured in the same run as the dumps beside them and therefore match the data being restored; pinned both directions.» | SHIPPED + PROVEN-LIVE (controller v0.188.0, 2026-07-30). Operator ruling 2026-07-30 on which secrets travel; manifest schema 2; live proof on scratch guest (AdventureLog secrets recovered 2/2, app read seeded row over TCP); 4 red-proofs. Related: R-127. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/D5-drive-alone-restore-2026-07-30.md | +| F-DIAG | **~~Four distinct offsite failure causes collapse into two operator-visible strings~~** **Reasoning kept:** «Unclassifiable says so rather than being folded into a neighbour.» | SHIPPED (controller v0.182.0, 2026-07-28). `ClassifyOffsiteFailure` → quota / orphaned / no_repo / no_units / transport / unknown; redaction by the target's actual host/user/path. Unit-proven; not yet exercised by a live offsite failure of each class. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| F-OPS | **~~A manual `pct restore` inherits the source guest's bind mounts — during a real DR, on a different host, under pressure~~** **Reasoning kept:** «Docs only by design — the agent already neutralises binds on its own restore paths, and a second implementation would drift» | DOCUMENTED (2026-07-28) in `runbooks/RUNBOOK-manual-guest-restore.md` (mpN volumes vs binds, the mp9 source-VMID trap, strip-and-re-add, positive pre-start verification). | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: runbooks/RUNBOOK-manual-guest-restore.md | +| F-REBOOT | **~~A guest rebooted during its backup does not come back — 9m47s outage with every alarm silent~~** **Reasoning kept:** «`onboot` is the deliberate-stop discriminator (already the stale-lock path's, and what `pve-guests` consults)» | SHIPPED + PROVEN-LIVE (agent v0.107.0, 2026-07-28). 60 s guest-power watchdog, retry 3x/1m-2m-4m then escalate. Live on demo-hp: 120 s unattended vs 587 s; `onboot:0` guest left stopped. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| F-LEAK | **~~A failed restore-test cannot destroy its own scratch guest (403 `VM.Allocate`); the 10-slot VMID band shrinks silently~~** **Reasoning kept:** «4th root-fenced exception, band enforced in sudoers **literally** (`pct destroy 99000[0-9] --purge`) + in code + at the caller; API destroy still tried first.» | SHIPPED + PROVEN-LIVE (agent v0.110.0 + host-install v1.21.0, 2026-07-28). Third attempt shipped: 4th root-fenced exception; live band permitted, 9201/9100/9999/990010/1 and `pct start 990000` refused. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| F-OBS | **~~`deadapp-check` leaves NO positive observable on a default (info-level) box~~** | SHIPPED + PROVEN-LIVE (controller v0.180.0 + agent v0.109.0, 2026-07-28). INFO summary every 20th scan; agent v0.109.0 fixes the same shape in the guest-power watchdog. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| E-2 | **~~Drive-role machinery around the moved vzdump target~~** | CLOSED — PARTIALLY PROVEN (Session C, 2026-07-29). C1/C2 proven in E-2d; C3, C4 proven live (R-114, R-112); C5 FAILED → R-116. Broken legs R-112, R-113, R-114, R-116 all since CLOSED (CLOSED-ITEMS). Arc's definition of done: R-106+R-109, R-108, D5. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`; evidence: audits/SESSION-C-2026-07-29.md, audits/E2D-fresh-vm-2026-07-29.md | +| E-2a | **~~The target move needs a root-fenced wrapper — the agent cannot do it~~** **Reasoning kept:** «the agent's PVE role was NOT widened.» «has NO storage-removal path (grep-assertable), is idempotent and refuses to repoint.» | SHIPPED + PROVEN-LIVE (agent v0.113.0 + host-install v1.22.0, 2026-07-29). `felhom-backup-target-apply` behind literal `FELHOM_BACKUPTARGET` sudoers alias; all five laws proven live as root on demo-hp, 0 stray storages. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| E-2b | **~~`NotifyStorageDisconnected`/`Reconnected` defined and called NOWHERE — a drive going absent emitted no event~~** | SHIPPED + PROVEN-LIVE (controller v0.184.1 + agent v0.112.0 + hub v0.81.0, 2026-07-29). Seam wired in `ReconcileDriveGates`; keying bug (guest path vs host MountPath) caught before deploy. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | +| E-2c | **~~E-1 put the whole-guest backups on a drive `POST /disks/eject` would eject~~** **Reasoning kept:** «NOT a role reclassification — `RoleForStorage` untouched, because on both boxes that drive is ALSO the enrolled user-data drive» | SHIPPED + PROVEN-LIVE (agent v0.112.0, 2026-07-29). Eject + decommission refuse 409 on the backup-target mount; refused live on both boxes. | full text: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` | + ## 2026-09-17 — nothing decides by reading a Hungarian word (controller v0.251.0) | Row | What | Closed | Full text | diff --git a/documentation/backlog/FOLLOWUP-nas-automount-guest-reboot-reassert.md b/documentation/backlog/FOLLOWUP-nas-automount-guest-reboot-reassert.md index 5038df6c..f015ac09 100644 --- a/documentation/backlog/FOLLOWUP-nas-automount-guest-reboot-reassert.md +++ b/documentation/backlog/FOLLOWUP-nas-automount-guest-reboot-reassert.md @@ -1,3 +1,6 @@ +> **Status 2026-10-03 (backlog triage): SHIPPED** — agent v0.84/v0.85 (CAMPAIGN-3). Kept here as history because +> `CONTEXT.md`, `controller/network-storage-nas.md` and the controller's `CONTEXT.md` link to this path. + # FOLLOW-UP — NAS automount trigger does not survive a guest reboot (agent reassert gap) **Opened:** 2026-07-11 · **Severity:** HIGH (data-safety-adjacent for media apps) · **Class:** diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 9fcb407a..1a9f79ec 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -5,7 +5,16 @@ page keeps only what is **open**, and it is the file to read first. Root `REPORT and **overwritten** — nothing durable may live only there; a session that must not clobber it writes a non-overwritten `REPORT-.md` sibling instead (`CLAUDE.md:82-87`), of which 14 now exist. -State: `BLOCKED` · `READY` · `WAITING-ON-OPERATOR` · `WATCHING`. Every row has an owner. + +State: `READY` · `OPEN` · `BLOCKED` · `WATCHING` · `WAITING-ON-OPERATOR` · `NARROWED` · `DEFERRED` · +`VERIFY`. Every row has an owner. **A row whose state LEADS with a finished word (CLOSED, SHIPPED, FIXED, +DECIDED, …) does not belong here** — it moves to `CLOSED-ITEMS.md` in the same commit that finishes it, and +`scripts/closed_register_gate.py` RULE 3 refuses the push otherwise (2026-10-03). + +**2026-10-03 — the triage.** 125 finished rows moved to `CLOSED-ITEMS.md` (and 20 old rows that never had +an id); the campaign write-ups, the 2026-08-04 rulings and the old ranking paragraphs that sat between the +tables moved, word for word, to `documentation/archive/OPEN-ITEMS-narratives-2026-10-03.md`. The full +previous text of this file: `git show 9e2786c:documentation/backlog/OPEN-ITEMS.md`. ## DECIDED — the backup and restore arc is CLOSED FOR BETA (2026-09-01) @@ -67,510 +76,104 @@ match what the reader sees is how an instrument stops being believed (R-421). It **nothing was proven on the day this line was drawn** and a stopping line that moves a status is a stopping line that lies. -## Operator rulings — 2026-08-04 - -Recorded here because a ruling that lives only in a conversation binds nobody (the R-96 standing rule). - -1. **Run the recovery drill, after R-198.** R-198 shipped in hub v0.93.0; **the drill is the next - session.** Design: `audits/RECON-offsite-dr-chain-2026-08-04.md` §10 — demo-hp, ~3–4 h, a recovery - code created and KEPT, a sentinel file, wipe, reinstall, recover, and **pass = a byte-identical - sha256, not "the repository opened"**. → R-201 -2. **Delete the orphaned ciphertext** (~1.2 GB across the two demo boxes, in set-aside restic stores - nothing prunes). **STILL OWED — not done in v0.93.0.** It is a destructive act on a protected - endpoint and belongs to a session that is scoped for it, not to a release that ships a schema - change. → R-193 -3. **Accept the risk on R-193 candidate (c)** — no repository password retained on the Proxmox host. - **This is what makes R-198 load-bearing rather than tidy:** with no host-retained copy, the - customer-present recovery path is the ONLY way back from a rebuild, and that path runs entirely - through the retained identity blob. → R-193, R-199, R-200, R-201 - -~~**Still open and untouched by v0.93.0:** R-199, R-200, R-201.~~ **UPDATED 2026-08-04 (evening), after -hub v0.94.0 + agent v0.125.0 + controller v0.195.0:** - -- **R-199 — CLOSED, proven on hardware.** Chain links 6–8 are assembled and walked. The offsite - repository password came back out of the sealed bundle **byte-identical** to the one on disk - (`c60c8bc737a6…` from three independent sources: the box's file, the recovered bundle, and the hash - the hub already stored). -- **R-200 — the plumbing half shipped**, the customer-facing form did not, deliberately. -- **R-201 — PASSED 2026-08-04 (night run).** A customer's file survived a machine rebuild and came - back byte-identical, through the customer's own restore flow. **It passed only because a person was - there:** four manual interventions stood between the recovered key and the restored file, none of - them in any design document → R-204. -- **R-202 — untouched.** The orphan card still promises recoverability unconditionally. -- **The orphaned ciphertext deletion (~1.2 GB) is STILL OWED** — ruling 2, above. - -**UPDATED 2026-08-05, after controller v0.198.0 + hub v0.95.0 (R-204):** - -- **R-204 items 1–3 — CLOSED.** The reset code works without a restart; a re-issue no longer marks a - healthy escrow stale (→ **R-196 CLOSED**); a unit restore states what it did NOT restore. -- **R-204 item 4 — OPEN and unstarted:** a rebuilt box cannot obtain an off-site credential unaided. - It needs an operator ruling on the one-shot credential design → **R-193**. -- **Still open and untouched by this session, stated so nothing is presumed closed by association:** - **R-202** (the orphan card's unconditional promise), **the ~1.2 GB orphaned-ciphertext deletion** - (ruling 2 — still owed, still needs its own scoped session), and **R-198's retention, which remains - UNIT-PROVEN ONLY** — nothing has superseded a key in production, and proving it needs a SECOND - deliberate wipe. That retention drill is the next item, and it is not this session's. - -v0.93.0 made the key survive. v0.94.0/v0.125.0/v0.195.0 make it come back. v0.197.0 got the file into -the snapshot and the drill got it out again. **v0.198.0/v0.95.0 remove three of the four crutches the -drill needed — the fourth is R-193, and until it goes the recovery is still operator-assisted.** - -## R-201 — THE RE-WALK, 2026-08-06 (attended) - -**The question was asked a second time, on the fixed build, on a brand-new appliance built from the -published ISO. The answer is still no — but it is a nearer no.** - -| half | verdict | -|---|---| -| **the data** | **PASS** — all three sentinels byte-identical, including a 12 MB binary and an accented Hungarian filename **whose NAME BYTES are also byte-identical** (verified as hex, not as rendered text). Restored in **23 s** out of the pre-destruction snapshot `a7bc23bd`, through the customer's own two-step full-restore flow | -| **the journey** | **FAIL — two dead ends**, against Phase 1's four. One needed a command line **inside the guest**, one a **Proxmox-host** action | - -**RTO: the unaided figure is STILL UNDEFINED**, because the unaided journey still does not complete. -Attended: login 11:42:22 → key placed **+45 s** → tier up **+24 m 12 s** (after intervention 1) → -all three sentinels restored and verified **+30 m 13 s** (after intervention 2). **The 30 m figure must -not be quoted as the customer number.** The only segment that reflects the product working alone is -**23 seconds** to pull 12.8 MB back once everything was in place. - -**The two dead ends:** **R-218's consume half** (row corrected above) and **R-220** (drives -unenrollable after a rebuild — reproduced and red-proved again; without it no app can be redeployed, -and without a redeployed app the restore page is empty, which is R-213's territory and follows from -R-220 rather than being separate). - -**What PASSED and is worth keeping:** the recovery screen **appeared without being sought** -(`/` → `/launcher` → `/recovery`); it answered all three of its questions and its seal date matched the -hub's `created_at` exactly; the emailed reset code worked **first try**; the unlock took **1.528 s** — -a real unseal — and placed the key; and **R-225's fix was seen working in the wild** (the store read -„a pillanatképek száma még ismeretlen" rather than a false zero). - -**R-216 part 4 reproduced live:** the reinstall **downgraded the agent 0.126.0 → 0.125.0**, back to the -vouched version — an operator's hand-fix undone by the very event that makes recovery necessary. - -⚠ **THE DELIVERY GAP, and it is owed.** A fresh install landed on controller **0.201.0** / agent -**0.125.0** — the vouched versions, **neither carrying the fixes**. They were installed **by hand**. -Fleet delivery needs a golden carrying 0.202.0 **and** a vouched agent 0.126.0. **Nothing was vouched; -that is the operator's act.** **This re-walk proves the JOURNEY on the fixed build; it does NOT prove a -customer would receive that build.** - -Evidence: `tests/rewalk-r201-2026-08-06/journal.md`. - -## CAMPAIGN 11 — the recovery journey, 2026-08-05 - -**The whole journey was walked end to end for the first time, on a throwaway appliance built from the -published ISO. The data came back byte-identical; the journey did not exist.** **Ten findings from -Phases 1 and 3, R-214 … R-223** (seven fixed in controller **v0.201.0** + hub **v0.97.0/0.97.1**; -**three deliberately still open**, each blocking a real flow), **plus five from Phase 2's injected -faults, R-224 … R-228.** Evidence: `tests/campaign11-evidence-2026-08-05/journal.md` (Phases 0/1/3) -and `journal-phase24.md` (Phases 2/4). Campaign document: -`audits/CAMPAIGN-11-recovery-journey-2026-08-05.md`. - -> **Phase 2's verdict in one line.** The cryptography, the retention and the transport all work and -> are now proven live. **What fails is being told the truth:** a mistyped code, a hub outage, a -> stopped agent and a correct code for a retained earlier package all produce **one** message, and -> three of the four are wrong. - -| ID | What | State | -|---|---|---| - -### Phase 2 — the injected faults, 2026-08-05/06 (unattended) - -> **ALL FIVE CLOSED 2026-08-06** in controller **v0.202.0** + agent **v0.126.0**. The rule they now -> enforce, stated so it outlives them: **on the unlock path the customer is blamed only after a real -> attempt REFUSED their code; every other outcome, including an unclassifiable one, says something -> else.** -> -> **STILL OPEN AND DELIBERATELY UNTOUCHED BY THAT WORK — said explicitly rather than left to -> inference:** **R-214** (the console never stops showing a stale pairing code), **R-220** (a rebuilt -> box's drives cannot be re-enrolled — *currently worked around BY HAND on the campaign venue*, which -> is the only reason an app could be deployed there at all), **R-221** (a rebuilt box cannot run the -> escrow ceremony), **R-213** (putting files back), **R-202** (the orphan card's unconditional -> promise — now the **last** place on that surface still promising recoverability, two doors from -> where R-228 removed the same promise). - -Eleven faults, each judged on **the message, not the outcome**, each with a positive control proving -the fault was real. Full observables: `tests/campaign11-evidence-2026-08-05/journal-phase24.md`. - -| ID | What | State | -|---|---|---| - -## Instruction files — deferred half, 2026-08-06 - -| ID | What | State | -|---|---|---| -| **R-398** | **`resticStep` is not a seam, so no test can drive any restic-backed path.** Every off-site operation funnels through it, and it shells out — so `RestoreOffboxScratch`, `PlaceOffsiteRestore` and the whole capture side can only be unit-tested up to the point restic would run. Felt directly in v0.226.0: R-358's safety property is an ORDER (clear the marker before restic, write it after), and with no seam that order could only be pinned by an **AST walk** of the function rather than by executing it. That works and is honest about what it proves, but it is a structural workaround for a missing seam and would not catch a reordering introduced through a helper. **Contrast, and it is the argument for the row:** `offboxLatestSnapshot` gained a seam in this same release (`SetOffboxLatestSnapshotFn`) precisely because a correctness gate could not otherwise be proven, and that one took four lines. **⚠ THIS ROW WAS WRONG AND I FILED IT. Corrected rather than deleted, because a row that quietly disappears teaches nobody.** It said 'resticStep is not a seam, so no test can drive any restic-backed path'. **The first half is true and the conclusion is false:** `resticStep` is not itself overridable, but the layer it calls — `offboxRunner`, injected by `SetOffboxRunner` (`offbox.go:52`) — **has been a seam since the off-site tier shipped**, and other tests in that package have been driving restic-backed paths through it all along (`offbox_3a_test.go` uses it five times). I read one function and generalised from it. **And the proposed fix would have been actively worse:** a `resticStepFn` seam REPLACES `resticStep`, which would have hidden its `unlock --remove-all` escalation from exactly the assertions that must observe it — R-359's lock-safety test asserts `unlock` never appears in any argv, and it can only do that because the runner seam sees every command. **What the row asked for that WAS real is done:** R-358's AST ordering test is now an execution test through the existing seam (`TestR358_MarkerOrderingIsExecuted`), which immediately surfaced something the AST walk could not — `unlockStale` legitimately runs before the restore. **Nothing is owed. The row survives as the record that the seam EXISTS, so the next session does not re-file it.** | **CORRECTED, NOT CLOSED — the premise was wrong (2026-08-30)** | — | Add `resticStepFn` beside the existing `offboxFreeFn` / `offboxSizer` / `offboxLatestSnapFn` seams, nil in production, and convert R-358's AST test to an execution test. **Do it as its own change, not folded into a bugfix** — a seam added under pressure is how a test ends up asserting the shape of the thing it was written beside. | CC | -| **R-385** | **A controller was built, baked AND vouched with no CHANGELOG entry of its own, and every gate stayed green.** Controller **0.221.1** shipped on 2026-08-23 while the newest heading in `felhom-controller/CHANGELOG.md` still read `v0.221.0` — the prune-ordering fix (commit `810b18a`) had been written INSIDE the v0.221.0 entry instead of getting its own. The image was never in question; the RECORD was, and the fleet ran a version the record did not name. **`scripts/golden_currency_gate.py` could not catch it by construction:** it failed only on `released > baked`, so a golden AHEAD of the record passed silently. Measured on the real history: `newest released 0.221.0 / newest golden baked 0.221.1 → OK, exit 0`. | **CLOSED — 2026-08-23** | — | **Both halves fixed, both directions red-proofed.** The record: `v0.221.1` has its own heading carrying the MOVED (not duplicated, not deleted) reasoning — commit `da75603`, pushed alone before anything else. The gate now asks *"is the baked version WRITTEN DOWN?"* — the baked version must have its own `## vX.Y.Z` heading **anywhere** in the CHANGELOG. **Membership, not `baked > released`, deliberately:** a comparison against the newest heading alone goes green the moment any later entry is written, leaving the unrecorded version permanently unrecorded and the gate permanently silent about it. INCONCLUSIVE (exit 2) preserved. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/evidence/gate-0*.txt` — old gate/old record `exit 0`, new gate/old record `exit 1`, new gate/fixed record `exit 0`. | CC | -| **R-387** | **The hub REWRITES an unknown severity and says nothing, and the guard built to catch that sits downstream of the rewrite.** One handler, two fields, opposite discipline: an unknown `event_type` is rejected with a loud `400`, while an unknown `severity` was silently coerced to `info` — after which `severityNotifies` drops it and NEITHER delivery leg runs. **Two shipped features went out that way**: `DiskAlertKind.Severity` emitted `"warn"` until controller v0.215.0, `app_start_failed` until v0.223.0. **Measured on the live hub DB 2026-08-23: 91 `app_start_failed` events stored all-time and ZERO `notification_log` rows before that day** — not one, on any channel, while every POST returned 200. **The dispatcher's `unrecognized severity` line could never execute** for an API event, because the coercion one line upstream guarantees the value it looks for cannot arrive. | **CLOSED — hub v0.107.0, 2026-08-23** | — | **The coercion STAYS; only the silence is fixed** — a rejected event is a LOST event, and losing an alarm is worse than mis-routing one. A `WARN` now names the customer, the event type, the rejected value and the consequence. **The dispatcher branch was KEPT, on evidence not caution:** `cmd/hub/main.go` wires `dispatcher.ProcessEvent` DIRECTLY as the `monitor.EventNotifyFunc` for the staleness, host-staleness and offsite-box checkers, which never pass through the handler — for them it is the only severity guard there is; deleting it as "dead" would have removed the live half while the dead half supplied the justification. All 90 severity literals in `internal/monitor` verified already valid. Proven live: `[WARN] [api] Event from demo-hp: severity "warn" is not in {info,warning,error,critical}…`, with an `error` control silent. Evidence: `audits/DRILL-r329-r386-2026-08-23/evidence/live-19-scenarioH-after.txt`. | CC | -| **R-391** | **Gate 11 (observations) is registered in three of the four runners; `app-catalog-felhom.eu` is the exception.** The controller and agent runners already carried a shared-gate mechanism (`SHARED_REUSE`, `SHARED_INSTRUCTIONS` pointing into `felhom.eu/scripts/`), so registering there was one constant and one `GATES` line each. **`catalog_gates.py` has no such mechanism:** its `run_gate` joins every entry against its OWN `scripts/` directory, so it cannot invoke a sibling repo's script at all; and its loop appends `--all` to every gate unconditionally, which the observations gate would read as a path. Registering there therefore needs `run_gate`'s contract widened AND the argument handling changed — a refactor of a runner whose shape is deliberately different (per-app scoping, network/runtime gates excluded from `--fast`), in a repo this task marked out of scope. **The exposure today is nil** — `app-catalog-felhom.eu/REPORT.md` has no observations section, and the gate passes quietly on that — but a future catalog session could write one and nothing would read it. **Filed rather than left as a sentence in a report, which is the exact failure R-389 records.** | **OPEN — LOW** | — | Either give `catalog_gates.py` the `SHARED_*` absolute-path mechanism the other two runners already have and stop appending `--all` to gates that do not take it, or state in that repo's CLAUDE.md that its REPORT.md carries no observations section by convention. **Do not copy the gate script** — the shared checker lives in ONE place (`felhom.eu/scripts/`) and copying it is the drift the shared pattern exists to prevent. | CC | -| **R-390** | **The golden-bake runbook omits `pveam update`, and the failure it produces names the wrong cause.** `documentation/runbooks/RUNBOOK-manual-build.md` §4.1 step 2 says to list the current Debian template because "the exact point release rots" — but on the drill VM's `virgin` snapshot **the `pveam` INDEX is stale too**, so `pveam available` offers an old point release and `pveam download local ` fails with **`400 Parameter verification failed. template: no such template`**. That reads as a typo or a bad argument, not as an old index, and it costs a diagnosis every time. **Hit on two consecutive bakes** (golden 0.222.0 and 0.223.0, both 2026-08-23). The runbook is otherwise correct verbatim — the qemu launch line, the token-read-inside-the-VM pattern and the acceptance markers all worked unchanged. | **OPEN — LOW** | — | Add `pveam update` as its own numbered step before the listing, and say WHY: a snapshot that never changes carries an index that never updates, so the rot warning already in the step applies to the index as well as to the release. Recorded meanwhile in the workspace memory `golden-bake-needs-pveam-update` and in `documentation/tests/golden-0.223.0-2026-08-23/README.md`. | CC | -| **R-392** | **No architecture document covers the two-AI workflow.** `documentation/architecture/` holds eight documents and **all eight cover the product** — topology, host agent, control-plane authorization, hub, off-site connectivity, backup, controller modules, capability map. Nothing records how the Claude.ai / Claude Code split works, what each side owns, how skills and `.claude/rules/` are scoped, or why. **The absence was found by trying to fill the template field, not by a survey:** the task that added the five process skills (2026-08-25) had to name an owning architecture document and could not, and the template requires that be recorded rather than passed over. The exposure today is low — the split is stable and both sides work — but it lives entirely in the operator's head and in chat, which is precisely the shape of a commitment nothing enforces. | **OPEN — LOW** | — | Write one architecture document for the agent-tooling layer: which AI owns which artifact class (`TASK-*.md`, `RUNBOOK-*.md`, validation), how skills are scoped and installed, what belongs in a `CLAUDE.md` versus a skill versus a rules file, and the reasoning for each boundary. **The rules themselves already exist** in `skills/felhom-doc-authoring/SKILL.md`; what is missing is the map of who owns what. Do not restate the doc-authoring rules there — point at that skill. | CC | -| **R-393** | **A decision-log skill for unattended runs was considered and deliberately deferred.** Filed 2026-08-25 by the session that added the five process skills, so the deferral is a decision on the record rather than a thing that was dropped. **The gap it would close:** an overnight or unattended run makes dozens of decisions and the operator can only reconstruct them by reading the whole transcript, which is exactly what nobody does. The proposal is an appended row per decision — what was chosen, why, the evidence pointer, and the result — so a long run is reconstructable in a page. **Why it was NOT built with the other five:** the other five are text files that need nothing but the existing installer glob. This one needs a helper script to append rows and a storage convention for where the log lives and when it is rotated, which makes it an implementation task with its own acceptance criteria, not a skill file. | **OPEN — LOW** | — | Decide the storage convention FIRST — most likely a per-session file beside the session's evidence directory, never `REPORT.md`, which is overwritten every session (the R-341 shape). Then the skill, then the helper. **Check it does not duplicate `felhom-handoff`**, which already owns the end-of-session note; a decision log is the during-the-run half and the two must point at each other rather than overlap. | CC | -| **R-394** | **`felhom-build-deploy/SKILL.md` is 179 lines, over the 150-line limit its own repo now enforces.** Found 2026-08-25 by `scripts/check_skills.py` on its first run — the over-length was discovered BY the new checker, on the day the limit was written down, which is the checker working as intended. **It is not edited and not trimmed here:** the task that introduced the limit explicitly scoped the four pre-existing skills out, and trimming a build-and-deploy skill without exercising its commands is how a wrong command ships to a live host. **It is a named single-entry exception in `GRANDFATHERED` in `scripts/check_skills.py`, printed as a WARN on every run**, so it cannot fade; a NEW skill over the limit is convicted normally, and growing the set requires editing that file in a commit with a row to name. **The rationale for the limit** — attention thins across the excess, so the lines that matter are not the ones that survive — is in `skills/felhom-doc-authoring/SKILL.md` §5. | **OPEN — LOW** | — | Trim `felhom-build-deploy/SKILL.md` under 150 lines in a session that can VERIFY the commands it keeps, then delete its `GRANDFATHERED` entry in the same commit. The likely trim is the per-artifact command blocks moving behind a pointer to the runbooks, keeping the gotchas inline — but that is a judgement for a session with a build to run, not a line-count exercise. | CC | -| **R-388** | **PRODUCT DECISION (not a defect): the customer notification model is the wrong shape, and the settings page grows by one toggle per detector.** The operator's framing, recorded verbatim 2026-08-23: *"A customer should be notified only about things they can act on or are responsible for — the drive they unplugged, the storage they filled. **A failed backup is our incident, not theirs.** The intended shape is that we detect it, we tell them we noticed and are dealing with it, and they are not handed an error they cannot solve. The subscription should feel like being looked after, not like being on call."* Today's page is the opposite shape — one switch per detector, and it **grew from 12 to 15 in a single session** (one new alarm plus two compound toggles split into four). That growth is the argument, not an aside: a page that grows per detector keeps asking a household to make engineering decisions. | **OPEN — DIRECTION, operator's call** | a decision on scope; nothing here is a bug | Recorded as a dated **[DESIGN — DIRECTION]** entry at `documentation/architecture/08-alarm-ladder.md` §8, marked plainly as *not current behaviour*. **Deliberately NOT implemented in the session that recorded it.** `app_start_failed` defaulting OFF is consistent with the direction and reversible either way, but was ruled on its own merits and does not pre-judge the redesign. | Viktor | -| **R-229** | **The instruction-file rightsizing landed for `felhom-controller` and the workspace root; three pieces were deliberately deferred.** Done 2026-08-06: controller split into a 92-effective-line core plus four `paths:`-scoped `.claude/rules/*.md`; workspace root 208→142 effective lines with its versioned copy kept byte-identical; surgical corrections to `felhom-agent` and `felhom.eu` (expired TEMPORARY block, every version literal, the Legacy-Windows copies, the duplicated health-check rule); five contradictions resolved — including a drill-VM claim **measured live** (`qm list` on demo-hp shows VM 300 `drill-r50`; `felhom-agent` was right, `felhom-controller` was wrong); new shared `felhom.eu/scripts/instructions_gate.py` registered in `controller_gates.py` and `agent_gates.py`, 20 fixture tests + red-proof. **Leg (a) CLOSED 2026-08-06 (part 2):** `felhom.eu/CLAUDE.md` **227 → 115 effective lines**, split into a core plus `.claude/rules/{hub,website,manifests,docs}.md`; `instructions_gate` **registered in `scripts/repo_gates.py`** (six gates, all OK) in the required order — trim first, register second, because a registered-but-failing gate refuses every push. Scoping proven from the `InstructionsLoaded` hook log in two fresh sessions, not from frontmatter. **Still deferred:** (b) **CLOSED 2026-08-06 (close-out)** — `felhom-agent/CLAUDE.md` **175 → 99 effective lines** (measured 175, not 173: the CI correction added two), split into a core plus `.claude/rules/{proxmox,localapi,backup,storage}.md` beside the existing `health-checks.md`. The release section now points at the `felhom-build-deploy` skill instead of restating a table that drifts from the script. **Every `CLAUDE.md` in the workspace is now ≤120 effective lines except the workspace root at 142, which is deliberate — it is the only file re-injected after `/compact`.** (c) **CLOSED 2026-08-06 (part 2)** — all 44 orphans resolved with **zero deletions** (file count 158 before and after): 4 durable `reference`-type files indexed, 40 dated episode records moved to `.claude-memory/archive/`. `MEMORY.md` 145 → **150 lines / 17,977 bytes**, and `instructions_gate` check 6 now watches it (over-limit FAILS, orphan WARNS, absent store PASSES *printing its reason*). (d) **The spec-as-failing-test pilot** — moved to R-230. Full accounting: `audits/LEDGER-instruction-trim-2026-08-06.md` + `audits/LEDGER-instruction-trim-part2-2026-08-06.md` | **READY** — owner Viktor | -| **R-230** | **Three instruction/memory follow-ups deliberately left by the part-2 session (2026-08-06), each needing a decision rather than an implementation.** (a) **A ruling is owed on auto-written staleness.** The hand-written `CLAUDE.md` files are now clean of version literals and expired blocks — the gate enforces it — but `MEMORY.md`, which Claude writes and which is the LARGER half of what loads (8.4k tokens vs the root file's 6.6k), carries **21 lines with component version literals**, **5 with bare host addresses**, and an entry still reading *"demo boxes REMOTE till ~08-02"* — the same expired-TEMPORARY class the gate was built to kill, now surviving in the one file the gate's content rules do not cover. **Partly actioned 2026-08-06 (close-out), and the ruling is STILL OWED:** the **three statements that were actively false** were corrected — `R-193 decision open` (closed 2026-08-05), `demo boxes REMOTE till ~08-02` (the box answers on the home LAN), `OPEN R-25b` (shipped 2026-07-21) — and gate check 6 now **WARNs** on version literals, host addresses, expired statements and stale-open citations in the index. WARN, never FAIL: Claude writes that file between sessions, so a hard failure would refuse a human's push over a line no human typed, and the warning is read by the model that will next edit it. **The remaining 32 version literals and 4 host addresses were deliberately left** for that loop. What is still owed is the bulk-correction ruling. **Correcting the premise:** the earlier report's "three expired statements" were all FALSE POSITIVES — each matched an ISO date inside a markdown link target, i.e. a filename — while the one real expired claim carried no ISO date at all. (b) **CLOSED 2026-08-06 (close-out)** — the workspace-root `CLAUDE.md` **is now a relative symlink** to the versioned copy, so the divergence class is gone rather than policed. Check 5 learned two shapes: for a link it asserts the target resolves to a real file (**a dangling link is worse than a diverged copy — the instructions load NOTHING and there is no content left to notice is wrong**), for two files byte-identity as before, so a clone elsewhere is unaffected. **Proven, not assumed:** three fresh sessions logged `session_start` for the link path, and a fourth **with no tools at all** quoted standing rule 1 verbatim — the content reaches the model, not just the path. (c) **The spec-as-failing-test pilot**, approved in principle and not started (was R-229(d)). | **READY** — owner Viktor | -| **R-232** | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **READY** — owner Viktor | -| **R-231** | **`/opt/backup/scripts/` on DooPlex is unversioned host state** — found 2026-08-06 while adding the auto-memory store to the backup set. No repository tracks the scripts that protect the recovery chain, so the edit made that day (`CLAUDE_MEMORY_DIR` in `backup-config.sh`, multi-path restic call in `backup-data.sh`) exists only on the box. This is the same class the part-2 session was closing, found inside the fix for it; the change is transcribed in `felhom.eu/workspace/README.md` so it is at least *recorded*. **Two related facts, both understating current safety:** the backup destination (`/mnt/5_hdd/backup`) is on the **same physical disk** as the workspace it protects, and the DooPlex backup set has **no off-site leg** (`sync-hetzner-backups.sh` is jarrs.eu and pulls *from* Hetzner *to* DooPlex). Bringing a root-owned production backup script under version control, and deciding what installs it, is its own scoped change. | **READY** — owner Viktor | -| **R-233** | **The golden bake's acceptance checks were a list of strings the script does not print** — found 2026-08-06 while baking golden 0.203.0 by following `runbooks/RUNBOOK-manual-build.md` §4.1 verbatim. Two of the three named pass markers **cannot ever match**: `overlay2 OK` is not in `build-golden.sh` at all (the line it means is ` docker OK (overlay2; data-root /var/lib/docker)`), and `including mount point … mp1` refers to a volume that stopped existing in `build-golden.sh` **v3.0.0**, when R-165 collapsed the two data volumes into one. The 404 pre-gate's URL was also wrong — the published filename is `golden.tar.zst`, not `felhom-golden-.tar.zst`, so the pre-gate would 404 for the wrong reason and pass **even when the version already existed**. This is the *"an instrument that can silently drop results is not a measurement"* class landing on the bake's own acceptance check: a grep for an impossible string reads `0` forever, and `0` is indistinguishable from failure. **The bake was never actually unguarded** — the script's own `[ "$drv" = "overlay2" ] || { echo FATAL; exit 1; }` is fail-closed and the run exited 0 with no `FATAL`. The **document** was the broken part, which is why nothing had ever gone wrong and nobody had noticed. **FIXED in the same session:** all markers re-captured from the real log rather than paraphrased, the corrected pre-gate URL, the token handling moved off the command line into an in-VM runner script (the old `--setenv=GITEA_TOKEN=$GT` form put the value where `systemctl show` prints it), a required **positive control** on the token-leak grep, and the vouch step rewritten as the three-field change it actually is. **The general lesson:** a runbook's pass markers must be **copied from a captured log, never written from memory** — §4.0 of that same file already learned this for the qemu launch line and says so; §4.1 had not. | **CLOSED 2026-08-06** — fixed in `RUNBOOK-manual-build.md` | -| **R-235** | **The appliance console keeps telling an already-paired box to go and pair itself.** Measured 2026-08-06 on VM 323: **25 minutes after** the operator bind, with the guest provisioned, the controller reporting 0.203.0 and the agent ONLINE, the physical console still displayed „Felhom — a doboz készen áll, és a **párosításra vár**" together with the now-spent pairing code `US3-6GP` — and, in the same panel, the promise **„Ez a képernyő magától frissül — nincs teendő a doboznál"**. It does not refresh. A customer looking at their screen is told the setup has not happened, and is told the screen would have updated if it had. Cosmetic in mechanism, not in effect: it invites the customer to re-pair a working box, or to call for help about a box that is already fine. Same family as R-234 — a surface asserting a state that stopped being true. | **READY** — owner Viktor | -| **R-240** | **A backup that covered nothing calls itself „Sikeres".** On a configured box with no app selected for off-site backup, a run reports status `ok` with the warning „Sikeres — nincs mentésre jelölt alkalmazás" — *successful* immediately beside *nothing is selected*. Measured as T4 on the final walk, 2026-08-07; flagged once before (2026-08-06) and deliberately not touched then, because the task that noticed it forbade changing that path. **It is the same rhetorical shape the project has spent a fortnight removing** — R-203's *a warning beside a success is read as a success*, R-234's *„✓ Rendben" over an app that was skipped*, R-225's *unknown rendered as zero* — one notch weaker each time, and this is the weakest and last of them. The state itself is honest and must stay `ok`: an unconfigured box reporting `incomplete` forever is its own defect, pinned by a test. **The defect is the word „Sikeres", not the verdict.** Wording such as „Nincs mentésre jelölt alkalmazás — ez a futás semmit nem mentett" says the same thing without congratulating the customer on it. | **READY** — owner Viktor | -| **R-242** | **A controller release that changes customer-visible behaviour is not delivered until a golden carries it — and nothing enforces that.** R-239 is the symptom; this is the mechanism, recorded **2026-08-07** and deliberately **NOT built** (the task that found it scoped it as a record-only item). **Two releases went out without a golden and the gap was invisible until a walk measured it from the customer's side**: v0.204.0 (R-237) and v0.205.0 (R-234) were written, tested, pushed, and CHANGELOG'd, and every one of those steps passed while a machine installed that night received neither. The register said CLOSED; the fleet said otherwise. **Nothing in the release path knows a golden exists.** The version bump, the image push, the CHANGELOG entry and the register closure are all repo-local; the manifest's `golden_version` is edited by a separate operator act, in a different repo, with no link back. **Proposed shapes, cheapest first — the choice is the operator's and is not taken here.** (a) **A release-path checklist step** — one line in the controller's end-of-session checklist: *a release that changes customer-visible behaviour is not finished until a golden carries it or a register row says why not.* Costs nothing, catches nothing mechanically. (b) **A gate in `repo_gates.py`** comparing the manifest's `golden_version` against the newest released controller and FAILING (or warning) past a tolerance of one minor. Mechanical, runs on every push, and would have fired the morning after v0.204.0. (c) **A hub-side checker** — the hub already knows every box's running controller version from `/hosts` and the vouched golden from the manifest; a periodic comparison against the newest published image would catch drift the repo cannot see, including a vouch that was made and then rolled back. **Earliest catch: (b).** It fires on the push that creates the gap, before any box is installed, and it needs no live fleet. **(c) catches strictly more but only after boxes exist.** (a) is worth doing regardless because it is free. **Not built. No gate was written this session.** **⚠ IT RECURRED WITHIN A DAY, WHICH IS THE ARGUMENT FOR BUILDING IT.** Controller **v0.206.0** shipped the R-241 fixes on 2026-08-07 while the vouched golden still carried **0.205.0** — so a machine installed on the morning of 2026-08-08 would have received neither. Third occurrence of the shape in three days (R-111/R-115/R-120 are the older family). **✅ SHAPE (b) BUILT 2026-08-08 — `scripts/golden_currency_gate.py`, registered in `repo_gates.py` as gate 7.** **It was shown FAILING against that exact state before anything was baked**, which is its red-proof and the reason its own introducing push needed `--no-verify` (stated in the session report rather than worked around): `newest released controller : 0.206.0 / newest golden baked : 0.205.0 → CONVICTED`. **IT IS `--fast`, AND THAT FORCED ITS DESIGN:** both `.githooks/pre-push` AND CI run `repo_gates.py --fast`, so a non-fast gate would run in NEITHER — the R-29 census failure this runner exists to end. **THEREFORE IT CHECKS THE BAKE, NOT THE VOUCH**, because the vouched version lives only in the hub's `hub_settings` with no copy in git, and putting a copy there would create a second source of truth that can drift — a green gate over a false claim being the worst outcome available. **A bake without a vouch still passes: that half is NOT closed and stays on this row.** It also compares versions rather than behaviour, so a release changing nothing customer-visible trips it too — accepted deliberately, because judging that by hand is what failed three times and the cost of a false trip is one bake; a waiver belongs here, never in a habit of bypassing. **VOUCHED 2026-08-08 with the operator's approval** — golden `0.206.0` / sha `c85230b4…108e`; `agent_version` and `min_agent` both stayed `0.127.0`, and `wrapper_sha256` was carried through explicitly because the handler clears it when omitted. **The gate was CONVICTED before the bake and OK after it** — red→green on the same command, which is its proof that it measures something real. **⚠ THE GATE FIRED FOR REAL, 2026-08-08 — and it was right.** Controller **v0.207.0** (R-249/R-252/R-253) is released, tested and pushed, and **no golden carries it** — the newest bake is 0.206.0 — so `golden_currency_gate.py` FAILED, saying exactly the true thing: *a machine installed right now would receive v0.206.0*. **The `felhom.eu` push therefore used `git push --no-verify`, declared here, in the commit message and in the session report.** A bypass and NOT a waiver, deliberately: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **Owed: bake golden 0.207.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; the vouch is a three-field change). **This row's own remaining half is unchanged — nothing gates the VOUCH itself.** **✅ THE OWED BAKE IS DONE, SAME DAY — golden 0.207.0 baked, published, round-trip verified and VOUCHED (2026-08-08).** The gate went from red to **green**, and the `--no-verify` bypass declared above is now historical rather than standing. **Round trip is the evidence, not the build log:** the published bytes were downloaded back — 656 879 192 B, sha256 `20ec9602…22995`, both identical to what the bake reported — and **`./etc/felhom-controller-image` read OUT of the downloaded archive says `felhom-controller:0.207.0`**, which is the delivered artifact naming the controller it will start. **The vouch was a three-field change with all three checked deliberately** (`MinAgent 0.127.0` read from the golden's controller CHANGELOG header, not assumed; `agent_version` already ≥ it; `min_agent` not above `agent_version`, so not the R-216 shape) and verified by **re-reading the manifest rather than trusting the flash**. **This row's remaining half is UNCHANGED and is the whole of what is still open: nothing gates the VOUCH itself** — the currency gate's own docstring says it checks the bake, so a baked-but-unvouched golden still passes it silently. Evidence: `tests/golden-0.207.0-2026-08-08/`. **⚠ RED AGAIN, 2026-08-08 (second time in two days) — controller v0.208.0 (R-254) is released and the vouched golden is 0.207.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.207.0 and none of today's fixes. **The `felhom.eu` push used `git push --no-verify`, declared in the commit, the CHANGELOG and the session report** — **a bypass, not a waiver**, on the same reasoning as yesterday: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **Owed: bake golden 0.208.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change, `MinAgent 0.127.0` unchanged). **Note the cadence this is establishing: two releases, two bakes owed within 24 h.** That is the argument for this row's OTHER half — nothing gates the vouch, so the only thing standing between a release and an undelivered fleet is somebody remembering. **⚠ RED AGAIN, 2026-08-30 — controller v0.224.0 (R-330) and v0.225.0 (R-331) are released and the newest golden carries 0.223.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.223.0 and neither of today's fixes. **The `felhom.eu` push used `git push --no-verify`, declared in the commit message, in `hub/CHANGELOG.md` and in `REPORT.md` — a BYPASS, not a waiver**, on the same reasoning as the two 2026-08-08 entries above: the gate offers a waiver only for a release that *deliberately needs no golden*, and these need one. **The operator was asked and ruled bypass-now-bake-later on 2026-08-30**, on the stated ground that neither fix bites a DAY-0 box — R-330 is a nightly false alarm about apps a new box has not installed yet, and R-331 is a hub-side display over backups a new box has not taken yet — and both arrive by self-update afterwards. **That ground is recorded because it is the thing to re-check, not a general licence: the next release that changes first-boot behaviour cannot reuse it.** **OWED: bake a golden carrying 0.225.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change — `golden_version` + `agent_version` + `min_agent`; MinAgent is 0.129.0 per both CHANGELOG headers). **Cadence note, unchanged and now worse: this is the fourth bypass of this gate, and the gap it names is now two releases wide rather than one.** **⚠ WIDENED TO THREE THE SAME DAY — v0.226.0 (R-353/R-357/R-358/R-360) shipped 2026-08-30 and the golden still carries 0.223.0.** The `felhom.eu` push carrying that release's documentation used `git push --no-verify` on the operator's standing ruling from earlier the same day, declared in the commit and in `REPORT.md`. **The day-0 ground still holds for all three and was re-checked rather than assumed:** R-330 alarms about apps a new box has not installed; R-331 is a hub display over backups a new box has not taken; **R-353/357/358/360 are restore-surface fixes, and a day-0 box has nothing to restore.** **The ground expires the moment a release changes first-boot behaviour — that is the thing to re-check, not a licence.** **Owed: ONE bake carrying 0.226.0 covers all three** (`RUNBOOK-manual-build.md` §4.1; three-field vouch, MinAgent 0.129.0), then raise the floor. **✅ PAID THE SAME DAY — golden `0.226.1` baked, published, round-trip verified, VOUCHED, and the floor RAISED (2026-08-30).** Evidence: `documentation/tests/golden-0.226.1-2026-08-30/`. `golden_currency_gate.py` went **red → green** on the same command, which is its proof that it measures something real. **The three declared bypasses above are now HISTORICAL rather than standing.** **The round trip is the evidence, not the build log:** the published bytes were downloaded back — **657 197 592 B, sha256 `70ed8e93…baefe69`**, both identical to what the bake reported — and `./etc/felhom-controller-image` read **out of the downloaded archive** says `felhom-controller:0.226.1`, i.e. the delivered artifact naming the controller it will start. **The three-field vouch was checked deliberately, not assumed:** `MinAgent 0.129.0` read from the golden's controller CHANGELOG header, `agent_version 0.130.0 ≥ min_agent 0.129.0` (so NOT the R-216 shape), and the result verified by **re-reading the manifest** rather than trusting the flash — golden option `0.226.1 SELECTED`, all four shas matching. **The floor is proven ACTING, not merely set:** `demo-felhom` self-updated within 30 s, logging `[selfupdate] Post-update startup: update successful (0.225.0 → 0.226.1)`. **⚠ AND IT HAPPENED AGAIN THE SAME DAY, AND WAS PAID AGAIN.** v0.227.0/v0.227.1 (R-359/R-397) shipped after the 0.226.1 bake, the gate convicted a fifth time, that `felhom.eu` push used `--no-verify` and declared it, and golden **0.227.1** was baked, published, round-trip verified, **VOUCHED** and the floor **RAISED to 0.227.1** within the hour. Evidence: `documentation/tests/golden-0.227.1-2026-08-30/`. **THE CADENCE IS NOW MEASURED RATHER THAN ASSERTED: five convictions and two full bakes in one day.** Every bypass was declared and every debt was paid — but the pattern this row exists to name is exactly that a release and its delivery are separate acts, performed hours apart, by whoever remembers. **The floor was proven ACTING both times:** `demo-felhom` self-updated 0.225.0→0.226.1, then 0.226.1→0.227.1 — and the second time it also registered the new `offsite-integrity` job **by itself, on a box nobody deployed to**, which is the strongest evidence this row has ever carried that a floor delivers rather than merely records. **This row's OTHER half is still open and untouched: nothing gates the VOUCH itself** — the currency gate's own docstring says it checks the bake, so a baked-but-unvouched golden still passes it silently. **2026-08-31, the SEVENTH debt and it was paid the same day — twice in one day.** v0.230.0 shipped in the morning with the newest golden at 0.229.0, **which is the build R-403 says deletes a good copy**, so the gate was red across four commits (`dddcc80`, `6e550ae`, `130f7a6`, `32a4c35`). Golden **0.230.0** baked, published, round-trip verified, vouched, and the fleet floor raised 0.229.0 → 0.230.0; `demo-felhom` moved itself off the defective build unattended (`controller-swap: new controller healthy`, 16:21:40 CEST). Evidence: `documentation/tests/golden-0.230.0-2026-08-31/`. **The gate did its job and its own weakness surfaced doing it — R-410.** | **READY — the vouch half only** — owner Viktor **⚠ SIXTH CONVICTION, 2026-08-31 — controller v0.229.0 (R-102/R-103) is released and the newest golden carries 0.228.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.228.0 and neither of today's fixes. The `felhom.eu` push carrying this release's documentation used `git push --no-verify`, declared in the commit message and in `felhom-controller/REPORT.md` - **a BYPASS, not a waiver**, on the same reasoning as the five entries above: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **The day-0 ground was RE-CHECKED rather than reused:** R-102 and R-103 are restore-surface changes on the Tier-2 card, and a day-0 box has taken no Tier-2 copy and has nothing to restore from one; no first-boot behaviour changed, and `MinAgent` is unchanged at 0.129.0. **The ground still expires the moment a release changes first-boot behaviour.** **OWED: bake a golden carrying 0.229.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change - `golden_version` + `agent_version` + `min_agent`, MinAgent 0.129.0), then raise the floor. Fleet floor and golden are 0.228.0 today. **Golden and fleet delivery are the operator's (this row).** **✅ PAID THE SAME DAY — golden `0.229.0` baked, published, round-trip verified, VOUCHED, and the floor RAISED (2026-08-31).** Evidence: `documentation/tests/golden-0.229.0-2026-08-31/`. `golden_currency_gate.py` went **red to green** on the same command, which is its proof that it measures something real. **The `--no-verify` bypass declared above is now HISTORICAL rather than standing.** **The round trip is the evidence, not the build log:** the published bytes were downloaded back - **656 864 331 B, sha256 `39aa886d…d7bdae87`**, both identical to what the bake reported - and `./etc/felhom-controller-image` read **out of the downloaded archive** says `felhom-controller:0.229.0`. **A THIRD independent reader agreed before anything was vouched:** the hub's own Day-0 dropdown read the same sha straight from Gitea, a different code path from the round trip. **The three-field vouch was checked deliberately, not assumed** (`MinAgent 0.129.0` read from the golden's controller CHANGELOG header; `agent_version 0.130.0` >= `min_agent 0.129.0`, so NOT the R-216 shape; `agent_sha256` and `wrapper_sha256` carried through explicitly because the handler clears a field it is not sent), and verified by **re-reading the manifest** rather than trusting the flash. The **R-120 gate passed rather than being bypassed** - fleet newest 0.229.0, golden 0.229.0. **The floor is proven ACTING:** `demo-felhom` self-updated `0.228.0 -> 0.229.0` and logged `settle-gate: GO - at/above floor 0.229.0 (we are 0.229.0)` - **nobody deployed to that box.** **Cadence note: this is the SECOND bake in one day (0.228.0 then 0.229.0) and the sixth conviction, and both debts were paid within the hour.** This row's OTHER half is still open and untouched: **nothing gates the VOUCH itself** - the currency gate checks the bake, so a baked-but-unvouched golden still passes it silently. **⚠ SEVENTH CONVICTION, 2026-08-31 — controller v0.230.0 (R-403) is released and the newest golden carries 0.229.0. AND THIS ONE IS NOT LIKE THE OTHERS: the day-0 ground does NOT apply and must not be reused.** Every previous bypass rested on 'a day-0 box has nothing to restore / nothing to alarm about yet'. R-403 is a defect in the NIGHTLY TIER-2 COPY, which a day-0 box starts running on its first night: a machine installed on 0.229.0 can have a complete recovery package on its second drive replaced by an empty one, and that is measured, not suspected (120 082 104 B -> 7 036 B on demo-hp). **The row's own standing sentence - 'the ground expires the moment a release changes first-boot behaviour' - is what expires it here.** The `felhom.eu` push carrying this release's documentation used `git push --no-verify`, declared in the commit message and in `felhom-controller/REPORT.md` - a BYPASS, not a waiver. **OWED, and more urgent than the previous six: bake a golden carrying 0.230.0, vouch it (three fields, MinAgent 0.129.0 unchanged), and raise the floor.** `demo-hp` was updated by hand; `demo-felhom` is still on 0.229.0 and still carries the defect. **See also R-404, filed today: this is the seventh bypass and the habit is now the thing being reported.** **2026-09-01 (R-404): THE BAKE HALF IS UNCHANGED AND THE VOUCH HALF IS STILL OPEN.** R-404 moved WHO the bake check refuses and added a notice in the controller repo; it did NOT touch what is checked. **Nothing gates the VOUCH.** A baked-but-unvouched golden still passes both the gate and the new notice, and the reason is unchanged and forced: the vouched version lives only in the hub's `hub_settings` table, there is no copy in git, and a hub-reading gate could not be `--fast` so it would run in neither the hook nor CI. **Do not read R-404's closure as closing this.** **2026-09-13 — NARROWED: the WAIVER half is BUILT (R-468).** The docstring's *"honest fix is a recorded waiver in the register, never a habit of bypassing"* is now a mechanism: `golden_currency_gate.py` reads `documentation/tests/golden-waiver.yml` (dated, ≤ 14 days, row-bound), turns a BEHIND conviction into a loud advisory while valid, and is red again when it expires — the difference from this row's original rule, which recurred the next day, is that a dated waiver cannot be forgotten. It never covers an UNRECORDED golden (R-385). Operator ruling the same day: goldens weekly and before any install, not per release. **What stays open on THIS row is exactly one thing: nothing gates the VOUCH.** The waiver does not touch it, and the reason it is unbuilt is unchanged (the vouched version lives only in the hub). | -| **R-243** | **A box in the R-241 state silently stops backing up off-site, and NO ALARM OF ANY KIND FIRES.** Found by the R-241 spike (2026-08-07) as a by-product; **not part of the walk's finding and not previously filed.** The R-241 state is self-locking in a second, worse way than the recovery-journey dead end: `escrow_state` is stuck `pending` forever (the auto-confirm flips only on a hash match, and the hash cannot match a key the box minted itself), and `runOffboxBackup` returns at the escrow gate (`offbox.go:743`) before touching anything. **So off-site backups never run again — and the hub never notices.** All three signals that could catch it are excluded, each for its own individually-correct reason, verified in the hub this session: `offsite_stale` — `isStale` (`monitor/offsite.go:135`) returns false unless `EscrowState == "escrowed"`, and its own comment reads *"Pending/disabled = normal onboarding, never stale"*, so the box is classified as **still being set up, forever**; `offsite_delivery_stuck` — `monitor/offsite_delivery.go:91` skips the `applied` shape, and delivery genuinely IS applied (the credential was consumed and the target is in every report); `backup_failed` — never fires, because nothing fails: the run returns `nil` before it starts. **Three correct exclusions leaving one state unobserved.** This is the same class as the workspace `CLAUDE.md` "presence is not success" rule, one level up: **the absence of a failure is being read as the presence of a working tier.** **Partly subsumed by R-241's fix** — a box that recovers leaves this state — but **not for a box that does not**, and the alarm gap is what makes "does not" survivable indefinitely. **Not fixed; no code written.** **⚠ UPDATED 2026-08-07 (v0.206.0) — the STATE this row describes can no longer be entered, but the ALARM GAP is untouched and the row stays open.** R-241's mint guard means a box no longer mints a key over a sealed package, so it no longer arrives in the "escrow stuck pending against a self-minted key" state by itself. **What replaces it is a state that is VISIBLE rather than silent:** the box declares `offsite.state=awaiting_recovery_key` and the customer is offered the recovery screen. **But the hub still raises nothing for it**, and for the same three reasons: `isStale` needs `escrowed`, the delivery checker skips the `applied` shape, and `backup_failed` needs a run that never happens. **So a box whose customer never acts still stops backing up off-site with no operator signal** — the difference is that the customer can now see it and act, where before nobody could. **The remaining work is an operator-side signal for a box held in `awaiting_recovery_key` past some age**, and it is deliberately not bundled into R-241's fix. **⚠ MEASURED ON A REBUILD, 2026-08-07 (fifth walk) — the gap is real for the state this row describes, and NOT for the state a rebuild produces.** 88 seconds after the walk5 guest was destroyed and rebuilt, the hub emitted `offsite_delivery_stuck` (**warning**) and wrote an **operator-channel** `notification_log` row recording `offsite_credential_restaged` / status **REFUSED** with an accurate reason — *"the credential was applied and worked; the target was lost afterwards … a guest rebuild does, R-193"*. So on the **regressed-apply** shape the operator IS told, promptly and correctly, and this row's *"skips the applied shape"* does not apply. The gap stands for a box that reaches the held state **without** a prior working tier in its report history. **Recorded so the row is not read wider than it measures.** | **READY** — owner Viktor | -| **R-244** | **The customer DELETE cascade leaves `app_log_issues` behind, and it is systematic across every venue ever torn down.** Found **2026-08-07** while verifying the `finalwalk` teardown with a **full census** (every table, every column) rather than a per-table query. After a cascade that logged `COMPLETE … full teardown`, **61 rows still matched `finalwalk`**. Four of the five sources are **deliberate and correct** — the cascade's own header states *"Provenance/events are NEVER wiped — audit outlives every tier"*: `events` 16, `notification_log` 14, `host_deletions` 1, `customer_resets` 1. **The fifth is a gap:** `app_log_issues` 29 rows, which the residue purge does not touch (its logged leg covers `reports`/`app_telemetry`/`app_log_tails`/`log_tail_requests`/`notif_prefs`/`selfbind_tokens`/`appliance_registrations` — not this table). **It is not a `finalwalk` quirk:** rows still reference **`c11` 40, `rewalk` 20, `part4` 24** — all three torn down 2026-08-06, whose ledger recorded *"0 occurrences"*. **That prior claim was measured with a narrower query and does not survive a full census; the correction is recorded rather than the measurement quietly redone.** **Why it was probably never written, established rather than assumed:** the table is a **fleet-wide aggregate** keyed on `app_name`+`fingerprint` with an `affected_customers` JSON list — of the 29 `finalwalk` rows, **12 reference only `finalwalk`** (orphans, safely deletable) and **17 are shared with LIVE customers** (`demo-felhom`, `peti-felhom`, …) and **must not be deleted, only de-referenced.** A naive `DELETE … WHERE customer LIKE` would destroy a live customer's issue history — which is very likely why the leg does not exist, and is the reason this is not a one-line fix. **Severity is LOW and stated plainly: no secret material is involved** — app name, fingerprint, message text, counts, timestamps. What survives is a deleted customer's *identifier* inside an aggregate row. **Proposed shape:** a residue leg that (a) removes the customer id from `affected_customers`/`context_customer`, and (b) deletes rows whose `affected_customers` becomes empty; plus a one-off sweep for the four already-torn-down venues. **The general lesson is the reusable part:** *a per-table absence query is not a census.* The teardown verification is now a full-schema sweep, and that is what found this. **Not fixed** — a cascade change needs its own red-proof and this session was scoped as a spike plus two operations. Evidence: `tests/teardown-finalwalk-2026-08-07.md`. **⚠ STILL OWED, AND NOW MEASURED RATHER THAN ESTIMATED (2026-08-08 census, read-only, no truncation).** `app_log_issues` holds **1309 rows**; **71 reference a torn-down venue** (`finalwalk`, `c11`, `rewalk`, `part4`); of those **44 are ORPHANS** — they name only torn-down customers and are safely deletable — and **27 are SHARED with a live customer** (`demo-felhom`, `peti-felhom`, …) and **must be de-referenced, never deleted**. 1238 rows are untouched. **The 27 are exactly why the leg was never written**, and why a `DELETE … WHERE customer LIKE` would destroy a live customer's issue history. **What it needs, precisely:** a cascade leg that (a) removes the customer id from `affected_customers` / `context_customer`, and (b) deletes only rows whose `affected_customers` becomes empty; plus a one-off sweep for the four venues already gone. **Why it was NOT done on 2026-08-08:** the fix is hub code, and that session's scope forbade a hub version bump; a hand-run SQL mutation over 71 rows — 27 of them needing surgical de-referencing — with no tested code path and no red-proof is precisely the shape that goes wrong on a live database. **It accumulates one venue at a time, so the next walk adds to it**; the numbers above mean the next session starts from data rather than a guess. **⚠ IT GREW AGAIN, AS PREDICTED — walk5 teardown, 2026-08-08.** The fifth walk's venue was torn down with a full-schema census taken **before and after**: **168 rows → 67**. Of the 67, **37 are by design** (`events` 21, `notification_log` 14, `host_deletions` 1, `customer_resets` 1) and **30 are `app_log_issues`** — this row's gap, and the count was **predicted in the pre-run enumeration rather than discovered afterwards**, which is the difference from the ledger that once recorded *"0 occurrences"* from a narrower query. **The running total across torn-down venues therefore rises from 71 to ~101 rows** (`finalwalk`, `c11`, `rewalk`, `part4`, now `walk5`) — the shared-with-a-live-customer subset must still be **de-referenced, never deleted**. **It accumulates one venue at a time and it did so again.** Evidence: `tests/walk5-r201-2026-08-07/teardown-walk5-2026-08-08.md`. **2026-09-25:** `peti-felhom` is no longer a live customer (deleted through the cascade, journal #20); 8 `app_log_issues` rows still name it — the same gap. | **READY** — owner Viktor | -| **R-245** | **Should a customer who never decides be auto-abandoned after 30 days? RECORDED, NOT BUILT — and the reasoning against it is recorded with it so the decision can be revisited properly.** **The operator's proposal (2026-08-07):** a box that has been offered recovery for 30 days without the customer deciding is auto-abandoned, entering the 14-day grace, so an undecided box does not sit for ever holding history nobody has claimed. **What was built instead:** the escalating reminders (1/3/7/14 days) and the operator levers `--abandon-extend` / `--abandon-stop`. **The reasoning, as settled with the operator the same day:** (1) **nobody is absent** — a box does not reinstall itself, so whoever rebuilt it was standing there and met the recovery question; the "customer away for months" case does not arise from this situation, because **a reinstall implies a person**. (2) **A customer who cannot find their code will get in touch**, which is the moment to extend or disarm by hand — so the automation would be firing at people we are already talking to, which is why the levers were the thing worth building. (3) **The cost is theirs**: the old history sits in the customer's own storage allowance, and if they are paying to keep something they have not decided about, that is their call and they feel it before we do. (4) **The real harm, if it comes, is QUOTA** — old history blocking new backups — and **that is a condition, not a calendar**. An automatic ending should trigger on the harm, with a dated warning, never on a date alone. **If this is ever built, build it that way.** **✅ RE-FILED 2026-08-08 AS A DECISION TAKEN, not a question pending.** It sat in the operator's queue as `WAITING-ON-OPERATOR` for a day, and **nothing was actually pending** — the operator and the reviewer settled it on 2026-08-07: it is **not built**, the levers were built instead, and the whole reasoning above is the record of why. A settled decision parked in a queue is a queue nobody trusts, and an audit of every `WAITING-ON-OPERATOR` row the same day found this was the ONLY one — so the drift was caught while it was still a single row. **THE CONDITION THAT REOPENS IT, which the reasoning already names: QUOTA — old set-aside history blocking new backups.** Not a calendar. If a customer's retained history ever refuses a new backup, revisit this with a dated warning that triggers on the refusal; until then it stays decided. | **DECIDED 2026-08-07 — not built; reopens on quota** | -| **R-246** | **A leftover staleness flag has silently disabled the new recovery discriminator on `demo-hp` since 2026-08-04, and the flag is WRONG.** Found by a read-only spike, 2026-08-08. **Q1 — traced to an act, to the second:** at `2026-08-04 20:15:49` the hub emitted `offsite_reissued` and `escrow_stale` in the same second — an operator **Re-issue** pressed during the R-201 drill, three minutes after `escrow_blob_served` at 20:12:40/20:12:54. That was `offsite.ReissueCredentials`'s **precautionary** `MarkEscrowStale` call, which **hub v0.95.0 REMOVED the very next day** (R-196 / R-204 item 2) precisely because it marked healthy escrows stale. **Q2 — the flag is wrong, measured on both sides:** the hub's blob seals `restic_pw_sha256 = 8a9e33aa4da6769c…d080a`, and the key the box is actually using hashes to **the identical value**. The blob covers the key. **Q3 — nothing clears it by itself:** the ONLY writer of `stale_at = NULL` is `SaveHostEscrow`'s `ON CONFLICT` — i.e. a fresh escrow ceremony, **which is the one act that would supersede the good blob**. So the only exit from a false alarm is the destructive act the false alarm recommends. **Q5 — the blast radius, enumerated rather than assumed:** (1) `GetEscrowStatusForCustomer` withholds `restic_pw_sha256` from the ACK; (2) a **pending** box can never auto-confirm, so (3) **every off-site run is refused indefinitely** — neither bites `demo-hp`, which was already `escrowed` and is backing up healthily (12 snapshots, last success 2026-08-07T02:15:35Z); (4) the customer is told to create a new code; (5) **NEW — v0.206.0's shape (c) is inert**, because the box records an empty hub hash and falls back to (a)/(b), so the recovery screen would stay silent even if recovery were needed. **Q6 — a fresh box CANNOT reach this state:** `MarkEscrowStale` has **no production caller anywhere in the tree** (census: only its own definition, two comments and two test references). **The next walk cannot meet it.** **NOT CLEARED, deliberately** — Q2's answer says the fix is to clear this instance, but whether to also stop the column being settable at all is a separate ruling, and the spike was scoped read-only. **✅ THE FLAG IS CLEARED — operator-approved and applied 2026-08-08.** One row, identity-matched on `host_id` and guarded on `stale_at IS NOT NULL`; `changes()` returned **1**. **Verified end to end, not just in the database:** the hub now serves the hash again, the box recorded `hub_escrow_key_sha256 = 8a9e33aa4da6769c…d080a` at `11:10:19Z`, and that is **byte-identical to the key it is using** — so shape (c) compares, matches, and correctly stays silent. **The false stale warning is gone, proven with a positive control** rather than an absent line: **0** `escrow-confirm` lines since the restart while **5** scheduler lines in the same window prove the box was logging, and the recorded hash proves an ACK was processed. *(Method note: the hub pod is Alpine with no `sqlite3`; it was installed into the container's ephemeral writable layer — image and node untouched, gone on restart. SQLite's own file locking coordinated the write with the live hub; an earlier attempt failed cleanly on quoting and changed nothing, which is the fail-safe working.)* **STILL OPEN under this ID: the ruling on whether `stale_at` keeps a live setter at all.** It currently has NO production caller, so the column is write-only-by-accident — a field that changes behaviour, that nothing sets and nothing can see (R-248). Either give it an evidential setter or retire it; **do not leave it as a trap that only a database read can spring.** | **READY — flag cleared; the column ruling is still owed** — owner Viktor | -| **R-248** | **A flag that changes behaviour is visible to nobody who would look for it.** Q4 of the 2026-08-08 spike, answered plainly. **The customer** sees only a derived card stating a false reason (R-247). **The box** cannot see it at all (R-247's dropped field). **The operator** can see it on exactly ONE page — the **PBS-DR** view (`hub/internal/web/pbsdr.go:487`, `v.EscrowStale = escrow.StaleAt != ""`) — which is the wrong tier for this symptom: an operator investigating an OFF-SITE problem has no reason to open a PBS-DR page. **No alert, no report field, no off-site surface.** The one-shot `escrow_stale` event fired on 2026-08-04 and **was never notified** (a full census of `notification_log` for that customer that day returns 8 rows, none of them this one); it has fired twice ever, both on 4 August. **So the practical answer is: only a database read.** That is a finding in its own right — a flag that silently changes behaviour and cannot be seen is the shape this fortnight has been about (R-241's discarded comparison, R-228's unread field, R-243's unobserved state). **What it needs:** surface `stale_at` on the off-site operator surface and in the host report, or stop using a field nobody can observe to change what a customer is told. | **READY** — owner Viktor | -| **R-250** | **A customer create can fail fail-closed because the host-key scan ladder is shorter than the DNS/AAAA settle time.** Found 2026-08-07 creating the fifth walk's venue. `POST /configs/new` with off-site enabled provisions the Storage Box sub-account and then scans its SSH host key to pin it — **fail-closed by design** (`offsite.go:111-121`, *"don't serve a descriptor the controller can't verify"*), with `defaultScanBackoff` = 2+4+8+16+30 ≈ **60 s**, sized by its own comment to *"the observed DNS propagation lag"*. **Measured, both halves:** the first create exhausted the ladder — five `no such host`, then `dial tcp [2a01:4f8:bacc:2:200::d30]:23: connect: network is unreachable` (the name had just begun resolving, **AAAA-first, into a pod with no IPv6 route**) — and the hub logged `[ERROR] offsite provision for walk5`. An **identical second POST succeeded** ~70 s later on its own final rung (`shared already provisioned … (subaccount 285351)`). **Total settle ≈ 100 s against a 60 s budget.** The endpoint was never the problem: verified afterwards from both the node and the hub pod, by name, `SSH-2.0-OpenSSH_9.6p1`. **Two distinct things are wrong and should not be merged:** (1) the budget is sized against DNS *existence*, but what actually bit is the **AAAA-before-A** window, a different and longer phenomenon — lengthen the ladder and/or prefer the A record for the scan dial; (2) **the operator is told nothing actionable** — the create fails with a generic error, the remedy is "press it again", and nothing says so. The retry is genuinely safe (idempotent on the label; confirmed afterwards — `one_time_secrets` 1, one sub-account, no double-provision) **but that safety is invisible to the person deciding whether pressing again will double-charge them.** **Severity LOW-MEDIUM:** self-clearing, no data at risk — but it is the first thing a new customer's provisioning does and it fails looking like an outage. | **READY** — owner Viktor | -| **R-251** | **The recovery listing renders one row per restic TAG, so the customer is shown an "app" they never installed and their data counted twice.** Measured on the fifth walk, 2026-08-07, on the screen the customer reaches after entering R. The snapshot carries tags `felhom-offbox,calibre-web`; the listing renders **two rows** — `calibre-web · 2026-08-07 14:57 · 12.8 MB` and `felhom-offbox · 2026-08-07 14:57 · 12.8 MB`. `felhom-offbox` is the tier's own marker tag, not an application. **The screen's whole job is to let the customer check that what is in the store is what they expect** (*"Nézd át, hogy tényleg azt találod-e itt, amire számítasz"*), and it shows them a stranger's name beside their own data and a total that is double the truth. **Cosmetic, not a data defect** — the restore page correctly offers only `calibre-web`. **Fix:** filter the marker tag out of the listing, or key the rows on the app tag. | **READY** — owner Viktor | -| **R-254** | **The same render-then-hide pattern R-249 fixed is live in two more places, and one of them carries a real per-install secret.** Found by the §7.1 census that R-249's fix required — *a pattern found once is worth a census*, and it was. **(1) `app_info.html:185` — the serious one.** An app's auto-generated first-login password is rendered into `` beside a „Megjelenítés" button. `hidden` is the same class of control as R-249's `display:none`: it stops a browser drawing the value and leaves it in the response body, so a fetch of the app page returns it. **The value is a REAL per-install secret** — `ReadInitialCredentials` reads it live out of the deployed container (`internal/stacks/initialcreds.go`), it is not the catalog's published `default_creds`. **(2) `deploy.html:482` — the weaker one.** An auto-generated `type: secret` deploy field renders into a `` with a „Megjelenítés" toggle. On the PRE-deploy form this is close to unavoidable — the form must post the value, and it does, in a sibling hidden input — but on an **already-deployed** app's page (`$isDeployed`) the hidden input is correctly omitted while the readonly input still carries the value, and there the exposure is gratuitous. **Not fixed here, deliberately:** this session's scope was R-249/R-252/R-253, and each of these needs its own reveal endpoint and its own body-asserting test rather than a shared quick edit. **The fix shape already exists** — `POST /settings/retrieval-password/reveal` (v0.207.0) and `escrow_handlers.go`'s rule that a secret is revealed by an XHR and never templated server-side into HTML. **Severity MEDIUM for (1)** — a real credential in a page any logged-in customer opens, with no audit event, reaching caches, history and screen-shares; **LOW for (2)**. **Recommended next**, because R-249 proved the pattern is not theoretical: it was found by the value landing in a session transcript. | **✅ BOTH SITES CLOSED — controller v0.208.0, 2026-08-08, and they turned out to be two different problems.** | -| **R-255** | **The check that would catch a fourth secret-in-the-body covers 4 of 27 pages, and the cheap gate that covers all 36 templates is blind to the shape that actually shipped.** Filed 2026-08-08 while closing R-254, **because a partial guard reported as complete is worse than no guard — it stops the next person looking.** **Two nets, both measured.** **(1) `scripts/secret_in_markup_gate.py`** reads all 36 templates and convicts any `{{ … }}` naming a secret unless allowlisted with a reason. It catches `{{.RetrievalPassword}}` and `{{.InitialCreds.Password}}`, **and it catches a launder through a local variable** because the assignment itself names the secret (`{{$v := .InitialCreds.Password}}` is convicted — verified). **It is blind to a secret arriving under a NEUTRAL PAGE-DATA KEY** — `data["Tagline"] = creds.Password` then `{{.AppInfo.Tagline}}` passes it cleanly, also verified. **That is exactly the shape of R-254 site two** (`value="{{$val}}"` inside an `{{if eq .Type "secret"}}` branch), so the gate **would not have caught one of the three instances it was written for.** **(2) The runtime body assertion** — render the page and grep the response for a sentinel — catches every shape, including that one (demonstrated on the same planted leak the gate missed). But it needs each page's data to be constructible in a test, and **only 4 of 27 page templates have that today**: `settings_security`, `app_info`, `deploy`, `backups_restore` — the four that were touched by R-249/R-252/R-253/R-254 and therefore got their own tests. **The other 23 pages have no runtime coverage at all.** **What closing this needs, so the cost is not re-estimated:** a per-page data fixture for the remaining 23 (most need a wired `Server` — `stackMgr`, `backupMgr`, agent seams), then one table-driven test that renders each with a sentinel substituted for every string in its data and asserts the sentinel is absent. **That is real scaffolding, which is why it was NOT built inside R-254's session** rather than half-built and declared done. | **READY** — owner Viktor | - -**SITE ONE — the same defect, fixed the same way.** `app_info.html` no longer renders the value; the page carries the username, the note and a boolean, and the password comes from **`POST /apps//initial-credentials/reveal`**, which **re-reads the running container** rather than serving a cached copy (caching it in the handler would have put it straight back into the body one layer in). `no-store`, CSRF-covered, and **logged as an act**. Both controls — Megjelenítés *and* Másolás — go through it, and a reveal that cannot read the value **says why** instead of returning an empty string that would render as a blank password. - -**SITE TWO — NOT the defect the row described, and the difference is the finding.** The hidden input is **deliberate and was left alone**: it fires only on the PRE-DEPLOY form, and `README §318` documents why the value must round-trip — the customer is shown the generated secrets so they can note them down, and submitting them back is what makes the saved value the same one they saw (*"no silent re-generation on submit"*). **A form must carry what it submits.** What *was* indefensible is the neighbouring **readonly display input**: on an ALREADY-DEPLOYED app the hidden input is correctly omitted — nothing is being submitted — yet the secret was still rendered into a page the customer merely opens. Fixed by **`POST /stacks//auto-field/reveal`**, authorised by requiring the field to be a `type: secret` auto-generated field *of that stack's catalog metadata* — that check is what stops it becoming "read me any value out of any app". Both directions are pinned: the deployed page must not carry the value, **and the pre-deploy form must still submit it**. - -**⚠ THE PREMISE THAT THIS BROKE A REPO RULE DOES NOT HOLD, and it is recorded rather than quietly dropped.** The rule cited was *"Password fields require explicit user input or generation (no silent auto-fill)"*. **No such line exists anywhere in the repo.** What exists is `CONTEXT.md:2070` — *"Password fields require explicit input | Prevents accidental empty-password deployments"* — which is about **emptiness**, not auto-fill, and which the hidden input does not contradict. - -**§7.3 — HOW MUCH WAS ACTUALLY EXPOSED, measured on the fleet rather than assumed.** **Site one: nothing.** `crafty-controller` is the ONLY catalog app declaring `initial_credentials`, and it is **deployed nowhere** — the card renders only when `found.Deployed && found.Meta.InitialCreds != nil`, so that code path has never run in production. **Site two: nothing measurable either.** 26 catalog apps declare a generated `type: secret` field, but `demo-hp` has exactly **three** apps deployed — `calibre-web`, `opengist`, `privatebin` — and **none of the three declares one**. **THE HONEST LIMIT: this is a CURRENT-STATE measurement.** An app deployed and later removed would not appear in it, and **nothing anywhere recorded a read** — which is itself part of the defect being fixed. So: **no evidence of exposure, and no mechanism that could have produced evidence either way.** **Rotation is therefore not indicated by anything measured** — the decision is the operator's, and this note is the input to it. - -**Gate:** `scripts/secret_in_markup_gate.py`, registered in `controller_gates.py`. **Its blind spot is measured and in its docstring** — see R-255. | **CLOSED 2026-08-08** | - -**Recorded against existing rows by Phase 2:** - -- **R-216 — §4.1 is now MEASURED, not deduced.** The previous session could only offer two absences. - The box's own `/settings` renders „Minimális verzió (üzemeltető) **0.200.0**" (`GetFloor()`, whose - only writer is the report-ACK handler; **both hold branches serve `Floor=""`**, pinned by - `managed_floor_test.go:94`), and a **cold-started** controller logs - `settle-gate: GO — at/above floor 0.200.0 (we are 0.201.0)` against the same line reading - `floor still unknown after 1m30s` while the hold was in force. The hub's HELD lines ran every - 15 min to 22:12:06 and stopped, with a liveness control proving the hub kept logging. **The floor - is served.** -- **⚠ A correction to how that positive was to be taken.** `SetFloor`'s line is `u.dbg(...)`, gated on - `cfg.Logging.Level == "debug"` and written to the **logger** — it can **never** reach the logx debug - ring, so it cannot appear in `/api/debug/logs` at any level. A controller restart alone would not - have produced it. Confirmed with a level census on the ring first (1196 DEBUG / 2802 INFO / 2 WARN), - so the absence was known to be structural rather than evidential. -- **R-218 — the live half is STILL NOT MEASURED, deliberately.** The fix is present and correct - (`needsOffsiteCredential` now retires on the **target**, not the key), but the venue **has** a - target, so the box correctly does not declare; declaring here would be the bug. The state that - exercises it is shape (a), which the venue no longer holds. **Recorded as not measured rather than - inferred from the unit test.** -- **R-217 — its fix HELD under exactly its fault** (F5): with the store blocked after a successful - unlock, the page rendered M3 and **no listing block at all**; the three false-claim strings are - absent, verified in UTF-8 with accented positive controls present. -- **R-215 — its fix is present** and the `GET /recovery` gate consults the same predicate as the POST - sibling. -- **R-199's back-pointer in `architecture/00-capability-map.md` was already added** — the brief lists - it as owed; it is present on the escrow-recovery row, explicitly labelled as the omitted - back-pointer. **No action taken; the brief's assumption was stale.** - -**Untouched by this session, stated so nothing is presumed closed by association:** **R-213** (putting -files back — the half the recovery screen deliberately does not do) and **R-202** (the orphan card's -unconditional promise, which CAMPAIGN-11 §7 step 7 measured the customer-facing cost of). - +## Open rows | ID | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---| -| **R-213** | **Putting files back in place — the half the recovery screen deliberately does not do.** The screen (R-193, controller v0.200.0) unlocks the repository and LISTS what is in it; restoring is per-app and lives in the backups area, and the operator ruled the two separate on 2026-08-05: *a screen that unlocks and then offers to overwrite is two decisions wearing one button*. What is missing is the step after the listing — a customer who can now SEE their files still has to work out, per app, which restore to choose. **The operator named its requirement: a live-versus-backup comparison** — the customer must be able to see what would change before anything is overwritten | **OPEN — not started, deliberately** | the comparison design (nothing exists for it yet) | Design the live-vs-backup comparison, then the put-back flow on top of it. Do NOT fold it into the recovery screen | Operator + CC | -| **R-88a** | ~~Failing backup re-quiesces every 5 min, no backoff~~ | **SHIPPED** (controller v0.176.0, 2026-07-27) | — | Live on both boxes; breaker 15m→4h, per-tier, never permanent | — | -| **R-88b** | ~~`/backup/due` cannot say *unknown*~~ | **SHIPPED + PROVEN-LIVE** (agent v0.105.0 + controller v0.178.0, 2026-07-27) | — | `age_state=unknown` captured on real hardware during a deliberate ep0 outage; controller deferred, **zero app stacks stopped** | — | -| **E-2d** | **Prove E-2 on a fresh VM** — a real `felhom-host-install.sh` 1.22.0 run, Case B naturally, a claimable customer, then add a drive (the offer) and unplug it (`backup_target_absent` end-to-end) | **CLOSED — PARTIALLY PROVEN** (2026-07-29) | — | **C1, C2 proven** (`audits/E2D-fresh-vm-2026-07-29.md`); **C3, C4 proven live** (`audits/SESSION-C-2026-07-29.md`); **C5 FAILED → R-116** — the gate fires and an alarm reaches the hub, but it is the generic event, so the alarm and its recovery cannot be paired. **R-116 is the single named open leg**; per the Session-C runbook §9, decided in advance, a failed claim closes the item as partially proven rather than triggering a re-run. Both audits carry the full record — the `local-lvm` fence, the ISO/PAIRING derivation, the Phase 0 answers, the per-claim observables and the teardown evidence — and are the place to read it, not this cell. **The arc's actual definition of done is R-106 + R-109, R-108 and D5**, none of which this detour touched | CC | -| **R-121** | **A BOX's installed agent can sit releases behind the vouched one and nothing notices — the R-120 gate does not cover it.** demo-hp ran agent **0.113.0** while the hub vouched **0.116.0**, through the whole R-116/R-117 arc, and no signal existed on any channel | **READY (S) — NEW 2026-07-30** | — | **Fourth instance of the drift family** (R-111 golden's agent 17 releases behind, R-115 built+deployed but never published, R-120 golden a controller behind — and now installed-vs-vouched on a live box). **Confirmed at source that R-120's gate cannot catch it:** `hub/internal/web/configs.go:1165-1169` compares `goldenVer` against `store.NewestReportedControllerVersion()` — it is a **golden-artifact vs fleet-CONTROLLER** check and says nothing about the agent installed on a box. **`MinAgent` does not cover it either:** it is used to HOLD the controller floor for a box whose agent is too old (`hub/internal/api/handler.go:530-538`, `store.go:1857`) — protective, not an alarm — and demo-hp's 0.113.0 **equalled** `min_agent` 0.113.0, so even a floor comparison was satisfied. **The cost, measured:** R-117's whole subject is the R-113 conjunction, which landed in **0.114.0** — so the designated drill host could not exercise the code under investigation at all, and the R-117 spike had to route every predicate result through an out-of-repo probe built from `main` instead of the installed agent (`audits/SPIKE-r117-bind-liveness-2026-07-30.md` §1, §2.3). Discovered because the R-117 task made bringing the box current an explicit prerequisite. **Fix shape (not implemented):** the hub already receives `AgentVersion` on every host report, and already has semver comparison in Go — the missing piece is a checker comparing reported agent vs the **vouched** agent and surfacing it, operator-tier. Note the honest tension: a box legitimately lags between publish and deploy, so this wants a staleness window rather than an instant alarm | CC | -| **R-118** | **An absent drive's union row advertises the ROOT filesystem's capacity as its own.** In the absent-state payload the registry-union row reports `total_bytes: 49675956224 / used_bytes: 4584579072` — **byte-identical to the `local` row** (`durable_id: path:/var/lib/vz`, i.e. `pve-root`) in the same response. The real drive is **4 GB** | **READY (XS) — NEW 2026-07-30** | — | Cause: `statfsCapacity(d.MountPath)` (`disks.go:335-338`) statfs's `/mnt/cel`, which with the device gone is a **bare directory on the root filesystem**. `observe.go:176-183`'s comment warns about exactly this trap and guards the Observe path ("*an unmounted removable dir-storage's mountpoint reverts to a bare directory on root … catastrophic DR mis-id*"); **the union path has no equivalent guard.** **Not a DR mis-id** — `durable_id` on that row is still the correct `uuid:…`, so re-attach identity is safe. It is a **false capacity** reaching every consumer of `total_bytes`/`used_fraction` (fill monitors, storage cards): a detached 4 GB drive advertises 46 GiB at 9.2 % used. Same class as `role.go:180-181` — an absent drive's fields decaying to the root filesystem's. Evidence: `audits/DIAG-r116-disks-payload-2026-07-30.md` §12 | CC | -| **R-95** | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` **-- UNBLOCKED 2026-09-22, not closed.** The premise in this row's own title - *SFTP cannot express append-only* - is answered by Hetzner (ticket #2026090103040671, recorded in R-433 and R-436): on a Storage Box the `rclone serve restic --stdio --append-only` backend can be pinned to an SSH key as a FORCED COMMAND, so append-only IS expressible without a new machine. **The row stays open because nothing has been measured**: no forced-command key exists on `u629488`, no `forget --prune` has been refused through one, and the protection only holds if EVERY key that can reach the repository carries the prefix. The next step is R-436's, and it is small. | CC **SPIKE 2026-09-01 — `audits/SPIKE-r95-offsite-delete-2026-09-01.md`. THE WORD "ARMED" ABOVE IS NOT SUPPORTED AND IS WITHDRAWN PENDING R-429:** no `.snapshots` is visible to either box's sub-account (measured, both machines, with controls), so the seven-day bound is unverified and unverifiable from the product side. **Q3 (documented, `hub/internal/hetznerapi/hetznerapi.go:38-45`): the sub-account API has ONE permission axis, `readonly` — there is no append-only, so the PBS shape does NOT transfer** (PBS is a server that can refuse; a Storage Box is a filesystem that runs nothing). **Q5 (measured): withdrawing delete does NOT wedge the store** — restic treats a dead owner's lock as stale and proceeds — so the constraint everyone feared is not the blocker; but `unlock --remove-all` lies about success (R-430), and the crash-lock window is UNKNOWN. **Q6 (measured): restic 0.14.0 DOES speak `rest:`** (control: `banana:` → invalid backend), and append-only is a rest-server flag, not a restic one — reachable, but it needs a machine in the recovery path and ep0 is protected. **Q7 (measured): detection is nearly free** — `snapshot_count` already reaches the hub and the hub APPENDS reports, so the comparison needs no box change. **RECOMMENDATION: answer R-429 first (Viktor, ten minutes), then build detection, then move retention off the box; defer the transport change.** Two `forget --prune` sites must be disarmed together — `offbox.go:1388` AND `offbox.go:1759` — or R-191 repeats. **RE-SCOPED 2026-09-01 — THE STORY WAS WORSE THAN THE TRUTH FOR TWO MONTHS.** The box can delete its own LIVE repository, **but it cannot write to the daily snapshots of it** — MEASURED, not cited: a write into `/.zfs/snapshot` is refused on both boxes while the same write to the account home succeeds (R-432). Seven daily snapshots are confirmed in the panel (R-429). **So a deletion costs at most the data written since the last daily snapshot, and the rest is recoverable — file by file, one customer at a time, with no effect on anyone else** (vendor: *"You can download individual files or entire directories as usual"*; *"It is not possible to write to the `/.zfs` directory or its subfolder"*). **NOT open-ended loss.** Two caveats kept honest: a panel-driven snapshot restore rolls back the WHOLE Storage Box and deletes newer snapshots, which is why the per-file route matters; and per-file recovery is **operator-only today** (R-432). **DETECTION SHIPPED hub v0.111.0 (R-431)** — an unexplained fall is noticed within a day. **THE RANKING IS VIKTOR'S:** this has been #1 since July on the old story. On the new facts I would rank it below the items that can still lose data outright, but I am not re-ranking it myself. `audits/SPIKE-r95-offsite-delete-2026-09-01.md` **DRILL 2026-09-01, LATER THE SAME DAY — THE RE-SCOPE'S SECOND HALF IS WITHDRAWN.** The drill that was to walk the recovery found there is no route to walk: **no snapshot is reachable from a sub-account by ANY name** (R-433) — 777,600 exact names in the vendor format over nine days, zero hits, with a passing control, plus the structural reason (`/home` st_dev 0,82 vs `/.zfs/snapshot` st_dev 0,276, and `/home/.zfs` absent). **Clause (a) stands: the box can delete the live repo and cannot write into the snapshot area.** Clause (b) — *"the rest is recoverable file by file"* — is NOT SUPPORTED. **The drill was STOPPED before its destructive phase on the operator's ruling**, because with no recovery route the deletion would have destroyed real history to buy only an alarm test that could not fire at the specified size (R-435). Nothing was deleted; the store is verified untouched at 69 snapshots. **So the comfort that lowered this row rested on an unwalked route, and the route does not exist.** Two new leads decide what happens next: **R-433** (can the MAIN account see them? nobody here holds that credential) and **R-436** (`rclone serve restic --stdio` is offered server-side and restic speaks `rclone:` — measured — which could make real prevention cheap, IF the provider pins `--append-only`). **The rank stays Viktor's. Plainly: the argument that moved this row down is the argument the drill removed.** `audits/evidence-drill-r95-recovery-2026-09-01/` **BLOCKED-ON-PROVIDER 2026-09-01, and the record of the demotion is kept deliberately: this row spent ONE DAY demoted on a clause that did not hold.** It was re-scoped down on the morning of 2026-09-01 on the strength of *"recoverable file by file"*, and that clause was measured false the same afternoon (R-433). **On today's evidence it belongs back near the top — that is a proposal, not an action; CC has not re-ranked it and will not.** Both questions that can settle it are drafted in `documentation/runbooks/provider-questions-2026-09-01.md`: **Q1** decides how urgent this is, **Q2** (R-436) could remove the root cause cheaply. **Precondition on any build here: R-430**, which is harmless only while the box can still delete. | -| **R-191** | **Every weekly offsite backup UPLOADS successfully and then FAILS the job on a prune the box is deliberately not allowed to do — on both demo boxes.** Measured on demo-felhom 2026-08-04 06:49–06:53: the upload completed (223 s, 629 MiB of 1.874 GiB, **67.2 % reused incrementally**), then `ERROR: prune 'ct/9201': proxmox-backup-client failed: Error: permission check failed - missing Datastore.Modify\|Datastore.Prune on /datastore/felhom-offsite/demo-felhom` → `ERROR: Backup of VM 9201 failed - error pruning backups` → `TASK ERROR: job errors`. The hub raised `whole_guest_backup_failed` | **CLOSED — SHIPPED 2026-08-04** (installer **1.25.0**; both live boxes corrected) | — | **This is R-89's rule not reaching the config.** R-89 moved PBS pruning SERVER-SIDE — *"boxes set `keep_last: 0`, ep0 runs prune jobs; box tokens stay write-only, never widen the grant"*. The token behaves exactly as designed: it refuses. But **both** demo boxes still arm the offsite tier with `keep_last=2 prune_pbs_allowed=true` (`backup_targets: [{target_id: felhom-pbs, cadence_seconds: 604800, keep_last: 2}]`), so every run asks for a prune that must fail. **The data is SAFE and that is why this is not a P1:** the snapshot lands before the prune is attempted; what is wrong is the job's VERDICT and the weekly operator e-mail it produces. **But it is corrosive in the specific way this project keeps finding:** a backup that reports FAILED while succeeding trains the operator to discount `whole_guest_backup_failed`, which is the same alert that would carry a real one — and it is exactly the failure the R-100 corollary warns about, an alarm whose text is true and whose trigger is not the thing you would act on. **Fix is one config line per box** (`keep_last: 0` on the PBS tier) plus whatever writes it on a fresh install; **deliberately NOT applied in this session** — the session was a runbook with an explicit "change nothing, and if a change appears necessary, stop and report" rule, and a retention field on a live backup tier is not a change to slip into an observation run. **Check before fixing:** whether ep0's prune jobs actually cover these two namespaces, or the snapshots simply accumulate once the box stops asking **THE GATE WAS RUN FIRST, AND IT MATTERED.** Before disabling anything, ep0 was read (read-only, Tier 2): prune jobs `prune-demo-felhom` and `prune-demo-hp` exist on datastore `felhom-offsite`, one per namespace, `schedule 03:30`, `keep-last 2`, comment *"R-82 retention keep-last=2, server-side (box tokens are write-only)"* — and they have run **every day since 2026-07-27: 18 tasks, all `status=OK`**. The newest task log reads `retention options: --ns demo-felhom --max-depth 0 --keep-last 2` / `keep ct/9201/2026-07-27…` / `keep ct/9201/2026-07-28…` / `TASK OK`. Retention happens, and it happens there. **A METHODOLOGICAL WARNING WORTH MORE THAN THE FIX.** Three separate queries said the OPPOSITE — *no prune jobs have ever run* — and **all three were broken instruments**: `worker-type` where the field is `worker_type`; the value `prune` where the worker type is `prunejob`; and `journalctl -u proxmox-backup` where the unit is `proxmox-backup-proxy`. A fourth reading (3 snapshots under keep-last 2) was mis-framed by CC and self-corrected — the third snapshot had landed AFTER that day's 03:30 window. Acting on any of them would have disabled the only pruning ATTEMPT while reporting that nothing prunes: a weekly false alarm traded for unbounded growth on the protected endpoint, invisible for months. **The gate is what caught it, and only because it demanded evidence rather than a verdict.** **Shipped:** installer **1.25.0** writes `keep_last: 0` on the offsite tier (the agent's existing guard `allowPBSPrune = !primary && keep_last > 0` already reads that as *never prune from the box* — no agent change), the justifying paragraph is rewritten to say where retention lives and cite R-89, and `hostinstall_gates.py` asserts it (red-proved: pinning `keep_last: 2` back fails the gate). **Both live boxes corrected in their own config** — `backup tier armed target=felhom-pbs … keep_last=0 … prune_pbs_allowed=false` on demo-felhom and demo-hp, with the local tier untouched at `keep_last=3`. Served over HTTPS at `1.25.0` with `"keep_last":0` in the served bytes. **STILL TO OBSERVE:** the next weekly offsite run completing OK end-to-end. The change removes the failing step; the *schedule* proving it is next week's event, and this row should carry that line when it happens. | CC | -| **R-194** | **PVE's permission cache delays every grant-state verdict by an unknown amount, so "the agent can read it" and "the ACL exists" are not the same measurement.** Observed twice while validating R-190's self-repair on demo-felhom 2026-08-04: both ACL rows for `/storage/felhom-backup` were deleted, and `GET /access/permissions` continued to report `Datastore.AllocateSpace` present — for **~40 s** in one run and **~16 minutes** in another. During that window the capability probe reads healthy and the self-repair does not fire | **OPEN** | — | **Why it matters beyond the delay:** it puts a floor under how fast a lost grant can be noticed, it makes any single permission read a lagging indicator, and — the interesting part — **it is a candidate contributor to R-190's own timeline**: a grant removed at an unknown moment could keep working until a cache expiry, which is exactly the shape of *worked at 04:44, refused at 09:24*. That does not explain what removed it, but it may explain when the refusal SURFACED, and the two have been treated as the same instant. **Not a defect in our code** — it is PVE behaviour, and the mitigation already tolerates it (the repair fires on the next probe after the cache clears). **What is worth deciding:** whether the store-grant probe should ALSO consult the storage content listing as a second signal, since that appeared to reflect the loss immediately (`{"data":[]}` while the permission read still said present) — two signals disagreeing is itself information, and today only one of them is read | CC | -| **R-200** | **The DR password-injection seam has a handler, a route and tests — and no form.** `POST /backup/offbox/inject-password` is routed (`controller/internal/web/server.go:510`) to `offboxInjectPasswordHandler` (`offbox_handlers.go:174-196`) → `InjectOffboxPassword` (`backup/offbox.go:541`). **No template in the repository contains that path or any form posting to it** (grep over `internal/web/templates/`: one unrelated hit, an XSS comment) | **PLUMBING COMPLETE** (controller **v0.196.0**); **the FORM is not built — still open** | — | **The tenth instance of this project's built-but-never-wired class, and the exact shape `CLAUDE.md` and `felhom.eu/CLAUDE.md`'s seam-wiring rule were written for:** handler tests that POST directly (`offbox_escrow_test.go:167,180`) prove nothing about reachability. To use the only implemented recovery seam today, a person must hand-craft an authenticated POST with a session cookie and CSRF token. **Note the layering while fixing it:** this form takes a **64-hex repo password**, not a recovery code (`offboxRepoPwPattern`, `offbox.go:543`) — they are different secrets at different layers, and the operator's 2026-08-04 ruling asks for a form that takes **R**. Build the R form and treat this one as the operator/DR fallback it was written as, but **ship it with a render test per branch of whatever gate it sits behind**. Source: `audits/RECON-offsite-dr-chain-2026-08-04.md` §3 link 9 **THE DIAGNOSTIC HALF IS DONE AND IT ANSWERED THE QUESTION.** `--recover-offsite-check` is a `docker exec` escape hatch in the shape of `--print-reset-code`: R on **STDIN** (never argv, never `ps`, never shell history, never a transcript), fetch+unseal via the agent, and a verdict of **two sha256 hashes**. **It compares and never installs** — the recovered password is not written to `offbox/repo_password`; a test asserts the data dir is byte-unchanged and its red-proof (adding the install call) fails it. Confirmed live: `repo_password` mtime still `2026-08-03 07:18:02` after the successful check at `2026-08-04 11:49`. **Exit codes are load-bearing** — `0` match, `2` a clean MISMATCH, `1` a step failed; "it failed" and "it worked and disagreed" must never share a status because only one is a finding about the system. A box with no local password reports distinctly (the rebuilt-box shape, where the next step is to INSTALL rather than compare). **WHAT IS NOT BUILT, deliberately:** no card, no form, no preview, no customer-facing text — building an interface on top of a chain nobody had walked is how the preceding three weeks went wrong. **What remains for this row:** link 9 (the recovered password placed so `WriteOffboxSecrets` keeps it) and the customer-facing shape the operator ruled on 2026-08-04 (yell → R form → preview → proceed), which is now priced against a chain that exists rather than one that is assumed **PART 0 SHIPPED 2026-08-04 (v0.196.0):** `--recover-offsite-install` is the sibling of the check — same fetch/unseal path, same STDIN discipline for R — and it **places** the recovered password via `InjectOffboxPassword`. The confirmation is a SECOND invocation (`--confirm-install`): without it, both hashes print and nothing is written, so the operator sees the comparison before a write is possible. Three outcomes named distinctly: **installed** (no local password — the rebuilt-box shape), **unchanged** (identical key, nothing written), **refused** (a DIFFERENT key present — installing would clobber the key the current repository is encrypted under; exit 2, no force offered). It re-reads the file after writing rather than trusting the call. **Red-proof observed:** removing the confirmation gate makes the dry run write the password. The R-persistence test carries a **positive control** (a planted copy found, then removed and not found). **NOT YET EXERCISED AGAINST A LIVE RECOVERY** — the R-201 drill halted before step 9, so this is unit-proven only. **What remains for this row:** the customer-facing shape the operator ruled on 2026-08-04 (yell → recovery-code form → preview → proceed) | CC | -| **R-201** | **Nothing in the offsite DR chain has ever been exercised past the ceremony — and the one live proof that exists predates the field it is cited for.** `slice10d-identity-restore-spike-findings.md` §1 (2026-06-10) proved a real wrap→unwrap of an identity bundle with a real R on a secret-less box, byte-identical, wrong-R failing closed. **That bundle was `{tunnel_token, pbs_token}`.** `ResticRepoPassword` was added in agent **v0.77.0 on 2026-07-09** — a month later — and is **unit-proven only** | **BOTH HALVES PASSED — DATA 2026-08-04, JOURNEY 2026-08-07** — the customer file came back byte-identical (four times now), and on the fifth walk the customer's own journey completed with **zero guest command lines**. **It does NOT claim:** that the journey is smooth (R-252/R-253 are two unsignposted stops on it), nor that the fingerprint discriminator's POSITIVE half is proven (the mint guard fired, so there was no divergent key for shape (c) to catch — only its negative half was measured) | — | **Never exercised, in these words:** a blob has never been served to a box; a fork-4 bundle has never been unsealed with a real R outside a unit test (the only production caller is `--selftest=identity-consume`, `cmd/felhom-agent/main.go:2845`, reading R from `FELHOM_RECOVERY_CODE`); a recovered repo password has never been injected; an existing offsite repository has never been reopened with one; **no restore of any kind has ever been performed from a recovered secret.** The one live consume ever prepared (S5 Part 4-B, 2026-07-04) is still marked *"DEFERRED, not run"*. **`_recovery-inventory-2026-07-28.md` already recorded this accurately** (A.2.7, C.1 row 1, C-2) — this row exists so the gap has an owner and a closing condition, not only a description. **Closing condition = the drill** designed in `audits/RECON-offsite-dr-chain-2026-08-04.md` §10, whose pass condition is a **byte-identical sha256 of a sentinel file restored after a wipe** — explicitly NOT "the repository opened". **Do not run it before R-198 is fixed**: the drill would walk a chain that is missing a link everyone believed was there **MATERIALLY ADVANCED 2026-08-04, AND THE DISTINCTION MATTERS.** What was proven on hardware is that the offsite repository **password** comes back out of the sealed bundle byte-identical (R-199). What has STILL never happened: a recovered password **installed**, an existing repository **reopened** under one, and a **file restored** from it. **The drill's pass condition is unchanged and is not this** — it is a byte-identical sha256 of a sentinel file after a wipe and reinstall, and "the store opened" was explicitly ruled insufficient. Design: `audits/RECON-offsite-dr-chain-2026-08-04.md` §10. **The drill is now cheaper and better-founded than when it was designed:** links 6–8 are walked, so a failure during the drill can be localised instead of being an undifferentiated "recovery did not work", and R-198 means the ceremony that creates the drill's recovery code no longer destroys the key it is meant to protect **THE DRILL RAN AND DID NOT REACH ITS VERDICT, and stopping was the correct call** (`audits/DRILL-r201-offsite-recovery-2026-08-04.md`). Steps 1–4 completed; the wipe never happened; **nothing irreversible was done**. It halted because **the sentinel file was not in the off-site snapshot** (R-203) — wiping would have destroyed the only copy and proven nothing. **Sentinel sha256 `643166269103a25c…`, still on the box.** **What the attempt established live, all of it new:** (1) a rebuilt box's off-site run **refuses** with the orphan card and pushes `offbox_repo_orphaned` — the spike's predicted third outcome, measured for the first time, and it does NOT silently start a fresh history (this closes R-193's Q3); (2) the **orphan reset works** — move-aside to `/home/felhom-repo.orphaned-20260804`, never delete, fresh repo initialised, `offbox_repo_reset` pushed (operator-authorised during the session); (3) demo-hp's pre-rebuild history is **permanently unrecoverable** — its key sits in superseded row id 3 with `identity_blob` NULL, superseded 07:15:36, **four hours before v0.93.0 fixed the retention**; (4) **a precondition the runbook did not contain:** neither pre-existing off-site-toggled app has a restorable file leg — both are named-volume-only, which the off-site tier tars but the customer restore never unpacks — so a file-leg app (`calibre-web`, mandatory `userdata: media/books`) had to be deployed, and it is now in place as the fixture. **TO RESUME:** fix or scope R-203, re-run steps 4–5 and confirm the sentinel IS in the snapshot **by listing it, not by a green status**, then P4 + the STOP + steps 6–11. Everything else is already staged: versions, recovery code, working repository, file-leg app, sentinel. **The pass condition is unchanged** — a byte-identical sentinel sha256, not "the store opened" **UNBLOCKED 2026-08-04 by R-203 (controller v0.197.0).** The sentinel now lands in the off-site snapshot and is listed there by name and size — the thing whose absence halted the drill. **Everything the resumed run needs is already in place on demo-hp:** agent v0.125.0, controller v0.197.0, the recovery code held by the operator (`R_DEMO-HP`), a working off-site repository (3 snapshots), `calibre-web` deployed with a mandatory userdata path, and the sentinel at sha256 `643166269103a25c…` — **verified byte-identical after the R-203 migration moved it to the corrected directory**. **What remains is exactly steps 4–11 of the drill**: re-verify the snapshot listing, take the deliberate rollback archive (P4), the §7 STOP, then wipe, reinstall, recover, install, restore and compare. **The pass condition is unchanged** — a byte-identical sentinel sha256, not "the store opened" **NIGHT RUN 2026-08-04 (`audits/DRILL-r201-night-run-2026-08-04.md`). THE HEADLINE: after a real rebuild — controller data volume destroyed, sentinel deleted from disk — the customer's recovery code produced `8a9e33aa4da6769c5aea1831f87759e10930e2ec1dea0062576484e0598d080a`, BYTE-IDENTICAL to the pre-wipe on-disk key and to the hub's independent record. The off-site backup key is recoverable after a machine is rebuilt, and that had never been shown.** **Step 7's assertion PASSED:** `identity_blob` 572 B and `restic_pw_sha256` unchanged across the wipe (`updated_at` still 11:11:37) — nothing re-escrowed itself. **Step 9a passed:** the key installed cleanly on a bare box (the "installed" branch's first real run); **9b:** the apply kept it. **THE VERDICT WAS NOT REACHED** — step 10 never ran, so there is no post-restore sha256 and no snapshot count. **That is not a FAIL** (nothing came back wrong and no fresh history was started); it is a wall, and the wall is **R-204**. **The wipe was faithful to the incident, deliberately:** the 2026-08-03 rebuild R-193 is filed against was NOT a guest reprovision — the journal shows guest 9201 running continuously with no `pct destroy`/`pct restore`/`--selftest=provision` — so a controller-data-volume wipe reproduces it, and an unrehearsed provisioning chain improvised unattended is what §8.10 exists to prevent. **A precondition had drifted and was repaired, not worked around:** the staged snapshot had lost the sentinel because the afternoon's experiment produced a later same-day snapshot and `forget --keep-daily 7 --group-by host,tags` pruned the good one — **a good snapshot is not durable against a later bad run on the same day.** **TO FINISH (~5 min, operator present):** re-claim the box, confirm the escrow (**NOT a new ceremony** — it would supersede the identity blob and destroy the key under test), run a backup, restore the sentinel and compare to `643166269103a25c…`. Rollback available: the verified archive `vzdump-lxc-9201-2026_08_04-21_44_24.tar.zst` **THE DRILL PASSED (`audits/DRILL-r201-night-run-2026-08-04.md`).** demo-hp's controller data volume was destroyed and the sentinel deleted from disk; the customer's recovery code recovered the key (`8a9e33aa4da6…`, byte-identical to the pre-wipe on-disk key AND to the hub's independent record); it installed on the bare box; the **existing repository OPENED** (`repo_state: null`, **3 snapshots — NOT 1**, 42 026 B = the pre-wipe size exactly); and the customer restore flow returned **`643166269103a25cf41d34a26b75fd6ebba0a837bbcd7e0d2b5aba7106bcbe7c` — byte-identical to the pre-wipe sentinel.** **The Felhom backup story is proved end to end for the first time.** **Step 7's assertion held:** `identity_blob` 572 B and `restic_pw_sha256` unchanged across the wipe — nothing re-escrowed itself, and no ceremony was run at any point (superseded rows still 2). **The wipe was faithful to the incident:** the 2026-08-03 rebuild R-193 is filed against was a controller-DATA-VOLUME loss, not a guest reprovision — the journal shows guest 9201 up throughout with no `pct destroy`/`pct restore`/`--selftest=provision`. **IT TOOK FOUR UNDOCUMENTED STEPS (→ R-204):** an operator Re-issue; a re-claim (whose escape hatch needs a controller restart to work at all); a manual escrow confirm; and `mode=full` on the restore, because the default `mode=unit` returns the app definition and NOT the customer's files. **A customer hitting this alone today would not get their data back.** Box left healthy, claimed, re-armed, sentinel restored to its live location. **⚠ WALKED TO COMPLETION 2026-08-06/07 — DATA PASS, JOURNEY FAIL, and the row stays open.** The full walk ran overnight on a new venue: built from the published ISO, fixtured, soaked through a real scheduled cycle, destroyed, rebuilt, and finished in the morning with the operator's emailed code. **DATA: PASS** — all three sentinels byte-identical out of snapshot `f5c53b03`, accented filename bytes included. **JOURNEY: FAIL** — the customer had **no route at all** to enter the code they hold: the recovery screen had retired itself, and the remote page offered to CREATE a new code instead. Root cause **R-241** (the credential self-heal writes a fresh repository key and moves the box out of the recovery-offer's pristine case, while orphan detection is unreachable behind `escrow_state: pending`). Recovery needed three guest command lines. **Also established:** the credential chain runs end to end unaided on an unclaimed box (first live sighting of its success line), and **R-239** — a fresh install lands on controller 0.203.0 while 0.205.0 is released, so R-234 and R-237 are undelivered. Full record: `tests/finalwalk-r201-2026-08-07/journal.md` **✅ CLOSED 2026-08-07 — BOTH HALVES PASS, on the fifth walk** (`documentation/tests/walk5-r201-2026-08-07/journal.md`). A fresh appliance was installed from the published ISO on `demo-hp` (VM 325), given three sentinels, escrowed, backed up off-site, then **destroyed on purpose** — guest and data volumes — and rebuilt through the documented day-0 path. **THE DATA: PASS** — all three sentinels byte-identical out of snapshot `5b0f20f7`, including the accented filename's *bytes*, read back with `os.listdir` on a bytes path so no decode round trip could launder a `U+FFFD`. **THE JOURNEY: PASS — zero guest command lines were needed to progress**, against three on the previous walk; the reset-code hatch was used once, in Phase A only, where §3 permits it. **RTO 71.7 s** from login to an open store (12.44 s of it the unseal itself; ~22 s a harness retry of mine). **The recovery screen appeared without being sought** (`/` → `/launcher` → `/recovery`) and answered all three questions, with a sealed-at timestamp matching `host_escrow.created_at` exactly. **What made the difference is R-241's mint guard, exercised live for the first time:** at 14:58:52Z the rebuilt box collected its re-staged credential, configured the transport and **refused to mint a repository password** over the sealed package — where the previous walk minted one and lost the journey silently. **Two obstacles remain and are filed rather than absorbed into this close — R-252 and R-253** — neither needing a shell, both cleared from the dashboard, and neither signposted; the journey succeeds and is not yet smooth. **This close does NOT claim** that shape (c) fired positively (it did not — with no local key the offer comes from shape (a); shape (c) was measured in its negative half in Phase A), nor that a customer would clear R-252/R-253 unaided. | **CLOSED 2026-08-07** | -| **R-202** | **NOW EVIDENCED, NOT ARGUED (2026-08-10).** demo-felhom’s orphan card told the customer that the set-aside backups may be restorable later *with their corresponding recovery code*. **For those 1.2 GB that is false and unfixable:** they were written under `48741892f0ef4d59…`, whose `identity_blob` is NULL, and the restic password exists nowhere else. The card promises a route that does not exist, to exactly the customers most likely to read it — the ones who have just lost their history. **The orphan card promises the customer their old backups "may later be restorable with the matching recovery code" — and after v0.93.0 that is true for supersessions from now on and FALSE for anything already orphaned.** `controller/internal/web/templates/backups_remote.html:66,69` states it unconditionally, in Hungarian, on the one surface where being wrong costs most | **OPEN — Part 5 hit its gate 2026-08-04; the card is UNTOUCHED and the sentence is still live** | R-199/R-201 (which generation an orphaned repo belongs to is not knowable to the box today) | — | **THE GATE, and why it was hit rather than squeezed past.** The condition was: ship it iff the hub can tell a box what it needs with **one** additional boolean on the escrow ACK it already sends. The hub *can* cheaply compute *"≥1 retained blob for this host carries an identity blob"* — one correlated predicate in `GetEscrowStatusForCustomer`, and the controller even has the right seam already (`SetEscrowStale`/`StaleBlob` is exactly this shape). **But that boolean does not answer the card's question.** The card renders on `RepoState == "orphaned"`, and the promise is about *the key THIS orphaned repository was written under*. A box does not know which escrow generation the orphaned remote belongs to; a box with a pre-v0.93.0 orphan and a post-v0.93.0 supersession would read the boolean TRUE and the promise would still be false — a **conditional** falsehood that looks verified, which is strictly worse on a customer-facing card than today's hedged one. **What would actually make it truthful** is knowing the orphaned repo's generation, which is the same knowledge R-199/R-201's unassembled chain needs. **Interim exposure, stated rather than buried:** the sentence remains live and remains false for both demo boxes. **Cheapest honest interim** (not taken here — it is a customer-copy change and the gate said leave it alone): drop the recoverability clause and say only that the old history is set aside and not deleted, which is true unconditionally | CC + operator | -| — | Storage Box **snapshots** on `storage-box-pool-1` — plan SET (daily 00:00, keep 7) but **0 taken yet** | WATCHING | first run tonight 00:00 | Confirm `size_snapshots > 0` tomorrow; until then the mitigation is armed, not proven | CC | -| — | `PBS-storage-1` (u629193, box 611421) still `status=active`, 19.9 MB | WAITING-ON-OPERATOR | operator console | Delete the box | operator | -| **R-91** | Old 13 GB datastore copy at `/srv/pbs-felhom` on ep0's root disk | WATCHING | demo-felhom's first **post-migration** PBS backup | Delete once it lands; fix `CONTEXT.md:1018` same commit | CC | -| — | First-ever **GC** on `felhom-offsite` (armed today 13:11 UTC, never run) | WATCHING | schedule | **Sun 2026-08-02 04:30 UTC** — confirm it completes | CC | -| — | demo-felhom's next weekly PBS backup (newest is 2026-07-26) | WATCHING | schedule | ~2026-08-02; also releases R-91 | CC | -| — | demo-felhom's next restore-test (84 h cadence, last 2026-07-27 06:38 UTC) | WATCHING | schedule | ~2026-07-30 18:38 UTC | CC | -| **F-CRIT-2** | ~~A failed offsite backup left a phantom snapshot (1 B, manifest-less, NEWEST) that RESET the tier's freshness clock — 7 days silent on the real 168h cadence, invisible to both the R-88 breaker and the hub deadline monitor~~ | **SHIPPED + PROVEN-LIVE** (agent v0.106.0, 2026-07-28) | — | `NewestArchiveTime` now counts only plausibly-complete entries (measured 1 MiB floor; undecidable ⇒ not counted). Campaign fault 2 replayed on demo-hp: phantom rejected + logged once, tier correctly DUE and backed up, and **no thrash** on the inverse | — | -| **R-99** | Server-side prune **never removes** a phantom snapshot. Confirmed it does NOT count them toward `keep-last` (dry-run kept 2 real + the phantom) so there is **no retention/data-loss bug** — but one accumulates per aborted upload, forever | READY (S) | — | Decide a cleanup path. Deletion on a **customer** datastore is a separate ruling — detection shipped, removal deliberately not automated | CC | -| **F-CRIT-1** | ~~An app that **fails to restart** after a quiesce never alarms on any channel — `restartAll` discarded the error AND `StateStopped` was whitelisted on invariant I1, which the quiesce path had made false~~ | **SHIPPED + PROVEN-LIVE** (controller v0.179.0, 2026-07-28) | — | Both causes fixed. Live on demo-hp: alarmed 9s after grace expiry, banner shows `(stopped)`; a deliberate user stop stayed silent through 9 dead-app scans | — | -| **F-A1** | ~~A restore-test in flight made a healthy backup report as FAILED (HTTP 409 read as a tier failure): breaker armed + operator emailed, on both boxes~~ | **SHIPPED + PROVEN-LIVE** (controller v0.179.0, 2026-07-28) | — | 409 → contention: tier stays DUE, dropped before anything stops (15m), and BLOCKED alarm if contention outlives the agent's 120m ceiling (3h). Hub DB: 409 → **0** operator emails, real failure → **1** | — | -| **C9-F1** | ~~Tier-2 „Fájlok visszaállítása" is offered for apps whose copy has no restorable file leg; stops the app, restores 0 files, reports „Nincs hiányzó fájl — minden fájl megvan a helyén."~~ | **SHIPPED + PROVEN-LIVE** (controller v0.183.0, 2026-07-28) | — | **Phase 0 sized it: 43 of 53 catalog apps read NOTHING, 9 read file legs but never their DB/volumes, 1 stateless.** Honesty half shipped: `Tier2RestoreCoverage` refuses UP FRONT without stopping the app and NAMES the working action; a run that proceeds claims only what it **examined** and discloses that the database and volumes are not covered. Live on demo-felhom: bookstack refused, uptime stayed „Up About an hour" (was „Up 25 seconds"); paperless A1 re-run still byte-identical, 16/16 docs clean | — | -| **C9-F2** | ~~An app in a Docker crash loop never alarms on any channel; `StateRestarting` is in no down-set~~ | **SHIPPED + PROVEN-LIVE** (controller v0.183.0, 2026-07-28) | — | `StateRestarting` deliberately NOT added to `IsDownState` (that alarms on every deploy fleet-wide); a SUSTAINED run becomes down after `crashLoopAfter`=5m, set above the 120s deploy timeout, Mealie's 60s start_period and R-97b's 180s grace. Dashboard counter uses the same predicate so it no longer contradicts the alarm. Red-proof that matters: the naive `IsDownState` change fails the brief-restart test | — | -| **C9-F3** → **R-104** | An **interrupted offsite run leaves an exclusive restic lock the self-heal cannot reach**: `resticStep` (`offbox.go:634-648`) has `unlock --remove-all`, but `ensureOffboxRepo`'s probe fails first, `classifyResticProbe` (`offbox.go:77-93`) has no lock case → `"other"` → fail-fast. Tier dead until a human unlocks; `ClassifyOffsiteFailure` likewise has no lock case so the operator is told **„A távoli mentés ismeretlen okból nem sikerült"** for a precisely-known, self-healable condition | **READY (MEDIUM)** | — | Add a lock case to both classifiers and let the probe path escalate to `unlock --remove-all`. Answers Phase C item 8: the repo is NOT usable after a killed run. Cleared manually this run; tier proven working again (`ok`, 1m35s). Reachable by any interruption — container restart, OOM, **host reboot mid-backup** | CC | -| **D5** | ~~**Move app secrets into the LOCAL recovery unit** so Tier-1/Tier-2 restore stop needing the guest~~ | **SHIPPED + PROVEN-LIVE** (controller v0.188.0, 2026-07-30) | — | **CLOSED — the arc's architectural centrepiece is done, and Tier-1/2 no longer depend on the whole-guest tier.** A customer now needs **the drive and nothing else**. **Part 0 overturned the brief's own recommendation, on evidence gathered before any code** — that is the substantive part of this row. It proposed that only `data_key`-flagged secrets travel; two findings killed that: (1) the flag is **unreliable** — only 5 fields across 4 apps carry it, yet `n8n/N8N_ENCRYPTION_KEY` („Titkosítási kulcs"), `wanderer/POCKETBASE_ENCRYPTION_KEY` („Adatbázis titkosítási kulcs"), `calcom/CALENDSO_ENCRYPTION_KEY` and `bookstack/APP_KEY` carry the SAME labels as flagged `adventurelog/SECRET_KEY` and are unflagged (→ **R-127**), so data-keys-only would omit real data keys and the fail-closed gate would not fire for them; (2) a **DB password is not resettable in practice** — proven on a throwaway `postgres:16-alpine`: with PGDATA restored from the volume tar, `POSTGRES_PASSWORD` is **ignored** (initdb skipped), so a regenerated value fails over the compose network (`FATAL: password authentication failed`) while the old one still works AND the dump replay still SUCCEEDS via the container's local trust socket — a restore that reports success onto data the app cannot reach. 18 DB/root-password fields affected; MariaDB fails louder (`getMariaDBPassword` reads the new value against a datadir holding the old hash → Access denied). **Operator ruling 2026-07-30: `type: secret` travels (45 fields), `type: password` NEVER (7) plus a code register (`vaultwarden/ADMIN_TOKEN`); plaintext.** The exclusion is what LICENSES the plaintext — coupled, not independent. `stacks.PortableSecretEnvVars` is the single boundary; the register is **code, not a catalog flag** (a boundary a catalog push can move is not a boundary — R-97a). **Precedence: the UNIT WINS** over the guest, because the unit's secrets were captured in the same run as the dumps beside them and therefore match the data being restored; pinned both directions. **Fail-closed data-key gate UNCHANGED.** Manifest → schema 2 + `portable_secret_env_vars` (names only); schema-1 units still restore from the guest. **Live proof** on a scratch drill guest through the real endpoints: AdventureLog restored with the guest `app.yaml` moved aside → `secrets recovered=2/2`, 27.6 s, then **the app read the seeded row over TCP with its own credential** (the observable that matters), pre-backup row back / post-backup row gone, **no `.sql` dump** so the DB came from the volume tar. Withheld half proven with Grafana: sentinel live in the container, `ENC:` in the guest, **0 files** under the whole backup namespace. 4 red-proofs, each verified to land. `audits/D5-drive-alone-restore-2026-07-30.md` Flips `07` §3, §7.1, §7.3, §7.4 (new), §8 rows 3/3c/13, §10.1; new capability-map row. **Consequence recorded, not changed:** the unit already travels to Tier-2 (another customer drive, plaintext, same reasoning) and offsite via restic (encrypted at rest under the customer-owned repo password) — no tier code touched | — | -| **R-127** | **The catalog's `data_key: true` flag is UNRELIABLE — at least four data-encrypting keys the catalog itself labels as encryption keys are unflagged; and the O4 restore path can regenerate a DB password that then does not match the restored data directory** | **READY (S/M)** | — | **Found by D5's Part 0, and it is why D5's boundary is `type: secret` rather than `data_key`.** Two separable legs. **(a) The misclassification.** Only 5 fields across 4 apps set `data_key: true` (`adventurelog/SECRET_KEY`, `homebox/HBOX_AUTH_API_KEY_PEPPER`, `papra/AUTH_SECRET`, `sparkyfitness/{API_ENCRYPTION_KEY,BETTER_AUTH_SECRET}`), yet `n8n/N8N_ENCRYPTION_KEY` („Titkosítási kulcs"), `wanderer/POCKETBASE_ENCRYPTION_KEY` („Adatbázis titkosítási kulcs"), `calcom/CALENDSO_ENCRYPTION_KEY` and `bookstack/APP_KEY` are unflagged — the catalog's own Hungarian labels contradict the flag. **D5 makes this non-urgent but not harmless:** everything `type: secret` now travels, so the keys DO reach the drive; what stays wrong is the **fail-closed gate**, which only refuses for `data_key` names — so if one of these is missing from both sources the restore proceeds onto data it cannot decrypt instead of refusing. Fix = flag them (`app-catalog-felhom.eu`, a catalog-only change) + a gate/test that the flag set and the label set agree. **(b) The regenerated-DB-password trap.** `internal/backup/restore_unit.go` O4 generates a replacement for any missing non-data-key secret. Proven on `postgres:16-alpine`: with PGDATA restored from the volume tar, `POSTGRES_PASSWORD` is ignored (initdb skipped), so the app fails over the compose network while the dump replay still succeeds through the container's local **trust** socket — success reported, data unreachable. v0.188.0 corrected the WARN's false claim that *"stored data is unaffected"* and scoped it, but did **not** add a guard: D5 shrinks this to the rare case (the secret was empty at capture AND absent from the guest). Real fix = either treat a DB password as fail-closed like a data key, or `ALTER USER` to the regenerated value after the volume restore. 18 DB/root-password fields are in scope; MariaDB fails loudly instead (Access denied), which is the safer half | CC | -| **R-126** | **A `.fab` bundle — plaintext secrets, OPTIONAL password — can be exported ONTO a NAS.** `storageDriveList()` (`internal/web/handler_export.go`) does not filter network paths | READY (S) | — | Split out of R-108, which closed without it: this is an explicit customer-chosen **export destination**, not a browsing surface reaching a backup tree, so it was never part of D5's precondition (`07` §7.3 records that reasoning). Was recorded inside R-108's row as its "second effect, independent of D5"; promoted to its own row so it does not vanish with R-108's closure. Fix = filter network paths out of the export destination list, or require the bundle password when the destination is a share | CC | -| **F-DIAG** | ~~Four distinct offsite failure causes collapse into two operator-visible strings~~ | **SHIPPED** (controller v0.182.0, 2026-07-28) | — | `ClassifyOffsiteFailure` → quota / orphaned / no_repo / no_units / transport / **unknown**, each with its own Hungarian message. Unclassifiable says so rather than being folded into a neighbour. **Secrets:** the old message was a raw `err.Error()` passthrough carrying `sftp:@:`; redaction is now by the target's **actual** host/user/path (a first regex-only attempt leaked on a bare hostname and its own test caught it). Unit-proven; **not** yet exercised by a live offsite failure of each class | — | -| **F-OPS** | ~~A manual `pct restore` inherits the source guest's bind mounts — during a real DR, on a different host, under pressure~~ | **DOCUMENTED** (2026-07-28) | — | `documentation/runbooks/RUNBOOK-manual-guest-restore.md`: which `mpN` are volumes vs host binds, the `mp9` source-VMID trap (it can bind **another guest's bootstrap credentials**), strip-and-re-add before first boot, and a positive pre-start verification. Docs only by design — the agent already neutralises binds on its own restore paths, and a second implementation would drift | — | -| **F-REBOOT** | ~~A guest rebooted during its backup does not come back — shutdown completes, start never happens, no self-heal; 9m47s total appliance outage with every alarm silent~~ | **SHIPPED + PROVEN-LIVE** (agent v0.107.0, 2026-07-28) | — | 60 s guest-power watchdog; `onboot` is the deliberate-stop discriminator (already the stale-lock path's, and what `pve-guests` consults), retry bounded 3x/1m-2m-4m then escalates once. Live on demo-hp: **120 s unattended** vs the incident's 587 s with a human; Scenario B proven (an `onboot:0` guest left stopped) | — | -| **F-LEAK** | ~~A failed restore-test cannot destroy its own scratch guest (403 `VM.Allocate`); the 10-slot VMID band shrinks silently~~ | **SHIPPED + PROVEN-LIVE** (agent v0.110.0 + host-install v1.21.0, 2026-07-28) | — | **Three attempts, two refuted live.** (1) Pool adoption: `PUT /pools/{pool}` also needs `VM.Allocate` on the VM — membership cannot bootstrap its own authority. (2) Per-path `/vms/990000..990009` ACLs: work, but PVE's destroy calls `remove_vm_access` (`LXC.pm:906`) which deletes every ACL at `/vms/` — **consumed by the op it authorises**, one use per slot. (3) SHIPPED: 4th root-fenced exception, band enforced in sudoers **literally** (`pct destroy 99000[0-9] --purge`) + in code + at the caller; API destroy still tried first. Live: band PERMITTED, `9201`/`9100`/`9999`/`990010`/`1` REFUSED, and `pct start 990000` REFUSED too | — | -| **F-OBS** | ~~`deadapp-check` leaves NO positive observable on a default (info-level) box — "no alarms" was indistinguishable from "never ran"~~ | **SHIPPED + PROVEN-LIVE** (controller v0.180.0 + agent v0.109.0, 2026-07-28) | — | INFO summary every 20th scan carrying scans/evaluated/down. **Agent v0.109.0 fixes the same shape in the guest-power watchdog shipped hours earlier in v0.107.0** — it logged only at startup and when it acted, so its health could be read only from absence | — | -| **E-2** | ~~Drive-role machinery around the moved vzdump target~~ | **CLOSED — PARTIALLY PROVEN** (Session C, 2026-07-29) | — | **CLOSED by `audits/SESSION-C-2026-07-29.md`.** C1/C2 proven in E-2d; **C3 and C4 PROVEN LIVE** this session (R-114, R-112); **C5 FAILED** — the gate fires and an alarm reaches the hub, but it is the generic event, not `backup_target_absent` (→ **R-116**, the one named open leg). Per the runbook's §9, decided in advance: a failed claim closes E-2 as partially proven with a named leg rather than re-running. The arc's stated definition of done is R-106+R-109, R-108 and D5 — none of which this detour touched. Parts 1–5 complete. Role model + offer-only assignment + Hungarian degraded banner + absent-target signal + installer Case A/B. **NOT yet live-proven:** the DEGRADED banner and the offer acceptance (both demo boxes are healthy, so neither state occurs naturally) and `backup_target_absent` end-to-end. Installer is ~~installer-logic-tested, not install-tested~~ — **INSTALL-TESTED 2026-07-29** on a fresh nested box via the real ISO/PAIRING route, rc=0 (`audits/E2D-fresh-vm-2026-07-29.md` §3). **Of the "NOT yet live-proven" list: Case B + the degraded state are now PROVEN at the installer and API level; the OFFER ACCEPTANCE is PROVEN at the API level** (decline path, `restart_required:true`, E-2a wrapper, healthy-renders-nothing). **Still NOT proven, and now known to be BROKEN rather than merely untested:** the banner/offer never reach a customer (**R-112**) and `backup_target_absent` cannot fire on device loss (**R-113**), with the absent-state message itself wrong (**R-114**) | CC | -| **E-2a** | ~~The target move needs a root-fenced wrapper — the agent cannot do it~~ | **SHIPPED + PROVEN-LIVE** (agent v0.113.0 + host-install v1.22.0, 2026-07-29) | — | `felhom-backup-target-apply` behind a literal `FELHOM_BACKUPTARGET` sudoers alias; the agent's PVE role was NOT widened. Enforces F-1 (`mountpoint -q`) and F-2 (`is_mountpoint 1` hardcoded), refuses a root-device target, has NO storage-removal path (grep-assertable), is idempotent and refuses to repoint. All five laws proven live as root on demo-hp with 0 stray storages | — | -| **E-2b** | ~~`NotifyStorageDisconnected`/`Reconnected` defined and called NOWHERE — a drive going absent emitted no event on any channel~~ | **SHIPPED + PROVEN-LIVE** (controller v0.184.1 + agent v0.112.0 + hub v0.81.0, 2026-07-29) | — | Seam wired in `ReconcileDriveGates`; a target drive raises the specific `backup_target_absent` instead. **A keying bug was caught before deploy:** `a.Path` is the registered GUEST path, not the agent's host `MountPath`, so the target branch was unreachable — every absent drive, target included, fell through to the generic event (v0.184.1). Tests observe the WIRE (httptest hub), not a mock | — | -| **E-2c** | ~~E-1 put the whole-guest backups on a drive `POST /disks/eject` would eject~~ | **SHIPPED + PROVEN-LIVE** (agent v0.112.0, 2026-07-29) | — | Eject + decommission refuse 409 on the backup-target mount, naming the storage and the remedy. **Live on BOTH boxes:** demo-hp `/mnt/nvme-1tb` and demo-felhom `/mnt/hdd_1` both refused, drives unmoved. NOT a role reclassification — `RoleForStorage` untouched, because on both boxes that drive is ALSO the enrolled user-data drive; `TestEjectStillAllowedOnANonTargetDrive` pins the non-over-correction and `/var/lib/vz` is still refused by the PRE-EXISTING role gate, not this one | — | -| **R-124** | **The recipe spells PBS's root namespace `"root"`, but the PBS API spells it `""`** and no namespace is literally named `root` — an operator pasting the field into `pct restore --ns root` gets a failure | READY (XS) | — | Pre-existing wire convention (`ToHub` has normalised empty→`"root"` since slice 6), deliberately NOT changed under R-106 so the field's meaning did not shift mid-fix. Documented at `hub.PBSRootNamespace`. Affects only a box with no `namespace` line — **no real customer today**, all three are per-customer. Fix = emit `""` + rely on `namespace_state`, or emit a `--ns`-ready form | CC | +| **R-10** | T-6E-1: DB-dump dir-fsync asymmetry (LOW, confirmed in 6E) **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-15, size XS, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **T-6E-1, confirmed in CAMPAIGN-6E.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | One-line hardening; batch with the next controller task | CC | +| **R-25** | **Device-node TOCTOU hardening (drive init).** Graduate the controller v0.141.0 Observation: the `format → resolveEnrollUUID(path) → AssignDisk(uuid)` sequence has a narrow /dev-re-enumeration window (agent-guarded on the destructive format via anti-retarget durable-id; benign fs-UUID mount). Bind resolve+assign to the format's durable-id so the mount can't target a moved node. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-19, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | From the v0.141.0 F6 commit's security-review finding (`felhom-controller` REPORT). Low real risk (single-operator, agent-guarded), but cheap to close | CC | +| **R-30** | **[P2-HIGH] Liveness presence should come from the wait channel, not the report clock.** The box was powered off at the start of the rehearsal, yet the hub carried it as healthy until the staleness threshold expired ~30 min later (`host_stale` 16:05:24 "no report for 30m"; cleared 16:33:24 "was stale for 27m"). The host-delete guard compounds it: RESET refuses while any host row exists, so a stale-but-"Online" host stalls a forced teardown. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: operator-side presence delay; alarms still fire after 30 minutes and no household data is at risk.** | — | Direction: derive presence from **Dir-2 long-poll connectedness (~90 s grace)**, decoupled from notification hysteresis (the hysteresis is right for *alerting*, wrong for *presence*); an agent/ep0 analog can follow. Pairs with R-13/R-23 — the transport already exists, this is about believing it. *(Discussed in-session as "R-29"; that number was already taken by the gate-rot item earlier the same day, so it is R-30.)* | CC | +| **R-31** | **[P2-HIGH] Offsite provisioning is synchronous with no status affordance.** Save runs the Hetzner sync in-request, so the request can hit the nginx 504 **while succeeding server-side**: the operator cannot tell failed from slow, and a retry races the first attempt. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: operator-only; a known workaround (click once, wait, verify) exists.** | — | Direction: make it async + a status card, reusing the proven **awaiting-card/poll idiom** (v0.138.0 escrow card). **Interim mitigation belongs in R-3 as an operator note: click once, wait, verify — do not re-click.** | CC | +| **R-32** | **[P2-HIGH] RESET must purge the customer base dir; the orphan card must stay honest; unattributed bytes must be visible.** The rehearsal's S7 said in advance that an orphan card would BE a finding — and one appeared (16:58:14). Cause: RESET's `"hetzner":"ok"` leg destroys the sub-account, but **a Hetzner sub-account is an access-control object, not a data object** — its directory survives, so re-enabling offsite recreated an account over the previous lifecycle's ciphertext, encrypted under a key that same RESET had destroyed. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | **Ruling from the run (three parts, deliberately separate):** (1) because RESET destroys custody, the ciphertext it leaves behind is unrecoverable **BY DESIGN** → RESET gains a **main-account purge of the customer base dir** (the existing operator ack already covers it); (2) the **move-aside guard STAYS** for reinstall-*without*-RESET — there custody survives and the card's "history recoverable" promise is true (R-26 depends on exactly that); (3) the operator **Restic tab shows per-customer directory bytes vs attributed snapshot bytes**, so dead data cannot hide. Measured on the pool box that night: **49 M attributed** (2 snapshots, 48.717 MiB) against **1.4 G + 3.0 M unattributed** across TWO `.orphaned-*` dirs. Evidence `restic-and-pool.txt` | CC | +| **R-35** | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` | CC | +| **R-49** | **[P2] The offsite capture set is ~90% cache and duplication — 1.1 GB of a 1.2 GB immich "photo backup".** Measured 2026-07-19: `immich_ml_cache.tar` **823 660 032 B (~60%)** — re-downloadable ML model weights; `immich_postgres_data.tar` **308 251 136 B (~23%)** — a raw tar of the postgres data dir that DUPLICATES the logical `.sql` dump captured beside it; `upload/backups/` **18 MB** — immich's own nightly dump, a backup inside the backup, growing daily; plus the stranded pre-v3 `dccc13fe…` tree (~36 MB) no DB has ever referenced. Actual irreplaceable content: **72 MB of originals**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-19, size S–M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: wasted space and transfer, no data risk; needs a capture-set ruling, not a sale blocker.** | — | **Evidence: `audits/DIAG-immich-restore-round2-2026-07-19.md` §4 (full byte breakdown).** This is the customer's offsite quota and transfer cost, and it lands on the Hetzner sub-account they are billed for. **Recorded, deliberately not changed** — a capture-set exclusion is a data-loss-shaped decision and gets its own ruling, not a drive-by edit. Candidates in priority order: (a) `immich_ml_cache` — pure cache, strongest case; (b) the `postgres_data` volume tar where a logical dump of the same DB is already captured (the dump is what the restore path actually replays); (c) `upload/backups/`. Likely generalises past immich into a template-classification rule about cache volumes and self-backup directories, so it should be specified against the catalog, not one app | CC | +| **R-76** | **FileBrowser-created folders break the setgid chain, and a drop-zone's mode is not stable** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-26, size S, roadmap state `idea (surfaced by the R-75 spike, 2026-07-26)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | Two related findings from `audits/SPIKE-catalog-data-paths-2026-07-26.md` P3/P5, both **pre-existing** and deliberately left alone by that spike. **(a)** FileBrowser Quantum 1.3.3 creates files `0644` and folders `0755` and does **not** propagate the setgid bit — even though the entrypoint wrapper's `umask 002` really is in effect (`/proc/1/status` `Umask: 0002`). Group inheritance itself works (a file uploaded into a 2775 group-100 dir landed group 100, not the process gid 1000), so the convention's *group* half holds and only its *mode* half is lost. The consequence is proven with a control: inside a UI-created `0755` folder a gid-1000 process's file landed group **1000**, while the identical write into the 2775 parent landed group **100**. So **any folder a customer creates through FileBrowser breaks the shared-group chain one level down.** Latent today — every userdata-touching catalog app that declares an identity declares uid/gid **1000**, the same uid FileBrowser runs as, so owner permissions mask it; it bites the day a content app runs as a different non-root uid with gid 1000. The comment at `infra/infra.go:156` is right that the image ignores `-e UMASK` but does not say t | CC | +| **R-78** | **`local_api` authority ruling — auto-reconcile vs detect-only** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-25, size M, roadmap state `idea (deferred OUT of R-77 on purpose)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **An OWED OPERATOR DECISION, not a defect — moved because an owed ruling hidden among feature ideas is the shape this session exists to remove.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | R-77 ships detection because the fix is genuinely undecided, and **both directions can lose customer-visible function**. **Direction 1 (today):** `controller.yaml` wins and drift is silent → the 2026-07-25 island migration blinded the whole fleet's control plane for 17.5 h (drive gate, guest-reboot recovery, quiesce/backup all degrade). R-77 makes that loud but does not stop it recurring. **Direction 2 (`bootstrap.json` wins, auto-reconcile on boot):** a guest whose `controller.yaml` is CORRECT and whose `bootstrap.json` is stale — a half-completed re-provision, a hand-repaired guest, a setup-wizard box — gets a **working channel clobbered on the next restart**, fleet-wide and silently, during a routine deploy. That is not obviously better than the bug. Needs a spike: which writer is authoritative per field (endpoint vs fingerprint vs token — `mergeLocalAPI` replaces the whole block, so they cannot be reconciled independently today), whether the agent side should stamp a generation/mtime so 'newer wins' is even expressible, and whether reconcile should require an operator ack. Until then the drift alert plus a manual edit is the supported path. | CC | +| **R-79** | **`report.Issues` / `report.Warnings` are English on customer-facing surfaces** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-26, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | **Whole-surface, not a one-off** (DIAG §6): every producer is English — `"SSD/HDD disk usage critical"`, `"Docker: %v"`, `"Protected container not running: %s"`, and all six `Warnings` strings. They render on the customer's Hungarian dashboard, and the `health_critical` path has reached the **customer** email channel three times historically. Deliberately NOT bundled into R-77: a copy sweep across every producer would have buried two safety fixes in string churn, and the seam is not obvious — translate at the producer, or at the render/notification boundary where operator-English and customer-Hungarian already diverge? Pick the seam in a spike; the strings are mechanical after. | CC | | **R-89** | Retention as a per-customer **commercial** policy on the hub | READY (increment 2) | — | Policy object + reconciler → ep0 prune job; keep box tokens write-only | CC | +| **R-91** | Old 13 GB datastore copy at `/srv/pbs-felhom` on ep0's root disk | WATCHING | demo-felhom's first **post-migration** PBS backup | Delete once it lands; fix `CONTEXT.md:1018` same commit | CC | | **R-92** | Hub PBS-DR gauge is 0.1 GB-granular — small deltas unverifiable | READY (XS) | — | Widen precision when retention becomes customer-visible | CC | | **R-93** | `drill-r50` is both a blocked customer and the only drift fixture **FACT 2026-09-13 (R-461): the fixture is GONE — `qm list` is empty on both demo boxes, so neither option is available and the row's premise no longer holds; the operator decides whether that closes it or reopens it as "build a drift fixture".** | READY (XS) | — | Retire it for a synthetic fixture, or unblock + silence per-customer | CC | +| **R-95** | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` **-- UNBLOCKED 2026-09-22, not closed.** The premise in this row's own title - *SFTP cannot express append-only* - is answered by Hetzner (ticket #2026090103040671, recorded in R-433 and R-436): on a Storage Box the `rclone serve restic --stdio --append-only` backend can be pinned to an SSH key as a FORCED COMMAND, so append-only IS expressible without a new machine. **The row stays open because nothing has been measured**: no forced-command key exists on `u629488`, no `forget --prune` has been refused through one, and the protection only holds if EVERY key that can reach the repository carries the prefix. The next step is R-436's, and it is small. | CC **SPIKE 2026-09-01 — `audits/SPIKE-r95-offsite-delete-2026-09-01.md`. THE WORD "ARMED" ABOVE IS NOT SUPPORTED AND IS WITHDRAWN PENDING R-429:** no `.snapshots` is visible to either box's sub-account (measured, both machines, with controls), so the seven-day bound is unverified and unverifiable from the product side. **Q3 (documented, `hub/internal/hetznerapi/hetznerapi.go:38-45`): the sub-account API has ONE permission axis, `readonly` — there is no append-only, so the PBS shape does NOT transfer** (PBS is a server that can refuse; a Storage Box is a filesystem that runs nothing). **Q5 (measured): withdrawing delete does NOT wedge the store** — restic treats a dead owner's lock as stale and proceeds — so the constraint everyone feared is not the blocker; but `unlock --remove-all` lies about success (R-430), and the crash-lock window is UNKNOWN. **Q6 (measured): restic 0.14.0 DOES speak `rest:`** (control: `banana:` → invalid backend), and append-only is a rest-server flag, not a restic one — reachable, but it needs a machine in the recovery path and ep0 is protected. **Q7 (measured): detection is nearly free** — `snapshot_count` already reaches the hub and the hub APPENDS reports, so the comparison needs no box change. **RECOMMENDATION: answer R-429 first (Viktor, ten minutes), then build detection, then move retention off the box; defer the transport change.** Two `forget --prune` sites must be disarmed together — `offbox.go:1388` AND `offbox.go:1759` — or R-191 repeats. **RE-SCOPED 2026-09-01 — THE STORY WAS WORSE THAN THE TRUTH FOR TWO MONTHS.** The box can delete its own LIVE repository, **but it cannot write to the daily snapshots of it** — MEASURED, not cited: a write into `/.zfs/snapshot` is refused on both boxes while the same write to the account home succeeds (R-432). Seven daily snapshots are confirmed in the panel (R-429). **So a deletion costs at most the data written since the last daily snapshot, and the rest is recoverable — file by file, one customer at a time, with no effect on anyone else** (vendor: *"You can download individual files or entire directories as usual"*; *"It is not possible to write to the `/.zfs` directory or its subfolder"*). **NOT open-ended loss.** Two caveats kept honest: a panel-driven snapshot restore rolls back the WHOLE Storage Box and deletes newer snapshots, which is why the per-file route matters; and per-file recovery is **operator-only today** (R-432). **DETECTION SHIPPED hub v0.111.0 (R-431)** — an unexplained fall is noticed within a day. **THE RANKING IS VIKTOR'S:** this has been #1 since July on the old story. On the new facts I would rank it below the items that can still lose data outright, but I am not re-ranking it myself. `audits/SPIKE-r95-offsite-delete-2026-09-01.md` **DRILL 2026-09-01, LATER THE SAME DAY — THE RE-SCOPE'S SECOND HALF IS WITHDRAWN.** The drill that was to walk the recovery found there is no route to walk: **no snapshot is reachable from a sub-account by ANY name** (R-433) — 777,600 exact names in the vendor format over nine days, zero hits, with a passing control, plus the structural reason (`/home` st_dev 0,82 vs `/.zfs/snapshot` st_dev 0,276, and `/home/.zfs` absent). **Clause (a) stands: the box can delete the live repo and cannot write into the snapshot area.** Clause (b) — *"the rest is recoverable file by file"* — is NOT SUPPORTED. **The drill was STOPPED before its destructive phase on the operator's ruling**, because with no recovery route the deletion would have destroyed real history to buy only an alarm test that could not fire at the specified size (R-435). Nothing was deleted; the store is verified untouched at 69 snapshots. **So the comfort that lowered this row rested on an unwalked route, and the route does not exist.** Two new leads decide what happens next: **R-433** (can the MAIN account see them? nobody here holds that credential) and **R-436** (`rclone serve restic --stdio` is offered server-side and restic speaks `rclone:` — measured — which could make real prevention cheap, IF the provider pins `--append-only`). **The rank stays Viktor's. Plainly: the argument that moved this row down is the argument the drill removed.** `audits/evidence-drill-r95-recovery-2026-09-01/` **BLOCKED-ON-PROVIDER 2026-09-01, and the record of the demotion is kept deliberately: this row spent ONE DAY demoted on a clause that did not hold.** It was re-scoped down on the morning of 2026-09-01 on the strength of *"recoverable file by file"*, and that clause was measured false the same afternoon (R-433). **On today's evidence it belongs back near the top — that is a proposal, not an action; CC has not re-ranked it and will not.** Both questions that can settle it are drafted in `documentation/runbooks/provider-questions-2026-09-01.md`: **Q1** decides how urgent this is, **Q2** (R-436) could remove the root cause cheaply. **Precondition on any build here: R-430**, which is harmless only while the box can still delete. | +| **R-99** | Server-side prune **never removes** a phantom snapshot. Confirmed it does NOT count them toward `keep-last` (dry-run kept 2 real + the phantom) so there is **no retention/data-loss bug** — but one accumulates per aborted upload, forever | READY (S) | — | Decide a cleanup path. Deletion on a **customer** datastore is a separate ruling — detection shipped, removal deliberately not automated | CC | +| **R-104** | **An interrupted offsite run leaves an exclusive restic lock the existing self-heal cannot reach.** `resticStep` has `unlock --remove-all` (`internal/backup/offbox.go:634-648`) but `ensureOffboxRepo`'s probe fails first, `classifyResticProbe` (`:77-93`) has no lock case → `"other"` → fail-fast; `ClassifyOffsiteFailure` likewise, so the operator is told *„A távoli mentés ismeretlen okból nem sikerült"* for a precisely-known, self-healable condition **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-28, size S, roadmap state `READY — 2026-07-28`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. ****PARTLY STALE, checked against live source 2026-08-22 — migrated as written per the task rule, with the staleness named rather than edited away.** The self-heal this row calls unreachable was built: `resticStep` escalates to `unlock --remove-all` and retries once (`internal/backup/offbox.go:763-768`), and `unlockStale` runs before every off-site run and restore (`:1274`, `offbox_restore.go:261`). Its premise that the probe fails first is also doubtful: the probe is `restic cat config`, a read that takes no lock. **What REMAINS true:** `ClassifyOffsiteFailure` (`offbox.go:179-193`) still has no lock case, so if a lock ever did survive both layers the customer would still be told an unknown reason. Re-rank on that basis, not on the original text.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | **Was C9-F3.** Reachable by any interruption — container restart, OOM, network drop, host reboot mid-backup. The tier stays dead until a human runs `restic unlock --remove-all`. Flips: the offsite row in map §C; `07` §8 row 15 | CC | +| **R-105** | **Three hub-held DR records are empty on the entire live fleet.** `hosts.dr_record_json` = `{}` on all 3 hosts; `host_escrow.directive_json` = `{}` on both escrowed hosts; `dr_recipe.host_half.drives` = `[]` on every customer **including two with enrolled data drives** (916 GB USB on demo-felhom, 938 GB NVMe on demo-hp) **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-28, size M, roadmap state `READY — 2026-07-28`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. ****PARTLY FIXED BY ITS OWN UPDATE.** The `drives` third was traced and populated on both demo boxes on 2026-07-28 (the enrolled drives were never PVE storages, so `isUserDataDrive` never saw them). The other two thirds — `hosts.dr_record_json` and `host_escrow.directive_json` — were NOT re-verified this session and are carried as written.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | These are exactly the fields a host-loss recovery reads: `05-hub-architecture.md:175-176,186` names the slim DR record as one of four durable sources; `06-offsite-connectivity.md:148-150` says the escrow upload carried the DR directive; `felhom-agent/internal/dr/plan.go:34-35` makes `PlannedDrive` the re-attach-by-`durable_id` wrong-disk guard. **The three may have different causes** — `isUserDataDrive` (`internal/hub/dr_recipe.go:129-136`) requires type `usb`/`local-dir` **and** a non-empty `DurableID` **and** `MountPath`, and which of the three fails was not traced. Evidence: `architecture/_recovery-inventory-2026-07-28.md` Part D2.3. **UPDATE 2026-07-28 (vzdump-target move): the `drives` third is TRACED and now POPULATED on both demo boxes.** Cause: the enrolled data drives were never PVE storages at all — only agent-generated systemd mounts — so they never entered `report.StorageTargets` and `isUserDataDrive` never saw them. Giving each drive a `dir` storage at its own mountpoint supplied all three required fields at once (type `local-dir`, fs-UUID durable id, mount path), and the recipe now emits `uuid:91d2dc2d-…`/`/mnt/nvme-1tb` on demo-hp and `uuid:47a3361a-…`/`/mnt/hdd_1` o | CC | +| **R-118** | **An absent drive's union row advertises the ROOT filesystem's capacity as its own.** In the absent-state payload the registry-union row reports `total_bytes: 49675956224 / used_bytes: 4584579072` — **byte-identical to the `local` row** (`durable_id: path:/var/lib/vz`, i.e. `pve-root`) in the same response. The real drive is **4 GB** | **READY (XS) — NEW 2026-07-30** | — | Cause: `statfsCapacity(d.MountPath)` (`disks.go:335-338`) statfs's `/mnt/cel`, which with the device gone is a **bare directory on the root filesystem**. `observe.go:176-183`'s comment warns about exactly this trap and guards the Observe path ("*an unmounted removable dir-storage's mountpoint reverts to a bare directory on root … catastrophic DR mis-id*"); **the union path has no equivalent guard.** **Not a DR mis-id** — `durable_id` on that row is still the correct `uuid:…`, so re-attach identity is safe. It is a **false capacity** reaching every consumer of `total_bytes`/`used_fraction` (fill monitors, storage cards): a detached 4 GB drive advertises 46 GiB at 9.2 % used. Same class as `role.go:180-181` — an absent drive's fields decaying to the root filesystem's. Evidence: `audits/DIAG-r116-disks-payload-2026-07-30.md` §12 | CC | +| **R-121** | **A BOX's installed agent can sit releases behind the vouched one and nothing notices — the R-120 gate does not cover it.** demo-hp ran agent **0.113.0** while the hub vouched **0.116.0**, through the whole R-116/R-117 arc, and no signal existed on any channel | **READY (S) — NEW 2026-07-30** | — | **Fourth instance of the drift family** (R-111 golden's agent 17 releases behind, R-115 built+deployed but never published, R-120 golden a controller behind — and now installed-vs-vouched on a live box). **Confirmed at source that R-120's gate cannot catch it:** `hub/internal/web/configs.go:1165-1169` compares `goldenVer` against `store.NewestReportedControllerVersion()` — it is a **golden-artifact vs fleet-CONTROLLER** check and says nothing about the agent installed on a box. **`MinAgent` does not cover it either:** it is used to HOLD the controller floor for a box whose agent is too old (`hub/internal/api/handler.go:530-538`, `store.go:1857`) — protective, not an alarm — and demo-hp's 0.113.0 **equalled** `min_agent` 0.113.0, so even a floor comparison was satisfied. **The cost, measured:** R-117's whole subject is the R-113 conjunction, which landed in **0.114.0** — so the designated drill host could not exercise the code under investigation at all, and the R-117 spike had to route every predicate result through an out-of-repo probe built from `main` instead of the installed agent (`audits/SPIKE-r117-bind-liveness-2026-07-30.md` §1, §2.3). Discovered because the R-117 task made bringing the box current an explicit prerequisite. **Fix shape (not implemented):** the hub already receives `AgentVersion` on every host report, and already has semver comparison in Go — the missing piece is a checker comparing reported agent vs the **vouched** agent and surfacing it, operator-tier. Note the honest tension: a box legitimately lags between publish and deploy, so this wants a staleness window rather than an instant alarm | CC | +| **R-124** | **The recipe spells PBS's root namespace `"root"`, but the PBS API spells it `""`** and no namespace is literally named `root` — an operator pasting the field into `pct restore --ns root` gets a failure | READY (XS) | — | Pre-existing wire convention (`ToHub` has normalised empty→`"root"` since slice 6), deliberately NOT changed under R-106 so the field's meaning did not shift mid-fix. Documented at `hub.PBSRootNamespace`. Affects only a box with no `namespace` line — **no real customer today**, all three are per-customer. Fix = emit `""` + rely on `namespace_state`, or emit a `--ns`-ready form | CC | +| **R-126** | **A `.fab` bundle — plaintext secrets, OPTIONAL password — can be exported ONTO a NAS.** `storageDriveList()` (`internal/web/handler_export.go`) does not filter network paths | READY (S) | — | Split out of R-108, which closed without it: this is an explicit customer-chosen **export destination**, not a browsing surface reaching a backup tree, so it was never part of D5's precondition (`07` §7.3 records that reasoning). Was recorded inside R-108's row as its "second effect, independent of D5"; promoted to its own row so it does not vanish with R-108's closure. Fix = filter network paths out of the export destination list, or require the bundle password when the destination is a share | CC | +| **R-127** | **The catalog's `data_key: true` flag is UNRELIABLE — at least four data-encrypting keys the catalog itself labels as encryption keys are unflagged; and the O4 restore path can regenerate a DB password that then does not match the restored data directory** | **READY (S/M)** | — | **Found by D5's Part 0, and it is why D5's boundary is `type: secret` rather than `data_key`.** Two separable legs. **(a) The misclassification.** Only 5 fields across 4 apps set `data_key: true` (`adventurelog/SECRET_KEY`, `homebox/HBOX_AUTH_API_KEY_PEPPER`, `papra/AUTH_SECRET`, `sparkyfitness/{API_ENCRYPTION_KEY,BETTER_AUTH_SECRET}`), yet `n8n/N8N_ENCRYPTION_KEY` („Titkosítási kulcs"), `wanderer/POCKETBASE_ENCRYPTION_KEY` („Adatbázis titkosítási kulcs"), `calcom/CALENDSO_ENCRYPTION_KEY` and `bookstack/APP_KEY` are unflagged — the catalog's own Hungarian labels contradict the flag. **D5 makes this non-urgent but not harmless:** everything `type: secret` now travels, so the keys DO reach the drive; what stays wrong is the **fail-closed gate**, which only refuses for `data_key` names — so if one of these is missing from both sources the restore proceeds onto data it cannot decrypt instead of refusing. Fix = flag them (`app-catalog-felhom.eu`, a catalog-only change) + a gate/test that the flag set and the label set agree. **(b) The regenerated-DB-password trap.** `internal/backup/restore_unit.go` O4 generates a replacement for any missing non-data-key secret. Proven on `postgres:16-alpine`: with PGDATA restored from the volume tar, `POSTGRES_PASSWORD` is ignored (initdb skipped), so the app fails over the compose network while the dump replay still succeeds through the container's local **trust** socket — success reported, data unreachable. v0.188.0 corrected the WARN's false claim that *"stored data is unaffected"* and scoped it, but did **not** add a guard: D5 shrinks this to the rare case (the secret was empty at capture AND absent from the guest). Real fix = either treat a DB password as fail-closed like a data key, or `ALTER USER` to the regenerated value after the volume restore. 18 DB/root-password fields are in scope; MariaDB fails loudly instead (Access denied), which is the safer half | CC | | **R-129** | **Every doc says demo-hp has "no baked SSH key"** and needs the G1 break-glass password — but `ssh -o BatchMode=yes demo-hp` authenticated **by key**, first try, 2026-07-31 | READY (XS) | — | Stale in the expensive direction: a session that believes it sends itself to the hub vault for a credential it does not need. Verify who owns the key and when it landed, then correct `CLAUDE.md`, `runbooks/target-selection.md:41-42`, `runbooks/workspace-CLAUDE.md` and `felhom-agent/CLAUDE.md` together — or remove the key if it was not deliberate | CC | | **R-130** | **A "hard min" that only warns.** A fresh box's `local-lvm` was ~75 GiB against `HARD_MIN_LVM_GIB=120` (`scripts/felhom-host-install.sh`); the installer logged `[WARN] local-lvm free ~75 GiB < hard min 120 GiB` and went on to a **fully successful** install | READY (S) | — | Either the minimum is not hard (rename it and state the real floor) or it is wrong (and 120 GiB is not what a working appliance needs). Leaving it is the R-29 shape: a check that reads as coverage while providing none. Evidence: same audit §8 | CC | -| **R-131** | **`sess-f` is a fourth orphaned scratch customer** on the hub ("R-120 golden 0.186.0 proof", DOWN), left by the 2026-07-30 session | READY (XS) | — | After `drill-r50`, `sess-c`, `sess-d` — the accumulation `runbooks/target-selection.md:86-87` and `PROMPT-TEMPLATE.md` §13 both warn about, now on its fourth instance. Delete it (see the recorded command in `audits/tester-gate-golden-0.188.0-2026-07-31.md` §7.1); the recurrence itself argues for a periodic scratch-customer sweep rather than another reminder | CC | -| **R-132** | **`curl -w '%{redirect_url}'` reconstructs the request URL WITH its basic-auth credential** — so a `-u ":$HUB_PW"` call that never put the password in a URL still printed it | **ACTION: rotate `HUB_PW`** | — | Happened on 2026-07-31 while red-proofing the R-120 gate: the hub operator password was written to the session transcript by the write-out format, not by the request. `-u` is safe; the *reporting* was not. Rule: read the redirect from `-D -` and grep `^Location:`, never `%{redirect_url}`, on any authenticated call. Rotate the hub password (`/configuration` → Login password; ConfigMap `auth.password_hash` is the reset path) and update `~/.config/credentials` | Viktor | -| **R-415** (was R-133) | **The hub enforces uniqueness on `customer_id` only** — `domain` is `TEXT NOT NULL DEFAULT ''` with no UNIQUE/CHECK (`hub/internal/store/store.go:114`) and the create path only rejects a duplicate id (`hub/internal/web/configs.go:644`), so two customers can be given the identical domain silently | READY (XS) | — | Harmless while every customer owns their own zone; a real footgun the moment customers share one (the subdomain-onboarding plan). Fix = reject a duplicate domain on create/edit, or warn. Evidence: `audits/RECON-subdomain-onboarding-2026-07-31.md` §2.2. **RENUMBERED from R-133 on 2026-09-01 (R-406):** two unrelated findings shared that id. This one kept the SHORTER citation trail (3 references, all inside `RECON-subdomain-onboarding-2026-07-31.md`), so it moved and the plaintext-credential row kept R-133 with its 5 references across `CONTEXT.md`, `break-glass.md`, `hub/CHANGELOG.md`, the capability map and a spike | CC | +| **R-132** | **`curl -w '%{redirect_url}'` reconstructs the request URL WITH its basic-auth credential** — so a `-u ":$HUB_PW"` call that never put the password in a URL still printed it | **WAITING-ON-OPERATOR** — **ACTION: rotate `HUB_PW`** **Folded R-580 2026-10-03** (the same `curl -w %{redirect_url}` credential echo, seen again 2026-09-18). | — | Happened on 2026-07-31 while red-proofing the R-120 gate: the hub operator password was written to the session transcript by the write-out format, not by the request. `-u` is safe; the *reporting* was not. Rule: read the redirect from `-D -` and grep `^Location:`, never `%{redirect_url}`, on any authenticated call. Rotate the hub password (`/configuration` → Login password; ConfigMap `auth.password_hash` is the reset path) and update `~/.config/credentials` | Viktor | +| **R-133** | **The vaulted break-glass console credential is PLAINTEXT AT REST — every hub DB backup is a fleet-wide console-credential dump.** `host_recovery.secret` holds each managed box's `root@pam` password verbatim, so any copy of the SQLite DB (Longhorn snapshot, PBS backup of the hub PVC, a hand-taken copy during a diagnosis) carries root console access to every Felhom host in one file | **READY (M) — NEW 2026-07-31** | — | **The deferred leg of hub v0.84.0** (Console access card), filed separately because v0.84.0 changed only WHO can retrieve the secret, never how it is stored. v0.84.0 makes it more worth doing, not more broken: retrieval now rides the hub SESSION, so the DB and the login password are jointly the whole protection (ruling **S-4**, `CONTEXT.md`). Fix shape: **envelope-encrypt the `host_recovery.secret` column under a KEK held outside the DB** — the hub already proves it can hold something it cannot itself read (escrow blobs), and that contrast is the argument. Two constraints the design must respect: the credential must stay retrievable **when the box is unreachable** (that is the whole point of break-glass), so the KEK cannot live on the box or depend on the agent; and the global-key API path must keep working with the hub UI down. Would flip the capability-map row **"Break-glass management-plane recovery"**, which today reads IMPLEMENTED with this as its caveat | CC | | **R-134** | **Two zone-resolvers disagree on depth.** The controller strips labels progressively (`controller/internal/cloudflare/zone.go:18`); the hub's `resolveZone` tries the exact name then `parentDomain`, which strips exactly ONE label (`hub/internal/cloudflare/unblock.go:115,136`) | READY (XS) | — | For a one-label Felhom-issued subdomain both work; for anything deeper the hub silently fails to find the zone while the controller succeeds — the geo-unblock would then no-op with a "no active zone found" error. One concept, two implementations. Same audit §2.6 | CC | | **R-135** | **`validateCSRF` returns TRUE when there is no session cookie** (`hub/internal/web/server.go:678-683`) — measured live: `POST` with Basic auth and no cookie goes straight past the CSRF gate (404, not 403), while the same POST with a cookie and no token is 403 | READY (S) — **security** | — | Browsers cache HTTP Basic credentials per origin and resend them automatically on cross-origin requests, and `SameSite` does not govern the `Authorization` header. So if the operator has ever Basic-authed to the hub in a browser, any attacker page can POST to every mutating route. Latent on the condition, not guaranteed absent. Fix = require the token whenever the request is not provably programmatic, or drop browser-usable Basic auth. Same audit §4.3 | CC | | **R-136** | **Rename `hub_session` → `__Host-hub_session`** — makes cookie tossing structurally impossible | READY (XS, one line) | — | Verified on the live production response that all three prefix preconditions already hold: `Path=/`, `Secure`, no `Domain`. **Caveat for the ticket:** browsers reject a `__Host-` cookie without `Secure`, and `isSecure` is conditional on `r.TLS`/`X-Forwarded-Proto`, so plain-HTTP *browser* access to the hub would stop working (non-browser access uses Basic auth, unaffected). Tested consequence: `r.Cookie` returns the FIRST match and never tries the others, so a tossed cookie wins outright. Same audit §4.1-4.2 | CC | | **R-137** | **Cloudflare geo-WAF rules are zone-scoped and non-namespaced — four cross-tenant faults.** `globalRuleDesc = "[felhom-geo] Global"` (`waf.go:18`) is one literal description per ZONE; `appRuleDescPrefix` keys by app name with no customer (`waf.go:21`); `BuildGlobalExpression` has no positive hostname scoping (`waf.go:241`); `applyDiff` deletes every `[felhom-geo]` rule not in THIS box's desired set (`geosync.go:320`) | READY (M) — **blocks shared-zone onboarding** | — | With two customers in one zone: they overwrite each other's Global rule forever; one customer's country policy applies zone-wide; per-app rules collide by name; and disabling the feature for one (or the hub's `RemoveGeoRules`) wipes them all. Interim mitigation, no code: keep geo-restriction OFF for every shared-zone customer. Fix = namespace descriptions by `customer_id` + add `http.host ends_with ""` to both expressions — a TWO-REPO change (controller + hub `RemoveGeoRules`). Same audit §5.1 | CC | | **R-138** | **A shared-zone `cf_api_token` is a zone-wide DNS-write capability on a customer's box** — written 0600 to `/opt/docker/stacks/traefik/.env` (`controller/internal/infra/infra.go:123`) | READY (S) | — | Today each box holds a token for a zone nobody else uses, so the blast radius is one customer. Under a shared customer zone, one compromised tester box could repoint every other tester's DNS. The ACME path is already switchable — an empty token selects HTTP-01 (`traefik.yml.tmpl`) — so the fix is policy plus a guard that refuses to hand a shared-zone customer a zone-scoped token. Same audit §5.2 | CC | -| **R-133** | **The vaulted break-glass console credential is PLAINTEXT AT REST — every hub DB backup is a fleet-wide console-credential dump.** `host_recovery.secret` holds each managed box's `root@pam` password verbatim, so any copy of the SQLite DB (Longhorn snapshot, PBS backup of the hub PVC, a hand-taken copy during a diagnosis) carries root console access to every Felhom host in one file | **READY (M) — NEW 2026-07-31** | — | **The deferred leg of hub v0.84.0** (Console access card), filed separately because v0.84.0 changed only WHO can retrieve the secret, never how it is stored. v0.84.0 makes it more worth doing, not more broken: retrieval now rides the hub SESSION, so the DB and the login password are jointly the whole protection (ruling **S-4**, `CONTEXT.md`). Fix shape: **envelope-encrypt the `host_recovery.secret` column under a KEK held outside the DB** — the hub already proves it can hold something it cannot itself read (escrow blobs), and that contrast is the argument. Two constraints the design must respect: the credential must stay retrievable **when the box is unreachable** (that is the whole point of break-glass), so the KEK cannot live on the box or depend on the agent; and the global-key API path must keep working with the hub UI down. Would flip the capability-map row **"Break-glass management-plane recovery"**, which today reads IMPLEMENTED with this as its caveat | CC | -| **R-176** | **Two prerequisites for the R-165 merge are UNMEASURED, and both are cheap.** (a) Whether a **pre-merge archive** (carrying `mp1`) restore-tests cleanly into a **merged-layout** guest — reading `mountParity` (`felhom-agent/internal/reconcile/restoretest.go:347`) says it should, because the restore recreates `mp1` from the archive so archive and restored guest agree; **that was reasoned from source and never executed.** (b) The in-place per-box migration (move `/felhom-data` onto `mp0`, drop the slot, verify) has **never been rehearsed even once**, so "is the box restorable at every point of it?" is currently unknown | **(a) ANSWERED 2026-08-03 (P1: PASS). (b) NOT REQUIRED — operator ruling: every node is reinstalled, none migrated** | blocks R-165 landing safely | **Filed because this project's own record is that FOUR production designs specced against unvalidated mechanisms were all wrong** — which is exactly why R-165's own spike refused to design. Both are one command on a **Tier-0** box (D-d: both demo boxes are disposable). (b) is only required work if Peti's box turns out to need migrating rather than reinstalling — the hub cannot answer that (M5: `peti-felhom` exists as a customer with **no host in the register**), so it is the operator's input. **ID established free:** `grep -ro "R-176\b" documentation/ *.md` → 0 hits **UPDATE 2026-08-03.** **(a) is measured and passed** — `audits/SPIKE-r165-phase0-2026-08-03.md` P1: a real pre-merge archive (`mp0+mp1`, confirmed from its own vzdump log) restore-tested on demo-hp, `pass: true`, `mount_parity: ok`, 84 s, with `mountParity` untouched. One limit stated rather than glossed: it ran with the pre-merge agent because the merged one did not exist yet, and the comparison is archive-vs-its-own-restore which never consults the host layout — **re-run it once against agent v0.120.0**, which is one command. **(b) is withdrawn, not deferred:** the operator ruled that every node is REINSTALLED rather than migrated in place (both demo boxes are Tier 0; the colleague's box carries none of our customer data and is clean-installed in a few weeks), so the in-place migration rehearsal has no consumer. Recorded explicitly rather than silently skipped | CC | -| **R-184** | **Nothing prevents the hub from vouching an agent version that was never released.** The R-115 gate proves every RELEASED version is installable, but it works from git tags — so a hub artifact-manifest entry naming a version with no tag and no package is invisible to it. The installer would then die at step 5 on a virgin machine, as root | **READY (S) — NEW 2026-08-03** | — | **Filed BECAUSE the R-115 gate deliberately does not cover it, rather than leaving the gap unstated.** CI cannot check it: the hub's `/api/v1/artifacts/` answers **401** without a per-customer retrieval passphrase and Gitea's package **listing** api answers **401** without a token (both measured 2026-08-03, P-C), so a credential-free gate can ask *"is this version installable"* but never *"which version is vouched"*. **Two shapes, and the second is better:** (a) give CI a hub credential — expands what CI can reach, and is the operator's call not a gate author's; (b) **validate at vouch time, in the hub**: the operator UI's Day-0 artifact form refuses a version whose package is not downloadable. (b) fails closed at the moment of the decision, needs no new credential anywhere, and puts the check where the mistake is actually made. **Exposure is low and should be said so:** vouching is a deliberate operator action against a version they have just released, and R-115's release path now makes released-but-unpublished nearly impossible. This is the residue, not the main risk | CC | -| **R-180** | **`--archive-storage` is accepted without checking the agent's token will ever be granted on it, and the failure lands at step 8/8 — after the root@pam password has already been rotated.** `felhom-host-install.sh` validates the archive storage EXISTS (`pvesm status --storage`, `:1583`) and that the golden volid RESOLVES on it (`:1661`), both in pre-flight. It never checks that storage against the ACL set it is about to grant, which is the fixed default `local local-lvm felhom-pbs` (`--acl-storages`, which `runbooks/day0-install.md` tells the operator **not** to pass). A storage outside that set therefore passes every pre-flight gate and dies at the last step | **READY (S) — NEW 2026-08-03** | — | **Hit live on demo-hp 2026-08-03** during R-178 Phase A, self-inflicted and therefore a clean demonstration: the golden was staged on `felhom-backup` (the enrolled NVMe, where the box's vzdumps live) and `--archive-storage felhom-backup` passed. Pre-flight passed; steps 1–7 ran; step 8 returned `reconcile: bring-up restore: proxmox: POST /nodes/felhom-host/lxc -> HTTP 403: permission denied at /storage/felhom-backup (missing privilege Datastore.AllocateSpace)`. **The cost is the ORDER, not the error** — by the time it fires, step 2 has minted the PVE token, step 4b has **rotated root@pam and vaulted it** (so the old console password is already dead), and step 5 has installed the agent. Recovery was `--resume` after moving the golden to `local`, which worked cleanly. **This is statically checkable in pre-flight**: `ARCHIVE_STORAGE ∈ PVE_STORAGES` is a one-line assertion over two variables both known at `:1583`. Same class as R-29 — the checkable thing that nothing checks | CC | -| **R-179** | **`--uninstall` leaves the NAS network-storage systemd units behind, with the automount in `failed` state and the parent bind still mounted.** The teardown's residue-diff provenance (`day0-install.md` Part E: *"a full-filesystem diff against the pre-install baseline showed zero Felhom-named leftovers"*) is from **v1.9.1**, which predates the NAS network-storage feature. A box that has ever had a network share configured keeps `/etc/systemd/system/mnt-felhom\x2ddrives-.mount` and `.automount` after a full uninstall | **READY (S) — NEW 2026-08-03** | — | **Observed on demo-hp 2026-08-03** after `--uninstall --vmid 9201`: `mnt-felhom\x2ddrives-Felhom\x2dShare.automount` **loaded failed failed**, its `.mount` `loaded inactive dead`, and `mnt-felhom\x2ddrives.mount` still `active mounted` — the uninstall's own output had warned `/mnt/felhom-drives/Felhom-Share is busy — NOT forcing` and `/mnt/felhom-drives root bind left mounted`, which is correct behaviour (it never forces an unmount) but is not teardown. Cleared by hand before the reinstall: stop both units, remove both unit files, `daemon-reload`, unmount the autofs then the parent. **NEGATIVE CONTROL, same day:** demo-felhom's uninstall left **nothing** (`ls /etc/systemd/system | grep -i felhom` → only the unrelated `felhom-bootstrap.service`; no felhom mounts) — because that box had no network share configured. **So the residue is conditional on the feature having been used, which is exactly why a diff taken on a box that never used it reported clean.** `felhom-bootstrap.service` is NOT residue — it is the ISO first-boot unit, `disabled`+`inactive`, exactly-once and already fired | CC | -| **R-177** | **There is no operator-triggerable "run the fill check now" path.** `fill-watch` is reachable only on its daily 03:30 schedule plus the once-at-startup run added in controller v0.191.1 — so the only way to exercise it on demand is to restart the controller | **READY (S) — NEW 2026-08-02** | — | **Noticed while live-validating R-167 on 9201, not by a failure.** It cost a controller restart per observation during validation, and it costs the same on a support call: after a customer frees space, nobody can confirm the warning has cleared without restarting their controller or waiting until 03:30. **Partially mitigated already** — v0.191.2 makes every run log a positive observable (`checked N filesystem(s), M unreadable/skipped, K notification(s)`), so at least a run that DID happen is visible; the gap is triggering one. The scheduler has `GetJobs` but no run-now, so this is a general affordance, not a fill-watch one — **scope it as "run a named scheduler job now", operator-gated.** **ID established free:** `grep -ro "R-177\b" documentation/ *.md` → 0 hits | CC | -| **R-173** | **The hub's SQLite PVC is excluded from every Longhorn backup job.** `pvc/hub-data` carries `recurring-job-group.longhorn.io/default: disabled`, and `backup-daily` + `backup-weekly` (04:00 / Sun 05:00) are the ONLY recurring jobs and both target the `default` group — so the 128 MB `/data/hub.db` has **no volume-level backup**. That database holds `host_recovery` (every managed box's break-glass root password), `host_escrow` + `host_escrow_superseded` (escrow custody), `host_pbs_secrets`, `customer_configs`, `dr_recipe` and the wg endpoints/peers — i.e. the material several documented recovery routes depend on | **READY (M) — NEW 2026-08-02** | — | **Noticed while checking the blast radius of the R-172 WAL change, not by a failure** — the WAL work needed to know who copies this file, and the answer turned out to be nobody on a schedule. **Establish before designing:** (a) whether the exclusion is deliberate (a 1 Gi RWO Longhorn volume snapshotting a 128 MB SQLite file is cheap, so the label looks like a leftover rather than a decision) and by whom; (b) whether anything else backs it up out-of-band that this census missed — the `_recovery-inventory-2026-07-28.md` records a MANUAL hot copy, which is not a backup. **When it is designed, it must be WAL-aware** (R-172): a volume snapshot of a live WAL database is crash-consistent and replays on open, which is fine, but any file-level copy must take `hub.db-wal` too or it silently loses the newest writes. **Grep establishing the ID was free:** `grep -ro "R-173\b" documentation/ *.md` → 0 hits | CC | -| **R-161** | **The volume-persistence gate is enforced by CONVENTION, not automatically.** The catalog repo has no CI of any kind (`.gitea/workflows`, `.github`, drone/woodpecker — searched, none exists). | **REDUCED SCOPE — open** (operator ruling 2026-08-02) | a second person touching templates | **RULED. Both obvious enforcement points were rejected for measured reasons.** *Controller-side at template load:* rejected because such a check can only read the file, and a static audit of all 53 templates reports the catalog clean **including papra** — **it would pass on the exact defect it exists to catch**; the property is decidable only at runtime. *CI:* rejected for now — neither repo has any, and there are no users yet. **SHIPPED instead** (`app-catalog-felhom.eu` `fd7747d`): `scripts/catalog_gates.py`, ONE entry point running all three gates, non-zero exit on any failure, **mandated in the catalog's `CLAUDE.md`** the way `site_gates.py` is. Rationale for the record: of this project's gates, the only ones that ever get run are those with a single entry point named in a CLAUDE.md — `site_gates.py` is run, R-29's three orphans are named nowhere and have stopped nothing. **What remains open is only the automatic half:** this is convention, run by a person, and that is sufficient while one person touches templates. Revisit when a second does **UPDATE 2026-08-02:** `catalog_gates.py` gained `--fast` (gate 1 only — the network and runtime gates are deliberately NOT in a hook: a push that pulls images and starts containers gets bypassed within a week and the bypass becomes the habit) and `.githooks/pre-push` now runs it. **The automatic half now has a designated successor row: R-168** (Gitea Actions runner). This row stays open at its reduced scope — the runtime gate remains a deliberate periodic run **UPDATE 2026-08-02 (second):** the automatic half now EXISTS — R-168's runner executes `catalog_gates.py --fast` on every push to this repo (measured: run #1, `image-pin gate OK — 53 templates`, with the two runtime gates announced as skipped and their own output absent from the log). This row's *original* scope — the RUNTIME volume-persistence gate — is deliberately still NOT automatic and should stay that way: CI that pulls 53 images on every push gets disabled. It remains a periodic run | operator | +| **R-161** | **The volume-persistence gate is enforced by CONVENTION, not automatically.** The catalog repo has no CI of any kind (`.gitea/workflows`, `.github`, drone/woodpecker — searched, none exists). | **OPEN** — **REDUCED SCOPE — open** (operator ruling 2026-08-02) | a second person touching templates | **RULED. Both obvious enforcement points were rejected for measured reasons.** *Controller-side at template load:* rejected because such a check can only read the file, and a static audit of all 53 templates reports the catalog clean **including papra** — **it would pass on the exact defect it exists to catch**; the property is decidable only at runtime. *CI:* rejected for now — neither repo has any, and there are no users yet. **SHIPPED instead** (`app-catalog-felhom.eu` `fd7747d`): `scripts/catalog_gates.py`, ONE entry point running all three gates, non-zero exit on any failure, **mandated in the catalog's `CLAUDE.md`** the way `site_gates.py` is. Rationale for the record: of this project's gates, the only ones that ever get run are those with a single entry point named in a CLAUDE.md — `site_gates.py` is run, R-29's three orphans are named nowhere and have stopped nothing. **What remains open is only the automatic half:** this is convention, run by a person, and that is sufficient while one person touches templates. Revisit when a second does **UPDATE 2026-08-02:** `catalog_gates.py` gained `--fast` (gate 1 only — the network and runtime gates are deliberately NOT in a hook: a push that pulls images and starts containers gets bypassed within a week and the bypass becomes the habit) and `.githooks/pre-push` now runs it. **The automatic half now has a designated successor row: R-168** (Gitea Actions runner). This row stays open at its reduced scope — the runtime gate remains a deliberate periodic run **UPDATE 2026-08-02 (second):** the automatic half now EXISTS — R-168's runner executes `catalog_gates.py --fast` on every push to this repo (measured: run #1, `image-pin gate OK — 53 templates`, with the two runtime gates announced as skipped and their own output absent from the log). This row's *original* scope — the RUNTIME volume-persistence gate — is deliberately still NOT automatic and should stay that way: CI that pulls 53 images on every push gets disabled. It remains a periodic run | operator | | **R-162** | **`docker diff` is the gate's only witness, and its failure mode is quiet.** The gate's power comes from `docker diff` excluding mounted paths, which makes "in the writable layer" mechanically decidable — an implementation detail of the overlay driver. On a driver where `docker diff` is unsupported or lies, the gate degrades to the mount-occupancy and writability legs **and would not say so**. | **WATCHING** — a limitation, not a defect | — | It **fails closed**: the canary self-test would stop reporting BROKEN and the gate would then refuse to report at all. What is wrong is the message — it would blame the prober rather than the driver. Revisit only if a non-overlay storage driver ever ships | CC | | **R-164** | **C2's chain: the DB volume tar cannot be dropped until a SOUND dump predicate exists.** The unit carries both a volume tar and a SQL dump; the restore uses **both** — the dump is authoritative and replayed *after* the tar so it WINS (F17), with only the DB service up (R-47) — `internal/backup/restore_unit.go:262-266`. Dropping the DB container's tar would halve DB-app units **and** close the R-127(b) initdb-skip password trap (restored PGDATA ⇒ `POSTGRES_PASSWORD` ignored). | **BLOCKED** — on the predicate | a dump-validity predicate that is not `accounts has rows` | **The obvious gate is DEAD, measured:** `ValidateDump` warns when the `accounts` table is empty, and that warning was **correct** — the live DB genuinely had 0 accounts, and seeding one stopped the warning and put the row in the dump. But **a fresh appliance legitimately has zero accounts**, so promoting that predicate to a gate would **block every new customer's first backup**. Order: (1) a sound predicate — dump vs **live** per-table counts, not an absolute expectation; (2) warn→gate; (3) tar-drop. **Until (1), the tar is load-bearing** — not because dumps are bad, but because nothing can yet prove one is good. Pairs with **R-127** | CC | | **R-169** | **CI can only report, because there is no gate in the road.** Every felhom repo pushes straight to `main` with no pull request, so there is no merge for a status check to stand at. R-168's runner therefore notices a broken push *after* it has landed | **WAITING-ON-OPERATOR** (a working-style decision, not a defect) | an operator ruling | Making CI *blocking* requires two things this task deliberately did NOT do, because both change how the operator works and that is not a task's call: **(a)** branch protection on `main`, and **(b)** a pull-request workflow instead of direct-to-`main` pushes. The cost is real — every change would need a PR, which for a single-operator project may be worse than the disease. **The current arrangement is two nets, and it is not nothing**: `.githooks/pre-push` REFUSES locally, and R-168's runner NOTICES when that hook was skipped or was never armed in a clone, and emails. The honest gap is the window between a `--no-verify` push landing and the operator reading the alarm. Decide only if that window ever actually costs something | operator | +| **R-173** | **The hub's SQLite PVC is excluded from every Longhorn backup job.** `pvc/hub-data` carries `recurring-job-group.longhorn.io/default: disabled`, and `backup-daily` + `backup-weekly` (04:00 / Sun 05:00) are the ONLY recurring jobs and both target the `default` group — so the 128 MB `/data/hub.db` has **no volume-level backup**. That database holds `host_recovery` (every managed box's break-glass root password), `host_escrow` + `host_escrow_superseded` (escrow custody), `host_pbs_secrets`, `customer_configs`, `dr_recipe` and the wg endpoints/peers — i.e. the material several documented recovery routes depend on | **READY (M) — NEW 2026-08-02** | — | **Noticed while checking the blast radius of the R-172 WAL change, not by a failure** — the WAL work needed to know who copies this file, and the answer turned out to be nobody on a schedule. **Establish before designing:** (a) whether the exclusion is deliberate (a 1 Gi RWO Longhorn volume snapshotting a 128 MB SQLite file is cheap, so the label looks like a leftover rather than a decision) and by whom; (b) whether anything else backs it up out-of-band that this census missed — the `_recovery-inventory-2026-07-28.md` records a MANUAL hot copy, which is not a backup. **When it is designed, it must be WAL-aware** (R-172): a volume snapshot of a live WAL database is crash-consistent and replays on open, which is fine, but any file-level copy must take `hub.db-wal` too or it silently loses the newest writes. **Grep establishing the ID was free:** `grep -ro "R-173\b" documentation/ *.md` → 0 hits | CC | +| **R-177** | **There is no operator-triggerable "run the fill check now" path.** `fill-watch` is reachable only on its daily 03:30 schedule plus the once-at-startup run added in controller v0.191.1 — so the only way to exercise it on demand is to restart the controller | **READY (S) — NEW 2026-08-02** | — | **Noticed while live-validating R-167 on 9201, not by a failure.** It cost a controller restart per observation during validation, and it costs the same on a support call: after a customer frees space, nobody can confirm the warning has cleared without restarting their controller or waiting until 03:30. **Partially mitigated already** — v0.191.2 makes every run log a positive observable (`checked N filesystem(s), M unreadable/skipped, K notification(s)`), so at least a run that DID happen is visible; the gap is triggering one. The scheduler has `GetJobs` but no run-now, so this is a general affordance, not a fill-watch one — **scope it as "run a named scheduler job now", operator-gated.** **ID established free:** `grep -ro "R-177\b" documentation/ *.md` → 0 hits | CC | +| **R-179** | **`--uninstall` leaves the NAS network-storage systemd units behind, with the automount in `failed` state and the parent bind still mounted.** The teardown's residue-diff provenance (`day0-install.md` Part E: *"a full-filesystem diff against the pre-install baseline showed zero Felhom-named leftovers"*) is from **v1.9.1**, which predates the NAS network-storage feature. A box that has ever had a network share configured keeps `/etc/systemd/system/mnt-felhom\x2ddrives-.mount` and `.automount` after a full uninstall | **READY (S) — NEW 2026-08-03** | — | **Observed on demo-hp 2026-08-03** after `--uninstall --vmid 9201`: `mnt-felhom\x2ddrives-Felhom\x2dShare.automount` **loaded failed failed**, its `.mount` `loaded inactive dead`, and `mnt-felhom\x2ddrives.mount` still `active mounted` — the uninstall's own output had warned `/mnt/felhom-drives/Felhom-Share is busy — NOT forcing` and `/mnt/felhom-drives root bind left mounted`, which is correct behaviour (it never forces an unmount) but is not teardown. Cleared by hand before the reinstall: stop both units, remove both unit files, `daemon-reload`, unmount the autofs then the parent. **NEGATIVE CONTROL, same day:** demo-felhom's uninstall left **nothing** (`ls /etc/systemd/system | grep -i felhom` → only the unrelated `felhom-bootstrap.service`; no felhom mounts) — because that box had no network share configured. **So the residue is conditional on the feature having been used, which is exactly why a diff taken on a box that never used it reported clean.** `felhom-bootstrap.service` is NOT residue — it is the ISO first-boot unit, `disabled`+`inactive`, exactly-once and already fired | CC | +| **R-180** | **`--archive-storage` is accepted without checking the agent's token will ever be granted on it, and the failure lands at step 8/8 — after the root@pam password has already been rotated.** `felhom-host-install.sh` validates the archive storage EXISTS (`pvesm status --storage`, `:1583`) and that the golden volid RESOLVES on it (`:1661`), both in pre-flight. It never checks that storage against the ACL set it is about to grant, which is the fixed default `local local-lvm felhom-pbs` (`--acl-storages`, which `runbooks/day0-install.md` tells the operator **not** to pass). A storage outside that set therefore passes every pre-flight gate and dies at the last step | **READY (S) — NEW 2026-08-03** | — | **Hit live on demo-hp 2026-08-03** during R-178 Phase A, self-inflicted and therefore a clean demonstration: the golden was staged on `felhom-backup` (the enrolled NVMe, where the box's vzdumps live) and `--archive-storage felhom-backup` passed. Pre-flight passed; steps 1–7 ran; step 8 returned `reconcile: bring-up restore: proxmox: POST /nodes/felhom-host/lxc -> HTTP 403: permission denied at /storage/felhom-backup (missing privilege Datastore.AllocateSpace)`. **The cost is the ORDER, not the error** — by the time it fires, step 2 has minted the PVE token, step 4b has **rotated root@pam and vaulted it** (so the old console password is already dead), and step 5 has installed the agent. Recovery was `--resume` after moving the golden to `local`, which worked cleanly. **This is statically checkable in pre-flight**: `ARCHIVE_STORAGE ∈ PVE_STORAGES` is a one-line assertion over two variables both known at `:1583`. Same class as R-29 — the checkable thing that nothing checks | CC | +| **R-184** | **Nothing prevents the hub from vouching an agent version that was never released.** The R-115 gate proves every RELEASED version is installable, but it works from git tags — so a hub artifact-manifest entry naming a version with no tag and no package is invisible to it. The installer would then die at step 5 on a virgin machine, as root | **READY (S) — NEW 2026-08-03** | — | **Filed BECAUSE the R-115 gate deliberately does not cover it, rather than leaving the gap unstated.** CI cannot check it: the hub's `/api/v1/artifacts/` answers **401** without a per-customer retrieval passphrase and Gitea's package **listing** api answers **401** without a token (both measured 2026-08-03, P-C), so a credential-free gate can ask *"is this version installable"* but never *"which version is vouched"*. **Two shapes, and the second is better:** (a) give CI a hub credential — expands what CI can reach, and is the operator's call not a gate author's; (b) **validate at vouch time, in the hub**: the operator UI's Day-0 artifact form refuses a version whose package is not downloadable. (b) fails closed at the moment of the decision, needs no new credential anywhere, and puts the check where the mistake is actually made. **Exposure is low and should be said so:** vouching is a deliberate operator action against a version they have just released, and R-115's release path now makes released-but-unpublished nearly impossible. This is the residue, not the main risk | CC | +| **R-190** | **A storage ACL that demonstrably WORKED in the morning was gone by mid-morning, and nothing recorded its removal.** On demo-felhom, a `vzdump` by `felhom-agent@pve!agent` with `--storage felhom-backup` completed **OK at 04:44:50 CEST 2026-08-03** (task log read in full). From **09:24:56** the same path returned `HTTP 403 … missing privilege Datastore.Allocate at /storage/felhom-backup`, six times through the day, until the grant was re-applied by hand at 18:54. By ~14:50 `pveum acl list` showed **no row at all** for that path | **NARROWED** — **MITIGATION SHIPPED 2026-08-04** (agent **v0.124.0 → v0.124.1**) — **MECHANISM STILL OPEN** | — | **Why this is not just R-185 restated:** R-185's mechanism (the installer's Scenario-F arm resolves a pre-existing target without granting) explains a box that NEVER had the grant. This box HAD it and lost it, inside five hours, with the machine up throughout. **Ruled out, each by measurement:** a host reinstall (`uptime` = 12 days); any `pveum`/ACL/`user.cfg` activity in syslog between 04:00 and 10:00 (none); any ACL entry in `/cluster/log` (none). **Correlated, not established:** `host_leaf_changed` at 09:15 and `controller_started` at 09:19 — guest 9201 was reprovisioned nine minutes before the first 403. PVE removes ACLs at `/vms/` when a guest is destroyed (`AccessControl::remove_vm_access`, the F-LEAK mechanism); whether any path can take a `/storage/` row with it has NOT been established and is the first thing to check. **Why it matters more than the grant did:** a permission that can vanish silently makes every ACL-based guarantee on these hosts provisional, and the agent's new store-grant probe (v0.123.0) now detects the STATE but says nothing about the TRANSITION. **Worth pairing with:** whether the probe should report a grant it once had and no longer has as a distinct, louder signal than one it never had **THE ROW NOW REFLECTS THE MITIGATION, NOT THE CAUSE — stated plainly because the two are different things.** The box is resilient; the loss is still unexplained. **Mitigation:** when the store-grant probe finds the grant absent, the agent runs the EXISTING root wrapper (`felhom-backup-target-apply grant `) and **re-reads once** to confirm — the pbsdr R-22 self-grant shape, including its restraint. **No new privileged surface:** the sudoers vector `grant *` already covers any storage id (confirmed in `configs/felhom-agent.sudoers`, not assumed), and the verb already grants BOTH user and token. The verb existed, was permitted, and had only ever been called at storage CREATION — the *built but never wired* shape in a verb rather than a seam, this project's seventh instance. Bounded at one attempt per tier per hour (a storage can be unreadable for reasons an ACL cannot fix; re-granting every cycle is a repair loop wearing a fix's clothes). **THE RECORD IS THE HALF THIS ROW IS ABOUT, and v0.124.0 got it wrong in production while every unit test passed.** It reported degraded for *one cycle* — meaning the probe call that repaired. But `probeAll` is invoked INDEPENDENTLY by the self-check log and by the collector building a host-report: on the box the repairing call was the log's (`09:39:34`, journal shows the repair and `degraded=1`) and the report three seconds later found the grant present and sent **`ok`**. The agent's journal had the record, the hub had nothing, and the operator would have learned nothing — the exact silence this row exists for, re-created inside its own mitigation. **v0.124.1** replaces it with a latch on TIME (20 min > the 900 s report interval), so at least one report must carry it. **PROVEN LIVE, twice, on demo-felhom** (grant deleted by hand, both rows): agent logs `store-grant: GRANT WAS MISSING AND HAS BEEN SELF-REPAIRED — investigate the loss (R-190) target=felhom-backup privilege=Datastore.AllocateSpace action="felhom-backup-target-apply grant felhom-backup" confirmed_by=re-read`; the ACL rows return; and on v0.124.1 the **host-report at 08:00:30Z carried `status=degraded`** with the explanation and the hub raised `agent_capability_degraded` and **e-mailed the operator** at 08:00:40. Nothing new was built to carry it — the hub's existing ok→degraded→ok edge is the channel, and the text rides `Feature` because that is the field the hub interpolates into the e-mail (`Reason` does not travel). **PART 3 — the single bounded pass at the mechanism, with the negatives named.** The lead §8.6 nominated is real as a CLASS and is documented in our own installer: *"`pveum user token remove` purges the token's ACL, so re-applying post-rotate is mandatory"*. **It does NOT fit this box.** A rotation purges ALL of the token's ACLs and mints a NEW secret; demo-felhom's token still authenticates with the same secret (`--selftest` OK), it retained its other three storage grants throughout, and only `felhom-backup` was refused. No installer run is evidenced (no 2026-08-03 install log; host uptime 12 days at the time). Previously ruled out and unchanged: a host reinstall, any `pveum`/ACL/`user.cfg` activity in syslog 04:00–10:00, any cluster-log ACL entry. **Ruled out on THIS box; NOT ruled out fleet-wide** — any installer run still purges and re-grants only the hardcoded `PVE_STORAGES` set, though installer 1.24.0's reuse-arm fix now re-grants the backup target on that path. **A NEW OBSERVATION FROM THE LIVE RUNS, relevant to the timeline:** PVE **caches permissions** — after deleting both ACL rows the probe still read the privilege as present for ~40 s in one run and ~16 min in another. Detection is only as prompt as that cache, and a cache expiry could equally explain why a box kept working for hours after a grant was removed. → **R-194**. **The alert pair CLOSED on its own at 10:30:40** (`degraded → ok`, `agent_capability_recovered`) once the 20-minute latch expired — one lost grant, one e-mail, one recovery, nothing further. | CC | +| **R-194** | **PVE's permission cache delays every grant-state verdict by an unknown amount, so "the agent can read it" and "the ACL exists" are not the same measurement.** Observed twice while validating R-190's self-repair on demo-felhom 2026-08-04: both ACL rows for `/storage/felhom-backup` were deleted, and `GET /access/permissions` continued to report `Datastore.AllocateSpace` present — for **~40 s** in one run and **~16 minutes** in another. During that window the capability probe reads healthy and the self-repair does not fire | **OPEN** | — | **Why it matters beyond the delay:** it puts a floor under how fast a lost grant can be noticed, it makes any single permission read a lagging indicator, and — the interesting part — **it is a candidate contributor to R-190's own timeline**: a grant removed at an unknown moment could keep working until a cache expiry, which is exactly the shape of *worked at 04:44, refused at 09:24*. That does not explain what removed it, but it may explain when the refusal SURFACED, and the two have been treated as the same instant. **Not a defect in our code** — it is PVE behaviour, and the mitigation already tolerates it (the repair fires on the next probe after the cache clears). **What is worth deciding:** whether the store-grant probe should ALSO consult the storage content listing as a second signal, since that appeared to reflect the loss immediately (`{"data":[]}` while the permission read still said present) — two signals disagreeing is itself information, and today only one of them is read | CC | +| **R-200** | **The DR password-injection seam has a handler, a route and tests — and no form.** `POST /backup/offbox/inject-password` is routed (`controller/internal/web/server.go:510`) to `offboxInjectPasswordHandler` (`offbox_handlers.go:174-196`) → `InjectOffboxPassword` (`backup/offbox.go:541`). **No template in the repository contains that path or any form posting to it** (grep over `internal/web/templates/`: one unrelated hit, an XSS comment) | **NARROWED** — **PLUMBING COMPLETE** (controller **v0.196.0**); **the FORM is not built — still open** | — | **The tenth instance of this project's built-but-never-wired class, and the exact shape `CLAUDE.md` and `felhom.eu/CLAUDE.md`'s seam-wiring rule were written for:** handler tests that POST directly (`offbox_escrow_test.go:167,180`) prove nothing about reachability. To use the only implemented recovery seam today, a person must hand-craft an authenticated POST with a session cookie and CSRF token. **Note the layering while fixing it:** this form takes a **64-hex repo password**, not a recovery code (`offboxRepoPwPattern`, `offbox.go:543`) — they are different secrets at different layers, and the operator's 2026-08-04 ruling asks for a form that takes **R**. Build the R form and treat this one as the operator/DR fallback it was written as, but **ship it with a render test per branch of whatever gate it sits behind**. Source: `audits/RECON-offsite-dr-chain-2026-08-04.md` §3 link 9 **THE DIAGNOSTIC HALF IS DONE AND IT ANSWERED THE QUESTION.** `--recover-offsite-check` is a `docker exec` escape hatch in the shape of `--print-reset-code`: R on **STDIN** (never argv, never `ps`, never shell history, never a transcript), fetch+unseal via the agent, and a verdict of **two sha256 hashes**. **It compares and never installs** — the recovered password is not written to `offbox/repo_password`; a test asserts the data dir is byte-unchanged and its red-proof (adding the install call) fails it. Confirmed live: `repo_password` mtime still `2026-08-03 07:18:02` after the successful check at `2026-08-04 11:49`. **Exit codes are load-bearing** — `0` match, `2` a clean MISMATCH, `1` a step failed; "it failed" and "it worked and disagreed" must never share a status because only one is a finding about the system. A box with no local password reports distinctly (the rebuilt-box shape, where the next step is to INSTALL rather than compare). **WHAT IS NOT BUILT, deliberately:** no card, no form, no preview, no customer-facing text — building an interface on top of a chain nobody had walked is how the preceding three weeks went wrong. **What remains for this row:** link 9 (the recovered password placed so `WriteOffboxSecrets` keeps it) and the customer-facing shape the operator ruled on 2026-08-04 (yell → R form → preview → proceed), which is now priced against a chain that exists rather than one that is assumed **PART 0 SHIPPED 2026-08-04 (v0.196.0):** `--recover-offsite-install` is the sibling of the check — same fetch/unseal path, same STDIN discipline for R — and it **places** the recovered password via `InjectOffboxPassword`. The confirmation is a SECOND invocation (`--confirm-install`): without it, both hashes print and nothing is written, so the operator sees the comparison before a write is possible. Three outcomes named distinctly: **installed** (no local password — the rebuilt-box shape), **unchanged** (identical key, nothing written), **refused** (a DIFFERENT key present — installing would clobber the key the current repository is encrypted under; exit 2, no force offered). It re-reads the file after writing rather than trusting the call. **Red-proof observed:** removing the confirmation gate makes the dry run write the password. The R-persistence test carries a **positive control** (a planted copy found, then removed and not found). **NOT YET EXERCISED AGAINST A LIVE RECOVERY** — the R-201 drill halted before step 9, so this is unit-proven only. **What remains for this row:** the customer-facing shape the operator ruled on 2026-08-04 (yell → recovery-code form → preview → proceed) | CC | | **R-206** | **The build-cache cap and the weekly prune exist only as a hand-edited `/etc/docker/daemon.json` on DooPlex — not in Ansible, so a rebuild loses them.** The `node_housekeeping` role must also carry the prune, which today it is forbidden to run | **READY (M) — NEW 2026-08-05** | — | **The spike validated the recipe; this row builds it.** Three parts. **(a) Template `/etc/docker/daemon.json`** with the **`policy` array** form — **the flat form (`{"gc":{"reservedSpace":…}}`) is SILENTLY IGNORED**, measured: the daemon starts, logs nothing, and `docker buildx inspect` still reports the built-in defaults. **The oracle is `docker buildx inspect`, never `dockerd --validate`** — the validator returned `configuration OK` for a bogus key AND for the config that then **crashed the daemon** (`filter` takes one value per policy entry, not an array; `error initializing buildkit: filters expect only one value`). **(b) Narrow the role's Docker ban** (`node-housekeeping.sh.j2:10-14`) to permit exactly `docker builder prune -af` and nothing else — the ban's stated premise ("Docker here runs only unrelated jarr-* dev containers") is obsolete: the growth is Felhom Go build cache. **The measured prune is SYNCHRONOUS** (150.35 GB back at t+0, two consecutive polls <1 MB apart within 60 s) — **unlike containerd's image GC, so it needs no `settle_imagefs` equivalent**, but it MUST measure the filesystem rather than trust the command: `prune` claimed **156.9 GB** and the filesystem returned **150.35 GB**, the 6.5 GB gap being layers still shared with images. **(c) A restart-safety note in the role:** a bad `daemon.json` takes the daemon down AND leaves the `unless-stopped` dev containers stopped — they needed a manual `docker start` — so the role must restart-and-verify, not validate-and-assume. Recipe + every measurement: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` | CC | | **R-207** | **`DRY_RUN=1` on `node-housekeeping.sh` is NOT non-mutating — it destroys the metric history it is supposed to let you inspect** | **READY (S) — NEW 2026-08-05** | — | `write_metrics()` (`node-housekeeping.sh.j2:119-150`) has **no `DRY_RUN` guard at all** — `DRY_RUN` appears in it only inside a log line — and it is called from an **unconditional `EXIT` trap** (`:150`). A dry run therefore atomically renames over the live node_exporter textfile, overwriting `node_housekeeping_last_success_timestamp_seconds` with *now* and `reclaimed_bytes` with ~0 — **resetting the staleness clock `HousekeepingStale` watches and erasing the 8-week reclaim history**. Confirmed by reading the source in both the 2026-08-05 audit and this spike; **neither run executed it**, so the history survives. Fix: guard `write_metrics` on `DRY_RUN`, or have the trap skip it. Pairs naturally with R-206 (same file, same role) | CC | | **R-208** | **Every Felhom Go build re-downloads its modules because `ARG VERSION` sits ABOVE the module-download layer — ~440 MB of dead cache per build, 90.5 GB of the 157 GB** | **READY (S) — NEW 2026-08-05, ROOT CAUSE PROVEN** | — | **Measured, not inferred.** All **208** retained `go mod download` records carried **`Usage count: 1`** — not one was ever reused in ~a month of builds. The mechanism was isolated by four controlled builds: an unchanged tree rebuilt with the **same** `--build-arg VERSION` → `RUN go mod download` **CACHED**; the **same** tree with a **new** `VERSION` → **executed**. `COPY go.mod ./` stays CACHED either way, which is the tell: a `COPY`'s key is content-based, while a `RUN`'s key includes the stage **environment**, and `ARG VERSION`/`ARG GIT_COMMIT` are declared *before* the download in `felhom-controller/controller/Dockerfile`. Since every real build passes a fresh version, the layer is invalidated **every single time**. **`felhom.eu/hub/Dockerfile` has the identical defect** (`ARG VERSION`/`ARG BUILD_TIME` above `COPY go.mod go.sum*` → `RUN go mod download`) — and because both Dockerfiles produce byte-identical `buildx du` description strings, the 208 records are a COMBINED count and must not be attributed to one project. **Fix shape (one line each, not applied here):** move the `ARG VERSION`/`ARG GIT_COMMIT`/`ARG BUILD_TIME` declarations down to just above the final `go build`. **Worth more than the cap and the move combined** — the cap bounds the symptom, this removes the source. `build.sh`'s `rm -rf` + `cp -a` and its host-side `go mod tidy` were **ruled out by fingerprinting**: the tree is byte-identical across runs and `tidy` is a no-op | CC | -| **R-209** | ~~**Should the containerd store move to SSD2 at all?**~~ | **EXECUTED 2026-08-05 on operator ruling — reboot validation DEFERRED → R-209a** | — | **Operator ruled "proceed" having read the pre-analysis; CC's `storageReserved` condition was applied with it.** Moved with **zero loss, verified on four independent observables BEFORE the original was touched** (550,891 entries = 550,891; **448 = 448 `trusted.overlay` xattrs**; 37,243 = 37,243 hardlinks; byte-identical `meta.db` sha256) and again after (identical image/tag/volume ID **sets**, cache 2.782 GB/38 records, pg 4 DBs / 31 tables / 175,135,767 B, redis 2437). **`-X` is load-bearing** — overlayfs stacking rides `trusted.overlay.*`. End-to-end proof was a **real build** on the relocated store, `rc=0`. **k3s was never at risk and this was established before stopping anything:** it runs a separate containerd, so Gitea, the registry, the hub, PBS, Longhorn and ~160 pods stayed up; only the two `jarr-*` dev containers were affected. `storageReserved` on SSD2 **0 → 80 GB**, still `Schedulable=True` at 76.34%. **A TRAP was found while proving the guard, and it is the reusable part: `RequiresMountsFor` on a path with NO mount unit is a SILENT NO-OP** — containerd started normally against an absent-but-unmounted path, so a typo'd guard buys nothing and says nothing (the *built-but-never-wired* shape again). The guard was therefore verified positively at the unit level (`Requires=` **and** `After=mnt-ssd_2.mount` on both units), and refusal was then proven with a genuinely absent **device** — via a temporary synthetic `.mount` unit, because `/mnt/ssd_2` hosts 12 live Longhorn replicas and must never be unmounted, and editing `fstab` on a production host risks emergency mode at boot: `Job containerd.service/start failed with result 'dependency'`, `is-active: inactive`. **Rollback is one documented sequence** (audit §11.8); the pre-move tree is **moved aside, not deleted**. Evidence: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` §11 | — | -| **R-209a** | **The SSD2 move has NOT survived a reboot, so by this project's own standard it is not fully validated** | **WATCHING — NEW 2026-08-05** | the next DooPlex reboot | **Operator ruled explicitly: do NOT reboot DooPlex.** Uptime verified unbroken (7 weeks 6 days, since 2026-06-10). **The distinction is stated rather than glossed: the MECHANISM is proven** — the guard is wired into both units and containerd refuses to start when a required mount's device is absent — **but the CONSEQUENCE is not**: that a real boot mounts `/mnt/ssd_2` before containerd starts, in this host's actual ordering. Mount-ordering reasoning is precisely the class this project has been burned by (`RequiresMountsFor` RE-MOUNTS rather than refusing — the ep0 lesson), and `CLAUDE.md` prefers a consequence assertion over a mechanism one. **Two deliberate consequences: (1)** the rollback copy `/var/lib/containerd.pre-move-2026-08-05` (**34.3 GB on `/`**) **STAYS** until a reboot validates — which is why `/` sits at 54% and not lower; deleting it now would trade a cheap 34 GB for the only cheap way back. **(2)** validation is **automatic and needs no one to remember it**: `felhom-store-postboot-check.service` (oneshot, enabled, dry-run PASS at install) runs at **every** boot and writes `RESULT: PASS`/`FAIL` to `/var/log/felhom-store-postboot-check.log`, asserting positively that `/mnt/ssd_2` is mounted, that containerd's root is on it, that **`/var/lib/containerd` does NOT exist** (the empty-store trap), that ≥100 images are visible and that both dev containers run. **Next action: after the next reboot — planned or not — read that file; on PASS, `rm -rf /var/lib/containerd.pre-move-2026-08-05` returns ~34 GB to `/`** | operator + CC | **P3's prune already removed the urgency: `/` went 86% → 53% used and SSD1's Longhorn disk went `Schedulable=False (DiskPressure)` → `Schedulable=True` (18.85% → 50.32% available).** The move was ruled "cap then move"; the cap is in and the pressure is gone, so this is now a deliberate choice rather than a rescue. **The numbers, measured (`Crucial-SSD-240G`, `/mnt/ssd_2/data/longhorn`, `storageMaximum` 235,148,750,848):** available today **214,958,080,000 (91.41%)**; 25% floor **58,787,187,712**. Moving the whole containerd tree at steady state (~31.5 GB images + ≤30 GB cache ≈ 65 GB) leaves **63.77%, i.e. +38.8 pp above the floor — comfortably safe as measured.** **But `storageScheduled` on SSD2 is 139,586,437,120 while `df` says only 20,094,939,136 is actually used** — Longhorn has overcommitted 6.9× — and if those volumes ever inflate to their scheduled size, the same disk lands at **13.00%, i.e. 12 pp BELOW the floor → `Schedulable=False`**, which is exactly the failure that just took SSD1 out. **Recommendation: do the move only together with setting `storageReserved` on SSD2 to cover the containerd tree (~80 GB); SSD2 reserving zero while HDD2 and HDD4 each reserve 500 GB is an anomaly in its own right.** Mechanism, if it goes ahead: **containerd's `root` in `/etc/containerd/config.toml`** (the key is present but commented out) — **not** Docker's `data-root`, which would move only 0.62 GB. Guard: `RequiresMountsFor=/mnt/ssd_2` on `containerd.service` **and** `docker.service`, remembering that **`RequiresMountsFor` RE-MOUNTS rather than refusing** ([[ep0-datastore-volume-move-2026-07-27]]) — so it must be tested with a genuinely absent device, and the move is not validated until it has survived a **reboot**. Full pre-analysis: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` §P6 | operator | +| **R-209a** | **The SSD2 move has NOT survived a reboot, so by this project's own standard it is not fully validated** | **WATCHING — NEW 2026-08-05** | the next DooPlex reboot | **Operator ruled explicitly: do NOT reboot DooPlex.** Uptime verified unbroken (7 weeks 6 days, since 2026-06-10). **The distinction is stated rather than glossed: the MECHANISM is proven** — the guard is wired into both units and containerd refuses to start when a required mount's device is absent — **but the CONSEQUENCE is not**: that a real boot mounts `/mnt/ssd_2` before containerd starts, in this host's actual ordering. Mount-ordering reasoning is precisely the class this project has been burned by (`RequiresMountsFor` RE-MOUNTS rather than refusing — the ep0 lesson), and `CLAUDE.md` prefers a consequence assertion over a mechanism one. **Two deliberate consequences: (1)** the rollback copy `/var/lib/containerd.pre-move-2026-08-05` (**34.3 GB on `/`**) **STAYS** until a reboot validates — which is why `/` sits at 54% and not lower; deleting it now would trade a cheap 34 GB for the only cheap way back. **(2)** validation is **automatic and needs no one to remember it**: `felhom-store-postboot-check.service` (oneshot, enabled, dry-run PASS at install) runs at **every** boot and writes `RESULT: PASS`/`FAIL` to `/var/log/felhom-store-postboot-check.log`, asserting positively that `/mnt/ssd_2` is mounted, that containerd's root is on it, that **`/var/lib/containerd` does NOT exist** (the empty-store trap), that ≥100 images are visible and that both dev containers run. **Next action: after the next reboot — planned or not — read that file; on PASS, `rm -rf /var/lib/containerd.pre-move-2026-08-05` returns ~34 GB to `/`** **P3's prune already removed the urgency: `/` went 86% → 53% used and SSD1's Longhorn disk went `Schedulable=False (DiskPressure)` → `Schedulable=True` (18.85% → 50.32% available).** The move was ruled "cap then move"; the cap is in and the pressure is gone, so this is now a deliberate choice rather than a rescue. **The numbers, measured (`Crucial-SSD-240G`, `/mnt/ssd_2/data/longhorn`, `storageMaximum` 235,148,750,848):** available today **214,958,080,000 (91.41%)**; 25% floor **58,787,187,712**. Moving the whole containerd tree at steady state (~31.5 GB images + ≤30 GB cache ≈ 65 GB) leaves **63.77%, i.e. +38.8 pp above the floor — comfortably safe as measured.** **But `storageScheduled` on SSD2 is 139,586,437,120 while `df` says only 20,094,939,136 is actually used** — Longhorn has overcommitted 6.9× — and if those volumes ever inflate to their scheduled size, the same disk lands at **13.00%, i.e. 12 pp BELOW the floor → `Schedulable=False`**, which is exactly the failure that just took SSD1 out. **Recommendation: do the move only together with setting `storageReserved` on SSD2 to cover the containerd tree (~80 GB); SSD2 reserving zero while HDD2 and HDD4 each reserve 500 GB is an anomaly in its own right.** Mechanism, if it goes ahead: **containerd's `root` in `/etc/containerd/config.toml`** (the key is present but commented out) — **not** Docker's `data-root`, which would move only 0.62 GB. Guard: `RequiresMountsFor=/mnt/ssd_2` on `containerd.service` **and** `docker.service`, remembering that **`RequiresMountsFor` RE-MOUNTS rather than refusing** ([[ep0-datastore-volume-move-2026-07-27]]) — so it must be tested with a genuinely absent device, and the move is not validated until it has survived a **reboot**. Full pre-analysis: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` §P6 | operator + CC | | **R-210** | **Which of the 345 local images may be deleted — 131 controller tags and 62 hub tags exist ONLY on this box and are not recoverable by `docker pull`** | **WAITING-ON-OPERATOR** — NEW 2026-08-05 | an operator ruling | **Nothing was deleted; this is a list, not an action.** The registry was queried directly: `felhom-controller` has **76** tags in Gitea vs **207** locally, `felhom-hub` **45** vs **107**. The **131 + 62 local-only tags are all OLD** — controller `0.39.0`–`0.135.0` plus `v0.35.0`–`v0.39.0`, hub `0.9.0`–`0.57.0` plus `v0.7.2`–`v0.13.0` — while everything from controller `0.136.0` and hub `0.58.0` upward IS in the registry and therefore re-pullable. **Size the prize honestly before spending a decision on it:** per-tag sizes sum to 139.29 GB, but that double-counts shared layers — `docker system df` puts the **real** dedup'd image footprint at **31.02 GB with 27.02 GB reclaimable**, i.e. an order of magnitude less than the build cache P3 already returned. `docker image prune -a` would remove 343 of 345 (only `redis:7-alpine` and `postgres:16-alpine` are held by running containers). **CC's view: not worth doing for the space** — it buys ~27 GB against 199 GB now free, and its only real benefit is dropping unrecoverable clutter | operator | | **R-211** | **Prometheus has no config-reloader — a rules change reaches the pod and is never read** | **READY (S) — NEW 2026-08-05** | — | Found while verifying R-205 rather than by looking for it. The `mon-system/prometheus` Deployment runs **one** container (`prom/prometheus:v3.12.0`) with **no `configmap-reload`/`prometheus-config-reloader` sidecar**. After the ArgoCD sync the updated `node-housekeeping-alerts.yml` was present **inside the pod** (`grep -c "and on(instance)"` → 3 on the mounted symlink) while the Prometheus **rules API still served the old expression** — for **4+ minutes**, with no error anywhere. It only took effect after an explicit `POST /-/reload`. **The consequence is general, not specific to R-205: every rule edit in this repo since the stack was built has silently not applied until something happened to restart the pod** — so "committed and synced" has never meant "in force", and ArgoCD reporting `Synced/Healthy` is true and beside the point. `--web.enable-lifecycle` IS already set, so the fix is small: add a reloader sidecar watching the ConfigMap, or a `checksum/config` pod annotation so a rules change rolls the pod. **Same class as the four *built-but-never-wired* seams** — the control exists, nothing walks it | CC | - -## CAMPAIGN 12 — the class sweep, 2026-08-08 (unattended) - -Eight rows, **grouped by class so the classes are visible as classes**. Full method, controls, blind -spots and the Part-4 gating ranking: `audits/CAMPAIGN-12-class-sweep-2026-08-08.md`. Gating candidates -are in `ROADMAP.md`, not here. **C1 produced no new instance** and has no row, deliberately. - -**⚠ A correction the campaign owed to its own brief:** the task described C5's `escrow_stale` instance -as "closed individually". **It is not — it is R-247, `READY`.** The live repo is the source. - -| ID | What | State | -|---|---|---| -| **R-256** | **C2 — „A mentéskezelő nem elérhető." names no route at all.** `web/offbox_handlers.go:47` and `:181` flash this to the customer on the off-site backup surface. It states an internal component's unavailability in the operator's vocabulary („mentéskezelő" = the backup Manager object), gives no reason the customer can act on, and names no next step — not "try again in a few minutes", not "contact support", not a page to go to. **Contrast, in the same subsystem and shipped the same week:** R-252's fix reads *„Meghajtók", „Meglévő meghajtó csatolása". Utána gyere vissza ide.* — a route. **Severity is low and stated so it is not over-ranked:** the condition is a nil backup manager, which on a healthy box does not occur; this is about the copy, not a broken path. Found by the C2 sample (19 refusals on the recovery/restore/offbox surface; **~202 of the repo's 221 refusal strings were NOT examined**) | **READY** — owner Viktor | -| **R-257** | **C2 — „Az offsite tároló nincs elárvult állapotban." puts an English loanword and an internal state name in front of a Hungarian household customer, and names no route.** `web/offbox_handlers.go:270` (the Go error it mirrors is `backup/offbox.go:343`). „Offsite" is untranslated; „elárvult állapot" is the codebase's own `OffboxOrphaned()` predicate surfacing verbatim. A customer who pressed a button and got this cannot tell whether something failed, whether they did something wrong, or what to do instead. **This is a refusal that is CORRECT and fail-closed and still a dead end** — the same shape R-241 recorded for `--recover-offsite-install`. **Fix shape, not a decision:** say what the customer tried to do, why it does not apply right now, and where to look — or, since this is a state they cannot reach deliberately, do not offer the action at all | **READY** — owner Viktor | -| **R-261** | **C6 — `CountSelfBindTokens` exists so that callers can assert an invariant, and no production caller asserts it.** `hub/internal/store/selfbind.go:106-111`. Its doc comment: *"it exists so callers can assert the 'after this runs, the only live link is one we just issued — or none' invariant that the auto-mint at customer-create / RESET-completion depends on."* **Census: only its own declaration in production; the two callers are `selfbind_automint_test.go:29` and `customer_delete_test.go:510`.** Tests are not callers (the campaign's rule), so the invariant the auto-mint *depends on* is checked in the test suite and never at the moment it matters. **This is the smallest of the eight rows and is filed at its true size, because the rest of the C6 sweep found INERT dead accessors rather than defects:** `OffboxOrphanedRenamedTo` and `OffboxEscrowState` have no caller but their data reaches the card another way (the template reads the settings field directly, `backups_remote.html:80`) — **R-228 is genuinely closed, and the sweep's first reading that it had regressed was wrong.** **The more consequential C6 result is a method result and is in the report, not here:** `golang.org/x/tools/cmd/deadcode` re-finds **neither** known instance, and a planted probe measured why — it reports an unreachable exported FUNCTION and not an unreachable exported METHOD on a widely-used type, and both known instances are methods | **READY** — owner Viktor | -| **R-262** | **C7 — a comment claims a cross-repo contract is mirrored „field-for-field" and „the key-set tests guard drift"; it is two fields short, AND THE FIXTURE THE TEST READS OMITS THE SAME TWO FIELDS.** `hub/internal/api/handler.go:682-687` covers `hostBackup` **and** `hostRestoreTest`. **It is TRUE of `hostBackup`** (verified field-for-field against `agent/internal/hub/Backup`). **It is FALSE of `hostRestoreTest`:** the agent emits `mount_parity` and `mount_inventory` (`hub/report.go:432-433`, populated in production from `reconcile/restoretest.go:277-283` via `backup/runner.go:517`), and the hub has no field for either — **0 occurrences in the entire hub repo** outside the CHANGELOG. **The guard is blind in exactly the place the drift is:** `TestHostReport_GoldenContract` reads `testdata/host-report.golden.json`, the two copies of which are byte-identical as required — and **neither contains `mount_parity` or `mount_inventory` at all**, so the key sets agree on a shape that is not the shape the agent sends. A test that cannot fail on the drift it names is the R-97b lesson (*prove the consequence, not the mechanism*) landing on a contract test. **Consequence, stated precisely:** the verdict is not lost (a parity mismatch fails the test before `Pass` is set), but the hub cannot distinguish a full-fidelity restore-test pass from a boot-only one, for any agent, ever. **Fix shape, not a decision:** add the two fields and put them in the fixture — or narrow the comment to name `hostBackup` only and say plainly that `hostRestoreTest` is a subset. **Attached observation:** the same fixture carries `cpu_temp_c` and `loadavg`, which no hub struct decodes — a fixture carrying keys the receiver cannot read is the same shape one level down | **READY** — owner Viktor | -| **R-263** | **C7 — „This is the ONLY writer of `StoragePath.BackupTarget`" is false, and nothing pins it.** `settings/settings.go:1317-1319`, on `SetBackupTarget`. **`ClearBackupTarget` (`:1358-1363`) also writes the field**, 17 lines below, in the same file. **The GUARANTEE the comment protects is intact and that is why this is filed small:** the sentence continues *"registration must never set it (E-2 §3: a drive never acquires a role by appearing)"*, and `ClearBackupTarget` only ever writes `false`, so no path other than `SetBackupTarget` **grants** the role. **What is wrong is the claim as written, and the absence of anything holding it:** `backup_target_role_test.go` exercises the behaviour and asserts nothing about writer uniqueness, so if a third writer appeared tomorrow — one that granted — the comment would still read as settled and the suite would still be green. **This is the class's own definition:** an invariant asserted in prose with no test pinning it. **Fix shape:** one word (*"the only writer that GRANTS the role"*) plus a test that fails when a second granting writer appears. **Method honesty:** found in a sample of **60 of 2652** production invariant comments — the "is the only" form only, chosen because a uniqueness claim is the one form a grep can falsify. **~2592 production and all 1440 test invariant comments were NOT examined**, so C7 has the weakest coverage of the seven classes and the task's instruction to check the tests' own claims is **owed, not discharged** | **READY** — owner Viktor | - -## CAMPAIGN 12 follow-through — G-1 built, R-260 and R-247 closed, 2026-08-08 - -**The gate was built BEFORE the fixes and was seen failing on 40 fields**, captured verbatim in -`documentation/tests/wire-contract-gate-2026-08-08/BEFORE.md`. That order was the method, not -bureaucracy: the night before, `deadcode` was rejected for class C6 precisely because it was made to -prove itself first and found neither of the two defects it was meant for. - -**⚠ A COUNT THIS SESSION'S OWN PROMPT GOT WRONG, corrected against the repo rather than quoted.** The -prompt said *"465 emitted tags, eight unreachable"*. R-260's wording was "at least eight -DECISION-BEARING facts", never eight tags in total, and its own census already listed more. Measured -on the three declared wires: **210 tags checked, 51 skipped, 40 convicted.** Two prompt claims were -wrong this week and both were caught the same way. - -**Two things the gate's CONTROL caught before it was trusted**, each a defect in the instrument: - -1. **A substring false negative.** `grep -F healed_at` also matches `privsep_healed_at`, so a - genuinely dropped field read as received. R-260 named `healed_at`, so its absence from the output - was the tell. Now a whole-token regex. -2. **`dr_recipe` is not wholly opaque.** The hub stores each half as `json.RawMessage` and re-emits - nested shapes verbatim — but the TOP-LEVEL section keys are decoded by `hostHalfShape` / - `appHalfShape`, and **those are allow-lists**: a section an emitter adds is silently dropped until - named in both. That already cost `offsite_restic` (R-122). The gate is therefore opaque BELOW - depth 1, so the sections are checked; treating the whole subtree as opaque would have put R-122's - shape back outside its reach. - -**THE FORTY, BY DISPOSITION.** Full per-field reasons live in the gate's own `ALLOWLIST`, where each -entry is a claim someone can re-check. - -| # | field(s) | direction | decision | what changed | -|---|---|---|---|---| -| 1 | `oob.operator_key_configured` | agent → hub | **RECEIVE AND ACT ON IT** | decoded as a pointer; `oobDegraded` now fails when the key is absent, and the alert NAMES it. Hub v0.99.0 | -| 2 | `oob.wg_handshake_age_s`, `oob.healed_at` | agent → hub | **RECEIVE, for the message only** | decoded into `HostOOBRow` and put in the event payload; deliberately NOT in the predicate — widening the check beyond the fact that is now arriving is how a check stops being read | -| 3 | `escrow_stale` | hub → controller | **RECEIVE AND ACT ON IT** | `report.EscrowStatus.Stale`; the box tells a withheld hash from a hash-less one. Controller v0.209.0. **This is R-247** | -| 4 | `host.{cpu_temp_c,loadavg,memory_total_bytes,memory_used_bytes,uptime_seconds}`, `system.{load_avg_1,5,15, memory_total_mb, memory_used_mb, temperature_celsius, uptime_seconds}` | agent/controller → hub | **NO CONSUMER WANTED — redundant** | allowlisted: the hub decodes `cpu_percent` / `memory_percent` / `disk_percent` from the same stanzas and every threshold is expressed on those | -| 5 | `guests.spec.{disk_bytes,memory_bytes}` | agent → hub | **NO CONSUMER WANTED — redundant** | allowlisted: guest sizing is hub-owned INTENT (the manifest), not mirrored reality | -| 6 | `storage_targets.smart.model_name` | agent → hub | **NO CONSUMER WANTED — redundant** | allowlisted: a display label with no threshold on it; `smart.health` and every counter the hub bands on ARE decoded | -| 7 | `wireguard.last_handshake_age_s` | agent → hub | **NO CONSUMER WANTED — redundant** | allowlisted: wgsync reconciles from its own state, and the OOB path's own handshake age is now decoded | -| 8 | `guest_net` + its 7 children, `selfupdate_pending(_version)`, `mgmt_plane.healed_recently`, `pbs_dr.applied_at`, `restore_tests.mount_{parity,inventory}`, `config_hash`, `reporting_disabled`, `stacks`, `storage.migrated_to`, `backup.last_db_dump`, `backup.last_integrity_check` | agent/controller → hub | **NO CONSUMER TODAY, AND ONE IS ARGUABLY OWED** | allowlisted **against R-264, which stays OPEN**. Allowlisting is not deciding, and the entries say so | - -**A C7 instance found while doing it, and corrected.** `HostReport.SelfUpdatePending`'s own comment -claimed *"The hub reads an absent field as pending=false, the correct default."* The hub has no field -for it and reads nothing either way. Comment corrected in the agent (no version bump — comment only). - -**The hub's OOB test fixture was part of why this survived.** `oobReport()` omitted -`operator_key_configured` entirely, so every pre-existing scenario ran against a report shape **no -released agent produces**. Fixed, and the new tests drive the raw JSON decode boundary — a test that -builds the receiving struct by hand cannot see a field that never decodes, which is the whole class. - -| ID | What | State | -|---|---|---| - - -**Explicitly still open, untouched by this session:** R-246 (the wrong stale flag on `demo-hp` — -clearing it is an operator act hub-side), R-255, R-256, R-257, R-258, R-259, R-261, R-262, R-263, and -**C7's test-comment half**, which Campaign 12 recorded as *owed, not done* (60 of 2652 production -invariant comments sampled; none of the 1440 test comments). - -## The seed that never ran twice, and three pictures that were not true — 2026-08-08 - -Four defects of one family: something the box already knows, either thrown away or drawn as its -opposite. Agent **v0.128.0**, controller **v0.210.0**, `gates.yml` (no hub change, no hub bump). - -**R-221's writer, ESTABLISHED at `file:line` rather than assumed** — the prompt asked for this and it -was owed. `step_agent_config` (`felhom.eu/scripts/felhom-host-install.sh:2396`) renders `agent.json` -from `base = {}` unless an explicit `--preserve-from` is passed (flag `:1246`, defaulting empty at -`:256`), and writes it with `O_TRUNC` (`:2579`). **The render never writes an `escrow` section at -all** — grep over the whole heredoc returns zero hits. The pbsdr marker lives host-side -(`/pbsdr/marker.json`) and survives. So a rebuild keeps the marker and takes the key: -same descriptor, same hash, early return, seed never re-runs. **The attribution in R-221 was -correct.** A rebuild is nonetheless only the case that was measured — the same hole opens for a -hand-edited or restored config, which is the honest reason the fix is at the seam and not in the -installer. - -**The idempotent early return was KEPT**, and that is load-bearing: it stops a converged box -re-running Proxmox operations every 60 s. -`TestSeedReasserted_OnConvergedTick_WithZeroProxmoxCalls` asserts **zero** recorded runner calls on -that tick, so a "fix" that simply deleted the return fails the test. Verified by mutation. - -**The §7.3 truth table as implemented** (R-258): - -| this app's own most recent dump result | restore point | verdict | -|---|---|---| -| any of its databases failed | yes | `error` — cross | -| all clean | yes | `ok` — tick | -| none recorded (no database / no run yet) | yes | **no icon**, time only, title „Erről a mentésről nincs eredményünk." | -| any | no | no tier-1 row at all, unchanged | - -**An existing test was asserting the defect and was corrected, not deleted.** -`TestBuildAppBackupRows_Tier1FromRestorePoints` expected `"ok"` for a `FullBackupStatus` with **no -`LastDBDump` at all** — a green tick derived from nothing but a file's existence, i.e. Scenario G. -Its real subject, the `Tier1LastRun` time, is unchanged. - -**The convention is now ruled** (§7.2, `CONTEXT.md` S-39): a `…Known bool` companion beside the -figures. `ROADMAP.md` G-3 was blocked on that decision and is unblocked. - -**Six red-proofs, every one demonstrated failing and restored, each with the mutation asserted -applied.** The one that matters: Scenario A **fails against today's tree** with the intended message -— so the test tests the defect. - -| ID | What | State | -|---|---|---| -| **R-266** | **A failed root `statfs` still reaches the hub as a 0-of-0 disk, and the hub cannot tell that from an empty one.** Split out of R-259 on 2026-08-08 so that fixing the CUSTOMER-facing half could not be mistaken for fixing the wire. `report/builder.go:93-95` copies `sysInfo.DiskTotalGB` / `DiskUsedGB` / `DiskPercent` into `r.Storage[0]` (`Mount: "/"`), and those are exactly the zeros a failed `statfs` leaves behind — the controller now KNOWS the measurement failed (`SystemInfo.DiskKnown`, controller v0.210.0) and the report still does not carry it. **Deliberately not fixed here, for a reason that is now structural rather than a preference:** adding a field to that report is a change to a declared wire, which since G-1 means the receiving side must model it in the same session (`scripts/wire_contract_gate.py` refuses otherwise) — a two-repo change with a hub bump, and this session deliberately touched no hub code. **RANKED LOW, and the reason is that the consequence is bounded:** the hub bands host storage on `disk_percent`, so a failed read presents as 0% used — the *quiet* direction. It cannot raise a false "nearly full" alarm; it can only fail to raise a true one, and only while the root filesystem is unreadable, which is a state with louder symptoms of its own. **Fix shape when it is taken:** carry `disk_known` on the storage entry and have the hub's fill checker skip an unknown reading rather than band it — never treat absent as 0 | **READY** — owner Viktor | +| **R-213** | **Putting files back in place — the half the recovery screen deliberately does not do.** The screen (R-193, controller v0.200.0) unlocks the repository and LISTS what is in it; restoring is per-app and lives in the backups area, and the operator ruled the two separate on 2026-08-05: *a screen that unlocks and then offers to overwrite is two decisions wearing one button*. What is missing is the step after the listing — a customer who can now SEE their files still has to work out, per app, which restore to choose. **The operator named its requirement: a live-versus-backup comparison** — the customer must be able to see what would change before anything is overwritten | **OPEN — not started, deliberately** | the comparison design (nothing exists for it yet) | Design the live-vs-backup comparison, then the put-back flow on top of it. Do NOT fold it into the recovery screen | Operator + CC | +| **R-230** | **Three instruction/memory follow-ups deliberately left by the part-2 session (2026-08-06), each needing a decision rather than an implementation.** (a) **A ruling is owed on auto-written staleness.** The hand-written `CLAUDE.md` files are now clean of version literals and expired blocks — the gate enforces it — but `MEMORY.md`, which Claude writes and which is the LARGER half of what loads (8.4k tokens vs the root file's 6.6k), carries **21 lines with component version literals**, **5 with bare host addresses**, and an entry still reading *"demo boxes REMOTE till ~08-02"* — the same expired-TEMPORARY class the gate was built to kill, now surviving in the one file the gate's content rules do not cover. **Partly actioned 2026-08-06 (close-out), and the ruling is STILL OWED:** the **three statements that were actively false** were corrected — `R-193 decision open` (closed 2026-08-05), `demo boxes REMOTE till ~08-02` (the box answers on the home LAN), `OPEN R-25b` (shipped 2026-07-21) — and gate check 6 now **WARNs** on version literals, host addresses, expired statements and stale-open citations in the index. WARN, never FAIL: Claude writes that file between sessions, so a hard failure would refuse a human's push over a line no human typed, and the warning is read by the model that will next edit it. **The remaining 32 version literals and 4 host addresses were deliberately left** for that loop. What is still owed is the bulk-correction ruling. **Correcting the premise:** the earlier report's "three expired statements" were all FALSE POSITIVES — each matched an ISO date inside a markdown link target, i.e. a filename — while the one real expired claim carried no ISO date at all. (b) **CLOSED 2026-08-06 (close-out)** — the workspace-root `CLAUDE.md` **is now a relative symlink** to the versioned copy, so the divergence class is gone rather than policed. Check 5 learned two shapes: for a link it asserts the target resolves to a real file (**a dangling link is worse than a diverged copy — the instructions load NOTHING and there is no content left to notice is wrong**), for two files byte-identity as before, so a clone elsewhere is unaffected. **Proven, not assumed:** three fresh sessions logged `session_start` for the link path, and a fourth **with no tools at all** quoted standing rule 1 verbatim — the content reaches the model, not just the path. (c) **The spec-as-failing-test pilot**, approved in principle and not started (was R-229(d)). | **READY** — owner Viktor | — | — | operator | +| **R-231** | **`/opt/backup/scripts/` on DooPlex is unversioned host state** — found 2026-08-06 while adding the auto-memory store to the backup set. No repository tracks the scripts that protect the recovery chain, so the edit made that day (`CLAUDE_MEMORY_DIR` in `backup-config.sh`, multi-path restic call in `backup-data.sh`) exists only on the box. This is the same class the part-2 session was closing, found inside the fix for it; the change is transcribed in `felhom.eu/workspace/README.md` so it is at least *recorded*. **Two related facts, both understating current safety:** the backup destination (`/mnt/5_hdd/backup`) is on the **same physical disk** as the workspace it protects, and the DooPlex backup set has **no off-site leg** (`sync-hetzner-backups.sh` is jarrs.eu and pulls *from* Hetzner *to* DooPlex). Bringing a root-owned production backup script under version control, and deciding what installs it, is its own scoped change. | **READY** — owner Viktor | — | — | operator | +| **R-232** | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **READY** — owner Viktor | — | — | operator | +| **R-235** | **The appliance console keeps telling an already-paired box to go and pair itself.** Measured 2026-08-06 on VM 323: **25 minutes after** the operator bind, with the guest provisioned, the controller reporting 0.203.0 and the agent ONLINE, the physical console still displayed „Felhom — a doboz készen áll, és a **párosításra vár**" together with the now-spent pairing code `US3-6GP` — and, in the same panel, the promise **„Ez a képernyő magától frissül — nincs teendő a doboznál"**. It does not refresh. A customer looking at their screen is told the setup has not happened, and is told the screen would have updated if it had. Cosmetic in mechanism, not in effect: it invites the customer to re-pair a working box, or to call for help about a box that is already fine. Same family as R-234 — a surface asserting a state that stopped being true. | **READY** — owner Viktor | — | — | operator | +| **R-240** | **A backup that covered nothing calls itself „Sikeres".** On a configured box with no app selected for off-site backup, a run reports status `ok` with the warning „Sikeres — nincs mentésre jelölt alkalmazás" — *successful* immediately beside *nothing is selected*. Measured as T4 on the final walk, 2026-08-07; flagged once before (2026-08-06) and deliberately not touched then, because the task that noticed it forbade changing that path. **It is the same rhetorical shape the project has spent a fortnight removing** — R-203's *a warning beside a success is read as a success*, R-234's *„✓ Rendben" over an app that was skipped*, R-225's *unknown rendered as zero* — one notch weaker each time, and this is the weakest and last of them. The state itself is honest and must stay `ok`: an unconfigured box reporting `incomplete` forever is its own defect, pinned by a test. **The defect is the word „Sikeres", not the verdict.** Wording such as „Nincs mentésre jelölt alkalmazás — ez a futás semmit nem mentett" says the same thing without congratulating the customer on it. | **READY** — owner Viktor | — | — | operator | +| **R-242** | **A controller release that changes customer-visible behaviour is not delivered until a golden carries it — and nothing enforces that.** R-239 is the symptom; this is the mechanism, recorded **2026-08-07** and deliberately **NOT built** (the task that found it scoped it as a record-only item). **Two releases went out without a golden and the gap was invisible until a walk measured it from the customer's side**: v0.204.0 (R-237) and v0.205.0 (R-234) were written, tested, pushed, and CHANGELOG'd, and every one of those steps passed while a machine installed that night received neither. The register said CLOSED; the fleet said otherwise. **Nothing in the release path knows a golden exists.** The version bump, the image push, the CHANGELOG entry and the register closure are all repo-local; the manifest's `golden_version` is edited by a separate operator act, in a different repo, with no link back. **Proposed shapes, cheapest first — the choice is the operator's and is not taken here.** (a) **A release-path checklist step** — one line in the controller's end-of-session checklist: *a release that changes customer-visible behaviour is not finished until a golden carries it or a register row says why not.* Costs nothing, catches nothing mechanically. (b) **A gate in `repo_gates.py`** comparing the manifest's `golden_version` against the newest released controller and FAILING (or warning) past a tolerance of one minor. Mechanical, runs on every push, and would have fired the morning after v0.204.0. (c) **A hub-side checker** — the hub already knows every box's running controller version from `/hosts` and the vouched golden from the manifest; a periodic comparison against the newest published image would catch drift the repo cannot see, including a vouch that was made and then rolled back. **Earliest catch: (b).** It fires on the push that creates the gap, before any box is installed, and it needs no live fleet. **(c) catches strictly more but only after boxes exist.** (a) is worth doing regardless because it is free. **Not built. No gate was written this session.** **⚠ IT RECURRED WITHIN A DAY, WHICH IS THE ARGUMENT FOR BUILDING IT.** Controller **v0.206.0** shipped the R-241 fixes on 2026-08-07 while the vouched golden still carried **0.205.0** — so a machine installed on the morning of 2026-08-08 would have received neither. Third occurrence of the shape in three days (R-111/R-115/R-120 are the older family). **✅ SHAPE (b) BUILT 2026-08-08 — `scripts/golden_currency_gate.py`, registered in `repo_gates.py` as gate 7.** **It was shown FAILING against that exact state before anything was baked**, which is its red-proof and the reason its own introducing push needed `--no-verify` (stated in the session report rather than worked around): `newest released controller : 0.206.0 / newest golden baked : 0.205.0 → CONVICTED`. **IT IS `--fast`, AND THAT FORCED ITS DESIGN:** both `.githooks/pre-push` AND CI run `repo_gates.py --fast`, so a non-fast gate would run in NEITHER — the R-29 census failure this runner exists to end. **THEREFORE IT CHECKS THE BAKE, NOT THE VOUCH**, because the vouched version lives only in the hub's `hub_settings` with no copy in git, and putting a copy there would create a second source of truth that can drift — a green gate over a false claim being the worst outcome available. **A bake without a vouch still passes: that half is NOT closed and stays on this row.** It also compares versions rather than behaviour, so a release changing nothing customer-visible trips it too — accepted deliberately, because judging that by hand is what failed three times and the cost of a false trip is one bake; a waiver belongs here, never in a habit of bypassing. **VOUCHED 2026-08-08 with the operator's approval** — golden `0.206.0` / sha `c85230b4…108e`; `agent_version` and `min_agent` both stayed `0.127.0`, and `wrapper_sha256` was carried through explicitly because the handler clears it when omitted. **The gate was CONVICTED before the bake and OK after it** — red→green on the same command, which is its proof that it measures something real. **⚠ THE GATE FIRED FOR REAL, 2026-08-08 — and it was right.** Controller **v0.207.0** (R-249/R-252/R-253) is released, tested and pushed, and **no golden carries it** — the newest bake is 0.206.0 — so `golden_currency_gate.py` FAILED, saying exactly the true thing: *a machine installed right now would receive v0.206.0*. **The `felhom.eu` push therefore used `git push --no-verify`, declared here, in the commit message and in the session report.** A bypass and NOT a waiver, deliberately: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **Owed: bake golden 0.207.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; the vouch is a three-field change). **This row's own remaining half is unchanged — nothing gates the VOUCH itself.** **✅ THE OWED BAKE IS DONE, SAME DAY — golden 0.207.0 baked, published, round-trip verified and VOUCHED (2026-08-08).** The gate went from red to **green**, and the `--no-verify` bypass declared above is now historical rather than standing. **Round trip is the evidence, not the build log:** the published bytes were downloaded back — 656 879 192 B, sha256 `20ec9602…22995`, both identical to what the bake reported — and **`./etc/felhom-controller-image` read OUT of the downloaded archive says `felhom-controller:0.207.0`**, which is the delivered artifact naming the controller it will start. **The vouch was a three-field change with all three checked deliberately** (`MinAgent 0.127.0` read from the golden's controller CHANGELOG header, not assumed; `agent_version` already ≥ it; `min_agent` not above `agent_version`, so not the R-216 shape) and verified by **re-reading the manifest rather than trusting the flash**. **This row's remaining half is UNCHANGED and is the whole of what is still open: nothing gates the VOUCH itself** — the currency gate's own docstring says it checks the bake, so a baked-but-unvouched golden still passes it silently. Evidence: `tests/golden-0.207.0-2026-08-08/`. **⚠ RED AGAIN, 2026-08-08 (second time in two days) — controller v0.208.0 (R-254) is released and the vouched golden is 0.207.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.207.0 and none of today's fixes. **The `felhom.eu` push used `git push --no-verify`, declared in the commit, the CHANGELOG and the session report** — **a bypass, not a waiver**, on the same reasoning as yesterday: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **Owed: bake golden 0.208.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change, `MinAgent 0.127.0` unchanged). **Note the cadence this is establishing: two releases, two bakes owed within 24 h.** That is the argument for this row's OTHER half — nothing gates the vouch, so the only thing standing between a release and an undelivered fleet is somebody remembering. **⚠ RED AGAIN, 2026-08-30 — controller v0.224.0 (R-330) and v0.225.0 (R-331) are released and the newest golden carries 0.223.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.223.0 and neither of today's fixes. **The `felhom.eu` push used `git push --no-verify`, declared in the commit message, in `hub/CHANGELOG.md` and in `REPORT.md` — a BYPASS, not a waiver**, on the same reasoning as the two 2026-08-08 entries above: the gate offers a waiver only for a release that *deliberately needs no golden*, and these need one. **The operator was asked and ruled bypass-now-bake-later on 2026-08-30**, on the stated ground that neither fix bites a DAY-0 box — R-330 is a nightly false alarm about apps a new box has not installed yet, and R-331 is a hub-side display over backups a new box has not taken yet — and both arrive by self-update afterwards. **That ground is recorded because it is the thing to re-check, not a general licence: the next release that changes first-boot behaviour cannot reuse it.** **OWED: bake a golden carrying 0.225.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change — `golden_version` + `agent_version` + `min_agent`; MinAgent is 0.129.0 per both CHANGELOG headers). **Cadence note, unchanged and now worse: this is the fourth bypass of this gate, and the gap it names is now two releases wide rather than one.** **⚠ WIDENED TO THREE THE SAME DAY — v0.226.0 (R-353/R-357/R-358/R-360) shipped 2026-08-30 and the golden still carries 0.223.0.** The `felhom.eu` push carrying that release's documentation used `git push --no-verify` on the operator's standing ruling from earlier the same day, declared in the commit and in `REPORT.md`. **The day-0 ground still holds for all three and was re-checked rather than assumed:** R-330 alarms about apps a new box has not installed; R-331 is a hub display over backups a new box has not taken; **R-353/357/358/360 are restore-surface fixes, and a day-0 box has nothing to restore.** **The ground expires the moment a release changes first-boot behaviour — that is the thing to re-check, not a licence.** **Owed: ONE bake carrying 0.226.0 covers all three** (`RUNBOOK-manual-build.md` §4.1; three-field vouch, MinAgent 0.129.0), then raise the floor. **✅ PAID THE SAME DAY — golden `0.226.1` baked, published, round-trip verified, VOUCHED, and the floor RAISED (2026-08-30).** Evidence: `documentation/tests/golden-0.226.1-2026-08-30/`. `golden_currency_gate.py` went **red → green** on the same command, which is its proof that it measures something real. **The three declared bypasses above are now HISTORICAL rather than standing.** **The round trip is the evidence, not the build log:** the published bytes were downloaded back — **657 197 592 B, sha256 `70ed8e93…baefe69`**, both identical to what the bake reported — and `./etc/felhom-controller-image` read **out of the downloaded archive** says `felhom-controller:0.226.1`, i.e. the delivered artifact naming the controller it will start. **The three-field vouch was checked deliberately, not assumed:** `MinAgent 0.129.0` read from the golden's controller CHANGELOG header, `agent_version 0.130.0 ≥ min_agent 0.129.0` (so NOT the R-216 shape), and the result verified by **re-reading the manifest** rather than trusting the flash — golden option `0.226.1 SELECTED`, all four shas matching. **The floor is proven ACTING, not merely set:** `demo-felhom` self-updated within 30 s, logging `[selfupdate] Post-update startup: update successful (0.225.0 → 0.226.1)`. **⚠ AND IT HAPPENED AGAIN THE SAME DAY, AND WAS PAID AGAIN.** v0.227.0/v0.227.1 (R-359/R-397) shipped after the 0.226.1 bake, the gate convicted a fifth time, that `felhom.eu` push used `--no-verify` and declared it, and golden **0.227.1** was baked, published, round-trip verified, **VOUCHED** and the floor **RAISED to 0.227.1** within the hour. Evidence: `documentation/tests/golden-0.227.1-2026-08-30/`. **THE CADENCE IS NOW MEASURED RATHER THAN ASSERTED: five convictions and two full bakes in one day.** Every bypass was declared and every debt was paid — but the pattern this row exists to name is exactly that a release and its delivery are separate acts, performed hours apart, by whoever remembers. **The floor was proven ACTING both times:** `demo-felhom` self-updated 0.225.0→0.226.1, then 0.226.1→0.227.1 — and the second time it also registered the new `offsite-integrity` job **by itself, on a box nobody deployed to**, which is the strongest evidence this row has ever carried that a floor delivers rather than merely records. **This row's OTHER half is still open and untouched: nothing gates the VOUCH itself** — the currency gate's own docstring says it checks the bake, so a baked-but-unvouched golden still passes it silently. **2026-08-31, the SEVENTH debt and it was paid the same day — twice in one day.** v0.230.0 shipped in the morning with the newest golden at 0.229.0, **which is the build R-403 says deletes a good copy**, so the gate was red across four commits (`dddcc80`, `6e550ae`, `130f7a6`, `32a4c35`). Golden **0.230.0** baked, published, round-trip verified, vouched, and the fleet floor raised 0.229.0 → 0.230.0; `demo-felhom` moved itself off the defective build unattended (`controller-swap: new controller healthy`, 16:21:40 CEST). Evidence: `documentation/tests/golden-0.230.0-2026-08-31/`. **The gate did its job and its own weakness surfaced doing it — R-410.** | **READY — the vouch half only** — owner Viktor **⚠ SIXTH CONVICTION, 2026-08-31 — controller v0.229.0 (R-102/R-103) is released and the newest golden carries 0.228.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.228.0 and neither of today's fixes. The `felhom.eu` push carrying this release's documentation used `git push --no-verify`, declared in the commit message and in `felhom-controller/REPORT.md` - **a BYPASS, not a waiver**, on the same reasoning as the five entries above: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **The day-0 ground was RE-CHECKED rather than reused:** R-102 and R-103 are restore-surface changes on the Tier-2 card, and a day-0 box has taken no Tier-2 copy and has nothing to restore from one; no first-boot behaviour changed, and `MinAgent` is unchanged at 0.129.0. **The ground still expires the moment a release changes first-boot behaviour.** **OWED: bake a golden carrying 0.229.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change - `golden_version` + `agent_version` + `min_agent`, MinAgent 0.129.0), then raise the floor. Fleet floor and golden are 0.228.0 today. **Golden and fleet delivery are the operator's (this row).** **✅ PAID THE SAME DAY — golden `0.229.0` baked, published, round-trip verified, VOUCHED, and the floor RAISED (2026-08-31).** Evidence: `documentation/tests/golden-0.229.0-2026-08-31/`. `golden_currency_gate.py` went **red to green** on the same command, which is its proof that it measures something real. **The `--no-verify` bypass declared above is now HISTORICAL rather than standing.** **The round trip is the evidence, not the build log:** the published bytes were downloaded back - **656 864 331 B, sha256 `39aa886d…d7bdae87`**, both identical to what the bake reported - and `./etc/felhom-controller-image` read **out of the downloaded archive** says `felhom-controller:0.229.0`. **A THIRD independent reader agreed before anything was vouched:** the hub's own Day-0 dropdown read the same sha straight from Gitea, a different code path from the round trip. **The three-field vouch was checked deliberately, not assumed** (`MinAgent 0.129.0` read from the golden's controller CHANGELOG header; `agent_version 0.130.0` >= `min_agent 0.129.0`, so NOT the R-216 shape; `agent_sha256` and `wrapper_sha256` carried through explicitly because the handler clears a field it is not sent), and verified by **re-reading the manifest** rather than trusting the flash. The **R-120 gate passed rather than being bypassed** - fleet newest 0.229.0, golden 0.229.0. **The floor is proven ACTING:** `demo-felhom` self-updated `0.228.0 -> 0.229.0` and logged `settle-gate: GO - at/above floor 0.229.0 (we are 0.229.0)` - **nobody deployed to that box.** **Cadence note: this is the SECOND bake in one day (0.228.0 then 0.229.0) and the sixth conviction, and both debts were paid within the hour.** This row's OTHER half is still open and untouched: **nothing gates the VOUCH itself** - the currency gate checks the bake, so a baked-but-unvouched golden still passes it silently. **⚠ SEVENTH CONVICTION, 2026-08-31 — controller v0.230.0 (R-403) is released and the newest golden carries 0.229.0. AND THIS ONE IS NOT LIKE THE OTHERS: the day-0 ground does NOT apply and must not be reused.** Every previous bypass rested on 'a day-0 box has nothing to restore / nothing to alarm about yet'. R-403 is a defect in the NIGHTLY TIER-2 COPY, which a day-0 box starts running on its first night: a machine installed on 0.229.0 can have a complete recovery package on its second drive replaced by an empty one, and that is measured, not suspected (120 082 104 B -> 7 036 B on demo-hp). **The row's own standing sentence - 'the ground expires the moment a release changes first-boot behaviour' - is what expires it here.** The `felhom.eu` push carrying this release's documentation used `git push --no-verify`, declared in the commit message and in `felhom-controller/REPORT.md` - a BYPASS, not a waiver. **OWED, and more urgent than the previous six: bake a golden carrying 0.230.0, vouch it (three fields, MinAgent 0.129.0 unchanged), and raise the floor.** `demo-hp` was updated by hand; `demo-felhom` is still on 0.229.0 and still carries the defect. **See also R-404, filed today: this is the seventh bypass and the habit is now the thing being reported.** **2026-09-01 (R-404): THE BAKE HALF IS UNCHANGED AND THE VOUCH HALF IS STILL OPEN.** R-404 moved WHO the bake check refuses and added a notice in the controller repo; it did NOT touch what is checked. **Nothing gates the VOUCH.** A baked-but-unvouched golden still passes both the gate and the new notice, and the reason is unchanged and forced: the vouched version lives only in the hub's `hub_settings` table, there is no copy in git, and a hub-reading gate could not be `--fast` so it would run in neither the hook nor CI. **Do not read R-404's closure as closing this.** **2026-09-13 — NARROWED: the WAIVER half is BUILT (R-468).** The docstring's *"honest fix is a recorded waiver in the register, never a habit of bypassing"* is now a mechanism: `golden_currency_gate.py` reads `documentation/tests/golden-waiver.yml` (dated, ≤ 14 days, row-bound), turns a BEHIND conviction into a loud advisory while valid, and is red again when it expires — the difference from this row's original rule, which recurred the next day, is that a dated waiver cannot be forgotten. It never covers an UNRECORDED golden (R-385). Operator ruling the same day: goldens weekly and before any install, not per release. **What stays open on THIS row is exactly one thing: nothing gates the VOUCH.** The waiver does not touch it, and the reason it is unbuilt is unchanged (the vouched version lives only in the hub). | — | — | operator | +| **R-243** | **A box in the R-241 state silently stops backing up off-site, and NO ALARM OF ANY KIND FIRES.** Found by the R-241 spike (2026-08-07) as a by-product; **not part of the walk's finding and not previously filed.** The R-241 state is self-locking in a second, worse way than the recovery-journey dead end: `escrow_state` is stuck `pending` forever (the auto-confirm flips only on a hash match, and the hash cannot match a key the box minted itself), and `runOffboxBackup` returns at the escrow gate (`offbox.go:743`) before touching anything. **So off-site backups never run again — and the hub never notices.** All three signals that could catch it are excluded, each for its own individually-correct reason, verified in the hub this session: `offsite_stale` — `isStale` (`monitor/offsite.go:135`) returns false unless `EscrowState == "escrowed"`, and its own comment reads *"Pending/disabled = normal onboarding, never stale"*, so the box is classified as **still being set up, forever**; `offsite_delivery_stuck` — `monitor/offsite_delivery.go:91` skips the `applied` shape, and delivery genuinely IS applied (the credential was consumed and the target is in every report); `backup_failed` — never fires, because nothing fails: the run returns `nil` before it starts. **Three correct exclusions leaving one state unobserved.** This is the same class as the workspace `CLAUDE.md` "presence is not success" rule, one level up: **the absence of a failure is being read as the presence of a working tier.** **Partly subsumed by R-241's fix** — a box that recovers leaves this state — but **not for a box that does not**, and the alarm gap is what makes "does not" survivable indefinitely. **Not fixed; no code written.** **⚠ UPDATED 2026-08-07 (v0.206.0) — the STATE this row describes can no longer be entered, but the ALARM GAP is untouched and the row stays open.** R-241's mint guard means a box no longer mints a key over a sealed package, so it no longer arrives in the "escrow stuck pending against a self-minted key" state by itself. **What replaces it is a state that is VISIBLE rather than silent:** the box declares `offsite.state=awaiting_recovery_key` and the customer is offered the recovery screen. **But the hub still raises nothing for it**, and for the same three reasons: `isStale` needs `escrowed`, the delivery checker skips the `applied` shape, and `backup_failed` needs a run that never happens. **So a box whose customer never acts still stops backing up off-site with no operator signal** — the difference is that the customer can now see it and act, where before nobody could. **The remaining work is an operator-side signal for a box held in `awaiting_recovery_key` past some age**, and it is deliberately not bundled into R-241's fix. **⚠ MEASURED ON A REBUILD, 2026-08-07 (fifth walk) — the gap is real for the state this row describes, and NOT for the state a rebuild produces.** 88 seconds after the walk5 guest was destroyed and rebuilt, the hub emitted `offsite_delivery_stuck` (**warning**) and wrote an **operator-channel** `notification_log` row recording `offsite_credential_restaged` / status **REFUSED** with an accurate reason — *"the credential was applied and worked; the target was lost afterwards … a guest rebuild does, R-193"*. So on the **regressed-apply** shape the operator IS told, promptly and correctly, and this row's *"skips the applied shape"* does not apply. The gap stands for a box that reaches the held state **without** a prior working tier in its report history. **Recorded so the row is not read wider than it measures.** | **READY** — owner Viktor | — | — | operator | +| **R-244** | **The customer DELETE cascade leaves `app_log_issues` behind, and it is systematic across every venue ever torn down.** Found **2026-08-07** while verifying the `finalwalk` teardown with a **full census** (every table, every column) rather than a per-table query. After a cascade that logged `COMPLETE … full teardown`, **61 rows still matched `finalwalk`**. Four of the five sources are **deliberate and correct** — the cascade's own header states *"Provenance/events are NEVER wiped — audit outlives every tier"*: `events` 16, `notification_log` 14, `host_deletions` 1, `customer_resets` 1. **The fifth is a gap:** `app_log_issues` 29 rows, which the residue purge does not touch (its logged leg covers `reports`/`app_telemetry`/`app_log_tails`/`log_tail_requests`/`notif_prefs`/`selfbind_tokens`/`appliance_registrations` — not this table). **It is not a `finalwalk` quirk:** rows still reference **`c11` 40, `rewalk` 20, `part4` 24** — all three torn down 2026-08-06, whose ledger recorded *"0 occurrences"*. **That prior claim was measured with a narrower query and does not survive a full census; the correction is recorded rather than the measurement quietly redone.** **Why it was probably never written, established rather than assumed:** the table is a **fleet-wide aggregate** keyed on `app_name`+`fingerprint` with an `affected_customers` JSON list — of the 29 `finalwalk` rows, **12 reference only `finalwalk`** (orphans, safely deletable) and **17 are shared with LIVE customers** (`demo-felhom`, `peti-felhom`, …) and **must not be deleted, only de-referenced.** A naive `DELETE … WHERE customer LIKE` would destroy a live customer's issue history — which is very likely why the leg does not exist, and is the reason this is not a one-line fix. **Severity is LOW and stated plainly: no secret material is involved** — app name, fingerprint, message text, counts, timestamps. What survives is a deleted customer's *identifier* inside an aggregate row. **Proposed shape:** a residue leg that (a) removes the customer id from `affected_customers`/`context_customer`, and (b) deletes rows whose `affected_customers` becomes empty; plus a one-off sweep for the four already-torn-down venues. **The general lesson is the reusable part:** *a per-table absence query is not a census.* The teardown verification is now a full-schema sweep, and that is what found this. **Not fixed** — a cascade change needs its own red-proof and this session was scoped as a spike plus two operations. Evidence: `tests/teardown-finalwalk-2026-08-07.md`. **⚠ STILL OWED, AND NOW MEASURED RATHER THAN ESTIMATED (2026-08-08 census, read-only, no truncation).** `app_log_issues` holds **1309 rows**; **71 reference a torn-down venue** (`finalwalk`, `c11`, `rewalk`, `part4`); of those **44 are ORPHANS** — they name only torn-down customers and are safely deletable — and **27 are SHARED with a live customer** (`demo-felhom`, `peti-felhom`, …) and **must be de-referenced, never deleted**. 1238 rows are untouched. **The 27 are exactly why the leg was never written**, and why a `DELETE … WHERE customer LIKE` would destroy a live customer's issue history. **What it needs, precisely:** a cascade leg that (a) removes the customer id from `affected_customers` / `context_customer`, and (b) deletes only rows whose `affected_customers` becomes empty; plus a one-off sweep for the four venues already gone. **Why it was NOT done on 2026-08-08:** the fix is hub code, and that session's scope forbade a hub version bump; a hand-run SQL mutation over 71 rows — 27 of them needing surgical de-referencing — with no tested code path and no red-proof is precisely the shape that goes wrong on a live database. **It accumulates one venue at a time, so the next walk adds to it**; the numbers above mean the next session starts from data rather than a guess. **⚠ IT GREW AGAIN, AS PREDICTED — walk5 teardown, 2026-08-08.** The fifth walk's venue was torn down with a full-schema census taken **before and after**: **168 rows → 67**. Of the 67, **37 are by design** (`events` 21, `notification_log` 14, `host_deletions` 1, `customer_resets` 1) and **30 are `app_log_issues`** — this row's gap, and the count was **predicted in the pre-run enumeration rather than discovered afterwards**, which is the difference from the ledger that once recorded *"0 occurrences"* from a narrower query. **The running total across torn-down venues therefore rises from 71 to ~101 rows** (`finalwalk`, `c11`, `rewalk`, `part4`, now `walk5`) — the shared-with-a-live-customer subset must still be **de-referenced, never deleted**. **It accumulates one venue at a time and it did so again.** Evidence: `tests/walk5-r201-2026-08-07/teardown-walk5-2026-08-08.md`. **2026-09-25:** `peti-felhom` is no longer a live customer (deleted through the cascade, journal #20); 8 `app_log_issues` rows still name it — the same gap. | **READY** — owner Viktor | — | — | operator | +| **R-246** | **A leftover staleness flag has silently disabled the new recovery discriminator on `demo-hp` since 2026-08-04, and the flag is WRONG.** Found by a read-only spike, 2026-08-08. **Q1 — traced to an act, to the second:** at `2026-08-04 20:15:49` the hub emitted `offsite_reissued` and `escrow_stale` in the same second — an operator **Re-issue** pressed during the R-201 drill, three minutes after `escrow_blob_served` at 20:12:40/20:12:54. That was `offsite.ReissueCredentials`'s **precautionary** `MarkEscrowStale` call, which **hub v0.95.0 REMOVED the very next day** (R-196 / R-204 item 2) precisely because it marked healthy escrows stale. **Q2 — the flag is wrong, measured on both sides:** the hub's blob seals `restic_pw_sha256 = 8a9e33aa4da6769c…d080a`, and the key the box is actually using hashes to **the identical value**. The blob covers the key. **Q3 — nothing clears it by itself:** the ONLY writer of `stale_at = NULL` is `SaveHostEscrow`'s `ON CONFLICT` — i.e. a fresh escrow ceremony, **which is the one act that would supersede the good blob**. So the only exit from a false alarm is the destructive act the false alarm recommends. **Q5 — the blast radius, enumerated rather than assumed:** (1) `GetEscrowStatusForCustomer` withholds `restic_pw_sha256` from the ACK; (2) a **pending** box can never auto-confirm, so (3) **every off-site run is refused indefinitely** — neither bites `demo-hp`, which was already `escrowed` and is backing up healthily (12 snapshots, last success 2026-08-07T02:15:35Z); (4) the customer is told to create a new code; (5) **NEW — v0.206.0's shape (c) is inert**, because the box records an empty hub hash and falls back to (a)/(b), so the recovery screen would stay silent even if recovery were needed. **Q6 — a fresh box CANNOT reach this state:** `MarkEscrowStale` has **no production caller anywhere in the tree** (census: only its own definition, two comments and two test references). **The next walk cannot meet it.** **NOT CLEARED, deliberately** — Q2's answer says the fix is to clear this instance, but whether to also stop the column being settable at all is a separate ruling, and the spike was scoped read-only. **✅ THE FLAG IS CLEARED — operator-approved and applied 2026-08-08.** One row, identity-matched on `host_id` and guarded on `stale_at IS NOT NULL`; `changes()` returned **1**. **Verified end to end, not just in the database:** the hub now serves the hash again, the box recorded `hub_escrow_key_sha256 = 8a9e33aa4da6769c…d080a` at `11:10:19Z`, and that is **byte-identical to the key it is using** — so shape (c) compares, matches, and correctly stays silent. **The false stale warning is gone, proven with a positive control** rather than an absent line: **0** `escrow-confirm` lines since the restart while **5** scheduler lines in the same window prove the box was logging, and the recorded hash proves an ACK was processed. *(Method note: the hub pod is Alpine with no `sqlite3`; it was installed into the container's ephemeral writable layer — image and node untouched, gone on restart. SQLite's own file locking coordinated the write with the live hub; an earlier attempt failed cleanly on quoting and changed nothing, which is the fail-safe working.)* **STILL OPEN under this ID: the ruling on whether `stale_at` keeps a live setter at all.** It currently has NO production caller, so the column is write-only-by-accident — a field that changes behaviour, that nothing sets and nothing can see (R-248). Either give it an evidential setter or retire it; **do not leave it as a trap that only a database read can spring.** | **READY — flag cleared; the column ruling is still owed** — owner Viktor **Folded R-248 2026-10-03** (the same ruling: give `stale_at` a visible, evidence-bearing setter, or retire it). | — | — | operator | +| **R-250** | **A customer create can fail fail-closed because the host-key scan ladder is shorter than the DNS/AAAA settle time.** Found 2026-08-07 creating the fifth walk's venue. `POST /configs/new` with off-site enabled provisions the Storage Box sub-account and then scans its SSH host key to pin it — **fail-closed by design** (`offsite.go:111-121`, *"don't serve a descriptor the controller can't verify"*), with `defaultScanBackoff` = 2+4+8+16+30 ≈ **60 s**, sized by its own comment to *"the observed DNS propagation lag"*. **Measured, both halves:** the first create exhausted the ladder — five `no such host`, then `dial tcp [2a01:4f8:bacc:2:200::d30]:23: connect: network is unreachable` (the name had just begun resolving, **AAAA-first, into a pod with no IPv6 route**) — and the hub logged `[ERROR] offsite provision for walk5`. An **identical second POST succeeded** ~70 s later on its own final rung (`shared already provisioned … (subaccount 285351)`). **Total settle ≈ 100 s against a 60 s budget.** The endpoint was never the problem: verified afterwards from both the node and the hub pod, by name, `SSH-2.0-OpenSSH_9.6p1`. **Two distinct things are wrong and should not be merged:** (1) the budget is sized against DNS *existence*, but what actually bit is the **AAAA-before-A** window, a different and longer phenomenon — lengthen the ladder and/or prefer the A record for the scan dial; (2) **the operator is told nothing actionable** — the create fails with a generic error, the remedy is "press it again", and nothing says so. The retry is genuinely safe (idempotent on the label; confirmed afterwards — `one_time_secrets` 1, one sub-account, no double-provision) **but that safety is invisible to the person deciding whether pressing again will double-charge them.** **Severity LOW-MEDIUM:** self-clearing, no data at risk — but it is the first thing a new customer's provisioning does and it fails looking like an outage. | **READY** — owner Viktor | — | — | operator | +| **R-251** | **The recovery listing renders one row per restic TAG, so the customer is shown an "app" they never installed and their data counted twice.** Measured on the fifth walk, 2026-08-07, on the screen the customer reaches after entering R. The snapshot carries tags `felhom-offbox,calibre-web`; the listing renders **two rows** — `calibre-web · 2026-08-07 14:57 · 12.8 MB` and `felhom-offbox · 2026-08-07 14:57 · 12.8 MB`. `felhom-offbox` is the tier's own marker tag, not an application. **The screen's whole job is to let the customer check that what is in the store is what they expect** (*"Nézd át, hogy tényleg azt találod-e itt, amire számítasz"*), and it shows them a stranger's name beside their own data and a total that is double the truth. **Cosmetic, not a data defect** — the restore page correctly offers only `calibre-web`. **Fix:** filter the marker tag out of the listing, or key the rows on the app tag. | **READY** — owner Viktor | — | — | operator | +| **R-255** | **The check that would catch a fourth secret-in-the-body covers 4 of 27 pages, and the cheap gate that covers all 36 templates is blind to the shape that actually shipped.** Filed 2026-08-08 while closing R-254, **because a partial guard reported as complete is worse than no guard — it stops the next person looking.** **Two nets, both measured.** **(1) `scripts/secret_in_markup_gate.py`** reads all 36 templates and convicts any `{{ … }}` naming a secret unless allowlisted with a reason. It catches `{{.RetrievalPassword}}` and `{{.InitialCreds.Password}}`, **and it catches a launder through a local variable** because the assignment itself names the secret (`{{$v := .InitialCreds.Password}}` is convicted — verified). **It is blind to a secret arriving under a NEUTRAL PAGE-DATA KEY** — `data["Tagline"] = creds.Password` then `{{.AppInfo.Tagline}}` passes it cleanly, also verified. **That is exactly the shape of R-254 site two** (`value="{{$val}}"` inside an `{{if eq .Type "secret"}}` branch), so the gate **would not have caught one of the three instances it was written for.** **(2) The runtime body assertion** — render the page and grep the response for a sentinel — catches every shape, including that one (demonstrated on the same planted leak the gate missed). But it needs each page's data to be constructible in a test, and **only 4 of 27 page templates have that today**: `settings_security`, `app_info`, `deploy`, `backups_restore` — the four that were touched by R-249/R-252/R-253/R-254 and therefore got their own tests. **The other 23 pages have no runtime coverage at all.** **What closing this needs, so the cost is not re-estimated:** a per-page data fixture for the remaining 23 (most need a wired `Server` — `stackMgr`, `backupMgr`, agent seams), then one table-driven test that renders each with a sentinel substituted for every string in its data and asserts the sentinel is absent. **That is real scaffolding, which is why it was NOT built inside R-254's session** rather than half-built and declared done. | **READY** — owner Viktor | — | — | operator | +| **R-256** | **C2 — „A mentéskezelő nem elérhető." names no route at all.** `web/offbox_handlers.go:47` and `:181` flash this to the customer on the off-site backup surface. It states an internal component's unavailability in the operator's vocabulary („mentéskezelő" = the backup Manager object), gives no reason the customer can act on, and names no next step — not "try again in a few minutes", not "contact support", not a page to go to. **Contrast, in the same subsystem and shipped the same week:** R-252's fix reads *„Meghajtók", „Meglévő meghajtó csatolása". Utána gyere vissza ide.* — a route. **Severity is low and stated so it is not over-ranked:** the condition is a nil backup manager, which on a healthy box does not occur; this is about the copy, not a broken path. Found by the C2 sample (19 refusals on the recovery/restore/offbox surface; **~202 of the repo's 221 refusal strings were NOT examined**) | **READY** — owner Viktor | — | — | operator | +| **R-257** | **C2 — „Az offsite tároló nincs elárvult állapotban." puts an English loanword and an internal state name in front of a Hungarian household customer, and names no route.** `web/offbox_handlers.go:270` (the Go error it mirrors is `backup/offbox.go:343`). „Offsite" is untranslated; „elárvult állapot" is the codebase's own `OffboxOrphaned()` predicate surfacing verbatim. A customer who pressed a button and got this cannot tell whether something failed, whether they did something wrong, or what to do instead. **This is a refusal that is CORRECT and fail-closed and still a dead end** — the same shape R-241 recorded for `--recover-offsite-install`. **Fix shape, not a decision:** say what the customer tried to do, why it does not apply right now, and where to look — or, since this is a state they cannot reach deliberately, do not offer the action at all | **READY** — owner Viktor | — | — | operator | +| **R-261** | **C6 — `CountSelfBindTokens` exists so that callers can assert an invariant, and no production caller asserts it.** `hub/internal/store/selfbind.go:106-111`. Its doc comment: *"it exists so callers can assert the 'after this runs, the only live link is one we just issued — or none' invariant that the auto-mint at customer-create / RESET-completion depends on."* **Census: only its own declaration in production; the two callers are `selfbind_automint_test.go:29` and `customer_delete_test.go:510`.** Tests are not callers (the campaign's rule), so the invariant the auto-mint *depends on* is checked in the test suite and never at the moment it matters. **This is the smallest of the eight rows and is filed at its true size, because the rest of the C6 sweep found INERT dead accessors rather than defects:** `OffboxOrphanedRenamedTo` and `OffboxEscrowState` have no caller but their data reaches the card another way (the template reads the settings field directly, `backups_remote.html:80`) — **R-228 is genuinely closed, and the sweep's first reading that it had regressed was wrong.** **The more consequential C6 result is a method result and is in the report, not here:** `golang.org/x/tools/cmd/deadcode` re-finds **neither** known instance, and a planted probe measured why — it reports an unreachable exported FUNCTION and not an unreachable exported METHOD on a widely-used type, and both known instances are methods | **READY** — owner Viktor | — | — | operator | +| **R-262** | **C7 — a comment claims a cross-repo contract is mirrored „field-for-field" and „the key-set tests guard drift"; it is two fields short, AND THE FIXTURE THE TEST READS OMITS THE SAME TWO FIELDS.** `hub/internal/api/handler.go:682-687` covers `hostBackup` **and** `hostRestoreTest`. **It is TRUE of `hostBackup`** (verified field-for-field against `agent/internal/hub/Backup`). **It is FALSE of `hostRestoreTest`:** the agent emits `mount_parity` and `mount_inventory` (`hub/report.go:432-433`, populated in production from `reconcile/restoretest.go:277-283` via `backup/runner.go:517`), and the hub has no field for either — **0 occurrences in the entire hub repo** outside the CHANGELOG. **The guard is blind in exactly the place the drift is:** `TestHostReport_GoldenContract` reads `testdata/host-report.golden.json`, the two copies of which are byte-identical as required — and **neither contains `mount_parity` or `mount_inventory` at all**, so the key sets agree on a shape that is not the shape the agent sends. A test that cannot fail on the drift it names is the R-97b lesson (*prove the consequence, not the mechanism*) landing on a contract test. **Consequence, stated precisely:** the verdict is not lost (a parity mismatch fails the test before `Pass` is set), but the hub cannot distinguish a full-fidelity restore-test pass from a boot-only one, for any agent, ever. **Fix shape, not a decision:** add the two fields and put them in the fixture — or narrow the comment to name `hostBackup` only and say plainly that `hostRestoreTest` is a subset. **Attached observation:** the same fixture carries `cpu_temp_c` and `loadavg`, which no hub struct decodes — a fixture carrying keys the receiver cannot read is the same shape one level down | **READY** — owner Viktor | — | — | operator | +| **R-263** | **C7 — „This is the ONLY writer of `StoragePath.BackupTarget`" is false, and nothing pins it.** `settings/settings.go:1317-1319`, on `SetBackupTarget`. **`ClearBackupTarget` (`:1358-1363`) also writes the field**, 17 lines below, in the same file. **The GUARANTEE the comment protects is intact and that is why this is filed small:** the sentence continues *"registration must never set it (E-2 §3: a drive never acquires a role by appearing)"*, and `ClearBackupTarget` only ever writes `false`, so no path other than `SetBackupTarget` **grants** the role. **What is wrong is the claim as written, and the absence of anything holding it:** `backup_target_role_test.go` exercises the behaviour and asserts nothing about writer uniqueness, so if a third writer appeared tomorrow — one that granted — the comment would still read as settled and the suite would still be green. **This is the class's own definition:** an invariant asserted in prose with no test pinning it. **Fix shape:** one word (*"the only writer that GRANTS the role"*) plus a test that fails when a second granting writer appears. **Method honesty:** found in a sample of **60 of 2652** production invariant comments — the "is the only" form only, chosen because a uniqueness claim is the one form a grep can falsify. **~2592 production and all 1440 test invariant comments were NOT examined**, so C7 has the weakest coverage of the seven classes and the task's instruction to check the tests' own claims is **owed, not discharged** | **READY** — owner Viktor | — | — | operator | +| **R-264** | **Twenty-one facts the boxes report that the hub can now decode nowhere, each allowlisted with a reason rather than silently skipped — and for these the reason is "no consumer today, and one is arguably owed".** Split out of R-260 on 2026-08-08 so that closing the CLASS (gated) and fixing its sharpest instance (`operator_key_configured`) could not be mistaken for having decided what the hub should do with the rest. **The list, grouped by what a consumer would be for.** **(a) Guest-network health — `guest_net` and its seven children** (`checked_at`, `has_route`, `dhclient_alive`, `heal_succeeded`, `heals_last_hour`, `last_heal_at`, `damped`). The R-54 watchdog reports per-guest network state and self-heal counts every cycle and the hub — the component that emails the operator — models none of it. There is a live incident in this project's own record where a killed `dhclient` took a tunnel down for 1 h 15 m (`audits/INCIDENT-guest-dhclient-killed-2026-07-20.md`); a recurring-heal signal is exactly what would have surfaced it. **This is the strongest candidate of the twenty-one.** **(b) `selfupdate_pending` + `selfupdate_pending_version`** — an agent that has flipped its binary and never committed reports pending on every heartbeat so that "the operator sees WHY the version isn't advancing", and no operator can see it. **(c) `mgmt_plane.healed_recently`** — bounded: the hub DOES alarm on the `privsep_healed_at` timestamp beside it, so the recurring-clobber signal is not lost, only this flag. **(d) `restore_tests.mount_parity` + `mount_inventory`** — R-262's subject; the verdict is not lost (a mismatch fails before `Pass` is set) but the hub cannot tell a full-fidelity pass from a boot-only one. **(e) `pbs_dr.applied_at`.** **(f) Controller-side: `config_hash`, `reporting_disabled`, `stacks`, `storage.migrated_to`, `backup.last_db_dump`, `backup.last_integrity_check`** — the last two are backup-integrity timestamps, which is the "presence is not success" neighbourhood. **For each the question is the same and is NOT answered here: is it wanted? If the hub should act on it, model it and name what consults it. If it should not, the honest end is that the emitter stops sending it** — a fact emitted forever and consumed nowhere is a future false green waiting for someone to write a check against it. **Deliberately not decided in the G-1 session**, whose scope was the gate plus the operator-access instance; unilaterally removing emitters would also break the byte-identical cross-repo host-report golden and is a coordinated two-repo change. **⚠ DISPOSITIONS RECORDED 2026-08-13 — and the first thing to say is that they were NOT in this register before today.** The rulings were made on 2026-08-12; this row still read **READY — owner Viktor** and the gate's twenty entries still all said *"arguably owed"*, so a session told to *"re-read the dispositions from the register"* would have found none. They are written down now, which is the point of writing them down. **THE COUNT WAS ALSO WRONG:** this row says *twenty-one*; the gate's allowlist held **twenty**, measured. Twenty is the number the dispositions below account for, exactly. **(1) BUILD A READER — four groups, fourteen facts.** (a) guest-network health, (b) the staged-update-pending pair, (d) the restore-test depth pair, (f-part) the two backup-integrity timestamps. **(2) NO READER WANTED — five facts, now recorded as `not consumed, DELIBERATELY` with the ruling and its date** in `scripts/wire_contract_gate.py`, each with its own reason rather than a bare refusal: `mgmt_plane.healed_recently` (the hub already alarms on the timestamp beside it), `pbs_dr.applied_at` (`pbs_dr.state` is the verdict; the timestamp alone is the attempt-read-as-result trap), `config_hash` (the hub authors the config and knows its own generation), `stacks` (the app view is built from the purpose-built `app_telemetry` wire), `storage.migrated_to` (box-local bookkeeping with no hub-side intent to reconcile against). **The emitters are deliberately left alone** — removing one is a coordinated two-repo change and breaks the host-report golden; the honest end here is a recorded decision, not a deletion. The gate grew a THIRD entry kind to carry them, because an undecided fact and a decided one must not read alike. **(3) `reporting_disabled`, decided on its own merits: RECLASSIFIED `redundant`** — `health.status = "disabled"` travels in the same minimal report, is decoded into `reports.health_status`, and IS rendered. **The decision surfaced a real defect the flag would not have fixed: the staleness checker is age-only, so a deliberately-silent box still alarms → R-321.** **PROGRESS, 2026-08-13: the first reader is BUILT — guest-network health, R-319.** Its eight allowlist entries are **removed** (an allowlisted tag is skipped, so leaving them would have meant the new reader's fields were never checked); the gate's checked-tag count rose **182 → 190** and skipped fell **88 → 80**, which is the positive control that the wiring is real. **WHERE THE TWENTY NOW STAND: 8 read · 5 deliberately unread · 1 redundant · 6 still owed a reader** (`selfupdate_pending`, `selfupdate_pending_version`, `restore_tests.mount_parity`, `restore_tests.mount_inventory`, `backup.last_db_dump`, `backup.last_integrity_check`) — counts measured from the allowlist, not estimated. **Only ONE reader was built on purpose:** four at once is a design session pretending to be an implementation, and this one now tells us what the other three cost | **OPEN — 6 of 20 still owed a reader; dispositions recorded 2026-08-13, first reader shipped (R-319)** — owner Viktor | — | — | operator | +| **R-266** | **A failed root `statfs` still reaches the hub as a 0-of-0 disk, and the hub cannot tell that from an empty one.** Split out of R-259 on 2026-08-08 so that fixing the CUSTOMER-facing half could not be mistaken for fixing the wire. `report/builder.go:93-95` copies `sysInfo.DiskTotalGB` / `DiskUsedGB` / `DiskPercent` into `r.Storage[0]` (`Mount: "/"`), and those are exactly the zeros a failed `statfs` leaves behind — the controller now KNOWS the measurement failed (`SystemInfo.DiskKnown`, controller v0.210.0) and the report still does not carry it. **Deliberately not fixed here, for a reason that is now structural rather than a preference:** adding a field to that report is a change to a declared wire, which since G-1 means the receiving side must model it in the same session (`scripts/wire_contract_gate.py` refuses otherwise) — a two-repo change with a hub bump, and this session deliberately touched no hub code. **RANKED LOW, and the reason is that the consequence is bounded:** the hub bands host storage on `disk_percent`, so a failed read presents as 0% used — the *quiet* direction. It cannot raise a false "nearly full" alarm; it can only fail to raise a true one, and only while the root filesystem is unreadable, which is a state with louder symptoms of its own. **Fix shape when it is taken:** carry `disk_known` on the storage entry and have the hub's fill checker skip an unknown reading rather than band it — never treat absent as 0 | **READY** — owner Viktor | — | — | operator | | **R-269** | **A rotated-out per-guest local-API token still authorises, and the test that appears to pin the opposite passes only because of its lookup ORDER.** `localapi.TokenStore.Mint` documents *"last-write wins — any previous token for this guest is revoked"*. Across processes that is FALSE until something unrelated forces a reload: the long-lived agent serves `Lookup` from an in-memory index and re-reads the store **only on a miss** (the B3 reload-on-miss optimisation, `tokenstore.go`), so a superseded token is a direct map **hit** and returns `(vmid, true)`. **Red-proved twice.** (a) A unit probe — `TestTokenStore_ReloadOnMiss_RemintCoherence` with the two lookups swapped, i.e. present the rotated-out token FIRST — **fails**; the shipped test passes only because it looks up the NEW token first, and that miss is what evicts the old hash. (b) On hardware, 2026-08-09: after the on-disk rotation the old token returned **HTTP 200**, then 401 only once a new-token lookup had forced the reload, and reliably 401 after `systemctl restart felhom-agent`. **This is the `CLAUDE.md` invariant-comment case exactly** — the comment reads as settled and the test that looks like its pin is order-dependent. **Fix options:** pin the reversed order with a test, or make eviction not depend on an unrelated miss. Until then, **an operator rotating a leaked token MUST restart the agent** — the runbook step is not optional | **READY (S) — NEW 2026-08-09** | — | Found by doing R-268's rotation rather than reading about it | CC | | **R-270** | **R-268's stated rotation recipe is incomplete: the controller never re-reads `bootstrap.json`'s `local_api`, so a rotation leaves the agent channel dead across restarts.** `bootstrap.ensureLocalAPI` returns early when `cfg.LocalAPI.Endpoint != ""` — by design it FILLS an absent block and never refreshes a present one — so the token the controller uses lives in its own `controller.yaml`, not in the mount. Proved live 2026-08-09: two controller restarts after a correct `bootstrap.json` rotation, still `HTTP 401`; the channel came up only once `local_api.token` was written into `controller.yaml`. The neighbouring `DetectEndpointDrift` compares the ENDPOINT and deliberately does not compare the token (*"a token mismatch is a different failure"*), so this shape is knowingly unmodelled. Parent question — which file is authoritative — is **R-78** | **READY (S) — NEW 2026-08-09** | — | Either teach the drift detector the token, or make the rotation path write both files. Correct the R-268 row's recipe either way | CC | | **R-271** | **The `agent_channel_unauthorized` alarm can never be closed, because its own prescribed remedy is what silences the recovery.** `channelhealth.Checker.Check`'s UP branch notifies only when `prev != "" && prev != "up"`; a controller restart resets `state` to `""`, so an unseeded→up transition is silent by construction. The alert text says *"token stale/rotated (**re-bootstrap**)"* — i.e. restart the controller — so **following the instruction guarantees no recovery event.** Observed live 2026-08-09: two `agent_channel_unauthorized` errors on the hub (one `sent`, one `suppressed` by the 1 h operator cooldown) and **nothing afterwards**, though the channel came up 3 minutes later and stayed up. The down side is deliberately asymmetric (F2: a born-down channel alerts on cycle 1); the up side never got the matching treatment. Customer dashboard is fine — `SetDashboard` reflects current state every cycle. It is the OPERATOR's trail that ends on "down" | **READY (S) — NEW 2026-08-09** | — | Notify on unseeded→up when the previous *persisted* state was down, or seed from the hub's last event | CC | -| **R-272** | **RANK 1 — Felhom's `--uninstall` leaves the exact condition that makes Felhom's own reinstall REFUSE.** Chain, fully evidenced on demo-hp 2026-08-09: Felhom installs `dnsmasq` at day-0 (`/var/lib/dpkg/info/dnsmasq.list` dated **2026-07-21 18:24 CEST**, demo-hp's day-0) and constrains it with a snippet in `/etc/dnsmasq.d/`; `--uninstall` removes the snippet and **restarts the daemon** (running process start time **2026-08-09 10:37:39 CEST — inside the 10:37:23–10:38:23 uninstall window**) but leaves the package installed and the unit **enabled**; unconstrained, dnsmasq binds `0.0.0.0:53`; the next install's preflight then hard-refuses with *"a resolver is already bound to :53"*. It is **not** PVE SDN's (`/etc/pve/sdn/` empty; stock unit). The teardown mentions it only as *"the 'sudo' and 'dnsmasq' packages were left installed (**system packages**)"* — **dnsmasq is not a system package here, Felhom installed it.** **What a customer does next:** reads a message blaming a resolver, concludes their own LAN DNS is at fault, and debugs something they never configured. **Counterfactual confirmed:** `systemctl stop dnsmasq && systemctl disable dnsmasq` → `host DNS (:53): free` → PRE-FLIGHT PASS, nothing else changed. **The refusal MESSAGE is good** (finding, evidence, two routes, and an explicit promise not to touch DNS on a host it does not own) — the defect is that Felhom caused the condition and does not say so | **READY (M) — NEW 2026-08-09** | — | Either stop+disable dnsmasq on uninstall when Felhom installed it, or have the preflight recognise its own leftover and say so | CC | -| **R-274** | **A local golden is adopted with NO version and NO checksum check, so a reinstall can silently come up releases behind.** `felhom-host-install.sh` step 7: `if [[ -n "$GOLDEN_VOLID" ]] && ! $FORCE_GITEA_GOLDEN; then log_skip "using local golden"; return 0; fi` — **the hub manifest's `golden.sha256`, whose entire purpose is to vouch from a different trust root than Gitea, is consulted only on the FETCH path.** A locally-present archive bypasses the vouch: no version compare, no digest, no warning. On demo-hp 2026-08-09 the preflight selected `local:backup/vzdump-lxc-9100-2026_08_03-07_33_00.tar.zst`, whose baked marker reads `felhom-controller:`**`0.192.0`**, against a vouched golden of **0.210.0** — 18 releases stale. **The sharp consequence:** 0.192.0 is **below 0.200.0, where R-193's off-site recovery SCREEN shipped**, so a customer reinstalled today returns on a controller that cannot run the recovery ceremony their data depends on; it is also born below the managed-update floor (0.200.0), and the updater's auto-target is the floor, never the newest. **It compounds with the teardown**, which deliberately keeps the old golden (*"golden vzdump left in place"*). This is the R-111/R-115/R-120 drift family one layer down: the R-120 gate guards what may be VOUCHED, nothing guards what an install TAKES. **OBSERVED 2026-08-09, AND THE RESULT NARROWS THE ROW — recorded because it partly refutes what was written above.** On the RESUME path step 7 **fetched the vouched 0.210.0 correctly** (`fetching golden v0.210.0 from Gitea`), because `--resume` skips preflight and preflight is where local auto-discovery sets `GOLDEN_VOLID` (the GL6-F4 comment says so). **So the fresh-install and resume paths disagree on golden selection, and the resume path is the safe one.** Discovery is `… | sort | tail -1`, i.e. the NEWEST local archive by filename — a sensible heuristic, **and still no comparison against the manifest's version or sha**. The defect therefore stands as: *a box whose newest local golden predates the vouched one installs stale, silently* — which is exactly the state demo-hp was in before this run (newest local 0.192.0 vs vouched 0.210.0). It is now masked on this box because the freshly fetched 0.210.0 is the newest — **correct by recency, not by verification**. There are now **three** goldens on `local` (07-21, 08-03, 08-09), because the teardown keeps them. **Still not observed: a FRESH (non-resume) install taking a stale local golden.** | **READY (S) — NEW 2026-08-09, NARROWED same day** | — | Compare the local golden's version/sha against the manifest and refuse or re-fetch on mismatch; say so in the BYO disclosure, which today lists only what the install CREATES, never what it REUSES | CC | +| **R-274** | **A local golden is adopted with NO version and NO checksum check, so a reinstall can silently come up releases behind.** `felhom-host-install.sh` step 7: `if [[ -n "$GOLDEN_VOLID" ]] && ! $FORCE_GITEA_GOLDEN; then log_skip "using local golden"; return 0; fi` — **the hub manifest's `golden.sha256`, whose entire purpose is to vouch from a different trust root than Gitea, is consulted only on the FETCH path.** A locally-present archive bypasses the vouch: no version compare, no digest, no warning. On demo-hp 2026-08-09 the preflight selected `local:backup/vzdump-lxc-9100-2026_08_03-07_33_00.tar.zst`, whose baked marker reads `felhom-controller:`**`0.192.0`**, against a vouched golden of **0.210.0** — 18 releases stale. **The sharp consequence:** 0.192.0 is **below 0.200.0, where R-193's off-site recovery SCREEN shipped**, so a customer reinstalled today returns on a controller that cannot run the recovery ceremony their data depends on; it is also born below the managed-update floor (0.200.0), and the updater's auto-target is the floor, never the newest. **It compounds with the teardown**, which deliberately keeps the old golden (*"golden vzdump left in place"*). This is the R-111/R-115/R-120 drift family one layer down: the R-120 gate guards what may be VOUCHED, nothing guards what an install TAKES. **OBSERVED 2026-08-09, AND THE RESULT NARROWS THE ROW — recorded because it partly refutes what was written above.** On the RESUME path step 7 **fetched the vouched 0.210.0 correctly** (`fetching golden v0.210.0 from Gitea`), because `--resume` skips preflight and preflight is where local auto-discovery sets `GOLDEN_VOLID` (the GL6-F4 comment says so). **So the fresh-install and resume paths disagree on golden selection, and the resume path is the safe one.** Discovery is `… | sort | tail -1`, i.e. the NEWEST local archive by filename — a sensible heuristic, **and still no comparison against the manifest's version or sha**. The defect therefore stands as: *a box whose newest local golden predates the vouched one installs stale, silently* — which is exactly the state demo-hp was in before this run (newest local 0.192.0 vs vouched 0.210.0). It is now masked on this box because the freshly fetched 0.210.0 is the newest — **correct by recency, not by verification**. There are now **three** goldens on `local` (07-21, 08-03, 08-09), because the teardown keeps them. **Still not observed: a FRESH (non-resume) install taking a stale local golden.** | **VERIFY** (2026-10-03 triage: Local golden is now checked against the hub manifest's version and sha before use (GOLDEN_CHECK_WHY, R-297) — felhom.eu/scripts/felhom-host-install.sh:2855-2905) — **READY (S) — NEW 2026-08-09, NARROWED same day** | — | Compare the local golden's version/sha against the manifest and refuse or re-fetch on mismatch; say so in the BYO disclosure, which today lists only what the install CREATES, never what it REUSES | CC | | **R-275** | **`--uninstall` leaves five 0600 `agent.json.*` credential backups, and the reinstall hands them to the new service account.** `/etc/felhom-agent/` survives with `agent.json.{campaign8-before,campaign9-before,campaign9-prev,pre-e-target-move,pre-prunegate.bak}`, each carrying a 64-char `hub.api_key` and a 59-char `proxmox.token`. The teardown claims to remove *"config (+ its .bak backups)"* and `scripts/CHANGELOG` F1 records *"uninstall now purges the agent config's `.bak*` siblings (one held a live hub api_key)"* — **that fix does not match the filenames in use, and it misses `agent.json.pre-prunegate.bak`, a file that literally ends in `.bak`.** **Exposure assessed, not assumed:** these are SUPERSEDED — the orphaned key hashes to `a5d2222a…`, the hub's current demo-hp key to `8c59d1b6…`, and the Proxmox token was deleted by the same uninstall. **But the reinstall recreates `felhom-agent` at uid 999, the same uid the deleted account had**, so three of the backups become the new account's files — verified readable as `felhom-agent`. A fresh install's service account inherits read access to the prior install's credentials; superseded today, live if the backups were recent (R-179's precedent). **Also left, undeclared:** `/etc/felhom/{.bootstrap-done,appliance-pairing-code}`, `felhom-bootstrap.service` + `/usr/local/sbin/felhom-bootstrap.sh`, the `vmbr9` stanza in `/etc/network/interfaces`, and `/etc/sudoers.d/felhom-agent.bak-pre-e2a` (21 KB — **INERT: sudo skips dotted filenames, verified with `sudo -l -U felhom-agent`; `visudo -c -f` parsing it OK is NOT evidence sudo loads it**) | **READY (S) — NEW 2026-08-09** | — | Purge by directory, not by glob; and do not let a new service account reuse a uid that owns old secrets | CC | | **R-276** | **RANK 2 — an uninstalled box keeps a live WireGuard tunnel into Felhom's off-site endpoint, and the teardown says nothing.** After `--uninstall` on demo-hp, `wg-quick@wg-felhom` is **enabled and active**, `/etc/wireguard/wg-felhom.conf` present, handshake to `167.233.158.164:443` **52 s old**, counters 5.86 GiB in / 2.48 GiB sent. It appears in **neither** the WIPED nor the KEPT list, though the BYO install disclosure names it prominently on the way in (*"an OUTBOUND WireGuard tunnel to the Felhom hub"*). A host told to leave Felhom retains a live network path into Felhom infrastructure, its hub-side peer registration intact, and nobody is told | **READY (S) — NEW 2026-08-09** | — | Tear the tunnel down and deregister the peer, or list it under KEPT with the reason and the removal command | CC | | **R-277** | **Three hub surfaces jointly present a HEALTHY off-site tier as an absent one — and it produced a wrong operator statement during this run.** For demo-hp on 2026-08-09 the box was pushing off-site daily without a gap (18 restic snapshots, `last_status: ok`), yet: (a) the customer page's Backup panel read `Snapshots 0 / Repo Size 0 MB / Integrity Unknown` — it renders the **local disk tier**, while the healthy `offsite` object sits **in the same report** unrendered on that panel; (b) the Offsite page read `0.0 GB` — true, but a 162 KB repo rounds to nothing; (c) a stale `offsite_delivery_stuck` event from **2026-08-07 10:19** (not recurring) reads as current state. **Three independent surfaces agreeing on a wrong picture is how a working backup gets "fixed".** It did exactly that here: the rehearsal reported a fleet-wide off-site outage to the operator and had to retract it. **Note the true half:** demo-felhom IS genuinely stuck (`offsite.state=needs_credential`, no run has ever succeeded) → **R-278** | **READY (S) — NEW 2026-08-09** | — | Render the offsite object on the offsite row; show bytes not rounded GB; distinguish a live alarm from event history | CC | | **R-279** | **There is no operator-triggerable off-site backup.** The only route to `POST /backup/offbox/run` is the customer's own dashboard session; `signed_jobs` carries opaque operator-SIGNED blobs and the hub holds no signing key. This cost the rehearsal a stop: preparing the run needed one off-site push and there was no operator path to it. Sibling of **R-177** (no operator-triggerable fill check) | **READY (XS) — NEW 2026-08-09** | — | Same shape as R-177; solve both together | CC | -| **R-282** | **One secret, three different Hungarian names, and the email sends the customer to a page their box is not showing.** Sending it from the hub is „**Visszaállító** kód küldése"; the email that arrives is subject „Jelszó-**visszaállítási** kód", body „**Visszaállító** kód: …", and it instructs *„Add meg a vezérlőpult »**Elfelejtett jelszó**« oldalán"*; the page the box actually serves is „A szerver **beállítása**" asking for a „**Beállító** kód". **A rebuilt box shows a SETUP page and the hub can only send a RESET mail** (because hub-side the customer is still `claimed_at 2026-07-21`), so the instruction names a route that does not exist on screen. **It does work if you ignore the instructions** — the reset code was accepted on the setup page (302 + session), so this is naming, not function. **It cost this session real time and one wasted code:** the operator supplied a 3-word Hungarian code believing it was the recovery code, because the hub calls the claim code „Visszaállító kód" and the ESCROW code is also „Visszaállító kód" — the only reliable discriminator is length (claim = 3 Hungarian words; recovery = **10** EFF-list words, and the recovery screen does say „(tíz szó)") | **READY (S) — NEW 2026-08-09** | — | Pick one name per secret and use it on all three surfaces; make the mail's page reference match what a rebuilt box actually shows | CC | +| **R-282** | **One secret, three different Hungarian names, and the email sends the customer to a page their box is not showing.** Sending it from the hub is „**Visszaállító** kód küldése"; the email that arrives is subject „Jelszó-**visszaállítási** kód", body „**Visszaállító** kód: …", and it instructs *„Add meg a vezérlőpult »**Elfelejtett jelszó**« oldalán"*; the page the box actually serves is „A szerver **beállítása**" asking for a „**Beállító** kód". **A rebuilt box shows a SETUP page and the hub can only send a RESET mail** (because hub-side the customer is still `claimed_at 2026-07-21`), so the instruction names a route that does not exist on screen. **It does work if you ignore the instructions** — the reset code was accepted on the setup page (302 + session), so this is naming, not function. **It cost this session real time and one wasted code:** the operator supplied a 3-word Hungarian code believing it was the recovery code, because the hub calls the claim code „Visszaállító kód" and the ESCROW code is also „Visszaállító kód" — the only reliable discriminator is length (claim = 3 Hungarian words; recovery = **10** EFF-list words, and the recovery screen does say „(tíz szó)") | **VERIFY** (2026-10-03 triage: Naming fixed on box (R-295), hub and mail page reference (R-295 hub, reenroll mail) and R-323 — OPEN-ITEMS.md R-327 row (line 578)) — **READY (S) — NEW 2026-08-09** | — | Pick one name per secret and use it on all three surfaces; make the mail's page reference match what a rebuilt box actually shows | CC | | **R-283** | **After a rebuild the hub says "Claimed 18d ago" while the box serves its first-run setup page.** `customer_claims` for demo-hp still read `claimed_at 2026-07-21 16:29:25`, `generation 2`, `issued_at 2026-08-03` while the freshly provisioned guest — whose `settings.json` is new — correctly showed „A szerver beállítása". The two sides never reconcile: the hub's claim state survives a guest rebuild and the box's does not. Consequences: the operator's screen says the box is claimed when it is not, a resend produces a RESET code instead of a SETUP code (→ **R-282**), and any previously issued code fails with *„Hibás vagy lejárt kód"* — a message that is technically true and tells the customer nothing about the real cause, namely their own reinstall. Mirror image of **R-214/R-235** (an already-paired box still told to pair itself) | **READY (S) — NEW 2026-08-09** | — | Let a report from a box carrying no claim state clear the hub's, or show both sides on the operator page | CC | -| **R-284** | **„A kiválasztott tárhely majdnem megtelt." on a store that is 93 % FREE — an apparent inverted threshold.** Calibre-Web's deploy page rendered `