From 3f4fb3825f0658038285ffe69205e5838d1318b0 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Fri, 7 Aug 2026 17:24:27 +0200 Subject: [PATCH] =?UTF-8?q?R-201=20CLOSED=20=E2=80=94=20the=20unaided=20re?= =?UTF-8?q?covery=20journey=20passes,=20both=20halves,=20on=20the=20fifth?= =?UTF-8?q?=20walk?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Capability map: the unaided-recovery row turns FAILED -> PROVEN-LIVE, scoped, with what it still does not claim stated in the row itself: shape (c) did not fire positively (with the mint guard holding there is no local key, so the offer comes from shape (a)); and 'unaided' here means possible-without-a-shell, not obvious, because two obstacles are unsignposted. OPEN-ITEMS: R-201 closed with its evidence. Five new rows R-249..R-253 (the retrieval passphrase in page HTML; the host-key scan ladder vs AAAA settle; the listing's per-tag rows; the two unsignposted restore steps). R-243 annotated rather than re-filed: on a REBUILD offsite_delivery_stuck does not skip, so the row's gap is narrower than it reads. STATUS.md: headline changed, and trimmed 97 -> 92 lines rather than extended, per its own header. Teardown recorded as OWED with its before-measurements, the stop-and-age gate, and the positive controls that must survive. --- REPORT-walk5-r201-2026-08-07.md | 221 ++++++++++++++++++ STATUS.md | 67 +++--- .../architecture/00-capability-map.md | 2 +- documentation/backlog/OPEN-ITEMS.md | 9 +- .../walk5-r201-2026-08-07/teardown-owed.md | 73 ++++++ 5 files changed, 338 insertions(+), 34 deletions(-) create mode 100644 REPORT-walk5-r201-2026-08-07.md create mode 100644 documentation/tests/walk5-r201-2026-08-07/teardown-owed.md diff --git a/REPORT-walk5-r201-2026-08-07.md b/REPORT-walk5-r201-2026-08-07.md new file mode 100644 index 0000000..e83a5b1 --- /dev/null +++ b/REPORT-walk5-r201-2026-08-07.md @@ -0,0 +1,221 @@ +# REPORT — the fifth walk (R-201), 2026-08-07 + +**Runbook-style validation, supervised. One machine destroyed on purpose, with the operator's +confirmation at the §6 STOP. No product code written.** Venue: `demo-hp` VM **325 `walk5-appliance`**, +customer `walk5`, host `walk5-4bada5`. Full evidence: `documentation/tests/walk5-r201-2026-08-07/`. + +--- + +## 1. THE VERDICT — both halves, separately + +### THE DATA: **PASS** + +All three sentinels came back **byte-identical** out of snapshot `5b0f20f7`, under the key recovered +from the sealed package with R: + +``` +11eb7fb2d1891f62a685e3d9a9fda44e0a342d437d4aebaa15fcb4c0fd79d539 61 WALK5-SENTINEL-A.txt +6b504d1e83bf4c673334d0b9e701f3e5f07311d954fdc4463a6b0a55b261ae32 66 WALK5-őrszem-ékezetes-árvíztűrő.txt +0baaf402733e6a85f3d9ce8c4c3a57b05692f5f6733da6a23da76d2cd7635210 12582912 WALK5-SENTINEL-C-12MB.bin +name hex 57414c4b352d c591 72737a656d2d c3a9 6b657a657465732d c3a1 7276 c3ad 7a74 c5b1 72 c591 2e747874 +``` + +Identical to Phase A in every byte **including the accented filename's name bytes**. Read back with +`os.listdir` on a **bytes** path, so no decode/encode round trip could launder a `U+FFFD` into looking +correct — the check that caught this three times before. + +### THE JOURNEY: **PASS — the first time in five walks** + +**No step needed a command line inside the guest.** Everything that *progressed* the journey was an +HTTP request a browser makes. The previous walk needed **three** guest command lines; this needed +**zero**. The reset-code hatch was used **once, in Phase A**, where §3 permits it. + +**Named per §3 so the claim is not read wider than it is** — the guest command lines used were the +`w5watch.log` sampler, `docker logs`, the settings reads, the restic listing and the final sentinel +verification. **All instrumentation:** none changed state, none was needed to progress, and removing +them all would have changed nothing but my ability to describe what happened. + +--- + +## 2. §5's OBSERVATION — the first live exercise of R-241's fix, and it stands on its own + +**At 14:58:52Z, unaided, before anyone had logged in:** + +``` +[WARN] [offbox] NOT minting a repository password: the hub holds a sealed recovery package for this + box, and a fresh key would orphan the history that package protects (R-241). The transport is + configured; the tier stays down until the customer's recovery code places the escrowed key. +[INFO] [offbox] apply-offsite: transport configured …, tier HELD awaiting the escrowed key +``` + +| §5 asks | answer | +|---|---| +| does it declare a need, and when is it staged/collected? | **yes** — declared `needs_credential` 14:38:49Z and 14:53:49Z; hub staged **14:56:34Z**; collected + applied **14:58:52Z**. **Zero human actions**, on a box not yet claimed | +| **is any repository key written?** | **NO** — sampled every ~20 s from 14:24:52Z; `repo_password` absent at every sample. The directory holds `applied_marker`, `known_hosts`, `ssh_key` and nothing else | +| what state does it report instead? | `enabled=true`, `escrow_state=pending`, no key → the derived **`awaiting_recovery_key`** holding state | +| the two fingerprints, before login | hub's package seals `eabf427c7274…144f`; **the box holds NONE** | + +**Positive control, because an absent line is not evidence:** the scheduler logged +`agent-channel-health` ×5, `stack-scan` ×2, `system-health`, `backup-cache` and +`offsite-credential-retry` in the same window. The absence of a mint is **explained**, not merely +observed. + +> **HONEST SCOPE.** With the mint guard holding there is **no local key**, so the offer fires on +> **shape (a)**, not shape (c). Shape (c) was measured in **Phase A, in its negative half** — hub hash +> == local hash, correctly silent. **This walk proves the mint guard positively and the discriminator +> negatively.** A positive shape-(c) firing needs a box holding a *different* key, which v0.206.0 now +> prevents from arising by itself. + +--- + +## 3. THE RTO + +**71.7 s**, login (15:04:25.417Z) → open store (15:05:37.156Z). Of that, **12.44 s was the unseal +itself**; ~22 s was **my own harness retry** (I scraped the CSRF token from a `` tag the recovery +page does not carry, got a 403, re-read it from the form). **A customer clicking the button sees +≈50 s.** Both numbers are given because 71.7 s is what was measured. + +--- + +## 4. DEAD ENDS, in the customer's terms + +**By §3's definition — something needing a shell inside the guest — there were ZERO.** Two obstacles +were met, both cleared **from the dashboard**, and **neither is signposted**: + +| # | what the customer sees | what got past it | known? | +|---|---|---|---| +| 1 | „nincs elérhető adatmeghajtó a visszaállításhoz" | Tárhely → Meghajtók → „Meglévő meghajtó csatolása" re-registers both surviving disks | **NEW — R-252** | +| 2 | „a(z) calibre-web nincs telepítve — előbb állítsd helyre az alkalmazást" — **on a page that says three lines above „Nincs telepítve — a visszaállítás előbb újratelepíti"** | redeploy from the catalogue (~90 s), then re-run the restore | **NEW — R-253** | + +**So: the machinery works end to end and the data is provably safe. The unaided journey now succeeds, +and it succeeds through two obstacles a customer must guess their way past.** + +--- + +## 5. WHAT EACH INSTALL LANDED ON — and delivery is part of the pass + +| | vouched | first install | after the rebuild | +|---|---|---|---| +| agent | 0.127.0 | **0.127.0** | **0.127.0** | +| controller | golden 0.206.0 | **0.206.0** | **0.206.0** | + +**No hand upgrade either time, and no downgrade on the reinstall.** This is the first walk of the five +where the box under test **is the box a customer receives** — R-239's delivery gap, the headline of both +previous walks, is closed for this run. + +--- + +## 6. THE RECOVERY SCREEN, QUOTED + +Appeared **without being sought**: `/` → 302 `/launcher` → 302 **`/recovery`**. + +> „Ezt a gépet újratelepítették. **A korábbi, házon kívüli mentéseid megvannak** — a Felhom központi +> rendszere őriz hozzájuk egy lezárt csomagot, amelyet **2026-08-07T12:51:02Z** zártunk le. […] +> **A helyreállítási kódot senki nem tudja pótolni** — sem a Felhom, sem az ügyfélszolgálat, sem az +> üzemeltető. […] Ha megadod a kódot, **feloldjuk a mentéseid zárolását és megmutatjuk, mi van +> bennük**. **Ebben a lépésben semmit nem állítunk vissza és semmi nem változik.**" + +The sealed-at timestamp **matches `host_escrow.created_at` exactly**. *(Copy wart, recorded not filed: +it is a raw ISO-8601 string on a Hungarian customer screen where every other date reads `2026-08-07 14:57`.)* + +--- + +## 7. THE LISTING, against Phase A + +| | Phase A | the screen | +|---|---|---| +| app | `calibre-web` | **`calibre-web`** ✅ | +| when | `5b0f20f7` @ 12:57:41Z | **2026-08-07 14:57** ✅ (CEST) | +| size | 12.784 MiB | **12.8 MB** ✅ | + +A second row `felhom-offbox · 12.8 MB` also appears — the tier's marker tag rendered as an app, and the +total doubled. **R-251.** + +--- + +## 8. §4.6's TWO PRE-DESTRUCTION CHECKS — both pass, neither previously exercised on a clean box + +- **The recovery offer is NOT shown**: `/` → `/launcher`, **zero** recovery mentions on either landing + page, `/recovery` 302s. And the reason is measured, not assumed — box key, box's ACK-cached hub hash + and hub `restic_pw_sha256` are all `eabf427c7274…144f`, so **shape (c) compares equal and stays + silent**. +- **The restore page lists the app with the future-backup toggle OFF** — identical rendering both ways. + **R-237's fix, live**; the last walk measured 0 entries and a 302 here. + +--- + +## 9. PHASE A's SEVEN RECORDS + +Controller `0.206.0` · agent `0.127.0` · PBS wrapper **matches vouched** · guests 1/1 · DR recipe +**present** · key escrow **present** · snapshot `5b0f20f7` · 1 snapshot · **12.0 MB** (12 611 563 B) · +`host_escrow` blob **383 B**, `identity_blob` **572 B**, `stale_at` **NULL**, 0 superseded rows · +escrowed key fingerprint `a6:86:f7:fb:…:4c:f9` · box key == hub hash == `eabf427c7274…144f`. + +**Sentinels listed BY NAME** out of the snapshot with `restic ls latest --long` — see §1. + +> **The §4.5 gate earned its place again, and this time it caught MY fault.** The first off-site run +> reported **`ok` in 28 s with 0 snapshots**: I had sent the per-app toggle as `enabled=1`, and the +> handler accepts only `on`/`true`, so it recorded *off* and the run correctly backed up nothing. +> Re-toggled, selection verified in the rendered page, re-run → 1 snapshot, 12.0 MB. **A green tick is +> not evidence a file is in a snapshot.** + +--- + +## 10. R — SHREDDED, with a working control + +One `0600` file on **DooPlex only**, never rendered, never an argument, never a log line. Shape only: +**82 characters, 10 hyphen-separated tokens**. + +``` +plant → ~/.config/walk5/R_PLANTED_CONTROL.txt +sweep → 2 hits (the real file + the planted control) ← the control PROVES the sweep works +shred → both, then the pattern file itself +sweep → 0 hits +``` + +**Every sweep path was asserted to exist first** — a sweep pointed at a missing path returns zero for +the wrong reason. The appliance's copy was `shred -u`'d mid-walk and its absence verified. + +--- + +## 11. NEW FINDINGS — the highest register ID moved **R-248 → R-253** + +| ID | | +|---|---| +| **R-249** | **The retrieval passphrase ships in the customer page's HTML** (`data-secret`), so any headless read puts it in a transcript — with no reveal action and **no audit event**, where the break-glass credential emits one. Found by doing it. **MEDIUM** | +| **R-250** | **A customer create can fail fail-closed** because the host-key scan ladder (~60 s) is shorter than the fresh sub-account's DNS/**AAAA-before-A** settle (~100 s measured). Retry is safe and idempotent; nothing says so. **LOW-MEDIUM** | +| **R-251** | The recovery listing renders **one row per restic tag**, showing the customer an "app" they never installed and their data counted twice. **Cosmetic** | +| **R-252** | After a rebuild the restore refuses — **the drives lost their registration** — and nothing on the recovery path says to re-attach them | +| **R-253** | The restore refuses because the app is not installed, **on a page that says the restore reinstalls it**. Two shipped sentences that contradict each other, in the customer's language, at the last step of a recovery | + +**Recorded against an existing row rather than minted:** **R-243** claims `offsite_delivery_stuck` +"skips the applied shape". **On a rebuild it does not skip** — 88 s after the destruction the hub +emitted the warning and wrote an **operator-channel** `notification_log` row naming a guest rebuild as +the cause, correctly. The gap is real for the state R-243 describes and **not** for the state a rebuild +produces; the row is annotated so it is not read wider than it measures. + +--- + +## 12. TEARDOWN — OWED, not done + +The machine is the evidence until the verdict is written. Full enumeration, the "before" measurements, +the stop-and-age gate, and the positive controls that must survive: +`documentation/tests/walk5-r201-2026-08-07/teardown-owed.md`. **R-244's residue will grow by this +venue.** + +--- + +## 13. WHAT DID NOT RUN, AND WHY + +- **A positive shape-(c) firing.** Structurally unreachable on a healthy v0.206.0 rebuild — see §2. +- **A soak / scheduled cycle.** The previous walk covered it; §7 does not ask for one and adding it + would have delayed the destruction past the operator's window. +- **Any product code.** §0 forbids it: five findings were filed and the walk continued. +- **Teardown.** §10 defers it deliberately. +- **`felhom-offbox`'s second listing row and the raw ISO date** were observed, not chased. + +## Documents updated + +`00-capability-map.md` (the unaided-recovery row → **PROVEN-LIVE, scoped**), `OPEN-ITEMS.md` +(**R-201 CLOSED**; R-249…R-253 filed; R-243 annotated), `STATUS.md` (headline changed; trimmed 97 → 92 +lines rather than extended), and the journal + teardown ledger. diff --git a/STATUS.md b/STATUS.md index c346bdf..d647ae4 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,16 +1,14 @@ # STATUS — what works, what's broken, what's next -**Updated 2026-08-08.** +**Updated 2026-08-07.** > **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates > part of it in plain words, and **nothing may exist only here**. Not `CONTEXT.md`, which is technical > state written for Claude Code. **Items, not paragraphs. One screen.** If it does not fit, something > belongs in the register instead. > -> *Rebuilt from the register on 2026-08-08. It had reached 258 lines; its "waiting on you" list asked -> for two things already shipped and carried a stray line reading only "Nothing."; and it mixed the -> DooPlex infrastructure work in with the product. The old "what shipped recently" log — 100 lines of -> it — is what the per-repo `CHANGELOG.md` files and the register are for, and is not restated here.* +> *Rebuilt from the register on 2026-08-07, from 258 lines. The old "what shipped recently" log is what +> the per-repo `CHANGELOG.md` files and the register are for, and is not restated here.* ## What works @@ -19,21 +17,31 @@ sets their own password. They install apps from a catalogue of fifty-three, shar network, and open apps from a launcher or a shared link. Backups run on their own to three places — the machine's drive, a second drive, and an encrypted off-site copy. -**The backup promise is proved.** A machine has been destroyed on purpose and its files came back byte -for byte identical — three separate times, including a filename with Hungarian accents. +**The backup promise is proved, and so is getting the data back yourself.** A machine has been +destroyed on purpose and its files came back byte for byte identical — four times now, including a +filename with Hungarian accents. On **2026-08-07 the household's own journey passed for the first +time**: someone with a browser and their recovery code got everything back with **no command line +inside the machine at any point**. From logging in to seeing what is in the store took **72 seconds**. +*(R-201 — closed. Two rough edges remain, below.)* ## What's broken -- **A household still cannot get their own data back unaided.** Every individual link now works; no - single walk has completed end to end without someone stepping in. *(R-201)* +- **The recovery works, but two steps are unsignposted, and one of them the machine gets wrong.** After + a rebuild the restore stops with "no data drive available" and never says the drives must be + re-attached *(R-252)*; then it refuses because the app is not installed — on a page that says, three + lines above, that the restore will reinstall it *(R-253)*. Both are fixable from the dashboard in a + couple of minutes. Neither is something a household would work out on its own. +- **The passphrase that fetches a customer's whole configuration is sitting in the page's HTML**, behind + a Reveal button that only hides it visually. Anything that reads the page rather than looking at it + gets it, with no record that it was read. *(R-249)* - **The machine's own screen keeps telling an already-paired box to pair itself** — 25 minutes after it was paired, on a screen that promises it refreshes itself. *(R-214, R-235)* - **A rebuilt machine cannot create a new recovery code at all.** *(R-221)* - **A backup that covered nothing still calls itself „Sikeres".** The state is honest; the word is not. *(R-240)* -- **A machine waiting for its recovery code raises no alarm to us.** It quietly stops making off-site - backups, and three separate safety nets each correctly decide it is not their business. The - household can see it; we cannot. *(R-243)* +- **A machine waiting for its recovery code can stop making off-site backups without alarming us.** + Measured on 7 August: after a *rebuild* we ARE told, promptly and correctly. The gap is narrower than + it read — it is a machine that reaches the state without a working tier behind it. *(R-243)* - **The card offering to reopen set-aside backups promises more than we can deliver** — we keep the old sealed package, but nothing can open it. *(R-202)* - **Deleting a customer leaves rows behind** on every test machine ever torn down, while reporting a @@ -42,31 +50,29 @@ for byte identical — three separate times, including a filename with Hungarian ## Found today -- **A leftover flag had been silently switching off the new recovery detection on one demo machine - since 4 August — found, and cleared with your approval.** A Re-issue during the recovery drill set it, using code we removed the next day. - **The flag is wrong** — the sealed package does cover the key that machine is using, and the two - fingerprints match exactly. Because of it the hub withholds a figure the machine needs, so the - machine tells its owner *"create a new recovery code"* — the one act that would put their old backups - beyond reach. **A freshly installed machine cannot reach this state**, because nothing has set that - flag since 5 August. **Cleared the same day; the machine now compares its key correctly again.** What - is still owed is a ruling on the flag itself: nothing sets it, nothing can see it, and it changes - what a household is told. *(R-246, R-247, R-248)* +- **The recovery walk passed** — see above. It also turned up five things: the retrieval passphrase + sitting in the page HTML *(R-249)*, a customer create that can fail on a fresh sub-account's DNS and + only says "try again" by not saying anything *(R-250)*, a recovery listing that shows the household + an "app" they never installed *(R-251)*, and the two unsignposted restore steps *(R-252, R-253)*. +- **The leftover flag that had been switching off recovery detection on one demo machine since + 4 August was cleared with your approval**, and the fresh machine built for the walk confirmed it + cannot reach that state. Still owed: a ruling on the flag itself — nothing sets it, nothing can see + it, and it changes what a household is told. *(R-246, R-247, R-248)* ## What we're working on -- **The walk that either finishes the arc or says why not.** It needs the base image current (done - today) and a machine whose recovery detection is not silently switched off (answered today). *(R-201)* +- **Smoothing the two rough edges the walk found**, so the recovery reads as one path rather than + three. *(R-252, R-253)* - **Proving the hub really keeps the old sealed key** when a machine re-seals. Never run outside a test; needs a second deliberate wipe and its own session. *(R-198)* ## Waiting on you -- **Nothing.** The new base image was approved and is live — every future installation now carries - this week's fixes. *(R-239, R-242)* +- **Nothing blocking.** One thing is owed by us, not you: the walk's machine (`walk5`, VM 325) is + still standing as the evidence and needs tearing down. -*Nothing else is pending. R-245 — whether an undecided household is auto-abandoned after 30 days — was -**settled on 7 August** (we do not build it, and the reasoning is recorded). It has been re-filed as a -decision taken rather than a question sitting in your queue.* +*R-245 — whether an undecided household is auto-abandoned after 30 days — was settled on 7 August: we +do not build it, and the reasoning is recorded.* ## DooPlex infrastructure — separate from the product @@ -80,8 +86,7 @@ being readable.* free — clutter, not space. Nothing deleted. *(R-210)* - **The hub password needs rotating** — a diagnostic printed it into a session log; nothing suggests anyone else saw it. *(R-132)* -- **One thing to read after DooPlex next restarts.** The move to the second SSD has never been through - a reboot; it now writes PASS or FAIL to `/var/log/felhom-store-postboot-check.log` on every start. - On PASS, 34 GB comes back. *(R-209a)* +- **One thing to read after DooPlex next restarts** — the second-SSD move has never survived a reboot; + it writes PASS/FAIL to `/var/log/felhom-store-postboot-check.log`. On PASS, 34 GB comes back. *(R-209a)* - **Backup scripts on DooPlex are unversioned host state.** *(R-231)* - **Instruction-file follow-ups**, each needing a decision rather than an edit. *(R-229, R-230)* diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index d202f28..8c926ff 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -41,7 +41,7 @@ | **The offsite repository password can be RECOVERED from the sealed escrow with the customer's recovery code** | hub v0.94.0, agent v0.125.0, controller v0.195.0 | **PROVEN-LIVE (2026-08-04)** | On demo-felhom, through the real endpoints end to end: the box fetched its own sealed blob from the hub with its own per-host credential (hub log: *escrow blob SERVED … 572 opaque bytes, self_scope=true*), the agent unsealed it with the customer's recovery code, and the extracted repository password's sha256 was **byte-identical** to the one on disk — `c60c8bc737a6…`, which is ALSO the hash the hub had independently stored, so three sources agree. Five minutes earlier the same path with a WRONG code failed closed at age's KDF with nothing written, which proves links 6 and 7 ran independently of the success. R was searched for afterwards and found in 0 log lines and 0 files, with a positive control confirming the search would have found it. Evidence: per-repo CHANGELOGs; `audits/RECON-offsite-dr-chain-2026-08-04.md` §3 links 6–8 | **WHAT THIS ROW DOES NOT CLAIM, stated because the previous over-claim here was struck out four hours earlier.** It covers the KEY, not the DATA. **R-199 BACK-POINTER (omitted when this row was written): the recovery chain's link inventory and the per-link status live in `audits/RECON-offsite-dr-chain-2026-08-04.md` §3; links 1–8 are walked, 9–11 are not.** **Re-confirmed 2026-08-04 evening by an attempt to prove the DATA half:** the R-201 drill was prepared on demo-hp and **halted before the wipe** — the sentinel file was not in the off-site snapshot (R-203), so the wipe would have destroyed it and proven nothing. **No file has still ever been restored from an off-site backup after a wipe** (`audits/DRILL-r201-offsite-recovery-2026-08-04.md`). The install half of the chain (controller v0.196.0 `--recover-offsite-install`, R-200) is likewise unit-proven only — it has never run against a live recovery. A recovered password has never been **installed** (the diagnostic compares and refuses to write, by design), no existing repository has ever been **reopened** under one, and **no file has ever been restored** from an off-site history via a recovered key. Links 9–11 of the chain are open (R-200's remaining half, R-201). The proof also used a box whose local key still exists — the rebuilt-box case, where there is nothing to compare against, is exactly what the drill covers and it has not run || DR tier by default: PBS + WireGuard base infra on every install, hub-controlled activation | installer v1.15, agent v0.86, hub v0.51 | **PROVEN-LIVE** (2026-07-21) | `DRILL-day0-take2-2026-07-12` §2 (WG enabled both modes, PBS-DR descriptor auto-provisioned ~1s after WG registration, zero operator steps); ships installer v1.15/agent v0.86/hub v0.51 | Live only on demo/drill fleet. (Cited spike was slice-0 mechanics — shipped nothing; corrected.) **⚠ The candidate upgrade to PROVEN-LIVE is WITHDRAWN — the 2026-07-18 rehearsal produced a live counter-example (R-39).** On the reborn N100 the descriptor auto-provisioned and the agent reported `converged state=applied` (16:45:53), yet **the storage is dead**: `pvesm status` → `felhom-pbs: error fetching datastores - 401 Unauthorized` / `inactive`, and a direct probe with the stored credential returns **401 on every endpoint including `/version`** while the WG transport is healthy (handshake 9 s, 27.9 ms RTT) — i.e. authentication failure, not ACL scope. Root cause in the evidence: **the hub minted a SECOND token secret at 16:47:52, two minutes after the agent had applied the first, and `consumed_at` is still NULL**; the converged state machine will not re-apply, and the agent's 15-minute verify loop **cannot even read the credential to notice** (`open /etc/pve/priv/storage/felhom-pbs.pw: permission denied` — non-root agent reading a file it writes through a root wrapper). A tier that reports `applied` while silently unable to authenticate is exactly the shape that must not carry a PROVEN-LIVE badge. See `tests/VALIDATION-n100-rehearsal-2026-07-18.md` F2 and `pbs-dr-state.txt`. **agent v0.89.0 closes the F4 non-default-storage-id gap (R-22) — PROVEN-LIVE 2026-07-17:** the reconcile self-grants the ACL through the root wrapper on a pre-check 403 instead of dead-locking. Reproduced F4 on the demo (marker moved aside = reinstall fresh-state + felhom-offsite ACLs revoked) → next reconcile tick `pbsdr: pre-check 403 … self-granting … (R-22)` → `converged state=adopted` in ~3 s, ACLs self-restored, `pvesm status felhom-offsite`=active, zero operator action. No more one-shot `pveum` grant **2026-07-21 — the R-39 fleet fix SHIPPED (hub v0.68.0 + agent v0.91.2), closing the self-heal chain end to end.** The three defects that let a box be `applied` and dead simultaneously are each addressed: the hub stamps a monotonic `secret_generation` into the descriptor so a credential re-key finally MOVES the content hash the agent re-applies on; the wrapper gains a narrow `read` verb so the non-root agent can read the credential it writes (it never could — `/etc/pve/priv` is 0700 root:www-data, which made the verify loop blind by construction); and `pbs.ProbeAuth` turns a 401 into a loud `auth_failed` that the existing `pbsdrheal` damper escalates to a fresh mint. Plus a consumed_at honesty gauge for the disagreement no single tier can see (box says `applied`, hub's staged secret never consumed). Proven live on felhom-pve: the agent read its credential through the wrapper (`rc=0`) and probed successfully (`credential probe OK storage=felhom-pbs`). **STOP-2 RAN 2026-07-21 AND THE CHAIN CLOSED — 13 SECONDS, operator click to converged.** The operator pressed **Re-issue PBS credentials**; the identical click on 2026-07-18 did nothing at all. Full chain (hub UTC / host CEST = UTC+2): `08:39:31Z` hub mints a fresh secret, **generation 0 → 1**, and the descriptor gains `"secret_generation": 1` — with `token_id` and `fingerprint` **byte-identical**, i.e. exactly the re-key shape that used to be invisible → `10:39:34` the agent READS its credential through the wrapper (leg b — the read that was impossible until v0.91.0) → `10:39:38` **`ERROR pbsdr: the DR endpoint REJECTED this box's credential — the tier is applied and DEAD` `previous_state=applied`** (leg c: the exact R-39 failure state, detected out loud for the first time ever) → `10:39:45` **`one-time token secret consumed`** `secret_len=36` (leg a: **NO short-circuit** — this is the line that never appeared on 2026-07-18) → `10:39:45` `felhom-pbs-apply reconcile` (the set-only wrapper, no `--server`) → `10:39:47` **`pbsdr: converged state=applied`**. Corroboration: the agent marker hash moved to `afbb3b41…` (it was byte-identical to the pre-reissue marker in the failure); `consumed_at` stamped `08:39:45Z`; the on-disk secret's mtime moved `2026-07-18 20:28:52` → `2026-07-21 10:39:45`; a live probe with the NEW credential returns **200**; three consecutive hub reports trace the whole state machine `applied → auth_failed → applied`; and **zero** `pbsdr_selfheal` escalations fired — the box healed through the descriptor path before the damper was ever needed, with exactly ONE mint and ONE consume and no `consumed-failed.json`. **Row upgraded to PROVEN-LIVE (2026-07-21).** Evidence: `felhom-agent/REPORT.md` (2026-07-21). | | **A customer's file survives a machine rebuild and comes back — the whole off-site story, end to end** | hub v0.94.0, agent v0.125.0, controller v0.197.0 | **PROVEN-LIVE (2026-08-04 night drill)** | demo-hp's controller data volume was destroyed and the sentinel deleted from disk. The customer's recovery code then produced `8a9e33aa4da6…` — **byte-identical** to the pre-wipe on-disk key and to the hub's independent record — and installed cleanly on the bare box. `identity_blob` was unchanged across the wipe (572 B, `updated_at` still 11:11:37). Evidence: `audits/DRILL-r201-night-run-2026-08-04.md` §2 | **WHAT IT DOES AND DOES NOT CLAIM.** PROVEN: after a real rebuild the key recovers byte-identical, the EXISTING repository opens (3 snapshots, 42 026 B — the pre-wipe size; not a fresh history), and the customer restore flow returns the file byte-identical. **ALL FOUR MANUAL INTERVENTIONS ARE CLOSED, AND THE CUSTOMER IS NOW OFFERED THE RECOVERY (2026-08-05). ONE QUALIFIER REMAINS AND IT IS NOT THE OLD ONE — state it precisely.** The 2026-08-04 drill needed four interventions between "the key is recoverable" and "the file is back", none of them in any design document (R-204). **Items 1–3 closed in controller v0.198.0 + hub v0.95.0** (the reset code works first time; a Re-issue no longer marks a healthy escrow stale; a unit restore states it returned the app's definition and database and NOT the customer's files). **Item 4 closed in controller v0.199.0 + hub v0.96.0** (a rebuilt box DECLARES `offsite.state=needs_credential` and `internal/offsiteheal` re-arms the stored credential before minting). **AND THE "needs someone who knows to look" HALF IS NOW GONE TOO (controller v0.200.0, R-193):** a full-page recovery screen takes over the landing pages while the hub holds a sealed package this box cannot open, explains that nobody can replace a lost recovery code, takes the code, opens the repository and lists what is in it — proven live on demo-felhom, which is genuinely in that state (`/launcher` → 302 `/recovery`; a wrong code refused with nothing written, three times, no lockout; the code found in no file, no container log and no debug ring, with a planted-copy positive control that first exposed a mis-aimed sweep). **WHAT THE QUALIFIER IS NOW, and it is narrower:** (1) **the final unlock has never been driven with a CORRECT code through the page** — no recovery code was kept for demo-felhom's orphaned history and demo-hp's is operator-held out of band, so the live run exercised the whole chain (handler → agent → hub fetch → age KDF) and stopped at the unseal; the install and listing rest on tests. (2) **putting files back in place is deliberately NOT part of this** — restore stays per-app, and the step after the listing is R-213. (3) **the journey has still not been re-walked end to end since these fixes** — the closures are proven individually, not as one uninterrupted run. That re-walk is one more drill and it is what is owed. Scope: `demo-hp`, a controller-data-volume rebuild — **NOT** a total host loss, and **NOT** a guest reprovision. **⚠ THE RE-WALK RAN ON 2026-08-05 (CAMPAIGN-11) AND IT FAILED. THIS ROW STAYS AS IT IS — it records a key recovering and a file returning, which remains true — but the JOURNEY row below is the honest verdict and this row must not be read as covering it.** The campaign built a throwaway appliance from the published ISO, destroyed it, and walked the customer's route: **the data half PASSED** (three sentinels byte-identical, including a 12 MB binary and a non-ASCII Hungarian filename, out of the pre-wipe snapshot in 16 s) and **the journey half FAILED** — four operator interventions, three needing root on the appliance, and the first of them was the machine telling the customer their perfectly correct recovery code was wrong (R-216). **R-198's retention is NO LONGER unit-proven — it was PROVEN IN PRODUCTION on 2026-08-05** (CAMPAIGN-11 Phase 3): the first supersession since the fix retained `identity_blob` at **572 B, byte-length exact**, carrying the OLD key `626e424670248db3` while the current row moved to the newly minted `e11a6c542b73477a`; `offsite_repo_key_changed` and `escrow_superseded` both fired at the instant of supersession, and the next off-site run REFUSED as `orphaned` rather than starting a fresh history. Evidence: `tests/campaign11-evidence-2026-08-05/journal.md` | -| **The customer's own UNAIDED recovery journey, end to end** | controller v0.201.0, hub v0.97.1, agent v0.125.0 | **FAILED (2026-08-05, CAMPAIGN-11 Phase 1) — and it stays FAILED until a re-walk passes** | A throwaway appliance was built from the published ISO on demo-hp, claimed, given three sentinels, escrowed, backed up off-site, then destroyed — guest purged and both enrolled drives wiped. The rebuild and the customer's route were then walked with **no command line inside the guest**. **Data: PASS.** All three sentinels byte-identical (`beb9175d…`, `7c8cb0ad…`, `e012e76f…`), including a 12 MB binary and `C11-őrszem-ékezetes-árvíztűrő.txt` — the Hungarian filename survived disk → restic → SFTP → Storage Box → restore → disk intact. Restored in **16 s** out of the pre-wipe snapshot `f3d9cd67`, non-destructively. **Journey: FAIL.** Four interventions, three of them root on the appliance: R-216 (a correct code called wrong), R-218 (the store never opened; the box stopped asking), R-219 (the promised listing structurally could not render), R-220 (the app could not be redeployed — its drives were unenrollable). Evidence: `tests/campaign11-evidence-2026-08-05/journal.md` | **WHAT IT DOES NOT CLAIM, and what the v0.201.0/v0.97.1 fixes do and do not change.** Six of the nine findings are fixed and each is red-proofed (R-216, R-218, R-219, R-217, R-222, R-215) — but **fixes are not a journey.** This row goes green only when the walk is repeated end to end and completes with no operator intervention. **Three findings are deliberately still open and each is a live blocker for some flow:** **R-214** (the console never stops showing a stale pairing code), **R-220** (after a rebuild the customer's drives cannot be re-enrolled — the deploy refuses and the wizard's list is empty, so no app can be redeployed onto its own data drive), **R-221** (a rebuilt box cannot run the escrow ceremony at all, because the pbsdr convergence marker survives on the host while the seeded `escrow.pbs_storage_id` is rewritten out of `agent.json`). **And R-216's fix does not by itself make a NEW box work:** with the hub guard corrected but the Day-0 manifest still vouching agent 0.120.0, a new box is **HELD, not served** — it stops being lied to, which is better, but the feature does not work for it until agent 0.125.0 is vouched. **⚠ CAMPAIGN-11 PHASES 2 + 4 RAN OVERNIGHT 2026-08-05/06 AND THIS ROW STAYS **FAIL**.** **What they add:** eleven injected faults against the recovery journey, each with a positive control, and a full unattended scheduled cycle. **The backup promise strengthened** — the off-site tier ran itself at 04:15 on a box rebuilt twice and set aside hours earlier (`snapshot_count` 1 → 2), all five daily jobs fired exactly once, and nothing on the must-not list fired. **R-217's and R-215's fixes were proven live under exactly their faults**, the set-aside proved it does not delete (12 535 KB byte-exact at the far end), and a wrong code was refused three times with nothing written and no lockout. **What they do NOT add: any progress on the JOURNEY.** These faults are **not a re-walk** — no customer route was walked end to end — and four new findings say the journey got no better where it matters: **R-224** (a hub outage and a stopped agent are BOTH reported to the customer as a bad recovery code, in 0.056 s and 0.030 s — no unseal attempted; the agent's own `err` distinguishes them and it is discarded at the HTTP boundary, and R-216's gate answers `source=version` so it cannot see reachability — **Phase 1's headline defect relocated from the version channel to the transport**), **R-226** (M1, the only message telling a customer to check their typing, is unreachable on any box that has re-escrowed), **R-225** (an unread store renders `0 pillanatkép · 0 GB` above a card saying it holds backups — 1 snapshot and 12 535 KB were really there), **R-228** (the set-aside history is recorded in `orphaned_renamed_to` and shown by nobody — `OrphanedRenamedTo` has zero references in any template or handler). **§4.1 is now MEASURED rather than deduced** — the floor IS served, from the box's own rendered `GetFloor()` and a cold-started controller's settle-gate line. **§4.2's positive half is still NOT measured** and needs a rebuild. Evidence: `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md`, `tests/campaign11-evidence-2026-08-05/journal-phase24.md` **2026-08-06 — THE FIVE PHASE-2 FINDINGS ARE FIXED AND THIS ROW STILL SAYS FAIL.** controller **v0.202.0** + agent **v0.126.0** close R-224, R-226, R-225, R-227, R-228. **What changed:** the unlock path now classifies WHY it failed, from the value and never the text, and the message that mentions typing is reachable from exactly one class — a real refusal — with **unknown defaulting to neutral**. Proven live on the venue with the same wrong code and only the hub's reachability changed: `400 → 502 → 400`. **What it does NOT change: this row.** Fixes are not a journey; nothing here walked a customer end to end, and three of the blockers that made Phase 1 fail are still open (**R-220** in particular is worked around BY HAND on the venue — without that unmount no app can be deployed on a rebuilt box). **What still needs proving:** a re-walk with no operator intervention, and specifically the customer-facing messages end-to-end — they were NOT re-driven, because `/recovery` correctly retires itself once the old data is set aside and restoring that state would be the reconfiguration the fix task forbade. Evidence: `felhom-controller/REPORT.md` **⚠ THE RE-WALK RAN 2026-08-06 (R-201, attended) AND THIS ROW STAYS FAIL — but it is a nearer no.** A brand-new appliance was built from the published ISO, given three sentinels, escrowed, backed up off-site, destroyed (guest purged AND both drives wiped), reinstalled and walked. **THE DATA: PASS** — all three sentinels byte-identical, including a 12 MB binary and an accented Hungarian filename **whose NAME BYTES are also identical**, restored in **23 s** out of the pre-destruction snapshot through the customer's own flow. **THE JOURNEY: FAIL — TWO dead ends against Phase 1's four.** (1) **R-218's CONSUME half**: the hub re-staged the credential and said "the box re-consumes on its next cycle"; a full cycle ran (positive control) and it did not — no customer-reachable action fetches it, and the only lever is a command line INSIDE THE GUEST. (2) **R-220**: the drives are still unenrollable after a rebuild, needing a Proxmox-host unmount, without which no app can be redeployed and the restore page stays empty. **The unaided RTO is therefore STILL UNDEFINED**; the attended figure was 30 m 13 s and must not be quoted as the customer number. **What passed and is new:** the recovery screen appeared WITHOUT being sought, answered all three of its questions with a seal date matching the hub exactly, the emailed reset code worked first try, and the unlock was a real 1.528 s unseal that placed the key. **⚠ AND IT DOES NOT CLAIM DELIVERY:** the fresh install landed on controller 0.201.0 / agent 0.125.0 — the vouched versions, neither carrying the fixes — which were installed BY HAND. Fleet delivery needs a golden carrying the fixed controller and a vouched agent, and nothing was vouched. Evidence: `tests/rewalk-r201-2026-08-06/journal.md`   **⚠ 2026-08-06 (later the same day) — THE TWO DEAD ENDS ARE CLOSED AND VOUCHED, AND THIS ROW STILL SAYS FAIL.** controller **v0.203.0** + agent **v0.127.0** close **R-218's consume half** and **R-220**, and golden **0.203.0** / agent **0.127.0** / `min_agent` **0.127.0** are now VOUCHED — so the delivery gap the re-walk recorded (“the fresh install landed on versions neither carrying the fixes”) is gone. **Both were then proven live on a genuinely rebuilt box** (Part 4 venue, VM 323): the reinstall landed on agent 0.127.0 with **no downgrade and no hand upgrade** (the previous re-walk's reinstall downgraded 0.126.0→0.125.0), `/disks/candidates` returned **both drives** after the guest was purged with the raw mounts still on the surviving host — before the fix, two empty lists — and both **re-attached through the customer endpoint**; the off-site credential was re-staged by `offsiteheal` at 13:24:57Z and **collected by the box on a tick, unaided**. **Why the row is still FAIL:** the walk did not finish. It stopped at the restore surface — **R-237**, the list was keyed on apps that are currently INSTALLED and currently TOGGLED ON for future backups, so a rebuilt box was shown nothing to restore while the repository held its snapshots. Fixed in **controller v0.204.0** (the store is now the source of the list) and proven live on that same venue with the toggle switched OFF, but **no sentinel was restored in that walk**, so the data half is unproven in EITHER direction for this venue. **R-238 was RECLASSIFIED, not fixed as filed:** `mode=full` without `confirm=1` is step 1 of a deliberate two-step that starts no job by design; the endpoint driver did not carry `?full_prep=` forward. The operator drove the same restore to completion in a browser. Its real residue — that step logging nothing at all, including when the headroom gate refuses — is fixed. **R-236 was WITHDRAWN: not a defect.** **This row goes green only when a walk completes end to end with no operator intervention, restoring a byte-identical sentinel.** Evidence: `tests/part4-rewalk-2026-08-06/journal.md`, `tests/teardown-2026-08-06.md`, `felhom-controller/REPORT.md` **⚠ THE FINAL WALK RAN OVERNIGHT 2026-08-06/07 AND THIS ROW STAYS FAIL — the data half passed again, the journey half failed further from the line than before.** A brand-new appliance was built from the published ISO, fixtured, soaked through a real scheduled cycle, destroyed and rebuilt. **THE DATA: PASS** — all three sentinels byte-identical out of snapshot `f5c53b03`, including a 12 MB binary and an accented Hungarian filename **whose name BYTES are identical too**. **THE JOURNEY: FAIL, and this time the customer has NO route at all** — `/` lands on the launcher with no recovery pointer, `/recovery` **302s away**, and the remote-backup page offers to **CREATE a new recovery code**, which would orphan the very history the customer's code protects. **There is no field anywhere to enter the code they hold**, and the operator's documented remedy refuses too. Cause (**R-241**): the credential self-heal — proven working unaided in the same walk, six hours earlier, as its best result — writes a FRESH repository password, which moves the box out of `OffsiteRecoveryOffer()`'s (a) pristine case while (b) orphan-detected is unreachable because runs are blocked by `escrow_state: pending`. **R-218's shape one level up.** **What DID pass, and is new:** the whole credential chain ran end to end with zero human action on an unclaimed box — declare, `offsiteheal` re-stages after two reports, the box's 5-minute retry collects it, tier applied — the **first live sighting** of that success line, settling R-218's consume half and R-236's withdrawal. Also new: **R-239** — a fresh install lands on controller **0.203.0** while 0.205.0 is released, so R-234 and R-237 are written but **not delivered**, proven from the customer's side (T2, T3). **This row goes green only when a walk completes with no operator intervention AND a byte-identical sentinel.** Evidence: `tests/finalwalk-r201-2026-08-07/journal.md` **⚠ 2026-08-07 — R-241 WAS DIAGNOSED BY A READ-ONLY SPIKE ON THE STANDING VENUE, AND THE DIAGNOSIS REVERSES THE FIX. This row is unchanged: still FAIL, and no code was written.** The question was whether R-241 is a screen-predicate defect or a minting defect. **It is a MINTING defect.** The screen was telling the truth — there genuinely was nothing recoverable under the key the box held, because the box minted that key itself over the top of a sealed package it already knew the hub was holding. **Three measurements, from the venue rather than from the earlier report:** (1) `WriteOffboxSecrets` (`offbox.go:411`) mints on ONE input — does the file exist — while `OffsiteRecoveryOffer()` and `needsOffsiteCredential()`, both in the same file, consult `GetHubEscrowIdentityPresent()`; the same fact is available on three paths and used on two. (2) That flag was the **precondition of the chain that reached the minting**: the retry job logs only when the declaration is live, and the venue logged `credential retry: … (the box still declares a need; retrying)` at **02:48:03Z** — thirty minutes and six ticks before the mint at 03:18:06Z. (3) **The box computed the answer and discarded it**: at **03:28:03Z**, thirty-five minutes before the customer looked, `EscrowAutoConfirmer.Reconcile` (`escrow_confirm.go:154`) logged `the hub's escrow blob does not cover the CURRENT repo password (hub hash 30ef574fe492… != local 9b4a9a9dcec7…)`. It is recomputed every report cycle and never persisted. **And the hub explicitly disclaims doing this** — `offsiteheal`'s package doc: *"credential automatic, key customer-present — the ruling this session implements and must not quietly widen."* **Shape (b) is structurally unreachable on this box** (the escrow gate at `offbox.go:743` sits upstream of `ensureOffboxRepo`, the only producer of `RepoState="orphaned"`; positive control: the scheduler was alive, 241 `agent-channel-health` ticks, and `offbox-backup` is a `sched.Daily` leg whose slot fell before the destruction). **Two by-products:** **R-243** — a box in this state silently stops backing up off-site and **no alarm fires** (`isStale` requires `escrowed`, `offsite_delivery_stuck` skips the `applied` shape, `backup_failed` needs a run that never happens); and the trap in the obvious fix — `ResetOrphanedRepo` clears the orphan flag without a ceremony, so a hash-mismatch discriminator alone would re-offer the screen forever to a customer who declined the old data. **Q7:** the „Helyreállítási kód létrehozása" button does NOT destroy the data — R-198's retention copies the `identity_blob` — but it converts a self-service recovery into one needing an unbuilt read path (R-199), and it re-enables the screen while invalidating the code that screen accepts. Evidence: `audits/SPIKE-r241-recovery-offer-2026-08-07.md` **⚠ 2026-08-07 (later) — R-241 IS FIXED (controller v0.206.0 + hub v0.98.0) AND THIS ROW STILL SAYS FAIL.** Three changes, following the spike's ruling rather than the obvious reading: the box **no longer mints a repository key while the hub holds a sealed package** (a conjunction, so a first-time box is untouched; the refusal is a HOLDING state that still writes the transport, declared as `offsite.state=awaiting_recovery_key` and shown inert to every existing hub reader); the **hub-vs-local key comparison that was computed every cycle and discarded is now persisted** and drives the offer as **shape (c)**; and **abandoning starts a 14-day countdown** whose terminal step removes the set-aside store and the sealed package **together**, so the offer ends because the state is right rather than because a flag suppresses it. Surface: the full page appears **once per ENTRY into the offered state**, three dismissal levers with three scopes, and **none removes the entry point**. Q7's trap is closed — „Helyreállítási kód létrehozása" is **unavailable** while a recovery is outstanding. **WHY THE ROW STAYS FAIL: these are fixes, not a walk.** Nothing here walked a customer end to end, and this row goes green only when one completes with **no operator intervention AND a byte-identical sentinel**. Two of Phase 1's blockers are also still open (**R-214**, **R-202**), and **R-240** is untouched. Evidence: `felhom-controller/CHANGELOG.md` v0.206.0, `hub/CHANGELOG.md` v0.98.0. **DELIVERY, separately: R-239 is CLOSED 2026-08-07** — golden **0.205.0** baked, published, round-trip verified (`./etc/felhom-controller-image` read OUT of the downloaded archive) and **VOUCHED**, so a fresh install now lands on 0.205.0 carrying R-234 and R-237. In the event it was a ONE-field change: `agent_version` and `min_agent` both stayed 0.127.0. Evidence: `tests/golden-0.205.0-2026-08-07/`. **The row still says FAIL** — delivery is not a journey, and R-241 is diagnosed, not fixed. | +| **The customer's own UNAIDED recovery journey, end to end** | controller v0.206.0, hub v0.98.0, agent v0.127.0 | **PROVEN-LIVE (2026-08-07, the fifth walk) — an unaided customer journey on a real installation, SCOPED: see what it does not claim** | A throwaway appliance was built from the published ISO on demo-hp, claimed, given three sentinels, escrowed, backed up off-site, then destroyed — guest purged and both enrolled drives wiped. The rebuild and the customer's route were then walked with **no command line inside the guest**. **Data: PASS.** All three sentinels byte-identical (`beb9175d…`, `7c8cb0ad…`, `e012e76f…`), including a 12 MB binary and `C11-őrszem-ékezetes-árvíztűrő.txt` — the Hungarian filename survived disk → restic → SFTP → Storage Box → restore → disk intact. Restored in **16 s** out of the pre-wipe snapshot `f3d9cd67`, non-destructively. **Journey: FAIL.** Four interventions, three of them root on the appliance: R-216 (a correct code called wrong), R-218 (the store never opened; the box stopped asking), R-219 (the promised listing structurally could not render), R-220 (the app could not be redeployed — its drives were unenrollable). Evidence: `tests/campaign11-evidence-2026-08-05/journal.md` | **WHAT IT DOES NOT CLAIM, and what the v0.201.0/v0.97.1 fixes do and do not change.** Six of the nine findings are fixed and each is red-proofed (R-216, R-218, R-219, R-217, R-222, R-215) — but **fixes are not a journey.** This row goes green only when the walk is repeated end to end and completes with no operator intervention. **Three findings are deliberately still open and each is a live blocker for some flow:** **R-214** (the console never stops showing a stale pairing code), **R-220** (after a rebuild the customer's drives cannot be re-enrolled — the deploy refuses and the wizard's list is empty, so no app can be redeployed onto its own data drive), **R-221** (a rebuilt box cannot run the escrow ceremony at all, because the pbsdr convergence marker survives on the host while the seeded `escrow.pbs_storage_id` is rewritten out of `agent.json`). **And R-216's fix does not by itself make a NEW box work:** with the hub guard corrected but the Day-0 manifest still vouching agent 0.120.0, a new box is **HELD, not served** — it stops being lied to, which is better, but the feature does not work for it until agent 0.125.0 is vouched. **⚠ CAMPAIGN-11 PHASES 2 + 4 RAN OVERNIGHT 2026-08-05/06 AND THIS ROW STAYS **FAIL**.** **What they add:** eleven injected faults against the recovery journey, each with a positive control, and a full unattended scheduled cycle. **The backup promise strengthened** — the off-site tier ran itself at 04:15 on a box rebuilt twice and set aside hours earlier (`snapshot_count` 1 → 2), all five daily jobs fired exactly once, and nothing on the must-not list fired. **R-217's and R-215's fixes were proven live under exactly their faults**, the set-aside proved it does not delete (12 535 KB byte-exact at the far end), and a wrong code was refused three times with nothing written and no lockout. **What they do NOT add: any progress on the JOURNEY.** These faults are **not a re-walk** — no customer route was walked end to end — and four new findings say the journey got no better where it matters: **R-224** (a hub outage and a stopped agent are BOTH reported to the customer as a bad recovery code, in 0.056 s and 0.030 s — no unseal attempted; the agent's own `err` distinguishes them and it is discarded at the HTTP boundary, and R-216's gate answers `source=version` so it cannot see reachability — **Phase 1's headline defect relocated from the version channel to the transport**), **R-226** (M1, the only message telling a customer to check their typing, is unreachable on any box that has re-escrowed), **R-225** (an unread store renders `0 pillanatkép · 0 GB` above a card saying it holds backups — 1 snapshot and 12 535 KB were really there), **R-228** (the set-aside history is recorded in `orphaned_renamed_to` and shown by nobody — `OrphanedRenamedTo` has zero references in any template or handler). **§4.1 is now MEASURED rather than deduced** — the floor IS served, from the box's own rendered `GetFloor()` and a cold-started controller's settle-gate line. **§4.2's positive half is still NOT measured** and needs a rebuild. Evidence: `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md`, `tests/campaign11-evidence-2026-08-05/journal-phase24.md` **2026-08-06 — THE FIVE PHASE-2 FINDINGS ARE FIXED AND THIS ROW STILL SAYS FAIL.** controller **v0.202.0** + agent **v0.126.0** close R-224, R-226, R-225, R-227, R-228. **What changed:** the unlock path now classifies WHY it failed, from the value and never the text, and the message that mentions typing is reachable from exactly one class — a real refusal — with **unknown defaulting to neutral**. Proven live on the venue with the same wrong code and only the hub's reachability changed: `400 → 502 → 400`. **What it does NOT change: this row.** Fixes are not a journey; nothing here walked a customer end to end, and three of the blockers that made Phase 1 fail are still open (**R-220** in particular is worked around BY HAND on the venue — without that unmount no app can be deployed on a rebuilt box). **What still needs proving:** a re-walk with no operator intervention, and specifically the customer-facing messages end-to-end — they were NOT re-driven, because `/recovery` correctly retires itself once the old data is set aside and restoring that state would be the reconfiguration the fix task forbade. Evidence: `felhom-controller/REPORT.md` **⚠ THE RE-WALK RAN 2026-08-06 (R-201, attended) AND THIS ROW STAYS FAIL — but it is a nearer no.** A brand-new appliance was built from the published ISO, given three sentinels, escrowed, backed up off-site, destroyed (guest purged AND both drives wiped), reinstalled and walked. **THE DATA: PASS** — all three sentinels byte-identical, including a 12 MB binary and an accented Hungarian filename **whose NAME BYTES are also identical**, restored in **23 s** out of the pre-destruction snapshot through the customer's own flow. **THE JOURNEY: FAIL — TWO dead ends against Phase 1's four.** (1) **R-218's CONSUME half**: the hub re-staged the credential and said "the box re-consumes on its next cycle"; a full cycle ran (positive control) and it did not — no customer-reachable action fetches it, and the only lever is a command line INSIDE THE GUEST. (2) **R-220**: the drives are still unenrollable after a rebuild, needing a Proxmox-host unmount, without which no app can be redeployed and the restore page stays empty. **The unaided RTO is therefore STILL UNDEFINED**; the attended figure was 30 m 13 s and must not be quoted as the customer number. **What passed and is new:** the recovery screen appeared WITHOUT being sought, answered all three of its questions with a seal date matching the hub exactly, the emailed reset code worked first try, and the unlock was a real 1.528 s unseal that placed the key. **⚠ AND IT DOES NOT CLAIM DELIVERY:** the fresh install landed on controller 0.201.0 / agent 0.125.0 — the vouched versions, neither carrying the fixes — which were installed BY HAND. Fleet delivery needs a golden carrying the fixed controller and a vouched agent, and nothing was vouched. Evidence: `tests/rewalk-r201-2026-08-06/journal.md`   **⚠ 2026-08-06 (later the same day) — THE TWO DEAD ENDS ARE CLOSED AND VOUCHED, AND THIS ROW STILL SAYS FAIL.** controller **v0.203.0** + agent **v0.127.0** close **R-218's consume half** and **R-220**, and golden **0.203.0** / agent **0.127.0** / `min_agent` **0.127.0** are now VOUCHED — so the delivery gap the re-walk recorded (“the fresh install landed on versions neither carrying the fixes”) is gone. **Both were then proven live on a genuinely rebuilt box** (Part 4 venue, VM 323): the reinstall landed on agent 0.127.0 with **no downgrade and no hand upgrade** (the previous re-walk's reinstall downgraded 0.126.0→0.125.0), `/disks/candidates` returned **both drives** after the guest was purged with the raw mounts still on the surviving host — before the fix, two empty lists — and both **re-attached through the customer endpoint**; the off-site credential was re-staged by `offsiteheal` at 13:24:57Z and **collected by the box on a tick, unaided**. **Why the row is still FAIL:** the walk did not finish. It stopped at the restore surface — **R-237**, the list was keyed on apps that are currently INSTALLED and currently TOGGLED ON for future backups, so a rebuilt box was shown nothing to restore while the repository held its snapshots. Fixed in **controller v0.204.0** (the store is now the source of the list) and proven live on that same venue with the toggle switched OFF, but **no sentinel was restored in that walk**, so the data half is unproven in EITHER direction for this venue. **R-238 was RECLASSIFIED, not fixed as filed:** `mode=full` without `confirm=1` is step 1 of a deliberate two-step that starts no job by design; the endpoint driver did not carry `?full_prep=` forward. The operator drove the same restore to completion in a browser. Its real residue — that step logging nothing at all, including when the headroom gate refuses — is fixed. **R-236 was WITHDRAWN: not a defect.** **This row goes green only when a walk completes end to end with no operator intervention, restoring a byte-identical sentinel.** Evidence: `tests/part4-rewalk-2026-08-06/journal.md`, `tests/teardown-2026-08-06.md`, `felhom-controller/REPORT.md` **⚠ THE FINAL WALK RAN OVERNIGHT 2026-08-06/07 AND THIS ROW STAYS FAIL — the data half passed again, the journey half failed further from the line than before.** A brand-new appliance was built from the published ISO, fixtured, soaked through a real scheduled cycle, destroyed and rebuilt. **THE DATA: PASS** — all three sentinels byte-identical out of snapshot `f5c53b03`, including a 12 MB binary and an accented Hungarian filename **whose name BYTES are identical too**. **THE JOURNEY: FAIL, and this time the customer has NO route at all** — `/` lands on the launcher with no recovery pointer, `/recovery` **302s away**, and the remote-backup page offers to **CREATE a new recovery code**, which would orphan the very history the customer's code protects. **There is no field anywhere to enter the code they hold**, and the operator's documented remedy refuses too. Cause (**R-241**): the credential self-heal — proven working unaided in the same walk, six hours earlier, as its best result — writes a FRESH repository password, which moves the box out of `OffsiteRecoveryOffer()`'s (a) pristine case while (b) orphan-detected is unreachable because runs are blocked by `escrow_state: pending`. **R-218's shape one level up.** **What DID pass, and is new:** the whole credential chain ran end to end with zero human action on an unclaimed box — declare, `offsiteheal` re-stages after two reports, the box's 5-minute retry collects it, tier applied — the **first live sighting** of that success line, settling R-218's consume half and R-236's withdrawal. Also new: **R-239** — a fresh install lands on controller **0.203.0** while 0.205.0 is released, so R-234 and R-237 are written but **not delivered**, proven from the customer's side (T2, T3). **This row goes green only when a walk completes with no operator intervention AND a byte-identical sentinel.** Evidence: `tests/finalwalk-r201-2026-08-07/journal.md` **⚠ 2026-08-07 — R-241 WAS DIAGNOSED BY A READ-ONLY SPIKE ON THE STANDING VENUE, AND THE DIAGNOSIS REVERSES THE FIX. This row is unchanged: still FAIL, and no code was written.** The question was whether R-241 is a screen-predicate defect or a minting defect. **It is a MINTING defect.** The screen was telling the truth — there genuinely was nothing recoverable under the key the box held, because the box minted that key itself over the top of a sealed package it already knew the hub was holding. **Three measurements, from the venue rather than from the earlier report:** (1) `WriteOffboxSecrets` (`offbox.go:411`) mints on ONE input — does the file exist — while `OffsiteRecoveryOffer()` and `needsOffsiteCredential()`, both in the same file, consult `GetHubEscrowIdentityPresent()`; the same fact is available on three paths and used on two. (2) That flag was the **precondition of the chain that reached the minting**: the retry job logs only when the declaration is live, and the venue logged `credential retry: … (the box still declares a need; retrying)` at **02:48:03Z** — thirty minutes and six ticks before the mint at 03:18:06Z. (3) **The box computed the answer and discarded it**: at **03:28:03Z**, thirty-five minutes before the customer looked, `EscrowAutoConfirmer.Reconcile` (`escrow_confirm.go:154`) logged `the hub's escrow blob does not cover the CURRENT repo password (hub hash 30ef574fe492… != local 9b4a9a9dcec7…)`. It is recomputed every report cycle and never persisted. **And the hub explicitly disclaims doing this** — `offsiteheal`'s package doc: *"credential automatic, key customer-present — the ruling this session implements and must not quietly widen."* **Shape (b) is structurally unreachable on this box** (the escrow gate at `offbox.go:743` sits upstream of `ensureOffboxRepo`, the only producer of `RepoState="orphaned"`; positive control: the scheduler was alive, 241 `agent-channel-health` ticks, and `offbox-backup` is a `sched.Daily` leg whose slot fell before the destruction). **Two by-products:** **R-243** — a box in this state silently stops backing up off-site and **no alarm fires** (`isStale` requires `escrowed`, `offsite_delivery_stuck` skips the `applied` shape, `backup_failed` needs a run that never happens); and the trap in the obvious fix — `ResetOrphanedRepo` clears the orphan flag without a ceremony, so a hash-mismatch discriminator alone would re-offer the screen forever to a customer who declined the old data. **Q7:** the „Helyreállítási kód létrehozása" button does NOT destroy the data — R-198's retention copies the `identity_blob` — but it converts a self-service recovery into one needing an unbuilt read path (R-199), and it re-enables the screen while invalidating the code that screen accepts. Evidence: `audits/SPIKE-r241-recovery-offer-2026-08-07.md` **⚠ 2026-08-07 (later) — R-241 IS FIXED (controller v0.206.0 + hub v0.98.0) AND THIS ROW STILL SAYS FAIL.** Three changes, following the spike's ruling rather than the obvious reading: the box **no longer mints a repository key while the hub holds a sealed package** (a conjunction, so a first-time box is untouched; the refusal is a HOLDING state that still writes the transport, declared as `offsite.state=awaiting_recovery_key` and shown inert to every existing hub reader); the **hub-vs-local key comparison that was computed every cycle and discarded is now persisted** and drives the offer as **shape (c)**; and **abandoning starts a 14-day countdown** whose terminal step removes the set-aside store and the sealed package **together**, so the offer ends because the state is right rather than because a flag suppresses it. Surface: the full page appears **once per ENTRY into the offered state**, three dismissal levers with three scopes, and **none removes the entry point**. Q7's trap is closed — „Helyreállítási kód létrehozása" is **unavailable** while a recovery is outstanding. **WHY THE ROW STAYS FAIL: these are fixes, not a walk.** Nothing here walked a customer end to end, and this row goes green only when one completes with **no operator intervention AND a byte-identical sentinel**. Two of Phase 1's blockers are also still open (**R-214**, **R-202**), and **R-240** is untouched. Evidence: `felhom-controller/CHANGELOG.md` v0.206.0, `hub/CHANGELOG.md` v0.98.0. **DELIVERY, separately: R-239 is CLOSED 2026-08-07** — golden **0.205.0** baked, published, round-trip verified (`./etc/felhom-controller-image` read OUT of the downloaded archive) and **VOUCHED**, so a fresh install now lands on 0.205.0 carrying R-234 and R-237. In the event it was a ONE-field change: `agent_version` and `min_agent` both stayed 0.127.0. Evidence: `tests/golden-0.205.0-2026-08-07/`. **The row still says FAIL** — delivery is not a journey, and R-241 is diagnosed, not fixed. **✅ 2026-08-07 — THE FIFTH WALK PASSED, BOTH HALVES, AND THIS ROW TURNS.** A brand-new appliance (VM **325**, customer `walk5`) was installed from the published ISO on `demo-hp`, claimed, given three sentinels, escrowed, backed up off-site, then **destroyed on purpose** — guest purged, both data drives wiped — and rebuilt through the documented day-0 path. **THE DATA: PASS** — all three sentinels byte-identical out of snapshot `5b0f20f7` (`11eb7fb2…`, `6b504d1e…`, `0baaf402…`), including a 12 MB binary and `WALK5-őrszem-ékezetes-árvíztűrő.txt` **whose name BYTES are identical too**, read back with `os.listdir` on a bytes path so no decode round trip could launder a `U+FFFD`. **THE JOURNEY: PASS — ZERO guest command lines were needed to progress**, against three on the previous walk. **RTO 71.7 s** from login to an open store (12.44 s of it the unseal; ~22 s a harness retry). The recovery screen **appeared without being sought** (`/` → `/launcher` → `/recovery`), answered all three of its questions, and its sealed-at timestamp matched `host_escrow.created_at` exactly. **WHAT MADE THE DIFFERENCE — R-241's mint guard, exercised live for the first time:** at **14:58:52Z**, unaided and before anyone logged in, the rebuilt box collected its re-staged credential, configured the transport and **refused to mint a repository password** over the sealed package — `NOT minting a repository password: the hub holds a sealed recovery package for this box`. Sampled every 20 s from T0: **no key at any moment**, with the scheduler proven alive throughout. At the equivalent moment the previous walk minted one and lost the journey silently. **AND DELIVERY IS PART OF THE PASS:** the fresh install landed on **controller 0.206.0 + agent 0.127.0 — the vouched set, no hand upgrade**, both times, so the box under test is the box a customer receives. **WHAT THIS ROW STILL DOES NOT CLAIM.** (1) **Putting files back in place is built and worked here** — `reconstitute` placed 6 files — but **only after two obstacles the customer must guess past**: **R-252** (the restore refuses with „nincs elérhető adatmeghajtó" because a rebuild loses the drive *registration*, and nothing on the recovery path says to re-attach) and **R-253** (the restore refuses because the app is not installed, on a page that says three lines above that the restore reinstalls it). Both were cleared **from the dashboard with no shell** — which is why the journey passes — but neither is signposted, so *unaided* here means *possible without a shell*, not *obvious*. (2) **Shape (c) did NOT fire positively.** With the mint guard holding there is no local key, so the offer comes from **shape (a)**; shape (c) was measured in Phase A in its **negative** half (hub hash == local hash, correctly silent). The mint guard is proven positively; the discriminator only negatively. (3) **R-214, R-202 and R-240 are untouched.** (4) Teardown of this venue is **owed**. Evidence: `tests/walk5-r201-2026-08-07/journal.md` | | **Off-site app-data capture covers the paths an app declares MANDATORY — on BOTH drive layouts** | controller v0.197.0 | **PROVEN-LIVE (2026-08-04) — and this row was OPTIMISTIC before it** | On demo-hp, an app on the **system-data fallback** bound `/mnt/sys_drive/userdata/media/books` while the capture set looked in `/mnt/sys_drive/felhom-data/userdata/media/books`: the declared-mandatory directory was in **no** snapshot and the run reported `ok` (R-203). After v0.197.0 the capture log reads `1 mandatory path(s)` and the file is **listed inside the snapshot** — `restic ls -l latest --tag calibre-web` → `-rw-r--r-- 1000 1000 181 … /DRILL-SENTINEL.txt`. Evidence: `audits/DRILL-r201-offsite-recovery-2026-08-04.md` §2 and the v0.197.0 CHANGELOG | **WHAT IT DOES NOT CLAIM.** It covers CAPTURE, not RESTORE: **no file has ever been restored from an off-site backup after a wipe** (R-201, ready to resume). It also does not claim the enrolled-drive layout was ever wrong — it was not, which is precisely why this survived. And a run that still misses a mandatory directory now reports **`incomplete`** rather than `ok`, so this row's guarantee is one the status can express. **Widened 2026-08-06 (controller v0.205.0, R-234):** the same verdict now also covers an app skipped ENTIRELY — until then a missing declared FOLDER made the run incomplete while an app with no recovery unit at all still reported `ok`, so the smaller gap moved the verdict and the bigger one did not. **This row still claims CAPTURE, not that a newly-selected app is protected by the next run** — for a DEPLOYED app it is (the run's own pre-dump phase writes the unit, measured 2026-08-06), and for an undeployed one it is not and the card now says so || **Recurring offsite (PBS) whole-guest backups actually LAND, and RESTORE** — local daily + offsite weekly as scheduled work | agent v0.97–0.103, controller v0.174/0.175, hub v0.76.0, host-install 1.20.0 | **PROVEN-LIVE (2026-07-26)** | `audits/SPIKE-r82-phase0-2026-07-26.md`; per-repo CHANGELOGs/REPORTs. **Restore round-trip on demo-hp:** `--selftest=restore-test` against `felhom-pbs:backup/ct/9201/2026-07-26T15:42:42Z` → `pass:true`, `verified:"boot+running"`, **`mount_parity:"ok"`** (`mp0=/var/lib/docker 50G`, `mp1=/mnt/sys_drive 20G`, mp8/mp9 throwaway stand-ins for the archived binds), `source_tier:"pbs"`, 4m5s restore+boot+verify+teardown, scratch band clean afterwards and the live guest untouched. | **This row is distinct from the DR-tier row above, which proves ACTIVATION, not ARRIVAL.** That tier was PROVEN-LIVE as `applied` since 2026-07-21 while demo-felhom held ONE snapshot (a healing artifact) and demo-hp held **zero, ever** — "applied and empty", the R-39 shape one level quieter. **What earns PROVEN-LIVE here:** (a) demo-hp's FIRST EVER offsite backup landed (4.25 GB into a verifiably empty namespace); (b) it **restores into a bootable, mount-complete guest** — `mount_parity` is the non-hollow half, since a boot-only verify cannot see a missing data volume; (c) **the multi-tier quiesce ran through the real UI endpoint** (`POST /api/guest-backup/trigger`, authed+CSRF) and produced **exactly ONE stop/start pair with BOTH backups inside it** — `quiescing 1 stack(s)` 17:01:39 → local done 17:02:56 *"next tier may start (app still quiesced)"* → felhom-pbs snapshotted 17:03:06 → `unquiescing` 17:03:06. **App downtime 1m27s for both tiers**, and the app came back healthy. **Known gaps, recorded not hidden:** the SCHEDULED restore-test still only selects the PRIMARY tier, so the offsite tier is never AUTOMATICALLY restore-tested (the manual/selftest path is proven, the unattended one is not); the hub infers "PBS ⇒ weekly" from storage TYPE rather than a reported cadence; the installer-default fleet flip awaits a full weekly cycle. → **R-82** | | **Restore-proof is UNATTENDED — the scheduler covers EVERY tier, follows the BACKUP rather than the clock, and a failure is heard** | agent v0.104.0 → **v0.121.0**, hub v0.77.0 → **v0.91.0** | **PROVEN-LIVE (2026-08-03)** | per-repo CHANGELOGs; `backlog/SPEC-r85-phase4-5-2026-07-26.md`. Unit red-proofs for tier rotation, restart-survival, the one-heavy-op gate, failure-emits-an-event, and newborn silence. | **The row above is earned by a MANUAL `--selftest=restore-test`; this one is about the SCHEDULED path, and the distinction is the whole point.** Before R-85 the scheduler could only ever see `cfg.Backup.BackupTarget()`, so the offsite tier was never a candidate — and a failed restore-test was a `[WARN]` line with no event at all, which was true for the LOCAL tier that WAS being tested. Now: oldest-first rotation across every configured tier (operator ruling 2026-07-26), persisted so it survives a restart; a restore-test joins the host-wide one-heavy-operation gate so it never contends with a backup over the same tunnel; and the hub raises two DISTINCT operator-tier signals — `restore_test_failed` (broken now) and `restore_test_stale` (unverified, not known-broken), anchored on R-81 so a newborn box never alarms. **Why this was NOT PROVEN-LIVE until now:** rotation had not been observed selecting both tiers across consecutive UNATTENDED cadences — at a 24h cadence a multi-day window — and a single passing run proves the code path, not the schedule. → **R-85**. **R-86 (agent v0.121.0 + hub v0.91.0, 2026-08-03) replaced the schedule and the observation became possible in one afternoon**, because what has to be observed is no longer a multi-day rotation but a RULE: a tier is due when its newest archive that has settled ~24 h has not been proven. **UPGRADE EVIDENCE — a real unattended run on demo-felhom (2026-08-03), triggered by DUE-NESS, not by a timer.** The scheduler's own log: `15:14:38 restore-test tier is DUE … target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z … reason="newest settled archive … has not been proven"` → `proxmox-backup-client restore --crypt-mode=encrypt` under the agent's own token → `15:25:08 gate decision class=guest_destroy guest=990000 allowed=true` → `15:25:14 scratch guest torn down` → **`15:25:14 backup: scheduled restore-test passed archive=felhom-pbs:… duration_s=635.1`**. A **14.5 GB encrypted offsite archive pulled from ep0 over the WAN**, restored, booted, verified and destroyed in 635 s, unattended. **The three things a timer could not show**, all verified after it: the state names THAT archive (`{"felhom-pbs":{"archive":"…2026-07-28T04:49:43Z","proven_at":"2026-08-03T13:25:14Z"}}`); a second evaluation reports `due=false … is already proven` and runs nothing; and **an agent restart runs nothing**, which is the defect a person actually noticed — every deploy used to restart the timer. Teardown verified at all three layers: guest absent from `pct list`, **zero** `990000` volumes in `lvs`, and the hub-side `restore_tests[]` entry deliberately RETAINED (it IS the proof the staleness check reads). **R-189 (agent v0.122.0, 2026-08-03) closed the reporting half of this row, and it was a REAL gap in the evidence path:** the proof above reached the hub only because no restart intervened — `restore_tests[]` came solely from an in-memory store, so the 15:25:14 PASS was in fact LOST when the agent restarted 2 m 43 s later for a deploy (`0 restore-tests` on the next two host-reports). Under per-archive due-ness the box would not have repeated the work for a week. The persisted per-tier proof (with the archive, and now the tier) is merged into the report, so a proof survives a restart — one entry per tier, newest wins, and a record that cannot be described honestly is not emitted. **THE HOST TIER'S HALF OF THIS ROW WAS OPTIMISTIC UNTIL 2026-08-03, and it is worth saying plainly.** Every live restore-test cited here is on the OFFSITE tier. The HOST tier was not merely unproven on demo-felhom — it was **unprovable**, and on demo-hp equally: the agent's token had no ACL on `/storage/felhom-backup`, the storage both boxes configure as `local_backup_target`, so the content API answered `{"data":[]}` through the token while root listed three archives. The scheduler skipped it as *"no settled archive yet"* — which is exactly what a brand-new tier reports — so nothing ever said so (**R-185**). **CLOSED 2026-08-03 — agent v0.123.0 + installer 1.24.0.** The grant is applied on both demo boxes (the token now lists 3 and 4 archives), the installer's reuse arm grants on a pre-existing target, and the agent now **asks whether it may read each tier it depends on** instead of inferring it from an empty list: a critical degraded capability naming the storage and the missing role, which the hub alerted and emailed on **while the box was still blind**. **THE HOST TIER IS NOW PROVEN-LIVE, UNATTENDED, ON BOTH DEMO BOXES (2026-08-04) — this row's optimistic half is cashed.** Four SCHEDULED runs, nothing triggered by hand: **demo-felhom** host tier `…2026_08_02-04_42_14.tar.zst` **passed in 83.8 s** at 00:55, offsite `…2026-07-28T04:49:43Z` **passed in 540.4 s** at 06:55; **demo-hp** host tier `…2026_08_02-04_49_29.tar.zst` **passed in 109.3 s** at 02:05, offsite `…2026-07-28T19:19:45Z` **passed in 300.1 s** at 08:05. Each restored into a scratch guest, booted, verified and destroyed itself; `pct list` and `lvs` show zero `990000` afterwards on both, and both boxes' `local-lvm` returned to their pre-run figures (1.95 % and 30.83 %). **Both boxes had BOTH tiers due at once and the ordering was observed live for the first time under R-86's rule:** never-proven sorted first, so each box took its HOST tier, deferred the offsite one, and picked it up on the following evaluation six hours later — one heavy operation at a time, per box, without anyone sequencing it. **The proofs reached the hub**, which is R-189 carrying a host-tier entry for the first time: demo-felhom's report holds **two** entries, one per tier, and the `local` one can only have come from the persisted state because the in-memory store held only that morning's offsite run. **SCOPE, stated because one box proving something does not make it a fleet property:** this covers `demo-felhom` and `demo-hp`. The tester's box is untested and untouched. Evidence: `felhom-agent/REPORT.md`, `felhom.eu/REPORT.md` | | Customer RESET (middle lifecycle tier: host delete < RESET < customer Delete): one operator action → pre-first-install; all operational state destroyed, identity + basic config survive | hub v0.61.0, felhom-tenantsync v1.1.0 | **PROVEN-LIVE (external teardown, incl. two real firings)** | **`tests/VALIDATION-n100-rehearsal-2026-07-18.md` — two live firings, both host-delete-first, on two different customers** (`demo-vm-felhom` 15:49:57, `demo-felhom` 16:08:51): every leg `ok` (`claim`, `db_purge`, `descriptor`, `hetzner`, `pbs`), escrow acked separately, each completing in 8–9 s (`hub-state.txt` `customer_resets`). The **Hetzner sub-account destruction is now verified against the live pool box** — and produced the run's sharpest lesson: **a sub-account is an access-control object, not a data object.** Deleting it left its `/home` intact, so re-enabling offsite recreated an account over the previous lifecycle's ciphertext under a key this same RESET had destroyed — which is why the orphan guard fired at 16:58:14 (**a finding by S7's own criterion**) and why RESET now needs a base-dir purge → **R-32**. Prior: hub v0.61.0 REPORT; **ep0 live drill 2026-07-17** (throwaway `drill-reset-01` with a real backup: deprovision `deleted:true` destroyed the namespace + backup group + token, idempotent re-run `deleted:false`, all 3 real tenants + shared user survived); red-proofs (ack-gate, partial-failure resumability) + orchestration/store/offsite/render tests | External teardown FIRST, DB purge LAST, every leg idempotent; refuses while any host row exists; separate escrow-custody ack; clears claim (fresh code next onboarding); keeps the offsite tier CHOICE, drops provisioned fields. **Live-clicked 2026-07-18** (twice, by Viktor) — this supersedes the earlier "not live-clicked / Hetzner delete unit-tested only" note. **R-25b CLOSED (hub v0.69.0, 2026-07-21):** the Danger-zone DELETE is now the guided full-teardown cascade that runs this very sequence as its middle leg — see the row below | diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 5a2c8b8..312f143 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -170,7 +170,7 @@ the fault was real. Full observables: `tests/campaign11-evidence-2026-08-05/jour | **R-241** | **The credential self-heal, succeeding, locks the customer out of their own recovery.** Measured end to end on the final walk (2026-08-07, `tests/finalwalk-r201-2026-08-07/journal.md`). `OffsiteRecoveryOffer()` shows the recovery screen on exactly two conditions: **(a)** the box has **no** repository password — the pristine rebuilt shape — or **(b)** it has one but the inherited history will not open under it (`OffboxOrphaned()`). Overnight, unaided and exactly as designed, `offsiteheal` re-staged the one-time credential and the box's 5-minute retry **collected it and applied the tier**, writing a **fresh repository password** at 03:18Z. That makes **(a) false**. **(b)** is false too, because orphan detection only fires when a run actually tries the repository — and runs are blocked by `escrow_state: pending`. **The box therefore sits in the gap between the two conditions, and the gap is self-locking:** it cannot detect the orphan without running, cannot run without escrow, and cannot escrow without minting a NEW recovery code — which would orphan the history the customer's existing code protects. **What the customer sees:** `/` is „Indítópult" with no recovery pointer; `/recovery` **302s away**; `/backups/remote` offers „Helyreállítási kód **létrehozása**". **There is no field anywhere to enter the code they hold.** **And the operator's documented remedy also refuses** — `--recover-offsite-install` returns *„[REFUSED] a DIFFERENT repository password is already present… which history to keep is not a decision this command may take. Nothing written."*, which is correct and fail-closed and still a dead end. Recovery required moving the fresh key aside by hand and re-running the install: **three guest command lines**. **The two keys, measured:** on-disk `9b4a9a9d…` (self-heal) vs recovered-from-R `30ef574f…`. **THE DATA WAS NEVER AT RISK** — all three sentinels restored byte-identical once the right key was in place. **This is R-218's shape one level up:** that finding read *"succeeding at recovery stopped the box asking for what it still needed"*; here, succeeding at the credential self-heal stopped the box **offering** the recovery it still needed. The same walk proved the self-heal working unaided six hours earlier, and that success is what causes this. **Likely shape of the fix, not yet a decision:** the offer needs a third condition — a box holding a password it has never successfully used, while the hub holds a sealed package, is a recovery candidate — or the self-heal must not install a credential on a box whose escrow is still `pending` and whose hub blob is unconsumed. **Which of those is right is a design decision, deliberately not taken here.** **⚠ RULED 2026-08-07 by a read-only spike on the standing venue — `audits/SPIKE-r241-recovery-offer-2026-08-07.md`. IT IS A MINTING DEFECT, NOT A SCREEN-PREDICATE DEFECT, and that reverses the fix.** The screen was telling the truth: there genuinely was nothing recoverable under the key the box held, because the box minted that key itself over the top of a sealed package it already knew the hub was holding. **Three measurements carry it.** (1) **`WriteOffboxSecrets` (`offbox.go:411`) mints on ONE input — does the file exist.** It never reads `GetHubEscrowIdentityPresent()`, while its two neighbours in the same file, `OffsiteRecoveryOffer()` (`:1412`) and `needsOffsiteCredential()` (`:1377`), both do. **The same fact is available on three paths and used on two.** (2) **The flag was not merely available — it was the precondition of the chain that reached the minting.** The 5-minute retry job only logs when `RetryIfDeclared` fires, which requires the declaration, which requires that flag; the venue logged `credential retry: … (the box still declares a need; retrying)` at **02:48:03Z** and five times after — **thirty minutes and six ticks before the mint at 03:18:06Z**. (3) **The box KNEW and threw it away:** at **03:28:03Z**, thirty-five minutes before the customer looked, `EscrowAutoConfirmer.Reconcile` (`escrow_confirm.go:154`) computed the exact discriminator and logged `[WARN] [escrow-confirm] the hub's escrow blob does not cover the CURRENT repo password (hub hash 30ef574fe492… != local 9b4a9a9dcec7…) … staying pending`. **It is computed on every report cycle, never persisted, never surfaced.** **And the hub explicitly disclaims doing this** — `offsiteheal`'s package doc: *"it never runs, or asks for, an escrow ceremony… **credential automatic, key customer-present** — the ruling this session implements and must not quietly widen."* The repository key is minted in the seam between two sides that each honoured their contract, by a helper doing exactly what its doc comment says. **Fixing the predicate would paper over a box quietly making its own history unopenable.** **Recommended fix (not started, no code written): persist the discriminator the ACK already carries and add it as shape (c)** — needs no hub change and no new protocol field — **plus a `decided` latch**, because `ResetOrphanedRepo` clears `RepoState` without running a ceremony, so an `H`-mismatch discriminator alone would re-offer the screen forever to a customer who explicitly declined the old data. **Four operator decisions are stated and left unanswered in §"THE OPERATOR'S DECISION".** **✅ FIXED 2026-08-07 — controller v0.206.0 + hub v0.98.0, and the ruling above is what the fix follows.** **(1) It stops minting:** the guard is a CONJUNCTION (a package held AND no key present), so a first-time box mints exactly as before; the refusal is a HOLDING state, not a failure — the transport is still written so the recovery screen can bring the tier up the instant the key arrives (R-219), and returning an error instead would have left the hub re-staging a consumed credential for ever. New declared state `offsite.state=awaiting_recovery_key`, shown INERT to every existing hub reader from their code rather than assumed. **(2) The discriminator is persisted and drives the offer as shape (c).** §7.2 resolved deliberately: **a known difference offers however old the reading** (age is NOT gated on — gating would make a box offline from the hub silently stop offering, the very failure this removes), and **a hash never learned falls back to (a)/(b)**, because an empty hash is the hub positively saying its package seals no key rather than an unknown. **(3) Abandoning ends the question** — a 14-day countdown, visible and reversible, whose terminal step removes the set-aside store AND the sealed package together, after which shape (c) has nothing to compare and the offer falls silent **because the state is right, not because something remembers it once was not**. The two halves cannot be atomic across two machines, so it is a two-phase commit whose confirmation rides the SAME ACK that carries the request. **(4) The surface:** the full page appears **once per ENTRY into the offered state, not once ever** (an epoch — a box rebuilt months later is a new situation); three dismissal levers with three scopes, and **none removes the entry point on the backups page**. **Q7's trap does not survive:** while a recovery is outstanding „Helyreállítási kód létrehozása" is **unavailable**, not merely captioned. **§2.4 honoured:** the abandon confirmation no longer promises *„félretesszük — nem töröljük"* — it states the deletion date. **TWO REAL BUGS WERE CAUGHT BY TESTS RATHER THAN REVIEW, and both are recorded because the shape matters:** `OffboxAwaitingRecoveryKey` omitted `t.Enabled`, so a customer who had switched off-site OFF would have declared a holding state (caught by the EXISTING `TestOffsiteDeclare_DisabledTargetIsNotStranded`); and `recoveryInterrupts` returned early when the offer was false, so the FALLING edge was never recorded and the full page never came back — **the exact defect the epoch exists to fix, reintroduced inside the fix**. **Nine red-proofs, each with the mutation confirmed present in the file before its result was trusted**, including the ships-inert one (unwiring `RecordEscrowKeyHash`, which leaves everything compiling and every test passing while shape (c) reads an empty hash for ever). **NOTHING WAS DELETED ANYWHERE** — the terminal step has only ever run against injected fakes and an injected clock (§7.4). **Still open and NOT built by this:** R-242 (the release-to-golden gate) and R-245 (the automatic 30-day ending). | **FIXED 2026-08-07 — v0.206.0 / hub v0.98.0** | | **R-242** | **A controller release that changes customer-visible behaviour is not delivered until a golden carries it — and nothing enforces that.** R-239 is the symptom; this is the mechanism, recorded **2026-08-07** and deliberately **NOT built** (the task that found it scoped it as a record-only item). **Two releases went out without a golden and the gap was invisible until a walk measured it from the customer's side**: v0.204.0 (R-237) and v0.205.0 (R-234) were written, tested, pushed, and CHANGELOG'd, and every one of those steps passed while a machine installed that night received neither. The register said CLOSED; the fleet said otherwise. **Nothing in the release path knows a golden exists.** The version bump, the image push, the CHANGELOG entry and the register closure are all repo-local; the manifest's `golden_version` is edited by a separate operator act, in a different repo, with no link back. **Proposed shapes, cheapest first — the choice is the operator's and is not taken here.** (a) **A release-path checklist step** — one line in the controller's end-of-session checklist: *a release that changes customer-visible behaviour is not finished until a golden carries it or a register row says why not.* Costs nothing, catches nothing mechanically. (b) **A gate in `repo_gates.py`** comparing the manifest's `golden_version` against the newest released controller and FAILING (or warning) past a tolerance of one minor. Mechanical, runs on every push, and would have fired the morning after v0.204.0. (c) **A hub-side checker** — the hub already knows every box's running controller version from `/hosts` and the vouched golden from the manifest; a periodic comparison against the newest published image would catch drift the repo cannot see, including a vouch that was made and then rolled back. **Earliest catch: (b).** It fires on the push that creates the gap, before any box is installed, and it needs no live fleet. **(c) catches strictly more but only after boxes exist.** (a) is worth doing regardless because it is free. **Not built. No gate was written this session.** **⚠ IT RECURRED WITHIN A DAY, WHICH IS THE ARGUMENT FOR BUILDING IT.** Controller **v0.206.0** shipped the R-241 fixes on 2026-08-07 while the vouched golden still carried **0.205.0** — so a machine installed on the morning of 2026-08-08 would have received neither. Third occurrence of the shape in three days (R-111/R-115/R-120 are the older family). **✅ SHAPE (b) BUILT 2026-08-08 — `scripts/golden_currency_gate.py`, registered in `repo_gates.py` as gate 7.** **It was shown FAILING against that exact state before anything was baked**, which is its red-proof and the reason its own introducing push needed `--no-verify` (stated in the session report rather than worked around): `newest released controller : 0.206.0 / newest golden baked : 0.205.0 → CONVICTED`. **IT IS `--fast`, AND THAT FORCED ITS DESIGN:** both `.githooks/pre-push` AND CI run `repo_gates.py --fast`, so a non-fast gate would run in NEITHER — the R-29 census failure this runner exists to end. **THEREFORE IT CHECKS THE BAKE, NOT THE VOUCH**, because the vouched version lives only in the hub's `hub_settings` with no copy in git, and putting a copy there would create a second source of truth that can drift — a green gate over a false claim being the worst outcome available. **A bake without a vouch still passes: that half is NOT closed and stays on this row.** It also compares versions rather than behaviour, so a release changing nothing customer-visible trips it too — accepted deliberately, because judging that by hand is what failed three times and the cost of a false trip is one bake; a waiver belongs here, never in a habit of bypassing. **VOUCHED 2026-08-08 with the operator's approval** — golden `0.206.0` / sha `c85230b4…108e`; `agent_version` and `min_agent` both stayed `0.127.0`, and `wrapper_sha256` was carried through explicitly because the handler clears it when omitted. **The gate was CONVICTED before the bake and OK after it** — red→green on the same command, which is its proof that it measures something real. | **PARTLY BUILT 2026-08-08 — the bake half is gated; the VOUCH half is not** | -| **R-243** | **A box in the R-241 state silently stops backing up off-site, and NO ALARM OF ANY KIND FIRES.** Found by the R-241 spike (2026-08-07) as a by-product; **not part of the walk's finding and not previously filed.** The R-241 state is self-locking in a second, worse way than the recovery-journey dead end: `escrow_state` is stuck `pending` forever (the auto-confirm flips only on a hash match, and the hash cannot match a key the box minted itself), and `runOffboxBackup` returns at the escrow gate (`offbox.go:743`) before touching anything. **So off-site backups never run again — and the hub never notices.** All three signals that could catch it are excluded, each for its own individually-correct reason, verified in the hub this session: `offsite_stale` — `isStale` (`monitor/offsite.go:135`) returns false unless `EscrowState == "escrowed"`, and its own comment reads *"Pending/disabled = normal onboarding, never stale"*, so the box is classified as **still being set up, forever**; `offsite_delivery_stuck` — `monitor/offsite_delivery.go:91` skips the `applied` shape, and delivery genuinely IS applied (the credential was consumed and the target is in every report); `backup_failed` — never fires, because nothing fails: the run returns `nil` before it starts. **Three correct exclusions leaving one state unobserved.** This is the same class as the workspace `CLAUDE.md` "presence is not success" rule, one level up: **the absence of a failure is being read as the presence of a working tier.** **Partly subsumed by R-241's fix** — a box that recovers leaves this state — but **not for a box that does not**, and the alarm gap is what makes "does not" survivable indefinitely. **Not fixed; no code written.** **⚠ UPDATED 2026-08-07 (v0.206.0) — the STATE this row describes can no longer be entered, but the ALARM GAP is untouched and the row stays open.** R-241's mint guard means a box no longer mints a key over a sealed package, so it no longer arrives in the "escrow stuck pending against a self-minted key" state by itself. **What replaces it is a state that is VISIBLE rather than silent:** the box declares `offsite.state=awaiting_recovery_key` and the customer is offered the recovery screen. **But the hub still raises nothing for it**, and for the same three reasons: `isStale` needs `escrowed`, the delivery checker skips the `applied` shape, and `backup_failed` needs a run that never happens. **So a box whose customer never acts still stops backing up off-site with no operator signal** — the difference is that the customer can now see it and act, where before nobody could. **The remaining work is an operator-side signal for a box held in `awaiting_recovery_key` past some age**, and it is deliberately not bundled into R-241's fix. | **READY** — owner Viktor | +| **R-243** | **A box in the R-241 state silently stops backing up off-site, and NO ALARM OF ANY KIND FIRES.** Found by the R-241 spike (2026-08-07) as a by-product; **not part of the walk's finding and not previously filed.** The R-241 state is self-locking in a second, worse way than the recovery-journey dead end: `escrow_state` is stuck `pending` forever (the auto-confirm flips only on a hash match, and the hash cannot match a key the box minted itself), and `runOffboxBackup` returns at the escrow gate (`offbox.go:743`) before touching anything. **So off-site backups never run again — and the hub never notices.** All three signals that could catch it are excluded, each for its own individually-correct reason, verified in the hub this session: `offsite_stale` — `isStale` (`monitor/offsite.go:135`) returns false unless `EscrowState == "escrowed"`, and its own comment reads *"Pending/disabled = normal onboarding, never stale"*, so the box is classified as **still being set up, forever**; `offsite_delivery_stuck` — `monitor/offsite_delivery.go:91` skips the `applied` shape, and delivery genuinely IS applied (the credential was consumed and the target is in every report); `backup_failed` — never fires, because nothing fails: the run returns `nil` before it starts. **Three correct exclusions leaving one state unobserved.** This is the same class as the workspace `CLAUDE.md` "presence is not success" rule, one level up: **the absence of a failure is being read as the presence of a working tier.** **Partly subsumed by R-241's fix** — a box that recovers leaves this state — but **not for a box that does not**, and the alarm gap is what makes "does not" survivable indefinitely. **Not fixed; no code written.** **⚠ UPDATED 2026-08-07 (v0.206.0) — the STATE this row describes can no longer be entered, but the ALARM GAP is untouched and the row stays open.** R-241's mint guard means a box no longer mints a key over a sealed package, so it no longer arrives in the "escrow stuck pending against a self-minted key" state by itself. **What replaces it is a state that is VISIBLE rather than silent:** the box declares `offsite.state=awaiting_recovery_key` and the customer is offered the recovery screen. **But the hub still raises nothing for it**, and for the same three reasons: `isStale` needs `escrowed`, the delivery checker skips the `applied` shape, and `backup_failed` needs a run that never happens. **So a box whose customer never acts still stops backing up off-site with no operator signal** — the difference is that the customer can now see it and act, where before nobody could. **The remaining work is an operator-side signal for a box held in `awaiting_recovery_key` past some age**, and it is deliberately not bundled into R-241's fix. **⚠ MEASURED ON A REBUILD, 2026-08-07 (fifth walk) — the gap is real for the state this row describes, and NOT for the state a rebuild produces.** 88 seconds after the walk5 guest was destroyed and rebuilt, the hub emitted `offsite_delivery_stuck` (**warning**) and wrote an **operator-channel** `notification_log` row recording `offsite_credential_restaged` / status **REFUSED** with an accurate reason — *"the credential was applied and worked; the target was lost afterwards … a guest rebuild does, R-193"*. So on the **regressed-apply** shape the operator IS told, promptly and correctly, and this row's *"skips the applied shape"* does not apply. The gap stands for a box that reaches the held state **without** a prior working tier in its report history. **Recorded so the row is not read wider than it measures.** | **READY** — owner Viktor | | **R-244** | **The customer DELETE cascade leaves `app_log_issues` behind, and it is systematic across every venue ever torn down.** Found **2026-08-07** while verifying the `finalwalk` teardown with a **full census** (every table, every column) rather than a per-table query. After a cascade that logged `COMPLETE … full teardown`, **61 rows still matched `finalwalk`**. Four of the five sources are **deliberate and correct** — the cascade's own header states *"Provenance/events are NEVER wiped — audit outlives every tier"*: `events` 16, `notification_log` 14, `host_deletions` 1, `customer_resets` 1. **The fifth is a gap:** `app_log_issues` 29 rows, which the residue purge does not touch (its logged leg covers `reports`/`app_telemetry`/`app_log_tails`/`log_tail_requests`/`notif_prefs`/`selfbind_tokens`/`appliance_registrations` — not this table). **It is not a `finalwalk` quirk:** rows still reference **`c11` 40, `rewalk` 20, `part4` 24** — all three torn down 2026-08-06, whose ledger recorded *"0 occurrences"*. **That prior claim was measured with a narrower query and does not survive a full census; the correction is recorded rather than the measurement quietly redone.** **Why it was probably never written, established rather than assumed:** the table is a **fleet-wide aggregate** keyed on `app_name`+`fingerprint` with an `affected_customers` JSON list — of the 29 `finalwalk` rows, **12 reference only `finalwalk`** (orphans, safely deletable) and **17 are shared with LIVE customers** (`demo-felhom`, `peti-felhom`, …) and **must not be deleted, only de-referenced.** A naive `DELETE … WHERE customer LIKE` would destroy a live customer's issue history — which is very likely why the leg does not exist, and is the reason this is not a one-line fix. **Severity is LOW and stated plainly: no secret material is involved** — app name, fingerprint, message text, counts, timestamps. What survives is a deleted customer's *identifier* inside an aggregate row. **Proposed shape:** a residue leg that (a) removes the customer id from `affected_customers`/`context_customer`, and (b) deletes rows whose `affected_customers` becomes empty; plus a one-off sweep for the four already-torn-down venues. **The general lesson is the reusable part:** *a per-table absence query is not a census.* The teardown verification is now a full-schema sweep, and that is what found this. **Not fixed** — a cascade change needs its own red-proof and this session was scoped as a spike plus two operations. Evidence: `tests/teardown-finalwalk-2026-08-07.md`. **⚠ STILL OWED, AND NOW MEASURED RATHER THAN ESTIMATED (2026-08-08 census, read-only, no truncation).** `app_log_issues` holds **1309 rows**; **71 reference a torn-down venue** (`finalwalk`, `c11`, `rewalk`, `part4`); of those **44 are ORPHANS** — they name only torn-down customers and are safely deletable — and **27 are SHARED with a live customer** (`demo-felhom`, `peti-felhom`, …) and **must be de-referenced, never deleted**. 1238 rows are untouched. **The 27 are exactly why the leg was never written**, and why a `DELETE … WHERE customer LIKE` would destroy a live customer's issue history. **What it needs, precisely:** a cascade leg that (a) removes the customer id from `affected_customers` / `context_customer`, and (b) deletes only rows whose `affected_customers` becomes empty; plus a one-off sweep for the four venues already gone. **Why it was NOT done on 2026-08-08:** the fix is hub code, and that session's scope forbade a hub version bump; a hand-run SQL mutation over 71 rows — 27 of them needing surgical de-referencing — with no tested code path and no red-proof is precisely the shape that goes wrong on a live database. **It accumulates one venue at a time, so the next walk adds to it**; the numbers above mean the next session starts from data rather than a guess. | **READY** — owner Viktor | @@ -179,6 +179,11 @@ the fault was real. Full observables: `tests/campaign11-evidence-2026-08-05/jour | **R-246** | **A leftover staleness flag has silently disabled the new recovery discriminator on `demo-hp` since 2026-08-04, and the flag is WRONG.** Found by a read-only spike, 2026-08-08. **Q1 — traced to an act, to the second:** at `2026-08-04 20:15:49` the hub emitted `offsite_reissued` and `escrow_stale` in the same second — an operator **Re-issue** pressed during the R-201 drill, three minutes after `escrow_blob_served` at 20:12:40/20:12:54. That was `offsite.ReissueCredentials`'s **precautionary** `MarkEscrowStale` call, which **hub v0.95.0 REMOVED the very next day** (R-196 / R-204 item 2) precisely because it marked healthy escrows stale. **Q2 — the flag is wrong, measured on both sides:** the hub's blob seals `restic_pw_sha256 = 8a9e33aa4da6769c…d080a`, and the key the box is actually using hashes to **the identical value**. The blob covers the key. **Q3 — nothing clears it by itself:** the ONLY writer of `stale_at = NULL` is `SaveHostEscrow`'s `ON CONFLICT` — i.e. a fresh escrow ceremony, **which is the one act that would supersede the good blob**. So the only exit from a false alarm is the destructive act the false alarm recommends. **Q5 — the blast radius, enumerated rather than assumed:** (1) `GetEscrowStatusForCustomer` withholds `restic_pw_sha256` from the ACK; (2) a **pending** box can never auto-confirm, so (3) **every off-site run is refused indefinitely** — neither bites `demo-hp`, which was already `escrowed` and is backing up healthily (12 snapshots, last success 2026-08-07T02:15:35Z); (4) the customer is told to create a new code; (5) **NEW — v0.206.0's shape (c) is inert**, because the box records an empty hub hash and falls back to (a)/(b), so the recovery screen would stay silent even if recovery were needed. **Q6 — a fresh box CANNOT reach this state:** `MarkEscrowStale` has **no production caller anywhere in the tree** (census: only its own definition, two comments and two test references). **The next walk cannot meet it.** **NOT CLEARED, deliberately** — Q2's answer says the fix is to clear this instance, but whether to also stop the column being settable at all is a separate ruling, and the spike was scoped read-only. **✅ THE FLAG IS CLEARED — operator-approved and applied 2026-08-08.** One row, identity-matched on `host_id` and guarded on `stale_at IS NOT NULL`; `changes()` returned **1**. **Verified end to end, not just in the database:** the hub now serves the hash again, the box recorded `hub_escrow_key_sha256 = 8a9e33aa4da6769c…d080a` at `11:10:19Z`, and that is **byte-identical to the key it is using** — so shape (c) compares, matches, and correctly stays silent. **The false stale warning is gone, proven with a positive control** rather than an absent line: **0** `escrow-confirm` lines since the restart while **5** scheduler lines in the same window prove the box was logging, and the recorded hash proves an ACK was processed. *(Method note: the hub pod is Alpine with no `sqlite3`; it was installed into the container's ephemeral writable layer — image and node untouched, gone on restart. SQLite's own file locking coordinated the write with the live hub; an earlier attempt failed cleanly on quoting and changed nothing, which is the fail-safe working.)* **STILL OPEN under this ID: the ruling on whether `stale_at` keeps a live setter at all.** It currently has NO production caller, so the column is write-only-by-accident — a field that changes behaviour, that nothing sets and nothing can see (R-248). Either give it an evidential setter or retire it; **do not leave it as a trap that only a database read can spring.** | **READY — flag cleared; the column ruling is still owed** — owner Viktor | | **R-247** | **The box is being told something false, in its own words, and it recommends the destructive act.** `demo-hp`, live, on controller v0.206.0: *"STALE escrow: the hub's current blob carries NO password hash (hash-less supersession) — the stored recovery bundle does not cover the offsite password; create a new recovery code"*. **Every clause is false.** The hub HAS the hash and is *withholding* it (R-246); there was no supersession (`host_escrow_superseded` has no row for this host); and the bundle **does** cover the password — the hashes match exactly. It raises `EscrowStale`, which renders the customer-facing card *„A letétben lévő helyreállítási csomag nem fedi a jelenlegi távoli mentési jelszót. Hozzon létre új helyreállítási kódot."* **THE CAUSE IS A FACT ON THE WIRE THAT THE BOX THROWS AWAY.** The hub already sends `escrow_stale` in the ACK (`json:"escrow_stale,omitempty"`), and the controller's `report.EscrowStatus` **has no matching field**, so `encoding/json` drops it silently. The box therefore cannot distinguish *withheld because flagged stale* from *genuinely hash-less*, and guesses the latter. **This is R-241's shape for the third time: the answer is available, and it is discarded at the boundary.** **The fix is small and is NOT made here** (§0 forbids a controller change this session): add the field, and say the true thing — or say nothing, since on an ESCROWED box with matching hashes there is nothing wrong to report. | **READY** — owner Viktor | | **R-248** | **A flag that changes behaviour is visible to nobody who would look for it.** Q4 of the 2026-08-08 spike, answered plainly. **The customer** sees only a derived card stating a false reason (R-247). **The box** cannot see it at all (R-247's dropped field). **The operator** can see it on exactly ONE page — the **PBS-DR** view (`hub/internal/web/pbsdr.go:487`, `v.EscrowStale = escrow.StaleAt != ""`) — which is the wrong tier for this symptom: an operator investigating an OFF-SITE problem has no reason to open a PBS-DR page. **No alert, no report field, no off-site surface.** The one-shot `escrow_stale` event fired on 2026-08-04 and **was never notified** (a full census of `notification_log` for that customer that day returns 8 rows, none of them this one); it has fired twice ever, both on 4 August. **So the practical answer is: only a database read.** That is a finding in its own right — a flag that silently changes behaviour and cannot be seen is the shape this fortnight has been about (R-241's discarded comparison, R-228's unread field, R-243's unobserved state). **What it needs:** surface `stale_at` on the off-site operator surface and in the host report, or stop using a field nobody can observe to change what a customer is told. | **READY** — owner Viktor | +| **R-249** | **The retrieval passphrase ships in the customer page's HTML, so any headless read puts it in a transcript.** Found 2026-08-07 on the fifth walk **by doing it**: driving the documented rebuild path, the cleartext value reached the session transcript from a `grep` over the saved page — **not** from pressing Reveal. The hub renders the per-customer **Retrieval Password** masked (`••••••••`) behind a Reveal/Copy control, and carries the cleartext in a **`data-secret="…"` attribute** of that control; `hub/internal/web/render_test.go:169` pins exactly this (*"data-secret not populated for the reveal control"*). **The masking is a presentation control, not a containment one.** **Why it is a defect and not a design note:** the same page's own hint reads *"never place it on a command line (the installer reads it at a no-echo prompt)"* and the generated install command deliberately omits it — **the page states the threat model and then violates it in its own markup**. Anything that reads the page rather than looking at it (curl, a scrape, an agent, a saved HAR, a support bundle) gets the secret with no reveal action **and no audit event** — where the break-glass credential, which has the same sensitivity, emits `recovery_credential_revealed` when revealed. **Proposed shape:** serve it from a POST endpoint the button calls — the break-glass credential already works exactly this way (`POST /hosts/{id}/reveal-recovery-credential`) — and render the page with no value in it; that also gets the reveal audited, which it is not today. **Severity MEDIUM:** it is a live per-customer secret that fetches the whole config (`GET /api/v1/config/` with `X-Retrieval-Password`), but the exposure is to someone who can already read the operator page — a defence-in-depth failure, not a boundary crossed. **The walk5 instance is compromised (it is in a transcript) and dies with that customer's teardown, which is owed.** | **READY** — owner Viktor | +| **R-250** | **A customer create can fail fail-closed because the host-key scan ladder is shorter than the DNS/AAAA settle time.** Found 2026-08-07 creating the fifth walk's venue. `POST /configs/new` with off-site enabled provisions the Storage Box sub-account and then scans its SSH host key to pin it — **fail-closed by design** (`offsite.go:111-121`, *"don't serve a descriptor the controller can't verify"*), with `defaultScanBackoff` = 2+4+8+16+30 ≈ **60 s**, sized by its own comment to *"the observed DNS propagation lag"*. **Measured, both halves:** the first create exhausted the ladder — five `no such host`, then `dial tcp [2a01:4f8:bacc:2:200::d30]:23: connect: network is unreachable` (the name had just begun resolving, **AAAA-first, into a pod with no IPv6 route**) — and the hub logged `[ERROR] offsite provision for walk5`. An **identical second POST succeeded** ~70 s later on its own final rung (`shared already provisioned … (subaccount 285351)`). **Total settle ≈ 100 s against a 60 s budget.** The endpoint was never the problem: verified afterwards from both the node and the hub pod, by name, `SSH-2.0-OpenSSH_9.6p1`. **Two distinct things are wrong and should not be merged:** (1) the budget is sized against DNS *existence*, but what actually bit is the **AAAA-before-A** window, a different and longer phenomenon — lengthen the ladder and/or prefer the A record for the scan dial; (2) **the operator is told nothing actionable** — the create fails with a generic error, the remedy is "press it again", and nothing says so. The retry is genuinely safe (idempotent on the label; confirmed afterwards — `one_time_secrets` 1, one sub-account, no double-provision) **but that safety is invisible to the person deciding whether pressing again will double-charge them.** **Severity LOW-MEDIUM:** self-clearing, no data at risk — but it is the first thing a new customer's provisioning does and it fails looking like an outage. | **READY** — owner Viktor | +| **R-251** | **The recovery listing renders one row per restic TAG, so the customer is shown an "app" they never installed and their data counted twice.** Measured on the fifth walk, 2026-08-07, on the screen the customer reaches after entering R. The snapshot carries tags `felhom-offbox,calibre-web`; the listing renders **two rows** — `calibre-web · 2026-08-07 14:57 · 12.8 MB` and `felhom-offbox · 2026-08-07 14:57 · 12.8 MB`. `felhom-offbox` is the tier's own marker tag, not an application. **The screen's whole job is to let the customer check that what is in the store is what they expect** (*"Nézd át, hogy tényleg azt találod-e itt, amire számítasz"*), and it shows them a stranger's name beside their own data and a total that is double the truth. **Cosmetic, not a data defect** — the restore page correctly offers only `calibre-web`. **Fix:** filter the marker tag out of the listing, or key the rows on the app tag. | **READY** — owner Viktor | +| **R-252** | **After a rebuild the restore refuses because the data drives are not registered, and nothing on the recovery path says so.** Measured on the fifth walk, 2026-08-07, at the last step of a successful recovery. The customer enters R, sees the listing, presses through to the restore — and gets **„nincs elérhető adatmeghajtó a visszaállításhoz"**. The drives physically survived (the raw mounts are the R-220 condition and were deliberately left in place); what did not survive is their **registration**, which lived in the destroyed guest's settings. **It is recoverable without a shell** — `GET /api/disks/candidates` offers both disks (`mountable: true`, `data_bearing: true`) and Tárhely → Meghajtók → „Meglévő meghajtó csatolása" re-registers them — **but nothing tells the customer that, and the recovery screen's own hand-off ("Tovább a visszaállításhoz") walks them straight into it.** **This is why the walk's journey half passes on the letter and not the spirit:** no guest command line was needed, and a customer who did not already know the product would stop here. **Fix:** detect the unregistered-drive state on the recovery/restore path and say what to do, or offer the re-attach inline. **Related:** R-253, which is the very next step and worse. | **READY** — owner Viktor | +| **R-253** | **The restore page promises it will reinstall the app, and the restore then refuses because the app is not installed — in the customer's own language, three lines apart.** Measured on the fifth walk, 2026-08-07. `/backups/restore` lists the app with the note **„Nincs telepítve — a visszaállítás előbb újratelepíti."** (*not installed — the restore will reinstall it first*). Pressing through to the full restore returns **„A teljes visszaállítás sikertelen: a(z) calibre-web nincs telepítve — előbb állítsd helyre az alkalmazást, utána az adatokat"** (*…not installed — restore the application first, then the data*). **Both sentences are on the recovery path, both are addressed to the same customer, and they say opposite things about the same fact.** The customer clears it by redeploying from the catalog (customer-facing, ~90 s) and re-running the restore — which then works — but they must work out for themselves that the first sentence was wrong. **This is the R-203/R-234/R-240 class one level up:** not a warning misread as a success, but a promise the next screen contradicts. **Fix:** either make the reconstitute path deploy the app when it is absent (which is what the list already claims), or change the list's note to say the app must be deployed first. **Do not leave the two sentences both shipped.** | **READY** — owner Viktor | **Recorded against existing rows by Phase 2:** @@ -252,7 +257,7 @@ unconditional promise, which CAMPAIGN-11 §7 step 7 measured the customer-facing | **R-198** | **The hub's superseded-escrow retention does NOT retain the offsite repository password — and the ceremony the system tells the customer to run is what destroys the last copy.** `host_escrow_superseded` has **no `identity_blob` column**, and `demoteCurrentEscrowTx` (`hub/internal/store/store.go:2547-2556`) copies only `host_id, blob, key_fingerprint, posture, created_at, restic_pw_sha256`. `blob` is the **K-escrow** (the PBS datastore key, PBS-native scrypt); the **restic repo password lives in `identity_blob`** (`felhom-agent/internal/escrow/identity.go:34-39`, age-wrapped `IdentityBundle`). Measured live: both hosts' current rows hold `blob`=383 B **and** `identity_blob`=572 B; both superseded rows hold `blob`=383 B and nothing else | **SHIPPED** (hub **v0.93.0**, 2026-08-04) | — | **This is the NINTH entry in `CLAUDE.md`'s table of comments asserting an invariant the code does not provide, and the first that is ALSO customer-facing copy.** The claim appears three times: the schema comment (`store.go:370-375`, *"so the old passphrase stays customer-R-recoverable … turns the reinstall-orphan incident from 'history destroyed' into 'history recoverable'"*), the capability map's escrow row, and — in Hungarian, to the customer, on the orphan card — `controller/internal/web/templates/backups_remote.html:66,69` (*„a hozzájuk tartozó helyreállítási kóddal később visszaállíthatók lehetnek"*). **For the offsite restic repository, the incident it names, all three are false.** **Why it is worse than a missing column:** a rebuilt box lands in `EscrowState: pending`, offsite runs are blocked (`OffboxRunnable`, `offbox.go:570`), the card says „Helyreállítási kód szükséges", and the controller's own detector logs *"run the escrow ceremony"* (`report/escrow_confirm.go:100`) — so **the prescribed remedy is the act that overwrites `host_escrow.identity_blob` and loses the old password forever**. Both demo boxes crossed that line on **2026-08-04 at 07:15:36 (demo-hp) and 07:20:08 (demo-felhom)**. **This is an independent, stronger reason the 51 orphaned demo snapshots are unrecoverable than "nobody kept the recovery codes" — keeping R would not have helped.** **Fix shape (small):** add `identity_blob` to the superseded table and to `demoteCurrentEscrowTx`'s SELECT list; pin it with a test that asserts the CONSEQUENCE (a superseded row can still yield a repo password) rather than the mechanism. **Then correct all three claims in the same commit** — including the Hungarian card, which must not promise what the system cannot do. **Prerequisite for R-193's (d)-alone branch** and for the drill. Source: `audits/RECON-offsite-dr-chain-2026-08-04.md` §7 **SHIPPED.** `host_escrow_superseded` gains `identity_blob` (CREATE + additive `ALTER TABLE`) and `demoteCurrentEscrowTx` carries it, so **both** callers — re-escrow and host-delete demotion — are fixed by one change to the routine its own comment calls *"THE ONE escrow row-copy routine"*. `ListSupersededEscrow` reads it back and `store.HostEscrow` gains `IdentityBlob`, so a retained blob is reachable from Go at all. The table comment now records that the ruling stated above it was not met and what that cost. **Tests assert the CONSEQUENCE, which is why the existing one stayed green:** `TestSaveHostEscrow_RetainsSuperseded` asserted that a retained row exists carrying the old K-blob and passed throughout; `TestSaveHostEscrow_RetainsIdentityBlob` asserts the retained row can still yield a repository password, and pins the load-bearing ordering (the identity blob is written by `SaveHostDRBundle` AFTER `SaveHostEscrow`, so the demote sees the PREVIOUS generation — if that inverts, the retained bytes would be the new blob filed under the old hash, recoverable-looking and wrong). `TestDeleteHost_DemotesIdentityBlob` proves the shared routine through its other caller. **Red-proofs, both observed failing:** dropping `identity_blob` from the copy (production behaviour ≤ v0.92.0) fails BOTH scenarios; fixing only the re-escrow caller fails the delete scenario while the re-escrow one passes — the §8.2 mistake, demonstrated rather than asserted. **Nothing was backfillable and it was CHECKED, not deduced:** rows superseded before v0.93.0 were written without the blob and their source rows are already overwritten; the live database holds exactly 2 retained rows, both from 2026-08-04, both `identity_blob` NULL. **Who the fix protects, measured live:** 2 of 2 hosts with a current escrow carry an identity blob (`demo-felhom-8363b5`, `demo-hp-bb76ea`) — their NEXT ceremony now retains a recoverable off-site key instead of destroying one. **STILL UNIT-PROVEN ONLY after the 2026-08-04 night drill.** Part 2 (wipe again, do NOT recover, let a ceremony seal a DIFFERENT password, then inspect the superseded row's `identity_blob`) was gated on the first drill passing and **did not run** — the verdict was not reached, and a second wipe would have destroyed the state that makes the first one finishable in five minutes. **Nothing has yet superseded a key in production**, so the retention's live behaviour is unobserved. That check remains the cheapest way to prove or disprove it. | CC | | **R-199** | **The hub serves recovery blobs on two endpoints that have no client anywhere in the system.** `handleReEnroll` (`hub/internal/api/dr.go:101`) and `handleGetRestoreDirective` (`:155`) return `identity_escrow_b64` + `k_escrow_b64`, gated on operator-armed recovery mode. Census: **zero** callers in `felhom-agent` (no `ReEnroll` symbol at all; `internal/hub/client.go` reaches only `desired-state`, `wg`, `pbs/consume-token`, `jobs`), **zero** in the hub UI or any template, **zero** in `scripts/` or any runbook | **SHIPPED + PROVEN-LIVE 2026-08-04** (hub **v0.94.0**, agent **v0.125.0**) | — | **The documented retrieval path is a human with `sqlite3`:** `SELECT writefile('/root/idblob', identity_blob) FROM host_escrow …` on a `kubectl cp`-ed `hub.db` — recorded in project memory as the 2026-07-04 S5 prep, where hub-side blob serving was explicitly deemed *"Part 3 NOT needed"*. **This is the built-but-never-wired class at the DR capstone**, and it is why the chain from a dead node to an open repository has no automatable middle. **What a recovery flow actually needs is smaller than what exists:** a narrow `GET /hosts//escrow` authed with the box's own per-host key, serving opaque bytes to the box that owns them — zero-knowledge untouched, and far lighter than `re-enroll`, which **rotates the host API key** and returns the new key in the response body (`dr.go:130,148`). **Decide before building:** whether `re-enroll`/`restore-directive` should get a client, be replaced by the narrow GET, or be retired. Source: `audits/RECON-offsite-dr-chain-2026-08-04.md` §3 link 6 **LINKS 6, 7 AND 8 ARE ASSEMBLED AND WALKED.** **Link 6:** `GET /api/v1/hosts/{host_id}/escrow` — the box-authenticated MIRROR of the PUT that stores the blob, self-scoped by the per-host key (global may read any, the same asymmetry the PUT has). The operator-driven DR endpoints are UNTOUCHED and pinned by a test that exercises them with recovery mode off and on. **Link 7:** `POST /escrow/recover-offsite-password` on the agent's pinned local API gives `UnwrapIdentityBundle` its first production caller in two months. **Link 8:** it extracts and returns **only** the repository password (not the tunnel token, not the PBS token, not the WG key — the controller is a trust tier down). **PROVEN ON HARDWARE, demo-felhom, 2026-08-04 13:49 CEST:** on-disk `c60c8bc737a6…` vs recovered `c60c8bc737a6…` — **MATCH**, and the same hash the hub independently stores as `restic_pw_sha256`, so three sources agree. **Scenario B proven live 5 minutes earlier** with a deliberately wrong code: hub served the blob (572 B, `self_scope=true`), agent logged *the recovery code did not unwrap the identity escrow … exit status 1*, nothing written — which also proves links 6 and 7 ran independently of the success. **Scenario E proven live:** both retrievals raised `escrow_blob_served` (warning, operator-only); the first mailed the operator, the second was cooldown-suppressed AND that suppression is itself recorded; the customer leg reads `skipped/operator_only` on both. **R persisted nowhere, searched not claimed:** 0 lines in the agent journal, 0 in the controller log, 0 files under `/tmp`, `/var/tmp`, `/var/lib/felhom-agent`, `/root`, 0 leftover `felhom-idesc-*` staging dirs, and the staged-secret dir empty — with a **positive control** (a planted copy was found: 1, then 0 after removal) so the sweep is a measurement rather than an unfalsifiable absence. **THE §8.2 TRADE, made deliberately and recorded in the handler:** obtaining the blob used to require the operator to arm recovery mode; it now needs only the box's own credential. They still cannot open it (the hub never held R; a wrong code fails closed at age's scrypt KDF). `escrowSelfServiceRetrieval` is a single named constant — flipping it to false re-imposes recovery mode and changes nothing else, so the operator can overrule the trade for the cost of a boolean. **Red-proofs observed:** removing the ownership check served host B's blob to host A; removing the audit record made the retrieval silent; returning `PBSToken` instead of `ResticRepoPassword` yielded a plausible bundle with a non-matching key; commenting the `Options.EscrowRecovery` wiring failed the AST seam test. **Seam discipline:** the wiring is asserted by walking `main` → `runDaemon` → `buildLocalAPIServer` and checking the composite literal, not by `strings.Contains` — links 6 and 7 were two of this project's six built-but-never-wired instances and the fix must not become the seventh | CC | | **R-200** | **The DR password-injection seam has a handler, a route and tests — and no form.** `POST /backup/offbox/inject-password` is routed (`controller/internal/web/server.go:510`) to `offboxInjectPasswordHandler` (`offbox_handlers.go:174-196`) → `InjectOffboxPassword` (`backup/offbox.go:541`). **No template in the repository contains that path or any form posting to it** (grep over `internal/web/templates/`: one unrelated hit, an XSS comment) | **PLUMBING COMPLETE** (controller **v0.196.0**); **the FORM is not built — still open** | — | **The tenth instance of this project's built-but-never-wired class, and the exact shape `CLAUDE.md` and `felhom.eu/CLAUDE.md`'s seam-wiring rule were written for:** handler tests that POST directly (`offbox_escrow_test.go:167,180`) prove nothing about reachability. To use the only implemented recovery seam today, a person must hand-craft an authenticated POST with a session cookie and CSRF token. **Note the layering while fixing it:** this form takes a **64-hex repo password**, not a recovery code (`offboxRepoPwPattern`, `offbox.go:543`) — they are different secrets at different layers, and the operator's 2026-08-04 ruling asks for a form that takes **R**. Build the R form and treat this one as the operator/DR fallback it was written as, but **ship it with a render test per branch of whatever gate it sits behind**. Source: `audits/RECON-offsite-dr-chain-2026-08-04.md` §3 link 9 **THE DIAGNOSTIC HALF IS DONE AND IT ANSWERED THE QUESTION.** `--recover-offsite-check` is a `docker exec` escape hatch in the shape of `--print-reset-code`: R on **STDIN** (never argv, never `ps`, never shell history, never a transcript), fetch+unseal via the agent, and a verdict of **two sha256 hashes**. **It compares and never installs** — the recovered password is not written to `offbox/repo_password`; a test asserts the data dir is byte-unchanged and its red-proof (adding the install call) fails it. Confirmed live: `repo_password` mtime still `2026-08-03 07:18:02` after the successful check at `2026-08-04 11:49`. **Exit codes are load-bearing** — `0` match, `2` a clean MISMATCH, `1` a step failed; "it failed" and "it worked and disagreed" must never share a status because only one is a finding about the system. A box with no local password reports distinctly (the rebuilt-box shape, where the next step is to INSTALL rather than compare). **WHAT IS NOT BUILT, deliberately:** no card, no form, no preview, no customer-facing text — building an interface on top of a chain nobody had walked is how the preceding three weeks went wrong. **What remains for this row:** link 9 (the recovered password placed so `WriteOffboxSecrets` keeps it) and the customer-facing shape the operator ruled on 2026-08-04 (yell → R form → preview → proceed), which is now priced against a chain that exists rather than one that is assumed **PART 0 SHIPPED 2026-08-04 (v0.196.0):** `--recover-offsite-install` is the sibling of the check — same fetch/unseal path, same STDIN discipline for R — and it **places** the recovered password via `InjectOffboxPassword`. The confirmation is a SECOND invocation (`--confirm-install`): without it, both hashes print and nothing is written, so the operator sees the comparison before a write is possible. Three outcomes named distinctly: **installed** (no local password — the rebuilt-box shape), **unchanged** (identical key, nothing written), **refused** (a DIFFERENT key present — installing would clobber the key the current repository is encrypted under; exit 2, no force offered). It re-reads the file after writing rather than trusting the call. **Red-proof observed:** removing the confirmation gate makes the dry run write the password. The R-persistence test carries a **positive control** (a planted copy found, then removed and not found). **NOT YET EXERCISED AGAINST A LIVE RECOVERY** — the R-201 drill halted before step 9, so this is unit-proven only. **What remains for this row:** the customer-facing shape the operator ruled on 2026-08-04 (yell → recovery-code form → preview → proceed) | CC | -| **R-201** | **Nothing in the offsite DR chain has ever been exercised past the ceremony — and the one live proof that exists predates the field it is cited for.** `slice10d-identity-restore-spike-findings.md` §1 (2026-06-10) proved a real wrap→unwrap of an identity bundle with a real R on a secret-less box, byte-identical, wrong-R failing closed. **That bundle was `{tunnel_token, pbs_token}`.** `ResticRepoPassword` was added in agent **v0.77.0 on 2026-07-09** — a month later — and is **unit-proven only** | **PASSED + PROVEN-LIVE 2026-08-04** — the customer file came back byte-identical | — | **Never exercised, in these words:** a blob has never been served to a box; a fork-4 bundle has never been unsealed with a real R outside a unit test (the only production caller is `--selftest=identity-consume`, `cmd/felhom-agent/main.go:2845`, reading R from `FELHOM_RECOVERY_CODE`); a recovered repo password has never been injected; an existing offsite repository has never been reopened with one; **no restore of any kind has ever been performed from a recovered secret.** The one live consume ever prepared (S5 Part 4-B, 2026-07-04) is still marked *"DEFERRED, not run"*. **`_recovery-inventory-2026-07-28.md` already recorded this accurately** (A.2.7, C.1 row 1, C-2) — this row exists so the gap has an owner and a closing condition, not only a description. **Closing condition = the drill** designed in `audits/RECON-offsite-dr-chain-2026-08-04.md` §10, whose pass condition is a **byte-identical sha256 of a sentinel file restored after a wipe** — explicitly NOT "the repository opened". **Do not run it before R-198 is fixed**: the drill would walk a chain that is missing a link everyone believed was there **MATERIALLY ADVANCED 2026-08-04, AND THE DISTINCTION MATTERS.** What was proven on hardware is that the offsite repository **password** comes back out of the sealed bundle byte-identical (R-199). What has STILL never happened: a recovered password **installed**, an existing repository **reopened** under one, and a **file restored** from it. **The drill's pass condition is unchanged and is not this** — it is a byte-identical sha256 of a sentinel file after a wipe and reinstall, and "the store opened" was explicitly ruled insufficient. Design: `audits/RECON-offsite-dr-chain-2026-08-04.md` §10. **The drill is now cheaper and better-founded than when it was designed:** links 6–8 are walked, so a failure during the drill can be localised instead of being an undifferentiated "recovery did not work", and R-198 means the ceremony that creates the drill's recovery code no longer destroys the key it is meant to protect **THE DRILL RAN AND DID NOT REACH ITS VERDICT, and stopping was the correct call** (`audits/DRILL-r201-offsite-recovery-2026-08-04.md`). Steps 1–4 completed; the wipe never happened; **nothing irreversible was done**. It halted because **the sentinel file was not in the off-site snapshot** (R-203) — wiping would have destroyed the only copy and proven nothing. **Sentinel sha256 `643166269103a25c…`, still on the box.** **What the attempt established live, all of it new:** (1) a rebuilt box's off-site run **refuses** with the orphan card and pushes `offbox_repo_orphaned` — the spike's predicted third outcome, measured for the first time, and it does NOT silently start a fresh history (this closes R-193's Q3); (2) the **orphan reset works** — move-aside to `/home/felhom-repo.orphaned-20260804`, never delete, fresh repo initialised, `offbox_repo_reset` pushed (operator-authorised during the session); (3) demo-hp's pre-rebuild history is **permanently unrecoverable** — its key sits in superseded row id 3 with `identity_blob` NULL, superseded 07:15:36, **four hours before v0.93.0 fixed the retention**; (4) **a precondition the runbook did not contain:** neither pre-existing off-site-toggled app has a restorable file leg — both are named-volume-only, which the off-site tier tars but the customer restore never unpacks — so a file-leg app (`calibre-web`, mandatory `userdata: media/books`) had to be deployed, and it is now in place as the fixture. **TO RESUME:** fix or scope R-203, re-run steps 4–5 and confirm the sentinel IS in the snapshot **by listing it, not by a green status**, then P4 + the STOP + steps 6–11. Everything else is already staged: versions, recovery code, working repository, file-leg app, sentinel. **The pass condition is unchanged** — a byte-identical sentinel sha256, not "the store opened" **UNBLOCKED 2026-08-04 by R-203 (controller v0.197.0).** The sentinel now lands in the off-site snapshot and is listed there by name and size — the thing whose absence halted the drill. **Everything the resumed run needs is already in place on demo-hp:** agent v0.125.0, controller v0.197.0, the recovery code held by the operator (`R_DEMO-HP`), a working off-site repository (3 snapshots), `calibre-web` deployed with a mandatory userdata path, and the sentinel at sha256 `643166269103a25c…` — **verified byte-identical after the R-203 migration moved it to the corrected directory**. **What remains is exactly steps 4–11 of the drill**: re-verify the snapshot listing, take the deliberate rollback archive (P4), the §7 STOP, then wipe, reinstall, recover, install, restore and compare. **The pass condition is unchanged** — a byte-identical sentinel sha256, not "the store opened" **NIGHT RUN 2026-08-04 (`audits/DRILL-r201-night-run-2026-08-04.md`). THE HEADLINE: after a real rebuild — controller data volume destroyed, sentinel deleted from disk — the customer's recovery code produced `8a9e33aa4da6769c5aea1831f87759e10930e2ec1dea0062576484e0598d080a`, BYTE-IDENTICAL to the pre-wipe on-disk key and to the hub's independent record. The off-site backup key is recoverable after a machine is rebuilt, and that had never been shown.** **Step 7's assertion PASSED:** `identity_blob` 572 B and `restic_pw_sha256` unchanged across the wipe (`updated_at` still 11:11:37) — nothing re-escrowed itself. **Step 9a passed:** the key installed cleanly on a bare box (the "installed" branch's first real run); **9b:** the apply kept it. **THE VERDICT WAS NOT REACHED** — step 10 never ran, so there is no post-restore sha256 and no snapshot count. **That is not a FAIL** (nothing came back wrong and no fresh history was started); it is a wall, and the wall is **R-204**. **The wipe was faithful to the incident, deliberately:** the 2026-08-03 rebuild R-193 is filed against was NOT a guest reprovision — the journal shows guest 9201 running continuously with no `pct destroy`/`pct restore`/`--selftest=provision` — so a controller-data-volume wipe reproduces it, and an unrehearsed provisioning chain improvised unattended is what §8.10 exists to prevent. **A precondition had drifted and was repaired, not worked around:** the staged snapshot had lost the sentinel because the afternoon's experiment produced a later same-day snapshot and `forget --keep-daily 7 --group-by host,tags` pruned the good one — **a good snapshot is not durable against a later bad run on the same day.** **TO FINISH (~5 min, operator present):** re-claim the box, confirm the escrow (**NOT a new ceremony** — it would supersede the identity blob and destroy the key under test), run a backup, restore the sentinel and compare to `643166269103a25c…`. Rollback available: the verified archive `vzdump-lxc-9201-2026_08_04-21_44_24.tar.zst` **THE DRILL PASSED (`audits/DRILL-r201-night-run-2026-08-04.md`).** demo-hp's controller data volume was destroyed and the sentinel deleted from disk; the customer's recovery code recovered the key (`8a9e33aa4da6…`, byte-identical to the pre-wipe on-disk key AND to the hub's independent record); it installed on the bare box; the **existing repository OPENED** (`repo_state: null`, **3 snapshots — NOT 1**, 42 026 B = the pre-wipe size exactly); and the customer restore flow returned **`643166269103a25cf41d34a26b75fd6ebba0a837bbcd7e0d2b5aba7106bcbe7c` — byte-identical to the pre-wipe sentinel.** **The Felhom backup story is proved end to end for the first time.** **Step 7's assertion held:** `identity_blob` 572 B and `restic_pw_sha256` unchanged across the wipe — nothing re-escrowed itself, and no ceremony was run at any point (superseded rows still 2). **The wipe was faithful to the incident:** the 2026-08-03 rebuild R-193 is filed against was a controller-DATA-VOLUME loss, not a guest reprovision — the journal shows guest 9201 up throughout with no `pct destroy`/`pct restore`/`--selftest=provision`. **IT TOOK FOUR UNDOCUMENTED STEPS (→ R-204):** an operator Re-issue; a re-claim (whose escape hatch needs a controller restart to work at all); a manual escrow confirm; and `mode=full` on the restore, because the default `mode=unit` returns the app definition and NOT the customer's files. **A customer hitting this alone today would not get their data back.** Box left healthy, claimed, re-armed, sentinel restored to its live location. **⚠ WALKED TO COMPLETION 2026-08-06/07 — DATA PASS, JOURNEY FAIL, and the row stays open.** The full walk ran overnight on a new venue: built from the published ISO, fixtured, soaked through a real scheduled cycle, destroyed, rebuilt, and finished in the morning with the operator's emailed code. **DATA: PASS** — all three sentinels byte-identical out of snapshot `f5c53b03`, accented filename bytes included. **JOURNEY: FAIL** — the customer had **no route at all** to enter the code they hold: the recovery screen had retired itself, and the remote page offered to CREATE a new code instead. Root cause **R-241** (the credential self-heal writes a fresh repository key and moves the box out of the recovery-offer's pristine case, while orphan detection is unreachable behind `escrow_state: pending`). Recovery needed three guest command lines. **Also established:** the credential chain runs end to end unaided on an unclaimed box (first live sighting of its success line), and **R-239** — a fresh install lands on controller 0.203.0 while 0.205.0 is released, so R-234 and R-237 are undelivered. Full record: `tests/finalwalk-r201-2026-08-07/journal.md` | **READY** — owner Viktor (R-241 blocks the journey half) | +| **R-201** | **Nothing in the offsite DR chain has ever been exercised past the ceremony — and the one live proof that exists predates the field it is cited for.** `slice10d-identity-restore-spike-findings.md` §1 (2026-06-10) proved a real wrap→unwrap of an identity bundle with a real R on a secret-less box, byte-identical, wrong-R failing closed. **That bundle was `{tunnel_token, pbs_token}`.** `ResticRepoPassword` was added in agent **v0.77.0 on 2026-07-09** — a month later — and is **unit-proven only** | **PASSED + PROVEN-LIVE 2026-08-04** — the customer file came back byte-identical | — | **Never exercised, in these words:** a blob has never been served to a box; a fork-4 bundle has never been unsealed with a real R outside a unit test (the only production caller is `--selftest=identity-consume`, `cmd/felhom-agent/main.go:2845`, reading R from `FELHOM_RECOVERY_CODE`); a recovered repo password has never been injected; an existing offsite repository has never been reopened with one; **no restore of any kind has ever been performed from a recovered secret.** The one live consume ever prepared (S5 Part 4-B, 2026-07-04) is still marked *"DEFERRED, not run"*. **`_recovery-inventory-2026-07-28.md` already recorded this accurately** (A.2.7, C.1 row 1, C-2) — this row exists so the gap has an owner and a closing condition, not only a description. **Closing condition = the drill** designed in `audits/RECON-offsite-dr-chain-2026-08-04.md` §10, whose pass condition is a **byte-identical sha256 of a sentinel file restored after a wipe** — explicitly NOT "the repository opened". **Do not run it before R-198 is fixed**: the drill would walk a chain that is missing a link everyone believed was there **MATERIALLY ADVANCED 2026-08-04, AND THE DISTINCTION MATTERS.** What was proven on hardware is that the offsite repository **password** comes back out of the sealed bundle byte-identical (R-199). What has STILL never happened: a recovered password **installed**, an existing repository **reopened** under one, and a **file restored** from it. **The drill's pass condition is unchanged and is not this** — it is a byte-identical sha256 of a sentinel file after a wipe and reinstall, and "the store opened" was explicitly ruled insufficient. Design: `audits/RECON-offsite-dr-chain-2026-08-04.md` §10. **The drill is now cheaper and better-founded than when it was designed:** links 6–8 are walked, so a failure during the drill can be localised instead of being an undifferentiated "recovery did not work", and R-198 means the ceremony that creates the drill's recovery code no longer destroys the key it is meant to protect **THE DRILL RAN AND DID NOT REACH ITS VERDICT, and stopping was the correct call** (`audits/DRILL-r201-offsite-recovery-2026-08-04.md`). Steps 1–4 completed; the wipe never happened; **nothing irreversible was done**. It halted because **the sentinel file was not in the off-site snapshot** (R-203) — wiping would have destroyed the only copy and proven nothing. **Sentinel sha256 `643166269103a25c…`, still on the box.** **What the attempt established live, all of it new:** (1) a rebuilt box's off-site run **refuses** with the orphan card and pushes `offbox_repo_orphaned` — the spike's predicted third outcome, measured for the first time, and it does NOT silently start a fresh history (this closes R-193's Q3); (2) the **orphan reset works** — move-aside to `/home/felhom-repo.orphaned-20260804`, never delete, fresh repo initialised, `offbox_repo_reset` pushed (operator-authorised during the session); (3) demo-hp's pre-rebuild history is **permanently unrecoverable** — its key sits in superseded row id 3 with `identity_blob` NULL, superseded 07:15:36, **four hours before v0.93.0 fixed the retention**; (4) **a precondition the runbook did not contain:** neither pre-existing off-site-toggled app has a restorable file leg — both are named-volume-only, which the off-site tier tars but the customer restore never unpacks — so a file-leg app (`calibre-web`, mandatory `userdata: media/books`) had to be deployed, and it is now in place as the fixture. **TO RESUME:** fix or scope R-203, re-run steps 4–5 and confirm the sentinel IS in the snapshot **by listing it, not by a green status**, then P4 + the STOP + steps 6–11. Everything else is already staged: versions, recovery code, working repository, file-leg app, sentinel. **The pass condition is unchanged** — a byte-identical sentinel sha256, not "the store opened" **UNBLOCKED 2026-08-04 by R-203 (controller v0.197.0).** The sentinel now lands in the off-site snapshot and is listed there by name and size — the thing whose absence halted the drill. **Everything the resumed run needs is already in place on demo-hp:** agent v0.125.0, controller v0.197.0, the recovery code held by the operator (`R_DEMO-HP`), a working off-site repository (3 snapshots), `calibre-web` deployed with a mandatory userdata path, and the sentinel at sha256 `643166269103a25c…` — **verified byte-identical after the R-203 migration moved it to the corrected directory**. **What remains is exactly steps 4–11 of the drill**: re-verify the snapshot listing, take the deliberate rollback archive (P4), the §7 STOP, then wipe, reinstall, recover, install, restore and compare. **The pass condition is unchanged** — a byte-identical sentinel sha256, not "the store opened" **NIGHT RUN 2026-08-04 (`audits/DRILL-r201-night-run-2026-08-04.md`). THE HEADLINE: after a real rebuild — controller data volume destroyed, sentinel deleted from disk — the customer's recovery code produced `8a9e33aa4da6769c5aea1831f87759e10930e2ec1dea0062576484e0598d080a`, BYTE-IDENTICAL to the pre-wipe on-disk key and to the hub's independent record. The off-site backup key is recoverable after a machine is rebuilt, and that had never been shown.** **Step 7's assertion PASSED:** `identity_blob` 572 B and `restic_pw_sha256` unchanged across the wipe (`updated_at` still 11:11:37) — nothing re-escrowed itself. **Step 9a passed:** the key installed cleanly on a bare box (the "installed" branch's first real run); **9b:** the apply kept it. **THE VERDICT WAS NOT REACHED** — step 10 never ran, so there is no post-restore sha256 and no snapshot count. **That is not a FAIL** (nothing came back wrong and no fresh history was started); it is a wall, and the wall is **R-204**. **The wipe was faithful to the incident, deliberately:** the 2026-08-03 rebuild R-193 is filed against was NOT a guest reprovision — the journal shows guest 9201 running continuously with no `pct destroy`/`pct restore`/`--selftest=provision` — so a controller-data-volume wipe reproduces it, and an unrehearsed provisioning chain improvised unattended is what §8.10 exists to prevent. **A precondition had drifted and was repaired, not worked around:** the staged snapshot had lost the sentinel because the afternoon's experiment produced a later same-day snapshot and `forget --keep-daily 7 --group-by host,tags` pruned the good one — **a good snapshot is not durable against a later bad run on the same day.** **TO FINISH (~5 min, operator present):** re-claim the box, confirm the escrow (**NOT a new ceremony** — it would supersede the identity blob and destroy the key under test), run a backup, restore the sentinel and compare to `643166269103a25c…`. Rollback available: the verified archive `vzdump-lxc-9201-2026_08_04-21_44_24.tar.zst` **THE DRILL PASSED (`audits/DRILL-r201-night-run-2026-08-04.md`).** demo-hp's controller data volume was destroyed and the sentinel deleted from disk; the customer's recovery code recovered the key (`8a9e33aa4da6…`, byte-identical to the pre-wipe on-disk key AND to the hub's independent record); it installed on the bare box; the **existing repository OPENED** (`repo_state: null`, **3 snapshots — NOT 1**, 42 026 B = the pre-wipe size exactly); and the customer restore flow returned **`643166269103a25cf41d34a26b75fd6ebba0a837bbcd7e0d2b5aba7106bcbe7c` — byte-identical to the pre-wipe sentinel.** **The Felhom backup story is proved end to end for the first time.** **Step 7's assertion held:** `identity_blob` 572 B and `restic_pw_sha256` unchanged across the wipe — nothing re-escrowed itself, and no ceremony was run at any point (superseded rows still 2). **The wipe was faithful to the incident:** the 2026-08-03 rebuild R-193 is filed against was a controller-DATA-VOLUME loss, not a guest reprovision — the journal shows guest 9201 up throughout with no `pct destroy`/`pct restore`/`--selftest=provision`. **IT TOOK FOUR UNDOCUMENTED STEPS (→ R-204):** an operator Re-issue; a re-claim (whose escape hatch needs a controller restart to work at all); a manual escrow confirm; and `mode=full` on the restore, because the default `mode=unit` returns the app definition and NOT the customer's files. **A customer hitting this alone today would not get their data back.** Box left healthy, claimed, re-armed, sentinel restored to its live location. **⚠ WALKED TO COMPLETION 2026-08-06/07 — DATA PASS, JOURNEY FAIL, and the row stays open.** The full walk ran overnight on a new venue: built from the published ISO, fixtured, soaked through a real scheduled cycle, destroyed, rebuilt, and finished in the morning with the operator's emailed code. **DATA: PASS** — all three sentinels byte-identical out of snapshot `f5c53b03`, accented filename bytes included. **JOURNEY: FAIL** — the customer had **no route at all** to enter the code they hold: the recovery screen had retired itself, and the remote page offered to CREATE a new code instead. Root cause **R-241** (the credential self-heal writes a fresh repository key and moves the box out of the recovery-offer's pristine case, while orphan detection is unreachable behind `escrow_state: pending`). Recovery needed three guest command lines. **Also established:** the credential chain runs end to end unaided on an unclaimed box (first live sighting of its success line), and **R-239** — a fresh install lands on controller 0.203.0 while 0.205.0 is released, so R-234 and R-237 are undelivered. Full record: `tests/finalwalk-r201-2026-08-07/journal.md` **✅ CLOSED 2026-08-07 — BOTH HALVES PASS, on the fifth walk** (`documentation/tests/walk5-r201-2026-08-07/journal.md`). A fresh appliance was installed from the published ISO on `demo-hp` (VM 325), given three sentinels, escrowed, backed up off-site, then **destroyed on purpose** — guest and data volumes — and rebuilt through the documented day-0 path. **THE DATA: PASS** — all three sentinels byte-identical out of snapshot `5b0f20f7`, including the accented filename's *bytes*, read back with `os.listdir` on a bytes path so no decode round trip could launder a `U+FFFD`. **THE JOURNEY: PASS — zero guest command lines were needed to progress**, against three on the previous walk; the reset-code hatch was used once, in Phase A only, where §3 permits it. **RTO 71.7 s** from login to an open store (12.44 s of it the unseal itself; ~22 s a harness retry of mine). **The recovery screen appeared without being sought** (`/` → `/launcher` → `/recovery`) and answered all three questions, with a sealed-at timestamp matching `host_escrow.created_at` exactly. **What made the difference is R-241's mint guard, exercised live for the first time:** at 14:58:52Z the rebuilt box collected its re-staged credential, configured the transport and **refused to mint a repository password** over the sealed package — where the previous walk minted one and lost the journey silently. **Two obstacles remain and are filed rather than absorbed into this close — R-252 and R-253** — neither needing a shell, both cleared from the dashboard, and neither signposted; the journey succeeds and is not yet smooth. **This close does NOT claim** that shape (c) fired positively (it did not — with no local key the offer comes from shape (a); shape (c) was measured in its negative half in Phase A), nor that a customer would clear R-252/R-253 unaided. | **CLOSED 2026-08-07** | | **R-202** | **The orphan card promises the customer their old backups "may later be restorable with the matching recovery code" — and after v0.93.0 that is true for supersessions from now on and FALSE for anything already orphaned.** `controller/internal/web/templates/backups_remote.html:66,69` states it unconditionally, in Hungarian, on the one surface where being wrong costs most | **OPEN — Part 5 hit its gate 2026-08-04; the card is UNTOUCHED and the sentence is still live** | R-199/R-201 (which generation an orphaned repo belongs to is not knowable to the box today) | — | **THE GATE, and why it was hit rather than squeezed past.** The condition was: ship it iff the hub can tell a box what it needs with **one** additional boolean on the escrow ACK it already sends. The hub *can* cheaply compute *"≥1 retained blob for this host carries an identity blob"* — one correlated predicate in `GetEscrowStatusForCustomer`, and the controller even has the right seam already (`SetEscrowStale`/`StaleBlob` is exactly this shape). **But that boolean does not answer the card's question.** The card renders on `RepoState == "orphaned"`, and the promise is about *the key THIS orphaned repository was written under*. A box does not know which escrow generation the orphaned remote belongs to; a box with a pre-v0.93.0 orphan and a post-v0.93.0 supersession would read the boolean TRUE and the promise would still be false — a **conditional** falsehood that looks verified, which is strictly worse on a customer-facing card than today's hedged one. **What would actually make it truthful** is knowing the orphaned repo's generation, which is the same knowledge R-199/R-201's unassembled chain needs. **Interim exposure, stated rather than buried:** the sentence remains live and remains false for both demo boxes. **Cheapest honest interim** (not taken here — it is a customer-copy change and the gate said leave it alone): drop the recoverability clause and say only that the old history is set aside and not deleted, which is true unconditionally | CC + operator | | **R-203** | **A customer-declared MANDATORY data directory was silently absent from the off-site snapshot while the run reported `ok`.** Measured live on demo-hp 2026-08-04: `calibre-web` declares `userdata: media/books class: mandatory`; its live bind is `/mnt/sys_drive/userdata/media/books` (the sentinel file was there), while the off-site capture set looked for `/mnt/sys_drive/felhom-data/userdata/media/books`, which does not exist. Result: `[WARN] mandatory data path missing on disk, skipped from offsite`, then `backed up calibre-web (…, 0 mandatory path(s))` and `backup OK: 3 app(s), 3 snapshot(s)` — `last_status: ok`, `last_success` stamped, nothing customer-visible, nothing hub-visible | **SHIPPED + PROVEN-LIVE 2026-08-04** (controller **v0.197.0**) | — | **THE MECHANISM, from source.** `NamespaceRoot(drivePath, inGuestDrive)` (`appbackup/paths.go:28-33`) appends the `felhom-data` segment **when the drive IS the system data path** (`m.namespaceRoot` = `NamespaceRoot(drivePath, drivePath != m.systemDataPath)`, `backup/backup.go:331`). The deploy-time bind does not: `${USERDATA_PATH}` is `/userdata` (`stacks/classify_binds.go:14`). With `system_data_path: /mnt/sys_drive` and an app deployed at `HDD_PATH=/mnt/sys_drive`, the two produce different directories. **The same compose used BOTH roots**, from one deploy: `${IMPORT_PATH}` → `/mnt/sys_drive/felhom-data/userdata/import/calibre` (with the segment), `${USERDATA_PATH}` → `/mnt/sys_drive/userdata/media/books` (without). **WHAT IS MEASURED vs NOT, because it changes the fix.** MEASURED: the paths disagree, the mandatory directory is absent from the snapshot, the run says `ok`, and the only signal is a container-log WARN. **NOT ESTABLISHED:** whether `HDD_PATH=/mnt/sys_drive` is a SUPPORTED choice — it was used because demo-hp's only registered drive (`Felhom-Share`) is a NAS and was **correctly refused** as an app namespace (R-108 working as designed), while `/mnt/sys_drive` was **accepted** (HTTP 202). **EITHER BRANCH IS A DEFECT:** if the system drive is a supported app namespace, userdata resolution is wrong for every app on it and their mandatory directories are silently unprotected; if it is not supported, the deploy accepted a namespace it should have refused one call after refusing the NAS. **NOT a general off-site failure:** `opengist` and `privatebin` declare no mandatory userdata paths (everything of theirs is in named volumes), so they are unaffected and their snapshots are real. **Fix shape:** make the two roots one function, whichever is right — and make a skipped MANDATORY path a customer/hub-visible signal rather than a WARN, because `ok` with a missing mandatory directory is this project's own *a path the customer thinks is protected is not in the snapshot* shape. Source: `audits/DRILL-r201-offsite-recovery-2026-08-04.md` §2 **BOTH HALVES SHIPPED.** **(1) The paths.** `appbackup`'s helpers take a NAMESPACE ROOT; the census found **FIVE** bare-drive-path callers, not the four the spec named — the fifth is the **FileBrowser mount builder** (`web/handlers.go`), i.e. the customer's own file browser would have shown the wrong directory on a non-enrolled path (latent: the system drive is deliberately never a registered `StoragePath`). The rule now has **ONE expression** (`appbackup.NamespaceRootFor` / `IsEnrolledDrive`); there were already **two** copies and **they differed** — `backup.Manager.namespaceRoot` compared without `filepath.Clean`, `stacks.Manager.inGuest` with it, so a trailing slash from config would have flipped the mode in one package and not the other. `ComputeFabBuckets` now receives the namespace root, which is what `ComputeCaptureSet` has always received, so the export and the backup describe the same directories by construction. **(2) The verdict.** `last_status` gains **`incomplete`** — minted, because `ok`|`error`|`running` had nothing meaning *"it ran, and this app is not fully protected"*. **Not `error`:** the rest of the run worked, so `SnapshotCount` and the `LastSuccess` anchor still record what WAS captured. It reaches the operator through the **existing** per-run digest (`backup_run_failures`) — a new event type would be a two-repo change and the hub drops anything outside `allowedEventTypes`. **§8.4's narrowing is a NO-OP and no customer warning disappears:** `TierOffsite`'s `tierKeeps()` already admits mandatory only, demonstrated by widening the tier filter alone and watching the class check hold the line. **THE SPEC'S §8.3 RISK DOES NOT EXIST, and this is the correction owed:** `ExportDataMounts` lives in `delete.go` but is **export-only** — its single production caller is the `.fab` adapter, nothing deletes on its result, and the delete path's own guard `ProtectedHDDPaths` is layout-agnostic by construction (it protects BOTH `/…` and `/felhom-data/…`). It shipped as its own commit anyway. **PROVEN LIVE on demo-hp:** the bind moved `/mnt/sys_drive/userdata/media/books` → `/mnt/sys_drive/felhom-data/userdata/media/books`, the capture log went `0 mandatory path(s)` → **`1 mandatory path(s)`**, and **the sentinel is in the snapshot's own file listing** — `-rw-r--r-- 1000 1000 181 … /mnt/sys_drive/felhom-data/userdata/media/books/DRILL-SENTINEL.txt` — not a green status. **Red-proofs:** the bare-path call makes the two paths differ; inverting the drive-kind comparison breaks every enrolled row; leaving the export site bare emits the short path; and the verdict fails under both an unreachable gap-recording and an unconditional `ok`. **One red-proof PASSED and the test was wrong, not the code** — the first Scenario-C test only reached `offboxCaptureSet` while the mutation lives in `runOffboxInternal`; a run-level test replaced it | CC | | **R-204** | **A rebuilt box can recover its off-site key and still cannot use it: the remedy that reconfigures the tier is the thing that blocks the recovery.** Measured end to end on demo-hp during the 2026-08-04 night drill, after a real controller-data wipe | **ALL FOUR ITEMS CLOSED 2026-08-05** (items 1–3 controller v0.198.0 + hub v0.95.0; item 4 controller v0.199.0 + hub v0.96.0) | — | **THE CHAIN, each link measured.** (1) A rebuilt controller **cannot configure its off-site target at all**: `offsite-apply: consume one-time password: no unconsumed offsite password` — the previous controller consumed it (ledger: created `07:11:51`, consumed `07:12:06`). That is R-193, reconfirmed live. (2) The documented remedy is an operator **Re-issue**, which works — and **sets `stale_at` on the escrow** (measured: `2026-08-04 20:15:49`) **while `restic_pw_sha256` is unchanged**, i.e. R-196's false staleness. (3) A stale escrow makes the hub **withhold the hash from the report ACK**, so `EscrowAutoConfirmer` can never flip `pending → escrowed`, and `OffboxRunnable` (`configured && escrowed`) **refuses every off-site run**. (4) The only documented way to clear a stale escrow is **a fresh ceremony — which supersedes the identity blob and destroys the key being recovered.** **So the recovery and its precondition are mutually exclusive as built.** The key comes back (proven — `8a9e33aa4da6…` recovered byte-identical after the wipe, and installed) and then cannot be used to open the repository. **(5) A FOURTH link, undocumented anywhere:** a rebuilt box is **unclaimed**, and the claim gate correctly intercepts every non-claim route (`claim gate: intercepting POST /backup/offbox/confirm-escrow (unclaimed)`), so **no controller endpoint responds at all** until the customer re-claims. That step appears in no design document and comes first. **Fix shape, not decided here:** either the Re-issue stops marking the escrow stale on evidence it does not have (R-196's own fix), or a recovered-and-verified key is allowed to confirm the escrow without a ceremony — the hash comparison that `EscrowAutoConfirmer` already performs is exactly the evidence needed, and it is being withheld precisely when it would be conclusive. **Do NOT fix by widening `OffboxRunnable`** — the atomicity guarantee it enforces (no un-recoverable ciphertext) is the reason the escrow exists. Source: `audits/DRILL-r201-night-run-2026-08-04.md` §3 **ALL FOUR WALKED AND MEASURED during the passing 2026-08-04 drill, so this row is now a known-good manual runbook AND the gap list.** (1) **R-193:** a rebuilt controller cannot configure its off-site tier — `no unconsumed offsite password` (ledger: created `07:11:51`, consumed `07:12:06` by its predecessor). Remedy: operator Re-issue. (2) **The claim gate:** a rebuilt box is unclaimed, so the gate intercepts EVERY controller endpoint (`claim gate: intercepting POST /backup/offbox/confirm-escrow (unclaimed)`) — a first step of every recovery that appears in no design document. **And the local escape hatch does not work unaided:** `--print-reset-code` writes the new hash to `settings.json` while the RUNNING controller keeps its old copy in memory, so `effectiveClaimCode()` never sees it and the claim fails with *"Hibás vagy lejárt kód"*. **The controller must be restarted between minting and claiming** — two attempts failed before this was diagnosed. (3) **R-196:** the Re-issue sets `stale_at` (`20:15:49`) while `restic_pw_sha256` is unchanged → the hub withholds the hash → auto-confirm can never fire → `OffboxRunnable` refuses every run. Cleared here with the **manual** confirm (`/backup/offbox/confirm-escrow`) — NOT a ceremony, which would have superseded the identity blob and destroyed the recovered key. (4) **`mode=unit` is the restore default and returns the recovery unit, NOT the userdata leg.** A customer told to "restore from off-site" gets their app definition back and not their documents, and nothing in that outcome says so. **Fix priorities, in the order they hurt:** (4) is a customer-facing trap on the last step; (3) is R-196's fix; (2) needs the escape hatch to reload settings (or the claim state to survive a rebuild); (1) is R-193's restage. Source: `audits/DRILL-r201-night-run-2026-08-04.md` §3–§4 | CC + operator **OUTCOME 2026-08-05 — controller v0.198.0 + hub v0.95.0.** **Item 1 (the reset code needs a restart) — CLOSED, proven live.** `effectiveClaimCode` reads through to the persisted claim state, so a code minted by the separate `--print-reset-code` process is seen without a restart; the precedence rule between settings and config is unchanged. Read-through, not a TTL: a TTL leaves a window in which a superseded code still works, and that is the mutation `TestClaimCode_SupersededByASecondMint_RefusedImmediately` kills. Fails closed on an unreadable state. **Live on demo-felhom 9201, nothing restarted (`restarts=0`, container older than both mints): the superseded code returned „Hibás vagy lejárt kód" and the current one was accepted first time.** **Item 2 (a re-issue marks a healthy escrow stale) — CLOSED, → R-196.** Test-proven; deliberately NOT fired live on demo-hp. **Item 3 (the restore's default silently returns the wrong thing) — CLOSED, proven live.** A `mode=unit` restore now names what came back, what did not and the step that gets it; the wizard's intent card states its scope BEFORE the choice; the full-restore size gate is untouched and pinned as unchanged. **The default stays `unit`** — all three wizard forms set `mode` explicitly, so changing it would alter nothing the customer sees while silently changing a mode-less POST. **ITEM 4 REMAINS AND IS THE WHOLE OF WHAT IS LEFT HERE: a rebuilt box cannot obtain an off-site credential unaided**, because the one-time password was spent by its predecessor, so an operator Re-issue is still required. **Its dependency is the one-shot credential design decision — it needs an operator ruling and belongs to → R-193.** Not begun in this session, deliberately. **ITEM 4 CLOSED 2026-08-05 — controller v0.199.0 + hub v0.96.0.** The box now DECLARES that it needs a credential (`offsite.state=needs_credential`) instead of reporting an absence the hub cannot interpret; the hub's new `internal/offsiteheal` answers it. **Operator ruling, recorded because a ruling that lives only in a conversation binds nobody (R-96): the trigger is a state the BOX DECLARES, not an inference.** An absent off-site object has FOUR meanings — never configured, mid-restart, a transient config read failure, rebuilt-and-stranded — and the hub cannot tell them apart; the box can, from two local facts (a fresh data area AND a hub-held recovery package). **Both halves are required:** freshness alone would make every un-configured box in the fleet ask for a credential, which is what `TestOffsiteDeclare_NeverHadOffsiteSaysNothing` exists to catch. The reconciler mirrors `pbsdrheal`: declared states only, a two-DISTINCT-REPORT debounce (derived from the ~15-min report cadence), **restage before mint**, an event per remediation, and a healthy box is a pure no-op. **PROVEN LIVE:** demo-felhom 9201 arranged (reversibly) into the stranded shape produced report id=16743 carrying `{enabled:false, state:needs_credential, quota_gb:0, repo_size_bytes:0}`, and the single declaration was **absorbed by the debounce** — no self-heal event fired — with the box restored the same minute. **What is deliberately NOT automated: the escrow ceremony.** A credential is replaceable; the recovery code is not. **Credential automatic, key customer-present.** **Second ruling recorded: the dashboard-password exposure on the recovery preview is METADATA (backup dates, app names), not content, and is ACCEPTED.**| diff --git a/documentation/tests/walk5-r201-2026-08-07/teardown-owed.md b/documentation/tests/walk5-r201-2026-08-07/teardown-owed.md new file mode 100644 index 0000000..7499e93 --- /dev/null +++ b/documentation/tests/walk5-r201-2026-08-07/teardown-owed.md @@ -0,0 +1,73 @@ +# Teardown — OWED, not done (walk5, the fifth walk, 2026-08-07) + +**Deliberately not performed in this session, per §10: the machine is the evidence until the verdict +is written.** The verdict is written (`journal.md`); the venue still stands so it can be re-read if +anything in this report is questioned. **Predecessor ledgers:** `teardown-finalwalk-2026-08-07.md`, +`teardown-2026-08-06.md`. + +## The enumeration, matched on IDENTITY — never on size + +Size is not a key here: on 2026-08-06 an item that measured exactly the expected size turned out to be +a working store. + +| Layer | Item | Identified by | +|---|---|---| +| **machine** | `demo-hp` VM **325** `walk5-appliance` — 4 disks (efidisk + 200 G + 50 G + 50 G), **17 G actual** on `/mnt/nvme-1tb/images/325` | `qm config 325` → `name: walk5-appliance` | +| **hub** | customer **`walk5`**, host **`walk5-4bada5`** | `GET /configs/walk5/delete` preview + a FULL-SCHEMA census | +| **off-site** | Storage Box sub-account **285351**, user `u629488-sub4`, home naming the customer | Hetzner API, matched on the home directory | +| **off-site** | `ep0` PBS namespace **`walk5`** in datastore `felhom-offsite` | `ls /mnt/pbs-datastore/ns` — **the live store is `/mnt/pbs-datastore`; `/srv/pbs-felhom` is STALE and reading it gives a wrong answer in both directions** | +| **network** | WireGuard peer **10.77.0.5** | `wg_peers.host_id = walk5-4bada5`; verify removal on ep0's **live `wg show`**, not only in the hub DB | +| **DooPlex** | `~/.config/walk5/` — `dashboard_pw.txt`, `root_pw.txt`, `root_pw_rotated.txt`, `breakglass.json`, `retrieval_passphrase.txt`, `claim_code.txt` (all `0600`) | **`R_walk5.txt` is already `shred -u`'d, with a planted-copy control proving the sweep works** | +| **DooPlex** | the `walk5` block in `~/.ssh/config`, and the appliance's host key in `~/.ssh/known_hosts` | | + +## `pvesm status` BEFORE (2026-08-07, venue standing) + +``` +c11-scratch dir active 983379700 KiB total 23680040 KiB used 909673048 KiB avail 2.41% +``` + +Take it again after; the delta should be ≈ 17 G. (`c11-scratch` and `felhom-backup` are two `dir` +entries over the **same** path `/mnt/nvme-1tb`, so they move together — do not read that as double +counting.) + +## The gate that shapes the operation + +**The cascade REFUSES to delete a live host** — there is no hub decommission endpoint; the word appears +only in the refusal. So VM 325 must be **stopped first** (guarded on `qm config 325` reading +`name: walk5-appliance` — `demo-hp` also carries a guest 9201, and the appliance carries its own) and +the hub allowed to age it past its `stale_threshold` (**30 m**, read from the deployed `hub-config`, +not assumed). Poll `GET /hosts/walk5-4bada5/delete-impact` for `deletable` — and **treat an empty +response as retry, not as success**. + +Then `POST /configs/walk5/delete` with all six gates: `ack_hosts=1 ack_reset=1 ack_purge=1 +confirm_id=walk5 expect_hosts=1`. + +## What will NOT be gone, and must be checked rather than assumed — R-244 + +A full-schema census after the cascade will return rows, not zero. Four sources are **deliberate** +(`events`, `notification_log`, `host_deletions`, `customer_resets` — *"provenance/events are NEVER +wiped"*). The fifth is the open gap: **`app_log_issues`**, which the residue purge does not touch. + +**Measured 2026-08-08 across the whole table:** 1309 rows, of which **71 reference a torn-down venue** +(`finalwalk`, `c11`, `rewalk`, `part4`) — **44 orphans** (safely deletable) and **27 shared with a live +customer** (`demo-felhom`, `peti-felhom`, …) which must be **de-referenced, never deleted**. **This +walk will add to that count.** Do not attempt a `DELETE … WHERE customer LIKE` — it would destroy a +live customer's issue history. R-244 carries the proposed shape. + +**A per-table absence query is not a census.** The verification is a full-schema sweep, and that is +what found this. + +## Positive controls the teardown must keep (each must SURVIVE) + +- `qm list` still shows VM **300 `drill-r50`** — the protected drift fixture. +- Guest **9201 on `demo-hp` itself** still running (the host was never the target). +- Storage Box sub-accounts **`u629488-sub1/2/3`** (demo-felhom, peti-felhom, demo-hp) still present. +- `ep0` namespaces **`demo-felhom`** and **`demo-hp`** still present. +- WireGuard peers **10.77.0.2/.3/.4/.250** still on `wg0`. +- Hub rows for `demo-felhom`, `demo-hp`, `peti-felhom`, `david` untouched. + +## One extra item this walk adds + +**The `walk5` retrieval passphrase is compromised** — it reached a session transcript from the customer +page's `data-secret` attribute (**R-249**). It dies with this customer's deletion, which is the reason +the teardown should not be left indefinitely. No other walk5 secret left its `0600` file.