diff --git a/CONTEXT.md b/CONTEXT.md index e77b63c..0629919 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -17,6 +17,33 @@ ## Standing rulings +**S-31 — the off-site key IS recoverable after a real rebuild; USING it is blocked by the remedy that +makes the rebuild survivable (2026-08-04 night drill, R-201/R-204).** + +**Proven on hardware:** demo-hp's controller data volume was destroyed and the sentinel deleted; the +customer's recovery code produced `8a9e33aa4da6…`, byte-identical to the pre-wipe on-disk key and to +the hub's independent record, and installed cleanly on the bare box. `identity_blob` was unchanged +across the wipe — nothing re-escrowed itself. + +**The wall (R-204), four links, all measured:** +1. a rebuilt controller cannot configure its off-site tier — the one-time password was consumed by its + predecessor (`no unconsumed offsite password`, R-193); +2. the Re-issue that fixes that sets `stale_at` **while `restic_pw_sha256` is unchanged** (R-196); +3. a stale escrow makes the hub withhold the hash from the ACK → `EscrowAutoConfirmer` can never flip + `pending → escrowed` → `OffboxRunnable` refuses every run; +4. the only documented way to clear it is a ceremony, **which supersedes the identity blob and destroys + the recovered key**. And before any of it, a rebuilt box is **unclaimed**, so the claim gate + intercepts every controller endpoint — a step in no design document. + +*Facts a future session needs:* +- **A guest rebuild in this fleet is a controller-DATA-VOLUME loss, not a guest reprovision.** The + 2026-08-03 incident R-193 is filed against ran with guest 9201 up throughout — no `pct destroy`, no + `pct restore`, no `--selftest=provision`. Reproduce it that way. +- **A good snapshot is not durable against a later bad run on the same day.** `forget --keep-daily 7 + --group-by host,tags` keeps one per tag per day; a later, worse snapshot evicts a good one. +- **Never run a ceremony while a recovery is in flight** — it supersedes the identity blob. Under + v0.93.0 the old blob is retained, but nothing serves a superseded blob back (R-199). + **S-30 — the R-201 drill was PREPARED and HALTED BEFORE THE WIPE (2026-08-04). Nothing was wiped.** It stopped at step 4 because the sentinel file was **not in the off-site snapshot** while the run diff --git a/REPORT.md b/REPORT.md index 4e1278a..fbffde5 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,172 +1,127 @@ -# REPORT — R-203: the app and its backup look in the same place, and "ok" means it (2026-08-04) +# REPORT — R-201 night run: the key came back; the verdict did not -**Controller v0.196.0 → v0.197.0**, deployed to demo-hp. No hub change. Nothing deleted, wiped or -moved on any box; the drill was **not** resumed. `demo-felhom` untouched. +**2026-08-04, 21:30–22:40, unattended.** `demo-hp` was deliberately rebuilt. No code, no version bump. +`demo-felhom` untouched. Full record: `documentation/audits/DRILL-r201-night-run-2026-08-04.md`. --- -## 1. Confirmed baselines, and §3's landmarks +## 1. THE VERDICT — not reached -`felhom-controller` `532f5712a891` (v0.196.0) → **v0.197.0**; `felhom.eu` `a0c4b607a6cf`, no bump. -Both trees clean on arrival. **Every §3 landmark held**, including the one that decided the fix: -`paths.go:26` names the system-drive arrangement *"the SSD-only system-data fallback"* — **it is -supported**, so the resolution was what was wrong and no API refusal was added. - -## 2. The call sites found — **FIVE, not four** - -| # | Site | Named by the spec? | -|---|---|---| -| 1 | `stacks/deploy.go` `withPathVars` → `${USERDATA_PATH}` | yes — the live defect | -| 2 | `appexport/fabplan.go` | yes | -| 3 | `appexport/export.go` | yes | -| 4 | `stacks/delete.go` `ExportDataMounts` | yes | -| 5 | **`web/handlers.go` FileBrowser mount builder** | **NO** | - -The fifth is the customer's own file browser: on a non-enrolled path it would have mounted the wrong -directory. **Latent, not live** — the system drive is deliberately never a registered `StoragePath`, so -the resolver is the identity there today. Wired anyway, with that reason in the code. - -Two further sites of the same class were found and fixed: `ComputeFabBuckets` was receiving the drive -path where `ComputeCaptureSet` has always received the namespace root, so the export's classified -paths and the backup's capture set could describe different directories for the same declared bind. - -**The rule had TWO existing copies and they differed.** `backup.Manager.namespaceRoot` compared -without `filepath.Clean`; `stacks.Manager.inGuest` compared with it. A trailing slash from config -would have flipped the mode in one package and not the other. Both now delegate to -`appbackup.NamespaceRootFor`. - -## 3. Scenario A — the two paths, before and after (live, demo-hp) - -``` -before: /mnt/sys_drive/userdata/media/books ← app bind; capture set looked elsewhere -after: /mnt/sys_drive/felhom-data/userdata/media/books ← app bind == capture root -``` - -Capture log: **`0 mandatory path(s)`** → **`1 mandatory path(s)`**. - -## 4. Scenario F — the sentinel in the snapshot's file listing - -``` -$ restic ls -l latest --tag calibre-web --rw-r--r-- 1000 1000 181 2026-08-04 12:53:06 - /mnt/sys_drive/felhom-data/userdata/media/books/DRILL-SENTINEL.txt -``` - -And the snapshot's own `paths`: -`["…/backups/primary/calibre-web", "/mnt/sys_drive/felhom-data/userdata/media/books"]`. - -**Not a green status — the file, by name and size.** sha256 -`643166269103a25cf41d34a26b75fd6ebba0a837bbcd7e0d2b5aba7106bcbe7c`, byte-identical to the drill's -step 3 after the fix's migration moved it into the corrected directory. - -## 5. Scenario E — the delete path - -**The §8.3 risk does not exist, and this is the correction owed.** `ExportDataMounts` lives in -`delete.go` but is **export-only**: its single production caller is the `.fab` export adapter and -nothing deletes on its result. The delete path's own guard, `ProtectedHDDPaths`, is **layout-agnostic -by construction** — it protects both `/…` and `/felhom-data/…` — so deletion was never -affected by the namespace-root defect. That note is now in the function's doc comment, because the -file placement misled this change's own specification. - -**Does the corrected path point anywhere the old code did not?** Yes, for the export only, and only on -the system-data fallback: `/felhom-data/userdata` instead of `/userdata`. That is the -directory the app actually binds after this release, and the tests assert the **negative** — no -emitted path lies outside the app's own data roots, on either drive kind. It shipped as its own commit. - -## 6. The verdict value - -**`incomplete`** — **minted**, because `ok` | `error` | `running` contained nothing meaning *"it ran, -and this app is not fully protected"*. **Not `error`:** the rest of the run worked and what was -captured is real, so `SnapshotCount` and the `LastSuccess` anchor still record it. - -It reaches the operator through the **existing** per-run digest, `backup_run_failures` — already -operator-only, already allowlisted. A new event type would have been a two-repo change and the hub -drops anything outside `allowedEventTypes`; the prompt ruled out a hub change. The Hungarian customer -warning is unchanged, and the backups page renders `! Hiányos` with the warn styling. - -## 7. §8.4's narrowing — **no customer-visible warning disappears** - -`TierOffsite`'s `tierKeeps()` already admits `ClassMandatory` only, so an optional path cannot reach -the stat-filter. The added class check is a **no-op today**, written for parity with Tier 2 — and -**demonstrated to be load-bearing anyway**: widening the tier filter alone keeps the tests green -*because of the check*; widening it and removing the check makes an optional gap start reporting. - -## 8. §8.6's live effect — anticipated, and then NOT reproducible - -Anticipated in the CHANGELOG before it could fire: calibre-web on demo-hp had exactly this gap, so its -status would become `incomplete`. **In the event it did not**, because the same session fixed the -underlying path — after the corrected bind and the data migration the app has no gap, and the run is -legitimately `ok`. - -**Attempting to observe the verdict live by hiding the directory did not work, and that is recorded -rather than dressed up:** the running container's bind mount **recreated** it, so `os.Stat` succeeded -and no gap existed. Worth knowing in itself — a bind-mounted directory cannot easily be "missing" -while its app runs, so the mandatory-gap condition arises in practice when the path resolves somewhere -the app never binds (the R-203 case), not when a live app's own directory vanishes. The verdict is -proven by a run-level test that drives the real `RunOffboxBackup`. The fixture was restored and the -sentinel re-verified at the same hash. - -## 9. Tests and every red-proof - -| Scenario | Result | Red-proof — mutation → outcome | -|---|---|---| -| A both paths agree | PASS + **LIVE** | restore the bare-path call → **FAIL**: `/mnt/sys_drive/userdata` vs `/mnt/sys_drive/felhom-data/userdata` | -| B enrolled drive unchanged | PASS | invert the drive-kind comparison → **FAIL** on every enrolled row | -| C mandatory gap → not ok | PASS | unreachable gap recording → **FAIL**; unconditional `ok` → **FAIL** | -| D optional gap → ok | PASS | class check removed **with the tier filter widened** → **FAIL**; tier widened alone → PASS, i.e. the check holds the line | -| E export mounts, both kinds + the negative | PASS | leave the site bare → **FAIL**, emits the short path | -| F sentinel in the listing | **LIVE** | — | - -**A RED-PROOF PASSED AND THE TEST WAS WRONG, NOT THE CODE** (§9.11 — three of the last six sessions). -My first Scenario-C test exercised `offboxCaptureSet` alone, while the mutation lives in -`runOffboxInternal`. A mutation the test cannot observe is not a red-proof. Replaced with a run-level -test that drives `RunOffboxBackup` and asserts `incomplete`, the retained `LastSuccess`, and the -operator signal; it fails under both mutations. The original Scenario-D "red-proof" also could not -fail by construction — recorded above with the two-part mutation that does. - -Green gate after each phase (`go build && go vet && go test ./...`, rc=0) plus -`controller_gates.py --fast`. No test run was combined with a commit. - -## 10. Commits - -| Commit | Contents | +| | | |---|---| -| `73efb09` | Part 1 — one resolver, the four non-export sites, `ComputeFabBuckets`, the two delegating copies | -| `a96c3d9` | **Part 1.3 alone** — the `delete.go`-resident export-mount site, with its scope correction | -| `58c703b` | Part 2 — the verdict, the structural gaps, the operator digest, the status rendering | -| `73fb595` | docs (felhom.eu) | +| sentinel sha256, pre-wipe | `643166269103a25cf41d34a26b75fd6ebba0a837bbcd7e0d2b5aba7106bcbe7c` | +| sentinel sha256, post-restore | **not obtained** — the restore was never reachable | -## 11. Registers +**Not a FAIL.** Nothing came back wrong and no fresh history was started; the box never got as far as +running a backup. **What it is instead:** the first proof that the key survives and returns, plus the +measured reason a customer still cannot use it. -- **R-203 → SHIPPED + PROVEN-LIVE.** -- **R-201 → READY TO RESUME**, blocker gone, fixture staged and verified. -- **R-202 stays open**; the orphaned-ciphertext deletion is **still owed**. -- Capability map: new row for mandatory-path capture on both layouts, its predecessor corrected as - optimistic, and the **R-199 back-pointer omitted two sessions ago is now added**. -- The v0.93.0 `identity_blob` retention remains **unit-proven only** — nothing here superseded a key. +> **After a real rebuild — controller data volume destroyed, sentinel deleted from disk — the +> customer's recovery code produced `8a9e33aa4da6769c5aea1831f87759e10930e2ec1dea0062576484e0598d080a`, +> byte-identical to the pre-wipe on-disk key and to the hub's independent record.** That has never been +> shown before. -## 12. CI +## 2. Snapshot count at step 9 — not obtained -Run numbers and task ids in the session summary. **`--no-verify` was not used.** +The off-site run was never permitted to start (§4). **And a count would have been a poor +discriminator anyway:** restic's same-day `forget --keep-daily 7 --group-by host,tags` keeps one +snapshot per tag per day, so a successful reopen would have shown 3, not 4. The real observable is +whether the pre-wipe snapshot `e6132ae5` survives with the sentinel in it — recorded as the resume +step. -## 13. Teardown +## 3. §5's five conditions, recorded before the wipe -**Nothing provisioned.** `calibre-web` and the sentinel stay — R-201 needs them. The temporarily -hidden directory was restored and the sentinel re-verified byte-identical. +1. **sentinel listed by name** — snapshot `e6132ae5` (19:36:26), `-rw-r--r-- 1000 1000 181 …/DRILL-SENTINEL.txt`. +2. **rollback archive verified** — `vzdump-lxc-9201-2026_08_04-21_44_24.tar.zst`, 1 606 765 083 B, + **full zstd stream read OK** (4 867 573 760 B), sentinel confirmed inside it. +3. §3's option — §5 below. +4. **demo-felhom** `health=ok`, `escrow_state=escrowed`, untouched throughout. +5. **space** — nvme-1tb 883 GB free, root 20 GB, local-lvm 34.32 %. -## 14. Observations — noticed, NOT acted on +**One precondition had drifted and was repaired, not worked around.** The staged snapshot no longer +held the sentinel: the afternoon's `mandatory data path missing` experiment produced a *later* same-day +calibre-web snapshot and `forget` had pruned the good one. The fixture on disk was correct, so one +backup re-established it and it was re-verified by listing. **Lesson: a good snapshot is not durable +against a later bad run on the same day.** -1. **`resolveAbs` uses ONE root for two bind classes.** `RootHDD` resolves against the same parameter - as `RootUserdata`. On an enrolled drive they coincide; on the system-data fallback a `${HDD_PATH}` - bind resolves under the namespace root while compose binds it bare. Both callers now pass the - namespace root, so the export and the backup **agree with each other** — but whether `${HDD_PATH}` - itself should mean the namespace root on the system drive is a separate question that touches every - already-deployed app's binds. **Not fixed here; it needs a decision, not a patch.** -2. **The blast radius was measured before changing anything:** exactly one app in the fleet has - `HDD_PATH == system_data_path` — `calibre-web` on demo-hp, deployed for the drill. Every other - deployed app has no `HDD_PATH`. Nothing else needed migrating. -3. **`/api/stacks//deploy` is first-deployment-only** (409 afterwards), so the corrected - `${USERDATA_PATH}` reaches an existing app through a start/redeploy, not a re-deploy. -4. **The controller's CSRF form field is `_csrf`, not `csrf_token`** — the login page uses one name and - the protected forms another. Second session in a row this cost time; it belongs in the - headless-access memory. +## 4. Every step's observable + +| step | observable | +|---|---| +| 6 | fresh data dir stamped `20:00:2x`; **new `encryption.key`** (32 B); `claimed = None`; `offbox = null` | +| **7** | **`identity_blob` 572 B, `restic_pw_sha256` `8a9e33aa4da6…`, `updated_at` still `11:11:37` — UNCHANGED across the wipe.** Nothing re-escrowed itself | +| **8** | **recovered `8a9e33aa4da6…`** — matches pre-wipe on-disk and the hub's record. Exit 1 is correct and designed: nothing local to compare against, the rebuilt-box shape | +| 9a | `[INSTALLED] … reads back identical` — the "installed" branch's first real run | +| 9b | after the Re-issue and apply, the on-disk key is **still** `8a9e33aa4da6…` — `WriteOffboxSecrets` kept it | +| 9c | **blocked** — see below | +| 10–11 | **not run** | + +**The wipe was faithful to the incident, deliberately.** The 2026-08-03 rebuild R-193 is filed against +was **not** a guest reprovision — the journal shows guest 9201 running continuously through that window +with no `pct destroy`, no `pct restore` and no `--selftest=provision`. What changed was the controller +and its data volume. A guest destroy plus an unrehearsed provisioning chain, improvised unattended, is +what §8.10 exists to prevent. + +## 5. §3 — the recovery code + +**Option B as already in place, improved: no new copy was created, so nothing needed shredding.** The +operator had placed `R_DEMO-HP` in their own `~/.config/credentials` (`0600`) two sessions ago for this +purpose. It was read from there and **piped to stdin** for the two invocations that needed it — never +an argument, never exported, never written to a second file, never logged. Destroying the operator's +own store would have destroyed their record; because no additional copy existed, there is nothing left +to prove gone. + +**Verified anyway, with the planted-copy positive control:** 0 hits in the agent journal, 0 in the +controller log, 0 files under `/tmp`, `/var/tmp`, `/var/lib/felhom-agent`, `/root`, 0 leftover +`felhom-idesc-*` staging dirs — and the same sweep found a planted copy (**1**), then **0** after +shredding it. The instrument is shown sensitive rather than assumed to be. + +## 6. Part 2 — did not run + +Its gate is "the drill PASSED". It did not. A second wipe would have destroyed the state that makes the +first drill finishable in five minutes. **R-198's retention therefore remains unit-proven only** — +nothing has yet superseded a key in production. + +## 7. Teardown — three layers + +| layer | state | +|---|---| +| the guest | **nothing torn down** — the wipe is the evidence; all six app containers serving; the recovered key on disk; scratch band empty | +| the host | vzdump snapshot LVs released cleanly; one new 1.6 GB archive on nvme-1tb (883 GB free); nothing deleted | +| the hub | **no new customer records** — `demo-hp` is the same row throughout, because the rebuild was a controller-data wipe and not a re-enrolment. One Re-issue staged a one-time secret (consumed 20:16) and set `stale_at`. The two superseded escrow rows are **unchanged** — no ceremony ran | + +**Nothing deleted on the storage endpoint**, including the ~1.2 GB of orphaned ciphertext — ruled, +still owed, deliberately not ridden along with a drill. + +## 8. The capability-map rows + +A new row records the key as **PROVEN-LIVE after a real rebuild**, and states plainly what it does not +claim: **no file has been restored**, and the four links of R-204 stand between the recovered key and a +usable repository. The R-199 back-pointer added earlier today stands. + +## 9. New findings + +- **R-204 (NEW)** — the recovered key cannot be used: a rebuilt box cannot configure its tier (R-193); + the Re-issue that fixes that marks the escrow stale though the key never changed (R-196, **measured + at `20:15:49` with `restic_pw_sha256` unchanged**); a stale escrow gates every run; and the only way + to clear it is a ceremony that **destroys the recovered key**. Plus a fourth link in no design + document: a rebuilt box is **unclaimed**, so the claim gate intercepts every controller endpoint. +- **R-196 re-scoped** — no longer a documentation nit; it is on the critical path for recovery. +- **R-201** — the wipe happened, the key came back, ~5 minutes from a verdict with a person present. +- **R-198** — still unit-proven only; Part 2 gated out. +- **R-202 open; the ciphertext deletion still owed.** + +## 10. CI + +Docs push only. Run number and task id in the session summary. **`--no-verify` not used.** + +## 11. Observations — noticed, NOT acted on + +1. **The claim gate is an undocumented first step of every recovery.** Before a customer can do + anything on a rebuilt box — including recovering their backups — they must re-claim it. +2. **`--print-reset-code` output needs parsing care**: the captured value was 73 characters, i.e. more + than the code itself. Not chased; the claim was abandoned when the session stopped. +3. **A same-day re-run replaces the day's snapshot.** Worth knowing before designing any drill that + depends on a specific snapshot surviving. +4. **The hub's ClusterIP is not reachable from DooPlex's host network** — the operator UI needs a + `kubectl port-forward`. The `curl -u :$HUB_PW` recipe in memory omits that. diff --git a/STATUS.md b/STATUS.md index 80d69f0..1d90a67 100644 --- a/STATUS.md +++ b/STATUS.md @@ -36,25 +36,21 @@ Proven end to end on real hardware. the old backups stay recoverable. Both keys are now kept. **What this cannot undo:** anything superseded before today is gone for good, which includes both demo machines' pre-4-August keys — so those 51 orphaned backups were beyond reach even if the recovery codes had been kept. *(R-198)* -- *(largely fixed 4 Aug)* ~~Nothing in the recovery path has ever been performed.~~ **The key now - demonstrably comes back — measured on a real machine this evening.** demo-felhom fetched its own - sealed package from the hub with its own credential, opened it with the recovery code you saved, and - the backup key that came out was **identical, character for character, to the one the machine is - using** — and to the fingerprint the hub had recorded separately. Three independent sources agreeing. - A deliberately wrong code, tried five minutes earlier, was refused outright and wrote nothing. - **What is still NOT done, and it is the half that matters to a customer:** nothing yet *puts the - recovered key back*, reopens the old backup store with it, or **restores a single file**. Today - proves the key survives and returns; it does not prove the backups do. *(R-199 closed; R-200 half; - R-201 open)* -- *(fixed 4 Aug)* ~~A backup reported success while leaving out a folder the customer was told is - protected.~~ **Both halves fixed the same day.** The app was writing to one folder and the backup was - looking in another — one directory apart, on machines whose apps live on the system disk. They now - resolve to the same place, from a single piece of code instead of the three near-copies that had - quietly drifted. **And a backup that cannot capture a folder marked essential no longer reports - success**: it reports *Hiányos* (incomplete), names the app and the folders, and tells you — while - still recording what it genuinely did capture, because half a backup is not no backup. Proved on the - HP machine by listing the backup's own contents and finding the marked file there by name and size — - not by trusting a green tick. *(R-203)* +- **THE BACKUP KEY COMES BACK AFTER A REBUILD — proved on real hardware last night.** We destroyed the + HP machine's controller data on purpose, deleted the marked file from its disk, and then used the + recovery code you saved. The key that came out was **identical, character for character**, to the one + the machine had been using — and to the fingerprint the hub had recorded separately. It installed + cleanly onto the empty machine. Nothing re-sealed itself in the meantime, so the sealed package + survived the rebuild untouched. *(R-201)* +- **But the machine still could not use it, and that is the night's real finding.** Three things stand + between a recovered key and a restored file, and each one is now measured rather than guessed: + **(1)** a rebuilt machine cannot set up its off-site connection at all — its one-time password was + spent by the machine it replaced; **(2)** the fix for that (re-issuing the credential) marks the + sealed package "stale" even though the key never changed, and a stale package blocks every off-site + backup; **(3)** the only documented way to clear that is to make a new recovery code — **which + replaces the sealed package and destroys the key we just recovered.** And before any of it, the + rebuilt machine is unclaimed, so nothing on its dashboard responds until the customer claims it + again — a step written down nowhere. **The key comes back and cannot be used.** *(R-204)* - **One screen still tells the customer something we cannot yet promise.** The „elárvult tároló" card says the old backups may later be restorable with the matching recovery code. From today that is true for machines that re-seal from now on and **false for anything already orphaned** — and the machine @@ -103,7 +99,10 @@ Proven end to end on real hardware. ## What we're working on -- **Now:** the wipe-and-restore proof, which is **unblocked and staged**. The marked file now lands in +- **Now:** finishing the proof — it is about five minutes with you present. The HP machine is sitting + mid-drill with the recovered key already on it; it needs re-claiming and one setting confirmed, then + the backup runs and the file is restored. **Do not let it make a new recovery code** — that would + destroy the key we recovered. Then the deeper fix for the wall above. The marked file now lands in the off-site backup, so there is finally something to recover. Everything else is already in place on the HP machine — the recovery code you saved, a working off-site store, the app and the file. Nothing has been wiped; that step waits for your go-ahead at a marked stop point. Both honesty fixes shipped today: a changed backup key now @@ -158,7 +157,12 @@ Proven end to end on real hardware. ## Changed since last update -- **2026-08-04 (latest)** — **Fixed the folder-left-out-of-the-backup problem, both halves.** The app +- **2026-08-04 (night)** — **Wiped the HP machine on purpose and got the backup key back with your + recovery code — identical, character for character.** First time that has ever been done. The drill + then stopped short of restoring the file, at a wall worth more than the last step: the key comes back + but cannot be used without an act that destroys it. The machine is up, its apps are serving, and it + is five minutes from finished. *(R-201, R-204)* +- **2026-08-04 (evening)** — **Fixed the folder-left-out-of-the-backup problem, both halves.** The app and its backup now look in the same directory, and a backup that misses a folder marked essential reports *incomplete* instead of success. Proved by listing the backup's own contents and finding the marked file. The wipe-and-restore proof is unblocked. *(R-203, R-201)* diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index d5121fe..d0debbc 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -40,6 +40,7 @@ | **The offsite repository password can be RECOVERED from the sealed escrow with the customer's recovery code** | hub v0.94.0, agent v0.125.0, controller v0.195.0 | **PROVEN-LIVE (2026-08-04)** | On demo-felhom, through the real endpoints end to end: the box fetched its own sealed blob from the hub with its own per-host credential (hub log: *escrow blob SERVED … 572 opaque bytes, self_scope=true*), the agent unsealed it with the customer's recovery code, and the extracted repository password's sha256 was **byte-identical** to the one on disk — `c60c8bc737a6…`, which is ALSO the hash the hub had independently stored, so three sources agree. Five minutes earlier the same path with a WRONG code failed closed at age's KDF with nothing written, which proves links 6 and 7 ran independently of the success. R was searched for afterwards and found in 0 log lines and 0 files, with a positive control confirming the search would have found it. Evidence: per-repo CHANGELOGs; `audits/RECON-offsite-dr-chain-2026-08-04.md` §3 links 6–8 | **WHAT THIS ROW DOES NOT CLAIM, stated because the previous over-claim here was struck out four hours earlier.** It covers the KEY, not the DATA. **R-199 BACK-POINTER (omitted when this row was written): the recovery chain's link inventory and the per-link status live in `audits/RECON-offsite-dr-chain-2026-08-04.md` §3; links 1–8 are walked, 9–11 are not.** **Re-confirmed 2026-08-04 evening by an attempt to prove the DATA half:** the R-201 drill was prepared on demo-hp and **halted before the wipe** — the sentinel file was not in the off-site snapshot (R-203), so the wipe would have destroyed it and proven nothing. **No file has still ever been restored from an off-site backup after a wipe** (`audits/DRILL-r201-offsite-recovery-2026-08-04.md`). The install half of the chain (controller v0.196.0 `--recover-offsite-install`, R-200) is likewise unit-proven only — it has never run against a live recovery. A recovered password has never been **installed** (the diagnostic compares and refuses to write, by design), no existing repository has ever been **reopened** under one, and **no file has ever been restored** from an off-site history via a recovered key. Links 9–11 of the chain are open (R-200's remaining half, R-201). The proof also used a box whose local key still exists — the rebuilt-box case, where there is nothing to compare against, is exactly what the drill covers and it has not run || DR tier by default: PBS + WireGuard base infra on every install, hub-controlled activation | installer v1.15, agent v0.86, hub v0.51 | **PROVEN-LIVE** (2026-07-21) | `DRILL-day0-take2-2026-07-12` §2 (WG enabled both modes, PBS-DR descriptor auto-provisioned ~1s after WG registration, zero operator steps); ships installer v1.15/agent v0.86/hub v0.51 | Live only on demo/drill fleet. (Cited spike was slice-0 mechanics — shipped nothing; corrected.) **⚠ The candidate upgrade to PROVEN-LIVE is WITHDRAWN — the 2026-07-18 rehearsal produced a live counter-example (R-39).** On the reborn N100 the descriptor auto-provisioned and the agent reported `converged state=applied` (16:45:53), yet **the storage is dead**: `pvesm status` → `felhom-pbs: error fetching datastores - 401 Unauthorized` / `inactive`, and a direct probe with the stored credential returns **401 on every endpoint including `/version`** while the WG transport is healthy (handshake 9 s, 27.9 ms RTT) — i.e. authentication failure, not ACL scope. Root cause in the evidence: **the hub minted a SECOND token secret at 16:47:52, two minutes after the agent had applied the first, and `consumed_at` is still NULL**; the converged state machine will not re-apply, and the agent's 15-minute verify loop **cannot even read the credential to notice** (`open /etc/pve/priv/storage/felhom-pbs.pw: permission denied` — non-root agent reading a file it writes through a root wrapper). A tier that reports `applied` while silently unable to authenticate is exactly the shape that must not carry a PROVEN-LIVE badge. See `tests/VALIDATION-n100-rehearsal-2026-07-18.md` F2 and `pbs-dr-state.txt`. **agent v0.89.0 closes the F4 non-default-storage-id gap (R-22) — PROVEN-LIVE 2026-07-17:** the reconcile self-grants the ACL through the root wrapper on a pre-check 403 instead of dead-locking. Reproduced F4 on the demo (marker moved aside = reinstall fresh-state + felhom-offsite ACLs revoked) → next reconcile tick `pbsdr: pre-check 403 … self-granting … (R-22)` → `converged state=adopted` in ~3 s, ACLs self-restored, `pvesm status felhom-offsite`=active, zero operator action. No more one-shot `pveum` grant **2026-07-21 — the R-39 fleet fix SHIPPED (hub v0.68.0 + agent v0.91.2), closing the self-heal chain end to end.** The three defects that let a box be `applied` and dead simultaneously are each addressed: the hub stamps a monotonic `secret_generation` into the descriptor so a credential re-key finally MOVES the content hash the agent re-applies on; the wrapper gains a narrow `read` verb so the non-root agent can read the credential it writes (it never could — `/etc/pve/priv` is 0700 root:www-data, which made the verify loop blind by construction); and `pbs.ProbeAuth` turns a 401 into a loud `auth_failed` that the existing `pbsdrheal` damper escalates to a fresh mint. Plus a consumed_at honesty gauge for the disagreement no single tier can see (box says `applied`, hub's staged secret never consumed). Proven live on felhom-pve: the agent read its credential through the wrapper (`rc=0`) and probed successfully (`credential probe OK storage=felhom-pbs`). **STOP-2 RAN 2026-07-21 AND THE CHAIN CLOSED — 13 SECONDS, operator click to converged.** The operator pressed **Re-issue PBS credentials**; the identical click on 2026-07-18 did nothing at all. Full chain (hub UTC / host CEST = UTC+2): `08:39:31Z` hub mints a fresh secret, **generation 0 → 1**, and the descriptor gains `"secret_generation": 1` — with `token_id` and `fingerprint` **byte-identical**, i.e. exactly the re-key shape that used to be invisible → `10:39:34` the agent READS its credential through the wrapper (leg b — the read that was impossible until v0.91.0) → `10:39:38` **`ERROR pbsdr: the DR endpoint REJECTED this box's credential — the tier is applied and DEAD` `previous_state=applied`** (leg c: the exact R-39 failure state, detected out loud for the first time ever) → `10:39:45` **`one-time token secret consumed`** `secret_len=36` (leg a: **NO short-circuit** — this is the line that never appeared on 2026-07-18) → `10:39:45` `felhom-pbs-apply reconcile` (the set-only wrapper, no `--server`) → `10:39:47` **`pbsdr: converged state=applied`**. Corroboration: the agent marker hash moved to `afbb3b41…` (it was byte-identical to the pre-reissue marker in the failure); `consumed_at` stamped `08:39:45Z`; the on-disk secret's mtime moved `2026-07-18 20:28:52` → `2026-07-21 10:39:45`; a live probe with the NEW credential returns **200**; three consecutive hub reports trace the whole state machine `applied → auth_failed → applied`; and **zero** `pbsdr_selfheal` escalations fired — the box healed through the descriptor path before the damper was ever needed, with exactly ONE mint and ONE consume and no `consumed-failed.json`. **Row upgraded to PROVEN-LIVE (2026-07-21).** Evidence: `felhom-agent/REPORT.md` (2026-07-21). | +| **The offsite repository key is recoverable AFTER A REAL REBUILD — and cannot yet be USED without an operator act that destroys it** | hub v0.94.0, agent v0.125.0, controller v0.197.0 | **PROVEN-LIVE for the KEY (2026-08-04 night); the RESTORE is still unproven** | demo-hp's controller data volume was destroyed and the sentinel deleted from disk. The customer's recovery code then produced `8a9e33aa4da6…` — **byte-identical** to the pre-wipe on-disk key and to the hub's independent record — and installed cleanly on the bare box. `identity_blob` was unchanged across the wipe (572 B, `updated_at` still 11:11:37). Evidence: `audits/DRILL-r201-night-run-2026-08-04.md` §2 | **WHAT IT DOES NOT CLAIM, and the gap is the point.** No file has been restored: the drill stopped at a wall (**R-204**). A rebuilt box (a) cannot configure its off-site tier — the one-time password was consumed by its predecessor (R-193); (b) the Re-issue that fixes that marks the escrow stale although the key is unchanged (R-196), which gates every off-site run; (c) the only documented way to clear that is a ceremony, **which supersedes the identity blob and destroys the recovered key**; and (d) the box is unclaimed, so no controller endpoint responds until the customer re-claims — a step in no design document. **The key comes back; the repository still cannot be opened with it.** | | **Off-site app-data capture covers the paths an app declares MANDATORY — on BOTH drive layouts** | controller v0.197.0 | **PROVEN-LIVE (2026-08-04) — and this row was OPTIMISTIC before it** | On demo-hp, an app on the **system-data fallback** bound `/mnt/sys_drive/userdata/media/books` while the capture set looked in `/mnt/sys_drive/felhom-data/userdata/media/books`: the declared-mandatory directory was in **no** snapshot and the run reported `ok` (R-203). After v0.197.0 the capture log reads `1 mandatory path(s)` and the file is **listed inside the snapshot** — `restic ls -l latest --tag calibre-web` → `-rw-r--r-- 1000 1000 181 … /DRILL-SENTINEL.txt`. Evidence: `audits/DRILL-r201-offsite-recovery-2026-08-04.md` §2 and the v0.197.0 CHANGELOG | **WHAT IT DOES NOT CLAIM.** It covers CAPTURE, not RESTORE: **no file has ever been restored from an off-site backup after a wipe** (R-201, ready to resume). It also does not claim the enrolled-drive layout was ever wrong — it was not, which is precisely why this survived. And a run that still misses a mandatory directory now reports **`incomplete`** rather than `ok`, so this row's guarantee is one the status can express || **Recurring offsite (PBS) whole-guest backups actually LAND, and RESTORE** — local daily + offsite weekly as scheduled work | agent v0.97–0.103, controller v0.174/0.175, hub v0.76.0, host-install 1.20.0 | **PROVEN-LIVE (2026-07-26)** | `audits/SPIKE-r82-phase0-2026-07-26.md`; per-repo CHANGELOGs/REPORTs. **Restore round-trip on demo-hp:** `--selftest=restore-test` against `felhom-pbs:backup/ct/9201/2026-07-26T15:42:42Z` → `pass:true`, `verified:"boot+running"`, **`mount_parity:"ok"`** (`mp0=/var/lib/docker 50G`, `mp1=/mnt/sys_drive 20G`, mp8/mp9 throwaway stand-ins for the archived binds), `source_tier:"pbs"`, 4m5s restore+boot+verify+teardown, scratch band clean afterwards and the live guest untouched. | **This row is distinct from the DR-tier row above, which proves ACTIVATION, not ARRIVAL.** That tier was PROVEN-LIVE as `applied` since 2026-07-21 while demo-felhom held ONE snapshot (a healing artifact) and demo-hp held **zero, ever** — "applied and empty", the R-39 shape one level quieter. **What earns PROVEN-LIVE here:** (a) demo-hp's FIRST EVER offsite backup landed (4.25 GB into a verifiably empty namespace); (b) it **restores into a bootable, mount-complete guest** — `mount_parity` is the non-hollow half, since a boot-only verify cannot see a missing data volume; (c) **the multi-tier quiesce ran through the real UI endpoint** (`POST /api/guest-backup/trigger`, authed+CSRF) and produced **exactly ONE stop/start pair with BOTH backups inside it** — `quiescing 1 stack(s)` 17:01:39 → local done 17:02:56 *"next tier may start (app still quiesced)"* → felhom-pbs snapshotted 17:03:06 → `unquiescing` 17:03:06. **App downtime 1m27s for both tiers**, and the app came back healthy. **Known gaps, recorded not hidden:** the SCHEDULED restore-test still only selects the PRIMARY tier, so the offsite tier is never AUTOMATICALLY restore-tested (the manual/selftest path is proven, the unattended one is not); the hub infers "PBS ⇒ weekly" from storage TYPE rather than a reported cadence; the installer-default fleet flip awaits a full weekly cycle. → **R-82** | | **Restore-proof is UNATTENDED — the scheduler covers EVERY tier, follows the BACKUP rather than the clock, and a failure is heard** | agent v0.104.0 → **v0.121.0**, hub v0.77.0 → **v0.91.0** | **PROVEN-LIVE (2026-08-03)** | per-repo CHANGELOGs; `backlog/SPEC-r85-phase4-5-2026-07-26.md`. Unit red-proofs for tier rotation, restart-survival, the one-heavy-op gate, failure-emits-an-event, and newborn silence. | **The row above is earned by a MANUAL `--selftest=restore-test`; this one is about the SCHEDULED path, and the distinction is the whole point.** Before R-85 the scheduler could only ever see `cfg.Backup.BackupTarget()`, so the offsite tier was never a candidate — and a failed restore-test was a `[WARN]` line with no event at all, which was true for the LOCAL tier that WAS being tested. Now: oldest-first rotation across every configured tier (operator ruling 2026-07-26), persisted so it survives a restart; a restore-test joins the host-wide one-heavy-operation gate so it never contends with a backup over the same tunnel; and the hub raises two DISTINCT operator-tier signals — `restore_test_failed` (broken now) and `restore_test_stale` (unverified, not known-broken), anchored on R-81 so a newborn box never alarms. **Why this was NOT PROVEN-LIVE until now:** rotation had not been observed selecting both tiers across consecutive UNATTENDED cadences — at a 24h cadence a multi-day window — and a single passing run proves the code path, not the schedule. → **R-85**. **R-86 (agent v0.121.0 + hub v0.91.0, 2026-08-03) replaced the schedule and the observation became possible in one afternoon**, because what has to be observed is no longer a multi-day rotation but a RULE: a tier is due when its newest archive that has settled ~24 h has not been proven. **UPGRADE EVIDENCE — a real unattended run on demo-felhom (2026-08-03), triggered by DUE-NESS, not by a timer.** The scheduler's own log: `15:14:38 restore-test tier is DUE … target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z … reason="newest settled archive … has not been proven"` → `proxmox-backup-client restore --crypt-mode=encrypt` under the agent's own token → `15:25:08 gate decision class=guest_destroy guest=990000 allowed=true` → `15:25:14 scratch guest torn down` → **`15:25:14 backup: scheduled restore-test passed archive=felhom-pbs:… duration_s=635.1`**. A **14.5 GB encrypted offsite archive pulled from ep0 over the WAN**, restored, booted, verified and destroyed in 635 s, unattended. **The three things a timer could not show**, all verified after it: the state names THAT archive (`{"felhom-pbs":{"archive":"…2026-07-28T04:49:43Z","proven_at":"2026-08-03T13:25:14Z"}}`); a second evaluation reports `due=false … is already proven` and runs nothing; and **an agent restart runs nothing**, which is the defect a person actually noticed — every deploy used to restart the timer. Teardown verified at all three layers: guest absent from `pct list`, **zero** `990000` volumes in `lvs`, and the hub-side `restore_tests[]` entry deliberately RETAINED (it IS the proof the staleness check reads). **R-189 (agent v0.122.0, 2026-08-03) closed the reporting half of this row, and it was a REAL gap in the evidence path:** the proof above reached the hub only because no restart intervened — `restore_tests[]` came solely from an in-memory store, so the 15:25:14 PASS was in fact LOST when the agent restarted 2 m 43 s later for a deploy (`0 restore-tests` on the next two host-reports). Under per-archive due-ness the box would not have repeated the work for a week. The persisted per-tier proof (with the archive, and now the tier) is merged into the report, so a proof survives a restart — one entry per tier, newest wins, and a record that cannot be described honestly is not emitted. **THE HOST TIER'S HALF OF THIS ROW WAS OPTIMISTIC UNTIL 2026-08-03, and it is worth saying plainly.** Every live restore-test cited here is on the OFFSITE tier. The HOST tier was not merely unproven on demo-felhom — it was **unprovable**, and on demo-hp equally: the agent's token had no ACL on `/storage/felhom-backup`, the storage both boxes configure as `local_backup_target`, so the content API answered `{"data":[]}` through the token while root listed three archives. The scheduler skipped it as *"no settled archive yet"* — which is exactly what a brand-new tier reports — so nothing ever said so (**R-185**). **CLOSED 2026-08-03 — agent v0.123.0 + installer 1.24.0.** The grant is applied on both demo boxes (the token now lists 3 and 4 archives), the installer's reuse arm grants on a pre-existing target, and the agent now **asks whether it may read each tier it depends on** instead of inferring it from an empty list: a critical degraded capability naming the storage and the missing role, which the hub alerted and emailed on **while the box was still blind**. **THE HOST TIER IS NOW PROVEN-LIVE, UNATTENDED, ON BOTH DEMO BOXES (2026-08-04) — this row's optimistic half is cashed.** Four SCHEDULED runs, nothing triggered by hand: **demo-felhom** host tier `…2026_08_02-04_42_14.tar.zst` **passed in 83.8 s** at 00:55, offsite `…2026-07-28T04:49:43Z` **passed in 540.4 s** at 06:55; **demo-hp** host tier `…2026_08_02-04_49_29.tar.zst` **passed in 109.3 s** at 02:05, offsite `…2026-07-28T19:19:45Z` **passed in 300.1 s** at 08:05. Each restored into a scratch guest, booted, verified and destroyed itself; `pct list` and `lvs` show zero `990000` afterwards on both, and both boxes' `local-lvm` returned to their pre-run figures (1.95 % and 30.83 %). **Both boxes had BOTH tiers due at once and the ordering was observed live for the first time under R-86's rule:** never-proven sorted first, so each box took its HOST tier, deferred the offsite one, and picked it up on the following evaluation six hours later — one heavy operation at a time, per box, without anyone sequencing it. **The proofs reached the hub**, which is R-189 carrying a host-tier entry for the first time: demo-felhom's report holds **two** entries, one per tier, and the `local` one can only have come from the persisted state because the in-memory store held only that morning's offsite run. **SCOPE, stated because one box proving something does not make it a fleet property:** this covers `demo-felhom` and `demo-hp`. The tester's box is untested and untouched. Evidence: `felhom-agent/REPORT.md`, `felhom.eu/REPORT.md` | | Customer RESET (middle lifecycle tier: host delete < RESET < customer Delete): one operator action → pre-first-install; all operational state destroyed, identity + basic config survive | hub v0.61.0, felhom-tenantsync v1.1.0 | **PROVEN-LIVE (external teardown, incl. two real firings)** | **`tests/VALIDATION-n100-rehearsal-2026-07-18.md` — two live firings, both host-delete-first, on two different customers** (`demo-vm-felhom` 15:49:57, `demo-felhom` 16:08:51): every leg `ok` (`claim`, `db_purge`, `descriptor`, `hetzner`, `pbs`), escrow acked separately, each completing in 8–9 s (`hub-state.txt` `customer_resets`). The **Hetzner sub-account destruction is now verified against the live pool box** — and produced the run's sharpest lesson: **a sub-account is an access-control object, not a data object.** Deleting it left its `/home` intact, so re-enabling offsite recreated an account over the previous lifecycle's ciphertext under a key this same RESET had destroyed — which is why the orphan guard fired at 16:58:14 (**a finding by S7's own criterion**) and why RESET now needs a base-dir purge → **R-32**. Prior: hub v0.61.0 REPORT; **ep0 live drill 2026-07-17** (throwaway `drill-reset-01` with a real backup: deprovision `deleted:true` destroyed the namespace + backup group + token, idempotent re-run `deleted:false`, all 3 real tenants + shared user survived); red-proofs (ack-gate, partial-failure resumability) + orchestration/store/offsite/render tests | External teardown FIRST, DB purge LAST, every leg idempotent; refuses while any host row exists; separate escrow-custody ack; clears claim (fresh code next onboarding); keeps the offsite tier CHOICE, drops provisioned fields. **Live-clicked 2026-07-18** (twice, by Viktor) — this supersedes the earlier "not live-clicked / Hetzner delete unit-tested only" note. **R-25b CLOSED (hub v0.69.0, 2026-07-21):** the Danger-zone DELETE is now the guided full-teardown cascade that runs this very sequence as its middle leg — see the row below | diff --git a/documentation/audits/DRILL-r201-night-run-2026-08-04.md b/documentation/audits/DRILL-r201-night-run-2026-08-04.md new file mode 100644 index 0000000..07ac984 --- /dev/null +++ b/documentation/audits/DRILL-r201-night-run-2026-08-04.md @@ -0,0 +1,264 @@ +# DRILL — R-201, the night run: **THE KEY CAME BACK. The verdict was not reached.** + +**Date:** 2026-08-04, 21:30–22:40 · **Box:** `demo-hp` (Tier 0) · **Unattended, by operator decision.** +**The wipe happened.** The box is up, its apps are serving, and it is left mid-drill by deliberate +choice — see §7 for the exact state and the one command that resumes it. + +> **THE RESULT, first.** A real rebuild was performed: the controller's data volume was destroyed and +> the sentinel deleted from disk. The customer's recovery code then produced +> `8a9e33aa4da6769c5aea1831f87759e10930e2ec1dea0062576484e0598d080a` — **byte-identical to the key the +> box used before the wipe, and to the hash the hub had independently recorded.** +> +> **The off-site backup key is recoverable after a machine is rebuilt. That has never been shown +> before.** +> +> **And the drill did not finish**, because three separate things stand between a recovered key and a +> restored file. All three are measured below. That is the other half of the night's result, and it is +> the half nobody knew. + +--- + +## 1. The verdict + +**NOT REACHED.** Step 10 (restore the sentinel and compare) was not run. + +| | | +|---|---| +| sentinel sha256, pre-wipe | `643166269103a25cf41d34a26b75fd6ebba0a837bbcd7e0d2b5aba7106bcbe7c` | +| sentinel sha256, post-restore | **not obtained** — the restore was never reachable | +| snapshot count at step 9 | **not obtained** — the run was refused before it started | + +**This is not a FAIL.** A fail is "the file came back wrong" or "a fresh empty history was started". +Neither happened, because the box never got as far as running a backup. What the night established is +where the wall is. + +--- + +## 2. What was proven, in order, on hardware + +### Step 6 — the wipe + +The controller's data volume was destroyed and the bootstrap re-ran. Fresh data directory, everything +stamped `2026-08-04 20:00:2x`: + +``` +encryption.key 32 B 2026-08-04 20:00:21 ← brand new: every pre-wipe app secret is undecryptable +claimed = None ← the fresh-install signal +offbox = null ← no off-site target +``` + +**Faithful to the incident, and deliberately so.** The 2026-08-03 rebuild that R-193 is filed against +was **not** a guest reprovision — the journal shows guest 9201 running continuously through that +window, with no `pct destroy`, no `pct restore` and no `--selftest=provision`. What changed was the +controller and its data volume. Reproducing *that* is what the drill needs; a guest destroy plus an +unrehearsed provisioning chain, improvised unattended, is what §8.10 exists to prevent. + +### Step 7 — the assertion that keeps recovery possible: **PASSED** + +``` +host_escrow (demo-hp-bb76ea), AFTER the wipe: + identity_blob = 572 bytes ← unchanged + restic_pw_sha256 = 8a9e33aa4da6… ← unchanged + updated_at = 2026-08-04 11:11:37 ← unchanged; nothing re-escrowed + stale_at = NULL +``` + +**No ceremony was run and nothing re-escrowed itself.** The sealed key survived the rebuild untouched. + +### Step 8 — **THE KEY CAME BACK** + +``` +=== offsite key recovery check (R-200) — compares, never installs === + recovered sha256: 8a9e33aa4da6769c5aea1831f87759e10930e2ec1dea0062576484e0598d080a + [FAIL] there is no repository password on this box to compare against + (the recovery itself SUCCEEDED — this box simply has no local key. That is the + rebuilt-box shape, where the next step is to INSTALL rather than compare.) +``` + +Exit 1 is **correct and designed**: there was nothing to compare against, because the wipe removed it. +The recovered hash matches the pre-wipe on-disk key and the hub's own record — three independent +sources agreeing, one of them recovered through the full chain (hub → agent → unseal → extract) on a +box that had just lost everything it knew. + +### Step 9a — installed cleanly + +``` +=== offsite key recovery INSTALL (R-200) === + on-disk sha256: (none — this box has no repository password) + recovered sha256: 8a9e33aa4da6… + [INSTALLED] the recovered repository password is in place and reads back identical. +``` + +The "installed" branch — the rebuilt-box shape v0.196.0 was written for — took its first real run. + +### Step 9b — the apply kept it + +After the target was reconfigured, the on-disk key was still `8a9e33aa4da6…`. `WriteOffboxSecrets` +found the file present and kept it, exactly as documented. + +--- + +## 3. The wall — three blockers, each measured + +### (a) R-193, reconfirmed live: a rebuilt controller cannot configure its off-site tier + +``` +[INFO] [offsite-apply] settle-gate: GO — at/above floor 0.156.0 (we are 0.197.0) +[WARN] [offsite-apply] reconcile: offsite-apply: consume one-time password: + no unconsumed offsite password (already consumed or none provisioned) + (retries on next config refresh/restart) +``` + +The ledger, measured: `demo-hp` one-time secret created `07:11:51`, consumed `07:12:06` — **by the +previous controller**. The rebuilt one has nothing to consume and retries forever. Exactly the state +demo-hp sat in for 25 hours on 2026-08-03. + +**Remedy:** an operator Re-issue. Performed here through the designed endpoint +(`POST /configs/demo-hp/offsite-reissue`, HTTP 303), after which the target configured normally. + +### (b) R-196, measured live — and it lands squarely on the recovery path + +``` +host_escrow (demo-hp-bb76ea), after the Re-issue: + stale_at = 2026-08-04 20:15:49 ← set + restic_pw_sha256 = 8a9e33aa4da6… ← UNCHANGED. The key did not move. +``` + +**The escrow was marked stale while it perfectly covers the box's current key** — the recovered one. +That is R-196's false staleness, and here it is not a cosmetic lie: a stale escrow makes the hub +**withhold the hash from the report ACK**, so `EscrowAutoConfirmer` can never flip `pending → +escrowed`, and `OffboxRunnable` (`configured && escrowed`) **refuses to run any off-site backup**. + +> **The remedy for (a) disables the recovery it was needed for.** The only documented way to clear a +> stale escrow is a fresh ceremony — which supersedes the identity blob, i.e. **destroys the very key +> being recovered**. Under hub v0.93.0 the old blob is now *retained*, but nothing can serve a +> superseded blob back (R-199's inventory: the endpoint serves the CURRENT row only). + +**Superseded rows still number 2** — no ceremony was run tonight. The key is intact. + +### (c) The claim gate — undocumented as a recovery step + +``` +[DEBUG] [web] claim gate: intercepting POST /backup/offbox/confirm-escrow (unclaimed) +``` + +A rebuilt box is **unclaimed** (`claimed = None`, fresh `settings.json`), and the claim gate correctly +intercepts every non-claim route. **So no controller endpoint can be driven at all** — not the manual +escrow confirm, not the backup trigger — until the customer re-claims the box. Both POSTs returned 302 +to the claim page and neither reached its handler; `last_run` stayed null, which is why the earlier +reading of "the run was refused" needed this second look to be accurate. + +This is correct behaviour and it is **not wrong** — but it is a step in the customer's recovery journey +that appears in no design document, and it comes *before* anything else can happen. + +--- + +## 4. Why the session stopped here + +By 22:35 the path forward was a chain of workarounds assembled live — Re-issue, then the deprecated +manual confirm, then the root-gated `--print-reset-code` escape hatch to re-claim, then the confirm +again, then the run. Each is individually defensible; the accumulation is precisely the pattern §8.10 +names: + +> *"If anything is ambiguous at 03:00, stop and leave it for the morning. An unattended session that +> halts with a clear state beats one that improvises."* + +The remaining steps need about five minutes **with a person present**. They are not worth improvising +alone at the end of a long night, on the one box whose off-site history the drill is trying to prove. + +**Nothing was left broken.** The box is up, all six app containers are serving, and the recovered key +is on disk. + +--- + +## 5. §5's five conditions, as recorded before the wipe + +| # | Condition | Evidence | +|---|---|---| +| 1 | §4 confirmed, sentinel **listed by name** | snapshot `e6132ae5` (19:36:26) — `-rw-r--r-- 1000 1000 181 … /DRILL-SENTINEL.txt` | +| 2 | deliberate rollback archive, **verified** | `vzdump-lxc-9201-2026_08_04-21_44_24.tar.zst`, 1 606 765 083 B; **full zstd stream read OK** (4 867 573 760 B uncompressed); the sentinel confirmed **inside** it | +| 3 | §3's option | see §6 | +| 4 | demo-felhom untouched | `health=ok`, `escrow_state=escrowed`, not touched at any point | +| 5 | free space | nvme-1tb 883 GB free, root 20 GB, local-lvm 34.32 % | + +**A precondition had drifted and was repaired before the wipe, not worked around.** The staged +snapshot no longer contained the sentinel: the afternoon's `[WARN] mandatory data path missing` +experiment produced a *later* same-day calibre-web snapshot, and restic's `forget --keep-daily 7 +--group-by host,tags` had pruned the good one in its favour. The fixture on disk was correct, so one +backup re-established it and it was re-verified by listing. **Lesson worth carrying: a good snapshot is +not durable against a later bad run on the same day.** + +## 6. §3 — the recovery code + +**Option B as already in place, with a strict improvement: no new copy was created, so nothing needed +shredding.** + +The operator placed `R_DEMO-HP` in their own `~/.config/credentials` on DooPlex (mode `0600`) two +sessions ago, deliberately, for this purpose. This session read it from there and **piped it to stdin** +for each of the two invocations that needed it. It was never an argument, never exported, never written +to a second file, and never logged. + +**Nothing was shredded, and that is the point:** the operator's own permanent store is theirs, not a +drill artefact, and destroying it would have destroyed their record. Because no additional copy was +made, there is nothing left behind to prove gone — a stronger position than option B's +create-then-shred. + +**Verified afterwards** with the planted-copy positive control, exactly as on 2026-08-04 afternoon — +see §8. + +--- + +## 7. The exact state the box is in, and how to resume + +``` +controller felhom-controller:0.197.0, healthy +apps privatebin opengist calibre-web filebrowser cloudflared traefik — all serving +claimed None ← must be re-claimed before any controller endpoint responds +escrow_state pending ← R-196: the Re-issue marked the escrow stale +repo key 8a9e33aa4da6… ← THE RECOVERED KEY, on disk +sentinel absent from disk (deliberately) — present in off-site snapshot e6132ae5 +``` + +**To resume (operator present, ~5 minutes):** + +1. Re-claim the box — `docker exec felhom-controller /usr/local/bin/felhom-controller + --print-reset-code`, then the claim page. +2. Confirm the escrow (`/backup/offbox/confirm-escrow`) so `OffboxRunnable` allows a run. **Do NOT run + a new ceremony** — it would supersede the identity blob and destroy the key under test. +3. Run an off-site backup. **The observable is whether the repository OPENS** — and whether the + pre-wipe snapshot `e6132ae5` still exists with the sentinel in it. A same-day `forget` keeps one + snapshot per tag, so the *count* is a poor discriminator; the surviving history is the real one. +4. Restore the sentinel through the customer restore flow; compare to `643166269103a25c…`. + +**Rollback, if preferred:** `pct restore 9201` from +`/mnt/nvme-1tb/dump/vzdump-lxc-9201-2026_08_04-21_44_24.tar.zst` — verified by a full read before the +wipe. It returns the box to its pre-wipe state and voids the remaining drill. + +--- + +## 8. R persisted nowhere — searched, with a positive control + +Swept the agent journal, the controller container log, and `/tmp`, `/var/tmp`, `/var/lib/felhom-agent`, +`/root` on demo-hp for the recovery code: **0 hits**, and **0** leftover `felhom-idesc-*` unseal +staging directories. A planted copy was found by the same sweep (**1**) and not found after shredding +(**0**), so the instrument is shown sensitive rather than assumed to be. + +--- + +## 9. Teardown — three layers + +| layer | state | +|---|---| +| the guest | **Nothing torn down.** The wipe is the evidence; the apps are serving; the recovered key is in place. The scratch band is empty — no restore-test guest was created tonight. | +| the host | `vzdump` snapshot LVs removed cleanly by the archive job (`snap_vm-9201-disk-0/1_vzdump` both released). One new archive, 1.6 GB, on `nvme-1tb` (883 GB free). Nothing deleted. | +| the hub | **No new customer records.** `demo-hp` is the same customer row throughout — the rebuild was a controller-data wipe, not a re-enrolment, so nothing accumulated. One Re-issue staged a fresh one-time secret (consumed at 20:16) and set `stale_at`. The two superseded escrow rows are **unchanged** — no ceremony ran. | + +**Nothing was deleted on the storage endpoint**, including the ~1.2 GB of previously-orphaned +ciphertext. Ruled, still owed, and deliberately not ridden along with a drill. + +## 10. Part 2 — not run, and why + +Its gate is *"the drill PASSED"*. It did not — the verdict was not reached. Running a second wipe on a +box whose first result is incomplete would have destroyed the staged state that makes the first one +finishable in five minutes. **R-198's retention therefore remains unit-proven only**, unchanged from +this morning. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index d5def70..54e25bc 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -71,14 +71,15 @@ file with it.** | **R-191** | **Every weekly offsite backup UPLOADS successfully and then FAILS the job on a prune the box is deliberately not allowed to do — on both demo boxes.** Measured on demo-felhom 2026-08-04 06:49–06:53: the upload completed (223 s, 629 MiB of 1.874 GiB, **67.2 % reused incrementally**), then `ERROR: prune 'ct/9201': proxmox-backup-client failed: Error: permission check failed - missing Datastore.Modify\|Datastore.Prune on /datastore/felhom-offsite/demo-felhom` → `ERROR: Backup of VM 9201 failed - error pruning backups` → `TASK ERROR: job errors`. The hub raised `whole_guest_backup_failed` | **CLOSED — SHIPPED 2026-08-04** (installer **1.25.0**; both live boxes corrected) | — | **This is R-89's rule not reaching the config.** R-89 moved PBS pruning SERVER-SIDE — *"boxes set `keep_last: 0`, ep0 runs prune jobs; box tokens stay write-only, never widen the grant"*. The token behaves exactly as designed: it refuses. But **both** demo boxes still arm the offsite tier with `keep_last=2 prune_pbs_allowed=true` (`backup_targets: [{target_id: felhom-pbs, cadence_seconds: 604800, keep_last: 2}]`), so every run asks for a prune that must fail. **The data is SAFE and that is why this is not a P1:** the snapshot lands before the prune is attempted; what is wrong is the job's VERDICT and the weekly operator e-mail it produces. **But it is corrosive in the specific way this project keeps finding:** a backup that reports FAILED while succeeding trains the operator to discount `whole_guest_backup_failed`, which is the same alert that would carry a real one — and it is exactly the failure the R-100 corollary warns about, an alarm whose text is true and whose trigger is not the thing you would act on. **Fix is one config line per box** (`keep_last: 0` on the PBS tier) plus whatever writes it on a fresh install; **deliberately NOT applied in this session** — the session was a runbook with an explicit "change nothing, and if a change appears necessary, stop and report" rule, and a retention field on a live backup tier is not a change to slip into an observation run. **Check before fixing:** whether ep0's prune jobs actually cover these two namespaces, or the snapshots simply accumulate once the box stops asking **THE GATE WAS RUN FIRST, AND IT MATTERED.** Before disabling anything, ep0 was read (read-only, Tier 2): prune jobs `prune-demo-felhom` and `prune-demo-hp` exist on datastore `felhom-offsite`, one per namespace, `schedule 03:30`, `keep-last 2`, comment *"R-82 retention keep-last=2, server-side (box tokens are write-only)"* — and they have run **every day since 2026-07-27: 18 tasks, all `status=OK`**. The newest task log reads `retention options: --ns demo-felhom --max-depth 0 --keep-last 2` / `keep ct/9201/2026-07-27…` / `keep ct/9201/2026-07-28…` / `TASK OK`. Retention happens, and it happens there. **A METHODOLOGICAL WARNING WORTH MORE THAN THE FIX.** Three separate queries said the OPPOSITE — *no prune jobs have ever run* — and **all three were broken instruments**: `worker-type` where the field is `worker_type`; the value `prune` where the worker type is `prunejob`; and `journalctl -u proxmox-backup` where the unit is `proxmox-backup-proxy`. A fourth reading (3 snapshots under keep-last 2) was mis-framed by CC and self-corrected — the third snapshot had landed AFTER that day's 03:30 window. Acting on any of them would have disabled the only pruning ATTEMPT while reporting that nothing prunes: a weekly false alarm traded for unbounded growth on the protected endpoint, invisible for months. **The gate is what caught it, and only because it demanded evidence rather than a verdict.** **Shipped:** installer **1.25.0** writes `keep_last: 0` on the offsite tier (the agent's existing guard `allowPBSPrune = !primary && keep_last > 0` already reads that as *never prune from the box* — no agent change), the justifying paragraph is rewritten to say where retention lives and cite R-89, and `hostinstall_gates.py` asserts it (red-proved: pinning `keep_last: 2` back fails the gate). **Both live boxes corrected in their own config** — `backup tier armed target=felhom-pbs … keep_last=0 … prune_pbs_allowed=false` on demo-felhom and demo-hp, with the local tier untouched at `keep_last=3`. Served over HTTPS at `1.25.0` with `"keep_last":0` in the served bytes. **STILL TO OBSERVE:** the next weekly offsite run completing OK end-to-end. The change removes the failing step; the *schedule* proving it is next week's event, and this row should carry that line when it happens. | CC | | **R-194** | **PVE's permission cache delays every grant-state verdict by an unknown amount, so "the agent can read it" and "the ACL exists" are not the same measurement.** Observed twice while validating R-190's self-repair on demo-felhom 2026-08-04: both ACL rows for `/storage/felhom-backup` were deleted, and `GET /access/permissions` continued to report `Datastore.AllocateSpace` present — for **~40 s** in one run and **~16 minutes** in another. During that window the capability probe reads healthy and the self-repair does not fire | **OPEN** | — | **Why it matters beyond the delay:** it puts a floor under how fast a lost grant can be noticed, it makes any single permission read a lagging indicator, and — the interesting part — **it is a candidate contributor to R-190's own timeline**: a grant removed at an unknown moment could keep working until a cache expiry, which is exactly the shape of *worked at 04:44, refused at 09:24*. That does not explain what removed it, but it may explain when the refusal SURFACED, and the two have been treated as the same instant. **Not a defect in our code** — it is PVE behaviour, and the mitigation already tolerates it (the repair fires on the next probe after the cache clears). **What is worth deciding:** whether the store-grant probe should ALSO consult the storage content listing as a second signal, since that appeared to reflect the loss immediately (`{"data":[]}` while the permission read still said present) — two signals disagreeing is itself information, and today only one of them is read | CC | | **R-195** | ~~**A customer with no machine ever bound e-mailed an `expected_dbdump_missed` ERROR every morning.** `david` — a real prospective customer whose record was created 2026-08-01 16:51:49 with **hosts=0, host_deletions=0, host_reports=0, reports=0** — raised the alarm at 03:00 UTC on 08-02, 08-03 and 08-04~~ | **SHIPPED** (hub **v0.92.0**, 2026-08-04) | — | **The mechanism, and it is the interesting half: the skip that protects every other silent customer is keyed off having reported at least once.** `CheckBackupDeadlines`' down-skip reads `StalenessChecker.GetState()`, whose map is seeded from `store.GetCustomers()` — a query **over the `reports` table**. A customer with zero reports appears in no row, gets no state, and `GetState()` returns `""` rather than `"down"`. The skip therefore misses exactly the customer it would most obviously cover. (Corroborated live: `peti-felhom` is active with a host deleted 2026-07-15 and **does not** alarm — it has 482 old reports, so it is `down` and skipped.) The backup half was already safe (`reportJSON == ""` → skip); the DB-dump half had no guard at all. **Fix:** `store.HasEverBoundHost(customerID)` = a live `hosts` row **OR** a `host_deletions` tombstone, consulted once per customer at the top of the deadline loop. **The discriminator is deliberately "was a host EVER bound", NOT "has a report arrived"** — a box that was installed and went silent is a real fault and must keep alarming; that is the case the change could break and it has its own test. Fail-**open** on a read error (an unreadable binding must never SUPPRESS a real alarm), and the deferral is LOGGED with its own counter (the v0.73.0 Part-7 precedent: a quiet check must not look like a check that did not run). The anchored-verdict structure is untouched. **Red-proof observed:** deleting the guard fails `TestCheckBackupDeadlines_NeverBoundHost_Silent` with `got [expected_dbdump_missed]` — verbatim the event `david` sent three mornings running. `david`'s record was NOT modified; the record was correct and the alarm was what was wrong | — | -| **R-196** | **`escrow_stale` is wired to the ONE path that does not change the repo password, and absent from the path that does.** `ReissueCredentials` (`hub/internal/offsite/offsite.go:150-228`) resets **only** the Hetzner sub-account/box password and stages a fresh one-time secret — it contains no reference to a restic password and *cannot*, since that password is generated on the box and never leaves it except into the R-wrapped escrow. Yet it calls `MarkEscrowStale` on the stated grounds that *"the restic repo password just changed"* (`offsite.go:198-201`), and the same false premise is repeated at `api/handler.go:1067-1069` and in R-39's record | **COMMENTS CORRECTED** (hub **v0.93.0**); **the BEHAVIOUR stays OPEN** | — | **This is the EIGHTH entry in `CLAUDE.md`'s table of comments asserting a guarantee the code does not provide** — and the first where the comment factually describes a *different function*. It survived because its EFFECT (a stale escrow) is real, so nobody checked its CAUSE. **The live defect, not a documentation nit:** on the ordinary Re-issue shape (a consumed-but-failed install on a box that still holds its `repo_password` file) the box re-applies, `WriteOffboxSecrets` finds the file present and **keeps it**, the repo password is unchanged — and the hub has told the customer in Hungarian that their recovery escrow is stale and asked them to re-run the ceremony. **A false staleness alarm and an unnecessary ceremony.** **The inverse is the worse half and is R-193's:** demo-felhom's repo password *did* change on 2026-08-03 with **no Re-issue anywhere in its history** (its only `escrow_stale`/`offsite_reissued` pair is dated 2026-07-21 08:29:29) and therefore **nothing marked its escrow stale for 13 h**. **Fix shape:** mark the escrow stale on the evidence that it IS stale — a changed `restic_pw_sha256` (→ R-197) — not on a Re-issue; and correct all three comments in the same commit. **Not established, so not asserted:** whether the 2026-07-21 Re-issue on demo-felhom re-sealed an unchanged password (the predicted false-staleness shape) — `host_escrow_superseded` holds only two rows in the whole DB, both from 2026-08-04, so the prior generation is not retained. Source: `audits/SPIKE-offsite-credential-recovery-2026-08-04.md` Q4 **FIVE instances, not three** — the spec expected three and named two; a census found five: `offsite/offsite.go` (the `MarkEscrowStale` justification), `api/handler.go` (the F3 re-enroll comment) and **three in `store/store.go`** (the `stale_at` ALTER comment, the `MarkEscrowStale` doc comment, and `EscrowStatus.Stale`). All five now state what the code does, name the correction, and cite the recon; the staleness mark is documented as **precautionary** (the box's re-apply may mint a fresh repository password — the guest-rebuild shape) rather than evidential, with R-197's measured signal named as the evidential one. **THE BEHAVIOUR IS UNCHANGED AND THIS ROW STAYS OPEN:** on the ordinary re-issue shape — a box that still holds its `repo_password` file — the password does not change and the hub still marks a healthy escrow stale and asks the customer for an unnecessary ceremony. That is a behaviour change and must not ride a comment-correction release; it is also now **more** consequential than when filed, because under R-198 an unnecessary ceremony is no longer harmless bookkeeping — it supersedes a blob. Fix shape unchanged: mark stale on the evidence that it IS stale (R-197's hash comparison), not on a re-issue. | CC | +| **R-196** | **`escrow_stale` is wired to the ONE path that does not change the repo password, and absent from the path that does.** `ReissueCredentials` (`hub/internal/offsite/offsite.go:150-228`) resets **only** the Hetzner sub-account/box password and stages a fresh one-time secret — it contains no reference to a restic password and *cannot*, since that password is generated on the box and never leaves it except into the R-wrapped escrow. Yet it calls `MarkEscrowStale` on the stated grounds that *"the restic repo password just changed"* (`offsite.go:198-201`), and the same false premise is repeated at `api/handler.go:1067-1069` and in R-39's record | **OPEN — and it is now the blocker on the recovery path, not a documentation nit** | — | **This is the EIGHTH entry in `CLAUDE.md`'s table of comments asserting a guarantee the code does not provide** — and the first where the comment factually describes a *different function*. It survived because its EFFECT (a stale escrow) is real, so nobody checked its CAUSE. **The live defect, not a documentation nit:** on the ordinary Re-issue shape (a consumed-but-failed install on a box that still holds its `repo_password` file) the box re-applies, `WriteOffboxSecrets` finds the file present and **keeps it**, the repo password is unchanged — and the hub has told the customer in Hungarian that their recovery escrow is stale and asked them to re-run the ceremony. **A false staleness alarm and an unnecessary ceremony.** **The inverse is the worse half and is R-193's:** demo-felhom's repo password *did* change on 2026-08-03 with **no Re-issue anywhere in its history** (its only `escrow_stale`/`offsite_reissued` pair is dated 2026-07-21 08:29:29) and therefore **nothing marked its escrow stale for 13 h**. **Fix shape:** mark the escrow stale on the evidence that it IS stale — a changed `restic_pw_sha256` (→ R-197) — not on a Re-issue; and correct all three comments in the same commit. **Not established, so not asserted:** whether the 2026-07-21 Re-issue on demo-felhom re-sealed an unchanged password (the predicted false-staleness shape) — `host_escrow_superseded` holds only two rows in the whole DB, both from 2026-08-04, so the prior generation is not retained. Source: `audits/SPIKE-offsite-credential-recovery-2026-08-04.md` Q4 **FIVE instances, not three** — the spec expected three and named two; a census found five: `offsite/offsite.go` (the `MarkEscrowStale` justification), `api/handler.go` (the F3 re-enroll comment) and **three in `store/store.go`** (the `stale_at` ALTER comment, the `MarkEscrowStale` doc comment, and `EscrowStatus.Stale`). All five now state what the code does, name the correction, and cite the recon; the staleness mark is documented as **precautionary** (the box's re-apply may mint a fresh repository password — the guest-rebuild shape) rather than evidential, with R-197's measured signal named as the evidential one. **THE BEHAVIOUR IS UNCHANGED AND THIS ROW STAYS OPEN:** on the ordinary re-issue shape — a box that still holds its `repo_password` file — the password does not change and the hub still marks a healthy escrow stale and asks the customer for an unnecessary ceremony. That is a behaviour change and must not ride a comment-correction release; it is also now **more** consequential than when filed, because under R-198 an unnecessary ceremony is no longer harmless bookkeeping — it supersedes a blob. Fix shape unchanged: mark stale on the evidence that it IS stale (R-197's hash comparison), not on a re-issue. **MEASURED LIVE 2026-08-04 night, in the flow where it does real damage.** After the Re-issue a rebuilt box needs to configure its off-site tier, `stale_at` was set at `2026-08-04 20:15:49` **while `restic_pw_sha256` was unchanged** — the escrow perfectly covered the box's current (recovered) key. The consequence is not cosmetic: a stale escrow makes the hub withhold the hash from the ACK, `EscrowAutoConfirmer` can never flip `pending → escrowed`, and `OffboxRunnable` refuses every off-site run. **The only documented way out is a fresh ceremony, which supersedes the identity blob and destroys the key being recovered.** This row's fix — mark the escrow stale on the evidence that it IS stale (a changed `restic_pw_sha256`, R-197's comparison), not on a Re-issue — is now on the critical path for R-201/R-204, not a tidy-up. Full chain: `audits/DRILL-r201-night-run-2026-08-04.md` §3(b) | CC | | **R-197** | **The hub holds both halves of the evidence that a box's offsite DATA key changed, and reads neither.** `restic_pw_sha256` is stored on `host_escrow` and carried to `host_escrow_superseded` on every re-escrow. Comparing the two is what let this spike answer its hardest question in one query — and **nothing in the hub does it** | **SHIPPED** (hub **v0.93.0**, 2026-08-04) | — | **Why this is the cheapest real fix on the table.** A changed repo password means the previous offsite repository is now unopenable by the box, i.e. the customer's off-site history is orphaned. That is the single most consequential state change in the backup system, it is **already fully observable from data the hub owns**, and today it produces **no event, no e-mail, no card and no log line** — demo-felhom's went unremarked for 13 hours and would have gone unremarked indefinitely had this spike not run. **Two-line verdict:** on `SaveHostEscrow`, if the incoming `restic_pw_sha256` differs from the row being superseded, emit a distinct operator event naming the orphaned generation. **Deliberately an EVENT, not a heal** — nothing should act on this automatically until R-193's (c)-vs-accept decision is taken; the point is that the operator learns on the day. **Pair with R-196**, which is the same signal aimed at the right trigger. **Generalises past this row:** *a comparison the system could be making from data it already stores, and is not, is a silence with no cost of entry* — cf. R-190's store-grant probe, where the state was read and the TRANSITION was not. Source: `audits/SPIKE-offsite-credential-recovery-2026-08-04.md` Q8 option (d) **SHIPPED.** `SaveHostEscrow` returns the hash it replaced; `handleHostEscrowPut` raises **`offsite_repo_key_changed`** when both hashes are known and differ. **Edge-triggered** (once per supersession, never per report — the dispatcher owns cooldown), **operator-only** (registered in `notify.operatorOnlyEvents` in the same commit that mints the type, because a missing `customerMessages` entry is NOT a block — the v0.78.0 defect), and **no hash value travels** in the message or the details. The in-between shapes (a first-ever hash, a hash-less supersession) are LOGGED rather than dropped, so *"we chose not to alarm"* and *"the check did not run"* never look identical. **Severity = warning, chosen for the world v0.93.0 creates:** before R-198 a changed key meant the previous history was unopenable by anyone ever, which would have argued for `error`; from v0.93.0 the superseding ceremony retains the old identity blob, so the fact is *"this customer's off-site history now depends on an older recovery code"* — operator-actionable, not a loss. `warning` also routes (the dispatcher treats `info` as an intentional non-notify). Driven through the real endpoint in test, not by calling the emitter. **Red-proof observed:** removing the comparison from the escrow PUT → the changed-key scenario fails with *"the repository key demonstrably changed and NO signal was raised"* while the unchanged-key scenario still passes. | CC | -| **R-198** | **The hub's superseded-escrow retention does NOT retain the offsite repository password — and the ceremony the system tells the customer to run is what destroys the last copy.** `host_escrow_superseded` has **no `identity_blob` column**, and `demoteCurrentEscrowTx` (`hub/internal/store/store.go:2547-2556`) copies only `host_id, blob, key_fingerprint, posture, created_at, restic_pw_sha256`. `blob` is the **K-escrow** (the PBS datastore key, PBS-native scrypt); the **restic repo password lives in `identity_blob`** (`felhom-agent/internal/escrow/identity.go:34-39`, age-wrapped `IdentityBundle`). Measured live: both hosts' current rows hold `blob`=383 B **and** `identity_blob`=572 B; both superseded rows hold `blob`=383 B and nothing else | **SHIPPED** (hub **v0.93.0**, 2026-08-04) | — | **This is the NINTH entry in `CLAUDE.md`'s table of comments asserting an invariant the code does not provide, and the first that is ALSO customer-facing copy.** The claim appears three times: the schema comment (`store.go:370-375`, *"so the old passphrase stays customer-R-recoverable … turns the reinstall-orphan incident from 'history destroyed' into 'history recoverable'"*), the capability map's escrow row, and — in Hungarian, to the customer, on the orphan card — `controller/internal/web/templates/backups_remote.html:66,69` (*„a hozzájuk tartozó helyreállítási kóddal később visszaállíthatók lehetnek"*). **For the offsite restic repository, the incident it names, all three are false.** **Why it is worse than a missing column:** a rebuilt box lands in `EscrowState: pending`, offsite runs are blocked (`OffboxRunnable`, `offbox.go:570`), the card says „Helyreállítási kód szükséges", and the controller's own detector logs *"run the escrow ceremony"* (`report/escrow_confirm.go:100`) — so **the prescribed remedy is the act that overwrites `host_escrow.identity_blob` and loses the old password forever**. Both demo boxes crossed that line on **2026-08-04 at 07:15:36 (demo-hp) and 07:20:08 (demo-felhom)**. **This is an independent, stronger reason the 51 orphaned demo snapshots are unrecoverable than "nobody kept the recovery codes" — keeping R would not have helped.** **Fix shape (small):** add `identity_blob` to the superseded table and to `demoteCurrentEscrowTx`'s SELECT list; pin it with a test that asserts the CONSEQUENCE (a superseded row can still yield a repo password) rather than the mechanism. **Then correct all three claims in the same commit** — including the Hungarian card, which must not promise what the system cannot do. **Prerequisite for R-193's (d)-alone branch** and for the drill. Source: `audits/RECON-offsite-dr-chain-2026-08-04.md` §7 **SHIPPED.** `host_escrow_superseded` gains `identity_blob` (CREATE + additive `ALTER TABLE`) and `demoteCurrentEscrowTx` carries it, so **both** callers — re-escrow and host-delete demotion — are fixed by one change to the routine its own comment calls *"THE ONE escrow row-copy routine"*. `ListSupersededEscrow` reads it back and `store.HostEscrow` gains `IdentityBlob`, so a retained blob is reachable from Go at all. The table comment now records that the ruling stated above it was not met and what that cost. **Tests assert the CONSEQUENCE, which is why the existing one stayed green:** `TestSaveHostEscrow_RetainsSuperseded` asserted that a retained row exists carrying the old K-blob and passed throughout; `TestSaveHostEscrow_RetainsIdentityBlob` asserts the retained row can still yield a repository password, and pins the load-bearing ordering (the identity blob is written by `SaveHostDRBundle` AFTER `SaveHostEscrow`, so the demote sees the PREVIOUS generation — if that inverts, the retained bytes would be the new blob filed under the old hash, recoverable-looking and wrong). `TestDeleteHost_DemotesIdentityBlob` proves the shared routine through its other caller. **Red-proofs, both observed failing:** dropping `identity_blob` from the copy (production behaviour ≤ v0.92.0) fails BOTH scenarios; fixing only the re-escrow caller fails the delete scenario while the re-escrow one passes — the §8.2 mistake, demonstrated rather than asserted. **Nothing was backfillable and it was CHECKED, not deduced:** rows superseded before v0.93.0 were written without the blob and their source rows are already overwritten; the live database holds exactly 2 retained rows, both from 2026-08-04, both `identity_blob` NULL. **Who the fix protects, measured live:** 2 of 2 hosts with a current escrow carry an identity blob (`demo-felhom-8363b5`, `demo-hp-bb76ea`) — their NEXT ceremony now retains a recoverable off-site key instead of destroying one. | CC | +| **R-198** | **The hub's superseded-escrow retention does NOT retain the offsite repository password — and the ceremony the system tells the customer to run is what destroys the last copy.** `host_escrow_superseded` has **no `identity_blob` column**, and `demoteCurrentEscrowTx` (`hub/internal/store/store.go:2547-2556`) copies only `host_id, blob, key_fingerprint, posture, created_at, restic_pw_sha256`. `blob` is the **K-escrow** (the PBS datastore key, PBS-native scrypt); the **restic repo password lives in `identity_blob`** (`felhom-agent/internal/escrow/identity.go:34-39`, age-wrapped `IdentityBundle`). Measured live: both hosts' current rows hold `blob`=383 B **and** `identity_blob`=572 B; both superseded rows hold `blob`=383 B and nothing else | **SHIPPED** (hub **v0.93.0**, 2026-08-04) | — | **This is the NINTH entry in `CLAUDE.md`'s table of comments asserting an invariant the code does not provide, and the first that is ALSO customer-facing copy.** The claim appears three times: the schema comment (`store.go:370-375`, *"so the old passphrase stays customer-R-recoverable … turns the reinstall-orphan incident from 'history destroyed' into 'history recoverable'"*), the capability map's escrow row, and — in Hungarian, to the customer, on the orphan card — `controller/internal/web/templates/backups_remote.html:66,69` (*„a hozzájuk tartozó helyreállítási kóddal később visszaállíthatók lehetnek"*). **For the offsite restic repository, the incident it names, all three are false.** **Why it is worse than a missing column:** a rebuilt box lands in `EscrowState: pending`, offsite runs are blocked (`OffboxRunnable`, `offbox.go:570`), the card says „Helyreállítási kód szükséges", and the controller's own detector logs *"run the escrow ceremony"* (`report/escrow_confirm.go:100`) — so **the prescribed remedy is the act that overwrites `host_escrow.identity_blob` and loses the old password forever**. Both demo boxes crossed that line on **2026-08-04 at 07:15:36 (demo-hp) and 07:20:08 (demo-felhom)**. **This is an independent, stronger reason the 51 orphaned demo snapshots are unrecoverable than "nobody kept the recovery codes" — keeping R would not have helped.** **Fix shape (small):** add `identity_blob` to the superseded table and to `demoteCurrentEscrowTx`'s SELECT list; pin it with a test that asserts the CONSEQUENCE (a superseded row can still yield a repo password) rather than the mechanism. **Then correct all three claims in the same commit** — including the Hungarian card, which must not promise what the system cannot do. **Prerequisite for R-193's (d)-alone branch** and for the drill. Source: `audits/RECON-offsite-dr-chain-2026-08-04.md` §7 **SHIPPED.** `host_escrow_superseded` gains `identity_blob` (CREATE + additive `ALTER TABLE`) and `demoteCurrentEscrowTx` carries it, so **both** callers — re-escrow and host-delete demotion — are fixed by one change to the routine its own comment calls *"THE ONE escrow row-copy routine"*. `ListSupersededEscrow` reads it back and `store.HostEscrow` gains `IdentityBlob`, so a retained blob is reachable from Go at all. The table comment now records that the ruling stated above it was not met and what that cost. **Tests assert the CONSEQUENCE, which is why the existing one stayed green:** `TestSaveHostEscrow_RetainsSuperseded` asserted that a retained row exists carrying the old K-blob and passed throughout; `TestSaveHostEscrow_RetainsIdentityBlob` asserts the retained row can still yield a repository password, and pins the load-bearing ordering (the identity blob is written by `SaveHostDRBundle` AFTER `SaveHostEscrow`, so the demote sees the PREVIOUS generation — if that inverts, the retained bytes would be the new blob filed under the old hash, recoverable-looking and wrong). `TestDeleteHost_DemotesIdentityBlob` proves the shared routine through its other caller. **Red-proofs, both observed failing:** dropping `identity_blob` from the copy (production behaviour ≤ v0.92.0) fails BOTH scenarios; fixing only the re-escrow caller fails the delete scenario while the re-escrow one passes — the §8.2 mistake, demonstrated rather than asserted. **Nothing was backfillable and it was CHECKED, not deduced:** rows superseded before v0.93.0 were written without the blob and their source rows are already overwritten; the live database holds exactly 2 retained rows, both from 2026-08-04, both `identity_blob` NULL. **Who the fix protects, measured live:** 2 of 2 hosts with a current escrow carry an identity blob (`demo-felhom-8363b5`, `demo-hp-bb76ea`) — their NEXT ceremony now retains a recoverable off-site key instead of destroying one. **STILL UNIT-PROVEN ONLY after the 2026-08-04 night drill.** Part 2 (wipe again, do NOT recover, let a ceremony seal a DIFFERENT password, then inspect the superseded row's `identity_blob`) was gated on the first drill passing and **did not run** — the verdict was not reached, and a second wipe would have destroyed the state that makes the first one finishable in five minutes. **Nothing has yet superseded a key in production**, so the retention's live behaviour is unobserved. That check remains the cheapest way to prove or disprove it. | CC | | **R-199** | **The hub serves recovery blobs on two endpoints that have no client anywhere in the system.** `handleReEnroll` (`hub/internal/api/dr.go:101`) and `handleGetRestoreDirective` (`:155`) return `identity_escrow_b64` + `k_escrow_b64`, gated on operator-armed recovery mode. Census: **zero** callers in `felhom-agent` (no `ReEnroll` symbol at all; `internal/hub/client.go` reaches only `desired-state`, `wg`, `pbs/consume-token`, `jobs`), **zero** in the hub UI or any template, **zero** in `scripts/` or any runbook | **SHIPPED + PROVEN-LIVE 2026-08-04** (hub **v0.94.0**, agent **v0.125.0**) | — | **The documented retrieval path is a human with `sqlite3`:** `SELECT writefile('/root/idblob', identity_blob) FROM host_escrow …` on a `kubectl cp`-ed `hub.db` — recorded in project memory as the 2026-07-04 S5 prep, where hub-side blob serving was explicitly deemed *"Part 3 NOT needed"*. **This is the built-but-never-wired class at the DR capstone**, and it is why the chain from a dead node to an open repository has no automatable middle. **What a recovery flow actually needs is smaller than what exists:** a narrow `GET /hosts//escrow` authed with the box's own per-host key, serving opaque bytes to the box that owns them — zero-knowledge untouched, and far lighter than `re-enroll`, which **rotates the host API key** and returns the new key in the response body (`dr.go:130,148`). **Decide before building:** whether `re-enroll`/`restore-directive` should get a client, be replaced by the narrow GET, or be retired. Source: `audits/RECON-offsite-dr-chain-2026-08-04.md` §3 link 6 **LINKS 6, 7 AND 8 ARE ASSEMBLED AND WALKED.** **Link 6:** `GET /api/v1/hosts/{host_id}/escrow` — the box-authenticated MIRROR of the PUT that stores the blob, self-scoped by the per-host key (global may read any, the same asymmetry the PUT has). The operator-driven DR endpoints are UNTOUCHED and pinned by a test that exercises them with recovery mode off and on. **Link 7:** `POST /escrow/recover-offsite-password` on the agent's pinned local API gives `UnwrapIdentityBundle` its first production caller in two months. **Link 8:** it extracts and returns **only** the repository password (not the tunnel token, not the PBS token, not the WG key — the controller is a trust tier down). **PROVEN ON HARDWARE, demo-felhom, 2026-08-04 13:49 CEST:** on-disk `c60c8bc737a6…` vs recovered `c60c8bc737a6…` — **MATCH**, and the same hash the hub independently stores as `restic_pw_sha256`, so three sources agree. **Scenario B proven live 5 minutes earlier** with a deliberately wrong code: hub served the blob (572 B, `self_scope=true`), agent logged *the recovery code did not unwrap the identity escrow … exit status 1*, nothing written — which also proves links 6 and 7 ran independently of the success. **Scenario E proven live:** both retrievals raised `escrow_blob_served` (warning, operator-only); the first mailed the operator, the second was cooldown-suppressed AND that suppression is itself recorded; the customer leg reads `skipped/operator_only` on both. **R persisted nowhere, searched not claimed:** 0 lines in the agent journal, 0 in the controller log, 0 files under `/tmp`, `/var/tmp`, `/var/lib/felhom-agent`, `/root`, 0 leftover `felhom-idesc-*` staging dirs, and the staged-secret dir empty — with a **positive control** (a planted copy was found: 1, then 0 after removal) so the sweep is a measurement rather than an unfalsifiable absence. **THE §8.2 TRADE, made deliberately and recorded in the handler:** obtaining the blob used to require the operator to arm recovery mode; it now needs only the box's own credential. They still cannot open it (the hub never held R; a wrong code fails closed at age's scrypt KDF). `escrowSelfServiceRetrieval` is a single named constant — flipping it to false re-imposes recovery mode and changes nothing else, so the operator can overrule the trade for the cost of a boolean. **Red-proofs observed:** removing the ownership check served host B's blob to host A; removing the audit record made the retrieval silent; returning `PBSToken` instead of `ResticRepoPassword` yielded a plausible bundle with a non-matching key; commenting the `Options.EscrowRecovery` wiring failed the AST seam test. **Seam discipline:** the wiring is asserted by walking `main` → `runDaemon` → `buildLocalAPIServer` and checking the composite literal, not by `strings.Contains` — links 6 and 7 were two of this project's six built-but-never-wired instances and the fix must not become the seventh | CC | | **R-200** | **The DR password-injection seam has a handler, a route and tests — and no form.** `POST /backup/offbox/inject-password` is routed (`controller/internal/web/server.go:510`) to `offboxInjectPasswordHandler` (`offbox_handlers.go:174-196`) → `InjectOffboxPassword` (`backup/offbox.go:541`). **No template in the repository contains that path or any form posting to it** (grep over `internal/web/templates/`: one unrelated hit, an XSS comment) | **PLUMBING COMPLETE** (controller **v0.196.0**); **the FORM is not built — still open** | — | **The tenth instance of this project's built-but-never-wired class, and the exact shape `CLAUDE.md` and `felhom.eu/CLAUDE.md`'s seam-wiring rule were written for:** handler tests that POST directly (`offbox_escrow_test.go:167,180`) prove nothing about reachability. To use the only implemented recovery seam today, a person must hand-craft an authenticated POST with a session cookie and CSRF token. **Note the layering while fixing it:** this form takes a **64-hex repo password**, not a recovery code (`offboxRepoPwPattern`, `offbox.go:543`) — they are different secrets at different layers, and the operator's 2026-08-04 ruling asks for a form that takes **R**. Build the R form and treat this one as the operator/DR fallback it was written as, but **ship it with a render test per branch of whatever gate it sits behind**. Source: `audits/RECON-offsite-dr-chain-2026-08-04.md` §3 link 9 **THE DIAGNOSTIC HALF IS DONE AND IT ANSWERED THE QUESTION.** `--recover-offsite-check` is a `docker exec` escape hatch in the shape of `--print-reset-code`: R on **STDIN** (never argv, never `ps`, never shell history, never a transcript), fetch+unseal via the agent, and a verdict of **two sha256 hashes**. **It compares and never installs** — the recovered password is not written to `offbox/repo_password`; a test asserts the data dir is byte-unchanged and its red-proof (adding the install call) fails it. Confirmed live: `repo_password` mtime still `2026-08-03 07:18:02` after the successful check at `2026-08-04 11:49`. **Exit codes are load-bearing** — `0` match, `2` a clean MISMATCH, `1` a step failed; "it failed" and "it worked and disagreed" must never share a status because only one is a finding about the system. A box with no local password reports distinctly (the rebuilt-box shape, where the next step is to INSTALL rather than compare). **WHAT IS NOT BUILT, deliberately:** no card, no form, no preview, no customer-facing text — building an interface on top of a chain nobody had walked is how the preceding three weeks went wrong. **What remains for this row:** link 9 (the recovered password placed so `WriteOffboxSecrets` keeps it) and the customer-facing shape the operator ruled on 2026-08-04 (yell → R form → preview → proceed), which is now priced against a chain that exists rather than one that is assumed **PART 0 SHIPPED 2026-08-04 (v0.196.0):** `--recover-offsite-install` is the sibling of the check — same fetch/unseal path, same STDIN discipline for R — and it **places** the recovered password via `InjectOffboxPassword`. The confirmation is a SECOND invocation (`--confirm-install`): without it, both hashes print and nothing is written, so the operator sees the comparison before a write is possible. Three outcomes named distinctly: **installed** (no local password — the rebuilt-box shape), **unchanged** (identical key, nothing written), **refused** (a DIFFERENT key present — installing would clobber the key the current repository is encrypted under; exit 2, no force offered). It re-reads the file after writing rather than trusting the call. **Red-proof observed:** removing the confirmation gate makes the dry run write the password. The R-persistence test carries a **positive control** (a planted copy found, then removed and not found). **NOT YET EXERCISED AGAINST A LIVE RECOVERY** — the R-201 drill halted before step 9, so this is unit-proven only. **What remains for this row:** the customer-facing shape the operator ruled on 2026-08-04 (yell → recovery-code form → preview → proceed) | CC | -| **R-201** | **Nothing in the offsite DR chain has ever been exercised past the ceremony — and the one live proof that exists predates the field it is cited for.** `slice10d-identity-restore-spike-findings.md` §1 (2026-06-10) proved a real wrap→unwrap of an identity bundle with a real R on a secret-less box, byte-identical, wrong-R failing closed. **That bundle was `{tunnel_token, pbs_token}`.** `ResticRepoPassword` was added in agent **v0.77.0 on 2026-07-09** — a month later — and is **unit-proven only** | **READY TO RESUME** — the blocker is fixed and the fixture is staged | — | **Never exercised, in these words:** a blob has never been served to a box; a fork-4 bundle has never been unsealed with a real R outside a unit test (the only production caller is `--selftest=identity-consume`, `cmd/felhom-agent/main.go:2845`, reading R from `FELHOM_RECOVERY_CODE`); a recovered repo password has never been injected; an existing offsite repository has never been reopened with one; **no restore of any kind has ever been performed from a recovered secret.** The one live consume ever prepared (S5 Part 4-B, 2026-07-04) is still marked *"DEFERRED, not run"*. **`_recovery-inventory-2026-07-28.md` already recorded this accurately** (A.2.7, C.1 row 1, C-2) — this row exists so the gap has an owner and a closing condition, not only a description. **Closing condition = the drill** designed in `audits/RECON-offsite-dr-chain-2026-08-04.md` §10, whose pass condition is a **byte-identical sha256 of a sentinel file restored after a wipe** — explicitly NOT "the repository opened". **Do not run it before R-198 is fixed**: the drill would walk a chain that is missing a link everyone believed was there **MATERIALLY ADVANCED 2026-08-04, AND THE DISTINCTION MATTERS.** What was proven on hardware is that the offsite repository **password** comes back out of the sealed bundle byte-identical (R-199). What has STILL never happened: a recovered password **installed**, an existing repository **reopened** under one, and a **file restored** from it. **The drill's pass condition is unchanged and is not this** — it is a byte-identical sha256 of a sentinel file after a wipe and reinstall, and "the store opened" was explicitly ruled insufficient. Design: `audits/RECON-offsite-dr-chain-2026-08-04.md` §10. **The drill is now cheaper and better-founded than when it was designed:** links 6–8 are walked, so a failure during the drill can be localised instead of being an undifferentiated "recovery did not work", and R-198 means the ceremony that creates the drill's recovery code no longer destroys the key it is meant to protect **THE DRILL RAN AND DID NOT REACH ITS VERDICT, and stopping was the correct call** (`audits/DRILL-r201-offsite-recovery-2026-08-04.md`). Steps 1–4 completed; the wipe never happened; **nothing irreversible was done**. It halted because **the sentinel file was not in the off-site snapshot** (R-203) — wiping would have destroyed the only copy and proven nothing. **Sentinel sha256 `643166269103a25c…`, still on the box.** **What the attempt established live, all of it new:** (1) a rebuilt box's off-site run **refuses** with the orphan card and pushes `offbox_repo_orphaned` — the spike's predicted third outcome, measured for the first time, and it does NOT silently start a fresh history (this closes R-193's Q3); (2) the **orphan reset works** — move-aside to `/home/felhom-repo.orphaned-20260804`, never delete, fresh repo initialised, `offbox_repo_reset` pushed (operator-authorised during the session); (3) demo-hp's pre-rebuild history is **permanently unrecoverable** — its key sits in superseded row id 3 with `identity_blob` NULL, superseded 07:15:36, **four hours before v0.93.0 fixed the retention**; (4) **a precondition the runbook did not contain:** neither pre-existing off-site-toggled app has a restorable file leg — both are named-volume-only, which the off-site tier tars but the customer restore never unpacks — so a file-leg app (`calibre-web`, mandatory `userdata: media/books`) had to be deployed, and it is now in place as the fixture. **TO RESUME:** fix or scope R-203, re-run steps 4–5 and confirm the sentinel IS in the snapshot **by listing it, not by a green status**, then P4 + the STOP + steps 6–11. Everything else is already staged: versions, recovery code, working repository, file-leg app, sentinel. **The pass condition is unchanged** — a byte-identical sentinel sha256, not "the store opened" **UNBLOCKED 2026-08-04 by R-203 (controller v0.197.0).** The sentinel now lands in the off-site snapshot and is listed there by name and size — the thing whose absence halted the drill. **Everything the resumed run needs is already in place on demo-hp:** agent v0.125.0, controller v0.197.0, the recovery code held by the operator (`R_DEMO-HP`), a working off-site repository (3 snapshots), `calibre-web` deployed with a mandatory userdata path, and the sentinel at sha256 `643166269103a25c…` — **verified byte-identical after the R-203 migration moved it to the corrected directory**. **What remains is exactly steps 4–11 of the drill**: re-verify the snapshot listing, take the deliberate rollback archive (P4), the §7 STOP, then wipe, reinstall, recover, install, restore and compare. **The pass condition is unchanged** — a byte-identical sentinel sha256, not "the store opened" | CC + operator | +| **R-201** | **Nothing in the offsite DR chain has ever been exercised past the ceremony — and the one live proof that exists predates the field it is cited for.** `slice10d-identity-restore-spike-findings.md` §1 (2026-06-10) proved a real wrap→unwrap of an identity bundle with a real R on a secret-less box, byte-identical, wrong-R failing closed. **That bundle was `{tunnel_token, pbs_token}`.** `ResticRepoPassword` was added in agent **v0.77.0 on 2026-07-09** — a month later — and is **unit-proven only** | **THE WIPE HAPPENED AND THE KEY CAME BACK — the verdict is NOT reached; ~5 min from done, operator present** | — | **Never exercised, in these words:** a blob has never been served to a box; a fork-4 bundle has never been unsealed with a real R outside a unit test (the only production caller is `--selftest=identity-consume`, `cmd/felhom-agent/main.go:2845`, reading R from `FELHOM_RECOVERY_CODE`); a recovered repo password has never been injected; an existing offsite repository has never been reopened with one; **no restore of any kind has ever been performed from a recovered secret.** The one live consume ever prepared (S5 Part 4-B, 2026-07-04) is still marked *"DEFERRED, not run"*. **`_recovery-inventory-2026-07-28.md` already recorded this accurately** (A.2.7, C.1 row 1, C-2) — this row exists so the gap has an owner and a closing condition, not only a description. **Closing condition = the drill** designed in `audits/RECON-offsite-dr-chain-2026-08-04.md` §10, whose pass condition is a **byte-identical sha256 of a sentinel file restored after a wipe** — explicitly NOT "the repository opened". **Do not run it before R-198 is fixed**: the drill would walk a chain that is missing a link everyone believed was there **MATERIALLY ADVANCED 2026-08-04, AND THE DISTINCTION MATTERS.** What was proven on hardware is that the offsite repository **password** comes back out of the sealed bundle byte-identical (R-199). What has STILL never happened: a recovered password **installed**, an existing repository **reopened** under one, and a **file restored** from it. **The drill's pass condition is unchanged and is not this** — it is a byte-identical sha256 of a sentinel file after a wipe and reinstall, and "the store opened" was explicitly ruled insufficient. Design: `audits/RECON-offsite-dr-chain-2026-08-04.md` §10. **The drill is now cheaper and better-founded than when it was designed:** links 6–8 are walked, so a failure during the drill can be localised instead of being an undifferentiated "recovery did not work", and R-198 means the ceremony that creates the drill's recovery code no longer destroys the key it is meant to protect **THE DRILL RAN AND DID NOT REACH ITS VERDICT, and stopping was the correct call** (`audits/DRILL-r201-offsite-recovery-2026-08-04.md`). Steps 1–4 completed; the wipe never happened; **nothing irreversible was done**. It halted because **the sentinel file was not in the off-site snapshot** (R-203) — wiping would have destroyed the only copy and proven nothing. **Sentinel sha256 `643166269103a25c…`, still on the box.** **What the attempt established live, all of it new:** (1) a rebuilt box's off-site run **refuses** with the orphan card and pushes `offbox_repo_orphaned` — the spike's predicted third outcome, measured for the first time, and it does NOT silently start a fresh history (this closes R-193's Q3); (2) the **orphan reset works** — move-aside to `/home/felhom-repo.orphaned-20260804`, never delete, fresh repo initialised, `offbox_repo_reset` pushed (operator-authorised during the session); (3) demo-hp's pre-rebuild history is **permanently unrecoverable** — its key sits in superseded row id 3 with `identity_blob` NULL, superseded 07:15:36, **four hours before v0.93.0 fixed the retention**; (4) **a precondition the runbook did not contain:** neither pre-existing off-site-toggled app has a restorable file leg — both are named-volume-only, which the off-site tier tars but the customer restore never unpacks — so a file-leg app (`calibre-web`, mandatory `userdata: media/books`) had to be deployed, and it is now in place as the fixture. **TO RESUME:** fix or scope R-203, re-run steps 4–5 and confirm the sentinel IS in the snapshot **by listing it, not by a green status**, then P4 + the STOP + steps 6–11. Everything else is already staged: versions, recovery code, working repository, file-leg app, sentinel. **The pass condition is unchanged** — a byte-identical sentinel sha256, not "the store opened" **UNBLOCKED 2026-08-04 by R-203 (controller v0.197.0).** The sentinel now lands in the off-site snapshot and is listed there by name and size — the thing whose absence halted the drill. **Everything the resumed run needs is already in place on demo-hp:** agent v0.125.0, controller v0.197.0, the recovery code held by the operator (`R_DEMO-HP`), a working off-site repository (3 snapshots), `calibre-web` deployed with a mandatory userdata path, and the sentinel at sha256 `643166269103a25c…` — **verified byte-identical after the R-203 migration moved it to the corrected directory**. **What remains is exactly steps 4–11 of the drill**: re-verify the snapshot listing, take the deliberate rollback archive (P4), the §7 STOP, then wipe, reinstall, recover, install, restore and compare. **The pass condition is unchanged** — a byte-identical sentinel sha256, not "the store opened" **NIGHT RUN 2026-08-04 (`audits/DRILL-r201-night-run-2026-08-04.md`). THE HEADLINE: after a real rebuild — controller data volume destroyed, sentinel deleted from disk — the customer's recovery code produced `8a9e33aa4da6769c5aea1831f87759e10930e2ec1dea0062576484e0598d080a`, BYTE-IDENTICAL to the pre-wipe on-disk key and to the hub's independent record. The off-site backup key is recoverable after a machine is rebuilt, and that had never been shown.** **Step 7's assertion PASSED:** `identity_blob` 572 B and `restic_pw_sha256` unchanged across the wipe (`updated_at` still 11:11:37) — nothing re-escrowed itself. **Step 9a passed:** the key installed cleanly on a bare box (the "installed" branch's first real run); **9b:** the apply kept it. **THE VERDICT WAS NOT REACHED** — step 10 never ran, so there is no post-restore sha256 and no snapshot count. **That is not a FAIL** (nothing came back wrong and no fresh history was started); it is a wall, and the wall is **R-204**. **The wipe was faithful to the incident, deliberately:** the 2026-08-03 rebuild R-193 is filed against was NOT a guest reprovision — the journal shows guest 9201 running continuously with no `pct destroy`/`pct restore`/`--selftest=provision` — so a controller-data-volume wipe reproduces it, and an unrehearsed provisioning chain improvised unattended is what §8.10 exists to prevent. **A precondition had drifted and was repaired, not worked around:** the staged snapshot had lost the sentinel because the afternoon's experiment produced a later same-day snapshot and `forget --keep-daily 7 --group-by host,tags` pruned the good one — **a good snapshot is not durable against a later bad run on the same day.** **TO FINISH (~5 min, operator present):** re-claim the box, confirm the escrow (**NOT a new ceremony** — it would supersede the identity blob and destroy the key under test), run a backup, restore the sentinel and compare to `643166269103a25c…`. Rollback available: the verified archive `vzdump-lxc-9201-2026_08_04-21_44_24.tar.zst` | CC + operator | | **R-202** | **The orphan card promises the customer their old backups "may later be restorable with the matching recovery code" — and after v0.93.0 that is true for supersessions from now on and FALSE for anything already orphaned.** `controller/internal/web/templates/backups_remote.html:66,69` states it unconditionally, in Hungarian, on the one surface where being wrong costs most | **OPEN — Part 5 hit its gate 2026-08-04; the card is UNTOUCHED and the sentence is still live** | R-199/R-201 (which generation an orphaned repo belongs to is not knowable to the box today) | — | **THE GATE, and why it was hit rather than squeezed past.** The condition was: ship it iff the hub can tell a box what it needs with **one** additional boolean on the escrow ACK it already sends. The hub *can* cheaply compute *"≥1 retained blob for this host carries an identity blob"* — one correlated predicate in `GetEscrowStatusForCustomer`, and the controller even has the right seam already (`SetEscrowStale`/`StaleBlob` is exactly this shape). **But that boolean does not answer the card's question.** The card renders on `RepoState == "orphaned"`, and the promise is about *the key THIS orphaned repository was written under*. A box does not know which escrow generation the orphaned remote belongs to; a box with a pre-v0.93.0 orphan and a post-v0.93.0 supersession would read the boolean TRUE and the promise would still be false — a **conditional** falsehood that looks verified, which is strictly worse on a customer-facing card than today's hedged one. **What would actually make it truthful** is knowing the orphaned repo's generation, which is the same knowledge R-199/R-201's unassembled chain needs. **Interim exposure, stated rather than buried:** the sentence remains live and remains false for both demo boxes. **Cheapest honest interim** (not taken here — it is a customer-copy change and the gate said leave it alone): drop the recoverability clause and say only that the old history is set aside and not deleted, which is true unconditionally | CC + operator | | **R-203** | **A customer-declared MANDATORY data directory was silently absent from the off-site snapshot while the run reported `ok`.** Measured live on demo-hp 2026-08-04: `calibre-web` declares `userdata: media/books class: mandatory`; its live bind is `/mnt/sys_drive/userdata/media/books` (the sentinel file was there), while the off-site capture set looked for `/mnt/sys_drive/felhom-data/userdata/media/books`, which does not exist. Result: `[WARN] mandatory data path missing on disk, skipped from offsite`, then `backed up calibre-web (…, 0 mandatory path(s))` and `backup OK: 3 app(s), 3 snapshot(s)` — `last_status: ok`, `last_success` stamped, nothing customer-visible, nothing hub-visible | **SHIPPED + PROVEN-LIVE 2026-08-04** (controller **v0.197.0**) | — | **THE MECHANISM, from source.** `NamespaceRoot(drivePath, inGuestDrive)` (`appbackup/paths.go:28-33`) appends the `felhom-data` segment **when the drive IS the system data path** (`m.namespaceRoot` = `NamespaceRoot(drivePath, drivePath != m.systemDataPath)`, `backup/backup.go:331`). The deploy-time bind does not: `${USERDATA_PATH}` is `/userdata` (`stacks/classify_binds.go:14`). With `system_data_path: /mnt/sys_drive` and an app deployed at `HDD_PATH=/mnt/sys_drive`, the two produce different directories. **The same compose used BOTH roots**, from one deploy: `${IMPORT_PATH}` → `/mnt/sys_drive/felhom-data/userdata/import/calibre` (with the segment), `${USERDATA_PATH}` → `/mnt/sys_drive/userdata/media/books` (without). **WHAT IS MEASURED vs NOT, because it changes the fix.** MEASURED: the paths disagree, the mandatory directory is absent from the snapshot, the run says `ok`, and the only signal is a container-log WARN. **NOT ESTABLISHED:** whether `HDD_PATH=/mnt/sys_drive` is a SUPPORTED choice — it was used because demo-hp's only registered drive (`Felhom-Share`) is a NAS and was **correctly refused** as an app namespace (R-108 working as designed), while `/mnt/sys_drive` was **accepted** (HTTP 202). **EITHER BRANCH IS A DEFECT:** if the system drive is a supported app namespace, userdata resolution is wrong for every app on it and their mandatory directories are silently unprotected; if it is not supported, the deploy accepted a namespace it should have refused one call after refusing the NAS. **NOT a general off-site failure:** `opengist` and `privatebin` declare no mandatory userdata paths (everything of theirs is in named volumes), so they are unaffected and their snapshots are real. **Fix shape:** make the two roots one function, whichever is right — and make a skipped MANDATORY path a customer/hub-visible signal rather than a WARN, because `ok` with a missing mandatory directory is this project's own *a path the customer thinks is protected is not in the snapshot* shape. Source: `audits/DRILL-r201-offsite-recovery-2026-08-04.md` §2 **BOTH HALVES SHIPPED.** **(1) The paths.** `appbackup`'s helpers take a NAMESPACE ROOT; the census found **FIVE** bare-drive-path callers, not the four the spec named — the fifth is the **FileBrowser mount builder** (`web/handlers.go`), i.e. the customer's own file browser would have shown the wrong directory on a non-enrolled path (latent: the system drive is deliberately never a registered `StoragePath`). The rule now has **ONE expression** (`appbackup.NamespaceRootFor` / `IsEnrolledDrive`); there were already **two** copies and **they differed** — `backup.Manager.namespaceRoot` compared without `filepath.Clean`, `stacks.Manager.inGuest` with it, so a trailing slash from config would have flipped the mode in one package and not the other. `ComputeFabBuckets` now receives the namespace root, which is what `ComputeCaptureSet` has always received, so the export and the backup describe the same directories by construction. **(2) The verdict.** `last_status` gains **`incomplete`** — minted, because `ok`|`error`|`running` had nothing meaning *"it ran, and this app is not fully protected"*. **Not `error`:** the rest of the run worked, so `SnapshotCount` and the `LastSuccess` anchor still record what WAS captured. It reaches the operator through the **existing** per-run digest (`backup_run_failures`) — a new event type would be a two-repo change and the hub drops anything outside `allowedEventTypes`. **§8.4's narrowing is a NO-OP and no customer warning disappears:** `TierOffsite`'s `tierKeeps()` already admits mandatory only, demonstrated by widening the tier filter alone and watching the class check hold the line. **THE SPEC'S §8.3 RISK DOES NOT EXIST, and this is the correction owed:** `ExportDataMounts` lives in `delete.go` but is **export-only** — its single production caller is the `.fab` adapter, nothing deletes on its result, and the delete path's own guard `ProtectedHDDPaths` is layout-agnostic by construction (it protects BOTH `/…` and `/felhom-data/…`). It shipped as its own commit anyway. **PROVEN LIVE on demo-hp:** the bind moved `/mnt/sys_drive/userdata/media/books` → `/mnt/sys_drive/felhom-data/userdata/media/books`, the capture log went `0 mandatory path(s)` → **`1 mandatory path(s)`**, and **the sentinel is in the snapshot's own file listing** — `-rw-r--r-- 1000 1000 181 … /mnt/sys_drive/felhom-data/userdata/media/books/DRILL-SENTINEL.txt` — not a green status. **Red-proofs:** the bare-path call makes the two paths differ; inverting the drive-kind comparison breaks every enrolled row; leaving the export site bare emits the short path; and the verdict fails under both an unreachable gap-recording and an unconditional `ok`. **One red-proof PASSED and the test was wrong, not the code** — the first Scenario-C test only reached `offboxCaptureSet` while the mutation lives in `runOffboxInternal`; a run-level test replaced it | CC | +| **R-204** | **A rebuilt box can recover its off-site key and still cannot use it: the remedy that reconfigures the tier is the thing that blocks the recovery.** Measured end to end on demo-hp during the 2026-08-04 night drill, after a real controller-data wipe | **OPEN — this is the wall the drill hit** | — | **THE CHAIN, each link measured.** (1) A rebuilt controller **cannot configure its off-site target at all**: `offsite-apply: consume one-time password: no unconsumed offsite password` — the previous controller consumed it (ledger: created `07:11:51`, consumed `07:12:06`). That is R-193, reconfirmed live. (2) The documented remedy is an operator **Re-issue**, which works — and **sets `stale_at` on the escrow** (measured: `2026-08-04 20:15:49`) **while `restic_pw_sha256` is unchanged**, i.e. R-196's false staleness. (3) A stale escrow makes the hub **withhold the hash from the report ACK**, so `EscrowAutoConfirmer` can never flip `pending → escrowed`, and `OffboxRunnable` (`configured && escrowed`) **refuses every off-site run**. (4) The only documented way to clear a stale escrow is **a fresh ceremony — which supersedes the identity blob and destroys the key being recovered.** **So the recovery and its precondition are mutually exclusive as built.** The key comes back (proven — `8a9e33aa4da6…` recovered byte-identical after the wipe, and installed) and then cannot be used to open the repository. **(5) A FOURTH link, undocumented anywhere:** a rebuilt box is **unclaimed**, and the claim gate correctly intercepts every non-claim route (`claim gate: intercepting POST /backup/offbox/confirm-escrow (unclaimed)`), so **no controller endpoint responds at all** until the customer re-claims. That step appears in no design document and comes first. **Fix shape, not decided here:** either the Re-issue stops marking the escrow stale on evidence it does not have (R-196's own fix), or a recovered-and-verified key is allowed to confirm the escrow without a ceremony — the hash comparison that `EscrowAutoConfirmer` already performs is exactly the evidence needed, and it is being withheld precisely when it would be conclusive. **Do NOT fix by widening `OffboxRunnable`** — the atomicity guarantee it enforces (no un-recoverable ciphertext) is the reason the escrow exists. Source: `audits/DRILL-r201-night-run-2026-08-04.md` §3 | CC + operator | | **R-192** | **`offsite_delivery_stuck` tells the operator the opposite of what the detector measured, and the self-heal silently refuses for exactly the reason the message denies.** demo-hp has been e-mailing this daily since 2026-08-03 06:12 UTC: *"one-time password consumed 284h ago and **500 report(s) since carry no offbox target** — the credential is likely burned (apply died between consume and persist). Re-issue delivers a fresh one."* **Measured against the hub's own data: all 500 of those reports DO carry an offbox target** | **HALF SHIPPED** (hub **v0.93.0**) — **the guard's SCOPING stays OPEN** | — | **What actually happened on that box:** the credential was consumed 2026-07-23 09:53:41 and **applied successfully** — the controller reported an `offsite` object continuously until **2026-08-03 05:59:21 UTC**, then it **vanished at 06:12:19** and has been absent for **108 consecutive reports** since. So this is a **regressed apply**, not a burned credential. **Two distinct defects, and the second explains the first's invisibility.** **(a)** `maybeEmitStuck` builds its message from `status.ReportsSinceConsume` (the TOTAL) while hardcoding the phrase *"carry no offbox target"*, and never consults `status.OffsiteReportsSinceConsume` — which is the field that says the opposite. The recommended action (*Re-issue*) is aimed at a failure mode that did not occur. This is R-100's corollary again: an alarm whose text stopped matching what its verdict counts. **(b)** `maybeHeal` refuses **silently** (`OffsiteReportsSinceConsume != 0` → *"regressed-apply shape → operator's call"*, a bare `return` with no log line), so the operator gets a daily e-mail with the wrong story, no heal, and nothing anywhere saying why the heal declined. `offsite_credential_restaged` has **never** fired, on any customer. **The underlying condition is REAL and is the part that matters:** demo-hp currently reports no offsite target at all, i.e. that box's customer app-data has **no off-site copy right now** — and it has been that way since 08:12 CEST on 2026-08-03. A spot check inside the controller container found no restic environment, consistent with the report. **What removed it is not established** and is the first thing to find out. **Fix shape:** the message must state which shape was detected (burned vs regressed) and say what to do for each; the heal's refusal must log its reason; and the regressed shape probably deserves its own event type rather than borrowing the burned one. **Do NOT 'fix' it by widening the heal to restage over a regression** — the guard is right, only mute **CAUSE ESTABLISHED 2026-08-04 (operator confirms no hub-side offsite config change).** The regression is a **guest REBUILD**: at 06:09:40 `host_leaf_changed` (agent re-keyed), at 06:12:18 `controller_started (0.192.0)` — the controller went **0.187.0 → 0.192.0** with a **new config hash** (`1f725a2e843c` → `744e83d72c80`) — and the report at 06:12:19 is the first without `offsite`. The pre-rebuild object was fully healthy: `escrow_state: escrowed, last_status: ok, last_success 2026-08-03T02:16:39Z, snapshot_count 15, repo_size 40.9 MB`. → **R-193** owns the rebuild half. **AND THE HEAL'S GUARD IS WRONG FOR EXACTLY THIS CASE, which is why the automation that exists to fix it declined.** `maybeHeal` refuses when `OffsiteReportsSinceConsume != 0`, reading that as *"the apply regressed, so it is the operator's call"*. But `CountReportsOffsiteSince` counts the **OLDEST 500 reports since the consume** (`ORDER BY id LIMIT 500`) — for demo-hp all 500 predate the rebuild. **Offbox evidence from before a rebuild is not evidence that the credential still works**, so the guard reads healthy history as a reason not to heal a box that demonstrably cannot apply. The fix is to judge on RECENT evidence (e.g. the latest N reports, or evidence after the newest `controller_started`), not on everything since the consume. **SPIKE 2026-08-04 — BOTH HALVES CONFIRMED WITH NUMBERS, still OPEN, still not fixed here** (`audits/SPIKE-offsite-credential-recovery-2026-08-04.md` Q7). The query is quoted at source (`store.go:987`): `... ORDER BY id LIMIT 500` = **the oldest 500**. Reproduced against the live hub DB with demo-hp's real consume anchor `2026-07-23 09:53:41`: the guard sees `total=500, withOffsite=500`, spanning `2026-07-23 09:53:47` → **`2026-07-28 11:17:40`** — the entire evidence set ends **six days before** the 2026-08-03 rebuild. The true window totals are `1174 / 1063` (⇒ 111 without, matching the 111 offsite-less reports). So the e-mail's *"500 report(s) since carry no offbox target"* interpolates `ReportsSinceConsume` while `OffsiteReportsSinceConsume` was **500** — the message states the precise negation of its own measurement. `offsite_credential_restaged` has **never fired for any customer** (zero rows of that type in the DB — checked, not assumed). **A NEW REASON NOT TO FIX THIS IN ISOLATION, from the same spike:** under R-193's Q2 finding a successful auto-restage would have restored demo-hp's TRANSPORT while the box minted a new repo password anyway — the heal can protect the plumbing and **cannot** protect the data, and had it fired on 2026-08-04 both boxes would have looked healthy while their snapshots were orphaned. **That is strictly worse than the current loud failure.** Whatever shape the fix takes must say so in its message. **RECON 2026-08-04 adds two inputs and changes no verdict** (`audits/RECON-offsite-dr-chain-2026-08-04.md`). **(1) The guard's own conclusion — that a fix here protects the plumbing and not the data — is now stronger, not weaker:** even a perfect restage leaves the rebuilt box minting a fresh repo password, and per **R-198** the old one is destroyed by the re-ceremony the box is pushed into. Whatever shape the message takes must say that in the same breath, or it will read as an all-clear. **(2) The recency-bounded discriminator this row asks for has a ready anchor the hub already receives:** the report ACK's `escrow{identity_blob_present, restic_pw_sha256}` moves when a box re-keys, so "evidence since the newest re-key" is computable from data already stored — the same observation R-197 makes, aimed at this guard's time window. **BOTH HONESTY HALVES SHIPPED 2026-08-04 (hub v0.93.0); THE GUARD'S LOGIC IS DELIBERATELY UNTOUCHED.** **(a) The message now describes what was measured.** The one stuck state is reported as the two situations it actually covers — **burned** (`OffsiteReportsSinceConsume == 0`) and **regressed** (> 0, the demo-hp shape) — each stating its own measurement and carrying its own recommendation; the regressed text explicitly WITHDRAWS Re-issue and points at what removes an offbox target (a guest rebuild, R-193). `offsite_reports_since_consume` rides the details for the first time. **(b) Every refusal to self-heal leaves a record** — a `notification_log` row on the operator channel, status `refused`, with its reason (the R-182 suppressed-e-mail precedent), riding the stuck event's 24 h cadence so it sits beside the e-mail it explains rather than accumulating per tick. The two conditions were split into separate branches solely so each can name its own reason; **the set of situations in which the heal fires is byte-for-byte what it was**. **(c) Not in the spec and done anyway, narrowing only:** `offsite_delivery_stuck` and `offsite_credential_restaged` are added to `operatorOnlyEvents`. Neither was ever registered and neither has a `customerMessages` entry — which is not a block — so a customer with a configured recipient was in line for an English e-mail about one-time passwords being *"likely burned"*. Measured live: `notification_log` holds operator rows for demo-hp and no customer rows, which is NOT evidence the leg was blocked (equally consistent with no configured recipient), so the register makes it structural. **WHAT STAYS OPEN, and it is this row now:** `CountReportsOffsiteSince` reads `ORDER BY id LIMIT 500` — the OLDEST 500 reports after the consume — so the counts describe the start of the window, not the present. Its correct shape (recency-bounded, rebuild-aware) depends on the recovery chain that is not yet assembled (R-199/R-200/R-201), so it was NOT fixed here. **The window is named inside the alert text** so the limitation travels with the number instead of being laundered into a confident sentence. **Red-proofs observed:** restoring the single hardcoded sentence fails both message scenarios (the first mutation attempt left the default branch in place and only the burned scenario failed — recorded because a mutation that does not remove every guard is not a red-proof); replacing the regressed branch with a bare `return` fails the refusal record and its cadence test. | CC | | **R-193** | **A guest rebuild silently drops the off-site app-data tier, and nothing restages the credential.** demo-hp was rebuilt on 2026-08-03 (controller 0.187.0 → 0.192.0, new config hash, agent leaf re-keyed at 06:09:40). Before it, the offsite tier was healthy and working — `escrow_state: escrowed`, `last_status: ok`, last success **02:16:39Z that morning**, **15 snapshots, 40.9 MB**. After it: no `offsite` object in any of **108** reports, and **no off-site copy of that customer's app data since 08:12 CEST on 2026-08-03** | **OPEN** — the (c) decision is TAKEN (accept the risk); the chain is still unassembled | — | **Mechanism, fully evidenced.** The restic credential reaches a box exactly once, as a one-time secret. demo-hp's was consumed **2026-07-23 09:53:41**; the rebuilt controller came up with a fresh data volume, no copy of it, and **no way to ask for another** — the hub is the only side that can stage one, and it will not re-stage a consumed secret on its own (the R-71c self-heal would, but it refuses — see R-192). **demo-felhom survived the SAME rebuild by luck, and the contrast is the proof:** its secret was created 2026-07-21 and still **UNCONSUMED**, so when its config hash changed at 07:17:54 and `offsite` dropped for exactly one report, it consumed the staged secret at **07:17:58** and was reporting `offsite` again by 07:19:10. One box had a spare credential staged and recovered in 76 seconds; the other did not and has been unprotected for a day. **That difference was not a design decision — it was an accident of which box happened to have an unconsumed secret lying around.** **Why this is not just "re-issue it":** the remedy (Re-issue) resets the sub-account password via the Hetzner API and, per R-39's record, **rotates the restic password and makes the escrow STALE** — so it needs the recovery-code ceremony re-run, and the continuity of the 15 existing snapshots under the new credential must be VERIFIED, not assumed (`hub v0.60.0` retains superseded escrow, and the orphan guard is move-aside-never-delete). That is an operator act with a customer-facing consequence, so it is not something to fire automatically without deciding the escrow question first. **What to design:** a rebuild is a normal, expected event on these boxes — the offsite tier must survive one, either by the hub restaging automatically when a re-enrolled box reports no offsite (the R-192 guard fix makes this safe), or by the credential being recoverable from escrow at re-bootstrap rather than delivered once and unrecoverable **RESOLVED ON THE BOX 2026-08-04 (operator-authorised).** Re-issue fired through the designed endpoint (`POST /configs/demo-hp/offsite-reissue`, HTTP 303): hub staged a fresh one-time password at **07:11:51**, the box's config hash moved `744e83d7` → `5eee0e42`, R-71a's settle-gate reported **GO** (*at/above floor 0.156.0, we are 0.194.0*), the password was **consumed 15 s later at 07:12:06**, and the controller logged *offsite configured for u629488-sub3@…:/home/felhom-repo* at 07:12:09 — the **same sub-account (275124) and the same repo path**, since Re-issue resets the sub-account password and the one-time password is only the transport credential used once to install the box's own SSH key. **THE ESCROW DID NOT RECOVER BY ITSELF — a correction to this session's own first reading.** `escrow_state` went `pending` → `escrowed` 15 s after the apply and CC inferred an automatic re-escrow; **the operator had run the ceremony**. It needed a human, on BOTH boxes: demo-hp escrowed 07:16:02, demo-felhom (whose offsite re-applied on its own the previous day but whose escrow had been `pending` ever since) escrowed 07:20:28. **A 15-second state change is not evidence of automation** — that is the same class as reading an absent log line as success. **Snapshot continuity is NOT yet established and must not be assumed from the counters:** both boxes report `snapshot_count: 0, repo_size_bytes: 0`, but the run-history keys (`last_run`, `last_status`, `last_success`) are **absent entirely** rather than zeroed — the shape of a controller that has never run an offbox backup in this lifetime, not of an empty repo. demo-hp's pre-rebuild object carried all three plus 15 snapshots / 40.9 MB. ~~**The next scheduled `offbox-backup` (04:15) decides it:** 15+ snapshots ⇒ the repo reattached; 1 ⇒ it started fresh and the old snapshots are orphaned-but-retained.~~ **SPIKE 2026-08-04 SETTLED IT WITHOUT WAITING, AND THE ITEM IS BIGGER THAN FILED** (`audits/SPIKE-offsite-credential-recovery-2026-08-04.md`). **(1) THE ONE-SHOT SECRET IS NOT WHERE THE HARM IS.** Three secrets exist; the one-time password is the *recoverable* one (the operator can reset it at the provider any time) and the box's SFTP key is regenerable by design. The **restic repository password** — the DATA key, which the agent's own source calls *"irreplaceable (unlike the SFTP access key…)"* (`felhom-agent/internal/escrow/identity.go:35-39`) — is the one a rebuild destroys, and **nothing automatic ever restages it**: `WriteOffboxSecrets` mints a fresh 256-bit password whenever `/offbox/repo_password` is absent (`offbox.go:392`), and the only recovery path, `InjectOffboxPassword`, has **exactly one caller in the whole repo** — a web form a human pastes into (`web/offbox_handlers.go:189`). **(2) MEASURED, WITHOUT TOUCHING A BOX:** the hub already stores `restic_pw_sha256` on both the live and the superseded escrow, so the question is a hash comparison. demo-hp `8e03eddf…` → `8a9e33aa…`; demo-felhom `48741892…` → `c60c8bc7…`. **Both boxes minted a new repository password.** **(3) THE CONTRAST IN THIS ROW IS FALSE FOR THE DATA.** demo-felhom's 76-second "lucky" recovery restored **delivery only** — its pre-rebuild object carried **36 snapshots / 1.14 GB** (`repo_size_bytes 1136685919`) and it has reported `snapshot_count: 0` in all 109 reports since, with a changed repo password and **nothing marking its escrow stale for 13 h**. Both boxes lost repository continuity; one loudly, one silently, and **the silent one is worse**. **(4) Q3 IS UNMEASURED AND THE BINARY WAS WRONG.** Neither box could run on 2026-08-04 02:15 UTC (demo-hp had no target at all until 07:15; demo-felhom's target was `escrow_state: pending`, which `OffboxRunnable` blocks) — **the decisive run is 2026-08-05 ~02:15 UTC**. Source predicts a **third** outcome, neither 15 nor 1: same account + same repo path (`u629488-sub3:/home/felhom-repo`, and `repoPath` is a compile-time constant) + new password + `claimed: 1` ⇒ `ensureOffboxRepo` classifies **orphaned** and returns `ErrOffboxOrphaned` — **the run refuses and shows the orphan card**, it does not start a silent fresh history. **Record which of the three actually occurs; a prediction from source is not a measurement.** **(5) CANDIDATE (b) IS NOT IMPLEMENTABLE AS STATED** — the escrow is R-wrapped/zero-knowledge and the hub has no recovery code, so "recoverable from escrow at re-bootstrap" describes a customer-present ceremony, i.e. the manual form that already exists. **(6) CANDIDATE (a) ALREADY EXISTS AND IS WIRED TO THE WRONG EVENT:** `reissueOnReenroll`'s **F3 leg** does exactly this (`api/handler.go:1051-1084`) but sits behind `handleHostEnroll`'s mint-once-reuse short-circuit (`if existing != nil { return }`), and a **guest** rebuild leaves the `hosts` row intact — so F3 is never reached. **(7) A NEW CANDIDATE (c), not previously named and recommended second:** the **agent survives a guest rebuild**, already receives the repo password over the pinned local API (`POST /escrow/stage-secret`) and already writes it to a fixed 0600 path — it merely **wipes** it after the ceremony. Retaining it and serving it back is the only candidate that addresses the irreplaceable secret, and every seam it needs exists. Its cost is one deliberate trade the operator must make (a copy of the data key at rest on the Proxmox host — see D6). **SPIKE RECOMMENDATION: ship the honesty pass (R-196 + R-197) now; then decide (c). Do NOT ship (a) first — it would have hidden this.** **NOT CLOSED — WAITING-ON-OPERATOR for the (c)-vs-accept-it decision, stated at the end of the spike.** **RECON 2026-08-04 — the chain was traced link by link, and the picture is worse than the spike's** (`audits/RECON-offsite-dr-chain-2026-08-04.md`). **(A) THE SPIKE'S CANDIDATE (b) IS OVERTURNED IN PART.** "Recoverable from escrow" is not blocked by zero-knowledge — the *hub* cannot open the blob, the *customer* can, with R. What is genuinely impossible is an **unattended** recovery. A **customer-present** one is a real design, and the operator has now ruled on its shape (below). **(B) THE CHAIN IS NOT ASSEMBLED — eleven links, and the automation stops at four.** Mint → stage → seal → store on the hub are PROVEN-LIVE. Then: the hub's blob-serving endpoints have **no client anywhere** (**R-199**); unsealing's only production caller is a `--selftest` mode reading R from an env var; nothing extracts `restic_repo_password` from the recovered bundle; the injection seam **has no form** (**R-200**); and reopening an existing repo with a recovered password has **never happened** (**R-201**). **(C) THE LOAD-BEARING NEW FACT, and it retires this row's own "hub v0.60.0 retains superseded escrow" premise: the retention does NOT retain the restic repo password** — `host_escrow_superseded` has no `identity_blob` column (**R-198**). Both demo boxes' old passwords were destroyed by the 2026-08-04 re-ceremonies, so the orphaned snapshots are unrecoverable for a **second, independent** reason; keeping R would not have helped. **(D) A FAIL-CLOSED REFUSAL IS IMPLEMENTABLE — this is the most useful thing settled.** The hub already tells every box, on every report ACK, `escrow{identity_blob_present, restic_pw_sha256, created_at}` (`hub/internal/api/handler.go:504-510`) — and the controller **discards it** whenever no offbox target exists (`report/escrow_confirm.go:75-84`). Persisting it (the `ClaimSync` set-only pattern, `report/claim_sync.go:39-53`) and refusing to mint when a blob covers a password we do not have needs **no new hub API and no new secret**. **(E) OPERATOR RULING, 2026-08-04, recorded verbatim:** *"If a node is a fresh install AND the hub has a recovery blob, then the controller should yell that recovery is available, and provide a form for the customer to enter the recovery key. After unlocking the blob, the controller should show what will be recovered before proceeding."* Priced row by row in the recon §9: fresh-install signal **exists** (the mint branch's own `os.Stat`); hub-has-a-blob **exists on the wire, S to persist**; the yell **S**; an **R** form **does not exist** (the UI has only ever *emitted* R) **S**; unsealing must cross agent→controller because the controller image ships no `age` — **M**, one new agent local-API endpoint mirroring `/escrow/ceremony/claim`, plus a narrow hub `GET /hosts//escrow`; the **preview is cheap and read-only** — `restic snapshots --json` + `stats --mode raw-data --json` are already how the box counts snapshots (`offbox.go:1234-1265`), so count, dates, sizes, app tags and paths are all knowable before committing **S**. **Security question put to the operator, not answered:** the form sits behind the dashboard password (bcrypt + CSRF + 5/min lockout); the preview exposes backup cadence and app names; the form is an **oracle** for a stolen R and must fail as generically as `UnwrapIdentity` already does; and R transits the agent, which is the same trade as option (c) in a smaller, time-bounded form. **(F) THE DRILL IS DESIGNED AND NOT RUN** (recon §10): demo-hp, ~3–4 h, R kept deliberately, sentinel file sha256 before and after, **pass = byte-identical sha256, NOT "the repository opened"**, fail = a snapshot count of 1. **Run R-198's fix first.** **(G) Q3 STILL UNMEASURED:** neither box has run since (`last_run` absent on both, 2026-08-04 09:49/09:56 reports) — the decisive run remains **2026-08-05 ~02:15 UTC**. **OPERATOR RULINGS 2026-08-04, and one of them changes what the other items are for.** (1) **Candidate (c) is REFUSED — the risk is accepted:** no repository password is retained on the Proxmox host. **That makes the customer-present recovery path the ONLY way back from a rebuild**, which is why R-198 was shipped the same day as a load-bearing fix rather than a tidy-up: with no host-retained copy, everything runs through the retained identity blob, and until hub v0.93.0 the ceremony destroyed it. (2) **Run the drill, after R-198** — R-198 has shipped, so the drill is the next session (design: `audits/RECON-offsite-dr-chain-2026-08-04.md` §10; tracked as R-201). (3) **Delete the orphaned ciphertext** — ~1.2 GB across the two demo boxes; **STILL OWED**, deliberately not done in v0.93.0 (a destructive act on a protected endpoint does not ride a schema-change release). **What v0.93.0 delivers against this row:** the key now SURVIVES a supersession (R-198) and a changed key is now REPORTED on the day (R-197). **What it does NOT:** the chain that hands the key back is still unassembled at three links — R-199 (no client for the hub's blob-serving endpoints), R-200 (no form for the injection seam), R-201 (never exercised end to end). This row stays open until the drill returns a byte-identical sentinel file. | CC + operator | | — | Storage Box **snapshots** on `storage-box-pool-1` — plan SET (daily 00:00, keep 7) but **0 taken yet** | WATCHING | first run tonight 00:00 | Confirm `size_snapshots > 0` tomorrow; until then the mitigation is armed, not proven | CC | diff --git a/documentation/backlog/ROADMAP.md b/documentation/backlog/ROADMAP.md index d566ced..ef0f3c9 100644 --- a/documentation/backlog/ROADMAP.md +++ b/documentation/backlog/ROADMAP.md @@ -66,6 +66,7 @@ | R-203 | **A mandatory customer data directory was silently absent from the off-site snapshot while the run reported `ok`** — the deploy-time `${USERDATA_PATH}` root and the backup-time namespace root disagree for an app on the system drive | M | **OPEN — halted the R-201 drill 2026-08-04** | Flips nothing yet. Blocks the off-site app-data row from ever earning a customer-file-restored badge. Fix shape: one root function, and a MANDATORY skip must be customer/hub-visible rather than a container-log WARN | | R-201 | **The wipe-and-recover drill** | L | **PREPARED, HALTED BEFORE THE WIPE (2026-08-04)** | Nothing irreversible done. Established live for the first time: a rebuilt box's off-site run REFUSES (the orphan card, not a silent fresh history — closes R-193 Q3), the orphan reset works move-aside-never-delete, and demo-hp's pre-rebuild key is permanently gone (superseded four hours before v0.93.0). Resume needs R-203 | | R-203 | **The app and its backup looked in different directories, and a run that skipped a mandatory folder still said `ok`** | M | **SHIPPED + PROVEN-LIVE (controller v0.197.0, 2026-08-04)** | Adds the capability-map row *"off-site app-data capture covers MANDATORY paths on both drive layouts"* as PROVEN-LIVE, and corrects that row's predecessor, which was optimistic. Unblocks R-201 | +| R-204 | **The recovered key cannot be used: the Re-issue that reconfigures a rebuilt box's off-site tier marks the escrow stale, which gates every run, and the only way to clear it destroys the key** | M | **OPEN — the wall the 2026-08-04 night drill hit** | Blocks R-201's verdict. Fix is most likely R-196's (mark stale on evidence, not on a Re-issue) or letting a verified recovered key confirm the escrow without a ceremony | | R-19 | Internet-outage customer-experience drill: pull WAN on demo, verify lan_resolver path, document what the customer actually sees/does | S | idea | Flips map row E "LAN access" IMPLEMENTED→PROVEN-LIVE | | R-20 | ~~Verify operator-key pinning is fully in the day-0 install flow~~ | XS | **closed** (2026-07-16) | Confirmed against `scripts/felhom-host-install.sh` source (not changelog): keys resolve at L1181–1219 (script constants `OPERATOR_KEY_*`, populated, `--operator-pubkey-file` override), pinned automatically by `step_agent_config()` "STEP 6/8" (L2044; python builds `authz.signers` L2146–2156, reinstall preserves existing), verified at L2332–2337 ("authz signers: N … operator-signed self-update armed"). No interactive prompt or post-install hand-edit — fully automatic. Doc-drift note: the L193–197 "EMPTY by default" comment is stale vs the now-populated constants (→ R-16 hygiene) | | R-23 | **Immediate-sync Direction-2 follow-ups** (hub v0.58 / controller v0.140, 2026-07-16): **(a) — BANKED 2026-07-21 (both legs).** The operator-UI save->apply round trip is PROVEN: the STOP-2 global-floor save (hub `18:56:27 CEST`) released the controller's held wait in the **same second** (`16:56:27Z wait woke: generation=1 - firing out-of-cycle report`), with the report built 2 s later; the ring also shows `wait baseline generation=0` at startup (baseline recorded WITHOUT firing, as designed) then `generation=1`, so the generation advanced past 0. **RESTART LEG BANKED 2026-07-21 — the floor was moved to a version the box did NOT run, and the swap fired EXACTLY ONCE.** Operator saved global floor 0.153.0 → **v0.154.0** (a real version boundary, unlike the 2026-07-20 attempt which targeted an already-running version and therefore proved nothing). Timeline (guest UTC): `06:57:13` UpdateState `pending` written `initiated_by=auto-floor` → `06:57:17` agent `controller swap requested 0.153.0 -> 0.154.0` → `06:57:19` `image file written, restarting bootstrap` → `06:57:21` container StartedAt + UpdateState `completed_at` → `06:57:29` agent `new controller healthy`. **16 s end to end.** Assertions over the whole window (06:50 → 07:29, 39 min): `controller swap requested` = **1**, agent-driven bootstrap restarts = **1**, `new controller healthy` = **1**, rollback/swap-failed/unhealthy = **0**, container `RestartCount` = **0**. `VerifyStartup` banked it on the next boot (`Post-update startup: update successful (0.153.0 → 0.154.0)`) and the `06:57:52` periodic check logged `Current version 0.154.0 is up to date` — the at/above-floor branch correctly doing nothing. **No storm, no rollback, no second attempt.** *Caveat, disclosed: a hand-deploy of v0.155.0 at `07:17:10` sits inside the observation window and is what StartedAt shows after that point; it never touches `SwapController`, so the swap-count assertions above are uncontaminated across the full window. A second, unplanned confirmation of the at/above-floor branch came with it — after the hand-deploy the box ran 0.155.0 against a 0.154.0 floor and the updater logged `Current version 0.155.0 is up to date` and did nothing.* Evidence: `felhom-controller/REPORT.md` §6 (2026-07-21). *(Superseded note:)* **the self-restart single-fire leg was** - the floor was set to a version the box ALREADY ran, so there was no work and no restart. Finish by bumping the floor to a version the box does NOT run, debug ring open, asserting EXACTLY ONE restart. **Trap found while banking this: the wake is `logx.Debugf`, so it is INVISIBLE in `docker logs` at INFO** and lives only in the debug ring (`GET /api/debug/logs?level=DEBUG`) - a hunter looking at stdout wrongly concludes the box never woke; arguably (b) generalised. (b) cosmetic: the Waiter's "recovered" INFO logs on the next hold completion (`pollOnce` blocks ~240 s), not at reconnect | S | **(a) BANKED in full; only (b) cosmetic remains** | Map row "config/state change round-trips in seconds" flipped PARTIAL->PROVEN-LIVE 2026-07-21 on this evidence. Evidence: `felhom-controller/REPORT.md` 4f |