Files
felhom.eu/documentation/audits/night-burndown-2026-10-06/design-R-366.md
T
2026-10-06 20:46:37 +02:00

6.6 KiB

R-366 — a reinstall orphans the whole-guest off-site archives — design proposal + first slice built (night 2026-10-06)

Baselines read: felhom.eu 8e2dc204 (hub v0.140.0), felhom-agent 74b5eae. Architecture: 07-backup-architecture.md §5 (encryption, „PBS … per-customer encryption-key") and §6.1 (off-site tier: server-side prune keep-last 2); 06-offsite-connectivity.md §3.5 (key custody: the escrow wraps the PBS key K under the recovery code R).

1. The problem, and what is still true today

  • Measured 2026-08-21 (hub event 3016): after demo-hp's reinstall the restore test failed on a pre-reinstall archive with wrong key … manifest's key 3f:4f:65:c0… does not match provided key dd:d1:d8:53…. Read as „a restore test failed", not as „every whole-guest archive from before the reinstall is unreadable here".
  • The wrong verdict is already gone. Agent v0.138.0 (R-727, decision 51) skips any archive written with another key before it picks a restore-test candidate (felhom-agent/internal/backup/runner.go:360-374). The skip is one INFO line per archive (runner.go:684-698). So the hub no longer sees a false failure — and now sees nothing.
  • The August archives no longer exist (inferred, not measured — ep0 is fenced tonight). The off-site tier is pruned on ep0 with keep-last 2 per group (07 §6.1 table); a reinstalled box writes the same group ct/9201 in the same namespace (measured for R-727 on 2026-09-30), and demo-hp has written weekly copies since 2026-08-21. So the old archives were pruned in early September. Question 1 of the row („are they recoverable?") is moot for demo-hp.
  • The real remaining gap, found tonight (read in source, then pinned by a test): the hub keeps an old escrow only when its restic password differs (hub/internal/store/store.go SaveHostEscrow, the rule before this fix). A reinstall mints a new K (felhom-agent/configs/felhom-pbs-apply:99, --encryption-key autogen). Since R-241 a rebuilt box keeps its restic password (memory retained-key-is-operator-only; not re-measured tonight), so the escrow PUT carries the same restic sha with a new key fingerprint — and the row holding the old K was overwritten. That made every pre-reinstall archive unopenable for good, not only for the new box. Red test observed: r366/red-key-change.txt („a new backup key with the same restic password must supersede (retain the old K)").

2. What the code does today (after the slice below)

  • Agent: foreign-key archives skipped from restore-test candidacy, logged by name (runner.go:360-374, 684-698).
  • Hub, on branch night-r366 (70b07fdf, not on main): SaveHostEscrow retains the current row when the restic sha or the key fingerprint changes (backupKeyChanged: case and space ignored; an empty side is unknown, never a change). The audit event escrow_superseded and the log line now name both causes. The repo-key-changed alarm (maybeEmitRepoKeyChanged) still fires only on a restic change — a K-only change does not raise it.
  • Nobody is told that „the previous install's whole-guest archives exist and this box cannot open them" in the ~14 days before ep0's prune removes them.

3. Options

A. Retain the old K (built tonight, slice 1) and stop there. One condition and two tests in the hub. Cost: one more retained row per reinstall (opaque, R-wrapped — the hub learns nothing). Wrong case: a producer that sends a different fingerprint format for the same key would add a retained row per ceremony; case/space are normalised, and today one producer (the agent's escrow PUT) sends it. Restoring from the old K stays operator-only (R-304).

B. A + tell the operator. The agent counts foreign-key archives per tier in the host report (a new additive field); the hub shows one line on the host page and raises one info operator event per new key: „N whole-guest archives on felhom-pbs were written with key X by a previous install; this box cannot open them; the hub retains key X: yes/no; ep0 prunes them after two new copies." Cost: agent + hub, an additive report field (wire contract gate), one event type (allow-list). Wrong case: noise on a returning customer every reinstall — once per key, so bounded.

C. Key continuity: a reinstall re-uses the escrowed K instead of autogen. The new box would read its own history. Cost: the reinstall then needs R at install time (the customer's code) — the DR consume path, made a normal path. Changes a promise (what a reinstall does with old backups) and custody → the operator's.

4. The pick — PROPOSAL

B, built in two slices; slice 1 (A) is done and waits on its branch. A closes the only irreversible part (the old key destroyed). B removes the silence without changing any promise. C is a product decision and stays a question.

5. First slice and its proof

  • Built: hub 70b07fdf on night-r366 — TestSaveHostEscrow_R366_NewBackupKeySameResticPasswordRetainsOldKey (red observed on the old rule) and TestSaveHostEscrow_R366_SameKeyOrUnknownFingerprintDoesNotSupersede (idempotence: same key in another case, or an empty fingerprint, retains nothing). Full hub go test ./... green. felhom.eu repo_gates.py --fast could not judge the worktree (sibling clones absent next to it — „controller clone not found at /mnt/5_hdd/felhom.eu/wt/…"); it must run on main after the cherry-pick.
  • Live proof (next day, not tonight): on scratch box 9202 only — re-run the escrow PUT with the same restic sha and a new fingerprint through the agent's ceremony (or a reinstall of 9202), then read the host page's „N superseded escrow blobs retained" (positive observable) and, as the control from another channel, the hub log line superseded an escrow with a different passphrase or backup key (R-366).
  • Slice 2 (B): red test first — a host report with one foreign-key archive on felhom-pbs must produce exactly one operator event naming the key, and a second identical report none.

6. Questions for the operator

  1. Deploy slice 1 (the hub keeps the old backup key after a reinstall)? My pick: yes, in tonight's or tomorrow's hub release. If you do nothing: the next reinstall of a box destroys the only key to its earlier whole-guest backups.
  2. Should a reinstall keep reading its old whole-guest backups (option C — it would need the household's recovery code during the reinstall)? If you do nothing: a reinstalled box starts a new off-site history; the old one is opened only by the operator, by hand, and only until ep0 prunes it (about two weeks).