Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
6.6 KiB
R-366 — a reinstall orphans the whole-guest off-site archives — design proposal + first slice built (night 2026-10-06)
Baselines read: felhom.eu 8e2dc204 (hub v0.140.0), felhom-agent 74b5eae. Architecture: 07-backup-architecture.md
§5 (encryption, „PBS … per-customer encryption-key") and §6.1 (off-site tier: server-side prune keep-last 2);
06-offsite-connectivity.md §3.5 (key custody: the escrow wraps the PBS key K under the recovery code R).
1. The problem, and what is still true today
- Measured 2026-08-21 (hub event 3016): after demo-hp's reinstall the restore test failed on a pre-reinstall archive
with
wrong key … manifest's key 3f:4f:65:c0… does not match provided key dd:d1:d8:53…. Read as „a restore test failed", not as „every whole-guest archive from before the reinstall is unreadable here". - The wrong verdict is already gone. Agent v0.138.0 (R-727, decision 51) skips any archive written with another key
before it picks a restore-test candidate (
felhom-agent/internal/backup/runner.go:360-374). The skip is one INFO line per archive (runner.go:684-698). So the hub no longer sees a false failure — and now sees nothing. - The August archives no longer exist (inferred, not measured — ep0 is fenced tonight). The off-site tier is pruned
on ep0 with
keep-last 2per group (07§6.1 table); a reinstalled box writes the same groupct/9201in the same namespace (measured for R-727 on 2026-09-30), and demo-hp has written weekly copies since 2026-08-21. So the old archives were pruned in early September. Question 1 of the row („are they recoverable?") is moot for demo-hp. - The real remaining gap, found tonight (read in source, then pinned by a test): the hub keeps an old escrow only
when its restic password differs (
hub/internal/store/store.goSaveHostEscrow, the rule before this fix). A reinstall mints a new K (felhom-agent/configs/felhom-pbs-apply:99,--encryption-key autogen). Since R-241 a rebuilt box keeps its restic password (memoryretained-key-is-operator-only; not re-measured tonight), so the escrow PUT carries the same restic sha with a new key fingerprint — and the row holding the old K was overwritten. That made every pre-reinstall archive unopenable for good, not only for the new box. Red test observed:r366/red-key-change.txt(„a new backup key with the same restic password must supersede (retain the old K)").
2. What the code does today (after the slice below)
- Agent: foreign-key archives skipped from restore-test candidacy, logged by name (
runner.go:360-374, 684-698). - Hub, on branch
night-r366(70b07fdf, not on main):SaveHostEscrowretains the current row when the restic sha or the key fingerprint changes (backupKeyChanged: case and space ignored; an empty side is unknown, never a change). The audit eventescrow_supersededand the log line now name both causes. The repo-key-changed alarm (maybeEmitRepoKeyChanged) still fires only on a restic change — a K-only change does not raise it. - Nobody is told that „the previous install's whole-guest archives exist and this box cannot open them" in the ~14 days before ep0's prune removes them.
3. Options
A. Retain the old K (built tonight, slice 1) and stop there. One condition and two tests in the hub. Cost: one more retained row per reinstall (opaque, R-wrapped — the hub learns nothing). Wrong case: a producer that sends a different fingerprint format for the same key would add a retained row per ceremony; case/space are normalised, and today one producer (the agent's escrow PUT) sends it. Restoring from the old K stays operator-only (R-304).
B. A + tell the operator. The agent counts foreign-key archives per tier in the host report (a new additive field);
the hub shows one line on the host page and raises one info operator event per new key: „N whole-guest archives on
felhom-pbs were written with key X by a previous install; this box cannot open them; the hub retains key X: yes/no;
ep0 prunes them after two new copies." Cost: agent + hub, an additive report field (wire contract gate), one event
type (allow-list). Wrong case: noise on a returning customer every reinstall — once per key, so bounded.
C. Key continuity: a reinstall re-uses the escrowed K instead of autogen. The new box would read its own history.
Cost: the reinstall then needs R at install time (the customer's code) — the DR consume path, made a normal path. Changes
a promise (what a reinstall does with old backups) and custody → the operator's.
4. The pick — PROPOSAL
B, built in two slices; slice 1 (A) is done and waits on its branch. A closes the only irreversible part (the old key destroyed). B removes the silence without changing any promise. C is a product decision and stays a question.
5. First slice and its proof
- Built: hub
70b07fdfonnight-r366—TestSaveHostEscrow_R366_NewBackupKeySameResticPasswordRetainsOldKey(red observed on the old rule) andTestSaveHostEscrow_R366_SameKeyOrUnknownFingerprintDoesNotSupersede(idempotence: same key in another case, or an empty fingerprint, retains nothing). Full hubgo test ./...green. felhom.eurepo_gates.py --fastcould not judge the worktree (sibling clones absent next to it — „controller clone not found at /mnt/5_hdd/felhom.eu/wt/…"); it must run onmainafter the cherry-pick. - Live proof (next day, not tonight): on scratch box 9202 only — re-run the escrow PUT with the same restic sha and
a new fingerprint through the agent's ceremony (or a reinstall of 9202), then read the host page's „N superseded
escrow blobs retained" (positive observable) and, as the control from another channel, the hub log line
superseded an escrow with a different passphrase or backup key (R-366). - Slice 2 (B): red test first — a host report with one foreign-key archive on felhom-pbs must produce exactly one operator event naming the key, and a second identical report none.
6. Questions for the operator
- Deploy slice 1 (the hub keeps the old backup key after a reinstall)? My pick: yes, in tonight's or tomorrow's hub release. If you do nothing: the next reinstall of a box destroys the only key to its earlier whole-guest backups.
- Should a reinstall keep reading its old whole-guest backups (option C — it would need the household's recovery code during the reinstall)? If you do nothing: a reinstalled box starts a new off-site history; the old one is opened only by the operator, by hand, and only until ep0 prunes it (about two weeks).