Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
7.1 KiB
R-32 — RESET leaves the household's off-site ciphertext behind — design proposal (burn-down night 2026-10-06, no code)
Baselines read: felhom.eu 8e2dc204 (hub v0.140.0), felhom-controller 5e7522e023 (v0.301.0). Architecture:
07-backup-architecture.md §246 (off-site deletion custody, decisions 68–69); 06-offsite-connectivity.md; the
row's own three-part ruling from the 2026-07-21 rehearsal.
1. The problem
RESET is meant to destroy a household's off-site copy. On the shared pool box it deletes the Hetzner SUB-ACCOUNT. A
sub-account is a login, not the data: its home directory stays. Re-enabling off-site for the same customer creates a
sub-account with the SAME home (felhom-<customer id>) over the old ciphertext, whose key that RESET destroyed. Measured
on the pool box the night of 2026-07-21: 49 MB attributed (2 snapshots, 48.7 MiB) against 1.4 GB + 3.0 MB
unattributed in two .orphaned-* folders (restic-and-pool.txt, R-32). The orphan card then appeared — the
rehearsal's S7 had said in advance that it would be a finding.
2. What the code does today (read in source)
- RESET's off-site leg calls
s.offsite.Deprovision(hub/internal/web/customer_reset.go:217-225) and journalshetzner: ok, logging "repo data destroyed" (:225). Deprovision, shared tier:DeleteSubaccountper labelled sub-account, nothing else (hub/internal/offsite/offsite.go:303-326). Its doc comment says "The offsite repo DATA dies with the sub-account/box" (:276-280) — true for the dedicated tier (the box is deleted,:283-300), false for the shared one. A comment that asserts an invariant the code does not provide.- The home directory is fixed per customer:
HomeDirectory: "felhom-" + customerID(offsite.go:349). So a new lifecycle lands on the old folder. - The hub already deletes off-site data in ONE place, through the sub-account's own password login (port 23):
Registrar.DeleteSetAside(hub/internal/offsitekeys/offsitekeys.go:424-445) —rm -rfof a<repo>.orphaned-…folder only, refusing anything else (IsSetAsidePath,:448-459), after the household's 7-day abandonment delay (decision 74,service.go„Decision 74"). - The move-aside for a reinstall WITHOUT RESET: the box asks the hub to rename the old repo to
<repo>.orphaned-<date>(felhom-controller/controller/internal/backup/offbox.go:329-345). The ruling keeps this — custody survives there. - The operator's Restic tab shows the pool box's totals from the provider API and, per customer, the usage the BOX
reports (
hub/internal/web/offsite_box.go:121-185). Bytes that no box reports (an.orphaned-*folder, an old lifecycle) are visible only in the pool total, not per customer.
3. Options
A. Purge through the sub-account, then delete it. In Deprovision (shared), before DeleteSubaccount: log in with
the hub's stored password for that sub-account (it already does this for authorized_keys and DeleteSetAside),
remove the live repo and every <repo>.orphaned-*, check the home holds no repository left, then delete the
sub-account. A purge that fails stops the leg (hetzner: failed, re-run resumes) — the sub-account is NOT deleted,
because after that only the main account can reach the folder.
- Costs: hub only, small; one new registrar method (
PurgeRepos) besideDeleteSetAside, same refusals style. - No new credential: the main-account password the ruling named is not needed.
- Can go wrong: the stored password no longer works (rotated, or the hub DB restored from an older copy) → the leg fails loudly and the operator must decide (main-account clean-up by hand). Deletes household data — that is RESET's purpose, and the existing RESET ack covers it per the ruling; still the operator's word on the route.
B. Purge with the pool box's MAIN account (the ruling's words). The hub gets the main-account SFTP password.
- Costs: a new credential that reaches every household's folder on the pool box. A hub compromise then deletes every household's history in one step. Bigger blast radius than A for the same result.
- Only advantage: also reaches folders of sub-accounts deleted BEFORE this fix (old lifecycles).
C. A new home folder per lifecycle (felhom-<id>-<n>), no purge. The new sub-account never sees old ciphertext,
so the orphan card cannot lie.
- Costs: small. But nothing is ever deleted: the ruling's part (1) is not met, and dead ciphertext fills the pool box for ever, unseen unless part (3) is built.
Part (3), the byte view, for every option: per customer, the bytes in its folder (all repos, set-aside copies
included) beside the bytes its box attributes. Needs a per-folder size from the provider. Not measured: whether the
password login on port 23 answers du -s (the shell has rm, mv, dd — memory storagebox-subaccount-shell…). So
not buildable tonight: it rests on a mechanism nobody has measured, and tonight's fence forbids touching the Storage Box.
4. The pick — PROPOSAL for the operator, not a decision
A, with part (3) after one read-only measurement. It meets the ruling's part (1) with no new credential, keeps the move-aside guard for reinstall-without-RESET (part 2 — untouched), and makes the journal's "repo data destroyed" true. Old lifecycles' folders (if any are left on the pool box) are a one-time clean-up by hand — the operator's (question 2).
5. First slice and its proof
- Build (hub):
offsitekeys.Registrar.PurgeRepos(ctx, t, pw)— removest.RepoPathand everyIsSetAsidePathmatch, nothing else; then lists and refuses success if any remains.offsite.Provisioner.Deprovision(shared) calls it through a seam BEFOREDeleteSubaccount. Fix the doc comment atoffsite.go:276-280. - Red test first (must FAIL today): a fake provider API and a fake shell; RESET's
Deprovisionfor a shared customer. Assert the shell sawrm -rf <repo>before the API sawDeleteSubaccount. Today it fails: no shell call at all. - Second red test: the purge fails →
Deprovisionreturns an error andDeleteSubaccountwas NOT called (else the folder becomes unreachable). - Keep green:
DeleteSetAside's refusals; the dedicated tier unchanged; RESET's journal re-run. - Live proof on a SCRATCH customer only (never a household): enable off-site, one run from scratch box 9202, RESET,
re-enable. Positive observable: the new sub-account's home lists no
restic/.orphaned-*folder, and no orphan card appears. Control from a different channel: the provider's pool-boxstatssize before and after (coarse, API). Evidence off the machine before teardown.
6. Open questions for the operator
- RESET deletes the household's off-site folder through the sub-account's own login before removing it (option A), instead of through the pool box's main account — agree? If you do nothing: RESET keeps leaving the ciphertext; a re-enabled customer sees an orphan card again.
- Folders left by RESETs done before this fix (the 2026-07-21 measurement found 1.4 GB): clean them up once by hand with the main account, or leave them? If you do nothing: they stay and use pool-box space; nothing reads them.