Files
felhom.eu/REPORT-offsite-lock-build-2026-10-03.md
T

7.4 KiB
Raw Blame History

REPORT — off-site backups a box cannot delete: built and live (decisions 68–70) — 2026-10-03 (evening)

Architecture read: 07-backup-architecture.md (custody, threat rows 9–12, §D), 06-offsite-connectivity.md §3/§5, 09 §3. Baselines (re-verified): controller 09453325d1b2 (0.288.0), agent d766666ff8cf (0.138.0), felhom.eu 9268d9933b8f (hub 0.126.0 deployed), catalog 917a779cca67. Register 327 rows, highest R-822. Rulings recorded first as decisions 68–70 (5188dbd). Evidence: documentation/audits/offsite-lock-build-2026-10-03/.

The Part table

Part Result Notes
A — migration spike done, passed An sftp-written repo is listed, extended, restored from (bytes identical), checked and check --read-dataed through the pinned rclone: key with restic 0.14.0; a delete is refused (403). The same measurements also settled several facts: one key on two lines → the first line wins (so the window = prepend a deleting line); an absolute pinned path works; a probe signal (exit 0 + rclone output = pinned; exit 8 = unpinned); the restricted shell's dd/mv/cat (no test).
B — hub registrar + sealed password done — hub v0.127.0 deployed internal/offsitekeys; consume-password → 410; password AES-256-GCM at rest (4 live rows sealed, read back as enc:v1:); daily key check 07:10 + on demand. 3 red-proofs.
C — box on the locked key done — controller v0.289.0, then v0.289.1 registrar client, pinned probe, rclone: transport (hub tier only — the household NAS stays sftp), all four deleting features off the box, window client + fake-snapshot guard. 4 red-proofs. v0.289.1 fixes a defect v0.289.0 put live (below).
D — live, both demo boxes done demo-felhom 11→13, demo-hp 91→100 (history kept); a delete from each box refused (403), count unchanged; one-file restore and check through the pinned key on each; the old endpoint answers 410 to demo-hp's own key. The hand-run key check is clean for both. Changed: the "unprefixed test line on a demo sub-account" decoy ran on tester-1's account instead, which already held 3 unpinned lines. The demo passwords are now sealed in the hub, and the hub DB is the only route to them.
— stop point passed
E — window + guard mechanics done; a real prune NOT done Window 1 on demo-hp ran live: the hub opened it (deleting line first), the guard refused, the window closed in 3 s, the operator was mailed, and the key file read back clean. The refusal is a design defect (R-824): any manual run makes the plan remove a same-day snapshot younger than 8 days. Weekly windows stay OFF.
F — ep0 → DooPlex copy done ep0: one read-only token (the only change there). DooPlex: an SSH forward (felhom-ep0-pbs-tunnel.service — operator ruling in-session, because ep0's PBS listens on wg0 only), PBS remote, datastore ep0-copy, nightly pull 05:00 with remove-vanished false, Saturday verify, failures to admin@ via Resend (the test mail arrived). First pull: 201 s, 12 GB, 4 of 4 snapshots, matching ep0. Runbook: runbooks/ep0-datastore-copy.md.
Golden + floor done Golden 0.289.1 baked (subagent, runbook §4.1, token-leak grep 0 with a positive control), round-trip sha matches, vouched (agent 0.138.0, min_agent 0.131.0), floor 0.289.1 SERVED to both boxes.

Claims in the brief that turned out wrong

  1. "An sftp-written repo reads through rclone" — confirmed: it was expected, and now it is measured.
  2. "The box stores no password today" — true on disk (it was only in an env var during install), but the box could fetch the password at will; that route is closed now.
  3. "One authorized_keys can hold two lines for the same key" — it can, but only the first line counts. That is what makes the window possible with a single key.
  4. "The integrity check works through the forced key" — confirmed (check, check --read-data, the exclusive lock).
  5. "DooPlex has room and a PBS that can pull from ep0" — it has the room (5.5 TB) and a PBS, but it cannot reach ep0's PBS (open on wg0 only). An SSH forward was added on an operator ruling.
  6. "Every deleting feature leaves the box" — only on the hub tier. The same code serves the household's own SFTP NAS, which keeps box-side retention. A Transport flag separates the two.
  7. "Abort when the plan exceeds a week's removal" — that would never prune after the interim. The box takes the oldest snapshots up to the cap instead (disagreement recorded in the code and the CHANGELOG).
  8. The guard as written is too strict (R-824). It also cannot see past-dated poisoning (R-822, residual).

Found and fixed in-session

  • R-825 (v0.289.1): the provider's rclone prints a NOTICE line on every connection, and restic forwards it into the output. Every --json parse failed, so demo-felhom recorded 0 snapshots as measured, and the hub mailed a false offsite_snapshots_dropped (11→0) at 17:17. The fix went live 15 min later. It strips the notice, and an unreadable count is never a measured zero. Red-proved.
  • A shadowed newPath in the NAS move-aside path was caught by the existing suite before release.
  • Stale unpinned keys sat in the sub-accounts: 4 on demo-felhom's and 5 on demo-hp's (every reinstall added one). The registrar removed them. tester-1's 3 remain, and the daily check alarms on them (R-826).

Deviations, stated

  • Two controller releases, against the one-release rule: v0.289.1 fixes a false zero that v0.289.0 put live.
  • The hub has a test-only commit after the release (window-sweep test + a test helper in the store). The deployed v0.127.0 image does not contain it; behaviour is unchanged.
  • I read tester-1's sealed-era password from the hub DB again for Part A (operator ruling from the morning session; the copy was deleted, the value never printed).
  • A Hetzner storage API token was printed into this session's transcript while I read manifests/storagebox.secret.yaml (a gitignored file; the redaction regex missed the quoted value). Rotate HETZNER_TOKEN — it is in Secret/storagebox.
  • The window's red-proofs ran in unit tests and on one live window. A live real prune did not happen (R-824).

Records

  • Closed: R-820, R-821, R-342 (+ R-825 opened and closed). Narrowed: R-95, R-822. Opened: R-823, R-824, R-826, R-827, R-828, R-830. Register 327 → 330.
  • Decisions 68–70 in 09 §3, CONTEXT, 07, 06. 07 threat rows 9/10/12 and 06 §3.6 carry [FACT] lines.
  • Runbooks: ep0-datastore-copy.md (new), secrets.md (offsite key, DooPlex PBS secrets), RUNBOOK-manual-build.md (pveam update).

Teardown, three layers

  • Machines: helper scripts were removed from both demo guests, their containers and hosts (0 left). tester-1's sub-account is back to .ssh, felhom-repo, with authorized_keys byte-identical to the start (sha256 795e7153…). The scratch dirs spike-r436, spike-migrate and the rclone .config were removed. The drill VM is reverted to virgin, build guest 9100 destroyed, the bake token shredded.
  • Host (DooPlex): kept on purpose: felhom-ep0-pbs-tunnel.service, the PBS remote/datastore/jobs/notification target (Part F). The scratchpad secret files (sub4 password, ep0 token, Resend key copy) were shredded.
  • Hub: kept: v0.127.0, Secret/offsite-secret-key, floor 0.289.1, golden 0.289.1 vouched. Weekly windows OFF. The one-shot grant for demo-hp was consumed.