Files
felhom.eu/REPORT-offsite-finish-2026-10-04.md

37 lines
5.0 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# REPORT — off-site safety finished (decisions 71–74) — 2026-10-04
Architecture: `07-backup-architecture.md` (custody block, threat rows 9/10), `06` §3.6, `09` §3 decisions 68–74.
Baselines (re-verified): controller `c4bf7306371a` (0.289.1), agent `d766666ff8cf`, felhom.eu `710a2505f9b4` (hub
0.127.0). Register 330, highest R-830. Rulings recorded first (decisions 71–73, R-831, R-832: `697c2a7`).
Evidence: `documentation/audits/offsite-finish-2026-10-04/`.
## The Part table
| Part | Result | Notes |
|---|---|---|
| A — guard fixed, one real window | **done; a window that removes something NOT yet observed** | Controller v0.290.0: a young snapshot superseded the same day is excluded instead of refusing; future-dated / newer-than-hub / above-the-week's-cap still refuse. 3 red-proofs (the demo-hp shape runs; the 13 future fakes refused; above-cap refused). Live: window 2 on demo-hp opened, guard ran with no refusal, closed in 3 s, **127→127 removed nothing**, because every candidate was a young same-day copy. The key file is clean and the hub's before/after check is quiet. Weekly windows **ON** (fleet switch). Next scheduled windows: demo-felhom at its next night run (never had one); demo-hp at the first night run after 2026-10-10 17:36 UTC. The first real removals are expected around 2026-10-11. |
| B — copy keeps 8 weekly | **done** | `prune-ep0-copy` keep-weekly 8, all namespaces, daily 07:30 (sync 05:00 — ran OK today); GC Sundays 08:30; `remove-vanished` false. Dry-run **by reasoning**, because PBS has no CLI dry-run for a prune job: nothing to remove (2 snapshots per group, 2 different weeks). 12 GB used. |
| C — tester-1's keys | **done** | Through the hub (`POST /offsite/remove-unpinned/tester-1` → `changed: 3`); the check reads 0 lines and raises no alarm. |
| D — restore from the copy | **done** | demo-hp, scratch VMID 9299 on `nvme-scratch`: list 2 s, restore **186 s** (15 GB logical, 14 GB on disk), data read by `pct mount` (not started — starting it would run a second demo-hp controller against the hub). Torn down: VMID, storage entry, DooPlex temporary token + ACL. **Trap found → R-834.** |
| E — set-aside deletion via the hub | **done** | Hub v0.128.0 + controller v0.290.0. Red-proofs: no deletion before the delay; a cancelled request deletes nothing; a recovery that does not cancel at the hub fails its test. Live on tester-1: naming the live repo was refused; a planted set-aside dir was deleted after the delay and read back absent. The delay was shortened to 3 min for that test only, by manifest config, logged at start-up, and reverted (the new pod logs no override). |
| F — releases, floor, golden | **done** | Hub v0.128.0 deployed. Controller v0.290.0 on both demo boxes. Golden 0.290.0 baked (subagent), round-trip sha matches, leak grep 0 with a control, vouched (agent 0.138.0, min_agent 0.131.0). Floor 0.290.0 SERVED. Golden gate OK. |
## Claims in the brief that turned out wrong
1. **"The young superseded copies are the only cause of the refusal"** — right for 2026-10-03. But the brief's own "refuse above the weekly cap" would also have refused every honest window: the hub's cap was 40% and an honest week removes about 41%. I raised the hub cap to half (red-proved), and the cap refusal can still wedge after a long gap (R-833).
2. **"PBS can prune `ep0-copy` without touching the sync"** — true. They are separate jobs at separate times. PBS has no dry-run for a prune JOB, so the dry-run was done by reasoning.
3. **"A demo box's whole-guest backup is in the copy and restorable with its own key"** — true (demo-hp's own `felhom-pbs.enc`). Two things the brief did not expect: `pvesm add pbs` without `--password` fails, and on failure it **deletes** the key files you placed; and the restored config is the production one (`onboot: 1`, the real drive binds) — R-834.
4. **"The household's page still names a deletion date"** — it did, during the countdown. After the date (since v0.289) the page showed nothing while nothing was deleted. It now shows the hub's date, and that date is true.
5. A live window that removes snapshots could not be shown today. Nothing was old enough. I did not fake the history.
## Rows
Closed: **R-823, R-824, R-826, R-827, R-828, R-830**. Opened: **R-831** (the token, waiting on the operator), **R-832**
(roadmap P4), **R-833**, **R-834**. R-95 narrowed further. Register **330 → 328**.
## Teardown, three layers
- **Machines:** demo-hp has no VMID 9299, no `tmp-dooplex-copy` entry, and only its own `felhom-pbs.*` priv files. tester-1's sub-account holds `.ssh` (an empty key file) and `felhom-repo`; the planted dir was deleted by the hub. Helper scripts were removed from demo-hp. The drill VM is back on `virgin`.
- **Host (DooPlex):** **kept on purpose:** the prune job and GC schedule (Part B). Removed: the temporary restore token and its ACL. Shredded in the scratchpad: tester-1's password, its API key, the seal key copy, the restore token.
- **Hub:** v0.128.0 at the 7-day delay (the test override was reverted). Weekly windows ON. Floor 0.290.0, golden 0.290.0 vouched.