Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
7.5 KiB
REPORT — the night's fixes, the power cut by day, a box that is off at night (2026-10-05)
Brief: "fix what the 2026-10-04 night found …" (operator, 2026-10-05). Evidence: documentation/audits/night-fixes-2026-10-05/.
Architecture read before the claims: 07-backup-architecture.md (§6.1, the Tier-3 row), 11-os-updates.md (§5.4.1, §8),
09 §3 (decisions 53, 68–74, 81, 86), audits/offsite-append-only-2026-10-03/DESIGN.md.
1. The Part table
| Part | State | Note |
|---|---|---|
| §1 rulings recorded first | done | 09 decisions 100–103; R-870 (tokens, like R-831); Tester 2's page in runbooks/target-selection.md |
| A — clean-up guard (R-867) | done | controller v0.294.0; tests run restic 0.14.0's policy (proven identical to the binary); 5 red-proofs; live windows on both demo boxes with counts; R-867 and R-95 closed (DUE-CHECKS entry removed) |
| B — image clean-up race (R-863, R-864) | done | one lock for every pulling compose verb; the night's shape reproduced in a test and live on 9202 (fired at minute 3, skipped, install OK first time) |
| C — R8 real download (R-865) | changed | fix + tests + red-proof done. Live: the INSTALLED wrapper's measure read 12 802 456 B on demo-hp. No R8 refusal line was produced: it needs < 500 MB free on demo-hp's guest = 28.5 GB written into a thin pool with 18.7 GB free — it would have stopped every guest |
| D — R-868, R-869, R-866 | done, with a second agent release | R-868's first fix (v0.144.0) was measured NOT to work live (broken stderr pipe); v0.144.1 fixed it and the A5 shape then delivered one applied report. R-866 live with the hub blackholed. R-869 test + red-proof |
| E — power cut mid-update (A1 by day) | done | operator's go; crash 06:13:55 UTC, back by itself in 37 s; the next pass failed → R-876 (P2); by-hand repair; demo-hp left with every package current |
| F — a box off at night (spike) | done | read only; partF/FINDINGS.md; no architecture covers it → R-871; R-872..R-874; STATUS decision A/B |
| G — releases, golden, records | done | controller v0.294.0; agent v0.144.0 + v0.144.1 (decision 108); golden 0.294.0 baked + vouched (gate OK); 07, 11, 00, 09 updated |
2. Claims in the brief that turned out wrong (or only partly true)
- "The policy's constants can drive the guard's line" — true, and done — but with a cost the brief did not name: a line derived from keep-daily (calendar days) can no longer catch a past-dated gap-fill that steers a 7–8-day-old keep, which the 8-day line caught while refusing every honest window. Past-dated fakes can never make the policy drop a snapshot inside the last keep-daily calendar days, so the line now catches only the skew-window shape. The rest is R-822's residual, bounded by the cap and the hub's count check (decision 104).
- "The same race exists in the update and restore paths" — the update path: NO, it was already guarded (no pass
while any app is updating, since v0.284). The restore and undo paths: exposed in principle (they pull inside
compose up), not measured; they now hold the same lock. - "Dropping
-sdownloads nothing" — TRUE, measured on 9202 (archive cache 10 → 10 files, versions unchanged). - "A box off at W runs nothing when it comes back" — mostly true: the database dumps, second copy, off-site copy and
app updates never catch up; but the whole-guest backup DOES catch up (48 h safety valve, about every 2 days for an
evening-only box), and the OS leg follows it by day under the
nightlabel. - "The night A5 showed the wrapper finished" (R-868's premise, and v0.144.0's design) — the packages were installed, but the wrapper process itself died on its first log line after the agent's death. Found only by running it live.
- Part E: "the next pass must repair dpkg first and finish" — it does NOT (R-876).
3. What was proven, with numbers
- Part A. Tests: the measured shape (7 d + 5 s → allowed), 60 nights × 3 apps with two same-day manual runs across a
month boundary (never refused, ≥ 45 removed), weekly windows, the skew-window fake (refused), the lab's 13 future fakes
(refused). Red-proofs: 8-day line back → 3 fail; each refusal dropped → its test fails. Lab: restic 0.14.0 binary vs
the test simulation over 92 snapshots: identical 72 removals. Live: demo-felhom 15 snapshots, predicted removals
92e49f62,343d57a5→ window 5 removed exactly those (16 → 14 incl. the run's new one); demo-hp 136 snapshots, predicted 18 → window 6 removed exactly those (145 → 127). Hub rowspruned, no event, no mail, key audit 0 findings. - Part B. Live on 9202: controller start 05:28:35, install 05:31:15, clean-up 05:31:36 "skipped — … pulling images now", BookStack deployed 05:31:54 (35.5 s), retry 05:33:36 ran (0 deleted). Teardown through the product verified.
- Part C.
download_bytes = 12802456 Bfrom the installed wrapper on demo-hp (was 0). - Part D. R-866:
block=SAVED(2026-10-05T05:23:26Z; hub unreachable: … connect: invalid argument). R-868 (v0.144.1): killed 06:04:00, wrapperDONE upgraded=13, kept copy, hub report id 63applied(13) at 06:09:05, once. - Part E. See
11§8.4: back in 37 s; apps healthy within 4 min; household saw one timeline line; one operator mailos_update_failed(true); next pass FAILED (R-876); afterdpkg --configure -athe pass installed 12; the guest's package list equals the pre-test one (279 lines). The first crash attempt did not fire (my watcher's pattern missed apt's dpkg call; the pass installed normally) — rolled back again and repeated. - Part F.
audits/night-fixes-2026-10-05/partF/FINDINGS.md; two claims re-checked in source by me.
4. Rows
Register before 341, after 340. Closed (8): R-95, R-863, R-864, R-865, R-866, R-867, R-868, R-869. Opened (7):
R-870 (tokens, operator), R-871 (a box off at night — the operator's decision), R-872 (P2, no missed-backup alarm for a
box down at 05:00), R-873 (nightly "cannot be reached" mail), R-874 (restore-test never runs on short sessions),
R-875 (P4, kept-report reason text), R-876 (P2, the repair misses dpkg's update journal). STATUS updated.
unproven.py --summary: unchanged — 55 claims, 35 not walked (no number moved).
5. Teardown, three layers
- Machines: 9202: BookStack removed through the product (no container, stack dir, volume); controller 0.294.0 (set by
hand, allowed there). demo-hp 9201: every package current, list identical to before; my scripts and logs in
/rootremoved after copying. The bake VM: CT 9100 destroyed, token shredded, reverted tovirgin, qemu gone. - Hosts: the blackhole routes on felhom-pve removed (in the same script). Nothing provisioned.
- Hub: two one-shot clean-up grants (consumed); floors for demo-hp, demo-felhom, tester-1 → 0.294.0; artifacts vouched (agent 0.144.1, golden 0.294.0, min_agent 0.131.0); 12 signed jobs (3× agent 0.144.0, 3× bundle 0.144.0, 3× agent 0.144.1, 3× bundle 0.144.1). Tester 2: read only, nothing sent (offline all session).
- Scratch secrets (hub password, hub key, controller password, deploy secrets) shredded at the end.
6. Notes
- The hub pod log prints a customer's e-mail address in "Customer email sent to …" lines (seen by the Part F helper; not recorded anywhere).
- demo-hp's System page read "no running customer guest" / "unknown" for several minutes after the crash although the guest ran — R-853 (facts late after a boot), not a new defect.