STATUS.md gets the morning note in the rules' order - decisions (none under the unattended rule), what was exercised, what broke (nothing in the product; three fixes worth making, all filed), rows (five opened, none closed), what could not be tested, cleanup, and what needs the operator with the cost of doing nothing. REPORT-chaos-night-2026-09-17.md follows template section 15 and names the prompt's wrong claims first. It is a separate file because REPORT.md holds the earlier session's write-up and this repo's rule forbids clobbering it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
9.3 KiB
REPORT — chaos night 2026-09-16/17
Full record: documentation/audits/DRILL-chaos-night-2026-09-17.md. Evidence (73+ files):
documentation/audits/evidence-chaos-night-2026-09-17/. Architecture read for the area:
documentation/architecture/00-capability-map.md (the journey and backup rows).
Interventions: 1 (round 6, a local backup leg that could never fit; the off-site leg then succeeded unaided). Ready for a volunteer: still yes. Worst pair: restore + hard reset.
Claims in the prompt that turned out wrong — named first
- „The automatic mail is waiting in the mailbox" — TRUE, checked: mail of 18:17:46Z; zero presses.
- „The WG hook provisions by itself after an acknowledged delete" — TRUE, measured live for the
first time:
pbsdr_auto_reissue, 20:19Z. - „Restore one DB-backed app from off-site onto 9202" — WRONG for this fixture. The box is a
rebuild; its restic repository is orphaned by design (restic:
wrong password or no key found, exit 1; product:orphaned:true, snapshots:0). Nothing to restore from. - „System disk + one data disk" — not what ran: a third 64 G disk was added by me in Phase 0.
- Round 7's drawn
update— not run; the catalog's own gates were INCONCLUSIVE.useran, logged. - „An internet cut tests hub unreachability" — false on this network (hub resolves to the LAN). Mine; fixed before round 9.
- The schedule's clock column was nominal; the twelve rounds ended 00:17Z. Order/apps/accidents unchanged.
1. Confirmed baselines (read live at 21:49 CEST 2026-09-16)
felhom-controller 714d5bce0920 v0.245.0 (MinAgent 0.131.0) · felhom-agent e98b857684f4 v0.131.0 ·
felhom.eu d124c77e176d hub v0.116.0, ISO 1.28.0 published · app-catalog 94bc5febaca2.
2. Files created / modified
89 files changed, 6252 insertions(+), 2 deletions(-). All under documentation/ plus STATUS.md and this file. No product code in any repo.
app-catalog: unchanged (bump reverted before push; verified level with origin, 0/0).
3. Commits pushed to main (49 before this report's own commit, oldest first)
9fae6dfCHAOS NIGHT phase 0: golden 0.245.0, a self-installing box, and R-5465f2ccecCHAOS NIGHT: household seeded, escrow done, round 1 measureda1a57eaCHAOS NIGHT: round 1 recorded, and three of my own conclusions correctedc3e1986CHAOS NIGHT: household repaired, headroom checked, round 2 armeda046db7CHAOS NIGHT: household verified, and a seventh error of mine found by controlcc87efaCHAOS NIGHT: household whole (11/12), injector fixed, round 2 running5b6e4b5CHAOS NIGHT: round 2 under way, and the household loop made to survive accidentsbca013eCHAOS NIGHT: round 2 passed, and the OOM finding goes on the row that owns it7221367CHAOS NIGHT round 3: the disk fills, and nothing is told about itfa1ddd9CHAOS NIGHT: the internet-block accident now cleans up unconditionallyee3da86CHAOS NIGHT round 3: interim reading, and a mislabel caught before it stoodaca0172CHAOS NIGHT: fences clean, firewall baseline taken, and a third mistimed label34d22a1CHAOS NIGHT round 3: the disk fills for ten minutes and nobody is toldec84eadCHAOS NIGHT round 4: the box repairs its own tunnel in 97 seconds3e66454CHAOS NIGHT: round 3 written up, and the household loop's blind spot statedeb638d3CHAOS NIGHT rounds 4-5: the box repairs its own tunnel, and survives docker dyingd431852CHAOS NIGHT round 6: what a whole-system backup costs the household36ae3b3CHAOS NIGHT round 6: I stopped a LEG, and the box carried on by itselfaaf0537CHAOS NIGHT round 6: the off-site leg takes no local space - measured, not arguedc3722e0CHAOS NIGHT round 6: the local backup tier cannot fit, and the box says so properlya103b62CHAOS NIGHT round 7: the block works, and "vzdump procs: 2" was my own command3abd25eCHAOS NIGHT round 7: a ten-minute outage falls between two reportse61aac1CHAOS NIGHT: two enumerated gaps become rows in the same session3129d4fCHAOS NIGHT round 7: the cut is invisible at home, ten minutes long from awayb917879CHAOS NIGHT round 7 closed: reporting resumed on time, nothing loste45fb5echaos night: draft alarm truth table for rounds 1-7418f3a2chaos night round 8: the accident did not do what its name said889310echaos night round 9: what a lost hub report actually costs, measured70f1e01chaos night: alarm truth table extended to rounds 1-9f973fd7chaos night: the hub link repaired itself on the next cycle, and a late ghost task9f40dc3chaos night round 10: a restore leaves no record, and four of my instruments failed51782a4chaos night: alarm truth table extended to rounds 1-103f844b7chaos night: interventions ledger, built from the evidence not from memory73ac9d7chaos night: pre-round-11 steadiness check, and a seventh instrument slip70bffb1chaos night round 11: the drive pulled for 20 minutes, and eight true alarmsd91822cchaos night: ep0 baseline and Phase 2 readiness, taken without touching the boxd335033chaos night round 12: the closing control round, and the truth table complete8e4365achaos night: the household loop summary for the whole nightbe99cf7chaos night: R-550 corrected - I guessed four endpoints and all four were wrong7c8a299chaos night Phase 2: the off-site restore cannot be done, and why - two instruments agree62f6b7bchaos night: a defect NOT filed, a worthless probe named, and the catalog verified clean9b44c44chaos night: teardown baseline, and the box's own logs copied off before anything stops0f65c81chaos night teardown: machine destroyed, host clean, 9202 untouched, ep0 unchanged6c450bcchaos night: the alarms were DELIVERED, and the first delete was correctly refused3d5c42cchaos night: the report headline, and the capability map52c54a0chaos night: Phase 2 and the interventions section written up5337c3bchaos night: the prompt's claims that turned out wrong, namedd6a0e7bchaos night: the before-picture of the records that must survive the delete69c08b1chaos night teardown complete: hub host record deleted, connect mail quoted, ep0 unchanged
4–5. Tests
N/A — no product code was written (the brief forbade it). No test count moved.
unproven.py --summary: 35 of 55 not walked — no number moved.
6. Deployed versions (the box under test)
golden 0.245.0 (baked and vouched in Phase 0, registry answered 200) · controller 0.245.0 · agent 0.131.0 · hub 0.116.0 · installed from ISO 1.28.0. No deploy to any standing box.
7. NOT live-validated
- Per-app off-site restore (orphaned repo by design on a rebuild box).
- Whole-guest off-site copy restorability — listed intact on ep0, not verified (a verify writes state).
- The event-drop path while the hub is unreachable — no event coincided with any of three outages.
- What a browser renders client-side (endpoint-level validation only; no browser on DooPlex).
8. Evidence copied off before each revert
Yes, per round (R-320). The box's own household log, disk-guard log, loop script and unit files were copied off before the units were stopped and before the machine was destroyed. Nothing was lost.
9. Teardown — three layers
- Machine: VM 336 destroyed with all three disks;
/mnt/hdd_1/images/336gone. - Host:
nvme-scratch6.78 % → 1.61 %;local-lvmunchanged 44.75 %; guests 9201/9202 running; household loop and disk guard stopped and disabled; firewall back to baseline (0 physdev rules). - Hub: host record
tester-1-022354DELETED through the acknowledged flow at 07:25:13Z (first attempt correctly refused 409 while the host was still live).drill-r50, both demo hosts and thetester-1customer still 200. Connect mail quoted inteardown-hub.txt(token redacted). - ep0: identical across three readings — 6 snapshots, 16 G. Nothing removed.
- Scratch 9202: nothing was ever placed on it; shown untouched.
Observations
- A restore interrupted by the machine stopping leaves no record the household can see. FILED: R-550
- The staleness alarm's budget is two report cycles; one failed push spends it (29 m 59 s measured). FILED: R-549
- A transient full disk between daily sweeps is never mentioned. FILED: R-547
- The whole-guest local tier cannot fit on a small-system-disk box and retries forever. FILED: R-548
- The first-hour guide asks for the recovery code ~17 min before the box can take it. FILED: R-546
- An OOM was detected and named on this box. NOT-A-FINDING: added as tonight's line on the existing row that owns it (R-528), not a new defect.
- The off-site orphan warning is absent from the static HTML of the remote-backup page. NOT-A-FINDING: the page renders it client-side from the status endpoint (a dedicated orphan card exists).
- Eleven faults in my own instruments (mistimed readings, a wrong hub-reachability model, guessed endpoints, a guard that could never pass). NOT-A-FINDING: harness errors, not product defects; each is recorded with its fix in the evidence.
CHANGELOG not updated: this repo's changelog is per product area (hub/scripts/website) and no product area changed tonight — the drill record, register and status note are the record.