chaos night: the report headline, and the capability map
gates / gates (push) Successful in 20s

Headline, three lines. Interventions: 1 - round 6's local backup leg, whose
off-site leg then succeeded unaided; both pre-declared presses went unused.
Ready for a volunteer: still yes - nothing cost a byte of customer data, the
box healed itself every time with no human, and all 17 alarms were true, none
missing, every one delivered. The pair that hurt most: restore + hard reset,
not because the box suffered (26/26 containers back in 150 s) but because it is
the only pair where the household is left not knowing what happened.

Capability map: a new PROVEN-LIVE row for a random night of household actions
under accidents, carrying what it does NOT claim - per-app off-site restore
untested (orphaned repo by design), the dropped-event path still unmeasured
because no event coincided with any hub outage, twelve rounds is a sample not
coverage, and the household loop's 2-minute sampling means ten rounds left no
mark in it.

The unaided-recovery-journey row gets a second scope note rather than a change:
tonight did not walk it and could not have, so its PROVEN-LIVE still stands on
0.206.0 only - neither re-proven nor contradicted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-17 02:43:02 +02:00
parent 6c450bca50
commit 3d5c42c846
2 changed files with 28 additions and 4 deletions
File diff suppressed because one or more lines are too long
@@ -1,8 +1,31 @@
# DRILL — CHAOS NIGHT: random actions on random apps while random things go wrong (2026-09-16/17)
**Interventions: PENDING — the run is in progress.**
**Ready for a volunteer: PENDING.**
**The accident-plus-action pair that hurt most: PENDING.**
**Interventions: 1.** One, at 21:59:45Z in round 6: I killed the **local leg** of a whole-guest
backup that could never have fit (a ~29 GB source into a 14 GB root filesystem, falling at ~16 MB/s).
The **off-site leg then ran by itself from the same snapshot and succeeded**, so the data still left
the house. Both pre-declared presses went **unused**: the automatic self-bind mail was already
waiting, and the acknowledged-delete path re-issued the PBS credentials on its own. Phase 0's
seeding repairs are listed separately in `evidence-chaos-night-2026-09-17/interventions.txt` — that
damage was mine, not the product's, and every repair went through the product's own endpoints.
**Ready for a volunteer: still yes.** Across twelve rounds — a power cut mid-restore, a hard reset
four seconds into another, a full disk, a killed tunnel, a restarted Docker, three severed networks
and the data drive pulled out of a running machine for twenty minutes — **nothing cost a byte of
customer data, and the box healed itself every single time with no human involved.** Seventeen
alarms fired, **all seventeen were true, none were missing**, and the mailbox proves every one was
**delivered** rather than merely stored. The honest qualifications: two P2 legibility gaps are filed
(an interrupted restore leaves no record; one failed report spends the whole 30-minute staleness
budget, measured at 29 m 59 s), and one thing this night could **not** test — per-app off-site
restore, because this box is a rebuild whose restic repository is orphaned **by design**, which the
product surfaced honestly within seconds.
**The accident-plus-action pair that hurt most: `restore` + hard reset (round 10).** Not because the
box suffered — it was back with 26 of 26 containers in **150 s**, boot reconciliation naming the app
it recovered, every front door serving. It hurt most because it is the **only** pair of the night
where the household is left not knowing what happened: they pressed restore, were told it had
started, the machine went dark four seconds later, and afterwards **nothing anywhere tells them
whether it finished.** The status surface exists and answers with the zero value; the record is
in-memory only and does not survive the machine stopping.
> **Baselines, verified live against Gitea at 21:49 CEST 2026-09-16 (not copied from the brief):**
> felhom-controller `714d5bce0920` v0.245.0 (MinAgent 0.131.0) · felhom-agent `e98b857684f4` v0.131.0 ·