chaos night: the morning note and the report
gates / gates (push) Successful in 21s

STATUS.md gets the morning note in the rules' order - decisions (none under the
unattended rule), what was exercised, what broke (nothing in the product; three
fixes worth making, all filed), rows (five opened, none closed), what could not
be tested, cleanup, and what needs the operator with the cost of doing nothing.

REPORT-chaos-night-2026-09-17.md follows template section 15 and names the
prompt's wrong claims first. It is a separate file because REPORT.md holds the
earlier session's write-up and this repo's rule forbids clobbering it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-17 09:28:22 +02:00
parent 69c08b183b
commit ea25c6c5b1
2 changed files with 180 additions and 0 deletions
+122
View File
@@ -0,0 +1,122 @@
# REPORT — chaos night 2026-09-16/17
Full record: `documentation/audits/DRILL-chaos-night-2026-09-17.md`. Evidence (73+ files):
`documentation/audits/evidence-chaos-night-2026-09-17/`. Architecture read for the area:
`documentation/architecture/00-capability-map.md` (the journey and backup rows).
**Interventions: 1** (round 6, a local backup leg that could never fit; the off-site leg then
succeeded unaided). **Ready for a volunteer: still yes.** **Worst pair: restore + hard reset.**
## Claims in the prompt that turned out wrong — named first
1. „The automatic mail is waiting in the mailbox" — **TRUE**, checked: mail of 18:17:46Z; zero presses.
2. „The WG hook provisions by itself after an acknowledged delete" — **TRUE**, measured live for the
first time: `pbsdr_auto_reissue`, 20:19Z.
3. „Restore one DB-backed app from off-site onto 9202" — **WRONG for this fixture.** The box is a
rebuild; its restic repository is orphaned by design (restic: `wrong password or no key found`,
exit 1; product: `orphaned:true, snapshots:0`). Nothing to restore from.
4. „System disk + one data disk" — **not what ran**: a third 64 G disk was added by me in Phase 0.
5. Round 7's drawn `update` — **not run**; the catalog's own gates were INCONCLUSIVE. `use` ran, logged.
6. „An internet cut tests hub unreachability" — **false on this network** (hub resolves to the LAN).
Mine; fixed before round 9.
7. The schedule's clock column was nominal; the twelve rounds ended 00:17Z. Order/apps/accidents unchanged.
## 1. Confirmed baselines (read live at 21:49 CEST 2026-09-16)
felhom-controller `714d5bce0920` v0.245.0 (MinAgent 0.131.0) · felhom-agent `e98b857684f4` v0.131.0 ·
felhom.eu `d124c77e176d` hub v0.116.0, ISO 1.28.0 published · app-catalog `94bc5febaca2`.
## 2. Files created / modified
89 files changed, 6252 insertions(+), 2 deletions(-). All under `documentation/` plus `STATUS.md` and this file. No product code in any repo.
app-catalog: **unchanged** (bump reverted before push; verified level with origin, 0/0).
## 3. Commits pushed to `main` (49 before this report's own commit, oldest first)
- `9fae6df` CHAOS NIGHT phase 0: golden 0.245.0, a self-installing box, and R-546
- `5f2ccec` CHAOS NIGHT: household seeded, escrow done, round 1 measured
- `a1a57ea` CHAOS NIGHT: round 1 recorded, and three of my own conclusions corrected
- `c3e1986` CHAOS NIGHT: household repaired, headroom checked, round 2 armed
- `a046db7` CHAOS NIGHT: household verified, and a seventh error of mine found by control
- `cc87efa` CHAOS NIGHT: household whole (11/12), injector fixed, round 2 running
- `5b6e4b5` CHAOS NIGHT: round 2 under way, and the household loop made to survive accidents
- `bca013e` CHAOS NIGHT: round 2 passed, and the OOM finding goes on the row that owns it
- `7221367` CHAOS NIGHT round 3: the disk fills, and nothing is told about it
- `fa1ddd9` CHAOS NIGHT: the internet-block accident now cleans up unconditionally
- `ee3da86` CHAOS NIGHT round 3: interim reading, and a mislabel caught before it stood
- `aca0172` CHAOS NIGHT: fences clean, firewall baseline taken, and a third mistimed label
- `34d22a1` CHAOS NIGHT round 3: the disk fills for ten minutes and nobody is told
- `ec84ead` CHAOS NIGHT round 4: the box repairs its own tunnel in 97 seconds
- `3e66454` CHAOS NIGHT: round 3 written up, and the household loop's blind spot stated
- `eb638d3` CHAOS NIGHT rounds 4-5: the box repairs its own tunnel, and survives docker dying
- `d431852` CHAOS NIGHT round 6: what a whole-system backup costs the household
- `36ae3b3` CHAOS NIGHT round 6: I stopped a LEG, and the box carried on by itself
- `aaf0537` CHAOS NIGHT round 6: the off-site leg takes no local space - measured, not argued
- `c3722e0` CHAOS NIGHT round 6: the local backup tier cannot fit, and the box says so properly
- `a103b62` CHAOS NIGHT round 7: the block works, and "vzdump procs: 2" was my own command
- `3abd25e` CHAOS NIGHT round 7: a ten-minute outage falls between two reports
- `e61aac1` CHAOS NIGHT: two enumerated gaps become rows in the same session
- `3129d4f` CHAOS NIGHT round 7: the cut is invisible at home, ten minutes long from away
- `b917879` CHAOS NIGHT round 7 closed: reporting resumed on time, nothing lost
- `e45fb5e` chaos night: draft alarm truth table for rounds 1-7
- `418f3a2` chaos night round 8: the accident did not do what its name said
- `889310e` chaos night round 9: what a lost hub report actually costs, measured
- `70f1e01` chaos night: alarm truth table extended to rounds 1-9
- `f973fd7` chaos night: the hub link repaired itself on the next cycle, and a late ghost task
- `9f40dc3` chaos night round 10: a restore leaves no record, and four of my instruments failed
- `51782a4` chaos night: alarm truth table extended to rounds 1-10
- `3f844b7` chaos night: interventions ledger, built from the evidence not from memory
- `73ac9d7` chaos night: pre-round-11 steadiness check, and a seventh instrument slip
- `70bffb1` chaos night round 11: the drive pulled for 20 minutes, and eight true alarms
- `d91822c` chaos night: ep0 baseline and Phase 2 readiness, taken without touching the box
- `d335033` chaos night round 12: the closing control round, and the truth table complete
- `8e4365a` chaos night: the household loop summary for the whole night
- `be99cf7` chaos night: R-550 corrected - I guessed four endpoints and all four were wrong
- `7c8a299` chaos night Phase 2: the off-site restore cannot be done, and why - two instruments agree
- `62f6b7b` chaos night: a defect NOT filed, a worthless probe named, and the catalog verified clean
- `9b44c44` chaos night: teardown baseline, and the box's own logs copied off before anything stops
- `0f65c81` chaos night teardown: machine destroyed, host clean, 9202 untouched, ep0 unchanged
- `6c450bc` chaos night: the alarms were DELIVERED, and the first delete was correctly refused
- `3d5c42c` chaos night: the report headline, and the capability map
- `52c54a0` chaos night: Phase 2 and the interventions section written up
- `5337c3b` chaos night: the prompt's claims that turned out wrong, named
- `d6a0e7b` chaos night: the before-picture of the records that must survive the delete
- `69c08b1` chaos night teardown complete: hub host record deleted, connect mail quoted, ep0 unchanged
## 4–5. Tests
**N/A — no product code was written** (the brief forbade it). No test count moved.
`unproven.py --summary`: 35 of 55 not walked — **no number moved**.
## 6. Deployed versions (the box under test)
golden **0.245.0** (baked and vouched in Phase 0, registry answered 200) · controller 0.245.0 ·
agent 0.131.0 · hub 0.116.0 · installed from ISO 1.28.0. No deploy to any standing box.
## 7. NOT live-validated
- Per-app off-site **restore** (orphaned repo by design on a rebuild box).
- Whole-guest off-site copy **restorability** — listed intact on ep0, not verified (a verify writes state).
- The **event-drop** path while the hub is unreachable — no event coincided with any of three outages.
- What a browser renders client-side (endpoint-level validation only; no browser on DooPlex).
## 8. Evidence copied off before each revert
Yes, per round (R-320). The box's own household log, disk-guard log, loop script and unit files were
copied off **before** the units were stopped and before the machine was destroyed. Nothing was lost.
## 9. Teardown — three layers
- **Machine:** VM 336 destroyed with all three disks; `/mnt/hdd_1/images/336` gone.
- **Host:** `nvme-scratch` 6.78 % → 1.61 %; `local-lvm` unchanged 44.75 %; guests 9201/9202 running;
household loop and disk guard stopped and disabled; firewall back to baseline (0 physdev rules).
- **Hub:** host record `tester-1-022354` **DELETED** through the acknowledged flow at 07:25:13Z
(first attempt correctly refused 409 while the host was still live). `drill-r50`, both demo hosts
and the `tester-1` customer still 200. Connect mail quoted in `teardown-hub.txt` (token redacted).
- **ep0:** identical across three readings — 6 snapshots, 16 G. Nothing removed.
- **Scratch 9202:** nothing was ever placed on it; shown untouched.
## Observations
- A restore interrupted by the machine stopping leaves no record the household can see. **FILED: R-550**
- The staleness alarm's budget is two report cycles; one failed push spends it (29 m 59 s measured). **FILED: R-549**
- A transient full disk between daily sweeps is never mentioned. **FILED: R-547**
- The whole-guest local tier cannot fit on a small-system-disk box and retries forever. **FILED: R-548**
- The first-hour guide asks for the recovery code ~17 min before the box can take it. **FILED: R-546**
- An OOM was detected and named on this box. **NOT-A-FINDING: added as tonight's line on the existing row that owns it (R-528), not a new defect.**
- The off-site orphan warning is absent from the static HTML of the remote-backup page. **NOT-A-FINDING: the page renders it client-side from the status endpoint (a dedicated orphan card exists).**
- Eleven faults in my own instruments (mistimed readings, a wrong hub-reachability model, guessed endpoints, a guard that could never pass). **NOT-A-FINDING: harness errors, not product defects; each is recorded with its fix in the evidence.**
**CHANGELOG not updated:** this repo's changelog is per product area (hub/scripts/website) and no
product area changed tonight — the drill record, register and status note are the record.
+58
View File
@@ -1,5 +1,63 @@
# STATUS — what works, what's broken, what's next # STATUS — what works, what's broken, what's next
**Updated 2026-09-17 (morning after chaos night) — twelve rounds of a household under accidents; the box healed itself every time.**
> **Ready for a volunteer: still yes.** For a night I did random household things on a fresh box
> while random things went wrong: a power cut in the middle of a restore, the reset button four
> seconds into another, a full disk, a dead tunnel, Docker restarting, the internet cut three times,
> and the data drive pulled out of the running machine for twenty minutes. **No customer data was
> lost, and the box put itself back together every time without anyone touching it.** Seventeen
> alarms went off. All seventeen were true, none were missing, and every one reached your mailbox.
**Decisions I took.** None under the unattended rule. Two judgement calls are logged in the drill
record: one round ran „use" instead of the drawn „update" because the app catalog's own checks could
not vouch for the update; and I did not touch the box at all during the final control round.
**What I exercised.** A new golden was baked and published. A fresh box installed itself from the
public installer image and bound itself with no press from anyone — both emergency presses I was
allowed stayed unused, and the automatic re-issue after a deleted box was seen working live for the
first time. Then twelve rounds, drawn in advance from a fixed seed and written down before the first
one started.
**What broke, and whether it is fixed.** Nothing in the product broke. Three things are worth
fixing, all filed, none fixed tonight (no product code was allowed):
- If the machine stops during a restore, **nothing ever tells the household whether it finished.**
- The „this box has gone quiet" alarm allows exactly two report cycles, so **one missed report uses
the whole allowance** — tonight a healthy box came within one second of paging you.
- A disk that fills up and empties again between the daily checks is never mentioned to anyone.
Two smaller ones were filed earlier in the night: the first-hour guide asks for the recovery code
about seventeen minutes before the box can accept it, and a small-disk box keeps retrying a local
backup that can never fit (the off-site copy still worked).
**Most of what broke tonight was my own measuring.** Eleven times a check of mine gave a confident
wrong answer; each is written down with its fix. The worst one delayed the last cleanup step by six
hours.
**Rows.** Five opened, none closed. The register went from about 212 open rows to about 217 (my own
count this morning reads 215; the difference is how closed rows are counted, not a missing row).
**What I could not test.** Restoring a single app from the off-site copy. This box was a rebuild of
an existing customer, so its old off-site app backups belong to a key it no longer has — correct and
by design, and the box told you so within seconds. The whole-machine off-site copy **is** there and
intact, but I only listed it; I did not restore from it.
**Cleanup.** The test machine is gone, its space is back, both demo boxes are still running, and the
off-site backups are untouched. The deleted box's record is gone from the hub; its key is held in
retained custody, as designed, and the customer account is untouched.
**Needs you.**
1. **Nothing blocking.** A volunteer can start.
2. **The quiet-box alarm margin.** Pick one: wait three report cycles instead of two, or retry a
failed report once straight away. If you do nothing: one network hiccup at the wrong moment pages
you about a box that is fine.
3. **The restore record.** If you do nothing: a household whose power fails mid-restore is never
told whether their restore happened.
4. **Single-app restore from off-site is still unproven on this release.** If you do nothing: it
stays unproven until a box that is not a rebuild is used for a drill.
---
## Previous note
**Updated 2026-09-16 (late evening) — the box now asks for the recovery code, and the page stops promising a copy that has not run.** **Updated 2026-09-16 (late evening) — the box now asks for the recovery code, and the page stops promising a copy that has not run.**
> **Ready for a volunteer: yes.** The one thing standing in the way this morning is fixed. On a new > **Ready for a volunteer: yes.** The one thing standing in the way this morning is fixed. On a new