diff --git a/documentation/audits/DRILL-chaos-night-2026-09-17.md b/documentation/audits/DRILL-chaos-night-2026-09-17.md index 184a40b5..dc2a11bb 100644 --- a/documentation/audits/DRILL-chaos-night-2026-09-17.md +++ b/documentation/audits/DRILL-chaos-night-2026-09-17.md @@ -541,18 +541,90 @@ cuts, resets, full disks, severed networks and a drive pulled out of a running m which nothing was done produced nothing — no alarm, no restart, no drift. The alarm feed is not simply noisy. -### Rounds 8-12 — *(written up individually above)* - -PENDING - - ## Phase 2 — the morning after -PENDING +**Every app answered through its own front door.** All twelve real names, measured at 00:19:16Z — +`cloud`, `inventory`, `media`, `paperless`, `paste`, `photos`, `recipes`, `share`, `status`, +`travel`, `vault`, `wiki` — **200 on the public path, every one**. And healthy is not inferred from +a door: all **26 containers** report `Up … (healthy)`, except `traefik` and `cloudflared`, which +carry no healthcheck and show a bare `Up`. The four showing seven minutes are the drive-backed apps +restarted after round 11. + +**The off-site restore could not be done, and two independent instruments agree why.** The brief asked +for one DB-backed app restored from off-site onto scratch 9202. It has nothing to restore from: + +| instrument | answer | +|---|---| +| `restic`, with the box's own key, password file and `known_hosts` | `Fatal: wrong password or no key found` — exit status **1**, read from restic itself rather than from the end of a pipeline | +| the product's own status surface, which is what a customer sees | `{"orphaned":true, … ,"snapshots":0,"status":"error"}` | + +The repository is orphaned **because this box is a rebuild for an existing customer**: on a rebuild +the restic password is minted fresh, so snapshots written under the old one can never be opened +again. That is a known, documented shape, and **the product surfaced it honestly** — the true +`offbox_repo_orphaned` alarm of round 1, mailed to the operator within two seconds of the run. + +**What did leave the house.** The *whole-guest* off-site copy is a different store and it worked: ep0 +holds two intact snapshots for this box, including tonight's at **21:59:54Z** — round 6's off-site +leg, the one that started by itself after I killed the local leg. Its file index is roughly four +times the afternoon copy's, consistent with a guest by then carrying twelve apps. **Stated as a +limit:** that is a *listing*, not a verification. A PBS verify job would prove restorability and it +**writes** verify state, so it was not run — ep0 is read-only for evidence tonight. + +**The household loop.** 204 probes from 21:06:30Z to 00:18:10Z. Seven lines flagged as failures, of +which **only two are real events** — `wiki` during round 2's power cut and `cloud` during round 10's +hard reset, each a single sample, each healed before the next probe. The other three were **my own +classifier** counting a 301 redirect as a dashboard failure; the log carries that correction in its +own words at 21:11:25Z and the wrong lines were left in place so the correction stays visible. Ten of +the twelve rounds left **no mark at all** in this log, including the twenty minutes with the drive +pulled — the loop samples each name every two minutes, so that silence is the instrument's sampling +rate and **not** evidence the household saw nothing. + +**The catalog bump: verified reverted, not remembered.** Clean tree, local exactly level with +`origin/main` (0 ahead, 0 behind), the nextcloud template still on its original `redis:7-alpine` pin +and `catalog_since: "2026-07-18"`, newest commit 2026-09-15. The bump was prepared, **refused by the +catalog's own gates** (image-resolvable and volume-persistence both INCONCLUSIVE — its own canary +failed, so the verdict was UNDETERMINED, never a pass) and reverted before any push. + +**And one delivery proof the truth table could not give.** The mailbox shows every alarm **arrived** +at the operator, not merely that it was stored — round 11's whole set, round 6's tier-naming backup +failure, round 1's orphan warning, the OOM, the deploy failure and my own thin-pool `storage_fill_critical`. +In a project where 91 events once sat in a database having e-mailed nobody, *raised* and *delivered* +are two different claims, and only one of them had evidence before tonight. ## Interventions — counted, with the reason for each verdict -PENDING +**One.** At 21:59:45Z in round 6 I killed the **local leg** of the whole-guest backup. The arithmetic +that forced it: a ~29 GB source being written into a 14 GB root filesystem at ~16 MB/s, i.e. under +four minutes to a full `/` on the nested host, mid-round. What it cost: a leg that could never have +succeeded. What happened next without me: **the off-site leg started by itself from the same snapshot +and succeeded in about eight and a half minutes.** Filed as **R-548**. + +My first note on it said „I stopped the backup". That was wrong and is corrected in the evidence: I +stopped **a leg** of it, and the box completed the other one unaided. + +**Both pre-declared presses went unused.** O1, the „Send self-bind link" button, was not needed — +the automatic mail was already waiting (18:17:46Z) and the box bound with **zero** operator presses. +O2, „Re-issue PBS credentials", was not needed either — the acknowledged-delete path re-issued them +by itself (`pbsdr_auto_reissue`, 20:19Z). **Both prompt claims they were insurance against turned out +to be true**, and the F-14 path was measured live for the first time. + +**Counted separately, because it is not a round result:** Phase 0's seeding repairs. I filled the LVM +thin pool to 100 % by firing twelve deploys at once, then repaired the damage — a guest restart to +clear an `emergency_ro` remount, dropping corrupt image layers, a remove-with-data and one consistent +re-deploy after a re-seed minted fresh database passwords over initialised volumes, and a rewritten +`APP_KEY`. **The damage was mine, not the product's**, the product's behaviour throughout was correct, +and every repair went through the product's own endpoints rather than by hand-running compose. Listed +in full in `evidence-chaos-night-2026-09-17/interventions.txt` so the distinction is visible rather +than convenient. + +**Not counted, and why:** acts on my **own instruments** — moving the dashboard password after a +power cut cleared `/tmp`, rewriting the injector and the runner, re-creating the disk guard as a real +unit. Counting those would flatter the night in one direction and pad the stop-rule count in the +other. **Round 4 is the clearest case of declining to intervene:** the tunnel was left dead on +purpose — „NOT restarting it by hand — whether it returns by itself IS the measurement". It +returned by itself in 97 s. + +**Standing against the stop rule: 1 of 4.** The night ran its full twelve rounds. ## Teardown — three layers, stated