diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/interventions.txt b/documentation/audits/evidence-chaos-night-2026-09-17/interventions.txt new file mode 100644 index 00000000..fdb99f8d --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/interventions.txt @@ -0,0 +1,55 @@ +# INTERVENTIONS LEDGER — CHAOS NIGHT 2026-09-17 +# The brief's stop rule ends the night at FOUR interventions. This is the count, built from the +# evidence files rather than from memory, with the boundary stated so the number cannot be gamed. + +## THE COUNT DURING THE MEASURED ROUNDS: 1 + + #1 Round 6, 21:59:45Z - I killed the LOCAL leg of the whole-guest backup. + Why: the local tier was writing a ~29 GB source into a 14 GB root filesystem at ~16 MB/s. + Left alone it would have filled `/` and wedged the nested PVE mid-round. + What it cost: it ended a leg that could never have succeeded. The OFF-SITE leg then started + BY ITSELF from the same snapshot and succeeded in ~8.5 minutes, so the data still left the + house. Recorded in round-6.txt, including my first wrong sentence ("I stopped the backup") + and its correction ("I stopped a LEG of the backup"). + Filed as a product finding: R-548. + +## THE PRE-DECLARED PRESSES: BOTH UNUSED, 0 counted + O1 "Send self-bind link" - NOT needed. The automatic mail was already waiting (18:17:46Z), + so the box bound with ZERO operator presses. + O2 "Re-issue PBS credentials" - NOT needed. The acknowledged-delete path re-issued by itself + (`pbsdr_auto_reissue`, 20:19Z) - the F-14 path, measured live + for the first time. + Both prompt claims they guarded against turned out to be TRUE. Recorded in phase0-notes.txt. + +## COUNTED SEPARATELY, BECAUSE IT IS NOT A ROUND RESULT: Phase 0 seeding repairs + During seeding (before round 1) I fired twelve deploys at once onto a thin pool that could not + hold them, filled it to 100%, and then repaired the damage: a guest restart to clear the ext4 + emergency_ro remount, dropping corrupt image layers so compose re-pulled them, a remove-with-data + and one consistent re-deploy after a re-seed minted fresh DB passwords over initialised volumes, + and a rewritten bookstack APP_KEY. + THE DAMAGE WAS MINE, NOT THE PRODUCT'S - the evidence says so in those words ("the corruption is + mine"), and the product's behaviour throughout was correct: it refused to route to unhealthy + containers, raised `app_start_failed` for the app that was down, and named that same app as the + one that failed to back up. Every repair went through the product's own endpoints, never by + hand-running compose, so its record of each app stayed consistent. + These are setup repairs, not rescues of a failing product during a measured round. They are + listed here in full so the distinction is visible rather than convenient. + +## NOT COUNTED, AND WHY + Acts on MY OWN INSTRUMENTS are not interventions in the product. Counting them would flatter the + night in one direction and pad the stop-rule count in the other. They are all recorded anyway: + * moving the dashboard password to /root/.pw after round 2's power cut cleared /tmp + * rewriting inject.sh (hub address) and run_round.sh (door timing, household counts) + * re-creating the disk guard as a real file-backed unit after finding it was transient + * the six mistimed clock readings and their fixes + Round 4 is the clearest case of me DECLINING to intervene: the tunnel was left dead on purpose - + "NOT restarting it by hand - whether it returns by itself IS the measurement". It returned by + itself in 97 s. + +## SWEEP METHOD, so the count is falsifiable + Every phase0-*.txt and round-*.txt file was searched for records of me acting on the box, then + each hit was read in context. Rounds 1, 2, 3, 5, 7, 8, 9 and 10 contain NO record of me touching + the product; their hits are all me correcting my own instruments or my own reasoning. + +## STANDING AGAINST THE STOP RULE + Interventions: 1 of 4. The night continues.