Files
felhom.eu/documentation/audits/evidence-chaos-night-2026-09-17/round_runner.md
T
admin 9fae6dfa98
gates / gates (push) Successful in 23s
CHAOS NIGHT phase 0: golden 0.245.0, a self-installing box, and R-546
The schedule was drawn from seed 20260917 and written into the findings document
BEFORE round 1, with its re-draw log.

Phase 0 measured:
- golden 0.245.0 baked, published (registry 200, not an exit code) and vouched;
  the box installed itself from the published ISO 1.28.0 and landed on it with
  no hand upgrade (controller 0.245.0, agent 0.131.0).
- ZERO operator presses: the waiting self-bind mail worked, and the acknowledged
  -delete path re-issued off-site AND PBS-DR credentials by itself
  (pbsdr_auto_reissue) - the F-14 half nobody had watched happen live.
- R-546 filed (P2): tonight's own guide sends the household to create the
  recovery code ~17 minutes before the box can do it. It self-heals; the bar
  urges them there the whole time. Measured on both sides, not inferred.
- R-543 proven through its whole lifecycle on a fresh box: bar present while
  paused, gone for good once escrowed.
- Known rows met and recorded, not re-filed: R-542, R-536's failure events.

Also recorded honestly: three harness errors of mine (a script that announced
"all twelve deploys ACCEPTED" without checking, a "login ok (csrf 0)" that
turned eleven of my own 401s into what looked like product refusals, and a
head -12 that hid a disk), and a near-miss where I almost filed a defect
against a drive gate that was working and logging at DEBUG.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 22:56:08 +02:00

2.2 KiB

The per-round record — CHAOS NIGHT

Every round records the SAME five things, in this order, and nothing is written from memory:

  1. what the customer saw — screens quoted verbatim (Hungarian in „ ", searched with ASCII fragments, with a positive and a negative control; accented grep dies with a complexity error here, so searching is done in Python)
  2. what the box did by itself — no shell, no help; anything I had to do is an intervention
  3. time to steady — measured from the accident to the moment every app is running, the hub says ONLINE and no alarm is active. — when the box never got there on its own, never a guess
  4. which alarm fired, and was it TRUE
  5. which alarm SHOULD have fired (per 08-alarm-ladder.md) and did not

plus the background loop's failures inside the round's window, counted from its own log.

Two things known BEFORE the night that change how rounds are scored

  • Rounds 7, 8 and 9 all block the box's network. An event generated while the hub is unreachable is retried 3 times over ~6 s and then dropped permanently (PushEvent, no queue). So a missing alarm in those rounds is not evidence the alarm did not fire — it may have been posted into a blocked path. Scored as LOST-IN-BLOCK, never as MISSED.
  • The dedupe is not one window. 5 minutes applies ONLY to the node/host liveness events; the default operator cooldown is 1 hour, keyed customer:type[...]. A second identical alarm inside an hour is suppressed and written to notification_log with status suppressed — so "suppressed" and "never fired" are distinguishable, and must be distinguished.

Steady-state check (the same command set every round)

  • every app: docker ps shows it running AND its front door answers
  • the hub: the host row reads ONLINE, guests 1/1, agent version present
  • no active alarm on the dashboard; notification_log read for the round's window
  • the data drive: the storage page reads „Aktív", not „Leválasztva"

Evidence discipline

Evidence is copied off the box at the end of each round, before the next accident (R-320) — the one that gets forgotten is the middle one, never the last.