# Alarm truth table — CHAOS NIGHT 2026-09-17 (draft, rounds 1-7; 8-12 appended as they close)
# Built from the verdicts recorded in each round's write-up, not from a fresh grep.
# Columns: round | accident | alarms that FIRED | all true? | any that SHOULD have fired and did not

R1 | none (control)        | app_start_failed; backup_run_failures; offbox_repo_orphaned | YES (3/3) | none
R2 | power cut mid-restore | controller_started (info)                                   | YES (1/1) | none - 60s outage is under both thresholds (node_stale 30m, boot grace 90s)
R3 | system disk 96% / 10m | none                                                        | n/a       | none MISSING, but see note: disk_critical is defined at >=95% and never fired, because
   |                       |                                                             |           | the fill-watch is a DAILY sweep (03:30) + one check ~90s after a controller start.
   |                       |                                                             |           | Predicted BEFORE the round, confirmed after. Filed as R-547 (P3), not a false alarm.
R4 | tunnel killed 10m     | health_critical (error) -> health_recovered (info)          | YES (2/2) | none
R5 | docker restarted      | controller_started (info)                                   | YES (1/1) | none - 90s boot grace covers the app restarts
R6 | none (control)        | whole_guest_backup_failed (error)                           | YES (1/1) | none - and it named the TIER ("local tier"), not "the backup".
   |                       |                                                             |           | No alarm for the 22 apps the backup stopped: correct, those stops are suppressed.
R7 | internet cut 10m      | none                                                        | n/a       | none - node_stale threshold is 30m, the cut was 10m. Nothing false raised either.
R8 | internet cut 10m      | none                                                        | n/a       | none - node_stale threshold is 30m, the cut was 10m. BUT the accident did not do what its
   | (public path only)    |                                                             |           | name said: the hub sits on the LAN here, so the box never lost it. Instrument fault, mine.
R9 | internet cut 10m      | none                                                        | n/a       | none - and this time the hub REALLY was cut. The report was built, pushed 3x over 1m40.8s,
   | (hub cut too)         |                                                             |           | then given up. Nothing queued - correct, a report is a snapshot. Gap to next report 29m59s
   |                       |                                                             |           | against a 30m threshold: one second inside the alarm. Filed as R-549 (P2).
R10| hard reset, 4 s into  | controller_started (info)                                   | YES (1/1) | none from the ladder. But the RESTORE left no record anywhere: 4 status endpoints 404,
   | an app restore        |                                                             |           | no restore field in the status JSON, only a button label on the pages, and no file at
   |                       |                                                             |           | all modified in the reset window. Interrupted and never-happened look identical to the
   |                       |                                                             |           | customer. Filed as R-550 (P2). The box itself recovered 26/26 containers in 150 s.
R11| data drive pulled     | storage_disconnected (error) -> app_start_failed x4 (warning)| YES (8/8) | none. The four apps named are EXACTLY the four whose data lives on the pulled drive;
   | out for 20 minutes    | -> health_degraded (warning) -> storage_reconnected (info)   |           | the other eleven kept serving. The drive is named by the household's own label
   |                       | -> health_recovered (info)                                  |           | ("Adatlemez"). Recovery unaided in 67 s after the drive returned.
   |                       |                                                             |           | NOTE: the round's own snapshot missed health_recovered by SECONDS. Re-read after
   |                       |                                                             |           | the precondition found it. Not a missing alarm - a premature reading, caught.
R12| none (control)        | none                                                        | n/a       | none - and none should have. The newest feed entry is still R11's recovery at 00:13.
   |                       |                                                             |           | Verified the control round runs its OWN path: inject.sh's default branch exits 2, so a
   |                       |                                                             |           | broken injector could not masquerade as a quiet round.

## Running totals, ALL TWELVE ROUNDS
  alarms fired:           17
  alarms TRUE:            17   (17/17 - NO FALSE ALARM IN TWELVE ROUNDS)
  alarms MISSING:          0
  design gaps found:       3   (R-547: a transient full disk is never mentioned to anyone)
                               (R-549: one failed report push spends the whole staleness budget)
                               (R-550: a restore leaves no record, so an interrupted one is invisible)

## STILL UNMEASURED after three internet cuts, and it must be said plainly
  Events pushed while the hub is unreachable are retried 3x and then DROPPED PERMANENTLY, with no
  queue. That path has never been exercised, because no event happened to be raised during any cut.
  Rounds 10-12 draw hard reset, drive pulled and nothing - none of them cuts the hub. So unless an
  event coincides with an outage by accident, this night will not answer it. Recorded, not hidden.

## Two traps avoided, recorded so the numbers can be trusted
  R3: "none fired" was checked TWICE, independently, after the fill was released.
  R5: the first alarm snapshot was taken 6 s after the controller start. An alarm cannot be
      called missing by a measurement taken before it could possibly have fired. Re-measured.
