## Alarm truth table — operator mails actually DELIVERED during the drill (read from the mailbox, not from the hub log)
## (to admin@felhom.eu, via the Gmail connector, 2026-09-16)
  1. 12:04:01 CEST  [Felhom] ✅ tester-1: node_recovered (severity INFO) — "Reports resumed (was down for 0m)".
     TRUE but odd: the box had never been down; it was newly enrolled. Also: severity INFO was MAILED, while the
     project's rule says severityNotifies DROPS info before both legs — to be checked against the dispatcher (below).
  2. 12:08:26 CEST  [Felhom] ⚠️ tester-1: backup_tier_skipped (severity warning) — "Whole-guest backup tier felhom-pbs
     skipped: its storage does not exist on the host (never provisioned or removed)". TRUE: R-534 blocked the tier, so
     the tier's storage is genuinely absent. This is R-518's cheap half firing on a real box, with its operator mail.
  No other operator mail arrived in the window (the third thread in the mailbox is a DMARC report, unrelated).
## Faults injected so far: none (Phase 2 starts with F9'). Alarms that SHOULD have fired and did not: none so far.
## Why the INFO mail is correct (checked in code, not assumed):
##   `node_recovered` is emitted with severity "info" (monitor/staleness.go emitTransition) and handed to
##   dispatcher.ProcessEvent, which returns before both legs for "info" — BUT the recovery branch
##   (`recoveredPairedDownTypes`, dispatcher.go:124) runs BEFORE that gate, deliberately, so the operator hears the
##   all-clear for a down it was told about. The customer leg stays pairing-gated.
##   In context the mail is also TRUE: customer `tester-1` had been reporting nothing since the BIGNIGHT teardown, so the
##   new box's first report IS a recovery, and "(was down for 0m)" is the age of the gap the checker could see.
##   No row filed.

## ALARM TRUTH TABLE - second pass, 2026-09-16 (hub 0.115.0), built from the hub's own log lines
## (pod hub-85478f77d7-mpxkr, times CEST) paired with the customer timeline on /customers/tester-1.
##
## event                         when      severity  operator mail?              true in context?
##  selfbind_link_sent           11:59:55  info      yes (self-bind e-mail)       yes - the operator pressed it
##  claim_reissued_reenroll      12:01:56  info      yes (claim code to customer) yes - 3rd generation after reinstall
##  controller_started           12:03:30  info      no                           correct - info, not paired
##  node_recovered               12:04:00  info      YES - mail sent              BY DESIGN: the recovery branch runs
##                                                                                before the severity gate; true here
##  backup_tier_skipped          12:08:25  warning   yes                          yes - the off-site tier has no storage
##  backup_tier_skipped          12:37:17  warning   SUPPRESSED - cooldown        correct - same key within the hour
##                                                    (key=tester-1:backup_tier_skipped:felhom-pbs)
##  app_deployed x5              12:21-12:31 info    no                           NOT always true -> R-536 (sent at accept)
##  app_start_failed (paperless) 12:48:47  warning   yes                          TRUE - and it was MY damage (the
##                                                                                hand-run compose, recorded in phase2-m1)
##  controller_restarted_by_agent 12:48    info      no (info)                    yes - F9' restart #1, timestamps match
##  backup_tier_skipped          12:58:13  warning   SUPPRESSED - cooldown        correct
##  claim_lockout                13:11:42  warning   YES, 1 s later               TRUE - F11 part 2, third wrong code
##  claim reset code (gen 4)     13:12:05  info      yes (to the customer)        yes - "Uj beallito kod kerese"
##
## What this pass adds to the first one:
##  1) The 5-minute node/host dedupe (R-529) was NOT exercised today: no node_* or host_* alarm fired
##     after 12:04, because the box stayed up and reporting. The two reboots of F12 were far too short
##     to reach the 30-minute staleness threshold. So the widened bypass is SHIPPED but UNPROVEN-LIVE
##     on this box; the proof from 2026-09-15 on the other box still stands.
##  2) The one-hour operator cooldown is PROVEN here twice, by its own log line, on backup_tier_skipped.
##  3) Every warning that fired today was true. No warning that should have fired stayed silent, EXCEPT
##     the class R-538 names: there is no alarm at all for "a restore finished and the app is now
##     inconsistent", so nothing could have fired.
##
## MAILBOX SIDE, read directly from the inbox (not from hub logs), 2026-09-16T11:2xZ:
##   09:59:56Z  to tester1@   "[Felhom] Kosd ossze a Felhom dobozodat"          (self-bind link)
##   10:01:57Z  to tester1@   "[Felhom] Uj beallito kod - ujratelepult a szervered"
##   10:04:01Z  to admin@     "[Felhom] tester-1: node_recovered" (info)        <- the by-design info mail
##   10:08:26Z  to admin@     "[Felhom] tester-1: backup_tier_skipped" (warning)
##   10:48:47Z  to admin@     "[Felhom] tester-1: app_start_failed" (warning)   <- true; my own damage
##   11:11:43Z  to admin@     "[Felhom] tester-1: claim_lockout" (warning)      <- F11, true
##   11:12:06Z  to tester1@   "[Felhom] Beallito kod a jelszavad visszaallitasahoz" (the reset code)
##   11:18:40Z  to admin@     "[Felhom] tester-1: backup_tier_skipped" (warning)
##
## The last line is the COOLDOWN EXPIRY measured end to end: the same event was suppressed at 12:37 and
## 12:58 CEST ("Operator email suppressed ... cooldown") and mailed again at 13:18 CEST - one hour after
## the 12:08 mail. So the one-hour operator cooldown both SUPPRESSES and RELEASES correctly, proven from
## the hub log and from the inbox independently.
##
## THE STALENESS ALARM, fired by the teardown itself (and it is TRUE - the box really is gone):
##   the box's last report: 13:56:14 CEST (11:56:14Z); the machine destroyed ~13:57:30 CEST.
##   14:22:00 CEST  host_stale  "tester-1-652049 ok -> stale"  -> OPERATOR EMAIL SENT, same second.
##   That is 25m46s after the last report, i.e. the 30-minute staleness threshold measured from the
##   LAST REPORT, not from the moment of death - exactly as the dead-man's-switch is designed.
##   The hub's delete-impact endpoint flipped to {"status":"stale","deletable":true} in the same minute
##   (12:21:27Z deletable=False -> 12:22:28Z deletable=True), which is what released the host delete.
## NOTE on R-529 (the widened host_* cooldown bypass): this fired ONCE, so the 5-minute dedupe was still
##   not exercised on this box - a second host_* within the hour would be needed, and the box was deleted
##   instead. The bypass remains proven only by its red-proofed test and the 2026-09-15 node_* run.
