5.0 KiB
Alarm truth table — BIGNIGHT 2026-09-14
Every alarm the hub recorded for tester-1 tonight, and every fault with the alarm it should have raised.
Source: hub log (alarms.sh, CEST times), the operator mailbox and the customer mailbox (Gmail connector).
„Mail" = the hub logged Operator email sent. The customer mailbox (tester1@felhom.eu) received no alarm mail all night
(checked after F4).
Delivery checked, not assumed (20:08Z): the operator mailbox holds each mailed alarm above as its own message from
monitoring@felhom.eu to admin@felhom.eu — node_recovered 19:59:44, whole_guest_backup_failed 21:11:14,
backup_failed 21:42:54, storage_disconnected 21:58:03, and four app_start_failed at 21:58:16 CEST.
A. Alarms that fired
| CEST | event (severity) | trigger | TRUE? | |
|---|---|---|---|---|
| 19:59:43 | node_recovered (staleness down → ok) |
new box reporting | operator | true |
| 20:14:35 … 20:45:42 | app_deployed (info) × 12 |
Phase 3 deploys | — | true |
| 21:02:50 … 21:03:06 | crossdrive_completed (info) × 13 |
Tier 2 run | — | true |
| 21:11:12 | whole_guest_backup_failed (error) |
PBS tier of „Mentés most" (storage absent, R-511) | operator | true (the tier cannot succeed); says nothing of the local success |
| 21:31:25 | controller_started (info) |
F1 power-on | — | true |
| 21:42:53 | backup_failed (error) „volume dump interrupted by a controller restart — 1 app(s) … restarted" |
F2 | operator | true |
| 21:42:58 | controller_started (info) |
F2 power-on | — | true |
| 21:54:50 | controller_started (info) |
F3 power-on | — | true |
| 21:58:02 | storage_disconnected (error) „Meghajtó váratlanul leválasztva: Adatlemez" |
F4 | operator | true |
| 21:58:15 | app_start_failed (warning) × 4 (Paperless-ngx, Jellyfin, Immich, Nextcloud) |
F4 consequence | operator × 4 | true but redundant (R-521) |
| 22:28:25 | storage_reconnected (info) „Meghajtó újra csatlakoztatva: Adatlemez" |
F5 | — | true |
| 22:38:33 | storage_disconnected (error) |
F6 (second unplug) | suppressed — cooldown | true, not delivered (R-521) |
| 22:38:45 | app_start_failed (warning) × 4 |
F6 consequence | suppressed — cooldown × 4 | true, not delivered |
| 22:39:45 | health_degraded (warning) „Rendszer állapot romlott (volt: ok)" |
F6 | operator | true |
| 22:54:46 | health_degraded (warning) |
F7 system disk 95 % | suppressed — cooldown | true, not delivered; no disk-specific event exists (R-521) |
| 23:25:43 | node_stale (staleness ok → stale) |
F8 internet cut (31 min after the last report) | operator | true |
| 23:26:43 | node_recovered |
F8 internet back | operator | true |
| 23:34:35 | app_deployed (info) Homebox |
F9 deploy (sent before the kill) | — | true |
| 00:04:43 (09-15) | node_stale |
F9 controller dead 30 min | suppressed — cooldown (F8's key) | true, not delivered (R-523, R-521) |
Hub WARN lines without an event or mail: host tester-1-a61396 backup FAILED: target=felhom-pbs … does not exist at
21:12:43 and 21:27:43 (the agent's own PBS attempts).
B. Faults and the alarm each should raise
| fault | should raise | fired? | verdict |
|---|---|---|---|
| Paperless OOM on 20 uploads (Phase 3) | a dead/unhealthy-worker alarm | no | MISSED (R-514) |
| Tunnel 502 for every public name (Phase 2) | tunnel has no working route | no | MISSED (R-510; box logs Request failed every request) |
| F1 power cut 61 s | nothing (under 30 min staleness) | controller_started only |
correct |
| F2 power cut during backup | interrupted backup | backup_failed (error) |
correct; customer page silent (R-519) |
| F3 power cut during update | interrupted update | controller_started only |
nothing recorded; untestable same-version (R-520) |
| F4 drive unplugged | drive lost | storage_disconnected (error) (+4 redundant) |
correct |
| F5 drive back after 30 min | drive back | storage_reconnected (info) |
correct; customer banner stale 2–10 min |
| F6 drive lost during backup | drive lost; backup incomplete | storage_disconnected logged, mail suppressed by cooldown; the run reported success:true |
MISSED on both counts (R-521, R-519) |
| F7 system disk 95 % | disk nearly full | health_degraded only, mail suppressed; customer banner in English |
MISSED for the operator (R-521); customer told, in English (R-516) |
| F8 internet gone 17½ min | box offline | node_stale + mail, then node_recovered + mail |
correct for the operator; the customer's dashboard says „Fut" (R-522) |
| F9 controller killed mid-deploy | the dashboard is down / the controller is not running | nothing for 30 min, then node_stale with its mail suppressed |
MISSED — the stop-rule finding (R-523) |
| F10 child deletes the photo folder | — | not run: stop rule met at F9 | — |
| F11 forgotten credentials | — | not run: stop rule met at F9 | — |
| F12 two reboots in two minutes | — | not run: stop rule met at F9 | — |