Files
felhom.eu/documentation/audits/evidence-bignight-2026-09-14/alarm-truth-table.md
T

5.0 KiB
Raw Blame History

Alarm truth table — BIGNIGHT 2026-09-14

Every alarm the hub recorded for tester-1 tonight, and every fault with the alarm it should have raised. Source: hub log (alarms.sh, CEST times), the operator mailbox and the customer mailbox (Gmail connector). „Mail" = the hub logged Operator email sent. The customer mailbox (tester1@felhom.eu) received no alarm mail all night (checked after F4).

Delivery checked, not assumed (20:08Z): the operator mailbox holds each mailed alarm above as its own message from monitoring@felhom.eu to admin@felhom.eu — node_recovered 19:59:44, whole_guest_backup_failed 21:11:14, backup_failed 21:42:54, storage_disconnected 21:58:03, and four app_start_failed at 21:58:16 CEST.

A. Alarms that fired

CEST event (severity) trigger mail TRUE?
19:59:43 node_recovered (staleness down → ok) new box reporting operator true
20:14:35 … 20:45:42 app_deployed (info) × 12 Phase 3 deploys — true
21:02:50 … 21:03:06 crossdrive_completed (info) × 13 Tier 2 run — true
21:11:12 whole_guest_backup_failed (error) PBS tier of „Mentés most" (storage absent, R-511) operator true (the tier cannot succeed); says nothing of the local success
21:31:25 controller_started (info) F1 power-on — true
21:42:53 backup_failed (error) „volume dump interrupted by a controller restart — 1 app(s) … restarted" F2 operator true
21:42:58 controller_started (info) F2 power-on — true
21:54:50 controller_started (info) F3 power-on — true
21:58:02 storage_disconnected (error) „Meghajtó váratlanul leválasztva: Adatlemez" F4 operator true
21:58:15 app_start_failed (warning) × 4 (Paperless-ngx, Jellyfin, Immich, Nextcloud) F4 consequence operator × 4 true but redundant (R-521)
22:28:25 storage_reconnected (info) „Meghajtó újra csatlakoztatva: Adatlemez" F5 — true
22:38:33 storage_disconnected (error) F6 (second unplug) suppressed — cooldown true, not delivered (R-521)
22:38:45 app_start_failed (warning) × 4 F6 consequence suppressed — cooldown × 4 true, not delivered
22:39:45 health_degraded (warning) „Rendszer állapot romlott (volt: ok)" F6 operator true
22:54:46 health_degraded (warning) F7 system disk 95 % suppressed — cooldown true, not delivered; no disk-specific event exists (R-521)
23:25:43 node_stale (staleness ok → stale) F8 internet cut (31 min after the last report) operator true
23:26:43 node_recovered F8 internet back operator true
23:34:35 app_deployed (info) Homebox F9 deploy (sent before the kill) — true
00:04:43 (09-15) node_stale F9 controller dead 30 min suppressed — cooldown (F8's key) true, not delivered (R-523, R-521)

Hub WARN lines without an event or mail: host tester-1-a61396 backup FAILED: target=felhom-pbs … does not exist at 21:12:43 and 21:27:43 (the agent's own PBS attempts).

B. Faults and the alarm each should raise

fault should raise fired? verdict
Paperless OOM on 20 uploads (Phase 3) a dead/unhealthy-worker alarm no MISSED (R-514)
Tunnel 502 for every public name (Phase 2) tunnel has no working route no MISSED (R-510; box logs Request failed every request)
F1 power cut 61 s nothing (under 30 min staleness) controller_started only correct
F2 power cut during backup interrupted backup backup_failed (error) correct; customer page silent (R-519)
F3 power cut during update interrupted update controller_started only nothing recorded; untestable same-version (R-520)
F4 drive unplugged drive lost storage_disconnected (error) (+4 redundant) correct
F5 drive back after 30 min drive back storage_reconnected (info) correct; customer banner stale 2–10 min
F6 drive lost during backup drive lost; backup incomplete storage_disconnected logged, mail suppressed by cooldown; the run reported success:true MISSED on both counts (R-521, R-519)
F7 system disk 95 % disk nearly full health_degraded only, mail suppressed; customer banner in English MISSED for the operator (R-521); customer told, in English (R-516)
F8 internet gone 17½ min box offline node_stale + mail, then node_recovered + mail correct for the operator; the customer's dashboard says „Fut" (R-522)
F9 controller killed mid-deploy the dashboard is down / the controller is not running nothing for 30 min, then node_stale with its mail suppressed MISSED — the stop-rule finding (R-523)
F10 child deletes the photo folder — not run: stop rule met at F9 —
F11 forgotten credentials — not run: stop rule met at F9 —
F12 two reboots in two minutes — not run: stop rule met at F9 —