Files
felhom.eu/documentation/runbooks/crash-guard.md
T

1.3 KiB

The crash guard — read it, re-arm it

Design: architecture/11-os-updates.md §5.9 (decision 88). When: the hub mailed host_crash_guard_tripped, or the System page shows TRIPPED for a box.

A tripped guard means: the box stopped uncleanly twice within an hour (a crash, a power cut or a hard reset), so it set kernel.panic = 0 — the next crash leaves it off until someone switches it on. It re-arms by itself after 24 h of normal running.

  1. Read why (as root on the box's host): felhom-crash-guard status — unclean_boots, tripped_at, tripped_reason. Then the end of each crashed boot: journalctl --list-boots and journalctl -b -1 -n 50 (a crashed boot ends with no shutdown lines). journalctl -k -b -1 | grep -iE "panic|oops|BUG:" for a kernel message.
  2. Fix the cause first if you found one (a bad kernel → boot the previous one; a power problem → the PSU, the cable).
  3. Re-arm (operator's choice): felhom-crash-guard rearm → kernel.panic = 10 now; the history stays; a fresh 60-minute window starts. The hub announces host_crash_guard_rearmed with the next report that carries the facts (up to ~15 min, R-853).
  4. Check: sysctl kernel.panic = 10; the System page shows "armed".

The numbers live in /etc/felhom/crash-guard.conf (LIMIT, WINDOW_MINUTES, PANIC_SECONDS, REARM_HOURS).