Files
felhom.eu/documentation/runbooks/crash-guard.md
T

20 lines
1.3 KiB
Markdown

# The crash guard — read it, re-arm it
> **Design:** `architecture/11-os-updates.md` §5.9 (decision 88). **When:** the hub mailed `host_crash_guard_tripped`,
> or the System page shows **TRIPPED** for a box.
A tripped guard means: the box stopped uncleanly twice within an hour (a crash, a power cut or a hard reset), so it set
`kernel.panic = 0` — **the next crash leaves it off** until someone switches it on. It re-arms by itself after 24 h of
normal running.
1. **Read why** (as root on the box's host): `felhom-crash-guard status` — `unclean_boots`, `tripped_at`,
`tripped_reason`. Then the end of each crashed boot: `journalctl --list-boots` and `journalctl -b -1 -n 50` (a crashed
boot ends with no shutdown lines). `journalctl -k -b -1 | grep -iE "panic|oops|BUG:"` for a kernel message.
2. **Fix the cause first** if you found one (a bad kernel → boot the previous one; a power problem → the PSU, the cable).
3. **Re-arm** (operator's choice): `felhom-crash-guard rearm` → `kernel.panic = 10` now; the history stays; a fresh
60-minute window starts. The hub announces `host_crash_guard_rearmed` with the next report that carries the facts
(up to ~15 min, R-853).
4. **Check:** `sysctl kernel.panic` = 10; the System page shows "armed".
The numbers live in `/etc/felhom/crash-guard.conf` (LIMIT, WINDOW_MINUTES, PANIC_SECONDS, REARM_HOURS).