3.4 KiB
REPORT — Power-outage recovery audit, vacation site (2026-07-22)
Task: RUNBOOK-driven, strictly read-only forensic + recovery audit after the 14:41–42 CEST
site-wide power cut at the vacation site (felhom-pve N100 + demo-hp t740, ~4.5 h down, manual
power-on ~19:13). No code, no fixes, no restarts — diagnosis only. Deliverable:
documentation/audits/AUDIT-power-outage-recovery-2026-07-22.md; raw evidence on DooPlex
~/outage-20260722/evidence/.
Verdict
- Cause: simultaneous external power cut on both boxes — both prev-boot journals end abruptly
mid-agent-routine,
last -xsayscrash, zero non-agent sudo/sshd entries in the window. The earlier CC session is exonerated. - Recovery: every layer self-healed unaided in ~3 m 15 s after power-on — fs journal
replay, agent 0.93.0 (+26 s), wg + PBS-over-wg (+32 s), tailscale, guest 9201 (onboot=1,
+2 m 38 s), all stacks healthy on the controllers' own surfaces, cloudflared 4/4 (+2 m 48 s),
hub
*_recovered(+3 m 17 s). Edge-pinned public probes 302. Felhom-Share autofs woke + CIFS mounted on first access (first unplanned re-run of the R-67 dead-NAS proof — passed). - Dead-man's-switch live fire: host/node stale + down fired exactly on the 30 m/60 m design
schedule (measured from last received report, not from the outage instant), all 9 notification
dispatches
sent(8 operator + 1 customer). Recovery events are severityinfo→ intentionally no email (flagged as product question F11). - F1 regression: did NOT recur — static .162 held, demo-hp re-acquired .87, local API bound cleanly on both.
- Scars: only OS-self-healed ones (journald rotate, ESP dirty bit, orphan inodes). Metrics DBs
quick_checkok, gap 14:40→19:16. Offbox restic: 14 snapshots, latest pre-outage 04:16, zero stale locks. No backup window fell inside the outage; catch-up is tonight 03:30/04:15.
Findings (F8+ continuing the vacation arc)
- F8 (HIGH, mitigated-on-site): no auto-power-on — power was back in minutes, boxes sat off 4.5 h. BIOS AC-power-on now set (attested); roadmap-candidate: provisioning-checklist item.
- F9 (MEDIUM, needs-ruling): H1 OOB belt (
felhom-sshd+felhom-oob-nft) installed on NEITHER fleet box — only the mgmt-watchdog timer; access is stock sshd :22 + break-glass. - F10 (MEDIUM, roadmap): demo-hp has tier-2 only — no offbox target, 0 PBS DR snapshots.
- F11 (LOW, needs-ruling): recovery is silent (
infoseverity never emails). - F12 (LOW, roadmap): demo-hp has no customer notification prefs row — node_down customer email impossible there.
- F13 (LOW, needs-ruling): PBS DR tier shows no snapshot cadence (jobs.cfg empty; single 07-18 snapshot on N100, none on demo-hp) — intended?
Observations
demo-vm tunnel Down is permanent-correct (nested drill VMs 9310–9312 destroyed pre-outage this morning; record discard pending). peti-felhom silent since 07-15 08:39 UTC (his site, untouched). Controllers had been on 0.160.0 for only ~28 min when the cut hit — first dirty shutdown on the new version came back clean.
Open question for Viktor: did the ~14:57–15:30 CEST alert emails actually land in the inbox?
Hub-side they are all sent; delivery is inbox/Resend-dashboard-side.
Not validated (needs-mutation or operator-side): BIOS setting itself, inbox delivery, Cloudflare dashboard state, UI click-throughs.