Files
felhom.eu/REPORT.md
T

3.4 KiB
Raw Blame History

REPORT — Power-outage recovery audit, vacation site (2026-07-22)

Task: RUNBOOK-driven, strictly read-only forensic + recovery audit after the 14:4142 CEST site-wide power cut at the vacation site (felhom-pve N100 + demo-hp t740, ~4.5 h down, manual power-on ~19:13). No code, no fixes, no restarts — diagnosis only. Deliverable: documentation/audits/AUDIT-power-outage-recovery-2026-07-22.md; raw evidence on DooPlex ~/outage-20260722/evidence/.

Verdict

  • Cause: simultaneous external power cut on both boxes — both prev-boot journals end abruptly mid-agent-routine, last -x says crash, zero non-agent sudo/sshd entries in the window. The earlier CC session is exonerated.
  • Recovery: every layer self-healed unaided in ~3 m 15 s after power-on — fs journal replay, agent 0.93.0 (+26 s), wg + PBS-over-wg (+32 s), tailscale, guest 9201 (onboot=1, +2 m 38 s), all stacks healthy on the controllers' own surfaces, cloudflared 4/4 (+2 m 48 s), hub *_recovered (+3 m 17 s). Edge-pinned public probes 302. Felhom-Share autofs woke + CIFS mounted on first access (first unplanned re-run of the R-67 dead-NAS proof — passed).
  • Dead-man's-switch live fire: host/node stale + down fired exactly on the 30 m/60 m design schedule (measured from last received report, not from the outage instant), all 9 notification dispatches sent (8 operator + 1 customer). Recovery events are severity info → intentionally no email (flagged as product question F11).
  • F1 regression: did NOT recur — static .162 held, demo-hp re-acquired .87, local API bound cleanly on both.
  • Scars: only OS-self-healed ones (journald rotate, ESP dirty bit, orphan inodes). Metrics DBs quick_check ok, gap 14:40→19:16. Offbox restic: 14 snapshots, latest pre-outage 04:16, zero stale locks. No backup window fell inside the outage; catch-up is tonight 03:30/04:15.

Findings (F8+ continuing the vacation arc)

  • F8 (HIGH, mitigated-on-site): no auto-power-on — power was back in minutes, boxes sat off 4.5 h. BIOS AC-power-on now set (attested); roadmap-candidate: provisioning-checklist item.
  • F9 (MEDIUM, needs-ruling): H1 OOB belt (felhom-sshd + felhom-oob-nft) installed on NEITHER fleet box — only the mgmt-watchdog timer; access is stock sshd :22 + break-glass.
  • F10 (MEDIUM, roadmap): demo-hp has tier-2 only — no offbox target, 0 PBS DR snapshots.
  • F11 (LOW, needs-ruling): recovery is silent (info severity never emails).
  • F12 (LOW, roadmap): demo-hp has no customer notification prefs row — node_down customer email impossible there.
  • F13 (LOW, needs-ruling): PBS DR tier shows no snapshot cadence (jobs.cfg empty; single 07-18 snapshot on N100, none on demo-hp) — intended?

Observations

demo-vm tunnel Down is permanent-correct (nested drill VMs 93109312 destroyed pre-outage this morning; record discard pending). peti-felhom silent since 07-15 08:39 UTC (his site, untouched). Controllers had been on 0.160.0 for only ~28 min when the cut hit — first dirty shutdown on the new version came back clean.

Open question for Viktor: did the ~14:5715:30 CEST alert emails actually land in the inbox? Hub-side they are all sent; delivery is inbox/Resend-dashboard-side.

Not validated (needs-mutation or operator-side): BIOS setting itself, inbox delivery, Cloudflare dashboard state, UI click-throughs.