Files
felhom.eu/REPORT.md
T

54 lines
3.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# REPORT — Power-outage recovery audit, vacation site (2026-07-22)
**Task:** RUNBOOK-driven, strictly read-only forensic + recovery audit after the 14:4142 CEST
site-wide power cut at the vacation site (felhom-pve N100 + demo-hp t740, ~4.5 h down, manual
power-on ~19:13). No code, no fixes, no restarts — diagnosis only. Deliverable:
`documentation/audits/AUDIT-power-outage-recovery-2026-07-22.md`; raw evidence on DooPlex
`~/outage-20260722/evidence/`.
## Verdict
- **Cause:** simultaneous external power cut on both boxes — both prev-boot journals end abruptly
mid-agent-routine, `last -x` says `crash`, zero non-agent sudo/sshd entries in the window. The
earlier CC session is exonerated.
- **Recovery:** every layer self-healed unaided in **~3 m 15 s** after power-on — fs journal
replay, agent 0.93.0 (+26 s), wg + PBS-over-wg (+32 s), tailscale, guest 9201 (onboot=1,
+2 m 38 s), all stacks healthy on the controllers' own surfaces, cloudflared 4/4 (+2 m 48 s),
hub `*_recovered` (+3 m 17 s). Edge-pinned public probes 302. Felhom-Share autofs woke + CIFS
mounted on first access (first unplanned re-run of the R-67 dead-NAS proof — passed).
- **Dead-man's-switch live fire:** host/node stale + down fired **exactly on the 30 m/60 m design
schedule** (measured from last received report, not from the outage instant), all 9 notification
dispatches `sent` (8 operator + 1 customer). Recovery events are severity `info` → intentionally
no email (flagged as product question F11).
- **F1 regression:** did NOT recur — static .162 held, demo-hp re-acquired .87, local API bound
cleanly on both.
- **Scars:** only OS-self-healed ones (journald rotate, ESP dirty bit, orphan inodes). Metrics DBs
`quick_check` ok, gap 14:40→19:16. Offbox restic: 14 snapshots, latest pre-outage 04:16, zero
stale locks. No backup window fell inside the outage; catch-up is tonight 03:30/04:15.
## Findings (F8+ continuing the vacation arc)
- **F8 (HIGH, mitigated-on-site):** no auto-power-on — power was back in minutes, boxes sat off
4.5 h. BIOS AC-power-on now set (attested); roadmap-candidate: provisioning-checklist item.
- **F9 (MEDIUM, needs-ruling):** H1 OOB belt (`felhom-sshd` + `felhom-oob-nft`) installed on
NEITHER fleet box — only the mgmt-watchdog timer; access is stock sshd :22 + break-glass.
- **F10 (MEDIUM, roadmap):** demo-hp has tier-2 only — no offbox target, 0 PBS DR snapshots.
- **F11 (LOW, needs-ruling):** recovery is silent (`info` severity never emails).
- **F12 (LOW, roadmap):** demo-hp has no customer notification prefs row — node_down customer
email impossible there.
- **F13 (LOW, needs-ruling):** PBS DR tier shows no snapshot cadence (jobs.cfg empty; single
07-18 snapshot on N100, none on demo-hp) — intended?
## Observations
demo-vm tunnel Down is permanent-correct (nested drill VMs 93109312 destroyed pre-outage this
morning; record discard pending). peti-felhom silent since 07-15 08:39 UTC (his site, untouched).
Controllers had been on 0.160.0 for only ~28 min when the cut hit — first dirty shutdown on the
new version came back clean.
**Open question for Viktor:** did the ~14:5715:30 CEST alert emails actually land in the inbox?
Hub-side they are all `sent`; delivery is inbox/Resend-dashboard-side.
**Not validated (needs-mutation or operator-side):** BIOS setting itself, inbox delivery,
Cloudflare dashboard state, UI click-throughs.