54 lines
3.4 KiB
Markdown
54 lines
3.4 KiB
Markdown
# REPORT — Power-outage recovery audit, vacation site (2026-07-22)
|
||
|
||
**Task:** RUNBOOK-driven, strictly read-only forensic + recovery audit after the 14:41–42 CEST
|
||
site-wide power cut at the vacation site (felhom-pve N100 + demo-hp t740, ~4.5 h down, manual
|
||
power-on ~19:13). No code, no fixes, no restarts — diagnosis only. Deliverable:
|
||
`documentation/audits/AUDIT-power-outage-recovery-2026-07-22.md`; raw evidence on DooPlex
|
||
`~/outage-20260722/evidence/`.
|
||
|
||
## Verdict
|
||
|
||
- **Cause:** simultaneous external power cut on both boxes — both prev-boot journals end abruptly
|
||
mid-agent-routine, `last -x` says `crash`, zero non-agent sudo/sshd entries in the window. The
|
||
earlier CC session is exonerated.
|
||
- **Recovery:** every layer self-healed unaided in **~3 m 15 s** after power-on — fs journal
|
||
replay, agent 0.93.0 (+26 s), wg + PBS-over-wg (+32 s), tailscale, guest 9201 (onboot=1,
|
||
+2 m 38 s), all stacks healthy on the controllers' own surfaces, cloudflared 4/4 (+2 m 48 s),
|
||
hub `*_recovered` (+3 m 17 s). Edge-pinned public probes 302. Felhom-Share autofs woke + CIFS
|
||
mounted on first access (first unplanned re-run of the R-67 dead-NAS proof — passed).
|
||
- **Dead-man's-switch live fire:** host/node stale + down fired **exactly on the 30 m/60 m design
|
||
schedule** (measured from last received report, not from the outage instant), all 9 notification
|
||
dispatches `sent` (8 operator + 1 customer). Recovery events are severity `info` → intentionally
|
||
no email (flagged as product question F11).
|
||
- **F1 regression:** did NOT recur — static .162 held, demo-hp re-acquired .87, local API bound
|
||
cleanly on both.
|
||
- **Scars:** only OS-self-healed ones (journald rotate, ESP dirty bit, orphan inodes). Metrics DBs
|
||
`quick_check` ok, gap 14:40→19:16. Offbox restic: 14 snapshots, latest pre-outage 04:16, zero
|
||
stale locks. No backup window fell inside the outage; catch-up is tonight 03:30/04:15.
|
||
|
||
## Findings (F8+ continuing the vacation arc)
|
||
|
||
- **F8 (HIGH, mitigated-on-site):** no auto-power-on — power was back in minutes, boxes sat off
|
||
4.5 h. BIOS AC-power-on now set (attested); roadmap-candidate: provisioning-checklist item.
|
||
- **F9 (MEDIUM, needs-ruling):** H1 OOB belt (`felhom-sshd` + `felhom-oob-nft`) installed on
|
||
NEITHER fleet box — only the mgmt-watchdog timer; access is stock sshd :22 + break-glass.
|
||
- **F10 (MEDIUM, roadmap):** demo-hp has tier-2 only — no offbox target, 0 PBS DR snapshots.
|
||
- **F11 (LOW, needs-ruling):** recovery is silent (`info` severity never emails).
|
||
- **F12 (LOW, roadmap):** demo-hp has no customer notification prefs row — node_down customer
|
||
email impossible there.
|
||
- **F13 (LOW, needs-ruling):** PBS DR tier shows no snapshot cadence (jobs.cfg empty; single
|
||
07-18 snapshot on N100, none on demo-hp) — intended?
|
||
|
||
## Observations
|
||
|
||
demo-vm tunnel Down is permanent-correct (nested drill VMs 9310–9312 destroyed pre-outage this
|
||
morning; record discard pending). peti-felhom silent since 07-15 08:39 UTC (his site, untouched).
|
||
Controllers had been on 0.160.0 for only ~28 min when the cut hit — first dirty shutdown on the
|
||
new version came back clean.
|
||
|
||
**Open question for Viktor:** did the ~14:57–15:30 CEST alert emails actually land in the inbox?
|
||
Hub-side they are all `sent`; delivery is inbox/Resend-dashboard-side.
|
||
|
||
**Not validated (needs-mutation or operator-side):** BIOS setting itself, inbox delivery,
|
||
Cloudflare dashboard state, UI click-throughs.
|