Files
felhom.eu/documentation/audits/AUDIT-power-outage-recovery-2026-07-22.md

15 KiB
Raw Permalink Blame History

AUDIT — Power-outage recovery, vacation site, 2026-07-22

Run: read-only forensic + recovery audit, executed from DooPlex 2026-07-22 19:3820:00 CEST. Scope: felhom-pve (N100, demo-felhom-8363b5, tailnet 100.70.170.35) and demo-hp (HP t740, demo-hp-bb76ea, 100.76.96.79), their guests 9201, the hub plane, and the backup plane. Mode: strictly read-only on all live systems — nothing was restarted, redeployed, unlocked or triggered. The single state-affecting action was a designed one: a read-only ls that woke the idle Felhom-Share autofs trigger (its normal first-access path). Evidence: DooPlex ~/outage-20260722/evidence/ — every claim below cites a file there. Timestamps: normalized to CEST. Hosts log CEST; guests, the hub DB and Cloudflare log UTC (CEST2); the hub pod log prints CEST. Watch for this when re-reading raw evidence. Access paths used (p0-access.txt): felhom-pve = operator key, root@tailnet. demo-hp = key-less by design → hub-vaulted G1 break-glass (host_recovery/demo-hp-bb76ea, file→file, shredded after). Both worked on the first path; the OOB belt units however are absent — see F9.

Related: AUDIT-vacation-remote-ops-2026-07-20.md (F1F7 of this arc; F-numbers continue at F8).


1. Verdict on cause: simultaneous external power cut — CC session exonerated

  • Both previous-boot journals end abruptly mid-routine-agent-activity — felhom-pve at 14:41:31, demo-hp at 14:41:15 — with no Reached target Power-Off cascade (p1-felhom-pve-journal-prevboot-tail.txt, p1-demo-hp-journal-prevboot-tail.txt).
  • last -x marks the previous boot crash on both hosts (p1-*-forensic.txt).
  • The 14:0015:00 sudo/sshd window contains zero non-agent entries on either host — 4 569 resp. 4 533 journal lines, all felhom-agent sudo routine; no SSH login, no poweroff/shutdown/ pct/qm invocation (p1-*-forensic.txt).
  • Both boxes show the classic dirty-shutdown recovery trio on the next boot: journald system.journal corrupted or uncleanly shut down, renaming and replacing; ESP FAT Dirty bit is set … Automatically removing; ext4 orphan-inode cleanup (5 inodes on felhom-pve dm-7, 1 on demo-hp dm-7). NTP synchronized on both after boot.

Verdict: site-wide power cut at ~14:4114:42 CEST hitting both boxes (and the router and the NAS) simultaneously. No software or operator action involved; the earlier CC session is ruled out.

2. Timeline (CEST)

Time Event Evidence
~14:2714:29 Last reports received by hub (host_reports 14:26:54 / 14:27:54; controller reports 14:28:50 / 14:29:13 — N100 / hp) p5-report-gap.txt
14:4014:41 Last metrics rows (guest DBs): 14:40:49 N100, 14:40:12 hp p3-*-metrics-quickcheck.txt
14:4114:42 Power cut. Journals end 14:41:31 / 14:41:15; router reboots (operator-attested, uptime counter) — power itself returned within minutes; both miniPCs stayed OFF p1-*
14:57:09 / 14:58:09 hub host_stale (N100 / hp) + operator email sent p5-hub-log-grep.txt, p5-hub-notification-log.txt
14:59:09 / 15:00:09 hub node_stale + operator email sent
15:27:09 / 15:28:09 hub host_down + operator email sent
15:29:09 / 15:30:09 hub node_down + operator email; customer email (doodoo21@freemail.hu) for demo-felhom
~19:1019:13 Viktor powers both boxes on manually (attested)
19:13:52 / 19:13:53 T1: kernel boot (N100 / hp, uptime -s) p1-*-forensic.txt
19:14:0419:14:21 Host plane up on both: tailscaled, wg-quick, felhom-agent 0.93.0, first host_report at hub (hp 19:14:12, N100 19:14:21) p2-*, p4-*, p5-report-gap.txt
19:14:11 / 19:14:24 PBS-over-wg verify cycle complete (hp / N100) — WireGuard path proven live p6-*-backup.txt
19:15:09 hub host_recovered (both, same tick) p5-hub-events.txt
19:16:23 / 19:16:30 Guest 9201 boots (hp / N100; onboot: 1, PVE-started) p3-*-guests, p3-*-docker-ps
19:16:3319:16:40 cloudflared re-registers 4/4 tunnel connections on both guests p4-*-tunnels.txt
19:16:34 / 19:16:38 First post-outage metrics row (hp / N100) p3-*-metrics-quickcheck.txt
19:16:40 / 19:16:44 controller_started events; first controller reports at hub 19:16:41 / 19:16:44 p5-hub-events.txt, p5-report-gap.txt
19:17:09 hub node_recovered (both) — full platform recovery, Δ ≈ 3 m 15 s from power-on p5-hub-events.txt
19:3819:47 Audit probes: tailscale both active-direct; edge-pinned HTTPS 302 on both tunnels; wg handshakes fresh; NAS share wakes + mounts on first access p0-, p4-external-, p3-demo-hp-nas-wake.txt

Total service outage ≈ 4 h 35 m, of which ~4 h 30 m was boxes sitting powered off (F8) and ~3 m 15 s was actual recovery work.

3. Self-healing scorecard

Every component recovered unaided; Δ measured from T1 (19:13:52/53).

Component N100 Δ demo-hp Δ Unaided Evidence
Kernel boot + fs recovery (journal replay, orphan cleanup, ESP dirty bit) +3 s +5 s yes p1-*-forensic.txt
tailscaled (control + direct conns) ~+21 s ~+11 s yes p4-*, p0-tailscale-status.txt
wg-quick@wg-felhom (iface up) +22 s +12 s yes p4-*-tunnels.txt
felhom-agent 0.93.0 active +26 s +16 s yes p2-*-host.txt
First host_report at hub +29 s +19 s yes p5-report-gap.txt
PBS reachable over wg (verify cycle OK) +32 s +18 s yes p6-*-backup.txt
Agent local API bound (192.168.0.162 / .87:8443) +59 s +19 s yes p2-*-f1-check.txt
Guest 9201 (onboot=1) +2 m 38 s +2 m 30 s yes p3-*
Docker + all stacks (12 resp. 7 containers, 0 exited, 0 unhealthy; controller surface: every deployed stack running) ~+3 m ~+3 m yes p3-*-docker-ps.txt, p3-*-stacks-parsed.txt
cloudflared 4/4 connections +2 m 48 s +2 m 42 s yes p4-*-tunnels.txt
First controller report at hub +2 m 52 s +2 m 48 s yes p5-report-gap.txt
Hub *_recovered events +1 m 17 s (host) / +3 m 17 s (node) same ticks yes p5-hub-events.txt
Felhom-Share NAS (demo-hp): idle autofs trigger restored at boot; first access mounts CIFS //192.168.0.104/Share; FileBrowser bind live; content intact on first access yes p3-demo-hp-nas-mounts.txt, p3-demo-hp-nas-wake.txt
External URLs via pinned Cloudflare edge: 302 in 0.140.16 s ✔ at audit yes p4-external-edge-probes.txt

F1 regression check (load-bearing): F1 did NOT recur. felhom-pve came back on its static 192.168.0.162; demo-hp re-acquired its previous DHCP 192.168.0.87 despite the router reboot; agent local-API bound cleanly on both, zero bind errors in the journals (p2-*-f1-check.txt). The static-IP mitigation from AUDIT-vacation-remote-ops F1 held under its natural trigger.

Dead-man's-switch verdict: fired exactly on design schedule. The runbook's ~15:12/~15:42 expectation measured from the outage instant; the design (30 m stale / 60 m down, hub-config stale_threshold: "30m", staleness.go downAfter = 2×) measures from the last received report (~14:2714:29). Observed firings are correct to the minute (checker ticks at :09). All 9 dispatches logged status=sent, zero errors: 8 operator + 1 customer (demo-felhom/node_down — the only customer with notification prefs, see F12). *_recovered events carry severity info, which the dispatcher intentionally does not email (dispatcher.go: "info is an intentional non-notify (status/recovery events)") — see F11.

Data-integrity scars: confined to what the OS self-healed (§1). PRAGMA quick_check = ok on both metrics DBs (checked on copies); metrics gap cleanly bounded 14:40 → 19:16. Offbox restic repo consistent, 14 snapshots, latest 04:16 CEST (pre-outage), zero locks (read-only restic list locks). No backup job was running at the cut — guest backup windows are nightly (tier2 03:30, offbox 04:15 CEST), no window fell inside 14:4219:15, catch-up is automatic tonight (p6-*).

4. Findings (continuing the vacation arc; F1F7 in AUDIT-vacation-remote-ops-2026-07-20.md)

# Finding Severity Evidence Disposition
F8 Site power loss; no auto-power-on. Power itself returned within minutes (router rebooted at ~14:42 and stayed up); both miniPCs remained off ~4.5 h until manual power-on. For a paying customer this converts a power blip into a half-day outage ending only when someone is physically present. HIGH p1-*; router uptime operator-attested mitigated-on-site — Viktor set BIOS AC-power-on on both boxes (attested; not OS-verifiable). roadmap-candidate: make BIOS "restore on AC power" a provisioning-checklist item + a host-install doc requirement for every fleet box.
F9 H1 OOB belt only partially installed on the current fleet. felhom-mgmt-watchdog.timer is live on both boxes (heal marker absent = no heal needed), but felhom-sshd.service and felhom-oob-nft.service (TASK H1, present in felhom-agent/configs/) are installed on neither box; operator access rides stock sshd :22 + tailscale + G1 break-glass. Pre-existing (both boxes provisioned via the universal ISO), surfaced by this audit's access preflight. MEDIUM p2-felhom-pve-ssh-belt.txt, p2-*-host.txt RESOLVED 2026-07-23 (ISO train v1.25.0). Ruling: install everywhere. host-install grows the belt as a DEFAULT appliance leg (--no-oob opts out; byo still refuses — deliberate). Phase-0 confirmed the omission was NOT a coded exclusion, just --enable-oob never passed by the universal ISO. Belt installed + oob.enabled on BOTH live boxes; login PROVEN end-to-end on felhom-pve (felhom-op@demo-felhom over wg-felhom → belt); also re-anchored the orphaned operator identity to the operator's real machine. See REPORT.md (2026-07-23) + operations/nodes.md.
F10 demo-hp backup tiers incomplete. Tier-2 dump tree exists (paperless-ngx, fresh), but no offbox target is configured and the PBS DR datastore holds 0 snapshots for it. A power event with disk damage would have had no off-box recovery path for that guest. Pre-existing (box added 07-21), not outage-caused. MEDIUM p6-demo-hp-offbox-locks.txt (NO-OFFBOX), p6-demo-hp-backup.txt (snapshots=0), p6-demo-hp-tier2.txt OFFSITE LEG RESOLVED 2026-07-23 — root cause was NOT "provisioning unfinished": the day-0 managed update killed the apply-bridge after password-consume, burning the one-shot credential (full diagnosis, designed-path repair via operator Re-issue, escrow ceremony, and a byte-identical offsite restore round-trip: DIAG-f10-demo-hp-offsite-2026-07-23.md; product rows minted R-70 visibility + R-71 the race). The PBS-DR-snapshot half STAYS OPEN pending F13 (cadence ruling) and the deliberate DR ceremony R-moment on demo-hp.
F11 Recovery is silent. host_recovered/node_recovered fired correctly but carry severity info, which the dispatcher deliberately does not email — the operator/customer only learns of recovery by looking. During a real customer outage the "it's back" signal is arguably the second-most valuable email. Working-as-coded, so a product question, not a defect. LOW p5-hub-events.txt; hub/internal/notify/dispatcher.go severityNotifies needs-ruling — opt-in recovery notifications (operator at least)?
F12 demo-hp has no customer notification prefs (customer_notifications has no row for it) → its "customer" received no node_down email and never would. Only demo-felhom is wired (doodoo21@freemail.hu). Pre-existing demo-box config gap; on a real onboarding this must not be skippable. LOW p5-customer-notification-prefs.txt roadmap-candidate — make notification-prefs setup a claim/onboarding step, not an optional settings page.
F13 PBS DR tier has no visible backup cadence. /etc/pve/jobs.cfg is empty on both hosts; N100's DR datastore holds a single CT-9201 snapshot from 2026-07-18 (S8-era), demo-hp none. The agent's 6 h verify loop runs, but nothing appears to create periodic DR snapshots. If DR snapshots are meant to be on-demand only, fine — but then the "DR tier" freshness expectation should be documented; if they're meant to be periodic, the scheduler is missing. LOW p6-*-backup.txt needs-ruling — clarify intended DR-tier cadence; roadmap if periodic.

5. Explicitly NOT validated (would have required mutation, or is operator-side)

  • BIOS auto-power-on: not OS-verifiable; recorded as operator-attested. Its first real proof will be the next power event (or a deliberate drill — candidate for R-55's reboot leg).
  • Email delivery to the Gmail inbox: the hub logged all 9 sends as sent (Resend accepted); whether they landed at ~14:5715:30 CEST is Viktor-side (inbox + Resend dashboard). Question for Viktor: did the 8 operator mails + 1 customer mail arrive? If not, that is a delivery finding (dispatcher is proven good).
  • Cloudflare dashboard tunnel state: operator-attested; this audit proved the stronger fact (edge-pinned 302 end-to-end on both live tunnels).
  • demo-hp offbox restic: nothing to check — no offbox exists (F10).
  • UI click-throughs: no browser on DooPlex; controller state was validated on its own API surface (endpoint-level, per standing convention).

6. Observations (out of scope, not acted on)

  • demo-vm tunnel will not recover — and that is correct. Its owner (the nested appliance) no longer exists: the three ISO-train drill VMs 9310/9311/9312 were destroyed cleanly this morning 11:0111:43 CEST, before the outage (p3-felhom-pve-vm-absence.txt). The tunnel's Down state is permanent until the dead hub/Cloudflare records are discarded (already pending as operator cleanup from the ISO train; hub-side ghost-delete = R-62 arc).
  • peti-felhom is silent since 2026-07-15 08:39 UTC (last report in the hub DB; hub seeded it down at its 09:25 restart, so no fresh event fired). Matches the petifelhom tunnel Down in Viktor's screenshot. His site, his box — noted for the planned Friday reinstall, nothing touched.
  • Controllers restarted 14:1314:14 CEST (controller_started events 1701/1702) — that was the 0.160.0 publish train, 28 minutes before the cut; unrelated to the outage, but a reminder that the fleet had been on 0.160.0 for less than half an hour when it took its first dirty shutdown — and came back clean.
  • Hub pod log prints CEST while the hub DB stores UTC; this audit normalized everything to CEST.

7. Bottom line

The platform passed its first real, unplanned, site-wide dirty-shutdown drill. Every layer — filesystems, host plane, WireGuard, tailscale, guests, Docker stacks, cloudflared, NAS automount, hub staleness detection, notifications, backup integrity — recovered without a single manual intervention, in ~3 m 15 s of actual work after power returned. The entire 4.5 h of customer-visible downtime was attributable to one thing: nobody could press the power button (F8, now mitigated). The dead-man's-switch pipeline fired on schedule and the emails dispatched; the open questions are delivery-side (inbox) and product-side (silent recovery, F11).