Files
felhom.eu/documentation/audits/BIGNIGHT-household-month-2026-09-14.md
T

10 KiB
Raw Blame History

BIGNIGHT — a household's first month, compressed into one night (2026-09-14/15)

Interventions a customer could not have made: Phase 2 = 2, Phase 3 = 0. Ready for a volunteer: NO — the dashboard link does not open through the tunnel (R-510), a box installed for an existing customer gets no bind mail (R-509), and every box's file manager opens with admin / admin (R-513).

Brief: drills/BIGNIGHT-2026-09-14.md. Evidence: evidence-bignight-2026-09-14/ — journal.md (every observable in order), alarm-truth-table.md, screens/, box-logs-phase2..4/, phase3/, phase4/, phase5/F1..F12, phase6/, teardown-*.txt. Architecture read first: 00-capability-map.md, 07-backup-architecture.md §6 (the tiers), 09-update-architecture.md §3 (decisions 1–9).

Venue

VM 333 on demo-hp: ISO 1.27.1 (sha 25637007…, found in the build output, not rebuilt), q35/OVMF, 4 cores, 16 GB (the HP has 30 GB; 9201+9202 used ≈ 5.3 GB), system disk 200 G + data disk 100 G added after the install, both qcow2 on nvme-scratch (/mnt/hdd_1, its root). Customer „Tester 1" (tester-1, enkicsifelhom.hu, tester1@felhom.eu). Baselines: controller 406755fa v0.242.0 · agent 4586f0f7 v0.130.0 · felhom.eu a4d68441 hub v0.113.0 · catalog 6d6eec30. Harness substitutions as the earlier walks: U.S. keyboard (H2), auto-reboot unticked and ISO detached (H3), text-mode installer (H4).

Off-site, as found: the record has the DR tier (PBS on ep0) ticked and restic off-site off; ticking it provisions a Hetzner Storage Box (money, fenced), so it was not ticked. The DR tier then could not provision on the new box (R-511). This box had no off-site tier of any kind; nothing was written to ep0.

Phase 2 — the first hour (mail real this time)

step result
install (one disk) ≤ 2 m 45 s copying; the one-disk screen offers no choice; hostname and e-mail typed per the guide
first screen Felhom Hungarian text only (no :8006 line) ✓; pairing banner repeats 3× after bind (R-214 class)
data disk hot-added; nothing on dashboard, launcher or mail mentions it; found under Tárhely → „Új meghajtó inicializálása"; enrolled in 2 s. A household would not know to enrol it
bind mail none in 10 min for an existing customer → R-509 (P1), I1: operator pressed „Send self-bind link"; mail arrived in 1 s
self-bind link works end to end (first walk to exercise it): code + Tulajdonosi jelmondat → „Sikeres összekötés."
setup-code mail arrived 54 s after bind — the reinstall mail („újratelepült … A korábbi jelszavad már nem érvényes"), not the first-install mail the guide names; the mailed code worked
version controller 0.242.0 + agent 0.130.0, current, no self-update needed ✓
gate: dashboard via tunnel FAIL — 502 ×9: route reaches the box but lacks „No TLS Verify" → R-510 (P1), I2; LAN address used from then on

Phase 3 — twelve apps, seeded through their front doors

All 12 deployed on the default box — the memory guard never refused (guest 11 828 MB; ≈ 3.5 GB used with 12 apps). Every app „Fut · Naprakész" (12/12). Seeds: BookStack 5 Hungarian pages + 2 attachments (sha equal) + 2nd user + delete/undo; Docmost space + 5 docs + rename/delete/restore; PrivateBin 5 encrypted pastes incl. 10-min expiry; Gokapi 3 uploads incl. 50 MB (stranger download sha equal); Nextcloud 200 JPEG + 20 PDF, share with 2nd user, delete

  • trash restore (sha equal); Immich 200 photos, ML settled in ≈ 6 min; Vaultwarden 10 entries + attachment (client crypto); Paperless-ngx 20 PDFs → 0 documents (OOM, R-514); Jellyfin video via the product's SMB share → library scan 5 s → stream; Mealie 5 recipes + meal plan; Uptime Kuma 3 monitors (its first screen is English, R-516; the box's own names show DOWN through R-510); AdventureLog trip with visits (photos fail from a non-browser client — R-483 class). Interventions: 0.

Findings: R-512 Vaultwarden open signup with a read-only close control · R-513 (P1, security) FileBrowser admin/admin on every box, demo-hp's login page public · R-514 Paperless OOM silently · R-515 Paperless card's wrong login · R-516 English strings enumerated.

Phase 4 — a month of routines

Tier 1 run 2 m 08 s ✓ · Tier 2 12 apps in 17 s (drive apps state-only, as stated) ✓ · off-site: none to run or verify · whole-system „Mentés most": local 8.9 GB in 362 s ✓, then the absent PBS tier failed with every app stopped ≈ 7 m 45 s under „csak néhány másodpercre" (R-518) and the page afterwards claimed a current full backup and a remote copy that do not exist (R-517, P1). Hub: box ok, 24/24 containers, drive shown, true whole_guest_backup_failed mailed. Guarded Update on a real bump (privatebin 2.0.5 → 2.0.6, catalog d5d91e0, reverted a161ccb in the same phase): reached the box 14 m 25 s after the push, „Frissítés elérhető — ma", DONE in 11 s, data intact ✓.

Phase 5 — the accidents (five measures each: customer saw · box did · time · alarm true? · alarm missed)

# fault what the customer saw what the box did by itself steady state alarm fired / true? should have fired, did not
F1 power cut 61 s, family on 3 apps all apps down ≈ 3½ min, re-login all 12 back on same images; boot reconciler started paperless 4 m 03 s controller_started / true —
F2 power cut during nightly backup (adventurelog stopped for its dump) pages say „21:40 OK", no word of interruption app-stop guard restarted adventurelog ✓; all 12 same images 4 m 05 s backup_failed (error) + mail / true customer notice (R-519)
F3 power cut in „pulling" of a guarded Update (nextcloud, same version) „Fut · Naprakész", nothing about the update back on same images, pin consistent; no journal trace 3 m 50 s controller_started / true untestable same-version (R-520)
F4 data drive unplugged under running apps honest Hungarian: „Meghajtó leválasztva", „Hiányzó tárhely … Csatlakoztasd újra"; raw UTC time, banner ×2 drive apps stopped by +38 s; system-disk apps kept running ✓; path failed closed (I/O error) — storage_disconnected (error) + 4 × app_start_failed, 5 mails / true, redundant household mail (R-521)
F5 drive out 30 min, then back badge „Aktív" at once; stale banner cleared within 10 min re-bound as sdc by UUID, ext4 recovery, restarted the 4 apps 91 s storage_reconnected (info) / true —
F6 drive unplugged during a backup nothing about skipped apps skipped 4 volume dumps but success:true; torn .tmp not promoted; next run honest and complete 122 s after re-attach storage_disconnected + 4, all mails suppressed by cooldown the second drive loss (R-521); the incomplete run (R-519)
F7 system disk to 95 % dashboard „90 % · Kritikusan kevés hely"; banner English „SSD disk usage high: 90%" stayed reachable; uploads and a backup still succeeded; warnings cleared 3½ min after cleanup — health_degraded / true, mail suppressed by cooldown; no disk event operator told nothing (R-521)
F8 internet gone 17½ min (venue: hub on the LAN, so the hub's ingress was blocked too) LAN dashboard 200 throughout, no offline notice, tunnel tile „Fut" (R-522) report push failed and backed off; on return pushed at once, tunnel re-registered +9 s (all 4 +58 s) ≈ 1 min node_stale + mail at 31 min since last report, node_recovered + mail / true customer notice (R-522)
F9 docker kill the controller 4 s into a deploy dashboard 502 for 33 min, no page at all nothing restarted it (unless-stopped ignores a kill; bootstrap is one-shot; agent silent); the deployed app itself came up healthy — until power-cycle node_stale at 30 min, mail suppressed by cooldown everything: R-523 (P1)
F10–F12 photo folder deleted · forgotten credentials · two quick reboots not run — — — the brief's stop rule was met at F9

Stop rule, applied as written (22:08Z): F9 left the box in a state the product did not leave on its own and no customer screen can reach. Filed P1 (R-523), no further faults injected, the box recovered by a power-cycle.

Phase 6 — the morning after

All 12 apps (+ the throwaway homebox) running, none unhealthy. Labels: 11 true; privatebin „Frissítés elérhető" is false — the box runs 2.0.6, the reverted catalog 2.0.5, and the offered Update is a downgrade (R-524). Backup pages: DB copies and „Távoli rendszermentés nincs beállítva" honest; the whole-system tile „Naprakész" with no backup (R-517); the „a few seconds" promise (R-518). Off-site restore onto 9202: not walked — no off-site copy exists on this record. A local restore of BookStack after a deleted page: 24 s, page back, attachments sha equal. The alarm truth table is evidence-bignight-2026-09-14/alarm-truth-table.md.

Phase 7 — teardown, three layers

  • Machine (VM 333): destroyed with both disks at 22:25:58Z; ISO 1.27.1 and harness files removed from demo-hp; no firewall table left. pvesm status: nvme-scratch back to 10 140 556 KiB (10 134 820 before the drill), local 23 071 188 KiB (23 041 736 before), local-lvm 44.17 % unchanged all night. Evidence teardown-before.txt, teardown-layer1-machine.txt. pct fstrim not applied: only deleted qcow2 files on a dir storage were used.
  • Host (hub record + ep0 peer): host tester-1-a61396 deleted at 22:59:21Z once stale (the first attempt was refused for a missing confirm_host_id — harness); host page 404; after the 23:04:13Z peer sync ep0 lists no peer 10.77.0.5 (read-only check). Evidence teardown-layer3-hub.txt.
  • Hub customer: „Tester 1" is KEPT (the volunteer's record). Its off-site data on ep0 — namespace tester-1, one snapshot directory from the doorstep walk — is kept, stated, for the operator to rule (R-511 context).
  • Untouched, and checked: demo-hp 9201 and 9202 (running throughout; only read-only loopback probes, R-513); drill-r50 (does not exist, R-461); DooPlex (read-only plus git pushes); Peti's box (not contacted); ep0 (read only).