10 KiB
BIGNIGHT — a household's first month, compressed into one night (2026-09-14/15)
Interventions a customer could not have made: Phase 2 = 2, Phase 3 = 0.
Ready for a volunteer: NO — the dashboard link does not open through the tunnel (R-510), a box installed for an
existing customer gets no bind mail (R-509), and every box's file manager opens with admin / admin (R-513).
Brief: drills/BIGNIGHT-2026-09-14.md. Evidence: evidence-bignight-2026-09-14/ — journal.md (every observable in
order), alarm-truth-table.md, screens/, box-logs-phase2..4/, phase3/, phase4/, phase5/F1..F12, phase6/,
teardown-*.txt. Architecture read first: 00-capability-map.md, 07-backup-architecture.md §6 (the tiers),
09-update-architecture.md §3 (decisions 1–9).
Venue
VM 333 on demo-hp: ISO 1.27.1 (sha 25637007…, found in the build output, not rebuilt), q35/OVMF, 4 cores,
16 GB (the HP has 30 GB; 9201+9202 used ≈ 5.3 GB), system disk 200 G + data disk 100 G added after the install,
both qcow2 on nvme-scratch (/mnt/hdd_1, its root). Customer „Tester 1" (tester-1, enkicsifelhom.hu,
tester1@felhom.eu). Baselines: controller 406755fa v0.242.0 · agent 4586f0f7 v0.130.0 · felhom.eu a4d68441
hub v0.113.0 · catalog 6d6eec30. Harness substitutions as the earlier walks: U.S. keyboard (H2), auto-reboot
unticked and ISO detached (H3), text-mode installer (H4).
Off-site, as found: the record has the DR tier (PBS on ep0) ticked and restic off-site off; ticking it provisions a Hetzner Storage Box (money, fenced), so it was not ticked. The DR tier then could not provision on the new box (R-511). This box had no off-site tier of any kind; nothing was written to ep0.
Phase 2 — the first hour (mail real this time)
| step | result |
|---|---|
| install (one disk) | ≤ 2 m 45 s copying; the one-disk screen offers no choice; hostname and e-mail typed per the guide |
| first screen | Felhom Hungarian text only (no :8006 line) ✓; pairing banner repeats 3× after bind (R-214 class) |
| data disk | hot-added; nothing on dashboard, launcher or mail mentions it; found under Tárhely → „Új meghajtó inicializálása"; enrolled in 2 s. A household would not know to enrol it |
| bind mail | none in 10 min for an existing customer → R-509 (P1), I1: operator pressed „Send self-bind link"; mail arrived in 1 s |
| self-bind link | works end to end (first walk to exercise it): code + Tulajdonosi jelmondat → „Sikeres összekötés." |
| setup-code mail | arrived 54 s after bind — the reinstall mail („újratelepült … A korábbi jelszavad már nem érvényes"), not the first-install mail the guide names; the mailed code worked |
| version | controller 0.242.0 + agent 0.130.0, current, no self-update needed ✓ |
| gate: dashboard via tunnel | FAIL — 502 ×9: route reaches the box but lacks „No TLS Verify" → R-510 (P1), I2; LAN address used from then on |
Phase 3 — twelve apps, seeded through their front doors
All 12 deployed on the default box — the memory guard never refused (guest 11 828 MB; ≈ 3.5 GB used with 12 apps). Every app „Fut · Naprakész" (12/12). Seeds: BookStack 5 Hungarian pages + 2 attachments (sha equal) + 2nd user + delete/undo; Docmost space + 5 docs + rename/delete/restore; PrivateBin 5 encrypted pastes incl. 10-min expiry; Gokapi 3 uploads incl. 50 MB (stranger download sha equal); Nextcloud 200 JPEG + 20 PDF, share with 2nd user, delete
- trash restore (sha equal); Immich 200 photos, ML settled in ≈ 6 min; Vaultwarden 10 entries + attachment (client crypto); Paperless-ngx 20 PDFs → 0 documents (OOM, R-514); Jellyfin video via the product's SMB share → library scan 5 s → stream; Mealie 5 recipes + meal plan; Uptime Kuma 3 monitors (its first screen is English, R-516; the box's own names show DOWN through R-510); AdventureLog trip with visits (photos fail from a non-browser client — R-483 class). Interventions: 0.
Findings: R-512 Vaultwarden open signup with a read-only close control · R-513 (P1, security) FileBrowser
admin/admin on every box, demo-hp's login page public · R-514 Paperless OOM silently · R-515 Paperless card's
wrong login · R-516 English strings enumerated.
Phase 4 — a month of routines
Tier 1 run 2 m 08 s ✓ · Tier 2 12 apps in 17 s (drive apps state-only, as stated) ✓ · off-site: none to run or verify ·
whole-system „Mentés most": local 8.9 GB in 362 s ✓, then the absent PBS tier failed with every app stopped ≈ 7 m 45 s
under „csak néhány másodpercre" (R-518) and the page afterwards claimed a current full backup and a remote copy
that do not exist (R-517, P1). Hub: box ok, 24/24 containers, drive shown, true whole_guest_backup_failed mailed.
Guarded Update on a real bump (privatebin 2.0.5 → 2.0.6, catalog d5d91e0, reverted a161ccb in the same phase):
reached the box 14 m 25 s after the push, „Frissítés elérhető — ma", DONE in 11 s, data intact ✓.
Phase 5 — the accidents (five measures each: customer saw · box did · time · alarm true? · alarm missed)
| # | fault | what the customer saw | what the box did by itself | steady state | alarm fired / true? | should have fired, did not |
|---|---|---|---|---|---|---|
| F1 | power cut 61 s, family on 3 apps | all apps down ≈ 3½ min, re-login | all 12 back on same images; boot reconciler started paperless | 4 m 03 s | controller_started / true |
— |
| F2 | power cut during nightly backup (adventurelog stopped for its dump) | pages say „21:40 OK", no word of interruption | app-stop guard restarted adventurelog ✓; all 12 same images | 4 m 05 s | backup_failed (error) + mail / true |
customer notice (R-519) |
| F3 | power cut in „pulling" of a guarded Update (nextcloud, same version) | „Fut · Naprakész", nothing about the update | back on same images, pin consistent; no journal trace | 3 m 50 s | controller_started / true |
untestable same-version (R-520) |
| F4 | data drive unplugged under running apps | honest Hungarian: „Meghajtó leválasztva", „Hiányzó tárhely … Csatlakoztasd újra"; raw UTC time, banner ×2 | drive apps stopped by +38 s; system-disk apps kept running ✓; path failed closed (I/O error) | — | storage_disconnected (error) + 4 × app_start_failed, 5 mails / true, redundant |
household mail (R-521) |
| F5 | drive out 30 min, then back | badge „Aktív" at once; stale banner cleared within 10 min | re-bound as sdc by UUID, ext4 recovery, restarted the 4 apps |
91 s | storage_reconnected (info) / true |
— |
| F6 | drive unplugged during a backup | nothing about skipped apps | skipped 4 volume dumps but success:true; torn .tmp not promoted; next run honest and complete |
122 s after re-attach | storage_disconnected + 4, all mails suppressed by cooldown |
the second drive loss (R-521); the incomplete run (R-519) |
| F7 | system disk to 95 % | dashboard „90 % · Kritikusan kevés hely"; banner English „SSD disk usage high: 90%" | stayed reachable; uploads and a backup still succeeded; warnings cleared 3½ min after cleanup | — | health_degraded / true, mail suppressed by cooldown; no disk event |
operator told nothing (R-521) |
| F8 | internet gone 17½ min (venue: hub on the LAN, so the hub's ingress was blocked too) | LAN dashboard 200 throughout, no offline notice, tunnel tile „Fut" (R-522) | report push failed and backed off; on return pushed at once, tunnel re-registered +9 s (all 4 +58 s) | ≈ 1 min | node_stale + mail at 31 min since last report, node_recovered + mail / true |
customer notice (R-522) |
| F9 | docker kill the controller 4 s into a deploy |
dashboard 502 for 33 min, no page at all | nothing restarted it (unless-stopped ignores a kill; bootstrap is one-shot; agent silent); the deployed app itself came up healthy |
— until power-cycle | node_stale at 30 min, mail suppressed by cooldown |
everything: R-523 (P1) |
| F10–F12 | photo folder deleted · forgotten credentials · two quick reboots | not run | — | — | — | the brief's stop rule was met at F9 |
Stop rule, applied as written (22:08Z): F9 left the box in a state the product did not leave on its own and no customer screen can reach. Filed P1 (R-523), no further faults injected, the box recovered by a power-cycle.
Phase 6 — the morning after
All 12 apps (+ the throwaway homebox) running, none unhealthy. Labels: 11 true; privatebin „Frissítés elérhető" is
false — the box runs 2.0.6, the reverted catalog 2.0.5, and the offered Update is a downgrade (R-524). Backup
pages: DB copies and „Távoli rendszermentés nincs beállítva" honest; the whole-system tile „Naprakész" with no backup
(R-517); the „a few seconds" promise (R-518). Off-site restore onto 9202: not walked — no off-site copy exists on this
record. A local restore of BookStack after a deleted page: 24 s, page back, attachments sha equal. The alarm
truth table is evidence-bignight-2026-09-14/alarm-truth-table.md.
Phase 7 — teardown, three layers
- Machine (VM 333): destroyed with both disks at 22:25:58Z; ISO 1.27.1 and harness files removed from demo-hp; no
firewall table left.
pvesm status:nvme-scratchback to 10 140 556 KiB (10 134 820 before the drill),local23 071 188 KiB (23 041 736 before),local-lvm44.17 % unchanged all night. Evidenceteardown-before.txt,teardown-layer1-machine.txt.pct fstrimnot applied: only deleted qcow2 files on a dir storage were used. - Host (hub record + ep0 peer): host
tester-1-a61396deleted at 22:59:21Z once stale (the first attempt was refused for a missingconfirm_host_id— harness); host page 404; after the 23:04:13Z peer sync ep0 lists no peer10.77.0.5(read-only check). Evidenceteardown-layer3-hub.txt. - Hub customer: „Tester 1" is KEPT (the volunteer's record). Its off-site data on ep0 — namespace
tester-1, one snapshot directory from the doorstep walk — is kept, stated, for the operator to rule (R-511 context). - Untouched, and checked: demo-hp 9201 and 9202 (running throughout; only read-only loopback probes, R-513);
drill-r50(does not exist, R-461); DooPlex (read-only plus git pushes); Peti's box (not contacted); ep0 (read only).