Files
felhom.eu/REPORT-bignight-2026-09-14.md

5.5 KiB
Raw Permalink Blame History

REPORT — BIGNIGHT: a household's first month in one night (2026-09-14/15)

Unattended drill run from drills/BIGNIGHT-2026-09-14.md under .claude/rules/unprompted-work.md. No product code changed. Findings: documentation/audits/BIGNIGHT-household-month-2026-09-14.md. Every observable in order: documentation/audits/evidence-bignight-2026-09-14/journal.md. The alarm truth table: …/evidence-bignight-2026-09-14/alarm-truth-table.md. A parallel session may own root REPORT.md; this is a topic sibling.

0. Where the brief and the record disagreed — named first

  1. „Off-site (Tier 3) is ON for this customer — its own namespace on ep0." On the record the ep0 namespace is the DR tier (PBS); restic Tier 3 was off, and ticking it provisions a Hetzner Storage Box (money — fenced). Not ticked. The DR tier then could not provision on the new box (R-511). This box had no off-site tier of any kind; Phase 4's off-site integrity check and Phase 6's off-site restore onto 9202 were therefore not walked.
  2. „Expect the hub to issue a fresh claim, or to require its reset flow." Neither: the hub treated the box as a re-enrolment and mailed the reinstall setup code („újratelepült … A korábbi jelszavad már nem érvényes").
  3. „The operator fixed and tested the tunnel." The route now reaches the box, and still returns 502: it lacks „No TLS Verify" (R-510, with demo-hp's working route as the control).
  4. Faults F10–F12 were not run — the brief's own stop rule was met at F9 (R-523).

1. Baselines

controller 406755fa v0.242.0 · agent 4586f0f7 v0.130.0 · felhom.eu a4d68441 hub v0.113.0 · catalog 6d6eec30. ISO 1.27.1 sha 25637007… found in the build output, not rebuilt. Venue: VM 333 on demo-hp, 4 cores, 16 GB, 200 G + 100 G qcow2 on nvme-scratch (/mnt/hdd_1 root).

2. What changed in the repos

repo change
felhom.eu 16 register rows R-509 … R-524; amendments to R-516, R-517, R-519, R-521, R-523; audit page; evidence directory; capability-map annotations (self-bind row, first-hour row); STATUS morning note; this report. Documents only.
app-catalog-felhom.eu drill bump d5d91e0 privatebin 2.0.5 → 2.0.6 and its revert a161ccb in the same phase; two CHANGELOG entries. Net template change: none (catalog_since stays 2026-09-14 by the gate).
felhom-controller, felhom-agent nothing

3. Results in one table

phase result
2 first hour install ✓ · Hungarian first screen ✓ · self-bind via real mail ✓ · mailed setup code ✓ · version current ✓ · tunnel gate FAIL · data disk: no screen tells a household · interventions 2
3 twelve apps all deployed and seeded through their front doors · memory guard never refused · 12/12 „Naprakész" · interventions 0 · Paperless lost 20 uploads to OOM · FileBrowser admin/admin on every box
4 routines Tier 1 ✓ · Tier 2 ✓ · whole-system local ✓ but apps down 8 min and a false PBS claim · guarded Update on a real bump ✓ 11 s, data intact · catalog reverted
5 faults F1 F2 F3 power cuts heal ≈ 4 min · F4 F5 drive pull/return honest, heals 91 s · F6 drive lost in backup: skipped apps reported success, alarm mails silenced · F7 disk 95 % holds, English banner, operator not told · F8 internet gone: LAN works, tunnel self-heals 9 s · F9 controller killed: dead 33 min, nobody told — STOP
6 morning after apps healthy · 1 false label (downgrade offered as update) · local BookStack restore ✓ 24 s
7 teardown machine: VM 333 + disks, ISO, harness files removed, storage back to pre-drill levels · host: tester-1-a61396 deleted, ep0 peer gone · hub: customer tester-1 kept; its ep0 data (1 snapshot dir) kept, stated · 9201/9202 untouched · secrets shredded

4. Rows (register 221 → 237)

P1: R-509 no auto bind mail for an existing customer · R-510 tunnel route lacks No TLS Verify · R-513 FileBrowser admin/admin, demo-hp public · R-517 backup page claims a failed PBS tier current and present · R-523 killed controller never restarts. P2: R-511 DR tier stuck after a rebuild · R-512 Vaultwarden open signup, read-only control · R-514 Paperless OOM silent · R-518 whole-system backup stops apps 8 min · R-519 torn backup dated by its newest part · R-524 downgrade offered as update. P3: R-515 Paperless card's wrong login · R-516 English strings · R-520 interrupted update untestable same-version · R-521 alarm mail noise and cooldown silence · R-522 tunnel tile „Fut" while offline.

5. Harness slips, recorded

API key printed once into tool output (hub customer page read) · two quoted-string inserts into the register failed and were redone · first claim POST sent two CSRF tokens · several poll loops read a stale status and stopped early or ran long (guest backup, Tier 2, restore) · F8's first two attempts cut nothing (nft reserved word; the hub resolves to the LAN) and the measured cut lasted 17½ min, not 20 · time waiters broke at local midnight (date -d HH:MMZ) · AdventureLog account took five attempts. None changed a finding; each is in the journal where it happened.

6. Security note

A FileBrowser login (admin/admin) was tested against demo-hp 9201 and 9202 over loopback only; demo-hp's public login page was checked with a GET and no login. Nothing was changed on either guest. The operator was told by the morning note; the push notification was not sent because the terminal was active.