BIGNIGHT morning note (STATUS), topic report, capability-map annotations; teardown layer 1 done, layer 3 pending
gates / gates (push) Successful in 19s
gates / gates (push) Successful in 19s
This commit is contained in:
@@ -0,0 +1,66 @@
|
||||
# REPORT — BIGNIGHT: a household's first month in one night (2026-09-14/15)
|
||||
|
||||
**Unattended drill run from `drills/BIGNIGHT-2026-09-14.md` under `.claude/rules/unprompted-work.md`. No product code
|
||||
changed.** Findings: `documentation/audits/BIGNIGHT-household-month-2026-09-14.md`. Every observable in order:
|
||||
`documentation/audits/evidence-bignight-2026-09-14/journal.md`. The alarm truth table:
|
||||
`…/evidence-bignight-2026-09-14/alarm-truth-table.md`. A parallel session may own root `REPORT.md`; this is a topic sibling.
|
||||
|
||||
## 0. Where the brief and the record disagreed — named first
|
||||
|
||||
1. **„Off-site (Tier 3) is ON for this customer — its own namespace on ep0."** On the record the ep0 namespace is the
|
||||
**DR tier (PBS)**; restic Tier 3 was **off**, and ticking it provisions a Hetzner Storage Box (money — fenced). Not
|
||||
ticked. The DR tier then could not provision on the new box (R-511). This box had **no off-site tier of any kind**;
|
||||
Phase 4's off-site integrity check and Phase 6's off-site restore onto 9202 were therefore not walked.
|
||||
2. **„Expect the hub to issue a fresh claim, or to require its reset flow."** Neither: the hub treated the box as a
|
||||
**re-enrolment** and mailed the *reinstall* setup code („újratelepült … A korábbi jelszavad már nem érvényes").
|
||||
3. **„The operator fixed and tested the tunnel."** The route now reaches the box, and still returns 502: it lacks „No
|
||||
TLS Verify" (R-510, with demo-hp's working route as the control).
|
||||
4. **Faults F10–F12 were not run** — the brief's own stop rule was met at F9 (R-523).
|
||||
|
||||
## 1. Baselines
|
||||
|
||||
controller `406755fa` v0.242.0 · agent `4586f0f7` v0.130.0 · felhom.eu `a4d68441` hub v0.113.0 · catalog `6d6eec30`.
|
||||
ISO 1.27.1 sha `25637007…` found in the build output, not rebuilt. Venue: VM 333 on demo-hp, 4 cores, 16 GB, 200 G +
|
||||
100 G qcow2 on `nvme-scratch` (`/mnt/hdd_1` root).
|
||||
|
||||
## 2. What changed in the repos
|
||||
|
||||
| repo | change |
|
||||
|---|---|
|
||||
| felhom.eu | 16 register rows **R-509 … R-524**; amendments to R-516, R-517, R-519, R-521, R-523; audit page; evidence directory; capability-map annotations (self-bind row, first-hour row); STATUS morning note; this report. Documents only. |
|
||||
| app-catalog-felhom.eu | drill bump `d5d91e0` privatebin 2.0.5 → 2.0.6 and its revert `a161ccb` in the same phase; two CHANGELOG entries. Net template change: none (`catalog_since` stays 2026-09-14 by the gate). |
|
||||
| felhom-controller, felhom-agent | nothing |
|
||||
|
||||
## 3. Results in one table
|
||||
|
||||
| phase | result |
|
||||
|---|---|
|
||||
| 2 first hour | install ✓ · Hungarian first screen ✓ · self-bind via real mail ✓ · mailed setup code ✓ · version current ✓ · **tunnel gate FAIL** · data disk: no screen tells a household · **interventions 2** |
|
||||
| 3 twelve apps | all deployed and seeded through their front doors · memory guard never refused · 12/12 „Naprakész" · **interventions 0** · Paperless lost 20 uploads to OOM · FileBrowser `admin/admin` on every box |
|
||||
| 4 routines | Tier 1 ✓ · Tier 2 ✓ · whole-system local ✓ but apps down 8 min and a false PBS claim · guarded Update on a real bump ✓ 11 s, data intact · catalog reverted |
|
||||
| 5 faults | F1 F2 F3 power cuts heal ≈ 4 min · F4 F5 drive pull/return honest, heals 91 s · F6 drive lost in backup: skipped apps reported success, alarm mails silenced · F7 disk 95 % holds, English banner, operator not told · F8 internet gone: LAN works, tunnel self-heals 9 s · **F9 controller killed: dead 33 min, nobody told — STOP** |
|
||||
| 6 morning after | apps healthy · 1 false label (downgrade offered as update) · local BookStack restore ✓ 24 s |
|
||||
| 7 teardown | layer 1 done (VM, disks, ISO, harness files); layer 2 host delete + ep0 peer: see the audit's teardown section; customer `tester-1` **kept**; its ep0 data **kept, stated** |
|
||||
|
||||
## 4. Rows (register 221 → 237)
|
||||
|
||||
P1: **R-509** no auto bind mail for an existing customer · **R-510** tunnel route lacks No TLS Verify · **R-513** FileBrowser
|
||||
admin/admin, demo-hp public · **R-517** backup page claims a failed PBS tier current and present · **R-523** killed
|
||||
controller never restarts. P2: R-511 DR tier stuck after a rebuild · R-512 Vaultwarden open signup, read-only control ·
|
||||
R-514 Paperless OOM silent · R-518 whole-system backup stops apps 8 min · R-519 torn backup dated by its newest part ·
|
||||
R-524 downgrade offered as update. P3: R-515 Paperless card's wrong login · R-516 English strings · R-520 interrupted
|
||||
update untestable same-version · R-521 alarm mail noise and cooldown silence · R-522 tunnel tile „Fut" while offline.
|
||||
|
||||
## 5. Harness slips, recorded
|
||||
|
||||
API key printed once into tool output (hub customer page read) · two quoted-string inserts into the register failed
|
||||
and were redone · first claim POST sent two CSRF tokens · several poll loops read a stale status and stopped early or
|
||||
ran long (guest backup, Tier 2, restore) · F8's first two attempts cut nothing (nft reserved word; the hub resolves to
|
||||
the LAN) and the measured cut lasted 17½ min, not 20 · time waiters broke at local midnight (`date -d HH:MMZ`) ·
|
||||
AdventureLog account took five attempts. None changed a finding; each is in the journal where it happened.
|
||||
|
||||
## 6. Security note
|
||||
|
||||
A FileBrowser login (`admin`/`admin`) was tested against demo-hp 9201 and 9202 **over loopback only**; demo-hp's public
|
||||
login page was checked with a GET and no login. Nothing was changed on either guest. The operator was told by the
|
||||
morning note; the push notification was not sent because the terminal was active.
|
||||
@@ -1,5 +1,55 @@
|
||||
# STATUS — what works, what's broken, what's next
|
||||
|
||||
**Updated 2026-09-15 (morning note) — the big night: a household's first month on one fresh box.**
|
||||
|
||||
> **Ready for a volunteer: not yet — three things stop them.** The dashboard link still does not open through
|
||||
> the tunnel. A box installed for an existing customer gets no bind e-mail. And every box's file manager opens
|
||||
> with the login „admin" / „admin" — on the HP that login page is reachable from the internet.
|
||||
|
||||
**Decisions I took.** None under the unattended-decision rule. Two things I did not do, each with one reason: I did
|
||||
not switch on the paid off-site storage for „Tester 1" (it costs money), so this box had no off-site copy and the
|
||||
off-site restore could not be walked. I stopped injecting faults after the ninth, because the brief's stop rule was met.
|
||||
|
||||
**Interventions (first hour + moving in): 2.** (1) No bind e-mail came for 10 minutes; I pressed the operator's
|
||||
„send link" button. (2) The tunnel answered 502; I used the home-network address for the rest of the night.
|
||||
|
||||
**Alarm truth table, in five lines.**
|
||||
1. Every alarm that fired was true. None was false.
|
||||
2. Missed: the Paperless crash, the broken tunnel, the dead controller, the full disk, and the second drive loss.
|
||||
3. Three of those were missed because a 30–60 minute mail cooldown silenced a new incident.
|
||||
4. One drive loss sent five operator mails; the household got no mail for anything all night.
|
||||
5. The customer's pages were honest about drives, and wrong about backups twice.
|
||||
|
||||
**What broke, and whether the box healed itself.** Power cuts (three, one during a backup, one during an update): healed
|
||||
in about four minutes, same versions, data intact. Drive pulled and returned: healed in 91 seconds, data intact. Internet
|
||||
gone: the tunnel came back by itself in 9 seconds. Disk 95 % full: the box kept working. **Controller killed: it did not
|
||||
heal — no dashboard for 33 minutes, nobody told, only a reboot brought it back. That was the stop.** Also: Paperless
|
||||
silently lost 20 uploads to memory; the backup page claimed a remote backup that does not exist; the whole-system backup
|
||||
stopped every app for 8 minutes while promising „a few seconds"; after I reverted a test update in the catalog, the box
|
||||
offered the downgrade as an update. Nothing lost data that a restore could not bring back. I fixed nothing tonight.
|
||||
|
||||
**Rows.** Opened 16, closed 0. Register: 221 rows before, 237 after.
|
||||
|
||||
**The one sentence for recruiting.** Felhom survived power cuts, a pulled drive and a lost internet by itself tonight,
|
||||
but do not invite anyone until the tunnel works, the file-manager password is changed on every box, and a dead
|
||||
controller restarts itself.
|
||||
|
||||
**Needs you.**
|
||||
1. **Change the file manager's admin password on the HP now** (and on the N100). **If you do nothing:** anyone on
|
||||
the internet who tries „admin" / „admin" at the HP's files address can read its data drive.
|
||||
2. **Tick „No TLS Verify" on the „Tester 1" tunnel route in Cloudflare.** **If you do nothing:** no volunteer can open
|
||||
their dashboard from outside their home.
|
||||
3. **Decide whether a box installed for an existing customer should get the bind e-mail automatically**, or you press
|
||||
the button each time. **If you do nothing:** each new install waits for you.
|
||||
4. **Say whether installer 1.27.1 should be published.** My recommendation: **yes, publish it** — nothing tonight was
|
||||
the installer's fault; the install, first screen and bind all worked. The three blockers above are on the tunnel,
|
||||
the controller and the file manager, and they block inviting people, not the installer. **If you do nothing:** the
|
||||
old installer with the English admin line stays online.
|
||||
5. **Decide what to do with „Tester 1"'s old off-site data on ep0** (still there, kept tonight) and its stuck DR tier.
|
||||
**If you do nothing:** the data stays, and every new box for this customer gets a failed whole-system backup alarm.
|
||||
|
||||
---
|
||||
|
||||
**Updated 2026-09-14 (evening) — the doorstep: installer fixed, walked again, NOT published.**
|
||||
|
||||
> **Ready for a volunteer: not yet — one thing stops them.** On the test customer you chose, the
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -678,3 +678,18 @@ attachments sha equal.** PASS. (Harness: the poll loop missed the status's `last
|
||||
taken separately at 22:24:57Z.)
|
||||
|
||||
Phase 6 evidence pulled off the box 22:26Z (`box-logs-phase6/`), before any teardown.
|
||||
|
||||
## Phase 7 — teardown, three layers
|
||||
|
||||
**Layer 1 — the machine** (`teardown-before.txt`, `teardown-layer1-machine.txt`), 22:25:58Z: `qm stop 333`, `qm destroy
|
||||
333 --purge 1 --destroy-unreferenced-disks 1` (system and data qcow2 on `nvme-scratch` removed; `/mnt/hdd_1/images/`
|
||||
now holds only `9202`); ISO 1.27.1 removed from `local:iso`; `/root/bn` harness files removed; nft: only `inet
|
||||
felhom_oob` (the drill's bridge table was already deleted at the end of F8). `pvesm status` before → after:
|
||||
`nvme-scratch` 59 436 396 → **10 140 556 KiB** (before the drill 10 134 820); `local` 24 728 328 → **23 071 188 KiB**
|
||||
(before the drill 23 041 736); `local-lvm` 44.17 % unchanged all night. `qm list` empty; **9201 and 9202 running**.
|
||||
`pct fstrim`: not applied — the drill used only qcow2 files on a dir storage, now deleted; 9201 and 9202 were not
|
||||
touched and are not trimmed.
|
||||
|
||||
**ep0 read-only before the host delete** (`teardown-ep0-before-host-delete.txt`, 22:26:02Z): WireGuard peer
|
||||
`10.77.0.5/32` present; namespaces `demo-felhom demo-hp tester-1`; `tester-1/ct/9201` with 1 snapshot directory (the
|
||||
doorstep walk's data). Nothing was written to ep0 tonight.
|
||||
|
||||
Reference in New Issue
Block a user