Files
felhom.eu/documentation/audits/DRILL-new-household-2026-09-30.md
T
admin cdf974d186
gates / gates (push) Successful in 28s
DRILL new household on golden 0.282.0: 0 interventions; ready for a first real tester on a new record
Audit doc, STATUS one sentence, capability-map first-hour row (walk re-proven, day-one
off-site sentence narrowed: R-720/R-726/R-727), teardown in four layers, R-600 measured again.
Secret scan over all audits: 0 hits, positive control 1.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-30 09:25:17 +02:00

14 KiB
Raw Blame History

DRILL — a new household's first day on golden 0.282.0 (2026-09-29/30)

Interventions a volunteer could not have made: 0. Ready for a first real tester: YES, on a new customer record — nothing stopped the walk, from the download to a restored app; but send the guide only after its five stale lines are fixed (R-722), and know that their apps get NO off-site copy until someone switches each one on (R-720).

Evidence: evidence-drill-new-household-2026-09-30/ — journal.md (every observable, in order), golden/, phase0/, phase1/, phase2/, screens/ (codes blacked out), box-logs-phase1/, box-logs-final/, teardown/. Golden record: ../tests/golden-0.282.0-2026-09-29/. Architecture read first: 01 §5 (who may reach an app), 07 §6.5–§6.6, 09 §3 decisions 11–49, FIRST-ADMIN.md, the first-hour row of 00-capability-map.md.


Claims in the brief that turned out wrong (or right), named first

claim verdict
The box lands on 0.282.0 from the golden RIGHT. Controller 0.282.0 + agent 0.137.0, the vouched set; selfupdate: Current version 0.282.0 is up to date; no self-update.
A one-level-deeper app address (<app>.<sub>.felhom.eu) fails TLS at the edge CONSISTENT, NOT MEASURED END TO END. Every working customer domain is its own zone and the edge serves *.enkisfelhom.hu, enkisfelhom.hu only; felhom.eu's own names are DNS-only. A proxied deeper name would get *.felhom.eu, which cannot match. No record was made to prove it (no Cloudflare key; the operator ruled the drill onto an existing zone).
The gate never re-gates a set-up app after a kept-data reinstall THE PREMISE DID NOT HOLD for the app walked. A volume-only app has no „A megőrzött adataimat használom / Tiszta lappal kezdem" choice — the volume is always deleted and only the backup is kept. The reinstall came back empty and gated (correct); after the household restored its backup the data returned and the gate opened by itself 10 s later. The household never sees a set-up app gated over its own data. The drive-kept path was not walked (no data drive was set up).
A fresh customer's box runs an off-site leg on night one WRONG, twice. The leg ran and copied nothing it could keep: (1) apps start with off-site OFF and nothing tells the household to switch them on (R-720); (2) on this customer, who had a box before, the repository was found orphaned and every run skips until the household presses a reset nobody pointed them to (R-726).
The claim code is one-shot RIGHT. The spent setup code → „Hibás vagy lejárt kód"; the bind link survived one typo and then bound.

1. What held, in one table

# step result time
0.1 golden 0.282.0 baked, round-trip verified, vouched; R-120 refused 0.276.0 after PASS bake 7 min
0.2 operator's part customer tester-1 (operator ruling); tunnel already on the record; passphrase from the page; „Send self-bind link" pressed (R-719) —
1 download from felhom.eu/letoltes PASS — 1 705 324 544 B, sha == page == release record 46 s
2 install (guide answers, Hungarian keyboard kept) PASS — two screens differ from the guide (pre-selected disk; focus on „Previous") ≤ 7 m 27 s copying
3 first boot → the mailed link → self-bind page PASS — registered 31 s after power-on; one typo refused; bound; console turns to „a doboz össze van kötve" bind → controller 2 m 51 s
4 „ready", versions PASS — 0.282.0 / 0.137.0 from the golden —
5 dashboard through the public tunnel; recovery code PASS — no LAN detour (R-505 closed); code ready 12 min after the claim, made in 5 s, shown once, 410 after power-on → claimed 27 m 51 s
6 three apps: random first password / gated with a probe / gated with a sign-up lock PASS — 65 s / 23 s / 12 s; the stranger saw only the gate page; „Kész" before the setup refused 409; gates opened (probe ~12 s; press) —
7 use for ~10 min; a family member through the 15-minute window PASS — BookStack page + attachment, Vikunja project + tasks + file (sha equal); stranger 403 before and after the window, family member in —
8 backups page, „Mentés most" PASS — true dates on both pages, no customer alarm; every app reads „3. mentés Kikapcsolva" (R-720) 19 s
9 „Naprakész" TRUE — installed == catalog; no newer tested step offered —
10 remove keeping data, reinstall PASS with a defect — a Stop during the first whole-guest backup was undone by that backup (R-721); no kept-data choice for a volume app; restore brought the data back and the gate lifted itself —
11 restore the password app PASS — page + attachment byte-identical, the shown password works, same version 26 s
12 remove with „delete my data too" PASS — volume, backup, restore points gone; front door 404; „Nincs megőrzött adat" —
13 status pages as a week-old household small disagreements (R-724) —
A1 power cut (60 s dark) PASS — same versions, gates kept their state, data equal, no alarm; one re-login +123 s to dashboard
A2 typos: pairing code on the bind page; the setup code on „Elfelejtett jelszó" PASS — both refused in Hungarian, the right code accepted right after; spent code refused —
A3 a phone opens a gated app first PASS — Hungarian gate page → „Bejelentkezés" → dashboard login → straight back into the app; a native app's call gets English JSON (R-725) —

2. The operator's part — every step the volunteer depends on

  1. The customer: tester-1 (operator ruling during the run, instead of a new drill record) — domain enkicsifelhom.hu, tester1@felhom.eu, off-site shared 100 GB ON, DR tier ON.
  2. The domain and tunnel: already on the record (R-505, 2026-09-14). No Cloudflare API token exists in the operator's credentials file; nothing in Cloudflare was created or changed.
  3. The owner passphrase: read from the customer page's reveal field (5 words), handed over out of band.
  4. „Send self-bind link" — pressed at 18:54:51 UTC. The guide says no press is needed; for a customer whose last link expired (2026-09-24) nothing re-sends one when a new box registers (R-719).

Everything after that was the volunteer's own: both mails arrived in the tester's mailbox and were followed.

3. The first night, leg by leg (box time CEST; UTC in brackets)

leg what happened why
database dump 02:30 (00:30) RAN, OK — 1 database, 3 volume dumps, 15 s; recovery units for all three apps —
second-drive copy 03:30 (01:30) not run — „nincs másik fizikai meghajtó — a 2. mentéshez 2. meghajtó szükséges" not configured: no second drive assigned
off-site copy 04:15 (02:15) SKIPPED — offsite repo ORPHANED … runs will skip until reset; household timeline + operator mail the customer's previous box left its repository under a destroyed key (R-726); and only BookStack was switched on at all (R-720)
app update leg (after off-site) RAN — done=0 undone=0 held=0 failed=0 nothing newer in the catalog
whole-guest backup (agent) not run tonight — no new archive on either storage not due: local tier ran 21:27 CEST (24 h cadence), off-site PBS tier 21:37 CEST (7-day cadence)
restore test (agent) 03:52 (01:52) FAILED — picked a 2026-09-16 archive of an earlier drill box in the same namespace, wrong key; operator mail; the household sees „✗ Visszaállítás ellenőrizve — Helyi tároló (local)" R-727
off-site proof / integrity 05:30 / 06:00 checked nothing — the repository is orphaned (said so, not called a failure) follows R-726

So a new household's box does not fully protect itself on night one. The local layers did (database, volumes, whole-guest on the evening before); the off-site app copy did not, for two separate reasons, and the restore test proved nothing.

4. Findings

row rank one line
R-719 P2 An existing customer's new box gets no fresh connect link — the operator must press „Send self-bind link"
R-720 P2 Apps start with off-site OFF and nothing tells the household to switch them on — night one copies nothing off-site
R-721 P2 A household's Stop during a whole-guest backup is undone by that backup's resume
R-722 P2 The volunteer guide is stale in five places (BookStack's old default login is now refused, among them)
R-726 P2 A customer with an earlier box: the off-site repository is orphaned on night one; the fix is an unannounced button
R-727 P2 The restore test picks a previous box's archive, fails on its key; the page blames the local tier
R-723 P3 Two operator mails on day one that describe nothing wrong (node_recovered, backup_tier_skipped)
R-724 P3 Status pages disagree (schedule, raw timestamp, LAN unknown, „0 órája", R-500's UTC)
R-725 P3 Copy slips (bind page „a beállításkor", formal wizard, stray console glyph, English JSON to phone apps)
R-718 P3 measured again: the gate-open press restarts an app with its own switch (2 s 404) unannounced

Closed by this walk: R-505 — the box-side tunnel hop, proven end to end on a fresh box. Held: R-214 (console after bind), R-496/R-497 (console text). Still reproducing: R-498 (wiki.DOMAIN), R-500 (UTC). Register 350 → 359 rows.

No P1. Nothing stopped a volunteer; every stop above is either an operator press (R-719) or a backup promise not yet kept (R-720, R-726, R-727).

5. Harness substitutions — not interventions, and what each hides

what consequence
H3 auto-reboot unticked, ISO detached by qm set the "stick still in" reboot not exercised
H4 the Terminal-UI entry chosen over the graphical default the graphical installer not exercised
H5 a nested VM with three virtual disks disk choice and drive health not realistic („0 °C", „Nincs adat")
H6 no browser: every screen driven through the endpoint the page calls script-rendered state not observed

Gone since 2026-09-14: H1 (no mailbox — the tester1@ mailbox was read through the Gmail connector, both customer mails followed) and H2 (US keyboard — keystrokes sent as Hungarian-layout scancodes, @ as AltGr+V).

Harness slips, recorded: two joined form tokens on the first claim („Érvénytelen űrlap", no code judged); a POST without Origin and a -X POST carried across a redirect on the phone login (two 403s, both the harness); a hub-page search for the MAC that could never match; pkill -f ended my own shell once (redone).

6. Teardown — four layers

See evidence-drill-new-household-2026-09-30/teardown/.

  1. Machine: VM 340 destroyed --purge 06:48 UTC; qm list → only 9202 (and 9201 as a CT) remained.
  2. Host (demo-hp): the drill ISO removed (the three older ISOs were there before and stay); nvme-scratch back to 77.0 GB used, local to 24.07 GB; guests 9201 and 9202 running throughout; local-lvm never used.
  3. Hub: host tester-1-693e79 deleted with the escrow acknowledgement (key → retained custody). Customer tester-1 KEPT, not reset (operator ruling); its events are append-only and stay.
  4. Outside the hub: Cloudflare — nothing created, nothing removed (the tunnel is tester-1's own). ep0 (read-only): see §6b.

Untouched, and said: both demo boxes' guests, drill-r50, DooPlex beyond the bake / vouch / hub pages.

6a. Teardown, measured

layer UTC result
1. machine 06:48:25 qm destroy 340 --purge → qm list shows only 9202; /mnt/hdd_1/images/340 gone
2. host 06:48:3x drill ISO removed; nvme-scratch used 89.3 → 77.0 GB (76.98 GB before the VM existed); local 25.7 → 24.07 GB (24.06 before); 9201 and 9202 running before and after
3. hub 07:23:03 the host was „ONLINE" to the hub until 45 min after its last report (the configured stale threshold, R-549) — 26 refusals 409 Host is ONLINE, none changing anything; then host deleted: tester-1-693e79 with the escrow acknowledgement (key demoted to retained custody — the log's „escrow deleted: true" wording is R-544). Host page 404; hosts list back to the three it had before. Customer tester-1 KEPT (200). The delete sent the connect mail by itself (self-bind link auto-minted … on host delete) — R-509's trigger proven again; a valid 7-day link now sits in the tester mailbox with no box behind it
4. outside the hub 07:23:51 ep0, read-only: wgsync: pushed 4 peers → 10.77.0.5/32 gone, 48 s after the delete; peers 5 → 4; the four others unchanged. No manual step was needed (R-600 did not reproduce on a host delete). Cloudflare: nothing created, nothing removed

6b. Left on ep0 — listed, not touched, operator asked

/mnt/pbs-datastore/ns/tester-1/ct/9201/: 2026-09-16T17:27:32Z, 2026-09-16T21:59:54Z (earlier drill boxes, their keys destroyed on acknowledged deletes) and 2026-09-29T19:37:07Z (this drill's box, key now in retained custody). The older two are what the restore test tripped on (R-727). Per the brief, these are listed and left; STATUS asks the operator. The customer's Storage Box sub-account (311327) holds the orphaned restic repository of an earlier box; this drill's box wrote nothing to it (every run skipped) — also left.

6c. Local secrets

The session's 0600 files (passphrase, pairing and setup codes, dashboard and app passwords, recovery code, the vaulted break-glass credential, the bind token) are shredded at the end of this session. The break-glass credential belonged to the deleted host.