Files
felhom.eu/REPORT-drill-fresh-install-0242.md
T
admin 8c7f882d1c
gates / gates (push) Successful in 19s
drill 0242: teardown complete in three layers; R-501 filed
Machine: VM 330 destroyed with its disks. Host: ISO and scratch removed,
~8.5 GiB returned on nvme-scratch. Hub: customer drill0242 DELETED via the
cascade (journal #17), pages 404; its ep0 WireGuard peer dropped at the next
full-list push. Evidence pulled before the destroy. R-501: the documented
CI-check recipe reads only the last jobs page, which is not in id order.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-14 16:40:59 +02:00

9.3 KiB
Raw Blame History

REPORT — DRILL: a stranger's first hour on 0.242.0 (2026-09-14)

Runbook-style validation. No product code written. A parallel session owns root REPORT.md, so this is a topic sibling (CLAUDE.md). Findings doc: documentation/audits/DRILL-fresh-install-0242-2026-09-14.md; every observable: documentation/audits/evidence-drill-fresh-install-0242-2026-09-14/journal.md.

0. Claims in the brief that turned out wrong, or incomplete — named first

  1. "ep0 is not touched." Enrolment itself registers a WireGuard peer on ep0 for every box (wg registered … ip=10.77.0.5/32 … sync=ok), DR tier on or off. I avoided the part I could (DR tier off, which would have created an ep0 namespace and token); the peer is the product's own act and is removed by the host delete (§5).
  2. "This run re-proves or narrows the journey row (~L90)." That row is the rebuild-and-recover journey (walk 5). This drill walked the first hour, with off-site off. It can neither re-prove nor narrow that row. I added a scope note to it and a new first-hour row (PARTIAL).
  3. "Claim the box with the code." There are three secrets, not one: the console Párosító kód, the operator-held Tulajdonosi jelmondat (nothing delivers it — R-497) and the mailed Beállító kód. Two of the three arrive by mail, which this harness cannot read.
  4. The three claims marked "read, not measured" — measured now: website — true, no mention of the installer or iso.felhom.eu, and the ISO host has no index; instructions — true, none exist (R-493); golden landing — the box landed on 0.242.0, the new golden, with no self-update.
  5. "Newest baked golden 0.236.0; waiver to 2026-09-27; highest R-492; baselines" — all correct.

1. Baselines (re-verified 12:58 UTC)

controller 406755fa8fba v0.242.0 · agent 4586f0f7f6d1 v0.130.0 · felhom.eu 41590f8ee618 hub v0.112.0. Hub before: agent 0.130.0, golden 0.236.0, min_agent 0.129.0, floor 0.242.0. Architecture read for the area: 00-capability-map.md (journey row), 09-update-architecture.md §3.

2. The golden

0.242.0 baked, round-trip verified, vouched — sha 3ab480dd…e6d8, 653 288 425 B, all markers pass, token-leak 0 with a working control, three readers agree. Only golden_version moved. documentation/tests/golden-0.242.0-2026-09-14/README.md. The golden-currency gate is now plain OK. One slip: my first template pick was arm64; caught before the bake.

3. The verdict

Interventions: 1. Ready for a volunteer: no — no instructions exist (R-493), and the setup mail's dashboard link does not open for a new customer (R-494). Every mechanism after that passed: install, landing on the vouched set, deploy, use, backup, remove, byte-identical restore, power cut (same versions, no alarm), code typo and lockout.

intervention row what
I1 R-494 (filed before acting) dashboard reached by LAN address with the name forced — the mailed name has no DNS

Harness substitutions (a volunteer would not need them; each hides a part of the path): H1 no mailbox → operator bind instead of the self-bind page, and two box-printed setup codes via the vaulted break-glass; H2 US keyboard layout; H3 auto-reboot unticked, ISO detached; H4 Terminal UI entry. Full list with harness slips: findings doc §3.

4. Findings — every one a row

row rank
R-493 P1 no customer install instructions; ISO host has no index
R-494 P1 new customer's dashboard has no address (I1)
R-495 P2 installer's unanswered questions; refuses its own default hostname
R-496 P2 console sends a stranger to the Proxmox admin page; „a jelszavadat"
R-497 P2 the Tulajdonosi jelmondat is delivered by nothing
R-499 P2 „already in the PBS backup, nothing to do" on a box with no PBS
R-498 P3 52 of 53 app pages say literal wiki.DOMAIN
R-500 P3 dashboard backup time in UTC, backup pages in local time
R-501 P3 the documented CI-check recipe reads only the last jobs page and can miss a run (hygiene, found at push)

R-469 was not touched. R-214 reproduced (recorded, row unchanged). Register: 200 → 209 table rows. Opened 9 (R-493…R-501), closed 0.

5. Teardown — three layers

layer before after
1. machine qm list: VM 330 running, 8.6 G in /mnt/hdd_1/images/330 qm destroy 330 --purge → qm list empty; images/330 gone; images/9202 untouched
2. host nvme-scratch used 19 059 372 KiB · local used 24 768 200 KiB · ISO present · /root/drill0242 40 files nvme-scratch 10 130 532 KiB (≈8.5 GiB returned) · local 23 013 832 KiB (the 1.7 GiB ISO) · ISO 0 · scratch dir shredded and removed. nvme-scratch storage itself stays — it hosts 9202. vmbr9 pre-existed
3. hub customer drill0242, host drill0242-3f4b42 ONLINE, deletable:false, wg_peer_bound:true, recovery_present:true customer and host DELETED — customer page 404, host page 404, both lists 0 (control customer listed 1); ep0 peer 10.77.0.5 gone at the next peer push (control peer present) — §5b

5b. The hub record

Disposition: DELETED. Not retained as a fixture, not blocked.

UTC observable
14:34:45 hub: Host staleness: drill0242-3f4b42 ok → stale (host_stale) · Operator email sent for drill0242/host_stale — a true alarm, caused by the VM destroy
14:34:53 /hosts/drill0242-3f4b42/delete-impact → "deletable":true,"status":"stale","wg_peer_bound":true,"recovery_present":true
14:35:15 POST /configs/drill0242/delete ack_hosts=1 ack_reset=1 ack_purge=1 confirm_id=drill0242 expect_hosts=1 → 303 /configs?flash=deleted
14:35:16–17 customer DELETE cascade started … (journal #17, 1 host(s)) · host drill0242-3f4b42 deleted (escrow DEMOTED to retained custody) · tenantsync: deprovision ok for drill0242 (ns=drill0242, existed=false) · reset drill0242: PBS tenancy deprovisioned · [claim] reset to unclaimed · residue purged (reports=9 app_telemetry=16 … notif_prefs=1 selfbind_tokens=1 appliance_registrations=1) · customer DELETE cascade COMPLETE for drill0242 (journal #17) — full teardown
14:35:2x /customers/drill0242 404 · /hosts/drill0242-3f4b42 404 · drill0242 on /configs 0 (control enkisfelhom 1) · on /hosts 0
14:35:15 / 14:37:21 ep0 wg show all allowed-ips: 10.77.0.5/32 present 1, 1 (read-only; control 10.77.0.3/32 present)
14:39:30 hub: wgsync: pushed 4 peers to 167.233.158.164:22 (was 5)
14:40:00 ep0: 10.77.0.5/32 0, config files naming it 0; control 10.77.0.3/32 1

Observation, not filed: the delete does not trigger an immediate peer push, so the ep0 peer outlived the customer by 4 m 14 s, until the periodic full-list push (wgsync/reconciler.go:19-23, declarative by design). The retained escrow custody the host delete mentions is empty here — escrow_present:false; no ceremony ran — and the customer purge is the step that removes it anyway. tenantsync … existed=false confirms no ep0 PBS namespace was ever created (DR tier off).

Append-only and staying, by design: the hub's event stream (controller_started, app_removed, claim_lockout, the bind and enrol lines) and the operator e-mail for the lockout.

Untouched, and checked: demo-hp guests 9201 and 9202 (running before and after); local-lvm (44.17 % before and after); demo-felhom; DooPlex services (the bake VM, the accepted exception, back on virgin); Peti's box; ep0 beyond the product's own peer. drill-r50 does not exist (R-461).

6. Secrets

Hub password, BookStack and dashboard passwords, the passphrase, both setup codes, the break-glass credential and the installer root password lived only in 0600 files in the session scratchpad. Every committed file was swept for each, with a planted control that was found: 0 hits. The pairing code is redacted in two screenshots and the text. The break-glass reveal emitted its audit event, by design.

7. unproven.py --summary

Unchanged: 55 claims, NOT WALKED 35 of 55. It reads its own claim list, not the new map row.

7b. CI for my push

38848ff (felhom.eu main): job id 581 gates — completed, conclusion success, 14:18:23Z (run id 582, run_number 335 on actions/tasks). Found by scanning every page of actions/jobs and matching head_sha: the documented last-page recipe did not list it for 8 minutes because the list is not in id order — R-501. Register now 200 → 209 table rows (opened 9, closed 0).

8. Observations

  • The deploy page's poll has no degraded branch; 16 s of stale step text on BookStack.
  • The drive-attach list offers the guest's own system volume as an existing drive.
  • The removal dialog names „éjszakai restic pillanatképek" on a box with no restic.
  • „0 °C" beside „Nincs adat" for a virtual disk; an English „Debug" menu item.
  • The claim lockout is global as well as per source; its mail reaches the operator only.
  • Two of my first readings were of script-rendered elements from server HTML (host-metrics banner, restore-finish message); both corrected in the journal. No browser here — strict UI coverage is the operator's click-through.