P1 uninstall, P2 preflight, P3 install. The install FAILED at step 5/8 in 44s, and the two rank-1 findings are both on the setting-up path a tester's visit is made of. Eleven register rows minted (R-269..R-279); ceiling moves 268 -> 279. R-273 (RANK 1) — the hub vouches agent 0.128.0; that version was published as a Gitea PACKAGE but never git-tagged. Since R-183 the installer correctly pins its config fetches to raw/tag/v<vouched>, so every fresh install and every reinstall now 404s as root, mid-install. Measured: main 200, v0.127.0 200, v0.128.0 404. This is R-184 arriving; release-agent.sh:23 already documents the exact hazard. Existing boxes are fine (self-update takes the binary from the registry). NOT fixed here — publishing a release tag is outward-facing and the runbook says stop and report. One command unblocks it; it is in STATUS.md. R-272 (RANK 1) — Felhom's own uninstall leaves the condition that makes Felhom's own reinstall refuse. It installs dnsmasq at day-0, then on teardown removes the snippet and RESTARTS the daemon unconstrained (process start time lands inside the uninstall window), which grabs 0.0.0.0:53; the next preflight then refuses, and the message reads as though the owner's LAN DNS is at fault. R-274 — a local golden is adopted with no version and no sha check; the manifest vouch is consulted only on the fetch path. demo-hp's local copy is controller 0.192.0 against a vouched 0.210.0, and below the 0.200.0 where the recovery screen shipped. Not yet observed end-to-end (R-273 killed step 5 first). Also: R-275 orphaned credential backups + uid reuse, R-276 the wg tunnel outlives the uninstall, R-269/270/271 from the token rotation, R-277 three hub surfaces misreport a healthy off-site tier, R-278 demo-felhom six days unprotected, R-279 no operator-triggerable off-site run. Two hypotheses of mine were tested and REFUTED rather than shipped as findings: the leftover sudoers file is inert (sudo skips dotted filenames), and demo-hp's off-site tier was healthy all along - I had misread the hub and said so. STATUS.md records the three rulings §8.3 asked for, with the floor CORRECTED to its live value 0.200.0 and the count corrected to twenty.
8.1 KiB
REPORT — BYO reinstall rehearsal, 2026-08-09 (session report)
Written as REPORT-<topic>.md rather than REPORT.md per the repo's parallel-session rule.
The answer to the runbook's question, first
Not answerable today, and the reason is the answer. "If a person's machine is set up again from scratch, does their data come back through the product, in one sitting, with no command line inside the machine?" — the walk never reached the data. A machine wiped today cannot be set up again at all: the install fails at step 5 of 8, on every host, because the agent version the hub vouches was never given a git tag (R-273). The step before that had already refused, because removing Felhom leaves behind the very thing that blocks reinstalling it (R-272).
The data was never at risk and is not now: the planted dataset sits in off-site snapshot 41c830db
with its fingerprints recorded, plus a PBS snapshot and four local vzdump archives.
Venue
demo-hp (t740), operator-approved at STOP 1. It won every fidelity criterion that distinguishes
the two demo boxes: three customer apps against one, a registered storage path against none, and a
Secure-Boot shim install — the customer shape — against demo-felhom's SB-off mkimage firmware
workaround. drill-r50 (VM 300) was verified not at risk before proceeding: it is outside the
felhom pool, on local-lvm, and --uninstall removes no storage and no non-pool guest.
Findings, ranked by what they cost the person in front of you
1 — stops the visit
- R-273 · the vouched agent (0.128.0) has no git tag; every install and reinstall dies at 5/8.
- R-272 · Felhom's uninstall restarts its own dnsmasq unconstrained; it grabs
:53; the next install refuses and appears to blame the owner's network.
2 — costs the visit
- R-274 · a local golden is adopted with no version and no checksum check. The copy on demo-hp is controller 0.192.0 against a vouched 0.210.0 — and below 0.200.0, where the recovery screen the customer needs actually shipped.
- R-276 · an uninstalled box keeps a live WireGuard tunnel into the off-site endpoint; declared in neither the KEPT nor the WIPED list.
3 — misleads
- R-269 · a rotated-out local-API token still authorises until an unrelated lookup forces a reload. The shipped test passes only because of its lookup order.
- R-270 · R-268's own rotation recipe is a step short; the controller never re-reads the mount.
- R-271 · the
agent_channel_unauthorizedalarm can never close — its own advice silences the all-clear. - R-277 · three hub surfaces present a healthy off-site tier as absent. This one caught me.
- R-278 · demo-felhom has had no off-site backup for six days, waiting on a ceremony nobody ran.
4 — cosmetic / hygiene
- R-275 · five orphaned 0600 credential backups survive, and uid reuse hands them to the new service account. Superseded keys here; live if the backups were recent.
- R-279 · no operator-triggerable off-site backup exists.
What I got wrong, and corrected
I reported to the operator that the off-site tier had not run on either box since 2026-08-03. That was true of demo-felhom and false of demo-hp, which had 18 unbroken daily snapshots. I had read three hub surfaces that agreed with each other and none of which said what I took them to say (now R-277). I corrected it before it changed any decision, and the operator's "repair off-site first" ruling turned out to be unnecessary for the chosen venue.
I also raised R-275's sudoers half as a likely privilege-escalation on reinstall, then tested it
and refuted my own hypothesis: sudo skips filenames containing dots, so the leftover file is inert.
visudo -c -f parsing a file OK is not evidence that sudo loads it.
Integrity verdict
Not reached — P6 never ran. The instrument is proven and waiting: the comparator was watched
failing three ways, including an NFC→NFD rename that renders identically to the eye and would
have passed a rendered-text comparison. Four files, two with Hungarian accented names
(árvíztűrő-tükörfúrógép.txt, nested/őszibarack.md), one 3 MB binary; manifest keyed on raw name
bytes, committed as evidence-rehearsal-2026-08-09/GATE0-before-manifest.json; all four confirmed
present in snapshot 41c830db.
Wall clocks
| phase | duration |
|---|---|
| Pre-phase (R-268 rotation, proved both ways) | ~25 min |
| Gate 0 (venue, dataset, positive control, off-site run, capture) | ~55 min |
| P1 uninstall | 60 s (08:37:23 → 08:38:23 UTC) |
| P1 leave-behind measurement | ~12 min |
| P2 preflight (3 runs: 2 refusals, 1 pass) | ~6 min |
| P3 install — FAILED | 44 s (08:51:33 → 08:52:17 UTC) |
| bare machine → visible files | not reached |
Steps taken off-path, and what they cost
- R-268 rotation on demo-felhom — required by the runbook's pre-phase; not the venue.
- demo-hp's dashboard password re-set to the credentials-file value, on operator instruction. The customer-owned password was unknown to this session and no operator route to the off-site button exists (R-279). Prior hash preserved in-guest; destroyed with the guest at P1. Cost: none — P1 wiped it and P4 re-claims.
- The off-site run was started by a script pressing the dashboard's own endpoint with a real session and CSRF token, not by a person clicking. Identical server path; only the click synthetic.
--passphrase-fileinstead of the no-echo prompt — a first-class documented option with a permission check, so the secret still never touched argv. A person would type it.systemctl stop dnsmasq && systemctl disable dnsmasq— the action the refusal message tells the owner to take, used as the counterfactual that confirmed R-272.
Everything after the install command returned was read-only. No repair was attempted on the half-installed box.
Teardown — all four layers
- The machine — nothing created beyond the half-install itself, which is left in place
deliberately for inspection and resumption (
state.jsoncompleted: preflight, token, grows, enroll). demo-hp is not serving right now: no guest, agent installed but no unit. - The host —
local-lvm20 904 790 → 12 355 143 KiB (≈8.5 GiB returned);local≈64 MiB; NVMe unchanged, backups deliberately kept. Pre-existing leftovers found and not removed (not this run's, recorded instead): storagec11-scratch, the orphanedvzdump-lxc-9100archive, and/root/.dpw,.h,.sec.htmlin the old guest (now destroyed with it). - The hub — no customer or appliance record was created; the
demo-hpcustomer is retained deliberately, as the runbook requires. Nothing to delete. - The off-site side — one write, and it was the intended one: the Gate 0 backup that created
snapshots
41c830db,9e38b84c,78b93f04. No prune, no forget, no delete. How I know: every restic call wassnapshots,ls, or the product's ownPOST /backup/offbox/run; retention runs inside that product path and is ep0's server-side job (R-89, boxes keepkeep_last: 0).
Secrets handling
No secret reached stdout. Token values, the retrieval passphrase, the controller password and the hub DB copy were handled file→file at 0600 and shredded; the hub DB copy (which carries every host's break-glass credential) was shredded immediately after the one hash comparison it was taken for. Credential comparisons were done by sha256 prefix, never by value.
State for whoever resumes
- Blocked on one command —
git tag v0.128.0 28ba8593b8 && git push origin v0.128.0. - Nothing in Gate 0 needs redoing. Snapshot
41c830dbholds the dataset; the before-manifest is committed; the comparator is proven. - First thing to record at step 7/8 when the install resumes: which golden it takes (R-274 is the one finding still owed an end-to-end observation).
- P4 will need you — the claim code is emailed to
doodoo21@freemail.huand there is no operator-side route to it (R-119).
Full account: documentation/audits/REHEARSAL-byo-reinstall-2026-08-09.md.