The walk finished. All four planted files came back byte-identical out of snapshot 41c830db, including two Hungarian accented filenames verified as RAW NAME BYTES (NFC preserved) — the discriminator the Gate 0 positive control was built for, having been watched failing on an NFC->NFD rename that renders the same. Unlock 21s, restore 13.2s. It finished only because a terminal was available twice: - R-273 CLOSED. v0.128.0 was published as a package and never git-tagged, so every install died at 5/8. Tag pushed on operator instruction after an INDEPENDENT download proved the package sha equalled the vouched value; --resume then reached Day-0 SUCCESS in 3m49s on controller 0.210.0. The two guards that would stop the class recurring are still owed. - R-280 NEW, rank 1. A reinstalled box cannot re-attach its own data drive by any dashboard route: /api/disks/candidates returns empty because both lists are built from the UNCLAIMED-disk scan, and the drive is claimed precisely because it is also the backup target. Correct for "initialise", over-broad for "attach", which is non-destructive by definition. The restore page meanwhile says "Ez ket kattintas" and points at that empty list. Cleared by POSTing /mnt/sys_drive — an internal path no household could produce. Also new: R-281 the hub said NOTHING through the entire reinstall and the sealed-backup tripwire did not fire on a real unseal (positive control: 2 events all day fleet-wide); R-282 one code with three names and a mail pointing at a page the box does not show; R-283 hub reads "Claimed 18d ago" while the box serves its setup page; R-284 "almost full" over a 93%-free store. R-274 NARROWED by measurement rather than left as written: the resume path fetched the vouched golden correctly, because --resume skips the preflight that does local discovery. What survives is real — discovery is sort|tail -1 with no manifest comparison — but a FRESH install taking a stale golden is still not observed, and the row says so. Two of my own claims were refuted by test and are recorded as refuted, not quietly dropped: the leftover sudoers file is inert (sudo skips dotted names), and demo-hp's off-site tier was healthy all along.
11 KiB
REPORT — BYO reinstall rehearsal, 2026-08-09 (session report)
Written as REPORT-<topic>.md rather than REPORT.md per the repo's parallel-session rule.
The answer to the runbook's question, first
Yes, the data comes back byte for byte. No, not in one sitting, and not without a shell.
All four planted files returned BYTE-IDENTICAL — including two Hungarian accented filenames verified as raw name bytes, not as rendered text. The unlock took 21 s, the restore 13.2 s.
But the walk completed only because two hard stops were cleared by someone who could open a terminal
and read source. R-273: the install died at step 5/8 on an agent version that was published as a
package but never git-tagged — cleared by completing the release. R-280: the reinstalled machine
could not re-attach its own data drive through any dashboard route, while the restore page said
„Ez két kattintás" and pointed at an empty list — cleared by POSTing an internal path
(/mnt/sys_drive) that no household could produce.
Neither is a data-integrity problem. Both stop a household dead. This is the same shape the R-201 walks kept finding: the data half passes, the journey half fails.
Venue
demo-hp (t740), operator-approved at STOP 1. It won every fidelity criterion that distinguishes
the two demo boxes: three customer apps against one, a registered storage path against none, and a
Secure-Boot shim install — the customer shape — against demo-felhom's SB-off mkimage firmware
workaround. drill-r50 (VM 300) was verified not at risk before proceeding: it is outside the
felhom pool, on local-lvm, and --uninstall removes no storage and no non-pool guest.
Findings, ranked by what they cost the person in front of you
1 — stops the visit
- R-273 · the vouched agent (0.128.0) had no git tag; every install died at 5/8. CLOSED — tag pushed on your instruction after an independent sha check; install then succeeded in 3 m 49 s. The two guards that would prevent a recurrence are still owed.
- R-280 · a reinstalled machine cannot re-attach its data drive through any route, and the restore page promises „két kattintás" at an empty list. The one to fix before the tester's visit.
- R-272 · Felhom's uninstall restarts its own dnsmasq unconstrained; it grabs
:53; the next install refuses and appears to blame the owner's network.
2 — costs the visit
- R-274 · a local golden is adopted with no version and no checksum check. The copy on demo-hp is controller 0.192.0 against a vouched 0.210.0 — and below 0.200.0, where the recovery screen the customer needs actually shipped.
- R-276 · an uninstalled box keeps a live WireGuard tunnel into the off-site endpoint; declared in neither the KEPT nor the WIPED list.
3 — misleads
- R-281 · the hub said nothing at all through the whole reinstall, and the tripwire for a sealed-backup unseal did not fire on a real one.
- R-282 / R-283 · one code, three names; the mail points at a page the box is not showing; the hub reads "Claimed 18d ago" while the box serves its setup page.
- R-269 · a rotated-out local-API token still authorises until an unrelated lookup forces a reload. The shipped test passes only because of its lookup order.
- R-270 · R-268's own rotation recipe is a step short; the controller never re-reads the mount.
- R-271 · the
agent_channel_unauthorizedalarm can never close — its own advice silences the all-clear. - R-277 · three hub surfaces present a healthy off-site tier as absent. This one caught me.
- R-278 · demo-felhom has had no off-site backup for six days, waiting on a ceremony nobody ran.
4 — cosmetic / hygiene
- R-275 · five orphaned 0600 credential backups survive, and uid reuse hands them to the new service account. Superseded keys here; live if the backups were recent.
- R-279 · no operator-triggerable off-site backup exists.
What I got wrong, and corrected
I reported to the operator that the off-site tier had not run on either box since 2026-08-03. That was true of demo-felhom and false of demo-hp, which had 18 unbroken daily snapshots. I had read three hub surfaces that agreed with each other and none of which said what I took them to say (now R-277). I corrected it before it changed any decision, and the operator's "repair off-site first" ruling turned out to be unnecessary for the chosen venue.
I also raised R-275's sudoers half as a likely privilege-escalation on reinstall, then tested it
and refuted my own hypothesis: sudo skips filenames containing dots, so the leftover file is inert.
visudo -c -f parsing a file OK is not evidence that sudo loads it.
Integrity verdict — BYTE-IDENTICAL
expected 4 file(s); found 4
VERDICT: BYTE-IDENTICAL
Four expected, four restored, zero differences, compared against
evidence-rehearsal-2026-08-09/GATE0-before-manifest.json — a manifest keyed on raw name bytes.
árvíztűrő-tükörfúrógép.txt and nested/őszibarack.md came back with their name bytes intact (NFC
preserved, c3a1…), which is the discriminator the Gate 0 positive control was built to enforce: the
comparator had been watched failing on an NFC→NFD rename that renders identically to the eye.
Restored out of snapshot 41c830db into the verification folder the product names, with live data
untouched.
Wall clocks
| phase | duration |
|---|---|
| Pre-phase (R-268 rotation, proved both ways) | ~25 min |
| Gate 0 (venue, dataset, positive control, off-site run, capture) | ~55 min |
| P1 uninstall | 60 s (08:37:23 → 08:38:23 UTC) |
| P1 leave-behind measurement | ~12 min |
| P2 preflight (3 runs: 2 refusals, 1 pass) | ~6 min |
| P3 install — first attempt, FAILED | 44 s (08:51:33 → 08:52:17 UTC) |
| P3 install — resumed, SUCCESS | 3 m 49 s |
| P4 first contact (box live on its own URL) | within ~4 min of install |
| STOP 3 unlock | 21 s |
| app redeploy (calibre-web) | 1 m 36 s |
| restore prepare + execute | 8 s + 13.2 s |
| bare machine → verified files | 1 h 49 m 22 s (08:38:23 → 10:27:45 UTC) |
| — of which the product's own work | ≈ 7 m 47 s |
Neither figure is the customer number. The 1 h 49 m is dominated by the R-273 diagnosis and release fix (~38 min) and two waits on a human. The 7 m 47 s is what the product costs when the operator already knows every answer. The honest unaided figure is undefined, because an unaided household does not finish.
Steps taken off-path, and what they cost
- R-268 rotation on demo-felhom — required by the runbook's pre-phase; not the venue.
- demo-hp's dashboard password re-set to the credentials-file value, on operator instruction. The customer-owned password was unknown to this session and no operator route to the off-site button exists (R-279). Prior hash preserved in-guest; destroyed with the guest at P1. Cost: none — P1 wiped it and P4 re-claims.
- The off-site run was started by a script pressing the dashboard's own endpoint with a real session and CSRF token, not by a person clicking. Identical server path; only the click synthetic.
--passphrase-fileinstead of the no-echo prompt — a first-class documented option with a permission check, so the secret still never touched argv. A person would type it.systemctl stop dnsmasq && systemctl disable dnsmasq— the action the refusal message tells the owner to take, used as the counterfactual that confirmed R-272.- Pushed the
v0.128.0git tag — outward-facing, done on your "proceed", and only after an independent download proved the published package's sha256 equalled the hub's vouched value. It completes a half-finished release rather than changing code; the release script's own recovery text is the same line. Cost to the walk: the install that followed was a--resume, not a fresh run, which is why R-274 is only half-observed. POST /settings/storage/addwith/mnt/sys_drive— the manual escape hatch, typed. This is the R-280 wall; a customer could not produce that path. The biggest fidelity cost of the run.- SSH into the guest to fingerprint the restored tree. This is my instrument, not a customer step — the customer's step (the restore) finished at the dashboard. Byte-comparison inherently needs file access; nothing about the product was driven this way.
Everything else after the install returned was read-only, and no repair was attempted on the box.
Teardown — all four layers
- The machine — nothing created beyond the half-install itself, which is left in place
deliberately for inspection and resumption (
state.jsoncompleted: preflight, token, grows, enroll). demo-hp is not serving right now: no guest, agent installed but no unit. - The host —
local-lvm20 904 790 → 12 355 143 KiB (≈8.5 GiB returned);local≈64 MiB; NVMe unchanged, backups deliberately kept. Pre-existing leftovers found and not removed (not this run's, recorded instead): storagec11-scratch, the orphanedvzdump-lxc-9100archive, and/root/.dpw,.h,.sec.htmlin the old guest (now destroyed with it). - The hub — no customer or appliance record was created; the
demo-hpcustomer is retained deliberately, as the runbook requires. Nothing to delete. - The off-site side — one write, and it was the intended one: the Gate 0 backup that created
snapshots
41c830db,9e38b84c,78b93f04. No prune, no forget, no delete. How I know: every restic call wassnapshots,ls, or the product's ownPOST /backup/offbox/run; retention runs inside that product path and is ep0's server-side job (R-89, boxes keepkeep_last: 0).
Secrets handling
No secret reached stdout. Token values, the retrieval passphrase, the controller password and the hub DB copy were handled file→file at 0600 and shredded; the hub DB copy (which carries every host's break-glass credential) was shredded immediately after the one hash comparison it was taken for. Credential comparisons were done by sha256 prefix, never by value.
State demo-hp was left in
Back in service and healthy — agent 0.128.0, controller 0.210.0, guest 9201 running and onboot,
claimed, storage path registered, calibre-web deployed, off-site repository unlocked and intact at
18 snapshots. drill-r50 (VM 300) untouched throughout.
Deliberately left alone, and named rather than tidied: the pre-existing c11-scratch storage and
the three vzdump-lxc-9100 golden archives on local (the teardown keeps goldens by design, and they
now number three). The restored files sit in the product's verification folder, not back in place —
that is R-213 and the product says so.
Still owed
- R-280 — the drive wall. The one finding that would stop the tester's visit outright.
- R-273's two guards — refuse a vouch whose tag does not resolve; check that a package and its tag ship together. The tag push fixed one box, not the class.
- R-274's missing observation — a fresh (non-resume) install taking a stale local golden.
Full account: documentation/audits/REHEARSAL-byo-reinstall-2026-08-09.md.