Files
felhom.eu/REPORT-rehearsal-2026-08-09.md
T
admin 1d6f1c522d
gates / gates (push) Successful in 23s
Rehearsal 2026-08-09 COMPLETE: data BYTE-IDENTICAL, journey needs a shell twice
The walk finished. All four planted files came back byte-identical out of
snapshot 41c830db, including two Hungarian accented filenames verified as RAW
NAME BYTES (NFC preserved) — the discriminator the Gate 0 positive control was
built for, having been watched failing on an NFC->NFD rename that renders the
same. Unlock 21s, restore 13.2s.

It finished only because a terminal was available twice:

- R-273 CLOSED. v0.128.0 was published as a package and never git-tagged, so
  every install died at 5/8. Tag pushed on operator instruction after an
  INDEPENDENT download proved the package sha equalled the vouched value;
  --resume then reached Day-0 SUCCESS in 3m49s on controller 0.210.0. The two
  guards that would stop the class recurring are still owed.
- R-280 NEW, rank 1. A reinstalled box cannot re-attach its own data drive by
  any dashboard route: /api/disks/candidates returns empty because both lists
  are built from the UNCLAIMED-disk scan, and the drive is claimed precisely
  because it is also the backup target. Correct for "initialise", over-broad for
  "attach", which is non-destructive by definition. The restore page meanwhile
  says "Ez ket kattintas" and points at that empty list. Cleared by POSTing
  /mnt/sys_drive — an internal path no household could produce.

Also new: R-281 the hub said NOTHING through the entire reinstall and the
sealed-backup tripwire did not fire on a real unseal (positive control: 2 events
all day fleet-wide); R-282 one code with three names and a mail pointing at a
page the box does not show; R-283 hub reads "Claimed 18d ago" while the box
serves its setup page; R-284 "almost full" over a 93%-free store.

R-274 NARROWED by measurement rather than left as written: the resume path
fetched the vouched golden correctly, because --resume skips the preflight that
does local discovery. What survives is real — discovery is sort|tail -1 with no
manifest comparison — but a FRESH install taking a stale golden is still not
observed, and the row says so.

Two of my own claims were refuted by test and are recorded as refuted, not
quietly dropped: the leftover sudoers file is inert (sudo skips dotted names),
and demo-hp's off-site tier was healthy all along.
2026-08-09 12:35:26 +02:00

184 lines
11 KiB
Markdown

# REPORT — BYO reinstall rehearsal, 2026-08-09 (session report)
*Written as `REPORT-<topic>.md` rather than `REPORT.md` per the repo's parallel-session rule.*
## The answer to the runbook's question, first
**Yes, the data comes back byte for byte. No, not in one sitting, and not without a shell.**
All four planted files returned **BYTE-IDENTICAL** — including two Hungarian accented filenames
verified as *raw name bytes*, not as rendered text. The unlock took **21 s**, the restore **13.2 s**.
But the walk completed only because two hard stops were cleared by someone who could open a terminal
and read source. **R-273**: the install died at step 5/8 on an agent version that was published as a
package but never git-tagged — cleared by completing the release. **R-280**: the reinstalled machine
could not re-attach its own data drive through any dashboard route, while the restore page said
„**Ez két kattintás**" and pointed at an empty list — cleared by POSTing an internal path
(`/mnt/sys_drive`) that no household could produce.
Neither is a data-integrity problem. Both stop a household dead. **This is the same shape the R-201
walks kept finding: the data half passes, the journey half fails.**
## Venue
**demo-hp (t740)**, operator-approved at STOP 1. It won every fidelity criterion that distinguishes
the two demo boxes: three customer apps against one, a registered storage path against none, and a
Secure-Boot `shim` install — the customer shape — against demo-felhom's SB-off `mkimage` firmware
workaround. `drill-r50` (VM 300) was verified not at risk before proceeding: it is outside the
`felhom` pool, on `local-lvm`, and `--uninstall` removes no storage and no non-pool guest.
## Findings, ranked by what they cost the person in front of you
**1 — stops the visit**
- **R-273** · the vouched agent (0.128.0) had no git tag; every install died at 5/8. **CLOSED** — tag
pushed on your instruction after an independent sha check; install then succeeded in 3 m 49 s. The
two guards that would prevent a recurrence are still owed.
- **R-280** · a reinstalled machine cannot re-attach its data drive through any route, and the restore
page promises „két kattintás" at an empty list. **The one to fix before the tester's visit.**
- **R-272** · Felhom's uninstall restarts its own dnsmasq unconstrained; it grabs `:53`; the next
install refuses and appears to blame the owner's network.
**2 — costs the visit**
- **R-274** · a local golden is adopted with no version and no checksum check. The copy on demo-hp is
controller **0.192.0** against a vouched **0.210.0** — and below 0.200.0, where the recovery screen
the customer needs actually shipped.
- **R-276** · an uninstalled box keeps a live WireGuard tunnel into the off-site endpoint; declared
in neither the KEPT nor the WIPED list.
**3 — misleads**
- **R-281** · the hub said nothing at all through the whole reinstall, and the tripwire for a
sealed-backup unseal did not fire on a real one.
- **R-282 / R-283** · one code, three names; the mail points at a page the box is not showing; the hub
reads "Claimed 18d ago" while the box serves its setup page.
- **R-269** · a rotated-out local-API token still authorises until an unrelated lookup forces a
reload. The shipped test passes only because of its lookup order.
- **R-270** · R-268's own rotation recipe is a step short; the controller never re-reads the mount.
- **R-271** · the `agent_channel_unauthorized` alarm can never close — its own advice silences the
all-clear.
- **R-277** · three hub surfaces present a healthy off-site tier as absent. **This one caught me.**
- **R-278** · demo-felhom has had no off-site backup for six days, waiting on a ceremony nobody ran.
**4 — cosmetic / hygiene**
- **R-275** · five orphaned 0600 credential backups survive, and uid reuse hands them to the new
service account. Superseded keys here; live if the backups were recent.
- **R-279** · no operator-triggerable off-site backup exists.
## What I got wrong, and corrected
I reported to the operator that the off-site tier had not run on **either** box since 2026-08-03.
That was true of demo-felhom and **false of demo-hp**, which had 18 unbroken daily snapshots. I had
read three hub surfaces that agreed with each other and none of which said what I took them to say
(now **R-277**). I corrected it before it changed any decision, and the operator's "repair off-site
first" ruling turned out to be unnecessary for the chosen venue.
I also raised **R-275**'s sudoers half as a likely privilege-escalation on reinstall, then **tested it
and refuted my own hypothesis**: sudo skips filenames containing dots, so the leftover file is inert.
`visudo -c -f` parsing a file OK is not evidence that sudo loads it.
## Integrity verdict — BYTE-IDENTICAL
```
expected 4 file(s); found 4
VERDICT: BYTE-IDENTICAL
```
Four expected, four restored, zero differences, compared against
`evidence-rehearsal-2026-08-09/GATE0-before-manifest.json` — a manifest keyed on **raw name bytes**.
`árvíztűrő-tükörfúrógép.txt` and `nested/őszibarack.md` came back with their name bytes intact (NFC
preserved, `c3a1…`), which is the discriminator the Gate 0 positive control was built to enforce: the
comparator had been watched **failing** on an NFC→NFD rename that renders identically to the eye.
Restored out of snapshot `41c830db` into the verification folder the product names, with live data
untouched.
## Wall clocks
| phase | duration |
|---|---|
| Pre-phase (R-268 rotation, proved both ways) | ~25 min |
| Gate 0 (venue, dataset, positive control, off-site run, capture) | ~55 min |
| P1 uninstall | **60 s** (08:37:23 → 08:38:23 UTC) |
| P1 leave-behind measurement | ~12 min |
| P2 preflight (3 runs: 2 refusals, 1 pass) | ~6 min |
| P3 install — first attempt, FAILED | 44 s (08:51:33 → 08:52:17 UTC) |
| P3 install — resumed, SUCCESS | **3 m 49 s** |
| P4 first contact (box live on its own URL) | within ~4 min of install |
| STOP 3 unlock | **21 s** |
| app redeploy (calibre-web) | **1 m 36 s** |
| restore prepare + execute | **8 s + 13.2 s** |
| **bare machine → verified files** | **1 h 49 m 22 s** (08:38:23 → 10:27:45 UTC) |
| — of which the product's own work | **≈ 7 m 47 s** |
**Neither figure is the customer number.** The 1 h 49 m is dominated by the R-273 diagnosis and release
fix (~38 min) and two waits on a human. The 7 m 47 s is what the product costs when the operator
already knows every answer. **The honest unaided figure is undefined, because an unaided household does
not finish.**
## Steps taken off-path, and what they cost
1. **R-268 rotation on demo-felhom** — required by the runbook's pre-phase; not the venue.
2. **demo-hp's dashboard password re-set to the credentials-file value**, on operator instruction.
The customer-owned password was unknown to this session and no operator route to the off-site
button exists (R-279). Prior hash preserved in-guest; destroyed with the guest at P1. Cost: none —
P1 wiped it and P4 re-claims.
3. **The off-site run was started by a script pressing the dashboard's own endpoint** with a real
session and CSRF token, not by a person clicking. Identical server path; only the click synthetic.
4. **`--passphrase-file` instead of the no-echo prompt** — a first-class documented option with a
permission check, so the secret still never touched argv. A person would type it.
5. **`systemctl stop dnsmasq && systemctl disable dnsmasq`** — the action the refusal message tells
the owner to take, used as the counterfactual that confirmed R-272.
6. **Pushed the `v0.128.0` git tag** — outward-facing, done on your "proceed", and only after an
independent download proved the published package's sha256 equalled the hub's vouched value. It
completes a half-finished release rather than changing code; the release script's own recovery text
is the same line. **Cost to the walk: the install that followed was a `--resume`, not a fresh run,
which is why R-274 is only half-observed.**
7. **`POST /settings/storage/add` with `/mnt/sys_drive`** — the manual escape hatch, typed. This is the
R-280 wall; a customer could not produce that path. **The biggest fidelity cost of the run.**
8. **SSH into the guest to fingerprint the restored tree.** This is my *instrument*, not a customer
step — the customer's step (the restore) finished at the dashboard. Byte-comparison inherently needs
file access; nothing about the product was driven this way.
Everything else after the install returned was read-only, and no repair was attempted on the box.
## Teardown — all four layers
1. **The machine** — nothing created beyond the half-install itself, which is **left in place
deliberately** for inspection and resumption (`state.json` completed: preflight, token, grows,
enroll). demo-hp is **not serving** right now: no guest, agent installed but no unit.
2. **The host**`local-lvm` 20 904 790 → 12 355 143 KiB (**≈8.5 GiB returned**); `local` ≈64 MiB;
NVMe unchanged, backups deliberately kept. Pre-existing leftovers found and **not** removed
(not this run's, recorded instead): storage `c11-scratch`, the orphaned `vzdump-lxc-9100` archive,
and `/root/.dpw`, `.h`, `.sec.html` in the old guest (now destroyed with it).
3. **The hub****no customer or appliance record was created**; the `demo-hp` customer is retained
deliberately, as the runbook requires. Nothing to delete.
4. **The off-site side****one write, and it was the intended one**: the Gate 0 backup that created
snapshots `41c830db`, `9e38b84c`, `78b93f04`. **No prune, no forget, no delete.** How I know: every
restic call was `snapshots`, `ls`, or the product's own `POST /backup/offbox/run`; retention runs
inside that product path and is ep0's server-side job (R-89, boxes keep `keep_last: 0`).
## Secrets handling
No secret reached stdout. Token values, the retrieval passphrase, the controller password and the hub
DB copy were handled file→file at 0600 and shredded; the hub DB copy (which carries every host's
break-glass credential) was shredded immediately after the one hash comparison it was taken for.
Credential comparisons were done by sha256 prefix, never by value.
## State demo-hp was left in
**Back in service and healthy** — agent 0.128.0, controller 0.210.0, guest 9201 running and onboot,
claimed, storage path registered, `calibre-web` deployed, off-site repository unlocked and intact at
18 snapshots. `drill-r50` (VM 300) untouched throughout.
**Deliberately left alone, and named rather than tidied:** the pre-existing `c11-scratch` storage and
the three `vzdump-lxc-9100` golden archives on `local` (the teardown keeps goldens by design, and they
now number three). The restored files sit in the product's verification folder, not back in place —
that is R-213 and the product says so.
## Still owed
- **R-280** — the drive wall. The one finding that would stop the tester's visit outright.
- **R-273's two guards** — refuse a vouch whose tag does not resolve; check that a package and its tag
ship together. The tag push fixed one box, not the class.
- **R-274's missing observation** — a *fresh* (non-resume) install taking a stale local golden.
Full account: `documentation/audits/REHEARSAL-byo-reinstall-2026-08-09.md`.