diff --git a/REPORT-rehearsal-2026-08-09.md b/REPORT-rehearsal-2026-08-09.md new file mode 100644 index 0000000..1aa1148 --- /dev/null +++ b/REPORT-rehearsal-2026-08-09.md @@ -0,0 +1,183 @@ +# REPORT — BYO reinstall rehearsal, 2026-08-09 (session report) + +*Written as `REPORT-.md` rather than `REPORT.md` per the repo's parallel-session rule.* + +## The answer to the runbook's question, first + +**Yes, the data comes back byte for byte. No, not in one sitting, and not without a shell.** + +All four planted files returned **BYTE-IDENTICAL** — including two Hungarian accented filenames +verified as *raw name bytes*, not as rendered text. The unlock took **21 s**, the restore **13.2 s**. + +But the walk completed only because two hard stops were cleared by someone who could open a terminal +and read source. **R-273**: the install died at step 5/8 on an agent version that was published as a +package but never git-tagged — cleared by completing the release. **R-280**: the reinstalled machine +could not re-attach its own data drive through any dashboard route, while the restore page said +„**Ez két kattintás**" and pointed at an empty list — cleared by POSTing an internal path +(`/mnt/sys_drive`) that no household could produce. + +Neither is a data-integrity problem. Both stop a household dead. **This is the same shape the R-201 +walks kept finding: the data half passes, the journey half fails.** + +## Venue + +**demo-hp (t740)**, operator-approved at STOP 1. It won every fidelity criterion that distinguishes +the two demo boxes: three customer apps against one, a registered storage path against none, and a +Secure-Boot `shim` install — the customer shape — against demo-felhom's SB-off `mkimage` firmware +workaround. `drill-r50` (VM 300) was verified not at risk before proceeding: it is outside the +`felhom` pool, on `local-lvm`, and `--uninstall` removes no storage and no non-pool guest. + +## Findings, ranked by what they cost the person in front of you + +**1 — stops the visit** +- **R-273** · the vouched agent (0.128.0) had no git tag; every install died at 5/8. **CLOSED** — tag + pushed on your instruction after an independent sha check; install then succeeded in 3 m 49 s. The + two guards that would prevent a recurrence are still owed. +- **R-280** · a reinstalled machine cannot re-attach its data drive through any route, and the restore + page promises „két kattintás" at an empty list. **The one to fix before the tester's visit.** +- **R-272** · Felhom's uninstall restarts its own dnsmasq unconstrained; it grabs `:53`; the next + install refuses and appears to blame the owner's network. + +**2 — costs the visit** +- **R-274** · a local golden is adopted with no version and no checksum check. The copy on demo-hp is + controller **0.192.0** against a vouched **0.210.0** — and below 0.200.0, where the recovery screen + the customer needs actually shipped. +- **R-276** · an uninstalled box keeps a live WireGuard tunnel into the off-site endpoint; declared + in neither the KEPT nor the WIPED list. + +**3 — misleads** +- **R-281** · the hub said nothing at all through the whole reinstall, and the tripwire for a + sealed-backup unseal did not fire on a real one. +- **R-282 / R-283** · one code, three names; the mail points at a page the box is not showing; the hub + reads "Claimed 18d ago" while the box serves its setup page. +- **R-269** · a rotated-out local-API token still authorises until an unrelated lookup forces a + reload. The shipped test passes only because of its lookup order. +- **R-270** · R-268's own rotation recipe is a step short; the controller never re-reads the mount. +- **R-271** · the `agent_channel_unauthorized` alarm can never close — its own advice silences the + all-clear. +- **R-277** · three hub surfaces present a healthy off-site tier as absent. **This one caught me.** +- **R-278** · demo-felhom has had no off-site backup for six days, waiting on a ceremony nobody ran. + +**4 — cosmetic / hygiene** +- **R-275** · five orphaned 0600 credential backups survive, and uid reuse hands them to the new + service account. Superseded keys here; live if the backups were recent. +- **R-279** · no operator-triggerable off-site backup exists. + +## What I got wrong, and corrected + +I reported to the operator that the off-site tier had not run on **either** box since 2026-08-03. +That was true of demo-felhom and **false of demo-hp**, which had 18 unbroken daily snapshots. I had +read three hub surfaces that agreed with each other and none of which said what I took them to say +(now **R-277**). I corrected it before it changed any decision, and the operator's "repair off-site +first" ruling turned out to be unnecessary for the chosen venue. + +I also raised **R-275**'s sudoers half as a likely privilege-escalation on reinstall, then **tested it +and refuted my own hypothesis**: sudo skips filenames containing dots, so the leftover file is inert. +`visudo -c -f` parsing a file OK is not evidence that sudo loads it. + +## Integrity verdict — BYTE-IDENTICAL + +``` +expected 4 file(s); found 4 +VERDICT: BYTE-IDENTICAL +``` + +Four expected, four restored, zero differences, compared against +`evidence-rehearsal-2026-08-09/GATE0-before-manifest.json` — a manifest keyed on **raw name bytes**. +`árvíztűrő-tükörfúrógép.txt` and `nested/őszibarack.md` came back with their name bytes intact (NFC +preserved, `c3a1…`), which is the discriminator the Gate 0 positive control was built to enforce: the +comparator had been watched **failing** on an NFC→NFD rename that renders identically to the eye. +Restored out of snapshot `41c830db` into the verification folder the product names, with live data +untouched. + +## Wall clocks + +| phase | duration | +|---|---| +| Pre-phase (R-268 rotation, proved both ways) | ~25 min | +| Gate 0 (venue, dataset, positive control, off-site run, capture) | ~55 min | +| P1 uninstall | **60 s** (08:37:23 → 08:38:23 UTC) | +| P1 leave-behind measurement | ~12 min | +| P2 preflight (3 runs: 2 refusals, 1 pass) | ~6 min | +| P3 install — first attempt, FAILED | 44 s (08:51:33 → 08:52:17 UTC) | +| P3 install — resumed, SUCCESS | **3 m 49 s** | +| P4 first contact (box live on its own URL) | within ~4 min of install | +| STOP 3 unlock | **21 s** | +| app redeploy (calibre-web) | **1 m 36 s** | +| restore prepare + execute | **8 s + 13.2 s** | +| **bare machine → verified files** | **1 h 49 m 22 s** (08:38:23 → 10:27:45 UTC) | +| — of which the product's own work | **≈ 7 m 47 s** | + +**Neither figure is the customer number.** The 1 h 49 m is dominated by the R-273 diagnosis and release +fix (~38 min) and two waits on a human. The 7 m 47 s is what the product costs when the operator +already knows every answer. **The honest unaided figure is undefined, because an unaided household does +not finish.** + +## Steps taken off-path, and what they cost + +1. **R-268 rotation on demo-felhom** — required by the runbook's pre-phase; not the venue. +2. **demo-hp's dashboard password re-set to the credentials-file value**, on operator instruction. + The customer-owned password was unknown to this session and no operator route to the off-site + button exists (R-279). Prior hash preserved in-guest; destroyed with the guest at P1. Cost: none — + P1 wiped it and P4 re-claims. +3. **The off-site run was started by a script pressing the dashboard's own endpoint** with a real + session and CSRF token, not by a person clicking. Identical server path; only the click synthetic. +4. **`--passphrase-file` instead of the no-echo prompt** — a first-class documented option with a + permission check, so the secret still never touched argv. A person would type it. +5. **`systemctl stop dnsmasq && systemctl disable dnsmasq`** — the action the refusal message tells + the owner to take, used as the counterfactual that confirmed R-272. +6. **Pushed the `v0.128.0` git tag** — outward-facing, done on your "proceed", and only after an + independent download proved the published package's sha256 equalled the hub's vouched value. It + completes a half-finished release rather than changing code; the release script's own recovery text + is the same line. **Cost to the walk: the install that followed was a `--resume`, not a fresh run, + which is why R-274 is only half-observed.** +7. **`POST /settings/storage/add` with `/mnt/sys_drive`** — the manual escape hatch, typed. This is the + R-280 wall; a customer could not produce that path. **The biggest fidelity cost of the run.** +8. **SSH into the guest to fingerprint the restored tree.** This is my *instrument*, not a customer + step — the customer's step (the restore) finished at the dashboard. Byte-comparison inherently needs + file access; nothing about the product was driven this way. + +Everything else after the install returned was read-only, and no repair was attempted on the box. + +## Teardown — all four layers + +1. **The machine** — nothing created beyond the half-install itself, which is **left in place + deliberately** for inspection and resumption (`state.json` completed: preflight, token, grows, + enroll). demo-hp is **not serving** right now: no guest, agent installed but no unit. +2. **The host** — `local-lvm` 20 904 790 → 12 355 143 KiB (**≈8.5 GiB returned**); `local` ≈64 MiB; + NVMe unchanged, backups deliberately kept. Pre-existing leftovers found and **not** removed + (not this run's, recorded instead): storage `c11-scratch`, the orphaned `vzdump-lxc-9100` archive, + and `/root/.dpw`, `.h`, `.sec.html` in the old guest (now destroyed with it). +3. **The hub** — **no customer or appliance record was created**; the `demo-hp` customer is retained + deliberately, as the runbook requires. Nothing to delete. +4. **The off-site side** — **one write, and it was the intended one**: the Gate 0 backup that created + snapshots `41c830db`, `9e38b84c`, `78b93f04`. **No prune, no forget, no delete.** How I know: every + restic call was `snapshots`, `ls`, or the product's own `POST /backup/offbox/run`; retention runs + inside that product path and is ep0's server-side job (R-89, boxes keep `keep_last: 0`). + +## Secrets handling + +No secret reached stdout. Token values, the retrieval passphrase, the controller password and the hub +DB copy were handled file→file at 0600 and shredded; the hub DB copy (which carries every host's +break-glass credential) was shredded immediately after the one hash comparison it was taken for. +Credential comparisons were done by sha256 prefix, never by value. + +## State demo-hp was left in + +**Back in service and healthy** — agent 0.128.0, controller 0.210.0, guest 9201 running and onboot, +claimed, storage path registered, `calibre-web` deployed, off-site repository unlocked and intact at +18 snapshots. `drill-r50` (VM 300) untouched throughout. + +**Deliberately left alone, and named rather than tidied:** the pre-existing `c11-scratch` storage and +the three `vzdump-lxc-9100` golden archives on `local` (the teardown keeps goldens by design, and they +now number three). The restored files sit in the product's verification folder, not back in place — +that is R-213 and the product says so. + +## Still owed + +- **R-280** — the drive wall. The one finding that would stop the tester's visit outright. +- **R-273's two guards** — refuse a vouch whose tag does not resolve; check that a package and its tag + ship together. The tag push fixed one box, not the class. +- **R-274's missing observation** — a *fresh* (non-resume) install taking a stale local golden. + +Full account: `documentation/audits/REHEARSAL-byo-reinstall-2026-08-09.md`. diff --git a/STATUS.md b/STATUS.md index 6e536c7..675ba93 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,6 +1,6 @@ # STATUS — what works, what's broken, what's next -**Updated 2026-08-08.** +**Updated 2026-08-09.** > **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates > part of it in plain words, and **nothing may exist only here**. Not `CONTEXT.md`, which is technical @@ -10,6 +10,25 @@ > *Rebuilt from the register on 2026-08-07, from 258 lines. The old "what shipped recently" log is what > the per-repo `CHANGELOG.md` files and the register are for, and is not restated here.* +## ⚠ BOTH DEMO MACHINES ARE OFF AND MUTED — unmute them when they are home + +**Powered down 2026-08-09 14:08 CEST** for the move back from the vacation home. Guests stopped +cleanly first (no vzdump was running, no locks), then the hosts. Confirmed off at the fabric, not +merely unreachable: the tailnet is healthy and both peers report *"offline, last seen 1m ago"*. + +**Both customers are BLOCKED on the hub, deliberately, to stop four false alarms an hour into the +drive.** Blocking gates every monitor and the notification intake; it does **not** gate config pull or +report intake, so the boxes come back normally on power-up. + +> **THE TAIL, and it is the reason this banner exists: while they are blocked, a box that FAILS to +> come back up is also silent.** When the machines are home and powered on, unblock them and confirm +> both report: +> +> Hub → Customers → **demo-hp** → Unblock, and **demo-felhom** → Unblock. +> +> Then check both read ONLINE on Hosts. **Until that is done, the hub cannot tell you either box is +> in trouble.** + ## What works A blank machine boots the Felhom disc, installs itself unattended, and is claimed by the customer, who @@ -23,9 +42,46 @@ destroyed on purpose and its files came back byte for byte identical — four ti their recovery code got everything back with **no command line inside the machine at any point**, in 72 seconds. The two rough edges that walk found are also gone. *(R-201, R-252, R-253 — closed.)* +## The rehearsal finished. The data came back byte for byte; the journey did not. + +**We wiped a working demo machine and put it back. All four test files returned identical — including +the two with Hungarian accents, checked as raw bytes, not as text on screen.** The unlock took 21 +seconds and the restore 13. **But it only finished because I could open a terminal twice.** A +household would have stopped, twice, and the second time the screen would have told them it was easy. + +**The two walls, both fixed-or-fixable, neither about the data:** + +- **The install died four steps in** — the agent version you approved had been published as a download + but never given its version label, and the installer looks it up by that label. **Now unblocked** — + I pushed the label after checking the published file matched what you vouched. *(R-273 — closed. The + two guards that would stop it recurring are still owed.)* +- **A reinstalled machine cannot re-attach its own data drive.** Every route is a dead end, and the + restore page cheerfully says „**Ez két kattintás**" while pointing at an empty list. The drive is + fine and the machine can see it — it just is not offered, because the same drive is also the backup + target. I got past it by typing an internal path no customer could know. **This is the one to fix + before the tester's visit.** *(R-280)* + +**Also broken, found on the way:** + +- **Taking Felhom off a machine leaves the one thing that stops it going back on.** We install a small + network service at setup; removing Felhom restarts it without its settings, it seizes the port the + next install needs, and the next install then refuses — appearing to blame the owner's network. + *(R-272)* +- **A machine we removed keeps its private line to us open.** *(R-276)* +- **The hub said nothing at all** while a machine was wiped, rebuilt, re-claimed and had its sealed + backups opened. No false alarm — but also no word, and the alarm that exists for "someone is opening + this customer's backups" stayed silent through a real one. *(R-281)* +- **A rebuilt machine may still come back on software from last week** — narrower than I first wrote: + the resumed install fetched the right version, but a fresh one takes whatever copy is newest on the + disk without checking it against what you approved. *(R-274)* +- **demo-felhom has not had an off-site backup in six days** and is waiting for a recovery code nobody + has entered. demo-hp, rebuilt the same day, recovered by itself. *(R-278)* +- **One code, three different names**, and the email points at a page the machine is not showing — + this cost us a wasted code today. *(R-282, R-283)* + ## What's broken -- **Nothing new is broken.** The *check* against a fourth secret-in-a-page covers 4 pages of 27, and +- **Nothing else new is broken.** The *check* against a fourth secret-in-a-page covers 4 pages of 27, and the cheap one covering all of them is blind to the shape that shipped. *(R-255)* - **An already-paired box is still told to pair itself**, 25 minutes on. *(R-214, R-235)* - **A backup that covered nothing still calls itself „Sikeres".** *(R-240)* @@ -35,48 +91,71 @@ in 72 seconds. The two rough edges that walk found are also gone. *(R-201, R-252 accumulates. *(R-244)* - **Putting restored files back where they belong is still manual.** *(R-213)* -## Fixed today — four things the machine knew and did not say +## The rest of what the rehearsal found -All one family: something the box already knows, thrown away or drawn as its opposite. +Sixteen findings in one afternoon, **none of them visible from reading the code** — three sessions of +review had not seen any. -- **A rebuilt machine can set up its own recovery again.** The one fact the setup needs was written - only the first time, and a rebuild replaced the configuration while leaving the note saying - "already done". It is now checked and re-written every minute instead of remembered once, so a - hand-edited or restored configuration heals too. **This was the last item blocking a customer from - something we promise them.** *(R-221 — agent 0.128.0.)* -- **A disk we failed to read is no longer drawn as a healthy empty one.** No figures, no bar, and it - says so: „A tárhely mérete most nem olvasható ki." *(R-259 — controller 0.210.0.)* -- **A backup tick now answers about that app.** It went green because *some* backup file existed and - *some other* app's database dump had succeeded most recently. Now: that app's own result, and - **no mark at all** when we have none. *(R-258 — controller 0.210.0.)* -- **Our own alarm no longer points at a page that may not exist.** A check run now gives up after - five minutes rather than hanging until something else kills it, and the mail says how long it ran. - *(R-265.)* +- Removing Felhom leaves five files holding old keys *(R-275)*, and rotating a leaked key does not + revoke the old one until the service restarts *(R-269)* — the written recipe for it is a step short + *(R-270)*, and the alarm it raises can never be closed because the fix it recommends is what + silences the all-clear *(R-271)*. +- **It caught me being wrong twice, and that matters more than the count.** I told you the fleet's + off-site backups were down; demo-hp had eighteen snapshots and I had read three misleading screens + instead of asking the machine *(R-277)*. And I raised a leftover permissions file as a security + hole, then tested it and refuted myself — it is inert. -**Not fixed, and said rather than glossed:** that failed disk reading still reaches us as "0 of -0 GB". It is the quiet direction — it can only miss a true alarm, never raise a false one. *(R-266)* +**What worked, and should not be lost in the count:** the machine came up on its own at the approved +version; the setup page appeared unprompted, in Hungarian, naming the customer; **the recovery screen +appeared without being looked for** and said plainly that unlocking changes nothing; the restore told +the truth about putting files in a checking folder rather than back in place; and no false alarm fired. -**And R-221 is now proven on hardware, not just in tests** — on a demo machine we removed the one -line, watched the setup screen refuse, waited one minute, and watched it go green by itself with -nothing restarted. Every other line of that file came back identical. +Full account: `documentation/audits/REHEARSAL-byo-reinstall-2026-08-09.md`. + +## Three rulings, written down so they stop living in a conversation + +- **The managed-update floor.** It is deliberately parked, and the trigger to raise it is **the first + machine that is not ours**; after that it moves with the publish train. Worth knowing alongside it: + the updater always aims at the floor, never at the newest, so a machine at or above the floor + updates to nothing. **Correction to the number that was going round: the floor is live at 0.200.0, + not 0.156.0** — checked twice today, on the hub page and in both machines' own logs. +- **R-264 is decided.** Build a reader for guest-network health, the staged-update pair, the + restore-test depth pair, and the two backup-integrity timestamps. Record a stated "no reader + wanted" for the repaired-recently flag, the tier-applied timestamp, the config fingerprint, the + per-stack object and the drive-migration marker. Decide the reporting-disabled flag on its own + merits — if that is a state we support, it must be visible or a staleness alarm will one day fire + on a machine that is fine. A "no" ends by changing the allowlist reason from *arguably owed* to + *deliberately not consumed* — not by ripping out an emitter, which is a two-repo change that also + breaks a shared fixture. **The implementation is its own session.** And **it is twenty facts, not + twenty-one**. +- **R-268 is closed** — the leaked key is rotated, and the rotation is proved in both directions + rather than assumed. + +## Fixed 2026-08-08 — four things the machine knew and did not say + +A rebuilt machine can set up its own recovery again *(R-221, agent 0.128.0 — proved on hardware)*; an +unreadable disk is no longer drawn as a healthy empty one *(R-259)*; a backup tick now answers about +*that* app *(R-258)*; our own alarm no longer points at a log that may not exist *(R-265)*. Still true +and not glossed: a failed disk reading still reaches us as "0 of 0 GB" — the quiet direction, it can +only miss a true alarm, never raise a false one *(R-266)*. ## What we're working on -- **Widening the check** so a fourth secret-in-a-page is caught by a machine. *(R-255)* · **Deciding - the twenty-one** — each gets a reader, or stops being sent. *(R-264)* +- **Widening the check** so a fourth secret-in-a-page is caught by a machine. *(R-255)* · **R-264 is + now decided** (above); building the readers is a session of its own. - **Proving the hub really keeps the old sealed key** when a machine re-seals. *(R-198)* · Still open, none urgent: *(R-256, R-257, R-261…R-263, R-266)* ## Waiting on you -- **A watching moment, five minutes.** Today's recovery fix is proved by removing one line from a - demo machine's config — backed up first, disposable machine, no customer data near it — and - watching the setup screen go green on its own. Nothing is destroyed. Say when. -- **One approval, three values this time.** Hub → Configuration → Day-0 artifacts: Golden - **0.210.0**, Agent **0.128.0**, minimum agent **0.127.0** (unchanged) → Save. Each was checked to - be downloadable and selectable before being written here. **Agent 0.128.0 is the one that carries - today's recovery fix**, so a new machine needs both, not just the image. It supersedes the 0.209.0 - approval you already gave, and it is reversible. *(R-242)* +- **Nothing blocking.** The rehearsal is finished and demo-hp is back in service: agent 0.128.0, + controller 0.210.0, claimed, off-site backups unlocked and intact. +- **One decision worth taking before the tester comes:** whether to fix the drive wall *(R-280)* now. + It is the only finding that would stop his visit outright, and it is the difference between "his + data comes back" and "his data comes back if someone types a path for him." +- **Two guards are still owed** so the install cannot break the same way twice: refuse to vouch a + version whose label does not resolve, and check that a published version and its label ship + together. *(R-273's tail.)* ## DooPlex infrastructure — separate from the product @@ -97,11 +176,8 @@ being readable.* connections are kept, the answer is held for a minute, and the old artifacts are gone. Worst case is 5 s, once a minute at most. **Your instinct to prune was right and my measurement said otherwise** — trimming to ten of each halved the slow path. *(R-267 — closed.)* -- **I printed a live access token into a session log** while setting up today's drill, and I am - telling you rather than quietly rotating it. It only opens the agent's private channel to one - demo guest, on a wire that exists solely between that host and that guest — not reachable from - your network or the internet, on a disposable machine with no customer data. Rotating it also - means updating the guest, so it is a deliberate act, not a background one. *(R-268)* +- **The access token I printed into a log yesterday is rotated**, and I checked it both ways: the old + one is refused, the new one works, and the machine's own channel is back up. *(R-268 — closed.)* - **Our build-check alarm has one gap left.** A run that hangs is now cut off after five minutes and the mail says how long it took — but **whether the alarm fires at all when the machinery kills a run outright is still unverified**, and we have not claimed otherwise. *(R-265)* diff --git a/documentation/audits/REHEARSAL-byo-reinstall-2026-08-09.md b/documentation/audits/REHEARSAL-byo-reinstall-2026-08-09.md new file mode 100644 index 0000000..a6324c5 --- /dev/null +++ b/documentation/audits/REHEARSAL-byo-reinstall-2026-08-09.md @@ -0,0 +1,712 @@ +# REHEARSAL — the BYO reinstall walk (2026-08-09) + +> **Status: COMPLETE.** All phases walked; the integrity verdict is **BYTE-IDENTICAL**. §1–§8 were +> written before the destructive phase deliberately — a finding that exists only in a session that +> later crashes is a finding nobody has — and are left as written, including one claim later refuted +> by test (§8.3 → §9a) and one by measurement (§7.2 F-9). + +**Venue: `demo-hp` (HP t740, `felhom-host`, guest 9201, customer `demo-hp`).** Operator-approved at +STOP 1. **Driven from DooPlex.** All times UTC unless marked; the host runs CEST (UTC+2). + +--- + +## 1. Baselines — re-confirmed live on arrival, not taken from the spec + +| | spec said | live reading | source | +|---|---|---|---| +| `felhom-agent` | v0.128.0 @ `28ba8593b8` | **0.128.0** on demo-felhom, **0.127.0** on demo-hp; HEAD == `origin/main` == `28ba8593b8` | `felhom-agent --version` on both nodes; `git rev-parse` | +| `felhom-controller` | v0.210.0 @ `c732fe1283` | **0.210.0** demo-felhom, **0.208.0** demo-hp; HEAD == `origin/main` == `c732fe1283` | hub `/configs`; `git rev-parse` | +| hub | v0.101.0 @ `56f8aa611c` | **0.101.0** (deployed image tag matches) | `kubectl get deploy hub`; page footer | +| Day-0 manifest | golden 0.210.0 / agent 0.128.0 / min 0.127.0 | **all three already saved** — golden `0.210.0` (`b9f701fa…`), agent `0.128.0` (`c6eba73b…`), min agent `0.127.0`, wrapper `104db0a4…` | hub `/configuration`, selected `