teardown of two venues; R-218/R-220 CLOSED, R-236 WITHDRAWN, R-238 reclassified
gates / gates (push) Successful in 18s
gates / gates (push) Successful in 18s
Part 0 — c11 and rewalk destroyed, three layers each plus the off-site side and
the WireGuard peer, via the hub's own cascade (external teardown FIRST, DB purge
LAST). 37.3 GB reclaimed on c11-scratch, matching the 20G+17G measured. Positive
control after each: part4 must still be found, and was. ep0 namespaces now exactly
demo-felhom, demo-hp, part4.
The cascade refuses to delete a live host and there is no decommission endpoint,
so both boxes were stopped and aged past the hub's 30m stale_threshold first.
Register corrections — the durable record was wrong about two shipped fixes:
R-218 REOPENED -> CLOSED. Shipped controller v0.203.0, proven live on a
genuinely rebuilt box: the hub re-staged at 13:24:57Z and the box
collected it on a tick, no guest command line, no operator action.
R-220 "OPEN — NOT FIXED" -> CLOSED. Shipped agent v0.127.0, proven live after a
real guest purge with both raw mounts still on the surviving host:
/disks/candidates returned both drives (before: two empty lists) and both
re-attached through the customer endpoint.
R-236 WITHDRAWN — I FILED THIS WRONGLY. The hub log shows offsiteheal re-staged
the stored secret at 13:24:57Z after its documented two-report debounce, with no
provider credential minted. My Re-issue at 13:26:39Z came 102s LATER, was
redundant, and minted an unnecessary provider credential (subaccount 284735) —
the very double-issue the offsite-delivery guard warns about once a minute in the
log. "Nincs teendod" is true; I did not wait ~16 minutes. Operational lesson, not
a product defect.
R-238 reclassified as a harness artifact (mode=full without confirm=1 is step 1 of
a deliberate two-step and starts no job by design); its real residue — the total
silence of that step — is fixed in controller v0.204.0.
R-237 CLOSED by controller v0.204.0.
This commit is contained in:
@@ -111,13 +111,13 @@ and `journal-phase24.md` (Phases 2/4). Campaign document:
|
||||
| ID | What | State |
|
||||
|---|---|---|
|
||||
| **R-216** | **A correct recovery code was reported to the customer as wrong.** The unseal needs agent v0.125.0; on an older agent the route 404s and the unlock was attempted anyway, producing *„…nem fogadtuk el. Ellenőrizd, hogy mind a tíz szót pontosan…"* in 0.134 s. **The default state** — the Day-0 manifest vouches 0.120.0, and a reinstall actively DOWNGRADES a hand-fixed box back to it. The hub's own guard could not catch it: `ResolveManagedFloor` compared against the GOLDEN's MinAgent while serving a FLOOR that pointed elsewhere (the 8th entry in `CLAUDE.md`'s comment-vs-code table, and the first where the false invariant was a guard) | **SHIPPED** (controller v0.201.0 + hub v0.97.0/0.97.1) — **but see R-223**: the feature does not work on a NEW box until the manifest vouches agent 0.125.0. Until then such a box is correctly HELD, not lied to |
|
||||
| **R-218** | **Succeeding at recovery stopped the box asking for what it still needed.** `needsOffsiteCredential` short-circuited on a repository password existing — and installing one is the recovery screen's whole job. 32 s after the hub re-staged the credential, the customer's success switched the mechanism off; the hub held an unconsumed credential the box had no reason to collect, and nothing ever asked again | **REOPENED 2026-08-06 — the fix covers the DECLARATION half only, and this row over-claimed it.** Measured on the R-201 re-walk: the box declared, and **`offsiteheal` re-staged the secret at 11:44:57** saying *"the box re-consumes on its next cycle"* — **the next cycle came and went** (`host-report` 11:55:46, `Received report` 11:55:54, a full cycle **with a positive control that it ran**) **and the credential was still not consumed.** 23 minutes after the re-stage the box's last off-site-apply attempt was still the pre-re-stage one. A census of the customer-reachable actions on `/backups/remote` (`config`, `reset`, `run`, `toggle`) found **none that fetches a staged credential**, and the only lever is `systemctl restart felhom-controller-bootstrap.service` **inside the guest** — which worked in **18 s** (Campaign 11 measured 17), confirming nothing was wrong with the credential, the target or the key: **the only thing missing is anything at all to trigger a retry.** **This is the FIRST of the two dead ends that keep the recovery journey failing** |
|
||||
| **R-218** | **Succeeding at recovery stopped the box asking for what it still needed.** `needsOffsiteCredential` short-circuited on a repository password existing — and installing one is the recovery screen's whole job. 32 s after the hub re-staged the credential, the customer's success switched the mechanism off; the hub held an unconsumed credential the box had no reason to collect, and nothing ever asked again | **REOPENED 2026-08-06 — the fix covers the DECLARATION half only, and this row over-claimed it.** Measured on the R-201 re-walk: the box declared, and **`offsiteheal` re-staged the secret at 11:44:57** saying *"the box re-consumes on its next cycle"* — **the next cycle came and went** (`host-report` 11:55:46, `Received report` 11:55:54, a full cycle **with a positive control that it ran**) **and the credential was still not consumed.** 23 minutes after the re-stage the box's last off-site-apply attempt was still the pre-re-stage one. A census of the customer-reachable actions on `/backups/remote` (`config`, `reset`, `run`, `toggle`) found **none that fetches a staged credential**, and the only lever is `systemctl restart felhom-controller-bootstrap.service` **inside the guest** — which worked in **18 s** (Campaign 11 measured 17), confirming nothing was wrong with the credential, the target or the key: **the only thing missing is anything at all to trigger a retry.** **This is the FIRST of the two dead ends that keep the recovery journey failing** | **CLOSED 2026-08-06 — shipped in controller v0.203.0 and PROVEN LIVE on a genuinely rebuilt box.** `Bridge.RetryIfDeclared` re-runs the SAME reconcile on a 5-minute tick for exactly as long as the box's own published declaration (`OffboxReportStatus().State`) says it needs a credential — the very statement the hub acts on, so the two can never disagree. Measured on the Part 4 venue (`documentation/tests/part4-rewalk-2026-08-06/journal.md`): after the rebuild the retry ticked and logged its reason honestly, the hub's `offsiteheal` re-staged at **13:24:57Z**, and the box **collected it on a tick and configured the tier** — `[offsite-apply] offsite configured for u629488-sub6@…:/home/felhom-repo` — **with no guest command line and no operator action**. On a healthy box the job registers, ticks, completes in 0 s and says nothing, which is asserted. |
|
||||
| **R-219** | **The listing the screen promises could never render on the shape it exists for.** Listing needs a target; a target cannot exist without a repository password; shape (a) is defined by having none. And placing the key flipped the offer false, so the unlock response was the only chance and was guaranteed not to contain it | **SHIPPED** (controller v0.201.0) — the unlock now places the key, brings the tier up, then lists |
|
||||
| **R-217** | **An unreadable store reported as "opened, with unattributable content".** The failure path passed `backup.OffsiteInventory{}`, whose `Empty=false` the template read as `InvUntagged`. The field built to prevent exactly this names the hazard in its own doc comment | **SHIPPED** (controller v0.201.0) — opened / empty / unreadable are three distinguishable states |
|
||||
| **R-222** | **Reaching for a RETAINED earlier package read as a wrong code.** The engine is right (it fails closed against the current package); the message was not. Proven live: the correct code for the orphaned history got *„check your ten words"* | **SHIPPED** (controller v0.201.0 + hub v0.97.0) — the ACK carries `superseded_present`/`superseded_at` and the screen names the situation. **It states what the hub knows and promises nothing** — the read path is still unbuilt (R-199's inventory) |
|
||||
| **R-215** | **`GET /recovery` rendered the recovery story on a box that never had off-site backups.** The predicate was right and the page never asked it; the POST sibling and the backups-area template both gated the same sentence correctly | **SHIPPED** (controller v0.201.0) |
|
||||
| **R-214** | **The physical console never stops asking to be paired.** Half an hour after `Day-0 provision SUCCESS`, with the host ONLINE, the console still showed the pairing banner and a stale code — on a screen whose own text promises *„Ez a képernyő magától frissül"*. Census: exactly two `/dev/console` writers in the whole day-0 path, both in the pairing loop; `felhom-host-install.sh` writes to the console not at all | **OPEN — NOT FIXED** |
|
||||
| **R-220** | **After a rebuild the customer's drives cannot be re-enrolled, and the refusal names an impossible action.** The deploy refuses (*„Válasszon a listából csatlakoztatott meghajtót"*) and the list is empty: `claim.go:84` treats a device mounted outside `/mnt/felhom-drives` as claimed, and the raw `/mnt/<name>` mount that enrolment itself creates survives the guest rebuild while the controller's registry does not. **Red-proved**: unmounting only the raw mounts flipped `attach: []` → both drives. **This is the state Campaign 10 reached by hand and recorded as its own harness error; the product's rebuild path now arrives there.** Breaches I3 | **OPEN — NOT FIXED** |
|
||||
| **R-220** | **After a rebuild the customer's drives cannot be re-enrolled, and the refusal names an impossible action.** The deploy refuses (*„Válasszon a listából csatlakoztatott meghajtót"*) and the list is empty: `claim.go:84` treats a device mounted outside `/mnt/felhom-drives` as claimed, and the raw `/mnt/<name>` mount that enrolment itself creates survives the guest rebuild while the controller's registry does not. **Red-proved**: unmounting only the raw mounts flipped `attach: []` → both drives. **This is the state Campaign 10 reached by hand and recorded as its own harness error; the product's rebuild path now arrives there.** Breaches I3 | **CLOSED 2026-08-06 — shipped in agent v0.127.0 and PROVEN LIVE on a genuinely rebuilt box.** The fix is **corroborated, not a widened prefix**: a mountpoint outside `/mnt/felhom-drives` is forgiven only when the SAME device is also mounted under the managed path — a pairing only Felhom's own enrolment produces, so a disk another system is using at `/srv/data` or even `/mnt/someone-elses-disk` is still refused (own test + red-proof). Read from `/proc/mounts` deliberately: the lsblk invocation is pinned verbatim in the sudoers file, so switching to plural `MOUNTPOINTS` would have shipped a sudoers change with the binary. Fail-safe: an unreadable mount table corroborates nothing. **Measured on the Part 4 venue after a real guest purge, with both raw mounts still present on the surviving host:** `/disks/candidates` returned both drives in `attach` and `initialize` (before the fix: two empty lists), and both **re-attached through the customer endpoint** (`registered: true`). The customer-facing refusal was corrected in controller v0.203.0. |
|
||||
| **R-221** | **A rebuilt box cannot run the escrow ceremony at all.** The preflight refuses on `escrow.pbs_storage_id`, which the pbsdr bridge seeds into `agent.json` only via `finishConverged`. The convergence marker lives on the HOST and survives a guest rebuild; `agent.json` is rewritten by the installer. Unchanged descriptor → same hash → early return → the seed never runs into a config that no longer has it. **Red-proved**: moving only the marker aside seeded it instantly (`grep -c escrow`: 0 → 1). A real blocker for re-escrow, which is exactly what a rebuilt box must do | **OPEN — NOT FIXED** |
|
||||
| **R-223** | **The Day-0 manifest vouched agent 0.120.0 while the recovery feature needs 0.125.0 — and a reinstall DOWNGRADES a box that was fixed by hand.** Verbatim from the second reinstall: `agent (existing): felhom-agent 0.125.0` → `manifest: agent v0.120.0` → `installed /usr/local/bin/felhom-agent (felhom-agent 0.120.0)`. So every rebuild re-broke the recovery path — the one event that makes the feature necessary. **⚠ AND IT WAS NOT A DROPDOWN.** The first vouch attempt was REFUSED by R-120's gate (`configs.go:1162`): *"golden 0.192.0 is older than the newest controller the fleet reports (0.201.0)"*. **The artifacts form saves as a unit, so the agent could not be vouched while the golden was stale — and the golden had been stale since controller 0.193.0, meaning the Day-0 manifest had been effectively UNVOUCHABLE for days and nobody had cause to notice.** The real remedy was a golden rebake | **CLOSED 2026-08-05.** Golden **0.201.0** baked in the drill VM (658 165 766 B, sha `e730d7cab343eb35…f007654`, **round-trip verified from Gitea**), then manifest set in one save: `agent=0.125.0 golden=0.201.0 min_agent=0.125.0`. A fresh install now lands on current agent AND current controller |
|
||||
|
||||
@@ -160,9 +160,9 @@ the fault was real. Full observables: `tests/campaign11-evidence-2026-08-05/jour
|
||||
| **R-234** | **An off-site run reports success while silently omitting an app the customer just switched on.** Found 2026-08-06 on the Part 4 venue (`part4`, VM 323), and found ONLY because the pre-destruction verification restore was run instead of trusting the green tick. Sequence, measured: a run with no app selected produced **1 snapshot**; `POST /backup/offbox/toggle` enabled `calibre-web` (HTTP 302, and the „Nincs távoli mentésre jelölt alkalmazás" warning disappeared, so the selection HAD landed); the next run finished in 30 s and reported **„✓ Rendben · 12.0 MB · 1 pillanatkép"** — still one snapshot. The app restore then refused: **„offbox: nincs pillanatkép a(z) calibre-web alkalmazáshoz"**. A THIRD run took the snapshot count to 2 and the same restore then succeeded. So a run that the customer sees as a green success did not carry the app they had just enabled, and **nothing in the card distinguishes that from a run that did**. The customer's belief ("my app is off-site") and the truth diverge silently, and they would discover it only at restore — the worst possible moment. The snapshot count is on the same card, which is what makes the omission detectable in hindsight and invisible in the moment. **Not yet root-caused**: the likely shape is that the run captured the app-selection set before the toggle committed, but that is a hypothesis, not a measurement. **This is the exact class the project already has a rule for** — "presence is not success": the run's timestamp and tick record that a run HAPPENED, not that it carried what the customer asked for. | **READY** — owner Viktor |
|
||||
| **R-235** | **The appliance console keeps telling an already-paired box to go and pair itself.** Measured 2026-08-06 on VM 323: **25 minutes after** the operator bind, with the guest provisioned, the controller reporting 0.203.0 and the agent ONLINE, the physical console still displayed „Felhom — a doboz készen áll, és a **párosításra vár**" together with the now-spent pairing code `US3-6GP` — and, in the same panel, the promise **„Ez a képernyő magától frissül — nincs teendő a doboznál"**. It does not refresh. A customer looking at their screen is told the setup has not happened, and is told the screen would have updated if it had. Cosmetic in mechanism, not in effect: it invites the customer to re-pair a working box, or to call for help about a box that is already fine. Same family as R-234 — a surface asserting a state that stopped being true. | **READY** — owner Viktor |
|
||||
|
||||
| **R-236** | **After a guest rebuild the hub never re-stages the off-site credential, so the customer's "automatic" recovery stalls until an operator notices.** Measured 2026-08-06 on the Part 4 venue. The unlock screen promises „A gép még várja a házon kívüli tárhely kapcsolódási adatait — amint megvannak, a mentéseid listája megjelenik… **Nincs teendőd**". What actually happens: the rebuilt box declares `needs_credential`, R-218's retry job ticks every 5 min exactly as designed and logs the honest reason — `[offsite-apply] credential retry: consume one-time password: no unconsumed offsite password (already consumed or none provisioned) (the box still declares a need; retrying)` — **forever**, because the one-time password was consumed by the guest that no longer exists and nothing mints a new one. The fix is a single operator action that already exists (`/configs/<id>/offsite-reissue`), and pressing it resolved the stall within one tick. **R-218's consume half is not at fault — it is the half that works**; the gap is upstream, in who re-stages after a rebuild. The customer-facing copy is honest in its fallback („Ha egy napon belül nem áll be, jelezd az üzemeltetőnek") but the headline „Nincs teendőd" is not true for the rebuild case, which is precisely the case the recovery feature exists for. | **READY** — owner Viktor |
|
||||
| **R-237** | **After a successful recovery the customer is shown no backups at all, because the restore surface is keyed on apps that are currently installed and currently marked for future remote backup.** Measured 2026-08-06. Post-rebuild, with the key recovered, the tier configured and the escrow re-sealed, `/backups/restore` said „**Nincs telepített alkalmazás**" and „Nincs távoli mentésre jelölt alkalmazás — a kijelölés a Távoli mentés oldalon történik" — pointing at a page which itself said there were no installed applications. A **circular dead end**: to restore an app you must select it; to select it, it must be installed; to know what to install, you must see the backup you cannot see. Reaching the restore actually required three undocumented steps in order — re-attach the drives, redeploy the app, and toggle it on for **future** remote backups (`/backups/restore/app` refuses with „Ez az alkalmazás nincs távoli mentésre kijelölve", i.e. a *forward-looking* setting gates a *backward-looking* action). None of this is hinted at by the recovery screen, which says „Nincs teendőd". A household that has just lost its box does not know which apps it used to run. | **READY** — owner Viktor |
|
||||
| **R-238** | **„Teljes visszaállítás előkészítése" — the only path to the customer's own files — accepts the click and does nothing, silently.** Measured 2026-08-06, repeatedly. `POST /backup/offbox/restore` with `mode=full` (the exact form the page renders) returns **302**, and then: no job is ever recorded (`/api/backup/restore-status` `last` stays `null` across ~9 minutes of polling), the wizard stays on step 1 „Előkészítés" with the same two forms, no error is shown to the customer, and **the controller's own debug ring contains no line for it at all** — zero `offbox`, `prepare` or `snapshot` entries. `mode=unit` on the identical form works and reports properly, so the plumbing and the session are fine. This is the terminal dead end of the recovery journey: everything upstream succeeded — code accepted, key recovered, tier configured, escrow re-sealed, drives re-attached, app redeployed — and the customer still cannot get their files back. **PASS for the R-201 re-walk was defined as a sentinel's sha256 byte-identical after recovery; this is why that criterion was NOT met.** Related but distinct from R-237: that one is about not being able to *find* the backup, this one is about the button not working once found. | **READY** — owner Viktor |
|
||||
| **R-236** | ~~**After a guest rebuild the hub never re-stages the off-site credential.**~~ **WITHDRAWN 2026-08-06 — THIS WAS WRONG, and the mechanism works.** Measured 2026-08-06 on the Part 4 venue. The unlock screen promises „A gép még várja a házon kívüli tárhely kapcsolódási adatait — amint megvannak, a mentéseid listája megjelenik… **Nincs teendőd**". What actually happens: the rebuilt box declares `needs_credential`, R-218's retry job ticks every 5 min exactly as designed and logs the honest reason — `[offsite-apply] credential retry: consume one-time password: no unconsumed offsite password (already consumed or none provisioned) (the box still declares a need; retrying)` — **forever**, because the one-time password was consumed by the guest that no longer exists and nothing mints a new one. The fix is a single operator action that already exists (`/configs/<id>/offsite-reissue`), and pressing it resolved the stall within one tick. **R-218's consume half is not at fault — it is the half that works**; the gap is upstream, in who re-stages after a rebuild. **What the hub log actually shows** (`kubectl logs`, 2026-08-06): `offsite-delivery` refused to self-heal once a minute, correctly, naming `internal/offsiteheal` as the owner of the remediation and warning that a second mechanism minting there would double-issue; then at **15:24:57 local (13:24:57Z)** — `offsiteheal: re-staged the stored one-time offsite secret for customer part4 (declared needs_credential across 2 reports) — no provider credential was minted`. **The reconciler fired on schedule and succeeded.** The operator Re-issue at **13:26:39Z** came **102 seconds LATER**, was redundant, and **minted an unnecessary provider credential** (`re-issued shared credentials for part4 (subaccount 284735)`) — exactly the double-issue the guard exists to prevent. **The error was mine: I did not wait for the documented two-report debounce** (~16 minutes after a rebuild, since reports are ~15 min apart and the reconciler ticks every 5), and attributed the recovery to my own button press. „Nincs teendőd" is TRUE, and the wait is far inside the „egy napon belül" the card promises. **Operational lesson, not a product defect:** after a rebuild, wait for the debounce — the hub says once a minute, in its own log, that it owns this remediation. | **CLOSED 2026-08-06 — not a defect** |
|
||||
| **R-237** | **After a successful recovery the customer is shown no backups at all, because the restore surface is keyed on apps that are currently installed and currently marked for future remote backup.** Measured 2026-08-06. Post-rebuild, with the key recovered, the tier configured and the escrow re-sealed, `/backups/restore` said „**Nincs telepített alkalmazás**" and „Nincs távoli mentésre jelölt alkalmazás — a kijelölés a Távoli mentés oldalon történik" — pointing at a page which itself said there were no installed applications. A **circular dead end**: to restore an app you must select it; to select it, it must be installed; to know what to install, you must see the backup you cannot see. Reaching the restore actually required three undocumented steps in order — re-attach the drives, redeploy the app, and toggle it on for **future** remote backups (`/backups/restore/app` refuses with „Ez az alkalmazás nincs távoli mentésre kijelölve", i.e. a *forward-looking* setting gates a *backward-looking* action). None of this is hinted at by the recovery screen, which says „Nincs teendőd". A household that has just lost its box does not know which apps it used to run. | **CLOSED 2026-08-06** — controller v0.204.0: the list is now built from `OffsiteInventoryList` (the repository's own snapshot tags). Installed-ness became a property OF a row, never a filter; an unreadable store renders as UNKNOWN **and keeps the action offered**; `felhom-offbox` and `_shares` are excluded. 7 new tests incl. a rendered-page test for the rebuilt shape, and a red-proof that keys the list back on installed-and-toggled apps. |
|
||||
| **R-238** | ~~**„Teljes visszaállítás előkészítése" accepts the click and does nothing.**~~ **RECLASSIFIED 2026-08-06 — a HARNESS ARTIFACT, with a real residue that IS fixed.** Measured 2026-08-06, repeatedly. `POST /backup/offbox/restore` with `mode=full` (the exact form the page renders) returns **302**, and then: no job is ever recorded (`/api/backup/restore-status` `last` stays `null` across ~9 minutes of polling), the wizard stays on step 1 „Előkészítés" with the same two forms, no error is shown to the customer, and **the controller's own debug ring contains no line for it at all** — zero `offbox`, `prepare` or `snapshot` entries. `mode=unit` on the identical form works and reports properly, so the plumbing and the session are fine. This is the terminal dead end of the recovery journey: everything upstream succeeded — code accepted, key recovered, tier configured, escrow re-sealed, drives re-attached, app redeployed — and the customer still cannot get their files back. **PASS for the R-201 re-walk was defined as a sentinel's sha256 byte-identical after recovery; this is why that criterion was NOT met.** **The diagnosis** (`restore_wizard.go`, `offbox_handlers.go:314`): `mode=full` without `confirm=1` is **step 1 of a deliberate two-step** — it computes size + headroom, **starts no job by design**, and redirects carrying `&full_prep=<app>&full_size=<size>`. `deriveWizardStep` reveals the commit **only** when `FullPrepApp == App`. The endpoint driver posted step 1 and re-fetched the wizard **without** that parameter, so the pure function correctly returned the intent step, and `restore-status.last == null` was correct too. **The operator drove the same restore to completion in a browser.** The wizard was NOT re-keyed. **The residue was real and is fixed in controller v0.204.0:** neither branch of step 1 wrote anything to the log — `offboxRedirectTo` only flashes to the page — so a refusal, **including by the headroom gate**, left no trace on the box. Both branches now log, as does the concurrent-op refusal; pinned by a handler-level test with a red-proof. | **CLOSED 2026-08-06** — controller v0.204.0 |
|
||||
|
||||
**Recorded against existing rows by Phase 2:**
|
||||
|
||||
|
||||
Reference in New Issue
Block a user