docs: CAMPAIGN-11 — the journey FAILED, R-198's retention PROVEN, six findings fixed
gates / gates (push) Successful in 8s

Registers and evidence for the campaign and its fix pass.

OPEN-ITEMS: R-214..R-223. Six SHIPPED (R-215/216/217/218/219/222); three deliberately
still open and each blocks a real flow (R-214 console banner, R-220 drives unenrollable
after a rebuild, R-221 a rebuilt box cannot run the escrow ceremony); R-223 minted and
WAITING-ON-OPERATOR (vouch agent 0.125.0). R-213 and R-202 untouched.

Capability map: a new row for the customer's UNAIDED journey, recorded FAILED and staying
failed until a re-walk passes — fixes are not a journey. The existing rebuild row is
corrected where it said R-198's retention was unit-proven only: it was proven in production
on the first supersession since the fix, identity_blob retained at 572 B byte-length exact.

CLAUDE.md comment-vs-code table: eighth entry — ResolveManagedFloor, the first where the
false invariant was a GUARD rather than a comment alone.

STATUS: the headline is now "the backup promise is proved, the recovery journey is not",
and the one thing waiting on the operator.
This commit is contained in:
2026-08-05 18:04:13 +02:00
parent 79e31ac24b
commit 1a0f7db92f
5 changed files with 1200 additions and 30 deletions
+41 -28
View File
@@ -15,41 +15,46 @@ who sets their own password. They install apps from a catalogue of fifty-three,
home network, and open apps from a launcher or a shared link. Backups run on their own to three
places — the machine's drive, a second drive, and an encrypted off-site copy.
**And the whole backup promise is now proved.** On 4 August we destroyed a machine on purpose and
deleted a marked file from its disk. Using the recovery code you saved: the backup key came back
**identical, character for character**; the existing off-site store **opened** rather than starting
over; and the file was restored **byte for byte identical**. *(R-201)*
**The backup promise is proved. The recovery JOURNEY is not.** On 5 August we built a brand-new
machine from the published disc, gave it three marked files, destroyed it, and tried to get them back
**the way a household would** — no shortcuts, no command line. The files came back **byte for byte
identical**, all three, including one with Hungarian accents in its name. But **the journey needed us
four times**, and the very first thing the machine did was tell the customer their correct recovery
code was wrong. *(CAMPAIGN 11)*
## What's broken
- **A customer can now get their backups open on their own — but not yet put the files back.** Every
step from a rebuilt machine to an open backup store is done, and today the last piece landed: a
**full screen** meets the owner of a rebuilt machine, explains that the backups are still there,
says plainly that **nobody can replace a lost recovery code**, takes the code, and shows what is in
the store — which apps, from when, how big. Nothing needs you, and nothing needs a command line.
**What it deliberately does not do is put files back.** That is per-app, in the backups area, and
the piece that would guide it — showing what would change before anything is overwritten — is not
built yet. *(R-193 closed; the put-back is R-213)*
- **Two things still owed as proof.** The final unlock has never been done with a **correct** code
through the new screen: no recovery code was kept for the N100 machine's orphaned history, and the
HP machine's is in your hands, not ours — so the live test ran the whole chain and stopped at the
last step. **And the whole journey has not been re-run end to end since these fixes** — the pieces
are proved one at a time, not as a single walk. That re-run is one more drill. *(R-201)*
- **A machine installed today would tell its owner their correct recovery code is wrong.** The
recovery screen needs a newer in-house service than a new machine is given; when it asked and got
nothing, it blamed the customer's typing. **We fixed the lie today** — it now says plainly that the
*machine* cannot do this yet, and never accuses anyone. **It still needs one click from you to
actually work on new machines** (below). *(R-216, R-223)*
- **Three things a rebuilt machine still cannot do by itself.** Its owner cannot re-attach their own
drives, so no app can be put back on its data; it cannot create a new recovery code at all; and the
screen at the machine itself never stops showing a stale pairing code. Each is understood, measured
and written down — none is fixed yet. *(R-220, R-221, R-214)*
- **Rebuilding a machine still throws away its off-site backup HISTORY.** The machine invents the key
that encrypts its own off-site backups, and a rebuilt machine invents a brand-new one. Both demo
machines did this on 34 August — 51 backups (~1.2 GB) between them. The old key now survives the
recovery ceremony, and a changed key now raises an alarm the same day, but **the rebuild itself
still starts a fresh history**. *(R-193)*
- **One screen still tells the customer something we cannot yet promise.** The „elárvult tároló" card
says the old backups may later be restorable with the matching recovery code. That is true for
machines that re-seal from now on and **false for anything already orphaned** — and the machine
cannot tell which case it is in. We deliberately did **not** patch the sentence: a conditional
promise that can still be wrong is worse there than a vague one. *(R-202)*
- **The off-site copy can be erased by the machine that made it.** The credential that writes it can
also delete it. A daily snapshot is armed as a stopgap. *(R-95, R-87)*
that encrypts its own off-site backups, and a rebuilt machine invents a new one. **The good news:
the old key really is kept now — we proved it on a real machine today, for the first time**, and a
changed key raises an alarm the same day. *(R-193, R-198)*
- **The kept older backups cannot be opened yet.** We keep the previous sealed package, and there is
no way to open it. A customer holding exactly the right code for it used to be told they had
mistyped; today the screen names the situation honestly instead — but it still cannot open it, and
it does not pretend otherwise. *(R-222, R-202)*
- **The off-site copy can be erased by the machine that made it.** A daily snapshot is armed as a
stopgap. *(R-95, R-87)*
## What shipped recently
- **2026-08-05 (evening)** — **Four sentences where there was one, and none of them blames you.**
"We did not accept your recovery code" used to appear when the code was wrong, when the machine
could not ask, when the store could not be read, and when the customer held the code for an older
backup we still keep. Only the first is the customer's doing. Also today: succeeding at recovery no
longer switches off the machine's own request for the thing it still needs; the screen now finishes
the job and shows what is in the backups instead of promising a list it could never produce; and the
recovery page can no longer be reached on a machine that never had backups.
*(R-216, R-217, R-218, R-219, R-222, R-215)*
- **2026-08-05** — **The recovery screen: a customer whose machine was rebuilt is now told, and shown
how.** Until today they had everything needed to get their data back and no way to find out — the
only route was a command line. The screen unlocks the backups and lists what is in them; it does
@@ -113,6 +118,14 @@ over; and the file was restored **byte for byte identical**. *(R-201)*
## Waiting on you
- **One click, and it is the most valuable one available: approve host-service version 0.125.0 for new
machines.** Hub → Configuration → Day-0 artifacts → agent. Today new machines get 0.120.0, which
**cannot** open a recovery package — and a reinstall actively puts the older one back over a machine
we fixed by hand, so every rebuild re-breaks the very thing a rebuild needs. 0.125.0 has run on both
demo machines since 4 August and through the entire campaign. **Until you do this, new machines are
correctly held back rather than lied to — which is better, but the feature does not work for them.**
*(R-223)*
- **One thing to read after the machine next restarts — and nothing to do until then.** You told me
not to restart DooPlex, so I did not, and the move to the second SSD has therefore never been
through a restart. It works right now and nothing was lost, but a restart is the one test that
File diff suppressed because one or more lines are too long
+26
View File
@@ -55,6 +55,32 @@ v0.93.0 made the key survive. v0.94.0/v0.125.0/v0.195.0 make it come back. v0.19
the snapshot and the drill got it out again. **v0.198.0/v0.95.0 remove three of the four crutches the
drill needed — the fourth is R-193, and until it goes the recovery is still operator-assisted.**
## CAMPAIGN 11 — the recovery journey, 2026-08-05
**The whole journey was walked end to end for the first time, on a throwaway appliance built from the
published ISO. The data came back byte-identical; the journey did not exist.** Nine findings, R-214 …
R-222. Six are fixed in controller **v0.201.0** + hub **v0.97.1**; **three are deliberately still
open** and each blocks a real flow. Evidence: `tests/campaign11-evidence-2026-08-05/journal.md`.
| ID | What | State |
|---|---|---|
| **R-216** | **A correct recovery code was reported to the customer as wrong.** The unseal needs agent v0.125.0; on an older agent the route 404s and the unlock was attempted anyway, producing *„…nem fogadtuk el. Ellenőrizd, hogy mind a tíz szót pontosan…"* in 0.134 s. **The default state** — the Day-0 manifest vouches 0.120.0, and a reinstall actively DOWNGRADES a hand-fixed box back to it. The hub's own guard could not catch it: `ResolveManagedFloor` compared against the GOLDEN's MinAgent while serving a FLOOR that pointed elsewhere (the 8th entry in `CLAUDE.md`'s comment-vs-code table, and the first where the false invariant was a guard) | **SHIPPED** (controller v0.201.0 + hub v0.97.0/0.97.1) — **but see R-223**: the feature does not work on a NEW box until the manifest vouches agent 0.125.0. Until then such a box is correctly HELD, not lied to |
| **R-218** | **Succeeding at recovery stopped the box asking for what it still needed.** `needsOffsiteCredential` short-circuited on a repository password existing — and installing one is the recovery screen's whole job. 32 s after the hub re-staged the credential, the customer's success switched the mechanism off; the hub held an unconsumed credential the box had no reason to collect, and nothing ever asked again | **SHIPPED** (controller v0.201.0) — the declaration now stops when the TIER WORKS, not when a key exists |
| **R-219** | **The listing the screen promises could never render on the shape it exists for.** Listing needs a target; a target cannot exist without a repository password; shape (a) is defined by having none. And placing the key flipped the offer false, so the unlock response was the only chance and was guaranteed not to contain it | **SHIPPED** (controller v0.201.0) — the unlock now places the key, brings the tier up, then lists |
| **R-217** | **An unreadable store reported as "opened, with unattributable content".** The failure path passed `backup.OffsiteInventory{}`, whose `Empty=false` the template read as `InvUntagged`. The field built to prevent exactly this names the hazard in its own doc comment | **SHIPPED** (controller v0.201.0) — opened / empty / unreadable are three distinguishable states |
| **R-222** | **Reaching for a RETAINED earlier package read as a wrong code.** The engine is right (it fails closed against the current package); the message was not. Proven live: the correct code for the orphaned history got *„check your ten words"* | **SHIPPED** (controller v0.201.0 + hub v0.97.0) — the ACK carries `superseded_present`/`superseded_at` and the screen names the situation. **It states what the hub knows and promises nothing** — the read path is still unbuilt (R-199's inventory) |
| **R-215** | **`GET /recovery` rendered the recovery story on a box that never had off-site backups.** The predicate was right and the page never asked it; the POST sibling and the backups-area template both gated the same sentence correctly | **SHIPPED** (controller v0.201.0) |
| **R-214** | **The physical console never stops asking to be paired.** Half an hour after `Day-0 provision SUCCESS`, with the host ONLINE, the console still showed the pairing banner and a stale code — on a screen whose own text promises *„Ez a képernyő magától frissül"*. Census: exactly two `/dev/console` writers in the whole day-0 path, both in the pairing loop; `felhom-host-install.sh` writes to the console not at all | **OPEN — NOT FIXED** |
| **R-220** | **After a rebuild the customer's drives cannot be re-enrolled, and the refusal names an impossible action.** The deploy refuses (*„Válasszon a listából csatlakoztatott meghajtót"*) and the list is empty: `claim.go:84` treats a device mounted outside `/mnt/felhom-drives` as claimed, and the raw `/mnt/<name>` mount that enrolment itself creates survives the guest rebuild while the controller's registry does not. **Red-proved**: unmounting only the raw mounts flipped `attach: []` → both drives. **This is the state Campaign 10 reached by hand and recorded as its own harness error; the product's rebuild path now arrives there.** Breaches I3 | **OPEN — NOT FIXED** |
| **R-221** | **A rebuilt box cannot run the escrow ceremony at all.** The preflight refuses on `escrow.pbs_storage_id`, which the pbsdr bridge seeds into `agent.json` only via `finishConverged`. The convergence marker lives on the HOST and survives a guest rebuild; `agent.json` is rewritten by the installer. Unchanged descriptor → same hash → early return → the seed never runs into a config that no longer has it. **Red-proved**: moving only the marker aside seeded it instantly (`grep -c escrow`: 0 → 1). A real blocker for re-escrow, which is exactly what a rebuilt box must do | **OPEN — NOT FIXED** |
| **R-223** | **The Day-0 manifest vouches agent 0.120.0 while the recovery feature needs 0.125.0 — and a reinstall DOWNGRADES a box that was fixed by hand.** Verbatim from the second reinstall: `agent (existing): felhom-agent 0.125.0``manifest: agent v0.120.0``installed /usr/local/bin/felhom-agent (felhom-agent 0.120.0)`. So every rebuild re-breaks the recovery path — the one event that makes the feature necessary | **WAITING-ON-OPERATOR** — vouch agent 0.125.0 in Day-0 artifacts. One dropdown; it has run on both demo boxes since 2026-08-04 and through the whole campaign |
**Untouched by this session, stated so nothing is presumed closed by association:** **R-213** (putting
files back — the half the recovery screen deliberately does not do) and **R-202** (the orphan card's
unconditional promise, which CAMPAIGN-11 §7 step 7 measured the customer-facing cost of).
| ID | What | State | Blocked on | Next action | Owner |
|---|---|---|---|---|---|
| **R-213** | **Putting files back in place — the half the recovery screen deliberately does not do.** The screen (R-193, controller v0.200.0) unlocks the repository and LISTS what is in it; restoring is per-app and lives in the backups area, and the operator ruled the two separate on 2026-08-05: *a screen that unlocks and then offers to overwrite is two decisions wearing one button*. What is missing is the step after the listing — a customer who can now SEE their files still has to work out, per app, which restore to choose. **The operator named its requirement: a live-versus-backup comparison** — the customer must be able to see what would change before anything is overwritten | **OPEN — not started, deliberately** | the comparison design (nothing exists for it yet) | Design the live-vs-backup comparison, then the put-back flow on top of it. Do NOT fold it into the recovery screen | Operator + CC |
+2 -1
View File
@@ -178,7 +178,7 @@ turns a true alarm into one the operator dismisses.
### A comment asserting an invariant needs a test pinning it, or it is a wish
**Seven instances in this project have shipped guarantees the code did not provide** — each survived
**Eight instances in this project have shipped guarantees the code did not provide** — each survived
review because the comment read as settled:
| # | Comment | What it claimed | What the code did |
@@ -190,6 +190,7 @@ review because the comment read as settled:
| 5 | `inflight.go` | *"a caller that cannot acquire DEFERS"* | the backup caller recorded a failure and paged the operator (F-A1) |
| 6 | `quiesce.go` | the agent's 409 *prevents* "a spurious failure" | on the start path it produced one (F-A1) |
| 7 | `recovery_unit.go` B2 refusal (R-181) | *"the previous unit is untouched and NOTHING was deleted"* | *nothing deleted* held; **untouched was measured false** — the floor was checked ONLY in `captureAllRecoveryUnits`, while the two dump legs wrote the bulk into the same tree first and unguarded, so a 182,272 B tar became 2,147,666,432 B under a manifest that had not moved |
| 8 | `ResolveManagedFloor` (R-216) | *"never push a controller past the agent it depends on"* | it compared the box's agent against the **golden's** MinAgent while serving a **floor** that could point elsewhere. Raise a floor above the vouched golden — which the day-0 runbook recommends and a per-customer override makes trivial — and the guard checks a version it is not serving. Measured live 2026-08-05: golden 0.192.0/MinAgent 0.113.0, floor 0.200.0, agent 0.120.0 → served, and the box landed on a controller needing 0.125.0. Its customer was then told their correct recovery code was wrong. **The first entry in this table where the false invariant was a GUARD, not a comment alone.** Fixed hub v0.97.0: a floor above the vouched golden is HELD, with its own reason |
Three of these (4, 5/6 and 7) were found **on live hardware**, not by review or unit tests — #4 had a
green, red-proofed test suite over a production path that was broken two independent ways, and #7
File diff suppressed because it is too large Load Diff