R-201 re-walk: the data PASSES again, the journey still FAILS — two dead ends, down from four
gates / gates (push) Successful in 9s

Asked Campaign 11 Phase 1's question a second time, on the fixed build, on a
NEW appliance (VM 322, customer rewalk). The Campaign 11 venue was untouched.

THE DATA: PASS. All three sentinels byte-identical out of the pre-destruction
snapshot a7bc23bd in 23s through the customer's own restore flow — including a
12 MB binary and an accented Hungarian filename whose NAME BYTES are identical
too (verified as hex, not as rendered text).

THE JOURNEY: FAIL, two dead ends against Phase 1's four.
 1. R-218's CONSUME half. The hub re-staged the credential at 11:44:57 saying
    'the box re-consumes on its next cycle'; a full cycle ran at 11:55:46/54
    (with a positive control that it ran) and it did not. A census of the
    customer-reachable actions found none that fetches it. Only a command line
    INSIDE THE GUEST moved it — 18s, confirming nothing was wrong with the
    credential, target or key: only the trigger. R-218's row said SHIPPED and
    over-claimed; it is corrected to REOPENED for the consume half.
 2. R-220. Drives still unenrollable after a rebuild, needing a Proxmox-host
    unmount; without it no app redeploys and the restore page stays empty.

Unaided RTO STILL UNDEFINED. Attended: +45s key placed, +24m12s tier up,
+30m13s data verified. The 30m must not be quoted as the customer number.

What passed and is new: the recovery screen appeared WITHOUT being sought,
answered all three questions with a seal date matching the hub exactly, the
emailed reset code worked first try, the unlock was a real 1.528s unseal, and
R-225's fix was seen working in the wild (unknown, not a false zero).

R-216 part 4 reproduced live: the reinstall downgraded the hand-installed agent
0.126.0 -> 0.125.0.

DELIVERY GAP recorded as owed and NOT conflated with the journey: a fresh
install landed on controller 0.201.0 / agent 0.125.0 — the vouched versions,
neither carrying the fixes — installed by hand. Nothing was vouched.

Capability map row STAYS FAIL. Campaign 11 doc gets a dated ADDENDUM, not a
rewrite.
This commit is contained in:
2026-08-06 12:18:29 +02:00
parent a1a542b9a7
commit 0c4411e54b
4 changed files with 223 additions and 26 deletions
+37 -25
View File
@@ -44,35 +44,47 @@ code was wrong. *(CAMPAIGN 11)*
- **The off-site copy can be erased by the machine that made it.** A daily snapshot is armed as a
stopgap. *(R-95, R-87)*
## What last night's stress test found — and what we fixed this morning
## Can a household get their data back on their own? Asked again today — still no, but nearer
We spent the night trying to break the recovery journey, then left the machine alone and watched it
run. **Nothing we did lost a byte.** When the customer chose "I do not want the old data", the old
backups were **set aside and not deleted** — we checked the far end of the wire and the 12.5 MB was
still there, to the byte. A wrong code was refused three times with nothing written and no lockout.
The machine's alarm fired when we switched it off and cleared itself when it came back. Overnight it
ran a full cycle on its own and made a fresh off-site copy without being asked.
We built a **brand-new machine** from the published disc, gave it three marked files, destroyed it
guest and both drives, as a hardware loss would — and tried to get them back the way a household
would. *(R-201, the re-walk)*
**What it found: the machine still blamed the customer for failures that were not theirs.** Pull the
plug on our own central system and the customer was told their recovery code was bad — in three
hundredths of a second, when actually checking a code takes about one. The machine had not even
tried. **All of that is fixed and deployed** *(R-224, R-226, R-225, R-227, R-228)*:
**The files came back perfectly.** All three, byte for byte, including a 12 MB file and one whose
Hungarian accented filename came back **letter-for-letter identical**. Out of the pre-destruction
backup, in **23 seconds**, through the customer's own restore screen.
- **When something on our side is down, we say so** — and we say plainly that the code was **not**
used, so it is still good. Proven on the real machine: with our hub unreachable the answer changed
from "your code is wrong" to "we could not reach the central system".
- **A customer who mistypes is told to check their typing again.** That message had become
unreachable on any machine that had been given a new code — exactly the machine that just recovered.
- **When we do not know why something failed, we say that**, and never guess the customer.
- **"0 snapshots · 0 GB" is gone** where the truth is "we have not read it yet".
- **The set-aside backups are visible again** — the machine says they are kept and not deleted, and
does **not** pretend they can be reopened, because today they cannot be.
**And much of the journey now works.** The machine showed the recovery screen **without being asked**,
told the customer what was waiting and when it was sealed, said plainly that nobody can replace a lost
code, and accepted the real code first time. The emailed claim code worked first try.
**Still open, and worth knowing:** a rebuilt machine still cannot re-attach its own drives without us
*(R-220 — we are working around it by hand on the test machine right now)*, cannot create a new
recovery code *(R-221)*, and the screen at the machine still shows a stale pairing code *(R-214)*.
**The recovery journey is still recorded as FAILED** — these are fixes, not a re-walk, and it stays
failed until someone walks it end to end with no help from us.
**But it still needed us twice**, and a household has neither hand:
- **The machine never picks up its own storage connection.** Our hub hands it over and says "the box
will collect this on its next cycle" — the cycle came and went and it did not. Nothing the customer
can click fixes it; it took a command inside the machine. *(R-218 — we had recorded this as fixed;
only half of it was)*
- **A rebuilt machine still cannot re-attach its own drives** — and without them no app can be put
back, so the restore screen stays empty. *(R-220)*
**Two dead ends, down from four.** The verdict stays **FAILED** until a walk needs us zero times.
**One thing to decide.** A machine installed today still gets the older software — **the fixes are
built and published but not approved for new machines**. We installed them by hand for this test. So
this proves the journey works on the fixed build; it does **not** prove a customer would receive it.
## What we fixed this morning, and what it did not fix
Overnight we tried to break the recovery journey with eleven faults and then left the machine alone
for a full cycle. Nothing lost a byte; the set-aside backups really were kept; the alarm fired and
cleared itself. What it found was that **the machine blamed the customer for failures that were not
theirs** — our hub being unreachable came back as "your recovery code is wrong", in three hundredths
of a second, without the machine even trying. **That is fixed and deployed** *(R-224, R-226)*, along
with three smaller truths: "0 snapshots" where the answer is "we have not looked yet" *(R-225)*, a raw
English error mid-recovery *(R-227)*, and set-aside backups that had become invisible *(R-228)*.
**None of that shortened the journey**, which is why today's re-walk above still says FAILED — the two
remaining dead ends are different ones.
## What shipped recently