R-201 re-walk: the data PASSES again, the journey still FAILS — two dead ends, down from four
gates / gates (push) Successful in 9s
gates / gates (push) Successful in 9s
Asked Campaign 11 Phase 1's question a second time, on the fixed build, on a
NEW appliance (VM 322, customer rewalk). The Campaign 11 venue was untouched.
THE DATA: PASS. All three sentinels byte-identical out of the pre-destruction
snapshot a7bc23bd in 23s through the customer's own restore flow — including a
12 MB binary and an accented Hungarian filename whose NAME BYTES are identical
too (verified as hex, not as rendered text).
THE JOURNEY: FAIL, two dead ends against Phase 1's four.
1. R-218's CONSUME half. The hub re-staged the credential at 11:44:57 saying
'the box re-consumes on its next cycle'; a full cycle ran at 11:55:46/54
(with a positive control that it ran) and it did not. A census of the
customer-reachable actions found none that fetches it. Only a command line
INSIDE THE GUEST moved it — 18s, confirming nothing was wrong with the
credential, target or key: only the trigger. R-218's row said SHIPPED and
over-claimed; it is corrected to REOPENED for the consume half.
2. R-220. Drives still unenrollable after a rebuild, needing a Proxmox-host
unmount; without it no app redeploys and the restore page stays empty.
Unaided RTO STILL UNDEFINED. Attended: +45s key placed, +24m12s tier up,
+30m13s data verified. The 30m must not be quoted as the customer number.
What passed and is new: the recovery screen appeared WITHOUT being sought,
answered all three questions with a seal date matching the hub exactly, the
emailed reset code worked first try, the unlock was a real 1.528s unseal, and
R-225's fix was seen working in the wild (unknown, not a false zero).
R-216 part 4 reproduced live: the reinstall downgraded the hand-installed agent
0.126.0 -> 0.125.0.
DELIVERY GAP recorded as owed and NOT conflated with the journey: a fresh
install landed on controller 0.201.0 / agent 0.125.0 — the vouched versions,
neither carrying the fixes — installed by hand. Nothing was vouched.
Capability map row STAYS FAIL. Campaign 11 doc gets a dated ADDENDUM, not a
rewrite.
This commit is contained in:
@@ -44,35 +44,47 @@ code was wrong. *(CAMPAIGN 11)*
|
||||
- **The off-site copy can be erased by the machine that made it.** A daily snapshot is armed as a
|
||||
stopgap. *(R-95, R-87)*
|
||||
|
||||
## What last night's stress test found — and what we fixed this morning
|
||||
## Can a household get their data back on their own? Asked again today — still no, but nearer
|
||||
|
||||
We spent the night trying to break the recovery journey, then left the machine alone and watched it
|
||||
run. **Nothing we did lost a byte.** When the customer chose "I do not want the old data", the old
|
||||
backups were **set aside and not deleted** — we checked the far end of the wire and the 12.5 MB was
|
||||
still there, to the byte. A wrong code was refused three times with nothing written and no lockout.
|
||||
The machine's alarm fired when we switched it off and cleared itself when it came back. Overnight it
|
||||
ran a full cycle on its own and made a fresh off-site copy without being asked.
|
||||
We built a **brand-new machine** from the published disc, gave it three marked files, destroyed it —
|
||||
guest and both drives, as a hardware loss would — and tried to get them back the way a household
|
||||
would. *(R-201, the re-walk)*
|
||||
|
||||
**What it found: the machine still blamed the customer for failures that were not theirs.** Pull the
|
||||
plug on our own central system and the customer was told their recovery code was bad — in three
|
||||
hundredths of a second, when actually checking a code takes about one. The machine had not even
|
||||
tried. **All of that is fixed and deployed** *(R-224, R-226, R-225, R-227, R-228)*:
|
||||
**The files came back perfectly.** All three, byte for byte, including a 12 MB file and one whose
|
||||
Hungarian accented filename came back **letter-for-letter identical**. Out of the pre-destruction
|
||||
backup, in **23 seconds**, through the customer's own restore screen.
|
||||
|
||||
- **When something on our side is down, we say so** — and we say plainly that the code was **not**
|
||||
used, so it is still good. Proven on the real machine: with our hub unreachable the answer changed
|
||||
from "your code is wrong" to "we could not reach the central system".
|
||||
- **A customer who mistypes is told to check their typing again.** That message had become
|
||||
unreachable on any machine that had been given a new code — exactly the machine that just recovered.
|
||||
- **When we do not know why something failed, we say that**, and never guess the customer.
|
||||
- **"0 snapshots · 0 GB" is gone** where the truth is "we have not read it yet".
|
||||
- **The set-aside backups are visible again** — the machine says they are kept and not deleted, and
|
||||
does **not** pretend they can be reopened, because today they cannot be.
|
||||
**And much of the journey now works.** The machine showed the recovery screen **without being asked**,
|
||||
told the customer what was waiting and when it was sealed, said plainly that nobody can replace a lost
|
||||
code, and accepted the real code first time. The emailed claim code worked first try.
|
||||
|
||||
**Still open, and worth knowing:** a rebuilt machine still cannot re-attach its own drives without us
|
||||
*(R-220 — we are working around it by hand on the test machine right now)*, cannot create a new
|
||||
recovery code *(R-221)*, and the screen at the machine still shows a stale pairing code *(R-214)*.
|
||||
**The recovery journey is still recorded as FAILED** — these are fixes, not a re-walk, and it stays
|
||||
failed until someone walks it end to end with no help from us.
|
||||
**But it still needed us twice**, and a household has neither hand:
|
||||
|
||||
- **The machine never picks up its own storage connection.** Our hub hands it over and says "the box
|
||||
will collect this on its next cycle" — the cycle came and went and it did not. Nothing the customer
|
||||
can click fixes it; it took a command inside the machine. *(R-218 — we had recorded this as fixed;
|
||||
only half of it was)*
|
||||
- **A rebuilt machine still cannot re-attach its own drives** — and without them no app can be put
|
||||
back, so the restore screen stays empty. *(R-220)*
|
||||
|
||||
**Two dead ends, down from four.** The verdict stays **FAILED** until a walk needs us zero times.
|
||||
|
||||
**One thing to decide.** A machine installed today still gets the older software — **the fixes are
|
||||
built and published but not approved for new machines**. We installed them by hand for this test. So
|
||||
this proves the journey works on the fixed build; it does **not** prove a customer would receive it.
|
||||
|
||||
## What we fixed this morning, and what it did not fix
|
||||
|
||||
Overnight we tried to break the recovery journey with eleven faults and then left the machine alone
|
||||
for a full cycle. Nothing lost a byte; the set-aside backups really were kept; the alarm fired and
|
||||
cleared itself. What it found was that **the machine blamed the customer for failures that were not
|
||||
theirs** — our hub being unreachable came back as "your recovery code is wrong", in three hundredths
|
||||
of a second, without the machine even trying. **That is fixed and deployed** *(R-224, R-226)*, along
|
||||
with three smaller truths: "0 snapshots" where the answer is "we have not looked yet" *(R-225)*, a raw
|
||||
English error mid-recovery *(R-227)*, and set-aside backups that had become invisible *(R-228)*.
|
||||
|
||||
**None of that shortened the journey**, which is why today's re-walk above still says FAILED — the two
|
||||
remaining dead ends are different ones.
|
||||
|
||||
## What shipped recently
|
||||
|
||||
|
||||
Reference in New Issue
Block a user