R-201 re-walk: the data PASSES again, the journey still FAILS — two dead ends, down from four
gates / gates (push) Successful in 9s

Asked Campaign 11 Phase 1's question a second time, on the fixed build, on a
NEW appliance (VM 322, customer rewalk). The Campaign 11 venue was untouched.

THE DATA: PASS. All three sentinels byte-identical out of the pre-destruction
snapshot a7bc23bd in 23s through the customer's own restore flow — including a
12 MB binary and an accented Hungarian filename whose NAME BYTES are identical
too (verified as hex, not as rendered text).

THE JOURNEY: FAIL, two dead ends against Phase 1's four.
 1. R-218's CONSUME half. The hub re-staged the credential at 11:44:57 saying
    'the box re-consumes on its next cycle'; a full cycle ran at 11:55:46/54
    (with a positive control that it ran) and it did not. A census of the
    customer-reachable actions found none that fetches it. Only a command line
    INSIDE THE GUEST moved it — 18s, confirming nothing was wrong with the
    credential, target or key: only the trigger. R-218's row said SHIPPED and
    over-claimed; it is corrected to REOPENED for the consume half.
 2. R-220. Drives still unenrollable after a rebuild, needing a Proxmox-host
    unmount; without it no app redeploys and the restore page stays empty.

Unaided RTO STILL UNDEFINED. Attended: +45s key placed, +24m12s tier up,
+30m13s data verified. The 30m must not be quoted as the customer number.

What passed and is new: the recovery screen appeared WITHOUT being sought,
answered all three questions with a seal date matching the hub exactly, the
emailed reset code worked first try, the unlock was a real 1.528s unseal, and
R-225's fix was seen working in the wild (unknown, not a false zero).

R-216 part 4 reproduced live: the reinstall downgraded the hand-installed agent
0.126.0 -> 0.125.0.

DELIVERY GAP recorded as owed and NOT conflated with the journey: a fresh
install landed on controller 0.201.0 / agent 0.125.0 — the vouched versions,
neither carrying the fixes — installed by hand. Nothing was vouched.

Capability map row STAYS FAIL. Campaign 11 doc gets a dated ADDENDUM, not a
rewrite.
This commit is contained in:
2026-08-06 12:18:29 +02:00
parent a1a542b9a7
commit 0c4411e54b
4 changed files with 223 additions and 26 deletions
+37 -25
View File
@@ -44,35 +44,47 @@ code was wrong. *(CAMPAIGN 11)*
- **The off-site copy can be erased by the machine that made it.** A daily snapshot is armed as a - **The off-site copy can be erased by the machine that made it.** A daily snapshot is armed as a
stopgap. *(R-95, R-87)* stopgap. *(R-95, R-87)*
## What last night's stress test found — and what we fixed this morning ## Can a household get their data back on their own? Asked again today — still no, but nearer
We spent the night trying to break the recovery journey, then left the machine alone and watched it We built a **brand-new machine** from the published disc, gave it three marked files, destroyed it
run. **Nothing we did lost a byte.** When the customer chose "I do not want the old data", the old guest and both drives, as a hardware loss would — and tried to get them back the way a household
backups were **set aside and not deleted** — we checked the far end of the wire and the 12.5 MB was would. *(R-201, the re-walk)*
still there, to the byte. A wrong code was refused three times with nothing written and no lockout.
The machine's alarm fired when we switched it off and cleared itself when it came back. Overnight it
ran a full cycle on its own and made a fresh off-site copy without being asked.
**What it found: the machine still blamed the customer for failures that were not theirs.** Pull the **The files came back perfectly.** All three, byte for byte, including a 12 MB file and one whose
plug on our own central system and the customer was told their recovery code was bad — in three Hungarian accented filename came back **letter-for-letter identical**. Out of the pre-destruction
hundredths of a second, when actually checking a code takes about one. The machine had not even backup, in **23 seconds**, through the customer's own restore screen.
tried. **All of that is fixed and deployed** *(R-224, R-226, R-225, R-227, R-228)*:
- **When something on our side is down, we say so** — and we say plainly that the code was **not** **And much of the journey now works.** The machine showed the recovery screen **without being asked**,
used, so it is still good. Proven on the real machine: with our hub unreachable the answer changed told the customer what was waiting and when it was sealed, said plainly that nobody can replace a lost
from "your code is wrong" to "we could not reach the central system". code, and accepted the real code first time. The emailed claim code worked first try.
- **A customer who mistypes is told to check their typing again.** That message had become
unreachable on any machine that had been given a new code — exactly the machine that just recovered.
- **When we do not know why something failed, we say that**, and never guess the customer.
- **"0 snapshots · 0 GB" is gone** where the truth is "we have not read it yet".
- **The set-aside backups are visible again** — the machine says they are kept and not deleted, and
does **not** pretend they can be reopened, because today they cannot be.
**Still open, and worth knowing:** a rebuilt machine still cannot re-attach its own drives without us **But it still needed us twice**, and a household has neither hand:
*(R-220 — we are working around it by hand on the test machine right now)*, cannot create a new
recovery code *(R-221)*, and the screen at the machine still shows a stale pairing code *(R-214)*. - **The machine never picks up its own storage connection.** Our hub hands it over and says "the box
**The recovery journey is still recorded as FAILED** — these are fixes, not a re-walk, and it stays will collect this on its next cycle" — the cycle came and went and it did not. Nothing the customer
failed until someone walks it end to end with no help from us. can click fixes it; it took a command inside the machine. *(R-218 — we had recorded this as fixed;
only half of it was)*
- **A rebuilt machine still cannot re-attach its own drives** — and without them no app can be put
back, so the restore screen stays empty. *(R-220)*
**Two dead ends, down from four.** The verdict stays **FAILED** until a walk needs us zero times.
**One thing to decide.** A machine installed today still gets the older software — **the fixes are
built and published but not approved for new machines**. We installed them by hand for this test. So
this proves the journey works on the fixed build; it does **not** prove a customer would receive it.
## What we fixed this morning, and what it did not fix
Overnight we tried to break the recovery journey with eleven faults and then left the machine alone
for a full cycle. Nothing lost a byte; the set-aside backups really were kept; the alarm fired and
cleared itself. What it found was that **the machine blamed the customer for failures that were not
theirs** — our hub being unreachable came back as "your recovery code is wrong", in three hundredths
of a second, without the machine even trying. **That is fixed and deployed** *(R-224, R-226)*, along
with three smaller truths: "0 snapshots" where the answer is "we have not looked yet" *(R-225)*, a raw
English error mid-recovery *(R-227)*, and set-aside backups that had become invisible *(R-228)*.
**None of that shortened the journey**, which is why today's re-walk above still says FAILED — the two
remaining dead ends are different ones.
## What shipped recently ## What shipped recently
File diff suppressed because one or more lines are too long
@@ -35,6 +35,22 @@ Evidence: `../tests/campaign11-evidence-2026-08-05/` — `journal.md` (Phases 0,
> which stays **FAIL** until a re-walk passes. **R-214, R-220, R-221 remain open**, and R-220 is still > which stays **FAIL** until a re-walk passes. **R-214, R-220, R-221 remain open**, and R-220 is still
> worked around by hand on this venue. > worked around by hand on this venue.
> **ADDENDUM 2026-08-06 — THE RE-WALK (R-201). The body below is NOT rewritten.**
>
> Phase 1's question was asked again on the fixed build, on a **new** appliance (VM 322, customer
> `rewalk`) — this campaign's venue was left untouched. **The data half PASSED again**: all three
> sentinels byte-identical, including an accented Hungarian filename whose **name bytes** are also
> identical, restored in **23 s**. **The journey half still FAILS, with TWO dead ends instead of
> four**: R-218's *consume* half (the hub re-stages, the box never collects — only a guest command
> line moves it) and R-220 (drives unenrollable after a rebuild). **The unaided RTO remains
> undefined.**
>
> Two of this campaign's findings were reproduced live: **R-216 part 4** (the reinstall downgraded the
> hand-installed agent back to the vouched version) and **R-220**. One of last night's fixes was seen
> working in the wild: **R-225** (an unread store said "unknown", not a false zero).
>
> Evidence: `../tests/rewalk-r201-2026-08-06/journal.md`.
## 1. Venue and baselines ## 1. Venue and baselines
| | | | | |
@@ -234,3 +234,172 @@ DR Recipe: present · Key Escrow: present
``` ```
**Phase A gate: PASSED.** All seven records taken, sentinels listed **by name**. **Phase A gate: PASSED.** All seven records taken, sentinels listed **by name**.
---
## Phase B — the journey
**The rule: no command line inside the guest, at any point.** After the destruction the only things
that reached the guest were HTTP requests a browser could have made — plus the interventions counted
below, which is exactly why they are counted.
| # | step | result |
|---|---|---|
| 1 | **Destroy** — 11:21:12 | guest 9201 purged (both LVs), **both drives wiped to 4.0 K**. Host identity `rewalk-1ab77d` survived |
| 2 | **Reinstall** | `felhom-host-install.sh` **v1.25.0** fetched live from `felhom.eu/scripts/`; **Day-0 provision SUCCESS 11:25:03**, 2 m 25 s |
| 3 | **Fixed versions** | **hand step — and it was needed a SECOND time** (below) |
| 4 | **Claim back** | the emailed reset code (generation 2) worked **first try**, accents and all |
| 5 | **Log in** | **the recovery screen appeared without being sought**: `/``/launcher`**`/recovery`** |
| 6 | **Read the screen** | all three questions answered (below) |
| 7 | **Enter the code** | HTTP 200 in **1.528 s** — a real unseal; key **recovered and placed** |
| 8 | **The listing** | **did not render** — the tier was not up. Dead end 1 |
| 9 | **Restore** | **all three sentinels byte-identical** |
### The reinstall DOWNGRADED the agent — R-216 part 4, live again
```
agent BEFORE the rebuild : 0.126.0 (hand-installed in Phase A)
agent AFTER the rebuild : 0.125.0 (the vouched version)
```
**An operator who fixes a box by hand has it re-broken by the next rebuild — which is precisely the
event that makes the recovery feature necessary.** Re-applied by hand, as §3.2 directs.
### Step 6 — the screen, read as a customer
> „Ezt a gépet újratelepítették. A korábbi, **házon kívüli mentéseid megvannak** — a Felhom központi
> rendszere őriz hozzájuk egy lezárt csomagot, amelyet **2026-08-06T09:14:10Z** zártunk le."
>
> „**A helyreállítási kódot senki nem tudja pótolni** — sem a Felhom, sem az ügyfélszolgálat, sem az
> üzemeltető… Ha a kód elveszett, a korábbi mentések nem nyithatók meg többé."
>
> „Ebben a lépésben **semmit nem állítunk vissza és semmi nem változik**."
All three of step 6's questions answered, and the seal date **matches the hub's `created_at` exactly**.
The set-aside option was correctly **withheld**, with its reason stated rather than the button merely
hidden. *(The seal date still renders as a raw RFC3339 string to a Hungarian household — the copy
defect Phase 1 recorded, still unfixed.)*
---
## The dead ends — TWO, against Phase 1's four
### Dead end 1 — the off-site tier never came up on its own (R-218's SECOND half)
**The declaration half works** — that part of R-218's fix is confirmed live:
```
11:43:07 recovery: the offsite repository key was recovered and placed (outcome=installed)
11:43:07 recovery: the offsite tier could not be brought up yet:
consume one-time password: no unconsumed offsite password (already consumed…)
11:44:57 (hub) offsiteheal: re-staged the stored one-time offsite secret for rewalk
(declared needs_credential across 2 reports) — the box re-consumes on its next cycle
```
**The consume half does not.** The hub re-staged at 11:44:57 and said *"the box re-consumes on its
next cycle"*. **The next cycle came and went**`host-report from rewalk-1ab77d` at **11:55:46** and
`Received report from rewalk` at **11:55:54**, a full cycle, **with a positive control that the cycle
ran** — and the credential was still not consumed. Twenty-three minutes after the re-stage the box's
last off-site-apply attempt was still **11:43:07**, before it.
**What the customer sees meanwhile is honest but does not unblock them:** clicking the only relevant
control returns „**A távoli mentési cél nincs beállítva**", and the page says „*Felhom offsite tárhely
kiépítve — a beállítás automatikus, folyamatban. Ha egy napon belül nem áll be, jelezd az
üzemeltetőnek.*" **A census of the customer-reachable actions on that page**`config`, `reset`,
`run`, `toggle` — **found none that fetches a staged credential.**
**The lever, and its cost:** `systemctl restart felhom-controller-bootstrap.service` **inside the
guest** — which breaks the journey's pass condition. It worked in **18 seconds**
(Campaign 11 measured 17):
```
12:06:16 restart
12:06:34 [offsite-apply] offsite configured for u629488-sub5@…:/home/felhom-repo
```
**Which confirms R-218 exactly: nothing was wrong with the credential, the target or the key — the
only thing missing was anything at all to trigger a retry.**
### Dead end 2 — R-220, the drives, reproduced and red-proved
`GET /api/disks/candidates``initialize: [], attach: []`, while both drives sat mounted at **both**
`/mnt/felhom-drives/<name>` **and** the raw `/mnt/<name>` — the mount that enrolling them created.
```
before: initialize: [] attach: []
after : initialize: [/dev/sdb, /dev/sdc] (fstype ext4, data_bearing true)
```
Unmounting only the raw mounts flipped it. **That is a Proxmox-host action a customer cannot perform**,
so it counts. Without it no app can be redeployed, and **without a redeployed app the restore page is
empty** — „Nincs távoli mentésre jelölt alkalmazás" — which is R-213's territory and follows from this
one rather than being separate.
---
## THE VERDICT — both halves, separately
### The data: **PASS**
Restored in **23 seconds** out of snapshot **`a7bc23bd`** — the pre-destruction snapshot — through the
customer's own two-step full-restore flow (size gate `12.8 MB`, then confirm), non-destructively.
| # | file | bytes | expected = restored |
|---|---|---|---|
| A | `REWALK-SENTINEL-A.txt` | 62 | `1573b0e1…bacc705` **BYTE-IDENTICAL** |
| B | `REWALK-őrszem-ékezetes-árvíztűrő.txt` | 66 | `57676fcf…4fe74430` **BYTE-IDENTICAL** |
| C | `REWALK-SENTINEL-C-12MB.bin` | 12 582 912 | `c0faacd7…d9aa14e` **BYTE-IDENTICAL** |
**And the accented filename's BYTES are byte-identical too** — verified as hex, not as rendered text:
```
expected 524557414c4b2dc59172737a656d2dc3a96b657a657465732dc3a17276c3ad7a74c5b172c5912e747874
restored 524557414c4b2dc59172737a656d2dc3a96b657a657465732dc3a17276c3ad7a74c5b172c5912e747874
```
### The journey: **FAIL**
**Two steps needed a hand a customer does not have** — one inside the guest, one on the Proxmox host.
**Better than Phase 1's four, and not zero.**
### The RTO
| | |
|---|---|
| login (clock start) | **11:42:22** |
| recovery code accepted, key placed | 11:43:07 (**+45 s**) |
| off-site tier up — **after intervention 1** | 12:06:34 (**+24 m 12 s**) |
| all three sentinels restored and verified — **after intervention 2** | 12:12:35 (**+30 m 13 s**) |
**The unaided RTO remains UNDEFINED**, because the unaided journey still does not complete. **30 m 13 s
is the attended figure** and must not be quoted as the customer number. The only segment that reflects
the product working alone is the last one: **23 seconds to pull 12.8 MB back out of the off-site
repository once everything was in place.**
---
## Harness faults, separated from the product's
1. **The accented sentinel's filename was destroyed at creation** by the `base64 → bash → pct exec`
chain (U+FFFD), and **a Python `decode('utf-8')` check called it valid** — U+FFFD *is* valid UTF-8.
Only a hex dump exposed it. Rewritten from explicit bytes.
2. **The same trap bit twice more**, in the verification script: a non-ASCII Python literal was mangled
in transit and reported the accented sentinel as **MISSING**. Re-verified keyed on **hashes with no
non-ASCII anywhere in the script**. **Three occurrences in one session: never put non-ASCII inside a
script that crosses this chain.**
3. **`/api/storage/init` needs `fstype`** — omitting it failed with the honest
„nem támogatott fájlrendszer" and I read the first failure as the product's.
4. **Wrong field names** on two endpoints (`app` not `stack`; `path` not `mount_name`), each caught by
the endpoint's own refusal.
5. **A ping alone could not tell a collision from the box's own DHCP lease**`192.168.0.140` answered
and looked taken; the **MAC** showed it was VM 322 itself.
## Venue constraints, recorded so they do not inflate the dead-end count
- The hub's host page shows the guest's LAN address as `—` **by design** (R-66), and this box is
LAN-only with no Cloudflare tunnel, so the address was found from the host's ARP table. A real
customer reaches `felhom.<domain>` through the tunnel. **Not a dead end.**
- The claim code arrives **by email**, which is R-119's recorded single human step. The operator
relayed it and it worked **first try**. **Not a dead end.**
- The hub still read "Claimed, generation 1" after the rebuild, because the box's claim state lived in
the destroyed guest; the reset-code path exists for exactly this and worked.