R-201 re-walk: the data PASSES again, the journey still FAILS — two dead ends, down from four
gates / gates (push) Successful in 9s

Asked Campaign 11 Phase 1's question a second time, on the fixed build, on a
NEW appliance (VM 322, customer rewalk). The Campaign 11 venue was untouched.

THE DATA: PASS. All three sentinels byte-identical out of the pre-destruction
snapshot a7bc23bd in 23s through the customer's own restore flow — including a
12 MB binary and an accented Hungarian filename whose NAME BYTES are identical
too (verified as hex, not as rendered text).

THE JOURNEY: FAIL, two dead ends against Phase 1's four.
 1. R-218's CONSUME half. The hub re-staged the credential at 11:44:57 saying
    'the box re-consumes on its next cycle'; a full cycle ran at 11:55:46/54
    (with a positive control that it ran) and it did not. A census of the
    customer-reachable actions found none that fetches it. Only a command line
    INSIDE THE GUEST moved it — 18s, confirming nothing was wrong with the
    credential, target or key: only the trigger. R-218's row said SHIPPED and
    over-claimed; it is corrected to REOPENED for the consume half.
 2. R-220. Drives still unenrollable after a rebuild, needing a Proxmox-host
    unmount; without it no app redeploys and the restore page stays empty.

Unaided RTO STILL UNDEFINED. Attended: +45s key placed, +24m12s tier up,
+30m13s data verified. The 30m must not be quoted as the customer number.

What passed and is new: the recovery screen appeared WITHOUT being sought,
answered all three questions with a seal date matching the hub exactly, the
emailed reset code worked first try, the unlock was a real 1.528s unseal, and
R-225's fix was seen working in the wild (unknown, not a false zero).

R-216 part 4 reproduced live: the reinstall downgraded the hand-installed agent
0.126.0 -> 0.125.0.

DELIVERY GAP recorded as owed and NOT conflated with the journey: a fresh
install landed on controller 0.201.0 / agent 0.125.0 — the vouched versions,
neither carrying the fixes — installed by hand. Nothing was vouched.

Capability map row STAYS FAIL. Campaign 11 doc gets a dated ADDENDUM, not a
rewrite.
This commit is contained in:
2026-08-06 12:18:29 +02:00
parent a1a542b9a7
commit 0c4411e54b
4 changed files with 223 additions and 26 deletions
File diff suppressed because one or more lines are too long
@@ -35,6 +35,22 @@ Evidence: `../tests/campaign11-evidence-2026-08-05/` — `journal.md` (Phases 0,
> which stays **FAIL** until a re-walk passes. **R-214, R-220, R-221 remain open**, and R-220 is still
> worked around by hand on this venue.
> **ADDENDUM 2026-08-06 — THE RE-WALK (R-201). The body below is NOT rewritten.**
>
> Phase 1's question was asked again on the fixed build, on a **new** appliance (VM 322, customer
> `rewalk`) — this campaign's venue was left untouched. **The data half PASSED again**: all three
> sentinels byte-identical, including an accented Hungarian filename whose **name bytes** are also
> identical, restored in **23 s**. **The journey half still FAILS, with TWO dead ends instead of
> four**: R-218's *consume* half (the hub re-stages, the box never collects — only a guest command
> line moves it) and R-220 (drives unenrollable after a rebuild). **The unaided RTO remains
> undefined.**
>
> Two of this campaign's findings were reproduced live: **R-216 part 4** (the reinstall downgraded the
> hand-installed agent back to the vouched version) and **R-220**. One of last night's fixes was seen
> working in the wild: **R-225** (an unread store said "unknown", not a false zero).
>
> Evidence: `../tests/rewalk-r201-2026-08-06/journal.md`.
## 1. Venue and baselines
| | |
@@ -234,3 +234,172 @@ DR Recipe: present · Key Escrow: present
```
**Phase A gate: PASSED.** All seven records taken, sentinels listed **by name**.
---
## Phase B — the journey
**The rule: no command line inside the guest, at any point.** After the destruction the only things
that reached the guest were HTTP requests a browser could have made — plus the interventions counted
below, which is exactly why they are counted.
| # | step | result |
|---|---|---|
| 1 | **Destroy** — 11:21:12 | guest 9201 purged (both LVs), **both drives wiped to 4.0 K**. Host identity `rewalk-1ab77d` survived |
| 2 | **Reinstall** | `felhom-host-install.sh` **v1.25.0** fetched live from `felhom.eu/scripts/`; **Day-0 provision SUCCESS 11:25:03**, 2 m 25 s |
| 3 | **Fixed versions** | **hand step — and it was needed a SECOND time** (below) |
| 4 | **Claim back** | the emailed reset code (generation 2) worked **first try**, accents and all |
| 5 | **Log in** | **the recovery screen appeared without being sought**: `/``/launcher`**`/recovery`** |
| 6 | **Read the screen** | all three questions answered (below) |
| 7 | **Enter the code** | HTTP 200 in **1.528 s** — a real unseal; key **recovered and placed** |
| 8 | **The listing** | **did not render** — the tier was not up. Dead end 1 |
| 9 | **Restore** | **all three sentinels byte-identical** |
### The reinstall DOWNGRADED the agent — R-216 part 4, live again
```
agent BEFORE the rebuild : 0.126.0 (hand-installed in Phase A)
agent AFTER the rebuild : 0.125.0 (the vouched version)
```
**An operator who fixes a box by hand has it re-broken by the next rebuild — which is precisely the
event that makes the recovery feature necessary.** Re-applied by hand, as §3.2 directs.
### Step 6 — the screen, read as a customer
> „Ezt a gépet újratelepítették. A korábbi, **házon kívüli mentéseid megvannak** — a Felhom központi
> rendszere őriz hozzájuk egy lezárt csomagot, amelyet **2026-08-06T09:14:10Z** zártunk le."
>
> „**A helyreállítási kódot senki nem tudja pótolni** — sem a Felhom, sem az ügyfélszolgálat, sem az
> üzemeltető… Ha a kód elveszett, a korábbi mentések nem nyithatók meg többé."
>
> „Ebben a lépésben **semmit nem állítunk vissza és semmi nem változik**."
All three of step 6's questions answered, and the seal date **matches the hub's `created_at` exactly**.
The set-aside option was correctly **withheld**, with its reason stated rather than the button merely
hidden. *(The seal date still renders as a raw RFC3339 string to a Hungarian household — the copy
defect Phase 1 recorded, still unfixed.)*
---
## The dead ends — TWO, against Phase 1's four
### Dead end 1 — the off-site tier never came up on its own (R-218's SECOND half)
**The declaration half works** — that part of R-218's fix is confirmed live:
```
11:43:07 recovery: the offsite repository key was recovered and placed (outcome=installed)
11:43:07 recovery: the offsite tier could not be brought up yet:
consume one-time password: no unconsumed offsite password (already consumed…)
11:44:57 (hub) offsiteheal: re-staged the stored one-time offsite secret for rewalk
(declared needs_credential across 2 reports) — the box re-consumes on its next cycle
```
**The consume half does not.** The hub re-staged at 11:44:57 and said *"the box re-consumes on its
next cycle"*. **The next cycle came and went**`host-report from rewalk-1ab77d` at **11:55:46** and
`Received report from rewalk` at **11:55:54**, a full cycle, **with a positive control that the cycle
ran** — and the credential was still not consumed. Twenty-three minutes after the re-stage the box's
last off-site-apply attempt was still **11:43:07**, before it.
**What the customer sees meanwhile is honest but does not unblock them:** clicking the only relevant
control returns „**A távoli mentési cél nincs beállítva**", and the page says „*Felhom offsite tárhely
kiépítve — a beállítás automatikus, folyamatban. Ha egy napon belül nem áll be, jelezd az
üzemeltetőnek.*" **A census of the customer-reachable actions on that page**`config`, `reset`,
`run`, `toggle` — **found none that fetches a staged credential.**
**The lever, and its cost:** `systemctl restart felhom-controller-bootstrap.service` **inside the
guest** — which breaks the journey's pass condition. It worked in **18 seconds**
(Campaign 11 measured 17):
```
12:06:16 restart
12:06:34 [offsite-apply] offsite configured for u629488-sub5@…:/home/felhom-repo
```
**Which confirms R-218 exactly: nothing was wrong with the credential, the target or the key — the
only thing missing was anything at all to trigger a retry.**
### Dead end 2 — R-220, the drives, reproduced and red-proved
`GET /api/disks/candidates``initialize: [], attach: []`, while both drives sat mounted at **both**
`/mnt/felhom-drives/<name>` **and** the raw `/mnt/<name>` — the mount that enrolling them created.
```
before: initialize: [] attach: []
after : initialize: [/dev/sdb, /dev/sdc] (fstype ext4, data_bearing true)
```
Unmounting only the raw mounts flipped it. **That is a Proxmox-host action a customer cannot perform**,
so it counts. Without it no app can be redeployed, and **without a redeployed app the restore page is
empty** — „Nincs távoli mentésre jelölt alkalmazás" — which is R-213's territory and follows from this
one rather than being separate.
---
## THE VERDICT — both halves, separately
### The data: **PASS**
Restored in **23 seconds** out of snapshot **`a7bc23bd`** — the pre-destruction snapshot — through the
customer's own two-step full-restore flow (size gate `12.8 MB`, then confirm), non-destructively.
| # | file | bytes | expected = restored |
|---|---|---|---|
| A | `REWALK-SENTINEL-A.txt` | 62 | `1573b0e1…bacc705` **BYTE-IDENTICAL** |
| B | `REWALK-őrszem-ékezetes-árvíztűrő.txt` | 66 | `57676fcf…4fe74430` **BYTE-IDENTICAL** |
| C | `REWALK-SENTINEL-C-12MB.bin` | 12 582 912 | `c0faacd7…d9aa14e` **BYTE-IDENTICAL** |
**And the accented filename's BYTES are byte-identical too** — verified as hex, not as rendered text:
```
expected 524557414c4b2dc59172737a656d2dc3a96b657a657465732dc3a17276c3ad7a74c5b172c5912e747874
restored 524557414c4b2dc59172737a656d2dc3a96b657a657465732dc3a17276c3ad7a74c5b172c5912e747874
```
### The journey: **FAIL**
**Two steps needed a hand a customer does not have** — one inside the guest, one on the Proxmox host.
**Better than Phase 1's four, and not zero.**
### The RTO
| | |
|---|---|
| login (clock start) | **11:42:22** |
| recovery code accepted, key placed | 11:43:07 (**+45 s**) |
| off-site tier up — **after intervention 1** | 12:06:34 (**+24 m 12 s**) |
| all three sentinels restored and verified — **after intervention 2** | 12:12:35 (**+30 m 13 s**) |
**The unaided RTO remains UNDEFINED**, because the unaided journey still does not complete. **30 m 13 s
is the attended figure** and must not be quoted as the customer number. The only segment that reflects
the product working alone is the last one: **23 seconds to pull 12.8 MB back out of the off-site
repository once everything was in place.**
---
## Harness faults, separated from the product's
1. **The accented sentinel's filename was destroyed at creation** by the `base64 → bash → pct exec`
chain (U+FFFD), and **a Python `decode('utf-8')` check called it valid** — U+FFFD *is* valid UTF-8.
Only a hex dump exposed it. Rewritten from explicit bytes.
2. **The same trap bit twice more**, in the verification script: a non-ASCII Python literal was mangled
in transit and reported the accented sentinel as **MISSING**. Re-verified keyed on **hashes with no
non-ASCII anywhere in the script**. **Three occurrences in one session: never put non-ASCII inside a
script that crosses this chain.**
3. **`/api/storage/init` needs `fstype`** — omitting it failed with the honest
„nem támogatott fájlrendszer" and I read the first failure as the product's.
4. **Wrong field names** on two endpoints (`app` not `stack`; `path` not `mount_name`), each caught by
the endpoint's own refusal.
5. **A ping alone could not tell a collision from the box's own DHCP lease**`192.168.0.140` answered
and looked taken; the **MAC** showed it was VM 322 itself.
## Venue constraints, recorded so they do not inflate the dead-end count
- The hub's host page shows the guest's LAN address as `—` **by design** (R-66), and this box is
LAN-only with no Cloudflare tunnel, so the address was found from the host's ARP table. A real
customer reaches `felhom.<domain>` through the tunnel. **Not a dead end.**
- The claim code arrives **by email**, which is R-119's recorded single human step. The operator
relayed it and it worked **first try**. **Not a dead end.**
- The hub still read "Claimed, generation 1" after the rebuild, because the box's claim state lived in
the destroyed guest; the reset-code path exists for exactly this and worked.