R-201 re-walk: the data PASSES again, the journey still FAILS — two dead ends, down from four
gates / gates (push) Successful in 9s
gates / gates (push) Successful in 9s
Asked Campaign 11 Phase 1's question a second time, on the fixed build, on a
NEW appliance (VM 322, customer rewalk). The Campaign 11 venue was untouched.
THE DATA: PASS. All three sentinels byte-identical out of the pre-destruction
snapshot a7bc23bd in 23s through the customer's own restore flow — including a
12 MB binary and an accented Hungarian filename whose NAME BYTES are identical
too (verified as hex, not as rendered text).
THE JOURNEY: FAIL, two dead ends against Phase 1's four.
1. R-218's CONSUME half. The hub re-staged the credential at 11:44:57 saying
'the box re-consumes on its next cycle'; a full cycle ran at 11:55:46/54
(with a positive control that it ran) and it did not. A census of the
customer-reachable actions found none that fetches it. Only a command line
INSIDE THE GUEST moved it — 18s, confirming nothing was wrong with the
credential, target or key: only the trigger. R-218's row said SHIPPED and
over-claimed; it is corrected to REOPENED for the consume half.
2. R-220. Drives still unenrollable after a rebuild, needing a Proxmox-host
unmount; without it no app redeploys and the restore page stays empty.
Unaided RTO STILL UNDEFINED. Attended: +45s key placed, +24m12s tier up,
+30m13s data verified. The 30m must not be quoted as the customer number.
What passed and is new: the recovery screen appeared WITHOUT being sought,
answered all three questions with a seal date matching the hub exactly, the
emailed reset code worked first try, the unlock was a real 1.528s unseal, and
R-225's fix was seen working in the wild (unknown, not a false zero).
R-216 part 4 reproduced live: the reinstall downgraded the hand-installed agent
0.126.0 -> 0.125.0.
DELIVERY GAP recorded as owed and NOT conflated with the journey: a fresh
install landed on controller 0.201.0 / agent 0.125.0 — the vouched versions,
neither carrying the fixes — installed by hand. Nothing was vouched.
Capability map row STAYS FAIL. Campaign 11 doc gets a dated ADDENDUM, not a
rewrite.
This commit is contained in:
@@ -44,35 +44,47 @@ code was wrong. *(CAMPAIGN 11)*
|
|||||||
- **The off-site copy can be erased by the machine that made it.** A daily snapshot is armed as a
|
- **The off-site copy can be erased by the machine that made it.** A daily snapshot is armed as a
|
||||||
stopgap. *(R-95, R-87)*
|
stopgap. *(R-95, R-87)*
|
||||||
|
|
||||||
## What last night's stress test found — and what we fixed this morning
|
## Can a household get their data back on their own? Asked again today — still no, but nearer
|
||||||
|
|
||||||
We spent the night trying to break the recovery journey, then left the machine alone and watched it
|
We built a **brand-new machine** from the published disc, gave it three marked files, destroyed it —
|
||||||
run. **Nothing we did lost a byte.** When the customer chose "I do not want the old data", the old
|
guest and both drives, as a hardware loss would — and tried to get them back the way a household
|
||||||
backups were **set aside and not deleted** — we checked the far end of the wire and the 12.5 MB was
|
would. *(R-201, the re-walk)*
|
||||||
still there, to the byte. A wrong code was refused three times with nothing written and no lockout.
|
|
||||||
The machine's alarm fired when we switched it off and cleared itself when it came back. Overnight it
|
|
||||||
ran a full cycle on its own and made a fresh off-site copy without being asked.
|
|
||||||
|
|
||||||
**What it found: the machine still blamed the customer for failures that were not theirs.** Pull the
|
**The files came back perfectly.** All three, byte for byte, including a 12 MB file and one whose
|
||||||
plug on our own central system and the customer was told their recovery code was bad — in three
|
Hungarian accented filename came back **letter-for-letter identical**. Out of the pre-destruction
|
||||||
hundredths of a second, when actually checking a code takes about one. The machine had not even
|
backup, in **23 seconds**, through the customer's own restore screen.
|
||||||
tried. **All of that is fixed and deployed** *(R-224, R-226, R-225, R-227, R-228)*:
|
|
||||||
|
|
||||||
- **When something on our side is down, we say so** — and we say plainly that the code was **not**
|
**And much of the journey now works.** The machine showed the recovery screen **without being asked**,
|
||||||
used, so it is still good. Proven on the real machine: with our hub unreachable the answer changed
|
told the customer what was waiting and when it was sealed, said plainly that nobody can replace a lost
|
||||||
from "your code is wrong" to "we could not reach the central system".
|
code, and accepted the real code first time. The emailed claim code worked first try.
|
||||||
- **A customer who mistypes is told to check their typing again.** That message had become
|
|
||||||
unreachable on any machine that had been given a new code — exactly the machine that just recovered.
|
|
||||||
- **When we do not know why something failed, we say that**, and never guess the customer.
|
|
||||||
- **"0 snapshots · 0 GB" is gone** where the truth is "we have not read it yet".
|
|
||||||
- **The set-aside backups are visible again** — the machine says they are kept and not deleted, and
|
|
||||||
does **not** pretend they can be reopened, because today they cannot be.
|
|
||||||
|
|
||||||
**Still open, and worth knowing:** a rebuilt machine still cannot re-attach its own drives without us
|
**But it still needed us twice**, and a household has neither hand:
|
||||||
*(R-220 — we are working around it by hand on the test machine right now)*, cannot create a new
|
|
||||||
recovery code *(R-221)*, and the screen at the machine still shows a stale pairing code *(R-214)*.
|
- **The machine never picks up its own storage connection.** Our hub hands it over and says "the box
|
||||||
**The recovery journey is still recorded as FAILED** — these are fixes, not a re-walk, and it stays
|
will collect this on its next cycle" — the cycle came and went and it did not. Nothing the customer
|
||||||
failed until someone walks it end to end with no help from us.
|
can click fixes it; it took a command inside the machine. *(R-218 — we had recorded this as fixed;
|
||||||
|
only half of it was)*
|
||||||
|
- **A rebuilt machine still cannot re-attach its own drives** — and without them no app can be put
|
||||||
|
back, so the restore screen stays empty. *(R-220)*
|
||||||
|
|
||||||
|
**Two dead ends, down from four.** The verdict stays **FAILED** until a walk needs us zero times.
|
||||||
|
|
||||||
|
**One thing to decide.** A machine installed today still gets the older software — **the fixes are
|
||||||
|
built and published but not approved for new machines**. We installed them by hand for this test. So
|
||||||
|
this proves the journey works on the fixed build; it does **not** prove a customer would receive it.
|
||||||
|
|
||||||
|
## What we fixed this morning, and what it did not fix
|
||||||
|
|
||||||
|
Overnight we tried to break the recovery journey with eleven faults and then left the machine alone
|
||||||
|
for a full cycle. Nothing lost a byte; the set-aside backups really were kept; the alarm fired and
|
||||||
|
cleared itself. What it found was that **the machine blamed the customer for failures that were not
|
||||||
|
theirs** — our hub being unreachable came back as "your recovery code is wrong", in three hundredths
|
||||||
|
of a second, without the machine even trying. **That is fixed and deployed** *(R-224, R-226)*, along
|
||||||
|
with three smaller truths: "0 snapshots" where the answer is "we have not looked yet" *(R-225)*, a raw
|
||||||
|
English error mid-recovery *(R-227)*, and set-aside backups that had become invisible *(R-228)*.
|
||||||
|
|
||||||
|
**None of that shortened the journey**, which is why today's re-walk above still says FAILED — the two
|
||||||
|
remaining dead ends are different ones.
|
||||||
|
|
||||||
## What shipped recently
|
## What shipped recently
|
||||||
|
|
||||||
|
|||||||
File diff suppressed because one or more lines are too long
@@ -35,6 +35,22 @@ Evidence: `../tests/campaign11-evidence-2026-08-05/` — `journal.md` (Phases 0,
|
|||||||
> which stays **FAIL** until a re-walk passes. **R-214, R-220, R-221 remain open**, and R-220 is still
|
> which stays **FAIL** until a re-walk passes. **R-214, R-220, R-221 remain open**, and R-220 is still
|
||||||
> worked around by hand on this venue.
|
> worked around by hand on this venue.
|
||||||
|
|
||||||
|
> **ADDENDUM 2026-08-06 — THE RE-WALK (R-201). The body below is NOT rewritten.**
|
||||||
|
>
|
||||||
|
> Phase 1's question was asked again on the fixed build, on a **new** appliance (VM 322, customer
|
||||||
|
> `rewalk`) — this campaign's venue was left untouched. **The data half PASSED again**: all three
|
||||||
|
> sentinels byte-identical, including an accented Hungarian filename whose **name bytes** are also
|
||||||
|
> identical, restored in **23 s**. **The journey half still FAILS, with TWO dead ends instead of
|
||||||
|
> four**: R-218's *consume* half (the hub re-stages, the box never collects — only a guest command
|
||||||
|
> line moves it) and R-220 (drives unenrollable after a rebuild). **The unaided RTO remains
|
||||||
|
> undefined.**
|
||||||
|
>
|
||||||
|
> Two of this campaign's findings were reproduced live: **R-216 part 4** (the reinstall downgraded the
|
||||||
|
> hand-installed agent back to the vouched version) and **R-220**. One of last night's fixes was seen
|
||||||
|
> working in the wild: **R-225** (an unread store said "unknown", not a false zero).
|
||||||
|
>
|
||||||
|
> Evidence: `../tests/rewalk-r201-2026-08-06/journal.md`.
|
||||||
|
|
||||||
## 1. Venue and baselines
|
## 1. Venue and baselines
|
||||||
|
|
||||||
| | |
|
| | |
|
||||||
|
|||||||
@@ -234,3 +234,172 @@ DR Recipe: present · Key Escrow: present
|
|||||||
```
|
```
|
||||||
|
|
||||||
**Phase A gate: PASSED.** All seven records taken, sentinels listed **by name**.
|
**Phase A gate: PASSED.** All seven records taken, sentinels listed **by name**.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Phase B — the journey
|
||||||
|
|
||||||
|
**The rule: no command line inside the guest, at any point.** After the destruction the only things
|
||||||
|
that reached the guest were HTTP requests a browser could have made — plus the interventions counted
|
||||||
|
below, which is exactly why they are counted.
|
||||||
|
|
||||||
|
| # | step | result |
|
||||||
|
|---|---|---|
|
||||||
|
| 1 | **Destroy** — 11:21:12 | guest 9201 purged (both LVs), **both drives wiped to 4.0 K**. Host identity `rewalk-1ab77d` survived |
|
||||||
|
| 2 | **Reinstall** | `felhom-host-install.sh` **v1.25.0** fetched live from `felhom.eu/scripts/`; **Day-0 provision SUCCESS 11:25:03**, 2 m 25 s |
|
||||||
|
| 3 | **Fixed versions** | **hand step — and it was needed a SECOND time** (below) |
|
||||||
|
| 4 | **Claim back** | the emailed reset code (generation 2) worked **first try**, accents and all |
|
||||||
|
| 5 | **Log in** | **the recovery screen appeared without being sought**: `/` → `/launcher` → **`/recovery`** |
|
||||||
|
| 6 | **Read the screen** | all three questions answered (below) |
|
||||||
|
| 7 | **Enter the code** | HTTP 200 in **1.528 s** — a real unseal; key **recovered and placed** |
|
||||||
|
| 8 | **The listing** | **did not render** — the tier was not up. Dead end 1 |
|
||||||
|
| 9 | **Restore** | **all three sentinels byte-identical** |
|
||||||
|
|
||||||
|
### The reinstall DOWNGRADED the agent — R-216 part 4, live again
|
||||||
|
|
||||||
|
```
|
||||||
|
agent BEFORE the rebuild : 0.126.0 (hand-installed in Phase A)
|
||||||
|
agent AFTER the rebuild : 0.125.0 (the vouched version)
|
||||||
|
```
|
||||||
|
|
||||||
|
**An operator who fixes a box by hand has it re-broken by the next rebuild — which is precisely the
|
||||||
|
event that makes the recovery feature necessary.** Re-applied by hand, as §3.2 directs.
|
||||||
|
|
||||||
|
### Step 6 — the screen, read as a customer
|
||||||
|
|
||||||
|
> „Ezt a gépet újratelepítették. A korábbi, **házon kívüli mentéseid megvannak** — a Felhom központi
|
||||||
|
> rendszere őriz hozzájuk egy lezárt csomagot, amelyet **2026-08-06T09:14:10Z** zártunk le."
|
||||||
|
>
|
||||||
|
> „**A helyreállítási kódot senki nem tudja pótolni** — sem a Felhom, sem az ügyfélszolgálat, sem az
|
||||||
|
> üzemeltető… Ha a kód elveszett, a korábbi mentések nem nyithatók meg többé."
|
||||||
|
>
|
||||||
|
> „Ebben a lépésben **semmit nem állítunk vissza és semmi nem változik**."
|
||||||
|
|
||||||
|
All three of step 6's questions answered, and the seal date **matches the hub's `created_at` exactly**.
|
||||||
|
The set-aside option was correctly **withheld**, with its reason stated rather than the button merely
|
||||||
|
hidden. *(The seal date still renders as a raw RFC3339 string to a Hungarian household — the copy
|
||||||
|
defect Phase 1 recorded, still unfixed.)*
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## The dead ends — TWO, against Phase 1's four
|
||||||
|
|
||||||
|
### Dead end 1 — the off-site tier never came up on its own (R-218's SECOND half)
|
||||||
|
|
||||||
|
**The declaration half works** — that part of R-218's fix is confirmed live:
|
||||||
|
|
||||||
|
```
|
||||||
|
11:43:07 recovery: the offsite repository key was recovered and placed (outcome=installed)
|
||||||
|
11:43:07 recovery: the offsite tier could not be brought up yet:
|
||||||
|
consume one-time password: no unconsumed offsite password (already consumed…)
|
||||||
|
11:44:57 (hub) offsiteheal: re-staged the stored one-time offsite secret for rewalk
|
||||||
|
(declared needs_credential across 2 reports) — the box re-consumes on its next cycle
|
||||||
|
```
|
||||||
|
|
||||||
|
**The consume half does not.** The hub re-staged at 11:44:57 and said *"the box re-consumes on its
|
||||||
|
next cycle"*. **The next cycle came and went** — `host-report from rewalk-1ab77d` at **11:55:46** and
|
||||||
|
`Received report from rewalk` at **11:55:54**, a full cycle, **with a positive control that the cycle
|
||||||
|
ran** — and the credential was still not consumed. Twenty-three minutes after the re-stage the box's
|
||||||
|
last off-site-apply attempt was still **11:43:07**, before it.
|
||||||
|
|
||||||
|
**What the customer sees meanwhile is honest but does not unblock them:** clicking the only relevant
|
||||||
|
control returns „**A távoli mentési cél nincs beállítva**", and the page says „*Felhom offsite tárhely
|
||||||
|
kiépítve — a beállítás automatikus, folyamatban. Ha egy napon belül nem áll be, jelezd az
|
||||||
|
üzemeltetőnek.*" **A census of the customer-reachable actions on that page** — `config`, `reset`,
|
||||||
|
`run`, `toggle` — **found none that fetches a staged credential.**
|
||||||
|
|
||||||
|
**The lever, and its cost:** `systemctl restart felhom-controller-bootstrap.service` **inside the
|
||||||
|
guest** — which breaks the journey's pass condition. It worked in **18 seconds**
|
||||||
|
(Campaign 11 measured 17):
|
||||||
|
|
||||||
|
```
|
||||||
|
12:06:16 restart
|
||||||
|
12:06:34 [offsite-apply] offsite configured for u629488-sub5@…:/home/felhom-repo
|
||||||
|
```
|
||||||
|
|
||||||
|
**Which confirms R-218 exactly: nothing was wrong with the credential, the target or the key — the
|
||||||
|
only thing missing was anything at all to trigger a retry.**
|
||||||
|
|
||||||
|
### Dead end 2 — R-220, the drives, reproduced and red-proved
|
||||||
|
|
||||||
|
`GET /api/disks/candidates` → `initialize: [], attach: []`, while both drives sat mounted at **both**
|
||||||
|
`/mnt/felhom-drives/<name>` **and** the raw `/mnt/<name>` — the mount that enrolling them created.
|
||||||
|
|
||||||
|
```
|
||||||
|
before: initialize: [] attach: []
|
||||||
|
after : initialize: [/dev/sdb, /dev/sdc] (fstype ext4, data_bearing true)
|
||||||
|
```
|
||||||
|
|
||||||
|
Unmounting only the raw mounts flipped it. **That is a Proxmox-host action a customer cannot perform**,
|
||||||
|
so it counts. Without it no app can be redeployed, and **without a redeployed app the restore page is
|
||||||
|
empty** — „Nincs távoli mentésre jelölt alkalmazás" — which is R-213's territory and follows from this
|
||||||
|
one rather than being separate.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## THE VERDICT — both halves, separately
|
||||||
|
|
||||||
|
### The data: **PASS**
|
||||||
|
|
||||||
|
Restored in **23 seconds** out of snapshot **`a7bc23bd`** — the pre-destruction snapshot — through the
|
||||||
|
customer's own two-step full-restore flow (size gate `12.8 MB`, then confirm), non-destructively.
|
||||||
|
|
||||||
|
| # | file | bytes | expected = restored |
|
||||||
|
|---|---|---|---|
|
||||||
|
| A | `REWALK-SENTINEL-A.txt` | 62 | `1573b0e1…bacc705` **BYTE-IDENTICAL** |
|
||||||
|
| B | `REWALK-őrszem-ékezetes-árvíztűrő.txt` | 66 | `57676fcf…4fe74430` **BYTE-IDENTICAL** |
|
||||||
|
| C | `REWALK-SENTINEL-C-12MB.bin` | 12 582 912 | `c0faacd7…d9aa14e` **BYTE-IDENTICAL** |
|
||||||
|
|
||||||
|
**And the accented filename's BYTES are byte-identical too** — verified as hex, not as rendered text:
|
||||||
|
|
||||||
|
```
|
||||||
|
expected 524557414c4b2dc59172737a656d2dc3a96b657a657465732dc3a17276c3ad7a74c5b172c5912e747874
|
||||||
|
restored 524557414c4b2dc59172737a656d2dc3a96b657a657465732dc3a17276c3ad7a74c5b172c5912e747874
|
||||||
|
```
|
||||||
|
|
||||||
|
### The journey: **FAIL**
|
||||||
|
|
||||||
|
**Two steps needed a hand a customer does not have** — one inside the guest, one on the Proxmox host.
|
||||||
|
**Better than Phase 1's four, and not zero.**
|
||||||
|
|
||||||
|
### The RTO
|
||||||
|
|
||||||
|
| | |
|
||||||
|
|---|---|
|
||||||
|
| login (clock start) | **11:42:22** |
|
||||||
|
| recovery code accepted, key placed | 11:43:07 (**+45 s**) |
|
||||||
|
| off-site tier up — **after intervention 1** | 12:06:34 (**+24 m 12 s**) |
|
||||||
|
| all three sentinels restored and verified — **after intervention 2** | 12:12:35 (**+30 m 13 s**) |
|
||||||
|
|
||||||
|
**The unaided RTO remains UNDEFINED**, because the unaided journey still does not complete. **30 m 13 s
|
||||||
|
is the attended figure** and must not be quoted as the customer number. The only segment that reflects
|
||||||
|
the product working alone is the last one: **23 seconds to pull 12.8 MB back out of the off-site
|
||||||
|
repository once everything was in place.**
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Harness faults, separated from the product's
|
||||||
|
|
||||||
|
1. **The accented sentinel's filename was destroyed at creation** by the `base64 → bash → pct exec`
|
||||||
|
chain (U+FFFD), and **a Python `decode('utf-8')` check called it valid** — U+FFFD *is* valid UTF-8.
|
||||||
|
Only a hex dump exposed it. Rewritten from explicit bytes.
|
||||||
|
2. **The same trap bit twice more**, in the verification script: a non-ASCII Python literal was mangled
|
||||||
|
in transit and reported the accented sentinel as **MISSING**. Re-verified keyed on **hashes with no
|
||||||
|
non-ASCII anywhere in the script**. **Three occurrences in one session: never put non-ASCII inside a
|
||||||
|
script that crosses this chain.**
|
||||||
|
3. **`/api/storage/init` needs `fstype`** — omitting it failed with the honest
|
||||||
|
„nem támogatott fájlrendszer" and I read the first failure as the product's.
|
||||||
|
4. **Wrong field names** on two endpoints (`app` not `stack`; `path` not `mount_name`), each caught by
|
||||||
|
the endpoint's own refusal.
|
||||||
|
5. **A ping alone could not tell a collision from the box's own DHCP lease** — `192.168.0.140` answered
|
||||||
|
and looked taken; the **MAC** showed it was VM 322 itself.
|
||||||
|
|
||||||
|
## Venue constraints, recorded so they do not inflate the dead-end count
|
||||||
|
|
||||||
|
- The hub's host page shows the guest's LAN address as `—` **by design** (R-66), and this box is
|
||||||
|
LAN-only with no Cloudflare tunnel, so the address was found from the host's ARP table. A real
|
||||||
|
customer reaches `felhom.<domain>` through the tunnel. **Not a dead end.**
|
||||||
|
- The claim code arrives **by email**, which is R-119's recorded single human step. The operator
|
||||||
|
relayed it and it worked **first try**. **Not a dead end.**
|
||||||
|
- The hub still read "Claimed, generation 1" after the rebuild, because the box's claim state lived in
|
||||||
|
the destroyed guest; the reset-code path exists for exactly this and worked.
|
||||||
|
|||||||
Reference in New Issue
Block a user