# RE-WALK of R-201 / CAMPAIGN-11 Phase 1 — journal Every observable in the order taken. **Attended, 2026-08-06.** Clocks: demo-hp and the appliance = CEST; the guest = UTC; DooPlex = CEST. --- ## Pre-flight — baselines re-read on arrival, and a drift | Component | Runbook says | **Actual on arrival** | |---|---|---| | `felhom-controller` | `a62bb3874b25` | **`7db42c5fec3b`** | | `felhom-agent` | `a2e914f683bd` | **`062a7027abff`** | | `felhom.eu` | `d30c2a51ed2a` | **`c21bcf84f709`** | | highest register | R-228 | **R-229** | **The drift is benign and was checked rather than assumed:** exactly one commit per repo, all of them **R-229, documentation-only** (a `CLAUDE.md` restructuring plus a gate). No product code, no version change — controller **v0.202.0** and agent **v0.126.0** stand. **The highest register ID is R-229, not R-228**, which is what matters for minting. **Also stale in the runbook, same class Campaigns 10 and 11 both caught:** it names installer **1.25.0**; the published artifact is **1.26.1** (since 2026-07-31). The local copy on demo-hp was verified byte-identical to the published one: ``` f3cc86d5f0ec68bba4155c994b4fa84e208d50209bb6e815636c99e5441059a6 felhom-installer-1.26.1-pve9.2-1.iso ``` ### §2 — credentials: discovered, not assumed Key names found in `~/.config/credentials` (**names only, values never read**): ``` HETZNER_API · PASSWORD · TS_KEY · HUB_PW · ISO_S3_CLIENT_AK · ISO_S3_CLIENT_SK · ISO_S3_URL R_DEMO-FELHOM · R_DEMO-HP ``` **Mapped: `HUB_PW` → the hub operator login**, verified live (`/hosts` and `/configuration` both 200) rather than assumed from the name. **Nothing else was needed from the file** — the appliance's root credential comes from the hub's own reveal endpoint, and the dashboard password is created during the claim and stored in `~/.config/rewalk/` (0600, DooPlex only). No key was guessed and none was carried from memory. --- ## §3.1 — what a fresh install ACTUALLY landed on **This is a result in its own right: it is what a customer receives today.** | | vouched in the Day-0 manifest | the box landed on | |---|---|---| | golden | **0.201.0** | — | | controller | (baked into the golden) | **0.201.0** | | agent | **0.125.0** | **0.125.0** | | `min_agent` | 0.125.0 | — | **Neither carries the fixes this re-walk exists to exercise** (controller v0.202.0, agent v0.126.0). Verbatim from the box's own day-0 log: ``` [OK] controller: Up 20 seconds (healthy) (after ~0s) [INFO] controller image: gitea.dooplex.hu/admin/felhom-controller:0.201.0 [OK] Day-0 provision SUCCESS — vmid=9201 host_id=rewalk-1ab77d customer=rewalk golden=local:backup/vzdump-lxc-9100-2026_08_06-10_44_36.tar.zst [INFO] root@pam was rotated + vaulted at step 4b ``` **§3.2 — brought to the fixed versions BY HAND**, and it is a hand step, not a delivery: ``` BEFORE felhom-agent 0.125.0 · controller 0.201.0 AFTER felhom-agent 0.126.0 · controller 0.202.0 (healthy) ``` The agent binary was verified against the published sha (`7ecf8e9cdba237bc…`) before installing. **§3.3 — THE DELIVERY GAP, recorded as owed.** Fleet delivery of these versions needs a golden carrying controller 0.202.0 **and** a vouched agent 0.126.0. **Nothing was vouched** — that is the operator's act. **This re-walk proves the JOURNEY on the fixed build; it does NOT prove that a real customer would receive that build, and the two must not be read as one.** --- ## Venue | | | |---|---| | Host | `demo-hp` (HP t740), Tier 0 | | VM | **322 `rewalk-appliance`** — q35/OVMF (`pre-enrolled-keys=0`), 4 cores, 8 GB, `cpu=host` | | Disks | `scsi0` 200 G · `scsi1` 50 G · `scsi2` 50 G, qcow2 on `c11-scratch` (dir at the **mount root** `/mnt/nvme-1tb`) | | Appliance | `rewalk.felhom.eu` @ **192.168.0.140/24**, gw/DNS 192.168.0.1 | | Guest | LXC **9201** @ **192.168.0.119** | | Hub customer | **`rewalk`** "Re-walk R-201", DR tier ON, off-site ON (shared, 50 GB) | | Host id | **`rewalk-1ab77d`** · appliance uuid `8feb5727-2992-4b9a-a919-071e73dddeb6` | | Off-site | Storage Box sub-account **284605**, user `u629488-sub5` | | **Untouched** | the **Campaign 11 venue (VM 321)**, `drill-r50` (VM 300), guest 9201 on both demo boxes, DooPlex, ep0 | **Storage naming, stated so teardown is unambiguous:** the VM's disks live on the existing `c11-scratch` storage (a `dir` at the mount root, which is what the agent's `exactMount` check requires). Teardown is by **VM id 322**, not by storage name. --- ## Phase A — the fixture ### A1 — installed from the published ISO, through the Terminal UI Driven blind (`qm monitor screendump` → PNG → read visually; `qm sendkey` for input). **All three of Campaign 11's traps reproduced and handled:** 1. **GRUB defaults to the graphical entry.** `down`+`ret` sent **inside one remote command** to hit the ~15 s window — the text installer came up first try. 2. **The keymap defaults to Hungarian while `sendkey` emits US scancodes.** Changed to **U.S. English** before any typing. **Positive control:** the administrator email was typed through the identical path and rendered **`rewalk@felhom.eu`** — `@` is `shift-2` on a US layout and `AltGr+V` on a Hungarian one, so a correct `@` proves the mapping for the 24 masked password characters that cannot be read back. 3. **`--boot` set in its own `qm set` after the disks existed**, and **verified from `qm config` before the first boot** (`boot: order=scsi0`, ISO detached). `Automatically reboot` was **unchecked** and confirmed `[ ]` with the focus moved away, so the reboot was deliberate. Summary screen, verbatim: `ext4` · `/dev/sda` · `Europe/Budapest` · **`U.S. English`** · `rewalk@felhom.eu` · `nic0` · `rewalk.felhom.eu` · `192.168.0.140/24` · `192.168.0.1` · `192.168.0.1`. **One reading corrected by a second instrument:** `192.168.0.140` answered a ping and looked like a collision. The **MAC** was `bc:24:11:d6:e3:93` — VM 322's own DHCP lease. Not a collision; a ping alone could not have told the difference. **Day-0 fired on first boot, unaided.** The console showed the Hungarian pairing banner with code **`WD6-BQG`**, and the hub's unclaimed table carried the same code, the same MAC and three SSH host keys within a minute. Bound through the real endpoint (`POST /appliances/21/bind`, HTTP 303). **Day-0 provision SUCCESS 10:47:04 — 3 m 36 s after the bind** (Campaign 11 took ~7 min). ### A2–A3 — fixed versions, then claimed Claimed through the real `/claim` form with a 24-character password (stored `0600` in `~/.config/rewalk/`). **The claim code came from the documented `--print-reset-code` escape hatch (R-204 item 1) — a guest command line, used deliberately as FIXTURE CONSTRUCTION.** Phase B's claim must not use it; that is the journey and it is measured. ### A4 — the app and the three sentinels `calibre-web` deployed through the real API with `HDD_PATH=/mnt/felhom-drives/adatok` (a real enrolled drive), healthy in 42 s. Both drives were enrolled through the customer endpoints and the backup target assigned to `mentes` — which reported `restart_required: true` and flipped only after the agent restart it asked for: ``` before: {"degraded":true, …"A rendszermentés jelenleg ugyanazon a lemezen van…"} after: {"degraded":false,"known":true,"label":"mentes","target":"felhom-backup"} ``` **THE THREE SENTINELS — and the accented one had to be written twice.** | # | file | bytes | sha256 | |---|---|---|---| | A | `REWALK-SENTINEL-A.txt` | 62 | `1573b0e1bad10c41c393ff690bfad0702d77ea0697f9cc7ef99403fd5bacc705` | | B | `REWALK-őrszem-ékezetes-árvíztűrő.txt` | 66 | `57676fcfb90f9695a84ddb6c9e656e7f9ff772fa20624d35d0adb34a4fe74430` | | C | `REWALK-SENTINEL-C-12MB.bin` | 12 582 912 | `c0faacd716cf92e8a6ef93f8295377b61566783dfabb8562599b53601d9aa14e` | > **HARNESS FAULT, caught by the one reading that cannot lie.** The first write of sentinel B produced > a filename full of `efbfbd` — **U+FFFD replacement characters**: the accents were destroyed by my own > `base64 → bash → pct exec` chain *before any backup happened*, which would have made the encoding > canary worthless while looking fine. **A Python `decode('utf-8')` check called it "valid UTF-8"**, > because U+FFFD *is* valid UTF-8; only the **hex dump of the filename bytes** exposed it. Rewritten > from explicit bytes inside Python on the guest, bypassing every shell layer: > `524557414c4b2d c591 72737a656d2d c3a9 …` = a genuine `ő`, `é`, `á`, `í`, `ű`, `ő`. ### A5 — the escrow ceremony Preflight **6 of 6 green** (`pbs_storage_id` · `dr_tier` · `age_binary` · `hub_upload` · `staged_secret` · `sudo_grant`). Ceremony through the customer wizard's own endpoints: ``` phase: done · restic_pw_sealed: TRUE · uploaded: true · entropy_bits: 129.24 key_fingerprint: d7:d3:4e:52:61:63:ac:e3:…:2a:ad:6f:ce · claimable: true ``` **`restic_pw_sealed: true` is the field the whole exercise rests on.** **R was claimed ONE-SHOT and streamed file→file** into `~/.config/rewalk/R_rewalk.txt` (0600, DooPlex only) **without touching any intermediate disk and without ever being rendered.** Its shape was verified without printing it: **ten words, 75 characters**. > **A tension in the runbook, resolved deliberately rather than silently.** §5.5 says the operator > records R "and where it lives"; §9.4 says R is "never a file on any machine". Campaign 11's > precedent — which this runbook cites approvingly — was a `0600` file on DooPlex that the operator > then moves into their own store. That is what was done, and it is flagged here rather than chosen > quietly. **The operator should move it into their own store and confirm.** ### A6 — the off-site backup, and the sentinels BY NAME `ok`, 55 s, **1 snapshot, 12 611 969 B**. **The gate is not the green tick** — `restic snapshots` + `ls -l latest`, run against the repository with its own credentials: ``` a7bc23bd 2026-08-06 09:15:22 rewalk [felhom-offbox, calibre-web] /mnt/felhom-drives/adatok/backups/primary/calibre-web /mnt/felhom-drives/adatok/userdata/media/books -rw-r--r-- 1000 1000 62 …/userdata/media/books/REWALK-SENTINEL-A.txt -rw-r--r-- 1000 1000 12582912 …/userdata/media/books/REWALK-SENTINEL-C-12MB.bin -rw-r--r-- 1000 1000 66 …/userdata/media/books/REWALK-őrszem-ékezetes-árvíztűrő.txt + the recovery unit: compose/{.felhom.yml,app.yaml,docker-compose.yml}, manifest.json ``` **All three sentinels are in the snapshot, by name, at the right sizes — and the accented filename survived into restic intact.** ### A7 — the pre-destruction truth **Box** (`settings.json`, secrets stripped): ``` offbox: enabled true · escrow_state "escrowed" · last_status "ok" · last_duration 55s last_run/last_success 2026-08-06T09:16:05Z · snapshot_count 1 repo_size_bytes 12 611 969 ("12.0 MB") · stats_known true · quota_gb 50 host u629488-sub5.your-storagebox.de · repo_path /home/felhom-repo hub_escrow_identity_present: true · claimed: true agent 0.126.0 · controller 0.202.0 (healthy) ``` **Hub** (SQLite snapshot taken **with its `-wal` and `-shm`**; `PRAGMA integrity_check` → `ok`; freshness by positive observable — newest `host_reports.received_at` `09:14:44` against `datetime('now')` `09:17:36`, **2 m 52 s old**): ``` host_escrow(rewalk-1ab77d): blob 383 B · identity_blob 572 B · stale_at NULL restic_pw_sha256 68182837607c93f4… · created 2026-08-06T09:14:10Z host_escrow_superseded: 0 rows for rewalk DR Recipe: present · Key Escrow: present ``` **Phase A gate: PASSED.** All seven records taken, sentinels listed **by name**. --- ## Phase B — the journey **The rule: no command line inside the guest, at any point.** After the destruction the only things that reached the guest were HTTP requests a browser could have made — plus the interventions counted below, which is exactly why they are counted. | # | step | result | |---|---|---| | 1 | **Destroy** — 11:21:12 | guest 9201 purged (both LVs), **both drives wiped to 4.0 K**. Host identity `rewalk-1ab77d` survived | | 2 | **Reinstall** | `felhom-host-install.sh` **v1.25.0** fetched live from `felhom.eu/scripts/`; **Day-0 provision SUCCESS 11:25:03**, 2 m 25 s | | 3 | **Fixed versions** | **hand step — and it was needed a SECOND time** (below) | | 4 | **Claim back** | the emailed reset code (generation 2) worked **first try**, accents and all | | 5 | **Log in** | **the recovery screen appeared without being sought**: `/` → `/launcher` → **`/recovery`** | | 6 | **Read the screen** | all three questions answered (below) | | 7 | **Enter the code** | HTTP 200 in **1.528 s** — a real unseal; key **recovered and placed** | | 8 | **The listing** | **did not render** — the tier was not up. Dead end 1 | | 9 | **Restore** | **all three sentinels byte-identical** | ### The reinstall DOWNGRADED the agent — R-216 part 4, live again ``` agent BEFORE the rebuild : 0.126.0 (hand-installed in Phase A) agent AFTER the rebuild : 0.125.0 (the vouched version) ``` **An operator who fixes a box by hand has it re-broken by the next rebuild — which is precisely the event that makes the recovery feature necessary.** Re-applied by hand, as §3.2 directs. ### Step 6 — the screen, read as a customer > „Ezt a gépet újratelepítették. A korábbi, **házon kívüli mentéseid megvannak** — a Felhom központi > rendszere őriz hozzájuk egy lezárt csomagot, amelyet **2026-08-06T09:14:10Z** zártunk le." > > „**A helyreállítási kódot senki nem tudja pótolni** — sem a Felhom, sem az ügyfélszolgálat, sem az > üzemeltető… Ha a kód elveszett, a korábbi mentések nem nyithatók meg többé." > > „Ebben a lépésben **semmit nem állítunk vissza és semmi nem változik**." All three of step 6's questions answered, and the seal date **matches the hub's `created_at` exactly**. The set-aside option was correctly **withheld**, with its reason stated rather than the button merely hidden. *(The seal date still renders as a raw RFC3339 string to a Hungarian household — the copy defect Phase 1 recorded, still unfixed.)* --- ## The dead ends — TWO, against Phase 1's four ### Dead end 1 — the off-site tier never came up on its own (R-218's SECOND half) **The declaration half works** — that part of R-218's fix is confirmed live: ``` 11:43:07 recovery: the offsite repository key was recovered and placed (outcome=installed) 11:43:07 recovery: the offsite tier could not be brought up yet: consume one-time password: no unconsumed offsite password (already consumed…) 11:44:57 (hub) offsiteheal: re-staged the stored one-time offsite secret for rewalk (declared needs_credential across 2 reports) — the box re-consumes on its next cycle ``` **The consume half does not.** The hub re-staged at 11:44:57 and said *"the box re-consumes on its next cycle"*. **The next cycle came and went** — `host-report from rewalk-1ab77d` at **11:55:46** and `Received report from rewalk` at **11:55:54**, a full cycle, **with a positive control that the cycle ran** — and the credential was still not consumed. Twenty-three minutes after the re-stage the box's last off-site-apply attempt was still **11:43:07**, before it. **What the customer sees meanwhile is honest but does not unblock them:** clicking the only relevant control returns „**A távoli mentési cél nincs beállítva**", and the page says „*Felhom offsite tárhely kiépítve — a beállítás automatikus, folyamatban. Ha egy napon belül nem áll be, jelezd az üzemeltetőnek.*" **A census of the customer-reachable actions on that page** — `config`, `reset`, `run`, `toggle` — **found none that fetches a staged credential.** **The lever, and its cost:** `systemctl restart felhom-controller-bootstrap.service` **inside the guest** — which breaks the journey's pass condition. It worked in **18 seconds** (Campaign 11 measured 17): ``` 12:06:16 restart 12:06:34 [offsite-apply] offsite configured for u629488-sub5@…:/home/felhom-repo ``` **Which confirms R-218 exactly: nothing was wrong with the credential, the target or the key — the only thing missing was anything at all to trigger a retry.** ### Dead end 2 — R-220, the drives, reproduced and red-proved `GET /api/disks/candidates` → `initialize: [], attach: []`, while both drives sat mounted at **both** `/mnt/felhom-drives/` **and** the raw `/mnt/` — the mount that enrolling them created. ``` before: initialize: [] attach: [] after : initialize: [/dev/sdb, /dev/sdc] (fstype ext4, data_bearing true) ``` Unmounting only the raw mounts flipped it. **That is a Proxmox-host action a customer cannot perform**, so it counts. Without it no app can be redeployed, and **without a redeployed app the restore page is empty** — „Nincs távoli mentésre jelölt alkalmazás" — which is R-213's territory and follows from this one rather than being separate. --- ## THE VERDICT — both halves, separately ### The data: **PASS** Restored in **23 seconds** out of snapshot **`a7bc23bd`** — the pre-destruction snapshot — through the customer's own two-step full-restore flow (size gate `12.8 MB`, then confirm), non-destructively. | # | file | bytes | expected = restored | |---|---|---|---| | A | `REWALK-SENTINEL-A.txt` | 62 | `1573b0e1…bacc705` **BYTE-IDENTICAL** | | B | `REWALK-őrszem-ékezetes-árvíztűrő.txt` | 66 | `57676fcf…4fe74430` **BYTE-IDENTICAL** | | C | `REWALK-SENTINEL-C-12MB.bin` | 12 582 912 | `c0faacd7…d9aa14e` **BYTE-IDENTICAL** | **And the accented filename's BYTES are byte-identical too** — verified as hex, not as rendered text: ``` expected 524557414c4b2dc59172737a656d2dc3a96b657a657465732dc3a17276c3ad7a74c5b172c5912e747874 restored 524557414c4b2dc59172737a656d2dc3a96b657a657465732dc3a17276c3ad7a74c5b172c5912e747874 ``` ### The journey: **FAIL** **Two steps needed a hand a customer does not have** — one inside the guest, one on the Proxmox host. **Better than Phase 1's four, and not zero.** ### The RTO | | | |---|---| | login (clock start) | **11:42:22** | | recovery code accepted, key placed | 11:43:07 (**+45 s**) | | off-site tier up — **after intervention 1** | 12:06:34 (**+24 m 12 s**) | | all three sentinels restored and verified — **after intervention 2** | 12:12:35 (**+30 m 13 s**) | **The unaided RTO remains UNDEFINED**, because the unaided journey still does not complete. **30 m 13 s is the attended figure** and must not be quoted as the customer number. The only segment that reflects the product working alone is the last one: **23 seconds to pull 12.8 MB back out of the off-site repository once everything was in place.** --- ## Harness faults, separated from the product's 1. **The accented sentinel's filename was destroyed at creation** by the `base64 → bash → pct exec` chain (U+FFFD), and **a Python `decode('utf-8')` check called it valid** — U+FFFD *is* valid UTF-8. Only a hex dump exposed it. Rewritten from explicit bytes. 2. **The same trap bit twice more**, in the verification script: a non-ASCII Python literal was mangled in transit and reported the accented sentinel as **MISSING**. Re-verified keyed on **hashes with no non-ASCII anywhere in the script**. **Three occurrences in one session: never put non-ASCII inside a script that crosses this chain.** 3. **`/api/storage/init` needs `fstype`** — omitting it failed with the honest „nem támogatott fájlrendszer" and I read the first failure as the product's. 4. **Wrong field names** on two endpoints (`app` not `stack`; `path` not `mount_name`), each caught by the endpoint's own refusal. 5. **A ping alone could not tell a collision from the box's own DHCP lease** — `192.168.0.140` answered and looked taken; the **MAC** showed it was VM 322 itself. ## Venue constraints, recorded so they do not inflate the dead-end count - The hub's host page shows the guest's LAN address as `—` **by design** (R-66), and this box is LAN-only with no Cloudflare tunnel, so the address was found from the host's ARP table. A real customer reaches `felhom.` through the tunnel. **Not a dead end.** - The claim code arrives **by email**, which is R-119's recorded single human step. The operator relayed it and it worked **first try**. **Not a dead end.** - The hub still read "Claimed, generation 1" after the rebuild, because the box's claim state lived in the destroyed guest; the reset-code path exists for exactly this and worked.