Files
felhom.eu/documentation/tests/rewalk-r201-2026-08-06/journal.md
T
admin 0c4411e54b
gates / gates (push) Successful in 9s
R-201 re-walk: the data PASSES again, the journey still FAILS — two dead ends, down from four
Asked Campaign 11 Phase 1's question a second time, on the fixed build, on a
NEW appliance (VM 322, customer rewalk). The Campaign 11 venue was untouched.

THE DATA: PASS. All three sentinels byte-identical out of the pre-destruction
snapshot a7bc23bd in 23s through the customer's own restore flow — including a
12 MB binary and an accented Hungarian filename whose NAME BYTES are identical
too (verified as hex, not as rendered text).

THE JOURNEY: FAIL, two dead ends against Phase 1's four.
 1. R-218's CONSUME half. The hub re-staged the credential at 11:44:57 saying
    'the box re-consumes on its next cycle'; a full cycle ran at 11:55:46/54
    (with a positive control that it ran) and it did not. A census of the
    customer-reachable actions found none that fetches it. Only a command line
    INSIDE THE GUEST moved it — 18s, confirming nothing was wrong with the
    credential, target or key: only the trigger. R-218's row said SHIPPED and
    over-claimed; it is corrected to REOPENED for the consume half.
 2. R-220. Drives still unenrollable after a rebuild, needing a Proxmox-host
    unmount; without it no app redeploys and the restore page stays empty.

Unaided RTO STILL UNDEFINED. Attended: +45s key placed, +24m12s tier up,
+30m13s data verified. The 30m must not be quoted as the customer number.

What passed and is new: the recovery screen appeared WITHOUT being sought,
answered all three questions with a seal date matching the hub exactly, the
emailed reset code worked first try, the unlock was a real 1.528s unseal, and
R-225's fix was seen working in the wild (unknown, not a false zero).

R-216 part 4 reproduced live: the reinstall downgraded the hand-installed agent
0.126.0 -> 0.125.0.

DELIVERY GAP recorded as owed and NOT conflated with the journey: a fresh
install landed on controller 0.201.0 / agent 0.125.0 — the vouched versions,
neither carrying the fixes — installed by hand. Nothing was vouched.

Capability map row STAYS FAIL. Campaign 11 doc gets a dated ADDENDUM, not a
rewrite.
2026-08-06 12:18:29 +02:00

20 KiB
Raw Blame History

RE-WALK of R-201 / CAMPAIGN-11 Phase 1 — journal

Every observable in the order taken. Attended, 2026-08-06. Clocks: demo-hp and the appliance = CEST; the guest = UTC; DooPlex = CEST.


Pre-flight — baselines re-read on arrival, and a drift

Component Runbook says Actual on arrival
felhom-controller a62bb3874b25 7db42c5fec3b
felhom-agent a2e914f683bd 062a7027abff
felhom.eu d30c2a51ed2a c21bcf84f709
highest register R-228 R-229

The drift is benign and was checked rather than assumed: exactly one commit per repo, all of them R-229, documentation-only (a CLAUDE.md restructuring plus a gate). No product code, no version change — controller v0.202.0 and agent v0.126.0 stand. The highest register ID is R-229, not R-228, which is what matters for minting.

Also stale in the runbook, same class Campaigns 10 and 11 both caught: it names installer 1.25.0; the published artifact is 1.26.1 (since 2026-07-31). The local copy on demo-hp was verified byte-identical to the published one:

f3cc86d5f0ec68bba4155c994b4fa84e208d50209bb6e815636c99e5441059a6  felhom-installer-1.26.1-pve9.2-1.iso

§2 — credentials: discovered, not assumed

Key names found in ~/.config/credentials (names only, values never read):

HETZNER_API · PASSWORD · TS_KEY · HUB_PW · ISO_S3_CLIENT_AK · ISO_S3_CLIENT_SK · ISO_S3_URL
R_DEMO-FELHOM · R_DEMO-HP

Mapped: HUB_PW → the hub operator login, verified live (/hosts and /configuration both 200) rather than assumed from the name. Nothing else was needed from the file — the appliance's root credential comes from the hub's own reveal endpoint, and the dashboard password is created during the claim and stored in ~/.config/rewalk/ (0600, DooPlex only). No key was guessed and none was carried from memory.


§3.1 — what a fresh install ACTUALLY landed on

This is a result in its own right: it is what a customer receives today.

vouched in the Day-0 manifest the box landed on
golden 0.201.0
controller (baked into the golden) 0.201.0
agent 0.125.0 0.125.0
min_agent 0.125.0

Neither carries the fixes this re-walk exists to exercise (controller v0.202.0, agent v0.126.0). Verbatim from the box's own day-0 log:

[OK]   controller: Up 20 seconds (healthy) (after ~0s)
[INFO]   controller image: gitea.dooplex.hu/admin/felhom-controller:0.201.0
[OK] Day-0 provision SUCCESS — vmid=9201 host_id=rewalk-1ab77d customer=rewalk
     golden=local:backup/vzdump-lxc-9100-2026_08_06-10_44_36.tar.zst
[INFO] root@pam was rotated + vaulted at step 4b

§3.2 — brought to the fixed versions BY HAND, and it is a hand step, not a delivery:

BEFORE  felhom-agent 0.125.0 · controller 0.201.0
AFTER   felhom-agent 0.126.0 · controller 0.202.0 (healthy)

The agent binary was verified against the published sha (7ecf8e9cdba237bc…) before installing.

§3.3 — THE DELIVERY GAP, recorded as owed. Fleet delivery of these versions needs a golden carrying controller 0.202.0 and a vouched agent 0.126.0. Nothing was vouched — that is the operator's act. This re-walk proves the JOURNEY on the fixed build; it does NOT prove that a real customer would receive that build, and the two must not be read as one.


Venue

Host demo-hp (HP t740), Tier 0
VM 322 rewalk-appliance — q35/OVMF (pre-enrolled-keys=0), 4 cores, 8 GB, cpu=host
Disks scsi0 200 G · scsi1 50 G · scsi2 50 G, qcow2 on c11-scratch (dir at the mount root /mnt/nvme-1tb)
Appliance rewalk.felhom.eu @ 192.168.0.140/24, gw/DNS 192.168.0.1
Guest LXC 9201 @ 192.168.0.119
Hub customer rewalk "Re-walk R-201", DR tier ON, off-site ON (shared, 50 GB)
Host id rewalk-1ab77d · appliance uuid 8feb5727-2992-4b9a-a919-071e73dddeb6
Off-site Storage Box sub-account 284605, user u629488-sub5
Untouched the Campaign 11 venue (VM 321), drill-r50 (VM 300), guest 9201 on both demo boxes, DooPlex, ep0

Storage naming, stated so teardown is unambiguous: the VM's disks live on the existing c11-scratch storage (a dir at the mount root, which is what the agent's exactMount check requires). Teardown is by VM id 322, not by storage name.


Phase A — the fixture

A1 — installed from the published ISO, through the Terminal UI

Driven blind (qm monitor screendump → PNG → read visually; qm sendkey for input). All three of Campaign 11's traps reproduced and handled:

  1. GRUB defaults to the graphical entry. down+ret sent inside one remote command to hit the ~15 s window — the text installer came up first try.
  2. The keymap defaults to Hungarian while sendkey emits US scancodes. Changed to U.S. English before any typing. Positive control: the administrator email was typed through the identical path and rendered rewalk@felhom.eu@ is shift-2 on a US layout and AltGr+V on a Hungarian one, so a correct @ proves the mapping for the 24 masked password characters that cannot be read back.
  3. --boot set in its own qm set after the disks existed, and verified from qm config before the first boot (boot: order=scsi0, ISO detached). Automatically reboot was unchecked and confirmed [ ] with the focus moved away, so the reboot was deliberate.

Summary screen, verbatim: ext4 · /dev/sda · Europe/Budapest · U.S. English · rewalk@felhom.eu · nic0 · rewalk.felhom.eu · 192.168.0.140/24 · 192.168.0.1 · 192.168.0.1.

One reading corrected by a second instrument: 192.168.0.140 answered a ping and looked like a collision. The MAC was bc:24:11:d6:e3:93 — VM 322's own DHCP lease. Not a collision; a ping alone could not have told the difference.

Day-0 fired on first boot, unaided. The console showed the Hungarian pairing banner with code WD6-BQG, and the hub's unclaimed table carried the same code, the same MAC and three SSH host keys within a minute. Bound through the real endpoint (POST /appliances/21/bind, HTTP 303).

Day-0 provision SUCCESS 10:47:04 — 3 m 36 s after the bind (Campaign 11 took ~7 min).

A2A3 — fixed versions, then claimed

Claimed through the real /claim form with a 24-character password (stored 0600 in ~/.config/rewalk/). The claim code came from the documented --print-reset-code escape hatch (R-204 item 1) — a guest command line, used deliberately as FIXTURE CONSTRUCTION. Phase B's claim must not use it; that is the journey and it is measured.

A4 — the app and the three sentinels

calibre-web deployed through the real API with HDD_PATH=/mnt/felhom-drives/adatok (a real enrolled drive), healthy in 42 s. Both drives were enrolled through the customer endpoints and the backup target assigned to mentes — which reported restart_required: true and flipped only after the agent restart it asked for:

before: {"degraded":true, …"A rendszermentés jelenleg ugyanazon a lemezen van…"}
after:  {"degraded":false,"known":true,"label":"mentes","target":"felhom-backup"}

THE THREE SENTINELS — and the accented one had to be written twice.

# file bytes sha256
A REWALK-SENTINEL-A.txt 62 1573b0e1bad10c41c393ff690bfad0702d77ea0697f9cc7ef99403fd5bacc705
B REWALK-őrszem-ékezetes-árvíztűrő.txt 66 57676fcfb90f9695a84ddb6c9e656e7f9ff772fa20624d35d0adb34a4fe74430
C REWALK-SENTINEL-C-12MB.bin 12 582 912 c0faacd716cf92e8a6ef93f8295377b61566783dfabb8562599b53601d9aa14e

HARNESS FAULT, caught by the one reading that cannot lie. The first write of sentinel B produced a filename full of efbfbdU+FFFD replacement characters: the accents were destroyed by my own base64 → bash → pct exec chain before any backup happened, which would have made the encoding canary worthless while looking fine. A Python decode('utf-8') check called it "valid UTF-8", because U+FFFD is valid UTF-8; only the hex dump of the filename bytes exposed it. Rewritten from explicit bytes inside Python on the guest, bypassing every shell layer: 524557414c4b2d c591 72737a656d2d c3a9 … = a genuine ő, é, á, í, ű, ő.

A5 — the escrow ceremony

Preflight 6 of 6 green (pbs_storage_id · dr_tier · age_binary · hub_upload · staged_secret · sudo_grant). Ceremony through the customer wizard's own endpoints:

phase: done · restic_pw_sealed: TRUE · uploaded: true · entropy_bits: 129.24
key_fingerprint: d7:d3:4e:52:61:63:ac:e3:…:2a:ad:6f:ce · claimable: true

restic_pw_sealed: true is the field the whole exercise rests on.

R was claimed ONE-SHOT and streamed file→file into ~/.config/rewalk/R_rewalk.txt (0600, DooPlex only) without touching any intermediate disk and without ever being rendered. Its shape was verified without printing it: ten words, 75 characters.

A tension in the runbook, resolved deliberately rather than silently. §5.5 says the operator records R "and where it lives"; §9.4 says R is "never a file on any machine". Campaign 11's precedent — which this runbook cites approvingly — was a 0600 file on DooPlex that the operator then moves into their own store. That is what was done, and it is flagged here rather than chosen quietly. The operator should move it into their own store and confirm.

A6 — the off-site backup, and the sentinels BY NAME

ok, 55 s, 1 snapshot, 12 611 969 B. The gate is not the green tickrestic snapshots + ls -l latest, run against the repository with its own credentials:

a7bc23bd  2026-08-06 09:15:22  rewalk  [felhom-offbox, calibre-web]
   /mnt/felhom-drives/adatok/backups/primary/calibre-web
   /mnt/felhom-drives/adatok/userdata/media/books

-rw-r--r-- 1000 1000       62  …/userdata/media/books/REWALK-SENTINEL-A.txt
-rw-r--r-- 1000 1000 12582912  …/userdata/media/books/REWALK-SENTINEL-C-12MB.bin
-rw-r--r-- 1000 1000       66  …/userdata/media/books/REWALK-őrszem-ékezetes-árvíztűrő.txt
+ the recovery unit: compose/{.felhom.yml,app.yaml,docker-compose.yml}, manifest.json

All three sentinels are in the snapshot, by name, at the right sizes — and the accented filename survived into restic intact.

A7 — the pre-destruction truth

Box (settings.json, secrets stripped):

offbox: enabled true · escrow_state "escrowed" · last_status "ok" · last_duration 55s
        last_run/last_success 2026-08-06T09:16:05Z · snapshot_count 1
        repo_size_bytes 12 611 969 ("12.0 MB") · stats_known true · quota_gb 50
        host u629488-sub5.your-storagebox.de · repo_path /home/felhom-repo
hub_escrow_identity_present: true · claimed: true
agent 0.126.0 · controller 0.202.0 (healthy)

Hub (SQLite snapshot taken with its -wal and -shm; PRAGMA integrity_checkok; freshness by positive observable — newest host_reports.received_at 09:14:44 against datetime('now') 09:17:36, 2 m 52 s old):

host_escrow(rewalk-1ab77d):  blob 383 B · identity_blob 572 B · stale_at NULL
                             restic_pw_sha256 68182837607c93f4… · created 2026-08-06T09:14:10Z
host_escrow_superseded:      0 rows for rewalk
DR Recipe: present · Key Escrow: present

Phase A gate: PASSED. All seven records taken, sentinels listed by name.


Phase B — the journey

The rule: no command line inside the guest, at any point. After the destruction the only things that reached the guest were HTTP requests a browser could have made — plus the interventions counted below, which is exactly why they are counted.

# step result
1 Destroy — 11:21:12 guest 9201 purged (both LVs), both drives wiped to 4.0 K. Host identity rewalk-1ab77d survived
2 Reinstall felhom-host-install.sh v1.25.0 fetched live from felhom.eu/scripts/; Day-0 provision SUCCESS 11:25:03, 2 m 25 s
3 Fixed versions hand step — and it was needed a SECOND time (below)
4 Claim back the emailed reset code (generation 2) worked first try, accents and all
5 Log in the recovery screen appeared without being sought: //launcher/recovery
6 Read the screen all three questions answered (below)
7 Enter the code HTTP 200 in 1.528 s — a real unseal; key recovered and placed
8 The listing did not render — the tier was not up. Dead end 1
9 Restore all three sentinels byte-identical

The reinstall DOWNGRADED the agent — R-216 part 4, live again

agent BEFORE the rebuild : 0.126.0   (hand-installed in Phase A)
agent AFTER  the rebuild : 0.125.0   (the vouched version)

An operator who fixes a box by hand has it re-broken by the next rebuild — which is precisely the event that makes the recovery feature necessary. Re-applied by hand, as §3.2 directs.

Step 6 — the screen, read as a customer

„Ezt a gépet újratelepítették. A korábbi, házon kívüli mentéseid megvannak — a Felhom központi rendszere őriz hozzájuk egy lezárt csomagot, amelyet 2026-08-06T09:14:10Z zártunk le."

A helyreállítási kódot senki nem tudja pótolni — sem a Felhom, sem az ügyfélszolgálat, sem az üzemeltető… Ha a kód elveszett, a korábbi mentések nem nyithatók meg többé."

„Ebben a lépésben semmit nem állítunk vissza és semmi nem változik."

All three of step 6's questions answered, and the seal date matches the hub's created_at exactly. The set-aside option was correctly withheld, with its reason stated rather than the button merely hidden. (The seal date still renders as a raw RFC3339 string to a Hungarian household — the copy defect Phase 1 recorded, still unfixed.)


The dead ends — TWO, against Phase 1's four

Dead end 1 — the off-site tier never came up on its own (R-218's SECOND half)

The declaration half works — that part of R-218's fix is confirmed live:

11:43:07  recovery: the offsite repository key was recovered and placed (outcome=installed)
11:43:07  recovery: the offsite tier could not be brought up yet:
          consume one-time password: no unconsumed offsite password (already consumed…)
11:44:57  (hub) offsiteheal: re-staged the stored one-time offsite secret for rewalk
          (declared needs_credential across 2 reports) — the box re-consumes on its next cycle

The consume half does not. The hub re-staged at 11:44:57 and said "the box re-consumes on its next cycle". The next cycle came and wenthost-report from rewalk-1ab77d at 11:55:46 and Received report from rewalk at 11:55:54, a full cycle, with a positive control that the cycle ran — and the credential was still not consumed. Twenty-three minutes after the re-stage the box's last off-site-apply attempt was still 11:43:07, before it.

What the customer sees meanwhile is honest but does not unblock them: clicking the only relevant control returns „A távoli mentési cél nincs beállítva", and the page says „Felhom offsite tárhely kiépítve — a beállítás automatikus, folyamatban. Ha egy napon belül nem áll be, jelezd az üzemeltetőnek." A census of the customer-reachable actions on that pageconfig, reset, run, togglefound none that fetches a staged credential.

The lever, and its cost: systemctl restart felhom-controller-bootstrap.service inside the guest — which breaks the journey's pass condition. It worked in 18 seconds (Campaign 11 measured 17):

12:06:16  restart
12:06:34  [offsite-apply] offsite configured for u629488-sub5@…:/home/felhom-repo

Which confirms R-218 exactly: nothing was wrong with the credential, the target or the key — the only thing missing was anything at all to trigger a retry.

Dead end 2 — R-220, the drives, reproduced and red-proved

GET /api/disks/candidatesinitialize: [], attach: [], while both drives sat mounted at both /mnt/felhom-drives/<name> and the raw /mnt/<name> — the mount that enrolling them created.

before:  initialize: []                    attach: []
after :  initialize: [/dev/sdb, /dev/sdc]  (fstype ext4, data_bearing true)

Unmounting only the raw mounts flipped it. That is a Proxmox-host action a customer cannot perform, so it counts. Without it no app can be redeployed, and without a redeployed app the restore page is empty — „Nincs távoli mentésre jelölt alkalmazás" — which is R-213's territory and follows from this one rather than being separate.


THE VERDICT — both halves, separately

The data: PASS

Restored in 23 seconds out of snapshot a7bc23bd — the pre-destruction snapshot — through the customer's own two-step full-restore flow (size gate 12.8 MB, then confirm), non-destructively.

# file bytes expected = restored
A REWALK-SENTINEL-A.txt 62 1573b0e1…bacc705 BYTE-IDENTICAL
B REWALK-őrszem-ékezetes-árvíztűrő.txt 66 57676fcf…4fe74430 BYTE-IDENTICAL
C REWALK-SENTINEL-C-12MB.bin 12 582 912 c0faacd7…d9aa14e BYTE-IDENTICAL

And the accented filename's BYTES are byte-identical too — verified as hex, not as rendered text:

expected 524557414c4b2dc59172737a656d2dc3a96b657a657465732dc3a17276c3ad7a74c5b172c5912e747874
restored 524557414c4b2dc59172737a656d2dc3a96b657a657465732dc3a17276c3ad7a74c5b172c5912e747874

The journey: FAIL

Two steps needed a hand a customer does not have — one inside the guest, one on the Proxmox host. Better than Phase 1's four, and not zero.

The RTO

login (clock start) 11:42:22
recovery code accepted, key placed 11:43:07 (+45 s)
off-site tier up — after intervention 1 12:06:34 (+24 m 12 s)
all three sentinels restored and verified — after intervention 2 12:12:35 (+30 m 13 s)

The unaided RTO remains UNDEFINED, because the unaided journey still does not complete. 30 m 13 s is the attended figure and must not be quoted as the customer number. The only segment that reflects the product working alone is the last one: 23 seconds to pull 12.8 MB back out of the off-site repository once everything was in place.


Harness faults, separated from the product's

  1. The accented sentinel's filename was destroyed at creation by the base64 → bash → pct exec chain (U+FFFD), and a Python decode('utf-8') check called it valid — U+FFFD is valid UTF-8. Only a hex dump exposed it. Rewritten from explicit bytes.
  2. The same trap bit twice more, in the verification script: a non-ASCII Python literal was mangled in transit and reported the accented sentinel as MISSING. Re-verified keyed on hashes with no non-ASCII anywhere in the script. Three occurrences in one session: never put non-ASCII inside a script that crosses this chain.
  3. /api/storage/init needs fstype — omitting it failed with the honest „nem támogatott fájlrendszer" and I read the first failure as the product's.
  4. Wrong field names on two endpoints (app not stack; path not mount_name), each caught by the endpoint's own refusal.
  5. A ping alone could not tell a collision from the box's own DHCP lease192.168.0.140 answered and looked taken; the MAC showed it was VM 322 itself.

Venue constraints, recorded so they do not inflate the dead-end count

  • The hub's host page shows the guest's LAN address as by design (R-66), and this box is LAN-only with no Cloudflare tunnel, so the address was found from the host's ARP table. A real customer reaches felhom.<domain> through the tunnel. Not a dead end.
  • The claim code arrives by email, which is R-119's recorded single human step. The operator relayed it and it worked first try. Not a dead end.
  • The hub still read "Claimed, generation 1" after the rebuild, because the box's claim state lived in the destroyed guest; the reset-code path exists for exactly this and worked.