Files
felhom.eu/documentation/tests/finalwalk-r201-2026-08-07/journal.md
T
admin 2228c0bff6
gates / gates (push) Successful in 15s
final walk COMPLETE — data PASS, journey FAIL; R-241 filed
THE DATA: PASS. All three sentinels byte-identical out of snapshot f5c53b03,
including the 12 MB binary and the accented Hungarian filename whose NAME BYTES
are identical too. Disk -> restic -> SFTP -> Storage Box -> rebuilt machine ->
disk, intact.

THE JOURNEY: FAIL, and further from the line than the previous walk. The claim
worked first try (302 in 0.164s). Then: / lands on the launcher with no recovery
pointer, /recovery 302s away, and the remote page offers to CREATE a new recovery
code — which would orphan the history the customer's code protects. There is no
field anywhere to enter the code they hold. The operator's documented remedy also
refuses, correctly and fail-closed. Recovery needed three guest command lines.

R-241 — and the cause is a success this same walk proved six hours earlier.
OffsiteRecoveryOffer() shows the screen only when (a) there is NO repository
password (pristine rebuild) or (b) one exists but the history will not open under
it. Overnight the credential self-heal collected the staged credential and applied
the tier, writing a FRESH key at 03:18Z — so (a) is false; and (b) is unreachable
because orphan detection needs a run, and runs are blocked by escrow_state=pending.
The gap is self-locking. Measured keys: on-disk 9b4a9a9d... vs recovered-from-R
30ef574f... This is R-218's shape one level up: succeeding at the self-heal stopped
the box OFFERING the recovery it still needed.

Registers: R-201 moved to its outcome; R-241 filed; capability map's recovery row
stays FAIL with both halves and the cause named; STATUS rewritten for the operator.
Highest ID R-238 -> R-241.

The venue is left with the recovered key in place and the self-heal key moved
aside, never deleted. Teardown still owed.
2026-08-07 06:45:36 +02:00

22 KiB
Raw Blame History

THE FINAL WALK (R-201) — overnight, unattended — journal

Venue: demo-hp VM 324 finalwalk-appliance (finalwalk.felhom.eu @ 192.168.0.142), guest LXC 9201, hub customer finalwalk, host id finalwalk-ed05d6. All three earlier venues were torn down on 2026-08-06; nothing was reused except a freed Storage Box sub-account number.

Written before the destruction, per §9.9.


THE HEADLINE FINDING — a fresh install does NOT get the fixes

landed on vouched
agent 0.127.0 0.127.0
controller 0.203.0 golden 0.203.0
newest released controller 0.205.0

No hand upgrade was needed and none was applied — that part works. But the vouched golden still bakes controller 0.203.0, so a box installed tonight is two releases behind: it has neither R-237 (v0.204.0, the restore list keyed on the store) nor R-234 (v0.205.0, the skipped-app verdict and the single-flight message).

This is not a regression — it is a delivery gap. The fixes exist, are tested and are pushed; what is missing is a golden carrying them and a vouch. Three of tonight's five checks measure exactly those fixes, and they measure the OLD behaviour because that is what a customer receives.

Timeline: bind 21:56:42Z → agent 0.127.0 ONLINE 21:59:01Z (2 m 19 s) → controller 0.203.0 reporting 22:01:03Z → guest 9201 running.


Phase A — the fixture (six records, §4)

1. Installed from the published ISO. felhom-installer-1.26.1-pve9.2-1.iso, verified byte-identical to the published copy (sha256 f3cc86d5f0ec…59a6 local == iso.felhom.eu). Served installer script tag confirmed installer-v1.25.0 on both git-syncs.

All three known TUI traps handled: GRUB's graphical default (down+ret inside one command, terminal installer first try); the Hungarian keymap switched to U.S. Englishpositive control: the administrator email rendered finalwalk@felhom.eu, and @ is shift-2 on US vs AltGr+V on HU, the only available evidence for the 24 masked password characters; and --boot set in its own qm set with the ISO detached, both verified from qm config before first boot (boot: order=scsi0, ide2 lines = 0), with auto-reboot unchecked and confirmed [ ] with focus moved away.

Day-0 fired unaided: pairing code T63-485, matching MAC bc:24:11:c6:73:d0.

2. Claimed through the real /claim form. The reset-code hatch was used — permitted in Phase A by §3, and it is a guest command line, so it is counted as such and does not touch the Phase-B rule.

3. App + three sentinels. calibre-web deployed with HDD_PATH=/mnt/felhom-drives/adatok, state: running and health_probe.healthy: true.

# file bytes sha256
A FINALWALK-SENTINEL-A.txt 57 863fa61c091c64488d8224b12f3915bfe264c46f9cad3d3d874b2ce725a1e5ee
B FINALWALK-őrszem-ékezetes-árvíztűrő.txt 80 b56668663035a332ae9b8a77b5847c8310c7ff48c5c2068114060561389ccf23
C FINALWALK-SENTINEL-C-12MB.bin 12 582 912 28630aa93119af0790b749671ef3896dbab88f8d239da8313ef5631fa068efda

Sentinel B's filename as hex, identical at source and on the box: 46494e414c57414c4b2d c591 72737a656d2d c3a9 6b657a657465732d c3a1 7276 c3ad 7a74 c5b1 72 c591 2e747874ő é á í ű ő, no efbfbd. Written from explicit bytes via pct push + a Python placer; no shell chain ever saw the name.

4. Escrow ceremony. Preflight 6 of 6 green (pbs_storage_id · dr_tier · age_binary · hub_upload · staged_secret · sudo_grant), agent_supported: true. Result: phase: done · restic_pw_sealed: TRUE · uploaded: true · entropy_bits: 129.24 · key_fingerprint: 81:dc:91:ce:a1:d0:50:3a:…:68:d5.

R was claimed ONE-SHOT and streamed file→file to ~/.config/finalwalk/R_finalwalk.txt (0600, DooPlex only); the guest and jump-host copies were shred -u'd and the raw response deleted. It was never rendered, never an argument, never a log line. Shape only: 10 words, 85 characters.

5. Off-site backup, and the sentinels proved BY NAME.

1da4f80d  2026-08-06 22:24:33  finalwalk  [felhom-offbox, calibre-web]
  /mnt/felhom-drives/adatok/userdata/media/books/FINALWALK-SENTINEL-A.txt                  57
  /mnt/felhom-drives/adatok/userdata/media/books/FINALWALK-SENTINEL-C-12MB.bin       12582912
  /mnt/felhom-drives/adatok/userdata/media/books/FINALWALK-őrszem-ékezetes-árvíztűrő.txt   80

HARNESS FAULT, and the §4 gate is what caught it. The FIRST run reported ok with a 26.6 KB repository — impossible for a 12 MB incompressible sentinel. Listing showed why: I had placed the files under …/adatok/**felhom-data**/userdata/…, while this box's namespace root is /mnt/felhom-drives/adatok directly, so they were never in the app's data at all. The product was correct throughout — it captured the unit and the declared mandatory path, including calibre-web's real metadata.db. Moved to the right path, re-run: 12.0 MB and all three listed. The gate's rule — prove by listing, never by a green status — is exactly what stopped a destruction that would have proven nothing.

6. Pre-destruction truth. Controller 0.203.0 · agent 0.127.0 · PBS wrapper matches vouched · guests 1/1 · DR recipe present · key escrow present · snapshot 1da4f80d · 1 snapshot · 12.0 MB · repo sftp:u629488-sub4@…:/home/felhom-repo port 23.


The five checks (§5) — on controller 0.203.0, which is what a customer gets

check observable verdict
T1 selected app that cannot be captured (opengist, not deployed) → run verdict ok; warning „Figyelmeztetés: 1 alkalmazásnak nincs elérhető mentése, ezek kimaradtak: opengist"; counters intact (1 snapshot, 12.0 MB, last_success advanced) old behaviour. v0.205.0 also keeps ok for an undeployed app — deliberately — but says which, why and what to do. Here the message is a bare count with no next step
T2 two off-site runs back to back run #1 „A távoli mentés elindult…"; run #2 the same message, flash_error count 0 FAILS. This is R-234's measured cause, live. v0.205.0 answers „Már fut egy távoli mentés — ez a kérés nem indított újat…" as an error
T3 toggle future-backup off for an app that HAS a snapshot, then open the restore page restore entries for calibre-web: 0; wizard 302 away FAILS. R-237 exactly: the customer's existing backup is hidden by a setting about the future
T4 off-site run with nothing selected verdict ok; warning „Sikeres — nincs mentésre jelölt alkalmazás" unchanged, as required — **and the wording is still "successful" beside "nothing is covered". Filed (§8)
T5 full-restore wizard driven as a browser does prepare → &full_prep=calibre-web&full_size=12.8 MB → following it revealed the confirm → commit → ok: true, „A(z) calibre-web teljes mentése visszaállítva ellenőrző mappába" PASSES. R-238 confirmed a harness artifact, not a product defect — carrying the state the wizard hands back makes it complete

Live sentinels re-verified after T5: all three MATCH (the verification restore writes to a separate folder and left live data untouched).

T2 and T3 are the same finding as the headline, seen from the customer's side: the fixes are written and pushed but not delivered, so tonight's box still exhibits both defects.


Phase B.1 — the soak

Window 1: 22:44:05Z → 01:54:03Z (3 h 10 m), untouched.

Every periodic job fired at exactly its declared cadence, which is the positive control that the scheduler was running at all:

job cadence fired expected in 190 min
agent-channel-health 1 m 185 ~190
stack-scan 2 m 108 ~95
system-health · backup-cache · offsite-credential-retry 5 m 44 each ~38
hub-report 15 m 14 ~13
tier · fill-watch · db-dump 1 each

offsite-credential-retry ran 44 times and did no work and said nothing. That is R-218's asserted healthy-box behaviour — the job exists, ticks, completes in 0 s, and stays silent because the declaration it keys on is false. A box that had never seen the defect behaves exactly as designed.

No alert, notification or digest fired. Nothing on the must-not list fired. The off-site state was unchanged throughout (last_run fixed at 22:38:54Z, snaps=1, 12.0 MB).

AND THAT LAST LINE IS NOT A FINDING — IT IS MY PLANNING ERROR, STATED AS SUCH. The daily jobs run on the CONTROLLER's clock, and the guest is UTC while the appliance is CEST (date +%Z: guest UTC, appliance CEST). The nightly local backup (~02:30) and off-site (~04:15) therefore fall at 02:30Z and 04:15Z, and I sized the window against CEST — so it closed at 01:54Z, before either. Reporting "the nightly did not fire" as a defect would have been a false finding produced by a badly-chosen window, which is precisely the shape §6.1 warns about in the other direction.

Window 2 (corrective): 01:56Z → ~02:35Z, to cover the 02:30Z local backup. The 04:15Z off-site nightly is deliberately NOT covered: finishing the walk and leaving the machine at the claim screen before 07:00 is the primary deliverable (§11.1), and waiting for it would have put the destruction at ~06:35 CEST with no margin. Recorded as not run, with the reason — the off-site tier was exercised four times manually tonight instead, including a full listing by name.

Window 2: 01:56:28Z → 02:36:45Z — and the nightly DID fire, unprompted.

00:30:25Z  db-dump
01:30:19Z  tier + fill-watch          ← the local tier legs
02:00:35Z  metrics-prune
02:15:03Z  offbox-backup              ← the off-site nightly
02:16:05Z  offbox-backup
           offbox: snaps 1 → 2, last_run 22:38:54Z → 02:15:24Z

So the soak covered a genuine scheduled cycle after all: the local legs in window 1 and the off-site nightly in window 2. My 04:15 prediction was wrong in the other direction — it fires at ~02:15 controller time. Both the prediction and the correction are recorded rather than quietly fixed.

BOTH DIRECTIONS, as §6.1 demands:

What fired and should have: every periodic job at its cadence; the DB-dump, tier and fill-watch legs; metrics-prune; the off-site nightly, which added the second snapshot with no prompting.

What should have fired and did not: nothing. Six registered jobs were never seen in the log — health-probes, status-refresh, ring-spill, deadapp-check, disk-health-check, selfupdate-check — and none of them is a finding. The first four are quiet by construction (scheduler.go:267quiet := job.Interval <= 30*time.Second, so a ≤30 s job never emits Running job:), and the last two run every 6 h, outside a ~4 h window. An absent log line is not evidence cuts this way too: I checked the source rather than filing four phantom defects.

What fired and should not have: nothing. No alert, notification, email or digest. The only WARN lines in five hours were three of mine — a failed login attempt, a CSRF token I mangled, and T1's deliberate opengist skip, which logged correctly.

One observation worth keeping: offbox-backup ticked twice, 62 s apart, and the snapshot count went 1 → 2, not 1 → 3. The second run was silently dropped by the single-flight — which for the NIGHTLY path is correct and deliberate (nobody asked; the next run retries). It is the same mechanism that, on the MANUAL path, produced R-234; v0.205.0 changes only the manual half and leaves this one silent, and this soak is live corroboration that the nightly half genuinely needs to stay quiet.


Phase B.2 — the destruction and the rebuild

All five §6.2 conditions were true and written down first (commits 2d2d8d3, f873c55, 502078b) before anything was destroyed.

02:40:31Z  pct destroy 9201 --purge     (guarded on hostname = finalwalk;
                                         demo-hp also has a guest 9201)
           both logical volumes removed; pct list empty
           /mnt/adatok and /mnt/mentes wiped to 20 K, MOUNTS LEFT IN PLACE
           — deliberately: the surviving raw mount IS the R-220 condition
02:40:56Z  felhom-host-install.sh v1.25.0, fetched live from felhom.eu/scripts/
02:43:28Z  Day-0 provision SUCCESS — 2 m 32 s

What the rebuild landed on — the second measurement of R-239:

before after
agent 0.127.0 0.127.0 — no downgrade, no hand upgrade
controller 0.203.0 0.203.0 — the same two-release gap

The agent half is exactly right: the reinstall neither downgraded nor needed a hand. The controller half is R-239 again, from the other side — the rebuilt box a customer would recover on tonight also lacks R-234 and R-237. Golden used: vzdump-lxc-9100-2026_08_06-23_58_32.tar.zst.

Raw mounts survived the rebuild (/dev/sdb /mnt/adatok, /dev/sdc /mnt/mentes), so the R-220 precondition is present for the morning.

Phase B.3 — THE HALT (§7)

The machine is at the claim screen and is waiting for the operator. Verified over HTTP from the appliance, with no guest shell:

GET /  ->  200,  <title>A szerver beállítása — Felhom
forms:  POST /claim   ·   POST /claim/request-new-code
claimed: null   ·   offbox config: absent (pristine rebuilt state)

A claim code has already been requested and emailed, through the customer-facing „Új kód kérése" path (POST /claim/request-new-code → 200, „Ha az e-mail cím regisztrálva van, elküldtük a kódot"), so the morning is paste a code, not request one and then paste it.

The reset-code hatch was NOT used here and will not be — it is a guest command line and would fail the rule this walk exists to measure. It was used once, in Phase A, where §3 permits it.

Honesty about the no-guest-command-line rule

The customer-journey steps from the destruction onward are driven over HTTP from the appliance to the guest's island address, which is what a browser would do. Some instrumentation readspct exec … python3 against settings.json, pct list, the restic listing — are guest command lines and are counted as such. They are not steps of the journey: none of them changed state, and none was needed to progress it. The distinction matters because conflating the two is how a walk claims a property it does not have.

§7's observation — the rebuilt, still-UNCLAIMED box

Three questions, all answered from the hub's own log rather than inferred:

1. Does it report while unclaimed? YES. host-report from finalwalk-ed05d6 (1 guests, 4 storage targets, 1 backups, 0 restore-tests, 1 pbs-snapshots, 12907 bytes) at 05:12:04 local, and Received report from finalwalk at 05:13:03.

2. Does it declare the credential need? YES, and the hub's other mechanism correctly declines to act on it — logged once a minute, from 02:59Z onward:

offsite-delivery: finalwalk: self-heal REFUSED — the box DECLARES offsite.state=needs_credential; internal/offsiteheal owns this remediation (it re-stages the stored credential before minting). A second mechanism minting here would double-issue.

3. Did the hub stage one unaided? YES, at 05:14:57 local (03:15Z):

offsiteheal: re-staged the stored one-time offsite secret for customer finalwalk (declared needs_credential across 2 reports) — the box re-consumes on its next cycle; no provider credential was minted

That is ~32 minutes after the rebuild, consistent with the documented 2 × 15-minute report debounce, and nobody touched anything.

This is a stronger case than yesterday's. On 2026-08-06 I pressed Re-issue 102 seconds after the self-heal had already fired and mistook my own button for the cause — the error that produced R-236 and forced its withdrawal. Here the box was pristine, rebuilt and not even claimed, no operator action of any kind was taken, and the chain ran end to end on its own. R-236's withdrawal is now confirmed on the exact shape it was filed against.

4. And then the box COLLECTED it, on its own tick — the full loop, unaided:

[offsite-apply] settle-gate: GO — at/above floor 0.156.0 (we are 0.203.0), no managed update running
[offsite-apply] offsite configured for u629488-sub4@…:/home/felhom-repo (pending key escrow)
[offsite-apply] credential retry: the staged credential was collected and the tier applied

The box's own settings went from no offbox object at all to enabled: true, host, user and repo_path populated, escrow_state: pending.

This is the first time that success line has ever been observed live. It shipped in controller v0.203.0; the 2026-08-06 walk only ever produced its sibling — "no unconsumed offsite password … (the box still declares a need; retrying)" — before an operator intervened. Here the whole chain ran with zero human action, on a box that has not even been claimed yet:

declare → offsite-delivery declines and names the owner → offsiteheal re-stages after two reports → the box's 5-minute retry collects it → the tier is applied.

That is R-218's consume half, R-236's withdrawal, and the previous walk's dead end 1, all settled by one unattended observation. By the time the operator claims this machine in the morning, its off-site tier is already up — which is exactly the property the recovery journey needed and never had.


THE MORNING REMAINDER, AND THE VERDICT

# step result
1 claim with the emailed code worked FIRST TRY — 302 in 0.164 s, accents intact (md5 identical source→box)
2 log in landed on „Indítópult"the recovery screen did NOT appear, and /recovery 302s away
3 read the screen as a customer there is nothing to read: no recovery pointer anywhere on the landing page
4 enter the recovery code impossible through the UI — there is no field to enter it in
5 read the listing not reachable by a customer
6 restore the sentinels byte-identical — but only via an operator command line

The verdict, both halves separately

THE DATA: PASS

All three sentinels came back byte-identical, out of snapshot f5c53b03, under the key recovered from the sealed package with R:

863fa61c091c64488d8224b12f3915bfe264c46f9cad3d3d874b2ce725a1e5ee  FINALWALK-SENTINEL-A.txt
28630aa93119af0790b749671ef3896dbab88f8d239da8313ef5631fa068efda  FINALWALK-SENTINEL-C-12MB.bin
b56668663035a332ae9b8a77b5847c8310c7ff48c5c2068114060561389ccf23  FINALWALK-őrszem-ékezetes-árvíztűrő.txt
name hex: 46494e414c57414c4b2d c591 72737a656d2d c3a9 6b657a657465732d c3a1 7276 c3ad 7a74 c5b1 72 c591 2e747874

The accented filename's bytes are identical too. The backup promise holds: disk → restic → SFTP → Storage Box → rebuilt machine → disk, intact.

THE JOURNEY: FAIL — and further from passing than the last walk

The customer has no route to their data at all. Not a slow one, not a confusing one — none.

  • / lands on „Indítópult" with no recovery pointer.
  • /recovery 302s away — the screen has retired itself.
  • /backups/remote says „Helyreállítási kód szükséges … Helyreállítási kód létrehozása" — it offers to create a NEW code, which would mint a new key and orphan the very backups R protects.
  • Even the operator's documented remedy refuses: --recover-offsite-install returned „[REFUSED] a DIFFERENT repository password is already present… which history to keep is not a decision this command may take. Nothing written." — correct, fail-closed, and a dead end.

The only way through was to move the fresh key aside by hand and re-run the install. That is three guest command lines, and it is what the walk exists to measure the absence of.

WHY — and it is the success I praised six hours earlier

OffsiteRecoveryOffer() offers the screen on exactly two conditions: (a) no repository password at all — the pristine rebuilt box — or (b) a password exists but the history will not open under it (OffboxOrphaned()).

Overnight, unaided and exactly as designed, the credential chain gave this box a fresh repository password at 03:18Z. That made (a) false. And (b) is false too, because orphan detection only fires when a run tries the repo — and runs are blocked by escrow_state: pending.

The box sits in the gap between the two conditions, and the gap is self-locking: it cannot detect the orphan without running, it cannot run without escrow, and it cannot escrow without minting a new code that destroys what R protects.

Measured, not deduced — the two keys:

on-disk (self-heal, 03:18Z) : 9b4a9a9dcec7898e7544f35b18470aac77c3d9064e5d3a302897617fa62edd65
recovered from R            : 30ef574fe492a43f89bf1a5071c44e89f51c44c6b08ebcb184b320a2634fad75

This is R-218's shape one level up. That finding read "succeeding at recovery stopped the box asking for what it still needed." Here: succeeding at the credential self-heal stopped the box offering the recovery it still needed — and the self-heal is the very mechanism this same walk proved working, six hours earlier, as its best result. Both things are true, and reporting only the first would have been the more flattering half of one night.

Filed as R-241.

State the machine is in

The venue is left with the recovered key in place (30ef574f…) and the self-heal key moved aside, not deleted (repo_password.selfheal-aside). The old repository opens; two snapshots; the restored copy sits in the controller container at /tmp/fwrestore. Teardown remains owed.