Files
felhom.eu/documentation/tests/finalwalk-r201-2026-08-07/journal.md
T
admin 1ff6f8f8e0
gates / gates (push) Successful in 12s
final walk §7: the credential chain runs end to end, unaided, on an UNCLAIMED box
Three questions answered from the hub's own log, not inferred:
  1. the rebuilt, still-unclaimed box DOES report (host-report + Received report)
  2. it DOES declare offsite.state=needs_credential, and offsite-delivery correctly
     declines once a minute, naming internal/offsiteheal as the owner
  3. offsiteheal re-staged UNAIDED at 03:15Z, after two reports carried the
     declaration, with no provider credential minted — about 32 minutes after the
     rebuild, matching the documented 2x15-minute debounce

And then the box COLLECTED it on its own 5-minute tick:
  [offsite-apply] credential retry: the staged credential was collected and the
  tier applied

That success line shipped in v0.203.0 and this is the FIRST time it has been seen
live: yesterday's walk only produced its sibling before I intervened at 102s and
mistook my own button press for the cause — the error that produced R-236 and
forced its withdrawal. Here nobody touched anything and the box was not even
claimed. R-218's consume half, R-236's withdrawal and the previous walk's dead
end 1 are all settled by one unattended observation.
2026-08-07 05:23:03 +02:00

18 KiB
Raw Blame History

THE FINAL WALK (R-201) — overnight, unattended — journal

Venue: demo-hp VM 324 finalwalk-appliance (finalwalk.felhom.eu @ 192.168.0.142), guest LXC 9201, hub customer finalwalk, host id finalwalk-ed05d6. All three earlier venues were torn down on 2026-08-06; nothing was reused except a freed Storage Box sub-account number.

Written before the destruction, per §9.9.


THE HEADLINE FINDING — a fresh install does NOT get the fixes

landed on vouched
agent 0.127.0 0.127.0
controller 0.203.0 golden 0.203.0
newest released controller 0.205.0

No hand upgrade was needed and none was applied — that part works. But the vouched golden still bakes controller 0.203.0, so a box installed tonight is two releases behind: it has neither R-237 (v0.204.0, the restore list keyed on the store) nor R-234 (v0.205.0, the skipped-app verdict and the single-flight message).

This is not a regression — it is a delivery gap. The fixes exist, are tested and are pushed; what is missing is a golden carrying them and a vouch. Three of tonight's five checks measure exactly those fixes, and they measure the OLD behaviour because that is what a customer receives.

Timeline: bind 21:56:42Z → agent 0.127.0 ONLINE 21:59:01Z (2 m 19 s) → controller 0.203.0 reporting 22:01:03Z → guest 9201 running.


Phase A — the fixture (six records, §4)

1. Installed from the published ISO. felhom-installer-1.26.1-pve9.2-1.iso, verified byte-identical to the published copy (sha256 f3cc86d5f0ec…59a6 local == iso.felhom.eu). Served installer script tag confirmed installer-v1.25.0 on both git-syncs.

All three known TUI traps handled: GRUB's graphical default (down+ret inside one command, terminal installer first try); the Hungarian keymap switched to U.S. Englishpositive control: the administrator email rendered finalwalk@felhom.eu, and @ is shift-2 on US vs AltGr+V on HU, the only available evidence for the 24 masked password characters; and --boot set in its own qm set with the ISO detached, both verified from qm config before first boot (boot: order=scsi0, ide2 lines = 0), with auto-reboot unchecked and confirmed [ ] with focus moved away.

Day-0 fired unaided: pairing code T63-485, matching MAC bc:24:11:c6:73:d0.

2. Claimed through the real /claim form. The reset-code hatch was used — permitted in Phase A by §3, and it is a guest command line, so it is counted as such and does not touch the Phase-B rule.

3. App + three sentinels. calibre-web deployed with HDD_PATH=/mnt/felhom-drives/adatok, state: running and health_probe.healthy: true.

# file bytes sha256
A FINALWALK-SENTINEL-A.txt 57 863fa61c091c64488d8224b12f3915bfe264c46f9cad3d3d874b2ce725a1e5ee
B FINALWALK-őrszem-ékezetes-árvíztűrő.txt 80 b56668663035a332ae9b8a77b5847c8310c7ff48c5c2068114060561389ccf23
C FINALWALK-SENTINEL-C-12MB.bin 12 582 912 28630aa93119af0790b749671ef3896dbab88f8d239da8313ef5631fa068efda

Sentinel B's filename as hex, identical at source and on the box: 46494e414c57414c4b2d c591 72737a656d2d c3a9 6b657a657465732d c3a1 7276 c3ad 7a74 c5b1 72 c591 2e747874ő é á í ű ő, no efbfbd. Written from explicit bytes via pct push + a Python placer; no shell chain ever saw the name.

4. Escrow ceremony. Preflight 6 of 6 green (pbs_storage_id · dr_tier · age_binary · hub_upload · staged_secret · sudo_grant), agent_supported: true. Result: phase: done · restic_pw_sealed: TRUE · uploaded: true · entropy_bits: 129.24 · key_fingerprint: 81:dc:91:ce:a1:d0:50:3a:…:68:d5.

R was claimed ONE-SHOT and streamed file→file to ~/.config/finalwalk/R_finalwalk.txt (0600, DooPlex only); the guest and jump-host copies were shred -u'd and the raw response deleted. It was never rendered, never an argument, never a log line. Shape only: 10 words, 85 characters.

5. Off-site backup, and the sentinels proved BY NAME.

1da4f80d  2026-08-06 22:24:33  finalwalk  [felhom-offbox, calibre-web]
  /mnt/felhom-drives/adatok/userdata/media/books/FINALWALK-SENTINEL-A.txt                  57
  /mnt/felhom-drives/adatok/userdata/media/books/FINALWALK-SENTINEL-C-12MB.bin       12582912
  /mnt/felhom-drives/adatok/userdata/media/books/FINALWALK-őrszem-ékezetes-árvíztűrő.txt   80

HARNESS FAULT, and the §4 gate is what caught it. The FIRST run reported ok with a 26.6 KB repository — impossible for a 12 MB incompressible sentinel. Listing showed why: I had placed the files under …/adatok/**felhom-data**/userdata/…, while this box's namespace root is /mnt/felhom-drives/adatok directly, so they were never in the app's data at all. The product was correct throughout — it captured the unit and the declared mandatory path, including calibre-web's real metadata.db. Moved to the right path, re-run: 12.0 MB and all three listed. The gate's rule — prove by listing, never by a green status — is exactly what stopped a destruction that would have proven nothing.

6. Pre-destruction truth. Controller 0.203.0 · agent 0.127.0 · PBS wrapper matches vouched · guests 1/1 · DR recipe present · key escrow present · snapshot 1da4f80d · 1 snapshot · 12.0 MB · repo sftp:u629488-sub4@…:/home/felhom-repo port 23.


The five checks (§5) — on controller 0.203.0, which is what a customer gets

check observable verdict
T1 selected app that cannot be captured (opengist, not deployed) → run verdict ok; warning „Figyelmeztetés: 1 alkalmazásnak nincs elérhető mentése, ezek kimaradtak: opengist"; counters intact (1 snapshot, 12.0 MB, last_success advanced) old behaviour. v0.205.0 also keeps ok for an undeployed app — deliberately — but says which, why and what to do. Here the message is a bare count with no next step
T2 two off-site runs back to back run #1 „A távoli mentés elindult…"; run #2 the same message, flash_error count 0 FAILS. This is R-234's measured cause, live. v0.205.0 answers „Már fut egy távoli mentés — ez a kérés nem indított újat…" as an error
T3 toggle future-backup off for an app that HAS a snapshot, then open the restore page restore entries for calibre-web: 0; wizard 302 away FAILS. R-237 exactly: the customer's existing backup is hidden by a setting about the future
T4 off-site run with nothing selected verdict ok; warning „Sikeres — nincs mentésre jelölt alkalmazás" unchanged, as required — **and the wording is still "successful" beside "nothing is covered". Filed (§8)
T5 full-restore wizard driven as a browser does prepare → &full_prep=calibre-web&full_size=12.8 MB → following it revealed the confirm → commit → ok: true, „A(z) calibre-web teljes mentése visszaállítva ellenőrző mappába" PASSES. R-238 confirmed a harness artifact, not a product defect — carrying the state the wizard hands back makes it complete

Live sentinels re-verified after T5: all three MATCH (the verification restore writes to a separate folder and left live data untouched).

T2 and T3 are the same finding as the headline, seen from the customer's side: the fixes are written and pushed but not delivered, so tonight's box still exhibits both defects.


Phase B.1 — the soak

Window 1: 22:44:05Z → 01:54:03Z (3 h 10 m), untouched.

Every periodic job fired at exactly its declared cadence, which is the positive control that the scheduler was running at all:

job cadence fired expected in 190 min
agent-channel-health 1 m 185 ~190
stack-scan 2 m 108 ~95
system-health · backup-cache · offsite-credential-retry 5 m 44 each ~38
hub-report 15 m 14 ~13
tier · fill-watch · db-dump 1 each

offsite-credential-retry ran 44 times and did no work and said nothing. That is R-218's asserted healthy-box behaviour — the job exists, ticks, completes in 0 s, and stays silent because the declaration it keys on is false. A box that had never seen the defect behaves exactly as designed.

No alert, notification or digest fired. Nothing on the must-not list fired. The off-site state was unchanged throughout (last_run fixed at 22:38:54Z, snaps=1, 12.0 MB).

AND THAT LAST LINE IS NOT A FINDING — IT IS MY PLANNING ERROR, STATED AS SUCH. The daily jobs run on the CONTROLLER's clock, and the guest is UTC while the appliance is CEST (date +%Z: guest UTC, appliance CEST). The nightly local backup (~02:30) and off-site (~04:15) therefore fall at 02:30Z and 04:15Z, and I sized the window against CEST — so it closed at 01:54Z, before either. Reporting "the nightly did not fire" as a defect would have been a false finding produced by a badly-chosen window, which is precisely the shape §6.1 warns about in the other direction.

Window 2 (corrective): 01:56Z → ~02:35Z, to cover the 02:30Z local backup. The 04:15Z off-site nightly is deliberately NOT covered: finishing the walk and leaving the machine at the claim screen before 07:00 is the primary deliverable (§11.1), and waiting for it would have put the destruction at ~06:35 CEST with no margin. Recorded as not run, with the reason — the off-site tier was exercised four times manually tonight instead, including a full listing by name.

Window 2: 01:56:28Z → 02:36:45Z — and the nightly DID fire, unprompted.

00:30:25Z  db-dump
01:30:19Z  tier + fill-watch          ← the local tier legs
02:00:35Z  metrics-prune
02:15:03Z  offbox-backup              ← the off-site nightly
02:16:05Z  offbox-backup
           offbox: snaps 1 → 2, last_run 22:38:54Z → 02:15:24Z

So the soak covered a genuine scheduled cycle after all: the local legs in window 1 and the off-site nightly in window 2. My 04:15 prediction was wrong in the other direction — it fires at ~02:15 controller time. Both the prediction and the correction are recorded rather than quietly fixed.

BOTH DIRECTIONS, as §6.1 demands:

What fired and should have: every periodic job at its cadence; the DB-dump, tier and fill-watch legs; metrics-prune; the off-site nightly, which added the second snapshot with no prompting.

What should have fired and did not: nothing. Six registered jobs were never seen in the log — health-probes, status-refresh, ring-spill, deadapp-check, disk-health-check, selfupdate-check — and none of them is a finding. The first four are quiet by construction (scheduler.go:267quiet := job.Interval <= 30*time.Second, so a ≤30 s job never emits Running job:), and the last two run every 6 h, outside a ~4 h window. An absent log line is not evidence cuts this way too: I checked the source rather than filing four phantom defects.

What fired and should not have: nothing. No alert, notification, email or digest. The only WARN lines in five hours were three of mine — a failed login attempt, a CSRF token I mangled, and T1's deliberate opengist skip, which logged correctly.

One observation worth keeping: offbox-backup ticked twice, 62 s apart, and the snapshot count went 1 → 2, not 1 → 3. The second run was silently dropped by the single-flight — which for the NIGHTLY path is correct and deliberate (nobody asked; the next run retries). It is the same mechanism that, on the MANUAL path, produced R-234; v0.205.0 changes only the manual half and leaves this one silent, and this soak is live corroboration that the nightly half genuinely needs to stay quiet.


Phase B.2 — the destruction and the rebuild

All five §6.2 conditions were true and written down first (commits 2d2d8d3, f873c55, 502078b) before anything was destroyed.

02:40:31Z  pct destroy 9201 --purge     (guarded on hostname = finalwalk;
                                         demo-hp also has a guest 9201)
           both logical volumes removed; pct list empty
           /mnt/adatok and /mnt/mentes wiped to 20 K, MOUNTS LEFT IN PLACE
           — deliberately: the surviving raw mount IS the R-220 condition
02:40:56Z  felhom-host-install.sh v1.25.0, fetched live from felhom.eu/scripts/
02:43:28Z  Day-0 provision SUCCESS — 2 m 32 s

What the rebuild landed on — the second measurement of R-239:

before after
agent 0.127.0 0.127.0 — no downgrade, no hand upgrade
controller 0.203.0 0.203.0 — the same two-release gap

The agent half is exactly right: the reinstall neither downgraded nor needed a hand. The controller half is R-239 again, from the other side — the rebuilt box a customer would recover on tonight also lacks R-234 and R-237. Golden used: vzdump-lxc-9100-2026_08_06-23_58_32.tar.zst.

Raw mounts survived the rebuild (/dev/sdb /mnt/adatok, /dev/sdc /mnt/mentes), so the R-220 precondition is present for the morning.

Phase B.3 — THE HALT (§7)

The machine is at the claim screen and is waiting for the operator. Verified over HTTP from the appliance, with no guest shell:

GET /  ->  200,  <title>A szerver beállítása — Felhom
forms:  POST /claim   ·   POST /claim/request-new-code
claimed: null   ·   offbox config: absent (pristine rebuilt state)

A claim code has already been requested and emailed, through the customer-facing „Új kód kérése" path (POST /claim/request-new-code → 200, „Ha az e-mail cím regisztrálva van, elküldtük a kódot"), so the morning is paste a code, not request one and then paste it.

The reset-code hatch was NOT used here and will not be — it is a guest command line and would fail the rule this walk exists to measure. It was used once, in Phase A, where §3 permits it.

Honesty about the no-guest-command-line rule

The customer-journey steps from the destruction onward are driven over HTTP from the appliance to the guest's island address, which is what a browser would do. Some instrumentation readspct exec … python3 against settings.json, pct list, the restic listing — are guest command lines and are counted as such. They are not steps of the journey: none of them changed state, and none was needed to progress it. The distinction matters because conflating the two is how a walk claims a property it does not have.

§7's observation — the rebuilt, still-UNCLAIMED box

Three questions, all answered from the hub's own log rather than inferred:

1. Does it report while unclaimed? YES. host-report from finalwalk-ed05d6 (1 guests, 4 storage targets, 1 backups, 0 restore-tests, 1 pbs-snapshots, 12907 bytes) at 05:12:04 local, and Received report from finalwalk at 05:13:03.

2. Does it declare the credential need? YES, and the hub's other mechanism correctly declines to act on it — logged once a minute, from 02:59Z onward:

offsite-delivery: finalwalk: self-heal REFUSED — the box DECLARES offsite.state=needs_credential; internal/offsiteheal owns this remediation (it re-stages the stored credential before minting). A second mechanism minting here would double-issue.

3. Did the hub stage one unaided? YES, at 05:14:57 local (03:15Z):

offsiteheal: re-staged the stored one-time offsite secret for customer finalwalk (declared needs_credential across 2 reports) — the box re-consumes on its next cycle; no provider credential was minted

That is ~32 minutes after the rebuild, consistent with the documented 2 × 15-minute report debounce, and nobody touched anything.

This is a stronger case than yesterday's. On 2026-08-06 I pressed Re-issue 102 seconds after the self-heal had already fired and mistook my own button for the cause — the error that produced R-236 and forced its withdrawal. Here the box was pristine, rebuilt and not even claimed, no operator action of any kind was taken, and the chain ran end to end on its own. R-236's withdrawal is now confirmed on the exact shape it was filed against.

4. And then the box COLLECTED it, on its own tick — the full loop, unaided:

[offsite-apply] settle-gate: GO — at/above floor 0.156.0 (we are 0.203.0), no managed update running
[offsite-apply] offsite configured for u629488-sub4@…:/home/felhom-repo (pending key escrow)
[offsite-apply] credential retry: the staged credential was collected and the tier applied

The box's own settings went from no offbox object at all to enabled: true, host, user and repo_path populated, escrow_state: pending.

This is the first time that success line has ever been observed live. It shipped in controller v0.203.0; the 2026-08-06 walk only ever produced its sibling — "no unconsumed offsite password … (the box still declares a need; retrying)" — before an operator intervened. Here the whole chain ran with zero human action, on a box that has not even been claimed yet:

declare → offsite-delivery declines and names the owner → offsiteheal re-stages after two reports → the box's 5-minute retry collects it → the tier is applied.

That is R-218's consume half, R-236's withdrawal, and the previous walk's dead end 1, all settled by one unattended observation. By the time the operator claims this machine in the morning, its off-site tier is already up — which is exactly the property the recovery journey needed and never had.