Three questions answered from the hub's own log, not inferred:
1. the rebuilt, still-unclaimed box DOES report (host-report + Received report)
2. it DOES declare offsite.state=needs_credential, and offsite-delivery correctly
declines once a minute, naming internal/offsiteheal as the owner
3. offsiteheal re-staged UNAIDED at 03:15Z, after two reports carried the
declaration, with no provider credential minted — about 32 minutes after the
rebuild, matching the documented 2x15-minute debounce
And then the box COLLECTED it on its own 5-minute tick:
[offsite-apply] credential retry: the staged credential was collected and the
tier applied
That success line shipped in v0.203.0 and this is the FIRST time it has been seen
live: yesterday's walk only produced its sibling before I intervened at 102s and
mistook my own button press for the cause — the error that produced R-236 and
forced its withdrawal. Here nobody touched anything and the box was not even
claimed. R-218's consume half, R-236's withdrawal and the previous walk's dead
end 1 are all settled by one unattended observation.
18 KiB
THE FINAL WALK (R-201) — overnight, unattended — journal
Venue: demo-hp VM 324 finalwalk-appliance (finalwalk.felhom.eu @ 192.168.0.142), guest
LXC 9201, hub customer finalwalk, host id finalwalk-ed05d6. All three earlier venues
were torn down on 2026-08-06; nothing was reused except a freed Storage Box sub-account number.
Written before the destruction, per §9.9.
THE HEADLINE FINDING — a fresh install does NOT get the fixes
| landed on | vouched | |
|---|---|---|
| agent | 0.127.0 ✅ | 0.127.0 |
| controller | 0.203.0 ❌ | golden 0.203.0 |
| newest released controller | 0.205.0 | — |
No hand upgrade was needed and none was applied — that part works. But the vouched golden still bakes controller 0.203.0, so a box installed tonight is two releases behind: it has neither R-237 (v0.204.0, the restore list keyed on the store) nor R-234 (v0.205.0, the skipped-app verdict and the single-flight message).
This is not a regression — it is a delivery gap. The fixes exist, are tested and are pushed; what is missing is a golden carrying them and a vouch. Three of tonight's five checks measure exactly those fixes, and they measure the OLD behaviour because that is what a customer receives.
Timeline: bind 21:56:42Z → agent 0.127.0 ONLINE 21:59:01Z (2 m 19 s) → controller 0.203.0 reporting 22:01:03Z → guest 9201 running.
Phase A — the fixture (six records, §4)
1. Installed from the published ISO. felhom-installer-1.26.1-pve9.2-1.iso, verified
byte-identical to the published copy (sha256 f3cc86d5f0ec…59a6 local == iso.felhom.eu). Served
installer script tag confirmed installer-v1.25.0 on both git-syncs.
All three known TUI traps handled: GRUB's graphical default (down+ret inside one command, terminal
installer first try); the Hungarian keymap switched to U.S. English — positive control: the
administrator email rendered finalwalk@felhom.eu, and @ is shift-2 on US vs AltGr+V on HU, the
only available evidence for the 24 masked password characters; and --boot set in its own qm set
with the ISO detached, both verified from qm config before first boot (boot: order=scsi0,
ide2 lines = 0), with auto-reboot unchecked and confirmed [ ] with focus moved away.
Day-0 fired unaided: pairing code T63-485, matching MAC bc:24:11:c6:73:d0.
2. Claimed through the real /claim form. The reset-code hatch was used — permitted in Phase A
by §3, and it is a guest command line, so it is counted as such and does not touch the Phase-B rule.
3. App + three sentinels. calibre-web deployed with HDD_PATH=/mnt/felhom-drives/adatok,
state: running and health_probe.healthy: true.
| # | file | bytes | sha256 |
|---|---|---|---|
| A | FINALWALK-SENTINEL-A.txt |
57 | 863fa61c091c64488d8224b12f3915bfe264c46f9cad3d3d874b2ce725a1e5ee |
| B | FINALWALK-őrszem-ékezetes-árvíztűrő.txt |
80 | b56668663035a332ae9b8a77b5847c8310c7ff48c5c2068114060561389ccf23 |
| C | FINALWALK-SENTINEL-C-12MB.bin |
12 582 912 | 28630aa93119af0790b749671ef3896dbab88f8d239da8313ef5631fa068efda |
Sentinel B's filename as hex, identical at source and on the box:
46494e414c57414c4b2d c591 72737a656d2d c3a9 6b657a657465732d c3a1 7276 c3ad 7a74 c5b1 72 c591 2e747874
— ő é á í ű ő, no efbfbd. Written from explicit bytes via pct push + a Python placer; no
shell chain ever saw the name.
4. Escrow ceremony. Preflight 6 of 6 green (pbs_storage_id · dr_tier · age_binary ·
hub_upload · staged_secret · sudo_grant), agent_supported: true.
Result: phase: done · restic_pw_sealed: TRUE · uploaded: true · entropy_bits: 129.24 ·
key_fingerprint: 81:dc:91:ce:a1:d0:50:3a:…:68:d5.
R was claimed ONE-SHOT and streamed file→file to ~/.config/finalwalk/R_finalwalk.txt (0600,
DooPlex only); the guest and jump-host copies were shred -u'd and the raw response deleted. It
was never rendered, never an argument, never a log line. Shape only: 10 words, 85 characters.
5. Off-site backup, and the sentinels proved BY NAME.
1da4f80d 2026-08-06 22:24:33 finalwalk [felhom-offbox, calibre-web]
/mnt/felhom-drives/adatok/userdata/media/books/FINALWALK-SENTINEL-A.txt 57
/mnt/felhom-drives/adatok/userdata/media/books/FINALWALK-SENTINEL-C-12MB.bin 12582912
/mnt/felhom-drives/adatok/userdata/media/books/FINALWALK-őrszem-ékezetes-árvíztűrő.txt 80
HARNESS FAULT, and the §4 gate is what caught it. The FIRST run reported
okwith a 26.6 KB repository — impossible for a 12 MB incompressible sentinel. Listing showed why: I had placed the files under…/adatok/**felhom-data**/userdata/…, while this box's namespace root is/mnt/felhom-drives/adatokdirectly, so they were never in the app's data at all. The product was correct throughout — it captured the unit and the declared mandatory path, including calibre-web's realmetadata.db. Moved to the right path, re-run: 12.0 MB and all three listed. The gate's rule — prove by listing, never by a green status — is exactly what stopped a destruction that would have proven nothing.
6. Pre-destruction truth. Controller 0.203.0 · agent 0.127.0 · PBS wrapper matches
vouched · guests 1/1 · DR recipe present · key escrow present · snapshot 1da4f80d ·
1 snapshot · 12.0 MB · repo sftp:u629488-sub4@…:/home/felhom-repo port 23.
The five checks (§5) — on controller 0.203.0, which is what a customer gets
| check | observable | verdict | |
|---|---|---|---|
| T1 | selected app that cannot be captured (opengist, not deployed) → run |
verdict ok; warning „Figyelmeztetés: 1 alkalmazásnak nincs elérhető mentése, ezek kimaradtak: opengist"; counters intact (1 snapshot, 12.0 MB, last_success advanced) |
old behaviour. v0.205.0 also keeps ok for an undeployed app — deliberately — but says which, why and what to do. Here the message is a bare count with no next step |
| T2 | two off-site runs back to back | run #1 „A távoli mentés elindult…"; run #2 the same message, flash_error count 0 |
FAILS. This is R-234's measured cause, live. v0.205.0 answers „Már fut egy távoli mentés — ez a kérés nem indított újat…" as an error |
| T3 | toggle future-backup off for an app that HAS a snapshot, then open the restore page | restore entries for calibre-web: 0; wizard 302 away | FAILS. R-237 exactly: the customer's existing backup is hidden by a setting about the future |
| T4 | off-site run with nothing selected | verdict ok; warning „Sikeres — nincs mentésre jelölt alkalmazás" |
unchanged, as required — **and the wording is still "successful" beside "nothing is covered". Filed (§8) |
| T5 | full-restore wizard driven as a browser does | prepare → &full_prep=calibre-web&full_size=12.8 MB → following it revealed the confirm → commit → ok: true, „A(z) calibre-web teljes mentése visszaállítva ellenőrző mappába" |
PASSES. R-238 confirmed a harness artifact, not a product defect — carrying the state the wizard hands back makes it complete |
Live sentinels re-verified after T5: all three MATCH (the verification restore writes to a separate folder and left live data untouched).
T2 and T3 are the same finding as the headline, seen from the customer's side: the fixes are written and pushed but not delivered, so tonight's box still exhibits both defects.
Phase B.1 — the soak
Window 1: 22:44:05Z → 01:54:03Z (3 h 10 m), untouched.
Every periodic job fired at exactly its declared cadence, which is the positive control that the scheduler was running at all:
| job | cadence | fired | expected in 190 min |
|---|---|---|---|
agent-channel-health |
1 m | 185 | ~190 |
stack-scan |
2 m | 108 | ~95 |
system-health · backup-cache · offsite-credential-retry |
5 m | 44 each | ~38 |
hub-report |
15 m | 14 | ~13 |
tier · fill-watch · db-dump |
— | 1 each | — |
offsite-credential-retry ran 44 times and did no work and said nothing. That is R-218's asserted
healthy-box behaviour — the job exists, ticks, completes in 0 s, and stays silent because the
declaration it keys on is false. A box that had never seen the defect behaves exactly as designed.
No alert, notification or digest fired. Nothing on the must-not list fired. The off-site state was
unchanged throughout (last_run fixed at 22:38:54Z, snaps=1, 12.0 MB).
AND THAT LAST LINE IS NOT A FINDING — IT IS MY PLANNING ERROR, STATED AS SUCH. The daily jobs run on the CONTROLLER's clock, and the guest is UTC while the appliance is CEST (
date +%Z: guestUTC, applianceCEST). The nightly local backup (~02:30) and off-site (~04:15) therefore fall at 02:30Z and 04:15Z, and I sized the window against CEST — so it closed at 01:54Z, before either. Reporting "the nightly did not fire" as a defect would have been a false finding produced by a badly-chosen window, which is precisely the shape §6.1 warns about in the other direction.Window 2 (corrective): 01:56Z → ~02:35Z, to cover the 02:30Z local backup. The 04:15Z off-site nightly is deliberately NOT covered: finishing the walk and leaving the machine at the claim screen before 07:00 is the primary deliverable (§11.1), and waiting for it would have put the destruction at ~06:35 CEST with no margin. Recorded as not run, with the reason — the off-site tier was exercised four times manually tonight instead, including a full listing by name.
Window 2: 01:56:28Z → 02:36:45Z — and the nightly DID fire, unprompted.
00:30:25Z db-dump
01:30:19Z tier + fill-watch ← the local tier legs
02:00:35Z metrics-prune
02:15:03Z offbox-backup ← the off-site nightly
02:16:05Z offbox-backup
offbox: snaps 1 → 2, last_run 22:38:54Z → 02:15:24Z
So the soak covered a genuine scheduled cycle after all: the local legs in window 1 and the off-site nightly in window 2. My 04:15 prediction was wrong in the other direction — it fires at ~02:15 controller time. Both the prediction and the correction are recorded rather than quietly fixed.
BOTH DIRECTIONS, as §6.1 demands:
What fired and should have: every periodic job at its cadence; the DB-dump, tier and fill-watch legs; metrics-prune; the off-site nightly, which added the second snapshot with no prompting.
What should have fired and did not: nothing. Six registered jobs were never seen in the log —
health-probes, status-refresh, ring-spill, deadapp-check, disk-health-check,
selfupdate-check — and none of them is a finding. The first four are quiet by construction
(scheduler.go:267 — quiet := job.Interval <= 30*time.Second, so a ≤30 s job never emits
Running job:), and the last two run every 6 h, outside a ~4 h window. An absent log line is not
evidence cuts this way too: I checked the source rather than filing four phantom defects.
What fired and should not have: nothing. No alert, notification, email or digest. The only WARN
lines in five hours were three of mine — a failed login attempt, a CSRF token I mangled, and T1's
deliberate opengist skip, which logged correctly.
One observation worth keeping: offbox-backup ticked twice, 62 s apart, and the snapshot count
went 1 → 2, not 1 → 3. The second run was silently dropped by the single-flight — which for the
NIGHTLY path is correct and deliberate (nobody asked; the next run retries). It is the same mechanism
that, on the MANUAL path, produced R-234; v0.205.0 changes only the manual half and leaves this one
silent, and this soak is live corroboration that the nightly half genuinely needs to stay quiet.
Phase B.2 — the destruction and the rebuild
All five §6.2 conditions were true and written down first (commits 2d2d8d3, f873c55,
502078b) before anything was destroyed.
02:40:31Z pct destroy 9201 --purge (guarded on hostname = finalwalk;
demo-hp also has a guest 9201)
both logical volumes removed; pct list empty
/mnt/adatok and /mnt/mentes wiped to 20 K, MOUNTS LEFT IN PLACE
— deliberately: the surviving raw mount IS the R-220 condition
02:40:56Z felhom-host-install.sh v1.25.0, fetched live from felhom.eu/scripts/
02:43:28Z Day-0 provision SUCCESS — 2 m 32 s
What the rebuild landed on — the second measurement of R-239:
| before | after | |
|---|---|---|
| agent | 0.127.0 | 0.127.0 — no downgrade, no hand upgrade |
| controller | 0.203.0 | 0.203.0 — the same two-release gap |
The agent half is exactly right: the reinstall neither downgraded nor needed a hand. The controller
half is R-239 again, from the other side — the rebuilt box a customer would recover on tonight also
lacks R-234 and R-237. Golden used: vzdump-lxc-9100-2026_08_06-23_58_32.tar.zst.
Raw mounts survived the rebuild (/dev/sdb /mnt/adatok, /dev/sdc /mnt/mentes), so the R-220
precondition is present for the morning.
Phase B.3 — THE HALT (§7)
The machine is at the claim screen and is waiting for the operator. Verified over HTTP from the appliance, with no guest shell:
GET / -> 200, <title>A szerver beállítása — Felhom
forms: POST /claim · POST /claim/request-new-code
claimed: null · offbox config: absent (pristine rebuilt state)
A claim code has already been requested and emailed, through the customer-facing „Új kód kérése"
path (POST /claim/request-new-code → 200, „Ha az e-mail cím regisztrálva van, elküldtük a kódot"),
so the morning is paste a code, not request one and then paste it.
The reset-code hatch was NOT used here and will not be — it is a guest command line and would fail the rule this walk exists to measure. It was used once, in Phase A, where §3 permits it.
Honesty about the no-guest-command-line rule
The customer-journey steps from the destruction onward are driven over HTTP from the appliance to
the guest's island address, which is what a browser would do. Some instrumentation reads —
pct exec … python3 against settings.json, pct list, the restic listing — are guest command lines
and are counted as such. They are not steps of the journey: none of them changed state, and none was
needed to progress it. The distinction matters because conflating the two is how a walk claims a
property it does not have.
§7's observation — the rebuilt, still-UNCLAIMED box
Three questions, all answered from the hub's own log rather than inferred:
1. Does it report while unclaimed? YES.
host-report from finalwalk-ed05d6 (1 guests, 4 storage targets, 1 backups, 0 restore-tests, 1 pbs-snapshots, 12907 bytes) at 05:12:04 local, and Received report from finalwalk at 05:13:03.
2. Does it declare the credential need? YES, and the hub's other mechanism correctly declines to act on it — logged once a minute, from 02:59Z onward:
offsite-delivery: finalwalk: self-heal REFUSED — the box DECLARES offsite.state=needs_credential; internal/offsiteheal owns this remediation (it re-stages the stored credential before minting). A second mechanism minting here would double-issue.
3. Did the hub stage one unaided? YES, at 05:14:57 local (03:15Z):
offsiteheal: re-staged the stored one-time offsite secret for customer finalwalk (declared needs_credential across 2 reports) — the box re-consumes on its next cycle; no provider credential was minted
That is ~32 minutes after the rebuild, consistent with the documented 2 × 15-minute report debounce, and nobody touched anything.
This is a stronger case than yesterday's. On 2026-08-06 I pressed Re-issue 102 seconds after the self-heal had already fired and mistook my own button for the cause — the error that produced R-236 and forced its withdrawal. Here the box was pristine, rebuilt and not even claimed, no operator action of any kind was taken, and the chain ran end to end on its own. R-236's withdrawal is now confirmed on the exact shape it was filed against.
4. And then the box COLLECTED it, on its own tick — the full loop, unaided:
[offsite-apply] settle-gate: GO — at/above floor 0.156.0 (we are 0.203.0), no managed update running
[offsite-apply] offsite configured for u629488-sub4@…:/home/felhom-repo (pending key escrow)
[offsite-apply] credential retry: the staged credential was collected and the tier applied
The box's own settings went from no offbox object at all to enabled: true, host, user and
repo_path populated, escrow_state: pending.
This is the first time that success line has ever been observed live. It shipped in controller v0.203.0; the 2026-08-06 walk only ever produced its sibling — "no unconsumed offsite password … (the box still declares a need; retrying)" — before an operator intervened. Here the whole chain ran with zero human action, on a box that has not even been claimed yet:
declare →
offsite-deliverydeclines and names the owner →offsitehealre-stages after two reports → the box's 5-minute retry collects it → the tier is applied.
That is R-218's consume half, R-236's withdrawal, and the previous walk's dead end 1, all settled by one unattended observation. By the time the operator claims this machine in the morning, its off-site tier is already up — which is exactly the property the recovery journey needed and never had.