Files
felhom.eu/documentation/tests/finalwalk-r201-2026-08-07/journal.md
T
admin 1b490c8cbf
gates / gates (push) Successful in 27s
final walk: destroyed, rebuilt, and HALTED at the claim screen
Destroyed 02:40:31Z (guarded on hostname — demo-hp also has a guest 9201), drives
wiped to 20K with the mounts deliberately left in place because the surviving raw
mount IS the R-220 condition. Reinstalled through the published day-0 path,
installer v1.25.0 fetched live; Day-0 provision SUCCESS in 2m32s.

R-239 measured a second time, from the other side: the rebuild landed on agent
0.127.0 (no downgrade, no hand upgrade — that half is right) and controller
0.203.0. The box a customer would recover on tonight also lacks R-234 and R-237.

The machine is AT THE CLAIM SCREEN awaiting the operator. A claim code has already
been requested through the customer-facing path and emailed, so the morning is
paste-a-code rather than request-then-paste. The reset-code hatch was NOT used and
will not be: it is a guest command line and would fail the rule the walk measures.

Stated plainly in the journal: journey steps from the destruction onward are driven
over HTTP from the appliance to the guest's island address, as a browser would;
some instrumentation reads are guest command lines and are counted as such, but
none changed state or was needed to progress the journey.
2026-08-07 04:46:11 +02:00

14 KiB

THE FINAL WALK (R-201) — overnight, unattended — journal

Venue: demo-hp VM 324 finalwalk-appliance (finalwalk.felhom.eu @ 192.168.0.142), guest LXC 9201, hub customer finalwalk, host id finalwalk-ed05d6. All three earlier venues were torn down on 2026-08-06; nothing was reused except a freed Storage Box sub-account number.

Written before the destruction, per §9.9.


THE HEADLINE FINDING — a fresh install does NOT get the fixes

landed on vouched
agent 0.127.0 0.127.0
controller 0.203.0 golden 0.203.0
newest released controller 0.205.0

No hand upgrade was needed and none was applied — that part works. But the vouched golden still bakes controller 0.203.0, so a box installed tonight is two releases behind: it has neither R-237 (v0.204.0, the restore list keyed on the store) nor R-234 (v0.205.0, the skipped-app verdict and the single-flight message).

This is not a regression — it is a delivery gap. The fixes exist, are tested and are pushed; what is missing is a golden carrying them and a vouch. Three of tonight's five checks measure exactly those fixes, and they measure the OLD behaviour because that is what a customer receives.

Timeline: bind 21:56:42Z → agent 0.127.0 ONLINE 21:59:01Z (2 m 19 s) → controller 0.203.0 reporting 22:01:03Z → guest 9201 running.


Phase A — the fixture (six records, §4)

1. Installed from the published ISO. felhom-installer-1.26.1-pve9.2-1.iso, verified byte-identical to the published copy (sha256 f3cc86d5f0ec…59a6 local == iso.felhom.eu). Served installer script tag confirmed installer-v1.25.0 on both git-syncs.

All three known TUI traps handled: GRUB's graphical default (down+ret inside one command, terminal installer first try); the Hungarian keymap switched to U.S. Englishpositive control: the administrator email rendered finalwalk@felhom.eu, and @ is shift-2 on US vs AltGr+V on HU, the only available evidence for the 24 masked password characters; and --boot set in its own qm set with the ISO detached, both verified from qm config before first boot (boot: order=scsi0, ide2 lines = 0), with auto-reboot unchecked and confirmed [ ] with focus moved away.

Day-0 fired unaided: pairing code T63-485, matching MAC bc:24:11:c6:73:d0.

2. Claimed through the real /claim form. The reset-code hatch was used — permitted in Phase A by §3, and it is a guest command line, so it is counted as such and does not touch the Phase-B rule.

3. App + three sentinels. calibre-web deployed with HDD_PATH=/mnt/felhom-drives/adatok, state: running and health_probe.healthy: true.

# file bytes sha256
A FINALWALK-SENTINEL-A.txt 57 863fa61c091c64488d8224b12f3915bfe264c46f9cad3d3d874b2ce725a1e5ee
B FINALWALK-őrszem-ékezetes-árvíztűrő.txt 80 b56668663035a332ae9b8a77b5847c8310c7ff48c5c2068114060561389ccf23
C FINALWALK-SENTINEL-C-12MB.bin 12 582 912 28630aa93119af0790b749671ef3896dbab88f8d239da8313ef5631fa068efda

Sentinel B's filename as hex, identical at source and on the box: 46494e414c57414c4b2d c591 72737a656d2d c3a9 6b657a657465732d c3a1 7276 c3ad 7a74 c5b1 72 c591 2e747874ő é á í ű ő, no efbfbd. Written from explicit bytes via pct push + a Python placer; no shell chain ever saw the name.

4. Escrow ceremony. Preflight 6 of 6 green (pbs_storage_id · dr_tier · age_binary · hub_upload · staged_secret · sudo_grant), agent_supported: true. Result: phase: done · restic_pw_sealed: TRUE · uploaded: true · entropy_bits: 129.24 · key_fingerprint: 81:dc:91:ce:a1:d0:50:3a:…:68:d5.

R was claimed ONE-SHOT and streamed file→file to ~/.config/finalwalk/R_finalwalk.txt (0600, DooPlex only); the guest and jump-host copies were shred -u'd and the raw response deleted. It was never rendered, never an argument, never a log line. Shape only: 10 words, 85 characters.

5. Off-site backup, and the sentinels proved BY NAME.

1da4f80d  2026-08-06 22:24:33  finalwalk  [felhom-offbox, calibre-web]
  /mnt/felhom-drives/adatok/userdata/media/books/FINALWALK-SENTINEL-A.txt                  57
  /mnt/felhom-drives/adatok/userdata/media/books/FINALWALK-SENTINEL-C-12MB.bin       12582912
  /mnt/felhom-drives/adatok/userdata/media/books/FINALWALK-őrszem-ékezetes-árvíztűrő.txt   80

HARNESS FAULT, and the §4 gate is what caught it. The FIRST run reported ok with a 26.6 KB repository — impossible for a 12 MB incompressible sentinel. Listing showed why: I had placed the files under …/adatok/**felhom-data**/userdata/…, while this box's namespace root is /mnt/felhom-drives/adatok directly, so they were never in the app's data at all. The product was correct throughout — it captured the unit and the declared mandatory path, including calibre-web's real metadata.db. Moved to the right path, re-run: 12.0 MB and all three listed. The gate's rule — prove by listing, never by a green status — is exactly what stopped a destruction that would have proven nothing.

6. Pre-destruction truth. Controller 0.203.0 · agent 0.127.0 · PBS wrapper matches vouched · guests 1/1 · DR recipe present · key escrow present · snapshot 1da4f80d · 1 snapshot · 12.0 MB · repo sftp:u629488-sub4@…:/home/felhom-repo port 23.


The five checks (§5) — on controller 0.203.0, which is what a customer gets

check observable verdict
T1 selected app that cannot be captured (opengist, not deployed) → run verdict ok; warning „Figyelmeztetés: 1 alkalmazásnak nincs elérhető mentése, ezek kimaradtak: opengist"; counters intact (1 snapshot, 12.0 MB, last_success advanced) old behaviour. v0.205.0 also keeps ok for an undeployed app — deliberately — but says which, why and what to do. Here the message is a bare count with no next step
T2 two off-site runs back to back run #1 „A távoli mentés elindult…"; run #2 the same message, flash_error count 0 FAILS. This is R-234's measured cause, live. v0.205.0 answers „Már fut egy távoli mentés — ez a kérés nem indított újat…" as an error
T3 toggle future-backup off for an app that HAS a snapshot, then open the restore page restore entries for calibre-web: 0; wizard 302 away FAILS. R-237 exactly: the customer's existing backup is hidden by a setting about the future
T4 off-site run with nothing selected verdict ok; warning „Sikeres — nincs mentésre jelölt alkalmazás" unchanged, as required — **and the wording is still "successful" beside "nothing is covered". Filed (§8)
T5 full-restore wizard driven as a browser does prepare → &full_prep=calibre-web&full_size=12.8 MB → following it revealed the confirm → commit → ok: true, „A(z) calibre-web teljes mentése visszaállítva ellenőrző mappába" PASSES. R-238 confirmed a harness artifact, not a product defect — carrying the state the wizard hands back makes it complete

Live sentinels re-verified after T5: all three MATCH (the verification restore writes to a separate folder and left live data untouched).

T2 and T3 are the same finding as the headline, seen from the customer's side: the fixes are written and pushed but not delivered, so tonight's box still exhibits both defects.


Phase B.1 — the soak

Window 1: 22:44:05Z → 01:54:03Z (3 h 10 m), untouched.

Every periodic job fired at exactly its declared cadence, which is the positive control that the scheduler was running at all:

job cadence fired expected in 190 min
agent-channel-health 1 m 185 ~190
stack-scan 2 m 108 ~95
system-health · backup-cache · offsite-credential-retry 5 m 44 each ~38
hub-report 15 m 14 ~13
tier · fill-watch · db-dump 1 each

offsite-credential-retry ran 44 times and did no work and said nothing. That is R-218's asserted healthy-box behaviour — the job exists, ticks, completes in 0 s, and stays silent because the declaration it keys on is false. A box that had never seen the defect behaves exactly as designed.

No alert, notification or digest fired. Nothing on the must-not list fired. The off-site state was unchanged throughout (last_run fixed at 22:38:54Z, snaps=1, 12.0 MB).

AND THAT LAST LINE IS NOT A FINDING — IT IS MY PLANNING ERROR, STATED AS SUCH. The daily jobs run on the CONTROLLER's clock, and the guest is UTC while the appliance is CEST (date +%Z: guest UTC, appliance CEST). The nightly local backup (~02:30) and off-site (~04:15) therefore fall at 02:30Z and 04:15Z, and I sized the window against CEST — so it closed at 01:54Z, before either. Reporting "the nightly did not fire" as a defect would have been a false finding produced by a badly-chosen window, which is precisely the shape §6.1 warns about in the other direction.

Window 2 (corrective): 01:56Z → ~02:35Z, to cover the 02:30Z local backup. The 04:15Z off-site nightly is deliberately NOT covered: finishing the walk and leaving the machine at the claim screen before 07:00 is the primary deliverable (§11.1), and waiting for it would have put the destruction at ~06:35 CEST with no margin. Recorded as not run, with the reason — the off-site tier was exercised four times manually tonight instead, including a full listing by name.

Window 2: 01:56:28Z → 02:36:45Z — and the nightly DID fire, unprompted.

00:30:25Z  db-dump
01:30:19Z  tier + fill-watch          ← the local tier legs
02:00:35Z  metrics-prune
02:15:03Z  offbox-backup              ← the off-site nightly
02:16:05Z  offbox-backup
           offbox: snaps 1 → 2, last_run 22:38:54Z → 02:15:24Z

So the soak covered a genuine scheduled cycle after all: the local legs in window 1 and the off-site nightly in window 2. My 04:15 prediction was wrong in the other direction — it fires at ~02:15 controller time. Both the prediction and the correction are recorded rather than quietly fixed.

BOTH DIRECTIONS, as §6.1 demands:

What fired and should have: every periodic job at its cadence; the DB-dump, tier and fill-watch legs; metrics-prune; the off-site nightly, which added the second snapshot with no prompting.

What should have fired and did not: nothing. Six registered jobs were never seen in the log — health-probes, status-refresh, ring-spill, deadapp-check, disk-health-check, selfupdate-check — and none of them is a finding. The first four are quiet by construction (scheduler.go:267quiet := job.Interval <= 30*time.Second, so a ≤30 s job never emits Running job:), and the last two run every 6 h, outside a ~4 h window. An absent log line is not evidence cuts this way too: I checked the source rather than filing four phantom defects.

What fired and should not have: nothing. No alert, notification, email or digest. The only WARN lines in five hours were three of mine — a failed login attempt, a CSRF token I mangled, and T1's deliberate opengist skip, which logged correctly.

One observation worth keeping: offbox-backup ticked twice, 62 s apart, and the snapshot count went 1 → 2, not 1 → 3. The second run was silently dropped by the single-flight — which for the NIGHTLY path is correct and deliberate (nobody asked; the next run retries). It is the same mechanism that, on the MANUAL path, produced R-234; v0.205.0 changes only the manual half and leaves this one silent, and this soak is live corroboration that the nightly half genuinely needs to stay quiet.


Phase B.2 — the destruction and the rebuild

All five §6.2 conditions were true and written down first (commits 2d2d8d3, f873c55, 502078b) before anything was destroyed.

02:40:31Z  pct destroy 9201 --purge     (guarded on hostname = finalwalk;
                                         demo-hp also has a guest 9201)
           both logical volumes removed; pct list empty
           /mnt/adatok and /mnt/mentes wiped to 20 K, MOUNTS LEFT IN PLACE
           — deliberately: the surviving raw mount IS the R-220 condition
02:40:56Z  felhom-host-install.sh v1.25.0, fetched live from felhom.eu/scripts/
02:43:28Z  Day-0 provision SUCCESS — 2 m 32 s

What the rebuild landed on — the second measurement of R-239:

before after
agent 0.127.0 0.127.0 — no downgrade, no hand upgrade
controller 0.203.0 0.203.0 — the same two-release gap

The agent half is exactly right: the reinstall neither downgraded nor needed a hand. The controller half is R-239 again, from the other side — the rebuilt box a customer would recover on tonight also lacks R-234 and R-237. Golden used: vzdump-lxc-9100-2026_08_06-23_58_32.tar.zst.

Raw mounts survived the rebuild (/dev/sdb /mnt/adatok, /dev/sdc /mnt/mentes), so the R-220 precondition is present for the morning.

Phase B.3 — THE HALT (§7)

The machine is at the claim screen and is waiting for the operator. Verified over HTTP from the appliance, with no guest shell:

GET /  ->  200,  <title>A szerver beállítása — Felhom
forms:  POST /claim   ·   POST /claim/request-new-code
claimed: null   ·   offbox config: absent (pristine rebuilt state)

A claim code has already been requested and emailed, through the customer-facing „Új kód kérése" path (POST /claim/request-new-code → 200, „Ha az e-mail cím regisztrálva van, elküldtük a kódot"), so the morning is paste a code, not request one and then paste it.

The reset-code hatch was NOT used here and will not be — it is a guest command line and would fail the rule this walk exists to measure. It was used once, in Phase A, where §3 permits it.

Honesty about the no-guest-command-line rule

The customer-journey steps from the destruction onward are driven over HTTP from the appliance to the guest's island address, which is what a browser would do. Some instrumentation readspct exec … python3 against settings.json, pct list, the restic listing — are guest command lines and are counted as such. They are not steps of the journey: none of them changed state, and none was needed to progress it. The distinction matters because conflating the two is how a walk claims a property it does not have.