1ff6f8f8e0
gates / gates (push) Successful in 12s
Three questions answered from the hub's own log, not inferred:
1. the rebuilt, still-unclaimed box DOES report (host-report + Received report)
2. it DOES declare offsite.state=needs_credential, and offsite-delivery correctly
declines once a minute, naming internal/offsiteheal as the owner
3. offsiteheal re-staged UNAIDED at 03:15Z, after two reports carried the
declaration, with no provider credential minted — about 32 minutes after the
rebuild, matching the documented 2x15-minute debounce
And then the box COLLECTED it on its own 5-minute tick:
[offsite-apply] credential retry: the staged credential was collected and the
tier applied
That success line shipped in v0.203.0 and this is the FIRST time it has been seen
live: yesterday's walk only produced its sibling before I intervened at 102s and
mistook my own button press for the cause — the error that produced R-236 and
forced its withdrawal. Here nobody touched anything and the box was not even
claimed. R-218's consume half, R-236's withdrawal and the previous walk's dead
end 1 are all settled by one unattended observation.
298 lines
18 KiB
Markdown
298 lines
18 KiB
Markdown
# THE FINAL WALK (R-201) — overnight, unattended — journal
|
||
|
||
**Venue:** `demo-hp` VM **324 `finalwalk-appliance`** (`finalwalk.felhom.eu` @ 192.168.0.142), guest
|
||
LXC **9201**, hub customer **`finalwalk`**, host id **`finalwalk-ed05d6`**. All three earlier venues
|
||
were torn down on 2026-08-06; nothing was reused except a **freed** Storage Box sub-account number.
|
||
|
||
**Written before the destruction, per §9.9.**
|
||
|
||
---
|
||
|
||
## THE HEADLINE FINDING — a fresh install does NOT get the fixes
|
||
|
||
| | landed on | vouched |
|
||
|---|---|---|
|
||
| agent | **0.127.0** ✅ | 0.127.0 |
|
||
| controller | **0.203.0** ❌ | golden 0.203.0 |
|
||
| newest released controller | **0.205.0** | — |
|
||
|
||
**No hand upgrade was needed and none was applied — that part works.** But the vouched golden still
|
||
bakes controller **0.203.0**, so a box installed tonight is **two releases behind**: it has neither
|
||
**R-237** (v0.204.0, the restore list keyed on the store) nor **R-234** (v0.205.0, the skipped-app
|
||
verdict and the single-flight message).
|
||
|
||
**This is not a regression — it is a delivery gap.** The fixes exist, are tested and are pushed; what
|
||
is missing is a golden carrying them and a vouch. Three of tonight's five checks measure exactly those
|
||
fixes, and they measure the OLD behaviour because that is what a customer receives.
|
||
|
||
Timeline: bind **21:56:42Z** → agent 0.127.0 ONLINE **21:59:01Z** (2 m 19 s) → controller 0.203.0
|
||
reporting **22:01:03Z** → guest 9201 running.
|
||
|
||
---
|
||
|
||
## Phase A — the fixture (six records, §4)
|
||
|
||
**1. Installed from the published ISO.** `felhom-installer-1.26.1-pve9.2-1.iso`, verified
|
||
**byte-identical to the published copy** (`sha256 f3cc86d5f0ec…59a6` local == iso.felhom.eu). Served
|
||
installer script tag confirmed `installer-v1.25.0` on both git-syncs.
|
||
|
||
All three known TUI traps handled: GRUB's graphical default (down+ret inside one command, terminal
|
||
installer first try); the Hungarian keymap switched to **U.S. English** — **positive control:** the
|
||
administrator email rendered `finalwalk@felhom.eu`, and `@` is `shift-2` on US vs `AltGr+V` on HU, the
|
||
only available evidence for the 24 masked password characters; and `--boot` set in its own `qm set`
|
||
with the ISO detached, both verified from `qm config` **before first boot** (`boot: order=scsi0`,
|
||
`ide2` lines = 0), with auto-reboot unchecked and confirmed `[ ]` with focus moved away.
|
||
|
||
Day-0 fired unaided: pairing code **`T63-485`**, matching MAC `bc:24:11:c6:73:d0`.
|
||
|
||
**2. Claimed** through the real `/claim` form. The reset-code hatch was used — **permitted in Phase A
|
||
by §3**, and it is a guest command line, so it is counted as such and does not touch the Phase-B rule.
|
||
|
||
**3. App + three sentinels.** `calibre-web` deployed with `HDD_PATH=/mnt/felhom-drives/adatok`,
|
||
`state: running` **and `health_probe.healthy: true`**.
|
||
|
||
| # | file | bytes | sha256 |
|
||
|---|---|---|---|
|
||
| A | `FINALWALK-SENTINEL-A.txt` | 57 | `863fa61c091c64488d8224b12f3915bfe264c46f9cad3d3d874b2ce725a1e5ee` |
|
||
| B | `FINALWALK-őrszem-ékezetes-árvíztűrő.txt` | 80 | `b56668663035a332ae9b8a77b5847c8310c7ff48c5c2068114060561389ccf23` |
|
||
| C | `FINALWALK-SENTINEL-C-12MB.bin` | 12 582 912 | `28630aa93119af0790b749671ef3896dbab88f8d239da8313ef5631fa068efda` |
|
||
|
||
Sentinel B's filename **as hex**, identical at source and on the box:
|
||
`46494e414c57414c4b2d c591 72737a656d2d c3a9 6b657a657465732d c3a1 7276 c3ad 7a74 c5b1 72 c591 2e747874`
|
||
— `ő é á í ű ő`, **no `efbfbd`**. Written from explicit bytes via `pct push` + a Python placer; no
|
||
shell chain ever saw the name.
|
||
|
||
**4. Escrow ceremony.** Preflight **6 of 6 green** (`pbs_storage_id` · `dr_tier` · `age_binary` ·
|
||
`hub_upload` · `staged_secret` · `sudo_grant`), `agent_supported: true`.
|
||
Result: `phase: done` · **`restic_pw_sealed: TRUE`** · `uploaded: true` · `entropy_bits: 129.24` ·
|
||
`key_fingerprint: 81:dc:91:ce:a1:d0:50:3a:…:68:d5`.
|
||
|
||
**R** was claimed ONE-SHOT and streamed file→file to `~/.config/finalwalk/R_finalwalk.txt` (`0600`,
|
||
**DooPlex only**); the guest and jump-host copies were `shred -u`'d and the raw response deleted. It
|
||
was never rendered, never an argument, never a log line. Shape only: **10 words, 85 characters**.
|
||
|
||
**5. Off-site backup, and the sentinels proved BY NAME.**
|
||
|
||
```
|
||
1da4f80d 2026-08-06 22:24:33 finalwalk [felhom-offbox, calibre-web]
|
||
/mnt/felhom-drives/adatok/userdata/media/books/FINALWALK-SENTINEL-A.txt 57
|
||
/mnt/felhom-drives/adatok/userdata/media/books/FINALWALK-SENTINEL-C-12MB.bin 12582912
|
||
/mnt/felhom-drives/adatok/userdata/media/books/FINALWALK-őrszem-ékezetes-árvíztűrő.txt 80
|
||
```
|
||
|
||
> **HARNESS FAULT, and the §4 gate is what caught it.** The FIRST run reported **`ok` with a 26.6 KB
|
||
> repository** — impossible for a 12 MB incompressible sentinel. Listing showed why: I had placed the
|
||
> files under `…/adatok/**felhom-data**/userdata/…`, while this box's namespace root is
|
||
> `/mnt/felhom-drives/adatok` **directly**, so they were never in the app's data at all. **The product
|
||
> was correct throughout** — it captured the unit and the declared mandatory path, including
|
||
> calibre-web's real `metadata.db`. Moved to the right path, re-run: 12.0 MB and all three listed. The
|
||
> gate's rule — *prove by listing, never by a green status* — is exactly what stopped a destruction
|
||
> that would have proven nothing.
|
||
|
||
**6. Pre-destruction truth.** Controller `0.203.0` · agent `0.127.0` · PBS wrapper **matches
|
||
vouched** · guests **1/1** · DR recipe **present** · key escrow **present** · snapshot `1da4f80d` ·
|
||
1 snapshot · **12.0 MB** · repo `sftp:u629488-sub4@…:/home/felhom-repo` port 23.
|
||
|
||
---
|
||
|
||
## The five checks (§5) — on controller 0.203.0, which is what a customer gets
|
||
|
||
| | check | observable | verdict |
|
||
|---|---|---|---|
|
||
| **T1** | selected app that cannot be captured (`opengist`, not deployed) → run | verdict **`ok`**; warning „Figyelmeztetés: 1 alkalmazásnak nincs elérhető mentése, ezek kimaradtak: opengist"; counters intact (1 snapshot, 12.0 MB, `last_success` advanced) | **old behaviour.** v0.205.0 also keeps `ok` for an *undeployed* app — deliberately — but says which, why and what to do. Here the message is a bare count with **no next step** |
|
||
| **T2** | two off-site runs back to back | run #1 „A távoli mentés elindult…"; run #2 **the same message**, `flash_error` count **0** | **FAILS.** This is R-234's measured cause, live. v0.205.0 answers „Már fut egy távoli mentés — ez a kérés nem indított újat…" as an error |
|
||
| **T3** | toggle future-backup **off** for an app that HAS a snapshot, then open the restore page | restore entries for calibre-web: **0**; wizard **302** away | **FAILS.** R-237 exactly: the customer's existing backup is hidden by a setting about the future |
|
||
| **T4** | off-site run with nothing selected | verdict `ok`; warning **„Sikeres — nincs mentésre jelölt alkalmazás"** | unchanged, as required — **and the wording is still *"successful"* beside *"nothing is covered"*. Filed (§8) |
|
||
| **T5** | full-restore wizard driven **as a browser does** | prepare → `&full_prep=calibre-web&full_size=12.8 MB` → following it revealed the confirm → commit → **`ok: true`**, „A(z) calibre-web teljes mentése visszaállítva ellenőrző mappába" | **PASSES.** R-238 confirmed a **harness artifact**, not a product defect — carrying the state the wizard hands back makes it complete |
|
||
|
||
**Live sentinels re-verified after T5: all three MATCH** (the verification restore writes to a
|
||
separate folder and left live data untouched).
|
||
|
||
**T2 and T3 are the same finding as the headline**, seen from the customer's side: the fixes are
|
||
written and pushed but not delivered, so tonight's box still exhibits both defects.
|
||
|
||
---
|
||
|
||
## Phase B.1 — the soak
|
||
|
||
**Window 1: 22:44:05Z → 01:54:03Z (3 h 10 m), untouched.**
|
||
|
||
Every periodic job fired at exactly its declared cadence, which is the positive control that the
|
||
scheduler was running at all:
|
||
|
||
| job | cadence | fired | expected in 190 min |
|
||
|---|---|---|---|
|
||
| `agent-channel-health` | 1 m | 185 | ~190 |
|
||
| `stack-scan` | 2 m | 108 | ~95 |
|
||
| `system-health` · `backup-cache` · **`offsite-credential-retry`** | 5 m | 44 each | ~38 |
|
||
| `hub-report` | 15 m | 14 | ~13 |
|
||
| `tier` · `fill-watch` · `db-dump` | — | 1 each | — |
|
||
|
||
**`offsite-credential-retry` ran 44 times and did no work and said nothing.** That is R-218's asserted
|
||
healthy-box behaviour — the job exists, ticks, completes in 0 s, and stays silent because the
|
||
declaration it keys on is false. A box that had never seen the defect behaves exactly as designed.
|
||
|
||
**No alert, notification or digest fired. Nothing on the must-not list fired.** The off-site state was
|
||
unchanged throughout (`last_run` fixed at `22:38:54Z`, `snaps=1`, `12.0 MB`).
|
||
|
||
> **AND THAT LAST LINE IS NOT A FINDING — IT IS MY PLANNING ERROR, STATED AS SUCH.** The daily jobs run
|
||
> on the CONTROLLER's clock, and **the guest is UTC while the appliance is CEST** (`date +%Z`:
|
||
> guest `UTC`, appliance `CEST`). The nightly local backup (~02:30) and off-site (~04:15) therefore
|
||
> fall at **02:30Z and 04:15Z**, and I sized the window against CEST — so it closed at 01:54Z, before
|
||
> either. Reporting "the nightly did not fire" as a defect would have been a false finding produced by
|
||
> a badly-chosen window, which is precisely the shape §6.1 warns about in the other direction.
|
||
>
|
||
> **Window 2 (corrective): 01:56Z → ~02:35Z**, to cover the 02:30Z local backup. The 04:15Z off-site
|
||
> nightly is deliberately **NOT** covered: finishing the walk and leaving the machine at the claim
|
||
> screen before 07:00 is the primary deliverable (§11.1), and waiting for it would have put the
|
||
> destruction at ~06:35 CEST with no margin. **Recorded as not run, with the reason** — the off-site
|
||
> tier was exercised four times manually tonight instead, including a full listing by name.
|
||
|
||
**Window 2: 01:56:28Z → 02:36:45Z — and the nightly DID fire, unprompted.**
|
||
|
||
```
|
||
00:30:25Z db-dump
|
||
01:30:19Z tier + fill-watch ← the local tier legs
|
||
02:00:35Z metrics-prune
|
||
02:15:03Z offbox-backup ← the off-site nightly
|
||
02:16:05Z offbox-backup
|
||
offbox: snaps 1 → 2, last_run 22:38:54Z → 02:15:24Z
|
||
```
|
||
|
||
So the soak covered a genuine scheduled cycle after all: the local legs in window 1 and the off-site
|
||
nightly in window 2. **My 04:15 prediction was wrong in the other direction — it fires at ~02:15
|
||
controller time.** Both the prediction and the correction are recorded rather than quietly fixed.
|
||
|
||
**BOTH DIRECTIONS, as §6.1 demands:**
|
||
|
||
*What fired and should have:* every periodic job at its cadence; the DB-dump, tier and fill-watch
|
||
legs; metrics-prune; the off-site nightly, which added the second snapshot with no prompting.
|
||
|
||
*What should have fired and did not:* **nothing.** Six registered jobs were never seen in the log —
|
||
`health-probes`, `status-refresh`, `ring-spill`, `deadapp-check`, `disk-health-check`,
|
||
`selfupdate-check` — and **none of them is a finding.** The first four are **quiet by construction**
|
||
(`scheduler.go:267` — `quiet := job.Interval <= 30*time.Second`, so a ≤30 s job never emits
|
||
`Running job:`), and the last two run every 6 h, outside a ~4 h window. *An absent log line is not
|
||
evidence* cuts this way too: I checked the source rather than filing four phantom defects.
|
||
|
||
*What fired and should not have:* nothing. No alert, notification, email or digest. The only `WARN`
|
||
lines in five hours were **three of mine** — a failed login attempt, a CSRF token I mangled, and T1's
|
||
deliberate `opengist` skip, which logged correctly.
|
||
|
||
**One observation worth keeping:** `offbox-backup` ticked **twice, 62 s apart**, and the snapshot count
|
||
went 1 → 2, not 1 → 3. The second run was silently dropped by the single-flight — which for the
|
||
NIGHTLY path is correct and deliberate (nobody asked; the next run retries). It is the same mechanism
|
||
that, on the MANUAL path, produced R-234; v0.205.0 changes only the manual half and leaves this one
|
||
silent, and this soak is live corroboration that the nightly half genuinely needs to stay quiet.
|
||
|
||
---
|
||
|
||
## Phase B.2 — the destruction and the rebuild
|
||
|
||
All five §6.2 conditions were true and **written down first** (commits `2d2d8d3`, `f873c55`,
|
||
`502078b`) before anything was destroyed.
|
||
|
||
```
|
||
02:40:31Z pct destroy 9201 --purge (guarded on hostname = finalwalk;
|
||
demo-hp also has a guest 9201)
|
||
both logical volumes removed; pct list empty
|
||
/mnt/adatok and /mnt/mentes wiped to 20 K, MOUNTS LEFT IN PLACE
|
||
— deliberately: the surviving raw mount IS the R-220 condition
|
||
02:40:56Z felhom-host-install.sh v1.25.0, fetched live from felhom.eu/scripts/
|
||
02:43:28Z Day-0 provision SUCCESS — 2 m 32 s
|
||
```
|
||
|
||
**What the rebuild landed on — the second measurement of R-239:**
|
||
|
||
| | before | after |
|
||
|---|---|---|
|
||
| agent | 0.127.0 | **0.127.0** — no downgrade, no hand upgrade |
|
||
| controller | 0.203.0 | **0.203.0** — the same two-release gap |
|
||
|
||
The agent half is exactly right: the reinstall neither downgraded nor needed a hand. The controller
|
||
half is R-239 again, from the other side — **the rebuilt box a customer would recover on tonight also
|
||
lacks R-234 and R-237.** Golden used: `vzdump-lxc-9100-2026_08_06-23_58_32.tar.zst`.
|
||
|
||
Raw mounts survived the rebuild (`/dev/sdb /mnt/adatok`, `/dev/sdc /mnt/mentes`), so the R-220
|
||
precondition is present for the morning.
|
||
|
||
## Phase B.3 — THE HALT (§7)
|
||
|
||
**The machine is at the claim screen and is waiting for the operator.** Verified over HTTP **from the
|
||
appliance**, with no guest shell:
|
||
|
||
```
|
||
GET / -> 200, <title>A szerver beállítása — Felhom
|
||
forms: POST /claim · POST /claim/request-new-code
|
||
claimed: null · offbox config: absent (pristine rebuilt state)
|
||
```
|
||
|
||
**A claim code has already been requested and emailed**, through the customer-facing „Új kód kérése"
|
||
path (`POST /claim/request-new-code` → 200, „Ha az e-mail cím regisztrálva van, elküldtük a kódot"),
|
||
so the morning is *paste a code*, not *request one and then paste it*.
|
||
|
||
**The reset-code hatch was NOT used here and will not be** — it is a guest command line and would
|
||
fail the rule this walk exists to measure. It was used once, in Phase A, where §3 permits it.
|
||
|
||
### Honesty about the no-guest-command-line rule
|
||
|
||
The customer-journey steps from the destruction onward are driven over **HTTP from the appliance to
|
||
the guest's island address**, which is what a browser would do. Some **instrumentation reads** —
|
||
`pct exec … python3` against `settings.json`, `pct list`, the restic listing — are guest command lines
|
||
and are counted as such. They are not steps of the journey: none of them changed state, and none was
|
||
needed to progress it. The distinction matters because conflating the two is how a walk claims a
|
||
property it does not have.
|
||
|
||
### §7's observation — the rebuilt, still-UNCLAIMED box
|
||
|
||
Three questions, all answered from the hub's own log rather than inferred:
|
||
|
||
**1. Does it report while unclaimed? YES.**
|
||
`host-report from finalwalk-ed05d6 (1 guests, 4 storage targets, 1 backups, 0 restore-tests,
|
||
1 pbs-snapshots, 12907 bytes)` at 05:12:04 local, and `Received report from finalwalk` at 05:13:03.
|
||
|
||
**2. Does it declare the credential need? YES**, and the hub's other mechanism correctly declines to
|
||
act on it — logged once a minute, from 02:59Z onward:
|
||
|
||
> `offsite-delivery: finalwalk: self-heal REFUSED — the box DECLARES offsite.state=needs_credential;
|
||
> internal/offsiteheal owns this remediation (it re-stages the stored credential before minting).
|
||
> A second mechanism minting here would double-issue.`
|
||
|
||
**3. Did the hub stage one unaided? YES, at 05:14:57 local (03:15Z):**
|
||
|
||
> `offsiteheal: re-staged the stored one-time offsite secret for customer finalwalk (declared
|
||
> needs_credential across 2 reports) — the box re-consumes on its next cycle; no provider credential
|
||
> was minted`
|
||
|
||
That is **~32 minutes after the rebuild**, consistent with the documented 2 × 15-minute report
|
||
debounce, and **nobody touched anything**.
|
||
|
||
**This is a stronger case than yesterday's.** On 2026-08-06 I pressed Re-issue 102 seconds after the
|
||
self-heal had already fired and mistook my own button for the cause — the error that produced R-236
|
||
and forced its withdrawal. Here the box was **pristine, rebuilt and not even claimed**, no operator
|
||
action of any kind was taken, and the chain ran end to end on its own. R-236's withdrawal is now
|
||
confirmed on the exact shape it was filed against.
|
||
|
||
**4. And then the box COLLECTED it, on its own tick — the full loop, unaided:**
|
||
|
||
```
|
||
[offsite-apply] settle-gate: GO — at/above floor 0.156.0 (we are 0.203.0), no managed update running
|
||
[offsite-apply] offsite configured for u629488-sub4@…:/home/felhom-repo (pending key escrow)
|
||
[offsite-apply] credential retry: the staged credential was collected and the tier applied
|
||
```
|
||
|
||
The box's own settings went from **no `offbox` object at all** to `enabled: true`, host, user and
|
||
`repo_path` populated, `escrow_state: pending`.
|
||
|
||
**This is the first time that success line has ever been observed live.** It shipped in controller
|
||
v0.203.0; the 2026-08-06 walk only ever produced its sibling — *"no unconsumed offsite password …
|
||
(the box still declares a need; retrying)"* — before an operator intervened. Here the whole chain ran
|
||
with **zero human action, on a box that has not even been claimed yet**:
|
||
|
||
> declare → `offsite-delivery` declines and names the owner → `offsiteheal` re-stages after two
|
||
> reports → the box's 5-minute retry collects it → the tier is applied.
|
||
|
||
**That is R-218's consume half, R-236's withdrawal, and the previous walk's dead end 1, all settled by
|
||
one unattended observation.** By the time the operator claims this machine in the morning, its
|
||
off-site tier is already up — which is exactly the property the recovery journey needed and never had.
|