1b490c8cbf
gates / gates (push) Successful in 27s
Destroyed 02:40:31Z (guarded on hostname — demo-hp also has a guest 9201), drives wiped to 20K with the mounts deliberately left in place because the surviving raw mount IS the R-220 condition. Reinstalled through the published day-0 path, installer v1.25.0 fetched live; Day-0 provision SUCCESS in 2m32s. R-239 measured a second time, from the other side: the rebuild landed on agent 0.127.0 (no downgrade, no hand upgrade — that half is right) and controller 0.203.0. The box a customer would recover on tonight also lacks R-234 and R-237. The machine is AT THE CLAIM SCREEN awaiting the operator. A claim code has already been requested through the customer-facing path and emailed, so the morning is paste-a-code rather than request-then-paste. The reset-code hatch was NOT used and will not be: it is a guest command line and would fail the rule the walk measures. Stated plainly in the journal: journey steps from the destruction onward are driven over HTTP from the appliance to the guest's island address, as a browser would; some instrumentation reads are guest command lines and are counted as such, but none changed state or was needed to progress the journey.
245 lines
14 KiB
Markdown
245 lines
14 KiB
Markdown
# THE FINAL WALK (R-201) — overnight, unattended — journal
|
|
|
|
**Venue:** `demo-hp` VM **324 `finalwalk-appliance`** (`finalwalk.felhom.eu` @ 192.168.0.142), guest
|
|
LXC **9201**, hub customer **`finalwalk`**, host id **`finalwalk-ed05d6`**. All three earlier venues
|
|
were torn down on 2026-08-06; nothing was reused except a **freed** Storage Box sub-account number.
|
|
|
|
**Written before the destruction, per §9.9.**
|
|
|
|
---
|
|
|
|
## THE HEADLINE FINDING — a fresh install does NOT get the fixes
|
|
|
|
| | landed on | vouched |
|
|
|---|---|---|
|
|
| agent | **0.127.0** ✅ | 0.127.0 |
|
|
| controller | **0.203.0** ❌ | golden 0.203.0 |
|
|
| newest released controller | **0.205.0** | — |
|
|
|
|
**No hand upgrade was needed and none was applied — that part works.** But the vouched golden still
|
|
bakes controller **0.203.0**, so a box installed tonight is **two releases behind**: it has neither
|
|
**R-237** (v0.204.0, the restore list keyed on the store) nor **R-234** (v0.205.0, the skipped-app
|
|
verdict and the single-flight message).
|
|
|
|
**This is not a regression — it is a delivery gap.** The fixes exist, are tested and are pushed; what
|
|
is missing is a golden carrying them and a vouch. Three of tonight's five checks measure exactly those
|
|
fixes, and they measure the OLD behaviour because that is what a customer receives.
|
|
|
|
Timeline: bind **21:56:42Z** → agent 0.127.0 ONLINE **21:59:01Z** (2 m 19 s) → controller 0.203.0
|
|
reporting **22:01:03Z** → guest 9201 running.
|
|
|
|
---
|
|
|
|
## Phase A — the fixture (six records, §4)
|
|
|
|
**1. Installed from the published ISO.** `felhom-installer-1.26.1-pve9.2-1.iso`, verified
|
|
**byte-identical to the published copy** (`sha256 f3cc86d5f0ec…59a6` local == iso.felhom.eu). Served
|
|
installer script tag confirmed `installer-v1.25.0` on both git-syncs.
|
|
|
|
All three known TUI traps handled: GRUB's graphical default (down+ret inside one command, terminal
|
|
installer first try); the Hungarian keymap switched to **U.S. English** — **positive control:** the
|
|
administrator email rendered `finalwalk@felhom.eu`, and `@` is `shift-2` on US vs `AltGr+V` on HU, the
|
|
only available evidence for the 24 masked password characters; and `--boot` set in its own `qm set`
|
|
with the ISO detached, both verified from `qm config` **before first boot** (`boot: order=scsi0`,
|
|
`ide2` lines = 0), with auto-reboot unchecked and confirmed `[ ]` with focus moved away.
|
|
|
|
Day-0 fired unaided: pairing code **`T63-485`**, matching MAC `bc:24:11:c6:73:d0`.
|
|
|
|
**2. Claimed** through the real `/claim` form. The reset-code hatch was used — **permitted in Phase A
|
|
by §3**, and it is a guest command line, so it is counted as such and does not touch the Phase-B rule.
|
|
|
|
**3. App + three sentinels.** `calibre-web` deployed with `HDD_PATH=/mnt/felhom-drives/adatok`,
|
|
`state: running` **and `health_probe.healthy: true`**.
|
|
|
|
| # | file | bytes | sha256 |
|
|
|---|---|---|---|
|
|
| A | `FINALWALK-SENTINEL-A.txt` | 57 | `863fa61c091c64488d8224b12f3915bfe264c46f9cad3d3d874b2ce725a1e5ee` |
|
|
| B | `FINALWALK-őrszem-ékezetes-árvíztűrő.txt` | 80 | `b56668663035a332ae9b8a77b5847c8310c7ff48c5c2068114060561389ccf23` |
|
|
| C | `FINALWALK-SENTINEL-C-12MB.bin` | 12 582 912 | `28630aa93119af0790b749671ef3896dbab88f8d239da8313ef5631fa068efda` |
|
|
|
|
Sentinel B's filename **as hex**, identical at source and on the box:
|
|
`46494e414c57414c4b2d c591 72737a656d2d c3a9 6b657a657465732d c3a1 7276 c3ad 7a74 c5b1 72 c591 2e747874`
|
|
— `ő é á í ű ő`, **no `efbfbd`**. Written from explicit bytes via `pct push` + a Python placer; no
|
|
shell chain ever saw the name.
|
|
|
|
**4. Escrow ceremony.** Preflight **6 of 6 green** (`pbs_storage_id` · `dr_tier` · `age_binary` ·
|
|
`hub_upload` · `staged_secret` · `sudo_grant`), `agent_supported: true`.
|
|
Result: `phase: done` · **`restic_pw_sealed: TRUE`** · `uploaded: true` · `entropy_bits: 129.24` ·
|
|
`key_fingerprint: 81:dc:91:ce:a1:d0:50:3a:…:68:d5`.
|
|
|
|
**R** was claimed ONE-SHOT and streamed file→file to `~/.config/finalwalk/R_finalwalk.txt` (`0600`,
|
|
**DooPlex only**); the guest and jump-host copies were `shred -u`'d and the raw response deleted. It
|
|
was never rendered, never an argument, never a log line. Shape only: **10 words, 85 characters**.
|
|
|
|
**5. Off-site backup, and the sentinels proved BY NAME.**
|
|
|
|
```
|
|
1da4f80d 2026-08-06 22:24:33 finalwalk [felhom-offbox, calibre-web]
|
|
/mnt/felhom-drives/adatok/userdata/media/books/FINALWALK-SENTINEL-A.txt 57
|
|
/mnt/felhom-drives/adatok/userdata/media/books/FINALWALK-SENTINEL-C-12MB.bin 12582912
|
|
/mnt/felhom-drives/adatok/userdata/media/books/FINALWALK-őrszem-ékezetes-árvíztűrő.txt 80
|
|
```
|
|
|
|
> **HARNESS FAULT, and the §4 gate is what caught it.** The FIRST run reported **`ok` with a 26.6 KB
|
|
> repository** — impossible for a 12 MB incompressible sentinel. Listing showed why: I had placed the
|
|
> files under `…/adatok/**felhom-data**/userdata/…`, while this box's namespace root is
|
|
> `/mnt/felhom-drives/adatok` **directly**, so they were never in the app's data at all. **The product
|
|
> was correct throughout** — it captured the unit and the declared mandatory path, including
|
|
> calibre-web's real `metadata.db`. Moved to the right path, re-run: 12.0 MB and all three listed. The
|
|
> gate's rule — *prove by listing, never by a green status* — is exactly what stopped a destruction
|
|
> that would have proven nothing.
|
|
|
|
**6. Pre-destruction truth.** Controller `0.203.0` · agent `0.127.0` · PBS wrapper **matches
|
|
vouched** · guests **1/1** · DR recipe **present** · key escrow **present** · snapshot `1da4f80d` ·
|
|
1 snapshot · **12.0 MB** · repo `sftp:u629488-sub4@…:/home/felhom-repo` port 23.
|
|
|
|
---
|
|
|
|
## The five checks (§5) — on controller 0.203.0, which is what a customer gets
|
|
|
|
| | check | observable | verdict |
|
|
|---|---|---|---|
|
|
| **T1** | selected app that cannot be captured (`opengist`, not deployed) → run | verdict **`ok`**; warning „Figyelmeztetés: 1 alkalmazásnak nincs elérhető mentése, ezek kimaradtak: opengist"; counters intact (1 snapshot, 12.0 MB, `last_success` advanced) | **old behaviour.** v0.205.0 also keeps `ok` for an *undeployed* app — deliberately — but says which, why and what to do. Here the message is a bare count with **no next step** |
|
|
| **T2** | two off-site runs back to back | run #1 „A távoli mentés elindult…"; run #2 **the same message**, `flash_error` count **0** | **FAILS.** This is R-234's measured cause, live. v0.205.0 answers „Már fut egy távoli mentés — ez a kérés nem indított újat…" as an error |
|
|
| **T3** | toggle future-backup **off** for an app that HAS a snapshot, then open the restore page | restore entries for calibre-web: **0**; wizard **302** away | **FAILS.** R-237 exactly: the customer's existing backup is hidden by a setting about the future |
|
|
| **T4** | off-site run with nothing selected | verdict `ok`; warning **„Sikeres — nincs mentésre jelölt alkalmazás"** | unchanged, as required — **and the wording is still *"successful"* beside *"nothing is covered"*. Filed (§8) |
|
|
| **T5** | full-restore wizard driven **as a browser does** | prepare → `&full_prep=calibre-web&full_size=12.8 MB` → following it revealed the confirm → commit → **`ok: true`**, „A(z) calibre-web teljes mentése visszaállítva ellenőrző mappába" | **PASSES.** R-238 confirmed a **harness artifact**, not a product defect — carrying the state the wizard hands back makes it complete |
|
|
|
|
**Live sentinels re-verified after T5: all three MATCH** (the verification restore writes to a
|
|
separate folder and left live data untouched).
|
|
|
|
**T2 and T3 are the same finding as the headline**, seen from the customer's side: the fixes are
|
|
written and pushed but not delivered, so tonight's box still exhibits both defects.
|
|
|
|
---
|
|
|
|
## Phase B.1 — the soak
|
|
|
|
**Window 1: 22:44:05Z → 01:54:03Z (3 h 10 m), untouched.**
|
|
|
|
Every periodic job fired at exactly its declared cadence, which is the positive control that the
|
|
scheduler was running at all:
|
|
|
|
| job | cadence | fired | expected in 190 min |
|
|
|---|---|---|---|
|
|
| `agent-channel-health` | 1 m | 185 | ~190 |
|
|
| `stack-scan` | 2 m | 108 | ~95 |
|
|
| `system-health` · `backup-cache` · **`offsite-credential-retry`** | 5 m | 44 each | ~38 |
|
|
| `hub-report` | 15 m | 14 | ~13 |
|
|
| `tier` · `fill-watch` · `db-dump` | — | 1 each | — |
|
|
|
|
**`offsite-credential-retry` ran 44 times and did no work and said nothing.** That is R-218's asserted
|
|
healthy-box behaviour — the job exists, ticks, completes in 0 s, and stays silent because the
|
|
declaration it keys on is false. A box that had never seen the defect behaves exactly as designed.
|
|
|
|
**No alert, notification or digest fired. Nothing on the must-not list fired.** The off-site state was
|
|
unchanged throughout (`last_run` fixed at `22:38:54Z`, `snaps=1`, `12.0 MB`).
|
|
|
|
> **AND THAT LAST LINE IS NOT A FINDING — IT IS MY PLANNING ERROR, STATED AS SUCH.** The daily jobs run
|
|
> on the CONTROLLER's clock, and **the guest is UTC while the appliance is CEST** (`date +%Z`:
|
|
> guest `UTC`, appliance `CEST`). The nightly local backup (~02:30) and off-site (~04:15) therefore
|
|
> fall at **02:30Z and 04:15Z**, and I sized the window against CEST — so it closed at 01:54Z, before
|
|
> either. Reporting "the nightly did not fire" as a defect would have been a false finding produced by
|
|
> a badly-chosen window, which is precisely the shape §6.1 warns about in the other direction.
|
|
>
|
|
> **Window 2 (corrective): 01:56Z → ~02:35Z**, to cover the 02:30Z local backup. The 04:15Z off-site
|
|
> nightly is deliberately **NOT** covered: finishing the walk and leaving the machine at the claim
|
|
> screen before 07:00 is the primary deliverable (§11.1), and waiting for it would have put the
|
|
> destruction at ~06:35 CEST with no margin. **Recorded as not run, with the reason** — the off-site
|
|
> tier was exercised four times manually tonight instead, including a full listing by name.
|
|
|
|
**Window 2: 01:56:28Z → 02:36:45Z — and the nightly DID fire, unprompted.**
|
|
|
|
```
|
|
00:30:25Z db-dump
|
|
01:30:19Z tier + fill-watch ← the local tier legs
|
|
02:00:35Z metrics-prune
|
|
02:15:03Z offbox-backup ← the off-site nightly
|
|
02:16:05Z offbox-backup
|
|
offbox: snaps 1 → 2, last_run 22:38:54Z → 02:15:24Z
|
|
```
|
|
|
|
So the soak covered a genuine scheduled cycle after all: the local legs in window 1 and the off-site
|
|
nightly in window 2. **My 04:15 prediction was wrong in the other direction — it fires at ~02:15
|
|
controller time.** Both the prediction and the correction are recorded rather than quietly fixed.
|
|
|
|
**BOTH DIRECTIONS, as §6.1 demands:**
|
|
|
|
*What fired and should have:* every periodic job at its cadence; the DB-dump, tier and fill-watch
|
|
legs; metrics-prune; the off-site nightly, which added the second snapshot with no prompting.
|
|
|
|
*What should have fired and did not:* **nothing.** Six registered jobs were never seen in the log —
|
|
`health-probes`, `status-refresh`, `ring-spill`, `deadapp-check`, `disk-health-check`,
|
|
`selfupdate-check` — and **none of them is a finding.** The first four are **quiet by construction**
|
|
(`scheduler.go:267` — `quiet := job.Interval <= 30*time.Second`, so a ≤30 s job never emits
|
|
`Running job:`), and the last two run every 6 h, outside a ~4 h window. *An absent log line is not
|
|
evidence* cuts this way too: I checked the source rather than filing four phantom defects.
|
|
|
|
*What fired and should not have:* nothing. No alert, notification, email or digest. The only `WARN`
|
|
lines in five hours were **three of mine** — a failed login attempt, a CSRF token I mangled, and T1's
|
|
deliberate `opengist` skip, which logged correctly.
|
|
|
|
**One observation worth keeping:** `offbox-backup` ticked **twice, 62 s apart**, and the snapshot count
|
|
went 1 → 2, not 1 → 3. The second run was silently dropped by the single-flight — which for the
|
|
NIGHTLY path is correct and deliberate (nobody asked; the next run retries). It is the same mechanism
|
|
that, on the MANUAL path, produced R-234; v0.205.0 changes only the manual half and leaves this one
|
|
silent, and this soak is live corroboration that the nightly half genuinely needs to stay quiet.
|
|
|
|
---
|
|
|
|
## Phase B.2 — the destruction and the rebuild
|
|
|
|
All five §6.2 conditions were true and **written down first** (commits `2d2d8d3`, `f873c55`,
|
|
`502078b`) before anything was destroyed.
|
|
|
|
```
|
|
02:40:31Z pct destroy 9201 --purge (guarded on hostname = finalwalk;
|
|
demo-hp also has a guest 9201)
|
|
both logical volumes removed; pct list empty
|
|
/mnt/adatok and /mnt/mentes wiped to 20 K, MOUNTS LEFT IN PLACE
|
|
— deliberately: the surviving raw mount IS the R-220 condition
|
|
02:40:56Z felhom-host-install.sh v1.25.0, fetched live from felhom.eu/scripts/
|
|
02:43:28Z Day-0 provision SUCCESS — 2 m 32 s
|
|
```
|
|
|
|
**What the rebuild landed on — the second measurement of R-239:**
|
|
|
|
| | before | after |
|
|
|---|---|---|
|
|
| agent | 0.127.0 | **0.127.0** — no downgrade, no hand upgrade |
|
|
| controller | 0.203.0 | **0.203.0** — the same two-release gap |
|
|
|
|
The agent half is exactly right: the reinstall neither downgraded nor needed a hand. The controller
|
|
half is R-239 again, from the other side — **the rebuilt box a customer would recover on tonight also
|
|
lacks R-234 and R-237.** Golden used: `vzdump-lxc-9100-2026_08_06-23_58_32.tar.zst`.
|
|
|
|
Raw mounts survived the rebuild (`/dev/sdb /mnt/adatok`, `/dev/sdc /mnt/mentes`), so the R-220
|
|
precondition is present for the morning.
|
|
|
|
## Phase B.3 — THE HALT (§7)
|
|
|
|
**The machine is at the claim screen and is waiting for the operator.** Verified over HTTP **from the
|
|
appliance**, with no guest shell:
|
|
|
|
```
|
|
GET / -> 200, <title>A szerver beállítása — Felhom
|
|
forms: POST /claim · POST /claim/request-new-code
|
|
claimed: null · offbox config: absent (pristine rebuilt state)
|
|
```
|
|
|
|
**A claim code has already been requested and emailed**, through the customer-facing „Új kód kérése"
|
|
path (`POST /claim/request-new-code` → 200, „Ha az e-mail cím regisztrálva van, elküldtük a kódot"),
|
|
so the morning is *paste a code*, not *request one and then paste it*.
|
|
|
|
**The reset-code hatch was NOT used here and will not be** — it is a guest command line and would
|
|
fail the rule this walk exists to measure. It was used once, in Phase A, where §3 permits it.
|
|
|
|
### Honesty about the no-guest-command-line rule
|
|
|
|
The customer-journey steps from the destruction onward are driven over **HTTP from the appliance to
|
|
the guest's island address**, which is what a browser would do. Some **instrumentation reads** —
|
|
`pct exec … python3` against `settings.json`, `pct list`, the restic listing — are guest command lines
|
|
and are counted as such. They are not steps of the journey: none of them changed state, and none was
|
|
needed to progress it. The distinction matters because conflating the two is how a walk claims a
|
|
property it does not have.
|