Registers and evidence for the campaign and its fix pass. OPEN-ITEMS: R-214..R-223. Six SHIPPED (R-215/216/217/218/219/222); three deliberately still open and each blocks a real flow (R-214 console banner, R-220 drives unenrollable after a rebuild, R-221 a rebuilt box cannot run the escrow ceremony); R-223 minted and WAITING-ON-OPERATOR (vouch agent 0.125.0). R-213 and R-202 untouched. Capability map: a new row for the customer's UNAIDED journey, recorded FAILED and staying failed until a re-walk passes — fixes are not a journey. The existing rebuild row is corrected where it said R-198's retention was unit-proven only: it was proven in production on the first supersession since the fix, identity_blob retained at 572 B byte-length exact. CLAUDE.md comment-vs-code table: eighth entry — ResolveManagedFloor, the first where the false invariant was a GUARD rather than a comment alone. STATUS: the headline is now "the backup promise is proved, the recovery journey is not", and the one thing waiting on the operator.
59 KiB
CAMPAIGN 11 — step-by-step journal
Every observable, in the order it was taken. Times are CEST (demo-hp's clock) unless marked UTC.
Pre-flight — baselines re-read on arrival (2026-08-05)
| What | Value | How read |
|---|---|---|
felhom-controller main |
v0.200.0 @ a315d623b857 |
git rev-parse HEAD == origin/main, tree clean |
felhom-agent main |
v0.125.0 @ 0404f60e6a7b |
same |
felhom.eu main |
hub v0.96.0 @ f45b1f6761ff |
same |
| hub, LIVE | gitea.dooplex.hu/admin/felhom-hub:0.96.0 |
kubectl -n felhom-system get deploy hub -o jsonpath=… |
| ISO, published | felhom-installer-1.26.1-pve9.2-1.iso · 1 705 322 496 B · sha f3cc86d5f0ec…59a6 |
curl -I on iso.felhom.eu + published .sha256 |
| ISO, local copy on demo-hp | sha f3cc86d5f0ec…59a6, 1 705 322 496 B — identical |
sha256sum on demo-hp |
| Highest register ID | R-213 | fresh grep -rhoE 'R-[0-9]{1,3}' --include=*.md over all four repos |
All three cited commits match the brief exactly. Two brief assumptions were stale and reading fresh caught both:
- The brief says installer 1.25.0; the published artifact is 1.26.1 (since 2026-07-31). Same class Campaign 10 caught, same file.
- The brief's §11 says "Five scratch customers have accumulated". There are none. The hub holds
exactly five customer records and every one is legitimate:
david(a real prospect),demo-felhom,demo-hp,drill-r50(the fixture, BLOCKED),peti-felhom(the external tester).c10-soak,sess-c,sess-d,sess-fare all gone.
The finding that decided the campaign's shape, found before anything was built
Day-0 artifact manifest, read from the hub's /configuration:
| Vouched | Newest that exists | main |
|
|---|---|---|---|
| golden (bakes the controller) | 0.192.0 | 0.192.0 (no newer golden exists) | controller 0.200.0 |
| agent | 0.120.0 | 0.125.0 (published in Gitea) | 0.125.0 |
| global controller floor | 0.156.0 | — | — |
The controller's auto-update target is always the floor, never latest
(selfupdate/updater.go:64). So a box installed from today's published ISO lands on controller
0.192.0 and stays there — which is before R-203's fix (0.197.0), before the credential self-heal
(0.199.0), and before the recovery screen this campaign exists to test (0.200.0).
Both demo boxes run agent 0.125.0 while the manifest vouches 0.120.0 — R-121's drift family, live, on the day-0 channel rather than on a box.
Operator ruling, 2026-08-05: install stock, record what a customer gets today, then raise the
per-customer floor (configs.go:1229) to 0.200.0 and let the box converge. Per-customer means the
other four customers are untouched.
Two further operator rulings taken before building
- ep0 + the Hetzner Storage Box are written to, additively. Off-site hard-requires the DR tier
(
hub/internal/web/configs.go:1298), and the DR tier provisions a PBS namespace + token on ep0 via tenantsync plus a WireGuard peer; enabling off-site provisions a Storage Box sub-account. All three are creates on the ordinary customer path — nothing existing is modified or deleted — and teardown removes the campaign's own. Campaign 10 avoided this only by running with Tier 3 OFF, which is not available to a campaign about off-site recovery. The brief's I7 wording ("ep0 read-only") was therefore relaxed deliberately, by ruling, and is recorded here rather than quietly widened. - Recovery codes travel file→file into a
0600file, are moved by the operator into their own store, and are then shredded with a planted-copy positive control.
Venue
| Host | demo-hp (HP t740), Tier 0, the designated drill host. Reached by SSH key, first try — R-129 stands |
| VM | 321 c11-appliance — q35/OVMF (pre-enrolled-keys=0), 4 cores, 8 GB, cpu=host |
| Disks | scsi0 200 G system · scsi1 50 G · scsi2 50 G, qcow2 |
| Storage | c11-scratch, dir at /mnt/nvme-1tb — the mount ROOT (a subdirectory fails the agent's exactMount check; Campaign 10 §1) |
| Box | c11.felhom.eu @ 192.168.0.105/24, gw/DNS 192.168.0.1, on demo-hp's vmbr0 |
| Hub customer | c11 "Campaign 11", domain c11.felhom.eu, DR tier ON, off-site initially OFF |
| Host id | c11-36d660 |
| Untouched | drill-r50 (VM 300, stopped), guest 9201 on both demo boxes, local-lvm, felhom-backup, every other hub customer |
Harness constraint — the appliance is NOT reachable from DooPlex
DooPlex is 192.168.0.180/24 on eno1 and demo-hp is a different site with an overlapping
192.168.0.0/24 — ssh demo-hp resolves to Tailscale 100.76.96.79, and DooPlex's ARP for
192.168.0.105 is INCOMPLETE while demo-hp pings it in 0.4 ms. Every HTTP call to the appliance
therefore goes DooPlex → (Tailscale) → demo-hp → vmbr0 → appliance, through a SOCKS proxy
(ssh -N -D 11080 demo-hp). No server logic is skipped — this is a transport, not a shortcut.
Harness, not product.
Phase 0
Step 1 — install from the published ISO, through the real day-0 path
Driven blind through the Terminal UI (qm monitor screendump → PNG → read visually; qm sendkey
for input). Campaign 10's three traps, all handled:
- GRUB's default is the graphical entry and the graphical entry is undrivable (monitor mouse
events do not move the guest cursor). The menu appears ~6 s in with a 15 s countdown; the first
attempt sent
downfrom a separate round trip and missed the window, booting graphical (a grey screen with a cursor). Reset and sentdown+retinside the same remote command — the console then read "A Felhom telepítése indul szöveges módban…". - The guest keymap defaulted to Hungarian while
sendkeyemits US scancodes. Changed to U.S. English before typing anything. Positive control: the administrator email was typed through the identical code path and renderedc11@felhom.euon screen — the@isshift-2on a US layout andAltGr+Von a Hungarian one, so a correct@proves the mapping for the 24-char root password that is masked and cannot be read back. (Deviation from a real Hungarian install, recorded: the installed console keymap is US. The customer journey is the web dashboard, so this is cosmetic — but it is a deviation and it is not buried.) --bootset in its ownqm setafter the disks existed, and verified fromqm config(boot: order=scsi0) before the first boot.Automatically rebootwas unchecked so the reboot into the freshly installed system was deliberate rather than a race with the ISO.
Only scsi0 was attached during the install, so the installer offered exactly one target
(/dev/sda (QEMU HARDDISK) (200.00 GiB)) and could not reach a data drive. The two 50 G drives were
attached afterwards.
Summary screen, verbatim: ext4 · /dev/sda · Europe/Budapest · U.S. English ·
c11@felhom.eu · nic0 · c11.felhom.eu · 192.168.0.105/24 · 192.168.0.1 · 192.168.0.1.
Day-0 fired on first boot, unaided. The console showed, in Hungarian:
Felhom — a doboz készen áll, és a párosításra vár.
Párosító kód: RPM-YNF
Nyisd meg az e-mailben kapott linket, és add meg ezt a kódot és a jelszavadat.
Ez a képernyő magától frissül — nincs teendő a doboznál, és nyugodtan itt hagyhatod bekapcsolva.
and the hub's Unclaimed appliances table carried it within a minute: uuid
27f35841-0609-410d-af52-76d5c0dbf801 (matching the VM's smbios uuid), pairing code RPM-YNF,
MAC bc:24:11:c8:4e:ce, hardware Standard PC (Q35 + ICH9, 2009) / AMD Ryzen Embedded V1756B /
7.7 GB, and three SSH host keys.
Bound to c11 through the real endpoint (POST /appliances/20/bind, HTTP 303) → the row reads
bound → Campaign 11. Host c11-36d660 enrolled ~3 minutes later: ONLINE, agent 0.120.0,
PBS wrapper matches vouched, desired generation 2.
Day-0 provision SUCCESS, 14:12:37 CEST (journalctl -u felhom-bootstrap), 7 minutes after the
bind:
[OK] controller: Up 19 seconds (healthy) (after ~0s)
[INFO] controller image: gitea.dooplex.hu/admin/felhom-controller:0.192.0
[OK] Day-0 provision SUCCESS — vmid=9201 host_id=c11-36d660 customer=c11
golden=local:backup/vzdump-lxc-9100-2026_08_05-14_10_02.tar.zst
[INFO] root@pam was rotated + vaulted at step 4b — retrieve at hub → host page
felhom-bootstrap: host-install SUCCESS — writing done-flag, disabling unit, scrubbing secrets
/etc/felhom/.bootstrap-done written 14:12, unit inactive (dead) and disabled, 1 min 35 s CPU.
The installer's root password stopped working exactly as Campaign 10's A2 recorded — the first
SSH attempt with it was refused, and the hub's vaulted managed credential
(POST /hosts/c11-36d660/reveal-recovery-credential, set_at 2026-08-05T12:09:34Z,
username root@pam) authenticated first try. That refusal is the positive observable that the
day-0 credential handover completed.
The controller landed on 0.192.0 — the golden's baked version, exactly as predicted from the manifest.
Step 1b — the per-customer floor, and the convergence it drives
POST /customers/c11/floor with min_controller_version=0.200.0 (HTTP 303). The customers table then
reads c11 … 0.192.0 … v0.200.0 (override) while every other customer still reads v0.156.0 —
the blast radius is zero, measured rather than asserted.
The box converged in under two minutes, unaided:
gitea.dooplex.hu/admin/felhom-controller:0.200.0 Up 9 seconds (healthy)
So the recovery screen IS reachable on a fresh box — but only behind an operator floor override, because no golden carrying it exists.
FINDING (candidate R-214) — the physical console never stops asking to be paired
Thirty minutes after Day-0 provision SUCCESS, with the host ONLINE in the hub and the customer
guest running, the appliance's physical console still read:
Felhom — a doboz készen áll, és a párosításra vár.
Párosító kód: RPM-YNF
...
Ez a képernyő magától frissül — nincs teendő a doboznál, és nyugodtan itt hagyhatod bekapcsolva.
Established from source, not inferred. print_pairing_banner is called from exactly one place —
the pairing wait loop (scripts/iso/felhom-bootstrap.sh:456), every 10 cycles (5 min). A census of
/dev/console writers across the entire day-0 path returns two, both in felhom-bootstrap.sh:
the banner (:76) and the R-59 network-refusal screen (:252). scripts/felhom-host-install.sh
writes to the console not at all. So after the bind lands, nothing repaints, and the banner is the
terminal state of the screen.
Two things are wrong at once, and the second is the one that makes it a defect rather than a cosmetic
gap: the screen shows a stale pairing code for a box that is already bound, and the banner
asserts of itself that it refreshes. That is CLAUDE.md's "a comment asserting an invariant needs
a test pinning it" family, in customer-facing Hungarian copy on the one surface a customer looks at
when they are already unsure.
Mitigating, and stated so the severity is not overclaimed: the same banner says "nincs teendő a doboznál" — the customer's real route is the e-mailed link, not the box.
Step 2 — the claim, and the constraint it exposed
The dashboard is served by traefik on 443 in the guest (192.168.0.207), not by the controller
directly; Host: felhom.c11.felhom.eu is mandatory. GET / → 302 → /claim, titled
"A szerver beállítása — Campaign 11", asking for the beállító kód and a new password (min. 12
characters).
The claim code is e-mail-only by design and the campaign has to live with it. The hub stores a
hash (store.RotateClaimCode, store.go:1667) and the resend path states it outright:
"the hub stores a hash — no plaintext code exists to display" (configs.go:806). It went first to
c11@felhom.eu, a mailbox that does not exist.
This is not a Phase-0 inconvenience — it decides whether Phase 1 can be claimed as unaided. Phase 1
and Phase 3 each need a further claim, and the controller's --print-reset-code escape hatch is a
command line inside the guest. Operator ruling, 2026-08-05: point c11's registered address at the
operator's own mailbox and have them relay each code — which is precisely what a real customer does,
and what R-119 already records as "the single human step".
POST /configs/c11/edit with email=<the operator's address> (HTTP 303), then
POST /configs/c11/claim-resend (HTTP 303) → "Nyitott — kód kiküldve just now, generation 2".
Claim driven through the real customer form (POST /claim, code + a 20-char password, HTTP 302 → /),
landing on /launcher ("Indítópult"), controller 0.200.0. The box is claimed and the
household password is customer-owned.
F9, taken here rather than in Phase 2 — and it FAILS
F9's stated right answer for a box that never had off-site backups is "no screen, no entry point,
nothing". c11 is exactly that box: installed today, off-site never configured. Two of the three
hold. The third does not.
| Expected | Observed | |
|---|---|---|
| full-page interruption | none | none — GET / → /launcher, not /recovery ✅ |
| entry point in the backups area | none | none — /backups/remote reads "Még nincs beállítva távoli mentési cél." and contains zero /recovery links ✅ |
| the page itself | should not exist for this box | GET /recovery → HTTP 200, and it asserts the machine was reinstalled and that sealed backups exist ❌ |
Verbatim, from the rendered page on a box with no backups of any kind:
„Ezt a gépet újratelepítették. A korábbi, házon kívüli mentéseid megvannak — a Felhom központi rendszere őriz hozzájuk egy lezárt csomagot."
Ground truth, measured, not inferred. Hub SQLite snapshot taken WITH its -wal and -shm
(PRAGMA integrity_check → ok; freshness by positive observable — newest host_reports.received_at
2026-08-05 12:55:05 against datetime('now') 12:55:39, 34 s old):
SELECT host_id, length(identity_blob) FROM host_escrow;
demo-felhom-8363b5|572
demo-hp-bb76ea|572
-- c11-36d660: NO ROW AT ALL
So identity_blob_present is false in every ACK c11 has ever received, GetHubEscrowIdentityPresent()
is false, and OffsiteRecoveryOffer() correctly returns false. The predicate is right. The page
never asks it.
The mechanism, from source.
recoveryPageHandler(internal/web/recovery_handlers.go:65-67) is three lines and callsrenderRecoverywith no gate.renderRecoverysetsdata["Offer"] = s.recoveryOffer()(:76) andrecovery.htmlnever reads it —grep -c '\.Offer'over the template returns 0. The template's only top-level branch is{{if .Unlocked}}, so every non-unlocked render prints the headline above.- Its POST sibling IS gated:
recoveryUnlockHandler:141opens withif !s.recoveryOffer() { redirect to /backups/remote }. The action is protected; the page is not. - And the other template gets it right, in the same release, for the same claim:
backups_remote.html:11wraps the entry point in{{if .RecoveryOffer}}— which is why no entry point appeared. One gated, one not, same sentence.
Why the green suite missed it. TestRecovery_B_DoesNotAppearForAnyoneElse/never had off-site backups (no hub package) (recovery_test.go:200-208) asserts f.s.recoveryOffer() and issues no
GET /recovery. It pins the mechanism; the consequence is unpinned. That is CLAUDE.md's own rule —
prefer a test that asserts the consequence over one that asserts the mechanism — and the seam-wiring
rule's inverse case: usually the gate is missing, here the gate exists and the render does not consult
it.
Severity, stated without inflation. There is no link to /recovery on such a box, so a customer
reaches it only by typing the URL or following a stale bookmark. But what they are then told is false
in two ways at once — your machine was reinstalled and your off-site backups exist — on the one
screen whose entire purpose is to be believed about backups. It is I6: an absence reported as a fact.
→ R-215
Ground state for teardown — ep0, captured BEFORE anything else
The DR tier provisioned itself on ep0 the moment the customer was saved with dr_tier=on and the WG
peer registered. Recorded so teardown removes exactly the campaign's own and nothing else:
namespaces: c11 demo-felhom demo-hp ← c11 is the campaign's
tokens: felhom@pbs!c11 (+ the two demo) ← c11 is the campaign's
acl: felhom@pbs → /datastore/felhom-offsite/c11 DatastoreBackup
felhom@pbs!c11 → /datastore/felhom-offsite/c11 DatastoreBackup
df /mnt/pbs-datastore: 98G total, 9.9G used, 11%
Step 2b — the drives, through the real endpoints
Both 50 G drives appeared as raw candidates (GET /api/disks/candidates → initialize: [/dev/sdb, /dev/sdc])
and were enrolled through the customer flow (POST /api/storage/init, polled on
/api/storage/init/status to phase: done — Campaign 10's F6 detached-job path):
| Drive | Device | Stable path | Role |
|---|---|---|---|
adatok |
/dev/sdb |
/mnt/felhom-drives/adatok |
app data, set_default |
mentes |
/dev/sdc |
/mnt/felhom-drives/mentes |
backup target |
The backup-target offer behaved exactly as Campaign 10 recorded — offered, accepted through
POST /api/storage/backup-target/assign, restart_required: true, and the state flipped only after
the agent restart it asked for:
before: {"degraded":true,"target":"local","message":"A rendszermentés jelenleg ugyanazon a lemezen van…"}
after: {"degraded":false,"known":true,"label":"Mentes","target":"felhom-backup"}
That healthy read is this campaign's I5/I6 baseline.
Step 3 — the app and the three sentinels
calibre-web — the drill's own fixture app, and the only sensible choice: it declares
userdata: media/books class: mandatory (app-catalog-felhom.eu/templates/calibre-web/.felhom.yml:21-23),
it is single-container, and its file leg is a bind rather than a named volume (the 2026-08-04 drill
established that a named-volume-only app has no restorable file leg — the off-site tier tars it and
the customer restore flow never unpacks it).
Deployed through the real API (POST /api/stacks/calibre-web/deploy, HTTP 202) with
HDD_PATH=/mnt/felhom-drives/adatok — a real enrolled drive, which is the normal customer shape and
deliberately not R-203's /mnt/sys_drive case. Reached running / Up About a minute (healthy).
Three sentinels, written into the mandatory path, hashed after sync:
| # | File | Bytes | sha256 |
|---|---|---|---|
| A | C11-SENTINEL-A.txt |
54 | beb9175d52325bf583039beb6db1ecc80ea0781720e262c83230ed46a166df27 |
| B | C11-őrszem-ékezetes-árvíztűrő.txt |
73 | 7c8cb0ad040856440cc96bb1ad8f336d04f458138d85e6904ee759e690743a87 |
| C | C11-SENTINEL-C-12MB.bin |
12 582 912 | e012e76f3502de7c608d5b35f67b89c2404edad49eea92a1fc35a1bf02cbb9c0 |
(Method, stated: the sentinels were written on the box, not through the dashboard. Phase 1's no-command-line rule governs Phase 1; fixture construction is not part of what it measures.)
Step 4 — the off-site tier, and what enabling it actually did
POST /configs/c11/edit with offsite_enabled / shared / 50 GB / bx11. The hub provisioned a Storage
Box sub-account and the box picked up the target:
[offsite] shared provisioned for c11 (subaccount 284166, user u629488-sub4)
[offsite] host-key scan u629488-sub4.your-storagebox.de failed (attempt 1/6 … 5/6):
lookup … no such host ← DNS for a just-created sub-account
[INFO] Customer config updated: c11 ← succeeded on attempt 6
Worth recording rather than treating as noise: first-provision blocks on DNS propagation for the new sub-account hostname and took over 60 s of backoff (2+4+8+16+30). The retry ladder is doing exactly its job — but the HTTP request outlived a 120 s client timeout while the hub completed the work anyway, so a client that gives up here would be looking at a succeeded provision it thinks failed. No R minted: the hub's own copy already warns "offsite provisioning can take up to a minute".
Target on the box: u629488-sub4@u629488-sub4.your-storagebox.de:/home/felhom-repo, quota 50 GB.
calibre-web toggled on for off-site (POST /backup/offbox/toggle).
Step 5 — the escrow ceremony, and R
Preflight (GET /api/escrow/preflight) — six of six green, including the one that matters:
pbs_storage_id ok (felhom-pbs) · dr_tier ok (DR tier applied) · age_binary ok (/usr/bin/age)
hub_upload ok · staged_secret ok (staged secret present) · sudo_grant ok
Ceremony driven through the customer wizard's own endpoints (/api/escrow/start with a password
re-auth → /api/escrow/status → /api/escrow/claim):
phase: done · restic_pw_sealed: TRUE · uploaded: true · entropy_bits: 129.24
key_fingerprint: 3c:f7:76:82:e3:0c:85:d7:…:1d:04:43:b8
claimable: true → claimed: true, claim_expires_in_sec 585 → 0
restic_pw_sealed: true is the field this whole campaign rests on — the off-site repository
password is inside the sealed bundle.
R was captured file→file into a 0600 file and never rendered. Its shape was verified without
printing it: ten hyphen-separated alphabetic words, 85 characters, consistent with the ~129-bit EFF
phrase the ceremony reports.
Hub-side, immediately after (SQLite snapshot with -wal, freshness 9 s):
host_escrow(c11-36d660): blob 383 B · identity_blob 572 B · stale_at NULL
restic_pw_sha256 626e424670248db3a25b0a48aa8a114da9061529823ce9cea7d95d9e0ee3c303
created_at 2026-08-05T13:11:11Z
The box then read „A helyreállítási kód letétbe helyezve." and escrow_state: escrowed.
Step 5b — the off-site backup, and the sentinels listed BY NAME
POST /backup/offbox/run → Fut… → ✓ Rendben, 1 m 6 s.
The gate is not the green tick. restic snapshots + restic ls -l latest, run against the
repository with its own credentials:
snapshot f3d9cd67d539c00359f0454ea7a782f45beac8406c55275bee5a6886afa8d791 (short f3d9cd67)
time 2026-08-05T13:13:32.326542714Z
hostname c11 tags [felhom-offbox, calibre-web]
paths /mnt/felhom-drives/adatok/backups/primary/calibre-web
/mnt/felhom-drives/adatok/userdata/media/books
-rw-r--r-- 1000 1000 54 …/userdata/media/books/C11-SENTINEL-A.txt
-rw-r--r-- 1000 1000 12582912 …/userdata/media/books/C11-SENTINEL-C-12MB.bin
-rw-r--r-- 1000 1000 73 …/userdata/media/books/C11-őrszem-ékezetes-árvíztűrő.txt
-rw-r--r-- 1000 1000 413696 …/userdata/media/books/metadata.db
+ the recovery unit: compose/{.felhom.yml,app.yaml,docker-compose.yml}, manifest.json,
volume-dumps/calibre-web_calibre_web_config.tar (389 632 B)
All three sentinels are in the snapshot, by name, at the right sizes — and the Hungarian filename survived the whole chain into restic byte-intact. The mojibake gate the brief warned about does not fire here.
Independent live confirmation of R-203's fix, on a shape the drill never tested. The capture set's
path is /mnt/felhom-drives/adatok/userdata/media/books and the app's live bind is the same
directory (docker inspect: bind /mnt/felhom-drives/adatok/userdata/media/books -> /calibre-library).
On a non-system drive NamespaceRoot correctly omits the felhom-data segment for both, so the two
roots agree. (The import leg still resolves under /mnt/sys_drive/felhom-data/userdata/import/calibre
— a different drive AND the felhom-data root — which is correct for an excluded-class path but is
the same asymmetry R-203 was filed about. It is harmless here because nothing mandatory lives there.)
Step 6 — the pre-destruction truth
Box's own state (settings.json, secrets stripped):
offbox: enabled true · escrow_state "escrowed"
last_run/last_success 2026-08-05T13:14:26Z · last_status "ok" · last_duration 1m6s
snapshot_count 1 · repo_size_bytes 12 611 522 ("12.0 MB") · quota_gb 50
hub_escrow_identity_present: true
Hub's state: identity_blob 572 B, restic_pw_sha256 626e4246…c303, stale_at NULL,
Key Escrow present, DR Recipe present.
I5 at the Phase 0 → Phase 1 boundary — investigated, and DISPROVED
At 15:20 CEST the hub's operator view read Snapshots 0 · Repo Size 0 MB · Integrity Unknown
while the box held 1 snapshot and 12 611 522 B. That is exactly the shape I5 exists to catch, and it
would have been easy to file.
It was cadence, not disagreement. c11's legacy controller report last landed at 13:11:43 UTC — three minutes before the backup finished at 13:14:26. Rather than judge from a single stale read, a poll ran until either convergence or a full interval elapsed. It converged at 13:23:10 UTC, 8 m 44 s after the backup:
{"enabled":true,"escrow_state":"escrowed","last_run":"2026-08-05T13:14:26Z","last_status":"ok",
"last_success":"2026-08-05T13:14:26Z","snapshot_count":1,"repo_size_bytes":12611522,"quota_gb":50}
Byte-for-byte what the box's own settings.json holds. I5 PASS. No R-n minted.
The general lesson, and the reason this is written down rather than dropped: on this fleet the hub's view lags the box's by up to one report interval, so "the two views disagree" is only a finding after an interval has passed. A verdict taken from one read at the wrong instant is Campaign 10 §5.3's mistake — reading the two halves at different instants of a transition.
Phase 1 — the clean journey, unaided
Method note, stated before the result. There is no browser on DooPlex, so every customer action is
driven at the endpoint the UI itself invokes (felhom-controller/CLAUDE.md's standard method), through
the SOCKS transport described in the venue section. The rule Phase 1 actually tests — no command
line INSIDE THE GUEST — is enforced literally: after the destruction, nothing reaches guest 9201
except HTTP requests a browser could have made. Actions on the appliance host (the Proxmox layer) are
the customer's installer, not the guest.
No rollback archive was taken, deliberately. qm snapshot 321 refused — "snapshot feature is not
available", because the OVMF efidisk0 is raw and a raw disk blocks qcow2 snapshots. Rather than
convert a disk mid-campaign, the run proceeded without one: the venue is disposable and its loss is
authorised, and the only failure that matters — the off-site repository failing to open — is not
protected by a local archive anyway.
Step 1 — the destruction, 2026-08-05T13:35:19Z
Recorded immediately before, so the loss is measured rather than assumed:
rootfs local-lvm:vm-9201-disk-0 32G · mp0 local-lvm:vm-9201-disk-1 70G (/var/lib/felhom)
containers: felhom-controller 0.200.0 · calibre-web · filebrowser · traefik
/mnt/felhom-drives/adatok/userdata/media/books: all three sentinels + metadata.db (13M total)
/mnt/felhom-drives/adatok/backups/primary/calibre-web: the Tier-1 recovery unit
Then:
pct stop 9201 ; pct destroy 9201 --purge
Logical volume "vm-9201-disk-0" successfully removed.
Logical volume "vm-9201-disk-1" successfully removed.
purging CT 9201 from related configurations..
pct list → (empty)
rm -rf /mnt/felhom-drives/{adatok,mentes}/* ← a hardware loss takes the drives too
du -sh → 4.0K / 4.0K
The guest, its 102 GB of volumes, both enrolled drives and the entire local backup tier are gone.
The customer's three sentinels now exist in exactly one place on Earth: the off-site restic repository
at u629488-sub4.your-storagebox.de:/home/felhom-repo.
This is a deliberately harder destruction than the 2026-08-04 night drill, which wiped only the
controller's data volume. The brief asks for "the guest and its data volume, as a hardware loss would",
and being attended rather than unattended makes the guest-destroy + reprovision chain safe to walk.
The host identity c11-36d660 survives — which is the point: the hub's sealed package is keyed to
the HOST, so destroying the host as well would make recovery impossible by construction rather than by
defect.
Step 2 — reinstall through the same day-0 path, started 13:36:07Z
The hub's own Setup Command, felhom-host-install.sh v1.25.0 fetched live from
https://felhom.eu/scripts/ — the documented path that "does what the ISO's package would have
done". The retrieval passphrase went in via --passphrase-file (a 0600 file), never on argv.
Completed 13:38:06Z — Day-0 provision SUCCESS — vmid=9201 host_id=c11-36d660 customer=c11,
guest back at 192.168.0.227, controller 0.192.0 healthy. root@pam rotated + re-vaulted
again. Then, unprompted, the per-customer floor did its job:
13:35:19Z destroy
13:36:07Z reinstall start
13:38:06Z controller 0.192.0 healthy ← the GOLDEN's version
13:38:42Z controller 0.200.0 healthy ← floor-driven self-update, 36 s later
The assertion that keeps recovery possible — PASSED. Read from the hub with the -wal, after the
destruction:
host_escrow(c11-36d660): identity_blob 572 B · restic_pw_sha256 626e424670248db3… (UNCHANGED)
created/updated 2026-08-05 13:11:11 (UNCHANGED) · stale_at NULL
host_escrow_superseded: rows 3 and 4 only — the two demo boxes. NOTHING re-escrowed itself.
And the rebuilt box already knew. Its fresh settings.json, before any human touched it:
hub_escrow_identity_present: true, no offbox key, no claimed key — precisely shape (a) of
OffsiteRecoveryOffer().
Steps 3–4 — claim, and the first thing the customer sees
The rebuild left the box unclaimed, so the hub rotated a reset code (generation 3) to the registered address; the operator relayed it, as R-119 records is the one human step.
13:42:41.101Z POST /claim (code + a new password) → 302 → /
GET / → 200 https://192.168.0.227/recovery ← NOT /launcher
The recovery screen appears without being sought. The customer's very first landing after claiming their rebuilt machine is the screen that tells them their backups survived. ✅
Step 5 — the screen, read as a customer
It says what is waiting, when it was sealed, and that nobody can replace a lost code:
„Ezt a gépet újratelepítették. A korábbi, házon kívüli mentéseid megvannak — a Felhom központi rendszere őriz hozzájuk egy lezárt csomagot, amelyet 2026-08-05T13:11:11Z zártunk le."
„A helyreállítási kódot senki nem tudja pótolni — sem a Felhom, sem az ügyfélszolgálat, sem az üzemeltető. Ez szándékos… Ha a kód elveszett, a korábbi mentések nem nyithatók meg többé."
„Ebben a lépésben semmit nem állítunk vissza és semmi nem változik."
All three of step 5's questions answered, in plain Hungarian. The set-aside option was correctly
not offered, with its reason stated rather than the button merely hidden: „ez a lehetőség akkor
válik elérhetővé, ha a gép már újra kapcsolódott a házon kívüli tárhelyhez" — the CanSetAside
gate (ResetOrphanedRepo refuses unless orphaned) working, and explaining itself.
One copy defect on this screen: the seal date renders as the raw RFC3339 UTC string
2026-08-05T13:11:11Z to a Hungarian household. recoverySealedAt() is documented as "the human
date the hub says the sealed package was created"; it is not humanised anywhere. Minor beside what
follows, and recorded rather than dropped.
Step 6 — the recovery code was entered, and REFUSED. Phase 1 FAILS HERE.
13:43:30.800Z POST /recovery/unlock (the code exactly as issued)
13:43:30.953Z HTTP 200 in 0.134 s
„A megadott helyreállítási kódot nem fogadtuk el. Ellenőrizd, hogy mind a tíz szót pontosan, szóközökkel elválasztva írtad be — a kis- és nagybetűk nem számítanak."
0.134 s is the tell. age's scrypt is deliberately slow; a real unseal cannot finish that fast.
The controller's own log names the truth:
[WARN] [web] recovery: unlock failed: agentapi: POST /escrow/recover-offsite-password: HTTP 404
The agent does not have the endpoint. POST /escrow/recover-offsite-password ships in agent
v0.125.0 ("the agent opens the sealed bundle and returns one field", 2026-08-04, R-199 links
7–8). This box runs the vouched day-0 agent, 0.120.0.
So the machine told the customer their recovery code was wrong, when the code was perfect and the machine simply could not ask. That is the single worst outcome the brief names — an unrelated failure reported as a wrong code — and it happened in the CLEAN journey, not under an injected fault.
→ R-216
The counterfactual — measured, so R-216 is a fact and not an inference
Agent 0.125.0 installed on the campaign box only (published artifact, sha
f7d8339b53d92a6c45be7eaf189469a041b6b00b758a64511c0479beae7016b3, verified on the box against the
hub's own vouched value; no global manifest change, so no other customer moved). The identical
code was then re-entered:
| agent 0.120.0 | agent 0.125.0 | |
|---|---|---|
| verdict | „A megadott helyreállítási kódot nem fogadtuk el" | „A mentéseid elérhetők — A kulcs visszakerült" |
| elapsed | 0.134 s | 1.113 s — the scrypt cost, visible |
| outcome | — | recovery: the offsite repository key was recovered and placed (outcome=installed) |
Nothing else changed. The code was always correct; the machine could not ask, and blamed the customer.
R-217 — the same screen then asserted three things it could not know
The successful unlock rendered the honest error and a contradiction beneath it:
„A kulcs visszakerült. A gép még most kapcsolódik újra a házon kívüli tárhelyhez…" and „A tároló megnyílt, és van benne tartalom, de nem tudtuk alkalmazásokhoz rendelni."
The store did not open, its content is unknown, and attribution is not the problem.
recovery_handlers.go:184 builds backup.OffsiteInventory{} on the failure path — Empty defaults to
false — so the template's InvUntagged (!Empty && len(Apps)==0) fires. The type's own doc
comment names precisely this hazard:
"Empty is true when the repository opened cleanly and holds no snapshots… it is named rather than inferred from
len(Apps)==0, which is also what a failed read looks like."
The field built to prevent it was defaulted past. → R-217
R-218 — succeeding at recovery is what stops the machine asking for what it still needs
The store never opened. 47 minutes after the unlock, against a screen promising "néhány perc". Not slow — deadlocked, and each half is independently sufficient.
Measured, from the hub's own report stream:
13:38:06 (offsite absent)
13:38:46 (offsite absent)
13:39:54 {"state":"needs_credential", …} ← the box asking. v0.199.0 working.
13:42:43 {"state":"needs_credential", …} ← second report = the debounce satisfied
13:47:03 hub: "offsiteheal: re-staged the stored one-time offsite secret for customer c11
(declared needs_credential across 2 reports) — the box re-consumes on its next cycle"
13:47:35 THE CUSTOMER'S RECOVERY UNLOCK PLACES THE REPOSITORY PASSWORD
13:53:42 (offsite absent) ← the box has stopped asking
14:08:42 (offsite absent)
14:23:42 (offsite absent)
Half 1 — the box stops asking the moment recovery succeeds. needsOffsiteCredential
(backup/offbox.go) short-circuits on the repository password:
if _, ok := m.OffboxRepoPasswordHash(); ok {
return false // we still hold our repository password: not a fresh data area
}
Placing that password is the recovery screen's entire job. Doing it makes
OffboxReportStatus() return nil, the offsite object vanishes from the report, and
offsiteheal's reconciler — which acts on that one declaration and nothing else
(reconciler.go:39-52) — can never fire again. The box's state now:
/data/offbox/ repo_password 64 B (13:47) ← the recovered key, present
no known_hosts, no ssh_key, no applied_marker
settings.json offbox key ABSENT ← no coordinates, no transport credential
hub_escrow_identity_present true
A correct key, and nowhere to use it.
Half 2 — the re-stage nudges nothing. Reconciler.heal re-arms the secret, logs "the box
re-consumes on its next cycle", emits an event, and returns. A census of the whole offsiteheal
package for poke|bumpIntent|SetHostDesired|generation returns zero hits — no poke, no
generation bump, nothing that would make the box re-pull. And the controller's apply-bridge retries
only "on next config refresh/restart" (its own 13:38:51 line). Last Config downloaded for customer c11: 13:38:00 — nine minutes BEFORE the re-stage, and none since.
So the hub holds an unconsumed credential (one_time_secrets: one c11 row, created
2026-08-05 13:07:03, consumed_at NULL after the re-arm) that the box has no reason to come and
collect, while the box no longer says it needs one.
The remedy fires correctly, 32 seconds before the customer succeeds — and the customer's success is what closes the door on it. Two shipped features, each correct alone, cancel each other on the exact path they were both built for.
→ R-218. This is the direct cause of Phase 1 not reaching an open store.
The intervention (deviation 3, operator-authorised) — and what it proved
systemctl restart felhom-controller-bootstrap.service at 14:31:33Z — the apply-bridge's own
documented retry trigger, not a code change. It worked in 17 seconds:
14:31:46 [offsite-apply] settle-gate: GO — at/above floor 0.200.0
14:31:50 [offsite-apply] offsite configured for u629488-sub4@…:/home/felhom-repo (pending key escrow)
Which confirms R-218 exactly: nothing was wrong with the credential, the target or the key — the only thing missing was anything at all to trigger a retry.
And the recovered key survived the re-apply, byte-exact — three independent sources agreeing:
| source | sha256 |
|---|---|
hub restic_pw_sha256, recorded BEFORE the destruction |
626e424670248db3a25b0a48aa8a114da9061529823ce9cea7d95d9e0ee3c303 |
| the file on the rebuilt box | 626e424670248db3a25b0a48aa8a114da9061529823ce9cea7d95d9e0ee3c303 |
| mtime of that file | 13:47 — the unlock. Untouched by the 14:31 re-apply |
The re-apply wrote known_hosts, ssh_key and applied_marker at 14:31 and left
repo_password alone, exactly as WriteOffboxSecrets documents. R-193's orphaning mechanism did
not fire.
R-219 — the listing the screen promises can never render on the shape it exists for
The screen's headline promise is explicit:
„Ha megadod a kódot, feloldjuk a mentéseid zárolását és megmutatjuk, mi van bennük — melyik alkalmazás, mikorról, mekkora."
On shape (a) — the pristine rebuilt box, the shape OffsiteRecoveryOffer's own header calls the
customer's actual state — that listing cannot render. Ever. By construction:
ApplyOffsiteTarget→WriteOffboxSecretsauto-generates the repository password whenever none is present (offbox.go:400-403). So a target cannot exist without a repository password.- Shape (a) is defined by no repository password. Contrapositive of (1): no password ⇒ no target.
OffsiteInventoryListneeds a target → returnsErrNoOffsiteTarget→ the handler takes the failure branch every time.
Measured live at 13:47:35: the unlock succeeded and the customer got „a mentéseid listája néhány perc múlva jelenik meg" — plus R-217's false „A tároló megnyílt, és van benne tartalom". Not once did the promised listing appear.
And there is no second chance. After the unlock places the password, recoveryOffer() evaluates
OffboxRepoPasswordHash() → ok, then OffboxOrphaned() → t != nil && … → false with no target.
So the offer retires itself, and recoveryUnlockHandler's opening guard redirects. Measured, after
the re-apply:
POST /recovery/unlock → 302 /backups/remote ← the listing is permanently unreachable
GET / → 200 /launcher ← the screen has stepped aside
The customer's single opportunity to see what is in their backups is the unlock response itself — and on shape (a) that response is structurally guaranteed not to contain it. → R-219
(The screen stepping aside once the job is done is CORRECT and is recorded as a pass: GET / went
back to /launcher. The defect is that the one thing it promised was never delivered before it went.)
Step 8 blocked — the restore page is empty exactly when it is needed
/backups/restore, immediately after a successful recovery on a rebuilt box:
„Nincs távoli mentésre jelölt alkalmazás — a kijelölés a … oldalon történik."
„Nincs ellenőrző másolat a meghajtón."
„Nincs telepített alkalmazás."
The post-unlock screen's own button — „Tovább a visszaállításhoz" — lands here. The restore flow is
keyed on locally installed, off-site-toggled apps, and a rebuilt box by definition has none; the
snapshot in the repository (calibre-web, with its recovery unit) is invisible to it. This is R-213's
territory ("putting files back is not built"), but sharper than not built: the entry point exists,
is reachable, and is empty precisely on the path the recovery screen hands the customer to.
R-220 — after a rebuild the customer's own drives cannot be re-enrolled, and the refusal names an impossible action
Redeploying the app was the next step, so the app needs its data drive. The deploy refused, in Hungarian, through the real endpoint:
„A megadott tárhely nem azonosítható regisztrált meghajtóként, ezért alkalmazás adatkönyvtáraként nem használható. Válasszon a listából csatlakoztatott meghajtót."
The list is empty. GET /api/disks/candidates → initialize: [], attach: [], while the
customer's two drives sit mounted and visible in /api/disks. The wizard at /storage/attach
renders „Betöltés…" over nothing.
The seam, named: felhom-agent/internal/storage/claim.go:84-85
if n.mountpoint != "" && !underFelhomDrives(n.mountpoint) {
return false, "device is mounted at " + n.mountpoint + " (" + n.name + ")"
}
A device mounted anywhere outside /mnt/felhom-drives is claimed and never offered as a candidate.
After the rebuild, /dev/sdb was mounted at both /mnt/felhom-drives/adatok (the bind the guest
sees) and /mnt/adatok — the raw enrol mount that disks.go:606 leaves behind by design,
which eject also deliberately preserves. The controller's registry, however, lives in settings.json
and went with the wipe.
The drive is invisible to the re-enrolment wizard because of the mount that enrolling it created.
RED-PROOF, both directions. Unmounting only /mnt/adatok and /mnt/mentes — changing nothing
else — flipped the scan immediately:
before: initialize: [] attach: []
after : initialize: [/dev/sdb, /dev/sdc] attach: [(/dev/sdb, ext4), (/dev/sdc, ext4)]
Both drives then re-enrolled through the real endpoint with their data intact
(POST /api/storage/attach → {"registered":true,"where":"/mnt/felhom-drives/adatok"}), and the
deploy that had returned 409 returned 202.
Why this is a defect and not an edge case. (a) It is on the rebuild path, which is the whole
subject of the recovery feature. (b) It breaches I3 — the refusal names an action ("choose a
connected drive from the list") that the customer cannot perform, because the list it points at is
empty for a reason unrelated to anything they did. (c) Campaign 10 reached this exact state by hand
and recorded it as its own harness error (§4a: "hand-driving format + assign first left the drives
mounted but unregistered — invisible to both the raw-candidate scan and /disks"). Here the
product's own rebuild path arrives there. What was a harness mistake a week ago is now a product
state. → R-220
(Deviation 4, recorded: the unmount was an operator action on the appliance host. It doubles as the red-proof, and everything after it went through the customer's real endpoints.)
Step 8 — the sentinels, restored and compared
calibre-web redeployed (70.1 s) onto the re-enrolled drive, toggled for off-site, and restored
through the customer's own two-step full-restore flow:
step 1 (size gate) POST /backup/offbox/restore mode=full → 302 …&full_size=12.8+MB
step 2 (confirm) POST /backup/offbox/restore mode=full&confirm=1
14:43:27 → 14:43:43 [offbox] restored calibre-web (f3d9cd67, full=true)
→ /mnt/felhom-drives/adatok/backups/offsite-restore/calibre-web
16 seconds, out of snapshot f3d9cd67 — the one taken before the wipe. Non-destructive: the
live data was untouched, exactly as the message says („A meglévő adatok változatlanok.").
| # | file | expected sha256 | restored sha256 | |
|---|---|---|---|---|
| A | C11-SENTINEL-A.txt |
beb9175d…6df27 |
beb9175d52325bf583039beb6db1ecc80ea0781720e262c83230ed46a166df27 |
✅ |
| B | C11-őrszem-ékezetes-árvíztűrő.txt |
7c8cb0ad…43a87 |
7c8cb0ad040856440cc96bb1ad8f336d04f458138d85e6904ee759e690743a87 |
✅ |
| C | C11-SENTINEL-C-12MB.bin |
e012e76f…bb9c0 |
e012e76f3502de7c608d5b35f67b89c2404edad49eea92a1fc35a1bf02cbb9c0 |
✅ |
All three byte-identical, sizes intact (54 / 73 / 12 582 912 B), and the Hungarian filename survived the entire round trip — disk → restic → SFTP → Storage Box → restic restore → disk — without a single mangled character. The mojibake gate never fired.
Phase 1 — VERDICT: FAIL
The pass condition is "a sentinel's sha256 is byte-identical AND no step needed a command line inside the guest." The brief asks which of the two failed, and it matters enormously:
| half | result |
|---|---|
| the data | PASS — all three sentinels byte-identical, including a 12 MB binary and a non-ASCII Hungarian name |
| the journey | FAIL — four operator interventions, none available to a customer |
The backup promise holds. The recovery journey does not.
The four interventions, each one a customer dead end
| # | Where it stopped | What was needed | Finding |
|---|---|---|---|
| 1 | the correct recovery code was refused as wrong | install agent 0.125.0 | R-216 |
| 2 | the store never opened; the box stopped asking for its credential | restart the controller | R-218 |
| 3 | the promised listing never rendered, and became unreachable | (none — structurally impossible) | R-219 |
| 4 | the app could not be redeployed; its drives were unenrollable | unmount the raw mounts on the host | R-220 |
The RTO the business gets to quote
Undefined — the unaided journey does not complete. The attended elapsed time, stated so it is not mistaken for the customer number:
13:42:41 login (claim submitted)
13:43:30 recovery code entered → REFUSED (R-216)
13:47:35 recovery code accepted (after intervention 1) +4m 54s
14:31:50 off-site target applied (after intervention 2) +49m 09s
14:42:42 app redeployed (after intervention 4) +59m 61s
14:43:43 sentinels restored + verified +61m 02s
61 minutes with an operator, three of whose four blockers required root on the appliance. The only segment that reflects the product working is the last one: 16 seconds to pull 12.8 MB back out of the off-site repository once everything was in place.
For comparison, the 2026-08-04 night drill measured the same restore leg on a smaller fixture and also needed four undocumented manual steps (R-204). Three of those four were closed on 2026-08-05. This run found four new ones.
Phase 3 — the retention proof (R-198)
Step 1 — the pre-supersession record
Hub SQLite with its -wal, freshness 2 s (MAX(received_at) 14:47:19 vs now 14:47:21):
host_escrow (c11-36d660): blob 383 B · identity_blob 572 B · stale_at NULL
restic_pw_sha256 626e424670248db3a25b0a48aa8a114da9061529823ce9cea7d95d9e0ee3c303
created 2026-08-05T13:11:11Z
host_escrow_superseded: id 3 demo-hp-bb76ea 383 identity_blob NULL 2026-08-04 07:15:36
id 4 demo-felhom-… 383 identity_blob NULL 2026-08-04 07:20:08
Two superseded rows, neither for c11, and both carrying identity_blob NULL — the pre-v0.93.0
shape the RECON documented. This is the "before" the whole phase turns on.
Step 2 — destroy and reinstall, WITHOUT recovering (14:47:43Z)
Same destruction as Phase 1 (guest purged, both drives wiped to 4.0K), same day-0 path, and the
recovery screen deliberately not used so the box mints a fresh key.
R-216, part 4 — the day-0 path REVERTS the fix on every rebuild
The reinstall put the broken agent back, and said so:
[INFO] agent (existing): felhom-agent 0.125.0
[INFO] manifest: agent v0.120.0 (sha a7763d31b55b5ce7…), golden v0.192.0
[OK] installed /usr/local/bin/felhom-agent (felhom-agent 0.120.0)
Not a skip — a deliberate downgrade to the vouched version. The first reinstall had logged
[SKIP] agent v0.120.0 already installed; this one found 0.125.0, compared it against the manifest,
and installed 0.120.0 over it.
So an operator who fixes a box by hand has it re-broken by the next rebuild — which is precisely the event that makes the recovery feature necessary in the first place. The vouched-artifact channel does not merely default to a stale agent; it actively restores it. Until the manifest vouches ≥ 0.125.0, no box can hold the fix across the only event that needs it.
Steps 2b–2c — the box mints a fresh key, exactly as R-193 predicts
The rebuilt box came up claimed-less with no offbox directory at all — no repository password —
and hub_escrow_identity_present: true. It declared its need on schedule:
14:50:43 {"state":"needs_credential", …}
14:55:42 {"state":"needs_credential", …} ← debounceReportsDefault = 2 satisfied
14:57:03 hub: offsiteheal re-staged the stored one-time offsite secret for c11
The reconciler's own constants explain the timing: debounceReportsDefault = 2,
tickIntervalDefault = 5 * time.Minute (offsiteheal/reconciler.go:70,74). The hub side works.
And R-218 reproduced exactly. The box again did not consume the re-staged credential on its own;
the same lever was applied (systemctl restart felhom-controller-bootstrap.service, 14:57:18Z) and the
tier configured 17 seconds later — the identical 17 s as in Phase 1:
14:57:31 [offsite-apply] settle-gate: GO
14:57:35 [offsite-apply] offsite configured for u629488-sub4@…:/home/felhom-repo (pending key escrow)
Second occurrence, same signature, same fix. R-218 is not a one-off.
The fresh key, which is the whole premise of Phase 3:
| sha256 of the repository password | |
|---|---|
| Phase 0/1 (recovered, sealed under the FIRST code) | 626e424670248db3a25b0a48aa8a114da9061529823ce9cea7d95d9e0ee3c303 |
| after this rebuild (newly minted) | e11a6c542b73477ad02635673286f22f56ba8e1d5d9706efb4b0a0ff94473d57 |
Different keys. The Phase 0 history (snapshot f3d9cd67, holding all three sentinels) is now encrypted
under a password this box no longer has — R-193's orphaning mechanism, live, on demand. That is
precisely the state whose recoverability the retention is supposed to preserve.
A third instance of the "nothing triggers a refresh" family
The escrow preflight refused immediately after the tier configured:
{"id":"pbs_storage_id","ok":false,"detail":"escrow.pbs_storage_id not configured"} ← ok:false overall
(the other five green: dr_tier, age_binary, hub_upload, staged_secret, sudo_grant)
The ceremony cannot start until the controller's config carries escrow.pbs_storage_id, which arrives
on a config refresh — the same channel R-218 showed nothing reliably triggers. Recorded as an
observation on the same family rather than as a separate finding; it resolved on the next refresh.
R-221 — a rebuilt box cannot run the escrow ceremony at all, and the cause is a lifetime mismatch
With the tier configured, the ceremony's preflight refused:
{"id":"pbs_storage_id","ok":false,"detail":"escrow.pbs_storage_id not configured"} ← overall ok:false
(the other five green)
escrow.pbs_storage_id is agent-side and is seeded by the pbsdr bridge —
finishConverged → seedEscrowStorageID (felhom-agent/internal/pbsdr/manager.go:427-438), whose own
comment says "pbsdr convergence that just seeded escrow.pbs_storage_id flips this row green with no
restart." Measured on the box: grep -c escrow /etc/felhom-agent/agent.json → 0.
The two things have different lifetimes, and that is the whole defect.
| lives | survives a guest rebuild? | |
|---|---|---|
| the convergence marker | /var/lib/felhom-agent/pbsdr/marker.json — on the HOST |
YES (applied_at 2026-08-05T12:25:07Z, from Phase 0) |
| the seeded key | /etc/felhom-agent/agent.json — rewritten by the installer |
NO |
And the idempotence guard reads only the marker (manager.go:285-288):
if mk := m.loadMarker(); mk != nil && mk.Hash == h && (cf == nil || cf.Hash != h) {
m.setStatus(...)
return // idempotent: this exact descriptor already converged
}
Unchanged descriptor → same hash → early return → finishConverged never runs → the seed never
happens — into a config file that no longer has it. The directory even carries
agent.json.pre-seed (14:25) proving the seeder ran exactly once, in the Phase 0 window.
RED-PROOF. Moving only the marker aside and restarting the agent — nothing else — produced the seed instantly:
17:02:26 pbsdr: seeded escrow.pbs_storage_id (the ceremony one-liner needs no flags) storage_id=felhom-pbs
17:02:26 pbsdr: converged state=adopted storage_id=felhom-pbs
grep -c escrow agent.json: 0 → 1 preflight overall ok: false → TRUE
Consequence: on any rebuilt box whose PBS-DR descriptor has not changed, the customer cannot run the escrow ceremony at all — which is precisely when they must, because the rebuild just changed their repository key. → R-221 (deviation 5: the marker move-aside. It doubles as the red-proof.)
★ Phase 3 step 4 — THE RESULT. R-198's retention WORKS IN PRODUCTION.
Second ceremony run through the customer wizard at 15:03:12Z (restic_pw_sealed: true,
uploaded: true), its code claimed one-shot into a separate 0600 file — ten words, 90
characters, verified different from the Phase 0 code without either being rendered.
Hub SQLite with -wal, freshness 99 s:
=========== host_escrow_superseded ===========
id host_id k_blob identity_blob_len restic_pw_sha256 superseded_at
3 demo-hp-bb76ea 383 -1 (NULL) 8e03eddf9ff7fde2 2026-08-04 07:15:36
4 demo-felhom-8363b5 383 -1 (NULL) 48741892f0ef4d59 2026-08-04 07:20:08
5 c11-36d660 383 572 626e424670248db3 2026-08-05 15:03:14 ★
=========== host_escrow (current) ===========
c11-36d660 blob 383 identity_blob 572 e11a6c542b73477a 2026-08-05T15:03:14Z
| the question | the answer |
|---|---|
| does the superseded row carry the previous sealed package? | YES |
| is its length equal to step 1's? | YES — 572 B, exactly |
| is it the OLD key? | YES — 626e424670248db3, the Phase 0 key, sealed under the FIRST code |
| and rows 3–4, superseded before hub v0.93.0? | identity_blob NULL — the old, lossy shape |
This is the first supersession that has ever occurred in production since the fix, and the retention held. What was unit-proven only is now proven live. The blob that R-198 exists to keep was kept, byte-length exact, on the first real opportunity to lose it.
Step 5 — both signals fired, at the instant of supersession:
15:03:14 offsite_repo_key_changed warning "Offsite repository key CHANGED for host c11-36d660…" (R-197)
15:03:14 escrow_superseded info "A korábbi helyreállítási csomag megőrizve (új kulcs érkezett)."
14:57:03 offsite_selfheal_restaged info (R-204 item 4's event, also correct)
Step 6 — the run REFUSED rather than starting a fresh history:
15:03:41 [escrow-confirm] hub-verified: the escrow covers the current repo password (hash e11a6c542b73…)
— EscrowState auto-confirmed escrowed; offsite runs enabled
15:04:43 [offbox] offsite repo ORPHANED — remote holds backups written under a previous,
no-longer-available key; runs will skip until reset
15:04:43 Event pushed: offbox_repo_orphaned (warning)
The Phase 0 history — snapshot f3d9cd67, holding all three sentinels — is intact and untouched on
the remote. Nothing was overwritten. The escrow auto-confirm also worked with no manual
confirm-escrow, which the 2026-08-04 drill had needed (R-204 item 3, closed, confirmed live here).
R-222 — step 7: the retained package has no read path, and reaching for it is reported as a wrong code
With agent 0.125.0 restored so the answer would not be confounded by R-216, the first recovery code — the one that genuinely opens the orphaned history — was entered:
15:05:59 unlock HTTP 200 in 1.111 s ← a REAL unseal attempt, not a 404
[WARN] recovery: unlock failed: agentapi: POST /escrow/recover-offsite-password: HTTP 400:
the recovery code did not open the sealed bundle, or the bundle could not be fetched
— nothing was written
The engine is correct: the retrieval serves the CURRENT row, the first code cannot open a bundle
sealed under the second, it fails closed, and nothing is written. Exactly as the RECON predicted
(ListSupersededEscrow: zero production callers).
The customer-facing message is not:
„A megadott helyreállítási kódot nem fogadtuk el. Ellenőrizd, hogy mind a tíz szót pontosan…"
— and the same screen simultaneously states the hub holds a package sealed at 2026-08-05T15:03:14Z
(the NEW one), with no hint that an older package exists, is deliberately retained, or corresponds to
the code just typed.
So the customer holding exactly the right code for their orphaned history is told they mistyped it. And this is the same sentence R-216 produces for a missing agent capability, and the same sentence a genuinely wrong code produces. Three different situations — wrong code / the machine cannot ask / the machine asked about the wrong package — one message.
This is also the live proof of R-202's open concern: the orphan card promises the old history „a hozzá tartozó helyreállítási kóddal később visszaállítható lehet". Here is a customer with that code, and the answer is "wrong code". → R-222. Per the brief: the gap is recorded, nothing was built.