From 0c4411e54b9dca0296888c279cadc49fc9db67c2 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Thu, 6 Aug 2026 12:18:29 +0200 Subject: [PATCH] =?UTF-8?q?R-201=20re-walk:=20the=20data=20PASSES=20again,?= =?UTF-8?q?=20the=20journey=20still=20FAILS=20=E2=80=94=20two=20dead=20end?= =?UTF-8?q?s,=20down=20from=20four?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Asked Campaign 11 Phase 1's question a second time, on the fixed build, on a NEW appliance (VM 322, customer rewalk). The Campaign 11 venue was untouched. THE DATA: PASS. All three sentinels byte-identical out of the pre-destruction snapshot a7bc23bd in 23s through the customer's own restore flow — including a 12 MB binary and an accented Hungarian filename whose NAME BYTES are identical too (verified as hex, not as rendered text). THE JOURNEY: FAIL, two dead ends against Phase 1's four. 1. R-218's CONSUME half. The hub re-staged the credential at 11:44:57 saying 'the box re-consumes on its next cycle'; a full cycle ran at 11:55:46/54 (with a positive control that it ran) and it did not. A census of the customer-reachable actions found none that fetches it. Only a command line INSIDE THE GUEST moved it — 18s, confirming nothing was wrong with the credential, target or key: only the trigger. R-218's row said SHIPPED and over-claimed; it is corrected to REOPENED for the consume half. 2. R-220. Drives still unenrollable after a rebuild, needing a Proxmox-host unmount; without it no app redeploys and the restore page stays empty. Unaided RTO STILL UNDEFINED. Attended: +45s key placed, +24m12s tier up, +30m13s data verified. The 30m must not be quoted as the customer number. What passed and is new: the recovery screen appeared WITHOUT being sought, answered all three questions with a seal date matching the hub exactly, the emailed reset code worked first try, the unlock was a real 1.528s unseal, and R-225's fix was seen working in the wild (unknown, not a false zero). R-216 part 4 reproduced live: the reinstall downgraded the hand-installed agent 0.126.0 -> 0.125.0. DELIVERY GAP recorded as owed and NOT conflated with the journey: a fresh install landed on controller 0.201.0 / agent 0.125.0 — the vouched versions, neither carrying the fixes — installed by hand. Nothing was vouched. Capability map row STAYS FAIL. Campaign 11 doc gets a dated ADDENDUM, not a rewrite. --- STATUS.md | 62 ++++--- .../architecture/00-capability-map.md | 2 +- ...CAMPAIGN-11-recovery-journey-2026-08-05.md | 16 ++ .../tests/rewalk-r201-2026-08-06/journal.md | 169 ++++++++++++++++++ 4 files changed, 223 insertions(+), 26 deletions(-) diff --git a/STATUS.md b/STATUS.md index f258916..c14db7c 100644 --- a/STATUS.md +++ b/STATUS.md @@ -44,35 +44,47 @@ code was wrong. *(CAMPAIGN 11)* - **The off-site copy can be erased by the machine that made it.** A daily snapshot is armed as a stopgap. *(R-95, R-87)* -## What last night's stress test found — and what we fixed this morning +## Can a household get their data back on their own? Asked again today — still no, but nearer -We spent the night trying to break the recovery journey, then left the machine alone and watched it -run. **Nothing we did lost a byte.** When the customer chose "I do not want the old data", the old -backups were **set aside and not deleted** — we checked the far end of the wire and the 12.5 MB was -still there, to the byte. A wrong code was refused three times with nothing written and no lockout. -The machine's alarm fired when we switched it off and cleared itself when it came back. Overnight it -ran a full cycle on its own and made a fresh off-site copy without being asked. +We built a **brand-new machine** from the published disc, gave it three marked files, destroyed it — +guest and both drives, as a hardware loss would — and tried to get them back the way a household +would. *(R-201, the re-walk)* -**What it found: the machine still blamed the customer for failures that were not theirs.** Pull the -plug on our own central system and the customer was told their recovery code was bad — in three -hundredths of a second, when actually checking a code takes about one. The machine had not even -tried. **All of that is fixed and deployed** *(R-224, R-226, R-225, R-227, R-228)*: +**The files came back perfectly.** All three, byte for byte, including a 12 MB file and one whose +Hungarian accented filename came back **letter-for-letter identical**. Out of the pre-destruction +backup, in **23 seconds**, through the customer's own restore screen. -- **When something on our side is down, we say so** — and we say plainly that the code was **not** - used, so it is still good. Proven on the real machine: with our hub unreachable the answer changed - from "your code is wrong" to "we could not reach the central system". -- **A customer who mistypes is told to check their typing again.** That message had become - unreachable on any machine that had been given a new code — exactly the machine that just recovered. -- **When we do not know why something failed, we say that**, and never guess the customer. -- **"0 snapshots · 0 GB" is gone** where the truth is "we have not read it yet". -- **The set-aside backups are visible again** — the machine says they are kept and not deleted, and - does **not** pretend they can be reopened, because today they cannot be. +**And much of the journey now works.** The machine showed the recovery screen **without being asked**, +told the customer what was waiting and when it was sealed, said plainly that nobody can replace a lost +code, and accepted the real code first time. The emailed claim code worked first try. -**Still open, and worth knowing:** a rebuilt machine still cannot re-attach its own drives without us -*(R-220 — we are working around it by hand on the test machine right now)*, cannot create a new -recovery code *(R-221)*, and the screen at the machine still shows a stale pairing code *(R-214)*. -**The recovery journey is still recorded as FAILED** — these are fixes, not a re-walk, and it stays -failed until someone walks it end to end with no help from us. +**But it still needed us twice**, and a household has neither hand: + +- **The machine never picks up its own storage connection.** Our hub hands it over and says "the box + will collect this on its next cycle" — the cycle came and went and it did not. Nothing the customer + can click fixes it; it took a command inside the machine. *(R-218 — we had recorded this as fixed; + only half of it was)* +- **A rebuilt machine still cannot re-attach its own drives** — and without them no app can be put + back, so the restore screen stays empty. *(R-220)* + +**Two dead ends, down from four.** The verdict stays **FAILED** until a walk needs us zero times. + +**One thing to decide.** A machine installed today still gets the older software — **the fixes are +built and published but not approved for new machines**. We installed them by hand for this test. So +this proves the journey works on the fixed build; it does **not** prove a customer would receive it. + +## What we fixed this morning, and what it did not fix + +Overnight we tried to break the recovery journey with eleven faults and then left the machine alone +for a full cycle. Nothing lost a byte; the set-aside backups really were kept; the alarm fired and +cleared itself. What it found was that **the machine blamed the customer for failures that were not +theirs** — our hub being unreachable came back as "your recovery code is wrong", in three hundredths +of a second, without the machine even trying. **That is fixed and deployed** *(R-224, R-226)*, along +with three smaller truths: "0 snapshots" where the answer is "we have not looked yet" *(R-225)*, a raw +English error mid-recovery *(R-227)*, and set-aside backups that had become invisible *(R-228)*. + +**None of that shortened the journey**, which is why today's re-walk above still says FAILED — the two +remaining dead ends are different ones. ## What shipped recently diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 2beac2f..f4e0df0 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -41,7 +41,7 @@ | **The offsite repository password can be RECOVERED from the sealed escrow with the customer's recovery code** | hub v0.94.0, agent v0.125.0, controller v0.195.0 | **PROVEN-LIVE (2026-08-04)** | On demo-felhom, through the real endpoints end to end: the box fetched its own sealed blob from the hub with its own per-host credential (hub log: *escrow blob SERVED … 572 opaque bytes, self_scope=true*), the agent unsealed it with the customer's recovery code, and the extracted repository password's sha256 was **byte-identical** to the one on disk — `c60c8bc737a6…`, which is ALSO the hash the hub had independently stored, so three sources agree. Five minutes earlier the same path with a WRONG code failed closed at age's KDF with nothing written, which proves links 6 and 7 ran independently of the success. R was searched for afterwards and found in 0 log lines and 0 files, with a positive control confirming the search would have found it. Evidence: per-repo CHANGELOGs; `audits/RECON-offsite-dr-chain-2026-08-04.md` §3 links 6–8 | **WHAT THIS ROW DOES NOT CLAIM, stated because the previous over-claim here was struck out four hours earlier.** It covers the KEY, not the DATA. **R-199 BACK-POINTER (omitted when this row was written): the recovery chain's link inventory and the per-link status live in `audits/RECON-offsite-dr-chain-2026-08-04.md` §3; links 1–8 are walked, 9–11 are not.** **Re-confirmed 2026-08-04 evening by an attempt to prove the DATA half:** the R-201 drill was prepared on demo-hp and **halted before the wipe** — the sentinel file was not in the off-site snapshot (R-203), so the wipe would have destroyed it and proven nothing. **No file has still ever been restored from an off-site backup after a wipe** (`audits/DRILL-r201-offsite-recovery-2026-08-04.md`). The install half of the chain (controller v0.196.0 `--recover-offsite-install`, R-200) is likewise unit-proven only — it has never run against a live recovery. A recovered password has never been **installed** (the diagnostic compares and refuses to write, by design), no existing repository has ever been **reopened** under one, and **no file has ever been restored** from an off-site history via a recovered key. Links 9–11 of the chain are open (R-200's remaining half, R-201). The proof also used a box whose local key still exists — the rebuilt-box case, where there is nothing to compare against, is exactly what the drill covers and it has not run || DR tier by default: PBS + WireGuard base infra on every install, hub-controlled activation | installer v1.15, agent v0.86, hub v0.51 | **PROVEN-LIVE** (2026-07-21) | `DRILL-day0-take2-2026-07-12` §2 (WG enabled both modes, PBS-DR descriptor auto-provisioned ~1s after WG registration, zero operator steps); ships installer v1.15/agent v0.86/hub v0.51 | Live only on demo/drill fleet. (Cited spike was slice-0 mechanics — shipped nothing; corrected.) **⚠ The candidate upgrade to PROVEN-LIVE is WITHDRAWN — the 2026-07-18 rehearsal produced a live counter-example (R-39).** On the reborn N100 the descriptor auto-provisioned and the agent reported `converged state=applied` (16:45:53), yet **the storage is dead**: `pvesm status` → `felhom-pbs: error fetching datastores - 401 Unauthorized` / `inactive`, and a direct probe with the stored credential returns **401 on every endpoint including `/version`** while the WG transport is healthy (handshake 9 s, 27.9 ms RTT) — i.e. authentication failure, not ACL scope. Root cause in the evidence: **the hub minted a SECOND token secret at 16:47:52, two minutes after the agent had applied the first, and `consumed_at` is still NULL**; the converged state machine will not re-apply, and the agent's 15-minute verify loop **cannot even read the credential to notice** (`open /etc/pve/priv/storage/felhom-pbs.pw: permission denied` — non-root agent reading a file it writes through a root wrapper). A tier that reports `applied` while silently unable to authenticate is exactly the shape that must not carry a PROVEN-LIVE badge. See `tests/VALIDATION-n100-rehearsal-2026-07-18.md` F2 and `pbs-dr-state.txt`. **agent v0.89.0 closes the F4 non-default-storage-id gap (R-22) — PROVEN-LIVE 2026-07-17:** the reconcile self-grants the ACL through the root wrapper on a pre-check 403 instead of dead-locking. Reproduced F4 on the demo (marker moved aside = reinstall fresh-state + felhom-offsite ACLs revoked) → next reconcile tick `pbsdr: pre-check 403 … self-granting … (R-22)` → `converged state=adopted` in ~3 s, ACLs self-restored, `pvesm status felhom-offsite`=active, zero operator action. No more one-shot `pveum` grant **2026-07-21 — the R-39 fleet fix SHIPPED (hub v0.68.0 + agent v0.91.2), closing the self-heal chain end to end.** The three defects that let a box be `applied` and dead simultaneously are each addressed: the hub stamps a monotonic `secret_generation` into the descriptor so a credential re-key finally MOVES the content hash the agent re-applies on; the wrapper gains a narrow `read` verb so the non-root agent can read the credential it writes (it never could — `/etc/pve/priv` is 0700 root:www-data, which made the verify loop blind by construction); and `pbs.ProbeAuth` turns a 401 into a loud `auth_failed` that the existing `pbsdrheal` damper escalates to a fresh mint. Plus a consumed_at honesty gauge for the disagreement no single tier can see (box says `applied`, hub's staged secret never consumed). Proven live on felhom-pve: the agent read its credential through the wrapper (`rc=0`) and probed successfully (`credential probe OK storage=felhom-pbs`). **STOP-2 RAN 2026-07-21 AND THE CHAIN CLOSED — 13 SECONDS, operator click to converged.** The operator pressed **Re-issue PBS credentials**; the identical click on 2026-07-18 did nothing at all. Full chain (hub UTC / host CEST = UTC+2): `08:39:31Z` hub mints a fresh secret, **generation 0 → 1**, and the descriptor gains `"secret_generation": 1` — with `token_id` and `fingerprint` **byte-identical**, i.e. exactly the re-key shape that used to be invisible → `10:39:34` the agent READS its credential through the wrapper (leg b — the read that was impossible until v0.91.0) → `10:39:38` **`ERROR pbsdr: the DR endpoint REJECTED this box's credential — the tier is applied and DEAD` `previous_state=applied`** (leg c: the exact R-39 failure state, detected out loud for the first time ever) → `10:39:45` **`one-time token secret consumed`** `secret_len=36` (leg a: **NO short-circuit** — this is the line that never appeared on 2026-07-18) → `10:39:45` `felhom-pbs-apply reconcile` (the set-only wrapper, no `--server`) → `10:39:47` **`pbsdr: converged state=applied`**. Corroboration: the agent marker hash moved to `afbb3b41…` (it was byte-identical to the pre-reissue marker in the failure); `consumed_at` stamped `08:39:45Z`; the on-disk secret's mtime moved `2026-07-18 20:28:52` → `2026-07-21 10:39:45`; a live probe with the NEW credential returns **200**; three consecutive hub reports trace the whole state machine `applied → auth_failed → applied`; and **zero** `pbsdr_selfheal` escalations fired — the box healed through the descriptor path before the damper was ever needed, with exactly ONE mint and ONE consume and no `consumed-failed.json`. **Row upgraded to PROVEN-LIVE (2026-07-21).** Evidence: `felhom-agent/REPORT.md` (2026-07-21). | | **A customer's file survives a machine rebuild and comes back — the whole off-site story, end to end** | hub v0.94.0, agent v0.125.0, controller v0.197.0 | **PROVEN-LIVE (2026-08-04 night drill)** | demo-hp's controller data volume was destroyed and the sentinel deleted from disk. The customer's recovery code then produced `8a9e33aa4da6…` — **byte-identical** to the pre-wipe on-disk key and to the hub's independent record — and installed cleanly on the bare box. `identity_blob` was unchanged across the wipe (572 B, `updated_at` still 11:11:37). Evidence: `audits/DRILL-r201-night-run-2026-08-04.md` §2 | **WHAT IT DOES AND DOES NOT CLAIM.** PROVEN: after a real rebuild the key recovers byte-identical, the EXISTING repository opens (3 snapshots, 42 026 B — the pre-wipe size; not a fresh history), and the customer restore flow returns the file byte-identical. **ALL FOUR MANUAL INTERVENTIONS ARE CLOSED, AND THE CUSTOMER IS NOW OFFERED THE RECOVERY (2026-08-05). ONE QUALIFIER REMAINS AND IT IS NOT THE OLD ONE — state it precisely.** The 2026-08-04 drill needed four interventions between "the key is recoverable" and "the file is back", none of them in any design document (R-204). **Items 1–3 closed in controller v0.198.0 + hub v0.95.0** (the reset code works first time; a Re-issue no longer marks a healthy escrow stale; a unit restore states it returned the app's definition and database and NOT the customer's files). **Item 4 closed in controller v0.199.0 + hub v0.96.0** (a rebuilt box DECLARES `offsite.state=needs_credential` and `internal/offsiteheal` re-arms the stored credential before minting). **AND THE "needs someone who knows to look" HALF IS NOW GONE TOO (controller v0.200.0, R-193):** a full-page recovery screen takes over the landing pages while the hub holds a sealed package this box cannot open, explains that nobody can replace a lost recovery code, takes the code, opens the repository and lists what is in it — proven live on demo-felhom, which is genuinely in that state (`/launcher` → 302 `/recovery`; a wrong code refused with nothing written, three times, no lockout; the code found in no file, no container log and no debug ring, with a planted-copy positive control that first exposed a mis-aimed sweep). **WHAT THE QUALIFIER IS NOW, and it is narrower:** (1) **the final unlock has never been driven with a CORRECT code through the page** — no recovery code was kept for demo-felhom's orphaned history and demo-hp's is operator-held out of band, so the live run exercised the whole chain (handler → agent → hub fetch → age KDF) and stopped at the unseal; the install and listing rest on tests. (2) **putting files back in place is deliberately NOT part of this** — restore stays per-app, and the step after the listing is R-213. (3) **the journey has still not been re-walked end to end since these fixes** — the closures are proven individually, not as one uninterrupted run. That re-walk is one more drill and it is what is owed. Scope: `demo-hp`, a controller-data-volume rebuild — **NOT** a total host loss, and **NOT** a guest reprovision. **⚠ THE RE-WALK RAN ON 2026-08-05 (CAMPAIGN-11) AND IT FAILED. THIS ROW STAYS AS IT IS — it records a key recovering and a file returning, which remains true — but the JOURNEY row below is the honest verdict and this row must not be read as covering it.** The campaign built a throwaway appliance from the published ISO, destroyed it, and walked the customer's route: **the data half PASSED** (three sentinels byte-identical, including a 12 MB binary and a non-ASCII Hungarian filename, out of the pre-wipe snapshot in 16 s) and **the journey half FAILED** — four operator interventions, three needing root on the appliance, and the first of them was the machine telling the customer their perfectly correct recovery code was wrong (R-216). **R-198's retention is NO LONGER unit-proven — it was PROVEN IN PRODUCTION on 2026-08-05** (CAMPAIGN-11 Phase 3): the first supersession since the fix retained `identity_blob` at **572 B, byte-length exact**, carrying the OLD key `626e424670248db3` while the current row moved to the newly minted `e11a6c542b73477a`; `offsite_repo_key_changed` and `escrow_superseded` both fired at the instant of supersession, and the next off-site run REFUSED as `orphaned` rather than starting a fresh history. Evidence: `tests/campaign11-evidence-2026-08-05/journal.md` | -| **The customer's own UNAIDED recovery journey, end to end** | controller v0.201.0, hub v0.97.1, agent v0.125.0 | **FAILED (2026-08-05, CAMPAIGN-11 Phase 1) — and it stays FAILED until a re-walk passes** | A throwaway appliance was built from the published ISO on demo-hp, claimed, given three sentinels, escrowed, backed up off-site, then destroyed — guest purged and both enrolled drives wiped. The rebuild and the customer's route were then walked with **no command line inside the guest**. **Data: PASS.** All three sentinels byte-identical (`beb9175d…`, `7c8cb0ad…`, `e012e76f…`), including a 12 MB binary and `C11-őrszem-ékezetes-árvíztűrő.txt` — the Hungarian filename survived disk → restic → SFTP → Storage Box → restore → disk intact. Restored in **16 s** out of the pre-wipe snapshot `f3d9cd67`, non-destructively. **Journey: FAIL.** Four interventions, three of them root on the appliance: R-216 (a correct code called wrong), R-218 (the store never opened; the box stopped asking), R-219 (the promised listing structurally could not render), R-220 (the app could not be redeployed — its drives were unenrollable). Evidence: `tests/campaign11-evidence-2026-08-05/journal.md` | **WHAT IT DOES NOT CLAIM, and what the v0.201.0/v0.97.1 fixes do and do not change.** Six of the nine findings are fixed and each is red-proofed (R-216, R-218, R-219, R-217, R-222, R-215) — but **fixes are not a journey.** This row goes green only when the walk is repeated end to end and completes with no operator intervention. **Three findings are deliberately still open and each is a live blocker for some flow:** **R-214** (the console never stops showing a stale pairing code), **R-220** (after a rebuild the customer's drives cannot be re-enrolled — the deploy refuses and the wizard's list is empty, so no app can be redeployed onto its own data drive), **R-221** (a rebuilt box cannot run the escrow ceremony at all, because the pbsdr convergence marker survives on the host while the seeded `escrow.pbs_storage_id` is rewritten out of `agent.json`). **And R-216's fix does not by itself make a NEW box work:** with the hub guard corrected but the Day-0 manifest still vouching agent 0.120.0, a new box is **HELD, not served** — it stops being lied to, which is better, but the feature does not work for it until agent 0.125.0 is vouched. **⚠ CAMPAIGN-11 PHASES 2 + 4 RAN OVERNIGHT 2026-08-05/06 AND THIS ROW STAYS **FAIL**.** **What they add:** eleven injected faults against the recovery journey, each with a positive control, and a full unattended scheduled cycle. **The backup promise strengthened** — the off-site tier ran itself at 04:15 on a box rebuilt twice and set aside hours earlier (`snapshot_count` 1 → 2), all five daily jobs fired exactly once, and nothing on the must-not list fired. **R-217's and R-215's fixes were proven live under exactly their faults**, the set-aside proved it does not delete (12 535 KB byte-exact at the far end), and a wrong code was refused three times with nothing written and no lockout. **What they do NOT add: any progress on the JOURNEY.** These faults are **not a re-walk** — no customer route was walked end to end — and four new findings say the journey got no better where it matters: **R-224** (a hub outage and a stopped agent are BOTH reported to the customer as a bad recovery code, in 0.056 s and 0.030 s — no unseal attempted; the agent's own `err` distinguishes them and it is discarded at the HTTP boundary, and R-216's gate answers `source=version` so it cannot see reachability — **Phase 1's headline defect relocated from the version channel to the transport**), **R-226** (M1, the only message telling a customer to check their typing, is unreachable on any box that has re-escrowed), **R-225** (an unread store renders `0 pillanatkép · 0 GB` above a card saying it holds backups — 1 snapshot and 12 535 KB were really there), **R-228** (the set-aside history is recorded in `orphaned_renamed_to` and shown by nobody — `OrphanedRenamedTo` has zero references in any template or handler). **§4.1 is now MEASURED rather than deduced** — the floor IS served, from the box's own rendered `GetFloor()` and a cold-started controller's settle-gate line. **§4.2's positive half is still NOT measured** and needs a rebuild. Evidence: `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md`, `tests/campaign11-evidence-2026-08-05/journal-phase24.md` **2026-08-06 — THE FIVE PHASE-2 FINDINGS ARE FIXED AND THIS ROW STILL SAYS FAIL.** controller **v0.202.0** + agent **v0.126.0** close R-224, R-226, R-225, R-227, R-228. **What changed:** the unlock path now classifies WHY it failed, from the value and never the text, and the message that mentions typing is reachable from exactly one class — a real refusal — with **unknown defaulting to neutral**. Proven live on the venue with the same wrong code and only the hub's reachability changed: `400 → 502 → 400`. **What it does NOT change: this row.** Fixes are not a journey; nothing here walked a customer end to end, and three of the blockers that made Phase 1 fail are still open (**R-220** in particular is worked around BY HAND on the venue — without that unmount no app can be deployed on a rebuilt box). **What still needs proving:** a re-walk with no operator intervention, and specifically the customer-facing messages end-to-end — they were NOT re-driven, because `/recovery` correctly retires itself once the old data is set aside and restoring that state would be the reconfiguration the fix task forbade. Evidence: `felhom-controller/REPORT.md` | +| **The customer's own UNAIDED recovery journey, end to end** | controller v0.201.0, hub v0.97.1, agent v0.125.0 | **FAILED (2026-08-05, CAMPAIGN-11 Phase 1) — and it stays FAILED until a re-walk passes** | A throwaway appliance was built from the published ISO on demo-hp, claimed, given three sentinels, escrowed, backed up off-site, then destroyed — guest purged and both enrolled drives wiped. The rebuild and the customer's route were then walked with **no command line inside the guest**. **Data: PASS.** All three sentinels byte-identical (`beb9175d…`, `7c8cb0ad…`, `e012e76f…`), including a 12 MB binary and `C11-őrszem-ékezetes-árvíztűrő.txt` — the Hungarian filename survived disk → restic → SFTP → Storage Box → restore → disk intact. Restored in **16 s** out of the pre-wipe snapshot `f3d9cd67`, non-destructively. **Journey: FAIL.** Four interventions, three of them root on the appliance: R-216 (a correct code called wrong), R-218 (the store never opened; the box stopped asking), R-219 (the promised listing structurally could not render), R-220 (the app could not be redeployed — its drives were unenrollable). Evidence: `tests/campaign11-evidence-2026-08-05/journal.md` | **WHAT IT DOES NOT CLAIM, and what the v0.201.0/v0.97.1 fixes do and do not change.** Six of the nine findings are fixed and each is red-proofed (R-216, R-218, R-219, R-217, R-222, R-215) — but **fixes are not a journey.** This row goes green only when the walk is repeated end to end and completes with no operator intervention. **Three findings are deliberately still open and each is a live blocker for some flow:** **R-214** (the console never stops showing a stale pairing code), **R-220** (after a rebuild the customer's drives cannot be re-enrolled — the deploy refuses and the wizard's list is empty, so no app can be redeployed onto its own data drive), **R-221** (a rebuilt box cannot run the escrow ceremony at all, because the pbsdr convergence marker survives on the host while the seeded `escrow.pbs_storage_id` is rewritten out of `agent.json`). **And R-216's fix does not by itself make a NEW box work:** with the hub guard corrected but the Day-0 manifest still vouching agent 0.120.0, a new box is **HELD, not served** — it stops being lied to, which is better, but the feature does not work for it until agent 0.125.0 is vouched. **⚠ CAMPAIGN-11 PHASES 2 + 4 RAN OVERNIGHT 2026-08-05/06 AND THIS ROW STAYS **FAIL**.** **What they add:** eleven injected faults against the recovery journey, each with a positive control, and a full unattended scheduled cycle. **The backup promise strengthened** — the off-site tier ran itself at 04:15 on a box rebuilt twice and set aside hours earlier (`snapshot_count` 1 → 2), all five daily jobs fired exactly once, and nothing on the must-not list fired. **R-217's and R-215's fixes were proven live under exactly their faults**, the set-aside proved it does not delete (12 535 KB byte-exact at the far end), and a wrong code was refused three times with nothing written and no lockout. **What they do NOT add: any progress on the JOURNEY.** These faults are **not a re-walk** — no customer route was walked end to end — and four new findings say the journey got no better where it matters: **R-224** (a hub outage and a stopped agent are BOTH reported to the customer as a bad recovery code, in 0.056 s and 0.030 s — no unseal attempted; the agent's own `err` distinguishes them and it is discarded at the HTTP boundary, and R-216's gate answers `source=version` so it cannot see reachability — **Phase 1's headline defect relocated from the version channel to the transport**), **R-226** (M1, the only message telling a customer to check their typing, is unreachable on any box that has re-escrowed), **R-225** (an unread store renders `0 pillanatkép · 0 GB` above a card saying it holds backups — 1 snapshot and 12 535 KB were really there), **R-228** (the set-aside history is recorded in `orphaned_renamed_to` and shown by nobody — `OrphanedRenamedTo` has zero references in any template or handler). **§4.1 is now MEASURED rather than deduced** — the floor IS served, from the box's own rendered `GetFloor()` and a cold-started controller's settle-gate line. **§4.2's positive half is still NOT measured** and needs a rebuild. Evidence: `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md`, `tests/campaign11-evidence-2026-08-05/journal-phase24.md` **2026-08-06 — THE FIVE PHASE-2 FINDINGS ARE FIXED AND THIS ROW STILL SAYS FAIL.** controller **v0.202.0** + agent **v0.126.0** close R-224, R-226, R-225, R-227, R-228. **What changed:** the unlock path now classifies WHY it failed, from the value and never the text, and the message that mentions typing is reachable from exactly one class — a real refusal — with **unknown defaulting to neutral**. Proven live on the venue with the same wrong code and only the hub's reachability changed: `400 → 502 → 400`. **What it does NOT change: this row.** Fixes are not a journey; nothing here walked a customer end to end, and three of the blockers that made Phase 1 fail are still open (**R-220** in particular is worked around BY HAND on the venue — without that unmount no app can be deployed on a rebuilt box). **What still needs proving:** a re-walk with no operator intervention, and specifically the customer-facing messages end-to-end — they were NOT re-driven, because `/recovery` correctly retires itself once the old data is set aside and restoring that state would be the reconfiguration the fix task forbade. Evidence: `felhom-controller/REPORT.md` **⚠ THE RE-WALK RAN 2026-08-06 (R-201, attended) AND THIS ROW STAYS FAIL — but it is a nearer no.** A brand-new appliance was built from the published ISO, given three sentinels, escrowed, backed up off-site, destroyed (guest purged AND both drives wiped), reinstalled and walked. **THE DATA: PASS** — all three sentinels byte-identical, including a 12 MB binary and an accented Hungarian filename **whose NAME BYTES are also identical**, restored in **23 s** out of the pre-destruction snapshot through the customer's own flow. **THE JOURNEY: FAIL — TWO dead ends against Phase 1's four.** (1) **R-218's CONSUME half**: the hub re-staged the credential and said "the box re-consumes on its next cycle"; a full cycle ran (positive control) and it did not — no customer-reachable action fetches it, and the only lever is a command line INSIDE THE GUEST. (2) **R-220**: the drives are still unenrollable after a rebuild, needing a Proxmox-host unmount, without which no app can be redeployed and the restore page stays empty. **The unaided RTO is therefore STILL UNDEFINED**; the attended figure was 30 m 13 s and must not be quoted as the customer number. **What passed and is new:** the recovery screen appeared WITHOUT being sought, answered all three of its questions with a seal date matching the hub exactly, the emailed reset code worked first try, and the unlock was a real 1.528 s unseal that placed the key. **⚠ AND IT DOES NOT CLAIM DELIVERY:** the fresh install landed on controller 0.201.0 / agent 0.125.0 — the vouched versions, neither carrying the fixes — which were installed BY HAND. Fleet delivery needs a golden carrying the fixed controller and a vouched agent, and nothing was vouched. Evidence: `tests/rewalk-r201-2026-08-06/journal.md` | | **Off-site app-data capture covers the paths an app declares MANDATORY — on BOTH drive layouts** | controller v0.197.0 | **PROVEN-LIVE (2026-08-04) — and this row was OPTIMISTIC before it** | On demo-hp, an app on the **system-data fallback** bound `/mnt/sys_drive/userdata/media/books` while the capture set looked in `/mnt/sys_drive/felhom-data/userdata/media/books`: the declared-mandatory directory was in **no** snapshot and the run reported `ok` (R-203). After v0.197.0 the capture log reads `1 mandatory path(s)` and the file is **listed inside the snapshot** — `restic ls -l latest --tag calibre-web` → `-rw-r--r-- 1000 1000 181 … /DRILL-SENTINEL.txt`. Evidence: `audits/DRILL-r201-offsite-recovery-2026-08-04.md` §2 and the v0.197.0 CHANGELOG | **WHAT IT DOES NOT CLAIM.** It covers CAPTURE, not RESTORE: **no file has ever been restored from an off-site backup after a wipe** (R-201, ready to resume). It also does not claim the enrolled-drive layout was ever wrong — it was not, which is precisely why this survived. And a run that still misses a mandatory directory now reports **`incomplete`** rather than `ok`, so this row's guarantee is one the status can express || **Recurring offsite (PBS) whole-guest backups actually LAND, and RESTORE** — local daily + offsite weekly as scheduled work | agent v0.97–0.103, controller v0.174/0.175, hub v0.76.0, host-install 1.20.0 | **PROVEN-LIVE (2026-07-26)** | `audits/SPIKE-r82-phase0-2026-07-26.md`; per-repo CHANGELOGs/REPORTs. **Restore round-trip on demo-hp:** `--selftest=restore-test` against `felhom-pbs:backup/ct/9201/2026-07-26T15:42:42Z` → `pass:true`, `verified:"boot+running"`, **`mount_parity:"ok"`** (`mp0=/var/lib/docker 50G`, `mp1=/mnt/sys_drive 20G`, mp8/mp9 throwaway stand-ins for the archived binds), `source_tier:"pbs"`, 4m5s restore+boot+verify+teardown, scratch band clean afterwards and the live guest untouched. | **This row is distinct from the DR-tier row above, which proves ACTIVATION, not ARRIVAL.** That tier was PROVEN-LIVE as `applied` since 2026-07-21 while demo-felhom held ONE snapshot (a healing artifact) and demo-hp held **zero, ever** — "applied and empty", the R-39 shape one level quieter. **What earns PROVEN-LIVE here:** (a) demo-hp's FIRST EVER offsite backup landed (4.25 GB into a verifiably empty namespace); (b) it **restores into a bootable, mount-complete guest** — `mount_parity` is the non-hollow half, since a boot-only verify cannot see a missing data volume; (c) **the multi-tier quiesce ran through the real UI endpoint** (`POST /api/guest-backup/trigger`, authed+CSRF) and produced **exactly ONE stop/start pair with BOTH backups inside it** — `quiescing 1 stack(s)` 17:01:39 → local done 17:02:56 *"next tier may start (app still quiesced)"* → felhom-pbs snapshotted 17:03:06 → `unquiescing` 17:03:06. **App downtime 1m27s for both tiers**, and the app came back healthy. **Known gaps, recorded not hidden:** the SCHEDULED restore-test still only selects the PRIMARY tier, so the offsite tier is never AUTOMATICALLY restore-tested (the manual/selftest path is proven, the unattended one is not); the hub infers "PBS ⇒ weekly" from storage TYPE rather than a reported cadence; the installer-default fleet flip awaits a full weekly cycle. → **R-82** | | **Restore-proof is UNATTENDED — the scheduler covers EVERY tier, follows the BACKUP rather than the clock, and a failure is heard** | agent v0.104.0 → **v0.121.0**, hub v0.77.0 → **v0.91.0** | **PROVEN-LIVE (2026-08-03)** | per-repo CHANGELOGs; `backlog/SPEC-r85-phase4-5-2026-07-26.md`. Unit red-proofs for tier rotation, restart-survival, the one-heavy-op gate, failure-emits-an-event, and newborn silence. | **The row above is earned by a MANUAL `--selftest=restore-test`; this one is about the SCHEDULED path, and the distinction is the whole point.** Before R-85 the scheduler could only ever see `cfg.Backup.BackupTarget()`, so the offsite tier was never a candidate — and a failed restore-test was a `[WARN]` line with no event at all, which was true for the LOCAL tier that WAS being tested. Now: oldest-first rotation across every configured tier (operator ruling 2026-07-26), persisted so it survives a restart; a restore-test joins the host-wide one-heavy-operation gate so it never contends with a backup over the same tunnel; and the hub raises two DISTINCT operator-tier signals — `restore_test_failed` (broken now) and `restore_test_stale` (unverified, not known-broken), anchored on R-81 so a newborn box never alarms. **Why this was NOT PROVEN-LIVE until now:** rotation had not been observed selecting both tiers across consecutive UNATTENDED cadences — at a 24h cadence a multi-day window — and a single passing run proves the code path, not the schedule. → **R-85**. **R-86 (agent v0.121.0 + hub v0.91.0, 2026-08-03) replaced the schedule and the observation became possible in one afternoon**, because what has to be observed is no longer a multi-day rotation but a RULE: a tier is due when its newest archive that has settled ~24 h has not been proven. **UPGRADE EVIDENCE — a real unattended run on demo-felhom (2026-08-03), triggered by DUE-NESS, not by a timer.** The scheduler's own log: `15:14:38 restore-test tier is DUE … target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z … reason="newest settled archive … has not been proven"` → `proxmox-backup-client restore --crypt-mode=encrypt` under the agent's own token → `15:25:08 gate decision class=guest_destroy guest=990000 allowed=true` → `15:25:14 scratch guest torn down` → **`15:25:14 backup: scheduled restore-test passed archive=felhom-pbs:… duration_s=635.1`**. A **14.5 GB encrypted offsite archive pulled from ep0 over the WAN**, restored, booted, verified and destroyed in 635 s, unattended. **The three things a timer could not show**, all verified after it: the state names THAT archive (`{"felhom-pbs":{"archive":"…2026-07-28T04:49:43Z","proven_at":"2026-08-03T13:25:14Z"}}`); a second evaluation reports `due=false … is already proven` and runs nothing; and **an agent restart runs nothing**, which is the defect a person actually noticed — every deploy used to restart the timer. Teardown verified at all three layers: guest absent from `pct list`, **zero** `990000` volumes in `lvs`, and the hub-side `restore_tests[]` entry deliberately RETAINED (it IS the proof the staleness check reads). **R-189 (agent v0.122.0, 2026-08-03) closed the reporting half of this row, and it was a REAL gap in the evidence path:** the proof above reached the hub only because no restart intervened — `restore_tests[]` came solely from an in-memory store, so the 15:25:14 PASS was in fact LOST when the agent restarted 2 m 43 s later for a deploy (`0 restore-tests` on the next two host-reports). Under per-archive due-ness the box would not have repeated the work for a week. The persisted per-tier proof (with the archive, and now the tier) is merged into the report, so a proof survives a restart — one entry per tier, newest wins, and a record that cannot be described honestly is not emitted. **THE HOST TIER'S HALF OF THIS ROW WAS OPTIMISTIC UNTIL 2026-08-03, and it is worth saying plainly.** Every live restore-test cited here is on the OFFSITE tier. The HOST tier was not merely unproven on demo-felhom — it was **unprovable**, and on demo-hp equally: the agent's token had no ACL on `/storage/felhom-backup`, the storage both boxes configure as `local_backup_target`, so the content API answered `{"data":[]}` through the token while root listed three archives. The scheduler skipped it as *"no settled archive yet"* — which is exactly what a brand-new tier reports — so nothing ever said so (**R-185**). **CLOSED 2026-08-03 — agent v0.123.0 + installer 1.24.0.** The grant is applied on both demo boxes (the token now lists 3 and 4 archives), the installer's reuse arm grants on a pre-existing target, and the agent now **asks whether it may read each tier it depends on** instead of inferring it from an empty list: a critical degraded capability naming the storage and the missing role, which the hub alerted and emailed on **while the box was still blind**. **THE HOST TIER IS NOW PROVEN-LIVE, UNATTENDED, ON BOTH DEMO BOXES (2026-08-04) — this row's optimistic half is cashed.** Four SCHEDULED runs, nothing triggered by hand: **demo-felhom** host tier `…2026_08_02-04_42_14.tar.zst` **passed in 83.8 s** at 00:55, offsite `…2026-07-28T04:49:43Z` **passed in 540.4 s** at 06:55; **demo-hp** host tier `…2026_08_02-04_49_29.tar.zst` **passed in 109.3 s** at 02:05, offsite `…2026-07-28T19:19:45Z` **passed in 300.1 s** at 08:05. Each restored into a scratch guest, booted, verified and destroyed itself; `pct list` and `lvs` show zero `990000` afterwards on both, and both boxes' `local-lvm` returned to their pre-run figures (1.95 % and 30.83 %). **Both boxes had BOTH tiers due at once and the ordering was observed live for the first time under R-86's rule:** never-proven sorted first, so each box took its HOST tier, deferred the offsite one, and picked it up on the following evaluation six hours later — one heavy operation at a time, per box, without anyone sequencing it. **The proofs reached the hub**, which is R-189 carrying a host-tier entry for the first time: demo-felhom's report holds **two** entries, one per tier, and the `local` one can only have come from the persisted state because the in-memory store held only that morning's offsite run. **SCOPE, stated because one box proving something does not make it a fleet property:** this covers `demo-felhom` and `demo-hp`. The tester's box is untested and untouched. Evidence: `felhom-agent/REPORT.md`, `felhom.eu/REPORT.md` | | Customer RESET (middle lifecycle tier: host delete < RESET < customer Delete): one operator action → pre-first-install; all operational state destroyed, identity + basic config survive | hub v0.61.0, felhom-tenantsync v1.1.0 | **PROVEN-LIVE (external teardown, incl. two real firings)** | **`tests/VALIDATION-n100-rehearsal-2026-07-18.md` — two live firings, both host-delete-first, on two different customers** (`demo-vm-felhom` 15:49:57, `demo-felhom` 16:08:51): every leg `ok` (`claim`, `db_purge`, `descriptor`, `hetzner`, `pbs`), escrow acked separately, each completing in 8–9 s (`hub-state.txt` `customer_resets`). The **Hetzner sub-account destruction is now verified against the live pool box** — and produced the run's sharpest lesson: **a sub-account is an access-control object, not a data object.** Deleting it left its `/home` intact, so re-enabling offsite recreated an account over the previous lifecycle's ciphertext under a key this same RESET had destroyed — which is why the orphan guard fired at 16:58:14 (**a finding by S7's own criterion**) and why RESET now needs a base-dir purge → **R-32**. Prior: hub v0.61.0 REPORT; **ep0 live drill 2026-07-17** (throwaway `drill-reset-01` with a real backup: deprovision `deleted:true` destroyed the namespace + backup group + token, idempotent re-run `deleted:false`, all 3 real tenants + shared user survived); red-proofs (ack-gate, partial-failure resumability) + orchestration/store/offsite/render tests | External teardown FIRST, DB purge LAST, every leg idempotent; refuses while any host row exists; separate escrow-custody ack; clears claim (fresh code next onboarding); keeps the offsite tier CHOICE, drops provisioned fields. **Live-clicked 2026-07-18** (twice, by Viktor) — this supersedes the earlier "not live-clicked / Hetzner delete unit-tested only" note. **R-25b CLOSED (hub v0.69.0, 2026-07-21):** the Danger-zone DELETE is now the guided full-teardown cascade that runs this very sequence as its middle leg — see the row below | diff --git a/documentation/audits/CAMPAIGN-11-recovery-journey-2026-08-05.md b/documentation/audits/CAMPAIGN-11-recovery-journey-2026-08-05.md index f5a5411..cd750b5 100644 --- a/documentation/audits/CAMPAIGN-11-recovery-journey-2026-08-05.md +++ b/documentation/audits/CAMPAIGN-11-recovery-journey-2026-08-05.md @@ -35,6 +35,22 @@ Evidence: `../tests/campaign11-evidence-2026-08-05/` — `journal.md` (Phases 0, > which stays **FAIL** until a re-walk passes. **R-214, R-220, R-221 remain open**, and R-220 is still > worked around by hand on this venue. +> **ADDENDUM 2026-08-06 — THE RE-WALK (R-201). The body below is NOT rewritten.** +> +> Phase 1's question was asked again on the fixed build, on a **new** appliance (VM 322, customer +> `rewalk`) — this campaign's venue was left untouched. **The data half PASSED again**: all three +> sentinels byte-identical, including an accented Hungarian filename whose **name bytes** are also +> identical, restored in **23 s**. **The journey half still FAILS, with TWO dead ends instead of +> four**: R-218's *consume* half (the hub re-stages, the box never collects — only a guest command +> line moves it) and R-220 (drives unenrollable after a rebuild). **The unaided RTO remains +> undefined.** +> +> Two of this campaign's findings were reproduced live: **R-216 part 4** (the reinstall downgraded the +> hand-installed agent back to the vouched version) and **R-220**. One of last night's fixes was seen +> working in the wild: **R-225** (an unread store said "unknown", not a false zero). +> +> Evidence: `../tests/rewalk-r201-2026-08-06/journal.md`. + ## 1. Venue and baselines | | | diff --git a/documentation/tests/rewalk-r201-2026-08-06/journal.md b/documentation/tests/rewalk-r201-2026-08-06/journal.md index ec8b24d..460436c 100644 --- a/documentation/tests/rewalk-r201-2026-08-06/journal.md +++ b/documentation/tests/rewalk-r201-2026-08-06/journal.md @@ -234,3 +234,172 @@ DR Recipe: present · Key Escrow: present ``` **Phase A gate: PASSED.** All seven records taken, sentinels listed **by name**. + +--- + +## Phase B — the journey + +**The rule: no command line inside the guest, at any point.** After the destruction the only things +that reached the guest were HTTP requests a browser could have made — plus the interventions counted +below, which is exactly why they are counted. + +| # | step | result | +|---|---|---| +| 1 | **Destroy** — 11:21:12 | guest 9201 purged (both LVs), **both drives wiped to 4.0 K**. Host identity `rewalk-1ab77d` survived | +| 2 | **Reinstall** | `felhom-host-install.sh` **v1.25.0** fetched live from `felhom.eu/scripts/`; **Day-0 provision SUCCESS 11:25:03**, 2 m 25 s | +| 3 | **Fixed versions** | **hand step — and it was needed a SECOND time** (below) | +| 4 | **Claim back** | the emailed reset code (generation 2) worked **first try**, accents and all | +| 5 | **Log in** | **the recovery screen appeared without being sought**: `/` → `/launcher` → **`/recovery`** | +| 6 | **Read the screen** | all three questions answered (below) | +| 7 | **Enter the code** | HTTP 200 in **1.528 s** — a real unseal; key **recovered and placed** | +| 8 | **The listing** | **did not render** — the tier was not up. Dead end 1 | +| 9 | **Restore** | **all three sentinels byte-identical** | + +### The reinstall DOWNGRADED the agent — R-216 part 4, live again + +``` +agent BEFORE the rebuild : 0.126.0 (hand-installed in Phase A) +agent AFTER the rebuild : 0.125.0 (the vouched version) +``` + +**An operator who fixes a box by hand has it re-broken by the next rebuild — which is precisely the +event that makes the recovery feature necessary.** Re-applied by hand, as §3.2 directs. + +### Step 6 — the screen, read as a customer + +> „Ezt a gépet újratelepítették. A korábbi, **házon kívüli mentéseid megvannak** — a Felhom központi +> rendszere őriz hozzájuk egy lezárt csomagot, amelyet **2026-08-06T09:14:10Z** zártunk le." +> +> „**A helyreállítási kódot senki nem tudja pótolni** — sem a Felhom, sem az ügyfélszolgálat, sem az +> üzemeltető… Ha a kód elveszett, a korábbi mentések nem nyithatók meg többé." +> +> „Ebben a lépésben **semmit nem állítunk vissza és semmi nem változik**." + +All three of step 6's questions answered, and the seal date **matches the hub's `created_at` exactly**. +The set-aside option was correctly **withheld**, with its reason stated rather than the button merely +hidden. *(The seal date still renders as a raw RFC3339 string to a Hungarian household — the copy +defect Phase 1 recorded, still unfixed.)* + +--- + +## The dead ends — TWO, against Phase 1's four + +### Dead end 1 — the off-site tier never came up on its own (R-218's SECOND half) + +**The declaration half works** — that part of R-218's fix is confirmed live: + +``` +11:43:07 recovery: the offsite repository key was recovered and placed (outcome=installed) +11:43:07 recovery: the offsite tier could not be brought up yet: + consume one-time password: no unconsumed offsite password (already consumed…) +11:44:57 (hub) offsiteheal: re-staged the stored one-time offsite secret for rewalk + (declared needs_credential across 2 reports) — the box re-consumes on its next cycle +``` + +**The consume half does not.** The hub re-staged at 11:44:57 and said *"the box re-consumes on its +next cycle"*. **The next cycle came and went** — `host-report from rewalk-1ab77d` at **11:55:46** and +`Received report from rewalk` at **11:55:54**, a full cycle, **with a positive control that the cycle +ran** — and the credential was still not consumed. Twenty-three minutes after the re-stage the box's +last off-site-apply attempt was still **11:43:07**, before it. + +**What the customer sees meanwhile is honest but does not unblock them:** clicking the only relevant +control returns „**A távoli mentési cél nincs beállítva**", and the page says „*Felhom offsite tárhely +kiépítve — a beállítás automatikus, folyamatban. Ha egy napon belül nem áll be, jelezd az +üzemeltetőnek.*" **A census of the customer-reachable actions on that page** — `config`, `reset`, +`run`, `toggle` — **found none that fetches a staged credential.** + +**The lever, and its cost:** `systemctl restart felhom-controller-bootstrap.service` **inside the +guest** — which breaks the journey's pass condition. It worked in **18 seconds** +(Campaign 11 measured 17): + +``` +12:06:16 restart +12:06:34 [offsite-apply] offsite configured for u629488-sub5@…:/home/felhom-repo +``` + +**Which confirms R-218 exactly: nothing was wrong with the credential, the target or the key — the +only thing missing was anything at all to trigger a retry.** + +### Dead end 2 — R-220, the drives, reproduced and red-proved + +`GET /api/disks/candidates` → `initialize: [], attach: []`, while both drives sat mounted at **both** +`/mnt/felhom-drives/` **and** the raw `/mnt/` — the mount that enrolling them created. + +``` +before: initialize: [] attach: [] +after : initialize: [/dev/sdb, /dev/sdc] (fstype ext4, data_bearing true) +``` + +Unmounting only the raw mounts flipped it. **That is a Proxmox-host action a customer cannot perform**, +so it counts. Without it no app can be redeployed, and **without a redeployed app the restore page is +empty** — „Nincs távoli mentésre jelölt alkalmazás" — which is R-213's territory and follows from this +one rather than being separate. + +--- + +## THE VERDICT — both halves, separately + +### The data: **PASS** + +Restored in **23 seconds** out of snapshot **`a7bc23bd`** — the pre-destruction snapshot — through the +customer's own two-step full-restore flow (size gate `12.8 MB`, then confirm), non-destructively. + +| # | file | bytes | expected = restored | +|---|---|---|---| +| A | `REWALK-SENTINEL-A.txt` | 62 | `1573b0e1…bacc705` **BYTE-IDENTICAL** | +| B | `REWALK-őrszem-ékezetes-árvíztűrő.txt` | 66 | `57676fcf…4fe74430` **BYTE-IDENTICAL** | +| C | `REWALK-SENTINEL-C-12MB.bin` | 12 582 912 | `c0faacd7…d9aa14e` **BYTE-IDENTICAL** | + +**And the accented filename's BYTES are byte-identical too** — verified as hex, not as rendered text: + +``` +expected 524557414c4b2dc59172737a656d2dc3a96b657a657465732dc3a17276c3ad7a74c5b172c5912e747874 +restored 524557414c4b2dc59172737a656d2dc3a96b657a657465732dc3a17276c3ad7a74c5b172c5912e747874 +``` + +### The journey: **FAIL** + +**Two steps needed a hand a customer does not have** — one inside the guest, one on the Proxmox host. +**Better than Phase 1's four, and not zero.** + +### The RTO + +| | | +|---|---| +| login (clock start) | **11:42:22** | +| recovery code accepted, key placed | 11:43:07 (**+45 s**) | +| off-site tier up — **after intervention 1** | 12:06:34 (**+24 m 12 s**) | +| all three sentinels restored and verified — **after intervention 2** | 12:12:35 (**+30 m 13 s**) | + +**The unaided RTO remains UNDEFINED**, because the unaided journey still does not complete. **30 m 13 s +is the attended figure** and must not be quoted as the customer number. The only segment that reflects +the product working alone is the last one: **23 seconds to pull 12.8 MB back out of the off-site +repository once everything was in place.** + +--- + +## Harness faults, separated from the product's + +1. **The accented sentinel's filename was destroyed at creation** by the `base64 → bash → pct exec` + chain (U+FFFD), and **a Python `decode('utf-8')` check called it valid** — U+FFFD *is* valid UTF-8. + Only a hex dump exposed it. Rewritten from explicit bytes. +2. **The same trap bit twice more**, in the verification script: a non-ASCII Python literal was mangled + in transit and reported the accented sentinel as **MISSING**. Re-verified keyed on **hashes with no + non-ASCII anywhere in the script**. **Three occurrences in one session: never put non-ASCII inside a + script that crosses this chain.** +3. **`/api/storage/init` needs `fstype`** — omitting it failed with the honest + „nem támogatott fájlrendszer" and I read the first failure as the product's. +4. **Wrong field names** on two endpoints (`app` not `stack`; `path` not `mount_name`), each caught by + the endpoint's own refusal. +5. **A ping alone could not tell a collision from the box's own DHCP lease** — `192.168.0.140` answered + and looked taken; the **MAC** showed it was VM 322 itself. + +## Venue constraints, recorded so they do not inflate the dead-end count + +- The hub's host page shows the guest's LAN address as `—` **by design** (R-66), and this box is + LAN-only with no Cloudflare tunnel, so the address was found from the host's ARP table. A real + customer reaches `felhom.` through the tunnel. **Not a dead end.** +- The claim code arrives **by email**, which is R-119's recorded single human step. The operator + relayed it and it worked **first try**. **Not a dead end.** +- The hub still read "Claimed, generation 1" after the rebuild, because the box's claim state lived in + the destroyed guest; the reset-code path exists for exactly this and worked.