# CAMPAIGN 11 — the recovery journey (2026-08-05 → 08-06) **Four phases. Phase 1 (the clean journey) FAILED on the journey and PASSED on the data. Phase 3 (key supersession) PASSED on its central question. Phase 2 (eleven injected faults) and Phase 4 (an unattended soak) ran overnight on 2026-08-05/06 and are reported here for the first time.** **Findings: R-214 … R-223 from Phases 1 and 3** (seven since fixed and shipped), **plus R-224 … R-228 from Phase 2.** **Three suspicions investigated — two confirmed, one DISPROVED.** **Five harness faults, separated from the product's.** Evidence: `../tests/campaign11-evidence-2026-08-05/` — `journal.md` (Phases 0, 1, 3) and `journal-phase24.md` (Phases 2 and 4, every observable in the order taken). > **The one-line answer to the question this campaign was built to ask.** A Hungarian household whose > machine is rebuilt **gets their data back only if an operator is standing next to them.** The > cryptography, the retention and the transport all work and are now proven live. **What fails is > being told the truth**: on this box, four different situations — a mistyped code, a hub outage, a > stopped agent, and a correct code for a retained earlier package — produce **one** message, and > three of the four are wrong. --- > **ANNOTATION 2026-08-06 — what has since been fixed. The body below is NOT rewritten.** This > document records what was true when the campaign ran, and that is its value; the fixes are recorded > here and in `felhom-controller/REPORT.md`. > > **R-224, R-226, R-225, R-227, R-228 are CLOSED** in controller **v0.202.0** + agent **v0.126.0**. > The unlock path now classifies why it failed — from the value, never the text — and the message > that mentions typing is reachable only after a real refusal; anything unclassifiable renders a > neutral message rather than an accusation. Proven live on this venue: same wrong code, hub up → > `400`, hub REJECTed → `502` naming the connection and stating the code was **not used**, hub > restored → `400`. > > **Unchanged by that work:** the campaign's verdict, the RTO, and the capability map's recovery row, > which stays **FAIL** until a re-walk passes. **R-214, R-220, R-221 remain open**, and R-220 is still > worked around by hand on this venue. > **ADDENDUM 2026-08-06 — THE RE-WALK (R-201). The body below is NOT rewritten.** > > Phase 1's question was asked again on the fixed build, on a **new** appliance (VM 322, customer > `rewalk`) — this campaign's venue was left untouched. **The data half PASSED again**: all three > sentinels byte-identical, including an accented Hungarian filename whose **name bytes** are also > identical, restored in **23 s**. **The journey half still FAILS, with TWO dead ends instead of > four**: R-218's *consume* half (the hub re-stages, the box never collects — only a guest command > line moves it) and R-220 (drives unenrollable after a rebuild). **The unaided RTO remains > undefined.** > > Two of this campaign's findings were reproduced live: **R-216 part 4** (the reinstall downgraded the > hand-installed agent back to the vouched version) and **R-220**. One of last night's fixes was seen > working in the wild: **R-225** (an unread store said "unknown", not a false zero). > > Evidence: `../tests/rewalk-r201-2026-08-06/journal.md`. ## 1. Venue and baselines | | | |---|---| | Host | `demo-hp` (HP t740), **Tier 0**, the designated drill host. Reached **by SSH key, first try** — R-129 stands | | VM | **321 `c11-appliance`** — q35/OVMF, 4 cores, 8 GB, `cpu=host` | | Disks | `scsi0` 200 G system · `scsi1` 50 G · `scsi2` 50 G, qcow2 on `c11-scratch` | | Storage | **`c11-scratch`**, `dir` at **`/mnt/nvme-1tb` — the mount ROOT** (a subdirectory fails the agent's `exactMount` check; Campaign 10 §1) | | Appliance | `192.168.0.105`, hostname `c11`, `felhom-agent 0.125.0` | | Guest | LXC **9201**, **`192.168.0.106`** — DHCP, and it MOVED between phases (`.207` → `.227` → `.106`) | | Hub customer | **`c11`** "Campaign 11", domain `c11.felhom.eu`, DR tier ON, off-site ON | | Host id | **`c11-36d660`** | | Untouched | `drill-r50` (VM 300, stopped), guest 9201 on both demo boxes, `local-lvm`, `felhom-backup`, every other hub customer | ### Baselines — every value re-read fresh at the start of Phase 2 | What | Value | How | |---|---|---| | `felhom-controller` `main` | **v0.201.0** @ `05cf352a2f17` | `HEAD` == `origin/main`, tree clean | | `felhom-agent` `main` | **v0.125.0** @ `0404f60e6a7b` | same | | `felhom.eu` `main` | @ `3a539ea5306a` | same | | hub, LIVE | **`felhom-hub:0.97.1`** | `kubectl … get deploy hub -o jsonpath` | | Day-0 manifest | agent **0.125.0** · golden **0.201.0** · `min_agent` **0.125.0** | hub `/configuration`, `selected` options | | c11 floor | **`v0.200.0 (override)`**; every other customer `v0.156.0` | hub `/configs` | | Highest register ID | **R-223** | fresh `grep -rhoE 'R-[0-9]{1,3}'` over all four repos | **All four cited commits match the brief exactly**, and the manifest matches the golden rebake the Phase-1/3 journal records. **Blast radius of the per-customer floor is still zero, measured** — the other five customers all read `v0.156.0`. ### The hub CHANGELOG/deployed mismatch — diagnosed more precisely than filed The brief records *"`0.97.1` shipped without an entry"*. **The content was not missing; its heading was.** Commit `a7f1d27` wrote the change into the **v0.97.0** entry and `79e31ac` bumped the manifest to `0.97.1`, so a reader matching the running tag against `hub/CHANGELOG.md` found no `v0.97.1` heading at all. **Fixed in this session** — the paragraph now has its own entry, marked as added retroactively. Second occurrence of the class (the first was agent `0.90.1`). --- ## 2. Scope, and what was deliberately not isolated Campaign 10 could run with Tier 3 OFF; **a campaign about off-site recovery cannot.** Off-site hard-requires the DR tier, which provisions on **ep0** (Tier 2). The Phase-0 operator ruling stands and is restated here rather than quietly inherited: **ep0 and the Hetzner Storage Box are written to, additively** — a PBS namespace and token, a WireGuard peer, and a Storage Box sub-account, all created on the ordinary customer path. **Nothing existing is modified or deleted.** The brief's I7 wording ("ep0 read-only") was relaxed by that ruling, not widened by this session. **Phase 2 added no new external writes.** Its faults are network blocks, a service stop, a container restart and a VM shutdown — all on the campaign's own appliance, all reverted, each with a positive control proving the fault was real. --- ## 3. The two positives that were owed (brief §4) ### §4.1 — is the floor actually SERVED? **MEASURED. YES. Twice, independently.** The previous session recorded *"no HELD line and no held reason"* — **two absences** — and correctly refused to call that a measurement. Both positives were taken here. **(a) The box's own rendered state.** `/settings` → „Automatikus frissítés – **Minimális verzió (üzemeltető) `0.200.0`**". That value is `s.updater.GetFloor()`, and `u.floor` has exactly one writer — `SetFloor`, whose only non-test caller is the report-ACK handler. **Both hold branches of `ResolveManagedFloor` set `Floor = ""`** (`store.go:2086`, `:2095`), pinned by `managed_floor_test.go:94`. A non-empty floor on the box therefore proves an ACK carried one. **(b) A cold-started process saying it out loud.** F8's restart produced the decisive A/B — the same log line, same box, same code, before and after the golden rebake: ``` 15:58:35 settle-gate: GO — floor still unknown after 1m30s ← hold IN FORCE 21:06:16 settle-gate: GO — at/above floor 0.200.0 (we are 0.201.0) ← hold RELEASED ``` **The hub half, as a positive rather than an absence:** `managed floor HELD for c11 …` fired every 15 minutes from 17:57:11 to **22:12:06** and then stopped — and the silence is backed by a liveness control (the hub logged `host-report from demo-hp-bb76ea` at 22:44:41 and `wgsync: pushed 5 peers` at 22:44:57, so it was demonstrably still logging). > **A correction to the brief's own plan.** It says to take this at DEBUG during F8's restart. > **A restart alone would not have produced it**: `SetFloor`'s line is `u.dbg(...)`, gated on a private > flag set from `cfg.Logging.Level == "debug"`, and it writes to the logger — **never to the logx debug > ring**, so it cannot appear in `/api/debug/logs` at any level. Confirmed live: 4 000 ring entries > spanning the release window contain no `SetFloor` line, **with a level census run first** (1196 > DEBUG / 2802 INFO / 2 WARN) so the absence was known to be structural rather than evidential. ### §4.2 — R-218's live half: **still NOT measured, deliberately, with the reason** The fix is present and readable — `needsOffsiteCredential` now retires the declaration on the **target**, not on the key: ```go if t != nil { return false } // a TARGET exists — not a rebuild ``` **The venue cannot exercise it.** c11 has a target (`applied_marker` present since 14:57), so `t != nil` and the box correctly does **not** declare; declaring here would be the bug. The state that exercises the fix — hub identity blob present, key placed, **no target** — is shape (a), which the venue held during Phase 1 and does not hold now. F7's set-aside was examined as a route to it and **does not produce it either** (§5, F7). **Recorded as still not measured rather than inferred from the unit test.** It needs one rebuild, which the brief forbids before Phase 4. --- ## 4. Phase 2 — the eleven faults **Judged on the message, not the outcome.** Every fault carries a positive control proving the fault was real; **every control that failed is reported as a failed control, not as a result.** ### The four messages (`internal/web/recovery_handlers.go`, v0.201.0) | | fires when | first words | |---|---|---| | **M1** | unseal failed **and** no superseded package | „A megadott helyreállítási kódot **nem fogadtuk el**. Ellenőrizd, hogy mind a tíz szót…" | | **M2** | the capability gate says the agent cannot do it (R-216) | „Ez a gép **még nem tudja megnyitni** a mentéseidet… **A kódoddal semmi baj**…" | | **M3** | unlocked, inventory unreadable (R-217) | „**A kulcs visszakerült**, de a mentések listáját most nem sikerült beolvasni…" | | **M4** | unseal failed **and** a superseded package is retained (R-222) | „**Ez a kód nem nyitja meg azt a csomagot, amit most őrzünk** ehhez a géphez…" | **The venue retains a superseded package**, so M4's branch — tested **before** M1 on the failure path — is live for every failed unlock. That single ordering fact produces two of the four new findings. ### Results | | Fault | Right answer | Result | |---|---|---|---| | **F1** | wrong code ×3 | refused, **M1 only**, no lockout, nothing written | **PARTIAL** — refused ✅, no lockout ✅, nothing written ✅, but **M4, not M1** → **R-226** | | **F2** | no code at all | states nobody can recover it; offers set-aside | **PASS** | | **F3** | hub unreachable | names **the hub**, never the code | **FAIL** — M4. → **R-224** | | **F4** | agent stopped | **M2**, the R-216 fix under pressure | **FAIL** — M4. → **R-224** | | **F5** | store unreachable after a successful unlock | unlock counts; „could not be read", not „opened with content" | **PASS** — R-217's fix holds | | **F6** | „Most nem", return later | entry point survives; unlock still works | **PASS** | | **F7** | set aside, then a change of mind | set aside, never deleted; honest afterwards | **SPLIT** — set-aside **PASS**, verified byte-for-byte; the afterwards **FAIL** → **R-228** | | **F8** | controller restarted mid-unlock | no half-state; the screen says which | **PARTIAL** — no half-state ✅, but a raw English `Bad Gateway` → **R-227** | | **F9** | a box that never had off-site | nothing, incl. a direct `GET /recovery` | **PARTIAL PASS** — R-215's gate proven live on a narrower shape; the literal precondition was not staged | | **F10** | app with a mandatory data path missing | run reports `incomplete`, names the app, digest carries it | **NOT INJECTED** — harness. Three attempts, each self-healed | | **F11** | box offline a whole reporting window | hub and box agree once it returns | **PASS** — both directions fired, each with an operator mail | ### F7 — the set-aside, and the afterwards **The move-aside is exactly what it claims.** Both confirmations state the consequences first (*„félretesszük — nem töröljük"*, *„a gép új, üres mentési tárolót kezd"*, *„ez az oldal többé nem jelenik meg"*), and the result measured against the Storage Box: ``` BEFORE /home/felhom-repo mtime 13:13 snapshot f3d9cd67 du -s 12535 AFTER /home/felhom-repo.orphaned-20260805 mtime 13:13 snapshot f3d9cd67 du -s 12535 ← untouched /home/felhom-repo mtime 21:32 (empty) ← fresh ``` **Nothing deleted.** Then the customer changes their mind — and finds nothing: `GET /recovery` 302s, `POST /recovery/unlock` 302s in 0.028 s with no message, and `/backups/remote` never mentions the set-aside history at all, though the box records its exact path. → **R-228** ### F9 — what was and was not proven The literal precondition (a box that never had off-site backups) needs a rebuild, which the brief forbids before Phase 4, so it was **not staged**. But **the assertion that failed in Phase 1 was tested and passes**: Phase 1's defect was that `GET /recovery` never asked the predicate. After F7 drives `recoveryOffer()` false, `GET /recovery` → **302 `/backups/remote`**. **R-215's fix is proven live.** The `GetHubEscrowIdentityPresent()==false` arm remains covered by tests only. ### F10 — not injected, and why that is a harness result Three attempts, each with a control confirming the directory was genuinely absent — and each self-healed before capture: (1) the app recreated it; (2) `docker stop` was undone by **the controller's own monitor restarting the stack**; (3) with the stack stopped through the controller API, the directory reappeared at **21:51:38.65**, coincident with the run's own start. All three runs reported `ok` with `1 mandatory path(s)`, and R-203's stat-gap correctly never fired, because by the time `os.Stat` ran the path existed. **The state F10 describes is not reachable on a deployed app of this kind.** Recorded as **harness, not product**. **One observation kept, with its evidence:** at capture the directory held only a recreated `metadata.db` and **not** the customer's `F10-SENTINEL.txt`, and the run still said `ok`. The verdict is about a path's *presence*, not its *content* — correct as designed, and it means an `ok` off-site run can immediately follow the loss of everything that path contained. ### F11 — the dead-man's switch, both ways ``` 23:28:11 ok → stale (host_stale) + Operator email SENT 23:29:59 Received report from c11 ← the box returns unaided 23:30:11 stale → ok (host_recovered) + Operator email SENT 23:30:11 customer mail skipped — no unanswered customer down mail (pairing miss) ``` The customer mail was correctly **withheld** by the pairing gate under a real outage. **I5 holds** — box and hub agree after the return, judged after a full report cycle rather than from one read. **Not reached:** `STALE → DOWN` (>1 h); the outage was ended once both transitions had fired because Phase 4 needed the venue back. ### The timing discriminator used throughout `age`'s scrypt makes a real unseal cost ~1 s. **Phase 1's headline was diagnosed by a 0.134 s response**, and the same instrument separates every fault below: | | elapsed | what it means | |---|---|---| | F1 wrong code | **1.194 / 1.004 / 1.014 s** | a real unseal was attempted and failed | | F3 hub down | **0.0556 s** | **no unseal attempted** — it failed before the KDF | | F4 agent down | **0.0299 s** | **no unseal attempted** | | F5 store down | **1.198 s** | a real unseal, which SUCCEEDED | **The customer sees the same sentence for the 1.0 s case and the 0.03 s case.** --- ## 5. Findings ### R-224 — every non-code failure on the unlock path is reported as a statement about the code **F3 and F4 are one defect with two faces**, and it is Phase 1's headline finding relocated from the version channel to the transport. | | injected (control) | customer sees | machine's own log | |---|---|---|---| | **F3** | hub REJECTed (`302` → `exit 7`) | **M4** | `fetching the sealed bundle: hub: transport error: … no route to host` | | **F4** | `systemctl stop felhom-agent` (`:8443` gone) | **M4** | `dial tcp 169.254.253.1:8443: connect: connection refused` | In both the **correct, current** recovery code was entered, so the only possible cause of failure was the injected fault. **The discriminator exists and is discarded at the HTTP boundary.** The agent's own `err` separates the cases exactly — ``` wrong code : "escrow: the recovery code did not unwrap the identity escrow …" hub down : "escrow: fetching the sealed bundle: hub: transport error: … no route to host" ``` — but both return **HTTP 400** under one merged sentence (*"the recovery code did not open the sealed bundle, **or the bundle could not be fetched**"*), the agent's own `msg=` collapses them too, and the controller's failure path has **no branch for "could not ask / could not reach"** at all. `rerr != nil` falls straight into the code/package messages. **Why R-216's gate did not catch F4**, measured rather than reasoned — from the box's debug ring: ``` [web] recovery capability gate: offsite_key_recovery=yes (source=version) ``` **`source=version`.** The gate answers from the *known agent version* without probe traffic, which is correct for the question it was built for (*is this agent too old?*) and cannot answer the question it is being used for (*can this machine ask right now?*). **A dead agent of the right version sails through it** — and the failure that follows is attributed to the code, which R-216's own comment says must never happen: *"An attempt that cannot succeed must never be made, because its failure is attributed to the code."* **Consequence.** During any hub outage or agent restart, a customer holding a perfect recovery code is told it does not open their package and is routed to customer support about *older* backups. **I6** — an unreachable service reported as a fact about the code. ### R-225 — the store reports `0 snapshots · 0 GB` when it cannot read it, beside a card saying it holds backups Found while checking F6's *"the listing is coherent"* clause. `/backups/remote` renders, **on one screen**: ``` Tároló méret · 0 pillanatkép Tárhelykeret: 0 / 50 GB (0%) ``` immediately above: > „**A távoli tároló másik kulccsal készült mentéseket tartalmaz.** … **A meglévő mentések nem > sérültek**…" **Ground truth, measured directly against the Storage Box over SFTP** — a read-only listing, no decryption, using the box's own transport credential: ``` /home/felhom-repo/snapshots: f3d9cd67d539c00359f0454ea7a782f45beac8406c55275bee5a6886afa8d791 (Aug 5 13:14) /home/felhom-repo: du -s → 12535 KB /home/felhom-repo/keys: exactly ONE key ``` That is the Phase 0 snapshot holding **all three sentinels** — the customer's only surviving copy — matching the journal's `repo_size_bytes 12 611 522`. **Mechanism, from the box's own state:** after the Phase 3 rebuild the `offbox` block in `settings.json` carries **no `snapshot_count` and no `repo_size_bytes` key at all**. The values are *unknown*, and unknown renders as the zero value. **This is R-217's defect class in a second location** — a field whose zero is indistinguishable from a real measurement, defaulted past on an unknown path. `OffsiteInventory.Empty` exists precisely because *"len(Apps)==0 is also what a failed read looks like"*. **I6.** **I5 checked and NOT breached:** the hub's `/offsite` shows `Campaign 11 · 0.0 GB`, but that is 12.5 MB rounded to one decimal of a GB and the pool total (`Used 3.8 GB`) is consistent. The two views do not disagree; **both understate, for different reasons**, and only the box's snapshot **count** — an integer — is false. ### R-226 — M1, the only message that tells a customer to check their typing, is unreachable on any box that has re-escrowed The failure path tests M4's condition **before** M1's: ```go if present, at := s.recoverySuperseded(); present { …M4…; return } …M1… ``` So on every box the hub keeps an earlier package for, a **genuinely mistyped code** produces M4 — which is hedged (*„**Ha** egy korábbi kódot adtál meg"*) and states two true facts, but offers no hint to re-check the ten words and routes the customer to support about older backups. **Measured:** F1's three wrong-code attempts each returned M4, each after a real ~1 s unseal. **Why it matters rather than being a nicety:** the population that has re-escrowed is exactly the population that has just been handed a new recovery code and is most likely to be typing one. R-222's fix removed one conflation (correct-earlier-code read as a mistype) and introduced another (mistype read as a correct-earlier-code) **on the same branch**. ### R-227 — a restart mid-unlock returns a raw English `Bad Gateway` F8 restarted the controller at T+0.7 s, inside the unseal window (control: the container's `StartedAt` moved). The customer got: ``` HTTP 502 · "Bad Gateway" ``` A raw upstream error, in **English**, from traefik. It names no reason, offers no action, and says nothing about whether the key was installed. **I3** — *"every refusal names a reason a person can act on, in Hungarian, with no raw error."* Low severity (the window is ~1 s wide) and recorded rather than inflated. ### R-228 — the set-aside history becomes invisible the moment it is set aside The move-aside is correct and was verified byte-for-byte (§F7). What follows it is not. ```json "orphaned_renamed_to": "/home/felhom-repo.orphaned-20260805" ``` The box records the exact path. And: ``` grep -rn "OrphanedRenamedTo" internal/web/templates/ internal/web/*.go → (no hits) ``` The field is **written and read by nobody**. `/backups/remote` after the set-aside contains no occurrence of the path, „félretéve", „régi előzmény" or any equivalent — with instrument controls passing (`felhom-repo` → 2, „letétbe helyezve" → 1), so the page and the matcher both work. > **12.5 MB of the customer's deliberately retained data sits at a path the box knows and never > shows.** Its only mention is a flash message on the redirect, gone on the next click. **The project's own "seam built but never wired" pattern** — the fifth recorded instance — landing on the one promise the set-aside screen makes. The fix is a plain statement that an earlier history is set aside and not deleted; **it must not promise the history can be reopened**, because R-222 means it cannot be, and that is exactly the conditional promise R-202's gate exists to prevent. --- ## 6. What behaved correctly **A campaign that reports only what broke is half a campaign.** These were tested and held. | | | |---|---| | **R-217's fix** | F5: with the store blocked after a successful unlock, the page rendered M3 and **no listing block at all**. The four false-claim strings („A tároló megnyílt", „van benne tartalom", „nem tudtuk alkalmazásokhoz rendelni") are absent — verified in UTF-8 with two accented positive controls present, after the documented accented-substring false-zero trap was accounted for | | **No lockout** | F1: three wrong codes, three identical responses, no rate limit, no refusal to try again — as `§8.4 NO LOCKOUT` documents | | **Nothing written on failure** | F1: all four `/data/offbox` files byte-identical, **mtimes frozen at 14:57:35** | | **I4** | the entered code appears in **no** log, file or page. Sweep run with a planted-canary control first; the only hits were the harness's own script | | **The „Most nem" asymmetry** | F6: the full page stops interrupting `/launcher` and `/dashboard`, while `/recovery` stays 200 and the backups-area entry point survives — bound to the offer, never to the postpone flag | | **No half-state** | F8: a restart mid-unlock left the key files and `settings.json` coherent; the controller returned healthy in 40 s | | **R-215's fix** | the `GET /recovery` gate is present in v0.201.0 and consults the same predicate as the POST sibling | | **R-198's retention, in the UI** | the hub host page reads „Key Escrow: present · **1 superseded escrow blob(s) retained**" | | **The empty submission** | F2: „Add meg a helyreállítási kódot." in 0.027 s — no unseal, and no blame attached to a code never given | --- ## 7. Harness faults, separated from the product's **Four, all mine, none a product defect** — recorded because Campaign 10's §4d is the format and because two of them nearly produced false findings. 1. **`source ~/.config/credentials` echoes secrets.** The file holds keys with hyphens (`R_DEMO-HP=…`) that bash cannot assign, and the `command not found` error **prints the value**. Two demo-box recovery codes were printed this way before the helper was rewritten to `grep` the one key it needs. **A live I4 hazard for any session that sources that file** — worth a memory, not a register row. 2. **Quoted values in that file** — a bare `cut -d=` keeps the quotes and yields the wrong secret; the hub returned `302` until they were stripped. Already in project memory; re-confirmed. 3. **The recovery page carries no ``** — it renders a hidden `_csrf` input (`data["CSRFField"]`). The first F1 run produced three `403`s and `CSRF-LEN-0`; **the harness's own length check caught it**, and the controller's log named the reason word (`token mismatch`). Harness, not product. 4. **Two failed reachability controls, reported as failures.** The F5 probe first used `nc`, which the controller container does not have, so blocked and unblocked printed the same fallback; the second attempt used the IPv6 address `getent hosts` prefers, and **the container has no IPv6 route at all**. Only a `/dev/tcp` probe against the IPv4 address (`91.98.242.176`) discriminated `TCP-OPEN` → `TCP-CLOSED`. **Two readings that looked like results and were instrument failures.** --- ## 8. Suspicions investigated | | verdict | |---|---| | *"The floor is still held — there is no HELD line and no held reason"* | **DISPROVED.** Both absences were real, and both are explained: the hub stops logging when it stops holding, and `SetFloor`'s DEBUG line cannot reach the debug ring at all. The floor is served, measured two ways (§3) | | *"The hub and the box disagree about the store's contents (I5)"* | **DISPROVED.** Both understate; the hub's `0.0 GB` is rounding of 12.5 MB. Only the box's snapshot **count** is false → R-225, which is an I6 finding, not an I5 one | | *"`GET /recovery` still renders on a box it should not"* | **CONFIRMED FIXED** — the gate is present in v0.201.0 (§9, F9) | --- *(Sections 9–12 — F7, F9, F10, F11, Phase 4, invariants, teardown and hygiene — follow below.)* --- ## 8b. Phase 4 — the unattended soak **23:56 → 04:35, venue untouched.** 30 five-minute samples plus a full log census. **Expectations were pre-registered in the journal BEFORE the window**, in both directions, so nothing here is fitted afterwards. ### Everything scheduled fired, exactly once, on time | job | due | result | |---|---|---| | `db-dump` | 02:30 | ✅ 5.237 s | | `tier2-backup` | 03:30 | ✅ 118 ms — **and the copy is real** (below) | | `fill-watch` | 03:30 | ✅ | | `metrics-prune` | 04:00 | ✅ | | **`offbox-backup`** | **04:15** | ✅ **24.884 s — `snapshot_count` 1 → 2, `last_status: ok`** | **The off-site tier runs itself, unprompted, on a box that was rebuilt twice and had its repository set aside four hours earlier.** That is the strongest positive of the whole campaign for the backup promise, as distinct from the recovery journey. ### Nothing fired that should not have `needs_credential` 0 · `offsiteheal` 0 · `offbox_repo_orphaned` 0 in-window · `offsite_repo_key_changed` 0 · `escrow blob SERVED` 0 · `host_stale`/`host_recovered` 0 · self-update 0 · operator emails 0. > **§4.2's negative control passes.** A box that HAS a target did not declare `needs_credential` once > in five hours. R-218's fix is not over-firing in the other direction — which does **not** substitute > for its positive half, still owed. ### Investigated and DISPROVED — `tier2-backup` in 118 ms It looked like a scheduled backup that silently no-ops. It is not: `Tier 2 copied calibre-web → /mnt/felhom-drives/mentes/backups/secondary/calibre-web (818.5 KB, 1 leg(s), 0s)`, and the copy is on the backup drive. 388 KB on local NVMe in 118 ms is honest. **No finding.** ### A correction to my own pre-registration I pre-registered a `backup_run_digest` event. **No such event type exists** — that is a *test filename*, and the real one is **`backup_run_failures`**, a failures digest whose silence on a clean night is correct. Reporting it as a miss would have been a finding invented by a bad reading. **What survives the correction:** the off-site run emitted **no hub event at all**, while both lesser tiers announced success (`db_dump_completed`, `crossdrive_completed`). Failures are covered by `backup_run_failures` and staleness by the hub's tier deadline monitor (`offsiteBackupStaleAfter = 8 days`), so this is **a consistency wrinkle, not a blind spot** — recorded, not filed. ### What should have fired and did not — answered in both directions - **The restore-test did not run, and that is CORRECT** — `eval_interval=6h` (armed 23:29:43, first evaluation ~05:29) with a **24 h settle** on archives hours old. **Pre-registered as "expected not to run"**; without that, its silence would have read as a gap. - **The agent's whole-guest tier did not run, and whether it should have is NOT RESOLVABLE.** The box reports `0 backups` all night; the agent is emphatically alive (**2 091 of its own lines** since 00:00, polling the guest at 04:17:57); **but routine local-api requests are not logged at INFO** — a five-hour search for `local-api` returns **0** on a box that demonstrably served such calls earlier — so "no `/backup/due` poll" is **not evidence**. The hub's deadline monitor is equally invisible at INFO. **Recorded as an open question, not scored as a pass.** --- ## 9. Invariants at every phase boundary | | | verdict | |---|---|---| | **I1** | no customer data destroyed without an explicit confirmation naming what is lost | **HELD.** The only destructive act was F7's set-aside, behind **two** confirmations that name the consequences first; and it does not destroy — `du -s 12535` before and after | | **I2** | a green status never coexists with missing mandatory data | **HELD as tested, with a caveat.** F10 could not be injected, so the intended test never ran. The caveat is recorded rather than scored: an `ok` run captured a recreated directory that no longer held the customer's file | | **I3** | every refusal names a reason a person can act on, in Hungarian, with no raw error | **BREACHED twice.** F8's `Bad Gateway` (raw, English) → R-227. And R-220's refusal still names an action the customer cannot perform (the list it points at is empty) — reproduced live a third time | | **I4** | no secret anywhere — code, repository password, blob, credential | **HELD on the product.** The F1 sweep found the entered code in no log, file or page, with a planted-canary control passing first. **Breached by the HARNESS**, not the product: `source ~/.config/credentials` echoed two demo-box recovery codes into the session transcript (§7) | | **I5** | the hub's view and the box's never disagree about protection | **HELD.** Checked at the F11 boundary after a full report cycle. The apparent `0.0 GB` disagreement was investigated and DISPROVED (rounding) | | **I6** | an absence is never reported as a fact | **BREACHED twice.** R-224 (an unreachable hub/agent reported as a fact about the code) and R-225 (an unread store reported as `0 pillanatkép · 0 GB`) | | **I7** | nothing reaches a machine outside the venue | **HELD.** Every fault was applied to VM 321 or its guest. Both demo boxes were read, never written. Nothing was deleted on ep0 or the Storage Box — F7's move-aside is a rename **inside the campaign's own sub-account home** | --- ## 10. RTO — unchanged, and why this campaign does not move it **Phase 1's number stands: undefined for an unaided customer**, because the unaided journey does not complete. The attended figure — **61 minutes, three of whose four blockers needed root on the appliance** — is Phase 1's and is not re-measured here. **Phase 2's faults are not a re-walk**, and nothing in them shortens or lengthens that path. **The only segment that reflects the product working remains Phase 1's last one: 16 seconds to pull 12.8 MB back out of the off-site repository once everything was in place.** Phase 2 does add one measurement to the picture: **the off-site tier, once healthy, works unremarkably.** Three manual runs completed in 18 s, 59 s and 1 m 4 s, and the scheduled run is reported in §11. --- ## 11. Teardown — OWED, nothing removed **Deliberately not torn down.** Phase 4 needed the venue, this document had to be written first, and **deleting evidence unattended is worse than leaving a VM running.** All three layers are owed, plus the off-site side. | layer | what it will need | |---|---| | **1. The machine** | VM **321 `c11-appliance`** on `demo-hp` and its three disks (`scsi0` 200 G, `scsi1` 50 G, `scsi2` 50 G) on `c11-scratch`. `qm stop 321 && qm destroy 321 --purge` | | **2. The host** | `pvesm status` on `demo-hp` **before and after**, and the space returned recorded. The `c11-scratch` storage definition itself, once empty | | **3. The hub** | customer **`c11`** and host **`c11-36d660`**: the customer Danger-zone Delete is the one true purge point (it cascades both escrow tables). **⚠ It destroys the retained escrow custody** — which is the point, but say so before pressing it | **The off-site side — and this one needs the most care:** - The campaign's **Storage Box sub-account 284166 / `u629488-sub4`**, holding **two** repositories: `/home/felhom-repo` (the fresh one, ~26 KB) and **`/home/felhom-repo.orphaned-20260805` (12 535 KB — the Phase 0 history with all three sentinels)**. - On **ep0**: namespace `c11`, token `felhom@pbs!c11`, and the two ACL lines on `/datastore/felhom-offsite/c11`. **`demo-felhom` and `demo-hp` namespaces and tokens must not be touched** — the Phase 0 capture recorded all three so teardown can tell them apart. - The WireGuard peer for `c11-36d660` (`10.77.0.5/32`). **Delete nothing that is not the campaign's own.** The brief's §8 warns that scratch customers have accumulated before; the Phase 0 census found **none** outstanding, and c11 must not become the next. --- ## 12. Hygiene - **The recovery codes are SHREDDED**, with the plant→find→shred→fail-to-find control the brief asks for. **The control paid for itself on its first run**: it found the Phase 0 code in `~/.config/credentials` as **`R_CAMPAIGN_11`** — a copy this session did not create and would never have looked for, which would have made a "codes shredded" claim **false**. That key was removed from the shared file carefully (backup → exact-match removal → `diff` proving every other line identical → `HUB_PW` re-verified at `hub:200` → backup shredded). Everything else was `shred -u`'d on both hosts and the absence re-swept; only the planted controls remained, and they were shredded too. **⚠ Consequence, stated plainly: `/home/felhom-repo.orphaned-20260805` (12 535 KB, the three Phase 0 sentinels) is now permanently unopenable.** That is what the set-aside screen promises will happen, teardown removes the repository anyway, and R-222 means no read path existed for it regardless — but the door is now shut for good. - **No secret is written into any committed file.** Every hash quoted here is a sha256 prefix; the codes' contents appear nowhere. - **`git add -A` was never used** — every commit staged explicit paths, and `git status --porcelain` was checked before each one to confirm no foreign file was swept (a parallel session shares this clone). - **The pre-push hook is armed on this clone** (`core.hooksPath=.githooks`) and ran `repo_gates.py --fast` green before every push. **No `--no-verify` was used.** - **No product code was changed.** Every finding was filed and the run continued, per the brief's rule 1. - **No version was bumped** in any repo. --- ## 13. What did not run, and why | | why | |---|---| | **F10 as specified** | **Not injectable.** Three attempts, each with a control: the app, then the controller's monitor, then the run itself recreate the mandatory directory within ~1 s. The state does not exist on a deployed app of this kind. **Harness, not product** | | **F9's literal precondition** | A box that never had off-site backups needs a **rebuild**, which the brief forbids before Phase 4. The assertion that failed in Phase 1 (the page not consulting its predicate) was tested instead, and passes | | **§4.2's positive half** | Needs shape (a) — hub blob present, key placed, **no target**. The venue has a target, so the box correctly does not declare. Also needs a rebuild | | **`STALE → DOWN` escalation** | F11 was ended once `stale` and `recovered` had both fired, because Phase 4 needed the venue back. The >1 h arm is untested | | **The whole-guest tier's due-ness** | **Not resolvable with existing instruments** — routine local-api calls and the hub's deadline monitor are both invisible at INFO. Recorded as an open question, not scored | | **The retained package's read path** | Does not exist (R-199's inventory is unbuilt). R-222 was re-confirmed live rather than re-tested | | **A re-walk of the journey** | **Deliberately out of scope.** These faults are not a re-walk, and the capability map's recovery row stays **FAIL** until one passes | | **Any fix** | Brief rule 1 — file and continue. **No product code was changed and no version bumped** | --- ## 14. Venue state at the end — **WORKING** | | | |---|---| | Hub | `c11-36d660` **ONLINE**, agent `0.125.0`, 1/1 guests, reporting on schedule (last seen 04:29:54) | | Containers | `calibre-web` · `filebrowser` · `felhom-controller:0.201.0` · `traefik` — all healthy | | Drives | both enrolled; backup target `{"degraded":false,"label":"mentes","target":"felhom-backup"}` | | Off-site | **2 snapshots**, `last_status: ok`, `last_success 2026-08-06T02:15:24Z`, 51 206 B, `escrow_state: escrowed` | | Set aside | `/home/felhom-repo.orphaned-20260805` — 12 535 KB, intact, and now **permanently unopenable** (§12) | | Left in place | the raw `/mnt/adatok` and `/mnt/mentes` mounts remain **unmounted** — R-220's workaround, without which no app can be deployed on a rebuilt box. The stable `/mnt/felhom-drives/*` mounts are what everything uses | | Access | the appliance's vaulted root credential was **shredded with the codes**, so a future session must re-fetch it from the hub (`POST /hosts/c11-36d660/reveal-recovery-credential`) — the designed path | **Nothing is left broken, and nothing is left running that should not be.**