Files
felhom.eu/documentation/audits/CAMPAIGN-11-recovery-journey-2026-08-05.md
admin 0c4411e54b
gates / gates (push) Successful in 9s
R-201 re-walk: the data PASSES again, the journey still FAILS — two dead ends, down from four
Asked Campaign 11 Phase 1's question a second time, on the fixed build, on a
NEW appliance (VM 322, customer rewalk). The Campaign 11 venue was untouched.

THE DATA: PASS. All three sentinels byte-identical out of the pre-destruction
snapshot a7bc23bd in 23s through the customer's own restore flow — including a
12 MB binary and an accented Hungarian filename whose NAME BYTES are identical
too (verified as hex, not as rendered text).

THE JOURNEY: FAIL, two dead ends against Phase 1's four.
 1. R-218's CONSUME half. The hub re-staged the credential at 11:44:57 saying
    'the box re-consumes on its next cycle'; a full cycle ran at 11:55:46/54
    (with a positive control that it ran) and it did not. A census of the
    customer-reachable actions found none that fetches it. Only a command line
    INSIDE THE GUEST moved it — 18s, confirming nothing was wrong with the
    credential, target or key: only the trigger. R-218's row said SHIPPED and
    over-claimed; it is corrected to REOPENED for the consume half.
 2. R-220. Drives still unenrollable after a rebuild, needing a Proxmox-host
    unmount; without it no app redeploys and the restore page stays empty.

Unaided RTO STILL UNDEFINED. Attended: +45s key placed, +24m12s tier up,
+30m13s data verified. The 30m must not be quoted as the customer number.

What passed and is new: the recovery screen appeared WITHOUT being sought,
answered all three questions with a seal date matching the hub exactly, the
emailed reset code worked first try, the unlock was a real 1.528s unseal, and
R-225's fix was seen working in the wild (unknown, not a false zero).

R-216 part 4 reproduced live: the reinstall downgraded the hand-installed agent
0.126.0 -> 0.125.0.

DELIVERY GAP recorded as owed and NOT conflated with the journey: a fresh
install landed on controller 0.201.0 / agent 0.125.0 — the vouched versions,
neither carrying the fixes — installed by hand. Nothing was vouched.

Capability map row STAYS FAIL. Campaign 11 doc gets a dated ADDENDUM, not a
rewrite.
2026-08-06 12:18:29 +02:00

641 lines
39 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# CAMPAIGN 11 — the recovery journey (2026-08-05 → 08-06)
**Four phases. Phase 1 (the clean journey) FAILED on the journey and PASSED on the data. Phase 3 (key
supersession) PASSED on its central question. Phase 2 (eleven injected faults) and Phase 4 (an
unattended soak) ran overnight on 2026-08-05/06 and are reported here for the first time.**
**Findings: R-214 … R-223 from Phases 1 and 3** (seven since fixed and shipped), **plus R-224 … R-228
from Phase 2.** **Three suspicions investigated — two confirmed, one DISPROVED.** **Five harness
faults, separated from the product's.**
Evidence: `../tests/campaign11-evidence-2026-08-05/``journal.md` (Phases 0, 1, 3) and
`journal-phase24.md` (Phases 2 and 4, every observable in the order taken).
> **The one-line answer to the question this campaign was built to ask.** A Hungarian household whose
> machine is rebuilt **gets their data back only if an operator is standing next to them.** The
> cryptography, the retention and the transport all work and are now proven live. **What fails is
> being told the truth**: on this box, four different situations — a mistyped code, a hub outage, a
> stopped agent, and a correct code for a retained earlier package — produce **one** message, and
> three of the four are wrong.
---
> **ANNOTATION 2026-08-06 — what has since been fixed. The body below is NOT rewritten.** This
> document records what was true when the campaign ran, and that is its value; the fixes are recorded
> here and in `felhom-controller/REPORT.md`.
>
> **R-224, R-226, R-225, R-227, R-228 are CLOSED** in controller **v0.202.0** + agent **v0.126.0**.
> The unlock path now classifies why it failed — from the value, never the text — and the message
> that mentions typing is reachable only after a real refusal; anything unclassifiable renders a
> neutral message rather than an accusation. Proven live on this venue: same wrong code, hub up →
> `400`, hub REJECTed → `502` naming the connection and stating the code was **not used**, hub
> restored → `400`.
>
> **Unchanged by that work:** the campaign's verdict, the RTO, and the capability map's recovery row,
> which stays **FAIL** until a re-walk passes. **R-214, R-220, R-221 remain open**, and R-220 is still
> worked around by hand on this venue.
> **ADDENDUM 2026-08-06 — THE RE-WALK (R-201). The body below is NOT rewritten.**
>
> Phase 1's question was asked again on the fixed build, on a **new** appliance (VM 322, customer
> `rewalk`) — this campaign's venue was left untouched. **The data half PASSED again**: all three
> sentinels byte-identical, including an accented Hungarian filename whose **name bytes** are also
> identical, restored in **23 s**. **The journey half still FAILS, with TWO dead ends instead of
> four**: R-218's *consume* half (the hub re-stages, the box never collects — only a guest command
> line moves it) and R-220 (drives unenrollable after a rebuild). **The unaided RTO remains
> undefined.**
>
> Two of this campaign's findings were reproduced live: **R-216 part 4** (the reinstall downgraded the
> hand-installed agent back to the vouched version) and **R-220**. One of last night's fixes was seen
> working in the wild: **R-225** (an unread store said "unknown", not a false zero).
>
> Evidence: `../tests/rewalk-r201-2026-08-06/journal.md`.
## 1. Venue and baselines
| | |
|---|---|
| Host | `demo-hp` (HP t740), **Tier 0**, the designated drill host. Reached **by SSH key, first try** — R-129 stands |
| VM | **321 `c11-appliance`** — q35/OVMF, 4 cores, 8 GB, `cpu=host` |
| Disks | `scsi0` 200 G system · `scsi1` 50 G · `scsi2` 50 G, qcow2 on `c11-scratch` |
| Storage | **`c11-scratch`**, `dir` at **`/mnt/nvme-1tb` — the mount ROOT** (a subdirectory fails the agent's `exactMount` check; Campaign 10 §1) |
| Appliance | `192.168.0.105`, hostname `c11`, `felhom-agent 0.125.0` |
| Guest | LXC **9201**, **`192.168.0.106`** — DHCP, and it MOVED between phases (`.207``.227``.106`) |
| Hub customer | **`c11`** "Campaign 11", domain `c11.felhom.eu`, DR tier ON, off-site ON |
| Host id | **`c11-36d660`** |
| Untouched | `drill-r50` (VM 300, stopped), guest 9201 on both demo boxes, `local-lvm`, `felhom-backup`, every other hub customer |
### Baselines — every value re-read fresh at the start of Phase 2
| What | Value | How |
|---|---|---|
| `felhom-controller` `main` | **v0.201.0** @ `05cf352a2f17` | `HEAD` == `origin/main`, tree clean |
| `felhom-agent` `main` | **v0.125.0** @ `0404f60e6a7b` | same |
| `felhom.eu` `main` | @ `3a539ea5306a` | same |
| hub, LIVE | **`felhom-hub:0.97.1`** | `kubectl … get deploy hub -o jsonpath` |
| Day-0 manifest | agent **0.125.0** · golden **0.201.0** · `min_agent` **0.125.0** | hub `/configuration`, `selected` options |
| c11 floor | **`v0.200.0 (override)`**; every other customer `v0.156.0` | hub `/configs` |
| Highest register ID | **R-223** | fresh `grep -rhoE 'R-[0-9]{1,3}'` over all four repos |
**All four cited commits match the brief exactly**, and the manifest matches the golden rebake the
Phase-1/3 journal records. **Blast radius of the per-customer floor is still zero, measured** — the
other five customers all read `v0.156.0`.
### The hub CHANGELOG/deployed mismatch — diagnosed more precisely than filed
The brief records *"`0.97.1` shipped without an entry"*. **The content was not missing; its heading
was.** Commit `a7f1d27` wrote the change into the **v0.97.0** entry and `79e31ac` bumped the manifest
to `0.97.1`, so a reader matching the running tag against `hub/CHANGELOG.md` found no `v0.97.1`
heading at all. **Fixed in this session** — the paragraph now has its own entry, marked as added
retroactively. Second occurrence of the class (the first was agent `0.90.1`).
---
## 2. Scope, and what was deliberately not isolated
Campaign 10 could run with Tier 3 OFF; **a campaign about off-site recovery cannot.** Off-site
hard-requires the DR tier, which provisions on **ep0** (Tier 2). The Phase-0 operator ruling stands
and is restated here rather than quietly inherited: **ep0 and the Hetzner Storage Box are written to,
additively** — a PBS namespace and token, a WireGuard peer, and a Storage Box sub-account, all created
on the ordinary customer path. **Nothing existing is modified or deleted.** The brief's I7 wording
("ep0 read-only") was relaxed by that ruling, not widened by this session.
**Phase 2 added no new external writes.** Its faults are network blocks, a service stop, a container
restart and a VM shutdown — all on the campaign's own appliance, all reverted, each with a positive
control proving the fault was real.
---
## 3. The two positives that were owed (brief §4)
### §4.1 — is the floor actually SERVED? **MEASURED. YES. Twice, independently.**
The previous session recorded *"no HELD line and no held reason"***two absences** — and correctly
refused to call that a measurement. Both positives were taken here.
**(a) The box's own rendered state.** `/settings` → „Automatikus frissítés **Minimális verzió
(üzemeltető) `0.200.0`**". That value is `s.updater.GetFloor()`, and `u.floor` has exactly one writer
`SetFloor`, whose only non-test caller is the report-ACK handler. **Both hold branches of
`ResolveManagedFloor` set `Floor = ""`** (`store.go:2086`, `:2095`), pinned by
`managed_floor_test.go:94`. A non-empty floor on the box therefore proves an ACK carried one.
**(b) A cold-started process saying it out loud.** F8's restart produced the decisive A/B — the same
log line, same box, same code, before and after the golden rebake:
```
15:58:35 settle-gate: GO — floor still unknown after 1m30s ← hold IN FORCE
21:06:16 settle-gate: GO — at/above floor 0.200.0 (we are 0.201.0) ← hold RELEASED
```
**The hub half, as a positive rather than an absence:** `managed floor HELD for c11 …` fired every
15 minutes from 17:57:11 to **22:12:06** and then stopped — and the silence is backed by a liveness
control (the hub logged `host-report from demo-hp-bb76ea` at 22:44:41 and `wgsync: pushed 5 peers` at
22:44:57, so it was demonstrably still logging).
> **A correction to the brief's own plan.** It says to take this at DEBUG during F8's restart.
> **A restart alone would not have produced it**: `SetFloor`'s line is `u.dbg(...)`, gated on a private
> flag set from `cfg.Logging.Level == "debug"`, and it writes to the logger — **never to the logx debug
> ring**, so it cannot appear in `/api/debug/logs` at any level. Confirmed live: 4 000 ring entries
> spanning the release window contain no `SetFloor` line, **with a level census run first** (1196
> DEBUG / 2802 INFO / 2 WARN) so the absence was known to be structural rather than evidential.
### §4.2 — R-218's live half: **still NOT measured, deliberately, with the reason**
The fix is present and readable — `needsOffsiteCredential` now retires the declaration on the
**target**, not on the key:
```go
if t != nil { return false } // a TARGET exists — not a rebuild
```
**The venue cannot exercise it.** c11 has a target (`applied_marker` present since 14:57), so
`t != nil` and the box correctly does **not** declare; declaring here would be the bug. The state that
exercises the fix — hub identity blob present, key placed, **no target** — is shape (a), which the
venue held during Phase 1 and does not hold now. F7's set-aside was examined as a route to it and
**does not produce it either** (§5, F7).
**Recorded as still not measured rather than inferred from the unit test.** It needs one rebuild,
which the brief forbids before Phase 4.
---
## 4. Phase 2 — the eleven faults
**Judged on the message, not the outcome.** Every fault carries a positive control proving the fault
was real; **every control that failed is reported as a failed control, not as a result.**
### The four messages (`internal/web/recovery_handlers.go`, v0.201.0)
| | fires when | first words |
|---|---|---|
| **M1** | unseal failed **and** no superseded package | „A megadott helyreállítási kódot **nem fogadtuk el**. Ellenőrizd, hogy mind a tíz szót…" |
| **M2** | the capability gate says the agent cannot do it (R-216) | „Ez a gép **még nem tudja megnyitni** a mentéseidet… **A kódoddal semmi baj**…" |
| **M3** | unlocked, inventory unreadable (R-217) | „**A kulcs visszakerült**, de a mentések listáját most nem sikerült beolvasni…" |
| **M4** | unseal failed **and** a superseded package is retained (R-222) | „**Ez a kód nem nyitja meg azt a csomagot, amit most őrzünk** ehhez a géphez…" |
**The venue retains a superseded package**, so M4's branch — tested **before** M1 on the failure path
— is live for every failed unlock. That single ordering fact produces two of the four new findings.
### Results
| | Fault | Right answer | Result |
|---|---|---|---|
| **F1** | wrong code ×3 | refused, **M1 only**, no lockout, nothing written | **PARTIAL** — refused ✅, no lockout ✅, nothing written ✅, but **M4, not M1****R-226** |
| **F2** | no code at all | states nobody can recover it; offers set-aside | **PASS** |
| **F3** | hub unreachable | names **the hub**, never the code | **FAIL** — M4. → **R-224** |
| **F4** | agent stopped | **M2**, the R-216 fix under pressure | **FAIL** — M4. → **R-224** |
| **F5** | store unreachable after a successful unlock | unlock counts; „could not be read", not „opened with content" | **PASS** — R-217's fix holds |
| **F6** | „Most nem", return later | entry point survives; unlock still works | **PASS** |
| **F7** | set aside, then a change of mind | set aside, never deleted; honest afterwards | **SPLIT** — set-aside **PASS**, verified byte-for-byte; the afterwards **FAIL****R-228** |
| **F8** | controller restarted mid-unlock | no half-state; the screen says which | **PARTIAL** — no half-state ✅, but a raw English `Bad Gateway`**R-227** |
| **F9** | a box that never had off-site | nothing, incl. a direct `GET /recovery` | **PARTIAL PASS** — R-215's gate proven live on a narrower shape; the literal precondition was not staged |
| **F10** | app with a mandatory data path missing | run reports `incomplete`, names the app, digest carries it | **NOT INJECTED** — harness. Three attempts, each self-healed |
| **F11** | box offline a whole reporting window | hub and box agree once it returns | **PASS** — both directions fired, each with an operator mail |
### F7 — the set-aside, and the afterwards
**The move-aside is exactly what it claims.** Both confirmations state the consequences first
(*„félretesszük — nem töröljük"*, *„a gép új, üres mentési tárolót kezd"*, *„ez az oldal többé nem
jelenik meg"*), and the result measured against the Storage Box:
```
BEFORE /home/felhom-repo mtime 13:13 snapshot f3d9cd67 du -s 12535
AFTER /home/felhom-repo.orphaned-20260805 mtime 13:13 snapshot f3d9cd67 du -s 12535 ← untouched
/home/felhom-repo mtime 21:32 (empty) ← fresh
```
**Nothing deleted.** Then the customer changes their mind — and finds nothing: `GET /recovery` 302s,
`POST /recovery/unlock` 302s in 0.028 s with no message, and `/backups/remote` never mentions the
set-aside history at all, though the box records its exact path. → **R-228**
### F9 — what was and was not proven
The literal precondition (a box that never had off-site backups) needs a rebuild, which the brief
forbids before Phase 4, so it was **not staged**. But **the assertion that failed in Phase 1 was
tested and passes**: Phase 1's defect was that `GET /recovery` never asked the predicate. After F7
drives `recoveryOffer()` false, `GET /recovery`**302 `/backups/remote`**. **R-215's fix is proven
live.** The `GetHubEscrowIdentityPresent()==false` arm remains covered by tests only.
### F10 — not injected, and why that is a harness result
Three attempts, each with a control confirming the directory was genuinely absent — and each
self-healed before capture: (1) the app recreated it; (2) `docker stop` was undone by **the
controller's own monitor restarting the stack**; (3) with the stack stopped through the controller
API, the directory reappeared at **21:51:38.65**, coincident with the run's own start. All three runs
reported `ok` with `1 mandatory path(s)`, and R-203's stat-gap correctly never fired, because by the
time `os.Stat` ran the path existed.
**The state F10 describes is not reachable on a deployed app of this kind.** Recorded as **harness,
not product**.
**One observation kept, with its evidence:** at capture the directory held only a recreated
`metadata.db` and **not** the customer's `F10-SENTINEL.txt`, and the run still said `ok`. The verdict
is about a path's *presence*, not its *content* — correct as designed, and it means an `ok` off-site
run can immediately follow the loss of everything that path contained.
### F11 — the dead-man's switch, both ways
```
23:28:11 ok → stale (host_stale) + Operator email SENT
23:29:59 Received report from c11 ← the box returns unaided
23:30:11 stale → ok (host_recovered) + Operator email SENT
23:30:11 customer mail skipped — no unanswered customer down mail (pairing miss)
```
The customer mail was correctly **withheld** by the pairing gate under a real outage. **I5 holds**
box and hub agree after the return, judged after a full report cycle rather than from one read.
**Not reached:** `STALE → DOWN` (>1 h); the outage was ended once both transitions had fired because
Phase 4 needed the venue back.
### The timing discriminator used throughout
`age`'s scrypt makes a real unseal cost ~1 s. **Phase 1's headline was diagnosed by a 0.134 s
response**, and the same instrument separates every fault below:
| | elapsed | what it means |
|---|---|---|
| F1 wrong code | **1.194 / 1.004 / 1.014 s** | a real unseal was attempted and failed |
| F3 hub down | **0.0556 s** | **no unseal attempted** — it failed before the KDF |
| F4 agent down | **0.0299 s** | **no unseal attempted** |
| F5 store down | **1.198 s** | a real unseal, which SUCCEEDED |
**The customer sees the same sentence for the 1.0 s case and the 0.03 s case.**
---
## 5. Findings
### R-224 — every non-code failure on the unlock path is reported as a statement about the code
**F3 and F4 are one defect with two faces**, and it is Phase 1's headline finding relocated from the
version channel to the transport.
| | injected (control) | customer sees | machine's own log |
|---|---|---|---|
| **F3** | hub REJECTed (`302``exit 7`) | **M4** | `fetching the sealed bundle: hub: transport error: … no route to host` |
| **F4** | `systemctl stop felhom-agent` (`:8443` gone) | **M4** | `dial tcp 169.254.253.1:8443: connect: connection refused` |
In both the **correct, current** recovery code was entered, so the only possible cause of failure was
the injected fault.
**The discriminator exists and is discarded at the HTTP boundary.** The agent's own `err` separates
the cases exactly —
```
wrong code : "escrow: the recovery code did not unwrap the identity escrow …"
hub down : "escrow: fetching the sealed bundle: hub: transport error: … no route to host"
```
— but both return **HTTP 400** under one merged sentence (*"the recovery code did not open the sealed
bundle, **or the bundle could not be fetched**"*), the agent's own `msg=` collapses them too, and the
controller's failure path has **no branch for "could not ask / could not reach"** at all. `rerr != nil`
falls straight into the code/package messages.
**Why R-216's gate did not catch F4**, measured rather than reasoned — from the box's debug ring:
```
[web] recovery capability gate: offsite_key_recovery=yes (source=version)
```
**`source=version`.** The gate answers from the *known agent version* without probe traffic, which is
correct for the question it was built for (*is this agent too old?*) and cannot answer the question it
is being used for (*can this machine ask right now?*). **A dead agent of the right version sails
through it** — and the failure that follows is attributed to the code, which R-216's own comment says
must never happen: *"An attempt that cannot succeed must never be made, because its failure is
attributed to the code."*
**Consequence.** During any hub outage or agent restart, a customer holding a perfect recovery code is
told it does not open their package and is routed to customer support about *older* backups. **I6**
an unreachable service reported as a fact about the code.
### R-225 — the store reports `0 snapshots · 0 GB` when it cannot read it, beside a card saying it holds backups
Found while checking F6's *"the listing is coherent"* clause. `/backups/remote` renders, **on one
screen**:
```
Tároló méret · 0 pillanatkép Tárhelykeret: 0 / 50 GB (0%)
```
immediately above:
> „**A távoli tároló másik kulccsal készült mentéseket tartalmaz.** … **A meglévő mentések nem
> sérültek**…"
**Ground truth, measured directly against the Storage Box over SFTP** — a read-only listing, no
decryption, using the box's own transport credential:
```
/home/felhom-repo/snapshots:
f3d9cd67d539c00359f0454ea7a782f45beac8406c55275bee5a6886afa8d791 (Aug 5 13:14)
/home/felhom-repo: du -s → 12535 KB /home/felhom-repo/keys: exactly ONE key
```
That is the Phase 0 snapshot holding **all three sentinels** — the customer's only surviving copy —
matching the journal's `repo_size_bytes 12 611 522`.
**Mechanism, from the box's own state:** after the Phase 3 rebuild the `offbox` block in
`settings.json` carries **no `snapshot_count` and no `repo_size_bytes` key at all**. The values are
*unknown*, and unknown renders as the zero value.
**This is R-217's defect class in a second location** — a field whose zero is indistinguishable from a
real measurement, defaulted past on an unknown path. `OffsiteInventory.Empty` exists precisely because
*"len(Apps)==0 is also what a failed read looks like"*. **I6.**
**I5 checked and NOT breached:** the hub's `/offsite` shows `Campaign 11 · 0.0 GB`, but that is 12.5 MB
rounded to one decimal of a GB and the pool total (`Used 3.8 GB`) is consistent. The two views do not
disagree; **both understate, for different reasons**, and only the box's snapshot **count** — an
integer — is false.
### R-226 — M1, the only message that tells a customer to check their typing, is unreachable on any box that has re-escrowed
The failure path tests M4's condition **before** M1's:
```go
if present, at := s.recoverySuperseded(); present { M4; return }
M1
```
So on every box the hub keeps an earlier package for, a **genuinely mistyped code** produces M4 —
which is hedged (*„**Ha** egy korábbi kódot adtál meg"*) and states two true facts, but offers no hint
to re-check the ten words and routes the customer to support about older backups.
**Measured:** F1's three wrong-code attempts each returned M4, each after a real ~1 s unseal.
**Why it matters rather than being a nicety:** the population that has re-escrowed is exactly the
population that has just been handed a new recovery code and is most likely to be typing one. R-222's
fix removed one conflation (correct-earlier-code read as a mistype) and introduced another
(mistype read as a correct-earlier-code) **on the same branch**.
### R-227 — a restart mid-unlock returns a raw English `Bad Gateway`
F8 restarted the controller at T+0.7 s, inside the unseal window (control: the container's `StartedAt`
moved). The customer got:
```
HTTP 502 · "Bad Gateway"
```
A raw upstream error, in **English**, from traefik. It names no reason, offers no action, and says
nothing about whether the key was installed. **I3** — *"every refusal names a reason a person can act
on, in Hungarian, with no raw error."* Low severity (the window is ~1 s wide) and recorded rather than
inflated.
### R-228 — the set-aside history becomes invisible the moment it is set aside
The move-aside is correct and was verified byte-for-byte (§F7). What follows it is not.
```json
"orphaned_renamed_to": "/home/felhom-repo.orphaned-20260805"
```
The box records the exact path. And:
```
grep -rn "OrphanedRenamedTo" internal/web/templates/ internal/web/*.go → (no hits)
```
The field is **written and read by nobody**. `/backups/remote` after the set-aside contains no
occurrence of the path, „félretéve", „régi előzmény" or any equivalent — with instrument controls
passing (`felhom-repo` → 2, „letétbe helyezve" → 1), so the page and the matcher both work.
> **12.5 MB of the customer's deliberately retained data sits at a path the box knows and never
> shows.** Its only mention is a flash message on the redirect, gone on the next click.
**The project's own "seam built but never wired" pattern** — the fifth recorded instance — landing on
the one promise the set-aside screen makes. The fix is a plain statement that an earlier history is
set aside and not deleted; **it must not promise the history can be reopened**, because R-222 means it
cannot be, and that is exactly the conditional promise R-202's gate exists to prevent.
---
## 6. What behaved correctly
**A campaign that reports only what broke is half a campaign.** These were tested and held.
| | |
|---|---|
| **R-217's fix** | F5: with the store blocked after a successful unlock, the page rendered M3 and **no listing block at all**. The four false-claim strings („A tároló megnyílt", „van benne tartalom", „nem tudtuk alkalmazásokhoz rendelni") are absent — verified in UTF-8 with two accented positive controls present, after the documented accented-substring false-zero trap was accounted for |
| **No lockout** | F1: three wrong codes, three identical responses, no rate limit, no refusal to try again — as `§8.4 NO LOCKOUT` documents |
| **Nothing written on failure** | F1: all four `/data/offbox` files byte-identical, **mtimes frozen at 14:57:35** |
| **I4** | the entered code appears in **no** log, file or page. Sweep run with a planted-canary control first; the only hits were the harness's own script |
| **The „Most nem" asymmetry** | F6: the full page stops interrupting `/launcher` and `/dashboard`, while `/recovery` stays 200 and the backups-area entry point survives — bound to the offer, never to the postpone flag |
| **No half-state** | F8: a restart mid-unlock left the key files and `settings.json` coherent; the controller returned healthy in 40 s |
| **R-215's fix** | the `GET /recovery` gate is present in v0.201.0 and consults the same predicate as the POST sibling |
| **R-198's retention, in the UI** | the hub host page reads „Key Escrow: present · **1 superseded escrow blob(s) retained**" |
| **The empty submission** | F2: „Add meg a helyreállítási kódot." in 0.027 s — no unseal, and no blame attached to a code never given |
---
## 7. Harness faults, separated from the product's
**Four, all mine, none a product defect** — recorded because Campaign 10's §4d is the format and
because two of them nearly produced false findings.
1. **`source ~/.config/credentials` echoes secrets.** The file holds keys with hyphens
(`R_DEMO-HP=…`) that bash cannot assign, and the `command not found` error **prints the value**.
Two demo-box recovery codes were printed this way before the helper was rewritten to `grep` the one
key it needs. **A live I4 hazard for any session that sources that file** — worth a memory, not a
register row.
2. **Quoted values in that file** — a bare `cut -d=` keeps the quotes and yields the wrong secret; the
hub returned `302` until they were stripped. Already in project memory; re-confirmed.
3. **The recovery page carries no `<meta name="csrf-token">`** — it renders a hidden `_csrf` input
(`data["CSRFField"]`). The first F1 run produced three `403`s and `CSRF-LEN-0`; **the harness's own
length check caught it**, and the controller's log named the reason word (`token mismatch`).
Harness, not product.
4. **Two failed reachability controls, reported as failures.** The F5 probe first used `nc`, which the
controller container does not have, so blocked and unblocked printed the same fallback; the second
attempt used the IPv6 address `getent hosts` prefers, and **the container has no IPv6 route at
all**. Only a `/dev/tcp` probe against the IPv4 address (`91.98.242.176`) discriminated
`TCP-OPEN``TCP-CLOSED`. **Two readings that looked like results and were instrument failures.**
---
## 8. Suspicions investigated
| | verdict |
|---|---|
| *"The floor is still held — there is no HELD line and no held reason"* | **DISPROVED.** Both absences were real, and both are explained: the hub stops logging when it stops holding, and `SetFloor`'s DEBUG line cannot reach the debug ring at all. The floor is served, measured two ways (§3) |
| *"The hub and the box disagree about the store's contents (I5)"* | **DISPROVED.** Both understate; the hub's `0.0 GB` is rounding of 12.5 MB. Only the box's snapshot **count** is false → R-225, which is an I6 finding, not an I5 one |
| *"`GET /recovery` still renders on a box it should not"* | **CONFIRMED FIXED** — the gate is present in v0.201.0 (§9, F9) |
---
*(Sections 912 — F7, F9, F10, F11, Phase 4, invariants, teardown and hygiene — follow below.)*
---
## 8b. Phase 4 — the unattended soak
**23:56 → 04:35, venue untouched.** 30 five-minute samples plus a full log census. **Expectations were
pre-registered in the journal BEFORE the window**, in both directions, so nothing here is fitted
afterwards.
### Everything scheduled fired, exactly once, on time
| job | due | result |
|---|---|---|
| `db-dump` | 02:30 | ✅ 5.237 s |
| `tier2-backup` | 03:30 | ✅ 118 ms — **and the copy is real** (below) |
| `fill-watch` | 03:30 | ✅ |
| `metrics-prune` | 04:00 | ✅ |
| **`offbox-backup`** | **04:15** | ✅ **24.884 s — `snapshot_count` 1 → 2, `last_status: ok`** |
**The off-site tier runs itself, unprompted, on a box that was rebuilt twice and had its repository set
aside four hours earlier.** That is the strongest positive of the whole campaign for the backup
promise, as distinct from the recovery journey.
### Nothing fired that should not have
`needs_credential` 0 · `offsiteheal` 0 · `offbox_repo_orphaned` 0 in-window · `offsite_repo_key_changed`
0 · `escrow blob SERVED` 0 · `host_stale`/`host_recovered` 0 · self-update 0 · operator emails 0.
> **§4.2's negative control passes.** A box that HAS a target did not declare `needs_credential` once
> in five hours. R-218's fix is not over-firing in the other direction — which does **not** substitute
> for its positive half, still owed.
### Investigated and DISPROVED — `tier2-backup` in 118 ms
It looked like a scheduled backup that silently no-ops. It is not: `Tier 2 copied calibre-web →
/mnt/felhom-drives/mentes/backups/secondary/calibre-web (818.5 KB, 1 leg(s), 0s)`, and the copy is on
the backup drive. 388 KB on local NVMe in 118 ms is honest. **No finding.**
### A correction to my own pre-registration
I pre-registered a `backup_run_digest` event. **No such event type exists** — that is a *test
filename*, and the real one is **`backup_run_failures`**, a failures digest whose silence on a clean
night is correct. Reporting it as a miss would have been a finding invented by a bad reading.
**What survives the correction:** the off-site run emitted **no hub event at all**, while both lesser
tiers announced success (`db_dump_completed`, `crossdrive_completed`). Failures are covered by
`backup_run_failures` and staleness by the hub's tier deadline monitor
(`offsiteBackupStaleAfter = 8 days`), so this is **a consistency wrinkle, not a blind spot**
recorded, not filed.
### What should have fired and did not — answered in both directions
- **The restore-test did not run, and that is CORRECT** — `eval_interval=6h` (armed 23:29:43, first
evaluation ~05:29) with a **24 h settle** on archives hours old. **Pre-registered as "expected not to
run"**; without that, its silence would have read as a gap.
- **The agent's whole-guest tier did not run, and whether it should have is NOT RESOLVABLE.** The box
reports `0 backups` all night; the agent is emphatically alive (**2 091 of its own lines** since
00:00, polling the guest at 04:17:57); **but routine local-api requests are not logged at INFO** — a
five-hour search for `local-api` returns **0** on a box that demonstrably served such calls earlier —
so "no `/backup/due` poll" is **not evidence**. The hub's deadline monitor is equally invisible at
INFO. **Recorded as an open question, not scored as a pass.**
---
## 9. Invariants at every phase boundary
| | | verdict |
|---|---|---|
| **I1** | no customer data destroyed without an explicit confirmation naming what is lost | **HELD.** The only destructive act was F7's set-aside, behind **two** confirmations that name the consequences first; and it does not destroy — `du -s 12535` before and after |
| **I2** | a green status never coexists with missing mandatory data | **HELD as tested, with a caveat.** F10 could not be injected, so the intended test never ran. The caveat is recorded rather than scored: an `ok` run captured a recreated directory that no longer held the customer's file |
| **I3** | every refusal names a reason a person can act on, in Hungarian, with no raw error | **BREACHED twice.** F8's `Bad Gateway` (raw, English) → R-227. And R-220's refusal still names an action the customer cannot perform (the list it points at is empty) — reproduced live a third time |
| **I4** | no secret anywhere — code, repository password, blob, credential | **HELD on the product.** The F1 sweep found the entered code in no log, file or page, with a planted-canary control passing first. **Breached by the HARNESS**, not the product: `source ~/.config/credentials` echoed two demo-box recovery codes into the session transcript (§7) |
| **I5** | the hub's view and the box's never disagree about protection | **HELD.** Checked at the F11 boundary after a full report cycle. The apparent `0.0 GB` disagreement was investigated and DISPROVED (rounding) |
| **I6** | an absence is never reported as a fact | **BREACHED twice.** R-224 (an unreachable hub/agent reported as a fact about the code) and R-225 (an unread store reported as `0 pillanatkép · 0 GB`) |
| **I7** | nothing reaches a machine outside the venue | **HELD.** Every fault was applied to VM 321 or its guest. Both demo boxes were read, never written. Nothing was deleted on ep0 or the Storage Box — F7's move-aside is a rename **inside the campaign's own sub-account home** |
---
## 10. RTO — unchanged, and why this campaign does not move it
**Phase 1's number stands: undefined for an unaided customer**, because the unaided journey does not
complete. The attended figure — **61 minutes, three of whose four blockers needed root on the
appliance** — is Phase 1's and is not re-measured here. **Phase 2's faults are not a re-walk**, and
nothing in them shortens or lengthens that path.
**The only segment that reflects the product working remains Phase 1's last one: 16 seconds to pull
12.8 MB back out of the off-site repository once everything was in place.**
Phase 2 does add one measurement to the picture: **the off-site tier, once healthy, works
unremarkably.** Three manual runs completed in 18 s, 59 s and 1 m 4 s, and the scheduled run is
reported in §11.
---
## 11. Teardown — OWED, nothing removed
**Deliberately not torn down.** Phase 4 needed the venue, this document had to be written first, and
**deleting evidence unattended is worse than leaving a VM running.** All three layers are owed, plus
the off-site side.
| layer | what it will need |
|---|---|
| **1. The machine** | VM **321 `c11-appliance`** on `demo-hp` and its three disks (`scsi0` 200 G, `scsi1` 50 G, `scsi2` 50 G) on `c11-scratch`. `qm stop 321 && qm destroy 321 --purge` |
| **2. The host** | `pvesm status` on `demo-hp` **before and after**, and the space returned recorded. The `c11-scratch` storage definition itself, once empty |
| **3. The hub** | customer **`c11`** and host **`c11-36d660`**: the customer Danger-zone Delete is the one true purge point (it cascades both escrow tables). **⚠ It destroys the retained escrow custody** — which is the point, but say so before pressing it |
**The off-site side — and this one needs the most care:**
- The campaign's **Storage Box sub-account 284166 / `u629488-sub4`**, holding **two** repositories:
`/home/felhom-repo` (the fresh one, ~26 KB) and **`/home/felhom-repo.orphaned-20260805`
(12 535 KB — the Phase 0 history with all three sentinels)**.
- On **ep0**: namespace `c11`, token `felhom@pbs!c11`, and the two ACL lines on
`/datastore/felhom-offsite/c11`. **`demo-felhom` and `demo-hp` namespaces and tokens must not be
touched** — the Phase 0 capture recorded all three so teardown can tell them apart.
- The WireGuard peer for `c11-36d660` (`10.77.0.5/32`).
**Delete nothing that is not the campaign's own.** The brief's §8 warns that scratch customers have
accumulated before; the Phase 0 census found **none** outstanding, and c11 must not become the next.
---
## 12. Hygiene
- **The recovery codes are SHREDDED**, with the plant→find→shred→fail-to-find control the brief asks
for. **The control paid for itself on its first run**: it found the Phase 0 code in
`~/.config/credentials` as **`R_CAMPAIGN_11`** — a copy this session did not create and would never
have looked for, which would have made a "codes shredded" claim **false**. That key was removed from
the shared file carefully (backup → exact-match removal → `diff` proving every other line identical →
`HUB_PW` re-verified at `hub:200` → backup shredded). Everything else was `shred -u`'d on both hosts
and the absence re-swept; only the planted controls remained, and they were shredded too.
**⚠ Consequence, stated plainly: `/home/felhom-repo.orphaned-20260805` (12 535 KB, the three Phase 0
sentinels) is now permanently unopenable.** That is what the set-aside screen promises will happen,
teardown removes the repository anyway, and R-222 means no read path existed for it regardless — but
the door is now shut for good.
- **No secret is written into any committed file.** Every hash quoted here is a sha256 prefix; the
codes' contents appear nowhere.
- **`git add -A` was never used** — every commit staged explicit paths, and `git status --porcelain`
was checked before each one to confirm no foreign file was swept (a parallel session shares this
clone).
- **The pre-push hook is armed on this clone** (`core.hooksPath=.githooks`) and ran
`repo_gates.py --fast` green before every push. **No `--no-verify` was used.**
- **No product code was changed.** Every finding was filed and the run continued, per the brief's
rule 1.
- **No version was bumped** in any repo.
---
## 13. What did not run, and why
| | why |
|---|---|
| **F10 as specified** | **Not injectable.** Three attempts, each with a control: the app, then the controller's monitor, then the run itself recreate the mandatory directory within ~1 s. The state does not exist on a deployed app of this kind. **Harness, not product** |
| **F9's literal precondition** | A box that never had off-site backups needs a **rebuild**, which the brief forbids before Phase 4. The assertion that failed in Phase 1 (the page not consulting its predicate) was tested instead, and passes |
| **§4.2's positive half** | Needs shape (a) — hub blob present, key placed, **no target**. The venue has a target, so the box correctly does not declare. Also needs a rebuild |
| **`STALE → DOWN` escalation** | F11 was ended once `stale` and `recovered` had both fired, because Phase 4 needed the venue back. The >1 h arm is untested |
| **The whole-guest tier's due-ness** | **Not resolvable with existing instruments** — routine local-api calls and the hub's deadline monitor are both invisible at INFO. Recorded as an open question, not scored |
| **The retained package's read path** | Does not exist (R-199's inventory is unbuilt). R-222 was re-confirmed live rather than re-tested |
| **A re-walk of the journey** | **Deliberately out of scope.** These faults are not a re-walk, and the capability map's recovery row stays **FAIL** until one passes |
| **Any fix** | Brief rule 1 — file and continue. **No product code was changed and no version bumped** |
---
## 14. Venue state at the end — **WORKING**
| | |
|---|---|
| Hub | `c11-36d660` **ONLINE**, agent `0.125.0`, 1/1 guests, reporting on schedule (last seen 04:29:54) |
| Containers | `calibre-web` · `filebrowser` · `felhom-controller:0.201.0` · `traefik` — all healthy |
| Drives | both enrolled; backup target `{"degraded":false,"label":"mentes","target":"felhom-backup"}` |
| Off-site | **2 snapshots**, `last_status: ok`, `last_success 2026-08-06T02:15:24Z`, 51 206 B, `escrow_state: escrowed` |
| Set aside | `/home/felhom-repo.orphaned-20260805` — 12 535 KB, intact, and now **permanently unopenable** (§12) |
| Left in place | the raw `/mnt/adatok` and `/mnt/mentes` mounts remain **unmounted** — R-220's workaround, without which no app can be deployed on a rebuilt box. The stable `/mnt/felhom-drives/*` mounts are what everything uses |
| Access | the appliance's vaulted root credential was **shredded with the codes**, so a future session must re-fetch it from the hub (`POST /hosts/c11-36d660/reveal-recovery-credential`) — the designed path |
**Nothing is left broken, and nothing is left running that should not be.**