hub v0.103.0 — a host can read the packages we kept for it (R-311)
gates / gates (push) Successful in 37s

ListSupersededEscrow had zero production callers for nineteen days. It is the only
reader of a retained identity_blob, so the retention shipped in v0.93.0 was
material the product could not reach - proven on the fixture 2026-08-12, where a
code that opens a retained package was answered as a code that opened nothing.

New GET /api/v1/hosts/<id>/escrow/retained: self-scoped exactly as the current-row
GET, same recovery-mode gate, same audit event written BEFORE the bytes leave,
capped at 16. Rows with a NULL identity_blob are WITHHELD and returned as
unopenable_count - they retain the PBS key, not the repository password, so they
can never open what the caller is asking about, and serving them would let the
screen claim an earlier package is openable on exactly the boxes the original
defect hurt. The count is returned because their existence is load-bearing and
underivable by the caller.

The trade, stated rather than waved through: the hub still cannot read any of it -
sealed bytes in, sealed bytes out, no decrypt path, no recovery code ever held.
What widens is volume, bounded by self-scope, the recovery-mode gate and the cap.

The response is a NAMED TYPE, not a map, so the wire-contract gate can resolve it;
the wire is declared as a fourth ROOT and the gate now checks 182 tags rather than
174. A positive control shows that check is name-presence, not decodability -
filed as R-315 rather than reported as coverage.

Also: golden 0.214.0 baked, published and round-trip verified; the countdown on
demo-felhom cancelled on the operator's ruling (R-307); the spike that halted
Part 3 recorded as R-312; the set-aside store found unrecoverable as R-313.
Six hub tests through the real endpoint; four red-proofs asserted applied.
This commit is contained in:
2026-08-12 18:56:44 +02:00
parent 6362bb6cb6
commit 8b188bea68
5 changed files with 511 additions and 69 deletions
+116 -68
View File
@@ -1,87 +1,135 @@
# REPORT — DRILL: the retained key, and the two fixes nobody had watched work (2026-08-12)
# REPORT — The door, part one: a correct code stops being called wrong (2026-08-12 night)
**Class:** drill (unattended, destructive on Tier 0) + spike for Phase C's first step
**Venues:** `drill-r50` (nested PVE on DooPlex), `demo-felhom` (guest 9201) — both Tier 0.
**`demo-hp` was never touched. `peti-felhom` was never contacted. No abandon countdown was started,
shortened or triggered.**
**Full record:** `documentation/audits/DRILL-retained-key-2026-08-12.md`
**Three repos.** hub **v0.103.0** · agent **v0.129.0** · controller **v0.214.0** · golden **0.214.0**.
Register: **R-311 CLOSED**, R-307 CLOSED, **R-312 / R-313 / R-314 / R-315 opened**. Ceiling
R-310 → **R-315**.
---
## The answer to the question this drill existed to answer
## 1. The spike's answer, first and in plain language
**(a) Is the old key kept? YES** — proven for the first time in the fleet's history.
**(b) Does the kept key open the old backups? YES** — three planted files, including a Hungarian
accented filename verified as raw bytes, restored **byte-identical** from a store the machine itself
could no longer open.
**(c) Can the customer get there through the product? NO — and they are told their correct code is
wrong.**
**Can a customer restore from a set-aside store with the machinery that already exists? NO — and
building it is new surface, not wiring.** Established read-only, at `file:line`, before a line was
written.
The brief said to be ready for the answer to be no, and our own records predicted retention would be
*"a box we fill and cannot open"*. **That was half right, and the wrong half was the one nobody had
checked.** The box opens. What does not exist is the door: `ListSupersededEscrow`
(`hub/internal/store/store.go:2841`) is the only reader of a retained key and has **zero production
callers**; the recovery path selects `FROM host_escrow` — the current row only. Asked with the very
code that had just opened the retained row by hand, the product answered *"the recovery code did not
open the sealed bundle — nothing was written"*. → **R-304, rank 1**
Every restore entry point resolves the repository from `settings.GetOffboxTarget()` and the password
from the single `offboxPwPath()` file: `offboxLatestSnapshot` (`offbox_restore.go:85-86`),
`offboxSnapshotSize` (`:139-140`), `RestoreOffboxScratch` (`:206+`). **A grep for a repo-path
parameter anywhere in the restore chain returns nothing.** The only seam that installs a recovered
password, `InjectOffboxPassword` (`offbox.go:667`), writes that same one file — i.e. **adoption**.
Consequences: the census answer **stands**; the countdown banner's promise is **true in substance,
false in practice**; the capability map's recovery claim **has been moved** with today's evidence.
What the drill did to read the set-aside store was `restic` **by hand**, with `-r <alt repo>` and an
overridden `RESTIC_PASSWORD_FILE`. **That distance is exactly what (b)-to-(c) costs.**
## What shipped
**So the session halted at Part 3 by its own rule, shipped Part 2, and hands back options → R-312.**
Cost to find out: ~35 minutes, read-only.
**`installer-v1.27.0` published** — tag cut and **both** `--ref`s in `manifests/webpage.yaml` bumped
(sidecar line 327, init container line 372). Publication was earned: both faults were watched
happening first, from a machine reset to factory state.
## 2. Part 0 — the countdown, cancelled on your ruling
- **R-300 CLOSED** — pre-fix uninstall left dnsmasq `enabled`/`active` on `0.0.0.0:53`; the next byo
install refused, exit 1. Fixed path: recorded `not present before Felhom``stopping + disabling
it``:53 FREE` → preflight PASS. The owner's side proven too (record `yes` → left running).
- **R-297 CLOSED** — a stale `golden-0.98.3.tar.zst` planted as newest-by-filename; v1.25.0 took it
with no comparison and **the box came up on controller 0.98.3** against a vouched 0.213.0 — below
the floor and below v0.206.0 where the off-site recovery screen exists. Fixed path re-fetched and
sha-verified the vouched golden (landed 0.213.0); an operator-named stale archive was **refused**.
Through the product's own operator path (`--abandon-stop`, which refuses rather than silently
no-opping), with the container **stopped first** so the running controller could not overwrite
`settings.json` from memory. **Proved, not trusted to the exit code:**
## Findings opened — ceiling R-303 → R-310
- `abandon_started_at` and `abandon_at`**gone**. `AbandonStatus` returns `Active=false` when
`AbandonAt` is empty (`offbox_abandon.go:111-113`), so no countdown renders.
- `abandon_repo_path`**deliberately kept**, as the pointer to the preserved store.
- The set-aside store — **still there**: 36 snapshot objects, full `config/data/index/keys/locks/
snapshots` structure. Both repositories still on the endpoint. **Nothing deleted anywhere.**
| # | Rank | What |
**And the thing you should know about what was preserved (R-313):** it holds **36 snapshots and one
key slot**, and it does **not** open with the box's current password (`Fatal: wrong password or no key
found`, exit 1 — measured). Its key is the one hashed `48741892f0ef…` — retained row id 4,
`identity_blob` **NULL**, a pre-v0.93.0 row. **The material was dropped by the R-198 defect during its
two-month window, so no recovery code in existence opens that store.** Keeping it is still the right
call — deleting is irreversible and a decision, not an accumulation — but it is 36 unreadable
snapshots, and that is the concrete, still-present cost of R-198 sitting on the endpoint.
## 3. Part 2 — what shipped, and a correction to the premise
**The premise needed correcting first.** The task described the customer being told *"the recovery code
did not open the sealed bundle"*. That is the **agent's local-API** reply. The **customer-facing
screen already hedged** (R-222/R-226) — it named both causes, named the kept package and its date, and
said it could not tell them apart. That was **honest**; it could not tell them apart **because nothing
ever looked**. So what shipped is smaller and more precise than "stop the lie": **the hedge becomes an
answer.**
- **hub v0.103.0** — `GET /hosts/<id>/escrow/retained`, the **first production caller
`ListSupersededEscrow` has ever had**. Self-scoped identically, same recovery-mode gate, same audit
event written before the bytes leave, capped at 16. Rows with a NULL `identity_blob` are **withheld
and counted** (`unopenable_count`): they can never open what the caller is asking about, and serving
them would let the screen promise recovery on exactly the boxes the original defect hurt.
- **agent v0.129.0** — retained packages tried **only after** the current one refuses; `422` with
`superseded_at`; bounded at 6 attempts (~1 s of scrypt each); fail-safe in every direction.
- **controller v0.214.0** — class `RecoveryCodeOpensRetained`, gated on MinAgent **0.129.0** via a
**second, separate** trust flag (a box can sit between 0.126.0 and 0.129.0). The message says the
code is correct, names the date, says the package is kept, says the **current** backups are
unaffected, and **promises no restore** — it routes to support, which can do it.
**The trade you should see stated:** the hub still cannot read any of it — sealed bytes in, sealed
bytes out, no decrypt path, no recovery code ever held. What widens is **volume**: a host key that
could fetch one opaque package can now fetch N, bounded by self-scope, the recovery-mode gate and the
cap.
## 4. Red-proofs — and where the lie actually lives
Every mutation asserted to have applied before its run.
| Repo | Mutation | Outcome |
|---|---|---|
| **R-304** | **1** | Retained key has no product route; the correct old code is reported as wrong |
| **R-305** | 2 | The R-300 cleanup fires **once per machine** — the leftover returns on the second reinstall (proven, cycles 2/3) |
| **R-308** | 2 | Stored controller `PASSWORD` no longer opens demo-felhom (`Hibás jelszó`) — not the quoting trap |
| **R-306** | 3 | `--preflight-only` says *"no state written"* and writes `state.json` — with an ownership answer that can be wrong |
| **R-309** | 3 | The day-0 runbook says pushing publishes the installer; false since R-110 (measured: public URL served 1.25.0 while `main` had 1.27.0) |
| **R-310** | 4 | Duplicated sentence in the golden refusal; `--uninstall` needs a pty and `--force` does not bypass it |
| **R-307** | — | **Operator decision, deadline 2026-08-24** — see below |
| hub | serve the CURRENT row instead of retained | FAILS (count 2→1) |
| hub | drop the unopenable guard | FAILS (count 1→2, unopenable 1→0) |
| hub | drop self-scope | FAILS (403→200) |
| hub | collapse the route suffix | FAILS (count 1→0) |
| **agent** | **remove the retained lookup** | **FAILS — the fail-closed wrong-code error returns. THE LIE COMES BACK.** |
| agent | + 6 more (nil fetcher, wrong code, fetch failure, bounded attempts, predates-field, success path) | all pinned |
| controller | delete the new case | FAILS — but the customer gets the **neutral** message, because R-224's safe default catches it |
| controller | make 422 unconditional | FAILS — an agent that never looked is read as having looked |
| controller | route 400 to the new class | FAILS — a mistype is congratulated |
## What needs you
**Answering the question directly:** the lie returns when the **agent's** retained lookup is removed,
not when the controller's case is. R-224's safe default is doing its job one layer up.
**`demo-felhom` carries a live abandon countdown** — started 2026-08-10, **firing 2026-08-24**, for
`/home/felhom-repo.orphaned-20260810`. This drill did **not** start it and deliberately did **not**
cancel it. The brief's end state asked for no countdown anywhere; satisfying that means choosing:
**cancel it** (copy kept indefinitely, storage cost, no data risk) or **let it run** (copy deleted,
irreversibly). **Doing nothing selects deletion.****R-307**
## 5. The claim guard, and a gate whose positive control failed
## End state
**The claim guard had a blind spot the size of the recovery screen** — it scanned templates only,
while every recovery message is a Go string in a handler. It now scans `recovery_handlers.go` too, and
**on its first run convicted a pre-existing unregistered claim**. 8 → 10 registered claims.
- **`demo-felhom`** — up, reporting, healthy, on the vouched pair; `repo_password` restored to the
original (`sha c60c8bc737a6b7c6…`), escrow re-sealed and uploaded, off-site repo reachable
(`restic snapshots` exit 0). Its recovery code was rotated by the final ceremony and
`R_DEMO-FELHOM` updated in place (prior file backed up alongside). Planted data removed; eight
secret-bearing files **shredded**.
- **`demo-hp`** — untouched, reporting.
- **`drill-r50`** — **reverted to snapshot `virgin`, powered off.**
- **Hub** — two new retained rows (the P1 and P2 blobs), deliberately kept as the fixture proving the
retention works. `drill-r50-0a4f9a` re-used, not duplicated: no new scratch customer.
- **Off-site** — only demo-felhom's own repository path touched, `backup` the only mutating verb used.
**No prune, no forget, no delete, no rename anywhere.** One snapshot added and deliberately left:
`6ea85413`, 66 KiB, tagged `drill-retained-key-20260812` — removable by ID if you want it gone.
**The wire-contract gate: declared, and honestly weaker than it looks (R-315).** The hub response was
made a **named type** so the gate could resolve it; the wire is declared as a fourth ROOT and the tag
count rose **174 → 182**, so the fields are inspected. But a positive control — renaming the
agent-side `superseded_at` tag — **still passed**, because the check is repo-wide name-presence and
the string also occurs as a map key elsewhere. The gate documents this ("name-reachability is not
use"), so it is a known limit, not a regression — **but declaring this wire bought documentation, not
enforcement**, and saying otherwise would have been false.
## Honest gaps
## 6. Live state
- **The Phase A logs did not survive** the intermediate revert to `virgin`. Every quotation in the
audit is verbatim from the live run, but the raw files are gone. Procedural lesson, recorded.
- The planted data reached the store via `restic` directly, not the dashboard button, because of
R-308 — so the app-backup→unit→offsite chain went unexercised. Not what this drill measured.
- Wall clock **≈ 1 h 03 min** against a 45 h envelope. Nothing was dropped; Phase C ran concurrently
with Phase B on a different machine.
| | |
|---|---|
| controller | **0.214.0** on guest 9201, `Up … (healthy)` |
| agent | **0.129.0** on `felhom-pve`, unit active, journal clean |
| hub | **0.103.0** — see §7 |
| golden | **0.214.0** baked + published |
**Live validation was endpoint/handler-level, not a click-through**, and the reason is a finding:
**the dashboard password of record no longer opens `demo-felhom` (R-308)**.
## 7. What was dropped, named plainly
- **Part 3 (the route) — HALTED at the spike, by the task's own rule.** → R-312.
- **§7's fixture walk was not re-run end-to-end.** Yesterday's drill already proved the byte-identical
restore from a set-aside store; today's change is upstream of it (which sentence is shown), the
dashboard is unreachable headlessly (R-308), and the restore route does not exist (R-312). What was
proved live is the 422 itself.
- Explicitly out of scope and still open, so it does not read as forgotten: **R-305** (the removal fix
helps a machine once — the tester's second reinstall still hits it), the hub emails naming the
retired secret, **R-309** (the runbook's publication claim), the CI runs that fail with no log, the
twenty unread facts, the nine grey claims, **R-303**.
## 8. Bypass, stated as required
`git push --no-verify` was used **once**, on `felhom-agent`. The `release-complete` gate refuses a
CHANGELOG entry whose tag and package do not exist; `release-agent.sh` refuses a tree that is not
pushed. Circular by construction. The bypass was immediately followed by the real release
(`release-agent.sh 0.129.0`), and the gates were re-run afterwards: **green**.