13 KiB
DRILL — R-201, the night run: THE KEY CAME BACK. The verdict was not reached.
Date: 2026-08-04, 21:30–22:40 · Box: demo-hp (Tier 0) · Unattended, by operator decision.
The wipe happened. The box is up, its apps are serving, and it is left mid-drill by deliberate
choice — see §7 for the exact state and the one command that resumes it.
THE RESULT, first. A real rebuild was performed: the controller's data volume was destroyed and the sentinel deleted from disk. The customer's recovery code then produced
8a9e33aa4da6769c5aea1831f87759e10930e2ec1dea0062576484e0598d080a— byte-identical to the key the box used before the wipe, and to the hash the hub had independently recorded.The off-site backup key is recoverable after a machine is rebuilt. That has never been shown before.
And the drill did not finish, because three separate things stand between a recovered key and a restored file. All three are measured below. That is the other half of the night's result, and it is the half nobody knew.
1. The verdict
NOT REACHED. Step 10 (restore the sentinel and compare) was not run.
| sentinel sha256, pre-wipe | 643166269103a25cf41d34a26b75fd6ebba0a837bbcd7e0d2b5aba7106bcbe7c |
| sentinel sha256, post-restore | not obtained — the restore was never reachable |
| snapshot count at step 9 | not obtained — the run was refused before it started |
This is not a FAIL. A fail is "the file came back wrong" or "a fresh empty history was started". Neither happened, because the box never got as far as running a backup. What the night established is where the wall is.
2. What was proven, in order, on hardware
Step 6 — the wipe
The controller's data volume was destroyed and the bootstrap re-ran. Fresh data directory, everything
stamped 2026-08-04 20:00:2x:
encryption.key 32 B 2026-08-04 20:00:21 ← brand new: every pre-wipe app secret is undecryptable
claimed = None ← the fresh-install signal
offbox = null ← no off-site target
Faithful to the incident, and deliberately so. The 2026-08-03 rebuild that R-193 is filed against
was not a guest reprovision — the journal shows guest 9201 running continuously through that
window, with no pct destroy, no pct restore and no --selftest=provision. What changed was the
controller and its data volume. Reproducing that is what the drill needs; a guest destroy plus an
unrehearsed provisioning chain, improvised unattended, is what §8.10 exists to prevent.
Step 7 — the assertion that keeps recovery possible: PASSED
host_escrow (demo-hp-bb76ea), AFTER the wipe:
identity_blob = 572 bytes ← unchanged
restic_pw_sha256 = 8a9e33aa4da6… ← unchanged
updated_at = 2026-08-04 11:11:37 ← unchanged; nothing re-escrowed
stale_at = NULL
No ceremony was run and nothing re-escrowed itself. The sealed key survived the rebuild untouched.
Step 8 — THE KEY CAME BACK
=== offsite key recovery check (R-200) — compares, never installs ===
recovered sha256: 8a9e33aa4da6769c5aea1831f87759e10930e2ec1dea0062576484e0598d080a
[FAIL] there is no repository password on this box to compare against
(the recovery itself SUCCEEDED — this box simply has no local key. That is the
rebuilt-box shape, where the next step is to INSTALL rather than compare.)
Exit 1 is correct and designed: there was nothing to compare against, because the wipe removed it. The recovered hash matches the pre-wipe on-disk key and the hub's own record — three independent sources agreeing, one of them recovered through the full chain (hub → agent → unseal → extract) on a box that had just lost everything it knew.
Step 9a — installed cleanly
=== offsite key recovery INSTALL (R-200) ===
on-disk sha256: (none — this box has no repository password)
recovered sha256: 8a9e33aa4da6…
[INSTALLED] the recovered repository password is in place and reads back identical.
The "installed" branch — the rebuilt-box shape v0.196.0 was written for — took its first real run.
Step 9b — the apply kept it
After the target was reconfigured, the on-disk key was still 8a9e33aa4da6…. WriteOffboxSecrets
found the file present and kept it, exactly as documented.
3. The wall — three blockers, each measured
(a) R-193, reconfirmed live: a rebuilt controller cannot configure its off-site tier
[INFO] [offsite-apply] settle-gate: GO — at/above floor 0.156.0 (we are 0.197.0)
[WARN] [offsite-apply] reconcile: offsite-apply: consume one-time password:
no unconsumed offsite password (already consumed or none provisioned)
(retries on next config refresh/restart)
The ledger, measured: demo-hp one-time secret created 07:11:51, consumed 07:12:06 — by the
previous controller. The rebuilt one has nothing to consume and retries forever. Exactly the state
demo-hp sat in for 25 hours on 2026-08-03.
Remedy: an operator Re-issue. Performed here through the designed endpoint
(POST /configs/demo-hp/offsite-reissue, HTTP 303), after which the target configured normally.
(b) R-196, measured live — and it lands squarely on the recovery path
host_escrow (demo-hp-bb76ea), after the Re-issue:
stale_at = 2026-08-04 20:15:49 ← set
restic_pw_sha256 = 8a9e33aa4da6… ← UNCHANGED. The key did not move.
The escrow was marked stale while it perfectly covers the box's current key — the recovered one.
That is R-196's false staleness, and here it is not a cosmetic lie: a stale escrow makes the hub
withhold the hash from the report ACK, so EscrowAutoConfirmer can never flip pending → escrowed, and OffboxRunnable (configured && escrowed) refuses to run any off-site backup.
The remedy for (a) disables the recovery it was needed for. The only documented way to clear a stale escrow is a fresh ceremony — which supersedes the identity blob, i.e. destroys the very key being recovered. Under hub v0.93.0 the old blob is now retained, but nothing can serve a superseded blob back (R-199's inventory: the endpoint serves the CURRENT row only).
Superseded rows still number 2 — no ceremony was run tonight. The key is intact.
(c) The claim gate — undocumented as a recovery step
[DEBUG] [web] claim gate: intercepting POST /backup/offbox/confirm-escrow (unclaimed)
A rebuilt box is unclaimed (claimed = None, fresh settings.json), and the claim gate correctly
intercepts every non-claim route. So no controller endpoint can be driven at all — not the manual
escrow confirm, not the backup trigger — until the customer re-claims the box. Both POSTs returned 302
to the claim page and neither reached its handler; last_run stayed null, which is why the earlier
reading of "the run was refused" needed this second look to be accurate.
This is correct behaviour and it is not wrong — but it is a step in the customer's recovery journey that appears in no design document, and it comes before anything else can happen.
4. Why the session stopped here
By 22:35 the path forward was a chain of workarounds assembled live — Re-issue, then the deprecated
manual confirm, then the root-gated --print-reset-code escape hatch to re-claim, then the confirm
again, then the run. Each is individually defensible; the accumulation is precisely the pattern §8.10
names:
"If anything is ambiguous at 03:00, stop and leave it for the morning. An unattended session that halts with a clear state beats one that improvises."
The remaining steps need about five minutes with a person present. They are not worth improvising alone at the end of a long night, on the one box whose off-site history the drill is trying to prove.
Nothing was left broken. The box is up, all six app containers are serving, and the recovered key is on disk.
5. §5's five conditions, as recorded before the wipe
| # | Condition | Evidence |
|---|---|---|
| 1 | §4 confirmed, sentinel listed by name | snapshot e6132ae5 (19:36:26) — -rw-r--r-- 1000 1000 181 … /DRILL-SENTINEL.txt |
| 2 | deliberate rollback archive, verified | vzdump-lxc-9201-2026_08_04-21_44_24.tar.zst, 1 606 765 083 B; full zstd stream read OK (4 867 573 760 B uncompressed); the sentinel confirmed inside it |
| 3 | §3's option | see §6 |
| 4 | demo-felhom untouched | health=ok, escrow_state=escrowed, not touched at any point |
| 5 | free space | nvme-1tb 883 GB free, root 20 GB, local-lvm 34.32 % |
A precondition had drifted and was repaired before the wipe, not worked around. The staged
snapshot no longer contained the sentinel: the afternoon's [WARN] mandatory data path missing
experiment produced a later same-day calibre-web snapshot, and restic's forget --keep-daily 7 --group-by host,tags had pruned the good one in its favour. The fixture on disk was correct, so one
backup re-established it and it was re-verified by listing. Lesson worth carrying: a good snapshot is
not durable against a later bad run on the same day.
6. §3 — the recovery code
Option B as already in place, with a strict improvement: no new copy was created, so nothing needed shredding.
The operator placed R_DEMO-HP in their own ~/.config/credentials on DooPlex (mode 0600) two
sessions ago, deliberately, for this purpose. This session read it from there and piped it to stdin
for each of the two invocations that needed it. It was never an argument, never exported, never written
to a second file, and never logged.
Nothing was shredded, and that is the point: the operator's own permanent store is theirs, not a drill artefact, and destroying it would have destroyed their record. Because no additional copy was made, there is nothing left behind to prove gone — a stronger position than option B's create-then-shred.
Verified afterwards with the planted-copy positive control, exactly as on 2026-08-04 afternoon — see §8.
7. The exact state the box is in, and how to resume
controller felhom-controller:0.197.0, healthy
apps privatebin opengist calibre-web filebrowser cloudflared traefik — all serving
claimed None ← must be re-claimed before any controller endpoint responds
escrow_state pending ← R-196: the Re-issue marked the escrow stale
repo key 8a9e33aa4da6… ← THE RECOVERED KEY, on disk
sentinel absent from disk (deliberately) — present in off-site snapshot e6132ae5
To resume (operator present, ~5 minutes):
- Re-claim the box —
docker exec felhom-controller /usr/local/bin/felhom-controller --print-reset-code, then the claim page. - Confirm the escrow (
/backup/offbox/confirm-escrow) soOffboxRunnableallows a run. Do NOT run a new ceremony — it would supersede the identity blob and destroy the key under test. - Run an off-site backup. The observable is whether the repository OPENS — and whether the
pre-wipe snapshot
e6132ae5still exists with the sentinel in it. A same-dayforgetkeeps one snapshot per tag, so the count is a poor discriminator; the surviving history is the real one. - Restore the sentinel through the customer restore flow; compare to
643166269103a25c….
Rollback, if preferred: pct restore 9201 from
/mnt/nvme-1tb/dump/vzdump-lxc-9201-2026_08_04-21_44_24.tar.zst — verified by a full read before the
wipe. It returns the box to its pre-wipe state and voids the remaining drill.
8. R persisted nowhere — searched, with a positive control
Swept the agent journal, the controller container log, and /tmp, /var/tmp, /var/lib/felhom-agent,
/root on demo-hp for the recovery code: 0 hits, and 0 leftover felhom-idesc-* unseal
staging directories. A planted copy was found by the same sweep (1) and not found after shredding
(0), so the instrument is shown sensitive rather than assumed to be.
9. Teardown — three layers
| layer | state |
|---|---|
| the guest | Nothing torn down. The wipe is the evidence; the apps are serving; the recovered key is in place. The scratch band is empty — no restore-test guest was created tonight. |
| the host | vzdump snapshot LVs removed cleanly by the archive job (snap_vm-9201-disk-0/1_vzdump both released). One new archive, 1.6 GB, on nvme-1tb (883 GB free). Nothing deleted. |
| the hub | No new customer records. demo-hp is the same customer row throughout — the rebuild was a controller-data wipe, not a re-enrolment, so nothing accumulated. One Re-issue staged a fresh one-time secret (consumed at 20:16) and set stale_at. The two superseded escrow rows are unchanged — no ceremony ran. |
Nothing was deleted on the storage endpoint, including the ~1.2 GB of previously-orphaned ciphertext. Ruled, still owed, and deliberately not ridden along with a drill.
10. Part 2 — not run, and why
Its gate is "the drill PASSED". It did not — the verdict was not reached. Running a second wipe on a box whose first result is incomplete would have destroyed the staged state that makes the first one finishable in five minutes. R-198's retention therefore remains unit-proven only, unchanged from this morning.