R-201 night run: the off-site key IS recoverable after a real rebuild (proven); the verdict is blocked by R-204
gates / gates (push) Successful in 6s
gates / gates (push) Successful in 6s
This commit is contained in:
File diff suppressed because one or more lines are too long
@@ -0,0 +1,264 @@
|
||||
# DRILL — R-201, the night run: **THE KEY CAME BACK. The verdict was not reached.**
|
||||
|
||||
**Date:** 2026-08-04, 21:30–22:40 · **Box:** `demo-hp` (Tier 0) · **Unattended, by operator decision.**
|
||||
**The wipe happened.** The box is up, its apps are serving, and it is left mid-drill by deliberate
|
||||
choice — see §7 for the exact state and the one command that resumes it.
|
||||
|
||||
> **THE RESULT, first.** A real rebuild was performed: the controller's data volume was destroyed and
|
||||
> the sentinel deleted from disk. The customer's recovery code then produced
|
||||
> `8a9e33aa4da6769c5aea1831f87759e10930e2ec1dea0062576484e0598d080a` — **byte-identical to the key the
|
||||
> box used before the wipe, and to the hash the hub had independently recorded.**
|
||||
>
|
||||
> **The off-site backup key is recoverable after a machine is rebuilt. That has never been shown
|
||||
> before.**
|
||||
>
|
||||
> **And the drill did not finish**, because three separate things stand between a recovered key and a
|
||||
> restored file. All three are measured below. That is the other half of the night's result, and it is
|
||||
> the half nobody knew.
|
||||
|
||||
---
|
||||
|
||||
## 1. The verdict
|
||||
|
||||
**NOT REACHED.** Step 10 (restore the sentinel and compare) was not run.
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| sentinel sha256, pre-wipe | `643166269103a25cf41d34a26b75fd6ebba0a837bbcd7e0d2b5aba7106bcbe7c` |
|
||||
| sentinel sha256, post-restore | **not obtained** — the restore was never reachable |
|
||||
| snapshot count at step 9 | **not obtained** — the run was refused before it started |
|
||||
|
||||
**This is not a FAIL.** A fail is "the file came back wrong" or "a fresh empty history was started".
|
||||
Neither happened, because the box never got as far as running a backup. What the night established is
|
||||
where the wall is.
|
||||
|
||||
---
|
||||
|
||||
## 2. What was proven, in order, on hardware
|
||||
|
||||
### Step 6 — the wipe
|
||||
|
||||
The controller's data volume was destroyed and the bootstrap re-ran. Fresh data directory, everything
|
||||
stamped `2026-08-04 20:00:2x`:
|
||||
|
||||
```
|
||||
encryption.key 32 B 2026-08-04 20:00:21 ← brand new: every pre-wipe app secret is undecryptable
|
||||
claimed = None ← the fresh-install signal
|
||||
offbox = null ← no off-site target
|
||||
```
|
||||
|
||||
**Faithful to the incident, and deliberately so.** The 2026-08-03 rebuild that R-193 is filed against
|
||||
was **not** a guest reprovision — the journal shows guest 9201 running continuously through that
|
||||
window, with no `pct destroy`, no `pct restore` and no `--selftest=provision`. What changed was the
|
||||
controller and its data volume. Reproducing *that* is what the drill needs; a guest destroy plus an
|
||||
unrehearsed provisioning chain, improvised unattended, is what §8.10 exists to prevent.
|
||||
|
||||
### Step 7 — the assertion that keeps recovery possible: **PASSED**
|
||||
|
||||
```
|
||||
host_escrow (demo-hp-bb76ea), AFTER the wipe:
|
||||
identity_blob = 572 bytes ← unchanged
|
||||
restic_pw_sha256 = 8a9e33aa4da6… ← unchanged
|
||||
updated_at = 2026-08-04 11:11:37 ← unchanged; nothing re-escrowed
|
||||
stale_at = NULL
|
||||
```
|
||||
|
||||
**No ceremony was run and nothing re-escrowed itself.** The sealed key survived the rebuild untouched.
|
||||
|
||||
### Step 8 — **THE KEY CAME BACK**
|
||||
|
||||
```
|
||||
=== offsite key recovery check (R-200) — compares, never installs ===
|
||||
recovered sha256: 8a9e33aa4da6769c5aea1831f87759e10930e2ec1dea0062576484e0598d080a
|
||||
[FAIL] there is no repository password on this box to compare against
|
||||
(the recovery itself SUCCEEDED — this box simply has no local key. That is the
|
||||
rebuilt-box shape, where the next step is to INSTALL rather than compare.)
|
||||
```
|
||||
|
||||
Exit 1 is **correct and designed**: there was nothing to compare against, because the wipe removed it.
|
||||
The recovered hash matches the pre-wipe on-disk key and the hub's own record — three independent
|
||||
sources agreeing, one of them recovered through the full chain (hub → agent → unseal → extract) on a
|
||||
box that had just lost everything it knew.
|
||||
|
||||
### Step 9a — installed cleanly
|
||||
|
||||
```
|
||||
=== offsite key recovery INSTALL (R-200) ===
|
||||
on-disk sha256: (none — this box has no repository password)
|
||||
recovered sha256: 8a9e33aa4da6…
|
||||
[INSTALLED] the recovered repository password is in place and reads back identical.
|
||||
```
|
||||
|
||||
The "installed" branch — the rebuilt-box shape v0.196.0 was written for — took its first real run.
|
||||
|
||||
### Step 9b — the apply kept it
|
||||
|
||||
After the target was reconfigured, the on-disk key was still `8a9e33aa4da6…`. `WriteOffboxSecrets`
|
||||
found the file present and kept it, exactly as documented.
|
||||
|
||||
---
|
||||
|
||||
## 3. The wall — three blockers, each measured
|
||||
|
||||
### (a) R-193, reconfirmed live: a rebuilt controller cannot configure its off-site tier
|
||||
|
||||
```
|
||||
[INFO] [offsite-apply] settle-gate: GO — at/above floor 0.156.0 (we are 0.197.0)
|
||||
[WARN] [offsite-apply] reconcile: offsite-apply: consume one-time password:
|
||||
no unconsumed offsite password (already consumed or none provisioned)
|
||||
(retries on next config refresh/restart)
|
||||
```
|
||||
|
||||
The ledger, measured: `demo-hp` one-time secret created `07:11:51`, consumed `07:12:06` — **by the
|
||||
previous controller**. The rebuilt one has nothing to consume and retries forever. Exactly the state
|
||||
demo-hp sat in for 25 hours on 2026-08-03.
|
||||
|
||||
**Remedy:** an operator Re-issue. Performed here through the designed endpoint
|
||||
(`POST /configs/demo-hp/offsite-reissue`, HTTP 303), after which the target configured normally.
|
||||
|
||||
### (b) R-196, measured live — and it lands squarely on the recovery path
|
||||
|
||||
```
|
||||
host_escrow (demo-hp-bb76ea), after the Re-issue:
|
||||
stale_at = 2026-08-04 20:15:49 ← set
|
||||
restic_pw_sha256 = 8a9e33aa4da6… ← UNCHANGED. The key did not move.
|
||||
```
|
||||
|
||||
**The escrow was marked stale while it perfectly covers the box's current key** — the recovered one.
|
||||
That is R-196's false staleness, and here it is not a cosmetic lie: a stale escrow makes the hub
|
||||
**withhold the hash from the report ACK**, so `EscrowAutoConfirmer` can never flip `pending →
|
||||
escrowed`, and `OffboxRunnable` (`configured && escrowed`) **refuses to run any off-site backup**.
|
||||
|
||||
> **The remedy for (a) disables the recovery it was needed for.** The only documented way to clear a
|
||||
> stale escrow is a fresh ceremony — which supersedes the identity blob, i.e. **destroys the very key
|
||||
> being recovered**. Under hub v0.93.0 the old blob is now *retained*, but nothing can serve a
|
||||
> superseded blob back (R-199's inventory: the endpoint serves the CURRENT row only).
|
||||
|
||||
**Superseded rows still number 2** — no ceremony was run tonight. The key is intact.
|
||||
|
||||
### (c) The claim gate — undocumented as a recovery step
|
||||
|
||||
```
|
||||
[DEBUG] [web] claim gate: intercepting POST /backup/offbox/confirm-escrow (unclaimed)
|
||||
```
|
||||
|
||||
A rebuilt box is **unclaimed** (`claimed = None`, fresh `settings.json`), and the claim gate correctly
|
||||
intercepts every non-claim route. **So no controller endpoint can be driven at all** — not the manual
|
||||
escrow confirm, not the backup trigger — until the customer re-claims the box. Both POSTs returned 302
|
||||
to the claim page and neither reached its handler; `last_run` stayed null, which is why the earlier
|
||||
reading of "the run was refused" needed this second look to be accurate.
|
||||
|
||||
This is correct behaviour and it is **not wrong** — but it is a step in the customer's recovery journey
|
||||
that appears in no design document, and it comes *before* anything else can happen.
|
||||
|
||||
---
|
||||
|
||||
## 4. Why the session stopped here
|
||||
|
||||
By 22:35 the path forward was a chain of workarounds assembled live — Re-issue, then the deprecated
|
||||
manual confirm, then the root-gated `--print-reset-code` escape hatch to re-claim, then the confirm
|
||||
again, then the run. Each is individually defensible; the accumulation is precisely the pattern §8.10
|
||||
names:
|
||||
|
||||
> *"If anything is ambiguous at 03:00, stop and leave it for the morning. An unattended session that
|
||||
> halts with a clear state beats one that improvises."*
|
||||
|
||||
The remaining steps need about five minutes **with a person present**. They are not worth improvising
|
||||
alone at the end of a long night, on the one box whose off-site history the drill is trying to prove.
|
||||
|
||||
**Nothing was left broken.** The box is up, all six app containers are serving, and the recovered key
|
||||
is on disk.
|
||||
|
||||
---
|
||||
|
||||
## 5. §5's five conditions, as recorded before the wipe
|
||||
|
||||
| # | Condition | Evidence |
|
||||
|---|---|---|
|
||||
| 1 | §4 confirmed, sentinel **listed by name** | snapshot `e6132ae5` (19:36:26) — `-rw-r--r-- 1000 1000 181 … /DRILL-SENTINEL.txt` |
|
||||
| 2 | deliberate rollback archive, **verified** | `vzdump-lxc-9201-2026_08_04-21_44_24.tar.zst`, 1 606 765 083 B; **full zstd stream read OK** (4 867 573 760 B uncompressed); the sentinel confirmed **inside** it |
|
||||
| 3 | §3's option | see §6 |
|
||||
| 4 | demo-felhom untouched | `health=ok`, `escrow_state=escrowed`, not touched at any point |
|
||||
| 5 | free space | nvme-1tb 883 GB free, root 20 GB, local-lvm 34.32 % |
|
||||
|
||||
**A precondition had drifted and was repaired before the wipe, not worked around.** The staged
|
||||
snapshot no longer contained the sentinel: the afternoon's `[WARN] mandatory data path missing`
|
||||
experiment produced a *later* same-day calibre-web snapshot, and restic's `forget --keep-daily 7
|
||||
--group-by host,tags` had pruned the good one in its favour. The fixture on disk was correct, so one
|
||||
backup re-established it and it was re-verified by listing. **Lesson worth carrying: a good snapshot is
|
||||
not durable against a later bad run on the same day.**
|
||||
|
||||
## 6. §3 — the recovery code
|
||||
|
||||
**Option B as already in place, with a strict improvement: no new copy was created, so nothing needed
|
||||
shredding.**
|
||||
|
||||
The operator placed `R_DEMO-HP` in their own `~/.config/credentials` on DooPlex (mode `0600`) two
|
||||
sessions ago, deliberately, for this purpose. This session read it from there and **piped it to stdin**
|
||||
for each of the two invocations that needed it. It was never an argument, never exported, never written
|
||||
to a second file, and never logged.
|
||||
|
||||
**Nothing was shredded, and that is the point:** the operator's own permanent store is theirs, not a
|
||||
drill artefact, and destroying it would have destroyed their record. Because no additional copy was
|
||||
made, there is nothing left behind to prove gone — a stronger position than option B's
|
||||
create-then-shred.
|
||||
|
||||
**Verified afterwards** with the planted-copy positive control, exactly as on 2026-08-04 afternoon —
|
||||
see §8.
|
||||
|
||||
---
|
||||
|
||||
## 7. The exact state the box is in, and how to resume
|
||||
|
||||
```
|
||||
controller felhom-controller:0.197.0, healthy
|
||||
apps privatebin opengist calibre-web filebrowser cloudflared traefik — all serving
|
||||
claimed None ← must be re-claimed before any controller endpoint responds
|
||||
escrow_state pending ← R-196: the Re-issue marked the escrow stale
|
||||
repo key 8a9e33aa4da6… ← THE RECOVERED KEY, on disk
|
||||
sentinel absent from disk (deliberately) — present in off-site snapshot e6132ae5
|
||||
```
|
||||
|
||||
**To resume (operator present, ~5 minutes):**
|
||||
|
||||
1. Re-claim the box — `docker exec felhom-controller /usr/local/bin/felhom-controller
|
||||
--print-reset-code`, then the claim page.
|
||||
2. Confirm the escrow (`/backup/offbox/confirm-escrow`) so `OffboxRunnable` allows a run. **Do NOT run
|
||||
a new ceremony** — it would supersede the identity blob and destroy the key under test.
|
||||
3. Run an off-site backup. **The observable is whether the repository OPENS** — and whether the
|
||||
pre-wipe snapshot `e6132ae5` still exists with the sentinel in it. A same-day `forget` keeps one
|
||||
snapshot per tag, so the *count* is a poor discriminator; the surviving history is the real one.
|
||||
4. Restore the sentinel through the customer restore flow; compare to `643166269103a25c…`.
|
||||
|
||||
**Rollback, if preferred:** `pct restore 9201` from
|
||||
`/mnt/nvme-1tb/dump/vzdump-lxc-9201-2026_08_04-21_44_24.tar.zst` — verified by a full read before the
|
||||
wipe. It returns the box to its pre-wipe state and voids the remaining drill.
|
||||
|
||||
---
|
||||
|
||||
## 8. R persisted nowhere — searched, with a positive control
|
||||
|
||||
Swept the agent journal, the controller container log, and `/tmp`, `/var/tmp`, `/var/lib/felhom-agent`,
|
||||
`/root` on demo-hp for the recovery code: **0 hits**, and **0** leftover `felhom-idesc-*` unseal
|
||||
staging directories. A planted copy was found by the same sweep (**1**) and not found after shredding
|
||||
(**0**), so the instrument is shown sensitive rather than assumed to be.
|
||||
|
||||
---
|
||||
|
||||
## 9. Teardown — three layers
|
||||
|
||||
| layer | state |
|
||||
|---|---|
|
||||
| the guest | **Nothing torn down.** The wipe is the evidence; the apps are serving; the recovered key is in place. The scratch band is empty — no restore-test guest was created tonight. |
|
||||
| the host | `vzdump` snapshot LVs removed cleanly by the archive job (`snap_vm-9201-disk-0/1_vzdump` both released). One new archive, 1.6 GB, on `nvme-1tb` (883 GB free). Nothing deleted. |
|
||||
| the hub | **No new customer records.** `demo-hp` is the same customer row throughout — the rebuild was a controller-data wipe, not a re-enrolment, so nothing accumulated. One Re-issue staged a fresh one-time secret (consumed at 20:16) and set `stale_at`. The two superseded escrow rows are **unchanged** — no ceremony ran. |
|
||||
|
||||
**Nothing was deleted on the storage endpoint**, including the ~1.2 GB of previously-orphaned
|
||||
ciphertext. Ruled, still owed, and deliberately not ridden along with a drill.
|
||||
|
||||
## 10. Part 2 — not run, and why
|
||||
|
||||
Its gate is *"the drill PASSED"*. It did not — the verdict was not reached. Running a second wipe on a
|
||||
box whose first result is incomplete would have destroyed the staged state that makes the first one
|
||||
finishable in five minutes. **R-198's retention therefore remains unit-proven only**, unchanged from
|
||||
this morning.
|
||||
File diff suppressed because one or more lines are too long
@@ -66,6 +66,7 @@
|
||||
| R-203 | **A mandatory customer data directory was silently absent from the off-site snapshot while the run reported `ok`** — the deploy-time `${USERDATA_PATH}` root and the backup-time namespace root disagree for an app on the system drive | M | **OPEN — halted the R-201 drill 2026-08-04** | Flips nothing yet. Blocks the off-site app-data row from ever earning a customer-file-restored badge. Fix shape: one root function, and a MANDATORY skip must be customer/hub-visible rather than a container-log WARN |
|
||||
| R-201 | **The wipe-and-recover drill** | L | **PREPARED, HALTED BEFORE THE WIPE (2026-08-04)** | Nothing irreversible done. Established live for the first time: a rebuilt box's off-site run REFUSES (the orphan card, not a silent fresh history — closes R-193 Q3), the orphan reset works move-aside-never-delete, and demo-hp's pre-rebuild key is permanently gone (superseded four hours before v0.93.0). Resume needs R-203 |
|
||||
| R-203 | **The app and its backup looked in different directories, and a run that skipped a mandatory folder still said `ok`** | M | **SHIPPED + PROVEN-LIVE (controller v0.197.0, 2026-08-04)** | Adds the capability-map row *"off-site app-data capture covers MANDATORY paths on both drive layouts"* as PROVEN-LIVE, and corrects that row's predecessor, which was optimistic. Unblocks R-201 |
|
||||
| R-204 | **The recovered key cannot be used: the Re-issue that reconfigures a rebuilt box's off-site tier marks the escrow stale, which gates every run, and the only way to clear it destroys the key** | M | **OPEN — the wall the 2026-08-04 night drill hit** | Blocks R-201's verdict. Fix is most likely R-196's (mark stale on evidence, not on a Re-issue) or letting a verified recovered key confirm the escrow without a ceremony |
|
||||
| R-19 | Internet-outage customer-experience drill: pull WAN on demo, verify lan_resolver path, document what the customer actually sees/does | S | idea | Flips map row E "LAN access" IMPLEMENTED→PROVEN-LIVE |
|
||||
| R-20 | ~~Verify operator-key pinning is fully in the day-0 install flow~~ | XS | **closed** (2026-07-16) | Confirmed against `scripts/felhom-host-install.sh` source (not changelog): keys resolve at L1181–1219 (script constants `OPERATOR_KEY_*`, populated, `--operator-pubkey-file` override), pinned automatically by `step_agent_config()` "STEP 6/8" (L2044; python builds `authz.signers` L2146–2156, reinstall preserves existing), verified at L2332–2337 ("authz signers: N … operator-signed self-update armed"). No interactive prompt or post-install hand-edit — fully automatic. Doc-drift note: the L193–197 "EMPTY by default" comment is stale vs the now-populated constants (→ R-16 hygiene) |
|
||||
| R-23 | **Immediate-sync Direction-2 follow-ups** (hub v0.58 / controller v0.140, 2026-07-16): **(a) — BANKED 2026-07-21 (both legs).** The operator-UI save->apply round trip is PROVEN: the STOP-2 global-floor save (hub `18:56:27 CEST`) released the controller's held wait in the **same second** (`16:56:27Z wait woke: generation=1 - firing out-of-cycle report`), with the report built 2 s later; the ring also shows `wait baseline generation=0` at startup (baseline recorded WITHOUT firing, as designed) then `generation=1`, so the generation advanced past 0. **RESTART LEG BANKED 2026-07-21 — the floor was moved to a version the box did NOT run, and the swap fired EXACTLY ONCE.** Operator saved global floor 0.153.0 → **v0.154.0** (a real version boundary, unlike the 2026-07-20 attempt which targeted an already-running version and therefore proved nothing). Timeline (guest UTC): `06:57:13` UpdateState `pending` written `initiated_by=auto-floor` → `06:57:17` agent `controller swap requested 0.153.0 -> 0.154.0` → `06:57:19` `image file written, restarting bootstrap` → `06:57:21` container StartedAt + UpdateState `completed_at` → `06:57:29` agent `new controller healthy`. **16 s end to end.** Assertions over the whole window (06:50 → 07:29, 39 min): `controller swap requested` = **1**, agent-driven bootstrap restarts = **1**, `new controller healthy` = **1**, rollback/swap-failed/unhealthy = **0**, container `RestartCount` = **0**. `VerifyStartup` banked it on the next boot (`Post-update startup: update successful (0.153.0 → 0.154.0)`) and the `06:57:52` periodic check logged `Current version 0.154.0 is up to date` — the at/above-floor branch correctly doing nothing. **No storm, no rollback, no second attempt.** *Caveat, disclosed: a hand-deploy of v0.155.0 at `07:17:10` sits inside the observation window and is what StartedAt shows after that point; it never touches `SwapController`, so the swap-count assertions above are uncontaminated across the full window. A second, unplanned confirmation of the at/above-floor branch came with it — after the hand-deploy the box ran 0.155.0 against a 0.154.0 floor and the updater logged `Current version 0.155.0 is up to date` and did nothing.* Evidence: `felhom-controller/REPORT.md` §6 (2026-07-21). *(Superseded note:)* **the self-restart single-fire leg was** - the floor was set to a version the box ALREADY ran, so there was no work and no restart. Finish by bumping the floor to a version the box does NOT run, debug ring open, asserting EXACTLY ONE restart. **Trap found while banking this: the wake is `logx.Debugf`, so it is INVISIBLE in `docker logs` at INFO** and lives only in the debug ring (`GET /api/debug/logs?level=DEBUG`) - a hunter looking at stdout wrongly concludes the box never woke; arguably (b) generalised. (b) cosmetic: the Waiter's "recovered" INFO logs on the next hold completion (`pollOnce` blocks ~240 s), not at reconnect | S | **(a) BANKED in full; only (b) cosmetic remains** | Map row "config/state change round-trips in seconds" flipped PARTIAL->PROVEN-LIVE 2026-07-21 on this evidence. Evidence: `felhom-controller/REPORT.md` 4f |
|
||||
|
||||
Reference in New Issue
Block a user