R-201 PASSED: a customer's file survived a machine rebuild and came back byte-identical
gates / gates (push) Successful in 7s

This commit is contained in:
2026-08-04 23:18:47 +02:00
parent b228fd102d
commit 2a7ac03c47
7 changed files with 194 additions and 188 deletions
+19 -10
View File
@@ -17,23 +17,32 @@
## Standing rulings ## Standing rulings
**S-31 — the off-site key IS recoverable after a real rebuild; USING it is blocked by the remedy that **S-31 — THE DRILL PASSED: a customer's file survives a machine rebuild and comes back byte-identical.
makes the rebuild survivable (2026-08-04 night drill, R-201/R-204).** The capability is proven; the customer JOURNEY is four undocumented manual steps (2026-08-04, R-201/R-204).**
**Proven on hardware:** demo-hp's controller data volume was destroyed and the sentinel deleted; the **Proven on hardware, end to end:** demo-hp's controller data volume was destroyed and the sentinel
customer's recovery code produced `8a9e33aa4da6…`, byte-identical to the pre-wipe on-disk key and to deleted from disk. The recovery code produced `8a9e33aa4da6…` (byte-identical to the pre-wipe on-disk
the hub's independent record, and installed cleanly on the bare box. `identity_blob` was unchanged key AND the hub's independent record); it installed on the bare box; the **existing repository opened**
across the wipe — nothing re-escrowed itself. (`repo_state: null`, 3 snapshots, 42 026 B = the pre-wipe size — not a fresh history); and the customer
restore flow returned the sentinel at `643166269103a25c…`, **byte-identical**. `identity_blob` was
unchanged throughout and **no ceremony was run at any point** (superseded rows still 2).
**The wall (R-204), four links, all measured:** **But it took four manual steps that are in no design document (R-204), all measured:**
1. a rebuilt controller cannot configure its off-site tier — the one-time password was consumed by its 1. a rebuilt controller cannot configure its off-site tier — the one-time password was consumed by its
predecessor (`no unconsumed offsite password`, R-193); predecessor (`no unconsumed offsite password`, R-193);
2. the Re-issue that fixes that sets `stale_at` **while `restic_pw_sha256` is unchanged** (R-196); 2. the Re-issue that fixes that sets `stale_at` **while `restic_pw_sha256` is unchanged** (R-196);
3. a stale escrow makes the hub withhold the hash from the ACK → `EscrowAutoConfirmer` can never flip 3. a stale escrow makes the hub withhold the hash from the ACK → `EscrowAutoConfirmer` can never flip
`pending → escrowed``OffboxRunnable` refuses every run; `pending → escrowed``OffboxRunnable` refuses every run;
4. the only documented way to clear it is a ceremony, **which supersedes the identity blob and destroys 4. it was cleared with the **manual** confirm (`/backup/offbox/confirm-escrow`), NOT a ceremony — a
the recovered key**. And before any of it, a rebuilt box is **unclaimed**, so the claim gate ceremony would have superseded the identity blob and destroyed the recovered key;
intercepts every controller endpoint — a step in no design document. 5. and the restore's **default `mode=unit` returns the recovery unit, not the customer's files**
`mode=full` is what returns documents. A customer following the obvious path gets the wrong thing
and is told nothing.
**Before any of it, a rebuilt box is unclaimed** and the claim gate intercepts every controller
endpoint. **The local escape hatch does not work unaided:** `--print-reset-code` writes the new hash to
`settings.json` while the running controller keeps its old copy in memory, so the code never matches —
**restart the controller between minting and claiming.**
*Facts a future session needs:* *Facts a future session needs:*
- **A guest rebuild in this fleet is a controller-DATA-VOLUME loss, not a guest reprovision.** The - **A guest rebuild in this fleet is a controller-DATA-VOLUME loss, not a guest reprovision.** The
+86 -89
View File
@@ -1,127 +1,124 @@
# REPORT — R-201 night run: the key came back; the verdict did not # REPORT — R-201: **THE DRILL PASSED.** A customer's file survived a machine rebuild
**2026-08-04, 21:3022:40, unattended.** `demo-hp` was deliberately rebuilt. No code, no version bump. **2026-08-04, 21:3023:15** · `demo-hp` deliberately rebuilt · no code, no version bump ·
`demo-felhom` untouched. Full record: `documentation/audits/DRILL-r201-night-run-2026-08-04.md`. `demo-felhom` untouched. Record: `documentation/audits/DRILL-r201-night-run-2026-08-04.md`.
--- ---
## 1. THE VERDICT — not reached ## 1. THE VERDICT — PASS
| | | ```
|---|---| pre-wipe sha256: 643166269103a25cf41d34a26b75fd6ebba0a837bbcd7e0d2b5aba7106bcbe7c
| sentinel sha256, pre-wipe | `643166269103a25cf41d34a26b75fd6ebba0a837bbcd7e0d2b5aba7106bcbe7c` | restored sha256: 643166269103a25cf41d34a26b75fd6ebba0a837bbcd7e0d2b5aba7106bcbe7c
| sentinel sha256, post-restore | **not obtained** — the restore was never reachable | ```
**Not a FAIL.** Nothing came back wrong and no fresh history was started; the box never got as far as **Byte-identical.** The controller's data volume was destroyed and the file deleted from disk; the
running a backup. **What it is instead:** the first proof that the key survives and returns, plus the customer's recovery code, the customer's restore flow, and the file came back. **The Felhom backup
measured reason a customer still cannot use it. story is proved end to end for the first time.**
> **After a real rebuild — controller data volume destroyed, sentinel deleted from disk — the ## 2. Snapshot count at step 9 — **3, not 1**
> customer's recovery code produced `8a9e33aa4da6769c5aea1831f87759e10930e2ec1dea0062576484e0598d080a`,
> byte-identical to the pre-wipe on-disk key and to the hub's independent record.** That has never been
> shown before.
## 2. Snapshot count at step 9 — not obtained `repo_state: null` (the repository **opened**, not orphaned), `snapshot_count: 3`, `repo_size_bytes:
42 026` — the pre-wipe size exactly. The pre-wipe snapshot `e6132ae5` was still present with the file
The off-site run was never permitted to start (§4). **And a count would have been a poor inside it. **No fresh history was started over the old one** — the failure this drill was designed to
discriminator anyway:** restic's same-day `forget --keep-daily 7 --group-by host,tags` keeps one catch did not occur.
snapshot per tag per day, so a successful reopen would have shown 3, not 4. The real observable is
whether the pre-wipe snapshot `e6132ae5` survives with the sentinel in it — recorded as the resume
step.
## 3. §5's five conditions, recorded before the wipe ## 3. §5's five conditions, recorded before the wipe
1. **sentinel listed by name** snapshot `e6132ae5` (19:36:26), `-rw-r--r-- 1000 1000 181 …/DRILL-SENTINEL.txt`. 1. sentinel **listed by name** in snapshot `e6132ae5` (19:36:26), 181 B.
2. **rollback archive verified** `vzdump-lxc-9201-2026_08_04-21_44_24.tar.zst`, 1 606 765 083 B, 2. rollback archive `vzdump-lxc-9201-2026_08_04-21_44_24.tar.zst` (1 606 765 083 B) **verified by a
**full zstd stream read OK** (4 867 573 760 B), sentinel confirmed inside it. full zstd stream read** (4 867 573 760 B) with the sentinel confirmed inside.
3. §3's option — §5 below. 3. §3's option — §4 below.
4. **demo-felhom** `health=ok`, `escrow_state=escrowed`, untouched throughout. 4. `demo-felhom` `health=ok`, `escrow_state=escrowed`, untouched throughout.
5. **space** nvme-1tb 883 GB free, root 20 GB, local-lvm 34.32 %. 5. nvme-1tb 883 GB free, root 20 GB, local-lvm 34.32 %.
**One precondition had drifted and was repaired, not worked around.** The staged snapshot no longer **One precondition had drifted and was repaired, not worked around:** the staged snapshot had lost the
held the sentinel: the afternoon's `mandatory data path missing` experiment produced a *later* same-day sentinel because the afternoon's experiment produced a *later* same-day snapshot and `forget
calibre-web snapshot and `forget` had pruned the good one. The fixture on disk was correct, so one --keep-daily 7 --group-by host,tags` pruned the good one. One backup re-established it, re-verified by
backup re-established it and it was re-verified by listing. **Lesson: a good snapshot is not durable listing. **A good snapshot is not durable against a later bad run on the same day.**
against a later bad run on the same day.**
## 4. Every step's observable ## 4. §3 — the recovery code
| step | observable | **Option B as already in place, improved: no new copy was made, so nothing needed shredding.** The
|---|---| operator's own `~/.config/credentials` (`0600`) was the source; the code was **piped to stdin** for
| 6 | fresh data dir stamped `20:00:2x`; **new `encryption.key`** (32 B); `claimed = None`; `offbox = null` | each invocation — never an argument, never exported, never a second file, never logged. Destroying the
| **7** | **`identity_blob` 572 B, `restic_pw_sha256` `8a9e33aa4da6…`, `updated_at` still `11:11:37` — UNCHANGED across the wipe.** Nothing re-escrowed itself | operator's own store would have destroyed their record; because no extra copy existed, there is nothing
| **8** | **recovered `8a9e33aa4da6…`** — matches pre-wipe on-disk and the hub's record. Exit 1 is correct and designed: nothing local to compare against, the rebuilt-box shape | left to prove gone.
| 9a | `[INSTALLED] … reads back identical` — the "installed" branch's first real run |
| 9b | after the Re-issue and apply, the on-disk key is **still** `8a9e33aa4da6…``WriteOffboxSecrets` kept it |
| 9c | **blocked** — see below |
| 1011 | **not run** |
**The wipe was faithful to the incident, deliberately.** The 2026-08-03 rebuild R-193 is filed against
was **not** a guest reprovision — the journal shows guest 9201 running continuously through that window
with no `pct destroy`, no `pct restore` and no `--selftest=provision`. What changed was the controller
and its data volume. A guest destroy plus an unrehearsed provisioning chain, improvised unattended, is
what §8.10 exists to prevent.
## 5. §3 — the recovery code
**Option B as already in place, improved: no new copy was created, so nothing needed shredding.** The
operator had placed `R_DEMO-HP` in their own `~/.config/credentials` (`0600`) two sessions ago for this
purpose. It was read from there and **piped to stdin** for the two invocations that needed it — never
an argument, never exported, never written to a second file, never logged. Destroying the operator's
own store would have destroyed their record; because no additional copy existed, there is nothing left
to prove gone.
**Verified anyway, with the planted-copy positive control:** 0 hits in the agent journal, 0 in the **Verified anyway, with the planted-copy positive control:** 0 hits in the agent journal, 0 in the
controller log, 0 files under `/tmp`, `/var/tmp`, `/var/lib/felhom-agent`, `/root`, 0 leftover controller log, 0 files under `/tmp`, `/var/tmp`, `/var/lib/felhom-agent`, `/root`, 0 leftover
`felhom-idesc-*` staging dirs — and the same sweep found a planted copy (**1**), then **0** after `felhom-idesc-*` dirs — and the same sweep found a planted copy (**1**), then **0** after shredding it.
shredding it. The instrument is shown sensitive rather than assumed to be.
## 6. Part 2 — did not run ## 5. Every step's observable
Its gate is "the drill PASSED". It did not. A second wipe would have destroyed the state that makes the | step | observable |
first drill finishable in five minutes. **R-198's retention therefore remains unit-proven only** |---|---|
nothing has yet superseded a key in production. | 6 | fresh data dir `20:00:2x`; new `encryption.key`; `claimed = None`; `offbox = null` |
| **7** | **`identity_blob` 572 B, `restic_pw_sha256` `8a9e33aa4da6…`, `updated_at` still `11:11:37` — UNCHANGED.** Nothing re-escrowed itself |
| **8** | **recovered `8a9e33aa4da6…`** — matches the pre-wipe on-disk key and the hub's record |
| 9a | `[INSTALLED] … reads back identical` |
| 9b | after Re-issue + apply, the on-disk key is **still** the recovered one |
| 9c | **repository OPENED**`repo_state: null`, 3 snapshots, 42 026 B |
| **10** | **restored sha256 byte-identical** |
| 11 | this record |
Superseded escrow rows remained **2** throughout — **no ceremony was run at any point.**
## 6. Part 2 — not run
Gate: *"the drill PASSED; its evidence is written down; and there is time."* The first two are met; the
third is not — 23:15, after a full destructive cycle, and Part 2 is a second wipe. Deliberately left
for a fresh session with drill 1's result already recorded, so a second wipe cannot overwrite it.
**R-198's retention remains unit-proven only** — nothing has yet superseded a key in production, and it
is now the last unproven link in this chain.
## 7. Teardown — three layers ## 7. Teardown — three layers
| layer | state | | layer | state |
|---|---| |---|---|
| the guest | **nothing torn down** — the wipe is the evidence; all six app containers serving; the recovered key on disk; scratch band empty | | the guest | **healthy and re-armed**: controller v0.197.0 claimed, six app containers serving, all three apps re-toggled for off-site, the sentinel restored to its live location and verified. The verification-copy tree is **kept as evidence**. Scratch band empty |
| the host | vzdump snapshot LVs released cleanly; one new 1.6 GB archive on nvme-1tb (883 GB free); nothing deleted | | the host | vzdump snapshot LVs released cleanly; one new 1.6 GB archive on nvme-1tb (883 GB free); nothing deleted |
| the hub | **no new customer records**`demo-hp` is the same row throughout, because the rebuild was a controller-data wipe and not a re-enrolment. One Re-issue staged a one-time secret (consumed 20:16) and set `stale_at`. The two superseded escrow rows are **unchanged** — no ceremony ran | | the hub | **no new customer records**`demo-hp` is the same row throughout (a controller-data rebuild, not a re-enrolment). One Re-issue staged a one-time secret (consumed 20:16) and set `stale_at`. **The two superseded escrow rows are unchanged** |
**Nothing deleted on the storage endpoint**, including the ~1.2 GB of orphaned ciphertext — ruled, **Nothing deleted on the storage endpoint**, including the ~1.2 GB of orphaned ciphertext — ruled,
still owed, deliberately not ridden along with a drill. still owed, deliberately not ridden along with a drill.
## 8. The capability-map rows ## 8. The capability-map row
A new row records the key as **PROVEN-LIVE after a real rebuild**, and states plainly what it does not Now reads *"A customer's file survives a machine rebuild and comes back — the whole off-site story, end
claim: **no file has been restored**, and the four links of R-204 stand between the recovered key and a to end"*, **PROVEN-LIVE**. What it deliberately does **not** claim: the journey is manual and
usable repository. The R-199 back-pointer added earlier today stands. undocumented (R-204); the scope is `demo-hp` and a **controller-data** rebuild, **not** a total host
loss and **not** a guest reprovision; and R-198's retention is still unit-proven.
## 9. New findings ## 9. New findings — R-204, expanded into the gap list
- **R-204 (NEW)** — the recovered key cannot be used: a rebuilt box cannot configure its tier (R-193); The four steps between a recovered key and a restored file, all measured while walking them:
the Re-issue that fixes that marks the escrow stale though the key never changed (R-196, **measured
at `20:15:49` with `restic_pw_sha256` unchanged**); a stale escrow gates every run; and the only way 1. **R-193** — a rebuilt controller cannot configure its off-site tier (`no unconsumed offsite
to clear it is a ceremony that **destroys the recovered key**. Plus a fourth link in no design password`; the ledger shows its predecessor consumed it at 07:12:06). Needs an operator Re-issue.
document: a rebuilt box is **unclaimed**, so the claim gate intercepts every controller endpoint. 2. **The claim gate** — a rebuilt box is unclaimed, so *every* controller endpoint is intercepted. And
- **R-196 re-scoped** — no longer a documentation nit; it is on the critical path for recovery. **the local escape hatch does not work unaided**: `--print-reset-code` writes the new hash to
- **R-201** — the wipe happened, the key came back, ~5 minutes from a verdict with a person present. `settings.json` while the running controller keeps its old copy in memory, so `effectiveClaimCode()`
- **R-198** — still unit-proven only; Part 2 gated out. never sees it. **Restart the controller between minting and claiming** — two claim attempts failed
- **R-202 open; the ciphertext deletion still owed.** before this was diagnosed.
3. **R-196** — the Re-issue sets `stale_at` (`20:15:49`) while `restic_pw_sha256` is unchanged, which
gates every off-site run. Cleared with the **manual** confirm, never a ceremony.
4. **`mode=unit` is the restore default** and returns the recovery unit, **not** the userdata leg. A
customer told to "restore from off-site" gets their app definition and not their documents, and
nothing in the outcome says so. **The worst of the four**, because it fails silently at the last
step.
## 10. CI ## 10. CI
Docs push only. Run number and task id in the session summary. **`--no-verify` not used.** Docs push only; run number and task id in the session summary. **`--no-verify` not used.**
## 11. Observations — noticed, NOT acted on ## 11. Observations — noticed, NOT acted on
1. **The claim gate is an undocumented first step of every recovery.** Before a customer can do 1. **A same-day re-run replaces the day's snapshot** (`forget --keep-daily 7 --group-by host,tags`).
anything on a rebuilt box — including recovering their backups — they must re-claim it. Any drill depending on a specific snapshot surviving must account for it.
2. **`--print-reset-code` output needs parsing care**: the captured value was 73 characters, i.e. more 2. **The hub's ClusterIP is not reachable from DooPlex's host network** — the operator UI needs a
than the code itself. Not chased; the claim was abandoned when the session stopped. `kubectl port-forward`; the memory's `curl -u :$HUB_PW` recipe omits that.
3. **A same-day re-run replaces the day's snapshot.** Worth knowing before designing any drill that 3. **`--print-reset-code` mints a new generation on every invocation** — three wasted generations here
depends on a specific snapshot surviving. (4, 5, 6) while diagnosing the in-memory staleness.
4. **The hub's ClusterIP is not reachable from DooPlex's host network** — the operator UI needs a 4. **The restore wrote into `backups/offsite-restore/<app>/` mirroring the full absolute path** — deep
`kubectl port-forward`. The `curl -u :$HUB_PW` recipe in memory omits that. but unambiguous; worth knowing before writing customer-facing copy about where files land.
+20 -22
View File
@@ -36,21 +36,19 @@ Proven end to end on real hardware.
the old backups stay recoverable. Both keys are now kept. **What this cannot undo:** anything the old backups stay recoverable. Both keys are now kept. **What this cannot undo:** anything
superseded before today is gone for good, which includes both demo machines' pre-4-August keys — so superseded before today is gone for good, which includes both demo machines' pre-4-August keys — so
those 51 orphaned backups were beyond reach even if the recovery codes had been kept. *(R-198)* those 51 orphaned backups were beyond reach even if the recovery codes had been kept. *(R-198)*
- **THE BACKUP KEY COMES BACK AFTER A REBUILD — proved on real hardware last night.** We destroyed the - **IT WORKS. A customer's file survived a machine being destroyed and came back — proved last night on
HP machine's controller data on purpose, deleted the marked file from its disk, and then used the real hardware, end to end.** We wiped the HP machine's controller data on purpose and deleted the
recovery code you saved. The key that came out was **identical, character for character**, to the one marked file from its disk. Using the recovery code you saved: the backup key came back **identical,
the machine had been using — and to the fingerprint the hub had recorded separately. It installed character for character**; the existing off-site store **opened** rather than starting over (the same
cleanly onto the empty machine. Nothing re-sealed itself in the meantime, so the sealed package three backups, the same 42 026 bytes as before the wipe); and the file was restored **byte for byte
survived the rebuild untouched. *(R-201)* identical**. That is the whole backup promise, demonstrated for the first time. *(R-201)*
- **But the machine still could not use it, and that is the night's real finding.** Three things stand - **But a customer could not have done it alone, and that is now the open work.** Getting from "the key
between a recovered key and a restored file, and each one is now measured rather than guessed: is recoverable" to "the file is back" took four steps that appear in no instructions: re-issuing the
**(1)** a rebuilt machine cannot set up its off-site connection at all — its one-time password was storage credential; re-claiming the machine (whose local reset-code tool **does not work until the
spent by the machine it replaced; **(2)** the fix for that (re-issuing the credential) marks the controller is restarted** — two attempts failed before we spotted it); manually clearing a "stale"
sealed package "stale" even though the key never changed, and a stale package blocks every off-site mark that the re-issue sets even though the key never changed; and choosing the *full* restore,
backup; **(3)** the only documented way to clear that is to make a new recovery code — **which because the default one returns the app's settings and **not** the customer's documents — with
replaces the sealed package and destroys the key we just recovered.** And before any of it, the nothing saying so. **The capability is real; the experience is not built yet.** *(R-204)*
rebuilt machine is unclaimed, so nothing on its dashboard responds until the customer claims it
again — a step written down nowhere. **The key comes back and cannot be used.** *(R-204)*
- **One screen still tells the customer something we cannot yet promise.** The „elárvult tároló" card - **One screen still tells the customer something we cannot yet promise.** The „elárvult tároló" card
says the old backups may later be restorable with the matching recovery code. From today that is true says the old backups may later be restorable with the matching recovery code. From today that is true
for machines that re-seal from now on and **false for anything already orphaned** — and the machine for machines that re-seal from now on and **false for anything already orphaned** — and the machine
@@ -99,10 +97,11 @@ Proven end to end on real hardware.
## What we're working on ## What we're working on
- **Now:** finishing the proof — it is about five minutes with you present. The HP machine is sitting - **Now:** making that journey something a customer can actually follow — the four steps above, worst
first: the restore default that silently returns the wrong thing. The HP machine is sitting
mid-drill with the recovered key already on it; it needs re-claiming and one setting confirmed, then mid-drill with the recovered key already on it; it needs re-claiming and one setting confirmed, then
the backup runs and the file is restored. **Do not let it make a new recovery code** — that would the backup runs and the file is restored. **Do not let it make a new recovery code** — that would
destroy the key we recovered. Then the deeper fix for the wall above. The marked file now lands in The marked file now lands in
the off-site backup, so there is finally something to recover. Everything else is already in place on the off-site backup, so there is finally something to recover. Everything else is already in place on
the HP machine — the recovery code you saved, a working off-site store, the app and the file. Nothing the HP machine — the recovery code you saved, a working off-site store, the app and the file. Nothing
has been wiped; that step waits for your go-ahead at a marked stop point. Both honesty fixes shipped today: a changed backup key now has been wiped; that step waits for your go-ahead at a marked stop point. Both honesty fixes shipped today: a changed backup key now
@@ -157,11 +156,10 @@ Proven end to end on real hardware.
## Changed since last update ## Changed since last update
- **2026-08-04 (night)** — **Wiped the HP machine on purpose and got the backup key back with your - **2026-08-04 (night)** — **The drill PASSED.** We destroyed a machine on purpose and the customer's
recovery code — identical, character for character.** First time that has ever been done. The drill file came back byte-for-byte identical, using the recovery code you saved. First time the backup
then stopped short of restoring the file, at a wall worth more than the last step: the key comes back story has been proved end to end. It needed four undocumented manual steps to get there, which are
but cannot be used without an act that destroys it. The machine is up, its apps are serving, and it now the next piece of work. The machine is up, healthy and re-armed. *(R-201, R-204)*
is five minutes from finished. *(R-201, R-204)*
- **2026-08-04 (evening)** — **Fixed the folder-left-out-of-the-backup problem, both halves.** The app - **2026-08-04 (evening)** — **Fixed the folder-left-out-of-the-backup problem, both halves.** The app
and its backup now look in the same directory, and a backup that misses a folder marked essential and its backup now look in the same directory, and a backup that misses a folder marked essential
reports *incomplete* instead of success. Proved by listing the backup's own contents and finding the reports *incomplete* instead of success. Proved by listing the backup's own contents and finding the
File diff suppressed because one or more lines are too long
@@ -1,38 +1,47 @@
# DRILL — R-201, the night run: **THE KEY CAME BACK. The verdict was not reached.** # DRILL — R-201, the night run: **PASS. The customer's file came back, byte-identical.**
**Date:** 2026-08-04, 21:3022:40 · **Box:** `demo-hp` (Tier 0) · **Unattended, by operator decision.** **Date:** 2026-08-04, 21:3023:15 · **Box:** `demo-hp` (Tier 0) · unattended, then **finished with the
**The wipe happened.** The box is up, its apps are serving, and it is left mid-drill by deliberate operator present**. The box is up, healthy, and fully re-armed.
choice — see §7 for the exact state and the one command that resumes it.
> **THE RESULT, first.** A real rebuild was performed: the controller's data volume was destroyed and > **THE VERDICT: PASS.**
> the sentinel deleted from disk. The customer's recovery code then produced
> `8a9e33aa4da6769c5aea1831f87759e10930e2ec1dea0062576484e0598d080a` — **byte-identical to the key the
> box used before the wipe, and to the hash the hub had independently recorded.**
> >
> **The off-site backup key is recoverable after a machine is rebuilt. That has never been shown > ```
> before.** > pre-wipe sha256: 643166269103a25cf41d34a26b75fd6ebba0a837bbcd7e0d2b5aba7106bcbe7c
> restored sha256: 643166269103a25cf41d34a26b75fd6ebba0a837bbcd7e0d2b5aba7106bcbe7c
> ```
> >
> **And the drill did not finish**, because three separate things stand between a recovered key and a > A machine was destroyed on purpose — controller data volume gone, the customer's file deleted from
> restored file. All three are measured below. That is the other half of the night's result, and it is > disk — and **the file came back byte-identical**, through the customer's own recovery code and the
> the half nobody knew. > customer's own restore flow. **This is the first time the Felhom backup story has been proved end to
> end.**
>
> The key recovered byte-identical (`8a9e33aa4da6…`), the existing repository **opened rather than
> being re-initialised** (`repo_state: null`, 3 snapshots, 42 026 B — the pre-wipe size exactly), and
> the pre-wipe snapshot `e6132ae5` was still there with the file inside it.
>
> **It took four undocumented steps to get there**, each measured, and they are the other half of the
> result — see §3. A customer hitting this alone today would not get their data back.
--- ---
## 1. The verdict ## 1. The verdict — **PASS**
**NOT REACHED.** Step 10 (restore the sentinel and compare) was not run.
| | | | | |
|---|---| |---|---|
| sentinel sha256, pre-wipe | `643166269103a25cf41d34a26b75fd6ebba0a837bbcd7e0d2b5aba7106bcbe7c` | | sentinel sha256, pre-wipe (step 3) | `643166269103a25cf41d34a26b75fd6ebba0a837bbcd7e0d2b5aba7106bcbe7c` |
| sentinel sha256, post-restore | **not obtained** — the restore was never reachable | | sentinel sha256, restored (step 10) | **`643166269103a25cf41d34a26b75fd6ebba0a837bbcd7e0d2b5aba7106bcbe7c`** |
| snapshot count at step 9 | **not obtained** — the run was refused before it started | | **match** | **YES — byte-identical** |
| snapshot count at step 9 | **3** — the pre-wipe count. **NOT 1.** No fresh history was started |
| repo size at step 9 | **42 026 B** — the pre-wipe size exactly |
| `repo_state` | **`null`** — the repository OPENED. Not orphaned |
**This is not a FAIL.** A fail is "the file came back wrong" or "a fresh empty history was started". The restore ran through the customer flow (`POST /backup/offbox/restore`, `mode=full`, two-step
Neither happened, because the box never got as far as running a backup. What the night established is confirm), out of snapshot `e6132ae5` — the one taken **before** the wipe.
where the wall is.
--- **A detail worth keeping:** the first restore attempt used the default `mode=unit`, which recovers the
recovery unit only and **not** the userdata leg. `mode=full` is what returns the customer's files. A
customer told to "restore from off-site" would, on the default, get their app definition back and not
their documents — and nothing in that outcome says so.
## 2. What was proven, in order, on hardware ## 2. What was proven, in order, on hardware
@@ -153,23 +162,25 @@ that appears in no design document, and it comes *before* anything else can happ
--- ---
## 4. Why the session stopped here ## 4. How it was finished, and the fourth undocumented step
By 22:35 the path forward was a chain of workarounds assembled live — Re-issue, then the deprecated The unattended run stopped at the wall in §3 rather than improvising further at 22:35 (§8.10). With the
manual confirm, then the root-gated `--print-reset-code` escape hatch to re-claim, then the confirm operator present the remaining path was walked, and it took **four steps a customer would have to
again, then the run. Each is individually defensible; the accumulation is precisely the pattern §8.10 discover**:
names:
> *"If anything is ambiguous at 03:00, stop and leave it for the morning. An unattended session that 1. **Re-issue** the offsite credential (hub, operator-only) — otherwise the box can never configure its
> halts with a clear state beats one that improvises."* tier. §3(a).
2. **Re-claim the box.** The local escape hatch (`--print-reset-code`) mints a code — **and does not
work on its own**: it writes the new hash to `settings.json` while the RUNNING controller keeps its
old copy in memory, so `effectiveClaimCode()` never sees it and every attempt returns
*"Hibás vagy lejárt kód"*. **The controller must be restarted between minting and claiming.** Two
claim attempts failed on this before it was diagnosed; the generation counter reached 6.
3. **Confirm the escrow** manually (`/backup/offbox/confirm-escrow`) → `escrowed`, so `OffboxRunnable`
permits a run. **Not a ceremony** — the identity blob was never superseded (still 2 retained rows).
4. **Restore with `mode=full`.** The default `mode=unit` returns the recovery unit and **not** the
customer's files.
The remaining steps need about five minutes **with a person present**. They are not worth improvising **None of these four is in any design document, and the drill needed all four.**
alone at the end of a long night, on the one box whose off-site history the drill is trying to prove.
**Nothing was left broken.** The box is up, all six app containers are serving, and the recovered key
is on disk.
---
## 5. §5's five conditions, as recorded before the wipe ## 5. §5's five conditions, as recorded before the wipe
@@ -208,33 +219,18 @@ see §8.
--- ---
## 7. The exact state the box is in, and how to resume ## 7. The state the box is left in — healthy and re-armed
``` ```
controller felhom-controller:0.197.0, healthy controller felhom-controller:0.197.0, healthy, CLAIMED
apps privatebin opengist calibre-web filebrowser cloudflared traefik — all serving apps privatebin opengist calibre-web filebrowser cloudflared traefik — all serving
claimed None ← must be re-claimed before any controller endpoint responds escrow_state escrowed (no ceremony was run; identity_blob untouched)
escrow_state pending ← R-196: the Re-issue marked the escrow stale repo key 8a9e33aa4da6… THE RECOVERED KEY — and the repository opens with it
repo key 8a9e33aa4da6… ← THE RECOVERED KEY, on disk off-site 3 snapshots, 42 026 B, repo_state null; all three apps re-toggled for off-site
sentinel absent from disk (deliberately) — present in off-site snapshot e6132ae5 sentinel restored to its live location, sha256 643166269103a25c… — verified
``` ```
**To resume (operator present, ~5 minutes):** The verification copy under `backups/offsite-restore/calibre-web/` is **left in place as evidence**.
1. Re-claim the box — `docker exec felhom-controller /usr/local/bin/felhom-controller
--print-reset-code`, then the claim page.
2. Confirm the escrow (`/backup/offbox/confirm-escrow`) so `OffboxRunnable` allows a run. **Do NOT run
a new ceremony** — it would supersede the identity blob and destroy the key under test.
3. Run an off-site backup. **The observable is whether the repository OPENS** — and whether the
pre-wipe snapshot `e6132ae5` still exists with the sentinel in it. A same-day `forget` keeps one
snapshot per tag, so the *count* is a poor discriminator; the surviving history is the real one.
4. Restore the sentinel through the customer restore flow; compare to `643166269103a25c…`.
**Rollback, if preferred:** `pct restore 9201` from
`/mnt/nvme-1tb/dump/vzdump-lxc-9201-2026_08_04-21_44_24.tar.zst` — verified by a full read before the
wipe. It returns the box to its pre-wipe state and voids the remaining drill.
---
## 8. R persisted nowhere — searched, with a positive control ## 8. R persisted nowhere — searched, with a positive control
@@ -256,9 +252,14 @@ staging directories. A planted copy was found by the same sweep (**1**) and not
**Nothing was deleted on the storage endpoint**, including the ~1.2 GB of previously-orphaned **Nothing was deleted on the storage endpoint**, including the ~1.2 GB of previously-orphaned
ciphertext. Ruled, still owed, and deliberately not ridden along with a drill. ciphertext. Ruled, still owed, and deliberately not ridden along with a drill.
## 10. Part 2 — not run, and why ## 10. Part 2 — not run
Its gate is *"the drill PASSED"*. It did not — the verdict was not reached. Running a second wipe on a Its gate is *"the drill PASSED; its evidence is written down; and there is time."* The first two are
box whose first result is incomplete would have destroyed the staged state that makes the first one now met. **The third was not** — it was 23:15 after a full destructive cycle, and Part 2 is a second
finishable in five minutes. **R-198's retention therefore remains unit-proven only**, unchanged from wipe. It is deliberately left for a session that starts fresh, with drill 1's result already recorded
this morning. so a second wipe cannot overwrite it.
**R-198's retention therefore remains unit-proven only.** Nothing has yet superseded a key in
production, so the retention's live behaviour is still unobserved — and it is now the last unproven
link in the chain this drill exercised. The check is cheap: wipe again, do NOT recover, let a ceremony
seal a different password, then read the superseded row's `identity_blob`.
File diff suppressed because one or more lines are too long
+1
View File
@@ -67,6 +67,7 @@
| R-201 | **The wipe-and-recover drill** | L | **PREPARED, HALTED BEFORE THE WIPE (2026-08-04)** | Nothing irreversible done. Established live for the first time: a rebuilt box's off-site run REFUSES (the orphan card, not a silent fresh history — closes R-193 Q3), the orphan reset works move-aside-never-delete, and demo-hp's pre-rebuild key is permanently gone (superseded four hours before v0.93.0). Resume needs R-203 | | R-201 | **The wipe-and-recover drill** | L | **PREPARED, HALTED BEFORE THE WIPE (2026-08-04)** | Nothing irreversible done. Established live for the first time: a rebuilt box's off-site run REFUSES (the orphan card, not a silent fresh history — closes R-193 Q3), the orphan reset works move-aside-never-delete, and demo-hp's pre-rebuild key is permanently gone (superseded four hours before v0.93.0). Resume needs R-203 |
| R-203 | **The app and its backup looked in different directories, and a run that skipped a mandatory folder still said `ok`** | M | **SHIPPED + PROVEN-LIVE (controller v0.197.0, 2026-08-04)** | Adds the capability-map row *"off-site app-data capture covers MANDATORY paths on both drive layouts"* as PROVEN-LIVE, and corrects that row's predecessor, which was optimistic. Unblocks R-201 | | R-203 | **The app and its backup looked in different directories, and a run that skipped a mandatory folder still said `ok`** | M | **SHIPPED + PROVEN-LIVE (controller v0.197.0, 2026-08-04)** | Adds the capability-map row *"off-site app-data capture covers MANDATORY paths on both drive layouts"* as PROVEN-LIVE, and corrects that row's predecessor, which was optimistic. Unblocks R-201 |
| R-204 | **The recovered key cannot be used: the Re-issue that reconfigures a rebuilt box's off-site tier marks the escrow stale, which gates every run, and the only way to clear it destroys the key** | M | **OPEN — the wall the 2026-08-04 night drill hit** | Blocks R-201's verdict. Fix is most likely R-196's (mark stale on evidence, not on a Re-issue) or letting a verified recovered key confirm the escrow without a ceremony | | R-204 | **The recovered key cannot be used: the Re-issue that reconfigures a rebuilt box's off-site tier marks the escrow stale, which gates every run, and the only way to clear it destroys the key** | M | **OPEN — the wall the 2026-08-04 night drill hit** | Blocks R-201's verdict. Fix is most likely R-196's (mark stale on evidence, not on a Re-issue) or letting a verified recovered key confirm the escrow without a ceremony |
| R-201 | **The wipe-and-recover drill** | L | **PASSED + PROVEN-LIVE (2026-08-04 night)** | Flips the off-site app-data row to *"a customer's file survives a machine rebuild and comes back"* — PROVEN-LIVE, scoped to demo-hp and to a controller-data rebuild. The journey took four undocumented manual steps (R-204); the capability is real, the customer experience is not yet |
| R-19 | Internet-outage customer-experience drill: pull WAN on demo, verify lan_resolver path, document what the customer actually sees/does | S | idea | Flips map row E "LAN access" IMPLEMENTED→PROVEN-LIVE | | R-19 | Internet-outage customer-experience drill: pull WAN on demo, verify lan_resolver path, document what the customer actually sees/does | S | idea | Flips map row E "LAN access" IMPLEMENTED→PROVEN-LIVE |
| R-20 | ~~Verify operator-key pinning is fully in the day-0 install flow~~ | XS | **closed** (2026-07-16) | Confirmed against `scripts/felhom-host-install.sh` source (not changelog): keys resolve at L11811219 (script constants `OPERATOR_KEY_*`, populated, `--operator-pubkey-file` override), pinned automatically by `step_agent_config()` "STEP 6/8" (L2044; python builds `authz.signers` L21462156, reinstall preserves existing), verified at L23322337 ("authz signers: N … operator-signed self-update armed"). No interactive prompt or post-install hand-edit — fully automatic. Doc-drift note: the L193197 "EMPTY by default" comment is stale vs the now-populated constants (→ R-16 hygiene) | | R-20 | ~~Verify operator-key pinning is fully in the day-0 install flow~~ | XS | **closed** (2026-07-16) | Confirmed against `scripts/felhom-host-install.sh` source (not changelog): keys resolve at L11811219 (script constants `OPERATOR_KEY_*`, populated, `--operator-pubkey-file` override), pinned automatically by `step_agent_config()` "STEP 6/8" (L2044; python builds `authz.signers` L21462156, reinstall preserves existing), verified at L23322337 ("authz signers: N … operator-signed self-update armed"). No interactive prompt or post-install hand-edit — fully automatic. Doc-drift note: the L193197 "EMPTY by default" comment is stale vs the now-populated constants (→ R-16 hygiene) |
| R-23 | **Immediate-sync Direction-2 follow-ups** (hub v0.58 / controller v0.140, 2026-07-16): **(a) — BANKED 2026-07-21 (both legs).** The operator-UI save->apply round trip is PROVEN: the STOP-2 global-floor save (hub `18:56:27 CEST`) released the controller's held wait in the **same second** (`16:56:27Z wait woke: generation=1 - firing out-of-cycle report`), with the report built 2 s later; the ring also shows `wait baseline generation=0` at startup (baseline recorded WITHOUT firing, as designed) then `generation=1`, so the generation advanced past 0. **RESTART LEG BANKED 2026-07-21 — the floor was moved to a version the box did NOT run, and the swap fired EXACTLY ONCE.** Operator saved global floor 0.153.0 → **v0.154.0** (a real version boundary, unlike the 2026-07-20 attempt which targeted an already-running version and therefore proved nothing). Timeline (guest UTC): `06:57:13` UpdateState `pending` written `initiated_by=auto-floor``06:57:17` agent `controller swap requested 0.153.0 -> 0.154.0``06:57:19` `image file written, restarting bootstrap``06:57:21` container StartedAt + UpdateState `completed_at``06:57:29` agent `new controller healthy`. **16 s end to end.** Assertions over the whole window (06:50 → 07:29, 39 min): `controller swap requested` = **1**, agent-driven bootstrap restarts = **1**, `new controller healthy` = **1**, rollback/swap-failed/unhealthy = **0**, container `RestartCount` = **0**. `VerifyStartup` banked it on the next boot (`Post-update startup: update successful (0.153.0 → 0.154.0)`) and the `06:57:52` periodic check logged `Current version 0.154.0 is up to date` — the at/above-floor branch correctly doing nothing. **No storm, no rollback, no second attempt.** *Caveat, disclosed: a hand-deploy of v0.155.0 at `07:17:10` sits inside the observation window and is what StartedAt shows after that point; it never touches `SwapController`, so the swap-count assertions above are uncontaminated across the full window. A second, unplanned confirmation of the at/above-floor branch came with it — after the hand-deploy the box ran 0.155.0 against a 0.154.0 floor and the updater logged `Current version 0.155.0 is up to date` and did nothing.* Evidence: `felhom-controller/REPORT.md` §6 (2026-07-21). *(Superseded note:)* **the self-restart single-fire leg was** - the floor was set to a version the box ALREADY ran, so there was no work and no restart. Finish by bumping the floor to a version the box does NOT run, debug ring open, asserting EXACTLY ONE restart. **Trap found while banking this: the wake is `logx.Debugf`, so it is INVISIBLE in `docker logs` at INFO** and lives only in the debug ring (`GET /api/debug/logs?level=DEBUG`) - a hunter looking at stdout wrongly concludes the box never woke; arguably (b) generalised. (b) cosmetic: the Waiter's "recovered" INFO logs on the next hold completion (`pollOnce` blocks ~240 s), not at reconnect | S | **(a) BANKED in full; only (b) cosmetic remains** | Map row "config/state change round-trips in seconds" flipped PARTIAL->PROVEN-LIVE 2026-07-21 on this evidence. Evidence: `felhom-controller/REPORT.md` 4f | | R-23 | **Immediate-sync Direction-2 follow-ups** (hub v0.58 / controller v0.140, 2026-07-16): **(a) — BANKED 2026-07-21 (both legs).** The operator-UI save->apply round trip is PROVEN: the STOP-2 global-floor save (hub `18:56:27 CEST`) released the controller's held wait in the **same second** (`16:56:27Z wait woke: generation=1 - firing out-of-cycle report`), with the report built 2 s later; the ring also shows `wait baseline generation=0` at startup (baseline recorded WITHOUT firing, as designed) then `generation=1`, so the generation advanced past 0. **RESTART LEG BANKED 2026-07-21 — the floor was moved to a version the box did NOT run, and the swap fired EXACTLY ONCE.** Operator saved global floor 0.153.0 → **v0.154.0** (a real version boundary, unlike the 2026-07-20 attempt which targeted an already-running version and therefore proved nothing). Timeline (guest UTC): `06:57:13` UpdateState `pending` written `initiated_by=auto-floor``06:57:17` agent `controller swap requested 0.153.0 -> 0.154.0``06:57:19` `image file written, restarting bootstrap``06:57:21` container StartedAt + UpdateState `completed_at``06:57:29` agent `new controller healthy`. **16 s end to end.** Assertions over the whole window (06:50 → 07:29, 39 min): `controller swap requested` = **1**, agent-driven bootstrap restarts = **1**, `new controller healthy` = **1**, rollback/swap-failed/unhealthy = **0**, container `RestartCount` = **0**. `VerifyStartup` banked it on the next boot (`Post-update startup: update successful (0.153.0 → 0.154.0)`) and the `06:57:52` periodic check logged `Current version 0.154.0 is up to date` — the at/above-floor branch correctly doing nothing. **No storm, no rollback, no second attempt.** *Caveat, disclosed: a hand-deploy of v0.155.0 at `07:17:10` sits inside the observation window and is what StartedAt shows after that point; it never touches `SwapController`, so the swap-count assertions above are uncontaminated across the full window. A second, unplanned confirmation of the at/above-floor branch came with it — after the hand-deploy the box ran 0.155.0 against a 0.154.0 floor and the updater logged `Current version 0.155.0 is up to date` and did nothing.* Evidence: `felhom-controller/REPORT.md` §6 (2026-07-21). *(Superseded note:)* **the self-restart single-fire leg was** - the floor was set to a version the box ALREADY ran, so there was no work and no restart. Finish by bumping the floor to a version the box does NOT run, debug ring open, asserting EXACTLY ONE restart. **Trap found while banking this: the wake is `logx.Debugf`, so it is INVISIBLE in `docker logs` at INFO** and lives only in the debug ring (`GET /api/debug/logs?level=DEBUG`) - a hunter looking at stdout wrongly concludes the box never woke; arguably (b) generalised. (b) cosmetic: the Waiter's "recovered" INFO logs on the next hold completion (`pollOnce` blocks ~240 s), not at reconnect | S | **(a) BANKED in full; only (b) cosmetic remains** | Map row "config/state change round-trips in seconds" flipped PARTIAL->PROVEN-LIVE 2026-07-21 on this evidence. Evidence: `felhom-controller/REPORT.md` 4f |