# DRILL RECORD — R-201, the wipe-and-recover proof: **PREPARED, HALTED BEFORE THE WIPE** **Date:** 2026-08-04 · **Box:** `demo-hp` (Tier 0, the designated drill host) · **Nothing was wiped.** **Outcome:** the drill did not reach its verdict. It was halted at step 4 by a defect that makes the verdict unobtainable — and that defect is worth more than the drill. > **THE HEADLINE.** A customer-declared **mandatory** data directory was **silently absent from the > off-site snapshot**, while the backup reported `ok` with three snapshots. The only trace is one > `[WARN]` line inside the controller container. The hub, the card and the counters all say the backup > succeeded. → **R-203** > > **Nothing irreversible happened.** The wipe never ran. `demo-hp` is left in a BETTER state than it > started: it now has a working off-site repository (it had an orphaned one), three snapshots, and a > sentinel file on disk. --- ## 1. The verdict — not reached, and why that is the correct outcome The drill's pass condition is a byte-identical sentinel sha256 after a wipe. **Step 4 established that the sentinel is not in the off-site snapshot at all.** Wiping the box would therefore have: - destroyed the sentinel, which exists only on that box; - proven nothing about recovery, because there would be nothing to recover; - and done so *after* the point of no return. The runbook's own rule applies: *"If this session finds itself writing code beyond Part 0, stop. That means a precondition was wrong, and the finding outranks the drill."* A precondition was wrong. It was one the runbook's P1–P6 table did not contain, because nobody knew to look for it. **Sentinel sha256 (step 3), recorded and still on the box:** `643166269103a25cf41d34a26b75fd6ebba0a837bbcd7e0d2b5aba7106bcbe7c` at `/mnt/sys_drive/userdata/media/books/DRILL-SENTINEL.txt` (181 B). --- ## 2. R-203 — the defect that halted the drill **Measured, twice, on the live box.** | what | path | exists? | |---|---|---| | the app's live bind, where the customer's files actually land | `/mnt/sys_drive/userdata/media/books` | **YES** (the sentinel is here) | | the path the off-site capture set treats as the mandatory directory | `/mnt/sys_drive/felhom-data/userdata/media/books` | **NO** | The controller's own log, verbatim: ``` [WARN] [offbox] calibre-web: mandatory data path missing on disk, skipped from offsite: /mnt/sys_drive/felhom-data/userdata/media/books [INFO] [offbox] backed up calibre-web (…/backups/primary/calibre-web, 0 mandatory path(s)) [INFO] [offbox] backup OK: 3 app(s) backed up, 3 snapshot(s), 37s ``` **The run reported `ok`.** `last_status: ok`, `last_success` stamped, `snapshot_count: 3`. Nothing customer-visible, nothing hub-visible and nothing in any counter says the mandatory directory was dropped. This is `CLAUDE.md`'s recurring shape — *a path the customer thinks is protected is not in the snapshot* — and the code even has the right words for it in a WARN nobody reads. **The mechanism, from source.** `NamespaceRoot(drivePath, inGuestDrive)` (`appbackup/paths.go:28-33`) appends `felhom-data` **when the drive IS the system data path** (`m.namespaceRoot` = `NamespaceRoot(drivePath, drivePath != m.systemDataPath)`, `backup/backup.go:331`). The deploy-time bind does not: `${USERDATA_PATH}` is documented and computed as **`/userdata`** (`stacks/classify_binds.go:14`). With `system_data_path: /mnt/sys_drive` and an app deployed at `HDD_PATH=/mnt/sys_drive`, those two produce different directories. **And the same compose file used BOTH roots.** From `docker inspect calibre-web`: ``` bind /mnt/sys_drive/felhom-data/userdata/import/calibre -> /cwa-book-ingest ← felhom-data root bind /mnt/sys_drive/userdata/media/books -> /calibre-library ← NO felhom-data root ``` `${IMPORT_PATH}` resolved *with* the segment; `${USERDATA_PATH}` resolved *without* it. One deploy, one template, two roots. ### What is measured and what is not — stated because the scope changes the fix - **MEASURED:** the two paths disagree; the mandatory directory is absent from the snapshot; the run reports `ok`; the only signal is a container-log WARN. - **NOT ESTABLISHED:** whether `HDD_PATH=/mnt/sys_drive` is a *supported* choice. It was chosen because demo-hp's only registered drive, `Felhom-Share`, is a NAS and was **correctly refused** as an app namespace (R-108's `RefuseAsAppNamespace`, working as designed). `/mnt/sys_drive` was **accepted** (HTTP 202) rather than refused. **Either branch is a defect, which is why this is filed regardless:** - if the system drive **is** a supported app namespace → the userdata path resolution is wrong for every app deployed on it, and their mandatory directories are silently unprotected; - if it is **not** supported → the deploy accepted a namespace it should have refused, exactly as it refused the NAS one call earlier, and the refusal that exists is not reaching this case. **What must NOT be concluded from this drill:** that off-site backups are broken generally. The two pre-existing apps (`opengist`, `privatebin`) declare **no** mandatory userdata paths — everything they own is in named volumes — so they are unaffected, and their snapshots are real. --- ## 3. Preconditions, each measured | # | Precondition | Result | |---|---|---| | **P1** | operator holds the recovery code | **PASS** — `R_DEMO-HP`, recorded by the operator in the DooPlex credentials file (`~/.config/credentials`). *The code itself appears nowhere in this record.* | | **P2** | demo-hp on agent v0.125.0 + controller v0.196.0 | **FAILED ON ARRIVAL, FIXED** — the box was on agent **v0.124.1**. Deployed the published v0.125.0 (sha `f7d8339b53d9…`, verified against the release) and controller v0.196.0. `age` present at `/usr/bin/age`. | | **P3** | the current `identity_blob` seals the repository under test | **PASS with a caveat that reshaped the drill** — `identity_blob` = 572 B, `restic_pw_sha256` = `8a9e33aa4da6…`. But **no repository existed under that key** (see §4). | | **P4** | a deliberate, verified whole-guest archive as rollback | **NOT TAKEN** — deliberately. It is only needed for the wipe, and the wipe did not happen. | | **P5** | demo-felhom untouched and healthy | **PASS** — not touched at any point. | | **P6** | space | **PASS** — `felhom-backup` 927 GB free, `local-lvm` 30.9 % used, guest 64 GB free. | --- ## 4. Step-by-step, with every observable ### Step 1 — starting state (hub, read-only) `demo-hp-bb76ea`: `identity_blob` 572 B, `restic_pw_sha256` `8a9e33aa4da6769c5aea1831f87759e10930e2ec1dea0062576484e0598d080a`, created `2026-08-04T11:11:37Z`, `escrow_state: escrowed`. Report: `snapshot_count: 0`, `repo_size_bytes: 0`, **no `last_run`/`last_status` at all** — the shape of a controller that has never run an off-site backup in this lifetime. ### Step 1b — the repository was ORPHANED, and this is the first live proof of a prediction The 2026-08-04 spike predicted from source — and could not measure — that a rebuilt box's next off-site run would hit a **third** outcome: neither reattaching the old snapshots nor silently starting a fresh history, but **refusing**. Measured here, twice over. **Read-only probe first** (`restic cat config` with the current key, writes nothing): ``` Fatal: wrong password or no key found ``` — the exact string `classifyResticProbe` maps to `"orphaned"`. **Then the real run, through the customer UI endpoint** (`POST /backup/offbox/run`): ``` [WARN] [offbox] offsite repo ORPHANED — remote holds backups written under a previous, no-longer-available key; runs will skip until reset [WARN] [offbox] run skipped — offsite repo orphaned (card shown; awaiting reset) [INFO] Event pushed: offbox_repo_orphaned (warning) — A távoli mentési tároló elárvult … ``` `repo_state: "orphaned"`, `orphaned_at: 2026-08-04T12:37:30Z`, `last_status: "error"`, the orphan card rendered on `/backups/remote`, and the event reached the hub (HTTP 200). > **The system stopped and said so. It did not quietly start a new history over the old one.** > That closes R-193's open Q3 — and it is the good half of this month's story. **Why the repository was orphaned:** the 15 snapshots / 40.9 MB at `/home/felhom-repo` were written under key `8e03eddf9ff7…` before the 2026-08-03 rebuild. That key survives only in superseded escrow row **id 3**, which carries **`identity_blob` = NULL** — it was superseded at `2026-08-04 07:15:36`, **four hours before** hub v0.93.0 fixed the retention. Unrecoverable, permanently, with or without a recovery code. ### Step 1c — the reset (operator-authorised) The orphan card's own reset, confirmed by the operator during the session: ``` [WARN] [offbox] resetting orphaned repo (operator-confirmed (claimed)): move-aside /home/felhom-repo -> /home/felhom-repo.orphaned-20260804, then re-init [INFO] [offbox] orphaned repo reset complete — old history set aside at /home/felhom-repo.orphaned-20260804 (move-aside, not deleted); fresh repo initialized [INFO] Event pushed: offbox_repo_reset (info) ``` **Nothing was deleted.** The reset path had never run in anger before; it works. ### Steps 2–3 — the recovery code and the sentinel The operator ran the ceremonies for **both** boxes earlier the same day and saved the codes. Measured hub-side: the new escrow rows are stamped `2026-08-04T11:11:37Z` (demo-hp) and `11:13:06Z` (demo-felhom), and `restic_pw_sha256` is **unchanged** on both — so `SaveHostEscrow` correctly treated them as same-password re-ceremonies: **no superseded row, no `offsite_repo_key_changed`**. That is v0.93.0's Scenario E, live. **A file-leg app had to be deployed**, and this is a precondition the runbook did not anticipate: neither off-site-toggled app on the box (`opengist`, `privatebin`) has a restorable file leg — both keep everything in **named volumes**, which the off-site tier backs up as tars but the customer restore flow **never unpacks** (`offbox_reconstitute.go`). A sentinel in either would have been unrecoverable by design. `calibre-web` was chosen: it declares `userdata: media/books class: mandatory`, and it is a single-container app. Deployed through the real API (`POST /api/stacks/calibre-web/deploy`, HTTP 202), toggled for off-site, and a Tier-1 recovery unit captured (`Recovery unit captured for calibre-web → …/backups/primary/calibre-web`). Sentinel written, `sync`ed, hashed: **`643166269103a25c…`**, 181 B. ### Step 4 — the pre-wipe off-site backup: **`ok`, and wrong** ``` last_run 2026-08-04T12:54:50Z · last_status ok · last_success stamped snapshot_count 3 · repo_size_bytes 30 636 · repo_state null ``` Three snapshots, three apps, 37 s — and **the sentinel is in none of them**, per §2. ### Steps 5–11 — NOT RUN Step 5 (verify the pre-wipe snapshot contents) would have confirmed the absence a second way; it is moot given §2. Steps 6–11 (archive, wipe, reinstall, recover, install, restore, compare) were **not started**. The §7 STOP was never reached, because the drill failed its own precondition first. --- ## 5. Part 0 — shipped, tested, deployed (R-200's plumbing half) `--recover-offsite-install` (controller **v0.196.0**, commit `1b1366b`): same fetch → unseal → extract path as `--recover-offsite-check`, same STDIN discipline for R, but it **places** the recovered password via `InjectOffboxPassword`. - **Confirmation is a second invocation.** Without `--confirm-install` it prints both hashes and writes nothing. A single interactive prompt would have had to share stdin with R. - **Three outcomes, named distinctly:** *installed* (no local password — the rebuilt-box shape), *unchanged* (identical key already present, nothing written), *refused* (a DIFFERENT key present; installing would clobber the key the current repository is encrypted under — exit 2, no force offered). - It re-reads the file after writing rather than trusting the call's return. **Tests + red-proof.** `go build && go vet && go test ./...` rc=0; `controller_gates.py --fast` OK. Removing the confirmation gate makes the dry run write the password and fails `TestRecoverAndInstall_InstallsOnABareBox` with *"the DRY RUN wrote the password"*. The R-persistence test carries a **positive control** — a planted copy of the code is found by the sweep, then removed and not found — because an absence check is worth only what its sensitivity is. **It was deployed to demo-hp and never exercised against a live recovery**, because the drill halted before step 9. Its unit proof stands; its live proof does not exist. --- ## 6. What this drill did and did not establish **Established, live, for the first time:** 1. A rebuilt box's off-site repository is **orphaned and the run refuses** — the spike's predicted third outcome, measured. It does not silently start a fresh history. 2. The **orphan reset works**: move-aside, never delete, fresh repo initialised, event pushed. 3. **A customer-declared mandatory data directory can be silently absent from the off-site snapshot while the run reports `ok`** (R-203). 4. demo-hp's pre-rebuild off-site history is **permanently unrecoverable** — its key was destroyed four hours before the fix that would have kept it. **NOT established — and unchanged from before this session:** - **No file has ever been restored from an off-site backup after a wipe.** R-201's question is still open, and its pass condition is unchanged. - Part 0's install path has never run against a live recovery. - The v0.93.0 `identity_blob` retention is **still unit-proven only** — nothing in this session superseded a key, so nothing exercised it. --- ## 7. State left behind, and teardown **Deliberately not torn down** — this is evidence, and the box is better off than it was: | layer | state | |---|---| | the guest | `calibre-web` deployed and running, off-site-toggled, sentinel file in place. **Kept** as the drill fixture for the resumed run — it is the only app on either demo box with a restorable file leg. | | the off-site repository | fresh, working, 3 snapshots, 30 636 B. The old 15-snapshot history is **set aside**, not deleted, at `/home/felhom-repo.orphaned-20260804`. | | the host | agent v0.125.0, controller v0.196.0, `pvesm` unchanged apart from normal usage. | | the hub | `offbox_repo_orphaned` + `offbox_repo_reset` events recorded for `demo-hp`. No scratch customer records were created — **nothing was reinstalled**. | **The ~1.2 GB of previously-orphaned ciphertext was NOT deleted** — that act is ruled and still owed, and §8.3 of the runbook forbids riding it along with a drill. The reset added `/home/felhom-repo.orphaned-20260804` (≈41 MB) to what is set aside. **No `R` persisted anywhere** — the recovery code was never used in this session. It was read only to confirm the key exists in the credentials file; no unseal was performed on demo-hp. --- ## 7b. UPDATE, same day — the blocker is FIXED and the drill is ready to resume Controller **v0.197.0** shipped both halves of R-203: - **the paths agree.** On demo-hp the live bind moved from `/mnt/sys_drive/userdata/media/books` to `/mnt/sys_drive/felhom-data/userdata/media/books` — the directory the capture set looks in. The capture log went from `0 mandatory path(s)` to **`1 mandatory path(s)`**. - **`ok` means it.** A run that cannot capture a MANDATORY directory now reports **`incomplete`**, names the app and the folders, and raises the operator digest — instead of `ok` with a warning beside it. **And the sentinel is in the snapshot, listed by name:** ``` $ restic ls -l latest --tag calibre-web -rw-r--r-- 1000 1000 181 2026-08-04 12:53:06 /mnt/sys_drive/felhom-data/userdata/media/books/DRILL-SENTINEL.txt ``` sha256 **`643166269103a25cf41d34a26b75fd6ebba0a837bbcd7e0d2b5aba7106bcbe7c`** — byte-identical to §1, verified after the fix's migration moved the file to the corrected directory. **One live check could NOT be reproduced, and is recorded rather than claimed.** Hiding the books directory to watch the `incomplete` verdict fire on hardware did not work: the running container's bind mount **recreated** the directory, so `os.Stat` succeeded and there was no gap. That is itself worth knowing — a bind-mounted directory cannot easily be "missing" while its app runs, so the mandatory-gap condition arises in practice when the path resolves somewhere the app never binds (the R-203 case), not when a live app's own directory vanishes. The verdict is proven by the run-level test, which drives the real `RunOffboxBackup` and asserts both `incomplete` and the operator signal. The fixture was restored and the sentinel re-verified at the same hash. ## 8. To resume the drill 1. **Fix or scope R-203.** Until the mandatory userdata path lands in the snapshot, no sentinel can survive a wipe and the drill cannot reach its verdict. 2. Re-run steps 4–5 and confirm the sentinel IS in the snapshot — by listing it, not by a green status. 3. Then P4 (the deliberate archive), the §7 STOP, and steps 6–11 as written. Everything else is already in place: the code, the versions, the recovery code, the working repository, the file-leg app and the sentinel.