Files
felhom.eu/documentation/audits/DRILL-r201-offsite-recovery-2026-08-04.md

312 lines
17 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# DRILL RECORD — R-201, the wipe-and-recover proof: **PREPARED, HALTED BEFORE THE WIPE**
**Date:** 2026-08-04 · **Box:** `demo-hp` (Tier 0, the designated drill host) · **Nothing was wiped.**
**Outcome:** the drill did not reach its verdict. It was halted at step 4 by a defect that makes the
verdict unobtainable — and that defect is worth more than the drill.
> **THE HEADLINE.** A customer-declared **mandatory** data directory was **silently absent from the
> off-site snapshot**, while the backup reported `ok` with three snapshots. The only trace is one
> `[WARN]` line inside the controller container. The hub, the card and the counters all say the backup
> succeeded. → **R-203**
>
> **Nothing irreversible happened.** The wipe never ran. `demo-hp` is left in a BETTER state than it
> started: it now has a working off-site repository (it had an orphaned one), three snapshots, and a
> sentinel file on disk.
---
## 1. The verdict — not reached, and why that is the correct outcome
The drill's pass condition is a byte-identical sentinel sha256 after a wipe. **Step 4 established that
the sentinel is not in the off-site snapshot at all.** Wiping the box would therefore have:
- destroyed the sentinel, which exists only on that box;
- proven nothing about recovery, because there would be nothing to recover;
- and done so *after* the point of no return.
The runbook's own rule applies: *"If this session finds itself writing code beyond Part 0, stop. That
means a precondition was wrong, and the finding outranks the drill."* A precondition was wrong. It was
one the runbook's P1P6 table did not contain, because nobody knew to look for it.
**Sentinel sha256 (step 3), recorded and still on the box:**
`643166269103a25cf41d34a26b75fd6ebba0a837bbcd7e0d2b5aba7106bcbe7c`
at `/mnt/sys_drive/userdata/media/books/DRILL-SENTINEL.txt` (181 B).
---
## 2. R-203 — the defect that halted the drill
**Measured, twice, on the live box.**
| what | path | exists? |
|---|---|---|
| the app's live bind, where the customer's files actually land | `/mnt/sys_drive/userdata/media/books` | **YES** (the sentinel is here) |
| the path the off-site capture set treats as the mandatory directory | `/mnt/sys_drive/felhom-data/userdata/media/books` | **NO** |
The controller's own log, verbatim:
```
[WARN] [offbox] calibre-web: mandatory data path missing on disk, skipped from offsite:
/mnt/sys_drive/felhom-data/userdata/media/books
[INFO] [offbox] backed up calibre-web (…/backups/primary/calibre-web, 0 mandatory path(s))
[INFO] [offbox] backup OK: 3 app(s) backed up, 3 snapshot(s), 37s
```
**The run reported `ok`.** `last_status: ok`, `last_success` stamped, `snapshot_count: 3`. Nothing
customer-visible, nothing hub-visible and nothing in any counter says the mandatory directory was
dropped. This is `CLAUDE.md`'s recurring shape — *a path the customer thinks is protected is not in the
snapshot* — and the code even has the right words for it in a WARN nobody reads.
**The mechanism, from source.** `NamespaceRoot(drivePath, inGuestDrive)`
(`appbackup/paths.go:28-33`) appends `felhom-data` **when the drive IS the system data path**
(`m.namespaceRoot` = `NamespaceRoot(drivePath, drivePath != m.systemDataPath)`,
`backup/backup.go:331`). The deploy-time bind does not: `${USERDATA_PATH}` is documented and computed
as **`<HDD_PATH>/userdata`** (`stacks/classify_binds.go:14`). With `system_data_path: /mnt/sys_drive`
and an app deployed at `HDD_PATH=/mnt/sys_drive`, those two produce different directories.
**And the same compose file used BOTH roots.** From `docker inspect calibre-web`:
```
bind /mnt/sys_drive/felhom-data/userdata/import/calibre -> /cwa-book-ingest ← felhom-data root
bind /mnt/sys_drive/userdata/media/books -> /calibre-library ← NO felhom-data root
```
`${IMPORT_PATH}` resolved *with* the segment; `${USERDATA_PATH}` resolved *without* it. One deploy, one
template, two roots.
### What is measured and what is not — stated because the scope changes the fix
- **MEASURED:** the two paths disagree; the mandatory directory is absent from the snapshot; the run
reports `ok`; the only signal is a container-log WARN.
- **NOT ESTABLISHED:** whether `HDD_PATH=/mnt/sys_drive` is a *supported* choice. It was chosen because
demo-hp's only registered drive, `Felhom-Share`, is a NAS and was **correctly refused** as an app
namespace (R-108's `RefuseAsAppNamespace`, working as designed). `/mnt/sys_drive` was **accepted**
(HTTP 202) rather than refused.
**Either branch is a defect, which is why this is filed regardless:**
- if the system drive **is** a supported app namespace → the userdata path resolution is wrong for
every app deployed on it, and their mandatory directories are silently unprotected;
- if it is **not** supported → the deploy accepted a namespace it should have refused, exactly as it
refused the NAS one call earlier, and the refusal that exists is not reaching this case.
**What must NOT be concluded from this drill:** that off-site backups are broken generally. The two
pre-existing apps (`opengist`, `privatebin`) declare **no** mandatory userdata paths — everything they
own is in named volumes — so they are unaffected, and their snapshots are real.
---
## 3. Preconditions, each measured
| # | Precondition | Result |
|---|---|---|
| **P1** | operator holds the recovery code | **PASS**`R_DEMO-HP`, recorded by the operator in the DooPlex credentials file (`~/.config/credentials`). *The code itself appears nowhere in this record.* |
| **P2** | demo-hp on agent v0.125.0 + controller v0.196.0 | **FAILED ON ARRIVAL, FIXED** — the box was on agent **v0.124.1**. Deployed the published v0.125.0 (sha `f7d8339b53d9…`, verified against the release) and controller v0.196.0. `age` present at `/usr/bin/age`. |
| **P3** | the current `identity_blob` seals the repository under test | **PASS with a caveat that reshaped the drill**`identity_blob` = 572 B, `restic_pw_sha256` = `8a9e33aa4da6…`. But **no repository existed under that key** (see §4). |
| **P4** | a deliberate, verified whole-guest archive as rollback | **NOT TAKEN** — deliberately. It is only needed for the wipe, and the wipe did not happen. |
| **P5** | demo-felhom untouched and healthy | **PASS** — not touched at any point. |
| **P6** | space | **PASS**`felhom-backup` 927 GB free, `local-lvm` 30.9 % used, guest 64 GB free. |
---
## 4. Step-by-step, with every observable
### Step 1 — starting state (hub, read-only)
`demo-hp-bb76ea`: `identity_blob` 572 B, `restic_pw_sha256` `8a9e33aa4da6769c5aea1831f87759e10930e2ec1dea0062576484e0598d080a`, created `2026-08-04T11:11:37Z`, `escrow_state: escrowed`.
Report: `snapshot_count: 0`, `repo_size_bytes: 0`, **no `last_run`/`last_status` at all** — the shape of
a controller that has never run an off-site backup in this lifetime.
### Step 1b — the repository was ORPHANED, and this is the first live proof of a prediction
The 2026-08-04 spike predicted from source — and could not measure — that a rebuilt box's next
off-site run would hit a **third** outcome: neither reattaching the old snapshots nor silently starting
a fresh history, but **refusing**. Measured here, twice over.
**Read-only probe first** (`restic cat config` with the current key, writes nothing):
```
Fatal: wrong password or no key found
```
— the exact string `classifyResticProbe` maps to `"orphaned"`.
**Then the real run, through the customer UI endpoint** (`POST /backup/offbox/run`):
```
[WARN] [offbox] offsite repo ORPHANED — remote holds backups written under a previous,
no-longer-available key; runs will skip until reset
[WARN] [offbox] run skipped — offsite repo orphaned (card shown; awaiting reset)
[INFO] Event pushed: offbox_repo_orphaned (warning) — A távoli mentési tároló elárvult …
```
`repo_state: "orphaned"`, `orphaned_at: 2026-08-04T12:37:30Z`, `last_status: "error"`, the orphan card
rendered on `/backups/remote`, and the event reached the hub (HTTP 200).
> **The system stopped and said so. It did not quietly start a new history over the old one.**
> That closes R-193's open Q3 — and it is the good half of this month's story.
**Why the repository was orphaned:** the 15 snapshots / 40.9 MB at `/home/felhom-repo` were written
under key `8e03eddf9ff7…` before the 2026-08-03 rebuild. That key survives only in superseded escrow
row **id 3**, which carries **`identity_blob` = NULL** — it was superseded at `2026-08-04 07:15:36`,
**four hours before** hub v0.93.0 fixed the retention. Unrecoverable, permanently, with or without a
recovery code.
### Step 1c — the reset (operator-authorised)
The orphan card's own reset, confirmed by the operator during the session:
```
[WARN] [offbox] resetting orphaned repo (operator-confirmed (claimed)):
move-aside /home/felhom-repo -> /home/felhom-repo.orphaned-20260804, then re-init
[INFO] [offbox] orphaned repo reset complete — old history set aside at
/home/felhom-repo.orphaned-20260804 (move-aside, not deleted); fresh repo initialized
[INFO] Event pushed: offbox_repo_reset (info)
```
**Nothing was deleted.** The reset path had never run in anger before; it works.
### Steps 23 — the recovery code and the sentinel
The operator ran the ceremonies for **both** boxes earlier the same day and saved the codes. Measured
hub-side: the new escrow rows are stamped `2026-08-04T11:11:37Z` (demo-hp) and `11:13:06Z`
(demo-felhom), and `restic_pw_sha256` is **unchanged** on both — so `SaveHostEscrow` correctly treated
them as same-password re-ceremonies: **no superseded row, no `offsite_repo_key_changed`**. That is
v0.93.0's Scenario E, live.
**A file-leg app had to be deployed**, and this is a precondition the runbook did not anticipate:
neither off-site-toggled app on the box (`opengist`, `privatebin`) has a restorable file leg — both
keep everything in **named volumes**, which the off-site tier backs up as tars but the customer restore
flow **never unpacks** (`offbox_reconstitute.go`). A sentinel in either would have been unrecoverable
by design.
`calibre-web` was chosen: it declares `userdata: media/books class: mandatory`, and it is a
single-container app. Deployed through the real API (`POST /api/stacks/calibre-web/deploy`, HTTP 202),
toggled for off-site, and a Tier-1 recovery unit captured
(`Recovery unit captured for calibre-web → …/backups/primary/calibre-web`).
Sentinel written, `sync`ed, hashed: **`643166269103a25c…`**, 181 B.
### Step 4 — the pre-wipe off-site backup: **`ok`, and wrong**
```
last_run 2026-08-04T12:54:50Z · last_status ok · last_success stamped
snapshot_count 3 · repo_size_bytes 30 636 · repo_state null
```
Three snapshots, three apps, 37 s — and **the sentinel is in none of them**, per §2.
### Steps 511 — NOT RUN
Step 5 (verify the pre-wipe snapshot contents) would have confirmed the absence a second way; it is
moot given §2. Steps 611 (archive, wipe, reinstall, recover, install, restore, compare) were **not
started**. The §7 STOP was never reached, because the drill failed its own precondition first.
---
## 5. Part 0 — shipped, tested, deployed (R-200's plumbing half)
`--recover-offsite-install` (controller **v0.196.0**, commit `1b1366b`): same fetch → unseal → extract
path as `--recover-offsite-check`, same STDIN discipline for R, but it **places** the recovered
password via `InjectOffboxPassword`.
- **Confirmation is a second invocation.** Without `--confirm-install` it prints both hashes and writes
nothing. A single interactive prompt would have had to share stdin with R.
- **Three outcomes, named distinctly:** *installed* (no local password — the rebuilt-box shape),
*unchanged* (identical key already present, nothing written), *refused* (a DIFFERENT key present;
installing would clobber the key the current repository is encrypted under — exit 2, no force offered).
- It re-reads the file after writing rather than trusting the call's return.
**Tests + red-proof.** `go build && go vet && go test ./...` rc=0; `controller_gates.py --fast` OK.
Removing the confirmation gate makes the dry run write the password and fails
`TestRecoverAndInstall_InstallsOnABareBox` with *"the DRY RUN wrote the password"*. The R-persistence
test carries a **positive control** — a planted copy of the code is found by the sweep, then removed
and not found — because an absence check is worth only what its sensitivity is.
**It was deployed to demo-hp and never exercised against a live recovery**, because the drill halted
before step 9. Its unit proof stands; its live proof does not exist.
---
## 6. What this drill did and did not establish
**Established, live, for the first time:**
1. A rebuilt box's off-site repository is **orphaned and the run refuses** — the spike's predicted
third outcome, measured. It does not silently start a fresh history.
2. The **orphan reset works**: move-aside, never delete, fresh repo initialised, event pushed.
3. **A customer-declared mandatory data directory can be silently absent from the off-site snapshot
while the run reports `ok`** (R-203).
4. demo-hp's pre-rebuild off-site history is **permanently unrecoverable** — its key was destroyed four
hours before the fix that would have kept it.
**NOT established — and unchanged from before this session:**
- **No file has ever been restored from an off-site backup after a wipe.** R-201's question is still
open, and its pass condition is unchanged.
- Part 0's install path has never run against a live recovery.
- The v0.93.0 `identity_blob` retention is **still unit-proven only** — nothing in this session
superseded a key, so nothing exercised it.
---
## 7. State left behind, and teardown
**Deliberately not torn down** — this is evidence, and the box is better off than it was:
| layer | state |
|---|---|
| the guest | `calibre-web` deployed and running, off-site-toggled, sentinel file in place. **Kept** as the drill fixture for the resumed run — it is the only app on either demo box with a restorable file leg. |
| the off-site repository | fresh, working, 3 snapshots, 30 636 B. The old 15-snapshot history is **set aside**, not deleted, at `/home/felhom-repo.orphaned-20260804`. |
| the host | agent v0.125.0, controller v0.196.0, `pvesm` unchanged apart from normal usage. |
| the hub | `offbox_repo_orphaned` + `offbox_repo_reset` events recorded for `demo-hp`. No scratch customer records were created — **nothing was reinstalled**. |
**The ~1.2 GB of previously-orphaned ciphertext was NOT deleted** — that act is ruled and still owed,
and §8.3 of the runbook forbids riding it along with a drill. The reset added
`/home/felhom-repo.orphaned-20260804` (≈41 MB) to what is set aside.
**No `R` persisted anywhere** — the recovery code was never used in this session. It was read only to
confirm the key exists in the credentials file; no unseal was performed on demo-hp.
---
## 7b. UPDATE, same day — the blocker is FIXED and the drill is ready to resume
Controller **v0.197.0** shipped both halves of R-203:
- **the paths agree.** On demo-hp the live bind moved from `/mnt/sys_drive/userdata/media/books` to
`/mnt/sys_drive/felhom-data/userdata/media/books` — the directory the capture set looks in. The
capture log went from `0 mandatory path(s)` to **`1 mandatory path(s)`**.
- **`ok` means it.** A run that cannot capture a MANDATORY directory now reports **`incomplete`**,
names the app and the folders, and raises the operator digest — instead of `ok` with a warning
beside it.
**And the sentinel is in the snapshot, listed by name:**
```
$ restic ls -l latest --tag calibre-web
-rw-r--r-- 1000 1000 181 2026-08-04 12:53:06 /mnt/sys_drive/felhom-data/userdata/media/books/DRILL-SENTINEL.txt
```
sha256 **`643166269103a25cf41d34a26b75fd6ebba0a837bbcd7e0d2b5aba7106bcbe7c`** — byte-identical to §1,
verified after the fix's migration moved the file to the corrected directory.
**One live check could NOT be reproduced, and is recorded rather than claimed.** Hiding the books
directory to watch the `incomplete` verdict fire on hardware did not work: the running container's
bind mount **recreated** the directory, so `os.Stat` succeeded and there was no gap. That is itself
worth knowing — a bind-mounted directory cannot easily be "missing" while its app runs, so the
mandatory-gap condition arises in practice when the path resolves somewhere the app never binds (the
R-203 case), not when a live app's own directory vanishes. The verdict is proven by the run-level
test, which drives the real `RunOffboxBackup` and asserts both `incomplete` and the operator signal.
The fixture was restored and the sentinel re-verified at the same hash.
## 8. To resume the drill
1. **Fix or scope R-203.** Until the mandatory userdata path lands in the snapshot, no sentinel can
survive a wipe and the drill cannot reach its verdict.
2. Re-run steps 45 and confirm the sentinel IS in the snapshot — by listing it, not by a green status.
3. Then P4 (the deliberate archive), the §7 STOP, and steps 611 as written.
Everything else is already in place: the code, the versions, the recovery code, the working repository,
the file-leg app and the sentinel.