17 KiB
DRILL RECORD — R-201, the wipe-and-recover proof: PREPARED, HALTED BEFORE THE WIPE
Date: 2026-08-04 · Box: demo-hp (Tier 0, the designated drill host) · Nothing was wiped.
Outcome: the drill did not reach its verdict. It was halted at step 4 by a defect that makes the
verdict unobtainable — and that defect is worth more than the drill.
THE HEADLINE. A customer-declared mandatory data directory was silently absent from the off-site snapshot, while the backup reported
okwith three snapshots. The only trace is one[WARN]line inside the controller container. The hub, the card and the counters all say the backup succeeded. → R-203Nothing irreversible happened. The wipe never ran.
demo-hpis left in a BETTER state than it started: it now has a working off-site repository (it had an orphaned one), three snapshots, and a sentinel file on disk.
1. The verdict — not reached, and why that is the correct outcome
The drill's pass condition is a byte-identical sentinel sha256 after a wipe. Step 4 established that the sentinel is not in the off-site snapshot at all. Wiping the box would therefore have:
- destroyed the sentinel, which exists only on that box;
- proven nothing about recovery, because there would be nothing to recover;
- and done so after the point of no return.
The runbook's own rule applies: "If this session finds itself writing code beyond Part 0, stop. That means a precondition was wrong, and the finding outranks the drill." A precondition was wrong. It was one the runbook's P1–P6 table did not contain, because nobody knew to look for it.
Sentinel sha256 (step 3), recorded and still on the box:
643166269103a25cf41d34a26b75fd6ebba0a837bbcd7e0d2b5aba7106bcbe7c
at /mnt/sys_drive/userdata/media/books/DRILL-SENTINEL.txt (181 B).
2. R-203 — the defect that halted the drill
Measured, twice, on the live box.
| what | path | exists? |
|---|---|---|
| the app's live bind, where the customer's files actually land | /mnt/sys_drive/userdata/media/books |
YES (the sentinel is here) |
| the path the off-site capture set treats as the mandatory directory | /mnt/sys_drive/felhom-data/userdata/media/books |
NO |
The controller's own log, verbatim:
[WARN] [offbox] calibre-web: mandatory data path missing on disk, skipped from offsite:
/mnt/sys_drive/felhom-data/userdata/media/books
[INFO] [offbox] backed up calibre-web (…/backups/primary/calibre-web, 0 mandatory path(s))
[INFO] [offbox] backup OK: 3 app(s) backed up, 3 snapshot(s), 37s
The run reported ok. last_status: ok, last_success stamped, snapshot_count: 3. Nothing
customer-visible, nothing hub-visible and nothing in any counter says the mandatory directory was
dropped. This is CLAUDE.md's recurring shape — a path the customer thinks is protected is not in the
snapshot — and the code even has the right words for it in a WARN nobody reads.
The mechanism, from source. NamespaceRoot(drivePath, inGuestDrive)
(appbackup/paths.go:28-33) appends felhom-data when the drive IS the system data path
(m.namespaceRoot = NamespaceRoot(drivePath, drivePath != m.systemDataPath),
backup/backup.go:331). The deploy-time bind does not: ${USERDATA_PATH} is documented and computed
as <HDD_PATH>/userdata (stacks/classify_binds.go:14). With system_data_path: /mnt/sys_drive
and an app deployed at HDD_PATH=/mnt/sys_drive, those two produce different directories.
And the same compose file used BOTH roots. From docker inspect calibre-web:
bind /mnt/sys_drive/felhom-data/userdata/import/calibre -> /cwa-book-ingest ← felhom-data root
bind /mnt/sys_drive/userdata/media/books -> /calibre-library ← NO felhom-data root
${IMPORT_PATH} resolved with the segment; ${USERDATA_PATH} resolved without it. One deploy, one
template, two roots.
What is measured and what is not — stated because the scope changes the fix
- MEASURED: the two paths disagree; the mandatory directory is absent from the snapshot; the run
reports
ok; the only signal is a container-log WARN. - NOT ESTABLISHED: whether
HDD_PATH=/mnt/sys_driveis a supported choice. It was chosen because demo-hp's only registered drive,Felhom-Share, is a NAS and was correctly refused as an app namespace (R-108'sRefuseAsAppNamespace, working as designed)./mnt/sys_drivewas accepted (HTTP 202) rather than refused.
Either branch is a defect, which is why this is filed regardless:
- if the system drive is a supported app namespace → the userdata path resolution is wrong for every app deployed on it, and their mandatory directories are silently unprotected;
- if it is not supported → the deploy accepted a namespace it should have refused, exactly as it refused the NAS one call earlier, and the refusal that exists is not reaching this case.
What must NOT be concluded from this drill: that off-site backups are broken generally. The two
pre-existing apps (opengist, privatebin) declare no mandatory userdata paths — everything they
own is in named volumes — so they are unaffected, and their snapshots are real.
3. Preconditions, each measured
| # | Precondition | Result |
|---|---|---|
| P1 | operator holds the recovery code | PASS — R_DEMO-HP, recorded by the operator in the DooPlex credentials file (~/.config/credentials). The code itself appears nowhere in this record. |
| P2 | demo-hp on agent v0.125.0 + controller v0.196.0 | FAILED ON ARRIVAL, FIXED — the box was on agent v0.124.1. Deployed the published v0.125.0 (sha f7d8339b53d9…, verified against the release) and controller v0.196.0. age present at /usr/bin/age. |
| P3 | the current identity_blob seals the repository under test |
PASS with a caveat that reshaped the drill — identity_blob = 572 B, restic_pw_sha256 = 8a9e33aa4da6…. But no repository existed under that key (see §4). |
| P4 | a deliberate, verified whole-guest archive as rollback | NOT TAKEN — deliberately. It is only needed for the wipe, and the wipe did not happen. |
| P5 | demo-felhom untouched and healthy | PASS — not touched at any point. |
| P6 | space | PASS — felhom-backup 927 GB free, local-lvm 30.9 % used, guest 64 GB free. |
4. Step-by-step, with every observable
Step 1 — starting state (hub, read-only)
demo-hp-bb76ea: identity_blob 572 B, restic_pw_sha256 8a9e33aa4da6769c5aea1831f87759e10930e2ec1dea0062576484e0598d080a, created 2026-08-04T11:11:37Z, escrow_state: escrowed.
Report: snapshot_count: 0, repo_size_bytes: 0, no last_run/last_status at all — the shape of
a controller that has never run an off-site backup in this lifetime.
Step 1b — the repository was ORPHANED, and this is the first live proof of a prediction
The 2026-08-04 spike predicted from source — and could not measure — that a rebuilt box's next off-site run would hit a third outcome: neither reattaching the old snapshots nor silently starting a fresh history, but refusing. Measured here, twice over.
Read-only probe first (restic cat config with the current key, writes nothing):
Fatal: wrong password or no key found
— the exact string classifyResticProbe maps to "orphaned".
Then the real run, through the customer UI endpoint (POST /backup/offbox/run):
[WARN] [offbox] offsite repo ORPHANED — remote holds backups written under a previous,
no-longer-available key; runs will skip until reset
[WARN] [offbox] run skipped — offsite repo orphaned (card shown; awaiting reset)
[INFO] Event pushed: offbox_repo_orphaned (warning) — A távoli mentési tároló elárvult …
repo_state: "orphaned", orphaned_at: 2026-08-04T12:37:30Z, last_status: "error", the orphan card
rendered on /backups/remote, and the event reached the hub (HTTP 200).
The system stopped and said so. It did not quietly start a new history over the old one. That closes R-193's open Q3 — and it is the good half of this month's story.
Why the repository was orphaned: the 15 snapshots / 40.9 MB at /home/felhom-repo were written
under key 8e03eddf9ff7… before the 2026-08-03 rebuild. That key survives only in superseded escrow
row id 3, which carries identity_blob = NULL — it was superseded at 2026-08-04 07:15:36,
four hours before hub v0.93.0 fixed the retention. Unrecoverable, permanently, with or without a
recovery code.
Step 1c — the reset (operator-authorised)
The orphan card's own reset, confirmed by the operator during the session:
[WARN] [offbox] resetting orphaned repo (operator-confirmed (claimed)):
move-aside /home/felhom-repo -> /home/felhom-repo.orphaned-20260804, then re-init
[INFO] [offbox] orphaned repo reset complete — old history set aside at
/home/felhom-repo.orphaned-20260804 (move-aside, not deleted); fresh repo initialized
[INFO] Event pushed: offbox_repo_reset (info)
Nothing was deleted. The reset path had never run in anger before; it works.
Steps 2–3 — the recovery code and the sentinel
The operator ran the ceremonies for both boxes earlier the same day and saved the codes. Measured
hub-side: the new escrow rows are stamped 2026-08-04T11:11:37Z (demo-hp) and 11:13:06Z
(demo-felhom), and restic_pw_sha256 is unchanged on both — so SaveHostEscrow correctly treated
them as same-password re-ceremonies: no superseded row, no offsite_repo_key_changed. That is
v0.93.0's Scenario E, live.
A file-leg app had to be deployed, and this is a precondition the runbook did not anticipate:
neither off-site-toggled app on the box (opengist, privatebin) has a restorable file leg — both
keep everything in named volumes, which the off-site tier backs up as tars but the customer restore
flow never unpacks (offbox_reconstitute.go). A sentinel in either would have been unrecoverable
by design.
calibre-web was chosen: it declares userdata: media/books class: mandatory, and it is a
single-container app. Deployed through the real API (POST /api/stacks/calibre-web/deploy, HTTP 202),
toggled for off-site, and a Tier-1 recovery unit captured
(Recovery unit captured for calibre-web → …/backups/primary/calibre-web).
Sentinel written, synced, hashed: 643166269103a25c…, 181 B.
Step 4 — the pre-wipe off-site backup: ok, and wrong
last_run 2026-08-04T12:54:50Z · last_status ok · last_success stamped
snapshot_count 3 · repo_size_bytes 30 636 · repo_state null
Three snapshots, three apps, 37 s — and the sentinel is in none of them, per §2.
Steps 5–11 — NOT RUN
Step 5 (verify the pre-wipe snapshot contents) would have confirmed the absence a second way; it is moot given §2. Steps 6–11 (archive, wipe, reinstall, recover, install, restore, compare) were not started. The §7 STOP was never reached, because the drill failed its own precondition first.
5. Part 0 — shipped, tested, deployed (R-200's plumbing half)
--recover-offsite-install (controller v0.196.0, commit 1b1366b): same fetch → unseal → extract
path as --recover-offsite-check, same STDIN discipline for R, but it places the recovered
password via InjectOffboxPassword.
- Confirmation is a second invocation. Without
--confirm-installit prints both hashes and writes nothing. A single interactive prompt would have had to share stdin with R. - Three outcomes, named distinctly: installed (no local password — the rebuilt-box shape), unchanged (identical key already present, nothing written), refused (a DIFFERENT key present; installing would clobber the key the current repository is encrypted under — exit 2, no force offered).
- It re-reads the file after writing rather than trusting the call's return.
Tests + red-proof. go build && go vet && go test ./... rc=0; controller_gates.py --fast OK.
Removing the confirmation gate makes the dry run write the password and fails
TestRecoverAndInstall_InstallsOnABareBox with "the DRY RUN wrote the password". The R-persistence
test carries a positive control — a planted copy of the code is found by the sweep, then removed
and not found — because an absence check is worth only what its sensitivity is.
It was deployed to demo-hp and never exercised against a live recovery, because the drill halted before step 9. Its unit proof stands; its live proof does not exist.
6. What this drill did and did not establish
Established, live, for the first time:
- A rebuilt box's off-site repository is orphaned and the run refuses — the spike's predicted third outcome, measured. It does not silently start a fresh history.
- The orphan reset works: move-aside, never delete, fresh repo initialised, event pushed.
- A customer-declared mandatory data directory can be silently absent from the off-site snapshot
while the run reports
ok(R-203). - demo-hp's pre-rebuild off-site history is permanently unrecoverable — its key was destroyed four hours before the fix that would have kept it.
NOT established — and unchanged from before this session:
- No file has ever been restored from an off-site backup after a wipe. R-201's question is still open, and its pass condition is unchanged.
- Part 0's install path has never run against a live recovery.
- The v0.93.0
identity_blobretention is still unit-proven only — nothing in this session superseded a key, so nothing exercised it.
7. State left behind, and teardown
Deliberately not torn down — this is evidence, and the box is better off than it was:
| layer | state |
|---|---|
| the guest | calibre-web deployed and running, off-site-toggled, sentinel file in place. Kept as the drill fixture for the resumed run — it is the only app on either demo box with a restorable file leg. |
| the off-site repository | fresh, working, 3 snapshots, 30 636 B. The old 15-snapshot history is set aside, not deleted, at /home/felhom-repo.orphaned-20260804. |
| the host | agent v0.125.0, controller v0.196.0, pvesm unchanged apart from normal usage. |
| the hub | offbox_repo_orphaned + offbox_repo_reset events recorded for demo-hp. No scratch customer records were created — nothing was reinstalled. |
The ~1.2 GB of previously-orphaned ciphertext was NOT deleted — that act is ruled and still owed,
and §8.3 of the runbook forbids riding it along with a drill. The reset added
/home/felhom-repo.orphaned-20260804 (≈41 MB) to what is set aside.
No R persisted anywhere — the recovery code was never used in this session. It was read only to
confirm the key exists in the credentials file; no unseal was performed on demo-hp.
7b. UPDATE, same day — the blocker is FIXED and the drill is ready to resume
Controller v0.197.0 shipped both halves of R-203:
- the paths agree. On demo-hp the live bind moved from
/mnt/sys_drive/userdata/media/booksto/mnt/sys_drive/felhom-data/userdata/media/books— the directory the capture set looks in. The capture log went from0 mandatory path(s)to1 mandatory path(s). okmeans it. A run that cannot capture a MANDATORY directory now reportsincomplete, names the app and the folders, and raises the operator digest — instead ofokwith a warning beside it.
And the sentinel is in the snapshot, listed by name:
$ restic ls -l latest --tag calibre-web
-rw-r--r-- 1000 1000 181 2026-08-04 12:53:06 /mnt/sys_drive/felhom-data/userdata/media/books/DRILL-SENTINEL.txt
sha256 643166269103a25cf41d34a26b75fd6ebba0a837bbcd7e0d2b5aba7106bcbe7c — byte-identical to §1,
verified after the fix's migration moved the file to the corrected directory.
One live check could NOT be reproduced, and is recorded rather than claimed. Hiding the books
directory to watch the incomplete verdict fire on hardware did not work: the running container's
bind mount recreated the directory, so os.Stat succeeded and there was no gap. That is itself
worth knowing — a bind-mounted directory cannot easily be "missing" while its app runs, so the
mandatory-gap condition arises in practice when the path resolves somewhere the app never binds (the
R-203 case), not when a live app's own directory vanishes. The verdict is proven by the run-level
test, which drives the real RunOffboxBackup and asserts both incomplete and the operator signal.
The fixture was restored and the sentinel re-verified at the same hash.
8. To resume the drill
- Fix or scope R-203. Until the mandatory userdata path lands in the snapshot, no sentinel can survive a wipe and the drill cannot reach its verdict.
- Re-run steps 4–5 and confirm the sentinel IS in the snapshot — by listing it, not by a green status.
- Then P4 (the deliberate archive), the §7 STOP, and steps 6–11 as written.
Everything else is already in place: the code, the versions, the recovery code, the working repository, the file-leg app and the sentinel.