10 KiB
REPORT — R-201 drill: PREPARED, HALTED BEFORE THE WIPE (2026-08-04)
Nothing was wiped. Nothing irreversible happened. The drill stopped at step 4 because the sentinel file was not in the off-site snapshot — wiping would have destroyed the only copy and proven nothing.
1. THE VERDICT — not reached, and that is the correct outcome
The pass condition is a byte-identical sentinel sha256 after a wipe. Step 4 established the sentinel is not in the off-site snapshot at all, so there was nothing to recover and the verdict is unobtainable. The runbook's own rule applies: a precondition was wrong, and the finding outranks the drill.
Sentinel (step 3), still on the box:
643166269103a25cf41d34a26b75fd6ebba0a837bbcd7e0d2b5aba7106bcbe7c
at /mnt/sys_drive/userdata/media/books/DRILL-SENTINEL.txt, 181 B.
2. Snapshot count at step 9 — N/A; at step 4 it was 3, and the sentinel was in none of them
last_run 2026-08-04T12:54:50Z · last_status ok · last_success stamped
snapshot_count 3 · repo_size_bytes 30 636
[WARN] [offbox] calibre-web: mandatory data path missing on disk, skipped from offsite:
/mnt/sys_drive/felhom-data/userdata/media/books
[INFO] [offbox] backed up calibre-web (…, 0 mandatory path(s))
[INFO] [offbox] backup OK: 3 app(s) backed up, 3 snapshot(s), 37s
The run reported ok. → R-203, §9.
3. Preconditions P1–P6
| # | Result |
|---|---|
| P1 recovery code held | PASS — R_DEMO-HP, recorded by the operator in the DooPlex credentials file. The code appears nowhere in this report and was never used — no unseal ran on demo-hp. |
| P2 versions | FAILED ON ARRIVAL, FIXED — the box was on agent v0.124.1. Deployed published v0.125.0 (sha f7d8339b53d9…, verified against the release output) + controller v0.196.0. age present at /usr/bin/age. |
| P3 the blob seals the repo under test | PASS with the caveat that reshaped the drill — identity_blob 572 B, restic_pw_sha256 8a9e33aa4da6…; but no repository existed under that key (§5). |
| P4 verified rollback archive | NOT TAKEN, deliberately — only needed for the wipe, and the wipe did not happen. |
| P5 demo-felhom untouched | PASS — not touched at any point. |
| P6 space | PASS — felhom-backup 927 GB free, local-lvm 30.9 %, guest 64 GB free. |
4. Every step's observable
| step | observable |
|---|---|
| 1 | identity_blob 572 B; restic_pw_sha256 8a9e33aa4da6…; snapshot_count: 0, no last_run/last_status at all — never ran an off-site backup in this lifetime |
| 1b | repository ORPHANED — see §5 |
| 1c | reset (operator-authorised) — move-aside to /home/felhom-repo.orphaned-20260804, fresh repo initialised, offbox_repo_reset pushed |
| 2 | recovery codes for both boxes created and saved by the operator earlier the same day; escrow rows stamped 11:11:37Z / 11:13:06Z with restic_pw_sha256 unchanged on both — correctly treated as same-password re-ceremonies (no superseded row, no offsite_repo_key_changed): v0.93.0's Scenario E, live |
| 3 | calibre-web deployed through the real API (HTTP 202), toggled off-site, Tier-1 recovery unit captured; sentinel written + synced + hashed |
| 4 | ok, 3 snapshots, 30 636 B — and the mandatory directory missing |
| 5–11 | NOT RUN. The §7 STOP was never reached: the drill failed its own precondition first |
5. Step 1b — the first live proof of a prediction, and it is the good news
The 2026-08-04 spike predicted from source, and could not measure, that a rebuilt box's next off-site run would hit a third outcome — neither reattaching the old snapshots nor silently starting fresh, but refusing. Measured two ways:
Read-only probe (restic cat config with the current key, writes nothing):
Fatal: wrong password or no key found — the exact string classifyResticProbe maps to orphaned.
Then the real customer path (POST /backup/offbox/run):
[WARN] [offbox] offsite repo ORPHANED — remote holds backups written under a previous,
no-longer-available key; runs will skip until reset
[WARN] [offbox] run skipped — offsite repo orphaned (card shown; awaiting reset)
[INFO] Event pushed: offbox_repo_orphaned (warning) …
repo_state: orphaned, orphan card rendered, event at the hub (HTTP 200).
The system stopped and said so. It did not quietly start a new history. Closes R-193's Q3.
Why it was orphaned: the 15 snapshots / 40.9 MB were written under key 8e03eddf9ff7… before the
2026-08-03 rebuild. That key survives only in superseded escrow row id 3, whose identity_blob is
NULL — superseded at 2026-08-04 07:15:36, four hours before hub v0.93.0 fixed the retention.
Permanently unrecoverable, with or without a recovery code.
6. Step 7's assertion (identity_blob unchanged across the rebuild) — NOT REACHED
There was no rebuild. Unmeasured, and stated as such.
7. Did the built chain carry steps 8–9? — NOT REACHED
Neither the check nor the install ran against a live recovery on demo-hp. The manual substitute was not needed and was not used. Part 0 remains unit-proven only.
8. Part 0 — files, commit, tests, deployment
Commit 1b1366b, controller v0.196.0, deployed to demo-hp guest 9201 (Up (healthy)).
internal/backup/offbox_recovery_cli.go (RecoverAndInstall), cmd/controller/main.go
(--recover-offsite-install + --confirm-install), internal/backup/offbox_recovery_check_test.go.
go build && go vet && go test ./... rc=0; controller_gates.py --fast OK. Test run and commit were
never combined.
Red-proof observed: removing the confirmation gate makes the dry run write the password —
TestRecoverAndInstall_InstallsOnABareBox fails with "the DRY RUN wrote the password".
9. R-203 — the finding that halted the drill
| what | path | exists? |
|---|---|---|
| the app's live bind (where the sentinel is) | /mnt/sys_drive/userdata/media/books |
YES |
| what the off-site capture set looked for | /mnt/sys_drive/felhom-data/userdata/media/books |
NO |
Mechanism, from source: NamespaceRoot (appbackup/paths.go:28-33) appends felhom-data when
the drive IS the system data path (backup/backup.go:331); the deploy-time ${USERDATA_PATH} is
<HDD_PATH>/userdata (stacks/classify_binds.go:14). With system_data_path: /mnt/sys_drive and an
app at HDD_PATH=/mnt/sys_drive, they differ. The same compose used both roots —
${IMPORT_PATH} → …/felhom-data/userdata/import/calibre, ${USERDATA_PATH} →
/mnt/sys_drive/userdata/media/books.
Measured vs not, because it changes the fix. MEASURED: the disagreement, the absence from the
snapshot, the ok status, and that the only signal is a container-log WARN. NOT ESTABLISHED:
whether HDD_PATH=/mnt/sys_drive is supported — it was chosen because demo-hp's only registered drive
(Felhom-Share) is a NAS and was correctly refused as an app namespace (R-108 working), while
/mnt/sys_drive was accepted (202). Either branch is a defect and the row says so.
Not a general off-site failure: opengist and privatebin declare no mandatory userdata paths, so
they are unaffected and their snapshots are real.
10. Teardown — three layers, nothing torn down that proves anything
| layer | state |
|---|---|
| the guest | calibre-web deployed, running, off-site-toggled, sentinel in place. Kept — it is the only app on either demo box with a restorable file leg, and the resumed drill needs it |
| the off-site store | fresh, working, 3 snapshots, 30 636 B. Old 15-snapshot history set aside, not deleted, at /home/felhom-repo.orphaned-20260804 (≈41 MB) |
| the host | agent v0.125.0, controller v0.196.0; pvesm unchanged beyond normal usage |
| the hub | offbox_repo_orphaned + offbox_repo_reset recorded for demo-hp. No scratch customer records were created — nothing was reinstalled |
The ~1.2 GB of previously-orphaned ciphertext was NOT deleted — ruled, still owed, and §8.3 forbids riding it along with a drill.
11. The capability-map row as it now reads
The R-199 row (the KEY is recoverable, PROVEN-LIVE) keeps its scope and gains the back-pointer the previous update omitted, plus: the attempt to prove the DATA half was halted before the wipe, and no file has still ever been restored from an off-site backup after a wipe. The install half (v0.196.0) is likewise unit-proven only.
12. Registers
- R-203 NEW — the missing mandatory directory. Highest ID before this session was R-202.
- R-201 — prepared, halted, with everything it established and exactly what resuming needs.
- R-200 — plumbing complete (v0.196.0); the form is not built; the install has never run live.
- R-202 stays open. The ciphertext deletion is still owed.
- The v0.93.0
identity_blobretention is still unit-proven only — nothing here superseded a key.
13. CI
Run numbers and task ids in the session summary; --no-verify was not used.
14. Observations — noticed, NOT acted on
${IMPORT_PATH}and${USERDATA_PATH}resolved to different roots in one deploy — the sharpest single expression of R-203, and the place a fix should probably start.- The controller's CSRF form field is
_csrf, notcsrf_token— the login page carries one name and the protected forms another; an hour went into that. Worth a line in the headless-access memory. /api/stacks/<name>/deployacceptedHDD_PATH=/mnt/sys_driveone call after refusing the NAS path with a clear Hungarian error. If the system drive is not a valid app namespace, that refusal is missing.- demo-hp has no registered drive usable as an app namespace — its only drive is a NAS. Any drill needing a file-leg app on that box is constrained by this.
- The
felhom-r-leak-style positive control is now used in three separate absence checks across two sessions. It is cheap and it has already caught one wrong test; worth making a documented habit.