R-201 night run: the off-site key IS recoverable after a real rebuild (proven); the verdict is blocked by R-204
gates / gates (push) Successful in 6s
gates / gates (push) Successful in 6s
This commit is contained in:
@@ -1,172 +1,127 @@
|
||||
# REPORT — R-203: the app and its backup look in the same place, and "ok" means it (2026-08-04)
|
||||
# REPORT — R-201 night run: the key came back; the verdict did not
|
||||
|
||||
**Controller v0.196.0 → v0.197.0**, deployed to demo-hp. No hub change. Nothing deleted, wiped or
|
||||
moved on any box; the drill was **not** resumed. `demo-felhom` untouched.
|
||||
**2026-08-04, 21:30–22:40, unattended.** `demo-hp` was deliberately rebuilt. No code, no version bump.
|
||||
`demo-felhom` untouched. Full record: `documentation/audits/DRILL-r201-night-run-2026-08-04.md`.
|
||||
|
||||
---
|
||||
|
||||
## 1. Confirmed baselines, and §3's landmarks
|
||||
## 1. THE VERDICT — not reached
|
||||
|
||||
`felhom-controller` `532f5712a891` (v0.196.0) → **v0.197.0**; `felhom.eu` `a0c4b607a6cf`, no bump.
|
||||
Both trees clean on arrival. **Every §3 landmark held**, including the one that decided the fix:
|
||||
`paths.go:26` names the system-drive arrangement *"the SSD-only system-data fallback"* — **it is
|
||||
supported**, so the resolution was what was wrong and no API refusal was added.
|
||||
|
||||
## 2. The call sites found — **FIVE, not four**
|
||||
|
||||
| # | Site | Named by the spec? |
|
||||
|---|---|---|
|
||||
| 1 | `stacks/deploy.go` `withPathVars` → `${USERDATA_PATH}` | yes — the live defect |
|
||||
| 2 | `appexport/fabplan.go` | yes |
|
||||
| 3 | `appexport/export.go` | yes |
|
||||
| 4 | `stacks/delete.go` `ExportDataMounts` | yes |
|
||||
| 5 | **`web/handlers.go` FileBrowser mount builder** | **NO** |
|
||||
|
||||
The fifth is the customer's own file browser: on a non-enrolled path it would have mounted the wrong
|
||||
directory. **Latent, not live** — the system drive is deliberately never a registered `StoragePath`, so
|
||||
the resolver is the identity there today. Wired anyway, with that reason in the code.
|
||||
|
||||
Two further sites of the same class were found and fixed: `ComputeFabBuckets` was receiving the drive
|
||||
path where `ComputeCaptureSet` has always received the namespace root, so the export's classified
|
||||
paths and the backup's capture set could describe different directories for the same declared bind.
|
||||
|
||||
**The rule had TWO existing copies and they differed.** `backup.Manager.namespaceRoot` compared
|
||||
without `filepath.Clean`; `stacks.Manager.inGuest` compared with it. A trailing slash from config
|
||||
would have flipped the mode in one package and not the other. Both now delegate to
|
||||
`appbackup.NamespaceRootFor`.
|
||||
|
||||
## 3. Scenario A — the two paths, before and after (live, demo-hp)
|
||||
|
||||
```
|
||||
before: /mnt/sys_drive/userdata/media/books ← app bind; capture set looked elsewhere
|
||||
after: /mnt/sys_drive/felhom-data/userdata/media/books ← app bind == capture root
|
||||
```
|
||||
|
||||
Capture log: **`0 mandatory path(s)`** → **`1 mandatory path(s)`**.
|
||||
|
||||
## 4. Scenario F — the sentinel in the snapshot's file listing
|
||||
|
||||
```
|
||||
$ restic ls -l latest --tag calibre-web
|
||||
-rw-r--r-- 1000 1000 181 2026-08-04 12:53:06
|
||||
/mnt/sys_drive/felhom-data/userdata/media/books/DRILL-SENTINEL.txt
|
||||
```
|
||||
|
||||
And the snapshot's own `paths`:
|
||||
`["…/backups/primary/calibre-web", "/mnt/sys_drive/felhom-data/userdata/media/books"]`.
|
||||
|
||||
**Not a green status — the file, by name and size.** sha256
|
||||
`643166269103a25cf41d34a26b75fd6ebba0a837bbcd7e0d2b5aba7106bcbe7c`, byte-identical to the drill's
|
||||
step 3 after the fix's migration moved it into the corrected directory.
|
||||
|
||||
## 5. Scenario E — the delete path
|
||||
|
||||
**The §8.3 risk does not exist, and this is the correction owed.** `ExportDataMounts` lives in
|
||||
`delete.go` but is **export-only**: its single production caller is the `.fab` export adapter and
|
||||
nothing deletes on its result. The delete path's own guard, `ProtectedHDDPaths`, is **layout-agnostic
|
||||
by construction** — it protects both `<hdd>/…` and `<hdd>/felhom-data/…` — so deletion was never
|
||||
affected by the namespace-root defect. That note is now in the function's doc comment, because the
|
||||
file placement misled this change's own specification.
|
||||
|
||||
**Does the corrected path point anywhere the old code did not?** Yes, for the export only, and only on
|
||||
the system-data fallback: `<hdd>/felhom-data/userdata` instead of `<hdd>/userdata`. That is the
|
||||
directory the app actually binds after this release, and the tests assert the **negative** — no
|
||||
emitted path lies outside the app's own data roots, on either drive kind. It shipped as its own commit.
|
||||
|
||||
## 6. The verdict value
|
||||
|
||||
**`incomplete`** — **minted**, because `ok` | `error` | `running` contained nothing meaning *"it ran,
|
||||
and this app is not fully protected"*. **Not `error`:** the rest of the run worked and what was
|
||||
captured is real, so `SnapshotCount` and the `LastSuccess` anchor still record it.
|
||||
|
||||
It reaches the operator through the **existing** per-run digest, `backup_run_failures` — already
|
||||
operator-only, already allowlisted. A new event type would have been a two-repo change and the hub
|
||||
drops anything outside `allowedEventTypes`; the prompt ruled out a hub change. The Hungarian customer
|
||||
warning is unchanged, and the backups page renders `! Hiányos` with the warn styling.
|
||||
|
||||
## 7. §8.4's narrowing — **no customer-visible warning disappears**
|
||||
|
||||
`TierOffsite`'s `tierKeeps()` already admits `ClassMandatory` only, so an optional path cannot reach
|
||||
the stat-filter. The added class check is a **no-op today**, written for parity with Tier 2 — and
|
||||
**demonstrated to be load-bearing anyway**: widening the tier filter alone keeps the tests green
|
||||
*because of the check*; widening it and removing the check makes an optional gap start reporting.
|
||||
|
||||
## 8. §8.6's live effect — anticipated, and then NOT reproducible
|
||||
|
||||
Anticipated in the CHANGELOG before it could fire: calibre-web on demo-hp had exactly this gap, so its
|
||||
status would become `incomplete`. **In the event it did not**, because the same session fixed the
|
||||
underlying path — after the corrected bind and the data migration the app has no gap, and the run is
|
||||
legitimately `ok`.
|
||||
|
||||
**Attempting to observe the verdict live by hiding the directory did not work, and that is recorded
|
||||
rather than dressed up:** the running container's bind mount **recreated** it, so `os.Stat` succeeded
|
||||
and no gap existed. Worth knowing in itself — a bind-mounted directory cannot easily be "missing"
|
||||
while its app runs, so the mandatory-gap condition arises in practice when the path resolves somewhere
|
||||
the app never binds (the R-203 case), not when a live app's own directory vanishes. The verdict is
|
||||
proven by a run-level test that drives the real `RunOffboxBackup`. The fixture was restored and the
|
||||
sentinel re-verified at the same hash.
|
||||
|
||||
## 9. Tests and every red-proof
|
||||
|
||||
| Scenario | Result | Red-proof — mutation → outcome |
|
||||
|---|---|---|
|
||||
| A both paths agree | PASS + **LIVE** | restore the bare-path call → **FAIL**: `/mnt/sys_drive/userdata` vs `/mnt/sys_drive/felhom-data/userdata` |
|
||||
| B enrolled drive unchanged | PASS | invert the drive-kind comparison → **FAIL** on every enrolled row |
|
||||
| C mandatory gap → not ok | PASS | unreachable gap recording → **FAIL**; unconditional `ok` → **FAIL** |
|
||||
| D optional gap → ok | PASS | class check removed **with the tier filter widened** → **FAIL**; tier widened alone → PASS, i.e. the check holds the line |
|
||||
| E export mounts, both kinds + the negative | PASS | leave the site bare → **FAIL**, emits the short path |
|
||||
| F sentinel in the listing | **LIVE** | — |
|
||||
|
||||
**A RED-PROOF PASSED AND THE TEST WAS WRONG, NOT THE CODE** (§9.11 — three of the last six sessions).
|
||||
My first Scenario-C test exercised `offboxCaptureSet` alone, while the mutation lives in
|
||||
`runOffboxInternal`. A mutation the test cannot observe is not a red-proof. Replaced with a run-level
|
||||
test that drives `RunOffboxBackup` and asserts `incomplete`, the retained `LastSuccess`, and the
|
||||
operator signal; it fails under both mutations. The original Scenario-D "red-proof" also could not
|
||||
fail by construction — recorded above with the two-part mutation that does.
|
||||
|
||||
Green gate after each phase (`go build && go vet && go test ./...`, rc=0) plus
|
||||
`controller_gates.py --fast`. No test run was combined with a commit.
|
||||
|
||||
## 10. Commits
|
||||
|
||||
| Commit | Contents |
|
||||
| | |
|
||||
|---|---|
|
||||
| `73efb09` | Part 1 — one resolver, the four non-export sites, `ComputeFabBuckets`, the two delegating copies |
|
||||
| `a96c3d9` | **Part 1.3 alone** — the `delete.go`-resident export-mount site, with its scope correction |
|
||||
| `58c703b` | Part 2 — the verdict, the structural gaps, the operator digest, the status rendering |
|
||||
| `73fb595` | docs (felhom.eu) |
|
||||
| sentinel sha256, pre-wipe | `643166269103a25cf41d34a26b75fd6ebba0a837bbcd7e0d2b5aba7106bcbe7c` |
|
||||
| sentinel sha256, post-restore | **not obtained** — the restore was never reachable |
|
||||
|
||||
## 11. Registers
|
||||
**Not a FAIL.** Nothing came back wrong and no fresh history was started; the box never got as far as
|
||||
running a backup. **What it is instead:** the first proof that the key survives and returns, plus the
|
||||
measured reason a customer still cannot use it.
|
||||
|
||||
- **R-203 → SHIPPED + PROVEN-LIVE.**
|
||||
- **R-201 → READY TO RESUME**, blocker gone, fixture staged and verified.
|
||||
- **R-202 stays open**; the orphaned-ciphertext deletion is **still owed**.
|
||||
- Capability map: new row for mandatory-path capture on both layouts, its predecessor corrected as
|
||||
optimistic, and the **R-199 back-pointer omitted two sessions ago is now added**.
|
||||
- The v0.93.0 `identity_blob` retention remains **unit-proven only** — nothing here superseded a key.
|
||||
> **After a real rebuild — controller data volume destroyed, sentinel deleted from disk — the
|
||||
> customer's recovery code produced `8a9e33aa4da6769c5aea1831f87759e10930e2ec1dea0062576484e0598d080a`,
|
||||
> byte-identical to the pre-wipe on-disk key and to the hub's independent record.** That has never been
|
||||
> shown before.
|
||||
|
||||
## 12. CI
|
||||
## 2. Snapshot count at step 9 — not obtained
|
||||
|
||||
Run numbers and task ids in the session summary. **`--no-verify` was not used.**
|
||||
The off-site run was never permitted to start (§4). **And a count would have been a poor
|
||||
discriminator anyway:** restic's same-day `forget --keep-daily 7 --group-by host,tags` keeps one
|
||||
snapshot per tag per day, so a successful reopen would have shown 3, not 4. The real observable is
|
||||
whether the pre-wipe snapshot `e6132ae5` survives with the sentinel in it — recorded as the resume
|
||||
step.
|
||||
|
||||
## 13. Teardown
|
||||
## 3. §5's five conditions, recorded before the wipe
|
||||
|
||||
**Nothing provisioned.** `calibre-web` and the sentinel stay — R-201 needs them. The temporarily
|
||||
hidden directory was restored and the sentinel re-verified byte-identical.
|
||||
1. **sentinel listed by name** — snapshot `e6132ae5` (19:36:26), `-rw-r--r-- 1000 1000 181 …/DRILL-SENTINEL.txt`.
|
||||
2. **rollback archive verified** — `vzdump-lxc-9201-2026_08_04-21_44_24.tar.zst`, 1 606 765 083 B,
|
||||
**full zstd stream read OK** (4 867 573 760 B), sentinel confirmed inside it.
|
||||
3. §3's option — §5 below.
|
||||
4. **demo-felhom** `health=ok`, `escrow_state=escrowed`, untouched throughout.
|
||||
5. **space** — nvme-1tb 883 GB free, root 20 GB, local-lvm 34.32 %.
|
||||
|
||||
## 14. Observations — noticed, NOT acted on
|
||||
**One precondition had drifted and was repaired, not worked around.** The staged snapshot no longer
|
||||
held the sentinel: the afternoon's `mandatory data path missing` experiment produced a *later* same-day
|
||||
calibre-web snapshot and `forget` had pruned the good one. The fixture on disk was correct, so one
|
||||
backup re-established it and it was re-verified by listing. **Lesson: a good snapshot is not durable
|
||||
against a later bad run on the same day.**
|
||||
|
||||
1. **`resolveAbs` uses ONE root for two bind classes.** `RootHDD` resolves against the same parameter
|
||||
as `RootUserdata`. On an enrolled drive they coincide; on the system-data fallback a `${HDD_PATH}`
|
||||
bind resolves under the namespace root while compose binds it bare. Both callers now pass the
|
||||
namespace root, so the export and the backup **agree with each other** — but whether `${HDD_PATH}`
|
||||
itself should mean the namespace root on the system drive is a separate question that touches every
|
||||
already-deployed app's binds. **Not fixed here; it needs a decision, not a patch.**
|
||||
2. **The blast radius was measured before changing anything:** exactly one app in the fleet has
|
||||
`HDD_PATH == system_data_path` — `calibre-web` on demo-hp, deployed for the drill. Every other
|
||||
deployed app has no `HDD_PATH`. Nothing else needed migrating.
|
||||
3. **`/api/stacks/<name>/deploy` is first-deployment-only** (409 afterwards), so the corrected
|
||||
`${USERDATA_PATH}` reaches an existing app through a start/redeploy, not a re-deploy.
|
||||
4. **The controller's CSRF form field is `_csrf`, not `csrf_token`** — the login page uses one name and
|
||||
the protected forms another. Second session in a row this cost time; it belongs in the
|
||||
headless-access memory.
|
||||
## 4. Every step's observable
|
||||
|
||||
| step | observable |
|
||||
|---|---|
|
||||
| 6 | fresh data dir stamped `20:00:2x`; **new `encryption.key`** (32 B); `claimed = None`; `offbox = null` |
|
||||
| **7** | **`identity_blob` 572 B, `restic_pw_sha256` `8a9e33aa4da6…`, `updated_at` still `11:11:37` — UNCHANGED across the wipe.** Nothing re-escrowed itself |
|
||||
| **8** | **recovered `8a9e33aa4da6…`** — matches pre-wipe on-disk and the hub's record. Exit 1 is correct and designed: nothing local to compare against, the rebuilt-box shape |
|
||||
| 9a | `[INSTALLED] … reads back identical` — the "installed" branch's first real run |
|
||||
| 9b | after the Re-issue and apply, the on-disk key is **still** `8a9e33aa4da6…` — `WriteOffboxSecrets` kept it |
|
||||
| 9c | **blocked** — see below |
|
||||
| 10–11 | **not run** |
|
||||
|
||||
**The wipe was faithful to the incident, deliberately.** The 2026-08-03 rebuild R-193 is filed against
|
||||
was **not** a guest reprovision — the journal shows guest 9201 running continuously through that window
|
||||
with no `pct destroy`, no `pct restore` and no `--selftest=provision`. What changed was the controller
|
||||
and its data volume. A guest destroy plus an unrehearsed provisioning chain, improvised unattended, is
|
||||
what §8.10 exists to prevent.
|
||||
|
||||
## 5. §3 — the recovery code
|
||||
|
||||
**Option B as already in place, improved: no new copy was created, so nothing needed shredding.** The
|
||||
operator had placed `R_DEMO-HP` in their own `~/.config/credentials` (`0600`) two sessions ago for this
|
||||
purpose. It was read from there and **piped to stdin** for the two invocations that needed it — never
|
||||
an argument, never exported, never written to a second file, never logged. Destroying the operator's
|
||||
own store would have destroyed their record; because no additional copy existed, there is nothing left
|
||||
to prove gone.
|
||||
|
||||
**Verified anyway, with the planted-copy positive control:** 0 hits in the agent journal, 0 in the
|
||||
controller log, 0 files under `/tmp`, `/var/tmp`, `/var/lib/felhom-agent`, `/root`, 0 leftover
|
||||
`felhom-idesc-*` staging dirs — and the same sweep found a planted copy (**1**), then **0** after
|
||||
shredding it. The instrument is shown sensitive rather than assumed to be.
|
||||
|
||||
## 6. Part 2 — did not run
|
||||
|
||||
Its gate is "the drill PASSED". It did not. A second wipe would have destroyed the state that makes the
|
||||
first drill finishable in five minutes. **R-198's retention therefore remains unit-proven only** —
|
||||
nothing has yet superseded a key in production.
|
||||
|
||||
## 7. Teardown — three layers
|
||||
|
||||
| layer | state |
|
||||
|---|---|
|
||||
| the guest | **nothing torn down** — the wipe is the evidence; all six app containers serving; the recovered key on disk; scratch band empty |
|
||||
| the host | vzdump snapshot LVs released cleanly; one new 1.6 GB archive on nvme-1tb (883 GB free); nothing deleted |
|
||||
| the hub | **no new customer records** — `demo-hp` is the same row throughout, because the rebuild was a controller-data wipe and not a re-enrolment. One Re-issue staged a one-time secret (consumed 20:16) and set `stale_at`. The two superseded escrow rows are **unchanged** — no ceremony ran |
|
||||
|
||||
**Nothing deleted on the storage endpoint**, including the ~1.2 GB of orphaned ciphertext — ruled,
|
||||
still owed, deliberately not ridden along with a drill.
|
||||
|
||||
## 8. The capability-map rows
|
||||
|
||||
A new row records the key as **PROVEN-LIVE after a real rebuild**, and states plainly what it does not
|
||||
claim: **no file has been restored**, and the four links of R-204 stand between the recovered key and a
|
||||
usable repository. The R-199 back-pointer added earlier today stands.
|
||||
|
||||
## 9. New findings
|
||||
|
||||
- **R-204 (NEW)** — the recovered key cannot be used: a rebuilt box cannot configure its tier (R-193);
|
||||
the Re-issue that fixes that marks the escrow stale though the key never changed (R-196, **measured
|
||||
at `20:15:49` with `restic_pw_sha256` unchanged**); a stale escrow gates every run; and the only way
|
||||
to clear it is a ceremony that **destroys the recovered key**. Plus a fourth link in no design
|
||||
document: a rebuilt box is **unclaimed**, so the claim gate intercepts every controller endpoint.
|
||||
- **R-196 re-scoped** — no longer a documentation nit; it is on the critical path for recovery.
|
||||
- **R-201** — the wipe happened, the key came back, ~5 minutes from a verdict with a person present.
|
||||
- **R-198** — still unit-proven only; Part 2 gated out.
|
||||
- **R-202 open; the ciphertext deletion still owed.**
|
||||
|
||||
## 10. CI
|
||||
|
||||
Docs push only. Run number and task id in the session summary. **`--no-verify` not used.**
|
||||
|
||||
## 11. Observations — noticed, NOT acted on
|
||||
|
||||
1. **The claim gate is an undocumented first step of every recovery.** Before a customer can do
|
||||
anything on a rebuilt box — including recovering their backups — they must re-claim it.
|
||||
2. **`--print-reset-code` output needs parsing care**: the captured value was 73 characters, i.e. more
|
||||
than the code itself. Not chased; the claim was abandoned when the session stopped.
|
||||
3. **A same-day re-run replaces the day's snapshot.** Worth knowing before designing any drill that
|
||||
depends on a specific snapshot surviving.
|
||||
4. **The hub's ClusterIP is not reachable from DooPlex's host network** — the operator UI needs a
|
||||
`kubectl port-forward`. The `curl -u :$HUB_PW` recipe in memory omits that.
|
||||
|
||||
Reference in New Issue
Block a user