diff --git a/REPORT.md b/REPORT.md index 319ef548..4c0e7d41 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,72 +1,80 @@ -# REPORT — R-356 doc corrections, register housekeeping, and one gate fix (2026-08-22) +# REPORT — DRILL R-356b: the off-site restore for a driveless app that HAS a database (2026-08-22) -Companion to `felhom-controller` v0.219.0 (R-356). This repo carried the architecture correction, the -drill record, the register move and one genuine gate defect found on the way. +**A drill, not an implementation.** No production code was written, no version bumped, no CHANGELOG +entry made. The deliverables are a findings document, four register rows and a capability-map update. -## 1. `documentation/architecture/07-backup-architecture.md` — R-107 was closed and the doc said otherwise +Full record: `documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/` -Three places, each corrected with a dated **[FACT]** citing `offbox_reconstitute.go` `volReplay` and -controller v0.218.0: +## What was measured -- §6.3, the Tier-3 row — was *"no offsite action unpacks the named-volume tars it captures"*. -- §8, matrix row 4 — was *"the volume tars in either copy are unreachable"*. -- The R-107 index row. +Ten of the forty driveless apps carry a database. **I re-measured that count myself and got 10** — the +same ten the runbook names. For those ten, restoring is a five-leg operation that, until this week, +never ran at all: R-356 refused before any of it started. -**The old sentence's history is kept, not deleted:** each correction says what was true, until when, -and what closed it. A correction that erases what was believed leaves the next reader no way to tell -a fixed gap from one that was never noticed. +Both engines were walked end to end on `demo-hp`: `docmost` (Postgres 16) and `bookstack` +(MariaDB 12.3), each deployed for this drill, planted through the app's **own** interface, destroyed +for real, and restored through the exact endpoint the UI's button posts to. -**R-102 — the Tier-2 half — is NOT closed, and the correction says so explicitly** so it cannot be -read as covering both. The Tier-2 row stands exactly as written. +## The three answers -## 2. §6.3 — the R-356 reasoning recorded as reasoning, not as a closed row +**Q1 — does it complete? YES.** All five legs ran in order and all succeeded — 32 s for Postgres, +25 s for MariaDB. Data back, apps healthy, accented names byte-identical in both directions. -A new **[DESIGN]** paragraph: the restore destination is resolved by the same rule as the capture -destination (drive if the app has one, system data path otherwise); the refusal that protects a drive -app from being restored onto the wrong disk applies to apps that **have a drive to get wrong**. It -carries the 13/40 measurement and points at `felhom-controller/CONTEXT.md`. +**Q2 — which leg returned the data? The SQL dump.** A three-way discriminator (volume tar +`ORIGINAL-VALUE-A`, altered dump `ALTERED-VALUE-B`, live `LIVE-VALUE-C3`) returned **`ALTERED-VALUE-B`**. +The ordering the code comment asserts holds in practice. This **confirms R-164's F17 claim on a second +path** — R-164 cites `restore_unit.go`, the local restore; this measures `offbox_reconstitute.go`. +The mutation was applied to the prepared scratch only, and the store was proved unmutated afterwards +by re-preparing a fresh scratch (sha256 back to `c5414f24…`). -## 3. `STATUS.md` — the contradiction is gone, and the page is one screen +**Q3 — does a failure tell the truth? Partly.** The customer does see a failure and the undo copy is +named. But two things are wrong, and they are the drill's findings. -It said nothing was waiting while also saying one decision was waiting — to publish agent 0.130.0 — -which the same page recorded as already published (R-347, closed). **218 → 102 lines.** +## Findings filed — R-379 … R-382 -"Waiting on you" now lists **four** real items, and — per the exemption — **each says what happens if -the operator does nothing**. The two new ones are this release's hand-off: vouch a golden carrying -controller 0.219.0, then raise the floor last, in a separate save. +- **R-379 (HIGH)** — the undo copy is valid, is named, and **nothing in the product can apply it**. + Proven by applying it by hand on both engines and getting the exact prior state back. + `pre-restore-` files are deliberately skipped at three code sites; the filename appears only inside + an error string. +- **R-380 (HIGH)** — a failed **MariaDB** replay leaves a partially-applied database behind an app + reporting `health=healthy, running=true, restarts=0`. `bookstack`'s schema-version ledger was wiped + to 0 rows while its user data stayed intact and the dashboard said fine. Postgres, by contrast, + fails visibly (crash-loop). H3 fired — but not in its predicted shape: the prediction was a *quiet + success*; what happens is a loud error and a silent inconsistency. +- **R-381 (MEDIUM)** — the failure message pastes raw engine stderr into the Hungarian customer + surface: 407 bytes for Postgres, **615 for MariaDB, whose middle is an `INSERT INTO migrations + VALUES (…)` listing — actual table rows shown to the customer.** +- **R-382 (LOW)** — the reconstitution's summary log omits the volume count it already has. The + customer-facing flash names the volumes; the operator log does not. -## 4. `documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/` +**Register: 325 236 bytes before, 330 683 after.** Ceiling was R-378; next free id is now R-383. -The live walk on `demo-hp`: both classes, both messages verbatim with their byte counts, both accented -filenames as explicit hex, and 16 evidence files. **Copied off at the end of each phase.** Nothing was -reverted during this drill, so no intermediate teardown could have taken it. +## Also recorded -## 5. Register housekeeping (N.7) +- **R-361 reproduced independently** on a second app: after the first reconstitution docmost's + `db-dumps/` held only `pre-restore-*` files. Not re-filed — noted as corroboration. +- **`restic check` passed** at the end: `no errors were found`, 29 snapshots. +- **The `-db` suffix attribution is correct** for `bookstack-db` — the R-355 shape does not reproduce. +- **Observed, not filed:** a newly deployed app is absent from the off-site set until switched on by + hand. Plausibly deliberate; the consequence is stated so the default can be judged. +- **A flaw in the drill's own method, recorded rather than hidden:** the first accented title was + double-escaped by a shell chain and stored as literal ASCII. Caught by reading the stored bytes back + as hex, and re-measured properly in Phase 1b. -R-356 compressed out of `OPEN-ITEMS.md` into `CLOSED-ITEMS.md`, keeping its title, shipping version, -evidence path and every sentence stating a rule, plus the pointer -`git show e18668f9e19f:documentation/backlog/OPEN-ITEMS.md` for the full original text. +## Capability map -| file | before | after | -|---|---|---| -| `OPEN-ITEMS.md` | 327 109 bytes | **325 236 bytes** | -| `CLOSED-ITEMS.md` | 61 580 bytes | **63 507 bytes** | +The 2026-08-21 narrowing of *"A customer's file survives a machine rebuild and comes back"* is now +history: both defects it named (R-354, R-356) are closed and proven. The row records what is now +walked — including this drill — and states plainly what is still **not** claimed: the success path is +proven, the recovery-from-a-bad-restore path is not. -## 6. A real gate defect, found by CI going red +## Teardown, three layers -CI run **387** (job 386) failed `instructions_gate` on commit `08eb1a6` while run 73 on its parent had -passed. The cause was not the push: `ef6ac6f` (this repo, the same day) compressed closed rows out of -`OPEN-ITEMS.md` into `CLOSED-ITEMS.md`, and `register_state()` read only `OPEN-ITEMS.md` and -`ROADMAP.md`. **Every citation of a compressed item became "a reference to nothing"**, failing the -next push in a sibling repo for a rule file nobody had touched — and it would have fired again on this -task's own R-356 compression. +1. `docmost` and `bookstack` were deployed by this drill and are **RETAINED** with their planted data — + it is the evidence, and they are the only deployed members of this app class on the box. Both left + healthy and sane. +2. No `pvesm` "before" snapshot was taken — **said plainly rather than reconstructed.** Measured + directly: ~233 MB of volumes inside guest 9201 (16 % of 69 GB used). +3. **No hub-side record was created.** No customer, no appliance. Nothing to dispose of. -`CLOSED-ITEMS.md` is now the third source. It answers *does this ID exist* and answers *closed* for -the rows it owns; `OPEN-ITEMS.md` remains the sole authority on **openness**, so a row it claims as -open is not overridden. **Both controls still convict:** an ID present nowhere fails, and a citation -claiming a closed item is still open fails — each watched failing, then watched clearing. - -## Gate state - -`python3 scripts/repo_gates.py --fast` — all green **except `golden-currency`**, which is correct and -is operator item 1: controller 0.219.0 is released and no golden carries it. +All phases were run. Nothing was dropped. diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index fb00c91b..433114bf 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -55,7 +55,9 @@ | **The offsite repository password can be RECOVERED from the sealed escrow with the customer's recovery code** | hub v0.94.0, agent v0.125.0, controller v0.195.0 | **PROVEN-LIVE (2026-08-04)** | On demo-felhom, through the real endpoints end to end: the box fetched its own sealed blob from the hub with its own per-host credential (hub log: *escrow blob SERVED … 572 opaque bytes, self_scope=true*), the agent unsealed it with the customer's recovery code, and the extracted repository password's sha256 was **byte-identical** to the one on disk — `c60c8bc737a6…`, which is ALSO the hash the hub had independently stored, so three sources agree. Five minutes earlier the same path with a WRONG code failed closed at age's KDF with nothing written, which proves links 6 and 7 ran independently of the success. R was searched for afterwards and found in 0 log lines and 0 files, with a positive control confirming the search would have found it. Evidence: per-repo CHANGELOGs; `audits/RECON-offsite-dr-chain-2026-08-04.md` §3 links 6–8 | **WHAT THIS ROW DOES NOT CLAIM, stated because the previous over-claim here was struck out four hours earlier.** It covers the KEY, not the DATA. **R-199 BACK-POINTER (omitted when this row was written): the recovery chain's link inventory and the per-link status live in `audits/RECON-offsite-dr-chain-2026-08-04.md` §3; links 1–8 are walked, 9–11 are not.** **Re-confirmed 2026-08-04 evening by an attempt to prove the DATA half:** the R-201 drill was prepared on demo-hp and **halted before the wipe** — the sentinel file was not in the off-site snapshot (R-203), so the wipe would have destroyed it and proven nothing. **No file has still ever been restored from an off-site backup after a wipe** (`audits/DRILL-r201-offsite-recovery-2026-08-04.md`). The install half of the chain (controller v0.196.0 `--recover-offsite-install`, R-200) is likewise unit-proven only — it has never run against a live recovery. A recovered password has never been **installed** (the diagnostic compares and refuses to write, by design), no existing repository has ever been **reopened** under one, and **no file has ever been restored** from an off-site history via a recovered key. Links 9–11 of the chain are open (R-200's remaining half, R-201). The proof also used a box whose local key still exists — the rebuilt-box case, where there is nothing to compare against, is exactly what the drill covers and it has not run || DR tier by default: PBS + WireGuard base infra on every install, hub-controlled activation | installer v1.15, agent v0.86, hub v0.51 | **PROVEN-LIVE** (2026-07-21) | `DRILL-day0-take2-2026-07-12` §2 (WG enabled both modes, PBS-DR descriptor auto-provisioned ~1s after WG registration, zero operator steps); ships installer v1.15/agent v0.86/hub v0.51 | Live only on demo/drill fleet. (Cited spike was slice-0 mechanics — shipped nothing; corrected.) **⚠ The candidate upgrade to PROVEN-LIVE is WITHDRAWN — the 2026-07-18 rehearsal produced a live counter-example (R-39).** On the reborn N100 the descriptor auto-provisioned and the agent reported `converged state=applied` (16:45:53), yet **the storage is dead**: `pvesm status` → `felhom-pbs: error fetching datastores - 401 Unauthorized` / `inactive`, and a direct probe with the stored credential returns **401 on every endpoint including `/version`** while the WG transport is healthy (handshake 9 s, 27.9 ms RTT) — i.e. authentication failure, not ACL scope. Root cause in the evidence: **the hub minted a SECOND token secret at 16:47:52, two minutes after the agent had applied the first, and `consumed_at` is still NULL**; the converged state machine will not re-apply, and the agent's 15-minute verify loop **cannot even read the credential to notice** (`open /etc/pve/priv/storage/felhom-pbs.pw: permission denied` — non-root agent reading a file it writes through a root wrapper). A tier that reports `applied` while silently unable to authenticate is exactly the shape that must not carry a PROVEN-LIVE badge. See `tests/VALIDATION-n100-rehearsal-2026-07-18.md` F2 and `pbs-dr-state.txt`. **agent v0.89.0 closes the F4 non-default-storage-id gap (R-22) — PROVEN-LIVE 2026-07-17:** the reconcile self-grants the ACL through the root wrapper on a pre-check 403 instead of dead-locking. Reproduced F4 on the demo (marker moved aside = reinstall fresh-state + felhom-offsite ACLs revoked) → next reconcile tick `pbsdr: pre-check 403 … self-granting … (R-22)` → `converged state=adopted` in ~3 s, ACLs self-restored, `pvesm status felhom-offsite`=active, zero operator action. No more one-shot `pveum` grant **2026-07-21 — the R-39 fleet fix SHIPPED (hub v0.68.0 + agent v0.91.2), closing the self-heal chain end to end.** The three defects that let a box be `applied` and dead simultaneously are each addressed: the hub stamps a monotonic `secret_generation` into the descriptor so a credential re-key finally MOVES the content hash the agent re-applies on; the wrapper gains a narrow `read` verb so the non-root agent can read the credential it writes (it never could — `/etc/pve/priv` is 0700 root:www-data, which made the verify loop blind by construction); and `pbs.ProbeAuth` turns a 401 into a loud `auth_failed` that the existing `pbsdrheal` damper escalates to a fresh mint. Plus a consumed_at honesty gauge for the disagreement no single tier can see (box says `applied`, hub's staged secret never consumed). Proven live on felhom-pve: the agent read its credential through the wrapper (`rc=0`) and probed successfully (`credential probe OK storage=felhom-pbs`). **STOP-2 RAN 2026-07-21 AND THE CHAIN CLOSED — 13 SECONDS, operator click to converged.** The operator pressed **Re-issue PBS credentials**; the identical click on 2026-07-18 did nothing at all. Full chain (hub UTC / host CEST = UTC+2): `08:39:31Z` hub mints a fresh secret, **generation 0 → 1**, and the descriptor gains `"secret_generation": 1` — with `token_id` and `fingerprint` **byte-identical**, i.e. exactly the re-key shape that used to be invisible → `10:39:34` the agent READS its credential through the wrapper (leg b — the read that was impossible until v0.91.0) → `10:39:38` **`ERROR pbsdr: the DR endpoint REJECTED this box's credential — the tier is applied and DEAD` `previous_state=applied`** (leg c: the exact R-39 failure state, detected out loud for the first time ever) → `10:39:45` **`one-time token secret consumed`** `secret_len=36` (leg a: **NO short-circuit** — this is the line that never appeared on 2026-07-18) → `10:39:45` `felhom-pbs-apply reconcile` (the set-only wrapper, no `--server`) → `10:39:47` **`pbsdr: converged state=applied`**. Corroboration: the agent marker hash moved to `afbb3b41…` (it was byte-identical to the pre-reissue marker in the failure); `consumed_at` stamped `08:39:45Z`; the on-disk secret's mtime moved `2026-07-18 20:28:52` → `2026-07-21 10:39:45`; a live probe with the NEW credential returns **200**; three consecutive hub reports trace the whole state machine `applied → auth_failed → applied`; and **zero** `pbsdr_selfheal` escalations fired — the box healed through the descriptor path before the damper was ever needed, with exactly ONE mint and ONE consume and no `consumed-failed.json`. **Row upgraded to PROVEN-LIVE (2026-07-21).** Evidence: `felhom-agent/REPORT.md` (2026-07-21). | -| **A customer's file survives a machine rebuild and comes back — the whole off-site story, end to end** | hub v0.94.0, agent v0.125.0, controller v0.197.0 | **PROVEN-LIVE (2026-08-04 night drill)** | demo-hp's controller data volume was destroyed and the sentinel deleted from disk. The customer's recovery code then produced `8a9e33aa4da6…` — **byte-identical** to the pre-wipe on-disk key and to the hub's independent record — and installed cleanly on the bare box. `identity_blob` was unchanged across the wipe (572 B, `updated_at` still 11:11:37). Evidence: `audits/DRILL-r201-night-run-2026-08-04.md` §2 | **WHAT IT DOES AND DOES NOT CLAIM.** PROVEN: after a real rebuild the key recovers byte-identical, the EXISTING repository opens (3 snapshots, 42 026 B — the pre-wipe size; not a fresh history), and the customer restore flow returns the file byte-identical. **ALL FOUR MANUAL INTERVENTIONS ARE CLOSED, AND THE CUSTOMER IS NOW OFFERED THE RECOVERY (2026-08-05). ONE QUALIFIER REMAINS AND IT IS NOT THE OLD ONE — state it precisely.** The 2026-08-04 drill needed four interventions between "the key is recoverable" and "the file is back", none of them in any design document (R-204). **Items 1–3 closed in controller v0.198.0 + hub v0.95.0** (the reset code works first time; a Re-issue no longer marks a healthy escrow stale; a unit restore states it returned the app's definition and database and NOT the customer's files). **Item 4 closed in controller v0.199.0 + hub v0.96.0** (a rebuilt box DECLARES `offsite.state=needs_credential` and `internal/offsiteheal` re-arms the stored credential before minting). **AND THE "needs someone who knows to look" HALF IS NOW GONE TOO (controller v0.200.0, R-193):** a full-page recovery screen takes over the landing pages while the hub holds a sealed package this box cannot open, explains that nobody can replace a lost recovery code, takes the code, opens the repository and lists what is in it — proven live on demo-felhom, which is genuinely in that state (`/launcher` → 302 `/recovery`; a wrong code refused with nothing written, three times, no lockout; the code found in no file, no container log and no debug ring, with a planted-copy positive control that first exposed a mis-aimed sweep). **WHAT THE QUALIFIER IS NOW, and it is narrower:** (1) **the final unlock has never been driven with a CORRECT code through the page** — no recovery code was kept for demo-felhom's orphaned history and demo-hp's is operator-held out of band, so the live run exercised the whole chain (handler → agent → hub fetch → age KDF) and stopped at the unseal; the install and listing rest on tests. (2) **putting files back in place is deliberately NOT part of this** — restore stays per-app, and the step after the listing is R-213. (3) **the journey has still not been re-walked end to end since these fixes** — the closures are proven individually, not as one uninterrupted run. That re-walk is one more drill and it is what is owed. Scope: `demo-hp`, a controller-data-volume rebuild — **NOT** a total host loss, and **NOT** a guest reprovision. **⚠ THE RE-WALK RAN ON 2026-08-05 (CAMPAIGN-11) AND IT FAILED. THIS ROW STAYS AS IT IS — it records a key recovering and a file returning, which remains true — but the JOURNEY row below is the honest verdict and this row must not be read as covering it.** The campaign built a throwaway appliance from the published ISO, destroyed it, and walked the customer's route: **the data half PASSED** (three sentinels byte-identical, including a 12 MB binary and a non-ASCII Hungarian filename, out of the pre-wipe snapshot in 16 s) and **the journey half FAILED** — four operator interventions, three needing root on the appliance, and the first of them was the machine telling the customer their perfectly correct recovery code was wrong (R-216). **R-198's retention is NO LONGER unit-proven — it was PROVEN IN PRODUCTION on 2026-08-05** (CAMPAIGN-11 Phase 3): the first supersession since the fix retained `identity_blob` at **572 B, byte-length exact**, carrying the OLD key `626e424670248db3` while the current row moved to the newly minted `e11a6c542b73477a`; `offsite_repo_key_changed` and `escrow_superseded` both fired at the instant of supersession, and the next off-site run REFUSED as `orphaned` rather than starting a fresh history. Evidence: `tests/campaign11-evidence-2026-08-05/journal.md` **NARROWED 2026-08-21 by the backup-truth drill — the claim holds for the leg it was proven on and NOT for the others.** Proven that night on `demo-hp` with planted, hash-recorded files: the **declared-userdata** leg of a **drive-declaring** app does come back byte-identical (`calibre-web`, 5/5 including two Hungarian accented filenames). **Two legs of the same story do NOT:** (a) the off-site restore has **no named-volume leg at all**, so an app's volume tar sits in the unit, in the snapshot and in the checking folder and is never replayed (**R-354**); (b) for the **40 of 53** apps that declare no data drive the off-site restore **refuses outright**, saying a running app „nincs telepítve" (**R-356**). Since the 40-class keeps ALL its data in named volumes, the end-to-end story is **unproven for that class and disproven for the volume leg generally**. The escrow/key half of this row is untouched by that and still stands. Evidence: `audits/DRILL-backup-truth-2026-08-21/evidence/` and `REPORT.md` (2026-08-21). | +| **A customer's file survives a machine rebuild and comes back — the whole off-site story, end to end** | hub v0.94.0, agent v0.125.0, controller v0.197.0 | **PROVEN-LIVE (2026-08-04 night drill)** | demo-hp's controller data volume was destroyed and the sentinel deleted from disk. The customer's recovery code then produced `8a9e33aa4da6…` — **byte-identical** to the pre-wipe on-disk key and to the hub's independent record — and installed cleanly on the bare box. `identity_blob` was unchanged across the wipe (572 B, `updated_at` still 11:11:37). Evidence: `audits/DRILL-r201-night-run-2026-08-04.md` §2 | **WHAT IT DOES AND DOES NOT CLAIM.** PROVEN: after a real rebuild the key recovers byte-identical, the EXISTING repository opens (3 snapshots, 42 026 B — the pre-wipe size; not a fresh history), and the customer restore flow returns the file byte-identical. **ALL FOUR MANUAL INTERVENTIONS ARE CLOSED, AND THE CUSTOMER IS NOW OFFERED THE RECOVERY (2026-08-05). ONE QUALIFIER REMAINS AND IT IS NOT THE OLD ONE — state it precisely.** The 2026-08-04 drill needed four interventions between "the key is recoverable" and "the file is back", none of them in any design document (R-204). **Items 1–3 closed in controller v0.198.0 + hub v0.95.0** (the reset code works first time; a Re-issue no longer marks a healthy escrow stale; a unit restore states it returned the app's definition and database and NOT the customer's files). **Item 4 closed in controller v0.199.0 + hub v0.96.0** (a rebuilt box DECLARES `offsite.state=needs_credential` and `internal/offsiteheal` re-arms the stored credential before minting). **AND THE "needs someone who knows to look" HALF IS NOW GONE TOO (controller v0.200.0, R-193):** a full-page recovery screen takes over the landing pages while the hub holds a sealed package this box cannot open, explains that nobody can replace a lost recovery code, takes the code, opens the repository and lists what is in it — proven live on demo-felhom, which is genuinely in that state (`/launcher` → 302 `/recovery`; a wrong code refused with nothing written, three times, no lockout; the code found in no file, no container log and no debug ring, with a planted-copy positive control that first exposed a mis-aimed sweep). **WHAT THE QUALIFIER IS NOW, and it is narrower:** (1) **the final unlock has never been driven with a CORRECT code through the page** — no recovery code was kept for demo-felhom's orphaned history and demo-hp's is operator-held out of band, so the live run exercised the whole chain (handler → agent → hub fetch → age KDF) and stopped at the unseal; the install and listing rest on tests. (2) **putting files back in place is deliberately NOT part of this** — restore stays per-app, and the step after the listing is R-213. (3) **the journey has still not been re-walked end to end since these fixes** — the closures are proven individually, not as one uninterrupted run. That re-walk is one more drill and it is what is owed. Scope: `demo-hp`, a controller-data-volume rebuild — **NOT** a total host loss, and **NOT** a guest reprovision. **⚠ THE RE-WALK RAN ON 2026-08-05 (CAMPAIGN-11) AND IT FAILED. THIS ROW STAYS AS IT IS — it records a key recovering and a file returning, which remains true — but the JOURNEY row below is the honest verdict and this row must not be read as covering it.** The campaign built a throwaway appliance from the published ISO, destroyed it, and walked the customer's route: **the data half PASSED** (three sentinels byte-identical, including a 12 MB binary and a non-ASCII Hungarian filename, out of the pre-wipe snapshot in 16 s) and **the journey half FAILED** — four operator interventions, three needing root on the appliance, and the first of them was the machine telling the customer their perfectly correct recovery code was wrong (R-216). **R-198's retention is NO LONGER unit-proven — it was PROVEN IN PRODUCTION on 2026-08-05** (CAMPAIGN-11 Phase 3): the first supersession since the fix retained `identity_blob` at **572 B, byte-length exact**, carrying the OLD key `626e424670248db3` while the current row moved to the newly minted `e11a6c542b73477a`; `offsite_repo_key_changed` and `escrow_superseded` both fired at the instant of supersession, and the next off-site run REFUSED as `orphaned` rather than starting a fresh history. Evidence: `tests/campaign11-evidence-2026-08-05/journal.md` **RE-WIDENED 2026-08-22, and the narrowing below is now HISTORY — read both, in order.** Both defects named in the 2026-08-21 narrowing are closed and proven live: **R-354** (controller v0.218.0, `volReplay`) and **R-356** (controller v0.219.0). The **40-class end-to-end story is now WALKED**, including the hardest ten of it: `audits/DRILL-r356-hot-only-restore-2026-08-22/` proved a driveless app with NO database (`privatebin`: planted, off-sited, deleted, restored, **15/15 files byte-identical**, two Hungarian accented names), and `audits/DRILL-r356b-driveless-db-restore-2026-08-22/` proved a driveless app **WITH** a database on both engines — `docmost` (Postgres 16) and `bookstack` (MariaDB 12.3), each planted through the app's own interface, destroyed for real, and returned with accented names byte-identical. **Five legs that had never run in any combination all ran and all succeeded:** the undo copy, DB-service identification, the volume replay, the DB-only start window, and the dump replay on top. **A second claim was walked at the same time:** R-164's F17 ordering — *the logical dump wins over the volume tar's copy of the same database* — was recorded only for the LOCAL path (`restore_unit.go:262-266`) and is now measured on the **off-site** path too, by a three-way discriminator (volume tar `ORIGINAL-VALUE-A`, altered dump `ALTERED-VALUE-B`, live `LIVE-VALUE-C3`; result **`ALTERED-VALUE-B`**). **WHAT IS STILL NOT CLAIMED:** the FAILURE path is where this class is weak, not the success path — a corrupt dump leaves the customer with an emptied or partially-applied database and an undo copy **no product action can apply** (**R-379**), and on MariaDB it does so behind an app that reports `health=healthy` (**R-380**). The success story is proven; the recovery-from-a-bad-restore story is not. + +**NARROWED 2026-08-21 by the backup-truth drill — kept as history; both halves have since closed, see the entry immediately above.** Proven that night on `demo-hp` with planted, hash-recorded files: the **declared-userdata** leg of a **drive-declaring** app does come back byte-identical (`calibre-web`, 5/5 including two Hungarian accented filenames). **Two legs of the same story do NOT:** (a) the off-site restore has **no named-volume leg at all**, so an app's volume tar sits in the unit, in the snapshot and in the checking folder and is never replayed (**R-354**); (b) for the **40 of 53** apps that declare no data drive the off-site restore **refuses outright**, saying a running app „nincs telepítve" (**R-356**). Since the 40-class keeps ALL its data in named volumes, the end-to-end story is **unproven for that class and disproven for the volume leg generally**. The escrow/key half of this row is untouched by that and still stands. Evidence: `audits/DRILL-backup-truth-2026-08-21/evidence/` and `REPORT.md` (2026-08-21). | | **The customer's own UNAIDED recovery journey, end to end** | controller v0.206.0, hub v0.98.0, agent v0.127.0 | **PROVEN-LIVE (2026-08-07, the fifth walk) — an unaided customer journey on a real installation, SCOPED: see what it does not claim** | A throwaway appliance was built from the published ISO on demo-hp, claimed, given three sentinels, escrowed, backed up off-site, then destroyed — guest purged and both enrolled drives wiped. The rebuild and the customer's route were then walked with **no command line inside the guest**. **Data: PASS.** All three sentinels byte-identical (`beb9175d…`, `7c8cb0ad…`, `e012e76f…`), including a 12 MB binary and `C11-őrszem-ékezetes-árvíztűrő.txt` — the Hungarian filename survived disk → restic → SFTP → Storage Box → restore → disk intact. Restored in **16 s** out of the pre-wipe snapshot `f3d9cd67`, non-destructively. **Journey: FAIL.** Four interventions, three of them root on the appliance: R-216 (a correct code called wrong), R-218 (the store never opened; the box stopped asking), R-219 (the promised listing structurally could not render), R-220 (the app could not be redeployed — its drives were unenrollable). Evidence: `tests/campaign11-evidence-2026-08-05/journal.md` | **WHAT IT DOES NOT CLAIM, and what the v0.201.0/v0.97.1 fixes do and do not change.** Six of the nine findings are fixed and each is red-proofed (R-216, R-218, R-219, R-217, R-222, R-215) — but **fixes are not a journey.** This row goes green only when the walk is repeated end to end and completes with no operator intervention. **Three findings are deliberately still open and each is a live blocker for some flow:** **R-214** (the console never stops showing a stale pairing code), **R-220** (after a rebuild the customer's drives cannot be re-enrolled — the deploy refuses and the wizard's list is empty, so no app can be redeployed onto its own data drive), **R-221** (a rebuilt box cannot run the escrow ceremony at all, because the pbsdr convergence marker survives on the host while the seeded `escrow.pbs_storage_id` is rewritten out of `agent.json`). **And R-216's fix does not by itself make a NEW box work:** with the hub guard corrected but the Day-0 manifest still vouching agent 0.120.0, a new box is **HELD, not served** — it stops being lied to, which is better, but the feature does not work for it until agent 0.125.0 is vouched. **⚠ CAMPAIGN-11 PHASES 2 + 4 RAN OVERNIGHT 2026-08-05/06 AND THIS ROW STAYS **FAIL**.** **What they add:** eleven injected faults against the recovery journey, each with a positive control, and a full unattended scheduled cycle. **The backup promise strengthened** — the off-site tier ran itself at 04:15 on a box rebuilt twice and set aside hours earlier (`snapshot_count` 1 → 2), all five daily jobs fired exactly once, and nothing on the must-not list fired. **R-217's and R-215's fixes were proven live under exactly their faults**, the set-aside proved it does not delete (12 535 KB byte-exact at the far end), and a wrong code was refused three times with nothing written and no lockout. **What they do NOT add: any progress on the JOURNEY.** These faults are **not a re-walk** — no customer route was walked end to end — and four new findings say the journey got no better where it matters: **R-224** (a hub outage and a stopped agent are BOTH reported to the customer as a bad recovery code, in 0.056 s and 0.030 s — no unseal attempted; the agent's own `err` distinguishes them and it is discarded at the HTTP boundary, and R-216's gate answers `source=version` so it cannot see reachability — **Phase 1's headline defect relocated from the version channel to the transport**), **R-226** (M1, the only message telling a customer to check their typing, is unreachable on any box that has re-escrowed), **R-225** (an unread store renders `0 pillanatkép · 0 GB` above a card saying it holds backups — 1 snapshot and 12 535 KB were really there), **R-228** (the set-aside history is recorded in `orphaned_renamed_to` and shown by nobody — `OrphanedRenamedTo` has zero references in any template or handler). **§4.1 is now MEASURED rather than deduced** — the floor IS served, from the box's own rendered `GetFloor()` and a cold-started controller's settle-gate line. **§4.2's positive half is still NOT measured** and needs a rebuild. Evidence: `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md`, `tests/campaign11-evidence-2026-08-05/journal-phase24.md` **2026-08-06 — THE FIVE PHASE-2 FINDINGS ARE FIXED AND THIS ROW STILL SAYS FAIL.** controller **v0.202.0** + agent **v0.126.0** close R-224, R-226, R-225, R-227, R-228. **What changed:** the unlock path now classifies WHY it failed, from the value and never the text, and the message that mentions typing is reachable from exactly one class — a real refusal — with **unknown defaulting to neutral**. Proven live on the venue with the same wrong code and only the hub's reachability changed: `400 → 502 → 400`. **What it does NOT change: this row.** Fixes are not a journey; nothing here walked a customer end to end, and three of the blockers that made Phase 1 fail are still open (**R-220** in particular is worked around BY HAND on the venue — without that unmount no app can be deployed on a rebuilt box). **What still needs proving:** a re-walk with no operator intervention, and specifically the customer-facing messages end-to-end — they were NOT re-driven, because `/recovery` correctly retires itself once the old data is set aside and restoring that state would be the reconfiguration the fix task forbade. Evidence: `felhom-controller/REPORT.md` **⚠ THE RE-WALK RAN 2026-08-06 (R-201, attended) AND THIS ROW STAYS FAIL — but it is a nearer no.** A brand-new appliance was built from the published ISO, given three sentinels, escrowed, backed up off-site, destroyed (guest purged AND both drives wiped), reinstalled and walked. **THE DATA: PASS** — all three sentinels byte-identical, including a 12 MB binary and an accented Hungarian filename **whose NAME BYTES are also identical**, restored in **23 s** out of the pre-destruction snapshot through the customer's own flow. **THE JOURNEY: FAIL — TWO dead ends against Phase 1's four.** (1) **R-218's CONSUME half**: the hub re-staged the credential and said "the box re-consumes on its next cycle"; a full cycle ran (positive control) and it did not — no customer-reachable action fetches it, and the only lever is a command line INSIDE THE GUEST. (2) **R-220**: the drives are still unenrollable after a rebuild, needing a Proxmox-host unmount, without which no app can be redeployed and the restore page stays empty. **The unaided RTO is therefore STILL UNDEFINED**; the attended figure was 30 m 13 s and must not be quoted as the customer number. **What passed and is new:** the recovery screen appeared WITHOUT being sought, answered all three of its questions with a seal date matching the hub exactly, the emailed reset code worked first try, and the unlock was a real 1.528 s unseal that placed the key. **⚠ AND IT DOES NOT CLAIM DELIVERY:** the fresh install landed on controller 0.201.0 / agent 0.125.0 — the vouched versions, neither carrying the fixes — which were installed BY HAND. Fleet delivery needs a golden carrying the fixed controller and a vouched agent, and nothing was vouched. Evidence: `tests/rewalk-r201-2026-08-06/journal.md`   **⚠ 2026-08-06 (later the same day) — THE TWO DEAD ENDS ARE CLOSED AND VOUCHED, AND THIS ROW STILL SAYS FAIL.** controller **v0.203.0** + agent **v0.127.0** close **R-218's consume half** and **R-220**, and golden **0.203.0** / agent **0.127.0** / `min_agent` **0.127.0** are now VOUCHED — so the delivery gap the re-walk recorded (“the fresh install landed on versions neither carrying the fixes”) is gone. **Both were then proven live on a genuinely rebuilt box** (Part 4 venue, VM 323): the reinstall landed on agent 0.127.0 with **no downgrade and no hand upgrade** (the previous re-walk's reinstall downgraded 0.126.0→0.125.0), `/disks/candidates` returned **both drives** after the guest was purged with the raw mounts still on the surviving host — before the fix, two empty lists — and both **re-attached through the customer endpoint**; the off-site credential was re-staged by `offsiteheal` at 13:24:57Z and **collected by the box on a tick, unaided**. **Why the row is still FAIL:** the walk did not finish. It stopped at the restore surface — **R-237**, the list was keyed on apps that are currently INSTALLED and currently TOGGLED ON for future backups, so a rebuilt box was shown nothing to restore while the repository held its snapshots. Fixed in **controller v0.204.0** (the store is now the source of the list) and proven live on that same venue with the toggle switched OFF, but **no sentinel was restored in that walk**, so the data half is unproven in EITHER direction for this venue. **R-238 was RECLASSIFIED, not fixed as filed:** `mode=full` without `confirm=1` is step 1 of a deliberate two-step that starts no job by design; the endpoint driver did not carry `?full_prep=` forward. The operator drove the same restore to completion in a browser. Its real residue — that step logging nothing at all, including when the headroom gate refuses — is fixed. **R-236 was WITHDRAWN: not a defect.** **This row goes green only when a walk completes end to end with no operator intervention, restoring a byte-identical sentinel.** Evidence: `tests/part4-rewalk-2026-08-06/journal.md`, `tests/teardown-2026-08-06.md`, `felhom-controller/REPORT.md` **⚠ THE FINAL WALK RAN OVERNIGHT 2026-08-06/07 AND THIS ROW STAYS FAIL — the data half passed again, the journey half failed further from the line than before.** A brand-new appliance was built from the published ISO, fixtured, soaked through a real scheduled cycle, destroyed and rebuilt. **THE DATA: PASS** — all three sentinels byte-identical out of snapshot `f5c53b03`, including a 12 MB binary and an accented Hungarian filename **whose name BYTES are identical too**. **THE JOURNEY: FAIL, and this time the customer has NO route at all** — `/` lands on the launcher with no recovery pointer, `/recovery` **302s away**, and the remote-backup page offers to **CREATE a new recovery code**, which would orphan the very history the customer's code protects. **There is no field anywhere to enter the code they hold**, and the operator's documented remedy refuses too. Cause (**R-241**): the credential self-heal — proven working unaided in the same walk, six hours earlier, as its best result — writes a FRESH repository password, which moves the box out of `OffsiteRecoveryOffer()`'s (a) pristine case while (b) orphan-detected is unreachable because runs are blocked by `escrow_state: pending`. **R-218's shape one level up.** **What DID pass, and is new:** the whole credential chain ran end to end with zero human action on an unclaimed box — declare, `offsiteheal` re-stages after two reports, the box's 5-minute retry collects it, tier applied — the **first live sighting** of that success line, settling R-218's consume half and R-236's withdrawal. Also new: **R-239** — a fresh install lands on controller **0.203.0** while 0.205.0 is released, so R-234 and R-237 are written but **not delivered**, proven from the customer's side (T2, T3). **This row goes green only when a walk completes with no operator intervention AND a byte-identical sentinel.** Evidence: `tests/finalwalk-r201-2026-08-07/journal.md` **⚠ 2026-08-07 — R-241 WAS DIAGNOSED BY A READ-ONLY SPIKE ON THE STANDING VENUE, AND THE DIAGNOSIS REVERSES THE FIX. This row is unchanged: still FAIL, and no code was written.** The question was whether R-241 is a screen-predicate defect or a minting defect. **It is a MINTING defect.** The screen was telling the truth — there genuinely was nothing recoverable under the key the box held, because the box minted that key itself over the top of a sealed package it already knew the hub was holding. **Three measurements, from the venue rather than from the earlier report:** (1) `WriteOffboxSecrets` (`offbox.go:411`) mints on ONE input — does the file exist — while `OffsiteRecoveryOffer()` and `needsOffsiteCredential()`, both in the same file, consult `GetHubEscrowIdentityPresent()`; the same fact is available on three paths and used on two. (2) That flag was the **precondition of the chain that reached the minting**: the retry job logs only when the declaration is live, and the venue logged `credential retry: … (the box still declares a need; retrying)` at **02:48:03Z** — thirty minutes and six ticks before the mint at 03:18:06Z. (3) **The box computed the answer and discarded it**: at **03:28:03Z**, thirty-five minutes before the customer looked, `EscrowAutoConfirmer.Reconcile` (`escrow_confirm.go:154`) logged `the hub's escrow blob does not cover the CURRENT repo password (hub hash 30ef574fe492… != local 9b4a9a9dcec7…)`. It is recomputed every report cycle and never persisted. **And the hub explicitly disclaims doing this** — `offsiteheal`'s package doc: *"credential automatic, key customer-present — the ruling this session implements and must not quietly widen."* **Shape (b) is structurally unreachable on this box** (the escrow gate at `offbox.go:743` sits upstream of `ensureOffboxRepo`, the only producer of `RepoState="orphaned"`; positive control: the scheduler was alive, 241 `agent-channel-health` ticks, and `offbox-backup` is a `sched.Daily` leg whose slot fell before the destruction). **Two by-products:** **R-243** — a box in this state silently stops backing up off-site and **no alarm fires** (`isStale` requires `escrowed`, `offsite_delivery_stuck` skips the `applied` shape, `backup_failed` needs a run that never happens); and the trap in the obvious fix — `ResetOrphanedRepo` clears the orphan flag without a ceremony, so a hash-mismatch discriminator alone would re-offer the screen forever to a customer who declined the old data. **Q7:** the „Helyreállítási kód létrehozása" button does NOT destroy the data — R-198's retention copies the `identity_blob` — but it converts a self-service recovery into one needing an unbuilt read path (R-199), and it re-enables the screen while invalidating the code that screen accepts. Evidence: `audits/SPIKE-r241-recovery-offer-2026-08-07.md` **⚠ 2026-08-07 (later) — R-241 IS FIXED (controller v0.206.0 + hub v0.98.0) AND THIS ROW STILL SAYS FAIL.** Three changes, following the spike's ruling rather than the obvious reading: the box **no longer mints a repository key while the hub holds a sealed package** (a conjunction, so a first-time box is untouched; the refusal is a HOLDING state that still writes the transport, declared as `offsite.state=awaiting_recovery_key` and shown inert to every existing hub reader); the **hub-vs-local key comparison that was computed every cycle and discarded is now persisted** and drives the offer as **shape (c)**; and **abandoning starts a 14-day countdown** whose terminal step removes the set-aside store and the sealed package **together**, so the offer ends because the state is right rather than because a flag suppresses it. Surface: the full page appears **once per ENTRY into the offered state**, three dismissal levers with three scopes, and **none removes the entry point**. Q7's trap is closed — „Helyreállítási kód létrehozása" is **unavailable** while a recovery is outstanding. **WHY THE ROW STAYS FAIL: these are fixes, not a walk.** Nothing here walked a customer end to end, and this row goes green only when one completes with **no operator intervention AND a byte-identical sentinel**. Two of Phase 1's blockers are also still open (**R-214**, **R-202**), and **R-240** is untouched. Evidence: `felhom-controller/CHANGELOG.md` v0.206.0, `hub/CHANGELOG.md` v0.98.0. **DELIVERY, separately: R-239 is CLOSED 2026-08-07** — golden **0.205.0** baked, published, round-trip verified (`./etc/felhom-controller-image` read OUT of the downloaded archive) and **VOUCHED**, so a fresh install now lands on 0.205.0 carrying R-234 and R-237. In the event it was a ONE-field change: `agent_version` and `min_agent` both stayed 0.127.0. Evidence: `tests/golden-0.205.0-2026-08-07/`. **The row still says FAIL** — delivery is not a journey, and R-241 is diagnosed, not fixed. **✅ 2026-08-07 — THE FIFTH WALK PASSED, BOTH HALVES, AND THIS ROW TURNS.** A brand-new appliance (VM **325**, customer `walk5`) was installed from the published ISO on `demo-hp`, claimed, given three sentinels, escrowed, backed up off-site, then **destroyed on purpose** — guest purged, both data drives wiped — and rebuilt through the documented day-0 path. **THE DATA: PASS** — all three sentinels byte-identical out of snapshot `5b0f20f7` (`11eb7fb2…`, `6b504d1e…`, `0baaf402…`), including a 12 MB binary and `WALK5-őrszem-ékezetes-árvíztűrő.txt` **whose name BYTES are identical too**, read back with `os.listdir` on a bytes path so no decode round trip could launder a `U+FFFD`. **THE JOURNEY: PASS — ZERO guest command lines were needed to progress**, against three on the previous walk. **RTO 71.7 s** from login to an open store (12.44 s of it the unseal; ~22 s a harness retry). The recovery screen **appeared without being sought** (`/` → `/launcher` → `/recovery`), answered all three of its questions, and its sealed-at timestamp matched `host_escrow.created_at` exactly. **WHAT MADE THE DIFFERENCE — R-241's mint guard, exercised live for the first time:** at **14:58:52Z**, unaided and before anyone logged in, the rebuilt box collected its re-staged credential, configured the transport and **refused to mint a repository password** over the sealed package — `NOT minting a repository password: the hub holds a sealed recovery package for this box`. Sampled every 20 s from T0: **no key at any moment**, with the scheduler proven alive throughout. At the equivalent moment the previous walk minted one and lost the journey silently. **AND DELIVERY IS PART OF THE PASS:** the fresh install landed on **controller 0.206.0 + agent 0.127.0 — the vouched set, no hand upgrade**, both times, so the box under test is the box a customer receives. **WHAT THIS ROW STILL DOES NOT CLAIM.** (1) **Putting files back in place is built and worked here** — `reconstitute` placed 6 files — but **only after two obstacles the customer must guess past**: **R-252** (the restore refuses with „nincs elérhető adatmeghajtó" because a rebuild loses the drive *registration*, and nothing on the recovery path says to re-attach) and **R-253** (the restore refuses because the app is not installed, on a page that says three lines above that the restore reinstalls it). Both were cleared **from the dashboard with no shell** — which is why the journey passes — but neither is signposted, so *unaided* here means *possible without a shell*, not *obvious*. (2) **Shape (c) did NOT fire positively.** With the mint guard holding there is no local key, so the offer comes from **shape (a)**; shape (c) was measured in Phase A in its **negative** half (hub hash == local hash, correctly silent). The mint guard is proven positively; the discriminator only negatively. (3) **R-214, R-202 and R-240 are untouched.** (4) The venue is **torn down** (2026-08-08) — full-schema census 168 rows → 67, of which 37 are audit-by-design and **30 are R-244**, predicted before the run rather than found after; every layer verified absent against a surviving positive control, and **16.64 GiB** returned. `tests/walk5-r201-2026-08-07/teardown-walk5-2026-08-08.md`. Evidence: `tests/walk5-r201-2026-08-07/journal.md` | | **Off-site app-data capture covers the paths an app declares MANDATORY — on BOTH drive layouts** | controller v0.197.0 | **PROVEN-LIVE (2026-08-04) — and this row was OPTIMISTIC before it** | On demo-hp, an app on the **system-data fallback** bound `/mnt/sys_drive/userdata/media/books` while the capture set looked in `/mnt/sys_drive/felhom-data/userdata/media/books`: the declared-mandatory directory was in **no** snapshot and the run reported `ok` (R-203). After v0.197.0 the capture log reads `1 mandatory path(s)` and the file is **listed inside the snapshot** — `restic ls -l latest --tag calibre-web` → `-rw-r--r-- 1000 1000 181 … /DRILL-SENTINEL.txt`. Evidence: `audits/DRILL-r201-offsite-recovery-2026-08-04.md` §2 and the v0.197.0 CHANGELOG | **WHAT IT DOES NOT CLAIM.** It covers CAPTURE, not RESTORE: **no file has ever been restored from an off-site backup after a wipe** (R-201, ready to resume). It also does not claim the enrolled-drive layout was ever wrong — it was not, which is precisely why this survived. And a run that still misses a mandatory directory now reports **`incomplete`** rather than `ok`, so this row's guarantee is one the status can express. **Widened 2026-08-06 (controller v0.205.0, R-234):** the same verdict now also covers an app skipped ENTIRELY — until then a missing declared FOLDER made the run incomplete while an app with no recovery unit at all still reported `ok`, so the smaller gap moved the verdict and the bigger one did not. **This row still claims CAPTURE, not that a newly-selected app is protected by the next run** — for a DEPLOYED app it is (the run's own pre-dump phase writes the unit, measured 2026-08-06), and for an undeployed one it is not and the card now says so || **Recurring offsite (PBS) whole-guest backups actually LAND, and RESTORE** — local daily + offsite weekly as scheduled work | agent v0.97–0.103, controller v0.174/0.175, hub v0.76.0, host-install 1.20.0 | **PROVEN-LIVE (2026-07-26)** | `audits/SPIKE-r82-phase0-2026-07-26.md`; per-repo CHANGELOGs/REPORTs. **Restore round-trip on demo-hp:** `--selftest=restore-test` against `felhom-pbs:backup/ct/9201/2026-07-26T15:42:42Z` → `pass:true`, `verified:"boot+running"`, **`mount_parity:"ok"`** (`mp0=/var/lib/docker 50G`, `mp1=/mnt/sys_drive 20G`, mp8/mp9 throwaway stand-ins for the archived binds), `source_tier:"pbs"`, 4m5s restore+boot+verify+teardown, scratch band clean afterwards and the live guest untouched. | **This row is distinct from the DR-tier row above, which proves ACTIVATION, not ARRIVAL.** That tier was PROVEN-LIVE as `applied` since 2026-07-21 while demo-felhom held ONE snapshot (a healing artifact) and demo-hp held **zero, ever** — "applied and empty", the R-39 shape one level quieter. **What earns PROVEN-LIVE here:** (a) demo-hp's FIRST EVER offsite backup landed (4.25 GB into a verifiably empty namespace); (b) it **restores into a bootable, mount-complete guest** — `mount_parity` is the non-hollow half, since a boot-only verify cannot see a missing data volume; (c) **the multi-tier quiesce ran through the real UI endpoint** (`POST /api/guest-backup/trigger`, authed+CSRF) and produced **exactly ONE stop/start pair with BOTH backups inside it** — `quiescing 1 stack(s)` 17:01:39 → local done 17:02:56 *"next tier may start (app still quiesced)"* → felhom-pbs snapshotted 17:03:06 → `unquiescing` 17:03:06. **App downtime 1m27s for both tiers**, and the app came back healthy. **Known gaps, recorded not hidden:** the SCHEDULED restore-test still only selects the PRIMARY tier, so the offsite tier is never AUTOMATICALLY restore-tested (the manual/selftest path is proven, the unattended one is not); the hub infers "PBS ⇒ weekly" from storage TYPE rather than a reported cadence; the installer-default fleet flip awaits a full weekly cycle. → **R-82** | | **Restore-proof is UNATTENDED — the scheduler covers EVERY tier, follows the BACKUP rather than the clock, and a failure is heard** | agent v0.104.0 → **v0.121.0**, hub v0.77.0 → **v0.91.0** | **PROVEN-LIVE (2026-08-03)** | per-repo CHANGELOGs; `backlog/SPEC-r85-phase4-5-2026-07-26.md`. Unit red-proofs for tier rotation, restart-survival, the one-heavy-op gate, failure-emits-an-event, and newborn silence. | **The row above is earned by a MANUAL `--selftest=restore-test`; this one is about the SCHEDULED path, and the distinction is the whole point.** Before R-85 the scheduler could only ever see `cfg.Backup.BackupTarget()`, so the offsite tier was never a candidate — and a failed restore-test was a `[WARN]` line with no event at all, which was true for the LOCAL tier that WAS being tested. Now: oldest-first rotation across every configured tier (operator ruling 2026-07-26), persisted so it survives a restart; a restore-test joins the host-wide one-heavy-operation gate so it never contends with a backup over the same tunnel; and the hub raises two DISTINCT operator-tier signals — `restore_test_failed` (broken now) and `restore_test_stale` (unverified, not known-broken), anchored on R-81 so a newborn box never alarms. **Why this was NOT PROVEN-LIVE until now:** rotation had not been observed selecting both tiers across consecutive UNATTENDED cadences — at a 24h cadence a multi-day window — and a single passing run proves the code path, not the schedule. → **R-85**. **R-86 (agent v0.121.0 + hub v0.91.0, 2026-08-03) replaced the schedule and the observation became possible in one afternoon**, because what has to be observed is no longer a multi-day rotation but a RULE: a tier is due when its newest archive that has settled ~24 h has not been proven. **UPGRADE EVIDENCE — a real unattended run on demo-felhom (2026-08-03), triggered by DUE-NESS, not by a timer.** The scheduler's own log: `15:14:38 restore-test tier is DUE … target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z … reason="newest settled archive … has not been proven"` → `proxmox-backup-client restore --crypt-mode=encrypt` under the agent's own token → `15:25:08 gate decision class=guest_destroy guest=990000 allowed=true` → `15:25:14 scratch guest torn down` → **`15:25:14 backup: scheduled restore-test passed archive=felhom-pbs:… duration_s=635.1`**. A **14.5 GB encrypted offsite archive pulled from ep0 over the WAN**, restored, booted, verified and destroyed in 635 s, unattended. **The three things a timer could not show**, all verified after it: the state names THAT archive (`{"felhom-pbs":{"archive":"…2026-07-28T04:49:43Z","proven_at":"2026-08-03T13:25:14Z"}}`); a second evaluation reports `due=false … is already proven` and runs nothing; and **an agent restart runs nothing**, which is the defect a person actually noticed — every deploy used to restart the timer. Teardown verified at all three layers: guest absent from `pct list`, **zero** `990000` volumes in `lvs`, and the hub-side `restore_tests[]` entry deliberately RETAINED (it IS the proof the staleness check reads). **R-189 (agent v0.122.0, 2026-08-03) closed the reporting half of this row, and it was a REAL gap in the evidence path:** the proof above reached the hub only because no restart intervened — `restore_tests[]` came solely from an in-memory store, so the 15:25:14 PASS was in fact LOST when the agent restarted 2 m 43 s later for a deploy (`0 restore-tests` on the next two host-reports). Under per-archive due-ness the box would not have repeated the work for a week. The persisted per-tier proof (with the archive, and now the tier) is merged into the report, so a proof survives a restart — one entry per tier, newest wins, and a record that cannot be described honestly is not emitted. **THE HOST TIER'S HALF OF THIS ROW WAS OPTIMISTIC UNTIL 2026-08-03, and it is worth saying plainly.** Every live restore-test cited here is on the OFFSITE tier. The HOST tier was not merely unproven on demo-felhom — it was **unprovable**, and on demo-hp equally: the agent's token had no ACL on `/storage/felhom-backup`, the storage both boxes configure as `local_backup_target`, so the content API answered `{"data":[]}` through the token while root listed three archives. The scheduler skipped it as *"no settled archive yet"* — which is exactly what a brand-new tier reports — so nothing ever said so (**R-185**). **CLOSED 2026-08-03 — agent v0.123.0 + installer 1.24.0.** The grant is applied on both demo boxes (the token now lists 3 and 4 archives), the installer's reuse arm grants on a pre-existing target, and the agent now **asks whether it may read each tier it depends on** instead of inferring it from an empty list: a critical degraded capability naming the storage and the missing role, which the hub alerted and emailed on **while the box was still blind**. **THE HOST TIER IS NOW PROVEN-LIVE, UNATTENDED, ON BOTH DEMO BOXES (2026-08-04) — this row's optimistic half is cashed.** Four SCHEDULED runs, nothing triggered by hand: **demo-felhom** host tier `…2026_08_02-04_42_14.tar.zst` **passed in 83.8 s** at 00:55, offsite `…2026-07-28T04:49:43Z` **passed in 540.4 s** at 06:55; **demo-hp** host tier `…2026_08_02-04_49_29.tar.zst` **passed in 109.3 s** at 02:05, offsite `…2026-07-28T19:19:45Z` **passed in 300.1 s** at 08:05. Each restored into a scratch guest, booted, verified and destroyed itself; `pct list` and `lvs` show zero `990000` afterwards on both, and both boxes' `local-lvm` returned to their pre-run figures (1.95 % and 30.83 %). **Both boxes had BOTH tiers due at once and the ordering was observed live for the first time under R-86's rule:** never-proven sorted first, so each box took its HOST tier, deferred the offsite one, and picked it up on the following evaluation six hours later — one heavy operation at a time, per box, without anyone sequencing it. **The proofs reached the hub**, which is R-189 carrying a host-tier entry for the first time: demo-felhom's report holds **two** entries, one per tier, and the `local` one can only have come from the persisted state because the in-memory store held only that morning's offsite run. **SCOPE, stated because one box proving something does not make it a fleet property:** this covers `demo-felhom` and `demo-hp`. The tester's box is untested and untouched. Evidence: `felhom-agent/REPORT.md`, `felhom.eu/REPORT.md` | diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/README.md b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/README.md new file mode 100644 index 00000000..dfcaa912 --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/README.md @@ -0,0 +1,190 @@ +# DRILL — R-356b: the off-site restore for a driveless app that HAS a database (2026-08-22) + +**This was a drill, not an implementation. No production code was written, no version bumped, no +CHANGELOG entry made.** The output is this document and four register rows. + +**Subject:** `demo-hp` (Tier 0, disposable), guest 9201, controller **v0.219.0**, floor **0.219.0**. +**Method:** endpoint-level. `claude-in-chrome` is not available on DooPlex, so every product action was +the exact HTTP endpoint the UI's own form posts to — `/api/stacks//deploy`, +`/backup/offbox/{toggle,run,restore,reconstitute}` — driven with the session cookie and the session +CSRF token scraped from the page. App content was planted through each app's **own** interface +(docmost's JSON API; BookStack's login form + `POST /books`). + +## 1. Baselines, confirmed at drill start + +| repo | HEAD | note | +|---|---|---| +| felhom-controller | `0f3cf0dbb2c0` | v0.219.0 | +| felhom-agent | `40d857b52711` | v0.130.0 | +| felhom.eu | `8c9f1b798b12` | ahead of the runbook's hash — this session's own golden-bake commit | +| app-catalog | `459766cb1639` | | + +Read from the **hub** (`/configuration`, Basic auth, ClusterIP): `golden_version` **0.219.0**, +`agent_version` **0.130.0**, `min_agent` **0.129.0**, global controller floor **0.219.0**. All four as +expected — the operator's three-field vouch and the floor raise both landed. + +**The 10-app count, measured in this session and not taken from the runbook: 10.** +Method: `needs_hdd: false` in `.felhom.yml` **and** a `postgres|mariadb|mysql` image in the compose. +Same ten names as the runbook: bookstack, calcom, claper, docmost, kimai, outline, rallly, +sparkyfitness, tandoor, zipline. Evidence `01`. + +**Architecture document read:** `documentation/architecture/07-backup-architecture.md` §6.1, §6.2 and +§6.3 **as corrected on 2026-08-22** — including the corrected Tier-3 row, the note that R-102 is *not* +closed, and the `[DESIGN]` paragraph on destination resolution. + +Register ceiling before this drill: **R-378**. + +## 2. The three questions + +### Q1 — Does it complete at all? **YES.** + +A driveless app with a database restores end to end, comes back healthy, and returns the planted data. + +Measured on `docmost` (Postgres 16). All five legs that had **never run in this combination** executed +in order and every one succeeded, in **32 seconds**: + +| leg | evidence | +|---|---| +| undo copy of the live DB, taken while up | `pre-restore-20260822T135827Z-docmost-postgres.sql`, 135 624 B | +| database-service identification | `[docmost-postgres]` | +| stop → place files → **replay volumes** | "Restored **3** Docker volume(s)" — incl. the 52 MB Postgres data directory | +| start **only** the DB service | `docker compose up -d docmost-postgres`, 0.4 s | +| replay the SQL dump on top | "Imported DB dump docmost-postgres.sql", 2 s | + +3 pages planted through docmost's own API → destroyed (`count(*) FROM pages` = **0**, no soft-deleted +rows either) → **3 pages back**, same ids. The discriminator went `ORIGINAL-VALUE-A` → +`DESTROYED-VALUE-C` → **`ORIGINAL-VALUE-A`**. + +Repeated on `bookstack` (MariaDB 12.3): all five legs, **25 seconds**, 2 volumes replayed including a +**161 MB** MariaDB data directory, the planted book back with its accented name byte-identical. + +**Accented round-trip, byte-for-byte:** +`c3817276c3ad7a74c5b172c5912074c3bc6bc3b67266c3ba72c3b367c3a970203220e28094205233353662` +("Árvíztűrő tükörfúrógép 2 — R356b", docmost) and +`C3817276C3AD7A74C5B172C591206BC3B66E79766573706F6C6320E28094205233353662` +("Árvíztűrő könyvespolc — R356b", bookstack). Both identical before and after. Non-ASCII never crossed +a shell chain: strings were embedded in python files inside the guest, or passed with +`--data-urlencode "name@file"`. + +> **A flaw in this drill's own method, recorded rather than hidden.** The FIRST accented title was +> planted through a shell argument chain that double-escaped it, so it was stored as the literal ASCII +> text `Árv…` rather than as accented bytes. It was caught by reading the stored value back **as +> hex** — the same discipline that makes the rest of these numbers worth anything. The pages +> round-tripped exactly either way, so Q1's answer stands, but the accented coverage did not, and was +> re-measured in Phase 1b. The faulty page is deliberately left in place as the record. + +### Q2 — Which leg returned the data? **The SQL dump.** + +A three-way discriminator was built so the answer could not be ambiguous: + +| holder | value | +|---|---| +| live database before the restore | `LIVE-VALUE-C3` — if it survived, neither leg wrote | +| the volume tar (52 MB Postgres data dir) | `ORIGINAL-VALUE-A` | +| the SQL dump, **altered in the scratch only** | `ALTERED-VALUE-B` | + +**Result: `ALTERED-VALUE-B`.** The dump won. The ordering the code comment asserts — +*"volumes FIRST, database after, so a logical .sql dump still wins over whatever copy of the same +database a volume tar happens to contain"* (`offbox_reconstitute.go`, the R-354 block) — **holds in +practice on the off-site path.** It was an intention; it is now an observation. + +Mutation asserted: dump sha256 `c5414f24…` → `834c32b1…`, zero occurrences of the original left, one of +the altered. Evidence `13`. + +**The fence held.** Only the prepared scratch was mutated. Proved afterwards by re-preparing a fresh +scratch from the same snapshot: sha256 back to `c5414f24…`, reading `ORIGINAL-VALUE-A`. Evidence `14`. + +> **This confirms an existing claim on a new path rather than establishing a new one.** **R-164** +> already records *"the dump is authoritative and replayed after the tar so it WINS (F17)"* — but cites +> `internal/backup/restore_unit.go:262-266`, the **local** restore. Q2 extends that to +> `offbox_reconstitute.go`, the **off-site** path, which had never been exercised for this app class. + +### Q3 — Does a failure in the database leg tell the truth? **Partly. Two defects.** + +Measured by truncating the dump in the scratch mid-statement, on both engines. + +| | Postgres (docmost) | MariaDB (bookstack) | +|---|---|---| +| customer sees a failure | **yes** — `alert alert-error`, "sikertelen" | **yes** | +| message names the undo copy | **yes**, by filename | **yes** | +| app afterwards | **crash-loops** — visibly broken | **`health=healthy, running=true, restarts=0`** | +| database afterwards | **emptied** — 43 tables, 0 rows in pages/users/spaces | **partially applied** — book and users intact, `migrations` **wiped to 0 rows** | +| undo copy valid? | **yes** — proven by applying it | **yes** — proven by applying it | +| product can apply the undo copy | **NO** | **NO** | + +## 3. Hypotheses — what fired and what did not + +- **H1 (schema collision).** **DID NOT FIRE.** The dump replayed cleanly onto a freshly-replaced + Postgres data directory in 2 s under `ON_ERROR_STOP=1`. This is the notable negative: the two legs do + not collide in practice for this app class. +- **H2 (replayed data directory will not start).** **DID NOT FIRE.** Both engines started healthy on a + data directory that had just been replaced wholesale from a tar. +- **H3 (MariaDB and Postgres do not fail the same way).** **FIRED — but not in the predicted shape.** + The prediction was a quiet partial success with no error. What actually happens is a **visible + error** *and* a **silently inconsistent database behind a healthy-looking app**. See R-380. +- **H4 (credentials).** **DID NOT FIRE.** `ImportDump` discovered `dbUser=docmost/dbName=docmost` from + the **live** container and they matched the restored directory, exactly as predicted — secrets are + not regenerated on reconstitution. + +## 4. Findings filed + +| id | severity | what | +|---|---|---| +| **R-379** | HIGH | The undo copy is valid, is named to the customer, and **no product action can apply it** | +| **R-380** | HIGH | A failed **MariaDB** replay leaves a partially-applied database behind a **healthy** app | +| **R-381** | MEDIUM | The failure message pastes raw engine stderr — **including customer database rows** — into the Hungarian customer surface | +| **R-382** | LOW | The reconstitution's summary log line omits the volume count it already has | + +**An independent reproduction of an existing row.** After the first reconstitution, docmost's +`db-dumps/` held **only** `pre-restore-*` files — the app's own `docmost-postgres.sql` was gone. That is +**R-361** reproducing on a second app, unprompted. Not re-filed; noted here as corroboration. + +## 5. Observations — noticed, recorded, NOT acted on + +- **A newly deployed app is not in the off-site set.** `docmost` and `bookstack` were both absent from + `app_backup` entirely until switched on by hand. This is plausibly deliberate (the switch is the + customer's), and it is **not filed as a defect** — but the consequence is that an app can be deployed, + run, and never leave the box, while the nightly run reports "backup OK". Stated so the next reader can + decide whether the default is right. +- **`restic check` was run by hand and passed** — `no errors were found`, 29 snapshots. Nothing in the + product runs it (**R-359**, already filed). Reproducing the invocation took three attempts because the + key file is `ssh_key`, not `id_offbox` as the arg-builder's variable naming suggests. +- **The `-db` container-suffix attribution is correct here.** `bookstack-db`'s dump landed in + `bookstack`'s own unit as `bookstack-mariadb.sql` — the R-355 failure shape does **not** reproduce, + because v0.218.0's compose-project attribution handles it. + +## 6. Phases run + +| phase | status | +|---|---| +| 1 — Postgres, does it complete | **run** | +| 1b — accented round-trip, re-measured after a method flaw | **run** | +| 2 — which leg won | **run** | +| 3 — failure path, Postgres | **run** | +| 4 — MariaDB: Phase 1 and Phase 3 equivalents | **run** | + +Nothing was dropped. + +## 7. Evidence + +`evidence/`, files `00`–`31`, copied off `demo-hp` at the end of each phase. Nothing was reverted or +torn down mid-run, so no intermediate teardown could have taken any of it. + +## 8. Teardown — three layers + +1. **Apps.** `docmost` and `bookstack` were **deployed by this drill** and are **RETAINED**, with their + planted data. Reason: the planted data *is* the evidence for every claim above, and they are the only + deployed members of the 10-app class on the box, so the next drill in this area starts from a real + subject instead of rebuilding one. Both are healthy and sane at the end (docmost restored via its + undo copy; bookstack's `migrations` table restored 0 → 102 rows the same way). +2. **Storage.** No `pvesm` "before" snapshot was taken — **stated plainly rather than reconstructed.** + Measured directly instead: the two apps' volumes total ~233 MB + (docmost 67 MB + 228 KB + 4 KB, bookstack 166 MB + 224 KB) inside guest 9201, which sits at + 16 % of 69 GB. Host `local-lvm` 34.81 %, `local` 14.27 %. +3. **Hub-side record.** **None was created.** This drill created no customer and no appliance record; + it deployed two apps inside an existing guest of an existing customer. Nothing to dispose of. + +## 9. Store integrity at the end + +`restic check` against the live repository: **`no errors were found`**, 29 snapshots, exclusive lock +taken and released. Evidence `30`. diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/00-baselines.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/00-baselines.txt new file mode 100644 index 00000000..e5954955 --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/00-baselines.txt @@ -0,0 +1,17 @@ +Baselines re-verified at drill start, 2026-08-22. + +repo HEAD clean in-sync-with-origin/main +felhom-controller 0f3cf0dbb2c0 yes yes (v0.219.0) +felhom-agent 40d857b52711 yes yes (v0.130.0) +felhom.eu 8c9f1b798b12 yes yes (ahead of the runbook's f2edf7e54595: + this session's own golden-bake commit) +app-catalog-felhom.eu 459766cb1639 yes yes + +Read from the HUB (operator UI, Basic auth, ClusterIP 10.43.52.34:8080/configuration): + golden_version 0.219.0 <- vouched + agent_version 0.130.0 + min_agent 0.129.0 + global controller floor 0.219.0 <- raised + +Register ceiling before this drill: R-378 (highest id across OPEN-ITEMS.md + CLOSED-ITEMS.md). +Next free id: R-379. diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/01-driveless-with-db-measurement.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/01-driveless-with-db-measurement.txt new file mode 100644 index 00000000..e5fd26bb --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/01-driveless-with-db-measurement.txt @@ -0,0 +1,16 @@ +bookstack | image: mariadb:12.3 +calcom | image: postgres:16-alpine +claper | image: postgres:16-alpine +docmost | image: postgres:16-alpine +kimai | image: mariadb:11.6 +outline | image: postgres:16-alpine +rallly | image: postgres:16-alpine +sparkyfitness | image: postgres:15-alpine +tandoor | image: postgres:16-alpine +zipline | image: postgres:16-alpine + +Method: for each templates//, an app qualifies when BOTH + (a) .felhom.yml declares needs_hdd: false + (b) docker-compose.yml has an image: line matching postgres|mariadb|mysql +Catalog @ 459766cb1639. +COUNT MEASURED THIS SESSION: 10 -- agrees with the runbook's table. diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/02-findability-control.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/02-findability-control.txt new file mode 100644 index 00000000..5e105a66 --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/02-findability-control.txt @@ -0,0 +1,15 @@ +== 1. plant the control row +INSERT 0 1 +== 2. FIND it +found: R356B-FINDABILITY-CONTROL +== 3. remove it +DELETE 1 +== 4. fail to find it +absent, as required +== the planted state that must survive: +1|ORIGINAL-VALUE-A +== pages visible through the app's own API: + page: "\\u00c1rv\\u00edzt\\u0171r\\u0151 t\\u00fck\\u00f6rf\\u00far\\u00f3g\\u00e9p \\u2014 R356b drill" + page: "R356B-PAGE-2-sentinel" + page: "R356B-PAGE-1-sentinel" + count: 3 diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/03-planted-manifest.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/03-planted-manifest.txt new file mode 100644 index 00000000..2825992e --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/03-planted-manifest.txt @@ -0,0 +1,22 @@ +PLANTED STATE — docmost on demo-hp, 2026-08-22, before the off-site run. + +Through the app's OWN API (POST /api/pages/create, format=markdown), 3 pages: + id=01a029ad-81fe-792b-87c2-b940ffc80c0c slugId=F4G3YyHtFj title="R356B-PAGE-1-sentinel" + body: FELHOM-R356B-SENTINEL-ALPHA-2026-08-22 plain ascii body + id=01a029ad-8298-712b-bf80-e41d1186a746 slugId=87aKZR4rlm title="R356B-PAGE-2-sentinel" + body: FELHOM-R356B-SENTINEL-BETA-2026-08-22 second body + id=01a029ad-8314-76e9-bef0-2108aed8b7a6 slugId=fJOFUooa0N title= + body: FELHOM-R356B-SENTINEL-GAMMA-2026-08-22 akcentusos oldal torzse + +ACCENTED TITLE, as explicit UTF-8 hex bytes: + c3817276c3ad7a74c5b172c5912074c3bc6bc3b67266c3ba72c3b367c3a97020e28094205233353662206472696c6c + = "Árvíztűrő tükörfúrógép — R356b drill" +Non-ASCII never crossed a shell chain: the planting script was base64-encoded on DooPlex and +decoded inside the guest. + +DIRECTLY IN THE DATABASE — the Q2 discriminator: + table felhom_r356b_discriminator + row id=1, marker='ORIGINAL-VALUE-A' + +FINDABILITY CONTROL (evidence 02): a control row was planted, FOUND, removed, and then NOT found. +Every absence claim later in this drill rests on that control. diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/04-offsite-run.json b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/04-offsite-run.json new file mode 100644 index 00000000..611a9b6d --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/04-offsite-run.json @@ -0,0 +1 @@ +{"last_duration":"2m38s","last_error":"","last_run":"2026-08-22T13:37:41Z","orphaned":false,"progress":{"active":false,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":0,"current_file":"","elapsed_sec":0,"phase":""},"repo_size_human":"54.1 MB","snapshots":28,"status":"ok"} diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/05-scratch-shape.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/05-scratch-shape.txt new file mode 100644 index 00000000..05216bc9 --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/05-scratch-shape.txt @@ -0,0 +1,8 @@ +2332 /mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost/mnt/sys_drive/felhom-data/backups/primary/docmost/compose/.felhom.yml +445 /mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost/mnt/sys_drive/felhom-data/backups/primary/docmost/compose/app.yaml +3105 /mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost/mnt/sys_drive/felhom-data/backups/primary/docmost/compose/docker-compose.yml +139770 /mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/docmost-postgres.sql +1376 /mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost/mnt/sys_drive/felhom-data/backups/primary/docmost/manifest.json +52344320 /mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost/mnt/sys_drive/felhom-data/backups/primary/docmost/volume-dumps/docmost_docmost_postgres_data.tar +56320 /mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost/mnt/sys_drive/felhom-data/backups/primary/docmost/volume-dumps/docmost_docmost_redis_data.tar +1536 /mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost/mnt/sys_drive/felhom-data/backups/primary/docmost/volume-dumps/docmost_docmost_storage.tar diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/06-destruction.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/06-destruction.txt new file mode 100644 index 00000000..626853f1 --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/06-destruction.txt @@ -0,0 +1,16 @@ +== BEFORE — pages in the database: +01a029ad-81fe-792b-87c2-b940ffc80c0c | R356B-PAGE-1-sentinel +01a029ad-8298-712b-bf80-e41d1186a746 | R356B-PAGE-2-sentinel +01a029ad-8314-76e9-bef0-2108aed8b7a6 | \u00c1rv\u00edzt\u0171r\u0151 t\u00fck\u00f6rf\u00far\u00f3g\u00e9p \u2014 R356b drill +== deleting all three pages through the app's own API (force delete, not trash): + 01a029ad-81fe-792b-87c2-b940ffc80c0c -> {"success":true,"status":200} + 01a029ad-8298-712b-bf80-e41d1186a746 -> {"success":true,"status":200} + 01a029ad-8314-76e9-bef0-2108aed8b7a6 -> {"success":true,"status":200} +== move the discriminator to a POST-SNAPSHOT value, so a revert is visible: +UPDATE 1 +== AFTER — pages in the database (must be empty): + (none — genuinely gone) +== AFTER — any row at all in pages, incl. soft-deleted (a trash that still holds them is not a deletion): +0 +== AFTER — discriminator: +1|DESTROYED-VALUE-C diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/07-reconstitute-post.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/07-reconstitute-post.txt new file mode 100644 index 00000000..d24ebca4 --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/07-reconstitute-post.txt @@ -0,0 +1,2 @@ +HTTP/2 302 +location: /backups/restore/app?name=docmost&flash=A+teljes+visszaall%C3%ADtas+elindult+%E2%80%94+az+allapot+itt+friss%C3%BCl. diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/08-reconstitute-log.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/08-reconstitute-log.txt new file mode 100644 index 00000000..1627f51c --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/08-reconstitute-log.txt @@ -0,0 +1,40 @@ +2026/08/22 13:58:30 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container paperless-webserver (image=ghcr.io/paperless-ngx/paperless-ngx:2.20.15, not a database) +2026/08/22 13:58:30 dbdump.go:140: [DEBUG] DiscoverDatabases: found postgres container: paperless-postgres (id=5be5e049dab2) +2026/08/22 13:58:30 dbdump.go:805: [DEBUG] DiscoverDatabases: paperless-postgres → stack "paperless-ngx" from the compose project label (the container name would have given "paperless") +2026/08/22 13:58:30 dbdump.go:160: [DEBUG] DiscoverDatabases: paperless-postgres → stack=paperless-ngx, dbUser=paperless, dbName=paperless +2026/08/22 13:58:30 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container paperless-redis (image=redis:7-alpine, not a database) +2026/08/22 13:58:30 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container opengist (image=ghcr.io/thomiceli/opengist:1.13, not a database) +2026/08/22 13:58:30 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container kimai (image=kimai/kimai2:apache-2.57.0, not a database) +2026/08/22 13:58:30 dbdump.go:140: [DEBUG] DiscoverDatabases: found mariadb container: kimai-db (id=53e87f2a117d) +2026/08/22 13:58:30 dbdump.go:160: [DEBUG] DiscoverDatabases: kimai-db → stack=kimai, dbUser=root, dbName=kimai +2026/08/22 13:58:30 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container calibre-web (image=crocodilestick/calibre-web-automated:v4.0.6, not a database) +2026/08/22 13:58:30 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container felhom-controller (image=gitea.dooplex.hu/admin/felhom-controller:0.219.0, not a database) +2026/08/22 13:58:30 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container filebrowser (image=gtstef/filebrowser:1.3.3-stable, not a database) +2026/08/22 13:58:30 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container cloudflared (image=cloudflare/cloudflared:2026.6.0, not a database) +2026/08/22 13:58:30 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container traefik (image=traefik:v3.6.7, not a database) +2026/08/22 13:58:30 dbdump.go:167: [DEBUG] DiscoverDatabases: found 4 database(s), skipped 12 non-DB container(s) +2026/08/22 13:58:30 dbdump.go:170: [INFO] [backup] Discovered 4 databases +2026/08/22 13:58:30 restore_db.go:77: [INFO] [backup] Restore docmost: replaying DB dump into docmost-postgres (postgres) +2026/08/22 13:58:30 dbdump.go:700: [DEBUG] [backup] ImportDump: importing /mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/docmost-postgres.sql into docmost-postgres (postgres) +2026/08/22 13:58:32 dbdump.go:710: [INFO] [backup] Imported DB dump docmost-postgres.sql into docmost-postgres (postgres) +2026/08/22 13:58:32 restore_db.go:87: [INFO] [backup] Restore docmost: replayed 1 DB dump(s) +2026/08/22 13:58:32 manager.go:986: [DEBUG] [stacks] StartStack docmost: current state=starting deployed=true +2026/08/22 13:58:32 manager.go:989: [INFO] [stacks] Starting stack: docmost +2026/08/22 13:58:32 manager.go:996: [DEBUG] [stacks] StartStack docmost: prepared 10 env vars for compose +2026/08/22 13:58:32 manager.go:1261: [DEBUG] Env vars for compose: [PATH, HOSTNAME, FELHOM_BOOTSTRAP_PATH, HOME, DOMAIN, APP_SECRET, DB_PASSWORD, DOMAIN, SUBDOMAIN, IMPORT_PATH] (10 app + 0 system) +2026/08/22 13:58:32 manager.go:1270: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/docmost) +2026/08/22 13:58:43 manager.go:1004: [INFO] [stacks] Stack docmost started successfully (took 11.1s) +2026/08/22 13:58:46 manager.go:1261: [DEBUG] Env vars for compose: [PATH, HOSTNAME, FELHOM_BOOTSTRAP_PATH, HOME, DOMAIN, APP_SECRET, DB_PASSWORD, DOMAIN, SUBDOMAIN, IMPORT_PATH] (10 app + 0 system) +2026/08/22 13:58:46 manager.go:1270: [DEBUG] Running: docker compose ps -a --format table {{.Name}} {{.Image}} {{.State}} {{.Status}} (in /opt/docker/stacks/docmost) +2026/08/22 13:58:46 restore.go:204: [DEBUG] [backup] Post-restore health check: docmost not yet running, waiting... +2026/08/22 13:58:46 manager.go:1350: [INFO] [stacks] Stack docmost post-start status: +2026/08/22 13:58:46 manager.go:1353: [INFO] [stacks] docmost docmost/docmost:0.95.0 running Up 3 seconds (health: starting) +2026/08/22 13:58:46 manager.go:1353: [INFO] [stacks] docmost-postgres postgres:16-alpine running Up 16 seconds (healthy) +2026/08/22 13:58:46 manager.go:1353: [INFO] [stacks] docmost-redis redis:7-alpine running Up 14 seconds (healthy) +2026/08/22 13:58:46 auth.go:134: [DEBUG] [web] auth: valid session for GET /backup/offbox/status +2026/08/22 13:58:46 server.go:393: [DEBUG] [web] ServeHTTP: GET /backup/offbox/status from 172.18.0.4:52436 +2026/08/22 13:58:51 restore.go:204: [DEBUG] [backup] Post-restore health check: docmost not yet running, waiting... +2026/08/22 13:58:56 restore.go:199: [DEBUG] [backup] Post-restore health check: docmost is running +2026/08/22 13:58:56 offbox_reconstitute.go:452: [INFO] [offbox] reconstituted docmost from snapshot 239dd86c: 0 file(s) placed, 1 DB dump(s) replayed, safety dump=pre-restore-20260822T135827Z-docmost-postgres.sql, skewed=false +2026/08/22 13:58:56 offbox_handlers.go:455: [INFO] [web] off-box reconstitute docmost completed (async): files=0 dbs=1 snapshot=239dd86c +2026/08/22 13:59:02 healthprobe.go:153: [DEBUG] Health probe docmost: HTTP GET :3000/ → 200 (6ms) diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/09-verification.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/09-verification.txt new file mode 100644 index 00000000..bd942e4a --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/09-verification.txt @@ -0,0 +1,24 @@ +== pages in the database after the restore: +01a029ad-81fe-792b-87c2-b940ffc80c0c | R356B-PAGE-1-sentinel +01a029ad-8298-712b-bf80-e41d1186a746 | R356B-PAGE-2-sentinel +01a029ad-8314-76e9-bef0-2108aed8b7a6 | \u00c1rv\u00edzt\u0171r\u0151 t\u00fck\u00f6rf\u00far\u00f3g\u00e9p \u2014 R356b drill +== page count: +3 +== THE DISCRIMINATOR (Q2): +1|ORIGINAL-VALUE-A +== through the app's OWN interface — log in fresh and list: +HTTP/2 400 + count: 0 +== body of the accented page, read back through the app: + title hex: + content : null +login: HTTP/2 200 +== pages listed through the app's OWN interface: + title(hex)=5c753030633172765c75303065647a745c7530313731725c753031353120745c75303066636b5c753030663672665c7530306661725c7530306633675c753030653970205c7532303134205233353662206472696c6c id=01a029ad-8314-76e9-bef0-2108aed8b7a6 + title(hex)=52333536422d504147452d322d73656e74696e656c id=01a029ad-8298-712b-bf80-e41d1186a746 + title(hex)=52333536422d504147452d312d73656e74696e656c id=01a029ad-81fe-792b-87c2-b940ffc80c0c + count: 3 +== the accented page, body read back through the app: + title hex : 5c753030633172765c75303065647a745c7530313731725c753031353120745c75303066636b5c753030663672665c7530306661725c7530306633675c753030653970205c7532303134205233353662206472696c6c + title : "\\u00c1rv\\u00edzt\\u0171r\\u0151 t\\u00fck\\u00f6rf\\u00far\\u00f3g\\u00e9p \\u2014 R356b drill" + body has GAMMA sentinel: True diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/10-phase1-result.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/10-phase1-result.txt new file mode 100644 index 00000000..7e1b3964 --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/10-phase1-result.txt @@ -0,0 +1,63 @@ +PHASE 1 RESULT — docmost (driveless, Postgres 16), demo-hp, controller v0.219.0 + +THE FIVE LEGS, in order, from the controller log (evidence 08). Every one ran; every one succeeded. + + 13:58:24 POST /backup/offbox/reconstitute (the endpoint the UI's button posts to) + 13:58:27 leg 1 safety dump written -> pre-restore-20260822T135827Z-docmost-postgres.sql (135 624 B) + 13:58:2x leg 2 DB service identified -> [docmost-postgres] + 13:58:28 leg 3a stack stopped + 13:58:28 leg 3b volumes replayed -> docmost_docmost_postgres_data + 13:58:29 docmost_docmost_redis_data + 13:58:30 docmost_docmost_storage = "Restored 3 Docker volume(s)" + 13:58:30 leg 4 DB service ONLY started -> docker compose up -d docmost-postgres (0.4s) + 13:58:30 leg 5 ImportDump docmost-postgres.sql -> docmost-postgres (postgres) + 13:58:32 "Imported DB dump ... " / "replayed 1 DB dump(s)" (2 s) + 13:58:32 full StartStack + 13:58:56 reconstituted docmost from snapshot 239dd86c: + 0 file(s) placed, 1 DB dump(s) replayed, + safety dump=pre-restore-20260822T135827Z-docmost-postgres.sql, skewed=false + + Total 32 s. All three containers healthy afterwards. + +WHAT CAME BACK + pages before destruction : 3 + pages after destruction : 0 (count(*) FROM pages = 0 — not even soft-deleted rows) + pages after restore : 3 same ids, same titles + discriminator before : ORIGINAL-VALUE-A + discriminator moved to : DESTROYED-VALUE-C (deliberately, so a revert would be visible) + discriminator after : ORIGINAL-VALUE-A -> the snapshot's value was restored + + Read back through the APP'S OWN interface after the restore: login HTTP 200, /api/pages/recent + returned all 3 pages, and the body of the third still contained its GAMMA sentinel. + +HYPOTHESES — what fired and what did not + H1 (schema collision, dump replays over a restored data directory) — DID NOT FIRE. + The import completed in 2 s with no error under ON_ERROR_STOP=1. This is the notable result: + the two legs do not collide in practice for this app. + H2 (replayed data directory will not start) — DID NOT FIRE. docmost-postgres came up healthy + on a data directory that had just been replaced wholesale from a tar. + H4 (credentials) — DID NOT FIRE. ImportDump discovered dbUser=docmost/dbName=docmost from the + LIVE container and they matched the restored directory, as predicted: secrets are not + regenerated on reconstitution. + H3 (MariaDB vs Postgres error handling) — not applicable to this phase; Phase 4. + +A FLAW IN THIS DRILL'S OWN PLANTING, recorded rather than hidden + The first accented title was planted through a shell argument chain that double-escaped it, so + it was stored as the LITERAL ASCII text Árvízt... rather than as accented bytes. + Caught by reading the stored value back as hex. The pages round-tripped exactly, so Phase 1's + Q1 answer stands — but the ACCENTED-BYTES coverage did not, and was re-measured in Phase 1b + with the title embedded in a python source file inside the guest, never passed as a shell + argument. Wire bytes and stored bytes were compared and are identical. + +PHASE 1b — the accented round-trip, re-measured properly + planted (wire bytes) : c3817276c3ad7a74c5b172c5912074c3bc6bc3b67266c3ba72c3b367c3a970203220e28094205233353662 + stored at plant time : identical + destroyed : all 4 pages permanently; count(*) FROM pages = 0 + restored (snapshot a49de51d, 14:05:33): 4 pages back + accented title after : c3817276c3ad7a74c5b172c5912074c3bc6bc3b67266c3ba72c3b367c3a970203220e28094205233353662 + MATCH : TRUE ("Árvíztűrő tükörfúrógép 2 — R356b") + discriminator : DESTROYED-VALUE-C2 -> ORIGINAL-VALUE-A + + The first three pages also came back; page 01a029ad-8314-... still carries the double-escaped + ASCII title, which is what this drill's own faulty plant stored. It is left in place as the + record of that flaw rather than quietly re-planted. diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/11-phase1b-destruction.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/11-phase1b-destruction.txt new file mode 100644 index 00000000..a3e7f04f --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/11-phase1b-destruction.txt @@ -0,0 +1,12 @@ +== BEFORE: page count and the accented title as hex +4 + accented title hex: c3817276c3ad7a74c5b172c5912074c3bc6bc3b67266c3ba72c3b367c3a970203220e28094205233353662 +== destroy ALL pages through the app (permanent, not trash): + 01a029ad-81fe-792b-87c2-b940ffc80c0c -> {"success":true,"status":200} + 01a029ad-8298-712b-bf80-e41d1186a746 -> {"success":true,"status":200} + 01a029ad-8314-76e9-bef0-2108aed8b7a6 -> {"success":true,"status":200} + 01a029c6-47b8-7252-8ea2-4c2036f9e1d0 -> {"success":true,"status":200} +UPDATE 1 +== AFTER destruction — page count (must be 0) and discriminator: +0 +1|DESTROYED-VALUE-C2 diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/12-phase1b-verification.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/12-phase1b-verification.txt new file mode 100644 index 00000000..b269a78d --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/12-phase1b-verification.txt @@ -0,0 +1,13 @@ +== page count after restore (was 0 before): +4 +== all titles, as UTF-8 hex: + 01a029ad-81fe-792b-87c2-b940ffc80c0c hex=52333536422d504147452d312d73656e74696e656c + 01a029ad-8298-712b-bf80-e41d1186a746 hex=52333536422d504147452d322d73656e74696e656c + 01a029ad-8314-76e9-bef0-2108aed8b7a6 hex=5c753030633172765c75303065647a745c7530313731725c753031353120745c75303066636b5c753030663672665c7530306661725c7530306633675c753030653970205c7532303134205233353662206472696c6c + 01a029c6-47b8-7252-8ea2-4c2036f9e1d0 hex=c3817276c3ad7a74c5b172c5912074c3bc6bc3b67266c3ba72c3b367c3a970203220e28094205233353662 +== THE ACCENTED PAGE — expected hex: + c3817276c3ad7a74c5b172c5912074c3bc6bc3b67266c3ba72c3b367c3a970203220e28094205233353662 + actual hex: c3817276c3ad7a74c5b172c5912074c3bc6bc3b67266c3ba72c3b367c3a970203220e28094205233353662 + MATCH: True +== discriminator (was DESTROYED-VALUE-C2 before the restore): +1|ORIGINAL-VALUE-A diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/13-phase2-mutation.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/13-phase2-mutation.txt new file mode 100644 index 00000000..f693af4f --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/13-phase2-mutation.txt @@ -0,0 +1,17 @@ +BEFORE sha256: c5414f24426a4a0050becadeb564990465bf48772a7cb029daaebd22ef788441 +BEFORE line 1536: 1 ORIGINAL-VALUE-A 2026-08-22 15:34:22.455005+02 +AFTER line 1536: 1 ALTERED-VALUE-B 2026-08-22 15:34:22.455005+02 +AFTER sha256: 834c32b19b9ec053b2909f031751cbbcf8d14941b5af2fdf1f3779b113b82d36 +occurrences of ORIGINAL-VALUE-A left in the scratch dump: 0 +occurrences of ALTERED-VALUE-B: 1 +--- the volume tar still holds the ORIGINAL (it is the other candidate): +-rw-r--r-- 1 root root 53301760 Aug 22 14:01 /mnt/felhom-drives/hdd_1/backups/offsite-restore/docmost/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/../volume-dumps/docmost_docmost_postgres_data.tar + +THE THREE-WAY DISCRIMINATOR, set up before the reconstitution: + live database now holds : LIVE-VALUE-C3 <- if this survives, NEITHER leg wrote + volume tar holds : ORIGINAL-VALUE-A <- if this wins, the volume tar supplied the state + altered SQL dump holds : ALTERED-VALUE-B <- if this wins, the dump supplied it (as the comment intends) + +FENCE: the mutation was applied to the PREPARED SCRATCH ONLY. The off-site store and the live +recovery unit were not touched. Proved after the run by re-preparing a fresh scratch from the same +snapshot and confirming it still reads ORIGINAL-VALUE-A. diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/14-phase2-store-untouched.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/14-phase2-store-untouched.txt new file mode 100644 index 00000000..8e89c155 --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/14-phase2-store-untouched.txt @@ -0,0 +1,12 @@ +fresh scratch sha256: c5414f24426a4a0050becadeb564990465bf48772a7cb029daaebd22ef788441 +line 1536 : 1 ORIGINAL-VALUE-A 2026-08-22 15:34:22.455005+02 +ORIGINAL-VALUE-A present: 1 +ALTERED-VALUE-B present: 0 +0 +--- the LIVE recovery unit db-dumps dir (never touched by a restore): +total 420 +drwxr-xr-x 2 root root 4096 Aug 22 14:07 . +drwxr-xr-x 5 root root 4096 Aug 22 14:05 .. +-rw-r--r-- 1 root root 135624 Aug 22 13:58 pre-restore-20260822T135827Z-docmost-postgres.sql +-rw-r--r-- 1 root root 135847 Aug 22 14:05 pre-restore-20260822T140501Z-docmost-postgres.sql +-rw-r--r-- 1 root root 141361 Aug 22 14:07 pre-restore-20260822T140713Z-docmost-postgres.sql diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/15-phase3-corruption.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/15-phase3-corruption.txt new file mode 100644 index 00000000..895a4cc0 --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/15-phase3-corruption.txt @@ -0,0 +1,7 @@ +live discriminator before Phase 3: ALTERED-VALUE-B +live page count before Phase 3 : 4 +--- CORRUPTION METHOD: truncate the dump mid-COPY-block (a realistic partial/corrupt dump) +size before: 141364 +size after : 62000 +last line now: COPY public.felhom_r356b_discriminator +sha256 after: 30e84d2c0d0bcf3e7043459a0d135e96cdb65947f4138f181837b1ee35d2e087 diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/16-phase3-error-log.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/16-phase3-error-log.txt new file mode 100644 index 00000000..bc1eb5ba --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/16-phase3-error-log.txt @@ -0,0 +1 @@ +2026/08/22 14:09:40 offbox_handlers.go:451: [ERROR] [web] off-box reconstitute docmost (async): az adatbázis visszaállítása sikertelen: importing postgres dump for docmost: postgres import into docmost-postgres failed: ERROR: syntax error at end of input diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/17-phase3-customer-message.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/17-phase3-customer-message.txt new file mode 100644 index 00000000..2f1a8a04 --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/17-phase3-customer-message.txt @@ -0,0 +1,9 @@ +=== CUSTOMER-FACING MESSAGE, VERBATIM === +A teljes visszaállítás sikertelen: az adatbázis visszaállítása sikertelen: importing postgres dump for docmost: postgres import into docmost-postgres failed: ERROR: syntax error at end of input +LINE 1: COPY public.felhom_r356b_discriminator + ^ — exit status 3 — a korábbi állapot mentése megvan: pre-restore-20260822T140924Z-docmost-postgres.sql + +=== UTF-8 hex === +412074656c6a657320766973737a61c3a16c6cc3ad74c3a1732073696b657274656c656e3a20617a206164617462c3a17a697320766973737a61c3a16c6cc3ad74c3a173612073696b657274656c656e3a20696d706f7274696e6720706f7374677265732064756d7020666f7220646f636d6f73743a20706f73746772657320696d706f727420696e746f20646f636d6f73742d706f737467726573206661696c65643a204552524f523a202073796e746178206572726f7220617420656e64206f6620696e7075740a4c494e4520313a20434f5059207075626c69632e66656c686f6d5f72333536625f6469736372696d696e61746f72200a20202020202020202020202020202020202020202020202020202020202020202020202020202020202020202020205e20e28094206578697420737461747573203320e280942061206b6f72c3a162626920c3a16c6c61706f74206d656e74c3a97365206d656776616e3a207072652d726573746f72652d3230323630383232543134303932345a2d646f636d6f73742d706f7374677265732e73716c + +=== byte length: 407 diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/18-phase3-data-state.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/18-phase3-data-state.txt new file mode 100644 index 00000000..0ffbcc40 --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/18-phase3-data-state.txt @@ -0,0 +1,19 @@ +== tables left in the public schema (a healthy docmost has ~30): +43 +== does pages still exist? +pages +== rows in pages (was 4 before the failed restore): +0 +== rows in users (the workspace owner): +0 +== rows in spaces : +0 +== discriminator rows (was 1): +0 +== the undo copy that the message named: +-rw-r--r-- 1 root root 141363 Aug 22 14:09 /mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260822T140924Z-docmost-postgres.sql +== is that undo copy a VALID dump? (header + row count for pages) +43 +-- +-- PostgreSQL database dump +-- diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/19-phase3-undo-copy.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/19-phase3-undo-copy.txt new file mode 100644 index 00000000..f5e69306 --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/19-phase3-undo-copy.txt @@ -0,0 +1,12 @@ +== does the undo copy CONTAIN the lost data? + rows in its pages COPY block: 4 + rows in its users COPY block: 1 + contains the accented title : 1 +== APPLY it by hand — proving the undo copy WORKS (nothing in the product will do this): + +(1 row) + +== state after applying the undo by hand: + pages: 4 + users: 1 + discriminator: ALTERED-VALUE-B diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/20-phase3-result.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/20-phase3-result.txt new file mode 100644 index 00000000..64d7787f --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/20-phase3-result.txt @@ -0,0 +1,53 @@ +PHASE 3 RESULT — the failure path (docmost, Postgres) + +CORRUPTION METHOD: the prepared scratch's db-dumps/docmost-postgres.sql was TRUNCATED from +141 364 B to 62 000 B, landing mid-statement at "COPY public.felhom_r356b_discriminator ". +sha256 30e84d2c0d0bcf3e7043459a0d135e96cdb65947f4138f181837b1ee35d2e087. +Scratch only. Store and live unit untouched. + +DOES THE CUSTOMER SEE A FAILURE? YES. + Rendered in an `alert alert-error` block under the heading "Eredmény", beginning + "A teljes visszaállítás sikertelen:". Not a warning beside a success. + +DOES THE MESSAGE NAME THE UNDO COPY? YES. + "... — a korábbi állapot mentése megvan: pre-restore-20260822T140924Z-docmost-postgres.sql" + +IS THE APP RUNNING AFTERWARDS? NO — it crash-loops. + `docmost | Restarting (1)`; docmost-postgres and docmost-redis stay healthy. Login through + traefik returns 404 because there is no healthy backend. + This is recorded as HONEST rather than as a defect: the app is visibly broken, not falsely + green. An earlier reading in this session said "reports healthy" — that was wrong, taken from + a "health: starting" line during the restart, and is corrected here. + +WHAT STATE IS THE DATA IN? EMPTY, and that is inherent to the mechanism. + The truncated dump ran its DROP/CREATE sequence and died partway through the data: + tables in public schema : 43 (schema intact) + pages : 0 (was 4) + users : 0 (was 1 — the workspace owner) + spaces : 0 + felhom_r356b_discriminator: 0 (was 1) + +IS THE UNDO COPY VALID AND SUFFICIENT? YES — proven by using it. + 141 363 B, proper pg_dump header, 43 COPY blocks. + Its pages block holds 4 rows, its users block 1 row, and it contains the accented title. + Applied BY HAND (docker cp + psql -v ON_ERROR_STOP=1): pages 4, users 1, + discriminator ALTERED-VALUE-B — exactly the pre-restore state. + +THE GAP — and it is the finding of this phase: + NOTHING IN THE PRODUCT CAN APPLY THAT UNDO COPY. + - The filename appears ONLY inside the error text. There is no button, no list entry, no route. + - `preRestoreDumpPrefix` ("pre-restore-") is explicitly SKIPPED at three places so these files + are never offered as a restore source: + internal/backup/restore_unit.go:125 + internal/backup/offbox_reconstitute.go:489 + internal/backup/offbox_reconstitute.go:539 + - So a customer whose off-site dump is corrupt is left with a crash-looping app, an emptied + database, and a filename they cannot act on. The data is recoverable — by us, by hand. + +SECOND FINDING — the message leaks raw engine internals at a Hungarian customer. + 407 bytes, of which the middle ~250 are untranslated English psql output including a caret + diagram and "exit status 3": + "importing postgres dump for docmost: postgres import into docmost-postgres failed: + ERROR: syntax error at end of input + LINE 1: COPY public.felhom_r356b_discriminator + ^ — exit status 3" diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/21-phase4-planted.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/21-phase4-planted.txt new file mode 100644 index 00000000..bff7c82f --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/21-phase4-planted.txt @@ -0,0 +1,19 @@ +PHASE 4 PLANTED STATE — bookstack (driveless, MariaDB 12.3, container bookstack-db) + +Deployed by this drill (it was not on the box). Off-site switch turned ON and verified. + +Through the app's OWN web interface (login form -> POST /books, session cookie + Laravel _token): + entities: id=1 type=book + name (UTF-8 hex): C3817276C3AD7A74C5B172C591206BC3B66E79766573706F6C6320E28094205233353662 + = "Árvíztűrő könyvespolc — R356b" + description : FELHOM-R356B-BOOKSTACK-SENTINEL-2026-08-22 + The name was written from a python file inside the guest and passed to curl with + --data-urlencode "name@/tmp/bs.name", so it never crossed a shell argument. + + NOTE: this BookStack version has no `books` table — the 2025_09_15 migration + `drop_old_entity_tables` unified books/chapters/pages into `entities`. 42 tables total. + +Directly in the database — the discriminator: + felhom_r356b_disc: id=1, marker='ORIGINAL-VALUE-A' + +FINDABILITY CONTROL: a control row was planted (id=99), FOUND, removed, and then NOT found. diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/22-phase4-destruction.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/22-phase4-destruction.txt new file mode 100644 index 00000000..b6d7da3a --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/22-phase4-destruction.txt @@ -0,0 +1,9 @@ +== BEFORE: +1 book C3817276C3AD7A74C5B172C591206BC3B66E79766573706F6C6320E28094205233353662 +1 ORIGINAL-VALUE-A +== DESTROY (delete the book through the app is a soft-delete; this removes it outright): +== AFTER destruction — entities count (must be 0) and discriminator: +0 +1 DESTROYED-VALUE-C +== also check the deletions/recycle-bin table is not holding it: +0 diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/23-phase4-scratch-shape.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/23-phase4-scratch-shape.txt new file mode 100644 index 00000000..59ca372d --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/23-phase4-scratch-shape.txt @@ -0,0 +1,7 @@ + 2075 ./mnt/sys_drive/felhom-data/backups/primary/bookstack/compose/.felhom.yml + 426 ./mnt/sys_drive/felhom-data/backups/primary/bookstack/compose/app.yaml + 2330 ./mnt/sys_drive/felhom-data/backups/primary/bookstack/compose/docker-compose.yml + 58775 ./mnt/sys_drive/felhom-data/backups/primary/bookstack/db-dumps/bookstack-mariadb.sql + 1325 ./mnt/sys_drive/felhom-data/backups/primary/bookstack/manifest.json + 136704 ./mnt/sys_drive/felhom-data/backups/primary/bookstack/volume-dumps/bookstack_bookstack_config.tar + 160945152 ./mnt/sys_drive/felhom-data/backups/primary/bookstack/volume-dumps/bookstack_bookstack_db_data.tar diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/24-phase4-verification.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/24-phase4-verification.txt new file mode 100644 index 00000000..bd61b184 --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/24-phase4-verification.txt @@ -0,0 +1,9 @@ +== entities after restore (was 0): +1 book C3817276C3AD7A74C5B172C591206BC3B66E79766573706F6C6320E28094205233353662 +== expected accented hex: + C3817276C3AD7A74C5B172C591206BC3B66E79766573706F6C6320E28094205233353662 +== discriminator (was DESTROYED-VALUE-C): +1 ORIGINAL-VALUE-A +== containers: +bookstack | Up 41 seconds (healthy) +bookstack-db | Up 47 seconds (healthy) diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/25-phase4-corruption.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/25-phase4-corruption.txt new file mode 100644 index 00000000..6a42b8a3 --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/25-phase4-corruption.txt @@ -0,0 +1,8 @@ +size before : 30000 +sha before : e4575b65d84ef24a605c56883f54384c3cd0d91e75862bad9ae69ec7bff52da4 +size after : 30000 +sha after : e4575b65d84ef24a605c56883f54384c3cd0d91e75862bad9ae69ec7bff52da4 +tail (truncation point): +51_update_entity_relation_columns',1), +(96,'2025_09_15_134813_drop_old_entity_tables',1), +(97,'2025_ diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/26-phase4-error-log.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/26-phase4-error-log.txt new file mode 100644 index 00000000..6bf16a73 --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/26-phase4-error-log.txt @@ -0,0 +1 @@ +2026/08/22 14:24:30 offbox_handlers.go:451: [ERROR] [web] off-box reconstitute bookstack (async): az adatbázis visszaállítása sikertelen: importing mariadb dump for bookstack: mariadb import into bookstack-db failed: -------------- diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/27-phase4-customer-message.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/27-phase4-customer-message.txt new file mode 100644 index 00000000..8f96638d --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/27-phase4-customer-message.txt @@ -0,0 +1,14 @@ +=== CUSTOMER-FACING MESSAGE, VERBATIM === +A teljes visszaállítás sikertelen: az adatbázis visszaállítása sikertelen: importing mariadb dump for bookstack: mariadb import into bookstack-db failed: -------------- +INSERT INTO `migrations` VALUES +(1,'2014_10_12_000000_create_users_table',1), +(2,'2014_10_12_100000_create_password_resets_table',1), +(3,'2015_07_12_114933_create_books_table',1), +(4,'2015_07_12_190027_create_pages_table',1), +(5,'2015_07_13_172121_create_images_table',1), +(6,'2015_07_ — exit status 1 — a korábbi állapot mentése megvan: pre-restore-20260822T142418Z-bookstack-mariadb.sql + +=== UTF-8 hex === +412074656c6a657320766973737a61c3a16c6cc3ad74c3a1732073696b657274656c656e3a20617a206164617462c3a17a697320766973737a61c3a16c6cc3ad74c3a173612073696b657274656c656e3a20696d706f7274696e67206d6172696164622064756d7020666f7220626f6f6b737461636b3a206d61726961646220696d706f727420696e746f20626f6f6b737461636b2d6462206661696c65643a202d2d2d2d2d2d2d2d2d2d2d2d2d2d0a494e5345525420494e544f20606d6967726174696f6e73602056414c5545530a28312c262333393b323031345f31305f31325f3030303030305f6372656174655f75736572735f7461626c65262333393b2c31292c0a28322c262333393b323031345f31305f31325f3130303030305f6372656174655f70617373776f72645f7265736574735f7461626c65262333393b2c31292c0a28332c262333393b323031355f30375f31325f3131343933335f6372656174655f626f6f6b735f7461626c65262333393b2c31292c0a28342c262333393b323031355f30375f31325f3139303032375f6372656174655f70616765735f7461626c65262333393b2c31292c0a28352c262333393b323031355f30375f31335f3137323132315f6372656174655f696d616765735f7461626c65262333393b2c31292c0a28362c262333393b323031355f30375f20e28094206578697420737461747573203120e280942061206b6f72c3a162626920c3a16c6c61706f74206d656e74c3a97365206d656776616e3a207072652d726573746f72652d3230323630383232543134323431385a2d626f6f6b737461636b2d6d6172696164622e73716c + +=== byte length: 615 diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/28-phase4-data-state.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/28-phase4-data-state.txt new file mode 100644 index 00000000..0d818660 --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/28-phase4-data-state.txt @@ -0,0 +1,19 @@ +== tables left: +42 +== entities (was 1): +1 +== users: +2 +== discriminator table still there? +1 +== migrations rows (the table the dump died inside): +0 +== containers: +bookstack | Up 2 minutes (healthy) +bookstack-db | Up 3 minutes (healthy) +== undo copy on disk: +total 128 +drwxr-xr-x 2 root root 4096 Aug 22 14:24 . +drwxr-xr-x 5 root root 4096 Aug 22 14:25 .. +-rw-r--r-- 1 root root 58595 Aug 22 14:22 pre-restore-20260822T142223Z-bookstack-mariadb.sql +-rw-r--r-- 1 root root 58775 Aug 22 14:24 pre-restore-20260822T142418Z-bookstack-mariadb.sql diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/29-phase4-undo.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/29-phase4-undo.txt new file mode 100644 index 00000000..636889cb --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/29-phase4-undo.txt @@ -0,0 +1,7 @@ +== applying the undo copy by hand (nothing in the product will do this): +== state after: + entities : 1 + users : 2 + migrations: 102 + book name : C3817276C3AD7A74C5B172C591206BC3B66E79766573706F6C6320E28094205233353662 + disc : ORIGINAL-VALUE-A diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/30-restic-check.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/30-restic-check.txt new file mode 100644 index 00000000..c0d6f20a --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/30-restic-check.txt @@ -0,0 +1,9 @@ +repo: sftp:u629488-sub3@u629488-sub3.your-storagebox.de:/home/felhom-repo (port 23) +using temporary cache in /tmp/restic-check-cache-2896564091 +create exclusive lock for repository +load indexes +check all packs +check snapshots, trees and blobs +[0:01] 100.00% 29 / 29 snapshots + +no errors were found diff --git a/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/31-teardown-storage.txt b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/31-teardown-storage.txt new file mode 100644 index 00000000..ec14a1c7 --- /dev/null +++ b/documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/evidence/31-teardown-storage.txt @@ -0,0 +1,8 @@ +Name Type Status Total (KiB) Used (KiB) Available (KiB) % +felhom-pbs pbs active 0 0 0 0.00% +local dir active 40453376 5772488 32593772 14.27% +local-lvm lvmthin active 56487936 19663450 36824485 34.81% +--- guest 9201 disk: +Filesystem Size Used Avail Use% Mounted on +/dev/mapper/pve-vm--9201--disk--1 69G 10G 56G 16% /var/lib/felhom +/dev/mapper/pve-vm--9201--disk--1 69G 10G 56G 16% /mnt/sys_drive diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index e3d7ef6a..720b9b4b 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -136,6 +136,10 @@ the fault was real. Full observables: `tests/campaign11-evidence-2026-08-05/jour | ID | What | State | |---|---|---| +| **R-379** | **The pre-restore undo copy is taken, is valid, is named to the customer — and NOTHING IN THE PRODUCT CAN APPLY IT.** When an off-site DB replay fails, `reimportDBDumpsFrom` returns and the refusal names the safety dump by filename (`offbox_reconstitute.go:436-441`). That filename appears ONLY inside the error string: there is no button, no list entry, no route. `preRestoreDumpPrefix` ("pre-restore-") is deliberately SKIPPED at three sites so these files are never offered as a restore source — `internal/backup/restore_unit.go:125`, `internal/backup/offbox_reconstitute.go:489`, `internal/backup/offbox_reconstitute.go:539`. **PROVEN LIVE 2026-08-22 on `demo-hp`, both engines.** Postgres (`docmost`): after a truncated dump the live database held 43 tables and **0 rows** in `pages`, `users` and `spaces`, and the app crash-looped. The undo copy (141 363 B, 43 COPY blocks, 4 page rows, 1 user row, the accented title present) was applied BY HAND and restored the exact prior state. MariaDB (`bookstack`): same, `migrations` 0 -> 102 rows. **So the data is recoverable — by us, by hand, over a support conversation. The customer has a filename.** | **OPEN — HIGH** | — | Offer the undo copy as a restore source on the app's restore page when one exists, or state in the message that recovery needs support and how to ask. The skip at the three sites is correct for *normal* listing — the gap is that there is no deliberate second surface. **Do NOT widen the three skips**: they exist so a safety dump is never mistaken for the app's own backup (that confusion is R-361's neighbourhood). | CC | +| **R-380** | **A failed MariaDB replay leaves a PARTIALLY APPLIED database behind an app that reports HEALTHY — Postgres fails visibly, MariaDB does not.** `ImportDump` gives the Postgres branch `-v ON_ERROR_STOP=1`; the MariaDB branch is a plain `mariadb -u root -p ` with no equivalent (`internal/appbackup/dbdump.go:670-692`). Both DO surface the failure — H3's predicted 'quiet success' did NOT occur — but the STATE they leave differs, and that is the defect. **Measured 2026-08-22 on `demo-hp` with the same truncation on both engines.** Postgres: everything emptied, app crash-loops, `Restarting (1)` — visibly broken. MariaDB: the dump's DROP/CREATE/INSERT runs table by table, so tables it reached are rebuilt, tables it never reached keep their ORIGINAL data, and the table it died inside is left EMPTY. Result on `bookstack`: `entities` 1 (intact), `users` 2 (intact), **`migrations` 0 rows (wiped)** — the schema-version ledger — while `docker inspect` reported **`health=healthy running=true restarts=0`** and the app served HTTP. An empty `migrations` table means BookStack believes no migration has ever run; the next upgrade would re-run all 102 against an existing schema. **Nothing signals ongoing damage.** | **OPEN — HIGH** | — | Make a failed replay leave a KNOWN state rather than a partial one: wrap the MariaDB import so a failure is atomic, or re-apply the undo copy automatically on import failure (which needs R-379 first), or at minimum mark the app unhealthy so the dashboard stops saying it is fine. **The MariaDB client's default IS to abort on error — that was measured, not assumed — so this is not a missing flag; it is the absence of a transaction boundary.** | CC | +| **R-381** | **The restore-failure message pastes raw database-engine stderr — including the customer's own database rows — into a Hungarian customer-facing surface.** `ImportDump` truncates stderr to 300 chars and wraps it verbatim (`internal/appbackup/dbdump.go:700-706`); `offbox_reconstitute.go:436-441` wraps that again; the flash renders it in an `alert alert-error` block. **Measured verbatim 2026-08-22.** Postgres, **407 bytes**, of which ~250 are untranslated English psql output with a caret diagram and `exit status 3`. MariaDB, **615 bytes**, whose middle is an `INSERT INTO \`migrations\` VALUES (1,'2014_10_12_000000_create_users_table',1),(2,...` listing — i.e. **actual table contents, HTML-escaped, shown to the customer**. On a real app that statement could be any row the dump died inside. | **OPEN — MEDIUM** | — | Keep the engine text in the operator log where it belongs; give the customer the reason, the undo copy and the route. Same class as **R-79** (English on customer surfaces) and **R-257** (internal state names in customer copy), but a distinct producer and with a content-disclosure dimension neither has: this one can print rows. | CC | +| **R-382** | **The reconstitution's summary log line omits the volume count it already computed.** `offbox_reconstitute.go:452` logs `%d file(s) placed, %d DB dump(s) replayed, safety dump=%s, skewed=%v` — `res.VolumesReplayed` is set at line 412 and never printed. Measured 2026-08-22: `docmost` logged `0 file(s) placed, 1 DB dump(s) replayed` on a run that replayed **3** volumes including the entire 52 MB Postgres data directory; `bookstack` logged the same shape on a run that replayed 2 including a 161 MB one. The customer-facing flash DOES name the volumes („0 fájl és 1 adatkötet visszaállítva") — so the operator log is less informative than the customer message. This directly obstructed answering the 2026-08-22 drill's Q2 from the log and forced a planted discriminator instead. | **OPEN — LOW** | — | Add `%d volume(s) replayed` to the line. One format string. | CC | | **R-229** | **The instruction-file rightsizing landed for `felhom-controller` and the workspace root; three pieces were deliberately deferred.** Done 2026-08-06: controller split into a 92-effective-line core plus four `paths:`-scoped `.claude/rules/*.md`; workspace root 208→142 effective lines with its versioned copy kept byte-identical; surgical corrections to `felhom-agent` and `felhom.eu` (expired TEMPORARY block, every version literal, the Legacy-Windows copies, the duplicated health-check rule); five contradictions resolved — including a drill-VM claim **measured live** (`qm list` on demo-hp shows VM 300 `drill-r50`; `felhom-agent` was right, `felhom-controller` was wrong); new shared `felhom.eu/scripts/instructions_gate.py` registered in `controller_gates.py` and `agent_gates.py`, 20 fixture tests + red-proof. **Leg (a) CLOSED 2026-08-06 (part 2):** `felhom.eu/CLAUDE.md` **227 → 115 effective lines**, split into a core plus `.claude/rules/{hub,website,manifests,docs}.md`; `instructions_gate` **registered in `scripts/repo_gates.py`** (six gates, all OK) in the required order — trim first, register second, because a registered-but-failing gate refuses every push. Scoping proven from the `InstructionsLoaded` hook log in two fresh sessions, not from frontmatter. **Still deferred:** (b) **CLOSED 2026-08-06 (close-out)** — `felhom-agent/CLAUDE.md` **175 → 99 effective lines** (measured 175, not 173: the CI correction added two), split into a core plus `.claude/rules/{proxmox,localapi,backup,storage}.md` beside the existing `health-checks.md`. The release section now points at the `felhom-build-deploy` skill instead of restating a table that drifts from the script. **Every `CLAUDE.md` in the workspace is now ≤120 effective lines except the workspace root at 142, which is deliberate — it is the only file re-injected after `/compact`.** (c) **CLOSED 2026-08-06 (part 2)** — all 44 orphans resolved with **zero deletions** (file count 158 before and after): 4 durable `reference`-type files indexed, 40 dated episode records moved to `.claude-memory/archive/`. `MEMORY.md` 145 → **150 lines / 17,977 bytes**, and `instructions_gate` check 6 now watches it (over-limit FAILS, orphan WARNS, absent store PASSES *printing its reason*). (d) **The spec-as-failing-test pilot** — moved to R-230. Full accounting: `audits/LEDGER-instruction-trim-2026-08-06.md` + `audits/LEDGER-instruction-trim-part2-2026-08-06.md` | **READY** — owner Viktor | | **R-230** | **Three instruction/memory follow-ups deliberately left by the part-2 session (2026-08-06), each needing a decision rather than an implementation.** (a) **A ruling is owed on auto-written staleness.** The hand-written `CLAUDE.md` files are now clean of version literals and expired blocks — the gate enforces it — but `MEMORY.md`, which Claude writes and which is the LARGER half of what loads (8.4k tokens vs the root file's 6.6k), carries **21 lines with component version literals**, **5 with bare host addresses**, and an entry still reading *"demo boxes REMOTE till ~08-02"* — the same expired-TEMPORARY class the gate was built to kill, now surviving in the one file the gate's content rules do not cover. **Partly actioned 2026-08-06 (close-out), and the ruling is STILL OWED:** the **three statements that were actively false** were corrected — `R-193 decision open` (closed 2026-08-05), `demo boxes REMOTE till ~08-02` (the box answers on the home LAN), `OPEN R-25b` (shipped 2026-07-21) — and gate check 6 now **WARNs** on version literals, host addresses, expired statements and stale-open citations in the index. WARN, never FAIL: Claude writes that file between sessions, so a hard failure would refuse a human's push over a line no human typed, and the warning is read by the model that will next edit it. **The remaining 32 version literals and 4 host addresses were deliberately left** for that loop. What is still owed is the bulk-correction ruling. **Correcting the premise:** the earlier report's "three expired statements" were all FALSE POSITIVES — each matched an ISO date inside a markdown link target, i.e. a filename — while the one real expired claim carried no ISO date at all. (b) **CLOSED 2026-08-06 (close-out)** — the workspace-root `CLAUDE.md` **is now a relative symlink** to the versioned copy, so the divergence class is gone rather than policed. Check 5 learned two shapes: for a link it asserts the target resolves to a real file (**a dangling link is worse than a diverged copy — the instructions load NOTHING and there is no content left to notice is wrong**), for two files byte-identity as before, so a clone elsewhere is unaffected. **Proven, not assumed:** three fresh sessions logged `session_start` for the link path, and a fourth **with no tools at all** quoted standing rule 1 verbatim — the content reaches the model, not just the path. (c) **The spec-as-failing-test pilot**, approved in principle and not started (was R-229(d)). | **READY** — owner Viktor | | **R-232** | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **READY** — owner Viktor |