A drill, not an implementation. No code, no version bump, no CHANGELOG entry. Ten of the forty driveless apps carry a database; I re-measured that count and got 10. For those ten the restore is a five-leg operation that never ran at all until this week, because R-356 refused before any of it started. Walked end to end on demo-hp for both engines - docmost (Postgres 16) and bookstack (MariaDB 12.3) - each deployed for the drill, planted through the app's own interface, destroyed for real, restored through the endpoint the UI posts to. Q1 does it complete: YES. All five legs ran and succeeded, 32s / 25s. Accented names byte-identical both directions. Q2 which leg won: the SQL DUMP. Three-way discriminator returned the altered dump's value. This confirms R-164's F17 ordering on the OFF-SITE path; R-164 only ever cited the local one. Scratch-only mutation; store proved unmutated. Q3 does a failure tell the truth: partly, and two defects. Filed R-379 (HIGH, the undo copy is valid, named, and unappliable by any product action - proven by applying it by hand on both engines), R-380 (HIGH, a failed MariaDB replay leaves a partial database behind an app reporting healthy, where Postgres crash-loops visibly), R-381 (MEDIUM, the failure message pastes engine stderr including customer table rows into the Hungarian surface), R-382 (LOW, the summary log omits the volume count it already has). H1, H2 and H4 did NOT fire and that is recorded. H3 fired in a shape nobody predicted: not a quiet success, but a loud error over a silent inconsistency. R-361 reproduced independently on a second app. restic check: no errors, 29 snapshots. A flaw in the drill's own planting - a double-escaped accented title - was caught by reading stored bytes as hex, recorded, and re-measured in Phase 1b. Register 325236 -> 330683 bytes. Nothing dropped. Teardown: two apps retained with reason, no pvesm before-snapshot taken (said plainly), no hub-side record created.
4.9 KiB
REPORT — DRILL R-356b: the off-site restore for a driveless app that HAS a database (2026-08-22)
A drill, not an implementation. No production code was written, no version bumped, no CHANGELOG entry made. The deliverables are a findings document, four register rows and a capability-map update.
Full record: documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/
What was measured
Ten of the forty driveless apps carry a database. I re-measured that count myself and got 10 — the same ten the runbook names. For those ten, restoring is a five-leg operation that, until this week, never ran at all: R-356 refused before any of it started.
Both engines were walked end to end on demo-hp: docmost (Postgres 16) and bookstack
(MariaDB 12.3), each deployed for this drill, planted through the app's own interface, destroyed
for real, and restored through the exact endpoint the UI's button posts to.
The three answers
Q1 — does it complete? YES. All five legs ran in order and all succeeded — 32 s for Postgres, 25 s for MariaDB. Data back, apps healthy, accented names byte-identical in both directions.
Q2 — which leg returned the data? The SQL dump. A three-way discriminator (volume tar
ORIGINAL-VALUE-A, altered dump ALTERED-VALUE-B, live LIVE-VALUE-C3) returned ALTERED-VALUE-B.
The ordering the code comment asserts holds in practice. This confirms R-164's F17 claim on a second
path — R-164 cites restore_unit.go, the local restore; this measures offbox_reconstitute.go.
The mutation was applied to the prepared scratch only, and the store was proved unmutated afterwards
by re-preparing a fresh scratch (sha256 back to c5414f24…).
Q3 — does a failure tell the truth? Partly. The customer does see a failure and the undo copy is named. But two things are wrong, and they are the drill's findings.
Findings filed — R-379 … R-382
- R-379 (HIGH) — the undo copy is valid, is named, and nothing in the product can apply it.
Proven by applying it by hand on both engines and getting the exact prior state back.
pre-restore-files are deliberately skipped at three code sites; the filename appears only inside an error string. - R-380 (HIGH) — a failed MariaDB replay leaves a partially-applied database behind an app
reporting
health=healthy, running=true, restarts=0.bookstack's schema-version ledger was wiped to 0 rows while its user data stayed intact and the dashboard said fine. Postgres, by contrast, fails visibly (crash-loop). H3 fired — but not in its predicted shape: the prediction was a quiet success; what happens is a loud error and a silent inconsistency. - R-381 (MEDIUM) — the failure message pastes raw engine stderr into the Hungarian customer
surface: 407 bytes for Postgres, 615 for MariaDB, whose middle is an
INSERT INTO migrations VALUES (…)listing — actual table rows shown to the customer. - R-382 (LOW) — the reconstitution's summary log omits the volume count it already has. The customer-facing flash names the volumes; the operator log does not.
Register: 325 236 bytes before, 330 683 after. Ceiling was R-378; next free id is now R-383.
Also recorded
- R-361 reproduced independently on a second app: after the first reconstitution docmost's
db-dumps/held onlypre-restore-*files. Not re-filed — noted as corroboration. restic checkpassed at the end:no errors were found, 29 snapshots.- The
-dbsuffix attribution is correct forbookstack-db— the R-355 shape does not reproduce. - Observed, not filed: a newly deployed app is absent from the off-site set until switched on by hand. Plausibly deliberate; the consequence is stated so the default can be judged.
- A flaw in the drill's own method, recorded rather than hidden: the first accented title was double-escaped by a shell chain and stored as literal ASCII. Caught by reading the stored bytes back as hex, and re-measured properly in Phase 1b.
Capability map
The 2026-08-21 narrowing of "A customer's file survives a machine rebuild and comes back" is now history: both defects it named (R-354, R-356) are closed and proven. The row records what is now walked — including this drill — and states plainly what is still not claimed: the success path is proven, the recovery-from-a-bad-restore path is not.
Teardown, three layers
docmostandbookstackwere deployed by this drill and are RETAINED with their planted data — it is the evidence, and they are the only deployed members of this app class on the box. Both left healthy and sane.- No
pvesm"before" snapshot was taken — said plainly rather than reconstructed. Measured directly: ~233 MB of volumes inside guest 9201 (16 % of 69 GB used). - No hub-side record was created. No customer, no appliance. Nothing to dispose of.
All phases were run. Nothing was dropped.