DRILL R-356b: the off-site restore for a driveless app that HAS a database
gates / gates (push) Successful in 16s

A drill, not an implementation. No code, no version bump, no CHANGELOG entry.

Ten of the forty driveless apps carry a database; I re-measured that count and
got 10. For those ten the restore is a five-leg operation that never ran at all
until this week, because R-356 refused before any of it started.

Walked end to end on demo-hp for both engines - docmost (Postgres 16) and
bookstack (MariaDB 12.3) - each deployed for the drill, planted through the
app's own interface, destroyed for real, restored through the endpoint the UI
posts to.

Q1 does it complete: YES. All five legs ran and succeeded, 32s / 25s. Accented
names byte-identical both directions.

Q2 which leg won: the SQL DUMP. Three-way discriminator returned the altered
dump's value. This confirms R-164's F17 ordering on the OFF-SITE path; R-164
only ever cited the local one. Scratch-only mutation; store proved unmutated.

Q3 does a failure tell the truth: partly, and two defects.

Filed R-379 (HIGH, the undo copy is valid, named, and unappliable by any product
action - proven by applying it by hand on both engines), R-380 (HIGH, a failed
MariaDB replay leaves a partial database behind an app reporting healthy, where
Postgres crash-loops visibly), R-381 (MEDIUM, the failure message pastes engine
stderr including customer table rows into the Hungarian surface), R-382 (LOW,
the summary log omits the volume count it already has).

H1, H2 and H4 did NOT fire and that is recorded. H3 fired in a shape nobody
predicted: not a quiet success, but a loud error over a silent inconsistency.

R-361 reproduced independently on a second app. restic check: no errors, 29
snapshots. A flaw in the drill's own planting - a double-escaped accented title -
was caught by reading stored bytes as hex, recorded, and re-measured in Phase 1b.

Register 325236 -> 330683 bytes. Nothing dropped. Teardown: two apps retained
with reason, no pvesm before-snapshot taken (said plainly), no hub-side record
created.
This commit is contained in:
2026-08-22 16:33:45 +02:00
parent 8c9f1b798b
commit 4e488321bf
36 changed files with 747 additions and 54 deletions
+61 -53
View File
@@ -1,72 +1,80 @@
# REPORT — R-356 doc corrections, register housekeeping, and one gate fix (2026-08-22)
# REPORT — DRILL R-356b: the off-site restore for a driveless app that HAS a database (2026-08-22)
Companion to `felhom-controller` v0.219.0 (R-356). This repo carried the architecture correction, the
drill record, the register move and one genuine gate defect found on the way.
**A drill, not an implementation.** No production code was written, no version bumped, no CHANGELOG
entry made. The deliverables are a findings document, four register rows and a capability-map update.
## 1. `documentation/architecture/07-backup-architecture.md` — R-107 was closed and the doc said otherwise
Full record: `documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/`
Three places, each corrected with a dated **[FACT]** citing `offbox_reconstitute.go` `volReplay` and
controller v0.218.0:
## What was measured
- §6.3, the Tier-3 row — was *"no offsite action unpacks the named-volume tars it captures"*.
- §8, matrix row 4 — was *"the volume tars in either copy are unreachable"*.
- The R-107 index row.
Ten of the forty driveless apps carry a database. **I re-measured that count myself and got 10** — the
same ten the runbook names. For those ten, restoring is a five-leg operation that, until this week,
never ran at all: R-356 refused before any of it started.
**The old sentence's history is kept, not deleted:** each correction says what was true, until when,
and what closed it. A correction that erases what was believed leaves the next reader no way to tell
a fixed gap from one that was never noticed.
Both engines were walked end to end on `demo-hp`: `docmost` (Postgres 16) and `bookstack`
(MariaDB 12.3), each deployed for this drill, planted through the app's **own** interface, destroyed
for real, and restored through the exact endpoint the UI's button posts to.
**R-102 — the Tier-2 half — is NOT closed, and the correction says so explicitly** so it cannot be
read as covering both. The Tier-2 row stands exactly as written.
## The three answers
## 2. §6.3 — the R-356 reasoning recorded as reasoning, not as a closed row
**Q1 — does it complete? YES.** All five legs ran in order and all succeeded — 32 s for Postgres,
25 s for MariaDB. Data back, apps healthy, accented names byte-identical in both directions.
A new **[DESIGN]** paragraph: the restore destination is resolved by the same rule as the capture
destination (drive if the app has one, system data path otherwise); the refusal that protects a drive
app from being restored onto the wrong disk applies to apps that **have a drive to get wrong**. It
carries the 13/40 measurement and points at `felhom-controller/CONTEXT.md`.
**Q2 — which leg returned the data? The SQL dump.** A three-way discriminator (volume tar
`ORIGINAL-VALUE-A`, altered dump `ALTERED-VALUE-B`, live `LIVE-VALUE-C3`) returned **`ALTERED-VALUE-B`**.
The ordering the code comment asserts holds in practice. This **confirms R-164's F17 claim on a second
path** — R-164 cites `restore_unit.go`, the local restore; this measures `offbox_reconstitute.go`.
The mutation was applied to the prepared scratch only, and the store was proved unmutated afterwards
by re-preparing a fresh scratch (sha256 back to `c5414f24…`).
## 3. `STATUS.md` — the contradiction is gone, and the page is one screen
**Q3 — does a failure tell the truth? Partly.** The customer does see a failure and the undo copy is
named. But two things are wrong, and they are the drill's findings.
It said nothing was waiting while also saying one decision was waiting — to publish agent 0.130.0 —
which the same page recorded as already published (R-347, closed). **218 → 102 lines.**
## Findings filed — R-379 … R-382
"Waiting on you" now lists **four** real items, and — per the exemption — **each says what happens if
the operator does nothing**. The two new ones are this release's hand-off: vouch a golden carrying
controller 0.219.0, then raise the floor last, in a separate save.
- **R-379 (HIGH)** — the undo copy is valid, is named, and **nothing in the product can apply it**.
Proven by applying it by hand on both engines and getting the exact prior state back.
`pre-restore-` files are deliberately skipped at three code sites; the filename appears only inside
an error string.
- **R-380 (HIGH)** — a failed **MariaDB** replay leaves a partially-applied database behind an app
reporting `health=healthy, running=true, restarts=0`. `bookstack`'s schema-version ledger was wiped
to 0 rows while its user data stayed intact and the dashboard said fine. Postgres, by contrast,
fails visibly (crash-loop). H3 fired — but not in its predicted shape: the prediction was a *quiet
success*; what happens is a loud error and a silent inconsistency.
- **R-381 (MEDIUM)** — the failure message pastes raw engine stderr into the Hungarian customer
surface: 407 bytes for Postgres, **615 for MariaDB, whose middle is an `INSERT INTO migrations
VALUES (…)` listing — actual table rows shown to the customer.**
- **R-382 (LOW)** — the reconstitution's summary log omits the volume count it already has. The
customer-facing flash names the volumes; the operator log does not.
## 4. `documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/`
**Register: 325 236 bytes before, 330 683 after.** Ceiling was R-378; next free id is now R-383.
The live walk on `demo-hp`: both classes, both messages verbatim with their byte counts, both accented
filenames as explicit hex, and 16 evidence files. **Copied off at the end of each phase.** Nothing was
reverted during this drill, so no intermediate teardown could have taken it.
## Also recorded
## 5. Register housekeeping (N.7)
- **R-361 reproduced independently** on a second app: after the first reconstitution docmost's
`db-dumps/` held only `pre-restore-*` files. Not re-filed — noted as corroboration.
- **`restic check` passed** at the end: `no errors were found`, 29 snapshots.
- **The `-db` suffix attribution is correct** for `bookstack-db` — the R-355 shape does not reproduce.
- **Observed, not filed:** a newly deployed app is absent from the off-site set until switched on by
hand. Plausibly deliberate; the consequence is stated so the default can be judged.
- **A flaw in the drill's own method, recorded rather than hidden:** the first accented title was
double-escaped by a shell chain and stored as literal ASCII. Caught by reading the stored bytes back
as hex, and re-measured properly in Phase 1b.
R-356 compressed out of `OPEN-ITEMS.md` into `CLOSED-ITEMS.md`, keeping its title, shipping version,
evidence path and every sentence stating a rule, plus the pointer
`git show e18668f9e19f:documentation/backlog/OPEN-ITEMS.md` for the full original text.
## Capability map
| file | before | after |
|---|---|---|
| `OPEN-ITEMS.md` | 327 109 bytes | **325 236 bytes** |
| `CLOSED-ITEMS.md` | 61 580 bytes | **63 507 bytes** |
The 2026-08-21 narrowing of *"A customer's file survives a machine rebuild and comes back"* is now
history: both defects it named (R-354, R-356) are closed and proven. The row records what is now
walked — including this drill — and states plainly what is still **not** claimed: the success path is
proven, the recovery-from-a-bad-restore path is not.
## 6. A real gate defect, found by CI going red
## Teardown, three layers
CI run **387** (job 386) failed `instructions_gate` on commit `08eb1a6` while run 73 on its parent had
passed. The cause was not the push: `ef6ac6f` (this repo, the same day) compressed closed rows out of
`OPEN-ITEMS.md` into `CLOSED-ITEMS.md`, and `register_state()` read only `OPEN-ITEMS.md` and
`ROADMAP.md`. **Every citation of a compressed item became "a reference to nothing"**, failing the
next push in a sibling repo for a rule file nobody had touched — and it would have fired again on this
task's own R-356 compression.
1. `docmost` and `bookstack` were deployed by this drill and are **RETAINED** with their planted data —
it is the evidence, and they are the only deployed members of this app class on the box. Both left
healthy and sane.
2. No `pvesm` "before" snapshot was taken — **said plainly rather than reconstructed.** Measured
directly: ~233 MB of volumes inside guest 9201 (16 % of 69 GB used).
3. **No hub-side record was created.** No customer, no appliance. Nothing to dispose of.
`CLOSED-ITEMS.md` is now the third source. It answers *does this ID exist* and answers *closed* for
the rows it owns; `OPEN-ITEMS.md` remains the sole authority on **openness**, so a row it claims as
open is not overridden. **Both controls still convict:** an ID present nowhere fails, and a citation
claiming a closed item is still open fails — each watched failing, then watched clearing.
## Gate state
`python3 scripts/repo_gates.py --fast` — all green **except `golden-currency`**, which is correct and
is operator item 1: controller 0.219.0 is released and no golden carries it.
All phases were run. Nothing was dropped.