R-379/R-380 docs: the failure ladder, the drill record, register housekeeping
gates / gates (push) Successful in 17s

07-backup-architecture.md 6.3 gains a dated [DESIGN] paragraph on replay ->
rollback -> hold, including why no engine flag closes it: --single-transaction
makes Postgres atomic, MariaDB DDL is not transactional, so the rollback is the
fix and the flag is a belt.

Drill record for the live walk, including the TWO defects the walk found in the
fix itself (a rollback into a re-created container; an operator route that
cleared the file while the running controller kept refusing) and the ONE
red-proof that PASSED, which is reported rather than omitted.

R-379..R-382 compressed into CLOSED-ITEMS.md. OPEN-ITEMS 330683 -> 325236 bytes.

STATUS.md restates the outcome and names the next operator step.
This commit is contained in:
2026-08-22 18:31:23 +02:00
parent 4e488321bf
commit a8caa0fdde
23 changed files with 744 additions and 75 deletions
@@ -352,6 +352,36 @@ above stands exactly as written: the secondary unit mirror is still read by noth
where the primary drive is lost — and in exactly that case the primary unit is gone while this
mirror survives on the second drive, unreachable by any customer action.
**[DESIGN] 2026-08-22 — the failure ladder of a database restore: replay → rollback → hold.**
Recorded here rather than only in a closed register row, because a decision that survives only inside
a closed work item is a decision nobody will find.
1. **Replay.** The snapshot's `.sql` is imported into the app's database, with only the DB service up
(R-47). A pre-restore copy of the LIVE database was taken first and is on disk (R-43); the restore
refuses outright if it could not be taken.
2. **Rollback (controller v0.220.0, R-379).** If the replay fails, the product re-applies that undo
copy itself. **The whole set for this run** — an app with two databases gets two undo files, and
restoring only the first would leave the other half-written — matched on **the run's own stamp**,
never on the `pre-restore-` prefix, because several runs' copies coexist in the same directory. It
runs with the DB service still up and before any restart, so the app never observes the half
state, and into a **re-discovered** container: the DB-only start re-creates it, so the id captured
at dump time is dead by rollback time (v0.220.1, found by the first live run). The app then starts
and the customer is told **both** that the restore failed and that their data is as it was.
3. **Hold (v0.220.0, operator ruling 2026-08-22).** If the rollback ALSO fails, the app is **held
stopped**, not started. A running app on a half-written database lets the customer type into it and
makes the damage permanent. The hold is persisted, every start path refuses it with a reason and a
route, the app-stop marker is ended so nothing auto-restarts it at the next boot, and the app reads
**red** rather than green. An operator clears it with `--clear-restore-hold`, which requires a
controller restart.
**Why a rollback and not an engine flag.** Postgres gained `--single-transaction` in the same release
and that does make its replay all-or-nothing — but **MariaDB's DDL is not transactional**, so a
partial apply there is unavoidable at the engine. Measured 2026-08-22: the same truncated dump left
Postgres emptied and crash-looping, and left MariaDB with its user data intact, its schema-version
table wiped to zero rows, and the app reporting `health=healthy, running=true, restarts=0`. The flag
is a belt; the rollback is the fix. **R-379, R-380.** Evidence:
`audits/DRILL-r379-rollback-2026-08-22/`.
**[DESIGN] 2026-08-22 — the restore destination is resolved by the same rule as the capture
destination.** The drive if the app declares one (`HDD_PATH`), the system data path otherwise —
`Manager.GetAppDrivePath`, one expression, used by `CaptureRecoveryUnit` and, since controller