R-379/R-380 docs: the failure ladder, the drill record, register housekeeping
gates / gates (push) Successful in 17s
gates / gates (push) Successful in 17s
07-backup-architecture.md 6.3 gains a dated [DESIGN] paragraph on replay -> rollback -> hold, including why no engine flag closes it: --single-transaction makes Postgres atomic, MariaDB DDL is not transactional, so the rollback is the fix and the flag is a belt. Drill record for the live walk, including the TWO defects the walk found in the fix itself (a rollback into a re-created container; an operator route that cleared the file while the running controller kept refusing) and the ONE red-proof that PASSED, which is reported rather than omitted. R-379..R-382 compressed into CLOSED-ITEMS.md. OPEN-ITEMS 330683 -> 325236 bytes. STATUS.md restates the outcome and names the next operator step.
This commit is contained in:
@@ -352,6 +352,36 @@ above stands exactly as written: the secondary unit mirror is still read by noth
|
||||
where the primary drive is lost — and in exactly that case the primary unit is gone while this
|
||||
mirror survives on the second drive, unreachable by any customer action.
|
||||
|
||||
**[DESIGN] 2026-08-22 — the failure ladder of a database restore: replay → rollback → hold.**
|
||||
Recorded here rather than only in a closed register row, because a decision that survives only inside
|
||||
a closed work item is a decision nobody will find.
|
||||
|
||||
1. **Replay.** The snapshot's `.sql` is imported into the app's database, with only the DB service up
|
||||
(R-47). A pre-restore copy of the LIVE database was taken first and is on disk (R-43); the restore
|
||||
refuses outright if it could not be taken.
|
||||
2. **Rollback (controller v0.220.0, R-379).** If the replay fails, the product re-applies that undo
|
||||
copy itself. **The whole set for this run** — an app with two databases gets two undo files, and
|
||||
restoring only the first would leave the other half-written — matched on **the run's own stamp**,
|
||||
never on the `pre-restore-` prefix, because several runs' copies coexist in the same directory. It
|
||||
runs with the DB service still up and before any restart, so the app never observes the half
|
||||
state, and into a **re-discovered** container: the DB-only start re-creates it, so the id captured
|
||||
at dump time is dead by rollback time (v0.220.1, found by the first live run). The app then starts
|
||||
and the customer is told **both** that the restore failed and that their data is as it was.
|
||||
3. **Hold (v0.220.0, operator ruling 2026-08-22).** If the rollback ALSO fails, the app is **held
|
||||
stopped**, not started. A running app on a half-written database lets the customer type into it and
|
||||
makes the damage permanent. The hold is persisted, every start path refuses it with a reason and a
|
||||
route, the app-stop marker is ended so nothing auto-restarts it at the next boot, and the app reads
|
||||
**red** rather than green. An operator clears it with `--clear-restore-hold`, which requires a
|
||||
controller restart.
|
||||
|
||||
**Why a rollback and not an engine flag.** Postgres gained `--single-transaction` in the same release
|
||||
and that does make its replay all-or-nothing — but **MariaDB's DDL is not transactional**, so a
|
||||
partial apply there is unavoidable at the engine. Measured 2026-08-22: the same truncated dump left
|
||||
Postgres emptied and crash-looping, and left MariaDB with its user data intact, its schema-version
|
||||
table wiped to zero rows, and the app reporting `health=healthy, running=true, restarts=0`. The flag
|
||||
is a belt; the rollback is the fix. **R-379, R-380.** Evidence:
|
||||
`audits/DRILL-r379-rollback-2026-08-22/`.
|
||||
|
||||
**[DESIGN] 2026-08-22 — the restore destination is resolved by the same rule as the capture
|
||||
destination.** The drive if the app declares one (`HDD_PATH`), the system data path otherwise —
|
||||
`Manager.GetAppDrivePath`, one expression, used by `CaptureRecoveryUnit` and, since controller
|
||||
|
||||
Reference in New Issue
Block a user