R-379/R-380: put the customer's undo copy back when a database restore fails
gates / gates (push) Successful in 11s

R-379 and R-380 were one failure. Both ended with a half-restored database; the
only difference was whether it looked broken. Postgres emptied and crash-looped;
MariaDB applied part of the dump and reported health=healthy with a zero-row
schema-version table. Measured live on demo-hp 2026-08-22.

The undo copy was already taken and already good - proven by hand that day on
both engines. Nothing in the product could apply it. Now it does, with the same
ImportDump call, before any restart and inside the DB-only window.

The WHOLE undo set, matched on this run's stamp. writeSafetyDump returned one
path for an app with two databases; a rollback on that would restore one and
leave the other half-written.

When the rollback also fails the app is HELD STOPPED (operator ruling): a running
app on a half-written database lets the customer make the damage permanent. Every
start path refuses it - customer button, appstop Recover, boot sweep - via the
shared driveStartGate, checked ABOVE its driveless early return because these
apps have no drive. The marker is ended so nothing auto-restarts it. The row goes
red. Cleared with --clear-restore-hold, an operator CLI route.

--single-transaction is a belt on Postgres only; MariaDB DDL is not transactional
and that is why the rollback is the fix.

R-381: the engine's stderr stops reaching the customer (615 bytes on MariaDB, its
middle rows out of their own database) and starts reaching the operator log,
which never had it.
R-382: the summary log prints the volume count it already held.
Undo copies resolve to their own app, are marked IsUndo, and are capped at 3 per
app, pruned from the capture side. The reported render-as-an-app symptom did NOT
reproduce - the live page was read first and had zero occurrences.

Tests 1468 -> 1483. Eight red-proofs; ONE PASSED and is reported: the R-381
behavioural test injected below ImportDump. A guard at that layer now convicts.
This commit is contained in:
2026-08-22 18:03:18 +02:00
parent 0f3cf0dbb2
commit 2c724c9283
15 changed files with 1288 additions and 37 deletions
+63
View File
@@ -1,3 +1,66 @@
## v0.220.0 — when a database restore fails, the customer's own copy goes back (2026-08-22, R-379/R-380/R-381/R-382)
**MinAgent: 0.129.0** (unchanged — no new agent coupling)
**R-379 and R-380 were ONE failure and they get ONE fix.** Both ended with a half-restored database.
The only difference was whether it looked broken: on Postgres the database was emptied and the app
crash-looped; on MariaDB part of the dump applied, the rest did not, and the app reported
`health=healthy, running=true, restarts=0` with its schema-version table holding zero rows. Measured
live on `demo-hp` on 2026-08-22 (`audits/DRILL-r356b-driveless-db-restore-2026-08-22/`).
**The undo copy was already being taken, and was already good.** It was proven good by hand that day
on both engines — `docmost` recovered, `bookstack`'s `migrations` went 0 → 102 rows. What no product
action could do was apply it: `pre-restore-` files are skipped at three call sites so they are never
mistaken for a replay source, and the filename appeared only inside an error string. **The fix is the
product making the same `ImportDump` call a person made by hand.**
**When the replay fails, the undo set is now re-applied automatically**, before any restart and with
the DB service still up, so the app never observes the half state. The app then starts and the
message says **both** things: the restore failed, *and* the data is back as it was. A message that
reported only the failure would leave the customer believing their data was gone when it is not — the
omission of a gain misleads exactly as much as the omission of a loss.
**THE WHOLE undo set, not the first file.** `writeSafetyDump` returned one path for an app with two
databases; a rollback built on that would have restored one and left the other half-written — this
defect, one database over. It now returns the set, matched on **this run's stamp**: four `pre-restore-`
files accumulated on one app in one afternoon, so a prefix match would replay an arbitrary older state.
**When the rollback ALSO fails, the app is HELD STOPPED — operator ruling, 2026-08-22.** A running app
on a half-written database lets the customer type into it and turns a recoverable state into a
permanent one. The hold is persisted, every start path refuses it (the customer's button, the app-stop
`Recover()` starter, and the boot sweep — through the shared `driveStartGate`, checked **above** its
driveless early return because these apps have no drive), the app-stop marker is ended so nothing
auto-restarts it at the next boot, the row goes **red** rather than green, and the operator is told.
Clearing it is `--clear-restore-hold <app>`, an operator CLI route reached through `docker exec`.
**`--single-transaction` is a BELT, not the fix.** Added to the Postgres import so a replay is
all-or-nothing at the engine. **It does NOT make MariaDB atomic** — MariaDB's DDL is not
transactional, so a partial apply there is unavoidable at the engine level. That is precisely why the
rollback exists, and this flag does not make it optional.
**R-381 — the failure message stops pasting the engine's output at the customer.** It was 407 bytes on
Postgres (a caret diagram and `exit status 3`) and **615 on MariaDB, whose middle was an `INSERT INTO
migrations VALUES (…)` listing — rows out of the customer's own database, HTML-escaped, on their
dashboard.** The full engine text now goes to the operator log, which never had it before: the
diagnostic is **added**, not removed.
**R-382 — the summary log line prints the volume count it already held.** It said "0 file(s) placed,
1 DB dump(s) replayed" on a run that returned a 52 MB Postgres data directory.
**The undo copies are named for what they are, and bounded.** `pre-restore-20260822T140924Z-docmost-postgres.sql`
used to derive the phantom stack `pre-restore-20260822T140924Z-docmost`; it now resolves to `docmost`
and carries `IsUndo`. **The reported symptom — that they render as apps on the customer's backup page
— did NOT reproduce**: the live page was read first and contained zero `pre-restore` strings, because
`buildAppBackupRows` iterates deployed apps and only reads that map by key. The phantom name was real
as a map KEY; the row was not. Their VISIBILITY is unchanged and deliberate. A cap of **3 per app**
now applies, pruned from the capture side and never from the restore path — a delete on the failure
path is how an undo goes missing at the moment it is needed.
**Tests:** `internal/backup/r379_rollback_test.go`, `cmd/controller/r379_hold_gate_test.go`,
`internal/appbackup/r381_undo_naming_test.go`. Test count 1468 → 1483. **Eight red-proofs; one PASSED
and is reported rather than omitted** — the R-381 behavioural test injected below `ImportDump` and so
could not see a leak reintroduced inside it. A guard at that layer was added and the mutation then
convicted.
## v0.219.0 — the off-site restore refused every app that has no data drive (2026-08-22, R-356)
**MinAgent: 0.129.0** (unchanged — no new agent coupling)