R-379/R-380: put the customer's undo copy back when a database restore fails
gates / gates (push) Successful in 11s
gates / gates (push) Successful in 11s
R-379 and R-380 were one failure. Both ended with a half-restored database; the only difference was whether it looked broken. Postgres emptied and crash-looped; MariaDB applied part of the dump and reported health=healthy with a zero-row schema-version table. Measured live on demo-hp 2026-08-22. The undo copy was already taken and already good - proven by hand that day on both engines. Nothing in the product could apply it. Now it does, with the same ImportDump call, before any restart and inside the DB-only window. The WHOLE undo set, matched on this run's stamp. writeSafetyDump returned one path for an app with two databases; a rollback on that would restore one and leave the other half-written. When the rollback also fails the app is HELD STOPPED (operator ruling): a running app on a half-written database lets the customer make the damage permanent. Every start path refuses it - customer button, appstop Recover, boot sweep - via the shared driveStartGate, checked ABOVE its driveless early return because these apps have no drive. The marker is ended so nothing auto-restarts it. The row goes red. Cleared with --clear-restore-hold, an operator CLI route. --single-transaction is a belt on Postgres only; MariaDB DDL is not transactional and that is why the rollback is the fix. R-381: the engine's stderr stops reaching the customer (615 bytes on MariaDB, its middle rows out of their own database) and starts reaching the operator log, which never had it. R-382: the summary log prints the volume count it already held. Undo copies resolve to their own app, are marked IsUndo, and are capped at 3 per app, pruned from the capture side. The reported render-as-an-app symptom did NOT reproduce - the live page was read first and had zero occurrences. Tests 1468 -> 1483. Eight red-proofs; ONE PASSED and is reported: the R-381 behavioural test injected below ImportDump. A guard at that layer now convicts.
This commit is contained in:
@@ -1,3 +1,66 @@
|
||||
## v0.220.0 — when a database restore fails, the customer's own copy goes back (2026-08-22, R-379/R-380/R-381/R-382)
|
||||
**MinAgent: 0.129.0** (unchanged — no new agent coupling)
|
||||
|
||||
**R-379 and R-380 were ONE failure and they get ONE fix.** Both ended with a half-restored database.
|
||||
The only difference was whether it looked broken: on Postgres the database was emptied and the app
|
||||
crash-looped; on MariaDB part of the dump applied, the rest did not, and the app reported
|
||||
`health=healthy, running=true, restarts=0` with its schema-version table holding zero rows. Measured
|
||||
live on `demo-hp` on 2026-08-22 (`audits/DRILL-r356b-driveless-db-restore-2026-08-22/`).
|
||||
|
||||
**The undo copy was already being taken, and was already good.** It was proven good by hand that day
|
||||
on both engines — `docmost` recovered, `bookstack`'s `migrations` went 0 → 102 rows. What no product
|
||||
action could do was apply it: `pre-restore-` files are skipped at three call sites so they are never
|
||||
mistaken for a replay source, and the filename appeared only inside an error string. **The fix is the
|
||||
product making the same `ImportDump` call a person made by hand.**
|
||||
|
||||
**When the replay fails, the undo set is now re-applied automatically**, before any restart and with
|
||||
the DB service still up, so the app never observes the half state. The app then starts and the
|
||||
message says **both** things: the restore failed, *and* the data is back as it was. A message that
|
||||
reported only the failure would leave the customer believing their data was gone when it is not — the
|
||||
omission of a gain misleads exactly as much as the omission of a loss.
|
||||
|
||||
**THE WHOLE undo set, not the first file.** `writeSafetyDump` returned one path for an app with two
|
||||
databases; a rollback built on that would have restored one and left the other half-written — this
|
||||
defect, one database over. It now returns the set, matched on **this run's stamp**: four `pre-restore-`
|
||||
files accumulated on one app in one afternoon, so a prefix match would replay an arbitrary older state.
|
||||
|
||||
**When the rollback ALSO fails, the app is HELD STOPPED — operator ruling, 2026-08-22.** A running app
|
||||
on a half-written database lets the customer type into it and turns a recoverable state into a
|
||||
permanent one. The hold is persisted, every start path refuses it (the customer's button, the app-stop
|
||||
`Recover()` starter, and the boot sweep — through the shared `driveStartGate`, checked **above** its
|
||||
driveless early return because these apps have no drive), the app-stop marker is ended so nothing
|
||||
auto-restarts it at the next boot, the row goes **red** rather than green, and the operator is told.
|
||||
Clearing it is `--clear-restore-hold <app>`, an operator CLI route reached through `docker exec`.
|
||||
|
||||
**`--single-transaction` is a BELT, not the fix.** Added to the Postgres import so a replay is
|
||||
all-or-nothing at the engine. **It does NOT make MariaDB atomic** — MariaDB's DDL is not
|
||||
transactional, so a partial apply there is unavoidable at the engine level. That is precisely why the
|
||||
rollback exists, and this flag does not make it optional.
|
||||
|
||||
**R-381 — the failure message stops pasting the engine's output at the customer.** It was 407 bytes on
|
||||
Postgres (a caret diagram and `exit status 3`) and **615 on MariaDB, whose middle was an `INSERT INTO
|
||||
migrations VALUES (…)` listing — rows out of the customer's own database, HTML-escaped, on their
|
||||
dashboard.** The full engine text now goes to the operator log, which never had it before: the
|
||||
diagnostic is **added**, not removed.
|
||||
|
||||
**R-382 — the summary log line prints the volume count it already held.** It said "0 file(s) placed,
|
||||
1 DB dump(s) replayed" on a run that returned a 52 MB Postgres data directory.
|
||||
|
||||
**The undo copies are named for what they are, and bounded.** `pre-restore-20260822T140924Z-docmost-postgres.sql`
|
||||
used to derive the phantom stack `pre-restore-20260822T140924Z-docmost`; it now resolves to `docmost`
|
||||
and carries `IsUndo`. **The reported symptom — that they render as apps on the customer's backup page
|
||||
— did NOT reproduce**: the live page was read first and contained zero `pre-restore` strings, because
|
||||
`buildAppBackupRows` iterates deployed apps and only reads that map by key. The phantom name was real
|
||||
as a map KEY; the row was not. Their VISIBILITY is unchanged and deliberate. A cap of **3 per app**
|
||||
now applies, pruned from the capture side and never from the restore path — a delete on the failure
|
||||
path is how an undo goes missing at the moment it is needed.
|
||||
|
||||
**Tests:** `internal/backup/r379_rollback_test.go`, `cmd/controller/r379_hold_gate_test.go`,
|
||||
`internal/appbackup/r381_undo_naming_test.go`. Test count 1468 → 1483. **Eight red-proofs; one PASSED
|
||||
and is reported rather than omitted** — the R-381 behavioural test injected below `ImportDump` and so
|
||||
could not see a leak reintroduced inside it. A guard at that layer was added and the mutation then
|
||||
convicted.
|
||||
|
||||
## v0.219.0 — the off-site restore refused every app that has no data drive (2026-08-22, R-356)
|
||||
**MinAgent: 0.129.0** (unchanged — no new agent coupling)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user