R-379/R-380: put the customer's undo copy back when a database restore fails
gates / gates (push) Successful in 11s

R-379 and R-380 were one failure. Both ended with a half-restored database; the
only difference was whether it looked broken. Postgres emptied and crash-looped;
MariaDB applied part of the dump and reported health=healthy with a zero-row
schema-version table. Measured live on demo-hp 2026-08-22.

The undo copy was already taken and already good - proven by hand that day on
both engines. Nothing in the product could apply it. Now it does, with the same
ImportDump call, before any restart and inside the DB-only window.

The WHOLE undo set, matched on this run's stamp. writeSafetyDump returned one
path for an app with two databases; a rollback on that would restore one and
leave the other half-written.

When the rollback also fails the app is HELD STOPPED (operator ruling): a running
app on a half-written database lets the customer make the damage permanent. Every
start path refuses it - customer button, appstop Recover, boot sweep - via the
shared driveStartGate, checked ABOVE its driveless early return because these
apps have no drive. The marker is ended so nothing auto-restarts it. The row goes
red. Cleared with --clear-restore-hold, an operator CLI route.

--single-transaction is a belt on Postgres only; MariaDB DDL is not transactional
and that is why the rollback is the fix.

R-381: the engine's stderr stops reaching the customer (615 bytes on MariaDB, its
middle rows out of their own database) and starts reaching the operator log,
which never had it.
R-382: the summary log prints the volume count it already held.
Undo copies resolve to their own app, are marked IsUndo, and are capped at 3 per
app, pruned from the capture side. The reported render-as-an-app symptom did NOT
reproduce - the live page was read first and had zero occurrences.

Tests 1468 -> 1483. Eight red-proofs; ONE PASSED and is reported: the R-381
behavioural test injected below ImportDump. A guard at that layer now convicts.
This commit is contained in:
2026-08-22 18:03:18 +02:00
parent 0f3cf0dbb2
commit 2c724c9283
15 changed files with 1288 additions and 37 deletions
+20
View File
@@ -514,6 +514,16 @@ func (r *Router) deployStack(w http.ResponseWriter, req *http.Request, name stri
}
}
// restoreHoldFor reports whether an app is held after a failed restore whose rollback also failed,
// and returns the CUSTOMER-facing Hungarian reason. Delegates to the backup manager so the sentence
// has one source; the drive gate's own reason is operator-English and stays that way.
func (r *Router) restoreHoldFor(name string) (bool, string) {
if r.backupMgr == nil {
return false, ""
}
return r.backupMgr.RestoreHoldFor(name)
}
// startGatedByMissingDrive reports whether starting `name` must be BLOCKED because the drive its
// HDD_PATH points at is currently disconnected or decommissioned. Returns the storage path for the
// message. SSD-resident apps (no HDD_PATH) are never gated.
@@ -565,6 +575,16 @@ func (r *Router) actionStack(w http.ResponseWriter, action, name string) {
// Drive-absent gate: refuse to start an app whose data drive is currently disconnected/decommissioned
// (the intermediary-mount gate). Starting it would let it write to the empty fail-closed stable path
// or just crash-loop; block with a clear message until the drive returns (then the gate auto-restarts).
// R-379/R-380: an app held after a failed restore + failed rollback must not start from the
// customer's button either. Checked BEFORE the drive gate because it applies to driveless apps,
// which is the class the hold exists for.
if action == "start" || action == "restart" {
if held, why := r.restoreHoldFor(name); held {
writeJSON(w, http.StatusConflict, apiResponse{OK: false, Error: why})
return
}
}
if action == "start" {
if gated, hdd := r.startGatedByMissingDrive(name); gated {
writeJSON(w, http.StatusConflict, apiResponse{