R-379/R-380: put the customer's undo copy back when a database restore fails
gates / gates (push) Successful in 11s

R-379 and R-380 were one failure. Both ended with a half-restored database; the
only difference was whether it looked broken. Postgres emptied and crash-looped;
MariaDB applied part of the dump and reported health=healthy with a zero-row
schema-version table. Measured live on demo-hp 2026-08-22.

The undo copy was already taken and already good - proven by hand that day on
both engines. Nothing in the product could apply it. Now it does, with the same
ImportDump call, before any restart and inside the DB-only window.

The WHOLE undo set, matched on this run's stamp. writeSafetyDump returned one
path for an app with two databases; a rollback on that would restore one and
leave the other half-written.

When the rollback also fails the app is HELD STOPPED (operator ruling): a running
app on a half-written database lets the customer make the damage permanent. Every
start path refuses it - customer button, appstop Recover, boot sweep - via the
shared driveStartGate, checked ABOVE its driveless early return because these
apps have no drive. The marker is ended so nothing auto-restarts it. The row goes
red. Cleared with --clear-restore-hold, an operator CLI route.

--single-transaction is a belt on Postgres only; MariaDB DDL is not transactional
and that is why the rollback is the fix.

R-381: the engine's stderr stops reaching the customer (615 bytes on MariaDB, its
middle rows out of their own database) and starts reaching the operator log,
which never had it.
R-382: the summary log prints the volume count it already held.
Undo copies resolve to their own app, are marked IsUndo, and are capped at 3 per
app, pruned from the capture side. The reported render-as-an-app symptom did NOT
reproduce - the live page was read first and had zero occurrences.

Tests 1468 -> 1483. Eight red-proofs; ONE PASSED and is reported: the R-381
behavioural test injected below ImportDump. A guard at that layer now convicts.
This commit is contained in:
2026-08-22 18:03:18 +02:00
parent 0f3cf0dbb2
commit 2c724c9283
15 changed files with 1288 additions and 37 deletions
+67 -5
View File
@@ -75,6 +75,42 @@ type DumpFileInfo struct {
Size int64
ModTime time.Time
Validation DumpValidation
// IsUndo (R-379, v0.220.0) marks a `pre-restore-` safety copy: the undo taken immediately before
// a reconstitution, not the app's own backup. Its VISIBILITY is a recorded decision — see the
// comment on preRestoreDumpPrefix, "an undo the customer cannot see is not much of one" — so this
// flag names it rather than hiding it.
//
// It also fixes a real, if latent, naming defect: the filename parse below used to derive
// StackName by trimming only the engine suffix, so
// `pre-restore-20260822T140924Z-docmost-postgres.sql` yielded the phantom stack
// `pre-restore-20260822T140924Z-docmost`. That name reaches the `dbStacks` lookup in
// web.buildAppBackupRows as a KEY. It renders nothing today because that function iterates
// DEPLOYED apps and only reads the map by key — verified on the live page 2026-08-22, which
// contained zero `pre-restore` strings while four such files sat on disk — but any future code
// that RANGES over that map would surface a stack that does not exist.
IsUndo bool
// UndoAt is when the undo copy was taken, parsed from the stamp in its own filename. Zero for a
// normal dump.
UndoAt time.Time
}
// UndoDumpPrefix is the marker on a pre-restore safety copy. It is duplicated from
// backup.preRestoreDumpPrefix on purpose — appbackup must not import backup (that direction is the
// dependency, not this one) — and the two are pinned equal by a test.
const UndoDumpPrefix = "pre-restore-"
// trimUndoPrefix splits `pre-restore-<stamp>-<rest>` into (<rest>, <stamp>, true). Returns
// (base, "", false) for anything that is not an undo copy, so a normal dump flows unchanged.
func trimUndoPrefix(base string) (string, string, bool) {
if !strings.HasPrefix(base, UndoDumpPrefix) {
return base, "", false
}
after := strings.TrimPrefix(base, UndoDumpPrefix)
i := strings.Index(after, "-")
if i <= 0 {
return base, "", false // a prefix with no stamp is not the shape writeSafetyDump writes
}
return after[i+1:], after[:i], true
}
// DiscoverDatabases finds running database containers via docker ps.
@@ -571,6 +607,15 @@ func ListDumpFiles(dumpDir string, cached func(name string, size int64, mod time
// Parse stack name and DB type from filename: "paperless-ngx-postgres.sql"
base := strings.TrimSuffix(e.Name(), ".sql")
// R-379: strip the `pre-restore-<stamp>-` head FIRST, so an undo copy resolves to the app it
// belongs to instead of a phantom stack named after its own timestamp.
if rest, stamp, ok := trimUndoPrefix(base); ok {
base = rest
f.IsUndo = true
if t, err := time.Parse("20060102T150405Z", stamp); err == nil {
f.UndoAt = t
}
}
if strings.HasSuffix(base, "-postgres") {
f.StackName = strings.TrimSuffix(base, "-postgres")
f.DBType = DBTypePostgres
@@ -680,8 +725,15 @@ func ImportDump(ctx context.Context, db DiscoveredDB, dumpPath string, logger *l
dbName = user
}
// ON_ERROR_STOP=1: a real import error must FAIL (and surface), not silently half-apply.
//
// --single-transaction (R-380, v0.220.0) is a BELT, not the fix. It makes a Postgres replay
// all-or-nothing at the engine, so the emptied-and-crash-looping state measured on `docmost`
// on 2026-08-22 stops being reachable on this engine. It does NOT make MariaDB atomic —
// MariaDB's DDL is not transactional, so a partial apply there is unavoidable at the engine
// and is why the rollback in offbox_reconstitute.go exists and is the actual fix. Do not
// read this flag as making the rollback optional.
cmd = exec.CommandContext(impCtx, "docker", "exec", "-i", db.ContainerID,
"psql", "-v", "ON_ERROR_STOP=1", "-U", user, "-d", dbName)
"psql", "-v", "ON_ERROR_STOP=1", "--single-transaction", "-U", user, "-d", dbName)
case DBTypeMariaDB:
password := getMariaDBPassword(impCtx, db.ContainerID)
if password == "" {
@@ -700,11 +752,21 @@ func ImportDump(ctx context.Context, db DiscoveredDB, dumpPath string, logger *l
logger.Printf("[DEBUG] [backup] ImportDump: importing %s into %s (%s)", dumpPath, db.ContainerName, db.DBType)
}
if err := cmd.Run(); err != nil {
msg := strings.TrimSpace(stderr.String())
if len(msg) > 300 {
msg = msg[:300]
// R-381: the engine's stderr goes to the OPERATOR LOG, in full. It used to be truncated to
// 300 chars and folded into the returned error, which the restore surface then rendered to
// the customer — measured 2026-08-22 at 407 bytes (Postgres, a caret diagram and
// `exit status 3`) and 615 bytes (MariaDB, whose middle was an `INSERT INTO migrations
// VALUES (...)` listing, i.e. ROWS OUT OF THE CUSTOMER'S OWN DATABASE, HTML-escaped, on
// their dashboard).
//
// The diagnostic is ADDED here, not removed: previously nothing logged the full text at all,
// so this is strictly more for the operator and strictly less for the customer.
full := strings.TrimSpace(stderr.String())
if logger != nil {
logger.Printf("[ERROR] [backup] %s import into %s FAILED: %v — engine output follows:\n%s",
db.DBType, db.ContainerName, err, full)
}
return fmt.Errorf("%s import into %s failed: %s — %w", db.DBType, db.ContainerName, msg, err)
return fmt.Errorf("%s import into %s failed: %w", db.DBType, db.ContainerName, err)
}
if logger != nil {
logger.Printf("[INFO] [backup] Imported DB dump %s into %s (%s)", filepath.Base(dumpPath), db.ContainerName, db.DBType)