R-379/R-380: put the customer's undo copy back when a database restore fails
gates / gates (push) Successful in 11s

R-379 and R-380 were one failure. Both ended with a half-restored database; the
only difference was whether it looked broken. Postgres emptied and crash-looped;
MariaDB applied part of the dump and reported health=healthy with a zero-row
schema-version table. Measured live on demo-hp 2026-08-22.

The undo copy was already taken and already good - proven by hand that day on
both engines. Nothing in the product could apply it. Now it does, with the same
ImportDump call, before any restart and inside the DB-only window.

The WHOLE undo set, matched on this run's stamp. writeSafetyDump returned one
path for an app with two databases; a rollback on that would restore one and
leave the other half-written.

When the rollback also fails the app is HELD STOPPED (operator ruling): a running
app on a half-written database lets the customer make the damage permanent. Every
start path refuses it - customer button, appstop Recover, boot sweep - via the
shared driveStartGate, checked ABOVE its driveless early return because these
apps have no drive. The marker is ended so nothing auto-restarts it. The row goes
red. Cleared with --clear-restore-hold, an operator CLI route.

--single-transaction is a belt on Postgres only; MariaDB DDL is not transactional
and that is why the rollback is the fix.

R-381: the engine's stderr stops reaching the customer (615 bytes on MariaDB, its
middle rows out of their own database) and starts reaching the operator log,
which never had it.
R-382: the summary log prints the volume count it already held.
Undo copies resolve to their own app, are marked IsUndo, and are capped at 3 per
app, pruned from the capture side. The reported render-as-an-app symptom did NOT
reproduce - the live page was read first and had zero occurrences.

Tests 1468 -> 1483. Eight red-proofs; ONE PASSED and is reported: the R-381
behavioural test injected below ImportDump. A guard at that layer now convicts.
This commit is contained in:
2026-08-22 18:03:18 +02:00
parent 0f3cf0dbb2
commit 2c724c9283
15 changed files with 1288 additions and 37 deletions
+15
View File
@@ -1157,6 +1157,11 @@ type AppBackupRow struct {
// e.g., "DB + Konfiguráció + Adatok", "DB + Konfiguráció", "Konfiguráció"
BackupContents string
// RestoreHeld (R-379) — this app is deliberately stopped because a database restore failed AND
// the rollback failed. Distinct from any backup status: it is about the app's LIVE data, not its
// copies, which is why it drives the row red rather than yellow.
RestoreHeld bool
// Tier 1: Nightly backup (always exists)
Tier1LastRun string // RFC3339 time of the newest recovery-unit artifact ("" = no unit yet)
Tier1LastStatus string // "ok", "error", ""
@@ -1361,6 +1366,16 @@ func (s *Server) buildAppBackupRows(status *backup.FullBackupStatus) []AppBackup
row.Status = "yellow"
row.StatusText = "Adatbázis mentés sikertelen"
}
// R-379/R-380: a HELD app must never read as healthy. Last, so it wins over both branches
// above — a warning beside a green tick is read as a success, and this is the one state where
// the customer's data may not be intact. Measured 2026-08-22: after a failed MariaDB replay
// the app reported `health=healthy, running=true, restarts=0` while its schema-version table
// held zero rows.
if held, why := s.backupMgr.RestoreHoldFor(app.StackName); held {
row.Status = "red"
row.StatusText = why
row.RestoreHeld = true
}
// Tier 2 (off-drive copy) status, from the config the Tier 2 runner persists.
if cd := s.settings.GetCrossDriveConfig(app.StackName); cd != nil {