R-379/R-380: put the customer's undo copy back when a database restore fails
gates / gates (push) Successful in 11s
gates / gates (push) Successful in 11s
R-379 and R-380 were one failure. Both ended with a half-restored database; the only difference was whether it looked broken. Postgres emptied and crash-looped; MariaDB applied part of the dump and reported health=healthy with a zero-row schema-version table. Measured live on demo-hp 2026-08-22. The undo copy was already taken and already good - proven by hand that day on both engines. Nothing in the product could apply it. Now it does, with the same ImportDump call, before any restart and inside the DB-only window. The WHOLE undo set, matched on this run's stamp. writeSafetyDump returned one path for an app with two databases; a rollback on that would restore one and leave the other half-written. When the rollback also fails the app is HELD STOPPED (operator ruling): a running app on a half-written database lets the customer make the damage permanent. Every start path refuses it - customer button, appstop Recover, boot sweep - via the shared driveStartGate, checked ABOVE its driveless early return because these apps have no drive. The marker is ended so nothing auto-restarts it. The row goes red. Cleared with --clear-restore-hold, an operator CLI route. --single-transaction is a belt on Postgres only; MariaDB DDL is not transactional and that is why the rollback is the fix. R-381: the engine's stderr stops reaching the customer (615 bytes on MariaDB, its middle rows out of their own database) and starts reaching the operator log, which never had it. R-382: the summary log prints the volume count it already held. Undo copies resolve to their own app, are marked IsUndo, and are capped at 3 per app, pruned from the capture side. The reported render-as-an-app symptom did NOT reproduce - the live page was read first and had zero occurrences. Tests 1468 -> 1483. Eight red-proofs; ONE PASSED and is reported: the R-381 behavioural test injected below ImportDump. A guard at that layer now convicts.
This commit is contained in:
@@ -1157,6 +1157,11 @@ type AppBackupRow struct {
|
||||
// e.g., "DB + Konfiguráció + Adatok", "DB + Konfiguráció", "Konfiguráció"
|
||||
BackupContents string
|
||||
|
||||
// RestoreHeld (R-379) — this app is deliberately stopped because a database restore failed AND
|
||||
// the rollback failed. Distinct from any backup status: it is about the app's LIVE data, not its
|
||||
// copies, which is why it drives the row red rather than yellow.
|
||||
RestoreHeld bool
|
||||
|
||||
// Tier 1: Nightly backup (always exists)
|
||||
Tier1LastRun string // RFC3339 time of the newest recovery-unit artifact ("" = no unit yet)
|
||||
Tier1LastStatus string // "ok", "error", ""
|
||||
@@ -1361,6 +1366,16 @@ func (s *Server) buildAppBackupRows(status *backup.FullBackupStatus) []AppBackup
|
||||
row.Status = "yellow"
|
||||
row.StatusText = "Adatbázis mentés sikertelen"
|
||||
}
|
||||
// R-379/R-380: a HELD app must never read as healthy. Last, so it wins over both branches
|
||||
// above — a warning beside a green tick is read as a success, and this is the one state where
|
||||
// the customer's data may not be intact. Measured 2026-08-22: after a failed MariaDB replay
|
||||
// the app reported `health=healthy, running=true, restarts=0` while its schema-version table
|
||||
// held zero rows.
|
||||
if held, why := s.backupMgr.RestoreHoldFor(app.StackName); held {
|
||||
row.Status = "red"
|
||||
row.StatusText = why
|
||||
row.RestoreHeld = true
|
||||
}
|
||||
|
||||
// Tier 2 (off-drive copy) status, from the config the Tier 2 runner persists.
|
||||
if cd := s.settings.GetCrossDriveConfig(app.StackName); cd != nil {
|
||||
|
||||
Reference in New Issue
Block a user