R-379/R-380: put the customer's undo copy back when a database restore fails
gates / gates (push) Successful in 11s

R-379 and R-380 were one failure. Both ended with a half-restored database; the
only difference was whether it looked broken. Postgres emptied and crash-looped;
MariaDB applied part of the dump and reported health=healthy with a zero-row
schema-version table. Measured live on demo-hp 2026-08-22.

The undo copy was already taken and already good - proven by hand that day on
both engines. Nothing in the product could apply it. Now it does, with the same
ImportDump call, before any restart and inside the DB-only window.

The WHOLE undo set, matched on this run's stamp. writeSafetyDump returned one
path for an app with two databases; a rollback on that would restore one and
leave the other half-written.

When the rollback also fails the app is HELD STOPPED (operator ruling): a running
app on a half-written database lets the customer make the damage permanent. Every
start path refuses it - customer button, appstop Recover, boot sweep - via the
shared driveStartGate, checked ABOVE its driveless early return because these
apps have no drive. The marker is ended so nothing auto-restarts it. The row goes
red. Cleared with --clear-restore-hold, an operator CLI route.

--single-transaction is a belt on Postgres only; MariaDB DDL is not transactional
and that is why the rollback is the fix.

R-381: the engine's stderr stops reaching the customer (615 bytes on MariaDB, its
middle rows out of their own database) and starts reaching the operator log,
which never had it.
R-382: the summary log prints the volume count it already held.
Undo copies resolve to their own app, are marked IsUndo, and are capped at 3 per
app, pruned from the capture side. The reported render-as-an-app symptom did NOT
reproduce - the live page was read first and had zero occurrences.

Tests 1468 -> 1483. Eight red-proofs; ONE PASSED and is reported: the R-381
behavioural test injected below ImportDump. A guard at that layer now convicts.
This commit is contained in:
2026-08-22 18:03:18 +02:00
parent 0f3cf0dbb2
commit 2c724c9283
15 changed files with 1288 additions and 37 deletions
+19
View File
@@ -54,6 +54,18 @@ type Manager struct {
// precedent.
unitNotify func(stackName string, err error, usage *UnitSpace)
// restoreHoldNotify (R-379/R-380), if set, is called ONCE when an app is HELD after a database
// replay failed AND the rollback to the customer's own pre-restore copy also failed. Wired in
// cmd/controller/main.go. Same seam shape as unitNotify above and for the same reason: the
// manager must not import the notifier.
//
// OPERATOR-TIER, and this is the whole reason it is a seam rather than a direct call. A held app
// must NOT reach `NotifyBackupFailed` — that type is customer-enabled by default
// (`settings.DefaultEnabledEvents`) and carries the Hungarian "A biztonsági mentés sikertelen!",
// which would alarm a customer about an app we are DELIBERATELY holding. That is R-171's defect
// one path over, and it is the same distinction `ErrStartRefused` exists to keep.
restoreHoldNotify func(stack string, replayErr, rollbackErr error)
// unitSpaceFn (R-165 / B2), if set, replaces the real statfs behind the capture floor so a test
// can state a filesystem's occupancy as an input. Nil in production → `unitTargetSpace`.
unitSpaceFn func(stackName string) *UnitSpace
@@ -137,6 +149,13 @@ type Manager struct {
discoverDBs func(ctx context.Context) ([]DiscoveredDB, error)
importDBDump func(ctx context.Context, db DiscoveredDB, dumpPath string) error
// rollbackImport (R-379) — the ROLLBACK's ImportDump seam. Deliberately SEPARATE from
// importDBDump above even though both default to ImportDump: the whole point of the rollback is
// what happens when the replay fails, so a test must be able to make the replay fail and the
// rollback succeed (and the reverse). One shared seam cannot express that, and a test that
// cannot express the case cannot pin it.
rollbackImport func(ctx context.Context, db DiscoveredDB, dumpPath string) error
// F3 volume-dump seam — overridable in tests so runVolumeDumps' gating (protected / volume-less /
// disconnected) can be unit-tested without Docker. Nil → the real DumpAppVolumesSafe.
dumpVolumesSafe func(stackName string) error
+247 -25
View File
@@ -6,8 +6,11 @@ import (
"os"
"os/exec"
"path/filepath"
"sort"
"strings"
"time"
"gitea.dooplex.hu/admin/felhom-controller/internal/settings"
)
// Offsite reconstitution (R-43, v0.148.0) — the leg that was missing.
@@ -71,11 +74,16 @@ type OffsiteReconstituteResult struct {
// restore that had silently dropped a 1.4 MB volume archive — a true sentence leaving a false
// impression, which is the shape this surface keeps having removed from it.
VolumesReplayed int
SafetyDump string // path of the pre-restore dump (the undo), "" when the app has no DB
DumpsAt time.Time // when the snapshot's DB half was taken (zero = unknown/legacy unit)
OffsiteRunID string // "" for a pre-v0.148 snapshot — an unverified pair
Skewed bool // the snapshot carries no coherence stamp: files and DB may differ in age
LooksEmpty bool // R-44 sniff on the dump about to be replayed
SafetyDump string // path of the pre-restore dump (the undo), "" when the app has no DB
// RolledBack (R-379) is true when the database replay FAILED and this run put the customer's own
// pre-restore copy back. It is on the result rather than inferred from the error, because "the
// restore failed" and "your data is as it was" are two different facts and the surface has to be
// able to say both.
RolledBack bool
DumpsAt time.Time // when the snapshot's DB half was taken (zero = unknown/legacy unit)
OffsiteRunID string // "" for a pre-v0.148 snapshot — an unverified pair
Skewed bool // the snapshot carries no coherence stamp: files and DB may differ in age
LooksEmpty bool // R-44 sniff on the dump about to be replayed
// Placement (R-351) is what the backup recorded about where this app's data lived, compared
// against where this restore actually wrote. Carried on the RESULT and not only on the refusal,
// so a restore that proceeded into a different destination says so in its own outcome rather
@@ -112,13 +120,43 @@ func rsyncRestoreOverwrite(src, dst string) (int, error) {
return countRestoredFiles(string(out)), nil
}
// safetyDumpSet is what ONE reconstitution's undo consists of: the stamp that identifies this run's
// files, and one written path per database the app has.
//
// R-379: it exists because `writeSafetyDump` used to return only the FIRST path, and the rollback
// added in v0.220.0 must re-apply EVERY database's undo or it restores one and leaves the other
// half-written — the defect it exists to close, one database over. The stamp is the IDENTITY: three
// runs against `docmost` on 2026-08-22 left three `pre-restore-*` files in the same directory, so
// matching on the prefix would replay an arbitrary older state. Match on this stamp, never on the
// prefix, never on age or size.
type safetyDumpSet struct {
Stamp string // 20060102T150405Z — this run's, and only this run's
Files []safetyDumpFile // one per database, in discovery order
}
// safetyDumpFile pairs an undo file with the database it came from, so the rollback can hand each
// dump back to the container it belongs to instead of guessing from the filename.
type safetyDumpFile struct {
DB DiscoveredDB
Path string
}
// First returns the first written path, or "" — the value the pre-v0.220.0 signature returned, kept
// because the customer-facing message names one file and changing that is not this task.
func (s safetyDumpSet) First() string {
if len(s.Files) == 0 {
return ""
}
return s.Files[0].Path
}
// writeSafetyDump dumps every live database of stack into the app's unit db-dumps dir under the
// `pre-restore-` prefix, and returns the first dump's path. Returns ("", nil) when the app has no
// `pre-restore-` prefix, and returns the SET it wrote. Returns (zero, nil) when the app has no
// database at all — a no-DB app has nothing to undo and must flow exactly as it did before
// v0.148.0 (no dump, no replay, no behaviour change).
//
// A discovered database that CANNOT be dumped is a hard error: it means the undo would not exist.
func (m *Manager) writeSafetyDump(ctx context.Context, stackName, nsRoot string) (string, error) {
func (m *Manager) writeSafetyDump(ctx context.Context, stackName, nsRoot string) (safetyDumpSet, error) {
discover := m.discoverDBs
if discover == nil {
discover = func(ctx context.Context) ([]DiscoveredDB, error) {
@@ -127,7 +165,7 @@ func (m *Manager) writeSafetyDump(ctx context.Context, stackName, nsRoot string)
}
dbs, err := discover(ctx)
if err != nil {
return "", fmt.Errorf("a biztonsági mentés előtt nem sikerült felderíteni az adatbázisokat: %w", err)
return safetyDumpSet{}, fmt.Errorf("a biztonsági mentés előtt nem sikerült felderíteni az adatbázisokat: %w", err)
}
var mine []DiscoveredDB
for _, db := range dbs {
@@ -136,34 +174,185 @@ func (m *Manager) writeSafetyDump(ctx context.Context, stackName, nsRoot string)
}
}
if len(mine) == 0 {
return "", nil // no DB → nothing to undo → scenario E flows unchanged
return safetyDumpSet{}, nil // no DB → nothing to undo → scenario E flows unchanged
}
dumpDir := AppDBDumpPath(nsRoot, stackName)
if err := os.MkdirAll(dumpDir, 0755); err != nil {
return "", fmt.Errorf("a biztonsági mentés könyvtára nem hozható létre: %w", err)
return safetyDumpSet{}, fmt.Errorf("a biztonsági mentés könyvtára nem hozható létre: %w", err)
}
stamp := time.Now().UTC().Format("20060102T150405Z")
first := ""
set := safetyDumpSet{Stamp: time.Now().UTC().Format("20060102T150405Z")}
for _, db := range mine {
res := m.dumpForSafety(ctx, db, dumpDir)
if res.Error != nil {
return "", fmt.Errorf("a jelenlegi adatbázis biztonsági mentése sikertelen (%s): %w — a visszaállítás nem indult el", db.ContainerName, res.Error)
return safetyDumpSet{}, fmt.Errorf("a jelenlegi adatbázis biztonsági mentése sikertelen (%s): %w — a visszaállítás nem indult el", db.ContainerName, res.Error)
}
// DumpOne writes `<stack>-<dbtype>.sql`; rename it under the safety prefix so it can never be
// picked up as a replay SOURCE and can never overwrite the app's real dump.
safe := filepath.Join(dumpDir, fmt.Sprintf("%s%s-%s-%s.sql", preRestoreDumpPrefix, stamp, stackName, db.DBType))
safe := filepath.Join(dumpDir, fmt.Sprintf("%s%s-%s-%s.sql", preRestoreDumpPrefix, set.Stamp, stackName, db.DBType))
if res.FilePath != safe {
if err := os.Rename(res.FilePath, safe); err != nil {
return "", fmt.Errorf("a biztonsági mentés véglegesítése sikertelen: %w", err)
return safetyDumpSet{}, fmt.Errorf("a biztonsági mentés véglegesítése sikertelen: %w", err)
}
}
if first == "" {
first = safe
}
// EVERY file, not just the first — R-379, and the reason is on safetyDumpSet.
set.Files = append(set.Files, safetyDumpFile{DB: db, Path: safe})
m.logger.Printf("[INFO] [offbox] %s: pre-restore safety dump written → %s (%s)", stackName, filepath.Base(safe), humanizeBytes(res.Size))
}
return first, nil
return set, nil
}
// maxUndoCopiesPerApp is how many `pre-restore-` copies an app keeps.
//
// THREE, and the reasoning rather than a number pulled from the air. One is not enough: the case
// that needs an undo is a restore that went wrong, and the second-guess attempt is exactly when the
// customer reaches for the state before the FIRST attempt. Many is not free: they live inside the
// recovery unit, so every one is also mirrored to Tier 2 AND pushed off-site permanently — four
// accumulated on `docmost` in a single afternoon on 2026-08-22 (135 KB + 135 KB + 141 KB + 138 KB),
// each of them forever. Three keeps two prior attempts and bounds the off-site growth.
const maxUndoCopiesPerApp = 3
// pruneUndoCopies keeps the newest maxUndoCopiesPerApp undo copies for an app and removes the rest.
//
// DELIBERATELY NOT CALLED FROM THE RESTORE PATH. A delete on the failure path is how an undo goes
// missing at exactly the moment it is needed; this runs from the capture side, where nothing is
// depending on the files right now. It is called AFTER a successful capture, never before one.
//
// Ordering is by the stamp IN THE FILENAME, not by mtime and never by size: mtime moves when a file
// is copied or a filesystem is restored, and the stamp is the identity writeSafetyDump assigned.
// The newest is never a deletion candidate even if the list is somehow malformed.
func (m *Manager) pruneUndoCopies(dumpDir, stack string) {
entries, err := os.ReadDir(dumpDir)
if err != nil {
return
}
type undo struct{ name, stamp string }
var undos []undo
for _, e := range entries {
if e.IsDir() || !strings.HasSuffix(e.Name(), ".sql") {
continue
}
base := strings.TrimSuffix(e.Name(), ".sql")
if !strings.HasPrefix(base, preRestoreDumpPrefix) {
continue
}
after := strings.TrimPrefix(base, preRestoreDumpPrefix)
i := strings.Index(after, "-")
if i <= 0 {
continue // not the shape writeSafetyDump writes — leave it alone rather than guess
}
undos = append(undos, undo{name: e.Name(), stamp: after[:i]})
}
if len(undos) <= maxUndoCopiesPerApp {
return
}
sort.Slice(undos, func(a, b int) bool { return undos[a].stamp > undos[b].stamp }) // newest first
for _, u := range undos[maxUndoCopiesPerApp:] {
p := filepath.Join(dumpDir, u.name)
if err := os.Remove(p); err != nil {
m.logger.Printf("[WARN] [backup] %s: could not prune old undo copy %s: %v", stack, u.name, err)
continue
}
m.logger.Printf("[INFO] [backup] %s: pruned old undo copy %s (keeping the newest %d)", stack, u.name, maxUndoCopiesPerApp)
}
}
// RestoreHoldFor reports whether an app is being held stopped after a failed restore + failed
// rollback, and returns the customer-facing reason. Every start path consults this — the customer's
// button, the app-stop Recover() starter, and the boot reconciler — because a hold that only one
// path honours is not a hold.
//
// Nil settings ⇒ NOT held. That direction is deliberate and is the opposite of the usual fail-closed
// rule: with no settings there is no hold recorded, so refusing every start would strand every app
// on a misconfigured box. The write side logs loudly when it cannot persist (see
// holdAppAfterFailedRollback), which is where that case is caught.
func (m *Manager) RestoreHoldFor(stack string) (bool, string) {
if m == nil || m.settings == nil {
return false, ""
}
h, ok := m.settings.GetRestoreHold(stack)
if !ok {
return false, ""
}
when := h.At
if t, err := time.Parse(time.RFC3339, h.At); err == nil {
when = t.Format("2006-01-02 15:04")
}
return true, fmt.Sprintf("a(z) %s adatainak visszaállítása %s-kor megszakadt, és a korábbi állapotot sem sikerült visszatölteni. "+
"Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek tovább. Vedd fel velünk a kapcsolatot", stack, when)
}
// holdAppAfterFailedRollback records the R-379/R-380 hold and makes sure nothing restarts the app
// behind our back.
//
// OPERATOR RULING, 2026-08-22: when the replay fails AND the rollback fails, the app is HELD
// STOPPED rather than started. A running app on a half-written database lets the customer type into
// it, and that turns a recoverable state into a permanent one. The alternative — start it and mark
// it — was considered and declined.
//
// It ENDS the app-stop marker deliberately. The marker means "owed a restart"; a held app is not
// owed one, and leaving the marker active would have Recover() start the broken app at the next
// controller boot. The hold is the thing that persists, not the marker.
func (m *Manager) holdAppAfterFailedRollback(stack string, replayErr, rollbackErr error) {
if m.settings == nil {
m.logger.Printf("[ERROR] [offbox] %s: cannot persist the restore hold — no settings wired; the app is stopped but NOTHING will refuse a restart", stack)
return
}
h := settings.RestoreHold{
Stack: stack,
At: time.Now().UTC().Format(time.RFC3339),
}
if replayErr != nil {
h.ReplayError = replayErr.Error()
}
if rollbackErr != nil {
h.RollbackErr = rollbackErr.Error()
}
if err := m.settings.SetRestoreHold(h); err != nil {
m.logger.Printf("[ERROR] [offbox] %s: persisting the restore hold FAILED: %v — the app is stopped and unguarded", stack, err)
}
// The app is not owed a restart; it is deliberately held. See the doc comment.
if m.appStop != nil {
m.appStop.End()
}
if m.restoreHoldNotify != nil {
m.restoreHoldNotify(stack, replayErr, rollbackErr)
}
}
// SetRestoreHoldNotify wires the operator notification for a held app (cmd/controller/main.go).
func (m *Manager) SetRestoreHoldNotify(fn func(stack string, replayErr, rollbackErr error)) {
m.restoreHoldNotify = fn
}
// rollbackSafetyDump re-applies THIS RUN's undo set, database by database, and is the whole of
// R-379's fix: it is the same ImportDump call a person made by hand on 2026-08-22 to recover
// `docmost` and `bookstack` after a failed replay, moved into the product.
//
// It runs with the DB service still up (the replay's own window) and BEFORE any restart, so the
// app never observes the half-written state. An error here means the app cannot be trusted to run —
// see the hold in ReconstituteFromOffsite.
func (m *Manager) rollbackSafetyDump(ctx context.Context, stack string, set safetyDumpSet) error {
if len(set.Files) == 0 {
return nil
}
imp := m.rollbackImport
if imp == nil {
imp = func(ctx context.Context, db DiscoveredDB, path string) error {
return ImportDump(ctx, db, path, m.logger, m.isDebug())
}
}
for _, f := range set.Files {
if _, sErr := os.Stat(f.Path); sErr != nil {
return fmt.Errorf("a visszavonáshoz szükséges mentés nem található (%s): %w", filepath.Base(f.Path), sErr)
}
m.logger.Printf("[INFO] [offbox] %s: rolling back to the pre-restore state from %s", stack, filepath.Base(f.Path))
if err := imp(ctx, f.DB, f.Path); err != nil {
return fmt.Errorf("a korábbi állapot visszaállítása sikertelen (%s): %w", f.DB.ContainerName, err)
}
}
m.logger.Printf("[INFO] [offbox] %s: rollback complete — %d database(s) returned to the pre-restore state", stack, len(set.Files))
return nil
}
// dumpForSafety is the DumpOne seam for the safety dump (tests inject; nil → the real DumpOne).
@@ -325,10 +514,11 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string, ack
// --- THE UNDO, BEFORE THE ACT ---------------------------------------------------------------
// Taken while the stack is still UP (a stopped database cannot be dumped) and before a single
// byte is overwritten, so a failure here aborts with the live app completely untouched.
safety, err := m.writeSafetyDump(ctx, stack, liveNs)
safetySet, err := m.writeSafetyDump(ctx, stack, liveNs)
if err != nil {
return res, err
}
safety := safetySet.First()
res.SafetyDump = safety
hasDB := safety != ""
if hasDB {
@@ -436,10 +626,39 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string, ack
n, iErr := m.reimportDBDumpsFrom(ctx, stack, scratchDumpDir)
res.DBsReplayed = n
if iErr != nil {
if sErr := restartStack(); sErr != nil {
m.logger.Printf("[WARN] [offbox] %s: full start after failed replay also failed: %v", stack, sErr)
// --- R-379/R-380: PUT THE CUSTOMER'S OWN COPY BACK ---------------------------------
// Until v0.220.0 this branch restarted the app onto a HALF-WRITTEN database and named
// the undo file in the message. Measured 2026-08-22 on demo-hp: Postgres was left
// emptied and crash-looping; MariaDB was left partly applied while the app reported
// `health=healthy`. Both are the same failure — a half state — and the only difference
// was whether it looked broken. MariaDB's structural statements are not transactional,
// so no engine flag can prevent the half state; putting the undo back is what removes
// it. This is the same ImportDump call a person ran by hand that day to recover both
// apps, moved into the product.
//
// The rollback runs BEFORE any restart and with the DB service still up, so the app
// never observes the half state. The ORIGINAL replay error is never swallowed: it is
// logged here in full and named in the customer's sentence.
m.logger.Printf("[ERROR] [offbox] %s: database replay failed, rolling back to the pre-restore state: %v", stack, iErr)
if rbErr := m.rollbackSafetyDump(ctx, stack, safetySet); rbErr != nil {
// BOTH failed. Do NOT start the app: a running app on a half-written database lets
// the customer type into it and makes the damage permanent. Hold it instead —
// operator ruling, 2026-08-22.
m.logger.Printf("[ERROR] [offbox] %s: ROLLBACK ALSO FAILED (%v) — holding the app stopped; replay error was: %v", stack, rbErr, iErr)
m.holdAppAfterFailedRollback(stack, iErr, rbErr)
return res, fmt.Errorf("a(z) %s adatbázisának visszaállítása sikertelen, és a korábbi állapot visszatöltése sem sikerült. "+
"Az alkalmazást biztonsági okból LEÁLLÍTVA hagytuk, hogy az adatai ne sérüljenek tovább. "+
"Vedd fel velünk a kapcsolatot — a korábbi állapot mentése megvan: %s", stack, filepath.Base(safety))
}
return res, fmt.Errorf("az adatbázis visszaállítása sikertelen: %w — a korábbi állapot mentése megvan: %s", iErr, filepath.Base(safety))
res.RolledBack = true
if sErr := restartStack(); sErr != nil {
m.logger.Printf("[WARN] [offbox] %s: start after a successful rollback failed: %v", stack, sErr)
}
// Says BOTH things. A message that reported only the failure would leave the customer
// believing their data was gone when it is back — the omission of a GAIN is as
// misleading as the omission of a loss.
return res, fmt.Errorf("a(z) %s adatbázisának visszaállítása sikertelen — az adataid visszakerültek a visszaállítás előtti állapotba, "+
"az alkalmazás fut tovább. Ha újra megpróbálnád, előbb vedd fel velünk a kapcsolatot", stack)
}
}
if err := restartStack(); err != nil {
@@ -449,8 +668,11 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string, ack
m.logger.Printf("[WARN] [offbox] %s reconstituted but health check failed: %v", stack, err)
}
m.logger.Printf("[INFO] [offbox] reconstituted %s from snapshot %s: %d file(s) placed, %d DB dump(s) replayed, safety dump=%s, skewed=%v",
stack, id, res.FilesPlaced, res.DBsReplayed, filepath.Base(safety), res.Skewed)
// R-382: VolumesReplayed was set above and never printed, so the operator log said
// "0 file(s) placed, 1 DB dump(s) replayed" on a run that returned a 52 MB Postgres data
// directory — less informative than the customer's own flash, which already named the volumes.
m.logger.Printf("[INFO] [offbox] reconstituted %s from snapshot %s: %d file(s) placed, %d volume(s) replayed, %d DB dump(s) replayed, safety dump=%s, skewed=%v",
stack, id, res.FilesPlaced, res.VolumesReplayed, res.DBsReplayed, filepath.Base(safety), res.Skewed)
return res, nil
}
@@ -35,6 +35,12 @@ func (m *Manager) SetOffboxFullPlaceCopier(fn func(src, dst string) (int, error)
m.offboxFullPlaceCopier = fn
}
// SetRollbackImportFn overrides the ROLLBACK's ImportDump (tests; no Docker needed). Separate from
// the replay's own import seam on purpose — see the field comment on Manager.rollbackImport.
func (m *Manager) SetRollbackImportFn(fn func(ctx context.Context, db DiscoveredDB, dumpPath string) error) {
m.rollbackImport = fn
}
// SetSafetyDumpFn overrides the pre-restore safety dump (tests; no Docker needed).
func (m *Manager) SetSafetyDumpFn(fn func(ctx context.Context, db DiscoveredDB, dumpDir string) DumpResult) {
m.safetyDumpFn = fn
@@ -45,7 +45,11 @@ func TestR355_SafetyDumpIsTakenForTheCorrectlyAttributedApp(t *testing.T) {
return DumpResult{DB: db, FilePath: p, Size: 13}
}
safety, err := m.writeSafetyDump(context.Background(), "paperless-ngx", nsRoot)
// v0.220.0 (R-379): writeSafetyDump returns the SET it wrote, so a rollback can re-apply EVERY
// database's undo. `.First()` is the value this signature returned before; these assertions are
// unchanged in meaning.
set, err := m.writeSafetyDump(context.Background(), "paperless-ngx", nsRoot)
safety := set.First()
if err != nil {
t.Fatalf("writeSafetyDump: %v", err)
}
@@ -83,7 +87,11 @@ func TestR355_MisattributedAppGetsNoUndoCopy(t *testing.T) {
return DumpResult{DB: db}
}
safety, err := m.writeSafetyDump(context.Background(), "paperless-ngx", nsRoot)
// v0.220.0 (R-379): writeSafetyDump returns the SET it wrote, so a rollback can re-apply EVERY
// database's undo. `.First()` is the value this signature returned before; these assertions are
// unchanged in meaning.
set, err := m.writeSafetyDump(context.Background(), "paperless-ngx", nsRoot)
safety := set.First()
if err != nil {
t.Fatalf("writeSafetyDump: %v", err)
}
@@ -111,7 +119,11 @@ func TestR355_RestoreRefusesWhenTheUndoCannotBeTaken(t *testing.T) {
return DumpResult{DB: db, Error: os.ErrPermission}
}
safety, err := m.writeSafetyDump(context.Background(), "paperless-ngx", nsRoot)
// v0.220.0 (R-379): writeSafetyDump returns the SET it wrote, so a rollback can re-apply EVERY
// database's undo. `.First()` is the value this signature returned before; these assertions are
// unchanged in meaning.
set, err := m.writeSafetyDump(context.Background(), "paperless-ngx", nsRoot)
safety := set.First()
if err == nil {
t.Fatal("a database that cannot be dumped must be a hard error — the undo would not exist")
}
@@ -0,0 +1,330 @@
package backup
import (
"context"
"errors"
"os"
"path/filepath"
"strings"
"testing"
)
// R-379 / R-380. When an off-site database replay failed, the app was restarted onto a HALF-WRITTEN
// database and the customer was shown the undo copy's filename — a file nothing in the product could
// apply. Measured live on demo-hp 2026-08-22: Postgres left emptied and crash-looping, MariaDB left
// partly applied while `docker inspect` reported health=healthy. Both are the same failure — a half
// state — and the only difference was whether it looked broken.
//
// The fix puts the customer's own pre-restore copy back automatically. These tests pin that, the
// double-failure hold, and the two absence claims that go with them.
// rollbackFixture is reconFixture with the replay failing and the rollback injectable.
func rollbackFixture(t *testing.T, rollbackErr error) (*Manager, *recordingProvider, *int) {
t.Helper()
m, prov, _ := reconFixture(t, "run1", "2026-07-19T06:00:00Z", pgDump(1))
m.importDBDump = func(context.Context, DiscoveredDB, string) error {
return errors.New("replay blew up")
}
calls := 0
m.SetRollbackImportFn(func(_ context.Context, _ DiscoveredDB, path string) error {
calls++
// The rollback must be handed THIS RUN's undo file, not the scratch's dump.
if !strings.Contains(filepath.Base(path), preRestoreDumpPrefix) {
t.Errorf("rollback was handed %q — that is not an undo copy", filepath.Base(path))
}
return rollbackErr
})
return m, prov, &calls
}
// SCENARIO A — replay fails, rollback succeeds. The app comes back and the outcome records it.
//
// WRONG OUTCOMES PINNED: the app left down; or a message that reports the failure and omits that the
// data is back — the omission of a GAIN is as misleading as the omission of a loss, and here it is
// the difference between a customer who panics and one who does not.
func TestR379_ScenarioA_RollbackSucceeds_AppComesBack(t *testing.T) {
m, prov, calls := rollbackFixture(t, nil)
res, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err == nil {
t.Fatal("a failed replay must still be surfaced as a failure")
}
if *calls != 1 {
t.Fatalf("the undo copy must be re-applied exactly once, got %d", *calls)
}
if !res.RolledBack {
t.Error("the result must record that the rollback happened")
}
if !prov.fullStarted {
t.Fatal("after a successful rollback the app must be started — it is in a known good state")
}
// The customer sentence says BOTH things.
low := err.Error()
if !strings.Contains(low, "sikertelen") {
t.Errorf("the message must say the restore failed; got: %v", err)
}
if !strings.Contains(low, "visszakerültek") {
t.Errorf("the message must say the data is back as it was; got: %v", err)
}
// And no hold was written — the app is fine.
if held, _ := m.RestoreHoldFor("immich"); held {
t.Error("a recovered app must NOT be held")
}
}
// SCENARIO C — replay fails AND rollback fails. The app is held, not started.
//
// OPERATOR RULING 2026-08-22: a running app on a half-written database lets the customer type into
// it and makes the damage permanent.
//
// WRONG OUTCOMES PINNED: the app started anyway; or the app-stop marker left active, which would
// have Recover() start the broken app at the next controller boot, quietly, hours later.
func TestR379_ScenarioC_BothFail_AppIsHeldNotStarted(t *testing.T) {
m, prov, calls := rollbackFixture(t, errors.New("rollback blew up too"))
notified := 0
m.SetRestoreHoldNotify(func(string, error, error) { notified++ })
_, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err == nil {
t.Fatal("a double failure must be surfaced")
}
if *calls != 1 {
t.Fatalf("the rollback must have been attempted, got %d calls", *calls)
}
// THE OBSERVABLE THAT MATTERS.
if prov.fullStarted {
t.Fatal("the app was STARTED onto a half-written database — the exact outcome the hold exists to prevent")
}
held, why := m.RestoreHoldFor("immich")
if !held {
t.Fatal("the hold must be persisted, or nothing will refuse a restart later")
}
if why == "" || !strings.Contains(why, "kapcsolat") {
t.Errorf("the hold's reason must name a route the customer can take; got %q", why)
}
if notified != 1 {
t.Errorf("the operator must be told exactly once, got %d", notified)
}
if !strings.Contains(err.Error(), "LEÁLLÍTVA") {
t.Errorf("the customer must be told the app was deliberately stopped; got: %v", err)
}
}
// SCENARIO E — the replay SUCCEEDS. Nothing new may happen.
//
// WRONG OUTCOME PINNED: a rollback firing on a successful restore would overwrite the restored data
// with the pre-restore state — a silent, total loss of the thing the customer asked for.
func TestR379_ScenarioE_SuccessfulReplay_NoRollbackNoHold(t *testing.T) {
m, prov, _ := reconFixture(t, "run1", "2026-07-19T06:00:00Z", pgDump(1))
rolled := 0
m.SetRollbackImportFn(func(context.Context, DiscoveredDB, string) error { rolled++; return nil })
held := 0
m.SetRestoreHoldNotify(func(string, error, error) { held++ })
res, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err != nil {
t.Fatalf("a clean restore must succeed: %v", err)
}
if rolled != 0 {
t.Fatalf("a rollback fired on a SUCCESSFUL restore — the restored data would be overwritten; calls=%d", rolled)
}
if res.RolledBack {
t.Error("a successful restore must not report a rollback")
}
if held != 0 {
t.Errorf("no hold may be written on success, got %d", held)
}
if isHeld, _ := m.RestoreHoldFor("immich"); isHeld {
t.Error("a successful restore must leave no hold")
}
if !prov.fullStarted {
t.Error("a successful restore still starts the app")
}
}
// SCENARIO E, POSITIVE CONTROL. "No rollback fired" is an absence claim, so prove the counter can
// count: the same seam, driven through the failure path, must register.
func TestR379_ScenarioE_PositiveControl_TheCounterCanCount(t *testing.T) {
m, _, calls := rollbackFixture(t, nil)
if _, err := m.ReconstituteFromOffsite(context.Background(), "immich", false); err == nil {
t.Fatal("fixture: the replay must fail here")
}
if *calls == 0 {
t.Fatal("the rollback counter never increments — Scenario E's zero would have proven nothing")
}
}
// SCENARIO F — an app with no database. Untouched in every respect.
func TestR379_ScenarioF_NoDatabase_Unchanged(t *testing.T) {
m, prov, _ := reconFixture(t, "run1", "2026-07-19T06:00:00Z", "")
m.discoverDBs = func(context.Context) ([]DiscoveredDB, error) { return nil, nil }
rolled := 0
m.SetRollbackImportFn(func(context.Context, DiscoveredDB, string) error { rolled++; return nil })
res, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err != nil {
t.Fatalf("a no-DB app must restore: %v", err)
}
if rolled != 0 || res.RolledBack {
t.Error("a no-DB app has nothing to undo and nothing to roll back")
}
if res.SafetyDump != "" {
t.Errorf("a no-DB app takes no undo copy, got %q", res.SafetyDump)
}
if held, _ := m.RestoreHoldFor("immich"); held {
t.Error("a no-DB app must never be held")
}
if !prov.fullStarted {
t.Error("a no-DB app still starts")
}
}
// SCENARIO G — two databases, one replay fails. The WHOLE undo set is re-applied.
//
// WRONG OUTCOME PINNED: only the first. writeSafetyDump used to return one path for an app with two
// databases, so a rollback built on that value would restore one and leave the other half-written —
// this defect, one database over.
func TestR379_ScenarioG_TwoDatabases_WholeSetRolledBack(t *testing.T) {
m, _, _ := reconFixture(t, "run1", "2026-07-19T06:00:00Z", pgDump(1))
two := []DiscoveredDB{
{StackName: "immich", DBType: DBTypePostgres, ContainerName: "immich-postgres", ContainerID: "a"},
{StackName: "immich", DBType: DBTypeMariaDB, ContainerName: "immich-maria", ContainerID: "b"},
}
m.discoverDBs = func(context.Context) ([]DiscoveredDB, error) { return two, nil }
m.safetyDumpFn = func(_ context.Context, db DiscoveredDB, dumpDir string) DumpResult {
p := filepath.Join(dumpDir, "immich-"+string(db.DBType)+".sql")
if err := os.WriteFile(p, []byte(pgDump(1)), 0o644); err != nil {
t.Fatal(err)
}
return DumpResult{DB: db, FilePath: p, Size: 42}
}
m.importDBDump = func(context.Context, DiscoveredDB, string) error { return errors.New("replay failed") }
var rolled []string
m.SetRollbackImportFn(func(_ context.Context, db DiscoveredDB, path string) error {
rolled = append(rolled, db.ContainerName+":"+filepath.Base(path))
return nil
})
if _, err := m.ReconstituteFromOffsite(context.Background(), "immich", false); err == nil {
t.Fatal("the replay failure must surface")
}
if len(rolled) != 2 {
t.Fatalf("BOTH databases must be rolled back, got %d: %v", len(rolled), rolled)
}
// Each database got its OWN undo file, matched by identity rather than by prefix.
if !strings.Contains(rolled[0], "immich-postgres:") || !strings.Contains(rolled[1], "immich-maria:") {
t.Errorf("each database must get its own undo copy, got %v", rolled)
}
for _, r := range rolled {
if !strings.Contains(r, preRestoreDumpPrefix) {
t.Errorf("rollback used a non-undo file: %s", r)
}
}
}
// The undo set is matched on THIS RUN's stamp. Three undo copies from three different runs
// accumulated on docmost in one afternoon; a prefix match would replay an arbitrary older state.
func TestR379_UndoSetIsThisRunOnly(t *testing.T) {
nsRoot := t.TempDir()
m := newSafetyTestManager()
m.discoverDBs = func(context.Context) ([]DiscoveredDB, error) {
return []DiscoveredDB{{StackName: "app", DBType: DBTypePostgres, ContainerName: "app-postgres", ContainerID: "c"}}, nil
}
m.safetyDumpFn = func(_ context.Context, db DiscoveredDB, dumpDir string) DumpResult {
p := filepath.Join(dumpDir, "app-postgres.sql")
if err := os.WriteFile(p, []byte("x"), 0o644); err != nil {
t.Fatal(err)
}
return DumpResult{DB: db, FilePath: p, Size: 1}
}
// An OLDER undo copy from a previous run, already on disk.
dumpDir := AppDBDumpPath(nsRoot, "app")
if err := os.MkdirAll(dumpDir, 0o755); err != nil {
t.Fatal(err)
}
stale := filepath.Join(dumpDir, preRestoreDumpPrefix+"20200101T000000Z-app-postgres.sql")
if err := os.WriteFile(stale, []byte("STALE"), 0o644); err != nil {
t.Fatal(err)
}
set, err := m.writeSafetyDump(context.Background(), "app", nsRoot)
if err != nil {
t.Fatal(err)
}
if len(set.Files) != 1 {
t.Fatalf("one database → one undo file, got %d", len(set.Files))
}
if strings.Contains(set.Files[0].Path, "20200101") {
t.Fatal("the set picked up an OLDER run's undo copy — it must carry only this run's stamp")
}
if set.Stamp == "" || strings.Contains(set.Files[0].Path, set.Stamp) == false {
t.Errorf("the file must carry this run's stamp %q, got %q", set.Stamp, set.Files[0].Path)
}
}
// pruneUndoCopies keeps the newest N and never the oldest — ordered by the STAMP, not by mtime.
func TestR379_PruneKeepsTheNewestByStamp(t *testing.T) {
m := newSafetyTestManager()
dir := t.TempDir()
stamps := []string{"20260101T000000Z", "20260201T000000Z", "20260301T000000Z", "20260401T000000Z", "20260501T000000Z"}
for _, st := range stamps {
if err := os.WriteFile(filepath.Join(dir, preRestoreDumpPrefix+st+"-app-postgres.sql"), []byte("x"), 0o644); err != nil {
t.Fatal(err)
}
}
// The app's OWN dump must never be a candidate.
own := filepath.Join(dir, "app-postgres.sql")
if err := os.WriteFile(own, []byte("own"), 0o644); err != nil {
t.Fatal(err)
}
m.pruneUndoCopies(dir, "app")
if _, err := os.Stat(own); err != nil {
t.Fatal("the app's own dump was pruned — only undo copies are candidates")
}
for _, st := range stamps[len(stamps)-maxUndoCopiesPerApp:] {
if _, err := os.Stat(filepath.Join(dir, preRestoreDumpPrefix+st+"-app-postgres.sql")); err != nil {
t.Errorf("the newest %d must survive; %s is gone", maxUndoCopiesPerApp, st)
}
}
for _, st := range stamps[:len(stamps)-maxUndoCopiesPerApp] {
if _, err := os.Stat(filepath.Join(dir, preRestoreDumpPrefix+st+"-app-postgres.sql")); err == nil {
t.Errorf("the oldest must be pruned; %s survived", st)
}
}
}
// R-381 — the customer sentence must not carry the engine's output.
//
// MEASURED 2026-08-22: 407 bytes (Postgres, with a caret diagram and `exit status 3`) and 615 bytes
// (MariaDB, whose middle was an `INSERT INTO migrations VALUES (...)` listing — ROWS OUT OF THE
// CUSTOMER'S OWN DATABASE, HTML-escaped, on their dashboard).
func TestR381_CustomerMessageCarriesNoEngineOutput(t *testing.T) {
engineNoise := "ERROR: syntax error at end of input\nLINE 1: COPY public.felhom_r356b_discriminator \n ^ — exit status 3"
m, _, _ := reconFixture(t, "run1", "2026-07-19T06:00:00Z", pgDump(1))
m.importDBDump = func(context.Context, DiscoveredDB, string) error {
return errors.New(engineNoise)
}
m.SetRollbackImportFn(func(context.Context, DiscoveredDB, string) error { return nil })
_, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err == nil {
t.Fatal("the failure must surface")
}
msg := err.Error()
for _, leak := range []string{"ERROR: syntax", "LINE 1:", "COPY public.", "exit status"} {
if strings.Contains(msg, leak) {
t.Errorf("the customer message leaks engine output %q; got: %s", leak, msg)
}
}
if len(msg) > 320 {
t.Errorf("the customer message is %d bytes — it is a sentence, not a transcript: %s", len(msg), msg)
}
// POSITIVE CONTROL: the noise really was in the error the code received, so the absence above is
// the message being clean rather than the noise never existing.
if !strings.Contains(engineNoise, "exit status") {
t.Fatal("fixture: the planted noise does not contain the marker this test greps for")
}
}
@@ -264,11 +264,19 @@ func TestRestoreFromUnitIgnoresSafetyDumpsWhenDecidingToReplay(t *testing.T) {
// TestReconstituteReplayFailureStillBringsTheStackUp: the DB-only window is a deliberate half-started
// state, so EVERY exit from it must end in a full start. Otherwise a failed restore leaves the
// customer with a running database and no application — an outage caused by the recovery tool.
//
// UPDATED FOR v0.220.0 (R-379). The requirement is unchanged and still asserted; what changed is
// what happens BETWEEN the failure and the full start. A failed replay now re-applies the customer's
// own pre-restore copy first, and only then starts. The rollback seam is injected as SUCCEEDING here
// because that is this test's subject; the double-failure path — where the app is deliberately NOT
// started — is Scenario C and has its own test in r379_rollback_test.go.
func TestReconstituteReplayFailureStillBringsTheStackUp(t *testing.T) {
m, prov, _ := reconFixture(t, "run1", "2026-07-19T06:00:00Z", pgDump(1))
m.importDBDump = func(context.Context, DiscoveredDB, string) error {
return context.DeadlineExceeded
}
rolledBack := 0
m.SetRollbackImportFn(func(context.Context, DiscoveredDB, string) error { rolledBack++; return nil })
res, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err == nil {
@@ -280,9 +288,16 @@ func TestReconstituteReplayFailureStillBringsTheStackUp(t *testing.T) {
if got := strings.Join(prov.calls, ","); got != "stop,startsvc:immich-postgres,start" {
t.Fatalf("sequence = %q, want the best-effort full start after the failure", got)
}
// The existing message shape stays: the operator needs the undo's filename.
if !strings.Contains(err.Error(), filepath.Base(res.SafetyDump)) {
t.Fatalf("the error must name the safety dump so the operator can undo, got: %v", err)
if rolledBack != 1 {
t.Fatalf("the customer's pre-restore copy must be put back before the start; rollback calls = %d", rolledBack)
}
if !res.RolledBack {
t.Error("the outcome must record that a rollback happened — the surface has to be able to say the data is back")
}
// v0.220.0: the customer sentence now states the OUTCOME (their data is as it was) instead of a
// filename. The filename remains for the operator, in the log and on the result.
if res.SafetyDump == "" || !strings.Contains(filepath.Base(res.SafetyDump), preRestoreDumpPrefix) {
t.Fatalf("the result must still carry the undo copy for the operator, got %q", res.SafetyDump)
}
}
@@ -195,6 +195,11 @@ func (m *Manager) CaptureRecoveryUnit(stackName string) error {
stackName, RecoveryUnitPath(nsRoot, stackName), len(info.ImagePins), len(info.SecretEnvVars),
len(info.DataKeyEnvVars), len(info.PortableSecrets), len(info.PortableSecretEnvVars),
len(withheldSecretNames(info)))
// R-379: bound the undo copies. AFTER a successful capture and never before one — the capture is
// the point at which nothing is depending on those files, whereas the restore path is precisely
// where deleting one would remove the undo at the moment it is needed.
m.pruneUndoCopies(AppDBDumpPath(nsRoot, stackName), stackName)
return nil
}