R-379/R-380: put the customer's undo copy back when a database restore fails
gates / gates (push) Successful in 11s
gates / gates (push) Successful in 11s
R-379 and R-380 were one failure. Both ended with a half-restored database; the only difference was whether it looked broken. Postgres emptied and crash-looped; MariaDB applied part of the dump and reported health=healthy with a zero-row schema-version table. Measured live on demo-hp 2026-08-22. The undo copy was already taken and already good - proven by hand that day on both engines. Nothing in the product could apply it. Now it does, with the same ImportDump call, before any restart and inside the DB-only window. The WHOLE undo set, matched on this run's stamp. writeSafetyDump returned one path for an app with two databases; a rollback on that would restore one and leave the other half-written. When the rollback also fails the app is HELD STOPPED (operator ruling): a running app on a half-written database lets the customer make the damage permanent. Every start path refuses it - customer button, appstop Recover, boot sweep - via the shared driveStartGate, checked ABOVE its driveless early return because these apps have no drive. The marker is ended so nothing auto-restarts it. The row goes red. Cleared with --clear-restore-hold, an operator CLI route. --single-transaction is a belt on Postgres only; MariaDB DDL is not transactional and that is why the rollback is the fix. R-381: the engine's stderr stops reaching the customer (615 bytes on MariaDB, its middle rows out of their own database) and starts reaching the operator log, which never had it. R-382: the summary log prints the volume count it already held. Undo copies resolve to their own app, are marked IsUndo, and are capped at 3 per app, pruned from the capture side. The reported render-as-an-app symptom did NOT reproduce - the live page was read first and had zero occurrences. Tests 1468 -> 1483. Eight red-proofs; ONE PASSED and is reported: the R-381 behavioural test injected below ImportDump. A guard at that layer now convicts.
This commit is contained in:
@@ -54,6 +54,18 @@ type Manager struct {
|
||||
// precedent.
|
||||
unitNotify func(stackName string, err error, usage *UnitSpace)
|
||||
|
||||
// restoreHoldNotify (R-379/R-380), if set, is called ONCE when an app is HELD after a database
|
||||
// replay failed AND the rollback to the customer's own pre-restore copy also failed. Wired in
|
||||
// cmd/controller/main.go. Same seam shape as unitNotify above and for the same reason: the
|
||||
// manager must not import the notifier.
|
||||
//
|
||||
// OPERATOR-TIER, and this is the whole reason it is a seam rather than a direct call. A held app
|
||||
// must NOT reach `NotifyBackupFailed` — that type is customer-enabled by default
|
||||
// (`settings.DefaultEnabledEvents`) and carries the Hungarian "A biztonsági mentés sikertelen!",
|
||||
// which would alarm a customer about an app we are DELIBERATELY holding. That is R-171's defect
|
||||
// one path over, and it is the same distinction `ErrStartRefused` exists to keep.
|
||||
restoreHoldNotify func(stack string, replayErr, rollbackErr error)
|
||||
|
||||
// unitSpaceFn (R-165 / B2), if set, replaces the real statfs behind the capture floor so a test
|
||||
// can state a filesystem's occupancy as an input. Nil in production → `unitTargetSpace`.
|
||||
unitSpaceFn func(stackName string) *UnitSpace
|
||||
@@ -137,6 +149,13 @@ type Manager struct {
|
||||
discoverDBs func(ctx context.Context) ([]DiscoveredDB, error)
|
||||
importDBDump func(ctx context.Context, db DiscoveredDB, dumpPath string) error
|
||||
|
||||
// rollbackImport (R-379) — the ROLLBACK's ImportDump seam. Deliberately SEPARATE from
|
||||
// importDBDump above even though both default to ImportDump: the whole point of the rollback is
|
||||
// what happens when the replay fails, so a test must be able to make the replay fail and the
|
||||
// rollback succeed (and the reverse). One shared seam cannot express that, and a test that
|
||||
// cannot express the case cannot pin it.
|
||||
rollbackImport func(ctx context.Context, db DiscoveredDB, dumpPath string) error
|
||||
|
||||
// F3 volume-dump seam — overridable in tests so runVolumeDumps' gating (protected / volume-less /
|
||||
// disconnected) can be unit-tested without Docker. Nil → the real DumpAppVolumesSafe.
|
||||
dumpVolumesSafe func(stackName string) error
|
||||
|
||||
@@ -6,8 +6,11 @@ import (
|
||||
"os"
|
||||
"os/exec"
|
||||
"path/filepath"
|
||||
"sort"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-controller/internal/settings"
|
||||
)
|
||||
|
||||
// Offsite reconstitution (R-43, v0.148.0) — the leg that was missing.
|
||||
@@ -71,11 +74,16 @@ type OffsiteReconstituteResult struct {
|
||||
// restore that had silently dropped a 1.4 MB volume archive — a true sentence leaving a false
|
||||
// impression, which is the shape this surface keeps having removed from it.
|
||||
VolumesReplayed int
|
||||
SafetyDump string // path of the pre-restore dump (the undo), "" when the app has no DB
|
||||
DumpsAt time.Time // when the snapshot's DB half was taken (zero = unknown/legacy unit)
|
||||
OffsiteRunID string // "" for a pre-v0.148 snapshot — an unverified pair
|
||||
Skewed bool // the snapshot carries no coherence stamp: files and DB may differ in age
|
||||
LooksEmpty bool // R-44 sniff on the dump about to be replayed
|
||||
SafetyDump string // path of the pre-restore dump (the undo), "" when the app has no DB
|
||||
// RolledBack (R-379) is true when the database replay FAILED and this run put the customer's own
|
||||
// pre-restore copy back. It is on the result rather than inferred from the error, because "the
|
||||
// restore failed" and "your data is as it was" are two different facts and the surface has to be
|
||||
// able to say both.
|
||||
RolledBack bool
|
||||
DumpsAt time.Time // when the snapshot's DB half was taken (zero = unknown/legacy unit)
|
||||
OffsiteRunID string // "" for a pre-v0.148 snapshot — an unverified pair
|
||||
Skewed bool // the snapshot carries no coherence stamp: files and DB may differ in age
|
||||
LooksEmpty bool // R-44 sniff on the dump about to be replayed
|
||||
// Placement (R-351) is what the backup recorded about where this app's data lived, compared
|
||||
// against where this restore actually wrote. Carried on the RESULT and not only on the refusal,
|
||||
// so a restore that proceeded into a different destination says so in its own outcome rather
|
||||
@@ -112,13 +120,43 @@ func rsyncRestoreOverwrite(src, dst string) (int, error) {
|
||||
return countRestoredFiles(string(out)), nil
|
||||
}
|
||||
|
||||
// safetyDumpSet is what ONE reconstitution's undo consists of: the stamp that identifies this run's
|
||||
// files, and one written path per database the app has.
|
||||
//
|
||||
// R-379: it exists because `writeSafetyDump` used to return only the FIRST path, and the rollback
|
||||
// added in v0.220.0 must re-apply EVERY database's undo or it restores one and leaves the other
|
||||
// half-written — the defect it exists to close, one database over. The stamp is the IDENTITY: three
|
||||
// runs against `docmost` on 2026-08-22 left three `pre-restore-*` files in the same directory, so
|
||||
// matching on the prefix would replay an arbitrary older state. Match on this stamp, never on the
|
||||
// prefix, never on age or size.
|
||||
type safetyDumpSet struct {
|
||||
Stamp string // 20060102T150405Z — this run's, and only this run's
|
||||
Files []safetyDumpFile // one per database, in discovery order
|
||||
}
|
||||
|
||||
// safetyDumpFile pairs an undo file with the database it came from, so the rollback can hand each
|
||||
// dump back to the container it belongs to instead of guessing from the filename.
|
||||
type safetyDumpFile struct {
|
||||
DB DiscoveredDB
|
||||
Path string
|
||||
}
|
||||
|
||||
// First returns the first written path, or "" — the value the pre-v0.220.0 signature returned, kept
|
||||
// because the customer-facing message names one file and changing that is not this task.
|
||||
func (s safetyDumpSet) First() string {
|
||||
if len(s.Files) == 0 {
|
||||
return ""
|
||||
}
|
||||
return s.Files[0].Path
|
||||
}
|
||||
|
||||
// writeSafetyDump dumps every live database of stack into the app's unit db-dumps dir under the
|
||||
// `pre-restore-` prefix, and returns the first dump's path. Returns ("", nil) when the app has no
|
||||
// `pre-restore-` prefix, and returns the SET it wrote. Returns (zero, nil) when the app has no
|
||||
// database at all — a no-DB app has nothing to undo and must flow exactly as it did before
|
||||
// v0.148.0 (no dump, no replay, no behaviour change).
|
||||
//
|
||||
// A discovered database that CANNOT be dumped is a hard error: it means the undo would not exist.
|
||||
func (m *Manager) writeSafetyDump(ctx context.Context, stackName, nsRoot string) (string, error) {
|
||||
func (m *Manager) writeSafetyDump(ctx context.Context, stackName, nsRoot string) (safetyDumpSet, error) {
|
||||
discover := m.discoverDBs
|
||||
if discover == nil {
|
||||
discover = func(ctx context.Context) ([]DiscoveredDB, error) {
|
||||
@@ -127,7 +165,7 @@ func (m *Manager) writeSafetyDump(ctx context.Context, stackName, nsRoot string)
|
||||
}
|
||||
dbs, err := discover(ctx)
|
||||
if err != nil {
|
||||
return "", fmt.Errorf("a biztonsági mentés előtt nem sikerült felderíteni az adatbázisokat: %w", err)
|
||||
return safetyDumpSet{}, fmt.Errorf("a biztonsági mentés előtt nem sikerült felderíteni az adatbázisokat: %w", err)
|
||||
}
|
||||
var mine []DiscoveredDB
|
||||
for _, db := range dbs {
|
||||
@@ -136,34 +174,185 @@ func (m *Manager) writeSafetyDump(ctx context.Context, stackName, nsRoot string)
|
||||
}
|
||||
}
|
||||
if len(mine) == 0 {
|
||||
return "", nil // no DB → nothing to undo → scenario E flows unchanged
|
||||
return safetyDumpSet{}, nil // no DB → nothing to undo → scenario E flows unchanged
|
||||
}
|
||||
|
||||
dumpDir := AppDBDumpPath(nsRoot, stackName)
|
||||
if err := os.MkdirAll(dumpDir, 0755); err != nil {
|
||||
return "", fmt.Errorf("a biztonsági mentés könyvtára nem hozható létre: %w", err)
|
||||
return safetyDumpSet{}, fmt.Errorf("a biztonsági mentés könyvtára nem hozható létre: %w", err)
|
||||
}
|
||||
stamp := time.Now().UTC().Format("20060102T150405Z")
|
||||
first := ""
|
||||
set := safetyDumpSet{Stamp: time.Now().UTC().Format("20060102T150405Z")}
|
||||
for _, db := range mine {
|
||||
res := m.dumpForSafety(ctx, db, dumpDir)
|
||||
if res.Error != nil {
|
||||
return "", fmt.Errorf("a jelenlegi adatbázis biztonsági mentése sikertelen (%s): %w — a visszaállítás nem indult el", db.ContainerName, res.Error)
|
||||
return safetyDumpSet{}, fmt.Errorf("a jelenlegi adatbázis biztonsági mentése sikertelen (%s): %w — a visszaállítás nem indult el", db.ContainerName, res.Error)
|
||||
}
|
||||
// DumpOne writes `<stack>-<dbtype>.sql`; rename it under the safety prefix so it can never be
|
||||
// picked up as a replay SOURCE and can never overwrite the app's real dump.
|
||||
safe := filepath.Join(dumpDir, fmt.Sprintf("%s%s-%s-%s.sql", preRestoreDumpPrefix, stamp, stackName, db.DBType))
|
||||
safe := filepath.Join(dumpDir, fmt.Sprintf("%s%s-%s-%s.sql", preRestoreDumpPrefix, set.Stamp, stackName, db.DBType))
|
||||
if res.FilePath != safe {
|
||||
if err := os.Rename(res.FilePath, safe); err != nil {
|
||||
return "", fmt.Errorf("a biztonsági mentés véglegesítése sikertelen: %w", err)
|
||||
return safetyDumpSet{}, fmt.Errorf("a biztonsági mentés véglegesítése sikertelen: %w", err)
|
||||
}
|
||||
}
|
||||
if first == "" {
|
||||
first = safe
|
||||
}
|
||||
// EVERY file, not just the first — R-379, and the reason is on safetyDumpSet.
|
||||
set.Files = append(set.Files, safetyDumpFile{DB: db, Path: safe})
|
||||
m.logger.Printf("[INFO] [offbox] %s: pre-restore safety dump written → %s (%s)", stackName, filepath.Base(safe), humanizeBytes(res.Size))
|
||||
}
|
||||
return first, nil
|
||||
return set, nil
|
||||
}
|
||||
|
||||
// maxUndoCopiesPerApp is how many `pre-restore-` copies an app keeps.
|
||||
//
|
||||
// THREE, and the reasoning rather than a number pulled from the air. One is not enough: the case
|
||||
// that needs an undo is a restore that went wrong, and the second-guess attempt is exactly when the
|
||||
// customer reaches for the state before the FIRST attempt. Many is not free: they live inside the
|
||||
// recovery unit, so every one is also mirrored to Tier 2 AND pushed off-site permanently — four
|
||||
// accumulated on `docmost` in a single afternoon on 2026-08-22 (135 KB + 135 KB + 141 KB + 138 KB),
|
||||
// each of them forever. Three keeps two prior attempts and bounds the off-site growth.
|
||||
const maxUndoCopiesPerApp = 3
|
||||
|
||||
// pruneUndoCopies keeps the newest maxUndoCopiesPerApp undo copies for an app and removes the rest.
|
||||
//
|
||||
// DELIBERATELY NOT CALLED FROM THE RESTORE PATH. A delete on the failure path is how an undo goes
|
||||
// missing at exactly the moment it is needed; this runs from the capture side, where nothing is
|
||||
// depending on the files right now. It is called AFTER a successful capture, never before one.
|
||||
//
|
||||
// Ordering is by the stamp IN THE FILENAME, not by mtime and never by size: mtime moves when a file
|
||||
// is copied or a filesystem is restored, and the stamp is the identity writeSafetyDump assigned.
|
||||
// The newest is never a deletion candidate even if the list is somehow malformed.
|
||||
func (m *Manager) pruneUndoCopies(dumpDir, stack string) {
|
||||
entries, err := os.ReadDir(dumpDir)
|
||||
if err != nil {
|
||||
return
|
||||
}
|
||||
type undo struct{ name, stamp string }
|
||||
var undos []undo
|
||||
for _, e := range entries {
|
||||
if e.IsDir() || !strings.HasSuffix(e.Name(), ".sql") {
|
||||
continue
|
||||
}
|
||||
base := strings.TrimSuffix(e.Name(), ".sql")
|
||||
if !strings.HasPrefix(base, preRestoreDumpPrefix) {
|
||||
continue
|
||||
}
|
||||
after := strings.TrimPrefix(base, preRestoreDumpPrefix)
|
||||
i := strings.Index(after, "-")
|
||||
if i <= 0 {
|
||||
continue // not the shape writeSafetyDump writes — leave it alone rather than guess
|
||||
}
|
||||
undos = append(undos, undo{name: e.Name(), stamp: after[:i]})
|
||||
}
|
||||
if len(undos) <= maxUndoCopiesPerApp {
|
||||
return
|
||||
}
|
||||
sort.Slice(undos, func(a, b int) bool { return undos[a].stamp > undos[b].stamp }) // newest first
|
||||
for _, u := range undos[maxUndoCopiesPerApp:] {
|
||||
p := filepath.Join(dumpDir, u.name)
|
||||
if err := os.Remove(p); err != nil {
|
||||
m.logger.Printf("[WARN] [backup] %s: could not prune old undo copy %s: %v", stack, u.name, err)
|
||||
continue
|
||||
}
|
||||
m.logger.Printf("[INFO] [backup] %s: pruned old undo copy %s (keeping the newest %d)", stack, u.name, maxUndoCopiesPerApp)
|
||||
}
|
||||
}
|
||||
|
||||
// RestoreHoldFor reports whether an app is being held stopped after a failed restore + failed
|
||||
// rollback, and returns the customer-facing reason. Every start path consults this — the customer's
|
||||
// button, the app-stop Recover() starter, and the boot reconciler — because a hold that only one
|
||||
// path honours is not a hold.
|
||||
//
|
||||
// Nil settings ⇒ NOT held. That direction is deliberate and is the opposite of the usual fail-closed
|
||||
// rule: with no settings there is no hold recorded, so refusing every start would strand every app
|
||||
// on a misconfigured box. The write side logs loudly when it cannot persist (see
|
||||
// holdAppAfterFailedRollback), which is where that case is caught.
|
||||
func (m *Manager) RestoreHoldFor(stack string) (bool, string) {
|
||||
if m == nil || m.settings == nil {
|
||||
return false, ""
|
||||
}
|
||||
h, ok := m.settings.GetRestoreHold(stack)
|
||||
if !ok {
|
||||
return false, ""
|
||||
}
|
||||
when := h.At
|
||||
if t, err := time.Parse(time.RFC3339, h.At); err == nil {
|
||||
when = t.Format("2006-01-02 15:04")
|
||||
}
|
||||
return true, fmt.Sprintf("a(z) %s adatainak visszaállítása %s-kor megszakadt, és a korábbi állapotot sem sikerült visszatölteni. "+
|
||||
"Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek tovább. Vedd fel velünk a kapcsolatot", stack, when)
|
||||
}
|
||||
|
||||
// holdAppAfterFailedRollback records the R-379/R-380 hold and makes sure nothing restarts the app
|
||||
// behind our back.
|
||||
//
|
||||
// OPERATOR RULING, 2026-08-22: when the replay fails AND the rollback fails, the app is HELD
|
||||
// STOPPED rather than started. A running app on a half-written database lets the customer type into
|
||||
// it, and that turns a recoverable state into a permanent one. The alternative — start it and mark
|
||||
// it — was considered and declined.
|
||||
//
|
||||
// It ENDS the app-stop marker deliberately. The marker means "owed a restart"; a held app is not
|
||||
// owed one, and leaving the marker active would have Recover() start the broken app at the next
|
||||
// controller boot. The hold is the thing that persists, not the marker.
|
||||
func (m *Manager) holdAppAfterFailedRollback(stack string, replayErr, rollbackErr error) {
|
||||
if m.settings == nil {
|
||||
m.logger.Printf("[ERROR] [offbox] %s: cannot persist the restore hold — no settings wired; the app is stopped but NOTHING will refuse a restart", stack)
|
||||
return
|
||||
}
|
||||
h := settings.RestoreHold{
|
||||
Stack: stack,
|
||||
At: time.Now().UTC().Format(time.RFC3339),
|
||||
}
|
||||
if replayErr != nil {
|
||||
h.ReplayError = replayErr.Error()
|
||||
}
|
||||
if rollbackErr != nil {
|
||||
h.RollbackErr = rollbackErr.Error()
|
||||
}
|
||||
if err := m.settings.SetRestoreHold(h); err != nil {
|
||||
m.logger.Printf("[ERROR] [offbox] %s: persisting the restore hold FAILED: %v — the app is stopped and unguarded", stack, err)
|
||||
}
|
||||
// The app is not owed a restart; it is deliberately held. See the doc comment.
|
||||
if m.appStop != nil {
|
||||
m.appStop.End()
|
||||
}
|
||||
if m.restoreHoldNotify != nil {
|
||||
m.restoreHoldNotify(stack, replayErr, rollbackErr)
|
||||
}
|
||||
}
|
||||
|
||||
// SetRestoreHoldNotify wires the operator notification for a held app (cmd/controller/main.go).
|
||||
func (m *Manager) SetRestoreHoldNotify(fn func(stack string, replayErr, rollbackErr error)) {
|
||||
m.restoreHoldNotify = fn
|
||||
}
|
||||
|
||||
// rollbackSafetyDump re-applies THIS RUN's undo set, database by database, and is the whole of
|
||||
// R-379's fix: it is the same ImportDump call a person made by hand on 2026-08-22 to recover
|
||||
// `docmost` and `bookstack` after a failed replay, moved into the product.
|
||||
//
|
||||
// It runs with the DB service still up (the replay's own window) and BEFORE any restart, so the
|
||||
// app never observes the half-written state. An error here means the app cannot be trusted to run —
|
||||
// see the hold in ReconstituteFromOffsite.
|
||||
func (m *Manager) rollbackSafetyDump(ctx context.Context, stack string, set safetyDumpSet) error {
|
||||
if len(set.Files) == 0 {
|
||||
return nil
|
||||
}
|
||||
imp := m.rollbackImport
|
||||
if imp == nil {
|
||||
imp = func(ctx context.Context, db DiscoveredDB, path string) error {
|
||||
return ImportDump(ctx, db, path, m.logger, m.isDebug())
|
||||
}
|
||||
}
|
||||
for _, f := range set.Files {
|
||||
if _, sErr := os.Stat(f.Path); sErr != nil {
|
||||
return fmt.Errorf("a visszavonáshoz szükséges mentés nem található (%s): %w", filepath.Base(f.Path), sErr)
|
||||
}
|
||||
m.logger.Printf("[INFO] [offbox] %s: rolling back to the pre-restore state from %s", stack, filepath.Base(f.Path))
|
||||
if err := imp(ctx, f.DB, f.Path); err != nil {
|
||||
return fmt.Errorf("a korábbi állapot visszaállítása sikertelen (%s): %w", f.DB.ContainerName, err)
|
||||
}
|
||||
}
|
||||
m.logger.Printf("[INFO] [offbox] %s: rollback complete — %d database(s) returned to the pre-restore state", stack, len(set.Files))
|
||||
return nil
|
||||
}
|
||||
|
||||
// dumpForSafety is the DumpOne seam for the safety dump (tests inject; nil → the real DumpOne).
|
||||
@@ -325,10 +514,11 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string, ack
|
||||
// --- THE UNDO, BEFORE THE ACT ---------------------------------------------------------------
|
||||
// Taken while the stack is still UP (a stopped database cannot be dumped) and before a single
|
||||
// byte is overwritten, so a failure here aborts with the live app completely untouched.
|
||||
safety, err := m.writeSafetyDump(ctx, stack, liveNs)
|
||||
safetySet, err := m.writeSafetyDump(ctx, stack, liveNs)
|
||||
if err != nil {
|
||||
return res, err
|
||||
}
|
||||
safety := safetySet.First()
|
||||
res.SafetyDump = safety
|
||||
hasDB := safety != ""
|
||||
if hasDB {
|
||||
@@ -436,10 +626,39 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string, ack
|
||||
n, iErr := m.reimportDBDumpsFrom(ctx, stack, scratchDumpDir)
|
||||
res.DBsReplayed = n
|
||||
if iErr != nil {
|
||||
if sErr := restartStack(); sErr != nil {
|
||||
m.logger.Printf("[WARN] [offbox] %s: full start after failed replay also failed: %v", stack, sErr)
|
||||
// --- R-379/R-380: PUT THE CUSTOMER'S OWN COPY BACK ---------------------------------
|
||||
// Until v0.220.0 this branch restarted the app onto a HALF-WRITTEN database and named
|
||||
// the undo file in the message. Measured 2026-08-22 on demo-hp: Postgres was left
|
||||
// emptied and crash-looping; MariaDB was left partly applied while the app reported
|
||||
// `health=healthy`. Both are the same failure — a half state — and the only difference
|
||||
// was whether it looked broken. MariaDB's structural statements are not transactional,
|
||||
// so no engine flag can prevent the half state; putting the undo back is what removes
|
||||
// it. This is the same ImportDump call a person ran by hand that day to recover both
|
||||
// apps, moved into the product.
|
||||
//
|
||||
// The rollback runs BEFORE any restart and with the DB service still up, so the app
|
||||
// never observes the half state. The ORIGINAL replay error is never swallowed: it is
|
||||
// logged here in full and named in the customer's sentence.
|
||||
m.logger.Printf("[ERROR] [offbox] %s: database replay failed, rolling back to the pre-restore state: %v", stack, iErr)
|
||||
if rbErr := m.rollbackSafetyDump(ctx, stack, safetySet); rbErr != nil {
|
||||
// BOTH failed. Do NOT start the app: a running app on a half-written database lets
|
||||
// the customer type into it and makes the damage permanent. Hold it instead —
|
||||
// operator ruling, 2026-08-22.
|
||||
m.logger.Printf("[ERROR] [offbox] %s: ROLLBACK ALSO FAILED (%v) — holding the app stopped; replay error was: %v", stack, rbErr, iErr)
|
||||
m.holdAppAfterFailedRollback(stack, iErr, rbErr)
|
||||
return res, fmt.Errorf("a(z) %s adatbázisának visszaállítása sikertelen, és a korábbi állapot visszatöltése sem sikerült. "+
|
||||
"Az alkalmazást biztonsági okból LEÁLLÍTVA hagytuk, hogy az adatai ne sérüljenek tovább. "+
|
||||
"Vedd fel velünk a kapcsolatot — a korábbi állapot mentése megvan: %s", stack, filepath.Base(safety))
|
||||
}
|
||||
return res, fmt.Errorf("az adatbázis visszaállítása sikertelen: %w — a korábbi állapot mentése megvan: %s", iErr, filepath.Base(safety))
|
||||
res.RolledBack = true
|
||||
if sErr := restartStack(); sErr != nil {
|
||||
m.logger.Printf("[WARN] [offbox] %s: start after a successful rollback failed: %v", stack, sErr)
|
||||
}
|
||||
// Says BOTH things. A message that reported only the failure would leave the customer
|
||||
// believing their data was gone when it is back — the omission of a GAIN is as
|
||||
// misleading as the omission of a loss.
|
||||
return res, fmt.Errorf("a(z) %s adatbázisának visszaállítása sikertelen — az adataid visszakerültek a visszaállítás előtti állapotba, "+
|
||||
"az alkalmazás fut tovább. Ha újra megpróbálnád, előbb vedd fel velünk a kapcsolatot", stack)
|
||||
}
|
||||
}
|
||||
if err := restartStack(); err != nil {
|
||||
@@ -449,8 +668,11 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string, ack
|
||||
m.logger.Printf("[WARN] [offbox] %s reconstituted but health check failed: %v", stack, err)
|
||||
}
|
||||
|
||||
m.logger.Printf("[INFO] [offbox] reconstituted %s from snapshot %s: %d file(s) placed, %d DB dump(s) replayed, safety dump=%s, skewed=%v",
|
||||
stack, id, res.FilesPlaced, res.DBsReplayed, filepath.Base(safety), res.Skewed)
|
||||
// R-382: VolumesReplayed was set above and never printed, so the operator log said
|
||||
// "0 file(s) placed, 1 DB dump(s) replayed" on a run that returned a 52 MB Postgres data
|
||||
// directory — less informative than the customer's own flash, which already named the volumes.
|
||||
m.logger.Printf("[INFO] [offbox] reconstituted %s from snapshot %s: %d file(s) placed, %d volume(s) replayed, %d DB dump(s) replayed, safety dump=%s, skewed=%v",
|
||||
stack, id, res.FilesPlaced, res.VolumesReplayed, res.DBsReplayed, filepath.Base(safety), res.Skewed)
|
||||
return res, nil
|
||||
}
|
||||
|
||||
|
||||
@@ -35,6 +35,12 @@ func (m *Manager) SetOffboxFullPlaceCopier(fn func(src, dst string) (int, error)
|
||||
m.offboxFullPlaceCopier = fn
|
||||
}
|
||||
|
||||
// SetRollbackImportFn overrides the ROLLBACK's ImportDump (tests; no Docker needed). Separate from
|
||||
// the replay's own import seam on purpose — see the field comment on Manager.rollbackImport.
|
||||
func (m *Manager) SetRollbackImportFn(fn func(ctx context.Context, db DiscoveredDB, dumpPath string) error) {
|
||||
m.rollbackImport = fn
|
||||
}
|
||||
|
||||
// SetSafetyDumpFn overrides the pre-restore safety dump (tests; no Docker needed).
|
||||
func (m *Manager) SetSafetyDumpFn(fn func(ctx context.Context, db DiscoveredDB, dumpDir string) DumpResult) {
|
||||
m.safetyDumpFn = fn
|
||||
|
||||
@@ -45,7 +45,11 @@ func TestR355_SafetyDumpIsTakenForTheCorrectlyAttributedApp(t *testing.T) {
|
||||
return DumpResult{DB: db, FilePath: p, Size: 13}
|
||||
}
|
||||
|
||||
safety, err := m.writeSafetyDump(context.Background(), "paperless-ngx", nsRoot)
|
||||
// v0.220.0 (R-379): writeSafetyDump returns the SET it wrote, so a rollback can re-apply EVERY
|
||||
// database's undo. `.First()` is the value this signature returned before; these assertions are
|
||||
// unchanged in meaning.
|
||||
set, err := m.writeSafetyDump(context.Background(), "paperless-ngx", nsRoot)
|
||||
safety := set.First()
|
||||
if err != nil {
|
||||
t.Fatalf("writeSafetyDump: %v", err)
|
||||
}
|
||||
@@ -83,7 +87,11 @@ func TestR355_MisattributedAppGetsNoUndoCopy(t *testing.T) {
|
||||
return DumpResult{DB: db}
|
||||
}
|
||||
|
||||
safety, err := m.writeSafetyDump(context.Background(), "paperless-ngx", nsRoot)
|
||||
// v0.220.0 (R-379): writeSafetyDump returns the SET it wrote, so a rollback can re-apply EVERY
|
||||
// database's undo. `.First()` is the value this signature returned before; these assertions are
|
||||
// unchanged in meaning.
|
||||
set, err := m.writeSafetyDump(context.Background(), "paperless-ngx", nsRoot)
|
||||
safety := set.First()
|
||||
if err != nil {
|
||||
t.Fatalf("writeSafetyDump: %v", err)
|
||||
}
|
||||
@@ -111,7 +119,11 @@ func TestR355_RestoreRefusesWhenTheUndoCannotBeTaken(t *testing.T) {
|
||||
return DumpResult{DB: db, Error: os.ErrPermission}
|
||||
}
|
||||
|
||||
safety, err := m.writeSafetyDump(context.Background(), "paperless-ngx", nsRoot)
|
||||
// v0.220.0 (R-379): writeSafetyDump returns the SET it wrote, so a rollback can re-apply EVERY
|
||||
// database's undo. `.First()` is the value this signature returned before; these assertions are
|
||||
// unchanged in meaning.
|
||||
set, err := m.writeSafetyDump(context.Background(), "paperless-ngx", nsRoot)
|
||||
safety := set.First()
|
||||
if err == nil {
|
||||
t.Fatal("a database that cannot be dumped must be a hard error — the undo would not exist")
|
||||
}
|
||||
|
||||
@@ -0,0 +1,330 @@
|
||||
package backup
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// R-379 / R-380. When an off-site database replay failed, the app was restarted onto a HALF-WRITTEN
|
||||
// database and the customer was shown the undo copy's filename — a file nothing in the product could
|
||||
// apply. Measured live on demo-hp 2026-08-22: Postgres left emptied and crash-looping, MariaDB left
|
||||
// partly applied while `docker inspect` reported health=healthy. Both are the same failure — a half
|
||||
// state — and the only difference was whether it looked broken.
|
||||
//
|
||||
// The fix puts the customer's own pre-restore copy back automatically. These tests pin that, the
|
||||
// double-failure hold, and the two absence claims that go with them.
|
||||
|
||||
// rollbackFixture is reconFixture with the replay failing and the rollback injectable.
|
||||
func rollbackFixture(t *testing.T, rollbackErr error) (*Manager, *recordingProvider, *int) {
|
||||
t.Helper()
|
||||
m, prov, _ := reconFixture(t, "run1", "2026-07-19T06:00:00Z", pgDump(1))
|
||||
m.importDBDump = func(context.Context, DiscoveredDB, string) error {
|
||||
return errors.New("replay blew up")
|
||||
}
|
||||
calls := 0
|
||||
m.SetRollbackImportFn(func(_ context.Context, _ DiscoveredDB, path string) error {
|
||||
calls++
|
||||
// The rollback must be handed THIS RUN's undo file, not the scratch's dump.
|
||||
if !strings.Contains(filepath.Base(path), preRestoreDumpPrefix) {
|
||||
t.Errorf("rollback was handed %q — that is not an undo copy", filepath.Base(path))
|
||||
}
|
||||
return rollbackErr
|
||||
})
|
||||
return m, prov, &calls
|
||||
}
|
||||
|
||||
// SCENARIO A — replay fails, rollback succeeds. The app comes back and the outcome records it.
|
||||
//
|
||||
// WRONG OUTCOMES PINNED: the app left down; or a message that reports the failure and omits that the
|
||||
// data is back — the omission of a GAIN is as misleading as the omission of a loss, and here it is
|
||||
// the difference between a customer who panics and one who does not.
|
||||
func TestR379_ScenarioA_RollbackSucceeds_AppComesBack(t *testing.T) {
|
||||
m, prov, calls := rollbackFixture(t, nil)
|
||||
|
||||
res, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
|
||||
if err == nil {
|
||||
t.Fatal("a failed replay must still be surfaced as a failure")
|
||||
}
|
||||
if *calls != 1 {
|
||||
t.Fatalf("the undo copy must be re-applied exactly once, got %d", *calls)
|
||||
}
|
||||
if !res.RolledBack {
|
||||
t.Error("the result must record that the rollback happened")
|
||||
}
|
||||
if !prov.fullStarted {
|
||||
t.Fatal("after a successful rollback the app must be started — it is in a known good state")
|
||||
}
|
||||
// The customer sentence says BOTH things.
|
||||
low := err.Error()
|
||||
if !strings.Contains(low, "sikertelen") {
|
||||
t.Errorf("the message must say the restore failed; got: %v", err)
|
||||
}
|
||||
if !strings.Contains(low, "visszakerültek") {
|
||||
t.Errorf("the message must say the data is back as it was; got: %v", err)
|
||||
}
|
||||
// And no hold was written — the app is fine.
|
||||
if held, _ := m.RestoreHoldFor("immich"); held {
|
||||
t.Error("a recovered app must NOT be held")
|
||||
}
|
||||
}
|
||||
|
||||
// SCENARIO C — replay fails AND rollback fails. The app is held, not started.
|
||||
//
|
||||
// OPERATOR RULING 2026-08-22: a running app on a half-written database lets the customer type into
|
||||
// it and makes the damage permanent.
|
||||
//
|
||||
// WRONG OUTCOMES PINNED: the app started anyway; or the app-stop marker left active, which would
|
||||
// have Recover() start the broken app at the next controller boot, quietly, hours later.
|
||||
func TestR379_ScenarioC_BothFail_AppIsHeldNotStarted(t *testing.T) {
|
||||
m, prov, calls := rollbackFixture(t, errors.New("rollback blew up too"))
|
||||
notified := 0
|
||||
m.SetRestoreHoldNotify(func(string, error, error) { notified++ })
|
||||
|
||||
_, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
|
||||
if err == nil {
|
||||
t.Fatal("a double failure must be surfaced")
|
||||
}
|
||||
if *calls != 1 {
|
||||
t.Fatalf("the rollback must have been attempted, got %d calls", *calls)
|
||||
}
|
||||
// THE OBSERVABLE THAT MATTERS.
|
||||
if prov.fullStarted {
|
||||
t.Fatal("the app was STARTED onto a half-written database — the exact outcome the hold exists to prevent")
|
||||
}
|
||||
held, why := m.RestoreHoldFor("immich")
|
||||
if !held {
|
||||
t.Fatal("the hold must be persisted, or nothing will refuse a restart later")
|
||||
}
|
||||
if why == "" || !strings.Contains(why, "kapcsolat") {
|
||||
t.Errorf("the hold's reason must name a route the customer can take; got %q", why)
|
||||
}
|
||||
if notified != 1 {
|
||||
t.Errorf("the operator must be told exactly once, got %d", notified)
|
||||
}
|
||||
if !strings.Contains(err.Error(), "LEÁLLÍTVA") {
|
||||
t.Errorf("the customer must be told the app was deliberately stopped; got: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// SCENARIO E — the replay SUCCEEDS. Nothing new may happen.
|
||||
//
|
||||
// WRONG OUTCOME PINNED: a rollback firing on a successful restore would overwrite the restored data
|
||||
// with the pre-restore state — a silent, total loss of the thing the customer asked for.
|
||||
func TestR379_ScenarioE_SuccessfulReplay_NoRollbackNoHold(t *testing.T) {
|
||||
m, prov, _ := reconFixture(t, "run1", "2026-07-19T06:00:00Z", pgDump(1))
|
||||
rolled := 0
|
||||
m.SetRollbackImportFn(func(context.Context, DiscoveredDB, string) error { rolled++; return nil })
|
||||
held := 0
|
||||
m.SetRestoreHoldNotify(func(string, error, error) { held++ })
|
||||
|
||||
res, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
|
||||
if err != nil {
|
||||
t.Fatalf("a clean restore must succeed: %v", err)
|
||||
}
|
||||
if rolled != 0 {
|
||||
t.Fatalf("a rollback fired on a SUCCESSFUL restore — the restored data would be overwritten; calls=%d", rolled)
|
||||
}
|
||||
if res.RolledBack {
|
||||
t.Error("a successful restore must not report a rollback")
|
||||
}
|
||||
if held != 0 {
|
||||
t.Errorf("no hold may be written on success, got %d", held)
|
||||
}
|
||||
if isHeld, _ := m.RestoreHoldFor("immich"); isHeld {
|
||||
t.Error("a successful restore must leave no hold")
|
||||
}
|
||||
if !prov.fullStarted {
|
||||
t.Error("a successful restore still starts the app")
|
||||
}
|
||||
}
|
||||
|
||||
// SCENARIO E, POSITIVE CONTROL. "No rollback fired" is an absence claim, so prove the counter can
|
||||
// count: the same seam, driven through the failure path, must register.
|
||||
func TestR379_ScenarioE_PositiveControl_TheCounterCanCount(t *testing.T) {
|
||||
m, _, calls := rollbackFixture(t, nil)
|
||||
if _, err := m.ReconstituteFromOffsite(context.Background(), "immich", false); err == nil {
|
||||
t.Fatal("fixture: the replay must fail here")
|
||||
}
|
||||
if *calls == 0 {
|
||||
t.Fatal("the rollback counter never increments — Scenario E's zero would have proven nothing")
|
||||
}
|
||||
}
|
||||
|
||||
// SCENARIO F — an app with no database. Untouched in every respect.
|
||||
func TestR379_ScenarioF_NoDatabase_Unchanged(t *testing.T) {
|
||||
m, prov, _ := reconFixture(t, "run1", "2026-07-19T06:00:00Z", "")
|
||||
m.discoverDBs = func(context.Context) ([]DiscoveredDB, error) { return nil, nil }
|
||||
rolled := 0
|
||||
m.SetRollbackImportFn(func(context.Context, DiscoveredDB, string) error { rolled++; return nil })
|
||||
|
||||
res, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
|
||||
if err != nil {
|
||||
t.Fatalf("a no-DB app must restore: %v", err)
|
||||
}
|
||||
if rolled != 0 || res.RolledBack {
|
||||
t.Error("a no-DB app has nothing to undo and nothing to roll back")
|
||||
}
|
||||
if res.SafetyDump != "" {
|
||||
t.Errorf("a no-DB app takes no undo copy, got %q", res.SafetyDump)
|
||||
}
|
||||
if held, _ := m.RestoreHoldFor("immich"); held {
|
||||
t.Error("a no-DB app must never be held")
|
||||
}
|
||||
if !prov.fullStarted {
|
||||
t.Error("a no-DB app still starts")
|
||||
}
|
||||
}
|
||||
|
||||
// SCENARIO G — two databases, one replay fails. The WHOLE undo set is re-applied.
|
||||
//
|
||||
// WRONG OUTCOME PINNED: only the first. writeSafetyDump used to return one path for an app with two
|
||||
// databases, so a rollback built on that value would restore one and leave the other half-written —
|
||||
// this defect, one database over.
|
||||
func TestR379_ScenarioG_TwoDatabases_WholeSetRolledBack(t *testing.T) {
|
||||
m, _, _ := reconFixture(t, "run1", "2026-07-19T06:00:00Z", pgDump(1))
|
||||
two := []DiscoveredDB{
|
||||
{StackName: "immich", DBType: DBTypePostgres, ContainerName: "immich-postgres", ContainerID: "a"},
|
||||
{StackName: "immich", DBType: DBTypeMariaDB, ContainerName: "immich-maria", ContainerID: "b"},
|
||||
}
|
||||
m.discoverDBs = func(context.Context) ([]DiscoveredDB, error) { return two, nil }
|
||||
m.safetyDumpFn = func(_ context.Context, db DiscoveredDB, dumpDir string) DumpResult {
|
||||
p := filepath.Join(dumpDir, "immich-"+string(db.DBType)+".sql")
|
||||
if err := os.WriteFile(p, []byte(pgDump(1)), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return DumpResult{DB: db, FilePath: p, Size: 42}
|
||||
}
|
||||
m.importDBDump = func(context.Context, DiscoveredDB, string) error { return errors.New("replay failed") }
|
||||
|
||||
var rolled []string
|
||||
m.SetRollbackImportFn(func(_ context.Context, db DiscoveredDB, path string) error {
|
||||
rolled = append(rolled, db.ContainerName+":"+filepath.Base(path))
|
||||
return nil
|
||||
})
|
||||
|
||||
if _, err := m.ReconstituteFromOffsite(context.Background(), "immich", false); err == nil {
|
||||
t.Fatal("the replay failure must surface")
|
||||
}
|
||||
if len(rolled) != 2 {
|
||||
t.Fatalf("BOTH databases must be rolled back, got %d: %v", len(rolled), rolled)
|
||||
}
|
||||
// Each database got its OWN undo file, matched by identity rather than by prefix.
|
||||
if !strings.Contains(rolled[0], "immich-postgres:") || !strings.Contains(rolled[1], "immich-maria:") {
|
||||
t.Errorf("each database must get its own undo copy, got %v", rolled)
|
||||
}
|
||||
for _, r := range rolled {
|
||||
if !strings.Contains(r, preRestoreDumpPrefix) {
|
||||
t.Errorf("rollback used a non-undo file: %s", r)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// The undo set is matched on THIS RUN's stamp. Three undo copies from three different runs
|
||||
// accumulated on docmost in one afternoon; a prefix match would replay an arbitrary older state.
|
||||
func TestR379_UndoSetIsThisRunOnly(t *testing.T) {
|
||||
nsRoot := t.TempDir()
|
||||
m := newSafetyTestManager()
|
||||
m.discoverDBs = func(context.Context) ([]DiscoveredDB, error) {
|
||||
return []DiscoveredDB{{StackName: "app", DBType: DBTypePostgres, ContainerName: "app-postgres", ContainerID: "c"}}, nil
|
||||
}
|
||||
m.safetyDumpFn = func(_ context.Context, db DiscoveredDB, dumpDir string) DumpResult {
|
||||
p := filepath.Join(dumpDir, "app-postgres.sql")
|
||||
if err := os.WriteFile(p, []byte("x"), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return DumpResult{DB: db, FilePath: p, Size: 1}
|
||||
}
|
||||
// An OLDER undo copy from a previous run, already on disk.
|
||||
dumpDir := AppDBDumpPath(nsRoot, "app")
|
||||
if err := os.MkdirAll(dumpDir, 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
stale := filepath.Join(dumpDir, preRestoreDumpPrefix+"20200101T000000Z-app-postgres.sql")
|
||||
if err := os.WriteFile(stale, []byte("STALE"), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
set, err := m.writeSafetyDump(context.Background(), "app", nsRoot)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if len(set.Files) != 1 {
|
||||
t.Fatalf("one database → one undo file, got %d", len(set.Files))
|
||||
}
|
||||
if strings.Contains(set.Files[0].Path, "20200101") {
|
||||
t.Fatal("the set picked up an OLDER run's undo copy — it must carry only this run's stamp")
|
||||
}
|
||||
if set.Stamp == "" || strings.Contains(set.Files[0].Path, set.Stamp) == false {
|
||||
t.Errorf("the file must carry this run's stamp %q, got %q", set.Stamp, set.Files[0].Path)
|
||||
}
|
||||
}
|
||||
|
||||
// pruneUndoCopies keeps the newest N and never the oldest — ordered by the STAMP, not by mtime.
|
||||
func TestR379_PruneKeepsTheNewestByStamp(t *testing.T) {
|
||||
m := newSafetyTestManager()
|
||||
dir := t.TempDir()
|
||||
stamps := []string{"20260101T000000Z", "20260201T000000Z", "20260301T000000Z", "20260401T000000Z", "20260501T000000Z"}
|
||||
for _, st := range stamps {
|
||||
if err := os.WriteFile(filepath.Join(dir, preRestoreDumpPrefix+st+"-app-postgres.sql"), []byte("x"), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
// The app's OWN dump must never be a candidate.
|
||||
own := filepath.Join(dir, "app-postgres.sql")
|
||||
if err := os.WriteFile(own, []byte("own"), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
m.pruneUndoCopies(dir, "app")
|
||||
|
||||
if _, err := os.Stat(own); err != nil {
|
||||
t.Fatal("the app's own dump was pruned — only undo copies are candidates")
|
||||
}
|
||||
for _, st := range stamps[len(stamps)-maxUndoCopiesPerApp:] {
|
||||
if _, err := os.Stat(filepath.Join(dir, preRestoreDumpPrefix+st+"-app-postgres.sql")); err != nil {
|
||||
t.Errorf("the newest %d must survive; %s is gone", maxUndoCopiesPerApp, st)
|
||||
}
|
||||
}
|
||||
for _, st := range stamps[:len(stamps)-maxUndoCopiesPerApp] {
|
||||
if _, err := os.Stat(filepath.Join(dir, preRestoreDumpPrefix+st+"-app-postgres.sql")); err == nil {
|
||||
t.Errorf("the oldest must be pruned; %s survived", st)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// R-381 — the customer sentence must not carry the engine's output.
|
||||
//
|
||||
// MEASURED 2026-08-22: 407 bytes (Postgres, with a caret diagram and `exit status 3`) and 615 bytes
|
||||
// (MariaDB, whose middle was an `INSERT INTO migrations VALUES (...)` listing — ROWS OUT OF THE
|
||||
// CUSTOMER'S OWN DATABASE, HTML-escaped, on their dashboard).
|
||||
func TestR381_CustomerMessageCarriesNoEngineOutput(t *testing.T) {
|
||||
engineNoise := "ERROR: syntax error at end of input\nLINE 1: COPY public.felhom_r356b_discriminator \n ^ — exit status 3"
|
||||
m, _, _ := reconFixture(t, "run1", "2026-07-19T06:00:00Z", pgDump(1))
|
||||
m.importDBDump = func(context.Context, DiscoveredDB, string) error {
|
||||
return errors.New(engineNoise)
|
||||
}
|
||||
m.SetRollbackImportFn(func(context.Context, DiscoveredDB, string) error { return nil })
|
||||
|
||||
_, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
|
||||
if err == nil {
|
||||
t.Fatal("the failure must surface")
|
||||
}
|
||||
msg := err.Error()
|
||||
for _, leak := range []string{"ERROR: syntax", "LINE 1:", "COPY public.", "exit status"} {
|
||||
if strings.Contains(msg, leak) {
|
||||
t.Errorf("the customer message leaks engine output %q; got: %s", leak, msg)
|
||||
}
|
||||
}
|
||||
if len(msg) > 320 {
|
||||
t.Errorf("the customer message is %d bytes — it is a sentence, not a transcript: %s", len(msg), msg)
|
||||
}
|
||||
// POSITIVE CONTROL: the noise really was in the error the code received, so the absence above is
|
||||
// the message being clean rather than the noise never existing.
|
||||
if !strings.Contains(engineNoise, "exit status") {
|
||||
t.Fatal("fixture: the planted noise does not contain the marker this test greps for")
|
||||
}
|
||||
}
|
||||
@@ -264,11 +264,19 @@ func TestRestoreFromUnitIgnoresSafetyDumpsWhenDecidingToReplay(t *testing.T) {
|
||||
// TestReconstituteReplayFailureStillBringsTheStackUp: the DB-only window is a deliberate half-started
|
||||
// state, so EVERY exit from it must end in a full start. Otherwise a failed restore leaves the
|
||||
// customer with a running database and no application — an outage caused by the recovery tool.
|
||||
//
|
||||
// UPDATED FOR v0.220.0 (R-379). The requirement is unchanged and still asserted; what changed is
|
||||
// what happens BETWEEN the failure and the full start. A failed replay now re-applies the customer's
|
||||
// own pre-restore copy first, and only then starts. The rollback seam is injected as SUCCEEDING here
|
||||
// because that is this test's subject; the double-failure path — where the app is deliberately NOT
|
||||
// started — is Scenario C and has its own test in r379_rollback_test.go.
|
||||
func TestReconstituteReplayFailureStillBringsTheStackUp(t *testing.T) {
|
||||
m, prov, _ := reconFixture(t, "run1", "2026-07-19T06:00:00Z", pgDump(1))
|
||||
m.importDBDump = func(context.Context, DiscoveredDB, string) error {
|
||||
return context.DeadlineExceeded
|
||||
}
|
||||
rolledBack := 0
|
||||
m.SetRollbackImportFn(func(context.Context, DiscoveredDB, string) error { rolledBack++; return nil })
|
||||
|
||||
res, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
|
||||
if err == nil {
|
||||
@@ -280,9 +288,16 @@ func TestReconstituteReplayFailureStillBringsTheStackUp(t *testing.T) {
|
||||
if got := strings.Join(prov.calls, ","); got != "stop,startsvc:immich-postgres,start" {
|
||||
t.Fatalf("sequence = %q, want the best-effort full start after the failure", got)
|
||||
}
|
||||
// The existing message shape stays: the operator needs the undo's filename.
|
||||
if !strings.Contains(err.Error(), filepath.Base(res.SafetyDump)) {
|
||||
t.Fatalf("the error must name the safety dump so the operator can undo, got: %v", err)
|
||||
if rolledBack != 1 {
|
||||
t.Fatalf("the customer's pre-restore copy must be put back before the start; rollback calls = %d", rolledBack)
|
||||
}
|
||||
if !res.RolledBack {
|
||||
t.Error("the outcome must record that a rollback happened — the surface has to be able to say the data is back")
|
||||
}
|
||||
// v0.220.0: the customer sentence now states the OUTCOME (their data is as it was) instead of a
|
||||
// filename. The filename remains for the operator, in the log and on the result.
|
||||
if res.SafetyDump == "" || !strings.Contains(filepath.Base(res.SafetyDump), preRestoreDumpPrefix) {
|
||||
t.Fatalf("the result must still carry the undo copy for the operator, got %q", res.SafetyDump)
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -195,6 +195,11 @@ func (m *Manager) CaptureRecoveryUnit(stackName string) error {
|
||||
stackName, RecoveryUnitPath(nsRoot, stackName), len(info.ImagePins), len(info.SecretEnvVars),
|
||||
len(info.DataKeyEnvVars), len(info.PortableSecrets), len(info.PortableSecretEnvVars),
|
||||
len(withheldSecretNames(info)))
|
||||
|
||||
// R-379: bound the undo copies. AFTER a successful capture and never before one — the capture is
|
||||
// the point at which nothing is depending on those files, whereas the restore path is precisely
|
||||
// where deleting one would remove the undo at the moment it is needed.
|
||||
m.pruneUndoCopies(AppDBDumpPath(nsRoot, stackName), stackName)
|
||||
return nil
|
||||
}
|
||||
|
||||
|
||||
Reference in New Issue
Block a user