R-379/R-380: put the customer's undo copy back when a database restore fails
gates / gates (push) Successful in 11s

R-379 and R-380 were one failure. Both ended with a half-restored database; the
only difference was whether it looked broken. Postgres emptied and crash-looped;
MariaDB applied part of the dump and reported health=healthy with a zero-row
schema-version table. Measured live on demo-hp 2026-08-22.

The undo copy was already taken and already good - proven by hand that day on
both engines. Nothing in the product could apply it. Now it does, with the same
ImportDump call, before any restart and inside the DB-only window.

The WHOLE undo set, matched on this run's stamp. writeSafetyDump returned one
path for an app with two databases; a rollback on that would restore one and
leave the other half-written.

When the rollback also fails the app is HELD STOPPED (operator ruling): a running
app on a half-written database lets the customer make the damage permanent. Every
start path refuses it - customer button, appstop Recover, boot sweep - via the
shared driveStartGate, checked ABOVE its driveless early return because these
apps have no drive. The marker is ended so nothing auto-restarts it. The row goes
red. Cleared with --clear-restore-hold, an operator CLI route.

--single-transaction is a belt on Postgres only; MariaDB DDL is not transactional
and that is why the rollback is the fix.

R-381: the engine's stderr stops reaching the customer (615 bytes on MariaDB, its
middle rows out of their own database) and starts reaching the operator log,
which never had it.
R-382: the summary log prints the volume count it already held.
Undo copies resolve to their own app, are marked IsUndo, and are capped at 3 per
app, pruned from the capture side. The reported render-as-an-app symptom did NOT
reproduce - the live page was read first and had zero occurrences.

Tests 1468 -> 1483. Eight red-proofs; ONE PASSED and is reported: the R-381
behavioural test injected below ImportDump. A guard at that layer now convicts.
This commit is contained in:
2026-08-22 18:03:18 +02:00
parent 0f3cf0dbb2
commit 2c724c9283
15 changed files with 1288 additions and 37 deletions
+247 -25
View File
@@ -6,8 +6,11 @@ import (
"os"
"os/exec"
"path/filepath"
"sort"
"strings"
"time"
"gitea.dooplex.hu/admin/felhom-controller/internal/settings"
)
// Offsite reconstitution (R-43, v0.148.0) — the leg that was missing.
@@ -71,11 +74,16 @@ type OffsiteReconstituteResult struct {
// restore that had silently dropped a 1.4 MB volume archive — a true sentence leaving a false
// impression, which is the shape this surface keeps having removed from it.
VolumesReplayed int
SafetyDump string // path of the pre-restore dump (the undo), "" when the app has no DB
DumpsAt time.Time // when the snapshot's DB half was taken (zero = unknown/legacy unit)
OffsiteRunID string // "" for a pre-v0.148 snapshot — an unverified pair
Skewed bool // the snapshot carries no coherence stamp: files and DB may differ in age
LooksEmpty bool // R-44 sniff on the dump about to be replayed
SafetyDump string // path of the pre-restore dump (the undo), "" when the app has no DB
// RolledBack (R-379) is true when the database replay FAILED and this run put the customer's own
// pre-restore copy back. It is on the result rather than inferred from the error, because "the
// restore failed" and "your data is as it was" are two different facts and the surface has to be
// able to say both.
RolledBack bool
DumpsAt time.Time // when the snapshot's DB half was taken (zero = unknown/legacy unit)
OffsiteRunID string // "" for a pre-v0.148 snapshot — an unverified pair
Skewed bool // the snapshot carries no coherence stamp: files and DB may differ in age
LooksEmpty bool // R-44 sniff on the dump about to be replayed
// Placement (R-351) is what the backup recorded about where this app's data lived, compared
// against where this restore actually wrote. Carried on the RESULT and not only on the refusal,
// so a restore that proceeded into a different destination says so in its own outcome rather
@@ -112,13 +120,43 @@ func rsyncRestoreOverwrite(src, dst string) (int, error) {
return countRestoredFiles(string(out)), nil
}
// safetyDumpSet is what ONE reconstitution's undo consists of: the stamp that identifies this run's
// files, and one written path per database the app has.
//
// R-379: it exists because `writeSafetyDump` used to return only the FIRST path, and the rollback
// added in v0.220.0 must re-apply EVERY database's undo or it restores one and leaves the other
// half-written — the defect it exists to close, one database over. The stamp is the IDENTITY: three
// runs against `docmost` on 2026-08-22 left three `pre-restore-*` files in the same directory, so
// matching on the prefix would replay an arbitrary older state. Match on this stamp, never on the
// prefix, never on age or size.
type safetyDumpSet struct {
Stamp string // 20060102T150405Z — this run's, and only this run's
Files []safetyDumpFile // one per database, in discovery order
}
// safetyDumpFile pairs an undo file with the database it came from, so the rollback can hand each
// dump back to the container it belongs to instead of guessing from the filename.
type safetyDumpFile struct {
DB DiscoveredDB
Path string
}
// First returns the first written path, or "" — the value the pre-v0.220.0 signature returned, kept
// because the customer-facing message names one file and changing that is not this task.
func (s safetyDumpSet) First() string {
if len(s.Files) == 0 {
return ""
}
return s.Files[0].Path
}
// writeSafetyDump dumps every live database of stack into the app's unit db-dumps dir under the
// `pre-restore-` prefix, and returns the first dump's path. Returns ("", nil) when the app has no
// `pre-restore-` prefix, and returns the SET it wrote. Returns (zero, nil) when the app has no
// database at all — a no-DB app has nothing to undo and must flow exactly as it did before
// v0.148.0 (no dump, no replay, no behaviour change).
//
// A discovered database that CANNOT be dumped is a hard error: it means the undo would not exist.
func (m *Manager) writeSafetyDump(ctx context.Context, stackName, nsRoot string) (string, error) {
func (m *Manager) writeSafetyDump(ctx context.Context, stackName, nsRoot string) (safetyDumpSet, error) {
discover := m.discoverDBs
if discover == nil {
discover = func(ctx context.Context) ([]DiscoveredDB, error) {
@@ -127,7 +165,7 @@ func (m *Manager) writeSafetyDump(ctx context.Context, stackName, nsRoot string)
}
dbs, err := discover(ctx)
if err != nil {
return "", fmt.Errorf("a biztonsági mentés előtt nem sikerült felderíteni az adatbázisokat: %w", err)
return safetyDumpSet{}, fmt.Errorf("a biztonsági mentés előtt nem sikerült felderíteni az adatbázisokat: %w", err)
}
var mine []DiscoveredDB
for _, db := range dbs {
@@ -136,34 +174,185 @@ func (m *Manager) writeSafetyDump(ctx context.Context, stackName, nsRoot string)
}
}
if len(mine) == 0 {
return "", nil // no DB → nothing to undo → scenario E flows unchanged
return safetyDumpSet{}, nil // no DB → nothing to undo → scenario E flows unchanged
}
dumpDir := AppDBDumpPath(nsRoot, stackName)
if err := os.MkdirAll(dumpDir, 0755); err != nil {
return "", fmt.Errorf("a biztonsági mentés könyvtára nem hozható létre: %w", err)
return safetyDumpSet{}, fmt.Errorf("a biztonsági mentés könyvtára nem hozható létre: %w", err)
}
stamp := time.Now().UTC().Format("20060102T150405Z")
first := ""
set := safetyDumpSet{Stamp: time.Now().UTC().Format("20060102T150405Z")}
for _, db := range mine {
res := m.dumpForSafety(ctx, db, dumpDir)
if res.Error != nil {
return "", fmt.Errorf("a jelenlegi adatbázis biztonsági mentése sikertelen (%s): %w — a visszaállítás nem indult el", db.ContainerName, res.Error)
return safetyDumpSet{}, fmt.Errorf("a jelenlegi adatbázis biztonsági mentése sikertelen (%s): %w — a visszaállítás nem indult el", db.ContainerName, res.Error)
}
// DumpOne writes `<stack>-<dbtype>.sql`; rename it under the safety prefix so it can never be
// picked up as a replay SOURCE and can never overwrite the app's real dump.
safe := filepath.Join(dumpDir, fmt.Sprintf("%s%s-%s-%s.sql", preRestoreDumpPrefix, stamp, stackName, db.DBType))
safe := filepath.Join(dumpDir, fmt.Sprintf("%s%s-%s-%s.sql", preRestoreDumpPrefix, set.Stamp, stackName, db.DBType))
if res.FilePath != safe {
if err := os.Rename(res.FilePath, safe); err != nil {
return "", fmt.Errorf("a biztonsági mentés véglegesítése sikertelen: %w", err)
return safetyDumpSet{}, fmt.Errorf("a biztonsági mentés véglegesítése sikertelen: %w", err)
}
}
if first == "" {
first = safe
}
// EVERY file, not just the first — R-379, and the reason is on safetyDumpSet.
set.Files = append(set.Files, safetyDumpFile{DB: db, Path: safe})
m.logger.Printf("[INFO] [offbox] %s: pre-restore safety dump written → %s (%s)", stackName, filepath.Base(safe), humanizeBytes(res.Size))
}
return first, nil
return set, nil
}
// maxUndoCopiesPerApp is how many `pre-restore-` copies an app keeps.
//
// THREE, and the reasoning rather than a number pulled from the air. One is not enough: the case
// that needs an undo is a restore that went wrong, and the second-guess attempt is exactly when the
// customer reaches for the state before the FIRST attempt. Many is not free: they live inside the
// recovery unit, so every one is also mirrored to Tier 2 AND pushed off-site permanently — four
// accumulated on `docmost` in a single afternoon on 2026-08-22 (135 KB + 135 KB + 141 KB + 138 KB),
// each of them forever. Three keeps two prior attempts and bounds the off-site growth.
const maxUndoCopiesPerApp = 3
// pruneUndoCopies keeps the newest maxUndoCopiesPerApp undo copies for an app and removes the rest.
//
// DELIBERATELY NOT CALLED FROM THE RESTORE PATH. A delete on the failure path is how an undo goes
// missing at exactly the moment it is needed; this runs from the capture side, where nothing is
// depending on the files right now. It is called AFTER a successful capture, never before one.
//
// Ordering is by the stamp IN THE FILENAME, not by mtime and never by size: mtime moves when a file
// is copied or a filesystem is restored, and the stamp is the identity writeSafetyDump assigned.
// The newest is never a deletion candidate even if the list is somehow malformed.
func (m *Manager) pruneUndoCopies(dumpDir, stack string) {
entries, err := os.ReadDir(dumpDir)
if err != nil {
return
}
type undo struct{ name, stamp string }
var undos []undo
for _, e := range entries {
if e.IsDir() || !strings.HasSuffix(e.Name(), ".sql") {
continue
}
base := strings.TrimSuffix(e.Name(), ".sql")
if !strings.HasPrefix(base, preRestoreDumpPrefix) {
continue
}
after := strings.TrimPrefix(base, preRestoreDumpPrefix)
i := strings.Index(after, "-")
if i <= 0 {
continue // not the shape writeSafetyDump writes — leave it alone rather than guess
}
undos = append(undos, undo{name: e.Name(), stamp: after[:i]})
}
if len(undos) <= maxUndoCopiesPerApp {
return
}
sort.Slice(undos, func(a, b int) bool { return undos[a].stamp > undos[b].stamp }) // newest first
for _, u := range undos[maxUndoCopiesPerApp:] {
p := filepath.Join(dumpDir, u.name)
if err := os.Remove(p); err != nil {
m.logger.Printf("[WARN] [backup] %s: could not prune old undo copy %s: %v", stack, u.name, err)
continue
}
m.logger.Printf("[INFO] [backup] %s: pruned old undo copy %s (keeping the newest %d)", stack, u.name, maxUndoCopiesPerApp)
}
}
// RestoreHoldFor reports whether an app is being held stopped after a failed restore + failed
// rollback, and returns the customer-facing reason. Every start path consults this — the customer's
// button, the app-stop Recover() starter, and the boot reconciler — because a hold that only one
// path honours is not a hold.
//
// Nil settings ⇒ NOT held. That direction is deliberate and is the opposite of the usual fail-closed
// rule: with no settings there is no hold recorded, so refusing every start would strand every app
// on a misconfigured box. The write side logs loudly when it cannot persist (see
// holdAppAfterFailedRollback), which is where that case is caught.
func (m *Manager) RestoreHoldFor(stack string) (bool, string) {
if m == nil || m.settings == nil {
return false, ""
}
h, ok := m.settings.GetRestoreHold(stack)
if !ok {
return false, ""
}
when := h.At
if t, err := time.Parse(time.RFC3339, h.At); err == nil {
when = t.Format("2006-01-02 15:04")
}
return true, fmt.Sprintf("a(z) %s adatainak visszaállítása %s-kor megszakadt, és a korábbi állapotot sem sikerült visszatölteni. "+
"Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek tovább. Vedd fel velünk a kapcsolatot", stack, when)
}
// holdAppAfterFailedRollback records the R-379/R-380 hold and makes sure nothing restarts the app
// behind our back.
//
// OPERATOR RULING, 2026-08-22: when the replay fails AND the rollback fails, the app is HELD
// STOPPED rather than started. A running app on a half-written database lets the customer type into
// it, and that turns a recoverable state into a permanent one. The alternative — start it and mark
// it — was considered and declined.
//
// It ENDS the app-stop marker deliberately. The marker means "owed a restart"; a held app is not
// owed one, and leaving the marker active would have Recover() start the broken app at the next
// controller boot. The hold is the thing that persists, not the marker.
func (m *Manager) holdAppAfterFailedRollback(stack string, replayErr, rollbackErr error) {
if m.settings == nil {
m.logger.Printf("[ERROR] [offbox] %s: cannot persist the restore hold — no settings wired; the app is stopped but NOTHING will refuse a restart", stack)
return
}
h := settings.RestoreHold{
Stack: stack,
At: time.Now().UTC().Format(time.RFC3339),
}
if replayErr != nil {
h.ReplayError = replayErr.Error()
}
if rollbackErr != nil {
h.RollbackErr = rollbackErr.Error()
}
if err := m.settings.SetRestoreHold(h); err != nil {
m.logger.Printf("[ERROR] [offbox] %s: persisting the restore hold FAILED: %v — the app is stopped and unguarded", stack, err)
}
// The app is not owed a restart; it is deliberately held. See the doc comment.
if m.appStop != nil {
m.appStop.End()
}
if m.restoreHoldNotify != nil {
m.restoreHoldNotify(stack, replayErr, rollbackErr)
}
}
// SetRestoreHoldNotify wires the operator notification for a held app (cmd/controller/main.go).
func (m *Manager) SetRestoreHoldNotify(fn func(stack string, replayErr, rollbackErr error)) {
m.restoreHoldNotify = fn
}
// rollbackSafetyDump re-applies THIS RUN's undo set, database by database, and is the whole of
// R-379's fix: it is the same ImportDump call a person made by hand on 2026-08-22 to recover
// `docmost` and `bookstack` after a failed replay, moved into the product.
//
// It runs with the DB service still up (the replay's own window) and BEFORE any restart, so the
// app never observes the half-written state. An error here means the app cannot be trusted to run —
// see the hold in ReconstituteFromOffsite.
func (m *Manager) rollbackSafetyDump(ctx context.Context, stack string, set safetyDumpSet) error {
if len(set.Files) == 0 {
return nil
}
imp := m.rollbackImport
if imp == nil {
imp = func(ctx context.Context, db DiscoveredDB, path string) error {
return ImportDump(ctx, db, path, m.logger, m.isDebug())
}
}
for _, f := range set.Files {
if _, sErr := os.Stat(f.Path); sErr != nil {
return fmt.Errorf("a visszavonáshoz szükséges mentés nem található (%s): %w", filepath.Base(f.Path), sErr)
}
m.logger.Printf("[INFO] [offbox] %s: rolling back to the pre-restore state from %s", stack, filepath.Base(f.Path))
if err := imp(ctx, f.DB, f.Path); err != nil {
return fmt.Errorf("a korábbi állapot visszaállítása sikertelen (%s): %w", f.DB.ContainerName, err)
}
}
m.logger.Printf("[INFO] [offbox] %s: rollback complete — %d database(s) returned to the pre-restore state", stack, len(set.Files))
return nil
}
// dumpForSafety is the DumpOne seam for the safety dump (tests inject; nil → the real DumpOne).
@@ -325,10 +514,11 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string, ack
// --- THE UNDO, BEFORE THE ACT ---------------------------------------------------------------
// Taken while the stack is still UP (a stopped database cannot be dumped) and before a single
// byte is overwritten, so a failure here aborts with the live app completely untouched.
safety, err := m.writeSafetyDump(ctx, stack, liveNs)
safetySet, err := m.writeSafetyDump(ctx, stack, liveNs)
if err != nil {
return res, err
}
safety := safetySet.First()
res.SafetyDump = safety
hasDB := safety != ""
if hasDB {
@@ -436,10 +626,39 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string, ack
n, iErr := m.reimportDBDumpsFrom(ctx, stack, scratchDumpDir)
res.DBsReplayed = n
if iErr != nil {
if sErr := restartStack(); sErr != nil {
m.logger.Printf("[WARN] [offbox] %s: full start after failed replay also failed: %v", stack, sErr)
// --- R-379/R-380: PUT THE CUSTOMER'S OWN COPY BACK ---------------------------------
// Until v0.220.0 this branch restarted the app onto a HALF-WRITTEN database and named
// the undo file in the message. Measured 2026-08-22 on demo-hp: Postgres was left
// emptied and crash-looping; MariaDB was left partly applied while the app reported
// `health=healthy`. Both are the same failure — a half state — and the only difference
// was whether it looked broken. MariaDB's structural statements are not transactional,
// so no engine flag can prevent the half state; putting the undo back is what removes
// it. This is the same ImportDump call a person ran by hand that day to recover both
// apps, moved into the product.
//
// The rollback runs BEFORE any restart and with the DB service still up, so the app
// never observes the half state. The ORIGINAL replay error is never swallowed: it is
// logged here in full and named in the customer's sentence.
m.logger.Printf("[ERROR] [offbox] %s: database replay failed, rolling back to the pre-restore state: %v", stack, iErr)
if rbErr := m.rollbackSafetyDump(ctx, stack, safetySet); rbErr != nil {
// BOTH failed. Do NOT start the app: a running app on a half-written database lets
// the customer type into it and makes the damage permanent. Hold it instead —
// operator ruling, 2026-08-22.
m.logger.Printf("[ERROR] [offbox] %s: ROLLBACK ALSO FAILED (%v) — holding the app stopped; replay error was: %v", stack, rbErr, iErr)
m.holdAppAfterFailedRollback(stack, iErr, rbErr)
return res, fmt.Errorf("a(z) %s adatbázisának visszaállítása sikertelen, és a korábbi állapot visszatöltése sem sikerült. "+
"Az alkalmazást biztonsági okból LEÁLLÍTVA hagytuk, hogy az adatai ne sérüljenek tovább. "+
"Vedd fel velünk a kapcsolatot — a korábbi állapot mentése megvan: %s", stack, filepath.Base(safety))
}
return res, fmt.Errorf("az adatbázis visszaállítása sikertelen: %w — a korábbi állapot mentése megvan: %s", iErr, filepath.Base(safety))
res.RolledBack = true
if sErr := restartStack(); sErr != nil {
m.logger.Printf("[WARN] [offbox] %s: start after a successful rollback failed: %v", stack, sErr)
}
// Says BOTH things. A message that reported only the failure would leave the customer
// believing their data was gone when it is back — the omission of a GAIN is as
// misleading as the omission of a loss.
return res, fmt.Errorf("a(z) %s adatbázisának visszaállítása sikertelen — az adataid visszakerültek a visszaállítás előtti állapotba, "+
"az alkalmazás fut tovább. Ha újra megpróbálnád, előbb vedd fel velünk a kapcsolatot", stack)
}
}
if err := restartStack(); err != nil {
@@ -449,8 +668,11 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string, ack
m.logger.Printf("[WARN] [offbox] %s reconstituted but health check failed: %v", stack, err)
}
m.logger.Printf("[INFO] [offbox] reconstituted %s from snapshot %s: %d file(s) placed, %d DB dump(s) replayed, safety dump=%s, skewed=%v",
stack, id, res.FilesPlaced, res.DBsReplayed, filepath.Base(safety), res.Skewed)
// R-382: VolumesReplayed was set above and never printed, so the operator log said
// "0 file(s) placed, 1 DB dump(s) replayed" on a run that returned a 52 MB Postgres data
// directory — less informative than the customer's own flash, which already named the volumes.
m.logger.Printf("[INFO] [offbox] reconstituted %s from snapshot %s: %d file(s) placed, %d volume(s) replayed, %d DB dump(s) replayed, safety dump=%s, skewed=%v",
stack, id, res.FilesPlaced, res.VolumesReplayed, res.DBsReplayed, filepath.Base(safety), res.Skewed)
return res, nil
}