R-379/R-380: put the customer's undo copy back when a database restore fails
gates / gates (push) Successful in 11s

R-379 and R-380 were one failure. Both ended with a half-restored database; the
only difference was whether it looked broken. Postgres emptied and crash-looped;
MariaDB applied part of the dump and reported health=healthy with a zero-row
schema-version table. Measured live on demo-hp 2026-08-22.

The undo copy was already taken and already good - proven by hand that day on
both engines. Nothing in the product could apply it. Now it does, with the same
ImportDump call, before any restart and inside the DB-only window.

The WHOLE undo set, matched on this run's stamp. writeSafetyDump returned one
path for an app with two databases; a rollback on that would restore one and
leave the other half-written.

When the rollback also fails the app is HELD STOPPED (operator ruling): a running
app on a half-written database lets the customer make the damage permanent. Every
start path refuses it - customer button, appstop Recover, boot sweep - via the
shared driveStartGate, checked ABOVE its driveless early return because these
apps have no drive. The marker is ended so nothing auto-restarts it. The row goes
red. Cleared with --clear-restore-hold, an operator CLI route.

--single-transaction is a belt on Postgres only; MariaDB DDL is not transactional
and that is why the rollback is the fix.

R-381: the engine's stderr stops reaching the customer (615 bytes on MariaDB, its
middle rows out of their own database) and starts reaching the operator log,
which never had it.
R-382: the summary log prints the volume count it already held.
Undo copies resolve to their own app, are marked IsUndo, and are capped at 3 per
app, pruned from the capture side. The reported render-as-an-app symptom did NOT
reproduce - the live page was read first and had zero occurrences.

Tests 1468 -> 1483. Eight red-proofs; ONE PASSED and is reported: the R-381
behavioural test injected below ImportDump. A guard at that layer now convicts.
This commit is contained in:
2026-08-22 18:03:18 +02:00
parent 0f3cf0dbb2
commit 2c724c9283
15 changed files with 1288 additions and 37 deletions
+73
View File
@@ -164,6 +164,20 @@ type Settings struct {
// Storage paths registry
StoragePaths []StoragePath `json:"storage_paths,omitempty"`
// RestoreHolds (R-379/R-380, v0.220.0) records apps this controller is DELIBERATELY keeping
// stopped because a database restore failed AND the rollback to the customer's own pre-restore
// copy also failed. Keyed by stack name.
//
// WHY THIS AND NOT `AppConfig.DesiredState`: that field is the CUSTOMER'S stated intent — what
// they asked for. Writing our own failure into it would make our fault indistinguishable from
// their choice, which is the exact confusion its own comment says it exists to end. The fenced
// act is "writing DesiredState from the restore path"; reading it stays fine.
//
// WHY NOT the app-stop marker: that marker means "owed a restart". A held app is not owed one —
// leaving the marker active would have Recover() start the broken app at the next controller
// boot, hours later and quietly, which is the outcome the hold exists to prevent.
RestoreHolds map[string]RestoreHold `json:"restore_holds,omitempty"`
// Cross-drive restic repo password (auto-generated on first use)
CrossDriveResticPassword string `json:"cross_drive_restic_password,omitempty"`
@@ -1515,6 +1529,65 @@ func InferStorageLabel(path string) string {
return fmt.Sprintf("Tárhely (%s)", base)
}
// RestoreHold is one app held stopped after a failed restore whose rollback also failed. It carries
// the reason so every refusal can NAME it — a refusal without a reason and a route is how a customer
// is left with nothing to do.
type RestoreHold struct {
Stack string `json:"stack"`
At string `json:"at"` // RFC3339 UTC
ReplayError string `json:"replay_error,omitempty"` // what the restore hit
RollbackErr string `json:"rollback_error,omitempty"` // what the rollback then hit
SafetyDump string `json:"safety_dump,omitempty"` // basename of the undo copy that could not be applied
}
// SetRestoreHold records a hold. Modelled on SetDisconnected: a condition, plus what it is holding.
func (s *Settings) SetRestoreHold(h RestoreHold) error {
s.mu.Lock()
defer s.mu.Unlock()
if s.RestoreHolds == nil {
s.RestoreHolds = map[string]RestoreHold{}
}
s.RestoreHolds[h.Stack] = h
if s.log != nil {
s.log.Printf("[WARN] [settings] restore hold SET for %s — the app stays stopped until it is cleared", h.Stack)
}
return s.save()
}
// GetRestoreHold returns the hold for a stack, if one is in force.
func (s *Settings) GetRestoreHold(stack string) (RestoreHold, bool) {
s.mu.RLock()
defer s.mu.RUnlock()
h, ok := s.RestoreHolds[stack]
return h, ok
}
// ListRestoreHolds returns every hold in force.
func (s *Settings) ListRestoreHolds() []RestoreHold {
s.mu.RLock()
defer s.mu.RUnlock()
out := make([]RestoreHold, 0, len(s.RestoreHolds))
for _, h := range s.RestoreHolds {
out = append(out, h)
}
return out
}
// ClearRestoreHold removes a hold. Returns false when there was none, so a caller can tell "cleared"
// from "there was nothing to clear" rather than reporting success either way.
func (s *Settings) ClearRestoreHold(stack string) (bool, error) {
s.mu.Lock()
defer s.mu.Unlock()
if _, ok := s.RestoreHolds[stack]; !ok {
return false, nil
}
delete(s.RestoreHolds, stack)
if s.log != nil {
s.log.Printf("[INFO] [settings] restore hold CLEARED for %s", stack)
}
return true, s.save()
}
// SetDisconnected marks a storage path as disconnected (or connected) and records which stacks were stopped.
func (s *Settings) SetDisconnected(path string, disconnected bool, stoppedStacks []string) error {
s.mu.Lock()