R-379/R-380: put the customer's undo copy back when a database restore fails
gates / gates (push) Successful in 11s
gates / gates (push) Successful in 11s
R-379 and R-380 were one failure. Both ended with a half-restored database; the only difference was whether it looked broken. Postgres emptied and crash-looped; MariaDB applied part of the dump and reported health=healthy with a zero-row schema-version table. Measured live on demo-hp 2026-08-22. The undo copy was already taken and already good - proven by hand that day on both engines. Nothing in the product could apply it. Now it does, with the same ImportDump call, before any restart and inside the DB-only window. The WHOLE undo set, matched on this run's stamp. writeSafetyDump returned one path for an app with two databases; a rollback on that would restore one and leave the other half-written. When the rollback also fails the app is HELD STOPPED (operator ruling): a running app on a half-written database lets the customer make the damage permanent. Every start path refuses it - customer button, appstop Recover, boot sweep - via the shared driveStartGate, checked ABOVE its driveless early return because these apps have no drive. The marker is ended so nothing auto-restarts it. The row goes red. Cleared with --clear-restore-hold, an operator CLI route. --single-transaction is a belt on Postgres only; MariaDB DDL is not transactional and that is why the rollback is the fix. R-381: the engine's stderr stops reaching the customer (615 bytes on MariaDB, its middle rows out of their own database) and starts reaching the operator log, which never had it. R-382: the summary log prints the volume count it already held. Undo copies resolve to their own app, are marked IsUndo, and are capped at 3 per app, pruned from the capture side. The reported render-as-an-app symptom did NOT reproduce - the live page was read first and had zero occurrences. Tests 1468 -> 1483. Eight red-proofs; ONE PASSED and is reported: the R-381 behavioural test injected below ImportDump. A guard at that layer now convicts.
This commit is contained in:
@@ -164,6 +164,20 @@ type Settings struct {
|
||||
// Storage paths registry
|
||||
StoragePaths []StoragePath `json:"storage_paths,omitempty"`
|
||||
|
||||
// RestoreHolds (R-379/R-380, v0.220.0) records apps this controller is DELIBERATELY keeping
|
||||
// stopped because a database restore failed AND the rollback to the customer's own pre-restore
|
||||
// copy also failed. Keyed by stack name.
|
||||
//
|
||||
// WHY THIS AND NOT `AppConfig.DesiredState`: that field is the CUSTOMER'S stated intent — what
|
||||
// they asked for. Writing our own failure into it would make our fault indistinguishable from
|
||||
// their choice, which is the exact confusion its own comment says it exists to end. The fenced
|
||||
// act is "writing DesiredState from the restore path"; reading it stays fine.
|
||||
//
|
||||
// WHY NOT the app-stop marker: that marker means "owed a restart". A held app is not owed one —
|
||||
// leaving the marker active would have Recover() start the broken app at the next controller
|
||||
// boot, hours later and quietly, which is the outcome the hold exists to prevent.
|
||||
RestoreHolds map[string]RestoreHold `json:"restore_holds,omitempty"`
|
||||
|
||||
// Cross-drive restic repo password (auto-generated on first use)
|
||||
CrossDriveResticPassword string `json:"cross_drive_restic_password,omitempty"`
|
||||
|
||||
@@ -1515,6 +1529,65 @@ func InferStorageLabel(path string) string {
|
||||
return fmt.Sprintf("Tárhely (%s)", base)
|
||||
}
|
||||
|
||||
// RestoreHold is one app held stopped after a failed restore whose rollback also failed. It carries
|
||||
// the reason so every refusal can NAME it — a refusal without a reason and a route is how a customer
|
||||
// is left with nothing to do.
|
||||
type RestoreHold struct {
|
||||
Stack string `json:"stack"`
|
||||
At string `json:"at"` // RFC3339 UTC
|
||||
ReplayError string `json:"replay_error,omitempty"` // what the restore hit
|
||||
RollbackErr string `json:"rollback_error,omitempty"` // what the rollback then hit
|
||||
SafetyDump string `json:"safety_dump,omitempty"` // basename of the undo copy that could not be applied
|
||||
}
|
||||
|
||||
// SetRestoreHold records a hold. Modelled on SetDisconnected: a condition, plus what it is holding.
|
||||
func (s *Settings) SetRestoreHold(h RestoreHold) error {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
if s.RestoreHolds == nil {
|
||||
s.RestoreHolds = map[string]RestoreHold{}
|
||||
}
|
||||
s.RestoreHolds[h.Stack] = h
|
||||
if s.log != nil {
|
||||
s.log.Printf("[WARN] [settings] restore hold SET for %s — the app stays stopped until it is cleared", h.Stack)
|
||||
}
|
||||
return s.save()
|
||||
}
|
||||
|
||||
// GetRestoreHold returns the hold for a stack, if one is in force.
|
||||
func (s *Settings) GetRestoreHold(stack string) (RestoreHold, bool) {
|
||||
s.mu.RLock()
|
||||
defer s.mu.RUnlock()
|
||||
h, ok := s.RestoreHolds[stack]
|
||||
return h, ok
|
||||
}
|
||||
|
||||
// ListRestoreHolds returns every hold in force.
|
||||
func (s *Settings) ListRestoreHolds() []RestoreHold {
|
||||
s.mu.RLock()
|
||||
defer s.mu.RUnlock()
|
||||
out := make([]RestoreHold, 0, len(s.RestoreHolds))
|
||||
for _, h := range s.RestoreHolds {
|
||||
out = append(out, h)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// ClearRestoreHold removes a hold. Returns false when there was none, so a caller can tell "cleared"
|
||||
// from "there was nothing to clear" rather than reporting success either way.
|
||||
func (s *Settings) ClearRestoreHold(stack string) (bool, error) {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
if _, ok := s.RestoreHolds[stack]; !ok {
|
||||
return false, nil
|
||||
}
|
||||
delete(s.RestoreHolds, stack)
|
||||
if s.log != nil {
|
||||
s.log.Printf("[INFO] [settings] restore hold CLEARED for %s", stack)
|
||||
}
|
||||
return true, s.save()
|
||||
}
|
||||
|
||||
// SetDisconnected marks a storage path as disconnected (or connected) and records which stacks were stopped.
|
||||
func (s *Settings) SetDisconnected(path string, disconnected bool, stoppedStacks []string) error {
|
||||
s.mu.Lock()
|
||||
|
||||
Reference in New Issue
Block a user