R-353/R-357/R-358/R-360: the restore tells the truth (v0.226.0)
gates / gates (push) Successful in 11s
gates / gates (push) Successful in 11s
Four defects on the restore surface, all proven on demo-hp during the 2026-08-21 backup-truth drill, all still in shipped code. They share one acceptance idea: a restore surface must state what it actually did, and must refuse what it cannot do. VERSION NOTE. The task specifying this targeted v0.224.0 against baselinef8c9390. Both were consumed earlier the same day by R-330 (0.224.0) and R-331 (0.225.0). Drift re-confirmed against live Gitea before the first edit, operator authorised proceeding, every symbol the spec named re-verified present at the real baselinee5eee50. R-353 -- a restore that gave back nothing still said it worked. RestoreFromRecoveryUnit returned only error, so the surface printed "<app> visszaallitva (<snapshot>)." -- equally true of a run that returned an entire dataset and one that returned nothing. The count already existed and was discarded one line deep: restoreDockerVolumesFrom always returned it, the wrapper threw it away. Now (UnitRestoreResult, error), carrying replayed counts AND what the manifest LISTED, because zero-replayed has two causes that are opposite news. Three cases, three sentences, and EVERY one is a claim about the BACKUP, never about the app -- this path has no SafetyDump discriminator, and 07-backup-architecture 6.3 records that an absent dump says nothing about the app (R-361 destroyed canonical .sql files for four months). R-357 -- the destructive restore had no free-space gate. offbox_reconstitute.go contained ZERO references to offboxFree; all three existing gates guard non-destructive paths. The gate now sits before mapOffsiteRestorePaths, writeSafetyDump and StopStack, so a refusal costs nothing. Position IS the fix, which is why the test asserts StopStack was never called. No headroom multiplier (matches PlaceOffsiteRestore; the x1.1 elsewhere predicts a download). Fail closed on either probe <= 0 -- otherwise `free < need` with need==0 is FALSE and an unmeasurable scratch sails through: a gate present and inert. R-358 -- a failed download was offered as a good one. The gate answered "the directory exists and is non-empty", which is exactly what a part-way restic run leaves. Now a completion marker written 0600 atomically AFTER restic returns nil, with any stale one cleared BEFORE it starts; both orders pinned by an AST test because resticStep is not a seam. Both handlers refuse server-side: the wizard flags control a button, and a hidden button is not a guard. SCENARIO F ANSWERED, and worse than the question assumed: a unit-only scratch IS reachable through the real flow, by the most ordinary route. "Ellenorzo visszaallitas" (mode=unit, advertised non-destructive) writes the SAME directory -- offboxRestoreScratchDir ignores `full` and --include limits what restic extracts, never where -- so a customer who ran the SAFE restore was then offered the destructive one over a unit-only copy. Filed R-396; the marker closes it. R-360 -- the delete refused only while a BACKUP ran. IsRunning() is FALSE for the whole of a verification restore; the five sibling handlers all use restoreOpBlocked(). Its doc comment claimed it already did this, which is why nobody looked -- corrected in place. Red-proofs, each printing the pre-fix behaviour, in CHANGELOG and REPORT. The first R-357 red-proof exposed a hollow test OF MY OWN and it is recorded rather than quietly fixed: the fixture refused earlier at the placement stat pre-pass, so `stops == 0` passed against the pre-fix code. Fixture corrected, assertions reordered so a removed gate reports the outage rather than "no error returned". Green gate clean: 28 packages, rc 0. All 12 controller gates OK.
This commit is contained in:
@@ -129,6 +129,34 @@ func hasReplayableDump(dumpDir string) bool {
|
||||
return false
|
||||
}
|
||||
|
||||
// UnitRestoreResult is what a local recovery-unit restore actually did, so the surface can STATE it
|
||||
// rather than report a bare completion.
|
||||
//
|
||||
// It exists for the same reason OffsiteReconstituteResult does, and it is the same lesson arriving on
|
||||
// the other path: on 2026-08-21 an opengist restore reported "Restore-from-unit completed" over a unit
|
||||
// holding manifest.json and compose/ and nothing else, and no screen could have told the customer that
|
||||
// no data had been returned (R-353).
|
||||
//
|
||||
// The Manifest* counts are carried BECAUSE zero-replayed has two causes and they are not the same
|
||||
// fact. A unit that lists no dumps means THE BACKUP held no data. A unit that lists dumps none of which
|
||||
// replayed means something is wrong and the customer's live data was left untouched. R-355 is the
|
||||
// standing rule this obeys: a claim about the APP must never be inferred from a counter — and here it
|
||||
// is not merely unproven but unprovable, because 07-backup-architecture §6.3 records that an app's
|
||||
// canonical .sql could be absent from the unit for reasons that have nothing to do with whether the app
|
||||
// has a database (R-361 destroyed exactly that file for four months).
|
||||
type UnitRestoreResult struct {
|
||||
// VolumesReplayed is how many named-volume tars were unpacked into live Docker volumes.
|
||||
VolumesReplayed int
|
||||
// DBsReplayed is how many .sql dumps were imported. Never inferred from the presence of a database
|
||||
// service — only a completed import increments it.
|
||||
DBsReplayed int
|
||||
// ManifestVolumes is len(manifest.VolumeDumps): what the unit CLAIMS it captured. The gap between
|
||||
// this and VolumesReplayed is the whole of Scenario C.
|
||||
ManifestVolumes int
|
||||
// ManifestDBs is len(manifest.DBDumps): the same claim for the database leg.
|
||||
ManifestDBs int
|
||||
}
|
||||
|
||||
// RestoreFromRecoveryUnit recreates an app from its on-drive recovery unit.
|
||||
//
|
||||
// It reads the unit manifest, takes the portable secrets from the UNIT and the rest from the guest's
|
||||
@@ -140,15 +168,16 @@ func hasReplayableDump(dumpDir string) bool {
|
||||
// D5: this no longer needs the guest. A restore with the guest's app.yaml absent succeeds, which is
|
||||
// pinned by TestRestoreFromRecoveryUnitWithGuestAbsent — the withheld class is regenerated (O4) and
|
||||
// only a data key missing from BOTH sources still refuses.
|
||||
func (m *Manager) RestoreFromRecoveryUnit(stackName string) error {
|
||||
func (m *Manager) RestoreFromRecoveryUnit(stackName string) (UnitRestoreResult, error) {
|
||||
var res UnitRestoreResult
|
||||
if m.stackProvider == nil {
|
||||
return fmt.Errorf("stack provider not configured")
|
||||
return res, fmt.Errorf("stack provider not configured")
|
||||
}
|
||||
|
||||
m.mu.Lock()
|
||||
if m.running {
|
||||
m.mu.Unlock()
|
||||
return fmt.Errorf("backup or restore already in progress")
|
||||
return res, fmt.Errorf("backup or restore already in progress")
|
||||
}
|
||||
m.running = true
|
||||
m.mu.Unlock()
|
||||
@@ -160,7 +189,7 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) error {
|
||||
|
||||
drivePath := m.GetAppDrivePath(stackName)
|
||||
if drivePath == "" || !filepath.IsAbs(drivePath) {
|
||||
return fmt.Errorf("cannot determine drive path for %s", stackName)
|
||||
return res, fmt.Errorf("cannot determine drive path for %s", stackName)
|
||||
}
|
||||
nsRoot := m.namespaceRoot(drivePath)
|
||||
|
||||
@@ -170,9 +199,18 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) error {
|
||||
m.mu.Lock()
|
||||
m.running = false // RestoreApp re-acquires the running flag
|
||||
m.mu.Unlock()
|
||||
return m.RestoreApp(stackName, "")
|
||||
// The fallback path has no unit and therefore no manifest to count against: a ZERO result is
|
||||
// the honest answer, not a missing one. The surface must be able to tell "nothing came back"
|
||||
// from "we never looked", and it can — ManifestVolumes/ManifestDBs are zero too, which is
|
||||
// Scenario B's shape and reads as "the backup held no data", which is exactly true of a box
|
||||
// with no recovery unit.
|
||||
return res, m.RestoreApp(stackName, "")
|
||||
}
|
||||
|
||||
// R-353: what the unit CLAIMS it holds, recorded before any mutation. Read from the manifest that
|
||||
// was just parsed above, so the claim and the outcome are counted from the same document.
|
||||
res.ManifestVolumes, res.ManifestDBs = len(manifest.VolumeDumps), len(manifest.DBDumps)
|
||||
|
||||
composeDir := RecoveryUnitComposePath(nsRoot, stackName)
|
||||
nonSecretEnv, unitSecrets := readUnitEnv(filepath.Join(composeDir, "app.yaml"), manifest.PortableSecretEnvVars)
|
||||
|
||||
@@ -185,7 +223,7 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) error {
|
||||
fullEnv, missing, err := reconcileRestoreSecrets(nonSecretEnv, unitSecrets, guestSecrets, manifest.SecretEnvVars, manifest.DataKeyEnvVars)
|
||||
if err != nil {
|
||||
m.logger.Printf("[ERROR] [backup] Restore REFUSED for %s: %v", stackName, err)
|
||||
return err
|
||||
return res, err
|
||||
}
|
||||
// O4: a missing RESETTABLE secret used to redeploy blank (compose "Defaulting to a blank
|
||||
// string" → exit 1). Generate a replacement via the deploy flow's generator instead —
|
||||
@@ -242,7 +280,7 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) error {
|
||||
hasDumps := hasReplayableDump(AppDBDumpPath(nsRoot, stackName))
|
||||
if hasDumps && len(dbServices) == 0 {
|
||||
m.logger.Printf("[ERROR] [backup] Restore REFUSED for %s: a .sql dump exists but no database service is identifiable in the unit's compose", stackName)
|
||||
return fmt.Errorf("Az adatbázis-szolgáltatás nem azonosítható a(z) %s alkalmazásban — a visszaállítás biztonsági okból nem indult el.", stackName)
|
||||
return res, fmt.Errorf("Az adatbázis-szolgáltatás nem azonosítható a(z) %s alkalmazásban — a visszaállítás biztonsági okból nem indult el.", stackName)
|
||||
}
|
||||
|
||||
// Stop, restore named-volume data, recreate the definition, replay the DB with ONLY the database
|
||||
@@ -252,12 +290,17 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) error {
|
||||
if err := m.stackProvider.StopStack(stackName); err != nil {
|
||||
m.logger.Printf("[WARN] [backup] could not stop %s before restore: %v (continuing)", stackName, err)
|
||||
}
|
||||
if err := m.restoreDockerVolumes(stackName, drivePath); err != nil {
|
||||
m.logger.Printf("[ERROR] [backup] volume restore for %s: %v", stackName, err)
|
||||
dataErr = err
|
||||
// R-353: the count is captured even when the replay errors — a partial replay is a fact the
|
||||
// customer's sentence has to be built from, and discarding it on the error path is how Scenario C
|
||||
// would end up wearing Scenario B's wording.
|
||||
replayed, volErr := m.restoreDockerVolumes(stackName, drivePath)
|
||||
res.VolumesReplayed = replayed
|
||||
if volErr != nil {
|
||||
m.logger.Printf("[ERROR] [backup] volume restore for %s: %v", stackName, volErr)
|
||||
dataErr = volErr
|
||||
}
|
||||
if err := m.stackProvider.RecreateStackDefinitionFromUnit(stackName, composeDir, fullEnv); err != nil {
|
||||
return fmt.Errorf("recreating %s from unit: %w", stackName, err)
|
||||
return res, fmt.Errorf("recreating %s from unit: %w", stackName, err)
|
||||
}
|
||||
// F17: the captured .sql dump is the authoritative logical DB state — replay it AFTER the volume
|
||||
// restore, so the dump WINS over any volume-tar copy of the database.
|
||||
@@ -270,23 +313,27 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) error {
|
||||
if dataErr == nil {
|
||||
dataErr = err
|
||||
}
|
||||
} else if _, err := m.reimportDBDumpsCtx(stackName, nsRoot); err != nil {
|
||||
} else if n, err := m.reimportDBDumpsCtx(stackName, nsRoot); err != nil {
|
||||
res.DBsReplayed = n // partial credit: whatever imported before the failure really did import
|
||||
m.logger.Printf("[ERROR] [backup] DB re-import for %s: %v", stackName, err)
|
||||
if dataErr == nil {
|
||||
dataErr = err
|
||||
}
|
||||
} else {
|
||||
res.DBsReplayed = n
|
||||
}
|
||||
}
|
||||
if err := m.stackProvider.StartStack(stackName); err != nil {
|
||||
return fmt.Errorf("starting %s after restore from unit: %w", stackName, err)
|
||||
return res, fmt.Errorf("starting %s after restore from unit: %w", stackName, err)
|
||||
}
|
||||
if err := m.waitForHealthy(stackName, 90*time.Second); err != nil {
|
||||
m.logger.Printf("[WARN] [backup] %s restored but health check failed: %v", stackName, err)
|
||||
}
|
||||
|
||||
if dataErr != nil {
|
||||
return fmt.Errorf("restore of %s from unit completed with data errors: %w", stackName, dataErr)
|
||||
return res, fmt.Errorf("restore of %s from unit completed with data errors: %w", stackName, dataErr)
|
||||
}
|
||||
m.logger.Printf("[INFO] [backup] Restore-from-unit completed: %s", stackName)
|
||||
return nil
|
||||
m.logger.Printf("[INFO] [backup] Restore-from-unit completed: %s — %d volume(s) of %d listed, %d database(s) of %d listed",
|
||||
stackName, res.VolumesReplayed, res.ManifestVolumes, res.DBsReplayed, res.ManifestDBs)
|
||||
return res, nil
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user