R-353/R-357/R-358/R-360: the restore tells the truth (v0.226.0)
gates / gates (push) Successful in 11s

Four defects on the restore surface, all proven on demo-hp during the 2026-08-21
backup-truth drill, all still in shipped code. They share one acceptance idea: a
restore surface must state what it actually did, and must refuse what it cannot
do.

VERSION NOTE. The task specifying this targeted v0.224.0 against baseline
f8c9390. Both were consumed earlier the same day by R-330 (0.224.0) and R-331
(0.225.0). Drift re-confirmed against live Gitea before the first edit, operator
authorised proceeding, every symbol the spec named re-verified present at the
real baseline e5eee50.

R-353 -- a restore that gave back nothing still said it worked.
RestoreFromRecoveryUnit returned only error, so the surface printed
"<app> visszaallitva (<snapshot>)." -- equally true of a run that returned an
entire dataset and one that returned nothing. The count already existed and was
discarded one line deep: restoreDockerVolumesFrom always returned it, the
wrapper threw it away. Now (UnitRestoreResult, error), carrying replayed counts
AND what the manifest LISTED, because zero-replayed has two causes that are
opposite news. Three cases, three sentences, and EVERY one is a claim about the
BACKUP, never about the app -- this path has no SafetyDump discriminator, and
07-backup-architecture 6.3 records that an absent dump says nothing about the
app (R-361 destroyed canonical .sql files for four months).

R-357 -- the destructive restore had no free-space gate. offbox_reconstitute.go
contained ZERO references to offboxFree; all three existing gates guard
non-destructive paths. The gate now sits before mapOffsiteRestorePaths,
writeSafetyDump and StopStack, so a refusal costs nothing. Position IS the fix,
which is why the test asserts StopStack was never called. No headroom multiplier
(matches PlaceOffsiteRestore; the x1.1 elsewhere predicts a download). Fail
closed on either probe <= 0 -- otherwise `free < need` with need==0 is FALSE and
an unmeasurable scratch sails through: a gate present and inert.

R-358 -- a failed download was offered as a good one. The gate answered "the
directory exists and is non-empty", which is exactly what a part-way restic run
leaves. Now a completion marker written 0600 atomically AFTER restic returns
nil, with any stale one cleared BEFORE it starts; both orders pinned by an AST
test because resticStep is not a seam. Both handlers refuse server-side: the
wizard flags control a button, and a hidden button is not a guard.

SCENARIO F ANSWERED, and worse than the question assumed: a unit-only scratch IS
reachable through the real flow, by the most ordinary route. "Ellenorzo
visszaallitas" (mode=unit, advertised non-destructive) writes the SAME directory
-- offboxRestoreScratchDir ignores `full` and --include limits what restic
extracts, never where -- so a customer who ran the SAFE restore was then offered
the destructive one over a unit-only copy. Filed R-396; the marker closes it.

R-360 -- the delete refused only while a BACKUP ran. IsRunning() is FALSE for the
whole of a verification restore; the five sibling handlers all use
restoreOpBlocked(). Its doc comment claimed it already did this, which is why
nobody looked -- corrected in place.

Red-proofs, each printing the pre-fix behaviour, in CHANGELOG and REPORT. The
first R-357 red-proof exposed a hollow test OF MY OWN and it is recorded rather
than quietly fixed: the fixture refused earlier at the placement stat pre-pass,
so `stops == 0` passed against the pre-fix code. Fixture corrected, assertions
reordered so a removed gate reports the outage rather than "no error returned".

Green gate clean: 28 packages, rc 0. All 12 controller gates OK.
This commit is contained in:
2026-08-30 19:31:31 +02:00
parent e5eee501b5
commit b8af72764d
18 changed files with 1412 additions and 53 deletions
+63 -16
View File
@@ -129,6 +129,34 @@ func hasReplayableDump(dumpDir string) bool {
return false
}
// UnitRestoreResult is what a local recovery-unit restore actually did, so the surface can STATE it
// rather than report a bare completion.
//
// It exists for the same reason OffsiteReconstituteResult does, and it is the same lesson arriving on
// the other path: on 2026-08-21 an opengist restore reported "Restore-from-unit completed" over a unit
// holding manifest.json and compose/ and nothing else, and no screen could have told the customer that
// no data had been returned (R-353).
//
// The Manifest* counts are carried BECAUSE zero-replayed has two causes and they are not the same
// fact. A unit that lists no dumps means THE BACKUP held no data. A unit that lists dumps none of which
// replayed means something is wrong and the customer's live data was left untouched. R-355 is the
// standing rule this obeys: a claim about the APP must never be inferred from a counter — and here it
// is not merely unproven but unprovable, because 07-backup-architecture §6.3 records that an app's
// canonical .sql could be absent from the unit for reasons that have nothing to do with whether the app
// has a database (R-361 destroyed exactly that file for four months).
type UnitRestoreResult struct {
// VolumesReplayed is how many named-volume tars were unpacked into live Docker volumes.
VolumesReplayed int
// DBsReplayed is how many .sql dumps were imported. Never inferred from the presence of a database
// service — only a completed import increments it.
DBsReplayed int
// ManifestVolumes is len(manifest.VolumeDumps): what the unit CLAIMS it captured. The gap between
// this and VolumesReplayed is the whole of Scenario C.
ManifestVolumes int
// ManifestDBs is len(manifest.DBDumps): the same claim for the database leg.
ManifestDBs int
}
// RestoreFromRecoveryUnit recreates an app from its on-drive recovery unit.
//
// It reads the unit manifest, takes the portable secrets from the UNIT and the rest from the guest's
@@ -140,15 +168,16 @@ func hasReplayableDump(dumpDir string) bool {
// D5: this no longer needs the guest. A restore with the guest's app.yaml absent succeeds, which is
// pinned by TestRestoreFromRecoveryUnitWithGuestAbsent — the withheld class is regenerated (O4) and
// only a data key missing from BOTH sources still refuses.
func (m *Manager) RestoreFromRecoveryUnit(stackName string) error {
func (m *Manager) RestoreFromRecoveryUnit(stackName string) (UnitRestoreResult, error) {
var res UnitRestoreResult
if m.stackProvider == nil {
return fmt.Errorf("stack provider not configured")
return res, fmt.Errorf("stack provider not configured")
}
m.mu.Lock()
if m.running {
m.mu.Unlock()
return fmt.Errorf("backup or restore already in progress")
return res, fmt.Errorf("backup or restore already in progress")
}
m.running = true
m.mu.Unlock()
@@ -160,7 +189,7 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) error {
drivePath := m.GetAppDrivePath(stackName)
if drivePath == "" || !filepath.IsAbs(drivePath) {
return fmt.Errorf("cannot determine drive path for %s", stackName)
return res, fmt.Errorf("cannot determine drive path for %s", stackName)
}
nsRoot := m.namespaceRoot(drivePath)
@@ -170,9 +199,18 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) error {
m.mu.Lock()
m.running = false // RestoreApp re-acquires the running flag
m.mu.Unlock()
return m.RestoreApp(stackName, "")
// The fallback path has no unit and therefore no manifest to count against: a ZERO result is
// the honest answer, not a missing one. The surface must be able to tell "nothing came back"
// from "we never looked", and it can — ManifestVolumes/ManifestDBs are zero too, which is
// Scenario B's shape and reads as "the backup held no data", which is exactly true of a box
// with no recovery unit.
return res, m.RestoreApp(stackName, "")
}
// R-353: what the unit CLAIMS it holds, recorded before any mutation. Read from the manifest that
// was just parsed above, so the claim and the outcome are counted from the same document.
res.ManifestVolumes, res.ManifestDBs = len(manifest.VolumeDumps), len(manifest.DBDumps)
composeDir := RecoveryUnitComposePath(nsRoot, stackName)
nonSecretEnv, unitSecrets := readUnitEnv(filepath.Join(composeDir, "app.yaml"), manifest.PortableSecretEnvVars)
@@ -185,7 +223,7 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) error {
fullEnv, missing, err := reconcileRestoreSecrets(nonSecretEnv, unitSecrets, guestSecrets, manifest.SecretEnvVars, manifest.DataKeyEnvVars)
if err != nil {
m.logger.Printf("[ERROR] [backup] Restore REFUSED for %s: %v", stackName, err)
return err
return res, err
}
// O4: a missing RESETTABLE secret used to redeploy blank (compose "Defaulting to a blank
// string" → exit 1). Generate a replacement via the deploy flow's generator instead —
@@ -242,7 +280,7 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) error {
hasDumps := hasReplayableDump(AppDBDumpPath(nsRoot, stackName))
if hasDumps && len(dbServices) == 0 {
m.logger.Printf("[ERROR] [backup] Restore REFUSED for %s: a .sql dump exists but no database service is identifiable in the unit's compose", stackName)
return fmt.Errorf("Az adatbázis-szolgáltatás nem azonosítható a(z) %s alkalmazásban — a visszaállítás biztonsági okból nem indult el.", stackName)
return res, fmt.Errorf("Az adatbázis-szolgáltatás nem azonosítható a(z) %s alkalmazásban — a visszaállítás biztonsági okból nem indult el.", stackName)
}
// Stop, restore named-volume data, recreate the definition, replay the DB with ONLY the database
@@ -252,12 +290,17 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) error {
if err := m.stackProvider.StopStack(stackName); err != nil {
m.logger.Printf("[WARN] [backup] could not stop %s before restore: %v (continuing)", stackName, err)
}
if err := m.restoreDockerVolumes(stackName, drivePath); err != nil {
m.logger.Printf("[ERROR] [backup] volume restore for %s: %v", stackName, err)
dataErr = err
// R-353: the count is captured even when the replay errors — a partial replay is a fact the
// customer's sentence has to be built from, and discarding it on the error path is how Scenario C
// would end up wearing Scenario B's wording.
replayed, volErr := m.restoreDockerVolumes(stackName, drivePath)
res.VolumesReplayed = replayed
if volErr != nil {
m.logger.Printf("[ERROR] [backup] volume restore for %s: %v", stackName, volErr)
dataErr = volErr
}
if err := m.stackProvider.RecreateStackDefinitionFromUnit(stackName, composeDir, fullEnv); err != nil {
return fmt.Errorf("recreating %s from unit: %w", stackName, err)
return res, fmt.Errorf("recreating %s from unit: %w", stackName, err)
}
// F17: the captured .sql dump is the authoritative logical DB state — replay it AFTER the volume
// restore, so the dump WINS over any volume-tar copy of the database.
@@ -270,23 +313,27 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) error {
if dataErr == nil {
dataErr = err
}
} else if _, err := m.reimportDBDumpsCtx(stackName, nsRoot); err != nil {
} else if n, err := m.reimportDBDumpsCtx(stackName, nsRoot); err != nil {
res.DBsReplayed = n // partial credit: whatever imported before the failure really did import
m.logger.Printf("[ERROR] [backup] DB re-import for %s: %v", stackName, err)
if dataErr == nil {
dataErr = err
}
} else {
res.DBsReplayed = n
}
}
if err := m.stackProvider.StartStack(stackName); err != nil {
return fmt.Errorf("starting %s after restore from unit: %w", stackName, err)
return res, fmt.Errorf("starting %s after restore from unit: %w", stackName, err)
}
if err := m.waitForHealthy(stackName, 90*time.Second); err != nil {
m.logger.Printf("[WARN] [backup] %s restored but health check failed: %v", stackName, err)
}
if dataErr != nil {
return fmt.Errorf("restore of %s from unit completed with data errors: %w", stackName, dataErr)
return res, fmt.Errorf("restore of %s from unit completed with data errors: %w", stackName, dataErr)
}
m.logger.Printf("[INFO] [backup] Restore-from-unit completed: %s", stackName)
return nil
m.logger.Printf("[INFO] [backup] Restore-from-unit completed: %s — %d volume(s) of %d listed, %d database(s) of %d listed",
stackName, res.VolumesReplayed, res.ManifestVolumes, res.DBsReplayed, res.ManifestDBs)
return res, nil
}