R-102: the recovery unit on the second drive becomes a way back

Tier-2 mirrors each app's whole recovery unit to <dest>/backups/secondary/<app>/recovery-unit/ on
every run and has done for months. Nothing read it. In the one failure Tier-2 exists for - the
primary drive is lost, and the primary unit with it - the surviving copy could not be opened by any
action in the product (07-backup-architecture 6.3, 7.2).

Part 1.2: RestoreFromRecoveryUnitAt(stack, unitDir) holds the whole body; RestoreFromRecoveryUnit is
the thin caller naming the primary unit. ONE implementation, two callers. The SOURCE moves; the
DESTINATION does not - live Docker volumes, the live database container, the guest's definition, all
unchanged. The R-47 mutation order, the secret reconciliation with unit-over-guest precedence, the
fail-closed data-key gate and the no-unit fallback with CountsUnknown are untouched.
reimportDBDumpsAtCtx is the bounded-context twin of reimportDBDumpsFrom; the 35-minute bound is now
named once so the two paths cannot drift. The R-354 volume-replay seam is reused rather than a second
one invented, which is what lets the acceptance test assert the volume leg's source directory.

Part 1.3: RestoreTier2Unit resolves the recorded copy, refuses fail-closed unless the mirror carries a
parseable manifest - a directory is not a package - and delegates. The single-writer flag is taken
inside RestoreFromRecoveryUnitAt, not beside it.

Part 2.1: Tier2Coverage gains UnitRestorable and the copy's dates. CanRestore() is NOT widened; it
still answers only 'can the file restore run?'. One predicate answering two questions is R-356, which
refused 40 running apps for months.

Tests: A2-A6 and B1-B5, plus two non-regression guards. The Tier-2 fixtures build their mirror with
the production RunTier2, so the claim is 'the copy Tier-2 writes is the copy this restore reads'.
Red-proofs: A5 (swap volumes/recreate -> fails on the order), B2 (point the reader back at the
primary -> fails with the mirror never reaching the redeploy, and with permission denied once the
primary tree is unreadable).
This commit is contained in:
2026-08-31 11:30:34 +02:00
parent c732006d26
commit 0f9b796615
5 changed files with 856 additions and 16 deletions
+55 -9
View File
@@ -178,7 +178,36 @@ type UnitRestoreResult struct {
// D5: this no longer needs the guest. A restore with the guest's app.yaml absent succeeds, which is
// pinned by TestRestoreFromRecoveryUnitWithGuestAbsent — the withheld class is regenerated (O4) and
// only a data key missing from BOTH sources still refuses.
//
// R-102: this is now the thin caller. It names the PRIMARY unit — the app's own drive,
// backups/primary/<stack> — and hands it to RestoreFromRecoveryUnitAt, which holds the whole body.
// An unresolvable drive path is still refused inside …At, in the same place and with the same
// message, so the order of the checks a caller can observe is unchanged.
func (m *Manager) RestoreFromRecoveryUnit(stackName string) (UnitRestoreResult, error) {
return m.RestoreFromRecoveryUnitAt(stackName, RecoveryUnitPath(m.namespaceRoot(m.GetAppDrivePath(stackName)), stackName))
}
// RestoreFromRecoveryUnitAt is RestoreFromRecoveryUnit with an EXPLICIT recovery-unit directory.
//
// R-102. ONE implementation, two callers — the same rule restoreDockerVolumesFrom states beside
// itself in restore.go, and for the same reason: a second copy of this body is exactly how the local
// path and the off-site path drifted apart until nothing compared them.
//
// The reason it exists: Tier-2 mirrors the app's whole recovery unit to
// <dest>/backups/secondary/<stack>/recovery-unit/ on every run, and until now every reader of a unit
// could only name a path under backups/primary/. So in the one failure Tier-2 exists for — the
// primary drive is lost, taking the primary unit with it — the surviving copy could not be opened by
// any action in the product (07-backup-architecture §6.3, §7.2).
//
// THE SOURCE MOVES; THE DESTINATION DOES NOT. unitDir changes only where the manifest, the compose
// capture, the .sql dumps and the volume tars are READ from. The app's data is written back to the
// live Docker volumes and the live database container, and its definition to the guest, exactly as
// before — a restore that also relocated the app's data would be a migration, not a restore.
//
// Everything else is pinned and unchanged: the mutation order (stop → volumes → recreate → DB-only
// start → replay → start, R-47), the secret reconciliation with unit-over-guest precedence and the
// fail-closed data-key gate, and the no-unit fallback to RestoreApp with its CountsUnknown handling.
func (m *Manager) RestoreFromRecoveryUnitAt(stackName, unitDir string) (UnitRestoreResult, error) {
var res UnitRestoreResult
if m.stackProvider == nil {
return res, fmt.Errorf("stack provider not configured")
@@ -197,15 +226,18 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) (UnitRestoreResult,
m.mu.Unlock()
}()
// The DESTINATION side, and it is deliberately still resolved here: RestoreApp (the no-unit
// fallback below) needs it, and 07-backup-architecture §6.3 records that the restore destination
// is resolved by the same rule as the capture destination. It is no longer used to derive any
// SOURCE path — that is what unitDir is for.
drivePath := m.GetAppDrivePath(stackName)
if drivePath == "" || !filepath.IsAbs(drivePath) {
return res, fmt.Errorf("cannot determine drive path for %s", stackName)
}
nsRoot := m.namespaceRoot(drivePath)
manifest := readManifest(RecoveryUnitManifestPath(nsRoot, stackName))
manifest := readManifest(UnitManifestFile(unitDir))
if manifest == nil {
m.logger.Printf("[WARN] [backup] No recovery unit for %s — falling back to volume-only restore", stackName)
m.logger.Printf("[WARN] [backup] No readable recovery unit for %s at %s — falling back to volume-only restore", stackName, unitDir)
m.mu.Lock()
m.running = false // RestoreApp re-acquires the running flag
m.mu.Unlock()
@@ -220,7 +252,7 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) (UnitRestoreResult,
// was just parsed above, so the claim and the outcome are counted from the same document.
res.ManifestVolumes, res.ManifestDBs = len(manifest.VolumeDumps), len(manifest.DBDumps)
composeDir := RecoveryUnitComposePath(nsRoot, stackName)
composeDir := UnitComposeDir(unitDir)
nonSecretEnv, unitSecrets := readUnitEnv(filepath.Join(composeDir, "app.yaml"), manifest.PortableSecretEnvVars)
// D5: the unit carries the portable class, so this is the leg that no longer needs the guest. The
@@ -275,8 +307,11 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) (UnitRestoreResult,
stackName, len(unresolved), unresolved)
}
}
m.logger.Printf("[INFO] [backup] Restoring %s from recovery unit: images=%d, secrets recovered=%d/%d, data_keys=%d",
stackName, len(manifest.ImagePins), len(manifest.SecretEnvVars)-len(missing), len(manifest.SecretEnvVars), len(manifest.DataKeyEnvVars))
// R-102: the unit DIRECTORY is logged. Which copy a restore read from is now a real question with
// two answers, and "an absent log line is not evidence" — the drill reads this line to prove the
// secondary mirror, not the primary unit, was the source. It is a path, never a secret.
m.logger.Printf("[INFO] [backup] Restoring %s from recovery unit %s: images=%d, secrets recovered=%d/%d, data_keys=%d",
stackName, unitDir, len(manifest.ImagePins), len(manifest.SecretEnvVars)-len(missing), len(manifest.SecretEnvVars), len(manifest.DataKeyEnvVars))
// R-47: which compose service holds the database, and is there anything to replay? Resolved from
// the UNIT's compose, because that file is about to BECOME the live one. Both answers are needed
@@ -286,7 +321,8 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) (UnitRestoreResult,
// "cannot tell" is not "no database" — leave it empty and let the gate decide.
m.logger.Printf("[WARN] [backup] %s: could not read the unit's compose services: %v", stackName, dsErr)
}
hasDumps := hasReplayableDump(AppDBDumpPath(nsRoot, stackName))
dbDumpDir := UnitDBDumpDir(unitDir)
hasDumps := hasReplayableDump(dbDumpDir)
if hasDumps && len(dbServices) == 0 {
m.logger.Printf("[ERROR] [backup] Restore REFUSED for %s: a .sql dump exists but no database service is identifiable in the unit's compose", stackName)
return res, fmt.Errorf("Az adatbázis-szolgáltatás nem azonosítható a(z) %s alkalmazásban — a visszaállítás biztonsági okból nem indult el.", stackName)
@@ -302,7 +338,17 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) (UnitRestoreResult,
// R-353: the count is captured even when the replay errors — a partial replay is a fact the
// customer's sentence has to be built from, and discarding it on the error path is how Scenario C
// would end up wearing Scenario B's wording.
replayed, volErr := m.restoreDockerVolumes(stackName, drivePath)
// R-354's volume-REPLAY seam, reused here rather than a second one being invented. It defaults to
// the real restoreDockerVolumesFrom, so production behaviour is byte-for-byte what it was; what it
// buys is that R-102's acceptance test can assert WHICH directory the tars came out of without a
// Docker daemon. For 40 of the 53 catalogue apps that archive is the entire dataset, so "the
// mirror was the source" has to be provable for the volume leg too, not only for the env and the
// database.
volReplay := m.volumeReplayFrom
if volReplay == nil {
volReplay = m.restoreDockerVolumesFrom
}
replayed, volErr := volReplay(stackName, UnitVolumeDumpDir(unitDir))
res.VolumesReplayed = replayed
if volErr != nil {
m.logger.Printf("[ERROR] [backup] volume restore for %s: %v", stackName, volErr)
@@ -322,7 +368,7 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) (UnitRestoreResult,
if dataErr == nil {
dataErr = err
}
} else if n, err := m.reimportDBDumpsCtx(stackName, nsRoot); err != nil {
} else if n, err := m.reimportDBDumpsAtCtx(stackName, dbDumpDir); err != nil {
res.DBsReplayed = n // partial credit: whatever imported before the failure really did import
m.logger.Printf("[ERROR] [backup] DB re-import for %s: %v", stackName, err)
if dataErr == nil {