R-102: the recovery unit on the second drive becomes a way back
Tier-2 mirrors each app's whole recovery unit to <dest>/backups/secondary/<app>/recovery-unit/ on every run and has done for months. Nothing read it. In the one failure Tier-2 exists for - the primary drive is lost, and the primary unit with it - the surviving copy could not be opened by any action in the product (07-backup-architecture 6.3, 7.2). Part 1.2: RestoreFromRecoveryUnitAt(stack, unitDir) holds the whole body; RestoreFromRecoveryUnit is the thin caller naming the primary unit. ONE implementation, two callers. The SOURCE moves; the DESTINATION does not - live Docker volumes, the live database container, the guest's definition, all unchanged. The R-47 mutation order, the secret reconciliation with unit-over-guest precedence, the fail-closed data-key gate and the no-unit fallback with CountsUnknown are untouched. reimportDBDumpsAtCtx is the bounded-context twin of reimportDBDumpsFrom; the 35-minute bound is now named once so the two paths cannot drift. The R-354 volume-replay seam is reused rather than a second one invented, which is what lets the acceptance test assert the volume leg's source directory. Part 1.3: RestoreTier2Unit resolves the recorded copy, refuses fail-closed unless the mirror carries a parseable manifest - a directory is not a package - and delegates. The single-writer flag is taken inside RestoreFromRecoveryUnitAt, not beside it. Part 2.1: Tier2Coverage gains UnitRestorable and the copy's dates. CanRestore() is NOT widened; it still answers only 'can the file restore run?'. One predicate answering two questions is R-356, which refused 40 running apps for months. Tests: A2-A6 and B1-B5, plus two non-regression guards. The Tier-2 fixtures build their mirror with the production RunTier2, so the claim is 'the copy Tier-2 writes is the copy this restore reads'. Red-proofs: A5 (swap volumes/recreate -> fails on the order), B2 (point the reader back at the primary -> fails with the mirror never reaching the redeploy, and with permission denied once the primary tree is unreadable).
This commit is contained in:
@@ -178,7 +178,36 @@ type UnitRestoreResult struct {
|
||||
// D5: this no longer needs the guest. A restore with the guest's app.yaml absent succeeds, which is
|
||||
// pinned by TestRestoreFromRecoveryUnitWithGuestAbsent — the withheld class is regenerated (O4) and
|
||||
// only a data key missing from BOTH sources still refuses.
|
||||
//
|
||||
// R-102: this is now the thin caller. It names the PRIMARY unit — the app's own drive,
|
||||
// backups/primary/<stack> — and hands it to RestoreFromRecoveryUnitAt, which holds the whole body.
|
||||
// An unresolvable drive path is still refused inside …At, in the same place and with the same
|
||||
// message, so the order of the checks a caller can observe is unchanged.
|
||||
func (m *Manager) RestoreFromRecoveryUnit(stackName string) (UnitRestoreResult, error) {
|
||||
return m.RestoreFromRecoveryUnitAt(stackName, RecoveryUnitPath(m.namespaceRoot(m.GetAppDrivePath(stackName)), stackName))
|
||||
}
|
||||
|
||||
// RestoreFromRecoveryUnitAt is RestoreFromRecoveryUnit with an EXPLICIT recovery-unit directory.
|
||||
//
|
||||
// R-102. ONE implementation, two callers — the same rule restoreDockerVolumesFrom states beside
|
||||
// itself in restore.go, and for the same reason: a second copy of this body is exactly how the local
|
||||
// path and the off-site path drifted apart until nothing compared them.
|
||||
//
|
||||
// The reason it exists: Tier-2 mirrors the app's whole recovery unit to
|
||||
// <dest>/backups/secondary/<stack>/recovery-unit/ on every run, and until now every reader of a unit
|
||||
// could only name a path under backups/primary/. So in the one failure Tier-2 exists for — the
|
||||
// primary drive is lost, taking the primary unit with it — the surviving copy could not be opened by
|
||||
// any action in the product (07-backup-architecture §6.3, §7.2).
|
||||
//
|
||||
// THE SOURCE MOVES; THE DESTINATION DOES NOT. unitDir changes only where the manifest, the compose
|
||||
// capture, the .sql dumps and the volume tars are READ from. The app's data is written back to the
|
||||
// live Docker volumes and the live database container, and its definition to the guest, exactly as
|
||||
// before — a restore that also relocated the app's data would be a migration, not a restore.
|
||||
//
|
||||
// Everything else is pinned and unchanged: the mutation order (stop → volumes → recreate → DB-only
|
||||
// start → replay → start, R-47), the secret reconciliation with unit-over-guest precedence and the
|
||||
// fail-closed data-key gate, and the no-unit fallback to RestoreApp with its CountsUnknown handling.
|
||||
func (m *Manager) RestoreFromRecoveryUnitAt(stackName, unitDir string) (UnitRestoreResult, error) {
|
||||
var res UnitRestoreResult
|
||||
if m.stackProvider == nil {
|
||||
return res, fmt.Errorf("stack provider not configured")
|
||||
@@ -197,15 +226,18 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) (UnitRestoreResult,
|
||||
m.mu.Unlock()
|
||||
}()
|
||||
|
||||
// The DESTINATION side, and it is deliberately still resolved here: RestoreApp (the no-unit
|
||||
// fallback below) needs it, and 07-backup-architecture §6.3 records that the restore destination
|
||||
// is resolved by the same rule as the capture destination. It is no longer used to derive any
|
||||
// SOURCE path — that is what unitDir is for.
|
||||
drivePath := m.GetAppDrivePath(stackName)
|
||||
if drivePath == "" || !filepath.IsAbs(drivePath) {
|
||||
return res, fmt.Errorf("cannot determine drive path for %s", stackName)
|
||||
}
|
||||
nsRoot := m.namespaceRoot(drivePath)
|
||||
|
||||
manifest := readManifest(RecoveryUnitManifestPath(nsRoot, stackName))
|
||||
manifest := readManifest(UnitManifestFile(unitDir))
|
||||
if manifest == nil {
|
||||
m.logger.Printf("[WARN] [backup] No recovery unit for %s — falling back to volume-only restore", stackName)
|
||||
m.logger.Printf("[WARN] [backup] No readable recovery unit for %s at %s — falling back to volume-only restore", stackName, unitDir)
|
||||
m.mu.Lock()
|
||||
m.running = false // RestoreApp re-acquires the running flag
|
||||
m.mu.Unlock()
|
||||
@@ -220,7 +252,7 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) (UnitRestoreResult,
|
||||
// was just parsed above, so the claim and the outcome are counted from the same document.
|
||||
res.ManifestVolumes, res.ManifestDBs = len(manifest.VolumeDumps), len(manifest.DBDumps)
|
||||
|
||||
composeDir := RecoveryUnitComposePath(nsRoot, stackName)
|
||||
composeDir := UnitComposeDir(unitDir)
|
||||
nonSecretEnv, unitSecrets := readUnitEnv(filepath.Join(composeDir, "app.yaml"), manifest.PortableSecretEnvVars)
|
||||
|
||||
// D5: the unit carries the portable class, so this is the leg that no longer needs the guest. The
|
||||
@@ -275,8 +307,11 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) (UnitRestoreResult,
|
||||
stackName, len(unresolved), unresolved)
|
||||
}
|
||||
}
|
||||
m.logger.Printf("[INFO] [backup] Restoring %s from recovery unit: images=%d, secrets recovered=%d/%d, data_keys=%d",
|
||||
stackName, len(manifest.ImagePins), len(manifest.SecretEnvVars)-len(missing), len(manifest.SecretEnvVars), len(manifest.DataKeyEnvVars))
|
||||
// R-102: the unit DIRECTORY is logged. Which copy a restore read from is now a real question with
|
||||
// two answers, and "an absent log line is not evidence" — the drill reads this line to prove the
|
||||
// secondary mirror, not the primary unit, was the source. It is a path, never a secret.
|
||||
m.logger.Printf("[INFO] [backup] Restoring %s from recovery unit %s: images=%d, secrets recovered=%d/%d, data_keys=%d",
|
||||
stackName, unitDir, len(manifest.ImagePins), len(manifest.SecretEnvVars)-len(missing), len(manifest.SecretEnvVars), len(manifest.DataKeyEnvVars))
|
||||
|
||||
// R-47: which compose service holds the database, and is there anything to replay? Resolved from
|
||||
// the UNIT's compose, because that file is about to BECOME the live one. Both answers are needed
|
||||
@@ -286,7 +321,8 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) (UnitRestoreResult,
|
||||
// "cannot tell" is not "no database" — leave it empty and let the gate decide.
|
||||
m.logger.Printf("[WARN] [backup] %s: could not read the unit's compose services: %v", stackName, dsErr)
|
||||
}
|
||||
hasDumps := hasReplayableDump(AppDBDumpPath(nsRoot, stackName))
|
||||
dbDumpDir := UnitDBDumpDir(unitDir)
|
||||
hasDumps := hasReplayableDump(dbDumpDir)
|
||||
if hasDumps && len(dbServices) == 0 {
|
||||
m.logger.Printf("[ERROR] [backup] Restore REFUSED for %s: a .sql dump exists but no database service is identifiable in the unit's compose", stackName)
|
||||
return res, fmt.Errorf("Az adatbázis-szolgáltatás nem azonosítható a(z) %s alkalmazásban — a visszaállítás biztonsági okból nem indult el.", stackName)
|
||||
@@ -302,7 +338,17 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) (UnitRestoreResult,
|
||||
// R-353: the count is captured even when the replay errors — a partial replay is a fact the
|
||||
// customer's sentence has to be built from, and discarding it on the error path is how Scenario C
|
||||
// would end up wearing Scenario B's wording.
|
||||
replayed, volErr := m.restoreDockerVolumes(stackName, drivePath)
|
||||
// R-354's volume-REPLAY seam, reused here rather than a second one being invented. It defaults to
|
||||
// the real restoreDockerVolumesFrom, so production behaviour is byte-for-byte what it was; what it
|
||||
// buys is that R-102's acceptance test can assert WHICH directory the tars came out of without a
|
||||
// Docker daemon. For 40 of the 53 catalogue apps that archive is the entire dataset, so "the
|
||||
// mirror was the source" has to be provable for the volume leg too, not only for the env and the
|
||||
// database.
|
||||
volReplay := m.volumeReplayFrom
|
||||
if volReplay == nil {
|
||||
volReplay = m.restoreDockerVolumesFrom
|
||||
}
|
||||
replayed, volErr := volReplay(stackName, UnitVolumeDumpDir(unitDir))
|
||||
res.VolumesReplayed = replayed
|
||||
if volErr != nil {
|
||||
m.logger.Printf("[ERROR] [backup] volume restore for %s: %v", stackName, volErr)
|
||||
@@ -322,7 +368,7 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) (UnitRestoreResult,
|
||||
if dataErr == nil {
|
||||
dataErr = err
|
||||
}
|
||||
} else if n, err := m.reimportDBDumpsCtx(stackName, nsRoot); err != nil {
|
||||
} else if n, err := m.reimportDBDumpsAtCtx(stackName, dbDumpDir); err != nil {
|
||||
res.DBsReplayed = n // partial credit: whatever imported before the failure really did import
|
||||
m.logger.Printf("[ERROR] [backup] DB re-import for %s: %v", stackName, err)
|
||||
if dataErr == nil {
|
||||
|
||||
Reference in New Issue
Block a user