R-102: the recovery unit on the second drive becomes a way back

Tier-2 mirrors each app's whole recovery unit to <dest>/backups/secondary/<app>/recovery-unit/ on
every run and has done for months. Nothing read it. In the one failure Tier-2 exists for - the
primary drive is lost, and the primary unit with it - the surviving copy could not be opened by any
action in the product (07-backup-architecture 6.3, 7.2).

Part 1.2: RestoreFromRecoveryUnitAt(stack, unitDir) holds the whole body; RestoreFromRecoveryUnit is
the thin caller naming the primary unit. ONE implementation, two callers. The SOURCE moves; the
DESTINATION does not - live Docker volumes, the live database container, the guest's definition, all
unchanged. The R-47 mutation order, the secret reconciliation with unit-over-guest precedence, the
fail-closed data-key gate and the no-unit fallback with CountsUnknown are untouched.
reimportDBDumpsAtCtx is the bounded-context twin of reimportDBDumpsFrom; the 35-minute bound is now
named once so the two paths cannot drift. The R-354 volume-replay seam is reused rather than a second
one invented, which is what lets the acceptance test assert the volume leg's source directory.

Part 1.3: RestoreTier2Unit resolves the recorded copy, refuses fail-closed unless the mirror carries a
parseable manifest - a directory is not a package - and delegates. The single-writer flag is taken
inside RestoreFromRecoveryUnitAt, not beside it.

Part 2.1: Tier2Coverage gains UnitRestorable and the copy's dates. CanRestore() is NOT widened; it
still answers only 'can the file restore run?'. One predicate answering two questions is R-356, which
refused 40 running apps for months.

Tests: A2-A6 and B1-B5, plus two non-regression guards. The Tier-2 fixtures build their mirror with
the production RunTier2, so the claim is 'the copy Tier-2 writes is the copy this restore reads'.
Red-proofs: A5 (swap volumes/recreate -> fails on the order), B2 (point the reader back at the
primary -> fails with the mirror never reaching the redeploy, and with permission denied once the
primary tree is unreadable).
This commit is contained in:
2026-08-31 11:30:34 +02:00
parent c732006d26
commit 0f9b796615
5 changed files with 856 additions and 16 deletions
+16 -1
View File
@@ -89,10 +89,25 @@ func (m *Manager) reimportDBDumpsFrom(ctx context.Context, stackName, dumpDir st
return imported, nil
}
// dbReimportTimeout bounds a DB replay so a stuck import cannot hang a restore indefinitely. Named
// once because both bounded entry points below must agree: two paths that differ in how long they
// let a wedged import hold the restore are two different products (R-102 added the second one).
const dbReimportTimeout = 35 * time.Minute
// reimportDBDumpsCtx is a small helper that runs reimportDBDumps with a bounded context so a stuck DB
// import cannot hang the restore indefinitely.
func (m *Manager) reimportDBDumpsCtx(stackName, nsRoot string) (int, error) {
ctx, cancel := context.WithTimeout(context.Background(), 35*time.Minute)
ctx, cancel := context.WithTimeout(context.Background(), dbReimportTimeout)
defer cancel()
return m.reimportDBDumps(ctx, stackName, nsRoot)
}
// reimportDBDumpsAtCtx is reimportDBDumpsCtx with an EXPLICIT dump directory — the bounded-context
// twin of reimportDBDumpsFrom, added for R-102 so the Tier-2 unit restore can replay out of the
// secondary mirror. Same discovery/import seams, same failure semantics, same bound; only the source
// directory differs.
func (m *Manager) reimportDBDumpsAtCtx(stackName, dumpDir string) (int, error) {
ctx, cancel := context.WithTimeout(context.Background(), dbReimportTimeout)
defer cancel()
return m.reimportDBDumpsFrom(ctx, stackName, dumpDir)
}