R-102: the recovery unit on the second drive becomes a way back

Tier-2 mirrors each app's whole recovery unit to <dest>/backups/secondary/<app>/recovery-unit/ on
every run and has done for months. Nothing read it. In the one failure Tier-2 exists for - the
primary drive is lost, and the primary unit with it - the surviving copy could not be opened by any
action in the product (07-backup-architecture 6.3, 7.2).

Part 1.2: RestoreFromRecoveryUnitAt(stack, unitDir) holds the whole body; RestoreFromRecoveryUnit is
the thin caller naming the primary unit. ONE implementation, two callers. The SOURCE moves; the
DESTINATION does not - live Docker volumes, the live database container, the guest's definition, all
unchanged. The R-47 mutation order, the secret reconciliation with unit-over-guest precedence, the
fail-closed data-key gate and the no-unit fallback with CountsUnknown are untouched.
reimportDBDumpsAtCtx is the bounded-context twin of reimportDBDumpsFrom; the 35-minute bound is now
named once so the two paths cannot drift. The R-354 volume-replay seam is reused rather than a second
one invented, which is what lets the acceptance test assert the volume leg's source directory.

Part 1.3: RestoreTier2Unit resolves the recorded copy, refuses fail-closed unless the mirror carries a
parseable manifest - a directory is not a package - and delegates. The single-writer flag is taken
inside RestoreFromRecoveryUnitAt, not beside it.

Part 2.1: Tier2Coverage gains UnitRestorable and the copy's dates. CanRestore() is NOT widened; it
still answers only 'can the file restore run?'. One predicate answering two questions is R-356, which
refused 40 running apps for months.

Tests: A2-A6 and B1-B5, plus two non-regression guards. The Tier-2 fixtures build their mirror with
the production RunTier2, so the claim is 'the copy Tier-2 writes is the copy this restore reads'.
Red-proofs: A5 (swap volumes/recreate -> fails on the order), B2 (point the reader back at the
primary -> fails with the mirror never reaching the redeploy, and with permission denied once the
primary tree is unreadable).
This commit is contained in:
2026-08-31 11:30:34 +02:00
parent c732006d26
commit 0f9b796615
5 changed files with 856 additions and 16 deletions
+118 -6
View File
@@ -39,6 +39,12 @@ var (
// are in this class. Exported so the handler can refuse BEFORE stopping the app and name the action
// that does work, instead of taking an outage and reporting "no missing files".
ErrTier2NoRestorableData = errors.New("ennek az alkalmazásnak az adatai nem ebből a másolatból állíthatók vissza")
// ErrTier2NoUnitInCopy (R-102) — the recorded Tier-2 copy holds no OPENABLE recovery unit: either
// recovery-unit/ is absent, or it is a directory without a readable manifest.json. Exported so the
// handler can refuse before beginning any op. FAIL CLOSED is the whole point of the second half:
// a directory that exists is not a package, and reading a half-copied mirror as if it were one is
// how a restore would overwrite live data with nothing.
ErrTier2NoUnitInCopy = errors.New("a másodlagos másolatban nincs megnyitható mentési egység ehhez az alkalmazáshoz")
)
// Tier2Coverage says what a Tier-2 restore can and cannot return for one app — the asymmetry C9-F1
@@ -50,14 +56,64 @@ var (
// tarballs — which this restore path never opens. An app can have HasUnit && no Legs (43 of 53), in
// which case the restore is a guaranteed no-op no matter how much data was lost.
type Tier2Coverage struct {
Legs []string // subtrees this restore reads and that exist in the copy: "hdd", "userdata"
HasUnit bool // recovery-unit/ present — captured, but NOT restorable by this path
Legs []string // subtrees the FILE restore reads and that exist in the copy: "hdd", "userdata"
HasUnit bool // recovery-unit/ present as a DIRECTORY — the disclosure fact, see below
// UnitRestorable (R-102) — the mirror is a real PACKAGE, not merely a directory: recovery-unit/
// exists AND carries a manifest.json that parses. This is the gate for the UNIT restore.
//
// It is a SECOND FIELD and not a widening of HasUnit, and the distinction is load-bearing in both
// directions. HasUnit answers "is there captured data this FILE restore is not looking at?" — the
// question tier2UnitNotCoveredMsg is appended for, and the honest answer for a half-copied mirror
// is still yes. UnitRestorable answers "can the unit restore open this?" — and for that same
// half-copied mirror the answer is no. Collapsing them would either silence a true disclosure or
// arm a restore over an unopenable package.
UnitRestorable bool
// CopyLastRun / CopyLastSuccess — WHEN the copy this restore would read was written, so the
// surface can name the date before it overwrites anything with it (Scenario E). Filled only by
// Tier2RestoreCoverage, which is the path that holds the settings; tier2CoverageAt is a pure
// filesystem inspection and leaves them empty. They are STRINGS in the recorded RFC3339 form,
// carried verbatim — no formatting decision is taken in this package.
//
// R-101 applies here exactly as it does on the backup card: CopyLastRun is the ATTEMPT clock and
// CopyLastSuccess is the only evidence a copy was actually made. A surface that shows one must
// not present it as the other.
CopyLastRun string
CopyLastSuccess string
}
// CanRestore reports whether the restore has any subtree to read at all.
// CanRestore reports whether the FILE restore has any subtree to read at all.
//
// R-102/R-103: this answers exactly one question and must keep answering only that one. Widening it
// to include the unit is R-356 arriving a second time — there, ONE predicate meant both "has this app
// a drive?" and "is this app installed?", and it refused 40 running apps for months while telling
// their owners to reinstall them somewhere those apps never offer. Two questions, two predicates.
func (c Tier2Coverage) CanRestore() bool { return len(c.Legs) > 0 }
// tier2CoverageAt inspects a resolved copy directory. Pure filesystem stat — no side effects.
// CanRestoreUnit reports whether the UNIT restore can run from this copy — the second predicate.
func (c Tier2Coverage) CanRestoreUnit() bool { return c.UnitRestorable }
// tier2UnitDir returns the recovery-unit directory inside a resolved Tier-2 copy. ONE expression of
// where Tier-2 puts the mirror; tier2.go writes it at the same relative name ("Unit leg (always)").
func tier2UnitDir(destBase string) string {
return filepath.Join(destBase, "recovery-unit")
}
// tier2UnitIsOpenable reports whether a mirrored unit directory is a PACKAGE and not just a
// directory. Fail-closed by construction: the manifest must be present AND parse (readManifest
// returns nil for both a missing file and malformed JSON), because an unopenable unit that armed a
// restore would stop the app, replay nothing, and rewrite its definition from an empty capture.
//
// This is the R-358 lesson one tier over — the off-site scratch marker — arriving on Tier-2.
func tier2UnitIsOpenable(unitDir string) bool {
if fi, err := os.Stat(unitDir); err != nil || !fi.IsDir() {
return false
}
return readManifest(UnitManifestFile(unitDir)) != nil
}
// tier2CoverageAt inspects a resolved copy directory. Pure filesystem stat/read — no side effects.
func tier2CoverageAt(destBase string) Tier2Coverage {
var c Tier2Coverage
for _, leg := range []string{"hdd", "userdata"} {
@@ -65,9 +121,11 @@ func tier2CoverageAt(destBase string) Tier2Coverage {
c.Legs = append(c.Legs, leg)
}
}
if fi, err := os.Stat(filepath.Join(destBase, "recovery-unit")); err == nil && fi.IsDir() {
unitDir := tier2UnitDir(destBase)
if fi, err := os.Stat(unitDir); err == nil && fi.IsDir() {
c.HasUnit = true
}
c.UnitRestorable = tier2UnitIsOpenable(unitDir)
return c
}
@@ -79,7 +137,61 @@ func (m *Manager) Tier2RestoreCoverage(stackName string) (Tier2Coverage, error)
if err != nil {
return Tier2Coverage{}, err
}
return tier2CoverageAt(destBase), nil
cov := tier2CoverageAt(destBase)
// R-102: carry WHEN the copy was written, so the surface can name the date on an action that
// overwrites live data with it. Read from the same recorded config tier2RecordedCopyDir just
// resolved the path from, so the date and the directory cannot describe different runs.
if m.settings != nil {
if cfg := m.settings.GetCrossDriveConfig(stackName); cfg != nil {
cov.CopyLastRun, cov.CopyLastSuccess = cfg.LastRun, cfg.LastSuccess
}
}
return cov, nil
}
// RestoreTier2Unit runs the FULL recovery-unit restore from the app's Tier-2 copy on the SECOND
// DRIVE — R-102, and the reason this task exists.
//
// Tier-2 has mirrored each app's whole recovery unit to
// <dest>/backups/secondary/<app>/recovery-unit/ on every run for months, and no code path read it:
// every reader of a unit could only name a path under backups/primary/. So in the exact failure
// Tier-2 exists for — the primary drive is lost, and the primary unit with it — the surviving copy
// was unopenable by any customer action (07-backup-architecture §6.3, §7.2).
//
// It is NOT the additive file restore beside it. This one OVERWRITES: named volumes are recreated
// from the mirror's tars and the database is replayed from the mirror's dump. The surface must carry
// that difference; see tier2UnitConfirm in the web package.
//
// THE SINGLE-WRITER FLAG IS TAKEN INSIDE RestoreFromRecoveryUnitAt, exactly as on the primary path —
// do NOT add an acquireRunning() here. A second acquire would refuse the restore it is guarding.
func (m *Manager) RestoreTier2Unit(stackName string) (UnitRestoreResult, error) {
destBase, err := m.tier2RecordedCopyDir(stackName)
if err != nil {
return UnitRestoreResult{}, err
}
unitDir := tier2UnitDir(destBase)
if !tier2UnitIsOpenable(unitDir) {
m.logger.Printf("[WARN] [backup] Tier-2 unit restore refused for %s: no openable recovery unit in the recorded copy — the app was NOT stopped", stackName)
return UnitRestoreResult{}, ErrTier2NoUnitInCopy
}
// WARN, not INFO: this is the destructive one of the two Tier-2 restores. The unit directory is a
// path and never a secret, and naming it is what makes "the SECONDARY mirror was the source" a
// positive observable in the log rather than an absence to be argued from.
m.logger.Printf("[WARN] [backup] Tier-2 UNIT restore for %s from the secondary mirror %s — this OVERWRITES live app data", stackName, unitDir)
return m.RestoreFromRecoveryUnitAt(stackName, unitDir)
}
// Tier2CopyDate returns the date the surface should name for this app's Tier-2 copy, preferring the
// last SUCCESS over the last ATTEMPT (R-101: a timestamp that records "we tried" cannot answer "did
// it work"), and reports whether the returned value is a proven success.
//
// It exists so the confirm text and the outcome sentence cannot disagree about which copy is being
// restored: one resolver, two readers.
func (c Tier2Coverage) Tier2CopyDate() (date string, proven bool) {
if c.CopyLastSuccess != "" {
return c.CopyLastSuccess, true
}
return c.CopyLastRun, false
}
// tier2RecordedCopyDir resolves the RECORDED Tier-2 copy dir for a stack, applying every