R-102: the recovery unit on the second drive becomes a way back
Tier-2 mirrors each app's whole recovery unit to <dest>/backups/secondary/<app>/recovery-unit/ on every run and has done for months. Nothing read it. In the one failure Tier-2 exists for - the primary drive is lost, and the primary unit with it - the surviving copy could not be opened by any action in the product (07-backup-architecture 6.3, 7.2). Part 1.2: RestoreFromRecoveryUnitAt(stack, unitDir) holds the whole body; RestoreFromRecoveryUnit is the thin caller naming the primary unit. ONE implementation, two callers. The SOURCE moves; the DESTINATION does not - live Docker volumes, the live database container, the guest's definition, all unchanged. The R-47 mutation order, the secret reconciliation with unit-over-guest precedence, the fail-closed data-key gate and the no-unit fallback with CountsUnknown are untouched. reimportDBDumpsAtCtx is the bounded-context twin of reimportDBDumpsFrom; the 35-minute bound is now named once so the two paths cannot drift. The R-354 volume-replay seam is reused rather than a second one invented, which is what lets the acceptance test assert the volume leg's source directory. Part 1.3: RestoreTier2Unit resolves the recorded copy, refuses fail-closed unless the mirror carries a parseable manifest - a directory is not a package - and delegates. The single-writer flag is taken inside RestoreFromRecoveryUnitAt, not beside it. Part 2.1: Tier2Coverage gains UnitRestorable and the copy's dates. CanRestore() is NOT widened; it still answers only 'can the file restore run?'. One predicate answering two questions is R-356, which refused 40 running apps for months. Tests: A2-A6 and B1-B5, plus two non-regression guards. The Tier-2 fixtures build their mirror with the production RunTier2, so the claim is 'the copy Tier-2 writes is the copy this restore reads'. Red-proofs: A5 (swap volumes/recreate -> fails on the order), B2 (point the reader back at the primary -> fails with the mirror never reaching the redeploy, and with permission denied once the primary tree is unreadable).
This commit is contained in:
@@ -39,6 +39,12 @@ var (
|
||||
// are in this class. Exported so the handler can refuse BEFORE stopping the app and name the action
|
||||
// that does work, instead of taking an outage and reporting "no missing files".
|
||||
ErrTier2NoRestorableData = errors.New("ennek az alkalmazásnak az adatai nem ebből a másolatból állíthatók vissza")
|
||||
// ErrTier2NoUnitInCopy (R-102) — the recorded Tier-2 copy holds no OPENABLE recovery unit: either
|
||||
// recovery-unit/ is absent, or it is a directory without a readable manifest.json. Exported so the
|
||||
// handler can refuse before beginning any op. FAIL CLOSED is the whole point of the second half:
|
||||
// a directory that exists is not a package, and reading a half-copied mirror as if it were one is
|
||||
// how a restore would overwrite live data with nothing.
|
||||
ErrTier2NoUnitInCopy = errors.New("a másodlagos másolatban nincs megnyitható mentési egység ehhez az alkalmazáshoz")
|
||||
)
|
||||
|
||||
// Tier2Coverage says what a Tier-2 restore can and cannot return for one app — the asymmetry C9-F1
|
||||
@@ -50,14 +56,64 @@ var (
|
||||
// tarballs — which this restore path never opens. An app can have HasUnit && no Legs (43 of 53), in
|
||||
// which case the restore is a guaranteed no-op no matter how much data was lost.
|
||||
type Tier2Coverage struct {
|
||||
Legs []string // subtrees this restore reads and that exist in the copy: "hdd", "userdata"
|
||||
HasUnit bool // recovery-unit/ present — captured, but NOT restorable by this path
|
||||
Legs []string // subtrees the FILE restore reads and that exist in the copy: "hdd", "userdata"
|
||||
HasUnit bool // recovery-unit/ present as a DIRECTORY — the disclosure fact, see below
|
||||
|
||||
// UnitRestorable (R-102) — the mirror is a real PACKAGE, not merely a directory: recovery-unit/
|
||||
// exists AND carries a manifest.json that parses. This is the gate for the UNIT restore.
|
||||
//
|
||||
// It is a SECOND FIELD and not a widening of HasUnit, and the distinction is load-bearing in both
|
||||
// directions. HasUnit answers "is there captured data this FILE restore is not looking at?" — the
|
||||
// question tier2UnitNotCoveredMsg is appended for, and the honest answer for a half-copied mirror
|
||||
// is still yes. UnitRestorable answers "can the unit restore open this?" — and for that same
|
||||
// half-copied mirror the answer is no. Collapsing them would either silence a true disclosure or
|
||||
// arm a restore over an unopenable package.
|
||||
UnitRestorable bool
|
||||
|
||||
// CopyLastRun / CopyLastSuccess — WHEN the copy this restore would read was written, so the
|
||||
// surface can name the date before it overwrites anything with it (Scenario E). Filled only by
|
||||
// Tier2RestoreCoverage, which is the path that holds the settings; tier2CoverageAt is a pure
|
||||
// filesystem inspection and leaves them empty. They are STRINGS in the recorded RFC3339 form,
|
||||
// carried verbatim — no formatting decision is taken in this package.
|
||||
//
|
||||
// R-101 applies here exactly as it does on the backup card: CopyLastRun is the ATTEMPT clock and
|
||||
// CopyLastSuccess is the only evidence a copy was actually made. A surface that shows one must
|
||||
// not present it as the other.
|
||||
CopyLastRun string
|
||||
CopyLastSuccess string
|
||||
}
|
||||
|
||||
// CanRestore reports whether the restore has any subtree to read at all.
|
||||
// CanRestore reports whether the FILE restore has any subtree to read at all.
|
||||
//
|
||||
// R-102/R-103: this answers exactly one question and must keep answering only that one. Widening it
|
||||
// to include the unit is R-356 arriving a second time — there, ONE predicate meant both "has this app
|
||||
// a drive?" and "is this app installed?", and it refused 40 running apps for months while telling
|
||||
// their owners to reinstall them somewhere those apps never offer. Two questions, two predicates.
|
||||
func (c Tier2Coverage) CanRestore() bool { return len(c.Legs) > 0 }
|
||||
|
||||
// tier2CoverageAt inspects a resolved copy directory. Pure filesystem stat — no side effects.
|
||||
// CanRestoreUnit reports whether the UNIT restore can run from this copy — the second predicate.
|
||||
func (c Tier2Coverage) CanRestoreUnit() bool { return c.UnitRestorable }
|
||||
|
||||
// tier2UnitDir returns the recovery-unit directory inside a resolved Tier-2 copy. ONE expression of
|
||||
// where Tier-2 puts the mirror; tier2.go writes it at the same relative name ("Unit leg (always)").
|
||||
func tier2UnitDir(destBase string) string {
|
||||
return filepath.Join(destBase, "recovery-unit")
|
||||
}
|
||||
|
||||
// tier2UnitIsOpenable reports whether a mirrored unit directory is a PACKAGE and not just a
|
||||
// directory. Fail-closed by construction: the manifest must be present AND parse (readManifest
|
||||
// returns nil for both a missing file and malformed JSON), because an unopenable unit that armed a
|
||||
// restore would stop the app, replay nothing, and rewrite its definition from an empty capture.
|
||||
//
|
||||
// This is the R-358 lesson one tier over — the off-site scratch marker — arriving on Tier-2.
|
||||
func tier2UnitIsOpenable(unitDir string) bool {
|
||||
if fi, err := os.Stat(unitDir); err != nil || !fi.IsDir() {
|
||||
return false
|
||||
}
|
||||
return readManifest(UnitManifestFile(unitDir)) != nil
|
||||
}
|
||||
|
||||
// tier2CoverageAt inspects a resolved copy directory. Pure filesystem stat/read — no side effects.
|
||||
func tier2CoverageAt(destBase string) Tier2Coverage {
|
||||
var c Tier2Coverage
|
||||
for _, leg := range []string{"hdd", "userdata"} {
|
||||
@@ -65,9 +121,11 @@ func tier2CoverageAt(destBase string) Tier2Coverage {
|
||||
c.Legs = append(c.Legs, leg)
|
||||
}
|
||||
}
|
||||
if fi, err := os.Stat(filepath.Join(destBase, "recovery-unit")); err == nil && fi.IsDir() {
|
||||
unitDir := tier2UnitDir(destBase)
|
||||
if fi, err := os.Stat(unitDir); err == nil && fi.IsDir() {
|
||||
c.HasUnit = true
|
||||
}
|
||||
c.UnitRestorable = tier2UnitIsOpenable(unitDir)
|
||||
return c
|
||||
}
|
||||
|
||||
@@ -79,7 +137,61 @@ func (m *Manager) Tier2RestoreCoverage(stackName string) (Tier2Coverage, error)
|
||||
if err != nil {
|
||||
return Tier2Coverage{}, err
|
||||
}
|
||||
return tier2CoverageAt(destBase), nil
|
||||
cov := tier2CoverageAt(destBase)
|
||||
// R-102: carry WHEN the copy was written, so the surface can name the date on an action that
|
||||
// overwrites live data with it. Read from the same recorded config tier2RecordedCopyDir just
|
||||
// resolved the path from, so the date and the directory cannot describe different runs.
|
||||
if m.settings != nil {
|
||||
if cfg := m.settings.GetCrossDriveConfig(stackName); cfg != nil {
|
||||
cov.CopyLastRun, cov.CopyLastSuccess = cfg.LastRun, cfg.LastSuccess
|
||||
}
|
||||
}
|
||||
return cov, nil
|
||||
}
|
||||
|
||||
// RestoreTier2Unit runs the FULL recovery-unit restore from the app's Tier-2 copy on the SECOND
|
||||
// DRIVE — R-102, and the reason this task exists.
|
||||
//
|
||||
// Tier-2 has mirrored each app's whole recovery unit to
|
||||
// <dest>/backups/secondary/<app>/recovery-unit/ on every run for months, and no code path read it:
|
||||
// every reader of a unit could only name a path under backups/primary/. So in the exact failure
|
||||
// Tier-2 exists for — the primary drive is lost, and the primary unit with it — the surviving copy
|
||||
// was unopenable by any customer action (07-backup-architecture §6.3, §7.2).
|
||||
//
|
||||
// It is NOT the additive file restore beside it. This one OVERWRITES: named volumes are recreated
|
||||
// from the mirror's tars and the database is replayed from the mirror's dump. The surface must carry
|
||||
// that difference; see tier2UnitConfirm in the web package.
|
||||
//
|
||||
// THE SINGLE-WRITER FLAG IS TAKEN INSIDE RestoreFromRecoveryUnitAt, exactly as on the primary path —
|
||||
// do NOT add an acquireRunning() here. A second acquire would refuse the restore it is guarding.
|
||||
func (m *Manager) RestoreTier2Unit(stackName string) (UnitRestoreResult, error) {
|
||||
destBase, err := m.tier2RecordedCopyDir(stackName)
|
||||
if err != nil {
|
||||
return UnitRestoreResult{}, err
|
||||
}
|
||||
unitDir := tier2UnitDir(destBase)
|
||||
if !tier2UnitIsOpenable(unitDir) {
|
||||
m.logger.Printf("[WARN] [backup] Tier-2 unit restore refused for %s: no openable recovery unit in the recorded copy — the app was NOT stopped", stackName)
|
||||
return UnitRestoreResult{}, ErrTier2NoUnitInCopy
|
||||
}
|
||||
// WARN, not INFO: this is the destructive one of the two Tier-2 restores. The unit directory is a
|
||||
// path and never a secret, and naming it is what makes "the SECONDARY mirror was the source" a
|
||||
// positive observable in the log rather than an absence to be argued from.
|
||||
m.logger.Printf("[WARN] [backup] Tier-2 UNIT restore for %s from the secondary mirror %s — this OVERWRITES live app data", stackName, unitDir)
|
||||
return m.RestoreFromRecoveryUnitAt(stackName, unitDir)
|
||||
}
|
||||
|
||||
// Tier2CopyDate returns the date the surface should name for this app's Tier-2 copy, preferring the
|
||||
// last SUCCESS over the last ATTEMPT (R-101: a timestamp that records "we tried" cannot answer "did
|
||||
// it work"), and reports whether the returned value is a proven success.
|
||||
//
|
||||
// It exists so the confirm text and the outcome sentence cannot disagree about which copy is being
|
||||
// restored: one resolver, two readers.
|
||||
func (c Tier2Coverage) Tier2CopyDate() (date string, proven bool) {
|
||||
if c.CopyLastSuccess != "" {
|
||||
return c.CopyLastSuccess, true
|
||||
}
|
||||
return c.CopyLastRun, false
|
||||
}
|
||||
|
||||
// tier2RecordedCopyDir resolves the RECORDED Tier-2 copy dir for a stack, applying every
|
||||
|
||||
Reference in New Issue
Block a user