MEASURED FIRST, then fixed. On the shipped v0.229.0, on demo-hp, an app's Tier-2 copy went from 120 082 104 B (4 database dumps + 3 named-volume tars) to 7 036 B (none of either) in ONE nightly run, and the run recorded itself a success: 'Tier 2 copied docmost -> ... (14.9 KB, 0 leg(s), 0s)'. Evidence: felhom.eu/documentation/audits/DRILL-r403-tier2-delete-2026-08-31/. The mechanism was three individually-correct lines: RunTier2 guards the unit leg with os.Stat only (does the folder exist), rsyncMirror is rsync -a --delete, and nothing between them compared source to destination. An EMPTY unit is a folder that exists. THE GUARD. One predicate, unitCarriesData/unitIsHollow (r403_hollow.go), asking the MANIFEST and never the byte size - a big compose tree with no dumps is dangerous, a tiny unit for a tiny app is fine. Fail closed on an absent or unparseable manifest. RunTier2 skips the unit leg when the source is hollow AND the destination is not; the other legs still run, the run is not failed, and the skip is recorded for the SURFACE (CrossDriveBackup.UnitLegSkipped + UnitPackageDate) as well as logged. --delete STAYS and shrinking stays legal. 07 section 8 row 5's derived-copy rule is unchanged; the fence is exactly one shape. TestR403_DataLegShrinkIsUnaffected is the guard on the guard. THE HONESTY. A preserved package is older than the run that preserved it, so the card carries a notice and the unit-restore confirm names the PACKAGE's date - read from the mirrored manifest's own created_at, not from the status record - plus a clause saying why it is older. THE CAUSE. RestoreTier2Unit now refills a hollow or absent primary unit from the mirror it just restored from, INSIDE the call before returning. The hollow manifest was written two seconds after a restore by the 5-minute capture job; any follow-up job races it. The capture itself is NOT guarded: a capture describing an empty drive as empty is correct, and with the primary refilled there is no hollow state left to describe. Never over a complete primary, never after a failed restore. recordTier2Success and tier2UnitConfirmMsg keep their old signatures as thin callers, so no existing test needed editing. New seam unitRehydrate, separate from tier2Mirror on purpose. 22 new Go tests. Red-proofs run and reverted: A6 (predicate -> size threshold), B1 (guard removed -> the copy's 3 files are DELETED and the seam is called), B6 (a general never-shrink rule -> the shrink case fails), C2 (only-when-hollow dropped -> the complete primary is overwritten).
This commit is contained in:
@@ -81,6 +81,19 @@ type Tier2Coverage struct {
|
||||
// not present it as the other.
|
||||
CopyLastRun string
|
||||
CopyLastSuccess string
|
||||
|
||||
// UnitPackageDate / UnitLegPreserved (R-403) — WHEN the package in this copy was actually
|
||||
// captured, and whether the newest run PRESERVED it instead of refreshing it.
|
||||
//
|
||||
// They exist because after an R-403 skip, CopyLastRun and CopyLastSuccess stop describing the
|
||||
// package: the run really did succeed and really is from today, and the package in the copy is
|
||||
// from before it. A surface that names the run date as the package date would be trading a data
|
||||
// loss for a comforting lie, which is the failure family this project keeps finding.
|
||||
//
|
||||
// UnitPackageDate is read from the MIRRORED UNIT'S OWN MANIFEST, not from the recorded status, so
|
||||
// it is a fact about the artifact the restore will actually open. "" means UNKNOWN.
|
||||
UnitPackageDate string
|
||||
UnitLegPreserved bool
|
||||
}
|
||||
|
||||
// CanRestore reports whether the FILE restore has any subtree to read at all.
|
||||
@@ -126,6 +139,9 @@ func tier2CoverageAt(destBase string) Tier2Coverage {
|
||||
c.HasUnit = true
|
||||
}
|
||||
c.UnitRestorable = tier2UnitIsOpenable(unitDir)
|
||||
// R-403: ask the package itself when it was made. Reading the artifact rather than the status
|
||||
// record is what makes this date impossible to overstate.
|
||||
c.UnitPackageDate = unitPackageDate(unitDir)
|
||||
return c
|
||||
}
|
||||
|
||||
@@ -144,6 +160,7 @@ func (m *Manager) Tier2RestoreCoverage(stackName string) (Tier2Coverage, error)
|
||||
if m.settings != nil {
|
||||
if cfg := m.settings.GetCrossDriveConfig(stackName); cfg != nil {
|
||||
cov.CopyLastRun, cov.CopyLastSuccess = cfg.LastRun, cfg.LastSuccess
|
||||
cov.UnitLegPreserved = cfg.UnitLegSkipped
|
||||
}
|
||||
}
|
||||
return cov, nil
|
||||
@@ -178,7 +195,79 @@ func (m *Manager) RestoreTier2Unit(stackName string) (UnitRestoreResult, error)
|
||||
// path and never a secret, and naming it is what makes "the SECONDARY mirror was the source" a
|
||||
// positive observable in the log rather than an absence to be argued from.
|
||||
m.logger.Printf("[WARN] [backup] Tier-2 UNIT restore for %s from the secondary mirror %s — this OVERWRITES live app data", stackName, unitDir)
|
||||
return m.RestoreFromRecoveryUnitAt(stackName, unitDir)
|
||||
res, restoreErr := m.RestoreFromRecoveryUnitAt(stackName, unitDir)
|
||||
|
||||
// R-403, the CAUSE half. Refill the primary unit from the mirror we just restored from, INSIDE
|
||||
// this call, before it returns.
|
||||
//
|
||||
// THE TIMING IS THE REQUIREMENT, NOT A DETAIL. On 2026-08-31 the hollow primary manifest was
|
||||
// written TWO SECONDS after a restore of exactly this shape, by the 5-minute `backup-cache` job
|
||||
// (`backup.go` → `captureAllRecoveryUnits`). Any follow-up job, scheduled refresh or goroutine
|
||||
// races that capture and can lose. Doing it here is the only shape that cannot.
|
||||
//
|
||||
// The capture itself is NOT guarded and must not be: a capture that describes an empty drive as
|
||||
// empty is CORRECT. With the primary refilled there is no hollow state left for it to describe,
|
||||
// which is why the fix is here and not there. Guarding the capture would make the manifest lie.
|
||||
m.rehydratePrimaryUnit(stackName, unitDir, restoreErr)
|
||||
return res, restoreErr
|
||||
}
|
||||
|
||||
// rehydratePrimaryUnit copies a mirrored recovery unit back onto the app's own drive when the primary
|
||||
// unit is ABSENT or HOLLOW — the state a Tier-2 unit restore leaves behind, and the state that armed
|
||||
// R-403's delete on the following night.
|
||||
//
|
||||
// Three refusals, each earned:
|
||||
// - the restore FAILED → write nothing. A package written from a run that did not succeed is worse
|
||||
// than no package: it would look like a backup and describe data that never landed.
|
||||
// - the primary already CARRIES DATA → leave it byte-identical. It may be NEWER than the mirror
|
||||
// (the customer restored while their own drive was fine), and overwriting it with an older copy
|
||||
// is the very move this whole task exists to prevent, pointed the other way.
|
||||
// - anything goes wrong copying → WARN and carry on. The restore itself succeeded; the app is back.
|
||||
// Failing the restore because a convenience copy failed would report a success as a failure.
|
||||
//
|
||||
// It is best-effort by design and says so in the log either way, because an absent log line is not
|
||||
// evidence that it ran.
|
||||
func (m *Manager) rehydratePrimaryUnit(stackName, mirrorUnitDir string, restoreErr error) {
|
||||
if restoreErr != nil {
|
||||
m.logger.Printf("[INFO] [backup] %s: primary unit NOT refilled — the restore itself failed (R-403: a package from a failed run is worse than none)", stackName)
|
||||
return
|
||||
}
|
||||
drivePath := m.GetAppDrivePath(stackName)
|
||||
if drivePath == "" || !filepath.IsAbs(drivePath) {
|
||||
m.logger.Printf("[WARN] [backup] %s: primary unit NOT refilled — cannot resolve the app's drive", stackName)
|
||||
return
|
||||
}
|
||||
primaryUnit := RecoveryUnitPath(m.namespaceRoot(drivePath), stackName)
|
||||
if unitCarriesData(primaryUnit) {
|
||||
m.logger.Printf("[INFO] [backup] %s: primary unit already carries data — left untouched (R-403 never overwrites a richer package with a poorer one)", stackName)
|
||||
return
|
||||
}
|
||||
copier := m.unitRehydrate
|
||||
if copier == nil {
|
||||
copier = rsyncMirror
|
||||
}
|
||||
if err := copier(mirrorUnitDir, primaryUnit); err != nil {
|
||||
m.logger.Printf("[ERROR] [backup] %s: refilling the primary unit from the mirror FAILED: %v — the restore itself SUCCEEDED and the app is running; the local package stays incomplete until the next backup", stackName, err)
|
||||
return
|
||||
}
|
||||
m.logger.Printf("[INFO] [backup] %s: primary unit refilled from the secondary mirror (R-403) — %d volume tar(s), %d database dump(s) now on the app's own drive",
|
||||
stackName, countUnitFiles(UnitVolumeDumpDir(primaryUnit), ".tar"), countUnitFiles(UnitDBDumpDir(primaryUnit), ".sql"))
|
||||
}
|
||||
|
||||
// countUnitFiles counts files with a suffix in a unit leg directory — for the log line only, so the
|
||||
// refill states WHAT it put back rather than merely that it ran. Never a secret: counts, not names.
|
||||
func countUnitFiles(dir, suffix string) int {
|
||||
entries, err := os.ReadDir(dir)
|
||||
if err != nil {
|
||||
return 0
|
||||
}
|
||||
n := 0
|
||||
for _, e := range entries {
|
||||
if !e.IsDir() && strings.HasSuffix(e.Name(), suffix) {
|
||||
n++
|
||||
}
|
||||
}
|
||||
return n
|
||||
}
|
||||
|
||||
// Tier2CopyDate returns the date the surface should name for this app's Tier-2 copy, preferring the
|
||||
@@ -194,6 +283,21 @@ func (c Tier2Coverage) Tier2CopyDate() (date string, proven bool) {
|
||||
return c.CopyLastRun, false
|
||||
}
|
||||
|
||||
// UnitRestoreDate returns the date of the PACKAGE the unit restore would actually open, and whether
|
||||
// that package is OLDER than the copy's newest run.
|
||||
//
|
||||
// R-403. `Tier2CopyDate` answers "when was this copy last written to" and is right for the file
|
||||
// restore, whose legs really were refreshed by that run. It is the WRONG answer for the unit restore
|
||||
// after a preserved leg, because the package is then older than the run that reports success. This
|
||||
// asks the manifest first and falls back to the copy date only when the package cannot say.
|
||||
func (c Tier2Coverage) UnitRestoreDate() (date string, olderThanTheRun bool) {
|
||||
copyDate, _ := c.Tier2CopyDate()
|
||||
if c.UnitPackageDate == "" {
|
||||
return copyDate, c.UnitLegPreserved
|
||||
}
|
||||
return c.UnitPackageDate, c.UnitLegPreserved || (copyDate != "" && c.UnitPackageDate < copyDate)
|
||||
}
|
||||
|
||||
// tier2RecordedCopyDir resolves the RECORDED Tier-2 copy dir for a stack, applying every
|
||||
// source-side refusal in one place so the pre-flight check and the restore itself cannot drift.
|
||||
func (m *Manager) tier2RecordedCopyDir(stackName string) (string, error) {
|
||||
|
||||
Reference in New Issue
Block a user