MEASURED FIRST, then fixed. On the shipped v0.229.0, on demo-hp, an app's Tier-2 copy went from 120 082 104 B (4 database dumps + 3 named-volume tars) to 7 036 B (none of either) in ONE nightly run, and the run recorded itself a success: 'Tier 2 copied docmost -> ... (14.9 KB, 0 leg(s), 0s)'. Evidence: felhom.eu/documentation/audits/DRILL-r403-tier2-delete-2026-08-31/. The mechanism was three individually-correct lines: RunTier2 guards the unit leg with os.Stat only (does the folder exist), rsyncMirror is rsync -a --delete, and nothing between them compared source to destination. An EMPTY unit is a folder that exists. THE GUARD. One predicate, unitCarriesData/unitIsHollow (r403_hollow.go), asking the MANIFEST and never the byte size - a big compose tree with no dumps is dangerous, a tiny unit for a tiny app is fine. Fail closed on an absent or unparseable manifest. RunTier2 skips the unit leg when the source is hollow AND the destination is not; the other legs still run, the run is not failed, and the skip is recorded for the SURFACE (CrossDriveBackup.UnitLegSkipped + UnitPackageDate) as well as logged. --delete STAYS and shrinking stays legal. 07 section 8 row 5's derived-copy rule is unchanged; the fence is exactly one shape. TestR403_DataLegShrinkIsUnaffected is the guard on the guard. THE HONESTY. A preserved package is older than the run that preserved it, so the card carries a notice and the unit-restore confirm names the PACKAGE's date - read from the mirrored manifest's own created_at, not from the status record - plus a clause saying why it is older. THE CAUSE. RestoreTier2Unit now refills a hollow or absent primary unit from the mirror it just restored from, INSIDE the call before returning. The hollow manifest was written two seconds after a restore by the 5-minute capture job; any follow-up job races it. The capture itself is NOT guarded: a capture describing an empty drive as empty is correct, and with the primary refilled there is no hollow state left to describe. Never over a complete primary, never after a failed restore. recordTier2Success and tier2UnitConfirmMsg keep their old signatures as thin callers, so no existing test needed editing. New seam unitRehydrate, separate from tier2Mirror on purpose. 22 new Go tests. Red-proofs run and reverted: A6 (predicate -> size threshold), B1 (guard removed -> the copy's 3 files are DELETED and the seam is called), B6 (a general never-shrink rule -> the shrink case fails), C2 (only-when-hollow dropped -> the complete primary is overwritten).
This commit is contained in:
@@ -0,0 +1,50 @@
|
||||
package backup
|
||||
|
||||
// R-403 — a POORER copy must never delete a RICHER one.
|
||||
//
|
||||
// Measured on demo-hp 2026-08-31, on the shipped v0.229.0: an app's Tier-2 copy went from
|
||||
// 120 082 104 bytes (4 database dumps + 3 named-volume tars) to 7 036 bytes (none of either) in one
|
||||
// nightly run, and the run recorded itself as a SUCCESS — `Tier 2 copied docmost → … (14.9 KB,
|
||||
// 0 leg(s), 0s)`. Evidence: `felhom.eu/documentation/audits/DRILL-r403-tier2-delete-2026-08-31/`.
|
||||
//
|
||||
// THE MECHANISM, in three lines of existing code that were each individually correct:
|
||||
// 1. `RunTier2` guards the unit leg with `os.Stat(unitDir)` — *does the folder exist*.
|
||||
// 2. `rsyncMirror` is `rsync -a --delete` — an exact mirror, which is what a derived copy must be.
|
||||
// 3. Nothing between them compares the source to the destination.
|
||||
// An EMPTY recovery unit is a folder that exists. So a primary unit that had lost its dumps — after a
|
||||
// restore, a failed dump run, a crash mid-capture, a remount — was mirrored over a complete copy, and
|
||||
// `--delete` removed the customer's last surviving package.
|
||||
//
|
||||
// WHAT THIS FILE DELIBERATELY DOES **NOT** DO: it does not make the Tier-2 copy un-shrinkable.
|
||||
// `07-backup-architecture.md` §8 row 5 records that the secondary is a DERIVED copy, rebuilt on the
|
||||
// next run ("Migration = rebuild, not preserve"), and `tier2.go`'s own header records that a
|
||||
// classified app's copy legitimately shrinks as `export` drops out of its class set. Fencing
|
||||
// shrinkage would be calling a decision a defect. The fence here is exactly one shape: a source that
|
||||
// carries NO data replacing a destination that carries some.
|
||||
|
||||
// unitCarriesData reports whether a recovery-unit DIRECTORY holds RECOVERABLE DATA — the app's
|
||||
// database dumps or its named-volume tars.
|
||||
//
|
||||
// IT ASKS THE MANIFEST, NEVER THE BYTE SIZE, and that is the whole design of the predicate. A unit
|
||||
// with a large compose tree and no dumps is dangerous; a tiny unit belonging to a tiny app is fine.
|
||||
// Size answers "how big", and the question here is "is there anything to recover". `dirSizeBytes`
|
||||
// exists two files away and would have been the obvious wrong answer — TestR403_SizeIsNeverConsulted
|
||||
// is the guard that keeps it out.
|
||||
//
|
||||
// FAIL CLOSED on an absent or unparseable manifest: `readManifest` returns nil for both, and a unit
|
||||
// whose manifest cannot be read is a unit whose contents cannot be vouched for. Treating it as
|
||||
// data-bearing would let an unreadable source authorise a delete.
|
||||
//
|
||||
// ONE predicate, every caller. The mirror guard and the post-restore rehydrate both ask this
|
||||
// function; two copies of the definition is how the two halves of a fix drift apart.
|
||||
func unitCarriesData(unitDir string) bool {
|
||||
man := readManifest(UnitManifestFile(unitDir))
|
||||
if man == nil {
|
||||
return false
|
||||
}
|
||||
return len(man.DBDumps) > 0 || len(man.VolumeDumps) > 0
|
||||
}
|
||||
|
||||
// unitIsHollow is `unitCarriesData` negated, named for the way both callers actually ask it. It is a
|
||||
// separate function only so the call sites read as the question they are asking.
|
||||
func unitIsHollow(unitDir string) bool { return !unitCarriesData(unitDir) }
|
||||
Reference in New Issue
Block a user