Commit Graph

2 Commits

Author SHA1 Message Date
admin e43b5ec07d v0.231.0 - the box proves its own off-site copy still holds something (R-87)
gates / gates (push) Successful in 11s
R-87 re-scoped by its own spike and built as Option C. MinAgent 0.129.0 unchanged.

THE QUESTION NOTHING ASKED. The weekly check proves the stored bytes are the bytes we
stored; it cannot tell us we stored the WRONG thing. A hollow recovery unit backs up
cleanly, checks cleanly at 100 percent depth, restores cleanly and gives the customer
nothing back - measured on demo-hp 2026-08-31, 120082104 B to 7036 B in one nightly run
recorded as a success (R-403). No tier and no cadence asked it. Now offsite-proof does,
nightly, on one app.

IT DOES NOT prove a restore puts data back into a running app. That stays drill work and
07 section 8 matrix row 4 is NOT moved.

THE ACCEPTANCE RULE HAS TWO PARTS AND THE OBVIOUS ONE IS A TRAP. "Check the unit against
its own packing list" PASSES a hollow unit, because a hollow unit declares nothing. So:
(1) everything declared is present, AND (2) the manifest declares what the app is supposed
to have. Part 2 is the whole value. RED-PROOFED: the naive rule makes the hollow-unit test
read verdict "pass".

THE EXPECTATION COMES FROM INSIDE THE UNIT, never the live box - the snapshot may predate
the app's shape, and GetDockerVolumes describes the running app. Database half is
DBServiceNames, the same discriminator RestoreFromRecoveryUnit uses. Volume half is
ParseComposeNamedVolumes as an EXISTENCE check, not a name match: tars are
<project>_<volume>.tar and ResolveDockerVolumeNames derives the project from the compose
file's parent dir, which inside a unit is the literal string "compose". Measured on all
eight real units on demo-hp the counts match exactly and the naming held every time - but
"held on eight" is not "derivable" (R-355). Half a rule that is true beats a whole rule
that is invented.

THREE OUTCOMES: pass, fail (readable and empty), cannot judge. An app that legitimately
has neither a database nor volumes PASSES. RED-PROOFED: alarming on any empty unit makes
that test read verdict "fail".

IT NEVER WRITES TO THE REPOSITORY and that is asserted on the ARGV as a non-effect:
--no-lock, no unlockStale, and m.runner() rather than resticStep so the unlock --remove-all
escalation is unreachable. RED-PROOFED: routing it the customer path's way makes the test
fail on "unlock" appearing in the argv.

IT TAKES acquireRunning ITSELF and skips rather than waits, because RestoreOffboxScratch
does not take it (R-408) while offbox_integrity.go states that invariant as universal.

DUE-NESS IS PER SNAPSHOT (R-86's model), never per clock. RED-PROOFED: recording a
timestamp fails the stored-value test AND breaks the rotation - night 2 re-picks night 1's
app.

ITS SCRATCH IS A SEPARATE ROOT (backups/offsite-proof) and that is a safety decision, not
tidiness: the job deletes its copy on every path, and sharing backups/offsite-restore/<app>
would mean a nightly background job deleting the verification copy a CUSTOMER is looking
at. It is also invisible to placement, so a proof copy can never be pushed into a live app.

SHARED RATHER THAN FORKED: offboxScratchDirIn parameterises the scratch resolver on its
ROOT builder, and unitOnlyHeadroom extracts the free-space gate, so the customer path and
the proof refuse at the same floor with the same Hungarian sentence. RestoreOffboxScratch's
behaviour is unchanged.

NEW EVENT offsite_proof_empty, severity error, operator-only - deliberately NOT
backup_integrity_failed, whose hub template says the store is DAMAGED. Here the store is
sound and the content is absent: different cause, different action. The hub half shipped
FIRST, in felhom.eu 1aeaa30 (hub v0.110.0, live and verified), because an unallowlisted
type is 400'd and vanishes.

33 new tests, all groups green; full suite 1689 tests, 28 packages, rc=0. All 13 controller
gates OK. Five red-proofs run and recorded in REPORT.md.

A golden carrying 0.231.0 is OWED - the fleet is on 0.230.0. Viktor's call (R-242).
2026-08-31 20:55:34 +02:00
admin 2358e561b7 R-403: a poorer copy must never delete a richer one
gates / gates (push) Successful in 11s
MEASURED FIRST, then fixed. On the shipped v0.229.0, on demo-hp, an app's Tier-2 copy went from
120 082 104 B (4 database dumps + 3 named-volume tars) to 7 036 B (none of either) in ONE nightly
run, and the run recorded itself a success: 'Tier 2 copied docmost -> ... (14.9 KB, 0 leg(s), 0s)'.
Evidence: felhom.eu/documentation/audits/DRILL-r403-tier2-delete-2026-08-31/.

The mechanism was three individually-correct lines: RunTier2 guards the unit leg with os.Stat only
(does the folder exist), rsyncMirror is rsync -a --delete, and nothing between them compared source
to destination. An EMPTY unit is a folder that exists.

THE GUARD. One predicate, unitCarriesData/unitIsHollow (r403_hollow.go), asking the MANIFEST and
never the byte size - a big compose tree with no dumps is dangerous, a tiny unit for a tiny app is
fine. Fail closed on an absent or unparseable manifest. RunTier2 skips the unit leg when the source
is hollow AND the destination is not; the other legs still run, the run is not failed, and the skip
is recorded for the SURFACE (CrossDriveBackup.UnitLegSkipped + UnitPackageDate) as well as logged.

--delete STAYS and shrinking stays legal. 07 section 8 row 5's derived-copy rule is unchanged; the
fence is exactly one shape. TestR403_DataLegShrinkIsUnaffected is the guard on the guard.

THE HONESTY. A preserved package is older than the run that preserved it, so the card carries a
notice and the unit-restore confirm names the PACKAGE's date - read from the mirrored manifest's own
created_at, not from the status record - plus a clause saying why it is older.

THE CAUSE. RestoreTier2Unit now refills a hollow or absent primary unit from the mirror it just
restored from, INSIDE the call before returning. The hollow manifest was written two seconds after
a restore by the 5-minute capture job; any follow-up job races it. The capture itself is NOT guarded:
a capture describing an empty drive as empty is correct, and with the primary refilled there is no
hollow state left to describe. Never over a complete primary, never after a failed restore.

recordTier2Success and tier2UnitConfirmMsg keep their old signatures as thin callers, so no existing
test needed editing. New seam unitRehydrate, separate from tier2Mirror on purpose.

22 new Go tests. Red-proofs run and reverted: A6 (predicate -> size threshold), B1 (guard removed ->
the copy's 3 files are DELETED and the seam is called), B6 (a general never-shrink rule -> the shrink
case fails), C2 (only-when-hollow dropped -> the complete primary is overwritten).
2026-08-31 14:02:13 +02:00