v0.231.0 - the box proves its own off-site copy still holds something (R-87)
gates / gates (push) Successful in 11s
gates / gates (push) Successful in 11s
R-87 re-scoped by its own spike and built as Option C. MinAgent 0.129.0 unchanged. THE QUESTION NOTHING ASKED. The weekly check proves the stored bytes are the bytes we stored; it cannot tell us we stored the WRONG thing. A hollow recovery unit backs up cleanly, checks cleanly at 100 percent depth, restores cleanly and gives the customer nothing back - measured on demo-hp 2026-08-31, 120082104 B to 7036 B in one nightly run recorded as a success (R-403). No tier and no cadence asked it. Now offsite-proof does, nightly, on one app. IT DOES NOT prove a restore puts data back into a running app. That stays drill work and 07 section 8 matrix row 4 is NOT moved. THE ACCEPTANCE RULE HAS TWO PARTS AND THE OBVIOUS ONE IS A TRAP. "Check the unit against its own packing list" PASSES a hollow unit, because a hollow unit declares nothing. So: (1) everything declared is present, AND (2) the manifest declares what the app is supposed to have. Part 2 is the whole value. RED-PROOFED: the naive rule makes the hollow-unit test read verdict "pass". THE EXPECTATION COMES FROM INSIDE THE UNIT, never the live box - the snapshot may predate the app's shape, and GetDockerVolumes describes the running app. Database half is DBServiceNames, the same discriminator RestoreFromRecoveryUnit uses. Volume half is ParseComposeNamedVolumes as an EXISTENCE check, not a name match: tars are <project>_<volume>.tar and ResolveDockerVolumeNames derives the project from the compose file's parent dir, which inside a unit is the literal string "compose". Measured on all eight real units on demo-hp the counts match exactly and the naming held every time - but "held on eight" is not "derivable" (R-355). Half a rule that is true beats a whole rule that is invented. THREE OUTCOMES: pass, fail (readable and empty), cannot judge. An app that legitimately has neither a database nor volumes PASSES. RED-PROOFED: alarming on any empty unit makes that test read verdict "fail". IT NEVER WRITES TO THE REPOSITORY and that is asserted on the ARGV as a non-effect: --no-lock, no unlockStale, and m.runner() rather than resticStep so the unlock --remove-all escalation is unreachable. RED-PROOFED: routing it the customer path's way makes the test fail on "unlock" appearing in the argv. IT TAKES acquireRunning ITSELF and skips rather than waits, because RestoreOffboxScratch does not take it (R-408) while offbox_integrity.go states that invariant as universal. DUE-NESS IS PER SNAPSHOT (R-86's model), never per clock. RED-PROOFED: recording a timestamp fails the stored-value test AND breaks the rotation - night 2 re-picks night 1's app. ITS SCRATCH IS A SEPARATE ROOT (backups/offsite-proof) and that is a safety decision, not tidiness: the job deletes its copy on every path, and sharing backups/offsite-restore/<app> would mean a nightly background job deleting the verification copy a CUSTOMER is looking at. It is also invisible to placement, so a proof copy can never be pushed into a live app. SHARED RATHER THAN FORKED: offboxScratchDirIn parameterises the scratch resolver on its ROOT builder, and unitOnlyHeadroom extracts the free-space gate, so the customer path and the proof refuse at the same floor with the same Hungarian sentence. RestoreOffboxScratch's behaviour is unchanged. NEW EVENT offsite_proof_empty, severity error, operator-only - deliberately NOT backup_integrity_failed, whose hub template says the store is DAMAGED. Here the store is sound and the content is absent: different cause, different action. The hub half shipped FIRST, in felhom.eu 1aeaa30 (hub v0.110.0, live and verified), because an unallowlisted type is 400'd and vanishes. 33 new tests, all groups green; full suite 1689 tests, 28 packages, rc=0. All 13 controller gates OK. Five red-proofs run and recorded in REPORT.md. A golden carrying 0.231.0 is OWED - the fleet is on 0.230.0. Viktor's call (R-242).
This commit is contained in:
@@ -1140,6 +1140,26 @@ func main() {
|
||||
runOffsiteIntegrityCheck(ctx, backupMgr, notifier, logger, false)
|
||||
return nil
|
||||
})
|
||||
|
||||
// R-87 — the nightly off-site PROOF: one app, its newest snapshot, restored read-only and
|
||||
// judged, then deleted.
|
||||
//
|
||||
// 05:30 CHOSEN FROM THE LIVE SCHEDULE, read off demo-hp on 2026-08-31 and not from a
|
||||
// document: db-dump 02:30, tier2-backup 03:30, offbox-backup 04:15 (2m52s measured),
|
||||
// offsite-abandon-sweep 05:10, offsite-integrity 06:00 (40.3s measured at 100% depth). 05:30
|
||||
// sits in the measured empty gap — 20 min after the sweep starts and 30 min before the
|
||||
// integrity check, whose flag it shares.
|
||||
//
|
||||
// THE WHOLE RUN IS SECONDS, so the slot has room it does not need: all eight apps back to
|
||||
// back measured 25 s on this box, and this job does ONE. The backup WINDOW is
|
||||
// customer-configurable, so no fixed time is collision-proof on every box — but a collision
|
||||
// costs one skipped day rather than a missed proof, because due-ness is per SNAPSHOT and
|
||||
// tomorrow tries the same app again. That is the identical argument the integrity job's own
|
||||
// slot comment makes, and it is the reason skip-if-busy is the right policy here too.
|
||||
sched.Daily("offsite-proof", "05:30", func(ctx context.Context) error {
|
||||
runOffsiteProof(ctx, backupMgr, notifier, logger)
|
||||
return nil
|
||||
})
|
||||
}
|
||||
|
||||
// Metrics prune — daily at 04:00
|
||||
@@ -3165,6 +3185,66 @@ const (
|
||||
integrityOKBase = "A távoli mentés ellenőrzése rendben lezajlott."
|
||||
)
|
||||
|
||||
// ── R-87 — the nightly off-site PROOF's one caller ─────────────────────────────────────────────
|
||||
//
|
||||
// ONE function, like runOffsiteIntegrityCheck above, so a future debug button cannot drift from the
|
||||
// scheduled run. Every guard lives inside `ProveOffboxUnit`; this owns only what happens to the
|
||||
// verdict afterwards, because only the caller knows whether a verdict counts.
|
||||
//
|
||||
// THE FOUR NON-ALARMING OUTCOMES ARE NOT FAILURES and none of them notifies:
|
||||
// - Skipped — it yielded to a running backup. Correct behaviour; alarming would punish it.
|
||||
// - NoSnapshot — nothing is due. Not a fact about any backup.
|
||||
// - Err — the restore did not finish, so it SAW NOTHING and may not claim anything about
|
||||
// the backup. This is the same separation `IntegrityResult` draws between a failed
|
||||
// check and an unreachable store, and it is why a download error can never become
|
||||
// the "intact but empty" alarm.
|
||||
// - CannotJudge — recorded, never alarmed. Alarming on our own blind spot trains the operator to
|
||||
// discount the one alarm that means the customer's backup holds nothing.
|
||||
func runOffsiteProof(ctx context.Context, mgr *backup.Manager, n *notify.Notifier, logger *log.Logger) backup.ProofResult {
|
||||
if mgr == nil {
|
||||
return backup.ProofResult{Skipped: true, SkipReason: "backup manager not configured"}
|
||||
}
|
||||
res := mgr.ProveOffboxUnit(ctx)
|
||||
if res.Verdict() == "" {
|
||||
return res // skipped / nothing due / restore error — no verdict was reached
|
||||
}
|
||||
mgr.RecordProofVerdict(res)
|
||||
if res.Judgement.Verdict == backup.UnitProofFail {
|
||||
// ONE event. The customer-grade sentence goes in the message; the machine detail goes in the
|
||||
// detail field, where the operator can diagnose without a rebuild — the R-379 split.
|
||||
n.NotifyOffsiteProofEmpty(proofEmptyMsg(res.Stack), proofEmptyDetail(res))
|
||||
}
|
||||
if logger != nil && res.Judgement.Verdict == backup.UnitProofCannotJudge {
|
||||
logger.Printf("[WARN] [offbox] proof: %s could not be judged (%s) — recorded, deliberately NOT alarmed",
|
||||
res.Stack, res.Judgement.Reason)
|
||||
}
|
||||
return res
|
||||
}
|
||||
|
||||
// proofEmptyMsg is the R-87 customer-grade sentence. A named function because tests assert it
|
||||
// verbatim and because a silent edit is how an honest message drifts into a comforting one.
|
||||
//
|
||||
// IT MUST SAY THE TWO THINGS THAT MATTER AND NOT A THIRD. The backup is READABLE — the store is not
|
||||
// damaged, and saying so is load-bearing, because "sérült" (damaged) sends the operator and the
|
||||
// customer down the R-359 path, which is a different fault with a different remedy and a standing
|
||||
// instruction not to delete anything. And it names the ONE action that helps: the backup has to be
|
||||
// made again, which is a capture-side fix.
|
||||
func proofEmptyMsg(stack string) string {
|
||||
return "A(z) " + stack + " legutóbbi távoli mentése olvasható, de nem tartalmazza az alkalmazás adatait. " +
|
||||
"A tároló nem sérült — a mentés készült el üresen. A mentést újra el kell készíteni; addig ebből a mentésből nem lehet visszaállítani."
|
||||
}
|
||||
|
||||
// proofEmptyDetail is the OPERATOR half: the machine reason code, the snapshot it was proved on, and
|
||||
// what was expected. Never a file's content and never a path outside the unit — units carry portable
|
||||
// secrets.
|
||||
func proofEmptyDetail(res backup.ProofResult) string {
|
||||
d := "snapshot " + res.SnapshotID + ", reason " + string(res.Judgement.Reason)
|
||||
if len(res.Judgement.Missing) > 0 {
|
||||
d += ", expected: " + strings.Join(res.Judgement.Missing, " ")
|
||||
}
|
||||
return d
|
||||
}
|
||||
|
||||
// integrityOKMsg states what was actually checked, so a structure-only pass is never read as a
|
||||
// full data verification. The depth is a fact the customer's sentence has to carry: "checked" means
|
||||
// two different things depending on it.
|
||||
|
||||
Reference in New Issue
Block a user