v0.231.0 - the box proves its own off-site copy still holds something (R-87)
gates / gates (push) Successful in 11s

R-87 re-scoped by its own spike and built as Option C. MinAgent 0.129.0 unchanged.

THE QUESTION NOTHING ASKED. The weekly check proves the stored bytes are the bytes we
stored; it cannot tell us we stored the WRONG thing. A hollow recovery unit backs up
cleanly, checks cleanly at 100 percent depth, restores cleanly and gives the customer
nothing back - measured on demo-hp 2026-08-31, 120082104 B to 7036 B in one nightly run
recorded as a success (R-403). No tier and no cadence asked it. Now offsite-proof does,
nightly, on one app.

IT DOES NOT prove a restore puts data back into a running app. That stays drill work and
07 section 8 matrix row 4 is NOT moved.

THE ACCEPTANCE RULE HAS TWO PARTS AND THE OBVIOUS ONE IS A TRAP. "Check the unit against
its own packing list" PASSES a hollow unit, because a hollow unit declares nothing. So:
(1) everything declared is present, AND (2) the manifest declares what the app is supposed
to have. Part 2 is the whole value. RED-PROOFED: the naive rule makes the hollow-unit test
read verdict "pass".

THE EXPECTATION COMES FROM INSIDE THE UNIT, never the live box - the snapshot may predate
the app's shape, and GetDockerVolumes describes the running app. Database half is
DBServiceNames, the same discriminator RestoreFromRecoveryUnit uses. Volume half is
ParseComposeNamedVolumes as an EXISTENCE check, not a name match: tars are
<project>_<volume>.tar and ResolveDockerVolumeNames derives the project from the compose
file's parent dir, which inside a unit is the literal string "compose". Measured on all
eight real units on demo-hp the counts match exactly and the naming held every time - but
"held on eight" is not "derivable" (R-355). Half a rule that is true beats a whole rule
that is invented.

THREE OUTCOMES: pass, fail (readable and empty), cannot judge. An app that legitimately
has neither a database nor volumes PASSES. RED-PROOFED: alarming on any empty unit makes
that test read verdict "fail".

IT NEVER WRITES TO THE REPOSITORY and that is asserted on the ARGV as a non-effect:
--no-lock, no unlockStale, and m.runner() rather than resticStep so the unlock --remove-all
escalation is unreachable. RED-PROOFED: routing it the customer path's way makes the test
fail on "unlock" appearing in the argv.

IT TAKES acquireRunning ITSELF and skips rather than waits, because RestoreOffboxScratch
does not take it (R-408) while offbox_integrity.go states that invariant as universal.

DUE-NESS IS PER SNAPSHOT (R-86's model), never per clock. RED-PROOFED: recording a
timestamp fails the stored-value test AND breaks the rotation - night 2 re-picks night 1's
app.

ITS SCRATCH IS A SEPARATE ROOT (backups/offsite-proof) and that is a safety decision, not
tidiness: the job deletes its copy on every path, and sharing backups/offsite-restore/<app>
would mean a nightly background job deleting the verification copy a CUSTOMER is looking
at. It is also invisible to placement, so a proof copy can never be pushed into a live app.

SHARED RATHER THAN FORKED: offboxScratchDirIn parameterises the scratch resolver on its
ROOT builder, and unitOnlyHeadroom extracts the free-space gate, so the customer path and
the proof refuse at the same floor with the same Hungarian sentence. RestoreOffboxScratch's
behaviour is unchanged.

NEW EVENT offsite_proof_empty, severity error, operator-only - deliberately NOT
backup_integrity_failed, whose hub template says the store is DAMAGED. Here the store is
sound and the content is absent: different cause, different action. The hub half shipped
FIRST, in felhom.eu 1aeaa30 (hub v0.110.0, live and verified), because an unallowlisted
type is 400'd and vanishes.

33 new tests, all groups green; full suite 1689 tests, 28 packages, rc=0. All 13 controller
gates OK. Five red-proofs run and recorded in REPORT.md.

A golden carrying 0.231.0 is OWED - the fleet is on 0.230.0. Viktor's call (R-242).
This commit is contained in:
2026-08-31 20:55:34 +02:00
parent 2d802d75e8
commit e43b5ec07d
16 changed files with 1973 additions and 33 deletions
+22 -5
View File
@@ -397,6 +397,23 @@ func (n *Notifier) NotifyIntegrityFailed(message, errMsg string) {
n.PushEvent("backup_integrity_failed", "error", message, &BackupDetails{Error: errMsg})
}
// NotifyOffsiteProofEmpty (R-87) reports that the nightly off-site proof found a backup that is
// READABLE and holds none of the app's data.
//
// SEVERITY `error`, from the hub's exact vocabulary {info, warning, error, critical}. Anything else
// is silently coerced to `info` and mailed to nobody — that shipped twice (R-328 on
// `disk_health_degraded`, R-329 on `app_start_failed`, 91 events stored and zero delivered), and
// `r329_severity_contract_test.go` walks this file to keep it from shipping a third time.
//
// A SEPARATE TYPE FROM `backup_integrity_failed`, and that is the point rather than an oversight: the
// integrity check says THE STORE IS DAMAGED; this says the store is sound and the CONTENT is absent.
// The customer's action differs and so must the sentence.
//
// The message is passed through, not templated hub-side, so it can name the app and what is missing.
func (n *Notifier) NotifyOffsiteProofEmpty(message, detail string) {
n.PushEvent("offsite_proof_empty", "error", message, &BackupDetails{Error: detail})
}
// NotifyIntegrityOK sends a backup integrity check success event.
func (n *Notifier) NotifyIntegrityOK(message string) {
n.PushEvent("backup_integrity_ok", "info", message, nil)
@@ -577,11 +594,11 @@ type DiskHealthDetails struct {
type DiskAlertKind int
const (
DiskAlertWarn DiskAlertKind = iota // Figyelmeztetés — worth keeping an eye on
DiskAlertFailSelfReported // Hiba — the drive's own SMART verdict says FAILING
DiskAlertFailSectors // Hiba — reached from unreadable-sector counters
DiskAlertFailTemperature // Hiba — reached from heat
DiskAlertFailWorsened // Hiba — already reported, and still getting worse
DiskAlertWarn DiskAlertKind = iota // Figyelmeztetés — worth keeping an eye on
DiskAlertFailSelfReported // Hiba — the drive's own SMART verdict says FAILING
DiskAlertFailSectors // Hiba — reached from unreadable-sector counters
DiskAlertFailTemperature // Hiba — reached from heat
DiskAlertFailWorsened // Hiba — already reported, and still getting worse
)
// DiskAlert is the payload for one disk-health alert. It carries enough for the notifier to pick a