R-359 + R-397: the off-site store gets checked, and the advertised check becomes real
gates / gates (push) Successful in 12s

Nothing ever verified that the off-site copies are still readable. The
whole-guest tier has verify jobs; the tier holding the customer's documents and
photos had none -- the complete set of restic verbs this controller used
contained no `check`. We would have found out at restore time, with a customer
waiting. On 2026-08-21 a deliberately damaged pack was caught at once by plain
`restic check`; we had never run it.

R-397: NotifyIntegrityOK/NotifyIntegrityFailed existed with no caller, the hub
allowlists both event types and carries the Hungarian text for both, the
settings checkbox exists, and the debug button posts to /api/debug/backup/
integrity. Everything was built except the part that runs. SIXTH instance of
that shape in this project.

THE HAZARD SHAPES THE WHOLE DESIGN. resticStep self-heals a crash lock by
running `unlock --remove-all` and retrying, and its own comment records why that
is safe: every caller holds the in-process single-flight mutex, so any lock it
meets is stale. A check that did not take that flag could meet a LIVE prune's
lock from this same box, remove it, and retry over the top of it. So the check
TAKES THE FLAG and SKIPS rather than waits -- waiting would pin the nightly
backup behind it, and a skip costs nothing because due-ness makes tomorrow try
again. TestR359_SkipsWhenRunningFlagHeld asserts the NON-EFFECTS: restic never
invoked, `unlock` never in any argv. Its red-proof prints the real thing --
restic running `check` while the flag was held.

DUE-NESS, NOT A WEEKDAY. Daily job, weekly behaviour: "is the last successful
check older than 7 days?" not "is it Sunday?". R-341 is exactly the other shape,
a dated check quietly missed and never caught up. No Weekly primitive added.

THREE OUTCOMES, NOT TWO. Skipped, Unreachable and failed are different facts.
"I could not look" is not "I looked and it is broken" -- R-339 already owns
reachability, and a second alarm for the same fact trains the operator to
discount the one alarm that means the backups are damaged. A timeout is
unreachable, never damage. A failure advances due-ness (a broken store must not
be re-checked nightly); a skip and an unreachable store do not.

Success is severity `info`, which severityNotifies DROPS -- it mails NOBODY, by
design. A weekly success e-mail is how people stop reading their alerts.

The customer gets a SENTENCE; restic's words go to the log, truncated (R-379:
615 bytes of raw database text reached a customer once). read-data-subset ships
OFF and a malformed value is refused at read time rather than handed to restic,
where one typo would fail the whole check.

Published on OffboxReportStatus, NOT on report.BackupReport's IntegrityOK --
those were retired by R-331 YESTERDAY and TestBackupReport_DeadFieldsStayZero
still passes unmodified.

Also: the monitoring page stopped promising a Sunday job that never existed, and
the debug button got its dispatch case.

PART 0 WAS NOT BUILT, AND R-398 WAS MY OWN MISTAKE. The seam it asked for
already exists: offboxRunner/SetOffboxRunner/m.runner() has been injectable
since the off-site tier shipped, and other tests drive restic-backed paths
through it. A resticStepFn seam would have been WORSE here -- it would replace
the `unlock --remove-all` escalation and hide it from the assertions that must
see it. R-358's AST ordering test is converted to a real execution test instead,
which immediately surfaced something the AST walk could not: unlockStale
legitimately runs before the restore.

Four red-proofs, each printing the pre-fix behaviour. Green gate: 28 packages,
rc 0. All 12 controller gates OK.
This commit is contained in:
2026-08-30 21:03:29 +02:00
parent e64c84aef8
commit 0d52a42c17
13 changed files with 1365 additions and 52 deletions
+106
View File
@@ -1119,6 +1119,27 @@ func main() {
}
return nil
})
// R-359 — the off-site store gets checked. Nothing verified it before this: the whole-guest
// tier has verify jobs, the tier holding the customer's documents and photos had none, and we
// would have found out at restore time with a customer waiting.
//
// DAILY JOB, WEEKLY BEHAVIOUR, and that is the point. It asks "is the last successful check
// older than the max age?" rather than "is it Sunday?", so a box that was switched off on its
// check day is checked the next day it is on. R-341 is exactly the other shape: a dated check
// quietly missed and never caught up. No `Weekly` primitive was added to the scheduler —
// due-ness is smaller, catches up, and is what R-86 already chose for restore-tests.
//
// 06:00 CHOSEN FROM THE LIVE SCHEDULE, not from a document. Observed on demo-hp 2026-08-30:
// db-dump 02:30, tier2-backup 03:30, fill-watch 03:30, metrics-prune 04:00, offbox-backup
// 04:15, offsite-abandon-sweep 05:10, whole-guest gate [04:30, 08:30). 06:00 is 1h45 clear of
// the off-site leg (which runs ~2m20s, measured) and 50 min clear of the sweep. The backup
// WINDOW is customer-configurable, so no fixed time is collision-proof for every box — but a
// collision costs one skipped day, not a missed check, because due-ness makes tomorrow try
// again. What would be a real defect is a slot that collides EVERY night; this is not one.
sched.Daily("offsite-integrity", "06:00", func(ctx context.Context) error {
runOffsiteIntegrityCheck(ctx, backupMgr, notifier, logger, false)
return nil
})
}
// Metrics prune — daily at 04:00
@@ -1586,6 +1607,13 @@ func main() {
return hubPusher.Push(r)
}
}
// R-359/R-397: the debug button's caller. Same function as the scheduled job, `force=true` the
// only difference — see runOffsiteIntegrityCheck.
dc.RunIntegrityCheck = func(force bool) backup.IntegrityResult {
ctx, cancel := context.WithTimeout(context.Background(), 35*time.Minute)
defer cancel()
return runOffsiteIntegrityCheck(ctx, backupMgr, notifier, logger, force)
}
dc.HubConnectivityTest = func() (int, int64, error) {
start := time.Now()
client := &http.Client{Timeout: 10 * time.Second}
@@ -3069,3 +3097,81 @@ func discoverHDDPaths(stacksDir string, logger *log.Logger) []string {
}
return paths
}
// ── R-359 / R-397 — the integrity check's one caller, shared by the scheduler and the debug button ──
//
// ONE function for both so the hand-run cannot drift from the scheduled run. The only difference
// between them is `force`, which skips the due-ness question; every other guard — the single-writer
// flag above all — applies identically. A hand-run that bypassed the flag "because the operator asked
// for it" is precisely the hazard this whole feature is shaped around.
//
// R-397: `NotifyIntegrityOK` and `NotifyIntegrityFailed` have existed in internal/notify since the
// notifier did, the hub allowlists both event types, and the hub carries the Hungarian customer text
// for both. Everything was built except the caller. That is the SIXTH time this project has found that
// shape, and it is worth naming as a pattern rather than treating as a novelty each time.
func runOffsiteIntegrityCheck(ctx context.Context, mgr *backup.Manager, n *notify.Notifier, logger *log.Logger, force bool) backup.IntegrityResult {
if mgr == nil {
return backup.IntegrityResult{Skipped: true, SkipReason: "backup manager not configured"}
}
if !force {
due, last := mgr.IntegrityDue(time.Now())
if !due {
logger.Printf("[INFO] [offbox] integrity: not due (last successful check %s ago, at %s) — nothing was run",
time.Since(last).Round(time.Minute), last.Format("2006-01-02 15:04"))
return backup.IntegrityResult{Skipped: true, SkipReason: "not due"}
}
}
res := mgr.CheckOffboxIntegrity(ctx)
switch {
case res.Skipped:
// Not a failure and not an alarm. The log already said why, inside the check.
return res
case res.Unreachable:
// "I could not look" is not "I looked and it is broken". Deliberately NO notifier: R-339 owns
// reachability, and a second alarm for the same fact would train the operator to discount the
// one alarm that means the customer's backups are damaged.
return res
case res.OK:
mgr.RecordIntegrityOutcome(res.RanAt, true)
// severity `info`, which severityNotifies DROPS before either leg — so this mails NOBODY, by
// design (08 §6.1). A weekly success e-mail is how people stop reading their alerts. It is
// still pushed, because the hub stores it and the event stream is where "was it checked?" is
// answered.
n.NotifyIntegrityOK(integrityOKMsg(res))
return res
default:
mgr.RecordIntegrityOutcome(res.RanAt, false)
// A FAILING store advances due-ness too: re-checking a broken repository every night is load
// with no new information, and the hourly operator cooldown already governs the mail.
//
// The customer gets a SENTENCE. restic's own words go to the log, truncated, where the operator
// can diagnose without a rebuild. R-379 is the reason that split exists: 615 bytes of raw
// database text reached a customer once.
n.NotifyIntegrityFailed(integrityFailedMsg, "restic check reported repository errors")
return res
}
}
// R-359 customer-facing strings. Named constants because tests assert them verbatim and because a
// silent edit is how an honest message drifts into a comforting one.
const (
// integrityFailedMsg names what to do and what NOT to do. "Ne törölj semmit" is load-bearing: a
// customer who believes their backups are broken may try to "start fresh", which destroys the one
// copy that might still be partly recoverable.
integrityFailedMsg = "A távoli mentés ellenőrzése hibát talált a tárolóban. A mentések egy része sérült lehet. Ne törölj semmit, és vedd fel velünk a kapcsolatot."
integrityOKBase = "A távoli mentés ellenőrzése rendben lezajlott."
)
// integrityOKMsg states what was actually checked, so a structure-only pass is never read as a
// full data verification. The depth is a fact the customer's sentence has to carry: "checked" means
// two different things depending on it.
func integrityOKMsg(res backup.IntegrityResult) string {
msg := integrityOKBase + " (" + res.Duration.Round(time.Second).String()
if res.ReadDataSubset != "" {
msg += ", a mentett adatok " + res.ReadDataSubset + "-át újraolvasva"
}
return msg + ")"
}
@@ -0,0 +1,116 @@
package main
import (
"go/ast"
"go/parser"
"go/token"
"testing"
)
// R-359 — the scheduled job must be PROVEN WIRED.
//
// `func main()` cannot be called from a test, so this walks its AST. A `strings.Contains` would not
// do: a commented-out call still contains the string, and this project has a recorded case of a
// text-based wiring test passing the very red-proof it existed to fail (2026-07-21).
//
// This is the check for the class that has now bitten seven times here — a complete mechanism with no
// caller. R-397 is one instance (the notifiers), R-400 another (the debug button that posted to
// nothing). A job that exists and is never registered would be the third in this one release.
func TestR359_JobIsRegisteredInMain(t *testing.T) {
body := mainBody(t)
var found, named bool
ast.Inspect(body, func(n ast.Node) bool {
call, ok := n.(*ast.CallExpr)
if !ok {
return true
}
sel, ok := call.Fun.(*ast.SelectorExpr)
if !ok || sel.Sel.Name != "Daily" {
return true
}
if len(call.Args) == 0 {
return true
}
lit, ok := call.Args[0].(*ast.BasicLit)
if !ok || lit.Kind != token.STRING {
return true
}
if lit.Value == `"offsite-integrity"` {
found = true
// The schedule argument must be a literal time, not a variable that could resolve to "".
if len(call.Args) > 1 {
if s, ok := call.Args[1].(*ast.BasicLit); ok && s.Kind == token.STRING && len(s.Value) > 2 {
named = true
}
}
}
return true
})
if !found {
t.Fatal("func main() never registers the `offsite-integrity` daily job — the check exists and " +
"nothing would ever run it, which is the built-but-never-wired class this release closes " +
"two instances of")
}
if !named {
t.Error("the job is registered without a literal HH:MM schedule")
}
}
// TestR359_DebugCallbackIsWired pins the other caller. The debug button posted to a route that did not
// exist for its entire life (R-400); a route that exists but whose callback is nil is the same defect
// one layer down, and it would answer „Nem bekötött" forever.
func TestR359_DebugCallbackIsWired(t *testing.T) {
body := mainBody(t)
var wired bool
ast.Inspect(body, func(n ast.Node) bool {
as, ok := n.(*ast.AssignStmt)
if !ok {
return true
}
for _, lhs := range as.Lhs {
sel, ok := lhs.(*ast.SelectorExpr)
if ok && sel.Sel.Name == "RunIntegrityCheck" {
wired = true
}
}
return true
})
if !wired {
t.Fatal("DebugCallbacks.RunIntegrityCheck is never assigned in func main() — the debug route " +
"would answer „Nem bekötött" + " forever, which is exactly the shape of the button this " +
"release is fixing")
}
}
func TestR359_OutcomeMessagesCarryNoMachineDetail(t *testing.T) {
// The customer gets a sentence; restic's words go to the log. R-379: 615 bytes of raw database
// text reached a customer once.
for _, bad := range []string{"sftp:", "restic", "exit status", "@", "/srv/"} {
if contains(integrityFailedMsg, bad) {
t.Errorf("the failure sentence carries machine detail %q: %q", bad, integrityFailedMsg)
}
}
// It must still tell them what to do — and what NOT to do.
for _, want := range []string{"Ne törölj semmit", "vedd fel velünk a kapcsolatot"} {
if !contains(integrityFailedMsg, want) {
t.Errorf("the failure sentence is a dead end; missing %q", want)
}
}
}
func contains(s, sub string) bool {
return len(sub) > 0 && len(s) >= len(sub) && (func() bool {
for i := 0; i+len(sub) <= len(s); i++ {
if s[i:i+len(sub)] == sub {
return true
}
}
return false
})()
}
var _ = parser.ParseFile
var _ = token.NewFileSet