R-359 + R-397: the off-site store gets checked, and the advertised check becomes real
gates / gates (push) Successful in 12s

Nothing ever verified that the off-site copies are still readable. The
whole-guest tier has verify jobs; the tier holding the customer's documents and
photos had none -- the complete set of restic verbs this controller used
contained no `check`. We would have found out at restore time, with a customer
waiting. On 2026-08-21 a deliberately damaged pack was caught at once by plain
`restic check`; we had never run it.

R-397: NotifyIntegrityOK/NotifyIntegrityFailed existed with no caller, the hub
allowlists both event types and carries the Hungarian text for both, the
settings checkbox exists, and the debug button posts to /api/debug/backup/
integrity. Everything was built except the part that runs. SIXTH instance of
that shape in this project.

THE HAZARD SHAPES THE WHOLE DESIGN. resticStep self-heals a crash lock by
running `unlock --remove-all` and retrying, and its own comment records why that
is safe: every caller holds the in-process single-flight mutex, so any lock it
meets is stale. A check that did not take that flag could meet a LIVE prune's
lock from this same box, remove it, and retry over the top of it. So the check
TAKES THE FLAG and SKIPS rather than waits -- waiting would pin the nightly
backup behind it, and a skip costs nothing because due-ness makes tomorrow try
again. TestR359_SkipsWhenRunningFlagHeld asserts the NON-EFFECTS: restic never
invoked, `unlock` never in any argv. Its red-proof prints the real thing --
restic running `check` while the flag was held.

DUE-NESS, NOT A WEEKDAY. Daily job, weekly behaviour: "is the last successful
check older than 7 days?" not "is it Sunday?". R-341 is exactly the other shape,
a dated check quietly missed and never caught up. No Weekly primitive added.

THREE OUTCOMES, NOT TWO. Skipped, Unreachable and failed are different facts.
"I could not look" is not "I looked and it is broken" -- R-339 already owns
reachability, and a second alarm for the same fact trains the operator to
discount the one alarm that means the backups are damaged. A timeout is
unreachable, never damage. A failure advances due-ness (a broken store must not
be re-checked nightly); a skip and an unreachable store do not.

Success is severity `info`, which severityNotifies DROPS -- it mails NOBODY, by
design. A weekly success e-mail is how people stop reading their alerts.

The customer gets a SENTENCE; restic's words go to the log, truncated (R-379:
615 bytes of raw database text reached a customer once). read-data-subset ships
OFF and a malformed value is refused at read time rather than handed to restic,
where one typo would fail the whole check.

Published on OffboxReportStatus, NOT on report.BackupReport's IntegrityOK --
those were retired by R-331 YESTERDAY and TestBackupReport_DeadFieldsStayZero
still passes unmodified.

Also: the monitoring page stopped promising a Sunday job that never existed, and
the debug button got its dispatch case.

PART 0 WAS NOT BUILT, AND R-398 WAS MY OWN MISTAKE. The seam it asked for
already exists: offboxRunner/SetOffboxRunner/m.runner() has been injectable
since the off-site tier shipped, and other tests drive restic-backed paths
through it. A resticStepFn seam would have been WORSE here -- it would replace
the `unlock --remove-all` escalation and hide it from the assertions that must
see it. R-358's AST ordering test is converted to a real execution test instead,
which immediately surfaced something the AST walk could not: unlockStale
legitimately runs before the restore.

Four red-proofs, each printing the pre-fix behaviour. Green gate: 28 packages,
rc 0. All 12 controller gates OK.
This commit is contained in:
2026-08-30 21:03:29 +02:00
parent e64c84aef8
commit 0d52a42c17
13 changed files with 1365 additions and 52 deletions
+18
View File
@@ -1453,6 +1453,22 @@ type OffboxReportStatus struct {
// Absent on a pre-v0.225.0 controller, and a hub MUST degrade to "unknown" there rather than to
// "empty" — absence means the box cannot answer, never that the answer is no.
StatsKnown bool `json:"stats_known,omitempty"`
// LastIntegrityCheck / LastIntegrityOK (R-359) publish the off-site integrity verdict HERE, on the
// object the hub already trusts for this tier — and deliberately NOT on report.BackupReport's
// `IntegrityOK` / `LastIntegrityCheck`.
//
// Those were retired by R-331 on 2026-08-30, one day before this shipped, because the hub's
// operator Backup card rendered them and read `Integrity Unknown` for every customer forever while
// the boxes were backing up normally. They are kept only so historical reports still parse, and are
// pinned zero by TestBackupReport_DeadFieldsStayZero. **Giving them a live value now would
// resurrect a card that was deliberately removed and would break that test.**
//
// TWO FIELDS, following the StatsKnown precedent immediately above and for the identical reason:
// an empty timestamp means NEVER CHECKED, which is not the same news as checked-and-bad, and a
// single bool cannot carry the difference. `omitempty` on the bool means absent-or-false, so the
// timestamp is the field that says whether the bool means anything at all.
LastIntegrityCheck string `json:"last_integrity_check,omitempty"` // RFC3339
LastIntegrityOK bool `json:"last_integrity_ok,omitempty"`
}
// OffsiteStateNeedsCredential is the ONE declared state (v0.199.0, R-204 item 4 / R-193): this box
@@ -1644,6 +1660,8 @@ func (m *Manager) OffboxReportStatus() *OffboxReportStatus {
LastSuccess: t.LastSuccess,
SnapshotCount: t.SnapshotCount, RepoSizeBytes: t.RepoSizeBytes, QuotaGB: t.QuotaGB,
StatsKnown: t.StatsKnown, // R-331 — without it the hub cannot tell "empty" from "unmeasured"
LastIntegrityCheck: t.LastIntegrityCheck, // R-359
LastIntegrityOK: t.LastIntegrityOK,
AbandonPurgeRequested: t.AbandonPurgeRequested, // R-241: declared until the hub drops the package
}
}
@@ -0,0 +1,259 @@
package backup
import (
"context"
"fmt"
"regexp"
"strings"
"time"
"gitea.dooplex.hu/admin/felhom-controller/internal/settings"
)
// ── R-359 — nothing ever checked that the off-site copies are still readable ─────────────────────
//
// The whole-guest tier has verify jobs. The tier holding the customer's documents and photos had
// none: the complete set of restic verbs this controller used was
// `restore, snapshots, backup, unlock, stats, init, forget, prune, cat, config` — no `check`.
// We would have found out at restore time, with a customer waiting.
//
// On 2026-08-21 one packet in a set-aside store was deliberately damaged and plain `restic check`
// caught it at once. That is an empirical result, not an assumption — it is why this row exists.
//
// ── THE HAZARD THAT SHAPES THIS WHOLE FILE ──────────────────────────────────────────────────────
//
// `resticStep` self-heals a crash lock: when a command fails with "repository is already locked" it
// runs **`unlock --remove-all`** and retries once. Its own doc comment records why that is safe —
// *"the in-process single-flight mutex (held by every caller of this method) proves no sibling
// operation is live"*. Every off-site operation takes `acquireRunning` for exactly that reason.
//
// **An integrity check that did not take the flag would break that invariant.** It could meet the
// lock of a `forget --prune` running from this same box, remove it, and retry over the top of a live
// prune. So this check TAKES THE FLAG, and **skips rather than waits** when it cannot get it: waiting
// would pin a nightly backup behind a check, and a skipped check simply runs tomorrow. That is the
// single most important property in this file, and `TestR359_SkipsWhenRunningFlagHeld` asserts it as
// a NON-EFFECT — restic was never invoked, and `unlock` never appeared in any argument list.
// integrityCheckTimeout bounds one check so a hung repository cannot pin the single-writer flag.
//
// CHOSEN, not inherited: a structure check on a store this size is minutes (measured at ~4 s against
// demo-hp's 134 MB store, and it scales with the index rather than the data). 30 minutes is far above
// any plausible structure check and far below "forever", so the failure mode of a wedged SFTP mount is
// a released flag and a retry tomorrow, never a box whose backups stop because a check never returned.
//
// It is deliberately NOT sized for a `--read-data-subset` run, which downloads pack data and can take
// hours. That option ships OFF (R-399); whoever turns it on must revisit this number, and this comment
// is the note that says so.
const integrityCheckTimeout = 30 * time.Minute
// defaultIntegrityMaxAgeDays is the max age of a SUCCESSFUL check before one is due again.
const defaultIntegrityMaxAgeDays = 7
// readDataSubsetRe accepts the forms restic documents for --read-data-subset: "n/m", a percentage
// like "5%", or a size like "50M". Anything else is refused at read time rather than passed through —
// a typo must not fail the whole check, which is what handing restic an unparsed value would do.
var readDataSubsetRe = regexp.MustCompile(`^([0-9]+/[0-9]+|[0-9]+(\.[0-9]+)?%|[0-9]+[KMGT]?)$`)
// IntegrityResult is what one off-site integrity check did, so every surface — the log, the event, the
// report and the debug endpoint — states the same facts from one place.
//
// Skipped is a first-class outcome, not an error. A check that yielded to a running backup did the
// right thing, and reporting it as a failure would alarm on correct behaviour.
//
// Unreachable is separated from a failed check for the reason §8 of the task records and R-185
// established for a tier the box cannot read: **"I could not look" and "I looked and it is broken" are
// different facts** and must not produce the same alarm. R-339 already alarms on reachability; a
// second alarm for the same fact is noise, and worse, it would train the operator to discount the one
// alarm that means the customer's backups are damaged.
type IntegrityResult struct {
RanAt time.Time
Duration time.Duration
ReadDataSubset string // "" = structure only
OK bool
Skipped bool
SkipReason string
Unreachable bool
// Output is truncated restic output for the LOG. It never reaches a customer message — R-379 is
// the reason: 615 bytes of raw database text reached a customer once.
Output string
}
// integrityReadDataSubset resolves the configured subset spec, refusing anything malformed.
//
// Fail-safe direction: an unrecognised value becomes "" (structure check only) with a WARN, never a
// passthrough. Handing restic `--read-data-subset=banana` fails the entire check, which would turn a
// typo in a config file into a store that silently stops being verified.
func (m *Manager) integrityReadDataSubset() string {
spec := strings.TrimSpace(m.cfg.Monitoring.Integrity.ReadDataSubset)
if spec == "" {
return ""
}
if !readDataSubsetRe.MatchString(spec) {
m.logger.Printf("[WARN] [offbox] integrity: read_data_subset %q is not a form restic accepts (n/m, N%%, or a size like 50M) — running the STRUCTURE check only", spec)
return ""
}
return spec
}
// integrityMaxAge returns the configured max age of a successful check, defaulting to 7 days.
func (m *Manager) integrityMaxAge() time.Duration {
d := m.cfg.Monitoring.Integrity.MaxAgeDays
if d <= 0 {
d = defaultIntegrityMaxAgeDays
}
return time.Duration(d) * 24 * time.Hour
}
// IntegrityDue reports whether a check is due, and the age of the last successful one.
//
// DUE-NESS, NOT A WEEKDAY. A job that fires only on Sundays is a job that silently skips a week every
// time the box is off on a Sunday — which is R-341's exact failure, a dated check quietly missed and
// never caught up. Asking "is the last one older than the max age?" catches up on the next day the box
// is running, whatever day that is.
//
// A repository that has never been checked is DUE. That is the fail-safe direction: an unchecked store
// must not read as a fresh one.
func (m *Manager) IntegrityDue(now time.Time) (due bool, last time.Time) {
t := m.settings.GetOffboxTarget()
if t == nil || strings.TrimSpace(t.LastIntegrityCheck) == "" {
return true, time.Time{}
}
parsed, err := time.Parse(time.RFC3339, t.LastIntegrityCheck)
if err != nil {
m.logger.Printf("[WARN] [offbox] integrity: last-check stamp %q does not parse (%v) — treating the store as never checked", t.LastIntegrityCheck, err)
return true, time.Time{}
}
return now.Sub(parsed) >= m.integrityMaxAge(), parsed
}
// recordIntegrityOutcome persists the verdict and advances due-ness.
//
// Called ONLY when a check actually reached a verdict — pass or fail. A FAILING store advances
// due-ness deliberately: re-checking a broken repository every night is load with no new information,
// the hourly operator cooldown already governs the mail, and the failure is already recorded where a
// surface can read it. A skip or an unreachable repository does NOT reach here, so tomorrow tries
// again.
// RecordIntegrityOutcome is the exported entry point; the caller in main.go owns the decision of WHEN
// a verdict counts, because only it knows whether the run was forced or scheduled.
func (m *Manager) RecordIntegrityOutcome(at time.Time, ok bool) { m.recordIntegrityOutcome(at, ok) }
func (m *Manager) recordIntegrityOutcome(at time.Time, ok bool) {
if err := m.settings.UpdateOffboxStatus(func(o *settings.OffboxTarget) {
o.LastIntegrityCheck = at.UTC().Format(time.RFC3339)
o.LastIntegrityOK = ok
}); err != nil {
m.logger.Printf("[ERROR] [offbox] integrity: could not persist the check outcome: %v — the check RAN and its verdict was ok=%v, but due-ness did not advance, so it will run again tomorrow", err, ok)
}
}
// CheckOffboxIntegrity runs one off-site integrity check. It NEVER writes to the repository: `check`
// is a read verb, and nothing here prunes, forgets, unlocks or backs up.
//
// Due-ness is NOT consulted here — the caller decides. The scheduled job asks IntegrityDue first; the
// operator's debug button deliberately does not, because "run it now" is the whole point of a button.
// Every other guard applies to both, including the single-writer flag.
func (m *Manager) CheckOffboxIntegrity(ctx context.Context) IntegrityResult {
res := IntegrityResult{RanAt: time.Now()}
if !m.OffboxConfigured() {
// A box with no off-site tier has nothing to check. Silent, and not an alarm.
res.Skipped, res.SkipReason = true, "no off-site target configured"
return res
}
// THE HAZARD GATE. See the file header. Skip, never wait: waiting would pin the nightly backup
// behind this check, and a skipped check costs nothing because tomorrow's run tries again.
if err := m.acquireRunning(); err != nil {
res.Skipped, res.SkipReason = true, "a backup or restore is already running"
m.logger.Printf("[INFO] [offbox] integrity: skipped — %s; due-ness is NOT advanced, so this retries on the next run", res.SkipReason)
return res
}
defer m.releaseRunning()
t := m.settings.GetOffboxTarget()
base, env := m.offboxBaseArgs(t)
// Does the repository answer at all? A failure here is UNREACHABLE, not damage — the distinction
// the whole result type exists for.
if err := m.ensureOffboxRepo(ctx, base, env); err != nil {
res.Unreachable = true
res.Output = truncate([]byte(err.Error()))
m.logger.Printf("[WARN] [offbox] integrity: the repository could not be reached, so NOTHING was checked (this is not an integrity failure — R-339 owns reachability): %v", err)
return res
}
args := []string{"check"}
if subset := m.integrityReadDataSubset(); subset != "" {
res.ReadDataSubset = subset
args = append(args, "--read-data-subset="+subset)
}
cctx, cancel := context.WithTimeout(ctx, integrityCheckTimeout)
defer cancel()
start := time.Now()
out, err := m.resticStep(cctx, env, base, "integrity-check", args...)
res.Duration = time.Since(start)
res.Output = truncate(out)
if err == nil {
res.OK = true
m.logger.Printf("[INFO] [offbox] integrity: check PASSED in %s (%s)", res.Duration.Round(time.Second), integrityDepthLabel(res.ReadDataSubset))
return res
}
// A timeout is NOT damage. The check did not finish, so it saw nothing, so it may not claim the
// store is broken — and it must not advance due-ness either.
if cctx.Err() != nil {
res.Unreachable = true
m.logger.Printf("[WARN] [offbox] integrity: the check did not finish within %s, so nothing was concluded about the store (NOT reported as damage): %v", integrityCheckTimeout, err)
return res
}
// Connect / credential / lock failures are the repository being unavailable, not unreadable.
// classifyResticProbe already names those causes; reusing it avoids a second classifier that could
// disagree with the first.
if cls := classifyResticProbe(out, err); cls == "other" || cls == "orphaned" || cls == "norepo" {
if !looksLikeRepositoryDamage(out) {
res.Unreachable = true
m.logger.Printf("[WARN] [offbox] integrity: the check could not run to a verdict (%s) — NOT reported as damage: %v: %s", cls, err, res.Output)
return res
}
}
m.logger.Printf("[ERROR] [offbox] integrity: check FAILED after %s — restic reported: %s", res.Duration.Round(time.Second), res.Output)
return res
}
// looksLikeRepositoryDamage reports whether restic's output describes a store that is READABLE and
// WRONG, as opposed to one that could not be opened.
//
// The strings are restic's own, from the 2026-08-21 damaged-pack drill and restic's check source. The
// direction is deliberate: an output this does not recognise is treated as "could not look", because
// telling a customer their backups are damaged is the more expensive mistake of the two, and R-339
// already alarms when the store cannot be reached.
func looksLikeRepositoryDamage(out []byte) bool {
s := strings.ToLower(string(out))
for _, sig := range []string{
"pack ", // "pack 1234abcd: not found in index" / "... size mismatch"
"load index", // a broken index
"blob ", // "blob not found"
"tree ", // "tree 1234: file ... blob not found"
"snapshot ", // "error for snapshot ...: ..."
"ciphertext verification failed",
"integrity error",
"repository contains errors",
} {
if strings.Contains(s, sig) {
return true
}
}
return false
}
// integrityDepthLabel names how deep the check went, for a log line.
func integrityDepthLabel(subset string) string {
if subset == "" {
return "structure and index only — no pack data was downloaded"
}
return fmt.Sprintf("structure, index, and %s of the pack data re-read", subset)
}
@@ -2,9 +2,8 @@ package backup
import (
"bytes"
"go/ast"
"go/parser"
"go/token"
"context"
"encoding/json"
"log"
"os"
"path/filepath"
@@ -186,63 +185,123 @@ func TestR358_MarkerIsNeverPlaced(t *testing.T) {
}
}
// TestR358_MarkerIsClearedBeforeResticAndWrittenAfter walks the AST of RestoreOffboxScratch.
// TestR358_MarkerOrderingIsExecuted — R-398, and the row that prompted it was WRONG in a way worth
// recording rather than quietly fixing.
//
// It exists because `resticStep` is not a seam — a test cannot run the real download without restic,
// so the ORDER of the three calls cannot be proven by execution here. Order is the entire safety
// property: a marker written before restic certifies a download that has not happened, and a clear
// that runs after it leaves a stale certificate covering a fresh part-copy. A substring search would
// not do: a commented-out call satisfies strings.Contains, which a sibling test in this project
// records paying for.
func TestR358_MarkerIsClearedBeforeResticAndWrittenAfter(t *testing.T) {
fset := token.NewFileSet()
f, err := parser.ParseFile(fset, "offbox_restore.go", nil, 0)
if err != nil {
t.Fatalf("parse offbox_restore.go: %v", err)
}
var body *ast.BlockStmt
for _, d := range f.Decls {
if fn, ok := d.(*ast.FuncDecl); ok && fn.Name.Name == "RestoreOffboxScratch" && fn.Body != nil {
body = fn.Body
}
}
if body == nil {
t.Fatal("RestoreOffboxScratch not found")
// This test used to walk the AST of RestoreOffboxScratch, on the stated ground that "`resticStep` is
// not a seam, so the ORDER cannot be proven by execution". **The first half is true and the conclusion
// was false.** `resticStep` is not overridable, but the layer it calls — `offboxRunner`, injected by
// `SetOffboxRunner` — has been a seam since the off-site tier shipped, and other tests in this package
// have been driving restic-backed paths through it all along. R-398 was filed off my own mistaken
// reading; it is corrected, not closed.
//
// Executing it is strictly stronger than the AST walk, and not only because it runs the real function:
// the runner sees EVERY argv, including the `unlock --remove-all` escalation that `resticStep` performs
// on a lock error. A `resticStepFn` seam — which R-398 proposed and this test's existence argued for —
// would have REPLACED that escalation and hidden it from exactly the assertions that need to see it.
//
// The order is the whole safety property: a marker written before restic certifies a download that has
// not happened, and a clear that runs after it leaves a stale certificate over a fresh part-copy.
func TestR358_MarkerOrderingIsExecuted(t *testing.T) {
m, scratch := newR358Manager(t)
// A stale marker from a previous run, so "cleared before" is observable as a state change rather
// than as an absence that was always absent.
if err := m.writeScratchMarker(scratch, "snap-OLD", true); err != nil {
t.Fatal(err)
}
var order []string
ast.Inspect(body, func(n ast.Node) bool {
call, ok := n.(*ast.CallExpr)
if !ok {
return true
}
if sel, ok := call.Fun.(*ast.SelectorExpr); ok {
switch sel.Sel.Name {
case "clearScratchMarker", "resticStep", "writeScratchMarker":
order = append(order, sel.Sel.Name)
var events []string
markerAtRestic := "unset"
m.SetOffboxLatestSnapshotFn(func(context.Context, string) (string, []string, error) {
return "snap-NEW", []string{"/mnt/old/backups/primary/kimai"}, nil
})
m.SetOffboxFreeFn(func(string) int64 { return 100 << 30 })
m.SetOffboxRunner(func(_ context.Context, _ []string, args ...string) ([]byte, error) {
switch {
case containsArg(args, "cat") && containsArg(args, "config"):
return []byte(`{"version":2}`), nil
case containsArg(args, "stats"):
return []byte(`{"total_size":1024}`), nil
case containsArg(args, "restore"):
events = append(events, "restic-restore")
// THE OBSERVATION: what does the marker look like at the moment restic runs? If the clear
// happened first it is gone; if the write happened first it is already there, certifying a
// download that has not finished.
if _, err := os.Stat(filepath.Join(scratch, scratchMarkerName)); err == nil {
markerAtRestic = "present"
} else {
markerAtRestic = "absent"
}
return nil, nil
case containsArg(args, "unlock"):
// TWO DIFFERENT UNLOCKS, and only one of them is dangerous. Plain `unlock` is restic's
// stale-ONLY sweep, run as pre-run hygiene by unlockStale before every restore.
// `unlock --remove-all` is resticStep's crash-lock escalation, and it is the act that must
// never happen beside a live sibling operation. Recording them separately here is what
// lets the R-359 lock-safety tests assert the second is absent without tripping over the
// first — a distinction the AST walk this replaced could not have surfaced at all.
if containsArg(args, "--remove-all") {
events = append(events, "unlock--remove-all")
} else {
events = append(events, "unlock-stale")
}
return nil, nil
}
return true
return nil, nil
})
idx := func(name string) int {
for i, n := range order {
if n == name {
return i
}
}
return -1
if err := m.RestoreOffboxScratch(context.Background(), "kimai", true); err != nil {
t.Fatalf("RestoreOffboxScratch: %v", err)
}
clear, restic, write := idx("clearScratchMarker"), idx("resticStep"), idx("writeScratchMarker")
if clear < 0 || restic < 0 || write < 0 {
t.Fatalf("RestoreOffboxScratch does not call all three (order seen: %v) — the marker is not wired", order)
if markerAtRestic != "absent" {
t.Fatalf("the completion marker was %s while restic was still downloading — a marker that "+
"predates the download certifies a copy that does not exist yet (and here it was the "+
"PREVIOUS run's marker, covering a fresh part-copy)", markerAtRestic)
}
if !(clear < restic) {
t.Errorf("the stale marker is cleared AFTER restic runs (order %v) — a previous run's "+
"certificate would cover this run's part-copy", order)
// The restore must actually have run, or the ordering assertion above proves nothing. It is NOT
// asserted to be first: `unlockStale` legitimately precedes it as pre-run hygiene, which this
// execution test surfaced on its first run and the AST walk could never have shown.
if !containsEvent(events, "restic-restore") {
t.Fatalf("restic never ran, so this test proves nothing about ordering (events=%v)", events)
}
if !(restic < write) {
t.Errorf("the marker is written BEFORE restic returns (order %v) — that certifies a download "+
"that has not happened, which is the defect with an extra step", order)
// And the DANGEROUS unlock never appeared: nothing here met a lock, so resticStep never escalated.
if containsEvent(events, "unlock--remove-all") {
t.Fatalf("`unlock --remove-all` ran during an ordinary restore (events=%v) — that escalation "+
"is only safe because the single-writer flag proves no sibling is live", events)
}
// And after a successful run the marker is there, carrying THIS run's snapshot, not the old one.
data, err := os.ReadFile(filepath.Join(scratch, scratchMarkerName))
if err != nil {
t.Fatalf("no marker after a successful restore: %v", err)
}
var mk scratchMarker
if err := json.Unmarshal(data, &mk); err != nil {
t.Fatal(err)
}
if mk.SnapshotID != "snap-NEW" || !mk.Full {
t.Fatalf("marker = %+v, want snapshot snap-NEW and full=true — the stale one survived", mk)
}
if !m.OffboxFullScratchReady("kimai") {
t.Fatal("a completed full restore did not read as ready")
}
}
func containsEvent(ev []string, want string) bool {
for _, e := range ev {
if e == want {
return true
}
}
return false
}
// containsArg is a local helper so this file does not depend on another test file's ordering.
func containsArg(args []string, want string) bool {
for _, a := range args {
if a == want {
return true
}
}
return false
}
@@ -0,0 +1,136 @@
package backup
import (
"context"
"path/filepath"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-controller/internal/settings"
)
// ── Due-ness, NOT a weekday ──────────────────────────────────────────────────────────────────────
//
// A job that fires only on Sundays silently skips a week every time the box is off on a Sunday.
// **R-341 is exactly that failure** — a dated check quietly missed, five days overdue, and nothing
// asked again. Asking "is the last successful check older than the max age?" catches up on the next
// day the box is running, whatever day that is.
func seedLastCheck(t *testing.T, m *Manager, at time.Time, ok bool) {
t.Helper()
m.RecordIntegrityOutcome(at, ok)
}
func TestR359_NotDueRunsNothing(t *testing.T) {
m, _ := newIntegrityManager(t, okRepo(nil))
seedLastCheck(t, m, time.Now().Add(-2*24*time.Hour), true)
due, last := m.IntegrityDue(time.Now())
if due {
t.Fatalf("a check 2 days old was reported due against a 7-day max age (last=%s)", last)
}
}
func TestR359_OverdueRunsOnAnyWeekday(t *testing.T) {
// 21 days: the box was off. It must run on whatever day it next comes up — the assertion is made
// for EVERY weekday so a Sunday-gated implementation cannot pass by luck.
m, _ := newIntegrityManager(t, okRepo(nil))
seedLastCheck(t, m, time.Date(2026, 8, 1, 6, 0, 0, 0, time.UTC), true)
for d := 0; d < 7; d++ {
now := time.Date(2026, 8, 22, 6, 0, 0, 0, time.UTC).AddDate(0, 0, d)
due, _ := m.IntegrityDue(now)
if !due {
t.Fatalf("a 21-day-old check was NOT due on %s — a check that waits for one weekday is a "+
"check that skips a whole period every time the box is off that day (R-341)", now.Weekday())
}
}
}
func TestR359_FailureStillAdvancesDueness(t *testing.T) {
// A broken store must not be re-checked every night: that is load with no new information, and the
// hourly operator cooldown already governs the mail.
m, _ := newIntegrityManager(t, okRepo(nil))
m.RecordIntegrityOutcome(time.Now(), false)
due, _ := m.IntegrityDue(time.Now())
if due {
t.Fatal("a FAILED check left the store due immediately — it would be re-checked every night, " +
"telling nobody anything new")
}
if tgt := m.settings.GetOffboxTarget(); tgt.LastIntegrityOK {
t.Fatal("a failed check recorded a passing verdict")
}
}
func TestR359_SkipDoesNotAdvanceDueness(t *testing.T) {
// The complement of the failure rule, and the reason the two are separate calls: a skip looked at
// NOTHING, so it may not reset the clock.
m, cap := newIntegrityManager(t, okRepo(nil))
if err := m.AcquireRunningForTest(); err != nil {
t.Fatal(err)
}
res := m.CheckOffboxIntegrity(context.Background())
m.ReleaseRunningForTest()
if !res.Skipped || len(cap.argvs) != 0 {
t.Fatalf("fixture: expected a skip with no restic, got %+v / %v", res, cap.argvs)
}
due, _ := m.IntegrityDue(time.Now())
if !due {
t.Fatal("a SKIPPED check advanced due-ness — the store would wait a full period before " +
"anything looked at it again")
}
}
func TestR359_FirstEverRunIsDue(t *testing.T) {
// An unchecked store must never read as a fresh one. This is the fail-safe direction.
m, _ := newIntegrityManager(t, okRepo(nil))
due, last := m.IntegrityDue(time.Now())
if !due {
t.Fatal("a store that has NEVER been checked was reported as not due")
}
if !last.IsZero() {
t.Errorf("a never-checked store reported a last-check time of %s", last)
}
}
func TestR359_UnparseableStampIsTreatedAsNeverChecked(t *testing.T) {
// Fail-safe again: a corrupt stamp must not certify a store as recently verified.
m, _ := newIntegrityManager(t, okRepo(nil))
if err := m.settings.UpdateOffboxStatus(func(o *settings.OffboxTarget) { o.LastIntegrityCheck = "not-a-time" }); err != nil {
t.Fatal(err)
}
due, _ := m.IntegrityDue(time.Now())
if !due {
t.Fatal("an unparseable last-check stamp was read as a recent check — a corrupt field must " +
"never buy the store a free period")
}
}
func TestR359_DuenessSurvivesRestart(t *testing.T) {
// Written and read back through the REAL store, not an in-memory field: a due-ness anchor that
// does not survive a controller restart would re-check on every boot, or never.
m, _ := newIntegrityManager(t, okRepo(nil))
at := time.Now().Add(-3 * 24 * time.Hour).UTC().Truncate(time.Second)
m.RecordIntegrityOutcome(at, true)
reloaded, err := settings.Load(filepath.Join(m.cfg.Paths.DataDir, "settings.json"), m.logger)
if err != nil {
t.Fatalf("reload: %v", err)
}
tgt := reloaded.GetOffboxTarget()
if tgt == nil || tgt.LastIntegrityCheck == "" {
t.Fatal("the last-check stamp did not survive a reload from disk")
}
got, perr := time.Parse(time.RFC3339, tgt.LastIntegrityCheck)
if perr != nil {
t.Fatalf("the persisted stamp does not parse: %v", perr)
}
if !got.Equal(at) {
t.Errorf("persisted stamp = %s, want %s", got, at)
}
if !tgt.LastIntegrityOK {
t.Error("the persisted verdict did not survive")
}
}
@@ -0,0 +1,261 @@
package backup
import (
"bytes"
"context"
"errors"
"log"
"strings"
"testing"
"gitea.dooplex.hu/admin/felhom-controller/internal/settings"
)
// ── R-359 — the off-site store was never checked ─────────────────────────────────────────────────
//
// The whole-guest tier has verify jobs; the tier holding the customer's documents and photos had none.
// The complete set of restic verbs this controller used contained no `check` — verified 2026-08-30.
//
// These drive the REAL CheckOffboxIntegrity through the EXISTING `offboxRunner` seam, which has been
// injectable since the off-site tier shipped. (R-398 claimed otherwise and was my own mistake; the
// seam sees every argv, including the `unlock --remove-all` escalation a `resticStepFn` would have
// hidden — which is exactly what the lock-safety tests must observe.)
// errFake is a plain non-nil error for seam replies; the classifier reads the OUTPUT, not the error
// type, so a synthetic error is faithful here.
var errFake = errors.New("restic exited non-zero")
// integrityCapture records every restic invocation so both the effects and the NON-effects are
// assertable. `argvs` is the whole point: a test that only checks the verdict cannot tell a check that
// ran from one that did not.
type integrityCapture struct {
argvs [][]string
reply func(args []string) ([]byte, error)
logBuf *bytes.Buffer
}
func (c *integrityCapture) runner() offboxRunner {
return func(_ context.Context, _ []string, args ...string) ([]byte, error) {
c.argvs = append(c.argvs, append([]string{}, args...))
if c.reply != nil {
return c.reply(args)
}
return nil, nil
}
}
func (c *integrityCapture) sawVerb(verb string) bool {
for _, a := range c.argvs {
for _, x := range a {
if x == verb {
return true
}
}
}
return false
}
func (c *integrityCapture) checkArgv() []string {
for _, a := range c.argvs {
for _, x := range a {
if x == "check" {
return a
}
}
}
return nil
}
// newIntegrityManager builds a manager with a configured off-site target and a captured runner.
func newIntegrityManager(t *testing.T, reply func(args []string) ([]byte, error)) (*Manager, *integrityCapture) {
t.Helper()
m, _ := newOffboxManager(t)
cap := &integrityCapture{reply: reply, logBuf: &bytes.Buffer{}}
m.logger = log.New(cap.logBuf, "", 0)
m.SetOffboxRunner(cap.runner())
return m, cap
}
// okRepo answers `cat config` so ensureOffboxRepo passes, then defers to `then` for everything else.
func okRepo(then func(args []string) ([]byte, error)) func(args []string) ([]byte, error) {
return func(args []string) ([]byte, error) {
for _, a := range args {
if a == "config" {
return []byte(`{"version":2}`), nil
}
}
if then != nil {
return then(args)
}
return nil, nil
}
}
func TestR359_HealthyRepoReportsOK(t *testing.T) {
m, cap := newIntegrityManager(t, okRepo(nil))
res := m.CheckOffboxIntegrity(context.Background())
if !res.OK || res.Skipped || res.Unreachable {
t.Fatalf("a healthy repo did not report OK: %+v", res)
}
if cap.checkArgv() == nil {
t.Fatal("`restic check` was never invoked — the check did not check anything")
}
}
func TestR359_RepositoryErrorReportsFailure(t *testing.T) {
// restic's own words from the 2026-08-21 damaged-pack drill.
const damaged = "pack 5b1f2c3d: not found in index\nrepository contains errors"
m, _ := newIntegrityManager(t, okRepo(func(args []string) ([]byte, error) {
return []byte(damaged), errFake
}))
res := m.CheckOffboxIntegrity(context.Background())
if res.OK {
t.Fatal("a repository restic said contains errors was reported as OK — this is the defect the " +
"whole feature exists to prevent")
}
if res.Unreachable {
t.Fatal("readable-and-damaged was misclassified as unreachable — those are different facts, " +
"and only one of them means the customer's backups are broken")
}
if !strings.Contains(res.Output, "not found in index") {
t.Errorf("restic's own words must reach the LOG so the operator can diagnose; got %q", res.Output)
}
}
func TestR359_UnreachableIsNotAnIntegrityFailure(t *testing.T) {
// The repo cannot even be opened. "I could not look" is not "I looked and it is broken".
m, _ := newIntegrityManager(t, func(args []string) ([]byte, error) {
return []byte("ssh: connect to host nas.local port 22: Connection refused"), errors.New("exit 1")
})
res := m.CheckOffboxIntegrity(context.Background())
if res.OK {
t.Fatal("an unreachable repository was reported as a passing check")
}
if !res.Unreachable {
t.Fatal("an unreachable repository was reported as DAMAGE — that would alarm the customer that " +
"their backups are corrupt when nothing was ever looked at, and R-339 already owns reachability")
}
}
func TestR359_TimeoutIsNotDamage(t *testing.T) {
m, _ := newIntegrityManager(t, okRepo(func(args []string) ([]byte, error) {
return nil, context.DeadlineExceeded
}))
ctx, cancel := context.WithCancel(context.Background())
cancel() // an already-dead context: the check cannot finish
res := m.CheckOffboxIntegrity(ctx)
if res.OK {
t.Fatal("a check that never finished reported OK")
}
if !res.Unreachable {
t.Fatalf("a check that timed out was reported as damage: %+v — it saw nothing, so it may not "+
"claim the store is broken", res)
}
}
func TestR359_StructureCheckPassesNoReadDataFlag(t *testing.T) {
m, cap := newIntegrityManager(t, okRepo(nil))
m.CheckOffboxIntegrity(context.Background())
argv := cap.checkArgv()
if argv == nil {
t.Fatal("no check ran")
}
for _, a := range argv {
if strings.HasPrefix(a, "--read-data") {
t.Fatalf("the DEFAULT check downloaded pack data (%q) — that is a bandwidth cost nobody "+
"chose, and R-399 exists precisely so it is not chosen here", a)
}
}
}
func TestR359_ReadDataSubsetIsPassedWhenConfigured(t *testing.T) {
m, cap := newIntegrityManager(t, okRepo(nil))
m.cfg.Monitoring.Integrity.ReadDataSubset = "5%"
res := m.CheckOffboxIntegrity(context.Background())
if res.ReadDataSubset != "5%" {
t.Errorf("result did not record the depth it ran at: %+v", res)
}
var found bool
for _, a := range cap.checkArgv() {
if a == "--read-data-subset=5%" {
found = true
}
}
if !found {
t.Fatalf("the configured subset did not reach restic; argv=%v", cap.checkArgv())
}
}
func TestR359_MalformedReadDataSubsetIsTreatedAsOff(t *testing.T) {
m, cap := newIntegrityManager(t, okRepo(nil))
m.cfg.Monitoring.Integrity.ReadDataSubset = "banana"
res := m.CheckOffboxIntegrity(context.Background())
for _, a := range cap.checkArgv() {
if strings.HasPrefix(a, "--read-data") {
t.Fatalf("a malformed value was handed to restic (%q) — restic rejects it and the WHOLE "+
"check fails, so one typo silently stops the store being verified at all", a)
}
}
if res.ReadDataSubset != "" {
t.Errorf("a refused value was still recorded as the depth: %q", res.ReadDataSubset)
}
if !strings.Contains(cap.logBuf.String(), "WARN") {
t.Error("a refused config value must say so — silence makes a typo indistinguishable from a " +
"deliberate structure-only setting")
}
}
func TestR359_MessageNeverCarriesResticOutputOrCredentials(t *testing.T) {
// R-379: 615 bytes of raw database text reached a customer once. And offboxBaseArgs builds the repo
// as `sftp:<user>@<host>:<path>`, so a raw passthrough leaks the credential shape too.
const secretish = "sftp:felhom@nas.local:/srv/repo pack 5b1f2c3d corrupt"
m, _ := newIntegrityManager(t, okRepo(func(args []string) ([]byte, error) {
return []byte(secretish), errFake
}))
res := m.CheckOffboxIntegrity(context.Background())
if res.OK {
t.Fatal("fixture wrong: this should be a failure")
}
// The customer sentence is a CONSTANT and contains none of it. Asserted here rather than only in
// the web package because this is where the output is captured.
for _, bad := range []string{"sftp:", "nas.local", "5b1f2c3d", "felhom@"} {
if strings.Contains(integrityFailedCustomerSentence, bad) {
t.Fatalf("the customer-facing failure sentence carries %q", bad)
}
}
// ...while the operator's log DOES get it, or the fault cannot be diagnosed without a rebuild.
if !strings.Contains(res.Output, "5b1f2c3d") {
t.Error("restic's output did not reach the result for the log")
}
}
// integrityFailedCustomerSentence mirrors the constant in cmd/controller. Duplicated deliberately and
// narrowly: this package cannot import main, and the property under test is that the SENTENCE carries
// no machine detail — a property of the words themselves.
const integrityFailedCustomerSentence = "A távoli mentés ellenőrzése hibát talált a tárolóban. A mentések egy része sérült lehet. Ne törölj semmit, és vedd fel velünk a kapcsolatot."
func TestR359_NoTargetConfiguredIsASilentSkip(t *testing.T) {
m, cap := newIntegrityManager(t, okRepo(nil))
if err := m.settings.SetOffboxTarget(&settings.OffboxTarget{Enabled: false}); err != nil {
t.Fatal(err)
}
res := m.CheckOffboxIntegrity(context.Background())
if !res.Skipped {
t.Fatalf("a box with no off-site tier did not skip: %+v", res)
}
if len(cap.argvs) != 0 {
t.Fatalf("restic ran on a box with no off-site target: %v", cap.argvs)
}
if res.OK {
t.Fatal("a skip was reported as a passing check — nothing was checked")
}
}
@@ -0,0 +1,114 @@
package backup
import (
"context"
"strings"
"testing"
)
// ── THE CONTROL — R-359's single most important property ─────────────────────────────────────────
//
// `resticStep` self-heals a crash lock: on "repository is already locked" it runs
// **`unlock --remove-all`** and retries. Its own doc comment states why that is safe — *"the in-process
// single-flight mutex (held by every caller of this method) proves no sibling operation is live"*.
//
// **An integrity check that did not take that flag would break the invariant.** It could meet the lock
// of a `forget --prune` running from this same box, remove it, and retry over the top of a live prune —
// on the tier holding the customer's documents and photos.
//
// So these assert NON-EFFECTS, which is the only way to test "it did not do the dangerous thing":
// restic was never invoked at all, and `unlock` never appeared in any argument list. A test that only
// checked `Skipped == true` would pass against an implementation that skipped AFTER running the check.
func TestR359_SkipsWhenRunningFlagHeld(t *testing.T) {
m, cap := newIntegrityManager(t, okRepo(nil))
// A sibling operation is live — exactly the state a nightly off-site backup produces.
if err := m.AcquireRunningForTest(); err != nil {
t.Fatalf("fixture: %v", err)
}
defer m.ReleaseRunningForTest()
res := m.CheckOffboxIntegrity(context.Background())
// THE TWO NON-EFFECTS. These are the assertions that fail against a missing guard; the verdict
// fields below would not.
if len(cap.argvs) != 0 {
t.Fatalf("RESTIC WAS INVOKED while another operation held the single-writer flag: %v — this is "+
"the hazard: the check can meet a live prune's lock, and resticStep escalates to "+
"`unlock --remove-all` and retries over the top of it", cap.argvs)
}
if cap.sawVerb("unlock") {
t.Fatal("`unlock` was invoked beside a live sibling operation — the exact act that must never happen")
}
if !res.Skipped {
t.Fatalf("the check did not report a skip: %+v", res)
}
if res.OK {
t.Fatal("a skip was reported as a passing check — nothing was checked, and recording it as a " +
"pass would let a store go unverified while the record says otherwise")
}
// Due-ness must NOT advance: tomorrow has to try again.
if tgt := m.settings.GetOffboxTarget(); tgt != nil && tgt.LastIntegrityCheck != "" {
t.Fatalf("a SKIPPED check advanced due-ness (%q) — the store would then wait a full period "+
"before anything looked at it again, which is R-341's failure with extra steps", tgt.LastIntegrityCheck)
}
}
func TestR359_ReleasesTheFlagOnEveryPath(t *testing.T) {
// A check that leaks the flag stops every backup on the box until restart. Each outcome is walked.
for _, tc := range []struct {
name string
reply func(args []string) ([]byte, error)
}{
{"ok", okRepo(nil)},
{"damaged", okRepo(func([]string) ([]byte, error) { return []byte("repository contains errors"), errFake })},
{"unreachable", func([]string) ([]byte, error) { return []byte("connection refused"), errFake }},
} {
t.Run(tc.name, func(t *testing.T) {
m, _ := newIntegrityManager(t, tc.reply)
m.CheckOffboxIntegrity(context.Background())
// If the flag leaked, this acquire fails.
if err := m.AcquireRunningForTest(); err != nil {
t.Fatalf("the single-writer flag was NOT released after a %s outcome (%v) — every "+
"backup and restore on this box would refuse until it restarts", tc.name, err)
}
m.ReleaseRunningForTest()
})
}
}
func TestR359_DoesNotBlockAConcurrentBackup(t *testing.T) {
// The complement: a check that skipped must leave the flag free for the operation it yielded to.
m, _ := newIntegrityManager(t, okRepo(nil))
if err := m.AcquireRunningForTest(); err != nil {
t.Fatal(err)
}
m.CheckOffboxIntegrity(context.Background()) // skips
m.ReleaseRunningForTest() // the sibling finishes
if err := m.AcquireRunningForTest(); err != nil {
t.Fatalf("after a skip the flag could not be acquired: %v — the check held something it "+
"never took", err)
}
m.ReleaseRunningForTest()
}
func TestR359_NeverWritesToTheRepository(t *testing.T) {
// `check` is a read verb. This pins that the check's argv carries no verb that could modify the
// store — the tier holds real customer data and the box's own credential can already delete from
// it (R-95, open and ranked).
m, cap := newIntegrityManager(t, okRepo(nil))
m.CheckOffboxIntegrity(context.Background())
for _, verb := range []string{"forget", "prune", "backup", "init", "unlock", "restore"} {
if cap.sawVerb(verb) {
t.Fatalf("the integrity check invoked `%s` — it must only ever READ (argv=%v)", verb, cap.argvs)
}
}
argv := strings.Join(cap.checkArgv(), " ")
if !strings.Contains(argv, "check") {
t.Fatalf("no check verb in %q", argv)
}
}