hub v0.111.0: notice a deletion within a day (R-431); correct R-429; re-scope R-95
gates / gates (push) Successful in 17s

THE RECORD WAS TELLING A WORSE STORY THAN THE TRUTH FOR TWO MONTHS, and my own probe is why.

R-429 CORRECTED. Yesterday's spike searched for a directory called `.snapshots`. The vendor documents
the path as /.zfs/snapshot. The probe's CONTROLS were sound and its SUBJECT was wrong, so "not found"
was true and meant nothing. Re-probed at the documented path on both boxes, with controls:

  - /.zfs lists (shares, snapshot) from inside the jail;
  - a write into /.zfs/snapshot is REFUSED - dest open ...: Failure - while the identical write to
    the account home SUCCEEDS and was cleaned up.

That is the append-only property PROVEN rather than cited, and it is the sentence the whole re-scope
rests on. Seven daily snapshots are confirmed in the panel. The mitigation works. What remains true,
and was always the actual finding: the row claiming it had no R-number, its "confirm tomorrow" went
36 days unanswered, and the DUE-CHECKS block built for that class was empty. THE FINDING WAS NEVER THE
SNAPSHOTS - IT WAS THAT NOBODY COULD TELL.

R-95 RE-SCOPED: the box can delete its LIVE repository but cannot write to the daily snapshots of it,
so a deletion costs at most one day plus a per-file recovery - not open-ended loss. The ranking is
Viktor's; it has been #1 since July on the old story.

R-432 FILED: a sub-account sees /.zfs/snapshot EMPTY while the same box holds seven snapshots, so
per-file recovery is operator-only today. One panel read settles whether a NAMED snapshot can still be
entered, which would make it product-reachable.

R-431 SHIPPED. Third signal in OffsiteChecker. On the hub deliberately: a detector on the box is one
the deletion can silence. Threshold REASONED, not invented - over 12 898 reports every decrease lands
on ZERO and predates stats_known, and in the stats_known window there are none, so observed churn gave
nothing to calibrate against. Retention cannot halve a total; a mass deletion goes to ~0. Hence: more
than half, and at least 5. Guarded by StatsKnown (R-331), the declared State (R-204) and run success
(R-100 - whose lesson lives in this very file).

ACCEPTANCE: 9 009 real report points replayed through the detector produced ZERO alarms.

Three red-proofs run. The escalation one only became real after the first version was found HOLLOW -
it re-swept the same report, so the baseline had already moved and the latch was never consulted.

07 row 10's status is NOT moved: the write-refusal is measured, but the recovery ROUTE has never been
walked, which is what PARTIAL means.
This commit is contained in:
2026-09-01 14:25:24 +02:00
parent 0476a8d8e6
commit 30681764cb
10 changed files with 511 additions and 5 deletions
+170
View File
@@ -17,6 +17,9 @@ import (
//
// - FILL: repo_size_bytes vs the shared-model soft quota (quota_gb>0) at warn 90% / crit 95% — the
// operator's early warning before the controller's own 100% run-refusal bites the customer.
// - SNAPSHOT-DROP (R-431): the reported snapshot count falls by more than retention can explain —
// the "something deleted this customer's off-site history" detector. See snapshotDrop below for
// why it lives HERE and not on the box, and for the measurement behind its threshold.
// - STALENESS: enabled + escrowed but no run in >48h (or never) — the silently-STUCK detector. A
// RECENTLY-failing offsite is NOT stale (backup_failed already alerts it); staleness is the
// complement: nothing is even trying. Pending/disabled targets are normal onboarding, never stale.
@@ -36,6 +39,10 @@ type OffsiteChecker struct {
mu sync.Mutex
fillStates map[string]string // customerID → fill band
staleStates map[string]string // customerID → "ok" | "stale"
// R-431. dropStates is the escalation latch ("ok" | "dropped"); lastCounts is the baseline the
// next sweep compares against. Only a TRUSTWORTHY report updates lastCounts — see snapshotDrop.
dropStates map[string]string
lastCounts map[string]int
}
const defaultOffsiteStaleAfter = 48 * time.Hour
@@ -52,6 +59,15 @@ type offsiteReport struct {
SnapshotCount int `json:"snapshot_count"`
RepoSizeBytes int64 `json:"repo_size_bytes"`
QuotaGB int `json:"quota_gb"`
// StatsKnown (R-331/R-225) — the ONLY thing separating "this repository holds nothing" from
// "nobody has ever measured this repository". Both are `snapshot_count: 0` on the wire and they
// are opposite news. ABSENT on a pre-v0.225.0 controller, and absence means the box CANNOT
// ANSWER, never that the answer is no. snapshotDrop refuses to judge without it.
StatsKnown bool `json:"stats_known,omitempty"`
// State (R-204/R-193) — a DECLARED condition, empty on every healthy box. A box that has
// re-initialised, lost its credential or been abandoned reports a real, correct, large drop.
// The box knows its own situation; believe it rather than alarming on it.
State string `json:"state,omitempty"`
}
// NewOffsiteChecker builds the checker. Same seeding philosophy as StorageFillChecker: already-breached
@@ -64,6 +80,7 @@ func NewOffsiteChecker(s *store.Store, staleAfter time.Duration, onEvent EventNo
oc := &OffsiteChecker{
store: s, logger: logger, onEvent: onEvent, staleAfter: staleAfter, now: time.Now,
fillStates: make(map[string]string), staleStates: make(map[string]string),
dropStates: make(map[string]string), lastCounts: make(map[string]int),
}
customers, err := s.GetCustomers()
if err != nil {
@@ -83,6 +100,13 @@ func NewOffsiteChecker(s *store.Store, staleAfter time.Duration, onEvent EventNo
if !oc.isStale(c.CustomerID, off) {
oc.staleStates[c.CustomerID] = "ok"
}
// R-431: seed the baseline from the current report so a hub restart does not read the first
// sweep as a drop from nothing. Only a trustworthy report seeds — an untrustworthy one leaves
// no baseline, and no baseline means no verdict.
if oc.countIsTrustworthy(off) {
oc.lastCounts[c.CustomerID] = off.SnapshotCount
oc.dropStates[c.CustomerID] = "ok"
}
}
logger.Printf("[INFO] Offsite checker initialized: fill warn=90%% crit=95%%, stale after %s, %d ok-seeded", staleAfter, seeded)
return oc
@@ -175,6 +199,124 @@ func (oc *OffsiteChecker) isStale(customerID string, off *offsiteReport) bool {
return oc.now().Sub(t) > oc.staleAfter
}
// ── SNAPSHOT-DROP (R-431) ───────────────────────────────────────────────────────────────────────
//
// WHY THIS LIVES ON THE HUB AND NOT ON THE BOX. The thing being detected is a box deleting its own
// off-site backups — so a detector living on that box is a detector the same event can silence. The
// hub already receives the count in every report and already keeps the history, so it can notice
// without trusting the box's judgement. The box still REPORTS the number, and a sophisticated
// attacker could lie about it; that is a far higher bar than deleting files, and the weekly integrity
// check (R-359) would then disagree with the lie.
//
// WHAT IT IS FOR. R-95: the restic credential can delete its own repository, and the sub-account API
// has no append-only axis, so prevention needs a transport change. The Storage Box's daily ZFS
// snapshots bound the loss — MEASURED 2026-09-01: a write into `/.zfs/snapshot` is REFUSED
// (`dest open …: Failure`) while the same write to the account home succeeds. So the remedy exists;
// what was missing was NOTICING, and this is that.
//
// THE THRESHOLD, AND THE MEASUREMENT IT CAME FROM. Measured over the hub's own `reports` table on
// 2026-09-01: 12 898 reports, 4 customers, 2026-06-05 → 2026-09-01.
//
// * In the whole history there are NINE decreases, and EVERY ONE of them lands exactly on ZERO
// (36→0, 18→0 ×2, 15→0, 12→0, 8→0, 3→0). There is not one gradual retention decrease anywhere.
// * Every one of those nine predates `stats_known`, i.e. they are the R-331 shape — a zero that
// means "could not measure", not "nothing is there". Several carry a declared State
// (`needs_credential`, `awaiting_recovery_key`), which says so outright.
// * In the window where `stats_known` is TRUE (380 reports across both live boxes) there are ZERO
// decreases: demo-felhom sat flat at 10, demo-hp moved 67→68→69. Only rises.
//
// So observed retention churn gives NOTHING to calibrate against, and saying so is the honest answer
// rather than inventing a number (R-401's lesson). The threshold is therefore reasoned from what
// retention CAN do, not from what it was seen to do:
//
// the box runs `forget --keep-daily 7 --keep-weekly 4 --keep-monthly 6 --group-by host,tags`
// over ~8 apps. On a boundary day several groups can expire at once, so a legitimate pass can
// plausibly remove low double digits. **What it can NEVER do is halve the total**: keeping 7 daily
// + 4 weekly + 6 monthly per group is a floor, and a mass deletion goes to ~0.
//
// Hence: **a fall of MORE THAN HALF the previous count, and at least 5 snapshots.** The 50% cannot be
// reached by retention; the floor of 5 stops a tiny-count box alarming on ordinary ageing. It is
// deliberately NOT sensitive — a detector that cries wolf is switched off within a fortnight, and
// this project has proved that twice in a week.
const (
snapshotDropFraction = 0.5 // more than half the history gone in one step
snapshotDropFloor = 5 // and at least this many, so small counts do not twitch
)
// countIsTrustworthy — the three pre-conditions, each with a scar behind it. A report failing ANY of
// them is not evidence of anything: it neither alarms NOR updates the baseline, because comparing
// against a number nobody could measure is how a detector invents its own findings.
func (oc *OffsiteChecker) countIsTrustworthy(off *offsiteReport) bool {
// 1. R-331 — a zero is not a zero when stats are unknown. Absent `stats_known` means the box
// cannot answer. Every historical decrease in this hub's data is this shape.
if !off.StatsKnown {
return false
}
// 2. R-204/R-193 — the box DECLARES its own situation. A re-initialised, credential-less or
// abandoned box reports a real, correct drop; alarming on it would be blaming the box for
// telling the truth.
if off.State != "" {
return false
}
// 3. R-100 — PRESENCE IS NOT SUCCESS, and this file is where that lesson was learned. A count
// from a failed or half-finished run is not a measurement of the store. "incomplete" (R-203)
// is deliberately excluded too: a partial run legitimately counts less.
if off.LastStatus != "ok" || off.LastSuccess == "" {
return false
}
return true
}
// snapshotDropped reports whether `off` is a drop worth alarming about, against the remembered
// baseline. Caller holds oc.mu.
func (oc *OffsiteChecker) snapshotDropped(customerID string, off *offsiteReport) (bool, int, int) {
prev, had := oc.lastCounts[customerID]
if !had || prev <= 0 {
return false, prev, off.SnapshotCount
}
drop := prev - off.SnapshotCount
if drop < snapshotDropFloor {
return false, prev, off.SnapshotCount
}
return float64(drop) > float64(prev)*snapshotDropFraction, prev, off.SnapshotCount
}
// emitSnapshotDrop — one event, at a severity inside the hub's exact vocabulary (08 §6.1).
//
// THE MESSAGE MUST NOT SAY THE DATA IS LOST, because after the 2026-09-01 measurement that is usually
// false: the daily Storage Box snapshots are read-only to every account (proven, not cited) and hold
// the older copy. It says what happened, what it means, and where the data still is.
func (oc *OffsiteChecker) emitSnapshotDrop(customerID string, off *offsiteReport, prev, cur int) {
message := fmt.Sprintf(
"Customer %s: off-site backup count fell from %d to %d snapshot(s) in one report — more than "+
"retention can explain. The daily Storage Box snapshots are read-only and still hold the "+
"older copy, so this is recoverable file-by-file; it is NOT confirmed data loss. Check "+
"whether a deletion ran on the box before restoring anything.",
customerID, prev, cur)
details, _ := json.Marshal(map[string]any{
"customer_id": customerID, "previous_count": prev, "current_count": cur,
"drop": prev - cur, "last_success": off.LastSuccess, "last_status": off.LastStatus,
})
oc.logger.Printf("[INFO] Offsite snapshot drop: %s %d -> %d (offsite_snapshots_dropped)", customerID, prev, cur)
if _, err := oc.store.SaveEvent(customerID, "offsite_snapshots_dropped", "error", message, string(details), "hub"); err != nil {
oc.logger.Printf("[WARN] Failed to save offsite snapshot-drop event for %s: %v", customerID, err)
return
}
if oc.onEvent != nil {
oc.onEvent(customerID, "offsite_snapshots_dropped", "error", message, string(details), "hub")
}
}
// GetDropState exposes the latch for tests.
func (oc *OffsiteChecker) GetDropState(customerID string) string {
oc.mu.Lock()
defer oc.mu.Unlock()
if s := oc.dropStates[customerID]; s != "" {
return s
}
return "unknown"
}
// warnLegacyOnce logs the legacy degrade a single time per customer. Once, because this is a
// steady-state condition until the box upgrades — a per-cycle line would be pure noise — but it must be
// logged at all, so a fleet silently running on the old anchor is visible rather than assumed.
@@ -228,11 +370,15 @@ func (oc *OffsiteChecker) Check() {
if off == nil {
delete(oc.fillStates, c.CustomerID) // vanished object (disabled / downgraded) → re-arm
delete(oc.staleStates, c.CustomerID)
delete(oc.dropStates, c.CustomerID)
delete(oc.lastCounts, c.CustomerID) // no baseline survives a vanished object
continue
}
if oc.store.IsCustomerBlocked(c.CustomerID) {
delete(oc.fillStates, c.CustomerID)
delete(oc.staleStates, c.CustomerID)
delete(oc.dropStates, c.CustomerID)
delete(oc.lastCounts, c.CustomerID)
continue
}
@@ -261,6 +407,24 @@ func (oc *OffsiteChecker) Check() {
}
}
oc.staleStates[c.CustomerID] = newStale
// SNAPSHOT-DROP (R-431). Escalation-only, exactly like the two signals above: one deletion
// produces ONE alarm, not one per report cycle. The baseline then moves to the new value, so a
// SECOND deletion later is still caught — from the new floor.
//
// An UNTRUSTWORTHY report is a no-op in both directions: no alarm, and the baseline is left
// alone rather than being overwritten with a number nobody could measure. That is what stops
// a `needs_credential` zero from becoming the baseline and making the RECOVERY look like a rise.
if oc.countIsTrustworthy(off) {
dropped, prev, cur := oc.snapshotDropped(c.CustomerID, off)
if dropped && oc.dropStates[c.CustomerID] != "dropped" {
oc.emitSnapshotDrop(c.CustomerID, off, prev, cur)
oc.dropStates[c.CustomerID] = "dropped"
} else if !dropped {
oc.dropStates[c.CustomerID] = "ok"
}
oc.lastCounts[c.CustomerID] = cur
}
}
for k := range oc.fillStates {
if !seen[k] {
@@ -272,6 +436,12 @@ func (oc *OffsiteChecker) Check() {
delete(oc.staleStates, k)
}
}
for k := range oc.dropStates {
if !seen[k] {
delete(oc.dropStates, k)
delete(oc.lastCounts, k)
}
}
}
// GetFillState / GetStaleState expose current states for tests.
+278
View File
@@ -0,0 +1,278 @@
package monitor
import (
"encoding/json"
"fmt"
"os"
"path/filepath"
"strings"
"testing"
"time"
)
// R-431 — the snapshot-drop signal, proven in BOTH directions.
//
// The acceptance test is TestR431_RealHistoryProducesZeroAlarms: a detector that fires on healthy
// boxes is the mistake this project caught twice in one week, and it is the one that would get this
// signal switched off within a fortnight.
// dropJSON builds a trustworthy offsite object (stats_known, no declared state, last run ok).
func dropJSON(count int, statsKnown bool, state, lastStatus string) string {
ts := time.Now().UTC().Add(-1 * time.Hour).Format(time.RFC3339)
success := ts
if lastStatus != "ok" {
success = ""
}
return fmt.Sprintf(
`{"enabled":true,"escrow_state":"escrowed","last_run":%q,"last_status":%q,"last_success":%q,`+
`"snapshot_count":%d,"repo_size_bytes":1073741824,"quota_gb":0,"stats_known":%v,"state":%q}`,
ts, lastStatus, success, count, statsKnown, state)
}
// TestR431_FiresOnAMassDeletion — direction 1. A drop past the threshold alarms EXACTLY once.
//
// RED-PROOF (run 2026-09-01, recorded in REPORT.md): setting snapshotDropFraction to 0.99 makes this
// fail — 69 → 4 is a 94% fall and would no longer qualify, which is what an over-loose threshold
// looks like in production.
func TestR431_FiresOnAMassDeletion(t *testing.T) {
st := newDiskStore(t)
var got []struct{ et, sev, msg string }
saveOffsiteReport(t, st, "victim", dropJSON(69, true, "", "ok"))
oc := NewOffsiteChecker(st, 48*time.Hour, func(_, et, sev, msg, _, _ string) {
got = append(got, struct{ et, sev, msg string }{et, sev, msg})
}, quietLog())
// the constructor seeded the baseline at 69; now the store is emptied
saveOffsiteReport(t, st, "victim", dropJSON(4, true, "", "ok"))
oc.Check()
var drops []struct{ et, sev, msg string }
for _, g := range got {
if g.et == "offsite_snapshots_dropped" {
drops = append(drops, g)
}
}
if len(drops) != 1 {
t.Fatalf("want exactly 1 offsite_snapshots_dropped, got %d (%v)", len(drops), got)
}
// The severity MUST be in the hub's exact vocabulary — anything else is coerced to info and
// mailed to nobody (08 §6.1, shipped twice).
switch drops[0].sev {
case "info", "warning", "error", "critical":
default:
t.Fatalf("severity %q is outside the hub vocabulary — it would be coerced to info and reach nobody", drops[0].sev)
}
for _, frag := range []string{"69", "4", "read-only", "NOT confirmed data loss"} {
if !strings.Contains(drops[0].msg, frag) {
t.Fatalf("message must contain %q; got: %s", frag, drops[0].msg)
}
}
if strings.Contains(drops[0].msg, "data is lost") || strings.Contains(drops[0].msg, "data lost") {
t.Fatalf("the message must NOT claim data loss — the snapshots usually still hold it: %s", drops[0].msg)
}
}
// TestR431_EscalationOnlyLatch — a CONTINUING deletion must not page on every report cycle.
//
// WRITTEN AFTER A HOLLOW FIRST ATTEMPT, and the failure is recorded because it is instructive: the
// original assertion re-swept the SAME report and called that "escalation-only". It proved nothing —
// the baseline had already moved to the new count, so the second sweep saw a drop of zero and the
// latch was never consulted. Its red-proof (removing the latch) PASSED, which is how it was caught.
//
// This drives a count that keeps FALLING, which is the only shape where the latch is load-bearing.
//
// RED-PROOF (run 2026-09-01): replacing `dropped && oc.dropStates[...] != "dropped"` with `dropped`
// makes this fail with 2 alarms — one per report cycle, which is the noise that trains an operator
// to ignore the alarm.
func TestR431_EscalationOnlyLatch(t *testing.T) {
st := newDiskStore(t)
var n int
saveOffsiteReport(t, st, "sliding", dropJSON(69, true, "", "ok"))
oc := NewOffsiteChecker(st, 48*time.Hour, func(_, et, _, _, _, _ string) {
if et == "offsite_snapshots_dropped" {
n++
}
}, quietLog())
saveOffsiteReport(t, st, "sliding", dropJSON(30, true, "", "ok")) // 69 -> 30: alarm
oc.Check()
if n != 1 {
t.Fatalf("the first large drop must alarm exactly once; got %d", n)
}
saveOffsiteReport(t, st, "sliding", dropJSON(2, true, "", "ok")) // 30 -> 2: still falling
oc.Check()
if n != 1 {
t.Fatalf("a CONTINUING deletion must not re-page while the latch is set; got %d alarms", n)
}
if oc.GetDropState("sliding") != "dropped" {
t.Fatalf("the latch must be held, got %q", oc.GetDropState("sliding"))
}
// RECOVERY RE-ARMS: a clean sweep clears the latch, so a LATER deletion is caught again.
saveOffsiteReport(t, st, "sliding", dropJSON(40, true, "", "ok")) // rebuilt, no drop
oc.Check()
if oc.GetDropState("sliding") != "ok" {
t.Fatalf("a clean sweep must re-arm the latch, got %q", oc.GetDropState("sliding"))
}
saveOffsiteReport(t, st, "sliding", dropJSON(1, true, "", "ok")) // deleted again
oc.Check()
if n != 2 {
t.Fatalf("after re-arming, a NEW deletion must alarm again; got %d", n)
}
}
// TestR431_SilentWhenNotTrustworthy — direction 2, the three pre-conditions, each with its scar.
func TestR431_SilentWhenNotTrustworthy(t *testing.T) {
cases := []struct {
name, first, second string
}{
{"stats_known absent (R-331: a zero that means UNMEASURED)",
dropJSON(69, true, "", "ok"), dropJSON(0, false, "", "ok")},
{"a DECLARED state (R-204: the box says what happened)",
dropJSON(69, true, "", "ok"), dropJSON(0, true, "needs_credential", "ok")},
{"the run FAILED (R-100: presence is not success)",
dropJSON(69, true, "", "ok"), dropJSON(0, true, "", "error")},
{"the run was INCOMPLETE (R-203: a partial run counts less)",
dropJSON(69, true, "", "ok"), dropJSON(0, true, "", "incomplete")},
}
for i, c := range cases {
t.Run(c.name, func(t *testing.T) {
st := newDiskStore(t)
cid := fmt.Sprintf("c%d", i)
var n int
saveOffsiteReport(t, st, cid, c.first)
oc := NewOffsiteChecker(st, 48*time.Hour, func(_, et, _, _, _, _ string) {
if et == "offsite_snapshots_dropped" {
n++
}
}, quietLog())
saveOffsiteReport(t, st, cid, c.second)
oc.Check()
if n != 0 {
t.Fatalf("%s: must NOT alarm; got %d", c.name, n)
}
// AND the baseline must be untouched, so the RECOVERY does not read as a rise-then-drop.
oc.mu.Lock()
base := oc.lastCounts[cid]
oc.mu.Unlock()
if base != 69 {
t.Fatalf("%s: an untrustworthy report must not overwrite the baseline; got %d", c.name, base)
}
})
}
}
// TestR431_OrdinaryRetentionIsSilent — a fall that retention CAN explain must not alarm.
func TestR431_OrdinaryRetentionIsSilent(t *testing.T) {
for _, c := range []struct {
name string
from, to int
wantAlarms int
}{
{"69 -> 60 (9 gone, under half)", 69, 60, 0},
{"10 -> 7 (3 gone, under the floor of 5)", 10, 7, 0},
{"10 -> 5 (5 gone, exactly half — NOT more than half)", 10, 5, 0},
{"69 -> 34 (35 gone, more than half)", 69, 34, 1},
} {
t.Run(c.name, func(t *testing.T) {
st := newDiskStore(t)
cid := fmt.Sprintf("r%d%d", c.from, c.to)
var n int
saveOffsiteReport(t, st, cid, dropJSON(c.from, true, "", "ok"))
oc := NewOffsiteChecker(st, 48*time.Hour, func(_, et, _, _, _, _ string) {
if et == "offsite_snapshots_dropped" {
n++
}
}, quietLog())
saveOffsiteReport(t, st, cid, dropJSON(c.to, true, "", "ok"))
oc.Check()
if n != c.wantAlarms {
t.Fatalf("%s: want %d alarm(s), got %d", c.name, c.wantAlarms, n)
}
})
}
}
// TestR431_RealHistoryProducesZeroAlarms — THE ACCEPTANCE STEP.
//
// Replays the ACTUAL snapshot-count history of both live boxes, exported from the hub's own reports
// table on 2026-09-01, through the real detector. It must produce ZERO alarms. A warning that fires
// on healthy boxes is worse than no warning at all.
//
// The fixture is committed beside this test so the assertion does not depend on a live database.
// If it is absent the test FAILS rather than skipping — a silent skip is how a green tick comes to
// mean nothing.
func TestR431_RealHistoryProducesZeroAlarms(t *testing.T) {
path := filepath.Join("testdata", "r431_real_history.json")
raw, err := os.ReadFile(path)
if err != nil {
t.Fatalf("the real-history fixture is missing (%v) — this is a FAILURE, never a skip: "+
"without it the acceptance step asserts nothing", err)
}
var hist map[string][]struct {
At string `json:"at"`
Count int `json:"count"`
StatsKnown bool `json:"stats_known"`
State string `json:"state"`
LastStatus string `json:"last_status"`
}
if err := json.Unmarshal(raw, &hist); err != nil {
t.Fatal(err)
}
if len(hist) == 0 {
t.Fatal("the fixture is empty — it would pass vacuously")
}
total := 0
for cust, points := range hist {
if len(points) < 2 {
t.Fatalf("%s: fewer than 2 points — nothing to compare", cust)
}
total += len(points)
st := newDiskStore(t)
var alarms int
var first = points[0]
saveOffsiteReport(t, st, cust, dropJSON(first.Count, first.StatsKnown, first.State, first.LastStatus))
oc := NewOffsiteChecker(st, 48*time.Hour, func(_, et, _, msg, _, _ string) {
if et == "offsite_snapshots_dropped" {
alarms++
t.Errorf("%s: FIRED ON REAL HISTORY: %s", cust, msg)
}
}, quietLog())
// Feed the remaining points straight through the detector. Driving Check() per point would
// need a 1.1s sleep each time (received_at is second-resolution) — hours for 5 000 points —
// so the sweep's own decision path is exercised directly instead, with the same guards.
for _, p := range points[1:] {
off := &offsiteReport{
SnapshotCount: p.Count, StatsKnown: p.StatsKnown, State: p.State,
LastStatus: p.LastStatus, LastSuccess: "2026-09-01T00:00:00Z",
}
if p.LastStatus != "ok" {
off.LastSuccess = ""
}
oc.mu.Lock()
if oc.countIsTrustworthy(off) {
dropped, prev, cur := oc.snapshotDropped(cust, off)
if dropped && oc.dropStates[cust] != "dropped" {
alarms++
t.Errorf("%s at %s: FIRED ON REAL HISTORY %d -> %d", cust, p.At, prev, cur)
oc.dropStates[cust] = "dropped"
} else if !dropped {
oc.dropStates[cust] = "ok"
}
oc.lastCounts[cust] = cur
}
oc.mu.Unlock()
}
if alarms != 0 {
t.Fatalf("%s: %d alarm(s) on real history — the threshold is wrong", cust, alarms)
}
}
// POSITIVE CONTROL: the replay must actually have looked at something. Without this the test
// passes when the fixture is a list of empty lists.
if total < 100 {
t.Fatalf("only %d points replayed — too few for this to mean anything", total)
}
t.Logf("replayed %d real report points across %d customers: ZERO alarms", total, len(hist))
}
File diff suppressed because one or more lines are too long