hub v0.111.0: notice a deletion within a day (R-431); correct R-429; re-scope R-95
gates / gates (push) Successful in 17s
gates / gates (push) Successful in 17s
THE RECORD WAS TELLING A WORSE STORY THAN THE TRUTH FOR TWO MONTHS, and my own probe is why.
R-429 CORRECTED. Yesterday's spike searched for a directory called `.snapshots`. The vendor documents
the path as /.zfs/snapshot. The probe's CONTROLS were sound and its SUBJECT was wrong, so "not found"
was true and meant nothing. Re-probed at the documented path on both boxes, with controls:
- /.zfs lists (shares, snapshot) from inside the jail;
- a write into /.zfs/snapshot is REFUSED - dest open ...: Failure - while the identical write to
the account home SUCCEEDS and was cleaned up.
That is the append-only property PROVEN rather than cited, and it is the sentence the whole re-scope
rests on. Seven daily snapshots are confirmed in the panel. The mitigation works. What remains true,
and was always the actual finding: the row claiming it had no R-number, its "confirm tomorrow" went
36 days unanswered, and the DUE-CHECKS block built for that class was empty. THE FINDING WAS NEVER THE
SNAPSHOTS - IT WAS THAT NOBODY COULD TELL.
R-95 RE-SCOPED: the box can delete its LIVE repository but cannot write to the daily snapshots of it,
so a deletion costs at most one day plus a per-file recovery - not open-ended loss. The ranking is
Viktor's; it has been #1 since July on the old story.
R-432 FILED: a sub-account sees /.zfs/snapshot EMPTY while the same box holds seven snapshots, so
per-file recovery is operator-only today. One panel read settles whether a NAMED snapshot can still be
entered, which would make it product-reachable.
R-431 SHIPPED. Third signal in OffsiteChecker. On the hub deliberately: a detector on the box is one
the deletion can silence. Threshold REASONED, not invented - over 12 898 reports every decrease lands
on ZERO and predates stats_known, and in the stats_known window there are none, so observed churn gave
nothing to calibrate against. Retention cannot halve a total; a mass deletion goes to ~0. Hence: more
than half, and at least 5. Guarded by StatsKnown (R-331), the declared State (R-204) and run success
(R-100 - whose lesson lives in this very file).
ACCEPTANCE: 9 009 real report points replayed through the detector produced ZERO alarms.
Three red-proofs run. The escalation one only became real after the first version was found HOLLOW -
it re-swept the same report, so the baseline had already moved and the latch was never consulted.
07 row 10's status is NOT moved: the write-refusal is measured, but the recovery ROUTE has never been
walked, which is what PARTIAL means.
This commit is contained in:
@@ -17,6 +17,9 @@ import (
|
||||
//
|
||||
// - FILL: repo_size_bytes vs the shared-model soft quota (quota_gb>0) at warn 90% / crit 95% — the
|
||||
// operator's early warning before the controller's own 100% run-refusal bites the customer.
|
||||
// - SNAPSHOT-DROP (R-431): the reported snapshot count falls by more than retention can explain —
|
||||
// the "something deleted this customer's off-site history" detector. See snapshotDrop below for
|
||||
// why it lives HERE and not on the box, and for the measurement behind its threshold.
|
||||
// - STALENESS: enabled + escrowed but no run in >48h (or never) — the silently-STUCK detector. A
|
||||
// RECENTLY-failing offsite is NOT stale (backup_failed already alerts it); staleness is the
|
||||
// complement: nothing is even trying. Pending/disabled targets are normal onboarding, never stale.
|
||||
@@ -36,6 +39,10 @@ type OffsiteChecker struct {
|
||||
mu sync.Mutex
|
||||
fillStates map[string]string // customerID → fill band
|
||||
staleStates map[string]string // customerID → "ok" | "stale"
|
||||
// R-431. dropStates is the escalation latch ("ok" | "dropped"); lastCounts is the baseline the
|
||||
// next sweep compares against. Only a TRUSTWORTHY report updates lastCounts — see snapshotDrop.
|
||||
dropStates map[string]string
|
||||
lastCounts map[string]int
|
||||
}
|
||||
|
||||
const defaultOffsiteStaleAfter = 48 * time.Hour
|
||||
@@ -52,6 +59,15 @@ type offsiteReport struct {
|
||||
SnapshotCount int `json:"snapshot_count"`
|
||||
RepoSizeBytes int64 `json:"repo_size_bytes"`
|
||||
QuotaGB int `json:"quota_gb"`
|
||||
// StatsKnown (R-331/R-225) — the ONLY thing separating "this repository holds nothing" from
|
||||
// "nobody has ever measured this repository". Both are `snapshot_count: 0` on the wire and they
|
||||
// are opposite news. ABSENT on a pre-v0.225.0 controller, and absence means the box CANNOT
|
||||
// ANSWER, never that the answer is no. snapshotDrop refuses to judge without it.
|
||||
StatsKnown bool `json:"stats_known,omitempty"`
|
||||
// State (R-204/R-193) — a DECLARED condition, empty on every healthy box. A box that has
|
||||
// re-initialised, lost its credential or been abandoned reports a real, correct, large drop.
|
||||
// The box knows its own situation; believe it rather than alarming on it.
|
||||
State string `json:"state,omitempty"`
|
||||
}
|
||||
|
||||
// NewOffsiteChecker builds the checker. Same seeding philosophy as StorageFillChecker: already-breached
|
||||
@@ -64,6 +80,7 @@ func NewOffsiteChecker(s *store.Store, staleAfter time.Duration, onEvent EventNo
|
||||
oc := &OffsiteChecker{
|
||||
store: s, logger: logger, onEvent: onEvent, staleAfter: staleAfter, now: time.Now,
|
||||
fillStates: make(map[string]string), staleStates: make(map[string]string),
|
||||
dropStates: make(map[string]string), lastCounts: make(map[string]int),
|
||||
}
|
||||
customers, err := s.GetCustomers()
|
||||
if err != nil {
|
||||
@@ -83,6 +100,13 @@ func NewOffsiteChecker(s *store.Store, staleAfter time.Duration, onEvent EventNo
|
||||
if !oc.isStale(c.CustomerID, off) {
|
||||
oc.staleStates[c.CustomerID] = "ok"
|
||||
}
|
||||
// R-431: seed the baseline from the current report so a hub restart does not read the first
|
||||
// sweep as a drop from nothing. Only a trustworthy report seeds — an untrustworthy one leaves
|
||||
// no baseline, and no baseline means no verdict.
|
||||
if oc.countIsTrustworthy(off) {
|
||||
oc.lastCounts[c.CustomerID] = off.SnapshotCount
|
||||
oc.dropStates[c.CustomerID] = "ok"
|
||||
}
|
||||
}
|
||||
logger.Printf("[INFO] Offsite checker initialized: fill warn=90%% crit=95%%, stale after %s, %d ok-seeded", staleAfter, seeded)
|
||||
return oc
|
||||
@@ -175,6 +199,124 @@ func (oc *OffsiteChecker) isStale(customerID string, off *offsiteReport) bool {
|
||||
return oc.now().Sub(t) > oc.staleAfter
|
||||
}
|
||||
|
||||
// ── SNAPSHOT-DROP (R-431) ───────────────────────────────────────────────────────────────────────
|
||||
//
|
||||
// WHY THIS LIVES ON THE HUB AND NOT ON THE BOX. The thing being detected is a box deleting its own
|
||||
// off-site backups — so a detector living on that box is a detector the same event can silence. The
|
||||
// hub already receives the count in every report and already keeps the history, so it can notice
|
||||
// without trusting the box's judgement. The box still REPORTS the number, and a sophisticated
|
||||
// attacker could lie about it; that is a far higher bar than deleting files, and the weekly integrity
|
||||
// check (R-359) would then disagree with the lie.
|
||||
//
|
||||
// WHAT IT IS FOR. R-95: the restic credential can delete its own repository, and the sub-account API
|
||||
// has no append-only axis, so prevention needs a transport change. The Storage Box's daily ZFS
|
||||
// snapshots bound the loss — MEASURED 2026-09-01: a write into `/.zfs/snapshot` is REFUSED
|
||||
// (`dest open …: Failure`) while the same write to the account home succeeds. So the remedy exists;
|
||||
// what was missing was NOTICING, and this is that.
|
||||
//
|
||||
// THE THRESHOLD, AND THE MEASUREMENT IT CAME FROM. Measured over the hub's own `reports` table on
|
||||
// 2026-09-01: 12 898 reports, 4 customers, 2026-06-05 → 2026-09-01.
|
||||
//
|
||||
// * In the whole history there are NINE decreases, and EVERY ONE of them lands exactly on ZERO
|
||||
// (36→0, 18→0 ×2, 15→0, 12→0, 8→0, 3→0). There is not one gradual retention decrease anywhere.
|
||||
// * Every one of those nine predates `stats_known`, i.e. they are the R-331 shape — a zero that
|
||||
// means "could not measure", not "nothing is there". Several carry a declared State
|
||||
// (`needs_credential`, `awaiting_recovery_key`), which says so outright.
|
||||
// * In the window where `stats_known` is TRUE (380 reports across both live boxes) there are ZERO
|
||||
// decreases: demo-felhom sat flat at 10, demo-hp moved 67→68→69. Only rises.
|
||||
//
|
||||
// So observed retention churn gives NOTHING to calibrate against, and saying so is the honest answer
|
||||
// rather than inventing a number (R-401's lesson). The threshold is therefore reasoned from what
|
||||
// retention CAN do, not from what it was seen to do:
|
||||
//
|
||||
// the box runs `forget --keep-daily 7 --keep-weekly 4 --keep-monthly 6 --group-by host,tags`
|
||||
// over ~8 apps. On a boundary day several groups can expire at once, so a legitimate pass can
|
||||
// plausibly remove low double digits. **What it can NEVER do is halve the total**: keeping 7 daily
|
||||
// + 4 weekly + 6 monthly per group is a floor, and a mass deletion goes to ~0.
|
||||
//
|
||||
// Hence: **a fall of MORE THAN HALF the previous count, and at least 5 snapshots.** The 50% cannot be
|
||||
// reached by retention; the floor of 5 stops a tiny-count box alarming on ordinary ageing. It is
|
||||
// deliberately NOT sensitive — a detector that cries wolf is switched off within a fortnight, and
|
||||
// this project has proved that twice in a week.
|
||||
const (
|
||||
snapshotDropFraction = 0.5 // more than half the history gone in one step
|
||||
snapshotDropFloor = 5 // and at least this many, so small counts do not twitch
|
||||
)
|
||||
|
||||
// countIsTrustworthy — the three pre-conditions, each with a scar behind it. A report failing ANY of
|
||||
// them is not evidence of anything: it neither alarms NOR updates the baseline, because comparing
|
||||
// against a number nobody could measure is how a detector invents its own findings.
|
||||
func (oc *OffsiteChecker) countIsTrustworthy(off *offsiteReport) bool {
|
||||
// 1. R-331 — a zero is not a zero when stats are unknown. Absent `stats_known` means the box
|
||||
// cannot answer. Every historical decrease in this hub's data is this shape.
|
||||
if !off.StatsKnown {
|
||||
return false
|
||||
}
|
||||
// 2. R-204/R-193 — the box DECLARES its own situation. A re-initialised, credential-less or
|
||||
// abandoned box reports a real, correct drop; alarming on it would be blaming the box for
|
||||
// telling the truth.
|
||||
if off.State != "" {
|
||||
return false
|
||||
}
|
||||
// 3. R-100 — PRESENCE IS NOT SUCCESS, and this file is where that lesson was learned. A count
|
||||
// from a failed or half-finished run is not a measurement of the store. "incomplete" (R-203)
|
||||
// is deliberately excluded too: a partial run legitimately counts less.
|
||||
if off.LastStatus != "ok" || off.LastSuccess == "" {
|
||||
return false
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
// snapshotDropped reports whether `off` is a drop worth alarming about, against the remembered
|
||||
// baseline. Caller holds oc.mu.
|
||||
func (oc *OffsiteChecker) snapshotDropped(customerID string, off *offsiteReport) (bool, int, int) {
|
||||
prev, had := oc.lastCounts[customerID]
|
||||
if !had || prev <= 0 {
|
||||
return false, prev, off.SnapshotCount
|
||||
}
|
||||
drop := prev - off.SnapshotCount
|
||||
if drop < snapshotDropFloor {
|
||||
return false, prev, off.SnapshotCount
|
||||
}
|
||||
return float64(drop) > float64(prev)*snapshotDropFraction, prev, off.SnapshotCount
|
||||
}
|
||||
|
||||
// emitSnapshotDrop — one event, at a severity inside the hub's exact vocabulary (08 §6.1).
|
||||
//
|
||||
// THE MESSAGE MUST NOT SAY THE DATA IS LOST, because after the 2026-09-01 measurement that is usually
|
||||
// false: the daily Storage Box snapshots are read-only to every account (proven, not cited) and hold
|
||||
// the older copy. It says what happened, what it means, and where the data still is.
|
||||
func (oc *OffsiteChecker) emitSnapshotDrop(customerID string, off *offsiteReport, prev, cur int) {
|
||||
message := fmt.Sprintf(
|
||||
"Customer %s: off-site backup count fell from %d to %d snapshot(s) in one report — more than "+
|
||||
"retention can explain. The daily Storage Box snapshots are read-only and still hold the "+
|
||||
"older copy, so this is recoverable file-by-file; it is NOT confirmed data loss. Check "+
|
||||
"whether a deletion ran on the box before restoring anything.",
|
||||
customerID, prev, cur)
|
||||
details, _ := json.Marshal(map[string]any{
|
||||
"customer_id": customerID, "previous_count": prev, "current_count": cur,
|
||||
"drop": prev - cur, "last_success": off.LastSuccess, "last_status": off.LastStatus,
|
||||
})
|
||||
oc.logger.Printf("[INFO] Offsite snapshot drop: %s %d -> %d (offsite_snapshots_dropped)", customerID, prev, cur)
|
||||
if _, err := oc.store.SaveEvent(customerID, "offsite_snapshots_dropped", "error", message, string(details), "hub"); err != nil {
|
||||
oc.logger.Printf("[WARN] Failed to save offsite snapshot-drop event for %s: %v", customerID, err)
|
||||
return
|
||||
}
|
||||
if oc.onEvent != nil {
|
||||
oc.onEvent(customerID, "offsite_snapshots_dropped", "error", message, string(details), "hub")
|
||||
}
|
||||
}
|
||||
|
||||
// GetDropState exposes the latch for tests.
|
||||
func (oc *OffsiteChecker) GetDropState(customerID string) string {
|
||||
oc.mu.Lock()
|
||||
defer oc.mu.Unlock()
|
||||
if s := oc.dropStates[customerID]; s != "" {
|
||||
return s
|
||||
}
|
||||
return "unknown"
|
||||
}
|
||||
|
||||
// warnLegacyOnce logs the legacy degrade a single time per customer. Once, because this is a
|
||||
// steady-state condition until the box upgrades — a per-cycle line would be pure noise — but it must be
|
||||
// logged at all, so a fleet silently running on the old anchor is visible rather than assumed.
|
||||
@@ -228,11 +370,15 @@ func (oc *OffsiteChecker) Check() {
|
||||
if off == nil {
|
||||
delete(oc.fillStates, c.CustomerID) // vanished object (disabled / downgraded) → re-arm
|
||||
delete(oc.staleStates, c.CustomerID)
|
||||
delete(oc.dropStates, c.CustomerID)
|
||||
delete(oc.lastCounts, c.CustomerID) // no baseline survives a vanished object
|
||||
continue
|
||||
}
|
||||
if oc.store.IsCustomerBlocked(c.CustomerID) {
|
||||
delete(oc.fillStates, c.CustomerID)
|
||||
delete(oc.staleStates, c.CustomerID)
|
||||
delete(oc.dropStates, c.CustomerID)
|
||||
delete(oc.lastCounts, c.CustomerID)
|
||||
continue
|
||||
}
|
||||
|
||||
@@ -261,6 +407,24 @@ func (oc *OffsiteChecker) Check() {
|
||||
}
|
||||
}
|
||||
oc.staleStates[c.CustomerID] = newStale
|
||||
|
||||
// SNAPSHOT-DROP (R-431). Escalation-only, exactly like the two signals above: one deletion
|
||||
// produces ONE alarm, not one per report cycle. The baseline then moves to the new value, so a
|
||||
// SECOND deletion later is still caught — from the new floor.
|
||||
//
|
||||
// An UNTRUSTWORTHY report is a no-op in both directions: no alarm, and the baseline is left
|
||||
// alone rather than being overwritten with a number nobody could measure. That is what stops
|
||||
// a `needs_credential` zero from becoming the baseline and making the RECOVERY look like a rise.
|
||||
if oc.countIsTrustworthy(off) {
|
||||
dropped, prev, cur := oc.snapshotDropped(c.CustomerID, off)
|
||||
if dropped && oc.dropStates[c.CustomerID] != "dropped" {
|
||||
oc.emitSnapshotDrop(c.CustomerID, off, prev, cur)
|
||||
oc.dropStates[c.CustomerID] = "dropped"
|
||||
} else if !dropped {
|
||||
oc.dropStates[c.CustomerID] = "ok"
|
||||
}
|
||||
oc.lastCounts[c.CustomerID] = cur
|
||||
}
|
||||
}
|
||||
for k := range oc.fillStates {
|
||||
if !seen[k] {
|
||||
@@ -272,6 +436,12 @@ func (oc *OffsiteChecker) Check() {
|
||||
delete(oc.staleStates, k)
|
||||
}
|
||||
}
|
||||
for k := range oc.dropStates {
|
||||
if !seen[k] {
|
||||
delete(oc.dropStates, k)
|
||||
delete(oc.lastCounts, k)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// GetFillState / GetStaleState expose current states for tests.
|
||||
|
||||
@@ -0,0 +1,278 @@
|
||||
package monitor
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// R-431 — the snapshot-drop signal, proven in BOTH directions.
|
||||
//
|
||||
// The acceptance test is TestR431_RealHistoryProducesZeroAlarms: a detector that fires on healthy
|
||||
// boxes is the mistake this project caught twice in one week, and it is the one that would get this
|
||||
// signal switched off within a fortnight.
|
||||
|
||||
// dropJSON builds a trustworthy offsite object (stats_known, no declared state, last run ok).
|
||||
func dropJSON(count int, statsKnown bool, state, lastStatus string) string {
|
||||
ts := time.Now().UTC().Add(-1 * time.Hour).Format(time.RFC3339)
|
||||
success := ts
|
||||
if lastStatus != "ok" {
|
||||
success = ""
|
||||
}
|
||||
return fmt.Sprintf(
|
||||
`{"enabled":true,"escrow_state":"escrowed","last_run":%q,"last_status":%q,"last_success":%q,`+
|
||||
`"snapshot_count":%d,"repo_size_bytes":1073741824,"quota_gb":0,"stats_known":%v,"state":%q}`,
|
||||
ts, lastStatus, success, count, statsKnown, state)
|
||||
}
|
||||
|
||||
// TestR431_FiresOnAMassDeletion — direction 1. A drop past the threshold alarms EXACTLY once.
|
||||
//
|
||||
// RED-PROOF (run 2026-09-01, recorded in REPORT.md): setting snapshotDropFraction to 0.99 makes this
|
||||
// fail — 69 → 4 is a 94% fall and would no longer qualify, which is what an over-loose threshold
|
||||
// looks like in production.
|
||||
func TestR431_FiresOnAMassDeletion(t *testing.T) {
|
||||
st := newDiskStore(t)
|
||||
var got []struct{ et, sev, msg string }
|
||||
saveOffsiteReport(t, st, "victim", dropJSON(69, true, "", "ok"))
|
||||
oc := NewOffsiteChecker(st, 48*time.Hour, func(_, et, sev, msg, _, _ string) {
|
||||
got = append(got, struct{ et, sev, msg string }{et, sev, msg})
|
||||
}, quietLog())
|
||||
|
||||
// the constructor seeded the baseline at 69; now the store is emptied
|
||||
saveOffsiteReport(t, st, "victim", dropJSON(4, true, "", "ok"))
|
||||
oc.Check()
|
||||
|
||||
var drops []struct{ et, sev, msg string }
|
||||
for _, g := range got {
|
||||
if g.et == "offsite_snapshots_dropped" {
|
||||
drops = append(drops, g)
|
||||
}
|
||||
}
|
||||
if len(drops) != 1 {
|
||||
t.Fatalf("want exactly 1 offsite_snapshots_dropped, got %d (%v)", len(drops), got)
|
||||
}
|
||||
// The severity MUST be in the hub's exact vocabulary — anything else is coerced to info and
|
||||
// mailed to nobody (08 §6.1, shipped twice).
|
||||
switch drops[0].sev {
|
||||
case "info", "warning", "error", "critical":
|
||||
default:
|
||||
t.Fatalf("severity %q is outside the hub vocabulary — it would be coerced to info and reach nobody", drops[0].sev)
|
||||
}
|
||||
for _, frag := range []string{"69", "4", "read-only", "NOT confirmed data loss"} {
|
||||
if !strings.Contains(drops[0].msg, frag) {
|
||||
t.Fatalf("message must contain %q; got: %s", frag, drops[0].msg)
|
||||
}
|
||||
}
|
||||
if strings.Contains(drops[0].msg, "data is lost") || strings.Contains(drops[0].msg, "data lost") {
|
||||
t.Fatalf("the message must NOT claim data loss — the snapshots usually still hold it: %s", drops[0].msg)
|
||||
}
|
||||
|
||||
}
|
||||
|
||||
// TestR431_EscalationOnlyLatch — a CONTINUING deletion must not page on every report cycle.
|
||||
//
|
||||
// WRITTEN AFTER A HOLLOW FIRST ATTEMPT, and the failure is recorded because it is instructive: the
|
||||
// original assertion re-swept the SAME report and called that "escalation-only". It proved nothing —
|
||||
// the baseline had already moved to the new count, so the second sweep saw a drop of zero and the
|
||||
// latch was never consulted. Its red-proof (removing the latch) PASSED, which is how it was caught.
|
||||
//
|
||||
// This drives a count that keeps FALLING, which is the only shape where the latch is load-bearing.
|
||||
//
|
||||
// RED-PROOF (run 2026-09-01): replacing `dropped && oc.dropStates[...] != "dropped"` with `dropped`
|
||||
// makes this fail with 2 alarms — one per report cycle, which is the noise that trains an operator
|
||||
// to ignore the alarm.
|
||||
func TestR431_EscalationOnlyLatch(t *testing.T) {
|
||||
st := newDiskStore(t)
|
||||
var n int
|
||||
saveOffsiteReport(t, st, "sliding", dropJSON(69, true, "", "ok"))
|
||||
oc := NewOffsiteChecker(st, 48*time.Hour, func(_, et, _, _, _, _ string) {
|
||||
if et == "offsite_snapshots_dropped" {
|
||||
n++
|
||||
}
|
||||
}, quietLog())
|
||||
|
||||
saveOffsiteReport(t, st, "sliding", dropJSON(30, true, "", "ok")) // 69 -> 30: alarm
|
||||
oc.Check()
|
||||
if n != 1 {
|
||||
t.Fatalf("the first large drop must alarm exactly once; got %d", n)
|
||||
}
|
||||
saveOffsiteReport(t, st, "sliding", dropJSON(2, true, "", "ok")) // 30 -> 2: still falling
|
||||
oc.Check()
|
||||
if n != 1 {
|
||||
t.Fatalf("a CONTINUING deletion must not re-page while the latch is set; got %d alarms", n)
|
||||
}
|
||||
if oc.GetDropState("sliding") != "dropped" {
|
||||
t.Fatalf("the latch must be held, got %q", oc.GetDropState("sliding"))
|
||||
}
|
||||
|
||||
// RECOVERY RE-ARMS: a clean sweep clears the latch, so a LATER deletion is caught again.
|
||||
saveOffsiteReport(t, st, "sliding", dropJSON(40, true, "", "ok")) // rebuilt, no drop
|
||||
oc.Check()
|
||||
if oc.GetDropState("sliding") != "ok" {
|
||||
t.Fatalf("a clean sweep must re-arm the latch, got %q", oc.GetDropState("sliding"))
|
||||
}
|
||||
saveOffsiteReport(t, st, "sliding", dropJSON(1, true, "", "ok")) // deleted again
|
||||
oc.Check()
|
||||
if n != 2 {
|
||||
t.Fatalf("after re-arming, a NEW deletion must alarm again; got %d", n)
|
||||
}
|
||||
}
|
||||
|
||||
// TestR431_SilentWhenNotTrustworthy — direction 2, the three pre-conditions, each with its scar.
|
||||
func TestR431_SilentWhenNotTrustworthy(t *testing.T) {
|
||||
cases := []struct {
|
||||
name, first, second string
|
||||
}{
|
||||
{"stats_known absent (R-331: a zero that means UNMEASURED)",
|
||||
dropJSON(69, true, "", "ok"), dropJSON(0, false, "", "ok")},
|
||||
{"a DECLARED state (R-204: the box says what happened)",
|
||||
dropJSON(69, true, "", "ok"), dropJSON(0, true, "needs_credential", "ok")},
|
||||
{"the run FAILED (R-100: presence is not success)",
|
||||
dropJSON(69, true, "", "ok"), dropJSON(0, true, "", "error")},
|
||||
{"the run was INCOMPLETE (R-203: a partial run counts less)",
|
||||
dropJSON(69, true, "", "ok"), dropJSON(0, true, "", "incomplete")},
|
||||
}
|
||||
for i, c := range cases {
|
||||
t.Run(c.name, func(t *testing.T) {
|
||||
st := newDiskStore(t)
|
||||
cid := fmt.Sprintf("c%d", i)
|
||||
var n int
|
||||
saveOffsiteReport(t, st, cid, c.first)
|
||||
oc := NewOffsiteChecker(st, 48*time.Hour, func(_, et, _, _, _, _ string) {
|
||||
if et == "offsite_snapshots_dropped" {
|
||||
n++
|
||||
}
|
||||
}, quietLog())
|
||||
saveOffsiteReport(t, st, cid, c.second)
|
||||
oc.Check()
|
||||
if n != 0 {
|
||||
t.Fatalf("%s: must NOT alarm; got %d", c.name, n)
|
||||
}
|
||||
// AND the baseline must be untouched, so the RECOVERY does not read as a rise-then-drop.
|
||||
oc.mu.Lock()
|
||||
base := oc.lastCounts[cid]
|
||||
oc.mu.Unlock()
|
||||
if base != 69 {
|
||||
t.Fatalf("%s: an untrustworthy report must not overwrite the baseline; got %d", c.name, base)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestR431_OrdinaryRetentionIsSilent — a fall that retention CAN explain must not alarm.
|
||||
func TestR431_OrdinaryRetentionIsSilent(t *testing.T) {
|
||||
for _, c := range []struct {
|
||||
name string
|
||||
from, to int
|
||||
wantAlarms int
|
||||
}{
|
||||
{"69 -> 60 (9 gone, under half)", 69, 60, 0},
|
||||
{"10 -> 7 (3 gone, under the floor of 5)", 10, 7, 0},
|
||||
{"10 -> 5 (5 gone, exactly half — NOT more than half)", 10, 5, 0},
|
||||
{"69 -> 34 (35 gone, more than half)", 69, 34, 1},
|
||||
} {
|
||||
t.Run(c.name, func(t *testing.T) {
|
||||
st := newDiskStore(t)
|
||||
cid := fmt.Sprintf("r%d%d", c.from, c.to)
|
||||
var n int
|
||||
saveOffsiteReport(t, st, cid, dropJSON(c.from, true, "", "ok"))
|
||||
oc := NewOffsiteChecker(st, 48*time.Hour, func(_, et, _, _, _, _ string) {
|
||||
if et == "offsite_snapshots_dropped" {
|
||||
n++
|
||||
}
|
||||
}, quietLog())
|
||||
saveOffsiteReport(t, st, cid, dropJSON(c.to, true, "", "ok"))
|
||||
oc.Check()
|
||||
if n != c.wantAlarms {
|
||||
t.Fatalf("%s: want %d alarm(s), got %d", c.name, c.wantAlarms, n)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestR431_RealHistoryProducesZeroAlarms — THE ACCEPTANCE STEP.
|
||||
//
|
||||
// Replays the ACTUAL snapshot-count history of both live boxes, exported from the hub's own reports
|
||||
// table on 2026-09-01, through the real detector. It must produce ZERO alarms. A warning that fires
|
||||
// on healthy boxes is worse than no warning at all.
|
||||
//
|
||||
// The fixture is committed beside this test so the assertion does not depend on a live database.
|
||||
// If it is absent the test FAILS rather than skipping — a silent skip is how a green tick comes to
|
||||
// mean nothing.
|
||||
func TestR431_RealHistoryProducesZeroAlarms(t *testing.T) {
|
||||
path := filepath.Join("testdata", "r431_real_history.json")
|
||||
raw, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
t.Fatalf("the real-history fixture is missing (%v) — this is a FAILURE, never a skip: "+
|
||||
"without it the acceptance step asserts nothing", err)
|
||||
}
|
||||
var hist map[string][]struct {
|
||||
At string `json:"at"`
|
||||
Count int `json:"count"`
|
||||
StatsKnown bool `json:"stats_known"`
|
||||
State string `json:"state"`
|
||||
LastStatus string `json:"last_status"`
|
||||
}
|
||||
if err := json.Unmarshal(raw, &hist); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if len(hist) == 0 {
|
||||
t.Fatal("the fixture is empty — it would pass vacuously")
|
||||
}
|
||||
|
||||
total := 0
|
||||
for cust, points := range hist {
|
||||
if len(points) < 2 {
|
||||
t.Fatalf("%s: fewer than 2 points — nothing to compare", cust)
|
||||
}
|
||||
total += len(points)
|
||||
st := newDiskStore(t)
|
||||
var alarms int
|
||||
var first = points[0]
|
||||
saveOffsiteReport(t, st, cust, dropJSON(first.Count, first.StatsKnown, first.State, first.LastStatus))
|
||||
oc := NewOffsiteChecker(st, 48*time.Hour, func(_, et, _, msg, _, _ string) {
|
||||
if et == "offsite_snapshots_dropped" {
|
||||
alarms++
|
||||
t.Errorf("%s: FIRED ON REAL HISTORY: %s", cust, msg)
|
||||
}
|
||||
}, quietLog())
|
||||
// Feed the remaining points straight through the detector. Driving Check() per point would
|
||||
// need a 1.1s sleep each time (received_at is second-resolution) — hours for 5 000 points —
|
||||
// so the sweep's own decision path is exercised directly instead, with the same guards.
|
||||
for _, p := range points[1:] {
|
||||
off := &offsiteReport{
|
||||
SnapshotCount: p.Count, StatsKnown: p.StatsKnown, State: p.State,
|
||||
LastStatus: p.LastStatus, LastSuccess: "2026-09-01T00:00:00Z",
|
||||
}
|
||||
if p.LastStatus != "ok" {
|
||||
off.LastSuccess = ""
|
||||
}
|
||||
oc.mu.Lock()
|
||||
if oc.countIsTrustworthy(off) {
|
||||
dropped, prev, cur := oc.snapshotDropped(cust, off)
|
||||
if dropped && oc.dropStates[cust] != "dropped" {
|
||||
alarms++
|
||||
t.Errorf("%s at %s: FIRED ON REAL HISTORY %d -> %d", cust, p.At, prev, cur)
|
||||
oc.dropStates[cust] = "dropped"
|
||||
} else if !dropped {
|
||||
oc.dropStates[cust] = "ok"
|
||||
}
|
||||
oc.lastCounts[cust] = cur
|
||||
}
|
||||
oc.mu.Unlock()
|
||||
}
|
||||
if alarms != 0 {
|
||||
t.Fatalf("%s: %d alarm(s) on real history — the threshold is wrong", cust, alarms)
|
||||
}
|
||||
}
|
||||
// POSITIVE CONTROL: the replay must actually have looked at something. Without this the test
|
||||
// passes when the fixture is a list of empty lists.
|
||||
if total < 100 {
|
||||
t.Fatalf("only %d points replayed — too few for this to mean anything", total)
|
||||
}
|
||||
t.Logf("replayed %d real report points across %d customers: ZERO alarms", total, len(hist))
|
||||
}
|
||||
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user