Files
felhom.eu/hub/internal/notify/dispatcher_box_reachability_test.go
T
admin ab2262c91c
gates / gates (push) Successful in 14s
hub v0.106.0: report loss of visibility into the off-site stores (R-339)
THE GAP, measured not supposed. On 2026-08-18 ep0's PBS proxy was wedged for
9 h 37 m and the hub emitted NOTHING on the operator channel. Both box
checkers hold their last snapshot and return silently on a failed fetch --
correct for a FILL signal, since a missing reading must never be read as 0%,
but it makes a dead off-site endpoint and a healthy one indistinguishable.
The only mails that morning came from the boxes' own backup failures, and
only because the WEEKLY offsite run happened to land inside the window. Two
days earlier nothing would have fired at all.

REACHABILITY is now a second, independent signal on both checkers:
consecutive failed fetch windows, reported past a default 3 windows
(~30-45 min) as pbsdr_box_unreachable / offsite_box_unreachable (warning) on
the customer-less pbsdr-box / pool-box scopes, each with a paired *_recovered
all-clear. Tunable via alerting.box_unreachable_windows (0/invalid -> 3).

THE FILL LOGIC IS UNTOUCHED. No threshold, throttle, band or escalate-once
behaviour changed; a degraded read still drives no transition.

Three decisions a later reader would otherwise "fix" back, so each is
argued in-code:
  - the unreachable event REPEATS rather than escalating once. The band shape
    would give exactly ONE mail at ~minute 30 of a nine-hour outage, and one
    mail is missable. It leans on the dispatcher's 1 h operator cooldown to
    become an hourly "still blind" heartbeat.
  - ErrUsageUnsupported is NOT blindness: an old ep0 answers "no such op",
    which means we reached it. Counting it would alert for days on a healthy
    pre-update endpoint.
  - born-blind is reported: the counter is not gated on having a snapshot, so
    a hub restarted INTO an outage still speaks. last_ok is OMITTED rather
    than zero-valued -- a fabricated timestamp reads as "it was fine until
    then".

Both recoveries are severity "info" and severityNotifies drops "info", so
they are registered in recoveredPairedDownTypes or the operator hears that
the tier broke and never that it healed. A cross-package test drives
ProcessEvent and asserts an actual operator MAIL, not a map entry -- a green
checker test proves nothing about the seam (agent v0.91.0 shipped fully green
with SetAuthSink never called).

Tests: box_reachability_test.go (Scenarios A-F) + dispatcher_box_reachability
_test.go (wiring). Three red-proofs run and reverted, each seen failing with a
message naming the right cause: threshold 3->1, the sentinel counter guard,
the pairing entry.

Register: R-339 filed and marked SHIPPED (PROVEN-LIVE still owed -- no real or
constructed outage has exercised the emit path, and one cannot be manufactured
against Tier-2 ep0). R-340 filed: the reachability read rides ep0's LOCAL API
daemon, which the incident explicitly cleared, so this check would have shown
GREEN for all 9 h 37 m -- the honest boundary, recorded rather than glossed.
R-336's next-step corrected: pvestatd's interval is NOT tunable (Proxmox staff
have said so); the only lever is disabling the storage entry, which collides
with the agent's consume-the-one-time-secret path. Doc-only, no agent code
touched.
2026-08-18 19:27:33 +02:00

101 lines
3.9 KiB
Go

package notify
import (
"io"
"log"
"testing"
)
// R-339, Group G — THE WIRING TEST, and it is not optional.
//
// The recovery leg spans two packages: internal/monitor emits, internal/notify routes. A fake-injected
// test in monitor proves the checker emits and proves NOTHING about whether the operator receives the
// mail. Both `*_box_recovered` events carry severity "info", and severityNotifies drops "info" — so
// without an entry in recoveredPairedDownTypes they are stored and never mailed, and the operator is
// told the off-site tier went blind and never told it came back.
//
// Precedent for why this test exists at all: agent v0.91.0 shipped fully green with SetAuthSink never
// called from main.go, and the entire auth-honesty leg was inert. A green unit test on one side of a
// seam is not evidence that the seam is connected.
//
// These drive ProcessEvent end-to-end and assert a MAIL, not a map entry. A map assertion would pass
// on a correctly-populated map that nothing reads.
func TestBoxRecovery_ReachesTheOperatorDespiteInfoSeverity(t *testing.T) {
if severityNotifies("info") {
t.Fatal("premise changed: \"info\" now notifies, so the pairing entries may be unnecessary — re-check")
}
for _, et := range []string{"pbsdr_box_recovered", "offsite_box_recovered"} {
st := newDispStore(t)
d := NewDispatcher(st, "test-key", "hub@felhom.eu", "op@felhom.eu", true, log.New(io.Discard, "", 0))
mails := captureSeam(d)
scope := "pbsdr-box"
if et == "offsite_box_recovered" {
scope = "pool-box"
}
d.ProcessEvent(scope, et, "info", "reachable again after 45m0s", `{"scope":"`+scope+`"}`, "hub")
op := mailsFor(*mails, "op@felhom.eu")
if len(op) != 1 {
t.Fatalf("%s: operator mails = %d, want 1 — the all-clear must reach the operator; "+
"0 means the recoveredPairedDownTypes entry is missing and \"info\" was dropped by the severity gate",
et, len(op))
}
// Customer leg is a no-op BY CONSTRUCTION: the scope is not a customer id, so no prefs row
// exists and no paired customer "sent" row can be found.
if len(*mails) != 1 {
t.Fatalf("%s: total mails = %d, want 1 — a customer-less scope must never produce a customer mail",
et, len(*mails))
}
}
}
// The DOWN edge needs no pairing entry — "warning" already notifies — but if that ever changed the
// operator would hear the all-clear for an outage they were never told about. Pin both edges.
func TestBoxUnreachable_ReachesTheOperatorOnItsOwnSeverity(t *testing.T) {
for _, c := range []struct{ scope, et string }{
{"pbsdr-box", "pbsdr_box_unreachable"},
{"pool-box", "offsite_box_unreachable"},
} {
st := newDispStore(t)
d := NewDispatcher(st, "test-key", "hub@felhom.eu", "op@felhom.eu", true, log.New(io.Discard, "", 0))
mails := captureSeam(d)
d.ProcessEvent(c.scope, c.et, "warning", "unreachable — 3 consecutive checks failed", `{"scope":"`+c.scope+`"}`, "hub")
op := mailsFor(*mails, "op@felhom.eu")
if len(op) != 1 {
t.Fatalf("%s: operator mails = %d, want 1", c.et, len(op))
}
if len(*mails) != 1 {
t.Fatalf("%s: total mails = %d, want 1 — operator-only", c.et, len(*mails))
}
}
}
// Both recovery types must be registered against the RIGHT down type. A recovery paired with the
// wrong down event would still mail the operator (processOperator runs unconditionally) while
// silently breaking the customer pairing rule for any future customer-scoped reuse.
func TestBoxRecovery_PairedWithTheCorrectDownType(t *testing.T) {
want := map[string]string{
"pbsdr_box_recovered": "pbsdr_box_unreachable",
"offsite_box_recovered": "offsite_box_unreachable",
}
for rec, down := range want {
paired, ok := recoveredPairedDownTypes[rec]
if !ok {
t.Fatalf("%s is not on the recovery branch — its \"info\" severity makes it silent", rec)
}
found := false
for _, p := range paired {
if p == down {
found = true
}
}
if !found {
t.Fatalf("%s must pair with %s; got %v", rec, down, paired)
}
}
}