hub v0.106.0: report loss of visibility into the off-site stores (R-339)
gates / gates (push) Successful in 14s

THE GAP, measured not supposed. On 2026-08-18 ep0's PBS proxy was wedged for
9 h 37 m and the hub emitted NOTHING on the operator channel. Both box
checkers hold their last snapshot and return silently on a failed fetch --
correct for a FILL signal, since a missing reading must never be read as 0%,
but it makes a dead off-site endpoint and a healthy one indistinguishable.
The only mails that morning came from the boxes' own backup failures, and
only because the WEEKLY offsite run happened to land inside the window. Two
days earlier nothing would have fired at all.

REACHABILITY is now a second, independent signal on both checkers:
consecutive failed fetch windows, reported past a default 3 windows
(~30-45 min) as pbsdr_box_unreachable / offsite_box_unreachable (warning) on
the customer-less pbsdr-box / pool-box scopes, each with a paired *_recovered
all-clear. Tunable via alerting.box_unreachable_windows (0/invalid -> 3).

THE FILL LOGIC IS UNTOUCHED. No threshold, throttle, band or escalate-once
behaviour changed; a degraded read still drives no transition.

Three decisions a later reader would otherwise "fix" back, so each is
argued in-code:
  - the unreachable event REPEATS rather than escalating once. The band shape
    would give exactly ONE mail at ~minute 30 of a nine-hour outage, and one
    mail is missable. It leans on the dispatcher's 1 h operator cooldown to
    become an hourly "still blind" heartbeat.
  - ErrUsageUnsupported is NOT blindness: an old ep0 answers "no such op",
    which means we reached it. Counting it would alert for days on a healthy
    pre-update endpoint.
  - born-blind is reported: the counter is not gated on having a snapshot, so
    a hub restarted INTO an outage still speaks. last_ok is OMITTED rather
    than zero-valued -- a fabricated timestamp reads as "it was fine until
    then".

Both recoveries are severity "info" and severityNotifies drops "info", so
they are registered in recoveredPairedDownTypes or the operator hears that
the tier broke and never that it healed. A cross-package test drives
ProcessEvent and asserts an actual operator MAIL, not a map entry -- a green
checker test proves nothing about the seam (agent v0.91.0 shipped fully green
with SetAuthSink never called).

Tests: box_reachability_test.go (Scenarios A-F) + dispatcher_box_reachability
_test.go (wiring). Three red-proofs run and reverted, each seen failing with a
message naming the right cause: threshold 3->1, the sentinel counter guard,
the pairing entry.

Register: R-339 filed and marked SHIPPED (PROVEN-LIVE still owed -- no real or
constructed outage has exercised the emit path, and one cannot be manufactured
against Tier-2 ep0). R-340 filed: the reachability read rides ep0's LOCAL API
daemon, which the incident explicitly cleared, so this check would have shown
GREEN for all 9 h 37 m -- the honest boundary, recorded rather than glossed.
R-336's next-step corrected: pvestatd's interval is NOT tunable (Proxmox staff
have said so); the only lever is disabling the storage entry, which collides
with the agent's consume-the-one-time-secret path. Doc-only, no agent code
touched.
This commit is contained in:
2026-08-18 19:27:33 +02:00
parent 78a244bf09
commit ab2262c91c
13 changed files with 765 additions and 21 deletions
+25
View File
@@ -15,6 +15,31 @@
> would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`**
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
## Box REACHABILITY is a separate signal from box FILL — and the cadence difference is deliberate (2026-08-18, R-339, hub v0.106.0)
Both off-site checkers now carry two independent signals, and conflating them is the mistake to avoid:
- **FILL** — how full is the store. Escalation-only (one emit on a band rise), and **a degraded or
missing reading drives NO band transition**, because missing ≠ 0%. That rule is unchanged by
v0.106.0 and must stay unchanged: it is what stops a dead endpoint from reading as an empty one.
- **REACHABILITY** — can the hub see the store at all. Counted in consecutive failed fetch windows,
emitted past a threshold (default 3 ≈ 30–45 min) with a paired all-clear.
**Why reachability REPEATS rather than escalating once.** `emitFill`/`emitOversub` fire once on a band
rise and stay silent while the condition persists, which is right for a capacity trend. Applied to
reachability it would produce exactly ONE mail at roughly minute 30 of a nine-hour outage — and one
mail is missable, which is the whole failure being fixed. So the unreachable event fires on every
failed window past the threshold and relies on the dispatcher's 1-hour per-type operator cooldown to
become an hourly "still blind" heartbeat. **If you find yourself "fixing" this back to the band shape,
this paragraph is why not.** The cost is a `suppressed` notification row every ~15 min during an
outage — honest bookkeeping.
**Two boundaries worth holding in mind.** `ErrUsageUnsupported` is *not* blindness — an ep0 on an old
tenantsync answers "no such op", which means we reached it; the counter is not advanced, or the alert
would fire for days on a healthy pre-update box. And the reachability read rides ep0's **local API
daemon**, not the HTTPS proxy on 8007 — so it would have shown green throughout the 2026-08-18 outage,
which is R-340 and is the honest limit of what this watches.
## The managed floor now tracks the vouched golden — and that is rule 2 ARRIVING, not an exception (2026-08-18, R-343)
Since 2026-08-18 12:36:58Z the global managed controller floor
+10
View File
@@ -43,6 +43,16 @@ record with no machine** — created 13 August, no host, no backups, nothing to
## Shipped
- **The system now tells you when it cannot see the off-site copies** (R-339). Until today, a
completely dead off-site store and a perfectly healthy one looked **identical** to you — the checks
only ever watched how full a store was getting, and a failed reading was written to a log nobody
reads. That is why Monday's nine-and-a-half-hour outage reached you only by accident, through the
weekly backup that happened to fall inside it. After about half an hour of being unable to see a
store you now get a mail, repeated hourly while it lasts, and one all-clear when it comes back.
**Caveat worth knowing:** this watches whether the machine answers at all — it would *not* have
caught Monday's exact fault, which was one service wedged while the machine stayed healthy. That
second check is written down as the next step.
- **The auto-update floor is current again** (R-343). You raised it to today's version this
afternoon. It was **not** left behind by accident — our own rule says the floor is raised *last*,
after the image is vouched, because it acts within seconds. It did: **one machine updated itself
@@ -157,5 +157,6 @@
| Per-customer offsite fill + staleness + freeze lever | hub v0.41 | **IMPLEMENTED** | `OffsiteChecker` (`hub/internal/monitor/offsite.go`): fill 90/95% vs soft quota, staleness >48h | No live-fired leg: `CAMPAIGN-offsite-overnight-2026-07-10` recorded no quota/fill/staleness emails, and the freeze write-block was **inconclusive** (only the Hetzner `readonly:true` API op succeeded). Demoted |
| **Box-level Storage Box aggregate (total fill, Σ quotas, oversubscription alert)** | hub v0.64.0 (R-5) | **IMPLEMENTED** (data pipeline PROVEN-LIVE) | `monitor.OffsiteBoxChecker` — fetch-throttled Hetzner GET (1/15 min), fill (used/`storage_box_type.size`, 80/90%) + oversubscription (Σ shared+enabled ConfigJSON quotas / capacity, 2.0×), escalation-only operator alert on the customer-less `"pool-box"` scope; Offsite-tab panel + dashboard tile. **Phase-0-pinned** live shape (box 611714) + **live-computed in-cluster:** `0.2% full (2.6 GB of 1.00 TB), Σ shared quota 150 GB, oversub 0.15x`. Tests + 4 red-proofs; hub v0.64.0 REPORT | Two open legs: the **UI render** is unit-verified only (hub UI password-gated → no screenshot); the **alert emails** are unit + red-proof verified but NOT fired live (real pool nominal — a live-fire emails Viktor). Thresholds pending Viktor's ruling (named config keys). READ-ONLY (GET) |
| **Operator sees PBS DR datastore fill at a glance (Offsite "PBS DR" tab + dashboard gauge)** | hub v0.65.0 + tenantsync v1.2.0 (R-5) | **IMPLEMENTED** (data pipeline PROVEN-LIVE) | The PBS DR datastore (`felhom-offsite` on ep0) fill — NOT a Hetzner box. **Option A:** a read-only `usage` op on the `felhom-tenantsync` ep0 forced command (twin of `fingerprint`; `df` on the datastore path — no customer_id, no admin token, NO mutation), polled by `monitor.PBSDRBoxChecker` (OffsiteBoxChecker clone; 15-min throttle; states ok/unavailable/degraded; fill 80/90% on the `"pbsdr-box"` operator scope). `/offsite` split into Restic + PBS DR tabs; two dashboard gauges. **Graceful: hub deploy ⟂ ep0 update** (ep0 ≤ v1.1.0 → gauge "n/a" until updated). **Phase-0-pinned** (`df` on ep0 PBS 4.2.3) + **live-computed in-cluster** (ep0 updated to v1.2.0 this session): `19.1% full (7.1 GB of 37.2 GB)`. 10 Go tests + a bash harness + 3 red-proofs; hub v0.65.0 REPORT | Open legs: **UI render** unit-verified only (hub UI password-gated); the **fill alert email** is unit + red-proof verified, NOT fired live (datastore nominal at 19%). Separate PBS threshold keys (default 80/90); no oversubscription (namespaces, not quotas). READ-ONLY |
| **The hub reports LOSS OF VISIBILITY into either off-site store (not just how full it is)** | hub **v0.106.0** (R-339) | **IMPLEMENTED — deliberately NOT proven-live** | Both box checkers count consecutive failed fetch windows and emit `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default 3 windows (≈30–45 min), each with a paired `*_recovered` all-clear routed via `recoveredPairedDownTypes` — required because the recoveries are severity `info`, which `severityNotifies` drops. Scopes stay customer-less (`pbsdr-box` / `pool-box`) → operator channel only. Fill logic untouched: a degraded read still drives no band transition. Evidence: `internal/monitor/box_reachability_test.go` + the cross-package wiring test in `internal/notify/`, which asserts an actual operator mail rather than a map entry. **Filed BECAUSE of a measured gap**, not a hypothesis: the 2026-08-18 ep0 outage ran 9 h 37 m with the hub silent | **The gap that remains is R-340**, and it is not small: the ep0 read is the `usage` op, which rides the LOCAL API daemon — the daemon that incident explicitly cleared — so this check would have shown GREEN for that entire outage. It closes "ep0 is unreachable as a host"; it does not close what actually happened. **No live or constructed outage has exercised the emit path**, and one cannot be manufactured against ep0 (Tier 2, protected) |
| Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | |
| Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | |
File diff suppressed because one or more lines are too long
+51
View File
@@ -1,3 +1,54 @@
## v0.106.0 — the hub says something when it loses sight of the off-site stores (2026-08-18, R-339)
**The gap this closes, measured rather than supposed.** On 2026-08-18 ep0's PBS proxy was wedged for
**9 h 37 m** and the hub emitted nothing on the operator channel. Both box checkers hold their last
snapshot and return silently when a fetch fails — correct for a *fill* signal, since a missing reading
must never be mistaken for 0%, but it means a dead off-site endpoint and a healthy one look identical.
The only mails that morning came from the boxes' own backup failures, and only because the **weekly**
offsite run happened to land inside the window. Two days earlier nothing would have fired at all.
(`documentation/audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md`.)
**Reachability is now a second, independent signal.** Both checkers count consecutive failed fetch
windows and, past a threshold, report on the operator channel with a paired all-clear:
| | |
|---|---|
| `pbsdr_box_unreachable` / `pbsdr_box_recovered` | scope `pbsdr-box` — `internal/monitor/pbsdr_box.go` |
| `offsite_box_unreachable` / `offsite_box_recovered` | scope `pool-box` — `internal/monitor/offsite_box.go` |
| default threshold | **3 consecutive failed 15-minute windows** (≈30–45 min); `alerting.box_unreachable_windows`, 0/invalid → 3 — `cmd/hub/main.go` |
| recovery routing | `recoveredPairedDownTypes` — `internal/notify/dispatcher.go` |
**The fill logic is untouched.** No threshold, no throttle, no band, no escalate-once behaviour
changed. A degraded read still drives no transition and the page still shows the last known number,
still marked stale.
**Three decisions worth keeping, because each is the kind a later reader would "fix" back:**
- **The unreachable event REPEATS; it does not escalate once.** Copying `emitFill`'s band-rank shape
would give exactly one mail at ~minute 30 of a nine-hour outage, and one mail is missable. It fires
every failed window past the threshold and leans on the dispatcher's 1-hour operator cooldown to
become an hourly "still blind" heartbeat. Cost: a `suppressed` notification row every ~15 min during
an outage — honest bookkeeping.
- **`ErrUsageUnsupported` is NOT blindness.** An ep0 on tenantsync ≤ v1.1.0 answers "no such op" — we
reached it. The counter is deliberately not advanced, or the alert would fire for days on a healthy
pre-update endpoint and train the operator to ignore it.
- **Born-blind is reported.** The counter is not gated on having a snapshot, so a hub restarted *into*
an outage still speaks — the case the previous code handled worst. `last_ok` is **omitted**, never
zero-valued: a fabricated timestamp reads as "it was fine until then".
**Both recovery events are severity `info`, and `severityNotifies` drops `info`** — so without their
`recoveredPairedDownTypes` entries they would be stored and never mailed, and the operator would be
told the tier broke and never told it healed. A cross-package test drives `ProcessEvent` and asserts
an actual operator mail, because a green checker test proves nothing about the seam.
**Tests:** `internal/monitor/box_reachability_test.go` (Scenarios A–F) and
`internal/notify/dispatcher_box_reachability_test.go` (the wiring). Three companion red-proofs run and
reverted: threshold 3→1, the sentinel counter guard, and the pairing entry — each seen failing with a
message naming the right cause.
**NOT proven live.** No real or constructed outage has exercised the emit path end to end; a real one
cannot be manufactured without making ep0 or the Hetzner API unreachable, and ep0 is Tier 2 protected.
## v0.105.0 — the third name, a machine told to be quiet, and a guard for the hub's own words (2026-08-13)
Hub only. **No controller change, no agent change, no wire change — nothing to bake.** R-323, R-324,
+7
View File
@@ -78,6 +78,11 @@ type Config struct {
// No oversubscription concept for PBS (namespaces, not quotas) — fill only.
PBSDRBoxFillWarnPercent float64 `yaml:"pbsdr_box_fill_warn_percent"`
PBSDRBoxFillCritPercent float64 `yaml:"pbsdr_box_fill_crit_percent"`
// Consecutive failed 15-minute box reads before the hub reports the store unreachable
// (v0.106.0, R-339). 0/invalid → 3 (≈30–45 min). Applies to BOTH the restic pool box and the
// PBS-DR datastore — one knob, because the two are the same operator question ("can I still
// see the off-site tier") and splitting it would invite them to drift apart.
BoxUnreachableWindows int `yaml:"box_unreachable_windows"`
} `yaml:"alerting"`
Registry struct {
Image string `yaml:"image"`
@@ -344,6 +349,7 @@ func main() {
if poolBoxID != 0 {
offsiteBoxChecker = monitor.NewOffsiteBoxChecker(client, poolBoxID, dataStore,
cfg.Alerting.OffsiteBoxFillWarnPercent, cfg.Alerting.OffsiteBoxFillCritPercent, cfg.Alerting.OffsiteOversubWarnRatio,
cfg.Alerting.BoxUnreachableWindows,
dispatcher.ProcessEvent, logger)
webServer.SetOffsiteBox(offsiteBoxChecker.Snapshot)
} else {
@@ -473,6 +479,7 @@ func main() {
// (read-only usage op). Graceful vs an ep0 still on script ≤ v1.1.0 (unavailable state).
pbsdrBoxChecker = monitor.NewPBSDRBoxChecker(tsClient,
cfg.Alerting.PBSDRBoxFillWarnPercent, cfg.Alerting.PBSDRBoxFillCritPercent,
cfg.Alerting.BoxUnreachableWindows,
dispatcher.ProcessEvent, logger)
webServer.SetPBSDRBox(pbsdrBoxChecker.Snapshot)
}
@@ -0,0 +1,346 @@
package monitor
import (
"encoding/json"
"errors"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-hub/internal/hetznerapi"
"gitea.dooplex.hu/admin/felhom-hub/internal/tenantsync"
)
// R-339 — box REACHABILITY. These tests cover the signal that was MISSING on 2026-08-18, when ep0's
// PBS proxy was wedged for 9 h 37 m and the hub said nothing on the operator channel.
//
// Every assertion here checks a VERDICT — the event type, its severity, and the specific details
// fields — not merely "an event was emitted" or "the count is N". That is deliberate and recent:
// this afternoon's due_checks_gate red-proof passed for the wrong reason because its only assertion
// was an exit code that a crash produced just as well as the logic under test. A bare count is the
// same trap in a different costume.
// capturedFull records ALL SIX EventNotifyFunc arguments, unlike capturedBox which keeps three.
// Reachability assertions need the details payload (consecutive_failures, last_ok, blind_for), and a
// recorder that discards it cannot tell a correct event from a plausible one.
type capturedFull struct {
cust, typ, sev, msg, details, src []string
}
func (c *capturedFull) fn(cust, et, sev, msg, details, src string) {
c.cust = append(c.cust, cust)
c.typ = append(c.typ, et)
c.sev = append(c.sev, sev)
c.msg = append(c.msg, msg)
c.details = append(c.details, details)
c.src = append(c.src, src)
}
func (c *capturedFull) n() int { return len(c.typ) }
// ofType returns the indices of events of a given type — so a test can assert "exactly these" rather
// than "at least one somewhere".
func (c *capturedFull) ofType(t string) []int {
var out []int
for i, et := range c.typ {
if et == t {
out = append(out, i)
}
}
return out
}
func (c *capturedFull) detail(i int) map[string]any {
var m map[string]any
_ = json.Unmarshal([]byte(c.details[i]), &m)
return m
}
// ── Group A — Scenario A: the endpoint goes dark and STAYS dark ──────────────────────────────
func TestPBSDRBox_Unreachable_SustainedOutage(t *testing.T) {
f := &fakeUsage{usage: usageBytes(40, 8)}
ev := &capturedFull{}
var cur time.Time
c := NewPBSDRBoxChecker(f, 80, 90, 0, ev.fn, quietLog()) // 0 → default threshold 3
c.now = func() time.Time { return cur }
base := time.Now().UTC()
cur = base
c.Check() // one healthy read
if ev.n() != 0 {
t.Fatalf("healthy read emitted %d event(s), want 0: %v", ev.n(), ev.typ)
}
bandBefore := c.FillState()
f.set(tenantsync.BoxUsage{}, errors.New("dial tcp 10.77.0.1:8007: i/o timeout"))
for i := 1; i <= 6; i++ {
cur = base.Add(time.Duration(i) * boxFetchInterval)
c.Check()
}
idx := ev.ofType("pbsdr_box_unreachable")
if len(idx) != 4 {
t.Fatalf("want 4 unreachable events across 6 failed windows (silent on 1 and 2), got %d (all: %v)", len(idx), ev.typ)
}
for _, i := range idx {
if ev.sev[i] != "warning" {
t.Fatalf("event %d severity = %q, want \"warning\"", i, ev.sev[i])
}
if ev.cust[i] != pbsdrBoxScope {
t.Fatalf("event %d scope = %q, want %q", i, ev.cust[i], pbsdrBoxScope)
}
if ev.src[i] != "hub" {
t.Fatalf("event %d source = %q, want \"hub\"", i, ev.src[i])
}
}
// The counter is the escalation signal — first alert at 3, last at 6.
if got := ev.detail(idx[0])["consecutive_failures"]; got != float64(3) {
t.Fatalf("first alert consecutive_failures = %v, want 3", got)
}
if got := ev.detail(idx[len(idx)-1])["consecutive_failures"]; got != float64(6) {
t.Fatalf("last alert consecutive_failures = %v, want 6", got)
}
if _, ok := ev.detail(idx[0])["last_ok"]; !ok {
t.Fatalf("first alert must carry last_ok — there WAS a successful read before the outage")
}
// The fill signal must be untouched throughout: missing data never moves a band.
if len(ev.ofType("pbsdr_box_fill")) != 0 {
t.Fatalf("an outage emitted a FILL event — missing data must never drive a band transition")
}
if c.FillState() != bandBefore {
t.Fatalf("fill band moved during an outage: %q → %q", bandBefore, c.FillState())
}
}
// ── Group B — Scenario B: a transient blip below threshold ───────────────────────────────────
func TestPBSDRBox_Unreachable_BlipBelowThreshold(t *testing.T) {
f := &fakeUsage{usage: usageBytes(40, 8)}
ev := &capturedFull{}
var cur time.Time
c := NewPBSDRBoxChecker(f, 80, 90, 0, ev.fn, quietLog())
c.now = func() time.Time { return cur }
base := time.Now().UTC()
cur = base
c.Check()
f.set(tenantsync.BoxUsage{}, errors.New("transient"))
cur = base.Add(boxFetchInterval)
c.Check()
cur = base.Add(2 * boxFetchInterval)
c.Check()
if ev.n() != 0 {
t.Fatalf("two failed windows emitted %v, want silence below the threshold", ev.typ)
}
f.set(usageBytes(40, 8), nil)
cur = base.Add(3 * boxFetchInterval)
c.Check()
if ev.n() != 0 {
t.Fatalf("recovery below threshold emitted %v — there was no alert to clear", ev.typ)
}
// The counter reset is proven by the NEGATIVE: one more failure must not reach the threshold.
// Reading the private field would prove the field, not the behaviour.
f.set(tenantsync.BoxUsage{}, errors.New("transient again"))
cur = base.Add(4 * boxFetchInterval)
c.Check()
if ev.n() != 0 {
t.Fatalf("a single failure after a success emitted %v — the counter did not reset", ev.typ)
}
}
// ── Group C — Scenario C: recovery after an alert ────────────────────────────────────────────
func TestPBSDRBox_Unreachable_Recovery(t *testing.T) {
f := &fakeUsage{usage: usageBytes(40, 8)}
ev := &capturedFull{}
var cur time.Time
c := NewPBSDRBoxChecker(f, 80, 90, 0, ev.fn, quietLog())
c.now = func() time.Time { return cur }
base := time.Now().UTC()
cur = base
c.Check()
f.set(tenantsync.BoxUsage{}, errors.New("down"))
for i := 1; i <= 3; i++ {
cur = base.Add(time.Duration(i) * boxFetchInterval)
c.Check()
}
if len(ev.ofType("pbsdr_box_unreachable")) == 0 {
t.Fatalf("setup failed: no unreachable event to recover from")
}
f.set(usageBytes(40, 8), nil)
cur = base.Add(4 * boxFetchInterval)
c.Check()
rec := ev.ofType("pbsdr_box_recovered")
if len(rec) != 1 {
t.Fatalf("want exactly 1 recovery event, got %d (all: %v)", len(rec), ev.typ)
}
if ev.sev[rec[0]] != "info" {
t.Fatalf("recovery severity = %q, want \"info\"", ev.sev[rec[0]])
}
if ev.cust[rec[0]] != pbsdrBoxScope {
t.Fatalf("recovery scope = %q, want %q", ev.cust[rec[0]], pbsdrBoxScope)
}
bf, _ := ev.detail(rec[0])["blind_for"].(string)
if bf == "" || bf == "unknown" {
t.Fatalf("recovery blind_for = %q, want a real duration", bf)
}
before := ev.n()
cur = base.Add(5 * boxFetchInterval)
c.Check()
if ev.n() != before {
t.Fatalf("a second consecutive success emitted %v — the alerted flag did not clear", ev.typ[before:])
}
}
// ── Group D — Scenario D: ErrUsageUnsupported is NOT blindness ───────────────────────────────
func TestPBSDRBox_UsageUnsupported_IsNotBlindness(t *testing.T) {
f := &fakeUsage{err: tenantsync.ErrUsageUnsupported}
ev := &capturedFull{}
var cur time.Time
c := NewPBSDRBoxChecker(f, 80, 90, 0, ev.fn, quietLog())
c.now = func() time.Time { return cur }
base := time.Now().UTC()
for i := 0; i < 10; i++ {
cur = base.Add(time.Duration(i) * boxFetchInterval)
c.Check()
}
if ev.n() != 0 {
t.Fatalf("ErrUsageUnsupported emitted %v — an expected pre-update condition must never alert", ev.typ)
}
snap, ok := c.Snapshot()
if !ok || snap.State != PBSStateUnavailable {
t.Fatalf("state = %q (have=%v), want %q", snap.State, ok, PBSStateUnavailable)
}
// The counter must never have advanced: one REAL error now must still be below the threshold.
f.set(tenantsync.BoxUsage{}, errors.New("real transport failure"))
cur = base.Add(10 * boxFetchInterval)
c.Check()
if ev.n() != 0 {
t.Fatalf("a single real error after 10 unsupported windows emitted %v — the unsupported "+
"branch advanced the counter, which would alert on a healthy pre-update endpoint", ev.typ)
}
}
// ── Group E — Scenario E: born blind (hub restarted INTO an outage) ──────────────────────────
func TestPBSDRBox_BornBlind_StillReports(t *testing.T) {
f := &fakeUsage{err: errors.New("connection refused")}
ev := &capturedFull{}
var cur time.Time
c := NewPBSDRBoxChecker(f, 80, 90, 0, ev.fn, quietLog()) // no snapshot, ever
c.now = func() time.Time { return cur }
base := time.Now().UTC()
for i := 0; i < 3; i++ {
cur = base.Add(time.Duration(i) * boxFetchInterval)
c.Check()
}
idx := ev.ofType("pbsdr_box_unreachable")
if len(idx) != 1 {
t.Fatalf("born-blind checker emitted %d unreachable events, want exactly 1 (all: %v)", len(idx), ev.typ)
}
d := ev.detail(idx[0])
if v, present := d["last_ok"]; present {
t.Fatalf("last_ok = %v, but there has NEVER been a successful read — a fabricated timestamp "+
"reads as \"it was fine until then\", which is a lie", v)
}
if got := d["consecutive_failures"]; got != float64(3) {
t.Fatalf("consecutive_failures = %v, want 3", got)
}
if ev.sev[idx[0]] != "warning" {
t.Fatalf("severity = %q, want \"warning\"", ev.sev[idx[0]])
}
}
// ── Group F — Scenario F: the storage box, same rules ────────────────────────────────────────
func TestOffsiteBox_Unreachable_AndRecovery(t *testing.T) {
st := newDiskStore(t)
fake := hetznerapi.NewFake()
fake.Boxes[1] = boxWith(1000*gib, 100*gib) // 10% — comfortably ok
ev := &capturedFull{}
var cur time.Time
c := NewOffsiteBoxChecker(fake, 1, st, 80, 90, 2.0, 0, ev.fn, quietLog())
c.now = func() time.Time { return cur }
base := time.Now().UTC()
cur = base
c.Check()
fillBefore, oversubBefore := c.FillState(), c.OversubState()
fake.FailGetBox = errors.New("hetzner api: 503 service unavailable")
for i := 1; i <= 3; i++ {
cur = base.Add(time.Duration(i) * boxFetchInterval)
c.Check()
}
idx := ev.ofType("offsite_box_unreachable")
if len(idx) != 1 {
t.Fatalf("want exactly 1 unreachable at the third window, got %d (all: %v)", len(idx), ev.typ)
}
if ev.sev[idx[0]] != "warning" || ev.cust[idx[0]] != offsiteBoxScope {
t.Fatalf("event severity/scope = %q/%q, want warning/%s", ev.sev[idx[0]], ev.cust[idx[0]], offsiteBoxScope)
}
if got := ev.detail(idx[0])["consecutive_failures"]; got != float64(3) {
t.Fatalf("consecutive_failures = %v, want 3", got)
}
fake.FailGetBox = nil
cur = base.Add(4 * boxFetchInterval)
c.Check()
rec := ev.ofType("offsite_box_recovered")
if len(rec) != 1 {
t.Fatalf("want exactly 1 recovery, got %d (all: %v)", len(rec), ev.typ)
}
if ev.sev[rec[0]] != "info" {
t.Fatalf("recovery severity = %q, want \"info\"", ev.sev[rec[0]])
}
// Neither capacity band may move across the whole outage.
if len(ev.ofType("offsite_box_fill")) != 0 || len(ev.ofType("offsite_box_oversub")) != 0 {
t.Fatalf("an outage emitted a capacity event: %v", ev.typ)
}
if c.FillState() != fillBefore || c.OversubState() != oversubBefore {
t.Fatalf("bands moved during an outage: fill %q→%q oversub %q→%q",
fillBefore, c.FillState(), oversubBefore, c.OversubState())
}
}
// A zero-capacity read is a SUCCESS for reachability: we reached the endpoint and it answered. The
// answer being unusable for a percentage is a different complaint, and conflating the two would
// leave a blindness alert outstanding forever on a box that is actually responding.
func TestPBSDRBox_ZeroCapacitySuccess_ClearsBlindness(t *testing.T) {
f := &fakeUsage{usage: usageBytes(40, 8)}
ev := &capturedFull{}
var cur time.Time
c := NewPBSDRBoxChecker(f, 80, 90, 0, ev.fn, quietLog())
c.now = func() time.Time { return cur }
base := time.Now().UTC()
cur = base
c.Check()
f.set(tenantsync.BoxUsage{}, errors.New("down"))
for i := 1; i <= 3; i++ {
cur = base.Add(time.Duration(i) * boxFetchInterval)
c.Check()
}
if len(ev.ofType("pbsdr_box_unreachable")) == 0 {
t.Fatal("setup failed: expected an unreachable alert")
}
f.set(tenantsync.BoxUsage{Total: 0, Used: 0}, nil) // reached, answered, unusable payload
cur = base.Add(4 * boxFetchInterval)
c.Check()
if len(ev.ofType("pbsdr_box_recovered")) != 1 {
t.Fatalf("a zero-capacity SUCCESS did not clear the blindness alert (events: %v)", ev.typ)
}
snap, _ := c.Snapshot()
if snap.State != PBSStateDegraded {
t.Fatalf("snapshot state = %q, want %q — the fill signal is still degraded, correctly",
snap.State, PBSStateDegraded)
}
}
+96 -6
View File
@@ -40,12 +40,24 @@ type OffsiteBoxChecker struct {
logger *log.Logger
now func() time.Time // injectable clock (tests)
reachFailThreshold int // consecutive failed fetch windows before reporting unreachable
mu sync.Mutex
snap BoxSnapshot
haveSnap bool
lastFetch time.Time
fillBand string // "" until first eval → born/persistent (a born-breach emits on the first sweep)
oversubBand string
// REACHABILITY (v0.106.0, R-339) — the sibling of PBSDRBoxChecker's, same rules over a different
// source. NOTE the one structural difference, stated so a reader does not go looking for a branch
// that was forgotten: there is NO ErrUsageUnsupported equivalent here. The Hetzner API has no
// version-gated op, so every non-nil error from GetStorageBox is genuine blindness and there is no
// "expected data gap" case to exclude.
reachFails int
reachAlerted bool
reachFirstFail time.Time
lastOK time.Time // zero = never (born blind — do NOT fabricate one)
}
// offsiteBoxScope is the customer-less cooldown/display key for pool-box events.
@@ -81,7 +93,7 @@ type BoxSnapshot struct {
// NewOffsiteBoxChecker builds the checker. Invalid thresholds fall back to the documented defaults
// (warn 80% / crit 90% / oversub 2.0×). Nothing is seeded → the first Check that finds a breach emits
// (born/persistent; the dispatcher's 1 h cooldown dedups a hub restart).
func NewOffsiteBoxChecker(api hetznerapi.CloudAPI, boxID int64, s *store.Store, fillWarn, fillCrit, oversubWarn float64, onEvent EventNotifyFunc, logger *log.Logger) *OffsiteBoxChecker {
func NewOffsiteBoxChecker(api hetznerapi.CloudAPI, boxID int64, s *store.Store, fillWarn, fillCrit, oversubWarn float64, reachWindows int, onEvent EventNotifyFunc, logger *log.Logger) *OffsiteBoxChecker {
if fillWarn <= 0 || fillWarn >= 100 {
fillWarn = defaultOffsiteBoxFillWarnPercent
}
@@ -94,13 +106,19 @@ func NewOffsiteBoxChecker(api hetznerapi.CloudAPI, boxID int64, s *store.Store,
if oversubWarn <= 0 {
oversubWarn = defaultOffsiteOversubWarnRatio
}
if reachWindows <= 0 {
reachWindows = defaultBoxUnreachableWindows
}
c := &OffsiteBoxChecker{
api: api, boxID: boxID, store: s,
fillWarn: fillWarn, fillCrit: fillCrit, oversubWarn: oversubWarn,
onEvent: onEvent, logger: logger, now: time.Now,
reachFailThreshold: reachWindows,
onEvent: onEvent, logger: logger, now: time.Now,
}
logger.Printf("[INFO] Offsite pool-box checker initialized: box=%d fill warn=%.0f%% crit=%.0f%%, oversub warn=%.2fx, refresh %s",
boxID, fillWarn, fillCrit, oversubWarn, boxFetchInterval)
// Reachability threshold in the init line for the same reason as the PBS-DR checker: its absence
// from the pod log is how you notice the parameter never arrived.
logger.Printf("[INFO] Offsite pool-box checker initialized: box=%d fill warn=%.0f%% crit=%.0f%%, oversub warn=%.2fx, unreachable after %d consecutive failed reads, refresh %s",
boxID, fillWarn, fillCrit, oversubWarn, reachWindows, boxFetchInterval)
return c
}
@@ -122,9 +140,28 @@ func (c *OffsiteBoxChecker) Check() {
if c.haveSnap {
c.snap.Degraded = true // serve the last-known, visibly stale; NO band transition (D)
}
// REACHABILITY — not gated on haveSnap (a hub restarted into an outage must still report).
if c.reachFails == 0 {
c.reachFirstFail = c.now()
}
c.reachFails++
if c.reachFails >= c.reachFailThreshold {
c.emitUnreachable(err)
}
return
}
// REACHABILITY: the API answered — a success for this signal even if the payload turns out to be
// unusable for a fill percentage (zero capacity below). Handled here, before the degraded early
// return, so a zero-capacity read still clears an outstanding blindness alert.
if c.reachAlerted {
c.emitRecovered()
c.reachAlerted = false
}
c.reachFails = 0
c.reachFirstFail = time.Time{}
c.lastOK = c.now()
sumQuotaGB, qerr := c.sumSharedQuotaGB()
if qerr != nil {
c.logger.Printf("[WARN] Offsite pool-box: quota sum failed (ratio omitted this cycle): %v", qerr)
@@ -182,8 +219,16 @@ func (c *OffsiteBoxChecker) Snapshot() (BoxSnapshot, bool) {
}
// FillState / OversubState expose the current bands (tests).
func (c *OffsiteBoxChecker) FillState() string { c.mu.Lock(); defer c.mu.Unlock(); return orUnknown(c.fillBand) }
func (c *OffsiteBoxChecker) OversubState() string { c.mu.Lock(); defer c.mu.Unlock(); return orUnknown(c.oversubBand) }
func (c *OffsiteBoxChecker) FillState() string {
c.mu.Lock()
defer c.mu.Unlock()
return orUnknown(c.fillBand)
}
func (c *OffsiteBoxChecker) OversubState() string {
c.mu.Lock()
defer c.mu.Unlock()
return orUnknown(c.oversubBand)
}
func orUnknown(b string) string {
if b == "" {
@@ -212,6 +257,51 @@ func (c *OffsiteBoxChecker) sumSharedQuotaGB() (int64, error) {
return sum, nil
}
// emitUnreachable reports that the hub cannot see the shared storage box.
//
// ⚠ CADENCE — THIS REPEATS BY DESIGN; see the twin in pbsdr_box.go for the full reasoning. In short:
// emitFill/emitOversub are escalation-only, which applied here would yield exactly one mail at ~30
// minutes into a multi-hour outage. This fires on every failed window past the threshold and relies on
// the dispatcher's 1-hour operator cooldown to become an hourly "still blind" heartbeat.
func (c *OffsiteBoxChecker) emitUnreachable(cause error) {
c.reachAlerted = true
lastOK := "never"
if !c.lastOK.IsZero() {
lastOK = c.lastOK.Format(time.RFC3339)
}
d := map[string]any{
"scope": offsiteBoxScope, "consecutive_failures": c.reachFails, "error": cause.Error(),
}
// Omitted, never zero-valued, when there has never been a successful read — see Scenario E.
if !c.lastOK.IsZero() {
d["last_ok"] = c.lastOK.Format(time.RFC3339)
}
details, _ := json.Marshal(d)
msg := fmt.Sprintf("Offsite pool box unreachable — %d consecutive 15-minute checks failed (last successful read: %s); the shared storage box cannot be seen from the hub",
c.reachFails, lastOK)
c.logger.Printf("[WARN] Offsite pool-box UNREACHABLE: %d consecutive failed reads (last ok: %s)", c.reachFails, lastOK)
if c.onEvent != nil {
c.onEvent(offsiteBoxScope, "offsite_box_unreachable", "warning", msg, string(details), "hub")
}
}
// emitRecovered is the paired all-clear (severity "info"; reaches the operator only via the
// dispatcher's recoveredPairedDownTypes entry, which short-circuits the severity gate).
func (c *OffsiteBoxChecker) emitRecovered() {
blindFor := "unknown"
if !c.reachFirstFail.IsZero() {
blindFor = c.now().Sub(c.reachFirstFail).Round(time.Second).String()
}
details, _ := json.Marshal(map[string]any{
"scope": offsiteBoxScope, "blind_for": blindFor, "consecutive_failures": c.reachFails,
})
msg := fmt.Sprintf("Offsite pool box reachable again after %s — the shared storage box is visible to the hub", blindFor)
c.logger.Printf("[INFO] Offsite pool-box RECOVERED after %s (%d failed reads)", blindFor, c.reachFails)
if c.onEvent != nil {
c.onEvent(offsiteBoxScope, "offsite_box_recovered", "info", msg, string(details), "hub")
}
}
func (c *OffsiteBoxChecker) emitFill(snap BoxSnapshot, band string) {
var severity, message string
switch band {
+6 -6
View File
@@ -39,7 +39,7 @@ func TestOffsiteBox_FetchThrottle(t *testing.T) {
fake := hetznerapi.NewFake()
fake.Boxes[1] = boxWith(1<<40, 300*gib) // 1 TiB, 300 GiB used
var cur time.Time
c := NewOffsiteBoxChecker(fake, 1, st, 80, 90, 2.0, noEvent, quietLog())
c := NewOffsiteBoxChecker(fake, 1, st, 80, 90, 2.0, 0, noEvent, quietLog())
c.now = func() time.Time { return cur }
base := time.Now().UTC()
@@ -68,7 +68,7 @@ func TestOffsiteBox_OversubMath(t *testing.T) {
fake := hetznerapi.NewFake()
fake.Boxes[1] = boxWith(1024*gib, 100*gib) // capacity 1024 GB
c := NewOffsiteBoxChecker(fake, 1, st, 80, 90, 2.0, noEvent, quietLog())
c := NewOffsiteBoxChecker(fake, 1, st, 80, 90, 2.0, 0, noEvent, quietLog())
c.Check()
snap, _ := c.Snapshot()
@@ -99,7 +99,7 @@ func TestOffsiteBox_FillBands(t *testing.T) {
fake.Boxes[1] = boxWith(cap, 750*gib) // 75%
ev := &capturedBox{}
var cur time.Time
c := NewOffsiteBoxChecker(fake, 1, st, 80, 90, 2.0, ev.fn, quietLog())
c := NewOffsiteBoxChecker(fake, 1, st, 80, 90, 2.0, 0, ev.fn, quietLog())
c.now = func() time.Time { return cur }
base := time.Now().UTC()
step := func(usedGiB int64) {
@@ -151,7 +151,7 @@ func TestOffsiteBox_OversubIndependent(t *testing.T) {
fake := hetznerapi.NewFake()
fake.Boxes[1] = boxWith(1000*gib, 500*gib) // 50% fill (nominal), ratio 2.3×
ev := &capturedBox{}
c := NewOffsiteBoxChecker(fake, 1, st, 80, 90, 2.0, ev.fn, quietLog())
c := NewOffsiteBoxChecker(fake, 1, st, 80, 90, 2.0, 0, ev.fn, quietLog())
c.Check()
if count(ev.typ, "offsite_box_oversub") != 1 {
t.Fatalf("2.3× oversub must warn once, got %v", ev.typ)
@@ -174,7 +174,7 @@ func TestOffsiteBox_FailedFetchHonesty(t *testing.T) {
fake.Boxes[1] = boxWith(cap, 920*gib) // 92% → critical
ev := &capturedBox{}
var cur time.Time
c := NewOffsiteBoxChecker(fake, 1, st, 80, 90, 2.0, ev.fn, quietLog())
c := NewOffsiteBoxChecker(fake, 1, st, 80, 90, 2.0, 0, ev.fn, quietLog())
c.now = func() time.Time { return cur }
base := time.Now().UTC()
cur = base
@@ -216,7 +216,7 @@ func TestOffsiteBox_ZeroCapacityGuard(t *testing.T) {
fake := hetznerapi.NewFake()
fake.Boxes[1] = boxWith(0, 0) // initializing box — no capacity yet
ev := &capturedBox{}
c := NewOffsiteBoxChecker(fake, 1, st, 80, 90, 2.0, ev.fn, quietLog())
c := NewOffsiteBoxChecker(fake, 1, st, 80, 90, 2.0, 0, ev.fn, quietLog())
c.Check()
snap, ok := c.Snapshot()
if !ok || !snap.Degraded {
+105 -4
View File
@@ -25,7 +25,7 @@ type usageReader interface {
// failure honesty. THREE snapshot states:
// - "ok" → bands drive alerts.
// - "unavailable" → the endpoint script predates the usage op (≤ v1.1.0) → ErrUsageUnsupported. An
// EXPECTED pre-update condition: neutral, NO alert, logged once. The gauge shows n/a.
// EXPECTED pre-update condition: neutral, NO alert, logged once. The gauge shows n/a.
// - "degraded" → exec failed/timed out → keep the last snapshot, NO band transition (missing ≠ 0%).
//
// Events carry the customer-less scope "pbsdr-box" → operator channel ONLY (processCustomer no-ops on it);
@@ -38,12 +38,23 @@ type PBSDRBoxChecker struct {
logger *log.Logger
now func() time.Time // injectable clock (tests)
reachFailThreshold int // consecutive failed fetch windows before reporting unreachable
mu sync.Mutex
snap PBSBoxSnapshot
haveSnap bool
lastFetch time.Time
fillBand string // "" until first ok eval → born/persistent; only an "ok" fetch updates it
unavailLogged bool // log the "unavailable" state ONCE, not per sweep
// REACHABILITY (v0.106.0, R-339) — a SECOND, independent signal, deliberately not part of the
// fill state above. Fill answers "how full is the store"; this answers "can the hub see the store
// at all". They are separate because a missing reading must never move a fill band (missing ≠ 0%),
// which is exactly what made a 9 h 37 m ep0 outage silent on 2026-08-18.
reachFails int // consecutive failed fetch windows (NOT sweeps — lastFetch throttles both)
reachAlerted bool // an unreachable event has been emitted for the CURRENT outage
reachFirstFail time.Time // start of the current outage → the honest blind_for duration
lastOK time.Time // last successful read; zero = never (born blind — do NOT fabricate one)
}
// pbsdrBoxScope is the customer-less cooldown/display key for PBS-DR box events.
@@ -59,6 +70,11 @@ const (
const (
defaultPBSDRBoxFillWarnPercent = 80.0
defaultPBSDRBoxFillCritPercent = 90.0
// defaultBoxUnreachableWindows is the consecutive failed 15-minute fetch windows before the hub
// reports a box unreachable. Three windows is ≈30–45 min depending on where the first failure
// lands relative to the last success — long enough that a single blip or a brief API wobble says
// nothing, short enough that a real outage is reported inside the hour.
defaultBoxUnreachableWindows = 3
)
// PBSBoxSnapshot is the cached PBS-DR datastore aggregate the web layer renders. Sizes in BYTES.
@@ -73,7 +89,7 @@ type PBSBoxSnapshot struct {
// NewPBSDRBoxChecker builds the checker (defaults 80/90 on invalid thresholds). Nothing seeded → the
// first ok Check that finds a breach emits (born/persistent; the dispatcher's 1 h cooldown dedups a restart).
func NewPBSDRBoxChecker(reader usageReader, fillWarn, fillCrit float64, onEvent EventNotifyFunc, logger *log.Logger) *PBSDRBoxChecker {
func NewPBSDRBoxChecker(reader usageReader, fillWarn, fillCrit float64, reachWindows int, onEvent EventNotifyFunc, logger *log.Logger) *PBSDRBoxChecker {
if fillWarn <= 0 || fillWarn >= 100 {
fillWarn = defaultPBSDRBoxFillWarnPercent
}
@@ -83,11 +99,19 @@ func NewPBSDRBoxChecker(reader usageReader, fillWarn, fillCrit float64, onEvent
if fillCrit <= fillWarn {
fillWarn, fillCrit = defaultPBSDRBoxFillWarnPercent, defaultPBSDRBoxFillCritPercent
}
if reachWindows <= 0 {
reachWindows = defaultBoxUnreachableWindows
}
c := &PBSDRBoxChecker{
reader: reader, fillWarn: fillWarn, fillCrit: fillCrit,
onEvent: onEvent, logger: logger, now: time.Now,
reachFailThreshold: reachWindows,
onEvent: onEvent, logger: logger, now: time.Now,
}
logger.Printf("[INFO] PBS-DR box checker initialized: fill warn=%.0f%% crit=%.0f%%, refresh %s", fillWarn, fillCrit, boxFetchInterval)
// The reachability threshold is in this line deliberately: a running hub can be asked what it is
// configured to do without anyone reading the config file, and its ABSENCE from the pod log is how
// you would notice the parameter never reached the checker.
logger.Printf("[INFO] PBS-DR box checker initialized: fill warn=%.0f%% crit=%.0f%%, unreachable after %d consecutive failed reads, refresh %s",
fillWarn, fillCrit, reachWindows, boxFetchInterval)
return c
}
@@ -111,16 +135,42 @@ func (c *PBSDRBoxChecker) Check() {
c.logger.Printf("[INFO] PBS-DR box: usage op unavailable — endpoint tenantsync update (v1.2.0) pending; gauge shows n/a until then")
c.unavailLogged = true
}
// REACHABILITY, Scenario D: the counter is deliberately NOT advanced here. This is an
// EXPECTED pre-update condition (endpoint tenantsync ≤ v1.1.0), not blindness — we reached
// the endpoint and it told us the op does not exist. Alerting on it would fire for days on
// a healthy box and train the operator to ignore the alert, which costs more than it buys.
return // fillBand untouched — an expected data gap never re-arms/transitions
}
c.logger.Printf("[WARN] PBS-DR box: usage read failed (keeping last snapshot): %v", err)
if c.haveSnap {
c.snap.State = PBSStateDegraded // serve last-known, visibly stale; NO band transition
}
// REACHABILITY: count the failed window and report once past the threshold. NOT gated on
// haveSnap — a hub restarted INTO an outage has no snapshot to mark degraded and is precisely
// the case the pre-v0.106.0 code handled worst: it would have stayed silent forever.
if c.reachFails == 0 {
c.reachFirstFail = c.now()
}
c.reachFails++
if c.reachFails >= c.reachFailThreshold {
c.emitUnreachable(err)
}
return
}
c.unavailLogged = false // recovered from unavailable
// REACHABILITY: we reached the endpoint and it answered — that is a success for this signal even
// if the answer is unusable for a fill percentage (zero capacity below). Handled HERE, before the
// `snap.State != PBSStateOK` early return further down, or a zero-capacity read would leave an
// outstanding blindness alert un-cleared forever.
if c.reachAlerted {
c.emitRecovered()
c.reachAlerted = false
}
c.reachFails = 0
c.reachFirstFail = time.Time{}
c.lastOK = c.now()
snap := PBSBoxSnapshot{CapacityBytes: usage.Total, UsedBytes: usage.Used, State: PBSStateOK, FetchedAt: c.now()}
if usage.Total > 0 {
snap.FillPercent = float64(usage.Used) * 100 / float64(usage.Total)
@@ -158,6 +208,57 @@ func (c *PBSDRBoxChecker) FillState() string {
return orUnknown(c.fillBand)
}
// emitUnreachable reports that the hub cannot see the PBS-DR datastore.
//
// ⚠ CADENCE — THIS REPEATS, AND THAT IS DELIBERATE. Do not "fix" it back to the band shape.
// emitFill/emitOversub are ESCALATION-ONLY: they fire once on a band rise and stay silent while the
// condition persists, which is right for a fill trend. Applied here it would produce exactly ONE mail
// at roughly minute 30 of a nine-hour outage, and one mail is missable — that is the failure this
// whole signal exists to remove. So this fires on EVERY failed window past the threshold and leans on
// the dispatcher's 1-hour per-type operator cooldown to throttle the mail down to an hourly "still
// blind" heartbeat, which is itself the escalation. The cost is a `suppressed` notification row about
// every 15 minutes during an outage — honest bookkeeping, and cheaper than a missed outage.
func (c *PBSDRBoxChecker) emitUnreachable(cause error) {
c.reachAlerted = true
lastOK := "never"
if !c.lastOK.IsZero() {
lastOK = c.lastOK.Format(time.RFC3339)
}
d := map[string]any{
"scope": pbsdrBoxScope, "consecutive_failures": c.reachFails, "error": cause.Error(),
}
// last_ok is OMITTED rather than zero-valued when there has never been a successful read — a
// fabricated timestamp reads as "it was fine until then", which would be a lie (Scenario E).
if !c.lastOK.IsZero() {
d["last_ok"] = c.lastOK.Format(time.RFC3339)
}
details, _ := json.Marshal(d)
msg := fmt.Sprintf("PBS DR endpoint unreachable — %d consecutive 15-minute checks failed (last successful read: %s); the off-site DR tier cannot be seen from the hub",
c.reachFails, lastOK)
c.logger.Printf("[WARN] PBS-DR box UNREACHABLE: %d consecutive failed reads (last ok: %s)", c.reachFails, lastOK)
if c.onEvent != nil {
c.onEvent(pbsdrBoxScope, "pbsdr_box_unreachable", "warning", msg, string(details), "hub")
}
}
// emitRecovered is the paired all-clear. Severity "info" — it reaches the operator only because
// `pbsdr_box_recovered` is registered in the dispatcher's recoveredPairedDownTypes, which short-
// circuits the severity gate. Without that entry this is stored and never mailed.
func (c *PBSDRBoxChecker) emitRecovered() {
blindFor := "unknown"
if !c.reachFirstFail.IsZero() {
blindFor = c.now().Sub(c.reachFirstFail).Round(time.Second).String()
}
details, _ := json.Marshal(map[string]any{
"scope": pbsdrBoxScope, "blind_for": blindFor, "consecutive_failures": c.reachFails,
})
msg := fmt.Sprintf("PBS DR endpoint reachable again after %s — the off-site DR tier is visible to the hub", blindFor)
c.logger.Printf("[INFO] PBS-DR box RECOVERED after %s (%d failed reads)", blindFor, c.reachFails)
if c.onEvent != nil {
c.onEvent(pbsdrBoxScope, "pbsdr_box_recovered", "info", msg, string(details), "hub")
}
}
func (c *PBSDRBoxChecker) emitFill(snap PBSBoxSnapshot, band string) {
var severity, message string
switch band {
+4 -4
View File
@@ -39,7 +39,7 @@ func usageBytes(totalGiB, usedGiB int64) tenantsync.BoxUsage {
func TestPBSDRBox_Throttle(t *testing.T) {
f := &fakeUsage{usage: usageBytes(40, 8)}
var cur time.Time
c := NewPBSDRBoxChecker(f, 80, 90, noEvent, quietLog())
c := NewPBSDRBoxChecker(f, 80, 90, 0, noEvent, quietLog())
c.now = func() time.Time { return cur }
base := time.Now().UTC()
for i := 0; i < 60; i++ {
@@ -61,7 +61,7 @@ func TestPBSDRBox_FillBands(t *testing.T) {
f := &fakeUsage{usage: usageBytes(1000, 750)} // 75%
ev := &capturedBox{}
var cur time.Time
c := NewPBSDRBoxChecker(f, 80, 90, ev.fn, quietLog())
c := NewPBSDRBoxChecker(f, 80, 90, 0, ev.fn, quietLog())
c.now = func() time.Time { return cur }
base := time.Now().UTC()
cur = base
@@ -110,7 +110,7 @@ func TestPBSDRBox_FillBands(t *testing.T) {
func TestPBSDRBox_Unavailable(t *testing.T) {
f := &fakeUsage{err: tenantsync.ErrUsageUnsupported}
ev := &capturedBox{}
c := NewPBSDRBoxChecker(f, 80, 90, ev.fn, quietLog())
c := NewPBSDRBoxChecker(f, 80, 90, 0, ev.fn, quietLog())
c.Check()
snap, ok := c.Snapshot()
if !ok || snap.State != PBSStateUnavailable {
@@ -129,7 +129,7 @@ func TestPBSDRBox_DegradedKeepsLast(t *testing.T) {
f := &fakeUsage{usage: usageBytes(1000, 920)} // 92% critical
ev := &capturedBox{}
var cur time.Time
c := NewPBSDRBoxChecker(f, 80, 90, ev.fn, quietLog())
c := NewPBSDRBoxChecker(f, 80, 90, 0, ev.fn, quietLog())
c.now = func() time.Time { return cur }
cur = time.Now().UTC()
+10
View File
@@ -81,6 +81,16 @@ var recoveredPairedDownTypes = map[string][]string{
// nobody's enabled_events — so no such row can exist, and the customer correctly hears neither
// edge. Operator hears both.
"whole_guest_backup_recovered": {"whole_guest_backup_failed"},
// R-339 (v0.106.0) — box REACHABILITY all-clears. Same reasoning as the entry above and the same
// necessity: both are severity "info", so without an entry here severityNotifies drops them and
// the operator is told the off-site tier went blind and never told it came back.
//
// The customer leg is a no-op BY CONSTRUCTION, not by luck: both scopes are customer-less
// ("pbsdr-box" / "pool-box" are not customer IDs), so GetNotificationPrefs finds no row and
// LastCustomerSentAt can never locate a paired down row for them. Operator hears both edges;
// no customer hears either, which is correct — neither store belongs to a customer.
"pbsdr_box_recovered": {"pbsdr_box_unreachable"},
"offsite_box_recovered": {"offsite_box_unreachable"},
}
// severityNotifies reports whether a severity triggers email notifications. warning / error / critical
@@ -0,0 +1,100 @@
package notify
import (
"io"
"log"
"testing"
)
// R-339, Group G — THE WIRING TEST, and it is not optional.
//
// The recovery leg spans two packages: internal/monitor emits, internal/notify routes. A fake-injected
// test in monitor proves the checker emits and proves NOTHING about whether the operator receives the
// mail. Both `*_box_recovered` events carry severity "info", and severityNotifies drops "info" — so
// without an entry in recoveredPairedDownTypes they are stored and never mailed, and the operator is
// told the off-site tier went blind and never told it came back.
//
// Precedent for why this test exists at all: agent v0.91.0 shipped fully green with SetAuthSink never
// called from main.go, and the entire auth-honesty leg was inert. A green unit test on one side of a
// seam is not evidence that the seam is connected.
//
// These drive ProcessEvent end-to-end and assert a MAIL, not a map entry. A map assertion would pass
// on a correctly-populated map that nothing reads.
func TestBoxRecovery_ReachesTheOperatorDespiteInfoSeverity(t *testing.T) {
if severityNotifies("info") {
t.Fatal("premise changed: \"info\" now notifies, so the pairing entries may be unnecessary — re-check")
}
for _, et := range []string{"pbsdr_box_recovered", "offsite_box_recovered"} {
st := newDispStore(t)
d := NewDispatcher(st, "test-key", "hub@felhom.eu", "op@felhom.eu", true, log.New(io.Discard, "", 0))
mails := captureSeam(d)
scope := "pbsdr-box"
if et == "offsite_box_recovered" {
scope = "pool-box"
}
d.ProcessEvent(scope, et, "info", "reachable again after 45m0s", `{"scope":"`+scope+`"}`, "hub")
op := mailsFor(*mails, "op@felhom.eu")
if len(op) != 1 {
t.Fatalf("%s: operator mails = %d, want 1 — the all-clear must reach the operator; "+
"0 means the recoveredPairedDownTypes entry is missing and \"info\" was dropped by the severity gate",
et, len(op))
}
// Customer leg is a no-op BY CONSTRUCTION: the scope is not a customer id, so no prefs row
// exists and no paired customer "sent" row can be found.
if len(*mails) != 1 {
t.Fatalf("%s: total mails = %d, want 1 — a customer-less scope must never produce a customer mail",
et, len(*mails))
}
}
}
// The DOWN edge needs no pairing entry — "warning" already notifies — but if that ever changed the
// operator would hear the all-clear for an outage they were never told about. Pin both edges.
func TestBoxUnreachable_ReachesTheOperatorOnItsOwnSeverity(t *testing.T) {
for _, c := range []struct{ scope, et string }{
{"pbsdr-box", "pbsdr_box_unreachable"},
{"pool-box", "offsite_box_unreachable"},
} {
st := newDispStore(t)
d := NewDispatcher(st, "test-key", "hub@felhom.eu", "op@felhom.eu", true, log.New(io.Discard, "", 0))
mails := captureSeam(d)
d.ProcessEvent(c.scope, c.et, "warning", "unreachable — 3 consecutive checks failed", `{"scope":"`+c.scope+`"}`, "hub")
op := mailsFor(*mails, "op@felhom.eu")
if len(op) != 1 {
t.Fatalf("%s: operator mails = %d, want 1", c.et, len(op))
}
if len(*mails) != 1 {
t.Fatalf("%s: total mails = %d, want 1 — operator-only", c.et, len(*mails))
}
}
}
// Both recovery types must be registered against the RIGHT down type. A recovery paired with the
// wrong down event would still mail the operator (processOperator runs unconditionally) while
// silently breaking the customer pairing rule for any future customer-scoped reuse.
func TestBoxRecovery_PairedWithTheCorrectDownType(t *testing.T) {
want := map[string]string{
"pbsdr_box_recovered": "pbsdr_box_unreachable",
"offsite_box_recovered": "offsite_box_unreachable",
}
for rec, down := range want {
paired, ok := recoveredPairedDownTypes[rec]
if !ok {
t.Fatalf("%s is not on the recovery branch — its \"info\" severity makes it silent", rec)
}
found := false
for _, p := range paired {
if p == down {
found = true
}
}
if !found {
t.Fatalf("%s must pair with %s; got %v", rec, down, paired)
}
}
}