hub v0.106.0: report loss of visibility into the off-site stores (R-339)
gates / gates (push) Successful in 14s

THE GAP, measured not supposed. On 2026-08-18 ep0's PBS proxy was wedged for
9 h 37 m and the hub emitted NOTHING on the operator channel. Both box
checkers hold their last snapshot and return silently on a failed fetch --
correct for a FILL signal, since a missing reading must never be read as 0%,
but it makes a dead off-site endpoint and a healthy one indistinguishable.
The only mails that morning came from the boxes' own backup failures, and
only because the WEEKLY offsite run happened to land inside the window. Two
days earlier nothing would have fired at all.

REACHABILITY is now a second, independent signal on both checkers:
consecutive failed fetch windows, reported past a default 3 windows
(~30-45 min) as pbsdr_box_unreachable / offsite_box_unreachable (warning) on
the customer-less pbsdr-box / pool-box scopes, each with a paired *_recovered
all-clear. Tunable via alerting.box_unreachable_windows (0/invalid -> 3).

THE FILL LOGIC IS UNTOUCHED. No threshold, throttle, band or escalate-once
behaviour changed; a degraded read still drives no transition.

Three decisions a later reader would otherwise "fix" back, so each is
argued in-code:
  - the unreachable event REPEATS rather than escalating once. The band shape
    would give exactly ONE mail at ~minute 30 of a nine-hour outage, and one
    mail is missable. It leans on the dispatcher's 1 h operator cooldown to
    become an hourly "still blind" heartbeat.
  - ErrUsageUnsupported is NOT blindness: an old ep0 answers "no such op",
    which means we reached it. Counting it would alert for days on a healthy
    pre-update endpoint.
  - born-blind is reported: the counter is not gated on having a snapshot, so
    a hub restarted INTO an outage still speaks. last_ok is OMITTED rather
    than zero-valued -- a fabricated timestamp reads as "it was fine until
    then".

Both recoveries are severity "info" and severityNotifies drops "info", so
they are registered in recoveredPairedDownTypes or the operator hears that
the tier broke and never that it healed. A cross-package test drives
ProcessEvent and asserts an actual operator MAIL, not a map entry -- a green
checker test proves nothing about the seam (agent v0.91.0 shipped fully green
with SetAuthSink never called).

Tests: box_reachability_test.go (Scenarios A-F) + dispatcher_box_reachability
_test.go (wiring). Three red-proofs run and reverted, each seen failing with a
message naming the right cause: threshold 3->1, the sentinel counter guard,
the pairing entry.

Register: R-339 filed and marked SHIPPED (PROVEN-LIVE still owed -- no real or
constructed outage has exercised the emit path, and one cannot be manufactured
against Tier-2 ep0). R-340 filed: the reachability read rides ep0's LOCAL API
daemon, which the incident explicitly cleared, so this check would have shown
GREEN for all 9 h 37 m -- the honest boundary, recorded rather than glossed.
R-336's next-step corrected: pvestatd's interval is NOT tunable (Proxmox staff
have said so); the only lever is disabling the storage entry, which collides
with the agent's consume-the-one-time-secret path. Doc-only, no agent code
touched.
This commit is contained in:
2026-08-18 19:27:33 +02:00
parent 78a244bf09
commit ab2262c91c
13 changed files with 765 additions and 21 deletions
+25
View File
@@ -15,6 +15,31 @@
> would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`**
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
## Box REACHABILITY is a separate signal from box FILL — and the cadence difference is deliberate (2026-08-18, R-339, hub v0.106.0)
Both off-site checkers now carry two independent signals, and conflating them is the mistake to avoid:
- **FILL** — how full is the store. Escalation-only (one emit on a band rise), and **a degraded or
missing reading drives NO band transition**, because missing ≠ 0%. That rule is unchanged by
v0.106.0 and must stay unchanged: it is what stops a dead endpoint from reading as an empty one.
- **REACHABILITY** — can the hub see the store at all. Counted in consecutive failed fetch windows,
emitted past a threshold (default 3 ≈ 30–45 min) with a paired all-clear.
**Why reachability REPEATS rather than escalating once.** `emitFill`/`emitOversub` fire once on a band
rise and stay silent while the condition persists, which is right for a capacity trend. Applied to
reachability it would produce exactly ONE mail at roughly minute 30 of a nine-hour outage — and one
mail is missable, which is the whole failure being fixed. So the unreachable event fires on every
failed window past the threshold and relies on the dispatcher's 1-hour per-type operator cooldown to
become an hourly "still blind" heartbeat. **If you find yourself "fixing" this back to the band shape,
this paragraph is why not.** The cost is a `suppressed` notification row every ~15 min during an
outage — honest bookkeeping.
**Two boundaries worth holding in mind.** `ErrUsageUnsupported` is *not* blindness — an ep0 on an old
tenantsync answers "no such op", which means we reached it; the counter is not advanced, or the alert
would fire for days on a healthy pre-update box. And the reachability read rides ep0's **local API
daemon**, not the HTTPS proxy on 8007 — so it would have shown green throughout the 2026-08-18 outage,
which is R-340 and is the honest limit of what this watches.
## The managed floor now tracks the vouched golden — and that is rule 2 ARRIVING, not an exception (2026-08-18, R-343)
Since 2026-08-18 12:36:58Z the global managed controller floor