Files
felhom.eu/hub/internal/monitor/offsite_box.go
T
admin ab2262c91c
gates / gates (push) Successful in 14s
hub v0.106.0: report loss of visibility into the off-site stores (R-339)
THE GAP, measured not supposed. On 2026-08-18 ep0's PBS proxy was wedged for
9 h 37 m and the hub emitted NOTHING on the operator channel. Both box
checkers hold their last snapshot and return silently on a failed fetch --
correct for a FILL signal, since a missing reading must never be read as 0%,
but it makes a dead off-site endpoint and a healthy one indistinguishable.
The only mails that morning came from the boxes' own backup failures, and
only because the WEEKLY offsite run happened to land inside the window. Two
days earlier nothing would have fired at all.

REACHABILITY is now a second, independent signal on both checkers:
consecutive failed fetch windows, reported past a default 3 windows
(~30-45 min) as pbsdr_box_unreachable / offsite_box_unreachable (warning) on
the customer-less pbsdr-box / pool-box scopes, each with a paired *_recovered
all-clear. Tunable via alerting.box_unreachable_windows (0/invalid -> 3).

THE FILL LOGIC IS UNTOUCHED. No threshold, throttle, band or escalate-once
behaviour changed; a degraded read still drives no transition.

Three decisions a later reader would otherwise "fix" back, so each is
argued in-code:
  - the unreachable event REPEATS rather than escalating once. The band shape
    would give exactly ONE mail at ~minute 30 of a nine-hour outage, and one
    mail is missable. It leans on the dispatcher's 1 h operator cooldown to
    become an hourly "still blind" heartbeat.
  - ErrUsageUnsupported is NOT blindness: an old ep0 answers "no such op",
    which means we reached it. Counting it would alert for days on a healthy
    pre-update endpoint.
  - born-blind is reported: the counter is not gated on having a snapshot, so
    a hub restarted INTO an outage still speaks. last_ok is OMITTED rather
    than zero-valued -- a fabricated timestamp reads as "it was fine until
    then".

Both recoveries are severity "info" and severityNotifies drops "info", so
they are registered in recoveredPairedDownTypes or the operator hears that
the tier broke and never that it healed. A cross-package test drives
ProcessEvent and asserts an actual operator MAIL, not a map entry -- a green
checker test proves nothing about the seam (agent v0.91.0 shipped fully green
with SetAuthSink never called).

Tests: box_reachability_test.go (Scenarios A-F) + dispatcher_box_reachability
_test.go (wiring). Three red-proofs run and reverted, each seen failing with a
message naming the right cause: threshold 3->1, the sentinel counter guard,
the pairing entry.

Register: R-339 filed and marked SHIPPED (PROVEN-LIVE still owed -- no real or
constructed outage has exercised the emit path, and one cannot be manufactured
against Tier-2 ep0). R-340 filed: the reachability read rides ep0's LOCAL API
daemon, which the incident explicitly cleared, so this check would have shown
GREEN for all 9 h 37 m -- the honest boundary, recorded rather than glossed.
R-336's next-step corrected: pvestatd's interval is NOT tunable (Proxmox staff
have said so); the only lever is disabling the storage entry, which collides
with the agent's consume-the-one-time-secret path. Doc-only, no agent code
touched.
2026-08-18 19:27:33 +02:00

351 lines
14 KiB
Go
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
package monitor
import (
"context"
"encoding/json"
"fmt"
"log"
"sync"
"time"
"gitea.dooplex.hu/admin/felhom-hub/internal/hetznerapi"
"gitea.dooplex.hu/admin/felhom-hub/internal/offsite"
"gitea.dooplex.hu/admin/felhom-hub/internal/store"
)
// OffsiteBoxChecker (v0.64.0, R-5) watches the SHARED POOL BOX as a whole — the aggregate the
// per-customer OffsiteChecker cannot see. Two independent box-level signals, one checker, mirroring the
// OffsiteChecker/StorageFillChecker shape (escalation-only emit, recovery re-arms silently, born/persistent
// via an unseeded band). It is DELIBERATELY separate from OffsiteChecker — a different data source (the
// Hetzner unified API, not the controller report) and a customer-less scope.
//
// - FILL: box used vs box capacity (from the box TYPE's size, never stats.size) at warn 80% / crit 90%.
// - OVERSUBSCRIPTION: Σ(shared+enabled customers' soft quotas) vs capacity, at a warn ratio (2.0×). The
// number that says how far the sold promises exceed the disk. Independent of fill (both can fire).
//
// Fetch is throttled (boxFetchInterval); between refreshes Check() serves the cached snapshot. A failed
// fetch keeps the last snapshot (marked Degraded) and never drives a band transition — absence of data is
// NEVER treated as 0%. Events carry the customer-less scope "pool-box": the dispatcher uses it as a
// cooldown key + display string only, and processCustomer no-ops on it (no prefs → no customer email;
// verified against dispatcher.go), so a box event reaches ONLY the operator channel. NO SaveEvent — the
// scope has no customer row to key it to.
type OffsiteBoxChecker struct {
api hetznerapi.CloudAPI
boxID int64
store *store.Store
fillWarn float64
fillCrit float64
oversubWarn float64
onEvent EventNotifyFunc
logger *log.Logger
now func() time.Time // injectable clock (tests)
reachFailThreshold int // consecutive failed fetch windows before reporting unreachable
mu sync.Mutex
snap BoxSnapshot
haveSnap bool
lastFetch time.Time
fillBand string // "" until first eval → born/persistent (a born-breach emits on the first sweep)
oversubBand string
// REACHABILITY (v0.106.0, R-339) — the sibling of PBSDRBoxChecker's, same rules over a different
// source. NOTE the one structural difference, stated so a reader does not go looking for a branch
// that was forgotten: there is NO ErrUsageUnsupported equivalent here. The Hetzner API has no
// version-gated op, so every non-nil error from GetStorageBox is genuine blindness and there is no
// "expected data gap" case to exclude.
reachFails int
reachAlerted bool
reachFirstFail time.Time
lastOK time.Time // zero = never (born blind — do NOT fabricate one)
}
// offsiteBoxScope is the customer-less cooldown/display key for pool-box events.
const offsiteBoxScope = "pool-box"
const (
defaultOffsiteBoxFillWarnPercent = 80.0
defaultOffsiteBoxFillCritPercent = 90.0
defaultOffsiteOversubWarnRatio = 2.0
// boxFetchInterval throttles the Hetzner read: Check() runs on the 60 s sweep but fetches at most
// once per this window (≈4 API reads/hour, not ≈60). A manual page load never forces a fetch.
boxFetchInterval = 15 * time.Minute
)
// BoxSnapshot is the cached pool-box aggregate the web layer renders. All sizes are BYTES. FillBand /
// OversubBand carry the checker's own band verdict so the UI colors bars EXACTLY as the alerts fire (no
// threshold re-derivation in the web layer).
type BoxSnapshot struct {
CapacityBytes int64
UsedBytes int64
DataBytes int64
SnapshotBytes int64
FillPercent float64
SumQuotaGB int64
Ratio float64 // Σ(shared quota bytes) / capacity
FillBand string // "" | ok | warning | critical
OversubBand string // "" | ok | warning
FetchedAt time.Time
BoxType string
Degraded bool // last fetch failed OR capacity==0 → render stale; no alert transitions
}
// NewOffsiteBoxChecker builds the checker. Invalid thresholds fall back to the documented defaults
// (warn 80% / crit 90% / oversub 2.0×). Nothing is seeded → the first Check that finds a breach emits
// (born/persistent; the dispatcher's 1 h cooldown dedups a hub restart).
func NewOffsiteBoxChecker(api hetznerapi.CloudAPI, boxID int64, s *store.Store, fillWarn, fillCrit, oversubWarn float64, reachWindows int, onEvent EventNotifyFunc, logger *log.Logger) *OffsiteBoxChecker {
if fillWarn <= 0 || fillWarn >= 100 {
fillWarn = defaultOffsiteBoxFillWarnPercent
}
if fillCrit <= 0 || fillCrit > 100 || fillCrit <= fillWarn {
fillCrit = defaultOffsiteBoxFillCritPercent
}
if fillCrit <= fillWarn {
fillWarn, fillCrit = defaultOffsiteBoxFillWarnPercent, defaultOffsiteBoxFillCritPercent
}
if oversubWarn <= 0 {
oversubWarn = defaultOffsiteOversubWarnRatio
}
if reachWindows <= 0 {
reachWindows = defaultBoxUnreachableWindows
}
c := &OffsiteBoxChecker{
api: api, boxID: boxID, store: s,
fillWarn: fillWarn, fillCrit: fillCrit, oversubWarn: oversubWarn,
reachFailThreshold: reachWindows,
onEvent: onEvent, logger: logger, now: time.Now,
}
// Reachability threshold in the init line for the same reason as the PBS-DR checker: its absence
// from the pod log is how you notice the parameter never arrived.
logger.Printf("[INFO] Offsite pool-box checker initialized: box=%d fill warn=%.0f%% crit=%.0f%%, oversub warn=%.2fx, unreachable after %d consecutive failed reads, refresh %s",
boxID, fillWarn, fillCrit, oversubWarn, reachWindows, boxFetchInterval)
return c
}
// Check runs on the 60 s sweep. It fetches (throttled), recomputes the snapshot, and emits on each
// band escalation. Recovery de-escalation re-arms silently. Degraded data drives no transition.
func (c *OffsiteBoxChecker) Check() {
c.mu.Lock()
defer c.mu.Unlock()
// Fetch throttle: between refreshes serve the cache — bands only move on a fresh fetch.
if c.haveSnap && c.now().Sub(c.lastFetch) < boxFetchInterval {
return
}
c.lastFetch = c.now() // updated even on failure → a failing API is retried once per window, not per sweep
box, err := c.api.GetStorageBox(context.Background(), c.boxID)
if err != nil {
c.logger.Printf("[WARN] Offsite pool-box: refresh failed (keeping last snapshot): %v", err)
if c.haveSnap {
c.snap.Degraded = true // serve the last-known, visibly stale; NO band transition (D)
}
// REACHABILITY — not gated on haveSnap (a hub restarted into an outage must still report).
if c.reachFails == 0 {
c.reachFirstFail = c.now()
}
c.reachFails++
if c.reachFails >= c.reachFailThreshold {
c.emitUnreachable(err)
}
return
}
// REACHABILITY: the API answered — a success for this signal even if the payload turns out to be
// unusable for a fill percentage (zero capacity below). Handled here, before the degraded early
// return, so a zero-capacity read still clears an outstanding blindness alert.
if c.reachAlerted {
c.emitRecovered()
c.reachAlerted = false
}
c.reachFails = 0
c.reachFirstFail = time.Time{}
c.lastOK = c.now()
sumQuotaGB, qerr := c.sumSharedQuotaGB()
if qerr != nil {
c.logger.Printf("[WARN] Offsite pool-box: quota sum failed (ratio omitted this cycle): %v", qerr)
sumQuotaGB = 0 // a wrong sum is worse than a nominal 0.00× — never emit oversub off bad data
}
capacity := box.StorageBoxType.Size
used := box.Stats.Size
snap := BoxSnapshot{
CapacityBytes: capacity, UsedBytes: used,
DataBytes: box.Stats.SizeData, SnapshotBytes: box.Stats.SizeSnapshots,
SumQuotaGB: sumQuotaGB, FetchedAt: c.now(), BoxType: box.StorageBoxType.Name,
}
if capacity > 0 {
snap.FillPercent = float64(used) * 100 / float64(capacity)
snap.Ratio = float64(sumQuotaGB<<30) / float64(capacity)
snap.FillBand = bandForPercent(snap.FillPercent, c.fillWarn, c.fillCrit)
snap.OversubBand = bandOK
if snap.Ratio >= c.oversubWarn {
snap.OversubBand = bandWarning // ratio is single-threshold — no "critical"
}
} else {
snap.Degraded = true // zero capacity (initializing box / probe surprise) — guard the division
}
c.snap = snap
c.haveSnap = true
if snap.Degraded {
c.logger.Printf("[INFO] Offsite pool-box refreshed: DEGRADED (capacity unavailable) — box %d", c.boxID)
return // no band transitions on degraded data
}
// Periodic operator visibility of the pool trend (one line per 15-min refresh; keys, no secrets).
c.logger.Printf("[INFO] Offsite pool-box refreshed: %.1f%% full (%s of %s), Σ shared quota %d GB, oversub %.2fx",
snap.FillPercent, fmtSize(snap.UsedBytes), fmtSize(snap.CapacityBytes), snap.SumQuotaGB, snap.Ratio)
// FILL band (escalation-only; recovery re-arms because bandOK has rank 0).
if bandRank(snap.FillBand) > bandRank(c.fillBand) {
c.emitFill(snap, snap.FillBand)
}
c.fillBand = snap.FillBand
// OVERSUBSCRIPTION band — independent signal (both can fire in one sweep; neither masks the other).
if bandRank(snap.OversubBand) > bandRank(c.oversubBand) {
c.emitOversub(snap)
}
c.oversubBand = snap.OversubBand
}
// Snapshot returns the current cached aggregate + whether one exists (mutex copy-out). The web layer's
// only source — it NEVER fetches from Hetzner in the request path.
func (c *OffsiteBoxChecker) Snapshot() (BoxSnapshot, bool) {
c.mu.Lock()
defer c.mu.Unlock()
return c.snap, c.haveSnap
}
// FillState / OversubState expose the current bands (tests).
func (c *OffsiteBoxChecker) FillState() string {
c.mu.Lock()
defer c.mu.Unlock()
return orUnknown(c.fillBand)
}
func (c *OffsiteBoxChecker) OversubState() string {
c.mu.Lock()
defer c.mu.Unlock()
return orUnknown(c.oversubBand)
}
func orUnknown(b string) string {
if b == "" {
return "unknown"
}
return b
}
// sumSharedQuotaGB sums the soft quotas of every ENABLED, SHARED customer from the authoritative
// ConfigJSON descriptor (never the report echo). Dedicated + disabled customers are excluded.
func (c *OffsiteBoxChecker) sumSharedQuotaGB() (int64, error) {
cfgs, err := c.store.ListCustomerConfigs()
if err != nil {
return 0, err
}
var sum int64
for _, cfg := range cfgs {
d, derr := offsite.ReadDescriptor(cfg.ConfigJSON)
if derr != nil || d == nil {
continue
}
if d.Enabled && d.Type == "shared" && d.QuotaGB > 0 {
sum += int64(d.QuotaGB)
}
}
return sum, nil
}
// emitUnreachable reports that the hub cannot see the shared storage box.
//
// ⚠ CADENCE — THIS REPEATS BY DESIGN; see the twin in pbsdr_box.go for the full reasoning. In short:
// emitFill/emitOversub are escalation-only, which applied here would yield exactly one mail at ~30
// minutes into a multi-hour outage. This fires on every failed window past the threshold and relies on
// the dispatcher's 1-hour operator cooldown to become an hourly "still blind" heartbeat.
func (c *OffsiteBoxChecker) emitUnreachable(cause error) {
c.reachAlerted = true
lastOK := "never"
if !c.lastOK.IsZero() {
lastOK = c.lastOK.Format(time.RFC3339)
}
d := map[string]any{
"scope": offsiteBoxScope, "consecutive_failures": c.reachFails, "error": cause.Error(),
}
// Omitted, never zero-valued, when there has never been a successful read — see Scenario E.
if !c.lastOK.IsZero() {
d["last_ok"] = c.lastOK.Format(time.RFC3339)
}
details, _ := json.Marshal(d)
msg := fmt.Sprintf("Offsite pool box unreachable — %d consecutive 15-minute checks failed (last successful read: %s); the shared storage box cannot be seen from the hub",
c.reachFails, lastOK)
c.logger.Printf("[WARN] Offsite pool-box UNREACHABLE: %d consecutive failed reads (last ok: %s)", c.reachFails, lastOK)
if c.onEvent != nil {
c.onEvent(offsiteBoxScope, "offsite_box_unreachable", "warning", msg, string(details), "hub")
}
}
// emitRecovered is the paired all-clear (severity "info"; reaches the operator only via the
// dispatcher's recoveredPairedDownTypes entry, which short-circuits the severity gate).
func (c *OffsiteBoxChecker) emitRecovered() {
blindFor := "unknown"
if !c.reachFirstFail.IsZero() {
blindFor = c.now().Sub(c.reachFirstFail).Round(time.Second).String()
}
details, _ := json.Marshal(map[string]any{
"scope": offsiteBoxScope, "blind_for": blindFor, "consecutive_failures": c.reachFails,
})
msg := fmt.Sprintf("Offsite pool box reachable again after %s — the shared storage box is visible to the hub", blindFor)
c.logger.Printf("[INFO] Offsite pool-box RECOVERED after %s (%d failed reads)", blindFor, c.reachFails)
if c.onEvent != nil {
c.onEvent(offsiteBoxScope, "offsite_box_recovered", "info", msg, string(details), "hub")
}
}
func (c *OffsiteBoxChecker) emitFill(snap BoxSnapshot, band string) {
var severity, message string
switch band {
case bandCritical:
severity = "critical"
message = fmt.Sprintf("Offsite pool box %.0f%% full (%s of %s) — approaching capacity; add capacity or reduce retention before customer offsite runs start failing",
snap.FillPercent, fmtSize(snap.UsedBytes), fmtSize(snap.CapacityBytes))
case bandWarning:
severity = "warning"
message = fmt.Sprintf("Offsite pool box %.0f%% full (%s of %s) — the shared pool is filling",
snap.FillPercent, fmtSize(snap.UsedBytes), fmtSize(snap.CapacityBytes))
default:
return
}
details, _ := json.Marshal(map[string]any{
"scope": offsiteBoxScope, "capacity_bytes": snap.CapacityBytes, "used_bytes": snap.UsedBytes,
"fill_percent": snap.FillPercent, "warn_percent": c.fillWarn, "crit_percent": c.fillCrit,
})
c.logger.Printf("[INFO] Offsite pool-box fill: %.0f%% (%s)", snap.FillPercent, band)
// NO SaveEvent — the scope is customer-less; emit straight to the operator dispatcher.
if c.onEvent != nil {
c.onEvent(offsiteBoxScope, "offsite_box_fill", severity, message, string(details), "hub")
}
}
func (c *OffsiteBoxChecker) emitOversub(snap BoxSnapshot) {
message := fmt.Sprintf("Offsite pool oversubscription %.2fx — Σ(shared soft quotas) %d GB against %s capacity exceeds the %.2fx threshold; a full-usage scenario would overrun the pool",
snap.Ratio, snap.SumQuotaGB, fmtSize(snap.CapacityBytes), c.oversubWarn)
details, _ := json.Marshal(map[string]any{
"scope": offsiteBoxScope, "sum_quota_gb": snap.SumQuotaGB, "capacity_bytes": snap.CapacityBytes,
"ratio": snap.Ratio, "oversub_warn_ratio": c.oversubWarn,
})
c.logger.Printf("[INFO] Offsite pool-box oversub: %.2fx (Σquota %d GB)", snap.Ratio, snap.SumQuotaGB)
if c.onEvent != nil {
c.onEvent(offsiteBoxScope, "offsite_box_oversub", "warning", message, string(details), "hub")
}
}
// fmtSize renders bytes as GB (one decimal) below 1 TB, TB above — for the operator email/log only.
func fmtSize(b int64) string {
const gb = 1 << 30
if b >= 1<<40 {
return fmt.Sprintf("%.2f TB", float64(b)/float64(int64(1)<<40))
}
return fmt.Sprintf("%.1f GB", float64(b)/float64(gb))
}