Files
felhom.eu/hub/internal/web/rollup.go
T
admin 37ae31fd44
gates / gates (push) Successful in 21s
hub v0.117.0: the status follows the configured threshold; slow crash loop and interrupted restore events
R-549 (operator ruling A): controllerStatus hardcoded 30m/1h while both
staleness checkers and hostStatus read alerting.stale_threshold. Moving the
threshold to 45m would have painted a customer amber 15 minutes before the
alarm could fire - the second definition rollup.go's header forbids. It now
reads the same value, down at 2x. Both 'checker initialized' log lines print
the threshold, which no line did before.

R-539 (ruling 3 of 2026-09-16): controller_slow_crashloop (warning,
operator-only), minted when the agent's slow_crashloop_since moves, with the
fast sibling's first-sight rule.

R-550: restore_interrupted (warning, for the household) allowlisted with a
Hungarian customer message.

Red-proofs, each seen failing then passing: the status test with the old
hardcoded numbers; the checker test with the movement branch removed; the
operator-only test with the registration removed; the household-message test
with the Hungarian entry removed (asserted on the SUBJECT - the body
legitimately repeats the raw message, which my first version of the test
mistook for a fallback).

go build/vet/test ./... green, 18 packages.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 10:20:32 +02:00

87 lines
3.9 KiB
Go

package web
// Dead-host roll-up honesty (v0.53.0, drill-1 observation; operator ruling 2026-07-13): a
// customer's status may never look better than its worst expected host. The customer roll-up
// derives from CONTROLLER reports, which reach the hub independently of the host agent — so a
// host DOWN for 23 hours hid behind a green customer row as long as the guest kept reporting
// (the live Peti-cluster shape: proxmox1 down 23h, fresh reports through proxmox2).
//
// foldHostStatus worsens the controller-derived status with per-host staleness via
// (*Server).hostStatus — THE single staleness definition (hosts.go; the same thresholds the
// HostStalenessChecker alerts on — no second definition anywhere). Display + derivation only:
// checker alerting is untouched.
import (
"time"
"gitea.dooplex.hu/admin/felhom-hub/internal/store"
)
// controllerStatus is the controller-report-derived customer status — the pre-roll-up chain the
// dashboard, the /configs list and the customer detail all inlined verbatim; this is now the
// ONE copy. Behavior-preserving: the branch order (incl. fail-after-warn) is the historical one.
func controllerStatus(c *store.CustomerSummary, threshold time.Duration) string {
// R-549 (operator ruling A, 2026-09-17): the threshold is CONFIGURATION (alerting.stale_threshold)
// and the checkers read it; this display must read the same value, "down" at 2x like them. It
// hardcoded 30 m / 1 h until hub v0.117.0 — a second definition the header above forbids, and one
// that would have painted a customer amber 15 minutes before the alarm could fire.
// Pinned by TestControllerStatus_FollowsConfiguredThreshold.
if threshold <= 0 {
threshold = 30 * time.Minute
}
switch {
case c.HealthStatus == "disabled":
return "disabled"
case c.TimeSinceReport > 2*threshold:
return "down"
case c.TimeSinceReport > threshold || c.HealthStatus == "warn":
return "warn"
case c.HealthStatus == "fail":
return "down"
default:
return "ok"
}
}
// hostFoldRank orders host states by badness for the worst-host pick. "ok" ranks 0 (never folds).
var hostFoldRank = map[string]int{"down": 3, "stale": 2, "pending": 1}
// hostFoldLabel is the operator-facing cause chip prefix per worst-host state.
var hostFoldLabel = map[string]string{"down": "host down", "stale": "host stale", "pending": "host pending"}
// foldHostStatus folds the customer's expected hosts into a controller-derived status:
// worst(controllerDerived, hostStatusOf(each host)). Any host down/stale caps the customer at
// WARN (a green row over a dead host is the masking bug); the returned cause names the state
// AND the host ("host down: <id>") so the detail header says WHICH host. "pending" hosts
// (enrolled, never reported) worsen only after initial onboarding — customerHasReported=false
// (the customer has never reported) excludes them, a half-installed box is not an incident.
// Statuses worse than warn (down) and administrative ones (disabled/blocked) keep their own
// token; the cause chip still surfaces the host signal. Read errors degrade to the unfolded
// status — the page must render.
func (s *Server) foldHostStatus(customerID, base string, customerHasReported bool) (status, cause string) {
hosts, err := s.store.ListHostsByCustomer(customerID)
if err != nil {
s.logger.Printf("[ERROR] roll-up: ListHostsByCustomer %s: %v", customerID, err)
return base, ""
}
worst, worstHost := "", ""
for i := range hosts {
hs := s.hostStatus(hosts[i].LastReportAt)
if hs == "pending" && !customerHasReported {
continue
}
if hostFoldRank[hs] > hostFoldRank[worst] {
worst, worstHost = hs, hosts[i].HostID
}
}
if worst == "" {
return base, ""
}
cause = hostFoldLabel[worst] + ": " + worstHost
// The fold worsens, never improves: ok / pending / no-report ("") cap at warn.
if base == "ok" || base == "pending" || base == "" {
return "warn", cause
}
return base, cause
}