37ae31fd44
gates / gates (push) Successful in 21s
R-549 (operator ruling A): controllerStatus hardcoded 30m/1h while both staleness checkers and hostStatus read alerting.stale_threshold. Moving the threshold to 45m would have painted a customer amber 15 minutes before the alarm could fire - the second definition rollup.go's header forbids. It now reads the same value, down at 2x. Both 'checker initialized' log lines print the threshold, which no line did before. R-539 (ruling 3 of 2026-09-16): controller_slow_crashloop (warning, operator-only), minted when the agent's slow_crashloop_since moves, with the fast sibling's first-sight rule. R-550: restore_interrupted (warning, for the household) allowlisted with a Hungarian customer message. Red-proofs, each seen failing then passing: the status test with the old hardcoded numbers; the checker test with the movement branch removed; the operator-only test with the registration removed; the household-message test with the Hungarian entry removed (asserted on the SUBJECT - the body legitimately repeats the raw message, which my first version of the test mistook for a fallback). go build/vet/test ./... green, 18 packages. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
87 lines
3.9 KiB
Go
87 lines
3.9 KiB
Go
package web
|
|
|
|
// Dead-host roll-up honesty (v0.53.0, drill-1 observation; operator ruling 2026-07-13): a
|
|
// customer's status may never look better than its worst expected host. The customer roll-up
|
|
// derives from CONTROLLER reports, which reach the hub independently of the host agent — so a
|
|
// host DOWN for 23 hours hid behind a green customer row as long as the guest kept reporting
|
|
// (the live Peti-cluster shape: proxmox1 down 23h, fresh reports through proxmox2).
|
|
//
|
|
// foldHostStatus worsens the controller-derived status with per-host staleness via
|
|
// (*Server).hostStatus — THE single staleness definition (hosts.go; the same thresholds the
|
|
// HostStalenessChecker alerts on — no second definition anywhere). Display + derivation only:
|
|
// checker alerting is untouched.
|
|
|
|
import (
|
|
"time"
|
|
|
|
"gitea.dooplex.hu/admin/felhom-hub/internal/store"
|
|
)
|
|
|
|
// controllerStatus is the controller-report-derived customer status — the pre-roll-up chain the
|
|
// dashboard, the /configs list and the customer detail all inlined verbatim; this is now the
|
|
// ONE copy. Behavior-preserving: the branch order (incl. fail-after-warn) is the historical one.
|
|
func controllerStatus(c *store.CustomerSummary, threshold time.Duration) string {
|
|
// R-549 (operator ruling A, 2026-09-17): the threshold is CONFIGURATION (alerting.stale_threshold)
|
|
// and the checkers read it; this display must read the same value, "down" at 2x like them. It
|
|
// hardcoded 30 m / 1 h until hub v0.117.0 — a second definition the header above forbids, and one
|
|
// that would have painted a customer amber 15 minutes before the alarm could fire.
|
|
// Pinned by TestControllerStatus_FollowsConfiguredThreshold.
|
|
if threshold <= 0 {
|
|
threshold = 30 * time.Minute
|
|
}
|
|
switch {
|
|
case c.HealthStatus == "disabled":
|
|
return "disabled"
|
|
case c.TimeSinceReport > 2*threshold:
|
|
return "down"
|
|
case c.TimeSinceReport > threshold || c.HealthStatus == "warn":
|
|
return "warn"
|
|
case c.HealthStatus == "fail":
|
|
return "down"
|
|
default:
|
|
return "ok"
|
|
}
|
|
}
|
|
|
|
// hostFoldRank orders host states by badness for the worst-host pick. "ok" ranks 0 (never folds).
|
|
var hostFoldRank = map[string]int{"down": 3, "stale": 2, "pending": 1}
|
|
|
|
// hostFoldLabel is the operator-facing cause chip prefix per worst-host state.
|
|
var hostFoldLabel = map[string]string{"down": "host down", "stale": "host stale", "pending": "host pending"}
|
|
|
|
// foldHostStatus folds the customer's expected hosts into a controller-derived status:
|
|
// worst(controllerDerived, hostStatusOf(each host)). Any host down/stale caps the customer at
|
|
// WARN (a green row over a dead host is the masking bug); the returned cause names the state
|
|
// AND the host ("host down: <id>") so the detail header says WHICH host. "pending" hosts
|
|
// (enrolled, never reported) worsen only after initial onboarding — customerHasReported=false
|
|
// (the customer has never reported) excludes them, a half-installed box is not an incident.
|
|
// Statuses worse than warn (down) and administrative ones (disabled/blocked) keep their own
|
|
// token; the cause chip still surfaces the host signal. Read errors degrade to the unfolded
|
|
// status — the page must render.
|
|
func (s *Server) foldHostStatus(customerID, base string, customerHasReported bool) (status, cause string) {
|
|
hosts, err := s.store.ListHostsByCustomer(customerID)
|
|
if err != nil {
|
|
s.logger.Printf("[ERROR] roll-up: ListHostsByCustomer %s: %v", customerID, err)
|
|
return base, ""
|
|
}
|
|
worst, worstHost := "", ""
|
|
for i := range hosts {
|
|
hs := s.hostStatus(hosts[i].LastReportAt)
|
|
if hs == "pending" && !customerHasReported {
|
|
continue
|
|
}
|
|
if hostFoldRank[hs] > hostFoldRank[worst] {
|
|
worst, worstHost = hs, hosts[i].HostID
|
|
}
|
|
}
|
|
if worst == "" {
|
|
return base, ""
|
|
}
|
|
cause = hostFoldLabel[worst] + ": " + worstHost
|
|
// The fold worsens, never improves: ok / pending / no-report ("") cap at warn.
|
|
if base == "ok" || base == "pending" || base == "" {
|
|
return "warn", cause
|
|
}
|
|
return base, cause
|
|
}
|