hub v0.107.0: the hub rewrote a severity and said nothing (R-387); golden 0.223.0
gates / gates (push) Successful in 17s

One handler, two fields, opposite discipline. An unknown event_type is rejected
with a loud 400. An unknown severity was rewritten to "info" without a word -
and severityNotifies drops "info" before BOTH legs, so the event was stored,
answered 200, and mailed to nobody.

Two shipped features went out that way: DiskAlertKind.Severity emitted "warn"
until controller v0.215.0, app_start_failed until v0.223.0. Measured on the live
hub DB today: 91 app_start_failed events stored all-time, ZERO notification_log
rows before this session - not one, on any channel.

The mechanism built to catch this class was structurally blind to it: the
dispatcher's `unrecognized severity` line cannot execute for anything arriving
over the API, because the coercion one line earlier guarantees the value it
looks for cannot arrive.

The coercion STAYS - a rejected event is a lost event, and losing an alarm is
worse than mis-routing one. Only the silence is fixed: a WARN naming the
customer, the event type and the rejected value.

The dispatcher branch is KEPT, not deleted as dead, and the reason is evidence
rather than caution: cmd/hub/main.go wires dispatcher.ProcessEvent DIRECTLY as
the monitor.EventNotifyFunc for the staleness, host-staleness and offsite-box
checkers, which never pass through the handler. For those it is the only
severity guard there is. All 90 severity literals in internal/monitor are
already valid, so the guard is silent because the producers are correct.

Test count 702 -> 709. Red-proof seen failing: delete the WARN line and the
coercion test fails with "the hub rewrote a severity and said nothing".

Golden 0.223.0 baked and published (sha 9eaf39ac3921...), round-trip HTTP 206.
Vouching is the operator's act and was not done here.
This commit is contained in:
2026-08-23 11:57:26 +02:00
parent 55274d5ef3
commit 68a9f5475c
8 changed files with 762 additions and 2 deletions
+15
View File
@@ -141,6 +141,21 @@ func (d *Dispatcher) ProcessEvent(customerID, eventType, severity, message, deta
// warning / error / critical trigger notifications. "info" is an intentional non-notify (status/
// recovery events). Anything else is UNRECOGNIZED — log it (don't silently drop), so a bad severity
// surfaces instead of vanishing (the felhom-pve-class lesson: a critical event must never be lost).
//
// R-387 — THIS BRANCH IS **KEPT DELIBERATELY**, and here is why, because the question was asked
// and a branch that cannot execute without a note is the thing to avoid.
//
// For an event arriving over the API it is genuinely unreachable: the ingest handler coerces any
// unknown severity to "info" before this is called, so the one value it looks for cannot arrive.
// **But the API is not the only producer.** `cmd/hub/main.go` wires `dispatcher.ProcessEvent`
// DIRECTLY as the `monitor.EventNotifyFunc` for the staleness, host-staleness and offsite-box
// checkers, and those hub-generated events never pass through the handler at all. For every one
// of them this line is the ONLY severity guard there is.
//
// Removing it as "dead" would therefore have deleted the live half while leaving the dead half
// looking like the reason. Verified 2026-08-23: every severity literal in `internal/monitor` (90
// of them) is already in the vocabulary — so the guard is currently silent because the producers
// are correct, which is exactly what a guard looks like when it is working.
if !severityNotifies(severity) {
if severity != "info" {
d.logger.Printf("[WARN] Dispatcher: unrecognized severity %q for %s/%s — not routing", severity, customerID, eventType)