hub v0.107.0: the hub rewrote a severity and said nothing (R-387); golden 0.223.0
gates / gates (push) Successful in 17s
gates / gates (push) Successful in 17s
One handler, two fields, opposite discipline. An unknown event_type is rejected with a loud 400. An unknown severity was rewritten to "info" without a word - and severityNotifies drops "info" before BOTH legs, so the event was stored, answered 200, and mailed to nobody. Two shipped features went out that way: DiskAlertKind.Severity emitted "warn" until controller v0.215.0, app_start_failed until v0.223.0. Measured on the live hub DB today: 91 app_start_failed events stored all-time, ZERO notification_log rows before this session - not one, on any channel. The mechanism built to catch this class was structurally blind to it: the dispatcher's `unrecognized severity` line cannot execute for anything arriving over the API, because the coercion one line earlier guarantees the value it looks for cannot arrive. The coercion STAYS - a rejected event is a lost event, and losing an alarm is worse than mis-routing one. Only the silence is fixed: a WARN naming the customer, the event type and the rejected value. The dispatcher branch is KEPT, not deleted as dead, and the reason is evidence rather than caution: cmd/hub/main.go wires dispatcher.ProcessEvent DIRECTLY as the monitor.EventNotifyFunc for the staleness, host-staleness and offsite-box checkers, which never pass through the handler. For those it is the only severity guard there is. All 90 severity literals in internal/monitor are already valid, so the guard is silent because the producers are correct. Test count 702 -> 709. Red-proof seen failing: delete the WARN line and the coercion test fails with "the hub rewrote a severity and said nothing". Golden 0.223.0 baked and published (sha 9eaf39ac3921...), round-trip HTTP 206. Vouching is the operator's act and was not done here.
This commit is contained in:
@@ -141,6 +141,21 @@ func (d *Dispatcher) ProcessEvent(customerID, eventType, severity, message, deta
|
||||
// warning / error / critical trigger notifications. "info" is an intentional non-notify (status/
|
||||
// recovery events). Anything else is UNRECOGNIZED — log it (don't silently drop), so a bad severity
|
||||
// surfaces instead of vanishing (the felhom-pve-class lesson: a critical event must never be lost).
|
||||
//
|
||||
// R-387 — THIS BRANCH IS **KEPT DELIBERATELY**, and here is why, because the question was asked
|
||||
// and a branch that cannot execute without a note is the thing to avoid.
|
||||
//
|
||||
// For an event arriving over the API it is genuinely unreachable: the ingest handler coerces any
|
||||
// unknown severity to "info" before this is called, so the one value it looks for cannot arrive.
|
||||
// **But the API is not the only producer.** `cmd/hub/main.go` wires `dispatcher.ProcessEvent`
|
||||
// DIRECTLY as the `monitor.EventNotifyFunc` for the staleness, host-staleness and offsite-box
|
||||
// checkers, and those hub-generated events never pass through the handler at all. For every one
|
||||
// of them this line is the ONLY severity guard there is.
|
||||
//
|
||||
// Removing it as "dead" would therefore have deleted the live half while leaving the dead half
|
||||
// looking like the reason. Verified 2026-08-23: every severity literal in `internal/monitor` (90
|
||||
// of them) is already in the vocabulary — so the guard is currently silent because the producers
|
||||
// are correct, which is exactly what a guard looks like when it is working.
|
||||
if !severityNotifies(severity) {
|
||||
if severity != "info" {
|
||||
d.logger.Printf("[WARN] Dispatcher: unrecognized severity %q for %s/%s — not routing", severity, customerID, eventType)
|
||||
|
||||
Reference in New Issue
Block a user