One handler, two fields, opposite discipline. An unknown event_type is rejected with a loud 400. An unknown severity was rewritten to "info" without a word - and severityNotifies drops "info" before BOTH legs, so the event was stored, answered 200, and mailed to nobody. Two shipped features went out that way: DiskAlertKind.Severity emitted "warn" until controller v0.215.0, app_start_failed until v0.223.0. Measured on the live hub DB today: 91 app_start_failed events stored all-time, ZERO notification_log rows before this session - not one, on any channel. The mechanism built to catch this class was structurally blind to it: the dispatcher's `unrecognized severity` line cannot execute for anything arriving over the API, because the coercion one line earlier guarantees the value it looks for cannot arrive. The coercion STAYS - a rejected event is a lost event, and losing an alarm is worse than mis-routing one. Only the silence is fixed: a WARN naming the customer, the event type and the rejected value. The dispatcher branch is KEPT, not deleted as dead, and the reason is evidence rather than caution: cmd/hub/main.go wires dispatcher.ProcessEvent DIRECTLY as the monitor.EventNotifyFunc for the staleness, host-staleness and offsite-box checkers, which never pass through the handler. For those it is the only severity guard there is. All 90 severity literals in internal/monitor are already valid, so the guard is silent because the producers are correct. Test count 702 -> 709. Red-proof seen failing: delete the WARN line and the coercion test fails with "the hub rewrote a severity and said nothing". Golden 0.223.0 baked and published (sha 9eaf39ac3921...), round-trip HTTP 206. Vouching is the operator's act and was not done here.
2.1 KiB
Golden bake — controller 0.223.0 (2026-08-23)
Baked and PUBLISHED by Claude Code. NOT vouched — vouching is the operator's act.
| Field | Value |
|---|---|
GOLDEN_VERSION |
0.223.0 |
GOLDEN_SHA256 |
9eaf39ac39219b42ec9e6cbf890275febcdcc6f53325fe0c0f591d3431044f17 |
| Controller image | gitea.dooplex.hu/admin/felhom-controller:0.223.0 |
| MinAgent | 0.129.0 (unchanged) |
| Package URL | https://gitea.dooplex.hu/api/packages/admin/generic/felhom-golden/0.223.0/golden.tar.zst |
| Archive size | 657,261,745 B |
| LXC template | debian-13-standard_13.6-1_amd64.tar.zst |
Acceptance markers, each counted from bake.log
| Marker | Required | Observed |
|---|---|---|
docker OK (overlay2 |
≥1 | 1 — docker OK (overlay2; data-root /var/lib/docker) |
including mount point (rootfs + mp0) |
2 | 2 |
upload OK (HTTP 201) |
1 | 1 |
excluding |
0 | 0 |
FATAL |
0 | 0 |
Round trip on the published package: HTTP 206 on a ranged GET.
Why this release needs a golden
app_start_failed is now deliverable and gains a customer-facing toggle — customer-visible behaviour
on a fresh install. A machine installed from the previous golden would receive a controller whose
app-down alarm reaches nobody.
The runbook step that is still missing
§4.1 does not say to run pveam update first. On the virgin snapshot the template index is
stale, so pveam available offers an old point release and downloading it fails with
400 Parameter verification failed. template: no such template — a confusing 400 rather than a
legible "your index is old". Second bake in a row to hit it; recorded in the workspace memory as
golden-bake-needs-pveam-update. The runbook itself is still not edited.
Vouching — the OPERATOR's step, not done here
Hub → Configuration → Day-0 artifacts:
golden_version→0.223.0golden_sha256→9eaf39ac39219b42ec9e6cbf890275febcdcc6f53325fe0c0f591d3431044f17min_agent→0.129.0(unchanged)- then, last and in its own save, the global controller floor →
0.223.0.