Files
felhom.eu/REPORT.md
T
admin 2f7c9a6ce5
gates / gates (push) Successful in 17s
docs(R-329/R-386/R-387): the severity contract, the intent ruling, and Part 5 recorded
The alarm ladder gains the severity contract (the hub's vocabulary is exact, it
coerces silently, and three things now hold it) and the intent test with its
three-way ruling on unknown. Both marked [DESIGN] with the live measurements.

Part 5 is RECORDED AND NOT IMPLEMENTED: the operator's notification philosophy,
verbatim, marked plainly as direction rather than current behaviour, with the
12 -> 15 toggle growth as the argument. Filed as R-388, a product decision.

R-329 and R-386 compressed into CLOSED-ITEMS with their rules kept and the
full-text commit named. R-387 filed closed - including WHY the dispatcher branch
was kept rather than deleted, which is evidence (three monitor checkers call
ProcessEvent directly) and not caution.

The drill record names three things that had to be re-run: an inert red-proof
mutation, Scenario G refused twice behind an HTTP 200, and the live Scenario A
NOT proving the customer gate because demo-hp has no prefs row at all.

Register: OPEN 328325 -> 328132 B, CLOSED 71441 -> 74642 B.
2026-08-23 12:03:49 +02:00

6.4 KiB

REPORT — felhom.eu: hub v0.107.0 (R-387), the alarm ladder, and golden 0.223.0

Session 2026-08-23. Companion to felhom-controller v0.223.0 (R-329, R-386) — see that repo's REPORT.md for the controller work and the full live walk.

1. Baselines, and the hub's four numbers as read

Repo at start at end
felhom.eu 55274d5e hub v0.107.0 deployed
felhom-controller 14137efa (v0.222.0) v0.223.0 deployed
felhom-agent 40d857b5 untouched

Hub's four numbers, read live from GET /configuration before starting: golden_version 0.222.0 · agent_version 0.130.0 · min_agent 0.129.0 · controller floor 0.222.0. All four as the task predicted; the operator's 0.222.0 vouch had landed.

2. The hub's deployment path — §6's premise was wrong, and here it is

felhom.eu/manifests/hub.yaml, line 128. ArgoCD Application/felhom tracks https://gitea.dooplex.hu/admin/felhom.eu.git, path manifests, with syncPolicy.automated.enabled = false. Bumped 0.106.0 → 0.107.0 in commit 68a9f54. There is no out-of-git deployment path — the finding §6 braced for does not exist. The image was built and pushed to the registry before the manifest landed, so a sync could never have pointed at a missing tag, and the sync was then requested deliberately (refresh=hard, then a patched operation). Never kubectl set image.

3. R-387 — what was wrong

One handler, two fields, opposite discipline: an unknown event_type is rejected with a loud 400; an unknown severity was rewritten to info without a word, and severityNotifies drops info before both legs. The guard built to catch exactly this sat downstream of the rewrite — the dispatcher's unrecognized severity line can never execute for an API event, because the coercion one line earlier guarantees the value it looks for cannot arrive.

Measured on the live hub DB: 91 app_start_failed events stored all-time, 0 notification_log rows before this session — not one, on any channel, while every POST returned 200.

The coercion stays. A rejected event is a lost event, and losing an alarm is worse than mis-routing one. Only the silence is fixed.

4. The dead-branch decision, and the reason

KEPT. Not caution — evidence. cmd/hub/main.go wires dispatcher.ProcessEvent directly as the monitor.EventNotifyFunc for the staleness, host-staleness and offsite-box checkers, and those hub-generated events never pass through the ingest handler at all. For every one of them that line is the only severity guard there is. Deleting it as "dead" would have removed the live half while the dead half supplied the justification.

Verified while deciding: all 90 severity literals in internal/monitor are already valid, so the guard is silent because the producers are correct. ("warn" in internal/web is UI badge vocabulary, not a severity.)

5. Files changed, commits, CI

Commit Contents
68a9f54 hub v0.107.0 (ingest WARN + kept-branch note + tests), manifest bump, golden 0.223.0 evidence
<docs> alarm ladder §6.1/§7/§8, register, STATUS.md, REPORT.md, drill record

Files: hub/internal/api/handler.go, hub/internal/notify/dispatcher.go, hub/CHANGELOG.md, manifests/hub.yaml, documentation/architecture/08-alarm-ladder.md, documentation/backlog/{OPEN,CLOSED}-ITEMS.md, documentation/tests/golden-0.223.0-2026-08-23/, documentation/audits/DRILL-r329-r386-2026-08-23/, plus two new test files.

CI runs confirmed BY ID (id and run_number diverge — both printed): see §5 of the controller REPORT for its runs; felhom.eu's are listed at the end of this file.

6. Tests and red-proofs

internal/api/r387_severity_visibility_test.go — the event is not lost, the stored severity is still info, and the WARN names customer + type + value; plus a guard that a valid severity stays silent, because an alarm on the normal path is one people learn to ignore. internal/notify/r329_app_start_failed_test.go — the routing consequence: operator emailed, customer not, unless opted in, in which case both legs deliver and the customer's copy carries the Hungarian template. Scenario B is also the positive control for Scenario A's absence claim.

Test count 702 → 709.

Red-proof (seen failing): delete the ingest WARN → the hub rewrote a severity and said nothing, with the log showing only the ordinary [INFO] Event from c1: backup_failed (info). The guard sits at ingest, because that is the last point at which the offending value still exists.

7. Golden

Baked and PUBLISHED: 0.223.0. GOLDEN_SHA256 = 9eaf39ac39219b42ec9e6cbf890275febcdcc6f53325fe0c0f591d3431044f17; upload OK (HTTP 201); round-trip HTTP 206; all five acceptance markers counted (docker OK (overlay2 1, including mount point 2, upload OK 1, excluding 0, FATAL 0). VOUCHING IS THE OPERATOR'S ACT AND WAS NOT DONE HERE.

Runbook deviation, second session running: §4.1 omits pveam update, so the virgin snapshot's stale template index fails as 400 … no such template.

8. Scenario H, live against v0.107.0

[WARN] [api] Event from demo-hp: severity "warn" is not in {info,warning,error,critical} — coercing
to "info", which severityNotifies DROPS, so this backup_failed alert will reach NOBODY. Fix the
emitting controller; this event is stored but not routed.

Both the bad-severity POST and the error control returned 200 (nothing lost), and the control produced no warning — the guard does not fire on the normal path.

9. Register size

File Before After
OPEN-ITEMS.md 328,325 B 328,132 B
CLOSED-ITEMS.md 71,441 B 74,642 B

R-329 and R-386 compressed into CLOSED with their rules kept; R-387 (closed) and R-388 (the notification-model product decision, open, operator's call) filed.

10. Observations — recorded, not acted on

  1. The operator cooldown key carries no app identifier. PrivateBin's alarm four minutes after BookStack's was logged suppressed — operator cooldown 1h, key=demo-hp:app_start_failed, so only the first app-down per hour reaches the operator by e-mail. R-182's known shape; harmless while the event was undeliverable, and no longer. Not fixed here.
  2. Two probe events remain as rows for demo-hp from Scenario H — inert, and named rather than left.