2fc4a15fa3
gates / gates (push) Successful in 16s
The cooldown key was customerID:eventType plus the tier and run suffixes, and none of them names an app, so every app going down inside the same hour collapsed onto one key and only the first was mailed. Measured on demo-hp: bookstack sent 09:27:51, privatebin suppressed 09:31:51 under key=demo-hp:app_start_failed. cooldownStackSuffix is the third sibling of cooldownTierSuffix and cooldownRunSuffix, and separate for the reason the second one's docstring already gives: the existing two keep byte-identical semantics for every type that uses them. It is ALLOW-LISTED to app_start_failed and takes the event type as well as the details, unlike its siblings, and that asymmetry is the safety property. The backup family's cooldown is coarse ON PURPOSE (R-97a, R-182) so one full disk sends one digest rather than one mail per app - and crossdrive_failed is severity error, reaches the operator leg, and carries stack_name through a DIFFERENT struct, so a payload-shape rule would have split it silently. The hour itself does not change. Gate 11 refuses a push whose REPORT.md carries an observation with neither `FILED: R-NNN` nor `NOT-A-FINDING: <reason>`. It deliberately does NOT accept a passing mention of some other R-number: the lost item cited R-182 as an analogy, so "cites a register row" would have passed the very item the gate exists to catch. That discrepancy with the spec is recorded in the gate's docstring. Registered here and in the controller and agent runners. NOT in the catalog runner - it has no shared-gate mechanism and appends --all to every gate; filed as R-391 rather than left as a sentence, which is this session's lesson. PROMPT-TEMPLATE.md §15.9 corrected: "documented, NOT acted on" was the wording that invited the gap, and it now names the markers and points at the gate. R-390 filed for the golden-bake runbook's missing `pveam update`. Hub tests 709 -> 716.
127 lines
6.8 KiB
Markdown
127 lines
6.8 KiB
Markdown
# REPORT — felhom.eu: hub v0.107.0 (R-387), the alarm ladder, and golden 0.223.0
|
|
|
|
**Session 2026-08-23.** Companion to `felhom-controller` v0.223.0 (R-329, R-386) — see that repo's
|
|
`REPORT.md` for the controller work and the full live walk.
|
|
|
|
## 1. Baselines, and the hub's four numbers as read
|
|
|
|
| Repo | at start | at end |
|
|
|---|---|---|
|
|
| felhom.eu | `55274d5e` | hub **v0.107.0** deployed |
|
|
| felhom-controller | `14137efa` (v0.222.0) | **v0.223.0** deployed |
|
|
| felhom-agent | `40d857b5` | untouched |
|
|
|
|
**Hub's four numbers, read live from `GET /configuration` before starting:**
|
|
`golden_version` **0.222.0** · `agent_version` **0.130.0** · `min_agent` **0.129.0** ·
|
|
controller floor **0.222.0**. All four as the task predicted; the operator's 0.222.0 vouch had landed.
|
|
|
|
## 2. The hub's deployment path — §6's premise was wrong, and here it is
|
|
|
|
**`felhom.eu/manifests/hub.yaml`, line 128.** ArgoCD `Application/felhom` tracks
|
|
`https://gitea.dooplex.hu/admin/felhom.eu.git`, path `manifests`, with `syncPolicy.automated.enabled
|
|
= false`. Bumped `0.106.0 → 0.107.0` in commit **`68a9f54`**. **There is no out-of-git deployment
|
|
path** — the finding §6 braced for does not exist. The image was built and pushed to the registry
|
|
**before** the manifest landed, so a sync could never have pointed at a missing tag, and the sync was
|
|
then requested deliberately (`refresh=hard`, then a patched `operation`). Never `kubectl set image`.
|
|
|
|
## 3. R-387 — what was wrong
|
|
|
|
One handler, two fields, opposite discipline: an unknown `event_type` is rejected with a loud `400`;
|
|
an unknown `severity` was rewritten to `info` **without a word**, and `severityNotifies` drops `info`
|
|
before *both* legs. **The guard built to catch exactly this sat downstream of the rewrite** — the
|
|
dispatcher's `unrecognized severity` line can never execute for an API event, because the coercion one
|
|
line earlier guarantees the value it looks for cannot arrive.
|
|
|
|
**Measured on the live hub DB:** `91` `app_start_failed` events stored all-time, **`0`
|
|
`notification_log` rows before this session** — not one, on any channel, while every POST returned 200.
|
|
|
|
**The coercion stays.** A rejected event is a *lost* event, and losing an alarm is worse than
|
|
mis-routing one. Only the silence is fixed.
|
|
|
|
## 4. The dead-branch decision, and the reason
|
|
|
|
**KEPT.** Not caution — evidence. `cmd/hub/main.go` wires `dispatcher.ProcessEvent` **directly** as
|
|
the `monitor.EventNotifyFunc` for the staleness, host-staleness and offsite-box checkers, and those
|
|
hub-generated events never pass through the ingest handler at all. For every one of them that line is
|
|
the **only** severity guard there is. Deleting it as "dead" would have removed the live half while the
|
|
dead half supplied the justification.
|
|
|
|
Verified while deciding: **all 90 severity literals in `internal/monitor` are already valid**, so the
|
|
guard is silent because the producers are correct. (`"warn"` in `internal/web` is UI badge vocabulary,
|
|
not a severity.)
|
|
|
|
## 5. Files changed, commits, CI
|
|
|
|
| Commit | Contents |
|
|
|---|---|
|
|
| **`68a9f54`** | hub v0.107.0 (ingest WARN + kept-branch note + tests), manifest bump, golden 0.223.0 evidence |
|
|
| **`<docs>`** | alarm ladder §6.1/§7/§8, register, `STATUS.md`, `REPORT.md`, drill record |
|
|
|
|
Files: `hub/internal/api/handler.go`, `hub/internal/notify/dispatcher.go`, `hub/CHANGELOG.md`,
|
|
`manifests/hub.yaml`, `documentation/architecture/08-alarm-ladder.md`,
|
|
`documentation/backlog/{OPEN,CLOSED}-ITEMS.md`, `documentation/tests/golden-0.223.0-2026-08-23/`,
|
|
`documentation/audits/DRILL-r329-r386-2026-08-23/`, plus two new test files.
|
|
|
|
**CI runs confirmed BY ID** (`id` and `run_number` diverge — both printed): see §5 of the controller
|
|
REPORT for its runs; felhom.eu's are listed at the end of this file.
|
|
|
|
## 6. Tests and red-proofs
|
|
|
|
`internal/api/r387_severity_visibility_test.go` — the event is **not lost**, the stored severity is
|
|
still `info`, and the WARN names customer + type + value; plus a guard that a **valid** severity stays
|
|
silent, because an alarm on the normal path is one people learn to ignore.
|
|
`internal/notify/r329_app_start_failed_test.go` — the routing consequence: operator emailed, customer
|
|
not, unless opted in, in which case both legs deliver and the customer's copy carries the Hungarian
|
|
template. Scenario B is also the **positive control** for Scenario A's absence claim.
|
|
|
|
Test count **702 → 709**.
|
|
|
|
**Red-proof (seen failing):** delete the ingest `WARN` → `the hub rewrote a severity and said
|
|
nothing`, with the log showing only the ordinary `[INFO] Event from c1: backup_failed (info)`. The
|
|
guard sits at **ingest**, because that is the last point at which the offending value still exists.
|
|
|
|
## 7. Golden
|
|
|
|
**Baked and PUBLISHED: 0.223.0.** `GOLDEN_SHA256 =
|
|
9eaf39ac39219b42ec9e6cbf890275febcdcc6f53325fe0c0f591d3431044f17`; `upload OK (HTTP 201)`; round-trip
|
|
**HTTP 206**; all five acceptance markers counted (`docker OK (overlay2` 1, `including mount point` 2,
|
|
`upload OK` 1, `excluding` 0, `FATAL` 0). **VOUCHING IS THE OPERATOR'S ACT AND WAS NOT DONE HERE.**
|
|
|
|
**Runbook deviation, second session running:** §4.1 omits `pveam update`, so the `virgin` snapshot's
|
|
stale template index fails as `400 … no such template`.
|
|
|
|
## 8. Scenario H, live against v0.107.0
|
|
|
|
```
|
|
[WARN] [api] Event from demo-hp: severity "warn" is not in {info,warning,error,critical} — coercing
|
|
to "info", which severityNotifies DROPS, so this backup_failed alert will reach NOBODY. Fix the
|
|
emitting controller; this event is stored but not routed.
|
|
```
|
|
|
|
Both the bad-severity POST and the `error` control returned **200** (nothing lost), and the control
|
|
produced **no** warning — the guard does not fire on the normal path.
|
|
|
|
## 9. Register size
|
|
|
|
| File | Before | After |
|
|
|---|---|---|
|
|
| `OPEN-ITEMS.md` | 328,325 B | **328,132 B** |
|
|
| `CLOSED-ITEMS.md` | 71,441 B | **74,642 B** |
|
|
|
|
R-329 and R-386 compressed into CLOSED with their rules kept; **R-387** (closed) and **R-388** (the
|
|
notification-model product decision, open, operator's call) filed.
|
|
|
|
## 10. Observations
|
|
|
|
> **Markers added 2026-08-24 (gate 11, R-389).** Item 1 is the finding that had no row; adding its
|
|
> marker is the first thing the gate ever asked for. The observations' text is unchanged.
|
|
|
|
1. **The operator cooldown key carries no app identifier.** PrivateBin's alarm four minutes after
|
|
BookStack's was logged `suppressed — operator cooldown 1h, key=demo-hp:app_start_failed`, so **only
|
|
the first app-down per hour reaches the operator by e-mail**. R-182's known shape; harmless while
|
|
the event was undeliverable, and no longer. Not fixed here.
|
|
FILED: R-389
|
|
2. Two probe events remain as rows for `demo-hp` from Scenario H — inert, and named rather than left.
|
|
NOT-A-FINDING: two inert event rows on a Tier 0 demo box, created deliberately as a live
|
|
control and named in that session's teardown; they carry no state and nothing reads them.
|