docs(hub v0.108.0): the delivery grain, the cooldown ruling, and gate 11's first subject
gates / gates (push) Successful in 15s

The alarm ladder gains §6.2 - which events are per-app, per-run, per-tier or
coarse, and why the default is coarse. CONTEXT records two rulings: the grain is
allow-listed rather than inferred from the payload, with crossdrive_failed as
the proof that a payload rule would have been wrong; and a finding recorded only
in REPORT.md has a lifetime of one session.

R-389 closed and compressed, keeping its rules and naming the commit whose
git show returns the full text. R-390 and R-391 left open.

REPORT.md is gate 11's first real subject and passes: six observations, two
FILED, four NOT-A-FINDING with their reasons. Three of those declarations are
things a tidier report would have omitted - the gate's own spec would have
passed the item it was built to catch, the burst has no ceiling, and ArgoCD
said "successfully rolled out" while still running the old image.

STATUS carries forward the one thing outstanding: the controller floor still
reads 0.222.0 while the golden reads 0.223.0.
This commit is contained in:
2026-08-23 14:12:23 +02:00
parent 45659bdc5a
commit ebdc04601d
24 changed files with 2514 additions and 111 deletions
+235 -95
View File
@@ -1,126 +1,266 @@
# REPORT — felhom.eu: hub v0.107.0 (R-387), the alarm ladder, and golden 0.223.0
# REPORT — hub v0.108.0 (R-389), gate 11, and the instruction that invited the gap
**Session 2026-08-23.** Companion to `felhom-controller` v0.223.0 (R-329, R-386) — see that repo's
`REPORT.md` for the controller work and the full live walk.
**Session 2026-08-23.** `felhom.eu` is the subject; the controller and agent were touched only to
register the shared gate. **No controller release — no golden bake, no vouch, no floor.**
**No halt condition fired.** Nothing was dropped.
## 1. Baselines, and the hub's four numbers as read
| Repo | at start | at end |
|---|---|---|
| felhom.eu | `55274d5e` | hub **v0.107.0** deployed |
| felhom-controller | `14137efa` (v0.222.0) | **v0.223.0** deployed |
| felhom-agent | `40d857b5` | untouched |
| felhom.eu | `2f7c9a6` (hub v0.107.0) | **hub v0.108.0** deployed |
| felhom-controller | `1da2c9c` (v0.223.0) | **unchanged** — runner registration only |
| felhom-agent | `40d857b` (v0.130.0) | **unchanged** — runner registration only |
**Hub's four numbers, read live from `GET /configuration` before starting:**
`golden_version` **0.222.0** · `agent_version` **0.130.0** · `min_agent` **0.129.0** ·
controller floor **0.222.0**. All four as the task predicted; the operator's 0.222.0 vouch had landed.
**Hub's four numbers, live from `GET /configuration` before starting:**
## 2. The hub's deployment path — §6's premise was wrong, and here it is
**`felhom.eu/manifests/hub.yaml`, line 128.** ArgoCD `Application/felhom` tracks
`https://gitea.dooplex.hu/admin/felhom.eu.git`, path `manifests`, with `syncPolicy.automated.enabled
= false`. Bumped `0.106.0 → 0.107.0` in commit **`68a9f54`**. **There is no out-of-git deployment
path** — the finding §6 braced for does not exist. The image was built and pushed to the registry
**before** the manifest landed, so a sync could never have pointed at a missing tag, and the sync was
then requested deliberately (`refresh=hard`, then a patched `operation`). Never `kubectl set image`.
## 3. R-387 — what was wrong
One handler, two fields, opposite discipline: an unknown `event_type` is rejected with a loud `400`;
an unknown `severity` was rewritten to `info` **without a word**, and `severityNotifies` drops `info`
before *both* legs. **The guard built to catch exactly this sat downstream of the rewrite** — the
dispatcher's `unrecognized severity` line can never execute for an API event, because the coercion one
line earlier guarantees the value it looks for cannot arrive.
**Measured on the live hub DB:** `91` `app_start_failed` events stored all-time, **`0`
`notification_log` rows before this session** — not one, on any channel, while every POST returned 200.
**The coercion stays.** A rejected event is a *lost* event, and losing an alarm is worse than
mis-routing one. Only the silence is fixed.
## 4. The dead-branch decision, and the reason
**KEPT.** Not caution — evidence. `cmd/hub/main.go` wires `dispatcher.ProcessEvent` **directly** as
the `monitor.EventNotifyFunc` for the staleness, host-staleness and offsite-box checkers, and those
hub-generated events never pass through the ingest handler at all. For every one of them that line is
the **only** severity guard there is. Deleting it as "dead" would have removed the live half while the
dead half supplied the justification.
Verified while deciding: **all 90 severity literals in `internal/monitor` are already valid**, so the
guard is silent because the producers are correct. (`"warn"` in `internal/web` is UI badge vocabulary,
not a severity.)
## 5. Files changed, commits, CI
| Commit | Contents |
| Field | Value |
|---|---|
| **`68a9f54`** | hub v0.107.0 (ingest WARN + kept-branch note + tests), manifest bump, golden 0.223.0 evidence |
| **`<docs>`** | alarm ladder §6.1/§7/§8, register, `STATUS.md`, `REPORT.md`, drill record |
| `golden_version` | **0.223.0** |
| `agent_version` | **0.130.0** |
| `min_agent` | **0.129.0** |
| controller floor (`min_controller_version`) | **0.222.0** |
Files: `hub/internal/api/handler.go`, `hub/internal/notify/dispatcher.go`, `hub/CHANGELOG.md`,
`manifests/hub.yaml`, `documentation/architecture/08-alarm-ladder.md`,
`documentation/backlog/{OPEN,CLOSED}-ITEMS.md`, `documentation/tests/golden-0.223.0-2026-08-23/`,
`documentation/audits/DRILL-r329-r386-2026-08-23/`, plus two new test files.
**The task predicted the floor at 0.223.0 and it reads 0.222.0.** The operator vouched the golden but
has not yet raised the floor — the last step of the previous release, which `STATUS.md` says to do
"last, in its own save". Carried forward as item 1 there. Not a halt; the correct reading is simply
different from the prediction.
**CI runs confirmed BY ID** (`id` and `run_number` diverge — both printed): see §5 of the controller
REPORT for its runs; felhom.eu's are listed at the end of this file.
**The golden-currency gate stayed green throughout** and no bake was needed, exactly as §1 said it
should be: the controller CHANGELOG's newest entry and the newest baked golden are both 0.223.0 and
this session moved neither.
## 6. Tests and red-proofs
## 2. Documents read
`internal/api/r387_severity_visibility_test.go` — the event is **not lost**, the stored severity is
still `info`, and the WARN names customer + type + value; plus a guard that a **valid** severity stays
silent, because an alarm on the normal path is one people learn to ignore.
`internal/notify/r329_app_start_failed_test.go` — the routing consequence: operator emailed, customer
not, unless opted in, in which case both legs deliver and the customer's copy carries the Hungarian
template. Scenario B is also the **positive control** for Scenario A's absence claim.
`hub/internal/notify/dispatcher.go` (`cooldownTierSuffix` + `cooldownRunSuffix` docstrings in full,
`processOperator` and its R-182 suppression-logging block, `operatorOnlyEvents`),
`scripts/repo_gates.py` (whole docstring, including "WHY 10 IS HERE" and "WHY 7 IS HERE"),
`scripts/due_checks_gate.py`, `documentation/PROMPT-TEMPLATE.md` §15, and the alarm ladder at
**`documentation/architecture/08-alarm-ladder.md`** — extended here with §6.2, the delivery grain.
Test count **702 → 709**.
## 3. The `AppDetails` emitter count, measured
**Red-proof (seen failing):** delete the ingest `WARN` → `the hub rewrote a severity and said
nothing`, with the log showing only the ordinary `[INFO] Event from c1: backup_failed (info)`. The
guard sits at **ingest**, because that is the last point at which the offending value still exists.
**Three, exactly as §3 said.** `grep -rn "AppDetails{" --include=*.go` over the controller, excluding
tests:
## 7. Golden
| Emitter | Event | Severity | Reaches the operator leg? |
|---|---|---|---|
| `notifier.go:502` | `app_deployed` | `info` | no — `info` is dropped by `severityNotifies` |
| `notifier.go:563` | `app_start_failed` | `warning` | **yes** |
| `notifier.go:661` | `app_removed` | `info` | no |
**Baked and PUBLISHED: 0.223.0.** `GOLDEN_SHA256 =
9eaf39ac39219b42ec9e6cbf890275febcdcc6f53325fe0c0f591d3431044f17`; `upload OK (HTTP 201)`; round-trip
**HTTP 206**; all five acceptance markers counted (`docker OK (overlay2` 1, `including mount point` 2,
`upload OK` 1, `excluding` 0, `FATAL` 0). **VOUCHING IS THE OPERATOR'S ACT AND WAS NOT DONE HERE.**
**No fourth emitter. No halt.**
**Runbook deviation, second session running:** §4.1 omits `pveam update`, so the `virgin` snapshot's
stale template index fails as `400 … no such template`.
**But the sweep found something the `AppDetails` question could not:** `stack_name` is also carried by
a **different struct**, `CrossDriveDetails` (`notifier.go:151-157`), used by `crossdrive_failed`
(severity **`error`**, so it *does* reach the operator leg) and `crossdrive_completed`. This is what
makes the allow-list load-bearing in fact rather than in principle — a payload-shape rule would have
split a backup-family event per app and silently undone R-182. It is Scenario C's live subject.
## 8. Scenario H, live against v0.107.0
## 4. Files, commits, CI
| Commit | Repo | Contents |
|---|---|---|
| **`f751aea`** | felhom.eu | R-389 filed — **alone, before any code** (Phase 1) |
| **`2fc4a15`** | felhom.eu | the suffix + allow-list + tests, gate 11, template fix, R-390/R-391 |
| **`45659bd`** | felhom.eu | hub v0.108.0 CHANGELOG + manifest bump |
| **`f8c9390`** | felhom-controller | gate 11 registered; its REPORT's observations marked up |
| **`058b945`** | felhom-agent | gate 11 registered |
**CI runs confirmed BY ID** — and the listing was checked for truncation rather than trusted, since a
silently truncated listing has already cost this arc a false claim:
| Commit | Repo | CI `id` | `run_number` | Result |
|---|---|---|---|---|
| `f751aea` | felhom.eu | **412** | 263 | success |
| `2fc4a15` | felhom.eu | **413** | 264 | success |
| `45659bd` | felhom.eu | **416** | 265 | success |
| `f8c9390` | felhom-controller | **414** | 90 | success |
| `058b945` | felhom-agent | **415** | 55 | success |
The API reported `total_count` 265 / 91 / 55 against 3 rows shown in each case — i.e. the pages were
known-partial and the newest rows are the ones quoted, not the whole set.
## 5. Red-proofs — two planted, both seen failing
| # | Mutation | Layer, and why that layer | Observed |
|---|---|---|---|
| 1 | `cooldownStackSuffix` dropped from the key expression | **`processOperator`'s KEY** — where the collapse physically happens | `2 apps down inside the hour produced 1 operator mail(s), want 2 (suppressed=1)`, and the suppression row reads `key=c1:app_start_failed` — the live shape reproduced in a unit test |
| 2 | the allow-list check removed from the suffix | **the REGISTER** — the fence that keeps the backup family coarse | `crossdrive_failed` … `= ":bookstack", want ""`, plus `app_deployed`, `app_removed`, `backup_failed` — the fence convicting exactly the types it was written for |
Both mutations asserted their pre-fix text was present before rewriting and printed `MUTATION
APPLIED`. **Neither passed first time**, and the check for that was explicit after yesterday's inert
mutation.
Gate 11 carries its own controls rather than a mutation, because the gate *is* the guard: ten cases
in §Part 3 of the drill record, including the historical red-proof against yesterday's real file.
## 6. Test counts
| Repo | Before | After |
|---|---|---|
| felhom.eu hub | 709 | **716** |
`go build ./... && go vet ./... && go test ./...` in the hub → **exit 0, zero failures**.
`python3 scripts/repo_gates.py --fast` → **12/12 OK** in felhom.eu; controller and agent runners both
OK with gate 11 registered.
## 7. Deployed hub version, and the manifest commit
**`gitea.dooplex.hu/admin/felhom-hub:0.108.0`**, ArgoCD `Synced` / `Healthy`, rolled out.
Deployed by **`45659bd`**, which bumped `manifests/hub.yaml:128`. The image was pushed to the registry
**before** that commit landed, so a sync could never have pointed at a missing tag. Never
`kubectl set image`.
**One thing worth stating because it looked like success and was not:** the first `refresh=hard` +
sync reported `successfully rolled out` while the deployment still read **0.107.0** and the app read
`OutOfSync` — ArgoCD had synced a pre-push revision. A second hard refresh took it to
`Synced rev=45659bd` and the image then read 0.108.0. **The rollout message alone would have been a
false confirmation**; the image tag is the observable that settles it.
## 8. The live walk
All counts filtered `created_at >= T0` (`2026-08-23 11:56:06Z`) so yesterday's two inert Scenario H
probe rows cannot contaminate them.
### Step 1 — Scenario A: two different apps, four minutes apart ✅
```
[WARN] [api] Event from demo-hp: severity "warn" is not in {info,warning,error,critical} — coercing
to "info", which severityNotifies DROPS, so this backup_failed alert will reach NOBODY. Fix the
emitting controller; this event is stored but not routed.
sent operator Telepített alkalmazás nem fut: OpenGist 2026-08-23 11:56:57
sent operator Telepített alkalmazás nem fut: Calibre-Web A 2026-08-23 12:00:57
sent: 2 suppressed: 0
```
Both the bad-severity POST and the `error` control returned **200** (nothing lost), and the control
produced **no** warning — the guard does not fire on the normal path.
Yesterday, the identical shape: `bookstack` **sent** 09:27:51, `privatebin` **suppressed** 09:31:51
under `key=demo-hp:app_start_failed`.
## 9. Register size
### Step 2 — Scenario B: each app again inside the hour ✅
```
suppressed OpenGist operator cooldown 1h, key=demo-hp:app_start_failed:opengist 12:02:48
suppressed Calibre-Web operator cooldown 1h, key=demo-hp:app_start_failed:calibre-web 12:02:48
```
One `sent` and one `suppressed` per app — **the hour is unchanged** — and the two keys **differ by the
app**, against v0.107.0's single shared key. The hub records the key only on a suppression, which is
why this step is what exposes it.
*Method note:* the repeats were posted through `/api/v1/event`, the exact endpoint the controller
invokes, with the controller's own `AppDetails` payload. The controller's own event is edge-triggered
per app, so a down→down cycle is silent **by design** and cannot re-fire from the box.
### Step 3 — Scenario C: a backup-family event carrying `stack_name` ✅
```
sent opengist
suppressed calibre-web operator cooldown 1h, key=demo-hp:crossdrive_failed 12:03:44
```
**Byte-identical to v0.107.0's key, with no app suffix.** The "before" value was obtained two
independent ways, both stated: derivation from the v0.107.0 expression (which has no app term), and
`TestR389_NoOtherEventTypeKeyChanges`, which models that expression inline for 12 event types **and
carries a positive control proving it can see a key change before reporting that none occurred**.
### Step 4 — Part 2's burst ✅ (see §9)
### Step 5 — Part 3's gate ✅ all controls plus the historical red-proof (see §10)
## 9. Part 2's three counts, and the judgement
Three apps stopped in one scan, after checking none carried a live cooldown — a stale one would have
halved the count and made the answer look better than it is:
| | |
|---|---|
| attempted | **3** |
| sent | **3** |
| suppressed | **0** |
**Plain judgement: per-app is the right grain and this volume is acceptable.** The reference box has 8
deployed apps, so a total outage is 8 mails; the boot grace (90 s), the quiesce grace (180 s) and the
per-app edge trigger absorb reboots, backup cycles and persistently-dead apps. **No burst-digest row
was filed**, and the condition that would reopen it is recorded rather than left implicit: the volume
scales linearly with app count and has no ceiling, so a box large enough that a total outage is
unreadable is the point at which the answer becomes a digest with a customer message — not a wider
cooldown.
## 10. Which repos gate 11 is registered in
| Repo | Registered | Note |
|---|---|---|
| `felhom.eu` | **yes** — gate 11 | its own `REPORT.md` is the gate's first real subject |
| `felhom-controller` | **yes** | already had `SHARED_*` constants; one constant + one `GATES` line |
| `felhom-agent` | **yes** | same; no observations section today, so it passes quietly |
| `app-catalog-felhom.eu` | **NO** | filed as **R-391** |
**Why not the catalog.** `catalog_gates.py` has no shared-gate mechanism at all: `run_gate` joins
every entry against its **own** `scripts/` directory, so it cannot invoke a sibling repo's script; and
the loop appends `--all` to every gate unconditionally, which the observations gate would read as a
path. Registering there needs `run_gate`'s contract widened **and** its argument handling changed — a
refactor of a runner whose shape is deliberately different, in a repo this task marked out of scope.
Exposure today is nil (that repo's `REPORT.md` has no observations section, and the gate passes
quietly on that), but a future session could write one. **Filed rather than left as a sentence in a
report, which is the exact failure this session exists to fix.**
## 11. Evidence
`documentation/audits/DRILL-cooldown-grain-2026-08-23/evidence/` — 17 files: 2 red-proof transcripts,
3 gate-11 control files covering 10 cases, 11 live-walk files, and a 1803-line controller-log window
**pulled off before the apps were restored**.
## 12. Teardown, three layers
1. **Guest 9201 / apps** — nothing provisioned. Five apps stopped across the walk (`opengist`,
`calibre-web`, `kimai`, `romm`, `paperless-ngx`); **all restarted and confirmed healthy**, 17
containers up. The three retained subjects (`docmost`, `bookstack`, `privatebin`) were not touched.
No app rebuilt, redeployed or restored.
2. **No VM, no bake** — this session built no golden and started no drill VM.
3. **Hub-side, stated explicitly.** The hub *was* written: deployed to v0.108.0 via the manifest, and
**six probe events POSTed for Scenarios B and C** (two `app_start_failed`, two `crossdrive_failed`,
plus yesterday's two, left in place deliberately). They are inert event rows for `demo-hp`, named
here rather than left to be found, and **every count in this report is `created_at`-filtered so
they cannot contaminate it**. Nothing else: no appliance registered, no customer created, no
artifact manifest changed, floor untouched.
## 13. Register size
| File | Before | After |
|---|---|---|
| `OPEN-ITEMS.md` | 328,325 B | **328,132 B** |
| `CLOSED-ITEMS.md` | 71,441 B | **74,642 B** |
| `OPEN-ITEMS.md` | 328,132 B | **331,024 B** |
| `CLOSED-ITEMS.md` | 74,642 B | **76,855 B** |
R-329 and R-386 compressed into CLOSED with their rules kept; **R-387** (closed) and **R-388** (the
notification-model product decision, open, operator's call) filed.
R-389 filed, then closed and compressed into `CLOSED-ITEMS.md` in the same session. **R-390** (the
golden-bake runbook's missing `pveam update`) and **R-391** (the catalog runner) filed open.
## 10. Observations
## 14. Observations
> **Markers added 2026-08-24 (gate 11, R-389).** Item 1 is the finding that had no row; adding its
> marker is the first thing the gate ever asked for. The observations' text is unchanged.
> **Gate 11's first real subject is this section.** Each item carries `FILED: R-NNN` naming a row
> opened this session, or `NOT-A-FINDING:` with its reason.
1. **The operator cooldown key carries no app identifier.** PrivateBin's alarm four minutes after
BookStack's was logged `suppressed — operator cooldown 1h, key=demo-hp:app_start_failed`, so **only
the first app-down per hour reaches the operator by e-mail**. R-182's known shape; harmless while
the event was undeliverable, and no longer. Not fixed here.
FILED: R-389
2. Two probe events remain as rows for `demo-hp` from Scenario H — inert, and named rather than left.
NOT-A-FINDING: two inert event rows on a Tier 0 demo box, created deliberately as a live
control and named in that session's teardown; they carry no state and nothing reads them.
1. **The gate's specification would have passed the item the gate exists to catch.** It said an
observation may "cite an `R-NNN` that resolves" — but yesterday's lost item cites `R-182`, which
resolves, as an **analogy** rather than as its own row. No parser can tell citation-as-precedent
from citation-as-filing by reading prose, so the marker is explicit instead. The discrepancy is
recorded in the gate's docstring and proven by `EDGE 7`.
NOT-A-FINDING: this is a design decision taken and documented inside the deliverable itself, not a
defect left behind — the gate ships with the stricter rule and its reasoning, so there is nothing
outstanding for a row to track.
2. **The burst has no ceiling.** Three apps in one scan produce three mails; the reference box's worst
case is 8, and it scales linearly with app count.
NOT-A-FINDING: measured and judged acceptable at today's scale in §9, with the reopening condition
stated there; filing a row for a digest would queue work the operator has not asked for and that
needs a customer message and their call on volume.
3. **ArgoCD reported "successfully rolled out" while still running the old image.** The first sync ran
against a pre-push revision; only a second hard refresh moved it. The rollout message alone was a
false confirmation and the image tag was the observable that settled it.
NOT-A-FINDING: the existing runbook already says to verify the image tag after a sync, and this run
followed it and caught the discrepancy — the procedure worked; recording the near-miss here is the
appropriate weight.
4. **The golden-bake runbook still omits `pveam update`**, and its failure names the wrong cause.
FILED: R-390
5. **Gate 11 is registered in three runners, not four** — `catalog_gates.py` cannot invoke a sibling
script and appends `--all` to every gate.
FILED: R-391
6. **Deliberately left open, untouched:** R-102, R-359, R-385, R-387, R-388's redesign.
NOT-A-FINDING: a pointer to rows that already exist, carried so their absence from this session
reads as deliberate rather than forgotten.