# REPORT — controller v0.223.0 (R-329, R-386, and the compound toggles) **Session 2026-08-23, UNATTENDED.** Live leg on `demo-hp` (Tier 0), guest 9201. **No halt condition fired.** Nothing was dropped. ## 1. Baselines, and the hub's four numbers as read | Repo | at start | at end | |---|---|---| | felhom-controller | `14137efa` (v0.222.0) | **v0.223.0** deployed | | felhom.eu | `55274d5e` (hub v0.106.0) | **hub v0.107.0** deployed | | felhom-agent | `40d857b5` (v0.130.0) | untouched | **Hub's four numbers, live from `GET /configuration` before starting:** `golden_version` **0.222.0**, `agent_version` **0.130.0**, `min_agent` **0.129.0**, controller floor **0.222.0** — all four as the task predicted. ## 2. Documents read `internal/notify/notifier.go:583-602` (`Severity()`'s doc comment — **it already stated the entire contract and named both hub locations**), `hub/internal/api/handler.go` (the type rejection and the severity coercion side by side), `hub/internal/notify/dispatcher.go` (`severityNotifies`, `ProcessEvent`, `operatorOnlyEvents`, `processOperator`), `internal/stacks/deploy.go` (the `DesiredState` comment and its one-owner rule), `cmd/controller/main.go` (`classifyRunStates`), `internal/web/handlers.go` (the compound toggles). **§2's conditional is answered: the alarm ladder DOES exist** — `felhom.eu/documentation/architecture/08-alarm-ladder.md`, written last session. It has been extended here with §6.1 (the severity contract), §7 (the intent test) and §8 (Part 5's direction). ## 3. The 1.1 sweep — the result in full **Exactly ONE bad severity in the whole controller: `notifier.go:546`, `"warn"`.** Nothing else. Verified across Go **and** templates **and** queued-event construction, because the task warned that reading zero from Go files while the answer sat in a template has produced three wrong conclusions here: - every `emit(...)` / `PushEvent(...)` literal — one offender, the rest valid; - `internal/channelhealth`'s classifier (the source of `NotifyAgentChannelDown`'s variable) — all `"warning"`/`"error"`; - `debug.html`'s operator-triggerable severity `` spans **three lines**, so a single-line grep read it as `""`. **Both times the hashes matched — because nothing was saved, not because nothing changed.** Fixed by asserting the refusal banner is absent. *A warning beside a success is read as a success.* ### Step 6 — Scenario H: a bad severity to the hub ✅ ``` [WARN] [api] Event from demo-hp: severity "warn" is not in {info,warning,error,critical} — coercing to "info", which severityNotifies DROPS, so this backup_failed alert will reach NOBODY. Fix the emitting controller; this event is stored but not routed. ``` Both the bad POST and an `error` control returned **200** (nothing lost); the control produced **no** warning. ## 10. The absent-intent count on `demo-hp`, in plain words **Zero.** All **8** deployed apps carry `desired_state: running`; none is `stopped` and none is absent. The unknown-intent fallback therefore suppresses nothing on this box today — the population is already empty on a machine that has been exercised through the interface. It will be larger on a box upgraded and left alone, which is why the log line exists rather than a one-off count. ## 11. The dead-branch decision, and the reason **KEPT.** `cmd/hub/main.go` wires `dispatcher.ProcessEvent` **directly** as the `monitor.EventNotifyFunc` for the staleness, host-staleness and offsite-box checkers — those events never pass the ingest handler, so for them that line is the only severity guard there is. Deleting it as "dead" would have removed the live half while the dead half supplied the justification. All 90 severity literals in `internal/monitor` were verified already valid, so the guard is silent because the producers are correct. ## 12. Evidence `felhom.eu/documentation/audits/DRILL-r329-r386-2026-08-23/evidence/` — 26 files: 5 red-proof transcripts, 20 live-walk files, two full controller-log windows (1041 and 4044 lines) **pulled off before each revert**, and the hub's own DB queries. ## 13. Teardown, three layers, and the end state 1. **Guest 9201 / apps** — nothing provisioned. **All 17 app containers healthy.** `privatebin`'s `app.yaml` restored from backup and the backup deleted; intent reads `running`. `demo-hp`'s notification settings restored to `enabled_events: null`, no e-mail — their pre-drill state. No app rebuilt, redeployed or restored; planted data untouched. 2. **Bake VM** — powered off, Gitea token and runner script **shredded**, `drill.qcow2` reverted to `virgin`. Build guest 9100 exists only inside that reverted snapshot. No storage added anywhere, so `pvesm status` has nothing to compare. 3. **Hub-side, stated explicitly.** The hub was **written** this session, unlike last: the deployment is v0.107.0 via the manifest, and **two probe events remain as rows for `demo-hp`** from Scenario H (`backup_failed`, "R-387 scenario H probe" and "…control"). They are inert records; named here rather than left for someone to find. Nothing else: no appliance registered, no customer created, no artifact manifest changed, floor untouched. **End state:** controller **0.223.0** and hub **0.107.0** deployed; golden **0.223.0** baked and published but **NOT vouched**; floor still **0.222.0**; all apps running; planted data present. ## 14. Register size | File | Before | After | |---|---|---| | `OPEN-ITEMS.md` | 328,325 B | **328,132 B** | | `CLOSED-ITEMS.md` | 71,441 B | **74,642 B** | R-329 and R-386 closed and compressed; **R-387** (closed) and **R-388** (the notification-model product decision — open, operator's call) filed. ## 15. Observations — noticed, documented, NOT acted on 1. **The operator cooldown key has no app identifier, and it now bites.** PrivateBin's alarm four minutes after BookStack's was logged `suppressed — operator cooldown 1h, key=demo-hp:app_start_failed`, so **only the first app-down per hour e-mails the operator**. This is R-182's known cooldown-key shape; it was harmless while `app_start_failed` was undeliverable and is not any more. **Same pattern as R-329 itself: a known-broken thing moved from unreachable to load-bearing.** Not fixed here. 2. **The settings page grew 12 → 15 toggles in one session** — one new alarm, plus two compound toggles split into four. Recorded as the argument inside R-388. 3. `internal/notify/notifier.go` carries **pre-existing** gofmt drift in an unrelated const block, confirmed by stashing this session's work and re-running `gofmt -l`. Not touched (§12). 4. **The golden-bake runbook still lacks `pveam update`** — second consecutive bake to hit the stale index on the `virgin` snapshot, presenting as `400 … no such template`. 5. **Deliberately left open, untouched:** R-102, R-359, R-385, and R-388's redesign. ### CI runs, confirmed by ID | Commit | Repo | CI `id` | `run_number` | Result | |---|---|---|---|---| | `9832760` | felhom-controller | **408** | 88 | success | | `68a9f54` | felhom.eu | **409** | 261 | success | | `2f7c9a6` | felhom.eu | **410** | 262 | success | (The controller docs commit's run is confirmed after its push and is the next `id` in that repo.)