Files
felhom-controller/REPORT.md
T
admin 1da2c9c6c6
gates / gates (push) Successful in 11s
docs(v0.223.0): REPORT, CONTEXT rulings, README severity contract
REPORT overwritten: the 1.1 sweep in full (one bad severity, nine legitimate
"warn" strings that are healthcheck statuses), the hub manifest's real location
since the task's premise was wrong, all five red-proofs with the layer each
guard sits at, the live walk in six steps with the hub's own records quoted, and
the absent-intent count (0 of 8).

Three things are reported that a tidier account would omit: red-proof 5 passed
first time because the mutation was INERT; Scenario G was silently refused twice
behind an HTTP 200; and the live Scenario A does NOT prove the customer gate,
because demo-hp has no prefs row at all.

CONTEXT records the severity vocabulary as a ruling with its mechanism, the
intent ruling with its three-way handling of unknown, both fences, and two traps
worth more than the fixes: a 200 can be a refusal, and a passing red-proof can
mean an inert mutation.

README: the event table said `app_start_failed | warn` - the defect, written
down as if correct. Now `warning`, with the vocabulary contract and who receives
what. `disk_critical` also corrected from `error` to `critical`, which is what
fillwatch has always sent.
2026-08-23 12:06:45 +02:00

267 lines
14 KiB
Markdown

# REPORT — controller v0.223.0 (R-329, R-386, and the compound toggles)
**Session 2026-08-23, UNATTENDED.** Live leg on `demo-hp` (Tier 0), guest 9201.
**No halt condition fired.** Nothing was dropped.
## 1. Baselines, and the hub's four numbers as read
| Repo | at start | at end |
|---|---|---|
| felhom-controller | `14137efa` (v0.222.0) | **v0.223.0** deployed |
| felhom.eu | `55274d5e` (hub v0.106.0) | **hub v0.107.0** deployed |
| felhom-agent | `40d857b5` (v0.130.0) | untouched |
**Hub's four numbers, live from `GET /configuration` before starting:** `golden_version` **0.222.0**,
`agent_version` **0.130.0**, `min_agent` **0.129.0**, controller floor **0.222.0** — all four as the
task predicted.
## 2. Documents read
`internal/notify/notifier.go:583-602` (`Severity()`'s doc comment — **it already stated the entire
contract and named both hub locations**), `hub/internal/api/handler.go` (the type rejection and the
severity coercion side by side), `hub/internal/notify/dispatcher.go` (`severityNotifies`,
`ProcessEvent`, `operatorOnlyEvents`, `processOperator`), `internal/stacks/deploy.go` (the
`DesiredState` comment and its one-owner rule), `cmd/controller/main.go` (`classifyRunStates`),
`internal/web/handlers.go` (the compound toggles).
**§2's conditional is answered: the alarm ladder DOES exist** —
`felhom.eu/documentation/architecture/08-alarm-ladder.md`, written last session. It has been extended
here with §6.1 (the severity contract), §7 (the intent test) and §8 (Part 5's direction).
## 3. The 1.1 sweep — the result in full
**Exactly ONE bad severity in the whole controller: `notifier.go:546`, `"warn"`.** Nothing else.
Verified across Go **and** templates **and** queued-event construction, because the task warned that
reading zero from Go files while the answer sat in a template has produced three wrong conclusions
here:
- every `emit(...)` / `PushEvent(...)` literal — one offender, the rest valid;
- `internal/channelhealth`'s classifier (the source of `NotifyAgentChannelDown`'s variable) — all
`"warning"`/`"error"`;
- `debug.html`'s operator-triggerable severity `<select>` — offers only `error`/`warning`/`info`;
- no `PendingEvent{}` literal construction exists anywhere.
Nine other `"warn"` strings exist and are **not** defects: `internal/monitor` and `internal/selftest`
use it as a *healthcheck status* vocabulary, and `statusRank`'s `case "warn"` maps it to the correct
`"warning"` severity. **This is why the guard is an AST walk and not grep.**
**One latent hazard found while sweeping, pinned rather than left:** `fillwatch.Band.Severity()`
returns `""` for `BandOK`. Unreachable, because `Check()` notifies only on an escalation — but that
safety lives in a *different function* from the one that looks unsafe, so the test asserts the
consequence.
**And the guard found two dynamic call sites the hand sweep missed** (`NotifyDRCompleted`, and
fillwatch's via `main.go`). All six are now registered by name with the values each can take.
## 4. The hub's manifest — §6's premise was wrong
**`felhom.eu/manifests/hub.yaml`, line 128.** ArgoCD `Application/felhom` tracks
`admin/felhom.eu.git` path `manifests`, `automated.enabled=false`. Bumped in **`68a9f54`**. **No
out-of-git deployment path exists.** The image was pushed to the registry *before* the manifest landed,
so a sync could never point at a missing tag; the sync was then requested deliberately. Never
`kubectl set image`.
## 5. Files, commits, CI
| Commit | Repo | Contents |
|---|---|---|
| **`9832760`** | controller | v0.223.0 — the severity word, the AST guard, the toggle, the intent test, the toggle split |
| **`<docs>`** | controller | REPORT / CONTEXT / README |
| **`68a9f54`** | felhom.eu | hub v0.107.0 + manifest bump + golden evidence |
| **`2f7c9a6`** | felhom.eu | alarm ladder, register, STATUS, drill record |
Modified: `internal/notify/notifier.go`, `cmd/controller/main.go`, `internal/web/handlers.go`,
`internal/web/templates/settings_notifications.html`, `CHANGELOG.md`.
Added: `internal/notify/r329_severity_contract_test.go`, `internal/fillwatch/r329_severity_test.go`,
`cmd/controller/r386_intent_test.go`, `internal/web/r329_toggle_split_test.go`.
**CI runs confirmed BY ID** (`id` and `run_number` diverge — both printed) — see §15 below.
## 6. Red-proofs — five planted, and ONE PASSED FIRST TIME
| # | Mutation | Layer, and why that layer | Observed |
|---|---|---|---|
| 1 | severity back to `"warn"` | **the EMITTER** — the last point at which the bad value still exists; the hub deliberately destroys it one line later | `notifier.go:561:30: emit(...) emits severity "warn", which is NOT in the hub's vocabulary` |
| 2 | `userStopped` back to the state guess | **`classifyRunStates`** — the single derivation point where the guess was made | `dead-app banner = [], want exactly privatebin` — the exact live symptom — plus `IntentUnknown = false` and the wiring check |
| 3 | no-op-save guard removed | **the SAVE handler** — where a render-then-save rewrites stored bytes | `a no-op save CHANGED the stored settings` on the `defaults` shape |
| 4 | hub ingest `WARN` removed | **INGEST** — the last point the offending value exists | `the hub rewrote a severity and said nothing`, log showing only `[INFO] … (info)` |
| 5 | fillwatch de-escalation guard removed | **`Check()`** — the invariant lives there, not in `Severity()` | `notified with band ok → severity "" (event type "")` |
**Red-proof 5 passed on the first attempt and that is reported, not omitted.** The first mutation —
`if next <= prev` → `if next < prev` — is **inert**: an earlier `if next == prev { continue }` had
already removed the equal case, so the code's behaviour did not change and the test was *right* to
pass. Removing the guard outright convicts it. **A red-proof that passes needs the mutation checked
before either verdict is believed.** Every other mutation asserted its pre-fix text was present before
rewriting and printed `MUTATION APPLIED`.
## 7. Test counts
| Repo | Before | After |
|---|---|---|
| felhom-controller | 1504 | **1522** |
| felhom.eu hub | 702 | **709** |
Full green gate `go build && go vet && go test ./...` → **exit 0, zero failures, both repos**. All 11
controller design gates OK; all 11 felhom.eu repo gates OK.
## 8. Deployed versions and the golden
```
gitea.dooplex.hu/admin/felhom-controller:0.223.0 Up 18 minutes (healthy)
gitea.dooplex.hu/admin/felhom-hub:0.107.0 Synced, rollout complete
```
**Golden BAKED and PUBLISHED: YES — 0.223.0.**
`sha256 9eaf39ac39219b42ec9e6cbf890275febcdcc6f53325fe0c0f591d3431044f17`, `upload OK (HTTP 201)`,
round-trip **HTTP 206**, all five acceptance markers counted.
**VOUCHING IS THE OPERATOR'S ACT AND WAS NOT DONE HERE.**
## 9. The live walk, all six steps
### Step 1 — Scenario A: a database dies, customer has not opted in ✅
**Quoted from the hub's own records, not the controller's:**
```
events demo-hp app_start_failed warning 2026-08-23 09:27:51 <- v0.223.0
demo-hp app_start_failed info 2026-08-23 05:30:14 <- v0.222.0, coerced
notification_log demo-hp app_start_failed warning sent operator 2026-08-23 09:27:51
(no customer row)
```
**The number that says it all: 91 `app_start_failed` events stored all-time, ZERO `notification_log`
rows before 09:00 today.** Not one, ever, on any channel.
**An honest limit, stated rather than implied:** this run does **not** prove the customer *gate*.
`demo-hp` has no `customer_notifications` row at all, so the customer leg could not have delivered
regardless. The gate is proven by the unit tests, which configure prefs both ways — and Scenario B
there is the positive control showing the customer leg *can* deliver for this event type.
### Step 2 — Scenario D: stopped out of band, intent `running` ✅
`docker compose stop privatebin` at 09:31:27Z → **`app_start_failed (warning)` at 09:31:51Z, 24
seconds later.** Heartbeat, new beside last session's:
```
2026/08/23 05:47:44 [deadapp] check alive: 40 scans since boot, 8 deployed app(s) evaluated, 0 currently down <- v0.222.0
2026/08/23 09:34:51 [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down <- v0.223.0
```
Hub-side the second alarm was logged `suppressed — operator cooldown 1h` (see §15.1).
### Step 3 — Scenario C: the customer presses Stop ✅
Through `POST /api/stacks/privatebin/stop`, the exact call the button makes. Intent moved
`running → stopped`. Four minutes and **9 dead-app scans later: 0 alarms, 0 unknown-intent lines.**
### Step 4 — Scenario E: intent absent ✅ (and the count)
demo-hp has **zero** such apps, so one was created: `desired_state` removed from `privatebin`'s
`app.yaml`, as a pre-R-166 box would look (backed up, restored, no code writer added). Result —
evaluated (`deployed=True state=stopped`, 8 apps evaluated), **suppressed (0 alarms)**, and:
```
[deadapp] 1 stopped app(s) have NO recorded customer intent, so their dead-app alarm is suppressed
by the unknown-intent fallback (R-386): privatebin. This closes itself as each app is started or
stopped through the interface.
```
### Step 5 — Scenario G: save changing nothing ✅ byte-identical, **on the third attempt**
`sha256(enabled_events)` = `10840f3a95bac168f0d7c79760998138` **before and after** a real save
(7 boxes ticked, refusal banner absent).
**The first two attempts were silently REFUSED behind an HTTP 200** — the empty-email wipe guard
declines and renders an error page, still `200`. Run 1: no prefs existed. Run 2: the email `<input>`
spans **three lines**, so a single-line grep read it as `""`. **Both times the hashes matched —
because nothing was saved, not because nothing changed.** Fixed by asserting the refusal banner is
absent. *A warning beside a success is read as a success.*
### Step 6 — Scenario H: a bad severity to the hub ✅
```
[WARN] [api] Event from demo-hp: severity "warn" is not in {info,warning,error,critical} — coercing
to "info", which severityNotifies DROPS, so this backup_failed alert will reach NOBODY. Fix the
emitting controller; this event is stored but not routed.
```
Both the bad POST and an `error` control returned **200** (nothing lost); the control produced **no**
warning.
## 10. The absent-intent count on `demo-hp`, in plain words
**Zero.** All **8** deployed apps carry `desired_state: running`; none is `stopped` and none is absent.
The unknown-intent fallback therefore suppresses nothing on this box today — the population is already
empty on a machine that has been exercised through the interface. It will be larger on a box upgraded
and left alone, which is why the log line exists rather than a one-off count.
## 11. The dead-branch decision, and the reason
**KEPT.** `cmd/hub/main.go` wires `dispatcher.ProcessEvent` **directly** as the
`monitor.EventNotifyFunc` for the staleness, host-staleness and offsite-box checkers — those events
never pass the ingest handler, so for them that line is the only severity guard there is. Deleting it
as "dead" would have removed the live half while the dead half supplied the justification. All 90
severity literals in `internal/monitor` were verified already valid, so the guard is silent because
the producers are correct.
## 12. Evidence
`felhom.eu/documentation/audits/DRILL-r329-r386-2026-08-23/evidence/` — 26 files: 5 red-proof
transcripts, 20 live-walk files, two full controller-log windows (1041 and 4044 lines) **pulled off
before each revert**, and the hub's own DB queries.
## 13. Teardown, three layers, and the end state
1. **Guest 9201 / apps** — nothing provisioned. **All 17 app containers healthy.** `privatebin`'s
`app.yaml` restored from backup and the backup deleted; intent reads `running`. `demo-hp`'s
notification settings restored to `enabled_events: null`, no e-mail — their pre-drill state. No app
rebuilt, redeployed or restored; planted data untouched.
2. **Bake VM** — powered off, Gitea token and runner script **shredded**, `drill.qcow2` reverted to
`virgin`. Build guest 9100 exists only inside that reverted snapshot. No storage added anywhere, so
`pvesm status` has nothing to compare.
3. **Hub-side, stated explicitly.** The hub was **written** this session, unlike last: the deployment
is v0.107.0 via the manifest, and **two probe events remain as rows for `demo-hp`** from Scenario H
(`backup_failed`, "R-387 scenario H probe" and "…control"). They are inert records; named here
rather than left for someone to find. Nothing else: no appliance registered, no customer created,
no artifact manifest changed, floor untouched.
**End state:** controller **0.223.0** and hub **0.107.0** deployed; golden **0.223.0** baked and
published but **NOT vouched**; floor still **0.222.0**; all apps running; planted data present.
## 14. Register size
| File | Before | After |
|---|---|---|
| `OPEN-ITEMS.md` | 328,325 B | **328,132 B** |
| `CLOSED-ITEMS.md` | 71,441 B | **74,642 B** |
R-329 and R-386 closed and compressed; **R-387** (closed) and **R-388** (the notification-model
product decision — open, operator's call) filed.
## 15. Observations — noticed, documented, NOT acted on
1. **The operator cooldown key has no app identifier, and it now bites.** PrivateBin's alarm four
minutes after BookStack's was logged `suppressed — operator cooldown 1h,
key=demo-hp:app_start_failed`, so **only the first app-down per hour e-mails the operator**. This is
R-182's known cooldown-key shape; it was harmless while `app_start_failed` was undeliverable and is
not any more. **Same pattern as R-329 itself: a known-broken thing moved from unreachable to
load-bearing.** Not fixed here.
2. **The settings page grew 12 → 15 toggles in one session** — one new alarm, plus two compound
toggles split into four. Recorded as the argument inside R-388.
3. `internal/notify/notifier.go` carries **pre-existing** gofmt drift in an unrelated const block,
confirmed by stashing this session's work and re-running `gofmt -l`. Not touched (§12).
4. **The golden-bake runbook still lacks `pveam update`** — second consecutive bake to hit the stale
index on the `virgin` snapshot, presenting as `400 … no such template`.
5. **Deliberately left open, untouched:** R-102, R-359, R-385, and R-388's redesign.
### CI runs, confirmed by ID
| Commit | Repo | CI `id` | `run_number` | Result |
|---|---|---|---|---|
| `9832760` | felhom-controller | **408** | 88 | success |
| `68a9f54` | felhom.eu | **409** | 261 | success |
| `2f7c9a6` | felhom.eu | **410** | 262 | success |
(The controller docs commit's run is confirmed after its push and is the next `id` in that repo.)