Files
felhom-controller/REPORT.md
T
admin f8c9390946
gates / gates (push) Successful in 13s
gate 11: register the observations gate, and mark up this repo's observations
The shared gate lives in felhom.eu/scripts/observations_gate.py and is invoked
across the workspace, exactly as reuse_refs_check.py and instructions_gate.py
already are. It is never copied.

REPORT.md's observations now carry their markers. Item 1 was the finding that
had no register row - only the first broken app per hour reaches the operator -
and it is now R-389. Item 4, the golden-bake runbook's missing `pveam update`,
is R-390. Items 3 and 5 are declared NOT-A-FINDING with their reasons. The
observations' text itself is unchanged; only the markers were added.
2026-08-23 13:53:18 +02:00

279 lines
15 KiB
Markdown

# REPORT — controller v0.223.0 (R-329, R-386, and the compound toggles)
**Session 2026-08-23, UNATTENDED.** Live leg on `demo-hp` (Tier 0), guest 9201.
**No halt condition fired.** Nothing was dropped.
## 1. Baselines, and the hub's four numbers as read
| Repo | at start | at end |
|---|---|---|
| felhom-controller | `14137efa` (v0.222.0) | **v0.223.0** deployed |
| felhom.eu | `55274d5e` (hub v0.106.0) | **hub v0.107.0** deployed |
| felhom-agent | `40d857b5` (v0.130.0) | untouched |
**Hub's four numbers, live from `GET /configuration` before starting:** `golden_version` **0.222.0**,
`agent_version` **0.130.0**, `min_agent` **0.129.0**, controller floor **0.222.0** — all four as the
task predicted.
## 2. Documents read
`internal/notify/notifier.go:583-602` (`Severity()`'s doc comment — **it already stated the entire
contract and named both hub locations**), `hub/internal/api/handler.go` (the type rejection and the
severity coercion side by side), `hub/internal/notify/dispatcher.go` (`severityNotifies`,
`ProcessEvent`, `operatorOnlyEvents`, `processOperator`), `internal/stacks/deploy.go` (the
`DesiredState` comment and its one-owner rule), `cmd/controller/main.go` (`classifyRunStates`),
`internal/web/handlers.go` (the compound toggles).
**§2's conditional is answered: the alarm ladder DOES exist** —
`felhom.eu/documentation/architecture/08-alarm-ladder.md`, written last session. It has been extended
here with §6.1 (the severity contract), §7 (the intent test) and §8 (Part 5's direction).
## 3. The 1.1 sweep — the result in full
**Exactly ONE bad severity in the whole controller: `notifier.go:546`, `"warn"`.** Nothing else.
Verified across Go **and** templates **and** queued-event construction, because the task warned that
reading zero from Go files while the answer sat in a template has produced three wrong conclusions
here:
- every `emit(...)` / `PushEvent(...)` literal — one offender, the rest valid;
- `internal/channelhealth`'s classifier (the source of `NotifyAgentChannelDown`'s variable) — all
`"warning"`/`"error"`;
- `debug.html`'s operator-triggerable severity `<select>` — offers only `error`/`warning`/`info`;
- no `PendingEvent{}` literal construction exists anywhere.
Nine other `"warn"` strings exist and are **not** defects: `internal/monitor` and `internal/selftest`
use it as a *healthcheck status* vocabulary, and `statusRank`'s `case "warn"` maps it to the correct
`"warning"` severity. **This is why the guard is an AST walk and not grep.**
**One latent hazard found while sweeping, pinned rather than left:** `fillwatch.Band.Severity()`
returns `""` for `BandOK`. Unreachable, because `Check()` notifies only on an escalation — but that
safety lives in a *different function* from the one that looks unsafe, so the test asserts the
consequence.
**And the guard found two dynamic call sites the hand sweep missed** (`NotifyDRCompleted`, and
fillwatch's via `main.go`). All six are now registered by name with the values each can take.
## 4. The hub's manifest — §6's premise was wrong
**`felhom.eu/manifests/hub.yaml`, line 128.** ArgoCD `Application/felhom` tracks
`admin/felhom.eu.git` path `manifests`, `automated.enabled=false`. Bumped in **`68a9f54`**. **No
out-of-git deployment path exists.** The image was pushed to the registry *before* the manifest landed,
so a sync could never point at a missing tag; the sync was then requested deliberately. Never
`kubectl set image`.
## 5. Files, commits, CI
| Commit | Repo | Contents |
|---|---|---|
| **`9832760`** | controller | v0.223.0 — the severity word, the AST guard, the toggle, the intent test, the toggle split |
| **`<docs>`** | controller | REPORT / CONTEXT / README |
| **`68a9f54`** | felhom.eu | hub v0.107.0 + manifest bump + golden evidence |
| **`2f7c9a6`** | felhom.eu | alarm ladder, register, STATUS, drill record |
Modified: `internal/notify/notifier.go`, `cmd/controller/main.go`, `internal/web/handlers.go`,
`internal/web/templates/settings_notifications.html`, `CHANGELOG.md`.
Added: `internal/notify/r329_severity_contract_test.go`, `internal/fillwatch/r329_severity_test.go`,
`cmd/controller/r386_intent_test.go`, `internal/web/r329_toggle_split_test.go`.
**CI runs confirmed BY ID** (`id` and `run_number` diverge — both printed) — see §15 below.
## 6. Red-proofs — five planted, and ONE PASSED FIRST TIME
| # | Mutation | Layer, and why that layer | Observed |
|---|---|---|---|
| 1 | severity back to `"warn"` | **the EMITTER** — the last point at which the bad value still exists; the hub deliberately destroys it one line later | `notifier.go:561:30: emit(...) emits severity "warn", which is NOT in the hub's vocabulary` |
| 2 | `userStopped` back to the state guess | **`classifyRunStates`** — the single derivation point where the guess was made | `dead-app banner = [], want exactly privatebin` — the exact live symptom — plus `IntentUnknown = false` and the wiring check |
| 3 | no-op-save guard removed | **the SAVE handler** — where a render-then-save rewrites stored bytes | `a no-op save CHANGED the stored settings` on the `defaults` shape |
| 4 | hub ingest `WARN` removed | **INGEST** — the last point the offending value exists | `the hub rewrote a severity and said nothing`, log showing only `[INFO] … (info)` |
| 5 | fillwatch de-escalation guard removed | **`Check()`** — the invariant lives there, not in `Severity()` | `notified with band ok → severity "" (event type "")` |
**Red-proof 5 passed on the first attempt and that is reported, not omitted.** The first mutation —
`if next <= prev` → `if next < prev` — is **inert**: an earlier `if next == prev { continue }` had
already removed the equal case, so the code's behaviour did not change and the test was *right* to
pass. Removing the guard outright convicts it. **A red-proof that passes needs the mutation checked
before either verdict is believed.** Every other mutation asserted its pre-fix text was present before
rewriting and printed `MUTATION APPLIED`.
## 7. Test counts
| Repo | Before | After |
|---|---|---|
| felhom-controller | 1504 | **1522** |
| felhom.eu hub | 702 | **709** |
Full green gate `go build && go vet && go test ./...` → **exit 0, zero failures, both repos**. All 11
controller design gates OK; all 11 felhom.eu repo gates OK.
## 8. Deployed versions and the golden
```
gitea.dooplex.hu/admin/felhom-controller:0.223.0 Up 18 minutes (healthy)
gitea.dooplex.hu/admin/felhom-hub:0.107.0 Synced, rollout complete
```
**Golden BAKED and PUBLISHED: YES — 0.223.0.**
`sha256 9eaf39ac39219b42ec9e6cbf890275febcdcc6f53325fe0c0f591d3431044f17`, `upload OK (HTTP 201)`,
round-trip **HTTP 206**, all five acceptance markers counted.
**VOUCHING IS THE OPERATOR'S ACT AND WAS NOT DONE HERE.**
## 9. The live walk, all six steps
### Step 1 — Scenario A: a database dies, customer has not opted in ✅
**Quoted from the hub's own records, not the controller's:**
```
events demo-hp app_start_failed warning 2026-08-23 09:27:51 <- v0.223.0
demo-hp app_start_failed info 2026-08-23 05:30:14 <- v0.222.0, coerced
notification_log demo-hp app_start_failed warning sent operator 2026-08-23 09:27:51
(no customer row)
```
**The number that says it all: 91 `app_start_failed` events stored all-time, ZERO `notification_log`
rows before 09:00 today.** Not one, ever, on any channel.
**An honest limit, stated rather than implied:** this run does **not** prove the customer *gate*.
`demo-hp` has no `customer_notifications` row at all, so the customer leg could not have delivered
regardless. The gate is proven by the unit tests, which configure prefs both ways — and Scenario B
there is the positive control showing the customer leg *can* deliver for this event type.
### Step 2 — Scenario D: stopped out of band, intent `running` ✅
`docker compose stop privatebin` at 09:31:27Z → **`app_start_failed (warning)` at 09:31:51Z, 24
seconds later.** Heartbeat, new beside last session's:
```
2026/08/23 05:47:44 [deadapp] check alive: 40 scans since boot, 8 deployed app(s) evaluated, 0 currently down <- v0.222.0
2026/08/23 09:34:51 [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down <- v0.223.0
```
Hub-side the second alarm was logged `suppressed — operator cooldown 1h` (see §15.1).
### Step 3 — Scenario C: the customer presses Stop ✅
Through `POST /api/stacks/privatebin/stop`, the exact call the button makes. Intent moved
`running → stopped`. Four minutes and **9 dead-app scans later: 0 alarms, 0 unknown-intent lines.**
### Step 4 — Scenario E: intent absent ✅ (and the count)
demo-hp has **zero** such apps, so one was created: `desired_state` removed from `privatebin`'s
`app.yaml`, as a pre-R-166 box would look (backed up, restored, no code writer added). Result —
evaluated (`deployed=True state=stopped`, 8 apps evaluated), **suppressed (0 alarms)**, and:
```
[deadapp] 1 stopped app(s) have NO recorded customer intent, so their dead-app alarm is suppressed
by the unknown-intent fallback (R-386): privatebin. This closes itself as each app is started or
stopped through the interface.
```
### Step 5 — Scenario G: save changing nothing ✅ byte-identical, **on the third attempt**
`sha256(enabled_events)` = `10840f3a95bac168f0d7c79760998138` **before and after** a real save
(7 boxes ticked, refusal banner absent).
**The first two attempts were silently REFUSED behind an HTTP 200** — the empty-email wipe guard
declines and renders an error page, still `200`. Run 1: no prefs existed. Run 2: the email `<input>`
spans **three lines**, so a single-line grep read it as `""`. **Both times the hashes matched —
because nothing was saved, not because nothing changed.** Fixed by asserting the refusal banner is
absent. *A warning beside a success is read as a success.*
### Step 6 — Scenario H: a bad severity to the hub ✅
```
[WARN] [api] Event from demo-hp: severity "warn" is not in {info,warning,error,critical} — coercing
to "info", which severityNotifies DROPS, so this backup_failed alert will reach NOBODY. Fix the
emitting controller; this event is stored but not routed.
```
Both the bad POST and an `error` control returned **200** (nothing lost); the control produced **no**
warning.
## 10. The absent-intent count on `demo-hp`, in plain words
**Zero.** All **8** deployed apps carry `desired_state: running`; none is `stopped` and none is absent.
The unknown-intent fallback therefore suppresses nothing on this box today — the population is already
empty on a machine that has been exercised through the interface. It will be larger on a box upgraded
and left alone, which is why the log line exists rather than a one-off count.
## 11. The dead-branch decision, and the reason
**KEPT.** `cmd/hub/main.go` wires `dispatcher.ProcessEvent` **directly** as the
`monitor.EventNotifyFunc` for the staleness, host-staleness and offsite-box checkers — those events
never pass the ingest handler, so for them that line is the only severity guard there is. Deleting it
as "dead" would have removed the live half while the dead half supplied the justification. All 90
severity literals in `internal/monitor` were verified already valid, so the guard is silent because
the producers are correct.
## 12. Evidence
`felhom.eu/documentation/audits/DRILL-r329-r386-2026-08-23/evidence/` — 26 files: 5 red-proof
transcripts, 20 live-walk files, two full controller-log windows (1041 and 4044 lines) **pulled off
before each revert**, and the hub's own DB queries.
## 13. Teardown, three layers, and the end state
1. **Guest 9201 / apps** — nothing provisioned. **All 17 app containers healthy.** `privatebin`'s
`app.yaml` restored from backup and the backup deleted; intent reads `running`. `demo-hp`'s
notification settings restored to `enabled_events: null`, no e-mail — their pre-drill state. No app
rebuilt, redeployed or restored; planted data untouched.
2. **Bake VM** — powered off, Gitea token and runner script **shredded**, `drill.qcow2` reverted to
`virgin`. Build guest 9100 exists only inside that reverted snapshot. No storage added anywhere, so
`pvesm status` has nothing to compare.
3. **Hub-side, stated explicitly.** The hub was **written** this session, unlike last: the deployment
is v0.107.0 via the manifest, and **two probe events remain as rows for `demo-hp`** from Scenario H
(`backup_failed`, "R-387 scenario H probe" and "…control"). They are inert records; named here
rather than left for someone to find. Nothing else: no appliance registered, no customer created,
no artifact manifest changed, floor untouched.
**End state:** controller **0.223.0** and hub **0.107.0** deployed; golden **0.223.0** baked and
published but **NOT vouched**; floor still **0.222.0**; all apps running; planted data present.
## 14. Register size
| File | Before | After |
|---|---|---|
| `OPEN-ITEMS.md` | 328,325 B | **328,132 B** |
| `CLOSED-ITEMS.md` | 71,441 B | **74,642 B** |
R-329 and R-386 closed and compressed; **R-387** (closed) and **R-388** (the notification-model
product decision — open, operator's call) filed.
## 15. Observations
> **Each item carries `FILED: R-NNN` or `NOT-A-FINDING: <reason>` (gate 11, added 2026-08-23).**
> The markers were added retrospectively on 2026-08-24 when the gate was built: item 1 was the very
> finding that had no row, and adding its marker is the first thing the gate ever asked for. The
> observations' text is unchanged.
1. **The operator cooldown key has no app identifier, and it now bites.** PrivateBin's alarm four
minutes after BookStack's was logged `suppressed — operator cooldown 1h,
key=demo-hp:app_start_failed`, so **only the first app-down per hour e-mails the operator**. This is
R-182's known cooldown-key shape; it was harmless while `app_start_failed` was undeliverable and is
not any more. **Same pattern as R-329 itself: a known-broken thing moved from unreachable to
load-bearing.** Not fixed here.
FILED: R-389
2. **The settings page grew 12 → 15 toggles in one session** — one new alarm, plus two compound
toggles split into four. Recorded as the argument inside R-388.
FILED: R-388
3. `internal/notify/notifier.go` carries **pre-existing** gofmt drift in an unrelated const block,
confirmed by stashing this session's work and re-running `gofmt -l`. Not touched (§12).
NOT-A-FINDING: cosmetic drift in an unrelated const block, predating this work and touching no
behaviour; §12 forbids nearby refactors, so filing a row would only queue a whitespace commit.
4. **The golden-bake runbook still lacks `pveam update`** — second consecutive bake to hit the stale
index on the `virgin` snapshot, presenting as `400 … no such template`.
FILED: R-390
5. **Deliberately left open, untouched:** R-102, R-359, R-385, and R-388's redesign.
NOT-A-FINDING: this item is a pointer to rows that already exist, not a new observation; it is
here so their absence from this session reads as deliberate rather than forgotten.
### CI runs, confirmed by ID
| Commit | Repo | CI `id` | `run_number` | Result |
|---|---|---|---|---|
| `9832760` | felhom-controller | **408** | 88 | success |
| `68a9f54` | felhom.eu | **409** | 261 | success |
| `2f7c9a6` | felhom.eu | **410** | 262 | success |
(The controller docs commit's run is confirmed after its push and is the next `id` in that repo.)