Files
felhom-controller/REPORT.md
T
admin f8c9390946
gates / gates (push) Successful in 13s
gate 11: register the observations gate, and mark up this repo's observations
The shared gate lives in felhom.eu/scripts/observations_gate.py and is invoked
across the workspace, exactly as reuse_refs_check.py and instructions_gate.py
already are. It is never copied.

REPORT.md's observations now carry their markers. Item 1 was the finding that
had no register row - only the first broken app per hour reaches the operator -
and it is now R-389. Item 4, the golden-bake runbook's missing `pveam update`,
is R-390. Items 3 and 5 are declared NOT-A-FINDING with their reasons. The
observations' text itself is unchanged; only the markers were added.
2026-08-23 13:53:18 +02:00

15 KiB

REPORT — controller v0.223.0 (R-329, R-386, and the compound toggles)

Session 2026-08-23, UNATTENDED. Live leg on demo-hp (Tier 0), guest 9201. No halt condition fired. Nothing was dropped.

1. Baselines, and the hub's four numbers as read

Repo at start at end
felhom-controller 14137efa (v0.222.0) v0.223.0 deployed
felhom.eu 55274d5e (hub v0.106.0) hub v0.107.0 deployed
felhom-agent 40d857b5 (v0.130.0) untouched

Hub's four numbers, live from GET /configuration before starting: golden_version 0.222.0, agent_version 0.130.0, min_agent 0.129.0, controller floor 0.222.0 — all four as the task predicted.

2. Documents read

internal/notify/notifier.go:583-602 (Severity()'s doc comment — it already stated the entire contract and named both hub locations), hub/internal/api/handler.go (the type rejection and the severity coercion side by side), hub/internal/notify/dispatcher.go (severityNotifies, ProcessEvent, operatorOnlyEvents, processOperator), internal/stacks/deploy.go (the DesiredState comment and its one-owner rule), cmd/controller/main.go (classifyRunStates), internal/web/handlers.go (the compound toggles).

§2's conditional is answered: the alarm ladder DOES exist — felhom.eu/documentation/architecture/08-alarm-ladder.md, written last session. It has been extended here with §6.1 (the severity contract), §7 (the intent test) and §8 (Part 5's direction).

3. The 1.1 sweep — the result in full

Exactly ONE bad severity in the whole controller: notifier.go:546, "warn". Nothing else.

Verified across Go and templates and queued-event construction, because the task warned that reading zero from Go files while the answer sat in a template has produced three wrong conclusions here:

  • every emit(...) / PushEvent(...) literal — one offender, the rest valid;
  • internal/channelhealth's classifier (the source of NotifyAgentChannelDown's variable) — all "warning"/"error";
  • debug.html's operator-triggerable severity <select> — offers only error/warning/info;
  • no PendingEvent{} literal construction exists anywhere.

Nine other "warn" strings exist and are not defects: internal/monitor and internal/selftest use it as a healthcheck status vocabulary, and statusRank's case "warn" maps it to the correct "warning" severity. This is why the guard is an AST walk and not grep.

One latent hazard found while sweeping, pinned rather than left: fillwatch.Band.Severity() returns "" for BandOK. Unreachable, because Check() notifies only on an escalation — but that safety lives in a different function from the one that looks unsafe, so the test asserts the consequence.

And the guard found two dynamic call sites the hand sweep missed (NotifyDRCompleted, and fillwatch's via main.go). All six are now registered by name with the values each can take.

4. The hub's manifest — §6's premise was wrong

felhom.eu/manifests/hub.yaml, line 128. ArgoCD Application/felhom tracks admin/felhom.eu.git path manifests, automated.enabled=false. Bumped in 68a9f54. No out-of-git deployment path exists. The image was pushed to the registry before the manifest landed, so a sync could never point at a missing tag; the sync was then requested deliberately. Never kubectl set image.

5. Files, commits, CI

Commit Repo Contents
9832760 controller v0.223.0 — the severity word, the AST guard, the toggle, the intent test, the toggle split
<docs> controller REPORT / CONTEXT / README
68a9f54 felhom.eu hub v0.107.0 + manifest bump + golden evidence
2f7c9a6 felhom.eu alarm ladder, register, STATUS, drill record

Modified: internal/notify/notifier.go, cmd/controller/main.go, internal/web/handlers.go, internal/web/templates/settings_notifications.html, CHANGELOG.md. Added: internal/notify/r329_severity_contract_test.go, internal/fillwatch/r329_severity_test.go, cmd/controller/r386_intent_test.go, internal/web/r329_toggle_split_test.go.

CI runs confirmed BY ID (id and run_number diverge — both printed) — see §15 below.

6. Red-proofs — five planted, and ONE PASSED FIRST TIME

# Mutation Layer, and why that layer Observed
1 severity back to "warn" the EMITTER — the last point at which the bad value still exists; the hub deliberately destroys it one line later notifier.go:561:30: emit(...) emits severity "warn", which is NOT in the hub's vocabulary
2 userStopped back to the state guess classifyRunStates — the single derivation point where the guess was made dead-app banner = [], want exactly privatebin — the exact live symptom — plus IntentUnknown = false and the wiring check
3 no-op-save guard removed the SAVE handler — where a render-then-save rewrites stored bytes a no-op save CHANGED the stored settings on the defaults shape
4 hub ingest WARN removed INGEST — the last point the offending value exists the hub rewrote a severity and said nothing, log showing only [INFO] … (info)
5 fillwatch de-escalation guard removed Check() — the invariant lives there, not in Severity() notified with band ok → severity "" (event type "")

Red-proof 5 passed on the first attempt and that is reported, not omitted. The first mutation — if next <= prev → if next < prev — is inert: an earlier if next == prev { continue } had already removed the equal case, so the code's behaviour did not change and the test was right to pass. Removing the guard outright convicts it. A red-proof that passes needs the mutation checked before either verdict is believed. Every other mutation asserted its pre-fix text was present before rewriting and printed MUTATION APPLIED.

7. Test counts

Repo Before After
felhom-controller 1504 1522
felhom.eu hub 702 709

Full green gate go build && go vet && go test ./... → exit 0, zero failures, both repos. All 11 controller design gates OK; all 11 felhom.eu repo gates OK.

8. Deployed versions and the golden

gitea.dooplex.hu/admin/felhom-controller:0.223.0   Up 18 minutes (healthy)
gitea.dooplex.hu/admin/felhom-hub:0.107.0          Synced, rollout complete

Golden BAKED and PUBLISHED: YES — 0.223.0. sha256 9eaf39ac39219b42ec9e6cbf890275febcdcc6f53325fe0c0f591d3431044f17, upload OK (HTTP 201), round-trip HTTP 206, all five acceptance markers counted. VOUCHING IS THE OPERATOR'S ACT AND WAS NOT DONE HERE.

9. The live walk, all six steps

Step 1 — Scenario A: a database dies, customer has not opted in ✅

Quoted from the hub's own records, not the controller's:

events            demo-hp  app_start_failed  warning  2026-08-23 09:27:51   <- v0.223.0
                  demo-hp  app_start_failed  info     2026-08-23 05:30:14   <- v0.222.0, coerced
notification_log  demo-hp  app_start_failed  warning  sent  operator  2026-08-23 09:27:51
                  (no customer row)

The number that says it all: 91 app_start_failed events stored all-time, ZERO notification_log rows before 09:00 today. Not one, ever, on any channel.

An honest limit, stated rather than implied: this run does not prove the customer gate. demo-hp has no customer_notifications row at all, so the customer leg could not have delivered regardless. The gate is proven by the unit tests, which configure prefs both ways — and Scenario B there is the positive control showing the customer leg can deliver for this event type.

Step 2 — Scenario D: stopped out of band, intent running ✅

docker compose stop privatebin at 09:31:27Z → app_start_failed (warning) at 09:31:51Z, 24 seconds later. Heartbeat, new beside last session's:

2026/08/23 05:47:44  [deadapp] check alive: 40 scans since boot, 8 deployed app(s) evaluated, 0 currently down   <- v0.222.0
2026/08/23 09:34:51  [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down   <- v0.223.0

Hub-side the second alarm was logged suppressed — operator cooldown 1h (see §15.1).

Step 3 — Scenario C: the customer presses Stop ✅

Through POST /api/stacks/privatebin/stop, the exact call the button makes. Intent moved running → stopped. Four minutes and 9 dead-app scans later: 0 alarms, 0 unknown-intent lines.

Step 4 — Scenario E: intent absent ✅ (and the count)

demo-hp has zero such apps, so one was created: desired_state removed from privatebin's app.yaml, as a pre-R-166 box would look (backed up, restored, no code writer added). Result — evaluated (deployed=True state=stopped, 8 apps evaluated), suppressed (0 alarms), and:

[deadapp] 1 stopped app(s) have NO recorded customer intent, so their dead-app alarm is suppressed
by the unknown-intent fallback (R-386): privatebin. This closes itself as each app is started or
stopped through the interface.

Step 5 — Scenario G: save changing nothing ✅ byte-identical, on the third attempt

sha256(enabled_events) = 10840f3a95bac168f0d7c79760998138 before and after a real save (7 boxes ticked, refusal banner absent).

The first two attempts were silently REFUSED behind an HTTP 200 — the empty-email wipe guard declines and renders an error page, still 200. Run 1: no prefs existed. Run 2: the email <input> spans three lines, so a single-line grep read it as "". Both times the hashes matched — because nothing was saved, not because nothing changed. Fixed by asserting the refusal banner is absent. A warning beside a success is read as a success.

Step 6 — Scenario H: a bad severity to the hub ✅

[WARN] [api] Event from demo-hp: severity "warn" is not in {info,warning,error,critical} — coercing
to "info", which severityNotifies DROPS, so this backup_failed alert will reach NOBODY. Fix the
emitting controller; this event is stored but not routed.

Both the bad POST and an error control returned 200 (nothing lost); the control produced no warning.

10. The absent-intent count on demo-hp, in plain words

Zero. All 8 deployed apps carry desired_state: running; none is stopped and none is absent. The unknown-intent fallback therefore suppresses nothing on this box today — the population is already empty on a machine that has been exercised through the interface. It will be larger on a box upgraded and left alone, which is why the log line exists rather than a one-off count.

11. The dead-branch decision, and the reason

KEPT. cmd/hub/main.go wires dispatcher.ProcessEvent directly as the monitor.EventNotifyFunc for the staleness, host-staleness and offsite-box checkers — those events never pass the ingest handler, so for them that line is the only severity guard there is. Deleting it as "dead" would have removed the live half while the dead half supplied the justification. All 90 severity literals in internal/monitor were verified already valid, so the guard is silent because the producers are correct.

12. Evidence

felhom.eu/documentation/audits/DRILL-r329-r386-2026-08-23/evidence/ — 26 files: 5 red-proof transcripts, 20 live-walk files, two full controller-log windows (1041 and 4044 lines) pulled off before each revert, and the hub's own DB queries.

13. Teardown, three layers, and the end state

  1. Guest 9201 / apps — nothing provisioned. All 17 app containers healthy. privatebin's app.yaml restored from backup and the backup deleted; intent reads running. demo-hp's notification settings restored to enabled_events: null, no e-mail — their pre-drill state. No app rebuilt, redeployed or restored; planted data untouched.
  2. Bake VM — powered off, Gitea token and runner script shredded, drill.qcow2 reverted to virgin. Build guest 9100 exists only inside that reverted snapshot. No storage added anywhere, so pvesm status has nothing to compare.
  3. Hub-side, stated explicitly. The hub was written this session, unlike last: the deployment is v0.107.0 via the manifest, and two probe events remain as rows for demo-hp from Scenario H (backup_failed, "R-387 scenario H probe" and "…control"). They are inert records; named here rather than left for someone to find. Nothing else: no appliance registered, no customer created, no artifact manifest changed, floor untouched.

End state: controller 0.223.0 and hub 0.107.0 deployed; golden 0.223.0 baked and published but NOT vouched; floor still 0.222.0; all apps running; planted data present.

14. Register size

File Before After
OPEN-ITEMS.md 328,325 B 328,132 B
CLOSED-ITEMS.md 71,441 B 74,642 B

R-329 and R-386 closed and compressed; R-387 (closed) and R-388 (the notification-model product decision — open, operator's call) filed.

15. Observations

Each item carries FILED: R-NNN or NOT-A-FINDING: <reason> (gate 11, added 2026-08-23). The markers were added retrospectively on 2026-08-24 when the gate was built: item 1 was the very finding that had no row, and adding its marker is the first thing the gate ever asked for. The observations' text is unchanged.

  1. The operator cooldown key has no app identifier, and it now bites. PrivateBin's alarm four minutes after BookStack's was logged suppressed — operator cooldown 1h, key=demo-hp:app_start_failed, so only the first app-down per hour e-mails the operator. This is R-182's known cooldown-key shape; it was harmless while app_start_failed was undeliverable and is not any more. Same pattern as R-329 itself: a known-broken thing moved from unreachable to load-bearing. Not fixed here. FILED: R-389
  2. The settings page grew 12 → 15 toggles in one session — one new alarm, plus two compound toggles split into four. Recorded as the argument inside R-388. FILED: R-388
  3. internal/notify/notifier.go carries pre-existing gofmt drift in an unrelated const block, confirmed by stashing this session's work and re-running gofmt -l. Not touched (§12). NOT-A-FINDING: cosmetic drift in an unrelated const block, predating this work and touching no behaviour; §12 forbids nearby refactors, so filing a row would only queue a whitespace commit.
  4. The golden-bake runbook still lacks pveam update — second consecutive bake to hit the stale index on the virgin snapshot, presenting as 400 … no such template. FILED: R-390
  5. Deliberately left open, untouched: R-102, R-359, R-385, and R-388's redesign. NOT-A-FINDING: this item is a pointer to rows that already exist, not a new observation; it is here so their absence from this session reads as deliberate rather than forgotten.

CI runs, confirmed by ID

Commit Repo CI id run_number Result
9832760 felhom-controller 408 88 success
68a9f54 felhom.eu 409 261 success
2f7c9a6 felhom.eu 410 262 success

(The controller docs commit's run is confirmed after its push and is the next id in that repo.)