REPORT overwritten: the 1.1 sweep in full (one bad severity, nine legitimate "warn" strings that are healthcheck statuses), the hub manifest's real location since the task's premise was wrong, all five red-proofs with the layer each guard sits at, the live walk in six steps with the hub's own records quoted, and the absent-intent count (0 of 8). Three things are reported that a tidier account would omit: red-proof 5 passed first time because the mutation was INERT; Scenario G was silently refused twice behind an HTTP 200; and the live Scenario A does NOT prove the customer gate, because demo-hp has no prefs row at all. CONTEXT records the severity vocabulary as a ruling with its mechanism, the intent ruling with its three-way handling of unknown, both fences, and two traps worth more than the fixes: a 200 can be a refusal, and a passing red-proof can mean an inert mutation. README: the event table said `app_start_failed | warn` - the defect, written down as if correct. Now `warning`, with the vocabulary contract and who receives what. `disk_critical` also corrected from `error` to `critical`, which is what fillwatch has always sent.
14 KiB
REPORT — controller v0.223.0 (R-329, R-386, and the compound toggles)
Session 2026-08-23, UNATTENDED. Live leg on demo-hp (Tier 0), guest 9201.
No halt condition fired. Nothing was dropped.
1. Baselines, and the hub's four numbers as read
| Repo | at start | at end |
|---|---|---|
| felhom-controller | 14137efa (v0.222.0) |
v0.223.0 deployed |
| felhom.eu | 55274d5e (hub v0.106.0) |
hub v0.107.0 deployed |
| felhom-agent | 40d857b5 (v0.130.0) |
untouched |
Hub's four numbers, live from GET /configuration before starting: golden_version 0.222.0,
agent_version 0.130.0, min_agent 0.129.0, controller floor 0.222.0 — all four as the
task predicted.
2. Documents read
internal/notify/notifier.go:583-602 (Severity()'s doc comment — it already stated the entire
contract and named both hub locations), hub/internal/api/handler.go (the type rejection and the
severity coercion side by side), hub/internal/notify/dispatcher.go (severityNotifies,
ProcessEvent, operatorOnlyEvents, processOperator), internal/stacks/deploy.go (the
DesiredState comment and its one-owner rule), cmd/controller/main.go (classifyRunStates),
internal/web/handlers.go (the compound toggles).
§2's conditional is answered: the alarm ladder DOES exist —
felhom.eu/documentation/architecture/08-alarm-ladder.md, written last session. It has been extended
here with §6.1 (the severity contract), §7 (the intent test) and §8 (Part 5's direction).
3. The 1.1 sweep — the result in full
Exactly ONE bad severity in the whole controller: notifier.go:546, "warn". Nothing else.
Verified across Go and templates and queued-event construction, because the task warned that reading zero from Go files while the answer sat in a template has produced three wrong conclusions here:
- every
emit(...)/PushEvent(...)literal — one offender, the rest valid; internal/channelhealth's classifier (the source ofNotifyAgentChannelDown's variable) — all"warning"/"error";debug.html's operator-triggerable severity<select>— offers onlyerror/warning/info;- no
PendingEvent{}literal construction exists anywhere.
Nine other "warn" strings exist and are not defects: internal/monitor and internal/selftest
use it as a healthcheck status vocabulary, and statusRank's case "warn" maps it to the correct
"warning" severity. This is why the guard is an AST walk and not grep.
One latent hazard found while sweeping, pinned rather than left: fillwatch.Band.Severity()
returns "" for BandOK. Unreachable, because Check() notifies only on an escalation — but that
safety lives in a different function from the one that looks unsafe, so the test asserts the
consequence.
And the guard found two dynamic call sites the hand sweep missed (NotifyDRCompleted, and
fillwatch's via main.go). All six are now registered by name with the values each can take.
4. The hub's manifest — §6's premise was wrong
felhom.eu/manifests/hub.yaml, line 128. ArgoCD Application/felhom tracks
admin/felhom.eu.git path manifests, automated.enabled=false. Bumped in 68a9f54. No
out-of-git deployment path exists. The image was pushed to the registry before the manifest landed,
so a sync could never point at a missing tag; the sync was then requested deliberately. Never
kubectl set image.
5. Files, commits, CI
| Commit | Repo | Contents |
|---|---|---|
9832760 |
controller | v0.223.0 — the severity word, the AST guard, the toggle, the intent test, the toggle split |
<docs> |
controller | REPORT / CONTEXT / README |
68a9f54 |
felhom.eu | hub v0.107.0 + manifest bump + golden evidence |
2f7c9a6 |
felhom.eu | alarm ladder, register, STATUS, drill record |
Modified: internal/notify/notifier.go, cmd/controller/main.go, internal/web/handlers.go,
internal/web/templates/settings_notifications.html, CHANGELOG.md.
Added: internal/notify/r329_severity_contract_test.go, internal/fillwatch/r329_severity_test.go,
cmd/controller/r386_intent_test.go, internal/web/r329_toggle_split_test.go.
CI runs confirmed BY ID (id and run_number diverge — both printed) — see §15 below.
6. Red-proofs — five planted, and ONE PASSED FIRST TIME
| # | Mutation | Layer, and why that layer | Observed |
|---|---|---|---|
| 1 | severity back to "warn" |
the EMITTER — the last point at which the bad value still exists; the hub deliberately destroys it one line later | notifier.go:561:30: emit(...) emits severity "warn", which is NOT in the hub's vocabulary |
| 2 | userStopped back to the state guess |
classifyRunStates — the single derivation point where the guess was made |
dead-app banner = [], want exactly privatebin — the exact live symptom — plus IntentUnknown = false and the wiring check |
| 3 | no-op-save guard removed | the SAVE handler — where a render-then-save rewrites stored bytes | a no-op save CHANGED the stored settings on the defaults shape |
| 4 | hub ingest WARN removed |
INGEST — the last point the offending value exists | the hub rewrote a severity and said nothing, log showing only [INFO] … (info) |
| 5 | fillwatch de-escalation guard removed | Check() — the invariant lives there, not in Severity() |
notified with band ok → severity "" (event type "") |
Red-proof 5 passed on the first attempt and that is reported, not omitted. The first mutation —
if next <= prev → if next < prev — is inert: an earlier if next == prev { continue } had
already removed the equal case, so the code's behaviour did not change and the test was right to
pass. Removing the guard outright convicts it. A red-proof that passes needs the mutation checked
before either verdict is believed. Every other mutation asserted its pre-fix text was present before
rewriting and printed MUTATION APPLIED.
7. Test counts
| Repo | Before | After |
|---|---|---|
| felhom-controller | 1504 | 1522 |
| felhom.eu hub | 702 | 709 |
Full green gate go build && go vet && go test ./... → exit 0, zero failures, both repos. All 11
controller design gates OK; all 11 felhom.eu repo gates OK.
8. Deployed versions and the golden
gitea.dooplex.hu/admin/felhom-controller:0.223.0 Up 18 minutes (healthy)
gitea.dooplex.hu/admin/felhom-hub:0.107.0 Synced, rollout complete
Golden BAKED and PUBLISHED: YES — 0.223.0.
sha256 9eaf39ac39219b42ec9e6cbf890275febcdcc6f53325fe0c0f591d3431044f17, upload OK (HTTP 201),
round-trip HTTP 206, all five acceptance markers counted.
VOUCHING IS THE OPERATOR'S ACT AND WAS NOT DONE HERE.
9. The live walk, all six steps
Step 1 — Scenario A: a database dies, customer has not opted in ✅
Quoted from the hub's own records, not the controller's:
events demo-hp app_start_failed warning 2026-08-23 09:27:51 <- v0.223.0
demo-hp app_start_failed info 2026-08-23 05:30:14 <- v0.222.0, coerced
notification_log demo-hp app_start_failed warning sent operator 2026-08-23 09:27:51
(no customer row)
The number that says it all: 91 app_start_failed events stored all-time, ZERO notification_log
rows before 09:00 today. Not one, ever, on any channel.
An honest limit, stated rather than implied: this run does not prove the customer gate.
demo-hp has no customer_notifications row at all, so the customer leg could not have delivered
regardless. The gate is proven by the unit tests, which configure prefs both ways — and Scenario B
there is the positive control showing the customer leg can deliver for this event type.
Step 2 — Scenario D: stopped out of band, intent running ✅
docker compose stop privatebin at 09:31:27Z → app_start_failed (warning) at 09:31:51Z, 24
seconds later. Heartbeat, new beside last session's:
2026/08/23 05:47:44 [deadapp] check alive: 40 scans since boot, 8 deployed app(s) evaluated, 0 currently down <- v0.222.0
2026/08/23 09:34:51 [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down <- v0.223.0
Hub-side the second alarm was logged suppressed — operator cooldown 1h (see §15.1).
Step 3 — Scenario C: the customer presses Stop ✅
Through POST /api/stacks/privatebin/stop, the exact call the button makes. Intent moved
running → stopped. Four minutes and 9 dead-app scans later: 0 alarms, 0 unknown-intent lines.
Step 4 — Scenario E: intent absent ✅ (and the count)
demo-hp has zero such apps, so one was created: desired_state removed from privatebin's
app.yaml, as a pre-R-166 box would look (backed up, restored, no code writer added). Result —
evaluated (deployed=True state=stopped, 8 apps evaluated), suppressed (0 alarms), and:
[deadapp] 1 stopped app(s) have NO recorded customer intent, so their dead-app alarm is suppressed
by the unknown-intent fallback (R-386): privatebin. This closes itself as each app is started or
stopped through the interface.
Step 5 — Scenario G: save changing nothing ✅ byte-identical, on the third attempt
sha256(enabled_events) = 10840f3a95bac168f0d7c79760998138 before and after a real save
(7 boxes ticked, refusal banner absent).
The first two attempts were silently REFUSED behind an HTTP 200 — the empty-email wipe guard
declines and renders an error page, still 200. Run 1: no prefs existed. Run 2: the email <input>
spans three lines, so a single-line grep read it as "". Both times the hashes matched —
because nothing was saved, not because nothing changed. Fixed by asserting the refusal banner is
absent. A warning beside a success is read as a success.
Step 6 — Scenario H: a bad severity to the hub ✅
[WARN] [api] Event from demo-hp: severity "warn" is not in {info,warning,error,critical} — coercing
to "info", which severityNotifies DROPS, so this backup_failed alert will reach NOBODY. Fix the
emitting controller; this event is stored but not routed.
Both the bad POST and an error control returned 200 (nothing lost); the control produced no
warning.
10. The absent-intent count on demo-hp, in plain words
Zero. All 8 deployed apps carry desired_state: running; none is stopped and none is absent.
The unknown-intent fallback therefore suppresses nothing on this box today — the population is already
empty on a machine that has been exercised through the interface. It will be larger on a box upgraded
and left alone, which is why the log line exists rather than a one-off count.
11. The dead-branch decision, and the reason
KEPT. cmd/hub/main.go wires dispatcher.ProcessEvent directly as the
monitor.EventNotifyFunc for the staleness, host-staleness and offsite-box checkers — those events
never pass the ingest handler, so for them that line is the only severity guard there is. Deleting it
as "dead" would have removed the live half while the dead half supplied the justification. All 90
severity literals in internal/monitor were verified already valid, so the guard is silent because
the producers are correct.
12. Evidence
felhom.eu/documentation/audits/DRILL-r329-r386-2026-08-23/evidence/ — 26 files: 5 red-proof
transcripts, 20 live-walk files, two full controller-log windows (1041 and 4044 lines) pulled off
before each revert, and the hub's own DB queries.
13. Teardown, three layers, and the end state
- Guest 9201 / apps — nothing provisioned. All 17 app containers healthy.
privatebin'sapp.yamlrestored from backup and the backup deleted; intent readsrunning.demo-hp's notification settings restored toenabled_events: null, no e-mail — their pre-drill state. No app rebuilt, redeployed or restored; planted data untouched. - Bake VM — powered off, Gitea token and runner script shredded,
drill.qcow2reverted tovirgin. Build guest 9100 exists only inside that reverted snapshot. No storage added anywhere, sopvesm statushas nothing to compare. - Hub-side, stated explicitly. The hub was written this session, unlike last: the deployment
is v0.107.0 via the manifest, and two probe events remain as rows for
demo-hpfrom Scenario H (backup_failed, "R-387 scenario H probe" and "…control"). They are inert records; named here rather than left for someone to find. Nothing else: no appliance registered, no customer created, no artifact manifest changed, floor untouched.
End state: controller 0.223.0 and hub 0.107.0 deployed; golden 0.223.0 baked and published but NOT vouched; floor still 0.222.0; all apps running; planted data present.
14. Register size
| File | Before | After |
|---|---|---|
OPEN-ITEMS.md |
328,325 B | 328,132 B |
CLOSED-ITEMS.md |
71,441 B | 74,642 B |
R-329 and R-386 closed and compressed; R-387 (closed) and R-388 (the notification-model product decision — open, operator's call) filed.
15. Observations — noticed, documented, NOT acted on
- The operator cooldown key has no app identifier, and it now bites. PrivateBin's alarm four
minutes after BookStack's was logged
suppressed — operator cooldown 1h, key=demo-hp:app_start_failed, so only the first app-down per hour e-mails the operator. This is R-182's known cooldown-key shape; it was harmless whileapp_start_failedwas undeliverable and is not any more. Same pattern as R-329 itself: a known-broken thing moved from unreachable to load-bearing. Not fixed here. - The settings page grew 12 → 15 toggles in one session — one new alarm, plus two compound toggles split into four. Recorded as the argument inside R-388.
internal/notify/notifier.gocarries pre-existing gofmt drift in an unrelated const block, confirmed by stashing this session's work and re-runninggofmt -l. Not touched (§12).- The golden-bake runbook still lacks
pveam update— second consecutive bake to hit the stale index on thevirginsnapshot, presenting as400 … no such template. - Deliberately left open, untouched: R-102, R-359, R-385, and R-388's redesign.
CI runs, confirmed by ID
| Commit | Repo | CI id |
run_number |
Result |
|---|---|---|---|---|
9832760 |
felhom-controller | 408 | 88 | success |
68a9f54 |
felhom.eu | 409 | 261 | success |
2f7c9a6 |
felhom.eu | 410 | 262 | success |
(The controller docs commit's run is confirmed after its push and is the next id in that repo.)