1da2c9c6c6
gates / gates (push) Successful in 11s
REPORT overwritten: the 1.1 sweep in full (one bad severity, nine legitimate "warn" strings that are healthcheck statuses), the hub manifest's real location since the task's premise was wrong, all five red-proofs with the layer each guard sits at, the live walk in six steps with the hub's own records quoted, and the absent-intent count (0 of 8). Three things are reported that a tidier account would omit: red-proof 5 passed first time because the mutation was INERT; Scenario G was silently refused twice behind an HTTP 200; and the live Scenario A does NOT prove the customer gate, because demo-hp has no prefs row at all. CONTEXT records the severity vocabulary as a ruling with its mechanism, the intent ruling with its three-way handling of unknown, both fences, and two traps worth more than the fixes: a 200 can be a refusal, and a passing red-proof can mean an inert mutation. README: the event table said `app_start_failed | warn` - the defect, written down as if correct. Now `warning`, with the vocabulary contract and who receives what. `disk_critical` also corrected from `error` to `critical`, which is what fillwatch has always sent.
267 lines
14 KiB
Markdown
267 lines
14 KiB
Markdown
# REPORT — controller v0.223.0 (R-329, R-386, and the compound toggles)
|
|
|
|
**Session 2026-08-23, UNATTENDED.** Live leg on `demo-hp` (Tier 0), guest 9201.
|
|
**No halt condition fired.** Nothing was dropped.
|
|
|
|
## 1. Baselines, and the hub's four numbers as read
|
|
|
|
| Repo | at start | at end |
|
|
|---|---|---|
|
|
| felhom-controller | `14137efa` (v0.222.0) | **v0.223.0** deployed |
|
|
| felhom.eu | `55274d5e` (hub v0.106.0) | **hub v0.107.0** deployed |
|
|
| felhom-agent | `40d857b5` (v0.130.0) | untouched |
|
|
|
|
**Hub's four numbers, live from `GET /configuration` before starting:** `golden_version` **0.222.0**,
|
|
`agent_version` **0.130.0**, `min_agent` **0.129.0**, controller floor **0.222.0** — all four as the
|
|
task predicted.
|
|
|
|
## 2. Documents read
|
|
|
|
`internal/notify/notifier.go:583-602` (`Severity()`'s doc comment — **it already stated the entire
|
|
contract and named both hub locations**), `hub/internal/api/handler.go` (the type rejection and the
|
|
severity coercion side by side), `hub/internal/notify/dispatcher.go` (`severityNotifies`,
|
|
`ProcessEvent`, `operatorOnlyEvents`, `processOperator`), `internal/stacks/deploy.go` (the
|
|
`DesiredState` comment and its one-owner rule), `cmd/controller/main.go` (`classifyRunStates`),
|
|
`internal/web/handlers.go` (the compound toggles).
|
|
|
|
**§2's conditional is answered: the alarm ladder DOES exist** —
|
|
`felhom.eu/documentation/architecture/08-alarm-ladder.md`, written last session. It has been extended
|
|
here with §6.1 (the severity contract), §7 (the intent test) and §8 (Part 5's direction).
|
|
|
|
## 3. The 1.1 sweep — the result in full
|
|
|
|
**Exactly ONE bad severity in the whole controller: `notifier.go:546`, `"warn"`.** Nothing else.
|
|
|
|
Verified across Go **and** templates **and** queued-event construction, because the task warned that
|
|
reading zero from Go files while the answer sat in a template has produced three wrong conclusions
|
|
here:
|
|
|
|
- every `emit(...)` / `PushEvent(...)` literal — one offender, the rest valid;
|
|
- `internal/channelhealth`'s classifier (the source of `NotifyAgentChannelDown`'s variable) — all
|
|
`"warning"`/`"error"`;
|
|
- `debug.html`'s operator-triggerable severity `<select>` — offers only `error`/`warning`/`info`;
|
|
- no `PendingEvent{}` literal construction exists anywhere.
|
|
|
|
Nine other `"warn"` strings exist and are **not** defects: `internal/monitor` and `internal/selftest`
|
|
use it as a *healthcheck status* vocabulary, and `statusRank`'s `case "warn"` maps it to the correct
|
|
`"warning"` severity. **This is why the guard is an AST walk and not grep.**
|
|
|
|
**One latent hazard found while sweeping, pinned rather than left:** `fillwatch.Band.Severity()`
|
|
returns `""` for `BandOK`. Unreachable, because `Check()` notifies only on an escalation — but that
|
|
safety lives in a *different function* from the one that looks unsafe, so the test asserts the
|
|
consequence.
|
|
|
|
**And the guard found two dynamic call sites the hand sweep missed** (`NotifyDRCompleted`, and
|
|
fillwatch's via `main.go`). All six are now registered by name with the values each can take.
|
|
|
|
## 4. The hub's manifest — §6's premise was wrong
|
|
|
|
**`felhom.eu/manifests/hub.yaml`, line 128.** ArgoCD `Application/felhom` tracks
|
|
`admin/felhom.eu.git` path `manifests`, `automated.enabled=false`. Bumped in **`68a9f54`**. **No
|
|
out-of-git deployment path exists.** The image was pushed to the registry *before* the manifest landed,
|
|
so a sync could never point at a missing tag; the sync was then requested deliberately. Never
|
|
`kubectl set image`.
|
|
|
|
## 5. Files, commits, CI
|
|
|
|
| Commit | Repo | Contents |
|
|
|---|---|---|
|
|
| **`9832760`** | controller | v0.223.0 — the severity word, the AST guard, the toggle, the intent test, the toggle split |
|
|
| **`<docs>`** | controller | REPORT / CONTEXT / README |
|
|
| **`68a9f54`** | felhom.eu | hub v0.107.0 + manifest bump + golden evidence |
|
|
| **`2f7c9a6`** | felhom.eu | alarm ladder, register, STATUS, drill record |
|
|
|
|
Modified: `internal/notify/notifier.go`, `cmd/controller/main.go`, `internal/web/handlers.go`,
|
|
`internal/web/templates/settings_notifications.html`, `CHANGELOG.md`.
|
|
Added: `internal/notify/r329_severity_contract_test.go`, `internal/fillwatch/r329_severity_test.go`,
|
|
`cmd/controller/r386_intent_test.go`, `internal/web/r329_toggle_split_test.go`.
|
|
|
|
**CI runs confirmed BY ID** (`id` and `run_number` diverge — both printed) — see §15 below.
|
|
|
|
## 6. Red-proofs — five planted, and ONE PASSED FIRST TIME
|
|
|
|
| # | Mutation | Layer, and why that layer | Observed |
|
|
|---|---|---|---|
|
|
| 1 | severity back to `"warn"` | **the EMITTER** — the last point at which the bad value still exists; the hub deliberately destroys it one line later | `notifier.go:561:30: emit(...) emits severity "warn", which is NOT in the hub's vocabulary` |
|
|
| 2 | `userStopped` back to the state guess | **`classifyRunStates`** — the single derivation point where the guess was made | `dead-app banner = [], want exactly privatebin` — the exact live symptom — plus `IntentUnknown = false` and the wiring check |
|
|
| 3 | no-op-save guard removed | **the SAVE handler** — where a render-then-save rewrites stored bytes | `a no-op save CHANGED the stored settings` on the `defaults` shape |
|
|
| 4 | hub ingest `WARN` removed | **INGEST** — the last point the offending value exists | `the hub rewrote a severity and said nothing`, log showing only `[INFO] … (info)` |
|
|
| 5 | fillwatch de-escalation guard removed | **`Check()`** — the invariant lives there, not in `Severity()` | `notified with band ok → severity "" (event type "")` |
|
|
|
|
**Red-proof 5 passed on the first attempt and that is reported, not omitted.** The first mutation —
|
|
`if next <= prev` → `if next < prev` — is **inert**: an earlier `if next == prev { continue }` had
|
|
already removed the equal case, so the code's behaviour did not change and the test was *right* to
|
|
pass. Removing the guard outright convicts it. **A red-proof that passes needs the mutation checked
|
|
before either verdict is believed.** Every other mutation asserted its pre-fix text was present before
|
|
rewriting and printed `MUTATION APPLIED`.
|
|
|
|
## 7. Test counts
|
|
|
|
| Repo | Before | After |
|
|
|---|---|---|
|
|
| felhom-controller | 1504 | **1522** |
|
|
| felhom.eu hub | 702 | **709** |
|
|
|
|
Full green gate `go build && go vet && go test ./...` → **exit 0, zero failures, both repos**. All 11
|
|
controller design gates OK; all 11 felhom.eu repo gates OK.
|
|
|
|
## 8. Deployed versions and the golden
|
|
|
|
```
|
|
gitea.dooplex.hu/admin/felhom-controller:0.223.0 Up 18 minutes (healthy)
|
|
gitea.dooplex.hu/admin/felhom-hub:0.107.0 Synced, rollout complete
|
|
```
|
|
|
|
**Golden BAKED and PUBLISHED: YES — 0.223.0.**
|
|
`sha256 9eaf39ac39219b42ec9e6cbf890275febcdcc6f53325fe0c0f591d3431044f17`, `upload OK (HTTP 201)`,
|
|
round-trip **HTTP 206**, all five acceptance markers counted.
|
|
**VOUCHING IS THE OPERATOR'S ACT AND WAS NOT DONE HERE.**
|
|
|
|
## 9. The live walk, all six steps
|
|
|
|
### Step 1 — Scenario A: a database dies, customer has not opted in ✅
|
|
|
|
**Quoted from the hub's own records, not the controller's:**
|
|
|
|
```
|
|
events demo-hp app_start_failed warning 2026-08-23 09:27:51 <- v0.223.0
|
|
demo-hp app_start_failed info 2026-08-23 05:30:14 <- v0.222.0, coerced
|
|
notification_log demo-hp app_start_failed warning sent operator 2026-08-23 09:27:51
|
|
(no customer row)
|
|
```
|
|
|
|
**The number that says it all: 91 `app_start_failed` events stored all-time, ZERO `notification_log`
|
|
rows before 09:00 today.** Not one, ever, on any channel.
|
|
|
|
**An honest limit, stated rather than implied:** this run does **not** prove the customer *gate*.
|
|
`demo-hp` has no `customer_notifications` row at all, so the customer leg could not have delivered
|
|
regardless. The gate is proven by the unit tests, which configure prefs both ways — and Scenario B
|
|
there is the positive control showing the customer leg *can* deliver for this event type.
|
|
|
|
### Step 2 — Scenario D: stopped out of band, intent `running` ✅
|
|
|
|
`docker compose stop privatebin` at 09:31:27Z → **`app_start_failed (warning)` at 09:31:51Z, 24
|
|
seconds later.** Heartbeat, new beside last session's:
|
|
|
|
```
|
|
2026/08/23 05:47:44 [deadapp] check alive: 40 scans since boot, 8 deployed app(s) evaluated, 0 currently down <- v0.222.0
|
|
2026/08/23 09:34:51 [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down <- v0.223.0
|
|
```
|
|
|
|
Hub-side the second alarm was logged `suppressed — operator cooldown 1h` (see §15.1).
|
|
|
|
### Step 3 — Scenario C: the customer presses Stop ✅
|
|
|
|
Through `POST /api/stacks/privatebin/stop`, the exact call the button makes. Intent moved
|
|
`running → stopped`. Four minutes and **9 dead-app scans later: 0 alarms, 0 unknown-intent lines.**
|
|
|
|
### Step 4 — Scenario E: intent absent ✅ (and the count)
|
|
|
|
demo-hp has **zero** such apps, so one was created: `desired_state` removed from `privatebin`'s
|
|
`app.yaml`, as a pre-R-166 box would look (backed up, restored, no code writer added). Result —
|
|
evaluated (`deployed=True state=stopped`, 8 apps evaluated), **suppressed (0 alarms)**, and:
|
|
|
|
```
|
|
[deadapp] 1 stopped app(s) have NO recorded customer intent, so their dead-app alarm is suppressed
|
|
by the unknown-intent fallback (R-386): privatebin. This closes itself as each app is started or
|
|
stopped through the interface.
|
|
```
|
|
|
|
### Step 5 — Scenario G: save changing nothing ✅ byte-identical, **on the third attempt**
|
|
|
|
`sha256(enabled_events)` = `10840f3a95bac168f0d7c79760998138` **before and after** a real save
|
|
(7 boxes ticked, refusal banner absent).
|
|
|
|
**The first two attempts were silently REFUSED behind an HTTP 200** — the empty-email wipe guard
|
|
declines and renders an error page, still `200`. Run 1: no prefs existed. Run 2: the email `<input>`
|
|
spans **three lines**, so a single-line grep read it as `""`. **Both times the hashes matched —
|
|
because nothing was saved, not because nothing changed.** Fixed by asserting the refusal banner is
|
|
absent. *A warning beside a success is read as a success.*
|
|
|
|
### Step 6 — Scenario H: a bad severity to the hub ✅
|
|
|
|
```
|
|
[WARN] [api] Event from demo-hp: severity "warn" is not in {info,warning,error,critical} — coercing
|
|
to "info", which severityNotifies DROPS, so this backup_failed alert will reach NOBODY. Fix the
|
|
emitting controller; this event is stored but not routed.
|
|
```
|
|
|
|
Both the bad POST and an `error` control returned **200** (nothing lost); the control produced **no**
|
|
warning.
|
|
|
|
## 10. The absent-intent count on `demo-hp`, in plain words
|
|
|
|
**Zero.** All **8** deployed apps carry `desired_state: running`; none is `stopped` and none is absent.
|
|
The unknown-intent fallback therefore suppresses nothing on this box today — the population is already
|
|
empty on a machine that has been exercised through the interface. It will be larger on a box upgraded
|
|
and left alone, which is why the log line exists rather than a one-off count.
|
|
|
|
## 11. The dead-branch decision, and the reason
|
|
|
|
**KEPT.** `cmd/hub/main.go` wires `dispatcher.ProcessEvent` **directly** as the
|
|
`monitor.EventNotifyFunc` for the staleness, host-staleness and offsite-box checkers — those events
|
|
never pass the ingest handler, so for them that line is the only severity guard there is. Deleting it
|
|
as "dead" would have removed the live half while the dead half supplied the justification. All 90
|
|
severity literals in `internal/monitor` were verified already valid, so the guard is silent because
|
|
the producers are correct.
|
|
|
|
## 12. Evidence
|
|
|
|
`felhom.eu/documentation/audits/DRILL-r329-r386-2026-08-23/evidence/` — 26 files: 5 red-proof
|
|
transcripts, 20 live-walk files, two full controller-log windows (1041 and 4044 lines) **pulled off
|
|
before each revert**, and the hub's own DB queries.
|
|
|
|
## 13. Teardown, three layers, and the end state
|
|
|
|
1. **Guest 9201 / apps** — nothing provisioned. **All 17 app containers healthy.** `privatebin`'s
|
|
`app.yaml` restored from backup and the backup deleted; intent reads `running`. `demo-hp`'s
|
|
notification settings restored to `enabled_events: null`, no e-mail — their pre-drill state. No app
|
|
rebuilt, redeployed or restored; planted data untouched.
|
|
2. **Bake VM** — powered off, Gitea token and runner script **shredded**, `drill.qcow2` reverted to
|
|
`virgin`. Build guest 9100 exists only inside that reverted snapshot. No storage added anywhere, so
|
|
`pvesm status` has nothing to compare.
|
|
3. **Hub-side, stated explicitly.** The hub was **written** this session, unlike last: the deployment
|
|
is v0.107.0 via the manifest, and **two probe events remain as rows for `demo-hp`** from Scenario H
|
|
(`backup_failed`, "R-387 scenario H probe" and "…control"). They are inert records; named here
|
|
rather than left for someone to find. Nothing else: no appliance registered, no customer created,
|
|
no artifact manifest changed, floor untouched.
|
|
|
|
**End state:** controller **0.223.0** and hub **0.107.0** deployed; golden **0.223.0** baked and
|
|
published but **NOT vouched**; floor still **0.222.0**; all apps running; planted data present.
|
|
|
|
## 14. Register size
|
|
|
|
| File | Before | After |
|
|
|---|---|---|
|
|
| `OPEN-ITEMS.md` | 328,325 B | **328,132 B** |
|
|
| `CLOSED-ITEMS.md` | 71,441 B | **74,642 B** |
|
|
|
|
R-329 and R-386 closed and compressed; **R-387** (closed) and **R-388** (the notification-model
|
|
product decision — open, operator's call) filed.
|
|
|
|
## 15. Observations — noticed, documented, NOT acted on
|
|
|
|
1. **The operator cooldown key has no app identifier, and it now bites.** PrivateBin's alarm four
|
|
minutes after BookStack's was logged `suppressed — operator cooldown 1h,
|
|
key=demo-hp:app_start_failed`, so **only the first app-down per hour e-mails the operator**. This is
|
|
R-182's known cooldown-key shape; it was harmless while `app_start_failed` was undeliverable and is
|
|
not any more. **Same pattern as R-329 itself: a known-broken thing moved from unreachable to
|
|
load-bearing.** Not fixed here.
|
|
2. **The settings page grew 12 → 15 toggles in one session** — one new alarm, plus two compound
|
|
toggles split into four. Recorded as the argument inside R-388.
|
|
3. `internal/notify/notifier.go` carries **pre-existing** gofmt drift in an unrelated const block,
|
|
confirmed by stashing this session's work and re-running `gofmt -l`. Not touched (§12).
|
|
4. **The golden-bake runbook still lacks `pveam update`** — second consecutive bake to hit the stale
|
|
index on the `virgin` snapshot, presenting as `400 … no such template`.
|
|
5. **Deliberately left open, untouched:** R-102, R-359, R-385, and R-388's redesign.
|
|
|
|
### CI runs, confirmed by ID
|
|
|
|
| Commit | Repo | CI `id` | `run_number` | Result |
|
|
|---|---|---|---|---|
|
|
| `9832760` | felhom-controller | **408** | 88 | success |
|
|
| `68a9f54` | felhom.eu | **409** | 261 | success |
|
|
| `2f7c9a6` | felhom.eu | **410** | 262 | success |
|
|
|
|
(The controller docs commit's run is confirmed after its push and is the next `id` in that repo.)
|