docs(R-329/R-386/R-387): the severity contract, the intent ruling, and Part 5 recorded
gates / gates (push) Successful in 17s

The alarm ladder gains the severity contract (the hub's vocabulary is exact, it
coerces silently, and three things now hold it) and the intent test with its
three-way ruling on unknown. Both marked [DESIGN] with the live measurements.

Part 5 is RECORDED AND NOT IMPLEMENTED: the operator's notification philosophy,
verbatim, marked plainly as direction rather than current behaviour, with the
12 -> 15 toggle growth as the argument. Filed as R-388, a product decision.

R-329 and R-386 compressed into CLOSED-ITEMS with their rules kept and the
full-text commit named. R-387 filed closed - including WHY the dispatcher branch
was kept rather than deleted, which is evidence (three monitor checkers call
ProcessEvent directly) and not caution.

The drill record names three things that had to be re-run: an inert red-proof
mutation, Scenario G refused twice behind an HTTP 200, and the live Scenario A
NOT proving the customer gate because demo-hp has no prefs row at all.

Register: OPEN 328325 -> 328132 B, CLOSED 71441 -> 74642 B.
This commit is contained in:
2026-08-23 12:03:49 +02:00
parent 68a9f5475c
commit 2f7c9a6ce5
33 changed files with 5774 additions and 104 deletions
+99 -14
View File
@@ -121,21 +121,106 @@ live 2026-08-23 — one event across 22 scans.
`[deadapp] check alive: N scans since boot, M deployed app(s) evaluated, K currently down`.
This line exists because an absent alarm and a stopped detector look identical in a log.
> ⚠ **R-329, OPEN and it bites here.** `app_start_failed` is pushed with severity **`warn`**, which is
> **not** in the hub's vocabulary (`{info, warning, error, critical}`) and is silently coerced to
> `info` — which e-mails nobody, while the POST still returns 200. Observed again on 2026-08-23:
> `PushEvent: type=app_start_failed severity=warn`. R-384 makes this event actually fire, so the
> severity bug now matters more than it did while the event was unreachable.
### 6.1 The severity contract [DESIGN, R-329 — CLOSED controller v0.223.0 / hub v0.107.0]
**The vocabulary is the HUB's and it is exact: `{info, warning, error, critical}`.** Anything else is
**coerced to `info` at ingest**, and `info` is dropped by `severityNotifies` before *both* delivery
legs. So a severity outside the set means the event is stored, answers `200`, shows on the dashboard —
and is e-mailed to **nobody**.
**This shipped twice.** `DiskAlertKind.Severity` emitted `"warn"` until controller v0.215.0;
`app_start_failed` emitted it until v0.223.0. Measured on the live hub DB 2026-08-23: **91
`app_start_failed` events stored all-time, ZERO `notification_log` rows before that day** — not one,
on any channel.
Three things now hold it:
1. **The emitter is pinned by an AST walk** over the whole controller
(`TestR329_EveryEmittedSeverityIsInTheHubVocabulary`). Not grep — "warn" is a legitimate
*healthcheck status* in `internal/monitor` and `internal/selftest`. The six call sites that pass a
variable are registered by name with the values each can take, so a new dynamic path fails.
2. **The hub SAYS SO** when it coerces (hub v0.107.0, R-387): a `WARN` naming the customer, the event
type and the rejected value. **The coercion stays** — a rejected event is a *lost* event, and
losing an alarm is worse than mis-routing one.
3. **The dispatcher's `unrecognized severity` branch is kept**, because the hub's own monitor checkers
call `ProcessEvent` directly and never pass the ingest handler. For them it is the only guard.
**Who gets it.** `processOperator` consults only `operatorOn`, the address and a 1-hour cooldown —
**never customer preferences** — so a valid severity always reaches the operator. `processCustomer`
consults `operatorOnlyEvents` and then the customer's `enabled_events`.
**`app_start_failed` is customer-switchable but OFF by default** [DESIGN, operator ruling 2026-08-23]:
it is deliberately absent from `DefaultEnabledEvents`, and deliberately **not** in `operatorOnlyEvents`
— being in that register would make the toggle visible, flickable and structurally unable to deliver.
---
## 7. Known gap, filed not fixed
## 7. The intent test [DESIGN, R-386 — CLOSED controller v0.223.0]
> **R-386 (filed 2026-08-23, OPEN).** An all-down stack aggregates to `stopped` — `StateExited` is
> folded into the same counter and never survives aggregation. `classifyRunStates` then whitelists
> `stopped` as a deliberate user stop. So a **single-container app stopped out of band raises no
> alarm at all**, which directly contradicts the comment at `cmd/controller/main.go`: *"An out-of-band
> `docker compose stop` leaves the containers present → StateExited → still alerts."*
> **Measured on `demo-hp` 2026-08-23:** `privatebin` stopped out of band, 9 dead-app scans over 4+
> minutes, `state=stopped`, **zero events and zero banner lines** — against a positive control from
> the same box 17 minutes earlier. Not fixed in v0.222.0 deliberately; it is a separate decision.
**"The customer stopped this" is asked of the FIELD THAT RECORDS IT, never inferred from the state.**
Until v0.223.0 `classifyRunStates` read `st.State == StateStopped` and assumed every stopped stack was
deliberate. It is not inferable: `aggregateState` folds `StateExited` into the stopped counter, so an
all-down stack returns `StateStopped` whatever killed it. Measured on `demo-hp` 2026-08-23:
`privatebin` stopped out of band, nine dead-app scans over four minutes, **zero events, zero banner
lines** — while a comment beside the code claimed an out-of-band stop *"still alerts"*.
`DesiredState` records the answer, has **exactly one writer** (the customer's own action), and is
tri-state:
| Intent | Verdict | Why |
|---|---|---|
| `Stopped` | **no alarm** | the customer asked |
| `Running` | **ALARM** | nobody asked — the R-386 case |
| absent (`""`) | **no alarm, and SAY SO** | UNKNOWN never means running |
**The absent case keeps the old behaviour deliberately.** Reading it as "nobody asked" would, on the
first cycle after upgrade, e-mail about every app any owner ever stopped — fleet-wide, from a field
that predates the intent being asked of it. The backfill cannot help: it seeds `Running` only from an
observed-**up** reading, so anything stopped at upgrade time stays unknown, which is precisely the
ambiguous population.
**The gap is BOUNDED, not silent.** Every such suppression sets `AppRunState.IntentUnknown`, and the
scheduler logs the names at `INFO` on the heartbeat cadence:
```
[deadapp] N stopped app(s) have NO recorded customer intent, so their dead-app alarm is
suppressed by the unknown-intent fallback (R-386): <names>. This closes itself as each app is
started or stopped through the interface.
```
**A rule without a mechanism is a wish.** Measured on `demo-hp` 2026-08-23: **0 of 8 deployed apps had
an absent intent** — the population is already empty on an exercised box; it will be larger on one
upgraded and left alone.
`failedRestart` still lifts a `Stopped` intent, and that ordering is load-bearing: the quiesce loop
stops stacks by the same path a customer does, so one it stopped and could not restart must alarm
whatever the intent says. Removing that term re-opens F-CRIT-1.
**Fenced act:** adding a `DesiredState` **writer**. Reading it anywhere is fine. Twelve of
`StopStack`'s fourteen callers are machines, so recording intent in the primitive would make a nightly
backup indistinguishable from the customer pressing Stop.
---
## 8. Direction — who a customer should be notified about at all
**[DESIGN — DIRECTION, NOT CURRENT BEHAVIOUR. Dated 2026-08-23, the operator's own framing.
Nothing in controller v0.223.0 / hub v0.107.0 implements this.]**
> **A customer should be notified only about things they can act on or are responsible for** — the
> drive they unplugged, the storage they filled. **A failed backup is our incident, not theirs.** The
> intended shape is that we detect it, we tell them we noticed and are dealing with it, and they are
> not handed an error they cannot solve. The subscription should feel like being looked after, not
> like being on call.
Today's settings page is the opposite shape: it exposes one toggle per detector and **grew from 12 to
15 in this session alone** (one new alarm, plus two compound toggles split into four). That growth is
the argument, not an aside — a page that grows by one per detector is a page that will keep asking a
household to make engineering decisions.
`app_start_failed` defaulting **off** is consistent with this direction and reversible either way; it
was ruled that way on its own merits and does not pre-judge the redesign.
**Filed as a PRODUCT DECISION, not a defect** — see the register. It is the operator's call to take
separately, and no part of it was implemented here.