docs(R-329/R-386/R-387): the severity contract, the intent ruling, and Part 5 recorded
gates / gates (push) Successful in 17s
gates / gates (push) Successful in 17s
The alarm ladder gains the severity contract (the hub's vocabulary is exact, it coerces silently, and three things now hold it) and the intent test with its three-way ruling on unknown. Both marked [DESIGN] with the live measurements. Part 5 is RECORDED AND NOT IMPLEMENTED: the operator's notification philosophy, verbatim, marked plainly as direction rather than current behaviour, with the 12 -> 15 toggle growth as the argument. Filed as R-388, a product decision. R-329 and R-386 compressed into CLOSED-ITEMS with their rules kept and the full-text commit named. R-387 filed closed - including WHY the dispatcher branch was kept rather than deleted, which is evidence (three monitor checkers call ProcessEvent directly) and not caution. The drill record names three things that had to be re-run: an inert red-proof mutation, Scenario G refused twice behind an HTTP 200, and the live Scenario A NOT proving the customer gate because demo-hp has no prefs row at all. Register: OPEN 328325 -> 328132 B, CLOSED 71441 -> 74642 B.
This commit is contained in:
@@ -121,21 +121,106 @@ live 2026-08-23 — one event across 22 scans.
|
||||
`[deadapp] check alive: N scans since boot, M deployed app(s) evaluated, K currently down`.
|
||||
This line exists because an absent alarm and a stopped detector look identical in a log.
|
||||
|
||||
> ⚠ **R-329, OPEN and it bites here.** `app_start_failed` is pushed with severity **`warn`**, which is
|
||||
> **not** in the hub's vocabulary (`{info, warning, error, critical}`) and is silently coerced to
|
||||
> `info` — which e-mails nobody, while the POST still returns 200. Observed again on 2026-08-23:
|
||||
> `PushEvent: type=app_start_failed severity=warn`. R-384 makes this event actually fire, so the
|
||||
> severity bug now matters more than it did while the event was unreachable.
|
||||
### 6.1 The severity contract [DESIGN, R-329 — CLOSED controller v0.223.0 / hub v0.107.0]
|
||||
|
||||
**The vocabulary is the HUB's and it is exact: `{info, warning, error, critical}`.** Anything else is
|
||||
**coerced to `info` at ingest**, and `info` is dropped by `severityNotifies` before *both* delivery
|
||||
legs. So a severity outside the set means the event is stored, answers `200`, shows on the dashboard —
|
||||
and is e-mailed to **nobody**.
|
||||
|
||||
**This shipped twice.** `DiskAlertKind.Severity` emitted `"warn"` until controller v0.215.0;
|
||||
`app_start_failed` emitted it until v0.223.0. Measured on the live hub DB 2026-08-23: **91
|
||||
`app_start_failed` events stored all-time, ZERO `notification_log` rows before that day** — not one,
|
||||
on any channel.
|
||||
|
||||
Three things now hold it:
|
||||
|
||||
1. **The emitter is pinned by an AST walk** over the whole controller
|
||||
(`TestR329_EveryEmittedSeverityIsInTheHubVocabulary`). Not grep — "warn" is a legitimate
|
||||
*healthcheck status* in `internal/monitor` and `internal/selftest`. The six call sites that pass a
|
||||
variable are registered by name with the values each can take, so a new dynamic path fails.
|
||||
2. **The hub SAYS SO** when it coerces (hub v0.107.0, R-387): a `WARN` naming the customer, the event
|
||||
type and the rejected value. **The coercion stays** — a rejected event is a *lost* event, and
|
||||
losing an alarm is worse than mis-routing one.
|
||||
3. **The dispatcher's `unrecognized severity` branch is kept**, because the hub's own monitor checkers
|
||||
call `ProcessEvent` directly and never pass the ingest handler. For them it is the only guard.
|
||||
|
||||
**Who gets it.** `processOperator` consults only `operatorOn`, the address and a 1-hour cooldown —
|
||||
**never customer preferences** — so a valid severity always reaches the operator. `processCustomer`
|
||||
consults `operatorOnlyEvents` and then the customer's `enabled_events`.
|
||||
|
||||
**`app_start_failed` is customer-switchable but OFF by default** [DESIGN, operator ruling 2026-08-23]:
|
||||
it is deliberately absent from `DefaultEnabledEvents`, and deliberately **not** in `operatorOnlyEvents`
|
||||
— being in that register would make the toggle visible, flickable and structurally unable to deliver.
|
||||
|
||||
---
|
||||
|
||||
## 7. Known gap, filed not fixed
|
||||
## 7. The intent test [DESIGN, R-386 — CLOSED controller v0.223.0]
|
||||
|
||||
> **R-386 (filed 2026-08-23, OPEN).** An all-down stack aggregates to `stopped` — `StateExited` is
|
||||
> folded into the same counter and never survives aggregation. `classifyRunStates` then whitelists
|
||||
> `stopped` as a deliberate user stop. So a **single-container app stopped out of band raises no
|
||||
> alarm at all**, which directly contradicts the comment at `cmd/controller/main.go`: *"An out-of-band
|
||||
> `docker compose stop` leaves the containers present → StateExited → still alerts."*
|
||||
> **Measured on `demo-hp` 2026-08-23:** `privatebin` stopped out of band, 9 dead-app scans over 4+
|
||||
> minutes, `state=stopped`, **zero events and zero banner lines** — against a positive control from
|
||||
> the same box 17 minutes earlier. Not fixed in v0.222.0 deliberately; it is a separate decision.
|
||||
**"The customer stopped this" is asked of the FIELD THAT RECORDS IT, never inferred from the state.**
|
||||
|
||||
Until v0.223.0 `classifyRunStates` read `st.State == StateStopped` and assumed every stopped stack was
|
||||
deliberate. It is not inferable: `aggregateState` folds `StateExited` into the stopped counter, so an
|
||||
all-down stack returns `StateStopped` whatever killed it. Measured on `demo-hp` 2026-08-23:
|
||||
`privatebin` stopped out of band, nine dead-app scans over four minutes, **zero events, zero banner
|
||||
lines** — while a comment beside the code claimed an out-of-band stop *"still alerts"*.
|
||||
|
||||
`DesiredState` records the answer, has **exactly one writer** (the customer's own action), and is
|
||||
tri-state:
|
||||
|
||||
| Intent | Verdict | Why |
|
||||
|---|---|---|
|
||||
| `Stopped` | **no alarm** | the customer asked |
|
||||
| `Running` | **ALARM** | nobody asked — the R-386 case |
|
||||
| absent (`""`) | **no alarm, and SAY SO** | UNKNOWN never means running |
|
||||
|
||||
**The absent case keeps the old behaviour deliberately.** Reading it as "nobody asked" would, on the
|
||||
first cycle after upgrade, e-mail about every app any owner ever stopped — fleet-wide, from a field
|
||||
that predates the intent being asked of it. The backfill cannot help: it seeds `Running` only from an
|
||||
observed-**up** reading, so anything stopped at upgrade time stays unknown, which is precisely the
|
||||
ambiguous population.
|
||||
|
||||
**The gap is BOUNDED, not silent.** Every such suppression sets `AppRunState.IntentUnknown`, and the
|
||||
scheduler logs the names at `INFO` on the heartbeat cadence:
|
||||
|
||||
```
|
||||
[deadapp] N stopped app(s) have NO recorded customer intent, so their dead-app alarm is
|
||||
suppressed by the unknown-intent fallback (R-386): <names>. This closes itself as each app is
|
||||
started or stopped through the interface.
|
||||
```
|
||||
|
||||
**A rule without a mechanism is a wish.** Measured on `demo-hp` 2026-08-23: **0 of 8 deployed apps had
|
||||
an absent intent** — the population is already empty on an exercised box; it will be larger on one
|
||||
upgraded and left alone.
|
||||
|
||||
`failedRestart` still lifts a `Stopped` intent, and that ordering is load-bearing: the quiesce loop
|
||||
stops stacks by the same path a customer does, so one it stopped and could not restart must alarm
|
||||
whatever the intent says. Removing that term re-opens F-CRIT-1.
|
||||
|
||||
**Fenced act:** adding a `DesiredState` **writer**. Reading it anywhere is fine. Twelve of
|
||||
`StopStack`'s fourteen callers are machines, so recording intent in the primitive would make a nightly
|
||||
backup indistinguishable from the customer pressing Stop.
|
||||
|
||||
---
|
||||
|
||||
## 8. Direction — who a customer should be notified about at all
|
||||
|
||||
**[DESIGN — DIRECTION, NOT CURRENT BEHAVIOUR. Dated 2026-08-23, the operator's own framing.
|
||||
Nothing in controller v0.223.0 / hub v0.107.0 implements this.]**
|
||||
|
||||
> **A customer should be notified only about things they can act on or are responsible for** — the
|
||||
> drive they unplugged, the storage they filled. **A failed backup is our incident, not theirs.** The
|
||||
> intended shape is that we detect it, we tell them we noticed and are dealing with it, and they are
|
||||
> not handed an error they cannot solve. The subscription should feel like being looked after, not
|
||||
> like being on call.
|
||||
|
||||
Today's settings page is the opposite shape: it exposes one toggle per detector and **grew from 12 to
|
||||
15 in this session alone** (one new alarm, plus two compound toggles split into four). That growth is
|
||||
the argument, not an aside — a page that grows by one per detector is a page that will keep asking a
|
||||
household to make engineering decisions.
|
||||
|
||||
`app_start_failed` defaulting **off** is consistent with this direction and reversible either way; it
|
||||
was ruled that way on its own merits and does not pre-judge the redesign.
|
||||
|
||||
**Filed as a PRODUCT DECISION, not a defect** — see the register. It is the operator's call to take
|
||||
separately, and no part of it was implemented here.
|
||||
|
||||
Reference in New Issue
Block a user