The alarm ladder gains the severity contract (the hub's vocabulary is exact, it coerces silently, and three things now hold it) and the intent test with its three-way ruling on unknown. Both marked [DESIGN] with the live measurements. Part 5 is RECORDED AND NOT IMPLEMENTED: the operator's notification philosophy, verbatim, marked plainly as direction rather than current behaviour, with the 12 -> 15 toggle growth as the argument. Filed as R-388, a product decision. R-329 and R-386 compressed into CLOSED-ITEMS with their rules kept and the full-text commit named. R-387 filed closed - including WHY the dispatcher branch was kept rather than deleted, which is evidence (three monitor checkers call ProcessEvent directly) and not caution. The drill record names three things that had to be re-run: an inert red-proof mutation, Scenario G refused twice behind an HTTP 200, and the live Scenario A NOT proving the customer gate because demo-hp has no prefs row at all. Register: OPEN 328325 -> 328132 B, CLOSED 71441 -> 74642 B.
12 KiB
08 — The app-down alarm ladder
Written 2026-08-23, with controller v0.222.0 (R-384).
The absence is the finding. Until this file existed, no document owned the question "when does a
customer's app being broken raise an alarm?" The rules were spread across four packages as comments,
each locally correct, and the ordering between them was legible only by reading
aggregateState top to bottom. That is exactly how R-384 survived: every individual rule was right,
and the composition was wrong. Three separate defects in this ladder (R-51, C9-F2, R-384) were each
found on live hardware rather than by review, and each is a case where a reader could not see the
whole ladder at once.
Everything below is [DESIGN] — deliberate, with the reason recorded — unless marked otherwise.
1. The two questions, and their order
Two different questions get asked about a multi-container app, and the order between them is load-bearing:
- Is a SUPERVISED member of this app dead? — a container Docker's restart policy says should be running, that is not.
- Is a RUNNING member failing its healthcheck?
Question 1 is asked FIRST. [DESIGN, R-384, v0.222.0]
Until v0.222.0 it was asked second, and the consequence was not subtle: a two-container app whose
database exits goes unhealthy seconds later because it cannot reach that database. So the symptom
the dead database causes was the thing that suppressed the alarm for it. Measured on demo-hp
2026-08-22 — bookstack-db stopped at 21:27:01 and the watcher reported 0 currently down
throughout.
"Some members are up" means any member NOT in the down bucket — running, unhealthy,
starting or restarting. [DESIGN, R-384] The earlier guard was running > 0, counting only
StateRunning, which made the supervised test unreachable in precisely the case it was written for:
an unhealthy survivor beside a dead database counted as nothing being up.
2. Where each decision is made
| Decision | Where | Notes |
|---|---|---|
| container → stack aggregate state | internal/stacks/manager.go aggregateState |
the ladder in §3 |
| is a down member supervised? | internal/stacks/manager.go supervisedPolicy |
no/on-failure benign; everything else, including unknown, supervised |
| which states mean "down" | internal/stacks/manager.go IsDownState |
{stopped, exited, degraded} |
| stack state → "this app is down" | cmd/controller/main.go classifyRunStates |
the single derivation point; all three suppressions live here |
| sustained restarting → down | internal/stacks/manager.go CrashLooping |
5-minute threshold |
| quiesce suppression | internal/quiesce/suppress.go |
cycle-keyed, 180 s grace |
| boot repair | internal/bootrecon/bootrecon.go |
consumes IsDownState |
3. The aggregation ladder, in order
aggregateState(containers, policyOf) — priority: degraded > unhealthy/starting > restarting >
all-running > stopped.
- no containers →
not_deployed - any DOWN member is supervised, and any member is up →
degraded← R-384 put this first - any
unhealthy→unhealthy - any
starting→starting - any
restarting→restarting - all running →
running - all down →
stopped - mix, every down member benign →
running
Step 2's policyOf is consulted only for the down members, and only when something is up. nil
is allowed; every down member then reads as supervised.
4. Which states alarm, and which deliberately do not
IsDownState = {stopped, exited, degraded}.
| State | Down? | Why |
|---|---|---|
stopped, exited |
yes | not running, will not recover alone |
degraded |
yes | [DESIGN, R-51] a dead supervised member is as unreachable as a single app that exited — immich-server sat Exited 18 h with the app 100 % dead and no alert |
unhealthy |
NO | [DESIGN] a running container whose healthcheck is failing. Folding it in reintroduces the flapping fix-3 was written to stop. R-384 did not change this — it asks a prior question instead |
restarting |
NO, until sustained | [DESIGN, C9-F2] restarting is on the normal deploy path, so folding it in would alarm fleet-wide on every update. Becomes down after 5 min (crashLoopAfter) |
starting, deploying |
no | mid-start |
paused |
no | a deliberate user action |
unknown |
no | [DESIGN] fail-OPEN — never manufacture a dead-app alert from an inconclusive read |
Note the two fail directions are deliberately opposite. IsDownState fails OPEN on unknown
(ambiguous state → do not alarm). supervisedPolicy fails CLOSED on unknown (we already KNOW a
member is dead; only the excuse is missing). Both are recorded at their sites.
5. The three suppressions, all at classifyRunStates
| Suppression | Rule | Expires? |
|---|---|---|
| deliberate user stop | StateStopped is not down unless the quiesce loop reports it failed to restart that stack |
n/a — lifted by failedRestart |
| quiesce cycle | a stack this backup cycle stopped is exempt | yes, 180 s after unquiescing |
| boot grace | no evaluation for 90 s after controller start | yes |
None of them latch. [DESIGN, R-97b + R-88 Scenario D] A permanent suppression trades a loud false alarm for a silent real one, which is the same error as an over-eager alarm, in the opposite direction. Every window expires; the cost is a bounded DELAY in reporting a real failure, never its loss.
The quiesce suppression is cycle-keyed, not state-keyed — an app caught mid-restart is
starting/unhealthy, not stopped, so no state test can see it. It is therefore state-blind,
which is why R-384 moving a stack from unhealthy to degraded cannot weaken it.
6. The alarm itself
Edge-triggered: app_start_failed, one event per transition into down, not per scan. Verified
live 2026-08-23 — one event across 22 scans.
- Operator/hub event + dashboard banner.
app_start_failedis not insettings.DefaultEnabledEvents, so by default it does not e-mail the customer. - F-OBS heartbeat, every 20 scans (~10 min), at
[INFO]:[deadapp] check alive: N scans since boot, M deployed app(s) evaluated, K currently down. This line exists because an absent alarm and a stopped detector look identical in a log.
6.1 The severity contract [DESIGN, R-329 — CLOSED controller v0.223.0 / hub v0.107.0]
The vocabulary is the HUB's and it is exact: {info, warning, error, critical}. Anything else is
coerced to info at ingest, and info is dropped by severityNotifies before both delivery
legs. So a severity outside the set means the event is stored, answers 200, shows on the dashboard —
and is e-mailed to nobody.
This shipped twice. DiskAlertKind.Severity emitted "warn" until controller v0.215.0;
app_start_failed emitted it until v0.223.0. Measured on the live hub DB 2026-08-23: 91
app_start_failed events stored all-time, ZERO notification_log rows before that day — not one,
on any channel.
Three things now hold it:
- The emitter is pinned by an AST walk over the whole controller
(
TestR329_EveryEmittedSeverityIsInTheHubVocabulary). Not grep — "warn" is a legitimate healthcheck status ininternal/monitorandinternal/selftest. The six call sites that pass a variable are registered by name with the values each can take, so a new dynamic path fails. - The hub SAYS SO when it coerces (hub v0.107.0, R-387): a
WARNnaming the customer, the event type and the rejected value. The coercion stays — a rejected event is a lost event, and losing an alarm is worse than mis-routing one. - The dispatcher's
unrecognized severitybranch is kept, because the hub's own monitor checkers callProcessEventdirectly and never pass the ingest handler. For them it is the only guard.
Who gets it. processOperator consults only operatorOn, the address and a 1-hour cooldown —
never customer preferences — so a valid severity always reaches the operator. processCustomer
consults operatorOnlyEvents and then the customer's enabled_events.
app_start_failed is customer-switchable but OFF by default [DESIGN, operator ruling 2026-08-23]:
it is deliberately absent from DefaultEnabledEvents, and deliberately not in operatorOnlyEvents
— being in that register would make the toggle visible, flickable and structurally unable to deliver.
7. The intent test [DESIGN, R-386 — CLOSED controller v0.223.0]
"The customer stopped this" is asked of the FIELD THAT RECORDS IT, never inferred from the state.
Until v0.223.0 classifyRunStates read st.State == StateStopped and assumed every stopped stack was
deliberate. It is not inferable: aggregateState folds StateExited into the stopped counter, so an
all-down stack returns StateStopped whatever killed it. Measured on demo-hp 2026-08-23:
privatebin stopped out of band, nine dead-app scans over four minutes, zero events, zero banner
lines — while a comment beside the code claimed an out-of-band stop "still alerts".
DesiredState records the answer, has exactly one writer (the customer's own action), and is
tri-state:
| Intent | Verdict | Why |
|---|---|---|
Stopped |
no alarm | the customer asked |
Running |
ALARM | nobody asked — the R-386 case |
absent ("") |
no alarm, and SAY SO | UNKNOWN never means running |
The absent case keeps the old behaviour deliberately. Reading it as "nobody asked" would, on the
first cycle after upgrade, e-mail about every app any owner ever stopped — fleet-wide, from a field
that predates the intent being asked of it. The backfill cannot help: it seeds Running only from an
observed-up reading, so anything stopped at upgrade time stays unknown, which is precisely the
ambiguous population.
The gap is BOUNDED, not silent. Every such suppression sets AppRunState.IntentUnknown, and the
scheduler logs the names at INFO on the heartbeat cadence:
[deadapp] N stopped app(s) have NO recorded customer intent, so their dead-app alarm is
suppressed by the unknown-intent fallback (R-386): <names>. This closes itself as each app is
started or stopped through the interface.
A rule without a mechanism is a wish. Measured on demo-hp 2026-08-23: 0 of 8 deployed apps had
an absent intent — the population is already empty on an exercised box; it will be larger on one
upgraded and left alone.
failedRestart still lifts a Stopped intent, and that ordering is load-bearing: the quiesce loop
stops stacks by the same path a customer does, so one it stopped and could not restart must alarm
whatever the intent says. Removing that term re-opens F-CRIT-1.
Fenced act: adding a DesiredState writer. Reading it anywhere is fine. Twelve of
StopStack's fourteen callers are machines, so recording intent in the primitive would make a nightly
backup indistinguishable from the customer pressing Stop.
8. Direction — who a customer should be notified about at all
[DESIGN — DIRECTION, NOT CURRENT BEHAVIOUR. Dated 2026-08-23, the operator's own framing. Nothing in controller v0.223.0 / hub v0.107.0 implements this.]
A customer should be notified only about things they can act on or are responsible for — the drive they unplugged, the storage they filled. A failed backup is our incident, not theirs. The intended shape is that we detect it, we tell them we noticed and are dealing with it, and they are not handed an error they cannot solve. The subscription should feel like being looked after, not like being on call.
Today's settings page is the opposite shape: it exposes one toggle per detector and grew from 12 to 15 in this session alone (one new alarm, plus two compound toggles split into four). That growth is the argument, not an aside — a page that grows by one per detector is a page that will keep asking a household to make engineering decisions.
app_start_failed defaulting off is consistent with this direction and reversible either way; it
was ruled that way on its own merits and does not pre-judge the redesign.
Filed as a PRODUCT DECISION, not a defect — see the register. It is the operator's call to take separately, and no part of it was implemented here.