# 08 — The app-down alarm ladder **Written 2026-08-23, with controller v0.222.0 (R-384).** **The absence is the finding.** Until this file existed, no document owned the question *"when does a customer's app being broken raise an alarm?"* The rules were spread across four packages as comments, each locally correct, and the ordering between them was legible only by reading `aggregateState` top to bottom. That is exactly how R-384 survived: every individual rule was right, and the composition was wrong. Three separate defects in this ladder (R-51, C9-F2, R-384) were each found on live hardware rather than by review, and each is a case where a reader could not see the whole ladder at once. Everything below is **[DESIGN]** — deliberate, with the reason recorded — unless marked otherwise. --- ## 1. The two questions, and their order Two different questions get asked about a multi-container app, and **the order between them is load-bearing**: 1. **Is a SUPERVISED member of this app dead?** — a container Docker's restart policy says should be running, that is not. 2. **Is a RUNNING member failing its healthcheck?** **Question 1 is asked FIRST.** [DESIGN, R-384, v0.222.0] Until v0.222.0 it was asked second, and the consequence was not subtle: a two-container app whose database exits goes `unhealthy` seconds later *because it cannot reach that database*. So the symptom the dead database causes was the thing that suppressed the alarm for it. Measured on `demo-hp` 2026-08-22 — `bookstack-db` stopped at 21:27:01 and the watcher reported `0 currently down` throughout. **"Some members are up" means any member NOT in the down bucket** — `running`, `unhealthy`, `starting` or `restarting`. [DESIGN, R-384] The earlier guard was `running > 0`, counting only `StateRunning`, which made the supervised test unreachable in precisely the case it was written for: an unhealthy survivor beside a dead database counted as nothing being up. --- ## 2. Where each decision is made | Decision | Where | Notes | |---|---|---| | container → stack aggregate state | `internal/stacks/manager.go` `aggregateState` | the ladder in §3 | | is a down member supervised? | `internal/stacks/manager.go` `supervisedPolicy` | `no`/`on-failure` benign; everything else, **including unknown**, supervised | | which states mean "down" | `internal/stacks/manager.go` `IsDownState` | `{stopped, exited, degraded}` | | stack state → "this app is down" | `cmd/controller/main.go` `classifyRunStates` | **the single derivation point**; all three suppressions live here | | sustained restarting → down | `internal/stacks/manager.go` `CrashLooping` | 5-minute threshold | | quiesce suppression | `internal/quiesce/suppress.go` | cycle-keyed, 180 s grace | | boot repair | `internal/bootrecon/bootrecon.go` | consumes `IsDownState` | --- ## 3. The aggregation ladder, in order `aggregateState(containers, policyOf)` — **priority: degraded > unhealthy/starting > restarting > all-running > stopped.** 1. no containers → `not_deployed` 2. **any DOWN member is supervised, and any member is up → `degraded`** ← R-384 put this first 3. any `unhealthy` → `unhealthy` 4. any `starting` → `starting` 5. any `restarting` → `restarting` 6. all running → `running` 7. all down → `stopped` 8. mix, every down member benign → `running` Step 2's `policyOf` is consulted **only** for the down members, and only when something is up. `nil` is allowed; every down member then reads as supervised. --- ## 4. Which states alarm, and which deliberately do not `IsDownState` = `{stopped, exited, degraded}`. | State | Down? | Why | |---|---|---| | `stopped`, `exited` | **yes** | not running, will not recover alone | | `degraded` | **yes** | [DESIGN, R-51] a dead supervised member is as unreachable as a single app that exited — immich-server sat Exited 18 h with the app 100 % dead and no alert | | `unhealthy` | **NO** | [DESIGN] a running container whose healthcheck is failing. Folding it in reintroduces the flapping fix-3 was written to stop. **R-384 did not change this** — it asks a prior question instead | | `restarting` | **NO**, until sustained | [DESIGN, C9-F2] `restarting` is on the normal deploy path, so folding it in would alarm fleet-wide on every update. Becomes down after **5 min** (`crashLoopAfter`) | | `starting`, `deploying` | no | mid-start | | `paused` | no | a deliberate user action | | `unknown` | no | [DESIGN] fail-OPEN — never manufacture a dead-app alert from an inconclusive read | **Note the two fail directions are deliberately opposite.** `IsDownState` fails OPEN on `unknown` (ambiguous *state* → do not alarm). `supervisedPolicy` fails CLOSED on unknown (we already KNOW a member is dead; only the excuse is missing). Both are recorded at their sites. --- ## 5. The three suppressions, all at `classifyRunStates` | Suppression | Rule | Expires? | |---|---|---| | **deliberate user stop** | `StateStopped` is not down **unless** the quiesce loop reports it failed to restart that stack | n/a — lifted by `failedRestart` | | **quiesce cycle** | a stack this backup cycle stopped is exempt | **yes**, 180 s after unquiescing | | **boot grace** | no evaluation for 90 s after controller start | **yes** | **None of them latch.** [DESIGN, R-97b + R-88 Scenario D] A permanent suppression trades a loud false alarm for a silent real one, which is the same error as an over-eager alarm, in the opposite direction. Every window expires; the cost is a bounded DELAY in reporting a real failure, never its loss. The quiesce suppression is **cycle-keyed, not state-keyed** — an app caught mid-restart is `starting`/`unhealthy`, not `stopped`, so no state test can see it. It is therefore **state-blind**, which is why R-384 moving a stack from `unhealthy` to `degraded` cannot weaken it. --- ## 6. The alarm itself Edge-triggered: `app_start_failed`, one event per transition into down, **not** per scan. Verified live 2026-08-23 — one event across 22 scans. - **Operator/hub event + dashboard banner.** `app_start_failed` is **not** in `settings.DefaultEnabledEvents`, so by default it does **not** e-mail the customer. - **F-OBS heartbeat**, every 20 scans (~10 min), at `[INFO]`: `[deadapp] check alive: N scans since boot, M deployed app(s) evaluated, K currently down`. This line exists because an absent alarm and a stopped detector look identical in a log. ### 6.1 The severity contract [DESIGN, R-329 — CLOSED controller v0.223.0 / hub v0.107.0] **The vocabulary is the HUB's and it is exact: `{info, warning, error, critical}`.** Anything else is **coerced to `info` at ingest**, and `info` is dropped by `severityNotifies` before *both* delivery legs. So a severity outside the set means the event is stored, answers `200`, shows on the dashboard — and is e-mailed to **nobody**. **This shipped twice.** `DiskAlertKind.Severity` emitted `"warn"` until controller v0.215.0; `app_start_failed` emitted it until v0.223.0. Measured on the live hub DB 2026-08-23: **91 `app_start_failed` events stored all-time, ZERO `notification_log` rows before that day** — not one, on any channel. Three things now hold it: 1. **The emitter is pinned by an AST walk** over the whole controller (`TestR329_EveryEmittedSeverityIsInTheHubVocabulary`). Not grep — "warn" is a legitimate *healthcheck status* in `internal/monitor` and `internal/selftest`. The six call sites that pass a variable are registered by name with the values each can take, so a new dynamic path fails. 2. **The hub SAYS SO** when it coerces (hub v0.107.0, R-387): a `WARN` naming the customer, the event type and the rejected value. **The coercion stays** — a rejected event is a *lost* event, and losing an alarm is worse than mis-routing one. 3. **The dispatcher's `unrecognized severity` branch is kept**, because the hub's own monitor checkers call `ProcessEvent` directly and never pass the ingest handler. For them it is the only guard. **Who gets it.** `processOperator` consults only `operatorOn`, the address and a 1-hour cooldown — **never customer preferences** — so a valid severity always reaches the operator. `processCustomer` consults `operatorOnlyEvents` and then the customer's `enabled_events`. **`app_start_failed` is customer-switchable but OFF by default** [DESIGN, operator ruling 2026-08-23]: it is deliberately absent from `DefaultEnabledEvents`, and deliberately **not** in `operatorOnlyEvents` — being in that register would make the toggle visible, flickable and structurally unable to deliver. --- ## 7. The intent test [DESIGN, R-386 — CLOSED controller v0.223.0] **"The customer stopped this" is asked of the FIELD THAT RECORDS IT, never inferred from the state.** Until v0.223.0 `classifyRunStates` read `st.State == StateStopped` and assumed every stopped stack was deliberate. It is not inferable: `aggregateState` folds `StateExited` into the stopped counter, so an all-down stack returns `StateStopped` whatever killed it. Measured on `demo-hp` 2026-08-23: `privatebin` stopped out of band, nine dead-app scans over four minutes, **zero events, zero banner lines** — while a comment beside the code claimed an out-of-band stop *"still alerts"*. `DesiredState` records the answer, has **exactly one writer** (the customer's own action), and is tri-state: | Intent | Verdict | Why | |---|---|---| | `Stopped` | **no alarm** | the customer asked | | `Running` | **ALARM** | nobody asked — the R-386 case | | absent (`""`) | **no alarm, and SAY SO** | UNKNOWN never means running | **The absent case keeps the old behaviour deliberately.** Reading it as "nobody asked" would, on the first cycle after upgrade, e-mail about every app any owner ever stopped — fleet-wide, from a field that predates the intent being asked of it. The backfill cannot help: it seeds `Running` only from an observed-**up** reading, so anything stopped at upgrade time stays unknown, which is precisely the ambiguous population. **The gap is BOUNDED, not silent.** Every such suppression sets `AppRunState.IntentUnknown`, and the scheduler logs the names at `INFO` on the heartbeat cadence: ``` [deadapp] N stopped app(s) have NO recorded customer intent, so their dead-app alarm is suppressed by the unknown-intent fallback (R-386): . This closes itself as each app is started or stopped through the interface. ``` **A rule without a mechanism is a wish.** Measured on `demo-hp` 2026-08-23: **0 of 8 deployed apps had an absent intent** — the population is already empty on an exercised box; it will be larger on one upgraded and left alone. `failedRestart` still lifts a `Stopped` intent, and that ordering is load-bearing: the quiesce loop stops stacks by the same path a customer does, so one it stopped and could not restart must alarm whatever the intent says. Removing that term re-opens F-CRIT-1. **Fenced act:** adding a `DesiredState` **writer**. Reading it anywhere is fine. Twelve of `StopStack`'s fourteen callers are machines, so recording intent in the primitive would make a nightly backup indistinguishable from the customer pressing Stop. --- ## 8. Direction — who a customer should be notified about at all **[DESIGN — DIRECTION, NOT CURRENT BEHAVIOUR. Dated 2026-08-23, the operator's own framing. Nothing in controller v0.223.0 / hub v0.107.0 implements this.]** > **A customer should be notified only about things they can act on or are responsible for** — the > drive they unplugged, the storage they filled. **A failed backup is our incident, not theirs.** The > intended shape is that we detect it, we tell them we noticed and are dealing with it, and they are > not handed an error they cannot solve. The subscription should feel like being looked after, not > like being on call. Today's settings page is the opposite shape: it exposes one toggle per detector and **grew from 12 to 15 in this session alone** (one new alarm, plus two compound toggles split into four). That growth is the argument, not an aside — a page that grows by one per detector is a page that will keep asking a household to make engineering decisions. `app_start_failed` defaulting **off** is consistent with this direction and reversible either way; it was ruled that way on its own merits and does not pre-judge the redesign. **Filed as a PRODUCT DECISION, not a defect** — see the register. It is the operator's call to take separately, and no part of it was implemented here.