The gate failed only on `released > baked`, so it could catch a forgotten bake and nothing else. A golden AHEAD of the record passed silently - and that is how controller 0.221.1 was built, baked AND vouched while the newest CHANGELOG heading still read v0.221.0, with every gate green. Reproduced on the real history: newest released 0.221.0 / newest golden baked 0.221.1 -> exit 0. The gate now asks whether the version being shipped is WRITTEN DOWN: the baked version must have its own `## vX.Y.Z` heading anywhere in the CHANGELOG. Membership rather than `baked > released` deliberately - a comparison against the newest heading alone goes green the moment any later entry is written, leaving the unrecorded version permanently unrecorded. INCONCLUSIVE (exit 2) preserved; every refusal names a reason and a route. Red-proofed both directions: old gate/old record exit 0, new gate/old record exit 1, new gate/fixed record exit 0, absent clone exit 2, post-bake exit 0. 08-alarm-ladder.md is new, and its absence was itself the finding: no document owned "when does a broken app raise an alarm?". The rules lived as comments in four packages, each locally correct, with the ordering between them legible only by reading one function top to bottom - which is how R-384 survived review. R-383 and R-384 closed into CLOSED-ITEMS with their rules kept. R-385 filed closed. R-386 filed OPEN: a single-container app stopped out of band raises no alarm, and a comment claims the opposite - measured live, 9 scans, 0 events, against a positive control from the same box 17 minutes earlier. Not fixed here. Golden 0.222.0 baked and published; vouching is the operator's act.
7.6 KiB
08 — The app-down alarm ladder
Written 2026-08-23, with controller v0.222.0 (R-384).
The absence is the finding. Until this file existed, no document owned the question "when does a
customer's app being broken raise an alarm?" The rules were spread across four packages as comments,
each locally correct, and the ordering between them was legible only by reading
aggregateState top to bottom. That is exactly how R-384 survived: every individual rule was right,
and the composition was wrong. Three separate defects in this ladder (R-51, C9-F2, R-384) were each
found on live hardware rather than by review, and each is a case where a reader could not see the
whole ladder at once.
Everything below is [DESIGN] — deliberate, with the reason recorded — unless marked otherwise.
1. The two questions, and their order
Two different questions get asked about a multi-container app, and the order between them is load-bearing:
- Is a SUPERVISED member of this app dead? — a container Docker's restart policy says should be running, that is not.
- Is a RUNNING member failing its healthcheck?
Question 1 is asked FIRST. [DESIGN, R-384, v0.222.0]
Until v0.222.0 it was asked second, and the consequence was not subtle: a two-container app whose
database exits goes unhealthy seconds later because it cannot reach that database. So the symptom
the dead database causes was the thing that suppressed the alarm for it. Measured on demo-hp
2026-08-22 — bookstack-db stopped at 21:27:01 and the watcher reported 0 currently down
throughout.
"Some members are up" means any member NOT in the down bucket — running, unhealthy,
starting or restarting. [DESIGN, R-384] The earlier guard was running > 0, counting only
StateRunning, which made the supervised test unreachable in precisely the case it was written for:
an unhealthy survivor beside a dead database counted as nothing being up.
2. Where each decision is made
| Decision | Where | Notes |
|---|---|---|
| container → stack aggregate state | internal/stacks/manager.go aggregateState |
the ladder in §3 |
| is a down member supervised? | internal/stacks/manager.go supervisedPolicy |
no/on-failure benign; everything else, including unknown, supervised |
| which states mean "down" | internal/stacks/manager.go IsDownState |
{stopped, exited, degraded} |
| stack state → "this app is down" | cmd/controller/main.go classifyRunStates |
the single derivation point; all three suppressions live here |
| sustained restarting → down | internal/stacks/manager.go CrashLooping |
5-minute threshold |
| quiesce suppression | internal/quiesce/suppress.go |
cycle-keyed, 180 s grace |
| boot repair | internal/bootrecon/bootrecon.go |
consumes IsDownState |
3. The aggregation ladder, in order
aggregateState(containers, policyOf) — priority: degraded > unhealthy/starting > restarting >
all-running > stopped.
- no containers →
not_deployed - any DOWN member is supervised, and any member is up →
degraded← R-384 put this first - any
unhealthy→unhealthy - any
starting→starting - any
restarting→restarting - all running →
running - all down →
stopped - mix, every down member benign →
running
Step 2's policyOf is consulted only for the down members, and only when something is up. nil
is allowed; every down member then reads as supervised.
4. Which states alarm, and which deliberately do not
IsDownState = {stopped, exited, degraded}.
| State | Down? | Why |
|---|---|---|
stopped, exited |
yes | not running, will not recover alone |
degraded |
yes | [DESIGN, R-51] a dead supervised member is as unreachable as a single app that exited — immich-server sat Exited 18 h with the app 100 % dead and no alert |
unhealthy |
NO | [DESIGN] a running container whose healthcheck is failing. Folding it in reintroduces the flapping fix-3 was written to stop. R-384 did not change this — it asks a prior question instead |
restarting |
NO, until sustained | [DESIGN, C9-F2] restarting is on the normal deploy path, so folding it in would alarm fleet-wide on every update. Becomes down after 5 min (crashLoopAfter) |
starting, deploying |
no | mid-start |
paused |
no | a deliberate user action |
unknown |
no | [DESIGN] fail-OPEN — never manufacture a dead-app alert from an inconclusive read |
Note the two fail directions are deliberately opposite. IsDownState fails OPEN on unknown
(ambiguous state → do not alarm). supervisedPolicy fails CLOSED on unknown (we already KNOW a
member is dead; only the excuse is missing). Both are recorded at their sites.
5. The three suppressions, all at classifyRunStates
| Suppression | Rule | Expires? |
|---|---|---|
| deliberate user stop | StateStopped is not down unless the quiesce loop reports it failed to restart that stack |
n/a — lifted by failedRestart |
| quiesce cycle | a stack this backup cycle stopped is exempt | yes, 180 s after unquiescing |
| boot grace | no evaluation for 90 s after controller start | yes |
None of them latch. [DESIGN, R-97b + R-88 Scenario D] A permanent suppression trades a loud false alarm for a silent real one, which is the same error as an over-eager alarm, in the opposite direction. Every window expires; the cost is a bounded DELAY in reporting a real failure, never its loss.
The quiesce suppression is cycle-keyed, not state-keyed — an app caught mid-restart is
starting/unhealthy, not stopped, so no state test can see it. It is therefore state-blind,
which is why R-384 moving a stack from unhealthy to degraded cannot weaken it.
6. The alarm itself
Edge-triggered: app_start_failed, one event per transition into down, not per scan. Verified
live 2026-08-23 — one event across 22 scans.
- Operator/hub event + dashboard banner.
app_start_failedis not insettings.DefaultEnabledEvents, so by default it does not e-mail the customer. - F-OBS heartbeat, every 20 scans (~10 min), at
[INFO]:[deadapp] check alive: N scans since boot, M deployed app(s) evaluated, K currently down. This line exists because an absent alarm and a stopped detector look identical in a log.
⚠ R-329, OPEN and it bites here.
app_start_failedis pushed with severitywarn, which is not in the hub's vocabulary ({info, warning, error, critical}) and is silently coerced toinfo— which e-mails nobody, while the POST still returns 200. Observed again on 2026-08-23:PushEvent: type=app_start_failed severity=warn. R-384 makes this event actually fire, so the severity bug now matters more than it did while the event was unreachable.
7. Known gap, filed not fixed
R-386 (filed 2026-08-23, OPEN). An all-down stack aggregates to
stopped—StateExitedis folded into the same counter and never survives aggregation.classifyRunStatesthen whitelistsstoppedas a deliberate user stop. So a single-container app stopped out of band raises no alarm at all, which directly contradicts the comment atcmd/controller/main.go: "An out-of-banddocker compose stopleaves the containers present → StateExited → still alerts." Measured ondemo-hp2026-08-23:privatebinstopped out of band, 9 dead-app scans over 4+ minutes,state=stopped, zero events and zero banner lines — against a positive control from the same box 17 minutes earlier. Not fixed in v0.222.0 deliberately; it is a separate decision.