Files
felhom.eu/documentation/architecture/08-alarm-ladder.md
T
admin 55274d5ef3
gates / gates (push) Successful in 17s
R-385: make an UNRECORDED golden fail the currency gate; file R-386; own the alarm ladder
The gate failed only on `released > baked`, so it could catch a forgotten bake
and nothing else. A golden AHEAD of the record passed silently - and that is
how controller 0.221.1 was built, baked AND vouched while the newest CHANGELOG
heading still read v0.221.0, with every gate green. Reproduced on the real
history: newest released 0.221.0 / newest golden baked 0.221.1 -> exit 0.

The gate now asks whether the version being shipped is WRITTEN DOWN: the baked
version must have its own `## vX.Y.Z` heading anywhere in the CHANGELOG.
Membership rather than `baked > released` deliberately - a comparison against
the newest heading alone goes green the moment any later entry is written,
leaving the unrecorded version permanently unrecorded. INCONCLUSIVE (exit 2)
preserved; every refusal names a reason and a route.

Red-proofed both directions: old gate/old record exit 0, new gate/old record
exit 1, new gate/fixed record exit 0, absent clone exit 2, post-bake exit 0.

08-alarm-ladder.md is new, and its absence was itself the finding: no document
owned "when does a broken app raise an alarm?". The rules lived as comments in
four packages, each locally correct, with the ordering between them legible only
by reading one function top to bottom - which is how R-384 survived review.

R-383 and R-384 closed into CLOSED-ITEMS with their rules kept. R-385 filed
closed. R-386 filed OPEN: a single-container app stopped out of band raises no
alarm, and a comment claims the opposite - measured live, 9 scans, 0 events,
against a positive control from the same box 17 minutes earlier. Not fixed here.

Golden 0.222.0 baked and published; vouching is the operator's act.
2026-08-23 07:59:52 +02:00

7.6 KiB

08 — The app-down alarm ladder

Written 2026-08-23, with controller v0.222.0 (R-384).

The absence is the finding. Until this file existed, no document owned the question "when does a customer's app being broken raise an alarm?" The rules were spread across four packages as comments, each locally correct, and the ordering between them was legible only by reading aggregateState top to bottom. That is exactly how R-384 survived: every individual rule was right, and the composition was wrong. Three separate defects in this ladder (R-51, C9-F2, R-384) were each found on live hardware rather than by review, and each is a case where a reader could not see the whole ladder at once.

Everything below is [DESIGN] — deliberate, with the reason recorded — unless marked otherwise.


1. The two questions, and their order

Two different questions get asked about a multi-container app, and the order between them is load-bearing:

  1. Is a SUPERVISED member of this app dead? — a container Docker's restart policy says should be running, that is not.
  2. Is a RUNNING member failing its healthcheck?

Question 1 is asked FIRST. [DESIGN, R-384, v0.222.0]

Until v0.222.0 it was asked second, and the consequence was not subtle: a two-container app whose database exits goes unhealthy seconds later because it cannot reach that database. So the symptom the dead database causes was the thing that suppressed the alarm for it. Measured on demo-hp 2026-08-22 — bookstack-db stopped at 21:27:01 and the watcher reported 0 currently down throughout.

"Some members are up" means any member NOT in the down bucket — running, unhealthy, starting or restarting. [DESIGN, R-384] The earlier guard was running > 0, counting only StateRunning, which made the supervised test unreachable in precisely the case it was written for: an unhealthy survivor beside a dead database counted as nothing being up.


2. Where each decision is made

Decision Where Notes
container → stack aggregate state internal/stacks/manager.go aggregateState the ladder in §3
is a down member supervised? internal/stacks/manager.go supervisedPolicy no/on-failure benign; everything else, including unknown, supervised
which states mean "down" internal/stacks/manager.go IsDownState {stopped, exited, degraded}
stack state → "this app is down" cmd/controller/main.go classifyRunStates the single derivation point; all three suppressions live here
sustained restarting → down internal/stacks/manager.go CrashLooping 5-minute threshold
quiesce suppression internal/quiesce/suppress.go cycle-keyed, 180 s grace
boot repair internal/bootrecon/bootrecon.go consumes IsDownState

3. The aggregation ladder, in order

aggregateState(containers, policyOf) — priority: degraded > unhealthy/starting > restarting > all-running > stopped.

  1. no containers → not_deployed
  2. any DOWN member is supervised, and any member is up → degraded ← R-384 put this first
  3. any unhealthy → unhealthy
  4. any starting → starting
  5. any restarting → restarting
  6. all running → running
  7. all down → stopped
  8. mix, every down member benign → running

Step 2's policyOf is consulted only for the down members, and only when something is up. nil is allowed; every down member then reads as supervised.


4. Which states alarm, and which deliberately do not

IsDownState = {stopped, exited, degraded}.

State Down? Why
stopped, exited yes not running, will not recover alone
degraded yes [DESIGN, R-51] a dead supervised member is as unreachable as a single app that exited — immich-server sat Exited 18 h with the app 100 % dead and no alert
unhealthy NO [DESIGN] a running container whose healthcheck is failing. Folding it in reintroduces the flapping fix-3 was written to stop. R-384 did not change this — it asks a prior question instead
restarting NO, until sustained [DESIGN, C9-F2] restarting is on the normal deploy path, so folding it in would alarm fleet-wide on every update. Becomes down after 5 min (crashLoopAfter)
starting, deploying no mid-start
paused no a deliberate user action
unknown no [DESIGN] fail-OPEN — never manufacture a dead-app alert from an inconclusive read

Note the two fail directions are deliberately opposite. IsDownState fails OPEN on unknown (ambiguous state → do not alarm). supervisedPolicy fails CLOSED on unknown (we already KNOW a member is dead; only the excuse is missing). Both are recorded at their sites.


5. The three suppressions, all at classifyRunStates

Suppression Rule Expires?
deliberate user stop StateStopped is not down unless the quiesce loop reports it failed to restart that stack n/a — lifted by failedRestart
quiesce cycle a stack this backup cycle stopped is exempt yes, 180 s after unquiescing
boot grace no evaluation for 90 s after controller start yes

None of them latch. [DESIGN, R-97b + R-88 Scenario D] A permanent suppression trades a loud false alarm for a silent real one, which is the same error as an over-eager alarm, in the opposite direction. Every window expires; the cost is a bounded DELAY in reporting a real failure, never its loss.

The quiesce suppression is cycle-keyed, not state-keyed — an app caught mid-restart is starting/unhealthy, not stopped, so no state test can see it. It is therefore state-blind, which is why R-384 moving a stack from unhealthy to degraded cannot weaken it.


6. The alarm itself

Edge-triggered: app_start_failed, one event per transition into down, not per scan. Verified live 2026-08-23 — one event across 22 scans.

  • Operator/hub event + dashboard banner. app_start_failed is not in settings.DefaultEnabledEvents, so by default it does not e-mail the customer.
  • F-OBS heartbeat, every 20 scans (~10 min), at [INFO]: [deadapp] check alive: N scans since boot, M deployed app(s) evaluated, K currently down. This line exists because an absent alarm and a stopped detector look identical in a log.

⚠ R-329, OPEN and it bites here. app_start_failed is pushed with severity warn, which is not in the hub's vocabulary ({info, warning, error, critical}) and is silently coerced to info — which e-mails nobody, while the POST still returns 200. Observed again on 2026-08-23: PushEvent: type=app_start_failed severity=warn. R-384 makes this event actually fire, so the severity bug now matters more than it did while the event was unreachable.


7. Known gap, filed not fixed

R-386 (filed 2026-08-23, OPEN). An all-down stack aggregates to stopped — StateExited is folded into the same counter and never survives aggregation. classifyRunStates then whitelists stopped as a deliberate user stop. So a single-container app stopped out of band raises no alarm at all, which directly contradicts the comment at cmd/controller/main.go: "An out-of-band docker compose stop leaves the containers present → StateExited → still alerts." Measured on demo-hp 2026-08-23: privatebin stopped out of band, 9 dead-app scans over 4+ minutes, state=stopped, zero events and zero banner lines — against a positive control from the same box 17 minutes earlier. Not fixed in v0.222.0 deliberately; it is a separate decision.