Files
felhom.eu/documentation/architecture/08-alarm-ladder.md
T
admin 3c1882a0a4
gates / gates (push) Successful in 21s
chaos-night fixes: rulings recorded, guide reordered, three rows closed, two filed
Architecture: 08 records ruling A (45 m, the round-9 arithmetic, the cost) and
the two new event types with their audiences; 03 records the slow counter as
built; 07 records the restore-record persistence as a REVERSED design for the
restore record only; CONTEXT.md carries the day's rulings.

Guide: the recovery code moves after the first apps and waits for the yellow bar.

Register: R-549 and R-550 closed PROVEN-LIVE; R-546 closed on red-proofed tests
with its live walk owed by R-551 (no Tier-0 box is paused AND agent-connected).
R-552 filed: an interrupted-restore notice for a removed app never clears -
found in my own v0.246.0 after the release was built.

Evidence: Part A (hub prints 45m/1h30m), Part C delivery on HP and N100, B.4(a)
live proof and its teardown.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 11:00:29 +02:00

20 KiB
Raw Blame History

08 — The app-down alarm ladder

Written 2026-08-23, with controller v0.222.0 (R-384).

The absence is the finding. Until this file existed, no document owned the question "when does a customer's app being broken raise an alarm?" The rules were spread across four packages as comments, each locally correct, and the ordering between them was legible only by reading aggregateState top to bottom. That is exactly how R-384 survived: every individual rule was right, and the composition was wrong. Three separate defects in this ladder (R-51, C9-F2, R-384) were each found on live hardware rather than by review, and each is a case where a reader could not see the whole ladder at once.

Everything below is [DESIGN] — deliberate, with the reason recorded — unless marked otherwise.


1. The two questions, and their order

Two different questions get asked about a multi-container app, and the order between them is load-bearing:

  1. Is a SUPERVISED member of this app dead? — a container Docker's restart policy says should be running, that is not.
  2. Is a RUNNING member failing its healthcheck?

Question 1 is asked FIRST. [DESIGN, R-384, v0.222.0]

Until v0.222.0 it was asked second, and the consequence was not subtle: a two-container app whose database exits goes unhealthy seconds later because it cannot reach that database. So the symptom the dead database causes was the thing that suppressed the alarm for it. Measured on demo-hp 2026-08-22 — bookstack-db stopped at 21:27:01 and the watcher reported 0 currently down throughout.

"Some members are up" means any member NOT in the down bucket — running, unhealthy, starting or restarting. [DESIGN, R-384] The earlier guard was running > 0, counting only StateRunning, which made the supervised test unreachable in precisely the case it was written for: an unhealthy survivor beside a dead database counted as nothing being up.


2. Where each decision is made

Decision Where Notes
container → stack aggregate state internal/stacks/manager.go aggregateState the ladder in §3
is a down member supervised? internal/stacks/manager.go supervisedPolicy no/on-failure benign; everything else, including unknown, supervised
which states mean "down" internal/stacks/manager.go IsDownState {stopped, exited, degraded}
stack state → "this app is down" cmd/controller/main.go classifyRunStates the single derivation point; all three suppressions live here
sustained restarting → down internal/stacks/manager.go CrashLooping 5-minute threshold
quiesce suppression internal/quiesce/suppress.go cycle-keyed, 180 s grace
boot repair internal/bootrecon/bootrecon.go consumes IsDownState

3. The aggregation ladder, in order

aggregateState(containers, policyOf) — priority: degraded > unhealthy/starting > restarting > all-running > stopped.

  1. no containers → not_deployed
  2. any DOWN member is supervised, and any member is up → degraded ← R-384 put this first
  3. any unhealthy → unhealthy
  4. any starting → starting
  5. any restarting → restarting
  6. all running → running
  7. all down → stopped
  8. mix, every down member benign → running

Step 2's policyOf is consulted only for the down members, and only when something is up. nil is allowed; every down member then reads as supervised.


4. Which states alarm, and which deliberately do not

IsDownState = {stopped, exited, degraded}.

State Down? Why
stopped, exited yes not running, will not recover alone
degraded yes [DESIGN, R-51] a dead supervised member is as unreachable as a single app that exited — immich-server sat Exited 18 h with the app 100 % dead and no alert
unhealthy NO [DESIGN] a running container whose healthcheck is failing. Folding it in reintroduces the flapping fix-3 was written to stop. R-384 did not change this — it asks a prior question instead
restarting NO, until sustained [DESIGN, C9-F2] restarting is on the normal deploy path, so folding it in would alarm fleet-wide on every update. Becomes down after 5 min (crashLoopAfter)
starting, deploying no mid-start
paused no a deliberate user action
unknown no [DESIGN] fail-OPEN — never manufacture a dead-app alert from an inconclusive read

Note the two fail directions are deliberately opposite. IsDownState fails OPEN on unknown (ambiguous state → do not alarm). supervisedPolicy fails CLOSED on unknown (we already KNOW a member is dead; only the excuse is missing). Both are recorded at their sites.


5. The three suppressions, all at classifyRunStates

Suppression Rule Expires?
deliberate user stop StateStopped is not down unless the quiesce loop reports it failed to restart that stack n/a — lifted by failedRestart
quiesce cycle a stack this backup cycle stopped is exempt yes, 180 s after unquiescing
boot grace no evaluation for 90 s after controller start yes

None of them latch. [DESIGN, R-97b + R-88 Scenario D] A permanent suppression trades a loud false alarm for a silent real one, which is the same error as an over-eager alarm, in the opposite direction. Every window expires; the cost is a bounded DELAY in reporting a real failure, never its loss.

The quiesce suppression is cycle-keyed, not state-keyed — an app caught mid-restart is starting/unhealthy, not stopped, so no state test can see it. It is therefore state-blind, which is why R-384 moving a stack from unhealthy to degraded cannot weaken it.


6. The alarm itself

Edge-triggered: app_start_failed, one event per transition into down, not per scan. Verified live 2026-08-23 — one event across 22 scans.

  • Operator/hub event + dashboard banner. app_start_failed is not in settings.DefaultEnabledEvents, so by default it does not e-mail the customer.
  • F-OBS heartbeat, every 20 scans (~10 min), at [INFO]: [deadapp] check alive: N scans since boot, M deployed app(s) evaluated, K currently down. This line exists because an absent alarm and a stopped detector look identical in a log.

6.1 The severity contract [DESIGN, R-329 — CLOSED controller v0.223.0 / hub v0.107.0]

The vocabulary is the HUB's and it is exact: {info, warning, error, critical}. Anything else is coerced to info at ingest, and info is dropped by severityNotifies before both delivery legs. So a severity outside the set means the event is stored, answers 200, shows on the dashboard — and is e-mailed to nobody.

This shipped twice. DiskAlertKind.Severity emitted "warn" until controller v0.215.0; app_start_failed emitted it until v0.223.0. Measured on the live hub DB 2026-08-23: 91 app_start_failed events stored all-time, ZERO notification_log rows before that day — not one, on any channel.

Three things now hold it:

  1. The emitter is pinned by an AST walk over the whole controller (TestR329_EveryEmittedSeverityIsInTheHubVocabulary). Not grep — "warn" is a legitimate healthcheck status in internal/monitor and internal/selftest. The six call sites that pass a variable are registered by name with the values each can take, so a new dynamic path fails.
  2. The hub SAYS SO when it coerces (hub v0.107.0, R-387): a WARN naming the customer, the event type and the rejected value. The coercion stays — a rejected event is a lost event, and losing an alarm is worse than mis-routing one.
  3. The dispatcher's unrecognized severity branch is kept, because the hub's own monitor checkers call ProcessEvent directly and never pass the ingest handler. For them it is the only guard.

Who gets it. processOperator consults only operatorOn, the address and a 1-hour cooldown — never customer preferences — so a valid severity always reaches the operator. processCustomer consults operatorOnlyEvents and then the customer's enabled_events.

app_start_failed is customer-switchable but OFF by default [DESIGN, operator ruling 2026-08-23]: it is deliberately absent from DefaultEnabledEvents, and deliberately not in operatorOnlyEvents — being in that register would make the toggle visible, flickable and structurally unable to deliver.

backup_integrity_ok / backup_integrity_failed [DESIGN, R-359/R-397 — controller v0.227.0]. Both existed in internal/notify with no caller until v0.227.0 wired them; the hub had allowlisted both and carried the Hungarian customer text for both the whole time.

event severity reaches why
backup_integrity_ok info NOBODY severityNotifies drops info before both legs, and that is the intended outcome, not an oversight. A weekly success e-mail is how people stop reading their alerts. It is still pushed and stored, because the event stream is where "was it checked?" is answered — the dashboard reads it, the inbox does not
backup_integrity_failed error operator always; customer if enabled the customer's backups may be damaged, which is the loudest fact this tier can produce

backup_integrity_failed is deliberately in NONE of the three registers, and all three were checked rather than assumed (2026-08-30):

  • not in perAppCooldownEvents — there is ONE store, not one per app. The coarse per-type hourly key is correct here, and adding it would be a fenced act under §6.2 for no gain.
  • not in operatorOnlyEvents — the operator leg ignores customer preferences anyway, so the operator is always mailed; putting it here would only remove the customer's ability to opt in.
  • not in DefaultEnabledEvents — customer-switchable, default OFF, the same ruling as app_start_failed. The checkbox already exists at settings_notifications.html:34.

And a caveat that belongs in the alarm ladder rather than only in the backup document: a backup_integrity_ok at the shipped depth means the index, the pack inventory and the snapshot graph are sound. It does not mean the stored bytes were re-read — measured 2026-08-30, a pack corrupted without a size change passes the structure check with no errors were found. An ok here is a real signal about a real class of failure, and it is narrower than the phrase suggests (R-399).


6.2 The delivery grain — how often, and per what [DESIGN, R-389 — hub v0.108.0]

"How loud" is a separate decision from "does it alarm", and it is made in one place: the operator cooldown key at processOperator. Everything sharing a key is collapsed for one hour.

Operator ruling 2026-09-15 (decision A) — „the box is down" skips the quiet hour. node_stale, node_down and node_recovered no longer share the one-hour operator cooldown; they keep a 5-minute dedupe on the same key, so a flapping link cannot mail every sweep (operatorCooldownFor, hub v0.114.0, pinned by TestOperatorCooldown_NodeLivenessBypassesQuietHour). This reverses the design above for those three types only, and it is a ruling, not a defect fix. The reason is BIGNIGHT F9 (2026-09-14): the controller was dead for 33 minutes; its node_stale mail was suppressed because F8's node_stale had used the hour 39 minutes earlier, and the later node_recovered mail was suppressed the same way. EXTENDED by operator ruling 2, 2026-09-16 (hub v0.115.0, R-529): host_stale, host_down and host_recovered join the bypass, with the same 5-minute dedupe. They are the same sentence about the same box — the agent's dead-man's-switch rather than the controller's — and leaving them on the hour would have kept exactly the F9 silence on the host plane. Everything else keeps the hour (pinned by TestOperatorCooldown_NodeLivenessBypassesQuietHour, which also asserts an unrelated type still waits).

Operator ruling A, 2026-09-17 (R-549) — „the box went quiet" waits THREE report cycles, not two. The staleness threshold (alerting.stale_threshold, configuration in manifests/hub.yaml) moves from 30 m to 45 m; node_down and host_down follow at 2× = 90 m. The reason is chaos night round 9 (audits/DRILL-chaos-night-2026-09-17.md): report cadence 15 m; a failed push is retried for ~100 s and then given up (correctly — a report is a snapshot); the measured gap to the next good report was 29 m 59 s against a 30-minute threshold. One missed push spent the entire budget, so ordinary jitter would page the operator about a box that was healthy and had already repaired itself. The cost, stated: a truly dead box now pages 15 minutes later (45 m instead of 30), and „the box is down" 30 minutes later (90 m instead of 60). One value, everywhere: both staleness checkers, the host status and — since hub v0.117.0 — the customer status on the dashboard read the same threshold (controllerStatus hardcoded 30 m / 1 h until then; pinned by TestControllerStatus_FollowsConfiguredThreshold). A running hub prints it at startup (node_stale after 45m0s, node_down after 1h30m0s).

Two event types added 2026-09-17, with who receives them:

event severity who minted by why that audience
controller_slow_crashloop warning operator only hub, from the agent's slow_crashloop_since moving (agent v0.132.0, hub v0.117.0) host ids and vmids; the household's side of it is the dashboard coming back each time
restore_interrupted warning the household (and the operator) controller v0.246.0 at startup, once per interruption they pressed restore, were told it started, and can run it again — Hungarian customerMessages entry
Family Grain Key carries Why
app down (app_start_failed) per APP …:<stack_name> no digest exists for it — see below
backup run (backup_run_failures) per RUN …:<run_id> a digest already lists every failing app; one per run
tiered backup (whole_guest_backup_failed, …) per TIER …:<tier> the tiers fail independently and mean different things
everything else, incl. crossdrive_failed and backup_integrity_failed per TYPE, per hour — coarse on purpose

The default is COARSE and that is deliberate. R-97a and R-182 exist so that one full disk produces one mail listing every affected app rather than one per app. Widening the grain is what makes an operator stop reading their alerts, which is the same failure as not sending them.

app_start_failed is the exception because it has no digest. There is no apps_down_run summarising a scan the way backup_run_failures summarises a run, so per-app is the only grain available that does not lose alarms. Until hub v0.108.0 it was keyed per type, and only the first broken app per hour reached the operator — measured 2026-08-23: bookstack sent at 09:27:51, privatebin suppressed at 09:31:51 under key=demo-hp:app_start_failed.

The mechanism is a named ALLOW-LIST (perAppCooldownEvents), not a payload rule, and the distinction is load-bearing rather than stylistic: crossdrive_failed is severity error, reaches the operator leg, and carries stack_name through CrossDriveDetails. A rule of the form "if the details carry a stack_name, split per app" would have split it, silently, and undone R-182. cooldownStackSuffix therefore takes the event type as well as the details — an asymmetry with its two siblings, and the reason for it is exactly this.

Fenced act: adding an entry to perAppCooldownEvents for a type whose family has a digest, or whose coarse cooldown is deliberate. Reading the register anywhere is fine.

Measured burst, so the volume is a number and not an impression: three apps stopped in one scan produced three attempted, three sent, zero suppressed (2026-08-23). The reference box has 8 deployed apps, so a total outage is 8 mails. The boot grace (90 s), the quiesce grace (180 s) and the per-app edge trigger absorb reboots, backup cycles and persistently-dead apps. This scales linearly with app count and has no ceiling — the condition that would reopen the question is a box large enough that a total outage is unreadable, at which point the answer is a digest with a customer message, not a wider cooldown.


7. The intent test [DESIGN, R-386 — CLOSED controller v0.223.0]

"The customer stopped this" is asked of the FIELD THAT RECORDS IT, never inferred from the state.

Until v0.223.0 classifyRunStates read st.State == StateStopped and assumed every stopped stack was deliberate. It is not inferable: aggregateState folds StateExited into the stopped counter, so an all-down stack returns StateStopped whatever killed it. Measured on demo-hp 2026-08-23: privatebin stopped out of band, nine dead-app scans over four minutes, zero events, zero banner lines — while a comment beside the code claimed an out-of-band stop "still alerts".

DesiredState records the answer, has exactly one writer (the customer's own action), and is tri-state:

Intent Verdict Why
Stopped no alarm the customer asked
Running ALARM nobody asked — the R-386 case
absent ("") no alarm, and SAY SO UNKNOWN never means running

The absent case keeps the old behaviour deliberately. Reading it as "nobody asked" would, on the first cycle after upgrade, e-mail about every app any owner ever stopped — fleet-wide, from a field that predates the intent being asked of it. The backfill cannot help: it seeds Running only from an observed-up reading, so anything stopped at upgrade time stays unknown, which is precisely the ambiguous population.

The gap is BOUNDED, not silent. Every such suppression sets AppRunState.IntentUnknown, and the scheduler logs the names at INFO on the heartbeat cadence:

[deadapp] N stopped app(s) have NO recorded customer intent, so their dead-app alarm is
suppressed by the unknown-intent fallback (R-386): <names>. This closes itself as each app is
started or stopped through the interface.

A rule without a mechanism is a wish. Measured on demo-hp 2026-08-23: 0 of 8 deployed apps had an absent intent — the population is already empty on an exercised box; it will be larger on one upgraded and left alone.

failedRestart still lifts a Stopped intent, and that ordering is load-bearing: the quiesce loop stops stacks by the same path a customer does, so one it stopped and could not restart must alarm whatever the intent says. Removing that term re-opens F-CRIT-1.

Fenced act: adding a DesiredState writer. Reading it anywhere is fine. Twelve of StopStack's fourteen callers are machines, so recording intent in the primitive would make a nightly backup indistinguishable from the customer pressing Stop.


8. Direction — who a customer should be notified about at all

[DESIGN — DIRECTION, NOT CURRENT BEHAVIOUR. Dated 2026-08-23, the operator's own framing. Nothing in controller v0.223.0 / hub v0.107.0 implements this.]

A customer should be notified only about things they can act on or are responsible for — the drive they unplugged, the storage they filled. A failed backup is our incident, not theirs. The intended shape is that we detect it, we tell them we noticed and are dealing with it, and they are not handed an error they cannot solve. The subscription should feel like being looked after, not like being on call.

Today's settings page is the opposite shape: it exposes one toggle per detector and grew from 12 to 15 in this session alone (one new alarm, plus two compound toggles split into four). That growth is the argument, not an aside — a page that grows by one per detector is a page that will keep asking a household to make engineering decisions.

app_start_failed defaulting off is consistent with this direction and reversible either way; it was ruled that way on its own merits and does not pre-judge the redesign.

Filed as a PRODUCT DECISION, not a defect — see the register. It is the operator's call to take separately, and no part of it was implemented here.