Files

443 lines
32 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 08 — The app-down alarm ladder
> **How to read this document.** Where a statement is marked, it is marked like this — the same wording as
> `07-backup-architecture.md:11-17`, carried here on 2026-10-05 (R-376, the three documents written after the
> 2026-08-22 pass):
>
> - **[DESIGN]** — a decision taken. Not derived from code; the code may not implement it yet.
> - **[FACT]** — an observed property, carrying a `file:line`, a live command output or a citation.
>
> **An unmarked statement means "not yet classified", never "observed"** (R-376).
**Written 2026-08-23, with controller v0.222.0 (R-384).**
**The absence is the finding.** Until this file existed, no document owned the question *"when does a
customer's app being broken raise an alarm?"* The rules were spread across four packages as comments,
each locally correct, and the ordering between them was legible only by reading
`aggregateState` top to bottom. That is exactly how R-384 survived: every individual rule was right,
and the composition was wrong. Three separate defects in this ladder (R-51, C9-F2, R-384) were each
found on live hardware rather than by review, and each is a case where a reader could not see the
whole ladder at once.
Everything below is **[DESIGN]** — deliberate, with the reason recorded — unless marked otherwise.
---
## 1. The two questions, and their order
Two different questions get asked about a multi-container app, and **the order between them is
load-bearing**:
1. **Is a SUPERVISED member of this app dead?** — a container Docker's restart policy says should be
running, that is not.
2. **Is a RUNNING member failing its healthcheck?**
**Question 1 is asked FIRST.** [DESIGN, R-384, v0.222.0]
Until v0.222.0 it was asked second, and the consequence was not subtle: a two-container app whose
database exits goes `unhealthy` seconds later *because it cannot reach that database*. So the symptom
the dead database causes was the thing that suppressed the alarm for it. Measured on `demo-hp`
2026-08-22 — `bookstack-db` stopped at 21:27:01 and the watcher reported `0 currently down`
throughout.
**"Some members are up" means any member NOT in the down bucket** — `running`, `unhealthy`,
`starting` or `restarting`. [DESIGN, R-384] The earlier guard was `running > 0`, counting only
`StateRunning`, which made the supervised test unreachable in precisely the case it was written for:
an unhealthy survivor beside a dead database counted as nothing being up.
---
## 2. Where each decision is made
| Decision | Where | Notes |
|---|---|---|
| container → stack aggregate state | `internal/stacks/manager.go` `aggregateState` | the ladder in §3 |
| is a down member supervised? | `internal/stacks/manager.go` `supervisedPolicy` | `no`/`on-failure` benign; everything else, **including unknown**, supervised |
| which states mean "down" | `internal/stacks/manager.go` `IsDownState` | `{stopped, exited, degraded}` |
| stack state → "this app is down" | `cmd/controller/main.go` `classifyRunStates` | **the single derivation point**; all three suppressions live here |
| sustained restarting → down | `internal/stacks/manager.go` `CrashLooping` | 5-minute threshold |
| quiesce suppression | `internal/quiesce/suppress.go` | cycle-keyed, 180 s grace |
| boot repair | `internal/bootrecon/bootrecon.go` | consumes `IsDownState` |
---
## 3. The aggregation ladder, in order
`aggregateState(containers, policyOf)` — **priority: degraded > unhealthy/starting > restarting >
all-running > stopped.**
1. no containers → `not_deployed`
2. **any DOWN member is supervised, and any member is up → `degraded`** ← R-384 put this first
3. any `unhealthy` → `unhealthy`
4. any `starting` → `starting`
5. any `restarting` → `restarting`
6. all running → `running`
7. all down → `stopped`
8. mix, every down member benign → `running`
Step 2's `policyOf` is consulted **only** for the down members, and only when something is up. `nil`
is allowed; every down member then reads as supervised.
---
## 4. Which states alarm, and which deliberately do not
`IsDownState` = `{stopped, exited, degraded}`.
| State | Down? | Why |
|---|---|---|
| `stopped`, `exited` | **yes** | not running, will not recover alone |
| `degraded` | **yes** | [DESIGN, R-51] a dead supervised member is as unreachable as a single app that exited — immich-server sat Exited 18 h with the app 100 % dead and no alert |
| `unhealthy` | **NO** | [DESIGN] a running container whose healthcheck is failing. Folding it in reintroduces the flapping fix-3 was written to stop. **R-384 did not change this** — it asks a prior question instead |
| `restarting` | **NO**, until sustained | [DESIGN, C9-F2] `restarting` is on the normal deploy path, so folding it in would alarm fleet-wide on every update. Becomes down after **5 min** (`crashLoopAfter`) of UNINTERRUPTED `restarting` — **measured 2026-09-24: a container that reads `running` for a moment between restarts resets the clock and never gets there (gokapi, 385 restarts, „0 currently down"; R-667)** **Since controller v0.269.0 the box counts `RestartCount` instead and STOPS the app (§6.2, decision 28).** |
| `starting`, `deploying` | no | mid-start |
| `paused` | no | a deliberate user action |
| `unknown` | no | [DESIGN] fail-OPEN — never manufacture a dead-app alert from an inconclusive read |
**Note the two fail directions are deliberately opposite.** `IsDownState` fails OPEN on `unknown`
(ambiguous *state* → do not alarm). `supervisedPolicy` fails CLOSED on unknown (we already KNOW a
member is dead; only the excuse is missing). Both are recorded at their sites.
---
## 5. The four suppressions, all at `classifyRunStates`
| Suppression | Rule | Expires? |
|---|---|---|
| **deliberate user stop** | `StateStopped` is not down **unless** the quiesce loop reports it failed to restart that stack | n/a — lifted by `failedRestart` |
| **quiesce cycle** | a stack this backup cycle stopped is exempt | **yes**, 180 s after unquiescing |
| **boot grace** | no evaluation for 90 s after controller start | **yes** |
| **update hold** (v0.268.0, R-660) | an app HELD after a failed update is stopped by the product and has its own event (`app_update_held`, and `app_hold_no_whole_copy` when no copy brings it back whole); it is not "down" | **yes** — lifted with the hold (a restore, or the operator). A RESTORE hold (R-379) is deliberately not in the set: it has no event of its own |
**None of them latch.** [DESIGN, R-97b + R-88 Scenario D] A permanent suppression trades a loud false
alarm for a silent real one, which is the same error as an over-eager alarm, in the opposite
direction. Every window expires; the cost is a bounded DELAY in reporting a real failure, never its
loss.
The quiesce suppression is **cycle-keyed, not state-keyed** — an app caught mid-restart is
`starting`/`unhealthy`, not `stopped`, so no state test can see it. It is therefore **state-blind**,
which is why R-384 moving a stack from `unhealthy` to `degraded` cannot weaken it.
---
## 6. The alarm itself
Edge-triggered: `app_start_failed`, one event per transition into down, **not** per scan. Verified
live 2026-08-23 — one event across 22 scans.
- **Operator/hub event + dashboard banner.** `app_start_failed` is **not** in
`settings.DefaultEnabledEvents`, so by default it does **not** e-mail the customer.
- **F-OBS heartbeat**, every 20 scans (~10 min), at `[INFO]`:
`[deadapp] check alive: N scans since boot, M deployed app(s) evaluated, K currently down`.
This line exists because an absent alarm and a stopped detector look identical in a log.
### 6.1 The severity contract [DESIGN, R-329 — CLOSED controller v0.223.0 / hub v0.107.0]
**The vocabulary is the HUB's and it is exact: `{info, warning, error, critical}`.** Anything else is
**coerced to `info` at ingest**, and `info` is dropped by `severityNotifies` before *both* delivery
legs. So a severity outside the set means the event is stored, answers `200`, shows on the dashboard —
and is e-mailed to **nobody**.
**This shipped twice.** `DiskAlertKind.Severity` emitted `"warn"` until controller v0.215.0;
`app_start_failed` emitted it until v0.223.0. Measured on the live hub DB 2026-08-23: **91
`app_start_failed` events stored all-time, ZERO `notification_log` rows before that day** — not one,
on any channel.
Three things now hold it:
1. **The emitter is pinned by an AST walk** over the whole controller
(`TestR329_EveryEmittedSeverityIsInTheHubVocabulary`). Not grep — "warn" is a legitimate
*healthcheck status* in `internal/monitor` and `internal/selftest`. The six call sites that pass a
variable are registered by name with the values each can take, so a new dynamic path fails.
2. **The hub SAYS SO** when it coerces (hub v0.107.0, R-387): a `WARN` naming the customer, the event
type and the rejected value. **The coercion stays** — a rejected event is a *lost* event, and
losing an alarm is worse than mis-routing one.
3. **The dispatcher's `unrecognized severity` branch is kept**, because the hub's own monitor checkers
call `ProcessEvent` directly and never pass the ingest handler. For them it is the only guard.
**Who gets it.** `processOperator` consults only `operatorOn`, the address and a 1-hour cooldown —
**never customer preferences** — so a valid severity always reaches the operator. `processCustomer`
consults `operatorOnlyEvents` and then the customer's `enabled_events`.
**`app_start_failed` is customer-switchable but OFF by default** [DESIGN, operator ruling 2026-08-23]:
it is deliberately absent from `DefaultEnabledEvents`, and deliberately **not** in `operatorOnlyEvents`
— being in that register would make the toggle visible, flickable and structurally unable to deliver.
**`backup_integrity_ok` / `backup_integrity_failed`** [DESIGN, R-359/R-397 — controller v0.227.0]. Both
existed in `internal/notify` with **no caller** until v0.227.0 wired them; the hub had allowlisted both
and carried the Hungarian customer text for both the whole time.
| event | severity | reaches | why |
|---|---|---|---|
| `backup_integrity_ok` | **`info`** | **NOBODY** | `severityNotifies` drops `info` before both legs, and that is the intended outcome, not an oversight. **A weekly success e-mail is how people stop reading their alerts.** It is still pushed and stored, because the event stream is where "was it checked?" is answered — the dashboard reads it, the inbox does not |
| `backup_integrity_failed` | `error` | operator always; customer if enabled | the customer's backups may be damaged, which is the loudest fact this tier can produce |
**`backup_integrity_failed` is deliberately in NONE of the three registers**, and all three were checked
rather than assumed (2026-08-30):
- **not** in `perAppCooldownEvents` — there is ONE store, not one per app. The coarse per-type hourly
key is correct here, and adding it would be a fenced act under §6.2 for no gain.
- **not** in `operatorOnlyEvents` — the operator leg ignores customer preferences anyway, so the
operator is always mailed; putting it here would only remove the customer's ability to opt in.
- **not** in `DefaultEnabledEvents` — customer-switchable, default OFF, the same ruling as
`app_start_failed`. The checkbox already exists at `settings_notifications.html:34`.
**And a caveat that belongs in the alarm ladder rather than only in the backup document:** a
`backup_integrity_ok` at the shipped depth means *the index, the pack inventory and the snapshot graph
are sound*. It does **not** mean the stored bytes were re-read — measured 2026-08-30, a pack corrupted
without a size change passes the structure check with `no errors were found`. An `ok` here is a real
signal about a real class of failure, and it is narrower than the phrase suggests (R-399).
---
## 6.2 The delivery grain — how often, and per what [DESIGN, R-389 — hub v0.108.0]
**"How loud" is a separate decision from "does it alarm", and it is made in one place**: the operator
cooldown key at `processOperator`. Everything sharing a key is collapsed for **one hour**.
**Operator ruling 2026-09-15 (decision A) — „the box is down" skips the quiet hour.** `node_stale`,
`node_down` and `node_recovered` no longer share the one-hour operator cooldown; they keep a **5-minute
dedupe** on the same key, so a flapping link cannot mail every sweep (`operatorCooldownFor`, hub
v0.114.0, pinned by `TestOperatorCooldown_NodeLivenessBypassesQuietHour`). **This reverses the design
above for those three types only**, and it is a ruling, not a defect fix. The reason is BIGNIGHT F9
(2026-09-14): the controller was dead for 33 minutes; its `node_stale` mail was suppressed because F8's
`node_stale` had used the hour 39 minutes earlier, and the later `node_recovered` mail was suppressed
the same way. **EXTENDED by operator ruling 2, 2026-09-16 (hub v0.115.0, R-529): `host_stale`, `host_down` and
`host_recovered` join the bypass**, with the same 5-minute dedupe. They are the same sentence about the
same box — the agent's dead-man's-switch rather than the controller's — and leaving them on the hour
would have kept exactly the F9 silence on the host plane. Everything else keeps the hour (pinned by
`TestOperatorCooldown_NodeLivenessBypassesQuietHour`, which also asserts an unrelated type still waits).
**Operator ruling A, 2026-09-17 (R-549) — „the box went quiet" waits THREE report cycles, not two.**
The staleness threshold (`alerting.stale_threshold`, configuration in `manifests/hub.yaml`) moves from
**30 m to 45 m**; `node_down` and `host_down` follow at 2× = **90 m**. The reason is chaos night round 9
(`audits/DRILL-chaos-night-2026-09-17.md`): report cadence 15 m; a failed push is retried for ~100 s and
then given up (correctly — a report is a snapshot); the measured gap to the next good report was
**29 m 59 s** against a 30-minute threshold. One missed push spent the entire budget, so ordinary jitter
would page the operator about a box that was healthy and had already repaired itself. **The cost,
stated:** a truly dead box now pages 15 minutes later (45 m instead of 30), and „the box is down" 30
minutes later (90 m instead of 60). **One value, everywhere:** both staleness checkers, the host status and
— since hub v0.117.0 — the customer status on the dashboard read the same threshold (`controllerStatus`
hardcoded 30 m / 1 h until then; pinned by `TestControllerStatus_FollowsConfiguredThreshold`). A running
hub prints it at startup (`node_stale after 45m0s, node_down after 1h30m0s`).
**An out-of-memory storm gets its own, louder rung (R-636, controller v0.265.0 / hub v0.121.0, 2026-09-23).**
`app_oom` stays exactly as it was: `warning`, operator-only, ONCE per container run (R-514) — the once
is what stops a crash loop from mailing thousands of times (R-629). Beside it, `app_oom_storm`:
| event | severity | who | minted by | why that audience |
|---|---|---|---|---|
| `app_oom_storm` | **error** | **operator only** | controller v0.265.0, when the kernel's `oom_kill` counter of the SAME container run rises by **≥ 20 within 30 min**; once per run | raw container names and memory figures; the household's side is the dashboard tag |
**Why a counter and not the flag:** Docker's `OOMKilled` is sticky — true for the whole run after ONE
kill — so "the key re-fired N times" measures only how long ago the first kill was. The kernel's
`memory.events` `oom_kill` counts kills. **Why 20 in 30 minutes:** RomM's measured rate on 2026-09-22
was 4,530 kills in six hours ≈ 375 per 30 min; a hiccup is 1–3. Live on 9202 (2026-09-23): RomM at 320M
reached 21 kills 2.5 min after start and sent ONE storm; at 49 kills still one. **Not an app-down state:**
the app still reads `running` and `IsDownState` is unchanged (§4). **Limit:** a container whose
`OOMKilled` flag stays false in an LXC guest (R-528) is never read, so it never storms.
**The box STOPS a crash loop or an out-of-memory storm (decision 28 of `09` §3, R-667, controller v0.269.0 /
hub v0.123.0, 2026-09-24).** Alarming was not enough: gokapi crash-looped for hours at 385 → 546 restarts
while the ladder said „0 currently down" (the `restarting` clock resets, §4). Now, besides the alarms above:
| rule | threshold | what the box does | event |
|---|---|---|---|
| crash loop | **≥ 6 restarts within 10 min**, from Docker's `RestartCount` summed per app, sampled every scan (a drop = the container was recreated → the window starts again) | stops the app, records an `unhealthy_stop` hold (the same store as every hold, so no start path revives it), the page says so with a Start button | `app_stopped_unhealthy`, **warning**, **the household AND the operator** (default on, seeded add-only) |
| out-of-memory storm | the `app_oom_storm` rule above (≥ 20 kernel kills within 30 min) | the same stop and hold | `app_oom_storm` (operator) + `app_stopped_unhealthy` |
| Start pressed | — | lifts the hold: **one more try** | — |
| stopped again within 24 h | — | stops it again; the sentence now says Felhom support is informed | `app_stopped_unhealthy` (repeat wording) |
**Why 6 in 10 minutes and not 10:** Docker's restart back-off caps a steady loop at about ONE restart a
minute (gokapi: 539 → 546 in 7 min), so 10 in 10 sits on the edge and misses a steady loop; a FRESH
container restarts fast (9 in 32 s after a Start, measured). No healthy app in any drill evidence (1,831
harness samples, 40 live containers) restarted more than once on a first start; the one borderline shape
is immich's first-start import (12 restarts, 2026-09-17 — broken that night; R-676 watches it).
**Never judged:** an app that is deploying, updating, held, or stopped by a backup or a quiesce.
**Precisely (read from source, controller v0.271.0, 2026-09-25):** "updating" covers an automatic update's
whole step — pull, start, verify AND its undo — because `Updating` stays true until the step ends
(`TestD28_NoCrashLoopStopDuringAnAutomaticStep`; the leg's own pages are therefore never stopped mid-step).
"Deploying" does **not** cover a deploy's FIRST START: the flag clears when `compose up -d` returns, and the
app is sampled from then on — a first start that restarts ≥ 6 times in 10 min is stopped (R-676). An automatic
update is a product stop in the §5 sense: a step's restarts are the product's own, never an alarm.
**Measured in the chaos hour (night 2026-09-24):** stopped at +185 s after a power cut mid-storm (the
counting starts again after the boot), +116 s for a crash loop under a backup run, +102 s for a storm with
the disk 1 GB above the floor.
**A thin pool is critical at 90 % (R-672, hub v0.124.0 + agent v0.133.0).** `storage_fill_*` judges an
`lvmthin` target on the worse of data and metadata fill, warning 85 %, **critical 90 %** (other storages keep
90/95 %). A full thin pool does not only refuse writes: every guest on it remounts read-only (demo-hp
2026-09-24 — 95 % → 100 % in about a minute, 9201 read-only five minutes later). The agent requests an
immediate host report when a pool crosses 90 %, so the alarm fires in seconds, not at the next 15-minute
report. Grain: **per pool, 6 hours** (`customer:type:host/storage`) — the old `customer:type`, 1 hour let one
pool silence another. The hub HAD alarmed on 2026-09-24 (operator mail at 100 %), but late and at the generic
bands.
**Two event types added 2026-09-17, with who receives them:**
| event | severity | who | minted by | why that audience |
|---|---|---|---|---|
| `controller_slow_crashloop` | warning | **operator only** | hub, from the agent's `slow_crashloop_since` moving (agent v0.132.0, hub v0.117.0) | host ids and vmids; the household's side of it is the dashboard coming back each time |
| `restore_interrupted` | warning | **the household** (and the operator) | controller v0.246.0 at startup, once per interruption | they pressed restore, were told it started, and can run it again — Hungarian `customerMessages` entry |
| Family | Grain | Key carries | Why |
|---|---|---|---|
| app down (`app_start_failed`) | **per APP** | `…:<stack_name>` | no digest exists for it — see below |
| update outcome (`app_update_undone`, `app_update_held`) | **per APP**, both legs | `…:<stack_name>` | hub v0.120.0 — one mail per app per outcome; the household leg has its own register (`perAppCustomerCooldownEvents`) |
| OOM storm (`app_oom_storm`) | **per APP**, operator | `…:<stack_name>` | hub v0.121.0 — no digest; the controller already sends it at most once per container run |
| backup run (`backup_run_failures`) | per RUN | `…:<run_id>` | a digest already lists every failing app; one per run |
| tiered backup (`whole_guest_backup_failed`, …) | per TIER | `…:<tier>` | the tiers fail independently and mean different things |
| everything else, incl. `crossdrive_failed` and `backup_integrity_failed` | per TYPE, per hour | — | coarse **on purpose** |
**The default is COARSE and that is deliberate.** R-97a and R-182 exist so that one full disk produces
**one** mail listing every affected app rather than one per app. Widening the grain is what makes an
operator stop reading their alerts, which is the same failure as not sending them.
**`app_start_failed` is the exception because it has no digest.** There is no `apps_down_run`
summarising a scan the way `backup_run_failures` summarises a run, so per-app is the only grain
available that does not lose alarms. Until hub v0.108.0 it was keyed per type, and **only the first
broken app per hour reached the operator** — measured 2026-08-23: `bookstack` sent at 09:27:51,
`privatebin` suppressed at 09:31:51 under `key=demo-hp:app_start_failed`.
**The mechanism is a named ALLOW-LIST (`perAppCooldownEvents`), not a payload rule**, and the
distinction is load-bearing rather than stylistic: **`crossdrive_failed` is severity `error`, reaches
the operator leg, and carries `stack_name`** through `CrossDriveDetails`. A rule of the form "if the
details carry a stack_name, split per app" would have split it, silently, and undone R-182.
`cooldownStackSuffix` therefore takes the **event type** as well as the details — an asymmetry with
its two siblings, and the reason for it is exactly this.
**Fenced act:** adding an entry to `perAppCooldownEvents` for a type whose family has a digest, or
whose coarse cooldown is deliberate. Reading the register anywhere is fine.
**Measured burst, so the volume is a number and not an impression:** three apps stopped in one scan
produced **three attempted, three sent, zero suppressed** (2026-08-23). The reference box has 8
deployed apps, so a total outage is 8 mails. The boot grace (90 s), the quiesce grace (180 s) and the
per-app edge trigger absorb reboots, backup cycles and persistently-dead apps. **This scales linearly
with app count and has no ceiling** — the condition that would reopen the question is a box large
enough that a total outage is unreadable, at which point the answer is a digest with a customer
message, not a wider cooldown.
---
## 6.3 Box alarms outside the app ladder: the tunnel and OS updates [DESIGN, hub v0.131.0, 2026-10-04]
These are **operator-only** (the household can act on none of them — except `host_restarted_after_crash`, the household's
one info line beside `host_crash_restart`; `operatorOnlyEvents`, pinned by
`TestOSUpdateEvents_OperatorOnlyExceptApplied`). They go through the same dispatcher and severity contract (§6.1):
`info` is recorded and never mailed; `warning` and `error` are mailed.
| Event | Severity | Raised when | Cleared | Pinned by |
|---|---|---|---|---|
| `tunnel_down` | error | the box's two newest host reports say the tunnel is `not_running` (more than one report cycle, 15 min) and the one before did not | the first `running` after it → `tunnel_recovered` (info) | `api/tunnel_test.go` |
| `os_update_stale` | warning | no successful OS leg for **7 days** while the switch is ON (agents that can run the leg only); the mail names the likely reason (box not reporting / the last leg's failure / no good night backup) | a successful leg | `TestAlarm_StaleLeg`, `TestAlarm_StaleNamesTheReason` |
| `os_reboot_needed` | warning | the host has needed a reboot for **14 days** (from the FIRST scanned report that said so) | a scanned pass that finds nothing (agent ≥ 0.141.1 scans the host every pass) | `TestAlarm_RebootNeeded`, `TestRebootNeeded_ClearedByAScannedPass` |
| `os_ring0_stalled` | error | ring 0 approved nothing for **7 days** in a layer while it has pending FAST-lane updates (a pending kernel does not count) | a new release | `TestAlarm_Ring0Stalled` |
| `os_not_covered` | warning | a ring-1 box has had fast-lane packages no approved release names for **14 days** | the packages are covered or gone | `TestAlarm_NotCovered` |
| `host_crash_restart` | warning | the box's crash guard reports a NEW unclean boot (a crash, a power cut or a hard reset; hub v0.132.0) | — (one per boot) | `api/crash_test.go` |
| `host_crash_guard_tripped` | error | the guard tripped: the next crash leaves the box OFF | the re-arm → `host_crash_guard_rearmed` (info) | `api/crash_test.go` |
| `host_kernel_oops` | warning | a kernel oops this boot (taint D) — the box keeps running | — (once per boot) | `api/crash_test.go` |
| `agent_behind` | warning | the box has run an agent OLDER than the vouched one for **7 days** (from when the hub first saw it behind; an unreadable version never counts; nothing vouched → nothing behind) — agents update only by a per-box signed job (R-530), so this is the "nobody signed for this box" alarm (hub v0.135.0) | the box reports the vouched agent (or newer) | `osupdates/r530_agent_alarm_test.go` |
| `floor_raise_skipped` | warning | a GLOBAL controller floor was raised and one or more boxes keep their own LOWER per-customer floor, so the raise does not move them — ONE mail naming them all (R-604, hub v0.135.0) | — (one per raise) | `web/r604_floor_held_back_test.go` |
- **`unknown` never alarms** (R-96 rule 3): a probe that could not ask is neither up nor down. An `unknown` report
breaks a `not_running` run.
- **A stopped cloudflared heals itself before the hub can see it** (measured 2026-10-04): the controller's
protected-container check recreates it within 5 minutes, and the host reports every 15. So `tunnel_down` catches
what the box cannot heal — a running container with no connection (wrong token, blocked network).
- The OS alarms are checked **hourly**, re-sent at most **once a week** while true, and forgotten when false, so the
next occurrence is announced again. The numbers are configuration (`OS_ALARM_STALE_AFTER`,
`OS_ALARM_REBOOT_AFTER`, `OS_ALARM_RING0_STALL_AFTER`, `OS_ALARM_NOT_COVERED_AFTER`, `OS_ALARM_BUNDLE_BEHIND_AFTER`,
`OS_ALARM_AGENT_BEHIND_AFTER`) — *decided by CC unattended, operator may reverse* (`11` §8.3; `09` decision 119).
---
## 6.4 A box that is not always on: the missed-backup deadline and the household's outage mail [DESIGN, hub v0.134.0, 2026-10-05]
Design home: `07` §6.1.1 (`09` decisions 109–110; CC decisions 115–116, *operator may reverse*).
- **R-872 — a box DOWN at the 05:00 deadline is judged, not skipped.** The check used to skip every customer whose node
is `down` ("they already have staleness events"), so a box down at EVERY deadline — a laptop off at night — was
never judged (measured 2026-10-05: `1 skipped (down)`). Now a down box is judged on longer lines:
`expected_dbdump_missed` after **48 h** without a `db_dump_completed`, `expected_backup_missed` after **72 h**
without a whole-guest backup in any retained host report — never for a box first seen less than 48 h ago. A box that
died last night still raises only its staleness alarm. A `disabled` box is still skipped (R-321). Pinned by
`monitor/r872_down_box_test.go`.
- **R-873 — "your server cannot be reached" reaches the HOUSEHOLD at most once per 7 days** (`node_stale`,
`node_down`, `host_stale`, `host_down`, read from the persisted notification log, so a hub restart does not reset
it). The operator still gets every edge. The recovery mail stays paired with a down mail the household actually
received (§6.2's pairing), so a held-back down mail also holds back its recovery. Pinned by
`notify/r873_liveness_weekly_test.go`.
- `backup_catchup_done` (info) is the box's "missed backup made now" line: recorded, never mailed.
---
## 7. The intent test [DESIGN, R-386 — CLOSED controller v0.223.0]
**"The customer stopped this" is asked of the FIELD THAT RECORDS IT, never inferred from the state.**
Until v0.223.0 `classifyRunStates` read `st.State == StateStopped` and assumed every stopped stack was
deliberate. It is not inferable: `aggregateState` folds `StateExited` into the stopped counter, so an
all-down stack returns `StateStopped` whatever killed it. Measured on `demo-hp` 2026-08-23:
`privatebin` stopped out of band, nine dead-app scans over four minutes, **zero events, zero banner
lines** — while a comment beside the code claimed an out-of-band stop *"still alerts"*.
`DesiredState` records the answer, has **exactly one writer** (the customer's own action), and is
tri-state:
| Intent | Verdict | Why |
|---|---|---|
| `Stopped` | **no alarm** | the customer asked |
| `Running` | **ALARM** | nobody asked — the R-386 case |
| absent (`""`) | **no alarm, and SAY SO** | UNKNOWN never means running |
**The absent case keeps the old behaviour deliberately.** Reading it as "nobody asked" would, on the
first cycle after upgrade, e-mail about every app any owner ever stopped — fleet-wide, from a field
that predates the intent being asked of it. The backfill cannot help: it seeds `Running` only from an
observed-**up** reading, so anything stopped at upgrade time stays unknown, which is precisely the
ambiguous population.
**The gap is BOUNDED, not silent.** Every such suppression sets `AppRunState.IntentUnknown`, and the
scheduler logs the names at `INFO` on the heartbeat cadence:
```
[deadapp] N stopped app(s) have NO recorded customer intent, so their dead-app alarm is
suppressed by the unknown-intent fallback (R-386): <names>. This closes itself as each app is
started or stopped through the interface.
```
**A rule without a mechanism is a wish.** Measured on `demo-hp` 2026-08-23: **0 of 8 deployed apps had
an absent intent** — the population is already empty on an exercised box; it will be larger on one
upgraded and left alone.
`failedRestart` still lifts a `Stopped` intent, and that ordering is load-bearing: the quiesce loop
stops stacks by the same path a customer does, so one it stopped and could not restart must alarm
whatever the intent says. Removing that term re-opens F-CRIT-1.
**Fenced act:** adding a `DesiredState` **writer**. Reading it anywhere is fine. Twelve of
`StopStack`'s fourteen callers are machines, so recording intent in the primitive would make a nightly
backup indistinguishable from the customer pressing Stop.
---
## 8. Direction — who a customer should be notified about at all
**[DESIGN — DIRECTION, NOT CURRENT BEHAVIOUR. Dated 2026-08-23, the operator's own framing.
Nothing in controller v0.223.0 / hub v0.107.0 implements this.]**
> **A customer should be notified only about things they can act on or are responsible for** — the
> drive they unplugged, the storage they filled. **A failed backup is our incident, not theirs.** The
> intended shape is that we detect it, we tell them we noticed and are dealing with it, and they are
> not handed an error they cannot solve. The subscription should feel like being looked after, not
> like being on call.
Today's settings page is the opposite shape: it exposes one toggle per detector and **grew from 12 to
15 in this session alone** (one new alarm, plus two compound toggles split into four). That growth is
the argument, not an aside — a page that grows by one per detector is a page that will keep asking a
household to make engineering decisions.
`app_start_failed` defaulting **off** is consistent with this direction and reversible either way; it
was ruled that way on its own merits and does not pre-judge the redesign.
**Filed as a PRODUCT DECISION, not a defect** — see the register. It is the operator's call to take
separately, and no part of it was implemented here.