f5a0aeb0b8
gates / gates (push) Failing after 1m16s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
443 lines
32 KiB
Markdown
443 lines
32 KiB
Markdown
# 08 — The app-down alarm ladder
|
||
|
||
> **How to read this document.** Where a statement is marked, it is marked like this — the same wording as
|
||
> `07-backup-architecture.md:11-17`, carried here on 2026-10-05 (R-376, the three documents written after the
|
||
> 2026-08-22 pass):
|
||
>
|
||
> - **[DESIGN]** — a decision taken. Not derived from code; the code may not implement it yet.
|
||
> - **[FACT]** — an observed property, carrying a `file:line`, a live command output or a citation.
|
||
>
|
||
> **An unmarked statement means "not yet classified", never "observed"** (R-376).
|
||
|
||
**Written 2026-08-23, with controller v0.222.0 (R-384).**
|
||
|
||
**The absence is the finding.** Until this file existed, no document owned the question *"when does a
|
||
customer's app being broken raise an alarm?"* The rules were spread across four packages as comments,
|
||
each locally correct, and the ordering between them was legible only by reading
|
||
`aggregateState` top to bottom. That is exactly how R-384 survived: every individual rule was right,
|
||
and the composition was wrong. Three separate defects in this ladder (R-51, C9-F2, R-384) were each
|
||
found on live hardware rather than by review, and each is a case where a reader could not see the
|
||
whole ladder at once.
|
||
|
||
Everything below is **[DESIGN]** — deliberate, with the reason recorded — unless marked otherwise.
|
||
|
||
---
|
||
|
||
## 1. The two questions, and their order
|
||
|
||
Two different questions get asked about a multi-container app, and **the order between them is
|
||
load-bearing**:
|
||
|
||
1. **Is a SUPERVISED member of this app dead?** — a container Docker's restart policy says should be
|
||
running, that is not.
|
||
2. **Is a RUNNING member failing its healthcheck?**
|
||
|
||
**Question 1 is asked FIRST.** [DESIGN, R-384, v0.222.0]
|
||
|
||
Until v0.222.0 it was asked second, and the consequence was not subtle: a two-container app whose
|
||
database exits goes `unhealthy` seconds later *because it cannot reach that database*. So the symptom
|
||
the dead database causes was the thing that suppressed the alarm for it. Measured on `demo-hp`
|
||
2026-08-22 — `bookstack-db` stopped at 21:27:01 and the watcher reported `0 currently down`
|
||
throughout.
|
||
|
||
**"Some members are up" means any member NOT in the down bucket** — `running`, `unhealthy`,
|
||
`starting` or `restarting`. [DESIGN, R-384] The earlier guard was `running > 0`, counting only
|
||
`StateRunning`, which made the supervised test unreachable in precisely the case it was written for:
|
||
an unhealthy survivor beside a dead database counted as nothing being up.
|
||
|
||
---
|
||
|
||
## 2. Where each decision is made
|
||
|
||
| Decision | Where | Notes |
|
||
|---|---|---|
|
||
| container → stack aggregate state | `internal/stacks/manager.go` `aggregateState` | the ladder in §3 |
|
||
| is a down member supervised? | `internal/stacks/manager.go` `supervisedPolicy` | `no`/`on-failure` benign; everything else, **including unknown**, supervised |
|
||
| which states mean "down" | `internal/stacks/manager.go` `IsDownState` | `{stopped, exited, degraded}` |
|
||
| stack state → "this app is down" | `cmd/controller/main.go` `classifyRunStates` | **the single derivation point**; all three suppressions live here |
|
||
| sustained restarting → down | `internal/stacks/manager.go` `CrashLooping` | 5-minute threshold |
|
||
| quiesce suppression | `internal/quiesce/suppress.go` | cycle-keyed, 180 s grace |
|
||
| boot repair | `internal/bootrecon/bootrecon.go` | consumes `IsDownState` |
|
||
|
||
---
|
||
|
||
## 3. The aggregation ladder, in order
|
||
|
||
`aggregateState(containers, policyOf)` — **priority: degraded > unhealthy/starting > restarting >
|
||
all-running > stopped.**
|
||
|
||
1. no containers → `not_deployed`
|
||
2. **any DOWN member is supervised, and any member is up → `degraded`** ← R-384 put this first
|
||
3. any `unhealthy` → `unhealthy`
|
||
4. any `starting` → `starting`
|
||
5. any `restarting` → `restarting`
|
||
6. all running → `running`
|
||
7. all down → `stopped`
|
||
8. mix, every down member benign → `running`
|
||
|
||
Step 2's `policyOf` is consulted **only** for the down members, and only when something is up. `nil`
|
||
is allowed; every down member then reads as supervised.
|
||
|
||
---
|
||
|
||
## 4. Which states alarm, and which deliberately do not
|
||
|
||
`IsDownState` = `{stopped, exited, degraded}`.
|
||
|
||
| State | Down? | Why |
|
||
|---|---|---|
|
||
| `stopped`, `exited` | **yes** | not running, will not recover alone |
|
||
| `degraded` | **yes** | [DESIGN, R-51] a dead supervised member is as unreachable as a single app that exited — immich-server sat Exited 18 h with the app 100 % dead and no alert |
|
||
| `unhealthy` | **NO** | [DESIGN] a running container whose healthcheck is failing. Folding it in reintroduces the flapping fix-3 was written to stop. **R-384 did not change this** — it asks a prior question instead |
|
||
| `restarting` | **NO**, until sustained | [DESIGN, C9-F2] `restarting` is on the normal deploy path, so folding it in would alarm fleet-wide on every update. Becomes down after **5 min** (`crashLoopAfter`) of UNINTERRUPTED `restarting` — **measured 2026-09-24: a container that reads `running` for a moment between restarts resets the clock and never gets there (gokapi, 385 restarts, „0 currently down"; R-667)** **Since controller v0.269.0 the box counts `RestartCount` instead and STOPS the app (§6.2, decision 28).** |
|
||
| `starting`, `deploying` | no | mid-start |
|
||
| `paused` | no | a deliberate user action |
|
||
| `unknown` | no | [DESIGN] fail-OPEN — never manufacture a dead-app alert from an inconclusive read |
|
||
|
||
**Note the two fail directions are deliberately opposite.** `IsDownState` fails OPEN on `unknown`
|
||
(ambiguous *state* → do not alarm). `supervisedPolicy` fails CLOSED on unknown (we already KNOW a
|
||
member is dead; only the excuse is missing). Both are recorded at their sites.
|
||
|
||
---
|
||
|
||
## 5. The four suppressions, all at `classifyRunStates`
|
||
|
||
| Suppression | Rule | Expires? |
|
||
|---|---|---|
|
||
| **deliberate user stop** | `StateStopped` is not down **unless** the quiesce loop reports it failed to restart that stack | n/a — lifted by `failedRestart` |
|
||
| **quiesce cycle** | a stack this backup cycle stopped is exempt | **yes**, 180 s after unquiescing |
|
||
| **boot grace** | no evaluation for 90 s after controller start | **yes** |
|
||
| **update hold** (v0.268.0, R-660) | an app HELD after a failed update is stopped by the product and has its own event (`app_update_held`, and `app_hold_no_whole_copy` when no copy brings it back whole); it is not "down" | **yes** — lifted with the hold (a restore, or the operator). A RESTORE hold (R-379) is deliberately not in the set: it has no event of its own |
|
||
|
||
**None of them latch.** [DESIGN, R-97b + R-88 Scenario D] A permanent suppression trades a loud false
|
||
alarm for a silent real one, which is the same error as an over-eager alarm, in the opposite
|
||
direction. Every window expires; the cost is a bounded DELAY in reporting a real failure, never its
|
||
loss.
|
||
|
||
The quiesce suppression is **cycle-keyed, not state-keyed** — an app caught mid-restart is
|
||
`starting`/`unhealthy`, not `stopped`, so no state test can see it. It is therefore **state-blind**,
|
||
which is why R-384 moving a stack from `unhealthy` to `degraded` cannot weaken it.
|
||
|
||
---
|
||
|
||
## 6. The alarm itself
|
||
|
||
Edge-triggered: `app_start_failed`, one event per transition into down, **not** per scan. Verified
|
||
live 2026-08-23 — one event across 22 scans.
|
||
|
||
- **Operator/hub event + dashboard banner.** `app_start_failed` is **not** in
|
||
`settings.DefaultEnabledEvents`, so by default it does **not** e-mail the customer.
|
||
- **F-OBS heartbeat**, every 20 scans (~10 min), at `[INFO]`:
|
||
`[deadapp] check alive: N scans since boot, M deployed app(s) evaluated, K currently down`.
|
||
This line exists because an absent alarm and a stopped detector look identical in a log.
|
||
|
||
### 6.1 The severity contract [DESIGN, R-329 — CLOSED controller v0.223.0 / hub v0.107.0]
|
||
|
||
**The vocabulary is the HUB's and it is exact: `{info, warning, error, critical}`.** Anything else is
|
||
**coerced to `info` at ingest**, and `info` is dropped by `severityNotifies` before *both* delivery
|
||
legs. So a severity outside the set means the event is stored, answers `200`, shows on the dashboard —
|
||
and is e-mailed to **nobody**.
|
||
|
||
**This shipped twice.** `DiskAlertKind.Severity` emitted `"warn"` until controller v0.215.0;
|
||
`app_start_failed` emitted it until v0.223.0. Measured on the live hub DB 2026-08-23: **91
|
||
`app_start_failed` events stored all-time, ZERO `notification_log` rows before that day** — not one,
|
||
on any channel.
|
||
|
||
Three things now hold it:
|
||
|
||
1. **The emitter is pinned by an AST walk** over the whole controller
|
||
(`TestR329_EveryEmittedSeverityIsInTheHubVocabulary`). Not grep — "warn" is a legitimate
|
||
*healthcheck status* in `internal/monitor` and `internal/selftest`. The six call sites that pass a
|
||
variable are registered by name with the values each can take, so a new dynamic path fails.
|
||
2. **The hub SAYS SO** when it coerces (hub v0.107.0, R-387): a `WARN` naming the customer, the event
|
||
type and the rejected value. **The coercion stays** — a rejected event is a *lost* event, and
|
||
losing an alarm is worse than mis-routing one.
|
||
3. **The dispatcher's `unrecognized severity` branch is kept**, because the hub's own monitor checkers
|
||
call `ProcessEvent` directly and never pass the ingest handler. For them it is the only guard.
|
||
|
||
**Who gets it.** `processOperator` consults only `operatorOn`, the address and a 1-hour cooldown —
|
||
**never customer preferences** — so a valid severity always reaches the operator. `processCustomer`
|
||
consults `operatorOnlyEvents` and then the customer's `enabled_events`.
|
||
|
||
**`app_start_failed` is customer-switchable but OFF by default** [DESIGN, operator ruling 2026-08-23]:
|
||
it is deliberately absent from `DefaultEnabledEvents`, and deliberately **not** in `operatorOnlyEvents`
|
||
— being in that register would make the toggle visible, flickable and structurally unable to deliver.
|
||
|
||
**`backup_integrity_ok` / `backup_integrity_failed`** [DESIGN, R-359/R-397 — controller v0.227.0]. Both
|
||
existed in `internal/notify` with **no caller** until v0.227.0 wired them; the hub had allowlisted both
|
||
and carried the Hungarian customer text for both the whole time.
|
||
|
||
| event | severity | reaches | why |
|
||
|---|---|---|---|
|
||
| `backup_integrity_ok` | **`info`** | **NOBODY** | `severityNotifies` drops `info` before both legs, and that is the intended outcome, not an oversight. **A weekly success e-mail is how people stop reading their alerts.** It is still pushed and stored, because the event stream is where "was it checked?" is answered — the dashboard reads it, the inbox does not |
|
||
| `backup_integrity_failed` | `error` | operator always; customer if enabled | the customer's backups may be damaged, which is the loudest fact this tier can produce |
|
||
|
||
**`backup_integrity_failed` is deliberately in NONE of the three registers**, and all three were checked
|
||
rather than assumed (2026-08-30):
|
||
|
||
- **not** in `perAppCooldownEvents` — there is ONE store, not one per app. The coarse per-type hourly
|
||
key is correct here, and adding it would be a fenced act under §6.2 for no gain.
|
||
- **not** in `operatorOnlyEvents` — the operator leg ignores customer preferences anyway, so the
|
||
operator is always mailed; putting it here would only remove the customer's ability to opt in.
|
||
- **not** in `DefaultEnabledEvents` — customer-switchable, default OFF, the same ruling as
|
||
`app_start_failed`. The checkbox already exists at `settings_notifications.html:34`.
|
||
|
||
**And a caveat that belongs in the alarm ladder rather than only in the backup document:** a
|
||
`backup_integrity_ok` at the shipped depth means *the index, the pack inventory and the snapshot graph
|
||
are sound*. It does **not** mean the stored bytes were re-read — measured 2026-08-30, a pack corrupted
|
||
without a size change passes the structure check with `no errors were found`. An `ok` here is a real
|
||
signal about a real class of failure, and it is narrower than the phrase suggests (R-399).
|
||
|
||
---
|
||
|
||
## 6.2 The delivery grain — how often, and per what [DESIGN, R-389 — hub v0.108.0]
|
||
|
||
**"How loud" is a separate decision from "does it alarm", and it is made in one place**: the operator
|
||
cooldown key at `processOperator`. Everything sharing a key is collapsed for **one hour**.
|
||
|
||
**Operator ruling 2026-09-15 (decision A) — „the box is down" skips the quiet hour.** `node_stale`,
|
||
`node_down` and `node_recovered` no longer share the one-hour operator cooldown; they keep a **5-minute
|
||
dedupe** on the same key, so a flapping link cannot mail every sweep (`operatorCooldownFor`, hub
|
||
v0.114.0, pinned by `TestOperatorCooldown_NodeLivenessBypassesQuietHour`). **This reverses the design
|
||
above for those three types only**, and it is a ruling, not a defect fix. The reason is BIGNIGHT F9
|
||
(2026-09-14): the controller was dead for 33 minutes; its `node_stale` mail was suppressed because F8's
|
||
`node_stale` had used the hour 39 minutes earlier, and the later `node_recovered` mail was suppressed
|
||
the same way. **EXTENDED by operator ruling 2, 2026-09-16 (hub v0.115.0, R-529): `host_stale`, `host_down` and
|
||
`host_recovered` join the bypass**, with the same 5-minute dedupe. They are the same sentence about the
|
||
same box — the agent's dead-man's-switch rather than the controller's — and leaving them on the hour
|
||
would have kept exactly the F9 silence on the host plane. Everything else keeps the hour (pinned by
|
||
`TestOperatorCooldown_NodeLivenessBypassesQuietHour`, which also asserts an unrelated type still waits).
|
||
|
||
**Operator ruling A, 2026-09-17 (R-549) — „the box went quiet" waits THREE report cycles, not two.**
|
||
The staleness threshold (`alerting.stale_threshold`, configuration in `manifests/hub.yaml`) moves from
|
||
**30 m to 45 m**; `node_down` and `host_down` follow at 2× = **90 m**. The reason is chaos night round 9
|
||
(`audits/DRILL-chaos-night-2026-09-17.md`): report cadence 15 m; a failed push is retried for ~100 s and
|
||
then given up (correctly — a report is a snapshot); the measured gap to the next good report was
|
||
**29 m 59 s** against a 30-minute threshold. One missed push spent the entire budget, so ordinary jitter
|
||
would page the operator about a box that was healthy and had already repaired itself. **The cost,
|
||
stated:** a truly dead box now pages 15 minutes later (45 m instead of 30), and „the box is down" 30
|
||
minutes later (90 m instead of 60). **One value, everywhere:** both staleness checkers, the host status and
|
||
— since hub v0.117.0 — the customer status on the dashboard read the same threshold (`controllerStatus`
|
||
hardcoded 30 m / 1 h until then; pinned by `TestControllerStatus_FollowsConfiguredThreshold`). A running
|
||
hub prints it at startup (`node_stale after 45m0s, node_down after 1h30m0s`).
|
||
|
||
**An out-of-memory storm gets its own, louder rung (R-636, controller v0.265.0 / hub v0.121.0, 2026-09-23).**
|
||
`app_oom` stays exactly as it was: `warning`, operator-only, ONCE per container run (R-514) — the once
|
||
is what stops a crash loop from mailing thousands of times (R-629). Beside it, `app_oom_storm`:
|
||
|
||
| event | severity | who | minted by | why that audience |
|
||
|---|---|---|---|---|
|
||
| `app_oom_storm` | **error** | **operator only** | controller v0.265.0, when the kernel's `oom_kill` counter of the SAME container run rises by **≥ 20 within 30 min**; once per run | raw container names and memory figures; the household's side is the dashboard tag |
|
||
|
||
**Why a counter and not the flag:** Docker's `OOMKilled` is sticky — true for the whole run after ONE
|
||
kill — so "the key re-fired N times" measures only how long ago the first kill was. The kernel's
|
||
`memory.events` `oom_kill` counts kills. **Why 20 in 30 minutes:** RomM's measured rate on 2026-09-22
|
||
was 4,530 kills in six hours ≈ 375 per 30 min; a hiccup is 1–3. Live on 9202 (2026-09-23): RomM at 320M
|
||
reached 21 kills 2.5 min after start and sent ONE storm; at 49 kills still one. **Not an app-down state:**
|
||
the app still reads `running` and `IsDownState` is unchanged (§4). **Limit:** a container whose
|
||
`OOMKilled` flag stays false in an LXC guest (R-528) is never read, so it never storms.
|
||
|
||
**The box STOPS a crash loop or an out-of-memory storm (decision 28 of `09` §3, R-667, controller v0.269.0 /
|
||
hub v0.123.0, 2026-09-24).** Alarming was not enough: gokapi crash-looped for hours at 385 → 546 restarts
|
||
while the ladder said „0 currently down" (the `restarting` clock resets, §4). Now, besides the alarms above:
|
||
|
||
| rule | threshold | what the box does | event |
|
||
|---|---|---|---|
|
||
| crash loop | **≥ 6 restarts within 10 min**, from Docker's `RestartCount` summed per app, sampled every scan (a drop = the container was recreated → the window starts again) | stops the app, records an `unhealthy_stop` hold (the same store as every hold, so no start path revives it), the page says so with a Start button | `app_stopped_unhealthy`, **warning**, **the household AND the operator** (default on, seeded add-only) |
|
||
| out-of-memory storm | the `app_oom_storm` rule above (≥ 20 kernel kills within 30 min) | the same stop and hold | `app_oom_storm` (operator) + `app_stopped_unhealthy` |
|
||
| Start pressed | — | lifts the hold: **one more try** | — |
|
||
| stopped again within 24 h | — | stops it again; the sentence now says Felhom support is informed | `app_stopped_unhealthy` (repeat wording) |
|
||
|
||
**Why 6 in 10 minutes and not 10:** Docker's restart back-off caps a steady loop at about ONE restart a
|
||
minute (gokapi: 539 → 546 in 7 min), so 10 in 10 sits on the edge and misses a steady loop; a FRESH
|
||
container restarts fast (9 in 32 s after a Start, measured). No healthy app in any drill evidence (1,831
|
||
harness samples, 40 live containers) restarted more than once on a first start; the one borderline shape
|
||
is immich's first-start import (12 restarts, 2026-09-17 — broken that night; R-676 watches it).
|
||
**Never judged:** an app that is deploying, updating, held, or stopped by a backup or a quiesce.
|
||
**Precisely (read from source, controller v0.271.0, 2026-09-25):** "updating" covers an automatic update's
|
||
whole step — pull, start, verify AND its undo — because `Updating` stays true until the step ends
|
||
(`TestD28_NoCrashLoopStopDuringAnAutomaticStep`; the leg's own pages are therefore never stopped mid-step).
|
||
"Deploying" does **not** cover a deploy's FIRST START: the flag clears when `compose up -d` returns, and the
|
||
app is sampled from then on — a first start that restarts ≥ 6 times in 10 min is stopped (R-676). An automatic
|
||
update is a product stop in the §5 sense: a step's restarts are the product's own, never an alarm.
|
||
**Measured in the chaos hour (night 2026-09-24):** stopped at +185 s after a power cut mid-storm (the
|
||
counting starts again after the boot), +116 s for a crash loop under a backup run, +102 s for a storm with
|
||
the disk 1 GB above the floor.
|
||
|
||
**A thin pool is critical at 90 % (R-672, hub v0.124.0 + agent v0.133.0).** `storage_fill_*` judges an
|
||
`lvmthin` target on the worse of data and metadata fill, warning 85 %, **critical 90 %** (other storages keep
|
||
90/95 %). A full thin pool does not only refuse writes: every guest on it remounts read-only (demo-hp
|
||
2026-09-24 — 95 % → 100 % in about a minute, 9201 read-only five minutes later). The agent requests an
|
||
immediate host report when a pool crosses 90 %, so the alarm fires in seconds, not at the next 15-minute
|
||
report. Grain: **per pool, 6 hours** (`customer:type:host/storage`) — the old `customer:type`, 1 hour let one
|
||
pool silence another. The hub HAD alarmed on 2026-09-24 (operator mail at 100 %), but late and at the generic
|
||
bands.
|
||
|
||
**Two event types added 2026-09-17, with who receives them:**
|
||
|
||
| event | severity | who | minted by | why that audience |
|
||
|---|---|---|---|---|
|
||
| `controller_slow_crashloop` | warning | **operator only** | hub, from the agent's `slow_crashloop_since` moving (agent v0.132.0, hub v0.117.0) | host ids and vmids; the household's side of it is the dashboard coming back each time |
|
||
| `restore_interrupted` | warning | **the household** (and the operator) | controller v0.246.0 at startup, once per interruption | they pressed restore, were told it started, and can run it again — Hungarian `customerMessages` entry |
|
||
|
||
| Family | Grain | Key carries | Why |
|
||
|---|---|---|---|
|
||
| app down (`app_start_failed`) | **per APP** | `…:<stack_name>` | no digest exists for it — see below |
|
||
| update outcome (`app_update_undone`, `app_update_held`) | **per APP**, both legs | `…:<stack_name>` | hub v0.120.0 — one mail per app per outcome; the household leg has its own register (`perAppCustomerCooldownEvents`) |
|
||
| OOM storm (`app_oom_storm`) | **per APP**, operator | `…:<stack_name>` | hub v0.121.0 — no digest; the controller already sends it at most once per container run |
|
||
| backup run (`backup_run_failures`) | per RUN | `…:<run_id>` | a digest already lists every failing app; one per run |
|
||
| tiered backup (`whole_guest_backup_failed`, …) | per TIER | `…:<tier>` | the tiers fail independently and mean different things |
|
||
| everything else, incl. `crossdrive_failed` and `backup_integrity_failed` | per TYPE, per hour | — | coarse **on purpose** |
|
||
|
||
**The default is COARSE and that is deliberate.** R-97a and R-182 exist so that one full disk produces
|
||
**one** mail listing every affected app rather than one per app. Widening the grain is what makes an
|
||
operator stop reading their alerts, which is the same failure as not sending them.
|
||
|
||
**`app_start_failed` is the exception because it has no digest.** There is no `apps_down_run`
|
||
summarising a scan the way `backup_run_failures` summarises a run, so per-app is the only grain
|
||
available that does not lose alarms. Until hub v0.108.0 it was keyed per type, and **only the first
|
||
broken app per hour reached the operator** — measured 2026-08-23: `bookstack` sent at 09:27:51,
|
||
`privatebin` suppressed at 09:31:51 under `key=demo-hp:app_start_failed`.
|
||
|
||
**The mechanism is a named ALLOW-LIST (`perAppCooldownEvents`), not a payload rule**, and the
|
||
distinction is load-bearing rather than stylistic: **`crossdrive_failed` is severity `error`, reaches
|
||
the operator leg, and carries `stack_name`** through `CrossDriveDetails`. A rule of the form "if the
|
||
details carry a stack_name, split per app" would have split it, silently, and undone R-182.
|
||
`cooldownStackSuffix` therefore takes the **event type** as well as the details — an asymmetry with
|
||
its two siblings, and the reason for it is exactly this.
|
||
|
||
**Fenced act:** adding an entry to `perAppCooldownEvents` for a type whose family has a digest, or
|
||
whose coarse cooldown is deliberate. Reading the register anywhere is fine.
|
||
|
||
**Measured burst, so the volume is a number and not an impression:** three apps stopped in one scan
|
||
produced **three attempted, three sent, zero suppressed** (2026-08-23). The reference box has 8
|
||
deployed apps, so a total outage is 8 mails. The boot grace (90 s), the quiesce grace (180 s) and the
|
||
per-app edge trigger absorb reboots, backup cycles and persistently-dead apps. **This scales linearly
|
||
with app count and has no ceiling** — the condition that would reopen the question is a box large
|
||
enough that a total outage is unreadable, at which point the answer is a digest with a customer
|
||
message, not a wider cooldown.
|
||
|
||
---
|
||
|
||
## 6.3 Box alarms outside the app ladder: the tunnel and OS updates [DESIGN, hub v0.131.0, 2026-10-04]
|
||
|
||
These are **operator-only** (the household can act on none of them — except `host_restarted_after_crash`, the household's
|
||
one info line beside `host_crash_restart`; `operatorOnlyEvents`, pinned by
|
||
`TestOSUpdateEvents_OperatorOnlyExceptApplied`). They go through the same dispatcher and severity contract (§6.1):
|
||
`info` is recorded and never mailed; `warning` and `error` are mailed.
|
||
|
||
| Event | Severity | Raised when | Cleared | Pinned by |
|
||
|---|---|---|---|---|
|
||
| `tunnel_down` | error | the box's two newest host reports say the tunnel is `not_running` (more than one report cycle, 15 min) and the one before did not | the first `running` after it → `tunnel_recovered` (info) | `api/tunnel_test.go` |
|
||
| `os_update_stale` | warning | no successful OS leg for **7 days** while the switch is ON (agents that can run the leg only); the mail names the likely reason (box not reporting / the last leg's failure / no good night backup) | a successful leg | `TestAlarm_StaleLeg`, `TestAlarm_StaleNamesTheReason` |
|
||
| `os_reboot_needed` | warning | the host has needed a reboot for **14 days** (from the FIRST scanned report that said so) | a scanned pass that finds nothing (agent ≥ 0.141.1 scans the host every pass) | `TestAlarm_RebootNeeded`, `TestRebootNeeded_ClearedByAScannedPass` |
|
||
| `os_ring0_stalled` | error | ring 0 approved nothing for **7 days** in a layer while it has pending FAST-lane updates (a pending kernel does not count) | a new release | `TestAlarm_Ring0Stalled` |
|
||
| `os_not_covered` | warning | a ring-1 box has had fast-lane packages no approved release names for **14 days** | the packages are covered or gone | `TestAlarm_NotCovered` |
|
||
| `host_crash_restart` | warning | the box's crash guard reports a NEW unclean boot (a crash, a power cut or a hard reset; hub v0.132.0) | — (one per boot) | `api/crash_test.go` |
|
||
| `host_crash_guard_tripped` | error | the guard tripped: the next crash leaves the box OFF | the re-arm → `host_crash_guard_rearmed` (info) | `api/crash_test.go` |
|
||
| `host_kernel_oops` | warning | a kernel oops this boot (taint D) — the box keeps running | — (once per boot) | `api/crash_test.go` |
|
||
| `agent_behind` | warning | the box has run an agent OLDER than the vouched one for **7 days** (from when the hub first saw it behind; an unreadable version never counts; nothing vouched → nothing behind) — agents update only by a per-box signed job (R-530), so this is the "nobody signed for this box" alarm (hub v0.135.0) | the box reports the vouched agent (or newer) | `osupdates/r530_agent_alarm_test.go` |
|
||
| `floor_raise_skipped` | warning | a GLOBAL controller floor was raised and one or more boxes keep their own LOWER per-customer floor, so the raise does not move them — ONE mail naming them all (R-604, hub v0.135.0) | — (one per raise) | `web/r604_floor_held_back_test.go` |
|
||
|
||
- **`unknown` never alarms** (R-96 rule 3): a probe that could not ask is neither up nor down. An `unknown` report
|
||
breaks a `not_running` run.
|
||
- **A stopped cloudflared heals itself before the hub can see it** (measured 2026-10-04): the controller's
|
||
protected-container check recreates it within 5 minutes, and the host reports every 15. So `tunnel_down` catches
|
||
what the box cannot heal — a running container with no connection (wrong token, blocked network).
|
||
- The OS alarms are checked **hourly**, re-sent at most **once a week** while true, and forgotten when false, so the
|
||
next occurrence is announced again. The numbers are configuration (`OS_ALARM_STALE_AFTER`,
|
||
`OS_ALARM_REBOOT_AFTER`, `OS_ALARM_RING0_STALL_AFTER`, `OS_ALARM_NOT_COVERED_AFTER`, `OS_ALARM_BUNDLE_BEHIND_AFTER`,
|
||
`OS_ALARM_AGENT_BEHIND_AFTER`) — *decided by CC unattended, operator may reverse* (`11` §8.3; `09` decision 119).
|
||
|
||
---
|
||
|
||
## 6.4 A box that is not always on: the missed-backup deadline and the household's outage mail [DESIGN, hub v0.134.0, 2026-10-05]
|
||
|
||
Design home: `07` §6.1.1 (`09` decisions 109–110; CC decisions 115–116, *operator may reverse*).
|
||
|
||
- **R-872 — a box DOWN at the 05:00 deadline is judged, not skipped.** The check used to skip every customer whose node
|
||
is `down` ("they already have staleness events"), so a box down at EVERY deadline — a laptop off at night — was
|
||
never judged (measured 2026-10-05: `1 skipped (down)`). Now a down box is judged on longer lines:
|
||
`expected_dbdump_missed` after **48 h** without a `db_dump_completed`, `expected_backup_missed` after **72 h**
|
||
without a whole-guest backup in any retained host report — never for a box first seen less than 48 h ago. A box that
|
||
died last night still raises only its staleness alarm. A `disabled` box is still skipped (R-321). Pinned by
|
||
`monitor/r872_down_box_test.go`.
|
||
- **R-873 — "your server cannot be reached" reaches the HOUSEHOLD at most once per 7 days** (`node_stale`,
|
||
`node_down`, `host_stale`, `host_down`, read from the persisted notification log, so a hub restart does not reset
|
||
it). The operator still gets every edge. The recovery mail stays paired with a down mail the household actually
|
||
received (§6.2's pairing), so a held-back down mail also holds back its recovery. Pinned by
|
||
`notify/r873_liveness_weekly_test.go`.
|
||
- `backup_catchup_done` (info) is the box's "missed backup made now" line: recorded, never mailed.
|
||
|
||
---
|
||
|
||
## 7. The intent test [DESIGN, R-386 — CLOSED controller v0.223.0]
|
||
|
||
**"The customer stopped this" is asked of the FIELD THAT RECORDS IT, never inferred from the state.**
|
||
|
||
Until v0.223.0 `classifyRunStates` read `st.State == StateStopped` and assumed every stopped stack was
|
||
deliberate. It is not inferable: `aggregateState` folds `StateExited` into the stopped counter, so an
|
||
all-down stack returns `StateStopped` whatever killed it. Measured on `demo-hp` 2026-08-23:
|
||
`privatebin` stopped out of band, nine dead-app scans over four minutes, **zero events, zero banner
|
||
lines** — while a comment beside the code claimed an out-of-band stop *"still alerts"*.
|
||
|
||
`DesiredState` records the answer, has **exactly one writer** (the customer's own action), and is
|
||
tri-state:
|
||
|
||
| Intent | Verdict | Why |
|
||
|---|---|---|
|
||
| `Stopped` | **no alarm** | the customer asked |
|
||
| `Running` | **ALARM** | nobody asked — the R-386 case |
|
||
| absent (`""`) | **no alarm, and SAY SO** | UNKNOWN never means running |
|
||
|
||
**The absent case keeps the old behaviour deliberately.** Reading it as "nobody asked" would, on the
|
||
first cycle after upgrade, e-mail about every app any owner ever stopped — fleet-wide, from a field
|
||
that predates the intent being asked of it. The backfill cannot help: it seeds `Running` only from an
|
||
observed-**up** reading, so anything stopped at upgrade time stays unknown, which is precisely the
|
||
ambiguous population.
|
||
|
||
**The gap is BOUNDED, not silent.** Every such suppression sets `AppRunState.IntentUnknown`, and the
|
||
scheduler logs the names at `INFO` on the heartbeat cadence:
|
||
|
||
```
|
||
[deadapp] N stopped app(s) have NO recorded customer intent, so their dead-app alarm is
|
||
suppressed by the unknown-intent fallback (R-386): <names>. This closes itself as each app is
|
||
started or stopped through the interface.
|
||
```
|
||
|
||
**A rule without a mechanism is a wish.** Measured on `demo-hp` 2026-08-23: **0 of 8 deployed apps had
|
||
an absent intent** — the population is already empty on an exercised box; it will be larger on one
|
||
upgraded and left alone.
|
||
|
||
`failedRestart` still lifts a `Stopped` intent, and that ordering is load-bearing: the quiesce loop
|
||
stops stacks by the same path a customer does, so one it stopped and could not restart must alarm
|
||
whatever the intent says. Removing that term re-opens F-CRIT-1.
|
||
|
||
**Fenced act:** adding a `DesiredState` **writer**. Reading it anywhere is fine. Twelve of
|
||
`StopStack`'s fourteen callers are machines, so recording intent in the primitive would make a nightly
|
||
backup indistinguishable from the customer pressing Stop.
|
||
|
||
---
|
||
|
||
## 8. Direction — who a customer should be notified about at all
|
||
|
||
**[DESIGN — DIRECTION, NOT CURRENT BEHAVIOUR. Dated 2026-08-23, the operator's own framing.
|
||
Nothing in controller v0.223.0 / hub v0.107.0 implements this.]**
|
||
|
||
> **A customer should be notified only about things they can act on or are responsible for** — the
|
||
> drive they unplugged, the storage they filled. **A failed backup is our incident, not theirs.** The
|
||
> intended shape is that we detect it, we tell them we noticed and are dealing with it, and they are
|
||
> not handed an error they cannot solve. The subscription should feel like being looked after, not
|
||
> like being on call.
|
||
|
||
Today's settings page is the opposite shape: it exposes one toggle per detector and **grew from 12 to
|
||
15 in this session alone** (one new alarm, plus two compound toggles split into four). That growth is
|
||
the argument, not an aside — a page that grows by one per detector is a page that will keep asking a
|
||
household to make engineering decisions.
|
||
|
||
`app_start_failed` defaulting **off** is consistent with this direction and reversible either way; it
|
||
was ruled that way on its own merits and does not pre-judge the redesign.
|
||
|
||
**Filed as a PRODUCT DECISION, not a defect** — see the register. It is the operator's call to take
|
||
separately, and no part of it was implemented here.
|