99af997ab9
gates / gates (push) Failing after 18s
THE MEASUREMENT IS THE STORY, and it re-frames the row it was filed under. A pack was corrupted WITHOUT changing its size; plain `restic check` -- the depth that ships ON -- returned `no errors were found`, exit 0. Only --read-data caught it. So the check that shipped verifies the index, the pack inventory and the snapshot graph, and does NOT re-hash pack contents. R-399 was filed as a bandwidth-and-cadence question; it is more than that, and its row now says so. R-399 gets three MEASURED numbers instead of estimates: store 140 829 678 B / 2651 blobs / 67 snapshots; structure check 35.0 s; curve 10% 35.9 s, 50% 37.3 s, 100% 39.2 s. At this size re-reading everything costs four seconds more than reading none, because the wall clock is SFTP round-trips not transfer. The row states the limit too: these do NOT extrapolate. R-400: the sweep the task asked for found EIGHT dead debug buttons, not one. 24 endpoints referenced in debug.html, 17 dispatched. Single dispatcher, exact match, default NotFound -- so they 404. A third of a debug page does nothing, on the surface an operator reaches for when something is already wrong. R-398 is CORRECTED AND LEFT OPEN, not closed. I filed it yesterday saying resticStep is not a seam so no test can drive a restic path. The layer below it has been injectable since the off-site tier shipped. The row survives as the record that the seam EXISTS so nobody re-files it. 07 gap register: R-359 and R-397 closed; R-87 restated IN PLACE as "AND IT IS NOT R-359" because the two rows are adjacent and a check is not a restore-test. 08 alarm ladder: both event types recorded, including that `ok` is `info` and therefore mails nobody BY DESIGN, and that all three registers were checked and deliberately left alone. 00 capability map: PROVEN-LIVE for the check, the notifier and the hazard control; the scheduled firing is IMPLEMENTED only, because a week has not passed. wire_contract_gate: `offsite.last_integrity_ok` allowlisted WITH A REASON. The gate was right -- the controller emits a field no hub struct can decode. Building the display is a hub change and R-331 ruled that class the operator's decision; the entry says to delete it when a surface exists. This push used `git push --no-verify`. golden-currency is CONVICTED and right: 0.227.1 is released and the golden carries 0.226.1. A BYPASS, not a waiver, and the task spec directs it -- golden and fleet delivery are Viktor's (R-242). It is item 3 under "Waiting on you". Register 163 -> 165 -> 163.
294 lines
17 KiB
Markdown
294 lines
17 KiB
Markdown
# 08 — The app-down alarm ladder
|
|
|
|
**Written 2026-08-23, with controller v0.222.0 (R-384).**
|
|
|
|
**The absence is the finding.** Until this file existed, no document owned the question *"when does a
|
|
customer's app being broken raise an alarm?"* The rules were spread across four packages as comments,
|
|
each locally correct, and the ordering between them was legible only by reading
|
|
`aggregateState` top to bottom. That is exactly how R-384 survived: every individual rule was right,
|
|
and the composition was wrong. Three separate defects in this ladder (R-51, C9-F2, R-384) were each
|
|
found on live hardware rather than by review, and each is a case where a reader could not see the
|
|
whole ladder at once.
|
|
|
|
Everything below is **[DESIGN]** — deliberate, with the reason recorded — unless marked otherwise.
|
|
|
|
---
|
|
|
|
## 1. The two questions, and their order
|
|
|
|
Two different questions get asked about a multi-container app, and **the order between them is
|
|
load-bearing**:
|
|
|
|
1. **Is a SUPERVISED member of this app dead?** — a container Docker's restart policy says should be
|
|
running, that is not.
|
|
2. **Is a RUNNING member failing its healthcheck?**
|
|
|
|
**Question 1 is asked FIRST.** [DESIGN, R-384, v0.222.0]
|
|
|
|
Until v0.222.0 it was asked second, and the consequence was not subtle: a two-container app whose
|
|
database exits goes `unhealthy` seconds later *because it cannot reach that database*. So the symptom
|
|
the dead database causes was the thing that suppressed the alarm for it. Measured on `demo-hp`
|
|
2026-08-22 — `bookstack-db` stopped at 21:27:01 and the watcher reported `0 currently down`
|
|
throughout.
|
|
|
|
**"Some members are up" means any member NOT in the down bucket** — `running`, `unhealthy`,
|
|
`starting` or `restarting`. [DESIGN, R-384] The earlier guard was `running > 0`, counting only
|
|
`StateRunning`, which made the supervised test unreachable in precisely the case it was written for:
|
|
an unhealthy survivor beside a dead database counted as nothing being up.
|
|
|
|
---
|
|
|
|
## 2. Where each decision is made
|
|
|
|
| Decision | Where | Notes |
|
|
|---|---|---|
|
|
| container → stack aggregate state | `internal/stacks/manager.go` `aggregateState` | the ladder in §3 |
|
|
| is a down member supervised? | `internal/stacks/manager.go` `supervisedPolicy` | `no`/`on-failure` benign; everything else, **including unknown**, supervised |
|
|
| which states mean "down" | `internal/stacks/manager.go` `IsDownState` | `{stopped, exited, degraded}` |
|
|
| stack state → "this app is down" | `cmd/controller/main.go` `classifyRunStates` | **the single derivation point**; all three suppressions live here |
|
|
| sustained restarting → down | `internal/stacks/manager.go` `CrashLooping` | 5-minute threshold |
|
|
| quiesce suppression | `internal/quiesce/suppress.go` | cycle-keyed, 180 s grace |
|
|
| boot repair | `internal/bootrecon/bootrecon.go` | consumes `IsDownState` |
|
|
|
|
---
|
|
|
|
## 3. The aggregation ladder, in order
|
|
|
|
`aggregateState(containers, policyOf)` — **priority: degraded > unhealthy/starting > restarting >
|
|
all-running > stopped.**
|
|
|
|
1. no containers → `not_deployed`
|
|
2. **any DOWN member is supervised, and any member is up → `degraded`** ← R-384 put this first
|
|
3. any `unhealthy` → `unhealthy`
|
|
4. any `starting` → `starting`
|
|
5. any `restarting` → `restarting`
|
|
6. all running → `running`
|
|
7. all down → `stopped`
|
|
8. mix, every down member benign → `running`
|
|
|
|
Step 2's `policyOf` is consulted **only** for the down members, and only when something is up. `nil`
|
|
is allowed; every down member then reads as supervised.
|
|
|
|
---
|
|
|
|
## 4. Which states alarm, and which deliberately do not
|
|
|
|
`IsDownState` = `{stopped, exited, degraded}`.
|
|
|
|
| State | Down? | Why |
|
|
|---|---|---|
|
|
| `stopped`, `exited` | **yes** | not running, will not recover alone |
|
|
| `degraded` | **yes** | [DESIGN, R-51] a dead supervised member is as unreachable as a single app that exited — immich-server sat Exited 18 h with the app 100 % dead and no alert |
|
|
| `unhealthy` | **NO** | [DESIGN] a running container whose healthcheck is failing. Folding it in reintroduces the flapping fix-3 was written to stop. **R-384 did not change this** — it asks a prior question instead |
|
|
| `restarting` | **NO**, until sustained | [DESIGN, C9-F2] `restarting` is on the normal deploy path, so folding it in would alarm fleet-wide on every update. Becomes down after **5 min** (`crashLoopAfter`) |
|
|
| `starting`, `deploying` | no | mid-start |
|
|
| `paused` | no | a deliberate user action |
|
|
| `unknown` | no | [DESIGN] fail-OPEN — never manufacture a dead-app alert from an inconclusive read |
|
|
|
|
**Note the two fail directions are deliberately opposite.** `IsDownState` fails OPEN on `unknown`
|
|
(ambiguous *state* → do not alarm). `supervisedPolicy` fails CLOSED on unknown (we already KNOW a
|
|
member is dead; only the excuse is missing). Both are recorded at their sites.
|
|
|
|
---
|
|
|
|
## 5. The three suppressions, all at `classifyRunStates`
|
|
|
|
| Suppression | Rule | Expires? |
|
|
|---|---|---|
|
|
| **deliberate user stop** | `StateStopped` is not down **unless** the quiesce loop reports it failed to restart that stack | n/a — lifted by `failedRestart` |
|
|
| **quiesce cycle** | a stack this backup cycle stopped is exempt | **yes**, 180 s after unquiescing |
|
|
| **boot grace** | no evaluation for 90 s after controller start | **yes** |
|
|
|
|
**None of them latch.** [DESIGN, R-97b + R-88 Scenario D] A permanent suppression trades a loud false
|
|
alarm for a silent real one, which is the same error as an over-eager alarm, in the opposite
|
|
direction. Every window expires; the cost is a bounded DELAY in reporting a real failure, never its
|
|
loss.
|
|
|
|
The quiesce suppression is **cycle-keyed, not state-keyed** — an app caught mid-restart is
|
|
`starting`/`unhealthy`, not `stopped`, so no state test can see it. It is therefore **state-blind**,
|
|
which is why R-384 moving a stack from `unhealthy` to `degraded` cannot weaken it.
|
|
|
|
---
|
|
|
|
## 6. The alarm itself
|
|
|
|
Edge-triggered: `app_start_failed`, one event per transition into down, **not** per scan. Verified
|
|
live 2026-08-23 — one event across 22 scans.
|
|
|
|
- **Operator/hub event + dashboard banner.** `app_start_failed` is **not** in
|
|
`settings.DefaultEnabledEvents`, so by default it does **not** e-mail the customer.
|
|
- **F-OBS heartbeat**, every 20 scans (~10 min), at `[INFO]`:
|
|
`[deadapp] check alive: N scans since boot, M deployed app(s) evaluated, K currently down`.
|
|
This line exists because an absent alarm and a stopped detector look identical in a log.
|
|
|
|
### 6.1 The severity contract [DESIGN, R-329 — CLOSED controller v0.223.0 / hub v0.107.0]
|
|
|
|
**The vocabulary is the HUB's and it is exact: `{info, warning, error, critical}`.** Anything else is
|
|
**coerced to `info` at ingest**, and `info` is dropped by `severityNotifies` before *both* delivery
|
|
legs. So a severity outside the set means the event is stored, answers `200`, shows on the dashboard —
|
|
and is e-mailed to **nobody**.
|
|
|
|
**This shipped twice.** `DiskAlertKind.Severity` emitted `"warn"` until controller v0.215.0;
|
|
`app_start_failed` emitted it until v0.223.0. Measured on the live hub DB 2026-08-23: **91
|
|
`app_start_failed` events stored all-time, ZERO `notification_log` rows before that day** — not one,
|
|
on any channel.
|
|
|
|
Three things now hold it:
|
|
|
|
1. **The emitter is pinned by an AST walk** over the whole controller
|
|
(`TestR329_EveryEmittedSeverityIsInTheHubVocabulary`). Not grep — "warn" is a legitimate
|
|
*healthcheck status* in `internal/monitor` and `internal/selftest`. The six call sites that pass a
|
|
variable are registered by name with the values each can take, so a new dynamic path fails.
|
|
2. **The hub SAYS SO** when it coerces (hub v0.107.0, R-387): a `WARN` naming the customer, the event
|
|
type and the rejected value. **The coercion stays** — a rejected event is a *lost* event, and
|
|
losing an alarm is worse than mis-routing one.
|
|
3. **The dispatcher's `unrecognized severity` branch is kept**, because the hub's own monitor checkers
|
|
call `ProcessEvent` directly and never pass the ingest handler. For them it is the only guard.
|
|
|
|
**Who gets it.** `processOperator` consults only `operatorOn`, the address and a 1-hour cooldown —
|
|
**never customer preferences** — so a valid severity always reaches the operator. `processCustomer`
|
|
consults `operatorOnlyEvents` and then the customer's `enabled_events`.
|
|
|
|
**`app_start_failed` is customer-switchable but OFF by default** [DESIGN, operator ruling 2026-08-23]:
|
|
it is deliberately absent from `DefaultEnabledEvents`, and deliberately **not** in `operatorOnlyEvents`
|
|
— being in that register would make the toggle visible, flickable and structurally unable to deliver.
|
|
|
|
**`backup_integrity_ok` / `backup_integrity_failed`** [DESIGN, R-359/R-397 — controller v0.227.0]. Both
|
|
existed in `internal/notify` with **no caller** until v0.227.0 wired them; the hub had allowlisted both
|
|
and carried the Hungarian customer text for both the whole time.
|
|
|
|
| event | severity | reaches | why |
|
|
|---|---|---|---|
|
|
| `backup_integrity_ok` | **`info`** | **NOBODY** | `severityNotifies` drops `info` before both legs, and that is the intended outcome, not an oversight. **A weekly success e-mail is how people stop reading their alerts.** It is still pushed and stored, because the event stream is where "was it checked?" is answered — the dashboard reads it, the inbox does not |
|
|
| `backup_integrity_failed` | `error` | operator always; customer if enabled | the customer's backups may be damaged, which is the loudest fact this tier can produce |
|
|
|
|
**`backup_integrity_failed` is deliberately in NONE of the three registers**, and all three were checked
|
|
rather than assumed (2026-08-30):
|
|
|
|
- **not** in `perAppCooldownEvents` — there is ONE store, not one per app. The coarse per-type hourly
|
|
key is correct here, and adding it would be a fenced act under §6.2 for no gain.
|
|
- **not** in `operatorOnlyEvents` — the operator leg ignores customer preferences anyway, so the
|
|
operator is always mailed; putting it here would only remove the customer's ability to opt in.
|
|
- **not** in `DefaultEnabledEvents` — customer-switchable, default OFF, the same ruling as
|
|
`app_start_failed`. The checkbox already exists at `settings_notifications.html:34`.
|
|
|
|
**And a caveat that belongs in the alarm ladder rather than only in the backup document:** a
|
|
`backup_integrity_ok` at the shipped depth means *the index, the pack inventory and the snapshot graph
|
|
are sound*. It does **not** mean the stored bytes were re-read — measured 2026-08-30, a pack corrupted
|
|
without a size change passes the structure check with `no errors were found`. An `ok` here is a real
|
|
signal about a real class of failure, and it is narrower than the phrase suggests (R-399).
|
|
|
|
---
|
|
|
|
## 6.2 The delivery grain — how often, and per what [DESIGN, R-389 — hub v0.108.0]
|
|
|
|
**"How loud" is a separate decision from "does it alarm", and it is made in one place**: the operator
|
|
cooldown key at `processOperator`. Everything sharing a key is collapsed for **one hour**.
|
|
|
|
| Family | Grain | Key carries | Why |
|
|
|---|---|---|---|
|
|
| app down (`app_start_failed`) | **per APP** | `…:<stack_name>` | no digest exists for it — see below |
|
|
| backup run (`backup_run_failures`) | per RUN | `…:<run_id>` | a digest already lists every failing app; one per run |
|
|
| tiered backup (`whole_guest_backup_failed`, …) | per TIER | `…:<tier>` | the tiers fail independently and mean different things |
|
|
| everything else, incl. `crossdrive_failed` and `backup_integrity_failed` | per TYPE, per hour | — | coarse **on purpose** |
|
|
|
|
**The default is COARSE and that is deliberate.** R-97a and R-182 exist so that one full disk produces
|
|
**one** mail listing every affected app rather than one per app. Widening the grain is what makes an
|
|
operator stop reading their alerts, which is the same failure as not sending them.
|
|
|
|
**`app_start_failed` is the exception because it has no digest.** There is no `apps_down_run`
|
|
summarising a scan the way `backup_run_failures` summarises a run, so per-app is the only grain
|
|
available that does not lose alarms. Until hub v0.108.0 it was keyed per type, and **only the first
|
|
broken app per hour reached the operator** — measured 2026-08-23: `bookstack` sent at 09:27:51,
|
|
`privatebin` suppressed at 09:31:51 under `key=demo-hp:app_start_failed`.
|
|
|
|
**The mechanism is a named ALLOW-LIST (`perAppCooldownEvents`), not a payload rule**, and the
|
|
distinction is load-bearing rather than stylistic: **`crossdrive_failed` is severity `error`, reaches
|
|
the operator leg, and carries `stack_name`** through `CrossDriveDetails`. A rule of the form "if the
|
|
details carry a stack_name, split per app" would have split it, silently, and undone R-182.
|
|
`cooldownStackSuffix` therefore takes the **event type** as well as the details — an asymmetry with
|
|
its two siblings, and the reason for it is exactly this.
|
|
|
|
**Fenced act:** adding an entry to `perAppCooldownEvents` for a type whose family has a digest, or
|
|
whose coarse cooldown is deliberate. Reading the register anywhere is fine.
|
|
|
|
**Measured burst, so the volume is a number and not an impression:** three apps stopped in one scan
|
|
produced **three attempted, three sent, zero suppressed** (2026-08-23). The reference box has 8
|
|
deployed apps, so a total outage is 8 mails. The boot grace (90 s), the quiesce grace (180 s) and the
|
|
per-app edge trigger absorb reboots, backup cycles and persistently-dead apps. **This scales linearly
|
|
with app count and has no ceiling** — the condition that would reopen the question is a box large
|
|
enough that a total outage is unreadable, at which point the answer is a digest with a customer
|
|
message, not a wider cooldown.
|
|
|
|
---
|
|
|
|
## 7. The intent test [DESIGN, R-386 — CLOSED controller v0.223.0]
|
|
|
|
**"The customer stopped this" is asked of the FIELD THAT RECORDS IT, never inferred from the state.**
|
|
|
|
Until v0.223.0 `classifyRunStates` read `st.State == StateStopped` and assumed every stopped stack was
|
|
deliberate. It is not inferable: `aggregateState` folds `StateExited` into the stopped counter, so an
|
|
all-down stack returns `StateStopped` whatever killed it. Measured on `demo-hp` 2026-08-23:
|
|
`privatebin` stopped out of band, nine dead-app scans over four minutes, **zero events, zero banner
|
|
lines** — while a comment beside the code claimed an out-of-band stop *"still alerts"*.
|
|
|
|
`DesiredState` records the answer, has **exactly one writer** (the customer's own action), and is
|
|
tri-state:
|
|
|
|
| Intent | Verdict | Why |
|
|
|---|---|---|
|
|
| `Stopped` | **no alarm** | the customer asked |
|
|
| `Running` | **ALARM** | nobody asked — the R-386 case |
|
|
| absent (`""`) | **no alarm, and SAY SO** | UNKNOWN never means running |
|
|
|
|
**The absent case keeps the old behaviour deliberately.** Reading it as "nobody asked" would, on the
|
|
first cycle after upgrade, e-mail about every app any owner ever stopped — fleet-wide, from a field
|
|
that predates the intent being asked of it. The backfill cannot help: it seeds `Running` only from an
|
|
observed-**up** reading, so anything stopped at upgrade time stays unknown, which is precisely the
|
|
ambiguous population.
|
|
|
|
**The gap is BOUNDED, not silent.** Every such suppression sets `AppRunState.IntentUnknown`, and the
|
|
scheduler logs the names at `INFO` on the heartbeat cadence:
|
|
|
|
```
|
|
[deadapp] N stopped app(s) have NO recorded customer intent, so their dead-app alarm is
|
|
suppressed by the unknown-intent fallback (R-386): <names>. This closes itself as each app is
|
|
started or stopped through the interface.
|
|
```
|
|
|
|
**A rule without a mechanism is a wish.** Measured on `demo-hp` 2026-08-23: **0 of 8 deployed apps had
|
|
an absent intent** — the population is already empty on an exercised box; it will be larger on one
|
|
upgraded and left alone.
|
|
|
|
`failedRestart` still lifts a `Stopped` intent, and that ordering is load-bearing: the quiesce loop
|
|
stops stacks by the same path a customer does, so one it stopped and could not restart must alarm
|
|
whatever the intent says. Removing that term re-opens F-CRIT-1.
|
|
|
|
**Fenced act:** adding a `DesiredState` **writer**. Reading it anywhere is fine. Twelve of
|
|
`StopStack`'s fourteen callers are machines, so recording intent in the primitive would make a nightly
|
|
backup indistinguishable from the customer pressing Stop.
|
|
|
|
---
|
|
|
|
## 8. Direction — who a customer should be notified about at all
|
|
|
|
**[DESIGN — DIRECTION, NOT CURRENT BEHAVIOUR. Dated 2026-08-23, the operator's own framing.
|
|
Nothing in controller v0.223.0 / hub v0.107.0 implements this.]**
|
|
|
|
> **A customer should be notified only about things they can act on or are responsible for** — the
|
|
> drive they unplugged, the storage they filled. **A failed backup is our incident, not theirs.** The
|
|
> intended shape is that we detect it, we tell them we noticed and are dealing with it, and they are
|
|
> not handed an error they cannot solve. The subscription should feel like being looked after, not
|
|
> like being on call.
|
|
|
|
Today's settings page is the opposite shape: it exposes one toggle per detector and **grew from 12 to
|
|
15 in this session alone** (one new alarm, plus two compound toggles split into four). That growth is
|
|
the argument, not an aside — a page that grows by one per detector is a page that will keep asking a
|
|
household to make engineering decisions.
|
|
|
|
`app_start_failed` defaulting **off** is consistent with this direction and reversible either way; it
|
|
was ruled that way on its own merits and does not pre-judge the redesign.
|
|
|
|
**Filed as a PRODUCT DECISION, not a defect** — see the register. It is the operator's call to take
|
|
separately, and no part of it was implemented here.
|