55274d5ef3
gates / gates (push) Successful in 17s
The gate failed only on `released > baked`, so it could catch a forgotten bake and nothing else. A golden AHEAD of the record passed silently - and that is how controller 0.221.1 was built, baked AND vouched while the newest CHANGELOG heading still read v0.221.0, with every gate green. Reproduced on the real history: newest released 0.221.0 / newest golden baked 0.221.1 -> exit 0. The gate now asks whether the version being shipped is WRITTEN DOWN: the baked version must have its own `## vX.Y.Z` heading anywhere in the CHANGELOG. Membership rather than `baked > released` deliberately - a comparison against the newest heading alone goes green the moment any later entry is written, leaving the unrecorded version permanently unrecorded. INCONCLUSIVE (exit 2) preserved; every refusal names a reason and a route. Red-proofed both directions: old gate/old record exit 0, new gate/old record exit 1, new gate/fixed record exit 0, absent clone exit 2, post-bake exit 0. 08-alarm-ladder.md is new, and its absence was itself the finding: no document owned "when does a broken app raise an alarm?". The rules lived as comments in four packages, each locally correct, with the ordering between them legible only by reading one function top to bottom - which is how R-384 survived review. R-383 and R-384 closed into CLOSED-ITEMS with their rules kept. R-385 filed closed. R-386 filed OPEN: a single-container app stopped out of band raises no alarm, and a comment claims the opposite - measured live, 9 scans, 0 events, against a positive control from the same box 17 minutes earlier. Not fixed here. Golden 0.222.0 baked and published; vouching is the operator's act.
142 lines
7.6 KiB
Markdown
142 lines
7.6 KiB
Markdown
# 08 — The app-down alarm ladder
|
|
|
|
**Written 2026-08-23, with controller v0.222.0 (R-384).**
|
|
|
|
**The absence is the finding.** Until this file existed, no document owned the question *"when does a
|
|
customer's app being broken raise an alarm?"* The rules were spread across four packages as comments,
|
|
each locally correct, and the ordering between them was legible only by reading
|
|
`aggregateState` top to bottom. That is exactly how R-384 survived: every individual rule was right,
|
|
and the composition was wrong. Three separate defects in this ladder (R-51, C9-F2, R-384) were each
|
|
found on live hardware rather than by review, and each is a case where a reader could not see the
|
|
whole ladder at once.
|
|
|
|
Everything below is **[DESIGN]** — deliberate, with the reason recorded — unless marked otherwise.
|
|
|
|
---
|
|
|
|
## 1. The two questions, and their order
|
|
|
|
Two different questions get asked about a multi-container app, and **the order between them is
|
|
load-bearing**:
|
|
|
|
1. **Is a SUPERVISED member of this app dead?** — a container Docker's restart policy says should be
|
|
running, that is not.
|
|
2. **Is a RUNNING member failing its healthcheck?**
|
|
|
|
**Question 1 is asked FIRST.** [DESIGN, R-384, v0.222.0]
|
|
|
|
Until v0.222.0 it was asked second, and the consequence was not subtle: a two-container app whose
|
|
database exits goes `unhealthy` seconds later *because it cannot reach that database*. So the symptom
|
|
the dead database causes was the thing that suppressed the alarm for it. Measured on `demo-hp`
|
|
2026-08-22 — `bookstack-db` stopped at 21:27:01 and the watcher reported `0 currently down`
|
|
throughout.
|
|
|
|
**"Some members are up" means any member NOT in the down bucket** — `running`, `unhealthy`,
|
|
`starting` or `restarting`. [DESIGN, R-384] The earlier guard was `running > 0`, counting only
|
|
`StateRunning`, which made the supervised test unreachable in precisely the case it was written for:
|
|
an unhealthy survivor beside a dead database counted as nothing being up.
|
|
|
|
---
|
|
|
|
## 2. Where each decision is made
|
|
|
|
| Decision | Where | Notes |
|
|
|---|---|---|
|
|
| container → stack aggregate state | `internal/stacks/manager.go` `aggregateState` | the ladder in §3 |
|
|
| is a down member supervised? | `internal/stacks/manager.go` `supervisedPolicy` | `no`/`on-failure` benign; everything else, **including unknown**, supervised |
|
|
| which states mean "down" | `internal/stacks/manager.go` `IsDownState` | `{stopped, exited, degraded}` |
|
|
| stack state → "this app is down" | `cmd/controller/main.go` `classifyRunStates` | **the single derivation point**; all three suppressions live here |
|
|
| sustained restarting → down | `internal/stacks/manager.go` `CrashLooping` | 5-minute threshold |
|
|
| quiesce suppression | `internal/quiesce/suppress.go` | cycle-keyed, 180 s grace |
|
|
| boot repair | `internal/bootrecon/bootrecon.go` | consumes `IsDownState` |
|
|
|
|
---
|
|
|
|
## 3. The aggregation ladder, in order
|
|
|
|
`aggregateState(containers, policyOf)` — **priority: degraded > unhealthy/starting > restarting >
|
|
all-running > stopped.**
|
|
|
|
1. no containers → `not_deployed`
|
|
2. **any DOWN member is supervised, and any member is up → `degraded`** ← R-384 put this first
|
|
3. any `unhealthy` → `unhealthy`
|
|
4. any `starting` → `starting`
|
|
5. any `restarting` → `restarting`
|
|
6. all running → `running`
|
|
7. all down → `stopped`
|
|
8. mix, every down member benign → `running`
|
|
|
|
Step 2's `policyOf` is consulted **only** for the down members, and only when something is up. `nil`
|
|
is allowed; every down member then reads as supervised.
|
|
|
|
---
|
|
|
|
## 4. Which states alarm, and which deliberately do not
|
|
|
|
`IsDownState` = `{stopped, exited, degraded}`.
|
|
|
|
| State | Down? | Why |
|
|
|---|---|---|
|
|
| `stopped`, `exited` | **yes** | not running, will not recover alone |
|
|
| `degraded` | **yes** | [DESIGN, R-51] a dead supervised member is as unreachable as a single app that exited — immich-server sat Exited 18 h with the app 100 % dead and no alert |
|
|
| `unhealthy` | **NO** | [DESIGN] a running container whose healthcheck is failing. Folding it in reintroduces the flapping fix-3 was written to stop. **R-384 did not change this** — it asks a prior question instead |
|
|
| `restarting` | **NO**, until sustained | [DESIGN, C9-F2] `restarting` is on the normal deploy path, so folding it in would alarm fleet-wide on every update. Becomes down after **5 min** (`crashLoopAfter`) |
|
|
| `starting`, `deploying` | no | mid-start |
|
|
| `paused` | no | a deliberate user action |
|
|
| `unknown` | no | [DESIGN] fail-OPEN — never manufacture a dead-app alert from an inconclusive read |
|
|
|
|
**Note the two fail directions are deliberately opposite.** `IsDownState` fails OPEN on `unknown`
|
|
(ambiguous *state* → do not alarm). `supervisedPolicy` fails CLOSED on unknown (we already KNOW a
|
|
member is dead; only the excuse is missing). Both are recorded at their sites.
|
|
|
|
---
|
|
|
|
## 5. The three suppressions, all at `classifyRunStates`
|
|
|
|
| Suppression | Rule | Expires? |
|
|
|---|---|---|
|
|
| **deliberate user stop** | `StateStopped` is not down **unless** the quiesce loop reports it failed to restart that stack | n/a — lifted by `failedRestart` |
|
|
| **quiesce cycle** | a stack this backup cycle stopped is exempt | **yes**, 180 s after unquiescing |
|
|
| **boot grace** | no evaluation for 90 s after controller start | **yes** |
|
|
|
|
**None of them latch.** [DESIGN, R-97b + R-88 Scenario D] A permanent suppression trades a loud false
|
|
alarm for a silent real one, which is the same error as an over-eager alarm, in the opposite
|
|
direction. Every window expires; the cost is a bounded DELAY in reporting a real failure, never its
|
|
loss.
|
|
|
|
The quiesce suppression is **cycle-keyed, not state-keyed** — an app caught mid-restart is
|
|
`starting`/`unhealthy`, not `stopped`, so no state test can see it. It is therefore **state-blind**,
|
|
which is why R-384 moving a stack from `unhealthy` to `degraded` cannot weaken it.
|
|
|
|
---
|
|
|
|
## 6. The alarm itself
|
|
|
|
Edge-triggered: `app_start_failed`, one event per transition into down, **not** per scan. Verified
|
|
live 2026-08-23 — one event across 22 scans.
|
|
|
|
- **Operator/hub event + dashboard banner.** `app_start_failed` is **not** in
|
|
`settings.DefaultEnabledEvents`, so by default it does **not** e-mail the customer.
|
|
- **F-OBS heartbeat**, every 20 scans (~10 min), at `[INFO]`:
|
|
`[deadapp] check alive: N scans since boot, M deployed app(s) evaluated, K currently down`.
|
|
This line exists because an absent alarm and a stopped detector look identical in a log.
|
|
|
|
> ⚠ **R-329, OPEN and it bites here.** `app_start_failed` is pushed with severity **`warn`**, which is
|
|
> **not** in the hub's vocabulary (`{info, warning, error, critical}`) and is silently coerced to
|
|
> `info` — which e-mails nobody, while the POST still returns 200. Observed again on 2026-08-23:
|
|
> `PushEvent: type=app_start_failed severity=warn`. R-384 makes this event actually fire, so the
|
|
> severity bug now matters more than it did while the event was unreachable.
|
|
|
|
---
|
|
|
|
## 7. Known gap, filed not fixed
|
|
|
|
> **R-386 (filed 2026-08-23, OPEN).** An all-down stack aggregates to `stopped` — `StateExited` is
|
|
> folded into the same counter and never survives aggregation. `classifyRunStates` then whitelists
|
|
> `stopped` as a deliberate user stop. So a **single-container app stopped out of band raises no
|
|
> alarm at all**, which directly contradicts the comment at `cmd/controller/main.go`: *"An out-of-band
|
|
> `docker compose stop` leaves the containers present → StateExited → still alerts."*
|
|
> **Measured on `demo-hp` 2026-08-23:** `privatebin` stopped out of band, 9 dead-app scans over 4+
|
|
> minutes, `state=stopped`, **zero events and zero banner lines** — against a positive control from
|
|
> the same box 17 minutes earlier. Not fixed in v0.222.0 deliberately; it is a separate decision.
|