R-385: make an UNRECORDED golden fail the currency gate; file R-386; own the alarm ladder
gates / gates (push) Successful in 17s

The gate failed only on `released > baked`, so it could catch a forgotten bake
and nothing else. A golden AHEAD of the record passed silently - and that is
how controller 0.221.1 was built, baked AND vouched while the newest CHANGELOG
heading still read v0.221.0, with every gate green. Reproduced on the real
history: newest released 0.221.0 / newest golden baked 0.221.1 -> exit 0.

The gate now asks whether the version being shipped is WRITTEN DOWN: the baked
version must have its own `## vX.Y.Z` heading anywhere in the CHANGELOG.
Membership rather than `baked > released` deliberately - a comparison against
the newest heading alone goes green the moment any later entry is written,
leaving the unrecorded version permanently unrecorded. INCONCLUSIVE (exit 2)
preserved; every refusal names a reason and a route.

Red-proofed both directions: old gate/old record exit 0, new gate/old record
exit 1, new gate/fixed record exit 0, absent clone exit 2, post-bake exit 0.

08-alarm-ladder.md is new, and its absence was itself the finding: no document
owned "when does a broken app raise an alarm?". The rules lived as comments in
four packages, each locally correct, with the ordering between them legible only
by reading one function top to bottom - which is how R-384 survived review.

R-383 and R-384 closed into CLOSED-ITEMS with their rules kept. R-385 filed
closed. R-386 filed OPEN: a single-container app stopped out of band raises no
alarm, and a comment claims the opposite - measured live, 9 scans, 0 events,
against a positive control from the same box 17 minutes earlier. Not fixed here.

Golden 0.222.0 baked and published; vouching is the operator's act.
This commit is contained in:
2026-08-23 07:59:52 +02:00
parent 1eb64bec51
commit 55274d5ef3
39 changed files with 4986 additions and 65 deletions
@@ -67,6 +67,23 @@ the banner, pinned by `TestClassifyRunStates_PositiveControl_ADownStackDoesAlarm
measurement DID expose is **R-384**: an app whose database has died is `unhealthy` too, and is
likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decision.txt`.
> **R-384 CLOSED in controller v0.222.0 (2026-08-23), proven live.** The defect was the ORDER of two
> questions, not the `unhealthy` exclusion: `aggregateState` now asks *"is a supervised member dead?"*
> **before** the `unhealthy`/`starting`/`restarting` returns, and "some members are up" counts any
> member not in the down bucket rather than `running` alone. `IsDownState` is byte-identical.
> **Measured on `demo-hp` 2026-08-23** with the same fixture that read `0 currently down` the day
> before: `bookstack-db` stopped 05:30:07Z → `app_start_failed` fired at **05:30:14Z**, the banner read
> *„Telepített alkalmazás nem fut: BookStack (degraded)"*, the stack read `state=degraded` **while its
> front end was `unhealthy`**, and the heartbeat printed **`1 currently down`** against the previous
> day's `0`. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/`.
>
> **The HELD-app half of the paragraph above is now also covered** — a held app keeps its database
> container, so it is the same shape and reaches the same `degraded` verdict.
>
> **The alarm ladder that decides all of this now has an owning document:** see
> `08-alarm-ladder.md` (written 2026-08-23 — before that date no document owned it, and that absence
> is why the ordering defect was legible only from source).
**WHAT IS STILL NOT CLAIMED:** the FAILURE path is where this class is weak, not the success path — a corrupt dump leaves the customer with an emptied or partially-applied database and an undo copy **no product action can apply** (**R-379**), and on MariaDB it does so behind an app that reports `health=healthy` (**R-380**). The success story is proven; the recovery-from-a-bad-restore story is not.
**NARROWED 2026-08-21 by the backup-truth drill — kept as history; both halves have since closed, see the entry immediately above.** Proven that night on `demo-hp` with planted, hash-recorded files: the **declared-userdata** leg of a **drive-declaring** app does come back byte-identical (`calibre-web`, 5/5 including two Hungarian accented filenames). **Two legs of the same story do NOT:** (a) the off-site restore has **no named-volume leg at all**, so an app's volume tar sits in the unit, in the snapshot and in the checking folder and is never replayed (**R-354**); (b) for the **40 of 53** apps that declare no data drive the off-site restore **refuses outright**, saying a running app „nincs telepítve" (**R-356**). Since the 40-class keeps ALL its data in named volumes, the end-to-end story is **unproven for that class and disproven for the volume leg generally**. The escrow/key half of this row is untouched by that and still stands. Evidence: `audits/DRILL-backup-truth-2026-08-21/evidence/` and `REPORT.md` (2026-08-21). |
@@ -0,0 +1,141 @@
# 08 — The app-down alarm ladder
**Written 2026-08-23, with controller v0.222.0 (R-384).**
**The absence is the finding.** Until this file existed, no document owned the question *"when does a
customer's app being broken raise an alarm?"* The rules were spread across four packages as comments,
each locally correct, and the ordering between them was legible only by reading
`aggregateState` top to bottom. That is exactly how R-384 survived: every individual rule was right,
and the composition was wrong. Three separate defects in this ladder (R-51, C9-F2, R-384) were each
found on live hardware rather than by review, and each is a case where a reader could not see the
whole ladder at once.
Everything below is **[DESIGN]** — deliberate, with the reason recorded — unless marked otherwise.
---
## 1. The two questions, and their order
Two different questions get asked about a multi-container app, and **the order between them is
load-bearing**:
1. **Is a SUPERVISED member of this app dead?** — a container Docker's restart policy says should be
running, that is not.
2. **Is a RUNNING member failing its healthcheck?**
**Question 1 is asked FIRST.** [DESIGN, R-384, v0.222.0]
Until v0.222.0 it was asked second, and the consequence was not subtle: a two-container app whose
database exits goes `unhealthy` seconds later *because it cannot reach that database*. So the symptom
the dead database causes was the thing that suppressed the alarm for it. Measured on `demo-hp`
2026-08-22 — `bookstack-db` stopped at 21:27:01 and the watcher reported `0 currently down`
throughout.
**"Some members are up" means any member NOT in the down bucket** — `running`, `unhealthy`,
`starting` or `restarting`. [DESIGN, R-384] The earlier guard was `running > 0`, counting only
`StateRunning`, which made the supervised test unreachable in precisely the case it was written for:
an unhealthy survivor beside a dead database counted as nothing being up.
---
## 2. Where each decision is made
| Decision | Where | Notes |
|---|---|---|
| container → stack aggregate state | `internal/stacks/manager.go` `aggregateState` | the ladder in §3 |
| is a down member supervised? | `internal/stacks/manager.go` `supervisedPolicy` | `no`/`on-failure` benign; everything else, **including unknown**, supervised |
| which states mean "down" | `internal/stacks/manager.go` `IsDownState` | `{stopped, exited, degraded}` |
| stack state → "this app is down" | `cmd/controller/main.go` `classifyRunStates` | **the single derivation point**; all three suppressions live here |
| sustained restarting → down | `internal/stacks/manager.go` `CrashLooping` | 5-minute threshold |
| quiesce suppression | `internal/quiesce/suppress.go` | cycle-keyed, 180 s grace |
| boot repair | `internal/bootrecon/bootrecon.go` | consumes `IsDownState` |
---
## 3. The aggregation ladder, in order
`aggregateState(containers, policyOf)` — **priority: degraded > unhealthy/starting > restarting >
all-running > stopped.**
1. no containers → `not_deployed`
2. **any DOWN member is supervised, and any member is up → `degraded`** ← R-384 put this first
3. any `unhealthy` → `unhealthy`
4. any `starting` → `starting`
5. any `restarting` → `restarting`
6. all running → `running`
7. all down → `stopped`
8. mix, every down member benign → `running`
Step 2's `policyOf` is consulted **only** for the down members, and only when something is up. `nil`
is allowed; every down member then reads as supervised.
---
## 4. Which states alarm, and which deliberately do not
`IsDownState` = `{stopped, exited, degraded}`.
| State | Down? | Why |
|---|---|---|
| `stopped`, `exited` | **yes** | not running, will not recover alone |
| `degraded` | **yes** | [DESIGN, R-51] a dead supervised member is as unreachable as a single app that exited — immich-server sat Exited 18 h with the app 100 % dead and no alert |
| `unhealthy` | **NO** | [DESIGN] a running container whose healthcheck is failing. Folding it in reintroduces the flapping fix-3 was written to stop. **R-384 did not change this** — it asks a prior question instead |
| `restarting` | **NO**, until sustained | [DESIGN, C9-F2] `restarting` is on the normal deploy path, so folding it in would alarm fleet-wide on every update. Becomes down after **5 min** (`crashLoopAfter`) |
| `starting`, `deploying` | no | mid-start |
| `paused` | no | a deliberate user action |
| `unknown` | no | [DESIGN] fail-OPEN — never manufacture a dead-app alert from an inconclusive read |
**Note the two fail directions are deliberately opposite.** `IsDownState` fails OPEN on `unknown`
(ambiguous *state* → do not alarm). `supervisedPolicy` fails CLOSED on unknown (we already KNOW a
member is dead; only the excuse is missing). Both are recorded at their sites.
---
## 5. The three suppressions, all at `classifyRunStates`
| Suppression | Rule | Expires? |
|---|---|---|
| **deliberate user stop** | `StateStopped` is not down **unless** the quiesce loop reports it failed to restart that stack | n/a — lifted by `failedRestart` |
| **quiesce cycle** | a stack this backup cycle stopped is exempt | **yes**, 180 s after unquiescing |
| **boot grace** | no evaluation for 90 s after controller start | **yes** |
**None of them latch.** [DESIGN, R-97b + R-88 Scenario D] A permanent suppression trades a loud false
alarm for a silent real one, which is the same error as an over-eager alarm, in the opposite
direction. Every window expires; the cost is a bounded DELAY in reporting a real failure, never its
loss.
The quiesce suppression is **cycle-keyed, not state-keyed** — an app caught mid-restart is
`starting`/`unhealthy`, not `stopped`, so no state test can see it. It is therefore **state-blind**,
which is why R-384 moving a stack from `unhealthy` to `degraded` cannot weaken it.
---
## 6. The alarm itself
Edge-triggered: `app_start_failed`, one event per transition into down, **not** per scan. Verified
live 2026-08-23 — one event across 22 scans.
- **Operator/hub event + dashboard banner.** `app_start_failed` is **not** in
`settings.DefaultEnabledEvents`, so by default it does **not** e-mail the customer.
- **F-OBS heartbeat**, every 20 scans (~10 min), at `[INFO]`:
`[deadapp] check alive: N scans since boot, M deployed app(s) evaluated, K currently down`.
This line exists because an absent alarm and a stopped detector look identical in a log.
> ⚠ **R-329, OPEN and it bites here.** `app_start_failed` is pushed with severity **`warn`**, which is
> **not** in the hub's vocabulary (`{info, warning, error, critical}`) and is silently coerced to
> `info` — which e-mails nobody, while the POST still returns 200. Observed again on 2026-08-23:
> `PushEvent: type=app_start_failed severity=warn`. R-384 makes this event actually fire, so the
> severity bug now matters more than it did while the event was unreachable.
---
## 7. Known gap, filed not fixed
> **R-386 (filed 2026-08-23, OPEN).** An all-down stack aggregates to `stopped` — `StateExited` is
> folded into the same counter and never survives aggregation. `classifyRunStates` then whitelists
> `stopped` as a deliberate user stop. So a **single-container app stopped out of band raises no
> alarm at all**, which directly contradicts the comment at `cmd/controller/main.go`: *"An out-of-band
> `docker compose stop` leaves the containers present → StateExited → still alerts."*
> **Measured on `demo-hp` 2026-08-23:** `privatebin` stopped out of band, 9 dead-app scans over 4+
> minutes, `state=stopped`, **zero events and zero banner lines** — against a positive control from
> the same box 17 minutes earlier. Not fixed in v0.222.0 deliberately; it is a separate decision.