R-385: make an UNRECORDED golden fail the currency gate; file R-386; own the alarm ladder
gates / gates (push) Successful in 17s
gates / gates (push) Successful in 17s
The gate failed only on `released > baked`, so it could catch a forgotten bake and nothing else. A golden AHEAD of the record passed silently - and that is how controller 0.221.1 was built, baked AND vouched while the newest CHANGELOG heading still read v0.221.0, with every gate green. Reproduced on the real history: newest released 0.221.0 / newest golden baked 0.221.1 -> exit 0. The gate now asks whether the version being shipped is WRITTEN DOWN: the baked version must have its own `## vX.Y.Z` heading anywhere in the CHANGELOG. Membership rather than `baked > released` deliberately - a comparison against the newest heading alone goes green the moment any later entry is written, leaving the unrecorded version permanently unrecorded. INCONCLUSIVE (exit 2) preserved; every refusal names a reason and a route. Red-proofed both directions: old gate/old record exit 0, new gate/old record exit 1, new gate/fixed record exit 0, absent clone exit 2, post-bake exit 0. 08-alarm-ladder.md is new, and its absence was itself the finding: no document owned "when does a broken app raise an alarm?". The rules lived as comments in four packages, each locally correct, with the ordering between them legible only by reading one function top to bottom - which is how R-384 survived review. R-383 and R-384 closed into CLOSED-ITEMS with their rules kept. R-385 filed closed. R-386 filed OPEN: a single-container app stopped out of band raises no alarm, and a comment claims the opposite - measured live, 9 scans, 0 events, against a positive control from the same box 17 minutes earlier. Not fixed here. Golden 0.222.0 baked and published; vouching is the operator's act.
This commit is contained in:
@@ -67,6 +67,23 @@ the banner, pinned by `TestClassifyRunStates_PositiveControl_ADownStackDoesAlarm
|
||||
measurement DID expose is **R-384**: an app whose database has died is `unhealthy` too, and is
|
||||
likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decision.txt`.
|
||||
|
||||
> **R-384 CLOSED in controller v0.222.0 (2026-08-23), proven live.** The defect was the ORDER of two
|
||||
> questions, not the `unhealthy` exclusion: `aggregateState` now asks *"is a supervised member dead?"*
|
||||
> **before** the `unhealthy`/`starting`/`restarting` returns, and "some members are up" counts any
|
||||
> member not in the down bucket rather than `running` alone. `IsDownState` is byte-identical.
|
||||
> **Measured on `demo-hp` 2026-08-23** with the same fixture that read `0 currently down` the day
|
||||
> before: `bookstack-db` stopped 05:30:07Z → `app_start_failed` fired at **05:30:14Z**, the banner read
|
||||
> *„Telepített alkalmazás nem fut: BookStack (degraded)"*, the stack read `state=degraded` **while its
|
||||
> front end was `unhealthy`**, and the heartbeat printed **`1 currently down`** against the previous
|
||||
> day's `0`. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/`.
|
||||
>
|
||||
> **The HELD-app half of the paragraph above is now also covered** — a held app keeps its database
|
||||
> container, so it is the same shape and reaches the same `degraded` verdict.
|
||||
>
|
||||
> **The alarm ladder that decides all of this now has an owning document:** see
|
||||
> `08-alarm-ladder.md` (written 2026-08-23 — before that date no document owned it, and that absence
|
||||
> is why the ordering defect was legible only from source).
|
||||
|
||||
**WHAT IS STILL NOT CLAIMED:** the FAILURE path is where this class is weak, not the success path — a corrupt dump leaves the customer with an emptied or partially-applied database and an undo copy **no product action can apply** (**R-379**), and on MariaDB it does so behind an app that reports `health=healthy` (**R-380**). The success story is proven; the recovery-from-a-bad-restore story is not.
|
||||
|
||||
**NARROWED 2026-08-21 by the backup-truth drill — kept as history; both halves have since closed, see the entry immediately above.** Proven that night on `demo-hp` with planted, hash-recorded files: the **declared-userdata** leg of a **drive-declaring** app does come back byte-identical (`calibre-web`, 5/5 including two Hungarian accented filenames). **Two legs of the same story do NOT:** (a) the off-site restore has **no named-volume leg at all**, so an app's volume tar sits in the unit, in the snapshot and in the checking folder and is never replayed (**R-354**); (b) for the **40 of 53** apps that declare no data drive the off-site restore **refuses outright**, saying a running app „nincs telepítve" (**R-356**). Since the 40-class keeps ALL its data in named volumes, the end-to-end story is **unproven for that class and disproven for the volume leg generally**. The escrow/key half of this row is untouched by that and still stands. Evidence: `audits/DRILL-backup-truth-2026-08-21/evidence/` and `REPORT.md` (2026-08-21). |
|
||||
|
||||
@@ -0,0 +1,141 @@
|
||||
# 08 — The app-down alarm ladder
|
||||
|
||||
**Written 2026-08-23, with controller v0.222.0 (R-384).**
|
||||
|
||||
**The absence is the finding.** Until this file existed, no document owned the question *"when does a
|
||||
customer's app being broken raise an alarm?"* The rules were spread across four packages as comments,
|
||||
each locally correct, and the ordering between them was legible only by reading
|
||||
`aggregateState` top to bottom. That is exactly how R-384 survived: every individual rule was right,
|
||||
and the composition was wrong. Three separate defects in this ladder (R-51, C9-F2, R-384) were each
|
||||
found on live hardware rather than by review, and each is a case where a reader could not see the
|
||||
whole ladder at once.
|
||||
|
||||
Everything below is **[DESIGN]** — deliberate, with the reason recorded — unless marked otherwise.
|
||||
|
||||
---
|
||||
|
||||
## 1. The two questions, and their order
|
||||
|
||||
Two different questions get asked about a multi-container app, and **the order between them is
|
||||
load-bearing**:
|
||||
|
||||
1. **Is a SUPERVISED member of this app dead?** — a container Docker's restart policy says should be
|
||||
running, that is not.
|
||||
2. **Is a RUNNING member failing its healthcheck?**
|
||||
|
||||
**Question 1 is asked FIRST.** [DESIGN, R-384, v0.222.0]
|
||||
|
||||
Until v0.222.0 it was asked second, and the consequence was not subtle: a two-container app whose
|
||||
database exits goes `unhealthy` seconds later *because it cannot reach that database*. So the symptom
|
||||
the dead database causes was the thing that suppressed the alarm for it. Measured on `demo-hp`
|
||||
2026-08-22 — `bookstack-db` stopped at 21:27:01 and the watcher reported `0 currently down`
|
||||
throughout.
|
||||
|
||||
**"Some members are up" means any member NOT in the down bucket** — `running`, `unhealthy`,
|
||||
`starting` or `restarting`. [DESIGN, R-384] The earlier guard was `running > 0`, counting only
|
||||
`StateRunning`, which made the supervised test unreachable in precisely the case it was written for:
|
||||
an unhealthy survivor beside a dead database counted as nothing being up.
|
||||
|
||||
---
|
||||
|
||||
## 2. Where each decision is made
|
||||
|
||||
| Decision | Where | Notes |
|
||||
|---|---|---|
|
||||
| container → stack aggregate state | `internal/stacks/manager.go` `aggregateState` | the ladder in §3 |
|
||||
| is a down member supervised? | `internal/stacks/manager.go` `supervisedPolicy` | `no`/`on-failure` benign; everything else, **including unknown**, supervised |
|
||||
| which states mean "down" | `internal/stacks/manager.go` `IsDownState` | `{stopped, exited, degraded}` |
|
||||
| stack state → "this app is down" | `cmd/controller/main.go` `classifyRunStates` | **the single derivation point**; all three suppressions live here |
|
||||
| sustained restarting → down | `internal/stacks/manager.go` `CrashLooping` | 5-minute threshold |
|
||||
| quiesce suppression | `internal/quiesce/suppress.go` | cycle-keyed, 180 s grace |
|
||||
| boot repair | `internal/bootrecon/bootrecon.go` | consumes `IsDownState` |
|
||||
|
||||
---
|
||||
|
||||
## 3. The aggregation ladder, in order
|
||||
|
||||
`aggregateState(containers, policyOf)` — **priority: degraded > unhealthy/starting > restarting >
|
||||
all-running > stopped.**
|
||||
|
||||
1. no containers → `not_deployed`
|
||||
2. **any DOWN member is supervised, and any member is up → `degraded`** ← R-384 put this first
|
||||
3. any `unhealthy` → `unhealthy`
|
||||
4. any `starting` → `starting`
|
||||
5. any `restarting` → `restarting`
|
||||
6. all running → `running`
|
||||
7. all down → `stopped`
|
||||
8. mix, every down member benign → `running`
|
||||
|
||||
Step 2's `policyOf` is consulted **only** for the down members, and only when something is up. `nil`
|
||||
is allowed; every down member then reads as supervised.
|
||||
|
||||
---
|
||||
|
||||
## 4. Which states alarm, and which deliberately do not
|
||||
|
||||
`IsDownState` = `{stopped, exited, degraded}`.
|
||||
|
||||
| State | Down? | Why |
|
||||
|---|---|---|
|
||||
| `stopped`, `exited` | **yes** | not running, will not recover alone |
|
||||
| `degraded` | **yes** | [DESIGN, R-51] a dead supervised member is as unreachable as a single app that exited — immich-server sat Exited 18 h with the app 100 % dead and no alert |
|
||||
| `unhealthy` | **NO** | [DESIGN] a running container whose healthcheck is failing. Folding it in reintroduces the flapping fix-3 was written to stop. **R-384 did not change this** — it asks a prior question instead |
|
||||
| `restarting` | **NO**, until sustained | [DESIGN, C9-F2] `restarting` is on the normal deploy path, so folding it in would alarm fleet-wide on every update. Becomes down after **5 min** (`crashLoopAfter`) |
|
||||
| `starting`, `deploying` | no | mid-start |
|
||||
| `paused` | no | a deliberate user action |
|
||||
| `unknown` | no | [DESIGN] fail-OPEN — never manufacture a dead-app alert from an inconclusive read |
|
||||
|
||||
**Note the two fail directions are deliberately opposite.** `IsDownState` fails OPEN on `unknown`
|
||||
(ambiguous *state* → do not alarm). `supervisedPolicy` fails CLOSED on unknown (we already KNOW a
|
||||
member is dead; only the excuse is missing). Both are recorded at their sites.
|
||||
|
||||
---
|
||||
|
||||
## 5. The three suppressions, all at `classifyRunStates`
|
||||
|
||||
| Suppression | Rule | Expires? |
|
||||
|---|---|---|
|
||||
| **deliberate user stop** | `StateStopped` is not down **unless** the quiesce loop reports it failed to restart that stack | n/a — lifted by `failedRestart` |
|
||||
| **quiesce cycle** | a stack this backup cycle stopped is exempt | **yes**, 180 s after unquiescing |
|
||||
| **boot grace** | no evaluation for 90 s after controller start | **yes** |
|
||||
|
||||
**None of them latch.** [DESIGN, R-97b + R-88 Scenario D] A permanent suppression trades a loud false
|
||||
alarm for a silent real one, which is the same error as an over-eager alarm, in the opposite
|
||||
direction. Every window expires; the cost is a bounded DELAY in reporting a real failure, never its
|
||||
loss.
|
||||
|
||||
The quiesce suppression is **cycle-keyed, not state-keyed** — an app caught mid-restart is
|
||||
`starting`/`unhealthy`, not `stopped`, so no state test can see it. It is therefore **state-blind**,
|
||||
which is why R-384 moving a stack from `unhealthy` to `degraded` cannot weaken it.
|
||||
|
||||
---
|
||||
|
||||
## 6. The alarm itself
|
||||
|
||||
Edge-triggered: `app_start_failed`, one event per transition into down, **not** per scan. Verified
|
||||
live 2026-08-23 — one event across 22 scans.
|
||||
|
||||
- **Operator/hub event + dashboard banner.** `app_start_failed` is **not** in
|
||||
`settings.DefaultEnabledEvents`, so by default it does **not** e-mail the customer.
|
||||
- **F-OBS heartbeat**, every 20 scans (~10 min), at `[INFO]`:
|
||||
`[deadapp] check alive: N scans since boot, M deployed app(s) evaluated, K currently down`.
|
||||
This line exists because an absent alarm and a stopped detector look identical in a log.
|
||||
|
||||
> ⚠ **R-329, OPEN and it bites here.** `app_start_failed` is pushed with severity **`warn`**, which is
|
||||
> **not** in the hub's vocabulary (`{info, warning, error, critical}`) and is silently coerced to
|
||||
> `info` — which e-mails nobody, while the POST still returns 200. Observed again on 2026-08-23:
|
||||
> `PushEvent: type=app_start_failed severity=warn`. R-384 makes this event actually fire, so the
|
||||
> severity bug now matters more than it did while the event was unreachable.
|
||||
|
||||
---
|
||||
|
||||
## 7. Known gap, filed not fixed
|
||||
|
||||
> **R-386 (filed 2026-08-23, OPEN).** An all-down stack aggregates to `stopped` — `StateExited` is
|
||||
> folded into the same counter and never survives aggregation. `classifyRunStates` then whitelists
|
||||
> `stopped` as a deliberate user stop. So a **single-container app stopped out of band raises no
|
||||
> alarm at all**, which directly contradicts the comment at `cmd/controller/main.go`: *"An out-of-band
|
||||
> `docker compose stop` leaves the containers present → StateExited → still alerts."*
|
||||
> **Measured on `demo-hp` 2026-08-23:** `privatebin` stopped out of band, 9 dead-app scans over 4+
|
||||
> minutes, `state=stopped`, **zero events and zero banner lines** — against a positive control from
|
||||
> the same box 17 minutes earlier. Not fixed in v0.222.0 deliberately; it is a separate decision.
|
||||
Reference in New Issue
Block a user