55274d5ef3
gates / gates (push) Successful in 17s
The gate failed only on `released > baked`, so it could catch a forgotten bake and nothing else. A golden AHEAD of the record passed silently - and that is how controller 0.221.1 was built, baked AND vouched while the newest CHANGELOG heading still read v0.221.0, with every gate green. Reproduced on the real history: newest released 0.221.0 / newest golden baked 0.221.1 -> exit 0. The gate now asks whether the version being shipped is WRITTEN DOWN: the baked version must have its own `## vX.Y.Z` heading anywhere in the CHANGELOG. Membership rather than `baked > released` deliberately - a comparison against the newest heading alone goes green the moment any later entry is written, leaving the unrecorded version permanently unrecorded. INCONCLUSIVE (exit 2) preserved; every refusal names a reason and a route. Red-proofed both directions: old gate/old record exit 0, new gate/old record exit 1, new gate/fixed record exit 0, absent clone exit 2, post-bake exit 0. 08-alarm-ladder.md is new, and its absence was itself the finding: no document owned "when does a broken app raise an alarm?". The rules lived as comments in four packages, each locally correct, with the ordering between them legible only by reading one function top to bottom - which is how R-384 survived review. R-383 and R-384 closed into CLOSED-ITEMS with their rules kept. R-385 filed closed. R-386 filed OPEN: a single-container app stopped out of band raises no alarm, and a comment claims the opposite - measured live, 9 scans, 0 events, against a positive control from the same box 17 minutes earlier. Not fixed here. Golden 0.222.0 baked and published; vouching is the operator's act.
97 lines
5.8 KiB
Markdown
97 lines
5.8 KiB
Markdown
# DRILL — R-384: an app whose database dies raised no alarm (2026-08-23)
|
||
|
||
**Controller v0.221.1 → v0.222.0. Live leg on `demo-hp` (Tier 0, disposable), guest 9201.**
|
||
**UNATTENDED.** Method: endpoint-level — no browser exists on DooPlex, so every read is either the
|
||
exact endpoint the UI calls or the controller's own log. Guest clock is UTC.
|
||
|
||
## Verdict
|
||
|
||
| Part | Outcome |
|
||
|---|---|
|
||
| Part 0 — the record | ✅ `v0.221.1` given its own heading, pushed ALONE (`da75603`) |
|
||
| Part 1 — the blind gate | ✅ fixed; both directions red-proofed against the real history |
|
||
| Part 2 — R-384 | ✅ shipped v0.222.0, **proven live** |
|
||
| Part 3 — R-383 | ✅ shipped v0.222.0 |
|
||
| §4 — the measurement | ⚠ **reproduced. Filed as R-386. NOT fixed — that was the instruction.** |
|
||
| Live walk step 4 (Scenario E) | **DROPPED** — consequence of the halt; drop-list item (3) |
|
||
|
||
## The one-line result
|
||
|
||
The same fixture that printed **`0 currently down`** on 2026-08-22 printed **`1 currently down`** on
|
||
2026-08-23, with 8 apps evaluated both times:
|
||
|
||
```
|
||
2026/08/22 21:13:48 [deadapp] check alive: 580 scans since boot, 8 deployed app(s) evaluated, 0 currently down
|
||
2026/08/23 05:37:44 [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down
|
||
```
|
||
|
||
## What was actually wrong
|
||
|
||
**The ORDER of two questions**, not the `unhealthy` exclusion. "Is a supervised member dead?" and "is
|
||
a running member failing its healthcheck?" are different questions, and the second was answering the
|
||
first — because a dying database drags its own front end `unhealthy`, **the symptom the fault causes
|
||
was what suppressed the alarm for it.** `IsDownState` was not touched; no state was minted.
|
||
|
||
The fix has **two halves and either alone leaves the defect standing**: the hoist, and widening "some
|
||
members are up" from `running > 0` to *any member not in the down bucket*. The old guard made the
|
||
R-51 block unreachable in precisely the case R-51 was written for.
|
||
|
||
## Evidence index (`evidence/`)
|
||
|
||
| File | What it shows |
|
||
|---|---|
|
||
| `gate-01-old-gate-old-changelog.txt` | the blindness: old gate, real history, **exit 0** |
|
||
| `gate-02-new-gate-old-changelog.txt` | new gate on the same history, **exit 1**, naming the fix |
|
||
| `gate-03-new-gate-new-changelog.txt` | with Part 0's heading, **exit 0** |
|
||
| `gate-04-inconclusive.txt` | INCONCLUSIVE (**exit 2**) preserved |
|
||
| `gate-05-new-gate-post-bake.txt` | v0.222.0 + golden 0.222.0, **exit 0** |
|
||
| `redproof-R384-1-order.txt` | mutation: hoist reverted → `"unhealthy", want "degraded"` |
|
||
| `redproof-R384-2-upguard.txt` | mutation: `up` narrowed → all three survivor shapes convict |
|
||
| `redproof-R384-3-classifier.txt` | mutation: classifier ignores `degraded` → empty banner |
|
||
| `redproof-R383-undo-phrase.txt` | mutation: unconditional claim → prints the false sentence |
|
||
| `live-00-the-old-heartbeat-2026-08-22.txt` | yesterday's `0 currently down`, carried in for contrast |
|
||
| `live-01`…`live-06` | Scenario A: pre-state, stop, log, decisive read, heartbeat, banner |
|
||
| `live-08-golden-bake-markers.txt` | the bake's acceptance markers, each counted |
|
||
| `live-09-…-full-controller-log.txt` | 1808 lines, pulled off **before** the app was restarted |
|
||
| `live-10`…`live-13` | Scenario B: unhealthy with nothing dead, no alarm, banner cleared |
|
||
| `live-14`…`live-16` | Scenario D: full stop→start cycle, **0 alarms across 9 scans** |
|
||
| `live-17`, `live-18` | §4: `privatebin` out-of-band stop, 9 scans, **0 events, 0 banner** |
|
||
| `live-19-…-full-log.txt` | 2139 lines covering Scenario D and §4 |
|
||
|
||
## §4 — what it found, in plain words
|
||
|
||
`aggregateState` folds `StateExited` into the `stopped` counter, so an all-down stack returns
|
||
`StateStopped` and **`StateExited` never survives aggregation** — that is the path the task suspected
|
||
and could not find in source. `classifyRunStates` then whitelists `StateStopped` as a deliberate user
|
||
stop. So this comment in `cmd/controller/main.go` is **false**:
|
||
|
||
> *"(An out-of-band `docker compose stop` leaves the containers present → StateExited → still alerts,
|
||
> which is correct: out-of-band tampering IS reportable.)"*
|
||
|
||
Measured: `privatebin` (1 container, `unless-stopped`) stopped 05:47:35Z; at 05:51:53Z it read
|
||
`state=stopped` with 9 dead-app scans behind it, **zero events and zero banner lines**.
|
||
|
||
**The absence is trustworthy because the detector was shown alive first** (standing rule 3):
|
||
`app_start_failed` fired for BookStack at 05:30:14Z on the same box 17 minutes earlier.
|
||
|
||
**Scoped honestly:** a genuine crash under `unless-stopped` is restarted by Docker and surfaces as
|
||
`restarting` → the 5-minute crash-loop path, which does alarm. The silent case is an explicit
|
||
out-of-band stop of a stack with no surviving member.
|
||
|
||
## Two things noticed that are NOT this drill's work
|
||
|
||
1. **R-329 moved from unreachable to load-bearing.** `app_start_failed` ships severity `warn`, which
|
||
is not in the hub's vocabulary and coerces silently to `info` — e-mailing nobody, POST still 200.
|
||
Observed again today. While the event never fired, this was harmless; it no longer is.
|
||
2. **The golden-bake runbook is missing `pveam update`.** On the `virgin` snapshot the template index
|
||
is stale, so `pveam available` offers `13.1-2` and downloading it fails with
|
||
`400 Parameter verification failed. template: no such template` — a confusing 400 rather than a
|
||
legible "your index is old". Recorded in `documentation/tests/golden-0.222.0-2026-08-23/README.md`.
|
||
|
||
## Teardown
|
||
|
||
Nothing was provisioned. All apps restored and confirmed healthy (`bookstack` + `bookstack-db`,
|
||
`docmost` ×3, `privatebin`); planted data untouched; no app rebuilt or restored. **Hub-side: nothing
|
||
to discard — the hub was READ ONLY this session** (`GET /configuration`, `GET /events`); no appliance
|
||
registered, no config written, no artifact manifest changed.
|