Files
felhom.eu/documentation/audits/DRILL-r384-dead-db-alarm-2026-08-23/README.md
T
admin 55274d5ef3
gates / gates (push) Successful in 17s
R-385: make an UNRECORDED golden fail the currency gate; file R-386; own the alarm ladder
The gate failed only on `released > baked`, so it could catch a forgotten bake
and nothing else. A golden AHEAD of the record passed silently - and that is
how controller 0.221.1 was built, baked AND vouched while the newest CHANGELOG
heading still read v0.221.0, with every gate green. Reproduced on the real
history: newest released 0.221.0 / newest golden baked 0.221.1 -> exit 0.

The gate now asks whether the version being shipped is WRITTEN DOWN: the baked
version must have its own `## vX.Y.Z` heading anywhere in the CHANGELOG.
Membership rather than `baked > released` deliberately - a comparison against
the newest heading alone goes green the moment any later entry is written,
leaving the unrecorded version permanently unrecorded. INCONCLUSIVE (exit 2)
preserved; every refusal names a reason and a route.

Red-proofed both directions: old gate/old record exit 0, new gate/old record
exit 1, new gate/fixed record exit 0, absent clone exit 2, post-bake exit 0.

08-alarm-ladder.md is new, and its absence was itself the finding: no document
owned "when does a broken app raise an alarm?". The rules lived as comments in
four packages, each locally correct, with the ordering between them legible only
by reading one function top to bottom - which is how R-384 survived review.

R-383 and R-384 closed into CLOSED-ITEMS with their rules kept. R-385 filed
closed. R-386 filed OPEN: a single-container app stopped out of band raises no
alarm, and a comment claims the opposite - measured live, 9 scans, 0 events,
against a positive control from the same box 17 minutes earlier. Not fixed here.

Golden 0.222.0 baked and published; vouching is the operator's act.
2026-08-23 07:59:52 +02:00

97 lines
5.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# DRILL — R-384: an app whose database dies raised no alarm (2026-08-23)
**Controller v0.221.1 → v0.222.0. Live leg on `demo-hp` (Tier 0, disposable), guest 9201.**
**UNATTENDED.** Method: endpoint-level — no browser exists on DooPlex, so every read is either the
exact endpoint the UI calls or the controller's own log. Guest clock is UTC.
## Verdict
| Part | Outcome |
|---|---|
| Part 0 — the record | ✅ `v0.221.1` given its own heading, pushed ALONE (`da75603`) |
| Part 1 — the blind gate | ✅ fixed; both directions red-proofed against the real history |
| Part 2 — R-384 | ✅ shipped v0.222.0, **proven live** |
| Part 3 — R-383 | ✅ shipped v0.222.0 |
| §4 — the measurement | ⚠ **reproduced. Filed as R-386. NOT fixed — that was the instruction.** |
| Live walk step 4 (Scenario E) | **DROPPED** — consequence of the halt; drop-list item (3) |
## The one-line result
The same fixture that printed **`0 currently down`** on 2026-08-22 printed **`1 currently down`** on
2026-08-23, with 8 apps evaluated both times:
```
2026/08/22 21:13:48 [deadapp] check alive: 580 scans since boot, 8 deployed app(s) evaluated, 0 currently down
2026/08/23 05:37:44 [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down
```
## What was actually wrong
**The ORDER of two questions**, not the `unhealthy` exclusion. "Is a supervised member dead?" and "is
a running member failing its healthcheck?" are different questions, and the second was answering the
first — because a dying database drags its own front end `unhealthy`, **the symptom the fault causes
was what suppressed the alarm for it.** `IsDownState` was not touched; no state was minted.
The fix has **two halves and either alone leaves the defect standing**: the hoist, and widening "some
members are up" from `running > 0` to *any member not in the down bucket*. The old guard made the
R-51 block unreachable in precisely the case R-51 was written for.
## Evidence index (`evidence/`)
| File | What it shows |
|---|---|
| `gate-01-old-gate-old-changelog.txt` | the blindness: old gate, real history, **exit 0** |
| `gate-02-new-gate-old-changelog.txt` | new gate on the same history, **exit 1**, naming the fix |
| `gate-03-new-gate-new-changelog.txt` | with Part 0's heading, **exit 0** |
| `gate-04-inconclusive.txt` | INCONCLUSIVE (**exit 2**) preserved |
| `gate-05-new-gate-post-bake.txt` | v0.222.0 + golden 0.222.0, **exit 0** |
| `redproof-R384-1-order.txt` | mutation: hoist reverted → `"unhealthy", want "degraded"` |
| `redproof-R384-2-upguard.txt` | mutation: `up` narrowed → all three survivor shapes convict |
| `redproof-R384-3-classifier.txt` | mutation: classifier ignores `degraded` → empty banner |
| `redproof-R383-undo-phrase.txt` | mutation: unconditional claim → prints the false sentence |
| `live-00-the-old-heartbeat-2026-08-22.txt` | yesterday's `0 currently down`, carried in for contrast |
| `live-01`…`live-06` | Scenario A: pre-state, stop, log, decisive read, heartbeat, banner |
| `live-08-golden-bake-markers.txt` | the bake's acceptance markers, each counted |
| `live-09-…-full-controller-log.txt` | 1808 lines, pulled off **before** the app was restarted |
| `live-10`…`live-13` | Scenario B: unhealthy with nothing dead, no alarm, banner cleared |
| `live-14`…`live-16` | Scenario D: full stop→start cycle, **0 alarms across 9 scans** |
| `live-17`, `live-18` | §4: `privatebin` out-of-band stop, 9 scans, **0 events, 0 banner** |
| `live-19-…-full-log.txt` | 2139 lines covering Scenario D and §4 |
## §4 — what it found, in plain words
`aggregateState` folds `StateExited` into the `stopped` counter, so an all-down stack returns
`StateStopped` and **`StateExited` never survives aggregation** — that is the path the task suspected
and could not find in source. `classifyRunStates` then whitelists `StateStopped` as a deliberate user
stop. So this comment in `cmd/controller/main.go` is **false**:
> *"(An out-of-band `docker compose stop` leaves the containers present → StateExited → still alerts,
> which is correct: out-of-band tampering IS reportable.)"*
Measured: `privatebin` (1 container, `unless-stopped`) stopped 05:47:35Z; at 05:51:53Z it read
`state=stopped` with 9 dead-app scans behind it, **zero events and zero banner lines**.
**The absence is trustworthy because the detector was shown alive first** (standing rule 3):
`app_start_failed` fired for BookStack at 05:30:14Z on the same box 17 minutes earlier.
**Scoped honestly:** a genuine crash under `unless-stopped` is restarted by Docker and surfaces as
`restarting` → the 5-minute crash-loop path, which does alarm. The silent case is an explicit
out-of-band stop of a stack with no surviving member.
## Two things noticed that are NOT this drill's work
1. **R-329 moved from unreachable to load-bearing.** `app_start_failed` ships severity `warn`, which
is not in the hub's vocabulary and coerces silently to `info` — e-mailing nobody, POST still 200.
Observed again today. While the event never fired, this was harmless; it no longer is.
2. **The golden-bake runbook is missing `pveam update`.** On the `virgin` snapshot the template index
is stale, so `pveam available` offers `13.1-2` and downloading it fails with
`400 Parameter verification failed. template: no such template` — a confusing 400 rather than a
legible "your index is old". Recorded in `documentation/tests/golden-0.222.0-2026-08-23/README.md`.
## Teardown
Nothing was provisioned. All apps restored and confirmed healthy (`bookstack` + `bookstack-db`,
`docmost` ×3, `privatebin`); planted data untouched; no app rebuilt or restored. **Hub-side: nothing
to discard — the hub was READ ONLY this session** (`GET /configuration`, `GET /events`); no appliance
registered, no config written, no artifact manifest changed.