Files
felhom.eu/documentation/audits/DRILL-r384-dead-db-alarm-2026-08-23
admin 55274d5ef3
gates / gates (push) Successful in 17s
R-385: make an UNRECORDED golden fail the currency gate; file R-386; own the alarm ladder
The gate failed only on `released > baked`, so it could catch a forgotten bake
and nothing else. A golden AHEAD of the record passed silently - and that is
how controller 0.221.1 was built, baked AND vouched while the newest CHANGELOG
heading still read v0.221.0, with every gate green. Reproduced on the real
history: newest released 0.221.0 / newest golden baked 0.221.1 -> exit 0.

The gate now asks whether the version being shipped is WRITTEN DOWN: the baked
version must have its own `## vX.Y.Z` heading anywhere in the CHANGELOG.
Membership rather than `baked > released` deliberately - a comparison against
the newest heading alone goes green the moment any later entry is written,
leaving the unrecorded version permanently unrecorded. INCONCLUSIVE (exit 2)
preserved; every refusal names a reason and a route.

Red-proofed both directions: old gate/old record exit 0, new gate/old record
exit 1, new gate/fixed record exit 0, absent clone exit 2, post-bake exit 0.

08-alarm-ladder.md is new, and its absence was itself the finding: no document
owned "when does a broken app raise an alarm?". The rules lived as comments in
four packages, each locally correct, with the ordering between them legible only
by reading one function top to bottom - which is how R-384 survived review.

R-383 and R-384 closed into CLOSED-ITEMS with their rules kept. R-385 filed
closed. R-386 filed OPEN: a single-container app stopped out of band raises no
alarm, and a comment claims the opposite - measured live, 9 scans, 0 events,
against a positive control from the same box 17 minutes earlier. Not fixed here.

Golden 0.222.0 baked and published; vouching is the operator's act.
2026-08-23 07:59:52 +02:00
..

DRILL — R-384: an app whose database dies raised no alarm (2026-08-23)

Controller v0.221.1 → v0.222.0. Live leg on demo-hp (Tier 0, disposable), guest 9201. UNATTENDED. Method: endpoint-level — no browser exists on DooPlex, so every read is either the exact endpoint the UI calls or the controller's own log. Guest clock is UTC.

Verdict

Part Outcome
Part 0 — the record ✅ v0.221.1 given its own heading, pushed ALONE (da75603)
Part 1 — the blind gate ✅ fixed; both directions red-proofed against the real history
Part 2 — R-384 ✅ shipped v0.222.0, proven live
Part 3 — R-383 ✅ shipped v0.222.0
§4 — the measurement ⚠ reproduced. Filed as R-386. NOT fixed — that was the instruction.
Live walk step 4 (Scenario E) DROPPED — consequence of the halt; drop-list item (3)

The one-line result

The same fixture that printed 0 currently down on 2026-08-22 printed 1 currently down on 2026-08-23, with 8 apps evaluated both times:

2026/08/22 21:13:48  [deadapp] check alive: 580 scans since boot, 8 deployed app(s) evaluated, 0 currently down
2026/08/23 05:37:44  [deadapp] check alive:  20 scans since boot, 8 deployed app(s) evaluated, 1 currently down

What was actually wrong

The ORDER of two questions, not the unhealthy exclusion. "Is a supervised member dead?" and "is a running member failing its healthcheck?" are different questions, and the second was answering the first — because a dying database drags its own front end unhealthy, the symptom the fault causes was what suppressed the alarm for it. IsDownState was not touched; no state was minted.

The fix has two halves and either alone leaves the defect standing: the hoist, and widening "some members are up" from running > 0 to any member not in the down bucket. The old guard made the R-51 block unreachable in precisely the case R-51 was written for.

Evidence index (evidence/)

File What it shows
gate-01-old-gate-old-changelog.txt the blindness: old gate, real history, exit 0
gate-02-new-gate-old-changelog.txt new gate on the same history, exit 1, naming the fix
gate-03-new-gate-new-changelog.txt with Part 0's heading, exit 0
gate-04-inconclusive.txt INCONCLUSIVE (exit 2) preserved
gate-05-new-gate-post-bake.txt v0.222.0 + golden 0.222.0, exit 0
redproof-R384-1-order.txt mutation: hoist reverted → "unhealthy", want "degraded"
redproof-R384-2-upguard.txt mutation: up narrowed → all three survivor shapes convict
redproof-R384-3-classifier.txt mutation: classifier ignores degraded → empty banner
redproof-R383-undo-phrase.txt mutation: unconditional claim → prints the false sentence
live-00-the-old-heartbeat-2026-08-22.txt yesterday's 0 currently down, carried in for contrast
live-01…live-06 Scenario A: pre-state, stop, log, decisive read, heartbeat, banner
live-08-golden-bake-markers.txt the bake's acceptance markers, each counted
live-09-…-full-controller-log.txt 1808 lines, pulled off before the app was restarted
live-10…live-13 Scenario B: unhealthy with nothing dead, no alarm, banner cleared
live-14…live-16 Scenario D: full stop→start cycle, 0 alarms across 9 scans
live-17, live-18 §4: privatebin out-of-band stop, 9 scans, 0 events, 0 banner
live-19-…-full-log.txt 2139 lines covering Scenario D and §4

§4 — what it found, in plain words

aggregateState folds StateExited into the stopped counter, so an all-down stack returns StateStopped and StateExited never survives aggregation — that is the path the task suspected and could not find in source. classifyRunStates then whitelists StateStopped as a deliberate user stop. So this comment in cmd/controller/main.go is false:

"(An out-of-band docker compose stop leaves the containers present → StateExited → still alerts, which is correct: out-of-band tampering IS reportable.)"

Measured: privatebin (1 container, unless-stopped) stopped 05:47:35Z; at 05:51:53Z it read state=stopped with 9 dead-app scans behind it, zero events and zero banner lines.

The absence is trustworthy because the detector was shown alive first (standing rule 3): app_start_failed fired for BookStack at 05:30:14Z on the same box 17 minutes earlier.

Scoped honestly: a genuine crash under unless-stopped is restarted by Docker and surfaces as restarting → the 5-minute crash-loop path, which does alarm. The silent case is an explicit out-of-band stop of a stack with no surviving member.

Two things noticed that are NOT this drill's work

  1. R-329 moved from unreachable to load-bearing. app_start_failed ships severity warn, which is not in the hub's vocabulary and coerces silently to info — e-mailing nobody, POST still 200. Observed again today. While the event never fired, this was harmless; it no longer is.
  2. The golden-bake runbook is missing pveam update. On the virgin snapshot the template index is stale, so pveam available offers 13.1-2 and downloading it fails with 400 Parameter verification failed. template: no such template — a confusing 400 rather than a legible "your index is old". Recorded in documentation/tests/golden-0.222.0-2026-08-23/README.md.

Teardown

Nothing was provisioned. All apps restored and confirmed healthy (bookstack + bookstack-db, docmost ×3, privatebin); planted data untouched; no app rebuilt or restored. Hub-side: nothing to discard — the hub was READ ONLY this session (GET /configuration, GET /events); no appliance registered, no config written, no artifact manifest changed.