The gate failed only on `released > baked`, so it could catch a forgotten bake and nothing else. A golden AHEAD of the record passed silently - and that is how controller 0.221.1 was built, baked AND vouched while the newest CHANGELOG heading still read v0.221.0, with every gate green. Reproduced on the real history: newest released 0.221.0 / newest golden baked 0.221.1 -> exit 0. The gate now asks whether the version being shipped is WRITTEN DOWN: the baked version must have its own `## vX.Y.Z` heading anywhere in the CHANGELOG. Membership rather than `baked > released` deliberately - a comparison against the newest heading alone goes green the moment any later entry is written, leaving the unrecorded version permanently unrecorded. INCONCLUSIVE (exit 2) preserved; every refusal names a reason and a route. Red-proofed both directions: old gate/old record exit 0, new gate/old record exit 1, new gate/fixed record exit 0, absent clone exit 2, post-bake exit 0. 08-alarm-ladder.md is new, and its absence was itself the finding: no document owned "when does a broken app raise an alarm?". The rules lived as comments in four packages, each locally correct, with the ordering between them legible only by reading one function top to bottom - which is how R-384 survived review. R-383 and R-384 closed into CLOSED-ITEMS with their rules kept. R-385 filed closed. R-386 filed OPEN: a single-container app stopped out of band raises no alarm, and a comment claims the opposite - measured live, 9 scans, 0 events, against a positive control from the same box 17 minutes earlier. Not fixed here. Golden 0.222.0 baked and published; vouching is the operator's act.
DRILL — R-384: an app whose database dies raised no alarm (2026-08-23)
Controller v0.221.1 → v0.222.0. Live leg on demo-hp (Tier 0, disposable), guest 9201.
UNATTENDED. Method: endpoint-level — no browser exists on DooPlex, so every read is either the
exact endpoint the UI calls or the controller's own log. Guest clock is UTC.
Verdict
| Part | Outcome |
|---|---|
| Part 0 — the record | ✅ v0.221.1 given its own heading, pushed ALONE (da75603) |
| Part 1 — the blind gate | ✅ fixed; both directions red-proofed against the real history |
| Part 2 — R-384 | ✅ shipped v0.222.0, proven live |
| Part 3 — R-383 | ✅ shipped v0.222.0 |
| §4 — the measurement | ⚠ reproduced. Filed as R-386. NOT fixed — that was the instruction. |
| Live walk step 4 (Scenario E) | DROPPED — consequence of the halt; drop-list item (3) |
The one-line result
The same fixture that printed 0 currently down on 2026-08-22 printed 1 currently down on
2026-08-23, with 8 apps evaluated both times:
2026/08/22 21:13:48 [deadapp] check alive: 580 scans since boot, 8 deployed app(s) evaluated, 0 currently down
2026/08/23 05:37:44 [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down
What was actually wrong
The ORDER of two questions, not the unhealthy exclusion. "Is a supervised member dead?" and "is
a running member failing its healthcheck?" are different questions, and the second was answering the
first — because a dying database drags its own front end unhealthy, the symptom the fault causes
was what suppressed the alarm for it. IsDownState was not touched; no state was minted.
The fix has two halves and either alone leaves the defect standing: the hoist, and widening "some
members are up" from running > 0 to any member not in the down bucket. The old guard made the
R-51 block unreachable in precisely the case R-51 was written for.
Evidence index (evidence/)
| File | What it shows |
|---|---|
gate-01-old-gate-old-changelog.txt |
the blindness: old gate, real history, exit 0 |
gate-02-new-gate-old-changelog.txt |
new gate on the same history, exit 1, naming the fix |
gate-03-new-gate-new-changelog.txt |
with Part 0's heading, exit 0 |
gate-04-inconclusive.txt |
INCONCLUSIVE (exit 2) preserved |
gate-05-new-gate-post-bake.txt |
v0.222.0 + golden 0.222.0, exit 0 |
redproof-R384-1-order.txt |
mutation: hoist reverted → "unhealthy", want "degraded" |
redproof-R384-2-upguard.txt |
mutation: up narrowed → all three survivor shapes convict |
redproof-R384-3-classifier.txt |
mutation: classifier ignores degraded → empty banner |
redproof-R383-undo-phrase.txt |
mutation: unconditional claim → prints the false sentence |
live-00-the-old-heartbeat-2026-08-22.txt |
yesterday's 0 currently down, carried in for contrast |
live-01…live-06 |
Scenario A: pre-state, stop, log, decisive read, heartbeat, banner |
live-08-golden-bake-markers.txt |
the bake's acceptance markers, each counted |
live-09-…-full-controller-log.txt |
1808 lines, pulled off before the app was restarted |
live-10…live-13 |
Scenario B: unhealthy with nothing dead, no alarm, banner cleared |
live-14…live-16 |
Scenario D: full stop→start cycle, 0 alarms across 9 scans |
live-17, live-18 |
§4: privatebin out-of-band stop, 9 scans, 0 events, 0 banner |
live-19-…-full-log.txt |
2139 lines covering Scenario D and §4 |
§4 — what it found, in plain words
aggregateState folds StateExited into the stopped counter, so an all-down stack returns
StateStopped and StateExited never survives aggregation — that is the path the task suspected
and could not find in source. classifyRunStates then whitelists StateStopped as a deliberate user
stop. So this comment in cmd/controller/main.go is false:
"(An out-of-band
docker compose stopleaves the containers present → StateExited → still alerts, which is correct: out-of-band tampering IS reportable.)"
Measured: privatebin (1 container, unless-stopped) stopped 05:47:35Z; at 05:51:53Z it read
state=stopped with 9 dead-app scans behind it, zero events and zero banner lines.
The absence is trustworthy because the detector was shown alive first (standing rule 3):
app_start_failed fired for BookStack at 05:30:14Z on the same box 17 minutes earlier.
Scoped honestly: a genuine crash under unless-stopped is restarted by Docker and surfaces as
restarting → the 5-minute crash-loop path, which does alarm. The silent case is an explicit
out-of-band stop of a stack with no surviving member.
Two things noticed that are NOT this drill's work
- R-329 moved from unreachable to load-bearing.
app_start_failedships severitywarn, which is not in the hub's vocabulary and coerces silently toinfo— e-mailing nobody, POST still 200. Observed again today. While the event never fired, this was harmless; it no longer is. - The golden-bake runbook is missing
pveam update. On thevirginsnapshot the template index is stale, sopveam availableoffers13.1-2and downloading it fails with400 Parameter verification failed. template: no such template— a confusing 400 rather than a legible "your index is old". Recorded indocumentation/tests/golden-0.222.0-2026-08-23/README.md.
Teardown
Nothing was provisioned. All apps restored and confirmed healthy (bookstack + bookstack-db,
docmost ×3, privatebin); planted data untouched; no app rebuilt or restored. Hub-side: nothing
to discard — the hub was READ ONLY this session (GET /configuration, GET /events); no appliance
registered, no config written, no artifact manifest changed.