Files
felhom.eu/documentation/audits/DRILL-cooldown-grain-2026-08-23

DRILL — R-389: only the first broken app per hour reached the operator (2026-08-23)

Hub v0.107.0 → v0.108.0. No controller release, so no golden and no floor. Live leg on demo-hp (Tier 0) plus the hub. UNATTENDED. Method: the hub's own SQLite records plus the controller log. Guest and hub DB are UTC; the hub pod logs CEST.

No halt condition fired.

The pair that says it

Same box, same shape — two apps down four minutes apart, inside one hour:

2026-08-23 09:27:51  sent        BookStack     <- v0.107.0
2026-08-23 09:31:51  suppressed  PrivateBin        operator cooldown 1h, key=demo-hp:app_start_failed

2026-08-23 11:56:57  sent        OpenGist      <- v0.108.0
2026-08-23 12:00:57  sent        Calibre-Web

2 sent / 0 suppressed, where the day before the identical shape gave one of each.

And the keys, from the hub's own suppression rows — it records the key only when it declines:

operator cooldown 1h, key=demo-hp:app_start_failed:opengist
operator cooldown 1h, key=demo-hp:app_start_failed:calibre-web

against v0.107.0's shared key=demo-hp:app_start_failed.

The fence, and why it is not decoration

crossdrive_failed is severity error, reaches the operator leg, and carries stack_name — through CrossDriveDetails, a different struct from AppDetails. A rule of the form "if the details carry a stack_name, split per app" would have split it and silently undone R-182.

Proven live: two different apps' crossdrive_failed, one minute apart →

sent        opengist
suppressed  calibre-web    operator cooldown 1h, key=demo-hp:crossdrive_failed

Byte-identical to the derived v0.107.0 key. No app suffix. That is why cooldownStackSuffix takes the event type as well as the details, unlike its two siblings.

Part 2 — the burst, measured

Three apps stopped in one scan (kimai, romm, paperless-ngx, none with a live cooldown — checked first, because a stale one would have halved the count and made the answer look better than it is):

attempted 3
sent 3
suppressed 0

Judgement: per-app is the right grain, and this volume is acceptable. The reference box has 8 deployed apps, so a total outage is 8 mails; the boot grace (90 s), the quiesce grace (180 s) and the per-app edge trigger absorb reboots, backup cycles and persistently-dead apps. No digest row was filed. The reopening condition is stated rather than left implicit: it scales linearly with app count and has no ceiling, so a box large enough that a total outage is unreadable is the point at which the answer becomes a digest with a customer message — not a wider cooldown.

Part 3 — gate 11, and the spec discrepancy it forced

The gate refuses a push whose REPORT.md carries an observation with neither FILED: R-NNN nor NOT-A-FINDING: <reason>.

The specification said an item may "cite an R-NNN that resolves". That rule would have passed the very item the gate was built to catch. Yesterday's lost observation reads "This is R-182's known cooldown-key shape…" — R-182 resolves, and it is cited as an analogy, not as the row that files it. No parser can tell citation-as-precedent from citation-as-filing by reading prose. The marker is therefore explicit, and the discrepancy is recorded in the gate's docstring rather than quietly resolved. EDGE 7 in the evidence is that exact case, convicted.

Control Expected Observed
historical: yesterday's real section, verbatim refuse exit 1, naming both items
plant an observation with no row refuse exit 1
add the row pass exit 0
remove the observation pass quietly exit 0
FILED: a row that does not resolve refuse exit 1
NOT-A-FINDING: with no reason refuse exit 1
both markers on one item refuse exit 1
section present, no numbered items inconclusive exit 2, saying what it could not read
no REPORT.md at all inconclusive exit 2
a bare R-182 mention (the trap) refuse exit 1

Evidence index (evidence/)

File What it shows
redproof-1-key.txt suffix dropped from the key → 1 operator mail(s), want 2, and the live key shape reproduced
redproof-2-allowlist.txt allow-list removed → crossdrive_failed, app_deployed, app_removed, backup_failed all split per app
gate11-01-historical-redproof.txt yesterday's actual observations section, refused
gate11-02-three-controls.txt plant → refuse, file → pass, remove → pass
gate11-03-edges.txt seven boundaries incl. the R-182 trap
live-01…live-05 Scenarios A and B: both sent, then each suppressed under its own key
live-06, live-07 Scenario C: crossdrive stays coarse
live-08, live-09 the burst, with its pre-check
live-10-full-controller-log.txt 1803 lines, pulled before the restore

Teardown

Nothing provisioned. Five apps were stopped across the walk (opengist, calibre-web, kimai, romm, paperless-ngx) and all were restarted and confirmed healthy — 17 containers up. No app was rebuilt, redeployed or restored; the three retained subjects (docmost, bookstack, privatebin) were not touched at all.

Hub-side, stated explicitly. The hub was written this session: deployment to v0.108.0 via the manifest, and six probe events were POSTed to the live hub for Scenarios B and C (two app_start_failed, two crossdrive_failed, plus the two from yesterday's Scenario H that were deliberately left in place). They are inert event rows for demo-hp and are named here rather than left to be found. Every count in this drill is filtered by created_at >= T0 precisely so those rows cannot contaminate it. Nothing else: no appliance registered, no customer created, no artifact manifest changed, floor untouched.