DRILL — R-389: only the first broken app per hour reached the operator (2026-08-23)
Hub v0.107.0 → v0.108.0. No controller release, so no golden and no floor. Live leg on demo-hp
(Tier 0) plus the hub. UNATTENDED. Method: the hub's own SQLite records plus the controller log.
Guest and hub DB are UTC; the hub pod logs CEST.
No halt condition fired.
The pair that says it
Same box, same shape — two apps down four minutes apart, inside one hour:
2026-08-23 09:27:51 sent BookStack <- v0.107.0
2026-08-23 09:31:51 suppressed PrivateBin operator cooldown 1h, key=demo-hp:app_start_failed
2026-08-23 11:56:57 sent OpenGist <- v0.108.0
2026-08-23 12:00:57 sent Calibre-Web
2 sent / 0 suppressed, where the day before the identical shape gave one of each.
And the keys, from the hub's own suppression rows — it records the key only when it declines:
operator cooldown 1h, key=demo-hp:app_start_failed:opengist
operator cooldown 1h, key=demo-hp:app_start_failed:calibre-web
against v0.107.0's shared key=demo-hp:app_start_failed.
The fence, and why it is not decoration
crossdrive_failed is severity error, reaches the operator leg, and carries stack_name — through
CrossDriveDetails, a different struct from AppDetails. A rule of the form "if the details
carry a stack_name, split per app" would have split it and silently undone R-182.
Proven live: two different apps' crossdrive_failed, one minute apart →
sent opengist
suppressed calibre-web operator cooldown 1h, key=demo-hp:crossdrive_failed
Byte-identical to the derived v0.107.0 key. No app suffix. That is why cooldownStackSuffix
takes the event type as well as the details, unlike its two siblings.
Part 2 — the burst, measured
Three apps stopped in one scan (kimai, romm, paperless-ngx, none with a live cooldown — checked
first, because a stale one would have halved the count and made the answer look better than it is):
| attempted | 3 |
| sent | 3 |
| suppressed | 0 |
Judgement: per-app is the right grain, and this volume is acceptable. The reference box has 8 deployed apps, so a total outage is 8 mails; the boot grace (90 s), the quiesce grace (180 s) and the per-app edge trigger absorb reboots, backup cycles and persistently-dead apps. No digest row was filed. The reopening condition is stated rather than left implicit: it scales linearly with app count and has no ceiling, so a box large enough that a total outage is unreadable is the point at which the answer becomes a digest with a customer message — not a wider cooldown.
Part 3 — gate 11, and the spec discrepancy it forced
The gate refuses a push whose REPORT.md carries an observation with neither FILED: R-NNN nor
NOT-A-FINDING: <reason>.
The specification said an item may "cite an R-NNN that resolves". That rule would have passed the
very item the gate was built to catch. Yesterday's lost observation reads "This is R-182's known
cooldown-key shape…" — R-182 resolves, and it is cited as an analogy, not as the row that files
it. No parser can tell citation-as-precedent from citation-as-filing by reading prose. The marker is
therefore explicit, and the discrepancy is recorded in the gate's docstring rather than quietly
resolved. EDGE 7 in the evidence is that exact case, convicted.
| Control | Expected | Observed |
|---|---|---|
| historical: yesterday's real section, verbatim | refuse | exit 1, naming both items |
| plant an observation with no row | refuse | exit 1 |
| add the row | pass | exit 0 |
| remove the observation | pass quietly | exit 0 |
FILED: a row that does not resolve |
refuse | exit 1 |
NOT-A-FINDING: with no reason |
refuse | exit 1 |
| both markers on one item | refuse | exit 1 |
| section present, no numbered items | inconclusive | exit 2, saying what it could not read |
no REPORT.md at all |
inconclusive | exit 2 |
a bare R-182 mention (the trap) |
refuse | exit 1 |
Evidence index (evidence/)
| File | What it shows |
|---|---|
redproof-1-key.txt |
suffix dropped from the key → 1 operator mail(s), want 2, and the live key shape reproduced |
redproof-2-allowlist.txt |
allow-list removed → crossdrive_failed, app_deployed, app_removed, backup_failed all split per app |
gate11-01-historical-redproof.txt |
yesterday's actual observations section, refused |
gate11-02-three-controls.txt |
plant → refuse, file → pass, remove → pass |
gate11-03-edges.txt |
seven boundaries incl. the R-182 trap |
live-01…live-05 |
Scenarios A and B: both sent, then each suppressed under its own key |
live-06, live-07 |
Scenario C: crossdrive stays coarse |
live-08, live-09 |
the burst, with its pre-check |
live-10-full-controller-log.txt |
1803 lines, pulled before the restore |
Teardown
Nothing provisioned. Five apps were stopped across the walk (opengist, calibre-web, kimai,
romm, paperless-ngx) and all were restarted and confirmed healthy — 17 containers up. No app
was rebuilt, redeployed or restored; the three retained subjects (docmost, bookstack,
privatebin) were not touched at all.
Hub-side, stated explicitly. The hub was written this session: deployment to v0.108.0 via the
manifest, and six probe events were POSTed to the live hub for Scenarios B and C (two
app_start_failed, two crossdrive_failed, plus the two from yesterday's Scenario H that were
deliberately left in place). They are inert event rows for demo-hp and are named here rather than
left to be found. Every count in this drill is filtered by created_at >= T0 precisely so those
rows cannot contaminate it. Nothing else: no appliance registered, no customer created, no artifact
manifest changed, floor untouched.