Files
felhom.eu/documentation/audits/DRILL-cooldown-grain-2026-08-23
admin ebdc04601d
gates / gates (push) Successful in 15s
docs(hub v0.108.0): the delivery grain, the cooldown ruling, and gate 11's first subject
The alarm ladder gains §6.2 - which events are per-app, per-run, per-tier or
coarse, and why the default is coarse. CONTEXT records two rulings: the grain is
allow-listed rather than inferred from the payload, with crossdrive_failed as
the proof that a payload rule would have been wrong; and a finding recorded only
in REPORT.md has a lifetime of one session.

R-389 closed and compressed, keeping its rules and naming the commit whose
git show returns the full text. R-390 and R-391 left open.

REPORT.md is gate 11's first real subject and passes: six observations, two
FILED, four NOT-A-FINDING with their reasons. Three of those declarations are
things a tidier report would have omitted - the gate's own spec would have
passed the item it was built to catch, the burst has no ceiling, and ArgoCD
said "successfully rolled out" while still running the old image.

STATUS carries forward the one thing outstanding: the controller floor still
reads 0.222.0 while the golden reads 0.223.0.
2026-08-23 14:12:23 +02:00
..

DRILL — R-389: only the first broken app per hour reached the operator (2026-08-23)

Hub v0.107.0 → v0.108.0. No controller release, so no golden and no floor. Live leg on demo-hp (Tier 0) plus the hub. UNATTENDED. Method: the hub's own SQLite records plus the controller log. Guest and hub DB are UTC; the hub pod logs CEST.

No halt condition fired.

The pair that says it

Same box, same shape — two apps down four minutes apart, inside one hour:

2026-08-23 09:27:51  sent        BookStack     <- v0.107.0
2026-08-23 09:31:51  suppressed  PrivateBin        operator cooldown 1h, key=demo-hp:app_start_failed

2026-08-23 11:56:57  sent        OpenGist      <- v0.108.0
2026-08-23 12:00:57  sent        Calibre-Web

2 sent / 0 suppressed, where the day before the identical shape gave one of each.

And the keys, from the hub's own suppression rows — it records the key only when it declines:

operator cooldown 1h, key=demo-hp:app_start_failed:opengist
operator cooldown 1h, key=demo-hp:app_start_failed:calibre-web

against v0.107.0's shared key=demo-hp:app_start_failed.

The fence, and why it is not decoration

crossdrive_failed is severity error, reaches the operator leg, and carries stack_name — through CrossDriveDetails, a different struct from AppDetails. A rule of the form "if the details carry a stack_name, split per app" would have split it and silently undone R-182.

Proven live: two different apps' crossdrive_failed, one minute apart →

sent        opengist
suppressed  calibre-web    operator cooldown 1h, key=demo-hp:crossdrive_failed

Byte-identical to the derived v0.107.0 key. No app suffix. That is why cooldownStackSuffix takes the event type as well as the details, unlike its two siblings.

Part 2 — the burst, measured

Three apps stopped in one scan (kimai, romm, paperless-ngx, none with a live cooldown — checked first, because a stale one would have halved the count and made the answer look better than it is):

attempted 3
sent 3
suppressed 0

Judgement: per-app is the right grain, and this volume is acceptable. The reference box has 8 deployed apps, so a total outage is 8 mails; the boot grace (90 s), the quiesce grace (180 s) and the per-app edge trigger absorb reboots, backup cycles and persistently-dead apps. No digest row was filed. The reopening condition is stated rather than left implicit: it scales linearly with app count and has no ceiling, so a box large enough that a total outage is unreadable is the point at which the answer becomes a digest with a customer message — not a wider cooldown.

Part 3 — gate 11, and the spec discrepancy it forced

The gate refuses a push whose REPORT.md carries an observation with neither FILED: R-NNN nor NOT-A-FINDING: <reason>.

The specification said an item may "cite an R-NNN that resolves". That rule would have passed the very item the gate was built to catch. Yesterday's lost observation reads "This is R-182's known cooldown-key shape…" — R-182 resolves, and it is cited as an analogy, not as the row that files it. No parser can tell citation-as-precedent from citation-as-filing by reading prose. The marker is therefore explicit, and the discrepancy is recorded in the gate's docstring rather than quietly resolved. EDGE 7 in the evidence is that exact case, convicted.

Control Expected Observed
historical: yesterday's real section, verbatim refuse exit 1, naming both items
plant an observation with no row refuse exit 1
add the row pass exit 0
remove the observation pass quietly exit 0
FILED: a row that does not resolve refuse exit 1
NOT-A-FINDING: with no reason refuse exit 1
both markers on one item refuse exit 1
section present, no numbered items inconclusive exit 2, saying what it could not read
no REPORT.md at all inconclusive exit 2
a bare R-182 mention (the trap) refuse exit 1

Evidence index (evidence/)

File What it shows
redproof-1-key.txt suffix dropped from the key → 1 operator mail(s), want 2, and the live key shape reproduced
redproof-2-allowlist.txt allow-list removed → crossdrive_failed, app_deployed, app_removed, backup_failed all split per app
gate11-01-historical-redproof.txt yesterday's actual observations section, refused
gate11-02-three-controls.txt plant → refuse, file → pass, remove → pass
gate11-03-edges.txt seven boundaries incl. the R-182 trap
live-01…live-05 Scenarios A and B: both sent, then each suppressed under its own key
live-06, live-07 Scenario C: crossdrive stays coarse
live-08, live-09 the burst, with its pre-check
live-10-full-controller-log.txt 1803 lines, pulled before the restore

Teardown

Nothing provisioned. Five apps were stopped across the walk (opengist, calibre-web, kimai, romm, paperless-ngx) and all were restarted and confirmed healthy — 17 containers up. No app was rebuilt, redeployed or restored; the three retained subjects (docmost, bookstack, privatebin) were not touched at all.

Hub-side, stated explicitly. The hub was written this session: deployment to v0.108.0 via the manifest, and six probe events were POSTed to the live hub for Scenarios B and C (two app_start_failed, two crossdrive_failed, plus the two from yesterday's Scenario H that were deliberately left in place). They are inert event rows for demo-hp and are named here rather than left to be found. Every count in this drill is filtered by created_at >= T0 precisely so those rows cannot contaminate it. Nothing else: no appliance registered, no customer created, no artifact manifest changed, floor untouched.