Files
felhom.eu/documentation/audits/DRILL-cooldown-grain-2026-08-23/README.md
T
admin ebdc04601d
gates / gates (push) Successful in 15s
docs(hub v0.108.0): the delivery grain, the cooldown ruling, and gate 11's first subject
The alarm ladder gains §6.2 - which events are per-app, per-run, per-tier or
coarse, and why the default is coarse. CONTEXT records two rulings: the grain is
allow-listed rather than inferred from the payload, with crossdrive_failed as
the proof that a payload rule would have been wrong; and a finding recorded only
in REPORT.md has a lifetime of one session.

R-389 closed and compressed, keeping its rules and naming the commit whose
git show returns the full text. R-390 and R-391 left open.

REPORT.md is gate 11's first real subject and passes: six observations, two
FILED, four NOT-A-FINDING with their reasons. Three of those declarations are
things a tidier report would have omitted - the gate's own spec would have
passed the item it was built to catch, the burst has no ceiling, and ArgoCD
said "successfully rolled out" while still running the old image.

STATUS carries forward the one thing outstanding: the controller floor still
reads 0.222.0 while the golden reads 0.223.0.
2026-08-23 14:12:23 +02:00

119 lines
5.8 KiB
Markdown

# DRILL — R-389: only the first broken app per hour reached the operator (2026-08-23)
**Hub v0.107.0 → v0.108.0. No controller release, so no golden and no floor.** Live leg on `demo-hp`
(Tier 0) plus the hub. UNATTENDED. Method: the hub's own SQLite records plus the controller log.
Guest and hub DB are UTC; the hub pod logs CEST.
**No halt condition fired.**
## The pair that says it
Same box, same shape — two apps down four minutes apart, inside one hour:
```
2026-08-23 09:27:51 sent BookStack <- v0.107.0
2026-08-23 09:31:51 suppressed PrivateBin operator cooldown 1h, key=demo-hp:app_start_failed
2026-08-23 11:56:57 sent OpenGist <- v0.108.0
2026-08-23 12:00:57 sent Calibre-Web
```
**2 sent / 0 suppressed**, where the day before the identical shape gave one of each.
And the keys, from the hub's own suppression rows — it records the key only when it declines:
```
operator cooldown 1h, key=demo-hp:app_start_failed:opengist
operator cooldown 1h, key=demo-hp:app_start_failed:calibre-web
```
against v0.107.0's shared `key=demo-hp:app_start_failed`.
## The fence, and why it is not decoration
`crossdrive_failed` is severity `error`, reaches the operator leg, and carries `stack_name` — through
`CrossDriveDetails`, **a different struct from `AppDetails`**. A rule of the form *"if the details
carry a stack_name, split per app"* would have split it and silently undone R-182.
Proven live: two different apps' `crossdrive_failed`, one minute apart →
```
sent opengist
suppressed calibre-web operator cooldown 1h, key=demo-hp:crossdrive_failed
```
**Byte-identical to the derived v0.107.0 key. No app suffix.** That is why `cooldownStackSuffix`
takes the event type as well as the details, unlike its two siblings.
## Part 2 — the burst, measured
Three apps stopped in one scan (`kimai`, `romm`, `paperless-ngx`, none with a live cooldown — checked
first, because a stale one would have halved the count and made the answer look better than it is):
| | |
|---|---|
| attempted | **3** |
| sent | **3** |
| suppressed | **0** |
**Judgement: per-app is the right grain, and this volume is acceptable.** The reference box has 8
deployed apps, so a total outage is 8 mails; the boot grace (90 s), the quiesce grace (180 s) and the
per-app edge trigger absorb reboots, backup cycles and persistently-dead apps. **No digest row was
filed.** The reopening condition is stated rather than left implicit: it scales linearly with app
count and has no ceiling, so a box large enough that a total outage is unreadable is the point at
which the answer becomes a digest with a customer message — not a wider cooldown.
## Part 3 — gate 11, and the spec discrepancy it forced
The gate refuses a push whose `REPORT.md` carries an observation with neither `FILED: R-NNN` nor
`NOT-A-FINDING: <reason>`.
**The specification said an item may "cite an R-NNN that resolves". That rule would have passed the
very item the gate was built to catch.** Yesterday's lost observation reads *"This is R-182's known
cooldown-key shape…"* — `R-182` resolves, and it is cited as an **analogy**, not as the row that files
it. No parser can tell citation-as-precedent from citation-as-filing by reading prose. The marker is
therefore explicit, and the discrepancy is recorded in the gate's docstring rather than quietly
resolved. `EDGE 7` in the evidence is that exact case, convicted.
| Control | Expected | Observed |
|---|---|---|
| historical: yesterday's real section, verbatim | refuse | **exit 1**, naming both items |
| plant an observation with no row | refuse | **exit 1** |
| add the row | pass | **exit 0** |
| remove the observation | pass quietly | **exit 0** |
| `FILED:` a row that does not resolve | refuse | **exit 1** |
| `NOT-A-FINDING:` with no reason | refuse | **exit 1** |
| both markers on one item | refuse | **exit 1** |
| section present, no numbered items | inconclusive | **exit 2**, saying what it could not read |
| no `REPORT.md` at all | inconclusive | **exit 2** |
| a bare `R-182` mention (the trap) | refuse | **exit 1** |
## Evidence index (`evidence/`)
| File | What it shows |
|---|---|
| `redproof-1-key.txt` | suffix dropped from the key → `1 operator mail(s), want 2`, and the live key shape reproduced |
| `redproof-2-allowlist.txt` | allow-list removed → `crossdrive_failed`, `app_deployed`, `app_removed`, `backup_failed` all split per app |
| `gate11-01-historical-redproof.txt` | yesterday's actual observations section, refused |
| `gate11-02-three-controls.txt` | plant → refuse, file → pass, remove → pass |
| `gate11-03-edges.txt` | seven boundaries incl. the R-182 trap |
| `live-01`…`live-05` | Scenarios A and B: both sent, then each suppressed under its own key |
| `live-06`, `live-07` | Scenario C: crossdrive stays coarse |
| `live-08`, `live-09` | the burst, with its pre-check |
| `live-10-full-controller-log.txt` | 1803 lines, pulled before the restore |
## Teardown
Nothing provisioned. Five apps were stopped across the walk (`opengist`, `calibre-web`, `kimai`,
`romm`, `paperless-ngx`) and **all were restarted and confirmed healthy** — 17 containers up. No app
was rebuilt, redeployed or restored; the three retained subjects (`docmost`, `bookstack`,
`privatebin`) were not touched at all.
**Hub-side, stated explicitly.** The hub was written this session: deployment to v0.108.0 via the
manifest, and **six probe events were POSTed to the live hub** for Scenarios B and C (two
`app_start_failed`, two `crossdrive_failed`, plus the two from yesterday's Scenario H that were
deliberately left in place). They are inert event rows for `demo-hp` and are named here rather than
left to be found. **Every count in this drill is filtered by `created_at >= T0`** precisely so those
rows cannot contaminate it. Nothing else: no appliance registered, no customer created, no artifact
manifest changed, floor untouched.