docs(hub v0.108.0): the delivery grain, the cooldown ruling, and gate 11's first subject
gates / gates (push) Successful in 15s

The alarm ladder gains §6.2 - which events are per-app, per-run, per-tier or
coarse, and why the default is coarse. CONTEXT records two rulings: the grain is
allow-listed rather than inferred from the payload, with crossdrive_failed as
the proof that a payload rule would have been wrong; and a finding recorded only
in REPORT.md has a lifetime of one session.

R-389 closed and compressed, keeping its rules and naming the commit whose
git show returns the full text. R-390 and R-391 left open.

REPORT.md is gate 11's first real subject and passes: six observations, two
FILED, four NOT-A-FINDING with their reasons. Three of those declarations are
things a tidier report would have omitted - the gate's own spec would have
passed the item it was built to catch, the burst has no ceiling, and ArgoCD
said "successfully rolled out" while still running the old image.

STATUS carries forward the one thing outstanding: the controller floor still
reads 0.222.0 while the golden reads 0.223.0.
This commit is contained in:
2026-08-23 14:12:23 +02:00
parent 45659bdc5a
commit ebdc04601d
24 changed files with 2514 additions and 111 deletions
@@ -155,6 +155,48 @@ it is deliberately absent from `DefaultEnabledEvents`, and deliberately **not**
---
## 6.2 The delivery grain — how often, and per what [DESIGN, R-389 — hub v0.108.0]
**"How loud" is a separate decision from "does it alarm", and it is made in one place**: the operator
cooldown key at `processOperator`. Everything sharing a key is collapsed for **one hour**.
| Family | Grain | Key carries | Why |
|---|---|---|---|
| app down (`app_start_failed`) | **per APP** | `…:<stack_name>` | no digest exists for it — see below |
| backup run (`backup_run_failures`) | per RUN | `…:<run_id>` | a digest already lists every failing app; one per run |
| tiered backup (`whole_guest_backup_failed`, …) | per TIER | `…:<tier>` | the tiers fail independently and mean different things |
| everything else, incl. `crossdrive_failed` | per TYPE, per hour | — | coarse **on purpose** |
**The default is COARSE and that is deliberate.** R-97a and R-182 exist so that one full disk produces
**one** mail listing every affected app rather than one per app. Widening the grain is what makes an
operator stop reading their alerts, which is the same failure as not sending them.
**`app_start_failed` is the exception because it has no digest.** There is no `apps_down_run`
summarising a scan the way `backup_run_failures` summarises a run, so per-app is the only grain
available that does not lose alarms. Until hub v0.108.0 it was keyed per type, and **only the first
broken app per hour reached the operator** — measured 2026-08-23: `bookstack` sent at 09:27:51,
`privatebin` suppressed at 09:31:51 under `key=demo-hp:app_start_failed`.
**The mechanism is a named ALLOW-LIST (`perAppCooldownEvents`), not a payload rule**, and the
distinction is load-bearing rather than stylistic: **`crossdrive_failed` is severity `error`, reaches
the operator leg, and carries `stack_name`** through `CrossDriveDetails`. A rule of the form "if the
details carry a stack_name, split per app" would have split it, silently, and undone R-182.
`cooldownStackSuffix` therefore takes the **event type** as well as the details — an asymmetry with
its two siblings, and the reason for it is exactly this.
**Fenced act:** adding an entry to `perAppCooldownEvents` for a type whose family has a digest, or
whose coarse cooldown is deliberate. Reading the register anywhere is fine.
**Measured burst, so the volume is a number and not an impression:** three apps stopped in one scan
produced **three attempted, three sent, zero suppressed** (2026-08-23). The reference box has 8
deployed apps, so a total outage is 8 mails. The boot grace (90 s), the quiesce grace (180 s) and the
per-app edge trigger absorb reboots, backup cycles and persistently-dead apps. **This scales linearly
with app count and has no ceiling** — the condition that would reopen the question is a box large
enough that a total outage is unreadable, at which point the answer is a digest with a customer
message, not a wider cooldown.
---
## 7. The intent test [DESIGN, R-386 — CLOSED controller v0.223.0]
**"The customer stopped this" is asked of the FIELD THAT RECORDS IT, never inferred from the state.**