docs(hub v0.108.0): the delivery grain, the cooldown ruling, and gate 11's first subject
gates / gates (push) Successful in 15s
gates / gates (push) Successful in 15s
The alarm ladder gains §6.2 - which events are per-app, per-run, per-tier or coarse, and why the default is coarse. CONTEXT records two rulings: the grain is allow-listed rather than inferred from the payload, with crossdrive_failed as the proof that a payload rule would have been wrong; and a finding recorded only in REPORT.md has a lifetime of one session. R-389 closed and compressed, keeping its rules and naming the commit whose git show returns the full text. R-390 and R-391 left open. REPORT.md is gate 11's first real subject and passes: six observations, two FILED, four NOT-A-FINDING with their reasons. Three of those declarations are things a tidier report would have omitted - the gate's own spec would have passed the item it was built to catch, the burst has no ceiling, and ArgoCD said "successfully rolled out" while still running the old image. STATUS carries forward the one thing outstanding: the controller floor still reads 0.222.0 while the golden reads 0.223.0.
This commit is contained in:
@@ -155,6 +155,48 @@ it is deliberately absent from `DefaultEnabledEvents`, and deliberately **not**
|
||||
|
||||
---
|
||||
|
||||
## 6.2 The delivery grain — how often, and per what [DESIGN, R-389 — hub v0.108.0]
|
||||
|
||||
**"How loud" is a separate decision from "does it alarm", and it is made in one place**: the operator
|
||||
cooldown key at `processOperator`. Everything sharing a key is collapsed for **one hour**.
|
||||
|
||||
| Family | Grain | Key carries | Why |
|
||||
|---|---|---|---|
|
||||
| app down (`app_start_failed`) | **per APP** | `…:<stack_name>` | no digest exists for it — see below |
|
||||
| backup run (`backup_run_failures`) | per RUN | `…:<run_id>` | a digest already lists every failing app; one per run |
|
||||
| tiered backup (`whole_guest_backup_failed`, …) | per TIER | `…:<tier>` | the tiers fail independently and mean different things |
|
||||
| everything else, incl. `crossdrive_failed` | per TYPE, per hour | — | coarse **on purpose** |
|
||||
|
||||
**The default is COARSE and that is deliberate.** R-97a and R-182 exist so that one full disk produces
|
||||
**one** mail listing every affected app rather than one per app. Widening the grain is what makes an
|
||||
operator stop reading their alerts, which is the same failure as not sending them.
|
||||
|
||||
**`app_start_failed` is the exception because it has no digest.** There is no `apps_down_run`
|
||||
summarising a scan the way `backup_run_failures` summarises a run, so per-app is the only grain
|
||||
available that does not lose alarms. Until hub v0.108.0 it was keyed per type, and **only the first
|
||||
broken app per hour reached the operator** — measured 2026-08-23: `bookstack` sent at 09:27:51,
|
||||
`privatebin` suppressed at 09:31:51 under `key=demo-hp:app_start_failed`.
|
||||
|
||||
**The mechanism is a named ALLOW-LIST (`perAppCooldownEvents`), not a payload rule**, and the
|
||||
distinction is load-bearing rather than stylistic: **`crossdrive_failed` is severity `error`, reaches
|
||||
the operator leg, and carries `stack_name`** through `CrossDriveDetails`. A rule of the form "if the
|
||||
details carry a stack_name, split per app" would have split it, silently, and undone R-182.
|
||||
`cooldownStackSuffix` therefore takes the **event type** as well as the details — an asymmetry with
|
||||
its two siblings, and the reason for it is exactly this.
|
||||
|
||||
**Fenced act:** adding an entry to `perAppCooldownEvents` for a type whose family has a digest, or
|
||||
whose coarse cooldown is deliberate. Reading the register anywhere is fine.
|
||||
|
||||
**Measured burst, so the volume is a number and not an impression:** three apps stopped in one scan
|
||||
produced **three attempted, three sent, zero suppressed** (2026-08-23). The reference box has 8
|
||||
deployed apps, so a total outage is 8 mails. The boot grace (90 s), the quiesce grace (180 s) and the
|
||||
per-app edge trigger absorb reboots, backup cycles and persistently-dead apps. **This scales linearly
|
||||
with app count and has no ceiling** — the condition that would reopen the question is a box large
|
||||
enough that a total outage is unreadable, at which point the answer is a digest with a customer
|
||||
message, not a wider cooldown.
|
||||
|
||||
---
|
||||
|
||||
## 7. The intent test [DESIGN, R-386 — CLOSED controller v0.223.0]
|
||||
|
||||
**"The customer stopped this" is asked of the FIELD THAT RECORDS IT, never inferred from the state.**
|
||||
|
||||
Reference in New Issue
Block a user