Clean-up evening: R-634 cause fixed, held badge, OOM storm; floor 0.265.0
gates / gates (push) Successful in 26s

R-634, R-625, R-636, R-647, R-648 closed; R-649 (operator question) and
R-650 opened. Open rows 334 -> 331. 08 §6.2 storm rung; 09 §6.4 parts
8-9 SHIPPED. Evidence: audits/cleanup-2026-09-23/.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-23 17:40:01 +02:00
parent 3add9fa678
commit 210394ec6e
38 changed files with 3121 additions and 115 deletions
@@ -211,6 +211,22 @@ minutes later (90 m instead of 60). **One value, everywhere:** both staleness ch
hardcoded 30 m / 1 h until then; pinned by `TestControllerStatus_FollowsConfiguredThreshold`). A running
hub prints it at startup (`node_stale after 45m0s, node_down after 1h30m0s`).
**An out-of-memory storm gets its own, louder rung (R-636, controller v0.265.0 / hub v0.121.0, 2026-09-23).**
`app_oom` stays exactly as it was: `warning`, operator-only, ONCE per container run (R-514) — the once
is what stops a crash loop from mailing thousands of times (R-629). Beside it, `app_oom_storm`:
| event | severity | who | minted by | why that audience |
|---|---|---|---|---|
| `app_oom_storm` | **error** | **operator only** | controller v0.265.0, when the kernel's `oom_kill` counter of the SAME container run rises by **≥ 20 within 30 min**; once per run | raw container names and memory figures; the household's side is the dashboard tag |
**Why a counter and not the flag:** Docker's `OOMKilled` is sticky — true for the whole run after ONE
kill — so "the key re-fired N times" measures only how long ago the first kill was. The kernel's
`memory.events` `oom_kill` counts kills. **Why 20 in 30 minutes:** RomM's measured rate on 2026-09-22
was 4,530 kills in six hours ≈ 375 per 30 min; a hiccup is 1–3. Live on 9202 (2026-09-23): RomM at 320M
reached 21 kills 2.5 min after start and sent ONE storm; at 49 kills still one. **Not an app-down state:**
the app still reads `running` and `IsDownState` is unchanged (§4). **Limit:** a container whose
`OOMKilled` flag stays false in an LXC guest (R-528) is never read, so it never storms.
**Two event types added 2026-09-17, with who receives them:**
| event | severity | who | minted by | why that audience |
@@ -221,6 +237,8 @@ hub prints it at startup (`node_stale after 45m0s, node_down after 1h30m0s`).
| Family | Grain | Key carries | Why |
|---|---|---|---|
| app down (`app_start_failed`) | **per APP** | `…:<stack_name>` | no digest exists for it — see below |
| update outcome (`app_update_undone`, `app_update_held`) | **per APP**, both legs | `…:<stack_name>` | hub v0.120.0 — one mail per app per outcome; the household leg has its own register (`perAppCustomerCooldownEvents`) |
| OOM storm (`app_oom_storm`) | **per APP**, operator | `…:<stack_name>` | hub v0.121.0 — no digest; the controller already sends it at most once per container run |
| backup run (`backup_run_failures`) | per RUN | `…:<run_id>` | a digest already lists every failing app; one per run |
| tiered backup (`whole_guest_backup_failed`, …) | per TIER | `…:<tier>` | the tiers fail independently and mean different things |
| everything else, incl. `crossdrive_failed` and `backup_integrity_failed` | per TYPE, per hour | — | coarse **on purpose** |
@@ -1049,8 +1049,8 @@ what the part can do to a household's data if it is wrong, not how likely that i
| **5** | **The ladder on the box.** Read `update_ladder:` from the clone, find the installed step, apply ONE step with its OWN definition (`steps/<to>.yml`, the last step the current template), repeat next night; a failed step stops the ladder for that app. `CatalogOrder` compares refs with the digest stripped (see part 7). | 14 | **2.5** | 1, 4 | medium — each step is the guarded update + undo; the new risk is rendering the wrong step's definition, pinned by a test per step shape |
| **6** | **Digests.** The catalog records `sha256` per pin at push time (`check-image-resolvable.py` already resolves it); the box compares it for the badge and renders `name:tag@sha256:…` when present. **Measured 2026-09-23 on 9202:** Docker and Compose both pull and run `redis:7-alpine@sha256:858f…`, and refuse a digest that does not exist (`audits/update-rulings-2026-09-23/70-…`). **Build trap, read from source:** `splitImageRef` returns "unorderable" for ANY ref containing `@` (`updateorder.go:134`), so the digest must be split off before ordering or every digest-pinned app reads Unknown. A digest gone upstream fails the PULL — Scenario E, pin back, nothing ran. | 17, R-446 | **2** | 4 (the entry carries the digest) | low |
| **7** | **The update leg in the chain + the automatic caller + the switch.** A leg that starts when the off-site leg has FINISHED (legs are clock-scheduled today, not chained — a completion signal is new), one app at a time (there is no single-flight, §3b Q4), `app_update.unattended` default ON, `stacks.update_window` removed, reads `UpdateRefusal.Reason`, remembers a failed step so it never re-presses it. **See the one open point below.** | 11, 12 | **3** | 1, 2, 5 | medium — the only part that acts with nobody watching; everything above is what makes it safe |
| **8** | **R-636** — the same OOM key re-firing escalates instead of staying one `warning` for six hours. | R-636 | **1** | — | none |
| **9** | **R-625** — a held app stops offering an Update it will refuse. With the undo, holds become rarer; the lie on the page does not go away by itself. | R-625 | **0.5** | 1 | none |
| **8** | **SHIPPED — controller v0.265.0 + hub v0.121.0, proven live on 9202 2026-09-23** (`audits/cleanup-2026-09-23/`; a kernel `oom_kill` counter, not the sticky flag — `08` §6.2). **R-636** — the same OOM key re-firing escalates instead of staying one `warning` for six hours. | R-636 | **1** | — | none |
| **9** | **SHIPPED — controller v0.265.0, proven live on 9202 2026-09-23** (badge „Megállítva — visszaállítás szükséges" / "Stopped — restore needed", no Update button, 409 `held` unchanged). **R-625** — a held app stops offering an Update it will refuse. With the undo, holds become rarer; the lie on the page does not go away by itself. | R-625 | **0.5** | 1 | none |
| **10** | **PostgreSQL majors converted by the box.** A guarded-update step: `pg_dumpall` from the old engine, a NEW datadir (the old one kept aside, never deleted, until the check passes), load, check; then each of the eleven apps proven on the bench before its catalog move. | 16, R-463 | **2 + 3** | 1 (the same load discipline), 4 | **HIGH** — it rebuilds the datadir; bounded by keeping the old datadir aside |
| **11** | **Fleet view** — per compose service: installed ref, catalog ref, badge state in the report; the hub lists boxes behind. | 18, R-451 | 2 | — | none — **deferred by the ruling** until the fleet grows |