controller v0.265.0: R-634 cause fixed, held apps say so, OOM storm alarm, R-647 leftovers
gates / gates (push) Successful in 27s

R-634: a whole-box backup no longer stops/restarts a DEPLOYING app (the
measured cause of containers running under 'not deployed'); StopStack
and StartStack refuse a deploying stack for every caller.
R-625: held badge 'Stopped - restore needed', no Update button.
R-636: kernel oom_kill counter; 20+ in 30 min -> one app_oom_storm.
R-647: held error per reader, copy_holds key, two log wordings.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-23 17:16:57 +02:00
parent 0a3026180a
commit 0054d4bd69
28 changed files with 832 additions and 48 deletions
+14
View File
@@ -650,6 +650,18 @@ the old version back; a power cut while undoing resumes the undo. Reasoning and
`felhom.eu/documentation/architecture/09-update-architecture.md` §6.1a,
`felhom.eu/documentation/audits/undo-bakeoff-2026-09-23/`.
**Held apps say so (v0.265.0, R-625).** While a hold stands, the update badge reads „Megállítva —
visszaállítás szükséges" / "Stopped — restore needed" (`tag-error`, title = the hold sentence's first
sentence, in the reader's language) and no Update button is rendered; the API still answers 409 `held`. A
held update's error is stored as the key `update.error.held` and rendered per reader, so a reader in the
other language than the box reads the hold in theirs (R-647).
**Deploys are not interrupted (v0.265.0, R-634).** A deploy still running is not in any backup leg's app
list, the volume leg asks again right before it stops an app, and `StopStack` / `StartStack` refuse a
deploying stack (`ErrStackDeploying`) for every caller. Measured cause: the whole-box backup's `compose down`
and second `compose up -d` in the middle of the deploy's own `up` — the deploy then recorded „not deployed"
over running containers.
**The household is told (v0.264.0).** An undone update sends `app_update_undone` (warning) ONCE; an
update that ends held sends `app_update_held` (error) ONCE — also when the hold itself could not be
saved. Details `{app, stack_name, from, to, at, copy_tier, copy_date, copy_holds}`. Both are in
@@ -1132,6 +1144,7 @@ The nightly backup has two phases that run sequentially. All paths are **per-dri
> **Absent-storage tier skip (v0.243.0, R-518).** A tier whose storage the agent (≥ 0.131.0) reports `absent` is dropped from the manual and the scheduled run before anything is stopped (`quiesce.skipAbsentTiers`), logged, and reported once as `backup_tier_skipped`. `unknown` or a legacy agent is never skipped.
> **OOM visibility (v0.243.0, R-514).** The 30 s dead-app check also reads `State.OOMKilled` for running app containers (`Manager.ScanOOMKilled`); the dashboard shows „Memória elfogyott" and the hub gets `app_oom` once per container run. The controller does not restart the app.
> **OOM storm (v0.265.0, R-636).** For a flagged container the scan also reads the kernel's own kill counter (`memory.events` `oom_kill`, plus `memory.max` / `memory.peak`) with one `docker exec … cat` — `OOMKilled` is sticky, so it cannot count. When the SAME container run's counter rises by **20 or more within 30 minutes**, the notifier sends **`app_oom_storm`** (severity `error`, operator-only, details `{app, container, kills, window_min, mem_limit, peak}`) — ONCE per container run. `app_oom` itself is unchanged. An unreadable counter (an image without `cat`) never escalates.
> **Multi-tier whole-guest backup (v0.174.0, R-82 Slice B).** The agent can serve SEVERAL whole-guest
> backup tiers with independent cadences — "local daily + PBS weekly" (agent >= v0.97.0,
@@ -2432,6 +2445,7 @@ The controller pushes structured events to the Hub's `/api/v1/event` endpoint. T
| `app_start_failed` | **warning** | A DEPLOYED app is not running (fix-3) — fired ONCE per running→down transition. **Customer-switchable („Alkalmazás nem fut"), OFF by default; the OPERATOR is e-mailed regardless.** Was `warn` until v0.223.0 — see the severity note below |
| `app_update_undone` | warning | v0.264.0 — a guarded update failed and the box put the previous version and its data back. Once per app per failed step. Household ON by default |
| `app_update_held` | **error** | v0.264.0 — a guarded update (or its undo) failed and the app is held until a restore. The mail's line is the hold sentence in the household's language. Household ON by default |
| `app_oom_storm` | **error** | v0.265.0 (R-636) — the same container run OOM-killed 20+ times in 30 min (the kernel's `oom_kill` counter). Once per container run. Operator-only |
| `disaster_recovery_started` | warning | DR restore begins |
| `disaster_recovery_completed` | info/error | DR restore finishes (success/partial) |