Files
felhom.eu/documentation/audits/night-burndown-2026-10-05/design-R-528.md
T
2026-10-05 23:32:12 +02:00

40 lines
4.5 KiB
Markdown

# R-528 — OOM kills not reported — design proposal (burn-down night 2026-10-05, no code)
Baselines read: felhom-controller `ef199c5`, felhom-agent `861d32a`, felhom.eu `b37902ce`. Architecture: `08-alarm-ladder.md` (OOM rung, lines 224-247). Memory note: `lxc-docker-oom-signals-unreliable`.
## 1. The problem
The out-of-memory alarm starts from one Docker flag, and inside a Felhom guest that flag sometimes stays false after a real kill. Measured 2026-09-15 on 9202: Paperless at 128M restarted 11 times and a memory hog was killed (rc 137), with `OOMKilled=false` and no `oom` event (`audits/evidence-p1fixes-2026-09-15/E2-oom-signal-measure-9202.txt`); the same on the drill box 2026-09-16 (`audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt`). Measured the other way: romm on demo-hp (2026-09-22) and on 9202 (2026-09-23) read `true`, and the alarm reached the operator's inbox.
## 2. What the code does today (read in source)
- Every 30 s the box runs one `docker inspect` and keeps only containers with `OOMKilled=true` (`stacks/oom.go:55-66`, scheduled at `cmd/controller/main.go:929, 940`).
- The kernel's real kill counter (`memory.events` `oom_kill`, read by `docker exec … cat` inside the container) is read ONLY for those flagged containers (`oom.go:71-77`).
- So when the flag is false, nothing reads the counter: no `app_oom`, no `app_oom_storm`. `08` line 238 states this limit.
- Partly covered since controller v0.269.0: a crash loop (≥ 6 restarts in 10 min) stops the app and alarms as `app_stopped_unhealthy` (`08` line 246). It does not say "memory".
- Why the flag is set on some runs and not others is **unknown**. Searched: the two evidence files above, the memory note, `08`.
## 3. Options
**A. Read the counter for every running container, not only flagged ones.** Same read, already proven live on 9202 (`oom_kill` 8 → 49).
- Changes: drop the flag gate in `oom.go`; read unflagged containers every 10th scan (5 min) to bound cost.
- Costs: one `docker exec` per container per read (≈ 20-40 on a full box; cost per exec not measured).
- Can go wrong: images with no `cat` return nothing (reads as unknown, not zero). Inferred: a container whose MAIN process is killed exits, its cgroup and counter go with it — that shape stays invisible here.
- Measure first: a memory hog killed inside a running container with the flag false — does the counter rise?
**B. The agent reads the guest's cgroups from the Proxmox host.** A parent cgroup's counter survives a container's death (inferred from cgroup v2 rules), so it also sees main-process kills.
- Costs: a new agent read, a new local-API field, a controller client, a `MinAgent` raise. A mechanism nobody has measured.
- Measure first: the host-side cgroup path of a guest's Docker container, and whether the guest-wide counter rises on each kill.
**C. Name the crash loop "probably memory".** The same `docker inspect` adds `ExitCode`; exit 137 with no stop from us → the crash-loop alarm text says "probably out of memory".
- Costs: small. Can go wrong: any other SIGKILL reads as memory too — the text must say "probably".
## 4. The pick — PROPOSAL for the operator, not a decision
A first, then C. A uses a read the box already does and proved live; it closes the worker-kill case (the BIGNIGHT Paperless shape) when the flag lies. C closes the main-process case cheaply, on an alarm that already fires. B is the complete answer but is a new mechanism on the operator-tier agent; it should wait for a measurement that shows A + C miss real kills.
## 5. First slice and its proof
- Red test first (must FAIL today): the fake `execCommand` answers `inspect` with `OOMKilled=false` and the container's `memory.events` with `oom_kill 3`. Assert `ScanOOMKilled` returns that container with `Kills=3`. Today it returns nothing (`oom.go:63`, the flag test).
- Second test: `oom_kill 0` on every container → nothing returned and no alarm (no false positive).
- Live proof on 9202, throwaway app only: run a memory hog in a child process of a running container under a low cap. Positive observable: `app_oom` in the hub's Events tab. Control from a different channel: `docker exec <c> cat /sys/fs/cgroup/memory.events` read by hand, and the `OOMKilled` flag recorded beside it. Evidence off the box before teardown; teardown on box, host and hub.
## 6. Open questions for the operator
1. Is one extra `docker exec` per container every 5 minutes acceptable on a small box? If you do nothing: kills the flag misses stay silent, as today.
2. Should the host-side agent read (B) be spiked now, or only after A + C have run a week?