Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
4.5 KiB
R-528 — OOM kills not reported — design proposal (burn-down night 2026-10-05, no code)
Baselines read: felhom-controller ef199c5, felhom-agent 861d32a, felhom.eu b37902ce. Architecture: 08-alarm-ladder.md (OOM rung, lines 224-247). Memory note: lxc-docker-oom-signals-unreliable.
1. The problem
The out-of-memory alarm starts from one Docker flag, and inside a Felhom guest that flag sometimes stays false after a real kill. Measured 2026-09-15 on 9202: Paperless at 128M restarted 11 times and a memory hog was killed (rc 137), with OOMKilled=false and no oom event (audits/evidence-p1fixes-2026-09-15/E2-oom-signal-measure-9202.txt); the same on the drill box 2026-09-16 (audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt). Measured the other way: romm on demo-hp (2026-09-22) and on 9202 (2026-09-23) read true, and the alarm reached the operator's inbox.
2. What the code does today (read in source)
- Every 30 s the box runs one
docker inspectand keeps only containers withOOMKilled=true(stacks/oom.go:55-66, scheduled atcmd/controller/main.go:929, 940). - The kernel's real kill counter (
memory.eventsoom_kill, read bydocker exec … catinside the container) is read ONLY for those flagged containers (oom.go:71-77). - So when the flag is false, nothing reads the counter: no
app_oom, noapp_oom_storm.08line 238 states this limit. - Partly covered since controller v0.269.0: a crash loop (≥ 6 restarts in 10 min) stops the app and alarms as
app_stopped_unhealthy(08line 246). It does not say "memory". - Why the flag is set on some runs and not others is unknown. Searched: the two evidence files above, the memory note,
08.
3. Options
A. Read the counter for every running container, not only flagged ones. Same read, already proven live on 9202 (oom_kill 8 → 49).
- Changes: drop the flag gate in
oom.go; read unflagged containers every 10th scan (5 min) to bound cost. - Costs: one
docker execper container per read (≈ 20-40 on a full box; cost per exec not measured). - Can go wrong: images with no
catreturn nothing (reads as unknown, not zero). Inferred: a container whose MAIN process is killed exits, its cgroup and counter go with it — that shape stays invisible here. - Measure first: a memory hog killed inside a running container with the flag false — does the counter rise?
B. The agent reads the guest's cgroups from the Proxmox host. A parent cgroup's counter survives a container's death (inferred from cgroup v2 rules), so it also sees main-process kills.
- Costs: a new agent read, a new local-API field, a controller client, a
MinAgentraise. A mechanism nobody has measured. - Measure first: the host-side cgroup path of a guest's Docker container, and whether the guest-wide counter rises on each kill.
C. Name the crash loop "probably memory". The same docker inspect adds ExitCode; exit 137 with no stop from us → the crash-loop alarm text says "probably out of memory".
- Costs: small. Can go wrong: any other SIGKILL reads as memory too — the text must say "probably".
4. The pick — PROPOSAL for the operator, not a decision
A first, then C. A uses a read the box already does and proved live; it closes the worker-kill case (the BIGNIGHT Paperless shape) when the flag lies. C closes the main-process case cheaply, on an alarm that already fires. B is the complete answer but is a new mechanism on the operator-tier agent; it should wait for a measurement that shows A + C miss real kills.
5. First slice and its proof
- Red test first (must FAIL today): the fake
execCommandanswersinspectwithOOMKilled=falseand the container'smemory.eventswithoom_kill 3. AssertScanOOMKilledreturns that container withKills=3. Today it returns nothing (oom.go:63, the flag test). - Second test:
oom_kill 0on every container → nothing returned and no alarm (no false positive). - Live proof on 9202, throwaway app only: run a memory hog in a child process of a running container under a low cap. Positive observable:
app_oomin the hub's Events tab. Control from a different channel:docker exec <c> cat /sys/fs/cgroup/memory.eventsread by hand, and theOOMKilledflag recorded beside it. Evidence off the box before teardown; teardown on box, host and hub.
6. Open questions for the operator
- Is one extra
docker execper container every 5 minutes acceptable on a small box? If you do nothing: kills the flag misses stay silent, as today. - Should the host-side agent read (B) be spiked now, or only after A + C have run a week?