Files
felhom.eu/documentation/audits/night-burndown-2026-10-05/design-R-528.md
T
2026-10-05 23:32:12 +02:00

4.5 KiB

R-528 — OOM kills not reported — design proposal (burn-down night 2026-10-05, no code)

Baselines read: felhom-controller ef199c5, felhom-agent 861d32a, felhom.eu b37902ce. Architecture: 08-alarm-ladder.md (OOM rung, lines 224-247). Memory note: lxc-docker-oom-signals-unreliable.

1. The problem

The out-of-memory alarm starts from one Docker flag, and inside a Felhom guest that flag sometimes stays false after a real kill. Measured 2026-09-15 on 9202: Paperless at 128M restarted 11 times and a memory hog was killed (rc 137), with OOMKilled=false and no oom event (audits/evidence-p1fixes-2026-09-15/E2-oom-signal-measure-9202.txt); the same on the drill box 2026-09-16 (audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt). Measured the other way: romm on demo-hp (2026-09-22) and on 9202 (2026-09-23) read true, and the alarm reached the operator's inbox.

2. What the code does today (read in source)

  • Every 30 s the box runs one docker inspect and keeps only containers with OOMKilled=true (stacks/oom.go:55-66, scheduled at cmd/controller/main.go:929, 940).
  • The kernel's real kill counter (memory.events oom_kill, read by docker exec … cat inside the container) is read ONLY for those flagged containers (oom.go:71-77).
  • So when the flag is false, nothing reads the counter: no app_oom, no app_oom_storm. 08 line 238 states this limit.
  • Partly covered since controller v0.269.0: a crash loop (≥ 6 restarts in 10 min) stops the app and alarms as app_stopped_unhealthy (08 line 246). It does not say "memory".
  • Why the flag is set on some runs and not others is unknown. Searched: the two evidence files above, the memory note, 08.

3. Options

A. Read the counter for every running container, not only flagged ones. Same read, already proven live on 9202 (oom_kill 8 → 49).

  • Changes: drop the flag gate in oom.go; read unflagged containers every 10th scan (5 min) to bound cost.
  • Costs: one docker exec per container per read (≈ 20-40 on a full box; cost per exec not measured).
  • Can go wrong: images with no cat return nothing (reads as unknown, not zero). Inferred: a container whose MAIN process is killed exits, its cgroup and counter go with it — that shape stays invisible here.
  • Measure first: a memory hog killed inside a running container with the flag false — does the counter rise?

B. The agent reads the guest's cgroups from the Proxmox host. A parent cgroup's counter survives a container's death (inferred from cgroup v2 rules), so it also sees main-process kills.

  • Costs: a new agent read, a new local-API field, a controller client, a MinAgent raise. A mechanism nobody has measured.
  • Measure first: the host-side cgroup path of a guest's Docker container, and whether the guest-wide counter rises on each kill.

C. Name the crash loop "probably memory". The same docker inspect adds ExitCode; exit 137 with no stop from us → the crash-loop alarm text says "probably out of memory".

  • Costs: small. Can go wrong: any other SIGKILL reads as memory too — the text must say "probably".

4. The pick — PROPOSAL for the operator, not a decision

A first, then C. A uses a read the box already does and proved live; it closes the worker-kill case (the BIGNIGHT Paperless shape) when the flag lies. C closes the main-process case cheaply, on an alarm that already fires. B is the complete answer but is a new mechanism on the operator-tier agent; it should wait for a measurement that shows A + C miss real kills.

5. First slice and its proof

  • Red test first (must FAIL today): the fake execCommand answers inspect with OOMKilled=false and the container's memory.events with oom_kill 3. Assert ScanOOMKilled returns that container with Kills=3. Today it returns nothing (oom.go:63, the flag test).
  • Second test: oom_kill 0 on every container → nothing returned and no alarm (no false positive).
  • Live proof on 9202, throwaway app only: run a memory hog in a child process of a running container under a low cap. Positive observable: app_oom in the hub's Events tab. Control from a different channel: docker exec <c> cat /sys/fs/cgroup/memory.events read by hand, and the OOMKilled flag recorded beside it. Evidence off the box before teardown; teardown on box, host and hub.

6. Open questions for the operator

  1. Is one extra docker exec per container every 5 minutes acceptable on a small box? If you do nothing: kills the flag misses stay silent, as today.
  2. Should the host-side agent read (B) be spiked now, or only after A + C have run a week?