CORRECTION: the app_oom alarm DID fire - R-635 was wrong, R-636 opened
gates / gates (push) Successful in 26s
gates / gates (push) Successful in 26s
The operator produced the mails. The controller HAS an OOM detector (main.go:821), it emits app_oom (notifier.go:726), the hub allow-lists it (dispatcher.go:636) and delivered it to the OPERATOR channel - two mails, 11:09 and 17:48 CEST, each naming the app and linking the dashboard. CUSTOMER skipped, correctly. I asserted an absence without opening the hub's Events or Notifications tab, reasoning instead from a memory note that said the signal was UNPROVEN - not that it was missing. That is R-628's shape again, from the same hand, four days later. R-636: the real defect is the signal's SHAPE. notifier.go:715-724 keys on container|startedAt and emits once per container lifetime, so 4530 worker kills over six hours produced exactly one warning-level mail - indistinguishable from one transient kill. The magnitude was already collected (App Telemetry: RomM 5023 errors, 632 warnings) but nothing turns it into a louder event. Memory note lxc-docker-oom-signals-unreliable corrected with the positive reading. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -29,7 +29,11 @@
|
||||
|
||||
**Evening addition, 2026-09-22 — you heard the fans, and you were right.**
|
||||
|
||||
**RomM was cooking the HP box and it was my fault.** I moved it to a new version that morning. The update said it worked, and it did — for two hours. Then it ran out of memory and spent six hours killing and restarting its own workers, about seven times a minute, burning five processor cores. Nothing warned anyone. **You found it by ear. That was the only detector that worked.**
|
||||
**RomM was cooking the HP box and it was my fault.** I moved it to a new version that morning. The update said it worked, and it did — for two hours. Then it ran out of memory and spent six hours killing and restarting its own workers, about seven times a minute, burning five processor cores.
|
||||
|
||||
**CORRECTION — the machine DID warn you, and I was wrong to say it did not.** You showed me the two e-mails: *"Alkalmazás memóriája elfogyott: romm"*, at 11:09 and again at 17:48. They are in the hub's Events and Notifications tabs too. **I wrote "nothing warned anyone" without opening either tab** — I went by an old note saying this signal was unproven and turned that into "it did not happen". That is the same mistake I made with the Hetzner tickets four days ago.
|
||||
|
||||
**What is actually wrong is smaller and real:** the machine sends **one** warning per app start. Six hours of trouble and 4,530 worker deaths produced **one e-mail** — the same e-mail a single harmless hiccup would send. It never gets louder, and the app keeps showing as running. The hub *did* have the full picture on the App Telemetry page (RomM: 5,023 errors, 632 warnings, while every other app showed zero), but nothing turns that into a second, louder alert. **That is why a correct warning still got missed, and it is now written down as its own item.**
|
||||
|
||||
**Fixed, and proven under real load.** Giving it more memory was not enough — I measured that rather than assuming it. The real cause was that the app starts **four web workers**, which is a server setting, on a box serving one household. It now starts **two**. I then drove **26,645 real requests** at it for five minutes: memory stayed between 416 and 614 MB against its 768 MB ceiling and **trended down**, with **zero** worker deaths. At rest it now uses **1.6% of a core** instead of 500%.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user