From f4eb94f19517ea26dd9f296b18a6782951e8dec7 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Tue, 22 Sep 2026 17:55:01 +0200 Subject: [PATCH] romm: two web workers, not four - the actual cause (R-635) WEB_SERVER_CONCURRENCY=2, limit left at 768M. No image: line moved. 768M was still a guess and it was wrong: kills slowed from ~12/min to ~7/min and stopped nothing. Measured instead - each warm uvicorn worker holds ~216 MiB, so four plus the master reach ~882 MiB, and the cgroup's memory.peak read exactly 768 MiB. /init:143 runs --workers "${WEB_SERVER_CONCURRENCY:-4}". Four is a server default; this is one household on one small box. Two measure ~450 MiB, fit 768M with headroom, and halve the CPU churn. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- CHANGELOG.md | 29 +++++++++++++++++++++++++++++ templates/romm/docker-compose.yml | 7 +++++++ 2 files changed, 36 insertions(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index 2481280..69b2cf8 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,3 +1,32 @@ +## romm runs TWO web workers, not four — the actual cause (2026-09-22, R-635) + +`WEB_SERVER_CONCURRENCY=2`, with the limit left at the 768M of the previous entry. **No `image:` +line moved.** + +**768M was still a guess, and the guess was wrong** — it slowed the kills from ~12/min to ~7/min and +stopped nothing. The number came from measuring instead: + +``` +pid 2318697 RSS 216 MiB <- warm uvicorn worker +pid 2318717 RSS 215 MiB <- warm uvicorn worker +pid 2319627 RSS 63 MiB <- just restarted after a kill +pid 2319629 RSS 62 MiB <- just restarted after a kill +pid 2315417 RSS 18 MiB <- gunicorn master +``` + +Four warm workers plus the master reach **~882 MiB** before nginx and the job runner in the same +container, and the cgroup's own `memory.peak` read **exactly 768 MiB** — it hit the new ceiling and +was killed there. + +**The lever was in the image all along:** `/init:143` runs +`--workers "${WEB_SERVER_CONCURRENCY:-4}"`. **Four workers is a server default on an appliance +serving one household.** Two measure ~450 MiB and fit 768M with headroom, and they halve the CPU +churn as well — the host's load average had been sitting at **5.2 while otherwise idle**. + +**The lesson is not about romm.** A version move is not only an `image:` line: the new version's +*shape* — worker counts, per-worker footprint — has to be measured too, and a walk that lasts +minutes cannot see a ceiling that is reached in two hours. + ## romm gets 768M — 5.3.0 does not fit in 512M (2026-09-22, R-635) **No `image:` line moved, so no `catalog_since` moved.** `romm` 512M -> **768M**; the header and diff --git a/templates/romm/docker-compose.yml b/templates/romm/docker-compose.yml index f09121f..a201e38 100644 --- a/templates/romm/docker-compose.yml +++ b/templates/romm/docker-compose.yml @@ -55,6 +55,13 @@ services: - REDIS_HOST=romm-redis - REDIS_PORT=6379 - ROMM_PORT=8080 + # 2, not the image's default of 4. MEASURED on demo-hp 2026-09-22 (R-635): each warm + # uvicorn worker holds ~216 MiB, so four of them plus the master reach ~882 MiB and the + # container was OOM-killed at both 512M and 768M — 37 worker SIGKILLs in five minutes, + # ~500% CPU, host load 5.2 while otherwise idle. Four workers is a SERVER default; this + # is one household on one small box. Two workers measure ~450 MiB and fit 768M with + # real headroom. `/init:143` reads this variable: --workers "${WEB_SERVER_CONCURRENCY:-4}". + - WEB_SERVER_CONCURRENCY=2 - IGDB_CLIENT_ID=${IGDB_CLIENT_ID:-} - IGDB_CLIENT_SECRET=${IGDB_CLIENT_SECRET:-} - STEAMGRIDDB_API_KEY=${STEAMGRIDDB_API_KEY:-}