romm: 768M, because 5.3.0 does not fit in 512M (R-635)
gates / gates (push) Successful in 1s

No image: line moved, so no catalog_since moved.

Measured on demo-hp after this morning's promotion: OOMKilled true, 4530 gunicorn worker SIGKILLs
in six hours, ~500% CPU in a permanent restart storm, host load 5.2 while otherwise idle. It ran
clean for two hours first, which is why the walk on the scratch guest did not catch it.

The update reported `done` and the app read `running` the whole time - nginx answers 200 while the
workers behind it die. Nothing alarmed; the operator heard the fans.

768M is a measured first step, not a final answer.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-22 17:46:24 +02:00
parent 1ad1f34b6e
commit 886956dcc7
3 changed files with 27 additions and 3 deletions
+20
View File
@@ -1,3 +1,23 @@
## romm gets 768M — 5.3.0 does not fit in 512M (2026-09-22, R-635)
**No `image:` line moved, so no `catalog_since` moved.** `romm` 512M -> **768M**; the header and
`.felhom.yml` totals follow (1024M -> 1280M).
**Measured on demo-hp, not reasoned about.** `romm 5.0.0 -> 5.3.0` was promoted that morning after
the edge was proven on the scratch guest, and the guarded Update on demo-hp reached `done` in 74.8 s.
It ran clean for two hours. Then: `OOMKilled: true`, **4,530** `Worker … was sent SIGKILL! Perhaps
out of memory?` in six hours, the container pinned at 457 MiB of its 512 MiB limit, **~500% CPU** in
a permanent restart storm, host load average **5.2 while otherwise idle**.
**Two things this cost that a bigger number alone does not fix, both recorded in R-635.** The update
read `done` and the app read `running` throughout, because nginx answers `GET /` with 200 while the
gunicorn workers behind it are being killed — the probe is right, the port is right, and the green is
still false. And nothing alarmed: OOM signals are invisible inside an LXC guest (R-528), and a
running-but-thrashing app is not "down". **The operator heard the fans. That was the detector.**
768M is a first measured step, not a final answer — the soak that follows this entry is what settles
whether it is enough.
## The probe gate now runs in CI's own degraded mode (2026-09-22, R-618)
**Caught by checking the push's CI run by job id, not by assuming it went green.** Job **877** on