8efd2d00df
gates / gates (push) Successful in 26s
Found because the operator heard the fans. romm 5.3.0 was promoted that morning; the update read `done` and the app ran clean for two hours, then OOM-crash-looped for six - 4530 worker SIGKILLs, ~500% CPU, host load 5.2 while otherwise idle, and nothing alarmed. Raising the limit to 768M was still a guess and fixed nothing (memory.peak hit exactly 768 MiB). Measured instead: ~216 MiB per warm uvicorn worker, so the image's default of 4 workers needs ~882 MiB. /init reads WEB_SERVER_CONCURRENCY; set to 2. Proven under load, not just at idle: 26,645 requests over 300 s, memory 416-614 MiB against 768, trending down, zero SIGKILLs, OOMKilled false. Idle CPU 500% -> 1.64%. The first soak measured nothing - it was pointed at the scratch-guest subdomain, every request 404'd at traefik in 9 ms, and the counter reported 14,026 successes. Positive and negative controls are now asserted before any load is driven. Carried into R-462: `proven` has meant "the update applied and the data survived", not "the new version runs". Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
41 lines
5.6 KiB
Markdown
41 lines
5.6 KiB
Markdown
# STATUS — what works, what's broken, what's next
|
|
|
|
**Updated 2026-09-22 (overnight) — I installed and tested all 28 apps that no test had ever touched. Every app in our catalogue has now been tried at least once. Three things are quietly wrong, and one of them stops a working app.**
|
|
|
|
**Decisions I took on my own: none.**
|
|
|
|
**What I did.** Each of the twenty-eight got the same walk: install it at the version our catalogue offers today, put real data in through the app's own front door, back it up, update it if a newer version really exists, **restore it from that backup and read the data back again**, then delete it and check a minute later that nothing came back. That restore step is new — the update night skipped it. **Twenty-six of the twenty-eight installed. Six are proven end to end.** Fourteen had no newer version to move to tonight. One failed honestly. Two would not install, and one of those is meant not to.
|
|
|
|
**The thing I would fix first — an app with no health check gets shut down by a successful update.** Paperless-ngx is never health-checked at all: its containers are named differently from the app, so the machine looks for one, finds nothing, and moves on without a word. I always thought that was just a missing badge. It is not. **The update waits five minutes for a health check that can never arrive, then declares failure and shuts the working app down.** All three of its containers were healthy the whole time. The machine says so in its own words: *"not healthy within 5m0s (last: no probe container) — stopping and HOLDING the app"*. Every household running Paperless who presses Update loses their app and is sent to a restore they do not need.
|
|
|
|
**Two more, both about the machine losing track of an app rather than its health.**
|
|
- **Deleting an app while it is being restored leaves a ghost.** Both buttons say they worked. The app vanishes from every screen, and a container keeps restarting on the machine, still holding a public web address. **The machine already knows how to refuse this** — it refuses an *update* while a backup runs, and refuses a second *restore* while one is going, and it even names which app is blocking. Delete has no such guard.
|
|
- **An app can be running perfectly while the machine records it as not installed — and then it cannot be deleted.** I saw this three times. Two only happened when several jobs ran at once; one happened on its own, repeatably. In that state there is no button that works.
|
|
|
|
**In all three cases I needed a command line to clean up what the product could not. A household has none.**
|
|
|
|
**The best thing I saw.** Two apps keep their files outside the database, and their local copy does not hold those files. When I asked to restore them, the machine **refused** — and said, in plain Hungarian, that it will not put an old database on top of files it does not have, that the files stay where they are, and which button does work. That is exactly right.
|
|
|
|
**What I got wrong.** My own test script had three bugs that cost nine apps their walk. I found them, fixed them, and walked those nine again one at a time — and that second pass is what corrected my conclusions and produced one of the six proofs. The report names all three.
|
|
|
|
**Rows opened and closed.** Three new (the two above, plus the one that raised Paperless to urgent). Two closed. The list went from 321 to 323.
|
|
|
|
**What needs you.**
|
|
1. **The second promotion list** — six app versions this night proved safe enough to move on the real catalogue, and six named that must not move, each with the reason. It is in the report. Moving a version is your call, never mine. *If you do nothing:* nothing breaks; those apps drift further from upstream each month.
|
|
|
|
**Nothing on your own machine, the tester's machine, or the off-site box was touched. No product code was written. The real catalogue was never changed — I checked its version lines against the start of the night and not one differs.**
|
|
|
|
---
|
|
|
|
**Evening addition, 2026-09-22 — you heard the fans, and you were right.**
|
|
|
|
**RomM was cooking the HP box and it was my fault.** I moved it to a new version that morning. The update said it worked, and it did — for two hours. Then it ran out of memory and spent six hours killing and restarting its own workers, about seven times a minute, burning five processor cores. Nothing warned anyone. **You found it by ear. That was the only detector that worked.**
|
|
|
|
**Fixed, and proven under real load.** Giving it more memory was not enough — I measured that rather than assuming it. The real cause was that the app starts **four web workers**, which is a server setting, on a box serving one household. It now starts **two**. I then drove **26,645 real requests** at it for five minutes: memory stayed between 416 and 614 MB against its 768 MB ceiling and **trended down**, with **zero** worker deaths. At rest it now uses **1.6% of a core** instead of 500%.
|
|
|
|
**Two things I got wrong on the way, both worth knowing.** My first memory number (768 MB) was a guess and it did not fix anything. And my first load test was pointed at the wrong web address — every request bounced off the front door in 9 milliseconds while the counter reported 14,026 successes. I caught that only because the app looked *too* idle. Both are written up.
|
|
|
|
**The lesson that outlives RomM.** Every "proven" result this week measured an app for the **minutes of the test**. RomM passed everything and broke two hours later. **Proven has meant "the update worked and the data survived", not "the new version runs".** That gap is now on the record.
|
|
|
|
**Still worth an eye:** RomM sits at 610 MB of its 768 MB, about 79%. It works, but it is not roomy, and nothing watches it.
|