Files
felhom.eu/STATUS.md
T
admin 737694c603
gates / gates (push) Successful in 26s
CORRECTION: the app_oom alarm DID fire - R-635 was wrong, R-636 opened
The operator produced the mails. The controller HAS an OOM detector (main.go:821), it emits
app_oom (notifier.go:726), the hub allow-lists it (dispatcher.go:636) and delivered it to the
OPERATOR channel - two mails, 11:09 and 17:48 CEST, each naming the app and linking the dashboard.
CUSTOMER skipped, correctly.

I asserted an absence without opening the hub's Events or Notifications tab, reasoning instead from
a memory note that said the signal was UNPROVEN - not that it was missing. That is R-628's shape
again, from the same hand, four days later.

R-636: the real defect is the signal's SHAPE. notifier.go:715-724 keys on container|startedAt and
emits once per container lifetime, so 4530 worker kills over six hours produced exactly one
warning-level mail - indistinguishable from one transient kill. The magnitude was already collected
(App Telemetry: RomM 5023 errors, 632 warnings) but nothing turns it into a louder event.

Memory note lxc-docker-oom-signals-unreliable corrected with the positive reading.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-22 20:13:41 +02:00

6.5 KiB

STATUS — what works, what's broken, what's next

Updated 2026-09-22 (overnight) — I installed and tested all 28 apps that no test had ever touched. Every app in our catalogue has now been tried at least once. Three things are quietly wrong, and one of them stops a working app.

Decisions I took on my own: none.

What I did. Each of the twenty-eight got the same walk: install it at the version our catalogue offers today, put real data in through the app's own front door, back it up, update it if a newer version really exists, restore it from that backup and read the data back again, then delete it and check a minute later that nothing came back. That restore step is new — the update night skipped it. Twenty-six of the twenty-eight installed. Six are proven end to end. Fourteen had no newer version to move to tonight. One failed honestly. Two would not install, and one of those is meant not to.

The thing I would fix first — an app with no health check gets shut down by a successful update. Paperless-ngx is never health-checked at all: its containers are named differently from the app, so the machine looks for one, finds nothing, and moves on without a word. I always thought that was just a missing badge. It is not. The update waits five minutes for a health check that can never arrive, then declares failure and shuts the working app down. All three of its containers were healthy the whole time. The machine says so in its own words: "not healthy within 5m0s (last: no probe container) — stopping and HOLDING the app". Every household running Paperless who presses Update loses their app and is sent to a restore they do not need.

Two more, both about the machine losing track of an app rather than its health.

  • Deleting an app while it is being restored leaves a ghost. Both buttons say they worked. The app vanishes from every screen, and a container keeps restarting on the machine, still holding a public web address. The machine already knows how to refuse this — it refuses an update while a backup runs, and refuses a second restore while one is going, and it even names which app is blocking. Delete has no such guard.
  • An app can be running perfectly while the machine records it as not installed — and then it cannot be deleted. I saw this three times. Two only happened when several jobs ran at once; one happened on its own, repeatably. In that state there is no button that works.

In all three cases I needed a command line to clean up what the product could not. A household has none.

The best thing I saw. Two apps keep their files outside the database, and their local copy does not hold those files. When I asked to restore them, the machine refused — and said, in plain Hungarian, that it will not put an old database on top of files it does not have, that the files stay where they are, and which button does work. That is exactly right.

What I got wrong. My own test script had three bugs that cost nine apps their walk. I found them, fixed them, and walked those nine again one at a time — and that second pass is what corrected my conclusions and produced one of the six proofs. The report names all three.

Rows opened and closed. Three new (the two above, plus the one that raised Paperless to urgent). Two closed. The list went from 321 to 323.

What needs you.

  1. The second promotion list — six app versions this night proved safe enough to move on the real catalogue, and six named that must not move, each with the reason. It is in the report. Moving a version is your call, never mine. If you do nothing: nothing breaks; those apps drift further from upstream each month.

Nothing on your own machine, the tester's machine, or the off-site box was touched. No product code was written. The real catalogue was never changed — I checked its version lines against the start of the night and not one differs.


Evening addition, 2026-09-22 — you heard the fans, and you were right.

RomM was cooking the HP box and it was my fault. I moved it to a new version that morning. The update said it worked, and it did — for two hours. Then it ran out of memory and spent six hours killing and restarting its own workers, about seven times a minute, burning five processor cores.

CORRECTION — the machine DID warn you, and I was wrong to say it did not. You showed me the two e-mails: "Alkalmazás memóriája elfogyott: romm", at 11:09 and again at 17:48. They are in the hub's Events and Notifications tabs too. I wrote "nothing warned anyone" without opening either tab — I went by an old note saying this signal was unproven and turned that into "it did not happen". That is the same mistake I made with the Hetzner tickets four days ago.

What is actually wrong is smaller and real: the machine sends one warning per app start. Six hours of trouble and 4,530 worker deaths produced one e-mail — the same e-mail a single harmless hiccup would send. It never gets louder, and the app keeps showing as running. The hub did have the full picture on the App Telemetry page (RomM: 5,023 errors, 632 warnings, while every other app showed zero), but nothing turns that into a second, louder alert. That is why a correct warning still got missed, and it is now written down as its own item.

Fixed, and proven under real load. Giving it more memory was not enough — I measured that rather than assuming it. The real cause was that the app starts four web workers, which is a server setting, on a box serving one household. It now starts two. I then drove 26,645 real requests at it for five minutes: memory stayed between 416 and 614 MB against its 768 MB ceiling and trended down, with zero worker deaths. At rest it now uses 1.6% of a core instead of 500%.

Two things I got wrong on the way, both worth knowing. My first memory number (768 MB) was a guess and it did not fix anything. And my first load test was pointed at the wrong web address — every request bounced off the front door in 9 milliseconds while the counter reported 14,026 successes. I caught that only because the app looked too idle. Both are written up.

The lesson that outlives RomM. Every "proven" result this week measured an app for the minutes of the test. RomM passed everything and broke two hours later. Proven has meant "the update worked and the data survived", not "the new version runs". That gap is now on the record.

Still worth an eye: RomM sits at 610 MB of its 768 MB, about 79%. It works, but it is not roomy, and nothing watches it.