v0.262.0 + v0.262.1 live on 9202: four scenarios proven, R-630/R-633/R-621/R-614 closed
gates / gates (push) Successful in 30s

A (R-630, the P1): paperless-ngx, same app same button - failed at +313.0s with the app stopped
under v0.261.0, done at +53.4s now, with no "no probe container" warning because the explicit
healthcheck.container resolved the target.

B (R-633): the busy guard fires - RemoveStack REFUSED (busy): a backup or restore is running - and
the live proof caught it answering HTTP 500, because router.go maps remove errors by grepping the
error TEXT. v0.262.1 makes it a typed error and a 409; re-proven live.

C (R-634 half): a half-state with app.yaml on disk and deployed=false answered 200, leftovers NONE.
Under v0.261.0 the same call said "not deployed".

F (R-614): phase done before the remove, no phase at all after redeploying the same name.

Also: the first B run proved NOTHING and nearly went down as a pass - the refusal came from the
pre-existing "still running" check, not the new guard. Recorded.

09 6.1 and 8.8, the capability map, and STATUS updated. R-625 and R-634's mechanism are named as
owed, not half-done. Register 325.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-22 21:53:53 +02:00
parent 1de6aaf904
commit 22439b0e43
12 changed files with 544 additions and 39 deletions
+15 -32
View File
@@ -1,44 +1,27 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-09-22 (overnight) — I installed and tested all 28 apps that no test had ever touched. Every app in our catalogue has now been tried at least once. Three things are quietly wrong, and one of them stops a working app.**
**Updated 2026-09-22 (late) — I fixed the six faults the two drill nights found in the update, delete and hold machinery, and shipped the six app versions you approved. One thing needs your word: whether the fleet moves to the new controller.**
**Decisions I took on my own: none.**
**What I did.** Each of the twenty-eight got the same walk: install it at the version our catalogue offers today, put real data in through the app's own front door, back it up, update it if a newer version really exists, **restore it from that backup and read the data back again**, then delete it and check a minute later that nothing came back. That restore step is new — the update night skipped it. **Twenty-six of the twenty-eight installed. Six are proven end to end.** Fourteen had no newer version to move to tonight. One failed honestly. Two would not install, and one of those is meant not to.
**The one that mattered most is fixed and proven.** An app with no health check used to be **shut down by a successful update** — the machine waited five minutes for a check that could never arrive, then stopped a working app. Paperless-ngx, same app, same button: **before, it failed after 5 minutes and the app went dark. Now it finishes in 53 seconds and keeps running.**
**The thing I would fix first — an app with no health check gets shut down by a successful update.** Paperless-ngx is never health-checked at all: its containers are named differently from the app, so the machine looks for one, finds nothing, and moves on without a word. I always thought that was just a missing badge. It is not. **The update waits five minutes for a health check that can never arrive, then declares failure and shuts the working app down.** All three of its containers were healthy the whole time. The machine says so in its own words: *"not healthy within 5m0s (last: no probe container) — stopping and HOLDING the app"*. Every household running Paperless who presses Update loses their app and is sent to a restore they do not need.
**Five more, all proven on the test machine.**
- **Deleting an app while it is being backed up or restored is now refused**, with a plain sentence telling you to wait — instead of quietly tearing it down and leaving a ghost behind.
- **A delete now checks its own work.** The machine watches for 25 seconds afterwards and removes anything that comes back, and says whether it verified.
- **An app the machine has lost track of can now be deleted.** Before, if its record went wrong, no button worked and only a command line could clear it.
- **A failed update now keeps the app's own log** before shutting it down. Twice we lost the only evidence of why.
- **Deleting an app clears its old update status**, so a fresh install of the same app no longer shows a stale "Updated".
**Two more, both about the machine losing track of an app rather than its health.**
- **Deleting an app while it is being restored leaves a ghost.** Both buttons say they worked. The app vanishes from every screen, and a container keeps restarting on the machine, still holding a public web address. **The machine already knows how to refuse this** — it refuses an *update* while a backup runs, and refuses a second *restore* while one is going, and it even names which app is blocking. Delete has no such guard.
- **An app can be running perfectly while the machine records it as not installed — and then it cannot be deleted.** I saw this three times. Two only happened when several jobs ran at once; one happened on its own, repeatably. In that state there is no button that works.
**The six versions you approved are live on the catalogue** — Emby, Ghost, Immich, Radarr, Sonarr, Termix. **None of them is installed on either demo machine**, so nothing updated; they simply show as available.
**In all three cases I needed a command line to clean up what the product could not. A household has none.**
**What I did not do, and it is on purpose.** Two items from the plan are untouched and named rather than half-finished: finding out *why* an app's record goes wrong in the first place (I fixed the consequence, not the cause), and making a held app stop offering an Update button it will refuse.
**The best thing I saw.** Two apps keep their files outside the database, and their local copy does not hold those files. When I asked to restore them, the machine **refused** — and said, in plain Hungarian, that it will not put an old database on top of files it does not have, that the files stay where they are, and which button does work. That is exactly right.
**What went wrong on my side.** I lost **44 minutes** to my own progress-watchers: they waited for a build that had already succeeded, because each was watching for a name its own command contained. The same bug cost me a pile of stuck watchers earlier in the day. It is now written down as a rule so it does not happen a third time. I also nearly recorded one test as passing when it had proved nothing — the refusal I saw came from an older rule, not the new one. I caught it and re-ran it properly.
**What I got wrong.** My own test script had three bugs that cost nine apps their walk. I found them, fixed them, and walked those nine again one at a time — and that second pass is what corrected my conclusions and produced one of the six proofs. The report names all three.
**Rows opened and closed.** Four closed, one narrowed to what is still unknown. The list stands at 325.
**Rows opened and closed.** Three new (the two above, plus the one that raised Paperless to urgent). Two closed. The list went from 321 to 323.
**What needs you — one question.**
1. **Shall the fleet move to controller 0.262.1?** Right now only the test machine has it. **My recommendation is yes:** every change here only refuses, waits, records, or removes what someone already asked to remove — none of them makes the machine do more on its own. *If you do nothing:* both demo machines stay on 0.261.0 and keep all six faults, including the one that shuts down a working app. Peti's machine is parked and would take it only if it ever comes back online.
**What needs you.**
1. **The second promotion list** — six app versions this night proved safe enough to move on the real catalogue, and six named that must not move, each with the reason. It is in the report. Moving a version is your call, never mine. *If you do nothing:* nothing breaks; those apps drift further from upstream each month.
**Nothing on your own machine, the tester's machine, or the off-site box was touched. No product code was written. The real catalogue was never changed — I checked its version lines against the start of the night and not one differs.**
---
**Evening addition, 2026-09-22 — you heard the fans, and you were right.**
**RomM was cooking the HP box and it was my fault.** I moved it to a new version that morning. The update said it worked, and it did — for two hours. Then it ran out of memory and spent six hours killing and restarting its own workers, about seven times a minute, burning five processor cores.
**CORRECTION — the machine DID warn you, and I was wrong to say it did not.** You showed me the two e-mails: *"Alkalmazás memóriája elfogyott: romm"*, at 11:09 and again at 17:48. They are in the hub's Events and Notifications tabs too. **I wrote "nothing warned anyone" without opening either tab** — I went by an old note saying this signal was unproven and turned that into "it did not happen". That is the same mistake I made with the Hetzner tickets four days ago.
**What is actually wrong is smaller and real:** the machine sends **one** warning per app start. Six hours of trouble and 4,530 worker deaths produced **one e-mail** — the same e-mail a single harmless hiccup would send. It never gets louder, and the app keeps showing as running. The hub *did* have the full picture on the App Telemetry page (RomM: 5,023 errors, 632 warnings, while every other app showed zero), but nothing turns that into a second, louder alert. **That is why a correct warning still got missed, and it is now written down as its own item.**
**Fixed, and proven under real load.** Giving it more memory was not enough — I measured that rather than assuming it. The real cause was that the app starts **four web workers**, which is a server setting, on a box serving one household. It now starts **two**. I then drove **26,645 real requests** at it for five minutes: memory stayed between 416 and 614 MB against its 768 MB ceiling and **trended down**, with **zero** worker deaths. At rest it now uses **1.6% of a core** instead of 500%.
**Two things I got wrong on the way, both worth knowing.** My first memory number (768 MB) was a guess and it did not fix anything. And my first load test was pointed at the wrong web address — every request bounced off the front door in 9 milliseconds while the counter reported 14,026 successes. I caught that only because the app looked *too* idle. Both are written up.
**The lesson that outlives RomM.** Every "proven" result this week measured an app for the **minutes of the test**. RomM passed everything and broke two hours later. **Proven has meant "the update worked and the data survived", not "the new version runs".** That gap is now on the record.
**Still worth an eye:** RomM sits at 610 MB of its 768 MB, about 79%. It works, but it is not roomy, and nothing watches it.
**Nothing on your own machine or the off-site box was touched. The demo machines were not touched — they only see the six new version badges.**