Files
felhom.eu/STATUS.md
T
admin a975cfde5b
gates / gates (push) Successful in 28s
probe fix, the gate, and the promotion train (R-618 closed, R-630..632 opened)
Part 1: tandoor/zipline/wger probes corrected in the catalog and red-proofed live on 9202 in both
directions - "Nem egeszseges" with the front door serving 200, then "Fut" after the real sync with
no redeploy. tandoor's failed edge re-walked: done at +41.1s where it was failed at +361.9s.

Part 2: fifteen proven versions on the live catalog, one commit per app; the guarded Update pressed
on four apps on demo-hp, all four done.

Opened: R-630 (paperless-ngx's probe has never run on any box - a silent absence, worse than the
wrong probe that was found in one night), R-631 (five templates no static rule can judge),
R-632 (28 of 53 templates never deployed by any drill). Closed: R-618.

Register 318 -> 321. No product code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-22 11:10:58 +02:00

4.6 KiB

STATUS — what works, what's broken, what's next

Updated 2026-09-22 (morning) — I fixed the three apps that shut themselves down after a good update, proved the fix on a real machine, built a check so it cannot happen again, and moved fifteen app versions onto the real catalog. One thing you asked for could not be done, and one thing I nearly got wrong was caught by running the check instead of trusting my reasoning.

Decisions I took on my own: one. I moved Nextcloud's database engine up a version. I had written it down as "dropped — the rules forbid it", then ran the rule instead of believing my memory of it, and the rule allows it by name: that permission was granted on 2026-09-21, the template already carries the setting that converts the data, and the move was proven end to end in under four minutes. You forbade PostgreSQL engine moves; this one is MariaDB. You can reverse it by reverting one commit.

The three broken apps are fixed. Tandoor, Zipline and Wger each had one wrong number or address, so the machine knocked on a door the app does not answer. I fixed all three, then proved it on the scratch machine in both directions: before the fix all three showed „Nem egészséges" on their own page while the app itself was serving customers normally; after the fix, with no restart and no reinstall, all three showed „Fut".

And the update that failed now works. I ran Tandoor's exact same update again — same app, same versions, same button, nothing changed but that one number. Yesterday it ran for six minutes and shut the app down. Today it finished in 41 seconds and the data was still there. That is the whole finding in one line.

There is now a check that catches this before it ships. The answer was always sitting in the same file, a few lines further down — each app already tells Docker where to knock. The check compares the two. It runs on every push, it refused all three apps before the fix, it passes now, and it is guarded by ten fake-out tests so it cannot quietly stop working.

Fifteen versions moved to the real catalog, one at a time, every check run before each one. Nothing was forced and nothing was dropped. Then I pressed the real Update button on four of them on the demo machine: all four finished cleanly — BookStack, Docmost, PrivateBin and RomM are running the new versions.

What I could not do. You asked me to press that button on both demo machines. The second machine has only one app installed and none of the four. I did not install them — installing apps on a demo machine is a change, not a test. Both machines did receive the new catalog, and I checked that.

What the check found that nobody was looking for — and it is quieter than the bug it was built for.

  • Paperless-ngx has never been health-checked at all. Not "checked wrongly" — never checked. Its containers are named differently from the app, so the machine looks for one and finds nothing, and moves on without a word. A wrong check is loud and we caught it in one night. A missing check looks exactly like a healthy app.
  • Five more apps cannot be checked this way. One of them, Home Assistant, is correct today only by luck: tighten its settings in the obvious way and it breaks the same way Tandoor did.
  • 28 of our 53 apps have never been installed by any test. The overnight run went from 3 apps to 21, which is a lot — but 21 is not 53. For those 28, we do not know whether updating works. That list is now the nightly queue.

Rows opened and closed. Three new, one closed. The list went from 318 to 321.

What needs you.

  1. Rotate the Gitea admin token — still open from yesterday. The machine stores it in plain text in its copy of the catalog. If you do nothing: the token keeps working and anyone with yesterday's session transcript has it.
  2. Clear 47 alarm e-mails from yesterday's drill, in one search: subject:"gates FAILED in admin/app-catalog-drill". The cause is fixed. If you do nothing: your alarm inbox stays noisy, and that is the inbox that must never be skimmed.
  3. Decide whether the Nextcloud engine move stays. I explained my reasoning above. If you do nothing: it stays, and Nextcloud households will be offered a database upgrade that was proven once on a scratch machine.
  4. The 28 untested apps. If you do nothing: the nightly rotation works through them at a few per night, and the catalog's update promise rests on nothing for those apps until it gets there.

Nothing on your own machine, the tester's machine, or the off-site box was touched. No product code was written.