THE TWENTY-EIGHT: every app no drill had touched, walked in one night
gates / gates (push) Successful in 27s
gates / gates (push) Successful in 27s
All 28 walked on scratch guest 9202 against the private drill catalog. 26 deployed, 6 proven, 5 inconclusive, 14 with no upstream edge, 1 failed honestly (outline 1.9.1->1.10.1, HELD with the right sentence), 2 undeployable - one (plant-it) by design, refused by the lifecycle gate, proven live for the first time. Each app also got the half the update night skipped: a restore from its own copy with the seed read back again - 21 restored, 2 correctly REFUSED per 07 6.2. R-630 RAISED TO P1 by measurement: a stack with NO probe container does not skip verifying - it waits out the full health timeout and HOLDS, stopping an app whose three containers read healthy. The controller's own words: "not healthy within 5m0s (last: no probe container)". R-633 opened: a remove sent during a restore reports success and leaves a container restarting with a live public route. The product already refuses that clash for update and for restore, naming the blocker; remove has no such guard. R-634 opened: an app can be running, healthy and serving while recorded as deployed=false, and is then unremovable. Reproducible alone on sparkyfitness; concurrency-linked on two others. R-631 and R-632 CLOSED. Register 321 -> 323. Seven interventions, six of them my own harness - named, with what each cost. No product code. The live catalog's image: lines are byte-identical to the start of the night. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -1,30 +1,26 @@
|
||||
# STATUS — what works, what's broken, what's next
|
||||
|
||||
**Updated 2026-09-22 (morning) — I fixed the three apps that shut themselves down after a good update, proved the fix on a real machine, built a check so it cannot happen again, and moved fifteen app versions onto the real catalog. One thing you asked for could not be done, and one thing I nearly got wrong was caught by running the check instead of trusting my reasoning.**
|
||||
**Updated 2026-09-22 (overnight) — I installed and tested all 28 apps that no test had ever touched. Every app in our catalogue has now been tried at least once. Three things are quietly wrong, and one of them stops a working app.**
|
||||
|
||||
**Decisions I took on my own: one.** I moved Nextcloud's database engine up a version. I had written it down as "dropped — the rules forbid it", then ran the rule instead of believing my memory of it, and the rule **allows** it by name: that permission was granted on 2026-09-21, the template already carries the setting that converts the data, and the move was proven end to end in under four minutes. You forbade PostgreSQL engine moves; this one is MariaDB. **You can reverse it by reverting one commit.**
|
||||
**Decisions I took on my own: none.**
|
||||
|
||||
**The three broken apps are fixed.** Tandoor, Zipline and Wger each had one wrong number or address, so the machine knocked on a door the app does not answer. I fixed all three, then proved it on the scratch machine **in both directions**: before the fix all three showed **„Nem egészséges"** on their own page while the app itself was serving customers normally; after the fix, with no restart and no reinstall, all three showed **„Fut"**.
|
||||
**What I did.** Each of the twenty-eight got the same walk: install it at the version our catalogue offers today, put real data in through the app's own front door, back it up, update it if a newer version really exists, **restore it from that backup and read the data back again**, then delete it and check a minute later that nothing came back. That restore step is new — the update night skipped it. **Twenty-six of the twenty-eight installed. Six are proven end to end.** Fourteen had no newer version to move to tonight. One failed honestly. Two would not install, and one of those is meant not to.
|
||||
|
||||
**And the update that failed now works.** I ran Tandoor's exact same update again — same app, same versions, same button, nothing changed but that one number. Yesterday it ran for **six minutes and shut the app down**. Today it finished in **41 seconds** and the data was still there. That is the whole finding in one line.
|
||||
**The thing I would fix first — an app with no health check gets shut down by a successful update.** Paperless-ngx is never health-checked at all: its containers are named differently from the app, so the machine looks for one, finds nothing, and moves on without a word. I always thought that was just a missing badge. It is not. **The update waits five minutes for a health check that can never arrive, then declares failure and shuts the working app down.** All three of its containers were healthy the whole time. The machine says so in its own words: *"not healthy within 5m0s (last: no probe container) — stopping and HOLDING the app"*. Every household running Paperless who presses Update loses their app and is sent to a restore they do not need.
|
||||
|
||||
**There is now a check that catches this before it ships.** The answer was always sitting in the same file, a few lines further down — each app already tells Docker where to knock. The check compares the two. It runs on every push, it refused all three apps before the fix, it passes now, and it is guarded by ten fake-out tests so it cannot quietly stop working.
|
||||
**Two more, both about the machine losing track of an app rather than its health.**
|
||||
- **Deleting an app while it is being restored leaves a ghost.** Both buttons say they worked. The app vanishes from every screen, and a container keeps restarting on the machine, still holding a public web address. **The machine already knows how to refuse this** — it refuses an *update* while a backup runs, and refuses a second *restore* while one is going, and it even names which app is blocking. Delete has no such guard.
|
||||
- **An app can be running perfectly while the machine records it as not installed — and then it cannot be deleted.** I saw this three times. Two only happened when several jobs ran at once; one happened on its own, repeatably. In that state there is no button that works.
|
||||
|
||||
**Fifteen versions moved to the real catalog**, one at a time, every check run before each one. Nothing was forced and nothing was dropped. Then I pressed the real Update button on four of them on the demo machine: **all four finished cleanly** — BookStack, Docmost, PrivateBin and RomM are running the new versions.
|
||||
**In all three cases I needed a command line to clean up what the product could not. A household has none.**
|
||||
|
||||
**What I could not do.** You asked me to press that button on **both** demo machines. The second machine has only one app installed and none of the four. I did not install them — installing apps on a demo machine is a change, not a test. Both machines did receive the new catalog, and I checked that.
|
||||
**The best thing I saw.** Two apps keep their files outside the database, and their local copy does not hold those files. When I asked to restore them, the machine **refused** — and said, in plain Hungarian, that it will not put an old database on top of files it does not have, that the files stay where they are, and which button does work. That is exactly right.
|
||||
|
||||
**What the check found that nobody was looking for — and it is quieter than the bug it was built for.**
|
||||
- **Paperless-ngx has never been health-checked at all.** Not "checked wrongly" — never checked. Its containers are named differently from the app, so the machine looks for one and finds nothing, and moves on without a word. A wrong check is loud and we caught it in one night. **A missing check looks exactly like a healthy app.**
|
||||
- **Five more apps cannot be checked this way.** One of them, Home Assistant, is correct today only by luck: tighten its settings in the obvious way and it breaks the same way Tandoor did.
|
||||
- **28 of our 53 apps have never been installed by any test.** The overnight run went from 3 apps to 21, which is a lot — but 21 is not 53. **For those 28, we do not know whether updating works.** That list is now the nightly queue.
|
||||
**What I got wrong.** My own test script had three bugs that cost nine apps their walk. I found them, fixed them, and walked those nine again one at a time — and that second pass is what corrected my conclusions and produced one of the six proofs. The report names all three.
|
||||
|
||||
**Rows opened and closed.** Three new, one closed. The list went from 318 to 321.
|
||||
**Rows opened and closed.** Three new (the two above, plus the one that raised Paperless to urgent). Two closed. The list went from 321 to 323.
|
||||
|
||||
**What needs you.**
|
||||
1. **Rotate the Gitea `admin` token** — still open from yesterday. The machine stores it in plain text in its copy of the catalog. *If you do nothing:* the token keeps working and anyone with yesterday's session transcript has it.
|
||||
2. **Clear 47 alarm e-mails** from yesterday's drill, in one search: `subject:"gates FAILED in admin/app-catalog-drill"`. The cause is fixed. *If you do nothing:* your alarm inbox stays noisy, and that is the inbox that must never be skimmed.
|
||||
3. **Decide whether the Nextcloud engine move stays.** I explained my reasoning above. *If you do nothing:* it stays, and Nextcloud households will be offered a database upgrade that was proven once on a scratch machine.
|
||||
4. **The 28 untested apps.** *If you do nothing:* the nightly rotation works through them at a few per night, and the catalog's update promise rests on nothing for those apps until it gets there.
|
||||
1. **The second promotion list** — six app versions this night proved safe enough to move on the real catalogue, and six named that must not move, each with the reason. It is in the report. Moving a version is your call, never mine. *If you do nothing:* nothing breaks; those apps drift further from upstream each month.
|
||||
|
||||
**Nothing on your own machine, the tester's machine, or the off-site box was touched. No product code was written.**
|
||||
**Nothing on your own machine, the tester's machine, or the off-site box was touched. No product code was written. The real catalogue was never changed — I checked its version lines against the start of the night and not one differs.**
|
||||
|
||||
Reference in New Issue
Block a user