All 28 walked on scratch guest 9202 against the private drill catalog. 26 deployed, 6 proven, 5 inconclusive, 14 with no upstream edge, 1 failed honestly (outline 1.9.1->1.10.1, HELD with the right sentence), 2 undeployable - one (plant-it) by design, refused by the lifecycle gate, proven live for the first time. Each app also got the half the update night skipped: a restore from its own copy with the seed read back again - 21 restored, 2 correctly REFUSED per 07 6.2. R-630 RAISED TO P1 by measurement: a stack with NO probe container does not skip verifying - it waits out the full health timeout and HOLDS, stopping an app whose three containers read healthy. The controller's own words: "not healthy within 5m0s (last: no probe container)". R-633 opened: a remove sent during a restore reports success and leaves a container restarting with a live public route. The product already refuses that clash for update and for restore, naming the blocker; remove has no such guard. R-634 opened: an app can be running, healthy and serving while recorded as deployed=false, and is then unremovable. Reproducible alone on sparkyfitness; concurrency-linked on two others. R-631 and R-632 CLOSED. Register 321 -> 323. Seven interventions, six of them my own harness - named, with what each cost. No product code. The live catalog's image: lines are byte-identical to the start of the night. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
3.9 KiB
STATUS — what works, what's broken, what's next
Updated 2026-09-22 (overnight) — I installed and tested all 28 apps that no test had ever touched. Every app in our catalogue has now been tried at least once. Three things are quietly wrong, and one of them stops a working app.
Decisions I took on my own: none.
What I did. Each of the twenty-eight got the same walk: install it at the version our catalogue offers today, put real data in through the app's own front door, back it up, update it if a newer version really exists, restore it from that backup and read the data back again, then delete it and check a minute later that nothing came back. That restore step is new — the update night skipped it. Twenty-six of the twenty-eight installed. Six are proven end to end. Fourteen had no newer version to move to tonight. One failed honestly. Two would not install, and one of those is meant not to.
The thing I would fix first — an app with no health check gets shut down by a successful update. Paperless-ngx is never health-checked at all: its containers are named differently from the app, so the machine looks for one, finds nothing, and moves on without a word. I always thought that was just a missing badge. It is not. The update waits five minutes for a health check that can never arrive, then declares failure and shuts the working app down. All three of its containers were healthy the whole time. The machine says so in its own words: "not healthy within 5m0s (last: no probe container) — stopping and HOLDING the app". Every household running Paperless who presses Update loses their app and is sent to a restore they do not need.
Two more, both about the machine losing track of an app rather than its health.
- Deleting an app while it is being restored leaves a ghost. Both buttons say they worked. The app vanishes from every screen, and a container keeps restarting on the machine, still holding a public web address. The machine already knows how to refuse this — it refuses an update while a backup runs, and refuses a second restore while one is going, and it even names which app is blocking. Delete has no such guard.
- An app can be running perfectly while the machine records it as not installed — and then it cannot be deleted. I saw this three times. Two only happened when several jobs ran at once; one happened on its own, repeatably. In that state there is no button that works.
In all three cases I needed a command line to clean up what the product could not. A household has none.
The best thing I saw. Two apps keep their files outside the database, and their local copy does not hold those files. When I asked to restore them, the machine refused — and said, in plain Hungarian, that it will not put an old database on top of files it does not have, that the files stay where they are, and which button does work. That is exactly right.
What I got wrong. My own test script had three bugs that cost nine apps their walk. I found them, fixed them, and walked those nine again one at a time — and that second pass is what corrected my conclusions and produced one of the six proofs. The report names all three.
Rows opened and closed. Three new (the two above, plus the one that raised Paperless to urgent). Two closed. The list went from 321 to 323.
What needs you.
- The second promotion list — six app versions this night proved safe enough to move on the real catalogue, and six named that must not move, each with the reason. It is in the report. Moving a version is your call, never mine. If you do nothing: nothing breaks; those apps drift further from upstream each month.
Nothing on your own machine, the tester's machine, or the off-site box was touched. No product code was written. The real catalogue was never changed — I checked its version lines against the start of the night and not one differs.