night shift 2026-09-23: the record, the register, the morning note
gates / gates (push) Successful in 27s

DRILL-night-2026-09-23.md: Parts A-E. 09 §3 decisions 21 (operator word),
22 and 23 (CC unattended, operator may reverse); §6.4 parts 4 and 6
(catalog half) shipped; §6.1a residuals R-658/R-659. Register 330 -> 336:
R-651..R-660 opened (R-658 and R-659 P1), R-650/R-640/R-499/R-626 closed.
Capability map, nightly rotation (opengist), STATUS (one question: the
floor), CONTEXT, REPORT. The floor stays 0.266.0.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-24 00:24:04 +02:00
parent 77335625fe
commit 3e58c184f6
140 changed files with 15399 additions and 316 deletions
+17 -17
View File
@@ -1,25 +1,25 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-09-23 (night) — clean-up evening, plus your install ruling. The "runs but not installed" fault has its cause found and fixed. The fleet has the new version.**
**Updated 2026-09-24 (morning) — the night shift. Eleven apps moved to newer versions, each with a written test result. The chaos hour found two serious faults in the undo's way back. The new controller is NOT yet on the demo boxes: one question for you below.**
**Decisions I took on my own: none.** Your install ruling is built (below).
**Decisions I took on my own (you may reverse them):**
1. The memory test now counts only the app's own memory, not the kernel's file cache. The cache made two healthy apps read "100 % full" with no memory kills, and would have forced bigger memory limits for no reason. The old figure is still recorded beside the new one.
2. The catalog's push check now asks the image registry, but only for images a push moves. Without it, a tested image could change before it reaches a box. Other pushes make no network call.
**The top fault: an app could run while the box said "not installed".** Found and fixed. The cause: the "back up now" button backs up every app. It also caught an app that was still installing. It stopped that app in the middle of the install and started it again. The install then failed and wrote "not installed", while the app's containers sometimes kept running. I made this happen on purpose on the scratch machine, twice. With the fix, the backup leaves an installing app alone. I proved that on the same machine: the backup stopped three other apps and did not touch the one installing, and the install finished normally. Sparkyfitness did not fail on its own this time. It installed in 47 seconds.
**What I exercised.** Tests can no longer run real Docker commands on your own machine by accident. A restore now refuses a cut-off database copy before it touches anything (proven with the real restore button). Two page texts now tell the truth: where this box's full backup really goes, and how long the backup button stops the apps (about 8 minutes). Every catalog version move now needs a written test result, and a check refuses a move without one. I moved 11 apps (12 steps), each tested twice: on a test bench with 10 minutes of memory watching, and on the scratch machine through the real Update button. The HP demo box then updated three of its own apps through the real button: all three finished and answer.
**A stopped app now says so.** When an update failed and the app is stopped, the badge says "Stopped — restore needed", in Hungarian or English. There is no Update button. Proven on the scratch machine in both languages.
**What broke, and whether it healed.**
- **After a restore, the automatic undo has nothing to put back.** A restore rebuilds an app's storage without the tag the undo uses to find it. The next failed update is then "undone" with no data copy. The app I tested happened to survive this. Not fixed tonight (one controller release per night). Serious.
- **A stopped app can be pointed at a restore that refuses it.** After a failed update and a failed undo, the page names a backup to restore. For an app with files on a drive, the restore refuses that backup, and on a box with no off-site copy nothing else brings the app back. The stop rule fired here; all chaos rounds had already run. Serious.
- Adventurelog's new version was not moved: our own template checks it the wrong way, and the new version needs an internet download at every start.
- Reinstalling Nextcloud over kept files never finishes, and the box only says "unhealthy".
- Smaller: a stopped app raises a second, extra alarm; the memory test once ran with no real load (fixed and re-run).
**A long memory problem is now louder.** The box now counts real memory kills. If one app is killed 20 times in 30 minutes, you get one louder alarm ("error"), once. Households do not get it. Proven with RomM on the scratch machine: the alarm came at 21 kills and stayed one alarm at 49.
**Rows.** Ten opened, four closed. The list went from 330 to 336. Wishlist and uptime-kuma were fixed in the catalog; their rows stay open, smaller.
**The three small mail and log leftovers are fixed.** A reader who picks the other language now reads a stopped app's message in their own language.
**What needs you — one question.** Should the demo boxes get controller 0.267.0 now?
- **Yes, raise the floor (my recommendation).** The two serious faults are also in the version the boxes run today. 0.267.0 does not cause them, and it makes restores safer.
- **No, wait.** Nothing changes on the boxes. The fixes above stay on the scratch machine only, until you say so.
If you do nothing, the demo boxes stay on 0.266.0.
**The test method is fixed.** My tests no longer press "back up now". An app's update makes its own backup instead.
**Rows.** Five closed, two new. The list went from 334 to 331. The floor is 0.265.0, and both demo machines updated themselves within 20 seconds.
**What was not clean.** While I wrote one test, it created an empty storage volume on your own machine. I saw it at once, checked it was new and unused, and removed it. Nothing else was touched. I filed a row so tests cannot do this again.
**Your ruling is built.** A failed install now removes the containers it started, and keeps the app's data. Proven on the scratch machine: an install that failed left no containers behind and kept its data. The floor is 0.266.0, and both demo machines updated themselves within 20 seconds. The list is now 330 rows.
**What needs you: nothing.**
**Nothing on Peti's machine or the off-site box was touched. On your own machine, only the hub was updated, plus the test mistake above. The demo-hp machine was not touched by hand. The scratch machine is back to its three standing apps and the real catalogue.**
**Nothing on Peti's machine, the off-site box or your own machine (beyond normal pushes) was touched. The test bench was deleted. The scratch machine is back on the real catalogue with its standing apps.**