DRILL-night-2026-09-23.md: Parts A-E. 09 §3 decisions 21 (operator word), 22 and 23 (CC unattended, operator may reverse); §6.4 parts 4 and 6 (catalog half) shipped; §6.1a residuals R-658/R-659. Register 330 -> 336: R-651..R-660 opened (R-658 and R-659 P1), R-650/R-640/R-499/R-626 closed. Capability map, nightly rotation (opengist), STATUS (one question: the floor), CONTEXT, REPORT. The floor stays 0.266.0. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
3.3 KiB
STATUS — what works, what's broken, what's next
Updated 2026-09-24 (morning) — the night shift. Eleven apps moved to newer versions, each with a written test result. The chaos hour found two serious faults in the undo's way back. The new controller is NOT yet on the demo boxes: one question for you below.
Decisions I took on my own (you may reverse them):
- The memory test now counts only the app's own memory, not the kernel's file cache. The cache made two healthy apps read "100 % full" with no memory kills, and would have forced bigger memory limits for no reason. The old figure is still recorded beside the new one.
- The catalog's push check now asks the image registry, but only for images a push moves. Without it, a tested image could change before it reaches a box. Other pushes make no network call.
What I exercised. Tests can no longer run real Docker commands on your own machine by accident. A restore now refuses a cut-off database copy before it touches anything (proven with the real restore button). Two page texts now tell the truth: where this box's full backup really goes, and how long the backup button stops the apps (about 8 minutes). Every catalog version move now needs a written test result, and a check refuses a move without one. I moved 11 apps (12 steps), each tested twice: on a test bench with 10 minutes of memory watching, and on the scratch machine through the real Update button. The HP demo box then updated three of its own apps through the real button: all three finished and answer.
What broke, and whether it healed.
- After a restore, the automatic undo has nothing to put back. A restore rebuilds an app's storage without the tag the undo uses to find it. The next failed update is then "undone" with no data copy. The app I tested happened to survive this. Not fixed tonight (one controller release per night). Serious.
- A stopped app can be pointed at a restore that refuses it. After a failed update and a failed undo, the page names a backup to restore. For an app with files on a drive, the restore refuses that backup, and on a box with no off-site copy nothing else brings the app back. The stop rule fired here; all chaos rounds had already run. Serious.
- Adventurelog's new version was not moved: our own template checks it the wrong way, and the new version needs an internet download at every start.
- Reinstalling Nextcloud over kept files never finishes, and the box only says "unhealthy".
- Smaller: a stopped app raises a second, extra alarm; the memory test once ran with no real load (fixed and re-run).
Rows. Ten opened, four closed. The list went from 330 to 336. Wishlist and uptime-kuma were fixed in the catalog; their rows stay open, smaller.
What needs you — one question. Should the demo boxes get controller 0.267.0 now?
- Yes, raise the floor (my recommendation). The two serious faults are also in the version the boxes run today. 0.267.0 does not cause them, and it makes restores safer.
- No, wait. Nothing changes on the boxes. The fixes above stay on the scratch machine only, until you say so. If you do nothing, the demo boxes stay on 0.266.0.
Nothing on Peti's machine, the off-site box or your own machine (beyond normal pushes) was touched. The test bench was deleted. The scratch machine is back on the real catalogue with its standing apps.