2026-09-24: 09 decisions 24-25, part 5 shipped; the whole-copy truth table; fourth suppression; rows; STATUS; evidence
gates / gates (push) Successful in 27s
gates / gates (push) Successful in 27s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -1,25 +1,30 @@
|
||||
# STATUS — what works, what's broken, what's next
|
||||
|
||||
**Updated 2026-09-24 (morning) — the night shift. Eleven apps moved to newer versions, each with a written test result. The chaos hour found two serious faults in the undo's way back. The new controller is NOT yet on the demo boxes: one question for you below.**
|
||||
**Updated 2026-09-24 (morning session). The fleet is on controller 0.268.0. The undo works after a restore. A stopped app's page tells the truth. A box that is several updates behind now climbs one tested step per press.**
|
||||
|
||||
**Decisions I took on my own (you may reverse them):**
|
||||
1. The memory test now counts only the app's own memory, not the kernel's file cache. The cache made two healthy apps read "100 % full" with no memory kills, and would have forced bigger memory limits for no reason. The old figure is still recorded beside the new one.
|
||||
2. The catalog's push check now asks the image registry, but only for images a push moves. Without it, a tested image could change before it reaches a box. Other pushes make no network call.
|
||||
**Decisions I took on my own:** none. Two small changes to your wording, both below.
|
||||
|
||||
**What I exercised.** Tests can no longer run real Docker commands on your own machine by accident. A restore now refuses a cut-off database copy before it touches anything (proven with the real restore button). Two page texts now tell the truth: where this box's full backup really goes, and how long the backup button stops the apps (about 8 minutes). Every catalog version move now needs a written test result, and a check refuses a move without one. I moved 11 apps (12 steps), each tested twice: on a test bench with 10 minutes of memory watching, and on the scratch machine through the real Update button. The HP demo box then updated three of its own apps through the real button: all three finished and answer.
|
||||
**What I did, and it worked.**
|
||||
- **The demo boxes got 0.267.0**, as you said yes to. Then, after every live test passed, they got **0.268.0**. Both boxes arrived healthy about 15 seconds after each change.
|
||||
- **The undo after a restore.** A restore used to rebuild an app's storage without the tag the undo looked for, so the undo copied nothing. Now the undo finds the storage by its name. A restore also puts the tag back. Tested on the scratch machine with an app restored by the OLD version, the way customer boxes are today: the failed update was undone, and data written before and after the restore came back.
|
||||
- **A stopped app's page (your option A).** When no copy on the box can bring an app back with its files, the page now says so, says support is informed, and shows no restore button. You get an urgent alarm. Tested by repeating last night's exact case, in Hungarian and English.
|
||||
- **Small fixes.** A stopped app no longer raises a second, false alarm. A removed app leaves no old files behind. The test bench now marks a memory test with no real load as "not proven", and it starts every run with an empty drive folder.
|
||||
- **The update ladder.** One press now does one tested step, never a jump. Tested with RomM on the scratch machine: first press moved only the app, second press moved only the database engine. The data came back after each press. The page says how many steps remain.
|
||||
|
||||
**What broke, and whether it healed.**
|
||||
- **After a restore, the automatic undo has nothing to put back.** A restore rebuilds an app's storage without the tag the undo uses to find it. The next failed update is then "undone" with no data copy. The app I tested happened to survive this. Not fixed tonight (one controller release per night). Serious.
|
||||
- **A stopped app can be pointed at a restore that refuses it.** After a failed update and a failed undo, the page names a backup to restore. For an app with files on a drive, the restore refuses that backup, and on a box with no off-site copy nothing else brings the app back. The stop rule fired here; all chaos rounds had already run. Serious.
|
||||
- Adventurelog's new version was not moved: our own template checks it the wrong way, and the new version needs an internet download at every start.
|
||||
- Reinstalling Nextcloud over kept files never finishes, and the box only says "unhealthy".
|
||||
- Smaller: a stopped app raises a second, extra alarm; the memory test once ran with no real load (fixed and re-run).
|
||||
**Your wording, changed in two small ways.**
|
||||
- English: I removed the word "please". The house rule for English copy forbids it.
|
||||
- Hungarian: I kept your text exactly. It uses the formal "Ön" form; the rest of the screens use the informal form. That mismatch is already on the list.
|
||||
|
||||
**Rows.** Ten opened, four closed. The list went from 330 to 336. Wishlist and uptime-kuma were fixed in the catalog; their rows stay open, smaller.
|
||||
**What I found (new, not fixed).**
|
||||
- **The second drive cannot bring back an app with files, even though it holds everything.** Its restore button refuses such apps, and the file restore brings back only files, not the database. So for Nextcloud, Immich, Paperless and Calibre, only the off-site copy counts. This needs your decision.
|
||||
- **A long crash loop never raises an alarm** if the app looks "up" for a moment between crashes. Seen on the scratch machine: 385 restarts, zero alarms.
|
||||
- The stopped-app sentence says "do not remove the app", but the Remove button is still there. This needs your word.
|
||||
- Smaller: a ladder step has no health check of its own yet; right after a restore, an update may be judged with the restored health check; one restore option ("database only") is described in the code but offered nowhere.
|
||||
|
||||
**What needs you — one question.** Should the demo boxes get controller 0.267.0 now?
|
||||
- **Yes, raise the floor (my recommendation).** The two serious faults are also in the version the boxes run today. 0.267.0 does not cause them, and it makes restores safer.
|
||||
- **No, wait.** Nothing changes on the boxes. The fixes above stay on the scratch machine only, until you say so.
|
||||
If you do nothing, the demo boxes stay on 0.266.0.
|
||||
**Rows.** Seven closed, six opened. The list went from 336 to 335.
|
||||
|
||||
**Nothing on Peti's machine, the off-site box or your own machine (beyond normal pushes) was touched. The test bench was deleted. The scratch machine is back on the real catalogue with its standing apps.**
|
||||
**What needs you — two questions.**
|
||||
1. **The second drive and apps with files.** (a) Build a "files plus database" restore for the second drive (my recommendation — then the second drive is a real way back, as customers would expect). (b) Rule that the second drive is files-only for these apps. If you do nothing, the page keeps naming only the off-site copy for these apps, which is true.
|
||||
2. **The Remove button on a stranded app.** (a) Hide it while support is informed (my recommendation — a removal destroys what support needs). (b) Keep it and soften the sentence. If you do nothing, both stay as they are.
|
||||
|
||||
**Not done.** No automatic updates yet (plan part 7). No customer box touched by hand. Nothing on Peti's machine, the off-site box or your own machine was touched, beyond the normal hub update. The scratch machine is back on the real catalog with its standing apps.
|
||||
|
||||
Reference in New Issue
Block a user