Files
felhom.eu/STATUS.md
T
admin 22439b0e43
gates / gates (push) Successful in 30s
v0.262.0 + v0.262.1 live on 9202: four scenarios proven, R-630/R-633/R-621/R-614 closed
A (R-630, the P1): paperless-ngx, same app same button - failed at +313.0s with the app stopped
under v0.261.0, done at +53.4s now, with no "no probe container" warning because the explicit
healthcheck.container resolved the target.

B (R-633): the busy guard fires - RemoveStack REFUSED (busy): a backup or restore is running - and
the live proof caught it answering HTTP 500, because router.go maps remove errors by grepping the
error TEXT. v0.262.1 makes it a typed error and a 409; re-proven live.

C (R-634 half): a half-state with app.yaml on disk and deployed=false answered 200, leftovers NONE.
Under v0.261.0 the same call said "not deployed".

F (R-614): phase done before the remove, no phase at all after redeploying the same name.

Also: the first B run proved NOTHING and nearly went down as a pass - the refusal came from the
pre-existing "still running" check, not the new guard. Recorded.

09 6.1 and 8.8, the capability map, and STATUS updated. R-625 and R-634's mechanism are named as
owed, not half-done. Register 325.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-22 21:53:53 +02:00

28 lines
3.2 KiB
Markdown

# STATUS — what works, what's broken, what's next
**Updated 2026-09-22 (late) — I fixed the six faults the two drill nights found in the update, delete and hold machinery, and shipped the six app versions you approved. One thing needs your word: whether the fleet moves to the new controller.**
**Decisions I took on my own: none.**
**The one that mattered most is fixed and proven.** An app with no health check used to be **shut down by a successful update** — the machine waited five minutes for a check that could never arrive, then stopped a working app. Paperless-ngx, same app, same button: **before, it failed after 5 minutes and the app went dark. Now it finishes in 53 seconds and keeps running.**
**Five more, all proven on the test machine.**
- **Deleting an app while it is being backed up or restored is now refused**, with a plain sentence telling you to wait — instead of quietly tearing it down and leaving a ghost behind.
- **A delete now checks its own work.** The machine watches for 25 seconds afterwards and removes anything that comes back, and says whether it verified.
- **An app the machine has lost track of can now be deleted.** Before, if its record went wrong, no button worked and only a command line could clear it.
- **A failed update now keeps the app's own log** before shutting it down. Twice we lost the only evidence of why.
- **Deleting an app clears its old update status**, so a fresh install of the same app no longer shows a stale "Updated".
**The six versions you approved are live on the catalogue** — Emby, Ghost, Immich, Radarr, Sonarr, Termix. **None of them is installed on either demo machine**, so nothing updated; they simply show as available.
**What I did not do, and it is on purpose.** Two items from the plan are untouched and named rather than half-finished: finding out *why* an app's record goes wrong in the first place (I fixed the consequence, not the cause), and making a held app stop offering an Update button it will refuse.
**What went wrong on my side.** I lost **44 minutes** to my own progress-watchers: they waited for a build that had already succeeded, because each was watching for a name its own command contained. The same bug cost me a pile of stuck watchers earlier in the day. It is now written down as a rule so it does not happen a third time. I also nearly recorded one test as passing when it had proved nothing — the refusal I saw came from an older rule, not the new one. I caught it and re-ran it properly.
**Rows opened and closed.** Four closed, one narrowed to what is still unknown. The list stands at 325.
**What needs you — one question.**
1. **Shall the fleet move to controller 0.262.1?** Right now only the test machine has it. **My recommendation is yes:** every change here only refuses, waits, records, or removes what someone already asked to remove — none of them makes the machine do more on its own. *If you do nothing:* both demo machines stay on 0.261.0 and keep all six faults, including the one that shuts down a working app. Peti's machine is parked and would take it only if it ever comes back online.
**Nothing on your own machine or the off-site box was touched. The demo machines were not touched — they only see the six new version badges.**