Update arc: the undo and ladder spiked, the build plan for the 2026-09-23 rulings
gates / gates (push) Successful in 26s
gates / gates (push) Successful in 26s
- Part 1 (09 §6.1a, audit): the undo performed by hand on 9202 for docmost (PostgreSQL), romm (MariaDB) and vikunja (SQLite volume) - all three came back with data written before AND after the backup. The product's loader cannot do it: over a migrated PG database it fails on the new tables' foreign keys; over MariaDB it leaves them behind. A truncated PG copy loads rc 0 into an empty database. No-DB apps have no last-second copy. - Part 2: one press jumps A -> C; the box's catalog clone is depth 1. Ladder format recommended: update_ladder in .felhom.yml, not git history. - Part 3: memory watch red-proof results (harness change in the catalog repo). - Part 4 (09 §6.4): ten parts, ~22 evenings; one open point (R-643). - Rows R-637..R-644 opened; R-446/450/451/462/463 updated. STATUS, CONTEXT. No product code. Live catalog untouched; 9202 back on it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -1,32 +1,21 @@
|
||||
# STATUS — what works, what's broken, what's next
|
||||
|
||||
**Updated 2026-09-23 (morning) — your seven answers on automatic updates are written into the design notes as rulings.** No later session will ask them again. The two new pieces your answers need — the automatic undo and the step-by-step climb — are being tested by hand on the scratch machine today, before anyone builds them. The build plan comes back to you part by part.
|
||||
|
||||
---
|
||||
|
||||
**Updated 2026-09-22 (late) — I fixed the six faults the two drill nights found in the update, delete and hold machinery, and shipped the six app versions you approved. One thing needs your word: whether the fleet moves to the new controller.**
|
||||
**Updated 2026-09-23 — your seven answers on automatic updates are now written rules. I tested the two new pieces by hand on the scratch machine. The build plan is ready for you, part by part.**
|
||||
|
||||
**Decisions I took on my own: none.**
|
||||
|
||||
**The one that mattered most is fixed and proven.** An app with no health check used to be **shut down by a successful update** — the machine waited five minutes for a check that could never arrive, then stopped a working app. Paperless-ngx, same app, same button: **before, it failed after 5 minutes and the app went dark. Now it finishes in 53 seconds and keeps running.**
|
||||
**The automatic undo works — but not with the tool the machine has today.** I broke three real updates on purpose. For two of the apps, the old version refused to start on the data the new version had changed. After I loaded the copy taken seconds before the update, all three apps came back with all their data, including what was written after the nightly backup. It took 16 and 38 seconds. **But the machine's current way of loading a copy fails on one app and leaves junk behind on another.** A fix is known and tested: empty the database first, then load, as one step.
|
||||
|
||||
**Five more, all proven on the test machine.**
|
||||
- **Deleting an app while it is being backed up or restored is now refused**, with a plain sentence telling you to wait — instead of quietly tearing it down and leaving a ghost behind.
|
||||
- **A delete now checks its own work.** The machine watches for 25 seconds afterwards and removes anything that comes back, and says whether it verified.
|
||||
- **An app the machine has lost track of can now be deleted.** Before, if its record went wrong, no button worked and only a command line could clear it.
|
||||
- **A failed update now keeps the app's own log** before shutting it down. Twice we lost the only evidence of why.
|
||||
- **Deleting an app clears its old update status**, so a fresh install of the same app no longer shows a stale "Updated".
|
||||
**One dangerous finding.** A cut-off database copy loads as a "success" — into an empty database. The machine's copy checker would accept it. The fix is simple: check that the copy has its end marker before loading. This also protects today's restores.
|
||||
|
||||
**The six versions you approved are live on the catalogue** — Emby, Ghost, Immich, Radarr, Sonarr, Termix. **None of them is installed on either demo machine**, so nothing updated; they simply show as available.
|
||||
**The step-by-step climb does not exist yet.** Today a machine two versions behind jumps straight to the newest one; the middle version never runs. The machine also cannot see old versions: it keeps only the newest copy of the catalogue. I propose a list of tested steps in each app's catalogue file.
|
||||
|
||||
**What I did not do, and it is on purpose.** Two items from the plan are untouched and named rather than half-finished: finding out *why* an app's record goes wrong in the first place (I fixed the consequence, not the cause), and making a held app stop offering an Update button it will refuse.
|
||||
**The memory lesson from RomM is now in the test bench.** After an update, the bench runs the app for 10 minutes and watches memory. RomM's old setup failed in 76 seconds, so the bench would have caught it.
|
||||
|
||||
**What went wrong on my side.** I lost **44 minutes** to my own progress-watchers: they waited for a build that had already succeeded, because each was watching for a name its own command contained. The same bug cost me a pile of stuck watchers earlier in the day. It is now written down as a rule so it does not happen a third time. I also nearly recorded one test as passing when it had proved nothing — the refusal I saw came from an older rule, not the new one. I caught it and re-ran it properly.
|
||||
**What needs you — two things.**
|
||||
1. **The build plan: about 22 evenings in 10 parts.** I recommend starting with the undo (4 evenings). It also makes the manual Update button safer on its own. If you do nothing, nothing is built and updates stay manual.
|
||||
2. **One question about the night schedule.** As ruled, updates run after the off-site copy and before the full-system backup. That gap is at most 15 minutes a night. Option A (my pick): the full-system backup waits for updates, still inside its own 4-hour window. Option B: keep the 15 minutes; a machine far behind takes weeks to catch up. If you do nothing, part 7 of the plan waits.
|
||||
|
||||
**Rows opened and closed.** Four closed, one narrowed to what is still unknown. The list stands at 325.
|
||||
**Rows opened and closed.** Eight opened, none closed. The list went from 329 to 337.
|
||||
|
||||
**The fleet is on the new controller — you said yes and it is done.** I raised the floor to **0.262.1** with the required agent version declared alongside it. **Both demo machines picked it up in under twenty seconds** and are running healthy. The drill machine is switched off and is **held back on purpose**: its helper software is older than the new controller needs, so the machine refuses to give it a version it cannot run. That is the guard working, not a failure. Peti's machine is parked and would take it only if it ever comes back online.
|
||||
|
||||
**What needs you: nothing.**
|
||||
|
||||
**Nothing on your own machine or the off-site box was touched. The demo machines were not touched — they only see the six new version badges.**
|
||||
**Nothing on your own machine, Peti's machine or the off-site box was touched. The demo machines were not touched. Only the scratch machine was used, and it is back on the real catalogue.**
|
||||
|
||||
Reference in New Issue
Block a user