The undo, built and proven live: controller v0.263.2 (09 decision 15)
gates / gates (push) Successful in 25s

- 09 §6.1 phase table (copying, undoing, undone), §6.1a SHIPPED with the two
  live-only defects, §6.4 part 1 SHIPPED.
- Capability map: a failed update is undone by the box - PROVEN-LIVE.
- Live evidence on 9202: three apps undone by the product with seeds before
  the backup, after it and seconds before the press read back; cut-off copy
  held honestly; power cut during the undo resumed; manual press after undo.
- Register: R-637, R-639, R-641, R-642 closed; R-638, R-640 narrowed; R-643
  ruled; R-646 opened. STATUS asks the floor question.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-23 12:25:13 +02:00
parent 5a349d9884
commit 05ea21e918
36 changed files with 2727 additions and 124 deletions
+16 -11
View File
@@ -1,21 +1,26 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-09-23 — your seven answers on automatic updates are now written rules. I tested the two new pieces by hand on the scratch machine. The build plan is ready for you, part by part.**
**Updated 2026-09-23 (afternoon) — a failed update now puts the app back by itself. Built, and proven on the scratch machine. One question for you: whether the fleet gets it.**
**Decisions I took on my own: none.**
**Decisions I took on my own: none.** Your two afternoon rulings are written down: the bake-off picks the copy method, and on update nights the full-system backup waits for the updates.
**The automatic undo works — but not with the tool the machine has today.** I broke three real updates on purpose. For two of the apps, the old version refused to start on the data the new version had changed. After I loaded the copy taken seconds before the update, all three apps came back with all their data, including what was written after the nightly backup. It took 16 and 38 seconds. **But the machine's current way of loading a copy fails on one app and leaves junk behind on another.** A fix is known and tested: empty the database first, then load, as one step.
**The bake-off.** Both ways of keeping the last-second copy passed every test on three apps. Copying the app's data folders won, because one of the apps has no database server and so gets no database copy at all. The extra downtime was 1 to 5 seconds.
**One dangerous finding.** A cut-off database copy loads as a "success" — into an empty database. The machine's copy checker would accept it. The fix is simple: check that the copy has its end marker before loading. This also protects today's restores.
**What the machine does now.** When an update fails its health check, the machine puts back the previous version and the data exactly as it was seconds before the update. It stops the app only if that undo fails too, and then it says so. It never touches the household's own folders (photos, documents).
**The step-by-step climb does not exist yet.** Today a machine two versions behind jumps straight to the newest one; the middle version never runs. The machine also cannot see old versions: it keeps only the newest copy of the catalogue. I propose a list of tested steps in each app's catalogue file.
**Proven on the scratch machine, through the same buttons the page uses:**
- Three apps undone by the machine. Data written before the backup, after it, and seconds before the update all came back. It took 30 to 52 seconds.
- The app page shows one line in Hungarian or English: the update failed, the machine put the app back, nothing was lost.
- A damaged copy is caught before anything is poured back, and the app is held with an honest sentence.
- A power cut in the middle of the undo: after restart, the machine finished the undo.
- A person pressed Update again after the catalogue was fixed, and it worked.
**The memory lesson from RomM is now in the test bench.** After an update, the bench runs the app for 10 minutes and watches memory. RomM's old setup failed in 76 seconds, so the bench would have caught it.
**What went wrong on my side.** My first build failed its own live test twice. Both times the app stayed stopped with an honest sentence, and the data was safe. The first fault: the machine never asked the old version the right health question. The second: it kept the wrong copy of that question. My unit tests had passed both times. I fixed both the same afternoon and proved the fixes. The released version is the third build.
**What needs you — two things.**
1. **The build plan: about 22 evenings in 10 parts.** I recommend starting with the undo (4 evenings). It also makes the manual Update button safer on its own. If you do nothing, nothing is built and updates stay manual.
2. **One question about the night schedule.** As ruled, updates run after the off-site copy and before the full-system backup. That gap is at most 15 minutes a night. Option A (my pick): the full-system backup waits for updates, still inside its own 4-hour window. Option B: keep the 15 minutes; a machine far behind takes weeks to catch up. If you do nothing, part 7 of the plan waits.
**Rows.** Four closed, two narrowed, one new. The list went from 337 to 335.
**Rows opened and closed.** Eight opened, none closed. The list went from 329 to 337.
**What needs you — one question.** Should the demo machines and the fleet get this version now?
- **Yes (my pick):** I raise the floor, and both demo machines update themselves within a minute. A failed update then puts the app back instead of stopping it.
- **Not yet:** nothing changes. A failed update keeps stopping the app until someone restores it.
**Nothing on your own machine, Peti's machine or the off-site box was touched. The demo machines were not touched. Only the scratch machine was used, and it is back on the real catalogue.**
**Nothing on your own machine, Peti's machine or the off-site box was touched. The demo machines were not touched. The scratch machine runs the new version and is back on the real catalogue.**