Update night 2026-09-21: the full record, twelve rows, and the answers to five of the seven questions
gates / gates (push) Successful in 28s

The drill is complete. Teardown done in three layers plus Gitea; the live catalog's every `image:`
line is proven identical to before.

WHAT WAS MEASURED. 21 edges across 19 apps, on scratch guest 9202 through the product's own
guarded Update, against a PRIVATE DRILL CATALOG so the live catalog carried no test reference at
any point: 14 proven, 3 failed, 4 inconclusive. Each app seeded and read back through its OWN
front door, with a negative control on every readback. Ten of the fourteen printed a verbatim
migration line. Up from the three apps this project had ever measured.

THE RESULT THAT MATTERS. R-618, P1: three of the 53 templates name a health probe the app does not
answer, and because the guarded update WAITS on that same probe, a SUCCESSFUL update ends by
STOPPING a working app. tandoor was measured serving HTTP 200 on the new version at four samples
across five minutes, docker's own healthcheck green, and was then stopped and the household sent
to a restore they did not need. zipline and wger are the same defect, both confirmed live. The
gate that catches all three is static and cheap: both health checks already sit in the same file.

WHAT THE NIGHT ANSWERED that was open. The UNATTENDED HOLD (312.9 s, pressed once, never again) —
which needed a purpose-built image store, because the rule that makes automatic updates safe is
the same rule that refuses the obvious way to break one. MariaDB across a major through the real
button, all four observables, first time. PostgreSQL across a major, refusing exactly as predicted,
with the conversion costed at ~9 s of engine work. There is NO single-flight: five updates ran at
once and all ended honest. And the two EARLY power-cut phases nobody had cut in.

TWELVE NEW ROWS (R-615..R-626), register 303 -> 315, and eight existing rows updated with what was
measured — including two CORRECTIONS: R-606 records the pre-flight refusals as reaching an English
household in English and they do not, and R-446/R-458 are both narrower than their rows state.

Two instrument fixes were needed before anything could be trusted: the unattended caller turned
every success into a timeout (R-623), and one of my own reproductions was wrong and is kept
labelled with what it actually measured.

Interventions: zero. No controller, agent or hub code written. The hub was never touched beyond
the floor the operator asked for.

Gates: repo_gates.py --fast, all 15 OK.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-21 22:34:24 +02:00
parent 9c69b3ff07
commit 8d786f7940
39 changed files with 1657 additions and 128 deletions
+31
View File
@@ -1,5 +1,36 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-09-21 (overnight) — I tested the "update my app" button on as many apps as fit in a night, on good days and bad ones. Fourteen updates are proven safe. Three apps are broken in a way that shuts down a working app, and I would fix that first.**
**Decisions I took on my own: none.** Nothing tonight needed a choice you had not already made.
**The fleet version is 0.261.0.** You asked for that. Both demo machines took it **thirteen seconds** after I saved it. The third machine is switched off and will take it when it comes back.
**What I did.** Twenty-one real updates on a scratch machine, each app installed at the version our catalog has today, filled with real data through the app's own front door, backed up, updated to the newer version that really exists upstream, then the data read back. **Fourteen proven, three failed, four I could not judge.** Ten of the fourteen printed their own "I am rewriting the database" line — so the data really was rewritten, and it still came back.
**The one I would fix first — three apps tell the machine they are broken when they are fine.** Tandoor, Zipline and Wger each have one wrong number or address in their settings file, so the machine knocks on the wrong door and hears nothing. That alone would only be a wrong label. **But the update also waits on that same check** — so when one of these apps updates *successfully*, the machine waits five minutes, decides it failed, **shuts the working app down**, and tells the household to restore from backup. I watched Tandoor serve customers for five minutes on its new version and then get switched off. Nothing is lost and the restore works, but the household loses their app and does work they did not need to do. **A cheap check would catch all three: each of those files already contains the right answer a few lines further down.**
**What broke, and whether the household could get out.**
- **Adventurelog's newer version rewrites the database and then never starts.** The machine did everything right: backup one minute before, waited the full five minutes, stopped the app so the data could not be hurt, and said in one sentence where the copy is, when it was made and what is inside it. I pressed that restore: **back in 75 seconds.** That app must not be moved to the newer version.
- **Tandoor, as above.** Restored in 32 seconds.
- **PostgreSQL will not jump a version.** Exactly as expected: the database engine refuses to start, the app stops honestly, the data is untouched, and the restore brings it back in 29 seconds. I also rehearsed the conversion that would let those eleven apps ever move: **about nine seconds of database work, under three minutes end to end.** That is a maintenance window, not a project.
- **MariaDB, by contrast, jumps a version cleanly** — and I pressed that through the real button for the first time. The engine converted the data, said so in its own words, and took its own backup first.
**The machine also passed every bad day I could invent.** A version that cannot be downloaded: refused in one second, app keeps running. Two updates at once, then five: all ran together and all ended honestly. Power cut in the middle of the backup: the machine recovered itself and said so. Disk nearly full: refused before touching anything. An app left stopped by a failed update stayed stopped after a power cut — *"whatever is holding it owns its recovery."*
**Four smaller faults, all written down.** A message that says "Updated" when nothing was updated. The failure message shown in Hungarian on the English page — including the sentence that tells a household where their files are. An app the household deleted that **came back by itself**, empty. And when an update fails, the machine deletes the broken app's log before anyone can read why.
**Rows opened and closed.** Twelve new, eight existing ones updated with what was measured. The list went from 303 to 315.
**What needs you.**
1. **Rotate the Gitea `admin` token.** The machine stores it in plain text inside its copy of the catalog, so an ordinary diagnostic printed it into my log. *If you do nothing:* the token keeps working and anyone with my session transcript has it.
2. **The promotion list** — fourteen updates proven safe enough to move on the real catalog, and three named that must not move. Moving a version is your call, never mine. *If you do nothing:* nothing breaks; those apps drift further from upstream each month.
3. **The seven questions** about automatic updates now have real facts beside them — including the two that had never been measured: what a stopped app looks like when nobody was watching, and what a database conversion costs. They are still yours. *If you do nothing:* the automatic-update work cannot start, because everything hangs off the first one — *may the machine update apps by itself at night?*
**The live catalog was never touched with a test change.** Not once, not for thirteen minutes. Everything ran against a private drill copy on a scratch machine. The only change to the real catalog is test code, and I have proved every app's version line is identical to before.
---
**Updated 2026-09-21 (evening) — I cut the power to a machine in the middle of an app update, three times, after the new version had already changed the data. It survived every time.**
**The fleet version is now 0.260.0.** You approved it. Both demo machines have it. Three machines are