R-379/R-380 docs: the failure ladder, the drill record, register housekeeping
gates / gates (push) Successful in 17s

07-backup-architecture.md 6.3 gains a dated [DESIGN] paragraph on replay ->
rollback -> hold, including why no engine flag closes it: --single-transaction
makes Postgres atomic, MariaDB DDL is not transactional, so the rollback is the
fix and the flag is a belt.

Drill record for the live walk, including the TWO defects the walk found in the
fix itself (a rollback into a re-created container; an operator route that
cleared the file while the running controller kept refusing) and the ONE
red-proof that PASSED, which is reported rather than omitted.

R-379..R-382 compressed into CLOSED-ITEMS.md. OPEN-ITEMS 330683 -> 325236 bytes.

STATUS.md restates the outcome and names the next operator step.
This commit is contained in:
2026-08-22 18:31:23 +02:00
parent 4e488321bf
commit a8caa0fdde
23 changed files with 744 additions and 75 deletions
+21 -10
View File
@@ -12,16 +12,14 @@ NOT yet delivered: two steps below are yours.**
*This section is allowed to be longer than one screen, and each item says what happens if you do
nothing.*
1. **Vouch the golden carrying controller 0.219.0** — Hub → Configuration → Day-0 artifacts.
**It is baked, published and round-trip verified** (`documentation/tests/golden-0.219.0-2026-08-22/`);
only the vouch is left, and only you can do it. **It is a THREE-field save, not one:**
`golden_version` → **0.219.0**, `agent_version` → **0.130.0**, `min_agent` → **0.129.0**. Moving
`golden_version` alone ships this controller onto an agent older than it declares it needs.
**If you do nothing:** a machine installed today still receives 0.218.0 — the image exists, on the
shelf, undelivered. Reversible: re-select the old values and Save.
2. **Then raise the auto-update floor to 0.219.0 — last, in a separate save.** It acts within seconds.
**If you do nothing:** every existing machine stays on 0.218.0, so the fix below reaches nobody and
40 of 53 apps stay un-restorable on the actual fleet. *(register: R-343's rule)*
1. **Vouch the golden carrying controller 0.220.2** — Hub → Configuration → Day-0 artifacts.
**It is already baked, published and round-trip verified**
(`documentation/tests/golden-0.220.2-2026-08-22/`); only the vouch is left, and only you can do it.
**It is a THREE-field save:** `golden_version` → **0.220.2**, `agent_version` → **0.130.0**,
`min_agent` → **0.129.0**. **Then** raise the floor to **0.220.2**, last, in its own save.
**If you do nothing:** the fleet stays on 0.219.0, so a failed database restore still leaves an app
broken with an unusable copy — the thing today's release fixes reaches nobody. New machines still
receive 0.219.0. The build system stays red about it and will mail you on every push.
3. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
@@ -52,6 +50,19 @@ off. **`peti-felhom` is a real machine we have not heard from since 15 July** an
## Shipped
- **A failed database restore now puts your data back by itself** (R-379/R-380, controller 0.220.2,
proven on `demo-hp`). Until today, if a restore of an app's database went wrong, the machine had
already taken a copy of your live database — a good copy — and **nothing in the product could put it
back.** You were shown a filename. On one of the two database types it was worse: part of the
restore applied, part did not, and **the dashboard said the app was healthy**. Now the machine puts
your own copy back automatically and says plainly: the restore failed, your data is as it was, the
app is running. Proven on both database types, byte-identical both times.
**If even that fails**, the app is deliberately **stopped and held** rather than started — a running
app on a half-written database lets you type into it and makes the damage permanent — and you are
told to contact us. That was your ruling this morning. **Two things also stopped:** the error no
longer pastes raw database text at you (it was 615 bytes once, including rows out of your own
database), and the undo copies no longer pile up forever — three per app, and they were being copied
off-site permanently.
- **The off-site restore now works for the other 40 apps** (R-356, controller 0.219.0, proven on
`demo-hp`). It used to refuse before starting, tell the customer a running app „nincs telepítve",
and send them to reinstall it "to the same place" — a place those 40 apps never offer, because they