Slice 4 shipped (R-448/R-443/R-439 CLOSED, proven live); R-472..R-476; the floor-between-bakes claim corrected
gates / gates (push) Successful in 19s

Controller v0.237.0-v0.238.1: the Update button is a guarded job — refusals, backup-first when the
proven Tier-2 copy is stale, safety dump, pin, pull (pin back on failure), health, HOLD on failure.
Proven live on demo-hp: A, B, E, F, H and the restore walk (audits/slice4-2026-09-13/).

Correction to this morning's pages: between golden bakes the hub HOLDS a floor above the vouched
golden, so a release does not reach the fleet by floor (R-472, operator decision). Corrected in the
runbook, STATUS, CONTEXT, R-468 and the gate docstring.

Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-13 12:30:17 +02:00
parent abe567e14d
commit 5ef0f52bcd
46 changed files with 1029 additions and 285 deletions
@@ -100,6 +100,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|---|---|---|---|---|
| Deploy an app from the catalog (env config, memory guard, health-aware progress) | controller, catalog (~52 apps, images pinned) | **PROVEN-LIVE** | `CAMPAIGN-2` T-DEPLOY-SET (7 apps, env config, health-aware); `RERUN-p1p3` (×4 PASS) | Memory-guard FIRING is not live-shown (T-RES-MEMGUARD never fired: ample RAM / auth-walled) — implemented + unit-level only |
| App lifecycle: start/stop/restart/update/logs/remove/redeploy | controller | **PROVEN-LIVE — the ACTIONS work. NARROWED 2026-09-13: `CAMPAIGN-3` proved `remove` removes the APP, not the DATA — the "delete my data" half was INERT on every box until controller v0.236.0 (R-442). RE-PROVEN 2026-09-13 on demo-hp: data written by the app itself (63 MB) gone after removal and listed; an unresolvable data location is REFUSED (409) with the app kept; an SSD app gets `[]` and a note.** | `CAMPAIGN-2` T-LIFECYCLE (stop/start/restart/update/logs); remove (app only) live in `CAMPAIGN-3`; **remove WITH data: `audits/R442-2026-09-13/`**; **data behaviour: `audits/SPIKE-app-update-2026-09-01.md` (2026-09-01)** | Redeploy-after-remove edge remains open (T-REMOVE-REDEPLOY never cleanly passed — stale dryrun journal); non-pilot-critical |
| **Update is GUARDED: it refuses without a restorable backup, backs up first when the copy is stale, and HOLDS an app that does not come up** | controller **v0.237.0 + v0.238.0 + v0.238.1** | **PROVEN-LIVE (2026-09-13)** — scenarios A (real upgrade, success only after health), B (stale copy → backup first), E (pull failure → pin back, app untouched), F (never healthy → held, hold text on API and page), H (start/restart/update and the boot sweep all refuse the held app) and **the restore walk** (Mentések unit restore → back on the old version, hold cleared), on demo-hp with a throwaway app | **`audits/slice4-2026-09-13/`** (live/, redproofs/, gates/); design `architecture/09-update-architecture.md` §6.1 | **Tier-2-only precondition** — an app with no Tier-2 copy cannot be updated (R-475); no automatic rollback, by measurement; a release does not reach the fleet by floor between golden bakes (R-472) |
| **What `restart` and `update` do to a deployed app whose compose file the catalog already moved** | controller **v0.235.0** | **CHANGED 2026-09-06 — they NO LONGER upgrade it.** The row below records what shipped; this text records what it replaced, because every box under v0.235.0 still behaves the old way. **Up to v0.234.0: PROVEN-LIVE (2026-09-01) — they UPGRADE it.** Every lifecycle action ends in `docker compose up -d`, which makes the container match the file and PULLS the image itself when it is missing (measured: 18.3 s with a pull, 0.5 s without; negative control with an unchanged file did not even recreate the container). This is DELIBERATE on the restart path — `Manager.RestartStack` says so in a comment — but the syncer moves the file under a deployed app on a 15-minute cycle with no deployed check (R-438), and NOTHING tells the customer. | `audits/SPIKE-app-update-2026-09-01.md` §2, §3 | **No safety copy is taken by any of them** — `writeSafetyDump` is DATABASE-ONLY and is not on the update path at all. R-438, R-440, R-443. |
| **Whether the box UPGRADES an app by itself, with nobody pressing anything** | controller | **PROVEN-LIVE (2026-09-01) — YES, but only when an app fails to come back.** A plain power cut does NOT upgrade: Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. When an app does NOT return, `Reconciler.Run` (`bootrecon.go:269`) calls `StartStack` -> `compose up -d` and the app comes back on the NEW version, unattended (measured). **13 non-API call sites across 9 files reach `up -d` this way** — not the five previously believed. | `audits/SPIKE-app-update-2026-09-01.md` §2, §8 | The drive-return gate (`intermediary.go:222`) and `AppStopGuard.Recover` (`appstop_marker.go:283`) call the same function; located by reading, **not exercised live** — stated as such. |
| **Whether an app UPGRADE can be undone** | controller + catalog | **PROVEN-LIVE (2026-09-01) — NO, and "rollback" is the wrong word for it.** Once a migration has RUN, putting the old image tag back yields a container that refuses to start: Nextcloud — *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"*. A 3-major jump is refused outright (*"only possible to upgrade one major version at a time"*) and IS recoverable, precisely because nothing migrated. Positive control: the data is not destroyed — returning to 32.0.9 restored both seeded markers byte-identical. | `audits/SPIKE-app-update-2026-09-01.md` §7 | The only route back is restoring DATA from a copy taken BEFORE the update — which no update path takes. And a restore's image-level rollback is itself overwritten by the syncer within 15 minutes (R-441). R-40 is confirmed live by the same measurement. |
@@ -301,11 +301,80 @@ earlier feature is the failure mode to look for whenever a file changes meaning.
| **1b** | **Seed the record for apps nobody touches** — a startup backfill, so the label is not restricted to apps that happen to get restarted. | **SHIPPED, controller v0.234.0 (2026-09-03)** |
| **2** | **One badge says whether the app is current** — „Naprakész" / „Frissítés elérhető — N napja", from `catalog_since`. No version number. | **SHIPPED, controller v0.233.0 + catalog `69761cf` (2026-09-02)** |
| **3** | **The compose file becomes DERIVED** — the pin in `app.yaml` wins; the syncer renders instead of copying. | **SHIPPED, controller v0.235.0 (2026-09-06)** — operator ruling §3.4 |
| **4** | **A guarded update** — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). | OPEN — R-448 |
| **4** | **A guarded update** — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). | **SHIPPED + PROVEN LIVE, controller v0.237.0 (job) + v0.238.0 (page) + v0.238.1 (2026-09-13)** — §6.1 |
| **5** | **An upgrade test that runs again** — a harness that upgrades a real app with real data in it and asks the app for the data back. | **SHIPPED, `app-catalog/scripts/upgrade-test.py` (2026-09-06)** — 7 edges, 3 apps; see §4.1 and §10 |
| **6** | **A version sequence** — updates automatic within a major, a human across one; **an engine change gets its own edge.** | OPEN — R-450 |
| **7** | **A fleet sweep pipeline** — the operator can see, and move, how far behind every box is. | OPEN — R-451 |
### 6.1 Slice 4 as SHIPPED (controller v0.237.0 + v0.238.0, 2026-09-13)
`POST /api/stacks/{name}/update` is a guarded job. It answers **202** at once; the outcome exists only
on `GET /api/stacks/{name}` (`updating`, `update_phase`, `update_phase_label`, `update_error`,
`hold_reason`), and `update_phase=done` is written only after the app's health is known. **R-443 is
closed by construction: nothing reports an update complete on the compose exit code.**
**The sequence, and the order is the design:**
| # | phase | what happens | on failure |
|---|---|---|---|
| 0 | refusals (409, before the intent is recorded) | held (R-439), busy (backup/restore/app-data op/quiesce), migration, already updating, deploying, memory (the deploy's `memoryVerdict`, releasing the app's own request), disk (**fixed 2 GB floor** — image size unknown without a registry), **no restorable Tier-2 copy** | nothing moves, nothing is recorded |
| 1 | `checking` | re-reads the precondition | nothing moves |
| 2 | `backing-up` — only when the proven copy is older than `update.backup_max_age` | `RunAppBackupNow`: this app's DB dump → volume dump → unit capture → Tier-2 copy | refused with the backup's own error; nothing moves |
| 3 | `safety-dump` | `WriteUpdateSafetyDump` (R-361's undo copy) — **before the pin moves** | refused; nothing moves |
| 4 | `pinning` | the previous definition is copied aside and journaled, then the pin advances | pin put back |
| 5 | `pulling` | `compose pull` | **pin and definition PUT BACK** — nothing ran (Scenario E) |
| 6 | `starting` | `compose up -d --remove-orphans` | stop + HOLD |
| 7 | `verifying` | the `.felhom.yml` health check through the existing probe, or 60 s of every container running and none restarting; bounded by `update.health_timeout` | **stop + HOLD; the pin STAYS** — the migration may have run (Scenario F) |
| 8 | `done` | installed images recorded, journal cleared | — |
**The two knobs** (`controller.yaml`, operator-owned): `update.backup_max_age` (default `24h`) and
`update.health_timeout` (default `5m`).
**The precondition is the existing verified backup, not a new copy** (§3 decision 1). It is
`backup.Tier2UnitRestorePoint` — the SAME predicate that permits the destructive „Teljes
visszaállítás", extracted from the backups page rather than copied. **The copy is aged by the last
SUCCESSFUL Tier-2 copy, not by the unit manifest's `created_at`**, and that was measured before it was
designed: a capture rewrites the manifest only when the app's DEFINITION changes, so on demo-hp
bookstack's mirror held a 2026-09-13T00:30Z dump under a manifest dated 2026-09-12T02:15:29Z. Aged by the
manifest, a quiet app would be "stale" forever and a backup-first would not fix it. **The predicate is
Tier-2-only, as specified — an app with no Tier-2 copy cannot be updated (R-475).**
**The hold** is `settings.RestoreHold` with `reason: update_failed` and `copy_date` — the SAME store and
gate as R-379, so every start path that already honoured a restore hold honours this one. A successful
unit restore lifts an update hold (only that kind). **Three unattended paths honoured no hold before
v0.237.0 and now do:** the drive-return gate's restart and boot recreate, and the nightly volume dump
(which ends in `StartStack`). The nightly capture and Tier-2 run skip a held app, so the restore point
the hold text names is never overwritten.
**Crash safety is a journal**, `<data>/update-journal.json`, written before every phase. `RecoverUpdates`
runs before the boot sweep: interrupted before the pin → dropped; while pinning/pulling → pin put back;
after `up` → marked Updating (the boot sweep and the dead-app alarm leave it alone) and resumed by
`ResumeInterruptedUpdates` once the backup side is wired, ending healthy or held.
**THE ABORT DECISION, restated so it is not reopened: NOT BUILT, BY MEASUREMENT.** Whether an old image
starts on data a new one migrated is per-app (§4.1: PrivateBin yes, Docmost and Nextcloud no) and cannot
be predicted. So the box never puts the old version back by itself. **The route back is the restore**,
and slice 4's whole purpose is that the restore exists before anything moves. Per-app abort data, where
the harness has proven it, is slice 6's.
**Not gated here:** a multi-major jump (R-40). It fails health and is held honestly; stepping is slice 6.
**The release could not reach the fleet by floor — R-472.** The hub holds a controller floor above the
vouched golden (publish-train rule 1), so under the weekly golden cadence (R-468) v0.237.0 and v0.238.0
were hand-deployed to the demo guests. That contradiction is an operator decision.
**Found live, fixed in v0.238.1: the nightly legs must leave an app alone WHILE it is updating, not
only once it is held.** In Scenario F the periodic unit capture ran at 10:17:09 — inside the 5-minute
health wait, 53 s before the hold — and wrote the never-started definition into the app's PRIMARY unit.
The Tier-2 mirror the hold names survived only because Tier 2 is daily. `backup.Manager.isHeld` now also
answers true for an app a guarded update is moving (`SetUpdatingCheck`).
**Proven live on demo-hp, 2026-09-13**, with a throwaway uptime-kuma and real catalog tag changes (each
reverted in the same phase): A (2.3.2→2.4.0, done after health), B (`backup_max_age: 2m` → backup first),
E (non-existent tag → pin back, container untouched), F (`alpine:3.20` → held), H (three buttons and the
boot sweep refuse the held app), and the restore walk (Mentések unit restore → back on 2.4.0, hold
cleared). Live evidence: `audits/slice4-2026-09-13/`.
### The verdict record — the contract Slice 6 carries
Decided here rather than invented twice. The harness writes one of these per edge, beside its
@@ -458,9 +527,11 @@ Version strings stay in the logs, the API and the hub.
§5.4. So a frozen app can receive a health check written for a NEWER version and read as degraded.
**The failure direction is a false alarm, never data loss**, and freezing `.felhom.yml` would break
the update badge by withholding `catalog_since`. **R-458.**
6. **The Update button is still unguarded.** It takes no backup, has no rollback, and can still
attempt a multi-major jump the app will refuse (R-40). **Slice 3 did not change that and must not
be read as having done so** — the precondition is slice 4 (R-448).
6. ~~**The Update button is still unguarded.**~~ **CLOSED 2026-09-13 by slice 4 (v0.237.0, §6.1).** It
refuses without a restorable, proven Tier-2 copy, backs up first when that copy is stale, takes a
safety dump, and holds an app that does not come up. **What stays true:** it still has no automatic
rollback (deliberately, §6.1) and can still attempt a multi-major jump the app will refuse (R-40) —
that now ends HELD rather than crash-looping behind a green button.
7. ~~**An engine major can be applied without its datadir upgrade, and nothing notices.**~~ **CLOSED
2026-09-13 for MariaDB (R-459):** every `mariadb:` sidecar carries `MARIADB_AUTO_UPGRADE=1`, and the
harness shows the conversion running on the E3/E3b edges (§3 decision 5). **What stays true:** the