v0.237.0: the Update button takes a backup first, and tells the truth (update arc slice 4 — R-448, R-443, R-439)
gates / gates (push) Successful in 13s

POST /api/stacks/{name}/update is now a guarded job answering 202:
cheap refusals (hold — R-439, busy, migration, deploying, memory via the
deploy's own memoryVerdict, a fixed 2 GB disk floor, and no restorable
Tier-2 copy) → backup-first when the proven copy is older than
update.backup_max_age (24h) → safety dump BEFORE the pin moves → pin →
pull (failure puts the pin back) → up → health (.felhom.yml check or 60 s
settle, update.health_timeout 5m). Not healthy → the app is stopped and
HELD (RestoreHold reason update_failed, same store and gate as R-379) and
the page names the backup to restore from; the pin stays. Success is only
ever update_phase=done after health (R-443). UpdateStack is deleted.

The restorable-unit predicate is EXTRACTED to backup.Tier2UnitRestorePoint
and shared with the backups page (row pinned unchanged). The copy is aged
by the last successful Tier-2 copy, not the manifest created_at — measured
on demo-hp that created_at moves only on definition changes.

Crash safety: update-journal.json before each phase; RecoverUpdates before
the boot sweep, ResumeInterruptedUpdates after the guards are wired.

Three unattended start paths ignored a hold and now honour it: the
drive-return gate (restart + boot recreate) and the nightly volume dump.
The nightly capture and Tier-2 run skip held apps so the restore point
survives. No automatic rollback — measured per-app; route back = restore.

Tests A–H across stacks/backup/api/web/cmd; six red-proofs seen to fail.

Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-13 11:41:31 +02:00
parent 1552716722
commit 0d402f711d
25 changed files with 2709 additions and 123 deletions
+48 -4
View File
@@ -472,7 +472,7 @@ through them: `EffectiveLifecycle()`, `CanInstall()`, `IsAbandoned()`.
Two additions, and **neither changes how an update behaves**. Slice 1 is a record; slice 2 is a label.
**Slice 1 — `app.yaml` gains `installed_images`.** After every successful `compose up` from
`StartStack`, `RestartStack`, `UpdateStack` and the deploy path, `Manager.recordInstalledImages`
`StartStack`, `RestartStack`, the guarded update (v0.237.0; `UpdateStack` before it) and the deploy path, `Manager.recordInstalledImages`
(`internal/stacks/installed.go`) reads what each container is ACTUALLY running and writes it down,
**keyed by compose SERVICE name**:
@@ -527,8 +527,8 @@ the current template pins and returns a `*MetaBadge` rendered by the existing `m
page. **Known limitation:** for the 23 floating pins (`postgres:16-alpine`, `mariadb:11.6`, …) the
reference can be identical while the image behind it has moved, so those apps can read „Naprakész"
when they may not be. Digest-level comparison needs a registry query and is deferred.
- **Information only.** The badge is wired to no action; `Frissítés`/`Újraindítás`/`Leállítás` are
byte-identical to before (`TestScenarioE_TheUpdateButtonIsUntouched`).
- **Information only.** The badge is wired to no action. (What the `Frissítés` button itself does
changed in v0.237.0 — see "The guarded update" below.)
Reasoning and the seven-slice plan: `felhom.eu/documentation/architecture/09-update-architecture.md`.
@@ -539,7 +539,7 @@ deliberate Update moves it. Everything else in a template — health checks, mem
fields — still arrives on the 15-minute cycle, and a broken definition still repairs itself.
- **`app.yaml` gains `pinned_images`** (service → ref): what the app is SUPPOSED to run. **Not**
`installed_images`, which is an observation. Written only by the deploy path, `UpdateStack`, a
`installed_images`, which is an observation. Written only by the deploy path, the guarded update, a
restore, and the one-time `AdoptPins`. **Absent = unpinned = pre-v0.235.0 behaviour.**
- **`applied-compose.yml`** in the stack dir stores the exact definition the pin came from. The
syncer copies only `docker-compose.yml` and `.felhom.yml`, so that name is safe.
@@ -555,6 +555,50 @@ fields — still arrives on the 15-minute cycle, and a broken definition still r
Reasoning: `felhom.eu/documentation/architecture/09-update-architecture.md` §3, §5.
#### The guarded update (v0.237.0 — update arc slice 4)
**`POST /api/stacks/{name}/update` no longer updates on the spot.** It refuses what it must, starts a
job, and answers **202**. The page polls `GET /api/stacks/{name}`.
| field | meaning |
|---|---|
| `updating` | a guarded update is in progress |
| `update_phase` / `update_phase_label` | `checking`, `backing-up`, `safety-dump`, `pinning`, `pulling`, `starting`, `verifying`, `done`, `failed` — and the Hungarian label for each |
| `update_error` | the customer sentence when the update did not complete |
| `hold_reason` | the hold's sentence while the app is held (failed update OR failed restore) |
**Refused with 409 before anything moves:** held; a backup, restore, app-data op or quiesce holding
the app; a migration; already updating; deploying; not enough memory for the NEW template's request
(the deploy's own `memoryVerdict`, releasing the app's current request); less than **2 GB** free on
the Docker data root (a fixed floor — image sizes are not known without a registry query); and **no
restorable backup** (no openable Tier-2 recovery unit with a proven copy).
**The sequence.** A proven copy older than `update.backup_max_age` is refreshed first
(`RunAppBackupNow`: this app's DB dump, volume dump, unit capture, Tier-2 copy). Then a database
safety dump, then the pin moves, then pull, `up`, and the health wait (`.felhom.yml` check, or 60 s of
every container running for an app with none; bounded by `update.health_timeout`). **A failed pull
puts the pin back. An app that does not become healthy is stopped and HELD** — the pin stays on the
new version, and the hold sentence names the backup it can be restored from. A successful unit
restore lifts an update hold.
**Config (`controller.yaml`):**
```yaml
update:
backup_max_age: 24h # a proven Tier-2 copy older than this is refreshed before the update
health_timeout: 5m # how long the new version has to become healthy before the app is held
```
**Crash safety:** `<data>/update-journal.json` is written before every phase; `RecoverUpdates` (before
the boot sweep) puts a pin back or marks an interrupted update for `ResumeInterruptedUpdates`.
**Every unattended start path honours a hold:** the boot sweep and the app-stop guard (as before), and
since v0.237.0 the drive-return gate and the nightly volume dump. The nightly capture and Tier-2 run
skip a held app so its restore point is not overwritten.
**Not done, deliberately:** the old version is never put back automatically — whether that works is
per-app and was measured unpredictable. Reasoning: `felhom.eu/documentation/architecture/09-update-architecture.md` §6.
#### App Info Pages
Each app can define rich metadata in `.felhom.yml`: