v0.237.0: the Update button takes a backup first, and tells the truth (update arc slice 4 — R-448, R-443, R-439)
gates / gates (push) Successful in 13s
gates / gates (push) Successful in 13s
POST /api/stacks/{name}/update is now a guarded job answering 202:
cheap refusals (hold — R-439, busy, migration, deploying, memory via the
deploy's own memoryVerdict, a fixed 2 GB disk floor, and no restorable
Tier-2 copy) → backup-first when the proven copy is older than
update.backup_max_age (24h) → safety dump BEFORE the pin moves → pin →
pull (failure puts the pin back) → up → health (.felhom.yml check or 60 s
settle, update.health_timeout 5m). Not healthy → the app is stopped and
HELD (RestoreHold reason update_failed, same store and gate as R-379) and
the page names the backup to restore from; the pin stays. Success is only
ever update_phase=done after health (R-443). UpdateStack is deleted.
The restorable-unit predicate is EXTRACTED to backup.Tier2UnitRestorePoint
and shared with the backups page (row pinned unchanged). The copy is aged
by the last successful Tier-2 copy, not the manifest created_at — measured
on demo-hp that created_at moves only on definition changes.
Crash safety: update-journal.json before each phase; RecoverUpdates before
the boot sweep, ResumeInterruptedUpdates after the guards are wired.
Three unattended start paths ignored a hold and now honour it: the
drive-return gate (restart + boot recreate) and the nightly volume dump.
The nightly capture and Tier-2 run skip held apps so the restore point
survives. No automatic rollback — measured per-app; route back = restore.
Tests A–H across stacks/backup/api/web/cmd; six red-proofs seen to fail.
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
+48
-4
@@ -472,7 +472,7 @@ through them: `EffectiveLifecycle()`, `CanInstall()`, `IsAbandoned()`.
|
||||
Two additions, and **neither changes how an update behaves**. Slice 1 is a record; slice 2 is a label.
|
||||
|
||||
**Slice 1 — `app.yaml` gains `installed_images`.** After every successful `compose up` from
|
||||
`StartStack`, `RestartStack`, `UpdateStack` and the deploy path, `Manager.recordInstalledImages`
|
||||
`StartStack`, `RestartStack`, the guarded update (v0.237.0; `UpdateStack` before it) and the deploy path, `Manager.recordInstalledImages`
|
||||
(`internal/stacks/installed.go`) reads what each container is ACTUALLY running and writes it down,
|
||||
**keyed by compose SERVICE name**:
|
||||
|
||||
@@ -527,8 +527,8 @@ the current template pins and returns a `*MetaBadge` rendered by the existing `m
|
||||
page. **Known limitation:** for the 23 floating pins (`postgres:16-alpine`, `mariadb:11.6`, …) the
|
||||
reference can be identical while the image behind it has moved, so those apps can read „Naprakész"
|
||||
when they may not be. Digest-level comparison needs a registry query and is deferred.
|
||||
- **Information only.** The badge is wired to no action; `Frissítés`/`Újraindítás`/`Leállítás` are
|
||||
byte-identical to before (`TestScenarioE_TheUpdateButtonIsUntouched`).
|
||||
- **Information only.** The badge is wired to no action. (What the `Frissítés` button itself does
|
||||
changed in v0.237.0 — see "The guarded update" below.)
|
||||
|
||||
Reasoning and the seven-slice plan: `felhom.eu/documentation/architecture/09-update-architecture.md`.
|
||||
|
||||
@@ -539,7 +539,7 @@ deliberate Update moves it. Everything else in a template — health checks, mem
|
||||
fields — still arrives on the 15-minute cycle, and a broken definition still repairs itself.
|
||||
|
||||
- **`app.yaml` gains `pinned_images`** (service → ref): what the app is SUPPOSED to run. **Not**
|
||||
`installed_images`, which is an observation. Written only by the deploy path, `UpdateStack`, a
|
||||
`installed_images`, which is an observation. Written only by the deploy path, the guarded update, a
|
||||
restore, and the one-time `AdoptPins`. **Absent = unpinned = pre-v0.235.0 behaviour.**
|
||||
- **`applied-compose.yml`** in the stack dir stores the exact definition the pin came from. The
|
||||
syncer copies only `docker-compose.yml` and `.felhom.yml`, so that name is safe.
|
||||
@@ -555,6 +555,50 @@ fields — still arrives on the 15-minute cycle, and a broken definition still r
|
||||
|
||||
Reasoning: `felhom.eu/documentation/architecture/09-update-architecture.md` §3, §5.
|
||||
|
||||
#### The guarded update (v0.237.0 — update arc slice 4)
|
||||
|
||||
**`POST /api/stacks/{name}/update` no longer updates on the spot.** It refuses what it must, starts a
|
||||
job, and answers **202**. The page polls `GET /api/stacks/{name}`.
|
||||
|
||||
| field | meaning |
|
||||
|---|---|
|
||||
| `updating` | a guarded update is in progress |
|
||||
| `update_phase` / `update_phase_label` | `checking`, `backing-up`, `safety-dump`, `pinning`, `pulling`, `starting`, `verifying`, `done`, `failed` — and the Hungarian label for each |
|
||||
| `update_error` | the customer sentence when the update did not complete |
|
||||
| `hold_reason` | the hold's sentence while the app is held (failed update OR failed restore) |
|
||||
|
||||
**Refused with 409 before anything moves:** held; a backup, restore, app-data op or quiesce holding
|
||||
the app; a migration; already updating; deploying; not enough memory for the NEW template's request
|
||||
(the deploy's own `memoryVerdict`, releasing the app's current request); less than **2 GB** free on
|
||||
the Docker data root (a fixed floor — image sizes are not known without a registry query); and **no
|
||||
restorable backup** (no openable Tier-2 recovery unit with a proven copy).
|
||||
|
||||
**The sequence.** A proven copy older than `update.backup_max_age` is refreshed first
|
||||
(`RunAppBackupNow`: this app's DB dump, volume dump, unit capture, Tier-2 copy). Then a database
|
||||
safety dump, then the pin moves, then pull, `up`, and the health wait (`.felhom.yml` check, or 60 s of
|
||||
every container running for an app with none; bounded by `update.health_timeout`). **A failed pull
|
||||
puts the pin back. An app that does not become healthy is stopped and HELD** — the pin stays on the
|
||||
new version, and the hold sentence names the backup it can be restored from. A successful unit
|
||||
restore lifts an update hold.
|
||||
|
||||
**Config (`controller.yaml`):**
|
||||
|
||||
```yaml
|
||||
update:
|
||||
backup_max_age: 24h # a proven Tier-2 copy older than this is refreshed before the update
|
||||
health_timeout: 5m # how long the new version has to become healthy before the app is held
|
||||
```
|
||||
|
||||
**Crash safety:** `<data>/update-journal.json` is written before every phase; `RecoverUpdates` (before
|
||||
the boot sweep) puts a pin back or marks an interrupted update for `ResumeInterruptedUpdates`.
|
||||
|
||||
**Every unattended start path honours a hold:** the boot sweep and the app-stop guard (as before), and
|
||||
since v0.237.0 the drive-return gate and the nightly volume dump. The nightly capture and Tier-2 run
|
||||
skip a held app so its restore point is not overwritten.
|
||||
|
||||
**Not done, deliberately:** the old version is never put back automatically — whether that works is
|
||||
per-app and was measured unpredictable. Reasoning: `felhom.eu/documentation/architecture/09-update-architecture.md` §6.
|
||||
|
||||
#### App Info Pages
|
||||
|
||||
Each app can define rich metadata in `.felhom.yml`:
|
||||
|
||||
Reference in New Issue
Block a user