v0.237.0: the Update button takes a backup first, and tells the truth (update arc slice 4 — R-448, R-443, R-439)
gates / gates (push) Successful in 13s

POST /api/stacks/{name}/update is now a guarded job answering 202:
cheap refusals (hold — R-439, busy, migration, deploying, memory via the
deploy's own memoryVerdict, a fixed 2 GB disk floor, and no restorable
Tier-2 copy) → backup-first when the proven copy is older than
update.backup_max_age (24h) → safety dump BEFORE the pin moves → pin →
pull (failure puts the pin back) → up → health (.felhom.yml check or 60 s
settle, update.health_timeout 5m). Not healthy → the app is stopped and
HELD (RestoreHold reason update_failed, same store and gate as R-379) and
the page names the backup to restore from; the pin stays. Success is only
ever update_phase=done after health (R-443). UpdateStack is deleted.

The restorable-unit predicate is EXTRACTED to backup.Tier2UnitRestorePoint
and shared with the backups page (row pinned unchanged). The copy is aged
by the last successful Tier-2 copy, not the manifest created_at — measured
on demo-hp that created_at moves only on definition changes.

Crash safety: update-journal.json before each phase; RecoverUpdates before
the boot sweep, ResumeInterruptedUpdates after the guards are wired.

Three unattended start paths ignored a hold and now honour it: the
drive-return gate (restart + boot recreate) and the nightly volume dump.
The nightly capture and Tier-2 run skip held apps so the restore point
survives. No automatic rollback — measured per-app; route back = restore.

Tests A–H across stacks/backup/api/web/cmd; six red-proofs seen to fail.

Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-13 11:41:31 +02:00
parent 1552716722
commit 0d402f711d
25 changed files with 2709 additions and 123 deletions
+65
View File
@@ -1,3 +1,68 @@
## v0.237.0 — the Update button takes a backup first, and tells the truth (2026-09-13, update arc slice 4 — R-448, R-443, R-439)
**MinAgent: 0.129.0** (unchanged)
**What it replaced.** `UpdateStack` advanced the pin, pulled, ran `up -d` and returned. No copy first,
no check of memory, disk, a running backup or a held app, and success the moment `up` returned —
measured in `SPIKE-app-update-2026-09-01` §4 as **HTTP 200 over an app that was already
crash-looping** (R-443). A held app could be updated at all (R-439). **`UpdateStack` is deleted**; its
only caller was the API.
**The guarded update** (`internal/stacks/update.go`), a job answering **202** at once, with phases
`checking → backing-up (only if stale) → safety-dump → pinning → pulling → starting → verifying →
done | failed` on `GET /api/stacks/{name}` (`updating`, `update_phase`, `update_phase_label`,
`update_error`, `hold_reason`):
1. **Cheap refusals first, each a 409 with a Hungarian sentence, before the intent is recorded:** held
(R-439 — `update` joins `start`/`restart` in the router's hold check), a backup/restore/app-data op
or quiesce holding it, a migration, already updating, deploying, memory (the deploy's own check,
extracted as `memoryVerdict`, releasing the app's current request), disk (**fixed 2 GB floor** on the
Docker data root — the image size is not known without a registry query).
2. **The precondition is the existing verified backup** (operator ruling 2026-09-02): an openable
Tier-2 recovery unit with a PROVEN copy date — `backup.Tier2UnitRestorePoint`, **extracted from the
backups page handler, not duplicated** (the page renders identically, pinned). No such copy →
refused. Copy older than `update.backup_max_age` (default `24h`) → `RunAppBackupNow` runs this app's
own legs (DB dump → volume dump → unit capture → Tier-2) first; it must succeed AND yield a fresh
restorable copy, or nothing moves.
3. **Safety dump before the pin moves** (`WriteUpdateSafetyDump` = R-361's `writeSafetyDump`).
4. **Pin before pull**; a **pull failure puts the pin and definition BACK** (nothing ran).
5. **Health, not the exit code:** the app's `.felhom.yml` check through the existing probe, or — with
none declared — every container running, none restarting, for 60 s; bounded by
`update.health_timeout` (default `5m`). Not healthy → **the app is stopped and HELD**
(`settings.RestoreHold`, new `reason: update_failed` + `copy_date`, same store and same gate as
R-379), the pin stays on the new version, and the page carries the hold sentence naming the backup
to restore from. A successful unit restore lifts an update hold (only that kind).
**Measured before it was designed, and the design changed because of it.** The copy's age is the last
SUCCESSFUL Tier-2 copy, not the unit manifest's `created_at`: on demo-hp bookstack's mirror held a
database dump from 2026-09-13T00:30Z under a manifest dated 2026-09-12T02:15:29Z, because a capture
rewrites the manifest only when the app's DEFINITION changes. Aged by the manifest, a quiet app would
be "stale" forever and "back up first" would not fix it.
**Crash safety is a journal** (`<data>/update-journal.json`, fsynced, written before each phase).
`RecoverUpdates` runs before the boot sweep: interrupted before the pin → dropped; while pinning or
pulling → pin put back; after `up` → the app is marked Updating (the boot sweep and the dead-app alarm
leave it alone) and `ResumeInterruptedUpdates` re-runs `up` + the health wait once the backup side is
wired, ending healthy or held.
**"A hold that only one path honours is not a hold" — three unattended paths did not honour one, and
now do:** the drive-return gate (`intermediary.go` restart + boot recreate), and the nightly volume
dump (`DumpAppVolumesSafe` ends in `StartStack`). The nightly capture and Tier-2 run also skip a held
app, so they cannot overwrite the restore point the hold text names.
**Deliberately NOT built:** putting the old version back automatically. Measured per-app
(`SPIKE-upgrade-test-2026-09-06`) and unpredictable; the route back is the restore. The multi-major
jump is not gated here (Slice 6); it fails health and is held honestly.
**Tests.** stacks: A–G (success only after health, stale → backup first, backup failure, no unit,
seven cheap refusals, pull failure pin-back, health failure hold, unsaved hold, three recovery
shapes), hold text on GetStacks, the exact phase labels. backup: proven copy time (incl. the measured
bookstack case), hold text, restore clears only update holds, busy, nightly legs skip held apps. api:
R-439, R-443 (202, never "completed"), no-backup 409 records no intent. web: backup row unchanged by the
extraction, drive-return skips a held app. cmd: guards wired, recovery before the boot sweep, resume
after the guards, boot gate + dead-app alarm know about updates. **Six companion red-proofs, each seen
to fail** — three first ran INERT or did not compile and were fixed before they counted (REPORT.md).
## v0.236.0 — "delete my data too" deletes the data, or says that it could not (2026-09-13, R-442)
**The defect.** A customer removes an app and ticks *also delete my data*. The box says it worked;