v0.237.0: the Update button takes a backup first, and tells the truth (update arc slice 4 — R-448, R-443, R-439)
gates / gates (push) Successful in 13s

POST /api/stacks/{name}/update is now a guarded job answering 202:
cheap refusals (hold — R-439, busy, migration, deploying, memory via the
deploy's own memoryVerdict, a fixed 2 GB disk floor, and no restorable
Tier-2 copy) → backup-first when the proven copy is older than
update.backup_max_age (24h) → safety dump BEFORE the pin moves → pin →
pull (failure puts the pin back) → up → health (.felhom.yml check or 60 s
settle, update.health_timeout 5m). Not healthy → the app is stopped and
HELD (RestoreHold reason update_failed, same store and gate as R-379) and
the page names the backup to restore from; the pin stays. Success is only
ever update_phase=done after health (R-443). UpdateStack is deleted.

The restorable-unit predicate is EXTRACTED to backup.Tier2UnitRestorePoint
and shared with the backups page (row pinned unchanged). The copy is aged
by the last successful Tier-2 copy, not the manifest created_at — measured
on demo-hp that created_at moves only on definition changes.

Crash safety: update-journal.json before each phase; RecoverUpdates before
the boot sweep, ResumeInterruptedUpdates after the guards are wired.

Three unattended start paths ignored a hold and now honour it: the
drive-return gate (restart + boot recreate) and the nightly volume dump.
The nightly capture and Tier-2 run skip held apps so the restore point
survives. No automatic rollback — measured per-app; route back = restore.

Tests A–H across stacks/backup/api/web/cmd; six red-proofs seen to fail.

Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-13 11:41:31 +02:00
parent 1552716722
commit 0d402f711d
25 changed files with 2709 additions and 123 deletions
+11 -1
View File
@@ -7,7 +7,17 @@
>
> Ask Claude Code: "Please update CONTEXT.md with what we did today"
Last updated: 2026-09-13 (v0.236.0 — R-442: "delete my data too" deletes it or refuses)
Last updated: 2026-09-13 (v0.237.0 — update arc slice 4: the guarded update)
> **2026-09-13 — v0.237.0 (slice 4: R-448, R-443, R-439).** Update is a guarded 202 job:
> refusals (hold, busy, memory, disk, no restorable Tier-2 copy) → backup-first if the proven copy is
> older than `update.backup_max_age` (24h) → safety dump → pin → pull (failure: pin BACK) → up →
> health (`.felhom.yml` check or 60 s settle, `update.health_timeout` 5m; failure: stop + HOLD,
> `RestoreHold.Reason=update_failed`, pin stays). `UpdateStack` deleted. **The copy is aged by the last
> successful Tier-2 copy, NOT the manifest `created_at`** — measured on demo-hp that `created_at` moves
> only on definition changes. Journal `update-journal.json`; `RecoverUpdates` before the boot sweep.
> Three unattended paths that ignored a hold now honour it (drive-return gate ×2, nightly volume dump);
> capture and Tier-2 skip held apps. **No automatic rollback** — measured per-app; route back = restore.
> **2026-09-13 — v0.236.0 (R-442).** Removal resolves the drive from the app's OWN `app.yaml`
> `HDD_PATH` (the `07` ~L437 rule), never the global `cfg.Paths.HDDPath` (set on no box); a data