v0.237.0: the Update button takes a backup first, and tells the truth (update arc slice 4 — R-448, R-443, R-439)
gates / gates (push) Successful in 13s

POST /api/stacks/{name}/update is now a guarded job answering 202:
cheap refusals (hold — R-439, busy, migration, deploying, memory via the
deploy's own memoryVerdict, a fixed 2 GB disk floor, and no restorable
Tier-2 copy) → backup-first when the proven copy is older than
update.backup_max_age (24h) → safety dump BEFORE the pin moves → pin →
pull (failure puts the pin back) → up → health (.felhom.yml check or 60 s
settle, update.health_timeout 5m). Not healthy → the app is stopped and
HELD (RestoreHold reason update_failed, same store and gate as R-379) and
the page names the backup to restore from; the pin stays. Success is only
ever update_phase=done after health (R-443). UpdateStack is deleted.

The restorable-unit predicate is EXTRACTED to backup.Tier2UnitRestorePoint
and shared with the backups page (row pinned unchanged). The copy is aged
by the last successful Tier-2 copy, not the manifest created_at — measured
on demo-hp that created_at moves only on definition changes.

Crash safety: update-journal.json before each phase; RecoverUpdates before
the boot sweep, ResumeInterruptedUpdates after the guards are wired.

Three unattended start paths ignored a hold and now honour it: the
drive-return gate (restart + boot recreate) and the nightly volume dump.
The nightly capture and Tier-2 run skip held apps so the restore point
survives. No automatic rollback — measured per-app; route back = restore.

Tests A–H across stacks/backup/api/web/cmd; six red-proofs seen to fail.

Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-13 11:41:31 +02:00
parent 1552716722
commit 0d402f711d
25 changed files with 2709 additions and 123 deletions
+22
View File
@@ -219,6 +219,12 @@ func (s *Server) stopAppsOnPath(storagePath string) []string {
// restartStacks starts each named stack (the gate-stopped set on drive return). Best-effort per app.
func (s *Server) restartStacks(names []string) {
for _, name := range names {
// Slice 4 / R-379: the drive-return gate starts apps UNATTENDED, so it must honour a hold
// exactly as the customer's button does. Until v0.237.0 it did not — a held app whose drive
// blinked would have been started again. "A hold that only one path honours is not a hold."
if s.appHeld(name) {
continue
}
if err := s.stackMgr.StartStack(name); err != nil {
s.logger.Printf("[WARN] [gate] restart %s: %v", name, err)
}
@@ -454,6 +460,9 @@ func (s *Server) processGuestBootChange() {
}
recreate := func(bs bootStack) {
s.logger.Printf("[INFO] [gate] boot %s: live bind confirmed — recreating drive-backed app %s (state=%s) onto %s", resp.GuestBootID, bs.name, bs.state, bs.hdd)
if s.appHeld(bs.name) {
return
}
_ = s.stackMgr.StopStack(bs.name)
if serr := s.stackMgr.StartStack(bs.name); serr != nil {
s.logger.Printf("[WARN] [gate] boot recreate %s: %v", bs.name, serr)
@@ -694,3 +703,16 @@ func (s *Server) notifyDriveReturned(path string, isTarget map[string]bool) {
}
s.notifier.NotifyStorageReconnected(label)
}
// appHeld reports whether an app carries a hold (failed update or failed restore) and logs the skip.
// Nil-safe: no backup manager means no hold store, so nothing is held.
func (s *Server) appHeld(name string) bool {
if s.backupMgr == nil {
return false
}
held, _ := s.backupMgr.RestoreHoldFor(name)
if held {
s.logger.Printf("[WARN] [gate] NOT starting %s — the app is HELD (a failed update or restore); a person releases it", name)
}
return held
}