v0.237.0: the Update button takes a backup first, and tells the truth (update arc slice 4 — R-448, R-443, R-439)
gates / gates (push) Successful in 13s
gates / gates (push) Successful in 13s
POST /api/stacks/{name}/update is now a guarded job answering 202:
cheap refusals (hold — R-439, busy, migration, deploying, memory via the
deploy's own memoryVerdict, a fixed 2 GB disk floor, and no restorable
Tier-2 copy) → backup-first when the proven copy is older than
update.backup_max_age (24h) → safety dump BEFORE the pin moves → pin →
pull (failure puts the pin back) → up → health (.felhom.yml check or 60 s
settle, update.health_timeout 5m). Not healthy → the app is stopped and
HELD (RestoreHold reason update_failed, same store and gate as R-379) and
the page names the backup to restore from; the pin stays. Success is only
ever update_phase=done after health (R-443). UpdateStack is deleted.
The restorable-unit predicate is EXTRACTED to backup.Tier2UnitRestorePoint
and shared with the backups page (row pinned unchanged). The copy is aged
by the last successful Tier-2 copy, not the manifest created_at — measured
on demo-hp that created_at moves only on definition changes.
Crash safety: update-journal.json before each phase; RecoverUpdates before
the boot sweep, ResumeInterruptedUpdates after the guards are wired.
Three unattended start paths ignored a hold and now honour it: the
drive-return gate (restart + boot recreate) and the nightly volume dump.
The nightly capture and Tier-2 run skip held apps so the restore point
survives. No automatic rollback — measured per-app; route back = restore.
Tests A–H across stacks/backup/api/web/cmd; six red-proofs seen to fail.
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -1585,8 +1585,27 @@ type RestoreHold struct {
|
||||
ReplayError string `json:"replay_error,omitempty"` // what the restore hit
|
||||
RollbackErr string `json:"rollback_error,omitempty"` // what the rollback then hit
|
||||
SafetyDump string `json:"safety_dump,omitempty"` // basename of the undo copy that could not be applied
|
||||
|
||||
// Reason (update arc slice 4, v0.237.0) says WHICH operation put the hold in place. Empty means
|
||||
// HoldReasonRestoreFailed — every hold written before this field existed was a restore hold, so
|
||||
// the zero value keeps their meaning and no migration is needed.
|
||||
//
|
||||
// ONE STORAGE, ONE GATE, TWO REASONS — deliberately not a second map. Every start path already
|
||||
// consults GetRestoreHold (the customer's button, the boot sweep, the app-stop guard's Recover),
|
||||
// and "a hold that only one path honours is not a hold". A second map would need every one of
|
||||
// those paths found and changed again, and the one that got missed would be the next R-439.
|
||||
Reason string `json:"reason,omitempty"`
|
||||
// CopyDate is the RFC3339 time of the proven backup the customer is told they can restore from.
|
||||
// Only set for HoldReasonUpdateFailed.
|
||||
CopyDate string `json:"copy_date,omitempty"`
|
||||
}
|
||||
|
||||
// Hold reasons. See RestoreHold.Reason.
|
||||
const (
|
||||
HoldReasonRestoreFailed = "" // R-379/R-380: a restore AND its rollback failed
|
||||
HoldReasonUpdateFailed = "update_failed" // slice 4: the new version did not come up healthy
|
||||
)
|
||||
|
||||
// SetRestoreHold records a hold. Modelled on SetDisconnected: a condition, plus what it is holding.
|
||||
func (s *Settings) SetRestoreHold(h RestoreHold) error {
|
||||
s.mu.Lock()
|
||||
|
||||
Reference in New Issue
Block a user