v0.263.0: a failed update puts the app back by itself (09 decision 15, R-637)
gates / gates (push) Successful in 26s

The guarded update gains a folder copy of the app's named volumes, taken
after the pull where the app stops anyway (decision 19, chosen by the
2026-09-23 bake-off). On a failed health check the box undoes: every copy
validated by its finished-marker first, volumes refilled, definition and pin
from the job's own pre-update copies, the old version checked with the OLD
.felhom.yml probe. It holds only if the undo fails, and the hold sentence
says so and what state the data is in. Bind-mounted folders are never
touched.

- R-637 built; R-638/R-640/R-641 do not arise with a folder copy; R-639
  (pre-update copies incl. .felhom.yml kept until the undo is over).
- journal phases copying/undoing with power-cut recovery.
- app.yaml last_update_undone + one line on the app page (hu/en).
- R-642: start/restart never answer "completed".
- Removal deletes kept undo copies.

MinAgent unchanged (0.131.0). Nine red-proofs in REPORT.md.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-23 11:12:49 +02:00
parent b9deec1907
commit 8fc2b4a1a9
26 changed files with 1326 additions and 69 deletions
+11 -6
View File
@@ -505,20 +505,25 @@ func fmtHoldTime(rfc3339 string) string {
// that the next restart button will quietly start again. The caller logs it at ERROR and keeps the
// failure on the page.
func (m *Manager) HoldAfterFailedUpdate(stackName string, at time.Time, copyDate time.Time, copyTier int) error {
return m.HoldAfterFailedUpdateHolding(stackName, at, copyDate, copyTier, "")
return m.HoldAfterFailedUpdateHolding(stackName, at, copyDate, copyTier, "", "")
}
// HoldAfterFailedUpdateHolding is HoldAfterFailedUpdate with the R-479 phrase for what the copy holds;
// "" records none (the tier-only sentence). The adapter in main.go computes the phrase with
// UpdateCopyHolds at hold time.
func (m *Manager) HoldAfterFailedUpdateHolding(stackName string, at time.Time, copyDate time.Time, copyTier int, copyHolds string) error {
//
// undoState (v0.263.0) is what a FAILED undo left the data as (stacks.UndoState*), "" when no undo was
// attempted. RestoreHoldFor puts it in front of the sentence, so the household reads that the box
// already tried to put the app back, and in what state that left the data.
func (m *Manager) HoldAfterFailedUpdateHolding(stackName string, at time.Time, copyDate time.Time, copyTier int, copyHolds, undoState string) error {
if m == nil || m.settings == nil {
return fmt.Errorf("no settings wired — the update hold for %s cannot be persisted", stackName)
}
h := settings.RestoreHold{
Stack: stackName,
At: at.UTC().Format(time.RFC3339),
Reason: settings.HoldReasonUpdateFailed,
Stack: stackName,
At: at.UTC().Format(time.RFC3339),
Reason: settings.HoldReasonUpdateFailed,
UndoState: undoState,
}
if !copyDate.IsZero() {
h.CopyDate = copyDate.UTC().Format(time.RFC3339)
@@ -528,7 +533,7 @@ func (m *Manager) HoldAfterFailedUpdateHolding(stackName string, at time.Time, c
if err := m.settings.SetRestoreHold(h); err != nil {
return fmt.Errorf("persisting the update hold for %s: %w", stackName, err)
}
m.logger.Printf("[WARN] [backup] %s is HELD STOPPED after a failed update (restore point: tier %d %q, %s; holds: %q)", stackName, h.CopyTier, UpdateTierLabel(h.CopyTier), h.CopyDate, h.CopyHolds)
m.logger.Printf("[WARN] [backup] %s is HELD STOPPED after a failed update (restore point: tier %d %q, %s; holds: %q; undo: %q)", stackName, h.CopyTier, UpdateTierLabel(h.CopyTier), h.CopyDate, h.CopyHolds, h.UndoState)
return nil
}