v0.237.0: the Update button takes a backup first, and tells the truth (update arc slice 4 — R-448, R-443, R-439)
gates / gates (push) Successful in 13s

POST /api/stacks/{name}/update is now a guarded job answering 202:
cheap refusals (hold — R-439, busy, migration, deploying, memory via the
deploy's own memoryVerdict, a fixed 2 GB disk floor, and no restorable
Tier-2 copy) → backup-first when the proven copy is older than
update.backup_max_age (24h) → safety dump BEFORE the pin moves → pin →
pull (failure puts the pin back) → up → health (.felhom.yml check or 60 s
settle, update.health_timeout 5m). Not healthy → the app is stopped and
HELD (RestoreHold reason update_failed, same store and gate as R-379) and
the page names the backup to restore from; the pin stays. Success is only
ever update_phase=done after health (R-443). UpdateStack is deleted.

The restorable-unit predicate is EXTRACTED to backup.Tier2UnitRestorePoint
and shared with the backups page (row pinned unchanged). The copy is aged
by the last successful Tier-2 copy, not the manifest created_at — measured
on demo-hp that created_at moves only on definition changes.

Crash safety: update-journal.json before each phase; RecoverUpdates before
the boot sweep, ResumeInterruptedUpdates after the guards are wired.

Three unattended start paths ignored a hold and now honour it: the
drive-return gate (restart + boot recreate) and the nightly volume dump.
The nightly capture and Tier-2 run skip held apps so the restore point
survives. No automatic rollback — measured per-app; route back = restore.

Tests A–H across stacks/backup/api/web/cmd; six red-proofs seen to fail.

Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-13 11:41:31 +02:00
parent 1552716722
commit 0d402f711d
25 changed files with 2709 additions and 123 deletions
+802
View File
@@ -0,0 +1,802 @@
package stacks
import (
"context"
"encoding/json"
"fmt"
"os"
"path/filepath"
"time"
"gitea.dooplex.hu/admin/felhom-controller/internal/system"
)
// ── The guarded update (update arc slice 4, controller v0.237.0) ─────────────────────────────────
//
// WHAT IT REPLACED. `UpdateStack` advanced the pin, pulled, ran `up -d` and returned — no copy first,
// no check of memory, disk, a running backup or a held app, and a success the moment `up` returned.
// SPIKE-app-update-2026-09-01 §4 measured that as HTTP 200 over an app that was already crash-looping
// (R-443), and R-439 is that a held app could be updated at all.
//
// THE SEQUENCE, and the order is the point:
//
// checking → backing-up (only if the proven copy is too old) → safety-dump → pinning → pulling
// → starting → verifying → done | failed
//
// 1. The PRECONDITION is the app's existing verified backup (operator ruling 2026-09-02,
// 09-update-architecture §3 decision 1): an openable Tier-2 unit with a PROVEN copy date. The
// same predicate that permits the destructive „Teljes visszaállítás" permits the update — the
// route back IS that restore, so an update without it has no route back.
// 2. The safety dump is taken BEFORE the pin moves: it is "the state the customer was in a minute ago",
// and a minute later the migration may have run.
// 3. The pin moves BEFORE the pull (v0.235.0's reason: pull and up act on the file on disk).
// 4. A PULL failure puts the pin BACK — nothing ran, so reverting is safe and honest (Scenario E).
// 5. A HEALTH failure leaves the pin where it is — the new version's migration may have run, and a
// pin claiming the old version would be a record of something untrue (Scenario F). The app is
// HELD STOPPED and the customer is told which backup it can be restored from.
//
// WHAT IT DELIBERATELY DOES NOT DO: put the old version back by itself. SPIKE-upgrade-test-2026-09-06
// measured that whether the old image starts on migrated data depends on the app (PrivateBin yes,
// Docmost and Nextcloud no) and cannot be predicted. The route back is the restore.
//
// CRASH SAFETY IS A JOURNAL, NOT A DEFER. A SIGKILL runs no deferred function (Campaign 8 fault 10),
// so every phase is written to `update-journal.json` BEFORE it starts, and RecoverUpdates reads it at
// the next startup (Scenario G) — the AppStopGuard pattern.
// Update phases, as recorded in the journal and served as Stack.UpdatePhase.
const (
UpdatePhaseChecking = "checking"
UpdatePhaseBackingUp = "backing-up"
UpdatePhaseSafetyDump = "safety-dump"
UpdatePhasePinning = "pinning"
UpdatePhasePulling = "pulling"
UpdatePhaseStarting = "starting"
UpdatePhaseVerifying = "verifying"
UpdatePhaseDone = "done"
UpdatePhaseFailed = "failed"
)
// updatePhaseLabels are the customer labels (slice 4 Part 4, exact). `pinning` has no row in the
// specification — it is instantaneous and is the first step of fetching the new version, so it shares
// the pull's label rather than inventing a sentence nobody would read.
var updatePhaseLabels = map[string]string{
UpdatePhaseChecking: "Ellenőrzés…",
UpdatePhaseBackingUp: "Biztonsági mentés készül a frissítés előtt…",
UpdatePhaseSafetyDump: "Adatbázis pillanatkép…",
UpdatePhasePinning: "Új verzió letöltése…",
UpdatePhasePulling: "Új verzió letöltése…",
UpdatePhaseStarting: "Indítás az új verzióval…",
UpdatePhaseVerifying: "Működés ellenőrzése…",
UpdatePhaseDone: "Frissítve",
UpdatePhaseFailed: "A frissítés nem sikerült",
}
// UpdatePhaseLabel returns the customer label for a phase, "" for an unknown one.
func UpdatePhaseLabel(phase string) string { return updatePhaseLabels[phase] }
// Customer sentences. Named so tests compare against the constant, never a retyped literal (R-364).
const (
MsgUpdateNoGuards = "A frissítés nem indítható: a frissítés előtti biztonsági ellenőrzés nem érhető el ezen a szerveren."
MsgUpdateNotDeployed = "Az alkalmazás nincs telepítve, ezért nem frissíthető."
MsgUpdateDeployingFmt = "A(z) %s telepítése még folyamatban van — a frissítés utána indítható."
MsgUpdateAlreadyFmt = "A(z) %s frissítése már folyamatban van."
MsgUpdateBusy = "A frissítés most nem indítható: mentés/visszaállítás folyamatban. Próbáld újra, ha befejeződött."
MsgUpdateMigrating = "A frissítés most nem indítható: adatáthelyezés folyamatban."
MsgUpdateNoBackupFmt = "A(z) %s nem frissíthető, mert nincs olyan biztonsági mentése, amelyből vissza lehetne állítani. Kapcsold be a 2. mentést az alkalmazás mentési beállításainál a Mentések oldalon, és várd meg az első sikeres másolatot — utána a frissítés elindítható."
MsgUpdateDiskFmt = "Nincs elég szabad hely a frissítéshez: %.1f GB szabad, az új verzió letöltéséhez legalább %.0f GB szükséges."
MsgUpdateBackupFailFmt = "A frissítés nem indult el, mert a frissítés előtti biztonsági mentés nem sikerült: %v. Az alkalmazás változatlanul fut tovább."
MsgUpdateBackupNoUnit = "A frissítés nem indult el: a frissítés előtti mentés lefutott, de nem jött létre friss, visszaállítható másolat. Az alkalmazás változatlanul fut tovább."
MsgUpdateDumpFailFmt = "A frissítés nem indult el, mert az adatbázis pillanatkép nem készült el: %v. Az alkalmazás változatlanul fut tovább."
MsgUpdatePinFailed = "A frissítés nem indult el: az új verzió leírása nem olvasható be. Az alkalmazás változatlanul fut tovább."
MsgUpdateJournalFailed = "A frissítés nem indult el: a frissítés naplója nem menthető. Az alkalmazás változatlanul fut tovább."
MsgUpdatePullFailed = "Az új verzió letöltése nem sikerült, ezért a frissítés elmaradt. Az alkalmazás a korábbi verzióval fut tovább."
MsgUpdateInterrupted = "A frissítés megszakadt, mert a vezérlő újraindult, mielőtt az új verzió elindult volna. Az alkalmazás a korábbi verzióval fut tovább."
MsgUpdateHoldUnsaved = "A frissítés nem sikerült, az alkalmazás le lett állítva, de a leállítás rögzítése nem sikerült. Ne indítsd újra — vedd fel velünk a kapcsolatot."
)
// updateDiskFloorGiB is the free space the Docker data root must have before a pull. A FIXED FLOOR,
// stated as such: the new image set's size is not known without a registry query (the catalog
// records tags, not sizes, and §8.1 of 09 already declines registry calls on the customer box), so
// the rule is "not less than 2 GB", not "enough for these images".
const updateDiskFloorGiB = 2.0
// updateSettleWindow is the rule for an app with no .felhom.yml health check: every container
// running, none restarting, for this long.
const updateSettleWindow = 60 * time.Second
// updatePollEvery is how often the health wait re-reads the stack.
const updatePollEvery = 5 * time.Second
// UpdateRestorePoint is the precondition answer, reduced to what the update needs.
type UpdateRestorePoint struct {
Restorable bool // an openable recovery unit exists in the Tier-2 copy
Proven bool // a copy actually succeeded (never an attempt clock)
ProvenAt time.Time // when the data in that copy was last proven copied
}
// UpdateGuards is everything the update needs from the backup side. The stacks package cannot import
// backup, so cmd/controller/main.go wires an adapter (TestSlice4_UpdateGuardsAreWiredAtStartup).
type UpdateGuards interface {
HoldFor(name string) (bool, string)
Busy(name string) (bool, string)
RestorePoint(name string) (UpdateRestorePoint, error)
BackupNow(ctx context.Context, name string) error
SafetyDump(ctx context.Context, name string) ([]string, error)
HoldAfterFailedUpdate(name string, at, provenCopyAt time.Time) error
}
// SetUpdateGuards wires the backup side. INIT-ONLY. Unwired, every update is refused (fail closed):
// an update that cannot see the backup cannot promise a route back.
func (m *Manager) SetUpdateGuards(g UpdateGuards) {
m.mu.Lock()
m.updateGuards = g
m.mu.Unlock()
}
func (m *Manager) guards() UpdateGuards {
m.mu.RLock()
defer m.mu.RUnlock()
return m.updateGuards
}
func fillHoldReason(g UpdateGuards, st *Stack) {
if g == nil || st == nil || !st.Deployed {
return
}
if held, why := g.HoldFor(st.Name); held {
st.HoldReason = why
}
}
// UpdateRefusal is a refusal taken before anything moved. Reason is a stable key for logs and tests;
// Message is the customer sentence.
type UpdateRefusal struct {
Reason string
Message string
}
func (r *UpdateRefusal) Error() string { return r.Message }
func (m *Manager) now() time.Time {
if m.updateNowFn != nil {
return m.updateNowFn()
}
return time.Now()
}
func (m *Manager) refuseUpdate(name, reason, msg, detail string) *UpdateRefusal {
m.logger.Printf("[ERROR] [stacks] update %s REFUSED (%s): %s", name, reason, detail)
return &UpdateRefusal{Reason: reason, Message: msg}
}
// UpdatePreflight runs every CHEAP refusal (slice 4 Part 1 + the precondition's existence), in order,
// and returns the first. Nothing is moved and nothing is recorded by it. The router calls it before
// recording the customer's intent, so an update that was never going to happen records nothing.
func (m *Manager) UpdatePreflight(name string) *UpdateRefusal {
st, ok := m.GetStack(name)
if !ok {
return m.refuseUpdate(name, "not_found", fmt.Sprintf("stack %q not found", name), "no such stack")
}
if !st.Deployed {
return m.refuseUpdate(name, "not_deployed", MsgUpdateNotDeployed, "not deployed")
}
g := m.guards()
if g == nil {
return m.refuseUpdate(name, "guards_unwired", MsgUpdateNoGuards, "no UpdateGuards wired — fail closed")
}
if st.Deploying {
return m.refuseUpdate(name, "deploying", fmt.Sprintf(MsgUpdateDeployingFmt, name), "a deploy is in progress")
}
if st.Updating {
return m.refuseUpdate(name, "updating", fmt.Sprintf(MsgUpdateAlreadyFmt, name), "an update is already in progress")
}
if held, why := g.HoldFor(name); held {
return m.refuseUpdate(name, "held", why, "the app is held")
}
if busy, why := g.Busy(name); busy {
return m.refuseUpdate(name, "busy", MsgUpdateBusy, why)
}
if m.IsMigrating() {
return m.refuseUpdate(name, "migrating", MsgUpdateMigrating, "a data migration is running")
}
rp, err := g.RestorePoint(name)
if err != nil || !rp.Restorable || !rp.Proven {
return m.refuseUpdate(name, "no_backup", fmt.Sprintf(MsgUpdateNoBackupFmt, name),
fmt.Sprintf("no restorable proven Tier-2 unit (restorable=%v proven=%v err=%v)", rp.Restorable, rp.Proven, err))
}
if ref := m.updateMemoryRefusal(name, st); ref != nil {
return ref
}
free, known := m.updateDiskFree()
switch {
case !known:
m.logger.Printf("[WARN] [stacks] update %s: free space on the Docker data root is unreadable — proceeding without the %.0f GB floor", name, updateDiskFloorGiB)
case free < updateDiskFloorGiB:
return m.refuseUpdate(name, "disk", fmt.Sprintf(MsgUpdateDiskFmt, free, updateDiskFloorGiB),
fmt.Sprintf("%.2f GiB free on the Docker data root, floor %.0f GiB (fixed floor — image size unknown)", free, updateDiskFloorGiB))
}
return nil
}
// updateMemoryRefusal applies the deploy's memory check to the NEW template's request, releasing the
// app's CURRENT request first (an update replaces it). An unknown new request proceeds with a WARN.
func (m *Manager) updateMemoryRefusal(name string, st *Stack) *UpdateRefusal {
catPath := m.CatalogTemplatePath(name, ".felhom.yml")
if _, err := os.Stat(catPath); err != nil {
m.logger.Printf("[WARN] [stacks] update %s: the new template's memory request is unknown (%v) — proceeding without the memory check", name, err)
return nil
}
newMeta := LoadMetadata(filepath.Dir(catPath))
newReq, newLim := ParseMemoryMB(newMeta.Resources.MemRequest), ParseMemoryMB(newMeta.Resources.MemLimit)
if newReq == 0 {
m.logger.Printf("[WARN] [stacks] update %s: the new template declares no memory request — proceeding without the memory check", name)
return nil
}
oldReq, oldLim := ParseMemoryMB(st.Meta.Resources.MemRequest), ParseMemoryMB(st.Meta.Resources.MemLimit)
verdict := m.updateMemoryFn
if verdict == nil {
verdict = m.memoryVerdict
}
if refusal, _ := verdict(newReq, newLim, oldReq, oldLim); refusal != "" {
return m.refuseUpdate(name, "memory", refusal, fmt.Sprintf("new_req=%dMB replacing %dMB does not fit", newReq, oldReq))
}
return nil
}
func (m *Manager) updateDiskFree() (float64, bool) {
if m.updateDiskFreeFn != nil {
return m.updateDiskFreeFn()
}
du := system.GetDiskUsage(system.DockerVolumePath)
if du == nil {
return 0, false
}
return du.AvailGB, true
}
// StartGuardedUpdate re-checks the cheap refusals, claims the Updating flag atomically and launches the
// job. It returns as soon as the job has STARTED — the result arrives on GET /api/stacks/{name}.
func (m *Manager) StartGuardedUpdate(name string) error {
if ref := m.UpdatePreflight(name); ref != nil {
return ref
}
m.mu.Lock()
s, ok := m.stacks[name]
if !ok {
m.mu.Unlock()
return &UpdateRefusal{Reason: "not_found", Message: fmt.Sprintf("stack %q not found", name)}
}
// A second press between the preflight and here is the race this lock closes.
if s.Updating || s.Deploying {
m.mu.Unlock()
return m.refuseUpdate(name, "updating", fmt.Sprintf(MsgUpdateAlreadyFmt, name), "lost the race for the Updating flag")
}
s.Updating, s.UpdateError = true, ""
s.UpdatePhase, s.UpdatePhaseLabel = UpdatePhaseChecking, UpdatePhaseLabel(UpdatePhaseChecking)
m.mu.Unlock()
m.logger.Printf("[INFO] [stacks] update %s: accepted — guarded update started", name)
go m.runGuardedUpdate(context.Background(), name)
return nil
}
// IsUpdating reports whether a guarded update is in progress for the app.
func (m *Manager) IsUpdating(name string) bool {
m.mu.RLock()
defer m.mu.RUnlock()
s, ok := m.stacks[name]
return ok && s.Updating
}
// UpdatingStacks is the set of apps an update is currently moving — for the dead-app alarm, which
// must not count an app the update itself is recreating (R-330's class, a third mechanism).
func (m *Manager) UpdatingStacks() map[string]bool {
m.mu.RLock()
defer m.mu.RUnlock()
var out map[string]bool
for name, s := range m.stacks {
if s.Updating {
if out == nil {
out = map[string]bool{}
}
out[name] = true
}
}
return out
}
func (m *Manager) setUpdatePhase(name, phase string) {
m.mu.Lock()
if s, ok := m.stacks[name]; ok {
s.UpdatePhase, s.UpdatePhaseLabel = phase, UpdatePhaseLabel(phase)
}
m.mu.Unlock()
}
// finishUpdate is the ONE place Updating goes false. msg is the customer sentence on failure.
func (m *Manager) finishUpdate(name, phase, msg string) {
m.mu.Lock()
if s, ok := m.stacks[name]; ok {
s.Updating = false
s.UpdatePhase, s.UpdatePhaseLabel = phase, UpdatePhaseLabel(phase)
s.UpdateError = msg
}
m.mu.Unlock()
}
func (m *Manager) updateCompose(dir string, env []string, args ...string) (string, error) {
if m.updateComposeFn != nil {
return m.updateComposeFn(dir, env, args...)
}
return m.composeExecCustomEnv(dir, env, args...)
}
func (m *Manager) updateHealth(ctx context.Context, name string, timeout time.Duration) (bool, string) {
if m.updateHealthFn != nil {
return m.updateHealthFn(ctx, name, timeout)
}
return m.waitUpdateHealthy(ctx, name, timeout)
}
func (m *Manager) healthTimeout() time.Duration {
if m.cfg == nil {
return 5 * time.Minute
}
return m.cfg.Update.HealthTimeoutDuration()
}
func (m *Manager) backupMaxAge() time.Duration {
if m.cfg == nil {
return 24 * time.Hour
}
return m.cfg.Update.BackupMaxAgeDuration()
}
// pre-update copies, kept in the stack dir so they travel with it. Neither name is one the syncer
// copies (it copies exactly docker-compose.yml and .felhom.yml).
const (
preUpdateComposeFile = "pre-update-compose.yml"
preUpdateAppliedFile = "pre-update-applied.yml"
)
func (m *Manager) runGuardedUpdate(ctx context.Context, name string) {
start := m.now()
st, ok := m.GetStack(name)
if !ok {
m.finishUpdate(name, UpdatePhaseFailed, fmt.Sprintf("stack %q not found", name))
return
}
dir := filepath.Dir(st.ComposePath)
g := m.guards()
entry := updateJournalEntry{StartedAt: start}
fail := func(msg, detail string) {
m.logger.Printf("[ERROR] [stacks] update %s FAILED in phase %s after %s — nothing was moved: %s", name, entry.Phase, m.now().Sub(start).Round(time.Millisecond), detail)
m.clearJournal(name)
m.finishUpdate(name, UpdatePhaseFailed, msg)
}
if !m.enterUpdatePhase(name, &entry, UpdatePhaseChecking) {
m.finishUpdate(name, UpdatePhaseFailed, MsgUpdateJournalFailed)
return
}
if g == nil {
fail(MsgUpdateNoGuards, "no UpdateGuards wired")
return
}
rp, err := g.RestorePoint(name)
if err != nil || !rp.Restorable || !rp.Proven {
fail(fmt.Sprintf(MsgUpdateNoBackupFmt, name), fmt.Sprintf("precondition vanished: restorable=%v proven=%v err=%v", rp.Restorable, rp.Proven, err))
return
}
maxAge := m.backupMaxAge()
if age := start.Sub(rp.ProvenAt); age > maxAge {
m.logger.Printf("[INFO] [stacks] update %s: the proven copy is %s old (limit %s) — backing up first", name, age.Round(time.Minute), maxAge)
if !m.enterUpdatePhase(name, &entry, UpdatePhaseBackingUp) {
fail(MsgUpdateJournalFailed, "journal write failed")
return
}
if err := g.BackupNow(ctx, name); err != nil {
fail(fmt.Sprintf(MsgUpdateBackupFailFmt, err), "pre-update backup: "+err.Error())
return
}
rp, err = g.RestorePoint(name)
if err != nil || !rp.Restorable || !rp.Proven || m.now().Sub(rp.ProvenAt) > maxAge {
fail(MsgUpdateBackupNoUnit, fmt.Sprintf("after the backup: restorable=%v proven=%v at=%s err=%v", rp.Restorable, rp.Proven, rp.ProvenAt.Format(time.RFC3339), err))
return
}
} else {
m.logger.Printf("[INFO] [stacks] update %s: precondition met — proven copy from %s (%s old, limit %s)", name, rp.ProvenAt.UTC().Format(time.RFC3339), age.Round(time.Minute), maxAge)
}
entry.ProvenCopyAt = rp.ProvenAt.UTC().Format(time.RFC3339)
// SAFETY DUMP BEFORE THE PIN MOVES — "a minute ago", before any migration can have run.
if !m.enterUpdatePhase(name, &entry, UpdatePhaseSafetyDump) {
fail(MsgUpdateJournalFailed, "journal write failed")
return
}
paths, err := g.SafetyDump(ctx, name)
if err != nil {
fail(fmt.Sprintf(MsgUpdateDumpFailFmt, err), "safety dump: "+err.Error())
return
}
m.logger.Printf("[INFO] [stacks] update %s: safety dump done (%d file(s)) %v", name, len(paths), paths)
// PINNING — the previous definition is copied aside and journaled BEFORE the pin moves, so a crash
// at any later instant can put it back (Scenario G).
prevLive, err := os.ReadFile(st.ComposePath)
if err != nil {
fail(MsgUpdatePinFailed, "reading the live compose file: "+err.Error())
return
}
if err := os.WriteFile(filepath.Join(dir, preUpdateComposeFile), prevLive, 0o644); err != nil {
fail(MsgUpdateJournalFailed, "saving the pre-update compose copy: "+err.Error())
return
}
entry.PrevCompose = filepath.Join(dir, preUpdateComposeFile)
if applied, aerr := LoadAppliedDefinition(dir); aerr == nil {
if err := os.WriteFile(filepath.Join(dir, preUpdateAppliedFile), applied, 0o644); err == nil {
entry.PrevApplied = filepath.Join(dir, preUpdateAppliedFile)
}
}
if cfg := LoadAppConfig(dir); cfg != nil && len(cfg.PinnedImages) > 0 {
entry.PrevPin = map[string]string{}
for k, v := range cfg.PinnedImages {
entry.PrevPin[k] = v
}
}
if !m.enterUpdatePhase(name, &entry, UpdatePhasePinning) {
m.removePreUpdateCopies(dir)
fail(MsgUpdateJournalFailed, "journal write failed")
return
}
if err := m.advancePinToCatalog(name, dir); err != nil {
m.pinBack(name, dir, entry)
fail(MsgUpdatePinFailed, "advancing the pin: "+err.Error())
return
}
env := m.stackEnv(dir)
if !m.enterUpdatePhase(name, &entry, UpdatePhasePulling) {
m.pinBack(name, dir, entry)
fail(MsgUpdateJournalFailed, "journal write failed")
return
}
if _, err := m.updateCompose(dir, env, "pull"); err != nil {
// Scenario E: NOTHING RAN. The containers are the old ones and still running, so the honest
// state is the old pin and the old file — put both back.
m.pinBack(name, dir, entry)
m.logger.Printf("[ERROR] [stacks] update %s: pull failed — pin and definition PUT BACK; the app was not touched. Docker said: %v", name, err)
fail(MsgUpdatePullFailed, "pull failed: "+err.Error())
return
}
if !m.enterUpdatePhase(name, &entry, UpdatePhaseStarting) {
m.failAndHold(ctx, name, dir, env, rp.ProvenAt, "journal write failed before up")
return
}
if _, err := m.updateCompose(dir, env, "up", "-d", "--remove-orphans"); err != nil {
// Containers may already have been recreated on the new image — something may have run.
m.failAndHold(ctx, name, dir, env, rp.ProvenAt, "compose up failed: "+err.Error())
return
}
m.verifyAndConclude(ctx, name, dir, env, rp.ProvenAt, start, &entry)
}
// verifyAndConclude is the TRUTH half (R-443): success is declared only after the app's health is
// known, and a failure holds the app.
func (m *Manager) verifyAndConclude(ctx context.Context, name, dir string, env []string, provenAt, start time.Time, entry *updateJournalEntry) {
if !m.enterUpdatePhase(name, entry, UpdatePhaseVerifying) {
m.logger.Printf("[ERROR] [stacks] update %s: could not journal the verifying phase — verifying anyway", name)
}
timeout := m.healthTimeout()
waitStart := m.now()
healthy, detail := m.updateHealth(ctx, name, timeout)
if !healthy {
m.failAndHold(ctx, name, dir, env, provenAt, "not healthy: "+detail)
return
}
m.logger.Printf("[INFO] [stacks] update %s: healthy after %s (%s)", name, m.now().Sub(waitStart).Round(time.Second), detail)
m.recordInstalledImages(name, dir, env)
_ = m.RefreshStatus()
m.clearJournal(name)
m.removePreUpdateCopies(dir)
m.finishUpdate(name, UpdatePhaseDone, "")
m.logger.Printf("[INFO] [stacks] update %s: DONE in %s", name, m.now().Sub(start).Round(time.Second))
}
// failAndHold is Scenario F: stop the app, record the hold, tell the customer the route back.
func (m *Manager) failAndHold(ctx context.Context, name, dir string, env []string, provenAt time.Time, why string) {
m.logger.Printf("[ERROR] [stacks] update %s FAILED after the new version was started: %s — stopping and HOLDING the app; the pin stays on the new version (its migration may have run)", name, why)
if _, err := m.updateCompose(dir, env, "down"); err != nil {
m.logger.Printf("[ERROR] [stacks] update %s: stopping the failed app also failed: %v", name, err)
}
msg := MsgUpdateHoldUnsaved
if g := m.guards(); g == nil {
m.logger.Printf("[ERROR] [stacks] update %s: no UpdateGuards — the hold CANNOT be recorded", name)
} else if err := g.HoldAfterFailedUpdate(name, m.now(), provenAt); err != nil {
m.logger.Printf("[ERROR] [stacks] update %s: %v", name, err)
} else if _, why := g.HoldFor(name); why != "" {
msg = why
}
_ = m.RefreshStatus()
m.clearJournal(name)
m.removePreUpdateCopies(dir)
m.finishUpdate(name, UpdatePhaseFailed, msg)
}
// pinBack restores the pin, the stored definition and the live file from the journaled copies.
func (m *Manager) pinBack(name, dir string, entry updateJournalEntry) {
prevLive, lerr := os.ReadFile(entry.PrevCompose)
if lerr != nil {
m.logger.Printf("[ERROR] [stacks] update %s: cannot read the pre-update compose copy (%v) — the definition could NOT be put back", name, lerr)
}
if len(entry.PrevPin) > 0 {
applied := prevLive
if entry.PrevApplied != "" {
if b, err := os.ReadFile(entry.PrevApplied); err == nil {
applied = b
}
}
if err := m.SetPin(name, dir, entry.PrevPin, applied); err != nil {
m.logger.Printf("[ERROR] [stacks] update %s: putting the pin back failed: %v", name, err)
}
}
if lerr == nil {
if err := os.WriteFile(ComposePathIn(dir), prevLive, 0o644); err != nil {
m.logger.Printf("[ERROR] [stacks] update %s: re-rendering the previous definition failed: %v", name, err)
}
}
m.removePreUpdateCopies(dir)
m.logger.Printf("[INFO] [stacks] update %s: pin and definition PUT BACK to the pre-update version (%s)", name, summarisePin(entry.PrevPin))
}
func (m *Manager) removePreUpdateCopies(dir string) {
_ = os.Remove(filepath.Join(dir, preUpdateComposeFile))
_ = os.Remove(filepath.Join(dir, preUpdateAppliedFile))
}
// waitUpdateHealthy is the production health wait: the app's own .felhom.yml health check through the
// existing probe, or — for an app with none — every container running and none restarting for
// updateSettleWindow. NEVER the compose exit code, and never logPostStartStatus's delayed log line.
func (m *Manager) waitUpdateHealthy(ctx context.Context, name string, timeout time.Duration) (bool, string) {
deadline := m.now().Add(timeout)
var runningSince time.Time
last := "no observation yet"
for {
_ = m.RefreshStatus()
st, ok := m.GetStack(name)
switch {
case !ok:
last = "stack vanished"
runningSince = time.Time{}
case st.State == StateRunning:
if hc := st.Meta.HealthCheck; hc != nil && len(hc.Checks) > 0 {
if c := findProbeContainer(name, st.Containers); c != "" {
res := m.runChecks(probeTarget{stackName: name, containerName: c, checks: hc.Checks})
m.mu.Lock()
if s, ok := m.stacks[name]; ok {
s.HealthProbe = res
}
m.mu.Unlock()
if res.Healthy {
return true, "the app's health check passed"
}
last = "health check failing"
} else {
last = "no probe container"
}
} else {
if runningSince.IsZero() {
runningSince = m.now()
}
if m.now().Sub(runningSince) >= updateSettleWindow {
return true, fmt.Sprintf("all containers running, none restarting, for %s (no health check declared)", updateSettleWindow)
}
last = "running, settling"
}
default:
runningSince = time.Time{}
last = "state " + string(st.State)
}
if !m.now().Before(deadline) {
return false, fmt.Sprintf("not healthy within %s (last: %s)", timeout, last)
}
select {
case <-ctx.Done():
return false, "cancelled: " + ctx.Err().Error()
case <-time.After(updatePollEvery):
}
}
}
// ── the journal ──────────────────────────────────────────────────────────────────────────────────
type updateJournalEntry struct {
Phase string `json:"phase"`
StartedAt time.Time `json:"started_at"`
PrevPin map[string]string `json:"prev_pin,omitempty"`
PrevCompose string `json:"prev_compose,omitempty"`
PrevApplied string `json:"prev_applied,omitempty"`
ProvenCopyAt string `json:"proven_copy_at,omitempty"`
}
type updateJournal struct {
Updates map[string]updateJournalEntry `json:"updates"`
}
func (m *Manager) updateJournalPath() string {
return filepath.Join(m.cfg.Paths.DataDir, "update-journal.json")
}
func (m *Manager) readUpdateJournal() updateJournal {
j := updateJournal{Updates: map[string]updateJournalEntry{}}
data, err := os.ReadFile(m.updateJournalPath())
if err != nil {
return j
}
if err := json.Unmarshal(data, &j); err != nil {
m.logger.Printf("[WARN] [stacks] update journal at %s is corrupt (%v) — quarantining", m.updateJournalPath(), err)
_ = os.Rename(m.updateJournalPath(), fmt.Sprintf("%s.corrupt-%d", m.updateJournalPath(), time.Now().Unix()))
return updateJournal{Updates: map[string]updateJournalEntry{}}
}
if j.Updates == nil {
j.Updates = map[string]updateJournalEntry{}
}
return j
}
// writeUpdateJournal is atomic and fsynced (the AppStopGuard shape): the point is surviving a power cut.
func (m *Manager) writeUpdateJournal(j updateJournal) error {
p := m.updateJournalPath()
if len(j.Updates) == 0 {
if err := os.Remove(p); err != nil && !os.IsNotExist(err) {
return err
}
return nil
}
data, err := json.MarshalIndent(j, "", " ")
if err != nil {
return err
}
if err := os.MkdirAll(filepath.Dir(p), 0o755); err != nil {
return err
}
tmp := p + ".tmp"
f, err := os.OpenFile(tmp, os.O_WRONLY|os.O_CREATE|os.O_TRUNC, 0o600)
if err != nil {
return err
}
if _, err := f.Write(data); err != nil {
f.Close()
os.Remove(tmp)
return err
}
if err := f.Sync(); err != nil {
f.Close()
os.Remove(tmp)
return err
}
if err := f.Close(); err != nil {
os.Remove(tmp)
return err
}
return os.Rename(tmp, p)
}
// enterUpdatePhase journals the phase BEFORE it starts and mirrors it for the UI. False means the
// journal could not be written — the caller must not perform a mutation it could not record.
func (m *Manager) enterUpdatePhase(name string, entry *updateJournalEntry, phase string) bool {
entry.Phase = phase
m.updateJournalMu.Lock()
j := m.readUpdateJournal()
j.Updates[name] = *entry
err := m.writeUpdateJournal(j)
m.updateJournalMu.Unlock()
m.setUpdatePhase(name, phase)
if err != nil {
m.logger.Printf("[ERROR] [stacks] update %s: journal write for phase %s failed: %v", name, phase, err)
return false
}
m.logger.Printf("[INFO] [stacks] update %s: phase %s", name, phase)
return true
}
func (m *Manager) clearJournal(name string) {
m.updateJournalMu.Lock()
defer m.updateJournalMu.Unlock()
j := m.readUpdateJournal()
delete(j.Updates, name)
if err := m.writeUpdateJournal(j); err != nil {
m.logger.Printf("[ERROR] [stacks] update %s: clearing the journal entry failed: %v", name, err)
}
}
// RecoverUpdates reads the journal at startup (Scenario G). Call it BEFORE the boot reconciler.
//
// - interrupted BEFORE the pin moved (checking, backing-up, safety-dump): nothing moved — the entry is
// dropped and the app carries the "interrupted" sentence;
// - interrupted while pinning or pulling: nothing RAN — the pin and definition are put back (as E);
// - interrupted while starting or verifying: something may have run — the app is marked Updating
// (so the boot sweep and the dead-app alarm leave it alone) and queued for ResumeInterruptedUpdates,
// which re-runs `up -d` and the health wait, ending in A or F.
//
// It needs no backup wiring, because nothing here holds an app — that is left to the resumed job.
func (m *Manager) RecoverUpdates() []string {
m.updateJournalMu.Lock()
j := m.readUpdateJournal()
m.updateJournalMu.Unlock()
if len(j.Updates) == 0 {
return nil
}
var resumed []string
for name, e := range j.Updates {
st, ok := m.GetStack(name)
if !ok {
m.logger.Printf("[WARN] [stacks] update recovery: %s is in the journal (phase %s) but no longer exists — dropping the entry", name, e.Phase)
m.clearJournal(name)
continue
}
dir := filepath.Dir(st.ComposePath)
switch e.Phase {
case UpdatePhaseChecking, UpdatePhaseBackingUp, UpdatePhaseSafetyDump:
m.logger.Printf("[WARN] [stacks] update recovery: %s was interrupted in %s (started %s) — nothing had moved; dropping it", name, e.Phase, e.StartedAt.Format(time.RFC3339))
m.clearJournal(name)
m.finishUpdate(name, UpdatePhaseFailed, MsgUpdateInterrupted)
case UpdatePhasePinning, UpdatePhasePulling:
m.logger.Printf("[WARN] [stacks] update recovery: %s was interrupted in %s (started %s) — nothing had run; putting the pin back", name, e.Phase, e.StartedAt.Format(time.RFC3339))
m.pinBack(name, dir, e)
m.clearJournal(name)
m.finishUpdate(name, UpdatePhaseFailed, MsgUpdateInterrupted)
case UpdatePhaseStarting, UpdatePhaseVerifying:
m.logger.Printf("[WARN] [stacks] update recovery: %s was interrupted in %s (started %s) — the new version may have run; marking it Updating and RESUMING the health wait", name, e.Phase, e.StartedAt.Format(time.RFC3339))
m.mu.Lock()
if s, ok := m.stacks[name]; ok {
s.Updating, s.UpdateError = true, ""
s.UpdatePhase, s.UpdatePhaseLabel = UpdatePhaseVerifying, UpdatePhaseLabel(UpdatePhaseVerifying)
}
m.updateResume = append(m.updateResume, name)
m.mu.Unlock()
resumed = append(resumed, name)
default:
m.logger.Printf("[WARN] [stacks] update recovery: %s has unknown phase %q — dropping the entry", name, e.Phase)
m.clearJournal(name)
}
}
return resumed
}
// ResumeInterruptedUpdates continues the updates RecoverUpdates queued, once the backup side is wired
// (a resumed update that fails must be able to HOLD). Returns how many were resumed.
func (m *Manager) ResumeInterruptedUpdates(ctx context.Context) int {
m.mu.Lock()
names := m.updateResume
m.updateResume = nil
m.mu.Unlock()
for _, name := range names {
st, ok := m.GetStack(name)
if !ok {
m.finishUpdate(name, UpdatePhaseFailed, MsgUpdateInterrupted)
continue
}
m.updateJournalMu.Lock()
e, ok := m.readUpdateJournal().Updates[name]
m.updateJournalMu.Unlock()
if !ok {
m.finishUpdate(name, UpdatePhaseFailed, MsgUpdateInterrupted)
continue
}
provenAt, _ := time.Parse(time.RFC3339, e.ProvenCopyAt)
dir := filepath.Dir(st.ComposePath)
go func(name, dir string, e updateJournalEntry, provenAt time.Time) {
env := m.stackEnv(dir)
m.logger.Printf("[INFO] [stacks] update %s: resuming after a controller restart — `up -d` then the health wait", name)
if _, err := m.updateCompose(dir, env, "up", "-d", "--remove-orphans"); err != nil {
m.failAndHold(ctx, name, dir, env, provenAt, "resumed compose up failed: "+err.Error())
return
}
m.verifyAndConclude(ctx, name, dir, env, provenAt, e.StartedAt, &e)
}(name, dir, e, provenAt)
}
return len(names)
}