package backup import "time" // ── R-330: the app-down alarm must not fire for an app the BACKUP ITSELF is holding down ───────── // // THE BUG, MEASURED LIVE ON demo-hp 2026-08-30 (controller 0.223.0). Every night both demo boxes // e-mailed the customer `app_start_failed — "Telepített alkalmazás nem fut: "` about apps that // were never broken. `DumpAppVolumesSafe` stops a stack (`docker compose down`), tars its volumes and // starts it again — ~13 s per stack — and the `deadapp-check` scheduler job runs every 30 s. It // caught whichever stack was mid-cycle. Observed that night, UTC: the `db-dump` leg at 00:30 produced // three events (Docmost, Paperless-ngx, RomM) and the `offbox-backup` leg at 02:15 produced two, while // every dead-app scan of the other 23 hours logged `0 currently down`. 61 e-mails had accumulated. // // WHY R-97b's WINDOW DID NOT COVER IT. quiesce/suppress.go solves exactly this problem and solves it // correctly — but it belongs to the QUIESCE LOOP, which stops stacks for the WHOLE-GUEST (vzdump/PBS) // backup. `classifyRunStates` consults only `Loop.SuppressedStacks()`. The nightly per-app legs stop // stacks through a different path (`Manager.DumpAppVolumesSafe`), which registered nothing with any // suppressor. Two mechanisms stop apps; only one told the alarm. // // WHY THIS LIVES ON AppStopGuard, and not in a fourth place. The guard already brackets EVERY // "we stopped this app on purpose" window in the product — `Begin` before the stop, `End` after a // successful restart — at all three call sites (volume dump, off-site reconstitute, .fab export), and // main.go hands the SAME guard object to the backup manager and the exporter. The fact the alarm needs // ("we stopped it, and we have not given it back yet") is already here and has exactly one writer. // Putting a fourth registry beside it would be the drift this codebase has paid for before. // // ── THE TENSION, WHICH IS THE WHOLE DESIGN (inherited from R-97b, and it still binds) ──────────── // // Suppress while we hold the app and for a grace period after we let go — but an app that GENUINELY // fails to come back MUST still alarm. Permanent suppression trades a loud false alarm for a silent // real one, which is F-CRIT-1 and R-88 Scenario D over again. Three independent things stop this // window from latching: // // 1. `ReleaseFailed` — a restart that was ATTEMPTED AND BROKE drops the entry immediately, so the // app alarms on the very next scan rather than after any delay at all. This is the primary // mechanism and every failure path calls it. // 2. `Begin` REPLACES the set. The marker file holds one operation, so a new Begin proves the // previous one is over; a set stranded by an earlier op cannot survive into a later one. // 3. `appStopMaxHold` — a backstop for a hold nothing ever released. See its comment. const ( // appStopAlarmGrace is how long after a successful restart a stack stays exempt. // // Deliberately the SAME 180 s as quiesce.quiesceAlarmGrace, and for the same derivation: the // deploy flow allows 120 s for a stack to come up healthy, and the slowest catalog healthcheck // start_period is Mealie's 60 s, after which a couple of check intervals must still elapse. Two // suppression windows over the same alarm that disagreed on how long a restart takes would be a // bug waiting to be found on whichever path used the shorter one. // // It is NOT longer than it needs to be: the dead-app scan runs on its own 30 s cadence, so an app // that is genuinely dead alarms on the first scan after the window closes. The cost of this // suppression is a BOUNDED DELAY in reporting a real failure, never its loss. appStopAlarmGrace = 180 * time.Second // appStopMaxHold caps an open-ended hold, and exists because of a hazard the quiesce loop does // not have. `Loop` always calls markUnquiesced; `AppStopGuard.End()` is called only on a restart // that SUCCEEDED, so a failure path that forgets to call `ReleaseFailed` would leave an entry // open-ended forever — a silently dead app, which is the exact defect this file must not create // while fixing a false alarm. // // Six hours is chosen to be longer than any real hold and far shorter than "forever": a volume // dump holds a stack ~13 s (measured), a .fab export minutes, and even a multi-gigabyte off-site // reconstitution is hours at the outside. Exceeding it means something is wrong, and the correct // behaviour when something is wrong is to let the alarm through. appStopMaxHold = 6 * time.Hour ) // markStopped records `names` as exempt from app-down alarms for the duration of the current // operation. The expiry is set at release; until then the entry is open-ended (bounded only by // appStopMaxHold), because an operation may legitimately run for a long time and an app we are // holding down that whole time must not alarm halfway through. // // It REPLACES the previous set rather than adding to it — see design note 2 above. func (g *AppStopGuard) markStopped(names []string) { if g == nil || len(names) == 0 { return } g.suppressMu.Lock() defer g.suppressMu.Unlock() g.suppressed = make(map[string]appStopHold, len(names)) now := g.now() for _, n := range names { g.suppressed[n] = appStopHold{since: now} // zero `until` = still held } } // releaseStarted starts the grace clock on every stack still held. Called from End(), which is the // single "we gave the app back and it started" signal on all three paths. func (g *AppStopGuard) releaseStarted() { if g == nil { return } until := g.now().Add(appStopAlarmGrace) g.suppressMu.Lock() defer g.suppressMu.Unlock() for n, h := range g.suppressed { if h.until.IsZero() { h.until = until g.suppressed[n] = h } } } // ReleaseFailed drops the suppression for stacks whose restart was ATTEMPTED AND BROKE, so they // alarm on the next dead-app scan instead of being silenced by a window that was only ever meant to // cover a restart in progress. // // Call it on EVERY path that stops an app and then fails to bring it back. It is the counterpart of // quiesce's noteRestartOutcome, and the same rule applies: the distinguishing fact is not in the // stack's state — a stack we stopped and could not restart is byte-identical on the Docker side to // one the customer stopped — it is that WE tried and could not, and only the caller knows that. // // Nil-safe and idempotent: releasing a stack that is not suppressed is a no-op, so a caller may call // it without first checking whether Begin ever ran. func (g *AppStopGuard) ReleaseFailed(names ...string) { if g == nil || len(names) == 0 { return } g.suppressMu.Lock() defer g.suppressMu.Unlock() for _, n := range names { delete(g.suppressed, n) } } // SuppressedStacks returns the set of stack names currently exempt from app-down alarms: those an // operation is holding stopped right now, plus those still inside the post-restart grace window. // // Nil-safe on a nil *AppStopGuard so the caller needs no branch — a controller with no guard // suppresses nothing, which is the correct default. func (g *AppStopGuard) SuppressedStacks() map[string]bool { if g == nil { return nil } now := g.now() g.suppressMu.Lock() defer g.suppressMu.Unlock() out := make(map[string]bool, len(g.suppressed)) for n, h := range g.suppressed { switch { case !h.until.IsZero() && !now.Before(h.until): delete(g.suppressed, n) // grace expired — reap so the map cannot grow without bound case h.until.IsZero() && now.Sub(h.since) >= appStopMaxHold: // Held open-ended past the backstop. Say so: an alarm that appears because a hold was // never released must be explainable, and standing rule 3 wants a positive observable. g.logger.Printf("[WARN] [appstop] %s has been held stopped for over %s with no release — dropping the alarm suppression so a genuine outage is not hidden", n, appStopMaxHold) delete(g.suppressed, n) default: out[n] = true } } return out } // appStopHold is one suppressed stack: when the hold started (for appStopMaxHold) and when it // expires (zero while the app is still held). type appStopHold struct { since time.Time until time.Time }