92cebb8c95
gates / gates (push) Successful in 11s
Measured live on demo-hp 2026-08-30 (controller 0.223.0): the nightly db-dump
and offbox-backup legs stop each stack ~13s to tar its volumes while the
deadapp-check job scans every 30s, so the scan caught whichever stack was
mid-cycle and pushed app_start_failed to the customer. 61 e-mails about apps
that were never broken.
The defect is not a missing mechanism. quiesce/suppress.go solved exactly this
in v0.179.0 and works -- but classifyRunStates read only the quiesce loop's set,
and that loop covers the WHOLE-GUEST backup. The per-app legs stop stacks
through Manager.DumpAppVolumesSafe, which registered with nothing. Two
mechanisms stop apps on purpose; only one told the alarm. Fifth instance of the
"seam built but never wired" class, and the first where the unwired half was a
consumer.
The suppression now rides AppStopGuard, which already brackets every deliberate
stop in the product (Begin before the stop, End after a successful restart) at
all three call sites, and which main.go hands as ONE object to the backup
manager and the exporter. scanDeployedAppRunStates takes the union of both sets.
All three per-app stop paths are covered, not only the reported nightly one.
It cannot latch -- End() runs only on a restart that SUCCEEDED, so unlike the
quiesce loop an open-ended hold is a real hazard here:
1. ReleaseFailed drops the entry IMMEDIATELY on a restart that broke, wired at
every failure path, so the app alarms on the next scan;
2. Begin REPLACES the set (one marker file = one operation);
3. appStopMaxHold (6h) caps a hold nothing released, logged at WARN.
Grace is 180s, deliberately quiesce's own constant and derivation. Suppression
is NOT persisted: after a crash the guard holds nothing and a down app must
alarm. ReleaseFailed keeps the durable crash marker; a test pins that.
Three companion red-proofs, each printing the pre-fix value (REPORT.md section 5):
- drop markStopped from Begin -> "suppressed at stop = map[]"
- drop ReleaseFailed from the dump -> "map[bookstack:true] after a restart that FAILED"
- pass nil instead of appStopGuard -> the AST wiring test fails
The third is load-bearing: the component was never the broken part, so a suite
that only injected it would have been green against the shipped defect.
Green gate clean: go build + go vet + go test ./... -- 28 packages, rc 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
161 lines
8.1 KiB
Go
161 lines
8.1 KiB
Go
package backup
|
|
|
|
import "time"
|
|
|
|
// ── R-330: the app-down alarm must not fire for an app the BACKUP ITSELF is holding down ─────────
|
|
//
|
|
// THE BUG, MEASURED LIVE ON demo-hp 2026-08-30 (controller 0.223.0). Every night both demo boxes
|
|
// e-mailed the customer `app_start_failed — "Telepített alkalmazás nem fut: <App>"` about apps that
|
|
// were never broken. `DumpAppVolumesSafe` stops a stack (`docker compose down`), tars its volumes and
|
|
// starts it again — ~13 s per stack — and the `deadapp-check` scheduler job runs every 30 s. It
|
|
// caught whichever stack was mid-cycle. Observed that night, UTC: the `db-dump` leg at 00:30 produced
|
|
// three events (Docmost, Paperless-ngx, RomM) and the `offbox-backup` leg at 02:15 produced two, while
|
|
// every dead-app scan of the other 23 hours logged `0 currently down`. 61 e-mails had accumulated.
|
|
//
|
|
// WHY R-97b's WINDOW DID NOT COVER IT. quiesce/suppress.go solves exactly this problem and solves it
|
|
// correctly — but it belongs to the QUIESCE LOOP, which stops stacks for the WHOLE-GUEST (vzdump/PBS)
|
|
// backup. `classifyRunStates` consults only `Loop.SuppressedStacks()`. The nightly per-app legs stop
|
|
// stacks through a different path (`Manager.DumpAppVolumesSafe`), which registered nothing with any
|
|
// suppressor. Two mechanisms stop apps; only one told the alarm.
|
|
//
|
|
// WHY THIS LIVES ON AppStopGuard, and not in a fourth place. The guard already brackets EVERY
|
|
// "we stopped this app on purpose" window in the product — `Begin` before the stop, `End` after a
|
|
// successful restart — at all three call sites (volume dump, off-site reconstitute, .fab export), and
|
|
// main.go hands the SAME guard object to the backup manager and the exporter. The fact the alarm needs
|
|
// ("we stopped it, and we have not given it back yet") is already here and has exactly one writer.
|
|
// Putting a fourth registry beside it would be the drift this codebase has paid for before.
|
|
//
|
|
// ── THE TENSION, WHICH IS THE WHOLE DESIGN (inherited from R-97b, and it still binds) ────────────
|
|
//
|
|
// Suppress while we hold the app and for a grace period after we let go — but an app that GENUINELY
|
|
// fails to come back MUST still alarm. Permanent suppression trades a loud false alarm for a silent
|
|
// real one, which is F-CRIT-1 and R-88 Scenario D over again. Three independent things stop this
|
|
// window from latching:
|
|
//
|
|
// 1. `ReleaseFailed` — a restart that was ATTEMPTED AND BROKE drops the entry immediately, so the
|
|
// app alarms on the very next scan rather than after any delay at all. This is the primary
|
|
// mechanism and every failure path calls it.
|
|
// 2. `Begin` REPLACES the set. The marker file holds one operation, so a new Begin proves the
|
|
// previous one is over; a set stranded by an earlier op cannot survive into a later one.
|
|
// 3. `appStopMaxHold` — a backstop for a hold nothing ever released. See its comment.
|
|
const (
|
|
// appStopAlarmGrace is how long after a successful restart a stack stays exempt.
|
|
//
|
|
// Deliberately the SAME 180 s as quiesce.quiesceAlarmGrace, and for the same derivation: the
|
|
// deploy flow allows 120 s for a stack to come up healthy, and the slowest catalog healthcheck
|
|
// start_period is Mealie's 60 s, after which a couple of check intervals must still elapse. Two
|
|
// suppression windows over the same alarm that disagreed on how long a restart takes would be a
|
|
// bug waiting to be found on whichever path used the shorter one.
|
|
//
|
|
// It is NOT longer than it needs to be: the dead-app scan runs on its own 30 s cadence, so an app
|
|
// that is genuinely dead alarms on the first scan after the window closes. The cost of this
|
|
// suppression is a BOUNDED DELAY in reporting a real failure, never its loss.
|
|
appStopAlarmGrace = 180 * time.Second
|
|
|
|
// appStopMaxHold caps an open-ended hold, and exists because of a hazard the quiesce loop does
|
|
// not have. `Loop` always calls markUnquiesced; `AppStopGuard.End()` is called only on a restart
|
|
// that SUCCEEDED, so a failure path that forgets to call `ReleaseFailed` would leave an entry
|
|
// open-ended forever — a silently dead app, which is the exact defect this file must not create
|
|
// while fixing a false alarm.
|
|
//
|
|
// Six hours is chosen to be longer than any real hold and far shorter than "forever": a volume
|
|
// dump holds a stack ~13 s (measured), a .fab export minutes, and even a multi-gigabyte off-site
|
|
// reconstitution is hours at the outside. Exceeding it means something is wrong, and the correct
|
|
// behaviour when something is wrong is to let the alarm through.
|
|
appStopMaxHold = 6 * time.Hour
|
|
)
|
|
|
|
// markStopped records `names` as exempt from app-down alarms for the duration of the current
|
|
// operation. The expiry is set at release; until then the entry is open-ended (bounded only by
|
|
// appStopMaxHold), because an operation may legitimately run for a long time and an app we are
|
|
// holding down that whole time must not alarm halfway through.
|
|
//
|
|
// It REPLACES the previous set rather than adding to it — see design note 2 above.
|
|
func (g *AppStopGuard) markStopped(names []string) {
|
|
if g == nil || len(names) == 0 {
|
|
return
|
|
}
|
|
g.suppressMu.Lock()
|
|
defer g.suppressMu.Unlock()
|
|
g.suppressed = make(map[string]appStopHold, len(names))
|
|
now := g.now()
|
|
for _, n := range names {
|
|
g.suppressed[n] = appStopHold{since: now} // zero `until` = still held
|
|
}
|
|
}
|
|
|
|
// releaseStarted starts the grace clock on every stack still held. Called from End(), which is the
|
|
// single "we gave the app back and it started" signal on all three paths.
|
|
func (g *AppStopGuard) releaseStarted() {
|
|
if g == nil {
|
|
return
|
|
}
|
|
until := g.now().Add(appStopAlarmGrace)
|
|
g.suppressMu.Lock()
|
|
defer g.suppressMu.Unlock()
|
|
for n, h := range g.suppressed {
|
|
if h.until.IsZero() {
|
|
h.until = until
|
|
g.suppressed[n] = h
|
|
}
|
|
}
|
|
}
|
|
|
|
// ReleaseFailed drops the suppression for stacks whose restart was ATTEMPTED AND BROKE, so they
|
|
// alarm on the next dead-app scan instead of being silenced by a window that was only ever meant to
|
|
// cover a restart in progress.
|
|
//
|
|
// Call it on EVERY path that stops an app and then fails to bring it back. It is the counterpart of
|
|
// quiesce's noteRestartOutcome, and the same rule applies: the distinguishing fact is not in the
|
|
// stack's state — a stack we stopped and could not restart is byte-identical on the Docker side to
|
|
// one the customer stopped — it is that WE tried and could not, and only the caller knows that.
|
|
//
|
|
// Nil-safe and idempotent: releasing a stack that is not suppressed is a no-op, so a caller may call
|
|
// it without first checking whether Begin ever ran.
|
|
func (g *AppStopGuard) ReleaseFailed(names ...string) {
|
|
if g == nil || len(names) == 0 {
|
|
return
|
|
}
|
|
g.suppressMu.Lock()
|
|
defer g.suppressMu.Unlock()
|
|
for _, n := range names {
|
|
delete(g.suppressed, n)
|
|
}
|
|
}
|
|
|
|
// SuppressedStacks returns the set of stack names currently exempt from app-down alarms: those an
|
|
// operation is holding stopped right now, plus those still inside the post-restart grace window.
|
|
//
|
|
// Nil-safe on a nil *AppStopGuard so the caller needs no branch — a controller with no guard
|
|
// suppresses nothing, which is the correct default.
|
|
func (g *AppStopGuard) SuppressedStacks() map[string]bool {
|
|
if g == nil {
|
|
return nil
|
|
}
|
|
now := g.now()
|
|
g.suppressMu.Lock()
|
|
defer g.suppressMu.Unlock()
|
|
out := make(map[string]bool, len(g.suppressed))
|
|
for n, h := range g.suppressed {
|
|
switch {
|
|
case !h.until.IsZero() && !now.Before(h.until):
|
|
delete(g.suppressed, n) // grace expired — reap so the map cannot grow without bound
|
|
case h.until.IsZero() && now.Sub(h.since) >= appStopMaxHold:
|
|
// Held open-ended past the backstop. Say so: an alarm that appears because a hold was
|
|
// never released must be explainable, and standing rule 3 wants a positive observable.
|
|
g.logger.Printf("[WARN] [appstop] %s has been held stopped for over %s with no release — dropping the alarm suppression so a genuine outage is not hidden", n, appStopMaxHold)
|
|
delete(g.suppressed, n)
|
|
default:
|
|
out[n] = true
|
|
}
|
|
}
|
|
return out
|
|
}
|
|
|
|
// appStopHold is one suppressed stack: when the hold started (for appStopMaxHold) and when it
|
|
// expires (zero while the app is still held).
|
|
type appStopHold struct {
|
|
since time.Time
|
|
until time.Time
|
|
}
|