R-330: stop the backup alarming about the apps it is holding down (v0.224.0)
gates / gates (push) Successful in 11s
gates / gates (push) Successful in 11s
Measured live on demo-hp 2026-08-30 (controller 0.223.0): the nightly db-dump
and offbox-backup legs stop each stack ~13s to tar its volumes while the
deadapp-check job scans every 30s, so the scan caught whichever stack was
mid-cycle and pushed app_start_failed to the customer. 61 e-mails about apps
that were never broken.
The defect is not a missing mechanism. quiesce/suppress.go solved exactly this
in v0.179.0 and works -- but classifyRunStates read only the quiesce loop's set,
and that loop covers the WHOLE-GUEST backup. The per-app legs stop stacks
through Manager.DumpAppVolumesSafe, which registered with nothing. Two
mechanisms stop apps on purpose; only one told the alarm. Fifth instance of the
"seam built but never wired" class, and the first where the unwired half was a
consumer.
The suppression now rides AppStopGuard, which already brackets every deliberate
stop in the product (Begin before the stop, End after a successful restart) at
all three call sites, and which main.go hands as ONE object to the backup
manager and the exporter. scanDeployedAppRunStates takes the union of both sets.
All three per-app stop paths are covered, not only the reported nightly one.
It cannot latch -- End() runs only on a restart that SUCCEEDED, so unlike the
quiesce loop an open-ended hold is a real hazard here:
1. ReleaseFailed drops the entry IMMEDIATELY on a restart that broke, wired at
every failure path, so the app alarms on the next scan;
2. Begin REPLACES the set (one marker file = one operation);
3. appStopMaxHold (6h) caps a hold nothing released, logged at WARN.
Grace is 180s, deliberately quiesce's own constant and derivation. Suppression
is NOT persisted: after a crash the guard holds nothing and a down app must
alarm. ReleaseFailed keeps the durable crash marker; a test pins that.
Three companion red-proofs, each printing the pre-fix value (REPORT.md section 5):
- drop markStopped from Begin -> "suppressed at stop = map[]"
- drop ReleaseFailed from the dump -> "map[bookstack:true] after a restart that FAILED"
- pass nil instead of appStopGuard -> the AST wiring test fails
The third is load-bearing: the component was never the broken part, so a suite
that only injected it would have been green against the shipped defect.
Green gate clean: go build + go vet + go test ./... -- 28 packages, rc 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
This commit is contained in:
@@ -2076,6 +2076,27 @@ invariant changes, revisit the suppression.
|
||||
callers rely on stopped counting as down). An out-of-band `docker compose stop` leaves the containers
|
||||
present → `StateExited` → still alerts, which is correct (out-of-band tampering is reportable).
|
||||
|
||||
> **R-330 (v0.224.0) — the alarm must be told by EVERY mechanism that stops an app, and until now it
|
||||
> was told by one.** The controller stops a customer's app on purpose in two quite separate places:
|
||||
> the **quiesce loop**, for the whole-guest (vzdump/PBS) backup, and the **`AppStopGuard`** paths —
|
||||
> the nightly volume dump, an off-site reconstitution and a `.fab` export. R-97b built the
|
||||
> suppression window for the first and it works. The second registered with nothing, so
|
||||
> `classifyRunStates` never knew, and the nightly backup e-mailed the customer
|
||||
> „Telepített alkalmazás nem fut" about apps it was holding down itself.
|
||||
>
|
||||
> Measured on `demo-hp` 2026-08-30 (controller 0.223.0): `DumpAppVolumesSafe` holds each stack down
|
||||
> **~13 s** while `deadapp-check` scans every **30 s**, so the scan caught whichever stack was
|
||||
> mid-cycle — 3 events at the 02:30 CEST `db-dump` leg, 2 more at the 04:15 `offbox-backup` leg,
|
||||
> every night, **61 e-mails**, while every other scan that day logged `0 currently down`.
|
||||
>
|
||||
> `classifyRunStates`'s `quiesced` argument is now the **union of both sets** (`unionSuppressed` in
|
||||
> `scanDeployedAppRunStates`). **A third way to stop an app means a third set here** — that omission
|
||||
> is the whole of this defect. The window cannot latch: a restart that was attempted and broke calls
|
||||
> `ReleaseFailed` and alarms on the **next** scan, `Begin` replaces the previous operation's set, and
|
||||
> a 6 h backstop covers a hold nothing released. Suppression is **not persisted** — after a crash the
|
||||
> guard holds nothing, `Recover()` either brings the app back or leaves it genuinely down, and a down
|
||||
> app must alarm.
|
||||
|
||||
**Boot desired-state reconciliation (R-52, v0.156.0, `internal/bootrecon`; rebuilt on recorded intent
|
||||
in R-166, v0.189.0).** A `deployed: true` app that missed its boot start used to stay down until a
|
||||
human noticed — the same shutdown that produced F4 left immich and calibre-web `Exited` while ten
|
||||
|
||||
@@ -727,7 +727,7 @@ func main() {
|
||||
if time.Since(startTime) < deadAppBootGrace {
|
||||
return nil // still inside the startup settle window
|
||||
}
|
||||
dead, states := scanDeployedAppRunStates(stackMgr, quiesceLoop)
|
||||
dead, states := scanDeployedAppRunStates(stackMgr, quiesceLoop, appStopGuard)
|
||||
alertMgr.SetDeadAppAlerts(dead)
|
||||
notifier.NotifyAppStartFailures(states)
|
||||
deadAppScans++
|
||||
@@ -2141,10 +2141,35 @@ func recordLateRecovery(logger *log.Logger, started time.Time, res bootrecon.Res
|
||||
// state-based dashboard banner) and EVERY deployed app's run state (for the notifier's one-event-per-
|
||||
// transition tracking). Deploying apps are skipped (mid-deploy is not a fault). Pure over GetStacks()
|
||||
// — the derivation itself lives in classifyRunStates so it is testable without a live Manager.
|
||||
func scanDeployedAppRunStates(mgr *stacks.Manager, q *quiesce.Loop) ([]web.DeadApp, []notify.AppRunState) {
|
||||
func scanDeployedAppRunStates(mgr *stacks.Manager, q *quiesce.Loop, g *backup.AppStopGuard) ([]web.DeadApp, []notify.AppRunState) {
|
||||
// R-97b: a stack THIS controller stopped for a backup is not a fault. q may be nil (unprovisioned
|
||||
// guest) — SuppressedStacks is nil-safe and returns nothing, i.e. suppress nothing.
|
||||
return classifyRunStates(mgr.GetStacks(), q.SuppressedStacks(), q.FailedRestarts(), time.Now())
|
||||
//
|
||||
// R-330: there are TWO mechanisms in this product that stop a customer's app on purpose, and until
|
||||
// v0.224.0 only one of them told the alarm. `q` covers the WHOLE-GUEST quiesce loop (vzdump/PBS);
|
||||
// `g` covers the per-app operations — the nightly volume dump, an off-site reconstitution and a
|
||||
// .fab export. Both are nil-safe, and the union is taken here rather than inside classifyRunStates
|
||||
// so that pure function keeps its single `quiesced` parameter and its existing tests.
|
||||
return classifyRunStates(mgr.GetStacks(), unionSuppressed(q.SuppressedStacks(), g.SuppressedStacks()), q.FailedRestarts(), time.Now())
|
||||
}
|
||||
|
||||
// unionSuppressed merges the suppression sets of the two mechanisms that stop apps on purpose.
|
||||
// Returns nil when both are empty so the common case allocates nothing.
|
||||
func unionSuppressed(a, b map[string]bool) map[string]bool {
|
||||
if len(a) == 0 {
|
||||
return b
|
||||
}
|
||||
if len(b) == 0 {
|
||||
return a
|
||||
}
|
||||
out := make(map[string]bool, len(a)+len(b))
|
||||
for n := range a {
|
||||
out[n] = true
|
||||
}
|
||||
for n := range b {
|
||||
out[n] = true
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// classifyRunStates is the pure fix-3 derivation over a plain stack slice. It splits the deployed
|
||||
@@ -2707,6 +2732,8 @@ func (a exportStopGuard) Begin(opID string, stacks []string) error {
|
||||
|
||||
func (a exportStopGuard) End() { a.g.End() }
|
||||
|
||||
func (a exportStopGuard) ReleaseFailed(stacks ...string) { a.g.ReleaseFailed(stacks...) }
|
||||
|
||||
func (a *exportAdapter) SaveEncryptedAppConfig(stackDir string, env map[string]string) error {
|
||||
meta := stacks.LoadMetadata(stackDir)
|
||||
sensitiveVars := stacks.SensitiveEnvVars(&meta)
|
||||
|
||||
@@ -0,0 +1,115 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"go/ast"
|
||||
"io"
|
||||
"log"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-controller/internal/backup"
|
||||
"gitea.dooplex.hu/admin/felhom-controller/internal/stacks"
|
||||
)
|
||||
|
||||
// ── R-330 — the per-app backup's suppression must actually REACH the alarm ───────────────────────
|
||||
//
|
||||
// Measured on demo-hp 2026-08-30 (controller 0.223.0): the nightly `db-dump` and `offbox-backup` legs
|
||||
// stop each stack for ~13 s to tar its volumes, the `deadapp-check` job scans every 30 s, and the
|
||||
// customer got `app_start_failed — "Telepített alkalmazás nem fut: <App>"` for apps that were never
|
||||
// broken. Sixty-one e-mails.
|
||||
//
|
||||
// THE SHAPE OF THE ORIGINAL DEFECT IS WHY THESE TESTS EXIST. `quiesce/suppress.go` already solved
|
||||
// this problem, correctly, in v0.179.0 — and the alarm still fired, because `classifyRunStates` read
|
||||
// only the QUIESCE loop's set while a SECOND mechanism (the AppStopGuard's per-app operations) was
|
||||
// stopping apps with nothing telling the alarm. A component that works and is not consulted is this
|
||||
// project's "seam built but never wired" class, hit five times now. So one test asserts the
|
||||
// CONSEQUENCE (a held app is not reported down) and one asserts the WIRING (main.go passes the
|
||||
// guard), because in this defect the component was never the broken part.
|
||||
|
||||
// TestHeldAppIsNotReportedDown is the consequence: whatever produces the suppression set, an app the
|
||||
// backup is holding must not reach the notifier's Down-set or the dashboard banner.
|
||||
func TestHeldAppIsNotReportedDown(t *testing.T) {
|
||||
g := backup.NewAppStopGuard(t.TempDir()+"/appstop.json", log.New(io.Discard, "", 0))
|
||||
if err := g.Begin("volume-dump:bookstack", backup.ReasonVolumeDump, []string{"bookstack"}); err != nil {
|
||||
t.Fatalf("Begin: %v", err)
|
||||
}
|
||||
|
||||
sts := []stacks.Stack{
|
||||
stack("bookstack", stacks.StateStopped, true, false), // held by the volume dump
|
||||
stack("romm", stacks.StateExited, true, false), // genuinely broken, must still alarm
|
||||
}
|
||||
// The union is exactly what scanDeployedAppRunStates builds. The nil first argument is the normal
|
||||
// state on a box with no whole-guest backup running — quiesce suppresses nothing, and the app-stop
|
||||
// guard's set has to carry the whole answer on its own.
|
||||
dead, states := classifyRunStates(sts, unionSuppressed(nil, g.SuppressedStacks()), nil, time.Now())
|
||||
|
||||
down := downByName(states)
|
||||
if down["bookstack"] {
|
||||
t.Fatal("the app the backup is holding stopped was reported DOWN — this is the false " +
|
||||
"`app_start_failed` e-mail the customer received every night")
|
||||
}
|
||||
if !down["romm"] {
|
||||
t.Fatal("a genuinely exited app stopped alarming — the fix silenced a real fault, which is " +
|
||||
"the over-correction (F-CRIT-1 / R-88 Scenario D) it must never make")
|
||||
}
|
||||
if deadNames(dead)["bookstack"] {
|
||||
t.Fatal("the held app still reached the dashboard dead-app banner")
|
||||
}
|
||||
if !deadNames(dead)["romm"] {
|
||||
t.Fatal("the genuinely exited app vanished from the dead-app banner")
|
||||
}
|
||||
}
|
||||
|
||||
func TestUnionSuppressed_KeepsBothMechanisms(t *testing.T) {
|
||||
// Two independent mechanisms stop apps on purpose. Dropping either set re-opens one of the two
|
||||
// false-alarm paths, and the bug shipped because only one was being read.
|
||||
got := unionSuppressed(map[string]bool{"quiesced-app": true}, map[string]bool{"dumped-app": true})
|
||||
if !got["quiesced-app"] || !got["dumped-app"] {
|
||||
t.Fatalf("union = %v, want both the quiesce loop's and the app-stop guard's stacks", got)
|
||||
}
|
||||
if got := unionSuppressed(nil, nil); len(got) != 0 {
|
||||
t.Fatalf("union of two empty sets = %v, want empty", got)
|
||||
}
|
||||
// Neither input may be mutated: both callers hold live maps that other code reads.
|
||||
a := map[string]bool{"a": true}
|
||||
b := map[string]bool{"b": true}
|
||||
unionSuppressed(a, b)
|
||||
if len(a) != 1 || len(b) != 1 {
|
||||
t.Fatalf("unionSuppressed mutated an input: a=%v b=%v", a, b)
|
||||
}
|
||||
}
|
||||
|
||||
// TestScanDeployedAppRunStatesIsGivenTheAppStopGuard walks main.go's AST. A substring search is not
|
||||
// enough — the sibling wiring tests in this package record why at first hand: a commented-out call
|
||||
// still satisfies strings.Contains, so the text version passed the very red-proof it existed to fail.
|
||||
func TestScanDeployedAppRunStatesIsGivenTheAppStopGuard(t *testing.T) {
|
||||
body := mainBody(t)
|
||||
|
||||
var found, withGuard bool
|
||||
ast.Inspect(body, func(n ast.Node) bool {
|
||||
call, ok := n.(*ast.CallExpr)
|
||||
if !ok {
|
||||
return true
|
||||
}
|
||||
id, ok := call.Fun.(*ast.Ident)
|
||||
if !ok || id.Name != "scanDeployedAppRunStates" {
|
||||
return true
|
||||
}
|
||||
found = true
|
||||
for _, arg := range call.Args {
|
||||
if a, ok := arg.(*ast.Ident); ok && a.Name == "appStopGuard" {
|
||||
withGuard = true
|
||||
}
|
||||
}
|
||||
return true
|
||||
})
|
||||
|
||||
if !found {
|
||||
t.Fatal("func main() never calls scanDeployedAppRunStates — the dead-app scan is not wired at all")
|
||||
}
|
||||
if !withGuard {
|
||||
t.Fatal("scanDeployedAppRunStates is called WITHOUT appStopGuard: the suppression set exists " +
|
||||
"but the alarm never reads it, which is precisely how R-330 shipped while R-97b's " +
|
||||
"identical mechanism sat working three lines away")
|
||||
}
|
||||
}
|
||||
@@ -114,6 +114,11 @@ type Exporter struct {
|
||||
type appStopGuard interface {
|
||||
Begin(opID string, stacks []string) error
|
||||
End()
|
||||
// ReleaseFailed (R-330) drops the app-down alarm suppression for a stack this export stopped and
|
||||
// then could not restart. Part of the seam rather than the exporter's own bookkeeping for the
|
||||
// same reason Begin is: the guard owns the "we are holding this app" fact, and a second owner
|
||||
// would drift from it.
|
||||
ReleaseFailed(stacks ...string)
|
||||
}
|
||||
|
||||
// NewExporter creates a new export/import engine.
|
||||
@@ -280,6 +285,12 @@ func (e *Exporter) executeExport(req ExportRequest, job *Job) {
|
||||
e.debugf("restarting stack %s after export", req.StackName)
|
||||
if err := e.provider.StartStack(req.StackName); err != nil {
|
||||
e.logger.Printf("[WARN] Export: could not restart %s: %v", req.StackName, err)
|
||||
// R-330: the app is genuinely down now, so the app-down alarm must see it on the very
|
||||
// next scan. The marker is still KEPT (above) so the next startup retries — the two
|
||||
// are independent: the marker is durable recovery, this is live alarm suppression.
|
||||
if e.stopGuard != nil {
|
||||
e.stopGuard.ReleaseFailed(req.StackName)
|
||||
}
|
||||
} else {
|
||||
e.debugf("stack %s restarted successfully", req.StackName)
|
||||
// Cleared only on a restart that succeeded — a failed one keeps the marker so the
|
||||
|
||||
@@ -8,6 +8,7 @@ import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"sort"
|
||||
"sync"
|
||||
"time"
|
||||
)
|
||||
|
||||
@@ -103,6 +104,14 @@ type AppStopGuard struct {
|
||||
now func() time.Time
|
||||
// starter is only needed by Recover; Begin/End work without one.
|
||||
starter AppStopStarter
|
||||
|
||||
// suppressed (R-330) is the in-memory "the alarm must not fire for these, we are holding them"
|
||||
// set. It is deliberately NOT persisted: after a restart the guard no longer holds anything —
|
||||
// Recover() either brings the apps back or leaves them genuinely down, and a down app must alarm.
|
||||
// Reviving a suppression across a crash would silence exactly the case the alarm exists for.
|
||||
// See appstop_suppress.go for the whole design.
|
||||
suppressMu sync.Mutex
|
||||
suppressed map[string]appStopHold
|
||||
}
|
||||
|
||||
// AppStopRecovery is what Recover found and did. Returned rather than pushed through a notifier
|
||||
@@ -202,13 +211,20 @@ func (g *AppStopGuard) Begin(opID string, reason AppStopReason, stackNames []str
|
||||
if len(stackNames) == 0 {
|
||||
return nil
|
||||
}
|
||||
return g.write(AppStopMarker{
|
||||
if err := g.write(AppStopMarker{
|
||||
Active: true,
|
||||
OpID: opID,
|
||||
Reason: reason,
|
||||
Stacks: append([]string(nil), stackNames...),
|
||||
StartedAt: g.now(),
|
||||
})
|
||||
}); err != nil {
|
||||
return err
|
||||
}
|
||||
// R-330: the marker is on disk, so the caller is now cleared to stop these apps — which is
|
||||
// exactly the moment the app-down alarm must stop counting them. Marked AFTER the write, so a
|
||||
// Begin that failed (and therefore stopped nothing) suppresses nothing either.
|
||||
g.markStopped(stackNames)
|
||||
return nil
|
||||
}
|
||||
|
||||
// End clears the marker after a successful restart. Best-effort by contract: a failure to clear is
|
||||
@@ -216,7 +232,15 @@ func (g *AppStopGuard) Begin(opID string, reason AppStopReason, stackNames []str
|
||||
// on the next boot, which is exactly D-b's "worst acceptable outcome" and far cheaper than failing
|
||||
// a backup that actually succeeded.
|
||||
func (g *AppStopGuard) End() {
|
||||
if g == nil || g.path == "" {
|
||||
if g == nil {
|
||||
return
|
||||
}
|
||||
// R-330: start the post-restart grace BEFORE the early return below, and unconditionally. End()
|
||||
// is the one "we gave the app back" signal on all three paths, and a guard with no marker path
|
||||
// still owes its suppressed stacks a release — otherwise an unwired-path guard would hold them
|
||||
// until appStopMaxHold.
|
||||
g.releaseStarted()
|
||||
if g.path == "" {
|
||||
return
|
||||
}
|
||||
if err := os.Remove(g.path); err != nil && !os.IsNotExist(err) {
|
||||
|
||||
@@ -0,0 +1,160 @@
|
||||
package backup
|
||||
|
||||
import "time"
|
||||
|
||||
// ── R-330: the app-down alarm must not fire for an app the BACKUP ITSELF is holding down ─────────
|
||||
//
|
||||
// THE BUG, MEASURED LIVE ON demo-hp 2026-08-30 (controller 0.223.0). Every night both demo boxes
|
||||
// e-mailed the customer `app_start_failed — "Telepített alkalmazás nem fut: <App>"` about apps that
|
||||
// were never broken. `DumpAppVolumesSafe` stops a stack (`docker compose down`), tars its volumes and
|
||||
// starts it again — ~13 s per stack — and the `deadapp-check` scheduler job runs every 30 s. It
|
||||
// caught whichever stack was mid-cycle. Observed that night, UTC: the `db-dump` leg at 00:30 produced
|
||||
// three events (Docmost, Paperless-ngx, RomM) and the `offbox-backup` leg at 02:15 produced two, while
|
||||
// every dead-app scan of the other 23 hours logged `0 currently down`. 61 e-mails had accumulated.
|
||||
//
|
||||
// WHY R-97b's WINDOW DID NOT COVER IT. quiesce/suppress.go solves exactly this problem and solves it
|
||||
// correctly — but it belongs to the QUIESCE LOOP, which stops stacks for the WHOLE-GUEST (vzdump/PBS)
|
||||
// backup. `classifyRunStates` consults only `Loop.SuppressedStacks()`. The nightly per-app legs stop
|
||||
// stacks through a different path (`Manager.DumpAppVolumesSafe`), which registered nothing with any
|
||||
// suppressor. Two mechanisms stop apps; only one told the alarm.
|
||||
//
|
||||
// WHY THIS LIVES ON AppStopGuard, and not in a fourth place. The guard already brackets EVERY
|
||||
// "we stopped this app on purpose" window in the product — `Begin` before the stop, `End` after a
|
||||
// successful restart — at all three call sites (volume dump, off-site reconstitute, .fab export), and
|
||||
// main.go hands the SAME guard object to the backup manager and the exporter. The fact the alarm needs
|
||||
// ("we stopped it, and we have not given it back yet") is already here and has exactly one writer.
|
||||
// Putting a fourth registry beside it would be the drift this codebase has paid for before.
|
||||
//
|
||||
// ── THE TENSION, WHICH IS THE WHOLE DESIGN (inherited from R-97b, and it still binds) ────────────
|
||||
//
|
||||
// Suppress while we hold the app and for a grace period after we let go — but an app that GENUINELY
|
||||
// fails to come back MUST still alarm. Permanent suppression trades a loud false alarm for a silent
|
||||
// real one, which is F-CRIT-1 and R-88 Scenario D over again. Three independent things stop this
|
||||
// window from latching:
|
||||
//
|
||||
// 1. `ReleaseFailed` — a restart that was ATTEMPTED AND BROKE drops the entry immediately, so the
|
||||
// app alarms on the very next scan rather than after any delay at all. This is the primary
|
||||
// mechanism and every failure path calls it.
|
||||
// 2. `Begin` REPLACES the set. The marker file holds one operation, so a new Begin proves the
|
||||
// previous one is over; a set stranded by an earlier op cannot survive into a later one.
|
||||
// 3. `appStopMaxHold` — a backstop for a hold nothing ever released. See its comment.
|
||||
const (
|
||||
// appStopAlarmGrace is how long after a successful restart a stack stays exempt.
|
||||
//
|
||||
// Deliberately the SAME 180 s as quiesce.quiesceAlarmGrace, and for the same derivation: the
|
||||
// deploy flow allows 120 s for a stack to come up healthy, and the slowest catalog healthcheck
|
||||
// start_period is Mealie's 60 s, after which a couple of check intervals must still elapse. Two
|
||||
// suppression windows over the same alarm that disagreed on how long a restart takes would be a
|
||||
// bug waiting to be found on whichever path used the shorter one.
|
||||
//
|
||||
// It is NOT longer than it needs to be: the dead-app scan runs on its own 30 s cadence, so an app
|
||||
// that is genuinely dead alarms on the first scan after the window closes. The cost of this
|
||||
// suppression is a BOUNDED DELAY in reporting a real failure, never its loss.
|
||||
appStopAlarmGrace = 180 * time.Second
|
||||
|
||||
// appStopMaxHold caps an open-ended hold, and exists because of a hazard the quiesce loop does
|
||||
// not have. `Loop` always calls markUnquiesced; `AppStopGuard.End()` is called only on a restart
|
||||
// that SUCCEEDED, so a failure path that forgets to call `ReleaseFailed` would leave an entry
|
||||
// open-ended forever — a silently dead app, which is the exact defect this file must not create
|
||||
// while fixing a false alarm.
|
||||
//
|
||||
// Six hours is chosen to be longer than any real hold and far shorter than "forever": a volume
|
||||
// dump holds a stack ~13 s (measured), a .fab export minutes, and even a multi-gigabyte off-site
|
||||
// reconstitution is hours at the outside. Exceeding it means something is wrong, and the correct
|
||||
// behaviour when something is wrong is to let the alarm through.
|
||||
appStopMaxHold = 6 * time.Hour
|
||||
)
|
||||
|
||||
// markStopped records `names` as exempt from app-down alarms for the duration of the current
|
||||
// operation. The expiry is set at release; until then the entry is open-ended (bounded only by
|
||||
// appStopMaxHold), because an operation may legitimately run for a long time and an app we are
|
||||
// holding down that whole time must not alarm halfway through.
|
||||
//
|
||||
// It REPLACES the previous set rather than adding to it — see design note 2 above.
|
||||
func (g *AppStopGuard) markStopped(names []string) {
|
||||
if g == nil || len(names) == 0 {
|
||||
return
|
||||
}
|
||||
g.suppressMu.Lock()
|
||||
defer g.suppressMu.Unlock()
|
||||
g.suppressed = make(map[string]appStopHold, len(names))
|
||||
now := g.now()
|
||||
for _, n := range names {
|
||||
g.suppressed[n] = appStopHold{since: now} // zero `until` = still held
|
||||
}
|
||||
}
|
||||
|
||||
// releaseStarted starts the grace clock on every stack still held. Called from End(), which is the
|
||||
// single "we gave the app back and it started" signal on all three paths.
|
||||
func (g *AppStopGuard) releaseStarted() {
|
||||
if g == nil {
|
||||
return
|
||||
}
|
||||
until := g.now().Add(appStopAlarmGrace)
|
||||
g.suppressMu.Lock()
|
||||
defer g.suppressMu.Unlock()
|
||||
for n, h := range g.suppressed {
|
||||
if h.until.IsZero() {
|
||||
h.until = until
|
||||
g.suppressed[n] = h
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// ReleaseFailed drops the suppression for stacks whose restart was ATTEMPTED AND BROKE, so they
|
||||
// alarm on the next dead-app scan instead of being silenced by a window that was only ever meant to
|
||||
// cover a restart in progress.
|
||||
//
|
||||
// Call it on EVERY path that stops an app and then fails to bring it back. It is the counterpart of
|
||||
// quiesce's noteRestartOutcome, and the same rule applies: the distinguishing fact is not in the
|
||||
// stack's state — a stack we stopped and could not restart is byte-identical on the Docker side to
|
||||
// one the customer stopped — it is that WE tried and could not, and only the caller knows that.
|
||||
//
|
||||
// Nil-safe and idempotent: releasing a stack that is not suppressed is a no-op, so a caller may call
|
||||
// it without first checking whether Begin ever ran.
|
||||
func (g *AppStopGuard) ReleaseFailed(names ...string) {
|
||||
if g == nil || len(names) == 0 {
|
||||
return
|
||||
}
|
||||
g.suppressMu.Lock()
|
||||
defer g.suppressMu.Unlock()
|
||||
for _, n := range names {
|
||||
delete(g.suppressed, n)
|
||||
}
|
||||
}
|
||||
|
||||
// SuppressedStacks returns the set of stack names currently exempt from app-down alarms: those an
|
||||
// operation is holding stopped right now, plus those still inside the post-restart grace window.
|
||||
//
|
||||
// Nil-safe on a nil *AppStopGuard so the caller needs no branch — a controller with no guard
|
||||
// suppresses nothing, which is the correct default.
|
||||
func (g *AppStopGuard) SuppressedStacks() map[string]bool {
|
||||
if g == nil {
|
||||
return nil
|
||||
}
|
||||
now := g.now()
|
||||
g.suppressMu.Lock()
|
||||
defer g.suppressMu.Unlock()
|
||||
out := make(map[string]bool, len(g.suppressed))
|
||||
for n, h := range g.suppressed {
|
||||
switch {
|
||||
case !h.until.IsZero() && !now.Before(h.until):
|
||||
delete(g.suppressed, n) // grace expired — reap so the map cannot grow without bound
|
||||
case h.until.IsZero() && now.Sub(h.since) >= appStopMaxHold:
|
||||
// Held open-ended past the backstop. Say so: an alarm that appears because a hold was
|
||||
// never released must be explainable, and standing rule 3 wants a positive observable.
|
||||
g.logger.Printf("[WARN] [appstop] %s has been held stopped for over %s with no release — dropping the alarm suppression so a genuine outage is not hidden", n, appStopMaxHold)
|
||||
delete(g.suppressed, n)
|
||||
default:
|
||||
out[n] = true
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// appStopHold is one suppressed stack: when the hold started (for appStopMaxHold) and when it
|
||||
// expires (zero while the app is still held).
|
||||
type appStopHold struct {
|
||||
since time.Time
|
||||
until time.Time
|
||||
}
|
||||
@@ -0,0 +1,214 @@
|
||||
package backup
|
||||
|
||||
import (
|
||||
"errors"
|
||||
"io"
|
||||
"log"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// ── R-330 — the nightly volume dump must not alarm about the app it is holding down ──────────────
|
||||
//
|
||||
// The defect these pin, measured on demo-hp 2026-08-30 with controller 0.223.0: `DumpAppVolumesSafe`
|
||||
// stopped a stack for ~13 s to tar its volumes while the `deadapp-check` job ran every 30 s, so the
|
||||
// scan caught the stack mid-cycle and pushed `app_start_failed` to the customer. 61 e-mails.
|
||||
//
|
||||
// The tests are split by what they must not lose:
|
||||
// - the SUPPRESSION exists and covers the whole window (the fix), and
|
||||
// - it can NEVER latch (the thing the fix must not break) — a restart that failed, a hold nothing
|
||||
// released, and a stranded set from an earlier operation each let the alarm through.
|
||||
//
|
||||
// RED-PROOF (run 2026-08-30, recorded in REPORT.md): removing the `g.markStopped(stackNames)` call
|
||||
// from `AppStopGuard.Begin` fails TestVolumeDump_SuppressesTheAlarmForTheAppItIsHolding with
|
||||
// `suppressed at stop = map[]` — the exact pre-fix shape that produced the e-mails.
|
||||
|
||||
// suppressWatchProvider drives the REAL DumpAppVolumesSafe and records the alarm-suppression set at
|
||||
// the two moments that matter: while the app is stopped, and at the restart call. Asserting on the
|
||||
// suppression set (what the dead-app scanner actually reads) rather than on a log line is the
|
||||
// standing-rule-3 positive observable — and it is the CONSEQUENCE, not the mechanism.
|
||||
type suppressWatchProvider struct {
|
||||
StackDataProvider
|
||||
guard *AppStopGuard
|
||||
|
||||
suppressedAtStop map[string]bool
|
||||
suppressedAtStart map[string]bool
|
||||
startErr error
|
||||
}
|
||||
|
||||
func (p *suppressWatchProvider) GetDockerVolumes(string) []string { return nil }
|
||||
|
||||
func (p *suppressWatchProvider) StopStack(string) error {
|
||||
p.suppressedAtStop = p.guard.SuppressedStacks()
|
||||
return nil
|
||||
}
|
||||
|
||||
func (p *suppressWatchProvider) StartStack(string) error {
|
||||
p.suppressedAtStart = p.guard.SuppressedStacks()
|
||||
return p.startErr
|
||||
}
|
||||
|
||||
// newSuppressManager builds a Manager over a real guard and returns both.
|
||||
func newSuppressManager(t *testing.T, p *suppressWatchProvider) *Manager {
|
||||
t.Helper()
|
||||
dir := t.TempDir()
|
||||
lg := log.New(io.Discard, "", 0)
|
||||
m := &Manager{logger: lg, stackProvider: p, systemDataPath: dir}
|
||||
m.appStop = NewAppStopGuard(markerPath(dir), lg)
|
||||
p.guard = m.appStop
|
||||
return m
|
||||
}
|
||||
|
||||
func TestVolumeDump_SuppressesTheAlarmForTheAppItIsHolding(t *testing.T) {
|
||||
// THE FIX. Through the production path, not by calling Begin from the test: an earlier sibling
|
||||
// test in this package was rewritten for exactly that reason — proving the guard works is not
|
||||
// proving DumpAppVolumesSafe uses it.
|
||||
p := &suppressWatchProvider{}
|
||||
m := newSuppressManager(t, p)
|
||||
|
||||
if err := m.DumpAppVolumesSafe("bookstack"); err != nil {
|
||||
t.Fatalf("DumpAppVolumesSafe: %v", err)
|
||||
}
|
||||
|
||||
if !p.suppressedAtStop["bookstack"] {
|
||||
t.Fatalf("suppressed at stop = %v, want bookstack — the dead-app scan runs every 30s and the "+
|
||||
"stack is down for ~13s, so an unsuppressed window is the false `app_start_failed` e-mail "+
|
||||
"the customer received nightly", p.suppressedAtStop)
|
||||
}
|
||||
if !p.suppressedAtStart["bookstack"] {
|
||||
t.Fatalf("suppressed at restart = %v, want bookstack — the app is STILL down at this point",
|
||||
p.suppressedAtStart)
|
||||
}
|
||||
// And it is still suppressed after the restart returns: R-97b's BookStack alarmed while `starting`
|
||||
// / `unhealthy`, which is neither stopped nor healthy, so a window that ends at the restart CALL
|
||||
// closes too early to fix anything.
|
||||
if !m.appStop.SuppressedStacks()["bookstack"] {
|
||||
t.Fatal("the suppression ended the instant the restart returned — an app that is up but not " +
|
||||
"yet healthy still reads as down, which is the exact shape R-97b was filed for")
|
||||
}
|
||||
}
|
||||
|
||||
func TestVolumeDump_SuppressionExpiresSoARealOutageStillAlarms(t *testing.T) {
|
||||
// The window is a BOUNDED DELAY in reporting a real failure, never its loss. Without this the fix
|
||||
// would trade a loud false alarm for a silent real one — F-CRIT-1 and R-88 Scenario D.
|
||||
p := &suppressWatchProvider{}
|
||||
m := newSuppressManager(t, p)
|
||||
|
||||
base := time.Now()
|
||||
m.appStop.now = func() time.Time { return base }
|
||||
|
||||
if err := m.DumpAppVolumesSafe("bookstack"); err != nil {
|
||||
t.Fatalf("DumpAppVolumesSafe: %v", err)
|
||||
}
|
||||
if !m.appStop.SuppressedStacks()["bookstack"] {
|
||||
t.Fatal("not suppressed immediately after the restart")
|
||||
}
|
||||
|
||||
// One second before the grace closes: still suppressed.
|
||||
m.appStop.now = func() time.Time { return base.Add(appStopAlarmGrace - time.Second) }
|
||||
if !m.appStop.SuppressedStacks()["bookstack"] {
|
||||
t.Fatalf("the grace closed early — a stack restarted %s ago must still be exempt", appStopAlarmGrace-time.Second)
|
||||
}
|
||||
|
||||
// At the grace boundary: the alarm owns it again.
|
||||
m.appStop.now = func() time.Time { return base.Add(appStopAlarmGrace) }
|
||||
if m.appStop.SuppressedStacks()["bookstack"] {
|
||||
t.Fatalf("still suppressed %s after the restart — an app that genuinely failed to come back "+
|
||||
"would never be reported", appStopAlarmGrace)
|
||||
}
|
||||
}
|
||||
|
||||
func TestVolumeDump_FailedRestartAlarmsImmediately(t *testing.T) {
|
||||
// The primary anti-latch mechanism. We stopped it, we could not give it back: it is genuinely
|
||||
// down, and it must alarm on the NEXT scan — not after any grace at all.
|
||||
p := &suppressWatchProvider{startErr: errors.New("compose up failed")}
|
||||
m := newSuppressManager(t, p)
|
||||
|
||||
if err := m.DumpAppVolumesSafe("bookstack"); err == nil {
|
||||
t.Fatal("a failed restart must surface as an error")
|
||||
}
|
||||
|
||||
if got := m.appStop.SuppressedStacks(); got["bookstack"] {
|
||||
t.Fatalf("suppressed = %v after a restart that FAILED — the app is down and the alarm is the "+
|
||||
"only thing that would tell anyone, which is F-CRIT-1 exactly", got)
|
||||
}
|
||||
// The durable marker is a SEPARATE concern and must be untouched by the release: it is what the
|
||||
// next startup retries from. Losing it here would trade a false alarm for a lost recovery.
|
||||
if !markerExists(t, m.systemDataPath) {
|
||||
t.Fatal("ReleaseFailed also cleared the crash marker — the next startup would not retry the restart")
|
||||
}
|
||||
}
|
||||
|
||||
func TestSuppression_CannotOutliveTheBackstop(t *testing.T) {
|
||||
// The belt-and-braces for a hold nothing ever released. `End()` runs only on a restart that
|
||||
// SUCCEEDED, so a future failure path that forgets ReleaseFailed would otherwise silence an app
|
||||
// forever. Exceeding the backstop means something is wrong, and the right answer when something
|
||||
// is wrong is to let the alarm through.
|
||||
base := time.Now()
|
||||
g := NewAppStopGuard(markerPath(t.TempDir()), log.New(io.Discard, "", 0))
|
||||
g.now = func() time.Time { return base }
|
||||
|
||||
if err := g.Begin("volume-dump:romm", ReasonVolumeDump, []string{"romm"}); err != nil {
|
||||
t.Fatalf("Begin: %v", err)
|
||||
}
|
||||
if !g.SuppressedStacks()["romm"] {
|
||||
t.Fatal("not suppressed while held")
|
||||
}
|
||||
|
||||
g.now = func() time.Time { return base.Add(appStopMaxHold - time.Minute) }
|
||||
if !g.SuppressedStacks()["romm"] {
|
||||
t.Fatalf("the backstop fired early — a legitimate long operation would start alarming mid-run")
|
||||
}
|
||||
|
||||
g.now = func() time.Time { return base.Add(appStopMaxHold) }
|
||||
if g.SuppressedStacks()["romm"] {
|
||||
t.Fatalf("a hold open-ended for %s is still suppressing the alarm — nothing released it and "+
|
||||
"nothing ever will", appStopMaxHold)
|
||||
}
|
||||
}
|
||||
|
||||
func TestBeginReplacesThePreviousOperationsSet(t *testing.T) {
|
||||
// The marker file holds ONE operation, so a new Begin proves the previous one is over. Without
|
||||
// this, a set stranded by an operation that died between Begin and End would suppress its stacks
|
||||
// for the whole appStopMaxHold, even though a later operation has since taken over.
|
||||
g := NewAppStopGuard(markerPath(t.TempDir()), log.New(io.Discard, "", 0))
|
||||
|
||||
if err := g.Begin("volume-dump:romm", ReasonVolumeDump, []string{"romm"}); err != nil {
|
||||
t.Fatalf("Begin romm: %v", err)
|
||||
}
|
||||
if err := g.Begin("volume-dump:kimai", ReasonVolumeDump, []string{"kimai"}); err != nil {
|
||||
t.Fatalf("Begin kimai: %v", err)
|
||||
}
|
||||
|
||||
got := g.SuppressedStacks()
|
||||
if got["romm"] {
|
||||
t.Fatalf("suppressed = %v — romm's hold survived into a later operation that is not holding it", got)
|
||||
}
|
||||
if !got["kimai"] {
|
||||
t.Fatalf("suppressed = %v, want kimai — the current operation's own stack is not exempt", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSuppression_FailedBeginSuppressesNothing(t *testing.T) {
|
||||
// A Begin whose write failed means the caller REFUSES to stop the app (DumpAppVolumesSafe returns
|
||||
// early). Nothing is stopped, so nothing may be exempt — suppressing here would hide a genuinely
|
||||
// dead app that this operation never touched.
|
||||
g := NewAppStopGuard("/proc/felhom-nonexistent-dir/appstop.json", log.New(io.Discard, "", 0))
|
||||
|
||||
if err := g.Begin("volume-dump:romm", ReasonVolumeDump, []string{"romm"}); err == nil {
|
||||
t.Skip("the unwritable path became writable — this environment cannot exercise the failure")
|
||||
}
|
||||
if got := g.SuppressedStacks(); got["romm"] {
|
||||
t.Fatalf("suppressed = %v after a Begin that FAILED — the app was never stopped", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestNilGuardSuppressesNothing(t *testing.T) {
|
||||
// An unprovisioned guest has no guard. Suppressing nothing is the correct default, and nil-safety
|
||||
// is what lets the caller in main.go stay branch-free.
|
||||
var g *AppStopGuard
|
||||
if got := g.SuppressedStacks(); len(got) != 0 {
|
||||
t.Fatalf("a nil guard suppressed %v", got)
|
||||
}
|
||||
g.ReleaseFailed("romm") // must not panic
|
||||
}
|
||||
@@ -879,6 +879,10 @@ func (m *Manager) DumpAppVolumesSafe(stackName string) error {
|
||||
startErr := m.stackProvider.StartStack(stackName)
|
||||
if startErr != nil {
|
||||
m.logger.Printf("[ERROR] [backup] Failed to restart %s after volume dump: %v", stackName, startErr)
|
||||
// R-330: we stopped it and could not give it back, so it is genuinely down and the app-down
|
||||
// alarm must own it from the very next scan. Dropping the suppression here is what keeps this
|
||||
// window from turning R-330's false alarm into F-CRIT-1's silent one.
|
||||
m.appStop.ReleaseFailed(stackName)
|
||||
} else {
|
||||
// Cleared ONLY on a restart that succeeded. A failed restart keeps the marker so the next
|
||||
// startup retries — the app really is still owed one.
|
||||
|
||||
@@ -654,6 +654,10 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string, ack
|
||||
err := m.stackProvider.StartStack(stack)
|
||||
if err == nil {
|
||||
m.appStop.End()
|
||||
} else {
|
||||
// R-330: same rule as the volume-dump path — a restart that was attempted and broke
|
||||
// leaves the app genuinely down, so the alarm suppression must go immediately.
|
||||
m.appStop.ReleaseFailed(stack)
|
||||
}
|
||||
return err
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user