R-856: after a crash boot of the host, app mails wait ~15 minutes; a normal boot keeps 90 s (09 decision 143)

The dead-app check (source of app_start_failed and app_stopped_unhealthy) now gates on a crash-aware
boot grace (internal/crashboot): 15 min when the host crash guard's last boot was UNCLEAN and within
30 min of the controller start, otherwise 90 s. The fact is read from the agent's local API
(GET /host/crash-guard, agentapi.Client.CrashGuard). UNKNOWN - no agent, an older agent's 404, no
crash-guard state - is a normal boot. The decision is logged once ("boot grace ...: ... (R-856)").

NEEDS AN AGENT CHANGE to take effect: GET /host/crash-guard serving the guard's state.json fields
(present, last_boot_at, last_boot_unclean, tripped). Until then every box keeps 90 s.

Tests: TestR856_CrashBootHoldsTheMailsForTheLongGrace, TestR856_NormalBootKeeps90s,
TestR856_FactReadLateInTheNormalGraceStillCounts, TestR856_AgentProbeReadsTheCrashGuardState,
TestR856_CrashGuardDecodesAndAnOlderAgentIs404, TestR856_DeadAppCheckWaitsOnTheCrashAwareGrace,
TestR856_NormalGraceIsTheDeadAppBootGrace.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-06 11:38:19 +02:00
parent 2d63714eca
commit c393d8529a
7 changed files with 411 additions and 2 deletions
+1
View File
@@ -353,6 +353,7 @@
| `Manager.execFn` (func seam) + `restartPolicyLookup` / `inspectRestartPolicyFn` (R-51, v0.156.0) | controller/internal/stacks/manager.go | nil → real `exec.Command` / `docker inspect -f {{.HostConfig.RestartPolicy.Name}}` | `scriptedDocker` in controller/internal/stacks/degraded_test.go drives the WHOLE production path (docker ps → aggregateState → docker inspect) — an aggregateState-only test proves the function, not the caller. Policy answers are cached per container+state and pruned to the live `docker ps` set; a FAILED inspect is deliberately never cached (a hiccup must not pin a container to "unknown") and reads as SUPERVISED, i.e. fail-closed — the opposite of `IsDownState`'s fail-open, because there the state is ambiguous while here a member is known dead |
| `bootrecon.StackProvider` (R-52, v0.156.0) | controller/internal/bootrecon/bootrecon.go | `*stacks.Manager` (GetStacks/StartStack/RefreshStatus) | `fakeStacks` counts StartStack per app; the load-bearing assertion is the NEGATIVE — a zero-container stack (a UI Stop = `compose down` = containers removed) must record **0** starts, while a boot orphan (containers present, Exited) records exactly 1. `Reconciler.sleep` is injected so the 30 s gap costs nothing |
| `bootReconcileFn` + `runBootReconcile` (package-main seam, v0.156.0) | controller/cmd/controller/main.go | `bootrecon.New(mgr, logger).Run` | controller/cmd/controller/bootrecon_wiring_test.go. **The wiring itself is asserted by an AST walk** over `func main()`, not a `strings.Contains` — the substring version passed its own red-proof because a commented-out call still contains the string. Comments are not callers |
| `crashboot.Probe` / `crashboot.Grace` (R-856) | controller/internal/crashboot/crashboot.go | `crashboot.AgentProbe(agentClient.CrashGuard)` (GET /host/crash-guard) | `fixed(Fact, err)` in crashboot_test.go. The dead-app check gates on `bootGrace.Within(now)` — 90 s normally, `CrashGrace` (15 min) when the HOST's last boot was unclean and within `BootWindow` of the controller start. **Unknown (no agent, 404, no state) is a normal boot** — never a longer silence on a guess. Asserted by AST walk in cmd/controller/r856_wiring_test.go |
| `classifyRunStates` (pure fix-3 derivation, v0.164.0) | controller/cmd/controller/main.go | `([]stacks.Stack, quiesced, failedRestart map[string]bool, now time.Time)` → `(dead []web.DeadApp, states []notify.AppRunState)` | classify_runstates_test.go. **THE single fix-3 rule: down = `(IsDownState(st.State) || st.CrashLooping(now)) && !userStopped && !quiesced`.** **`quiesced` is a UNION of TWO suppression sets** (R-330, v0.224.0): `quiesce.Loop.SuppressedStacks()` (whole-guest vzdump/PBS) and `backup.AppStopGuard.SuppressedStacks()` (per-app volume dump / offbox reconstitute / `.fab` export), merged by `unionSuppressed` in `scanDeployedAppRunStates`. **Adding a third way to stop an app means adding its set here** — R-330 was 61 false customer e-mails caused by exactly that omission, with a working suppressor sitting three lines away. C9-F2 (v0.183.0) added the crash-loop term: `restarting` is NOT in `IsDownState` and must not be — adding it alarms on every deploy and update fleet-wide — so a SUSTAINED restarting run (`stacks.crashLoopAfter` = 5 m, above the 120 s deploy timeout, Mealie's 60 s start_period AND R-97b's 180 s grace) becomes down instead. `now` is injected so the threshold is a testable contract. A deliberate UI stop (`compose down` → zero containers → StateStopped, I1) must not alarm — banner OR email — while faults (Exited/Degraded) alarm byte-identically; I2 (P2 census: all catalog services `unless-stopped`) is why a crash never rests at stopped. **Do NOT touch `IsDownState`** (other callers rely on stopped=down) and do NOT filter in `buildDeadAppAlerts`/`NotifyAppStartFailures` — one derivation point. If I1 or I2 changes, revisit the suppression |
| `report.SetPendingControllerLog` / `SetControllerLogSource` | controller/internal/report/selftail.go | ACK-armed consume-once self-log pull (the logtail.go shape) | selftail_test.go; source = `logBuffer.Lines`, wired once in main.go |
| `util.ParseVersion` / `util.Version.Compare` | controller/internal/util/version.go | THE one semver comparator (house rule: never a second) — selfupdate aliases it; agentapi's MinAgent comparison uses it | rejects pre-release/dev/latest (callers fall back, never trust); numeric compare (0.100 > 0.81) |