v0.190.0 — the boot settle window, both gates on intent, and R-171
gates / gates (push) Successful in 8s

R-171 (a regression v0.189.0 introduced, CONFIRMED on hardware before any fix
was written). Replacing isBootOrphan's container-count term with recorded intent
made a drive-gate-stopped app read as a boot orphan: the gate stops apps with
`compose down` (zero containers) and never touches desired_state, because it is
not the customer. Observed on 9201 with the drive held unmounted — the sweep
found and started it, burned both attempts, and handed it to the dead-app alarm.
The write hazard did not materialise (the unbound mountpoint is host-root-owned
and the guest is unprivileged) but that protection is accidental and untested.
New consumer-side seam bootrecon.StartGate, fail-safe (cannot determine ⇒ do not
start), wired in main.go. The rule is not new: the API's startGatedByMissingDrive
already refuses this; the sweep bypassed it.

R-157 mechanism A. The sweep looked once at T+5s, deriving candidates from a
fleet docker was still restoring — three of six hard resets. Now a settle-then-
sweep window: sample every 5s, settled after 3 identical samples, sweep ONCE at
the end; ends on settled or a 50s budget, and the log says which. The budget is
50s because settle+budget+one retry must stay under the 90s dead-app grace — a
test rejected 60s at 95s. A window that overruns emits a LATE RECOVERY warn
rather than the grace being widened to hide it.

Widening the window made two more holders reachable, so the one gate covers all
three: an absent drive, a quiesce, and an in-flight app-data operation — reusing
quiesce.SuppressedStacks() and a new read-only AppStopGuard.HeldStacks().

R-170. shouldRecreateOnBoot now reads desired_state with the identical three-way
table; absent keeps the old hasContainers behaviour exactly. Its comment argued
for the container count and was rewritten. presentStable is untouched. The two
gates' agreement is pinned from both sides against one fixture table.

27/27 packages green; 6 red-proofs observed FAIL then restored.
This commit is contained in:
2026-08-02 19:56:20 +02:00
parent 3446609420
commit 582135f861
13 changed files with 1272 additions and 43 deletions
+87 -7
View File
@@ -41,6 +41,45 @@ type StackProvider interface {
RefreshStatus() error
}
// StartGate answers the one question this package must ask before starting anything: **may this app
// be started right now?** Declared consumer-side, in the style of StackProvider, so `bootrecon`
// still imports `stacks` alone and knows nothing about settings, the agent, quiesce or the web layer.
//
// It is ONE seam rather than three because the three reasons a boot orphan must NOT be started share
// a shape — something else is deliberately holding this app — and differ only in the reason string:
//
// the drive-absent gate stopped it → its drive is not live (R-171, below)
// a quiesce is holding it for a backup → the quiesce loop restarts its own stacks
// an app-data operation stopped it → the app-stop guard's own Recover owns it
//
// A widened boot window (R-157 mechanism A) is what makes the last two reachable at all: the old
// T+5 s single sweep never overlapped them.
//
// R-171 — WHY THIS EXISTS, and it is a regression this package caused. Until v0.189.0 the sweep
// required a stack to still HAVE containers, and an app the drive-absent gate had stopped has zero,
// so such apps were skipped by accident. v0.189.0 replaced that term with the customer's recorded
// intent — correctly — and the drive gate does NOT change `desired_state` (it is not the customer),
// so a gate-stopped app now reads as `running` + zero containers, i.e. a boot orphan. Observed live
// on 2026-08-02: the sweep found and started an app whose drive was unmounted, burned both attempts,
// and handed it to the dead-app alarm — a false alarm about an app the drive gate is deliberately
// holding (audits/DIAG-bootrecon-drive-absent-2026-08-02.md).
//
// The rule itself is not new and is not invented here: the API's own start path already refuses this
// (`startGatedByMissingDrive`, internal/api/router.go) with a Hungarian message to the customer. The
// sweep simply bypassed it by calling Manager.StartStack directly. This seam gives the sweep the
// same question to ask.
//
// CONTRACT — the answer is fail-safe by design (§8.4): an implementation that CANNOT DETERMINE
// whether the drive is live must return false, not true. Not starting is recoverable — the drive
// gate's `Return` branch restarts the app when the drive comes back, and the dead-app alarm reports
// it meanwhile. Starting on an absent drive is not recoverable by anything automatic: compose
// creates the bind sources wherever the mountpoint currently points, which is the guest rootfs.
type StartGate interface {
// MayStart reports whether the named stack may be started. The reason is for the log line and is
// only read when may is false.
MayStart(stackName string) (may bool, reason string)
}
const (
// DefaultAttempts is the total number of start attempts per boot (not per app per retry-forever).
DefaultAttempts = 2
@@ -58,6 +97,13 @@ type Reconciler struct {
// sleep is the inter-attempt wait; injectable so tests never spend 30 real seconds.
sleep func(context.Context, time.Duration)
// startGate (R-171) refuses to start an app something else is deliberately holding. nil = NOT
// WIRED, which means "this caller has no such concept" and is permissive — the test fixtures'
// case. It is NOT the same as "cannot determine", which the gate itself answers with false (see
// StartGate's contract). Production MUST wire it; TestMainWiresBootDriveGate walks main.go's AST
// for the call, because an unwired seam here is silently the pre-v0.190.0 behaviour.
startGate StartGate
}
// Result is the outcome, returned for logging/testing (the hub learns about failures only through
@@ -67,6 +113,23 @@ type Result struct {
Recovered []string // running again by the end
StillDown []string // still down after the last attempt — the alarm's problem now
Attempts int // attempts actually made (0 when there was nothing to do)
// HeldByDrive (R-171) are apps that ARE boot orphans by intent but which something else is
// deliberately holding (an absent drive, a quiesce, an app-data operation), so they were not
// started. Reported separately from StillDown because they are not a fault this sweep failed to
// fix — the holder owns their recovery. Collapsing the two would put a deliberately-held app in
// the same bucket as a broken one, which is the false alarm R-171 removes.
HeldByDrive []string
}
// SetDriveGate wires the R-171 start refusal. INIT-ONLY — call once, before Run.
func (r *Reconciler) SetDriveGate(g StartGate) { r.startGate = g }
// mayStart asks the gate, or allows when none is wired (see the startGate field comment).
func (r *Reconciler) mayStart(stackName string) (bool, string) {
if r.startGate == nil {
return true, ""
}
return r.startGate.MayStart(stackName)
}
// New builds a Reconciler with the shipped defaults.
@@ -109,9 +172,9 @@ func sleepCtx(ctx context.Context, d time.Duration) {
// customer stopped this" and left alone. The safety goal was right and still holds. The SIGNAL was
// wrong, because zero containers has at least three causes and the count cannot tell them apart:
//
// a deliberate Stop → must stay down
// a power cut mid-compose, or an interrupted deploy → must come back
// a backup that stopped the app and died before restarting it → must come back
// a deliberate Stop → must stay down
// a power cut mid-compose, or an interrupted deploy → must come back
// a backup that stopped the app and died before restarting it → must come back
//
// Two of those three were silently unrecoverable: the app simply stayed gone until a human noticed.
// The count was never capable of separating them, so the fix is not a better inference — it is to
@@ -162,16 +225,33 @@ func (r *Reconciler) Run(ctx context.Context) Result {
pending := map[string]bool{}
for _, s := range r.stacks.GetStacks() {
if isBootOrphan(s) {
pending[s.Name] = true
res.Candidates = append(res.Candidates, s.Name)
if !isBootOrphan(s) {
continue
}
// R-171: intent says this app should be running and it is not — but if something else is
// deliberately holding it (absent drive, quiesce, an app-data operation), starting it is the
// wrong repair. Refuse, loudly, and let the holder own it.
if live, reason := r.mayStart(s.Name); !live {
res.HeldByDrive = append(res.HeldByDrive, s.Name)
r.logger.Printf("[INFO] [bootrecon] %q is a boot orphan by intent but is HELD (%s) — NOT starting it; whatever is holding it owns its recovery",
s.Name, reason)
continue
}
pending[s.Name] = true
res.Candidates = append(res.Candidates, s.Name)
}
sortStrings(res.Candidates)
sortStrings(res.HeldByDrive)
if len(pending) == 0 {
// The healthy path must be observable — "no alarms" and "never ran" have to be
// distinguishable in a log (the v0.91.2 lesson).
// distinguishable in a log (the v0.91.2 lesson). "Nothing to start" and "everything I found
// is held by an absent drive" must be distinguishable too, or the held case reads as healthy.
if len(res.HeldByDrive) > 0 {
r.logger.Printf("[INFO] [bootrecon] Boot reconciliation: nothing to start — %d app(s) held (absent drive / quiesce / app-data operation): %v",
len(res.HeldByDrive), res.HeldByDrive)
return res
}
r.logger.Printf("[INFO] [bootrecon] Boot reconciliation: no boot-orphaned apps (nothing to start)")
return res
}