diff --git a/documentation/audits/CAMPAIGN-10-two-storage-soak-2026-07-31.md b/documentation/audits/CAMPAIGN-10-two-storage-soak-2026-07-31.md index 6dc4f9f..f1af43b 100644 --- a/documentation/audits/CAMPAIGN-10-two-storage-soak-2026-07-31.md +++ b/documentation/audits/CAMPAIGN-10-two-storage-soak-2026-07-31.md @@ -290,6 +290,32 @@ defence is a catalog gate asserting each template's mounted volume is a path the **Not fixed here** — the fences forbid it, and a fix mid-run would prove a version that did not exist when the run started. +### R-157 — bootrecon's start-ONCE sweep acts on an unsettled snapshot; the boot orphan survives + +**New. Register grepped: R-52 is the feature, nothing covers its timing.** Reproduced twice, in two +independent runs, at the same cycle (seeded permutation). Full evidence: +`../tests/campaign10-evidence-2026-07-31/r157-bootrecon-start-once-race.md`. + +A `qm reset` mid-backup brought everything back **except the app half of the DB-backed stack**: +`rallly` left `Exited (255)` (`oom=false`, `restarts=0`, its own log ending `✓ Ready` — it died +healthy) while `rallly-postgres` returned healthy. + + 20:28:13 Status refresh: 8 containers across 55 stacks <-- docker ps -a shows NINE + 20:28:18 [bootrecon] Boot reconciliation: no boot-orphaned apps (nothing to start) + 20:28:25 Status refresh: 7 containers … 8 containers <-- still churning AFTER the sweep + 20:39:14 [deadapp] 20 scans since boot, 5 deployed app(s) evaluated, 1 currently down + +The predicate is not at fault — once settled the controller reports `rallly state=degraded +containers=2`, and `IsDownState` includes `StateDegraded` (`manager.go:55`), so +`len(Containers)>0 && IsDownState(State)` holds. **The snapshot was**: bootrecon fires as +`go runBootReconcile(...)` (`cmd/controller/main.go:236`) ~5 s after start, while docker is still +restoring containers, and is start-once by design so it never re-checks. + +**Consequence: the app stays down indefinitely.** Detection is perfect and recovery never happens — +exactly R-52's original shape, *an alarm with no recovery* (F5: two apps Exited for ~18 h). In run 2a +the app only returned because a later campaign atom redeployed it. **Not fixed** — the fences forbid +it, and the settle-condition fix belongs with a test that pins the *consequence*. + ### 5.2 — Investigated and DISPROVED: `/api/disks` reports `state=attached` for an absent drive Recorded because the *chase* is the useful artefact, not because anything is broken. diff --git a/documentation/tests/campaign10-evidence-2026-07-31/r157-bootrecon-start-once-race.md b/documentation/tests/campaign10-evidence-2026-07-31/r157-bootrecon-start-once-race.md new file mode 100644 index 0000000..71e8c24 --- /dev/null +++ b/documentation/tests/campaign10-evidence-2026-07-31/r157-bootrecon-start-once-race.md @@ -0,0 +1,97 @@ +# R-157 — bootrecon's start-ONCE sweep acts on an unsettled snapshot, so the boot orphan it exists to recover survives + +**Reproduced twice, in two independent runs, at the same cycle.** Controller **0.188.0**, agent 0.119.0. + +R-52 built `internal/bootrecon` for exactly this failure (audit F5: `immich-server` and `calibre-web` +left `Exited` after an interrupted boot, ten siblings back, those two **still down ~18 h later** — +*"the controller REPORTED them (the 30 s deadapp-check) but never started them"*). This is that +failure recurring through a timing hole in the fix. + +## What happened + +Cycle 10's permutation ran `hard_reset_mid_write` — a `qm reset` of the VM during a backup. +Everything came back **except the app half of the DB-backed stack**: + +| | | +|---|---| +| `rallly` | **`Exited (255)`**, `oom=false`, no error string, **`restarts=0`** | +| `rallly-postgres` | `Up (healthy)` | +| every other app | back and healthy, both drives bound | +| rallly's own last log line | `▲ Next.js 16.2.6 … ✓ Ready in 0ms` — it died **healthy** | + +The controller start-up sequence, verbatim: + + 20:28:13 [stacks] ScanStacks complete: 55 stacks found (5 deployed, 50 available) + 20:28:13 [stacks] Status refresh: 8 containers across 55 stacks + 20:28:18 [bootrecon] Boot reconciliation: no boot-orphaned apps (nothing to start) + 20:28:24 [stacks] Status refresh: 8 containers across 55 stacks + 20:28:25 [stacks] Status refresh: 7 containers across 55 stacks <-- still churning + 20:28:25 [stacks] Status refresh: 8 containers across 55 stacks + 20:29:44 Event pushed: app_start_failed (warn) — Telepített alkalmazás nem fut: Rallly + 20:39:14 [deadapp] check alive: 20 scans since boot, 5 deployed app(s) evaluated, 1 currently down + +**`docker ps -a` reports 9 containers; every `Status refresh` reports 8.** The controller's boot-time +snapshot is one short, and the missing one is the exited `rallly`. + +## Why this is a defect and not the documented safety boundary + +bootrecon deliberately never touches a stack the customer stopped, and the predicate for that is +sound (`bootrecon.go:100`): + + len(s.Containers) > 0 && stacks.IsDownState(s.State) + +Both terms hold for this stack **once the system settles** — the controller's own API now reports: + + rallly state=degraded deployed=true containers=2 + rallly exited Exited (255) + rallly-postgres running Up (healthy) + +`IsDownState` includes `StateDegraded` (`manager.go:55`), and `len(Containers)` is 2. So the predicate +was never the problem. **The snapshot was.** bootrecon fires as `go runBootReconcile(...)` at +`cmd/controller/main.go:236`, ~5 s after start, while docker is still bringing containers back — and +it is **start-once by design**, so it never re-evaluates. + +Two mechanisms are consistent with the evidence and this report does not claim to distinguish them: + +1. the exited container was not yet enumerated, so the stack looked `running` (1/1) rather than + `degraded` (1/2) — supported by the 8-of-9 count; or +2. the stack's state had not yet aggregated to `degraded` at that instant. + +Either way the root is the same: **a single observation taken before the system settles, with no +re-check.** + +## Consequence + +The app stays down **indefinitely**. Detection works perfectly and recovery never happens — the +deadapp check correctly counted `1 currently down` across 20 scans / 10 minutes, and +`app_start_failed` reached the customer. In run 2a the app only returned because a *later* campaign +atom happened to redeploy the stack; nothing in the product recovered it. + +This is the R-52 failure shape exactly: *an alarm with no recovery*. + +## Why it is not a harness artifact + +The preceding two violations in this campaign **were** harness defects and are documented as such +(`run2a-violations-were-harness.md`). This one is not: + +- both drives were bound (`adatok: True`, `mentes: True`), so no drive gate stopped anything; +- every other app returned, including `calibre-web`, whose data is on an enrolled drive; +- only the app half of a two-container stack is missing, while its database is healthy; +- the fixed harness reported it through the check written for precisely this + (*"apps did not return within 10 min — that IS the failure"*), not through a seed or canary path; +- it reproduced at cycle 10 in **run 2a and run 2b**. The runner's RNG is seeded (`Random(20260801)`), + so both runs drew the same permutation — the reproduction is deterministic, not a coincidence. + +## Shape of a fix — not applied (the fences forbid it, and a fix mid-run proves a version that did not exist) + +The bounded start-once property is the safety argument and should be kept. What is missing is a +**settle condition** before the one observation: wait for the container set to stop changing (or for +docker to report ready) before the sweep, or re-evaluate once at the end of `deadAppBootGrace` and +recover then, still bounded to the same total attempts. A test would pin the consequence — *an app +left Exited by an interrupted boot is running again before the grace expires* — rather than the +mechanism, per the workspace rule about invariants needing a test that fails when the dependency moves. + +## Register + +`R-52` is the feature; grep found **no existing item covering its timing**. `R-156` is this campaign's +papra finding. **R-157 is free and is claimed here.**