v0.222.0: ask whether a supervised member is DEAD before whether one is UNHEALTHY (R-384), and stop promising an undo copy nobody looked for (R-383)
gates / gates (push) Successful in 11s

R-384. aggregateState returned StateUnhealthy the moment unhealthy > 0, and the
R-51 mixed-case block that asks "is a supervised member dead?" sat below it. A
two-container app whose database exits goes unhealthy BECAUSE it cannot reach
that database - so the symptom the dead database causes was what suppressed the
alarm for it. unhealthy is not a down state, so classifyRunStates never marked
the app down and app_start_failed never fired.

Measured live on demo-hp 2026-08-22: bookstack-db stopped at 21:27:01 and the
F-OBS heartbeat printed "0 currently down" throughout. R-51's 18-hour immich
failure, back through a different door.

Two things moved, and either alone leaves the defect standing: the supervised
test is hoisted above the unhealthy/starting/restarting returns, and "some
members are up" now counts ANY member not in the down bucket. The old guard was
running > 0, which made the R-51 block unreachable in exactly the case it was
written for.

IsDownState is byte-identical - unhealthy stays excluded, because an unhealthy
container is running and folding it in reintroduces the flapping that exclusion
exists to stop. No new state was minted. Only the ORDER changed. The priority
comment was rewritten because it asserted an ordering the code no longer has.

Three subtests in TestAggregateState_UnchangedBranches were AMENDED: they
asserted an unhealthy/starting/restarting member beat an exited peer on
unless-stopped, which pinned the defect as settled behaviour. They keep their
intent with the down member given a benign policy.

R-383. The double-failure message said the previous state's backup EXISTS,
built from the returned path without asking the filesystem - and a missing file
is one of the two ways that rollback fails. undoCopyPhrase now describes the
copy from disk: present, partial, missing (still naming where it should be), or
never written. Zero-length counts as missing.

Test count 1494 -> 1504. Four red-proofs planted, four seen failing; the two
halves of R-384 convict independently.
This commit is contained in:
2026-08-23 07:25:00 +02:00
parent da75603553
commit 5da11c4480
7 changed files with 638 additions and 27 deletions
+85
View File
@@ -1,3 +1,88 @@
## v0.222.0 — an app whose database dies raised no alarm, because the wrong question answered first (2026-08-23, R-384 + R-383)
**MinAgent: 0.129.0** (unchanged — no new agent coupling)
**`IsDownState` did NOT change.** That is the first thing to say, because the obvious fix here is the
wrong one. `unhealthy` is still excluded from the down set, byte-identical, for the reason recorded at
`manager.go:41-53`: an unhealthy container is *running*, and folding it in reintroduces the flapping
that exclusion exists to stop. No new container state was minted either — `StateDegraded` already
means exactly this and every consumer already handles it. **The defect was the ORDER of two questions,
and only the order moved.**
### R-384 — a dead database hid behind its own unhealthy front end
"Is a SUPERVISED member of this app dead?" and "is a RUNNING member failing its healthcheck?" are two
different questions. `aggregateState` returned `StateUnhealthy` the moment `unhealthy > 0`, and the
R-51 mixed-case block that asks the first question sat **below** it — so the second question was
answering the first, and always won.
A two-container app whose database exits goes unhealthy seconds later *because* it cannot reach that
database. So the very symptom the dead database causes was what suppressed the alarm for it.
`unhealthy` is not a down state, so `classifyRunStates` never marked the app down, no banner appeared
and `app_start_failed` never fired.
**Measured live on `demo-hp` 2026-08-22:** `bookstack-db` stopped at 21:27:01 and the F-OBS heartbeat
printed **"0 currently down"** throughout. This is R-51's own 18-hour immich failure returning through
a different door — the dead member hiding behind an *unhealthy survivor* instead of behind three live
helpers.
**Two things had to move, and either one alone leaves the defect standing:**
1. **The order.** The supervised-down test is hoisted above the `unhealthy`/`starting`/`restarting`
returns.
2. **The guard.** "Some members are up" now means **any member not in the down bucket** — running,
unhealthy, starting or restarting. The old guard was `running > 0`, counting `StateRunning` alone,
which made the R-51 block **unreachable in exactly the case it was written for**: an unhealthy
survivor beside a dead database counted as nothing up.
**The benign case is untouched.** A one-shot init/migrate container that has finished has policy
`no`/`on-failure`, `supervisedPolicy` returns false, and the stack reads exactly as before. Without
that filter every app with a migration step would alarm on every start.
**The priority comment was rewritten**, because it asserted an ordering the code no longer has, and a
comment asserting an invariant the code does not provide is what cost this project four months one
version ago.
**Three existing subtests were AMENDED, and this is reported rather than buried.**
`TestAggregateState_UnchangedBranches` asserted that an unhealthy / starting / restarting member beat
an `exited` peer that was on `unless-stopped` — that is, it pinned the defect as if it were settled
behaviour. They keep their original intent ("the live member's state wins over a down member") with
the down member given a BENIGN policy, which is the only situation in which that sentence was ever
true. The supervised versions now assert `degraded`.
**Every `IsDownState` consumer was walked and is named in `REPORT.md`.** Two change deliberately
(`classifyRunStates`, the intended fix; and `bootrecon`, which will now repair a half-started stack at
boot instead of calling it recovered). `isObservedUp` is an allow-list of `{running, starting}` and is
unaffected — verified, not assumed. The quiesce suppression is cycle-keyed and state-blind, so
R-97b's guarantee is untouched.
### R-383 — the double-failure message promised an undo copy it never looked for
When BOTH a database replay and its rollback fail, the message ended „a korábbi állapot mentése
megvan: <file>" — *the previous state's backup exists*. It was built from the path `writeSafetyDump`
returned, **without ever asking the filesystem**. One of the two ways the rollback fails is that the
file is gone, so the sentence was most likely to be false in precisely the case it was printed.
Measured twice live, on v0.220.2 and v0.221.1.
`undoCopyPhrase` now describes the undo copy **from disk**: present (named), partially present (both
halves named), missing (says so, and still names where it should have been), or never written. A
zero-length dump counts as missing — a 0-byte file restores nothing. The filename is **not** simply
dropped: R-351's lesson is that a refusal naming nothing forces a person to remember what the product
already knows.
### Tests
`internal/stacks/degraded_test.go` (R-384 unit + production-path wiring),
`cmd/controller/r384_dead_db_alarm_test.go` (the classifier consequence + the quiesce suppression),
`internal/backup/r383_undo_phrase_test.go` (the phrase + an AST seam test that the message is still
wired to the builder). Test count **1494 → 1504**.
**Red-proofs: four planted, four SEEN FAILING.** The two halves of R-384 convict independently —
reverting the hoist reads `"unhealthy"`, the exact state the live box reported, and narrowing `up`
back to `running` reads `"unhealthy"`/`"starting"`/`"restarting"`. The classifier mutation returns an
empty banner. The R-383 mutation prints the false claim verbatim, and for an empty set printed
`megvan: .` — naming a file that never existed. Guards sit at the layer each defect lives in: the
ordering at `aggregateState`, the consequence at `classifyRunStates`, the claim at the phrase builder.
## v0.221.1 — the undo-copy prune stopped running because another fix made its guard reachable (2026-08-23, R-361 follow-on)
**MinAgent: 0.129.0** (unchanged — no new agent coupling)