5da11c4480
gates / gates (push) Successful in 11s
R-384. aggregateState returned StateUnhealthy the moment unhealthy > 0, and the R-51 mixed-case block that asks "is a supervised member dead?" sat below it. A two-container app whose database exits goes unhealthy BECAUSE it cannot reach that database - so the symptom the dead database causes was what suppressed the alarm for it. unhealthy is not a down state, so classifyRunStates never marked the app down and app_start_failed never fired. Measured live on demo-hp 2026-08-22: bookstack-db stopped at 21:27:01 and the F-OBS heartbeat printed "0 currently down" throughout. R-51's 18-hour immich failure, back through a different door. Two things moved, and either alone leaves the defect standing: the supervised test is hoisted above the unhealthy/starting/restarting returns, and "some members are up" now counts ANY member not in the down bucket. The old guard was running > 0, which made the R-51 block unreachable in exactly the case it was written for. IsDownState is byte-identical - unhealthy stays excluded, because an unhealthy container is running and folding it in reintroduces the flapping that exclusion exists to stop. No new state was minted. Only the ORDER changed. The priority comment was rewritten because it asserted an ordering the code no longer has. Three subtests in TestAggregateState_UnchangedBranches were AMENDED: they asserted an unhealthy/starting/restarting member beat an exited peer on unless-stopped, which pinned the defect as settled behaviour. They keep their intent with the down member given a benign policy. R-383. The double-failure message said the previous state's backup EXISTS, built from the returned path without asking the filesystem - and a missing file is one of the two ways that rollback fails. undoCopyPhrase now describes the copy from disk: present, partial, missing (still naming where it should be), or never written. Zero-length counts as missing. Test count 1494 -> 1504. Four red-proofs planted, four seen failing; the two halves of R-384 convict independently.
470 lines
20 KiB
Go
470 lines
20 KiB
Go
package stacks
|
|
|
|
import (
|
|
"fmt"
|
|
"io"
|
|
"log"
|
|
"strings"
|
|
"sync"
|
|
"testing"
|
|
|
|
"gitea.dooplex.hu/admin/felhom-controller/internal/config"
|
|
)
|
|
|
|
// R-51 (v0.156.0). The live defect these tests pin: on 2026-07-20 `immich-server` sat Exited for
|
|
// 18 hours behind three running helpers, the stack aggregated to StateRunning ("partial"), and
|
|
// because StateRunning is not a down state NOTHING fired — no dashboard banner, no
|
|
// `app_start_failed` hub event — while single-container Calibre-Web, down for the same reason,
|
|
// alerted in 90 s. Evidence: felhom.eu/documentation/audits/AUDIT-vacation-remote-ops-2026-07-20.md
|
|
// finding F4.
|
|
//
|
|
// RED-PROOF (recorded in REPORT.md): with the mix branch reverted to its pre-v0.156.0 body
|
|
//
|
|
// if running > 0 { return StateRunning }
|
|
//
|
|
// TestAggregateState_DeadSupervisedMemberIsDegraded fails with
|
|
// "aggregateState = running, want degraded", which is exactly the shape the audit observed.
|
|
|
|
// immichLike is the F4 fixture: the primary Exited, the helpers up.
|
|
func immichLike() []ContainerInfo {
|
|
return []ContainerInfo{
|
|
{Name: "immich-server", State: StateExited, Status: "Exited (137) 18 hours ago"},
|
|
{Name: "immich-machine-learning", State: StateRunning, Status: "Up 18 hours"},
|
|
{Name: "immich-redis", State: StateRunning, Status: "Up 18 hours (healthy)"},
|
|
{Name: "immich-postgres", State: StateRunning, Status: "Up 18 hours (healthy)"},
|
|
}
|
|
}
|
|
|
|
// policyMap builds a lookup over a name→policy table; an unlisted name reads as UNKNOWN ("").
|
|
func policyMap(t *testing.T, m map[string]string) restartPolicyLookup {
|
|
t.Helper()
|
|
return func(name string) string { return m[name] }
|
|
}
|
|
|
|
func TestAggregateState_DeadSupervisedMemberIsDegraded(t *testing.T) {
|
|
got := aggregateState(immichLike(), policyMap(t, map[string]string{
|
|
"immich-server": "unless-stopped",
|
|
"immich-machine-learning": "unless-stopped",
|
|
"immich-redis": "unless-stopped",
|
|
"immich-postgres": "unless-stopped",
|
|
}))
|
|
if got != StateDegraded {
|
|
t.Fatalf("aggregateState = %q, want %q (a dead supervised primary must not read as running)", got, StateDegraded)
|
|
}
|
|
if !IsDownState(got) {
|
|
t.Fatalf("IsDownState(%q) = false — the whole point of R-51 is that this state alerts", got)
|
|
}
|
|
}
|
|
|
|
// Scenario B: a one-shot init/migrate container that has legitimately finished must NOT alarm.
|
|
func TestAggregateState_OneShotExitedMemberIsBenign(t *testing.T) {
|
|
for _, policy := range []string{"no", "on-failure", ""} {
|
|
name := policy
|
|
if name == "" {
|
|
name = "(absent)"
|
|
}
|
|
t.Run(name, func(t *testing.T) {
|
|
containers := []ContainerInfo{
|
|
{Name: "app-migrate", State: StateExited, Status: "Exited (0) 2 minutes ago"},
|
|
{Name: "app-web", State: StateRunning, Status: "Up 2 minutes"},
|
|
}
|
|
got := aggregateState(containers, policyMap(t, map[string]string{
|
|
"app-migrate": policy,
|
|
"app-web": "unless-stopped",
|
|
}))
|
|
want := StateRunning
|
|
if policy == "" {
|
|
// UNKNOWN is deliberately fail-CLOSED — see supervisedPolicy. An absent policy in
|
|
// the compose file resolves to Docker's "no" at inspect time, so the "" case here
|
|
// is the INSPECT-FAILED case, not the no-restart-policy case.
|
|
want = StateDegraded
|
|
}
|
|
if got != want {
|
|
t.Fatalf("policy %q: aggregateState = %q, want %q", policy, got, want)
|
|
}
|
|
})
|
|
}
|
|
}
|
|
|
|
// The unchanged branches: R-51 must not move any state the pre-existing aggregation produced.
|
|
//
|
|
// R-384 (v0.222.0) AMENDED THREE OF THESE CASES, and the amendment is the fix, not an accommodation.
|
|
// They read `{a: unhealthy|starting|restarting, b: EXITED}` with BOTH members on `unless-stopped`,
|
|
// and asserted that the live member's state won. That is precisely the defect R-384 closes: member
|
|
// `b` is a dead SUPERVISED container, and the case was pinning the wrong answer as if it were
|
|
// settled. A dead database beside an unhealthy front end read `unhealthy`, which is not a down
|
|
// state, so nothing ever alarmed — measured live on `demo-hp` 2026-08-22.
|
|
//
|
|
// The cases keep their ORIGINAL INTENT — "the live member's state still wins over a down member" —
|
|
// by giving `b` a BENIGN policy, which is the only situation in which that sentence was ever true.
|
|
// The supervised versions moved to TestR384_DeadSupervisedMemberIsAskedAboutFirst, where they now
|
|
// assert `degraded`.
|
|
func TestAggregateState_UnchangedBranches(t *testing.T) {
|
|
all := policyMap(t, map[string]string{"a": "unless-stopped", "b": "unless-stopped"})
|
|
// `b` is a one-shot that finished: the live member's state must still win over it.
|
|
benignB := policyMap(t, map[string]string{"a": "unless-stopped", "b": "no"})
|
|
cases := []struct {
|
|
name string
|
|
containers []ContainerInfo
|
|
policy restartPolicyLookup
|
|
want ContainerState
|
|
}{
|
|
{"no containers", nil, all, StateNotDeployed},
|
|
{"all running", []ContainerInfo{{Name: "a", State: StateRunning}, {Name: "b", State: StateRunning}}, all, StateRunning},
|
|
{"all stopped", []ContainerInfo{{Name: "a", State: StateExited}, {Name: "b", State: StateStopped}}, all, StateStopped},
|
|
{"any unhealthy wins over a BENIGN exited", []ContainerInfo{{Name: "a", State: StateUnhealthy}, {Name: "b", State: StateExited}}, benignB, StateUnhealthy},
|
|
{"any starting wins over a BENIGN exited", []ContainerInfo{{Name: "a", State: StateStarting}, {Name: "b", State: StateExited}}, benignB, StateStarting},
|
|
{"any restarting wins over a BENIGN exited", []ContainerInfo{{Name: "a", State: StateRestarting}, {Name: "b", State: StateExited}}, benignB, StateRestarting},
|
|
{"single container exited", []ContainerInfo{{Name: "a", State: StateExited}}, all, StateStopped},
|
|
}
|
|
for _, tc := range cases {
|
|
t.Run(tc.name, func(t *testing.T) {
|
|
if got := aggregateState(tc.containers, tc.policy); got != tc.want {
|
|
t.Fatalf("aggregateState = %q, want %q", got, tc.want)
|
|
}
|
|
})
|
|
}
|
|
}
|
|
|
|
// R-384 (v0.222.0) — the ORDER was the defect, not the `unhealthy` exclusion.
|
|
//
|
|
// THE LIVE FAILURE THIS PINS. On `demo-hp` 2026-08-22 `bookstack-db` (MariaDB, `unless-stopped`) was
|
|
// stopped at 21:27:01. Its front end went `unhealthy` seconds later because it could not reach its
|
|
// database. `aggregateState` returned StateUnhealthy at the `unhealthy > 0` line and never reached
|
|
// the R-51 mixed-case block, so `classifyRunStates` never marked the app down and the F-OBS heartbeat
|
|
// printed "0 currently down" throughout. The dead database hid behind the unhealthy survivor.
|
|
//
|
|
// Two things had to move, and both are asserted here:
|
|
// 1. the supervised-down test runs BEFORE the unhealthy/starting/restarting returns;
|
|
// 2. "some members are up" counts ANY member not in the down bucket. The old guard was `running > 0`
|
|
// counting StateRunning alone, which made the block unreachable in exactly this case — an
|
|
// unhealthy survivor counted as nothing up.
|
|
//
|
|
// RED-PROOF (observed, see REPORT.md): move the hoisted block back below the `unhealthy > 0` return
|
|
// and the `unhealthy survivor` subtest fails with `aggregateState = "unhealthy", want "degraded"` —
|
|
// the exact state the live box reported. Narrowing `up` back to `running` alone fails the same
|
|
// subtest identically, so BOTH halves of the fix are convicted.
|
|
func TestR384_DeadSupervisedMemberIsAskedAboutFirst(t *testing.T) {
|
|
supervised := policyMap(t, map[string]string{"web": "unless-stopped", "db": "unless-stopped"})
|
|
cases := []struct {
|
|
name string
|
|
survivor ContainerState
|
|
}{
|
|
{"unhealthy survivor — the bookstack case measured live", StateUnhealthy},
|
|
{"starting survivor", StateStarting},
|
|
{"restarting survivor", StateRestarting},
|
|
{"running survivor — R-51's original case, must not regress", StateRunning},
|
|
}
|
|
for _, tc := range cases {
|
|
t.Run(tc.name, func(t *testing.T) {
|
|
got := aggregateState([]ContainerInfo{
|
|
{Name: "web", State: tc.survivor, Status: "Up 3 minutes (unhealthy)"},
|
|
{Name: "db", State: StateExited, Status: "Exited (0) 3 minutes ago"},
|
|
}, supervised)
|
|
if got != StateDegraded {
|
|
t.Fatalf("survivor %q beside a dead SUPERVISED member: aggregateState = %q, want %q "+
|
|
"— a dead database must not hide behind it", tc.survivor, got, StateDegraded)
|
|
}
|
|
// Non-hollow: the state is only worth anything if it actually alarms.
|
|
if !IsDownState(got) {
|
|
t.Fatalf("IsDownState(%q) = false — the state changed but nothing downstream alarms", got)
|
|
}
|
|
})
|
|
}
|
|
}
|
|
|
|
// The fenced act (§12): `IsDownState` must stay byte-identical. R-384 works by asking a PRIOR
|
|
// question, never by folding a running-but-failing container into the down set — that is the
|
|
// flapping the exclusion exists to stop. This is the consequence-level guard: if a future change
|
|
// widens IsDownState instead of reordering, this fails even though R-384's own tests still pass.
|
|
func TestR384_IsDownStateWasNotWidened(t *testing.T) {
|
|
for _, s := range []ContainerState{StateUnhealthy, StateRestarting, StateStarting, StatePaused, StateUnknown, StateDeploying} {
|
|
if IsDownState(s) {
|
|
t.Fatalf("IsDownState(%q) = true — R-384 must reorder the question, never widen the down set", s)
|
|
}
|
|
}
|
|
for _, s := range []ContainerState{StateStopped, StateExited, StateDegraded} {
|
|
if !IsDownState(s) {
|
|
t.Fatalf("IsDownState(%q) = false — R-384 must not narrow the down set either", s)
|
|
}
|
|
}
|
|
}
|
|
|
|
// Scenario C, at the layer the defect would live: a migration step that finished must never alarm,
|
|
// and R-384 hoisted the code that decides it. Every catalog app with an init container passes through
|
|
// this, so a regression here alarms fleet-wide on every start.
|
|
func TestR384_FinishedOneShotStaysBenignBehindAnUnhealthyMember(t *testing.T) {
|
|
for _, policy := range []string{"no", "on-failure"} {
|
|
t.Run(policy, func(t *testing.T) {
|
|
got := aggregateState([]ContainerInfo{
|
|
{Name: "app-web", State: StateUnhealthy, Status: "Up 1 minute (unhealthy)"},
|
|
{Name: "app-migrate", State: StateExited, Status: "Exited (0) 1 minute ago"},
|
|
}, policyMap(t, map[string]string{"app-web": "unless-stopped", "app-migrate": policy}))
|
|
if got != StateUnhealthy {
|
|
t.Fatalf("policy %q: aggregateState = %q, want %q — a finished migration is benign "+
|
|
"and must not turn an unhealthy app into a down one", policy, got, StateUnhealthy)
|
|
}
|
|
if IsDownState(got) {
|
|
t.Fatalf("policy %q: a finished one-shot made the stack alarm", policy)
|
|
}
|
|
})
|
|
}
|
|
}
|
|
|
|
// Scenario B, unchanged and byte-identical: unhealthy with NOTHING dead must stay unhealthy and must
|
|
// not alarm. This is the flapping case the IsDownState exclusion exists to stop, and re-creating it
|
|
// would be the over-correction rather than the fix.
|
|
func TestR384_UnhealthyWithNothingDeadDoesNotAlarm(t *testing.T) {
|
|
for _, containers := range [][]ContainerInfo{
|
|
{{Name: "solo", State: StateUnhealthy}},
|
|
{{Name: "a", State: StateUnhealthy}, {Name: "b", State: StateUnhealthy}},
|
|
{{Name: "a", State: StateUnhealthy}, {Name: "b", State: StateRunning}},
|
|
} {
|
|
got := aggregateState(containers, policyMap(t, map[string]string{"solo": "unless-stopped", "a": "unless-stopped", "b": "unless-stopped"}))
|
|
if got != StateUnhealthy {
|
|
t.Fatalf("%d container(s), none down: aggregateState = %q, want %q", len(containers), got, StateUnhealthy)
|
|
}
|
|
if IsDownState(got) {
|
|
t.Fatalf("%d container(s), none down: the stack alarmed", len(containers))
|
|
}
|
|
}
|
|
}
|
|
|
|
// Scenario F: every member down. R-384 is a MEASUREMENT here, not a change (§4 of the task) — the
|
|
// all-down path must be byte-identical, so that whatever the live walk finds about single-container
|
|
// crashes is a separate, later decision and not something this change quietly moved.
|
|
func TestR384_AllMembersDownIsUnchanged(t *testing.T) {
|
|
supervised := policyMap(t, map[string]string{"a": "unless-stopped", "b": "unless-stopped"})
|
|
cases := []struct {
|
|
name string
|
|
containers []ContainerInfo
|
|
}{
|
|
{"single exited", []ContainerInfo{{Name: "a", State: StateExited}}},
|
|
{"single stopped", []ContainerInfo{{Name: "a", State: StateStopped}}},
|
|
{"both down", []ContainerInfo{{Name: "a", State: StateExited}, {Name: "b", State: StateStopped}}},
|
|
}
|
|
for _, tc := range cases {
|
|
t.Run(tc.name, func(t *testing.T) {
|
|
if got := aggregateState(tc.containers, supervised); got != StateStopped {
|
|
t.Fatalf("aggregateState = %q, want %q — R-384 must not move the all-down path", got, StateStopped)
|
|
}
|
|
})
|
|
}
|
|
}
|
|
|
|
// The unhealthy/restarting/paused/unknown exclusions are the fix-3 contract (downstate_test.go owns
|
|
// them). This asserts the ONE addition, so a future reader can see R-51 widened the set by exactly
|
|
// one state and by nothing else.
|
|
func TestIsDownState_DegradedIsTheOnlyAddition(t *testing.T) {
|
|
if !IsDownState(StateDegraded) {
|
|
t.Fatalf("IsDownState(degraded) = false, want true")
|
|
}
|
|
for _, s := range []ContainerState{StateUnhealthy, StateRestarting, StatePaused, StateUnknown} {
|
|
if IsDownState(s) {
|
|
t.Fatalf("IsDownState(%q) = true — R-51 must not touch the fix-3 exclusions", s)
|
|
}
|
|
}
|
|
}
|
|
|
|
// --- production-path wiring test (§9 rule 6) ---------------------------------------------------
|
|
//
|
|
// Proves the chain the box actually runs: RefreshStatus → docker ps → aggregateState → docker
|
|
// inspect. An aggregateState-only test proves the function, not the caller.
|
|
|
|
type scriptedDocker struct {
|
|
mu sync.Mutex
|
|
ps string
|
|
policies map[string]string
|
|
inspects []string // every container name inspected, in order
|
|
}
|
|
|
|
func (s *scriptedDocker) exec(name string, args ...string) (string, error) {
|
|
s.mu.Lock()
|
|
defer s.mu.Unlock()
|
|
if name != "docker" {
|
|
return "", fmt.Errorf("unexpected command %q", name)
|
|
}
|
|
switch {
|
|
case len(args) > 0 && args[0] == "ps":
|
|
return s.ps, nil
|
|
case len(args) > 0 && args[0] == "inspect":
|
|
target := args[len(args)-1]
|
|
s.inspects = append(s.inspects, target)
|
|
p, ok := s.policies[target]
|
|
if !ok {
|
|
return "", fmt.Errorf("no such container: %s", target)
|
|
}
|
|
return p + "\n", nil
|
|
}
|
|
return "", fmt.Errorf("unexpected docker args %v", args)
|
|
}
|
|
|
|
func (s *scriptedDocker) inspectCount() int {
|
|
s.mu.Lock()
|
|
defer s.mu.Unlock()
|
|
return len(s.inspects)
|
|
}
|
|
|
|
func psLine(name, state, status, project string) string {
|
|
return strings.Join([]string{name, "img:1", state, status, project}, "\t")
|
|
}
|
|
|
|
func TestRefreshStatus_WiresDegradedThroughTheRealPath(t *testing.T) {
|
|
dock := &scriptedDocker{
|
|
ps: strings.Join([]string{
|
|
psLine("immich-server", "exited", "Exited (137) 18 hours ago", "immich"),
|
|
psLine("immich-redis", "running", "Up 18 hours (healthy)", "immich"),
|
|
psLine("calibre-web", "running", "Up 18 hours", "calibre-web"),
|
|
}, "\n"),
|
|
policies: map[string]string{"immich-server": "unless-stopped"},
|
|
}
|
|
m := &Manager{
|
|
cfg: &config.Config{},
|
|
logger: log.New(io.Discard, "", 0),
|
|
execFn: dock.exec,
|
|
stacks: map[string]*Stack{
|
|
"immich": {Name: "immich", Deployed: true},
|
|
"calibre-web": {Name: "calibre-web", Deployed: true},
|
|
},
|
|
}
|
|
|
|
if err := m.RefreshStatus(); err != nil {
|
|
t.Fatalf("RefreshStatus: %v", err)
|
|
}
|
|
if got := m.stacks["immich"].State; got != StateDegraded {
|
|
t.Fatalf("immich state = %q, want %q (the F4 shape must reach the stack map)", got, StateDegraded)
|
|
}
|
|
if got := m.stacks["calibre-web"].State; got != StateRunning {
|
|
t.Fatalf("calibre-web state = %q, want %q — a healthy app must be untouched", got, StateRunning)
|
|
}
|
|
|
|
// Only the DOWN member of the MIXED stack is inspected: never the running members, never the
|
|
// healthy stack. An inspect per container per 10 s refresh would be a real docker load.
|
|
if n := dock.inspectCount(); n != 1 {
|
|
t.Fatalf("docker inspect called %d times, want exactly 1 (%v)", n, dock.inspects)
|
|
}
|
|
|
|
// Second refresh: the answer comes from the cache, so the inspect count must NOT move.
|
|
if err := m.RefreshStatus(); err != nil {
|
|
t.Fatalf("RefreshStatus (2nd): %v", err)
|
|
}
|
|
if n := dock.inspectCount(); n != 1 {
|
|
t.Fatalf("docker inspect called %d times after a second refresh, want 1 — the cache is not being used", n)
|
|
}
|
|
if got := m.stacks["immich"].State; got != StateDegraded {
|
|
t.Fatalf("immich state after 2nd refresh = %q, want %q", got, StateDegraded)
|
|
}
|
|
}
|
|
|
|
// A container that vanishes must not leave its policy behind — an unbounded cache in a process that
|
|
// runs for months is a slow leak, and a stale entry would answer for a recreated container.
|
|
func TestRestartPolicyCache_PrunesVanishedContainers(t *testing.T) {
|
|
dock := &scriptedDocker{
|
|
ps: strings.Join([]string{
|
|
psLine("app-init", "exited", "Exited (0) 1 minute ago", "app"),
|
|
psLine("app-web", "running", "Up 1 minute", "app"),
|
|
}, "\n"),
|
|
policies: map[string]string{"app-init": "no"},
|
|
}
|
|
m := &Manager{
|
|
cfg: &config.Config{},
|
|
logger: log.New(io.Discard, "", 0),
|
|
execFn: dock.exec,
|
|
stacks: map[string]*Stack{"app": {Name: "app", Deployed: true}},
|
|
}
|
|
if err := m.RefreshStatus(); err != nil {
|
|
t.Fatalf("RefreshStatus: %v", err)
|
|
}
|
|
if got := m.stacks["app"].State; got != StateRunning {
|
|
t.Fatalf("state = %q, want running (a finished one-shot must not alarm)", got)
|
|
}
|
|
if len(m.restartPolicyCache) != 1 {
|
|
t.Fatalf("cache size = %d, want 1", len(m.restartPolicyCache))
|
|
}
|
|
|
|
// The one-shot container is reaped; only the web container remains.
|
|
dock.ps = psLine("app-web", "running", "Up 5 minutes", "app")
|
|
if err := m.RefreshStatus(); err != nil {
|
|
t.Fatalf("RefreshStatus (2nd): %v", err)
|
|
}
|
|
if len(m.restartPolicyCache) != 0 {
|
|
t.Fatalf("cache size = %d after the container vanished, want 0: %v", len(m.restartPolicyCache), m.restartPolicyCache)
|
|
}
|
|
}
|
|
|
|
// An inspect failure must not silence the alarm — see supervisedPolicy's fail-closed rationale.
|
|
func TestRefreshStatus_InspectFailureStillDegrades(t *testing.T) {
|
|
dock := &scriptedDocker{
|
|
ps: strings.Join([]string{
|
|
psLine("immich-server", "exited", "Exited (137) 1 hour ago", "immich"),
|
|
psLine("immich-redis", "running", "Up 1 hour", "immich"),
|
|
}, "\n"),
|
|
policies: map[string]string{}, // every inspect fails
|
|
}
|
|
m := &Manager{
|
|
cfg: &config.Config{},
|
|
logger: log.New(io.Discard, "", 0),
|
|
execFn: dock.exec,
|
|
stacks: map[string]*Stack{"immich": {Name: "immich", Deployed: true}},
|
|
}
|
|
if err := m.RefreshStatus(); err != nil {
|
|
t.Fatalf("RefreshStatus: %v", err)
|
|
}
|
|
if got := m.stacks["immich"].State; got != StateDegraded {
|
|
t.Fatalf("state = %q, want %q — an unreadable policy must not lose the alarm", got, StateDegraded)
|
|
}
|
|
// A failed inspect is deliberately NOT cached, so the next cycle retries.
|
|
if len(m.restartPolicyCache) != 0 {
|
|
t.Fatalf("failed inspect was cached: %v", m.restartPolicyCache)
|
|
}
|
|
}
|
|
|
|
// R-384 production-path wiring: the bookstack shape, through RefreshStatus.
|
|
//
|
|
// This is the SAME chain the box runs — docker ps → aggregateState → docker inspect → the stack map —
|
|
// with the fixture measured live on `demo-hp` 2026-08-22: `bookstack-db` stopped, `bookstack` itself
|
|
// still up but failing its healthcheck because it cannot reach its database. An aggregateState-only
|
|
// test proves the function and not the caller, and this defect lived in the caller's reading of it.
|
|
//
|
|
// The classifier half of the chain is pinned by
|
|
// TestR384_ADeadDatabaseBehindAnUnhealthyAppAlarms in cmd/controller — this test ends at the state,
|
|
// that one starts from it, and TestR384_IsDownStateWasNotWidened pins the join.
|
|
//
|
|
// RED-PROOF (observed): with the hoisted block moved back below the `unhealthy > 0` return, this
|
|
// fails with `bookstack state = "unhealthy", want "degraded"`.
|
|
func TestR384_WiresTheDeadDatabaseThroughTheRealPath(t *testing.T) {
|
|
dock := &scriptedDocker{
|
|
ps: strings.Join([]string{
|
|
psLine("bookstack", "running", "Up 3 minutes (unhealthy)", "bookstack"),
|
|
psLine("bookstack-db", "exited", "Exited (0) 3 minutes ago", "bookstack"),
|
|
psLine("docmost", "running", "Up 3 hours (healthy)", "docmost"),
|
|
}, "\n"),
|
|
policies: map[string]string{"bookstack-db": "unless-stopped"},
|
|
}
|
|
m := &Manager{
|
|
cfg: &config.Config{},
|
|
logger: log.New(io.Discard, "", 0),
|
|
execFn: dock.exec,
|
|
stacks: map[string]*Stack{
|
|
"bookstack": {Name: "bookstack", Deployed: true},
|
|
"docmost": {Name: "docmost", Deployed: true},
|
|
},
|
|
}
|
|
if err := m.RefreshStatus(); err != nil {
|
|
t.Fatalf("RefreshStatus: %v", err)
|
|
}
|
|
got := m.stacks["bookstack"].State
|
|
if got != StateDegraded {
|
|
t.Fatalf("bookstack state = %q, want %q — a dead database must not hide behind its own "+
|
|
"unhealthy front end (measured live 2026-08-22: this read %q and nothing alarmed)",
|
|
got, StateDegraded, StateUnhealthy)
|
|
}
|
|
// Non-hollow: the state only matters because it alarms.
|
|
if !IsDownState(got) {
|
|
t.Fatalf("IsDownState(%q) = false — the state moved but the app is still silently down", got)
|
|
}
|
|
if got := m.stacks["docmost"].State; got != StateRunning {
|
|
t.Fatalf("docmost state = %q, want %q — a healthy app must be untouched", got, StateRunning)
|
|
}
|
|
}
|