15 KiB
REPORT — TASK-D Parts 1–2: dead-primary alerting (R-51) + boot reconciliation (R-52)
Controller v0.155.0 → v0.156.0. Session 2026-07-21, on DooPlex as kisfenyo.
Status: SHIPPED, deployed via the floor save, and STOP-1 RAN — both legs passed (§6). P1 is
answered (§4). The validation also surfaced a real pre-existing gap, filed as R-55 (§6). Nothing
below claims live behaviour that was not observed.
1. Baselines (re-confirmed at session start)
| Repo | main @ start |
clean | now |
|---|---|---|---|
| felhom-controller | 0f9b29a |
yes | 285dd10 (v0.156.0) |
| felhom-agent | 08b55a1 |
yes | 98adb72 (v0.92.1) |
| felhom.eu | 50a7ffa |
yes | docs commit (see §9) |
Image: gitea.dooplex.hu/admin/felhom-controller:0.156.0, pushed and never hand-deployed — the
operator's floor save deployed it, which banked the R-23 single-fire datapoint for free:
update successful (0.155.0 → 0.156.0) then Current version 0.156.0 is up to date. 9201 runs
0.156.0.
2. The defect R-51 actually fixes (the ROADMAP row was wrong)
The row says aggregation classifies a dead-primary stack unhealthy, and that IsDownState's
deliberate unhealthy exclusion is therefore what suppresses the alarm. That is not what the code
did. aggregateState's final branch was:
// Mix (some running, some stopped) — report as running (partial)
if running > 0 { return StateRunning }
So the immich stack read running, not unhealthy. IsDownState was never consulted about
unhealthy at all, and the constraint the row protects was never in tension with the fix. The
correction is recorded in the ROADMAP row itself.
The fix. New StateDegraded. The mixed branch asks each DOWN member for its docker restart
policy: always / unless-stopped → docker was supposed to be keeping it up → the stack is
degraded; no / on-failure → a finished one-shot init/migrate container → benign, stack stays
running. IsDownState gains degraded and nothing else — downstate_test.go is untouched and
green, which is the fix-3 contract.
An unreadable policy counts as supervised — fail-CLOSED, deliberately the opposite of
IsDownState's documented fail-open. They are different questions: there the state is ambiguous,
here a member is known dead and only its excuse is missing. P2 backs it (§4).
Blast radius, decided per call site (StateDegraded is a stack-level aggregate and never a
container state, so container-level code is untouched by construction):
| Question the caller asks | Sites | Degraded counts as |
|---|---|---|
| "are there live containers to stop?" | quiesce RunningAppStacks, delete stop-first guard ×2, export IsStackRunning, telemetry, health-probe eligibility |
up |
| "is this app working?" | dashboard counter, stopped filter, dead-app banner + app_start_failed |
down |
| "did this come back healthy?" | waitHealthy (migration), post-restore RefreshAndIsRunning, integrations readiness |
not healthy (left strict) |
The health probe can no longer mask it: the probe result only ever overrides StateRunning, so a
failing probe on a degraded stack cannot downgrade it to unhealthy (which would silence the alarm
again). Asserted by the wiring test's second refresh.
3. R-52 — the recovery half
internal/bootrecon, one bounded sweep 5 s after startup (after the quiesce recovery, so the two
never race for the same stack): at most 2 attempts, 30 s apart, then it stops and the alarm owns
the problem. Never a loop. compose up -d exits 0 on a crash-loop, so success is decided by a fresh
RefreshStatus, never by a nil error.
The safety argument is the container gate. The UI's Stop is compose down, which REMOVES the
containers; a boot interruption leaves them behind as Exited. So "deployed, has containers, and
they are down" is the boot-orphan signature, and a zero-container stack is never touched.
Worst case 5 s + 30 s = 35 s < the 90 s deadAppBootGrace, so a successful recovery never alerts
and a failure alerts honestly. TestBootReconcileFitsInsideTheBootGrace asserts that arithmetic
instead of leaving it in a comment.
4. Phase-0 probes
P2 (R-51 one-shot census) — DONE, and the answer is clean. All 53 catalog templates,
78 services: every single one is unless-stopped. Zero services with no / on-failure /
absent. (The other 30 compose files under output/ are generated artifacts, not templates.) So the
restart-policy guard protects a case that does not exist in today's catalog — it is there for future
templates and for the fail-closed decision above, and it costs one cached docker inspect per down
member of a mixed stack.
P1 (why unless-stopped did not resurrect immich/calibre-web) — ANSWERED during STOP-1, and the
F5 hypothesis is CONFIRMED. It was never made a blocker (the reconciliation is correct either
way), and it fell out of the R-52 leg for free, as predicted. bookstack carries
restart=unless-stopped; the Docker daemon came up at ~10:53:15Z; the container's StartedAt is
10:53:28.05Z, the exact moment bootrecon's StartStack returned. Docker did not bring it
back — a container stopped before shutdown is recorded user-stopped and stays down. Details in §6.
P3/P4 were agent-side — see felhom-agent/REPORT.md.
5. Tests and red-proofs
Green gate: go build ./... && go vet ./... && go test ./... — all 25 packages ok (full run,
-count=1; internal/backup alone takes 174 s).
New: internal/stacks/degraded_test.go (8 cases incl. two production-path tests),
internal/bootrecon/bootrecon_test.go (8 cases), cmd/controller/bootrecon_wiring_test.go (4).
| # | Red-proof | Mutation | Result |
|---|---|---|---|
| A | dead primary must alert | mix branch → return StateRunning |
FAIL ×3 — aggregateState = "running", want "degraded"; both production-path tests report immich state = "running". Restored, green. |
| C | the boot sweep must be wired | go runBootReconcile(...) commented out in main() |
FAIL — func main() no longer starts the R-52 boot reconciliation. Restored, green. |
| D | a user-stopped app must stay stopped | zero-container gate dropped from isBootOrphan |
FAIL ×2 — StartStack("jellyfin") called 1 times, want 0 — a deliberate Stop must survive a reboot. Restored, green. |
A finding from red-proof C, worth more than the fix it guards. That wiring test was first
written as strings.Contains(src, "go runBootReconcile(ctx, stackMgr, logger)"). Its red-proof
passed — because commenting the call out leaves the string in the file. A test whose own
red-proof cannot fail it is decoration. Rewritten as an AST walk over func main()'s body for a
GoStmt calling runBootReconcile; the red-proof then failed as it must. Recorded in REUSE.md and
promoted into documentation/PROMPT-TEMPLATE.md §10 (Part-5 rider).
Seam discipline (§9 rule 6). Both features have a production-path test:
- R-51:
TestRefreshStatus_WiresDegradedThroughTheRealPathdrives the whole real chain —RefreshStatus→docker ps→aggregateState→docker inspect— through a newexecFnseam at the process boundary, and asserts the inspect COUNT (exactly 1: only the down member of the mixed stack, never the healthy stack; still 1 after a second refresh, proving the cache). - R-52: the sweep is asserted from
package main, plus the AST wiring test above.
6. Live validation — STOP-1 RAN, both legs PASSED (operator-present)
Method: the exact endpoints the UI calls, driven from inside guest 9201 through traefik
(--resolve felhom.demo-felhom.eu:443:127.0.0.1, session cookie + CSRF token), plus container-level
truth from docker ps -a and the controller's own log. No browser (none on DooPlex). The session
cookie jar was deleted from the guest afterwards.
Deploy first, per the resolved ordering: the operator saved the floor → 0.156.0, the box
self-updated, and the R-23 datapoint came free —
[selfupdate] Post-update startup: update successful (0.155.0 → 0.156.0) then
Current version 0.156.0 is up to date: one swap, no re-fire.
Leg A — R-51, the dead primary (12:50:40 → 12:51:52 CEST)
| Time | Event |
|---|---|
| 12:50:40 | docker stop immich-server → `exited |
| 12:50:53 | stack state = degraded (13 s). `immich-server -> exited |
| 10:51:11Z | exactly ONE Event pushed: app_start_failed (warn) — Telepített alkalmazás nem fut: Immich — single-fire verified by grep -c, = 1 |
| — | Dashboard: banner "Telepített alkalmazás nem fut: Immich (degraded)" + link Rendszermonitor, and the new „Részlegesen leállt" label present in the rendered page |
| 12:51:37 | docker start immich-server |
| 12:51:52 | state back to running; banner self-cleared (state-based, as designed) |
The (degraded) in the banner is the English-state wart predicted in §7 — visible now, pre-existing,
still a copy decision rather than a bug.
Leg B — R-52, the boot orphans (12:53:07 CEST reboot)
Fixture: docker stop on bookstack (2 containers) and calibre-web (containers left in
place); UI Stop on immich via POST /api/stacks/immich/stop (compose down → 0 containers).
Then pct reboot 9201.
10:53:22Z [bootrecon] Boot reconciliation: 1 boot-orphaned app(s) found: [bookstack] — up to 2 attempt(s)
10:53:28Z [bootrecon] attempt 1/2: started "bookstack" (took 6.0s)
10:53:28Z [bootrecon] complete: 1 app(s) recovered in 1 attempt(s): [bookstack]
Zero app_start_failed since the reboot — a successful recovery inside the boot grace is silent,
which was the design claim. And the quiet path is observable on the healthy box too: the earlier
0.156.0 boot logged Boot reconciliation: no boot-orphaned apps (nothing to start), so "nothing to
do" and "never ran" are distinguishable in production, not just in a test.
P1 — answered, and the F5 hypothesis CONFIRMED
bookstack carries restart=unless-stopped. The Docker daemon came up at ~10:53:15Z. The container's
StartedAt is 10:53:28.05Z — the exact moment bootrecon's StartStack returned. Docker's
own restart policy did not resurrect it: a container stopped before shutdown is recorded
user-stopped and stays down across the reboot. Only R-52 brought it back. That is the F5 hypothesis
the audit could not test, settled with direct evidence and at zero extra cost.
The finding this leg produced — new R-55 (it is the opposite of what the leg set out to prove)
Leg B was designed to show that a customer's deliberate Stop survives a reboot. It does not — for any drive-backed app. immich, stopped from the UI seconds earlier, came back running:
10:53:18Z [gate] boot …: live bind confirmed — recreating drive-backed app immich (state=stopped) onto /mnt/felhom-drives/hdd_1
10:53:18Z [stacks] Stopping stack: immich → Starting stack: immich
internal/web/intermediary.go's boot bind gate selects on cfg.Deployed && HDD_PATH alone and runs
StopStack + StartStack on every match; the stack's state is passed in for the log line only
and never consulted. calibre-web went the same way; bookstack, not drive-backed, fell through to
R-52 — which is why the sweep found exactly one orphan.
R-52 is not implicated. Its own gate behaved exactly as specified: immich, at zero containers, was never a candidate. But two things follow and need an operator ruling rather than a drive-by fix (the gate exists to fix a real bind-ordering bug and must keep doing that):
- A customer's Stop is silently undone on drive-backed apps at every reboot.
- R-52's practical scope on a real box is narrower than the brief assumed — the gate reaches every drive-backed app first, so the boot-orphan sweep only ever sees the remainder.
Filed as R-55; likely shape is that the gate should recreate only apps that were RUNNING at shutdown, i.e. the same running/stopped distinction R-52 already draws.
6b. What was NOT claimed live
Nothing beyond the above. What is verified without the box: the full green gate, all four design gates
(template_id_gate.py, emoji_gate.py, native_confirm_gate.py, offbox_rename_gate.py — all
OK), and the image build+push.
7. UI surfaces (Hungarian)
New state label „Részlegesen leállt", warn colour token, ◑ icon (shared with unhealthy),
counted with the stopped apps, filtered into the stopped bucket, and routeUnpublished → true (when
the dead member is the routed one Traefik withholds the route and the URL 404s — the card must not
imply the app is reachable). No new customer strings were needed for the alarm itself: the
existing app_start_failed sentence („Telepített alkalmazás nem fut: …") is reused unchanged, which
is why R-51 ships without touching the notifier.
One honest wart carried forward, not introduced: the dead-app banner appends the raw state in
English — it will now read (degraded) where it read (stopped). Pre-existing behaviour; changing
it is a copy decision, not part of this fix.
8. A workspace trap found the hard way (worth more than this task)
controller/.gitignore line 7 is controller — intended for the built binary. It also matches the
directory controller/cmd/controller/. Two consequences, both live today:
- ripgrep (and therefore the Grep tool) silently skips
cmd/controller/main.go— the entire production wiring file. Early in this session that produced a false "the whole fix-3 dead-app path has no production caller" reading — exactly the inert-seam conclusion this task is about.grep -rn --no-ignore/ plaingrep -rn(without--include) shows it. - New files under
cmd/controller/needgit add -f.bootrecon_wiring_test.gowas refused by a plaingit add— as part of a multi-path add, i.e. quietly. A wiring test that never reaches the repo is the same defect class it exists to prevent.
Recorded in the commit message and REUSE.md. Worth a .gitignore fix (anchored /controller) as
its own XS change — deliberately not smuggled into this task.
9. Deliverables
285dd10— v0.156.0 code + tests + CHANGELOG + REUSE + README.- Image
felhom-controller:0.156.0pushed; not deployed. - Docs in
felhom.eu: ROADMAP R-51/R-52 → SHIPPED (with the corrected R-51 diagnosis), new R-54 row, a new capability-map row (IMPLEMENTED — not PROVEN-LIVE, the live legs have not run), and the two Part-5 riders.
10. Operator actions outstanding
- STOP-1 (controller legs) — needs 0.156.0 live; see the ordering question in
felhom.eu/REPORT.md§5. - STOP-3 floor →
0.156.0, which is also what deploys it. - P1's answer falls out of STOP-1's reboot leg for free — worth recording either way.