Files
felhom.eu/documentation/tests/test-campaign-2-finding1-recovery-diagnosis.md
admin e4f49491a3 docs(tests): Finding #1 re-diagnosis — TEST ARTIFACT (docker kill != crash); controller auto-recovers
Read-only diagnosis: restart-manager works (natural-exit probe RestartCount climbs;
manualRestart=false); docker kill is a manual stop so restart is correctly suppressed.
Definitive: kill -9 the controller PID (real crash) -> unless-stopped auto-restarts it,
RestartCount=1 healthy in 5s. Daemon stable (NRestarts=0, no OOM). Finding #1 downgraded
in the campaign-2 doc. Separate minor obs: boot-time drive-ordering race. No fix applied.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 19:35:08 +02:00

6.0 KiB
Raw Permalink Blame History

Finding #1 diagnosis — does a crashed controller auto-recover? — 2026-06-22 (demo 9201)

Status: read-only diagnosis (no fix applied). Resolves campaign-#2 Finding #1 ("a crashed felhom-controller has no automatic recovery"). Verdict: Finding #1 is a TEST ARTIFACT. The controller's restart: unless-stopped policy does auto-recover it from a genuine crash — proven directly. The campaign tested with docker kill, which Docker treats as a manual stop (restart policy intentionally suppressed); that is not a crash. No real recovery gap. Downgrade Finding #1.

TL;DR

  • Docker's live restart-manager works on this host under a healthy daemon: a container that exits on its own is auto-restarted (probe --restart always, sleep 3 → RestartCount climbs 1→2→3→4; dockerd logs restarting container manualRestart=false restartPolicy="{always 0}").
  • docker kill is treated as a manual stop → restart policy is correctly suppressed (probe → after docker kill: Running=false RestartCount=0, no restart log). This is what the campaign's C1/C2 hit.
  • Definitive controller test: kill -9 the controller's container PID directly (a real process death, not the docker API) → docker auto-restarted it: RestartCount=1, health=healthy within 5 s.
  • Daemon is stable (NRestarts=0, up since the C8 reboot 17:01 UTC, no OOM, 8.7 GB available) — so this is not a daemon-state artifact and not a genuine restart-policy bug; it's the wrong-signal test.

Evidence

Step 1 — daemon stability

  • docker version: Server 29.5.3, Go 1.26.4, containerd v2.2.4, API 1.54. (Confirmed the unusual "29.5.3" — a very new/edge build; relevant only as context.)
  • docker info: overlay2, cgroup v2 (systemd), runc, Live Restore: false.
  • systemctl show docker: NRestarts=0, ExecMainStart 2026-06-22 17:01:03 UTC (= the C8 reboot). dockerd has NOT restarted since; daemon is stable.
  • The campaign's "stopping restart-manager" lines cluster at 17:01 (post-reboot bring-up) and 17:0517:08 (my C6/C7 drive-disconnect/quiesce tests, which stop containers). That log line is emitted whenever a container's restart-manager is cancelled (container stop or daemon shutdown) — it is not proof of a daemon restart. The daemon did not flap.

Step 2 — memory / OOM

free -h: 12 Gi total, 8.7 Gi available, swap unused. No oom/killed process in dmesg/journal. Memory pressure is not a factor.

Step 3 — decisive throwaway probe (healthy daemon)

probe after docker kill
--restart unless-stopped Running=false RestartCount=0 (no restart)
--restart unless-stopped (2nd kill) Running=false RestartCount=0
--restart always Running=false RestartCount=0 (no restart)

Both policies fail to restart after docker kill — which first looked like a host-wide bug…

Step 3b — natural-exit test (the disambiguator)

docker run --restart always busybox sh -c 'sleep 3' (exits on its own): → Running=true RestartCount=3 (and climbing). dockerd: restarting container … manualRestart=false restartPolicy="{always 0}" repeatedly. The restart-manager works for real exits. So docker kill is being recorded as a manual stop (manualRestart path), which Docker deliberately does not auto-restart — by design, an operator who stop/kills a container wants it stopped.

Step 6 — definitive controller test (real crash, not the docker API)

kill -9 <controller container PID 1878> (process death, bypassing docker stop/kill):

before:  Running=true  RestartCount=0  policy=unless-stopped
t+5s:    Running=true  RestartCount=1  health=healthy
… stable healthy thereafter

The controller auto-recovered from a real crash in ~5 s, with the data plane (apps + tunnel) unaffected throughout. (A 60 s timed auto-recovery guard was pre-armed and not needed; cleaned up.)

Step 4 — the real recovery chain

  • Running controller: RestartPolicy = unless-stopped (confirmed).
  • felhom-controller-bootstrap.service: Type=oneshot, no Restart= — provision/boot-time bring-up only (correct). No other unit watches the container.
  • What recovers a crashed controller: docker's restart policy — the live restart-manager on a crash (proven) and boot-time restore at daemon start (which is why campaign C8's reboot brought 26 containers back). Two independent, working paths. No secondary fallback is needed because the primary (restart policy) works.

Classification

Finding #1 = TEST ARTIFACT. Restart policy works; the campaign's docker kill is a manual stop, not a crash, so it correctly did not auto-restart. A genuine controller crash self-heals via unless-stopped. Downgrade Finding #1 (it is not a resilience gap).

Fix direction

  • No fix required for crash-recovery. The controller self-heals. Do not add a watchdog on this basis — it would patch a non-problem.
  • Optional, low priority: if defense-in-depth against an operator/tooling docker stop/kill (which intentionally won't auto-restart) is ever wanted, a tiny "controller absent ⇒ recreate" systemd timer could be considered — but it is not warranted by any observed failure.
  • Re-test hygiene (the actual lesson): simulate a crash with a real process death (kill -9 the PID, OOM, panic) — never docker kill/docker stop — when testing restart policies.

Separate observation (not Finding #1) — boot-time drive-ordering race

At the C8 reboot (17:01:05) dockerd logged many failed to start container … mkdir /mnt/felhom-drives/felhom-flash/userdata: permission denied — docker's boot-restore started drive-backed containers before the drives were mounted. They recovered shortly after (the controller's recreate-on-boot once the drives mounted; C8 ended at 26 running/0 unhealthy). Real but minor (transient boot race); worth a follow-up to order drive-backed app start after the mount converges. No change applied.