docs(tests): Finding #1 re-diagnosis — TEST ARTIFACT (docker kill != crash); controller auto-recovers
Read-only diagnosis: restart-manager works (natural-exit probe RestartCount climbs; manualRestart=false); docker kill is a manual stop so restart is correctly suppressed. Definitive: kill -9 the controller PID (real crash) -> unless-stopped auto-restarts it, RestartCount=1 healthy in 5s. Daemon stable (NRestarts=0, no OOM). Finding #1 downgraded in the campaign-2 doc. Separate minor obs: boot-time drive-ordering race. No fix applied. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -36,7 +36,17 @@ PBS `felhom-pbs:backup/ct/9201/2026-06-22T16:13:44Z` (verified). UI driven via a
|
||||
|
||||
## Findings (new)
|
||||
|
||||
### 🔴 1. docker restart policy not honored → no auto-recovery for a crashed controller
|
||||
### ✅ 1. [DOWNGRADED — TEST ARTIFACT] "docker restart policy not honored / no auto-recovery"
|
||||
> **Correction (2026-06-22, read-only re-diagnosis — see `test-campaign-2-finding1-recovery-diagnosis.md`):**
|
||||
> Finding #1 below was a **test artifact**. The controller **does** auto-recover from a genuine crash:
|
||||
> `kill -9` of the controller's container PID (a real process death) → docker `unless-stopped`
|
||||
> auto-restarted it (**RestartCount=1, healthy in ~5 s**). The campaign tested with `docker kill`, which
|
||||
> Docker treats as a **manual stop** (restart policy intentionally suppressed) — not a crash. A
|
||||
> natural-exit probe confirmed the restart-manager works (`restartCount` climbs, `manualRestart=false`).
|
||||
> Daemon is stable (NRestarts=0, no OOM). **No watchdog needed; no real recovery gap.** The original
|
||||
> (incorrect) text is kept below for the record.
|
||||
|
||||
#### (original, superseded) docker restart policy not honored → no auto-recovery for a crashed controller
|
||||
`restart: unless-stopped` containers do **not** auto-restart after `docker kill` (navidrome
|
||||
`RestartCount=0` over two trials; dockerd logged `stopping restart-manager`; Docker 29.5.3, Live
|
||||
Restore off). The controller does not reconcile-stop them (verified in logs/source). Consequences:
|
||||
|
||||
@@ -0,0 +1,94 @@
|
||||
# Finding #1 diagnosis — does a crashed controller auto-recover? — 2026-06-22 (demo 9201)
|
||||
|
||||
> **Status: read-only diagnosis (no fix applied).** Resolves campaign-#2 Finding #1 ("a crashed
|
||||
> felhom-controller has no automatic recovery"). **Verdict: Finding #1 is a TEST ARTIFACT.** The
|
||||
> controller's `restart: unless-stopped` policy **does** auto-recover it from a genuine crash — proven
|
||||
> directly. The campaign tested with `docker kill`, which Docker treats as a *manual* stop (restart
|
||||
> policy intentionally suppressed); that is not a crash. No real recovery gap. Downgrade Finding #1.
|
||||
|
||||
## TL;DR
|
||||
- Docker's **live restart-manager works** on this host under a healthy daemon: a container that exits
|
||||
**on its own** is auto-restarted (probe `--restart always`, `sleep 3` → RestartCount climbs 1→2→3→4;
|
||||
dockerd logs `restarting container manualRestart=false restartPolicy="{always 0}"`).
|
||||
- **`docker kill` is treated as a manual stop** → restart policy is correctly suppressed (probe → after
|
||||
`docker kill`: `Running=false RestartCount=0`, no restart log). This is what the campaign's C1/C2 hit.
|
||||
- **Definitive controller test:** `kill -9` the controller's container PID directly (a real process
|
||||
death, not the docker API) → docker **auto-restarted it: RestartCount=1, health=healthy within 5 s.**
|
||||
- Daemon is **stable** (NRestarts=0, up since the C8 reboot 17:01 UTC, no OOM, 8.7 GB available) — so
|
||||
this is not a daemon-state artifact and not a genuine restart-policy bug; it's the **wrong-signal
|
||||
test**.
|
||||
|
||||
## Evidence
|
||||
|
||||
### Step 1 — daemon stability
|
||||
- `docker version`: Server **29.5.3**, Go 1.26.4, containerd v2.2.4, API 1.54. (Confirmed the unusual
|
||||
"29.5.3" — a very new/edge build; relevant only as context.)
|
||||
- `docker info`: overlay2, cgroup v2 (systemd), runc, **Live Restore: false**.
|
||||
- `systemctl show docker`: **NRestarts=0**, ExecMainStart **2026-06-22 17:01:03 UTC** (= the C8 reboot).
|
||||
dockerd has NOT restarted since; daemon is stable.
|
||||
- The campaign's **"stopping restart-manager"** lines cluster at 17:01 (post-reboot bring-up) and
|
||||
17:05–17:08 (my C6/C7 drive-disconnect/quiesce tests, which *stop* containers). That log line is
|
||||
emitted whenever a container's restart-manager is cancelled (container stop **or** daemon shutdown) —
|
||||
it is **not** proof of a daemon restart. The daemon did not flap.
|
||||
|
||||
### Step 2 — memory / OOM
|
||||
`free -h`: 12 Gi total, **8.7 Gi available**, swap unused. No `oom`/`killed process` in dmesg/journal.
|
||||
Memory pressure is **not** a factor.
|
||||
|
||||
### Step 3 — decisive throwaway probe (healthy daemon)
|
||||
| probe | after `docker kill` |
|
||||
|---|---|
|
||||
| `--restart unless-stopped` | `Running=false RestartCount=0` (no restart) |
|
||||
| `--restart unless-stopped` (2nd kill) | `Running=false RestartCount=0` |
|
||||
| `--restart always` | `Running=false RestartCount=0` (no restart) |
|
||||
|
||||
Both policies fail to restart **after `docker kill`** — which first looked like a host-wide bug…
|
||||
|
||||
### Step 3b — natural-exit test (the disambiguator)
|
||||
`docker run --restart always busybox sh -c 'sleep 3'` (exits on its own):
|
||||
→ **`Running=true RestartCount=3`** (and climbing). dockerd:
|
||||
`restarting container … manualRestart=false restartPolicy="{always 0}"` repeatedly.
|
||||
**The restart-manager works for real exits.** So `docker kill` is being recorded as a **manual** stop
|
||||
(`manualRestart` path), which Docker deliberately does not auto-restart — by design, an operator who
|
||||
`stop`/`kill`s a container wants it stopped.
|
||||
|
||||
### Step 6 — definitive controller test (real crash, not the docker API)
|
||||
`kill -9 <controller container PID 1878>` (process death, bypassing docker stop/kill):
|
||||
```
|
||||
before: Running=true RestartCount=0 policy=unless-stopped
|
||||
t+5s: Running=true RestartCount=1 health=healthy
|
||||
… stable healthy thereafter
|
||||
```
|
||||
**The controller auto-recovered from a real crash in ~5 s**, with the data plane (apps + tunnel)
|
||||
unaffected throughout. (A 60 s timed auto-recovery guard was pre-armed and not needed; cleaned up.)
|
||||
|
||||
### Step 4 — the real recovery chain
|
||||
- Running controller: `RestartPolicy = unless-stopped` (confirmed).
|
||||
- `felhom-controller-bootstrap.service`: `Type=oneshot`, no `Restart=` — provision/boot-time bring-up
|
||||
only (correct). No other unit watches the container.
|
||||
- **What recovers a crashed controller:** docker's restart policy — the **live restart-manager** on a
|
||||
crash (proven) **and** boot-time restore at daemon start (which is why campaign C8's reboot brought
|
||||
26 containers back). Two independent, working paths. No secondary fallback is needed because the
|
||||
primary (restart policy) works.
|
||||
|
||||
## Classification
|
||||
**Finding #1 = TEST ARTIFACT.** Restart policy works; the campaign's `docker kill` is a manual stop, not
|
||||
a crash, so it correctly did not auto-restart. A genuine controller crash self-heals via
|
||||
`unless-stopped`. **Downgrade Finding #1** (it is not a resilience gap).
|
||||
|
||||
## Fix direction
|
||||
- **No fix required for crash-recovery.** The controller self-heals. Do **not** add a watchdog on this
|
||||
basis — it would patch a non-problem.
|
||||
- *Optional, low priority:* if defense-in-depth against an *operator/tooling* `docker stop`/`kill`
|
||||
(which intentionally won't auto-restart) is ever wanted, a tiny "controller absent ⇒ recreate"
|
||||
systemd timer could be considered — but it is **not** warranted by any observed failure.
|
||||
- **Re-test hygiene (the actual lesson):** simulate a crash with a real process death
|
||||
(`kill -9` the PID, OOM, panic) — never `docker kill`/`docker stop` — when testing restart policies.
|
||||
|
||||
## Separate observation (not Finding #1) — boot-time drive-ordering race
|
||||
At the C8 reboot (17:01:05) dockerd logged many `failed to start container … mkdir
|
||||
/mnt/felhom-drives/felhom-flash/userdata: permission denied` — docker's **boot-restore started
|
||||
drive-backed containers before the drives were mounted**. They recovered shortly after (the controller's
|
||||
recreate-on-boot once the drives mounted; C8 ended at 26 running/0 unhealthy). Real but minor
|
||||
(transient boot race); worth a follow-up to order drive-backed app start after the mount converges.
|
||||
No change applied.
|
||||
Reference in New Issue
Block a user