Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
1.1 KiB
REPORT — controller v0.293.0: self-heal after the Docker socket is re-created (R-860) — 2026-10-04
Measured first (9202): a docker.socket restart re-creates the socket file; with live-restore every container keeps
running, but the controller and traefik hold the deleted inode ("connection refused" forever, health still "healthy").
A dockerd crash or systemctl restart docker keeps the file — nothing to heal.
Built: internal/sockheal — 60 s of refusals (never timeouts; only after Docker answered once) → exit 75, Docker's
restart policy brings the controller back on the current socket (~1 s); every 5 min and 30 s after start it restarts any
other socket user on an older inode (traefik). MinAgent 0.131.0 unchanged.
Live: 9202 healed in 104 s; demo-hp 9201 in 120 s, 21 of 21 container ids unchanged. Red-proof 9 of 9. Full suite
go vet ./... && go test ./... rc=0. Delivered to the demo boxes by a per-customer floor 0.293.0; the global floor stays
0.292.0 (Tester 2 not moved this session). Golden 0.293.0 baked and vouched.
Evidence: felhom.eu/documentation/audits/r840-config-bundle-2026-10-04/partE/.