Files
felhom-controller/REPORT.md
T
admin 69e914534f
gates / gates (push) Successful in 33s
REPORT + CONTEXT: 2026-10-04 night (R-840 / R-860)
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 20:39:49 +02:00

1.1 KiB

REPORT — controller v0.293.0: self-heal after the Docker socket is re-created (R-860) — 2026-10-04

Measured first (9202): a docker.socket restart re-creates the socket file; with live-restore every container keeps running, but the controller and traefik hold the deleted inode ("connection refused" forever, health still "healthy"). A dockerd crash or systemctl restart docker keeps the file — nothing to heal.

Built: internal/sockheal — 60 s of refusals (never timeouts; only after Docker answered once) → exit 75, Docker's restart policy brings the controller back on the current socket (~1 s); every 5 min and 30 s after start it restarts any other socket user on an older inode (traefik). MinAgent 0.131.0 unchanged.

Live: 9202 healed in 104 s; demo-hp 9201 in 120 s, 21 of 21 container ids unchanged. Red-proof 9 of 9. Full suite go vet ./... && go test ./... rc=0. Delivered to the demo boxes by a per-customer floor 0.293.0; the global floor stays 0.292.0 (Tester 2 not moved this session). Golden 0.293.0 baked and vouched. Evidence: felhom.eu/documentation/audits/r840-config-bundle-2026-10-04/partE/.