R-858 closed (agent v0.142.1, ruling 95): the Docker step restarts the socket users; incident evidence; STATUS; report addendum
gates / gates (push) Successful in 34s
gates / gates (push) Successful in 34s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -26,6 +26,12 @@
|
||||
|
||||
---
|
||||
|
||||
## 2026-10-04 (~18:30) — R-858 incident (agent v0.142.1, ruling 95)
|
||||
|
||||
| Row | What | Closed | Evidence |
|
||||
|---|---|---|---|
|
||||
| **R-858** | **A Docker engine step left `felhom-controller` and `traefik` on the OLD docker socket** (found by the operator: demo-felhom DOWN 14:18–15:57 UTC). live-restore kept them running with the deleted, bind-mounted socket file; their own health stayed "healthy", so the step passed. Repaired by restarting the two; agent v0.142.1 restarts ONLY the socket-mounting containers after a step and fails the health rule when the controller cannot reach Docker. Proven live twice on demo-hp (signed undo and forward): both restarted, guest / controller / traefik on the same socket inode, `applied, healthy`. **Reasoning kept: a container that bind-mounts a socket FILE does not follow a restarted daemon — live-restore makes this worse, not better. Check the consequence (the controller reaches Docker), not the mechanism (the container still runs).** | CLOSED 2026-10-04 — FIXED | `audits/os-docker-crash-2026-10-04/partE-incident/` |
|
||||
|
||||
## 2026-10-04 (evening) — System page, Docker slow lane, crash guard (agent v0.142.0, hub v0.132.0, installer 1.30.0)
|
||||
|
||||
> Evidence: `audits/os-docker-crash-2026-10-04/`.
|
||||
|
||||
Reference in New Issue
Block a user