From 69e914534f918f5f62cfcd9b9413e30337fb41a2 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Sun, 4 Oct 2026 20:39:49 +0200 Subject: [PATCH] REPORT + CONTEXT: 2026-10-04 night (R-840 / R-860) Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- CONTEXT.md | 8 +++++++- REPORT.md | 19 ++++++++++++------- 2 files changed, 19 insertions(+), 8 deletions(-) diff --git a/CONTEXT.md b/CONTEXT.md index a027b5f..75591d1 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -7,7 +7,13 @@ > > Ask Claude Code: "Please update CONTEXT.md with what we did today" -Last updated: 2026-10-04 (v0.290.0 — the clean-up guard lets honest windows through, the set-aside deletion is the hub's) +Last updated: 2026-10-04 (v0.293.0 — the Docker-socket self-heal, R-860) + +> **2026-10-04 night — v0.293.0 (R-860; floor 0.293.0 on the demo customers only, golden 0.293.0 vouched):** +> `internal/sockheal` — the controller exits after 60 s of socket REFUSALS (never timeouts, only after Docker answered +> once) so Docker restarts it on the current socket, then restarts any other socket user on an older inode (traefik). +> Measured: only a `docker.socket` restart re-creates the file. Live: 9202 104 s, demo-hp 9201 120 s, ids unchanged. +> The global floor stays 0.292.0 until Tester 2 may move (operator ruling 97 limited Tester 2 to three acts). > **2026-10-04 — v0.290.0 (floored, golden 0.290.0 vouched, needs hub v0.128.0):** `offsiteGuard` EXCLUDES a young > snapshot superseded the same day in its group (R-824) and REFUSES above the hub's `max_remove` (half); a due diff --git a/REPORT.md b/REPORT.md index f0cd557..9282d3d 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,9 +1,14 @@ -# REPORT — v0.292.0: cloudflared readiness health check (R-841) — 2026-10-04 +# REPORT — controller v0.293.0: self-heal after the Docker socket is re-created (R-860) — 2026-10-04 -Full session report: `felhom.eu/REPORT-os-host-lane-2026-10-04.md`. +**Measured first (9202):** a `docker.socket` restart re-creates the socket file; with live-restore every container keeps +running, but the controller and traefik hold the deleted inode ("connection refused" forever, health still "healthy"). +A dockerd crash or `systemctl restart docker` keeps the file — nothing to heal. -- The cloudflared compose gains a Docker health check that asks cloudflared itself whether the tunnel is connected - (`/ready` on a fixed loopback metrics port). The host agent (v0.141.0) reads it and reports the tunnel as running, - not running or unknown; the hub (v0.131.0) alarms after two not-running reports. -- Tests: render test pins the health check and the command; red-proof recorded in the audit folder. -- Live: see the session report (both demo boxes, cloudflared recreated once, state `healthy`). +**Built:** `internal/sockheal` — 60 s of refusals (never timeouts; only after Docker answered once) → exit 75, Docker's +restart policy brings the controller back on the current socket (~1 s); every 5 min and 30 s after start it restarts any +other socket user on an older inode (traefik). MinAgent 0.131.0 unchanged. + +**Live:** 9202 healed in 104 s; demo-hp 9201 in 120 s, 21 of 21 container ids unchanged. Red-proof 9 of 9. Full suite +`go vet ./... && go test ./...` rc=0. Delivered to the demo boxes by a per-customer floor 0.293.0; the global floor stays +0.292.0 (Tester 2 not moved this session). Golden 0.293.0 baked and vouched. +Evidence: `felhom.eu/documentation/audits/r840-config-bundle-2026-10-04/partE/`.