Confirmed H5 (ephemeral-port exhaustion): agentClient() builds a new http.Transport per call (IdleConnTimeout:0, never CloseIdleConnections) -> ~5.8k leaked idle ESTABLISHED sockets/day to 162:8443, exhausts the 28k ephemeral range in ~5 days of controller uptime -> EADDRNOTAVAIL. :8006 immune (controller never dials pveproxy). Cleared by restart. H1/H2/H3/H4/H6 ruled out with positive evidence. Fix is controller-side (reuse one Client), NOT an agent rebind. No changes applied. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
9.9 KiB
8443 diagnosis — controller→agent local-API unreachability ROOT CAUSE — 2026-06-22
Status: read-only diagnosis (no fix applied). Follow-up to
unattended-test-campaign-2026-06-22-findings.mdfinding #1. The task rejected the campaign's "rebind agent to 0.0.0.0" guess (the bridge-IP bind is deliberate defense-in-depth) and asked to establish the real cause from evidence. It is established and verified.
TL;DR verdict
Root cause = a connection/http.Transport leak in the controller (client side), not a network,
firewall, agent-bind, or endpoint problem. felhom-controller's agentClient()
(internal/web/agent_disk_handlers.go:43) calls agentapi.New(...) on every agent API call, and
agentapi.New (internal/agentapi/client.go:83) builds a fresh &http.Client{Transport: &http.Transport{TLSClientConfig: …}} each time. That bare Transport has IdleConnTimeout: 0
(idle keep-alive connections never expire) and is discarded after the single call without
CloseIdleConnections(). Because the agent keeps connections alive, every call leaks one idle
ESTABLISHED socket to 192.168.0.162:8443. They accumulate monotonically over controller uptime
until the ephemeral source-port range for the (container → 192.168.0.162:8443) tuple is
exhausted → connect: cannot assign requested address (EADDRNOTAVAIL).
This is H5 (ephemeral-port / socket exhaustion), with a precise code cause. H1, H2, H3, H4, H6 are ruled out with positive evidence below.
Why :8006 always worked and :8443 "failed": the controller only ever dials :8443 (the agent).
It never connects to pveproxy :8006, so that destination never leaks and never exhausts. The failure
was never port-specific in the network sense — it was destination-specific exhaustion driven by the
leak.
Why it "broke persistently" in the campaign but works now: the leak accumulates with controller uptime. The campaign hit it at the controller's ~5-day uptime (pool exhausted). The Phase-7 guest reboot restarted the controller with a fresh socket table → it works again, and is already re-accumulating (measured below). This is an accumulating-state regression cleared by restart, NOT a static misconfiguration — exactly why the campaign's static theory (agent bind) didn't fit its own data.
Evidence
Step 1 — agent listener ground truth (host)
ss -tlnp 'sport = :8443' → LISTEN 192.168.0.162:8443 users:(("felhom-agent",pid=315505,fd=11))
ss -tlnp 'sport = :8006' → LISTEN *:8006 users:(("pveproxy"…))
ip -br addr → vmbr0 UP 192.168.0.162/24 ; veth9201i0@if2 UP ; (no IPv6 global)
Agent = a single IPv4 socket bound to 192.168.0.162 (vmbr0), one process. pveproxy binds wildcard
*:8006. Healthy listener. (The bind-to-bridge-IP is the intended defense-in-depth, and is fine.)
Step 2 — controller configured endpoint (redacted)
bootstrap.json local_api: endpoint: 192.168.0.162:8443, token: «REDACTED»,
fingerprint: «REDACTED». The endpoint is the vmbr0 bridge IP + correct port — matches the error
string. H2 (endpoint mismatch) ruled out.
Step 3 — connect matrix (the localizer) — ALL SUCCEED post-reboot
Raw socket.connect() (AF_INET) to 192.168.0.162, exact errno per layer × port:
| layer | :8443 | :8006 |
|---|---|---|
| controller container netns (pid 1917) | OK (local 172.17.0.2) | OK |
| guest root netns | OK (local 192.168.0.141) | OK |
| host | OK (local 192.168.0.162) | OK |
nsenter -t 1917 -n ip route get 192.168.0.162 → via 172.17.0.1 dev eth0 src 172.17.0.2 … cache
(resolves cleanly, source assignable). Container has eth0 172.17.0.2/16 + eth1 172.18.0.8/16, no
global IPv6. At low uptime the TCP path is fully healthy — H3 (route/source) and H4 (address-family)
ruled out. Live re-test of the real app path also succeeds now: controller /api/disks returns the
full disk JSON; in-container curl https://192.168.0.162:8443/ → 404 (TCP+TLS OK), :8006 → 200.
H6 (agent socket truth) ruled out — agent accepts past TCP+TLS.
Step 3b/core — the leak (the decisive artifact)
Controller container, 47 minutes after the Phase-7 reboot:
ss -s (container netns): TCP estab 194, closed 1049, timewait 9
estab → 192.168.0.162:8443: 194 (essentially ALL established sockets go to the agent)
ip_local_port_range: 32768 60999 (= 28231 ephemeral ports)
Monotonic growth (same netns, 30 s apart): 196 → 198, later samples 206. Rate ≈ 4/min ≈ ~5 800/day. Two-sided + idle confirmation:
HOST (agent) side, estab on :8443 from 192.168.0.141 (guest): 206 (Recv-Q/Send-Q = 0 → idle)
controller netns estab→8443: 206
controller process open fds: 217 (206 = leaked agent sockets)
Projection: ~5 800 leaked sockets/day ÷ 28 231 ports ⇒ exhaustion in ~4–5 days of controller uptime → EADDRNOTAVAIL. Matches the campaign hitting it at ~5-day uptime exactly.
Step 4 — guest-side nat/filter (H1, the campaign's prime suspect) — RULED OUT
Inside guest 9201:
iptables-save | grep -E '8443|8006' → NONE (no port-specific rule)
filter REJECT/DROP → only standard docker inter-bridge isolation
(-A DOCKER ! -i br-X -o br-X -j DROP), not outbound-to-LAN
docker subnets → 172.17/18/19/20/21.0.0/16 — none overlap 192.168.0.0/24
No rule treats 8443 differently from 8006; no subnet overlap. H1 ruled out (and the path works now, which alone disproves a static guest block).
Step 5 — tcpdump death-point — N/A at current uptime
The failure is not reproducible at low uptime (the path works — the SYN leaves and connects). A
capture now would only show successful handshakes. The "death point" only appears once the ephemeral
pool is exhausted, at which point connect() fails locally (no SYN emitted) — consistent with the
EADDRNOTAVAIL semantics. No capture taken (nothing to capture); the leak measurement is the decisive
artifact instead.
Step 6 — temporal verdict
Regression that accumulates with controller uptime; reset by restart. Definitive evidence: failed
persistently during the campaign at the controller's ~5-day uptime (6/6 EADDRNOTAVAIL); after the
Phase-7 guest reboot (controller Up 47 minutes) the path works and the leak is at 194→206 and
climbing. The current container's logs only span 47 min (post-reboot), so the pre-reboot failures are
gone with the old container — but the accumulate→reset mechanism is directly measured, not inferred. It
"works after every restart, breaks after ~5 days."
Step 7 — agent socket accept — confirmed reachable
curl -vk https://192.168.0.162:8443/ from in-container connects at TCP+TLS and returns HTTP 404; the
pinned-TLS app path (/api/disks) returns real data. The agent is NOT the problem.
Hypothesis scorecard
| Hypothesis | Verdict | Deciding evidence | |
|---|---|---|---|
| H1 | guest-side docker/nat block on 8443 | ❌ ruled out | no 8443/8006 rule; path works now |
| H2 | configured-endpoint mismatch | ❌ ruled out | endpoint = 192.168.0.162:8443 (the bridge IP) |
| H3 | source/route selection failure | ❌ ruled out | ip route get resolves, src assignable, raw connect OK |
| H4 | address-family / IPv6 | ❌ ruled out | agent + container are IPv4-only; no global v6 |
| H5 | ephemeral-port / socket exhaustion | ✅ CONFIRMED | 206 leaked idle ESTABLISHED→8443 in 47 min, growing ~5.8k/day vs 28 231 ports ⇒ exhausts in ~5 days; matches campaign timing |
| H6 | agent listener truth | ❌ ruled out | single healthy v4 socket; accepts TCP+TLS; returns data |
Precise mechanism within H5: agentClient() → agentapi.New() per call → new http.Transport
with IdleConnTimeout: 0, discarded without CloseIdleConnections(); agent keep-alive ⇒ one leaked
idle ESTABLISHED socket per call.
Proposed fix direction (controller-side; NOT an agent rebind)
The agent's bind to 192.168.0.162 (bridge IP) is correct and must stay. The fix is entirely in the
controller's client lifecycle:
- Build the
agentapi.Clientonce and reuse it (it is stateless config — endpoint/token/ fingerprint, all fromcfg.LocalAPI). Construct it at server init, store it onServer, and haveagentClient()return the shared instance. A single reused Transport pools/reuses connections (≈2 idle conns), eliminating the leak. Preferred. - If a per-call client is kept for any reason, the Transport must be tamed and closed: set
IdleConnTimeout(e.g. 30–90 s) +MaxIdleConnsPerHost, anddefer client.CloseIdleConnections()after use. (Reuse (#1) is cleaner and also avoids the per-call TLS handshake cost.) - Optionally add
MaxConnsPerHost/keep-alive tuning on the shared Transport as belt-and-suspenders.
These are pure felhom-controller changes (then build + bump + redeploy per the controller workflow);
no agent or firewall change.
Separate hardening gap (not the cause — note for later)
The defense-in-depth control the agent's own config comment calls for — "a host firewall rule should
limit [8443] to the guest bridge subnet" — is absent: pve-firewall is disabled and there is no
iptables rule scoping 8443. Close this once connectivity is fixed (a host rule allowing
192.168.0.0/24→:8443 and dropping others), independent of the leak fix.
Note on the campaign finding #1
unattended-test-campaign-2026-06-22-findings.md finding #1 correctly observed the symptom and
correctly rejected rebinding to 0.0.0.0 in this follow-up, but its first-pass theory (agent
bound to a specific IP ⇒ unreachable) was wrong — disproved by :8006 reachability to the same IP and
by the path working at low uptime. The real cause is the controller connection leak above. No code or
config was changed by this diagnosis.