docs(tests): 8443 unreachability ROOT CAUSE — controller http.Transport leak (read-only diagnosis)

Confirmed H5 (ephemeral-port exhaustion): agentClient() builds a new http.Transport
per call (IdleConnTimeout:0, never CloseIdleConnections) -> ~5.8k leaked idle ESTABLISHED
sockets/day to 162:8443, exhausts the 28k ephemeral range in ~5 days of controller uptime
-> EADDRNOTAVAIL. :8006 immune (controller never dials pveproxy). Cleared by restart.
H1/H2/H3/H4/H6 ruled out with positive evidence. Fix is controller-side (reuse one
Client), NOT an agent rebind. No changes applied.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-06-22 17:00:24 +02:00
parent f70c011749
commit a9ca5b4e68
@@ -0,0 +1,165 @@
# 8443 diagnosis — controller→agent local-API unreachability ROOT CAUSE — 2026-06-22
> **Status: read-only diagnosis (no fix applied).** Follow-up to
> `unattended-test-campaign-2026-06-22-findings.md` finding #1. The task rejected the campaign's
> "rebind agent to 0.0.0.0" guess (the bridge-IP bind is deliberate defense-in-depth) and asked to
> establish the real cause from evidence. **It is established and verified.**
## TL;DR verdict
**Root cause = a connection/`http.Transport` leak in the *controller* (client side), not a network,
firewall, agent-bind, or endpoint problem.** `felhom-controller`'s `agentClient()`
(`internal/web/agent_disk_handlers.go:43`) calls `agentapi.New(...)` **on every agent API call**, and
`agentapi.New` (`internal/agentapi/client.go:83`) builds a **fresh `&http.Client{Transport:
&http.Transport{TLSClientConfig: …}}`** each time. That bare Transport has **`IdleConnTimeout: 0`**
(idle keep-alive connections never expire) and is discarded after the single call **without
`CloseIdleConnections()`**. Because the agent keeps connections alive, every call **leaks one idle
ESTABLISHED socket** to `192.168.0.162:8443`. They accumulate monotonically over controller uptime
until the **ephemeral source-port range** for the `(container → 192.168.0.162:8443)` tuple is
exhausted → `connect: cannot assign requested address` (**EADDRNOTAVAIL**).
This is **H5 (ephemeral-port / socket exhaustion)**, with a precise code cause. **H1, H2, H3, H4, H6
are ruled out with positive evidence below.**
Why `:8006` always worked and `:8443` "failed": the controller **only ever dials `:8443`** (the agent).
It never connects to pveproxy `:8006`, so that destination never leaks and never exhausts. The failure
was never port-specific in the network sense — it was **destination-specific exhaustion** driven by the
leak.
Why it "broke persistently" in the campaign but **works now**: the leak accumulates with controller
**uptime**. The campaign hit it at the controller's ~5-day uptime (pool exhausted). The Phase-7 guest
reboot restarted the controller with a fresh socket table → it works again, and is **already
re-accumulating** (measured below). This is an accumulating-state regression cleared by restart, NOT a
static misconfiguration — exactly why the campaign's static theory (agent bind) didn't fit its own data.
---
## Evidence
### Step 1 — agent listener ground truth (host)
```
ss -tlnp 'sport = :8443' → LISTEN 192.168.0.162:8443 users:(("felhom-agent",pid=315505,fd=11))
ss -tlnp 'sport = :8006' → LISTEN *:8006 users:(("pveproxy"…))
ip -br addr → vmbr0 UP 192.168.0.162/24 ; veth9201i0@if2 UP ; (no IPv6 global)
```
Agent = a single **IPv4** socket bound to `192.168.0.162` (vmbr0), one process. pveproxy binds wildcard
`*:8006`. Healthy listener. (The bind-to-bridge-IP is the intended defense-in-depth, and is fine.)
### Step 2 — controller configured endpoint (redacted)
`bootstrap.json local_api`: `endpoint: 192.168.0.162:8443`, `token: «REDACTED»`,
`fingerprint: «REDACTED»`. The endpoint **is** the vmbr0 bridge IP + correct port — matches the error
string. **H2 (endpoint mismatch) ruled out.**
### Step 3 — connect matrix (the localizer) — ALL SUCCEED post-reboot
Raw `socket.connect()` (AF_INET) to `192.168.0.162`, exact errno per layer × port:
| layer | :8443 | :8006 |
|---|---|---|
| controller **container netns** (pid 1917) | **OK** (local 172.17.0.2) | OK |
| guest root netns | OK (local 192.168.0.141) | OK |
| host | OK (local 192.168.0.162) | OK |
`nsenter -t 1917 -n ip route get 192.168.0.162``via 172.17.0.1 dev eth0 src 172.17.0.2 … cache`
(resolves cleanly, source assignable). Container has eth0 `172.17.0.2/16` + eth1 `172.18.0.8/16`, no
global IPv6. **At low uptime the TCP path is fully healthy — H3 (route/source) and H4 (address-family)
ruled out.** Live re-test of the real app path also succeeds now: controller `/api/disks` returns the
full disk JSON; in-container `curl https://192.168.0.162:8443/`**404** (TCP+TLS OK), `:8006` → 200.
**H6 (agent socket truth) ruled out** — agent accepts past TCP+TLS.
### Step 3b/core — the leak (the decisive artifact)
Controller container, **47 minutes** after the Phase-7 reboot:
```
ss -s (container netns): TCP estab 194, closed 1049, timewait 9
estab → 192.168.0.162:8443: 194 (essentially ALL established sockets go to the agent)
ip_local_port_range: 32768 60999 (= 28231 ephemeral ports)
```
Monotonic growth (same netns, 30 s apart): **196 → 198**, later samples **206**. Rate ≈ 4/min ≈
**~5 800/day**. Two-sided + idle confirmation:
```
HOST (agent) side, estab on :8443 from 192.168.0.141 (guest): 206 (Recv-Q/Send-Q = 0 → idle)
controller netns estab→8443: 206
controller process open fds: 217 (206 = leaked agent sockets)
```
**Projection:** ~5 800 leaked sockets/day ÷ 28 231 ports ⇒ exhaustion in **~45 days** of controller
uptime → EADDRNOTAVAIL. Matches the campaign hitting it at ~5-day uptime exactly.
### Step 4 — guest-side nat/filter (H1, the campaign's prime suspect) — RULED OUT
Inside guest 9201:
```
iptables-save | grep -E '8443|8006' → NONE (no port-specific rule)
filter REJECT/DROP → only standard docker inter-bridge isolation
(-A DOCKER ! -i br-X -o br-X -j DROP), not outbound-to-LAN
docker subnets → 172.17/18/19/20/21.0.0/16 — none overlap 192.168.0.0/24
```
No rule treats 8443 differently from 8006; no subnet overlap. **H1 ruled out** (and the path works now,
which alone disproves a static guest block).
### Step 5 — tcpdump death-point — N/A at current uptime
The failure is **not reproducible at low uptime** (the path works — the SYN leaves and connects). A
capture now would only show successful handshakes. The "death point" only appears once the ephemeral
pool is exhausted, at which point `connect()` fails **locally** (no SYN emitted) — consistent with the
EADDRNOTAVAIL semantics. No capture taken (nothing to capture); the leak measurement is the decisive
artifact instead.
### Step 6 — temporal verdict
**Regression that accumulates with controller uptime; reset by restart.** Definitive evidence: failed
persistently during the campaign at the controller's ~5-day uptime (6/6 EADDRNOTAVAIL); after the
Phase-7 guest reboot (controller `Up 47 minutes`) the path works and the leak is at 194→206 and
climbing. The current container's logs only span 47 min (post-reboot), so the pre-reboot failures are
gone with the old container — but the accumulate→reset mechanism is directly measured, not inferred. It
"works after every restart, breaks after ~5 days."
### Step 7 — agent socket accept — confirmed reachable
`curl -vk https://192.168.0.162:8443/` from in-container connects at TCP+TLS and returns HTTP 404; the
pinned-TLS app path (`/api/disks`) returns real data. The agent is NOT the problem.
---
## Hypothesis scorecard
| | Hypothesis | Verdict | Deciding evidence |
|---|---|---|---|
| H1 | guest-side docker/nat block on 8443 | ❌ ruled out | no 8443/8006 rule; path works now |
| H2 | configured-endpoint mismatch | ❌ ruled out | endpoint = `192.168.0.162:8443` (the bridge IP) |
| H3 | source/route selection failure | ❌ ruled out | `ip route get` resolves, src assignable, raw connect OK |
| H4 | address-family / IPv6 | ❌ ruled out | agent + container are IPv4-only; no global v6 |
| **H5** | **ephemeral-port / socket exhaustion** | ✅ **CONFIRMED** | 206 leaked idle ESTABLISHED→8443 in 47 min, growing ~5.8k/day vs 28 231 ports ⇒ exhausts in ~5 days; matches campaign timing |
| H6 | agent listener truth | ❌ ruled out | single healthy v4 socket; accepts TCP+TLS; returns data |
**Precise mechanism within H5:** `agentClient()``agentapi.New()` per call → new `http.Transport`
with `IdleConnTimeout: 0`, discarded without `CloseIdleConnections()`; agent keep-alive ⇒ one leaked
idle ESTABLISHED socket per call.
---
## Proposed fix direction (controller-side; NOT an agent rebind)
The agent's bind to `192.168.0.162` (bridge IP) is correct and must stay. The fix is entirely in the
controller's client lifecycle:
1. **Build the `agentapi.Client` once and reuse it** (it is stateless config — endpoint/token/
fingerprint, all from `cfg.LocalAPI`). Construct it at server init, store it on `Server`, and have
`agentClient()` return the shared instance. A single reused Transport pools/reuses connections (≈2
idle conns), eliminating the leak. **Preferred.**
2. **If a per-call client is kept for any reason**, the Transport must be tamed and closed: set
`IdleConnTimeout` (e.g. 3090 s) + `MaxIdleConnsPerHost`, and `defer client.CloseIdleConnections()`
after use. (Reuse (#1) is cleaner and also avoids the per-call TLS handshake cost.)
3. Optionally add `MaxConnsPerHost`/keep-alive tuning on the shared Transport as belt-and-suspenders.
These are pure `felhom-controller` changes (then build + bump + redeploy per the controller workflow);
no agent or firewall change.
### Separate hardening gap (not the cause — note for later)
The defense-in-depth control the agent's own config comment calls for — *"a host firewall rule should
limit [8443] to the guest bridge subnet"* — is **absent**: `pve-firewall` is disabled and there is no
iptables rule scoping 8443. Close this once connectivity is fixed (a host rule allowing
`192.168.0.0/24``:8443` and dropping others), independent of the leak fix.
---
## Note on the campaign finding #1
`unattended-test-campaign-2026-06-22-findings.md` finding #1 correctly **observed** the symptom and
**correctly rejected** rebinding to 0.0.0.0 in this follow-up, but its first-pass theory (agent
bound to a specific IP ⇒ unreachable) was wrong — disproved by `:8006` reachability to the same IP and
by the path working at low uptime. The real cause is the controller connection leak above. No code or
config was changed by this diagnosis.