docs(tests): 8443 unreachability ROOT CAUSE — controller http.Transport leak (read-only diagnosis)
Confirmed H5 (ephemeral-port exhaustion): agentClient() builds a new http.Transport per call (IdleConnTimeout:0, never CloseIdleConnections) -> ~5.8k leaked idle ESTABLISHED sockets/day to 162:8443, exhausts the 28k ephemeral range in ~5 days of controller uptime -> EADDRNOTAVAIL. :8006 immune (controller never dials pveproxy). Cleared by restart. H1/H2/H3/H4/H6 ruled out with positive evidence. Fix is controller-side (reuse one Client), NOT an agent rebind. No changes applied. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,165 @@
|
||||
# 8443 diagnosis — controller→agent local-API unreachability ROOT CAUSE — 2026-06-22
|
||||
|
||||
> **Status: read-only diagnosis (no fix applied).** Follow-up to
|
||||
> `unattended-test-campaign-2026-06-22-findings.md` finding #1. The task rejected the campaign's
|
||||
> "rebind agent to 0.0.0.0" guess (the bridge-IP bind is deliberate defense-in-depth) and asked to
|
||||
> establish the real cause from evidence. **It is established and verified.**
|
||||
|
||||
## TL;DR verdict
|
||||
|
||||
**Root cause = a connection/`http.Transport` leak in the *controller* (client side), not a network,
|
||||
firewall, agent-bind, or endpoint problem.** `felhom-controller`'s `agentClient()`
|
||||
(`internal/web/agent_disk_handlers.go:43`) calls `agentapi.New(...)` **on every agent API call**, and
|
||||
`agentapi.New` (`internal/agentapi/client.go:83`) builds a **fresh `&http.Client{Transport:
|
||||
&http.Transport{TLSClientConfig: …}}`** each time. That bare Transport has **`IdleConnTimeout: 0`**
|
||||
(idle keep-alive connections never expire) and is discarded after the single call **without
|
||||
`CloseIdleConnections()`**. Because the agent keeps connections alive, every call **leaks one idle
|
||||
ESTABLISHED socket** to `192.168.0.162:8443`. They accumulate monotonically over controller uptime
|
||||
until the **ephemeral source-port range** for the `(container → 192.168.0.162:8443)` tuple is
|
||||
exhausted → `connect: cannot assign requested address` (**EADDRNOTAVAIL**).
|
||||
|
||||
This is **H5 (ephemeral-port / socket exhaustion)**, with a precise code cause. **H1, H2, H3, H4, H6
|
||||
are ruled out with positive evidence below.**
|
||||
|
||||
Why `:8006` always worked and `:8443` "failed": the controller **only ever dials `:8443`** (the agent).
|
||||
It never connects to pveproxy `:8006`, so that destination never leaks and never exhausts. The failure
|
||||
was never port-specific in the network sense — it was **destination-specific exhaustion** driven by the
|
||||
leak.
|
||||
|
||||
Why it "broke persistently" in the campaign but **works now**: the leak accumulates with controller
|
||||
**uptime**. The campaign hit it at the controller's ~5-day uptime (pool exhausted). The Phase-7 guest
|
||||
reboot restarted the controller with a fresh socket table → it works again, and is **already
|
||||
re-accumulating** (measured below). This is an accumulating-state regression cleared by restart, NOT a
|
||||
static misconfiguration — exactly why the campaign's static theory (agent bind) didn't fit its own data.
|
||||
|
||||
---
|
||||
|
||||
## Evidence
|
||||
|
||||
### Step 1 — agent listener ground truth (host)
|
||||
```
|
||||
ss -tlnp 'sport = :8443' → LISTEN 192.168.0.162:8443 users:(("felhom-agent",pid=315505,fd=11))
|
||||
ss -tlnp 'sport = :8006' → LISTEN *:8006 users:(("pveproxy"…))
|
||||
ip -br addr → vmbr0 UP 192.168.0.162/24 ; veth9201i0@if2 UP ; (no IPv6 global)
|
||||
```
|
||||
Agent = a single **IPv4** socket bound to `192.168.0.162` (vmbr0), one process. pveproxy binds wildcard
|
||||
`*:8006`. Healthy listener. (The bind-to-bridge-IP is the intended defense-in-depth, and is fine.)
|
||||
|
||||
### Step 2 — controller configured endpoint (redacted)
|
||||
`bootstrap.json local_api`: `endpoint: 192.168.0.162:8443`, `token: «REDACTED»`,
|
||||
`fingerprint: «REDACTED»`. The endpoint **is** the vmbr0 bridge IP + correct port — matches the error
|
||||
string. **H2 (endpoint mismatch) ruled out.**
|
||||
|
||||
### Step 3 — connect matrix (the localizer) — ALL SUCCEED post-reboot
|
||||
Raw `socket.connect()` (AF_INET) to `192.168.0.162`, exact errno per layer × port:
|
||||
|
||||
| layer | :8443 | :8006 |
|
||||
|---|---|---|
|
||||
| controller **container netns** (pid 1917) | **OK** (local 172.17.0.2) | OK |
|
||||
| guest root netns | OK (local 192.168.0.141) | OK |
|
||||
| host | OK (local 192.168.0.162) | OK |
|
||||
|
||||
`nsenter -t 1917 -n ip route get 192.168.0.162` → `via 172.17.0.1 dev eth0 src 172.17.0.2 … cache`
|
||||
(resolves cleanly, source assignable). Container has eth0 `172.17.0.2/16` + eth1 `172.18.0.8/16`, no
|
||||
global IPv6. **At low uptime the TCP path is fully healthy — H3 (route/source) and H4 (address-family)
|
||||
ruled out.** Live re-test of the real app path also succeeds now: controller `/api/disks` returns the
|
||||
full disk JSON; in-container `curl https://192.168.0.162:8443/` → **404** (TCP+TLS OK), `:8006` → 200.
|
||||
**H6 (agent socket truth) ruled out** — agent accepts past TCP+TLS.
|
||||
|
||||
### Step 3b/core — the leak (the decisive artifact)
|
||||
Controller container, **47 minutes** after the Phase-7 reboot:
|
||||
```
|
||||
ss -s (container netns): TCP estab 194, closed 1049, timewait 9
|
||||
estab → 192.168.0.162:8443: 194 (essentially ALL established sockets go to the agent)
|
||||
ip_local_port_range: 32768 60999 (= 28231 ephemeral ports)
|
||||
```
|
||||
Monotonic growth (same netns, 30 s apart): **196 → 198**, later samples **206**. Rate ≈ 4/min ≈
|
||||
**~5 800/day**. Two-sided + idle confirmation:
|
||||
```
|
||||
HOST (agent) side, estab on :8443 from 192.168.0.141 (guest): 206 (Recv-Q/Send-Q = 0 → idle)
|
||||
controller netns estab→8443: 206
|
||||
controller process open fds: 217 (206 = leaked agent sockets)
|
||||
```
|
||||
**Projection:** ~5 800 leaked sockets/day ÷ 28 231 ports ⇒ exhaustion in **~4–5 days** of controller
|
||||
uptime → EADDRNOTAVAIL. Matches the campaign hitting it at ~5-day uptime exactly.
|
||||
|
||||
### Step 4 — guest-side nat/filter (H1, the campaign's prime suspect) — RULED OUT
|
||||
Inside guest 9201:
|
||||
```
|
||||
iptables-save | grep -E '8443|8006' → NONE (no port-specific rule)
|
||||
filter REJECT/DROP → only standard docker inter-bridge isolation
|
||||
(-A DOCKER ! -i br-X -o br-X -j DROP), not outbound-to-LAN
|
||||
docker subnets → 172.17/18/19/20/21.0.0/16 — none overlap 192.168.0.0/24
|
||||
```
|
||||
No rule treats 8443 differently from 8006; no subnet overlap. **H1 ruled out** (and the path works now,
|
||||
which alone disproves a static guest block).
|
||||
|
||||
### Step 5 — tcpdump death-point — N/A at current uptime
|
||||
The failure is **not reproducible at low uptime** (the path works — the SYN leaves and connects). A
|
||||
capture now would only show successful handshakes. The "death point" only appears once the ephemeral
|
||||
pool is exhausted, at which point `connect()` fails **locally** (no SYN emitted) — consistent with the
|
||||
EADDRNOTAVAIL semantics. No capture taken (nothing to capture); the leak measurement is the decisive
|
||||
artifact instead.
|
||||
|
||||
### Step 6 — temporal verdict
|
||||
**Regression that accumulates with controller uptime; reset by restart.** Definitive evidence: failed
|
||||
persistently during the campaign at the controller's ~5-day uptime (6/6 EADDRNOTAVAIL); after the
|
||||
Phase-7 guest reboot (controller `Up 47 minutes`) the path works and the leak is at 194→206 and
|
||||
climbing. The current container's logs only span 47 min (post-reboot), so the pre-reboot failures are
|
||||
gone with the old container — but the accumulate→reset mechanism is directly measured, not inferred. It
|
||||
"works after every restart, breaks after ~5 days."
|
||||
|
||||
### Step 7 — agent socket accept — confirmed reachable
|
||||
`curl -vk https://192.168.0.162:8443/` from in-container connects at TCP+TLS and returns HTTP 404; the
|
||||
pinned-TLS app path (`/api/disks`) returns real data. The agent is NOT the problem.
|
||||
|
||||
---
|
||||
|
||||
## Hypothesis scorecard
|
||||
|
||||
| | Hypothesis | Verdict | Deciding evidence |
|
||||
|---|---|---|---|
|
||||
| H1 | guest-side docker/nat block on 8443 | ❌ ruled out | no 8443/8006 rule; path works now |
|
||||
| H2 | configured-endpoint mismatch | ❌ ruled out | endpoint = `192.168.0.162:8443` (the bridge IP) |
|
||||
| H3 | source/route selection failure | ❌ ruled out | `ip route get` resolves, src assignable, raw connect OK |
|
||||
| H4 | address-family / IPv6 | ❌ ruled out | agent + container are IPv4-only; no global v6 |
|
||||
| **H5** | **ephemeral-port / socket exhaustion** | ✅ **CONFIRMED** | 206 leaked idle ESTABLISHED→8443 in 47 min, growing ~5.8k/day vs 28 231 ports ⇒ exhausts in ~5 days; matches campaign timing |
|
||||
| H6 | agent listener truth | ❌ ruled out | single healthy v4 socket; accepts TCP+TLS; returns data |
|
||||
|
||||
**Precise mechanism within H5:** `agentClient()` → `agentapi.New()` per call → new `http.Transport`
|
||||
with `IdleConnTimeout: 0`, discarded without `CloseIdleConnections()`; agent keep-alive ⇒ one leaked
|
||||
idle ESTABLISHED socket per call.
|
||||
|
||||
---
|
||||
|
||||
## Proposed fix direction (controller-side; NOT an agent rebind)
|
||||
|
||||
The agent's bind to `192.168.0.162` (bridge IP) is correct and must stay. The fix is entirely in the
|
||||
controller's client lifecycle:
|
||||
|
||||
1. **Build the `agentapi.Client` once and reuse it** (it is stateless config — endpoint/token/
|
||||
fingerprint, all from `cfg.LocalAPI`). Construct it at server init, store it on `Server`, and have
|
||||
`agentClient()` return the shared instance. A single reused Transport pools/reuses connections (≈2
|
||||
idle conns), eliminating the leak. **Preferred.**
|
||||
2. **If a per-call client is kept for any reason**, the Transport must be tamed and closed: set
|
||||
`IdleConnTimeout` (e.g. 30–90 s) + `MaxIdleConnsPerHost`, and `defer client.CloseIdleConnections()`
|
||||
after use. (Reuse (#1) is cleaner and also avoids the per-call TLS handshake cost.)
|
||||
3. Optionally add `MaxConnsPerHost`/keep-alive tuning on the shared Transport as belt-and-suspenders.
|
||||
|
||||
These are pure `felhom-controller` changes (then build + bump + redeploy per the controller workflow);
|
||||
no agent or firewall change.
|
||||
|
||||
### Separate hardening gap (not the cause — note for later)
|
||||
The defense-in-depth control the agent's own config comment calls for — *"a host firewall rule should
|
||||
limit [8443] to the guest bridge subnet"* — is **absent**: `pve-firewall` is disabled and there is no
|
||||
iptables rule scoping 8443. Close this once connectivity is fixed (a host rule allowing
|
||||
`192.168.0.0/24`→`:8443` and dropping others), independent of the leak fix.
|
||||
|
||||
---
|
||||
|
||||
## Note on the campaign finding #1
|
||||
`unattended-test-campaign-2026-06-22-findings.md` finding #1 correctly **observed** the symptom and
|
||||
**correctly rejected** rebinding to 0.0.0.0 in this follow-up, but its first-pass theory (agent
|
||||
bound to a specific IP ⇒ unreachable) was wrong — disproved by `:8006` reachability to the same IP and
|
||||
by the path working at low uptime. The real cause is the controller connection leak above. No code or
|
||||
config was changed by this diagnosis.
|
||||
Reference in New Issue
Block a user