docs(tests): 8443 unreachability ROOT CAUSE — controller http.Transport leak (read-only diagnosis)
Confirmed H5 (ephemeral-port exhaustion): agentClient() builds a new http.Transport per call (IdleConnTimeout:0, never CloseIdleConnections) -> ~5.8k leaked idle ESTABLISHED sockets/day to 162:8443, exhausts the 28k ephemeral range in ~5 days of controller uptime -> EADDRNOTAVAIL. :8006 immune (controller never dials pveproxy). Cleared by restart. H1/H2/H3/H4/H6 ruled out with positive evidence. Fix is controller-side (reuse one Client), NOT an agent rebind. No changes applied. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,165 @@
|
|||||||
|
# 8443 diagnosis — controller→agent local-API unreachability ROOT CAUSE — 2026-06-22
|
||||||
|
|
||||||
|
> **Status: read-only diagnosis (no fix applied).** Follow-up to
|
||||||
|
> `unattended-test-campaign-2026-06-22-findings.md` finding #1. The task rejected the campaign's
|
||||||
|
> "rebind agent to 0.0.0.0" guess (the bridge-IP bind is deliberate defense-in-depth) and asked to
|
||||||
|
> establish the real cause from evidence. **It is established and verified.**
|
||||||
|
|
||||||
|
## TL;DR verdict
|
||||||
|
|
||||||
|
**Root cause = a connection/`http.Transport` leak in the *controller* (client side), not a network,
|
||||||
|
firewall, agent-bind, or endpoint problem.** `felhom-controller`'s `agentClient()`
|
||||||
|
(`internal/web/agent_disk_handlers.go:43`) calls `agentapi.New(...)` **on every agent API call**, and
|
||||||
|
`agentapi.New` (`internal/agentapi/client.go:83`) builds a **fresh `&http.Client{Transport:
|
||||||
|
&http.Transport{TLSClientConfig: …}}`** each time. That bare Transport has **`IdleConnTimeout: 0`**
|
||||||
|
(idle keep-alive connections never expire) and is discarded after the single call **without
|
||||||
|
`CloseIdleConnections()`**. Because the agent keeps connections alive, every call **leaks one idle
|
||||||
|
ESTABLISHED socket** to `192.168.0.162:8443`. They accumulate monotonically over controller uptime
|
||||||
|
until the **ephemeral source-port range** for the `(container → 192.168.0.162:8443)` tuple is
|
||||||
|
exhausted → `connect: cannot assign requested address` (**EADDRNOTAVAIL**).
|
||||||
|
|
||||||
|
This is **H5 (ephemeral-port / socket exhaustion)**, with a precise code cause. **H1, H2, H3, H4, H6
|
||||||
|
are ruled out with positive evidence below.**
|
||||||
|
|
||||||
|
Why `:8006` always worked and `:8443` "failed": the controller **only ever dials `:8443`** (the agent).
|
||||||
|
It never connects to pveproxy `:8006`, so that destination never leaks and never exhausts. The failure
|
||||||
|
was never port-specific in the network sense — it was **destination-specific exhaustion** driven by the
|
||||||
|
leak.
|
||||||
|
|
||||||
|
Why it "broke persistently" in the campaign but **works now**: the leak accumulates with controller
|
||||||
|
**uptime**. The campaign hit it at the controller's ~5-day uptime (pool exhausted). The Phase-7 guest
|
||||||
|
reboot restarted the controller with a fresh socket table → it works again, and is **already
|
||||||
|
re-accumulating** (measured below). This is an accumulating-state regression cleared by restart, NOT a
|
||||||
|
static misconfiguration — exactly why the campaign's static theory (agent bind) didn't fit its own data.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Evidence
|
||||||
|
|
||||||
|
### Step 1 — agent listener ground truth (host)
|
||||||
|
```
|
||||||
|
ss -tlnp 'sport = :8443' → LISTEN 192.168.0.162:8443 users:(("felhom-agent",pid=315505,fd=11))
|
||||||
|
ss -tlnp 'sport = :8006' → LISTEN *:8006 users:(("pveproxy"…))
|
||||||
|
ip -br addr → vmbr0 UP 192.168.0.162/24 ; veth9201i0@if2 UP ; (no IPv6 global)
|
||||||
|
```
|
||||||
|
Agent = a single **IPv4** socket bound to `192.168.0.162` (vmbr0), one process. pveproxy binds wildcard
|
||||||
|
`*:8006`. Healthy listener. (The bind-to-bridge-IP is the intended defense-in-depth, and is fine.)
|
||||||
|
|
||||||
|
### Step 2 — controller configured endpoint (redacted)
|
||||||
|
`bootstrap.json local_api`: `endpoint: 192.168.0.162:8443`, `token: «REDACTED»`,
|
||||||
|
`fingerprint: «REDACTED»`. The endpoint **is** the vmbr0 bridge IP + correct port — matches the error
|
||||||
|
string. **H2 (endpoint mismatch) ruled out.**
|
||||||
|
|
||||||
|
### Step 3 — connect matrix (the localizer) — ALL SUCCEED post-reboot
|
||||||
|
Raw `socket.connect()` (AF_INET) to `192.168.0.162`, exact errno per layer × port:
|
||||||
|
|
||||||
|
| layer | :8443 | :8006 |
|
||||||
|
|---|---|---|
|
||||||
|
| controller **container netns** (pid 1917) | **OK** (local 172.17.0.2) | OK |
|
||||||
|
| guest root netns | OK (local 192.168.0.141) | OK |
|
||||||
|
| host | OK (local 192.168.0.162) | OK |
|
||||||
|
|
||||||
|
`nsenter -t 1917 -n ip route get 192.168.0.162` → `via 172.17.0.1 dev eth0 src 172.17.0.2 … cache`
|
||||||
|
(resolves cleanly, source assignable). Container has eth0 `172.17.0.2/16` + eth1 `172.18.0.8/16`, no
|
||||||
|
global IPv6. **At low uptime the TCP path is fully healthy — H3 (route/source) and H4 (address-family)
|
||||||
|
ruled out.** Live re-test of the real app path also succeeds now: controller `/api/disks` returns the
|
||||||
|
full disk JSON; in-container `curl https://192.168.0.162:8443/` → **404** (TCP+TLS OK), `:8006` → 200.
|
||||||
|
**H6 (agent socket truth) ruled out** — agent accepts past TCP+TLS.
|
||||||
|
|
||||||
|
### Step 3b/core — the leak (the decisive artifact)
|
||||||
|
Controller container, **47 minutes** after the Phase-7 reboot:
|
||||||
|
```
|
||||||
|
ss -s (container netns): TCP estab 194, closed 1049, timewait 9
|
||||||
|
estab → 192.168.0.162:8443: 194 (essentially ALL established sockets go to the agent)
|
||||||
|
ip_local_port_range: 32768 60999 (= 28231 ephemeral ports)
|
||||||
|
```
|
||||||
|
Monotonic growth (same netns, 30 s apart): **196 → 198**, later samples **206**. Rate ≈ 4/min ≈
|
||||||
|
**~5 800/day**. Two-sided + idle confirmation:
|
||||||
|
```
|
||||||
|
HOST (agent) side, estab on :8443 from 192.168.0.141 (guest): 206 (Recv-Q/Send-Q = 0 → idle)
|
||||||
|
controller netns estab→8443: 206
|
||||||
|
controller process open fds: 217 (206 = leaked agent sockets)
|
||||||
|
```
|
||||||
|
**Projection:** ~5 800 leaked sockets/day ÷ 28 231 ports ⇒ exhaustion in **~4–5 days** of controller
|
||||||
|
uptime → EADDRNOTAVAIL. Matches the campaign hitting it at ~5-day uptime exactly.
|
||||||
|
|
||||||
|
### Step 4 — guest-side nat/filter (H1, the campaign's prime suspect) — RULED OUT
|
||||||
|
Inside guest 9201:
|
||||||
|
```
|
||||||
|
iptables-save | grep -E '8443|8006' → NONE (no port-specific rule)
|
||||||
|
filter REJECT/DROP → only standard docker inter-bridge isolation
|
||||||
|
(-A DOCKER ! -i br-X -o br-X -j DROP), not outbound-to-LAN
|
||||||
|
docker subnets → 172.17/18/19/20/21.0.0/16 — none overlap 192.168.0.0/24
|
||||||
|
```
|
||||||
|
No rule treats 8443 differently from 8006; no subnet overlap. **H1 ruled out** (and the path works now,
|
||||||
|
which alone disproves a static guest block).
|
||||||
|
|
||||||
|
### Step 5 — tcpdump death-point — N/A at current uptime
|
||||||
|
The failure is **not reproducible at low uptime** (the path works — the SYN leaves and connects). A
|
||||||
|
capture now would only show successful handshakes. The "death point" only appears once the ephemeral
|
||||||
|
pool is exhausted, at which point `connect()` fails **locally** (no SYN emitted) — consistent with the
|
||||||
|
EADDRNOTAVAIL semantics. No capture taken (nothing to capture); the leak measurement is the decisive
|
||||||
|
artifact instead.
|
||||||
|
|
||||||
|
### Step 6 — temporal verdict
|
||||||
|
**Regression that accumulates with controller uptime; reset by restart.** Definitive evidence: failed
|
||||||
|
persistently during the campaign at the controller's ~5-day uptime (6/6 EADDRNOTAVAIL); after the
|
||||||
|
Phase-7 guest reboot (controller `Up 47 minutes`) the path works and the leak is at 194→206 and
|
||||||
|
climbing. The current container's logs only span 47 min (post-reboot), so the pre-reboot failures are
|
||||||
|
gone with the old container — but the accumulate→reset mechanism is directly measured, not inferred. It
|
||||||
|
"works after every restart, breaks after ~5 days."
|
||||||
|
|
||||||
|
### Step 7 — agent socket accept — confirmed reachable
|
||||||
|
`curl -vk https://192.168.0.162:8443/` from in-container connects at TCP+TLS and returns HTTP 404; the
|
||||||
|
pinned-TLS app path (`/api/disks`) returns real data. The agent is NOT the problem.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Hypothesis scorecard
|
||||||
|
|
||||||
|
| | Hypothesis | Verdict | Deciding evidence |
|
||||||
|
|---|---|---|---|
|
||||||
|
| H1 | guest-side docker/nat block on 8443 | ❌ ruled out | no 8443/8006 rule; path works now |
|
||||||
|
| H2 | configured-endpoint mismatch | ❌ ruled out | endpoint = `192.168.0.162:8443` (the bridge IP) |
|
||||||
|
| H3 | source/route selection failure | ❌ ruled out | `ip route get` resolves, src assignable, raw connect OK |
|
||||||
|
| H4 | address-family / IPv6 | ❌ ruled out | agent + container are IPv4-only; no global v6 |
|
||||||
|
| **H5** | **ephemeral-port / socket exhaustion** | ✅ **CONFIRMED** | 206 leaked idle ESTABLISHED→8443 in 47 min, growing ~5.8k/day vs 28 231 ports ⇒ exhausts in ~5 days; matches campaign timing |
|
||||||
|
| H6 | agent listener truth | ❌ ruled out | single healthy v4 socket; accepts TCP+TLS; returns data |
|
||||||
|
|
||||||
|
**Precise mechanism within H5:** `agentClient()` → `agentapi.New()` per call → new `http.Transport`
|
||||||
|
with `IdleConnTimeout: 0`, discarded without `CloseIdleConnections()`; agent keep-alive ⇒ one leaked
|
||||||
|
idle ESTABLISHED socket per call.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Proposed fix direction (controller-side; NOT an agent rebind)
|
||||||
|
|
||||||
|
The agent's bind to `192.168.0.162` (bridge IP) is correct and must stay. The fix is entirely in the
|
||||||
|
controller's client lifecycle:
|
||||||
|
|
||||||
|
1. **Build the `agentapi.Client` once and reuse it** (it is stateless config — endpoint/token/
|
||||||
|
fingerprint, all from `cfg.LocalAPI`). Construct it at server init, store it on `Server`, and have
|
||||||
|
`agentClient()` return the shared instance. A single reused Transport pools/reuses connections (≈2
|
||||||
|
idle conns), eliminating the leak. **Preferred.**
|
||||||
|
2. **If a per-call client is kept for any reason**, the Transport must be tamed and closed: set
|
||||||
|
`IdleConnTimeout` (e.g. 30–90 s) + `MaxIdleConnsPerHost`, and `defer client.CloseIdleConnections()`
|
||||||
|
after use. (Reuse (#1) is cleaner and also avoids the per-call TLS handshake cost.)
|
||||||
|
3. Optionally add `MaxConnsPerHost`/keep-alive tuning on the shared Transport as belt-and-suspenders.
|
||||||
|
|
||||||
|
These are pure `felhom-controller` changes (then build + bump + redeploy per the controller workflow);
|
||||||
|
no agent or firewall change.
|
||||||
|
|
||||||
|
### Separate hardening gap (not the cause — note for later)
|
||||||
|
The defense-in-depth control the agent's own config comment calls for — *"a host firewall rule should
|
||||||
|
limit [8443] to the guest bridge subnet"* — is **absent**: `pve-firewall` is disabled and there is no
|
||||||
|
iptables rule scoping 8443. Close this once connectivity is fixed (a host rule allowing
|
||||||
|
`192.168.0.0/24`→`:8443` and dropping others), independent of the leak fix.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Note on the campaign finding #1
|
||||||
|
`unattended-test-campaign-2026-06-22-findings.md` finding #1 correctly **observed** the symptom and
|
||||||
|
**correctly rejected** rebinding to 0.0.0.0 in this follow-up, but its first-pass theory (agent
|
||||||
|
bound to a specific IP ⇒ unreachable) was wrong — disproved by `:8006` reachability to the same IP and
|
||||||
|
by the path working at low uptime. The real cause is the controller connection leak above. No code or
|
||||||
|
config was changed by this diagnosis.
|
||||||
Reference in New Issue
Block a user