R-344: restore the idle-connection timeout our hand-rolled transports lost
gates / gates (push) Successful in 7s
gates / gates (push) Successful in 7s
Every client here pins TLS, so none can use http.DefaultTransport and each hand-rolls its own. A composite literal takes IdleConnTimeout ZERO, which means retain idle connections forever -- not "use a sane default". pbsTargetsFromPVE builds a fresh pbs.Client every cycle and drops the previous one, and an abandoned http.Transport does not close its connections. One stranded socket per cycle, on both sides, forever. Measured: 388 established connections on ep0 over 46 h, 194 per box, zero closed in a 31-minute window. pvestatd and proxmox-backup-client made 162,404 requests in the same window and leaked none. New leaf package internal/httpx owns the default (90s, http.DefaultTransport's own value) and NewTransport, which returns a FRESH transport per call and treats <=0 as "use the default", never "no timeout". pbs.Config gains IdleConnTimeout for tests only. hub and proxmox carried the same missing default and are corrected here for consistency. Neither contributed to the ep0 leak -- both are built once per process and neither talks to ep0:8007. Tests count connections SERVER-side and model the abandonment, so they pin the consequence rather than the field. Two red-proofs, both seen failing: removing the timeout -> "still holds 5 open connection(s), want 0"; DisableKeepAlives -> "3 sequential requests over 3 connection(s), want 1" (the leak test PASSES under that one -- it is the worse-fix guard that catches it). Not released: hand-installed on demo-hp only so demo-felhom stays the control. CHANGELOG heading stays UNRELEASED until the publish is authorised.
This commit is contained in:
@@ -1,3 +1,72 @@
|
||||
## UNRELEASED — v0.130.0 candidate: the agent was the one leaking connections onto the off-site box (2026-08-20, R-344)
|
||||
|
||||
> **Deliberately not a release heading yet, and the `release-complete` gate is doing its job by
|
||||
> requiring that.** This build is **hand-installed on `demo-hp` only** so the fix can be proved
|
||||
> against `demo-felhom` as an untouched control. Publishing it — tag + Gitea package — would put it
|
||||
> on the control box through the self-update path and destroy the experiment. **When the operator
|
||||
> authorises the publish, this heading becomes `## v0.130.0` in the same commit as the tag and the
|
||||
> package**, and the gate then checks it for real. See R-347.
|
||||
|
||||
**What was measured, before anything was changed.** Between 2026-08-18 09:51:22Z and 2026-08-20
|
||||
08:02:13Z, ep0's PBS proxy accumulated **388 established connections** — 194 from each demo box, on a
|
||||
proxy whose descriptor ceiling is 65536 and whose runway at that rate was ~323 days. The connections
|
||||
were held open on **both** sides: ep0 showed 388 while the two boxes showed 194 + 194, at two separate
|
||||
instants, with the same source ports on each side, and **not one closed in a 31-minute window**.
|
||||
`ss -tnp` on the boxes named the holder: **`felhom-agent`**, 194 of 194 on each, one PID.
|
||||
|
||||
**`pvestatd` and `proxmox-backup-client` made 162,404 requests in that window and leaked zero.** They
|
||||
are 99.5% of the traffic to that endpoint and 0% of the leak. The agent made 811 requests — of which
|
||||
387 were `GET .../snapshots` — and leaked 388 sockets. One per call, within one.
|
||||
|
||||
**The defect, and it is two things compounding.**
|
||||
|
||||
- `internal/pbs/client.go` built its transport as a composite literal:
|
||||
`&http.Transport{TLSClientConfig: tlsCfg}`. That takes **`IdleConnTimeout` zero, which does not mean
|
||||
"use a sane default" — it means retain idle keep-alive connections FOREVER.**
|
||||
`http.DefaultTransport` sets 90s; hand-rolling the transport (which every client here must do,
|
||||
because they all pin TLS) silently discards it.
|
||||
- `pbsTargetsFromPVE` (`cmd/felhom-agent/main.go`) builds **a fresh `pbs.Client` every cycle**, as its
|
||||
own doc comment says, and drops the previous one. An abandoned `http.Transport` does **not** close
|
||||
its connections — it becomes unreachable while its `persistConn` read-loop goroutine keeps the
|
||||
socket alive. So each cycle stranded exactly one connection that nothing could ever close.
|
||||
|
||||
The cadences reconcile without fitting: a 900s live-snapshot collect (184.7 cycles in the window) plus
|
||||
a 6-hour verify loop (7.7) predicts 192.4 per box against **194 observed**.
|
||||
|
||||
**The fix is one field, restored to the standard library's own value.** New leaf package
|
||||
`internal/httpx` owns `DefaultIdleConnTimeout = 90 * time.Second` — 90s because that is what
|
||||
`http.DefaultTransport` uses, so there is nothing invented here to justify or tune — and
|
||||
`NewTransport(tlsCfg, idleConnTimeout)`, which returns a **fresh** transport (never shared: each caller
|
||||
pins a different endpoint) and treats a zero or negative timeout as **use the default, never "no
|
||||
timeout"**. `pbs.Config` gains an `IdleConnTimeout` field that production leaves unset; only tests set
|
||||
it, to avoid a 90-second wait.
|
||||
|
||||
**`internal/hub/client.go` and `internal/proxmox/client.go` carried the identical missing default and
|
||||
were corrected in the same pass — but neither contributed to the ep0 leak, and this entry must not be
|
||||
read as three leaks having been found.** Both are built **once per process**, so they held one idle
|
||||
connection for the life of the daemon rather than accumulating, and neither talks to ep0:8007.
|
||||
|
||||
**Tests, and what they deliberately do not assert.** `internal/pbs/client_leak_test.go` counts
|
||||
connections **server-side** and models what `pbsTargetsFromPVE` actually does — build a client, use it
|
||||
once, drop it on the floor — then asserts the connections go away. It does not assert `err == nil` and
|
||||
it does not assert that some field holds some value; both were true of the leaking code.
|
||||
|
||||
- **Red-proof 1 (the fix):** removing `IdleConnTimeout` from `NewTransport` fails the test with
|
||||
*"after 5s the server still holds 5 open connection(s), want 0 (5 dialled in total)"* — the count is
|
||||
in the message, so the failure cannot be mistaken for a timeout with another cause. Reverted.
|
||||
- **Red-proof 2 (the fix that would be worse than the bug):** setting `DisableKeepAlives: true` also
|
||||
makes the leak vanish — by dialling fresh for every request, which on a box polling ~40,000 times a
|
||||
day is strictly worse than what we started with. **The leak test PASSES under that mutation**;
|
||||
`TestPBSClient_KeepAliveStillReuses` is what catches it, failing with *"3 sequential requests over 3
|
||||
connection(s), want 1"*. Reverted.
|
||||
|
||||
**What this release does NOT do.** It does not reduce the poll rate (**R-336 stays open, but re-scoped
|
||||
— it was never the cause of this leak**), it does not refactor `pbsTargetsFromPVE` to cache or reuse
|
||||
clients (a one-line default restores the standard behaviour; a lifecycle refactor adds
|
||||
cache-invalidation questions for no measurable gain), it adds no `CloseIdleConnections` call, and it
|
||||
**does not clear the 388 descriptors already stuck on ep0** — those persist until that proxy restarts,
|
||||
which is not this change's to do.
|
||||
|
||||
## v0.129.0 — a correct code for an earlier package stops being called wrong (2026-08-12, R-311)
|
||||
|
||||
**The measurement this fixes.** On 2026-08-12 a recovery code that provably opens a RETAINED package
|
||||
|
||||
Reference in New Issue
Block a user