From 57dd62b097f6d825b52d8afc79e8ea7c8c0b1bb3 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Thu, 20 Aug 2026 12:39:32 +0200 Subject: [PATCH] R-344 fixed and proven on both boxes: ep0 is back to fd 17 from 415 P1, outcome (i) in one second: replacing the agent on demo-hp released exactly its 199 established connections (ep0 fd 415 -> 216). CLOSE-WAIT stayed 0, so outcome (ii) does not exist and gets no row -- ep0 reaps on peer FIN correctly, and the 543 CLOSE-WAIT at the 08-18 wedge has another explanation. P2, 1.03 h (operator closed the >=4 h window early, so no daily rate is extrapolated): control +4, fixed +0, with each box making exactly 4 /snapshots and 4 /version calls. Same cadence, same work: 4 cycles -> 4 leaks vs 4 cycles -> 0. The fixed box's cycles are in ep0's log, so the zero is the fix and not a stopped agent. P3: the second box took ep0 from 220 to 17 fd in under two seconds. 17 is precisely the t0 baseline of 2026-08-18 09:51:22Z. Corrects a claim this session made earlier the same day: the accumulated descriptors did NOT need an ep0 proxy restart. They were held on both sides. ep0 was read-only throughout; its PID never changed. R-344 updated and left OPEN (unpublished is not delivered). R-336 re-scoped -- its old next-step would have fixed nothing while looking like a failed fix, and it is now a scaling row (~25 req/s at fifty customers). R-347 filed for the delivery gap (Viktor decides). R-348 filed: an agent restart blanks the reported backup list for ~18 h and the Store comment calls it unaffected -- blinds no alarm, checked not assumed. --- REPORT-agent-transport-leak.md | 174 +++++++++++++++++ STATUS.md | 22 ++- ...-ep0-established-connections-2026-08-20.md | 179 ++++++++++++++++++ .../finding-zero-backups-hub-view.txt | 87 +++++++++ .../p1-restart-observation.txt | 41 ++++ .../p2-arithmetic.txt | 34 ++++ .../p2-divergence-window.txt | 9 + .../p2-final-and-positive-observable.txt | 19 ++ .../p4.3-fleet-sanity.txt | 20 ++ .../p4.3-hub-hosts.txt | 11 ++ .../p5-second-box.txt | 21 ++ .../p5-settle.txt | 11 ++ .../stop1-ep0-before.txt | 14 ++ documentation/backlog/OPEN-ITEMS.md | 6 +- 14 files changed, 641 insertions(+), 7 deletions(-) create mode 100644 REPORT-agent-transport-leak.md create mode 100644 documentation/audits/evidence-agent-transport-leak-2026-08-20/finding-zero-backups-hub-view.txt create mode 100644 documentation/audits/evidence-agent-transport-leak-2026-08-20/p1-restart-observation.txt create mode 100644 documentation/audits/evidence-agent-transport-leak-2026-08-20/p2-arithmetic.txt create mode 100644 documentation/audits/evidence-agent-transport-leak-2026-08-20/p2-divergence-window.txt create mode 100644 documentation/audits/evidence-agent-transport-leak-2026-08-20/p2-final-and-positive-observable.txt create mode 100644 documentation/audits/evidence-agent-transport-leak-2026-08-20/p4.3-fleet-sanity.txt create mode 100644 documentation/audits/evidence-agent-transport-leak-2026-08-20/p4.3-hub-hosts.txt create mode 100644 documentation/audits/evidence-agent-transport-leak-2026-08-20/p5-second-box.txt create mode 100644 documentation/audits/evidence-agent-transport-leak-2026-08-20/p5-settle.txt create mode 100644 documentation/audits/evidence-agent-transport-leak-2026-08-20/stop1-ep0-before.txt diff --git a/REPORT-agent-transport-leak.md b/REPORT-agent-transport-leak.md new file mode 100644 index 00000000..84888afe --- /dev/null +++ b/REPORT-agent-transport-leak.md @@ -0,0 +1,174 @@ +# REPORT — R-344: the agent's leaked PBS connections, fixed and proven on both boxes (2026-08-20) + +**Outcome: the fix works, measured three independent ways, and ep0 is back to its baseline of 17 file +descriptors from 415.** Both demo boxes now run agent **0.130.0**. **Nothing is published** — that is +the one decision left, filed as R-347. + +## 1. Confirmed baselines + +| repo | in | out | +|---|---|---| +| `felhom-agent` | `f17ed11` — v0.129.0 | **`ede49b6`** — 0.130.0, UNRELEASED | +| `felhom.eu` | `9299f85` | docs + register only | + +Both clean and equal to `origin/main` before each build. Agent version in: 0.129.0 on both boxes. +Out: **0.130.0 on both**, confirmed from the hub, not from the boxes' own `--version`. + +## 2. The diff + +| file | symbol | change | +|---|---|---| +| `internal/httpx/transport.go` | **new package** | `DefaultIdleConnTimeout = 90s`; `NewTransport(tlsCfg, idle)` — fresh transport per call, `<= 0` means **use the default, never "no timeout"** | +| `internal/httpx/transport_test.go` | new | zero/negative → default; the constant is read off `http.DefaultTransport`; freshness; TLS config preserved | +| `internal/pbs/client.go` | `Config`, `NewClient` | `IdleConnTimeout` field (tests only); transport via `httpx` | +| `internal/pbs/client_leak_test.go` | new | Scenarios A/C + the production-default pin | +| `internal/hub/client.go` | `NewClient` | via `httpx` — **consistency only, did not contribute to the leak** | +| `internal/proxmox/client.go` | `NewClient` | same | +| `cmd/felhom-agent/main.go` | `version` | 0.92.1 → 0.130.0 (ldflags default) | +| `REUSE.md`, `CHANGELOG.md` | — | `httpx.NewTransport` entry; the release note | + +Commit `ede49b6` on `main`. `grep '&http.Transport{'` now matches only `httpx` itself. + +**A check made before trusting the fix:** `doBody` already reads the body to completion and closes it, +so the connection genuinely reaches the idle pool. Had it not, the idle timeout would have been the +wrong fix entirely. + +## 3. Tests and the red-proofs + +`go build ./... && go vet ./... && go test ./...` → **30 packages, 0 failures.** +`python3 scripts/agent_gates.py` → **all 4 gates OK.** + +The leak test counts connections **server-side** and models what `pbsTargetsFromPVE` does — build a +client, use it once, drop it. It deliberately does **not** assert `err == nil` or that a field holds a +value; both were true of the leaking code (the R-224 lesson). + +| red-proof | mutation | seen failing with | +|---|---|---| +| 1 — the fix | remove `IdleConnTimeout` | *"abandoned pbs.Clients: after 5s the server still holds 5 open connection(s), want 0 (5 dialled in total)"* — the count is in the message, so it cannot be a timeout with another cause | +| 2 — the worse fix | `DisableKeepAlives: true` | **the leak test PASSES.** Caught only by `TestPBSClient_KeepAliveStillReuses`: *"3 sequential requests over 3 connection(s), want 1"* | + +**Red-proof 2 is the load-bearing one: Scenario A alone would have accepted a fix that made the problem +worse** — no leak, at the price of a fresh dial for every one of ~40,000 daily requests. Both mutations +reverted, tree re-verified clean. + +## 4. §4 quoted, beside the results + +> **P1** — *"(i) ep0's fd count falls by ≈194 within seconds … (ii) … ≈194 sockets convert ESTAB → +> CLOSE-WAIT … (iii) neither — the count barely moves. Then the ownership attribution is wrong and the +> finding must be withdrawn."* +> **P2** — *"demo-hp (fixed): ≈ 0 … demo-felhom (control, untouched): ≈ 4 per hour → ≈ 16 over 4 h."* +> **P3** — *"ep0's overall leak rate should fall from ≈ 200/day to ≈ 100/day while one box is fixed, and +> to ≈ 0/day after Part 5."* + +## 5. P1 — **outcome (i)**, in one second + +| | ep0 fd | ESTAB | demo-felhom | demo-hp | CLOSE-WAIT | +|---|---|---|---|---|---| +| T-0 `09:15:46Z` | 415 | 398 | 199 | **199** | 0 | +| T+1s `09:15:47Z` | **216** | **199** | 199 | **0** | **0** | +| T+60s `09:16:42Z` | 218 | 201 | 199 | 2 | 0 | + +**Outcome (ii) did not occur, so it gets no register row.** Not one socket converted to `CLOSE-WAIT`: +ep0 reaps on peer FIN correctly. That also means the 543 `CLOSE-WAIT` at the 2026-08-18 wedge has some +other explanation and is **not** evidence of a second defect on the protected machine — a finding in the +negative, worth the sixty seconds it cost. + +Ownership is now proven a **third** independent way: what dies with the process, agreeing with +`ss -tnp` and with the access-log user agent. At 133 s uptime the fixed box held **0** connections. + +## 6. P2 — divergence + +**Window 09:15:47Z → 10:17:52Z = 1.03 h. You closed the ≥4 h window early**, so no daily rate is +extrapolated and none is needed. + +| box | agent | start | end | delta | per hour | predicted | +|---|---|---|---|---|---|---| +| `demo-felhom` CONTROL | 0.129.0 | 199 | 203 | **+4** | 3.87 | ≈4 | +| `demo-hp` FIXED | 0.130.0 | 0 | 0 | **+0** | 0.00 | ≈0 | + +**The assumption-free statement.** ep0's access log counts the opportunities: each box made **exactly 4 +`/snapshots` and 4 `/version` calls** in the window. + +> **control: 4 cycles → 4 leaks. fixed: 4 cycles → 0 leaks.** + +**Positive observable (standing rule 3):** a zero leak is equally consistent with "the agent stopped +working" — it did not; its four cycles are in ep0's log. The boxes' other traffic is near-identical +(`libwww-perl` 924 vs 926, `proxmox-backup-client` 898 vs 898), so **the only difference between them +is the binary**. Poisson alone gives P(0 | λ=4) = **1.8%**, which is suggestive rather than conclusive +and is not relied on alone. + +## 7. P3 — the second box, and the backlog clearing itself + +`demo-felhom` upgraded `10:18:56Z` on your word. + +| | ep0 fd | ESTAB | CLOSE-WAIT | +|---|---|---|---| +| T-0 `10:18:55Z` | 220 | 203 | 0 | +| **T+2s** | **17** | **0** | 0 | +| settled 10:30–10:35Z | **17–19** | 0–2 | 0 | + +**17 is precisely ep0's `t0` baseline** (fd 17, ESTAB 0, 2026-08-18 09:51:22Z), and it returns to 17 +between poll cycles — the "healthy proxy near 20 fds" the incident document named. Predicted ≈0/day +residual; **observed the baseline itself.** + +**A correction, made within the hour it was written.** My STOP 1 report and the first CHANGELOG draft +said *"does not clear the 388 descriptors already stuck on ep0 — those persist until that proxy +restarts."* **Wrong.** They were held on both sides; restarting the agents released every one. **ep0 was +read-only throughout and its proxy PID never changed (551655).** Corrected in the CHANGELOG, the audit +document and R-344 rather than quietly edited. + +## 8. Fleet sanity + +Hub reports **0.130.0 on both** boxes. **No `floor held`** line (0.130.0 > golden MinAgent 0.129.0). +**No `pbsdr_box_unreachable` / `offsite_box_unreachable`** during any window. Positive observable +rather than the absent one: the PBS-DR gauge kept refreshing (`3.7% full (3.7 GB of 97.9 GB)`) and host +reports kept landing from both boxes throughout. + +## 9. What is NOT done + +- **Not published.** No package, no tag, no manifest or floor field touched, no self-update staged. + **A box installed from the current image still ships the leaking agent** — **R-347**, your call. +- The CHANGELOG heading is `## UNRELEASED — v0.130.0 candidate`. The `release-complete` gate convicted + on `## v0.130.0` because there is no tag and no package, and **it was right to**. I did not use + `--no-verify`; I made the heading stop claiming a release that has not happened. It flips to + `## v0.130.0` in the same commit as the tag. +- Poll rate unchanged (**R-336**, re-scoped). `pbsTargetsFromPVE` not refactored; no + `CloseIdleConnections` added. ep0 not touched. +- **The 388 descriptors ARE cleared** — see §7. This is the one item the prompt expected to remain + outstanding, and it did not. + +## 10. Register + +- **R-344** — updated with the fix, P1's outcome named, and P2/P3's numbers. **Left OPEN**, because a fix + on two boxes by hand is not delivered. +- **R-336 — re-scoped.** Its new next-step cell, verbatim: *"**NEW ACCEPTANCE CRITERION, since the old + one is void:** the fd count is NOT the observable for this row any more — that belongs to R-344 and is + already satisfied. Measure the REQUEST RATE at ep0's access log, and state the projected rate at the + target customer count."* The row now records explicitly that its old next-step **would have "fixed" + nothing while looking like a failed fix**, and re-scopes it to what it is: ~85,000 requests/day to a + weekly-write DR endpoint, ≈**25 requests/second at fifty customers** against a CX33. +- **R-347 (new)** — the delivery gap. Owner: **Viktor decides**, CC executes. +- **R-348 (new)** — an agent restart blanks the reported backup list for up to ~18 h, and the `Store` + comment calls backups *"unaffected"*. **Blinds no alarm** — checked, not assumed: the hub's + `backupEvidenceLookback` scans 7 days for exactly this case, and `pbs_snapshots` stayed populated. +- **No P1(ii) row**, because outcome (ii) did not occur. +- **R-346** — this run anchored on the measured `t0` (fd 17 at 2026-08-18 09:51:22Z), never on a systemd + timestamp, so the 5 h 56 m discrepancy did not touch these numbers. + +## 12. Observations + +- **The closure refactor is not worth doing — recommend leaving it.** With the idle timeout restored an + abandoned client's connection is gone in 90 s, so the standing population is bounded at about one + connection per box instead of growing without limit. Caching clients would add cache-invalidation + questions (a storage's fingerprint, token or namespace can change under it) for no observable gain. +- **Three other `http.Transport` defaults are still missing and were left alone:** `MaxIdleConns`, + `TLSHandshakeTimeout` (0 = no limit; `DefaultTransport` uses 10 s) and `ExpectContinueTimeout`. None + accumulates, and every client bounds its request with `http.Client.Timeout`. `TLSHandshakeTimeout` is + the only one with a plausible failure mode — a stalled handshake over the tunnel, bounded today only + by the outer 30 s. Not changed, because widening the diff would have made this measurement + unattributable. Worth a look on its own terms; not a defect. +- **A measurement error of mine, recorded because it nearly cost four hours.** The first P2 sampler + reported both per-box columns as 0 while the totals were right: `ss` prints `[::ffff:10.77.0.2]:port` + and my pattern expected `10.77.0.2:`. Caught 15 minutes in, because a 0/0 split cannot sum to 199. + Fixed, then **one sample proved by hand before committing the window** — which is what should have + happened first. diff --git a/STATUS.md b/STATUS.md index 93390d12..d44c772c 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,7 +1,7 @@ # STATUS — what works, what's broken, what's next -**Updated 2026-08-20 (morning — we looked at who was actually holding the off-site box's connections -open, and it turned out to be us).** +**Updated 2026-08-20 (midday — the leak was ours, and it is fixed and proven on both demo machines; +it is not yet published).** > **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates > part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.** @@ -124,6 +124,19 @@ record with no machine** — created 13 August, no host, no backups, nothing to rate down, then watch the count stop climbing — **would have changed nothing and looked like a failed fix.** The count still says about a year to the ceiling. **Nothing is broken today**; this is a deadline we now know where to aim at. *(register: R-344, and R-336 re-ranked)* +- **FIXED the same day, and proven on the machines** (R-344). One line of our code: the agent had the + idle-connection timer switched off, so nothing ever retired the connections it abandoned. Turning it + back on to the standard 90 seconds fixed it. **The proof cost nothing clever:** we put the fix on one + machine and left the other alone, and in the same hour the untouched one leaked 4 more connections + while the fixed one leaked none — with both doing exactly the same four rounds of work. **The + off-site box is back to 17 open connections, its normal resting number, down from 415.** All of the + built-up connections released themselves when the agents restarted; the off-site box was only ever + read from, never touched. *(register: R-344)* +- **The fix is on the two demo machines by hand and NOT published yet** (R-347). A machine installed + from today's image still gets the old, leaking agent. That was deliberate — publishing it mid-test + would have contaminated the comparison — and the reason has now expired. **It is not urgent:** a new + machine would take the better part of a year to matter, and any agent update clears the build-up. + **Publishing is your call**, and it needs the operator-only artifact screen at the end. - **The off-site box was updated, and it did not help — as expected** (R-341). On your ruling we installed the newer backup software for the practice, having first read its release notes and found **nothing** about the fault we have. The update went cleanly and everything works, but the leak @@ -163,9 +176,8 @@ record with no machine** — created 13 August, no host, no backups, nothing to ## Working on next -**One decision is waiting for you, at the top of `REPORT-spike-ep0-connections.md`:** whether to run the -overnight measurement window on `demo-hp` tonight, and which of two versions of it. Nothing was changed on -any machine in this session — every reading was read-only. After that: the three remaining +**One decision is waiting for you:** whether to publish agent 0.130.0 so machines other than the two +demo boxes get the fix (R-347). It needs the artifact screen, which only you can drive. After that: the three remaining R-264 readers, now that one has been built and we know what one costs; R-317 (one line in the agent); R-327 (decide what the naming claim's status should be); then the 2026-08-09 batch (R-279 … R-292), still untriaged against everything since. diff --git a/documentation/audits/SPIKE-ep0-established-connections-2026-08-20.md b/documentation/audits/SPIKE-ep0-established-connections-2026-08-20.md index ef1bc4b3..7ea82477 100644 --- a/documentation/audits/SPIKE-ep0-established-connections-2026-08-20.md +++ b/documentation/audits/SPIKE-ep0-established-connections-2026-08-20.md @@ -321,3 +321,182 @@ result.** That is the substantive change to R-336 this spike delivers. explain it beyond noting it. - **Whether a reader (restore-test) connection also leaks** was not separated out. At most 8 of 388 sockets could be involved (≤2%), which is inside the noise, so it was not pursued. + +--- + +# Fix and proof — 2026-08-20, later the same day (R-344) + +Run under `Claude Code Prompt: fix the agent's leaked PBS connections, then prove it on one box`, +which **cancelled Phase C**. Evidence: `evidence-agent-transport-leak-2026-08-20/`. + +**Phase C was never run, and cancelling it was right.** It would have quietened `pvestatd` on `demo-hp` +to test whether the leak tracks the request rate. The spike had already shown `pvestatd` leaks nothing, +so the honest prediction was a null result. The experiment that replaced it — deploy the fix to one box +and leave the other as a control — proves the fix and re-proves the mechanism in one window, and it +opens with a free, immediate, decisive observation Phase C could not have produced. + +## The fix + +One field, restored to the standard library's own value. New leaf package `felhom-agent/internal/httpx` +owns `DefaultIdleConnTimeout = 90 * time.Second` — 90s because that is `http.DefaultTransport`'s value, +so there is nothing invented to justify — and `NewTransport(tlsCfg, idle)`, which returns a **fresh** +transport every call (never shared: each caller pins a different endpoint) and treats a zero or negative +timeout as **use the default, never "no timeout"**. Wired into all three hand-rolled transports; +`grep '&http.Transport{'` now matches only `httpx` itself. + +`internal/hub/client.go` and `internal/proxmox/client.go` carried the identical missing default and were +corrected in the same pass. **Neither contributed to the ep0 leak** — both are built once per process, +so they held one idle connection for the life of the daemon rather than accumulating, and neither talks +to `ep0:8007`. This must not be read as three leaks having been found. + +**A check made before trusting the fix:** `doBody` already reads the response body to completion and +closes it, so the connection genuinely reaches the idle pool. Had it not, `IdleConnTimeout` would have +been the wrong fix and the leak would have been an un-drained body instead. + +### The tests, and the two red-proofs + +`internal/pbs/client_leak_test.go` counts connections **server-side** and models what +`pbsTargetsFromPVE` actually does — build a client, use it once, drop it on the floor. It deliberately +does not assert `err == nil` and does not assert that a field holds a value; **both were true of the +leaking code**. This is the R-224 lesson applied: assert the consequence, not the mechanism. + +| red-proof | mutation | result | +|---|---|---| +| 1 — the fix | remove `IdleConnTimeout` from `NewTransport` | **FAILS**: *"abandoned pbs.Clients: after 5s the server still holds 5 open connection(s), want 0 (5 dialled in total)"* — the count is in the message, so the failure cannot be a timeout with another cause | +| 2 — the worse fix | `DisableKeepAlives: true` | **the leak test PASSES.** Disabling keep-alive also stops the leak — by dialling fresh for every one of ~40,000 daily requests. `TestPBSClient_KeepAliveStillReuses` is what catches it: *"3 sequential requests over 3 connection(s), want 1"* | + +Red-proof 2 is the load-bearing one: **Scenario A alone would have accepted a fix that made the problem +worse.** Both mutations were reverted and the tree re-verified clean. + +Green gate: `go build ./... && go vet ./... && go test ./...` → 30 packages, 0 failures. +`python3 scripts/agent_gates.py` → all 4 gates OK. + +## P1 — the restart observation + +Predicted in advance as one of three outcomes. **Outcome (i), within one second.** + +| | ep0 fd | ESTAB | from `demo-felhom` | from `demo-hp` | CLOSE-WAIT | +|---|---|---|---|---|---| +| T-0 `09:15:46Z` | 415 | 398 | 199 | **199** | 0 | +| T+1s `09:15:47Z` | **216** | **199** | 199 | **0** | **0** | +| T+60s `09:16:42Z` | 218 | 201 | 199 | 2 | 0 | + +**Outcome (ii) did not occur and gets no register row.** Not one socket converted to `CLOSE-WAIT`: ep0 +reaps on the peer's FIN correctly, so the 543 `CLOSE-WAIT` seen at the 2026-08-18 wedge has some other +explanation and is not evidence of a second defect on the protected machine. That is a finding in the +negative and it was worth the sixty seconds it cost. + +**Ownership is now established by a third independent route.** Killing our process released exactly +`demo-hp`'s 199 descriptors — agreeing with `ss -tnp` on the box and with the access-log user agent. + +**The new code retires connections, observed directly:** at 133 s uptime `demo-hp` held **0** +connections; the 2 it opened after restart were closed at the 90 s idle mark. + +## P2 — the divergence window + +**Window 09:15:47Z → 10:17:52Z = 1.03 h. The task specified ≥ 4 h; the operator closed it early.** +The conclusion below therefore does not rest on an extrapolated daily rate, and none is used. + +| box | agent | start | end | delta | per hour | +|---|---|---|---|---|---| +| `demo-felhom` — CONTROL | 0.129.0 | 199 | 203 | **+4** | 3.87 | +| `demo-hp` — FIXED | 0.130.0 | 0 | 0 | **+0** | 0.00 | + +Predicted: control ≈4/hour, fixed ≈0. **Control observed 3.87/hour.** + +**The assumption-free statement, and the reason the short window still settles it.** ep0's access log +counts the poll cycles directly: in that window **each box made exactly 4 `GET .../snapshots` calls and +4 `GET /version` calls**. Same cadence, same work, the same four chances to leak. + +> **control: 4 cycles → 4 leaked connections. fixed: 4 cycles → 0 leaked connections.** + +**The positive observable, per standing rule 3.** A zero leak is equally consistent with "fixed" and +with "the agent stopped working". It is the former: the fixed box's four poll cycles are in ep0's log, +alongside the control's four. The rest of the two boxes' traffic is near-identical in the window — +`libwww-perl` 924 vs 926, `proxmox-backup-client` 898 vs 898 — so **the only difference between them is +the agent binary**. + +Poisson alone would give P(0 leaks | old rate, λ=4) = **1.8%**, which is suggestive rather than +conclusive. It does not stand alone: P1 released 199 descriptors instantly, the mechanism is identified +at `file:line` and pinned by a red-proofed test, and the fixed box was directly observed returning to 0. + +## P3 — the second box, and the accumulated leak clearing itself + +`demo-felhom` — the former control — was upgraded at `10:18:56Z` on the operator's word. + +| | ep0 fd | ESTAB | CLOSE-WAIT | +|---|---|---|---| +| T-0 `10:18:55Z` | 220 | 203 | 0 | +| **T+2s `10:18:56Z`** | **17** | **0** | **0** | +| T+65s `10:19:58Z` | 20 | 3 | 0 | +| settled, 10:30–10:35Z | **17–19** | 0–2 | 0 | + +**fd 17 is precisely ep0's `t0` baseline** — fd 17, ESTAB 0, recorded at 2026-08-18 09:51:22Z. The proxy +sits at 17–19 and returns to 17 between poll cycles, which is the "healthy proxy near 20 fds" the +incident document named as the positive observable. **ep0 was read-only throughout and its proxy PID +never changed (551655).** + +**This corrects a sentence written earlier the same day**, in the v0.130.0 CHANGELOG draft and in the +STOP 1 report: *"does not clear the 388 descriptors already stuck on ep0 — those persist until that +proxy restarts."* **Wrong, and measured wrong within the hour.** The descriptors were held on *both* +sides; closing either side ends them. Restarting the two agents released all of them. Nothing on the +protected machine had to be touched, and nothing was. + +## Fleet sanity + +- Hub reports `demo-hp` **0.130.0** and `demo-felhom` **0.130.0**; 0.130.0 is above the golden's + MinAgent of 0.129.0, and **no `floor held` line appeared** for either box. +- **No `pbsdr_box_unreachable` / `offsite_box_unreachable` event** fired during any window. +- Positive observable rather than the absent one: the hub's PBS-DR gauge kept refreshing on schedule + (`3.7% full (3.7 GB of 97.9 GB)`) and host reports kept landing from both boxes throughout. + +## An unrelated finding the deploy exposed — R-348 + +The **first two host reports after an agent restart carry `0 backups`**, while the box's own +`pvesm list` shows backups present on both tiers. `internal/backup/store.go`'s `Store` is in-memory and +its `byTarget` map is repopulated only when a backup **runs** — daily for the local tier, weekly for +offsite — so the field reads 0 for up to ~18 h after any restart. `restore_tests` did **not** blank, +because that half has a durable on-disk companion (`RestoreTestState`, R-189). + +**It blinds no alarm, and that was checked rather than assumed.** `hub/internal/monitor/deadline.go` +already scans back over stored reports with a 7-day `backupEvidenceLookback`, whose comment names this +exact case — *"when the LATEST report carries none... and against an agent that stayed restarted for +days"* — and `pbs_snapshots` stayed populated at 2 regardless. So this is an observability wart, not a +safety hole. + +**But the `Store` comment is misleading in a way this project has a rule about.** It reads *"Backups are +unaffected — their freshness has a ground truth on the storage (R-84)"*. That is true of the +**consequence** and false of the **field**, and a future reader may take it as a guarantee the field +stays populated. Filed as R-348. + +## What this run did NOT do + +- **Did not publish.** No Gitea package, no tag, no `artifact_agent_version` / `artifact_min_agent` / + `artifact_golden_version` / `min_controller_version` change, no staged self-update. **A box installed + from the current image still carries the leaking agent** — filed as **R-347**, and the CHANGELOG + heading stays `## UNRELEASED` until that decision is taken. +- **Did not reduce the poll rate.** R-336 stays open, **re-scoped**: it was never the cause of this leak. +- **Did not refactor `pbsTargetsFromPVE`** to cache or reuse clients, and added no + `CloseIdleConnections` call. See Observations. +- **Did not touch ep0** — no restart, no config, no package, no nftables. Reads only. +- Did not trigger a backup, restore or verify to generate traffic; did not contact `peti-felhom`. + +## Observations + +- **The closure refactor is not worth doing, and the measurement is why.** With the idle timeout + restored, an abandoned client's connection is gone in 90 s, so the standing population is bounded at + roughly one connection per box rather than growing without limit. Caching clients would add + cache-invalidation questions — a storage's fingerprint, token or namespace can change under it — for + no observable gain. Recommend leaving it. +- **Three other `http.Transport` defaults are still missing** and were deliberately left alone: + `MaxIdleConns` (0 = unlimited; `DefaultTransport` uses 100), `TLSHandshakeTimeout` (0 = no limit; + `DefaultTransport` uses 10 s) and `ExpectContinueTimeout`. None of them accumulates anything, and every + client bounds its whole request with `http.Client.Timeout`, so none is a leak. `TLSHandshakeTimeout` + is the only one with a plausible failure mode — a stalled handshake over the tunnel, bounded today + only by the outer 30 s client timeout. **Not changed here, because widening the diff would have made + this measurement unattributable.** Worth a look on its own terms; not filed as a defect. +- **A measurement error of mine, recorded because it nearly cost four hours.** The first P2 sampler + reported both per-box columns as 0 while the totals were right: `ss` prints + `[::ffff:10.77.0.2]:port`, and the pattern expected `10.77.0.2:`. Caught 15 minutes in by noticing + that a 0/0 split could not sum to 199. Fixed, then **one sample was proved by hand before committing + the window to it** — the check that should have happened first. diff --git a/documentation/audits/evidence-agent-transport-leak-2026-08-20/finding-zero-backups-hub-view.txt b/documentation/audits/evidence-agent-transport-leak-2026-08-20/finding-zero-backups-hub-view.txt new file mode 100644 index 00000000..58db60bc --- /dev/null +++ b/documentation/audits/evidence-agent-transport-leak-2026-08-20/finding-zero-backups-hub-view.txt @@ -0,0 +1,87 @@ +=== the 0-backups blip: what does the hub SHOW, and did anything ALARM? === +captured 2026-08-20T09:31:38+00:00 +demo-hp agent restarted 2026-08-20 09:15:46Z; two reports since, both '0 backups' + +--- hub host page, the DR/Backup region (tags stripped) --- + + + + + + DR / Backup + + + DR Recipe + present + + + Key Escrow + present · 1 superseded escrow blob(s) retained + + + + + + Console access + + + + User + root@pam + + + Password set + 29d ago + + + + •••••••••••••••• + Reveal + Copy + + + Break-glass credential for the PVE web console at https://<host-ip>:8006 (realm: Linux PAM standard authentication). Copy puts it straight on the clipboard without showing it; Reveal displays it for 60 s. Either one is recorded on the customer's event timeline. Last vaulted value — if root@pam was changed on the box without re-vaulting, this is stale. + + + + + + + + + + var consolePwState = typeof consolePwState !== 'undefined' ? consolePwState : {}; + var consolePwMask = '••••••••••••••••'; + function consoleHint(hostID, msg) { + var h = document.getElementById('console-hint-' + hostID); + if (h) { h.textContent = msg; } + } + function maskConsolePassword(hostID) { + var st = consolePwState[hostID]; + if (st && st.timer) { clearTimeout(st.timer); } + consolePwState[hostID] = null; + var code = document.getElementById('console-pw-' + hostID); + if (code) { code.textContent = consolePwMask; } + var reveal = document.getElementById('console-reveal-' + hostID); + if (reveal) { reveal.textContent = 'Reveal'; } + +--- CONTROL: same region for demo-felhom (still 0.129.0, never restarted) --- +763: DR / Backup +764- +765- +766- DR Recipe +767- present +768- +769- +770- Key Escrow +771- present · 3 superseded escrow blob(s) retained +772- +773- +774- +775- +776- +777- Console access +778- +779- +780- +781- User diff --git a/documentation/audits/evidence-agent-transport-leak-2026-08-20/p1-restart-observation.txt b/documentation/audits/evidence-agent-transport-leak-2026-08-20/p1-restart-observation.txt new file mode 100644 index 00000000..e40af894 --- /dev/null +++ b/documentation/audits/evidence-agent-transport-leak-2026-08-20/p1-restart-observation.txt @@ -0,0 +1,41 @@ +=== P1 — the restart observation. demo-hp gets 0.130.0; demo-felhom is the untouched control. === +DooPlex UTC now: 2026-08-20T09:15:45+00:00 + +--- T-0 ep0 IMMEDIATELY BEFORE --- +t=09:15:46Z pid=551655 fd=415 | 398 ESTAB 1 LISTEN | peers: 199 10.77.0.2 199 10.77.0.3 + +--- installing + restarting felhom-agent on demo-hp --- +old MainPID: 3199936 +RESTART AT: 09:15:46.964123161Z +new MainPID: 375451 +felhom-agent 0.130.0 + +--- T+~5s --- +t=09:15:47Z pid=551655 fd=216 | 199 ESTAB 1 LISTEN | peers: 199 10.77.0.2 +--- T+~15s --- +t=09:15:55Z pid=551655 fd=218 | 201 ESTAB 1 LISTEN | peers: 199 10.77.0.2 2 10.77.0.3 +--- T+~30s --- +t=09:16:11Z pid=551655 fd=218 | 201 ESTAB 1 LISTEN | peers: 199 10.77.0.2 2 10.77.0.3 +--- T+~60s --- +t=09:16:42Z pid=551655 fd=218 | 201 ESTAB 1 LISTEN | peers: 199 10.77.0.2 2 10.77.0.3 + +--- demo-hp client side after restart --- +ESTAB to 8007: 2 +active +felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/bin/lxc-info -n 9201 -p -H +pam_unix(sudo:session): session opened for user root(uid=0) by (uid=999) +pam_unix(sudo:session): session closed for user root +felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/sbin/blkid -p -o export /dev/nvme0n1 +pam_unix(sudo:session): session opened for user root(uid=0) by (uid=999) +pam_unix(sudo:session): session closed for user root +felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/bin/lsblk -J -o NAME,FSTYPE,PTTYPE,MOUNTPOINT /dev/nvme0n1 +pam_unix(sudo:session): session opened for user root(uid=0) by (uid=999) +pam_unix(sudo:session): session closed for user root +felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/bin/lxc-info -n 9201 -p -H +pam_unix(sudo:session): session opened for user root(uid=0) by (uid=999) +pam_unix(sudo:session): session closed for user root + +--- demo-felhom (CONTROL — untouched) client side --- +version: felhom-agent 0.129.0 +MainPID: 2596329 +ESTAB to 8007: 199 diff --git a/documentation/audits/evidence-agent-transport-leak-2026-08-20/p2-arithmetic.txt b/documentation/audits/evidence-agent-transport-leak-2026-08-20/p2-arithmetic.txt new file mode 100644 index 00000000..d20017b7 --- /dev/null +++ b/documentation/audits/evidence-agent-transport-leak-2026-08-20/p2-arithmetic.txt @@ -0,0 +1,34 @@ +P2 ARITHMETIC — the divergence window + + window 2026-08-20T09:15:47Z -> 10:17:52Z = 3725 s = 1.0347 h + NOTE: this is 1.03 h, NOT the >=4 h the task specified. The operator closed it early. + The conclusion below therefore does NOT rest on an extrapolated daily rate. + +box agent start end delta per hour +demo-felhom (CONTROL) 0.129.0 199 203 +4 3.87 +demo-hp (FIXED) 0.130.0 0 0 +0 0.00 + +PREDICTED (pre-registered, P2 table): control ~4/hour; fixed ~0. +OBSERVED: control 3.87/hour, fixed 0. Control matches the prediction. + +THE ASSUMPTION-FREE STATEMENT — matched opportunities, counted from ep0's access log: + Both agents made EXACTLY 4 GET .../snapshots calls and 4 GET /version calls in this window. + Same cadence, same work, same number of chances to leak. + control: 4 cycles -> 4 leaked connections + fixed : 4 cycles -> 0 leaked connections + No rate extrapolation is needed for that, and none is used. + + Poisson: P(observing 0 leaks | the OLD rate, lambda=4) = e^-4 = 1.83% + On its own that is suggestive, not proof. It is not on its own: + - P1 released exactly 199 descriptors the instant the old process died; + - the mechanism is identified at file:line and pinned by a red-proofed unit test; + - the fixed box was directly observed returning to 0 connections 133 s after restart. + +P3 — total rate with ONE box fixed + observed 92.8 fd/day (predicted ~100/day, from ~200/day with both boxes leaking) + Poisson on n=4 is +/-2, so the 2-sigma band is 0..186/day — the prediction sits inside it. + A 4-descriptor window cannot pin a daily rate tighter than that, and this does not pretend to. + +CONTROL INTEGRITY: the two boxes' OTHER traffic is unchanged and near-identical in this window -- + libwww-perl 924 (.2) vs 926 (.3); proxmox-backup-client 898 vs 898. + So the only thing that differs between them is the agent binary. diff --git a/documentation/audits/evidence-agent-transport-leak-2026-08-20/p2-divergence-window.txt b/documentation/audits/evidence-agent-transport-leak-2026-08-20/p2-divergence-window.txt new file mode 100644 index 00000000..9a13c127 --- /dev/null +++ b/documentation/audits/evidence-agent-transport-leak-2026-08-20/p2-divergence-window.txt @@ -0,0 +1,9 @@ +=== P2 — the divergence window === +fixed : demo-hp 10.77.0.3 agent 0.130.0 (restarted 2026-08-20 09:15:46Z) +control: demo-felhom 10.77.0.2 agent 0.129.0 (untouched) +ep0 read-only; sample every 900 s + +utc fd estab ctrl(.2) fix(.3) proxyPID CLOSEWAIT +2026-08-20T09:33:36Z 217 200 200 0 551655 0 +2026-08-20T09:48:37Z 218 201 201 0 551655 0 +2026-08-20T10:03:37Z 219 202 202 0 551655 0 diff --git a/documentation/audits/evidence-agent-transport-leak-2026-08-20/p2-final-and-positive-observable.txt b/documentation/audits/evidence-agent-transport-leak-2026-08-20/p2-final-and-positive-observable.txt new file mode 100644 index 00000000..d376172b --- /dev/null +++ b/documentation/audits/evidence-agent-transport-leak-2026-08-20/p2-final-and-positive-observable.txt @@ -0,0 +1,19 @@ +=== P2 FINAL SAMPLE + the positive observable === +UTC: 2026-08-20T10:17:52+00:00 +proxyPID=551655 fd=220 estab=203 ctrl(.2)=203 fix(.3)=0 CLOSE-WAIT=0 + +--- POSITIVE OBSERVABLE: is the FIXED agent still polling at all? --- +(a zero leak could equally mean 'fixed' or 'the agent stopped working' — standing rule 3) +agent (Go-http-client) requests since 09:16Z, per box: + 4 10.77.0.2 /api2/json/admin/datastore/felhom-offsite/snapshots + 4 10.77.0.2 /api2/json/version" + 4 10.77.0.3 /api2/json/admin/datastore/felhom-offsite/snapshots + 4 10.77.0.3 /api2/json/version" + +--- and the pvestatd/backup-client streams, to show BOTH boxes are otherwise identical --- + 8 10.77.0.2 Go-http-client/1.1 + 924 10.77.0.2 libwww-perl/6.78 + 898 10.77.0.2 proxmox-backup-client/1.0 + 8 10.77.0.3 Go-http-client/1.1 + 926 10.77.0.3 libwww-perl/6.78 + 898 10.77.0.3 proxmox-backup-client/1.0 diff --git a/documentation/audits/evidence-agent-transport-leak-2026-08-20/p4.3-fleet-sanity.txt b/documentation/audits/evidence-agent-transport-leak-2026-08-20/p4.3-fleet-sanity.txt new file mode 100644 index 00000000..f871a25e --- /dev/null +++ b/documentation/audits/evidence-agent-transport-leak-2026-08-20/p4.3-fleet-sanity.txt @@ -0,0 +1,20 @@ +=== 4.3 fleet sanity === +captured 2026-08-20T09:17:31+00:00 + +--- hub: agent version reported per host --- +2026/08/20 10:52:35 [INFO] host-report from demo-hp-bb76ea (1 guests, 5 storage targets, 2 backups, 2 restore-tests, 2 pbs-snapshots, 14945 bytes) +2026/08/20 11:00:33 [INFO] host-report from demo-felhom-8363b5 (1 guests, 4 storage targets, 2 backups, 2 restore-tests, 2 pbs-snapshots, 14341 bytes) +2026/08/20 11:07:35 [INFO] host-report from demo-hp-bb76ea (1 guests, 5 storage targets, 2 backups, 2 restore-tests, 2 pbs-snapshots, 14960 bytes) +2026/08/20 11:15:33 [INFO] host-report from demo-felhom-8363b5 (1 guests, 4 storage targets, 2 backups, 2 restore-tests, 2 pbs-snapshots, 14323 bytes) +2026/08/20 11:15:50 [INFO] host-report from demo-hp-bb76ea (1 guests, 5 storage targets, 0 backups, 2 restore-tests, 2 pbs-snapshots, 13784 bytes) + +--- hub: any floor-held line for demo-hp? --- +(empty = none) + +--- hub: any unreachable/recovered event since the deploy? --- +(empty = none) + +--- hub: PBS-DR gauge still refreshing (positive observable) --- +2026/08/20 10:38:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB) +2026/08/20 10:54:30 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB) +2026/08/20 11:10:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB) diff --git a/documentation/audits/evidence-agent-transport-leak-2026-08-20/p4.3-hub-hosts.txt b/documentation/audits/evidence-agent-transport-leak-2026-08-20/p4.3-hub-hosts.txt new file mode 100644 index 00000000..777115d9 --- /dev/null +++ b/documentation/audits/evidence-agent-transport-leak-2026-08-20/p4.3-hub-hosts.txt @@ -0,0 +1,11 @@ +=== 4.3 hub /hosts — agent version as the FLEET sees it === +captured 2026-08-20T09:17:45+00:00 +0.106.0 +demo-felhom-8363b5'" style="cursor: pointer;"> +demo-felhom-8363b5">demo-felhom-8363b5 +0.129.0 +demo-hp-bb76ea'" style="cursor: pointer;"> +demo-hp-bb76ea">demo-hp-bb76ea +0.130.0 +0.129.0 +0.106.0 diff --git a/documentation/audits/evidence-agent-transport-leak-2026-08-20/p5-second-box.txt b/documentation/audits/evidence-agent-transport-leak-2026-08-20/p5-second-box.txt new file mode 100644 index 00000000..1b8bf9a8 --- /dev/null +++ b/documentation/audits/evidence-agent-transport-leak-2026-08-20/p5-second-box.txt @@ -0,0 +1,21 @@ +=== PART 5 — the second box. demo-felhom (the former control) gets 0.130.0. === +DooPlex UTC: 2026-08-20T10:18:54+00:00 + +--- T-0 ep0 immediately before --- +t=10:18:55Z pid=551655 fd=220 estab=203 ctrl(.2)=203 fix(.3)=0 CLOSE-WAIT=0 + +old MainPID: 2596329 +RESTART AT : 10:18:56.139906374Z +new MainPID: 2540886 +felhom-agent 0.130.0 + +--- T+~2s --- +t=10:18:56Z pid=551655 fd=17 estab=0 ctrl(.2)=0 fix(.3)=0 CLOSE-WAIT=0 +--- T+~25s --- +t=10:19:17Z pid=551655 fd=19 estab=2 ctrl(.2)=2 fix(.3)=0 CLOSE-WAIT=0 +--- T+~65s --- +t=10:19:58Z pid=551655 fd=20 estab=3 ctrl(.2)=3 fix(.3)=0 CLOSE-WAIT=0 + +--- both boxes, client side --- +demo-felhom: agent 0.130.0, ESTAB to 8007 = 2, service active +felhom-host: agent 0.130.0, ESTAB to 8007 = 0, service active diff --git a/documentation/audits/evidence-agent-transport-leak-2026-08-20/p5-settle.txt b/documentation/audits/evidence-agent-transport-leak-2026-08-20/p5-settle.txt new file mode 100644 index 00000000..b496d57c --- /dev/null +++ b/documentation/audits/evidence-agent-transport-leak-2026-08-20/p5-settle.txt @@ -0,0 +1,11 @@ +=== ep0 settle check — both boxes on 0.130.0 since 10:18:56Z === +t=10:30:27Z pid=551655 fd=17 estab=0 ctrl(.2)=0 fix(.3)=0 CLOSE-WAIT=0 +t=10:32:08Z pid=551655 fd=19 estab=2 ctrl(.2)=1 fix(.3)=1 CLOSE-WAIT=0 +t=10:33:49Z pid=551655 fd=19 estab=1 ctrl(.2)=0 fix(.3)=1 CLOSE-WAIT=0 +t=10:35:29Z pid=551655 fd=19 estab=2 ctrl(.2)=2 fix(.3)=0 CLOSE-WAIT=0 + +--- agent poll activity in the settle period (proves both are still working) --- + 1 10.77.0.2 /api2/json/admin/datastore/felhom-offsite/snapshots + 1 10.77.0.2 /api2/json/version" + 1 10.77.0.3 /api2/json/admin/datastore/felhom-offsite/snapshots + 1 10.77.0.3 /api2/json/version" diff --git a/documentation/audits/evidence-agent-transport-leak-2026-08-20/stop1-ep0-before.txt b/documentation/audits/evidence-agent-transport-leak-2026-08-20/stop1-ep0-before.txt new file mode 100644 index 00000000..45a42520 --- /dev/null +++ b/documentation/audits/evidence-agent-transport-leak-2026-08-20/stop1-ep0-before.txt @@ -0,0 +1,14 @@ +=== ep0 state at STOP 1 — read-only, immediately before the deploy decision === +UTC: 2026-08-20T09:10:00+00:00 epoch=1787217000 +PID=551655 (must still be 551655) +Tue Aug 18 09:51:04 2026 +fd count: 414 +--- ESTAB / CLOSE-WAIT split (all states, sport 8007) --- + 397 ESTAB + 1 LISTEN +--- per-peer established --- + 198 10.77.0.2 + 199 10.77.0.3 +--- listener --- +State Recv-Q Send-Q Local Address:Port Peer Address:Port +LISTEN 0 1024 *:8007 *:* diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 1124a3e9..3a39e7de 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -636,16 +636,18 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-333** | **Two disk-health questions the deploy raised and did NOT act on.** **(a) The 55/60 °C bands are SPINNING-DISK bands applied to NVMe.** They were adopted unchanged from the operator's Prometheus config so the two systems cannot disagree — a deliberate, stated decision — but **measured on demo-hp 2026-08-14 the healthy Toshiba KXG50PNV1T02 NVMe idles at 53 °C, two degrees below Figyelmeztetés and seven below Hiba**, and NVMe routinely exceeds 60 °C under load with no fault whatever. As it stands a healthy customer NVMe under sustained write can be reported as **Hiba** — the single worst outcome this feature can produce. **(b) The agent runs bare `smartctl -a -j` with no `-n standby`** (`felhom-agent/internal/storage/hostops.go:368`), so every poll WAKES a spun-down drive; going 6h → hourly multiplies that by six. demo-hp is all-flash so the cadence measurement could not reveal it, and it was recorded rather than acted on per the task's own instruction. Mitigating datum from the fixture: the failing drive logged only **3375 load cycles in 60505 hours** (~one per 18h), i.e. that duty cycle barely spins down at all | **READY (S each) — NEW 2026-08-14** | — | (a) split the temperature bands by device class, or drop them for NVMe and rely on `critical_warning`; (b) add `-n standby` to the agent's smartctl invocation (an agent change, so fold it into R-330's session) | Viktor decides (a); CC does (b) | | **R-334** | **CLOSED 2026-08-18 — golden 0.216.0 baked, published and VOUCHED; CI green by run id.** ~~WAIVER + open item: controller v0.215.0 is released and deployed, and NO golden carries it.~~ Convicted by `golden_currency_gate.py` on the 2026-08-14 push: newest released controller **0.215.0**, newest golden bake **0.214.0** (`documentation/tests/golden-0.214.0-2026-08-12`). **A machine installed right now receives 0.214.0** — i.e. a brand-new box would ship WITHOUT the R-328 severity fix and would keep emailing nobody about a failing disk. The running fleet is unaffected (demo-hp guest 9201 is on 0.215.0 and healthy); this is purely the day-0 install path. **Not baked in this session deliberately:** the task scoped deployment to demo-hp only, and the second half of the fix — vouching the bake in the hub's day-0 artifact manifest — is **operator-password-gated, so CC cannot complete it**; a baked-but-unvouched golden is worse than none. **The push was made with `git push --no-verify` and it is stated here and in the session report**, per `.claude/rules/gates.md` — the gate has no waiver parser, so recording a waiver does not clear it. **STILL OPEN and now one version WIDER, 2026-08-18:** the newest released controller is **0.216.0** (v0.215.0's own follow-up fix, R-335) and the newest bake is still **0.214.0**, so a new install now misses *two* releases. Re-convicted on this date's documentation-only push, which was likewise made with `--no-verify`; the gate reads `felhom-controller/CHANGELOG.md` and `documentation/tests/golden-*`, **neither of which that session touched** — the conviction is inherited, not caused | **CLOSED 2026-08-18.** Baked from `RUNBOOK-manual-build.md` §4.0+§4.1 in the DooPlex drill VM and published: **`GOLDEN_VERSION=0.216.0`**, **`GOLDEN_SHA256=ac004dc90d8cefccc5448377892f9cff3a4c3e1e27d0e11129120e38ac31c34b`**, 656,970,239 bytes at `…/generic/felhom-golden/0.216.0/golden.tar.zst`. **The published bytes were verified, not just the script's print** — the artifact was downloaded back out of Gitea and hashed, and it matches. **Vouched by the operator, all THREE fields together**, confirmed by reading the hub's own store rather than the save: `artifact_golden_version=0.216.0`, `artifact_agent_version=0.129.0`, `artifact_min_agent=0.129.0` (2026-08-18 11:00:59–11:01:00), and the hub's recorded sha256 matches the downloaded artifact. The R-216 shape was checked on the machine: `MinAgent` 0.129.0 is **equal to**, not above, the newest **published** agent. **`golden_currency_gate.py` rc=0 and `repo_gates.py --fast` rc=0 — all nine gates — and CI is GREEN BY RUN ID: run **353**, `head_sha 7d81681d6`, conclusion `success`** (the two prior runs 351/352 on this same afternoon were red on exactly this row, which is the contrast). That push needed **no `--no-verify`** — the first of the day that did not. Evidence: `documentation/tests/golden-0.216.0-2026-08-18/`, report `REPORT-golden-0.216.0.md`. **Closed with the run id quoted deliberately**: this row was re-confirmed once and widened once, and closing it on a local green a third time would have left the same ambiguity | — | Bake a golden on **0.216.0** per `runbooks/RUNBOOK-manual-build.md` §4.1, then vouch it — a THREE-field change (`golden_version` + `agent_version` + `min_agent`). Until then every NEW install lacks the severity fix | CC bakes; **Viktor vouches** | | **R-335** | **One physical disk was walked TWICE per run, and the second walk sustained it against itself.** Found on live hardware ~2h after the v0.215.0 deploy, **by noticing the release's own positive observable disagreed with its own persisted artefact**: the hourly check logged *"3 disk(s) evaluated"* while `disk-health-state.json` held **two** records. Cause: demo-hp's `c11-scratch` and `felhom-backup` are the same physical NVMe (`/dev/nvme0n1`) and resolve to the same `diskKey`. **Not cosmetic** — `RunDiskHealthCheck` writes a disk's new record before the next entry reads it, so the SECOND copy consumed the FIRST copy's write as its prior: the disk **sustained against itself and reached Hiba on a FIRST sighting**, defeating truth-table row 6 — the exact rule separating a one-hour benign excursion from a false critical — and would have emitted **two identical events** for one drive. **Latent, not active, on demo-hp** (all three entries healthy, zero counters), but any aliased disk developing a single pending sector would have gone straight to Hiba. **This is the shape standing rule 3 warns about: an absent alarm was not evidence — the two artefacts had to be read AGAINST each other** | **CLOSED — controller v0.216.0, 2026-08-14.** Each `diskKey` is evaluated once per run; both entries stay marked `seen` so neither looks like a disappeared disk, and the card still renders both storage rows (the dedup is about state and alerts, not display). Pinned by `TestDiskCheck_SameDiskTwiceIsEvaluatedOnce`; companion red-proof run and reverted — deleting the guard makes the first sighting emit `Kind:2` (Hiba-from-sectors) at 8 sectors | — | — | CC | -| **R-336** | **The offsite DR endpoint is polled about once per second, and that is what turned a slow leak into an outage.** ep0's PBS proxy served **~85,000 requests/day** — a flat **3,538/hour**, every hour, from two boxes: `74,445 GET /api2/json/admin/datastore` (`libwww-perl`, i.e. PVE's `pvestatd`) and `73,171 GET /admin/datastore/felhom-offsite/status` (`proxmox-backup-client`). Two pollers asking substantially the same question at the same rate. On 2026-08-18 this walked a connection leak in the proxy to its 1024-fd soft limit in **14 days**, wedging the offsite tier for 9½ hours (`audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md`). The `LimitNOFILE=65536` drop-in applied that morning raises the ceiling **but does not fix the leak** — it converts a fortnightly outage into a multi-year one, which is mitigation, not a fix. A DR endpoint that is written to weekly does not need to be asked about every second | **READY (M) — NEW 2026-08-18** | — | **CORRECTED 2026-08-18 (evening) — the easy lever named here does not exist.** This cell used to read *"PVE storage status is the prime suspect, and its interval is tunable"*. **The first half is right and the second half is false.** `pvestatd` stats EVERY configured storage on each 10-second cycle, and Proxmox staff have stated the interval is not designed to be configurable — so there is no knob to turn down. The only lever PVE actually offers is disabling the storage entry (`pvesm set --disable 1`) around the backup window, and that is **substantially more than a tuning knob**: it collides with `felhom-agent/internal/pbsdr/manager.go`'s health model, where an inactive-but-existing entry drives the consume-the-one-time-secret recovery path. So the fix is a design question (does the hub still need a 15-minute fill reading at all, given R-339 now reports reachability separately?), not a config edit. **Doc-only correction — no agent code was changed.** The remaining step is unchanged: cut the poll rate by whatever means survives that question, then confirm the fd count between restarts stops climbing — the positive observable, per standing rule 3. **Baseline measured 2026-08-18, and the FIRST measurement published was WRONG.** The initial "~85/day, matching the ~73/day implied by the failure" came from a single 17-minute window whose delta was **one descriptor** — a sample of one cannot carry a daily rate, and the agreement that made it feel solid was coincidence. **Re-measured over two independent windows the same morning: 183/day (31 min) and 200/day (5.6 h)** — ~2.6x the published figure, putting the runway to the 65536 ceiling at **~357 days, not the ~2 years first claimed**. **And the named mechanism is the minority one:** across that window `CLOSE-WAIT` held flat at 1 while `ESTAB` grew 45→49 — *all* the growth was established connections, and at the wedge the split was 1011 ESTAB / 543 CLOSE-WAIT. **The fix must target connections the proxy never reaps, not just `CLOSE-WAIT` sockets.** The PBS 4.2.5-1 upgrade (2026-08-18) did NOT change the slope and was never expected to — see R-341 **SPIKE 2026-08-20 — THE PREMISE OF THIS ROW DOES NOT SURVIVE MEASUREMENT, and that is a change in what the row IS, not new evidence on it.** `audits/SPIKE-ep0-established-connections-2026-08-20.md`. **The leak is OURS, and the poll rate is not what feeds it.** Every one of the 388 leaked descriptors is an ESTABLISHED connection held open by **`felhom-agent`** on the boxes — 194 on each, `ss -tnp` naming a single PID per box, and **zero** held by `pvestatd` or `proxmox-backup-client`. Confirmed independently from ep0's access log over the same 46.18 h window: `libwww-perl` (pvestatd) **81,192 requests -> 0 descriptors**, `proxmox-backup-client` **80,061 requests -> 0 descriptors**, `Go-http-client/1.1` (the agent) **387 `/snapshots` calls -> 388 sockets — one per call, within one**. So **162,404 requests, 99.5% of the traffic, produce 0% of the leak.** **Mechanism, named from source:** `felhom-agent/internal/pbs/client.go:56-60` builds `&http.Transport{TLSClientConfig: tlsCfg}` — a composite literal, so `IdleConnTimeout` is the zero value = **no limit** (`http.DefaultTransport` sets 90 s; a literal does not inherit it) — and `cmd/felhom-agent/main.go:1486` (`pbsTargetsFromPVE`) builds **a fresh client every cycle**, as its own doc comment states. Each cycle therefore strands one idle keep-alive connection in a transport nothing ever closes; `CloseIdleConnections`/`IdleConnTimeout`/`MaxIdleConns` appear **nowhere** in the agent repo. Cadences reconcile without fitting: 900 s hub poll (184.7 cycles) + 6 h `DefaultVerifyCadence` (7.7 cycles) = 192.4 predicted vs **194 observed per box**. **CONSEQUENCE — RE-RANK.** The remaining step recorded above ("cut the poll rate, then confirm the fd count stops climbing") **would have produced a null result and read as a failed fix.** Cutting the Proxmox poll rate removes ~99.5% of ep0's request load and **zero** descriptors. The poll rate is still wrong on its own terms — 85,000 requests/day to a weekly-write DR endpoint — but it is now a **scaling/cost item, not the leak fix**, and the leak fix is **R-344**. **Q3 (is the leak proportional to the request rate?) is PREDICTED not-proportional and NOT YET MEASURED** — Phase C is held at STOP 1 with its prediction pre-registered in `evidence-ep0-established-connections-2026-08-20/phaseC-prediction.txt`. Do not record a proportionality verdict here until that window has run. | CC | +| **R-336** | **The offsite DR endpoint is polled about once per second, and that is what turned a slow leak into an outage.** ep0's PBS proxy served **~85,000 requests/day** — a flat **3,538/hour**, every hour, from two boxes: `74,445 GET /api2/json/admin/datastore` (`libwww-perl`, i.e. PVE's `pvestatd`) and `73,171 GET /admin/datastore/felhom-offsite/status` (`proxmox-backup-client`). Two pollers asking substantially the same question at the same rate. On 2026-08-18 this walked a connection leak in the proxy to its 1024-fd soft limit in **14 days**, wedging the offsite tier for 9½ hours (`audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md`). The `LimitNOFILE=65536` drop-in applied that morning raises the ceiling **but does not fix the leak** — it converts a fortnightly outage into a multi-year one, which is mitigation, not a fix. A DR endpoint that is written to weekly does not need to be asked about every second | **READY (M) — NEW 2026-08-18** | — | **CORRECTED 2026-08-18 (evening) — the easy lever named here does not exist.** This cell used to read *"PVE storage status is the prime suspect, and its interval is tunable"*. **The first half is right and the second half is false.** `pvestatd` stats EVERY configured storage on each 10-second cycle, and Proxmox staff have stated the interval is not designed to be configurable — so there is no knob to turn down. The only lever PVE actually offers is disabling the storage entry (`pvesm set --disable 1`) around the backup window, and that is **substantially more than a tuning knob**: it collides with `felhom-agent/internal/pbsdr/manager.go`'s health model, where an inactive-but-existing entry drives the consume-the-one-time-secret recovery path. So the fix is a design question (does the hub still need a 15-minute fill reading at all, given R-339 now reports reachability separately?), not a config edit. **Doc-only correction — no agent code was changed.** The remaining step is unchanged: cut the poll rate by whatever means survives that question, then confirm the fd count between restarts stops climbing — the positive observable, per standing rule 3. **Baseline measured 2026-08-18, and the FIRST measurement published was WRONG.** The initial "~85/day, matching the ~73/day implied by the failure" came from a single 17-minute window whose delta was **one descriptor** — a sample of one cannot carry a daily rate, and the agreement that made it feel solid was coincidence. **Re-measured over two independent windows the same morning: 183/day (31 min) and 200/day (5.6 h)** — ~2.6x the published figure, putting the runway to the 65536 ceiling at **~357 days, not the ~2 years first claimed**. **And the named mechanism is the minority one:** across that window `CLOSE-WAIT` held flat at 1 while `ESTAB` grew 45→49 — *all* the growth was established connections, and at the wedge the split was 1011 ESTAB / 543 CLOSE-WAIT. **The fix must target connections the proxy never reaps, not just `CLOSE-WAIT` sockets.** The PBS 4.2.5-1 upgrade (2026-08-18) did NOT change the slope and was never expected to — see R-341 **SPIKE 2026-08-20 — THE PREMISE OF THIS ROW DOES NOT SURVIVE MEASUREMENT, and that is a change in what the row IS, not new evidence on it.** `audits/SPIKE-ep0-established-connections-2026-08-20.md`. **The leak is OURS, and the poll rate is not what feeds it.** Every one of the 388 leaked descriptors is an ESTABLISHED connection held open by **`felhom-agent`** on the boxes — 194 on each, `ss -tnp` naming a single PID per box, and **zero** held by `pvestatd` or `proxmox-backup-client`. Confirmed independently from ep0's access log over the same 46.18 h window: `libwww-perl` (pvestatd) **81,192 requests -> 0 descriptors**, `proxmox-backup-client` **80,061 requests -> 0 descriptors**, `Go-http-client/1.1` (the agent) **387 `/snapshots` calls -> 388 sockets — one per call, within one**. So **162,404 requests, 99.5% of the traffic, produce 0% of the leak.** **Mechanism, named from source:** `felhom-agent/internal/pbs/client.go:56-60` builds `&http.Transport{TLSClientConfig: tlsCfg}` — a composite literal, so `IdleConnTimeout` is the zero value = **no limit** (`http.DefaultTransport` sets 90 s; a literal does not inherit it) — and `cmd/felhom-agent/main.go:1486` (`pbsTargetsFromPVE`) builds **a fresh client every cycle**, as its own doc comment states. Each cycle therefore strands one idle keep-alive connection in a transport nothing ever closes; `CloseIdleConnections`/`IdleConnTimeout`/`MaxIdleConns` appear **nowhere** in the agent repo. Cadences reconcile without fitting: 900 s hub poll (184.7 cycles) + 6 h `DefaultVerifyCadence` (7.7 cycles) = 192.4 predicted vs **194 observed per box**. **CONSEQUENCE — RE-RANK.** The remaining step recorded above ("cut the poll rate, then confirm the fd count stops climbing") **would have produced a null result and read as a failed fix.** Cutting the Proxmox poll rate removes ~99.5% of ep0's request load and **zero** descriptors. The poll rate is still wrong on its own terms — 85,000 requests/day to a weekly-write DR endpoint — but it is now a **scaling/cost item, not the leak fix**, and the leak fix is **R-344**. **Q3 (is the leak proportional to the request rate?) is PREDICTED not-proportional and NOT YET MEASURED** — Phase C is held at STOP 1 with its prediction pre-registered in `evidence-ep0-established-connections-2026-08-20/phaseC-prediction.txt`. Do not record a proportionality verdict here until that window has run. **RE-SCOPED 2026-08-20 — THIS ROW IS NO LONGER A LEAK FIX, AND ITS RECORDED NEXT-STEP WOULD HAVE "FIXED" NOTHING WHILE LOOKING LIKE A FAILED FIX.** That near-miss is the reason the spike-first rule exists and it is kept here deliberately. The old next-step read: *cut the poll rate by whatever means survives that question, then confirm the fd count between restarts stops climbing.* Had it been executed, the fd count would have kept climbing at the same ~200/day, the poll reduction would have been recorded as ineffective, and the real defect — **ours, in `felhom-agent`, R-344** — would have been further from being found, not closer. **Measured 2026-08-20:** `pvestatd` (`libwww-perl`) and `proxmox-backup-client` made **162,404 requests** in a 46 h window and leaked **zero** descriptors; the agent made 811 and leaked **388**. The fix (agent 0.130.0) took ep0 from 388 accumulated descriptors to its **baseline of 17**, with the poll rate completely unchanged — 85,000/day before and after. **WHAT THIS ROW ACTUALLY IS NOW — a SCALING concern, still worth fixing on its own merits:** ~85,000 requests/day to a DR endpoint that is WRITTEN TO WEEKLY, from two boxes. That is ~42,500/box/day, so **at fifty customers it is ~2.1 million requests/day — about 25 requests/second, constantly, against a CX33**. The design question is unchanged and is still the hard part: does the hub still need a 15-minute fill reading at all, given R-339 reports reachability separately? And the lever remains awkward — `pvestatd` stats every configured storage on each 10-second cycle with no tunable interval, so the only PVE-side lever is disabling the storage entry, which collides with `felhom-agent/internal/pbsdr/manager.go`'s health model. **NEW ACCEPTANCE CRITERION, since the old one is void:** the fd count is NOT the observable for this row any more — that belongs to R-344 and is already satisfied. Measure the REQUEST RATE at ep0's access log, and state the projected rate at the target customer count. | CC | | **R-337** | **`/backup/status` lagged a completed backup by minutes on one box and not the other — and it RESOLVED ITSELF, which is why this is WATCHING and not a defect.** During the R-336 recovery on 2026-08-18, `demo-hp`'s snapshot landed on ep0 at **03:58:43Z** (complete manifest; the host's own task index says `OK`) — yet `GET /backup/status` was **still serving the superseded 03:27:00Z failure at ~04:03Z**, four-plus minutes later. `demo-felhom` showed its new result within ~40 s of completion. **The lag cleared on its own:** demo-hp's 04:07:35Z host report carries `felhom-pbs success=true, 4.29 GB`, and the hub is green for both boxes. **The first draft of this row claimed the success was "still reported as failed" — that was written before the next report arrived and it was wrong; the corrected claim is a several-minute skew between the two boxes, not a stuck value.** It is recorded because a status field that can trail its own artifact by minutes will, during an incident, be read as a second failure — this session nearly did — and because the asymmetry between the two boxes is unexplained | **WATCHING — NEW 2026-08-18** | another observation, ideally during an incident rather than constructed | **Do not open a fix on this as written.** First establish the intended refresh path for `/backup/status` after an out-of-schedule run; only if the skew is not simply collection cadence is there anything to pin. If it is cadence, close this row and say so | CC | | **R-338** | **`demo-hp` is not on the R-50 island at all, and `operations/nodes.md` states that it is.** The page records both fleet boxes as island-migrated 2026-07-25. True of `felhom-pve`; **false of `demo-hp`**, whose `agent.json` has `listen_addr: 192.168.0.87:8443` — the customer LAN address — and **no `island_bridge`/`island_guest_addr` keys at all**, whose guest 9201 has `net0` only (no `eth1`), and whose `vmbr9` exists with **zero members**. The controller's `controller.yaml` points at the LAN address, so the box works; this is inventory drift, not breakage. **Two costs.** A session trusting the page addresses the wrong endpoint — that happened on 2026-08-18 and the resulting timeout was briefly read as a fault. And the agent's local API is **bound to the customer LAN on this box** rather than to a point-to-point island, which is the exposure R-50 was built to remove — so a documented security property is claimed for a box that does not have it | **READY (S) — NEW 2026-08-18** | — | Decide which is true: migrate `demo-hp` to the island, or correct `nodes.md`. Leaving both is the one option that keeps the doc lying | Viktor decides; CC executes | | **R-341** | **Does the fd slope change after the PBS 4.2.5 upgrade? — two dated checks, and the answer is expected to be NO.** ep0 was upgraded 4.2.2-1 → 4.2.5-1 on 2026-08-18 09:51Z on the operator's ruling, **for rehearsal value, not as a fix**: the full changelog range was read (128 lines, all three entries) and swept for connection-handling vocabulary, and it contains **no mechanism** by which descriptor reaping would change — the single keyword hit was `S3 … honor the node's proxy settings`, HTTP-proxy config for S3, not the PBS proxy daemon. **The 32-minute post-upgrade window is indistinguishable from the before window** (+5 fd/1919 s = 225/day vs +4 fd/1885 s = 183/day; the two differ by ONE descriptor and Poisson uncertainty on such counts is ±2, so both are consistent with one unchanged rate — the higher after-figure is noise, not a regression). **Thirty minutes cannot settle it in either direction and this row exists so nobody pretends it did.** New `t0` = **fd 17 at 2026-08-18 09:51:22Z, proxy PID 551655**; before-rate to beat = **183–200/day**. **Interpretation fixed in advance** (`evidence-ep0-pbs-upgrade-2026-08-18/stop1-ruling.txt`, written before any numbers existed): unchanged = EXPECTED, not a failed upgrade; changed = a SURPRISE needing explanation, not a confirmation | **WATCHING — FIRST CHECK TAKEN 2026-08-20: the rate is UNCHANGED, exactly as pre-registered. Second check still due 2026-08-25** | elapsed time only | **Two dated checks, both CC:** **+24 h — 2026-08-19 ~10:00Z** and **+7 d — 2026-08-25 ~10:00Z**. Command (the incident's own positive observable): `ssh root@ 'PID=$(systemctl show proxmox-backup-proxy -p MainPID --value); ls /proc/$PID/fd \| wc -l; ss -lnt "( sport = :8007 )"; ss -tn state all "( sport = :8007 )" \| awk "NR>1{print \$1}" \| sort \| uniq -c'`. **Record the ESTAB/CLOSE-WAIT split, not just the total** — the split is what says which leak it is. If PID ≠ 551655 the window is void: something restarted the proxy and the count began again **FIRST CHECK — TAKEN 2026-08-20 08:02:13Z, and it was taken at +46.2 h, NOT +24 h.** No session ran between 18 and 20 August, so the 2026-08-19 date passed as pure elapsed time; the delay is stated here rather than backfilled. **The longer window is a BETTER measurement, not a degraded one** — the leaked count is **388** against the 4 and 5 descriptors the original answer rested on, roughly **97x**, so the uncertainty falls from about +/-50% to about +/-5%. **Precondition PASSED:** PID still **551655**, `ps -o lstart` 2026-08-18 09:51:04, `NRestarts=0` on both units. **Result:** fd **17 -> 405** over **166,251 s** = **201.6 fd/day**, Poisson +/-10.2/day (1s), 2s band **181.2-222.1**. Pre-registered range was **370-450** (confirm-band 330-490); **observed 388**, near the centre. **VERDICT: unchanged — the EXPECTED result, and not a failed upgrade.** **Composition:** ESTAB 0 -> 388, **CLOSE-WAIT 0 — absent from the histogram entirely**, so 100% of the growth is established connections and `CLOSE-WAIT` is not merely the minority half. Runway from fd 405 at 201.6/day to the 65536 ceiling: **~323 days (~2027-07-09)**. Full working: `audits/SPIKE-ep0-established-connections-2026-08-20.md` + `evidence-ep0-established-connections-2026-08-20/step1-slope-computation.txt`. **MEASUREMENT TRAP for the 2026-08-25 check, found on this run — see R-346:** the anchor must be `ps -o lstart= -p $MainPID`, NOT `systemctl show -p ActiveEnterTimestamp`, which reads 03:54:54Z for this generation (the upgrade re-exec'd the proxy; systemd never saw a stop, `NRestarts` is still 0) and would put the rate ~15% low. **PERTURBATION NOTE for the 2026-08-25 check:** this spike's Phase C quietens `pvestatd` on demo-hp for one overnight window inside that 7-day interval. Quantified so nobody reads the shortfall as a change in the leak: even if the leak were fully proportional to the request rate (which Parts 1-2 predict it is NOT), a 10 h half-rate window inside 168 h shifts the 7-day slope by **~3%**, far inside the +/-10% Poisson band on a ~1,400-descriptor count. **The +7 d reading remains usable.** | CC on both dates | | **R-342** | **The ep0 snapshot covers less than it looks like it covers, and the next person will assume otherwise.** Quoting `audits/evidence-ep0-pbs-upgrade-2026-08-18/stop2-snapshot.txt` verbatim: *"covers — the 38 GB system disk /dev/sda (root), i.e. the PBS packages, unit files, /etc/systemd drop-ins, nftables and wg config. DOES NOT — /mnt/pbs-datastore. That is /dev/sdb, a separate 100 GB VOLUME, and Hetzner server snapshots do not include attached volumes. The backup data is therefore NOT protected by this snapshot."* Snapshot **421440873** (`felhom-hetzner-20260818`, 15.06 GB, Available) was taken as the rollback for the 4.2.2→4.2.5 PBS upgrade. **Rolling it back restores software state, not the datastore.** That was *acceptable for that change* — a package install writes no datastore content — and the file says so. **The problem is what happens next:** this fact lives in an evidence file nobody will open again, and a snapshot named as "the rollback" reads as protecting everything on the box. ep0 holds the only off-premises copy of a real customer's data | **READY (S) — NEW 2026-08-18** | — | **Decide the safeguard for any future ep0 procedure that could touch `/mnt/pbs-datastore` — it does not exist and has not been designed.** Candidates: a Hetzner **Volume** snapshot (a different object from the server snapshot), a PBS-level sync to a second location, or an explicit written acceptance that the datastore is unprotected for the duration. **Nothing may be added to a runbook implying a safeguard exists until one does** | **Viktor decides; CC executes** — a risk-to-customer-data question | | **R-343** | **The managed controller floor was raised 0.214.0 → 0.216.0 — and it was NOT the no-op it was expected to be: it moved a live box nine seconds later.** Raised by the operator 2026-08-18 **12:36:58Z**, in a **separate save after** the artifact vouch. **Read back from the store, not the form** (`hub_settings.min_controller_version`, `GetGlobalMinControllerVersion` — `hub/internal/store/store.go:1751`): `min_controller_version = 0.216.0`, updated 12:36:58. **WHY IT HAD BEEN BEHIND — this was NOT drift, and describing it as "two releases behind" without this context reads as a defect it was not.** `publish-train-rules.md` rule 1 is *manifest before floor*, and rule 2 requires the floor field to be filled **LAST, in a separate save**, because the DB row overrides the env floor and **acts immediately on the next report cycle**. That rule was earned: on the 2026-07-11 publish train the floor was saved together with the manifest, acted at once, and pushed controller 0.113.0 onto Peti's box **~9 minutes ahead of agent 0.81.0** — the exact forbidden skew, benign only because that box had no NAS shares. R-120's row records the same deliberate choice (*"Floor untouched per publish-train rule 2"*). So the floor sitting at 0.214.0 was **policy being followed**, not neglect. **What it was functionally while it sat there:** not a live problem — every reporting box was at or above it — but a **safety net set two versions low**. The floor is what drags a box forward if it ever falls behind (restored from an old backup, reinstalled, long offline), and at 0.214.0 it would have pulled such a box only to two versions back, missing R-328's severity fix and R-335's follow-up. **THE MEASURED BLAST RADIUS — five reads, and the third and fifth are the findings.** **(1) Floor:** `0.216.0` @ 12:36:58Z, from the store. **(2) Per-customer overrides:** **zero** — all five `customer_configs` rows carry an empty `min_controller_version`, so nothing hides behind a lower override and the global applies to everyone. **(3) Every box's controller version:** `demo-felhom` **0.216.0**, `demo-hp` **0.216.0** — both AT the floor **now**; `drill-r50` **0.213.0** (status `blocked`, last report 2026-08-12, powered off/reverted) and `peti-felhom` **0.115.0** (host row DELETED) are **below** it but are **not reporting boxes**. **(4) Directives/holds:** no `managed floor HELD` line exists; the hub logged `[INFO] Global controller-version floor set to "0.216.0"`. **(5) THE FINDING — a controller DID auto-update after the raise.** `demo-felhom` had been on **0.214.0** since 2026-08-12 16:44 and the hub recorded `controller_updated — Controller frissítve: 0.214.0 → 0.216.0` at **12:37:07Z**, then `controller_started (0.216.0)` at 12:37:12Z — **nine seconds after the raise**, exactly the immediate action rule 2 documents. `demo-hp` was already on 0.216.0 (hand-deployed 2026-08-14 08:31) and did not move. **No error, warning or critical event followed** — the update completed and the controller came back up. **So the change was real, not inert: "every reporting box is at or above the floor" is true BECAUSE of the raise, not independently of it.** **Why it is safe by construction**, cited rather than asserted: `ResolveManagedFloor` (`hub/internal/store/store.go:2068`) sets `Held` and clears the floor entirely when `Floor > GoldenVersion` (the R-216 shape) — floor 0.216.0 **equals** golden 0.216.0, so that guard does not trip — and holds per-box when the box's agent is below the manifest's `MinAgent`, or unknown, or unparseable; both boxes report agent 0.129.0 against `MinAgent` 0.129.0, so the floor was served rather than held. That second guard remains armed for any box reporting with an old agent. **No ISO rebuild is required:** the golden is fetched at first boot from the hub's manifest, which is 0.216.0 — at the floor, not below it — and rule 5's `assert_golden_ge_floor` is a **build-time** gate for *future* builds (`scripts/iso/build-felhom-iso.sh:77`, called at **:267**) which **fails open with a warning** when its inputs are absent (`:78-82`: `if [[ -z "$golden" \| \| -z "$floor" ]]` → `log_warn "… UNENFORCED …"` → `return 0`), both confirmed in the script. **`peti-felhom` was NOT contacted and needs no contact** — from the PETI row: its host row was deleted 2026-07-15 and *"a report from a deleted host 401s and is not persisted"*, so it cannot receive a floor directive at all and the raise cannot reach it | **OPEN — NEW 2026-08-18.** Deliberately NOT closed: the task's closing condition was *all five reads clean, no directive served*, and read 5 shows a live box updated. It went cleanly and is the floor working as designed — but a change recorded as a no-op when it moved a customer box is exactly the kind of record that misleads later | — | **Confirm the 0.216.0 update on `demo-felhom` is healthy in normal operation** (it reported and restarted clean, but it has not yet run a full backup cycle on 0.216.0 at the time of writing), then close. Separately: `drill-r50` at 0.213.0 will be dragged to 0.216.0 by this floor if it is ever booted and reports — that is the floor doing its job, noted so it is not read as a surprise | CC | -| **R-344** | **`felhom-agent` leaks one TCP connection to PBS per poll cycle, forever, on both sides — and it is the whole of the ep0 descriptor leak.** Found by the 2026-08-20 connections spike (`audits/SPIKE-ep0-established-connections-2026-08-20.md`), which was looking for a Proxmox poll-rate problem and found ours instead. **Two defects compounding.** **(a)** `internal/pbs/client.go:56-60` builds `&http.Transport{TLSClientConfig: tlsCfg}` as a composite literal, so **`IdleConnTimeout` is the zero value = no limit** — `http.DefaultTransport` sets 90 s and a literal does not inherit it. **(b)** `cmd/felhom-agent/main.go:1486` (`pbsTargetsFromPVE`) builds **a fresh `pbs.Client` every cycle**, as its own doc comment states, so each cycle strands one idle keep-alive connection in a transport that is then unreachable — and an unreachable `http.Transport` does **not** close its connections; the `persistConn` read-loop goroutine keeps the socket alive. `CloseIdleConnections` / `IdleConnTimeout` / `MaxIdleConns` appear **nowhere in the repo** (grep: no matches). **Measured live, both sides, twice:** ep0 held 388 ESTAB (194 from each box) at 08:02:39Z and 392 at 08:33:42Z; the boxes held 194+194 and 196+196 at the same instants, and the four new sockets carried **the same four source ports** on both sides. **Zero sockets closed in 31 minutes**, and all carried keepalive timers with `retrans=0` — mutually held live idle connections, not half-open ones. One socket per agent `/snapshots` call (387 calls vs 388 sockets). Rate **201.6/day** across the fleet; ep0 reaches its 65536 ceiling in **~323 days**. **It is bilateral and the box side is the under-watched half:** each agent holds **196 of its 208** descriptors in these sockets. Its limit is 524287, so the boxes are in no danger *today* — which is why this stayed invisible, not why it is harmless. | **READY (S) — NEW 2026-08-20** | none — it is a self-contained change in `felhom-agent` | **No fix is proposed here: the spike-first gate forbids it and the spec is a separate task.** What the spec must settle, and none of it is decided: whether the per-cycle client construction is the thing to remove or the transport is the thing to share; what `IdleConnTimeout` should be given the 900 s poll and the 6 h verify cadence; whether `internal/hub/client.go:53` and `internal/proxmox/client.go:69` — **the same composite-literal pattern, built once at start-up so not leaking by this route today** — should be changed in the same pass or left alone with a test pinning why. **The proof obligation is the fd count, not the diff:** per standing rule 3 the positive observable is ep0's ESTAB count going FLAT between proxy restarts, measured over a window long enough to matter — a green test suite proves nothing here, and a 30-minute window proves nothing here either (that error is already recorded twice in R-336 and R-341). | CC | +| **R-344** | **`felhom-agent` leaks one TCP connection to PBS per poll cycle, forever, on both sides — and it is the whole of the ep0 descriptor leak.** Found by the 2026-08-20 connections spike (`audits/SPIKE-ep0-established-connections-2026-08-20.md`), which was looking for a Proxmox poll-rate problem and found ours instead. **Two defects compounding.** **(a)** `internal/pbs/client.go:56-60` builds `&http.Transport{TLSClientConfig: tlsCfg}` as a composite literal, so **`IdleConnTimeout` is the zero value = no limit** — `http.DefaultTransport` sets 90 s and a literal does not inherit it. **(b)** `cmd/felhom-agent/main.go:1486` (`pbsTargetsFromPVE`) builds **a fresh `pbs.Client` every cycle**, as its own doc comment states, so each cycle strands one idle keep-alive connection in a transport that is then unreachable — and an unreachable `http.Transport` does **not** close its connections; the `persistConn` read-loop goroutine keeps the socket alive. `CloseIdleConnections` / `IdleConnTimeout` / `MaxIdleConns` appear **nowhere in the repo** (grep: no matches). **Measured live, both sides, twice:** ep0 held 388 ESTAB (194 from each box) at 08:02:39Z and 392 at 08:33:42Z; the boxes held 194+194 and 196+196 at the same instants, and the four new sockets carried **the same four source ports** on both sides. **Zero sockets closed in 31 minutes**, and all carried keepalive timers with `retrans=0` — mutually held live idle connections, not half-open ones. One socket per agent `/snapshots` call (387 calls vs 388 sockets). Rate **201.6/day** across the fleet; ep0 reaches its 65536 ceiling in **~323 days**. **It is bilateral and the box side is the under-watched half:** each agent holds **196 of its 208** descriptors in these sockets. Its limit is 524287, so the boxes are in no danger *today* — which is why this stayed invisible, not why it is harmless. | **READY (S) — NEW 2026-08-20** | none — it is a self-contained change in `felhom-agent` | **No fix is proposed here: the spike-first gate forbids it and the spec is a separate task.** What the spec must settle, and none of it is decided: whether the per-cycle client construction is the thing to remove or the transport is the thing to share; what `IdleConnTimeout` should be given the 900 s poll and the 6 h verify cadence; whether `internal/hub/client.go:53` and `internal/proxmox/client.go:69` — **the same composite-literal pattern, built once at start-up so not leaking by this route today** — should be changed in the same pass or left alone with a test pinning why. **The proof obligation is the fd count, not the diff:** per standing rule 3 the positive observable is ep0's ESTAB count going FLAT between proxy restarts, measured over a window long enough to matter — a green test suite proves nothing here, and a 30-minute window proves nothing here either (that error is already recorded twice in R-336 and R-341). **FIX SHIPPED TO THE TWO DEMO BOXES AND PROVEN LIVE, 2026-08-20 — agent 0.130.0. THIS ROW STAYS OPEN: see R-347.** `audits/SPIKE-ep0-established-connections-2026-08-20.md` §"Fix and proof" + `evidence-agent-transport-leak-2026-08-20/`. **The fix is one field restored to the standard library's own value.** New leaf package `internal/httpx` owns `DefaultIdleConnTimeout = 90s` (= `http.DefaultTransport`'s value, so there is no invented number to justify) and `NewTransport`, which returns a **fresh** transport per call and treats zero-or-negative as **use the default, never "no timeout"**. All three hand-rolled transports now go through it; `grep '&http.Transport{'` matches only `httpx` itself. `internal/hub/client.go` + `internal/proxmox/client.go` were corrected in the same pass and **neither contributed to the ep0 leak** — both are built once per process and neither talks to ep0:8007. **P1 — the restart, outcome (i) within ONE second:** ep0 fd **415 -> 216**, `demo-hp`'s 199 established connections gone, **CLOSE-WAIT stayed 0**. So **outcome (ii) does NOT exist and gets no row** — ep0 reaps on peer FIN correctly, and the 543 `CLOSE-WAIT` at the 2026-08-18 wedge has another explanation. Ownership thereby proven a THIRD independent way (socket owner, access-log user agent, and now what dies with the process). **P2 — divergence, 1.03 h (the operator closed the >=4 h window early; no daily rate is extrapolated and none is needed):** control `demo-felhom` 199->203 (**+4**), fixed `demo-hp` 0->0 (**+0**). ep0's access log counts the opportunities directly: **each box made exactly 4 `/snapshots` + 4 `/version` calls** — same cadence, same work. **control 4 cycles -> 4 leaks; fixed 4 cycles -> 0 leaks.** **Positive observable per standing rule 3** (a zero leak is equally consistent with "the agent stopped working"): the fixed box's four poll cycles are in ep0's log, and the boxes' other traffic is near-identical (libwww-perl 924 vs 926, proxmox-backup-client 898 vs 898), so **the only difference between them is the binary**. **P3 — second box, 10:18:56Z: ep0 fd 220 -> 17 in under two seconds**, settling at 17-19 and returning to 17 between cycles. **17 is precisely ep0's `t0` baseline** (fd 17, ESTAB 0, 2026-08-18 09:51:22Z). **CORRECTION, made within the hour it was written:** this session's own STOP 1 report and the first CHANGELOG draft said *"does not clear the 388 descriptors already stuck on ep0 — those persist until that proxy restarts."* **That is wrong.** The descriptors were held on BOTH sides; restarting the agents released every one. **ep0 was read-only throughout and its proxy PID never changed (551655)** — the protected machine was never touched and did not need to be. **Tests:** `internal/pbs/client_leak_test.go` counts connections SERVER-side and models the abandonment, so it pins the consequence, not the field — it deliberately does not assert `err == nil` (true of the leaking code). **Two red-proofs, both seen failing:** removing the timeout gives *"after 5s the server still holds 5 open connection(s), want 0"* (the count is in the message, so it cannot be a timeout with another cause); and `DisableKeepAlives: true` — **which the leak test PASSES** — is caught only by `TestPBSClient_KeepAliveStillReuses` (*"3 sequential requests over 3 connection(s), want 1"*). **Scenario A alone would have accepted a fix that made the problem worse.** Both reverted. **Fleet sanity:** hub reports 0.130.0 on both, no `floor held` (0.130.0 > golden MinAgent 0.129.0), no `*_unreachable` event, PBS-DR gauge refreshing throughout. **Why it is NOT closed:** the binary is hand-installed on two demo boxes and unpublished, so a fresh install still ships the leaking agent — **R-347**. A fix living on two boxes by hand is not delivered. | CC | | **R-345** | **`hub/Makefile` tags and pushes `:latest`, which the project's own rules forbid in two places.** Lines 21-22 of `docker-push`: `docker tag $(IMAGE):$(VERSION) $(IMAGE):latest` then `docker push $(IMAGE):latest`. `.claude/rules/hub.md:35` says *"Pin explicit versions, never `:latest`"* and `.claude/rules/manifests.md:15` repeats it. Verified present on `848368ec3`. Small, and the deployed manifests do pin a version, so nothing is currently broken by it — but a documented command that performs the prohibited action is a trap for whoever next reads the Makefile as the reference for how to publish, and a floating `:latest` on the registry is exactly the thing an emergency `kubectl set image` reaches for. Noticed while running the 2026-08-20 connections spike; unfiled until now. | **READY (XS) — NEW 2026-08-20** | — | Delete the two lines, or keep them behind an explicit opt-in target that says in a comment why it exists. Check whether a stale `:latest` tag already sits on `gitea.dooplex.hu/admin/felhom-hub` before deciding — an existing floating tag is the more dangerous half. | CC | | **R-346** | **`ActiveEnterTimestamp` answers a different question than the one a slope measurement asks, and on ep0 right now it is wrong by 5 h 56 m.** Found while taking R-341's first dated check. ep0's `proxmox-backup-proxy` has `MainPID=551655` started **2026-08-18 09:51:04Z** (`ps -o lstart=`), but `systemctl show -p ActiveEnterTimestamp` reads **03:54:54Z** and `NRestarts` reads **0** — because the 4.2.5-1 upgrade **re-exec'd** the daemon rather than restarting the unit, so systemd never observed a stop. Anyone anchoring "when did this proxy generation start" on `ActiveEnterTimestamp` would divide 388 descriptors by 52.1 h instead of 46.2 h and report **178/day instead of 201.6/day — ~12% low** — while every field consulted looks healthy and consistent. **This is the workspace rule's own case, in a new place:** ask of a timestamp *what exactly must have happened for this to be set?* Here the answer is "the unit entered active", which is not "this process started". `NRestarts=0` is the tell, and it reads like reassurance. | **READY (XS) — NEW 2026-08-20** | — | R-341's command already uses `ps -o lstart= -p $MainPID` and is correct; the risk is a future reader "improving" it to a systemd property. Add the reason as a comment beside that command in the R-341 row (done), and check whether any other slope or uptime check in the repo or in `scripts/felhom-tenantsync.sh` anchors on a systemd timestamp where it means a process start. | CC | +| **R-347** | **The R-344 fix exists on two demo boxes by hand and NOWHERE ELSE — a box installed from the current image still ships the leaking agent.** Agent **0.130.0** was built on DooPlex and hand-installed on `demo-hp` and `demo-felhom` on 2026-08-20, deliberately **without** publishing: no Gitea package, no `v0.130.0` tag, no `artifact_agent_version` / `artifact_min_agent` / `artifact_golden_version` change, no staged self-update. **That was correct at the time** — publishing would have pushed the fix onto `demo-felhom` through the self-update path and destroyed the control the whole experiment rested on. The experiment is now finished, so the reason has expired and only the gap remains. **The gap is real but not urgent:** the leak takes ~323 days to reach ep0's 65536 ceiling with two boxes on it, and any restart of the agent clears the whole accumulation (P3, measured: ep0 220 -> 17 fd in under two seconds). A newly installed box therefore leaks slowly and self-heals on every agent deploy. **The CHANGELOG heading is `## UNRELEASED — v0.130.0 candidate` for exactly this reason** — the `release-complete` gate would otherwise convict on a version claiming to be a release that has no tag and no package, and it was right to. | **READY (S) — NEW 2026-08-20** | R-344 (done) | **This is a publish-train decision with an ordering rule attached (`documentation/runbooks/publish-train-rules.md`), so the GO is the operator's; CC executes.** The sequence: flip the CHANGELOG heading to `## v0.130.0` **in the same commit as** the tag, `bash scripts/release-agent.sh 0.130.0`, then the hub's Day-0 artifact manifest must vouch it — **that UI is operator-password-gated and CC cannot drive it**. Confirm `release-complete` goes green afterwards, on the tag and the package, not on the heading. | **Viktor decides**, CC executes | +| **R-348** | **Every agent restart blanks the reported backup list for up to ~18 hours, and the comment that covers it says "unaffected".** Observed 2026-08-20 while deploying R-344: the first host reports after `demo-hp`'s agent restart carry **`0 backups`** (11:15:50 and 11:30:52 CEST, two consecutive), while the box's own `pvesm list` shows archives present on **both** tiers. `internal/backup/store.go`'s `Store` is in-memory and `byTarget` is repopulated only when a backup **runs** — daily for the local tier, weekly for offsite — so the field reads 0 until the next run. `restore_tests` did **not** blank, because that half has a durable on-disk companion (`RestoreTestState`, R-189). **It blinds no alarm, and that was CHECKED rather than assumed.** `hub/internal/monitor/deadline.go` scans back over stored reports with a 7-day `backupEvidenceLookback` whose own comment names this exact case — *"when the LATEST report carries none... and against an agent that stayed restarted for days"* — and `pbs_snapshots` stayed populated at 2 regardless. So this is an observability wart, **not** a safety hole, and it is filed at that severity deliberately. **What is actually wrong is the comment.** The `Store` doc says *"Backups are unaffected — their freshness has a ground truth on the storage (R-84)"*. That is true of the **consequence** and false of the **field**, and it sits three lines below a paragraph explaining that the very same sentence about restore-tests *"used to be here and it is now FALSE"* — so the file already carries one correction of this shape and invites the next reader to trust the surviving half. | **READY (XS) — NEW 2026-08-20** | — | Say what is measured: the field IS lost on restart and repopulates only when a backup runs; the freshness VERDICT is unaffected because the hub looks back 7 days. **Name `backupEvidenceLookback` in the comment** so the cross-repo dependency is visible from the agent side — today the agent's claim of safety rests on a hub constant it does not mention. Per the workspace rule, a comment asserting an invariant needs a test pinning it: the pin belongs on the HUB side, asserting the verdict survives a report carrying `backups: []`. | CC | | **R-339** | **The hub was SILENT when it lost sight of the off-site stores — and a 9 h 37 m outage proved it.** Both box checkers (`OffsiteBoxChecker` over the Hetzner API, `PBSDRBoxChecker` over ep0's `usage` op) held their last snapshot and returned quietly on a failed fetch. That is **correct for a fill signal** — a missing reading must never be mistaken for 0%, which is why degraded data drives no band transition — but the consequence was that a completely dead off-site endpoint and a healthy one were **indistinguishable on the operator channel**. During the 2026-08-18 ep0 incident the hub said nothing for the entire outage; the only mails came from the boxes' own backup failures, and **only because the WEEKLY offsite run happened to fall inside the window**. Two days earlier, nothing would have fired at all | **SHIPPED — hub v0.106.0, 2026-08-18.** Reachability is now a second, independent signal: consecutive failed fetch windows counted per checker, `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default **3 windows (≈30–45 min)**, with paired `*_recovered` all-clears wired into `recoveredPairedDownTypes` — necessary because both recoveries are severity `info` and `severityNotifies` drops `info`. Threshold tunable via `alerting.box_unreachable_windows`. **The fill logic is untouched**: no threshold, throttle, band or escalate-once behaviour changed. Evidence: `internal/monitor/box_reachability_test.go` (Scenarios A–F) + `internal/notify/dispatcher_box_reachability_test.go` (the cross-package wiring, asserting an actual operator mail), plus three companion red-proofs each seen failing with a message naming the right cause | — | **PROVEN-LIVE still owed.** No real or constructed outage has exercised the emit path end to end, and one cannot be manufactured without making ep0 or the Hetzner API unreachable — ep0 is Tier 2 protected, so that is forbidden. The honest route is a constructed outage against a scratch hub instance with the tenantsync client pointed at a blackholed address. **Do not close this row on the unit tests** | CC | | **R-340** | **The new reachability check does not touch the surface that actually failed.** R-339 reports when the hub cannot READ ep0 — but the read it performs is the `usage` op, which is `proxmox-backup-manager` plus `df` over SSH, and therefore rides the **local API daemon**. The 2026-08-18 incident explicitly CLEARED that daemon: `proxmox-backup.service` was healthy throughout, and it was the **HTTPS proxy on 8007** that was wedged with a full accept queue. **So R-339's check would have returned green for all 9 h 37 m of that outage.** It closes the case where ep0 is unreachable *as a host*; it does not close the case that actually happened. This is not a defect in R-339 — it is the honest boundary of what it watches, recorded so a future reader does not mistake a green box gauge for a working off-site tier | **READY (M) — NEW 2026-08-18** | a tenantsync endpoint-script version bump (the op is added on ep0, so it needs the same version-gated rollout `ErrUsageUnsupported` already models) | Add a **health op** to `scripts/felhom-tenantsync.sh` that probes `https://127.0.0.1:8007/` on ep0 and reports the proxy's fd count and listen-queue depth, then surface it as a third signal. **Overlaps the connections spike (R-336's remaining half): both want the same observations from ep0, so whichever runs SECOND must reuse the first's evidence rather than re-measuring a protected machine twice** **REUSE, per this row's own instruction — the connections spike ran FIRST (2026-08-20) and already produced most of what the health op wants; do not re-measure a protected machine a third time.** Available in `audits/evidence-ep0-established-connections-2026-08-20/`: the proxy **fd count** and its type breakdown (`lsof` + `/proc//fd`), the **listen-queue depth** (`ss -lnt` — `Recv-Q 0`, `Send-Q 1024`), the **ESTAB/CLOSE-WAIT split**, the **per-peer** connection histogram, a **31-minute persistence diff** of full 4-tuples, and a **46.18 h** slope with Poisson bounds. What the health op would still add beyond these: a loopback `GET https://127.0.0.1:8007/` probe — the observation that distinguished "process problem" from "network problem" on 2026-08-18 and the one thing this spike did NOT take, because it is the surface R-339 cannot see. **And this spike sharpens what the op should report:** a rising **ESTAB** count is the live signal (CLOSE-WAIT was **0**, not merely flat), and per **R-344** the fd ceiling that matters may be the **agent's**, not only ep0's. | CC |