From 910fd911244c113b259691ddef69090716ab501a Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Thu, 20 Aug 2026 12:54:54 +0200 Subject: [PATCH] agent 0.130.0 published and vouched; R-347 closed, R-349 + R-350 filed Released via scripts/release-agent.sh: tag v0.130.0 at 7569f34, sha256 a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3, verified by independent download and reproducible byte for byte with -trimpath -buildvcs=false. Vouched agent 0.129.0 -> 0.130.0 in the Day-0 manifest. Only the agent fields changed: min_agent stays 0.129.0 because it states what the GOLDEN CONTROLLER requires, and raising it would have HELD the floor for every box below 0.130.0. Global floor untouched at 0.216.0 -- and on hub v0.106.0 it is a separate form with its own action, so publish-train rule 2's hazard no longer exists in the shape its incident describes. No --no-verify: the CHANGELOG heading was flipped only after the tag and package existed, so release-complete passes on the real artifact. R-349: the fleet was running a DIFFERENT binary under the same version name -- the proof deploy was a hand build, the release is -trimpath. Self-update could never have corrected it, because every version check compares the string. Both boxes reinstalled from the downloaded package. The proper fix exists in miniature as wrapper_sha256 and was never extended to the agent's own binary. R-350: I printed the hub password into the session transcript via curl -w '%{redirect_url}' -- the hub answers 303 and curl re-attaches the credential. Not in git, not in any committed file, not in the evidence directory. Rotation is the operator's call. ep0 closes at fd 17, ESTAB 0, CLOSE-WAIT 0 -- its t0 baseline -- and was read-only for this entire arc. --- REPORT-agent-transport-leak.md | 54 ++++++++++++- STATUS.md | 16 ++-- ...-ep0-established-connections-2026-08-20.md | 81 +++++++++++++++++++ .../p6-final-verification.txt | 19 +++++ .../p6-fleet-matches-vouched-artifact.txt | 19 +++++ documentation/backlog/OPEN-ITEMS.md | 4 +- 6 files changed, 184 insertions(+), 9 deletions(-) create mode 100644 documentation/audits/evidence-agent-transport-leak-2026-08-20/p6-final-verification.txt create mode 100644 documentation/audits/evidence-agent-transport-leak-2026-08-20/p6-fleet-matches-vouched-artifact.txt diff --git a/REPORT-agent-transport-leak.md b/REPORT-agent-transport-leak.md index c87d4473..793109bb 100644 --- a/REPORT-agent-transport-leak.md +++ b/REPORT-agent-transport-leak.md @@ -1,8 +1,10 @@ # REPORT — R-344: the agent's leaked PBS connections, fixed and proven on both boxes (2026-08-20) -**Outcome: the fix works, measured three independent ways, and ep0 is back to its baseline of 17 file -descriptors from 415.** Both demo boxes now run agent **0.130.0**. **Nothing is published** — that is -the one decision left, filed as R-347. +**Outcome: the fix works, measured three independent ways; ep0 is back to its baseline of 17 file +descriptors from 415; and 0.130.0 is now published and vouched.** Both demo boxes run the **byte-exact +published artifact**. R-347 is CLOSED. Two new findings came out of the release itself — **R-349** +(the fleet was briefly running a different binary under the same version name) and **R-350** (I printed +the hub password into the transcript; rotation is your call). ## 1. Confirmed baselines @@ -179,6 +181,52 @@ t=10:40:23Z pid=551655 fd=17 estab=0 ctrl(.2)=0 fix(.3)=0 CLOSE-WAIT=0 `CLOSE-WAIT 0`, `ESTAB` in the low single digits, `fd` at the baseline, proxy PID **551655** — the same process that has been running since 2026-08-18 09:51:04, never restarted by this work. +## 11b. The release (R-347, CLOSED) + +`bash scripts/release-agent.sh 0.130.0` — the one documented way (R-115): build, tag, publish, and +**verify by independent download**. + +| | | +|---|---| +| tag | `v0.130.0` at `7569f34` | +| sha256 | **`a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3`** | +| size | 14,141,158 bytes | +| reproducible | **yes, checked** — `-trimpath -buildvcs=false` rebuild matches byte for byte | + +**Vouched** in the Day-0 artifact manifest: agent 0.129.0 → **0.130.0** + its sha. **Only the agent +fields changed.** + +- **`min_agent` left at 0.129.0** — it states what the **golden controller** needs. Raising it to + 0.130.0 would have made the hub **HOLD the floor** for every box below 0.130.0, which is the + opposite of shipping a fix. +- **Global floor never touched** (0.216.0). On hub v0.106.0 it is a *separate form with its own + action*, so publish-train rule 2's "save the floor last" hazard no longer exists in the shape its + incident describes — the rule's reasoning holds, its mechanism has moved. +- After: no `floor held`, no `*_unreachable`, both boxes reporting 0.130.0, artifact downloading + anonymously at the vouched sha. + +**No `--no-verify` in the train.** The heading was flipped to `## v0.130.0` only after tag and package +existed. Flipping first and bypassing would have produced a red CI run and an alarm mail for a release +that worked — R-168's failure mode. + +## 11c. Two findings from the release + +**R-349 — the fleet was running a different binary under the same version name.** The proof deploy was +a hand build; the release builds `-trimpath -buildvcs=false`. Same source, same version string, +different bytes (`256e0829…` vs `a56a92a7…`). **Self-update could never have corrected it** — the boxes +already reported 0.130.0, so the vouched version looked installed. Every version check in the system +compares the *string*. Fixed by installing the **downloaded** artifact on both. The proper fix already +exists in miniature: `wrapper_sha256` does exactly this drift detection for the PBS wrapper and was +never extended to the agent's own binary. + +**R-350 — I printed the hub password into the transcript.** Confirming the vouch used +`curl -w '%{redirect_url}'`; the hub answers 303 and curl re-attaches the basic-auth credential to the +redirect target it prints. **Not in git, not in any committed file** (checked by content), not in the +evidence directory — it is in the session transcript on DooPlex. Every other call printed only the +length; this came through curl's own formatting. **Rotation is your call** — I did not do it +unilaterally, and I can do it file-to-file without printing the new value if you want. The reusable +half: `%{redirect_url}`, `-v` and `--libcurl` all re-render a basic-auth credential. + ## 12. Observations - **The closure refactor is not worth doing — recommend leaving it.** With the idle timeout restored an diff --git a/STATUS.md b/STATUS.md index d44c772c..21cd7204 100644 --- a/STATUS.md +++ b/STATUS.md @@ -132,11 +132,17 @@ record with no machine** — created 13 August, no host, no backups, nothing to off-site box is back to 17 open connections, its normal resting number, down from 415.** All of the built-up connections released themselves when the agents restarted; the off-site box was only ever read from, never touched. *(register: R-344)* -- **The fix is on the two demo machines by hand and NOT published yet** (R-347). A machine installed - from today's image still gets the old, leaking agent. That was deliberate — publishing it mid-test - would have contaminated the comparison — and the reason has now expired. **It is not urgent:** a new - machine would take the better part of a year to matter, and any agent update clears the build-up. - **Publishing is your call**, and it needs the operator-only artifact screen at the end. +- **PUBLISHED the same day, on your word** (R-347, closed). Agent **0.130.0** is released, and the hub + now hands it to any new machine. Both demo machines run the exact published copy. Nothing else on + that screen was changed — in particular the controller floor was left alone, and the "minimum agent" + setting too, because raising that would have **stopped** machines getting updates rather than + helping them. +- **Two things the release itself turned up.** (1) The machines were briefly running a *different* + build of the same version number — harmless here, but nothing in the system would ever have noticed, + because everything compares the version *name*. Now corrected, and filed so it cannot repeat + (R-349). (2) **I printed the hub password into my own session log** while confirming the change + (R-350). It is not in git and not in any saved file — but it is in the log on this machine. + **Changing it is your call**; I can do it without ever showing the new one. Ask and I will. - **The off-site box was updated, and it did not help — as expected** (R-341). On your ruling we installed the newer backup software for the practice, having first read its release notes and found **nothing** about the fault we have. The update went cleanly and everything works, but the leak diff --git a/documentation/audits/SPIKE-ep0-established-connections-2026-08-20.md b/documentation/audits/SPIKE-ep0-established-connections-2026-08-20.md index 7ea82477..aaa09f59 100644 --- a/documentation/audits/SPIKE-ep0-established-connections-2026-08-20.md +++ b/documentation/audits/SPIKE-ep0-established-connections-2026-08-20.md @@ -500,3 +500,84 @@ stays populated. Filed as R-348. `[::ffff:10.77.0.2]:port`, and the pattern expected `10.77.0.2:`. Caught 15 minutes in by noticing that a 0/0 split could not sum to 199. Fixed, then **one sample was proved by hand before committing the window to it** — the check that should have happened first. + +--- + +# Published — 2026-08-20, on the operator's word (R-347 CLOSED) + +Released through `scripts/release-agent.sh 0.130.0`, the one documented way (R-115), which builds, +tags, publishes and **verifies by independent download** rather than by its own say-so. + +| | | +|---|---| +| version / tag | **0.130.0** / `v0.130.0` at `7569f34` | +| sha256 | **`a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3`** | +| size | 14,141,158 bytes | +| reproducible | **yes, checked** — a rebuild with `-trimpath -buildvcs=false` matches byte for byte (R-186) | +| anonymous fetch | matches the vouched sha | + +## The vouch, and what was deliberately NOT changed + +Day-0 artifact manifest: `agent_version` 0.129.0 → **0.130.0**, `agent_sha256` updated. **Everything +else re-sent unchanged**, and each for a reason: + +- **`min_agent` stays 0.129.0.** It expresses what the **golden controller** requires, not what the + newest agent is. Raising it to 0.130.0 would have made the hub **HOLD the controller floor** for + every box not yet on 0.130.0 — the exact opposite of shipping a fix, and the manifest's own help + text says so: *"The hub HOLDS the floor for any box whose agent is below this."* +- **`golden_version` / `golden_sha256` / `wrapper_sha256`** — untouched; no controller was released. +- **The global floor was never touched** (still 0.216.0). **And on hub v0.106.0 it could not have been + by accident:** it is a *separate form with its own action* (`/configuration/global-floor`), so + publish-train rule 2's "the manifest screen carries the live DB floor — save it last" hazard no + longer exists in the shape its incident describes. Worth knowing before the next train; the rule's + reasoning still holds, its mechanism has moved. + +Verified after: no `floor held` line for either box, no `*_unreachable`, both boxes reporting 0.130.0. + +**No `--no-verify` anywhere in the train.** The CHANGELOG heading was flipped from +`## UNRELEASED — v0.130.0 candidate` to `## v0.130.0` **only after** the tag and package existed, so +`release-complete` passes on the real artifact. The ordering was: release at the UNRELEASED-heading +commit, then flip. The alternative — flipping first and bypassing the gate — would have produced a +genuinely red CI run and an alarm mail for a release that worked, which is R-168's failure mode. + +## The trap this train exposed — R-349 + +**Both boxes were running a different binary under the same version name, and nothing would ever have +noticed.** The proof deploy used a hand build (`go build -ldflags …`); the release builds with +`-trimpath -buildvcs=false` for reproducibility. Same source, same version string, **different bytes**: +`256e0829…` on the boxes against `a56a92a7…` published. + +The sharp edge is that **self-update cannot correct it**: the boxes already reported `0.130.0`, so the +vouched version looked installed and nothing would have happened, indefinitely. Every version check in +the system — the hub, `--version`, the artifact manifest — compares the version **string**, so the +divergence is invisible to all of them. + +Corrected by installing the **downloaded** artifact (not a local rebuild — the boxes get the bytes a +fresh install would get) on both. Both now report `a56a92a7…`. + +**The proper fix already exists in miniature:** `wrapper_sha256` makes exactly this drift visible for +the PBS wrapper — *"agents report the installed file's hash and a mismatch is surfaced on the host +page"*. It was simply never extended to the agent's own binary. R-349. + +## A mistake of mine in this train — R-350 + +Confirming the vouch used `curl -w '%{redirect_url}'`. The hub answers the POST with a **303**, and +curl renders the redirect target **with the basic-auth credentials re-attached** — so the hub operator +password was printed in cleartext into the session transcript. + +It is **not** in git, not in any committed file (checked by content, not by assumption), and not in +this evidence directory; it is in the Claude Code transcript on DooPlex. Every other call in the +session printed only the password's length — this arrived through curl's output formatting, which is +why the usual discipline missed it. Rotation is recommended and is the operator's call; the reusable +half is that **`%{redirect_url}`, `-v` and `--libcurl` all re-render a basic-auth credential** — +confirm a redirect with `%{http_code}` and read the flash from a follow-up GET. + +## Closing state + +``` +ep0 10:52:26Z pid=551655 fd=17 estab=0 ctrl(.2)=0 fix(.3)=0 CLOSE-WAIT=0 +``` + +**fd 17 is ep0's `t0` baseline**, and it returns there between poll cycles. The proxy is the same +process that has run since 2026-08-18 09:51:04 — **ep0 was read-only for this entire arc**, from the +spike through the fix to the release, and was never restarted, reconfigured or upgraded by any of it. diff --git a/documentation/audits/evidence-agent-transport-leak-2026-08-20/p6-final-verification.txt b/documentation/audits/evidence-agent-transport-leak-2026-08-20/p6-final-verification.txt new file mode 100644 index 00000000..2ffc7c72 --- /dev/null +++ b/documentation/audits/evidence-agent-transport-leak-2026-08-20/p6-final-verification.txt @@ -0,0 +1,19 @@ +=== final verification, 2026-08-20T10:50:03+00:00 === +--- ep0 settle (expect fd back toward 17, CLOSE-WAIT 0) --- +t=10:50:04Z pid=551655 fd=21 estab=4 ctrl(.2)=2 fix(.3)=2 CLOSE-WAIT=0 +t=10:51:15Z pid=551655 fd=17 estab=0 ctrl(.2)=0 fix(.3)=0 CLOSE-WAIT=0 +t=10:52:26Z pid=551655 fd=17 estab=0 ctrl(.2)=0 fix(.3)=0 CLOSE-WAIT=0 + +--- hub: agent version per host, and any floor-held --- + demo-felhom-8363b5 + 0.130.0 + demo-hp-bb76ea + 0.130.0 + 0.129.0 + +--- hub log: floor held / unreachable since the vouch --- +(empty = none) + +--- published artifact is downloadable as an anonymous client would fetch it --- + anonymous GET sha256: a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3 + tag on origin : 7edfea9aa9 diff --git a/documentation/audits/evidence-agent-transport-leak-2026-08-20/p6-fleet-matches-vouched-artifact.txt b/documentation/audits/evidence-agent-transport-leak-2026-08-20/p6-fleet-matches-vouched-artifact.txt new file mode 100644 index 00000000..fbd80189 --- /dev/null +++ b/documentation/audits/evidence-agent-transport-leak-2026-08-20/p6-fleet-matches-vouched-artifact.txt @@ -0,0 +1,19 @@ +=== reconciling the fleet onto the VOUCHED artifact === +downloaded from the package registry, not rebuilt locally — the boxes get the bytes a fresh install would get +UTC: 2026-08-20T10:49:30+00:00 + +--- ep0 before --- +t=10:49:31Z pid=551655 fd=20 estab=3 ctrl(.2)=2 fix(.3)=1 CLOSE-WAIT=0 + +--- demo-hp --- + staged sha : a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3 + now running: felhom-agent 0.130.0 sha=a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3 + service : active + +--- felhom-pve --- + staged sha : a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3 + now running: felhom-agent 0.130.0 sha=a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3 + service : active + +--- ep0 after both restarts --- +t=10:49:37Z pid=551655 fd=21 estab=4 ctrl(.2)=2 fix(.3)=2 CLOSE-WAIT=0 diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 3a39e7de..03ab8fda 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -646,8 +646,10 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-344** | **`felhom-agent` leaks one TCP connection to PBS per poll cycle, forever, on both sides — and it is the whole of the ep0 descriptor leak.** Found by the 2026-08-20 connections spike (`audits/SPIKE-ep0-established-connections-2026-08-20.md`), which was looking for a Proxmox poll-rate problem and found ours instead. **Two defects compounding.** **(a)** `internal/pbs/client.go:56-60` builds `&http.Transport{TLSClientConfig: tlsCfg}` as a composite literal, so **`IdleConnTimeout` is the zero value = no limit** — `http.DefaultTransport` sets 90 s and a literal does not inherit it. **(b)** `cmd/felhom-agent/main.go:1486` (`pbsTargetsFromPVE`) builds **a fresh `pbs.Client` every cycle**, as its own doc comment states, so each cycle strands one idle keep-alive connection in a transport that is then unreachable — and an unreachable `http.Transport` does **not** close its connections; the `persistConn` read-loop goroutine keeps the socket alive. `CloseIdleConnections` / `IdleConnTimeout` / `MaxIdleConns` appear **nowhere in the repo** (grep: no matches). **Measured live, both sides, twice:** ep0 held 388 ESTAB (194 from each box) at 08:02:39Z and 392 at 08:33:42Z; the boxes held 194+194 and 196+196 at the same instants, and the four new sockets carried **the same four source ports** on both sides. **Zero sockets closed in 31 minutes**, and all carried keepalive timers with `retrans=0` — mutually held live idle connections, not half-open ones. One socket per agent `/snapshots` call (387 calls vs 388 sockets). Rate **201.6/day** across the fleet; ep0 reaches its 65536 ceiling in **~323 days**. **It is bilateral and the box side is the under-watched half:** each agent holds **196 of its 208** descriptors in these sockets. Its limit is 524287, so the boxes are in no danger *today* — which is why this stayed invisible, not why it is harmless. | **READY (S) — NEW 2026-08-20** | none — it is a self-contained change in `felhom-agent` | **No fix is proposed here: the spike-first gate forbids it and the spec is a separate task.** What the spec must settle, and none of it is decided: whether the per-cycle client construction is the thing to remove or the transport is the thing to share; what `IdleConnTimeout` should be given the 900 s poll and the 6 h verify cadence; whether `internal/hub/client.go:53` and `internal/proxmox/client.go:69` — **the same composite-literal pattern, built once at start-up so not leaking by this route today** — should be changed in the same pass or left alone with a test pinning why. **The proof obligation is the fd count, not the diff:** per standing rule 3 the positive observable is ep0's ESTAB count going FLAT between proxy restarts, measured over a window long enough to matter — a green test suite proves nothing here, and a 30-minute window proves nothing here either (that error is already recorded twice in R-336 and R-341). **FIX SHIPPED TO THE TWO DEMO BOXES AND PROVEN LIVE, 2026-08-20 — agent 0.130.0. THIS ROW STAYS OPEN: see R-347.** `audits/SPIKE-ep0-established-connections-2026-08-20.md` §"Fix and proof" + `evidence-agent-transport-leak-2026-08-20/`. **The fix is one field restored to the standard library's own value.** New leaf package `internal/httpx` owns `DefaultIdleConnTimeout = 90s` (= `http.DefaultTransport`'s value, so there is no invented number to justify) and `NewTransport`, which returns a **fresh** transport per call and treats zero-or-negative as **use the default, never "no timeout"**. All three hand-rolled transports now go through it; `grep '&http.Transport{'` matches only `httpx` itself. `internal/hub/client.go` + `internal/proxmox/client.go` were corrected in the same pass and **neither contributed to the ep0 leak** — both are built once per process and neither talks to ep0:8007. **P1 — the restart, outcome (i) within ONE second:** ep0 fd **415 -> 216**, `demo-hp`'s 199 established connections gone, **CLOSE-WAIT stayed 0**. So **outcome (ii) does NOT exist and gets no row** — ep0 reaps on peer FIN correctly, and the 543 `CLOSE-WAIT` at the 2026-08-18 wedge has another explanation. Ownership thereby proven a THIRD independent way (socket owner, access-log user agent, and now what dies with the process). **P2 — divergence, 1.03 h (the operator closed the >=4 h window early; no daily rate is extrapolated and none is needed):** control `demo-felhom` 199->203 (**+4**), fixed `demo-hp` 0->0 (**+0**). ep0's access log counts the opportunities directly: **each box made exactly 4 `/snapshots` + 4 `/version` calls** — same cadence, same work. **control 4 cycles -> 4 leaks; fixed 4 cycles -> 0 leaks.** **Positive observable per standing rule 3** (a zero leak is equally consistent with "the agent stopped working"): the fixed box's four poll cycles are in ep0's log, and the boxes' other traffic is near-identical (libwww-perl 924 vs 926, proxmox-backup-client 898 vs 898), so **the only difference between them is the binary**. **P3 — second box, 10:18:56Z: ep0 fd 220 -> 17 in under two seconds**, settling at 17-19 and returning to 17 between cycles. **17 is precisely ep0's `t0` baseline** (fd 17, ESTAB 0, 2026-08-18 09:51:22Z). **CORRECTION, made within the hour it was written:** this session's own STOP 1 report and the first CHANGELOG draft said *"does not clear the 388 descriptors already stuck on ep0 — those persist until that proxy restarts."* **That is wrong.** The descriptors were held on BOTH sides; restarting the agents released every one. **ep0 was read-only throughout and its proxy PID never changed (551655)** — the protected machine was never touched and did not need to be. **Tests:** `internal/pbs/client_leak_test.go` counts connections SERVER-side and models the abandonment, so it pins the consequence, not the field — it deliberately does not assert `err == nil` (true of the leaking code). **Two red-proofs, both seen failing:** removing the timeout gives *"after 5s the server still holds 5 open connection(s), want 0"* (the count is in the message, so it cannot be a timeout with another cause); and `DisableKeepAlives: true` — **which the leak test PASSES** — is caught only by `TestPBSClient_KeepAliveStillReuses` (*"3 sequential requests over 3 connection(s), want 1"*). **Scenario A alone would have accepted a fix that made the problem worse.** Both reverted. **Fleet sanity:** hub reports 0.130.0 on both, no `floor held` (0.130.0 > golden MinAgent 0.129.0), no `*_unreachable` event, PBS-DR gauge refreshing throughout. **Why it is NOT closed:** the binary is hand-installed on two demo boxes and unpublished, so a fresh install still ships the leaking agent — **R-347**. A fix living on two boxes by hand is not delivered. | CC | | **R-345** | **`hub/Makefile` tags and pushes `:latest`, which the project's own rules forbid in two places.** Lines 21-22 of `docker-push`: `docker tag $(IMAGE):$(VERSION) $(IMAGE):latest` then `docker push $(IMAGE):latest`. `.claude/rules/hub.md:35` says *"Pin explicit versions, never `:latest`"* and `.claude/rules/manifests.md:15` repeats it. Verified present on `848368ec3`. Small, and the deployed manifests do pin a version, so nothing is currently broken by it — but a documented command that performs the prohibited action is a trap for whoever next reads the Makefile as the reference for how to publish, and a floating `:latest` on the registry is exactly the thing an emergency `kubectl set image` reaches for. Noticed while running the 2026-08-20 connections spike; unfiled until now. | **READY (XS) — NEW 2026-08-20** | — | Delete the two lines, or keep them behind an explicit opt-in target that says in a comment why it exists. Check whether a stale `:latest` tag already sits on `gitea.dooplex.hu/admin/felhom-hub` before deciding — an existing floating tag is the more dangerous half. | CC | | **R-346** | **`ActiveEnterTimestamp` answers a different question than the one a slope measurement asks, and on ep0 right now it is wrong by 5 h 56 m.** Found while taking R-341's first dated check. ep0's `proxmox-backup-proxy` has `MainPID=551655` started **2026-08-18 09:51:04Z** (`ps -o lstart=`), but `systemctl show -p ActiveEnterTimestamp` reads **03:54:54Z** and `NRestarts` reads **0** — because the 4.2.5-1 upgrade **re-exec'd** the daemon rather than restarting the unit, so systemd never observed a stop. Anyone anchoring "when did this proxy generation start" on `ActiveEnterTimestamp` would divide 388 descriptors by 52.1 h instead of 46.2 h and report **178/day instead of 201.6/day — ~12% low** — while every field consulted looks healthy and consistent. **This is the workspace rule's own case, in a new place:** ask of a timestamp *what exactly must have happened for this to be set?* Here the answer is "the unit entered active", which is not "this process started". `NRestarts=0` is the tell, and it reads like reassurance. | **READY (XS) — NEW 2026-08-20** | — | R-341's command already uses `ps -o lstart= -p $MainPID` and is correct; the risk is a future reader "improving" it to a systemd property. Add the reason as a comment beside that command in the R-341 row (done), and check whether any other slope or uptime check in the repo or in `scripts/felhom-tenantsync.sh` anchors on a systemd timestamp where it means a process start. | CC | -| **R-347** | **The R-344 fix exists on two demo boxes by hand and NOWHERE ELSE — a box installed from the current image still ships the leaking agent.** Agent **0.130.0** was built on DooPlex and hand-installed on `demo-hp` and `demo-felhom` on 2026-08-20, deliberately **without** publishing: no Gitea package, no `v0.130.0` tag, no `artifact_agent_version` / `artifact_min_agent` / `artifact_golden_version` change, no staged self-update. **That was correct at the time** — publishing would have pushed the fix onto `demo-felhom` through the self-update path and destroyed the control the whole experiment rested on. The experiment is now finished, so the reason has expired and only the gap remains. **The gap is real but not urgent:** the leak takes ~323 days to reach ep0's 65536 ceiling with two boxes on it, and any restart of the agent clears the whole accumulation (P3, measured: ep0 220 -> 17 fd in under two seconds). A newly installed box therefore leaks slowly and self-heals on every agent deploy. **The CHANGELOG heading is `## UNRELEASED — v0.130.0 candidate` for exactly this reason** — the `release-complete` gate would otherwise convict on a version claiming to be a release that has no tag and no package, and it was right to. | **READY (S) — NEW 2026-08-20** | R-344 (done) | **This is a publish-train decision with an ordering rule attached (`documentation/runbooks/publish-train-rules.md`), so the GO is the operator's; CC executes.** The sequence: flip the CHANGELOG heading to `## v0.130.0` **in the same commit as** the tag, `bash scripts/release-agent.sh 0.130.0`, then the hub's Day-0 artifact manifest must vouch it — **that UI is operator-password-gated and CC cannot drive it**. Confirm `release-complete` goes green afterwards, on the tag and the package, not on the heading. | **Viktor decides**, CC executes | +| **R-347** | **The R-344 fix exists on two demo boxes by hand and NOWHERE ELSE — a box installed from the current image still ships the leaking agent.** Agent **0.130.0** was built on DooPlex and hand-installed on `demo-hp` and `demo-felhom` on 2026-08-20, deliberately **without** publishing: no Gitea package, no `v0.130.0` tag, no `artifact_agent_version` / `artifact_min_agent` / `artifact_golden_version` change, no staged self-update. **That was correct at the time** — publishing would have pushed the fix onto `demo-felhom` through the self-update path and destroyed the control the whole experiment rested on. The experiment is now finished, so the reason has expired and only the gap remains. **The gap is real but not urgent:** the leak takes ~323 days to reach ep0's 65536 ceiling with two boxes on it, and any restart of the agent clears the whole accumulation (P3, measured: ep0 220 -> 17 fd in under two seconds). A newly installed box therefore leaks slowly and self-heals on every agent deploy. **The CHANGELOG heading is `## UNRELEASED — v0.130.0 candidate` for exactly this reason** — the `release-complete` gate would otherwise convict on a version claiming to be a release that has no tag and no package, and it was right to. | **CLOSED 2026-08-20 — published, vouched, and the fleet reconciled onto the published bytes** | R-344 (done) | **This is a publish-train decision with an ordering rule attached (`documentation/runbooks/publish-train-rules.md`), so the GO is the operator's; CC executes.** The sequence: flip the CHANGELOG heading to `## v0.130.0` **in the same commit as** the tag, `bash scripts/release-agent.sh 0.130.0`, then the hub's Day-0 artifact manifest must vouch it — **that UI is operator-password-gated and CC cannot drive it**. Confirm `release-complete` goes green afterwards, on the tag and the package, not on the heading. **DONE 2026-08-20 on the operator's explicit word, including the artifact screen.** `bash scripts/release-agent.sh 0.130.0` — build, tag, publish, verify by independent download — produced **`sha256 a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3`**, 14,141,158 bytes, tag `v0.130.0` at `7569f34`. **Reproducible, checked rather than assumed** (R-186's property): a rebuild with `-trimpath -buildvcs=false` matches the published artifact byte for byte. **Vouched** in the hub Day-0 artifact manifest: `agent_version` 0.129.0 -> **0.130.0** and `agent_sha256` updated. **Publish-train rules honoured: ONLY the agent fields changed.** `golden_version` 0.216.0, `golden_sha256`, `wrapper_sha256` and **`min_agent` 0.129.0** were re-sent unchanged — `min_agent` expresses what the GOLDEN CONTROLLER requires, so raising it to 0.130.0 would have HELD the floor for every box not yet on 0.130.0, which is the opposite of shipping a fix. **The global floor was never touched** (still 0.216.0) and could not have been by accident: on hub v0.106.0 it is a **separate form with its own action** (`/configuration/global-floor`), so rule 2's "save the floor field last" hazard no longer exists in the shape its incident describes — worth knowing before the next train. Verified after: no `floor held` line for either box, no `*_unreachable`, both boxes reported 0.130.0, and the artifact downloads anonymously at the vouched sha. **The CHANGELOG heading was flipped to `## v0.130.0` only after the tag and package existed, so `release-complete` passes on the real thing and no `--no-verify` was used anywhere in the train.** **See R-349 for the trap this train exposed** — the fleet was briefly running a DIFFERENT binary under the same version name, and it has been reconciled. | **Viktor decides**, CC executes | | **R-348** | **Every agent restart blanks the reported backup list for up to ~18 hours, and the comment that covers it says "unaffected".** Observed 2026-08-20 while deploying R-344: the first host reports after `demo-hp`'s agent restart carry **`0 backups`** (11:15:50 and 11:30:52 CEST, two consecutive), while the box's own `pvesm list` shows archives present on **both** tiers. `internal/backup/store.go`'s `Store` is in-memory and `byTarget` is repopulated only when a backup **runs** — daily for the local tier, weekly for offsite — so the field reads 0 until the next run. `restore_tests` did **not** blank, because that half has a durable on-disk companion (`RestoreTestState`, R-189). **It blinds no alarm, and that was CHECKED rather than assumed.** `hub/internal/monitor/deadline.go` scans back over stored reports with a 7-day `backupEvidenceLookback` whose own comment names this exact case — *"when the LATEST report carries none... and against an agent that stayed restarted for days"* — and `pbs_snapshots` stayed populated at 2 regardless. So this is an observability wart, **not** a safety hole, and it is filed at that severity deliberately. **What is actually wrong is the comment.** The `Store` doc says *"Backups are unaffected — their freshness has a ground truth on the storage (R-84)"*. That is true of the **consequence** and false of the **field**, and it sits three lines below a paragraph explaining that the very same sentence about restore-tests *"used to be here and it is now FALSE"* — so the file already carries one correction of this shape and invites the next reader to trust the surviving half. | **READY (XS) — NEW 2026-08-20** | — | Say what is measured: the field IS lost on restart and repopulates only when a backup runs; the freshness VERDICT is unaffected because the hub looks back 7 days. **Name `backupEvidenceLookback` in the comment** so the cross-repo dependency is visible from the agent side — today the agent's claim of safety rests on a hub constant it does not mention. Per the workspace rule, a comment asserting an invariant needs a test pinning it: the pin belongs on the HUB side, asserting the verdict survives a report carrying `backups: []`. | CC | +| **R-349** | **"Prove it by hand, then publish" leaves the fleet running a DIFFERENT binary under the SAME version name — and self-update cannot notice.** Hit on 2026-08-20 during the R-344 train, caught and corrected the same hour, filed because the next prove-then-publish train will hit it identically. **The mechanism:** a proof deploy is a hand build (`go build -ldflags "-X main.version=0.130.0"`), while `scripts/release-agent.sh` deliberately builds with **`-trimpath -buildvcs=false`** so the published artifact is reproducible (R-186). Same source, same version string, **different bytes**: `256e0829...` on the boxes vs **`a56a92a7...`** published and vouched. **Nothing corrects it automatically**, and that is the sharp edge: the boxes already report `0.130.0`, so the self-update path sees the vouched version as already installed and does nothing, **forever**. The divergence is invisible to every version check in the system — the hub, `--version`, and the artifact manifest all agree, because they all compare the version STRING. **Consequence if unnoticed:** the binary a customer box runs is not the binary the operator vouched, and not the one a reinstall would fetch — so a bug reproduced on the fleet may not exist in the published artifact, or vice versa. It is the same "one version name, two binaries" hazard `publish-agent.sh` already carries a comment about for `CGO_ENABLED`; that comment fixed the two ENTRY POINTS and does not cover a hand build during a proof. **Corrected here** by downloading the published artifact from the registry (not rebuilding it locally — the boxes get the bytes a fresh install would get) and installing it on both; both now report `sha256 a56a92a7...`, matching the vouch. | **READY (S) — NEW 2026-08-20** | — | Make the reconciliation a step, not a memory: the honest fix is for the agent to REPORT the sha256 of its own binary in the host report, so the hub can compare it against the vouched `agent_sha256` and flag drift — **exactly the mechanism `wrapper_sha256` already implements for the PBS wrapper** (R-50b), whose manifest help text says it *"makes host drift visible: agents report the installed file's hash and a mismatch is surfaced on the host page"*. The pattern exists and is proven; it simply was never extended to the agent's own binary. Cheaper interim: end every prove-then-publish train by installing the DOWNLOADED artifact. | CC | +| **R-350** | **SECURITY — the hub operator password was printed in cleartext into a session transcript by CC, 2026-08-20. Rotation recommended.** **What happened:** vouching the artifact manifest used `curl -w '%{redirect_url}'` for confirmation. The hub answers `POST /configuration/artifacts` with a **303**, and curl renders the redirect target **with the basic-auth credentials re-attached** — so the URL it printed contained `http://:@10.43.52.34:8080/configuration?flash=artifacts_set`. The password was never read aloud from the credentials file, never echoed deliberately, and every other call in the session correctly printed only `${#HUB_PW}`; it arrived through curl's own output formatting, which is why the usual discipline did not catch it. **Blast radius, stated precisely rather than minimised:** the value is **not** in git, not in `CHANGELOG.md`/`REPORT*.md`/any committed file (checked), and not in the evidence directory — it is in the Claude Code session transcript under `~/.claude/projects/` on DooPlex, which is operator-readable and persists across sessions. The hub UI is reachable only on the k3s ClusterIP and via the operator's own routes, not from the internet. **The value is deliberately not recorded here; it is stored out-of-band in the usual credentials file.** | **READY (S) — NEW 2026-08-20** | — | **Operator decides whether to rotate.** The hub's own `/configuration` password form does it (`current_password`/`new_password`/`confirm_password`), and per `hub-password-ui-2026-07-13` the DB override wins over the ConfigMap, which stays break-glass. CC can perform the rotation **file-to-file without printing the new value** (the `operator-present-one-time-secrets` convention) if asked — it did not do so unilaterally, because rotating a credential the operator holds in their own head or notes is their call, not CC's. **The reusable half, which matters more than this one password:** never use curl's `%{redirect_url}` (or `-v`, or `--libcurl`) against a basic-auth endpoint — all three re-render the credential. Confirm a redirect with `%{http_code}` and read the flash from a follow-up GET. | **Viktor decides**, CC executes | | **R-339** | **The hub was SILENT when it lost sight of the off-site stores — and a 9 h 37 m outage proved it.** Both box checkers (`OffsiteBoxChecker` over the Hetzner API, `PBSDRBoxChecker` over ep0's `usage` op) held their last snapshot and returned quietly on a failed fetch. That is **correct for a fill signal** — a missing reading must never be mistaken for 0%, which is why degraded data drives no band transition — but the consequence was that a completely dead off-site endpoint and a healthy one were **indistinguishable on the operator channel**. During the 2026-08-18 ep0 incident the hub said nothing for the entire outage; the only mails came from the boxes' own backup failures, and **only because the WEEKLY offsite run happened to fall inside the window**. Two days earlier, nothing would have fired at all | **SHIPPED — hub v0.106.0, 2026-08-18.** Reachability is now a second, independent signal: consecutive failed fetch windows counted per checker, `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default **3 windows (≈30–45 min)**, with paired `*_recovered` all-clears wired into `recoveredPairedDownTypes` — necessary because both recoveries are severity `info` and `severityNotifies` drops `info`. Threshold tunable via `alerting.box_unreachable_windows`. **The fill logic is untouched**: no threshold, throttle, band or escalate-once behaviour changed. Evidence: `internal/monitor/box_reachability_test.go` (Scenarios A–F) + `internal/notify/dispatcher_box_reachability_test.go` (the cross-package wiring, asserting an actual operator mail), plus three companion red-proofs each seen failing with a message naming the right cause | — | **PROVEN-LIVE still owed.** No real or constructed outage has exercised the emit path end to end, and one cannot be manufactured without making ep0 or the Hetzner API unreachable — ep0 is Tier 2 protected, so that is forbidden. The honest route is a constructed outage against a scratch hub instance with the tenantsync client pointed at a blackholed address. **Do not close this row on the unit tests** | CC | | **R-340** | **The new reachability check does not touch the surface that actually failed.** R-339 reports when the hub cannot READ ep0 — but the read it performs is the `usage` op, which is `proxmox-backup-manager` plus `df` over SSH, and therefore rides the **local API daemon**. The 2026-08-18 incident explicitly CLEARED that daemon: `proxmox-backup.service` was healthy throughout, and it was the **HTTPS proxy on 8007** that was wedged with a full accept queue. **So R-339's check would have returned green for all 9 h 37 m of that outage.** It closes the case where ep0 is unreachable *as a host*; it does not close the case that actually happened. This is not a defect in R-339 — it is the honest boundary of what it watches, recorded so a future reader does not mistake a green box gauge for a working off-site tier | **READY (M) — NEW 2026-08-18** | a tenantsync endpoint-script version bump (the op is added on ep0, so it needs the same version-gated rollout `ErrUsageUnsupported` already models) | Add a **health op** to `scripts/felhom-tenantsync.sh` that probes `https://127.0.0.1:8007/` on ep0 and reports the proxy's fd count and listen-queue depth, then surface it as a third signal. **Overlaps the connections spike (R-336's remaining half): both want the same observations from ep0, so whichever runs SECOND must reuse the first's evidence rather than re-measuring a protected machine twice** **REUSE, per this row's own instruction — the connections spike ran FIRST (2026-08-20) and already produced most of what the health op wants; do not re-measure a protected machine a third time.** Available in `audits/evidence-ep0-established-connections-2026-08-20/`: the proxy **fd count** and its type breakdown (`lsof` + `/proc//fd`), the **listen-queue depth** (`ss -lnt` — `Recv-Q 0`, `Send-Q 1024`), the **ESTAB/CLOSE-WAIT split**, the **per-peer** connection histogram, a **31-minute persistence diff** of full 4-tuples, and a **46.18 h** slope with Poisson bounds. What the health op would still add beyond these: a loopback `GET https://127.0.0.1:8007/` probe — the observation that distinguished "process problem" from "network problem" on 2026-08-18 and the one thing this spike did NOT take, because it is the surface R-339 cannot see. **And this spike sharpens what the op should report:** a rising **ESTAB** count is the live signal (CLOSE-WAIT was **0**, not merely flat), and per **R-344** the fd ceiling that matters may be the **agent's**, not only ep0's. | CC |