diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 03ab8fda..78502532 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -643,7 +643,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-342** | **The ep0 snapshot covers less than it looks like it covers, and the next person will assume otherwise.** Quoting `audits/evidence-ep0-pbs-upgrade-2026-08-18/stop2-snapshot.txt` verbatim: *"covers — the 38 GB system disk /dev/sda (root), i.e. the PBS packages, unit files, /etc/systemd drop-ins, nftables and wg config. DOES NOT — /mnt/pbs-datastore. That is /dev/sdb, a separate 100 GB VOLUME, and Hetzner server snapshots do not include attached volumes. The backup data is therefore NOT protected by this snapshot."* Snapshot **421440873** (`felhom-hetzner-20260818`, 15.06 GB, Available) was taken as the rollback for the 4.2.2→4.2.5 PBS upgrade. **Rolling it back restores software state, not the datastore.** That was *acceptable for that change* — a package install writes no datastore content — and the file says so. **The problem is what happens next:** this fact lives in an evidence file nobody will open again, and a snapshot named as "the rollback" reads as protecting everything on the box. ep0 holds the only off-premises copy of a real customer's data | **READY (S) — NEW 2026-08-18** | — | **Decide the safeguard for any future ep0 procedure that could touch `/mnt/pbs-datastore` — it does not exist and has not been designed.** Candidates: a Hetzner **Volume** snapshot (a different object from the server snapshot), a PBS-level sync to a second location, or an explicit written acceptance that the datastore is unprotected for the duration. **Nothing may be added to a runbook implying a safeguard exists until one does** | **Viktor decides; CC executes** — a risk-to-customer-data question | | **R-343** | **The managed controller floor was raised 0.214.0 → 0.216.0 — and it was NOT the no-op it was expected to be: it moved a live box nine seconds later.** Raised by the operator 2026-08-18 **12:36:58Z**, in a **separate save after** the artifact vouch. **Read back from the store, not the form** (`hub_settings.min_controller_version`, `GetGlobalMinControllerVersion` — `hub/internal/store/store.go:1751`): `min_controller_version = 0.216.0`, updated 12:36:58. **WHY IT HAD BEEN BEHIND — this was NOT drift, and describing it as "two releases behind" without this context reads as a defect it was not.** `publish-train-rules.md` rule 1 is *manifest before floor*, and rule 2 requires the floor field to be filled **LAST, in a separate save**, because the DB row overrides the env floor and **acts immediately on the next report cycle**. That rule was earned: on the 2026-07-11 publish train the floor was saved together with the manifest, acted at once, and pushed controller 0.113.0 onto Peti's box **~9 minutes ahead of agent 0.81.0** — the exact forbidden skew, benign only because that box had no NAS shares. R-120's row records the same deliberate choice (*"Floor untouched per publish-train rule 2"*). So the floor sitting at 0.214.0 was **policy being followed**, not neglect. **What it was functionally while it sat there:** not a live problem — every reporting box was at or above it — but a **safety net set two versions low**. The floor is what drags a box forward if it ever falls behind (restored from an old backup, reinstalled, long offline), and at 0.214.0 it would have pulled such a box only to two versions back, missing R-328's severity fix and R-335's follow-up. **THE MEASURED BLAST RADIUS — five reads, and the third and fifth are the findings.** **(1) Floor:** `0.216.0` @ 12:36:58Z, from the store. **(2) Per-customer overrides:** **zero** — all five `customer_configs` rows carry an empty `min_controller_version`, so nothing hides behind a lower override and the global applies to everyone. **(3) Every box's controller version:** `demo-felhom` **0.216.0**, `demo-hp` **0.216.0** — both AT the floor **now**; `drill-r50` **0.213.0** (status `blocked`, last report 2026-08-12, powered off/reverted) and `peti-felhom` **0.115.0** (host row DELETED) are **below** it but are **not reporting boxes**. **(4) Directives/holds:** no `managed floor HELD` line exists; the hub logged `[INFO] Global controller-version floor set to "0.216.0"`. **(5) THE FINDING — a controller DID auto-update after the raise.** `demo-felhom` had been on **0.214.0** since 2026-08-12 16:44 and the hub recorded `controller_updated — Controller frissítve: 0.214.0 → 0.216.0` at **12:37:07Z**, then `controller_started (0.216.0)` at 12:37:12Z — **nine seconds after the raise**, exactly the immediate action rule 2 documents. `demo-hp` was already on 0.216.0 (hand-deployed 2026-08-14 08:31) and did not move. **No error, warning or critical event followed** — the update completed and the controller came back up. **So the change was real, not inert: "every reporting box is at or above the floor" is true BECAUSE of the raise, not independently of it.** **Why it is safe by construction**, cited rather than asserted: `ResolveManagedFloor` (`hub/internal/store/store.go:2068`) sets `Held` and clears the floor entirely when `Floor > GoldenVersion` (the R-216 shape) — floor 0.216.0 **equals** golden 0.216.0, so that guard does not trip — and holds per-box when the box's agent is below the manifest's `MinAgent`, or unknown, or unparseable; both boxes report agent 0.129.0 against `MinAgent` 0.129.0, so the floor was served rather than held. That second guard remains armed for any box reporting with an old agent. **No ISO rebuild is required:** the golden is fetched at first boot from the hub's manifest, which is 0.216.0 — at the floor, not below it — and rule 5's `assert_golden_ge_floor` is a **build-time** gate for *future* builds (`scripts/iso/build-felhom-iso.sh:77`, called at **:267**) which **fails open with a warning** when its inputs are absent (`:78-82`: `if [[ -z "$golden" \| \| -z "$floor" ]]` → `log_warn "… UNENFORCED …"` → `return 0`), both confirmed in the script. **`peti-felhom` was NOT contacted and needs no contact** — from the PETI row: its host row was deleted 2026-07-15 and *"a report from a deleted host 401s and is not persisted"*, so it cannot receive a floor directive at all and the raise cannot reach it | **OPEN — NEW 2026-08-18.** Deliberately NOT closed: the task's closing condition was *all five reads clean, no directive served*, and read 5 shows a live box updated. It went cleanly and is the floor working as designed — but a change recorded as a no-op when it moved a customer box is exactly the kind of record that misleads later | — | **Confirm the 0.216.0 update on `demo-felhom` is healthy in normal operation** (it reported and restarted clean, but it has not yet run a full backup cycle on 0.216.0 at the time of writing), then close. Separately: `drill-r50` at 0.213.0 will be dragged to 0.216.0 by this floor if it is ever booted and reports — that is the floor doing its job, noted so it is not read as a surprise | CC | -| **R-344** | **`felhom-agent` leaks one TCP connection to PBS per poll cycle, forever, on both sides — and it is the whole of the ep0 descriptor leak.** Found by the 2026-08-20 connections spike (`audits/SPIKE-ep0-established-connections-2026-08-20.md`), which was looking for a Proxmox poll-rate problem and found ours instead. **Two defects compounding.** **(a)** `internal/pbs/client.go:56-60` builds `&http.Transport{TLSClientConfig: tlsCfg}` as a composite literal, so **`IdleConnTimeout` is the zero value = no limit** — `http.DefaultTransport` sets 90 s and a literal does not inherit it. **(b)** `cmd/felhom-agent/main.go:1486` (`pbsTargetsFromPVE`) builds **a fresh `pbs.Client` every cycle**, as its own doc comment states, so each cycle strands one idle keep-alive connection in a transport that is then unreachable — and an unreachable `http.Transport` does **not** close its connections; the `persistConn` read-loop goroutine keeps the socket alive. `CloseIdleConnections` / `IdleConnTimeout` / `MaxIdleConns` appear **nowhere in the repo** (grep: no matches). **Measured live, both sides, twice:** ep0 held 388 ESTAB (194 from each box) at 08:02:39Z and 392 at 08:33:42Z; the boxes held 194+194 and 196+196 at the same instants, and the four new sockets carried **the same four source ports** on both sides. **Zero sockets closed in 31 minutes**, and all carried keepalive timers with `retrans=0` — mutually held live idle connections, not half-open ones. One socket per agent `/snapshots` call (387 calls vs 388 sockets). Rate **201.6/day** across the fleet; ep0 reaches its 65536 ceiling in **~323 days**. **It is bilateral and the box side is the under-watched half:** each agent holds **196 of its 208** descriptors in these sockets. Its limit is 524287, so the boxes are in no danger *today* — which is why this stayed invisible, not why it is harmless. | **READY (S) — NEW 2026-08-20** | none — it is a self-contained change in `felhom-agent` | **No fix is proposed here: the spike-first gate forbids it and the spec is a separate task.** What the spec must settle, and none of it is decided: whether the per-cycle client construction is the thing to remove or the transport is the thing to share; what `IdleConnTimeout` should be given the 900 s poll and the 6 h verify cadence; whether `internal/hub/client.go:53` and `internal/proxmox/client.go:69` — **the same composite-literal pattern, built once at start-up so not leaking by this route today** — should be changed in the same pass or left alone with a test pinning why. **The proof obligation is the fd count, not the diff:** per standing rule 3 the positive observable is ep0's ESTAB count going FLAT between proxy restarts, measured over a window long enough to matter — a green test suite proves nothing here, and a 30-minute window proves nothing here either (that error is already recorded twice in R-336 and R-341). **FIX SHIPPED TO THE TWO DEMO BOXES AND PROVEN LIVE, 2026-08-20 — agent 0.130.0. THIS ROW STAYS OPEN: see R-347.** `audits/SPIKE-ep0-established-connections-2026-08-20.md` §"Fix and proof" + `evidence-agent-transport-leak-2026-08-20/`. **The fix is one field restored to the standard library's own value.** New leaf package `internal/httpx` owns `DefaultIdleConnTimeout = 90s` (= `http.DefaultTransport`'s value, so there is no invented number to justify) and `NewTransport`, which returns a **fresh** transport per call and treats zero-or-negative as **use the default, never "no timeout"**. All three hand-rolled transports now go through it; `grep '&http.Transport{'` matches only `httpx` itself. `internal/hub/client.go` + `internal/proxmox/client.go` were corrected in the same pass and **neither contributed to the ep0 leak** — both are built once per process and neither talks to ep0:8007. **P1 — the restart, outcome (i) within ONE second:** ep0 fd **415 -> 216**, `demo-hp`'s 199 established connections gone, **CLOSE-WAIT stayed 0**. So **outcome (ii) does NOT exist and gets no row** — ep0 reaps on peer FIN correctly, and the 543 `CLOSE-WAIT` at the 2026-08-18 wedge has another explanation. Ownership thereby proven a THIRD independent way (socket owner, access-log user agent, and now what dies with the process). **P2 — divergence, 1.03 h (the operator closed the >=4 h window early; no daily rate is extrapolated and none is needed):** control `demo-felhom` 199->203 (**+4**), fixed `demo-hp` 0->0 (**+0**). ep0's access log counts the opportunities directly: **each box made exactly 4 `/snapshots` + 4 `/version` calls** — same cadence, same work. **control 4 cycles -> 4 leaks; fixed 4 cycles -> 0 leaks.** **Positive observable per standing rule 3** (a zero leak is equally consistent with "the agent stopped working"): the fixed box's four poll cycles are in ep0's log, and the boxes' other traffic is near-identical (libwww-perl 924 vs 926, proxmox-backup-client 898 vs 898), so **the only difference between them is the binary**. **P3 — second box, 10:18:56Z: ep0 fd 220 -> 17 in under two seconds**, settling at 17-19 and returning to 17 between cycles. **17 is precisely ep0's `t0` baseline** (fd 17, ESTAB 0, 2026-08-18 09:51:22Z). **CORRECTION, made within the hour it was written:** this session's own STOP 1 report and the first CHANGELOG draft said *"does not clear the 388 descriptors already stuck on ep0 — those persist until that proxy restarts."* **That is wrong.** The descriptors were held on BOTH sides; restarting the agents released every one. **ep0 was read-only throughout and its proxy PID never changed (551655)** — the protected machine was never touched and did not need to be. **Tests:** `internal/pbs/client_leak_test.go` counts connections SERVER-side and models the abandonment, so it pins the consequence, not the field — it deliberately does not assert `err == nil` (true of the leaking code). **Two red-proofs, both seen failing:** removing the timeout gives *"after 5s the server still holds 5 open connection(s), want 0"* (the count is in the message, so it cannot be a timeout with another cause); and `DisableKeepAlives: true` — **which the leak test PASSES** — is caught only by `TestPBSClient_KeepAliveStillReuses` (*"3 sequential requests over 3 connection(s), want 1"*). **Scenario A alone would have accepted a fix that made the problem worse.** Both reverted. **Fleet sanity:** hub reports 0.130.0 on both, no `floor held` (0.130.0 > golden MinAgent 0.129.0), no `*_unreachable` event, PBS-DR gauge refreshing throughout. **Why it is NOT closed:** the binary is hand-installed on two demo boxes and unpublished, so a fresh install still ships the leaking agent — **R-347**. A fix living on two boxes by hand is not delivered. | CC | +| **R-344** | **`felhom-agent` leaks one TCP connection to PBS per poll cycle, forever, on both sides — and it is the whole of the ep0 descriptor leak.** Found by the 2026-08-20 connections spike (`audits/SPIKE-ep0-established-connections-2026-08-20.md`), which was looking for a Proxmox poll-rate problem and found ours instead. **Two defects compounding.** **(a)** `internal/pbs/client.go:56-60` builds `&http.Transport{TLSClientConfig: tlsCfg}` as a composite literal, so **`IdleConnTimeout` is the zero value = no limit** — `http.DefaultTransport` sets 90 s and a literal does not inherit it. **(b)** `cmd/felhom-agent/main.go:1486` (`pbsTargetsFromPVE`) builds **a fresh `pbs.Client` every cycle**, as its own doc comment states, so each cycle strands one idle keep-alive connection in a transport that is then unreachable — and an unreachable `http.Transport` does **not** close its connections; the `persistConn` read-loop goroutine keeps the socket alive. `CloseIdleConnections` / `IdleConnTimeout` / `MaxIdleConns` appear **nowhere in the repo** (grep: no matches). **Measured live, both sides, twice:** ep0 held 388 ESTAB (194 from each box) at 08:02:39Z and 392 at 08:33:42Z; the boxes held 194+194 and 196+196 at the same instants, and the four new sockets carried **the same four source ports** on both sides. **Zero sockets closed in 31 minutes**, and all carried keepalive timers with `retrans=0` — mutually held live idle connections, not half-open ones. One socket per agent `/snapshots` call (387 calls vs 388 sockets). Rate **201.6/day** across the fleet; ep0 reaches its 65536 ceiling in **~323 days**. **It is bilateral and the box side is the under-watched half:** each agent holds **196 of its 208** descriptors in these sockets. Its limit is 524287, so the boxes are in no danger *today* — which is why this stayed invisible, not why it is harmless. | **CLOSED 2026-08-20 — fixed, proven live on both boxes, published and vouched** | none — it is a self-contained change in `felhom-agent` | **No fix is proposed here: the spike-first gate forbids it and the spec is a separate task.** What the spec must settle, and none of it is decided: whether the per-cycle client construction is the thing to remove or the transport is the thing to share; what `IdleConnTimeout` should be given the 900 s poll and the 6 h verify cadence; whether `internal/hub/client.go:53` and `internal/proxmox/client.go:69` — **the same composite-literal pattern, built once at start-up so not leaking by this route today** — should be changed in the same pass or left alone with a test pinning why. **The proof obligation is the fd count, not the diff:** per standing rule 3 the positive observable is ep0's ESTAB count going FLAT between proxy restarts, measured over a window long enough to matter — a green test suite proves nothing here, and a 30-minute window proves nothing here either (that error is already recorded twice in R-336 and R-341). **FIX SHIPPED TO THE TWO DEMO BOXES AND PROVEN LIVE, 2026-08-20 — agent 0.130.0. THIS ROW STAYS OPEN: see R-347.** `audits/SPIKE-ep0-established-connections-2026-08-20.md` §"Fix and proof" + `evidence-agent-transport-leak-2026-08-20/`. **The fix is one field restored to the standard library's own value.** New leaf package `internal/httpx` owns `DefaultIdleConnTimeout = 90s` (= `http.DefaultTransport`'s value, so there is no invented number to justify) and `NewTransport`, which returns a **fresh** transport per call and treats zero-or-negative as **use the default, never "no timeout"**. All three hand-rolled transports now go through it; `grep '&http.Transport{'` matches only `httpx` itself. `internal/hub/client.go` + `internal/proxmox/client.go` were corrected in the same pass and **neither contributed to the ep0 leak** — both are built once per process and neither talks to ep0:8007. **P1 — the restart, outcome (i) within ONE second:** ep0 fd **415 -> 216**, `demo-hp`'s 199 established connections gone, **CLOSE-WAIT stayed 0**. So **outcome (ii) does NOT exist and gets no row** — ep0 reaps on peer FIN correctly, and the 543 `CLOSE-WAIT` at the 2026-08-18 wedge has another explanation. Ownership thereby proven a THIRD independent way (socket owner, access-log user agent, and now what dies with the process). **P2 — divergence, 1.03 h (the operator closed the >=4 h window early; no daily rate is extrapolated and none is needed):** control `demo-felhom` 199->203 (**+4**), fixed `demo-hp` 0->0 (**+0**). ep0's access log counts the opportunities directly: **each box made exactly 4 `/snapshots` + 4 `/version` calls** — same cadence, same work. **control 4 cycles -> 4 leaks; fixed 4 cycles -> 0 leaks.** **Positive observable per standing rule 3** (a zero leak is equally consistent with "the agent stopped working"): the fixed box's four poll cycles are in ep0's log, and the boxes' other traffic is near-identical (libwww-perl 924 vs 926, proxmox-backup-client 898 vs 898), so **the only difference between them is the binary**. **P3 — second box, 10:18:56Z: ep0 fd 220 -> 17 in under two seconds**, settling at 17-19 and returning to 17 between cycles. **17 is precisely ep0's `t0` baseline** (fd 17, ESTAB 0, 2026-08-18 09:51:22Z). **CORRECTION, made within the hour it was written:** this session's own STOP 1 report and the first CHANGELOG draft said *"does not clear the 388 descriptors already stuck on ep0 — those persist until that proxy restarts."* **That is wrong.** The descriptors were held on BOTH sides; restarting the agents released every one. **ep0 was read-only throughout and its proxy PID never changed (551655)** — the protected machine was never touched and did not need to be. **Tests:** `internal/pbs/client_leak_test.go` counts connections SERVER-side and models the abandonment, so it pins the consequence, not the field — it deliberately does not assert `err == nil` (true of the leaking code). **Two red-proofs, both seen failing:** removing the timeout gives *"after 5s the server still holds 5 open connection(s), want 0"* (the count is in the message, so it cannot be a timeout with another cause); and `DisableKeepAlives: true` — **which the leak test PASSES** — is caught only by `TestPBSClient_KeepAliveStillReuses` (*"3 sequential requests over 3 connection(s), want 1"*). **Scenario A alone would have accepted a fix that made the problem worse.** Both reverted. **Fleet sanity:** hub reports 0.130.0 on both, no `floor held` (0.130.0 > golden MinAgent 0.129.0), no `*_unreachable` event, PBS-DR gauge refreshing throughout. **CLOSED the same day.** The only thing holding it open was delivery, and **R-347 closed that**: 0.130.0 is tagged, published (`sha256 a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3`, reproducible byte for byte) and **vouched** in the Day-0 manifest, so a fresh install now gets the fixed agent. Both boxes were then reinstalled from the DOWNLOADED artifact rather than the proof build — see **R-349** for why that step is not optional. **Closing evidence, the positive observable rather than the absent one:** ep0 sits at **fd 17, ESTAB 0, CLOSE-WAIT 0** — its `t0` baseline — and returns there between poll cycles, with both agents demonstrably still polling. **ep0 was read-only for the entire arc** (spike, fix, proof, release) and its proxy PID never changed from 551655. | CC | | **R-345** | **`hub/Makefile` tags and pushes `:latest`, which the project's own rules forbid in two places.** Lines 21-22 of `docker-push`: `docker tag $(IMAGE):$(VERSION) $(IMAGE):latest` then `docker push $(IMAGE):latest`. `.claude/rules/hub.md:35` says *"Pin explicit versions, never `:latest`"* and `.claude/rules/manifests.md:15` repeats it. Verified present on `848368ec3`. Small, and the deployed manifests do pin a version, so nothing is currently broken by it — but a documented command that performs the prohibited action is a trap for whoever next reads the Makefile as the reference for how to publish, and a floating `:latest` on the registry is exactly the thing an emergency `kubectl set image` reaches for. Noticed while running the 2026-08-20 connections spike; unfiled until now. | **READY (XS) — NEW 2026-08-20** | — | Delete the two lines, or keep them behind an explicit opt-in target that says in a comment why it exists. Check whether a stale `:latest` tag already sits on `gitea.dooplex.hu/admin/felhom-hub` before deciding — an existing floating tag is the more dangerous half. | CC | | **R-346** | **`ActiveEnterTimestamp` answers a different question than the one a slope measurement asks, and on ep0 right now it is wrong by 5 h 56 m.** Found while taking R-341's first dated check. ep0's `proxmox-backup-proxy` has `MainPID=551655` started **2026-08-18 09:51:04Z** (`ps -o lstart=`), but `systemctl show -p ActiveEnterTimestamp` reads **03:54:54Z** and `NRestarts` reads **0** — because the 4.2.5-1 upgrade **re-exec'd** the daemon rather than restarting the unit, so systemd never observed a stop. Anyone anchoring "when did this proxy generation start" on `ActiveEnterTimestamp` would divide 388 descriptors by 52.1 h instead of 46.2 h and report **178/day instead of 201.6/day — ~12% low** — while every field consulted looks healthy and consistent. **This is the workspace rule's own case, in a new place:** ask of a timestamp *what exactly must have happened for this to be set?* Here the answer is "the unit entered active", which is not "this process started". `NRestarts=0` is the tell, and it reads like reassurance. | **READY (XS) — NEW 2026-08-20** | — | R-341's command already uses `ps -o lstart= -p $MainPID` and is correct; the risk is a future reader "improving" it to a systemd property. Add the reason as a comment beside that command in the R-341 row (done), and check whether any other slope or uptime check in the repo or in `scripts/felhom-tenantsync.sh` anchors on a systemd timestamp where it means a process start. | CC | | **R-347** | **The R-344 fix exists on two demo boxes by hand and NOWHERE ELSE — a box installed from the current image still ships the leaking agent.** Agent **0.130.0** was built on DooPlex and hand-installed on `demo-hp` and `demo-felhom` on 2026-08-20, deliberately **without** publishing: no Gitea package, no `v0.130.0` tag, no `artifact_agent_version` / `artifact_min_agent` / `artifact_golden_version` change, no staged self-update. **That was correct at the time** — publishing would have pushed the fix onto `demo-felhom` through the self-update path and destroyed the control the whole experiment rested on. The experiment is now finished, so the reason has expired and only the gap remains. **The gap is real but not urgent:** the leak takes ~323 days to reach ep0's 65536 ceiling with two boxes on it, and any restart of the agent clears the whole accumulation (P3, measured: ep0 220 -> 17 fd in under two seconds). A newly installed box therefore leaks slowly and self-heals on every agent deploy. **The CHANGELOG heading is `## UNRELEASED — v0.130.0 candidate` for exactly this reason** — the `release-complete` gate would otherwise convict on a version claiming to be a release that has no tag and no package, and it was right to. | **CLOSED 2026-08-20 — published, vouched, and the fleet reconciled onto the published bytes** | R-344 (done) | **This is a publish-train decision with an ordering rule attached (`documentation/runbooks/publish-train-rules.md`), so the GO is the operator's; CC executes.** The sequence: flip the CHANGELOG heading to `## v0.130.0` **in the same commit as** the tag, `bash scripts/release-agent.sh 0.130.0`, then the hub's Day-0 artifact manifest must vouch it — **that UI is operator-password-gated and CC cannot drive it**. Confirm `release-complete` goes green afterwards, on the tag and the package, not on the heading. **DONE 2026-08-20 on the operator's explicit word, including the artifact screen.** `bash scripts/release-agent.sh 0.130.0` — build, tag, publish, verify by independent download — produced **`sha256 a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3`**, 14,141,158 bytes, tag `v0.130.0` at `7569f34`. **Reproducible, checked rather than assumed** (R-186's property): a rebuild with `-trimpath -buildvcs=false` matches the published artifact byte for byte. **Vouched** in the hub Day-0 artifact manifest: `agent_version` 0.129.0 -> **0.130.0** and `agent_sha256` updated. **Publish-train rules honoured: ONLY the agent fields changed.** `golden_version` 0.216.0, `golden_sha256`, `wrapper_sha256` and **`min_agent` 0.129.0** were re-sent unchanged — `min_agent` expresses what the GOLDEN CONTROLLER requires, so raising it to 0.130.0 would have HELD the floor for every box not yet on 0.130.0, which is the opposite of shipping a fix. **The global floor was never touched** (still 0.216.0) and could not have been by accident: on hub v0.106.0 it is a **separate form with its own action** (`/configuration/global-floor`), so rule 2's "save the floor field last" hazard no longer exists in the shape its incident describes — worth knowing before the next train. Verified after: no `floor held` line for either box, no `*_unreachable`, both boxes reported 0.130.0, and the artifact downloads anonymously at the vouched sha. **The CHANGELOG heading was flipped to `## v0.130.0` only after the tag and package existed, so `release-complete` passes on the real thing and no `--no-verify` was used anywhere in the train.** **See R-349 for the trap this train exposed** — the fleet was briefly running a DIFFERENT binary under the same version name, and it has been reconciled. | **Viktor decides**, CC executes |