From 36f86300209e9ce914f2a97dac16ee9eeea249e2 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Sun, 30 Aug 2026 18:49:53 +0200 Subject: [PATCH] R-341 check taken, and the golden-currency bypass declared Two gates blocked the R-331 hub push. One is FIXED, one is BYPASSED, and the difference is stated rather than blurred. FIXED -- due-checks (R-341, 5 days overdue). The +7d measurement was TAKEN on ep0 rather than deferred again. Precondition passed: proxy still MainPID 551655, ps -o lstart= still 2026-08-18 09:51:04, NRestarts=0, so this is the same proxy generation as t0 (anchor is ps, not ActiveEnterTimestamp, which reads 03:54:54Z here -- R-346's trap). Result: fd = 17. Not 17 more -- seventeen TOTAL, exactly the documented baseline, against 405 at the first check. Socket histogram: one LISTEN, ESTAB 0, CLOSE-WAIT 0. The verdict is UNANSWERABLE, not "the upgrade fixed it". R-341 asks whether the PBS 4.2.5-1 upgrade changed the fd slope; inside this interval we removed the leak OURSELVES (R-344, agent 0.130.0, now live on both boxes). A slope of ~0 measures our fix, not the upgrade, and reading it the other way would credit a changelog that was read in advance and found to contain no such mechanism. The perturbation pre-registered for this window was Phase C at ~3%; the actual perturbation was the removal of the entire phenomenon. Row closed as moot. What it DOES establish is worth more than the original question: twelve days after the R-344 fix, same proxy generation, no restart to hide behind, ep0 sits at baseline with zero established connections. R-336's ~323-day runway concern retires with it. BYPASSED -- golden-currency. Controller v0.224.0 and v0.225.0 are released and the newest golden bake carries 0.223.0, so a machine installed right now gets neither. The gate is RIGHT. This push therefore uses `git push --no-verify`, declared here, in hub/CHANGELOG.md, in REPORT.md and on R-242. A BYPASS, not a waiver: the gate offers a waiver only for a release that DELIBERATELY needs no golden, and these need one. The operator was asked and ruled bypass-now-bake-later, on the ground that neither fix bites a day-0 box -- R-330 is a nightly false alarm about apps a new box has not installed yet, R-331 is a hub display over backups a new box has not taken yet -- and both arrive by self-update. That ground is recorded because it is what to re-check: it does NOT extend to a release changing first-boot behaviour. OWED: bake a golden carrying 0.225.0 and vouch it (RUNBOOK-manual-build.md 4.1, three-field change, MinAgent 0.129.0). Fourth bypass of this gate, and the gap is now two releases wide rather than one. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM --- REPORT.md | 48 ++++++++++++++++++- .../step1-fd-and-sockets.txt | 14 ++++++ documentation/backlog/OPEN-ITEMS.md | 5 +- hub/CHANGELOG.md | 10 ++++ 4 files changed, 73 insertions(+), 4 deletions(-) create mode 100644 documentation/audits/evidence-r341-plus7d-2026-08-30/step1-fd-and-sockets.txt diff --git a/REPORT.md b/REPORT.md index f673939f..67a8f040 100644 --- a/REPORT.md +++ b/REPORT.md @@ -94,7 +94,53 @@ PENDING at the time of writing — see the follow-up commit. The two halves ship either order: the hub renders "unknown" for any box still on controller 0.224.0, which is correct rather than wrong. -## 7. Not done, and why +## 7. The push bypassed a gate, deliberately, and here is the declaration + +**`git push --no-verify` was used for this change.** `repo_gates.py`'s `golden-currency` gate was +CONVICTED and it was RIGHT: controller **v0.224.0** and **v0.225.0** are released and the newest golden +bake carries **0.223.0**, so a machine installed right now receives neither fix. + +**This is a BYPASS, not a waiver.** The gate offers a waiver only for a release that *deliberately needs +no golden*; these need one. **The operator was asked and ruled bypass-now-bake-later**, on the stated +ground that neither fix bites a day-0 box — R-330 is a nightly false alarm about apps a new box has not +installed yet, R-331 is a hub-side display over backups a new box has not taken yet — and both arrive by +self-update afterwards. That ground is recorded on R-242 precisely because it is the thing to re-check: +**it does not extend to a release that changes first-boot behaviour.** + +**A golden carrying 0.225.0 is OWED** (`RUNBOOK-manual-build.md` §4.1; the vouch is a three-field change, +`MinAgent 0.129.0`). This is the **fourth** bypass of this gate, and the gap it names is now two releases +wide rather than one. + +**The other failing gate was fixed, not bypassed.** `due-checks` was red on R-341's `+7 d` measurement, +five days overdue. It was **taken** during this session — see §8. + +## 8. R-341's overdue check was taken, and its premise did not survive + +Unrelated to R-331; it blocked the same push, so it was done rather than deferred. Evidence: +`documentation/audits/evidence-r341-plus7d-2026-08-30/step1-fd-and-sockets.txt`. + +**Precondition passed**, which is what makes the reading interpretable: `proxmox-backup-proxy` still +`MainPID 551655`, `ps -o lstart=` still `2026-08-18 09:51:04`, `NRestarts=0` — the same proxy generation +as t0, so nothing restarted and re-based the count. (The anchor is `ps`, not `ActiveEnterTimestamp`, +which reads 03:54:54Z here — R-346's trap, avoided.) + +**Result: fd = 17.** Not 17 more — seventeen total, exactly the documented baseline, against **405** at +the first check on 2026-08-20. The socket histogram holds **one LISTEN and nothing else**: ESTAB 0, +CLOSE-WAIT 0. + +**The verdict is "unanswerable", not "the upgrade fixed it".** R-341 asks whether the PBS 4.2.5-1 +upgrade changed the fd slope. Inside this interval **we removed the leak ourselves** (R-344, agent +0.130.0, now live on both boxes). A slope of ~0 measures our fix, not the upgrade — and reading it the +other way would credit a changelog that was read in advance and found to contain no such mechanism. The +perturbation pre-registered for this window was Phase C at ~3%; the actual perturbation was the removal +of the entire phenomenon. The question is now **moot**, and the row is closed as such. + +**What it does establish, which is worth more than the original question:** twelve days after the R-344 +fix, on the same proxy generation with no restart to hide behind, ep0 sits at baseline with zero +established connections. The 388-descriptor accumulation has not returned, and R-336's ~323-day runway +concern retires with it. + +## 9. Not done, and why - **No staleness verdict on the card.** `monitor.OffsiteChecker` already owns that and alarms on it. A second verdict over the same data is two things that can disagree — a shape this codebase has already diff --git a/documentation/audits/evidence-r341-plus7d-2026-08-30/step1-fd-and-sockets.txt b/documentation/audits/evidence-r341-plus7d-2026-08-30/step1-fd-and-sockets.txt new file mode 100644 index 00000000..6fe3ab78 --- /dev/null +++ b/documentation/audits/evidence-r341-plus7d-2026-08-30/step1-fd-and-sockets.txt @@ -0,0 +1,14 @@ +=== R-341 +7d check (taken late: 2026-08-30, due 2026-08-25) === +MainPID=551655 +ps lstart: Tue Aug 18 09:51:04 2026 +NRestarts=0 +ActiveEnterTimestamp=Tue 2026-08-18 03:54:54 UTC # R-346: NOT the anchor +fd_count=17 +now_utc=2026-08-30T16:40:43Z +--- ss -lnt sport :8007 --- +State Recv-Q Send-Q Local Address:Port Peer Address:Port +LISTEN 0 1024 *:8007 *:* +--- socket state histogram --- + 1 LISTEN +--- agent-side check: who holds connections now --- +Recv-Q Send-Q Local Address:Port Peer Address:PortProcess diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 0ac5436b..61f92eb6 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -156,7 +156,7 @@ the fault was real. Full observables: `tests/campaign11-evidence-2026-08-05/jour | **R-240** | **A backup that covered nothing calls itself „Sikeres".** On a configured box with no app selected for off-site backup, a run reports status `ok` with the warning „Sikeres — nincs mentésre jelölt alkalmazás" — *successful* immediately beside *nothing is selected*. Measured as T4 on the final walk, 2026-08-07; flagged once before (2026-08-06) and deliberately not touched then, because the task that noticed it forbade changing that path. **It is the same rhetorical shape the project has spent a fortnight removing** — R-203's *a warning beside a success is read as a success*, R-234's *„✓ Rendben" over an app that was skipped*, R-225's *unknown rendered as zero* — one notch weaker each time, and this is the weakest and last of them. The state itself is honest and must stay `ok`: an unconfigured box reporting `incomplete` forever is its own defect, pinned by a test. **The defect is the word „Sikeres", not the verdict.** Wording such as „Nincs mentésre jelölt alkalmazás — ez a futás semmit nem mentett" says the same thing without congratulating the customer on it. | **READY** — owner Viktor | -| **R-242** | **A controller release that changes customer-visible behaviour is not delivered until a golden carries it — and nothing enforces that.** R-239 is the symptom; this is the mechanism, recorded **2026-08-07** and deliberately **NOT built** (the task that found it scoped it as a record-only item). **Two releases went out without a golden and the gap was invisible until a walk measured it from the customer's side**: v0.204.0 (R-237) and v0.205.0 (R-234) were written, tested, pushed, and CHANGELOG'd, and every one of those steps passed while a machine installed that night received neither. The register said CLOSED; the fleet said otherwise. **Nothing in the release path knows a golden exists.** The version bump, the image push, the CHANGELOG entry and the register closure are all repo-local; the manifest's `golden_version` is edited by a separate operator act, in a different repo, with no link back. **Proposed shapes, cheapest first — the choice is the operator's and is not taken here.** (a) **A release-path checklist step** — one line in the controller's end-of-session checklist: *a release that changes customer-visible behaviour is not finished until a golden carries it or a register row says why not.* Costs nothing, catches nothing mechanically. (b) **A gate in `repo_gates.py`** comparing the manifest's `golden_version` against the newest released controller and FAILING (or warning) past a tolerance of one minor. Mechanical, runs on every push, and would have fired the morning after v0.204.0. (c) **A hub-side checker** — the hub already knows every box's running controller version from `/hosts` and the vouched golden from the manifest; a periodic comparison against the newest published image would catch drift the repo cannot see, including a vouch that was made and then rolled back. **Earliest catch: (b).** It fires on the push that creates the gap, before any box is installed, and it needs no live fleet. **(c) catches strictly more but only after boxes exist.** (a) is worth doing regardless because it is free. **Not built. No gate was written this session.** **⚠ IT RECURRED WITHIN A DAY, WHICH IS THE ARGUMENT FOR BUILDING IT.** Controller **v0.206.0** shipped the R-241 fixes on 2026-08-07 while the vouched golden still carried **0.205.0** — so a machine installed on the morning of 2026-08-08 would have received neither. Third occurrence of the shape in three days (R-111/R-115/R-120 are the older family). **✅ SHAPE (b) BUILT 2026-08-08 — `scripts/golden_currency_gate.py`, registered in `repo_gates.py` as gate 7.** **It was shown FAILING against that exact state before anything was baked**, which is its red-proof and the reason its own introducing push needed `--no-verify` (stated in the session report rather than worked around): `newest released controller : 0.206.0 / newest golden baked : 0.205.0 → CONVICTED`. **IT IS `--fast`, AND THAT FORCED ITS DESIGN:** both `.githooks/pre-push` AND CI run `repo_gates.py --fast`, so a non-fast gate would run in NEITHER — the R-29 census failure this runner exists to end. **THEREFORE IT CHECKS THE BAKE, NOT THE VOUCH**, because the vouched version lives only in the hub's `hub_settings` with no copy in git, and putting a copy there would create a second source of truth that can drift — a green gate over a false claim being the worst outcome available. **A bake without a vouch still passes: that half is NOT closed and stays on this row.** It also compares versions rather than behaviour, so a release changing nothing customer-visible trips it too — accepted deliberately, because judging that by hand is what failed three times and the cost of a false trip is one bake; a waiver belongs here, never in a habit of bypassing. **VOUCHED 2026-08-08 with the operator's approval** — golden `0.206.0` / sha `c85230b4…108e`; `agent_version` and `min_agent` both stayed `0.127.0`, and `wrapper_sha256` was carried through explicitly because the handler clears it when omitted. **The gate was CONVICTED before the bake and OK after it** — red→green on the same command, which is its proof that it measures something real. **⚠ THE GATE FIRED FOR REAL, 2026-08-08 — and it was right.** Controller **v0.207.0** (R-249/R-252/R-253) is released, tested and pushed, and **no golden carries it** — the newest bake is 0.206.0 — so `golden_currency_gate.py` FAILED, saying exactly the true thing: *a machine installed right now would receive v0.206.0*. **The `felhom.eu` push therefore used `git push --no-verify`, declared here, in the commit message and in the session report.** A bypass and NOT a waiver, deliberately: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **Owed: bake golden 0.207.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; the vouch is a three-field change). **This row's own remaining half is unchanged — nothing gates the VOUCH itself.** **✅ THE OWED BAKE IS DONE, SAME DAY — golden 0.207.0 baked, published, round-trip verified and VOUCHED (2026-08-08).** The gate went from red to **green**, and the `--no-verify` bypass declared above is now historical rather than standing. **Round trip is the evidence, not the build log:** the published bytes were downloaded back — 656 879 192 B, sha256 `20ec9602…22995`, both identical to what the bake reported — and **`./etc/felhom-controller-image` read OUT of the downloaded archive says `felhom-controller:0.207.0`**, which is the delivered artifact naming the controller it will start. **The vouch was a three-field change with all three checked deliberately** (`MinAgent 0.127.0` read from the golden's controller CHANGELOG header, not assumed; `agent_version` already ≥ it; `min_agent` not above `agent_version`, so not the R-216 shape) and verified by **re-reading the manifest rather than trusting the flash**. **This row's remaining half is UNCHANGED and is the whole of what is still open: nothing gates the VOUCH itself** — the currency gate's own docstring says it checks the bake, so a baked-but-unvouched golden still passes it silently. Evidence: `tests/golden-0.207.0-2026-08-08/`. **⚠ RED AGAIN, 2026-08-08 (second time in two days) — controller v0.208.0 (R-254) is released and the vouched golden is 0.207.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.207.0 and none of today's fixes. **The `felhom.eu` push used `git push --no-verify`, declared in the commit, the CHANGELOG and the session report** — **a bypass, not a waiver**, on the same reasoning as yesterday: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **Owed: bake golden 0.208.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change, `MinAgent 0.127.0` unchanged). **Note the cadence this is establishing: two releases, two bakes owed within 24 h.** That is the argument for this row's OTHER half — nothing gates the vouch, so the only thing standing between a release and an undelivered fleet is somebody remembering. | **READY — the vouch half only** — owner Viktor | +| **R-242** | **A controller release that changes customer-visible behaviour is not delivered until a golden carries it — and nothing enforces that.** R-239 is the symptom; this is the mechanism, recorded **2026-08-07** and deliberately **NOT built** (the task that found it scoped it as a record-only item). **Two releases went out without a golden and the gap was invisible until a walk measured it from the customer's side**: v0.204.0 (R-237) and v0.205.0 (R-234) were written, tested, pushed, and CHANGELOG'd, and every one of those steps passed while a machine installed that night received neither. The register said CLOSED; the fleet said otherwise. **Nothing in the release path knows a golden exists.** The version bump, the image push, the CHANGELOG entry and the register closure are all repo-local; the manifest's `golden_version` is edited by a separate operator act, in a different repo, with no link back. **Proposed shapes, cheapest first — the choice is the operator's and is not taken here.** (a) **A release-path checklist step** — one line in the controller's end-of-session checklist: *a release that changes customer-visible behaviour is not finished until a golden carries it or a register row says why not.* Costs nothing, catches nothing mechanically. (b) **A gate in `repo_gates.py`** comparing the manifest's `golden_version` against the newest released controller and FAILING (or warning) past a tolerance of one minor. Mechanical, runs on every push, and would have fired the morning after v0.204.0. (c) **A hub-side checker** — the hub already knows every box's running controller version from `/hosts` and the vouched golden from the manifest; a periodic comparison against the newest published image would catch drift the repo cannot see, including a vouch that was made and then rolled back. **Earliest catch: (b).** It fires on the push that creates the gap, before any box is installed, and it needs no live fleet. **(c) catches strictly more but only after boxes exist.** (a) is worth doing regardless because it is free. **Not built. No gate was written this session.** **⚠ IT RECURRED WITHIN A DAY, WHICH IS THE ARGUMENT FOR BUILDING IT.** Controller **v0.206.0** shipped the R-241 fixes on 2026-08-07 while the vouched golden still carried **0.205.0** — so a machine installed on the morning of 2026-08-08 would have received neither. Third occurrence of the shape in three days (R-111/R-115/R-120 are the older family). **✅ SHAPE (b) BUILT 2026-08-08 — `scripts/golden_currency_gate.py`, registered in `repo_gates.py` as gate 7.** **It was shown FAILING against that exact state before anything was baked**, which is its red-proof and the reason its own introducing push needed `--no-verify` (stated in the session report rather than worked around): `newest released controller : 0.206.0 / newest golden baked : 0.205.0 → CONVICTED`. **IT IS `--fast`, AND THAT FORCED ITS DESIGN:** both `.githooks/pre-push` AND CI run `repo_gates.py --fast`, so a non-fast gate would run in NEITHER — the R-29 census failure this runner exists to end. **THEREFORE IT CHECKS THE BAKE, NOT THE VOUCH**, because the vouched version lives only in the hub's `hub_settings` with no copy in git, and putting a copy there would create a second source of truth that can drift — a green gate over a false claim being the worst outcome available. **A bake without a vouch still passes: that half is NOT closed and stays on this row.** It also compares versions rather than behaviour, so a release changing nothing customer-visible trips it too — accepted deliberately, because judging that by hand is what failed three times and the cost of a false trip is one bake; a waiver belongs here, never in a habit of bypassing. **VOUCHED 2026-08-08 with the operator's approval** — golden `0.206.0` / sha `c85230b4…108e`; `agent_version` and `min_agent` both stayed `0.127.0`, and `wrapper_sha256` was carried through explicitly because the handler clears it when omitted. **The gate was CONVICTED before the bake and OK after it** — red→green on the same command, which is its proof that it measures something real. **⚠ THE GATE FIRED FOR REAL, 2026-08-08 — and it was right.** Controller **v0.207.0** (R-249/R-252/R-253) is released, tested and pushed, and **no golden carries it** — the newest bake is 0.206.0 — so `golden_currency_gate.py` FAILED, saying exactly the true thing: *a machine installed right now would receive v0.206.0*. **The `felhom.eu` push therefore used `git push --no-verify`, declared here, in the commit message and in the session report.** A bypass and NOT a waiver, deliberately: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **Owed: bake golden 0.207.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; the vouch is a three-field change). **This row's own remaining half is unchanged — nothing gates the VOUCH itself.** **✅ THE OWED BAKE IS DONE, SAME DAY — golden 0.207.0 baked, published, round-trip verified and VOUCHED (2026-08-08).** The gate went from red to **green**, and the `--no-verify` bypass declared above is now historical rather than standing. **Round trip is the evidence, not the build log:** the published bytes were downloaded back — 656 879 192 B, sha256 `20ec9602…22995`, both identical to what the bake reported — and **`./etc/felhom-controller-image` read OUT of the downloaded archive says `felhom-controller:0.207.0`**, which is the delivered artifact naming the controller it will start. **The vouch was a three-field change with all three checked deliberately** (`MinAgent 0.127.0` read from the golden's controller CHANGELOG header, not assumed; `agent_version` already ≥ it; `min_agent` not above `agent_version`, so not the R-216 shape) and verified by **re-reading the manifest rather than trusting the flash**. **This row's remaining half is UNCHANGED and is the whole of what is still open: nothing gates the VOUCH itself** — the currency gate's own docstring says it checks the bake, so a baked-but-unvouched golden still passes it silently. Evidence: `tests/golden-0.207.0-2026-08-08/`. **⚠ RED AGAIN, 2026-08-08 (second time in two days) — controller v0.208.0 (R-254) is released and the vouched golden is 0.207.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.207.0 and none of today's fixes. **The `felhom.eu` push used `git push --no-verify`, declared in the commit, the CHANGELOG and the session report** — **a bypass, not a waiver**, on the same reasoning as yesterday: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **Owed: bake golden 0.208.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change, `MinAgent 0.127.0` unchanged). **Note the cadence this is establishing: two releases, two bakes owed within 24 h.** That is the argument for this row's OTHER half — nothing gates the vouch, so the only thing standing between a release and an undelivered fleet is somebody remembering. **⚠ RED AGAIN, 2026-08-30 — controller v0.224.0 (R-330) and v0.225.0 (R-331) are released and the newest golden carries 0.223.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.223.0 and neither of today's fixes. **The `felhom.eu` push used `git push --no-verify`, declared in the commit message, in `hub/CHANGELOG.md` and in `REPORT.md` — a BYPASS, not a waiver**, on the same reasoning as the two 2026-08-08 entries above: the gate offers a waiver only for a release that *deliberately needs no golden*, and these need one. **The operator was asked and ruled bypass-now-bake-later on 2026-08-30**, on the stated ground that neither fix bites a DAY-0 box — R-330 is a nightly false alarm about apps a new box has not installed yet, and R-331 is a hub-side display over backups a new box has not taken yet — and both arrive by self-update afterwards. **That ground is recorded because it is the thing to re-check, not a general licence: the next release that changes first-boot behaviour cannot reuse it.** **OWED: bake a golden carrying 0.225.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change — `golden_version` + `agent_version` + `min_agent`; MinAgent is 0.129.0 per both CHANGELOG headers). **Cadence note, unchanged and now worse: this is the fourth bypass of this gate, and the gap it names is now two releases wide rather than one.** | **READY — the vouch half only** — owner Viktor | | **R-243** | **A box in the R-241 state silently stops backing up off-site, and NO ALARM OF ANY KIND FIRES.** Found by the R-241 spike (2026-08-07) as a by-product; **not part of the walk's finding and not previously filed.** The R-241 state is self-locking in a second, worse way than the recovery-journey dead end: `escrow_state` is stuck `pending` forever (the auto-confirm flips only on a hash match, and the hash cannot match a key the box minted itself), and `runOffboxBackup` returns at the escrow gate (`offbox.go:743`) before touching anything. **So off-site backups never run again — and the hub never notices.** All three signals that could catch it are excluded, each for its own individually-correct reason, verified in the hub this session: `offsite_stale` — `isStale` (`monitor/offsite.go:135`) returns false unless `EscrowState == "escrowed"`, and its own comment reads *"Pending/disabled = normal onboarding, never stale"*, so the box is classified as **still being set up, forever**; `offsite_delivery_stuck` — `monitor/offsite_delivery.go:91` skips the `applied` shape, and delivery genuinely IS applied (the credential was consumed and the target is in every report); `backup_failed` — never fires, because nothing fails: the run returns `nil` before it starts. **Three correct exclusions leaving one state unobserved.** This is the same class as the workspace `CLAUDE.md` "presence is not success" rule, one level up: **the absence of a failure is being read as the presence of a working tier.** **Partly subsumed by R-241's fix** — a box that recovers leaves this state — but **not for a box that does not**, and the alarm gap is what makes "does not" survivable indefinitely. **Not fixed; no code written.** **⚠ UPDATED 2026-08-07 (v0.206.0) — the STATE this row describes can no longer be entered, but the ALARM GAP is untouched and the row stays open.** R-241's mint guard means a box no longer mints a key over a sealed package, so it no longer arrives in the "escrow stuck pending against a self-minted key" state by itself. **What replaces it is a state that is VISIBLE rather than silent:** the box declares `offsite.state=awaiting_recovery_key` and the customer is offered the recovery screen. **But the hub still raises nothing for it**, and for the same three reasons: `isStale` needs `escrowed`, the delivery checker skips the `applied` shape, and `backup_failed` needs a run that never happens. **So a box whose customer never acts still stops backing up off-site with no operator signal** — the difference is that the customer can now see it and act, where before nobody could. **The remaining work is an operator-side signal for a box held in `awaiting_recovery_key` past some age**, and it is deliberately not bundled into R-241's fix. **⚠ MEASURED ON A REBUILD, 2026-08-07 (fifth walk) — the gap is real for the state this row describes, and NOT for the state a rebuild produces.** 88 seconds after the walk5 guest was destroyed and rebuilt, the hub emitted `offsite_delivery_stuck` (**warning**) and wrote an **operator-channel** `notification_log` row recording `offsite_credential_restaged` / status **REFUSED** with an accurate reason — *"the credential was applied and worked; the target was lost afterwards … a guest rebuild does, R-193"*. So on the **regressed-apply** shape the operator IS told, promptly and correctly, and this row's *"skips the applied shape"* does not apply. The gap stands for a box that reaches the held state **without** a prior working tier in its report history. **Recorded so the row is not read wider than it measures.** | **READY** — owner Viktor | | **R-244** | **The customer DELETE cascade leaves `app_log_issues` behind, and it is systematic across every venue ever torn down.** Found **2026-08-07** while verifying the `finalwalk` teardown with a **full census** (every table, every column) rather than a per-table query. After a cascade that logged `COMPLETE … full teardown`, **61 rows still matched `finalwalk`**. Four of the five sources are **deliberate and correct** — the cascade's own header states *"Provenance/events are NEVER wiped — audit outlives every tier"*: `events` 16, `notification_log` 14, `host_deletions` 1, `customer_resets` 1. **The fifth is a gap:** `app_log_issues` 29 rows, which the residue purge does not touch (its logged leg covers `reports`/`app_telemetry`/`app_log_tails`/`log_tail_requests`/`notif_prefs`/`selfbind_tokens`/`appliance_registrations` — not this table). **It is not a `finalwalk` quirk:** rows still reference **`c11` 40, `rewalk` 20, `part4` 24** — all three torn down 2026-08-06, whose ledger recorded *"0 occurrences"*. **That prior claim was measured with a narrower query and does not survive a full census; the correction is recorded rather than the measurement quietly redone.** **Why it was probably never written, established rather than assumed:** the table is a **fleet-wide aggregate** keyed on `app_name`+`fingerprint` with an `affected_customers` JSON list — of the 29 `finalwalk` rows, **12 reference only `finalwalk`** (orphans, safely deletable) and **17 are shared with LIVE customers** (`demo-felhom`, `peti-felhom`, …) and **must not be deleted, only de-referenced.** A naive `DELETE … WHERE customer LIKE` would destroy a live customer's issue history — which is very likely why the leg does not exist, and is the reason this is not a one-line fix. **Severity is LOW and stated plainly: no secret material is involved** — app name, fingerprint, message text, counts, timestamps. What survives is a deleted customer's *identifier* inside an aggregate row. **Proposed shape:** a residue leg that (a) removes the customer id from `affected_customers`/`context_customer`, and (b) deletes rows whose `affected_customers` becomes empty; plus a one-off sweep for the four already-torn-down venues. **The general lesson is the reusable part:** *a per-table absence query is not a census.* The teardown verification is now a full-schema sweep, and that is what found this. **Not fixed** — a cascade change needs its own red-proof and this session was scoped as a spike plus two operations. Evidence: `tests/teardown-finalwalk-2026-08-07.md`. **⚠ STILL OWED, AND NOW MEASURED RATHER THAN ESTIMATED (2026-08-08 census, read-only, no truncation).** `app_log_issues` holds **1309 rows**; **71 reference a torn-down venue** (`finalwalk`, `c11`, `rewalk`, `part4`); of those **44 are ORPHANS** — they name only torn-down customers and are safely deletable — and **27 are SHARED with a live customer** (`demo-felhom`, `peti-felhom`, …) and **must be de-referenced, never deleted**. 1238 rows are untouched. **The 27 are exactly why the leg was never written**, and why a `DELETE … WHERE customer LIKE` would destroy a live customer's issue history. **What it needs, precisely:** a cascade leg that (a) removes the customer id from `affected_customers` / `context_customer`, and (b) deletes only rows whose `affected_customers` becomes empty; plus a one-off sweep for the four venues already gone. **Why it was NOT done on 2026-08-08:** the fix is hub code, and that session's scope forbade a hub version bump; a hand-run SQL mutation over 71 rows — 27 of them needing surgical de-referencing — with no tested code path and no red-proof is precisely the shape that goes wrong on a live database. **It accumulates one venue at a time, so the next walk adds to it**; the numbers above mean the next session starts from data rather than a guess. **⚠ IT GREW AGAIN, AS PREDICTED — walk5 teardown, 2026-08-08.** The fifth walk's venue was torn down with a full-schema census taken **before and after**: **168 rows → 67**. Of the 67, **37 are by design** (`events` 21, `notification_log` 14, `host_deletions` 1, `customer_resets` 1) and **30 are `app_log_issues`** — this row's gap, and the count was **predicted in the pre-run enumeration rather than discovered afterwards**, which is the difference from the ledger that once recorded *"0 occurrences"* from a narrower query. **The running total across torn-down venues therefore rises from 71 to ~101 rows** (`finalwalk`, `c11`, `rewalk`, `part4`, now `walk5`) — the shared-with-a-live-customer subset must still be **de-referenced, never deleted**. **It accumulates one venue at a time and it did so again.** Evidence: `tests/walk5-r201-2026-08-07/teardown-walk5-2026-08-08.md`. | **READY** — owner Viktor | @@ -522,7 +522,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-336** | **The offsite DR endpoint is polled about once per second, and that is what turned a slow leak into an outage.** ep0's PBS proxy served **~85,000 requests/day** — a flat **3,538/hour**, every hour, from two boxes: `74,445 GET /api2/json/admin/datastore` (`libwww-perl`, i.e. PVE's `pvestatd`) and `73,171 GET /admin/datastore/felhom-offsite/status` (`proxmox-backup-client`). Two pollers asking substantially the same question at the same rate. On 2026-08-18 this walked a connection leak in the proxy to its 1024-fd soft limit in **14 days**, wedging the offsite tier for 9½ hours (`audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md`). The `LimitNOFILE=65536` drop-in applied that morning raises the ceiling **but does not fix the leak** — it converts a fortnightly outage into a multi-year one, which is mitigation, not a fix. A DR endpoint that is written to weekly does not need to be asked about every second | **READY (M) — NEW 2026-08-18** | — | **CORRECTED 2026-08-18 (evening) — the easy lever named here does not exist.** This cell used to read *"PVE storage status is the prime suspect, and its interval is tunable"*. **The first half is right and the second half is false.** `pvestatd` stats EVERY configured storage on each 10-second cycle, and Proxmox staff have stated the interval is not designed to be configurable — so there is no knob to turn down. The only lever PVE actually offers is disabling the storage entry (`pvesm set --disable 1`) around the backup window, and that is **substantially more than a tuning knob**: it collides with `felhom-agent/internal/pbsdr/manager.go`'s health model, where an inactive-but-existing entry drives the consume-the-one-time-secret recovery path. So the fix is a design question (does the hub still need a 15-minute fill reading at all, given R-339 now reports reachability separately?), not a config edit. **Doc-only correction — no agent code was changed.** The remaining step is unchanged: cut the poll rate by whatever means survives that question, then confirm the fd count between restarts stops climbing — the positive observable, per standing rule 3. **Baseline measured 2026-08-18, and the FIRST measurement published was WRONG.** The initial "~85/day, matching the ~73/day implied by the failure" came from a single 17-minute window whose delta was **one descriptor** — a sample of one cannot carry a daily rate, and the agreement that made it feel solid was coincidence. **Re-measured over two independent windows the same morning: 183/day (31 min) and 200/day (5.6 h)** — ~2.6x the published figure, putting the runway to the 65536 ceiling at **~357 days, not the ~2 years first claimed**. **And the named mechanism is the minority one:** across that window `CLOSE-WAIT` held flat at 1 while `ESTAB` grew 45→49 — *all* the growth was established connections, and at the wedge the split was 1011 ESTAB / 543 CLOSE-WAIT. **The fix must target connections the proxy never reaps, not just `CLOSE-WAIT` sockets.** The PBS 4.2.5-1 upgrade (2026-08-18) did NOT change the slope and was never expected to — see R-341 **SPIKE 2026-08-20 — THE PREMISE OF THIS ROW DOES NOT SURVIVE MEASUREMENT, and that is a change in what the row IS, not new evidence on it.** `audits/SPIKE-ep0-established-connections-2026-08-20.md`. **The leak is OURS, and the poll rate is not what feeds it.** Every one of the 388 leaked descriptors is an ESTABLISHED connection held open by **`felhom-agent`** on the boxes — 194 on each, `ss -tnp` naming a single PID per box, and **zero** held by `pvestatd` or `proxmox-backup-client`. Confirmed independently from ep0's access log over the same 46.18 h window: `libwww-perl` (pvestatd) **81,192 requests -> 0 descriptors**, `proxmox-backup-client` **80,061 requests -> 0 descriptors**, `Go-http-client/1.1` (the agent) **387 `/snapshots` calls -> 388 sockets — one per call, within one**. So **162,404 requests, 99.5% of the traffic, produce 0% of the leak.** **Mechanism, named from source:** `felhom-agent/internal/pbs/client.go:56-60` builds `&http.Transport{TLSClientConfig: tlsCfg}` — a composite literal, so `IdleConnTimeout` is the zero value = **no limit** (`http.DefaultTransport` sets 90 s; a literal does not inherit it) — and `cmd/felhom-agent/main.go:1486` (`pbsTargetsFromPVE`) builds **a fresh client every cycle**, as its own doc comment states. Each cycle therefore strands one idle keep-alive connection in a transport nothing ever closes; `CloseIdleConnections`/`IdleConnTimeout`/`MaxIdleConns` appear **nowhere** in the agent repo. Cadences reconcile without fitting: 900 s hub poll (184.7 cycles) + 6 h `DefaultVerifyCadence` (7.7 cycles) = 192.4 predicted vs **194 observed per box**. **CONSEQUENCE — RE-RANK.** The remaining step recorded above ("cut the poll rate, then confirm the fd count stops climbing") **would have produced a null result and read as a failed fix.** Cutting the Proxmox poll rate removes ~99.5% of ep0's request load and **zero** descriptors. The poll rate is still wrong on its own terms — 85,000 requests/day to a weekly-write DR endpoint — but it is now a **scaling/cost item, not the leak fix**, and the leak fix is **R-344**. **Q3 (is the leak proportional to the request rate?) is PREDICTED not-proportional and NOT YET MEASURED** — Phase C is held at STOP 1 with its prediction pre-registered in `evidence-ep0-established-connections-2026-08-20/phaseC-prediction.txt`. Do not record a proportionality verdict here until that window has run. **RE-SCOPED 2026-08-20 — THIS ROW IS NO LONGER A LEAK FIX, AND ITS RECORDED NEXT-STEP WOULD HAVE "FIXED" NOTHING WHILE LOOKING LIKE A FAILED FIX.** That near-miss is the reason the spike-first rule exists and it is kept here deliberately. The old next-step read: *cut the poll rate by whatever means survives that question, then confirm the fd count between restarts stops climbing.* Had it been executed, the fd count would have kept climbing at the same ~200/day, the poll reduction would have been recorded as ineffective, and the real defect — **ours, in `felhom-agent`, R-344** — would have been further from being found, not closer. **Measured 2026-08-20:** `pvestatd` (`libwww-perl`) and `proxmox-backup-client` made **162,404 requests** in a 46 h window and leaked **zero** descriptors; the agent made 811 and leaked **388**. The fix (agent 0.130.0) took ep0 from 388 accumulated descriptors to its **baseline of 17**, with the poll rate completely unchanged — 85,000/day before and after. **WHAT THIS ROW ACTUALLY IS NOW — a SCALING concern, still worth fixing on its own merits:** ~85,000 requests/day to a DR endpoint that is WRITTEN TO WEEKLY, from two boxes. That is ~42,500/box/day, so **at fifty customers it is ~2.1 million requests/day — about 25 requests/second, constantly, against a CX33**. The design question is unchanged and is still the hard part: does the hub still need a 15-minute fill reading at all, given R-339 reports reachability separately? And the lever remains awkward — `pvestatd` stats every configured storage on each 10-second cycle with no tunable interval, so the only PVE-side lever is disabling the storage entry, which collides with `felhom-agent/internal/pbsdr/manager.go`'s health model. **NEW ACCEPTANCE CRITERION, since the old one is void:** the fd count is NOT the observable for this row any more — that belongs to R-344 and is already satisfied. Measure the REQUEST RATE at ep0's access log, and state the projected rate at the target customer count. | CC | | **R-337** | **`/backup/status` lagged a completed backup by minutes on one box and not the other — and it RESOLVED ITSELF, which is why this is WATCHING and not a defect.** During the R-336 recovery on 2026-08-18, `demo-hp`'s snapshot landed on ep0 at **03:58:43Z** (complete manifest; the host's own task index says `OK`) — yet `GET /backup/status` was **still serving the superseded 03:27:00Z failure at ~04:03Z**, four-plus minutes later. `demo-felhom` showed its new result within ~40 s of completion. **The lag cleared on its own:** demo-hp's 04:07:35Z host report carries `felhom-pbs success=true, 4.29 GB`, and the hub is green for both boxes. **The first draft of this row claimed the success was "still reported as failed" — that was written before the next report arrived and it was wrong; the corrected claim is a several-minute skew between the two boxes, not a stuck value.** It is recorded because a status field that can trail its own artifact by minutes will, during an incident, be read as a second failure — this session nearly did — and because the asymmetry between the two boxes is unexplained | **WATCHING — NEW 2026-08-18** | another observation, ideally during an incident rather than constructed | **Do not open a fix on this as written.** First establish the intended refresh path for `/backup/status` after an out-of-schedule run; only if the skew is not simply collection cadence is there anything to pin. If it is cadence, close this row and say so | CC | | **R-338** | **`demo-hp` is not on the R-50 island at all, and `operations/nodes.md` states that it is.** The page records both fleet boxes as island-migrated 2026-07-25. True of `felhom-pve`; **false of `demo-hp`**, whose `agent.json` has `listen_addr: 192.168.0.87:8443` — the customer LAN address — and **no `island_bridge`/`island_guest_addr` keys at all**, whose guest 9201 has `net0` only (no `eth1`), and whose `vmbr9` exists with **zero members**. The controller's `controller.yaml` points at the LAN address, so the box works; this is inventory drift, not breakage. **Two costs.** A session trusting the page addresses the wrong endpoint — that happened on 2026-08-18 and the resulting timeout was briefly read as a fault. And the agent's local API is **bound to the customer LAN on this box** rather than to a point-to-point island, which is the exposure R-50 was built to remove — so a documented security property is claimed for a box that does not have it | **READY (S) — NEW 2026-08-18** | — | Decide which is true: migrate `demo-hp` to the island, or correct `nodes.md`. Leaving both is the one option that keeps the doc lying | Viktor decides; CC executes | -| **R-341** | **Does the fd slope change after the PBS 4.2.5 upgrade? — two dated checks, and the answer is expected to be NO.** ep0 was upgraded 4.2.2-1 → 4.2.5-1 on 2026-08-18 09:51Z on the operator's ruling, **for rehearsal value, not as a fix**: the full changelog range was read (128 lines, all three entries) and swept for connection-handling vocabulary, and it contains **no mechanism** by which descriptor reaping would change — the single keyword hit was `S3 … honor the node's proxy settings`, HTTP-proxy config for S3, not the PBS proxy daemon. **The 32-minute post-upgrade window is indistinguishable from the before window** (+5 fd/1919 s = 225/day vs +4 fd/1885 s = 183/day; the two differ by ONE descriptor and Poisson uncertainty on such counts is ±2, so both are consistent with one unchanged rate — the higher after-figure is noise, not a regression). **Thirty minutes cannot settle it in either direction and this row exists so nobody pretends it did.** New `t0` = **fd 17 at 2026-08-18 09:51:22Z, proxy PID 551655**; before-rate to beat = **183–200/day**. **Interpretation fixed in advance** (`evidence-ep0-pbs-upgrade-2026-08-18/stop1-ruling.txt`, written before any numbers existed): unchanged = EXPECTED, not a failed upgrade; changed = a SURPRISE needing explanation, not a confirmation | **WATCHING — FIRST CHECK TAKEN 2026-08-20: the rate is UNCHANGED, exactly as pre-registered. Second check still due 2026-08-25** | elapsed time only | **Two dated checks, both CC:** **+24 h — 2026-08-19 ~10:00Z** and **+7 d — 2026-08-25 ~10:00Z**. Command (the incident's own positive observable): `ssh root@ 'PID=$(systemctl show proxmox-backup-proxy -p MainPID --value); ls /proc/$PID/fd \| wc -l; ss -lnt "( sport = :8007 )"; ss -tn state all "( sport = :8007 )" \| awk "NR>1{print \$1}" \| sort \| uniq -c'`. **Record the ESTAB/CLOSE-WAIT split, not just the total** — the split is what says which leak it is. If PID ≠ 551655 the window is void: something restarted the proxy and the count began again **FIRST CHECK — TAKEN 2026-08-20 08:02:13Z, and it was taken at +46.2 h, NOT +24 h.** No session ran between 18 and 20 August, so the 2026-08-19 date passed as pure elapsed time; the delay is stated here rather than backfilled. **The longer window is a BETTER measurement, not a degraded one** — the leaked count is **388** against the 4 and 5 descriptors the original answer rested on, roughly **97x**, so the uncertainty falls from about +/-50% to about +/-5%. **Precondition PASSED:** PID still **551655**, `ps -o lstart` 2026-08-18 09:51:04, `NRestarts=0` on both units. **Result:** fd **17 -> 405** over **166,251 s** = **201.6 fd/day**, Poisson +/-10.2/day (1s), 2s band **181.2-222.1**. Pre-registered range was **370-450** (confirm-band 330-490); **observed 388**, near the centre. **VERDICT: unchanged — the EXPECTED result, and not a failed upgrade.** **Composition:** ESTAB 0 -> 388, **CLOSE-WAIT 0 — absent from the histogram entirely**, so 100% of the growth is established connections and `CLOSE-WAIT` is not merely the minority half. Runway from fd 405 at 201.6/day to the 65536 ceiling: **~323 days (~2027-07-09)**. Full working: `audits/SPIKE-ep0-established-connections-2026-08-20.md` + `evidence-ep0-established-connections-2026-08-20/step1-slope-computation.txt`. **MEASUREMENT TRAP for the 2026-08-25 check, found on this run — see R-346:** the anchor must be `ps -o lstart= -p $MainPID`, NOT `systemctl show -p ActiveEnterTimestamp`, which reads 03:54:54Z for this generation (the upgrade re-exec'd the proxy; systemd never saw a stop, `NRestarts` is still 0) and would put the rate ~15% low. **PERTURBATION NOTE for the 2026-08-25 check:** this spike's Phase C quietens `pvestatd` on demo-hp for one overnight window inside that 7-day interval. Quantified so nobody reads the shortfall as a change in the leak: even if the leak were fully proportional to the request rate (which Parts 1-2 predict it is NOT), a 10 h half-rate window inside 168 h shifts the 7-day slope by **~3%**, far inside the +/-10% Poisson band on a ~1,400-descriptor count. **The +7 d reading remains usable.** | CC on both dates | +| **R-341** | **Does the fd slope change after the PBS 4.2.5 upgrade? — two dated checks, and the answer is expected to be NO.** ep0 was upgraded 4.2.2-1 → 4.2.5-1 on 2026-08-18 09:51Z on the operator's ruling, **for rehearsal value, not as a fix**: the full changelog range was read (128 lines, all three entries) and swept for connection-handling vocabulary, and it contains **no mechanism** by which descriptor reaping would change — the single keyword hit was `S3 … honor the node's proxy settings`, HTTP-proxy config for S3, not the PBS proxy daemon. **The 32-minute post-upgrade window is indistinguishable from the before window** (+5 fd/1919 s = 225/day vs +4 fd/1885 s = 183/day; the two differ by ONE descriptor and Poisson uncertainty on such counts is ±2, so both are consistent with one unchanged rate — the higher after-figure is noise, not a regression). **Thirty minutes cannot settle it in either direction and this row exists so nobody pretends it did.** New `t0` = **fd 17 at 2026-08-18 09:51:22Z, proxy PID 551655**; before-rate to beat = **183–200/day**. **Interpretation fixed in advance** (`evidence-ep0-pbs-upgrade-2026-08-18/stop1-ruling.txt`, written before any numbers existed): unchanged = EXPECTED, not a failed upgrade; changed = a SURPRISE needing explanation, not a confirmation | **CLOSED 2026-08-30 — SECOND CHECK TAKEN, AND IT CANNOT ANSWER THE QUESTION: the leak this row measured was REMOVED mid-interval by our own fix (R-344, agent 0.130.0). Question moot; the fix is confirmed holding at 12 days** | elapsed time only | **Two dated checks, both CC:** **+24 h — 2026-08-19 ~10:00Z** and **+7 d — 2026-08-25 ~10:00Z**. Command (the incident's own positive observable): `ssh root@ 'PID=$(systemctl show proxmox-backup-proxy -p MainPID --value); ls /proc/$PID/fd \| wc -l; ss -lnt "( sport = :8007 )"; ss -tn state all "( sport = :8007 )" \| awk "NR>1{print \$1}" \| sort \| uniq -c'`. **Record the ESTAB/CLOSE-WAIT split, not just the total** — the split is what says which leak it is. If PID ≠ 551655 the window is void: something restarted the proxy and the count began again **FIRST CHECK — TAKEN 2026-08-20 08:02:13Z, and it was taken at +46.2 h, NOT +24 h.** No session ran between 18 and 20 August, so the 2026-08-19 date passed as pure elapsed time; the delay is stated here rather than backfilled. **The longer window is a BETTER measurement, not a degraded one** — the leaked count is **388** against the 4 and 5 descriptors the original answer rested on, roughly **97x**, so the uncertainty falls from about +/-50% to about +/-5%. **Precondition PASSED:** PID still **551655**, `ps -o lstart` 2026-08-18 09:51:04, `NRestarts=0` on both units. **Result:** fd **17 -> 405** over **166,251 s** = **201.6 fd/day**, Poisson +/-10.2/day (1s), 2s band **181.2-222.1**. Pre-registered range was **370-450** (confirm-band 330-490); **observed 388**, near the centre. **VERDICT: unchanged — the EXPECTED result, and not a failed upgrade.** **Composition:** ESTAB 0 -> 388, **CLOSE-WAIT 0 — absent from the histogram entirely**, so 100% of the growth is established connections and `CLOSE-WAIT` is not merely the minority half. Runway from fd 405 at 201.6/day to the 65536 ceiling: **~323 days (~2027-07-09)**. Full working: `audits/SPIKE-ep0-established-connections-2026-08-20.md` + `evidence-ep0-established-connections-2026-08-20/step1-slope-computation.txt`. **MEASUREMENT TRAP for the 2026-08-25 check, found on this run — see R-346:** the anchor must be `ps -o lstart= -p $MainPID`, NOT `systemctl show -p ActiveEnterTimestamp`, which reads 03:54:54Z for this generation (the upgrade re-exec'd the proxy; systemd never saw a stop, `NRestarts` is still 0) and would put the rate ~15% low. **PERTURBATION NOTE for the 2026-08-25 check:** this spike's Phase C quietens `pvestatd` on demo-hp for one overnight window inside that 7-day interval. Quantified so nobody reads the shortfall as a change in the leak: even if the leak were fully proportional to the request rate (which Parts 1-2 predict it is NOT), a 10 h half-rate window inside 168 h shifts the 7-day slope by **~3%**, far inside the +/-10% Poisson band on a ~1,400-descriptor count. **The +7 d reading remains usable.** **SECOND CHECK — TAKEN 2026-08-30 16:40Z, five days late, and its own premise did not survive the interval.** Evidence: `audits/evidence-r341-plus7d-2026-08-30/step1-fd-and-sockets.txt`. **Precondition PASSED and this is what makes the reading interpretable at all:** PID still **551655**, `ps -o lstart=` still **2026-08-18 09:51:04**, `NRestarts=0` — the SAME proxy generation as t0, so nothing restarted and re-based the count. **Result: fd = 17.** Not 17 more; **seventeen total** — exactly the documented baseline, against **405** at the first check. The socket histogram contains **one LISTEN and nothing else**: ESTAB **0**, CLOSE-WAIT **0**. **THE VERDICT IS "UNANSWERABLE", NOT "THE UPGRADE FIXED IT", AND THE DIFFERENCE MATTERS.** This row asks whether the PBS 4.2.5-1 upgrade changed the fd slope. Inside this interval **we removed the leak ourselves** — R-344 / agent 0.130.0 fixed the agent's unclosed keep-alive transport, and both boxes were confirmed on 0.130.0 on 2026-08-30. A slope of ~0 after that is a measurement of OUR fix, not of the upgrade, and reading it as "the upgrade worked" would credit a changelog that was read in advance and found to contain no such mechanism. The perturbation quantified in advance for this window was Phase C at ~3%; **the actual perturbation was the removal of the entire phenomenon**, which no pre-registration anticipated. **The question is now MOOT rather than open** — there is no leak left whose slope could differ. **What this check DOES establish, and it is worth more than the original question:** twelve days after the R-344 fix, on the same proxy generation and with no restart to hide behind, ep0 sits at its baseline with **zero** established connections. The 388-descriptor accumulation has not returned. R-336's runway concern (~323 days to the 65536 ceiling) is retired with it; what remains of R-336 is the REQUEST-RATE scaling item, which this check does not touch. | CC on both dates | | **R-342** | **The ep0 snapshot covers less than it looks like it covers, and the next person will assume otherwise.** Quoting `audits/evidence-ep0-pbs-upgrade-2026-08-18/stop2-snapshot.txt` verbatim: *"covers — the 38 GB system disk /dev/sda (root), i.e. the PBS packages, unit files, /etc/systemd drop-ins, nftables and wg config. DOES NOT — /mnt/pbs-datastore. That is /dev/sdb, a separate 100 GB VOLUME, and Hetzner server snapshots do not include attached volumes. The backup data is therefore NOT protected by this snapshot."* Snapshot **421440873** (`felhom-hetzner-20260818`, 15.06 GB, Available) was taken as the rollback for the 4.2.2→4.2.5 PBS upgrade. **Rolling it back restores software state, not the datastore.** That was *acceptable for that change* — a package install writes no datastore content — and the file says so. **The problem is what happens next:** this fact lives in an evidence file nobody will open again, and a snapshot named as "the rollback" reads as protecting everything on the box. ep0 holds the only off-premises copy of a real customer's data | **READY (S) — NEW 2026-08-18** | — | **Decide the safeguard for any future ep0 procedure that could touch `/mnt/pbs-datastore` — it does not exist and has not been designed.** Candidates: a Hetzner **Volume** snapshot (a different object from the server snapshot), a PBS-level sync to a second location, or an explicit written acceptance that the datastore is unprotected for the duration. **Nothing may be added to a runbook implying a safeguard exists until one does** | **Viktor decides; CC executes** — a risk-to-customer-data question | | **R-343** | **The managed controller floor was raised 0.214.0 → 0.216.0 — and it was NOT the no-op it was expected to be: it moved a live box nine seconds later.** Raised by the operator 2026-08-18 **12:36:58Z**, in a **separate save after** the artifact vouch. **Read back from the store, not the form** (`hub_settings.min_controller_version`, `GetGlobalMinControllerVersion` — `hub/internal/store/store.go:1751`): `min_controller_version = 0.216.0`, updated 12:36:58. **WHY IT HAD BEEN BEHIND — this was NOT drift, and describing it as "two releases behind" without this context reads as a defect it was not.** `publish-train-rules.md` rule 1 is *manifest before floor*, and rule 2 requires the floor field to be filled **LAST, in a separate save**, because the DB row overrides the env floor and **acts immediately on the next report cycle**. That rule was earned: on the 2026-07-11 publish train the floor was saved together with the manifest, acted at once, and pushed controller 0.113.0 onto Peti's box **~9 minutes ahead of agent 0.81.0** — the exact forbidden skew, benign only because that box had no NAS shares. R-120's row records the same deliberate choice (*"Floor untouched per publish-train rule 2"*). So the floor sitting at 0.214.0 was **policy being followed**, not neglect. **What it was functionally while it sat there:** not a live problem — every reporting box was at or above it — but a **safety net set two versions low**. The floor is what drags a box forward if it ever falls behind (restored from an old backup, reinstalled, long offline), and at 0.214.0 it would have pulled such a box only to two versions back, missing R-328's severity fix and R-335's follow-up. **THE MEASURED BLAST RADIUS — five reads, and the third and fifth are the findings.** **(1) Floor:** `0.216.0` @ 12:36:58Z, from the store. **(2) Per-customer overrides:** **zero** — all five `customer_configs` rows carry an empty `min_controller_version`, so nothing hides behind a lower override and the global applies to everyone. **(3) Every box's controller version:** `demo-felhom` **0.216.0**, `demo-hp` **0.216.0** — both AT the floor **now**; `drill-r50` **0.213.0** (status `blocked`, last report 2026-08-12, powered off/reverted) and `peti-felhom` **0.115.0** (host row DELETED) are **below** it but are **not reporting boxes**. **(4) Directives/holds:** no `managed floor HELD` line exists; the hub logged `[INFO] Global controller-version floor set to "0.216.0"`. **(5) THE FINDING — a controller DID auto-update after the raise.** `demo-felhom` had been on **0.214.0** since 2026-08-12 16:44 and the hub recorded `controller_updated — Controller frissítve: 0.214.0 → 0.216.0` at **12:37:07Z**, then `controller_started (0.216.0)` at 12:37:12Z — **nine seconds after the raise**, exactly the immediate action rule 2 documents. `demo-hp` was already on 0.216.0 (hand-deployed 2026-08-14 08:31) and did not move. **No error, warning or critical event followed** — the update completed and the controller came back up. **So the change was real, not inert: "every reporting box is at or above the floor" is true BECAUSE of the raise, not independently of it.** **Why it is safe by construction**, cited rather than asserted: `ResolveManagedFloor` (`hub/internal/store/store.go:2068`) sets `Held` and clears the floor entirely when `Floor > GoldenVersion` (the R-216 shape) — floor 0.216.0 **equals** golden 0.216.0, so that guard does not trip — and holds per-box when the box's agent is below the manifest's `MinAgent`, or unknown, or unparseable; both boxes report agent 0.129.0 against `MinAgent` 0.129.0, so the floor was served rather than held. That second guard remains armed for any box reporting with an old agent. **No ISO rebuild is required:** the golden is fetched at first boot from the hub's manifest, which is 0.216.0 — at the floor, not below it — and rule 5's `assert_golden_ge_floor` is a **build-time** gate for *future* builds (`scripts/iso/build-felhom-iso.sh:77`, called at **:267**) which **fails open with a warning** when its inputs are absent (`:78-82`: `if [[ -z "$golden" \| \| -z "$floor" ]]` → `log_warn "… UNENFORCED …"` → `return 0`), both confirmed in the script. **`peti-felhom` was NOT contacted and needs no contact** — from the PETI row: its host row was deleted 2026-07-15 and *"a report from a deleted host 401s and is not persisted"*, so it cannot receive a floor directive at all and the raise cannot reach it | **OPEN — NEW 2026-08-18.** Deliberately NOT closed: the task's closing condition was *all five reads clean, no directive served*, and read 5 shows a live box updated. It went cleanly and is the floor working as designed — but a change recorded as a no-op when it moved a customer box is exactly the kind of record that misleads later | — | **Confirm the 0.216.0 update on `demo-felhom` is healthy in normal operation** (it reported and restarted clean, but it has not yet run a full backup cycle on 0.216.0 at the time of writing), then close. Separately: `drill-r50` at 0.213.0 will be dragged to 0.216.0 by this floor if it is ever booted and reports — that is the floor doing its job, noted so it is not read as a surprise | CC | @@ -595,5 +595,4 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` the R-row. Duplicating them here would create the second source this design avoids. --> | item | due (UTC) | what to measure | |---|---|---| -| R-341 | 2026-08-25 | same, +7 d. Anchor the elapsed time on `ps -o lstart= -p $MainPID`, NOT `ActiveEnterTimestamp` (R-346). A Phase-C perturbation sits inside this interval and is quantified in the R-341 row | diff --git a/hub/CHANGELOG.md b/hub/CHANGELOG.md index 3fe237dc..9f402980 100644 --- a/hub/CHANGELOG.md +++ b/hub/CHANGELOG.md @@ -1,5 +1,15 @@ ## v0.109.0 — the Backup card told every operator that every customer had no backups (2026-08-30, R-331) +> **This push used `git push --no-verify`, and that is declared here rather than worked around.** +> `golden_currency_gate.py` was CONVICTED: controller v0.224.0 and v0.225.0 are released and the newest +> golden bake carries 0.223.0, so a machine installed right now receives neither. **A bypass, not a +> waiver** — the gate offers a waiver only for a release that *deliberately needs no golden*, and these +> need one. The operator ruled bypass-now-bake-later, on the ground that neither fix bites a day-0 box +> (R-330 is a nightly false alarm about apps a new box has not installed; R-331 is a hub display over +> backups a new box has not taken) and both arrive by self-update. **A golden carrying 0.225.0 is +> OWED** — tracked on R-242 in `documentation/backlog/OPEN-ITEMS.md`, which also records that this +> ground does not extend to a release changing first-boot behaviour. + ### R-331 — `Snapshots 0` over a repository holding 67 The customer page's **Backup** card read `Snapshots 0 · Repo Size 0 MB · Integrity Unknown` for