RUNBOOK ep0: PBS 4.2.2-1 -> 4.2.5-1, slope unchanged as predicted (R-341)
gates / gates (push) Failing after 14s

Both STOPs cleared by the operator. No code changed; documentation only.

STEP 3 (the run's primary deliverable): the full changelog range 4.2.2-1 ->
4.2.5-1 was read (128 lines, all three entries) and swept for
connection-handling vocabulary. Exactly one keyword hit, a false positive
("S3 ... honor the node's proxy settings" = HTTP proxy config for S3, not the
PBS proxy daemon). 4.2.5-1 is a manifest-hardening security release; 4.2.4-1
is S3 rate limits and a locking cache; 4.2.3-1 is UI/LDAP/tape. NOTHING
addresses descriptor lifetime or connection reaping. Recommendation was: do
not upgrade for this reason.

STOP 1: operator ruled to upgrade anyway for rehearsal value. Recorded as a
practice run, not a fix -- and the interpretation was fixed IN WRITING BEFORE
any numbers existed (stop1-ruling.txt): unchanged = expected; changed =
surprise. Neither outcome could then be rationalised into a success.

STOP 2: Hetzner snapshot 421440873, Available. Documented that it covers
/dev/sda ONLY -- /mnt/pbs-datastore is a separate Volume and is NOT in it, so
it is a software rollback and not a backup of the backup data.

UPGRADE: simulated first (0 to remove), then installed 09:51:00->09:51:06Z,
exit 0. Verified: 4.2.5-1 installed, both daemons active, effective open
files still 65536 (the drop-in survived the new package), Recv-Q 0, loopback
200, 200 from BOTH boxes over the tunnel with felhom-pbs active, and the hub
gauge refreshed post-upgrade at 11:59:31.

SLOPE: before +4 fd/1885 s = 183/day; after +5 fd/1919 s = 225/day. NOT
distinguishable -- one descriptor apart, Poisson +/-2 on such counts. The
higher after-figure is noise, not a regression and not an improvement. 30
minutes cannot settle it; R-341 files the +24 h and +7 d checks.

CORRECTIONS to this morning's own report, both published rather than quietly
fixed:
  - the "~85/day, ~2 years of runway" figures were WRONG. They came from a
    single 17-minute window with a delta of ONE descriptor. Real rate is
    183-200/day over two independent windows; runway ~357 days, not 2 years.
  - the leak was attributed to CLOSE-WAIT. It is mostly ESTAB: CLOSE-WAIT held
    flat at 1 while ESTAB grew 45->49, and at the wedge it was 1011 ESTAB vs
    543 CLOSE-WAIT. R-336's fix must target unreaped connections.
  - "proxmox-backup-api" reported inactive during verification; that unit does
    not exist. Bad query, not a fault, written down because it looked like one.

R-336 stays open: even a fixed leak would not make ~85k requests/day to a
weekly-write DR endpoint correct.

golden-currency still convicts (inherited R-334, controller 0.216.0 vs golden
0.214.0, untouched by this run), so this push is --no-verify per
.claude/rules/gates.md.
This commit is contained in:
2026-08-18 12:26:53 +02:00
parent 435e044cf1
commit 3e50902a98
22 changed files with 1026 additions and 113 deletions
@@ -168,8 +168,133 @@ A healthy proxy sits near 20 fds with `Recv-Q 0`. **A rising fd count between re
leak is still live** — an unchanging one after a poll-rate reduction would confirm the fix.
**It is still live, and it was measured rather than assumed.** At **16 m 43 s** after the restart the
proxy held **19 fds** (from 18) with **1** connection in `CLOSE-WAIT`. One descriptor per ~17 minutes
is **≈ 85/day** — which lands on the historical rate implied by the failure itself, 1016 sockets over
14 days ≈ **73/day**. Two independent estimates of the same slope agreeing is what makes this a
measurement instead of a story, and it puts the next ceiling at roughly **2 years** rather than the
fortnight the old limit gave. **The leak is unfixed; only its period changed.**
proxy held **19 fds** (from 18) with **1** connection in `CLOSE-WAIT`.
> ### ⚠ CORRECTED 2026-08-18 09:50 UTC — the two numbers first published here were wrong
>
> This section originally read *"one descriptor per ~17 minutes is ≈ 85/day … it puts the next
> ceiling at roughly 2 years"*. **Both figures are wrong, and the reason is worth keeping:** they were
> extrapolated from a single 17-minute window whose delta was **one descriptor**. A sample of one
> cannot carry a daily rate, and the agreement with the historical ~73/day that made it feel solid was
> coincidence.
>
> Re-measured on the same proxy generation (PID 542065, unrestarted) during
> `RUNBOOK ep0 PBS upgrade`, two independent windows:
>
> | window | from | to | rate |
> |---|---|---|---|
> | 31 min | 09:18:21Z fd=62 | 09:49:46Z fd=66 | **183/day** |
> | 5.64 h | 04:11:36Z fd=19 | 09:49:46Z fd=66 | **200/day** |
>
> **The real rate is ~185–200/day — about 2.6× what was published — and the runway to the 65536
> ceiling is ~357 days, not two years.** Still an enormous improvement on the fortnight the old limit
> gave, but under a year, so it is a deadline rather than a comfort.
>
> **And the mechanism named above is the minority one.** Across that window `CLOSE-WAIT` held flat at
> **1** while `ESTAB` grew **45 → 49**: *all* the growth was established connections. At the wedge the
> split was **1011 ESTAB / 543 CLOSE-WAIT**, so ESTAB was the larger half there too and this document
> put its emphasis on the wrong one. **R-336's fix must target connections the proxy never reaps, not
> only sockets left in `CLOSE-WAIT`.**
**The leak is unfixed; only its period changed.**
---
# Follow-up — 2026-08-18, later the same morning: the PBS upgrade
Run under `RUNBOOK — ep0: read the PBS changelog, then decide whether to upgrade`. Two supervised
STOPs, both cleared by the operator. Evidence: `evidence-ep0-pbs-upgrade-2026-08-18/`.
## The changelog said nothing relevant — and that was the finding
The runbook's first job was a pure read: does anything between the installed **4.2.2-1** and the
candidate **4.2.5-1** fix connection handling? All three intervening entries were read in full (128
lines) and swept for `connection|file descriptor|fd|accept(|close_wait|keep-alive|socket|EMFILE|
nofile|leak|proxy|listen|backlog|hyper|tokio`.
**Exactly one keyword hit, and it is a false positive:**
> `* S3: config: allow editing the use-node-config flag that controls whether requests S3 endpoints
> honor the node's proxy settings or not`
That is HTTP-proxy configuration *for S3 requests*, not the `proxmox-backup-proxy` daemon. What the
range actually contains: **4.2.5-1** is a security release hardening client-supplied backup manifests
(an archive name in a crafted manifest could make a sync job read or write outside the snapshot
directory as the `backup` user; manifests are now held in memory and only persisted at backup finish
with per-archive checksum verification), plus a sync/push chunk-reuse fix. **4.2.4-1** is S3 rate
limits, a file-locking user-lookup cache, docs. **4.2.3-1** is UI/journal work, an LDAP search-filter
escape, tape and timezone fixes.
**Nothing addresses descriptor lifetime or connection reaping.** The recommendation was therefore
**do not upgrade for this reason**.
## What was done anyway, and why that is fine
**Operator ruling at STOP 1: upgrade regardless, for rehearsal value** — *"see how that works for us,
we need practice with that too"*. So this was executed as **a practice run of the upgrade procedure on
a Tier-2 protected machine, not as a fix for the leak**, and the distinction was written into
`stop1-ruling.txt` *before* any numbers existed, precisely so the outcome could not be rationalised
afterwards in either direction.
STOP 2 cleared with Hetzner snapshot **421440873** `felhom-hetzner-20260818`, 15.06 GB, status
**Available**. **That snapshot covers `/dev/sda` only.** `/mnt/pbs-datastore` is `/dev/sdb`, a separate
100 GB Volume, and Hetzner server snapshots exclude attached volumes — so it is a rollback for the
software state and **not** a backup of the backup data. Acceptable here because a package install
writes no datastore content; it must not be remembered as datastore protection.
## Result
`apt-get install --only-upgrade proxmox-backup-server`, 09:51:00→09:51:06Z, exit 0. A `-s` simulation
was run first and reported **0 to remove**, so the runbook's abort condition never triggered. Upgraded
server/client/docs to **4.2.5-1**, plus one genuinely new dependency, `proxmox-enterprise-support-
keyring 1.1` — which the 4.2.4-1 changelog had declared, a small but real consistency check between
what was read and what apt did.
| check | result |
|---|---|
| installed (`dpkg -l`) | server / client / docs **4.2.5-1** |
| daemons | `proxmox-backup-proxy` **active running**, `proxmox-backup` **active running** |
| proxy restarted | PID 542065 → **551655** @ 09:51:04 |
| **effective `open files`** | **65536 / 65536** — the drop-in survived the new package |
| `Recv-Q` | 0 |
| loopback | `200` in 12 ms |
| from `felhom-pve` | `200` in 0.103 s, `felhom-pbs active` |
| from `demo-hp` | `200` in 0.096 s, `felhom-pbs active` |
| hub gauge, post-upgrade | `11:59:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)` |
`proxmox-backup-manager version` now reads `4.2.5-1 running version: 4.2.5`. **This independently
settles the confusion recorded in the original incident**: the string was never reporting a stale
daemon, and now that installed and running genuinely match, both halves agree.
**One false alarm, mine:** `systemctl is-active proxmox-backup-api` returned `inactive`. That unit
does not exist — `systemctl cat` says *"No files found for proxmox-backup-api.service"*. The real pair
is `proxmox-backup-proxy.service` ("API Proxy Server") and `proxmox-backup.service` ("API Server"),
both active. A bad query, not a fault, and it is written down because it looked exactly like a fault
for as long as it took to check.
## The slope: unchanged, as predicted
| | window | delta | rate |
|---|---|---|---|
| **before** (PID 542065) | 09:18:21Z fd=62 → 09:49:46Z fd=66 | +4 / 1885 s | **183/day** |
| **after** (PID 551655) | 09:51:22Z fd=17 → 10:23:21Z fd=22 | +5 / 1919 s | **225/day** |
**These are not distinguishable.** The two windows differ by a single descriptor; Poisson uncertainty
on n=4 is ±2 and on n=5 is ±2.2, so both are consistent with one unchanged underlying rate. **The
after-figure being numerically higher is noise, not a regression — and emphatically not an
improvement.** This is the expected outcome and it matches the changelog: no mechanism, no change.
Composition after the upgrade repeats the pattern that matters: **ESTAB 0 → 5, CLOSE-WAIT 0 → 1.**
**Thirty minutes cannot settle this**, in either direction, and this document does not claim it does.
The honest checks are **+24 h (2026-08-19 ~10:00Z)** and **+7 d (2026-08-25 ~10:00Z)** against the
new `t0` of **fd=17 at 09:51:22Z, PID 551655** — filed as **R-341**.
## What this run did not do
Did not reduce the poll rate (**R-336 stays open** — an upgrade that fixed the leak still would not
make ~85,000 requests/day to a weekly-write DR endpoint correct), did not touch the `LimitNOFILE`
drop-ins, did not run `full-upgrade` or touch the kernel, did not trigger a backup/restore/verify to
"prove" the endpoint, and did not change anything on either customer box. Nothing was provisioned, so
there is nothing to tear down; no `.deb` was downloaded, as `apt-get changelog` served the text
directly.