RUNBOOK ep0: PBS 4.2.2-1 -> 4.2.5-1, slope unchanged as predicted (R-341)
gates / gates (push) Failing after 14s
gates / gates (push) Failing after 14s
Both STOPs cleared by the operator. No code changed; documentation only.
STEP 3 (the run's primary deliverable): the full changelog range 4.2.2-1 ->
4.2.5-1 was read (128 lines, all three entries) and swept for
connection-handling vocabulary. Exactly one keyword hit, a false positive
("S3 ... honor the node's proxy settings" = HTTP proxy config for S3, not the
PBS proxy daemon). 4.2.5-1 is a manifest-hardening security release; 4.2.4-1
is S3 rate limits and a locking cache; 4.2.3-1 is UI/LDAP/tape. NOTHING
addresses descriptor lifetime or connection reaping. Recommendation was: do
not upgrade for this reason.
STOP 1: operator ruled to upgrade anyway for rehearsal value. Recorded as a
practice run, not a fix -- and the interpretation was fixed IN WRITING BEFORE
any numbers existed (stop1-ruling.txt): unchanged = expected; changed =
surprise. Neither outcome could then be rationalised into a success.
STOP 2: Hetzner snapshot 421440873, Available. Documented that it covers
/dev/sda ONLY -- /mnt/pbs-datastore is a separate Volume and is NOT in it, so
it is a software rollback and not a backup of the backup data.
UPGRADE: simulated first (0 to remove), then installed 09:51:00->09:51:06Z,
exit 0. Verified: 4.2.5-1 installed, both daemons active, effective open
files still 65536 (the drop-in survived the new package), Recv-Q 0, loopback
200, 200 from BOTH boxes over the tunnel with felhom-pbs active, and the hub
gauge refreshed post-upgrade at 11:59:31.
SLOPE: before +4 fd/1885 s = 183/day; after +5 fd/1919 s = 225/day. NOT
distinguishable -- one descriptor apart, Poisson +/-2 on such counts. The
higher after-figure is noise, not a regression and not an improvement. 30
minutes cannot settle it; R-341 files the +24 h and +7 d checks.
CORRECTIONS to this morning's own report, both published rather than quietly
fixed:
- the "~85/day, ~2 years of runway" figures were WRONG. They came from a
single 17-minute window with a delta of ONE descriptor. Real rate is
183-200/day over two independent windows; runway ~357 days, not 2 years.
- the leak was attributed to CLOSE-WAIT. It is mostly ESTAB: CLOSE-WAIT held
flat at 1 while ESTAB grew 45->49, and at the wedge it was 1011 ESTAB vs
543 CLOSE-WAIT. R-336's fix must target unreaped connections.
- "proxmox-backup-api" reported inactive during verification; that unit does
not exist. Bad query, not a fault, written down because it looked like one.
R-336 stays open: even a fixed leak would not make ~85k requests/day to a
weekly-write DR endpoint correct.
golden-currency still convicts (inherited R-334, controller 0.216.0 vs golden
0.214.0, untouched by this run), so this push is --no-verify per
.claude/rules/gates.md.
This commit is contained in:
@@ -168,8 +168,133 @@ A healthy proxy sits near 20 fds with `Recv-Q 0`. **A rising fd count between re
|
||||
leak is still live** — an unchanging one after a poll-rate reduction would confirm the fix.
|
||||
|
||||
**It is still live, and it was measured rather than assumed.** At **16 m 43 s** after the restart the
|
||||
proxy held **19 fds** (from 18) with **1** connection in `CLOSE-WAIT`. One descriptor per ~17 minutes
|
||||
is **≈ 85/day** — which lands on the historical rate implied by the failure itself, 1016 sockets over
|
||||
14 days ≈ **73/day**. Two independent estimates of the same slope agreeing is what makes this a
|
||||
measurement instead of a story, and it puts the next ceiling at roughly **2 years** rather than the
|
||||
fortnight the old limit gave. **The leak is unfixed; only its period changed.**
|
||||
proxy held **19 fds** (from 18) with **1** connection in `CLOSE-WAIT`.
|
||||
|
||||
> ### ⚠ CORRECTED 2026-08-18 09:50 UTC — the two numbers first published here were wrong
|
||||
>
|
||||
> This section originally read *"one descriptor per ~17 minutes is ≈ 85/day … it puts the next
|
||||
> ceiling at roughly 2 years"*. **Both figures are wrong, and the reason is worth keeping:** they were
|
||||
> extrapolated from a single 17-minute window whose delta was **one descriptor**. A sample of one
|
||||
> cannot carry a daily rate, and the agreement with the historical ~73/day that made it feel solid was
|
||||
> coincidence.
|
||||
>
|
||||
> Re-measured on the same proxy generation (PID 542065, unrestarted) during
|
||||
> `RUNBOOK ep0 PBS upgrade`, two independent windows:
|
||||
>
|
||||
> | window | from | to | rate |
|
||||
> |---|---|---|---|
|
||||
> | 31 min | 09:18:21Z fd=62 | 09:49:46Z fd=66 | **183/day** |
|
||||
> | 5.64 h | 04:11:36Z fd=19 | 09:49:46Z fd=66 | **200/day** |
|
||||
>
|
||||
> **The real rate is ~185–200/day — about 2.6× what was published — and the runway to the 65536
|
||||
> ceiling is ~357 days, not two years.** Still an enormous improvement on the fortnight the old limit
|
||||
> gave, but under a year, so it is a deadline rather than a comfort.
|
||||
>
|
||||
> **And the mechanism named above is the minority one.** Across that window `CLOSE-WAIT` held flat at
|
||||
> **1** while `ESTAB` grew **45 → 49**: *all* the growth was established connections. At the wedge the
|
||||
> split was **1011 ESTAB / 543 CLOSE-WAIT**, so ESTAB was the larger half there too and this document
|
||||
> put its emphasis on the wrong one. **R-336's fix must target connections the proxy never reaps, not
|
||||
> only sockets left in `CLOSE-WAIT`.**
|
||||
|
||||
**The leak is unfixed; only its period changed.**
|
||||
|
||||
---
|
||||
|
||||
# Follow-up — 2026-08-18, later the same morning: the PBS upgrade
|
||||
|
||||
Run under `RUNBOOK — ep0: read the PBS changelog, then decide whether to upgrade`. Two supervised
|
||||
STOPs, both cleared by the operator. Evidence: `evidence-ep0-pbs-upgrade-2026-08-18/`.
|
||||
|
||||
## The changelog said nothing relevant — and that was the finding
|
||||
|
||||
The runbook's first job was a pure read: does anything between the installed **4.2.2-1** and the
|
||||
candidate **4.2.5-1** fix connection handling? All three intervening entries were read in full (128
|
||||
lines) and swept for `connection|file descriptor|fd|accept(|close_wait|keep-alive|socket|EMFILE|
|
||||
nofile|leak|proxy|listen|backlog|hyper|tokio`.
|
||||
|
||||
**Exactly one keyword hit, and it is a false positive:**
|
||||
|
||||
> `* S3: config: allow editing the use-node-config flag that controls whether requests S3 endpoints
|
||||
> honor the node's proxy settings or not`
|
||||
|
||||
That is HTTP-proxy configuration *for S3 requests*, not the `proxmox-backup-proxy` daemon. What the
|
||||
range actually contains: **4.2.5-1** is a security release hardening client-supplied backup manifests
|
||||
(an archive name in a crafted manifest could make a sync job read or write outside the snapshot
|
||||
directory as the `backup` user; manifests are now held in memory and only persisted at backup finish
|
||||
with per-archive checksum verification), plus a sync/push chunk-reuse fix. **4.2.4-1** is S3 rate
|
||||
limits, a file-locking user-lookup cache, docs. **4.2.3-1** is UI/journal work, an LDAP search-filter
|
||||
escape, tape and timezone fixes.
|
||||
|
||||
**Nothing addresses descriptor lifetime or connection reaping.** The recommendation was therefore
|
||||
**do not upgrade for this reason**.
|
||||
|
||||
## What was done anyway, and why that is fine
|
||||
|
||||
**Operator ruling at STOP 1: upgrade regardless, for rehearsal value** — *"see how that works for us,
|
||||
we need practice with that too"*. So this was executed as **a practice run of the upgrade procedure on
|
||||
a Tier-2 protected machine, not as a fix for the leak**, and the distinction was written into
|
||||
`stop1-ruling.txt` *before* any numbers existed, precisely so the outcome could not be rationalised
|
||||
afterwards in either direction.
|
||||
|
||||
STOP 2 cleared with Hetzner snapshot **421440873** `felhom-hetzner-20260818`, 15.06 GB, status
|
||||
**Available**. **That snapshot covers `/dev/sda` only.** `/mnt/pbs-datastore` is `/dev/sdb`, a separate
|
||||
100 GB Volume, and Hetzner server snapshots exclude attached volumes — so it is a rollback for the
|
||||
software state and **not** a backup of the backup data. Acceptable here because a package install
|
||||
writes no datastore content; it must not be remembered as datastore protection.
|
||||
|
||||
## Result
|
||||
|
||||
`apt-get install --only-upgrade proxmox-backup-server`, 09:51:00→09:51:06Z, exit 0. A `-s` simulation
|
||||
was run first and reported **0 to remove**, so the runbook's abort condition never triggered. Upgraded
|
||||
server/client/docs to **4.2.5-1**, plus one genuinely new dependency, `proxmox-enterprise-support-
|
||||
keyring 1.1` — which the 4.2.4-1 changelog had declared, a small but real consistency check between
|
||||
what was read and what apt did.
|
||||
|
||||
| check | result |
|
||||
|---|---|
|
||||
| installed (`dpkg -l`) | server / client / docs **4.2.5-1** |
|
||||
| daemons | `proxmox-backup-proxy` **active running**, `proxmox-backup` **active running** |
|
||||
| proxy restarted | PID 542065 → **551655** @ 09:51:04 |
|
||||
| **effective `open files`** | **65536 / 65536** — the drop-in survived the new package |
|
||||
| `Recv-Q` | 0 |
|
||||
| loopback | `200` in 12 ms |
|
||||
| from `felhom-pve` | `200` in 0.103 s, `felhom-pbs active` |
|
||||
| from `demo-hp` | `200` in 0.096 s, `felhom-pbs active` |
|
||||
| hub gauge, post-upgrade | `11:59:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)` |
|
||||
|
||||
`proxmox-backup-manager version` now reads `4.2.5-1 running version: 4.2.5`. **This independently
|
||||
settles the confusion recorded in the original incident**: the string was never reporting a stale
|
||||
daemon, and now that installed and running genuinely match, both halves agree.
|
||||
|
||||
**One false alarm, mine:** `systemctl is-active proxmox-backup-api` returned `inactive`. That unit
|
||||
does not exist — `systemctl cat` says *"No files found for proxmox-backup-api.service"*. The real pair
|
||||
is `proxmox-backup-proxy.service` ("API Proxy Server") and `proxmox-backup.service` ("API Server"),
|
||||
both active. A bad query, not a fault, and it is written down because it looked exactly like a fault
|
||||
for as long as it took to check.
|
||||
|
||||
## The slope: unchanged, as predicted
|
||||
|
||||
| | window | delta | rate |
|
||||
|---|---|---|---|
|
||||
| **before** (PID 542065) | 09:18:21Z fd=62 → 09:49:46Z fd=66 | +4 / 1885 s | **183/day** |
|
||||
| **after** (PID 551655) | 09:51:22Z fd=17 → 10:23:21Z fd=22 | +5 / 1919 s | **225/day** |
|
||||
|
||||
**These are not distinguishable.** The two windows differ by a single descriptor; Poisson uncertainty
|
||||
on n=4 is ±2 and on n=5 is ±2.2, so both are consistent with one unchanged underlying rate. **The
|
||||
after-figure being numerically higher is noise, not a regression — and emphatically not an
|
||||
improvement.** This is the expected outcome and it matches the changelog: no mechanism, no change.
|
||||
|
||||
Composition after the upgrade repeats the pattern that matters: **ESTAB 0 → 5, CLOSE-WAIT 0 → 1.**
|
||||
|
||||
**Thirty minutes cannot settle this**, in either direction, and this document does not claim it does.
|
||||
The honest checks are **+24 h (2026-08-19 ~10:00Z)** and **+7 d (2026-08-25 ~10:00Z)** against the
|
||||
new `t0` of **fd=17 at 09:51:22Z, PID 551655** — filed as **R-341**.
|
||||
|
||||
## What this run did not do
|
||||
|
||||
Did not reduce the poll rate (**R-336 stays open** — an upgrade that fixed the leak still would not
|
||||
make ~85,000 requests/day to a weekly-write DR endpoint correct), did not touch the `LimitNOFILE`
|
||||
drop-ins, did not run `full-upgrade` or touch the kernel, did not trigger a backup/restore/verify to
|
||||
"prove" the endpoint, and did not change anything on either customer box. Nothing was provisioned, so
|
||||
there is nothing to tear down; no `.deb` was downloaded, as `apt-get changelog` served the text
|
||||
directly.
|
||||
|
||||
Reference in New Issue
Block a user