Both STOPs cleared by the operator. No code changed; documentation only.
STEP 3 (the run's primary deliverable): the full changelog range 4.2.2-1 ->
4.2.5-1 was read (128 lines, all three entries) and swept for
connection-handling vocabulary. Exactly one keyword hit, a false positive
("S3 ... honor the node's proxy settings" = HTTP proxy config for S3, not the
PBS proxy daemon). 4.2.5-1 is a manifest-hardening security release; 4.2.4-1
is S3 rate limits and a locking cache; 4.2.3-1 is UI/LDAP/tape. NOTHING
addresses descriptor lifetime or connection reaping. Recommendation was: do
not upgrade for this reason.
STOP 1: operator ruled to upgrade anyway for rehearsal value. Recorded as a
practice run, not a fix -- and the interpretation was fixed IN WRITING BEFORE
any numbers existed (stop1-ruling.txt): unchanged = expected; changed =
surprise. Neither outcome could then be rationalised into a success.
STOP 2: Hetzner snapshot 421440873, Available. Documented that it covers
/dev/sda ONLY -- /mnt/pbs-datastore is a separate Volume and is NOT in it, so
it is a software rollback and not a backup of the backup data.
UPGRADE: simulated first (0 to remove), then installed 09:51:00->09:51:06Z,
exit 0. Verified: 4.2.5-1 installed, both daemons active, effective open
files still 65536 (the drop-in survived the new package), Recv-Q 0, loopback
200, 200 from BOTH boxes over the tunnel with felhom-pbs active, and the hub
gauge refreshed post-upgrade at 11:59:31.
SLOPE: before +4 fd/1885 s = 183/day; after +5 fd/1919 s = 225/day. NOT
distinguishable -- one descriptor apart, Poisson +/-2 on such counts. The
higher after-figure is noise, not a regression and not an improvement. 30
minutes cannot settle it; R-341 files the +24 h and +7 d checks.
CORRECTIONS to this morning's own report, both published rather than quietly
fixed:
- the "~85/day, ~2 years of runway" figures were WRONG. They came from a
single 17-minute window with a delta of ONE descriptor. Real rate is
183-200/day over two independent windows; runway ~357 days, not 2 years.
- the leak was attributed to CLOSE-WAIT. It is mostly ESTAB: CLOSE-WAIT held
flat at 1 while ESTAB grew 45->49, and at the wedge it was 1011 ESTAB vs
543 CLOSE-WAIT. R-336's fix must target unreaped connections.
- "proxmox-backup-api" reported inactive during verification; that unit does
not exist. Bad query, not a fault, written down because it looked like one.
R-336 stays open: even a fixed leak would not make ~85k requests/day to a
weekly-write DR endpoint correct.
golden-currency still convicts (inherited R-334, controller 0.216.0 vs golden
0.214.0, untouched by this run), so this push is --no-verify per
.claude/rules/gates.md.
9.9 KiB
REPORT — RUNBOOK ep0: read the PBS changelog, then decide whether to upgrade (2026-08-18, midday)
Outcome: changelog read → no connection-handling fix in the range; operator ruled to upgrade anyway for rehearsal value; upgraded 4.2.2-1 → 4.2.5-1 cleanly; the fd slope did not change, which is the predicted result. Two dated checks filed as R-341.
Both STOPs cleared by the operator. No code changed in any repo; documentation/ only.
Evidence: documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/.
1. Baselines re-confirmed on the machine
Not from the audit file — from dpkg -l and apt-cache policy, per the runbook.
| Installed | proxmox-backup-server / -client 4.2.2-1 |
| Candidate | 4.2.5-1 (4.2.3-1, 4.2.4-1 also available) |
| Proxy PID / started | 542065, 03:54:53Z, unrestarted since the incident |
Effective open files |
65536 / 65536 — the morning's drop-in in force |
Recv-Q / loopback |
0 / 200 in 11 ms |
| Datastore | 3.7 G of 98 G, 4% |
felhom.eu main @ start |
435e044 — matches the runbook's stated baseline |
2. The before-slope — and a correction I owe the morning's report
Two independent windows on the same proxy generation:
| window | from → to | delta | rate |
|---|---|---|---|
| 31 min | 09:18:21Z fd=62 → 09:49:46Z fd=66 | +4 | 183/day |
| 5.64 h | 04:11:36Z fd=19 → 09:49:46Z fd=66 | +47 | 200/day |
This morning's incident note said ~85/day and "≈2 years of runway". Both were wrong. They were extrapolated from a single 17-minute window whose delta was one descriptor — a sample of one cannot carry a daily rate, and the agreement with the historical ~73/day that made it feel solid was coincidence. The real rate is ~185–200/day, ~2.6× what I published, and the runway is ~357 days, not two years. Corrected in the incident document and in R-336 rather than left standing.
The mechanism I named was also the minority one. CLOSE-WAIT held flat at 1 across the
window while ESTAB grew 45 → 49 — all the growth was established connections. At the wedge
the split was 1011 ESTAB / 543 CLOSE-WAIT, so ESTAB dominated there too. R-336's fix must target
connections the proxy never reaps, not just CLOSE-WAIT sockets.
3. The changelog, verbatim — the run's primary deliverable
All three entries between 4.2.2-1 and 4.2.5-1 read in full (128 lines), then swept for
connection|file descriptor|fd|accept(|close_wait|keep-alive|socket|EMFILE|nofile|leak|proxy|listen| backlog|hyper|tokio.
Exactly one hit, and it is a false positive:
* S3: config: allow editing the use-node-config flag that controls whether requests S3 endpointshonor the node's proxy settings or not
HTTP-proxy configuration for S3 requests — not the proxmox-backup-proxy daemon.
What the range does contain. 4.2.5-1 — a security release hardening client-supplied manifests:
* backup: harden the handling of client supplied backup manifests:- only accept archive names that are plain file names carrying a server side type extension. Acrafted name in a manifest could previously make a sync job read or write outside of thesnapshot directory, running as the unprivileged 'backup' user.- keep an uploaded manifest in memory and only persist it on backup finish, checking that everyarchive it lists was really uploaded during that session and that the checksums match
plus a sync/push chunk-reuse fix and a subscription-key architecture check. 4.2.4-1 — S3 rate
limits, a file-locking user-lookup cache, the new proxmox-enterprise-support-keyring dependency,
docs. 4.2.3-1 — UI/journal work, an LDAP search-filter escape, tape and timezone fixes.
Nothing addresses descriptor lifetime or connection reaping. My recommendation was: do not upgrade for this reason.
4. STOP 1 — the ruling
Operator (Viktor) ruled: upgrade anyway, for rehearsal value — "see how that works for us, we need practice with that too". Legitimate and recorded as such: this was a practice run of the upgrade procedure on a Tier-2 protected machine, not a fix for the leak.
The interpretation was fixed in writing before any numbers existed (stop1-ruling.txt):
unchanged slope = expected, not a failed upgrade; changed slope = a surprise needing
explanation, not a confirmation. That file was written at the ruling, not afterwards, so neither
outcome could be rationalised into a success.
5. STOP 2 — the snapshot, and what it does not cover
Snapshot 421440873 felhom-hetzner-20260818, 15.06 GB, status Available (complete, not
merely started), server #147604682, project 15217960.
It covers /dev/sda only. /mnt/pbs-datastore is /dev/sdb, a separate 100 GB Volume, and
Hetzner server snapshots exclude attached volumes — so this is a rollback for the software state
(packages, unit files, the LimitNOFILE drop-ins, nftables, wg) and not a backup of the backup
data. Fine for a package install that writes no datastore content; it must not be remembered as
datastore protection. Taken on a running server, deliberately: powering off ep0 to guard a userspace
package install would take the only off-premises copy offline.
6. The upgrade and its verification
Simulated first (-s): 0 to remove, so the abort condition never triggered. Then
apt-get install --only-upgrade -y proxmox-backup-server, 09:51:00→09:51:06Z, exit 0. Upgraded
server/client/docs to 4.2.5-1 plus one new dependency, proxmox-enterprise-support-keyring 1.1 —
which the 4.2.4-1 changelog had declared, a small real consistency check between what I read and
what apt did.
| check | result |
|---|---|
| installed | server / client / docs 4.2.5-1 |
| daemons | proxmox-backup-proxy active running, proxmox-backup active running |
| proxy restarted | 542065 → 551655 @ 09:51:04 |
effective open files |
65536 / 65536 — survived the new package |
| drop-ins on disk | both present, unmodified |
Recv-Q |
0 |
| loopback | 200 in 12 ms |
felhom-pve over tunnel |
200 in 0.103 s, felhom-pbs active |
demo-hp over tunnel |
200 in 0.096 s, felhom-pbs active |
| hub gauge, post-upgrade | 11:59:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB) |
proxmox-backup-manager version now reads 4.2.5-1 running version: 4.2.5. This independently
settles the morning's confusion: that string was never reporting a stale daemon, and now that
installed and running genuinely match, both halves agree.
One false alarm, mine. I queried systemctl is-active proxmox-backup-api and got inactive.
That unit does not exist — systemctl cat returns "No files found for
proxmox-backup-api.service". The real pair is proxmox-backup-proxy.service ("API Proxy Server")
and proxmox-backup.service ("API Server"), both active. A bad query, not a fault — recorded because
for as long as it took to check, it looked exactly like one.
No backup, restore or verify was triggered to "prove" the endpoint, per the runbook: the reads above answer it without mutating a protected datastore.
7. The after-slope — unchanged, as predicted
| window | delta | rate | |
|---|---|---|---|
| before (PID 542065) | 09:18:21Z fd=62 → 09:49:46Z fd=66 | +4 / 1885 s | 183/day |
| after (PID 551655) | 09:51:22Z fd=17 → 10:23:21Z fd=22 | +5 / 1919 s | 225/day |
These are not distinguishable. The windows differ by one descriptor; Poisson uncertainty on n=4 is ±2 and on n=5 is ±2.2, so both are consistent with a single unchanged rate. The after-figure being numerically higher is noise, not a regression — and certainly not an improvement. Composition repeats the pattern: ESTAB 0 → 5, CLOSE-WAIT 0 → 1.
Thirty minutes cannot settle this in either direction, and nothing here claims it does.
8. Register
- R-341 filed, WATCHING — the two dated checks: +24 h (2026-08-19 ~10:00Z) and
+7 d (2026-08-25 ~10:00Z), against new
t0fd=17 @ 09:51:22Z, PID 551655, with the exact command and the instruction to record the ESTAB/CLOSE-WAIT split, not just the total — the split is what identifies which leak it is. If the PID has changed, the window is void. - R-336 updated and STILL OPEN. Its wrong 85/day baseline is corrected to 183–200/day, the mechanism is re-pointed at ESTAB, and the upgrade's null result is recorded. It stays open on its own merits: an upgrade that had fixed the leak still would not make ~85,000 requests/day to a weekly-write DR endpoint correct.
documentation/runbooks/offsite-endpoint.md— no change made, correctly: it states no PBS version literal, only package names, so there was nothing to update (and perdocs.md, version literals do not belong in a current-state doc anyway).
9. Teardown
This run provisioned nothing. No .deb was downloaded — apt-get changelog served the text
directly, so the runbook's /tmp cleanup was never needed. The Hetzner snapshot is retained as
the rollback; deleting it is an operator decision, and it costs €0.018161/GB/month.
10. Observations, not acted on
pvesm statusreportsfelhom-pbswith Total/Used/Available all0on both boxes while status readsactive. Consistent before and after the upgrade, and the hub's own gauge reads the real 3.7 GB / 97.9 GB, so nothing is broken — but the PVE-side numbers are not usable as a capacity signal. Likely the namespace-scoped token lacking datastore-level audit. Not filed; noted here.- 8 other packages are held back (
8 not upgraded), untouched deliberately —full-upgradeon this machine is forbidden by the runbook and would change the kernel and WireGuard alongside the thing under test. - The morning's
--no-verifysituation is unchanged:golden-currencystill convicts on the inherited R-334 (controller 0.216.0 vs golden 0.214.0), untouched by this run.