Files
felhom.eu/REPORT.md
T
admin 3e50902a98
gates / gates (push) Failing after 14s
RUNBOOK ep0: PBS 4.2.2-1 -> 4.2.5-1, slope unchanged as predicted (R-341)
Both STOPs cleared by the operator. No code changed; documentation only.

STEP 3 (the run's primary deliverable): the full changelog range 4.2.2-1 ->
4.2.5-1 was read (128 lines, all three entries) and swept for
connection-handling vocabulary. Exactly one keyword hit, a false positive
("S3 ... honor the node's proxy settings" = HTTP proxy config for S3, not the
PBS proxy daemon). 4.2.5-1 is a manifest-hardening security release; 4.2.4-1
is S3 rate limits and a locking cache; 4.2.3-1 is UI/LDAP/tape. NOTHING
addresses descriptor lifetime or connection reaping. Recommendation was: do
not upgrade for this reason.

STOP 1: operator ruled to upgrade anyway for rehearsal value. Recorded as a
practice run, not a fix -- and the interpretation was fixed IN WRITING BEFORE
any numbers existed (stop1-ruling.txt): unchanged = expected; changed =
surprise. Neither outcome could then be rationalised into a success.

STOP 2: Hetzner snapshot 421440873, Available. Documented that it covers
/dev/sda ONLY -- /mnt/pbs-datastore is a separate Volume and is NOT in it, so
it is a software rollback and not a backup of the backup data.

UPGRADE: simulated first (0 to remove), then installed 09:51:00->09:51:06Z,
exit 0. Verified: 4.2.5-1 installed, both daemons active, effective open
files still 65536 (the drop-in survived the new package), Recv-Q 0, loopback
200, 200 from BOTH boxes over the tunnel with felhom-pbs active, and the hub
gauge refreshed post-upgrade at 11:59:31.

SLOPE: before +4 fd/1885 s = 183/day; after +5 fd/1919 s = 225/day. NOT
distinguishable -- one descriptor apart, Poisson +/-2 on such counts. The
higher after-figure is noise, not a regression and not an improvement. 30
minutes cannot settle it; R-341 files the +24 h and +7 d checks.

CORRECTIONS to this morning's own report, both published rather than quietly
fixed:
  - the "~85/day, ~2 years of runway" figures were WRONG. They came from a
    single 17-minute window with a delta of ONE descriptor. Real rate is
    183-200/day over two independent windows; runway ~357 days, not 2 years.
  - the leak was attributed to CLOSE-WAIT. It is mostly ESTAB: CLOSE-WAIT held
    flat at 1 while ESTAB grew 45->49, and at the wedge it was 1011 ESTAB vs
    543 CLOSE-WAIT. R-336's fix must target unreaped connections.
  - "proxmox-backup-api" reported inactive during verification; that unit does
    not exist. Bad query, not a fault, written down because it looked like one.

R-336 stays open: even a fixed leak would not make ~85k requests/day to a
weekly-write DR endpoint correct.

golden-currency still convicts (inherited R-334, controller 0.216.0 vs golden
0.214.0, untouched by this run), so this push is --no-verify per
.claude/rules/gates.md.
2026-08-18 12:26:53 +02:00

9.9 KiB
Raw Blame History

REPORT — RUNBOOK ep0: read the PBS changelog, then decide whether to upgrade (2026-08-18, midday)

Outcome: changelog read → no connection-handling fix in the range; operator ruled to upgrade anyway for rehearsal value; upgraded 4.2.2-1 → 4.2.5-1 cleanly; the fd slope did not change, which is the predicted result. Two dated checks filed as R-341.

Both STOPs cleared by the operator. No code changed in any repo; documentation/ only. Evidence: documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/.


1. Baselines re-confirmed on the machine

Not from the audit file — from dpkg -l and apt-cache policy, per the runbook.

Installed proxmox-backup-server / -client 4.2.2-1
Candidate 4.2.5-1 (4.2.3-1, 4.2.4-1 also available)
Proxy PID / started 542065, 03:54:53Z, unrestarted since the incident
Effective open files 65536 / 65536 — the morning's drop-in in force
Recv-Q / loopback 0 / 200 in 11 ms
Datastore 3.7 G of 98 G, 4%
felhom.eu main @ start 435e044 — matches the runbook's stated baseline

2. The before-slope — and a correction I owe the morning's report

Two independent windows on the same proxy generation:

window from → to delta rate
31 min 09:18:21Z fd=62 → 09:49:46Z fd=66 +4 183/day
5.64 h 04:11:36Z fd=19 → 09:49:46Z fd=66 +47 200/day

This morning's incident note said ~85/day and "≈2 years of runway". Both were wrong. They were extrapolated from a single 17-minute window whose delta was one descriptor — a sample of one cannot carry a daily rate, and the agreement with the historical ~73/day that made it feel solid was coincidence. The real rate is ~185–200/day, ~2.6× what I published, and the runway is ~357 days, not two years. Corrected in the incident document and in R-336 rather than left standing.

The mechanism I named was also the minority one. CLOSE-WAIT held flat at 1 across the window while ESTAB grew 45 → 49 — all the growth was established connections. At the wedge the split was 1011 ESTAB / 543 CLOSE-WAIT, so ESTAB dominated there too. R-336's fix must target connections the proxy never reaps, not just CLOSE-WAIT sockets.

3. The changelog, verbatim — the run's primary deliverable

All three entries between 4.2.2-1 and 4.2.5-1 read in full (128 lines), then swept for connection|file descriptor|fd|accept(|close_wait|keep-alive|socket|EMFILE|nofile|leak|proxy|listen| backlog|hyper|tokio.

Exactly one hit, and it is a false positive:

* S3: config: allow editing the use-node-config flag that controls whether requests S3 endpoints honor the node's proxy settings or not

HTTP-proxy configuration for S3 requests — not the proxmox-backup-proxy daemon.

What the range does contain. 4.2.5-1 — a security release hardening client-supplied manifests:

* backup: harden the handling of client supplied backup manifests: - only accept archive names that are plain file names carrying a server side type extension. A crafted name in a manifest could previously make a sync job read or write outside of the snapshot directory, running as the unprivileged 'backup' user. - keep an uploaded manifest in memory and only persist it on backup finish, checking that every archive it lists was really uploaded during that session and that the checksums match

plus a sync/push chunk-reuse fix and a subscription-key architecture check. 4.2.4-1 — S3 rate limits, a file-locking user-lookup cache, the new proxmox-enterprise-support-keyring dependency, docs. 4.2.3-1 — UI/journal work, an LDAP search-filter escape, tape and timezone fixes.

Nothing addresses descriptor lifetime or connection reaping. My recommendation was: do not upgrade for this reason.

4. STOP 1 — the ruling

Operator (Viktor) ruled: upgrade anyway, for rehearsal value — "see how that works for us, we need practice with that too". Legitimate and recorded as such: this was a practice run of the upgrade procedure on a Tier-2 protected machine, not a fix for the leak.

The interpretation was fixed in writing before any numbers existed (stop1-ruling.txt): unchanged slope = expected, not a failed upgrade; changed slope = a surprise needing explanation, not a confirmation. That file was written at the ruling, not afterwards, so neither outcome could be rationalised into a success.

5. STOP 2 — the snapshot, and what it does not cover

Snapshot 421440873 felhom-hetzner-20260818, 15.06 GB, status Available (complete, not merely started), server #147604682, project 15217960.

It covers /dev/sda only. /mnt/pbs-datastore is /dev/sdb, a separate 100 GB Volume, and Hetzner server snapshots exclude attached volumes — so this is a rollback for the software state (packages, unit files, the LimitNOFILE drop-ins, nftables, wg) and not a backup of the backup data. Fine for a package install that writes no datastore content; it must not be remembered as datastore protection. Taken on a running server, deliberately: powering off ep0 to guard a userspace package install would take the only off-premises copy offline.

6. The upgrade and its verification

Simulated first (-s): 0 to remove, so the abort condition never triggered. Then apt-get install --only-upgrade -y proxmox-backup-server, 09:51:00→09:51:06Z, exit 0. Upgraded server/client/docs to 4.2.5-1 plus one new dependency, proxmox-enterprise-support-keyring 1.1 — which the 4.2.4-1 changelog had declared, a small real consistency check between what I read and what apt did.

check result
installed server / client / docs 4.2.5-1
daemons proxmox-backup-proxy active running, proxmox-backup active running
proxy restarted 542065 → 551655 @ 09:51:04
effective open files 65536 / 65536 — survived the new package
drop-ins on disk both present, unmodified
Recv-Q 0
loopback 200 in 12 ms
felhom-pve over tunnel 200 in 0.103 s, felhom-pbs active
demo-hp over tunnel 200 in 0.096 s, felhom-pbs active
hub gauge, post-upgrade 11:59:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)

proxmox-backup-manager version now reads 4.2.5-1 running version: 4.2.5. This independently settles the morning's confusion: that string was never reporting a stale daemon, and now that installed and running genuinely match, both halves agree.

One false alarm, mine. I queried systemctl is-active proxmox-backup-api and got inactive. That unit does not exist — systemctl cat returns "No files found for proxmox-backup-api.service". The real pair is proxmox-backup-proxy.service ("API Proxy Server") and proxmox-backup.service ("API Server"), both active. A bad query, not a fault — recorded because for as long as it took to check, it looked exactly like one.

No backup, restore or verify was triggered to "prove" the endpoint, per the runbook: the reads above answer it without mutating a protected datastore.

7. The after-slope — unchanged, as predicted

window delta rate
before (PID 542065) 09:18:21Z fd=62 → 09:49:46Z fd=66 +4 / 1885 s 183/day
after (PID 551655) 09:51:22Z fd=17 → 10:23:21Z fd=22 +5 / 1919 s 225/day

These are not distinguishable. The windows differ by one descriptor; Poisson uncertainty on n=4 is ±2 and on n=5 is ±2.2, so both are consistent with a single unchanged rate. The after-figure being numerically higher is noise, not a regression — and certainly not an improvement. Composition repeats the pattern: ESTAB 0 → 5, CLOSE-WAIT 0 → 1.

Thirty minutes cannot settle this in either direction, and nothing here claims it does.

8. Register

  • R-341 filed, WATCHING — the two dated checks: +24 h (2026-08-19 ~10:00Z) and +7 d (2026-08-25 ~10:00Z), against new t0 fd=17 @ 09:51:22Z, PID 551655, with the exact command and the instruction to record the ESTAB/CLOSE-WAIT split, not just the total — the split is what identifies which leak it is. If the PID has changed, the window is void.
  • R-336 updated and STILL OPEN. Its wrong 85/day baseline is corrected to 183–200/day, the mechanism is re-pointed at ESTAB, and the upgrade's null result is recorded. It stays open on its own merits: an upgrade that had fixed the leak still would not make ~85,000 requests/day to a weekly-write DR endpoint correct.
  • documentation/runbooks/offsite-endpoint.md — no change made, correctly: it states no PBS version literal, only package names, so there was nothing to update (and per docs.md, version literals do not belong in a current-state doc anyway).

9. Teardown

This run provisioned nothing. No .deb was downloaded — apt-get changelog served the text directly, so the runbook's /tmp cleanup was never needed. The Hetzner snapshot is retained as the rollback; deleting it is an operator decision, and it costs €0.018161/GB/month.

10. Observations, not acted on

  • pvesm status reports felhom-pbs with Total/Used/Available all 0 on both boxes while status reads active. Consistent before and after the upgrade, and the hub's own gauge reads the real 3.7 GB / 97.9 GB, so nothing is broken — but the PVE-side numbers are not usable as a capacity signal. Likely the namespace-scoped token lacking datastore-level audit. Not filed; noted here.
  • 8 other packages are held back (8 not upgraded), untouched deliberately — full-upgrade on this machine is forbidden by the runbook and would change the kernel and WireGuard alongside the thing under test.
  • The morning's --no-verify situation is unchanged: golden-currency still convicts on the inherited R-334 (controller 0.216.0 vs golden 0.214.0), untouched by this run.