Files
felhom.eu/documentation/audits/REPORT-ep0-pbs-upgrade-2026-08-18.md
T
admin f5a4fceeeb
gates / gates (push) Successful in 16s
DRILL 2026-08-21: the off-site restore never replays named volumes (R-354..R-365)
Diagnostic only — no code changed, no version bumped, nothing deployed.

The verdict is a mixture. The unit and the off-site snapshot HOLD the data, proven
by identity in both storage classes including two Hungarian accented filenames. The
loss is in the last leg: ReconstituteFromOffsite skips every isUnit placement and the
volume tars live inside the unit, so the off-site full restore has no named-volume
leg at all — while the local restore-from-unit does, and returned the same tar
byte-identical minutes later.

Twelve rows opened, ceiling R-353 -> R-365. Three HIGH:
  R-354 off-site restore never replays volume dumps
  R-355 paperless-ngx's Postgres is dumped under a non-existent stack, so its unit
        has no DB dump, no safety dump is taken, and the customer is told it has none
  R-356 the off-site restore refuses for all 40 no-drive apps saying the running app
        "is not installed", with a remedy those apps make impossible

R-353's instruction (2) is satisfied and annotated: the 40-class DOES reach the
off-site tier. Its instruction (1) stands and is now larger. R-329 confirmed still
live and now the only bad-severity emit fleet-wide.

Evidence: documentation/audits/DRILL-backup-truth-2026-08-21/evidence/
2026-08-21 23:30:27 +02:00

11 KiB
Raw Blame History

REPORT — RUNBOOK ep0: read the PBS changelog, then decide whether to upgrade (2026-08-18, midday)

Outcome: changelog read → no connection-handling fix in the range; operator ruled to upgrade anyway for rehearsal value; upgraded 4.2.2-1 → 4.2.5-1 cleanly; the fd slope did not change, which is the predicted result. Two dated checks filed as R-341.

Both STOPs cleared by the operator. No code changed in any repo; documentation/ only. Evidence: documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/.


1. Baselines re-confirmed on the machine

Not from the audit file — from dpkg -l and apt-cache policy, per the runbook.

Installed proxmox-backup-server / -client 4.2.2-1
Candidate 4.2.5-1 (4.2.3-1, 4.2.4-1 also available)
Proxy PID / started 542065, 03:54:53Z, unrestarted since the incident
Effective open files 65536 / 65536 — the morning's drop-in in force
Recv-Q / loopback 0 / 200 in 11 ms
Datastore 3.7 G of 98 G, 4%
felhom.eu main @ start 435e044 — matches the runbook's stated baseline

2. The before-slope — and a correction I owe the morning's report

Two independent windows on the same proxy generation:

window from → to delta rate
31 min 09:18:21Z fd=62 → 09:49:46Z fd=66 +4 183/day
5.64 h 04:11:36Z fd=19 → 09:49:46Z fd=66 +47 200/day

This morning's incident note said ~85/day and "≈2 years of runway". Both were wrong. They were extrapolated from a single 17-minute window whose delta was one descriptor — a sample of one cannot carry a daily rate, and the agreement with the historical ~73/day that made it feel solid was coincidence. The real rate is ~185–200/day, ~2.6× what I published, and the runway is ~357 days, not two years. Corrected in the incident document and in R-336 rather than left standing.

The mechanism I named was also the minority one. CLOSE-WAIT held flat at 1 across the window while ESTAB grew 45 → 49 — all the growth was established connections. At the wedge the split was 1011 ESTAB / 543 CLOSE-WAIT, so ESTAB dominated there too. R-336's fix must target connections the proxy never reaps, not just CLOSE-WAIT sockets.

3. The changelog, verbatim — the run's primary deliverable

All three entries between 4.2.2-1 and 4.2.5-1 read in full (128 lines), then swept for connection|file descriptor|fd|accept(|close_wait|keep-alive|socket|EMFILE|nofile|leak|proxy|listen| backlog|hyper|tokio.

Exactly one hit, and it is a false positive:

* S3: config: allow editing the use-node-config flag that controls whether requests S3 endpoints honor the node's proxy settings or not

HTTP-proxy configuration for S3 requests — not the proxmox-backup-proxy daemon.

What the range does contain. 4.2.5-1 — a security release hardening client-supplied manifests:

* backup: harden the handling of client supplied backup manifests: - only accept archive names that are plain file names carrying a server side type extension. A crafted name in a manifest could previously make a sync job read or write outside of the snapshot directory, running as the unprivileged 'backup' user. - keep an uploaded manifest in memory and only persist it on backup finish, checking that every archive it lists was really uploaded during that session and that the checksums match

plus a sync/push chunk-reuse fix and a subscription-key architecture check. 4.2.4-1 — S3 rate limits, a file-locking user-lookup cache, the new proxmox-enterprise-support-keyring dependency, docs. 4.2.3-1 — UI/journal work, an LDAP search-filter escape, tape and timezone fixes.

Nothing addresses descriptor lifetime or connection reaping. My recommendation was: do not upgrade for this reason.

4. STOP 1 — the ruling

Operator (Viktor) ruled: upgrade anyway, for rehearsal value — "see how that works for us, we need practice with that too". Legitimate and recorded as such: this was a practice run of the upgrade procedure on a Tier-2 protected machine, not a fix for the leak.

The interpretation was fixed in writing before any numbers existed (stop1-ruling.txt): unchanged slope = expected, not a failed upgrade; changed slope = a surprise needing explanation, not a confirmation. That file was written at the ruling, not afterwards, so neither outcome could be rationalised into a success.

5. STOP 2 — the snapshot, and what it does not cover

Snapshot 421440873 felhom-hetzner-20260818, 15.06 GB, status Available (complete, not merely started), server #147604682, project 15217960.

It covers /dev/sda only. /mnt/pbs-datastore is /dev/sdb, a separate 100 GB Volume, and Hetzner server snapshots exclude attached volumes — so this is a rollback for the software state (packages, unit files, the LimitNOFILE drop-ins, nftables, wg) and not a backup of the backup data. Fine for a package install that writes no datastore content; it must not be remembered as datastore protection. Taken on a running server, deliberately: powering off ep0 to guard a userspace package install would take the only off-premises copy offline.

6. The upgrade and its verification

Simulated first (-s): 0 to remove, so the abort condition never triggered. Then apt-get install --only-upgrade -y proxmox-backup-server, 09:51:00→09:51:06Z, exit 0. Upgraded server/client/docs to 4.2.5-1 plus one new dependency, proxmox-enterprise-support-keyring 1.1 — which the 4.2.4-1 changelog had declared, a small real consistency check between what I read and what apt did.

check result
installed server / client / docs 4.2.5-1
daemons proxmox-backup-proxy active running, proxmox-backup active running
proxy restarted 542065 → 551655 @ 09:51:04
effective open files 65536 / 65536 — survived the new package
drop-ins on disk both present, unmodified
Recv-Q 0
loopback 200 in 12 ms
felhom-pve over tunnel 200 in 0.103 s, felhom-pbs active
demo-hp over tunnel 200 in 0.096 s, felhom-pbs active
hub gauge, post-upgrade 11:59:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)

proxmox-backup-manager version now reads 4.2.5-1 running version: 4.2.5. This independently settles the morning's confusion: that string was never reporting a stale daemon, and now that installed and running genuinely match, both halves agree.

One false alarm, mine. I queried systemctl is-active proxmox-backup-api and got inactive. That unit does not exist — systemctl cat returns "No files found for proxmox-backup-api.service". The real pair is proxmox-backup-proxy.service ("API Proxy Server") and proxmox-backup.service ("API Server"), both active. A bad query, not a fault — recorded because for as long as it took to check, it looked exactly like one.

No backup, restore or verify was triggered to "prove" the endpoint, per the runbook: the reads above answer it without mutating a protected datastore.

7. The after-slope — unchanged, as predicted

window delta rate
before (PID 542065) 09:18:21Z fd=62 → 09:49:46Z fd=66 +4 / 1885 s 183/day
after (PID 551655) 09:51:22Z fd=17 → 10:23:21Z fd=22 +5 / 1919 s 225/day

These are not distinguishable. The windows differ by one descriptor; Poisson uncertainty on n=4 is ±2 and on n=5 is ±2.2, so both are consistent with a single unchanged rate. The after-figure being numerically higher is noise, not a regression — and certainly not an improvement. Composition repeats the pattern: ESTAB 0 → 5, CLOSE-WAIT 0 → 1.

Thirty minutes cannot settle this in either direction, and nothing here claims it does.

8. Register

  • R-341 filed, WATCHING — the two dated checks: +24 h (2026-08-19 ~10:00Z) and +7 d (2026-08-25 ~10:00Z), against new t0 fd=17 @ 09:51:22Z, PID 551655, with the exact command and the instruction to record the ESTAB/CLOSE-WAIT split, not just the total — the split is what identifies which leak it is. If the PID has changed, the window is void.
  • R-336 updated and STILL OPEN. Its wrong 85/day baseline is corrected to 183–200/day, the mechanism is re-pointed at ESTAB, and the upgrade's null result is recorded. It stays open on its own merits: an upgrade that had fixed the leak still would not make ~85,000 requests/day to a weekly-write DR endpoint correct.
  • documentation/runbooks/offsite-endpoint.md — no change made, correctly: it states no PBS version literal, only package names, so there was nothing to update (and per docs.md, version literals do not belong in a current-state doc anyway).

9. Teardown

This run provisioned nothing. No .deb was downloaded — apt-get changelog served the text directly, so the runbook's /tmp cleanup was never needed. The Hetzner snapshot is retained as the rollback; deleting it is an operator decision, and it costs €0.018161/GB/month.

10. Observations, not acted on

  • pvesm status reports felhom-pbs with Total/Used/Available all 0 on both boxes while status reads active. Consistent before and after the upgrade, and the hub's own gauge reads the real 3.7 GB / 97.9 GB, so nothing is broken — but the PVE-side numbers are not usable as a capacity signal. Likely the namespace-scoped token lacking datastore-level audit. Not filed; noted here.
  • 8 other packages are held back (8 not upgraded), untouched deliberately — full-upgrade on this machine is forbidden by the runbook and would change the kernel and WireGuard alongside the thing under test.
  • The morning's --no-verify situation is unchanged: golden-currency still convicts on the inherited R-334 (controller 0.216.0 vs golden 0.214.0), untouched by this run.

11. CI, checked by run ID

Run 351, head_sha 3e50902a9, conclusion failure, elapsed 15 s (10:27:00→10:27:15Z). Inherited, not caused: CI's only step is python3 scripts/repo_gates.py --fast, which convicts golden-currency on R-334 (controller 0.216.0 vs newest golden 0.214.0) — files this run did not touch. 15 s is the workflow's own honest-failure band, not the R-265 reap band. The previous run 350 on 435e044 failed identically, before this run began.

Expect one [felhom CI] gates FAILED mail for run 351. Same cause as run 348 this morning; not a new fault, and not related to the upgrade.

python3 scripts/unproven.py --summary — unchanged: 23 walked, 32 not walked of 55. No number moved, correctly: this run proved an operational fact, not a product claim.