Files
felhom.eu/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step2-slope-computation.txt
T
admin 3e50902a98
gates / gates (push) Failing after 14s
RUNBOOK ep0: PBS 4.2.2-1 -> 4.2.5-1, slope unchanged as predicted (R-341)
Both STOPs cleared by the operator. No code changed; documentation only.

STEP 3 (the run's primary deliverable): the full changelog range 4.2.2-1 ->
4.2.5-1 was read (128 lines, all three entries) and swept for
connection-handling vocabulary. Exactly one keyword hit, a false positive
("S3 ... honor the node's proxy settings" = HTTP proxy config for S3, not the
PBS proxy daemon). 4.2.5-1 is a manifest-hardening security release; 4.2.4-1
is S3 rate limits and a locking cache; 4.2.3-1 is UI/LDAP/tape. NOTHING
addresses descriptor lifetime or connection reaping. Recommendation was: do
not upgrade for this reason.

STOP 1: operator ruled to upgrade anyway for rehearsal value. Recorded as a
practice run, not a fix -- and the interpretation was fixed IN WRITING BEFORE
any numbers existed (stop1-ruling.txt): unchanged = expected; changed =
surprise. Neither outcome could then be rationalised into a success.

STOP 2: Hetzner snapshot 421440873, Available. Documented that it covers
/dev/sda ONLY -- /mnt/pbs-datastore is a separate Volume and is NOT in it, so
it is a software rollback and not a backup of the backup data.

UPGRADE: simulated first (0 to remove), then installed 09:51:00->09:51:06Z,
exit 0. Verified: 4.2.5-1 installed, both daemons active, effective open
files still 65536 (the drop-in survived the new package), Recv-Q 0, loopback
200, 200 from BOTH boxes over the tunnel with felhom-pbs active, and the hub
gauge refreshed post-upgrade at 11:59:31.

SLOPE: before +4 fd/1885 s = 183/day; after +5 fd/1919 s = 225/day. NOT
distinguishable -- one descriptor apart, Poisson +/-2 on such counts. The
higher after-figure is noise, not a regression and not an improvement. 30
minutes cannot settle it; R-341 files the +24 h and +7 d checks.

CORRECTIONS to this morning's own report, both published rather than quietly
fixed:
  - the "~85/day, ~2 years of runway" figures were WRONG. They came from a
    single 17-minute window with a delta of ONE descriptor. Real rate is
    183-200/day over two independent windows; runway ~357 days, not 2 years.
  - the leak was attributed to CLOSE-WAIT. It is mostly ESTAB: CLOSE-WAIT held
    flat at 1 while ESTAB grew 45->49, and at the wedge it was 1011 ESTAB vs
    543 CLOSE-WAIT. R-336's fix must target unreaped connections.
  - "proxmox-backup-api" reported inactive during verification; that unit does
    not exist. Bad query, not a fault, written down because it looked like one.

R-336 stays open: even a fixed leak would not make ~85k requests/day to a
weekly-write DR endpoint correct.

golden-currency still convicts (inherited R-334, controller 0.216.0 vs golden
0.214.0, untouched by this run), so this push is --no-verify per
.claude/rules/gates.md.
2026-08-18 12:26:53 +02:00

19 lines
694 B
Plaintext

BEFORE-SLOPE — proxy PID 542065, started 2026-08-18 03:54:53 UTC, never restarted since
Step-2 window (31 min, this run)
09:18:21Z fd=62 -> 09:49:46Z fd=66
interval 0.5236 h (1885 s), delta 4 fd
= 7.64 fd/hour = 183.3 fd/DAY
Long window (incident t0+17m -> now)
04:11:36Z fd=19 -> 09:49:46Z fd=66
interval 5.6361 h (20290 s), delta 47 fd
= 8.34 fd/hour = 200.1 fd/DAY
Runway from fd=66 at 183/day to the 65536 soft limit:
357 days = 0.98 years
Historical rate implied by the outage: 1016 sockets over 14.29 d = 71.1/day
Socket composition: ESTAB 45 -> 49 (+4); CLOSE-WAIT 1 -> 1 (+0).
=> ALL of the growth in this window is ESTABLISHED connections, not CLOSE-WAIT.