Files
felhom.eu/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step6-slope-computation.txt
T
admin 3e50902a98
gates / gates (push) Failing after 14s
RUNBOOK ep0: PBS 4.2.2-1 -> 4.2.5-1, slope unchanged as predicted (R-341)
Both STOPs cleared by the operator. No code changed; documentation only.

STEP 3 (the run's primary deliverable): the full changelog range 4.2.2-1 ->
4.2.5-1 was read (128 lines, all three entries) and swept for
connection-handling vocabulary. Exactly one keyword hit, a false positive
("S3 ... honor the node's proxy settings" = HTTP proxy config for S3, not the
PBS proxy daemon). 4.2.5-1 is a manifest-hardening security release; 4.2.4-1
is S3 rate limits and a locking cache; 4.2.3-1 is UI/LDAP/tape. NOTHING
addresses descriptor lifetime or connection reaping. Recommendation was: do
not upgrade for this reason.

STOP 1: operator ruled to upgrade anyway for rehearsal value. Recorded as a
practice run, not a fix -- and the interpretation was fixed IN WRITING BEFORE
any numbers existed (stop1-ruling.txt): unchanged = expected; changed =
surprise. Neither outcome could then be rationalised into a success.

STOP 2: Hetzner snapshot 421440873, Available. Documented that it covers
/dev/sda ONLY -- /mnt/pbs-datastore is a separate Volume and is NOT in it, so
it is a software rollback and not a backup of the backup data.

UPGRADE: simulated first (0 to remove), then installed 09:51:00->09:51:06Z,
exit 0. Verified: 4.2.5-1 installed, both daemons active, effective open
files still 65536 (the drop-in survived the new package), Recv-Q 0, loopback
200, 200 from BOTH boxes over the tunnel with felhom-pbs active, and the hub
gauge refreshed post-upgrade at 11:59:31.

SLOPE: before +4 fd/1885 s = 183/day; after +5 fd/1919 s = 225/day. NOT
distinguishable -- one descriptor apart, Poisson +/-2 on such counts. The
higher after-figure is noise, not a regression and not an improvement. 30
minutes cannot settle it; R-341 files the +24 h and +7 d checks.

CORRECTIONS to this morning's own report, both published rather than quietly
fixed:
  - the "~85/day, ~2 years of runway" figures were WRONG. They came from a
    single 17-minute window with a delta of ONE descriptor. Real rate is
    183-200/day over two independent windows; runway ~357 days, not 2 years.
  - the leak was attributed to CLOSE-WAIT. It is mostly ESTAB: CLOSE-WAIT held
    flat at 1 while ESTAB grew 45->49, and at the wedge it was 1011 ESTAB vs
    543 CLOSE-WAIT. R-336's fix must target unreaped connections.
  - "proxmox-backup-api" reported inactive during verification; that unit does
    not exist. Bad query, not a fault, written down because it looked like one.

R-336 stays open: even a fixed leak would not make ~85k requests/day to a
weekly-write DR endpoint correct.

golden-currency still convicts (inherited R-334, controller 0.216.0 vs golden
0.214.0, untouched by this run), so this push is --no-verify per
.claude/rules/gates.md.
2026-08-18 12:26:53 +02:00

21 lines
1013 B
Plaintext

AFTER-SLOPE — proxy PID 551655, started 2026-08-18 09:51:04 UTC (the upgrade restart)
09:51:22Z fd=17 (ESTAB 0) -> 10:23:21Z fd=22 (ESTAB 5)
interval 0.5331 h (1919 s), delta 5 fd
= 9.38 fd/hour = 225.1 fd/DAY
BEFORE (31 min window): 183.3/day from delta=4 over 1885 s
AFTER (32 min window): 225.1/day from delta=5 over 1919 s
VERDICT: NOT DISTINGUISHABLE. The two windows differ by ONE descriptor.
With counts this small the Poisson uncertainty on n=4 is +/-2 and on n=5 is +/-2.2,
so both are consistent with a single unchanged underlying rate. The after-figure being
numerically HIGHER is noise, not a regression -- and it is certainly not an improvement.
This is the EXPECTED result: the 4.2.2->4.2.5 changelog contains no mechanism by which
connection reaping would change. Recorded before the numbers existed, in stop1-ruling.txt.
Composition again favours ESTAB: 0 -> 5 ESTAB, 0 -> 1 CLOSE-WAIT.
30 MINUTES CANNOT SETTLE THIS. The honest checks are +24 h and +7 d -> filed as R-341.