Files
felhom.eu/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step3-changelog-attempt1.txt
T
admin 3e50902a98
gates / gates (push) Failing after 14s
RUNBOOK ep0: PBS 4.2.2-1 -> 4.2.5-1, slope unchanged as predicted (R-341)
Both STOPs cleared by the operator. No code changed; documentation only.

STEP 3 (the run's primary deliverable): the full changelog range 4.2.2-1 ->
4.2.5-1 was read (128 lines, all three entries) and swept for
connection-handling vocabulary. Exactly one keyword hit, a false positive
("S3 ... honor the node's proxy settings" = HTTP proxy config for S3, not the
PBS proxy daemon). 4.2.5-1 is a manifest-hardening security release; 4.2.4-1
is S3 rate limits and a locking cache; 4.2.3-1 is UI/LDAP/tape. NOTHING
addresses descriptor lifetime or connection reaping. Recommendation was: do
not upgrade for this reason.

STOP 1: operator ruled to upgrade anyway for rehearsal value. Recorded as a
practice run, not a fix -- and the interpretation was fixed IN WRITING BEFORE
any numbers existed (stop1-ruling.txt): unchanged = expected; changed =
surprise. Neither outcome could then be rationalised into a success.

STOP 2: Hetzner snapshot 421440873, Available. Documented that it covers
/dev/sda ONLY -- /mnt/pbs-datastore is a separate Volume and is NOT in it, so
it is a software rollback and not a backup of the backup data.

UPGRADE: simulated first (0 to remove), then installed 09:51:00->09:51:06Z,
exit 0. Verified: 4.2.5-1 installed, both daemons active, effective open
files still 65536 (the drop-in survived the new package), Recv-Q 0, loopback
200, 200 from BOTH boxes over the tunnel with felhom-pbs active, and the hub
gauge refreshed post-upgrade at 11:59:31.

SLOPE: before +4 fd/1885 s = 183/day; after +5 fd/1919 s = 225/day. NOT
distinguishable -- one descriptor apart, Poisson +/-2 on such counts. The
higher after-figure is noise, not a regression and not an improvement. 30
minutes cannot settle it; R-341 files the +24 h and +7 d checks.

CORRECTIONS to this morning's own report, both published rather than quietly
fixed:
  - the "~85/day, ~2 years of runway" figures were WRONG. They came from a
    single 17-minute window with a delta of ONE descriptor. Real rate is
    183-200/day over two independent windows; runway ~357 days, not 2 years.
  - the leak was attributed to CLOSE-WAIT. It is mostly ESTAB: CLOSE-WAIT held
    flat at 1 while ESTAB grew 45->49, and at the wedge it was 1011 ESTAB vs
    543 CLOSE-WAIT. R-336's fix must target unreaped connections.
  - "proxmox-backup-api" reported inactive during verification; that unit does
    not exist. Bad query, not a fault, written down because it looked like one.

R-336 stays open: even a fixed leak would not make ~85k requests/day to a
weekly-write DR endpoint correct.

golden-currency still convicts (inherited R-334, controller 0.216.0 vs golden
0.214.0, untouched by this run), so this push is --no-verify per
.claude/rules/gates.md.
2026-08-18 12:26:53 +02:00

201 lines
9.1 KiB
Plaintext

Get:1 https://metadata.cdn.proxmox.com rust-proxmox-backup 4.2.5-1 Changelog [178 kB]
rust-proxmox-backup (4.2.5-1) trixie; urgency=medium
* backup: harden the handling of client supplied backup manifests:
- only accept archive names that are plain file names carrying a server
side type extension. A crafted name in a manifest could previously make
a sync job read or write outside of the snapshot directory, running as
the unprivileged 'backup' user. Reaching this needed a manifest from a
configured sync remote or from a client that already had backup access
to the datastore.
- keep an uploaded manifest in memory and only persist it on backup
finish, checking that every archive it lists was really uploaded during
that session and that the checksums match the ones computed server
side. A client uploading a manifest that references archives it did not
upload now gets an error on finish instead of such a snapshot being
created.
* fix #7878: sync: push: reuse the manifest of a previous snapshot on a
non-encrypting push if the source snapshot was encrypted with a matching
key, restoring chunk reuse and thus avoiding needlessly long sync runs.
* sync: push: keep the sign-only crypt mode of a source archive instead of
reducing it to unencrypted when pushing without server side encryption.
* subscription: reject a subscription key issued for a different
architecture than the host, as arm64 keys carry an explicit marker, so a
wrong key fails fast instead of only erroring during the online check.
* update to proxmox-upgrade-checks 1.1, which accepts the 7.0 kernel, tells
a bookworm backport apart from a trixie build and fixes the dkms check.
* docs: clarify in the backup protocol description that the manifest is
uploaded by the client and only persisted on backup finish.
-- Proxmox Support Team <support@proxmox.com> Wed, 05 Aug 2026 18:25:37 +0200
rust-proxmox-backup (4.2.4-1) trixie; urgency=medium
* docs: document the debug symbol repository
* datastore: fix wrong local path used for S3 bad chunk handling during
garbage collection
* refactor file creation/mode/ownership helpers to proxmox-product-config
crate
* fix #7642: avoid expensive user lookups on file locking by caching the
backup user/group ID
* depend on proxmox-enterprise-support-keyring, and track its version in the
package version API endpoint
* fix #5748: docs: add `catalog.pcat1` format specification
* docs: system requirements: document we recommend local storage
* S3: fix #6841: allow configuring request rate limits by updating to
proxmox-s3-client 1.4.1. these rate limits are split into active and
passive methods, allowing separate handling of POST/PUT/DELETE and GET/HEAD
request limits.
* S3: config: allow editing the use-node-config flag that controls whether
requests S3 endpoints honor the node's proxy settings or not
* sync: push: gracefully handle previous manifest signature mismatches, which
can happen when enabling or disabling push-encryption on an already synced
backup group
-- Proxmox Support Team <support@proxmox.com> Wed, 29 Jul 2026 15:08:15 +0200
rust-proxmox-backup (4.2.3-1) trixie; urgency=medium
* css: remove x-grid-row-loading class, replace it with non-blurry SVG
variant from proxmox-widget-toolkit
* pbs-client: add backoff log throttle, to ensure progress and similar output
appears quickly initially, but does not create overly long logs
* client: report progress during restore
* api: journal: adopt proxmox-syslog-api and stream the output, making the
implementation consistent with the one from Proxmox Datacenter Manager
* ui: enable the structured journal view and per-service logs, including
colored output and filtering capabilities
* ui: always use arrays for 'delete' property, instead of manually converting
* fix #5971: tape: don't warn on custom MAM attribute write failures
* fix #7175: api: time: use timedatectl instead of /etc/timezone
* fix #7187: report: add ethtool output for physical interfaces
* prune jobs: schedule jobs that do not prune anything, but warn during their
execution. such jobs allow testing scheduling options, but make no sense
for production use.
* fix #6691: allow search by comment in datastore content, make search
case-insensitive and correctly reset content view after empty searches
* ui: datastore: disable various action tooltips for actions which cannot be
triggered
* ldap: escape the user-provided user name when using it in the LDAP search
filter that looks up the user DN.
* ldap sync: log which user properties change when synchronizing an existing
user, instead of only reporting that the user was updated.
* api schema/section config: add support for declaring deprecated property
aliases, to allow renaming properties without showing the old name in the
documentation
* rest server: accept deprecated property aliases in JSON request bodies by
rewriting them to the canonical name before verification and dispatch, like
the CLI and query-string handling already do.
* fix #7690: fs: replace_file: close the temporary file before renaming or
unlinking it, fixing the replacement on WORM file systems and avoiding
leftover .fuse_hidden files on FUSE mounts.
* fs: make_tmp_file: append the temporary suffix instead of replacing the
file extension, keeping the original file name intact for easier
debugging of leftover temporary files.
-- Proxmox Support Team <support@proxmox.com> Tue, 14 Jul 2026 12:54:15 +0200
rust-proxmox-backup (4.2.2-1) trixie; urgency=medium
* api: backup: run synchronous chunk-insert operations off the asynchronous
runtime's worker threads. Blocking those threads, most notably during S3
uploads that can wait up to three hours for a chunk lock, could stall the
I/O and timer drivers of the entire runtime and starve other backup
workers.
* client: backup: make the file-based backup more robust against files that
cannot be accessed or that vanish while the backup is running:
- fix #7658: skip a file and log a warning instead of aborting the whole
backup when querying its metadata fails with a permission error,
matching the existing handling of permission errors when opening files.
Such files can also still be excluded explicitly.
- consistently ignore files that disappear during the backup and warn
about them, instead of treating this as a fatal error in some cases.
* tape: backup: fix the command-line group filter, which was passed to the
API under the wrong parameter name and thus had no effect.
* api: do not log a spurious error when listing the files of a snapshot
whose manifest does not exist yet, which is expected while a backup to
that snapshot is still running.
* ui: datastore summary: fix the per-datastore sync and prune job counts,
which were derived from an incorrectly parsed datastore ID; the prune
count in particular was always shown as zero.
* tape: fix two typos in log and informational messages.
-- Proxmox Support Team <support@proxmox.com> Thu, 18 Jun 2026 11:28:25 +0200
rust-proxmox-backup (4.2.1-1) trixie; urgency=medium
* fix #5076: api: support an 'audiences' property on OpenID realms, listing
additional trusted audience values besides the configured client-id.
Improves compatibility with providers that issue tokens with multiple
audiences.
* fix #7562: api/ui: tape: separate the format-media 'load-barcode'
parameter from the existing 'label-text' verification, so the web
interface can load and format empty or previously-unrelated tapes from a
changer slot in one step. The old shared parameter aborted formatting
after a successful load whenever the on-tape label did not match.
* sync: pull: refuse to overwrite a locally encrypted snapshot from an
unencrypted source or one using a different key, and detect content
differences between two unencrypted snapshots that share a backup time.
Previously such mismatches silently triggered a resync that overwrote the
local snapshot.
* datastore: fix tuning option changes not propagating the updated sync
level to the chunk store until the service was restarted.
* datastore: improve the error message when prune cannot acquire a snapshot
lock for deletion, by showing the snapshot directory and lock file paths
instead of an internal debug dump.
* api: backup: fix benchmark, finish-failed, and backup-failed cleanups
leaving an orphaned empty backup group behind on the datastore.
* api: node: tasks status: return the task end time as an optional field
once the task is finished, so the task viewer can render the correct
duration without an extra API call.
* api/ui: node: add a 'location' property to the node config, exposed
through the node options panel.
* subscription: reuse the server ID from an existing subscription info when
multiple candidates are detected, falling back to the first candidate only
when no prior info exists.