RUNBOOK ep0: PBS 4.2.2-1 -> 4.2.5-1, slope unchanged as predicted (R-341)
gates / gates (push) Failing after 14s

Both STOPs cleared by the operator. No code changed; documentation only.

STEP 3 (the run's primary deliverable): the full changelog range 4.2.2-1 ->
4.2.5-1 was read (128 lines, all three entries) and swept for
connection-handling vocabulary. Exactly one keyword hit, a false positive
("S3 ... honor the node's proxy settings" = HTTP proxy config for S3, not the
PBS proxy daemon). 4.2.5-1 is a manifest-hardening security release; 4.2.4-1
is S3 rate limits and a locking cache; 4.2.3-1 is UI/LDAP/tape. NOTHING
addresses descriptor lifetime or connection reaping. Recommendation was: do
not upgrade for this reason.

STOP 1: operator ruled to upgrade anyway for rehearsal value. Recorded as a
practice run, not a fix -- and the interpretation was fixed IN WRITING BEFORE
any numbers existed (stop1-ruling.txt): unchanged = expected; changed =
surprise. Neither outcome could then be rationalised into a success.

STOP 2: Hetzner snapshot 421440873, Available. Documented that it covers
/dev/sda ONLY -- /mnt/pbs-datastore is a separate Volume and is NOT in it, so
it is a software rollback and not a backup of the backup data.

UPGRADE: simulated first (0 to remove), then installed 09:51:00->09:51:06Z,
exit 0. Verified: 4.2.5-1 installed, both daemons active, effective open
files still 65536 (the drop-in survived the new package), Recv-Q 0, loopback
200, 200 from BOTH boxes over the tunnel with felhom-pbs active, and the hub
gauge refreshed post-upgrade at 11:59:31.

SLOPE: before +4 fd/1885 s = 183/day; after +5 fd/1919 s = 225/day. NOT
distinguishable -- one descriptor apart, Poisson +/-2 on such counts. The
higher after-figure is noise, not a regression and not an improvement. 30
minutes cannot settle it; R-341 files the +24 h and +7 d checks.

CORRECTIONS to this morning's own report, both published rather than quietly
fixed:
  - the "~85/day, ~2 years of runway" figures were WRONG. They came from a
    single 17-minute window with a delta of ONE descriptor. Real rate is
    183-200/day over two independent windows; runway ~357 days, not 2 years.
  - the leak was attributed to CLOSE-WAIT. It is mostly ESTAB: CLOSE-WAIT held
    flat at 1 while ESTAB grew 45->49, and at the wedge it was 1011 ESTAB vs
    543 CLOSE-WAIT. R-336's fix must target unreaped connections.
  - "proxmox-backup-api" reported inactive during verification; that unit does
    not exist. Bad query, not a fault, written down because it looked like one.

R-336 stays open: even a fixed leak would not make ~85k requests/day to a
weekly-write DR endpoint correct.

golden-currency still convicts (inherited R-334, controller 0.216.0 vs golden
0.214.0, untouched by this run), so this push is --no-verify per
.claude/rules/gates.md.
This commit is contained in:
2026-08-18 12:26:53 +02:00
parent 435e044cf1
commit 3e50902a98
22 changed files with 1026 additions and 113 deletions
@@ -168,8 +168,133 @@ A healthy proxy sits near 20 fds with `Recv-Q 0`. **A rising fd count between re
leak is still live** — an unchanging one after a poll-rate reduction would confirm the fix.
**It is still live, and it was measured rather than assumed.** At **16 m 43 s** after the restart the
proxy held **19 fds** (from 18) with **1** connection in `CLOSE-WAIT`. One descriptor per ~17 minutes
is **≈ 85/day** — which lands on the historical rate implied by the failure itself, 1016 sockets over
14 days ≈ **73/day**. Two independent estimates of the same slope agreeing is what makes this a
measurement instead of a story, and it puts the next ceiling at roughly **2 years** rather than the
fortnight the old limit gave. **The leak is unfixed; only its period changed.**
proxy held **19 fds** (from 18) with **1** connection in `CLOSE-WAIT`.
> ### ⚠ CORRECTED 2026-08-18 09:50 UTC — the two numbers first published here were wrong
>
> This section originally read *"one descriptor per ~17 minutes is ≈ 85/day … it puts the next
> ceiling at roughly 2 years"*. **Both figures are wrong, and the reason is worth keeping:** they were
> extrapolated from a single 17-minute window whose delta was **one descriptor**. A sample of one
> cannot carry a daily rate, and the agreement with the historical ~73/day that made it feel solid was
> coincidence.
>
> Re-measured on the same proxy generation (PID 542065, unrestarted) during
> `RUNBOOK ep0 PBS upgrade`, two independent windows:
>
> | window | from | to | rate |
> |---|---|---|---|
> | 31 min | 09:18:21Z fd=62 | 09:49:46Z fd=66 | **183/day** |
> | 5.64 h | 04:11:36Z fd=19 | 09:49:46Z fd=66 | **200/day** |
>
> **The real rate is ~185–200/day — about 2.6× what was published — and the runway to the 65536
> ceiling is ~357 days, not two years.** Still an enormous improvement on the fortnight the old limit
> gave, but under a year, so it is a deadline rather than a comfort.
>
> **And the mechanism named above is the minority one.** Across that window `CLOSE-WAIT` held flat at
> **1** while `ESTAB` grew **45 → 49**: *all* the growth was established connections. At the wedge the
> split was **1011 ESTAB / 543 CLOSE-WAIT**, so ESTAB was the larger half there too and this document
> put its emphasis on the wrong one. **R-336's fix must target connections the proxy never reaps, not
> only sockets left in `CLOSE-WAIT`.**
**The leak is unfixed; only its period changed.**
---
# Follow-up — 2026-08-18, later the same morning: the PBS upgrade
Run under `RUNBOOK — ep0: read the PBS changelog, then decide whether to upgrade`. Two supervised
STOPs, both cleared by the operator. Evidence: `evidence-ep0-pbs-upgrade-2026-08-18/`.
## The changelog said nothing relevant — and that was the finding
The runbook's first job was a pure read: does anything between the installed **4.2.2-1** and the
candidate **4.2.5-1** fix connection handling? All three intervening entries were read in full (128
lines) and swept for `connection|file descriptor|fd|accept(|close_wait|keep-alive|socket|EMFILE|
nofile|leak|proxy|listen|backlog|hyper|tokio`.
**Exactly one keyword hit, and it is a false positive:**
> `* S3: config: allow editing the use-node-config flag that controls whether requests S3 endpoints
> honor the node's proxy settings or not`
That is HTTP-proxy configuration *for S3 requests*, not the `proxmox-backup-proxy` daemon. What the
range actually contains: **4.2.5-1** is a security release hardening client-supplied backup manifests
(an archive name in a crafted manifest could make a sync job read or write outside the snapshot
directory as the `backup` user; manifests are now held in memory and only persisted at backup finish
with per-archive checksum verification), plus a sync/push chunk-reuse fix. **4.2.4-1** is S3 rate
limits, a file-locking user-lookup cache, docs. **4.2.3-1** is UI/journal work, an LDAP search-filter
escape, tape and timezone fixes.
**Nothing addresses descriptor lifetime or connection reaping.** The recommendation was therefore
**do not upgrade for this reason**.
## What was done anyway, and why that is fine
**Operator ruling at STOP 1: upgrade regardless, for rehearsal value** — *"see how that works for us,
we need practice with that too"*. So this was executed as **a practice run of the upgrade procedure on
a Tier-2 protected machine, not as a fix for the leak**, and the distinction was written into
`stop1-ruling.txt` *before* any numbers existed, precisely so the outcome could not be rationalised
afterwards in either direction.
STOP 2 cleared with Hetzner snapshot **421440873** `felhom-hetzner-20260818`, 15.06 GB, status
**Available**. **That snapshot covers `/dev/sda` only.** `/mnt/pbs-datastore` is `/dev/sdb`, a separate
100 GB Volume, and Hetzner server snapshots exclude attached volumes — so it is a rollback for the
software state and **not** a backup of the backup data. Acceptable here because a package install
writes no datastore content; it must not be remembered as datastore protection.
## Result
`apt-get install --only-upgrade proxmox-backup-server`, 09:51:00→09:51:06Z, exit 0. A `-s` simulation
was run first and reported **0 to remove**, so the runbook's abort condition never triggered. Upgraded
server/client/docs to **4.2.5-1**, plus one genuinely new dependency, `proxmox-enterprise-support-
keyring 1.1` — which the 4.2.4-1 changelog had declared, a small but real consistency check between
what was read and what apt did.
| check | result |
|---|---|
| installed (`dpkg -l`) | server / client / docs **4.2.5-1** |
| daemons | `proxmox-backup-proxy` **active running**, `proxmox-backup` **active running** |
| proxy restarted | PID 542065 → **551655** @ 09:51:04 |
| **effective `open files`** | **65536 / 65536** — the drop-in survived the new package |
| `Recv-Q` | 0 |
| loopback | `200` in 12 ms |
| from `felhom-pve` | `200` in 0.103 s, `felhom-pbs active` |
| from `demo-hp` | `200` in 0.096 s, `felhom-pbs active` |
| hub gauge, post-upgrade | `11:59:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)` |
`proxmox-backup-manager version` now reads `4.2.5-1 running version: 4.2.5`. **This independently
settles the confusion recorded in the original incident**: the string was never reporting a stale
daemon, and now that installed and running genuinely match, both halves agree.
**One false alarm, mine:** `systemctl is-active proxmox-backup-api` returned `inactive`. That unit
does not exist — `systemctl cat` says *"No files found for proxmox-backup-api.service"*. The real pair
is `proxmox-backup-proxy.service` ("API Proxy Server") and `proxmox-backup.service` ("API Server"),
both active. A bad query, not a fault, and it is written down because it looked exactly like a fault
for as long as it took to check.
## The slope: unchanged, as predicted
| | window | delta | rate |
|---|---|---|---|
| **before** (PID 542065) | 09:18:21Z fd=62 → 09:49:46Z fd=66 | +4 / 1885 s | **183/day** |
| **after** (PID 551655) | 09:51:22Z fd=17 → 10:23:21Z fd=22 | +5 / 1919 s | **225/day** |
**These are not distinguishable.** The two windows differ by a single descriptor; Poisson uncertainty
on n=4 is ±2 and on n=5 is ±2.2, so both are consistent with one unchanged underlying rate. **The
after-figure being numerically higher is noise, not a regression — and emphatically not an
improvement.** This is the expected outcome and it matches the changelog: no mechanism, no change.
Composition after the upgrade repeats the pattern that matters: **ESTAB 0 → 5, CLOSE-WAIT 0 → 1.**
**Thirty minutes cannot settle this**, in either direction, and this document does not claim it does.
The honest checks are **+24 h (2026-08-19 ~10:00Z)** and **+7 d (2026-08-25 ~10:00Z)** against the
new `t0` of **fd=17 at 09:51:22Z, PID 551655** — filed as **R-341**.
## What this run did not do
Did not reduce the poll rate (**R-336 stays open** — an upgrade that fixed the leak still would not
make ~85,000 requests/day to a weekly-write DR endpoint correct), did not touch the `LimitNOFILE`
drop-ins, did not run `full-upgrade` or touch the kernel, did not trigger a backup/restore/verify to
"prove" the endpoint, and did not change anything on either customer box. Nothing was provisioned, so
there is nothing to tear down; no `.deb` was downloaded, as `apt-get changelog` served the text
directly.
@@ -0,0 +1,105 @@
### capture-time
2026-08-18T09:18:21+00:00
### dpkg -l (INSTALLED — the truth)
ii proxmox-backup-client 4.2.2-1 amd64 Proxmox Backup Client tools
ii proxmox-backup-server 4.2.2-1 amd64 Proxmox Backup Server daemon with tools and GUI
### apt-cache policy proxmox-backup-server
proxmox-backup-server:
Installed: 4.2.2-1
Candidate: 4.2.5-1
Version table:
4.2.5-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.2.4-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.2.3-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
*** 4.2.2-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
100 /var/lib/dpkg/status
4.2.1-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.2.0-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.13-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.12-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.11-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.10-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.9-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.8-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.7-2 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.6-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.5-2 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.5-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.4-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.2-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.1-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.0-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.22-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.21-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.20-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.19-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.18-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.17-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.16-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.15-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.14-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.13-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.12-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.11-4 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.11-2 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.10-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.9-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.8-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.7-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.6-2 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
### proxmox-backup-manager version (available THEN running — misleading, see incident)
proxmox-backup-server 4.2.5-1 running version: 4.2.2
### proxy MainPID
542065
### fd count
62
### limits
Max open files 65536 65536 files
### listen queue
State Recv-Q Send-Q Local Address:Port Peer Address:Port
LISTEN 0 1024 *:8007 *:*
### close-wait count
1
### proxy start time
Tue Aug 18 03:54:53 2026
### datastore
Filesystem Size Used Avail Use% Mounted on
/dev/sdb 98G 3.7G 95G 4% /mnt/pbs-datastore
@@ -0,0 +1,15 @@
### loopback probe
2026-08-18T09:18:40+00:00
http=200 time_total=0.011537
### fd types
54 socket:anon
3 anon_inode:anon
1 0
1 /var/log/proxmox-backup/api/auth.log
1 /var/log/proxmox-backup/api/access.log
1 /var/lib/proxmox-backup/rrdb/rrd.journal
1 /mnt/pbs-datastore/.lock
1 /dev/null
### socket states on 8007
45 ESTAB
1 LISTEN
@@ -0,0 +1,16 @@
### step2-second-reading
2026-08-18T09:49:46+00:00
### MainPID
542065
### fd count
66
### close-wait
1
### socket states
49 ESTAB
1 LISTEN
### listen queue
State Recv-Q Send-Q Local Address:Port Peer Address:Port
LISTEN 0 1024 *:8007 *:*
### proxy start
Tue Aug 18 03:54:53 2026
@@ -0,0 +1,18 @@
BEFORE-SLOPE — proxy PID 542065, started 2026-08-18 03:54:53 UTC, never restarted since
Step-2 window (31 min, this run)
09:18:21Z fd=62 -> 09:49:46Z fd=66
interval 0.5236 h (1885 s), delta 4 fd
= 7.64 fd/hour = 183.3 fd/DAY
Long window (incident t0+17m -> now)
04:11:36Z fd=19 -> 09:49:46Z fd=66
interval 5.6361 h (20290 s), delta 47 fd
= 8.34 fd/hour = 200.1 fd/DAY
Runway from fd=66 at 183/day to the 65536 soft limit:
357 days = 0.98 years
Historical rate implied by the outage: 1016 sockets over 14.29 d = 71.1/day
Socket composition: ESTAB 45 -> 49 (+4); CLOSE-WAIT 1 -> 1 (+0).
=> ALL of the growth in this window is ESTABLISHED connections, not CLOSE-WAIT.
@@ -0,0 +1,200 @@
Get:1 https://metadata.cdn.proxmox.com rust-proxmox-backup 4.2.5-1 Changelog [178 kB]
rust-proxmox-backup (4.2.5-1) trixie; urgency=medium
* backup: harden the handling of client supplied backup manifests:
- only accept archive names that are plain file names carrying a server
side type extension. A crafted name in a manifest could previously make
a sync job read or write outside of the snapshot directory, running as
the unprivileged 'backup' user. Reaching this needed a manifest from a
configured sync remote or from a client that already had backup access
to the datastore.
- keep an uploaded manifest in memory and only persist it on backup
finish, checking that every archive it lists was really uploaded during
that session and that the checksums match the ones computed server
side. A client uploading a manifest that references archives it did not
upload now gets an error on finish instead of such a snapshot being
created.
* fix #7878: sync: push: reuse the manifest of a previous snapshot on a
non-encrypting push if the source snapshot was encrypted with a matching
key, restoring chunk reuse and thus avoiding needlessly long sync runs.
* sync: push: keep the sign-only crypt mode of a source archive instead of
reducing it to unencrypted when pushing without server side encryption.
* subscription: reject a subscription key issued for a different
architecture than the host, as arm64 keys carry an explicit marker, so a
wrong key fails fast instead of only erroring during the online check.
* update to proxmox-upgrade-checks 1.1, which accepts the 7.0 kernel, tells
a bookworm backport apart from a trixie build and fixes the dkms check.
* docs: clarify in the backup protocol description that the manifest is
uploaded by the client and only persisted on backup finish.
-- Proxmox Support Team <support@proxmox.com> Wed, 05 Aug 2026 18:25:37 +0200
rust-proxmox-backup (4.2.4-1) trixie; urgency=medium
* docs: document the debug symbol repository
* datastore: fix wrong local path used for S3 bad chunk handling during
garbage collection
* refactor file creation/mode/ownership helpers to proxmox-product-config
crate
* fix #7642: avoid expensive user lookups on file locking by caching the
backup user/group ID
* depend on proxmox-enterprise-support-keyring, and track its version in the
package version API endpoint
* fix #5748: docs: add `catalog.pcat1` format specification
* docs: system requirements: document we recommend local storage
* S3: fix #6841: allow configuring request rate limits by updating to
proxmox-s3-client 1.4.1. these rate limits are split into active and
passive methods, allowing separate handling of POST/PUT/DELETE and GET/HEAD
request limits.
* S3: config: allow editing the use-node-config flag that controls whether
requests S3 endpoints honor the node's proxy settings or not
* sync: push: gracefully handle previous manifest signature mismatches, which
can happen when enabling or disabling push-encryption on an already synced
backup group
-- Proxmox Support Team <support@proxmox.com> Wed, 29 Jul 2026 15:08:15 +0200
rust-proxmox-backup (4.2.3-1) trixie; urgency=medium
* css: remove x-grid-row-loading class, replace it with non-blurry SVG
variant from proxmox-widget-toolkit
* pbs-client: add backoff log throttle, to ensure progress and similar output
appears quickly initially, but does not create overly long logs
* client: report progress during restore
* api: journal: adopt proxmox-syslog-api and stream the output, making the
implementation consistent with the one from Proxmox Datacenter Manager
* ui: enable the structured journal view and per-service logs, including
colored output and filtering capabilities
* ui: always use arrays for 'delete' property, instead of manually converting
* fix #5971: tape: don't warn on custom MAM attribute write failures
* fix #7175: api: time: use timedatectl instead of /etc/timezone
* fix #7187: report: add ethtool output for physical interfaces
* prune jobs: schedule jobs that do not prune anything, but warn during their
execution. such jobs allow testing scheduling options, but make no sense
for production use.
* fix #6691: allow search by comment in datastore content, make search
case-insensitive and correctly reset content view after empty searches
* ui: datastore: disable various action tooltips for actions which cannot be
triggered
* ldap: escape the user-provided user name when using it in the LDAP search
filter that looks up the user DN.
* ldap sync: log which user properties change when synchronizing an existing
user, instead of only reporting that the user was updated.
* api schema/section config: add support for declaring deprecated property
aliases, to allow renaming properties without showing the old name in the
documentation
* rest server: accept deprecated property aliases in JSON request bodies by
rewriting them to the canonical name before verification and dispatch, like
the CLI and query-string handling already do.
* fix #7690: fs: replace_file: close the temporary file before renaming or
unlinking it, fixing the replacement on WORM file systems and avoiding
leftover .fuse_hidden files on FUSE mounts.
* fs: make_tmp_file: append the temporary suffix instead of replacing the
file extension, keeping the original file name intact for easier
debugging of leftover temporary files.
-- Proxmox Support Team <support@proxmox.com> Tue, 14 Jul 2026 12:54:15 +0200
rust-proxmox-backup (4.2.2-1) trixie; urgency=medium
* api: backup: run synchronous chunk-insert operations off the asynchronous
runtime's worker threads. Blocking those threads, most notably during S3
uploads that can wait up to three hours for a chunk lock, could stall the
I/O and timer drivers of the entire runtime and starve other backup
workers.
* client: backup: make the file-based backup more robust against files that
cannot be accessed or that vanish while the backup is running:
- fix #7658: skip a file and log a warning instead of aborting the whole
backup when querying its metadata fails with a permission error,
matching the existing handling of permission errors when opening files.
Such files can also still be excluded explicitly.
- consistently ignore files that disappear during the backup and warn
about them, instead of treating this as a fatal error in some cases.
* tape: backup: fix the command-line group filter, which was passed to the
API under the wrong parameter name and thus had no effect.
* api: do not log a spurious error when listing the files of a snapshot
whose manifest does not exist yet, which is expected while a backup to
that snapshot is still running.
* ui: datastore summary: fix the per-datastore sync and prune job counts,
which were derived from an incorrectly parsed datastore ID; the prune
count in particular was always shown as zero.
* tape: fix two typos in log and informational messages.
-- Proxmox Support Team <support@proxmox.com> Thu, 18 Jun 2026 11:28:25 +0200
rust-proxmox-backup (4.2.1-1) trixie; urgency=medium
* fix #5076: api: support an 'audiences' property on OpenID realms, listing
additional trusted audience values besides the configured client-id.
Improves compatibility with providers that issue tokens with multiple
audiences.
* fix #7562: api/ui: tape: separate the format-media 'load-barcode'
parameter from the existing 'label-text' verification, so the web
interface can load and format empty or previously-unrelated tapes from a
changer slot in one step. The old shared parameter aborted formatting
after a successful load whenever the on-tape label did not match.
* sync: pull: refuse to overwrite a locally encrypted snapshot from an
unencrypted source or one using a different key, and detect content
differences between two unencrypted snapshots that share a backup time.
Previously such mismatches silently triggered a resync that overwrote the
local snapshot.
* datastore: fix tuning option changes not propagating the updated sync
level to the chunk store until the service was restarted.
* datastore: improve the error message when prune cannot acquire a snapshot
lock for deletion, by showing the snapshot directory and lock file paths
instead of an internal debug dump.
* api: backup: fix benchmark, finish-failed, and backup-failed cleanups
leaving an orphaned empty backup group behind on the datastore.
* api: node: tasks status: return the task end time as an optional field
once the task is finished, so the task viewer can render the correct
duration without an extra API call.
* api/ui: node: add a 'location' property to the node config, exposed
through the node options panel.
* subscription: reuse the server ID from an existing subscription info when
multiple candidates are detected, falling back to the first candidate only
when no prior info exists.
@@ -0,0 +1,128 @@
rust-proxmox-backup (4.2.5-1) trixie; urgency=medium
* backup: harden the handling of client supplied backup manifests:
- only accept archive names that are plain file names carrying a server
side type extension. A crafted name in a manifest could previously make
a sync job read or write outside of the snapshot directory, running as
the unprivileged 'backup' user. Reaching this needed a manifest from a
configured sync remote or from a client that already had backup access
to the datastore.
- keep an uploaded manifest in memory and only persist it on backup
finish, checking that every archive it lists was really uploaded during
that session and that the checksums match the ones computed server
side. A client uploading a manifest that references archives it did not
upload now gets an error on finish instead of such a snapshot being
created.
* fix #7878: sync: push: reuse the manifest of a previous snapshot on a
non-encrypting push if the source snapshot was encrypted with a matching
key, restoring chunk reuse and thus avoiding needlessly long sync runs.
* sync: push: keep the sign-only crypt mode of a source archive instead of
reducing it to unencrypted when pushing without server side encryption.
* subscription: reject a subscription key issued for a different
architecture than the host, as arm64 keys carry an explicit marker, so a
wrong key fails fast instead of only erroring during the online check.
* update to proxmox-upgrade-checks 1.1, which accepts the 7.0 kernel, tells
a bookworm backport apart from a trixie build and fixes the dkms check.
* docs: clarify in the backup protocol description that the manifest is
uploaded by the client and only persisted on backup finish.
-- Proxmox Support Team <support@proxmox.com> Wed, 05 Aug 2026 18:25:37 +0200
rust-proxmox-backup (4.2.4-1) trixie; urgency=medium
* docs: document the debug symbol repository
* datastore: fix wrong local path used for S3 bad chunk handling during
garbage collection
* refactor file creation/mode/ownership helpers to proxmox-product-config
crate
* fix #7642: avoid expensive user lookups on file locking by caching the
backup user/group ID
* depend on proxmox-enterprise-support-keyring, and track its version in the
package version API endpoint
* fix #5748: docs: add `catalog.pcat1` format specification
* docs: system requirements: document we recommend local storage
* S3: fix #6841: allow configuring request rate limits by updating to
proxmox-s3-client 1.4.1. these rate limits are split into active and
passive methods, allowing separate handling of POST/PUT/DELETE and GET/HEAD
request limits.
* S3: config: allow editing the use-node-config flag that controls whether
requests S3 endpoints honor the node's proxy settings or not
* sync: push: gracefully handle previous manifest signature mismatches, which
can happen when enabling or disabling push-encryption on an already synced
backup group
-- Proxmox Support Team <support@proxmox.com> Wed, 29 Jul 2026 15:08:15 +0200
rust-proxmox-backup (4.2.3-1) trixie; urgency=medium
* css: remove x-grid-row-loading class, replace it with non-blurry SVG
variant from proxmox-widget-toolkit
* pbs-client: add backoff log throttle, to ensure progress and similar output
appears quickly initially, but does not create overly long logs
* client: report progress during restore
* api: journal: adopt proxmox-syslog-api and stream the output, making the
implementation consistent with the one from Proxmox Datacenter Manager
* ui: enable the structured journal view and per-service logs, including
colored output and filtering capabilities
* ui: always use arrays for 'delete' property, instead of manually converting
* fix #5971: tape: don't warn on custom MAM attribute write failures
* fix #7175: api: time: use timedatectl instead of /etc/timezone
* fix #7187: report: add ethtool output for physical interfaces
* prune jobs: schedule jobs that do not prune anything, but warn during their
execution. such jobs allow testing scheduling options, but make no sense
for production use.
* fix #6691: allow search by comment in datastore content, make search
case-insensitive and correctly reset content view after empty searches
* ui: datastore: disable various action tooltips for actions which cannot be
triggered
* ldap: escape the user-provided user name when using it in the LDAP search
filter that looks up the user DN.
* ldap sync: log which user properties change when synchronizing an existing
user, instead of only reporting that the user was updated.
* api schema/section config: add support for declaring deprecated property
aliases, to allow renaming properties without showing the old name in the
documentation
* rest server: accept deprecated property aliases in JSON request bodies by
rewriting them to the canonical name before verification and dispatch, like
the CLI and query-string handling already do.
* fix #7690: fs: replace_file: close the temporary file before renaming or
unlinking it, fixing the replacement on WORM file systems and avoiding
leftover .fuse_hidden files on FUSE mounts.
* fs: make_tmp_file: append the temporary suffix instead of replacing the
file extension, keeping the original file name intact for easier
debugging of leftover temporary files.
-- Proxmox Support Team <support@proxmox.com> Tue, 14 Jul 2026 12:54:15 +0200
rust-proxmox-backup (4.2.2-1) trixie; urgency=medium
@@ -0,0 +1,30 @@
### apt-get update start
2026-08-18T09:50:43+00:00
Hit:6 http://mirror.hetzner.com/debian/security trixie-security InRelease
Hit:7 http://download.proxmox.com/debian/pbs trixie InRelease
Get:8 http://deb.debian.org/debian trixie-backports InRelease [54.0 kB]
Get:9 http://deb.debian.org/debian-security trixie-security InRelease [43.4 kB]
Get:10 http://mirror.hetzner.com/debian/packages trixie-backports/main amd64 Packages [314 kB]
Get:11 http://deb.debian.org/debian trixie-backports/main Sources [297 kB]
Fetched 858 kB in 0s (6119 kB/s)
Reading package lists...
### SIMULATION (-s): nothing is installed by this
2026-08-18T09:50:44+00:00
Reading package lists...
Building dependency tree...
Reading state information...
The following additional packages will be installed:
proxmox-backup-client proxmox-backup-docs proxmox-enterprise-support-keyring
The following NEW packages will be installed:
proxmox-enterprise-support-keyring
The following packages will be upgraded:
proxmox-backup-client proxmox-backup-docs proxmox-backup-server
3 upgraded, 1 newly installed, 0 to remove and 8 not upgraded.
Inst proxmox-backup-client [4.2.2-1] (4.2.5-1 Proxmox Backup System Debian Repository:stable [amd64])
Inst proxmox-backup-docs [4.2.2-1] (4.2.5-1 Proxmox Backup System Debian Repository:stable [all])
Inst proxmox-enterprise-support-keyring (1.1 Proxmox Backup System Debian Repository:stable [all])
Inst proxmox-backup-server [4.2.2-1] (4.2.5-1 Proxmox Backup System Debian Repository:stable [amd64])
Conf proxmox-backup-client (4.2.5-1 Proxmox Backup System Debian Repository:stable [amd64])
Conf proxmox-backup-docs (4.2.5-1 Proxmox Backup System Debian Repository:stable [all])
Conf proxmox-enterprise-support-keyring (1.1 Proxmox Backup System Debian Repository:stable [all])
Conf proxmox-backup-server (4.2.5-1 Proxmox Backup System Debian Repository:stable [amd64])
@@ -0,0 +1,58 @@
### UPGRADE START
2026-08-18T09:51:00+00:00
Reading package lists...
Building dependency tree...
Reading state information...
The following additional packages will be installed:
proxmox-backup-client proxmox-backup-docs proxmox-enterprise-support-keyring
The following NEW packages will be installed:
proxmox-enterprise-support-keyring
The following packages will be upgraded:
proxmox-backup-client proxmox-backup-docs proxmox-backup-server
3 upgraded, 1 newly installed, 0 to remove and 8 not upgraded.
Need to get 47.8 MB of archives.
After this operation, 1404 kB of additional disk space will be used.
Get:1 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 proxmox-backup-client amd64 4.2.5-1 [3447 kB]
Get:2 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 proxmox-backup-docs all 4.2.5-1 [6351 kB]
Get:3 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 proxmox-enterprise-support-keyring all 1.1 [2988 B]
Get:4 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 proxmox-backup-server amd64 4.2.5-1 [38.0 MB]
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
locale: Cannot set LC_CTYPE to default locale: No such file or directory
locale: Cannot set LC_ALL to default locale: No such file or directory
Fetched 47.8 MB in 1s (39.3 MB/s)
(Reading database ...
(Reading database ... 5%
(Reading database ... 10%
(Reading database ... 15%
(Reading database ... 20%
(Reading database ... 25%
(Reading database ... 30%
(Reading database ... 35%
(Reading database ... 40%
(Reading database ... 45%
(Reading database ... 50%
(Reading database ... 55%
(Reading database ... 60%
(Reading database ... 65%
(Reading database ... 70%
(Reading database ... 75%
(Reading database ... 80%
(Reading database ... 85%
@@ -0,0 +1,48 @@
### verify-time
2026-08-18T09:51:22+00:00
### dpkg -l (INSTALLED)
ii proxmox-backup-client 4.2.5-1 amd64 Proxmox Backup Client tools
ii proxmox-backup-docs 4.2.5-1 all Proxmox Backup Documentation
ii proxmox-backup-server 4.2.5-1 amd64 Proxmox Backup Server daemon with tools and GUI
ii proxmox-enterprise-support-keyring 1.1 all Proxmox enterprise SSH support keyring
### daemons
active
inactive
active
### proxy MainPID (should DIFFER from 542065 = it restarted)
551655
### proxy start time
Tue Aug 18 09:51:04 2026
### EFFECTIVE open files on the RUNNING proxy (must be 65536)
Max open files 65536 65536 files
### drop-ins still present?
/etc/systemd/system/proxmox-backup-proxy.service.d/:
total 16
drwxr-xr-x 2 root root 4096 Aug 18 03:54 .
drwxr-xr-x 26 root root 4096 Jul 27 07:53 ..
-rw-r--r-- 1 root root 44 Jul 27 07:16 10-datastore-mount.conf
-rw-r--r-- 1 root root 337 Aug 18 03:54 20-nofile.conf
/etc/systemd/system/proxmox-backup.service.d/:
total 16
drwxr-xr-x 2 root root 4096 Aug 18 03:55 .
drwxr-xr-x 26 root root 4096 Jul 27 07:53 ..
-rw-r--r-- 1 root root 44 Jul 27 07:16 10-datastore-mount.conf
-rw-r--r-- 1 root root 237 Aug 18 03:55 20-nofile.conf
### fd count (new t0)
17
### listen queue (Recv-Q must be 0)
State Recv-Q Send-Q Local Address:Port Peer Address:Port
LISTEN 0 1024 *:8007 *:*
### socket states
1 LISTEN
### loopback
http=200 time=0.012245
### version string
proxmox-backup-server 4.2.5-1 running version: 4.2.5
### datastore
+================+====================+=========+
| name | path | comment |
+================+====================+=========+
| felhom-offsite | /mnt/pbs-datastore | |
+================+====================+=========+
@@ -0,0 +1,11 @@
### does proxmox-backup-api exist?
No files found for proxmox-backup-api.service.
### actual PBS units
proxmox-backup-banner.service loaded active exited Proxmox Backup Server Login Banner
proxmox-backup-daily-update.service loaded inactive dead Daily Proxmox Backup Server update and maintenance activities
proxmox-backup-proxy.service loaded active running Proxmox Backup API Proxy Server
proxmox-backup.service loaded active running Proxmox Backup API Server
### API server description
MainPID=551642
Description=Proxmox Backup API Server
ActiveState=active
@@ -0,0 +1,5 @@
### felhom-pve
2026-08-18T11:51:48+02:00
http=200 time=0.103326
Name Type Status Total (KiB) Used (KiB) Available (KiB) %
felhom-pbs pbs active 0 0 0 0.00%
@@ -0,0 +1,5 @@
### demo-hp
2026-08-18T11:51:50+02:00
http=200 time=0.095982
Name Type Status Total (KiB) Used (KiB) Available (KiB) %
felhom-pbs pbs active 0 0 0 0.00%
@@ -0,0 +1 @@
POST-UPGRADE REFRESH: 2026/08/18 11:59:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)
@@ -0,0 +1,3 @@
2026/08/18 11:12:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)
2026/08/18 11:28:30 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)
2026/08/18 11:43:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)
@@ -0,0 +1,18 @@
### step6-second-reading (post-upgrade)
2026-08-18T10:23:21+00:00
### MainPID
551655
### proxy start
Tue Aug 18 09:51:04 2026
### fd count
22
### socket states
5 ESTAB
1 LISTEN
### close-wait
1
### listen queue
State Recv-Q Send-Q Local Address:Port Peer Address:Port
LISTEN 0 1024 *:8007 *:*
### limits
Max open files 65536 65536 files
@@ -0,0 +1,20 @@
AFTER-SLOPE — proxy PID 551655, started 2026-08-18 09:51:04 UTC (the upgrade restart)
09:51:22Z fd=17 (ESTAB 0) -> 10:23:21Z fd=22 (ESTAB 5)
interval 0.5331 h (1919 s), delta 5 fd
= 9.38 fd/hour = 225.1 fd/DAY
BEFORE (31 min window): 183.3/day from delta=4 over 1885 s
AFTER (32 min window): 225.1/day from delta=5 over 1919 s
VERDICT: NOT DISTINGUISHABLE. The two windows differ by ONE descriptor.
With counts this small the Poisson uncertainty on n=4 is +/-2 and on n=5 is +/-2.2,
so both are consistent with a single unchanged underlying rate. The after-figure being
numerically HIGHER is noise, not a regression -- and it is certainly not an improvement.
This is the EXPECTED result: the 4.2.2->4.2.5 changelog contains no mechanism by which
connection reaping would change. Recorded before the numbers existed, in stop1-ruling.txt.
Composition again favours ESTAB: 0 -> 5 ESTAB, 0 -> 1 CLOSE-WAIT.
30 MINUTES CANNOT SETTLE THIS. The honest checks are +24 h and +7 d -> filed as R-341.
@@ -0,0 +1,25 @@
STOP 1 — operator ruling
========================
Recorded: 2026-08-18 ~09:25 UTC (11:25 CEST)
Given by: Viktor (operator), in session.
CC's recommendation was: DO NOT UPGRADE for the stated reason.
Basis: the full changelog range 4.2.2-1 -> 4.2.5-1 (128 lines, all three
entries read) contains NOTHING touching connection handling, descriptor
lifetime, accept(), CLOSE-WAIT, keep-alive or the proxy daemon's socket
lifecycle. An explicit keyword sweep returned exactly one hit, and it is a
false positive ("S3 ... honor the node's proxy settings" = HTTP proxy config
for S3 requests, not the PBS proxy daemon).
OPERATOR RULING: PROCEED WITH THE UPGRADE ANYWAY.
Stated reason: rehearsal value -- "see how that works for us, we need
practice with that too". The upgrade is therefore being performed as a
supervised practice run of the upgrade procedure on a Tier-2 protected
machine, NOT as a fix for R-336's leak.
CONSEQUENCE TO CARRY INTO THE REPORT, so it is not later misread:
this upgrade is NOT expected to change the fd slope. If the post-upgrade
slope differs, that is a surprise requiring explanation, not a confirmation
of anything -- the changelog gives no mechanism by which it should improve.
Equally, if the slope is unchanged, that is the EXPECTED result and is not
evidence the upgrade failed.
@@ -0,0 +1,31 @@
STOP 2 — Hetzner snapshot (the rollback)
========================================
Confirmed by the operator in the Hetzner console, 2026-08-18 ~09:27 UTC (11:27 CEST).
Snapshot ID : 421440873
Description : felhom-hetzner-20260818
Image size : 15.06 GB
Status : Available <-- COMPLETE, not merely started
Created : "less than a minute ago" as displayed at confirmation
Server : felhom-hetzner #147604682 (CX33, 167.233.158.164)
Hetzner project: 15217960 (Felhom.eu)
This is the rollback point for the 4.2.2-1 -> 4.2.5-1 upgrade performed in
Step 4. It was taken on a RUNNING server: Hetzner's own console recommends
powering off first for disk consistency, and that was not done -- a
deliberate trade, because powering off ep0 takes the only off-premises copy
offline and the upgrade being guarded is a userspace package install that
touches neither the datastore nor the boot path.
WHAT THIS SNAPSHOT DOES AND DOES NOT COVER:
covers - the 38 GB system disk /dev/sda (root), i.e. the PBS packages,
unit files, /etc/systemd drop-ins, nftables and wg config.
DOES NOT - /mnt/pbs-datastore. That is /dev/sdb, a separate 100 GB
VOLUME, and Hetzner server snapshots do not include attached
volumes. The backup data is therefore NOT protected by this
snapshot.
Consequence: rolling back restores the software state, not the datastore.
This is acceptable for THIS change because the upgrade writes no datastore
content -- but it must not be mistaken for a datastore backup, and any
future runbook step that could touch /mnt/pbs-datastore needs a different
safeguard than this one.
+2 -1
View File
@@ -636,6 +636,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-333** | **Two disk-health questions the deploy raised and did NOT act on.** **(a) The 55/60 °C bands are SPINNING-DISK bands applied to NVMe.** They were adopted unchanged from the operator's Prometheus config so the two systems cannot disagree — a deliberate, stated decision — but **measured on demo-hp 2026-08-14 the healthy Toshiba KXG50PNV1T02 NVMe idles at 53 °C, two degrees below Figyelmeztetés and seven below Hiba**, and NVMe routinely exceeds 60 °C under load with no fault whatever. As it stands a healthy customer NVMe under sustained write can be reported as **Hiba** — the single worst outcome this feature can produce. **(b) The agent runs bare `smartctl -a -j` with no `-n standby`** (`felhom-agent/internal/storage/hostops.go:368`), so every poll WAKES a spun-down drive; going 6h → hourly multiplies that by six. demo-hp is all-flash so the cadence measurement could not reveal it, and it was recorded rather than acted on per the task's own instruction. Mitigating datum from the fixture: the failing drive logged only **3375 load cycles in 60505 hours** (~one per 18h), i.e. that duty cycle barely spins down at all | **READY (S each) — NEW 2026-08-14** | — | (a) split the temperature bands by device class, or drop them for NVMe and rely on `critical_warning`; (b) add `-n standby` to the agent's smartctl invocation (an agent change, so fold it into R-330's session) | Viktor decides (a); CC does (b) |
| **R-334** | **WAIVER + open item: controller v0.215.0 is released and deployed, and NO golden carries it.** Convicted by `golden_currency_gate.py` on the 2026-08-14 push: newest released controller **0.215.0**, newest golden bake **0.214.0** (`documentation/tests/golden-0.214.0-2026-08-12`). **A machine installed right now receives 0.214.0** — i.e. a brand-new box would ship WITHOUT the R-328 severity fix and would keep emailing nobody about a failing disk. The running fleet is unaffected (demo-hp guest 9201 is on 0.215.0 and healthy); this is purely the day-0 install path. **Not baked in this session deliberately:** the task scoped deployment to demo-hp only, and the second half of the fix — vouching the bake in the hub's day-0 artifact manifest — is **operator-password-gated, so CC cannot complete it**; a baked-but-unvouched golden is worse than none. **The push was made with `git push --no-verify` and it is stated here and in the session report**, per `.claude/rules/gates.md` — the gate has no waiver parser, so recording a waiver does not clear it. **STILL OPEN and now one version WIDER, 2026-08-18:** the newest released controller is **0.216.0** (v0.215.0's own follow-up fix, R-335) and the newest bake is still **0.214.0**, so a new install now misses *two* releases. Re-convicted on this date's documentation-only push, which was likewise made with `--no-verify`; the gate reads `felhom-controller/CHANGELOG.md` and `documentation/tests/golden-*`, **neither of which that session touched** — the conviction is inherited, not caused | **READY (S) — NEW 2026-08-14, re-confirmed 2026-08-18** | operator availability for the vouch step | Bake a golden on **0.216.0** per `runbooks/RUNBOOK-manual-build.md` §4.1, then vouch it — a THREE-field change (`golden_version` + `agent_version` + `min_agent`). Until then every NEW install lacks the severity fix | CC bakes; **Viktor vouches** |
| **R-335** | **One physical disk was walked TWICE per run, and the second walk sustained it against itself.** Found on live hardware ~2h after the v0.215.0 deploy, **by noticing the release's own positive observable disagreed with its own persisted artefact**: the hourly check logged *"3 disk(s) evaluated"* while `disk-health-state.json` held **two** records. Cause: demo-hp's `c11-scratch` and `felhom-backup` are the same physical NVMe (`/dev/nvme0n1`) and resolve to the same `diskKey`. **Not cosmetic** — `RunDiskHealthCheck` writes a disk's new record before the next entry reads it, so the SECOND copy consumed the FIRST copy's write as its prior: the disk **sustained against itself and reached Hiba on a FIRST sighting**, defeating truth-table row 6 — the exact rule separating a one-hour benign excursion from a false critical — and would have emitted **two identical events** for one drive. **Latent, not active, on demo-hp** (all three entries healthy, zero counters), but any aliased disk developing a single pending sector would have gone straight to Hiba. **This is the shape standing rule 3 warns about: an absent alarm was not evidence — the two artefacts had to be read AGAINST each other** | **CLOSED — controller v0.216.0, 2026-08-14.** Each `diskKey` is evaluated once per run; both entries stay marked `seen` so neither looks like a disappeared disk, and the card still renders both storage rows (the dedup is about state and alerts, not display). Pinned by `TestDiskCheck_SameDiskTwiceIsEvaluatedOnce`; companion red-proof run and reverted — deleting the guard makes the first sighting emit `Kind:2` (Hiba-from-sectors) at 8 sectors | — | — | CC |
| **R-336** | **The offsite DR endpoint is polled about once per second, and that is what turned a slow leak into an outage.** ep0's PBS proxy served **~85,000 requests/day** — a flat **3,538/hour**, every hour, from two boxes: `74,445 GET /api2/json/admin/datastore` (`libwww-perl`, i.e. PVE's `pvestatd`) and `73,171 GET /admin/datastore/felhom-offsite/status` (`proxmox-backup-client`). Two pollers asking substantially the same question at the same rate. On 2026-08-18 this walked a connection leak in the proxy to its 1024-fd soft limit in **14 days**, wedging the offsite tier for 9½ hours (`audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md`). The `LimitNOFILE=65536` drop-in applied that morning raises the ceiling **but does not fix the leak** — it converts a fortnightly outage into a multi-year one, which is mitigation, not a fix. A DR endpoint that is written to weekly does not need to be asked about every second | **READY (M) — NEW 2026-08-18** | — | Find what polls `felhom-pbs` this hard (PVE storage status is the prime suspect, and its interval is tunable) and cut it; then confirm the fd count between restarts stops climbing — the positive observable, per standing rule 3. **Baseline measured 2026-08-18: 19 fds at 16m43s post-restart, ~85/day, matching the ~73/day implied by the original failure — the leak is confirmed live, not assumed** | CC |
| **R-336** | **The offsite DR endpoint is polled about once per second, and that is what turned a slow leak into an outage.** ep0's PBS proxy served **~85,000 requests/day** — a flat **3,538/hour**, every hour, from two boxes: `74,445 GET /api2/json/admin/datastore` (`libwww-perl`, i.e. PVE's `pvestatd`) and `73,171 GET /admin/datastore/felhom-offsite/status` (`proxmox-backup-client`). Two pollers asking substantially the same question at the same rate. On 2026-08-18 this walked a connection leak in the proxy to its 1024-fd soft limit in **14 days**, wedging the offsite tier for 9½ hours (`audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md`). The `LimitNOFILE=65536` drop-in applied that morning raises the ceiling **but does not fix the leak** — it converts a fortnightly outage into a multi-year one, which is mitigation, not a fix. A DR endpoint that is written to weekly does not need to be asked about every second | **READY (M) — NEW 2026-08-18** | — | Find what polls `felhom-pbs` this hard (PVE storage status is the prime suspect, and its interval is tunable) and cut it; then confirm the fd count between restarts stops climbing — the positive observable, per standing rule 3. **Baseline measured 2026-08-18, and the FIRST measurement published was WRONG.** The initial "~85/day, matching the ~73/day implied by the failure" came from a single 17-minute window whose delta was **one descriptor** — a sample of one cannot carry a daily rate, and the agreement that made it feel solid was coincidence. **Re-measured over two independent windows the same morning: 183/day (31 min) and 200/day (5.6 h)** — ~2.6x the published figure, putting the runway to the 65536 ceiling at **~357 days, not the ~2 years first claimed**. **And the named mechanism is the minority one:** across that window `CLOSE-WAIT` held flat at 1 while `ESTAB` grew 45→49 — *all* the growth was established connections, and at the wedge the split was 1011 ESTAB / 543 CLOSE-WAIT. **The fix must target connections the proxy never reaps, not just `CLOSE-WAIT` sockets.** The PBS 4.2.5-1 upgrade (2026-08-18) did NOT change the slope and was never expected to — see R-341 | CC |
| **R-337** | **`/backup/status` lagged a completed backup by minutes on one box and not the other — and it RESOLVED ITSELF, which is why this is WATCHING and not a defect.** During the R-336 recovery on 2026-08-18, `demo-hp`'s snapshot landed on ep0 at **03:58:43Z** (complete manifest; the host's own task index says `OK`) — yet `GET /backup/status` was **still serving the superseded 03:27:00Z failure at ~04:03Z**, four-plus minutes later. `demo-felhom` showed its new result within ~40 s of completion. **The lag cleared on its own:** demo-hp's 04:07:35Z host report carries `felhom-pbs success=true, 4.29 GB`, and the hub is green for both boxes. **The first draft of this row claimed the success was "still reported as failed" — that was written before the next report arrived and it was wrong; the corrected claim is a several-minute skew between the two boxes, not a stuck value.** It is recorded because a status field that can trail its own artifact by minutes will, during an incident, be read as a second failure — this session nearly did — and because the asymmetry between the two boxes is unexplained | **WATCHING — NEW 2026-08-18** | another observation, ideally during an incident rather than constructed | **Do not open a fix on this as written.** First establish the intended refresh path for `/backup/status` after an out-of-schedule run; only if the skew is not simply collection cadence is there anything to pin. If it is cadence, close this row and say so | CC |
| **R-338** | **`demo-hp` is not on the R-50 island at all, and `operations/nodes.md` states that it is.** The page records both fleet boxes as island-migrated 2026-07-25. True of `felhom-pve`; **false of `demo-hp`**, whose `agent.json` has `listen_addr: 192.168.0.87:8443` — the customer LAN address — and **no `island_bridge`/`island_guest_addr` keys at all**, whose guest 9201 has `net0` only (no `eth1`), and whose `vmbr9` exists with **zero members**. The controller's `controller.yaml` points at the LAN address, so the box works; this is inventory drift, not breakage. **Two costs.** A session trusting the page addresses the wrong endpoint — that happened on 2026-08-18 and the resulting timeout was briefly read as a fault. And the agent's local API is **bound to the customer LAN on this box** rather than to a point-to-point island, which is the exposure R-50 was built to remove — so a documented security property is claimed for a box that does not have it | **READY (S) — NEW 2026-08-18** | — | Decide which is true: migrate `demo-hp` to the island, or correct `nodes.md`. Leaving both is the one option that keeps the doc lying | Viktor decides; CC executes |
| **R-341** | **Does the fd slope change after the PBS 4.2.5 upgrade? — two dated checks, and the answer is expected to be NO.** ep0 was upgraded 4.2.2-1 → 4.2.5-1 on 2026-08-18 09:51Z on the operator's ruling, **for rehearsal value, not as a fix**: the full changelog range was read (128 lines, all three entries) and swept for connection-handling vocabulary, and it contains **no mechanism** by which descriptor reaping would change — the single keyword hit was `S3 … honor the node's proxy settings`, HTTP-proxy config for S3, not the PBS proxy daemon. **The 32-minute post-upgrade window is indistinguishable from the before window** (+5 fd/1919 s = 225/day vs +4 fd/1885 s = 183/day; the two differ by ONE descriptor and Poisson uncertainty on such counts is ±2, so both are consistent with one unchanged rate — the higher after-figure is noise, not a regression). **Thirty minutes cannot settle it in either direction and this row exists so nobody pretends it did.** New `t0` = **fd 17 at 2026-08-18 09:51:22Z, proxy PID 551655**; before-rate to beat = **183–200/day**. **Interpretation fixed in advance** (`evidence-ep0-pbs-upgrade-2026-08-18/stop1-ruling.txt`, written before any numbers existed): unchanged = EXPECTED, not a failed upgrade; changed = a SURPRISE needing explanation, not a confirmation | **WATCHING — NEW 2026-08-18** | elapsed time only | **Two dated checks, both CC:** **+24 h — 2026-08-19 ~10:00Z** and **+7 d — 2026-08-25 ~10:00Z**. Command (the incident's own positive observable): `ssh root@<ep0> 'PID=$(systemctl show proxmox-backup-proxy -p MainPID --value); ls /proc/$PID/fd \| wc -l; ss -lnt "( sport = :8007 )"; ss -tn state all "( sport = :8007 )" \| awk "NR>1{print \$1}" \| sort \| uniq -c'`. **Record the ESTAB/CLOSE-WAIT split, not just the total** — the split is what says which leak it is. If PID ≠ 551655 the window is void: something restarted the proxy and the count began again | CC on both dates |