RUNBOOK ep0: PBS 4.2.2-1 -> 4.2.5-1, slope unchanged as predicted (R-341)
gates / gates (push) Failing after 14s
gates / gates (push) Failing after 14s
Both STOPs cleared by the operator. No code changed; documentation only.
STEP 3 (the run's primary deliverable): the full changelog range 4.2.2-1 ->
4.2.5-1 was read (128 lines, all three entries) and swept for
connection-handling vocabulary. Exactly one keyword hit, a false positive
("S3 ... honor the node's proxy settings" = HTTP proxy config for S3, not the
PBS proxy daemon). 4.2.5-1 is a manifest-hardening security release; 4.2.4-1
is S3 rate limits and a locking cache; 4.2.3-1 is UI/LDAP/tape. NOTHING
addresses descriptor lifetime or connection reaping. Recommendation was: do
not upgrade for this reason.
STOP 1: operator ruled to upgrade anyway for rehearsal value. Recorded as a
practice run, not a fix -- and the interpretation was fixed IN WRITING BEFORE
any numbers existed (stop1-ruling.txt): unchanged = expected; changed =
surprise. Neither outcome could then be rationalised into a success.
STOP 2: Hetzner snapshot 421440873, Available. Documented that it covers
/dev/sda ONLY -- /mnt/pbs-datastore is a separate Volume and is NOT in it, so
it is a software rollback and not a backup of the backup data.
UPGRADE: simulated first (0 to remove), then installed 09:51:00->09:51:06Z,
exit 0. Verified: 4.2.5-1 installed, both daemons active, effective open
files still 65536 (the drop-in survived the new package), Recv-Q 0, loopback
200, 200 from BOTH boxes over the tunnel with felhom-pbs active, and the hub
gauge refreshed post-upgrade at 11:59:31.
SLOPE: before +4 fd/1885 s = 183/day; after +5 fd/1919 s = 225/day. NOT
distinguishable -- one descriptor apart, Poisson +/-2 on such counts. The
higher after-figure is noise, not a regression and not an improvement. 30
minutes cannot settle it; R-341 files the +24 h and +7 d checks.
CORRECTIONS to this morning's own report, both published rather than quietly
fixed:
- the "~85/day, ~2 years of runway" figures were WRONG. They came from a
single 17-minute window with a delta of ONE descriptor. Real rate is
183-200/day over two independent windows; runway ~357 days, not 2 years.
- the leak was attributed to CLOSE-WAIT. It is mostly ESTAB: CLOSE-WAIT held
flat at 1 while ESTAB grew 45->49, and at the wedge it was 1011 ESTAB vs
543 CLOSE-WAIT. R-336's fix must target unreaped connections.
- "proxmox-backup-api" reported inactive during verification; that unit does
not exist. Bad query, not a fault, written down because it looked like one.
R-336 stays open: even a fixed leak would not make ~85k requests/day to a
weekly-write DR endpoint correct.
golden-currency still convicts (inherited R-334, controller 0.216.0 vs golden
0.214.0, untouched by this run), so this push is --no-verify per
.claude/rules/gates.md.
This commit is contained in:
@@ -168,8 +168,133 @@ A healthy proxy sits near 20 fds with `Recv-Q 0`. **A rising fd count between re
|
||||
leak is still live** — an unchanging one after a poll-rate reduction would confirm the fix.
|
||||
|
||||
**It is still live, and it was measured rather than assumed.** At **16 m 43 s** after the restart the
|
||||
proxy held **19 fds** (from 18) with **1** connection in `CLOSE-WAIT`. One descriptor per ~17 minutes
|
||||
is **≈ 85/day** — which lands on the historical rate implied by the failure itself, 1016 sockets over
|
||||
14 days ≈ **73/day**. Two independent estimates of the same slope agreeing is what makes this a
|
||||
measurement instead of a story, and it puts the next ceiling at roughly **2 years** rather than the
|
||||
fortnight the old limit gave. **The leak is unfixed; only its period changed.**
|
||||
proxy held **19 fds** (from 18) with **1** connection in `CLOSE-WAIT`.
|
||||
|
||||
> ### ⚠ CORRECTED 2026-08-18 09:50 UTC — the two numbers first published here were wrong
|
||||
>
|
||||
> This section originally read *"one descriptor per ~17 minutes is ≈ 85/day … it puts the next
|
||||
> ceiling at roughly 2 years"*. **Both figures are wrong, and the reason is worth keeping:** they were
|
||||
> extrapolated from a single 17-minute window whose delta was **one descriptor**. A sample of one
|
||||
> cannot carry a daily rate, and the agreement with the historical ~73/day that made it feel solid was
|
||||
> coincidence.
|
||||
>
|
||||
> Re-measured on the same proxy generation (PID 542065, unrestarted) during
|
||||
> `RUNBOOK ep0 PBS upgrade`, two independent windows:
|
||||
>
|
||||
> | window | from | to | rate |
|
||||
> |---|---|---|---|
|
||||
> | 31 min | 09:18:21Z fd=62 | 09:49:46Z fd=66 | **183/day** |
|
||||
> | 5.64 h | 04:11:36Z fd=19 | 09:49:46Z fd=66 | **200/day** |
|
||||
>
|
||||
> **The real rate is ~185–200/day — about 2.6× what was published — and the runway to the 65536
|
||||
> ceiling is ~357 days, not two years.** Still an enormous improvement on the fortnight the old limit
|
||||
> gave, but under a year, so it is a deadline rather than a comfort.
|
||||
>
|
||||
> **And the mechanism named above is the minority one.** Across that window `CLOSE-WAIT` held flat at
|
||||
> **1** while `ESTAB` grew **45 → 49**: *all* the growth was established connections. At the wedge the
|
||||
> split was **1011 ESTAB / 543 CLOSE-WAIT**, so ESTAB was the larger half there too and this document
|
||||
> put its emphasis on the wrong one. **R-336's fix must target connections the proxy never reaps, not
|
||||
> only sockets left in `CLOSE-WAIT`.**
|
||||
|
||||
**The leak is unfixed; only its period changed.**
|
||||
|
||||
---
|
||||
|
||||
# Follow-up — 2026-08-18, later the same morning: the PBS upgrade
|
||||
|
||||
Run under `RUNBOOK — ep0: read the PBS changelog, then decide whether to upgrade`. Two supervised
|
||||
STOPs, both cleared by the operator. Evidence: `evidence-ep0-pbs-upgrade-2026-08-18/`.
|
||||
|
||||
## The changelog said nothing relevant — and that was the finding
|
||||
|
||||
The runbook's first job was a pure read: does anything between the installed **4.2.2-1** and the
|
||||
candidate **4.2.5-1** fix connection handling? All three intervening entries were read in full (128
|
||||
lines) and swept for `connection|file descriptor|fd|accept(|close_wait|keep-alive|socket|EMFILE|
|
||||
nofile|leak|proxy|listen|backlog|hyper|tokio`.
|
||||
|
||||
**Exactly one keyword hit, and it is a false positive:**
|
||||
|
||||
> `* S3: config: allow editing the use-node-config flag that controls whether requests S3 endpoints
|
||||
> honor the node's proxy settings or not`
|
||||
|
||||
That is HTTP-proxy configuration *for S3 requests*, not the `proxmox-backup-proxy` daemon. What the
|
||||
range actually contains: **4.2.5-1** is a security release hardening client-supplied backup manifests
|
||||
(an archive name in a crafted manifest could make a sync job read or write outside the snapshot
|
||||
directory as the `backup` user; manifests are now held in memory and only persisted at backup finish
|
||||
with per-archive checksum verification), plus a sync/push chunk-reuse fix. **4.2.4-1** is S3 rate
|
||||
limits, a file-locking user-lookup cache, docs. **4.2.3-1** is UI/journal work, an LDAP search-filter
|
||||
escape, tape and timezone fixes.
|
||||
|
||||
**Nothing addresses descriptor lifetime or connection reaping.** The recommendation was therefore
|
||||
**do not upgrade for this reason**.
|
||||
|
||||
## What was done anyway, and why that is fine
|
||||
|
||||
**Operator ruling at STOP 1: upgrade regardless, for rehearsal value** — *"see how that works for us,
|
||||
we need practice with that too"*. So this was executed as **a practice run of the upgrade procedure on
|
||||
a Tier-2 protected machine, not as a fix for the leak**, and the distinction was written into
|
||||
`stop1-ruling.txt` *before* any numbers existed, precisely so the outcome could not be rationalised
|
||||
afterwards in either direction.
|
||||
|
||||
STOP 2 cleared with Hetzner snapshot **421440873** `felhom-hetzner-20260818`, 15.06 GB, status
|
||||
**Available**. **That snapshot covers `/dev/sda` only.** `/mnt/pbs-datastore` is `/dev/sdb`, a separate
|
||||
100 GB Volume, and Hetzner server snapshots exclude attached volumes — so it is a rollback for the
|
||||
software state and **not** a backup of the backup data. Acceptable here because a package install
|
||||
writes no datastore content; it must not be remembered as datastore protection.
|
||||
|
||||
## Result
|
||||
|
||||
`apt-get install --only-upgrade proxmox-backup-server`, 09:51:00→09:51:06Z, exit 0. A `-s` simulation
|
||||
was run first and reported **0 to remove**, so the runbook's abort condition never triggered. Upgraded
|
||||
server/client/docs to **4.2.5-1**, plus one genuinely new dependency, `proxmox-enterprise-support-
|
||||
keyring 1.1` — which the 4.2.4-1 changelog had declared, a small but real consistency check between
|
||||
what was read and what apt did.
|
||||
|
||||
| check | result |
|
||||
|---|---|
|
||||
| installed (`dpkg -l`) | server / client / docs **4.2.5-1** |
|
||||
| daemons | `proxmox-backup-proxy` **active running**, `proxmox-backup` **active running** |
|
||||
| proxy restarted | PID 542065 → **551655** @ 09:51:04 |
|
||||
| **effective `open files`** | **65536 / 65536** — the drop-in survived the new package |
|
||||
| `Recv-Q` | 0 |
|
||||
| loopback | `200` in 12 ms |
|
||||
| from `felhom-pve` | `200` in 0.103 s, `felhom-pbs active` |
|
||||
| from `demo-hp` | `200` in 0.096 s, `felhom-pbs active` |
|
||||
| hub gauge, post-upgrade | `11:59:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)` |
|
||||
|
||||
`proxmox-backup-manager version` now reads `4.2.5-1 running version: 4.2.5`. **This independently
|
||||
settles the confusion recorded in the original incident**: the string was never reporting a stale
|
||||
daemon, and now that installed and running genuinely match, both halves agree.
|
||||
|
||||
**One false alarm, mine:** `systemctl is-active proxmox-backup-api` returned `inactive`. That unit
|
||||
does not exist — `systemctl cat` says *"No files found for proxmox-backup-api.service"*. The real pair
|
||||
is `proxmox-backup-proxy.service` ("API Proxy Server") and `proxmox-backup.service` ("API Server"),
|
||||
both active. A bad query, not a fault, and it is written down because it looked exactly like a fault
|
||||
for as long as it took to check.
|
||||
|
||||
## The slope: unchanged, as predicted
|
||||
|
||||
| | window | delta | rate |
|
||||
|---|---|---|---|
|
||||
| **before** (PID 542065) | 09:18:21Z fd=62 → 09:49:46Z fd=66 | +4 / 1885 s | **183/day** |
|
||||
| **after** (PID 551655) | 09:51:22Z fd=17 → 10:23:21Z fd=22 | +5 / 1919 s | **225/day** |
|
||||
|
||||
**These are not distinguishable.** The two windows differ by a single descriptor; Poisson uncertainty
|
||||
on n=4 is ±2 and on n=5 is ±2.2, so both are consistent with one unchanged underlying rate. **The
|
||||
after-figure being numerically higher is noise, not a regression — and emphatically not an
|
||||
improvement.** This is the expected outcome and it matches the changelog: no mechanism, no change.
|
||||
|
||||
Composition after the upgrade repeats the pattern that matters: **ESTAB 0 → 5, CLOSE-WAIT 0 → 1.**
|
||||
|
||||
**Thirty minutes cannot settle this**, in either direction, and this document does not claim it does.
|
||||
The honest checks are **+24 h (2026-08-19 ~10:00Z)** and **+7 d (2026-08-25 ~10:00Z)** against the
|
||||
new `t0` of **fd=17 at 09:51:22Z, PID 551655** — filed as **R-341**.
|
||||
|
||||
## What this run did not do
|
||||
|
||||
Did not reduce the poll rate (**R-336 stays open** — an upgrade that fixed the leak still would not
|
||||
make ~85,000 requests/day to a weekly-write DR endpoint correct), did not touch the `LimitNOFILE`
|
||||
drop-ins, did not run `full-upgrade` or touch the kernel, did not trigger a backup/restore/verify to
|
||||
"prove" the endpoint, and did not change anything on either customer box. Nothing was provisioned, so
|
||||
there is nothing to tear down; no `.deb` was downloaded, as `apt-get changelog` served the text
|
||||
directly.
|
||||
|
||||
@@ -0,0 +1,105 @@
|
||||
### capture-time
|
||||
2026-08-18T09:18:21+00:00
|
||||
### dpkg -l (INSTALLED — the truth)
|
||||
ii proxmox-backup-client 4.2.2-1 amd64 Proxmox Backup Client tools
|
||||
ii proxmox-backup-server 4.2.2-1 amd64 Proxmox Backup Server daemon with tools and GUI
|
||||
### apt-cache policy proxmox-backup-server
|
||||
proxmox-backup-server:
|
||||
Installed: 4.2.2-1
|
||||
Candidate: 4.2.5-1
|
||||
Version table:
|
||||
4.2.5-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.2.4-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.2.3-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
*** 4.2.2-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
100 /var/lib/dpkg/status
|
||||
4.2.1-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.2.0-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.1.13-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.1.12-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.1.11-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.1.10-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.1.9-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.1.8-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.1.7-2 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.1.6-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.1.5-2 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.1.5-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.1.4-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.1.2-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.1.1-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.1.0-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.0.22-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.0.21-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.0.20-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.0.19-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.0.18-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.0.17-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.0.16-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.0.15-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.0.14-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.0.13-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.0.12-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.0.11-4 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.0.11-2 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.0.10-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.0.9-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.0.8-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.0.7-1 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
4.0.6-2 500
|
||||
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
|
||||
### proxmox-backup-manager version (available THEN running — misleading, see incident)
|
||||
proxmox-backup-server 4.2.5-1 running version: 4.2.2
|
||||
### proxy MainPID
|
||||
542065
|
||||
### fd count
|
||||
62
|
||||
### limits
|
||||
Max open files 65536 65536 files
|
||||
### listen queue
|
||||
State Recv-Q Send-Q Local Address:Port Peer Address:Port
|
||||
LISTEN 0 1024 *:8007 *:*
|
||||
### close-wait count
|
||||
1
|
||||
### proxy start time
|
||||
Tue Aug 18 03:54:53 2026
|
||||
### datastore
|
||||
Filesystem Size Used Avail Use% Mounted on
|
||||
/dev/sdb 98G 3.7G 95G 4% /mnt/pbs-datastore
|
||||
@@ -0,0 +1,15 @@
|
||||
### loopback probe
|
||||
2026-08-18T09:18:40+00:00
|
||||
http=200 time_total=0.011537
|
||||
### fd types
|
||||
54 socket:anon
|
||||
3 anon_inode:anon
|
||||
1 0
|
||||
1 /var/log/proxmox-backup/api/auth.log
|
||||
1 /var/log/proxmox-backup/api/access.log
|
||||
1 /var/lib/proxmox-backup/rrdb/rrd.journal
|
||||
1 /mnt/pbs-datastore/.lock
|
||||
1 /dev/null
|
||||
### socket states on 8007
|
||||
45 ESTAB
|
||||
1 LISTEN
|
||||
@@ -0,0 +1,16 @@
|
||||
### step2-second-reading
|
||||
2026-08-18T09:49:46+00:00
|
||||
### MainPID
|
||||
542065
|
||||
### fd count
|
||||
66
|
||||
### close-wait
|
||||
1
|
||||
### socket states
|
||||
49 ESTAB
|
||||
1 LISTEN
|
||||
### listen queue
|
||||
State Recv-Q Send-Q Local Address:Port Peer Address:Port
|
||||
LISTEN 0 1024 *:8007 *:*
|
||||
### proxy start
|
||||
Tue Aug 18 03:54:53 2026
|
||||
@@ -0,0 +1,18 @@
|
||||
BEFORE-SLOPE — proxy PID 542065, started 2026-08-18 03:54:53 UTC, never restarted since
|
||||
|
||||
Step-2 window (31 min, this run)
|
||||
09:18:21Z fd=62 -> 09:49:46Z fd=66
|
||||
interval 0.5236 h (1885 s), delta 4 fd
|
||||
= 7.64 fd/hour = 183.3 fd/DAY
|
||||
|
||||
Long window (incident t0+17m -> now)
|
||||
04:11:36Z fd=19 -> 09:49:46Z fd=66
|
||||
interval 5.6361 h (20290 s), delta 47 fd
|
||||
= 8.34 fd/hour = 200.1 fd/DAY
|
||||
|
||||
Runway from fd=66 at 183/day to the 65536 soft limit:
|
||||
357 days = 0.98 years
|
||||
|
||||
Historical rate implied by the outage: 1016 sockets over 14.29 d = 71.1/day
|
||||
Socket composition: ESTAB 45 -> 49 (+4); CLOSE-WAIT 1 -> 1 (+0).
|
||||
=> ALL of the growth in this window is ESTABLISHED connections, not CLOSE-WAIT.
|
||||
+200
@@ -0,0 +1,200 @@
|
||||
Get:1 https://metadata.cdn.proxmox.com rust-proxmox-backup 4.2.5-1 Changelog [178 kB]
|
||||
rust-proxmox-backup (4.2.5-1) trixie; urgency=medium
|
||||
|
||||
* backup: harden the handling of client supplied backup manifests:
|
||||
- only accept archive names that are plain file names carrying a server
|
||||
side type extension. A crafted name in a manifest could previously make
|
||||
a sync job read or write outside of the snapshot directory, running as
|
||||
the unprivileged 'backup' user. Reaching this needed a manifest from a
|
||||
configured sync remote or from a client that already had backup access
|
||||
to the datastore.
|
||||
- keep an uploaded manifest in memory and only persist it on backup
|
||||
finish, checking that every archive it lists was really uploaded during
|
||||
that session and that the checksums match the ones computed server
|
||||
side. A client uploading a manifest that references archives it did not
|
||||
upload now gets an error on finish instead of such a snapshot being
|
||||
created.
|
||||
|
||||
* fix #7878: sync: push: reuse the manifest of a previous snapshot on a
|
||||
non-encrypting push if the source snapshot was encrypted with a matching
|
||||
key, restoring chunk reuse and thus avoiding needlessly long sync runs.
|
||||
|
||||
* sync: push: keep the sign-only crypt mode of a source archive instead of
|
||||
reducing it to unencrypted when pushing without server side encryption.
|
||||
|
||||
* subscription: reject a subscription key issued for a different
|
||||
architecture than the host, as arm64 keys carry an explicit marker, so a
|
||||
wrong key fails fast instead of only erroring during the online check.
|
||||
|
||||
* update to proxmox-upgrade-checks 1.1, which accepts the 7.0 kernel, tells
|
||||
a bookworm backport apart from a trixie build and fixes the dkms check.
|
||||
|
||||
* docs: clarify in the backup protocol description that the manifest is
|
||||
uploaded by the client and only persisted on backup finish.
|
||||
|
||||
-- Proxmox Support Team <support@proxmox.com> Wed, 05 Aug 2026 18:25:37 +0200
|
||||
|
||||
rust-proxmox-backup (4.2.4-1) trixie; urgency=medium
|
||||
|
||||
* docs: document the debug symbol repository
|
||||
|
||||
* datastore: fix wrong local path used for S3 bad chunk handling during
|
||||
garbage collection
|
||||
|
||||
* refactor file creation/mode/ownership helpers to proxmox-product-config
|
||||
crate
|
||||
|
||||
* fix #7642: avoid expensive user lookups on file locking by caching the
|
||||
backup user/group ID
|
||||
|
||||
* depend on proxmox-enterprise-support-keyring, and track its version in the
|
||||
package version API endpoint
|
||||
|
||||
* fix #5748: docs: add `catalog.pcat1` format specification
|
||||
|
||||
* docs: system requirements: document we recommend local storage
|
||||
|
||||
* S3: fix #6841: allow configuring request rate limits by updating to
|
||||
proxmox-s3-client 1.4.1. these rate limits are split into active and
|
||||
passive methods, allowing separate handling of POST/PUT/DELETE and GET/HEAD
|
||||
request limits.
|
||||
|
||||
* S3: config: allow editing the use-node-config flag that controls whether
|
||||
requests S3 endpoints honor the node's proxy settings or not
|
||||
|
||||
* sync: push: gracefully handle previous manifest signature mismatches, which
|
||||
can happen when enabling or disabling push-encryption on an already synced
|
||||
backup group
|
||||
|
||||
-- Proxmox Support Team <support@proxmox.com> Wed, 29 Jul 2026 15:08:15 +0200
|
||||
|
||||
rust-proxmox-backup (4.2.3-1) trixie; urgency=medium
|
||||
|
||||
* css: remove x-grid-row-loading class, replace it with non-blurry SVG
|
||||
variant from proxmox-widget-toolkit
|
||||
|
||||
* pbs-client: add backoff log throttle, to ensure progress and similar output
|
||||
appears quickly initially, but does not create overly long logs
|
||||
|
||||
* client: report progress during restore
|
||||
|
||||
* api: journal: adopt proxmox-syslog-api and stream the output, making the
|
||||
implementation consistent with the one from Proxmox Datacenter Manager
|
||||
|
||||
* ui: enable the structured journal view and per-service logs, including
|
||||
colored output and filtering capabilities
|
||||
|
||||
* ui: always use arrays for 'delete' property, instead of manually converting
|
||||
|
||||
* fix #5971: tape: don't warn on custom MAM attribute write failures
|
||||
|
||||
* fix #7175: api: time: use timedatectl instead of /etc/timezone
|
||||
|
||||
* fix #7187: report: add ethtool output for physical interfaces
|
||||
|
||||
* prune jobs: schedule jobs that do not prune anything, but warn during their
|
||||
execution. such jobs allow testing scheduling options, but make no sense
|
||||
for production use.
|
||||
|
||||
* fix #6691: allow search by comment in datastore content, make search
|
||||
case-insensitive and correctly reset content view after empty searches
|
||||
|
||||
* ui: datastore: disable various action tooltips for actions which cannot be
|
||||
triggered
|
||||
|
||||
* ldap: escape the user-provided user name when using it in the LDAP search
|
||||
filter that looks up the user DN.
|
||||
|
||||
* ldap sync: log which user properties change when synchronizing an existing
|
||||
user, instead of only reporting that the user was updated.
|
||||
|
||||
* api schema/section config: add support for declaring deprecated property
|
||||
aliases, to allow renaming properties without showing the old name in the
|
||||
documentation
|
||||
|
||||
* rest server: accept deprecated property aliases in JSON request bodies by
|
||||
rewriting them to the canonical name before verification and dispatch, like
|
||||
the CLI and query-string handling already do.
|
||||
|
||||
* fix #7690: fs: replace_file: close the temporary file before renaming or
|
||||
unlinking it, fixing the replacement on WORM file systems and avoiding
|
||||
leftover .fuse_hidden files on FUSE mounts.
|
||||
|
||||
* fs: make_tmp_file: append the temporary suffix instead of replacing the
|
||||
file extension, keeping the original file name intact for easier
|
||||
debugging of leftover temporary files.
|
||||
|
||||
-- Proxmox Support Team <support@proxmox.com> Tue, 14 Jul 2026 12:54:15 +0200
|
||||
|
||||
rust-proxmox-backup (4.2.2-1) trixie; urgency=medium
|
||||
|
||||
* api: backup: run synchronous chunk-insert operations off the asynchronous
|
||||
runtime's worker threads. Blocking those threads, most notably during S3
|
||||
uploads that can wait up to three hours for a chunk lock, could stall the
|
||||
I/O and timer drivers of the entire runtime and starve other backup
|
||||
workers.
|
||||
|
||||
* client: backup: make the file-based backup more robust against files that
|
||||
cannot be accessed or that vanish while the backup is running:
|
||||
- fix #7658: skip a file and log a warning instead of aborting the whole
|
||||
backup when querying its metadata fails with a permission error,
|
||||
matching the existing handling of permission errors when opening files.
|
||||
Such files can also still be excluded explicitly.
|
||||
- consistently ignore files that disappear during the backup and warn
|
||||
about them, instead of treating this as a fatal error in some cases.
|
||||
|
||||
* tape: backup: fix the command-line group filter, which was passed to the
|
||||
API under the wrong parameter name and thus had no effect.
|
||||
|
||||
* api: do not log a spurious error when listing the files of a snapshot
|
||||
whose manifest does not exist yet, which is expected while a backup to
|
||||
that snapshot is still running.
|
||||
|
||||
* ui: datastore summary: fix the per-datastore sync and prune job counts,
|
||||
which were derived from an incorrectly parsed datastore ID; the prune
|
||||
count in particular was always shown as zero.
|
||||
|
||||
* tape: fix two typos in log and informational messages.
|
||||
|
||||
-- Proxmox Support Team <support@proxmox.com> Thu, 18 Jun 2026 11:28:25 +0200
|
||||
|
||||
rust-proxmox-backup (4.2.1-1) trixie; urgency=medium
|
||||
|
||||
* fix #5076: api: support an 'audiences' property on OpenID realms, listing
|
||||
additional trusted audience values besides the configured client-id.
|
||||
Improves compatibility with providers that issue tokens with multiple
|
||||
audiences.
|
||||
|
||||
* fix #7562: api/ui: tape: separate the format-media 'load-barcode'
|
||||
parameter from the existing 'label-text' verification, so the web
|
||||
interface can load and format empty or previously-unrelated tapes from a
|
||||
changer slot in one step. The old shared parameter aborted formatting
|
||||
after a successful load whenever the on-tape label did not match.
|
||||
|
||||
* sync: pull: refuse to overwrite a locally encrypted snapshot from an
|
||||
unencrypted source or one using a different key, and detect content
|
||||
differences between two unencrypted snapshots that share a backup time.
|
||||
Previously such mismatches silently triggered a resync that overwrote the
|
||||
local snapshot.
|
||||
|
||||
* datastore: fix tuning option changes not propagating the updated sync
|
||||
level to the chunk store until the service was restarted.
|
||||
|
||||
* datastore: improve the error message when prune cannot acquire a snapshot
|
||||
lock for deletion, by showing the snapshot directory and lock file paths
|
||||
instead of an internal debug dump.
|
||||
|
||||
* api: backup: fix benchmark, finish-failed, and backup-failed cleanups
|
||||
leaving an orphaned empty backup group behind on the datastore.
|
||||
|
||||
* api: node: tasks status: return the task end time as an optional field
|
||||
once the task is finished, so the task viewer can render the correct
|
||||
duration without an extra API call.
|
||||
|
||||
* api/ui: node: add a 'location' property to the node config, exposed
|
||||
through the node options panel.
|
||||
|
||||
* subscription: reuse the server ID from an existing subscription info when
|
||||
multiple candidates are detected, falling back to the first candidate only
|
||||
when no prior info exists.
|
||||
|
||||
+128
@@ -0,0 +1,128 @@
|
||||
rust-proxmox-backup (4.2.5-1) trixie; urgency=medium
|
||||
|
||||
* backup: harden the handling of client supplied backup manifests:
|
||||
- only accept archive names that are plain file names carrying a server
|
||||
side type extension. A crafted name in a manifest could previously make
|
||||
a sync job read or write outside of the snapshot directory, running as
|
||||
the unprivileged 'backup' user. Reaching this needed a manifest from a
|
||||
configured sync remote or from a client that already had backup access
|
||||
to the datastore.
|
||||
- keep an uploaded manifest in memory and only persist it on backup
|
||||
finish, checking that every archive it lists was really uploaded during
|
||||
that session and that the checksums match the ones computed server
|
||||
side. A client uploading a manifest that references archives it did not
|
||||
upload now gets an error on finish instead of such a snapshot being
|
||||
created.
|
||||
|
||||
* fix #7878: sync: push: reuse the manifest of a previous snapshot on a
|
||||
non-encrypting push if the source snapshot was encrypted with a matching
|
||||
key, restoring chunk reuse and thus avoiding needlessly long sync runs.
|
||||
|
||||
* sync: push: keep the sign-only crypt mode of a source archive instead of
|
||||
reducing it to unencrypted when pushing without server side encryption.
|
||||
|
||||
* subscription: reject a subscription key issued for a different
|
||||
architecture than the host, as arm64 keys carry an explicit marker, so a
|
||||
wrong key fails fast instead of only erroring during the online check.
|
||||
|
||||
* update to proxmox-upgrade-checks 1.1, which accepts the 7.0 kernel, tells
|
||||
a bookworm backport apart from a trixie build and fixes the dkms check.
|
||||
|
||||
* docs: clarify in the backup protocol description that the manifest is
|
||||
uploaded by the client and only persisted on backup finish.
|
||||
|
||||
-- Proxmox Support Team <support@proxmox.com> Wed, 05 Aug 2026 18:25:37 +0200
|
||||
|
||||
rust-proxmox-backup (4.2.4-1) trixie; urgency=medium
|
||||
|
||||
* docs: document the debug symbol repository
|
||||
|
||||
* datastore: fix wrong local path used for S3 bad chunk handling during
|
||||
garbage collection
|
||||
|
||||
* refactor file creation/mode/ownership helpers to proxmox-product-config
|
||||
crate
|
||||
|
||||
* fix #7642: avoid expensive user lookups on file locking by caching the
|
||||
backup user/group ID
|
||||
|
||||
* depend on proxmox-enterprise-support-keyring, and track its version in the
|
||||
package version API endpoint
|
||||
|
||||
* fix #5748: docs: add `catalog.pcat1` format specification
|
||||
|
||||
* docs: system requirements: document we recommend local storage
|
||||
|
||||
* S3: fix #6841: allow configuring request rate limits by updating to
|
||||
proxmox-s3-client 1.4.1. these rate limits are split into active and
|
||||
passive methods, allowing separate handling of POST/PUT/DELETE and GET/HEAD
|
||||
request limits.
|
||||
|
||||
* S3: config: allow editing the use-node-config flag that controls whether
|
||||
requests S3 endpoints honor the node's proxy settings or not
|
||||
|
||||
* sync: push: gracefully handle previous manifest signature mismatches, which
|
||||
can happen when enabling or disabling push-encryption on an already synced
|
||||
backup group
|
||||
|
||||
-- Proxmox Support Team <support@proxmox.com> Wed, 29 Jul 2026 15:08:15 +0200
|
||||
|
||||
rust-proxmox-backup (4.2.3-1) trixie; urgency=medium
|
||||
|
||||
* css: remove x-grid-row-loading class, replace it with non-blurry SVG
|
||||
variant from proxmox-widget-toolkit
|
||||
|
||||
* pbs-client: add backoff log throttle, to ensure progress and similar output
|
||||
appears quickly initially, but does not create overly long logs
|
||||
|
||||
* client: report progress during restore
|
||||
|
||||
* api: journal: adopt proxmox-syslog-api and stream the output, making the
|
||||
implementation consistent with the one from Proxmox Datacenter Manager
|
||||
|
||||
* ui: enable the structured journal view and per-service logs, including
|
||||
colored output and filtering capabilities
|
||||
|
||||
* ui: always use arrays for 'delete' property, instead of manually converting
|
||||
|
||||
* fix #5971: tape: don't warn on custom MAM attribute write failures
|
||||
|
||||
* fix #7175: api: time: use timedatectl instead of /etc/timezone
|
||||
|
||||
* fix #7187: report: add ethtool output for physical interfaces
|
||||
|
||||
* prune jobs: schedule jobs that do not prune anything, but warn during their
|
||||
execution. such jobs allow testing scheduling options, but make no sense
|
||||
for production use.
|
||||
|
||||
* fix #6691: allow search by comment in datastore content, make search
|
||||
case-insensitive and correctly reset content view after empty searches
|
||||
|
||||
* ui: datastore: disable various action tooltips for actions which cannot be
|
||||
triggered
|
||||
|
||||
* ldap: escape the user-provided user name when using it in the LDAP search
|
||||
filter that looks up the user DN.
|
||||
|
||||
* ldap sync: log which user properties change when synchronizing an existing
|
||||
user, instead of only reporting that the user was updated.
|
||||
|
||||
* api schema/section config: add support for declaring deprecated property
|
||||
aliases, to allow renaming properties without showing the old name in the
|
||||
documentation
|
||||
|
||||
* rest server: accept deprecated property aliases in JSON request bodies by
|
||||
rewriting them to the canonical name before verification and dispatch, like
|
||||
the CLI and query-string handling already do.
|
||||
|
||||
* fix #7690: fs: replace_file: close the temporary file before renaming or
|
||||
unlinking it, fixing the replacement on WORM file systems and avoiding
|
||||
leftover .fuse_hidden files on FUSE mounts.
|
||||
|
||||
* fs: make_tmp_file: append the temporary suffix instead of replacing the
|
||||
file extension, keeping the original file name intact for easier
|
||||
debugging of leftover temporary files.
|
||||
|
||||
-- Proxmox Support Team <support@proxmox.com> Tue, 14 Jul 2026 12:54:15 +0200
|
||||
|
||||
rust-proxmox-backup (4.2.2-1) trixie; urgency=medium
|
||||
@@ -0,0 +1,30 @@
|
||||
### apt-get update start
|
||||
2026-08-18T09:50:43+00:00
|
||||
Hit:6 http://mirror.hetzner.com/debian/security trixie-security InRelease
|
||||
Hit:7 http://download.proxmox.com/debian/pbs trixie InRelease
|
||||
Get:8 http://deb.debian.org/debian trixie-backports InRelease [54.0 kB]
|
||||
Get:9 http://deb.debian.org/debian-security trixie-security InRelease [43.4 kB]
|
||||
Get:10 http://mirror.hetzner.com/debian/packages trixie-backports/main amd64 Packages [314 kB]
|
||||
Get:11 http://deb.debian.org/debian trixie-backports/main Sources [297 kB]
|
||||
Fetched 858 kB in 0s (6119 kB/s)
|
||||
Reading package lists...
|
||||
### SIMULATION (-s): nothing is installed by this
|
||||
2026-08-18T09:50:44+00:00
|
||||
Reading package lists...
|
||||
Building dependency tree...
|
||||
Reading state information...
|
||||
The following additional packages will be installed:
|
||||
proxmox-backup-client proxmox-backup-docs proxmox-enterprise-support-keyring
|
||||
The following NEW packages will be installed:
|
||||
proxmox-enterprise-support-keyring
|
||||
The following packages will be upgraded:
|
||||
proxmox-backup-client proxmox-backup-docs proxmox-backup-server
|
||||
3 upgraded, 1 newly installed, 0 to remove and 8 not upgraded.
|
||||
Inst proxmox-backup-client [4.2.2-1] (4.2.5-1 Proxmox Backup System Debian Repository:stable [amd64])
|
||||
Inst proxmox-backup-docs [4.2.2-1] (4.2.5-1 Proxmox Backup System Debian Repository:stable [all])
|
||||
Inst proxmox-enterprise-support-keyring (1.1 Proxmox Backup System Debian Repository:stable [all])
|
||||
Inst proxmox-backup-server [4.2.2-1] (4.2.5-1 Proxmox Backup System Debian Repository:stable [amd64])
|
||||
Conf proxmox-backup-client (4.2.5-1 Proxmox Backup System Debian Repository:stable [amd64])
|
||||
Conf proxmox-backup-docs (4.2.5-1 Proxmox Backup System Debian Repository:stable [all])
|
||||
Conf proxmox-enterprise-support-keyring (1.1 Proxmox Backup System Debian Repository:stable [all])
|
||||
Conf proxmox-backup-server (4.2.5-1 Proxmox Backup System Debian Repository:stable [amd64])
|
||||
@@ -0,0 +1,58 @@
|
||||
### UPGRADE START
|
||||
2026-08-18T09:51:00+00:00
|
||||
Reading package lists...
|
||||
Building dependency tree...
|
||||
Reading state information...
|
||||
The following additional packages will be installed:
|
||||
proxmox-backup-client proxmox-backup-docs proxmox-enterprise-support-keyring
|
||||
The following NEW packages will be installed:
|
||||
proxmox-enterprise-support-keyring
|
||||
The following packages will be upgraded:
|
||||
proxmox-backup-client proxmox-backup-docs proxmox-backup-server
|
||||
3 upgraded, 1 newly installed, 0 to remove and 8 not upgraded.
|
||||
Need to get 47.8 MB of archives.
|
||||
After this operation, 1404 kB of additional disk space will be used.
|
||||
Get:1 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 proxmox-backup-client amd64 4.2.5-1 [3447 kB]
|
||||
Get:2 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 proxmox-backup-docs all 4.2.5-1 [6351 kB]
|
||||
Get:3 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 proxmox-enterprise-support-keyring all 1.1 [2988 B]
|
||||
Get:4 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 proxmox-backup-server amd64 4.2.5-1 [38.0 MB]
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
locale: Cannot set LC_CTYPE to default locale: No such file or directory
|
||||
locale: Cannot set LC_ALL to default locale: No such file or directory
|
||||
Fetched 47.8 MB in 1s (39.3 MB/s)
|
||||
(Reading database ...
|
||||
(Reading database ... 5%
|
||||
(Reading database ... 10%
|
||||
(Reading database ... 15%
|
||||
(Reading database ... 20%
|
||||
(Reading database ... 25%
|
||||
(Reading database ... 30%
|
||||
(Reading database ... 35%
|
||||
(Reading database ... 40%
|
||||
(Reading database ... 45%
|
||||
(Reading database ... 50%
|
||||
(Reading database ... 55%
|
||||
(Reading database ... 60%
|
||||
(Reading database ... 65%
|
||||
(Reading database ... 70%
|
||||
(Reading database ... 75%
|
||||
(Reading database ... 80%
|
||||
(Reading database ... 85%
|
||||
@@ -0,0 +1,48 @@
|
||||
### verify-time
|
||||
2026-08-18T09:51:22+00:00
|
||||
### dpkg -l (INSTALLED)
|
||||
ii proxmox-backup-client 4.2.5-1 amd64 Proxmox Backup Client tools
|
||||
ii proxmox-backup-docs 4.2.5-1 all Proxmox Backup Documentation
|
||||
ii proxmox-backup-server 4.2.5-1 amd64 Proxmox Backup Server daemon with tools and GUI
|
||||
ii proxmox-enterprise-support-keyring 1.1 all Proxmox enterprise SSH support keyring
|
||||
### daemons
|
||||
active
|
||||
inactive
|
||||
active
|
||||
### proxy MainPID (should DIFFER from 542065 = it restarted)
|
||||
551655
|
||||
### proxy start time
|
||||
Tue Aug 18 09:51:04 2026
|
||||
### EFFECTIVE open files on the RUNNING proxy (must be 65536)
|
||||
Max open files 65536 65536 files
|
||||
### drop-ins still present?
|
||||
/etc/systemd/system/proxmox-backup-proxy.service.d/:
|
||||
total 16
|
||||
drwxr-xr-x 2 root root 4096 Aug 18 03:54 .
|
||||
drwxr-xr-x 26 root root 4096 Jul 27 07:53 ..
|
||||
-rw-r--r-- 1 root root 44 Jul 27 07:16 10-datastore-mount.conf
|
||||
-rw-r--r-- 1 root root 337 Aug 18 03:54 20-nofile.conf
|
||||
|
||||
/etc/systemd/system/proxmox-backup.service.d/:
|
||||
total 16
|
||||
drwxr-xr-x 2 root root 4096 Aug 18 03:55 .
|
||||
drwxr-xr-x 26 root root 4096 Jul 27 07:53 ..
|
||||
-rw-r--r-- 1 root root 44 Jul 27 07:16 10-datastore-mount.conf
|
||||
-rw-r--r-- 1 root root 237 Aug 18 03:55 20-nofile.conf
|
||||
### fd count (new t0)
|
||||
17
|
||||
### listen queue (Recv-Q must be 0)
|
||||
State Recv-Q Send-Q Local Address:Port Peer Address:Port
|
||||
LISTEN 0 1024 *:8007 *:*
|
||||
### socket states
|
||||
1 LISTEN
|
||||
### loopback
|
||||
http=200 time=0.012245
|
||||
### version string
|
||||
proxmox-backup-server 4.2.5-1 running version: 4.2.5
|
||||
### datastore
|
||||
+================+====================+=========+
|
||||
| name | path | comment |
|
||||
+================+====================+=========+
|
||||
| felhom-offsite | /mnt/pbs-datastore | |
|
||||
+================+====================+=========+
|
||||
@@ -0,0 +1,11 @@
|
||||
### does proxmox-backup-api exist?
|
||||
No files found for proxmox-backup-api.service.
|
||||
### actual PBS units
|
||||
proxmox-backup-banner.service loaded active exited Proxmox Backup Server Login Banner
|
||||
proxmox-backup-daily-update.service loaded inactive dead Daily Proxmox Backup Server update and maintenance activities
|
||||
proxmox-backup-proxy.service loaded active running Proxmox Backup API Proxy Server
|
||||
proxmox-backup.service loaded active running Proxmox Backup API Server
|
||||
### API server description
|
||||
MainPID=551642
|
||||
Description=Proxmox Backup API Server
|
||||
ActiveState=active
|
||||
@@ -0,0 +1,5 @@
|
||||
### felhom-pve
|
||||
2026-08-18T11:51:48+02:00
|
||||
http=200 time=0.103326
|
||||
Name Type Status Total (KiB) Used (KiB) Available (KiB) %
|
||||
felhom-pbs pbs active 0 0 0 0.00%
|
||||
@@ -0,0 +1,5 @@
|
||||
### demo-hp
|
||||
2026-08-18T11:51:50+02:00
|
||||
http=200 time=0.095982
|
||||
Name Type Status Total (KiB) Used (KiB) Available (KiB) %
|
||||
felhom-pbs pbs active 0 0 0 0.00%
|
||||
@@ -0,0 +1 @@
|
||||
POST-UPGRADE REFRESH: 2026/08/18 11:59:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)
|
||||
@@ -0,0 +1,3 @@
|
||||
2026/08/18 11:12:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)
|
||||
2026/08/18 11:28:30 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)
|
||||
2026/08/18 11:43:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)
|
||||
@@ -0,0 +1,18 @@
|
||||
### step6-second-reading (post-upgrade)
|
||||
2026-08-18T10:23:21+00:00
|
||||
### MainPID
|
||||
551655
|
||||
### proxy start
|
||||
Tue Aug 18 09:51:04 2026
|
||||
### fd count
|
||||
22
|
||||
### socket states
|
||||
5 ESTAB
|
||||
1 LISTEN
|
||||
### close-wait
|
||||
1
|
||||
### listen queue
|
||||
State Recv-Q Send-Q Local Address:Port Peer Address:Port
|
||||
LISTEN 0 1024 *:8007 *:*
|
||||
### limits
|
||||
Max open files 65536 65536 files
|
||||
@@ -0,0 +1,20 @@
|
||||
AFTER-SLOPE — proxy PID 551655, started 2026-08-18 09:51:04 UTC (the upgrade restart)
|
||||
|
||||
09:51:22Z fd=17 (ESTAB 0) -> 10:23:21Z fd=22 (ESTAB 5)
|
||||
interval 0.5331 h (1919 s), delta 5 fd
|
||||
= 9.38 fd/hour = 225.1 fd/DAY
|
||||
|
||||
BEFORE (31 min window): 183.3/day from delta=4 over 1885 s
|
||||
AFTER (32 min window): 225.1/day from delta=5 over 1919 s
|
||||
|
||||
VERDICT: NOT DISTINGUISHABLE. The two windows differ by ONE descriptor.
|
||||
With counts this small the Poisson uncertainty on n=4 is +/-2 and on n=5 is +/-2.2,
|
||||
so both are consistent with a single unchanged underlying rate. The after-figure being
|
||||
numerically HIGHER is noise, not a regression -- and it is certainly not an improvement.
|
||||
|
||||
This is the EXPECTED result: the 4.2.2->4.2.5 changelog contains no mechanism by which
|
||||
connection reaping would change. Recorded before the numbers existed, in stop1-ruling.txt.
|
||||
|
||||
Composition again favours ESTAB: 0 -> 5 ESTAB, 0 -> 1 CLOSE-WAIT.
|
||||
|
||||
30 MINUTES CANNOT SETTLE THIS. The honest checks are +24 h and +7 d -> filed as R-341.
|
||||
@@ -0,0 +1,25 @@
|
||||
STOP 1 — operator ruling
|
||||
========================
|
||||
Recorded: 2026-08-18 ~09:25 UTC (11:25 CEST)
|
||||
Given by: Viktor (operator), in session.
|
||||
|
||||
CC's recommendation was: DO NOT UPGRADE for the stated reason.
|
||||
Basis: the full changelog range 4.2.2-1 -> 4.2.5-1 (128 lines, all three
|
||||
entries read) contains NOTHING touching connection handling, descriptor
|
||||
lifetime, accept(), CLOSE-WAIT, keep-alive or the proxy daemon's socket
|
||||
lifecycle. An explicit keyword sweep returned exactly one hit, and it is a
|
||||
false positive ("S3 ... honor the node's proxy settings" = HTTP proxy config
|
||||
for S3 requests, not the PBS proxy daemon).
|
||||
|
||||
OPERATOR RULING: PROCEED WITH THE UPGRADE ANYWAY.
|
||||
Stated reason: rehearsal value -- "see how that works for us, we need
|
||||
practice with that too". The upgrade is therefore being performed as a
|
||||
supervised practice run of the upgrade procedure on a Tier-2 protected
|
||||
machine, NOT as a fix for R-336's leak.
|
||||
|
||||
CONSEQUENCE TO CARRY INTO THE REPORT, so it is not later misread:
|
||||
this upgrade is NOT expected to change the fd slope. If the post-upgrade
|
||||
slope differs, that is a surprise requiring explanation, not a confirmation
|
||||
of anything -- the changelog gives no mechanism by which it should improve.
|
||||
Equally, if the slope is unchanged, that is the EXPECTED result and is not
|
||||
evidence the upgrade failed.
|
||||
@@ -0,0 +1,31 @@
|
||||
STOP 2 — Hetzner snapshot (the rollback)
|
||||
========================================
|
||||
Confirmed by the operator in the Hetzner console, 2026-08-18 ~09:27 UTC (11:27 CEST).
|
||||
|
||||
Snapshot ID : 421440873
|
||||
Description : felhom-hetzner-20260818
|
||||
Image size : 15.06 GB
|
||||
Status : Available <-- COMPLETE, not merely started
|
||||
Created : "less than a minute ago" as displayed at confirmation
|
||||
Server : felhom-hetzner #147604682 (CX33, 167.233.158.164)
|
||||
Hetzner project: 15217960 (Felhom.eu)
|
||||
|
||||
This is the rollback point for the 4.2.2-1 -> 4.2.5-1 upgrade performed in
|
||||
Step 4. It was taken on a RUNNING server: Hetzner's own console recommends
|
||||
powering off first for disk consistency, and that was not done -- a
|
||||
deliberate trade, because powering off ep0 takes the only off-premises copy
|
||||
offline and the upgrade being guarded is a userspace package install that
|
||||
touches neither the datastore nor the boot path.
|
||||
|
||||
WHAT THIS SNAPSHOT DOES AND DOES NOT COVER:
|
||||
covers - the 38 GB system disk /dev/sda (root), i.e. the PBS packages,
|
||||
unit files, /etc/systemd drop-ins, nftables and wg config.
|
||||
DOES NOT - /mnt/pbs-datastore. That is /dev/sdb, a separate 100 GB
|
||||
VOLUME, and Hetzner server snapshots do not include attached
|
||||
volumes. The backup data is therefore NOT protected by this
|
||||
snapshot.
|
||||
Consequence: rolling back restores the software state, not the datastore.
|
||||
This is acceptable for THIS change because the upgrade writes no datastore
|
||||
content -- but it must not be mistaken for a datastore backup, and any
|
||||
future runbook step that could touch /mnt/pbs-datastore needs a different
|
||||
safeguard than this one.
|
||||
@@ -636,6 +636,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-333** | **Two disk-health questions the deploy raised and did NOT act on.** **(a) The 55/60 °C bands are SPINNING-DISK bands applied to NVMe.** They were adopted unchanged from the operator's Prometheus config so the two systems cannot disagree — a deliberate, stated decision — but **measured on demo-hp 2026-08-14 the healthy Toshiba KXG50PNV1T02 NVMe idles at 53 °C, two degrees below Figyelmeztetés and seven below Hiba**, and NVMe routinely exceeds 60 °C under load with no fault whatever. As it stands a healthy customer NVMe under sustained write can be reported as **Hiba** — the single worst outcome this feature can produce. **(b) The agent runs bare `smartctl -a -j` with no `-n standby`** (`felhom-agent/internal/storage/hostops.go:368`), so every poll WAKES a spun-down drive; going 6h → hourly multiplies that by six. demo-hp is all-flash so the cadence measurement could not reveal it, and it was recorded rather than acted on per the task's own instruction. Mitigating datum from the fixture: the failing drive logged only **3375 load cycles in 60505 hours** (~one per 18h), i.e. that duty cycle barely spins down at all | **READY (S each) — NEW 2026-08-14** | — | (a) split the temperature bands by device class, or drop them for NVMe and rely on `critical_warning`; (b) add `-n standby` to the agent's smartctl invocation (an agent change, so fold it into R-330's session) | Viktor decides (a); CC does (b) |
|
||||
| **R-334** | **WAIVER + open item: controller v0.215.0 is released and deployed, and NO golden carries it.** Convicted by `golden_currency_gate.py` on the 2026-08-14 push: newest released controller **0.215.0**, newest golden bake **0.214.0** (`documentation/tests/golden-0.214.0-2026-08-12`). **A machine installed right now receives 0.214.0** — i.e. a brand-new box would ship WITHOUT the R-328 severity fix and would keep emailing nobody about a failing disk. The running fleet is unaffected (demo-hp guest 9201 is on 0.215.0 and healthy); this is purely the day-0 install path. **Not baked in this session deliberately:** the task scoped deployment to demo-hp only, and the second half of the fix — vouching the bake in the hub's day-0 artifact manifest — is **operator-password-gated, so CC cannot complete it**; a baked-but-unvouched golden is worse than none. **The push was made with `git push --no-verify` and it is stated here and in the session report**, per `.claude/rules/gates.md` — the gate has no waiver parser, so recording a waiver does not clear it. **STILL OPEN and now one version WIDER, 2026-08-18:** the newest released controller is **0.216.0** (v0.215.0's own follow-up fix, R-335) and the newest bake is still **0.214.0**, so a new install now misses *two* releases. Re-convicted on this date's documentation-only push, which was likewise made with `--no-verify`; the gate reads `felhom-controller/CHANGELOG.md` and `documentation/tests/golden-*`, **neither of which that session touched** — the conviction is inherited, not caused | **READY (S) — NEW 2026-08-14, re-confirmed 2026-08-18** | operator availability for the vouch step | Bake a golden on **0.216.0** per `runbooks/RUNBOOK-manual-build.md` §4.1, then vouch it — a THREE-field change (`golden_version` + `agent_version` + `min_agent`). Until then every NEW install lacks the severity fix | CC bakes; **Viktor vouches** |
|
||||
| **R-335** | **One physical disk was walked TWICE per run, and the second walk sustained it against itself.** Found on live hardware ~2h after the v0.215.0 deploy, **by noticing the release's own positive observable disagreed with its own persisted artefact**: the hourly check logged *"3 disk(s) evaluated"* while `disk-health-state.json` held **two** records. Cause: demo-hp's `c11-scratch` and `felhom-backup` are the same physical NVMe (`/dev/nvme0n1`) and resolve to the same `diskKey`. **Not cosmetic** — `RunDiskHealthCheck` writes a disk's new record before the next entry reads it, so the SECOND copy consumed the FIRST copy's write as its prior: the disk **sustained against itself and reached Hiba on a FIRST sighting**, defeating truth-table row 6 — the exact rule separating a one-hour benign excursion from a false critical — and would have emitted **two identical events** for one drive. **Latent, not active, on demo-hp** (all three entries healthy, zero counters), but any aliased disk developing a single pending sector would have gone straight to Hiba. **This is the shape standing rule 3 warns about: an absent alarm was not evidence — the two artefacts had to be read AGAINST each other** | **CLOSED — controller v0.216.0, 2026-08-14.** Each `diskKey` is evaluated once per run; both entries stay marked `seen` so neither looks like a disappeared disk, and the card still renders both storage rows (the dedup is about state and alerts, not display). Pinned by `TestDiskCheck_SameDiskTwiceIsEvaluatedOnce`; companion red-proof run and reverted — deleting the guard makes the first sighting emit `Kind:2` (Hiba-from-sectors) at 8 sectors | — | — | CC |
|
||||
| **R-336** | **The offsite DR endpoint is polled about once per second, and that is what turned a slow leak into an outage.** ep0's PBS proxy served **~85,000 requests/day** — a flat **3,538/hour**, every hour, from two boxes: `74,445 GET /api2/json/admin/datastore` (`libwww-perl`, i.e. PVE's `pvestatd`) and `73,171 GET /admin/datastore/felhom-offsite/status` (`proxmox-backup-client`). Two pollers asking substantially the same question at the same rate. On 2026-08-18 this walked a connection leak in the proxy to its 1024-fd soft limit in **14 days**, wedging the offsite tier for 9½ hours (`audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md`). The `LimitNOFILE=65536` drop-in applied that morning raises the ceiling **but does not fix the leak** — it converts a fortnightly outage into a multi-year one, which is mitigation, not a fix. A DR endpoint that is written to weekly does not need to be asked about every second | **READY (M) — NEW 2026-08-18** | — | Find what polls `felhom-pbs` this hard (PVE storage status is the prime suspect, and its interval is tunable) and cut it; then confirm the fd count between restarts stops climbing — the positive observable, per standing rule 3. **Baseline measured 2026-08-18: 19 fds at 16m43s post-restart, ~85/day, matching the ~73/day implied by the original failure — the leak is confirmed live, not assumed** | CC |
|
||||
| **R-336** | **The offsite DR endpoint is polled about once per second, and that is what turned a slow leak into an outage.** ep0's PBS proxy served **~85,000 requests/day** — a flat **3,538/hour**, every hour, from two boxes: `74,445 GET /api2/json/admin/datastore` (`libwww-perl`, i.e. PVE's `pvestatd`) and `73,171 GET /admin/datastore/felhom-offsite/status` (`proxmox-backup-client`). Two pollers asking substantially the same question at the same rate. On 2026-08-18 this walked a connection leak in the proxy to its 1024-fd soft limit in **14 days**, wedging the offsite tier for 9½ hours (`audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md`). The `LimitNOFILE=65536` drop-in applied that morning raises the ceiling **but does not fix the leak** — it converts a fortnightly outage into a multi-year one, which is mitigation, not a fix. A DR endpoint that is written to weekly does not need to be asked about every second | **READY (M) — NEW 2026-08-18** | — | Find what polls `felhom-pbs` this hard (PVE storage status is the prime suspect, and its interval is tunable) and cut it; then confirm the fd count between restarts stops climbing — the positive observable, per standing rule 3. **Baseline measured 2026-08-18, and the FIRST measurement published was WRONG.** The initial "~85/day, matching the ~73/day implied by the failure" came from a single 17-minute window whose delta was **one descriptor** — a sample of one cannot carry a daily rate, and the agreement that made it feel solid was coincidence. **Re-measured over two independent windows the same morning: 183/day (31 min) and 200/day (5.6 h)** — ~2.6x the published figure, putting the runway to the 65536 ceiling at **~357 days, not the ~2 years first claimed**. **And the named mechanism is the minority one:** across that window `CLOSE-WAIT` held flat at 1 while `ESTAB` grew 45→49 — *all* the growth was established connections, and at the wedge the split was 1011 ESTAB / 543 CLOSE-WAIT. **The fix must target connections the proxy never reaps, not just `CLOSE-WAIT` sockets.** The PBS 4.2.5-1 upgrade (2026-08-18) did NOT change the slope and was never expected to — see R-341 | CC |
|
||||
| **R-337** | **`/backup/status` lagged a completed backup by minutes on one box and not the other — and it RESOLVED ITSELF, which is why this is WATCHING and not a defect.** During the R-336 recovery on 2026-08-18, `demo-hp`'s snapshot landed on ep0 at **03:58:43Z** (complete manifest; the host's own task index says `OK`) — yet `GET /backup/status` was **still serving the superseded 03:27:00Z failure at ~04:03Z**, four-plus minutes later. `demo-felhom` showed its new result within ~40 s of completion. **The lag cleared on its own:** demo-hp's 04:07:35Z host report carries `felhom-pbs success=true, 4.29 GB`, and the hub is green for both boxes. **The first draft of this row claimed the success was "still reported as failed" — that was written before the next report arrived and it was wrong; the corrected claim is a several-minute skew between the two boxes, not a stuck value.** It is recorded because a status field that can trail its own artifact by minutes will, during an incident, be read as a second failure — this session nearly did — and because the asymmetry between the two boxes is unexplained | **WATCHING — NEW 2026-08-18** | another observation, ideally during an incident rather than constructed | **Do not open a fix on this as written.** First establish the intended refresh path for `/backup/status` after an out-of-schedule run; only if the skew is not simply collection cadence is there anything to pin. If it is cadence, close this row and say so | CC |
|
||||
| **R-338** | **`demo-hp` is not on the R-50 island at all, and `operations/nodes.md` states that it is.** The page records both fleet boxes as island-migrated 2026-07-25. True of `felhom-pve`; **false of `demo-hp`**, whose `agent.json` has `listen_addr: 192.168.0.87:8443` — the customer LAN address — and **no `island_bridge`/`island_guest_addr` keys at all**, whose guest 9201 has `net0` only (no `eth1`), and whose `vmbr9` exists with **zero members**. The controller's `controller.yaml` points at the LAN address, so the box works; this is inventory drift, not breakage. **Two costs.** A session trusting the page addresses the wrong endpoint — that happened on 2026-08-18 and the resulting timeout was briefly read as a fault. And the agent's local API is **bound to the customer LAN on this box** rather than to a point-to-point island, which is the exposure R-50 was built to remove — so a documented security property is claimed for a box that does not have it | **READY (S) — NEW 2026-08-18** | — | Decide which is true: migrate `demo-hp` to the island, or correct `nodes.md`. Leaving both is the one option that keeps the doc lying | Viktor decides; CC executes |
|
||||
| **R-341** | **Does the fd slope change after the PBS 4.2.5 upgrade? — two dated checks, and the answer is expected to be NO.** ep0 was upgraded 4.2.2-1 → 4.2.5-1 on 2026-08-18 09:51Z on the operator's ruling, **for rehearsal value, not as a fix**: the full changelog range was read (128 lines, all three entries) and swept for connection-handling vocabulary, and it contains **no mechanism** by which descriptor reaping would change — the single keyword hit was `S3 … honor the node's proxy settings`, HTTP-proxy config for S3, not the PBS proxy daemon. **The 32-minute post-upgrade window is indistinguishable from the before window** (+5 fd/1919 s = 225/day vs +4 fd/1885 s = 183/day; the two differ by ONE descriptor and Poisson uncertainty on such counts is ±2, so both are consistent with one unchanged rate — the higher after-figure is noise, not a regression). **Thirty minutes cannot settle it in either direction and this row exists so nobody pretends it did.** New `t0` = **fd 17 at 2026-08-18 09:51:22Z, proxy PID 551655**; before-rate to beat = **183–200/day**. **Interpretation fixed in advance** (`evidence-ep0-pbs-upgrade-2026-08-18/stop1-ruling.txt`, written before any numbers existed): unchanged = EXPECTED, not a failed upgrade; changed = a SURPRISE needing explanation, not a confirmation | **WATCHING — NEW 2026-08-18** | elapsed time only | **Two dated checks, both CC:** **+24 h — 2026-08-19 ~10:00Z** and **+7 d — 2026-08-25 ~10:00Z**. Command (the incident's own positive observable): `ssh root@<ep0> 'PID=$(systemctl show proxmox-backup-proxy -p MainPID --value); ls /proc/$PID/fd \| wc -l; ss -lnt "( sport = :8007 )"; ss -tn state all "( sport = :8007 )" \| awk "NR>1{print \$1}" \| sort \| uniq -c'`. **Record the ESTAB/CLOSE-WAIT split, not just the total** — the split is what says which leak it is. If PID ≠ 551655 the window is void: something restarted the proxy and the count began again | CC on both dates |
|
||||
|
||||
Reference in New Issue
Block a user