RUNBOOK ep0: PBS 4.2.2-1 -> 4.2.5-1, slope unchanged as predicted (R-341)
gates / gates (push) Failing after 14s

Both STOPs cleared by the operator. No code changed; documentation only.

STEP 3 (the run's primary deliverable): the full changelog range 4.2.2-1 ->
4.2.5-1 was read (128 lines, all three entries) and swept for
connection-handling vocabulary. Exactly one keyword hit, a false positive
("S3 ... honor the node's proxy settings" = HTTP proxy config for S3, not the
PBS proxy daemon). 4.2.5-1 is a manifest-hardening security release; 4.2.4-1
is S3 rate limits and a locking cache; 4.2.3-1 is UI/LDAP/tape. NOTHING
addresses descriptor lifetime or connection reaping. Recommendation was: do
not upgrade for this reason.

STOP 1: operator ruled to upgrade anyway for rehearsal value. Recorded as a
practice run, not a fix -- and the interpretation was fixed IN WRITING BEFORE
any numbers existed (stop1-ruling.txt): unchanged = expected; changed =
surprise. Neither outcome could then be rationalised into a success.

STOP 2: Hetzner snapshot 421440873, Available. Documented that it covers
/dev/sda ONLY -- /mnt/pbs-datastore is a separate Volume and is NOT in it, so
it is a software rollback and not a backup of the backup data.

UPGRADE: simulated first (0 to remove), then installed 09:51:00->09:51:06Z,
exit 0. Verified: 4.2.5-1 installed, both daemons active, effective open
files still 65536 (the drop-in survived the new package), Recv-Q 0, loopback
200, 200 from BOTH boxes over the tunnel with felhom-pbs active, and the hub
gauge refreshed post-upgrade at 11:59:31.

SLOPE: before +4 fd/1885 s = 183/day; after +5 fd/1919 s = 225/day. NOT
distinguishable -- one descriptor apart, Poisson +/-2 on such counts. The
higher after-figure is noise, not a regression and not an improvement. 30
minutes cannot settle it; R-341 files the +24 h and +7 d checks.

CORRECTIONS to this morning's own report, both published rather than quietly
fixed:
  - the "~85/day, ~2 years of runway" figures were WRONG. They came from a
    single 17-minute window with a delta of ONE descriptor. Real rate is
    183-200/day over two independent windows; runway ~357 days, not 2 years.
  - the leak was attributed to CLOSE-WAIT. It is mostly ESTAB: CLOSE-WAIT held
    flat at 1 while ESTAB grew 45->49, and at the wedge it was 1011 ESTAB vs
    543 CLOSE-WAIT. R-336's fix must target unreaped connections.
  - "proxmox-backup-api" reported inactive during verification; that unit does
    not exist. Bad query, not a fault, written down because it looked like one.

R-336 stays open: even a fixed leak would not make ~85k requests/day to a
weekly-write DR endpoint correct.

golden-currency still convicts (inherited R-334, controller 0.216.0 vs golden
0.214.0, untouched by this run), so this push is --no-verify per
.claude/rules/gates.md.
This commit is contained in:
2026-08-18 12:26:53 +02:00
parent 435e044cf1
commit 3e50902a98
22 changed files with 1026 additions and 113 deletions
+146 -103
View File
@@ -1,133 +1,176 @@
# REPORT — a listening socket that served nobody (2026-08-18, early)
# REPORT — RUNBOOK ep0: read the PBS changelog, then decide whether to upgrade (2026-08-18, midday)
**Trigger:** two `whole_guest_backup_failed` alert mails, 04:30 and 04:32 CEST.
**Outcome:** root cause found on **ep0**, fixed, both missed backups re-driven and landed.
**No repo code changed** — this was an operational run. Documentation, register and evidence only.
**Outcome:** changelog read → **no connection-handling fix in the range**; operator ruled to upgrade
anyway **for rehearsal value**; upgraded **4.2.2-1 → 4.2.5-1** cleanly; **the fd slope did not change,
which is the predicted result.** Two dated checks filed as **R-341**.
**Both STOPs cleared by the operator.** No code changed in any repo; `documentation/` only.
Evidence: `documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/`.
---
## 1. What was wrong
## 1. Baselines re-confirmed on the machine
Both alerts were the same incident and neither was on a customer box. `demo-felhom` and `demo-hp`
each failed their `felhom-pbs` tier with `Can't connect to 10.77.0.1:8007 (Connection timed out)` —
the one thing they share, the Hetzner offsite PBS endpoint.
Not from the audit file — from `dpkg -l` and `apt-cache policy`, per the runbook.
On ep0, `proxmox-backup-proxy` was `active`, held its listening socket, and **served nobody**:
| | |
|---|---|
| Installed | `proxmox-backup-server` / `-client` **4.2.2-1** |
| Candidate | **4.2.5-1** (4.2.3-1, 4.2.4-1 also available) |
| Proxy PID / started | 542065, **03:54:53Z**, unrestarted since the incident |
| Effective `open files` | **65536 / 65536** — the morning's drop-in in force |
| `Recv-Q` / loopback | **0** / **`200` in 11 ms** |
| Datastore | 3.7 G of 98 G, 4% |
| felhom.eu `main` @ start | `435e044` — matches the runbook's stated baseline |
```
ss -lnt '( sport = :8007 )' → LISTEN Recv-Q 1025 Send-Q 1024
ls /proc/<proxy>/fd | wc -l → 1024 # == its soft RLIMIT_NOFILE
```
## 2. The before-slope — and a correction I owe the morning's report
`Send-Q` on a listener is the accept backlog; `Recv-Q` is the queue depth. At 1025 against 1024 the
queue had overflowed, because `accept()` was returning `EMFILE` on every call. **1016 of the 1024
descriptors were sockets and 547 connections sat in `CLOSE-WAIT`** — a connection leak, fed by
~85,000 requests/day, that reached the ceiling after 14 days of uptime. Last request served:
2026-08-17 18:15:32 UTC. Offsite DR was therefore down **≈ 9 h 37 m**.
Two independent windows on the same proxy generation:
**The observation that settled it:** the daemon was wedged from its own loopback too —
`curl https://127.0.0.1:8007/` on ep0 timed out. A listener that cannot serve `127.0.0.1` has no
network left to blame, and that check is cheap enough to make early.
| window | from → to | delta | rate |
|---|---|---|---|
| 31 min | 09:18:21Z fd=62 → 09:49:46Z fd=66 | +4 | **183/day** |
| 5.64 h | 04:11:36Z fd=19 → 09:49:46Z fd=66 | +47 | **200/day** |
Full record, including everything that was ruled out first (tunnel, nftables, disk, dead daemon):
**`documentation/audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md`**.
**This morning's incident note said ~85/day and "≈2 years of runway". Both were wrong.** They were
extrapolated from a single 17-minute window whose delta was **one descriptor** — a sample of one
cannot carry a daily rate, and the agreement with the historical ~73/day that made it feel solid was
coincidence. **The real rate is ~185–200/day, ~2.6× what I published, and the runway is ~357 days,
not two years.** Corrected in the incident document and in R-336 rather than left standing.
## 2. What was done
**The mechanism I named was also the minority one.** `CLOSE-WAIT` held flat at **1** across the
window while `ESTAB` grew **45 → 49** — *all* the growth was established connections. At the wedge
the split was **1011 ESTAB / 543 CLOSE-WAIT**, so ESTAB dominated there too. **R-336's fix must target
connections the proxy never reaps, not just `CLOSE-WAIT` sockets.**
1. `LimitNOFILE=65536` drop-ins for `proxmox-backup-proxy.service` and `proxmox-backup.service`, each
carrying its reason inline. **The API daemon was not implicated** (15 fds) and its drop-in says so
— a later reader must not mistake it for a second culprit.
2. Restarted both. After: soft limit 65536, fds back to 18, `Recv-Q 0`, loopback `200`.
3. **Verified from the customer side, not only from ep0** — both boxes got `200` in ~0.1 s and
`pvesm status` read `felhom-pbs pbs active`.
4. Re-drove the missed backups **through the product path** — `POST /backup?target=felhom-pbs` on
each agent's local API, issued from inside the guest's controller container with the controller's
own credentials, i.e. the same call the scheduler makes. Not a hand-run `vzdump`.
## 3. The changelog, verbatim — the run's primary deliverable
**Result:** `demo-felhom` → `ct/9201/2026-08-18T03:57:43Z` (4.10 GB, 36.4 s);
`demo-hp` → `ct/9201/2026-08-18T03:58:43Z` (4.29 GB, 41.5 s). Both directories carry a full manifest
on ep0, and both hosts log the re-run vzdump as `OK`. **And the hub agrees** — both boxes' next host
reports carry `felhom-pbs success=true` (04:00:33Z and 04:07:35Z), so the operator view went green on
the evidence rather than on the fact that a restart was performed.
**No data was lost and no backup was skipped** — the daily local
tier was never affected (it completed on both boxes at 05:00 and 05:02), and the PBS tier is weekly,
so the outage window cost exactly one attempt, which was re-driven the same morning.
All three entries between 4.2.2-1 and 4.2.5-1 read in full (128 lines), then swept for
`connection|file descriptor|fd|accept(|close_wait|keep-alive|socket|EMFILE|nofile|leak|proxy|listen|
backlog|hyper|tokio`.
Evidence copied off ep0 **before** the restart, per standing rule 5:
`documentation/audits/evidence-ep0-fd-2026-08-18/` — pre-restart state, post-fix state, access-log tail.
**Exactly one hit, and it is a false positive:**
## 3. What I got wrong, and corrected
> `* S3: config: allow editing the use-node-config flag that controls whether requests S3 endpoints`
> ` honor the node's proxy settings or not`
- **`proxmox-backup-manager version` prints *available* then *running*.** It read
`4.2.5-1 running version: 4.2.2`, which looks exactly like a daemon left behind by a package
upgrade. It is not: `dpkg -l` shows **4.2.2-1 installed**, 4.2.5-1 merely available in the repo,
and the on-disk binary is dated 2026-06-18. I restarted the API daemon on that mistaken reading;
harmless, and its drop-in is a genuine improvement, but it was not needed.
- **I addressed `demo-hp`'s agent at the island address** because `operations/nodes.md` says that box
is island-migrated. It is not (**R-338**), and the resulting timeout was briefly read as a fault.
HTTP-proxy configuration *for S3 requests* — not the `proxmox-backup-proxy` daemon.
## 4. Findings filed — R-336, R-337, R-338
What the range does contain. **4.2.5-1** — a security release hardening client-supplied manifests:
All three are in `documentation/backlog/OPEN-ITEMS.md` with numbers, per the registers-first rule.
> `* backup: harden the handling of client supplied backup manifests:`
> ` - only accept archive names that are plain file names carrying a server side type extension. A`
> ` crafted name in a manifest could previously make a sync job read or write outside of the`
> ` snapshot directory, running as the unprivileged 'backup' user.`
> ` - keep an uploaded manifest in memory and only persist it on backup finish, checking that every`
> ` archive it lists was really uploaded during that session and that the checksums match`
- **R-336 — the poll rate is the real defect.** ~1 request/second against a DR endpoint written to
weekly. `LimitNOFILE` raises the ceiling; **it does not fix the leak**, it converts a fortnightly
outage into a multi-year one.
- **R-337 — a status endpoint that trailed its own artifact, then caught up. WATCHING, not a defect.**
`demo-hp`'s `/backup/status` was still serving the superseded 03:27:00Z failure at ~04:03Z while the
snapshot sat on ep0 and the host logged `OK`; `demo-felhom` updated within ~40 s. **I filed this as
a defect and that was premature** — the next host report (04:07:35Z) carried the success and the
skew cleared with no intervention. Rewritten as WATCHING, with the explicit instruction not to open
a fix until someone establishes whether this is just collection cadence. Recorded at all because
during the recovery it read as a second failure, and it was not one.
- **R-338 — `demo-hp` is not on the R-50 island and `nodes.md` says it is.** No `island_bridge` keys,
guest has no `eth1`, `vmbr9` has zero members, and the agent's local API is bound to the customer
LAN — the exposure R-50 existed to remove.
plus a sync/push chunk-reuse fix and a subscription-key architecture check. **4.2.4-1** — S3 rate
limits, a file-locking user-lookup cache, the new `proxmox-enterprise-support-keyring` dependency,
docs. **4.2.3-1** — UI/journal work, an LDAP search-filter escape, tape and timezone fixes.
**Not done, deliberately:** PBS 4.2.5-1 was not applied. Upgrading a production offsite endpoint was
outside what this run was authorised to do, and its changelog should be read for the connection-
handling leak first.
**Nothing addresses descriptor lifetime or connection reaping. My recommendation was: do not upgrade
for this reason.**
## 5. Gates — one pre-existing conviction, and a stated bypass
## 4. STOP 1 — the ruling
`python3 scripts/repo_gates.py --fast` → **rc=1, CONVICTED: golden-currency.** Eight of nine gates
pass. The conviction is **R-334, inherited and not caused here**: newest released controller
**0.216.0**, newest golden bake **0.214.0**, so a new install misses two releases. The gate reads
`felhom-controller/CHANGELOG.md` and `documentation/tests/golden-*` — **this session touched
neither**, and its whole diff is documentation. Baking is possible; **vouching is
operator-password-gated, and a baked-but-unvouched golden is worse than none**, so it is not a
one-sided job CC can finish.
**Operator (Viktor) ruled: upgrade anyway, for rehearsal value** — *"see how that works for us, we
need practice with that too"*. Legitimate and recorded as such: this was **a practice run of the
upgrade procedure on a Tier-2 protected machine, not a fix for the leak**.
**This push therefore used `git push --no-verify`, stated here per `.claude/rules/gates.md`.**
R-334 is updated in the register with the new numbers rather than left reading 0.215.0.
**The interpretation was fixed in writing before any numbers existed** (`stop1-ruling.txt`):
unchanged slope = **expected**, not a failed upgrade; changed slope = a **surprise** needing
explanation, not a confirmation. That file was written at the ruling, not afterwards, so neither
outcome could be rationalised into a success.
**CI checked by run ID, as the checklist requires: run `348`, `head_sha ebfd0967c`, conclusion
`failure`, elapsed 13 s** (04:09:45→04:09:58Z). Expected and inherited — CI's only step is
`python3 scripts/repo_gates.py --fast`, the same entry point that convicts golden-currency locally,
with the sibling repos fetched. The 13 s runtime places it in the workflow's own "honest gate
failure" band rather than the R-265 reap band, so the result is the gate speaking, not the runner.
**I could not read the run log to name the gate from CI's own mouth** — `actions/runs/348/logs` and
`actions/tasks/348/logs` both 404, `actions/runs/348/jobs` returns an empty list, authenticated as
`admin`, and the web log endpoint 302s. So this is an inference from the local run plus the workflow
definition, not a direct reading, and it is stated as such.
## 5. STOP 2 — the snapshot, and what it does not cover
**Expect one `[felhom CI] gates FAILED in admin/felhom.eu` mail for run 348** — the workflow alarms
on failure by design. It is this push, and it is the golden-currency row, not a new fault; the same
mails on 12 and 14 August have the same cause.
Snapshot **421440873** `felhom-hetzner-20260818`, 15.06 GB, status **Available** (complete, not
merely started), server #147604682, project 15217960.
*(Noted for accuracy: the first gate run was piped to `tail`, which returned `rc=0` — `tail`'s exit
code, not the gate's. It was re-run unpiped to read the real `rc=1`. That is standing rule 1's trap
in its smaller form, and the number reported above is the unpiped one.)*
**It covers `/dev/sda` only.** `/mnt/pbs-datastore` is `/dev/sdb`, a separate 100 GB **Volume**, and
Hetzner server snapshots exclude attached volumes — so this is a rollback for the *software* state
(packages, unit files, the `LimitNOFILE` drop-ins, nftables, wg) and **not a backup of the backup
data**. Fine for a package install that writes no datastore content; **it must not be remembered as
datastore protection.** Taken on a running server, deliberately: powering off ep0 to guard a userspace
package install would take the only off-premises copy offline.
## 6. What to watch
## 6. The upgrade and its verification
The positive observable is the descriptor count, not the absence of an alert — an empty alert queue
is equally consistent with "healthy" and "wedged again":
Simulated first (`-s`): **0 to remove**, so the abort condition never triggered. Then
`apt-get install --only-upgrade -y proxmox-backup-server`, **09:51:00→09:51:06Z, exit 0**. Upgraded
server/client/docs to 4.2.5-1 plus one new dependency, `proxmox-enterprise-support-keyring 1.1` —
**which the 4.2.4-1 changelog had declared**, a small real consistency check between what I read and
what apt did.
```bash
ssh root@<ep0> 'PID=$(systemctl show proxmox-backup-proxy -p MainPID --value); \
ls /proc/$PID/fd | wc -l; ss -lnt "( sport = :8007 )"'
```
| check | result |
|---|---|
| installed | server / client / docs **4.2.5-1** |
| daemons | `proxmox-backup-proxy` **active running**, `proxmox-backup` **active running** |
| proxy restarted | 542065 → **551655** @ 09:51:04 |
| **effective `open files`** | **65536 / 65536** — survived the new package |
| drop-ins on disk | both present, unmodified |
| `Recv-Q` | **0** |
| loopback | **`200` in 12 ms** |
| `felhom-pve` over tunnel | **`200` in 0.103 s**, `felhom-pbs active` |
| `demo-hp` over tunnel | **`200` in 0.096 s**, `felhom-pbs active` |
| hub gauge, post-upgrade | `11:59:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)` |
Healthy is ~20 fds and `Recv-Q 0`. **A count climbing between restarts means R-336's leak is still
live.**
`proxmox-backup-manager version` now reads `4.2.5-1 running version: 4.2.5`. **This independently
settles the morning's confusion**: that string was never reporting a stale daemon, and now that
installed and running genuinely match, both halves agree.
**One false alarm, mine.** I queried `systemctl is-active proxmox-backup-api` and got `inactive`.
**That unit does not exist** — `systemctl cat` returns *"No files found for
proxmox-backup-api.service"*. The real pair is `proxmox-backup-proxy.service` ("API Proxy Server")
and `proxmox-backup.service` ("API Server"), both active. A bad query, not a fault — recorded because
for as long as it took to check, it looked exactly like one.
**No backup, restore or verify was triggered** to "prove" the endpoint, per the runbook: the reads
above answer it without mutating a protected datastore.
## 7. The after-slope — unchanged, as predicted
| | window | delta | rate |
|---|---|---|---|
| before (PID 542065) | 09:18:21Z fd=62 → 09:49:46Z fd=66 | +4 / 1885 s | **183/day** |
| after (PID 551655) | 09:51:22Z fd=17 → 10:23:21Z fd=22 | +5 / 1919 s | **225/day** |
**These are not distinguishable.** The windows differ by **one descriptor**; Poisson uncertainty on
n=4 is ±2 and on n=5 is ±2.2, so both are consistent with a single unchanged rate. **The after-figure
being numerically higher is noise, not a regression — and certainly not an improvement.** Composition
repeats the pattern: **ESTAB 0 → 5, CLOSE-WAIT 0 → 1.**
**Thirty minutes cannot settle this in either direction, and nothing here claims it does.**
## 8. Register
- **R-341 filed, WATCHING** — the two dated checks: **+24 h (2026-08-19 ~10:00Z)** and
**+7 d (2026-08-25 ~10:00Z)**, against new `t0` **fd=17 @ 09:51:22Z, PID 551655**, with the exact
command and the instruction to **record the ESTAB/CLOSE-WAIT split, not just the total** — the
split is what identifies which leak it is. If the PID has changed, the window is void.
- **R-336 updated and STILL OPEN.** Its wrong 85/day baseline is corrected to 183–200/day, the
mechanism is re-pointed at ESTAB, and the upgrade's null result is recorded. **It stays open on its
own merits:** an upgrade that *had* fixed the leak still would not make ~85,000 requests/day to a
weekly-write DR endpoint correct.
- `documentation/runbooks/offsite-endpoint.md` — **no change made, correctly**: it states no PBS
version literal, only package names, so there was nothing to update (and per `docs.md`, version
literals do not belong in a current-state doc anyway).
## 9. Teardown
**This run provisioned nothing.** No `.deb` was downloaded — `apt-get changelog` served the text
directly, so the runbook's `/tmp` cleanup was never needed. The Hetzner snapshot **is retained** as
the rollback; deleting it is an operator decision, and it costs €0.018161/GB/month.
## 10. Observations, not acted on
- **`pvesm status` reports `felhom-pbs` with Total/Used/Available all `0`** on both boxes while
status reads `active`. Consistent before and after the upgrade, and the hub's own gauge reads the
real 3.7 GB / 97.9 GB, so nothing is broken — but the PVE-side numbers are not usable as a capacity
signal. Likely the namespace-scoped token lacking datastore-level audit. Not filed; noted here.
- **8 other packages are held back** (`8 not upgraded`), untouched deliberately — `full-upgrade` on
this machine is forbidden by the runbook and would change the kernel and WireGuard alongside the
thing under test.
- **The morning's `--no-verify` situation is unchanged**: `golden-currency` still convicts on the
inherited R-334 (controller 0.216.0 vs golden 0.214.0), untouched by this run.
+11 -4
View File
@@ -1,6 +1,6 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-08-18 (early — a listening socket that served nobody; off-site backups were down for
**Updated 2026-08-18 (midday — a listening socket that served nobody, then a rehearsal upgrade that changed nothing; off-site backups were down for
9½ hours overnight and are back).**
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
@@ -50,8 +50,8 @@ record with no machine** — created 13 August, no host, no backups, nothing to
were re-run the same morning and are on the off-site box**. **Nothing was lost and nothing was
skipped:** the daily copies on the machines themselves were never affected, and the off-site copy is
weekly, so exactly one attempt fell in the window. The underlying cause — we ask that box a question
about once a second, all day — is filed and **not yet fixed**; the raised ceiling buys years, not a
cure.
about once a second, all day — is filed and **not yet fixed**; the raised ceiling buys **under a
year** on the corrected measurement, not a cure.
- **The removal now genuinely reverses the installation** (R-316, `installer-v1.28.0` published). The
second reinstall used to hit our own leftover; it was watched failing on the cycle that actually
@@ -86,7 +86,14 @@ record with no machine** — created 13 August, no host, no backups, nothing to
for a box we actually write to once a week. That volume is what turned a slow internal leak into
last night's outage in a fortnight. The higher ceiling makes it rare, not impossible — **the leak
itself is untouched.** The honest health check is the resource count climbing, not the absence of an
alarm.
alarm. **Corrected later the same morning: my first estimate of how fast it leaks was too
optimistic by about 2.5×** — measured properly it is under a year to the new ceiling, not two years.
A deadline, not a comfort.
- **The off-site box was updated, and it did not help — as expected** (R-341). On your ruling we
installed the newer backup software for the practice, having first read its release notes and found
**nothing** about the fault we have. The update went cleanly and everything works, but the leak
behaves exactly as before, which is the result the release notes predicted. **Two dated checks are
booked — 19 August and 25 August** — because half an hour of watching cannot honestly settle it.
- **One machine's status took several minutes to admit a backup had worked** (R-337). `demo-hp` was
still showing this morning's failure for four minutes after the copy was safely on the off-site box;
`demo-felhom` updated in under a minute. **It corrected itself** and both machines now read
@@ -168,8 +168,133 @@ A healthy proxy sits near 20 fds with `Recv-Q 0`. **A rising fd count between re
leak is still live** — an unchanging one after a poll-rate reduction would confirm the fix.
**It is still live, and it was measured rather than assumed.** At **16 m 43 s** after the restart the
proxy held **19 fds** (from 18) with **1** connection in `CLOSE-WAIT`. One descriptor per ~17 minutes
is **≈ 85/day** — which lands on the historical rate implied by the failure itself, 1016 sockets over
14 days ≈ **73/day**. Two independent estimates of the same slope agreeing is what makes this a
measurement instead of a story, and it puts the next ceiling at roughly **2 years** rather than the
fortnight the old limit gave. **The leak is unfixed; only its period changed.**
proxy held **19 fds** (from 18) with **1** connection in `CLOSE-WAIT`.
> ### ⚠ CORRECTED 2026-08-18 09:50 UTC — the two numbers first published here were wrong
>
> This section originally read *"one descriptor per ~17 minutes is ≈ 85/day … it puts the next
> ceiling at roughly 2 years"*. **Both figures are wrong, and the reason is worth keeping:** they were
> extrapolated from a single 17-minute window whose delta was **one descriptor**. A sample of one
> cannot carry a daily rate, and the agreement with the historical ~73/day that made it feel solid was
> coincidence.
>
> Re-measured on the same proxy generation (PID 542065, unrestarted) during
> `RUNBOOK ep0 PBS upgrade`, two independent windows:
>
> | window | from | to | rate |
> |---|---|---|---|
> | 31 min | 09:18:21Z fd=62 | 09:49:46Z fd=66 | **183/day** |
> | 5.64 h | 04:11:36Z fd=19 | 09:49:46Z fd=66 | **200/day** |
>
> **The real rate is ~185–200/day — about 2.6× what was published — and the runway to the 65536
> ceiling is ~357 days, not two years.** Still an enormous improvement on the fortnight the old limit
> gave, but under a year, so it is a deadline rather than a comfort.
>
> **And the mechanism named above is the minority one.** Across that window `CLOSE-WAIT` held flat at
> **1** while `ESTAB` grew **45 → 49**: *all* the growth was established connections. At the wedge the
> split was **1011 ESTAB / 543 CLOSE-WAIT**, so ESTAB was the larger half there too and this document
> put its emphasis on the wrong one. **R-336's fix must target connections the proxy never reaps, not
> only sockets left in `CLOSE-WAIT`.**
**The leak is unfixed; only its period changed.**
---
# Follow-up — 2026-08-18, later the same morning: the PBS upgrade
Run under `RUNBOOK — ep0: read the PBS changelog, then decide whether to upgrade`. Two supervised
STOPs, both cleared by the operator. Evidence: `evidence-ep0-pbs-upgrade-2026-08-18/`.
## The changelog said nothing relevant — and that was the finding
The runbook's first job was a pure read: does anything between the installed **4.2.2-1** and the
candidate **4.2.5-1** fix connection handling? All three intervening entries were read in full (128
lines) and swept for `connection|file descriptor|fd|accept(|close_wait|keep-alive|socket|EMFILE|
nofile|leak|proxy|listen|backlog|hyper|tokio`.
**Exactly one keyword hit, and it is a false positive:**
> `* S3: config: allow editing the use-node-config flag that controls whether requests S3 endpoints
> honor the node's proxy settings or not`
That is HTTP-proxy configuration *for S3 requests*, not the `proxmox-backup-proxy` daemon. What the
range actually contains: **4.2.5-1** is a security release hardening client-supplied backup manifests
(an archive name in a crafted manifest could make a sync job read or write outside the snapshot
directory as the `backup` user; manifests are now held in memory and only persisted at backup finish
with per-archive checksum verification), plus a sync/push chunk-reuse fix. **4.2.4-1** is S3 rate
limits, a file-locking user-lookup cache, docs. **4.2.3-1** is UI/journal work, an LDAP search-filter
escape, tape and timezone fixes.
**Nothing addresses descriptor lifetime or connection reaping.** The recommendation was therefore
**do not upgrade for this reason**.
## What was done anyway, and why that is fine
**Operator ruling at STOP 1: upgrade regardless, for rehearsal value** — *"see how that works for us,
we need practice with that too"*. So this was executed as **a practice run of the upgrade procedure on
a Tier-2 protected machine, not as a fix for the leak**, and the distinction was written into
`stop1-ruling.txt` *before* any numbers existed, precisely so the outcome could not be rationalised
afterwards in either direction.
STOP 2 cleared with Hetzner snapshot **421440873** `felhom-hetzner-20260818`, 15.06 GB, status
**Available**. **That snapshot covers `/dev/sda` only.** `/mnt/pbs-datastore` is `/dev/sdb`, a separate
100 GB Volume, and Hetzner server snapshots exclude attached volumes — so it is a rollback for the
software state and **not** a backup of the backup data. Acceptable here because a package install
writes no datastore content; it must not be remembered as datastore protection.
## Result
`apt-get install --only-upgrade proxmox-backup-server`, 09:51:00→09:51:06Z, exit 0. A `-s` simulation
was run first and reported **0 to remove**, so the runbook's abort condition never triggered. Upgraded
server/client/docs to **4.2.5-1**, plus one genuinely new dependency, `proxmox-enterprise-support-
keyring 1.1` — which the 4.2.4-1 changelog had declared, a small but real consistency check between
what was read and what apt did.
| check | result |
|---|---|
| installed (`dpkg -l`) | server / client / docs **4.2.5-1** |
| daemons | `proxmox-backup-proxy` **active running**, `proxmox-backup` **active running** |
| proxy restarted | PID 542065 → **551655** @ 09:51:04 |
| **effective `open files`** | **65536 / 65536** — the drop-in survived the new package |
| `Recv-Q` | 0 |
| loopback | `200` in 12 ms |
| from `felhom-pve` | `200` in 0.103 s, `felhom-pbs active` |
| from `demo-hp` | `200` in 0.096 s, `felhom-pbs active` |
| hub gauge, post-upgrade | `11:59:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)` |
`proxmox-backup-manager version` now reads `4.2.5-1 running version: 4.2.5`. **This independently
settles the confusion recorded in the original incident**: the string was never reporting a stale
daemon, and now that installed and running genuinely match, both halves agree.
**One false alarm, mine:** `systemctl is-active proxmox-backup-api` returned `inactive`. That unit
does not exist — `systemctl cat` says *"No files found for proxmox-backup-api.service"*. The real pair
is `proxmox-backup-proxy.service` ("API Proxy Server") and `proxmox-backup.service` ("API Server"),
both active. A bad query, not a fault, and it is written down because it looked exactly like a fault
for as long as it took to check.
## The slope: unchanged, as predicted
| | window | delta | rate |
|---|---|---|---|
| **before** (PID 542065) | 09:18:21Z fd=62 → 09:49:46Z fd=66 | +4 / 1885 s | **183/day** |
| **after** (PID 551655) | 09:51:22Z fd=17 → 10:23:21Z fd=22 | +5 / 1919 s | **225/day** |
**These are not distinguishable.** The two windows differ by a single descriptor; Poisson uncertainty
on n=4 is ±2 and on n=5 is ±2.2, so both are consistent with one unchanged underlying rate. **The
after-figure being numerically higher is noise, not a regression — and emphatically not an
improvement.** This is the expected outcome and it matches the changelog: no mechanism, no change.
Composition after the upgrade repeats the pattern that matters: **ESTAB 0 → 5, CLOSE-WAIT 0 → 1.**
**Thirty minutes cannot settle this**, in either direction, and this document does not claim it does.
The honest checks are **+24 h (2026-08-19 ~10:00Z)** and **+7 d (2026-08-25 ~10:00Z)** against the
new `t0` of **fd=17 at 09:51:22Z, PID 551655** — filed as **R-341**.
## What this run did not do
Did not reduce the poll rate (**R-336 stays open** — an upgrade that fixed the leak still would not
make ~85,000 requests/day to a weekly-write DR endpoint correct), did not touch the `LimitNOFILE`
drop-ins, did not run `full-upgrade` or touch the kernel, did not trigger a backup/restore/verify to
"prove" the endpoint, and did not change anything on either customer box. Nothing was provisioned, so
there is nothing to tear down; no `.deb` was downloaded, as `apt-get changelog` served the text
directly.
@@ -0,0 +1,105 @@
### capture-time
2026-08-18T09:18:21+00:00
### dpkg -l (INSTALLED — the truth)
ii proxmox-backup-client 4.2.2-1 amd64 Proxmox Backup Client tools
ii proxmox-backup-server 4.2.2-1 amd64 Proxmox Backup Server daemon with tools and GUI
### apt-cache policy proxmox-backup-server
proxmox-backup-server:
Installed: 4.2.2-1
Candidate: 4.2.5-1
Version table:
4.2.5-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.2.4-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.2.3-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
*** 4.2.2-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
100 /var/lib/dpkg/status
4.2.1-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.2.0-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.13-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.12-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.11-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.10-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.9-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.8-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.7-2 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.6-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.5-2 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.5-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.4-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.2-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.1-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.1.0-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.22-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.21-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.20-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.19-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.18-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.17-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.16-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.15-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.14-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.13-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.12-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.11-4 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.11-2 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.10-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.9-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.8-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.7-1 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
4.0.6-2 500
500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages
### proxmox-backup-manager version (available THEN running — misleading, see incident)
proxmox-backup-server 4.2.5-1 running version: 4.2.2
### proxy MainPID
542065
### fd count
62
### limits
Max open files 65536 65536 files
### listen queue
State Recv-Q Send-Q Local Address:Port Peer Address:Port
LISTEN 0 1024 *:8007 *:*
### close-wait count
1
### proxy start time
Tue Aug 18 03:54:53 2026
### datastore
Filesystem Size Used Avail Use% Mounted on
/dev/sdb 98G 3.7G 95G 4% /mnt/pbs-datastore
@@ -0,0 +1,15 @@
### loopback probe
2026-08-18T09:18:40+00:00
http=200 time_total=0.011537
### fd types
54 socket:anon
3 anon_inode:anon
1 0
1 /var/log/proxmox-backup/api/auth.log
1 /var/log/proxmox-backup/api/access.log
1 /var/lib/proxmox-backup/rrdb/rrd.journal
1 /mnt/pbs-datastore/.lock
1 /dev/null
### socket states on 8007
45 ESTAB
1 LISTEN
@@ -0,0 +1,16 @@
### step2-second-reading
2026-08-18T09:49:46+00:00
### MainPID
542065
### fd count
66
### close-wait
1
### socket states
49 ESTAB
1 LISTEN
### listen queue
State Recv-Q Send-Q Local Address:Port Peer Address:Port
LISTEN 0 1024 *:8007 *:*
### proxy start
Tue Aug 18 03:54:53 2026
@@ -0,0 +1,18 @@
BEFORE-SLOPE — proxy PID 542065, started 2026-08-18 03:54:53 UTC, never restarted since
Step-2 window (31 min, this run)
09:18:21Z fd=62 -> 09:49:46Z fd=66
interval 0.5236 h (1885 s), delta 4 fd
= 7.64 fd/hour = 183.3 fd/DAY
Long window (incident t0+17m -> now)
04:11:36Z fd=19 -> 09:49:46Z fd=66
interval 5.6361 h (20290 s), delta 47 fd
= 8.34 fd/hour = 200.1 fd/DAY
Runway from fd=66 at 183/day to the 65536 soft limit:
357 days = 0.98 years
Historical rate implied by the outage: 1016 sockets over 14.29 d = 71.1/day
Socket composition: ESTAB 45 -> 49 (+4); CLOSE-WAIT 1 -> 1 (+0).
=> ALL of the growth in this window is ESTABLISHED connections, not CLOSE-WAIT.
@@ -0,0 +1,200 @@
Get:1 https://metadata.cdn.proxmox.com rust-proxmox-backup 4.2.5-1 Changelog [178 kB]
rust-proxmox-backup (4.2.5-1) trixie; urgency=medium
* backup: harden the handling of client supplied backup manifests:
- only accept archive names that are plain file names carrying a server
side type extension. A crafted name in a manifest could previously make
a sync job read or write outside of the snapshot directory, running as
the unprivileged 'backup' user. Reaching this needed a manifest from a
configured sync remote or from a client that already had backup access
to the datastore.
- keep an uploaded manifest in memory and only persist it on backup
finish, checking that every archive it lists was really uploaded during
that session and that the checksums match the ones computed server
side. A client uploading a manifest that references archives it did not
upload now gets an error on finish instead of such a snapshot being
created.
* fix #7878: sync: push: reuse the manifest of a previous snapshot on a
non-encrypting push if the source snapshot was encrypted with a matching
key, restoring chunk reuse and thus avoiding needlessly long sync runs.
* sync: push: keep the sign-only crypt mode of a source archive instead of
reducing it to unencrypted when pushing without server side encryption.
* subscription: reject a subscription key issued for a different
architecture than the host, as arm64 keys carry an explicit marker, so a
wrong key fails fast instead of only erroring during the online check.
* update to proxmox-upgrade-checks 1.1, which accepts the 7.0 kernel, tells
a bookworm backport apart from a trixie build and fixes the dkms check.
* docs: clarify in the backup protocol description that the manifest is
uploaded by the client and only persisted on backup finish.
-- Proxmox Support Team <support@proxmox.com> Wed, 05 Aug 2026 18:25:37 +0200
rust-proxmox-backup (4.2.4-1) trixie; urgency=medium
* docs: document the debug symbol repository
* datastore: fix wrong local path used for S3 bad chunk handling during
garbage collection
* refactor file creation/mode/ownership helpers to proxmox-product-config
crate
* fix #7642: avoid expensive user lookups on file locking by caching the
backup user/group ID
* depend on proxmox-enterprise-support-keyring, and track its version in the
package version API endpoint
* fix #5748: docs: add `catalog.pcat1` format specification
* docs: system requirements: document we recommend local storage
* S3: fix #6841: allow configuring request rate limits by updating to
proxmox-s3-client 1.4.1. these rate limits are split into active and
passive methods, allowing separate handling of POST/PUT/DELETE and GET/HEAD
request limits.
* S3: config: allow editing the use-node-config flag that controls whether
requests S3 endpoints honor the node's proxy settings or not
* sync: push: gracefully handle previous manifest signature mismatches, which
can happen when enabling or disabling push-encryption on an already synced
backup group
-- Proxmox Support Team <support@proxmox.com> Wed, 29 Jul 2026 15:08:15 +0200
rust-proxmox-backup (4.2.3-1) trixie; urgency=medium
* css: remove x-grid-row-loading class, replace it with non-blurry SVG
variant from proxmox-widget-toolkit
* pbs-client: add backoff log throttle, to ensure progress and similar output
appears quickly initially, but does not create overly long logs
* client: report progress during restore
* api: journal: adopt proxmox-syslog-api and stream the output, making the
implementation consistent with the one from Proxmox Datacenter Manager
* ui: enable the structured journal view and per-service logs, including
colored output and filtering capabilities
* ui: always use arrays for 'delete' property, instead of manually converting
* fix #5971: tape: don't warn on custom MAM attribute write failures
* fix #7175: api: time: use timedatectl instead of /etc/timezone
* fix #7187: report: add ethtool output for physical interfaces
* prune jobs: schedule jobs that do not prune anything, but warn during their
execution. such jobs allow testing scheduling options, but make no sense
for production use.
* fix #6691: allow search by comment in datastore content, make search
case-insensitive and correctly reset content view after empty searches
* ui: datastore: disable various action tooltips for actions which cannot be
triggered
* ldap: escape the user-provided user name when using it in the LDAP search
filter that looks up the user DN.
* ldap sync: log which user properties change when synchronizing an existing
user, instead of only reporting that the user was updated.
* api schema/section config: add support for declaring deprecated property
aliases, to allow renaming properties without showing the old name in the
documentation
* rest server: accept deprecated property aliases in JSON request bodies by
rewriting them to the canonical name before verification and dispatch, like
the CLI and query-string handling already do.
* fix #7690: fs: replace_file: close the temporary file before renaming or
unlinking it, fixing the replacement on WORM file systems and avoiding
leftover .fuse_hidden files on FUSE mounts.
* fs: make_tmp_file: append the temporary suffix instead of replacing the
file extension, keeping the original file name intact for easier
debugging of leftover temporary files.
-- Proxmox Support Team <support@proxmox.com> Tue, 14 Jul 2026 12:54:15 +0200
rust-proxmox-backup (4.2.2-1) trixie; urgency=medium
* api: backup: run synchronous chunk-insert operations off the asynchronous
runtime's worker threads. Blocking those threads, most notably during S3
uploads that can wait up to three hours for a chunk lock, could stall the
I/O and timer drivers of the entire runtime and starve other backup
workers.
* client: backup: make the file-based backup more robust against files that
cannot be accessed or that vanish while the backup is running:
- fix #7658: skip a file and log a warning instead of aborting the whole
backup when querying its metadata fails with a permission error,
matching the existing handling of permission errors when opening files.
Such files can also still be excluded explicitly.
- consistently ignore files that disappear during the backup and warn
about them, instead of treating this as a fatal error in some cases.
* tape: backup: fix the command-line group filter, which was passed to the
API under the wrong parameter name and thus had no effect.
* api: do not log a spurious error when listing the files of a snapshot
whose manifest does not exist yet, which is expected while a backup to
that snapshot is still running.
* ui: datastore summary: fix the per-datastore sync and prune job counts,
which were derived from an incorrectly parsed datastore ID; the prune
count in particular was always shown as zero.
* tape: fix two typos in log and informational messages.
-- Proxmox Support Team <support@proxmox.com> Thu, 18 Jun 2026 11:28:25 +0200
rust-proxmox-backup (4.2.1-1) trixie; urgency=medium
* fix #5076: api: support an 'audiences' property on OpenID realms, listing
additional trusted audience values besides the configured client-id.
Improves compatibility with providers that issue tokens with multiple
audiences.
* fix #7562: api/ui: tape: separate the format-media 'load-barcode'
parameter from the existing 'label-text' verification, so the web
interface can load and format empty or previously-unrelated tapes from a
changer slot in one step. The old shared parameter aborted formatting
after a successful load whenever the on-tape label did not match.
* sync: pull: refuse to overwrite a locally encrypted snapshot from an
unencrypted source or one using a different key, and detect content
differences between two unencrypted snapshots that share a backup time.
Previously such mismatches silently triggered a resync that overwrote the
local snapshot.
* datastore: fix tuning option changes not propagating the updated sync
level to the chunk store until the service was restarted.
* datastore: improve the error message when prune cannot acquire a snapshot
lock for deletion, by showing the snapshot directory and lock file paths
instead of an internal debug dump.
* api: backup: fix benchmark, finish-failed, and backup-failed cleanups
leaving an orphaned empty backup group behind on the datastore.
* api: node: tasks status: return the task end time as an optional field
once the task is finished, so the task viewer can render the correct
duration without an extra API call.
* api/ui: node: add a 'location' property to the node config, exposed
through the node options panel.
* subscription: reuse the server ID from an existing subscription info when
multiple candidates are detected, falling back to the first candidate only
when no prior info exists.
@@ -0,0 +1,128 @@
rust-proxmox-backup (4.2.5-1) trixie; urgency=medium
* backup: harden the handling of client supplied backup manifests:
- only accept archive names that are plain file names carrying a server
side type extension. A crafted name in a manifest could previously make
a sync job read or write outside of the snapshot directory, running as
the unprivileged 'backup' user. Reaching this needed a manifest from a
configured sync remote or from a client that already had backup access
to the datastore.
- keep an uploaded manifest in memory and only persist it on backup
finish, checking that every archive it lists was really uploaded during
that session and that the checksums match the ones computed server
side. A client uploading a manifest that references archives it did not
upload now gets an error on finish instead of such a snapshot being
created.
* fix #7878: sync: push: reuse the manifest of a previous snapshot on a
non-encrypting push if the source snapshot was encrypted with a matching
key, restoring chunk reuse and thus avoiding needlessly long sync runs.
* sync: push: keep the sign-only crypt mode of a source archive instead of
reducing it to unencrypted when pushing without server side encryption.
* subscription: reject a subscription key issued for a different
architecture than the host, as arm64 keys carry an explicit marker, so a
wrong key fails fast instead of only erroring during the online check.
* update to proxmox-upgrade-checks 1.1, which accepts the 7.0 kernel, tells
a bookworm backport apart from a trixie build and fixes the dkms check.
* docs: clarify in the backup protocol description that the manifest is
uploaded by the client and only persisted on backup finish.
-- Proxmox Support Team <support@proxmox.com> Wed, 05 Aug 2026 18:25:37 +0200
rust-proxmox-backup (4.2.4-1) trixie; urgency=medium
* docs: document the debug symbol repository
* datastore: fix wrong local path used for S3 bad chunk handling during
garbage collection
* refactor file creation/mode/ownership helpers to proxmox-product-config
crate
* fix #7642: avoid expensive user lookups on file locking by caching the
backup user/group ID
* depend on proxmox-enterprise-support-keyring, and track its version in the
package version API endpoint
* fix #5748: docs: add `catalog.pcat1` format specification
* docs: system requirements: document we recommend local storage
* S3: fix #6841: allow configuring request rate limits by updating to
proxmox-s3-client 1.4.1. these rate limits are split into active and
passive methods, allowing separate handling of POST/PUT/DELETE and GET/HEAD
request limits.
* S3: config: allow editing the use-node-config flag that controls whether
requests S3 endpoints honor the node's proxy settings or not
* sync: push: gracefully handle previous manifest signature mismatches, which
can happen when enabling or disabling push-encryption on an already synced
backup group
-- Proxmox Support Team <support@proxmox.com> Wed, 29 Jul 2026 15:08:15 +0200
rust-proxmox-backup (4.2.3-1) trixie; urgency=medium
* css: remove x-grid-row-loading class, replace it with non-blurry SVG
variant from proxmox-widget-toolkit
* pbs-client: add backoff log throttle, to ensure progress and similar output
appears quickly initially, but does not create overly long logs
* client: report progress during restore
* api: journal: adopt proxmox-syslog-api and stream the output, making the
implementation consistent with the one from Proxmox Datacenter Manager
* ui: enable the structured journal view and per-service logs, including
colored output and filtering capabilities
* ui: always use arrays for 'delete' property, instead of manually converting
* fix #5971: tape: don't warn on custom MAM attribute write failures
* fix #7175: api: time: use timedatectl instead of /etc/timezone
* fix #7187: report: add ethtool output for physical interfaces
* prune jobs: schedule jobs that do not prune anything, but warn during their
execution. such jobs allow testing scheduling options, but make no sense
for production use.
* fix #6691: allow search by comment in datastore content, make search
case-insensitive and correctly reset content view after empty searches
* ui: datastore: disable various action tooltips for actions which cannot be
triggered
* ldap: escape the user-provided user name when using it in the LDAP search
filter that looks up the user DN.
* ldap sync: log which user properties change when synchronizing an existing
user, instead of only reporting that the user was updated.
* api schema/section config: add support for declaring deprecated property
aliases, to allow renaming properties without showing the old name in the
documentation
* rest server: accept deprecated property aliases in JSON request bodies by
rewriting them to the canonical name before verification and dispatch, like
the CLI and query-string handling already do.
* fix #7690: fs: replace_file: close the temporary file before renaming or
unlinking it, fixing the replacement on WORM file systems and avoiding
leftover .fuse_hidden files on FUSE mounts.
* fs: make_tmp_file: append the temporary suffix instead of replacing the
file extension, keeping the original file name intact for easier
debugging of leftover temporary files.
-- Proxmox Support Team <support@proxmox.com> Tue, 14 Jul 2026 12:54:15 +0200
rust-proxmox-backup (4.2.2-1) trixie; urgency=medium
@@ -0,0 +1,30 @@
### apt-get update start
2026-08-18T09:50:43+00:00
Hit:6 http://mirror.hetzner.com/debian/security trixie-security InRelease
Hit:7 http://download.proxmox.com/debian/pbs trixie InRelease
Get:8 http://deb.debian.org/debian trixie-backports InRelease [54.0 kB]
Get:9 http://deb.debian.org/debian-security trixie-security InRelease [43.4 kB]
Get:10 http://mirror.hetzner.com/debian/packages trixie-backports/main amd64 Packages [314 kB]
Get:11 http://deb.debian.org/debian trixie-backports/main Sources [297 kB]
Fetched 858 kB in 0s (6119 kB/s)
Reading package lists...
### SIMULATION (-s): nothing is installed by this
2026-08-18T09:50:44+00:00
Reading package lists...
Building dependency tree...
Reading state information...
The following additional packages will be installed:
proxmox-backup-client proxmox-backup-docs proxmox-enterprise-support-keyring
The following NEW packages will be installed:
proxmox-enterprise-support-keyring
The following packages will be upgraded:
proxmox-backup-client proxmox-backup-docs proxmox-backup-server
3 upgraded, 1 newly installed, 0 to remove and 8 not upgraded.
Inst proxmox-backup-client [4.2.2-1] (4.2.5-1 Proxmox Backup System Debian Repository:stable [amd64])
Inst proxmox-backup-docs [4.2.2-1] (4.2.5-1 Proxmox Backup System Debian Repository:stable [all])
Inst proxmox-enterprise-support-keyring (1.1 Proxmox Backup System Debian Repository:stable [all])
Inst proxmox-backup-server [4.2.2-1] (4.2.5-1 Proxmox Backup System Debian Repository:stable [amd64])
Conf proxmox-backup-client (4.2.5-1 Proxmox Backup System Debian Repository:stable [amd64])
Conf proxmox-backup-docs (4.2.5-1 Proxmox Backup System Debian Repository:stable [all])
Conf proxmox-enterprise-support-keyring (1.1 Proxmox Backup System Debian Repository:stable [all])
Conf proxmox-backup-server (4.2.5-1 Proxmox Backup System Debian Repository:stable [amd64])
@@ -0,0 +1,58 @@
### UPGRADE START
2026-08-18T09:51:00+00:00
Reading package lists...
Building dependency tree...
Reading state information...
The following additional packages will be installed:
proxmox-backup-client proxmox-backup-docs proxmox-enterprise-support-keyring
The following NEW packages will be installed:
proxmox-enterprise-support-keyring
The following packages will be upgraded:
proxmox-backup-client proxmox-backup-docs proxmox-backup-server
3 upgraded, 1 newly installed, 0 to remove and 8 not upgraded.
Need to get 47.8 MB of archives.
After this operation, 1404 kB of additional disk space will be used.
Get:1 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 proxmox-backup-client amd64 4.2.5-1 [3447 kB]
Get:2 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 proxmox-backup-docs all 4.2.5-1 [6351 kB]
Get:3 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 proxmox-enterprise-support-keyring all 1.1 [2988 B]
Get:4 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 proxmox-backup-server amd64 4.2.5-1 [38.0 MB]
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
locale: Cannot set LC_CTYPE to default locale: No such file or directory
locale: Cannot set LC_ALL to default locale: No such file or directory
Fetched 47.8 MB in 1s (39.3 MB/s)
(Reading database ... (Reading database ... 5% (Reading database ... 10% (Reading database ... 15% (Reading database ... 20% (Reading database ... 25% (Reading database ... 30% (Reading database ... 35% (Reading database ... 40% (Reading database ... 45% (Reading database ... 50% (Reading database ... 55% (Reading database ... 60% (Reading database ... 65% (Reading database ... 70% (Reading database ... 75% (Reading database ... 80% (Reading database ... 85% (Reading database ... 90% (Reading database ... 95% (Reading database ... 100% (Reading database ... 43107 files and directories currently installed.)
Preparing to unpack .../proxmox-backup-client_4.2.5-1_amd64.deb ...
Unpacking proxmox-backup-client (4.2.5-1) over (4.2.2-1) ...
Preparing to unpack .../proxmox-backup-docs_4.2.5-1_all.deb ...
Unpacking proxmox-backup-docs (4.2.5-1) over (4.2.2-1) ...
Selecting previously unselected package proxmox-enterprise-support-keyring.
Preparing to unpack .../proxmox-enterprise-support-keyring_1.1_all.deb ...
Unpacking proxmox-enterprise-support-keyring (1.1) ...
Preparing to unpack .../proxmox-backup-server_4.2.5-1_amd64.deb ...
Unpacking proxmox-backup-server (4.2.5-1) over (4.2.2-1) ...
Setting up proxmox-backup-docs (4.2.5-1) ...
Setting up proxmox-backup-client (4.2.5-1) ...
Setting up proxmox-enterprise-support-keyring (1.1) ...
Setting up proxmox-backup-server (4.2.5-1) ...
Processing triggers for man-db (2.13.1-1) ...
### apt exit: 0
### UPGRADE END
2026-08-18T09:51:06+00:00
@@ -0,0 +1,48 @@
### verify-time
2026-08-18T09:51:22+00:00
### dpkg -l (INSTALLED)
ii proxmox-backup-client 4.2.5-1 amd64 Proxmox Backup Client tools
ii proxmox-backup-docs 4.2.5-1 all Proxmox Backup Documentation
ii proxmox-backup-server 4.2.5-1 amd64 Proxmox Backup Server daemon with tools and GUI
ii proxmox-enterprise-support-keyring 1.1 all Proxmox enterprise SSH support keyring
### daemons
active
inactive
active
### proxy MainPID (should DIFFER from 542065 = it restarted)
551655
### proxy start time
Tue Aug 18 09:51:04 2026
### EFFECTIVE open files on the RUNNING proxy (must be 65536)
Max open files 65536 65536 files
### drop-ins still present?
/etc/systemd/system/proxmox-backup-proxy.service.d/:
total 16
drwxr-xr-x 2 root root 4096 Aug 18 03:54 .
drwxr-xr-x 26 root root 4096 Jul 27 07:53 ..
-rw-r--r-- 1 root root 44 Jul 27 07:16 10-datastore-mount.conf
-rw-r--r-- 1 root root 337 Aug 18 03:54 20-nofile.conf
/etc/systemd/system/proxmox-backup.service.d/:
total 16
drwxr-xr-x 2 root root 4096 Aug 18 03:55 .
drwxr-xr-x 26 root root 4096 Jul 27 07:53 ..
-rw-r--r-- 1 root root 44 Jul 27 07:16 10-datastore-mount.conf
-rw-r--r-- 1 root root 237 Aug 18 03:55 20-nofile.conf
### fd count (new t0)
17
### listen queue (Recv-Q must be 0)
State Recv-Q Send-Q Local Address:Port Peer Address:Port
LISTEN 0 1024 *:8007 *:*
### socket states
1 LISTEN
### loopback
http=200 time=0.012245
### version string
proxmox-backup-server 4.2.5-1 running version: 4.2.5
### datastore
+================+====================+=========+
| name | path | comment |
+================+====================+=========+
| felhom-offsite | /mnt/pbs-datastore | |
+================+====================+=========+
@@ -0,0 +1,11 @@
### does proxmox-backup-api exist?
No files found for proxmox-backup-api.service.
### actual PBS units
proxmox-backup-banner.service loaded active exited Proxmox Backup Server Login Banner
proxmox-backup-daily-update.service loaded inactive dead Daily Proxmox Backup Server update and maintenance activities
proxmox-backup-proxy.service loaded active running Proxmox Backup API Proxy Server
proxmox-backup.service loaded active running Proxmox Backup API Server
### API server description
MainPID=551642
Description=Proxmox Backup API Server
ActiveState=active
@@ -0,0 +1,5 @@
### felhom-pve
2026-08-18T11:51:48+02:00
http=200 time=0.103326
Name Type Status Total (KiB) Used (KiB) Available (KiB) %
felhom-pbs pbs active 0 0 0 0.00%
@@ -0,0 +1,5 @@
### demo-hp
2026-08-18T11:51:50+02:00
http=200 time=0.095982
Name Type Status Total (KiB) Used (KiB) Available (KiB) %
felhom-pbs pbs active 0 0 0 0.00%
@@ -0,0 +1 @@
POST-UPGRADE REFRESH: 2026/08/18 11:59:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)
@@ -0,0 +1,3 @@
2026/08/18 11:12:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)
2026/08/18 11:28:30 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)
2026/08/18 11:43:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)
@@ -0,0 +1,18 @@
### step6-second-reading (post-upgrade)
2026-08-18T10:23:21+00:00
### MainPID
551655
### proxy start
Tue Aug 18 09:51:04 2026
### fd count
22
### socket states
5 ESTAB
1 LISTEN
### close-wait
1
### listen queue
State Recv-Q Send-Q Local Address:Port Peer Address:Port
LISTEN 0 1024 *:8007 *:*
### limits
Max open files 65536 65536 files
@@ -0,0 +1,20 @@
AFTER-SLOPE — proxy PID 551655, started 2026-08-18 09:51:04 UTC (the upgrade restart)
09:51:22Z fd=17 (ESTAB 0) -> 10:23:21Z fd=22 (ESTAB 5)
interval 0.5331 h (1919 s), delta 5 fd
= 9.38 fd/hour = 225.1 fd/DAY
BEFORE (31 min window): 183.3/day from delta=4 over 1885 s
AFTER (32 min window): 225.1/day from delta=5 over 1919 s
VERDICT: NOT DISTINGUISHABLE. The two windows differ by ONE descriptor.
With counts this small the Poisson uncertainty on n=4 is +/-2 and on n=5 is +/-2.2,
so both are consistent with a single unchanged underlying rate. The after-figure being
numerically HIGHER is noise, not a regression -- and it is certainly not an improvement.
This is the EXPECTED result: the 4.2.2->4.2.5 changelog contains no mechanism by which
connection reaping would change. Recorded before the numbers existed, in stop1-ruling.txt.
Composition again favours ESTAB: 0 -> 5 ESTAB, 0 -> 1 CLOSE-WAIT.
30 MINUTES CANNOT SETTLE THIS. The honest checks are +24 h and +7 d -> filed as R-341.
@@ -0,0 +1,25 @@
STOP 1 — operator ruling
========================
Recorded: 2026-08-18 ~09:25 UTC (11:25 CEST)
Given by: Viktor (operator), in session.
CC's recommendation was: DO NOT UPGRADE for the stated reason.
Basis: the full changelog range 4.2.2-1 -> 4.2.5-1 (128 lines, all three
entries read) contains NOTHING touching connection handling, descriptor
lifetime, accept(), CLOSE-WAIT, keep-alive or the proxy daemon's socket
lifecycle. An explicit keyword sweep returned exactly one hit, and it is a
false positive ("S3 ... honor the node's proxy settings" = HTTP proxy config
for S3 requests, not the PBS proxy daemon).
OPERATOR RULING: PROCEED WITH THE UPGRADE ANYWAY.
Stated reason: rehearsal value -- "see how that works for us, we need
practice with that too". The upgrade is therefore being performed as a
supervised practice run of the upgrade procedure on a Tier-2 protected
machine, NOT as a fix for R-336's leak.
CONSEQUENCE TO CARRY INTO THE REPORT, so it is not later misread:
this upgrade is NOT expected to change the fd slope. If the post-upgrade
slope differs, that is a surprise requiring explanation, not a confirmation
of anything -- the changelog gives no mechanism by which it should improve.
Equally, if the slope is unchanged, that is the EXPECTED result and is not
evidence the upgrade failed.
@@ -0,0 +1,31 @@
STOP 2 — Hetzner snapshot (the rollback)
========================================
Confirmed by the operator in the Hetzner console, 2026-08-18 ~09:27 UTC (11:27 CEST).
Snapshot ID : 421440873
Description : felhom-hetzner-20260818
Image size : 15.06 GB
Status : Available <-- COMPLETE, not merely started
Created : "less than a minute ago" as displayed at confirmation
Server : felhom-hetzner #147604682 (CX33, 167.233.158.164)
Hetzner project: 15217960 (Felhom.eu)
This is the rollback point for the 4.2.2-1 -> 4.2.5-1 upgrade performed in
Step 4. It was taken on a RUNNING server: Hetzner's own console recommends
powering off first for disk consistency, and that was not done -- a
deliberate trade, because powering off ep0 takes the only off-premises copy
offline and the upgrade being guarded is a userspace package install that
touches neither the datastore nor the boot path.
WHAT THIS SNAPSHOT DOES AND DOES NOT COVER:
covers - the 38 GB system disk /dev/sda (root), i.e. the PBS packages,
unit files, /etc/systemd drop-ins, nftables and wg config.
DOES NOT - /mnt/pbs-datastore. That is /dev/sdb, a separate 100 GB
VOLUME, and Hetzner server snapshots do not include attached
volumes. The backup data is therefore NOT protected by this
snapshot.
Consequence: rolling back restores the software state, not the datastore.
This is acceptable for THIS change because the upgrade writes no datastore
content -- but it must not be mistaken for a datastore backup, and any
future runbook step that could touch /mnt/pbs-datastore needs a different
safeguard than this one.
+2 -1
View File
@@ -636,6 +636,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-333** | **Two disk-health questions the deploy raised and did NOT act on.** **(a) The 55/60 °C bands are SPINNING-DISK bands applied to NVMe.** They were adopted unchanged from the operator's Prometheus config so the two systems cannot disagree — a deliberate, stated decision — but **measured on demo-hp 2026-08-14 the healthy Toshiba KXG50PNV1T02 NVMe idles at 53 °C, two degrees below Figyelmeztetés and seven below Hiba**, and NVMe routinely exceeds 60 °C under load with no fault whatever. As it stands a healthy customer NVMe under sustained write can be reported as **Hiba** — the single worst outcome this feature can produce. **(b) The agent runs bare `smartctl -a -j` with no `-n standby`** (`felhom-agent/internal/storage/hostops.go:368`), so every poll WAKES a spun-down drive; going 6h → hourly multiplies that by six. demo-hp is all-flash so the cadence measurement could not reveal it, and it was recorded rather than acted on per the task's own instruction. Mitigating datum from the fixture: the failing drive logged only **3375 load cycles in 60505 hours** (~one per 18h), i.e. that duty cycle barely spins down at all | **READY (S each) — NEW 2026-08-14** | — | (a) split the temperature bands by device class, or drop them for NVMe and rely on `critical_warning`; (b) add `-n standby` to the agent's smartctl invocation (an agent change, so fold it into R-330's session) | Viktor decides (a); CC does (b) |
| **R-334** | **WAIVER + open item: controller v0.215.0 is released and deployed, and NO golden carries it.** Convicted by `golden_currency_gate.py` on the 2026-08-14 push: newest released controller **0.215.0**, newest golden bake **0.214.0** (`documentation/tests/golden-0.214.0-2026-08-12`). **A machine installed right now receives 0.214.0** — i.e. a brand-new box would ship WITHOUT the R-328 severity fix and would keep emailing nobody about a failing disk. The running fleet is unaffected (demo-hp guest 9201 is on 0.215.0 and healthy); this is purely the day-0 install path. **Not baked in this session deliberately:** the task scoped deployment to demo-hp only, and the second half of the fix — vouching the bake in the hub's day-0 artifact manifest — is **operator-password-gated, so CC cannot complete it**; a baked-but-unvouched golden is worse than none. **The push was made with `git push --no-verify` and it is stated here and in the session report**, per `.claude/rules/gates.md` — the gate has no waiver parser, so recording a waiver does not clear it. **STILL OPEN and now one version WIDER, 2026-08-18:** the newest released controller is **0.216.0** (v0.215.0's own follow-up fix, R-335) and the newest bake is still **0.214.0**, so a new install now misses *two* releases. Re-convicted on this date's documentation-only push, which was likewise made with `--no-verify`; the gate reads `felhom-controller/CHANGELOG.md` and `documentation/tests/golden-*`, **neither of which that session touched** — the conviction is inherited, not caused | **READY (S) — NEW 2026-08-14, re-confirmed 2026-08-18** | operator availability for the vouch step | Bake a golden on **0.216.0** per `runbooks/RUNBOOK-manual-build.md` §4.1, then vouch it — a THREE-field change (`golden_version` + `agent_version` + `min_agent`). Until then every NEW install lacks the severity fix | CC bakes; **Viktor vouches** |
| **R-335** | **One physical disk was walked TWICE per run, and the second walk sustained it against itself.** Found on live hardware ~2h after the v0.215.0 deploy, **by noticing the release's own positive observable disagreed with its own persisted artefact**: the hourly check logged *"3 disk(s) evaluated"* while `disk-health-state.json` held **two** records. Cause: demo-hp's `c11-scratch` and `felhom-backup` are the same physical NVMe (`/dev/nvme0n1`) and resolve to the same `diskKey`. **Not cosmetic** — `RunDiskHealthCheck` writes a disk's new record before the next entry reads it, so the SECOND copy consumed the FIRST copy's write as its prior: the disk **sustained against itself and reached Hiba on a FIRST sighting**, defeating truth-table row 6 — the exact rule separating a one-hour benign excursion from a false critical — and would have emitted **two identical events** for one drive. **Latent, not active, on demo-hp** (all three entries healthy, zero counters), but any aliased disk developing a single pending sector would have gone straight to Hiba. **This is the shape standing rule 3 warns about: an absent alarm was not evidence — the two artefacts had to be read AGAINST each other** | **CLOSED — controller v0.216.0, 2026-08-14.** Each `diskKey` is evaluated once per run; both entries stay marked `seen` so neither looks like a disappeared disk, and the card still renders both storage rows (the dedup is about state and alerts, not display). Pinned by `TestDiskCheck_SameDiskTwiceIsEvaluatedOnce`; companion red-proof run and reverted — deleting the guard makes the first sighting emit `Kind:2` (Hiba-from-sectors) at 8 sectors | — | — | CC |
| **R-336** | **The offsite DR endpoint is polled about once per second, and that is what turned a slow leak into an outage.** ep0's PBS proxy served **~85,000 requests/day** — a flat **3,538/hour**, every hour, from two boxes: `74,445 GET /api2/json/admin/datastore` (`libwww-perl`, i.e. PVE's `pvestatd`) and `73,171 GET /admin/datastore/felhom-offsite/status` (`proxmox-backup-client`). Two pollers asking substantially the same question at the same rate. On 2026-08-18 this walked a connection leak in the proxy to its 1024-fd soft limit in **14 days**, wedging the offsite tier for 9½ hours (`audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md`). The `LimitNOFILE=65536` drop-in applied that morning raises the ceiling **but does not fix the leak** — it converts a fortnightly outage into a multi-year one, which is mitigation, not a fix. A DR endpoint that is written to weekly does not need to be asked about every second | **READY (M) — NEW 2026-08-18** | — | Find what polls `felhom-pbs` this hard (PVE storage status is the prime suspect, and its interval is tunable) and cut it; then confirm the fd count between restarts stops climbing — the positive observable, per standing rule 3. **Baseline measured 2026-08-18: 19 fds at 16m43s post-restart, ~85/day, matching the ~73/day implied by the original failure — the leak is confirmed live, not assumed** | CC |
| **R-336** | **The offsite DR endpoint is polled about once per second, and that is what turned a slow leak into an outage.** ep0's PBS proxy served **~85,000 requests/day** — a flat **3,538/hour**, every hour, from two boxes: `74,445 GET /api2/json/admin/datastore` (`libwww-perl`, i.e. PVE's `pvestatd`) and `73,171 GET /admin/datastore/felhom-offsite/status` (`proxmox-backup-client`). Two pollers asking substantially the same question at the same rate. On 2026-08-18 this walked a connection leak in the proxy to its 1024-fd soft limit in **14 days**, wedging the offsite tier for 9½ hours (`audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md`). The `LimitNOFILE=65536` drop-in applied that morning raises the ceiling **but does not fix the leak** — it converts a fortnightly outage into a multi-year one, which is mitigation, not a fix. A DR endpoint that is written to weekly does not need to be asked about every second | **READY (M) — NEW 2026-08-18** | — | Find what polls `felhom-pbs` this hard (PVE storage status is the prime suspect, and its interval is tunable) and cut it; then confirm the fd count between restarts stops climbing — the positive observable, per standing rule 3. **Baseline measured 2026-08-18, and the FIRST measurement published was WRONG.** The initial "~85/day, matching the ~73/day implied by the failure" came from a single 17-minute window whose delta was **one descriptor** — a sample of one cannot carry a daily rate, and the agreement that made it feel solid was coincidence. **Re-measured over two independent windows the same morning: 183/day (31 min) and 200/day (5.6 h)** — ~2.6x the published figure, putting the runway to the 65536 ceiling at **~357 days, not the ~2 years first claimed**. **And the named mechanism is the minority one:** across that window `CLOSE-WAIT` held flat at 1 while `ESTAB` grew 45→49 — *all* the growth was established connections, and at the wedge the split was 1011 ESTAB / 543 CLOSE-WAIT. **The fix must target connections the proxy never reaps, not just `CLOSE-WAIT` sockets.** The PBS 4.2.5-1 upgrade (2026-08-18) did NOT change the slope and was never expected to — see R-341 | CC |
| **R-337** | **`/backup/status` lagged a completed backup by minutes on one box and not the other — and it RESOLVED ITSELF, which is why this is WATCHING and not a defect.** During the R-336 recovery on 2026-08-18, `demo-hp`'s snapshot landed on ep0 at **03:58:43Z** (complete manifest; the host's own task index says `OK`) — yet `GET /backup/status` was **still serving the superseded 03:27:00Z failure at ~04:03Z**, four-plus minutes later. `demo-felhom` showed its new result within ~40 s of completion. **The lag cleared on its own:** demo-hp's 04:07:35Z host report carries `felhom-pbs success=true, 4.29 GB`, and the hub is green for both boxes. **The first draft of this row claimed the success was "still reported as failed" — that was written before the next report arrived and it was wrong; the corrected claim is a several-minute skew between the two boxes, not a stuck value.** It is recorded because a status field that can trail its own artifact by minutes will, during an incident, be read as a second failure — this session nearly did — and because the asymmetry between the two boxes is unexplained | **WATCHING — NEW 2026-08-18** | another observation, ideally during an incident rather than constructed | **Do not open a fix on this as written.** First establish the intended refresh path for `/backup/status` after an out-of-schedule run; only if the skew is not simply collection cadence is there anything to pin. If it is cadence, close this row and say so | CC |
| **R-338** | **`demo-hp` is not on the R-50 island at all, and `operations/nodes.md` states that it is.** The page records both fleet boxes as island-migrated 2026-07-25. True of `felhom-pve`; **false of `demo-hp`**, whose `agent.json` has `listen_addr: 192.168.0.87:8443` — the customer LAN address — and **no `island_bridge`/`island_guest_addr` keys at all**, whose guest 9201 has `net0` only (no `eth1`), and whose `vmbr9` exists with **zero members**. The controller's `controller.yaml` points at the LAN address, so the box works; this is inventory drift, not breakage. **Two costs.** A session trusting the page addresses the wrong endpoint — that happened on 2026-08-18 and the resulting timeout was briefly read as a fault. And the agent's local API is **bound to the customer LAN on this box** rather than to a point-to-point island, which is the exposure R-50 was built to remove — so a documented security property is claimed for a box that does not have it | **READY (S) — NEW 2026-08-18** | — | Decide which is true: migrate `demo-hp` to the island, or correct `nodes.md`. Leaving both is the one option that keeps the doc lying | Viktor decides; CC executes |
| **R-341** | **Does the fd slope change after the PBS 4.2.5 upgrade? — two dated checks, and the answer is expected to be NO.** ep0 was upgraded 4.2.2-1 → 4.2.5-1 on 2026-08-18 09:51Z on the operator's ruling, **for rehearsal value, not as a fix**: the full changelog range was read (128 lines, all three entries) and swept for connection-handling vocabulary, and it contains **no mechanism** by which descriptor reaping would change — the single keyword hit was `S3 … honor the node's proxy settings`, HTTP-proxy config for S3, not the PBS proxy daemon. **The 32-minute post-upgrade window is indistinguishable from the before window** (+5 fd/1919 s = 225/day vs +4 fd/1885 s = 183/day; the two differ by ONE descriptor and Poisson uncertainty on such counts is ±2, so both are consistent with one unchanged rate — the higher after-figure is noise, not a regression). **Thirty minutes cannot settle it in either direction and this row exists so nobody pretends it did.** New `t0` = **fd 17 at 2026-08-18 09:51:22Z, proxy PID 551655**; before-rate to beat = **183–200/day**. **Interpretation fixed in advance** (`evidence-ep0-pbs-upgrade-2026-08-18/stop1-ruling.txt`, written before any numbers existed): unchanged = EXPECTED, not a failed upgrade; changed = a SURPRISE needing explanation, not a confirmation | **WATCHING — NEW 2026-08-18** | elapsed time only | **Two dated checks, both CC:** **+24 h — 2026-08-19 ~10:00Z** and **+7 d — 2026-08-25 ~10:00Z**. Command (the incident's own positive observable): `ssh root@<ep0> 'PID=$(systemctl show proxmox-backup-proxy -p MainPID --value); ls /proc/$PID/fd \| wc -l; ss -lnt "( sport = :8007 )"; ss -tn state all "( sport = :8007 )" \| awk "NR>1{print \$1}" \| sort \| uniq -c'`. **Record the ESTAB/CLOSE-WAIT split, not just the total** — the split is what says which leak it is. If PID ≠ 551655 the window is void: something restarted the proxy and the count began again | CC on both dates |