From 3e50902a985fa60d16e91b9ab55ce7e888d8e76e Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Tue, 18 Aug 2026 12:26:53 +0200 Subject: [PATCH] RUNBOOK ep0: PBS 4.2.2-1 -> 4.2.5-1, slope unchanged as predicted (R-341) Both STOPs cleared by the operator. No code changed; documentation only. STEP 3 (the run's primary deliverable): the full changelog range 4.2.2-1 -> 4.2.5-1 was read (128 lines, all three entries) and swept for connection-handling vocabulary. Exactly one keyword hit, a false positive ("S3 ... honor the node's proxy settings" = HTTP proxy config for S3, not the PBS proxy daemon). 4.2.5-1 is a manifest-hardening security release; 4.2.4-1 is S3 rate limits and a locking cache; 4.2.3-1 is UI/LDAP/tape. NOTHING addresses descriptor lifetime or connection reaping. Recommendation was: do not upgrade for this reason. STOP 1: operator ruled to upgrade anyway for rehearsal value. Recorded as a practice run, not a fix -- and the interpretation was fixed IN WRITING BEFORE any numbers existed (stop1-ruling.txt): unchanged = expected; changed = surprise. Neither outcome could then be rationalised into a success. STOP 2: Hetzner snapshot 421440873, Available. Documented that it covers /dev/sda ONLY -- /mnt/pbs-datastore is a separate Volume and is NOT in it, so it is a software rollback and not a backup of the backup data. UPGRADE: simulated first (0 to remove), then installed 09:51:00->09:51:06Z, exit 0. Verified: 4.2.5-1 installed, both daemons active, effective open files still 65536 (the drop-in survived the new package), Recv-Q 0, loopback 200, 200 from BOTH boxes over the tunnel with felhom-pbs active, and the hub gauge refreshed post-upgrade at 11:59:31. SLOPE: before +4 fd/1885 s = 183/day; after +5 fd/1919 s = 225/day. NOT distinguishable -- one descriptor apart, Poisson +/-2 on such counts. The higher after-figure is noise, not a regression and not an improvement. 30 minutes cannot settle it; R-341 files the +24 h and +7 d checks. CORRECTIONS to this morning's own report, both published rather than quietly fixed: - the "~85/day, ~2 years of runway" figures were WRONG. They came from a single 17-minute window with a delta of ONE descriptor. Real rate is 183-200/day over two independent windows; runway ~357 days, not 2 years. - the leak was attributed to CLOSE-WAIT. It is mostly ESTAB: CLOSE-WAIT held flat at 1 while ESTAB grew 45->49, and at the wedge it was 1011 ESTAB vs 543 CLOSE-WAIT. R-336's fix must target unreaped connections. - "proxmox-backup-api" reported inactive during verification; that unit does not exist. Bad query, not a fault, written down because it looked like one. R-336 stays open: even a fixed leak would not make ~85k requests/day to a weekly-write DR endpoint correct. golden-currency still convicts (inherited R-334, controller 0.216.0 vs golden 0.214.0, untouched by this run), so this push is --no-verify per .claude/rules/gates.md. --- REPORT.md | 249 ++++++++++-------- STATUS.md | 15 +- ...CIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md | 135 +++++++++- .../step1-state.txt | 105 ++++++++ .../step1b-loopback.txt | 15 ++ .../step2-second-reading.txt | 16 ++ .../step2-slope-computation.txt | 18 ++ .../step3-changelog-attempt1.txt | 200 ++++++++++++++ .../step3-range-4.2.2-to-4.2.5.txt | 128 +++++++++ .../step4a-simulate.txt | 30 +++ .../step4b-upgrade.txt | 58 ++++ .../step5a-verify-ep0.txt | 48 ++++ .../step5b-unit-check.txt | 11 + .../step5c-verify-felhom-pve.txt | 5 + .../step5d-verify-demo-hp.txt | 5 + .../step5e-hub-pbsdr-post.txt | 1 + .../step5e-hub-pbsdr.txt | 3 + .../step6-post-upgrade-slope.txt | 18 ++ .../step6-slope-computation.txt | 20 ++ .../stop1-ruling.txt | 25 ++ .../stop2-snapshot.txt | 31 +++ documentation/backlog/OPEN-ITEMS.md | 3 +- 22 files changed, 1026 insertions(+), 113 deletions(-) create mode 100644 documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step1-state.txt create mode 100644 documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step1b-loopback.txt create mode 100644 documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step2-second-reading.txt create mode 100644 documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step2-slope-computation.txt create mode 100644 documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step3-changelog-attempt1.txt create mode 100644 documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step3-range-4.2.2-to-4.2.5.txt create mode 100644 documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step4a-simulate.txt create mode 100644 documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step4b-upgrade.txt create mode 100644 documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step5a-verify-ep0.txt create mode 100644 documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step5b-unit-check.txt create mode 100644 documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step5c-verify-felhom-pve.txt create mode 100644 documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step5d-verify-demo-hp.txt create mode 100644 documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step5e-hub-pbsdr-post.txt create mode 100644 documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step5e-hub-pbsdr.txt create mode 100644 documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step6-post-upgrade-slope.txt create mode 100644 documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step6-slope-computation.txt create mode 100644 documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/stop1-ruling.txt create mode 100644 documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/stop2-snapshot.txt diff --git a/REPORT.md b/REPORT.md index 6ac6a822..396f1c51 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,133 +1,176 @@ -# REPORT — a listening socket that served nobody (2026-08-18, early) +# REPORT — RUNBOOK ep0: read the PBS changelog, then decide whether to upgrade (2026-08-18, midday) -**Trigger:** two `whole_guest_backup_failed` alert mails, 04:30 and 04:32 CEST. -**Outcome:** root cause found on **ep0**, fixed, both missed backups re-driven and landed. -**No repo code changed** — this was an operational run. Documentation, register and evidence only. +**Outcome:** changelog read → **no connection-handling fix in the range**; operator ruled to upgrade +anyway **for rehearsal value**; upgraded **4.2.2-1 → 4.2.5-1** cleanly; **the fd slope did not change, +which is the predicted result.** Two dated checks filed as **R-341**. + +**Both STOPs cleared by the operator.** No code changed in any repo; `documentation/` only. +Evidence: `documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/`. --- -## 1. What was wrong +## 1. Baselines re-confirmed on the machine -Both alerts were the same incident and neither was on a customer box. `demo-felhom` and `demo-hp` -each failed their `felhom-pbs` tier with `Can't connect to 10.77.0.1:8007 (Connection timed out)` — -the one thing they share, the Hetzner offsite PBS endpoint. +Not from the audit file — from `dpkg -l` and `apt-cache policy`, per the runbook. -On ep0, `proxmox-backup-proxy` was `active`, held its listening socket, and **served nobody**: +| | | +|---|---| +| Installed | `proxmox-backup-server` / `-client` **4.2.2-1** | +| Candidate | **4.2.5-1** (4.2.3-1, 4.2.4-1 also available) | +| Proxy PID / started | 542065, **03:54:53Z**, unrestarted since the incident | +| Effective `open files` | **65536 / 65536** — the morning's drop-in in force | +| `Recv-Q` / loopback | **0** / **`200` in 11 ms** | +| Datastore | 3.7 G of 98 G, 4% | +| felhom.eu `main` @ start | `435e044` — matches the runbook's stated baseline | -``` -ss -lnt '( sport = :8007 )' → LISTEN Recv-Q 1025 Send-Q 1024 -ls /proc//fd | wc -l → 1024 # == its soft RLIMIT_NOFILE -``` +## 2. The before-slope — and a correction I owe the morning's report -`Send-Q` on a listener is the accept backlog; `Recv-Q` is the queue depth. At 1025 against 1024 the -queue had overflowed, because `accept()` was returning `EMFILE` on every call. **1016 of the 1024 -descriptors were sockets and 547 connections sat in `CLOSE-WAIT`** — a connection leak, fed by -~85,000 requests/day, that reached the ceiling after 14 days of uptime. Last request served: -2026-08-17 18:15:32 UTC. Offsite DR was therefore down **≈ 9 h 37 m**. +Two independent windows on the same proxy generation: -**The observation that settled it:** the daemon was wedged from its own loopback too — -`curl https://127.0.0.1:8007/` on ep0 timed out. A listener that cannot serve `127.0.0.1` has no -network left to blame, and that check is cheap enough to make early. +| window | from → to | delta | rate | +|---|---|---|---| +| 31 min | 09:18:21Z fd=62 → 09:49:46Z fd=66 | +4 | **183/day** | +| 5.64 h | 04:11:36Z fd=19 → 09:49:46Z fd=66 | +47 | **200/day** | -Full record, including everything that was ruled out first (tunnel, nftables, disk, dead daemon): -**`documentation/audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md`**. +**This morning's incident note said ~85/day and "≈2 years of runway". Both were wrong.** They were +extrapolated from a single 17-minute window whose delta was **one descriptor** — a sample of one +cannot carry a daily rate, and the agreement with the historical ~73/day that made it feel solid was +coincidence. **The real rate is ~185–200/day, ~2.6× what I published, and the runway is ~357 days, +not two years.** Corrected in the incident document and in R-336 rather than left standing. -## 2. What was done +**The mechanism I named was also the minority one.** `CLOSE-WAIT` held flat at **1** across the +window while `ESTAB` grew **45 → 49** — *all* the growth was established connections. At the wedge +the split was **1011 ESTAB / 543 CLOSE-WAIT**, so ESTAB dominated there too. **R-336's fix must target +connections the proxy never reaps, not just `CLOSE-WAIT` sockets.** -1. `LimitNOFILE=65536` drop-ins for `proxmox-backup-proxy.service` and `proxmox-backup.service`, each - carrying its reason inline. **The API daemon was not implicated** (15 fds) and its drop-in says so - — a later reader must not mistake it for a second culprit. -2. Restarted both. After: soft limit 65536, fds back to 18, `Recv-Q 0`, loopback `200`. -3. **Verified from the customer side, not only from ep0** — both boxes got `200` in ~0.1 s and - `pvesm status` read `felhom-pbs pbs active`. -4. Re-drove the missed backups **through the product path** — `POST /backup?target=felhom-pbs` on - each agent's local API, issued from inside the guest's controller container with the controller's - own credentials, i.e. the same call the scheduler makes. Not a hand-run `vzdump`. +## 3. The changelog, verbatim — the run's primary deliverable -**Result:** `demo-felhom` → `ct/9201/2026-08-18T03:57:43Z` (4.10 GB, 36.4 s); -`demo-hp` → `ct/9201/2026-08-18T03:58:43Z` (4.29 GB, 41.5 s). Both directories carry a full manifest -on ep0, and both hosts log the re-run vzdump as `OK`. **And the hub agrees** — both boxes' next host -reports carry `felhom-pbs success=true` (04:00:33Z and 04:07:35Z), so the operator view went green on -the evidence rather than on the fact that a restart was performed. -**No data was lost and no backup was skipped** — the daily local -tier was never affected (it completed on both boxes at 05:00 and 05:02), and the PBS tier is weekly, -so the outage window cost exactly one attempt, which was re-driven the same morning. +All three entries between 4.2.2-1 and 4.2.5-1 read in full (128 lines), then swept for +`connection|file descriptor|fd|accept(|close_wait|keep-alive|socket|EMFILE|nofile|leak|proxy|listen| +backlog|hyper|tokio`. -Evidence copied off ep0 **before** the restart, per standing rule 5: -`documentation/audits/evidence-ep0-fd-2026-08-18/` — pre-restart state, post-fix state, access-log tail. +**Exactly one hit, and it is a false positive:** -## 3. What I got wrong, and corrected +> `* S3: config: allow editing the use-node-config flag that controls whether requests S3 endpoints` +> ` honor the node's proxy settings or not` -- **`proxmox-backup-manager version` prints *available* then *running*.** It read - `4.2.5-1 running version: 4.2.2`, which looks exactly like a daemon left behind by a package - upgrade. It is not: `dpkg -l` shows **4.2.2-1 installed**, 4.2.5-1 merely available in the repo, - and the on-disk binary is dated 2026-06-18. I restarted the API daemon on that mistaken reading; - harmless, and its drop-in is a genuine improvement, but it was not needed. -- **I addressed `demo-hp`'s agent at the island address** because `operations/nodes.md` says that box - is island-migrated. It is not (**R-338**), and the resulting timeout was briefly read as a fault. +HTTP-proxy configuration *for S3 requests* — not the `proxmox-backup-proxy` daemon. -## 4. Findings filed — R-336, R-337, R-338 +What the range does contain. **4.2.5-1** — a security release hardening client-supplied manifests: -All three are in `documentation/backlog/OPEN-ITEMS.md` with numbers, per the registers-first rule. +> `* backup: harden the handling of client supplied backup manifests:` +> ` - only accept archive names that are plain file names carrying a server side type extension. A` +> ` crafted name in a manifest could previously make a sync job read or write outside of the` +> ` snapshot directory, running as the unprivileged 'backup' user.` +> ` - keep an uploaded manifest in memory and only persist it on backup finish, checking that every` +> ` archive it lists was really uploaded during that session and that the checksums match` -- **R-336 — the poll rate is the real defect.** ~1 request/second against a DR endpoint written to - weekly. `LimitNOFILE` raises the ceiling; **it does not fix the leak**, it converts a fortnightly - outage into a multi-year one. -- **R-337 — a status endpoint that trailed its own artifact, then caught up. WATCHING, not a defect.** - `demo-hp`'s `/backup/status` was still serving the superseded 03:27:00Z failure at ~04:03Z while the - snapshot sat on ep0 and the host logged `OK`; `demo-felhom` updated within ~40 s. **I filed this as - a defect and that was premature** — the next host report (04:07:35Z) carried the success and the - skew cleared with no intervention. Rewritten as WATCHING, with the explicit instruction not to open - a fix until someone establishes whether this is just collection cadence. Recorded at all because - during the recovery it read as a second failure, and it was not one. -- **R-338 — `demo-hp` is not on the R-50 island and `nodes.md` says it is.** No `island_bridge` keys, - guest has no `eth1`, `vmbr9` has zero members, and the agent's local API is bound to the customer - LAN — the exposure R-50 existed to remove. +plus a sync/push chunk-reuse fix and a subscription-key architecture check. **4.2.4-1** — S3 rate +limits, a file-locking user-lookup cache, the new `proxmox-enterprise-support-keyring` dependency, +docs. **4.2.3-1** — UI/journal work, an LDAP search-filter escape, tape and timezone fixes. -**Not done, deliberately:** PBS 4.2.5-1 was not applied. Upgrading a production offsite endpoint was -outside what this run was authorised to do, and its changelog should be read for the connection- -handling leak first. +**Nothing addresses descriptor lifetime or connection reaping. My recommendation was: do not upgrade +for this reason.** -## 5. Gates — one pre-existing conviction, and a stated bypass +## 4. STOP 1 — the ruling -`python3 scripts/repo_gates.py --fast` → **rc=1, CONVICTED: golden-currency.** Eight of nine gates -pass. The conviction is **R-334, inherited and not caused here**: newest released controller -**0.216.0**, newest golden bake **0.214.0**, so a new install misses two releases. The gate reads -`felhom-controller/CHANGELOG.md` and `documentation/tests/golden-*` — **this session touched -neither**, and its whole diff is documentation. Baking is possible; **vouching is -operator-password-gated, and a baked-but-unvouched golden is worse than none**, so it is not a -one-sided job CC can finish. +**Operator (Viktor) ruled: upgrade anyway, for rehearsal value** — *"see how that works for us, we +need practice with that too"*. Legitimate and recorded as such: this was **a practice run of the +upgrade procedure on a Tier-2 protected machine, not a fix for the leak**. -**This push therefore used `git push --no-verify`, stated here per `.claude/rules/gates.md`.** -R-334 is updated in the register with the new numbers rather than left reading 0.215.0. +**The interpretation was fixed in writing before any numbers existed** (`stop1-ruling.txt`): +unchanged slope = **expected**, not a failed upgrade; changed slope = a **surprise** needing +explanation, not a confirmation. That file was written at the ruling, not afterwards, so neither +outcome could be rationalised into a success. -**CI checked by run ID, as the checklist requires: run `348`, `head_sha ebfd0967c`, conclusion -`failure`, elapsed 13 s** (04:09:45→04:09:58Z). Expected and inherited — CI's only step is -`python3 scripts/repo_gates.py --fast`, the same entry point that convicts golden-currency locally, -with the sibling repos fetched. The 13 s runtime places it in the workflow's own "honest gate -failure" band rather than the R-265 reap band, so the result is the gate speaking, not the runner. -**I could not read the run log to name the gate from CI's own mouth** — `actions/runs/348/logs` and -`actions/tasks/348/logs` both 404, `actions/runs/348/jobs` returns an empty list, authenticated as -`admin`, and the web log endpoint 302s. So this is an inference from the local run plus the workflow -definition, not a direct reading, and it is stated as such. +## 5. STOP 2 — the snapshot, and what it does not cover -**Expect one `[felhom CI] gates FAILED in admin/felhom.eu` mail for run 348** — the workflow alarms -on failure by design. It is this push, and it is the golden-currency row, not a new fault; the same -mails on 12 and 14 August have the same cause. +Snapshot **421440873** `felhom-hetzner-20260818`, 15.06 GB, status **Available** (complete, not +merely started), server #147604682, project 15217960. -*(Noted for accuracy: the first gate run was piped to `tail`, which returned `rc=0` — `tail`'s exit -code, not the gate's. It was re-run unpiped to read the real `rc=1`. That is standing rule 1's trap -in its smaller form, and the number reported above is the unpiped one.)* +**It covers `/dev/sda` only.** `/mnt/pbs-datastore` is `/dev/sdb`, a separate 100 GB **Volume**, and +Hetzner server snapshots exclude attached volumes — so this is a rollback for the *software* state +(packages, unit files, the `LimitNOFILE` drop-ins, nftables, wg) and **not a backup of the backup +data**. Fine for a package install that writes no datastore content; **it must not be remembered as +datastore protection.** Taken on a running server, deliberately: powering off ep0 to guard a userspace +package install would take the only off-premises copy offline. -## 6. What to watch +## 6. The upgrade and its verification -The positive observable is the descriptor count, not the absence of an alert — an empty alert queue -is equally consistent with "healthy" and "wedged again": +Simulated first (`-s`): **0 to remove**, so the abort condition never triggered. Then +`apt-get install --only-upgrade -y proxmox-backup-server`, **09:51:00→09:51:06Z, exit 0**. Upgraded +server/client/docs to 4.2.5-1 plus one new dependency, `proxmox-enterprise-support-keyring 1.1` — +**which the 4.2.4-1 changelog had declared**, a small real consistency check between what I read and +what apt did. -```bash -ssh root@ 'PID=$(systemctl show proxmox-backup-proxy -p MainPID --value); \ - ls /proc/$PID/fd | wc -l; ss -lnt "( sport = :8007 )"' -``` +| check | result | +|---|---| +| installed | server / client / docs **4.2.5-1** | +| daemons | `proxmox-backup-proxy` **active running**, `proxmox-backup` **active running** | +| proxy restarted | 542065 → **551655** @ 09:51:04 | +| **effective `open files`** | **65536 / 65536** — survived the new package | +| drop-ins on disk | both present, unmodified | +| `Recv-Q` | **0** | +| loopback | **`200` in 12 ms** | +| `felhom-pve` over tunnel | **`200` in 0.103 s**, `felhom-pbs active` | +| `demo-hp` over tunnel | **`200` in 0.096 s**, `felhom-pbs active` | +| hub gauge, post-upgrade | `11:59:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)` | -Healthy is ~20 fds and `Recv-Q 0`. **A count climbing between restarts means R-336's leak is still -live.** +`proxmox-backup-manager version` now reads `4.2.5-1 running version: 4.2.5`. **This independently +settles the morning's confusion**: that string was never reporting a stale daemon, and now that +installed and running genuinely match, both halves agree. + +**One false alarm, mine.** I queried `systemctl is-active proxmox-backup-api` and got `inactive`. +**That unit does not exist** — `systemctl cat` returns *"No files found for +proxmox-backup-api.service"*. The real pair is `proxmox-backup-proxy.service` ("API Proxy Server") +and `proxmox-backup.service` ("API Server"), both active. A bad query, not a fault — recorded because +for as long as it took to check, it looked exactly like one. + +**No backup, restore or verify was triggered** to "prove" the endpoint, per the runbook: the reads +above answer it without mutating a protected datastore. + +## 7. The after-slope — unchanged, as predicted + +| | window | delta | rate | +|---|---|---|---| +| before (PID 542065) | 09:18:21Z fd=62 → 09:49:46Z fd=66 | +4 / 1885 s | **183/day** | +| after (PID 551655) | 09:51:22Z fd=17 → 10:23:21Z fd=22 | +5 / 1919 s | **225/day** | + +**These are not distinguishable.** The windows differ by **one descriptor**; Poisson uncertainty on +n=4 is ±2 and on n=5 is ±2.2, so both are consistent with a single unchanged rate. **The after-figure +being numerically higher is noise, not a regression — and certainly not an improvement.** Composition +repeats the pattern: **ESTAB 0 → 5, CLOSE-WAIT 0 → 1.** + +**Thirty minutes cannot settle this in either direction, and nothing here claims it does.** + +## 8. Register + +- **R-341 filed, WATCHING** — the two dated checks: **+24 h (2026-08-19 ~10:00Z)** and + **+7 d (2026-08-25 ~10:00Z)**, against new `t0` **fd=17 @ 09:51:22Z, PID 551655**, with the exact + command and the instruction to **record the ESTAB/CLOSE-WAIT split, not just the total** — the + split is what identifies which leak it is. If the PID has changed, the window is void. +- **R-336 updated and STILL OPEN.** Its wrong 85/day baseline is corrected to 183–200/day, the + mechanism is re-pointed at ESTAB, and the upgrade's null result is recorded. **It stays open on its + own merits:** an upgrade that *had* fixed the leak still would not make ~85,000 requests/day to a + weekly-write DR endpoint correct. +- `documentation/runbooks/offsite-endpoint.md` — **no change made, correctly**: it states no PBS + version literal, only package names, so there was nothing to update (and per `docs.md`, version + literals do not belong in a current-state doc anyway). + +## 9. Teardown + +**This run provisioned nothing.** No `.deb` was downloaded — `apt-get changelog` served the text +directly, so the runbook's `/tmp` cleanup was never needed. The Hetzner snapshot **is retained** as +the rollback; deleting it is an operator decision, and it costs €0.018161/GB/month. + +## 10. Observations, not acted on + +- **`pvesm status` reports `felhom-pbs` with Total/Used/Available all `0`** on both boxes while + status reads `active`. Consistent before and after the upgrade, and the hub's own gauge reads the + real 3.7 GB / 97.9 GB, so nothing is broken — but the PVE-side numbers are not usable as a capacity + signal. Likely the namespace-scoped token lacking datastore-level audit. Not filed; noted here. +- **8 other packages are held back** (`8 not upgraded`), untouched deliberately — `full-upgrade` on + this machine is forbidden by the runbook and would change the kernel and WireGuard alongside the + thing under test. +- **The morning's `--no-verify` situation is unchanged**: `golden-currency` still convicts on the + inherited R-334 (controller 0.216.0 vs golden 0.214.0), untouched by this run. diff --git a/STATUS.md b/STATUS.md index 97a42e4f..d746c432 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,6 +1,6 @@ # STATUS — what works, what's broken, what's next -**Updated 2026-08-18 (early — a listening socket that served nobody; off-site backups were down for +**Updated 2026-08-18 (midday — a listening socket that served nobody, then a rehearsal upgrade that changed nothing; off-site backups were down for 9½ hours overnight and are back).** > **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates @@ -50,8 +50,8 @@ record with no machine** — created 13 August, no host, no backups, nothing to were re-run the same morning and are on the off-site box**. **Nothing was lost and nothing was skipped:** the daily copies on the machines themselves were never affected, and the off-site copy is weekly, so exactly one attempt fell in the window. The underlying cause — we ask that box a question - about once a second, all day — is filed and **not yet fixed**; the raised ceiling buys years, not a - cure. + about once a second, all day — is filed and **not yet fixed**; the raised ceiling buys **under a + year** on the corrected measurement, not a cure. - **The removal now genuinely reverses the installation** (R-316, `installer-v1.28.0` published). The second reinstall used to hit our own leftover; it was watched failing on the cycle that actually @@ -86,7 +86,14 @@ record with no machine** — created 13 August, no host, no backups, nothing to for a box we actually write to once a week. That volume is what turned a slow internal leak into last night's outage in a fortnight. The higher ceiling makes it rare, not impossible — **the leak itself is untouched.** The honest health check is the resource count climbing, not the absence of an - alarm. + alarm. **Corrected later the same morning: my first estimate of how fast it leaks was too + optimistic by about 2.5×** — measured properly it is under a year to the new ceiling, not two years. + A deadline, not a comfort. +- **The off-site box was updated, and it did not help — as expected** (R-341). On your ruling we + installed the newer backup software for the practice, having first read its release notes and found + **nothing** about the fault we have. The update went cleanly and everything works, but the leak + behaves exactly as before, which is the result the release notes predicted. **Two dated checks are + booked — 19 August and 25 August** — because half an hour of watching cannot honestly settle it. - **One machine's status took several minutes to admit a backup had worked** (R-337). `demo-hp` was still showing this morning's failure for four minutes after the copy was safely on the off-site box; `demo-felhom` updated in under a minute. **It corrected itself** and both machines now read diff --git a/documentation/audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md b/documentation/audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md index a8bfec52..bcd2afad 100644 --- a/documentation/audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md +++ b/documentation/audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md @@ -168,8 +168,133 @@ A healthy proxy sits near 20 fds with `Recv-Q 0`. **A rising fd count between re leak is still live** — an unchanging one after a poll-rate reduction would confirm the fix. **It is still live, and it was measured rather than assumed.** At **16 m 43 s** after the restart the -proxy held **19 fds** (from 18) with **1** connection in `CLOSE-WAIT`. One descriptor per ~17 minutes -is **≈ 85/day** — which lands on the historical rate implied by the failure itself, 1016 sockets over -14 days ≈ **73/day**. Two independent estimates of the same slope agreeing is what makes this a -measurement instead of a story, and it puts the next ceiling at roughly **2 years** rather than the -fortnight the old limit gave. **The leak is unfixed; only its period changed.** +proxy held **19 fds** (from 18) with **1** connection in `CLOSE-WAIT`. + +> ### ⚠ CORRECTED 2026-08-18 09:50 UTC — the two numbers first published here were wrong +> +> This section originally read *"one descriptor per ~17 minutes is ≈ 85/day … it puts the next +> ceiling at roughly 2 years"*. **Both figures are wrong, and the reason is worth keeping:** they were +> extrapolated from a single 17-minute window whose delta was **one descriptor**. A sample of one +> cannot carry a daily rate, and the agreement with the historical ~73/day that made it feel solid was +> coincidence. +> +> Re-measured on the same proxy generation (PID 542065, unrestarted) during +> `RUNBOOK ep0 PBS upgrade`, two independent windows: +> +> | window | from | to | rate | +> |---|---|---|---| +> | 31 min | 09:18:21Z fd=62 | 09:49:46Z fd=66 | **183/day** | +> | 5.64 h | 04:11:36Z fd=19 | 09:49:46Z fd=66 | **200/day** | +> +> **The real rate is ~185–200/day — about 2.6× what was published — and the runway to the 65536 +> ceiling is ~357 days, not two years.** Still an enormous improvement on the fortnight the old limit +> gave, but under a year, so it is a deadline rather than a comfort. +> +> **And the mechanism named above is the minority one.** Across that window `CLOSE-WAIT` held flat at +> **1** while `ESTAB` grew **45 → 49**: *all* the growth was established connections. At the wedge the +> split was **1011 ESTAB / 543 CLOSE-WAIT**, so ESTAB was the larger half there too and this document +> put its emphasis on the wrong one. **R-336's fix must target connections the proxy never reaps, not +> only sockets left in `CLOSE-WAIT`.** + +**The leak is unfixed; only its period changed.** + +--- + +# Follow-up — 2026-08-18, later the same morning: the PBS upgrade + +Run under `RUNBOOK — ep0: read the PBS changelog, then decide whether to upgrade`. Two supervised +STOPs, both cleared by the operator. Evidence: `evidence-ep0-pbs-upgrade-2026-08-18/`. + +## The changelog said nothing relevant — and that was the finding + +The runbook's first job was a pure read: does anything between the installed **4.2.2-1** and the +candidate **4.2.5-1** fix connection handling? All three intervening entries were read in full (128 +lines) and swept for `connection|file descriptor|fd|accept(|close_wait|keep-alive|socket|EMFILE| +nofile|leak|proxy|listen|backlog|hyper|tokio`. + +**Exactly one keyword hit, and it is a false positive:** + +> `* S3: config: allow editing the use-node-config flag that controls whether requests S3 endpoints +> honor the node's proxy settings or not` + +That is HTTP-proxy configuration *for S3 requests*, not the `proxmox-backup-proxy` daemon. What the +range actually contains: **4.2.5-1** is a security release hardening client-supplied backup manifests +(an archive name in a crafted manifest could make a sync job read or write outside the snapshot +directory as the `backup` user; manifests are now held in memory and only persisted at backup finish +with per-archive checksum verification), plus a sync/push chunk-reuse fix. **4.2.4-1** is S3 rate +limits, a file-locking user-lookup cache, docs. **4.2.3-1** is UI/journal work, an LDAP search-filter +escape, tape and timezone fixes. + +**Nothing addresses descriptor lifetime or connection reaping.** The recommendation was therefore +**do not upgrade for this reason**. + +## What was done anyway, and why that is fine + +**Operator ruling at STOP 1: upgrade regardless, for rehearsal value** — *"see how that works for us, +we need practice with that too"*. So this was executed as **a practice run of the upgrade procedure on +a Tier-2 protected machine, not as a fix for the leak**, and the distinction was written into +`stop1-ruling.txt` *before* any numbers existed, precisely so the outcome could not be rationalised +afterwards in either direction. + +STOP 2 cleared with Hetzner snapshot **421440873** `felhom-hetzner-20260818`, 15.06 GB, status +**Available**. **That snapshot covers `/dev/sda` only.** `/mnt/pbs-datastore` is `/dev/sdb`, a separate +100 GB Volume, and Hetzner server snapshots exclude attached volumes — so it is a rollback for the +software state and **not** a backup of the backup data. Acceptable here because a package install +writes no datastore content; it must not be remembered as datastore protection. + +## Result + +`apt-get install --only-upgrade proxmox-backup-server`, 09:51:00→09:51:06Z, exit 0. A `-s` simulation +was run first and reported **0 to remove**, so the runbook's abort condition never triggered. Upgraded +server/client/docs to **4.2.5-1**, plus one genuinely new dependency, `proxmox-enterprise-support- +keyring 1.1` — which the 4.2.4-1 changelog had declared, a small but real consistency check between +what was read and what apt did. + +| check | result | +|---|---| +| installed (`dpkg -l`) | server / client / docs **4.2.5-1** | +| daemons | `proxmox-backup-proxy` **active running**, `proxmox-backup` **active running** | +| proxy restarted | PID 542065 → **551655** @ 09:51:04 | +| **effective `open files`** | **65536 / 65536** — the drop-in survived the new package | +| `Recv-Q` | 0 | +| loopback | `200` in 12 ms | +| from `felhom-pve` | `200` in 0.103 s, `felhom-pbs active` | +| from `demo-hp` | `200` in 0.096 s, `felhom-pbs active` | +| hub gauge, post-upgrade | `11:59:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)` | + +`proxmox-backup-manager version` now reads `4.2.5-1 running version: 4.2.5`. **This independently +settles the confusion recorded in the original incident**: the string was never reporting a stale +daemon, and now that installed and running genuinely match, both halves agree. + +**One false alarm, mine:** `systemctl is-active proxmox-backup-api` returned `inactive`. That unit +does not exist — `systemctl cat` says *"No files found for proxmox-backup-api.service"*. The real pair +is `proxmox-backup-proxy.service` ("API Proxy Server") and `proxmox-backup.service` ("API Server"), +both active. A bad query, not a fault, and it is written down because it looked exactly like a fault +for as long as it took to check. + +## The slope: unchanged, as predicted + +| | window | delta | rate | +|---|---|---|---| +| **before** (PID 542065) | 09:18:21Z fd=62 → 09:49:46Z fd=66 | +4 / 1885 s | **183/day** | +| **after** (PID 551655) | 09:51:22Z fd=17 → 10:23:21Z fd=22 | +5 / 1919 s | **225/day** | + +**These are not distinguishable.** The two windows differ by a single descriptor; Poisson uncertainty +on n=4 is ±2 and on n=5 is ±2.2, so both are consistent with one unchanged underlying rate. **The +after-figure being numerically higher is noise, not a regression — and emphatically not an +improvement.** This is the expected outcome and it matches the changelog: no mechanism, no change. + +Composition after the upgrade repeats the pattern that matters: **ESTAB 0 → 5, CLOSE-WAIT 0 → 1.** + +**Thirty minutes cannot settle this**, in either direction, and this document does not claim it does. +The honest checks are **+24 h (2026-08-19 ~10:00Z)** and **+7 d (2026-08-25 ~10:00Z)** against the +new `t0` of **fd=17 at 09:51:22Z, PID 551655** — filed as **R-341**. + +## What this run did not do + +Did not reduce the poll rate (**R-336 stays open** — an upgrade that fixed the leak still would not +make ~85,000 requests/day to a weekly-write DR endpoint correct), did not touch the `LimitNOFILE` +drop-ins, did not run `full-upgrade` or touch the kernel, did not trigger a backup/restore/verify to +"prove" the endpoint, and did not change anything on either customer box. Nothing was provisioned, so +there is nothing to tear down; no `.deb` was downloaded, as `apt-get changelog` served the text +directly. diff --git a/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step1-state.txt b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step1-state.txt new file mode 100644 index 00000000..c15bc001 --- /dev/null +++ b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step1-state.txt @@ -0,0 +1,105 @@ +### capture-time +2026-08-18T09:18:21+00:00 +### dpkg -l (INSTALLED — the truth) +ii proxmox-backup-client 4.2.2-1 amd64 Proxmox Backup Client tools +ii proxmox-backup-server 4.2.2-1 amd64 Proxmox Backup Server daemon with tools and GUI +### apt-cache policy proxmox-backup-server +proxmox-backup-server: + Installed: 4.2.2-1 + Candidate: 4.2.5-1 + Version table: + 4.2.5-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.2.4-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.2.3-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + *** 4.2.2-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 100 /var/lib/dpkg/status + 4.2.1-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.2.0-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.1.13-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.1.12-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.1.11-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.1.10-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.1.9-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.1.8-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.1.7-2 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.1.6-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.1.5-2 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.1.5-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.1.4-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.1.2-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.1.1-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.1.0-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.0.22-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.0.21-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.0.20-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.0.19-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.0.18-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.0.17-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.0.16-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.0.15-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.0.14-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.0.13-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.0.12-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.0.11-4 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.0.11-2 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.0.10-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.0.9-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.0.8-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.0.7-1 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages + 4.0.6-2 500 + 500 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 Packages +### proxmox-backup-manager version (available THEN running — misleading, see incident) +proxmox-backup-server 4.2.5-1 running version: 4.2.2 +### proxy MainPID +542065 +### fd count +62 +### limits +Max open files 65536 65536 files +### listen queue +State Recv-Q Send-Q Local Address:Port Peer Address:Port +LISTEN 0 1024 *:8007 *:* +### close-wait count +1 +### proxy start time +Tue Aug 18 03:54:53 2026 +### datastore +Filesystem Size Used Avail Use% Mounted on +/dev/sdb 98G 3.7G 95G 4% /mnt/pbs-datastore diff --git a/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step1b-loopback.txt b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step1b-loopback.txt new file mode 100644 index 00000000..d0702366 --- /dev/null +++ b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step1b-loopback.txt @@ -0,0 +1,15 @@ +### loopback probe +2026-08-18T09:18:40+00:00 +http=200 time_total=0.011537 +### fd types + 54 socket:anon + 3 anon_inode:anon + 1 0 + 1 /var/log/proxmox-backup/api/auth.log + 1 /var/log/proxmox-backup/api/access.log + 1 /var/lib/proxmox-backup/rrdb/rrd.journal + 1 /mnt/pbs-datastore/.lock + 1 /dev/null +### socket states on 8007 + 45 ESTAB + 1 LISTEN diff --git a/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step2-second-reading.txt b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step2-second-reading.txt new file mode 100644 index 00000000..85a97ff6 --- /dev/null +++ b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step2-second-reading.txt @@ -0,0 +1,16 @@ +### step2-second-reading +2026-08-18T09:49:46+00:00 +### MainPID +542065 +### fd count +66 +### close-wait +1 +### socket states + 49 ESTAB + 1 LISTEN +### listen queue +State Recv-Q Send-Q Local Address:Port Peer Address:Port +LISTEN 0 1024 *:8007 *:* +### proxy start +Tue Aug 18 03:54:53 2026 diff --git a/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step2-slope-computation.txt b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step2-slope-computation.txt new file mode 100644 index 00000000..4faddeaa --- /dev/null +++ b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step2-slope-computation.txt @@ -0,0 +1,18 @@ +BEFORE-SLOPE — proxy PID 542065, started 2026-08-18 03:54:53 UTC, never restarted since + +Step-2 window (31 min, this run) + 09:18:21Z fd=62 -> 09:49:46Z fd=66 + interval 0.5236 h (1885 s), delta 4 fd + = 7.64 fd/hour = 183.3 fd/DAY + +Long window (incident t0+17m -> now) + 04:11:36Z fd=19 -> 09:49:46Z fd=66 + interval 5.6361 h (20290 s), delta 47 fd + = 8.34 fd/hour = 200.1 fd/DAY + +Runway from fd=66 at 183/day to the 65536 soft limit: + 357 days = 0.98 years + +Historical rate implied by the outage: 1016 sockets over 14.29 d = 71.1/day +Socket composition: ESTAB 45 -> 49 (+4); CLOSE-WAIT 1 -> 1 (+0). +=> ALL of the growth in this window is ESTABLISHED connections, not CLOSE-WAIT. diff --git a/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step3-changelog-attempt1.txt b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step3-changelog-attempt1.txt new file mode 100644 index 00000000..1dab59a9 --- /dev/null +++ b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step3-changelog-attempt1.txt @@ -0,0 +1,200 @@ +Get:1 https://metadata.cdn.proxmox.com rust-proxmox-backup 4.2.5-1 Changelog [178 kB] +rust-proxmox-backup (4.2.5-1) trixie; urgency=medium + + * backup: harden the handling of client supplied backup manifests: + - only accept archive names that are plain file names carrying a server + side type extension. A crafted name in a manifest could previously make + a sync job read or write outside of the snapshot directory, running as + the unprivileged 'backup' user. Reaching this needed a manifest from a + configured sync remote or from a client that already had backup access + to the datastore. + - keep an uploaded manifest in memory and only persist it on backup + finish, checking that every archive it lists was really uploaded during + that session and that the checksums match the ones computed server + side. A client uploading a manifest that references archives it did not + upload now gets an error on finish instead of such a snapshot being + created. + + * fix #7878: sync: push: reuse the manifest of a previous snapshot on a + non-encrypting push if the source snapshot was encrypted with a matching + key, restoring chunk reuse and thus avoiding needlessly long sync runs. + + * sync: push: keep the sign-only crypt mode of a source archive instead of + reducing it to unencrypted when pushing without server side encryption. + + * subscription: reject a subscription key issued for a different + architecture than the host, as arm64 keys carry an explicit marker, so a + wrong key fails fast instead of only erroring during the online check. + + * update to proxmox-upgrade-checks 1.1, which accepts the 7.0 kernel, tells + a bookworm backport apart from a trixie build and fixes the dkms check. + + * docs: clarify in the backup protocol description that the manifest is + uploaded by the client and only persisted on backup finish. + + -- Proxmox Support Team Wed, 05 Aug 2026 18:25:37 +0200 + +rust-proxmox-backup (4.2.4-1) trixie; urgency=medium + + * docs: document the debug symbol repository + + * datastore: fix wrong local path used for S3 bad chunk handling during + garbage collection + + * refactor file creation/mode/ownership helpers to proxmox-product-config + crate + + * fix #7642: avoid expensive user lookups on file locking by caching the + backup user/group ID + + * depend on proxmox-enterprise-support-keyring, and track its version in the + package version API endpoint + + * fix #5748: docs: add `catalog.pcat1` format specification + + * docs: system requirements: document we recommend local storage + + * S3: fix #6841: allow configuring request rate limits by updating to + proxmox-s3-client 1.4.1. these rate limits are split into active and + passive methods, allowing separate handling of POST/PUT/DELETE and GET/HEAD + request limits. + + * S3: config: allow editing the use-node-config flag that controls whether + requests S3 endpoints honor the node's proxy settings or not + + * sync: push: gracefully handle previous manifest signature mismatches, which + can happen when enabling or disabling push-encryption on an already synced + backup group + + -- Proxmox Support Team Wed, 29 Jul 2026 15:08:15 +0200 + +rust-proxmox-backup (4.2.3-1) trixie; urgency=medium + + * css: remove x-grid-row-loading class, replace it with non-blurry SVG + variant from proxmox-widget-toolkit + + * pbs-client: add backoff log throttle, to ensure progress and similar output + appears quickly initially, but does not create overly long logs + + * client: report progress during restore + + * api: journal: adopt proxmox-syslog-api and stream the output, making the + implementation consistent with the one from Proxmox Datacenter Manager + + * ui: enable the structured journal view and per-service logs, including + colored output and filtering capabilities + + * ui: always use arrays for 'delete' property, instead of manually converting + + * fix #5971: tape: don't warn on custom MAM attribute write failures + + * fix #7175: api: time: use timedatectl instead of /etc/timezone + + * fix #7187: report: add ethtool output for physical interfaces + + * prune jobs: schedule jobs that do not prune anything, but warn during their + execution. such jobs allow testing scheduling options, but make no sense + for production use. + + * fix #6691: allow search by comment in datastore content, make search + case-insensitive and correctly reset content view after empty searches + + * ui: datastore: disable various action tooltips for actions which cannot be + triggered + + * ldap: escape the user-provided user name when using it in the LDAP search + filter that looks up the user DN. + + * ldap sync: log which user properties change when synchronizing an existing + user, instead of only reporting that the user was updated. + + * api schema/section config: add support for declaring deprecated property + aliases, to allow renaming properties without showing the old name in the + documentation + + * rest server: accept deprecated property aliases in JSON request bodies by + rewriting them to the canonical name before verification and dispatch, like + the CLI and query-string handling already do. + + * fix #7690: fs: replace_file: close the temporary file before renaming or + unlinking it, fixing the replacement on WORM file systems and avoiding + leftover .fuse_hidden files on FUSE mounts. + + * fs: make_tmp_file: append the temporary suffix instead of replacing the + file extension, keeping the original file name intact for easier + debugging of leftover temporary files. + + -- Proxmox Support Team Tue, 14 Jul 2026 12:54:15 +0200 + +rust-proxmox-backup (4.2.2-1) trixie; urgency=medium + + * api: backup: run synchronous chunk-insert operations off the asynchronous + runtime's worker threads. Blocking those threads, most notably during S3 + uploads that can wait up to three hours for a chunk lock, could stall the + I/O and timer drivers of the entire runtime and starve other backup + workers. + + * client: backup: make the file-based backup more robust against files that + cannot be accessed or that vanish while the backup is running: + - fix #7658: skip a file and log a warning instead of aborting the whole + backup when querying its metadata fails with a permission error, + matching the existing handling of permission errors when opening files. + Such files can also still be excluded explicitly. + - consistently ignore files that disappear during the backup and warn + about them, instead of treating this as a fatal error in some cases. + + * tape: backup: fix the command-line group filter, which was passed to the + API under the wrong parameter name and thus had no effect. + + * api: do not log a spurious error when listing the files of a snapshot + whose manifest does not exist yet, which is expected while a backup to + that snapshot is still running. + + * ui: datastore summary: fix the per-datastore sync and prune job counts, + which were derived from an incorrectly parsed datastore ID; the prune + count in particular was always shown as zero. + + * tape: fix two typos in log and informational messages. + + -- Proxmox Support Team Thu, 18 Jun 2026 11:28:25 +0200 + +rust-proxmox-backup (4.2.1-1) trixie; urgency=medium + + * fix #5076: api: support an 'audiences' property on OpenID realms, listing + additional trusted audience values besides the configured client-id. + Improves compatibility with providers that issue tokens with multiple + audiences. + + * fix #7562: api/ui: tape: separate the format-media 'load-barcode' + parameter from the existing 'label-text' verification, so the web + interface can load and format empty or previously-unrelated tapes from a + changer slot in one step. The old shared parameter aborted formatting + after a successful load whenever the on-tape label did not match. + + * sync: pull: refuse to overwrite a locally encrypted snapshot from an + unencrypted source or one using a different key, and detect content + differences between two unencrypted snapshots that share a backup time. + Previously such mismatches silently triggered a resync that overwrote the + local snapshot. + + * datastore: fix tuning option changes not propagating the updated sync + level to the chunk store until the service was restarted. + + * datastore: improve the error message when prune cannot acquire a snapshot + lock for deletion, by showing the snapshot directory and lock file paths + instead of an internal debug dump. + + * api: backup: fix benchmark, finish-failed, and backup-failed cleanups + leaving an orphaned empty backup group behind on the datastore. + + * api: node: tasks status: return the task end time as an optional field + once the task is finished, so the task viewer can render the correct + duration without an extra API call. + + * api/ui: node: add a 'location' property to the node config, exposed + through the node options panel. + + * subscription: reuse the server ID from an existing subscription info when + multiple candidates are detected, falling back to the first candidate only + when no prior info exists. + diff --git a/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step3-range-4.2.2-to-4.2.5.txt b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step3-range-4.2.2-to-4.2.5.txt new file mode 100644 index 00000000..bc3a0f19 --- /dev/null +++ b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step3-range-4.2.2-to-4.2.5.txt @@ -0,0 +1,128 @@ +rust-proxmox-backup (4.2.5-1) trixie; urgency=medium + + * backup: harden the handling of client supplied backup manifests: + - only accept archive names that are plain file names carrying a server + side type extension. A crafted name in a manifest could previously make + a sync job read or write outside of the snapshot directory, running as + the unprivileged 'backup' user. Reaching this needed a manifest from a + configured sync remote or from a client that already had backup access + to the datastore. + - keep an uploaded manifest in memory and only persist it on backup + finish, checking that every archive it lists was really uploaded during + that session and that the checksums match the ones computed server + side. A client uploading a manifest that references archives it did not + upload now gets an error on finish instead of such a snapshot being + created. + + * fix #7878: sync: push: reuse the manifest of a previous snapshot on a + non-encrypting push if the source snapshot was encrypted with a matching + key, restoring chunk reuse and thus avoiding needlessly long sync runs. + + * sync: push: keep the sign-only crypt mode of a source archive instead of + reducing it to unencrypted when pushing without server side encryption. + + * subscription: reject a subscription key issued for a different + architecture than the host, as arm64 keys carry an explicit marker, so a + wrong key fails fast instead of only erroring during the online check. + + * update to proxmox-upgrade-checks 1.1, which accepts the 7.0 kernel, tells + a bookworm backport apart from a trixie build and fixes the dkms check. + + * docs: clarify in the backup protocol description that the manifest is + uploaded by the client and only persisted on backup finish. + + -- Proxmox Support Team Wed, 05 Aug 2026 18:25:37 +0200 + +rust-proxmox-backup (4.2.4-1) trixie; urgency=medium + + * docs: document the debug symbol repository + + * datastore: fix wrong local path used for S3 bad chunk handling during + garbage collection + + * refactor file creation/mode/ownership helpers to proxmox-product-config + crate + + * fix #7642: avoid expensive user lookups on file locking by caching the + backup user/group ID + + * depend on proxmox-enterprise-support-keyring, and track its version in the + package version API endpoint + + * fix #5748: docs: add `catalog.pcat1` format specification + + * docs: system requirements: document we recommend local storage + + * S3: fix #6841: allow configuring request rate limits by updating to + proxmox-s3-client 1.4.1. these rate limits are split into active and + passive methods, allowing separate handling of POST/PUT/DELETE and GET/HEAD + request limits. + + * S3: config: allow editing the use-node-config flag that controls whether + requests S3 endpoints honor the node's proxy settings or not + + * sync: push: gracefully handle previous manifest signature mismatches, which + can happen when enabling or disabling push-encryption on an already synced + backup group + + -- Proxmox Support Team Wed, 29 Jul 2026 15:08:15 +0200 + +rust-proxmox-backup (4.2.3-1) trixie; urgency=medium + + * css: remove x-grid-row-loading class, replace it with non-blurry SVG + variant from proxmox-widget-toolkit + + * pbs-client: add backoff log throttle, to ensure progress and similar output + appears quickly initially, but does not create overly long logs + + * client: report progress during restore + + * api: journal: adopt proxmox-syslog-api and stream the output, making the + implementation consistent with the one from Proxmox Datacenter Manager + + * ui: enable the structured journal view and per-service logs, including + colored output and filtering capabilities + + * ui: always use arrays for 'delete' property, instead of manually converting + + * fix #5971: tape: don't warn on custom MAM attribute write failures + + * fix #7175: api: time: use timedatectl instead of /etc/timezone + + * fix #7187: report: add ethtool output for physical interfaces + + * prune jobs: schedule jobs that do not prune anything, but warn during their + execution. such jobs allow testing scheduling options, but make no sense + for production use. + + * fix #6691: allow search by comment in datastore content, make search + case-insensitive and correctly reset content view after empty searches + + * ui: datastore: disable various action tooltips for actions which cannot be + triggered + + * ldap: escape the user-provided user name when using it in the LDAP search + filter that looks up the user DN. + + * ldap sync: log which user properties change when synchronizing an existing + user, instead of only reporting that the user was updated. + + * api schema/section config: add support for declaring deprecated property + aliases, to allow renaming properties without showing the old name in the + documentation + + * rest server: accept deprecated property aliases in JSON request bodies by + rewriting them to the canonical name before verification and dispatch, like + the CLI and query-string handling already do. + + * fix #7690: fs: replace_file: close the temporary file before renaming or + unlinking it, fixing the replacement on WORM file systems and avoiding + leftover .fuse_hidden files on FUSE mounts. + + * fs: make_tmp_file: append the temporary suffix instead of replacing the + file extension, keeping the original file name intact for easier + debugging of leftover temporary files. + + -- Proxmox Support Team Tue, 14 Jul 2026 12:54:15 +0200 + +rust-proxmox-backup (4.2.2-1) trixie; urgency=medium diff --git a/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step4a-simulate.txt b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step4a-simulate.txt new file mode 100644 index 00000000..77ea151a --- /dev/null +++ b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step4a-simulate.txt @@ -0,0 +1,30 @@ +### apt-get update start +2026-08-18T09:50:43+00:00 +Hit:6 http://mirror.hetzner.com/debian/security trixie-security InRelease +Hit:7 http://download.proxmox.com/debian/pbs trixie InRelease +Get:8 http://deb.debian.org/debian trixie-backports InRelease [54.0 kB] +Get:9 http://deb.debian.org/debian-security trixie-security InRelease [43.4 kB] +Get:10 http://mirror.hetzner.com/debian/packages trixie-backports/main amd64 Packages [314 kB] +Get:11 http://deb.debian.org/debian trixie-backports/main Sources [297 kB] +Fetched 858 kB in 0s (6119 kB/s) +Reading package lists... +### SIMULATION (-s): nothing is installed by this +2026-08-18T09:50:44+00:00 +Reading package lists... +Building dependency tree... +Reading state information... +The following additional packages will be installed: + proxmox-backup-client proxmox-backup-docs proxmox-enterprise-support-keyring +The following NEW packages will be installed: + proxmox-enterprise-support-keyring +The following packages will be upgraded: + proxmox-backup-client proxmox-backup-docs proxmox-backup-server +3 upgraded, 1 newly installed, 0 to remove and 8 not upgraded. +Inst proxmox-backup-client [4.2.2-1] (4.2.5-1 Proxmox Backup System Debian Repository:stable [amd64]) +Inst proxmox-backup-docs [4.2.2-1] (4.2.5-1 Proxmox Backup System Debian Repository:stable [all]) +Inst proxmox-enterprise-support-keyring (1.1 Proxmox Backup System Debian Repository:stable [all]) +Inst proxmox-backup-server [4.2.2-1] (4.2.5-1 Proxmox Backup System Debian Repository:stable [amd64]) +Conf proxmox-backup-client (4.2.5-1 Proxmox Backup System Debian Repository:stable [amd64]) +Conf proxmox-backup-docs (4.2.5-1 Proxmox Backup System Debian Repository:stable [all]) +Conf proxmox-enterprise-support-keyring (1.1 Proxmox Backup System Debian Repository:stable [all]) +Conf proxmox-backup-server (4.2.5-1 Proxmox Backup System Debian Repository:stable [amd64]) diff --git a/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step4b-upgrade.txt b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step4b-upgrade.txt new file mode 100644 index 00000000..6e03dc18 --- /dev/null +++ b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step4b-upgrade.txt @@ -0,0 +1,58 @@ +### UPGRADE START +2026-08-18T09:51:00+00:00 +Reading package lists... +Building dependency tree... +Reading state information... +The following additional packages will be installed: + proxmox-backup-client proxmox-backup-docs proxmox-enterprise-support-keyring +The following NEW packages will be installed: + proxmox-enterprise-support-keyring +The following packages will be upgraded: + proxmox-backup-client proxmox-backup-docs proxmox-backup-server +3 upgraded, 1 newly installed, 0 to remove and 8 not upgraded. +Need to get 47.8 MB of archives. +After this operation, 1404 kB of additional disk space will be used. +Get:1 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 proxmox-backup-client amd64 4.2.5-1 [3447 kB] +Get:2 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 proxmox-backup-docs all 4.2.5-1 [6351 kB] +Get:3 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 proxmox-enterprise-support-keyring all 1.1 [2988 B] +Get:4 http://download.proxmox.com/debian/pbs trixie/pbs-no-subscription amd64 proxmox-backup-server amd64 4.2.5-1 [38.0 MB] +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +locale: Cannot set LC_CTYPE to default locale: No such file or directory +locale: Cannot set LC_ALL to default locale: No such file or directory +Fetched 47.8 MB in 1s (39.3 MB/s) +(Reading database ... (Reading database ... 5% (Reading database ... 10% (Reading database ... 15% (Reading database ... 20% (Reading database ... 25% (Reading database ... 30% (Reading database ... 35% (Reading database ... 40% (Reading database ... 45% (Reading database ... 50% (Reading database ... 55% (Reading database ... 60% (Reading database ... 65% (Reading database ... 70% (Reading database ... 75% (Reading database ... 80% (Reading database ... 85% (Reading database ... 90% (Reading database ... 95% (Reading database ... 100% (Reading database ... 43107 files and directories currently installed.) +Preparing to unpack .../proxmox-backup-client_4.2.5-1_amd64.deb ... +Unpacking proxmox-backup-client (4.2.5-1) over (4.2.2-1) ... +Preparing to unpack .../proxmox-backup-docs_4.2.5-1_all.deb ... +Unpacking proxmox-backup-docs (4.2.5-1) over (4.2.2-1) ... +Selecting previously unselected package proxmox-enterprise-support-keyring. +Preparing to unpack .../proxmox-enterprise-support-keyring_1.1_all.deb ... +Unpacking proxmox-enterprise-support-keyring (1.1) ... +Preparing to unpack .../proxmox-backup-server_4.2.5-1_amd64.deb ... +Unpacking proxmox-backup-server (4.2.5-1) over (4.2.2-1) ... +Setting up proxmox-backup-docs (4.2.5-1) ... +Setting up proxmox-backup-client (4.2.5-1) ... +Setting up proxmox-enterprise-support-keyring (1.1) ... +Setting up proxmox-backup-server (4.2.5-1) ... +Processing triggers for man-db (2.13.1-1) ... +### apt exit: 0 +### UPGRADE END +2026-08-18T09:51:06+00:00 diff --git a/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step5a-verify-ep0.txt b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step5a-verify-ep0.txt new file mode 100644 index 00000000..7b4846cb --- /dev/null +++ b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step5a-verify-ep0.txt @@ -0,0 +1,48 @@ +### verify-time +2026-08-18T09:51:22+00:00 +### dpkg -l (INSTALLED) +ii proxmox-backup-client 4.2.5-1 amd64 Proxmox Backup Client tools +ii proxmox-backup-docs 4.2.5-1 all Proxmox Backup Documentation +ii proxmox-backup-server 4.2.5-1 amd64 Proxmox Backup Server daemon with tools and GUI +ii proxmox-enterprise-support-keyring 1.1 all Proxmox enterprise SSH support keyring +### daemons +active +inactive +active +### proxy MainPID (should DIFFER from 542065 = it restarted) +551655 +### proxy start time +Tue Aug 18 09:51:04 2026 +### EFFECTIVE open files on the RUNNING proxy (must be 65536) +Max open files 65536 65536 files +### drop-ins still present? +/etc/systemd/system/proxmox-backup-proxy.service.d/: +total 16 +drwxr-xr-x 2 root root 4096 Aug 18 03:54 . +drwxr-xr-x 26 root root 4096 Jul 27 07:53 .. +-rw-r--r-- 1 root root 44 Jul 27 07:16 10-datastore-mount.conf +-rw-r--r-- 1 root root 337 Aug 18 03:54 20-nofile.conf + +/etc/systemd/system/proxmox-backup.service.d/: +total 16 +drwxr-xr-x 2 root root 4096 Aug 18 03:55 . +drwxr-xr-x 26 root root 4096 Jul 27 07:53 .. +-rw-r--r-- 1 root root 44 Jul 27 07:16 10-datastore-mount.conf +-rw-r--r-- 1 root root 237 Aug 18 03:55 20-nofile.conf +### fd count (new t0) +17 +### listen queue (Recv-Q must be 0) +State Recv-Q Send-Q Local Address:Port Peer Address:Port +LISTEN 0 1024 *:8007 *:* +### socket states + 1 LISTEN +### loopback +http=200 time=0.012245 +### version string +proxmox-backup-server 4.2.5-1 running version: 4.2.5 +### datastore ++================+====================+=========+ +| name | path | comment | ++================+====================+=========+ +| felhom-offsite | /mnt/pbs-datastore | | ++================+====================+=========+ diff --git a/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step5b-unit-check.txt b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step5b-unit-check.txt new file mode 100644 index 00000000..7346b379 --- /dev/null +++ b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step5b-unit-check.txt @@ -0,0 +1,11 @@ +### does proxmox-backup-api exist? +No files found for proxmox-backup-api.service. +### actual PBS units + proxmox-backup-banner.service loaded active exited Proxmox Backup Server Login Banner + proxmox-backup-daily-update.service loaded inactive dead Daily Proxmox Backup Server update and maintenance activities + proxmox-backup-proxy.service loaded active running Proxmox Backup API Proxy Server + proxmox-backup.service loaded active running Proxmox Backup API Server +### API server description +MainPID=551642 +Description=Proxmox Backup API Server +ActiveState=active diff --git a/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step5c-verify-felhom-pve.txt b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step5c-verify-felhom-pve.txt new file mode 100644 index 00000000..b3781962 --- /dev/null +++ b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step5c-verify-felhom-pve.txt @@ -0,0 +1,5 @@ +### felhom-pve +2026-08-18T11:51:48+02:00 +http=200 time=0.103326 +Name Type Status Total (KiB) Used (KiB) Available (KiB) % +felhom-pbs pbs active 0 0 0 0.00% diff --git a/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step5d-verify-demo-hp.txt b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step5d-verify-demo-hp.txt new file mode 100644 index 00000000..7c6751e3 --- /dev/null +++ b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step5d-verify-demo-hp.txt @@ -0,0 +1,5 @@ +### demo-hp +2026-08-18T11:51:50+02:00 +http=200 time=0.095982 +Name Type Status Total (KiB) Used (KiB) Available (KiB) % +felhom-pbs pbs active 0 0 0 0.00% diff --git a/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step5e-hub-pbsdr-post.txt b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step5e-hub-pbsdr-post.txt new file mode 100644 index 00000000..a5d92b5b --- /dev/null +++ b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step5e-hub-pbsdr-post.txt @@ -0,0 +1 @@ +POST-UPGRADE REFRESH: 2026/08/18 11:59:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB) diff --git a/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step5e-hub-pbsdr.txt b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step5e-hub-pbsdr.txt new file mode 100644 index 00000000..dffa85c3 --- /dev/null +++ b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step5e-hub-pbsdr.txt @@ -0,0 +1,3 @@ +2026/08/18 11:12:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB) +2026/08/18 11:28:30 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB) +2026/08/18 11:43:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB) diff --git a/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step6-post-upgrade-slope.txt b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step6-post-upgrade-slope.txt new file mode 100644 index 00000000..6db5f00d --- /dev/null +++ b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step6-post-upgrade-slope.txt @@ -0,0 +1,18 @@ +### step6-second-reading (post-upgrade) +2026-08-18T10:23:21+00:00 +### MainPID +551655 +### proxy start +Tue Aug 18 09:51:04 2026 +### fd count +22 +### socket states + 5 ESTAB + 1 LISTEN +### close-wait +1 +### listen queue +State Recv-Q Send-Q Local Address:Port Peer Address:Port +LISTEN 0 1024 *:8007 *:* +### limits +Max open files 65536 65536 files diff --git a/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step6-slope-computation.txt b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step6-slope-computation.txt new file mode 100644 index 00000000..161605d3 --- /dev/null +++ b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/step6-slope-computation.txt @@ -0,0 +1,20 @@ +AFTER-SLOPE — proxy PID 551655, started 2026-08-18 09:51:04 UTC (the upgrade restart) + + 09:51:22Z fd=17 (ESTAB 0) -> 10:23:21Z fd=22 (ESTAB 5) + interval 0.5331 h (1919 s), delta 5 fd + = 9.38 fd/hour = 225.1 fd/DAY + +BEFORE (31 min window): 183.3/day from delta=4 over 1885 s +AFTER (32 min window): 225.1/day from delta=5 over 1919 s + +VERDICT: NOT DISTINGUISHABLE. The two windows differ by ONE descriptor. +With counts this small the Poisson uncertainty on n=4 is +/-2 and on n=5 is +/-2.2, +so both are consistent with a single unchanged underlying rate. The after-figure being +numerically HIGHER is noise, not a regression -- and it is certainly not an improvement. + +This is the EXPECTED result: the 4.2.2->4.2.5 changelog contains no mechanism by which +connection reaping would change. Recorded before the numbers existed, in stop1-ruling.txt. + +Composition again favours ESTAB: 0 -> 5 ESTAB, 0 -> 1 CLOSE-WAIT. + +30 MINUTES CANNOT SETTLE THIS. The honest checks are +24 h and +7 d -> filed as R-341. diff --git a/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/stop1-ruling.txt b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/stop1-ruling.txt new file mode 100644 index 00000000..84fd013b --- /dev/null +++ b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/stop1-ruling.txt @@ -0,0 +1,25 @@ +STOP 1 — operator ruling +======================== +Recorded: 2026-08-18 ~09:25 UTC (11:25 CEST) +Given by: Viktor (operator), in session. + +CC's recommendation was: DO NOT UPGRADE for the stated reason. +Basis: the full changelog range 4.2.2-1 -> 4.2.5-1 (128 lines, all three +entries read) contains NOTHING touching connection handling, descriptor +lifetime, accept(), CLOSE-WAIT, keep-alive or the proxy daemon's socket +lifecycle. An explicit keyword sweep returned exactly one hit, and it is a +false positive ("S3 ... honor the node's proxy settings" = HTTP proxy config +for S3 requests, not the PBS proxy daemon). + +OPERATOR RULING: PROCEED WITH THE UPGRADE ANYWAY. +Stated reason: rehearsal value -- "see how that works for us, we need +practice with that too". The upgrade is therefore being performed as a +supervised practice run of the upgrade procedure on a Tier-2 protected +machine, NOT as a fix for R-336's leak. + +CONSEQUENCE TO CARRY INTO THE REPORT, so it is not later misread: +this upgrade is NOT expected to change the fd slope. If the post-upgrade +slope differs, that is a surprise requiring explanation, not a confirmation +of anything -- the changelog gives no mechanism by which it should improve. +Equally, if the slope is unchanged, that is the EXPECTED result and is not +evidence the upgrade failed. diff --git a/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/stop2-snapshot.txt b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/stop2-snapshot.txt new file mode 100644 index 00000000..ae0f307d --- /dev/null +++ b/documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/stop2-snapshot.txt @@ -0,0 +1,31 @@ +STOP 2 — Hetzner snapshot (the rollback) +======================================== +Confirmed by the operator in the Hetzner console, 2026-08-18 ~09:27 UTC (11:27 CEST). + + Snapshot ID : 421440873 + Description : felhom-hetzner-20260818 + Image size : 15.06 GB + Status : Available <-- COMPLETE, not merely started + Created : "less than a minute ago" as displayed at confirmation + Server : felhom-hetzner #147604682 (CX33, 167.233.158.164) + Hetzner project: 15217960 (Felhom.eu) + +This is the rollback point for the 4.2.2-1 -> 4.2.5-1 upgrade performed in +Step 4. It was taken on a RUNNING server: Hetzner's own console recommends +powering off first for disk consistency, and that was not done -- a +deliberate trade, because powering off ep0 takes the only off-premises copy +offline and the upgrade being guarded is a userspace package install that +touches neither the datastore nor the boot path. + +WHAT THIS SNAPSHOT DOES AND DOES NOT COVER: + covers - the 38 GB system disk /dev/sda (root), i.e. the PBS packages, + unit files, /etc/systemd drop-ins, nftables and wg config. + DOES NOT - /mnt/pbs-datastore. That is /dev/sdb, a separate 100 GB + VOLUME, and Hetzner server snapshots do not include attached + volumes. The backup data is therefore NOT protected by this + snapshot. + Consequence: rolling back restores the software state, not the datastore. + This is acceptable for THIS change because the upgrade writes no datastore + content -- but it must not be mistaken for a datastore backup, and any + future runbook step that could touch /mnt/pbs-datastore needs a different + safeguard than this one. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index b125b51e..37adc290 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -636,6 +636,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-333** | **Two disk-health questions the deploy raised and did NOT act on.** **(a) The 55/60 °C bands are SPINNING-DISK bands applied to NVMe.** They were adopted unchanged from the operator's Prometheus config so the two systems cannot disagree — a deliberate, stated decision — but **measured on demo-hp 2026-08-14 the healthy Toshiba KXG50PNV1T02 NVMe idles at 53 °C, two degrees below Figyelmeztetés and seven below Hiba**, and NVMe routinely exceeds 60 °C under load with no fault whatever. As it stands a healthy customer NVMe under sustained write can be reported as **Hiba** — the single worst outcome this feature can produce. **(b) The agent runs bare `smartctl -a -j` with no `-n standby`** (`felhom-agent/internal/storage/hostops.go:368`), so every poll WAKES a spun-down drive; going 6h → hourly multiplies that by six. demo-hp is all-flash so the cadence measurement could not reveal it, and it was recorded rather than acted on per the task's own instruction. Mitigating datum from the fixture: the failing drive logged only **3375 load cycles in 60505 hours** (~one per 18h), i.e. that duty cycle barely spins down at all | **READY (S each) — NEW 2026-08-14** | — | (a) split the temperature bands by device class, or drop them for NVMe and rely on `critical_warning`; (b) add `-n standby` to the agent's smartctl invocation (an agent change, so fold it into R-330's session) | Viktor decides (a); CC does (b) | | **R-334** | **WAIVER + open item: controller v0.215.0 is released and deployed, and NO golden carries it.** Convicted by `golden_currency_gate.py` on the 2026-08-14 push: newest released controller **0.215.0**, newest golden bake **0.214.0** (`documentation/tests/golden-0.214.0-2026-08-12`). **A machine installed right now receives 0.214.0** — i.e. a brand-new box would ship WITHOUT the R-328 severity fix and would keep emailing nobody about a failing disk. The running fleet is unaffected (demo-hp guest 9201 is on 0.215.0 and healthy); this is purely the day-0 install path. **Not baked in this session deliberately:** the task scoped deployment to demo-hp only, and the second half of the fix — vouching the bake in the hub's day-0 artifact manifest — is **operator-password-gated, so CC cannot complete it**; a baked-but-unvouched golden is worse than none. **The push was made with `git push --no-verify` and it is stated here and in the session report**, per `.claude/rules/gates.md` — the gate has no waiver parser, so recording a waiver does not clear it. **STILL OPEN and now one version WIDER, 2026-08-18:** the newest released controller is **0.216.0** (v0.215.0's own follow-up fix, R-335) and the newest bake is still **0.214.0**, so a new install now misses *two* releases. Re-convicted on this date's documentation-only push, which was likewise made with `--no-verify`; the gate reads `felhom-controller/CHANGELOG.md` and `documentation/tests/golden-*`, **neither of which that session touched** — the conviction is inherited, not caused | **READY (S) — NEW 2026-08-14, re-confirmed 2026-08-18** | operator availability for the vouch step | Bake a golden on **0.216.0** per `runbooks/RUNBOOK-manual-build.md` §4.1, then vouch it — a THREE-field change (`golden_version` + `agent_version` + `min_agent`). Until then every NEW install lacks the severity fix | CC bakes; **Viktor vouches** | | **R-335** | **One physical disk was walked TWICE per run, and the second walk sustained it against itself.** Found on live hardware ~2h after the v0.215.0 deploy, **by noticing the release's own positive observable disagreed with its own persisted artefact**: the hourly check logged *"3 disk(s) evaluated"* while `disk-health-state.json` held **two** records. Cause: demo-hp's `c11-scratch` and `felhom-backup` are the same physical NVMe (`/dev/nvme0n1`) and resolve to the same `diskKey`. **Not cosmetic** — `RunDiskHealthCheck` writes a disk's new record before the next entry reads it, so the SECOND copy consumed the FIRST copy's write as its prior: the disk **sustained against itself and reached Hiba on a FIRST sighting**, defeating truth-table row 6 — the exact rule separating a one-hour benign excursion from a false critical — and would have emitted **two identical events** for one drive. **Latent, not active, on demo-hp** (all three entries healthy, zero counters), but any aliased disk developing a single pending sector would have gone straight to Hiba. **This is the shape standing rule 3 warns about: an absent alarm was not evidence — the two artefacts had to be read AGAINST each other** | **CLOSED — controller v0.216.0, 2026-08-14.** Each `diskKey` is evaluated once per run; both entries stay marked `seen` so neither looks like a disappeared disk, and the card still renders both storage rows (the dedup is about state and alerts, not display). Pinned by `TestDiskCheck_SameDiskTwiceIsEvaluatedOnce`; companion red-proof run and reverted — deleting the guard makes the first sighting emit `Kind:2` (Hiba-from-sectors) at 8 sectors | — | — | CC | -| **R-336** | **The offsite DR endpoint is polled about once per second, and that is what turned a slow leak into an outage.** ep0's PBS proxy served **~85,000 requests/day** — a flat **3,538/hour**, every hour, from two boxes: `74,445 GET /api2/json/admin/datastore` (`libwww-perl`, i.e. PVE's `pvestatd`) and `73,171 GET /admin/datastore/felhom-offsite/status` (`proxmox-backup-client`). Two pollers asking substantially the same question at the same rate. On 2026-08-18 this walked a connection leak in the proxy to its 1024-fd soft limit in **14 days**, wedging the offsite tier for 9½ hours (`audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md`). The `LimitNOFILE=65536` drop-in applied that morning raises the ceiling **but does not fix the leak** — it converts a fortnightly outage into a multi-year one, which is mitigation, not a fix. A DR endpoint that is written to weekly does not need to be asked about every second | **READY (M) — NEW 2026-08-18** | — | Find what polls `felhom-pbs` this hard (PVE storage status is the prime suspect, and its interval is tunable) and cut it; then confirm the fd count between restarts stops climbing — the positive observable, per standing rule 3. **Baseline measured 2026-08-18: 19 fds at 16m43s post-restart, ~85/day, matching the ~73/day implied by the original failure — the leak is confirmed live, not assumed** | CC | +| **R-336** | **The offsite DR endpoint is polled about once per second, and that is what turned a slow leak into an outage.** ep0's PBS proxy served **~85,000 requests/day** — a flat **3,538/hour**, every hour, from two boxes: `74,445 GET /api2/json/admin/datastore` (`libwww-perl`, i.e. PVE's `pvestatd`) and `73,171 GET /admin/datastore/felhom-offsite/status` (`proxmox-backup-client`). Two pollers asking substantially the same question at the same rate. On 2026-08-18 this walked a connection leak in the proxy to its 1024-fd soft limit in **14 days**, wedging the offsite tier for 9½ hours (`audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md`). The `LimitNOFILE=65536` drop-in applied that morning raises the ceiling **but does not fix the leak** — it converts a fortnightly outage into a multi-year one, which is mitigation, not a fix. A DR endpoint that is written to weekly does not need to be asked about every second | **READY (M) — NEW 2026-08-18** | — | Find what polls `felhom-pbs` this hard (PVE storage status is the prime suspect, and its interval is tunable) and cut it; then confirm the fd count between restarts stops climbing — the positive observable, per standing rule 3. **Baseline measured 2026-08-18, and the FIRST measurement published was WRONG.** The initial "~85/day, matching the ~73/day implied by the failure" came from a single 17-minute window whose delta was **one descriptor** — a sample of one cannot carry a daily rate, and the agreement that made it feel solid was coincidence. **Re-measured over two independent windows the same morning: 183/day (31 min) and 200/day (5.6 h)** — ~2.6x the published figure, putting the runway to the 65536 ceiling at **~357 days, not the ~2 years first claimed**. **And the named mechanism is the minority one:** across that window `CLOSE-WAIT` held flat at 1 while `ESTAB` grew 45→49 — *all* the growth was established connections, and at the wedge the split was 1011 ESTAB / 543 CLOSE-WAIT. **The fix must target connections the proxy never reaps, not just `CLOSE-WAIT` sockets.** The PBS 4.2.5-1 upgrade (2026-08-18) did NOT change the slope and was never expected to — see R-341 | CC | | **R-337** | **`/backup/status` lagged a completed backup by minutes on one box and not the other — and it RESOLVED ITSELF, which is why this is WATCHING and not a defect.** During the R-336 recovery on 2026-08-18, `demo-hp`'s snapshot landed on ep0 at **03:58:43Z** (complete manifest; the host's own task index says `OK`) — yet `GET /backup/status` was **still serving the superseded 03:27:00Z failure at ~04:03Z**, four-plus minutes later. `demo-felhom` showed its new result within ~40 s of completion. **The lag cleared on its own:** demo-hp's 04:07:35Z host report carries `felhom-pbs success=true, 4.29 GB`, and the hub is green for both boxes. **The first draft of this row claimed the success was "still reported as failed" — that was written before the next report arrived and it was wrong; the corrected claim is a several-minute skew between the two boxes, not a stuck value.** It is recorded because a status field that can trail its own artifact by minutes will, during an incident, be read as a second failure — this session nearly did — and because the asymmetry between the two boxes is unexplained | **WATCHING — NEW 2026-08-18** | another observation, ideally during an incident rather than constructed | **Do not open a fix on this as written.** First establish the intended refresh path for `/backup/status` after an out-of-schedule run; only if the skew is not simply collection cadence is there anything to pin. If it is cadence, close this row and say so | CC | | **R-338** | **`demo-hp` is not on the R-50 island at all, and `operations/nodes.md` states that it is.** The page records both fleet boxes as island-migrated 2026-07-25. True of `felhom-pve`; **false of `demo-hp`**, whose `agent.json` has `listen_addr: 192.168.0.87:8443` — the customer LAN address — and **no `island_bridge`/`island_guest_addr` keys at all**, whose guest 9201 has `net0` only (no `eth1`), and whose `vmbr9` exists with **zero members**. The controller's `controller.yaml` points at the LAN address, so the box works; this is inventory drift, not breakage. **Two costs.** A session trusting the page addresses the wrong endpoint — that happened on 2026-08-18 and the resulting timeout was briefly read as a fault. And the agent's local API is **bound to the customer LAN on this box** rather than to a point-to-point island, which is the exposure R-50 was built to remove — so a documented security property is claimed for a box that does not have it | **READY (S) — NEW 2026-08-18** | — | Decide which is true: migrate `demo-hp` to the island, or correct `nodes.md`. Leaving both is the one option that keeps the doc lying | Viktor decides; CC executes | +| **R-341** | **Does the fd slope change after the PBS 4.2.5 upgrade? — two dated checks, and the answer is expected to be NO.** ep0 was upgraded 4.2.2-1 → 4.2.5-1 on 2026-08-18 09:51Z on the operator's ruling, **for rehearsal value, not as a fix**: the full changelog range was read (128 lines, all three entries) and swept for connection-handling vocabulary, and it contains **no mechanism** by which descriptor reaping would change — the single keyword hit was `S3 … honor the node's proxy settings`, HTTP-proxy config for S3, not the PBS proxy daemon. **The 32-minute post-upgrade window is indistinguishable from the before window** (+5 fd/1919 s = 225/day vs +4 fd/1885 s = 183/day; the two differ by ONE descriptor and Poisson uncertainty on such counts is ±2, so both are consistent with one unchanged rate — the higher after-figure is noise, not a regression). **Thirty minutes cannot settle it in either direction and this row exists so nobody pretends it did.** New `t0` = **fd 17 at 2026-08-18 09:51:22Z, proxy PID 551655**; before-rate to beat = **183–200/day**. **Interpretation fixed in advance** (`evidence-ep0-pbs-upgrade-2026-08-18/stop1-ruling.txt`, written before any numbers existed): unchanged = EXPECTED, not a failed upgrade; changed = a SURPRISE needing explanation, not a confirmation | **WATCHING — NEW 2026-08-18** | elapsed time only | **Two dated checks, both CC:** **+24 h — 2026-08-19 ~10:00Z** and **+7 d — 2026-08-25 ~10:00Z**. Command (the incident's own positive observable): `ssh root@ 'PID=$(systemctl show proxmox-backup-proxy -p MainPID --value); ls /proc/$PID/fd \| wc -l; ss -lnt "( sport = :8007 )"; ss -tn state all "( sport = :8007 )" \| awk "NR>1{print \$1}" \| sort \| uniq -c'`. **Record the ESTAB/CLOSE-WAIT split, not just the total** — the split is what says which leak it is. If PID ≠ 551655 the window is void: something restarted the proxy and the count began again | CC on both dates |