Files
felhom.eu/REPORT.md
T
2026-08-18 12:27:36 +02:00

191 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# REPORT — RUNBOOK ep0: read the PBS changelog, then decide whether to upgrade (2026-08-18, midday)
**Outcome:** changelog read → **no connection-handling fix in the range**; operator ruled to upgrade
anyway **for rehearsal value**; upgraded **4.2.2-1 → 4.2.5-1** cleanly; **the fd slope did not change,
which is the predicted result.** Two dated checks filed as **R-341**.
**Both STOPs cleared by the operator.** No code changed in any repo; `documentation/` only.
Evidence: `documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/`.
---
## 1. Baselines re-confirmed on the machine
Not from the audit file — from `dpkg -l` and `apt-cache policy`, per the runbook.
| | |
|---|---|
| Installed | `proxmox-backup-server` / `-client` **4.2.2-1** |
| Candidate | **4.2.5-1** (4.2.3-1, 4.2.4-1 also available) |
| Proxy PID / started | 542065, **03:54:53Z**, unrestarted since the incident |
| Effective `open files` | **65536 / 65536** — the morning's drop-in in force |
| `Recv-Q` / loopback | **0** / **`200` in 11 ms** |
| Datastore | 3.7 G of 98 G, 4% |
| felhom.eu `main` @ start | `435e044` — matches the runbook's stated baseline |
## 2. The before-slope — and a correction I owe the morning's report
Two independent windows on the same proxy generation:
| window | from → to | delta | rate |
|---|---|---|---|
| 31 min | 09:18:21Z fd=62 → 09:49:46Z fd=66 | +4 | **183/day** |
| 5.64 h | 04:11:36Z fd=19 → 09:49:46Z fd=66 | +47 | **200/day** |
**This morning's incident note said ~85/day and "≈2 years of runway". Both were wrong.** They were
extrapolated from a single 17-minute window whose delta was **one descriptor** — a sample of one
cannot carry a daily rate, and the agreement with the historical ~73/day that made it feel solid was
coincidence. **The real rate is ~185–200/day, ~2.6× what I published, and the runway is ~357 days,
not two years.** Corrected in the incident document and in R-336 rather than left standing.
**The mechanism I named was also the minority one.** `CLOSE-WAIT` held flat at **1** across the
window while `ESTAB` grew **45 → 49** — *all* the growth was established connections. At the wedge
the split was **1011 ESTAB / 543 CLOSE-WAIT**, so ESTAB dominated there too. **R-336's fix must target
connections the proxy never reaps, not just `CLOSE-WAIT` sockets.**
## 3. The changelog, verbatim — the run's primary deliverable
All three entries between 4.2.2-1 and 4.2.5-1 read in full (128 lines), then swept for
`connection|file descriptor|fd|accept(|close_wait|keep-alive|socket|EMFILE|nofile|leak|proxy|listen|
backlog|hyper|tokio`.
**Exactly one hit, and it is a false positive:**
> `* S3: config: allow editing the use-node-config flag that controls whether requests S3 endpoints`
> ` honor the node's proxy settings or not`
HTTP-proxy configuration *for S3 requests* — not the `proxmox-backup-proxy` daemon.
What the range does contain. **4.2.5-1** — a security release hardening client-supplied manifests:
> `* backup: harden the handling of client supplied backup manifests:`
> ` - only accept archive names that are plain file names carrying a server side type extension. A`
> ` crafted name in a manifest could previously make a sync job read or write outside of the`
> ` snapshot directory, running as the unprivileged 'backup' user.`
> ` - keep an uploaded manifest in memory and only persist it on backup finish, checking that every`
> ` archive it lists was really uploaded during that session and that the checksums match`
plus a sync/push chunk-reuse fix and a subscription-key architecture check. **4.2.4-1** — S3 rate
limits, a file-locking user-lookup cache, the new `proxmox-enterprise-support-keyring` dependency,
docs. **4.2.3-1** — UI/journal work, an LDAP search-filter escape, tape and timezone fixes.
**Nothing addresses descriptor lifetime or connection reaping. My recommendation was: do not upgrade
for this reason.**
## 4. STOP 1 — the ruling
**Operator (Viktor) ruled: upgrade anyway, for rehearsal value** — *"see how that works for us, we
need practice with that too"*. Legitimate and recorded as such: this was **a practice run of the
upgrade procedure on a Tier-2 protected machine, not a fix for the leak**.
**The interpretation was fixed in writing before any numbers existed** (`stop1-ruling.txt`):
unchanged slope = **expected**, not a failed upgrade; changed slope = a **surprise** needing
explanation, not a confirmation. That file was written at the ruling, not afterwards, so neither
outcome could be rationalised into a success.
## 5. STOP 2 — the snapshot, and what it does not cover
Snapshot **421440873** `felhom-hetzner-20260818`, 15.06 GB, status **Available** (complete, not
merely started), server #147604682, project 15217960.
**It covers `/dev/sda` only.** `/mnt/pbs-datastore` is `/dev/sdb`, a separate 100 GB **Volume**, and
Hetzner server snapshots exclude attached volumes — so this is a rollback for the *software* state
(packages, unit files, the `LimitNOFILE` drop-ins, nftables, wg) and **not a backup of the backup
data**. Fine for a package install that writes no datastore content; **it must not be remembered as
datastore protection.** Taken on a running server, deliberately: powering off ep0 to guard a userspace
package install would take the only off-premises copy offline.
## 6. The upgrade and its verification
Simulated first (`-s`): **0 to remove**, so the abort condition never triggered. Then
`apt-get install --only-upgrade -y proxmox-backup-server`, **09:51:00→09:51:06Z, exit 0**. Upgraded
server/client/docs to 4.2.5-1 plus one new dependency, `proxmox-enterprise-support-keyring 1.1` —
**which the 4.2.4-1 changelog had declared**, a small real consistency check between what I read and
what apt did.
| check | result |
|---|---|
| installed | server / client / docs **4.2.5-1** |
| daemons | `proxmox-backup-proxy` **active running**, `proxmox-backup` **active running** |
| proxy restarted | 542065 → **551655** @ 09:51:04 |
| **effective `open files`** | **65536 / 65536** — survived the new package |
| drop-ins on disk | both present, unmodified |
| `Recv-Q` | **0** |
| loopback | **`200` in 12 ms** |
| `felhom-pve` over tunnel | **`200` in 0.103 s**, `felhom-pbs active` |
| `demo-hp` over tunnel | **`200` in 0.096 s**, `felhom-pbs active` |
| hub gauge, post-upgrade | `11:59:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)` |
`proxmox-backup-manager version` now reads `4.2.5-1 running version: 4.2.5`. **This independently
settles the morning's confusion**: that string was never reporting a stale daemon, and now that
installed and running genuinely match, both halves agree.
**One false alarm, mine.** I queried `systemctl is-active proxmox-backup-api` and got `inactive`.
**That unit does not exist** — `systemctl cat` returns *"No files found for
proxmox-backup-api.service"*. The real pair is `proxmox-backup-proxy.service` ("API Proxy Server")
and `proxmox-backup.service` ("API Server"), both active. A bad query, not a fault — recorded because
for as long as it took to check, it looked exactly like one.
**No backup, restore or verify was triggered** to "prove" the endpoint, per the runbook: the reads
above answer it without mutating a protected datastore.
## 7. The after-slope — unchanged, as predicted
| | window | delta | rate |
|---|---|---|---|
| before (PID 542065) | 09:18:21Z fd=62 → 09:49:46Z fd=66 | +4 / 1885 s | **183/day** |
| after (PID 551655) | 09:51:22Z fd=17 → 10:23:21Z fd=22 | +5 / 1919 s | **225/day** |
**These are not distinguishable.** The windows differ by **one descriptor**; Poisson uncertainty on
n=4 is ±2 and on n=5 is ±2.2, so both are consistent with a single unchanged rate. **The after-figure
being numerically higher is noise, not a regression — and certainly not an improvement.** Composition
repeats the pattern: **ESTAB 0 → 5, CLOSE-WAIT 0 → 1.**
**Thirty minutes cannot settle this in either direction, and nothing here claims it does.**
## 8. Register
- **R-341 filed, WATCHING** — the two dated checks: **+24 h (2026-08-19 ~10:00Z)** and
**+7 d (2026-08-25 ~10:00Z)**, against new `t0` **fd=17 @ 09:51:22Z, PID 551655**, with the exact
command and the instruction to **record the ESTAB/CLOSE-WAIT split, not just the total** — the
split is what identifies which leak it is. If the PID has changed, the window is void.
- **R-336 updated and STILL OPEN.** Its wrong 85/day baseline is corrected to 183–200/day, the
mechanism is re-pointed at ESTAB, and the upgrade's null result is recorded. **It stays open on its
own merits:** an upgrade that *had* fixed the leak still would not make ~85,000 requests/day to a
weekly-write DR endpoint correct.
- `documentation/runbooks/offsite-endpoint.md` — **no change made, correctly**: it states no PBS
version literal, only package names, so there was nothing to update (and per `docs.md`, version
literals do not belong in a current-state doc anyway).
## 9. Teardown
**This run provisioned nothing.** No `.deb` was downloaded — `apt-get changelog` served the text
directly, so the runbook's `/tmp` cleanup was never needed. The Hetzner snapshot **is retained** as
the rollback; deleting it is an operator decision, and it costs €0.018161/GB/month.
## 10. Observations, not acted on
- **`pvesm status` reports `felhom-pbs` with Total/Used/Available all `0`** on both boxes while
status reads `active`. Consistent before and after the upgrade, and the hub's own gauge reads the
real 3.7 GB / 97.9 GB, so nothing is broken — but the PVE-side numbers are not usable as a capacity
signal. Likely the namespace-scoped token lacking datastore-level audit. Not filed; noted here.
- **8 other packages are held back** (`8 not upgraded`), untouched deliberately — `full-upgrade` on
this machine is forbidden by the runbook and would change the kernel and WireGuard alongside the
thing under test.
- **The morning's `--no-verify` situation is unchanged**: `golden-currency` still convicts on the
inherited R-334 (controller 0.216.0 vs golden 0.214.0), untouched by this run.
## 11. CI, checked by run ID
**Run `351`, `head_sha 3e50902a9`, conclusion `failure`, elapsed 15 s** (10:27:00→10:27:15Z).
Inherited, not caused: CI's only step is `python3 scripts/repo_gates.py --fast`, which convicts
`golden-currency` on **R-334** (controller 0.216.0 vs newest golden 0.214.0) — files this run did not
touch. 15 s is the workflow's own honest-failure band, not the R-265 reap band. The previous run
`350` on `435e044` failed identically, before this run began.
**Expect one `[felhom CI] gates FAILED` mail for run 351.** Same cause as run 348 this morning; not a
new fault, and not related to the upgrade.
`python3 scripts/unproven.py --summary` — **unchanged: 23 walked, 32 not walked of 55.** No number
moved, correctly: this run proved an operational fact, not a product claim.