Files
felhom.eu/REPORT.md
T
admin 3e50902a98
gates / gates (push) Failing after 14s
RUNBOOK ep0: PBS 4.2.2-1 -> 4.2.5-1, slope unchanged as predicted (R-341)
Both STOPs cleared by the operator. No code changed; documentation only.

STEP 3 (the run's primary deliverable): the full changelog range 4.2.2-1 ->
4.2.5-1 was read (128 lines, all three entries) and swept for
connection-handling vocabulary. Exactly one keyword hit, a false positive
("S3 ... honor the node's proxy settings" = HTTP proxy config for S3, not the
PBS proxy daemon). 4.2.5-1 is a manifest-hardening security release; 4.2.4-1
is S3 rate limits and a locking cache; 4.2.3-1 is UI/LDAP/tape. NOTHING
addresses descriptor lifetime or connection reaping. Recommendation was: do
not upgrade for this reason.

STOP 1: operator ruled to upgrade anyway for rehearsal value. Recorded as a
practice run, not a fix -- and the interpretation was fixed IN WRITING BEFORE
any numbers existed (stop1-ruling.txt): unchanged = expected; changed =
surprise. Neither outcome could then be rationalised into a success.

STOP 2: Hetzner snapshot 421440873, Available. Documented that it covers
/dev/sda ONLY -- /mnt/pbs-datastore is a separate Volume and is NOT in it, so
it is a software rollback and not a backup of the backup data.

UPGRADE: simulated first (0 to remove), then installed 09:51:00->09:51:06Z,
exit 0. Verified: 4.2.5-1 installed, both daemons active, effective open
files still 65536 (the drop-in survived the new package), Recv-Q 0, loopback
200, 200 from BOTH boxes over the tunnel with felhom-pbs active, and the hub
gauge refreshed post-upgrade at 11:59:31.

SLOPE: before +4 fd/1885 s = 183/day; after +5 fd/1919 s = 225/day. NOT
distinguishable -- one descriptor apart, Poisson +/-2 on such counts. The
higher after-figure is noise, not a regression and not an improvement. 30
minutes cannot settle it; R-341 files the +24 h and +7 d checks.

CORRECTIONS to this morning's own report, both published rather than quietly
fixed:
  - the "~85/day, ~2 years of runway" figures were WRONG. They came from a
    single 17-minute window with a delta of ONE descriptor. Real rate is
    183-200/day over two independent windows; runway ~357 days, not 2 years.
  - the leak was attributed to CLOSE-WAIT. It is mostly ESTAB: CLOSE-WAIT held
    flat at 1 while ESTAB grew 45->49, and at the wedge it was 1011 ESTAB vs
    543 CLOSE-WAIT. R-336's fix must target unreaped connections.
  - "proxmox-backup-api" reported inactive during verification; that unit does
    not exist. Bad query, not a fault, written down because it looked like one.

R-336 stays open: even a fixed leak would not make ~85k requests/day to a
weekly-write DR endpoint correct.

golden-currency still convicts (inherited R-334, controller 0.216.0 vs golden
0.214.0, untouched by this run), so this push is --no-verify per
.claude/rules/gates.md.
2026-08-18 12:26:53 +02:00

177 lines
9.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# REPORT — RUNBOOK ep0: read the PBS changelog, then decide whether to upgrade (2026-08-18, midday)
**Outcome:** changelog read → **no connection-handling fix in the range**; operator ruled to upgrade
anyway **for rehearsal value**; upgraded **4.2.2-1 → 4.2.5-1** cleanly; **the fd slope did not change,
which is the predicted result.** Two dated checks filed as **R-341**.
**Both STOPs cleared by the operator.** No code changed in any repo; `documentation/` only.
Evidence: `documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/`.
---
## 1. Baselines re-confirmed on the machine
Not from the audit file — from `dpkg -l` and `apt-cache policy`, per the runbook.
| | |
|---|---|
| Installed | `proxmox-backup-server` / `-client` **4.2.2-1** |
| Candidate | **4.2.5-1** (4.2.3-1, 4.2.4-1 also available) |
| Proxy PID / started | 542065, **03:54:53Z**, unrestarted since the incident |
| Effective `open files` | **65536 / 65536** — the morning's drop-in in force |
| `Recv-Q` / loopback | **0** / **`200` in 11 ms** |
| Datastore | 3.7 G of 98 G, 4% |
| felhom.eu `main` @ start | `435e044` — matches the runbook's stated baseline |
## 2. The before-slope — and a correction I owe the morning's report
Two independent windows on the same proxy generation:
| window | from → to | delta | rate |
|---|---|---|---|
| 31 min | 09:18:21Z fd=62 → 09:49:46Z fd=66 | +4 | **183/day** |
| 5.64 h | 04:11:36Z fd=19 → 09:49:46Z fd=66 | +47 | **200/day** |
**This morning's incident note said ~85/day and "≈2 years of runway". Both were wrong.** They were
extrapolated from a single 17-minute window whose delta was **one descriptor** — a sample of one
cannot carry a daily rate, and the agreement with the historical ~73/day that made it feel solid was
coincidence. **The real rate is ~185–200/day, ~2.6× what I published, and the runway is ~357 days,
not two years.** Corrected in the incident document and in R-336 rather than left standing.
**The mechanism I named was also the minority one.** `CLOSE-WAIT` held flat at **1** across the
window while `ESTAB` grew **45 → 49** — *all* the growth was established connections. At the wedge
the split was **1011 ESTAB / 543 CLOSE-WAIT**, so ESTAB dominated there too. **R-336's fix must target
connections the proxy never reaps, not just `CLOSE-WAIT` sockets.**
## 3. The changelog, verbatim — the run's primary deliverable
All three entries between 4.2.2-1 and 4.2.5-1 read in full (128 lines), then swept for
`connection|file descriptor|fd|accept(|close_wait|keep-alive|socket|EMFILE|nofile|leak|proxy|listen|
backlog|hyper|tokio`.
**Exactly one hit, and it is a false positive:**
> `* S3: config: allow editing the use-node-config flag that controls whether requests S3 endpoints`
> ` honor the node's proxy settings or not`
HTTP-proxy configuration *for S3 requests* — not the `proxmox-backup-proxy` daemon.
What the range does contain. **4.2.5-1** — a security release hardening client-supplied manifests:
> `* backup: harden the handling of client supplied backup manifests:`
> ` - only accept archive names that are plain file names carrying a server side type extension. A`
> ` crafted name in a manifest could previously make a sync job read or write outside of the`
> ` snapshot directory, running as the unprivileged 'backup' user.`
> ` - keep an uploaded manifest in memory and only persist it on backup finish, checking that every`
> ` archive it lists was really uploaded during that session and that the checksums match`
plus a sync/push chunk-reuse fix and a subscription-key architecture check. **4.2.4-1** — S3 rate
limits, a file-locking user-lookup cache, the new `proxmox-enterprise-support-keyring` dependency,
docs. **4.2.3-1** — UI/journal work, an LDAP search-filter escape, tape and timezone fixes.
**Nothing addresses descriptor lifetime or connection reaping. My recommendation was: do not upgrade
for this reason.**
## 4. STOP 1 — the ruling
**Operator (Viktor) ruled: upgrade anyway, for rehearsal value** — *"see how that works for us, we
need practice with that too"*. Legitimate and recorded as such: this was **a practice run of the
upgrade procedure on a Tier-2 protected machine, not a fix for the leak**.
**The interpretation was fixed in writing before any numbers existed** (`stop1-ruling.txt`):
unchanged slope = **expected**, not a failed upgrade; changed slope = a **surprise** needing
explanation, not a confirmation. That file was written at the ruling, not afterwards, so neither
outcome could be rationalised into a success.
## 5. STOP 2 — the snapshot, and what it does not cover
Snapshot **421440873** `felhom-hetzner-20260818`, 15.06 GB, status **Available** (complete, not
merely started), server #147604682, project 15217960.
**It covers `/dev/sda` only.** `/mnt/pbs-datastore` is `/dev/sdb`, a separate 100 GB **Volume**, and
Hetzner server snapshots exclude attached volumes — so this is a rollback for the *software* state
(packages, unit files, the `LimitNOFILE` drop-ins, nftables, wg) and **not a backup of the backup
data**. Fine for a package install that writes no datastore content; **it must not be remembered as
datastore protection.** Taken on a running server, deliberately: powering off ep0 to guard a userspace
package install would take the only off-premises copy offline.
## 6. The upgrade and its verification
Simulated first (`-s`): **0 to remove**, so the abort condition never triggered. Then
`apt-get install --only-upgrade -y proxmox-backup-server`, **09:51:00→09:51:06Z, exit 0**. Upgraded
server/client/docs to 4.2.5-1 plus one new dependency, `proxmox-enterprise-support-keyring 1.1` —
**which the 4.2.4-1 changelog had declared**, a small real consistency check between what I read and
what apt did.
| check | result |
|---|---|
| installed | server / client / docs **4.2.5-1** |
| daemons | `proxmox-backup-proxy` **active running**, `proxmox-backup` **active running** |
| proxy restarted | 542065 → **551655** @ 09:51:04 |
| **effective `open files`** | **65536 / 65536** — survived the new package |
| drop-ins on disk | both present, unmodified |
| `Recv-Q` | **0** |
| loopback | **`200` in 12 ms** |
| `felhom-pve` over tunnel | **`200` in 0.103 s**, `felhom-pbs active` |
| `demo-hp` over tunnel | **`200` in 0.096 s**, `felhom-pbs active` |
| hub gauge, post-upgrade | `11:59:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)` |
`proxmox-backup-manager version` now reads `4.2.5-1 running version: 4.2.5`. **This independently
settles the morning's confusion**: that string was never reporting a stale daemon, and now that
installed and running genuinely match, both halves agree.
**One false alarm, mine.** I queried `systemctl is-active proxmox-backup-api` and got `inactive`.
**That unit does not exist** — `systemctl cat` returns *"No files found for
proxmox-backup-api.service"*. The real pair is `proxmox-backup-proxy.service` ("API Proxy Server")
and `proxmox-backup.service` ("API Server"), both active. A bad query, not a fault — recorded because
for as long as it took to check, it looked exactly like one.
**No backup, restore or verify was triggered** to "prove" the endpoint, per the runbook: the reads
above answer it without mutating a protected datastore.
## 7. The after-slope — unchanged, as predicted
| | window | delta | rate |
|---|---|---|---|
| before (PID 542065) | 09:18:21Z fd=62 → 09:49:46Z fd=66 | +4 / 1885 s | **183/day** |
| after (PID 551655) | 09:51:22Z fd=17 → 10:23:21Z fd=22 | +5 / 1919 s | **225/day** |
**These are not distinguishable.** The windows differ by **one descriptor**; Poisson uncertainty on
n=4 is ±2 and on n=5 is ±2.2, so both are consistent with a single unchanged rate. **The after-figure
being numerically higher is noise, not a regression — and certainly not an improvement.** Composition
repeats the pattern: **ESTAB 0 → 5, CLOSE-WAIT 0 → 1.**
**Thirty minutes cannot settle this in either direction, and nothing here claims it does.**
## 8. Register
- **R-341 filed, WATCHING** — the two dated checks: **+24 h (2026-08-19 ~10:00Z)** and
**+7 d (2026-08-25 ~10:00Z)**, against new `t0` **fd=17 @ 09:51:22Z, PID 551655**, with the exact
command and the instruction to **record the ESTAB/CLOSE-WAIT split, not just the total** — the
split is what identifies which leak it is. If the PID has changed, the window is void.
- **R-336 updated and STILL OPEN.** Its wrong 85/day baseline is corrected to 183–200/day, the
mechanism is re-pointed at ESTAB, and the upgrade's null result is recorded. **It stays open on its
own merits:** an upgrade that *had* fixed the leak still would not make ~85,000 requests/day to a
weekly-write DR endpoint correct.
- `documentation/runbooks/offsite-endpoint.md` — **no change made, correctly**: it states no PBS
version literal, only package names, so there was nothing to update (and per `docs.md`, version
literals do not belong in a current-state doc anyway).
## 9. Teardown
**This run provisioned nothing.** No `.deb` was downloaded — `apt-get changelog` served the text
directly, so the runbook's `/tmp` cleanup was never needed. The Hetzner snapshot **is retained** as
the rollback; deleting it is an operator decision, and it costs €0.018161/GB/month.
## 10. Observations, not acted on
- **`pvesm status` reports `felhom-pbs` with Total/Used/Available all `0`** on both boxes while
status reads `active`. Consistent before and after the upgrade, and the hub's own gauge reads the
real 3.7 GB / 97.9 GB, so nothing is broken — but the PVE-side numbers are not usable as a capacity
signal. Likely the namespace-scoped token lacking datastore-level audit. Not filed; noted here.
- **8 other packages are held back** (`8 not upgraded`), untouched deliberately — `full-upgrade` on
this machine is forbidden by the runbook and would change the kernel and WireGuard alongside the
thing under test.
- **The morning's `--no-verify` situation is unchanged**: `golden-currency` still convicts on the
inherited R-334 (controller 0.216.0 vs golden 0.214.0), untouched by this run.