RUNBOOK ep0: PBS 4.2.2-1 -> 4.2.5-1, slope unchanged as predicted (R-341)
gates / gates (push) Failing after 14s
gates / gates (push) Failing after 14s
Both STOPs cleared by the operator. No code changed; documentation only.
STEP 3 (the run's primary deliverable): the full changelog range 4.2.2-1 ->
4.2.5-1 was read (128 lines, all three entries) and swept for
connection-handling vocabulary. Exactly one keyword hit, a false positive
("S3 ... honor the node's proxy settings" = HTTP proxy config for S3, not the
PBS proxy daemon). 4.2.5-1 is a manifest-hardening security release; 4.2.4-1
is S3 rate limits and a locking cache; 4.2.3-1 is UI/LDAP/tape. NOTHING
addresses descriptor lifetime or connection reaping. Recommendation was: do
not upgrade for this reason.
STOP 1: operator ruled to upgrade anyway for rehearsal value. Recorded as a
practice run, not a fix -- and the interpretation was fixed IN WRITING BEFORE
any numbers existed (stop1-ruling.txt): unchanged = expected; changed =
surprise. Neither outcome could then be rationalised into a success.
STOP 2: Hetzner snapshot 421440873, Available. Documented that it covers
/dev/sda ONLY -- /mnt/pbs-datastore is a separate Volume and is NOT in it, so
it is a software rollback and not a backup of the backup data.
UPGRADE: simulated first (0 to remove), then installed 09:51:00->09:51:06Z,
exit 0. Verified: 4.2.5-1 installed, both daemons active, effective open
files still 65536 (the drop-in survived the new package), Recv-Q 0, loopback
200, 200 from BOTH boxes over the tunnel with felhom-pbs active, and the hub
gauge refreshed post-upgrade at 11:59:31.
SLOPE: before +4 fd/1885 s = 183/day; after +5 fd/1919 s = 225/day. NOT
distinguishable -- one descriptor apart, Poisson +/-2 on such counts. The
higher after-figure is noise, not a regression and not an improvement. 30
minutes cannot settle it; R-341 files the +24 h and +7 d checks.
CORRECTIONS to this morning's own report, both published rather than quietly
fixed:
- the "~85/day, ~2 years of runway" figures were WRONG. They came from a
single 17-minute window with a delta of ONE descriptor. Real rate is
183-200/day over two independent windows; runway ~357 days, not 2 years.
- the leak was attributed to CLOSE-WAIT. It is mostly ESTAB: CLOSE-WAIT held
flat at 1 while ESTAB grew 45->49, and at the wedge it was 1011 ESTAB vs
543 CLOSE-WAIT. R-336's fix must target unreaped connections.
- "proxmox-backup-api" reported inactive during verification; that unit does
not exist. Bad query, not a fault, written down because it looked like one.
R-336 stays open: even a fixed leak would not make ~85k requests/day to a
weekly-write DR endpoint correct.
golden-currency still convicts (inherited R-334, controller 0.216.0 vs golden
0.214.0, untouched by this run), so this push is --no-verify per
.claude/rules/gates.md.
This commit is contained in:
@@ -1,133 +1,176 @@
|
||||
# REPORT — a listening socket that served nobody (2026-08-18, early)
|
||||
# REPORT — RUNBOOK ep0: read the PBS changelog, then decide whether to upgrade (2026-08-18, midday)
|
||||
|
||||
**Trigger:** two `whole_guest_backup_failed` alert mails, 04:30 and 04:32 CEST.
|
||||
**Outcome:** root cause found on **ep0**, fixed, both missed backups re-driven and landed.
|
||||
**No repo code changed** — this was an operational run. Documentation, register and evidence only.
|
||||
**Outcome:** changelog read → **no connection-handling fix in the range**; operator ruled to upgrade
|
||||
anyway **for rehearsal value**; upgraded **4.2.2-1 → 4.2.5-1** cleanly; **the fd slope did not change,
|
||||
which is the predicted result.** Two dated checks filed as **R-341**.
|
||||
|
||||
**Both STOPs cleared by the operator.** No code changed in any repo; `documentation/` only.
|
||||
Evidence: `documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/`.
|
||||
|
||||
---
|
||||
|
||||
## 1. What was wrong
|
||||
## 1. Baselines re-confirmed on the machine
|
||||
|
||||
Both alerts were the same incident and neither was on a customer box. `demo-felhom` and `demo-hp`
|
||||
each failed their `felhom-pbs` tier with `Can't connect to 10.77.0.1:8007 (Connection timed out)` —
|
||||
the one thing they share, the Hetzner offsite PBS endpoint.
|
||||
Not from the audit file — from `dpkg -l` and `apt-cache policy`, per the runbook.
|
||||
|
||||
On ep0, `proxmox-backup-proxy` was `active`, held its listening socket, and **served nobody**:
|
||||
| | |
|
||||
|---|---|
|
||||
| Installed | `proxmox-backup-server` / `-client` **4.2.2-1** |
|
||||
| Candidate | **4.2.5-1** (4.2.3-1, 4.2.4-1 also available) |
|
||||
| Proxy PID / started | 542065, **03:54:53Z**, unrestarted since the incident |
|
||||
| Effective `open files` | **65536 / 65536** — the morning's drop-in in force |
|
||||
| `Recv-Q` / loopback | **0** / **`200` in 11 ms** |
|
||||
| Datastore | 3.7 G of 98 G, 4% |
|
||||
| felhom.eu `main` @ start | `435e044` — matches the runbook's stated baseline |
|
||||
|
||||
```
|
||||
ss -lnt '( sport = :8007 )' → LISTEN Recv-Q 1025 Send-Q 1024
|
||||
ls /proc/<proxy>/fd | wc -l → 1024 # == its soft RLIMIT_NOFILE
|
||||
```
|
||||
## 2. The before-slope — and a correction I owe the morning's report
|
||||
|
||||
`Send-Q` on a listener is the accept backlog; `Recv-Q` is the queue depth. At 1025 against 1024 the
|
||||
queue had overflowed, because `accept()` was returning `EMFILE` on every call. **1016 of the 1024
|
||||
descriptors were sockets and 547 connections sat in `CLOSE-WAIT`** — a connection leak, fed by
|
||||
~85,000 requests/day, that reached the ceiling after 14 days of uptime. Last request served:
|
||||
2026-08-17 18:15:32 UTC. Offsite DR was therefore down **≈ 9 h 37 m**.
|
||||
Two independent windows on the same proxy generation:
|
||||
|
||||
**The observation that settled it:** the daemon was wedged from its own loopback too —
|
||||
`curl https://127.0.0.1:8007/` on ep0 timed out. A listener that cannot serve `127.0.0.1` has no
|
||||
network left to blame, and that check is cheap enough to make early.
|
||||
| window | from → to | delta | rate |
|
||||
|---|---|---|---|
|
||||
| 31 min | 09:18:21Z fd=62 → 09:49:46Z fd=66 | +4 | **183/day** |
|
||||
| 5.64 h | 04:11:36Z fd=19 → 09:49:46Z fd=66 | +47 | **200/day** |
|
||||
|
||||
Full record, including everything that was ruled out first (tunnel, nftables, disk, dead daemon):
|
||||
**`documentation/audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md`**.
|
||||
**This morning's incident note said ~85/day and "≈2 years of runway". Both were wrong.** They were
|
||||
extrapolated from a single 17-minute window whose delta was **one descriptor** — a sample of one
|
||||
cannot carry a daily rate, and the agreement with the historical ~73/day that made it feel solid was
|
||||
coincidence. **The real rate is ~185–200/day, ~2.6× what I published, and the runway is ~357 days,
|
||||
not two years.** Corrected in the incident document and in R-336 rather than left standing.
|
||||
|
||||
## 2. What was done
|
||||
**The mechanism I named was also the minority one.** `CLOSE-WAIT` held flat at **1** across the
|
||||
window while `ESTAB` grew **45 → 49** — *all* the growth was established connections. At the wedge
|
||||
the split was **1011 ESTAB / 543 CLOSE-WAIT**, so ESTAB dominated there too. **R-336's fix must target
|
||||
connections the proxy never reaps, not just `CLOSE-WAIT` sockets.**
|
||||
|
||||
1. `LimitNOFILE=65536` drop-ins for `proxmox-backup-proxy.service` and `proxmox-backup.service`, each
|
||||
carrying its reason inline. **The API daemon was not implicated** (15 fds) and its drop-in says so
|
||||
— a later reader must not mistake it for a second culprit.
|
||||
2. Restarted both. After: soft limit 65536, fds back to 18, `Recv-Q 0`, loopback `200`.
|
||||
3. **Verified from the customer side, not only from ep0** — both boxes got `200` in ~0.1 s and
|
||||
`pvesm status` read `felhom-pbs pbs active`.
|
||||
4. Re-drove the missed backups **through the product path** — `POST /backup?target=felhom-pbs` on
|
||||
each agent's local API, issued from inside the guest's controller container with the controller's
|
||||
own credentials, i.e. the same call the scheduler makes. Not a hand-run `vzdump`.
|
||||
## 3. The changelog, verbatim — the run's primary deliverable
|
||||
|
||||
**Result:** `demo-felhom` → `ct/9201/2026-08-18T03:57:43Z` (4.10 GB, 36.4 s);
|
||||
`demo-hp` → `ct/9201/2026-08-18T03:58:43Z` (4.29 GB, 41.5 s). Both directories carry a full manifest
|
||||
on ep0, and both hosts log the re-run vzdump as `OK`. **And the hub agrees** — both boxes' next host
|
||||
reports carry `felhom-pbs success=true` (04:00:33Z and 04:07:35Z), so the operator view went green on
|
||||
the evidence rather than on the fact that a restart was performed.
|
||||
**No data was lost and no backup was skipped** — the daily local
|
||||
tier was never affected (it completed on both boxes at 05:00 and 05:02), and the PBS tier is weekly,
|
||||
so the outage window cost exactly one attempt, which was re-driven the same morning.
|
||||
All three entries between 4.2.2-1 and 4.2.5-1 read in full (128 lines), then swept for
|
||||
`connection|file descriptor|fd|accept(|close_wait|keep-alive|socket|EMFILE|nofile|leak|proxy|listen|
|
||||
backlog|hyper|tokio`.
|
||||
|
||||
Evidence copied off ep0 **before** the restart, per standing rule 5:
|
||||
`documentation/audits/evidence-ep0-fd-2026-08-18/` — pre-restart state, post-fix state, access-log tail.
|
||||
**Exactly one hit, and it is a false positive:**
|
||||
|
||||
## 3. What I got wrong, and corrected
|
||||
> `* S3: config: allow editing the use-node-config flag that controls whether requests S3 endpoints`
|
||||
> ` honor the node's proxy settings or not`
|
||||
|
||||
- **`proxmox-backup-manager version` prints *available* then *running*.** It read
|
||||
`4.2.5-1 running version: 4.2.2`, which looks exactly like a daemon left behind by a package
|
||||
upgrade. It is not: `dpkg -l` shows **4.2.2-1 installed**, 4.2.5-1 merely available in the repo,
|
||||
and the on-disk binary is dated 2026-06-18. I restarted the API daemon on that mistaken reading;
|
||||
harmless, and its drop-in is a genuine improvement, but it was not needed.
|
||||
- **I addressed `demo-hp`'s agent at the island address** because `operations/nodes.md` says that box
|
||||
is island-migrated. It is not (**R-338**), and the resulting timeout was briefly read as a fault.
|
||||
HTTP-proxy configuration *for S3 requests* — not the `proxmox-backup-proxy` daemon.
|
||||
|
||||
## 4. Findings filed — R-336, R-337, R-338
|
||||
What the range does contain. **4.2.5-1** — a security release hardening client-supplied manifests:
|
||||
|
||||
All three are in `documentation/backlog/OPEN-ITEMS.md` with numbers, per the registers-first rule.
|
||||
> `* backup: harden the handling of client supplied backup manifests:`
|
||||
> ` - only accept archive names that are plain file names carrying a server side type extension. A`
|
||||
> ` crafted name in a manifest could previously make a sync job read or write outside of the`
|
||||
> ` snapshot directory, running as the unprivileged 'backup' user.`
|
||||
> ` - keep an uploaded manifest in memory and only persist it on backup finish, checking that every`
|
||||
> ` archive it lists was really uploaded during that session and that the checksums match`
|
||||
|
||||
- **R-336 — the poll rate is the real defect.** ~1 request/second against a DR endpoint written to
|
||||
weekly. `LimitNOFILE` raises the ceiling; **it does not fix the leak**, it converts a fortnightly
|
||||
outage into a multi-year one.
|
||||
- **R-337 — a status endpoint that trailed its own artifact, then caught up. WATCHING, not a defect.**
|
||||
`demo-hp`'s `/backup/status` was still serving the superseded 03:27:00Z failure at ~04:03Z while the
|
||||
snapshot sat on ep0 and the host logged `OK`; `demo-felhom` updated within ~40 s. **I filed this as
|
||||
a defect and that was premature** — the next host report (04:07:35Z) carried the success and the
|
||||
skew cleared with no intervention. Rewritten as WATCHING, with the explicit instruction not to open
|
||||
a fix until someone establishes whether this is just collection cadence. Recorded at all because
|
||||
during the recovery it read as a second failure, and it was not one.
|
||||
- **R-338 — `demo-hp` is not on the R-50 island and `nodes.md` says it is.** No `island_bridge` keys,
|
||||
guest has no `eth1`, `vmbr9` has zero members, and the agent's local API is bound to the customer
|
||||
LAN — the exposure R-50 existed to remove.
|
||||
plus a sync/push chunk-reuse fix and a subscription-key architecture check. **4.2.4-1** — S3 rate
|
||||
limits, a file-locking user-lookup cache, the new `proxmox-enterprise-support-keyring` dependency,
|
||||
docs. **4.2.3-1** — UI/journal work, an LDAP search-filter escape, tape and timezone fixes.
|
||||
|
||||
**Not done, deliberately:** PBS 4.2.5-1 was not applied. Upgrading a production offsite endpoint was
|
||||
outside what this run was authorised to do, and its changelog should be read for the connection-
|
||||
handling leak first.
|
||||
**Nothing addresses descriptor lifetime or connection reaping. My recommendation was: do not upgrade
|
||||
for this reason.**
|
||||
|
||||
## 5. Gates — one pre-existing conviction, and a stated bypass
|
||||
## 4. STOP 1 — the ruling
|
||||
|
||||
`python3 scripts/repo_gates.py --fast` → **rc=1, CONVICTED: golden-currency.** Eight of nine gates
|
||||
pass. The conviction is **R-334, inherited and not caused here**: newest released controller
|
||||
**0.216.0**, newest golden bake **0.214.0**, so a new install misses two releases. The gate reads
|
||||
`felhom-controller/CHANGELOG.md` and `documentation/tests/golden-*` — **this session touched
|
||||
neither**, and its whole diff is documentation. Baking is possible; **vouching is
|
||||
operator-password-gated, and a baked-but-unvouched golden is worse than none**, so it is not a
|
||||
one-sided job CC can finish.
|
||||
**Operator (Viktor) ruled: upgrade anyway, for rehearsal value** — *"see how that works for us, we
|
||||
need practice with that too"*. Legitimate and recorded as such: this was **a practice run of the
|
||||
upgrade procedure on a Tier-2 protected machine, not a fix for the leak**.
|
||||
|
||||
**This push therefore used `git push --no-verify`, stated here per `.claude/rules/gates.md`.**
|
||||
R-334 is updated in the register with the new numbers rather than left reading 0.215.0.
|
||||
**The interpretation was fixed in writing before any numbers existed** (`stop1-ruling.txt`):
|
||||
unchanged slope = **expected**, not a failed upgrade; changed slope = a **surprise** needing
|
||||
explanation, not a confirmation. That file was written at the ruling, not afterwards, so neither
|
||||
outcome could be rationalised into a success.
|
||||
|
||||
**CI checked by run ID, as the checklist requires: run `348`, `head_sha ebfd0967c`, conclusion
|
||||
`failure`, elapsed 13 s** (04:09:45→04:09:58Z). Expected and inherited — CI's only step is
|
||||
`python3 scripts/repo_gates.py --fast`, the same entry point that convicts golden-currency locally,
|
||||
with the sibling repos fetched. The 13 s runtime places it in the workflow's own "honest gate
|
||||
failure" band rather than the R-265 reap band, so the result is the gate speaking, not the runner.
|
||||
**I could not read the run log to name the gate from CI's own mouth** — `actions/runs/348/logs` and
|
||||
`actions/tasks/348/logs` both 404, `actions/runs/348/jobs` returns an empty list, authenticated as
|
||||
`admin`, and the web log endpoint 302s. So this is an inference from the local run plus the workflow
|
||||
definition, not a direct reading, and it is stated as such.
|
||||
## 5. STOP 2 — the snapshot, and what it does not cover
|
||||
|
||||
**Expect one `[felhom CI] gates FAILED in admin/felhom.eu` mail for run 348** — the workflow alarms
|
||||
on failure by design. It is this push, and it is the golden-currency row, not a new fault; the same
|
||||
mails on 12 and 14 August have the same cause.
|
||||
Snapshot **421440873** `felhom-hetzner-20260818`, 15.06 GB, status **Available** (complete, not
|
||||
merely started), server #147604682, project 15217960.
|
||||
|
||||
*(Noted for accuracy: the first gate run was piped to `tail`, which returned `rc=0` — `tail`'s exit
|
||||
code, not the gate's. It was re-run unpiped to read the real `rc=1`. That is standing rule 1's trap
|
||||
in its smaller form, and the number reported above is the unpiped one.)*
|
||||
**It covers `/dev/sda` only.** `/mnt/pbs-datastore` is `/dev/sdb`, a separate 100 GB **Volume**, and
|
||||
Hetzner server snapshots exclude attached volumes — so this is a rollback for the *software* state
|
||||
(packages, unit files, the `LimitNOFILE` drop-ins, nftables, wg) and **not a backup of the backup
|
||||
data**. Fine for a package install that writes no datastore content; **it must not be remembered as
|
||||
datastore protection.** Taken on a running server, deliberately: powering off ep0 to guard a userspace
|
||||
package install would take the only off-premises copy offline.
|
||||
|
||||
## 6. What to watch
|
||||
## 6. The upgrade and its verification
|
||||
|
||||
The positive observable is the descriptor count, not the absence of an alert — an empty alert queue
|
||||
is equally consistent with "healthy" and "wedged again":
|
||||
Simulated first (`-s`): **0 to remove**, so the abort condition never triggered. Then
|
||||
`apt-get install --only-upgrade -y proxmox-backup-server`, **09:51:00→09:51:06Z, exit 0**. Upgraded
|
||||
server/client/docs to 4.2.5-1 plus one new dependency, `proxmox-enterprise-support-keyring 1.1` —
|
||||
**which the 4.2.4-1 changelog had declared**, a small real consistency check between what I read and
|
||||
what apt did.
|
||||
|
||||
```bash
|
||||
ssh root@<ep0> 'PID=$(systemctl show proxmox-backup-proxy -p MainPID --value); \
|
||||
ls /proc/$PID/fd | wc -l; ss -lnt "( sport = :8007 )"'
|
||||
```
|
||||
| check | result |
|
||||
|---|---|
|
||||
| installed | server / client / docs **4.2.5-1** |
|
||||
| daemons | `proxmox-backup-proxy` **active running**, `proxmox-backup` **active running** |
|
||||
| proxy restarted | 542065 → **551655** @ 09:51:04 |
|
||||
| **effective `open files`** | **65536 / 65536** — survived the new package |
|
||||
| drop-ins on disk | both present, unmodified |
|
||||
| `Recv-Q` | **0** |
|
||||
| loopback | **`200` in 12 ms** |
|
||||
| `felhom-pve` over tunnel | **`200` in 0.103 s**, `felhom-pbs active` |
|
||||
| `demo-hp` over tunnel | **`200` in 0.096 s**, `felhom-pbs active` |
|
||||
| hub gauge, post-upgrade | `11:59:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)` |
|
||||
|
||||
Healthy is ~20 fds and `Recv-Q 0`. **A count climbing between restarts means R-336's leak is still
|
||||
live.**
|
||||
`proxmox-backup-manager version` now reads `4.2.5-1 running version: 4.2.5`. **This independently
|
||||
settles the morning's confusion**: that string was never reporting a stale daemon, and now that
|
||||
installed and running genuinely match, both halves agree.
|
||||
|
||||
**One false alarm, mine.** I queried `systemctl is-active proxmox-backup-api` and got `inactive`.
|
||||
**That unit does not exist** — `systemctl cat` returns *"No files found for
|
||||
proxmox-backup-api.service"*. The real pair is `proxmox-backup-proxy.service` ("API Proxy Server")
|
||||
and `proxmox-backup.service` ("API Server"), both active. A bad query, not a fault — recorded because
|
||||
for as long as it took to check, it looked exactly like one.
|
||||
|
||||
**No backup, restore or verify was triggered** to "prove" the endpoint, per the runbook: the reads
|
||||
above answer it without mutating a protected datastore.
|
||||
|
||||
## 7. The after-slope — unchanged, as predicted
|
||||
|
||||
| | window | delta | rate |
|
||||
|---|---|---|---|
|
||||
| before (PID 542065) | 09:18:21Z fd=62 → 09:49:46Z fd=66 | +4 / 1885 s | **183/day** |
|
||||
| after (PID 551655) | 09:51:22Z fd=17 → 10:23:21Z fd=22 | +5 / 1919 s | **225/day** |
|
||||
|
||||
**These are not distinguishable.** The windows differ by **one descriptor**; Poisson uncertainty on
|
||||
n=4 is ±2 and on n=5 is ±2.2, so both are consistent with a single unchanged rate. **The after-figure
|
||||
being numerically higher is noise, not a regression — and certainly not an improvement.** Composition
|
||||
repeats the pattern: **ESTAB 0 → 5, CLOSE-WAIT 0 → 1.**
|
||||
|
||||
**Thirty minutes cannot settle this in either direction, and nothing here claims it does.**
|
||||
|
||||
## 8. Register
|
||||
|
||||
- **R-341 filed, WATCHING** — the two dated checks: **+24 h (2026-08-19 ~10:00Z)** and
|
||||
**+7 d (2026-08-25 ~10:00Z)**, against new `t0` **fd=17 @ 09:51:22Z, PID 551655**, with the exact
|
||||
command and the instruction to **record the ESTAB/CLOSE-WAIT split, not just the total** — the
|
||||
split is what identifies which leak it is. If the PID has changed, the window is void.
|
||||
- **R-336 updated and STILL OPEN.** Its wrong 85/day baseline is corrected to 183–200/day, the
|
||||
mechanism is re-pointed at ESTAB, and the upgrade's null result is recorded. **It stays open on its
|
||||
own merits:** an upgrade that *had* fixed the leak still would not make ~85,000 requests/day to a
|
||||
weekly-write DR endpoint correct.
|
||||
- `documentation/runbooks/offsite-endpoint.md` — **no change made, correctly**: it states no PBS
|
||||
version literal, only package names, so there was nothing to update (and per `docs.md`, version
|
||||
literals do not belong in a current-state doc anyway).
|
||||
|
||||
## 9. Teardown
|
||||
|
||||
**This run provisioned nothing.** No `.deb` was downloaded — `apt-get changelog` served the text
|
||||
directly, so the runbook's `/tmp` cleanup was never needed. The Hetzner snapshot **is retained** as
|
||||
the rollback; deleting it is an operator decision, and it costs €0.018161/GB/month.
|
||||
|
||||
## 10. Observations, not acted on
|
||||
|
||||
- **`pvesm status` reports `felhom-pbs` with Total/Used/Available all `0`** on both boxes while
|
||||
status reads `active`. Consistent before and after the upgrade, and the hub's own gauge reads the
|
||||
real 3.7 GB / 97.9 GB, so nothing is broken — but the PVE-side numbers are not usable as a capacity
|
||||
signal. Likely the namespace-scoped token lacking datastore-level audit. Not filed; noted here.
|
||||
- **8 other packages are held back** (`8 not upgraded`), untouched deliberately — `full-upgrade` on
|
||||
this machine is forbidden by the runbook and would change the kernel and WireGuard alongside the
|
||||
thing under test.
|
||||
- **The morning's `--no-verify` situation is unchanged**: `golden-currency` still convicts on the
|
||||
inherited R-334 (controller 0.216.0 vs golden 0.214.0), untouched by this run.
|
||||
|
||||
Reference in New Issue
Block a user