f5a4fceeeb
gates / gates (push) Successful in 16s
Diagnostic only — no code changed, no version bumped, nothing deployed.
The verdict is a mixture. The unit and the off-site snapshot HOLD the data, proven
by identity in both storage classes including two Hungarian accented filenames. The
loss is in the last leg: ReconstituteFromOffsite skips every isUnit placement and the
volume tars live inside the unit, so the off-site full restore has no named-volume
leg at all — while the local restore-from-unit does, and returned the same tar
byte-identical minutes later.
Twelve rows opened, ceiling R-353 -> R-365. Three HIGH:
R-354 off-site restore never replays volume dumps
R-355 paperless-ngx's Postgres is dumped under a non-existent stack, so its unit
has no DB dump, no safety dump is taken, and the customer is told it has none
R-356 the off-site restore refuses for all 40 no-drive apps saying the running app
"is not installed", with a remedy those apps make impossible
R-353's instruction (2) is satisfied and annotated: the 40-class DOES reach the
off-site tier. Its instruction (1) stands and is now larger. R-329 confirmed still
live and now the only bad-severity emit fleet-wide.
Evidence: documentation/audits/DRILL-backup-truth-2026-08-21/evidence/
191 lines
11 KiB
Markdown
191 lines
11 KiB
Markdown
# REPORT — RUNBOOK ep0: read the PBS changelog, then decide whether to upgrade (2026-08-18, midday)
|
||
|
||
**Outcome:** changelog read → **no connection-handling fix in the range**; operator ruled to upgrade
|
||
anyway **for rehearsal value**; upgraded **4.2.2-1 → 4.2.5-1** cleanly; **the fd slope did not change,
|
||
which is the predicted result.** Two dated checks filed as **R-341**.
|
||
|
||
**Both STOPs cleared by the operator.** No code changed in any repo; `documentation/` only.
|
||
Evidence: `documentation/audits/evidence-ep0-pbs-upgrade-2026-08-18/`.
|
||
|
||
---
|
||
|
||
## 1. Baselines re-confirmed on the machine
|
||
|
||
Not from the audit file — from `dpkg -l` and `apt-cache policy`, per the runbook.
|
||
|
||
| | |
|
||
|---|---|
|
||
| Installed | `proxmox-backup-server` / `-client` **4.2.2-1** |
|
||
| Candidate | **4.2.5-1** (4.2.3-1, 4.2.4-1 also available) |
|
||
| Proxy PID / started | 542065, **03:54:53Z**, unrestarted since the incident |
|
||
| Effective `open files` | **65536 / 65536** — the morning's drop-in in force |
|
||
| `Recv-Q` / loopback | **0** / **`200` in 11 ms** |
|
||
| Datastore | 3.7 G of 98 G, 4% |
|
||
| felhom.eu `main` @ start | `435e044` — matches the runbook's stated baseline |
|
||
|
||
## 2. The before-slope — and a correction I owe the morning's report
|
||
|
||
Two independent windows on the same proxy generation:
|
||
|
||
| window | from → to | delta | rate |
|
||
|---|---|---|---|
|
||
| 31 min | 09:18:21Z fd=62 → 09:49:46Z fd=66 | +4 | **183/day** |
|
||
| 5.64 h | 04:11:36Z fd=19 → 09:49:46Z fd=66 | +47 | **200/day** |
|
||
|
||
**This morning's incident note said ~85/day and "≈2 years of runway". Both were wrong.** They were
|
||
extrapolated from a single 17-minute window whose delta was **one descriptor** — a sample of one
|
||
cannot carry a daily rate, and the agreement with the historical ~73/day that made it feel solid was
|
||
coincidence. **The real rate is ~185–200/day, ~2.6× what I published, and the runway is ~357 days,
|
||
not two years.** Corrected in the incident document and in R-336 rather than left standing.
|
||
|
||
**The mechanism I named was also the minority one.** `CLOSE-WAIT` held flat at **1** across the
|
||
window while `ESTAB` grew **45 → 49** — *all* the growth was established connections. At the wedge
|
||
the split was **1011 ESTAB / 543 CLOSE-WAIT**, so ESTAB dominated there too. **R-336's fix must target
|
||
connections the proxy never reaps, not just `CLOSE-WAIT` sockets.**
|
||
|
||
## 3. The changelog, verbatim — the run's primary deliverable
|
||
|
||
All three entries between 4.2.2-1 and 4.2.5-1 read in full (128 lines), then swept for
|
||
`connection|file descriptor|fd|accept(|close_wait|keep-alive|socket|EMFILE|nofile|leak|proxy|listen|
|
||
backlog|hyper|tokio`.
|
||
|
||
**Exactly one hit, and it is a false positive:**
|
||
|
||
> `* S3: config: allow editing the use-node-config flag that controls whether requests S3 endpoints`
|
||
> ` honor the node's proxy settings or not`
|
||
|
||
HTTP-proxy configuration *for S3 requests* — not the `proxmox-backup-proxy` daemon.
|
||
|
||
What the range does contain. **4.2.5-1** — a security release hardening client-supplied manifests:
|
||
|
||
> `* backup: harden the handling of client supplied backup manifests:`
|
||
> ` - only accept archive names that are plain file names carrying a server side type extension. A`
|
||
> ` crafted name in a manifest could previously make a sync job read or write outside of the`
|
||
> ` snapshot directory, running as the unprivileged 'backup' user.`
|
||
> ` - keep an uploaded manifest in memory and only persist it on backup finish, checking that every`
|
||
> ` archive it lists was really uploaded during that session and that the checksums match`
|
||
|
||
plus a sync/push chunk-reuse fix and a subscription-key architecture check. **4.2.4-1** — S3 rate
|
||
limits, a file-locking user-lookup cache, the new `proxmox-enterprise-support-keyring` dependency,
|
||
docs. **4.2.3-1** — UI/journal work, an LDAP search-filter escape, tape and timezone fixes.
|
||
|
||
**Nothing addresses descriptor lifetime or connection reaping. My recommendation was: do not upgrade
|
||
for this reason.**
|
||
|
||
## 4. STOP 1 — the ruling
|
||
|
||
**Operator (Viktor) ruled: upgrade anyway, for rehearsal value** — *"see how that works for us, we
|
||
need practice with that too"*. Legitimate and recorded as such: this was **a practice run of the
|
||
upgrade procedure on a Tier-2 protected machine, not a fix for the leak**.
|
||
|
||
**The interpretation was fixed in writing before any numbers existed** (`stop1-ruling.txt`):
|
||
unchanged slope = **expected**, not a failed upgrade; changed slope = a **surprise** needing
|
||
explanation, not a confirmation. That file was written at the ruling, not afterwards, so neither
|
||
outcome could be rationalised into a success.
|
||
|
||
## 5. STOP 2 — the snapshot, and what it does not cover
|
||
|
||
Snapshot **421440873** `felhom-hetzner-20260818`, 15.06 GB, status **Available** (complete, not
|
||
merely started), server #147604682, project 15217960.
|
||
|
||
**It covers `/dev/sda` only.** `/mnt/pbs-datastore` is `/dev/sdb`, a separate 100 GB **Volume**, and
|
||
Hetzner server snapshots exclude attached volumes — so this is a rollback for the *software* state
|
||
(packages, unit files, the `LimitNOFILE` drop-ins, nftables, wg) and **not a backup of the backup
|
||
data**. Fine for a package install that writes no datastore content; **it must not be remembered as
|
||
datastore protection.** Taken on a running server, deliberately: powering off ep0 to guard a userspace
|
||
package install would take the only off-premises copy offline.
|
||
|
||
## 6. The upgrade and its verification
|
||
|
||
Simulated first (`-s`): **0 to remove**, so the abort condition never triggered. Then
|
||
`apt-get install --only-upgrade -y proxmox-backup-server`, **09:51:00→09:51:06Z, exit 0**. Upgraded
|
||
server/client/docs to 4.2.5-1 plus one new dependency, `proxmox-enterprise-support-keyring 1.1` —
|
||
**which the 4.2.4-1 changelog had declared**, a small real consistency check between what I read and
|
||
what apt did.
|
||
|
||
| check | result |
|
||
|---|---|
|
||
| installed | server / client / docs **4.2.5-1** |
|
||
| daemons | `proxmox-backup-proxy` **active running**, `proxmox-backup` **active running** |
|
||
| proxy restarted | 542065 → **551655** @ 09:51:04 |
|
||
| **effective `open files`** | **65536 / 65536** — survived the new package |
|
||
| drop-ins on disk | both present, unmodified |
|
||
| `Recv-Q` | **0** |
|
||
| loopback | **`200` in 12 ms** |
|
||
| `felhom-pve` over tunnel | **`200` in 0.103 s**, `felhom-pbs active` |
|
||
| `demo-hp` over tunnel | **`200` in 0.096 s**, `felhom-pbs active` |
|
||
| hub gauge, post-upgrade | `11:59:31 [INFO] PBS-DR box refreshed: 3.7% full (3.7 GB of 97.9 GB)` |
|
||
|
||
`proxmox-backup-manager version` now reads `4.2.5-1 running version: 4.2.5`. **This independently
|
||
settles the morning's confusion**: that string was never reporting a stale daemon, and now that
|
||
installed and running genuinely match, both halves agree.
|
||
|
||
**One false alarm, mine.** I queried `systemctl is-active proxmox-backup-api` and got `inactive`.
|
||
**That unit does not exist** — `systemctl cat` returns *"No files found for
|
||
proxmox-backup-api.service"*. The real pair is `proxmox-backup-proxy.service` ("API Proxy Server")
|
||
and `proxmox-backup.service` ("API Server"), both active. A bad query, not a fault — recorded because
|
||
for as long as it took to check, it looked exactly like one.
|
||
|
||
**No backup, restore or verify was triggered** to "prove" the endpoint, per the runbook: the reads
|
||
above answer it without mutating a protected datastore.
|
||
|
||
## 7. The after-slope — unchanged, as predicted
|
||
|
||
| | window | delta | rate |
|
||
|---|---|---|---|
|
||
| before (PID 542065) | 09:18:21Z fd=62 → 09:49:46Z fd=66 | +4 / 1885 s | **183/day** |
|
||
| after (PID 551655) | 09:51:22Z fd=17 → 10:23:21Z fd=22 | +5 / 1919 s | **225/day** |
|
||
|
||
**These are not distinguishable.** The windows differ by **one descriptor**; Poisson uncertainty on
|
||
n=4 is ±2 and on n=5 is ±2.2, so both are consistent with a single unchanged rate. **The after-figure
|
||
being numerically higher is noise, not a regression — and certainly not an improvement.** Composition
|
||
repeats the pattern: **ESTAB 0 → 5, CLOSE-WAIT 0 → 1.**
|
||
|
||
**Thirty minutes cannot settle this in either direction, and nothing here claims it does.**
|
||
|
||
## 8. Register
|
||
|
||
- **R-341 filed, WATCHING** — the two dated checks: **+24 h (2026-08-19 ~10:00Z)** and
|
||
**+7 d (2026-08-25 ~10:00Z)**, against new `t0` **fd=17 @ 09:51:22Z, PID 551655**, with the exact
|
||
command and the instruction to **record the ESTAB/CLOSE-WAIT split, not just the total** — the
|
||
split is what identifies which leak it is. If the PID has changed, the window is void.
|
||
- **R-336 updated and STILL OPEN.** Its wrong 85/day baseline is corrected to 183–200/day, the
|
||
mechanism is re-pointed at ESTAB, and the upgrade's null result is recorded. **It stays open on its
|
||
own merits:** an upgrade that *had* fixed the leak still would not make ~85,000 requests/day to a
|
||
weekly-write DR endpoint correct.
|
||
- `documentation/runbooks/offsite-endpoint.md` — **no change made, correctly**: it states no PBS
|
||
version literal, only package names, so there was nothing to update (and per `docs.md`, version
|
||
literals do not belong in a current-state doc anyway).
|
||
|
||
## 9. Teardown
|
||
|
||
**This run provisioned nothing.** No `.deb` was downloaded — `apt-get changelog` served the text
|
||
directly, so the runbook's `/tmp` cleanup was never needed. The Hetzner snapshot **is retained** as
|
||
the rollback; deleting it is an operator decision, and it costs €0.018161/GB/month.
|
||
|
||
## 10. Observations, not acted on
|
||
|
||
- **`pvesm status` reports `felhom-pbs` with Total/Used/Available all `0`** on both boxes while
|
||
status reads `active`. Consistent before and after the upgrade, and the hub's own gauge reads the
|
||
real 3.7 GB / 97.9 GB, so nothing is broken — but the PVE-side numbers are not usable as a capacity
|
||
signal. Likely the namespace-scoped token lacking datastore-level audit. Not filed; noted here.
|
||
- **8 other packages are held back** (`8 not upgraded`), untouched deliberately — `full-upgrade` on
|
||
this machine is forbidden by the runbook and would change the kernel and WireGuard alongside the
|
||
thing under test.
|
||
- **The morning's `--no-verify` situation is unchanged**: `golden-currency` still convicts on the
|
||
inherited R-334 (controller 0.216.0 vs golden 0.214.0), untouched by this run.
|
||
|
||
## 11. CI, checked by run ID
|
||
|
||
**Run `351`, `head_sha 3e50902a9`, conclusion `failure`, elapsed 15 s** (10:27:00→10:27:15Z).
|
||
Inherited, not caused: CI's only step is `python3 scripts/repo_gates.py --fast`, which convicts
|
||
`golden-currency` on **R-334** (controller 0.216.0 vs newest golden 0.214.0) — files this run did not
|
||
touch. 15 s is the workflow's own honest-failure band, not the R-265 reap band. The previous run
|
||
`350` on `435e044` failed identically, before this run began.
|
||
|
||
**Expect one `[felhom CI] gates FAILED` mail for run 351.** Same cause as run 348 this morning; not a
|
||
new fault, and not related to the upgrade.
|
||
|
||
`python3 scripts/unproven.py --summary` — **unchanged: 23 walked, 32 not walked of 55.** No number
|
||
moved, correctly: this run proved an operational fact, not a product claim.
|