From c60cd4234cf2036671c695c20aa44dc86ffaa4de Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Mon, 27 Jul 2026 10:53:38 +0200 Subject: [PATCH] =?UTF-8?q?docs(report):=20=C2=A711=20post-session=20live?= =?UTF-8?q?=20state=20=E2=80=94=20offsite=20PBS=20down,=20R-88=20loop=20le?= =?UTF-8?q?ft=20running?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Recorded after the session report was written. The offsite PBS service stopped listening on 8007 five minutes after this session's 14.46 GB restore-test read from it; the box is up and the tunnel is healthy, but no SSH key to it exists so the cause is unestablished — the restore load is a plausible mechanism on a cx23 and is recorded as correlation, not cause. The agent restart then exposed R-88: an unreachable target reads as 'no backup exists', so the offsite tier is perpetually due and the controller runs a full quiesce cycle every ~5 min. Operator ruling: leave it running, it self-heals when PBS returns and masking it would hide the fault. --- REPORT.md | 46 ++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 46 insertions(+) diff --git a/REPORT.md b/REPORT.md index d72dd12..909764f 100644 --- a/REPORT.md +++ b/REPORT.md @@ -233,6 +233,52 @@ An estimate extrapolated from a degraded measurement is not a measurement. --- +## 11. POST-SESSION — live state at hand-off (2026-07-27 ~07:10 UTC) + +Recorded after §1–§10 were written. Two facts, both still true at hand-off. + +### 11.1 The offsite PBS service is DOWN — cause unknown, box reachable + +`felhom-hetzner` (167.233.158.164) reports `status=running` via the Hetzner API; the `wg-felhom` +tunnel is healthy (handshake seconds old, ping 0% loss, ~40 ms); **SSH 22 answers but 8007 refuses**, +repeatedly over several minutes from felhom-pve. The box is up and `proxmox-backup-proxy` is not +listening. **No SSH key to that box exists from DooPlex or felhom-pve**, so diagnosis stopped there — +the box's own journal and `dmesg` are unread. + +**A cause I cannot rule out: this session's own restore-test.** The unattended PBS restore-test read +**14.46 GB** off that datastore 06:44–06:58 UTC and completed `OK`; PBS was refusing five minutes +later. On a **cx23 (2 vCPU / 4 GB)** an OOM of the proxy under that read is a plausible mechanism. +Correlation only — **not established**, and it must not be written up as though it were. First checks +for whoever gets into the box: `journalctl -u proxmox-backup-proxy` and `dmesg | grep -i oom`. +If it IS the restore load, it bears directly on **R-86**: restore-testing a tier weekly means putting +that read on a small offsite box on a schedule. + +### 11.2 An outage loop is RUNNING on demo-felhom, deliberately left running + +The agent restart that applied the reverted 3.5-day cadence exposed **R-88** (filed `eb3f0b8`, +severity corrected `5aca709`): an unreachable target reads as *no backup exists*, so the offsite tier +is perpetually "due". The controller re-polls every ~5 min and runs the **full quiesce cycle** each +time — `quiescing 4 stack(s): [bookstack calibre-web docmost immich]` → `unquiescing (backup failed)` +— roughly **19 s of app downtime per cycle, unbounded**, until PBS answers. + +**Operator ruling 2026-07-27: leave it running.** It is a demo box, the impact is contained, it +self-heals the instant PBS returns, and leaving it keeps the fault visible rather than masked. The +alternatives (disable the tier; ship the R-88 fix) were declined in favour of not masking it. + +**I recorded R-88 as bounded — "one spurious event per restart" — before measuring it.** It is +neither bounded nor event-only; it is a repeating availability fault. The roadmap entry carries the +correction. The error was the same shape as the ~2-hour estimate in §9: a severity asserted from the +mechanism I had reasoned about, before looking at what the mechanism actually did on the box. + +### 11.3 Verified clean at hand-off + +Agent `felhom-agent 0.104.0` active on demo-felhom, `restore-test scheduler starting cadence=84h0m0s`, +both tiers armed (`local` 24h / `felhom-pbs` 168h). The scratch guest leaked by the mid-test restart +was destroyed — **990000 band empty, zero leftover LVs**. `onboot: 0` held on that leaked scratch, +which is the **v0.101.0 fix working in exactly the scenario it was written for**. + +--- + ## 11. Observations - **drill-r50's 403** (`missing privilege VM.Backup at /vms/9100`) is real and was invisible. The box