docs(report): §11 post-session live state — offsite PBS down, R-88 loop left running

Recorded after the session report was written. The offsite PBS service stopped
listening on 8007 five minutes after this session's 14.46 GB restore-test read
from it; the box is up and the tunnel is healthy, but no SSH key to it exists so
the cause is unestablished — the restore load is a plausible mechanism on a cx23
and is recorded as correlation, not cause.

The agent restart then exposed R-88: an unreachable target reads as 'no backup
exists', so the offsite tier is perpetually due and the controller runs a full
quiesce cycle every ~5 min. Operator ruling: leave it running, it self-heals when
PBS returns and masking it would hide the fault.
This commit is contained in:
2026-07-27 10:53:38 +02:00
parent 2b24c70536
commit c60cd4234c
+46
View File
@@ -233,6 +233,52 @@ An estimate extrapolated from a degraded measurement is not a measurement.
---
## 11. POST-SESSION — live state at hand-off (2026-07-27 ~07:10 UTC)
Recorded after §1–§10 were written. Two facts, both still true at hand-off.
### 11.1 The offsite PBS service is DOWN — cause unknown, box reachable
`felhom-hetzner` (167.233.158.164) reports `status=running` via the Hetzner API; the `wg-felhom`
tunnel is healthy (handshake seconds old, ping 0% loss, ~40 ms); **SSH 22 answers but 8007 refuses**,
repeatedly over several minutes from felhom-pve. The box is up and `proxmox-backup-proxy` is not
listening. **No SSH key to that box exists from DooPlex or felhom-pve**, so diagnosis stopped there —
the box's own journal and `dmesg` are unread.
**A cause I cannot rule out: this session's own restore-test.** The unattended PBS restore-test read
**14.46 GB** off that datastore 06:4406:58 UTC and completed `OK`; PBS was refusing five minutes
later. On a **cx23 (2 vCPU / 4 GB)** an OOM of the proxy under that read is a plausible mechanism.
Correlation only — **not established**, and it must not be written up as though it were. First checks
for whoever gets into the box: `journalctl -u proxmox-backup-proxy` and `dmesg | grep -i oom`.
If it IS the restore load, it bears directly on **R-86**: restore-testing a tier weekly means putting
that read on a small offsite box on a schedule.
### 11.2 An outage loop is RUNNING on demo-felhom, deliberately left running
The agent restart that applied the reverted 3.5-day cadence exposed **R-88** (filed `eb3f0b8`,
severity corrected `5aca709`): an unreachable target reads as *no backup exists*, so the offsite tier
is perpetually "due". The controller re-polls every ~5 min and runs the **full quiesce cycle** each
time — `quiescing 4 stack(s): [bookstack calibre-web docmost immich]``unquiescing (backup failed)`
— roughly **19 s of app downtime per cycle, unbounded**, until PBS answers.
**Operator ruling 2026-07-27: leave it running.** It is a demo box, the impact is contained, it
self-heals the instant PBS returns, and leaving it keeps the fault visible rather than masked. The
alternatives (disable the tier; ship the R-88 fix) were declined in favour of not masking it.
**I recorded R-88 as bounded — "one spurious event per restart" — before measuring it.** It is
neither bounded nor event-only; it is a repeating availability fault. The roadmap entry carries the
correction. The error was the same shape as the ~2-hour estimate in §9: a severity asserted from the
mechanism I had reasoned about, before looking at what the mechanism actually did on the box.
### 11.3 Verified clean at hand-off
Agent `felhom-agent 0.104.0` active on demo-felhom, `restore-test scheduler starting cadence=84h0m0s`,
both tiers armed (`local` 24h / `felhom-pbs` 168h). The scratch guest leaked by the mid-test restart
was destroyed — **990000 band empty, zero leftover LVs**. `onboot: 0` held on that leaked scratch,
which is the **v0.101.0 fix working in exactly the scenario it was written for**.
---
## 11. Observations
- **drill-r50's 403** (`missing privilege VM.Backup at /vms/9100`) is real and was invisible. The box