docs(audits): R-80 DIAG — expected_backup_missed root-caused to in-memory agent backup store

Read-only diagnostic. No fix applied.

Verdict: FALSE ALARM for demo-felhom + demo-hp, true-but-premature for
drill-r50. Fired ONCE (2026-07-26 03:00 UTC), not nightly. No real
external customer notified.

Root cause: felhom-agent/internal/backup/store.go is in-memory by design;
an agent restart empties the host-report 'backups' array until the next
backup runs. The R-50 island migration restarted the fleet on 2026-07-25
12:44 UTC, opening an ~18h blind window that the 03:00 UTC deadline check
fell into. Hub's assessBackupFreshness reads the empty array as 'no
backup exists'.

Separate real finding: the PBS/offsite-DR tier has no schedule at all
(local_backup_target=local, no jobs.cfg) — demo-felhom has one snapshot
from 07-18, demo-hp has zero since pbs_dr was applied 07-21.
This commit is contained in:
Claude Code
2026-07-26 09:59:35 +02:00
parent 9cfa619ec3
commit add5b9bbbb
@@ -0,0 +1,412 @@
# DIAG — `expected_backup_missed` on demo-felhom / demo-hp / drill-r50 (2026-07-26)
**Class:** read-only diagnostic (R-80). Nothing was fixed, triggered, cleared or restarted.
**Author:** Claude Code, DooPlex. **Window investigated:** 2026-05-30 → 2026-07-26 09:55 UTC.
---
## 0. Headline
**The alarm is a FALSE ALARM for 2 of 3 customers, and premature for the third. It did not fire
nightly — it fired exactly once, at 2026-07-26 03:00 UTC.** No real external customer received
anything.
**Root cause (H2, empty-array variant):** the agent's backup record store is **in-memory**
`felhom-agent/internal/backup/store.go:15`, whose own doc-comment says *"In-memory (lost on restart;
the cadence re-populates)"*. An agent restart wipes it, so the host-report emits `backups: []` until
the next backup runs. The hub's `assessBackupFreshness` reads an empty array as *no backup exists*.
**Trigger:** the R-50 island migration restarted `felhom-agent` fleet-wide on 2026-07-25
~12:4112:44 UTC (agent v0.96.0). Backups run on a ~24 h cadence; the next one landed 07-26
07:03 UTC. The 03:00 UTC deadline check fell inside that ~18 h blind window.
**The one genuinely bad finding, unrelated to the alarm text:** the **PBS / offsite-DR tier has no
schedule at all**. demo-felhom's `felhom-offsite` datastore holds exactly one snapshot
(2026-07-18); demo-hp has `pbs_dr` applied since 07-21 and **zero snapshots ever**. See §7.
The premise that ~7.3 days of staleness predates the 07-25 channel outage is **correct but
misleading**: the 176 h figure is the age of the *PBS* snapshot, reached only as a fallback once the
vzdump array went empty. Local vzdump was never 7 days stale — it ran on 07-24, 07-25 and 07-26.
---
## 1. D0 — access inventory
| Target | Reachable | Path | User | Passwordless sudo | Notes |
|---|---|---|---|---|---|
| **DooPlex** (this host) | ✅ | local | `kisfenyo` | ✅ | hosts hub, hub DB, k3s, local PBS |
| **Hub** (k3s `felhom-system`) | ✅ | `sudo kubectl` | — | ✅ | pod `hub-9496fd7c9-kfpkx`, image `felhom-hub:0.74.0` |
| **Hub DB** | ✅ | `kubectl cp` → scratchpad, `sqlite3 -readonly` | — | ✅ | 113 MB copy; **live DB never opened writable** |
| **PVE `demo-felhom`** (N100) | ✅ | `ssh felhom-pve` (Tailscale 100.70.170.35) | `root` | n/a (is root) | full journal + `pvesm` |
| **PVE `demo-hp`** (t740) | ❌ | `ssh demo-hp``Permission denied (publickey,password)` | — | — | no baked key. Break-glass root exists in hub `host_recovery`**not used**, per scope ("do not escalate or hunt for credentials") |
| **PVE `drill-r50`** (host `drill-r50-0a4f9a`) | ❌ | no SSH alias configured | — | — | NOT COLLECTED |
| **PBS `felhom-offsite`** (ep0/Hetzner, via wg `10.77.0.1`) | ⚠️ indirect | read through PVE storage client + agent verify loop | — | — | no direct PBS API/root; free space + prune/GC history NOT COLLECTED |
| **Guest 9201 (demo-felhom)** | ✅ | `ssh felhom-pve "pct exec 9201 …"` | root | n/a | not needed for this diagnosis |
| **DooPlex-local PBS `felhom-spike`** | ✅ | local, `sudo` | — | ✅ | leftover spike store, newest snapshot 2026-07-10 — **not** the live offsite datastore |
**Exit-code discipline.** stderr was inspected on every rc=0. `pvesm status` / `pvesm list` first
emitted perl locale warnings to stderr that truncated the read — re-run with `LC_ALL=C` and stderr
captured to a separate file; all three commands rc=0 with **empty** stderr on the clean run.
`proxmox-backup-manager namespace list` returned rc≠0 (subcommand absent on this PBS version) and
`proxmox-backup-client list` failed with *"no password input mechanism available"* — neither was
retried, because `pvesm list felhom-pbs` yields the same ground truth via the configured credential.
---
## 2. Who the three customers are
| customer_id | Name | Status | DR tier | Customer notify channel | Real external customer? |
|---|---|---|---|---|---|
| `demo-felhom` | Demo Ügyfél | active | 1 | ✅ `doo…@freemail.hu` (operator's own mailbox) | **No** — operator demo box (N100) |
| `demo-hp` | Demo HP | active | 1 | ❌ none — operator channel only | **No** — operator demo box (HP t740) |
| `drill-r50` | drill-r50 | active | 0 | ❌ none — operator channel only | **No** — R-50 drill VM, bound 2026-07-25 09:49 UTC, ~19 h old |
**No real external customer was notified.** The one real external customer in the hub,
**`peti-felhom` ("Peti Proxmox", active, dr_tier 0)**, did **not** fire today and has **no**
`customer_notifications` row (operator channel only, always). It is silent because it has no agent
host-report at all → `deadline.go:176` takes the explicit *"legacy/defunct controller-only
customer"* branch, which deliberately raises no backup alarm.
`customer_notifications` holds exactly **one** real row (`demo-felhom`) plus a
`_resend-rotation-test` stub. So at most one mailbox — the operator's own — can ever receive this
event today.
---
## 3. D1 — the events, verbatim
### 3.1 Every `expected_backup_missed` in the current era (since the 2026-06-16 repoint)
| id | customer | created_at (UTC) | sev | verbatim message |
|---|---|---|---|---|
| 723 | demo-felhom | 2026-06-29 03:00:00 | error | `No fresh verified backup: no PBS snapshot or successful backup in the latest host-report` |
| 1129 | peti-felhom | 2026-07-13 03:00:00 | error | `No fresh verified backup: newest backup is 40h0m0s old (limit 26h0m0s)` |
| 1203 | peti-felhom | 2026-07-14 03:00:00 | error | `No fresh verified backup: no PBS snapshot or successful backup in the latest host-report` |
| 1275 | peti-felhom | 2026-07-15 03:00:00 | error | `No fresh verified backup: no PBS snapshot or successful backup in the latest host-report` |
| **1803** | **demo-felhom** | **2026-07-26 03:00:00** | error | `No fresh verified backup: newest backup is 176h0m0s old (limit 26h0m0s)` |
| **1804** | **demo-hp** | **2026-07-26 03:00:01** | error | `No fresh verified backup: no PBS snapshot or successful backup in the latest host-report` |
| **1805** | **drill-r50** | **2026-07-26 03:00:01** | error | `No fresh verified backup: no PBS snapshot or successful backup in the latest host-report` |
(Rows 339512, 2026-05-30 → 2026-06-16, are the **pre-repoint** era with the old message
`No backup completed or failed since midnight` — the v0.12.0 CHANGELOG names that as a known daily
false alarm, already fixed. Not relevant.)
### 3.2 Per-customer answers
| | demo-felhom | demo-hp | drill-r50 |
|---|---|---|---|
| **Reason branch fired** | `newest backup is Xh old` | `no PBS snapshot or successful backup` | `no PBS snapshot or successful backup` |
| **First occurrence** | 2026-07-26 03:00:00 UTC | 2026-07-26 03:00:01 UTC | 2026-07-26 03:00:01 UTC |
| **Total count** | 1 (this era: 2, incl. 06-29) | 1 | 1 |
| **Nightly?** | **No — single occurrence** | **No — single occurrence** | **No — single occurrence** |
| **Channels delivered** | operator **and customer** | operator only | operator only |
| **H2 status** | refuted by string alone → H1/H4/H5 | live → D2b decisive | live → D2b decisive |
**The "nightly" premise is refuted.** `notification_log` records **exactly one** customer-channel
send for this event since the 2026-06-16 repoint: id 245, `demo-felhom`, 2026-07-26 03:00:01,
status `sent`, Hungarian subject text **„A mai biztonsági mentés nem készült el a határidőig!"**
(`hub/internal/notify/templates.go:106`), body carrying the English reason string
`No fresh verified backup: newest backup is 176h0m0s old (limit 26h0m0s)`.
### 3.3 Co-firing signals in the same window
`expected_dbdump_missed`: **not fired** for any of the three (all three emitted `db_dump_completed`
at 00:30 UTC today). `host_staleness` / `node_stale` / `offsite_stale` / `pbsdr_*`: **none** on
2026-07-26. The only other events today are `agent_channel_unreachable` (demo-felhom 06:16, demo-hp
06:19, drill-r50 07:23 — transient) and `local_api_endpoint_drift` on drill-r50 (07:21, unrelated,
see §10).
---
## 4. D2 — what the agent reported
Latest host-report per customer (2026-07-26 07:4207:44 UTC), `pbs_snapshots` and `backups` verbatim:
**demo-felhom**`pbs_snapshots` 1 entry, `backups` 1 entry:
```json
"pbs_snapshots": [{ "namespace":"root","backup_type":"ct","backup_id":"9201",
"backup_time":"2026-07-18T18:31:06Z","size_bytes":9744319415,
"owner":"felhom@pbs!demo-felhom","protected":false,"encrypted":true,
"verify_state":"ok",
"verify_upid":"UPID:felhom-hetzner:000003D4:000005BE:00000048:6A65ACDA:verify:felhom\\x2doffsite\\x3ans-demo\\x2dfelhom:felhom@pbs!demo-felhom:" }]
"backups": [{ "target_id":"local","vmid":9201,
"archive":"local:backup/vzdump-lxc-9201-2026_07_26-09_03_19.tar.zst",
"mode":"snapshot","crash_consistent":true,"size_bytes":5690147837,
"success":true,"started_at":"2026-07-26T07:03:19Z","duration_seconds":216.556501101,
"uncovered_volumes":["/etc/felhom-bootstrap","/mnt/felhom-drives"] }]
```
**demo-hp**`pbs_snapshots: []` (0 entries), `backups` 1 entry:
```json
"backups": [{ "target_id":"local","vmid":9201,
"archive":"local:backup/vzdump-lxc-9201-2026_07_26-09_06_02.tar.zst",
"mode":"snapshot","crash_consistent":true,"size_bytes":1476771458,
"success":true,"started_at":"2026-07-26T07:06:02Z","duration_seconds":76.320400309,
"uncovered_volumes":["/etc/felhom-bootstrap","/mnt/felhom-drives"] }]
```
**drill-r50**`pbs_snapshots: []`, `backups: []`. Both empty. No `pbs_dr` stanza (`null`).
> Note the **timing trap**: these are the *07:44* reports. At **03:00**, when the check ran, all
> three `backups` arrays were **empty** — see §5.
### 4.1 D2b — the parse check (H2 parse variant)
Agent wire struct: `felhom-agent/internal/hub/report.go:27,29``Backups []Backup`,
`PBSSnapshots []PBSSnapshot`. Emitted values are `2026-07-26T07:03:19Z` / `2026-07-18T18:31:06Z`
plain RFC3339 with a `Z` offset. `parseBackupTime` (`deadline.go:107`) accepts exactly
`time.RFC3339`. **No format mismatch.**
**H2's parse variant is REFUTED. H2's empty-array variant is the root cause.** The two are conflated
behind one reason string — see fix #3.
### 4.2 The mechanism, proven
`felhom-agent/internal/backup/store.go:10-19`:
> *"Store holds the agent's LATEST backup result per target … **In-memory (lost on restart; the
> cadence re-populates)** and mutex-guarded…"*
`Collector.collectBackups` (`internal/hub/collect.go:393`) returns `[]Backup{}` when the store is
empty. Hub-side, `assessBackupFreshness` sees `!havePBS && !haveVzdump` → fires *"no PBS snapshot or
successful backup"*; or, if a stale PBS snapshot survives (demo-felhom), falls through to the age
branch and reports **the PBS snapshot's** age — 176 h.
**The alarm cannot distinguish "no backup was taken" from "the agent forgot".**
---
## 5. D4 — the discriminator: agent-reported vs ground truth
Ground truth gathered read-only on `felhom-pve` (`pvesm list`, rc=0, empty stderr — **no verify, no
prune, no GC, no backup triggered**):
```
PBS felhom-pbs (datastore felhom-offsite, ns demo-felhom, server 10.77.0.1):
felhom-pbs:backup/ct/9201/2026-07-18T18:31:06Z pbs-ct 9744319415 9201 <- ONLY entry
LOCAL local (content=backup):
local:backup/vzdump-lxc-9100-2026_07_18-18_30_52.tar.zst 612397450 9100
local:backup/vzdump-lxc-9100-2026_07_20-17_50_57.tar.zst 648549366 9100
local:backup/vzdump-lxc-9201-2026_07_24-21_12_03.tar.zst 5487897588 9201
local:backup/vzdump-lxc-9201-2026_07_25-08_30_14.tar.zst 5603727891 9201
local:backup/vzdump-lxc-9201-2026_07_26-09_03_19.tar.zst 5690147837 9201
```
### 5.1 What the check actually saw, night by night (demo-felhom)
The report nearest each 03:00 UTC — this is the decisive table:
| Check at | Report used | `backups` newest vzdump | `pbs_snapshots` newest | Age of newest | Fired? |
|---|---|---|---|---|---|
| 07-19 03:00 | 07-19 02:45 | 2026-07-18T18:20:11Z | none | 8.7 h | no ✅ correct |
| 07-20 03:00 | 07-19 11:30 | 2026-07-18T18:20:11Z | none | 32.7 h | no ⚠️ *(see note)* |
| 07-21 03:00 | 07-21 02:55 | 2026-07-20T06:40:45Z | none | 20.3 h | no ✅ correct |
| 07-22 03:00 | 07-22 02:47 | 2026-07-21T11:18:24Z | 2026-07-18T18:31:06Z | 15.7 h | no ✅ correct |
| 07-23 03:00 | 07-23 02:59 | 2026-07-22T17:21:46Z | 2026-07-18T18:31:06Z | 9.6 h | no ✅ correct |
| 07-24 03:00 | 07-24 02:55 | 2026-07-23T11:41:34Z | 2026-07-18T18:31:06Z | 15.3 h | no ✅ correct |
| 07-25 03:00 | 07-25 02:57 | 2026-07-24T19:12:03Z | 2026-07-18T18:31:06Z | 7.8 h | no ✅ correct |
| **07-26 03:00** | 07-26 02:59 | **NONE (array empty)** | 2026-07-18T18:31:06Z | **176 h** | **YES** ❌ false |
> *07-20 note:* the customer was `down` that morning (`node_down` 07-19 12:38) and would have been
> skipped by the `staleness.GetState(id) == "down"` guard at `deadline.go:160`. Either way no alarm
> was raised, and a real backup had run 07-20 06:40.
**The deadline check ran every night and was correctly silent on all of them.** The 07-26 fire is
caused solely by the array going empty. This also refutes any "the scheduler stopped" theory —
`scheduleDaily` (`cmd/hub/main.go:677`) recomputes the next 05:00 Budapest on every loop iteration
and is sound.
### 5.2 The array-emptying transitions (all three customers)
Only the change-points are shown:
```
demo-felhom 2026-07-24 19:11:58 nb=0 <- agent restart, refilled 30 min later
2026-07-25 06:23:55 nb=0 <- agent restart, refilled 15 min later
2026-07-25 12:44:05 nb=0 <- R-50 restart; STAYS EMPTY across 03:00
2026-07-26 07:12:xx nb=1 vzdump=2026-07-26T07:03:19Z <- self-cleared
demo-hp 2026-07-24 18:46:13 nb=1 vzdump=2026-07-24T18:38:30Z
2026-07-25 10:22:45 nb=0 <- agent restart, refilled 15 min later
2026-07-25 12:41:28 nb=0 <- R-50 restart; STAYS EMPTY across 03:00
2026-07-26 07:12:20 nb=1 vzdump=2026-07-26T07:06:02Z <- self-cleared
drill-r50 2026-07-25 09:50:55 nb=0 npbs=0 <- never had a backup; still nb=0
```
The two earlier restarts (07-24 19:11, 07-25 06:23) also emptied the array, but a backup followed
within 30 minutes, so no 03:00 boundary was crossed. **The R-50 restart is the one that landed in
the blind window.**
### 5.3 Verdict per customer
| Customer | Agent said | PBS/local ground truth | **Verdict** |
|---|---|---|---|
| **demo-felhom** | `backups: []` at 03:00; PBS 07-18 | vzdump ran **07-24 21:12, 07-25 08:30, 07-26 09:03 CEST** (3 archives retained, retention=3); PBS **1 snapshot, verify_state `ok`** | **H2 — FALSE ALARM.** Local tier fully healthy. PBS tier genuinely stale (§7), but that is a *different* fault than the message implies |
| **demo-hp** | `backups: []`, `pbs_snapshots: []` at 03:00 | vzdump **07-25 10:23, 07-26 07:06 UTC** both `success:true`; PBS **0 snapshots ever** | **H2 — FALSE ALARM.** Local tier healthy. PBS tier never ran (§7) |
| **drill-r50** | both empty | nothing, ever | **H1 — TRUE but premature.** Box bound 2026-07-25 09:49 UTC, ~17 h before the check; no backup was yet due |
---
## 6. Verdicts on H1H6
| | Verdict | The observation that decided it |
|---|---|---|
| **H1** genuine backup outage | **REFUTED for the host/vzdump tier; SUPPORTED for the PBS/offsite-DR tier; SUPPORTED-but-premature for drill-r50** | `pvesm list local` shows three retained vzdump archives (07-24, 07-25, 07-26) and the agent journal logs `backup: completed … target=local` at 09:06:56 CEST today. Conversely `pvesm list felhom-pbs` returns exactly one snapshot, from 07-18 |
| **H2** reporting/parse break | **SUPPORTED (empty-array variant); REFUTED (parse variant)** | `store.go:15` is in-memory by design; §5.2 shows the array emptying at each agent restart and refilling at the next backup. Timestamps are valid RFC3339 `Z`, so `parseBackupTime` never silently drops one |
| **H3** verification failure | **REFUTED** | demo-felhom's only snapshot carries `verify_state: "ok"`; the agent journal logs `pbs: verify cycle complete datastore=felhom-offsite snapshots=1` at 02:44 and 08:44 CEST today. demo-hp/drill-r50 have no snapshots to verify |
| **H4** transport | **REFUTED** | `wg show` on felhom-pve: peer `167.233.158.164:443`, **latest handshake 39 seconds ago**, 753.13 MiB rx / 481.85 MiB tx, keepalive 25 s. All three reports carry `wireguard.active: true` with `last_handshake_age_s` 3646. `pvesm status` shows `felhom-pbs` **active** |
| **H5** schedule stopped | **PARTIALLY SUPPORTED** | The vzdump cadence is alive but its 24 h phase **resets on agent restart** (`BackupCadence()` default 24 h, `config.go:390`), which opened the window. Separately, **no PBS schedule has ever existed**: `backup.local_backup_target = "local"` and no `/etc/pve/jobs.cfg` (§7) |
| **H6** stale host-report | **REFUTED** | Latest reports 2026-07-26 **07:42:19 / 07:44:41 / 07:44:53** UTC, all agent 0.96.0; 664 / 436 / 96 rows. No `host_staleness` fired |
---
## 7. The real finding — the PBS / offsite-DR tier has no schedule
This is **not** what the alarm text described, and it is the one thing here that is genuinely wrong.
- `felhom-agent`'s backup runner has exactly **one** target, resolved by
`BackupConfig.BackupTarget()` (`internal/config/config.go:415`): `local_backup_target` if set,
else the constant `defaultBackupTarget = "felhom-pbs"`.
- On demo-felhom, `/etc/felhom-agent/agent.json` sets `"local_backup_target": "local"`. **Therefore
no agent-driven backup on that host ever targets `felhom-pbs`.**
- There is **no `/etc/pve/jobs.cfg`** on the host — no PVE-native backup job either.
- The single PBS snapshot (2026-07-18T18:31:06Z) predates that and was not produced by a recurring
job.
- demo-hp reports `pbs_dr: {state:"applied", storage_id:"felhom-pbs", namespace:"demo-hp",
applied_at:"2026-07-21T16:35:44Z"}` and **`pbs_snapshots: []`** — the DR tier has been "applied"
for 5 days and has produced nothing.
- Today's vzdump was **controller-initiated**, not cadence-initiated:
`local-api: backup reached snapshotted (app may resume) vmid=9201 job=backup-9201-1785049399941273219`
→ `backup: completed … target=local`. The agent exposes `GET /backup/due` with a 24 h cadence
(`localapi/server.go:92-95`) and the in-guest controller drives it.
**So the offsite-DR promise is currently unbacked on both demo boxes.** Worth its own task; it is
not fixable as a one-liner.
---
## 8. D7 — timeline
| When (UTC) | What |
|---|---|
| 2026-06-16 | hub **v0.12.0** repoints `CheckBackupDeadlines` from `backup_completed` events to host-report freshness. **This is where the empty-array sensitivity is introduced** (`assessBackupFreshness`, commit `0635640`) |
| 2026-07-18 18:20 | last vzdump before the PBS snapshot |
| **2026-07-18 18:31:06** | **the one and only PBS snapshot on demo-felhom.** Nothing has been written to `felhom-offsite` since |
| 2026-07-21 08:39 / 16:35 | `pbs_dr` applied — demo-felhom / demo-hp |
| 2026-07-24 19:11 | agent restart → `backups` empties, refills 19:42 (no 03:00 crossed) |
| 2026-07-25 06:23 | agent restart → `backups` empties, refills 06:38 (no 03:00 crossed) |
| 2026-07-25 09:49 | **drill-r50** appliance bound — a brand-new customer with no backup history |
| **2026-07-25 12:41 / 12:44** | **R-50 island-bridge fleet migration — `felhom-agent` v0.96.0 restarted on every box.** `systemctl show felhom-agent` on felhom-pve: `ExecMainStartTimestamp=Sat 2026-07-25 14:44:37 CEST`, `NRestarts=0`. `backups` empties **and stays empty** |
| 2026-07-26 00:30 / 01:30 | `db_dump_completed` + `crossdrive_completed` — app-data tiers healthy |
| **2026-07-26 03:00:0003:00:01** | **deadline check fires on all three.** demo-felhom also mails the customer channel |
| 2026-07-26 07:03 / 07:06 | next vzdumps run → arrays repopulate → **condition self-clears** |
| 2026-07-26 07:15:58 | hub pod restarts onto **v0.74.0** — *after* the alarm, not causal |
**What changed at the outage start: the R-50 island migration's fleet-wide agent restart on
2026-07-25 12:44 UTC.** That is the named, correlated change. Nothing else in the window touches
backups.
Host uptime is 3 d 14 h (boot ~2026-07-22 19:14 UTC) — **no host reboot** in the window; the agent
restart was a service restart. hub v0.73.2 was, as stated in the task, a `hostInstallVersion` string
bump never deployed as an artifact (`manifests/hub.yaml` now pins 0.74.0) — **confirmed irrelevant**.
`deadline.go` has not been touched since 2026-06-16, so neither 0.73.x nor 0.74.0 changed this logic.
---
## 9. D8 — is the customer-facing message accurate?
**No, on two independent counts.**
The Hungarian template (`hub/internal/notify/templates.go:106`) is
**„A mai biztonsági mentés nem készült el a határidőig!"** — *"Today's backup was not completed by
the deadline!"*. A customer reads that as *all my backups failed*.
At the moment it was sent to demo-felhom (2026-07-26 03:00:01):
| Tier | State at send time | Healthy? |
|---|---|---|
| Local vzdump (host tier) | `vzdump-lxc-9201-2026_07_25-08_30_14.tar.zst` on disk, 5.6 GB, 20 h old | ✅ |
| App-data cross-drive (tier 2) | `crossdrive_completed` 2026-07-26 01:30:06 — 1.5 h before the mail | ✅ |
| DB dump | `db_dump_completed` 2026-07-26 00:30:56 — 2.5 h before the mail | ✅ |
| PBS / offsite-DR | last snapshot 2026-07-18, `verify_state: ok` | ❌ genuinely stale |
1. **The trigger was an artifact** — the local host-tier backup had run and was on disk; the agent
had merely forgotten it across a restart.
2. **Even when the signal is true, the scope is wrong** — this check covers only the PBS/vzdump host
tier. The controller's app-data tiers are separate signals (`backup_failed`, `db_dump_failed`,
`crossdrive_*`) and were all green.
Recorded as a finding. Not changed.
---
## 10. Proposed fixes — described, NOT applied
| # | Where | Fix | Class |
|---|---|---|---|
| **1** | `felhom-agent/internal/backup/store.go` (`Store`, `RecordBackup`, `Backups`) | Stop losing backup history on restart. Either **persist** the latest-per-target record to the agent's state dir, or have `Collector.collectBackups` **fall back to querying PVE** for the newest archive on the target storage (`pvesm list <target> --content backup` / the API equivalent) when the store is cold. The query path is arguably better — it is ground truth rather than remembered state, and it would also have surfaced the real PBS gap | **needs-a-spike** (persistence vs query-on-demand is a design call) |
| **2** | `felhom.eu/hub/internal/monitor/deadline.go` → `assessBackupFreshness` | **Do not alarm on absence of evidence.** The function's own doc says *"Only POSITIVE evidence of a problem fires an alarm"*, but `!havePBS && !haveVzdump` is precisely *absence* of evidence. Narrow it: empty+empty → **unknown**, log + operator-only, never the customer channel. Optionally require the condition to persist across ≥2 consecutive daily checks before escalating. **This is the same bug shape as the hub v0.73.0 `offsite_stale` fix** (never-ran branch with no time anchor) | **safe-obvious** for the operator-only narrowing; **needs-a-spike** for a full unknown-state model |
| **3** | same file, same function (`deadline.go:90`) | **Split the conflated reason string.** `"no PBS snapshot or successful backup in the latest host-report"` currently covers both *arrays empty* and *every timestamp unparseable*. Emit distinct strings so the next diagnosis is one query instead of a session | **safe-obvious** |
| **4** | `felhom.eu/hub/internal/notify/templates.go:106` | **Narrow the customer copy.** Name the host/DR tier explicitly and state that app data is covered by separate alerts, so the message stops implying total backup failure | **safe-obvious** — run `scripts/hub_confirm_gate.py` + the `felhom-ui-design` copy rules |
| **5** | agent config on both demo hosts + `internal/backup/runner.go` | **Wire an actual PBS/offsite-DR schedule** (§7). Today `local_backup_target: "local"` means nothing ever targets `felhom-pbs`, and the runner supports only one target per run. Decide: second target per run, alternating cadence, or a PVE backup job | **needs-a-spike** — this is the real defect and the largest piece of work |
| **6** | `customer_notifications` for `demo-felhom` | **RECOMMENDATION — operator decision, deliberately NOT applied.** `enabled_events` currently includes `expected_backup_missed`. Given that the check demonstrably false-alarms on any agent restart and the copy overstates scope, consider **removing `expected_backup_missed` from that list until fixes 14 land**. The operator channel keeps full visibility. Mitigating factor: the only subscribed mailbox is the operator's own, so the blast radius today is zero — this can wait for the proper fix | operator's call |
**Suppression urgency: LOW.** The condition self-cleared at 07:03 UTC, no real external customer is
subscribed, and the alarm requires an agent restart in the backup blind window to recur.
---
## 11. Not collectable
| Item | Why | What access would be needed |
|---|---|---|
| demo-hp host-level: agent journal, `wg show`, `agent.json` backup stanza, PBS attempt logs | `ssh demo-hp` → `Permission denied (publickey,password)`; break-glass root exists in hub `host_recovery/demo-hp-bb76ea` but was **not used** per scope | a baked SSH key, or explicit authorisation to use the break-glass credential |
| drill-r50 host-level (host `drill-r50-0a4f9a`) | no SSH alias configured; host not identified beyond the report | an SSH route to that host |
| PBS server side: `felhom-offsite` free space, prune/GC job history, per-namespace usage | no direct API/root on the ep0/Hetzner PBS; local `proxmox-backup-client` has no password input mechanism | PBS API token or root on the PBS box |
| Hub logs from 03:00 today | the hub pod restarted at 07:15:58 UTC onto v0.74.0; `kubectl logs` only reaches back to then, and there is no previous-container log | a log sink outside the pod |
| Whether a datastore-full condition contributed | `pvesm status` reports `felhom-pbs` Total/Used/Available all **0 KiB** while Status=`active` — see §12 | PBS-side status |
---
## 12. Observations — noticed, not acted on
1. **`pvesm status` reports `felhom-pbs` with Total = Used = Available = 0 KiB**, yet Status `active`
and `pvesm list` works. Either the PBS storage plugin isn't returning usage over the wg link, or
the datastore genuinely reports nothing. Worth a look when §7 is picked up — a full datastore is a
classic silent-write-refusal cause, and this reading cannot currently rule it out.
2. **`customer_notifications.updated_at` for `demo-felhom` = 2026-07-26 07:17:54** — that row was
touched **today, after the alarm**. Flagging in case the operator already adjusted something;
I did not change it and did not investigate further.
3. **drill-r50 is emitting `local_api_endpoint_drift`** (07:21 today):
`controller.yaml=192.168.0.176:8443 bootstrap.json=169.254.253.1:8443 (pin agrees: true)`, plus
`agent_channel_unreachable`. Unrelated to this diagnosis, but live and probably R-50 fallout.
4. **`agent_channel_unreachable` fired on demo-felhom (06:16) and demo-hp (06:19) today** and
recovered. The task's framing treated the 07-25 channel outage as a possible cause; it is not —
§5.1 shows backups ran normally throughout it. Coincidence confirmed.
5. **The DooPlex-local PBS datastore `felhom-spike`** (`/mnt/5_hdd/backup/demo-felhom`, the only
entry in `datastore.cfg`) is a leftover spike store with snapshots up to 2026-07-10. It is **not**
the live offsite datastore — that is `felhom-offsite` on `felhom-hetzner`, reached over wg at
`10.77.0.1`. Easy to confuse when hunting PBS ground truth on this host.
6. **`restore_test_cadence_seconds: 0`** in the agent config → `RestoreTestCadence()` maps 0 to the
24 h default, so restore tests should be running; `restore_tests` was not examined in this pass.
7. The pre-repoint era (2026-05-30 → 06-16) shows this same event false-alarming **daily** on
peti-felhom and demo-felhom for a different reason, fixed in v0.12.0. Combined with the v0.73.0
`offsite_stale` fix, **this is the third false-alarm in the same monitor family** — all three from
treating "no signal" as "bad signal". Fix #2 addresses the class, not just the instance.
---
## 13. What was NOT done (prime directive compliance)
No backup was run, triggered or tested. No PBS verify, prune or GC. No agent, hub, PBS or
`wg-felhom` restart. No alert cleared, acknowledged or muted. No notification preference or channel
setting changed. No schedule, timer or config file edited. No code changed, no version bumped.
The hub DB was **copied out** (`kubectl cp`) and every query ran against the **copy** with
`sqlite3 -readonly`; the live `/data/hub.db` was never opened writable. All host commands were
listings (`pvesm list`, `pvesm status`, `wg show`, `systemctl show`, `journalctl`).
The staleness measurement is intact: as of this writing the newest `felhom-offsite` snapshot is still
**2026-07-18T18:31:06Z**.