# REPORT — F-CRIT-2: a failed backup must not look like a fresh one (v0.106.0) (2026-07-28) **Overwritten** per the standing rule. Agent **v0.105.0 → v0.106.0**. Campaign 8 finding F-CRIT-2, fixed and **re-tested live by replaying the original injection**. Scope: `felhom-agent` only. No wire/contract change; the controller needed no change. ## The defect `NewestArchiveTime` (`internal/backup/runner.go`) selected the newest storage entry with `e.Content == "backup" && e.VMID == vmid` and **no completeness test**. A PBS daemon killed mid-upload leaves an aborted snapshot that PBS publishes into the same listing: manifest-less, **1 byte, and NEWEST**. Counting it made the tier report freshly backed up, so it went **not due** and was never retried — seven days of silence on the real 168h cadence, with the R-88 breaker (defers only *due* tiers) and the hub deadline monitor (reads the same freshness) both blind to it. R-84's *"ask the storage, it is ground truth"* was right. The bug is that **presence was taken for validity** — and a phantom is a more convincing lie than an empty array, because absence at least reads as absence. ## Phase 0 — the discriminator survey The campaign's own phantom was gone (Phase D deleted it deliberately to stop the suppression), so one was recreated on demo-hp by replaying fault 2. Raw PVE API, captured live: ``` PHANTOM: {"content":"backup","ctime":1785216674,"format":"pbs-ct","size":1,"subtype":"lxc", "vmid":9201,"volid":"felhom-pbs:backup/ct/9201/2026-07-28T05:31:14Z"} GOOD: {... "encrypted":"3f:4f:65:...","notes":"felhom local-api","size":4353457559, "verification":{"state":"ok","upid":"..."} ...} ``` | candidate | phantom | good PBS | good **local** (dir) | verdict | |---|---|---|---|---| | `verification` | **absent** | present | **absent** | UNUSABLE — would reject every local backup | | `encrypted` | absent | present | **absent** | UNUSABLE — same | | `notes` | absent | present | present | too fragile (agent-set only) | | `size` | **1 B** | 4.35 GB | 1.59 GB | **robust, tier-agnostic** | `verification` is absent on the phantom **and** on every local archive **and** on a good PBS snapshot until `verify-new` catches up — gating on it would have caused fleet-wide backup thrash. A PBS-level manifest query is not reachable from `BackupRunner` (PVE-only interface, tier-agnostic). ## The fix `minPlausibleArchiveBytes = 1 MiB`, measured not chosen — fleet survey 2026-07-28: | | bytes | |---|---| | smallest **real** backup anywhere | **612,397,450** (~584 MiB) | | demo-hp local / PBS | 1.59 GB / 4.35–4.37 GB | | demo-felhom local / PBS | 5.82–5.84 GB / 14.47–14.51 GB | | phantom | **1** | 1 MiB sits 584× below the smallest real backup and 1,048,576× above the phantom. `TestMinPlausibleArchiveBytes_LeavesHeadroomBelowTheSmallestRealBackup` asserts that headroom so the floor cannot be quietly raised into the thrash zone. **Fail-safe direction:** undecidable ⇒ **not** counted. Erring toward "less fresh" costs one extra backup; counting an undecidable entry is the defect. **Rejections are never silent:** WARN, once per distinct volid (not per 5-minute poll — ~288 lines/day would bury it), naming the snapshot, its size and the reason. ## Red-proofs — all four observed failing | red-proof | observed failure | |---|---| | remove the completeness test | `got ctime 1785216674, want 1785210042` · `the 1-byte manifest-less phantom set tier freshness — this is F-CRIT-2` | | reject everything (Scenario D) | `a REAL local dir vzdump ... backup was rejected — this is the backup-thrash path, not extra safety` | | flip the fail-safe (unknown ⇒ complete) | `size 0 counted as a complete backup — undecidable must fail safe` | | remove the once-per-volid dedupe | `rejection logged 5 times across 5 polls, want exactly 1` · `got 6 rejection lines for 2 distinct phantoms` | `go build`, `go vet ./...`, `go test ./...` — all green (run separately from the commit). ## §5 — LIVE re-test of the campaign injection (demo-hp) Fault 2 replayed against the **fixed** agent: PBS daemon killed mid-upload, phantom created (`2026-07-28T05:31:14Z`, 1 byte, no `index.json.blob`). The fix does not prevent the phantom — PBS creates it — it stops it *counting*. **Rejection logged, naming the snapshot:** ``` level=WARN msg="backup: ignoring an INCOMPLETE archive when computing tier freshness — it is not a successful backup" target=felhom-pbs vmid=9201 volid=felhom-pbs:backup/ct/9201/2026-07-28T05:31:14Z size_bytes=1 reason="size 1 B is below the 1048576 B plausibility floor — an aborted/incomplete archive, not a successful backup" ``` Logged **once across 8 due-checks**. (It appears twice in the day's journal only because the agent was restarted for the cadence change below — a process restart clears the in-memory dedupe, which is the documented behaviour.) **The discriminating test.** PBS cadence set to 1h so the two readings diverge: the phantom (24 min old) would read *not due*; the genuine backup (2h15m old) reads *due*. Pre-fix the tier stays silent; post-fix it must back up. ``` 05:53:17 [quiesce] backup due on 1 tier(s) — quiescing 1 stack(s): [paperless-ngx] 05:53:24 [quiesce] tier felhom-pbs: backup job backup-9201-felhom-pbs-1785218004218916301 started — polling 05:54:35 [quiesce] tier felhom-pbs: backup job ... done agent: backup: completed target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-07-28T05:53:24Z size_bytes=4359968099 ``` That single observation proves all three: the tier is **not** reported fresh, its age comes from the **genuine** backup, and it **is** due and attempted. **The inverse — no thrash.** After the successful backup, `backup job started` held at **3 → 3** across the following poll cycles, with **91 controller scheduler ticks** in the window proving the loop was alive (not silence from a dead loop). demo-felhom, which has no phantom, logged **0** rejections and **0** errors. ## Deployed `felhom-agent 0.106.0` on **demo-hp** and **demo-felhom**; both `active`, real cadences restored (`local 0` ⇒24h, `restore_test 302400`, `felhom-pbs 604800`). demo-felhom 15/15 containers healthy, demo-hp 8/8, zero unhealthy. Previous binaries kept as `/usr/local/bin/felhom-agent.bak-0.105.0`. ## Established while investigating — prune does NOT count phantoms A `keep-last 2` dry-run against three real snapshots plus the phantom retained **two real ones plus the phantom**: ``` 2026-07-27T21:00:03Z remove 2026-07-27T23:09:55Z keep 2026-07-28T03:40:42Z keep 2026-07-28T05:31:14Z keep ← PHANTOM ``` So the feared "one phantom ⇒ one real backup / two phantoms ⇒ zero" **does not occur — there is no retention or data-loss bug.** But prune never removes phantoms either, so they accumulate one per aborted upload, forever. Filed separately as **R-99** (hygiene, LOW), not as a retention bug. ## Left in place, deliberately The demo-hp phantom (`2026-07-28T05:31:14Z`) is **not** deleted. It is now demonstrably harmless (ignored, and announced once) and serves as a live regression fixture. No automatic phantom cleanup was added — deletion on a customer datastore is a separate ruling. ## NOT yet live-validated (carried forward from Campaign 8, minus what this closes) - **F-CRIT-1** — an app that fails to restart after a quiesce still never alarms. Untouched here. - Fault 4 — restic transport interruption; four injection approaches defeated by guest-bridged networking. **Still the most valuable follow-up**, because this fix answers the phantom question for PBS and leaves the analogous restic question open: does an interrupted restic run leave a partial pack that later reads as success? - Fault 12 (host reboot mid-backup); the three-way backup+restore-test+GC overlap; the agent's own DR bring-up path (`ModeDRGuestLoss`), still read-only-verified.