Files
felhom-agent/REPORT.md
T
admin af1c21abc4 docs: F-CRIT-2 fix — REPORT, CONTEXT (v0.106.0)
Phase 0 discriminator survey, the measured 1 MiB floor and its justification,
four red-proofs with observed failure text, and the live re-test of campaign
fault 2 on demo-hp (both directions). Records that server-side prune does NOT
count phantoms toward keep-last — no retention bug — but never removes them
either (filed as R-99).
2026-07-28 08:07:42 +02:00

7.8 KiB
Raw Blame History

REPORT — F-CRIT-2: a failed backup must not look like a fresh one (v0.106.0) (2026-07-28)

Overwritten per the standing rule. Agent v0.105.0 → v0.106.0. Campaign 8 finding F-CRIT-2, fixed and re-tested live by replaying the original injection. Scope: felhom-agent only. No wire/contract change; the controller needed no change.

The defect

NewestArchiveTime (internal/backup/runner.go) selected the newest storage entry with e.Content == "backup" && e.VMID == vmid and no completeness test. A PBS daemon killed mid-upload leaves an aborted snapshot that PBS publishes into the same listing: manifest-less, 1 byte, and NEWEST. Counting it made the tier report freshly backed up, so it went not due and was never retried — seven days of silence on the real 168h cadence, with the R-88 breaker (defers only due tiers) and the hub deadline monitor (reads the same freshness) both blind to it.

R-84's "ask the storage, it is ground truth" was right. The bug is that presence was taken for validity — and a phantom is a more convincing lie than an empty array, because absence at least reads as absence.

Phase 0 — the discriminator survey

The campaign's own phantom was gone (Phase D deleted it deliberately to stop the suppression), so one was recreated on demo-hp by replaying fault 2. Raw PVE API, captured live:

PHANTOM: {"content":"backup","ctime":1785216674,"format":"pbs-ct","size":1,"subtype":"lxc",
          "vmid":9201,"volid":"felhom-pbs:backup/ct/9201/2026-07-28T05:31:14Z"}
GOOD:    {... "encrypted":"3f:4f:65:...","notes":"felhom local-api","size":4353457559,
          "verification":{"state":"ok","upid":"..."} ...}
candidate phantom good PBS good local (dir) verdict
verification absent present absent UNUSABLE — would reject every local backup
encrypted absent present absent UNUSABLE — same
notes absent present present too fragile (agent-set only)
size 1 B 4.35 GB 1.59 GB robust, tier-agnostic

verification is absent on the phantom and on every local archive and on a good PBS snapshot until verify-new catches up — gating on it would have caused fleet-wide backup thrash. A PBS-level manifest query is not reachable from BackupRunner (PVE-only interface, tier-agnostic).

The fix

minPlausibleArchiveBytes = 1 MiB, measured not chosen — fleet survey 2026-07-28:

bytes
smallest real backup anywhere 612,397,450 (~584 MiB)
demo-hp local / PBS 1.59 GB / 4.354.37 GB
demo-felhom local / PBS 5.825.84 GB / 14.4714.51 GB
phantom 1

1 MiB sits 584× below the smallest real backup and 1,048,576× above the phantom. TestMinPlausibleArchiveBytes_LeavesHeadroomBelowTheSmallestRealBackup asserts that headroom so the floor cannot be quietly raised into the thrash zone.

Fail-safe direction: undecidable ⇒ not counted. Erring toward "less fresh" costs one extra backup; counting an undecidable entry is the defect.

Rejections are never silent: WARN, once per distinct volid (not per 5-minute poll — ~288 lines/day would bury it), naming the snapshot, its size and the reason.

Red-proofs — all four observed failing

red-proof observed failure
remove the completeness test got ctime 1785216674, want 1785210042 · the 1-byte manifest-less phantom set tier freshness — this is F-CRIT-2
reject everything (Scenario D) a REAL local dir vzdump ... backup was rejected — this is the backup-thrash path, not extra safety
flip the fail-safe (unknown ⇒ complete) size 0 counted as a complete backup — undecidable must fail safe
remove the once-per-volid dedupe rejection logged 5 times across 5 polls, want exactly 1 · got 6 rejection lines for 2 distinct phantoms

go build, go vet ./..., go test ./... — all green (run separately from the commit).

§5 — LIVE re-test of the campaign injection (demo-hp)

Fault 2 replayed against the fixed agent: PBS daemon killed mid-upload, phantom created (2026-07-28T05:31:14Z, 1 byte, no index.json.blob). The fix does not prevent the phantom — PBS creates it — it stops it counting.

Rejection logged, naming the snapshot:

level=WARN msg="backup: ignoring an INCOMPLETE archive when computing tier freshness — it is not a successful backup"
  target=felhom-pbs vmid=9201 volid=felhom-pbs:backup/ct/9201/2026-07-28T05:31:14Z size_bytes=1
  reason="size 1 B is below the 1048576 B plausibility floor — an aborted/incomplete archive, not a successful backup"

Logged once across 8 due-checks. (It appears twice in the day's journal only because the agent was restarted for the cadence change below — a process restart clears the in-memory dedupe, which is the documented behaviour.)

The discriminating test. PBS cadence set to 1h so the two readings diverge: the phantom (24 min old) would read not due; the genuine backup (2h15m old) reads due. Pre-fix the tier stays silent; post-fix it must back up.

05:53:17 [quiesce] backup due on 1 tier(s) — quiescing 1 stack(s): [paperless-ngx]
05:53:24 [quiesce] tier felhom-pbs: backup job backup-9201-felhom-pbs-1785218004218916301 started — polling
05:54:35 [quiesce] tier felhom-pbs: backup job ... done
agent: backup: completed target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-07-28T05:53:24Z size_bytes=4359968099

That single observation proves all three: the tier is not reported fresh, its age comes from the genuine backup, and it is due and attempted.

The inverse — no thrash. After the successful backup, backup job started held at 3 → 3 across the following poll cycles, with 91 controller scheduler ticks in the window proving the loop was alive (not silence from a dead loop). demo-felhom, which has no phantom, logged 0 rejections and 0 errors.

Deployed

felhom-agent 0.106.0 on demo-hp and demo-felhom; both active, real cadences restored (local 0 ⇒24h, restore_test 302400, felhom-pbs 604800). demo-felhom 15/15 containers healthy, demo-hp 8/8, zero unhealthy. Previous binaries kept as /usr/local/bin/felhom-agent.bak-0.105.0.

Established while investigating — prune does NOT count phantoms

A keep-last 2 dry-run against three real snapshots plus the phantom retained two real ones plus the phantom:

2026-07-27T21:00:03Z  remove      2026-07-27T23:09:55Z  keep
2026-07-28T03:40:42Z  keep        2026-07-28T05:31:14Z  keep   ← PHANTOM

So the feared "one phantom ⇒ one real backup / two phantoms ⇒ zero" does not occur — there is no retention or data-loss bug. But prune never removes phantoms either, so they accumulate one per aborted upload, forever. Filed separately as R-99 (hygiene, LOW), not as a retention bug.

Left in place, deliberately

The demo-hp phantom (2026-07-28T05:31:14Z) is not deleted. It is now demonstrably harmless (ignored, and announced once) and serves as a live regression fixture. No automatic phantom cleanup was added — deletion on a customer datastore is a separate ruling.

NOT yet live-validated (carried forward from Campaign 8, minus what this closes)

  • F-CRIT-1 — an app that fails to restart after a quiesce still never alarms. Untouched here.
  • Fault 4 — restic transport interruption; four injection approaches defeated by guest-bridged networking. Still the most valuable follow-up, because this fix answers the phantom question for PBS and leaves the analogous restic question open: does an interrupted restic run leave a partial pack that later reads as success?
  • Fault 12 (host reboot mid-backup); the three-way backup+restore-test+GC overlap; the agent's own DR bring-up path (ModeDRGuestLoss), still read-only-verified.