Files
felhom.eu/REPORT-fcrit2.md
T
admin d0cec9d3f3 docs: F-CRIT-2 shipped (agent v0.106.0) — OPEN-ITEMS, campaign doc, report
F-CRIT-2 marked SHIPPED + PROVEN-LIVE: NewestArchiveTime now counts only
plausibly-complete entries (measured 1 MiB floor), verified by replaying
campaign fault 2 on demo-hp in both directions.

Files two follow-ups:
- R-99: prune never REMOVES phantoms. It does NOT count them toward keep-last
  (dry-run kept 2 real + the phantom), so there is no retention/data-loss bug —
  but they accumulate one per aborted upload.
- F-CRIT-1: filed READY-HIGHEST. Campaign 8's other HIGH finding, untouched by
  this fix.
2026-07-28 08:08:56 +02:00

2.5 KiB

REPORT — F-CRIT-2 fixed: a failed backup no longer looks like a fresh one (2026-07-28)

Scope: felhom-agent v0.105.0 → v0.106.0. Docs here. Implementation detail and the full live re-test live in felhom-agent/REPORT.md; the campaign that found it is documentation/audits/CAMPAIGN-8-backup-restore-2026-07-27.md.

What changed

NewestArchiveTime counted an aborted PBS upload (1 byte, manifest-less, and NEWEST) as a successful backup, so the tier read fresh, went not due, and was never retried — seven days of silence on the real 168h cadence, invisible to both the R-88 breaker (defers only due tiers) and the hub deadline monitor (reads the same freshness). It now counts only plausibly complete entries via a measured 1 MiB floor; undecidable ⇒ not counted.

Size is the only tier-agnostic discriminator. verification and encrypted are absent on every local (dir) archive AND on a good PBS snapshot until verify-new catches up — gating on either would have rejected 100% of local backups and produced fleet-wide backup thrash. That inverse risk is a first-class test, red-proofed by making the filter reject everything.

Verified live, not just in unit tests

Campaign fault 2 was replayed against the fixed agent on demo-hp — phantom created, rejected and announced once; the tier correctly reported DUE and backed up (4,359,968,099 B landed); and the inverse showed no thrash, with 91 scheduler ticks as the positive observable that the loop was alive rather than dead.

Settled along the way — no retention bug

Server-side prune does not count phantoms toward keep-last: a dry-run against three real snapshots plus a phantom retained two real ones plus the phantom. The feared "two phantoms ⇒ zero real backups" does not occur. Prune never removes them either, so they accumulate one per aborted upload — filed as R-99 (LOW, hygiene), not as a retention bug.

Docs touched

  • documentation/backlog/OPEN-ITEMS.md — F-CRIT-2 → SHIPPED+PROVEN-LIVE; R-99 filed; F-CRIT-1 filed as READY-HIGHEST (Campaign 8's other HIGH finding, untouched here).
  • documentation/audits/CAMPAIGN-8-backup-restore-2026-07-27.md — F-CRIT-2 row struck through and a closing section added.

Still open, highest first

F-CRIT-1 — an app that fails to restart after a quiesce never alarms, for two independent reasons. Then fault 4 (restic transport interruption), which this fix makes more pointed: the phantom question is now answered for PBS and still open for restic.