Tail-of-campaign additions after the Phase D revert (both re-injections declared):
- fault 18 (delete a snapshot mid restore-test): detection PASS, and it ROOT-CAUSES
F-LEAK — a failed restore-test cannot destroy its own scratch guest (403,
missing VM.Allocate; the agent token is pool-scoped and a failed restore never
joins the felhom pool)
- fault 11 (guest reboot mid-backup): new finding F-REBOOT — the backup succeeds
but the guest never comes back; ~9m47s outage until a manual pct start
Two evidence corrections, both self-inflicted tooling errors:
- pgrep -cf <pattern> matches its own ssh command line, which invalidated fault
11's first two injections and put one unsound line in fault 9 (withdrawn; that
finding stands on the controller's own job state)
- ep0 runs Etc/UTC, so its 03:30 prune fires at 05:30 CEST — nearly misread as a
broken prune job
Nine findings now, still two HIGH. Fleet healthy.
Unattended 10h run against demo-felhom, demo-hp and ep0. No production code
changed; findings recorded and ranked, not fixed inline.
8 findings, 2 HIGH — both in the system's ability to report that a backup did
NOT happen:
- F-CRIT-1: an app failing to restart after a quiesce never alarms (invariant
I1 in main.go:1213 is false for the failed-restart path)
- F-CRIT-2: a failed offsite backup leaves a phantom snapshot that resets the
tier's freshness clock (NewestArchiveTime has no completeness check)
Retires several never-validated items, including R-87 (first restic restore
round-trip, byte-verified), the full R-88 backoff ladder, age_state=absent,
and the crash-recovery unquiesce under a real SIGKILL.
peti-felhom untouched; ep0 rollback copy intact; fleet healthy at end.