# REPORT — CAMPAIGN 8: the backup & restore subsystem (2026-07-27/28) Adversarial, destructive, unattended run against `demo-felhom`, `demo-hp` and `ep0`. Full report: `documentation/audits/CAMPAIGN-8-backup-restore-2026-07-27.md`. Evidence: `DooPlex:~/campaign8/evidence/` (103 files, 35 MB, written continuously by 11 collectors). **No production code was changed.** Findings are recorded and ranked, never fixed inline, per the campaign's own rules. ## Scope safety `peti-felhom`, its namespace and `u629488-sub2` were never touched. Phase 0 established with five documented probes that peti has **no data at all** in `felhom-offsite`, which is what made the operator-approved 100% datastore-fill safe. The 13 GB rollback copy `/srv/pbs-felhom` on ep0 is intact. ## Findings | # | Finding | Severity | Class | |---|---|---|---| | F-CRIT-1 | An app that fails to restart after a quiesce **never alarms**, on any channel | **HIGH** | DEFECT | | F-CRIT-2 | A failed offsite backup leaves a phantom snapshot that **resets tier freshness** (up to 7 days silent on real cadences) | **HIGH** | DEFECT | | F-A1 | A restore-test in progress makes a healthy backup report as FAILED, arms the breaker, pages the operator | MEDIUM | DEFECT (behaviour) / ARTIFACT (frequency) | | F-HUB | The hub dropped an event under concurrent load (`SQLITE_BUSY`), no retry, cause unnamed | MEDIUM | DEFECT | | F-LEAK | A **failed** restore-test cannot destroy its own scratch guest (403 `VM.Allocate`); leaks are never reclaimed | MEDIUM | DEFECT (root-caused by fault 18) | | F-REBOOT | A guest rebooted during its backup **does not come back** — shutdown completes, start never happens, no self-heal | MEDIUM | DEFECT | | F-DIAG | Four distinct offsite failure causes collapse into two operator-visible strings | LOW–MED | DEFECT | | F-OBS | `deadapp-check` leaves no positive observable on a default (info-level) box | LOW | DEFECT | | F-OPS | A manual `pct restore` inherits the source guest's binds (live data drive + another guest's credentials) | LOW | Operational | Both HIGH findings are in the same place: **the system's ability to tell you a backup did not happen.** Both cite the code and the comment that asserts the property the code does not provide. ## What is now proven that was not before R-88 breaker arming **and its full ladder** (15m/30m/1h/2h/4h/4h-cap) · per-tier isolation under a real one-tier-fails case · `whole_guest_backup_failed` end-to-end with correct tier attribution · R-97c operator-only routing verified against the hub DB (zero customer rows with `status='sent'`) · `whole_guest_backup_recovered` + the R-68 pairing gate firing live · **`age_state=absent`** · R-97b's suppression half · the crash-recovery unquiesce by an actual SIGKILL (1 s) · **R-87 — the first restic restore round-trip ever performed**, byte-verified (6/7 sha256 identical, the 7th explained) · R-82 one-quiesce-two-tiers · single-flight on two independent paths. ## Restore round-trips restic (R-87) · local vzdump → fresh CT · PBS offsite → fresh CT · corrupted snapshot → fails cleanly. `mount_parity` exact on both whole-guest tiers, `unprivileged: 1` preserved. ## Fleet state **Healthy. Nothing left broken.** All four compression knobs reverted and verified; every fault unwound; no leaked scratch guests, nft rules, ballast files or clock skew; ep0 datastore clean with zero `.bad` chunks. demo-felhom 15/15 containers healthy, demo-hp 8/8. ## Not tested (with reasons) Fault 4 (restic transport — four injection approaches defeated by guest-bridged networking; **the most valuable follow-up**, because F-CRIT-2 raises the same question for restic), fault 12 (host reboot — reasoned skip), and the agent's own DR bring-up path. Faults 6 and 8 were inconclusive for documented reasons. Faults 11 and 18 WERE run in the campaign's tail and both produced findings. **Campaign-caused outage, stated plainly:** fault 11 took demo-hp guest 9201 down for ~9m47s (the guest did not restart after a mid-backup reboot) until manually started. Fleet healthy after. ## Note on repo conventions This run touched no `hub/`, `scripts/` or `website/` code, so none of the per-area CHANGELOGs has an entry — there is nothing shipped to log. The deliverable is the audit document plus this report.