# REPORT — CAMPAIGN 9: the restore paths, proven (2026-07-28) **Overwritten** per the standing rule. **No production code shipped** — this was a proof campaign, and findings are recorded, never fixed inline. Full write-up: `documentation/audits/CAMPAIGN-9-restore-proof-2026-07-28.md`. Evidence: `DooPlex:~/campaign9/evidence/` (69 files, 221 MB, 7 collectors, written continuously). Fleet unchanged and healthy at close: hub v0.80.0, agent v0.110.0, controller v0.182.0 on both boxes. **`peti-felhom` was never touched.** The ep0 rollback copy `/srv/pbs-felhom` (13 G) is intact. ## The headline — two never-proven restore paths are now proven Driven through the **real endpoints the UI posts to**, over https through traefik with a real session and CSRF token, on live hardware. | proof | result | |---|---| | **A1** — Tier-2 restore of ordinary app data (`paperless-ngx`, demo-hp) | 6 deleted files back **byte-identical** (`sha256sum -c` all OK) | | A1 — „A meglévő fájlok NEM módosulnak és NEM törlődnek" | 2 created files survived; 1 locally-edited file **not overwritten** (edit marker intact) | | A1 — app stopped/restarted and healthy | stop→copy→start in 39 s, `paperless-webserver` healthy | | A1 — data **usable by the app**, not just on disk | paperless resolved all 3 docs, checksums matched its own DB, and **served the restored bytes over its own HTTP API** at the exact pre-deletion sha256 | | **A2** — Tier-1 recovery-unit restore is a **distinct** path | `POST /backup/restore` → `RestoreFromRecoveryUnit`; ran end-to-end in 18 s, 1 volume restored, app healthy | | **A3** — restore after **total loss** (whole appdata dir `rm -rf`) | loss proven by doc download going **200 → 404**; restore returned **43/43 files byte-identical**, `documents_ok 16 of 16`, downloads back to 200 | The honest boundary A1+A3 together establish: **existing files are untouched; destroyed files return at their last-backup state.** ## Findings — 3 defects, ranked (none fixed) | # | finding | severity | |---|---|---| | **C9-F1** | The Tier-2 restore button is offered for apps it can **never** restore (BookStack, Docmost). It takes a real app outage, restores 0 files, and reports „Nincs hiányzó fájl — minden fájl megvan a helyén." — while 156 MB of that app's data sits unread in the same copy | **HIGH** | | **C9-F2** | An app in a **crash loop never alarms on any channel**. `StateRestarting` is in no down-set, so the dead-app heartbeat printed *"180 scans … 0 currently down"* while the app had been looping for 9 minutes | **HIGH** | | **C9-F3** | An **interrupted offsite run** leaves an exclusive restic lock the existing self-heal cannot reach; the tier is dead until a human unlocks, and the operator is told *"unknown reason"* | **MEDIUM** | Two things were deliberately **not** filed as defects: a recovery-unit poisoning that the catalog sync self-healed within ~3 minutes (proven live — reporting it would have been reporting an artifact), and a `snapshot_id` that looked ignored but is documented as logging-only and confirmed so live. ## Mechanisms confirmed working, live R-82's one-quiesce rule under mixed outcomes (2 tiers due, apps stopped **once**, per-target breaker); R-88's breaker (edge-triggered, one WARN, one event, three silent DEBUG skips, **no app thrash**); F-A1's contention deferral (409 → no breaker, no event, prompt restart — both sides of the seam captured in the same second); **F-CRIT-2's size filter against a real 1-byte phantom** on demo-hp, confirmed independently on ep0's filesystem; R-100's success anchor twice; **F-DIAG's sanitiser on the exact bare-hostname case that defeated its first version** (nothing raw reaches the hub event or the report); F-OBS's positive observable — which is precisely what made C9-F2 provable; F-LEAK's fenced destroy (no leaked `990000` guests across ~10 restore-tests). ## Where it stopped, and what remains Stopped at the **end of Phase B**, plus Phase D item 10, then full recovery. Phase C item 6 (host reboot mid-backup) was deliberately not started — a large new fault class against boxes that are remote until ~08-02, and starting it would have meant rushing it or leaving the fleet unknown. **Approved but impossible:** Phase 0 cleared compressing the hub's `staleAfter` for R-100's threshold test. It is **not a knob** — `cmd/hub/main.go:552` passes `0`, selecting the compile-time `defaultOffsiteStaleAfter = 48h`. Compressing it needed a hub code change, which the campaign forbids. Reported rather than worked around. The no-code-change alternative (age the controller's reported `last_success` past 48 h and let the hub judge at its real threshold) is the recommended method next time. **The honest residue — still not proven:** Tier-1 **content** recovery after real loss (A2 ran on an intact app; A3 used Tier-2) — now the most valuable open item; host reboot mid-backup; three-way concurrency with GC; Scenario C live; `offsite_stale` actually firing; F-HUB `SQLITE_BUSY`. ## Recovery Every config reverted from `evidence/config-before/REVERT.md`, each verified with a **positive observable**: agent cadences back to `0 / 302400 / 604800` on both hosts (`is-active` = active), windows back to `02:30`, `pvesm` shows `felhom-pbs active` on both, 0 campaign iptables rules on either host or guest, 0 scratch guests in the `990000` band, all stacks healthy on both boxes, and the offsite tier not merely unblocked but **proven working again** (`ok`, 1m35s, 8 snapshots). One benign residue: the in-memory R-88 breaker still holds a `felhom-pbs` failure count on each box. Its `until` is long past so it blocks nothing; it clears on the next successful backup or any controller restart (by design, not persisted). Clearing it would have cost another app outage for no benefit. **One operational lesson worth a runbook line:** a hand-run `docker compose up -d` in `/opt/docker/stacks/` starts a Felhom app **without its secrets** — they are injected by the controller's `stackEnv` at start time, not stored in a `.env`. It turned a healthy docmost into a crash loop during recovery. Manual recovery must go through `POST /api/stacks//restart`.