Phase A is the headline and it passed on live hardware, through the real endpoints the UI
posts to: a customer who deletes files — or their entire app data directory — gets everything
back byte-identical, and the app works afterwards (paperless served the restored bytes over
its own API at the exact pre-deletion sha256). A1's two non-destruction promises both hold.
Three defects, recorded not fixed:
C9-F1 (HIGH) the Tier-2 restore button is offered for apps it can never restore, takes a
real outage, and reports "nothing was missing" — indistinguishable from a
genuine result, while 156 MB of that app's data sits unread in the same copy.
C9-F2 (HIGH) an app in a crash loop never alarms on any channel; StateRestarting is in no
down-set, so F-OBS's own heartbeat printed "0 currently down" for 9 minutes.
C9-F3 (MEDIUM) an interrupted offsite run leaves a lock the self-heal cannot reach; the tier
is dead until a human unlocks and the operator is told "unknown reason".
This answers Phase C item 8.
Two candidates were deliberately NOT filed: a recovery-unit poisoning the catalog sync healed
in ~3 min, and a snapshot_id that is documented as logging-only. Reporting either would have
been reporting an artifact.
Stopped at the end of Phase B (plus D10), then full recovery — both boxes healthy, real
cadences, offsite tier proven working again, no leaked scratch guests, peti untouched.
D11's approved staleAfter compression turned out not to be a knob; reported, not worked around.
6.1 KiB
REPORT — CAMPAIGN 9: the restore paths, proven (2026-07-28)
Overwritten per the standing rule. No production code shipped — this was a proof campaign,
and findings are recorded, never fixed inline. Full write-up:
documentation/audits/CAMPAIGN-9-restore-proof-2026-07-28.md.
Evidence: DooPlex:~/campaign9/evidence/ (69 files, 221 MB, 7 collectors, written continuously).
Fleet unchanged and healthy at close: hub v0.80.0, agent v0.110.0, controller v0.182.0 on both boxes.
peti-felhom was never touched. The ep0 rollback copy /srv/pbs-felhom (13 G) is intact.
The headline — two never-proven restore paths are now proven
Driven through the real endpoints the UI posts to, over https through traefik with a real session and CSRF token, on live hardware.
| proof | result |
|---|---|
A1 — Tier-2 restore of ordinary app data (paperless-ngx, demo-hp) |
6 deleted files back byte-identical (sha256sum -c all OK) |
| A1 — „A meglévő fájlok NEM módosulnak és NEM törlődnek" | 2 created files survived; 1 locally-edited file not overwritten (edit marker intact) |
| A1 — app stopped/restarted and healthy | stop→copy→start in 39 s, paperless-webserver healthy |
| A1 — data usable by the app, not just on disk | paperless resolved all 3 docs, checksums matched its own DB, and served the restored bytes over its own HTTP API at the exact pre-deletion sha256 |
| A2 — Tier-1 recovery-unit restore is a distinct path | POST /backup/restore → RestoreFromRecoveryUnit; ran end-to-end in 18 s, 1 volume restored, app healthy |
A3 — restore after total loss (whole appdata dir rm -rf) |
loss proven by doc download going 200 → 404; restore returned 43/43 files byte-identical, documents_ok 16 of 16, downloads back to 200 |
The honest boundary A1+A3 together establish: existing files are untouched; destroyed files return at their last-backup state.
Findings — 3 defects, ranked (none fixed)
| # | finding | severity |
|---|---|---|
| C9-F1 | The Tier-2 restore button is offered for apps it can never restore (BookStack, Docmost). It takes a real app outage, restores 0 files, and reports „Nincs hiányzó fájl — minden fájl megvan a helyén." — while 156 MB of that app's data sits unread in the same copy | HIGH |
| C9-F2 | An app in a crash loop never alarms on any channel. StateRestarting is in no down-set, so the dead-app heartbeat printed "180 scans … 0 currently down" while the app had been looping for 9 minutes |
HIGH |
| C9-F3 | An interrupted offsite run leaves an exclusive restic lock the existing self-heal cannot reach; the tier is dead until a human unlocks, and the operator is told "unknown reason" | MEDIUM |
Two things were deliberately not filed as defects: a recovery-unit poisoning that the catalog
sync self-healed within ~3 minutes (proven live — reporting it would have been reporting an
artifact), and a snapshot_id that looked ignored but is documented as logging-only and confirmed
so live.
Mechanisms confirmed working, live
R-82's one-quiesce rule under mixed outcomes (2 tiers due, apps stopped once, per-target
breaker); R-88's breaker (edge-triggered, one WARN, one event, three silent DEBUG skips, no app
thrash); F-A1's contention deferral (409 → no breaker, no event, prompt restart — both sides of
the seam captured in the same second); F-CRIT-2's size filter against a real 1-byte phantom on
demo-hp, confirmed independently on ep0's filesystem; R-100's success anchor twice; F-DIAG's
sanitiser on the exact bare-hostname case that defeated its first version (nothing raw reaches the
hub event or the report); F-OBS's positive observable — which is precisely what made C9-F2 provable;
F-LEAK's fenced destroy (no leaked 990000 guests across ~10 restore-tests).
Where it stopped, and what remains
Stopped at the end of Phase B, plus Phase D item 10, then full recovery. Phase C item 6 (host reboot mid-backup) was deliberately not started — a large new fault class against boxes that are remote until ~08-02, and starting it would have meant rushing it or leaving the fleet unknown.
Approved but impossible: Phase 0 cleared compressing the hub's staleAfter for R-100's
threshold test. It is not a knob — cmd/hub/main.go:552 passes 0, selecting the compile-time
defaultOffsiteStaleAfter = 48h. Compressing it needed a hub code change, which the campaign
forbids. Reported rather than worked around. The no-code-change alternative (age the controller's
reported last_success past 48 h and let the hub judge at its real threshold) is the recommended
method next time.
The honest residue — still not proven: Tier-1 content recovery after real loss (A2 ran on an
intact app; A3 used Tier-2) — now the most valuable open item; host reboot mid-backup; three-way
concurrency with GC; Scenario C live; offsite_stale actually firing; F-HUB SQLITE_BUSY.
Recovery
Every config reverted from evidence/config-before/REVERT.md, each verified with a positive
observable: agent cadences back to 0 / 302400 / 604800 on both hosts (is-active = active),
windows back to 02:30, pvesm shows felhom-pbs active on both, 0 campaign iptables rules on
either host or guest, 0 scratch guests in the 990000 band, all stacks healthy on both boxes, and
the offsite tier not merely unblocked but proven working again (ok, 1m35s, 8 snapshots).
One benign residue: the in-memory R-88 breaker still holds a felhom-pbs failure count on each box.
Its until is long past so it blocks nothing; it clears on the next successful backup or any
controller restart (by design, not persisted). Clearing it would have cost another app outage for no
benefit.
One operational lesson worth a runbook line: a hand-run docker compose up -d in
/opt/docker/stacks/<app> starts a Felhom app without its secrets — they are injected by the
controller's stackEnv at start time, not stored in a .env. It turned a healthy docmost into a
crash loop during recovery. Manual recovery must go through POST /api/stacks/<name>/restart.