Files
felhom.eu/REPORT-campaign9.md
T
admin 955083c0fc Campaign 9: the Tier-2 restore paths are PROVEN; 3 defects filed, none fixed
Phase A is the headline and it passed on live hardware, through the real endpoints the UI
posts to: a customer who deletes files — or their entire app data directory — gets everything
back byte-identical, and the app works afterwards (paperless served the restored bytes over
its own API at the exact pre-deletion sha256). A1's two non-destruction promises both hold.

Three defects, recorded not fixed:
  C9-F1 (HIGH)   the Tier-2 restore button is offered for apps it can never restore, takes a
                 real outage, and reports "nothing was missing" — indistinguishable from a
                 genuine result, while 156 MB of that app's data sits unread in the same copy.
  C9-F2 (HIGH)   an app in a crash loop never alarms on any channel; StateRestarting is in no
                 down-set, so F-OBS's own heartbeat printed "0 currently down" for 9 minutes.
  C9-F3 (MEDIUM) an interrupted offsite run leaves a lock the self-heal cannot reach; the tier
                 is dead until a human unlocks and the operator is told "unknown reason".
                 This answers Phase C item 8.

Two candidates were deliberately NOT filed: a recovery-unit poisoning the catalog sync healed
in ~3 min, and a snapshot_id that is documented as logging-only. Reporting either would have
been reporting an artifact.

Stopped at the end of Phase B (plus D10), then full recovery — both boxes healthy, real
cadences, offsite tier proven working again, no leaked scratch guests, peti untouched.
D11's approved staleAfter compression turned out not to be a knob; reported, not worked around.
2026-07-28 18:24:06 +02:00

6.1 KiB

REPORT — CAMPAIGN 9: the restore paths, proven (2026-07-28)

Overwritten per the standing rule. No production code shipped — this was a proof campaign, and findings are recorded, never fixed inline. Full write-up: documentation/audits/CAMPAIGN-9-restore-proof-2026-07-28.md. Evidence: DooPlex:~/campaign9/evidence/ (69 files, 221 MB, 7 collectors, written continuously).

Fleet unchanged and healthy at close: hub v0.80.0, agent v0.110.0, controller v0.182.0 on both boxes. peti-felhom was never touched. The ep0 rollback copy /srv/pbs-felhom (13 G) is intact.

The headline — two never-proven restore paths are now proven

Driven through the real endpoints the UI posts to, over https through traefik with a real session and CSRF token, on live hardware.

proof result
A1 — Tier-2 restore of ordinary app data (paperless-ngx, demo-hp) 6 deleted files back byte-identical (sha256sum -c all OK)
A1 — „A meglévő fájlok NEM módosulnak és NEM törlődnek" 2 created files survived; 1 locally-edited file not overwritten (edit marker intact)
A1 — app stopped/restarted and healthy stop→copy→start in 39 s, paperless-webserver healthy
A1 — data usable by the app, not just on disk paperless resolved all 3 docs, checksums matched its own DB, and served the restored bytes over its own HTTP API at the exact pre-deletion sha256
A2 — Tier-1 recovery-unit restore is a distinct path POST /backup/restoreRestoreFromRecoveryUnit; ran end-to-end in 18 s, 1 volume restored, app healthy
A3 — restore after total loss (whole appdata dir rm -rf) loss proven by doc download going 200 → 404; restore returned 43/43 files byte-identical, documents_ok 16 of 16, downloads back to 200

The honest boundary A1+A3 together establish: existing files are untouched; destroyed files return at their last-backup state.

Findings — 3 defects, ranked (none fixed)

# finding severity
C9-F1 The Tier-2 restore button is offered for apps it can never restore (BookStack, Docmost). It takes a real app outage, restores 0 files, and reports „Nincs hiányzó fájl — minden fájl megvan a helyén." — while 156 MB of that app's data sits unread in the same copy HIGH
C9-F2 An app in a crash loop never alarms on any channel. StateRestarting is in no down-set, so the dead-app heartbeat printed "180 scans … 0 currently down" while the app had been looping for 9 minutes HIGH
C9-F3 An interrupted offsite run leaves an exclusive restic lock the existing self-heal cannot reach; the tier is dead until a human unlocks, and the operator is told "unknown reason" MEDIUM

Two things were deliberately not filed as defects: a recovery-unit poisoning that the catalog sync self-healed within ~3 minutes (proven live — reporting it would have been reporting an artifact), and a snapshot_id that looked ignored but is documented as logging-only and confirmed so live.

Mechanisms confirmed working, live

R-82's one-quiesce rule under mixed outcomes (2 tiers due, apps stopped once, per-target breaker); R-88's breaker (edge-triggered, one WARN, one event, three silent DEBUG skips, no app thrash); F-A1's contention deferral (409 → no breaker, no event, prompt restart — both sides of the seam captured in the same second); F-CRIT-2's size filter against a real 1-byte phantom on demo-hp, confirmed independently on ep0's filesystem; R-100's success anchor twice; F-DIAG's sanitiser on the exact bare-hostname case that defeated its first version (nothing raw reaches the hub event or the report); F-OBS's positive observable — which is precisely what made C9-F2 provable; F-LEAK's fenced destroy (no leaked 990000 guests across ~10 restore-tests).

Where it stopped, and what remains

Stopped at the end of Phase B, plus Phase D item 10, then full recovery. Phase C item 6 (host reboot mid-backup) was deliberately not started — a large new fault class against boxes that are remote until ~08-02, and starting it would have meant rushing it or leaving the fleet unknown.

Approved but impossible: Phase 0 cleared compressing the hub's staleAfter for R-100's threshold test. It is not a knobcmd/hub/main.go:552 passes 0, selecting the compile-time defaultOffsiteStaleAfter = 48h. Compressing it needed a hub code change, which the campaign forbids. Reported rather than worked around. The no-code-change alternative (age the controller's reported last_success past 48 h and let the hub judge at its real threshold) is the recommended method next time.

The honest residue — still not proven: Tier-1 content recovery after real loss (A2 ran on an intact app; A3 used Tier-2) — now the most valuable open item; host reboot mid-backup; three-way concurrency with GC; Scenario C live; offsite_stale actually firing; F-HUB SQLITE_BUSY.

Recovery

Every config reverted from evidence/config-before/REVERT.md, each verified with a positive observable: agent cadences back to 0 / 302400 / 604800 on both hosts (is-active = active), windows back to 02:30, pvesm shows felhom-pbs active on both, 0 campaign iptables rules on either host or guest, 0 scratch guests in the 990000 band, all stacks healthy on both boxes, and the offsite tier not merely unblocked but proven working again (ok, 1m35s, 8 snapshots).

One benign residue: the in-memory R-88 breaker still holds a felhom-pbs failure count on each box. Its until is long past so it blocks nothing; it clears on the next successful backup or any controller restart (by design, not persisted). Clearing it would have cost another app outage for no benefit.

One operational lesson worth a runbook line: a hand-run docker compose up -d in /opt/docker/stacks/<app> starts a Felhom app without its secrets — they are injected by the controller's stackEnv at start time, not stored in a .env. It turned a healthy docmost into a crash loop during recovery. Manual recovery must go through POST /api/stacks/<name>/restart.