# REPORT — C9-F1 + C9-F2: a restore that restored nothing, and a crash loop nobody saw (2026-07-28) **Overwritten** per the standing rule. **Shipped: controller v0.183.0**, live on **both** demo boxes. Fleet: hub v0.80.0, agent v0.110.0, controller **0.183.0**. `peti-felhom` untouched. Both defects are the same shape — the system reporting healthy while the customer is not — and both live in the same status-derivation code. ## Phase 0 — the asymmetry, sized before designing anything Tier-2 writes **two** things on every run: the capture legs (`hdd/`, `userdata/`) and, always, a full `recovery-unit/` (DB dumps + named-volume tarballs). `RestoreTier2Files` reads **only the two legs** (`tier2_restore.go:101-104`) and has never opened `recovery-unit/`. All 53 catalog templates enumerated, cross-checked against both boxes' actual copies: | class | count | what the restore can return | |---|---|---| | **A** | **9** | the file legs only — **never** their database or named volumes | | **B** | **43** | **nothing at all** — a guaranteed no-op, forever | | C | 1 | `bentopdf`, stateless | Four apps (`plex`, `jellyfin`, `emby`, `navidrome`) are in B only because their single bind is a `:ro` media mount, which `ClassifyBinds` correctly excludes. **81% of the catalog.** **The asymmetry is Tier-2's alone.** Tier-1 (`RestoreFromRecoveryUnit`) and offsite both restore the unit and replay volume dumps — so BookStack already had a working restore; only this button lied. ## Shipped — Part 1a (honesty) - **Refuses UP FRONT.** `Tier2RestoreCoverage` is consulted before any op begins; a class-B app is refused **without being stopped**. Live: BookStack uptime stayed `Up About an hour` (Campaign 9 left it at `Up 25 seconds`). - **Names the working action** rather than dead-ending 81% of the catalog: > „Ennek az alkalmazásnak az adatai nem ebből a másolatból állíthatók vissza — az alkalmazás nem > állt le. Használd a Visszaállítás indítása gombot a Biztonsági mentés → Visszaállítás oldalon." - **Claims only what was examined** (the QUIET half — immich's 1.3 GB Postgres unit is not covered, so the old blanket sentence was a clean bill of health over data never opened): > „Minden vizsgált fájl megvan a helyén." + „Az alkalmazás adatbázisa és belső kötetei nem > tartoznak ebbe a visszaállításba." ## Shipped — Part 2 (C9-F2) `StateRestarting` is deliberately **NOT** added to `IsDownState` — that alarms on every deploy and update fleet-wide. A **sustained** run becomes down after `crashLoopAfter = 5m`, set above the three real numbers already in the codebase: the deploy flow's **120 s** health timeout, Mealie's **60 s** `start_period`, and R-97b's **180 s** quiesce grace (so the windows compose into one bounded delay instead of leaving a gap). Docker's backoff caps at 60 s, so a real loop registers ≥4 attempts inside it. Carried by `Stack.RestartingSince` (not persisted) + `Stack.CrashLooping(now)`, used by **both** the alarm and the dashboard counter — which previously counted `restarting` as running and so contradicted the alarm on the same screen. ## Live replay (demo-hp + demo-felhom) | # | scenario | result | |---|---|---| | 1 | **crash loop alarms** | Ten consecutive 30 s samples **silent** through the threshold window, then `17:05:40 Event pushed: app_start_failed (warn) — Telepített alkalmazás nem fut: Uptime Kuma` — **5m25s** after the loop began (5 min + one scan). | | 1b | **heartbeat COUNTS it** | `17:06:10 [deadapp] check alive: 20 scans since boot, 2 deployed app(s) evaluated, **1 currently down**` — Campaign 9's evidence was `0 currently down` while an app looped. | | 2 | **normal deploy is silent** | A real `uptime-kuma` deploy produced only `app_deployed (info)`; no alarm, with deadapp-check running every 30 s throughout. | | 3 | **restore refuses without an outage** | The honest message rendered; `[WARN] Tier-2 file restore refused up front: stack=bookstack has no restorable subtree in its copy (unit_present=true) — app NOT stopped`; BookStack uptime unbroken. | | 4 | **paperless still restores** (regression guard on Campaign 9's headline) | 3 files deleted → restored → **43/43 byte-identical to the pre-deletion sha256 set**, `documents_ok 16 of 16 problems []`. | ## Filed, not fixed - **C9-F1b** — route class-B apps to the Tier-1 unit restore from the card the customer already opened. Its own task **deliberately**: it puts a DESTRUCTIVE operation (overwrites live data with the backup state) behind a button reached via a NON-destructive one, so the confirm copy must carry that difference. - **C9-F4** — **nothing reads the Tier-2 copy's `recovery-unit/` mirror.** Written by every Tier-2 run (`tier2.go:369`), read by no path: `RecoveryUnitPath` resolves to `backups/**primary**/` (`appbackup/paths.go:46-48`), and the only reader of the secondary tree is `tier2_restore.go:79`. Tier-2 exists for the case where the PRIMARY drive is lost — and in exactly that case the primary unit is gone while this mirror survives on the second drive, unreachable by any customer action, leaving offsite as the only route. **Potentially larger than C9-F1.** ## Tests `go test ./...` **rc=0, 27 packages** — run and `rc` read *separately* from the commit. Six red-proofs all observed, including the one that matters most: adding `StateRestarting` to `IsDownState` fails the brief-restart test with *"every deploy and update would page the operator"*. ## Observations - **A seventh shipped-invariant-comment.** `controller/README.md` stated that faults "still surface as `exited`/`degraded`/**`restarting`**/`unhealthy`" — but `restarting` was in no down set at all. The sentence was a wish with no test pinning it. Corrected in place, with the threshold rule documented beside it. - **Pre-existing gate failure, not mine:** `scripts/docker_run_volume_path_gate.py` fails on `internal/appexport/estimate.go:179` (an unreviewed `docker run -v`). Verified it fails identically on clean HEAD; left alone as out of scope rather than silently "fixed". - The other six gates pass (`template_id`, `emoji`, `mojibake`, `native_confirm`, `app_row_dedup`, `offbox_rename`). ## NOT yet live-validated — carried forward - **Tier-1 content recovery after real loss** — still the most valuable unproven item (Campaign 9's A2 ran against an intact app; A3 used Tier-2). Unchanged by this work. - The C9-F2 threshold under a **quiesce** cycle (Scenario C) is unit-proven but was not replayed live; it needs a backup window plus an app that fails to come back. - C9-F1's refusal for the other 42 class-B apps is proven by enumeration and by BookStack live, not app-by-app. - Host reboot mid-backup, three-way concurrency with GC, Scenario C live, `offsite_stale` firing, F-HUB — all still open from Campaign 9.