Phase 0 sized C9-F1 properly before anything was designed: 43 of the 53 catalog apps have NO subtree the Tier-2 restore can read (not 2), 9 are covered only for their file legs and never their database or volumes, 1 is stateless. The asymmetry is Tier-2's alone — Tier-1 and offsite both restore the unit and replay volume dumps, so BookStack always had a working restore and only this button lied. Shipped: the restore refuses BEFORE stopping the app and names the action that does work; a run that proceeds claims only what it EXAMINED and discloses that the database and volumes are not covered. C9-F2 alarms after a 5-minute sustained-restarting threshold, set above the 120s deploy timeout, Mealie's 60s start_period and R-97b's 180s grace; StateRestarting is deliberately NOT added to IsDownState. Live: silent through ten 30s samples then app_start_failed at 5m25s, heartbeat now reads "1 currently down" where Campaign 9 recorded 0; a real deploy stayed silent; bookstack refused with its uptime unbroken; paperless re-restored 43/43 byte-identical, 16/16 docs clean. Filed, not fixed: C9-F1b (route to the Tier-1 restore — its own task because it puts a destructive operation behind a non-destructive button) and C9-F4 (nothing reads the Tier-2 copy's recovery-unit/ mirror, so the second local copy that exists for drive loss is unreachable by any customer action — potentially larger than C9-F1).
6.8 KiB
REPORT — C9-F1 + C9-F2: a restore that restored nothing, and a crash loop nobody saw (2026-07-28)
Overwritten per the standing rule. Shipped: controller v0.183.0, live on both demo boxes.
Fleet: hub v0.80.0, agent v0.110.0, controller 0.183.0. peti-felhom untouched.
Both defects are the same shape — the system reporting healthy while the customer is not — and both live in the same status-derivation code.
Phase 0 — the asymmetry, sized before designing anything
Tier-2 writes two things on every run: the capture legs (hdd/, userdata/) and, always, a full
recovery-unit/ (DB dumps + named-volume tarballs). RestoreTier2Files reads only the two legs
(tier2_restore.go:101-104) and has never opened recovery-unit/.
All 53 catalog templates enumerated, cross-checked against both boxes' actual copies:
| class | count | what the restore can return |
|---|---|---|
| A | 9 | the file legs only — never their database or named volumes |
| B | 43 | nothing at all — a guaranteed no-op, forever |
| C | 1 | bentopdf, stateless |
Four apps (plex, jellyfin, emby, navidrome) are in B only because their single bind is a
:ro media mount, which ClassifyBinds correctly excludes. 81% of the catalog.
The asymmetry is Tier-2's alone. Tier-1 (RestoreFromRecoveryUnit) and offsite both restore the
unit and replay volume dumps — so BookStack already had a working restore; only this button lied.
Shipped — Part 1a (honesty)
- Refuses UP FRONT.
Tier2RestoreCoverageis consulted before any op begins; a class-B app is refused without being stopped. Live: BookStack uptime stayedUp About an hour(Campaign 9 left it atUp 25 seconds). - Names the working action rather than dead-ending 81% of the catalog:
„Ennek az alkalmazásnak az adatai nem ebből a másolatból állíthatók vissza — az alkalmazás nem állt le. Használd a Visszaállítás indítása gombot a Biztonsági mentés → Visszaállítás oldalon."
- Claims only what was examined (the QUIET half — immich's 1.3 GB Postgres unit is not covered,
so the old blanket sentence was a clean bill of health over data never opened):
„Minden vizsgált fájl megvan a helyén." + „Az alkalmazás adatbázisa és belső kötetei nem tartoznak ebbe a visszaállításba."
Shipped — Part 2 (C9-F2)
StateRestarting is deliberately NOT added to IsDownState — that alarms on every deploy and
update fleet-wide. A sustained run becomes down after crashLoopAfter = 5m, set above the three
real numbers already in the codebase: the deploy flow's 120 s health timeout, Mealie's 60 s
start_period, and R-97b's 180 s quiesce grace (so the windows compose into one bounded delay
instead of leaving a gap). Docker's backoff caps at 60 s, so a real loop registers ≥4 attempts inside
it. Carried by Stack.RestartingSince (not persisted) + Stack.CrashLooping(now), used by both
the alarm and the dashboard counter — which previously counted restarting as running and so
contradicted the alarm on the same screen.
Live replay (demo-hp + demo-felhom)
| # | scenario | result |
|---|---|---|
| 1 | crash loop alarms | Ten consecutive 30 s samples silent through the threshold window, then 17:05:40 Event pushed: app_start_failed (warn) — Telepített alkalmazás nem fut: Uptime Kuma — 5m25s after the loop began (5 min + one scan). |
| 1b | heartbeat COUNTS it | 17:06:10 [deadapp] check alive: 20 scans since boot, 2 deployed app(s) evaluated, **1 currently down** — Campaign 9's evidence was 0 currently down while an app looped. |
| 2 | normal deploy is silent | A real uptime-kuma deploy produced only app_deployed (info); no alarm, with deadapp-check running every 30 s throughout. |
| 3 | restore refuses without an outage | The honest message rendered; [WARN] Tier-2 file restore refused up front: stack=bookstack has no restorable subtree in its copy (unit_present=true) — app NOT stopped; BookStack uptime unbroken. |
| 4 | paperless still restores (regression guard on Campaign 9's headline) | 3 files deleted → restored → 43/43 byte-identical to the pre-deletion sha256 set, documents_ok 16 of 16 problems []. |
Filed, not fixed
- C9-F1b — route class-B apps to the Tier-1 unit restore from the card the customer already opened. Its own task deliberately: it puts a DESTRUCTIVE operation (overwrites live data with the backup state) behind a button reached via a NON-destructive one, so the confirm copy must carry that difference.
- C9-F4 — nothing reads the Tier-2 copy's
recovery-unit/mirror. Written by every Tier-2 run (tier2.go:369), read by no path:RecoveryUnitPathresolves tobackups/**primary**/(appbackup/paths.go:46-48), and the only reader of the secondary tree istier2_restore.go:79. Tier-2 exists for the case where the PRIMARY drive is lost — and in exactly that case the primary unit is gone while this mirror survives on the second drive, unreachable by any customer action, leaving offsite as the only route. Potentially larger than C9-F1.
Tests
go test ./... rc=0, 27 packages — run and rc read separately from the commit. Six red-proofs
all observed, including the one that matters most: adding StateRestarting to IsDownState fails the
brief-restart test with "every deploy and update would page the operator".
Observations
- A seventh shipped-invariant-comment.
controller/README.mdstated that faults "still surface asexited/degraded/restarting/unhealthy" — butrestartingwas in no down set at all. The sentence was a wish with no test pinning it. Corrected in place, with the threshold rule documented beside it. - Pre-existing gate failure, not mine:
scripts/docker_run_volume_path_gate.pyfails oninternal/appexport/estimate.go:179(an unrevieweddocker run -v). Verified it fails identically on clean HEAD; left alone as out of scope rather than silently "fixed". - The other six gates pass (
template_id,emoji,mojibake,native_confirm,app_row_dedup,offbox_rename).
NOT yet live-validated — carried forward
- Tier-1 content recovery after real loss — still the most valuable unproven item (Campaign 9's A2 ran against an intact app; A3 used Tier-2). Unchanged by this work.
- The C9-F2 threshold under a quiesce cycle (Scenario C) is unit-proven but was not replayed live; it needs a backup window plus an app that fails to come back.
- C9-F1's refusal for the other 42 class-B apps is proven by enumeration and by BookStack live, not app-by-app.
- Host reboot mid-backup, three-way concurrency with GC, Scenario C live,
offsite_stalefiring, F-HUB — all still open from Campaign 9.