ff050cf409
Phase 0 sized C9-F1 properly before anything was designed: 43 of the 53 catalog apps have NO subtree the Tier-2 restore can read (not 2), 9 are covered only for their file legs and never their database or volumes, 1 is stateless. The asymmetry is Tier-2's alone — Tier-1 and offsite both restore the unit and replay volume dumps, so BookStack always had a working restore and only this button lied. Shipped: the restore refuses BEFORE stopping the app and names the action that does work; a run that proceeds claims only what it EXAMINED and discloses that the database and volumes are not covered. C9-F2 alarms after a 5-minute sustained-restarting threshold, set above the 120s deploy timeout, Mealie's 60s start_period and R-97b's 180s grace; StateRestarting is deliberately NOT added to IsDownState. Live: silent through ten 30s samples then app_start_failed at 5m25s, heartbeat now reads "1 currently down" where Campaign 9 recorded 0; a real deploy stayed silent; bookstack refused with its uptime unbroken; paperless re-restored 43/43 byte-identical, 16/16 docs clean. Filed, not fixed: C9-F1b (route to the Tier-1 restore — its own task because it puts a destructive operation behind a non-destructive button) and C9-F4 (nothing reads the Tier-2 copy's recovery-unit/ mirror, so the second local copy that exists for drive loss is unreachable by any customer action — potentially larger than C9-F1).
104 lines
6.8 KiB
Markdown
104 lines
6.8 KiB
Markdown
# REPORT — C9-F1 + C9-F2: a restore that restored nothing, and a crash loop nobody saw (2026-07-28)
|
|
|
|
**Overwritten** per the standing rule. **Shipped: controller v0.183.0**, live on **both** demo boxes.
|
|
Fleet: hub v0.80.0, agent v0.110.0, controller **0.183.0**. `peti-felhom` untouched.
|
|
|
|
Both defects are the same shape — the system reporting healthy while the customer is not — and both
|
|
live in the same status-derivation code.
|
|
|
|
## Phase 0 — the asymmetry, sized before designing anything
|
|
|
|
Tier-2 writes **two** things on every run: the capture legs (`hdd/`, `userdata/`) and, always, a full
|
|
`recovery-unit/` (DB dumps + named-volume tarballs). `RestoreTier2Files` reads **only the two legs**
|
|
(`tier2_restore.go:101-104`) and has never opened `recovery-unit/`.
|
|
|
|
All 53 catalog templates enumerated, cross-checked against both boxes' actual copies:
|
|
|
|
| class | count | what the restore can return |
|
|
|---|---|---|
|
|
| **A** | **9** | the file legs only — **never** their database or named volumes |
|
|
| **B** | **43** | **nothing at all** — a guaranteed no-op, forever |
|
|
| C | 1 | `bentopdf`, stateless |
|
|
|
|
Four apps (`plex`, `jellyfin`, `emby`, `navidrome`) are in B only because their single bind is a
|
|
`:ro` media mount, which `ClassifyBinds` correctly excludes. **81% of the catalog.**
|
|
|
|
**The asymmetry is Tier-2's alone.** Tier-1 (`RestoreFromRecoveryUnit`) and offsite both restore the
|
|
unit and replay volume dumps — so BookStack already had a working restore; only this button lied.
|
|
|
|
## Shipped — Part 1a (honesty)
|
|
|
|
- **Refuses UP FRONT.** `Tier2RestoreCoverage` is consulted before any op begins; a class-B app is
|
|
refused **without being stopped**. Live: BookStack uptime stayed `Up About an hour` (Campaign 9
|
|
left it at `Up 25 seconds`).
|
|
- **Names the working action** rather than dead-ending 81% of the catalog:
|
|
> „Ennek az alkalmazásnak az adatai nem ebből a másolatból állíthatók vissza — az alkalmazás nem
|
|
> állt le. Használd a Visszaállítás indítása gombot a Biztonsági mentés → Visszaállítás oldalon."
|
|
- **Claims only what was examined** (the QUIET half — immich's 1.3 GB Postgres unit is not covered,
|
|
so the old blanket sentence was a clean bill of health over data never opened):
|
|
> „Minden vizsgált fájl megvan a helyén." + „Az alkalmazás adatbázisa és belső kötetei nem
|
|
> tartoznak ebbe a visszaállításba."
|
|
|
|
## Shipped — Part 2 (C9-F2)
|
|
|
|
`StateRestarting` is deliberately **NOT** added to `IsDownState` — that alarms on every deploy and
|
|
update fleet-wide. A **sustained** run becomes down after `crashLoopAfter = 5m`, set above the three
|
|
real numbers already in the codebase: the deploy flow's **120 s** health timeout, Mealie's **60 s**
|
|
`start_period`, and R-97b's **180 s** quiesce grace (so the windows compose into one bounded delay
|
|
instead of leaving a gap). Docker's backoff caps at 60 s, so a real loop registers ≥4 attempts inside
|
|
it. Carried by `Stack.RestartingSince` (not persisted) + `Stack.CrashLooping(now)`, used by **both**
|
|
the alarm and the dashboard counter — which previously counted `restarting` as running and so
|
|
contradicted the alarm on the same screen.
|
|
|
|
## Live replay (demo-hp + demo-felhom)
|
|
|
|
| # | scenario | result |
|
|
|---|---|---|
|
|
| 1 | **crash loop alarms** | Ten consecutive 30 s samples **silent** through the threshold window, then `17:05:40 Event pushed: app_start_failed (warn) — Telepített alkalmazás nem fut: Uptime Kuma` — **5m25s** after the loop began (5 min + one scan). |
|
|
| 1b | **heartbeat COUNTS it** | `17:06:10 [deadapp] check alive: 20 scans since boot, 2 deployed app(s) evaluated, **1 currently down**` — Campaign 9's evidence was `0 currently down` while an app looped. |
|
|
| 2 | **normal deploy is silent** | A real `uptime-kuma` deploy produced only `app_deployed (info)`; no alarm, with deadapp-check running every 30 s throughout. |
|
|
| 3 | **restore refuses without an outage** | The honest message rendered; `[WARN] Tier-2 file restore refused up front: stack=bookstack has no restorable subtree in its copy (unit_present=true) — app NOT stopped`; BookStack uptime unbroken. |
|
|
| 4 | **paperless still restores** (regression guard on Campaign 9's headline) | 3 files deleted → restored → **43/43 byte-identical to the pre-deletion sha256 set**, `documents_ok 16 of 16 problems []`. |
|
|
|
|
## Filed, not fixed
|
|
|
|
- **C9-F1b** — route class-B apps to the Tier-1 unit restore from the card the customer already
|
|
opened. Its own task **deliberately**: it puts a DESTRUCTIVE operation (overwrites live data with
|
|
the backup state) behind a button reached via a NON-destructive one, so the confirm copy must carry
|
|
that difference.
|
|
- **C9-F4** — **nothing reads the Tier-2 copy's `recovery-unit/` mirror.** Written by every Tier-2 run
|
|
(`tier2.go:369`), read by no path: `RecoveryUnitPath` resolves to `backups/**primary**/`
|
|
(`appbackup/paths.go:46-48`), and the only reader of the secondary tree is `tier2_restore.go:79`.
|
|
Tier-2 exists for the case where the PRIMARY drive is lost — and in exactly that case the primary
|
|
unit is gone while this mirror survives on the second drive, unreachable by any customer action,
|
|
leaving offsite as the only route. **Potentially larger than C9-F1.**
|
|
|
|
## Tests
|
|
|
|
`go test ./...` **rc=0, 27 packages** — run and `rc` read *separately* from the commit. Six red-proofs
|
|
all observed, including the one that matters most: adding `StateRestarting` to `IsDownState` fails the
|
|
brief-restart test with *"every deploy and update would page the operator"*.
|
|
|
|
## Observations
|
|
|
|
- **A seventh shipped-invariant-comment.** `controller/README.md` stated that faults "still surface as
|
|
`exited`/`degraded`/**`restarting`**/`unhealthy`" — but `restarting` was in no down set at all. The
|
|
sentence was a wish with no test pinning it. Corrected in place, with the threshold rule documented
|
|
beside it.
|
|
- **Pre-existing gate failure, not mine:** `scripts/docker_run_volume_path_gate.py` fails on
|
|
`internal/appexport/estimate.go:179` (an unreviewed `docker run -v`). Verified it fails identically
|
|
on clean HEAD; left alone as out of scope rather than silently "fixed".
|
|
- The other six gates pass (`template_id`, `emoji`, `mojibake`, `native_confirm`, `app_row_dedup`,
|
|
`offbox_rename`).
|
|
|
|
## NOT yet live-validated — carried forward
|
|
|
|
- **Tier-1 content recovery after real loss** — still the most valuable unproven item (Campaign 9's A2
|
|
ran against an intact app; A3 used Tier-2). Unchanged by this work.
|
|
- The C9-F2 threshold under a **quiesce** cycle (Scenario C) is unit-proven but was not replayed live;
|
|
it needs a backup window plus an app that fails to come back.
|
|
- C9-F1's refusal for the other 42 class-B apps is proven by enumeration and by BookStack live, not
|
|
app-by-app.
|
|
- Host reboot mid-backup, three-way concurrency with GC, Scenario C live, `offsite_stale` firing,
|
|
F-HUB — all still open from Campaign 9.
|