C9-F1 + C9-F2: a restore that restored nothing, and a crash loop nobody saw (v0.183.0)
Both are the system reporting healthy while the customer is not, and both live in the same status-derivation code. Neither is fixed by making the system quieter. C9-F1 (HIGH) — Tier-2 writes recovery-unit/ on EVERY run and RestoreTier2Files has never read it (tier2_restore.go:101-104 reads hdd/ + userdata/ only). Phase 0 enumerated all 53 catalog templates against both demo boxes: 43 apps have NO readable subtree, so the button stopped the app, restored 0 files, restarted it and said "Nincs hiányzó fájl — minden fájl megvan a helyén." — at the moment the customer pressed it because files were missing, with 156 MB of BookStack's data unread in the same copy. 9 apps have file legs but never their DB or volumes, so the same sentence was also a clean bill of health over data never opened (immich: 1.3 GB Postgres unit). Honesty half shipped: a pre-flight coverage check refuses UP FRONT without stopping the app and NAMES the action that works; a run that proceeds claims only what it EXAMINED and discloses that the database and volumes are not covered. Completeness is filed as C9-F1b — routing to the Tier-1 unit restore puts a destructive operation behind a non-destructive button, so its confirm copy has to carry that difference. C9-F4 filed: nothing reads the Tier-2 recovery-unit/ mirror, so the second local copy that exists for drive loss is unreachable by any customer action. C9-F2 (HIGH) — a crash loop was counted as working. StateRestarting is deliberately NOT added to IsDownState (that alarms on every deploy fleet-wide, the over-correction F-A1 nearly cost us); a sustained run becomes down after crashLoopAfter = 5m, set above the 120s deploy timeout, Mealie's 60s start_period and R-97b's 180s grace. The dashboard counter uses the same predicate, so it no longer contradicts the alarm on the same screen. README's claim that faults "still surface as restarting" was a wish with no test — corrected in place; it is the seventh such instance. Six red-proofs observed, including the one that matters most: adding StateRestarting to IsDownState fails the brief-restart test with "every deploy and update would page the operator". go test ./... rc=0, 27 packages, run and read separately from this commit.
This commit is contained in:
+45
-5
@@ -853,10 +853,31 @@ customer edit after the last copy wins) and **nothing is ever deleted** — this
|
||||
`CrossDriveBackup.DestinationPath` (never a fresh target selection). Single-flight with
|
||||
backup/restore; refusals (no copy / never ran / copy dir gone / either drive disconnected /
|
||||
decommissioned) happen before the app is stopped, with customer-readable Hungarian reasons;
|
||||
stop → copy → start → health-wait; zero files copied is a success ("Nincs hiányzó fájl…"). Out of
|
||||
scope by design: overwrite/point-in-time restore (offbox + operator paths), per-file selection,
|
||||
`recovery-unit/`. Apps that index their data dir (e.g. Nextcloud) may need a rescan (occ
|
||||
files:scan) before restored files appear in their own UI.
|
||||
stop → copy → start → health-wait. Out of scope by design: overwrite/point-in-time restore (offbox +
|
||||
operator paths) and per-file selection. Apps that index their data dir (e.g. Nextcloud) may need a
|
||||
rescan (occ files:scan) before restored files appear in their own UI.
|
||||
|
||||
> **COVERAGE — read this before assuming an app is protected by this button (C9-F1, v0.183.0).**
|
||||
> This restore reads `hdd/` and `userdata/` **only**. It has never read `recovery-unit/`, which every
|
||||
> Tier-2 run also writes and which holds the app's DB dumps and named-volume tarballs. Enumerated
|
||||
> across all 53 catalog templates: **43 apps have no readable subtree at all** (their data is entirely
|
||||
> in named volumes — BookStack, Docmost, Vaultwarden, Gitea, …), **9** have file legs but never their
|
||||
> database or volumes, 1 is stateless. So the button is a guaranteed no-op for 81% of the catalog and
|
||||
> only ever partial for the rest.
|
||||
>
|
||||
> Since v0.183.0 it is HONEST about that instead of silently reporting success:
|
||||
> `Tier2RestoreCoverage` is consulted **before** anything starts, an app with no readable subtree is
|
||||
> refused **without being stopped** and told which action does work („…Használd a Visszaállítás
|
||||
> indítása gombot a Biztonsági mentés → Visszaállítás oldalon."), and a run that does proceed claims
|
||||
> only what it **examined** („Minden vizsgált fájl megvan a helyén.") plus a disclosure that the
|
||||
> database and internal volumes are not part of this restore.
|
||||
>
|
||||
> The action that DOES cover those apps is the keep-side recovery-unit restore
|
||||
> (`POST /backup/restore` → `RestoreFromRecoveryUnit`), which replays volume tarballs and DB dumps.
|
||||
> Routing customers there from the Tier-2 card is filed as **C9-F1b** — it puts a destructive
|
||||
> operation behind a button reached via a non-destructive one, so the confirm copy must carry that
|
||||
> difference. **C9-F4** is filed separately: nothing reads the Tier-2 copy's `recovery-unit/` mirror,
|
||||
> so the second local copy that exists precisely for drive loss is unreachable by any customer action.
|
||||
|
||||
**Per-app Tier-2 config panel (v0.57.0)** — `GET/POST /stacks/{name}/backup`
|
||||
(`internal/web/tier2_config_handler.go` + `templates/tier2_config.html`). The "2. mentés" row's
|
||||
@@ -1762,7 +1783,26 @@ display, not alarms, and are unchanged). This rests on two invariants: **I1**
|
||||
at `StateStopped` (compose down removes the containers); **I2** — the P2 restart-policy census
|
||||
(53 templates / 78 services, all `unless-stopped`) means a crashing app never comes to rest at
|
||||
`stopped`, so faults still surface as `exited`/`degraded`/`restarting`/`unhealthy`. If either
|
||||
invariant changes, revisit the suppression. `IsDownState` itself is deliberately UNCHANGED (other
|
||||
invariant changes, revisit the suppression.
|
||||
|
||||
> **C9-F2 (v0.183.0) — the `restarting` half of that sentence was a wish, not a fact.** `restarting`
|
||||
> was named above as a state through which faults "still surface", but it was in no down set at all:
|
||||
> `IsDownState` excludes it, so a crash-looping app raised no banner, no `app_start_failed`, no email
|
||||
> and no hub event — and `unless-stopped` means Docker retries forever, so the silence was permanent.
|
||||
> Campaign 9 watched docmost loop for nine minutes while the F-OBS heartbeat printed
|
||||
> „180 scans since boot, 4 deployed app(s) evaluated, **0 currently down**".
|
||||
>
|
||||
> The fix does **not** add `StateRestarting` to `IsDownState` — that alarms on every deploy and update
|
||||
> fleet-wide. A SUSTAINED restarting run becomes down after `stacks.crashLoopAfter` (**5 min**), chosen
|
||||
> above the deploy flow's 120 s health timeout, Mealie's 60 s `start_period` and R-97b's 180 s quiesce
|
||||
> grace, so the suppression windows compose into one bounded delay rather than leaving a gap. Carried
|
||||
> by `Stack.RestartingSince` (stamped in `refreshStatusLocked`, cleared on any other state, not
|
||||
> persisted) and read via `Stack.CrashLooping(now)` — used by BOTH the alarm and the dashboard
|
||||
> "how many of my apps work" counter, which previously counted `restarting` as running and so
|
||||
> contradicted the alarm on the same screen. Pinned by `crashloop_classify_test.go`; the test that a
|
||||
> brief restart stays silent is the one that fails against the naive fix.
|
||||
|
||||
`IsDownState` itself is deliberately UNCHANGED (other
|
||||
callers rely on stopped counting as down). An out-of-band `docker compose stop` leaves the containers
|
||||
present → `StateExited` → still alerts, which is correct (out-of-band tampering is reportable).
|
||||
|
||||
|
||||
@@ -1,6 +1,8 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"time"
|
||||
|
||||
"testing"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-controller/internal/notify"
|
||||
@@ -56,7 +58,7 @@ func TestClassifyRunStates_StoppedIsSuppressed(t *testing.T) {
|
||||
stack("nextcloud", stacks.StateDegraded, true, false),
|
||||
}
|
||||
|
||||
dead, states := classifyRunStates(sts, nil, nil)
|
||||
dead, states := classifyRunStates(sts, nil, nil, time.Now())
|
||||
|
||||
gotDead := deadNames(dead)
|
||||
if len(gotDead) != 2 || !gotDead["immich"] || !gotDead["nextcloud"] {
|
||||
@@ -90,7 +92,7 @@ func TestClassifyRunStates_FaultParity(t *testing.T) {
|
||||
stack("nextcloud", stacks.StateDegraded, true, false),
|
||||
}
|
||||
|
||||
dead, states := classifyRunStates(sts, nil, nil)
|
||||
dead, states := classifyRunStates(sts, nil, nil, time.Now())
|
||||
|
||||
gotDead := deadNames(dead)
|
||||
if len(gotDead) != 2 || !gotDead["immich"] || !gotDead["nextcloud"] {
|
||||
@@ -116,7 +118,7 @@ func TestClassifyRunStates_SkipsDeployingAndUndeployed(t *testing.T) {
|
||||
stack("mid", stacks.StateDeploying, true, true), // mid-deploy → skipped
|
||||
stack("gone", stacks.StateExited, false, false), // not deployed → skipped
|
||||
}
|
||||
dead, states := classifyRunStates(sts, nil, nil)
|
||||
dead, states := classifyRunStates(sts, nil, nil, time.Now())
|
||||
if len(dead) != 0 || len(states) != 0 {
|
||||
t.Fatalf("deploying and undeployed stacks must be skipped, got dead=%+v states=%+v", dead, states)
|
||||
}
|
||||
|
||||
@@ -0,0 +1,139 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-controller/internal/stacks"
|
||||
)
|
||||
|
||||
// C9-F2 — a SUSTAINED `restarting` is a crash loop and must alarm; a BRIEF one must not.
|
||||
//
|
||||
// The defect: `IsDownState` excludes `restarting` as "self-recovering", but for the catalog's
|
||||
// standard `restart: unless-stopped` Docker retries forever, so a crash loop sat in `restarting`
|
||||
// indefinitely and was counted as working. Campaign 9 watched docmost loop for nine minutes
|
||||
// (restartcount 18) while the F-OBS heartbeat printed "4 deployed app(s) evaluated, 0 currently down".
|
||||
//
|
||||
// The whole design tension is that B must keep passing while A does: an alarm that fires on every
|
||||
// deploy is one the operator learns to ignore.
|
||||
|
||||
// restartingSince builds a deployed stack that has been restarting since `since`.
|
||||
func restartingSince(name string, since time.Time) stacks.Stack {
|
||||
s := stack(name, stacks.StateRestarting, true, false)
|
||||
s.RestartingSince = since
|
||||
return s
|
||||
}
|
||||
|
||||
// SCENARIO A — a crash loop alarms. A stack restarting for longer than the threshold enters BOTH the
|
||||
// banner dead-list and the notifier Down-set, so app_start_failed can fire.
|
||||
//
|
||||
// RED-PROOF (observed): drop `|| crashLooping` from the `down` expression in classifyRunStates →
|
||||
//
|
||||
// crashloop_classify_test.go:52: docmost is NOT in the Down-set — a crash loop is silent (this is C9-F2)
|
||||
// crashloop_classify_test.go:55: docmost is NOT in the banner dead-list
|
||||
func TestClassifyRunStates_SustainedRestartingAlarms(t *testing.T) {
|
||||
now := time.Now()
|
||||
sts := []stacks.Stack{
|
||||
stack("paperless-ngx", stacks.StateRunning, true, false),
|
||||
restartingSince("docmost", now.Add(-9*time.Minute)), // the Campaign 9 observation, exactly
|
||||
}
|
||||
|
||||
dead, states := classifyRunStates(sts, nil, nil, now)
|
||||
|
||||
if !downByName(states)["docmost"] {
|
||||
t.Errorf("docmost is NOT in the Down-set — a crash loop is silent (this is C9-F2)")
|
||||
}
|
||||
if !deadNames(dead)["docmost"] {
|
||||
t.Errorf("docmost is NOT in the banner dead-list")
|
||||
}
|
||||
if downByName(states)["paperless-ngx"] {
|
||||
t.Errorf("a healthy app was dragged down with it")
|
||||
}
|
||||
}
|
||||
|
||||
// SCENARIO B — a normal deploy or update does NOT alarm. `docker compose up -d` passes through
|
||||
// restarting; alarming there would page the operator on every routine operation, fleet-wide.
|
||||
//
|
||||
// This is the test that must fail against the naive fix. RED-PROOF (observed): add StateRestarting
|
||||
// to IsDownState instead of using the threshold →
|
||||
//
|
||||
// crashloop_classify_test.go:78: a BRIEFLY restarting app alarms — every deploy and update would page the operator
|
||||
func TestClassifyRunStates_BriefRestartingIsSilent(t *testing.T) {
|
||||
now := time.Now()
|
||||
sts := []stacks.Stack{
|
||||
restartingSince("mealie", now.Add(-30*time.Second)), // mid-deploy
|
||||
restartingSince("ghost", now.Add(-2*time.Minute)), // slow image pull, still normal
|
||||
}
|
||||
|
||||
dead, states := classifyRunStates(sts, nil, nil, now)
|
||||
|
||||
for _, name := range []string{"mealie", "ghost"} {
|
||||
if downByName(states)[name] {
|
||||
t.Errorf("a BRIEFLY restarting app alarms (%s) — every deploy and update would page the operator", name)
|
||||
}
|
||||
}
|
||||
if len(dead) != 0 {
|
||||
t.Errorf("banner dead-list should be empty during normal restarts, got %v", deadNames(dead))
|
||||
}
|
||||
}
|
||||
|
||||
// The boundary itself, asserted from both sides so the threshold cannot drift silently.
|
||||
func TestCrashLooping_ThresholdBoundary(t *testing.T) {
|
||||
now := time.Now()
|
||||
for _, tc := range []struct {
|
||||
name string
|
||||
age time.Duration
|
||||
want bool
|
||||
}{
|
||||
{"just under the threshold", 4*time.Minute + 59*time.Second, false},
|
||||
{"exactly at the threshold", 5 * time.Minute, true},
|
||||
{"well past it", 30 * time.Minute, true},
|
||||
} {
|
||||
s := restartingSince("app", now.Add(-tc.age))
|
||||
if got := s.CrashLooping(now); got != tc.want {
|
||||
t.Errorf("%s: CrashLooping(age=%s) = %v, want %v", tc.name, tc.age, got, tc.want)
|
||||
}
|
||||
}
|
||||
|
||||
// A stack that is not restarting is never a crash loop, however old the stamp.
|
||||
s := stack("app", stacks.StateRunning, true, false)
|
||||
s.RestartingSince = now.Add(-time.Hour)
|
||||
if s.CrashLooping(now) {
|
||||
t.Error("a RUNNING stack reported as crash-looping — the state test is missing")
|
||||
}
|
||||
|
||||
// A zero stamp is "not yet observed restarting", never a crash loop — this is what makes the
|
||||
// first scan after a controller restart silent instead of alarming on everything at once.
|
||||
z := stack("app", stacks.StateRestarting, true, false)
|
||||
if z.CrashLooping(now) {
|
||||
t.Error("a zero RestartingSince reported as crash-looping — a controller restart would alarm fleet-wide")
|
||||
}
|
||||
}
|
||||
|
||||
// SCENARIO C — R-97b's quiesce suppression still wins inside its window. A stack the backup stopped
|
||||
// and is restarting must stay silent while suppressed, even if its restarting run is old enough to
|
||||
// qualify. The window EXPIRES, so a genuinely dead app still alarms afterwards — proven by the
|
||||
// second half of this test.
|
||||
//
|
||||
// RED-PROOF (observed): drop `&& !quiesced[st.Name]` from the `down` expression →
|
||||
//
|
||||
// crashloop_classify_test.go:129: a quiesced stack alarms — every backup would page the customer
|
||||
func TestClassifyRunStates_QuiesceSuppressionBeatsCrashLoop(t *testing.T) {
|
||||
now := time.Now()
|
||||
sts := []stacks.Stack{restartingSince("docmost", now.Add(-9*time.Minute))}
|
||||
|
||||
// Inside the R-97b window.
|
||||
_, states := classifyRunStates(sts, map[string]bool{"docmost": true}, nil, now)
|
||||
if downByName(states)["docmost"] {
|
||||
t.Errorf("a quiesced stack alarms — every backup would page the customer")
|
||||
}
|
||||
|
||||
// Window expired (the stack is no longer reported as suppressed): the same stack must now alarm.
|
||||
dead, states := classifyRunStates(sts, nil, nil, now)
|
||||
if !downByName(states)["docmost"] {
|
||||
t.Errorf("suppression outlived its window — a genuinely dead app stayed silent (R-97b's own warning)")
|
||||
}
|
||||
if !deadNames(dead)["docmost"] {
|
||||
t.Errorf("suppression outlived its window for the banner too")
|
||||
}
|
||||
}
|
||||
@@ -1,6 +1,8 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"time"
|
||||
|
||||
"testing"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-controller/internal/stacks"
|
||||
@@ -27,7 +29,7 @@ func TestClassifyRunStates_FailedRestartAlarmsDespiteStateStopped(t *testing.T)
|
||||
}
|
||||
failed := map[string]bool{"immich": true}
|
||||
|
||||
dead, states := classifyRunStates(sts, nil, failed)
|
||||
dead, states := classifyRunStates(sts, nil, failed, time.Now())
|
||||
|
||||
if !downByName(states)["immich"] {
|
||||
t.Error("a stack that FAILED to restart is silent (Down=false) — this is F-CRIT-1")
|
||||
@@ -54,7 +56,7 @@ func TestClassifyRunStates_UserStopStillSilent(t *testing.T) {
|
||||
// only immich failed to restart; cwa was never touched by a quiesce
|
||||
failed := map[string]bool{"immich": true}
|
||||
|
||||
dead, states := classifyRunStates(sts, nil, failed)
|
||||
dead, states := classifyRunStates(sts, nil, failed, time.Now())
|
||||
down := downByName(states)
|
||||
|
||||
if down["cwa"] || deadNames(dead)["cwa"] {
|
||||
@@ -75,7 +77,7 @@ func TestClassifyRunStates_NoFailedRestartsIsV0164Behaviour(t *testing.T) {
|
||||
stack("nextcloud", stacks.StateDegraded, true, false),
|
||||
}
|
||||
|
||||
dead, states := classifyRunStates(sts, nil, nil)
|
||||
dead, states := classifyRunStates(sts, nil, nil, time.Now())
|
||||
down := downByName(states)
|
||||
|
||||
if down["cwa"] {
|
||||
@@ -99,7 +101,7 @@ func TestClassifyRunStates_GraceWindowStillSuppresses(t *testing.T) {
|
||||
quiesced := map[string]bool{"immich": true} // still inside quiesceAlarmGrace
|
||||
failed := map[string]bool{"immich": true} // and we already know the restart failed
|
||||
|
||||
dead, states := classifyRunStates(sts, quiesced, failed)
|
||||
dead, states := classifyRunStates(sts, quiesced, failed, time.Now())
|
||||
|
||||
if downByName(states)["immich"] {
|
||||
t.Error("alarmed while still inside the grace window — R-97b Scenario E broken")
|
||||
@@ -115,7 +117,7 @@ func TestClassifyRunStates_UndeployedIgnored(t *testing.T) {
|
||||
stack("ghost", stacks.StateStopped, false, false),
|
||||
stack("deploying", stacks.StateStopped, true, true),
|
||||
}
|
||||
dead, states := classifyRunStates(sts, nil, map[string]bool{"ghost": true, "deploying": true})
|
||||
dead, states := classifyRunStates(sts, nil, map[string]bool{"ghost": true, "deploying": true}, time.Now())
|
||||
if len(dead) != 0 || len(states) != 0 {
|
||||
t.Errorf("undeployed/deploying stacks were classified: dead=%v states=%v", deadNames(dead), states)
|
||||
}
|
||||
|
||||
@@ -1242,7 +1242,7 @@ func runBootReconcile(ctx context.Context, mgr bootrecon.StackProvider, logger *
|
||||
func scanDeployedAppRunStates(mgr *stacks.Manager, q *quiesce.Loop) ([]web.DeadApp, []notify.AppRunState) {
|
||||
// R-97b: a stack THIS controller stopped for a backup is not a fault. q may be nil (unprovisioned
|
||||
// guest) — SuppressedStacks is nil-safe and returns nothing, i.e. suppress nothing.
|
||||
return classifyRunStates(mgr.GetStacks(), q.SuppressedStacks(), q.FailedRestarts())
|
||||
return classifyRunStates(mgr.GetStacks(), q.SuppressedStacks(), q.FailedRestarts(), time.Now())
|
||||
}
|
||||
|
||||
// classifyRunStates is the pure fix-3 derivation over a plain stack slice. It splits the deployed
|
||||
@@ -1274,7 +1274,9 @@ func scanDeployedAppRunStates(mgr *stacks.Manager, q *quiesce.Loop) ([]web.DeadA
|
||||
// which is correct: out-of-band tampering IS reportable.) IsDownState is intentionally left unchanged
|
||||
// — other callers rely on stopped counting as down; the suppression is a filter at this single
|
||||
// derivation point only.
|
||||
func classifyRunStates(sts []stacks.Stack, quiesced map[string]bool, failedRestart map[string]bool) ([]web.DeadApp, []notify.AppRunState) {
|
||||
// `now` is injected (C9-F2) so the crash-loop threshold is a unit-testable contract rather than a
|
||||
// property of the wall clock.
|
||||
func classifyRunStates(sts []stacks.Stack, quiesced map[string]bool, failedRestart map[string]bool, now time.Time) ([]web.DeadApp, []notify.AppRunState) {
|
||||
var dead []web.DeadApp
|
||||
var states []notify.AppRunState
|
||||
for _, st := range sts {
|
||||
@@ -1289,8 +1291,15 @@ func classifyRunStates(sts []stacks.Stack, quiesced map[string]bool, failedResta
|
||||
// F-CRIT-1: StateStopped is whitelisted as a deliberate user stop UNLESS the quiesce loop
|
||||
// reports that it stopped this stack and could not restart it. That single term is what turns
|
||||
// an indefinitely-silent dead app back into an alarm, without re-alarming genuine user stops.
|
||||
// C9-F2: a SUSTAINED restarting is a crash loop, and a crash loop is a dead app. Deliberately
|
||||
// NOT folded into IsDownState — that would alarm on every deploy and update fleet-wide, which
|
||||
// is the over-correction F-A1 nearly cost us. The threshold (stacks.crashLoopAfter, 5 min) sits
|
||||
// above the deploy health timeout, the slowest catalog start_period AND R-97b's grace, so a
|
||||
// brief restart never reaches it and the two suppression windows compose into one bounded
|
||||
// delay. Quiesce suppression below still wins inside its own window.
|
||||
crashLooping := st.CrashLooping(now)
|
||||
userStopped := st.State == stacks.StateStopped && !failedRestart[st.Name]
|
||||
down := stacks.IsDownState(st.State) && !userStopped && !quiesced[st.Name]
|
||||
down := (stacks.IsDownState(st.State) || crashLooping) && !userStopped && !quiesced[st.Name]
|
||||
states = append(states, notify.AppRunState{Name: st.Name, DisplayName: st.Meta.DisplayName, Down: down})
|
||||
if down {
|
||||
dead = append(dead, web.DeadApp{Name: st.Name, DisplayName: st.Meta.DisplayName, State: string(st.State)})
|
||||
|
||||
@@ -0,0 +1,143 @@
|
||||
package backup
|
||||
|
||||
import (
|
||||
"errors"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// C9-F1 — the Tier-2 restore reads `hdd/` and `userdata/` only. `recovery-unit/` — the app's DB dumps
|
||||
// and named-volume tarballs — is written by EVERY Tier-2 run and read by NOTHING on this path.
|
||||
//
|
||||
// For 43 of the 53 catalog apps (BookStack, Docmost, Vaultwarden, Gitea, …) that is the app's ENTIRE
|
||||
// dataset, so the restore was a guaranteed no-op that still took an outage and reported
|
||||
// „Nincs hiányzó fájl — minden fájl megvan a helyén."
|
||||
//
|
||||
// These tests pin the asymmetry itself, so a future change that alters what the restore reads must
|
||||
// either keep the coverage answer honest or fail here.
|
||||
|
||||
// unitOnlyCopy rewrites the fixture's copy into the BookStack shape: a recovery unit and nothing the
|
||||
// restore can read. Mirrors the live demo-felhom layout observed in Campaign 9
|
||||
// (`legs=[NONE] unit=156M`).
|
||||
func unitOnlyCopy(t *testing.T, destDrive string) string {
|
||||
t.Helper()
|
||||
destBase := filepath.Join(destDrive, "backups", "secondary", "app")
|
||||
if err := os.RemoveAll(filepath.Join(destBase, "hdd")); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
mustWrite(t, filepath.Join(destBase, "recovery-unit", "volume-dumps", "app_db_data.tar"), "TARBYTES")
|
||||
mustWrite(t, filepath.Join(destBase, "recovery-unit", "db-dumps", "app.sql"), "SQLDUMP")
|
||||
return destBase
|
||||
}
|
||||
|
||||
// SCENARIO D — a restore that cannot cover an app refuses BEFORE the outage.
|
||||
//
|
||||
// RED-PROOF (observed): remove the `!cov.CanRestore()` guard from RestoreTier2Files →
|
||||
//
|
||||
// tier2_coverage_test.go:63: RestoreTier2Files returned <nil> — a copy with nothing restorable was treated as success
|
||||
// tier2_coverage_test.go:69: THE APP WAS STOPPED for a restore that could not restore anything: [app]
|
||||
func TestRestoreTier2Files_NoRestorableSubtree_RefusesBeforeStopping(t *testing.T) {
|
||||
m, fake, _, destDrive := newT2RManager(t)
|
||||
unitOnlyCopy(t, destDrive)
|
||||
m.restoreFilesCopier = func(string, string) (int, error) {
|
||||
t.Fatal("the copier ran for an app with no restorable subtree")
|
||||
return 0, nil
|
||||
}
|
||||
|
||||
n, err := m.RestoreTier2Files("app")
|
||||
|
||||
if !errors.Is(err, ErrTier2NoRestorableData) {
|
||||
t.Errorf("RestoreTier2Files returned %v — a copy with nothing restorable was treated as success", err)
|
||||
}
|
||||
if n != 0 {
|
||||
t.Errorf("filesRestored = %d, want 0", n)
|
||||
}
|
||||
// The whole point: no outage was taken.
|
||||
if len(fake.stopped) != 0 {
|
||||
t.Errorf("THE APP WAS STOPPED for a restore that could not restore anything: %v", fake.stopped)
|
||||
}
|
||||
if len(fake.started) != 0 {
|
||||
t.Errorf("the app was restarted, so it must have been stopped: %v", fake.started)
|
||||
}
|
||||
}
|
||||
|
||||
// The coverage query itself — what the handler pre-flights on, so it can refuse without starting an
|
||||
// operation at all.
|
||||
func TestTier2RestoreCoverage_ReportsTheAsymmetry(t *testing.T) {
|
||||
m, _, _, destDrive := newT2RManager(t)
|
||||
|
||||
// Class A (paperless/immich shape): an hdd leg the restore reads, plus a unit it does not.
|
||||
mustWrite(t, filepath.Join(destDrive, "backups", "secondary", "app", "recovery-unit", "manifest.json"), "{}")
|
||||
cov, err := m.Tier2RestoreCoverage("app")
|
||||
if err != nil {
|
||||
t.Fatalf("coverage: %v", err)
|
||||
}
|
||||
if !cov.CanRestore() {
|
||||
t.Error("an app WITH an hdd leg reported as not restorable — this would refuse the one path that works")
|
||||
}
|
||||
if !cov.HasUnit {
|
||||
t.Error("the recovery unit was not detected — the disclosure would be omitted")
|
||||
}
|
||||
|
||||
// Class B (bookstack/docmost shape): unit only.
|
||||
unitOnlyCopy(t, destDrive)
|
||||
cov, err = m.Tier2RestoreCoverage("app")
|
||||
if err != nil {
|
||||
t.Fatalf("coverage: %v", err)
|
||||
}
|
||||
if cov.CanRestore() {
|
||||
t.Error("a unit-only copy reported as restorable — this is exactly C9-F1")
|
||||
}
|
||||
if !cov.HasUnit {
|
||||
t.Error("the unit that holds the app's whole dataset was not detected")
|
||||
}
|
||||
if len(cov.Legs) != 0 {
|
||||
t.Errorf("legs = %v, want none", cov.Legs)
|
||||
}
|
||||
}
|
||||
|
||||
// SCENARIO E — an app the restore CAN cover is completely unchanged. This is the regression guard on
|
||||
// Campaign 9's headline result (A1/A3, paperless-ngx): byte-identical restore, stop→copy→start, and
|
||||
// the additive-only promises intact. Breaking this to fix BookStack would be a straight regression on
|
||||
// the only restore path proven to work on live hardware.
|
||||
//
|
||||
// RED-PROOF (observed): make the coverage guard unconditional (`if true`) →
|
||||
//
|
||||
// tier2_coverage_test.go:118: a COVERED app was refused: ennek az alkalmazásnak az adatai nem ebből a másolatból állíthatók vissza
|
||||
func TestRestoreTier2Files_CoveredAppIsUnchanged(t *testing.T) {
|
||||
m, fake, liveDrive, destDrive := newT2RManager(t)
|
||||
// A unit is present too — a covered app has one as well; it must not change the outcome.
|
||||
mustWrite(t, filepath.Join(destDrive, "backups", "secondary", "app", "recovery-unit", "manifest.json"), "{}")
|
||||
|
||||
var copied [][2]string
|
||||
m.restoreFilesCopier = func(src, dst string) (int, error) {
|
||||
copied = append(copied, [2]string{src, dst})
|
||||
fake.order = append(fake.order, "copy")
|
||||
return 3, nil
|
||||
}
|
||||
|
||||
n, err := m.RestoreTier2Files("app")
|
||||
if err != nil {
|
||||
t.Fatalf("a COVERED app was refused: %v", err)
|
||||
}
|
||||
if n != 3 {
|
||||
t.Errorf("filesRestored = %d, want 3", n)
|
||||
}
|
||||
if len(fake.stopped) != 1 || len(fake.started) != 1 {
|
||||
t.Errorf("stop/start did not happen exactly once: %v / %v", fake.stopped, fake.started)
|
||||
}
|
||||
if got := fake.order; len(got) < 3 || got[0] != "stop" || got[len(got)-1] != "start" {
|
||||
t.Errorf("order = %v, want stop → copy → start", got)
|
||||
}
|
||||
if len(copied) == 0 {
|
||||
t.Fatal("nothing was copied for a covered app")
|
||||
}
|
||||
wantSrc := filepath.Join(destDrive, "backups", "secondary", "app", "hdd")
|
||||
if copied[0][0] != wantSrc {
|
||||
t.Errorf("src = %q, want %q", copied[0][0], wantSrc)
|
||||
}
|
||||
if copied[0][1] != liveDrive {
|
||||
t.Errorf("dst = %q, want the live namespace root %q", copied[0][1], liveDrive)
|
||||
}
|
||||
}
|
||||
@@ -33,8 +33,81 @@ var (
|
||||
// marker). Refuse rather than read a flat layout we no longer understand — safe, because tier-2
|
||||
// restore is missing-file recovery and the live data still exists in that scenario.
|
||||
errTier2OldLayout = errors.New("A 2. mentés régi formátumú — futtass előbb egy új másodlagos mentést.")
|
||||
// ErrTier2NoRestorableData (C9-F1) — this app HAS a Tier-2 copy, but that copy contains no subtree
|
||||
// this restore can read: its data lives entirely in Docker named volumes, which are captured into
|
||||
// recovery-unit/ (db-dumps + volume-dumps) and NEVER read by this path. 43 of the 53 catalog apps
|
||||
// are in this class. Exported so the handler can refuse BEFORE stopping the app and name the action
|
||||
// that does work, instead of taking an outage and reporting "no missing files".
|
||||
ErrTier2NoRestorableData = errors.New("ennek az alkalmazásnak az adatai nem ebből a másolatból állíthatók vissza")
|
||||
)
|
||||
|
||||
// Tier2Coverage says what a Tier-2 restore can and cannot return for one app — the asymmetry C9-F1
|
||||
// is about. Computed from the RECORDED copy on disk, never guessed from the catalog, so an app whose
|
||||
// template changed is judged by what its actual copy holds.
|
||||
//
|
||||
// The distinction that matters: Legs are the subtrees RestoreTier2Files reads (hdd/, userdata/);
|
||||
// HasUnit means the copy ALSO holds a full recovery unit — the app's database dumps and named-volume
|
||||
// tarballs — which this restore path never opens. An app can have HasUnit && no Legs (43 of 53), in
|
||||
// which case the restore is a guaranteed no-op no matter how much data was lost.
|
||||
type Tier2Coverage struct {
|
||||
Legs []string // subtrees this restore reads and that exist in the copy: "hdd", "userdata"
|
||||
HasUnit bool // recovery-unit/ present — captured, but NOT restorable by this path
|
||||
}
|
||||
|
||||
// CanRestore reports whether the restore has any subtree to read at all.
|
||||
func (c Tier2Coverage) CanRestore() bool { return len(c.Legs) > 0 }
|
||||
|
||||
// tier2CoverageAt inspects a resolved copy directory. Pure filesystem stat — no side effects.
|
||||
func tier2CoverageAt(destBase string) Tier2Coverage {
|
||||
var c Tier2Coverage
|
||||
for _, leg := range []string{"hdd", "userdata"} {
|
||||
if fi, err := os.Stat(filepath.Join(destBase, leg)); err == nil && fi.IsDir() {
|
||||
c.Legs = append(c.Legs, leg)
|
||||
}
|
||||
}
|
||||
if fi, err := os.Stat(filepath.Join(destBase, "recovery-unit")); err == nil && fi.IsDir() {
|
||||
c.HasUnit = true
|
||||
}
|
||||
return c
|
||||
}
|
||||
|
||||
// Tier2RestoreCoverage resolves the app's RECORDED Tier-2 copy and reports what a restore could
|
||||
// return from it. Errors are the same refusals RestoreTier2Files itself would raise, so the caller
|
||||
// can surface them before starting anything — this is what lets the handler refuse without an outage.
|
||||
func (m *Manager) Tier2RestoreCoverage(stackName string) (Tier2Coverage, error) {
|
||||
destBase, err := m.tier2RecordedCopyDir(stackName)
|
||||
if err != nil {
|
||||
return Tier2Coverage{}, err
|
||||
}
|
||||
return tier2CoverageAt(destBase), nil
|
||||
}
|
||||
|
||||
// tier2RecordedCopyDir resolves the RECORDED Tier-2 copy dir for a stack, applying every
|
||||
// source-side refusal in one place so the pre-flight check and the restore itself cannot drift.
|
||||
func (m *Manager) tier2RecordedCopyDir(stackName string) (string, error) {
|
||||
var destBase string
|
||||
if m.settings != nil {
|
||||
if cfg := m.settings.GetCrossDriveConfig(stackName); cfg != nil && cfg.LastRun != "" && cfg.DestinationPath != "" {
|
||||
if m.settings.IsDisconnected(cfg.DestinationPath) {
|
||||
return "", errTier2DriveGone
|
||||
}
|
||||
destBase = filepath.Join(cfg.DestinationPath, "backups", "secondary", stackName)
|
||||
}
|
||||
}
|
||||
if destBase == "" {
|
||||
return "", errNoTier2Copy
|
||||
}
|
||||
if _, statErr := os.Stat(destBase); statErr != nil {
|
||||
return "", errNoTier2Copy // recorded but the copy dir is gone — same honest refusal
|
||||
}
|
||||
// §7-G2 marker gate: a pre-v2 (flat) copy has no marker → refuse rather than read a layout we no
|
||||
// longer understand (live data still exists for missing-file recovery).
|
||||
if _, mErr := os.Stat(filepath.Join(destBase, tier2LayoutMarker)); mErr != nil {
|
||||
return "", errTier2OldLayout
|
||||
}
|
||||
return destBase, nil
|
||||
}
|
||||
|
||||
// RestoreTier2Files restores the app's MISSING user files in place from its recorded Tier-2 copy
|
||||
// (additive-only; see the package comment above). Returns how many regular files were copied back.
|
||||
//
|
||||
@@ -70,25 +143,21 @@ func (m *Manager) RestoreTier2Files(stackName string) (filesRestored int, err er
|
||||
liveNsRoot := m.namespaceRoot(drive)
|
||||
|
||||
// Source side: the RECORDED Tier-2 copy must exist, its drive connected, and it must be v2.
|
||||
var destBase string
|
||||
if m.settings != nil {
|
||||
if cfg := m.settings.GetCrossDriveConfig(stackName); cfg != nil && cfg.LastRun != "" && cfg.DestinationPath != "" {
|
||||
if m.settings.IsDisconnected(cfg.DestinationPath) {
|
||||
return 0, errTier2DriveGone
|
||||
}
|
||||
destBase = filepath.Join(cfg.DestinationPath, "backups", "secondary", stackName)
|
||||
}
|
||||
destBase, err := m.tier2RecordedCopyDir(stackName)
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
if destBase == "" {
|
||||
return 0, errNoTier2Copy
|
||||
}
|
||||
if _, statErr := os.Stat(destBase); statErr != nil {
|
||||
return 0, errNoTier2Copy // recorded but the copy dir is gone — same honest refusal
|
||||
}
|
||||
// §7-G2 marker gate: a pre-v2 (flat) copy has no marker → refuse rather than read a layout we no
|
||||
// longer understand (live data still exists for missing-file recovery).
|
||||
if _, mErr := os.Stat(filepath.Join(destBase, tier2LayoutMarker)); mErr != nil {
|
||||
return 0, errTier2OldLayout
|
||||
|
||||
// C9-F1: refuse BEFORE the app is stopped if this copy holds nothing this path can read. Without
|
||||
// this the app was stopped, zero files were copied, it was restarted, and the customer was told
|
||||
// "Nincs hiányzó fájl — minden fájl megvan a helyén." — an outage plus a claim about data the
|
||||
// restore never looked at. Placed with the other source-side refusals, all of which precede the
|
||||
// stop, so the promise "all refusals happen BEFORE the app is stopped" stays true.
|
||||
cov := tier2CoverageAt(destBase)
|
||||
if !cov.CanRestore() {
|
||||
m.logger.Printf("[WARN] [backup] Tier-2 file restore refused for %s: the recorded copy has no restorable subtree (unit_present=%v) — the app was NOT stopped",
|
||||
stackName, cov.HasUnit)
|
||||
return 0, ErrTier2NoRestorableData
|
||||
}
|
||||
|
||||
copier := m.restoreFilesCopier
|
||||
|
||||
@@ -55,6 +55,54 @@ func IsDownState(s ContainerState) bool {
|
||||
return s == StateStopped || s == StateExited || s == StateDegraded
|
||||
}
|
||||
|
||||
// C9-F2 — a SUSTAINED `restarting` is a crash loop, and a crash loop is a dead app.
|
||||
//
|
||||
// THE BUG THIS EXISTS TO KILL. `IsDownState` above excludes `restarting` as "self-recovering", and
|
||||
// for a brief restart that is exactly right. But Docker sets `restarting` while a container is being
|
||||
// restarted BY POLICY, and for the catalog's standard `restart: unless-stopped` that is precisely the
|
||||
// crash-loop signal — the retry count is unlimited, so "self-recovering" is a promise Docker never
|
||||
// made. Campaign 9 watched docmost loop for nine minutes (restartcount 18, policy `unless-stopped`)
|
||||
// while the F-OBS heartbeat printed "180 scans since boot, 4 deployed app(s) evaluated, 0 currently
|
||||
// down". No banner, no app_start_failed, no email, no hub event — indefinitely.
|
||||
//
|
||||
// This is CONTEXT.md's own lesson one state over: "Docker's .State says 'running' even for unhealthy
|
||||
// containers — must parse .Status". Same trap, different state, and this state means something worse.
|
||||
//
|
||||
// ── WHY A THRESHOLD AND NOT A DOWN-STATE ─────────────────────────────────────────────────────
|
||||
//
|
||||
// Adding StateRestarting to IsDownState would alarm on every deploy and every update, fleet-wide,
|
||||
// because the normal `docker compose up -d` path passes through `restarting`. An alarm that fires on
|
||||
// routine operations is one the operator learns to ignore — which is what F-A1 nearly cost us right
|
||||
// after R-97a built it. So `restarting` becomes down only once it has PERSISTED.
|
||||
//
|
||||
// ── WHERE 5 MINUTES COMES FROM ───────────────────────────────────────────────────────────────
|
||||
//
|
||||
// Measured against the three real numbers already in this codebase, not picked round:
|
||||
// - the deploy flow allows **120 s** for a stack to come up healthy — the project's own existing
|
||||
// answer to "how long is too long"; an app still restarting past it has failed deployment;
|
||||
// - the slowest catalog healthcheck start_period is Mealie's **60 s**, after which a couple of
|
||||
// check intervals must still elapse before any verdict is meaningful;
|
||||
// - R-97b's quiesce grace is **180 s**, and this must sit ABOVE it so the two windows compose into
|
||||
// one bounded delay rather than a gap where an app is un-suppressed but not yet sustained.
|
||||
//
|
||||
// 300 s clears all three with margin. It is also unambiguous against Docker's own backoff, which
|
||||
// grows 100 ms → 200 ms → … and caps at 60 s: a genuine crash loop registers at least four restart
|
||||
// attempts inside this window, so a stack that is still `restarting` at 5 minutes is not mid-deploy.
|
||||
//
|
||||
// The cost is a bounded DELAY in reporting a real crash loop, never its loss — the same trade R-97b
|
||||
// made deliberately, and the opposite of the indefinite silence this replaces.
|
||||
const crashLoopAfter = 5 * time.Minute
|
||||
|
||||
// CrashLooping reports whether the stack has been `restarting` for longer than crashLoopAfter.
|
||||
// `now` is injected so the rule is a unit-testable contract rather than a property of the clock.
|
||||
// A zero RestartingSince means "not restarting, or not yet observed restarting" — never a crash loop.
|
||||
func (s *Stack) CrashLooping(now time.Time) bool {
|
||||
if s == nil || s.State != StateRestarting || s.RestartingSince.IsZero() {
|
||||
return false
|
||||
}
|
||||
return now.Sub(s.RestartingSince) >= crashLoopAfter
|
||||
}
|
||||
|
||||
// ContainerInfo holds status info about a single container within a stack.
|
||||
type ContainerInfo struct {
|
||||
Name string `json:"name"`
|
||||
@@ -95,6 +143,13 @@ type Stack struct {
|
||||
DeployError string `json:"deploy_error,omitempty"` // last async deploy error
|
||||
HealthProbe *HealthProbeResult `json:"health_probe,omitempty"` // controller-side probe result
|
||||
LastUpdated time.Time `json:"last_updated"`
|
||||
// RestartingSince (C9-F2) is when this stack was FIRST observed in StateRestarting during the
|
||||
// current restarting run; zero whenever the stack is in any other state. It is what turns a brief
|
||||
// restart (normal: deploy, update, quiesce restart) into a distinguishable crash loop — see
|
||||
// CrashLooping. Not persisted: a controller restart re-observes the state within one refresh, and
|
||||
// forgetting costs at most one threshold window, whereas persisting could carry a stale
|
||||
// "this app is crash-looping" verdict across the restart that fixed it.
|
||||
RestartingSince time.Time `json:"restarting_since,omitempty"`
|
||||
}
|
||||
|
||||
// Manager handles all docker compose stack operations.
|
||||
@@ -587,6 +642,18 @@ func (m *Manager) refreshStatusLocked() error {
|
||||
stack.State = StateUnhealthy
|
||||
}
|
||||
|
||||
// C9-F2: stamp the start of a restarting RUN, and clear it the moment the stack is anything
|
||||
// else. Set AFTER the health-probe override above so the stamp always agrees with the state
|
||||
// that is actually stored. Clearing on any other state is what keeps a normal deploy — which
|
||||
// passes through restarting briefly — from ever accumulating toward the threshold.
|
||||
if stack.State == StateRestarting {
|
||||
if stack.RestartingSince.IsZero() {
|
||||
stack.RestartingSince = time.Now()
|
||||
}
|
||||
} else {
|
||||
stack.RestartingSince = time.Time{}
|
||||
}
|
||||
|
||||
if m.isDebug() {
|
||||
m.logger.Printf("[TRACE] [stacks] refreshStatusLocked: stack %q → state=%s containers=%d", name, stack.State, len(stack.Containers))
|
||||
}
|
||||
|
||||
@@ -3,6 +3,7 @@ package web
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"errors"
|
||||
"fmt"
|
||||
"log"
|
||||
"net/http"
|
||||
@@ -145,7 +146,17 @@ func (s *Server) dashboardHandler(w http.ResponseWriter, r *http.Request) {
|
||||
|
||||
// Count from the DISPLAYED set only
|
||||
running, stopped := 0, 0
|
||||
countNow := time.Now()
|
||||
for _, st := range deployedStacks {
|
||||
// C9-F2: a stack that has been `restarting` past the crash-loop threshold counts with STOPPED,
|
||||
// for the same reason R-51 moved `degraded` there — this counter answers "how many of my apps
|
||||
// work", and an app Docker has been restarting for five minutes does not. A BRIEF restart
|
||||
// still counts as running (deploys and updates pass through it), so the counter and the
|
||||
// dead-app alarm now agree instead of contradicting each other on the same screen.
|
||||
if st.CrashLooping(countNow) {
|
||||
stopped++
|
||||
continue
|
||||
}
|
||||
switch st.State {
|
||||
case stacks.StateRunning, stacks.StateStarting, stacks.StateUnhealthy, stacks.StateRestarting:
|
||||
running++
|
||||
@@ -1277,6 +1288,22 @@ func (s *Server) backupRestoreHandler(w http.ResponseWriter, r *http.Request) {
|
||||
http.Redirect(w, r, "/backups/restore?flash="+url.QueryEscape("Visszaállítás elindult — az állapot itt frissül."), http.StatusFound)
|
||||
}
|
||||
|
||||
// C9-F1 customer-facing strings. Kept as named constants, not inlined, because both are asserted
|
||||
// verbatim by tests — a silent edit to either is the way an honest message drifts back into a
|
||||
// comforting one.
|
||||
const (
|
||||
// tier2NoCoverageMsg is shown when this app's data cannot come from the secondary copy at all.
|
||||
// It NAMES the action that works rather than leaving a dead end: the keep-side recovery-unit
|
||||
// restore on /backups/restore, which does restore named volumes and DB dumps (proven live,
|
||||
// Campaign 9 A2). It also states plainly that no outage was taken, because the previous behaviour
|
||||
// took one.
|
||||
tier2NoCoverageMsg = "Ennek az alkalmazásnak az adatai nem ebből a másolatból állíthatók vissza — az alkalmazás nem állt le. Használd a Visszaállítás indítása gombot a Biztonsági mentés → Visszaállítás oldalon."
|
||||
|
||||
// tier2UnitNotCoveredMsg is appended wherever the restore DID run, so a clean result never reads
|
||||
// as a clean bill of health for data the operation never opened.
|
||||
tier2UnitNotCoveredMsg = "Az alkalmazás adatbázisa és belső kötetei nem tartoznak ebbe a visszaállításba."
|
||||
)
|
||||
|
||||
// backupTier2RestoreHandler (C2, closes F2) restores an app's MISSING user files in place from its
|
||||
// recorded Tier-2 copy — additive-only: existing live files are never overwritten and nothing is
|
||||
// ever deleted (see backup.RestoreTier2Files). Same handler shape as backupRestoreHandler.
|
||||
@@ -1303,20 +1330,50 @@ func (s *Server) backupTier2RestoreHandler(w http.ResponseWriter, r *http.Reques
|
||||
http.Redirect(w, r, "/backups/apps?flash_error="+url.QueryEscape("Egy mentési/visszaállítási művelet már fut."), http.StatusFound)
|
||||
return
|
||||
}
|
||||
|
||||
// C9-F1: refuse UP FRONT — before any op is begun and before the app is stopped — when this app's
|
||||
// Tier-2 copy holds nothing this restore can read (43 of the 53 catalog apps: their data lives in
|
||||
// Docker named volumes, captured into recovery-unit/ and never read here). Previously the customer
|
||||
// got an outage, zero files, and „Nincs hiányzó fájl — minden fájl megvan a helyén." — a claim
|
||||
// about data the restore never examined, at the exact moment they pressed it BECAUSE data was
|
||||
// missing. Only the no-coverage case is pre-flighted; every other refusal keeps its existing async
|
||||
// path so this change cannot alter behaviour anywhere else.
|
||||
cov, covErr := s.backupMgr.Tier2RestoreCoverage(stackName)
|
||||
if covErr == nil && !cov.CanRestore() {
|
||||
s.logger.Printf("[WARN] [web] Tier-2 file restore refused up front: stack=%s has no restorable subtree in its copy (unit_present=%v) — app NOT stopped", stackName, cov.HasUnit)
|
||||
http.Redirect(w, r, "/backups/apps?flash_error="+url.QueryEscape(tier2NoCoverageMsg), http.StatusFound)
|
||||
return
|
||||
}
|
||||
|
||||
s.logger.Printf("[WARN] [web] Tier-2 file restore requested (async): stack=%s from %s", stackName, r.RemoteAddr)
|
||||
s.backupMgr.BeginRestoreOp("tier2-restore", stackName)
|
||||
go func() {
|
||||
n, err := s.backupMgr.RestoreTier2Files(stackName)
|
||||
if err != nil {
|
||||
// The no-coverage refusal is not an operational failure — it means this action does not
|
||||
// apply to this app. Say that, and name the one that does, instead of "sikertelen".
|
||||
if errors.Is(err, backup.ErrTier2NoRestorableData) {
|
||||
s.logger.Printf("[WARN] [web] Tier-2 file restore not applicable: stack=%s", stackName)
|
||||
s.backupMgr.EndRestoreOp(false, tier2NoCoverageMsg)
|
||||
return
|
||||
}
|
||||
s.logger.Printf("[ERROR] [web] Tier-2 file restore failed (async): stack=%s: %v", stackName, err)
|
||||
s.backupMgr.EndRestoreOp(false, "Fájl-visszaállítás sikertelen: "+err.Error())
|
||||
return
|
||||
}
|
||||
msg := "Nincs hiányzó fájl — minden fájl megvan a helyén."
|
||||
// C9-F1 (the quiet half): even where the restore DOES cover something it covers only the
|
||||
// file-based legs — never the app's database or named volumes, which sit unread in the same
|
||||
// copy's recovery-unit/. „minden fájl megvan a helyén" was a blanket claim over data that was
|
||||
// never opened; immich's 1.3 GB Postgres unit is the case that makes it dangerous. Claim only
|
||||
// what was EXAMINED, and disclose the rest.
|
||||
msg := "Minden vizsgált fájl megvan a helyén."
|
||||
if n > 0 {
|
||||
msg = fmt.Sprintf("%s: %d fájl visszaállítva a másodlagos másolatból.", stackName, n)
|
||||
}
|
||||
s.logger.Printf("[INFO] [web] Tier-2 file restore completed (async): stack=%s (%d files)", stackName, n)
|
||||
if cov.HasUnit {
|
||||
msg += " " + tier2UnitNotCoveredMsg
|
||||
}
|
||||
s.logger.Printf("[INFO] [web] Tier-2 file restore completed (async): stack=%s (%d files, legs=%v)", stackName, n, cov.Legs)
|
||||
s.backupMgr.EndRestoreOp(true, msg)
|
||||
}()
|
||||
http.Redirect(w, r, "/backups/apps?flash="+url.QueryEscape("Fájl-visszaállítás elindult — az állapot itt frissül."), http.StatusFound)
|
||||
|
||||
@@ -0,0 +1,177 @@
|
||||
package web
|
||||
|
||||
import (
|
||||
"io"
|
||||
"log"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"sync/atomic"
|
||||
"testing"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-controller/internal/backup"
|
||||
"gitea.dooplex.hu/admin/felhom-controller/internal/config"
|
||||
"gitea.dooplex.hu/admin/felhom-controller/internal/settings"
|
||||
)
|
||||
|
||||
// C9-F1 at the customer surface. Two messages had to change, and they fail differently:
|
||||
//
|
||||
// - the LOUD lie: for the 43 apps whose data the restore cannot read, the customer got an outage
|
||||
// and „Nincs hiányzó fájl — minden fájl megvan a helyén." — pressed precisely BECAUSE files were
|
||||
// missing;
|
||||
// - the QUIET one: for the 9 apps it does cover, it covers only the file legs, never the database
|
||||
// or named volumes, so the same blanket sentence was a clean bill of health over data the
|
||||
// operation never opened (immich's 1.3 GB Postgres unit is the dangerous case).
|
||||
|
||||
// honestProvider is a minimal provider: the restore must never reach StopStack in the no-coverage
|
||||
// test, and `stops` is the assertion that proves it.
|
||||
type honestProvider struct {
|
||||
hdd string
|
||||
stops int32
|
||||
}
|
||||
|
||||
func (p *honestProvider) GetStackComposePath(string) (string, bool) { return "", false }
|
||||
func (p *honestProvider) ListDeployedStacks() []backup.StackSummary { return nil }
|
||||
func (p *honestProvider) GetStackHDDMounts(string) []string { return nil }
|
||||
func (p *honestProvider) GetStackHDDPath(string) string { return p.hdd }
|
||||
func (p *honestProvider) GetImportRoot() string { return "" }
|
||||
func (p *honestProvider) GetDockerVolumes(string) []string { return nil }
|
||||
func (p *honestProvider) StopStack(string) error { atomic.AddInt32(&p.stops, 1); return nil }
|
||||
func (p *honestProvider) StartStack(string) error { return nil }
|
||||
func (p *honestProvider) RefreshAndIsRunning(string) bool { return true }
|
||||
func (p *honestProvider) GetStackRecoveryInfo(string) (backup.RecoveryInfo, bool) {
|
||||
return backup.RecoveryInfo{}, false
|
||||
}
|
||||
func (p *honestProvider) GetStackClassifiedBinds(string) ([]backup.ClassifiedBind, bool) {
|
||||
return nil, false
|
||||
}
|
||||
func (p *honestProvider) RecoverStackSecrets(string, []string) map[string]string { return nil }
|
||||
func (p *honestProvider) RecreateStackDefinitionFromUnit(string, string, map[string]string) error {
|
||||
return nil
|
||||
}
|
||||
func (p *honestProvider) StartStackServices(string, []string) error { return nil }
|
||||
|
||||
// newHonestServer builds a server whose recorded Tier-2 copy has `legs` (each created as a dir) and,
|
||||
// optionally, a recovery unit — so one harness expresses both the class-A and class-B shapes.
|
||||
func newHonestServer(t *testing.T, legs []string, withUnit bool) (*Server, *honestProvider) {
|
||||
t.Helper()
|
||||
tmp := t.TempDir()
|
||||
live := filepath.Join(tmp, "usb")
|
||||
dest := filepath.Join(tmp, "flash")
|
||||
lg := log.New(io.Discard, "", 0)
|
||||
sett, err := settings.Load(filepath.Join(tmp, "settings.json"), lg)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
for _, p := range []string{live, dest} {
|
||||
if err := sett.AddStoragePath(settings.StoragePath{Path: p, Label: filepath.Base(p)}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
if err := sett.SetCrossDriveConfig("app", &settings.CrossDriveBackup{
|
||||
Enabled: true, Method: "rsync", DestinationPath: dest,
|
||||
LastRun: "2026-07-28T03:30:00Z", LastStatus: "ok",
|
||||
}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
destBase := filepath.Join(dest, "backups", "secondary", "app")
|
||||
for _, leg := range legs {
|
||||
if err := os.MkdirAll(filepath.Join(destBase, leg, "appdata"), 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
if withUnit {
|
||||
if err := os.MkdirAll(filepath.Join(destBase, "recovery-unit", "volume-dumps"), 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
if err := os.WriteFile(filepath.Join(destBase, ".felhom-tier2-layout"), []byte("2"), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
cfg := &config.Config{}
|
||||
cfg.Paths.DataDir = tmp
|
||||
m := backup.NewManager(cfg, sett, lg)
|
||||
prov := &honestProvider{hdd: live}
|
||||
m.SetStackProvider(prov)
|
||||
return &Server{cfg: cfg, backupMgr: m, logger: lg}, prov
|
||||
}
|
||||
|
||||
// SCENARIO D at the surface — the customer is told plainly, up front, and the app is NOT stopped.
|
||||
// The message must NAME the action that works; a dead end for 81% of the catalog is not honesty.
|
||||
//
|
||||
// RED-PROOF (observed): remove the pre-flight `!cov.CanRestore()` block from the handler →
|
||||
//
|
||||
// tier2_honest_message_test.go:118: no flash_error — the customer was told the restore STARTED
|
||||
// tier2_honest_message_test.go:129: THE APP WAS STOPPED (stops=1) for a restore that can never restore anything
|
||||
func TestTier2RestoreHandler_NoCoverage_RefusesUpFrontAndNamesTheAction(t *testing.T) {
|
||||
s, prov := newHonestServer(t, nil, true) // bookstack shape: unit only, no legs
|
||||
|
||||
req := httptest.NewRequest(http.MethodPost, "/backup/tier2/restore", strings.NewReader("stack_name=app"))
|
||||
req.Header.Set("Content-Type", "application/x-www-form-urlencoded")
|
||||
w := httptest.NewRecorder()
|
||||
s.backupTier2RestoreHandler(w, req)
|
||||
|
||||
loc := w.Header().Get("Location")
|
||||
if !strings.Contains(loc, "flash_error=") {
|
||||
t.Fatalf("no flash_error — the customer was told the restore STARTED: %s", loc)
|
||||
}
|
||||
// The refusal must not have started an operation at all.
|
||||
if st := s.backupMgr.RestoreStatus(); st.Running {
|
||||
t.Error("an async op was begun for a restore that cannot do anything")
|
||||
}
|
||||
if got := atomic.LoadInt32(&prov.stops); got != 0 {
|
||||
t.Errorf("THE APP WAS STOPPED (stops=%d) for a restore that can never restore anything", got)
|
||||
}
|
||||
// The message names the working action rather than dead-ending.
|
||||
for _, want := range []string{"nem ebből a másolatból", "nem állt le", "Visszaállítás indítása"} {
|
||||
if !strings.Contains(tier2NoCoverageMsg, want) {
|
||||
t.Errorf("the refusal message is missing %q:\n%s", want, tier2NoCoverageMsg)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// SCENARIO F — „nothing missing" must claim only what was EXAMINED, and must disclose what this
|
||||
// restore does not cover at all.
|
||||
//
|
||||
// RED-PROOF (observed): restore the old blanket string
|
||||
// (`msg := "Nincs hiányzó fájl — minden fájl megvan a helyén."` with no disclosure) →
|
||||
//
|
||||
// tier2_honest_message_test.go:154: the success message still claims ALL files: "Nincs hiányzó fájl — minden fájl megvan a helyén."
|
||||
// tier2_honest_message_test.go:161: the message does not disclose that the database and volumes were not covered
|
||||
func TestTier2RestoreHandler_CoveredApp_ClaimsOnlyWhatWasExamined(t *testing.T) {
|
||||
s, _ := newHonestServer(t, []string{"hdd"}, true) // paperless shape: a leg AND a unit
|
||||
|
||||
req := httptest.NewRequest(http.MethodPost, "/backup/tier2/restore", strings.NewReader("stack_name=app"))
|
||||
req.Header.Set("Content-Type", "application/x-www-form-urlencoded")
|
||||
w := httptest.NewRecorder()
|
||||
s.backupTier2RestoreHandler(w, req)
|
||||
|
||||
if !strings.Contains(w.Header().Get("Location"), "flash=") {
|
||||
t.Fatalf("a COVERED app was refused — this is the regression guard on Campaign 9's A1 result: %s", w.Header().Get("Location"))
|
||||
}
|
||||
|
||||
var last backup.RestoreOpResult
|
||||
waitFor(t, func() bool {
|
||||
st := s.backupMgr.RestoreStatus()
|
||||
if st.Running || st.Last == nil {
|
||||
return false
|
||||
}
|
||||
last = *st.Last
|
||||
return true
|
||||
}, "the restore to finish")
|
||||
|
||||
if !last.OK {
|
||||
t.Fatalf("a covered app's restore failed: %s", last.Message)
|
||||
}
|
||||
if strings.Contains(last.Message, "minden fájl megvan a helyén") {
|
||||
t.Errorf("the success message still claims ALL files: %q", last.Message)
|
||||
}
|
||||
if !strings.Contains(last.Message, "vizsgált") {
|
||||
t.Errorf("the message does not limit its claim to what was EXAMINED: %q", last.Message)
|
||||
}
|
||||
if !strings.Contains(last.Message, tier2UnitNotCoveredMsg) {
|
||||
t.Errorf("the message does not disclose that the database and volumes were not covered: %q", last.Message)
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user