diff --git a/CHANGELOG.md b/CHANGELOG.md index e5a205a..aca09d4 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,3 +1,54 @@ +## v0.239.0 — any backup tier lets an app update (2026-09-13, R-475) + +**MinAgent: 0.129.0** (unchanged) + +**Operator ruling 2026-09-13: every backup counts.** Until now the update's precondition was the +Tier-2 unit predicate alone, so an app with no second-drive copy could never be updated. On demo-hp +`gokapi` and `nextcloud` were in exactly that state, each with a fresh recovery unit on its own drive. + +**WHAT CHANGED.** + +- **`backup.Manager.UpdateRestorePoints`**, beside `Tier2UnitRestorePoint`. It walks the tiers in the + ruling's order: Tier 2 (second drive), Tier 1 (the app's own unit, `ListRestorePoints`, „helyi"), + Tier 3 (`OffsiteInventoryList`, bounded by 15 s). It returns the first copy the caller accepts and + stops there, so a fresh Tier-2 copy never reaches the network. An unreachable off-site repository + counts as ABSENT, with a WARN. A box with no off-site target is plainly absent, with no WARN. +- **The age rule is one rule.** stacks passes `freshRestorePoint` (`update.backup_max_age`) as the + acceptance test, so the limit applies to whichever tier is chosen. The first FRESH copy wins, so a + stale Tier-2 mirror never forces a backup while the app's own unit is minutes old. +- **No copy anywhere → back up first**, then re-read every tier. The preflight refuses `no_backup` + only when there is no copy on any tier AND `CanBackUpApp` says no backup can be taken now. The + Hungarian sentence changed to say that; it no longer tells the customer to switch on the 2nd backup. +- **`RunAppBackupNow` tolerates no Tier-2 target.** A Tier-2 failure after the unit capture is a WARN. + The captured unit's manifest is marked proven current, because the capture's checksum skip leaves it + untouched on a quiet app, and Tier 1 is aged by the newest artifact's mtime. Without that mark an + app with no database and no volume would be refused forever — the ProvenCopyTime trap, one tier down. +- **The hold names the tier.** `settings.RestoreHold.CopyTier`; the sentence now ends *„Visszaállítható + a Mentések oldalon ebből a biztonsági mentésből: , + ."* A hold written by v0.237.0–v0.238.1 has no tier and keeps its original sentence + (`UpdateHoldLegacyFmt`). The update journal records `proven_tier`, so a resumed update names the + right copy. +- **A successful off-site restore lifts an update hold.** That path never went through + `RestoreFromRecoveryUnitAt`, so a hold naming „távoli mentés" could otherwise never be cleared. +- **Unchanged:** the backups page still calls `Tier2UnitRestorePoint` for „Teljes visszaállítás". + +**TESTS.** `internal/stacks/update_tiers_test.go` — G (Tier 2 chosen when present), H (own unit alone), +I (off-site alone), K (nothing anywhere: backed up first and the update completes; a failing backup +moves nothing), L (an existing copy still carries an app that cannot be backed up; control: without it +the app is refused), M (a stale copy on each tier is backed up first; a stale Tier 2 does not block a +fresh Tier 1; the limit is inclusive on one clock). `internal/backup/update_tiers_test.go` — the tier +order and early stop, H/I at the source, J (unreachable → absent + WARN; no target → silent; a hanging +repository ends at the bound), the hold text per tier and the legacy text, the tolerant pre-backup +tail, `RunAppBackupNow` really calling it, `CanBackUpApp`, and the off-site restore clearing the hold. +`cmd/controller/r475_wiring_test.go` — the tier numbers agree across packages; the adapter reads every +tier and never `Tier2UnitRestorePoint`; the hold is told the tier. Slice 4's tests were moved onto the +new seam (`TestSlice4_C` is now the L refusal). + +**Red-proofs** (felhom.eu `documentation/audits/rulings-r472-r475-2026-09-13/`): **M** — age checked +only for Tier 2: three M cases fail (a 30-hour-old own unit and off-site copy carry the update with no +backup). **L** — refuse even with a copy: the L test fails. **Tail** — the proven-current mark misses: +the tail test fails on the mtime. **Off-site clear** removed: the hold stays and the test fails. + ## v0.238.1 — the nightly backup leaves an app alone WHILE it is being updated, not only once it is held (2026-09-13, slice 4 follow-up) **MinAgent: 0.129.0** (unchanged) diff --git a/REUSE.md b/REUSE.md index 4e62258..c20bc82 100644 --- a/REUSE.md +++ b/REUSE.md @@ -122,8 +122,9 @@ | `Manager.recordInstalledImages` (v0.233.0) | controller/internal/stacks/installed.go | `(name, stackDir string, env []string)` | writing down what each compose SERVICE is ACTUALLY running, into `app.yaml.installed_images` | Called after a successful compose up from `StartStack`/`RestartStack`/`UpdateStack`/`runComposeDeploy`. **Reads the CONTAINER, never `docker-compose.yml`** — that file is the value the syncer has already moved (spike §3: 25 minutes of disagreement). **A failed write NEVER refuses the action** — the deliberate OPPOSITE of `SetDesiredState`: intent refused, observation logged at ERROR. **NOT from `StartStackServices`** (the R-47 DB-only window would overwrite a complete record with a partial one). Skips the write when ref+digest are unchanged, and carries `at` forward so it means "running since". Its OWN seam (`installedExecFn`) with a **context + 30 s timeout** — the two existing exec helpers have neither | | `Manager.SetPin` / `AdoptPins` / `RenderPlanFor` / `AppliedComposePath` (v0.235.0) | controller/internal/stacks/pin.go | `SetPin(name, stackDir, pin, composeSrc) error` | THE version freeze — `app.yaml.pinned_images` + the stored `applied-compose.yml` | **`PinnedImages` is INTENT, `InstalledImages` is an OBSERVATION — never feed one from the other** (the R-166 category error, one field over). Four writers only: deploy, the guarded update (via `advancePinToCatalog`, which advances the pin and re-renders BEFORE the pull, and REFUSES the update if the pin cannot be written; a failed pull puts the pin BACK via `SetPin`), the restore adapter (this is what closes R-441), and `AdoptPins`. `AdoptPins` reuses `observationCoversTemplate` — do NOT write a second completeness rule — and skips loudly rather than inventing a pin. Absent pin = pre-v0.235.0 behaviour | | `Manager.UpdatePreflight` / `StartGuardedUpdate` / `RecoverUpdates` / `ResumeInterruptedUpdates` (v0.237.0) | controller/internal/stacks/update.go | `UpdatePreflight(name) *UpdateRefusal`; `StartGuardedUpdate(name) error` | THE update — refusals, then a 202 job with phases on `Stack.Updating/UpdatePhase/UpdateError` | **Never report an update complete before health is known (R-443).** Every cheap refusal runs BEFORE the intent write. Safety dump BEFORE the pin moves; pin BEFORE pull; pull failure → pin back; health failure → stop + HOLD, pin stays. Journal-before-mutate (`update-journal.json`); `RecoverUpdates` MUST run before the boot sweep and `ResumeInterruptedUpdates` AFTER `SetUpdateGuards`. Seams: `updateComposeFn`, `updateHealthFn`, `updateMemoryFn`, `updateDiskFreeFn`, `updateNowFn` (R-457: the age check and the test read ONE clock). Unwired guards ⇒ every update refused | -| `stacks.UpdateGuards` + `updateGuardsAdapter` (v0.237.0) | controller/internal/stacks/update.go, controller/cmd/controller/main.go | `HoldFor`, `Busy`, `RestorePoint`, `BackupNow`, `SafetyDump`, `HoldAfterFailedUpdate` | the ONLY bridge from the update job to the backup side (stacks cannot import backup) | Wired by `stackMgr.SetUpdateGuards` — pinned by `TestSlice4_UpdateGuardsAreWiredAtStartup`. Add a guard HERE, never by importing backup into stacks | -| `backup.Manager.Tier2UnitRestorePoint` + `Tier2RestorePoint.ProvenCopyTime` (v0.237.0) | controller/internal/backup/update_guard.go | `(stack) (Tier2RestorePoint, error)` | "can this app be restored from Tier 2, and from when" — the predicate that gates BOTH the „Teljes visszaállítás" action and an update | **ONE predicate, two callers** (extracted from `buildAppBackupRows`, not copied). `CopyDate` is what the page NAMES (the package date, R-403); `ProvenCopyTime` is how OLD the data is — the last successful copy, because the manifest's `created_at` moves only when the DEFINITION changes (measured: a fresh dump under a 22-h-older manifest). Do not age a copy by `CopyDate` | +| `stacks.UpdateGuards` + `updateGuardsAdapter` (v0.237.0; tiers v0.239.0) | controller/internal/stacks/update.go, controller/cmd/controller/main.go | `HoldFor`, `Busy`, `RestorePoints`, `CanBackUp`, `BackupNow`, `SafetyDump`, `HoldAfterFailedUpdate` | the ONLY bridge from the update job to the backup side (stacks cannot import backup) | Wired by `stackMgr.SetUpdateGuards` — pinned by `TestSlice4_UpdateGuardsAreWiredAtStartup`. Add a guard HERE, never by importing backup into stacks | +| `backup.Manager.Tier2UnitRestorePoint` + `Tier2RestorePoint.ProvenCopyTime` (v0.237.0) | controller/internal/backup/update_guard.go | `(stack) (Tier2RestorePoint, error)` | "can this app be restored from Tier 2, and from when" — the predicate that gates BOTH the „Teljes visszaállítás" action and an update | **ONE predicate, two callers** (extracted from `buildAppBackupRows`, not copied). `CopyDate` is what the page NAMES (the package date, R-403); `ProvenCopyTime` is how OLD the data is — the last successful copy, because the manifest's `created_at` moves only when the DEFINITION changes (measured: a fresh dump under a 22-h-older manifest). Do not age a copy by `CopyDate`. **Since v0.239.0 the UPDATE no longer calls it directly** — it goes through `UpdateRestorePoints` (next row); the page still does | +| `backup.Manager.UpdateRestorePoints` + `CanBackUpApp` (v0.239.0, R-475) | controller/internal/backup/update_guard.go | `(ctx, stack, accept func(UpdateTierPoint) bool) (UpdateTierPoint, bool, []UpdateTierPoint)` | "which backup can this update lean on" — walks Tier 2, 1, 3 and returns the first copy `accept` admits | **The age rule lives in the caller's `accept`** (stacks' `freshRestorePoint`), so it is ONE rule for every tier. Stops at the first accepted copy, so Tier 3 (restic, 15 s bound, unreachable = absent + WARN) is reached only when needed. Seams: `updateTier2PointFn`, `updateTier1PointsFn`, `updateOffsiteInvFn`. Tier numbers are pinned equal across stacks/backup by `TestR475_TierConstantsAgree` | | `backup.Manager.HoldAfterFailedUpdate` / `RunAppBackupNow` / `WriteUpdateSafetyDump` / `UpdateBusy` (v0.237.0) | controller/internal/backup/update_guard.go | see file | the update's hold, per-app backup-now, safety dump, busy check | The hold is `settings.RestoreHold` with `Reason: update_failed` — SAME store and gate as R-379, never a second map. `RunAppBackupNow` composes the nightly legs for ONE app (admission, DB dump, volume dump, capture, Tier-2) — do not write a second backup orchestration. A successful unit restore lifts an UPDATE hold only. **`isHeld` is ALSO true while a guarded update is moving the app (`SetUpdatingCheck`, v0.238.1)** — found live: the periodic capture overwrote a primary unit during a health wait | | `stacks.Manager.memoryVerdict` (v0.237.0) | controller/internal/stacks/deploy.go | `(newReq, newLimit, releasedReq, releasedLimit int) (refusal, warning string)` | the deploy's memory check, shared with the update | An update RELEASES the app's current request first. Deploy passes `0, 0` and is byte-identical in wording and log line | | `Syncer.SetRenderPlanFn` + `renderSource` (v0.235.0) | controller/internal/sync/sync.go | `func(appName string) stacks.RenderPlan` | the catalog render table | **NIL-SAFE: no seam = copy verbatim = the old product.** Catalog images == pin → verbatim (fixes flow + self-healing, both deliberately kept); differ → the WHOLE stored definition, **never a ref substitution into a newer template** (`wger 2.6`). `.felhom.yml` always verbatim (R-458). The syncer must NEVER read app.yaml. Re-reads the applied file before writing it — a test caught it writing an empty compose over a live app | diff --git a/controller/README.md b/controller/README.md index 187f115..0479e39 100644 --- a/controller/README.md +++ b/controller/README.md @@ -570,22 +570,31 @@ job, and answers **202**. The page polls `GET /api/stacks/{name}`. **Refused with 409 before anything moves:** held; a backup, restore, app-data op or quiesce holding the app; a migration; already updating; deploying; not enough memory for the NEW template's request (the deploy's own `memoryVerdict`, releasing the app's current request); less than **2 GB** free on -the Docker data root (a fixed floor — image sizes are not known without a registry query); and **no -restorable backup** (no openable Tier-2 recovery unit with a proven copy). +the Docker data root (a fixed floor — image sizes are not known without a registry query); and — since +v0.239.0 — **no copy on any backup tier AND no way to take one now** (drive unresolvable, disconnected, +or a migration running). -**The sequence.** A proven copy older than `update.backup_max_age` is refreshed first -(`RunAppBackupNow`: this app's DB dump, volume dump, unit capture, Tier-2 copy). Then a database +**Any backup tier counts (v0.239.0, R-475).** The update leans on the first FRESH copy in the order +second drive (Tier 2), the app's own recovery unit (Tier 1, „helyi"), off-site (Tier 3, looked up with +a 15 s bound — unreachable counts as absent, with a WARN). `update.backup_max_age` applies to whichever +tier is chosen. An app with no copy anywhere is backed up first. Tier 2 is required nowhere in the +update path; the backups page's „Teljes visszaállítás" still reads the Tier-2 predicate alone. + +**The sequence.** With no fresh copy on any tier the app is backed up first (`RunAppBackupNow`: this +app's DB dump, volume dump, unit capture, then a Tier-2 copy whose failure is only a WARN — the unit +just captured is marked proven current, so a quiet app's own unit counts as fresh). Then a database safety dump, then the pin moves, then pull, `up`, and the health wait (`.felhom.yml` check, or 60 s of every container running for an app with none; bounded by `update.health_timeout`). **A failed pull puts the pin back. An app that does not become healthy is stopped and HELD** — the pin stays on the -new version, and the hold sentence names the backup it can be restored from. A successful unit -restore lifts an update hold. +new version, and the hold sentence names the tier and the date of the copy it can be restored from +(„második meghajtó" / „saját meghajtó" / „távoli mentés"). A successful unit restore — and, since +v0.239.0, a successful off-site restore — lifts an update hold. **Config (`controller.yaml`):** ```yaml update: - backup_max_age: 24h # a proven Tier-2 copy older than this is refreshed before the update + backup_max_age: 24h # the chosen copy (any tier) must be younger than this, or the app is backed up first health_timeout: 5m # how long the new version has to become healthy before the app is held ``` diff --git a/controller/cmd/controller/main.go b/controller/cmd/controller/main.go index 00b6850..f3fe62d 100644 --- a/controller/cmd/controller/main.go +++ b/controller/cmd/controller/main.go @@ -3378,16 +3378,32 @@ func (a *updateGuardsAdapter) Busy(name string) (bool, string) { return a.b.UpdateBusy(name) } -func (a *updateGuardsAdapter) RestorePoint(name string) (stacks.UpdateRestorePoint, error) { +// RestorePoints (R-475) reads EVERY tier through backup.UpdateRestorePoints. It must never go back to +// Tier2UnitRestorePoint alone — pinned by TestR475_AdapterReadsEveryTier. +func (a *updateGuardsAdapter) RestorePoints(ctx context.Context, name string, accept func(stacks.UpdateRestorePoint) bool) (stacks.UpdateRestorePoint, bool, []stacks.UpdateRestorePoint) { if a.b == nil { - return stacks.UpdateRestorePoint{}, fmt.Errorf("backup is not enabled on this box") + return stacks.UpdateRestorePoint{}, false, nil } - rp, err := a.b.Tier2UnitRestorePoint(name) - if err != nil { - return stacks.UpdateRestorePoint{}, err + conv := func(p backup.UpdateTierPoint) stacks.UpdateRestorePoint { + return stacks.UpdateRestorePoint{Tier: p.Tier, ProvenAt: p.At} } - at, proven := rp.ProvenCopyTime() - return stacks.UpdateRestorePoint{Restorable: rp.Restorable, Proven: proven, ProvenAt: at}, nil + var acc func(backup.UpdateTierPoint) bool + if accept != nil { + acc = func(p backup.UpdateTierPoint) bool { return accept(conv(p)) } + } + chosen, ok, seen := a.b.UpdateRestorePoints(ctx, name, acc) + out := make([]stacks.UpdateRestorePoint, 0, len(seen)) + for _, p := range seen { + out = append(out, conv(p)) + } + return conv(chosen), ok, out +} + +func (a *updateGuardsAdapter) CanBackUp(name string) (bool, string) { + if a.b == nil { + return false, "backup is not enabled on this box" + } + return a.b.CanBackUpApp(name) } func (a *updateGuardsAdapter) BackupNow(ctx context.Context, name string) error { @@ -3404,9 +3420,9 @@ func (a *updateGuardsAdapter) SafetyDump(ctx context.Context, name string) ([]st return a.b.WriteUpdateSafetyDump(ctx, name) } -func (a *updateGuardsAdapter) HoldAfterFailedUpdate(name string, at, provenCopyAt time.Time) error { +func (a *updateGuardsAdapter) HoldAfterFailedUpdate(name string, at time.Time, rp stacks.UpdateRestorePoint) error { if a.b == nil { return fmt.Errorf("backup is not enabled on this box — the hold cannot be recorded") } - return a.b.HoldAfterFailedUpdate(name, at, provenCopyAt) + return a.b.HoldAfterFailedUpdate(name, at, rp.ProvenAt, rp.Tier) } diff --git a/controller/cmd/controller/r475_wiring_test.go b/controller/cmd/controller/r475_wiring_test.go new file mode 100644 index 0000000..f901af4 --- /dev/null +++ b/controller/cmd/controller/r475_wiring_test.go @@ -0,0 +1,71 @@ +package main + +import ( + "go/ast" + "go/parser" + "go/token" + "strings" + "testing" + + "gitea.dooplex.hu/admin/felhom-controller/internal/backup" + "gitea.dooplex.hu/admin/felhom-controller/internal/stacks" +) + +// R-475 — the update reads EVERY backup tier, through the one adapter. + +func TestR475_TierConstantsAgree(t *testing.T) { + if stacks.UpdateTierLocal != backup.UpdateTierLocal || + stacks.UpdateTierSecondDrive != backup.UpdateTierSecondDrive || + stacks.UpdateTierOffsite != backup.UpdateTierOffsite { + t.Fatal("stacks and backup number the tiers differently — a hold would name the wrong copy") + } +} + +// adapterMethodSelectors returns the selector names used in updateGuardsAdapter.'s body. +func adapterMethodSelectors(t *testing.T, name string) string { + t.Helper() + fset := token.NewFileSet() + f, err := parser.ParseFile(fset, "main.go", nil, 0) + if err != nil { + t.Fatal(err) + } + for _, d := range f.Decls { + fn, ok := d.(*ast.FuncDecl) + if !ok || fn.Recv == nil || fn.Name.Name != name || fn.Body == nil { + continue + } + star, ok := fn.Recv.List[0].Type.(*ast.StarExpr) + if !ok { + continue + } + if id, ok := star.X.(*ast.Ident); !ok || id.Name != "updateGuardsAdapter" { + continue + } + var names []string + ast.Inspect(fn.Body, func(n ast.Node) bool { + if sel, ok := n.(*ast.SelectorExpr); ok { + names = append(names, sel.Sel.Name) + } + return true + }) + return " " + strings.Join(names, " ") + " " + } + t.Fatalf("updateGuardsAdapter.%s not found in main.go", name) + return "" +} + +func TestR475_AdapterReadsEveryTier(t *testing.T) { + rp := adapterMethodSelectors(t, "RestorePoints") + if !strings.Contains(rp, " UpdateRestorePoints ") { + t.Errorf("the adapter must read every tier through backup.UpdateRestorePoints; selectors:%s", rp) + } + if strings.Contains(rp, " Tier2UnitRestorePoint ") { + t.Error("the Tier-2-only predicate must not come back into the update path (R-475)") + } + if cb := adapterMethodSelectors(t, "CanBackUp"); !strings.Contains(cb, " CanBackUpApp ") { + t.Errorf("CanBackUp must ask backup.CanBackUpApp; selectors:%s", cb) + } + if h := adapterMethodSelectors(t, "HoldAfterFailedUpdate"); !strings.Contains(h, " Tier ") { + t.Errorf("the hold must be told the chosen TIER, or it cannot name it; selectors:%s", h) + } +} diff --git a/controller/internal/api/slice4_update_test.go b/controller/internal/api/slice4_update_test.go index 2f594fa..ecb43a4 100644 --- a/controller/internal/api/slice4_update_test.go +++ b/controller/internal/api/slice4_update_test.go @@ -24,8 +24,9 @@ import ( // every path here refuses, or fails at the pin (the app's catalog template is absent on purpose). type apiFakeGuards struct { - b *backup.Manager - rp stacks.UpdateRestorePoint + b *backup.Manager + points []stacks.UpdateRestorePoint + cannotBackUp bool // blindToHolds makes the manager-side preflight NOT see holds, so a test can prove the ROUTER's // own hold check refuses — the two layers are each pinned separately (the preflight's by // TestSlice4_D_CheapRefusals/held). Without it, removing either layer passes inertly, because the @@ -40,14 +41,20 @@ func (g *apiFakeGuards) HoldFor(n string) (bool, string) { return g.b.RestoreHoldFor(n) } func (g *apiFakeGuards) Busy(string) (bool, string) { return false, "" } -func (g *apiFakeGuards) RestorePoint(string) (stacks.UpdateRestorePoint, error) { - return g.rp, nil +func (g *apiFakeGuards) RestorePoints(_ context.Context, _ string, accept func(stacks.UpdateRestorePoint) bool) (stacks.UpdateRestorePoint, bool, []stacks.UpdateRestorePoint) { + for _, p := range g.points { + if accept == nil || accept(p) { + return p, true, g.points + } + } + return stacks.UpdateRestorePoint{}, false, g.points } +func (g *apiFakeGuards) CanBackUp(string) (bool, string) { return !g.cannotBackUp, "fake: no drive" } func (g *apiFakeGuards) BackupNow(context.Context, string) error { return nil } func (g *apiFakeGuards) SafetyDump(context.Context, string) ([]string, error) { return nil, nil } -func (g *apiFakeGuards) HoldAfterFailedUpdate(string, time.Time, time.Time) error { return nil } +func (g *apiFakeGuards) HoldAfterFailedUpdate(string, time.Time, stacks.UpdateRestorePoint) error { return nil } const slice4AppYAML = "deployed: true\nenv: {}\npinned_images:\n app: nginx:1.27\n" @@ -82,7 +89,7 @@ func newSlice4Router(t *testing.T) (*Router, *settings.Settings, *apiFakeGuards, t.Fatal(err) } b := backup.NewManager(cfg, sett, lg) - g := &apiFakeGuards{b: b, rp: stacks.UpdateRestorePoint{Restorable: true, Proven: true, ProvenAt: time.Now().Add(-time.Hour)}} + g := &apiFakeGuards{b: b, points: []stacks.UpdateRestorePoint{{Tier: stacks.UpdateTierSecondDrive, ProvenAt: time.Now().Add(-time.Hour)}}} m.SetUpdateGuards(g) return &Router{cfg: cfg, stackMgr: m, backupMgr: b, logger: lg}, sett, g, dir } @@ -105,7 +112,7 @@ func postUpdate(t *testing.T, r *Router) (int, apiResponse) { // message — which is what proves the router line is the one doing it. func TestR439_UpdateOfAHeldAppIsRefused(t *testing.T) { r, sett, g, dir := newSlice4Router(t) - g.rp = stacks.UpdateRestorePoint{} // the preflight's own refusal would say "no backup" — not the hold + g.points, g.cannotBackUp = nil, true // the preflight's own refusal would say "no backup" — not the hold g.blindToHolds = true // only the router's line can produce the hold's sentence if err := sett.SetRestoreHold(settings.RestoreHold{Stack: "app", At: "2026-09-13T08:00:00Z", Reason: settings.HoldReasonUpdateFailed, CopyDate: "2026-09-13T01:30:00Z"}); err != nil { t.Fatal(err) @@ -124,7 +131,7 @@ func TestR439_UpdateOfAHeldAppIsRefused(t *testing.T) { func TestSlice4_Router_NoBackupIs409AndRecordsNothing(t *testing.T) { r, _, g, dir := newSlice4Router(t) - g.rp = stacks.UpdateRestorePoint{Restorable: false} + g.points, g.cannotBackUp = nil, true before, _ := os.ReadFile(filepath.Join(dir, "app.yaml")) code, resp := postUpdate(t, r) if code != http.StatusConflict || resp.Error != fmt.Sprintf(stacks.MsgUpdateNoBackupFmt, "app") { diff --git a/controller/internal/backup/backup.go b/controller/internal/backup/backup.go index 8d54106..1257de0 100644 --- a/controller/internal/backup/backup.go +++ b/controller/internal/backup/backup.go @@ -173,6 +173,13 @@ type Manager struct { // updatingCheck (slice 4) — nil-safe; see isHeld / SetUpdatingCheck. updatingCheck func(stackName string) bool + // R-475 update-precondition seams, one per tier. Nil → the real Tier2UnitRestorePoint / + // ListRestorePoints / OffsiteInventoryList. They let a test reach Tier 1 and Tier 3 without a drive + // or a restic repository; see UpdateRestorePoints. + updateTier2PointFn func(stackName string) (Tier2RestorePoint, error) + updateTier1PointsFn func(stackName string) ([]RestorePoint, bool) + updateOffsiteInvFn func(ctx context.Context) (OffsiteInventory, error) + // R-354 volume-REPLAY seam — the mirror of the F17 DB seams above, so the off-site path's new // volume leg is unit-testable without Docker. Nil → the real restoreDockerVolumesFrom. volumeReplayFrom func(stackName, dumpDir string) (int, error) diff --git a/controller/internal/backup/offbox_reconstitute.go b/controller/internal/backup/offbox_reconstitute.go index 2f91008..30f7413 100644 --- a/controller/internal/backup/offbox_reconstitute.go +++ b/controller/internal/backup/offbox_reconstitute.go @@ -344,7 +344,11 @@ func (m *Manager) RestoreHoldFor(stack string) (bool, string) { if h.CopyDate != "" { copyDate = fmtHoldTime(h.CopyDate) } - return true, fmt.Sprintf(UpdateHoldFmt, stack, fmtHoldTime(h.At), copyDate) + // R-475: name the tier when the hold recorded one; an older hold keeps its own sentence. + if label := UpdateTierLabel(h.CopyTier); label != "" && h.CopyDate != "" { + return true, fmt.Sprintf(UpdateHoldFmt, stack, fmtHoldTime(h.At), label, copyDate) + } + return true, fmt.Sprintf(UpdateHoldLegacyFmt, stack, fmtHoldTime(h.At), copyDate) } when := h.At if t, err := time.Parse(time.RFC3339, h.At); err == nil { @@ -824,6 +828,11 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string, ack if err := restartStack(); err != nil { return res, fmt.Errorf("a(z) %s újraindítása sikertelen a fájlok visszaállítása után: %w", stack, err) } + // R-475: an update hold may now name the OFF-SITE copy, and this is the route back it names. The + // unit restore clears it in RestoreFromRecoveryUnitAt; this path never went through that function, + // so without this line a successful off-site restore would leave the app refusing its next start. + // Pinned by TestR475_OffsiteRestoreClearsAnUpdateHold. + m.clearUpdateHoldAfterRestore(stack) if err := m.waitForHealthy(stack, 90*time.Second); err != nil { m.logger.Printf("[WARN] [offbox] %s reconstituted but health check failed: %v", stack, err) } diff --git a/controller/internal/backup/slice4_update_guard_test.go b/controller/internal/backup/slice4_update_guard_test.go index 7e765ec..c374e27 100644 --- a/controller/internal/backup/slice4_update_guard_test.go +++ b/controller/internal/backup/slice4_update_guard_test.go @@ -68,7 +68,7 @@ func TestSlice4_UpdateHoldTextNamesTheTimeAndTheCopy(t *testing.T) { m := &Manager{logger: log.New(io.Discard, "", 0), settings: sett} at := time.Date(2026, 9, 13, 8, 0, 0, 0, time.UTC) copyAt := time.Date(2026, 9, 13, 1, 30, 0, 0, time.UTC) - if err := m.HoldAfterFailedUpdate("bookstack", at, copyAt); err != nil { + if err := m.HoldAfterFailedUpdate("bookstack", at, copyAt, UpdateTierSecondDrive); err != nil { t.Fatal(err) } h, ok := sett.GetRestoreHold("bookstack") @@ -77,7 +77,7 @@ func TestSlice4_UpdateHoldTextNamesTheTimeAndTheCopy(t *testing.T) { } held, why := m.RestoreHoldFor("bookstack") // Budapest is UTC+2 in September: 08:00Z → 10:00, 01:30Z → 03:30. - want := fmt.Sprintf(UpdateHoldFmt, "bookstack", "2026-09-13 10:00", "2026-09-13 03:30") + want := fmt.Sprintf(UpdateHoldFmt, "bookstack", "2026-09-13 10:00", "második meghajtó", "2026-09-13 03:30") if !held || why != want { t.Errorf("hold text =\n%q\nwant\n%q", why, want) } @@ -98,7 +98,7 @@ func TestSlice4_RestoreHoldTextIsUnchanged(t *testing.T) { func TestSlice4_ASuccessfulRestoreClearsOnlyAnUpdateHold(t *testing.T) { sett := slice4Settings(t) m := &Manager{logger: log.New(io.Discard, "", 0), settings: sett} - _ = m.HoldAfterFailedUpdate("upd", time.Now(), time.Now()) + _ = m.HoldAfterFailedUpdate("upd", time.Now(), time.Now(), UpdateTierLocal) _ = sett.SetRestoreHold(settings.RestoreHold{Stack: "rst", At: "2026-08-22T14:00:00Z"}) m.clearUpdateHoldAfterRestore("upd") m.clearUpdateHoldAfterRestore("rst") @@ -136,7 +136,7 @@ func TestSlice4_UpdateBusy(t *testing.T) { func TestSlice4_NightlyLegsLeaveAHeldAppAlone(t *testing.T) { h := newAdmissionHarness(t, "held", "free") h.m.settings = slice4Settings(t) - if err := h.m.HoldAfterFailedUpdate("held", time.Now(), time.Now()); err != nil { + if err := h.m.HoldAfterFailedUpdate("held", time.Now(), time.Now(), UpdateTierSecondDrive); err != nil { t.Fatal(err) } h.m.runVolumeDumps() diff --git a/controller/internal/backup/update_guard.go b/controller/internal/backup/update_guard.go index d1e7085..728b011 100644 --- a/controller/internal/backup/update_guard.go +++ b/controller/internal/backup/update_guard.go @@ -4,6 +4,7 @@ import ( "context" "errors" "fmt" + "os" "path/filepath" "time" @@ -93,6 +94,149 @@ func (p Tier2RestorePoint) ProvenCopyTime() (time.Time, bool) { return t, true } +// ── R-475: any backup tier lets an app update (operator ruling 2026-09-13, controller v0.239.0) ──── +// +// Until v0.239.0 the update's precondition was Tier2UnitRestorePoint alone, so an app with no second +// drive could never be updated — even with a fresh recovery unit on its own drive and an off-site +// snapshot from last night. The ruling: every backup counts. Tier2UnitRestorePoint itself is NOT +// changed; the backups page still calls it for the „Teljes visszaállítás" action, which really does +// restore from the second drive only. + +// Backup tiers, as the update precondition and the hold sentence name them. +const ( + UpdateTierLocal = 1 // the app's own recovery unit on its drive — „helyi" on the restore page + UpdateTierSecondDrive = 2 // the Tier-2 mirror on another drive + UpdateTierOffsite = 3 // the off-site restic repository +) + +// updateTierOrder is the preference order the ruling set: the second drive, then the app's own unit, +// then off-site. The first tier holding a copy the caller ACCEPTS is chosen. +var updateTierOrder = []int{UpdateTierSecondDrive, UpdateTierLocal, UpdateTierOffsite} + +// UpdateTierLabel is a tier's name in the customer's hold sentence. "" for an unknown tier. +func UpdateTierLabel(tier int) string { + switch tier { + case UpdateTierSecondDrive: + return "második meghajtó" + case UpdateTierLocal: + return "saját meghajtó" + case UpdateTierOffsite: + return "távoli mentés" + } + return "" +} + +// updateOffsiteCheckTimeout bounds the off-site lookup. An update must not stall on an unreachable +// Storage Box: past this the off-site copy counts as ABSENT (with a WARN), and the update carries on +// with backing up first. A var only so a test can shorten it. +var updateOffsiteCheckTimeout = 15 * time.Second + +// UpdateTierPoint is one proven, restorable copy of an app on one tier. +type UpdateTierPoint struct { + Tier int + // At is when the data in that copy was last proven written: Tier 2 ProvenCopyTime, Tier 1 the + // newest artifact of the unit (ListRestorePoints), Tier 3 the newest snapshot for the app. + At time.Time +} + +// UpdateRestorePoints walks the tiers in preference order (2, 1, 3) and returns the FIRST copy that +// accept admits (nil accepts any), whether one was found, and every copy it looked at on the way. +// +// It stops at the first accepted copy, so a box with a fresh second-drive copy never touches the +// network. The AGE rule is the caller's (stacks applies backup_max_age through accept) — that is what +// makes "the age rule applies to whichever tier is chosen" one rule, not three (R-475 Scenario M). +func (m *Manager) UpdateRestorePoints(ctx context.Context, stackName string, accept func(UpdateTierPoint) bool) (UpdateTierPoint, bool, []UpdateTierPoint) { + var seen []UpdateTierPoint + for _, tier := range updateTierOrder { + p, ok := m.updateTierPoint(ctx, stackName, tier) + if !ok { + continue + } + seen = append(seen, p) + if accept == nil || accept(p) { + return p, true, seen + } + } + return UpdateTierPoint{}, false, seen +} + +func (m *Manager) updateTierPoint(ctx context.Context, stackName string, tier int) (UpdateTierPoint, bool) { + switch tier { + case UpdateTierSecondDrive: + get := m.updateTier2PointFn + if get == nil { + get = m.Tier2UnitRestorePoint + } + rp, err := get(stackName) + if err != nil { + if m.isDebug() { + m.logger.Printf("[DEBUG] [backup] update precondition for %s: no Tier-2 copy (%v)", stackName, err) + } + return UpdateTierPoint{}, false + } + at, ok := rp.ProvenCopyTime() + return UpdateTierPoint{Tier: tier, At: at}, ok + case UpdateTierLocal: + list := m.updateTier1PointsFn + if list == nil { + list = m.ListRestorePoints + } + pts, _ := list(stackName) + for _, rp := range pts { + if at, err := time.Parse(time.RFC3339, rp.Time); err == nil { + return UpdateTierPoint{Tier: tier, At: at}, true + } + } + return UpdateTierPoint{}, false + case UpdateTierOffsite: + inv := m.updateOffsiteInvFn + if inv == nil { + if m.settings == nil || !m.OffboxConfigured() { + return UpdateTierPoint{}, false + } + inv = m.OffsiteInventoryList + } + cctx, cancel := context.WithTimeout(ctx, updateOffsiteCheckTimeout) + defer cancel() + got, err := inv(cctx) + if err != nil { + if !errors.Is(err, errNoOffsiteTarget) { + m.logger.Printf("[WARN] [backup] update precondition for %s: the off-site copy could not be checked within %s (%v) — counted as ABSENT", stackName, updateOffsiteCheckTimeout, err) + } + return UpdateTierPoint{}, false + } + for _, a := range got.Apps { + if a.App == stackName && !a.LatestAt.IsZero() { + return UpdateTierPoint{Tier: tier, At: a.LatestAt}, true + } + } + } + return UpdateTierPoint{}, false +} + +// CanBackUpApp reports whether "back up first" can run for this app at all right now — the second +// half of R-475 Scenario L: an app with no copy anywhere is refused only when this is false too. +// Cheap and read-only; RunAppBackupNow re-checks everything when it actually runs. +func (m *Manager) CanBackUpApp(stackName string) (bool, string) { + if m == nil { + return false, "backup is not enabled on this box" + } + if m.stackProvider == nil { + return false, "stack provider not configured" + } + if m.migrationActive() { + return false, "a data migration is running" + } + drivePath := m.GetAppDrivePath(stackName) + if drivePath == "" || !filepath.IsAbs(drivePath) { + return false, "the app's drive cannot be resolved" + } + if m.settings != nil && (m.settings.IsDisconnected(drivePath) || m.settings.IsDecommissioned(drivePath)) { + return false, fmt.Sprintf("the app's drive %s is not available", drivePath) + } + return true, "" +} + // UpdateBusy reports whether something else is ALREADY touching this app's data, which refuses an // update before anything moves (slice 4 Scenario D). The reason is operator-English; the customer // sentence is chosen by the caller. @@ -142,6 +286,7 @@ func (m *Manager) RunAppBackupNow(ctx context.Context, stackName string) error { } m.logger.Printf("[INFO] [backup] update pre-backup for %s: starting (DB dump → volume dump → unit capture → Tier 2)", stackName) start := time.Now() + var nsRoot string legErr := func() error { defer m.releaseRunning() defer m.beginAdmissionRun()() @@ -156,7 +301,7 @@ func (m *Manager) RunAppBackupNow(ctx context.Context, stackName string) error { if !m.admitApp(stackName) { return fmt.Errorf("nincs elég szabad hely a mentéshez a(z) %s meghajtón", drivePath) } - nsRoot := m.namespaceRoot(drivePath) + nsRoot = m.namespaceRoot(drivePath) discover := m.discoverDBs if discover == nil { @@ -203,16 +348,35 @@ func (m *Manager) RunAppBackupNow(ctx context.Context, stackName string) error { return legErr } + m.updatePreBackupTail(stackName, nsRoot, time.Now()) + m.logger.Printf("[INFO] [backup] update pre-backup for %s: complete in %s", stackName, time.Since(start).Round(time.Millisecond)) + return nil +} + +// updatePreBackupTail is what "back up first" does after the capture succeeded (R-475). +// +// 1. It marks the app's OWN unit as proven current NOW. CaptureRecoveryUnit leaves the manifest alone +// when nothing changed (the checksum skip), and Tier 1's age is the newest artifact's mtime — so on +// an app with no database and no named volume a fresh "back up first" would leave Tier 1 as old as +// its last definition change, and the update would be refused forever. That is the trap +// ProvenCopyTime documents for Tier 2, one tier down. The capture has just compared the unit with +// the live definition, so "current as of now" is exactly what it established. +// 2. It runs the Tier-2 copy, and a Tier-2 failure is a WARN, not a failure: the update may lean on +// any tier, and the Tier-1 unit it can lean on was just written. Pinned by +// TestR475_PreBackupTail_Tier2FailureIsAWarnAndTheOwnUnitIsFresh. +func (m *Manager) updatePreBackupTail(stackName, nsRoot string, now time.Time) { + if nsRoot != "" { + if err := os.Chtimes(RecoveryUnitManifestPath(nsRoot, stackName), now, now); err != nil { + m.logger.Printf("[WARN] [backup] update pre-backup for %s: could not mark the recovery unit as proven current (%v) — its own-unit copy may read older than it is", stackName, err) + } + } runOne := m.perAppTier2 if runOne == nil { runOne = m.RunTier2 } if err := runOne(stackName); err != nil { - m.logger.Printf("[ERROR] [backup] update pre-backup for %s: Tier 2 copy FAILED: %v", stackName, err) - return fmt.Errorf("a másodlagos másolat elkészítése sikertelen: %w", err) + m.logger.Printf("[WARN] [backup] update pre-backup for %s: Tier 2 copy FAILED: %v — not fatal: the app's own recovery unit was just captured, and an update may lean on any tier (R-475)", stackName, err) } - m.logger.Printf("[INFO] [backup] update pre-backup for %s: complete in %s", stackName, time.Since(start).Round(time.Millisecond)) - return nil } // WriteUpdateSafetyDump takes the last-minute database copy an update makes just before it moves the @@ -244,9 +408,18 @@ func (m *Manager) WriteUpdateSafetyDump(ctx context.Context, stackName string) ( } // UpdateHoldFmt is the customer sentence for an app held after a failed update. Arguments: the app, -// the time of the failure, and the PROVEN date of the copy it can be restored from. One named string so -// a test asserts it verbatim instead of retyping Hungarian (R-364). +// the time of the failure, the TIER of the copy it can be restored from (UpdateTierLabel), and that +// copy's PROVEN date. One named string so a test asserts it verbatim instead of retyping Hungarian +// (R-364). Since v0.239.0 (R-475) it names the tier: the copy may be on any of three, and each is +// restored from a different place on the Mentések page. const UpdateHoldFmt = "A(z) %s frissítése %s-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. " + + "Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. " + + "Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: %s, %s." + +// UpdateHoldLegacyFmt is the v0.237.0–v0.238.1 sentence, kept for a hold written before the tier was +// recorded (CopyTier 0) — every such hold named a Tier-2 copy, but it did not SAY so, and rewriting +// it now would state a fact the record does not hold. +const UpdateHoldLegacyFmt = "A(z) %s frissítése %s-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. " + "Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. " + "Visszaállítható a(z) %s-i biztonsági mentésből a Mentések oldalon." @@ -276,7 +449,7 @@ func fmtHoldTime(rfc3339 string) string { // has just stopped the app on the strength of this record, and an unrecorded hold is a stopped app // that the next restart button will quietly start again. The caller logs it at ERROR and keeps the // failure on the page. -func (m *Manager) HoldAfterFailedUpdate(stackName string, at time.Time, copyDate time.Time) error { +func (m *Manager) HoldAfterFailedUpdate(stackName string, at time.Time, copyDate time.Time, copyTier int) error { if m == nil || m.settings == nil { return fmt.Errorf("no settings wired — the update hold for %s cannot be persisted", stackName) } @@ -287,11 +460,12 @@ func (m *Manager) HoldAfterFailedUpdate(stackName string, at time.Time, copyDate } if !copyDate.IsZero() { h.CopyDate = copyDate.UTC().Format(time.RFC3339) + h.CopyTier = copyTier } if err := m.settings.SetRestoreHold(h); err != nil { return fmt.Errorf("persisting the update hold for %s: %w", stackName, err) } - m.logger.Printf("[WARN] [backup] %s is HELD STOPPED after a failed update (restore point: %s)", stackName, h.CopyDate) + m.logger.Printf("[WARN] [backup] %s is HELD STOPPED after a failed update (restore point: tier %d %q, %s)", stackName, h.CopyTier, UpdateTierLabel(h.CopyTier), h.CopyDate) return nil } diff --git a/controller/internal/backup/update_tiers_test.go b/controller/internal/backup/update_tiers_test.go new file mode 100644 index 0000000..f0f9d03 --- /dev/null +++ b/controller/internal/backup/update_tiers_test.go @@ -0,0 +1,317 @@ +package backup + +import ( + "bytes" + "context" + "errors" + "fmt" + "go/ast" + "go/parser" + "go/token" + "io" + "log" + "os" + "path/filepath" + "strings" + "testing" + "time" + + "gitea.dooplex.hu/admin/felhom-controller/internal/settings" +) + +// R-475 — the backup side of "any backup tier lets an app update" (controller v0.239.0). + +var r475T0 = time.Date(2026, 9, 13, 12, 0, 0, 0, time.UTC) + +func r475Manager() (*Manager, *bytes.Buffer) { + var buf bytes.Buffer + return &Manager{logger: log.New(&buf, "", 0)}, &buf +} + +func tier2At(at time.Time) func(string) (Tier2RestorePoint, error) { + return func(string) (Tier2RestorePoint, error) { + ts := at.UTC().Format(time.RFC3339) + return Tier2RestorePoint{Restorable: true, CopyDateProven: true, CopyLastSuccess: ts, CopyDate: ts}, nil + } +} +func noTier2(string) (Tier2RestorePoint, error) { + return Tier2RestorePoint{}, errors.New("no Tier-2 copy recorded for this app") +} +func tier1At(at time.Time) func(string) ([]RestorePoint, bool) { + return func(string) ([]RestorePoint, bool) { + return []RestorePoint{{Time: at.UTC().Format(time.RFC3339), ShortID: "helyi", Tier: 1}}, true + } +} +func noTier1(string) ([]RestorePoint, bool) { return []RestorePoint{}, true } +func offsiteWith(app string, at time.Time) func(context.Context) (OffsiteInventory, error) { + return func(context.Context) (OffsiteInventory, error) { + return OffsiteInventory{Apps: []OffsiteInventoryApp{{App: app, LatestAt: at}}}, nil + } +} +func noOffsiteTarget(context.Context) (OffsiteInventory, error) { + return OffsiteInventory{}, errNoOffsiteTarget +} + +func tiersOf(ps []UpdateTierPoint) []int { + out := make([]int, 0, len(ps)) + for _, p := range ps { + out = append(out, p.Tier) + } + return out +} + +// G — the order is 2, 1, 3, and the walk STOPS at the first accepted copy (so a box with a fresh +// second-drive copy never reaches the network). +func TestR475_G_TierOrderIsSecondDriveThenOwnUnitThenOffsite(t *testing.T) { + m, _ := r475Manager() + offsiteCalls := 0 + m.updateTier2PointFn = tier2At(r475T0.Add(-1 * time.Hour)) + m.updateTier1PointsFn = tier1At(r475T0.Add(-2 * time.Hour)) + m.updateOffsiteInvFn = func(ctx context.Context) (OffsiteInventory, error) { + offsiteCalls++ + return offsiteWith("gokapi", r475T0.Add(-3*time.Hour))(ctx) + } + ctx := context.Background() + + p, ok, seen := m.UpdateRestorePoints(ctx, "gokapi", nil) + if !ok || p.Tier != UpdateTierSecondDrive || fmt.Sprint(tiersOf(seen)) != "[2]" || offsiteCalls != 0 { + t.Errorf("all three present: want tier 2 and a stop; got %+v ok=%v seen=%v offsite calls=%d", p, ok, tiersOf(seen), offsiteCalls) + } + p, ok, seen = m.UpdateRestorePoints(ctx, "gokapi", func(p UpdateTierPoint) bool { return p.Tier != UpdateTierSecondDrive }) + if !ok || p.Tier != UpdateTierLocal || fmt.Sprint(tiersOf(seen)) != "[2 1]" { + t.Errorf("tier 2 refused: want tier 1; got %+v seen=%v", p, tiersOf(seen)) + } + p, ok, seen = m.UpdateRestorePoints(ctx, "gokapi", func(p UpdateTierPoint) bool { return p.Tier == UpdateTierOffsite }) + if !ok || p.Tier != UpdateTierOffsite || !p.At.Equal(r475T0.Add(-3*time.Hour)) || fmt.Sprint(tiersOf(seen)) != "[2 1 3]" { + t.Errorf("tiers 2 and 1 refused: want tier 3; got %+v seen=%v", p, tiersOf(seen)) + } + if _, ok, seen = m.UpdateRestorePoints(ctx, "gokapi", func(UpdateTierPoint) bool { return false }); ok || len(seen) != 3 { + t.Errorf("nothing accepted: want not found with all three seen; ok=%v seen=%v", ok, tiersOf(seen)) + } +} + +func TestR475_H_OwnUnitOnly(t *testing.T) { + m, _ := r475Manager() + m.updateTier2PointFn = noTier2 + m.updateTier1PointsFn = tier1At(r475T0.Add(-2 * time.Hour)) + m.updateOffsiteInvFn = noOffsiteTarget + p, ok, _ := m.UpdateRestorePoints(context.Background(), "gokapi", nil) + if !ok || p.Tier != UpdateTierLocal || !p.At.Equal(r475T0.Add(-2*time.Hour)) { + t.Fatalf("got %+v ok=%v", p, ok) + } + // A Tier-2 record that was only ATTEMPTED is not a copy (R-101) — the own unit still wins. + m.updateTier2PointFn = func(string) (Tier2RestorePoint, error) { + return Tier2RestorePoint{Restorable: true, CopyDateProven: false}, nil + } + if p, ok, _ = m.UpdateRestorePoints(context.Background(), "gokapi", nil); !ok || p.Tier != UpdateTierLocal { + t.Errorf("an unproven Tier-2 record must be skipped; got %+v", p) + } + // No unit on disk (ListRestorePoints' empty list) is no copy. + m.updateTier1PointsFn = noTier1 + if _, ok, _ = m.UpdateRestorePoints(context.Background(), "gokapi", nil); ok { + t.Error("no copy on any tier must be not found") + } +} + +func TestR475_I_OffsiteOnly(t *testing.T) { + m, _ := r475Manager() + m.updateTier2PointFn, m.updateTier1PointsFn = noTier2, noTier1 + m.updateOffsiteInvFn = offsiteWith("gokapi", r475T0.Add(-5*time.Hour)) + if p, ok, _ := m.UpdateRestorePoints(context.Background(), "gokapi", nil); !ok || p.Tier != UpdateTierOffsite { + t.Fatalf("got %+v ok=%v", p, ok) + } + // Another app's snapshot is not this app's copy. + if _, ok, _ := m.UpdateRestorePoints(context.Background(), "nextcloud", nil); ok { + t.Error("a snapshot tagged for a different app must not count") + } +} + +// J — an unreachable off-site repository is ABSENT with a WARN, and it is bounded in time. +func TestR475_J_OffsiteUnreachableIsAbsentWithAWarn(t *testing.T) { + m, buf := r475Manager() + m.updateTier2PointFn, m.updateTier1PointsFn = noTier2, noTier1 + m.updateOffsiteInvFn = func(context.Context) (OffsiteInventory, error) { + return OffsiteInventory{}, errors.New("ssh: connect to host: connection timed out") + } + if _, ok, _ := m.UpdateRestorePoints(context.Background(), "gokapi", nil); ok { + t.Error("an unreachable off-site copy must count as absent") + } + if !strings.Contains(buf.String(), "[WARN]") || !strings.Contains(buf.String(), "counted as ABSENT") { + t.Errorf("an unreachable off-site copy must WARN; log = %q", buf.String()) + } + + // Control: a box with NO off-site target is plainly absent — that is not a fault, so no WARN. + buf.Reset() + m.updateOffsiteInvFn = noOffsiteTarget + if _, ok, _ := m.UpdateRestorePoints(context.Background(), "gokapi", nil); ok || strings.Contains(buf.String(), "WARN") { + t.Errorf("no off-site target: want absent and silent; ok=%v log=%q", ok, buf.String()) + } + + // The bound: a repository that never answers is given up on at the timeout. + old := updateOffsiteCheckTimeout + updateOffsiteCheckTimeout = 50 * time.Millisecond + defer func() { updateOffsiteCheckTimeout = old }() + buf.Reset() + m.updateOffsiteInvFn = func(ctx context.Context) (OffsiteInventory, error) { + <-ctx.Done() + return OffsiteInventory{}, ctx.Err() + } + start := time.Now() + _, ok, _ := m.UpdateRestorePoints(context.Background(), "gokapi", nil) + if took := time.Since(start); ok || took > 5*time.Second || !strings.Contains(buf.String(), "counted as ABSENT") { + t.Errorf("a hanging off-site check must end at its bound as absent with a WARN; ok=%v took=%s log=%q", ok, took, buf.String()) + } +} + +// The hold names the tier. The labels are the ruling's exact words. +func TestR475_HoldTextNamesTheTier(t *testing.T) { + at := time.Date(2026, 9, 13, 8, 0, 0, 0, time.UTC) + copyAt := time.Date(2026, 9, 13, 1, 30, 0, 0, time.UTC) + for tier, label := range map[int]string{ + UpdateTierSecondDrive: "második meghajtó", + UpdateTierLocal: "saját meghajtó", + UpdateTierOffsite: "távoli mentés", + } { + sett := slice4Settings(t) + m := &Manager{logger: log.New(io.Discard, "", 0), settings: sett} + if err := m.HoldAfterFailedUpdate("gokapi", at, copyAt, tier); err != nil { + t.Fatal(err) + } + held, why := m.RestoreHoldFor("gokapi") + want := fmt.Sprintf(UpdateHoldFmt, "gokapi", "2026-09-13 10:00", label, "2026-09-13 03:30") + if !held || why != want { + t.Errorf("tier %d: hold text =\n%q\nwant\n%q", tier, why, want) + } + if !strings.HasSuffix(why, "ebből a biztonsági mentésből: "+label+", 2026-09-13 03:30.") { + t.Errorf("tier %d: the sentence must END naming the tier and the date, got %q", tier, why) + } + } +} + +func TestR475_AHoldWrittenBeforeTheTierKeepsItsSentence(t *testing.T) { + sett := slice4Settings(t) + m := &Manager{logger: log.New(io.Discard, "", 0), settings: sett} + if err := sett.SetRestoreHold(settings.RestoreHold{Stack: "uptime-kuma", At: "2026-09-13T10:18:02Z", + Reason: settings.HoldReasonUpdateFailed, CopyDate: "2026-09-13T10:09:51Z"}); err != nil { + t.Fatal(err) + } + _, why := m.RestoreHoldFor("uptime-kuma") + // The exact sentence quoted in audits/slice4-2026-09-13 (13-H-refusals.txt), live on v0.238.1. + if want := fmt.Sprintf(UpdateHoldLegacyFmt, "uptime-kuma", "2026-09-13 12:18", "2026-09-13 12:09"); why != want { + t.Errorf("a v0.238.1 hold must keep its sentence:\n%q\nwant\n%q", why, want) + } +} + +// RunAppBackupNow's tail: a Tier-2 failure no longer fails "back up first", and the own unit it +// captured reads as fresh even when the capture found nothing to rewrite. +// +// COMPANION RED-PROOF (REPORT.md): aim the Chtimes at a path that does not exist — this test fails on +// the mtime: "back up first" would then leave a quiet app's own unit as old as its last definition change. +func TestR475_PreBackupTail_Tier2FailureIsAWarnAndTheOwnUnitIsFresh(t *testing.T) { + nsRoot := t.TempDir() + mp := RecoveryUnitManifestPath(nsRoot, "gokapi") + if err := os.MkdirAll(filepath.Dir(mp), 0o755); err != nil { + t.Fatal(err) + } + if err := os.WriteFile(mp, []byte(`{"app_name":"gokapi"}`), 0o644); err != nil { + t.Fatal(err) + } + old := time.Now().Add(-30 * time.Hour) + if err := os.Chtimes(mp, old, old); err != nil { + t.Fatal(err) + } + m, buf := r475Manager() + tier2Ran := false + m.perAppTier2 = func(string) error { tier2Ran = true; return errors.New("no second drive with room") } + + now := time.Now() + m.updatePreBackupTail("gokapi", nsRoot, now) + + if !tier2Ran { + t.Fatal("the Tier-2 copy must still be attempted") + } + if !strings.Contains(buf.String(), "[WARN]") || !strings.Contains(buf.String(), "Tier 2 copy FAILED") { + t.Errorf("a Tier-2 failure must be logged as a WARN; log = %q", buf.String()) + } + fi, err := os.Stat(mp) + if err != nil { + t.Fatal(err) + } + if d := fi.ModTime().Sub(now); d < -time.Second || d > time.Second { + t.Errorf("the own unit must read as proven NOW (mtime %s, now %s)", fi.ModTime(), now) + } + + // Control: the same unit through ListRestorePoints' own rule reads fresh — the newest artifact. + buf.Reset() + m.perAppTier2 = func(string) error { return nil } + m.updatePreBackupTail("gokapi", nsRoot, now) + if strings.Contains(buf.String(), "WARN") { + t.Errorf("a successful Tier-2 copy must not WARN; log = %q", buf.String()) + } +} + +// The tail is really what RunAppBackupNow runs — a helper tested and never called is the "seam built +// but never wired" shape this project has shipped repeatedly. +func TestR475_RunAppBackupNowUsesTheTolerantTail(t *testing.T) { + fset := token.NewFileSet() + f, err := parser.ParseFile(fset, "update_guard.go", nil, 0) + if err != nil { + t.Fatal(err) + } + var calls []string + found := false + for _, d := range f.Decls { + fn, ok := d.(*ast.FuncDecl) + if !ok || fn.Name.Name != "RunAppBackupNow" { + continue + } + found = true + ast.Inspect(fn.Body, func(n ast.Node) bool { + if sel, ok := n.(*ast.SelectorExpr); ok { + calls = append(calls, sel.Sel.Name) + } + return true + }) + } + joined := " " + strings.Join(calls, " ") + " " + if !found || !strings.Contains(joined, " updatePreBackupTail ") { + t.Fatalf("RunAppBackupNow must end in updatePreBackupTail; selectors = %s", joined) + } + if strings.Contains(joined, " RunTier2 ") || strings.Contains(joined, " perAppTier2 ") { + t.Error("RunAppBackupNow must not run the Tier-2 copy itself any more — only through the tolerant tail") + } +} + +func TestR475_CanBackUpApp(t *testing.T) { + var nilM *Manager + if ok, why := nilM.CanBackUpApp("gokapi"); ok || why == "" { + t.Error("no backup manager: cannot back up, with a reason") + } + m, _ := r475Manager() + if ok, why := m.CanBackUpApp("gokapi"); ok || !strings.Contains(why, "stack provider") { + t.Errorf("no stack provider: cannot back up; got ok=%v why=%q", ok, why) + } +} + +// Tier 3's route back is the off-site restore, and it never went through RestoreFromRecoveryUnitAt — +// so it must lift an update hold itself, or a hold naming „távoli mentés" could never be cleared. +// +// COMPANION RED-PROOF (REPORT.md): remove the clearUpdateHoldAfterRestore call from +// ReconstituteFromOffsite — this test fails with the hold still in place. +func TestR475_OffsiteRestoreClearsAnUpdateHold(t *testing.T) { + m, prov, _ := reconFixture(t, "run1", "2026-07-19T06:00:00Z", "") + prov.composePath = writeLiveCompose(t, noDBCompose) + m.discoverDBs = func(context.Context) ([]DiscoveredDB, error) { return nil, nil } + if err := m.HoldAfterFailedUpdate("immich", r475T0, r475T0.Add(-time.Hour), UpdateTierOffsite); err != nil { + t.Fatal(err) + } + if held, why := m.RestoreHoldFor("immich"); !held || !strings.Contains(why, "távoli mentés") { + t.Fatalf("fixture: the app must be held naming the off-site copy, got held=%v %q", held, why) + } + if _, err := m.ReconstituteFromOffsite(context.Background(), "immich", false); err != nil { + t.Fatalf("reconstitute: %v", err) + } + if held, _ := m.RestoreHoldFor("immich"); held { + t.Error("a successful off-site restore is the route back the hold names — the hold must be lifted") + } +} diff --git a/controller/internal/settings/settings.go b/controller/internal/settings/settings.go index a31bce0..1b427f5 100644 --- a/controller/internal/settings/settings.go +++ b/controller/internal/settings/settings.go @@ -1598,6 +1598,10 @@ type RestoreHold struct { // CopyDate is the RFC3339 time of the proven backup the customer is told they can restore from. // Only set for HoldReasonUpdateFailed. CopyDate string `json:"copy_date,omitempty"` + // CopyTier (R-475, v0.239.0) is WHICH backup tier CopyDate belongs to: 1 the app's own recovery + // unit, 2 the second drive, 3 off-site. 0 means a hold written before the field existed; it keeps + // its original sentence (backup.UpdateHoldLegacyFmt). Only set for HoldReasonUpdateFailed. + CopyTier int `json:"copy_tier,omitempty"` } // Hold reasons. See RestoreHold.Reason. diff --git a/controller/internal/stacks/update.go b/controller/internal/stacks/update.go index 67633ee..d05195c 100644 --- a/controller/internal/stacks/update.go +++ b/controller/internal/stacks/update.go @@ -6,6 +6,7 @@ import ( "fmt" "os" "path/filepath" + "strings" "time" "gitea.dooplex.hu/admin/felhom-controller/internal/system" @@ -82,7 +83,7 @@ const ( MsgUpdateAlreadyFmt = "A(z) %s frissítése már folyamatban van." MsgUpdateBusy = "A frissítés most nem indítható: mentés/visszaállítás folyamatban. Próbáld újra, ha befejeződött." MsgUpdateMigrating = "A frissítés most nem indítható: adatáthelyezés folyamatban." - MsgUpdateNoBackupFmt = "A(z) %s nem frissíthető, mert nincs olyan biztonsági mentése, amelyből vissza lehetne állítani. Kapcsold be a 2. mentést az alkalmazás mentési beállításainál a Mentések oldalon, és várd meg az első sikeres másolatot — utána a frissítés elindítható." + MsgUpdateNoBackupFmt = "A(z) %s nem frissíthető, mert nincs olyan biztonsági mentése, amelyből vissza lehetne állítani, és most új mentés sem készíthető róla. Ellenőrizd a Mentések oldalon, hogy az alkalmazás meghajtója elérhető-e — utána a frissítés elindítható." MsgUpdateDiskFmt = "Nincs elég szabad hely a frissítéshez: %.1f GB szabad, az új verzió letöltéséhez legalább %.0f GB szükséges." MsgUpdateBackupFailFmt = "A frissítés nem indult el, mert a frissítés előtti biztonsági mentés nem sikerült: %v. Az alkalmazás változatlanul fut tovább." MsgUpdateBackupNoUnit = "A frissítés nem indult el: a frissítés előtti mentés lefutott, de nem jött létre friss, visszaállítható másolat. Az alkalmazás változatlanul fut tovább." @@ -107,11 +108,55 @@ const updateSettleWindow = 60 * time.Second // updatePollEvery is how often the health wait re-reads the stack. const updatePollEvery = 5 * time.Second -// UpdateRestorePoint is the precondition answer, reduced to what the update needs. +// Backup tiers (R-475), mirroring backup.UpdateTier* — stacks cannot import backup, so +// TestR475_TierConstantsAgree (cmd/controller) pins the two sets equal. +const ( + UpdateTierLocal = 1 // the app's own recovery unit + UpdateTierSecondDrive = 2 // the Tier-2 mirror on another drive + UpdateTierOffsite = 3 // off-site +) + +// UpdateRestorePoint is one proven, restorable copy the update may lean on. The backup side returns +// only proven, restorable copies (never an attempt clock, never an unopenable unit), so there is no +// "maybe" field here to forget to check. type UpdateRestorePoint struct { - Restorable bool // an openable recovery unit exists in the Tier-2 copy - Proven bool // a copy actually succeeded (never an attempt clock) - ProvenAt time.Time // when the data in that copy was last proven copied + Tier int // UpdateTierSecondDrive / UpdateTierLocal / UpdateTierOffsite + ProvenAt time.Time // when the data in that copy was last proven written +} + +func updateTierName(tier int) string { + switch tier { + case UpdateTierSecondDrive: + return "Tier 2 (second drive)" + case UpdateTierLocal: + return "Tier 1 (own recovery unit)" + case UpdateTierOffsite: + return "Tier 3 (off-site)" + } + return fmt.Sprintf("tier %d", tier) +} + +// freshRestorePoint is THE age rule (backup_max_age), applied to whichever tier is being considered +// — R-475 Scenario M: a stale copy on ANY tier is stale. +// +// COMPANION RED-PROOF M (REPORT.md): check the age only for Tier 2 (let any other tier through +// whatever its age). TestR475_M_TheAgeRuleAppliesToTheChosenTier then fails: a 30-hour-old copy of +// the app's own unit carries the update with no backup first. +func freshRestorePoint(now time.Time, maxAge time.Duration) func(UpdateRestorePoint) bool { + return func(p UpdateRestorePoint) bool { + return !p.ProvenAt.IsZero() && now.Sub(p.ProvenAt) <= maxAge + } +} + +func describeRestorePoints(now time.Time, pts []UpdateRestorePoint) string { + if len(pts) == 0 { + return "none" + } + parts := make([]string, 0, len(pts)) + for _, p := range pts { + parts = append(parts, fmt.Sprintf("%s at %s (%s old)", updateTierName(p.Tier), p.ProvenAt.UTC().Format(time.RFC3339), now.Sub(p.ProvenAt).Round(time.Minute))) + } + return strings.Join(parts, "; ") } // UpdateGuards is everything the update needs from the backup side. The stacks package cannot import @@ -119,10 +164,14 @@ type UpdateRestorePoint struct { type UpdateGuards interface { HoldFor(name string) (bool, string) Busy(name string) (bool, string) - RestorePoint(name string) (UpdateRestorePoint, error) + // RestorePoints walks the tiers in preference order (2, 1, 3) and returns the first copy accept + // admits (nil = any), whether one was found, and every copy looked at (R-475). + RestorePoints(ctx context.Context, name string, accept func(UpdateRestorePoint) bool) (UpdateRestorePoint, bool, []UpdateRestorePoint) + // CanBackUp reports whether "back up first" can run for this app now (R-475 Scenario L). + CanBackUp(name string) (bool, string) BackupNow(ctx context.Context, name string) error SafetyDump(ctx context.Context, name string) ([]string, error) - HoldAfterFailedUpdate(name string, at, provenCopyAt time.Time) error + HoldAfterFailedUpdate(name string, at time.Time, rp UpdateRestorePoint) error } // SetUpdateGuards wires the backup side. INIT-ONLY. Unwired, every update is refused (fail closed): @@ -199,10 +248,19 @@ func (m *Manager) UpdatePreflight(name string) *UpdateRefusal { if m.IsMigrating() { return m.refuseUpdate(name, "migrating", MsgUpdateMigrating, "a data migration is running") } - rp, err := g.RestorePoint(name) - if err != nil || !rp.Restorable || !rp.Proven { - return m.refuseUpdate(name, "no_backup", fmt.Sprintf(MsgUpdateNoBackupFmt, name), - fmt.Sprintf("no restorable proven Tier-2 unit (restorable=%v proven=%v err=%v)", rp.Restorable, rp.Proven, err)) + // R-475: any tier counts, and an app with no copy at all is backed up first by the job. So the only + // refusal left here is Scenario L — no copy on any tier AND no way to make one now. (With a copy + // but no way to back up, the job still applies the age rule and refuses then if the copy is stale.) + // Tier 3 is looked at only on this branch, so an ordinary update never waits on the network here. + // + // COMPANION RED-PROOF (REPORT.md): drop the `!found` condition. TestR475_L then fails — an app + // with a copy but no way to back up is refused too. + if canBackUp, why := g.CanBackUp(name); !canBackUp { + if _, found, seen := g.RestorePoints(context.Background(), name, nil); !found { + return m.refuseUpdate(name, "no_backup", fmt.Sprintf(MsgUpdateNoBackupFmt, name), + fmt.Sprintf("no copy on any tier (found: %s) and no backup can be taken now: %s", describeRestorePoints(m.now(), seen), why)) + } + m.logger.Printf("[WARN] [stacks] update %s: no backup can be taken now (%s) — an existing copy must carry the update", name, why) } if ref := m.updateMemoryRefusal(name, st); ref != nil { return ref @@ -383,15 +441,15 @@ func (m *Manager) runGuardedUpdate(ctx context.Context, name string) { fail(MsgUpdateNoGuards, "no UpdateGuards wired") return } - rp, err := g.RestorePoint(name) - if err != nil || !rp.Restorable || !rp.Proven { - fail(fmt.Sprintf(MsgUpdateNoBackupFmt, name), fmt.Sprintf("precondition vanished: restorable=%v proven=%v err=%v", rp.Restorable, rp.Proven, err)) - return - } - + // R-475: the precondition is a copy on ANY tier, chosen in the order 2, 1, 3, and the age rule + // applies to whichever tier is chosen. The first FRESH copy wins — not merely the first copy — so a + // stale second-drive mirror never forces a backup while the app's own unit is minutes old. maxAge := m.backupMaxAge() - if age := start.Sub(rp.ProvenAt); age > maxAge { - m.logger.Printf("[INFO] [stacks] update %s: the proven copy is %s old (limit %s) — backing up first", name, age.Round(time.Minute), maxAge) + rp, ok, seen := g.RestorePoints(ctx, name, freshRestorePoint(start, maxAge)) + if ok { + m.logger.Printf("[INFO] [stacks] update %s: precondition met — %s copy from %s (%s old, limit %s)", name, updateTierName(rp.Tier), rp.ProvenAt.UTC().Format(time.RFC3339), start.Sub(rp.ProvenAt).Round(time.Minute), maxAge) + } else { + m.logger.Printf("[INFO] [stacks] update %s: no copy younger than %s on any tier (found: %s) — backing up first", name, maxAge, describeRestorePoints(start, seen)) if !m.enterUpdatePhase(name, &entry, UpdatePhaseBackingUp) { fail(MsgUpdateJournalFailed, "journal write failed") return @@ -400,15 +458,16 @@ func (m *Manager) runGuardedUpdate(ctx context.Context, name string) { fail(fmt.Sprintf(MsgUpdateBackupFailFmt, err), "pre-update backup: "+err.Error()) return } - rp, err = g.RestorePoint(name) - if err != nil || !rp.Restorable || !rp.Proven || m.now().Sub(rp.ProvenAt) > maxAge { - fail(MsgUpdateBackupNoUnit, fmt.Sprintf("after the backup: restorable=%v proven=%v at=%s err=%v", rp.Restorable, rp.Proven, rp.ProvenAt.Format(time.RFC3339), err)) + now := m.now() + rp, ok, seen = g.RestorePoints(ctx, name, freshRestorePoint(now, maxAge)) + if !ok { + fail(MsgUpdateBackupNoUnit, fmt.Sprintf("after the backup there is still no copy younger than %s on any tier (found: %s)", maxAge, describeRestorePoints(now, seen))) return } - } else { - m.logger.Printf("[INFO] [stacks] update %s: precondition met — proven copy from %s (%s old, limit %s)", name, rp.ProvenAt.UTC().Format(time.RFC3339), age.Round(time.Minute), maxAge) + m.logger.Printf("[INFO] [stacks] update %s: precondition met after the backup — %s copy from %s", name, updateTierName(rp.Tier), rp.ProvenAt.UTC().Format(time.RFC3339)) } entry.ProvenCopyAt = rp.ProvenAt.UTC().Format(time.RFC3339) + entry.ProvenTier = rp.Tier // SAFETY DUMP BEFORE THE PIN MOVES — "a minute ago", before any migration can have run. if !m.enterUpdatePhase(name, &entry, UpdatePhaseSafetyDump) { @@ -472,20 +531,20 @@ func (m *Manager) runGuardedUpdate(ctx context.Context, name string) { } if !m.enterUpdatePhase(name, &entry, UpdatePhaseStarting) { - m.failAndHold(ctx, name, dir, env, rp.ProvenAt, "journal write failed before up") + m.failAndHold(ctx, name, dir, env, rp, "journal write failed before up") return } if _, err := m.updateCompose(dir, env, "up", "-d", "--remove-orphans"); err != nil { // Containers may already have been recreated on the new image — something may have run. - m.failAndHold(ctx, name, dir, env, rp.ProvenAt, "compose up failed: "+err.Error()) + m.failAndHold(ctx, name, dir, env, rp, "compose up failed: "+err.Error()) return } - m.verifyAndConclude(ctx, name, dir, env, rp.ProvenAt, start, &entry) + m.verifyAndConclude(ctx, name, dir, env, rp, start, &entry) } // verifyAndConclude is the TRUTH half (R-443): success is declared only after the app's health is // known, and a failure holds the app. -func (m *Manager) verifyAndConclude(ctx context.Context, name, dir string, env []string, provenAt, start time.Time, entry *updateJournalEntry) { +func (m *Manager) verifyAndConclude(ctx context.Context, name, dir string, env []string, rp UpdateRestorePoint, start time.Time, entry *updateJournalEntry) { if !m.enterUpdatePhase(name, entry, UpdatePhaseVerifying) { m.logger.Printf("[ERROR] [stacks] update %s: could not journal the verifying phase — verifying anyway", name) } @@ -493,7 +552,7 @@ func (m *Manager) verifyAndConclude(ctx context.Context, name, dir string, env [ waitStart := m.now() healthy, detail := m.updateHealth(ctx, name, timeout) if !healthy { - m.failAndHold(ctx, name, dir, env, provenAt, "not healthy: "+detail) + m.failAndHold(ctx, name, dir, env, rp, "not healthy: "+detail) return } m.logger.Printf("[INFO] [stacks] update %s: healthy after %s (%s)", name, m.now().Sub(waitStart).Round(time.Second), detail) @@ -506,7 +565,7 @@ func (m *Manager) verifyAndConclude(ctx context.Context, name, dir string, env [ } // failAndHold is Scenario F: stop the app, record the hold, tell the customer the route back. -func (m *Manager) failAndHold(ctx context.Context, name, dir string, env []string, provenAt time.Time, why string) { +func (m *Manager) failAndHold(ctx context.Context, name, dir string, env []string, rp UpdateRestorePoint, why string) { m.logger.Printf("[ERROR] [stacks] update %s FAILED after the new version was started: %s — stopping and HOLDING the app; the pin stays on the new version (its migration may have run)", name, why) if _, err := m.updateCompose(dir, env, "down"); err != nil { m.logger.Printf("[ERROR] [stacks] update %s: stopping the failed app also failed: %v", name, err) @@ -514,7 +573,7 @@ func (m *Manager) failAndHold(ctx context.Context, name, dir string, env []strin msg := MsgUpdateHoldUnsaved if g := m.guards(); g == nil { m.logger.Printf("[ERROR] [stacks] update %s: no UpdateGuards — the hold CANNOT be recorded", name) - } else if err := g.HoldAfterFailedUpdate(name, m.now(), provenAt); err != nil { + } else if err := g.HoldAfterFailedUpdate(name, m.now(), rp); err != nil { m.logger.Printf("[ERROR] [stacks] update %s: %v", name, err) } else if _, why := g.HoldFor(name); why != "" { msg = why @@ -619,6 +678,9 @@ type updateJournalEntry struct { PrevCompose string `json:"prev_compose,omitempty"` PrevApplied string `json:"prev_applied,omitempty"` ProvenCopyAt string `json:"proven_copy_at,omitempty"` + // ProvenTier (R-475) — which tier ProvenCopyAt belongs to, so a resumed update that fails names + // the right copy. 0 in a journal written by v0.238.1 or older. + ProvenTier int `json:"proven_tier,omitempty"` } type updateJournal struct { @@ -787,16 +849,17 @@ func (m *Manager) ResumeInterruptedUpdates(ctx context.Context) int { continue } provenAt, _ := time.Parse(time.RFC3339, e.ProvenCopyAt) + rp := UpdateRestorePoint{Tier: e.ProvenTier, ProvenAt: provenAt} dir := filepath.Dir(st.ComposePath) - go func(name, dir string, e updateJournalEntry, provenAt time.Time) { + go func(name, dir string, e updateJournalEntry, rp UpdateRestorePoint) { env := m.stackEnv(dir) m.logger.Printf("[INFO] [stacks] update %s: resuming after a controller restart — `up -d` then the health wait", name) if _, err := m.updateCompose(dir, env, "up", "-d", "--remove-orphans"); err != nil { - m.failAndHold(ctx, name, dir, env, provenAt, "resumed compose up failed: "+err.Error()) + m.failAndHold(ctx, name, dir, env, rp, "resumed compose up failed: "+err.Error()) return } - m.verifyAndConclude(ctx, name, dir, env, provenAt, e.StartedAt, &e) - }(name, dir, e, provenAt) + m.verifyAndConclude(ctx, name, dir, env, rp, e.StartedAt, &e) + }(name, dir, e, rp) } return len(names) } diff --git a/controller/internal/stacks/update_test.go b/controller/internal/stacks/update_test.go index 47d0d7f..0e7f884 100644 --- a/controller/internal/stacks/update_test.go +++ b/controller/internal/stacks/update_test.go @@ -25,13 +25,15 @@ type fakeGuards struct { held bool holdWhy string busy bool - rp UpdateRestorePoint - rpErr error - rpAfterBackup *UpdateRestorePoint + // points are the copies the backup side holds, in tier order (R-475); pointsAfterBackup replaces + // them when BackupNow succeeds (nil = the backup changed nothing). + points []UpdateRestorePoint + pointsAfterBackup []UpdateRestorePoint + cannotBackUp bool backupErr error dumpErr error holdErr error - holdProvenAt time.Time + holdRP UpdateRestorePoint pinAtDump string stackDir string } @@ -48,18 +50,29 @@ func (f *fakeGuards) HoldFor(string) (bool, string) { return f.held, f.holdWhy } func (f *fakeGuards) Busy(string) (bool, string) { return f.busy, "fake busy" } -func (f *fakeGuards) RestorePoint(string) (UpdateRestorePoint, error) { - f.note("RestorePoint") +func (f *fakeGuards) RestorePoints(_ context.Context, _ string, accept func(UpdateRestorePoint) bool) (UpdateRestorePoint, bool, []UpdateRestorePoint) { + f.note("RestorePoints") f.mu.Lock() defer f.mu.Unlock() - return f.rp, f.rpErr + var seen []UpdateRestorePoint + for _, p := range f.points { + seen = append(seen, p) + if accept == nil || accept(p) { + return p, true, seen + } + } + return UpdateRestorePoint{}, false, seen +} +func (f *fakeGuards) CanBackUp(string) (bool, string) { + f.note("CanBackUp") + return !f.cannotBackUp, "fake: the drive is gone" } func (f *fakeGuards) BackupNow(context.Context, string) error { f.note("BackupNow") f.mu.Lock() defer f.mu.Unlock() - if f.backupErr == nil && f.rpAfterBackup != nil { - f.rp = *f.rpAfterBackup + if f.backupErr == nil && f.pointsAfterBackup != nil { + f.points = f.pointsAfterBackup } return f.backupErr } @@ -72,14 +85,14 @@ func (f *fakeGuards) SafetyDump(context.Context, string) ([]string, error) { } return []string{"/fake/pre-restore-x.sql"}, f.dumpErr } -func (f *fakeGuards) HoldAfterFailedUpdate(_ string, _ time.Time, provenAt time.Time) error { +func (f *fakeGuards) HoldAfterFailedUpdate(_ string, _ time.Time, rp UpdateRestorePoint) error { f.note("HoldAfterFailedUpdate") f.mu.Lock() defer f.mu.Unlock() if f.holdErr != nil { return f.holdErr } - f.held, f.holdWhy, f.holdProvenAt = true, "HELD-SENTENCE", provenAt + f.held, f.holdWhy, f.holdRP = true, "HELD-SENTENCE", rp return nil } @@ -111,7 +124,7 @@ func newSlice4Manager(t *testing.T) (*Manager, string, *fakeGuards, *composeRec) m, dir := newPinManager(t, pinTplOld, pinTplNew, "deployed: true\nenv: {}\npinned_images:\n web: nextcloud:31.0.14-apache\n") mustWrite(t, AppliedComposePath(dir), pinTplOld) - g := &fakeGuards{rp: UpdateRestorePoint{Restorable: true, Proven: true, ProvenAt: slice4T0.Add(-1 * time.Hour)}, stackDir: dir} + g := &fakeGuards{points: []UpdateRestorePoint{{Tier: UpdateTierSecondDrive, ProvenAt: slice4T0.Add(-1 * time.Hour)}}, stackDir: dir} c := &composeRec{fail: map[string]error{}} m.updateGuards = g m.updateComposeFn = c.fn @@ -212,9 +225,8 @@ func TestSlice4_A_SuccessIsDeclaredOnlyAfterHealth(t *testing.T) { if got, want := strings.Join(c.list(), " | "), "pull | up -d --remove-orphans"; got != want { t.Errorf("compose calls = %q, want %q", got, want) } - // RestorePoint twice by design: once in the preflight (the refusal), once inside the job (the - // precondition must still hold when the job actually starts). - if got := strings.Join(g.callList(), ","); got != "RestorePoint,RestorePoint,SafetyDump" { + // R-475: the preflight asks only whether a backup could be taken; the job reads the copies once. + if got := strings.Join(g.callList(), ","); got != "CanBackUp,RestorePoints,SafetyDump" { t.Errorf("a fresh copy needs no backup-first; guard calls = %s", got) } } @@ -223,8 +235,8 @@ func TestSlice4_A_SuccessIsDeclaredOnlyAfterHealth(t *testing.T) { func TestSlice4_B_StaleCopyIsRefreshedFirst(t *testing.T) { m, dir, g, _ := newSlice4Manager(t) - g.rp.ProvenAt = slice4T0.Add(-30 * time.Hour) // > 24 h default - g.rpAfterBackup = &UpdateRestorePoint{Restorable: true, Proven: true, ProvenAt: slice4T0.Add(-1 * time.Minute)} + g.points[0].ProvenAt = slice4T0.Add(-30 * time.Hour) // > 24 h default + g.pointsAfterBackup = []UpdateRestorePoint{{Tier: UpdateTierSecondDrive, ProvenAt: slice4T0.Add(-1 * time.Minute)}} if err := m.StartGuardedUpdate("nextcloud"); err != nil { t.Fatal(err) } @@ -233,7 +245,7 @@ func TestSlice4_B_StaleCopyIsRefreshedFirst(t *testing.T) { t.Fatalf("with a successful backup-first the update completes, got phase=%q err=%q", st.UpdatePhase, st.UpdateError) } calls := strings.Join(g.callList(), ",") - if !strings.HasPrefix(calls, "RestorePoint,RestorePoint,BackupNow,RestorePoint,SafetyDump") { + if !strings.HasPrefix(calls, "CanBackUp,RestorePoints,BackupNow,RestorePoints,SafetyDump") { t.Errorf("a stale copy must be backed up FIRST and the precondition re-read; calls = %s", calls) } if got := pinOf(t, dir); got != "nextcloud:34.0.1-apache" { @@ -243,7 +255,7 @@ func TestSlice4_B_StaleCopyIsRefreshedFirst(t *testing.T) { func TestSlice4_B_BackupFailureMovesNothing(t *testing.T) { m, dir, g, c := newSlice4Manager(t) - g.rp.ProvenAt = slice4T0.Add(-30 * time.Hour) + g.points[0].ProvenAt = slice4T0.Add(-30 * time.Hour) g.backupErr = errors.New("disk full") if err := m.StartGuardedUpdate("nextcloud"); err != nil { t.Fatal(err) @@ -265,7 +277,7 @@ func TestSlice4_B_BackupFailureMovesNothing(t *testing.T) { func TestSlice4_B_BackupThatYieldsNoFreshUnitRefuses(t *testing.T) { m, dir, g, c := newSlice4Manager(t) - g.rp.ProvenAt = slice4T0.Add(-30 * time.Hour) // stays stale: rpAfterBackup nil + g.points[0].ProvenAt = slice4T0.Add(-30 * time.Hour) // stays stale: pointsAfterBackup nil if err := m.StartGuardedUpdate("nextcloud"); err != nil { t.Fatal(err) } @@ -278,21 +290,18 @@ func TestSlice4_B_BackupThatYieldsNoFreshUnitRefuses(t *testing.T) { } } -// ── C: no backup exists that could restore this app ─────────────────────────────────────────────── +// ── C / R-475 L: no copy on any tier, and no way to make one ───────────────────────────────────── -// COMPANION RED-PROOF 2 (REPORT.md): make the precondition in UpdatePreflight proceed when -// !rp.Restorable. This test then fails with the update started. -func TestSlice4_C_NoRestorableCopyRefusesBeforeAnythingMoves(t *testing.T) { +// Until v0.239.0 this refused any app without a restorable Tier-2 unit. R-475: every tier counts and +// an app with nothing is backed up first, so the refusal is now only "nothing anywhere AND no backup +// can be taken". The K half (nothing, but a backup CAN be taken) is TestR475_K in update_tiers_test.go. +func TestSlice4_C_NoCopyAndNoWayToBackUpRefusesBeforeAnythingMoves(t *testing.T) { m, dir, g, c := newSlice4Manager(t) - // A PROVEN, FRESH copy whose unit cannot be opened — the realistic half-copied mirror. Proven and - // fresh on purpose: a fixture that is also unproven would be refused by the proven check alone, - // and a red-proof that drops the restorable check would then pass inertly (observed on the first - // run of red-proof 2, 2026-09-13). - g.rp = UpdateRestorePoint{Restorable: false, Proven: true, ProvenAt: slice4T0.Add(-time.Hour)} + g.points, g.cannotBackUp = nil, true err := m.StartGuardedUpdate("nextcloud") var ref *UpdateRefusal if !errors.As(err, &ref) || ref.Reason != "no_backup" { - t.Fatalf("an app with no restorable copy must be REFUSED (no_backup), got %v", err) + t.Fatalf("no copy anywhere and no way to back up must be REFUSED (no_backup), got %v", err) } if want := fmt.Sprintf(MsgUpdateNoBackupFmt, "nextcloud"); ref.Message != want { t.Errorf("message = %q", ref.Message) @@ -304,10 +313,10 @@ func TestSlice4_C_NoRestorableCopyRefusesBeforeAnythingMoves(t *testing.T) { if pinOf(t, dir) != "nextcloud:31.0.14-apache" || len(c.list()) != 0 { t.Error("a refused update must move nothing") } - // A copy that exists but was never PROVEN is not a copy (R-101). - g.rp = UpdateRestorePoint{Restorable: true, Proven: false} - if ref := m.UpdatePreflight("nextcloud"); ref == nil || ref.Reason != "no_backup" { - t.Errorf("an unproven copy must refuse too, got %v", ref) + for _, call := range g.callList() { + if call == "BackupNow" { + t.Error("a refused update must not try to back up") + } } } @@ -412,8 +421,8 @@ func TestSlice4_F_HealthFailureHoldsTheAppAndKeepsTheNewPin(t *testing.T) { if !held { t.Fatal("an app that did not come up must be HELD") } - if !g.holdProvenAt.Equal(g.rp.ProvenAt) { - t.Errorf("the hold must name the PROVEN copy date %s, got %s", g.rp.ProvenAt, g.holdProvenAt) + if !g.holdRP.ProvenAt.Equal(g.points[0].ProvenAt) || g.holdRP.Tier != UpdateTierSecondDrive { + t.Errorf("the hold must name the PROVEN copy it leans on (%+v), got %+v", g.points[0], g.holdRP) } if st.UpdateError != "HELD-SENTENCE" { t.Errorf("the page must carry the hold's own sentence, got %q", st.UpdateError) @@ -466,6 +475,7 @@ func simulateAdvanced(t *testing.T, m *Manager, dir string) updateJournalEntry { StartedAt: slice4T0, PrevPin: map[string]string{"web": "nextcloud:31.0.14-apache"}, PrevCompose: filepath.Join(dir, preUpdateComposeFile), PrevApplied: filepath.Join(dir, preUpdateAppliedFile), ProvenCopyAt: slice4T0.Add(-time.Hour).Format(time.RFC3339), + ProvenTier: UpdateTierLocal, } } @@ -530,8 +540,8 @@ func TestSlice4_G_InterruptedAfterUpResumesTheHealthWait(t *testing.T) { if got := strings.Join(c.list(), " | "); got != "up -d --remove-orphans | down" { t.Errorf("resumption re-runs `up` then stops the failed app; compose calls = %q", got) } - if !g.holdProvenAt.Equal(slice4T0.Add(-time.Hour)) { - t.Errorf("the resumed hold must name the journaled proven copy date, got %s", g.holdProvenAt) + if !g.holdRP.ProvenAt.Equal(slice4T0.Add(-time.Hour)) || g.holdRP.Tier != UpdateTierLocal { + t.Errorf("the resumed hold must name the journaled copy (tier %d at %s), got %+v", UpdateTierLocal, slice4T0.Add(-time.Hour), g.holdRP) } } diff --git a/controller/internal/stacks/update_tiers_test.go b/controller/internal/stacks/update_tiers_test.go new file mode 100644 index 0000000..33eb232 --- /dev/null +++ b/controller/internal/stacks/update_tiers_test.go @@ -0,0 +1,175 @@ +package stacks + +import ( + "context" + "errors" + "fmt" + "testing" + "time" +) + +// R-475 — any backup tier lets an app update (operator ruling 2026-09-13, controller v0.239.0). +// Tier order and the Tier-3 timeout are the backup side's (internal/backup/update_tiers_test.go); +// here is the update job's half: which copy it leans on, when it backs up first, when it refuses. +// Every age here is read against slice4T0, the same clock the job reads (R-457). + +func hasGuardCall(g *fakeGuards, name string) bool { + for _, c := range g.callList() { + if c == name { + return true + } + } + return false +} + +// failedUpdateHold runs an update whose new version never becomes healthy, so the hold records the +// copy the job chose. That is the observable consequence of the choice: the customer is told to +// restore from exactly that copy. +func failedUpdateHold(t *testing.T, points, afterBackup []UpdateRestorePoint) (*fakeGuards, *Stack) { + t.Helper() + m, _, g, _ := newSlice4Manager(t) + g.points, g.pointsAfterBackup = points, afterBackup + m.updateHealthFn = func(context.Context, string, time.Duration) (bool, string) { return false, "crash loop" } + if err := m.StartGuardedUpdate("nextcloud"); err != nil { + t.Fatalf("start: %v", err) + } + return g, waitUpdateDone(t, m, "nextcloud") +} + +func TestR475_G_TheSecondDriveIsChosenWhenPresent(t *testing.T) { + fresh := slice4T0.Add(-time.Hour) + g, st := failedUpdateHold(t, []UpdateRestorePoint{ + {Tier: UpdateTierSecondDrive, ProvenAt: fresh}, + {Tier: UpdateTierLocal, ProvenAt: fresh.Add(30 * time.Minute)}, + {Tier: UpdateTierOffsite, ProvenAt: fresh}, + }, nil) + if st.UpdatePhase != UpdatePhaseFailed || g.holdRP.Tier != UpdateTierSecondDrive || !g.holdRP.ProvenAt.Equal(fresh) { + t.Errorf("with a fresh second-drive copy that copy is chosen; hold = %+v phase=%q", g.holdRP, st.UpdatePhase) + } + if hasGuardCall(g, "BackupNow") { + t.Error("a fresh copy needs no backup first") + } +} + +func TestR475_H_TheOwnUnitAloneCarriesTheUpdate(t *testing.T) { + fresh := slice4T0.Add(-2 * time.Hour) + g, _ := failedUpdateHold(t, []UpdateRestorePoint{{Tier: UpdateTierLocal, ProvenAt: fresh}}, nil) + if g.holdRP.Tier != UpdateTierLocal || !g.holdRP.ProvenAt.Equal(fresh) { + t.Errorf("an app with only its own recovery unit must update against it; hold = %+v", g.holdRP) + } + if hasGuardCall(g, "BackupNow") { + t.Error("a fresh own unit needs no backup first — Tier 2 must not be required anywhere") + } +} + +func TestR475_I_TheOffsiteCopyAloneCarriesTheUpdate(t *testing.T) { + fresh := slice4T0.Add(-3 * time.Hour) + g, _ := failedUpdateHold(t, []UpdateRestorePoint{{Tier: UpdateTierOffsite, ProvenAt: fresh}}, nil) + if g.holdRP.Tier != UpdateTierOffsite || !g.holdRP.ProvenAt.Equal(fresh) || hasGuardCall(g, "BackupNow") { + t.Errorf("an app with only a fresh off-site copy must update against it; hold = %+v calls=%v", g.holdRP, g.callList()) + } +} + +func TestR475_K_NoCopyAnywhereIsBackedUpFirstAndTheUpdateProceeds(t *testing.T) { + m, dir, g, _ := newSlice4Manager(t) + g.points = nil + g.pointsAfterBackup = []UpdateRestorePoint{{Tier: UpdateTierLocal, ProvenAt: slice4T0.Add(-time.Minute)}} + if ref := m.UpdatePreflight("nextcloud"); ref != nil { + t.Fatalf("an app with no copy but a working backup must NOT be refused, got %q (%s)", ref.Message, ref.Reason) + } + if err := m.StartGuardedUpdate("nextcloud"); err != nil { + t.Fatal(err) + } + st := waitUpdateDone(t, m, "nextcloud") + if st.UpdatePhase != UpdatePhaseDone { + t.Fatalf("the backup made a Tier-1 copy, so the update completes; phase=%q err=%q", st.UpdatePhase, st.UpdateError) + } + if !hasGuardCall(g, "BackupNow") { + t.Error("with no copy anywhere the job must back up FIRST") + } + if got := pinOf(t, dir); got != "nextcloud:34.0.1-apache" { + t.Errorf("pin = %q", got) + } +} + +func TestR475_K_NoCopyAndTheBackupFailsMovesNothing(t *testing.T) { + m, dir, g, c := newSlice4Manager(t) + g.points = nil + g.backupErr = errors.New("nincs elég szabad hely") + if err := m.StartGuardedUpdate("nextcloud"); err != nil { + t.Fatal(err) + } + st := waitUpdateDone(t, m, "nextcloud") + if want := fmt.Sprintf(MsgUpdateBackupFailFmt, g.backupErr); st.UpdateError != want { + t.Errorf("UpdateError = %q, want %q", st.UpdateError, want) + } + if pinOf(t, dir) != "nextcloud:31.0.14-apache" || len(c.list()) != 0 { + t.Error("nothing may move when the backup-first failed") + } +} + +// L — nothing anywhere and no way to back up is refused (TestSlice4_C_… pins the refusal itself). The +// other half pins that the refusal is not wider than that: a copy that EXISTS still carries it. +func TestR475_L_AnExistingCopyStillCarriesAnAppThatCannotBeBackedUp(t *testing.T) { + m, _, g, _ := newSlice4Manager(t) + g.cannotBackUp = true + g.points = []UpdateRestorePoint{{Tier: UpdateTierOffsite, ProvenAt: slice4T0.Add(-time.Hour)}} + if ref := m.UpdatePreflight("nextcloud"); ref != nil { + t.Fatalf("a fresh off-site copy must carry the update even when no backup can be taken now, got %q", ref.Message) + } + g.points = nil + if ref := m.UpdatePreflight("nextcloud"); ref == nil || ref.Reason != "no_backup" { + t.Fatalf("positive control: with no copy the same app must be refused, got %v", ref) + } +} + +// M — a stale copy on ANY tier is stale. Each case would pass under a Tier-2-only age check except the +// first, which is the control proving the fixture's clock and limit are real. +func TestR475_M_TheAgeRuleAppliesToTheChosenTier(t *testing.T) { + stale, fresh, justNow := slice4T0.Add(-30*time.Hour), slice4T0.Add(-time.Hour), slice4T0.Add(-time.Minute) + cases := []struct { + name string + points []UpdateRestorePoint + afterBackup []UpdateRestorePoint + wantBackup bool + wantTier int + }{ + {"control: a stale second-drive copy is backed up first", []UpdateRestorePoint{{Tier: UpdateTierSecondDrive, ProvenAt: stale}}, + []UpdateRestorePoint{{Tier: UpdateTierSecondDrive, ProvenAt: justNow}}, true, UpdateTierSecondDrive}, + {"a stale own unit is backed up first", []UpdateRestorePoint{{Tier: UpdateTierLocal, ProvenAt: stale}}, + []UpdateRestorePoint{{Tier: UpdateTierLocal, ProvenAt: justNow}}, true, UpdateTierLocal}, + {"a stale off-site copy is backed up first", []UpdateRestorePoint{{Tier: UpdateTierOffsite, ProvenAt: stale}}, + []UpdateRestorePoint{{Tier: UpdateTierLocal, ProvenAt: justNow}, {Tier: UpdateTierOffsite, ProvenAt: stale}}, true, UpdateTierLocal}, + {"a stale second drive does not block a fresh own unit", []UpdateRestorePoint{{Tier: UpdateTierSecondDrive, ProvenAt: stale}, {Tier: UpdateTierLocal, ProvenAt: fresh}}, + nil, false, UpdateTierLocal}, + } + for _, tc := range cases { + t.Run(tc.name, func(t *testing.T) { + g, st := failedUpdateHold(t, tc.points, tc.afterBackup) + if got := hasGuardCall(g, "BackupNow"); got != tc.wantBackup { + t.Errorf("backed up first = %v, want %v (calls %v)", got, tc.wantBackup, g.callList()) + } + if g.holdRP.Tier != tc.wantTier { + t.Errorf("the chosen copy is tier %d, want %d (hold %+v, err %q)", g.holdRP.Tier, tc.wantTier, g.holdRP, st.UpdateError) + } + if age := slice4T0.Sub(g.holdRP.ProvenAt); age > 24*time.Hour { + t.Errorf("the chosen copy is %s old — past the age limit", age) + } + }) + } +} + +func TestR475_M_TheLimitIsInclusiveOnOneClock(t *testing.T) { + fresh := freshRestorePoint(slice4T0, 24*time.Hour) + for _, tier := range []int{UpdateTierSecondDrive, UpdateTierLocal, UpdateTierOffsite} { + if !fresh(UpdateRestorePoint{Tier: tier, ProvenAt: slice4T0.Add(-24 * time.Hour)}) { + t.Errorf("tier %d: exactly the limit old is still fresh", tier) + } + if fresh(UpdateRestorePoint{Tier: tier, ProvenAt: slice4T0.Add(-24*time.Hour - time.Second)}) { + t.Errorf("tier %d: a second past the limit is stale", tier) + } + if fresh(UpdateRestorePoint{Tier: tier}) { + t.Errorf("tier %d: a copy with no proven time is never fresh", tier) + } + } +} diff --git a/controller/internal/web/slice4_update_test.go b/controller/internal/web/slice4_update_test.go index aeb2d5d..6110d1b 100644 --- a/controller/internal/web/slice4_update_test.go +++ b/controller/internal/web/slice4_update_test.go @@ -62,7 +62,7 @@ func TestSlice4_BackupRowUnitFieldsAreUnchangedByTheExtraction(t *testing.T) { // manager here on purpose, so the unguarded StartStack call panics and this test fails. func TestSlice4_DriveReturnGateSkipsAHeldApp(t *testing.T) { s, _, m := newOffboxWebServer(t) - if err := m.HoldAfterFailedUpdate("held", time.Now(), time.Now()); err != nil { + if err := m.HoldAfterFailedUpdate("held", time.Now(), time.Now(), backup.UpdateTierLocal); err != nil { t.Fatal(err) } defer func() {