controller v0.241.0: a bind-data app leans on off-site before its own unit; the hold names what the copy holds (R-479)
gates / gates (push) Successful in 13s

Operator ruling 2026-09-13. An app with classified binds walks second
drive -> off-site -> own unit (its unit holds no files); volume apps keep
2 -> 1 -> 3. RestoreHold.CopyHolds records what the chosen copy holds and
the sentence ends with it; older holds keep their tier-only sentence.
Tests on both halves; red-proof: a layout-blind order fails the bind case.
This commit is contained in:
2026-09-13 21:47:33 +02:00
parent 3013a1cc93
commit 3e813307cc
11 changed files with 208 additions and 7 deletions
+1 -1
View File
@@ -124,7 +124,7 @@
| `Manager.UpdatePreflight` / `StartGuardedUpdate` / `RecoverUpdates` / `ResumeInterruptedUpdates` (v0.237.0) | controller/internal/stacks/update.go | `UpdatePreflight(name) *UpdateRefusal`; `StartGuardedUpdate(name) error` | THE update — refusals, then a 202 job with phases on `Stack.Updating/UpdatePhase/UpdateError` | **Never report an update complete before health is known (R-443).** Every cheap refusal runs BEFORE the intent write. Safety dump BEFORE the pin moves; pin BEFORE pull; pull failure → pin back; health failure → stop + HOLD, pin stays. Journal-before-mutate (`update-journal.json`); `RecoverUpdates` MUST run before the boot sweep and `ResumeInterruptedUpdates` AFTER `SetUpdateGuards`. Seams: `updateComposeFn`, `updateHealthFn`, `updateMemoryFn`, `updateDiskFreeFn`, `updateNowFn` (R-457: the age check and the test read ONE clock). Unwired guards ⇒ every update refused |
| `stacks.UpdateGuards` + `updateGuardsAdapter` (v0.237.0; tiers v0.239.0) | controller/internal/stacks/update.go, controller/cmd/controller/main.go | `HoldFor`, `Busy`, `RestorePoints`, `CanBackUp`, `BackupNow`, `SafetyDump`, `HoldAfterFailedUpdate` | the ONLY bridge from the update job to the backup side (stacks cannot import backup) | Wired by `stackMgr.SetUpdateGuards` — pinned by `TestSlice4_UpdateGuardsAreWiredAtStartup`. Add a guard HERE, never by importing backup into stacks |
| `backup.Manager.Tier2UnitRestorePoint` + `Tier2RestorePoint.ProvenCopyTime` (v0.237.0) | controller/internal/backup/update_guard.go | `(stack) (Tier2RestorePoint, error)` | "can this app be restored from Tier 2, and from when" — the predicate that gates BOTH the „Teljes visszaállítás" action and an update | **ONE predicate, two callers** (extracted from `buildAppBackupRows`, not copied). `CopyDate` is what the page NAMES (the package date, R-403); `ProvenCopyTime` is how OLD the data is — the last successful copy, because the manifest's `created_at` moves only when the DEFINITION changes (measured: a fresh dump under a 22-h-older manifest). Do not age a copy by `CopyDate`. **Since v0.239.0 the UPDATE no longer calls it directly** — it goes through `UpdateRestorePoints` (next row); the page still does |
| `backup.Manager.UpdateRestorePoints` + `CanBackUpApp` (v0.239.0, R-475) | controller/internal/backup/update_guard.go | `(ctx, stack, accept func(UpdateTierPoint) bool) (UpdateTierPoint, bool, []UpdateTierPoint)` | "which backup can this update lean on" — walks Tier 2, 1, 3 and returns the first copy `accept` admits | **The age rule lives in the caller's `accept`** (stacks' `freshRestorePoint`), so it is ONE rule for every tier. Stops at the first accepted copy, so Tier 3 (restic, 15 s bound, unreachable = absent + WARN) is reached only when needed. Seams: `updateTier2PointFn`, `updateTier1PointsFn`, `updateOffsiteTimesFn` (v0.240.0: Tier 3 reads `OffsiteSnapshotTimes` — snapshots only, never the inventory's per-app `stats`). stacks adds R-478's rule: a copy older than `deployed_at` does not count (`usableRestorePoint`). Tier numbers are pinned equal across stacks/backup by `TestR475_TierConstantsAgree` |
| `backup.Manager.UpdateRestorePoints` + `CanBackUpApp` (v0.239.0, R-475) | controller/internal/backup/update_guard.go | `(ctx, stack, accept func(UpdateTierPoint) bool) (UpdateTierPoint, bool, []UpdateTierPoint)` | "which backup can this update lean on" — walks Tier 2, 1, 3 and returns the first copy `accept` admits | **The age rule lives in the caller's `accept`** (stacks' `freshRestorePoint`), so it is ONE rule for every tier. **The ORDER is `UpdateTierOrderFor` (v0.241.0, R-479): 2 → 3 → 1 for an app with classified binds (`DataOutsideUnit`), 2 → 1 → 3 otherwise; `UpdateCopyHolds` is the matching phrase for the hold.** Stops at the first accepted copy, so Tier 3 (restic, 15 s bound, unreachable = absent + WARN) is reached only when needed. Seams: `updateTier2PointFn`, `updateTier1PointsFn`, `updateOffsiteTimesFn` (v0.240.0: Tier 3 reads `OffsiteSnapshotTimes` — snapshots only, never the inventory's per-app `stats`). stacks adds R-478's rule: a copy older than `deployed_at` does not count (`usableRestorePoint`). Tier numbers are pinned equal across stacks/backup by `TestR475_TierConstantsAgree` |
| `backup.Manager.Tier2MirrorDirsForApp` / `RemoveTier2Mirrors` + `settings.DeleteAppBackupPrefs` (v0.240.0, R-474 / R-486) | controller/internal/backup/r474_remove_mirrors.go | `(stack) []string`; `(stack, dirs) []string` | deleting an app's backups on removal — the Tier-2 mirror lives on ANOTHER drive, outside RemoveStack's per-app base; the same dirs size the backup card (R-485) | Read the mirror dirs BEFORE the prefs are forgotten. Deletes only `<root>/backups/secondary/<stack>` for a known root; never `_shares`, never a path that merely cleans to it. **The Tier-2 RECORD goes only with `remove_backups`** (R-486) — a removal that keeps the backups must keep the record, or the mirror is unrestorable. Pinned by `TestR474_RemoveHandlerDeletesUnitMirrorAndPrefs` + `TestR486_RemovalKeepsTheTier2RecordUnlessBackupsGo` |
| `appbackup.dbTypeForImage` (R-484, v0.240.0) | controller/internal/appbackup/dbservices.go | `(image) (DBType, bool)` | the ONE place an image is judged a database — nightly dumps, pre-update safety dump, DB-only replay | Derived Postgres images (`postgis`, `pgvector`, `timescaledb`) are Postgres. A new engine image goes HERE and in `dbservices_test.go`'s table, never in a second matcher |
| `backup.Manager.HoldAfterFailedUpdate` / `RunAppBackupNow` / `WriteUpdateSafetyDump` / `UpdateBusy` (v0.237.0) | controller/internal/backup/update_guard.go | see file | the update's hold, per-app backup-now, safety dump, busy check | The hold is `settings.RestoreHold` with `Reason: update_failed` — SAME store and gate as R-379, never a second map. `RunAppBackupNow` composes the nightly legs for ONE app (admission, DB dump, volume dump, capture, Tier-2) — do not write a second backup orchestration. A successful unit restore lifts an UPDATE hold only. **`isHeld` is ALSO true while a guarded update is moving the app (`SetUpdatingCheck`, v0.238.1)** — found live: the periodic capture overwrote a primary unit during a health wait |