controller v0.239.0: any backup tier lets an app update (R-475)
gates / gates (push) Successful in 14s
gates / gates (push) Successful in 14s
Operator ruling 2026-09-13. The update precondition walks Tier 2, Tier 1 (own recovery unit, "helyi") and Tier 3 (off-site, 15 s bound; unreachable counts as absent with a WARN) and leans on the first FRESH copy; the backup_max_age rule applies to whichever tier is chosen. No copy anywhere: back up first. Refused only when nothing exists and no backup can be taken. RunAppBackupNow tolerates a Tier-2 failure (WARN) and marks the captured unit proven current. The hold names the tier (második meghajtó / saját meghajtó / távoli mentés) and the date; pre-v0.239.0 holds keep their text. A successful off-site restore now lifts an update hold. The backups page still uses Tier2UnitRestorePoint unchanged. Scenarios G-M tested; red-proofs M, L, the tail and the off-site clear in felhom.eu documentation/audits/rulings-r472-r475-2026-09-13/.
This commit is contained in:
@@ -1,3 +1,54 @@
|
|||||||
|
## v0.239.0 — any backup tier lets an app update (2026-09-13, R-475)
|
||||||
|
|
||||||
|
**MinAgent: 0.129.0** (unchanged)
|
||||||
|
|
||||||
|
**Operator ruling 2026-09-13: every backup counts.** Until now the update's precondition was the
|
||||||
|
Tier-2 unit predicate alone, so an app with no second-drive copy could never be updated. On demo-hp
|
||||||
|
`gokapi` and `nextcloud` were in exactly that state, each with a fresh recovery unit on its own drive.
|
||||||
|
|
||||||
|
**WHAT CHANGED.**
|
||||||
|
|
||||||
|
- **`backup.Manager.UpdateRestorePoints`**, beside `Tier2UnitRestorePoint`. It walks the tiers in the
|
||||||
|
ruling's order: Tier 2 (second drive), Tier 1 (the app's own unit, `ListRestorePoints`, „helyi"),
|
||||||
|
Tier 3 (`OffsiteInventoryList`, bounded by 15 s). It returns the first copy the caller accepts and
|
||||||
|
stops there, so a fresh Tier-2 copy never reaches the network. An unreachable off-site repository
|
||||||
|
counts as ABSENT, with a WARN. A box with no off-site target is plainly absent, with no WARN.
|
||||||
|
- **The age rule is one rule.** stacks passes `freshRestorePoint` (`update.backup_max_age`) as the
|
||||||
|
acceptance test, so the limit applies to whichever tier is chosen. The first FRESH copy wins, so a
|
||||||
|
stale Tier-2 mirror never forces a backup while the app's own unit is minutes old.
|
||||||
|
- **No copy anywhere → back up first**, then re-read every tier. The preflight refuses `no_backup`
|
||||||
|
only when there is no copy on any tier AND `CanBackUpApp` says no backup can be taken now. The
|
||||||
|
Hungarian sentence changed to say that; it no longer tells the customer to switch on the 2nd backup.
|
||||||
|
- **`RunAppBackupNow` tolerates no Tier-2 target.** A Tier-2 failure after the unit capture is a WARN.
|
||||||
|
The captured unit's manifest is marked proven current, because the capture's checksum skip leaves it
|
||||||
|
untouched on a quiet app, and Tier 1 is aged by the newest artifact's mtime. Without that mark an
|
||||||
|
app with no database and no volume would be refused forever — the ProvenCopyTime trap, one tier down.
|
||||||
|
- **The hold names the tier.** `settings.RestoreHold.CopyTier`; the sentence now ends *„Visszaállítható
|
||||||
|
a Mentések oldalon ebből a biztonsági mentésből: <második meghajtó | saját meghajtó | távoli mentés>,
|
||||||
|
<dátum>."* A hold written by v0.237.0–v0.238.1 has no tier and keeps its original sentence
|
||||||
|
(`UpdateHoldLegacyFmt`). The update journal records `proven_tier`, so a resumed update names the
|
||||||
|
right copy.
|
||||||
|
- **A successful off-site restore lifts an update hold.** That path never went through
|
||||||
|
`RestoreFromRecoveryUnitAt`, so a hold naming „távoli mentés" could otherwise never be cleared.
|
||||||
|
- **Unchanged:** the backups page still calls `Tier2UnitRestorePoint` for „Teljes visszaállítás".
|
||||||
|
|
||||||
|
**TESTS.** `internal/stacks/update_tiers_test.go` — G (Tier 2 chosen when present), H (own unit alone),
|
||||||
|
I (off-site alone), K (nothing anywhere: backed up first and the update completes; a failing backup
|
||||||
|
moves nothing), L (an existing copy still carries an app that cannot be backed up; control: without it
|
||||||
|
the app is refused), M (a stale copy on each tier is backed up first; a stale Tier 2 does not block a
|
||||||
|
fresh Tier 1; the limit is inclusive on one clock). `internal/backup/update_tiers_test.go` — the tier
|
||||||
|
order and early stop, H/I at the source, J (unreachable → absent + WARN; no target → silent; a hanging
|
||||||
|
repository ends at the bound), the hold text per tier and the legacy text, the tolerant pre-backup
|
||||||
|
tail, `RunAppBackupNow` really calling it, `CanBackUpApp`, and the off-site restore clearing the hold.
|
||||||
|
`cmd/controller/r475_wiring_test.go` — the tier numbers agree across packages; the adapter reads every
|
||||||
|
tier and never `Tier2UnitRestorePoint`; the hold is told the tier. Slice 4's tests were moved onto the
|
||||||
|
new seam (`TestSlice4_C` is now the L refusal).
|
||||||
|
|
||||||
|
**Red-proofs** (felhom.eu `documentation/audits/rulings-r472-r475-2026-09-13/`): **M** — age checked
|
||||||
|
only for Tier 2: three M cases fail (a 30-hour-old own unit and off-site copy carry the update with no
|
||||||
|
backup). **L** — refuse even with a copy: the L test fails. **Tail** — the proven-current mark misses:
|
||||||
|
the tail test fails on the mtime. **Off-site clear** removed: the hold stays and the test fails.
|
||||||
|
|
||||||
## v0.238.1 — the nightly backup leaves an app alone WHILE it is being updated, not only once it is held (2026-09-13, slice 4 follow-up)
|
## v0.238.1 — the nightly backup leaves an app alone WHILE it is being updated, not only once it is held (2026-09-13, slice 4 follow-up)
|
||||||
|
|
||||||
**MinAgent: 0.129.0** (unchanged)
|
**MinAgent: 0.129.0** (unchanged)
|
||||||
|
|||||||
@@ -122,8 +122,9 @@
|
|||||||
| `Manager.recordInstalledImages` (v0.233.0) | controller/internal/stacks/installed.go | `(name, stackDir string, env []string)` | writing down what each compose SERVICE is ACTUALLY running, into `app.yaml.installed_images` | Called after a successful compose up from `StartStack`/`RestartStack`/`UpdateStack`/`runComposeDeploy`. **Reads the CONTAINER, never `docker-compose.yml`** — that file is the value the syncer has already moved (spike §3: 25 minutes of disagreement). **A failed write NEVER refuses the action** — the deliberate OPPOSITE of `SetDesiredState`: intent refused, observation logged at ERROR. **NOT from `StartStackServices`** (the R-47 DB-only window would overwrite a complete record with a partial one). Skips the write when ref+digest are unchanged, and carries `at` forward so it means "running since". Its OWN seam (`installedExecFn`) with a **context + 30 s timeout** — the two existing exec helpers have neither |
|
| `Manager.recordInstalledImages` (v0.233.0) | controller/internal/stacks/installed.go | `(name, stackDir string, env []string)` | writing down what each compose SERVICE is ACTUALLY running, into `app.yaml.installed_images` | Called after a successful compose up from `StartStack`/`RestartStack`/`UpdateStack`/`runComposeDeploy`. **Reads the CONTAINER, never `docker-compose.yml`** — that file is the value the syncer has already moved (spike §3: 25 minutes of disagreement). **A failed write NEVER refuses the action** — the deliberate OPPOSITE of `SetDesiredState`: intent refused, observation logged at ERROR. **NOT from `StartStackServices`** (the R-47 DB-only window would overwrite a complete record with a partial one). Skips the write when ref+digest are unchanged, and carries `at` forward so it means "running since". Its OWN seam (`installedExecFn`) with a **context + 30 s timeout** — the two existing exec helpers have neither |
|
||||||
| `Manager.SetPin` / `AdoptPins` / `RenderPlanFor` / `AppliedComposePath` (v0.235.0) | controller/internal/stacks/pin.go | `SetPin(name, stackDir, pin, composeSrc) error` | THE version freeze — `app.yaml.pinned_images` + the stored `applied-compose.yml` | **`PinnedImages` is INTENT, `InstalledImages` is an OBSERVATION — never feed one from the other** (the R-166 category error, one field over). Four writers only: deploy, the guarded update (via `advancePinToCatalog`, which advances the pin and re-renders BEFORE the pull, and REFUSES the update if the pin cannot be written; a failed pull puts the pin BACK via `SetPin`), the restore adapter (this is what closes R-441), and `AdoptPins`. `AdoptPins` reuses `observationCoversTemplate` — do NOT write a second completeness rule — and skips loudly rather than inventing a pin. Absent pin = pre-v0.235.0 behaviour |
|
| `Manager.SetPin` / `AdoptPins` / `RenderPlanFor` / `AppliedComposePath` (v0.235.0) | controller/internal/stacks/pin.go | `SetPin(name, stackDir, pin, composeSrc) error` | THE version freeze — `app.yaml.pinned_images` + the stored `applied-compose.yml` | **`PinnedImages` is INTENT, `InstalledImages` is an OBSERVATION — never feed one from the other** (the R-166 category error, one field over). Four writers only: deploy, the guarded update (via `advancePinToCatalog`, which advances the pin and re-renders BEFORE the pull, and REFUSES the update if the pin cannot be written; a failed pull puts the pin BACK via `SetPin`), the restore adapter (this is what closes R-441), and `AdoptPins`. `AdoptPins` reuses `observationCoversTemplate` — do NOT write a second completeness rule — and skips loudly rather than inventing a pin. Absent pin = pre-v0.235.0 behaviour |
|
||||||
| `Manager.UpdatePreflight` / `StartGuardedUpdate` / `RecoverUpdates` / `ResumeInterruptedUpdates` (v0.237.0) | controller/internal/stacks/update.go | `UpdatePreflight(name) *UpdateRefusal`; `StartGuardedUpdate(name) error` | THE update — refusals, then a 202 job with phases on `Stack.Updating/UpdatePhase/UpdateError` | **Never report an update complete before health is known (R-443).** Every cheap refusal runs BEFORE the intent write. Safety dump BEFORE the pin moves; pin BEFORE pull; pull failure → pin back; health failure → stop + HOLD, pin stays. Journal-before-mutate (`update-journal.json`); `RecoverUpdates` MUST run before the boot sweep and `ResumeInterruptedUpdates` AFTER `SetUpdateGuards`. Seams: `updateComposeFn`, `updateHealthFn`, `updateMemoryFn`, `updateDiskFreeFn`, `updateNowFn` (R-457: the age check and the test read ONE clock). Unwired guards ⇒ every update refused |
|
| `Manager.UpdatePreflight` / `StartGuardedUpdate` / `RecoverUpdates` / `ResumeInterruptedUpdates` (v0.237.0) | controller/internal/stacks/update.go | `UpdatePreflight(name) *UpdateRefusal`; `StartGuardedUpdate(name) error` | THE update — refusals, then a 202 job with phases on `Stack.Updating/UpdatePhase/UpdateError` | **Never report an update complete before health is known (R-443).** Every cheap refusal runs BEFORE the intent write. Safety dump BEFORE the pin moves; pin BEFORE pull; pull failure → pin back; health failure → stop + HOLD, pin stays. Journal-before-mutate (`update-journal.json`); `RecoverUpdates` MUST run before the boot sweep and `ResumeInterruptedUpdates` AFTER `SetUpdateGuards`. Seams: `updateComposeFn`, `updateHealthFn`, `updateMemoryFn`, `updateDiskFreeFn`, `updateNowFn` (R-457: the age check and the test read ONE clock). Unwired guards ⇒ every update refused |
|
||||||
| `stacks.UpdateGuards` + `updateGuardsAdapter` (v0.237.0) | controller/internal/stacks/update.go, controller/cmd/controller/main.go | `HoldFor`, `Busy`, `RestorePoint`, `BackupNow`, `SafetyDump`, `HoldAfterFailedUpdate` | the ONLY bridge from the update job to the backup side (stacks cannot import backup) | Wired by `stackMgr.SetUpdateGuards` — pinned by `TestSlice4_UpdateGuardsAreWiredAtStartup`. Add a guard HERE, never by importing backup into stacks |
|
| `stacks.UpdateGuards` + `updateGuardsAdapter` (v0.237.0; tiers v0.239.0) | controller/internal/stacks/update.go, controller/cmd/controller/main.go | `HoldFor`, `Busy`, `RestorePoints`, `CanBackUp`, `BackupNow`, `SafetyDump`, `HoldAfterFailedUpdate` | the ONLY bridge from the update job to the backup side (stacks cannot import backup) | Wired by `stackMgr.SetUpdateGuards` — pinned by `TestSlice4_UpdateGuardsAreWiredAtStartup`. Add a guard HERE, never by importing backup into stacks |
|
||||||
| `backup.Manager.Tier2UnitRestorePoint` + `Tier2RestorePoint.ProvenCopyTime` (v0.237.0) | controller/internal/backup/update_guard.go | `(stack) (Tier2RestorePoint, error)` | "can this app be restored from Tier 2, and from when" — the predicate that gates BOTH the „Teljes visszaállítás" action and an update | **ONE predicate, two callers** (extracted from `buildAppBackupRows`, not copied). `CopyDate` is what the page NAMES (the package date, R-403); `ProvenCopyTime` is how OLD the data is — the last successful copy, because the manifest's `created_at` moves only when the DEFINITION changes (measured: a fresh dump under a 22-h-older manifest). Do not age a copy by `CopyDate` |
|
| `backup.Manager.Tier2UnitRestorePoint` + `Tier2RestorePoint.ProvenCopyTime` (v0.237.0) | controller/internal/backup/update_guard.go | `(stack) (Tier2RestorePoint, error)` | "can this app be restored from Tier 2, and from when" — the predicate that gates BOTH the „Teljes visszaállítás" action and an update | **ONE predicate, two callers** (extracted from `buildAppBackupRows`, not copied). `CopyDate` is what the page NAMES (the package date, R-403); `ProvenCopyTime` is how OLD the data is — the last successful copy, because the manifest's `created_at` moves only when the DEFINITION changes (measured: a fresh dump under a 22-h-older manifest). Do not age a copy by `CopyDate`. **Since v0.239.0 the UPDATE no longer calls it directly** — it goes through `UpdateRestorePoints` (next row); the page still does |
|
||||||
|
| `backup.Manager.UpdateRestorePoints` + `CanBackUpApp` (v0.239.0, R-475) | controller/internal/backup/update_guard.go | `(ctx, stack, accept func(UpdateTierPoint) bool) (UpdateTierPoint, bool, []UpdateTierPoint)` | "which backup can this update lean on" — walks Tier 2, 1, 3 and returns the first copy `accept` admits | **The age rule lives in the caller's `accept`** (stacks' `freshRestorePoint`), so it is ONE rule for every tier. Stops at the first accepted copy, so Tier 3 (restic, 15 s bound, unreachable = absent + WARN) is reached only when needed. Seams: `updateTier2PointFn`, `updateTier1PointsFn`, `updateOffsiteInvFn`. Tier numbers are pinned equal across stacks/backup by `TestR475_TierConstantsAgree` |
|
||||||
| `backup.Manager.HoldAfterFailedUpdate` / `RunAppBackupNow` / `WriteUpdateSafetyDump` / `UpdateBusy` (v0.237.0) | controller/internal/backup/update_guard.go | see file | the update's hold, per-app backup-now, safety dump, busy check | The hold is `settings.RestoreHold` with `Reason: update_failed` — SAME store and gate as R-379, never a second map. `RunAppBackupNow` composes the nightly legs for ONE app (admission, DB dump, volume dump, capture, Tier-2) — do not write a second backup orchestration. A successful unit restore lifts an UPDATE hold only. **`isHeld` is ALSO true while a guarded update is moving the app (`SetUpdatingCheck`, v0.238.1)** — found live: the periodic capture overwrote a primary unit during a health wait |
|
| `backup.Manager.HoldAfterFailedUpdate` / `RunAppBackupNow` / `WriteUpdateSafetyDump` / `UpdateBusy` (v0.237.0) | controller/internal/backup/update_guard.go | see file | the update's hold, per-app backup-now, safety dump, busy check | The hold is `settings.RestoreHold` with `Reason: update_failed` — SAME store and gate as R-379, never a second map. `RunAppBackupNow` composes the nightly legs for ONE app (admission, DB dump, volume dump, capture, Tier-2) — do not write a second backup orchestration. A successful unit restore lifts an UPDATE hold only. **`isHeld` is ALSO true while a guarded update is moving the app (`SetUpdatingCheck`, v0.238.1)** — found live: the periodic capture overwrote a primary unit during a health wait |
|
||||||
| `stacks.Manager.memoryVerdict` (v0.237.0) | controller/internal/stacks/deploy.go | `(newReq, newLimit, releasedReq, releasedLimit int) (refusal, warning string)` | the deploy's memory check, shared with the update | An update RELEASES the app's current request first. Deploy passes `0, 0` and is byte-identical in wording and log line |
|
| `stacks.Manager.memoryVerdict` (v0.237.0) | controller/internal/stacks/deploy.go | `(newReq, newLimit, releasedReq, releasedLimit int) (refusal, warning string)` | the deploy's memory check, shared with the update | An update RELEASES the app's current request first. Deploy passes `0, 0` and is byte-identical in wording and log line |
|
||||||
| `Syncer.SetRenderPlanFn` + `renderSource` (v0.235.0) | controller/internal/sync/sync.go | `func(appName string) stacks.RenderPlan` | the catalog render table | **NIL-SAFE: no seam = copy verbatim = the old product.** Catalog images == pin → verbatim (fixes flow + self-healing, both deliberately kept); differ → the WHOLE stored definition, **never a ref substitution into a newer template** (`wger 2.6`). `.felhom.yml` always verbatim (R-458). The syncer must NEVER read app.yaml. Re-reads the applied file before writing it — a test caught it writing an empty compose over a live app |
|
| `Syncer.SetRenderPlanFn` + `renderSource` (v0.235.0) | controller/internal/sync/sync.go | `func(appName string) stacks.RenderPlan` | the catalog render table | **NIL-SAFE: no seam = copy verbatim = the old product.** Catalog images == pin → verbatim (fixes flow + self-healing, both deliberately kept); differ → the WHOLE stored definition, **never a ref substitution into a newer template** (`wger 2.6`). `.felhom.yml` always verbatim (R-458). The syncer must NEVER read app.yaml. Re-reads the applied file before writing it — a test caught it writing an empty compose over a live app |
|
||||||
|
|||||||
+16
-7
@@ -570,22 +570,31 @@ job, and answers **202**. The page polls `GET /api/stacks/{name}`.
|
|||||||
**Refused with 409 before anything moves:** held; a backup, restore, app-data op or quiesce holding
|
**Refused with 409 before anything moves:** held; a backup, restore, app-data op or quiesce holding
|
||||||
the app; a migration; already updating; deploying; not enough memory for the NEW template's request
|
the app; a migration; already updating; deploying; not enough memory for the NEW template's request
|
||||||
(the deploy's own `memoryVerdict`, releasing the app's current request); less than **2 GB** free on
|
(the deploy's own `memoryVerdict`, releasing the app's current request); less than **2 GB** free on
|
||||||
the Docker data root (a fixed floor — image sizes are not known without a registry query); and **no
|
the Docker data root (a fixed floor — image sizes are not known without a registry query); and — since
|
||||||
restorable backup** (no openable Tier-2 recovery unit with a proven copy).
|
v0.239.0 — **no copy on any backup tier AND no way to take one now** (drive unresolvable, disconnected,
|
||||||
|
or a migration running).
|
||||||
|
|
||||||
**The sequence.** A proven copy older than `update.backup_max_age` is refreshed first
|
**Any backup tier counts (v0.239.0, R-475).** The update leans on the first FRESH copy in the order
|
||||||
(`RunAppBackupNow`: this app's DB dump, volume dump, unit capture, Tier-2 copy). Then a database
|
second drive (Tier 2), the app's own recovery unit (Tier 1, „helyi"), off-site (Tier 3, looked up with
|
||||||
|
a 15 s bound — unreachable counts as absent, with a WARN). `update.backup_max_age` applies to whichever
|
||||||
|
tier is chosen. An app with no copy anywhere is backed up first. Tier 2 is required nowhere in the
|
||||||
|
update path; the backups page's „Teljes visszaállítás" still reads the Tier-2 predicate alone.
|
||||||
|
|
||||||
|
**The sequence.** With no fresh copy on any tier the app is backed up first (`RunAppBackupNow`: this
|
||||||
|
app's DB dump, volume dump, unit capture, then a Tier-2 copy whose failure is only a WARN — the unit
|
||||||
|
just captured is marked proven current, so a quiet app's own unit counts as fresh). Then a database
|
||||||
safety dump, then the pin moves, then pull, `up`, and the health wait (`.felhom.yml` check, or 60 s of
|
safety dump, then the pin moves, then pull, `up`, and the health wait (`.felhom.yml` check, or 60 s of
|
||||||
every container running for an app with none; bounded by `update.health_timeout`). **A failed pull
|
every container running for an app with none; bounded by `update.health_timeout`). **A failed pull
|
||||||
puts the pin back. An app that does not become healthy is stopped and HELD** — the pin stays on the
|
puts the pin back. An app that does not become healthy is stopped and HELD** — the pin stays on the
|
||||||
new version, and the hold sentence names the backup it can be restored from. A successful unit
|
new version, and the hold sentence names the tier and the date of the copy it can be restored from
|
||||||
restore lifts an update hold.
|
(„második meghajtó" / „saját meghajtó" / „távoli mentés"). A successful unit restore — and, since
|
||||||
|
v0.239.0, a successful off-site restore — lifts an update hold.
|
||||||
|
|
||||||
**Config (`controller.yaml`):**
|
**Config (`controller.yaml`):**
|
||||||
|
|
||||||
```yaml
|
```yaml
|
||||||
update:
|
update:
|
||||||
backup_max_age: 24h # a proven Tier-2 copy older than this is refreshed before the update
|
backup_max_age: 24h # the chosen copy (any tier) must be younger than this, or the app is backed up first
|
||||||
health_timeout: 5m # how long the new version has to become healthy before the app is held
|
health_timeout: 5m # how long the new version has to become healthy before the app is held
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|||||||
@@ -3378,16 +3378,32 @@ func (a *updateGuardsAdapter) Busy(name string) (bool, string) {
|
|||||||
return a.b.UpdateBusy(name)
|
return a.b.UpdateBusy(name)
|
||||||
}
|
}
|
||||||
|
|
||||||
func (a *updateGuardsAdapter) RestorePoint(name string) (stacks.UpdateRestorePoint, error) {
|
// RestorePoints (R-475) reads EVERY tier through backup.UpdateRestorePoints. It must never go back to
|
||||||
|
// Tier2UnitRestorePoint alone — pinned by TestR475_AdapterReadsEveryTier.
|
||||||
|
func (a *updateGuardsAdapter) RestorePoints(ctx context.Context, name string, accept func(stacks.UpdateRestorePoint) bool) (stacks.UpdateRestorePoint, bool, []stacks.UpdateRestorePoint) {
|
||||||
if a.b == nil {
|
if a.b == nil {
|
||||||
return stacks.UpdateRestorePoint{}, fmt.Errorf("backup is not enabled on this box")
|
return stacks.UpdateRestorePoint{}, false, nil
|
||||||
}
|
}
|
||||||
rp, err := a.b.Tier2UnitRestorePoint(name)
|
conv := func(p backup.UpdateTierPoint) stacks.UpdateRestorePoint {
|
||||||
if err != nil {
|
return stacks.UpdateRestorePoint{Tier: p.Tier, ProvenAt: p.At}
|
||||||
return stacks.UpdateRestorePoint{}, err
|
|
||||||
}
|
}
|
||||||
at, proven := rp.ProvenCopyTime()
|
var acc func(backup.UpdateTierPoint) bool
|
||||||
return stacks.UpdateRestorePoint{Restorable: rp.Restorable, Proven: proven, ProvenAt: at}, nil
|
if accept != nil {
|
||||||
|
acc = func(p backup.UpdateTierPoint) bool { return accept(conv(p)) }
|
||||||
|
}
|
||||||
|
chosen, ok, seen := a.b.UpdateRestorePoints(ctx, name, acc)
|
||||||
|
out := make([]stacks.UpdateRestorePoint, 0, len(seen))
|
||||||
|
for _, p := range seen {
|
||||||
|
out = append(out, conv(p))
|
||||||
|
}
|
||||||
|
return conv(chosen), ok, out
|
||||||
|
}
|
||||||
|
|
||||||
|
func (a *updateGuardsAdapter) CanBackUp(name string) (bool, string) {
|
||||||
|
if a.b == nil {
|
||||||
|
return false, "backup is not enabled on this box"
|
||||||
|
}
|
||||||
|
return a.b.CanBackUpApp(name)
|
||||||
}
|
}
|
||||||
|
|
||||||
func (a *updateGuardsAdapter) BackupNow(ctx context.Context, name string) error {
|
func (a *updateGuardsAdapter) BackupNow(ctx context.Context, name string) error {
|
||||||
@@ -3404,9 +3420,9 @@ func (a *updateGuardsAdapter) SafetyDump(ctx context.Context, name string) ([]st
|
|||||||
return a.b.WriteUpdateSafetyDump(ctx, name)
|
return a.b.WriteUpdateSafetyDump(ctx, name)
|
||||||
}
|
}
|
||||||
|
|
||||||
func (a *updateGuardsAdapter) HoldAfterFailedUpdate(name string, at, provenCopyAt time.Time) error {
|
func (a *updateGuardsAdapter) HoldAfterFailedUpdate(name string, at time.Time, rp stacks.UpdateRestorePoint) error {
|
||||||
if a.b == nil {
|
if a.b == nil {
|
||||||
return fmt.Errorf("backup is not enabled on this box — the hold cannot be recorded")
|
return fmt.Errorf("backup is not enabled on this box — the hold cannot be recorded")
|
||||||
}
|
}
|
||||||
return a.b.HoldAfterFailedUpdate(name, at, provenCopyAt)
|
return a.b.HoldAfterFailedUpdate(name, at, rp.ProvenAt, rp.Tier)
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,71 @@
|
|||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"go/ast"
|
||||||
|
"go/parser"
|
||||||
|
"go/token"
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
|
||||||
|
"gitea.dooplex.hu/admin/felhom-controller/internal/backup"
|
||||||
|
"gitea.dooplex.hu/admin/felhom-controller/internal/stacks"
|
||||||
|
)
|
||||||
|
|
||||||
|
// R-475 — the update reads EVERY backup tier, through the one adapter.
|
||||||
|
|
||||||
|
func TestR475_TierConstantsAgree(t *testing.T) {
|
||||||
|
if stacks.UpdateTierLocal != backup.UpdateTierLocal ||
|
||||||
|
stacks.UpdateTierSecondDrive != backup.UpdateTierSecondDrive ||
|
||||||
|
stacks.UpdateTierOffsite != backup.UpdateTierOffsite {
|
||||||
|
t.Fatal("stacks and backup number the tiers differently — a hold would name the wrong copy")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// adapterMethodSelectors returns the selector names used in updateGuardsAdapter.<name>'s body.
|
||||||
|
func adapterMethodSelectors(t *testing.T, name string) string {
|
||||||
|
t.Helper()
|
||||||
|
fset := token.NewFileSet()
|
||||||
|
f, err := parser.ParseFile(fset, "main.go", nil, 0)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
for _, d := range f.Decls {
|
||||||
|
fn, ok := d.(*ast.FuncDecl)
|
||||||
|
if !ok || fn.Recv == nil || fn.Name.Name != name || fn.Body == nil {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
star, ok := fn.Recv.List[0].Type.(*ast.StarExpr)
|
||||||
|
if !ok {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if id, ok := star.X.(*ast.Ident); !ok || id.Name != "updateGuardsAdapter" {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
var names []string
|
||||||
|
ast.Inspect(fn.Body, func(n ast.Node) bool {
|
||||||
|
if sel, ok := n.(*ast.SelectorExpr); ok {
|
||||||
|
names = append(names, sel.Sel.Name)
|
||||||
|
}
|
||||||
|
return true
|
||||||
|
})
|
||||||
|
return " " + strings.Join(names, " ") + " "
|
||||||
|
}
|
||||||
|
t.Fatalf("updateGuardsAdapter.%s not found in main.go", name)
|
||||||
|
return ""
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestR475_AdapterReadsEveryTier(t *testing.T) {
|
||||||
|
rp := adapterMethodSelectors(t, "RestorePoints")
|
||||||
|
if !strings.Contains(rp, " UpdateRestorePoints ") {
|
||||||
|
t.Errorf("the adapter must read every tier through backup.UpdateRestorePoints; selectors:%s", rp)
|
||||||
|
}
|
||||||
|
if strings.Contains(rp, " Tier2UnitRestorePoint ") {
|
||||||
|
t.Error("the Tier-2-only predicate must not come back into the update path (R-475)")
|
||||||
|
}
|
||||||
|
if cb := adapterMethodSelectors(t, "CanBackUp"); !strings.Contains(cb, " CanBackUpApp ") {
|
||||||
|
t.Errorf("CanBackUp must ask backup.CanBackUpApp; selectors:%s", cb)
|
||||||
|
}
|
||||||
|
if h := adapterMethodSelectors(t, "HoldAfterFailedUpdate"); !strings.Contains(h, " Tier ") {
|
||||||
|
t.Errorf("the hold must be told the chosen TIER, or it cannot name it; selectors:%s", h)
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -24,8 +24,9 @@ import (
|
|||||||
// every path here refuses, or fails at the pin (the app's catalog template is absent on purpose).
|
// every path here refuses, or fails at the pin (the app's catalog template is absent on purpose).
|
||||||
|
|
||||||
type apiFakeGuards struct {
|
type apiFakeGuards struct {
|
||||||
b *backup.Manager
|
b *backup.Manager
|
||||||
rp stacks.UpdateRestorePoint
|
points []stacks.UpdateRestorePoint
|
||||||
|
cannotBackUp bool
|
||||||
// blindToHolds makes the manager-side preflight NOT see holds, so a test can prove the ROUTER's
|
// blindToHolds makes the manager-side preflight NOT see holds, so a test can prove the ROUTER's
|
||||||
// own hold check refuses — the two layers are each pinned separately (the preflight's by
|
// own hold check refuses — the two layers are each pinned separately (the preflight's by
|
||||||
// TestSlice4_D_CheapRefusals/held). Without it, removing either layer passes inertly, because the
|
// TestSlice4_D_CheapRefusals/held). Without it, removing either layer passes inertly, because the
|
||||||
@@ -40,14 +41,20 @@ func (g *apiFakeGuards) HoldFor(n string) (bool, string) {
|
|||||||
return g.b.RestoreHoldFor(n)
|
return g.b.RestoreHoldFor(n)
|
||||||
}
|
}
|
||||||
func (g *apiFakeGuards) Busy(string) (bool, string) { return false, "" }
|
func (g *apiFakeGuards) Busy(string) (bool, string) { return false, "" }
|
||||||
func (g *apiFakeGuards) RestorePoint(string) (stacks.UpdateRestorePoint, error) {
|
func (g *apiFakeGuards) RestorePoints(_ context.Context, _ string, accept func(stacks.UpdateRestorePoint) bool) (stacks.UpdateRestorePoint, bool, []stacks.UpdateRestorePoint) {
|
||||||
return g.rp, nil
|
for _, p := range g.points {
|
||||||
|
if accept == nil || accept(p) {
|
||||||
|
return p, true, g.points
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return stacks.UpdateRestorePoint{}, false, g.points
|
||||||
}
|
}
|
||||||
|
func (g *apiFakeGuards) CanBackUp(string) (bool, string) { return !g.cannotBackUp, "fake: no drive" }
|
||||||
func (g *apiFakeGuards) BackupNow(context.Context, string) error { return nil }
|
func (g *apiFakeGuards) BackupNow(context.Context, string) error { return nil }
|
||||||
func (g *apiFakeGuards) SafetyDump(context.Context, string) ([]string, error) {
|
func (g *apiFakeGuards) SafetyDump(context.Context, string) ([]string, error) {
|
||||||
return nil, nil
|
return nil, nil
|
||||||
}
|
}
|
||||||
func (g *apiFakeGuards) HoldAfterFailedUpdate(string, time.Time, time.Time) error { return nil }
|
func (g *apiFakeGuards) HoldAfterFailedUpdate(string, time.Time, stacks.UpdateRestorePoint) error { return nil }
|
||||||
|
|
||||||
const slice4AppYAML = "deployed: true\nenv: {}\npinned_images:\n app: nginx:1.27\n"
|
const slice4AppYAML = "deployed: true\nenv: {}\npinned_images:\n app: nginx:1.27\n"
|
||||||
|
|
||||||
@@ -82,7 +89,7 @@ func newSlice4Router(t *testing.T) (*Router, *settings.Settings, *apiFakeGuards,
|
|||||||
t.Fatal(err)
|
t.Fatal(err)
|
||||||
}
|
}
|
||||||
b := backup.NewManager(cfg, sett, lg)
|
b := backup.NewManager(cfg, sett, lg)
|
||||||
g := &apiFakeGuards{b: b, rp: stacks.UpdateRestorePoint{Restorable: true, Proven: true, ProvenAt: time.Now().Add(-time.Hour)}}
|
g := &apiFakeGuards{b: b, points: []stacks.UpdateRestorePoint{{Tier: stacks.UpdateTierSecondDrive, ProvenAt: time.Now().Add(-time.Hour)}}}
|
||||||
m.SetUpdateGuards(g)
|
m.SetUpdateGuards(g)
|
||||||
return &Router{cfg: cfg, stackMgr: m, backupMgr: b, logger: lg}, sett, g, dir
|
return &Router{cfg: cfg, stackMgr: m, backupMgr: b, logger: lg}, sett, g, dir
|
||||||
}
|
}
|
||||||
@@ -105,7 +112,7 @@ func postUpdate(t *testing.T, r *Router) (int, apiResponse) {
|
|||||||
// message — which is what proves the router line is the one doing it.
|
// message — which is what proves the router line is the one doing it.
|
||||||
func TestR439_UpdateOfAHeldAppIsRefused(t *testing.T) {
|
func TestR439_UpdateOfAHeldAppIsRefused(t *testing.T) {
|
||||||
r, sett, g, dir := newSlice4Router(t)
|
r, sett, g, dir := newSlice4Router(t)
|
||||||
g.rp = stacks.UpdateRestorePoint{} // the preflight's own refusal would say "no backup" — not the hold
|
g.points, g.cannotBackUp = nil, true // the preflight's own refusal would say "no backup" — not the hold
|
||||||
g.blindToHolds = true // only the router's line can produce the hold's sentence
|
g.blindToHolds = true // only the router's line can produce the hold's sentence
|
||||||
if err := sett.SetRestoreHold(settings.RestoreHold{Stack: "app", At: "2026-09-13T08:00:00Z", Reason: settings.HoldReasonUpdateFailed, CopyDate: "2026-09-13T01:30:00Z"}); err != nil {
|
if err := sett.SetRestoreHold(settings.RestoreHold{Stack: "app", At: "2026-09-13T08:00:00Z", Reason: settings.HoldReasonUpdateFailed, CopyDate: "2026-09-13T01:30:00Z"}); err != nil {
|
||||||
t.Fatal(err)
|
t.Fatal(err)
|
||||||
@@ -124,7 +131,7 @@ func TestR439_UpdateOfAHeldAppIsRefused(t *testing.T) {
|
|||||||
|
|
||||||
func TestSlice4_Router_NoBackupIs409AndRecordsNothing(t *testing.T) {
|
func TestSlice4_Router_NoBackupIs409AndRecordsNothing(t *testing.T) {
|
||||||
r, _, g, dir := newSlice4Router(t)
|
r, _, g, dir := newSlice4Router(t)
|
||||||
g.rp = stacks.UpdateRestorePoint{Restorable: false}
|
g.points, g.cannotBackUp = nil, true
|
||||||
before, _ := os.ReadFile(filepath.Join(dir, "app.yaml"))
|
before, _ := os.ReadFile(filepath.Join(dir, "app.yaml"))
|
||||||
code, resp := postUpdate(t, r)
|
code, resp := postUpdate(t, r)
|
||||||
if code != http.StatusConflict || resp.Error != fmt.Sprintf(stacks.MsgUpdateNoBackupFmt, "app") {
|
if code != http.StatusConflict || resp.Error != fmt.Sprintf(stacks.MsgUpdateNoBackupFmt, "app") {
|
||||||
|
|||||||
@@ -173,6 +173,13 @@ type Manager struct {
|
|||||||
// updatingCheck (slice 4) — nil-safe; see isHeld / SetUpdatingCheck.
|
// updatingCheck (slice 4) — nil-safe; see isHeld / SetUpdatingCheck.
|
||||||
updatingCheck func(stackName string) bool
|
updatingCheck func(stackName string) bool
|
||||||
|
|
||||||
|
// R-475 update-precondition seams, one per tier. Nil → the real Tier2UnitRestorePoint /
|
||||||
|
// ListRestorePoints / OffsiteInventoryList. They let a test reach Tier 1 and Tier 3 without a drive
|
||||||
|
// or a restic repository; see UpdateRestorePoints.
|
||||||
|
updateTier2PointFn func(stackName string) (Tier2RestorePoint, error)
|
||||||
|
updateTier1PointsFn func(stackName string) ([]RestorePoint, bool)
|
||||||
|
updateOffsiteInvFn func(ctx context.Context) (OffsiteInventory, error)
|
||||||
|
|
||||||
// R-354 volume-REPLAY seam — the mirror of the F17 DB seams above, so the off-site path's new
|
// R-354 volume-REPLAY seam — the mirror of the F17 DB seams above, so the off-site path's new
|
||||||
// volume leg is unit-testable without Docker. Nil → the real restoreDockerVolumesFrom.
|
// volume leg is unit-testable without Docker. Nil → the real restoreDockerVolumesFrom.
|
||||||
volumeReplayFrom func(stackName, dumpDir string) (int, error)
|
volumeReplayFrom func(stackName, dumpDir string) (int, error)
|
||||||
|
|||||||
@@ -344,7 +344,11 @@ func (m *Manager) RestoreHoldFor(stack string) (bool, string) {
|
|||||||
if h.CopyDate != "" {
|
if h.CopyDate != "" {
|
||||||
copyDate = fmtHoldTime(h.CopyDate)
|
copyDate = fmtHoldTime(h.CopyDate)
|
||||||
}
|
}
|
||||||
return true, fmt.Sprintf(UpdateHoldFmt, stack, fmtHoldTime(h.At), copyDate)
|
// R-475: name the tier when the hold recorded one; an older hold keeps its own sentence.
|
||||||
|
if label := UpdateTierLabel(h.CopyTier); label != "" && h.CopyDate != "" {
|
||||||
|
return true, fmt.Sprintf(UpdateHoldFmt, stack, fmtHoldTime(h.At), label, copyDate)
|
||||||
|
}
|
||||||
|
return true, fmt.Sprintf(UpdateHoldLegacyFmt, stack, fmtHoldTime(h.At), copyDate)
|
||||||
}
|
}
|
||||||
when := h.At
|
when := h.At
|
||||||
if t, err := time.Parse(time.RFC3339, h.At); err == nil {
|
if t, err := time.Parse(time.RFC3339, h.At); err == nil {
|
||||||
@@ -824,6 +828,11 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string, ack
|
|||||||
if err := restartStack(); err != nil {
|
if err := restartStack(); err != nil {
|
||||||
return res, fmt.Errorf("a(z) %s újraindítása sikertelen a fájlok visszaállítása után: %w", stack, err)
|
return res, fmt.Errorf("a(z) %s újraindítása sikertelen a fájlok visszaállítása után: %w", stack, err)
|
||||||
}
|
}
|
||||||
|
// R-475: an update hold may now name the OFF-SITE copy, and this is the route back it names. The
|
||||||
|
// unit restore clears it in RestoreFromRecoveryUnitAt; this path never went through that function,
|
||||||
|
// so without this line a successful off-site restore would leave the app refusing its next start.
|
||||||
|
// Pinned by TestR475_OffsiteRestoreClearsAnUpdateHold.
|
||||||
|
m.clearUpdateHoldAfterRestore(stack)
|
||||||
if err := m.waitForHealthy(stack, 90*time.Second); err != nil {
|
if err := m.waitForHealthy(stack, 90*time.Second); err != nil {
|
||||||
m.logger.Printf("[WARN] [offbox] %s reconstituted but health check failed: %v", stack, err)
|
m.logger.Printf("[WARN] [offbox] %s reconstituted but health check failed: %v", stack, err)
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -68,7 +68,7 @@ func TestSlice4_UpdateHoldTextNamesTheTimeAndTheCopy(t *testing.T) {
|
|||||||
m := &Manager{logger: log.New(io.Discard, "", 0), settings: sett}
|
m := &Manager{logger: log.New(io.Discard, "", 0), settings: sett}
|
||||||
at := time.Date(2026, 9, 13, 8, 0, 0, 0, time.UTC)
|
at := time.Date(2026, 9, 13, 8, 0, 0, 0, time.UTC)
|
||||||
copyAt := time.Date(2026, 9, 13, 1, 30, 0, 0, time.UTC)
|
copyAt := time.Date(2026, 9, 13, 1, 30, 0, 0, time.UTC)
|
||||||
if err := m.HoldAfterFailedUpdate("bookstack", at, copyAt); err != nil {
|
if err := m.HoldAfterFailedUpdate("bookstack", at, copyAt, UpdateTierSecondDrive); err != nil {
|
||||||
t.Fatal(err)
|
t.Fatal(err)
|
||||||
}
|
}
|
||||||
h, ok := sett.GetRestoreHold("bookstack")
|
h, ok := sett.GetRestoreHold("bookstack")
|
||||||
@@ -77,7 +77,7 @@ func TestSlice4_UpdateHoldTextNamesTheTimeAndTheCopy(t *testing.T) {
|
|||||||
}
|
}
|
||||||
held, why := m.RestoreHoldFor("bookstack")
|
held, why := m.RestoreHoldFor("bookstack")
|
||||||
// Budapest is UTC+2 in September: 08:00Z → 10:00, 01:30Z → 03:30.
|
// Budapest is UTC+2 in September: 08:00Z → 10:00, 01:30Z → 03:30.
|
||||||
want := fmt.Sprintf(UpdateHoldFmt, "bookstack", "2026-09-13 10:00", "2026-09-13 03:30")
|
want := fmt.Sprintf(UpdateHoldFmt, "bookstack", "2026-09-13 10:00", "második meghajtó", "2026-09-13 03:30")
|
||||||
if !held || why != want {
|
if !held || why != want {
|
||||||
t.Errorf("hold text =\n%q\nwant\n%q", why, want)
|
t.Errorf("hold text =\n%q\nwant\n%q", why, want)
|
||||||
}
|
}
|
||||||
@@ -98,7 +98,7 @@ func TestSlice4_RestoreHoldTextIsUnchanged(t *testing.T) {
|
|||||||
func TestSlice4_ASuccessfulRestoreClearsOnlyAnUpdateHold(t *testing.T) {
|
func TestSlice4_ASuccessfulRestoreClearsOnlyAnUpdateHold(t *testing.T) {
|
||||||
sett := slice4Settings(t)
|
sett := slice4Settings(t)
|
||||||
m := &Manager{logger: log.New(io.Discard, "", 0), settings: sett}
|
m := &Manager{logger: log.New(io.Discard, "", 0), settings: sett}
|
||||||
_ = m.HoldAfterFailedUpdate("upd", time.Now(), time.Now())
|
_ = m.HoldAfterFailedUpdate("upd", time.Now(), time.Now(), UpdateTierLocal)
|
||||||
_ = sett.SetRestoreHold(settings.RestoreHold{Stack: "rst", At: "2026-08-22T14:00:00Z"})
|
_ = sett.SetRestoreHold(settings.RestoreHold{Stack: "rst", At: "2026-08-22T14:00:00Z"})
|
||||||
m.clearUpdateHoldAfterRestore("upd")
|
m.clearUpdateHoldAfterRestore("upd")
|
||||||
m.clearUpdateHoldAfterRestore("rst")
|
m.clearUpdateHoldAfterRestore("rst")
|
||||||
@@ -136,7 +136,7 @@ func TestSlice4_UpdateBusy(t *testing.T) {
|
|||||||
func TestSlice4_NightlyLegsLeaveAHeldAppAlone(t *testing.T) {
|
func TestSlice4_NightlyLegsLeaveAHeldAppAlone(t *testing.T) {
|
||||||
h := newAdmissionHarness(t, "held", "free")
|
h := newAdmissionHarness(t, "held", "free")
|
||||||
h.m.settings = slice4Settings(t)
|
h.m.settings = slice4Settings(t)
|
||||||
if err := h.m.HoldAfterFailedUpdate("held", time.Now(), time.Now()); err != nil {
|
if err := h.m.HoldAfterFailedUpdate("held", time.Now(), time.Now(), UpdateTierSecondDrive); err != nil {
|
||||||
t.Fatal(err)
|
t.Fatal(err)
|
||||||
}
|
}
|
||||||
h.m.runVolumeDumps()
|
h.m.runVolumeDumps()
|
||||||
|
|||||||
@@ -4,6 +4,7 @@ import (
|
|||||||
"context"
|
"context"
|
||||||
"errors"
|
"errors"
|
||||||
"fmt"
|
"fmt"
|
||||||
|
"os"
|
||||||
"path/filepath"
|
"path/filepath"
|
||||||
"time"
|
"time"
|
||||||
|
|
||||||
@@ -93,6 +94,149 @@ func (p Tier2RestorePoint) ProvenCopyTime() (time.Time, bool) {
|
|||||||
return t, true
|
return t, true
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// ── R-475: any backup tier lets an app update (operator ruling 2026-09-13, controller v0.239.0) ────
|
||||||
|
//
|
||||||
|
// Until v0.239.0 the update's precondition was Tier2UnitRestorePoint alone, so an app with no second
|
||||||
|
// drive could never be updated — even with a fresh recovery unit on its own drive and an off-site
|
||||||
|
// snapshot from last night. The ruling: every backup counts. Tier2UnitRestorePoint itself is NOT
|
||||||
|
// changed; the backups page still calls it for the „Teljes visszaállítás" action, which really does
|
||||||
|
// restore from the second drive only.
|
||||||
|
|
||||||
|
// Backup tiers, as the update precondition and the hold sentence name them.
|
||||||
|
const (
|
||||||
|
UpdateTierLocal = 1 // the app's own recovery unit on its drive — „helyi" on the restore page
|
||||||
|
UpdateTierSecondDrive = 2 // the Tier-2 mirror on another drive
|
||||||
|
UpdateTierOffsite = 3 // the off-site restic repository
|
||||||
|
)
|
||||||
|
|
||||||
|
// updateTierOrder is the preference order the ruling set: the second drive, then the app's own unit,
|
||||||
|
// then off-site. The first tier holding a copy the caller ACCEPTS is chosen.
|
||||||
|
var updateTierOrder = []int{UpdateTierSecondDrive, UpdateTierLocal, UpdateTierOffsite}
|
||||||
|
|
||||||
|
// UpdateTierLabel is a tier's name in the customer's hold sentence. "" for an unknown tier.
|
||||||
|
func UpdateTierLabel(tier int) string {
|
||||||
|
switch tier {
|
||||||
|
case UpdateTierSecondDrive:
|
||||||
|
return "második meghajtó"
|
||||||
|
case UpdateTierLocal:
|
||||||
|
return "saját meghajtó"
|
||||||
|
case UpdateTierOffsite:
|
||||||
|
return "távoli mentés"
|
||||||
|
}
|
||||||
|
return ""
|
||||||
|
}
|
||||||
|
|
||||||
|
// updateOffsiteCheckTimeout bounds the off-site lookup. An update must not stall on an unreachable
|
||||||
|
// Storage Box: past this the off-site copy counts as ABSENT (with a WARN), and the update carries on
|
||||||
|
// with backing up first. A var only so a test can shorten it.
|
||||||
|
var updateOffsiteCheckTimeout = 15 * time.Second
|
||||||
|
|
||||||
|
// UpdateTierPoint is one proven, restorable copy of an app on one tier.
|
||||||
|
type UpdateTierPoint struct {
|
||||||
|
Tier int
|
||||||
|
// At is when the data in that copy was last proven written: Tier 2 ProvenCopyTime, Tier 1 the
|
||||||
|
// newest artifact of the unit (ListRestorePoints), Tier 3 the newest snapshot for the app.
|
||||||
|
At time.Time
|
||||||
|
}
|
||||||
|
|
||||||
|
// UpdateRestorePoints walks the tiers in preference order (2, 1, 3) and returns the FIRST copy that
|
||||||
|
// accept admits (nil accepts any), whether one was found, and every copy it looked at on the way.
|
||||||
|
//
|
||||||
|
// It stops at the first accepted copy, so a box with a fresh second-drive copy never touches the
|
||||||
|
// network. The AGE rule is the caller's (stacks applies backup_max_age through accept) — that is what
|
||||||
|
// makes "the age rule applies to whichever tier is chosen" one rule, not three (R-475 Scenario M).
|
||||||
|
func (m *Manager) UpdateRestorePoints(ctx context.Context, stackName string, accept func(UpdateTierPoint) bool) (UpdateTierPoint, bool, []UpdateTierPoint) {
|
||||||
|
var seen []UpdateTierPoint
|
||||||
|
for _, tier := range updateTierOrder {
|
||||||
|
p, ok := m.updateTierPoint(ctx, stackName, tier)
|
||||||
|
if !ok {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
seen = append(seen, p)
|
||||||
|
if accept == nil || accept(p) {
|
||||||
|
return p, true, seen
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return UpdateTierPoint{}, false, seen
|
||||||
|
}
|
||||||
|
|
||||||
|
func (m *Manager) updateTierPoint(ctx context.Context, stackName string, tier int) (UpdateTierPoint, bool) {
|
||||||
|
switch tier {
|
||||||
|
case UpdateTierSecondDrive:
|
||||||
|
get := m.updateTier2PointFn
|
||||||
|
if get == nil {
|
||||||
|
get = m.Tier2UnitRestorePoint
|
||||||
|
}
|
||||||
|
rp, err := get(stackName)
|
||||||
|
if err != nil {
|
||||||
|
if m.isDebug() {
|
||||||
|
m.logger.Printf("[DEBUG] [backup] update precondition for %s: no Tier-2 copy (%v)", stackName, err)
|
||||||
|
}
|
||||||
|
return UpdateTierPoint{}, false
|
||||||
|
}
|
||||||
|
at, ok := rp.ProvenCopyTime()
|
||||||
|
return UpdateTierPoint{Tier: tier, At: at}, ok
|
||||||
|
case UpdateTierLocal:
|
||||||
|
list := m.updateTier1PointsFn
|
||||||
|
if list == nil {
|
||||||
|
list = m.ListRestorePoints
|
||||||
|
}
|
||||||
|
pts, _ := list(stackName)
|
||||||
|
for _, rp := range pts {
|
||||||
|
if at, err := time.Parse(time.RFC3339, rp.Time); err == nil {
|
||||||
|
return UpdateTierPoint{Tier: tier, At: at}, true
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return UpdateTierPoint{}, false
|
||||||
|
case UpdateTierOffsite:
|
||||||
|
inv := m.updateOffsiteInvFn
|
||||||
|
if inv == nil {
|
||||||
|
if m.settings == nil || !m.OffboxConfigured() {
|
||||||
|
return UpdateTierPoint{}, false
|
||||||
|
}
|
||||||
|
inv = m.OffsiteInventoryList
|
||||||
|
}
|
||||||
|
cctx, cancel := context.WithTimeout(ctx, updateOffsiteCheckTimeout)
|
||||||
|
defer cancel()
|
||||||
|
got, err := inv(cctx)
|
||||||
|
if err != nil {
|
||||||
|
if !errors.Is(err, errNoOffsiteTarget) {
|
||||||
|
m.logger.Printf("[WARN] [backup] update precondition for %s: the off-site copy could not be checked within %s (%v) — counted as ABSENT", stackName, updateOffsiteCheckTimeout, err)
|
||||||
|
}
|
||||||
|
return UpdateTierPoint{}, false
|
||||||
|
}
|
||||||
|
for _, a := range got.Apps {
|
||||||
|
if a.App == stackName && !a.LatestAt.IsZero() {
|
||||||
|
return UpdateTierPoint{Tier: tier, At: a.LatestAt}, true
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return UpdateTierPoint{}, false
|
||||||
|
}
|
||||||
|
|
||||||
|
// CanBackUpApp reports whether "back up first" can run for this app at all right now — the second
|
||||||
|
// half of R-475 Scenario L: an app with no copy anywhere is refused only when this is false too.
|
||||||
|
// Cheap and read-only; RunAppBackupNow re-checks everything when it actually runs.
|
||||||
|
func (m *Manager) CanBackUpApp(stackName string) (bool, string) {
|
||||||
|
if m == nil {
|
||||||
|
return false, "backup is not enabled on this box"
|
||||||
|
}
|
||||||
|
if m.stackProvider == nil {
|
||||||
|
return false, "stack provider not configured"
|
||||||
|
}
|
||||||
|
if m.migrationActive() {
|
||||||
|
return false, "a data migration is running"
|
||||||
|
}
|
||||||
|
drivePath := m.GetAppDrivePath(stackName)
|
||||||
|
if drivePath == "" || !filepath.IsAbs(drivePath) {
|
||||||
|
return false, "the app's drive cannot be resolved"
|
||||||
|
}
|
||||||
|
if m.settings != nil && (m.settings.IsDisconnected(drivePath) || m.settings.IsDecommissioned(drivePath)) {
|
||||||
|
return false, fmt.Sprintf("the app's drive %s is not available", drivePath)
|
||||||
|
}
|
||||||
|
return true, ""
|
||||||
|
}
|
||||||
|
|
||||||
// UpdateBusy reports whether something else is ALREADY touching this app's data, which refuses an
|
// UpdateBusy reports whether something else is ALREADY touching this app's data, which refuses an
|
||||||
// update before anything moves (slice 4 Scenario D). The reason is operator-English; the customer
|
// update before anything moves (slice 4 Scenario D). The reason is operator-English; the customer
|
||||||
// sentence is chosen by the caller.
|
// sentence is chosen by the caller.
|
||||||
@@ -142,6 +286,7 @@ func (m *Manager) RunAppBackupNow(ctx context.Context, stackName string) error {
|
|||||||
}
|
}
|
||||||
m.logger.Printf("[INFO] [backup] update pre-backup for %s: starting (DB dump → volume dump → unit capture → Tier 2)", stackName)
|
m.logger.Printf("[INFO] [backup] update pre-backup for %s: starting (DB dump → volume dump → unit capture → Tier 2)", stackName)
|
||||||
start := time.Now()
|
start := time.Now()
|
||||||
|
var nsRoot string
|
||||||
legErr := func() error {
|
legErr := func() error {
|
||||||
defer m.releaseRunning()
|
defer m.releaseRunning()
|
||||||
defer m.beginAdmissionRun()()
|
defer m.beginAdmissionRun()()
|
||||||
@@ -156,7 +301,7 @@ func (m *Manager) RunAppBackupNow(ctx context.Context, stackName string) error {
|
|||||||
if !m.admitApp(stackName) {
|
if !m.admitApp(stackName) {
|
||||||
return fmt.Errorf("nincs elég szabad hely a mentéshez a(z) %s meghajtón", drivePath)
|
return fmt.Errorf("nincs elég szabad hely a mentéshez a(z) %s meghajtón", drivePath)
|
||||||
}
|
}
|
||||||
nsRoot := m.namespaceRoot(drivePath)
|
nsRoot = m.namespaceRoot(drivePath)
|
||||||
|
|
||||||
discover := m.discoverDBs
|
discover := m.discoverDBs
|
||||||
if discover == nil {
|
if discover == nil {
|
||||||
@@ -203,16 +348,35 @@ func (m *Manager) RunAppBackupNow(ctx context.Context, stackName string) error {
|
|||||||
return legErr
|
return legErr
|
||||||
}
|
}
|
||||||
|
|
||||||
|
m.updatePreBackupTail(stackName, nsRoot, time.Now())
|
||||||
|
m.logger.Printf("[INFO] [backup] update pre-backup for %s: complete in %s", stackName, time.Since(start).Round(time.Millisecond))
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// updatePreBackupTail is what "back up first" does after the capture succeeded (R-475).
|
||||||
|
//
|
||||||
|
// 1. It marks the app's OWN unit as proven current NOW. CaptureRecoveryUnit leaves the manifest alone
|
||||||
|
// when nothing changed (the checksum skip), and Tier 1's age is the newest artifact's mtime — so on
|
||||||
|
// an app with no database and no named volume a fresh "back up first" would leave Tier 1 as old as
|
||||||
|
// its last definition change, and the update would be refused forever. That is the trap
|
||||||
|
// ProvenCopyTime documents for Tier 2, one tier down. The capture has just compared the unit with
|
||||||
|
// the live definition, so "current as of now" is exactly what it established.
|
||||||
|
// 2. It runs the Tier-2 copy, and a Tier-2 failure is a WARN, not a failure: the update may lean on
|
||||||
|
// any tier, and the Tier-1 unit it can lean on was just written. Pinned by
|
||||||
|
// TestR475_PreBackupTail_Tier2FailureIsAWarnAndTheOwnUnitIsFresh.
|
||||||
|
func (m *Manager) updatePreBackupTail(stackName, nsRoot string, now time.Time) {
|
||||||
|
if nsRoot != "" {
|
||||||
|
if err := os.Chtimes(RecoveryUnitManifestPath(nsRoot, stackName), now, now); err != nil {
|
||||||
|
m.logger.Printf("[WARN] [backup] update pre-backup for %s: could not mark the recovery unit as proven current (%v) — its own-unit copy may read older than it is", stackName, err)
|
||||||
|
}
|
||||||
|
}
|
||||||
runOne := m.perAppTier2
|
runOne := m.perAppTier2
|
||||||
if runOne == nil {
|
if runOne == nil {
|
||||||
runOne = m.RunTier2
|
runOne = m.RunTier2
|
||||||
}
|
}
|
||||||
if err := runOne(stackName); err != nil {
|
if err := runOne(stackName); err != nil {
|
||||||
m.logger.Printf("[ERROR] [backup] update pre-backup for %s: Tier 2 copy FAILED: %v", stackName, err)
|
m.logger.Printf("[WARN] [backup] update pre-backup for %s: Tier 2 copy FAILED: %v — not fatal: the app's own recovery unit was just captured, and an update may lean on any tier (R-475)", stackName, err)
|
||||||
return fmt.Errorf("a másodlagos másolat elkészítése sikertelen: %w", err)
|
|
||||||
}
|
}
|
||||||
m.logger.Printf("[INFO] [backup] update pre-backup for %s: complete in %s", stackName, time.Since(start).Round(time.Millisecond))
|
|
||||||
return nil
|
|
||||||
}
|
}
|
||||||
|
|
||||||
// WriteUpdateSafetyDump takes the last-minute database copy an update makes just before it moves the
|
// WriteUpdateSafetyDump takes the last-minute database copy an update makes just before it moves the
|
||||||
@@ -244,9 +408,18 @@ func (m *Manager) WriteUpdateSafetyDump(ctx context.Context, stackName string) (
|
|||||||
}
|
}
|
||||||
|
|
||||||
// UpdateHoldFmt is the customer sentence for an app held after a failed update. Arguments: the app,
|
// UpdateHoldFmt is the customer sentence for an app held after a failed update. Arguments: the app,
|
||||||
// the time of the failure, and the PROVEN date of the copy it can be restored from. One named string so
|
// the time of the failure, the TIER of the copy it can be restored from (UpdateTierLabel), and that
|
||||||
// a test asserts it verbatim instead of retyping Hungarian (R-364).
|
// copy's PROVEN date. One named string so a test asserts it verbatim instead of retyping Hungarian
|
||||||
|
// (R-364). Since v0.239.0 (R-475) it names the tier: the copy may be on any of three, and each is
|
||||||
|
// restored from a different place on the Mentések page.
|
||||||
const UpdateHoldFmt = "A(z) %s frissítése %s-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. " +
|
const UpdateHoldFmt = "A(z) %s frissítése %s-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. " +
|
||||||
|
"Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. " +
|
||||||
|
"Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: %s, %s."
|
||||||
|
|
||||||
|
// UpdateHoldLegacyFmt is the v0.237.0–v0.238.1 sentence, kept for a hold written before the tier was
|
||||||
|
// recorded (CopyTier 0) — every such hold named a Tier-2 copy, but it did not SAY so, and rewriting
|
||||||
|
// it now would state a fact the record does not hold.
|
||||||
|
const UpdateHoldLegacyFmt = "A(z) %s frissítése %s-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. " +
|
||||||
"Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. " +
|
"Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. " +
|
||||||
"Visszaállítható a(z) %s-i biztonsági mentésből a Mentések oldalon."
|
"Visszaállítható a(z) %s-i biztonsági mentésből a Mentések oldalon."
|
||||||
|
|
||||||
@@ -276,7 +449,7 @@ func fmtHoldTime(rfc3339 string) string {
|
|||||||
// has just stopped the app on the strength of this record, and an unrecorded hold is a stopped app
|
// has just stopped the app on the strength of this record, and an unrecorded hold is a stopped app
|
||||||
// that the next restart button will quietly start again. The caller logs it at ERROR and keeps the
|
// that the next restart button will quietly start again. The caller logs it at ERROR and keeps the
|
||||||
// failure on the page.
|
// failure on the page.
|
||||||
func (m *Manager) HoldAfterFailedUpdate(stackName string, at time.Time, copyDate time.Time) error {
|
func (m *Manager) HoldAfterFailedUpdate(stackName string, at time.Time, copyDate time.Time, copyTier int) error {
|
||||||
if m == nil || m.settings == nil {
|
if m == nil || m.settings == nil {
|
||||||
return fmt.Errorf("no settings wired — the update hold for %s cannot be persisted", stackName)
|
return fmt.Errorf("no settings wired — the update hold for %s cannot be persisted", stackName)
|
||||||
}
|
}
|
||||||
@@ -287,11 +460,12 @@ func (m *Manager) HoldAfterFailedUpdate(stackName string, at time.Time, copyDate
|
|||||||
}
|
}
|
||||||
if !copyDate.IsZero() {
|
if !copyDate.IsZero() {
|
||||||
h.CopyDate = copyDate.UTC().Format(time.RFC3339)
|
h.CopyDate = copyDate.UTC().Format(time.RFC3339)
|
||||||
|
h.CopyTier = copyTier
|
||||||
}
|
}
|
||||||
if err := m.settings.SetRestoreHold(h); err != nil {
|
if err := m.settings.SetRestoreHold(h); err != nil {
|
||||||
return fmt.Errorf("persisting the update hold for %s: %w", stackName, err)
|
return fmt.Errorf("persisting the update hold for %s: %w", stackName, err)
|
||||||
}
|
}
|
||||||
m.logger.Printf("[WARN] [backup] %s is HELD STOPPED after a failed update (restore point: %s)", stackName, h.CopyDate)
|
m.logger.Printf("[WARN] [backup] %s is HELD STOPPED after a failed update (restore point: tier %d %q, %s)", stackName, h.CopyTier, UpdateTierLabel(h.CopyTier), h.CopyDate)
|
||||||
return nil
|
return nil
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,317 @@
|
|||||||
|
package backup
|
||||||
|
|
||||||
|
import (
|
||||||
|
"bytes"
|
||||||
|
"context"
|
||||||
|
"errors"
|
||||||
|
"fmt"
|
||||||
|
"go/ast"
|
||||||
|
"go/parser"
|
||||||
|
"go/token"
|
||||||
|
"io"
|
||||||
|
"log"
|
||||||
|
"os"
|
||||||
|
"path/filepath"
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"gitea.dooplex.hu/admin/felhom-controller/internal/settings"
|
||||||
|
)
|
||||||
|
|
||||||
|
// R-475 — the backup side of "any backup tier lets an app update" (controller v0.239.0).
|
||||||
|
|
||||||
|
var r475T0 = time.Date(2026, 9, 13, 12, 0, 0, 0, time.UTC)
|
||||||
|
|
||||||
|
func r475Manager() (*Manager, *bytes.Buffer) {
|
||||||
|
var buf bytes.Buffer
|
||||||
|
return &Manager{logger: log.New(&buf, "", 0)}, &buf
|
||||||
|
}
|
||||||
|
|
||||||
|
func tier2At(at time.Time) func(string) (Tier2RestorePoint, error) {
|
||||||
|
return func(string) (Tier2RestorePoint, error) {
|
||||||
|
ts := at.UTC().Format(time.RFC3339)
|
||||||
|
return Tier2RestorePoint{Restorable: true, CopyDateProven: true, CopyLastSuccess: ts, CopyDate: ts}, nil
|
||||||
|
}
|
||||||
|
}
|
||||||
|
func noTier2(string) (Tier2RestorePoint, error) {
|
||||||
|
return Tier2RestorePoint{}, errors.New("no Tier-2 copy recorded for this app")
|
||||||
|
}
|
||||||
|
func tier1At(at time.Time) func(string) ([]RestorePoint, bool) {
|
||||||
|
return func(string) ([]RestorePoint, bool) {
|
||||||
|
return []RestorePoint{{Time: at.UTC().Format(time.RFC3339), ShortID: "helyi", Tier: 1}}, true
|
||||||
|
}
|
||||||
|
}
|
||||||
|
func noTier1(string) ([]RestorePoint, bool) { return []RestorePoint{}, true }
|
||||||
|
func offsiteWith(app string, at time.Time) func(context.Context) (OffsiteInventory, error) {
|
||||||
|
return func(context.Context) (OffsiteInventory, error) {
|
||||||
|
return OffsiteInventory{Apps: []OffsiteInventoryApp{{App: app, LatestAt: at}}}, nil
|
||||||
|
}
|
||||||
|
}
|
||||||
|
func noOffsiteTarget(context.Context) (OffsiteInventory, error) {
|
||||||
|
return OffsiteInventory{}, errNoOffsiteTarget
|
||||||
|
}
|
||||||
|
|
||||||
|
func tiersOf(ps []UpdateTierPoint) []int {
|
||||||
|
out := make([]int, 0, len(ps))
|
||||||
|
for _, p := range ps {
|
||||||
|
out = append(out, p.Tier)
|
||||||
|
}
|
||||||
|
return out
|
||||||
|
}
|
||||||
|
|
||||||
|
// G — the order is 2, 1, 3, and the walk STOPS at the first accepted copy (so a box with a fresh
|
||||||
|
// second-drive copy never reaches the network).
|
||||||
|
func TestR475_G_TierOrderIsSecondDriveThenOwnUnitThenOffsite(t *testing.T) {
|
||||||
|
m, _ := r475Manager()
|
||||||
|
offsiteCalls := 0
|
||||||
|
m.updateTier2PointFn = tier2At(r475T0.Add(-1 * time.Hour))
|
||||||
|
m.updateTier1PointsFn = tier1At(r475T0.Add(-2 * time.Hour))
|
||||||
|
m.updateOffsiteInvFn = func(ctx context.Context) (OffsiteInventory, error) {
|
||||||
|
offsiteCalls++
|
||||||
|
return offsiteWith("gokapi", r475T0.Add(-3*time.Hour))(ctx)
|
||||||
|
}
|
||||||
|
ctx := context.Background()
|
||||||
|
|
||||||
|
p, ok, seen := m.UpdateRestorePoints(ctx, "gokapi", nil)
|
||||||
|
if !ok || p.Tier != UpdateTierSecondDrive || fmt.Sprint(tiersOf(seen)) != "[2]" || offsiteCalls != 0 {
|
||||||
|
t.Errorf("all three present: want tier 2 and a stop; got %+v ok=%v seen=%v offsite calls=%d", p, ok, tiersOf(seen), offsiteCalls)
|
||||||
|
}
|
||||||
|
p, ok, seen = m.UpdateRestorePoints(ctx, "gokapi", func(p UpdateTierPoint) bool { return p.Tier != UpdateTierSecondDrive })
|
||||||
|
if !ok || p.Tier != UpdateTierLocal || fmt.Sprint(tiersOf(seen)) != "[2 1]" {
|
||||||
|
t.Errorf("tier 2 refused: want tier 1; got %+v seen=%v", p, tiersOf(seen))
|
||||||
|
}
|
||||||
|
p, ok, seen = m.UpdateRestorePoints(ctx, "gokapi", func(p UpdateTierPoint) bool { return p.Tier == UpdateTierOffsite })
|
||||||
|
if !ok || p.Tier != UpdateTierOffsite || !p.At.Equal(r475T0.Add(-3*time.Hour)) || fmt.Sprint(tiersOf(seen)) != "[2 1 3]" {
|
||||||
|
t.Errorf("tiers 2 and 1 refused: want tier 3; got %+v seen=%v", p, tiersOf(seen))
|
||||||
|
}
|
||||||
|
if _, ok, seen = m.UpdateRestorePoints(ctx, "gokapi", func(UpdateTierPoint) bool { return false }); ok || len(seen) != 3 {
|
||||||
|
t.Errorf("nothing accepted: want not found with all three seen; ok=%v seen=%v", ok, tiersOf(seen))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestR475_H_OwnUnitOnly(t *testing.T) {
|
||||||
|
m, _ := r475Manager()
|
||||||
|
m.updateTier2PointFn = noTier2
|
||||||
|
m.updateTier1PointsFn = tier1At(r475T0.Add(-2 * time.Hour))
|
||||||
|
m.updateOffsiteInvFn = noOffsiteTarget
|
||||||
|
p, ok, _ := m.UpdateRestorePoints(context.Background(), "gokapi", nil)
|
||||||
|
if !ok || p.Tier != UpdateTierLocal || !p.At.Equal(r475T0.Add(-2*time.Hour)) {
|
||||||
|
t.Fatalf("got %+v ok=%v", p, ok)
|
||||||
|
}
|
||||||
|
// A Tier-2 record that was only ATTEMPTED is not a copy (R-101) — the own unit still wins.
|
||||||
|
m.updateTier2PointFn = func(string) (Tier2RestorePoint, error) {
|
||||||
|
return Tier2RestorePoint{Restorable: true, CopyDateProven: false}, nil
|
||||||
|
}
|
||||||
|
if p, ok, _ = m.UpdateRestorePoints(context.Background(), "gokapi", nil); !ok || p.Tier != UpdateTierLocal {
|
||||||
|
t.Errorf("an unproven Tier-2 record must be skipped; got %+v", p)
|
||||||
|
}
|
||||||
|
// No unit on disk (ListRestorePoints' empty list) is no copy.
|
||||||
|
m.updateTier1PointsFn = noTier1
|
||||||
|
if _, ok, _ = m.UpdateRestorePoints(context.Background(), "gokapi", nil); ok {
|
||||||
|
t.Error("no copy on any tier must be not found")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestR475_I_OffsiteOnly(t *testing.T) {
|
||||||
|
m, _ := r475Manager()
|
||||||
|
m.updateTier2PointFn, m.updateTier1PointsFn = noTier2, noTier1
|
||||||
|
m.updateOffsiteInvFn = offsiteWith("gokapi", r475T0.Add(-5*time.Hour))
|
||||||
|
if p, ok, _ := m.UpdateRestorePoints(context.Background(), "gokapi", nil); !ok || p.Tier != UpdateTierOffsite {
|
||||||
|
t.Fatalf("got %+v ok=%v", p, ok)
|
||||||
|
}
|
||||||
|
// Another app's snapshot is not this app's copy.
|
||||||
|
if _, ok, _ := m.UpdateRestorePoints(context.Background(), "nextcloud", nil); ok {
|
||||||
|
t.Error("a snapshot tagged for a different app must not count")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// J — an unreachable off-site repository is ABSENT with a WARN, and it is bounded in time.
|
||||||
|
func TestR475_J_OffsiteUnreachableIsAbsentWithAWarn(t *testing.T) {
|
||||||
|
m, buf := r475Manager()
|
||||||
|
m.updateTier2PointFn, m.updateTier1PointsFn = noTier2, noTier1
|
||||||
|
m.updateOffsiteInvFn = func(context.Context) (OffsiteInventory, error) {
|
||||||
|
return OffsiteInventory{}, errors.New("ssh: connect to host: connection timed out")
|
||||||
|
}
|
||||||
|
if _, ok, _ := m.UpdateRestorePoints(context.Background(), "gokapi", nil); ok {
|
||||||
|
t.Error("an unreachable off-site copy must count as absent")
|
||||||
|
}
|
||||||
|
if !strings.Contains(buf.String(), "[WARN]") || !strings.Contains(buf.String(), "counted as ABSENT") {
|
||||||
|
t.Errorf("an unreachable off-site copy must WARN; log = %q", buf.String())
|
||||||
|
}
|
||||||
|
|
||||||
|
// Control: a box with NO off-site target is plainly absent — that is not a fault, so no WARN.
|
||||||
|
buf.Reset()
|
||||||
|
m.updateOffsiteInvFn = noOffsiteTarget
|
||||||
|
if _, ok, _ := m.UpdateRestorePoints(context.Background(), "gokapi", nil); ok || strings.Contains(buf.String(), "WARN") {
|
||||||
|
t.Errorf("no off-site target: want absent and silent; ok=%v log=%q", ok, buf.String())
|
||||||
|
}
|
||||||
|
|
||||||
|
// The bound: a repository that never answers is given up on at the timeout.
|
||||||
|
old := updateOffsiteCheckTimeout
|
||||||
|
updateOffsiteCheckTimeout = 50 * time.Millisecond
|
||||||
|
defer func() { updateOffsiteCheckTimeout = old }()
|
||||||
|
buf.Reset()
|
||||||
|
m.updateOffsiteInvFn = func(ctx context.Context) (OffsiteInventory, error) {
|
||||||
|
<-ctx.Done()
|
||||||
|
return OffsiteInventory{}, ctx.Err()
|
||||||
|
}
|
||||||
|
start := time.Now()
|
||||||
|
_, ok, _ := m.UpdateRestorePoints(context.Background(), "gokapi", nil)
|
||||||
|
if took := time.Since(start); ok || took > 5*time.Second || !strings.Contains(buf.String(), "counted as ABSENT") {
|
||||||
|
t.Errorf("a hanging off-site check must end at its bound as absent with a WARN; ok=%v took=%s log=%q", ok, took, buf.String())
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The hold names the tier. The labels are the ruling's exact words.
|
||||||
|
func TestR475_HoldTextNamesTheTier(t *testing.T) {
|
||||||
|
at := time.Date(2026, 9, 13, 8, 0, 0, 0, time.UTC)
|
||||||
|
copyAt := time.Date(2026, 9, 13, 1, 30, 0, 0, time.UTC)
|
||||||
|
for tier, label := range map[int]string{
|
||||||
|
UpdateTierSecondDrive: "második meghajtó",
|
||||||
|
UpdateTierLocal: "saját meghajtó",
|
||||||
|
UpdateTierOffsite: "távoli mentés",
|
||||||
|
} {
|
||||||
|
sett := slice4Settings(t)
|
||||||
|
m := &Manager{logger: log.New(io.Discard, "", 0), settings: sett}
|
||||||
|
if err := m.HoldAfterFailedUpdate("gokapi", at, copyAt, tier); err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
held, why := m.RestoreHoldFor("gokapi")
|
||||||
|
want := fmt.Sprintf(UpdateHoldFmt, "gokapi", "2026-09-13 10:00", label, "2026-09-13 03:30")
|
||||||
|
if !held || why != want {
|
||||||
|
t.Errorf("tier %d: hold text =\n%q\nwant\n%q", tier, why, want)
|
||||||
|
}
|
||||||
|
if !strings.HasSuffix(why, "ebből a biztonsági mentésből: "+label+", 2026-09-13 03:30.") {
|
||||||
|
t.Errorf("tier %d: the sentence must END naming the tier and the date, got %q", tier, why)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestR475_AHoldWrittenBeforeTheTierKeepsItsSentence(t *testing.T) {
|
||||||
|
sett := slice4Settings(t)
|
||||||
|
m := &Manager{logger: log.New(io.Discard, "", 0), settings: sett}
|
||||||
|
if err := sett.SetRestoreHold(settings.RestoreHold{Stack: "uptime-kuma", At: "2026-09-13T10:18:02Z",
|
||||||
|
Reason: settings.HoldReasonUpdateFailed, CopyDate: "2026-09-13T10:09:51Z"}); err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
_, why := m.RestoreHoldFor("uptime-kuma")
|
||||||
|
// The exact sentence quoted in audits/slice4-2026-09-13 (13-H-refusals.txt), live on v0.238.1.
|
||||||
|
if want := fmt.Sprintf(UpdateHoldLegacyFmt, "uptime-kuma", "2026-09-13 12:18", "2026-09-13 12:09"); why != want {
|
||||||
|
t.Errorf("a v0.238.1 hold must keep its sentence:\n%q\nwant\n%q", why, want)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// RunAppBackupNow's tail: a Tier-2 failure no longer fails "back up first", and the own unit it
|
||||||
|
// captured reads as fresh even when the capture found nothing to rewrite.
|
||||||
|
//
|
||||||
|
// COMPANION RED-PROOF (REPORT.md): aim the Chtimes at a path that does not exist — this test fails on
|
||||||
|
// the mtime: "back up first" would then leave a quiet app's own unit as old as its last definition change.
|
||||||
|
func TestR475_PreBackupTail_Tier2FailureIsAWarnAndTheOwnUnitIsFresh(t *testing.T) {
|
||||||
|
nsRoot := t.TempDir()
|
||||||
|
mp := RecoveryUnitManifestPath(nsRoot, "gokapi")
|
||||||
|
if err := os.MkdirAll(filepath.Dir(mp), 0o755); err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
if err := os.WriteFile(mp, []byte(`{"app_name":"gokapi"}`), 0o644); err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
old := time.Now().Add(-30 * time.Hour)
|
||||||
|
if err := os.Chtimes(mp, old, old); err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
m, buf := r475Manager()
|
||||||
|
tier2Ran := false
|
||||||
|
m.perAppTier2 = func(string) error { tier2Ran = true; return errors.New("no second drive with room") }
|
||||||
|
|
||||||
|
now := time.Now()
|
||||||
|
m.updatePreBackupTail("gokapi", nsRoot, now)
|
||||||
|
|
||||||
|
if !tier2Ran {
|
||||||
|
t.Fatal("the Tier-2 copy must still be attempted")
|
||||||
|
}
|
||||||
|
if !strings.Contains(buf.String(), "[WARN]") || !strings.Contains(buf.String(), "Tier 2 copy FAILED") {
|
||||||
|
t.Errorf("a Tier-2 failure must be logged as a WARN; log = %q", buf.String())
|
||||||
|
}
|
||||||
|
fi, err := os.Stat(mp)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
if d := fi.ModTime().Sub(now); d < -time.Second || d > time.Second {
|
||||||
|
t.Errorf("the own unit must read as proven NOW (mtime %s, now %s)", fi.ModTime(), now)
|
||||||
|
}
|
||||||
|
|
||||||
|
// Control: the same unit through ListRestorePoints' own rule reads fresh — the newest artifact.
|
||||||
|
buf.Reset()
|
||||||
|
m.perAppTier2 = func(string) error { return nil }
|
||||||
|
m.updatePreBackupTail("gokapi", nsRoot, now)
|
||||||
|
if strings.Contains(buf.String(), "WARN") {
|
||||||
|
t.Errorf("a successful Tier-2 copy must not WARN; log = %q", buf.String())
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The tail is really what RunAppBackupNow runs — a helper tested and never called is the "seam built
|
||||||
|
// but never wired" shape this project has shipped repeatedly.
|
||||||
|
func TestR475_RunAppBackupNowUsesTheTolerantTail(t *testing.T) {
|
||||||
|
fset := token.NewFileSet()
|
||||||
|
f, err := parser.ParseFile(fset, "update_guard.go", nil, 0)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
var calls []string
|
||||||
|
found := false
|
||||||
|
for _, d := range f.Decls {
|
||||||
|
fn, ok := d.(*ast.FuncDecl)
|
||||||
|
if !ok || fn.Name.Name != "RunAppBackupNow" {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
found = true
|
||||||
|
ast.Inspect(fn.Body, func(n ast.Node) bool {
|
||||||
|
if sel, ok := n.(*ast.SelectorExpr); ok {
|
||||||
|
calls = append(calls, sel.Sel.Name)
|
||||||
|
}
|
||||||
|
return true
|
||||||
|
})
|
||||||
|
}
|
||||||
|
joined := " " + strings.Join(calls, " ") + " "
|
||||||
|
if !found || !strings.Contains(joined, " updatePreBackupTail ") {
|
||||||
|
t.Fatalf("RunAppBackupNow must end in updatePreBackupTail; selectors = %s", joined)
|
||||||
|
}
|
||||||
|
if strings.Contains(joined, " RunTier2 ") || strings.Contains(joined, " perAppTier2 ") {
|
||||||
|
t.Error("RunAppBackupNow must not run the Tier-2 copy itself any more — only through the tolerant tail")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestR475_CanBackUpApp(t *testing.T) {
|
||||||
|
var nilM *Manager
|
||||||
|
if ok, why := nilM.CanBackUpApp("gokapi"); ok || why == "" {
|
||||||
|
t.Error("no backup manager: cannot back up, with a reason")
|
||||||
|
}
|
||||||
|
m, _ := r475Manager()
|
||||||
|
if ok, why := m.CanBackUpApp("gokapi"); ok || !strings.Contains(why, "stack provider") {
|
||||||
|
t.Errorf("no stack provider: cannot back up; got ok=%v why=%q", ok, why)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Tier 3's route back is the off-site restore, and it never went through RestoreFromRecoveryUnitAt —
|
||||||
|
// so it must lift an update hold itself, or a hold naming „távoli mentés" could never be cleared.
|
||||||
|
//
|
||||||
|
// COMPANION RED-PROOF (REPORT.md): remove the clearUpdateHoldAfterRestore call from
|
||||||
|
// ReconstituteFromOffsite — this test fails with the hold still in place.
|
||||||
|
func TestR475_OffsiteRestoreClearsAnUpdateHold(t *testing.T) {
|
||||||
|
m, prov, _ := reconFixture(t, "run1", "2026-07-19T06:00:00Z", "")
|
||||||
|
prov.composePath = writeLiveCompose(t, noDBCompose)
|
||||||
|
m.discoverDBs = func(context.Context) ([]DiscoveredDB, error) { return nil, nil }
|
||||||
|
if err := m.HoldAfterFailedUpdate("immich", r475T0, r475T0.Add(-time.Hour), UpdateTierOffsite); err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
if held, why := m.RestoreHoldFor("immich"); !held || !strings.Contains(why, "távoli mentés") {
|
||||||
|
t.Fatalf("fixture: the app must be held naming the off-site copy, got held=%v %q", held, why)
|
||||||
|
}
|
||||||
|
if _, err := m.ReconstituteFromOffsite(context.Background(), "immich", false); err != nil {
|
||||||
|
t.Fatalf("reconstitute: %v", err)
|
||||||
|
}
|
||||||
|
if held, _ := m.RestoreHoldFor("immich"); held {
|
||||||
|
t.Error("a successful off-site restore is the route back the hold names — the hold must be lifted")
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -1598,6 +1598,10 @@ type RestoreHold struct {
|
|||||||
// CopyDate is the RFC3339 time of the proven backup the customer is told they can restore from.
|
// CopyDate is the RFC3339 time of the proven backup the customer is told they can restore from.
|
||||||
// Only set for HoldReasonUpdateFailed.
|
// Only set for HoldReasonUpdateFailed.
|
||||||
CopyDate string `json:"copy_date,omitempty"`
|
CopyDate string `json:"copy_date,omitempty"`
|
||||||
|
// CopyTier (R-475, v0.239.0) is WHICH backup tier CopyDate belongs to: 1 the app's own recovery
|
||||||
|
// unit, 2 the second drive, 3 off-site. 0 means a hold written before the field existed; it keeps
|
||||||
|
// its original sentence (backup.UpdateHoldLegacyFmt). Only set for HoldReasonUpdateFailed.
|
||||||
|
CopyTier int `json:"copy_tier,omitempty"`
|
||||||
}
|
}
|
||||||
|
|
||||||
// Hold reasons. See RestoreHold.Reason.
|
// Hold reasons. See RestoreHold.Reason.
|
||||||
|
|||||||
@@ -6,6 +6,7 @@ import (
|
|||||||
"fmt"
|
"fmt"
|
||||||
"os"
|
"os"
|
||||||
"path/filepath"
|
"path/filepath"
|
||||||
|
"strings"
|
||||||
"time"
|
"time"
|
||||||
|
|
||||||
"gitea.dooplex.hu/admin/felhom-controller/internal/system"
|
"gitea.dooplex.hu/admin/felhom-controller/internal/system"
|
||||||
@@ -82,7 +83,7 @@ const (
|
|||||||
MsgUpdateAlreadyFmt = "A(z) %s frissítése már folyamatban van."
|
MsgUpdateAlreadyFmt = "A(z) %s frissítése már folyamatban van."
|
||||||
MsgUpdateBusy = "A frissítés most nem indítható: mentés/visszaállítás folyamatban. Próbáld újra, ha befejeződött."
|
MsgUpdateBusy = "A frissítés most nem indítható: mentés/visszaállítás folyamatban. Próbáld újra, ha befejeződött."
|
||||||
MsgUpdateMigrating = "A frissítés most nem indítható: adatáthelyezés folyamatban."
|
MsgUpdateMigrating = "A frissítés most nem indítható: adatáthelyezés folyamatban."
|
||||||
MsgUpdateNoBackupFmt = "A(z) %s nem frissíthető, mert nincs olyan biztonsági mentése, amelyből vissza lehetne állítani. Kapcsold be a 2. mentést az alkalmazás mentési beállításainál a Mentések oldalon, és várd meg az első sikeres másolatot — utána a frissítés elindítható."
|
MsgUpdateNoBackupFmt = "A(z) %s nem frissíthető, mert nincs olyan biztonsági mentése, amelyből vissza lehetne állítani, és most új mentés sem készíthető róla. Ellenőrizd a Mentések oldalon, hogy az alkalmazás meghajtója elérhető-e — utána a frissítés elindítható."
|
||||||
MsgUpdateDiskFmt = "Nincs elég szabad hely a frissítéshez: %.1f GB szabad, az új verzió letöltéséhez legalább %.0f GB szükséges."
|
MsgUpdateDiskFmt = "Nincs elég szabad hely a frissítéshez: %.1f GB szabad, az új verzió letöltéséhez legalább %.0f GB szükséges."
|
||||||
MsgUpdateBackupFailFmt = "A frissítés nem indult el, mert a frissítés előtti biztonsági mentés nem sikerült: %v. Az alkalmazás változatlanul fut tovább."
|
MsgUpdateBackupFailFmt = "A frissítés nem indult el, mert a frissítés előtti biztonsági mentés nem sikerült: %v. Az alkalmazás változatlanul fut tovább."
|
||||||
MsgUpdateBackupNoUnit = "A frissítés nem indult el: a frissítés előtti mentés lefutott, de nem jött létre friss, visszaállítható másolat. Az alkalmazás változatlanul fut tovább."
|
MsgUpdateBackupNoUnit = "A frissítés nem indult el: a frissítés előtti mentés lefutott, de nem jött létre friss, visszaállítható másolat. Az alkalmazás változatlanul fut tovább."
|
||||||
@@ -107,11 +108,55 @@ const updateSettleWindow = 60 * time.Second
|
|||||||
// updatePollEvery is how often the health wait re-reads the stack.
|
// updatePollEvery is how often the health wait re-reads the stack.
|
||||||
const updatePollEvery = 5 * time.Second
|
const updatePollEvery = 5 * time.Second
|
||||||
|
|
||||||
// UpdateRestorePoint is the precondition answer, reduced to what the update needs.
|
// Backup tiers (R-475), mirroring backup.UpdateTier* — stacks cannot import backup, so
|
||||||
|
// TestR475_TierConstantsAgree (cmd/controller) pins the two sets equal.
|
||||||
|
const (
|
||||||
|
UpdateTierLocal = 1 // the app's own recovery unit
|
||||||
|
UpdateTierSecondDrive = 2 // the Tier-2 mirror on another drive
|
||||||
|
UpdateTierOffsite = 3 // off-site
|
||||||
|
)
|
||||||
|
|
||||||
|
// UpdateRestorePoint is one proven, restorable copy the update may lean on. The backup side returns
|
||||||
|
// only proven, restorable copies (never an attempt clock, never an unopenable unit), so there is no
|
||||||
|
// "maybe" field here to forget to check.
|
||||||
type UpdateRestorePoint struct {
|
type UpdateRestorePoint struct {
|
||||||
Restorable bool // an openable recovery unit exists in the Tier-2 copy
|
Tier int // UpdateTierSecondDrive / UpdateTierLocal / UpdateTierOffsite
|
||||||
Proven bool // a copy actually succeeded (never an attempt clock)
|
ProvenAt time.Time // when the data in that copy was last proven written
|
||||||
ProvenAt time.Time // when the data in that copy was last proven copied
|
}
|
||||||
|
|
||||||
|
func updateTierName(tier int) string {
|
||||||
|
switch tier {
|
||||||
|
case UpdateTierSecondDrive:
|
||||||
|
return "Tier 2 (second drive)"
|
||||||
|
case UpdateTierLocal:
|
||||||
|
return "Tier 1 (own recovery unit)"
|
||||||
|
case UpdateTierOffsite:
|
||||||
|
return "Tier 3 (off-site)"
|
||||||
|
}
|
||||||
|
return fmt.Sprintf("tier %d", tier)
|
||||||
|
}
|
||||||
|
|
||||||
|
// freshRestorePoint is THE age rule (backup_max_age), applied to whichever tier is being considered
|
||||||
|
// — R-475 Scenario M: a stale copy on ANY tier is stale.
|
||||||
|
//
|
||||||
|
// COMPANION RED-PROOF M (REPORT.md): check the age only for Tier 2 (let any other tier through
|
||||||
|
// whatever its age). TestR475_M_TheAgeRuleAppliesToTheChosenTier then fails: a 30-hour-old copy of
|
||||||
|
// the app's own unit carries the update with no backup first.
|
||||||
|
func freshRestorePoint(now time.Time, maxAge time.Duration) func(UpdateRestorePoint) bool {
|
||||||
|
return func(p UpdateRestorePoint) bool {
|
||||||
|
return !p.ProvenAt.IsZero() && now.Sub(p.ProvenAt) <= maxAge
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func describeRestorePoints(now time.Time, pts []UpdateRestorePoint) string {
|
||||||
|
if len(pts) == 0 {
|
||||||
|
return "none"
|
||||||
|
}
|
||||||
|
parts := make([]string, 0, len(pts))
|
||||||
|
for _, p := range pts {
|
||||||
|
parts = append(parts, fmt.Sprintf("%s at %s (%s old)", updateTierName(p.Tier), p.ProvenAt.UTC().Format(time.RFC3339), now.Sub(p.ProvenAt).Round(time.Minute)))
|
||||||
|
}
|
||||||
|
return strings.Join(parts, "; ")
|
||||||
}
|
}
|
||||||
|
|
||||||
// UpdateGuards is everything the update needs from the backup side. The stacks package cannot import
|
// UpdateGuards is everything the update needs from the backup side. The stacks package cannot import
|
||||||
@@ -119,10 +164,14 @@ type UpdateRestorePoint struct {
|
|||||||
type UpdateGuards interface {
|
type UpdateGuards interface {
|
||||||
HoldFor(name string) (bool, string)
|
HoldFor(name string) (bool, string)
|
||||||
Busy(name string) (bool, string)
|
Busy(name string) (bool, string)
|
||||||
RestorePoint(name string) (UpdateRestorePoint, error)
|
// RestorePoints walks the tiers in preference order (2, 1, 3) and returns the first copy accept
|
||||||
|
// admits (nil = any), whether one was found, and every copy looked at (R-475).
|
||||||
|
RestorePoints(ctx context.Context, name string, accept func(UpdateRestorePoint) bool) (UpdateRestorePoint, bool, []UpdateRestorePoint)
|
||||||
|
// CanBackUp reports whether "back up first" can run for this app now (R-475 Scenario L).
|
||||||
|
CanBackUp(name string) (bool, string)
|
||||||
BackupNow(ctx context.Context, name string) error
|
BackupNow(ctx context.Context, name string) error
|
||||||
SafetyDump(ctx context.Context, name string) ([]string, error)
|
SafetyDump(ctx context.Context, name string) ([]string, error)
|
||||||
HoldAfterFailedUpdate(name string, at, provenCopyAt time.Time) error
|
HoldAfterFailedUpdate(name string, at time.Time, rp UpdateRestorePoint) error
|
||||||
}
|
}
|
||||||
|
|
||||||
// SetUpdateGuards wires the backup side. INIT-ONLY. Unwired, every update is refused (fail closed):
|
// SetUpdateGuards wires the backup side. INIT-ONLY. Unwired, every update is refused (fail closed):
|
||||||
@@ -199,10 +248,19 @@ func (m *Manager) UpdatePreflight(name string) *UpdateRefusal {
|
|||||||
if m.IsMigrating() {
|
if m.IsMigrating() {
|
||||||
return m.refuseUpdate(name, "migrating", MsgUpdateMigrating, "a data migration is running")
|
return m.refuseUpdate(name, "migrating", MsgUpdateMigrating, "a data migration is running")
|
||||||
}
|
}
|
||||||
rp, err := g.RestorePoint(name)
|
// R-475: any tier counts, and an app with no copy at all is backed up first by the job. So the only
|
||||||
if err != nil || !rp.Restorable || !rp.Proven {
|
// refusal left here is Scenario L — no copy on any tier AND no way to make one now. (With a copy
|
||||||
return m.refuseUpdate(name, "no_backup", fmt.Sprintf(MsgUpdateNoBackupFmt, name),
|
// but no way to back up, the job still applies the age rule and refuses then if the copy is stale.)
|
||||||
fmt.Sprintf("no restorable proven Tier-2 unit (restorable=%v proven=%v err=%v)", rp.Restorable, rp.Proven, err))
|
// Tier 3 is looked at only on this branch, so an ordinary update never waits on the network here.
|
||||||
|
//
|
||||||
|
// COMPANION RED-PROOF (REPORT.md): drop the `!found` condition. TestR475_L then fails — an app
|
||||||
|
// with a copy but no way to back up is refused too.
|
||||||
|
if canBackUp, why := g.CanBackUp(name); !canBackUp {
|
||||||
|
if _, found, seen := g.RestorePoints(context.Background(), name, nil); !found {
|
||||||
|
return m.refuseUpdate(name, "no_backup", fmt.Sprintf(MsgUpdateNoBackupFmt, name),
|
||||||
|
fmt.Sprintf("no copy on any tier (found: %s) and no backup can be taken now: %s", describeRestorePoints(m.now(), seen), why))
|
||||||
|
}
|
||||||
|
m.logger.Printf("[WARN] [stacks] update %s: no backup can be taken now (%s) — an existing copy must carry the update", name, why)
|
||||||
}
|
}
|
||||||
if ref := m.updateMemoryRefusal(name, st); ref != nil {
|
if ref := m.updateMemoryRefusal(name, st); ref != nil {
|
||||||
return ref
|
return ref
|
||||||
@@ -383,15 +441,15 @@ func (m *Manager) runGuardedUpdate(ctx context.Context, name string) {
|
|||||||
fail(MsgUpdateNoGuards, "no UpdateGuards wired")
|
fail(MsgUpdateNoGuards, "no UpdateGuards wired")
|
||||||
return
|
return
|
||||||
}
|
}
|
||||||
rp, err := g.RestorePoint(name)
|
// R-475: the precondition is a copy on ANY tier, chosen in the order 2, 1, 3, and the age rule
|
||||||
if err != nil || !rp.Restorable || !rp.Proven {
|
// applies to whichever tier is chosen. The first FRESH copy wins — not merely the first copy — so a
|
||||||
fail(fmt.Sprintf(MsgUpdateNoBackupFmt, name), fmt.Sprintf("precondition vanished: restorable=%v proven=%v err=%v", rp.Restorable, rp.Proven, err))
|
// stale second-drive mirror never forces a backup while the app's own unit is minutes old.
|
||||||
return
|
|
||||||
}
|
|
||||||
|
|
||||||
maxAge := m.backupMaxAge()
|
maxAge := m.backupMaxAge()
|
||||||
if age := start.Sub(rp.ProvenAt); age > maxAge {
|
rp, ok, seen := g.RestorePoints(ctx, name, freshRestorePoint(start, maxAge))
|
||||||
m.logger.Printf("[INFO] [stacks] update %s: the proven copy is %s old (limit %s) — backing up first", name, age.Round(time.Minute), maxAge)
|
if ok {
|
||||||
|
m.logger.Printf("[INFO] [stacks] update %s: precondition met — %s copy from %s (%s old, limit %s)", name, updateTierName(rp.Tier), rp.ProvenAt.UTC().Format(time.RFC3339), start.Sub(rp.ProvenAt).Round(time.Minute), maxAge)
|
||||||
|
} else {
|
||||||
|
m.logger.Printf("[INFO] [stacks] update %s: no copy younger than %s on any tier (found: %s) — backing up first", name, maxAge, describeRestorePoints(start, seen))
|
||||||
if !m.enterUpdatePhase(name, &entry, UpdatePhaseBackingUp) {
|
if !m.enterUpdatePhase(name, &entry, UpdatePhaseBackingUp) {
|
||||||
fail(MsgUpdateJournalFailed, "journal write failed")
|
fail(MsgUpdateJournalFailed, "journal write failed")
|
||||||
return
|
return
|
||||||
@@ -400,15 +458,16 @@ func (m *Manager) runGuardedUpdate(ctx context.Context, name string) {
|
|||||||
fail(fmt.Sprintf(MsgUpdateBackupFailFmt, err), "pre-update backup: "+err.Error())
|
fail(fmt.Sprintf(MsgUpdateBackupFailFmt, err), "pre-update backup: "+err.Error())
|
||||||
return
|
return
|
||||||
}
|
}
|
||||||
rp, err = g.RestorePoint(name)
|
now := m.now()
|
||||||
if err != nil || !rp.Restorable || !rp.Proven || m.now().Sub(rp.ProvenAt) > maxAge {
|
rp, ok, seen = g.RestorePoints(ctx, name, freshRestorePoint(now, maxAge))
|
||||||
fail(MsgUpdateBackupNoUnit, fmt.Sprintf("after the backup: restorable=%v proven=%v at=%s err=%v", rp.Restorable, rp.Proven, rp.ProvenAt.Format(time.RFC3339), err))
|
if !ok {
|
||||||
|
fail(MsgUpdateBackupNoUnit, fmt.Sprintf("after the backup there is still no copy younger than %s on any tier (found: %s)", maxAge, describeRestorePoints(now, seen)))
|
||||||
return
|
return
|
||||||
}
|
}
|
||||||
} else {
|
m.logger.Printf("[INFO] [stacks] update %s: precondition met after the backup — %s copy from %s", name, updateTierName(rp.Tier), rp.ProvenAt.UTC().Format(time.RFC3339))
|
||||||
m.logger.Printf("[INFO] [stacks] update %s: precondition met — proven copy from %s (%s old, limit %s)", name, rp.ProvenAt.UTC().Format(time.RFC3339), age.Round(time.Minute), maxAge)
|
|
||||||
}
|
}
|
||||||
entry.ProvenCopyAt = rp.ProvenAt.UTC().Format(time.RFC3339)
|
entry.ProvenCopyAt = rp.ProvenAt.UTC().Format(time.RFC3339)
|
||||||
|
entry.ProvenTier = rp.Tier
|
||||||
|
|
||||||
// SAFETY DUMP BEFORE THE PIN MOVES — "a minute ago", before any migration can have run.
|
// SAFETY DUMP BEFORE THE PIN MOVES — "a minute ago", before any migration can have run.
|
||||||
if !m.enterUpdatePhase(name, &entry, UpdatePhaseSafetyDump) {
|
if !m.enterUpdatePhase(name, &entry, UpdatePhaseSafetyDump) {
|
||||||
@@ -472,20 +531,20 @@ func (m *Manager) runGuardedUpdate(ctx context.Context, name string) {
|
|||||||
}
|
}
|
||||||
|
|
||||||
if !m.enterUpdatePhase(name, &entry, UpdatePhaseStarting) {
|
if !m.enterUpdatePhase(name, &entry, UpdatePhaseStarting) {
|
||||||
m.failAndHold(ctx, name, dir, env, rp.ProvenAt, "journal write failed before up")
|
m.failAndHold(ctx, name, dir, env, rp, "journal write failed before up")
|
||||||
return
|
return
|
||||||
}
|
}
|
||||||
if _, err := m.updateCompose(dir, env, "up", "-d", "--remove-orphans"); err != nil {
|
if _, err := m.updateCompose(dir, env, "up", "-d", "--remove-orphans"); err != nil {
|
||||||
// Containers may already have been recreated on the new image — something may have run.
|
// Containers may already have been recreated on the new image — something may have run.
|
||||||
m.failAndHold(ctx, name, dir, env, rp.ProvenAt, "compose up failed: "+err.Error())
|
m.failAndHold(ctx, name, dir, env, rp, "compose up failed: "+err.Error())
|
||||||
return
|
return
|
||||||
}
|
}
|
||||||
m.verifyAndConclude(ctx, name, dir, env, rp.ProvenAt, start, &entry)
|
m.verifyAndConclude(ctx, name, dir, env, rp, start, &entry)
|
||||||
}
|
}
|
||||||
|
|
||||||
// verifyAndConclude is the TRUTH half (R-443): success is declared only after the app's health is
|
// verifyAndConclude is the TRUTH half (R-443): success is declared only after the app's health is
|
||||||
// known, and a failure holds the app.
|
// known, and a failure holds the app.
|
||||||
func (m *Manager) verifyAndConclude(ctx context.Context, name, dir string, env []string, provenAt, start time.Time, entry *updateJournalEntry) {
|
func (m *Manager) verifyAndConclude(ctx context.Context, name, dir string, env []string, rp UpdateRestorePoint, start time.Time, entry *updateJournalEntry) {
|
||||||
if !m.enterUpdatePhase(name, entry, UpdatePhaseVerifying) {
|
if !m.enterUpdatePhase(name, entry, UpdatePhaseVerifying) {
|
||||||
m.logger.Printf("[ERROR] [stacks] update %s: could not journal the verifying phase — verifying anyway", name)
|
m.logger.Printf("[ERROR] [stacks] update %s: could not journal the verifying phase — verifying anyway", name)
|
||||||
}
|
}
|
||||||
@@ -493,7 +552,7 @@ func (m *Manager) verifyAndConclude(ctx context.Context, name, dir string, env [
|
|||||||
waitStart := m.now()
|
waitStart := m.now()
|
||||||
healthy, detail := m.updateHealth(ctx, name, timeout)
|
healthy, detail := m.updateHealth(ctx, name, timeout)
|
||||||
if !healthy {
|
if !healthy {
|
||||||
m.failAndHold(ctx, name, dir, env, provenAt, "not healthy: "+detail)
|
m.failAndHold(ctx, name, dir, env, rp, "not healthy: "+detail)
|
||||||
return
|
return
|
||||||
}
|
}
|
||||||
m.logger.Printf("[INFO] [stacks] update %s: healthy after %s (%s)", name, m.now().Sub(waitStart).Round(time.Second), detail)
|
m.logger.Printf("[INFO] [stacks] update %s: healthy after %s (%s)", name, m.now().Sub(waitStart).Round(time.Second), detail)
|
||||||
@@ -506,7 +565,7 @@ func (m *Manager) verifyAndConclude(ctx context.Context, name, dir string, env [
|
|||||||
}
|
}
|
||||||
|
|
||||||
// failAndHold is Scenario F: stop the app, record the hold, tell the customer the route back.
|
// failAndHold is Scenario F: stop the app, record the hold, tell the customer the route back.
|
||||||
func (m *Manager) failAndHold(ctx context.Context, name, dir string, env []string, provenAt time.Time, why string) {
|
func (m *Manager) failAndHold(ctx context.Context, name, dir string, env []string, rp UpdateRestorePoint, why string) {
|
||||||
m.logger.Printf("[ERROR] [stacks] update %s FAILED after the new version was started: %s — stopping and HOLDING the app; the pin stays on the new version (its migration may have run)", name, why)
|
m.logger.Printf("[ERROR] [stacks] update %s FAILED after the new version was started: %s — stopping and HOLDING the app; the pin stays on the new version (its migration may have run)", name, why)
|
||||||
if _, err := m.updateCompose(dir, env, "down"); err != nil {
|
if _, err := m.updateCompose(dir, env, "down"); err != nil {
|
||||||
m.logger.Printf("[ERROR] [stacks] update %s: stopping the failed app also failed: %v", name, err)
|
m.logger.Printf("[ERROR] [stacks] update %s: stopping the failed app also failed: %v", name, err)
|
||||||
@@ -514,7 +573,7 @@ func (m *Manager) failAndHold(ctx context.Context, name, dir string, env []strin
|
|||||||
msg := MsgUpdateHoldUnsaved
|
msg := MsgUpdateHoldUnsaved
|
||||||
if g := m.guards(); g == nil {
|
if g := m.guards(); g == nil {
|
||||||
m.logger.Printf("[ERROR] [stacks] update %s: no UpdateGuards — the hold CANNOT be recorded", name)
|
m.logger.Printf("[ERROR] [stacks] update %s: no UpdateGuards — the hold CANNOT be recorded", name)
|
||||||
} else if err := g.HoldAfterFailedUpdate(name, m.now(), provenAt); err != nil {
|
} else if err := g.HoldAfterFailedUpdate(name, m.now(), rp); err != nil {
|
||||||
m.logger.Printf("[ERROR] [stacks] update %s: %v", name, err)
|
m.logger.Printf("[ERROR] [stacks] update %s: %v", name, err)
|
||||||
} else if _, why := g.HoldFor(name); why != "" {
|
} else if _, why := g.HoldFor(name); why != "" {
|
||||||
msg = why
|
msg = why
|
||||||
@@ -619,6 +678,9 @@ type updateJournalEntry struct {
|
|||||||
PrevCompose string `json:"prev_compose,omitempty"`
|
PrevCompose string `json:"prev_compose,omitempty"`
|
||||||
PrevApplied string `json:"prev_applied,omitempty"`
|
PrevApplied string `json:"prev_applied,omitempty"`
|
||||||
ProvenCopyAt string `json:"proven_copy_at,omitempty"`
|
ProvenCopyAt string `json:"proven_copy_at,omitempty"`
|
||||||
|
// ProvenTier (R-475) — which tier ProvenCopyAt belongs to, so a resumed update that fails names
|
||||||
|
// the right copy. 0 in a journal written by v0.238.1 or older.
|
||||||
|
ProvenTier int `json:"proven_tier,omitempty"`
|
||||||
}
|
}
|
||||||
|
|
||||||
type updateJournal struct {
|
type updateJournal struct {
|
||||||
@@ -787,16 +849,17 @@ func (m *Manager) ResumeInterruptedUpdates(ctx context.Context) int {
|
|||||||
continue
|
continue
|
||||||
}
|
}
|
||||||
provenAt, _ := time.Parse(time.RFC3339, e.ProvenCopyAt)
|
provenAt, _ := time.Parse(time.RFC3339, e.ProvenCopyAt)
|
||||||
|
rp := UpdateRestorePoint{Tier: e.ProvenTier, ProvenAt: provenAt}
|
||||||
dir := filepath.Dir(st.ComposePath)
|
dir := filepath.Dir(st.ComposePath)
|
||||||
go func(name, dir string, e updateJournalEntry, provenAt time.Time) {
|
go func(name, dir string, e updateJournalEntry, rp UpdateRestorePoint) {
|
||||||
env := m.stackEnv(dir)
|
env := m.stackEnv(dir)
|
||||||
m.logger.Printf("[INFO] [stacks] update %s: resuming after a controller restart — `up -d` then the health wait", name)
|
m.logger.Printf("[INFO] [stacks] update %s: resuming after a controller restart — `up -d` then the health wait", name)
|
||||||
if _, err := m.updateCompose(dir, env, "up", "-d", "--remove-orphans"); err != nil {
|
if _, err := m.updateCompose(dir, env, "up", "-d", "--remove-orphans"); err != nil {
|
||||||
m.failAndHold(ctx, name, dir, env, provenAt, "resumed compose up failed: "+err.Error())
|
m.failAndHold(ctx, name, dir, env, rp, "resumed compose up failed: "+err.Error())
|
||||||
return
|
return
|
||||||
}
|
}
|
||||||
m.verifyAndConclude(ctx, name, dir, env, provenAt, e.StartedAt, &e)
|
m.verifyAndConclude(ctx, name, dir, env, rp, e.StartedAt, &e)
|
||||||
}(name, dir, e, provenAt)
|
}(name, dir, e, rp)
|
||||||
}
|
}
|
||||||
return len(names)
|
return len(names)
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -25,13 +25,15 @@ type fakeGuards struct {
|
|||||||
held bool
|
held bool
|
||||||
holdWhy string
|
holdWhy string
|
||||||
busy bool
|
busy bool
|
||||||
rp UpdateRestorePoint
|
// points are the copies the backup side holds, in tier order (R-475); pointsAfterBackup replaces
|
||||||
rpErr error
|
// them when BackupNow succeeds (nil = the backup changed nothing).
|
||||||
rpAfterBackup *UpdateRestorePoint
|
points []UpdateRestorePoint
|
||||||
|
pointsAfterBackup []UpdateRestorePoint
|
||||||
|
cannotBackUp bool
|
||||||
backupErr error
|
backupErr error
|
||||||
dumpErr error
|
dumpErr error
|
||||||
holdErr error
|
holdErr error
|
||||||
holdProvenAt time.Time
|
holdRP UpdateRestorePoint
|
||||||
pinAtDump string
|
pinAtDump string
|
||||||
stackDir string
|
stackDir string
|
||||||
}
|
}
|
||||||
@@ -48,18 +50,29 @@ func (f *fakeGuards) HoldFor(string) (bool, string) {
|
|||||||
return f.held, f.holdWhy
|
return f.held, f.holdWhy
|
||||||
}
|
}
|
||||||
func (f *fakeGuards) Busy(string) (bool, string) { return f.busy, "fake busy" }
|
func (f *fakeGuards) Busy(string) (bool, string) { return f.busy, "fake busy" }
|
||||||
func (f *fakeGuards) RestorePoint(string) (UpdateRestorePoint, error) {
|
func (f *fakeGuards) RestorePoints(_ context.Context, _ string, accept func(UpdateRestorePoint) bool) (UpdateRestorePoint, bool, []UpdateRestorePoint) {
|
||||||
f.note("RestorePoint")
|
f.note("RestorePoints")
|
||||||
f.mu.Lock()
|
f.mu.Lock()
|
||||||
defer f.mu.Unlock()
|
defer f.mu.Unlock()
|
||||||
return f.rp, f.rpErr
|
var seen []UpdateRestorePoint
|
||||||
|
for _, p := range f.points {
|
||||||
|
seen = append(seen, p)
|
||||||
|
if accept == nil || accept(p) {
|
||||||
|
return p, true, seen
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return UpdateRestorePoint{}, false, seen
|
||||||
|
}
|
||||||
|
func (f *fakeGuards) CanBackUp(string) (bool, string) {
|
||||||
|
f.note("CanBackUp")
|
||||||
|
return !f.cannotBackUp, "fake: the drive is gone"
|
||||||
}
|
}
|
||||||
func (f *fakeGuards) BackupNow(context.Context, string) error {
|
func (f *fakeGuards) BackupNow(context.Context, string) error {
|
||||||
f.note("BackupNow")
|
f.note("BackupNow")
|
||||||
f.mu.Lock()
|
f.mu.Lock()
|
||||||
defer f.mu.Unlock()
|
defer f.mu.Unlock()
|
||||||
if f.backupErr == nil && f.rpAfterBackup != nil {
|
if f.backupErr == nil && f.pointsAfterBackup != nil {
|
||||||
f.rp = *f.rpAfterBackup
|
f.points = f.pointsAfterBackup
|
||||||
}
|
}
|
||||||
return f.backupErr
|
return f.backupErr
|
||||||
}
|
}
|
||||||
@@ -72,14 +85,14 @@ func (f *fakeGuards) SafetyDump(context.Context, string) ([]string, error) {
|
|||||||
}
|
}
|
||||||
return []string{"/fake/pre-restore-x.sql"}, f.dumpErr
|
return []string{"/fake/pre-restore-x.sql"}, f.dumpErr
|
||||||
}
|
}
|
||||||
func (f *fakeGuards) HoldAfterFailedUpdate(_ string, _ time.Time, provenAt time.Time) error {
|
func (f *fakeGuards) HoldAfterFailedUpdate(_ string, _ time.Time, rp UpdateRestorePoint) error {
|
||||||
f.note("HoldAfterFailedUpdate")
|
f.note("HoldAfterFailedUpdate")
|
||||||
f.mu.Lock()
|
f.mu.Lock()
|
||||||
defer f.mu.Unlock()
|
defer f.mu.Unlock()
|
||||||
if f.holdErr != nil {
|
if f.holdErr != nil {
|
||||||
return f.holdErr
|
return f.holdErr
|
||||||
}
|
}
|
||||||
f.held, f.holdWhy, f.holdProvenAt = true, "HELD-SENTENCE", provenAt
|
f.held, f.holdWhy, f.holdRP = true, "HELD-SENTENCE", rp
|
||||||
return nil
|
return nil
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -111,7 +124,7 @@ func newSlice4Manager(t *testing.T) (*Manager, string, *fakeGuards, *composeRec)
|
|||||||
m, dir := newPinManager(t, pinTplOld, pinTplNew,
|
m, dir := newPinManager(t, pinTplOld, pinTplNew,
|
||||||
"deployed: true\nenv: {}\npinned_images:\n web: nextcloud:31.0.14-apache\n")
|
"deployed: true\nenv: {}\npinned_images:\n web: nextcloud:31.0.14-apache\n")
|
||||||
mustWrite(t, AppliedComposePath(dir), pinTplOld)
|
mustWrite(t, AppliedComposePath(dir), pinTplOld)
|
||||||
g := &fakeGuards{rp: UpdateRestorePoint{Restorable: true, Proven: true, ProvenAt: slice4T0.Add(-1 * time.Hour)}, stackDir: dir}
|
g := &fakeGuards{points: []UpdateRestorePoint{{Tier: UpdateTierSecondDrive, ProvenAt: slice4T0.Add(-1 * time.Hour)}}, stackDir: dir}
|
||||||
c := &composeRec{fail: map[string]error{}}
|
c := &composeRec{fail: map[string]error{}}
|
||||||
m.updateGuards = g
|
m.updateGuards = g
|
||||||
m.updateComposeFn = c.fn
|
m.updateComposeFn = c.fn
|
||||||
@@ -212,9 +225,8 @@ func TestSlice4_A_SuccessIsDeclaredOnlyAfterHealth(t *testing.T) {
|
|||||||
if got, want := strings.Join(c.list(), " | "), "pull | up -d --remove-orphans"; got != want {
|
if got, want := strings.Join(c.list(), " | "), "pull | up -d --remove-orphans"; got != want {
|
||||||
t.Errorf("compose calls = %q, want %q", got, want)
|
t.Errorf("compose calls = %q, want %q", got, want)
|
||||||
}
|
}
|
||||||
// RestorePoint twice by design: once in the preflight (the refusal), once inside the job (the
|
// R-475: the preflight asks only whether a backup could be taken; the job reads the copies once.
|
||||||
// precondition must still hold when the job actually starts).
|
if got := strings.Join(g.callList(), ","); got != "CanBackUp,RestorePoints,SafetyDump" {
|
||||||
if got := strings.Join(g.callList(), ","); got != "RestorePoint,RestorePoint,SafetyDump" {
|
|
||||||
t.Errorf("a fresh copy needs no backup-first; guard calls = %s", got)
|
t.Errorf("a fresh copy needs no backup-first; guard calls = %s", got)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -223,8 +235,8 @@ func TestSlice4_A_SuccessIsDeclaredOnlyAfterHealth(t *testing.T) {
|
|||||||
|
|
||||||
func TestSlice4_B_StaleCopyIsRefreshedFirst(t *testing.T) {
|
func TestSlice4_B_StaleCopyIsRefreshedFirst(t *testing.T) {
|
||||||
m, dir, g, _ := newSlice4Manager(t)
|
m, dir, g, _ := newSlice4Manager(t)
|
||||||
g.rp.ProvenAt = slice4T0.Add(-30 * time.Hour) // > 24 h default
|
g.points[0].ProvenAt = slice4T0.Add(-30 * time.Hour) // > 24 h default
|
||||||
g.rpAfterBackup = &UpdateRestorePoint{Restorable: true, Proven: true, ProvenAt: slice4T0.Add(-1 * time.Minute)}
|
g.pointsAfterBackup = []UpdateRestorePoint{{Tier: UpdateTierSecondDrive, ProvenAt: slice4T0.Add(-1 * time.Minute)}}
|
||||||
if err := m.StartGuardedUpdate("nextcloud"); err != nil {
|
if err := m.StartGuardedUpdate("nextcloud"); err != nil {
|
||||||
t.Fatal(err)
|
t.Fatal(err)
|
||||||
}
|
}
|
||||||
@@ -233,7 +245,7 @@ func TestSlice4_B_StaleCopyIsRefreshedFirst(t *testing.T) {
|
|||||||
t.Fatalf("with a successful backup-first the update completes, got phase=%q err=%q", st.UpdatePhase, st.UpdateError)
|
t.Fatalf("with a successful backup-first the update completes, got phase=%q err=%q", st.UpdatePhase, st.UpdateError)
|
||||||
}
|
}
|
||||||
calls := strings.Join(g.callList(), ",")
|
calls := strings.Join(g.callList(), ",")
|
||||||
if !strings.HasPrefix(calls, "RestorePoint,RestorePoint,BackupNow,RestorePoint,SafetyDump") {
|
if !strings.HasPrefix(calls, "CanBackUp,RestorePoints,BackupNow,RestorePoints,SafetyDump") {
|
||||||
t.Errorf("a stale copy must be backed up FIRST and the precondition re-read; calls = %s", calls)
|
t.Errorf("a stale copy must be backed up FIRST and the precondition re-read; calls = %s", calls)
|
||||||
}
|
}
|
||||||
if got := pinOf(t, dir); got != "nextcloud:34.0.1-apache" {
|
if got := pinOf(t, dir); got != "nextcloud:34.0.1-apache" {
|
||||||
@@ -243,7 +255,7 @@ func TestSlice4_B_StaleCopyIsRefreshedFirst(t *testing.T) {
|
|||||||
|
|
||||||
func TestSlice4_B_BackupFailureMovesNothing(t *testing.T) {
|
func TestSlice4_B_BackupFailureMovesNothing(t *testing.T) {
|
||||||
m, dir, g, c := newSlice4Manager(t)
|
m, dir, g, c := newSlice4Manager(t)
|
||||||
g.rp.ProvenAt = slice4T0.Add(-30 * time.Hour)
|
g.points[0].ProvenAt = slice4T0.Add(-30 * time.Hour)
|
||||||
g.backupErr = errors.New("disk full")
|
g.backupErr = errors.New("disk full")
|
||||||
if err := m.StartGuardedUpdate("nextcloud"); err != nil {
|
if err := m.StartGuardedUpdate("nextcloud"); err != nil {
|
||||||
t.Fatal(err)
|
t.Fatal(err)
|
||||||
@@ -265,7 +277,7 @@ func TestSlice4_B_BackupFailureMovesNothing(t *testing.T) {
|
|||||||
|
|
||||||
func TestSlice4_B_BackupThatYieldsNoFreshUnitRefuses(t *testing.T) {
|
func TestSlice4_B_BackupThatYieldsNoFreshUnitRefuses(t *testing.T) {
|
||||||
m, dir, g, c := newSlice4Manager(t)
|
m, dir, g, c := newSlice4Manager(t)
|
||||||
g.rp.ProvenAt = slice4T0.Add(-30 * time.Hour) // stays stale: rpAfterBackup nil
|
g.points[0].ProvenAt = slice4T0.Add(-30 * time.Hour) // stays stale: pointsAfterBackup nil
|
||||||
if err := m.StartGuardedUpdate("nextcloud"); err != nil {
|
if err := m.StartGuardedUpdate("nextcloud"); err != nil {
|
||||||
t.Fatal(err)
|
t.Fatal(err)
|
||||||
}
|
}
|
||||||
@@ -278,21 +290,18 @@ func TestSlice4_B_BackupThatYieldsNoFreshUnitRefuses(t *testing.T) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// ── C: no backup exists that could restore this app ───────────────────────────────────────────────
|
// ── C / R-475 L: no copy on any tier, and no way to make one ─────────────────────────────────────
|
||||||
|
|
||||||
// COMPANION RED-PROOF 2 (REPORT.md): make the precondition in UpdatePreflight proceed when
|
// Until v0.239.0 this refused any app without a restorable Tier-2 unit. R-475: every tier counts and
|
||||||
// !rp.Restorable. This test then fails with the update started.
|
// an app with nothing is backed up first, so the refusal is now only "nothing anywhere AND no backup
|
||||||
func TestSlice4_C_NoRestorableCopyRefusesBeforeAnythingMoves(t *testing.T) {
|
// can be taken". The K half (nothing, but a backup CAN be taken) is TestR475_K in update_tiers_test.go.
|
||||||
|
func TestSlice4_C_NoCopyAndNoWayToBackUpRefusesBeforeAnythingMoves(t *testing.T) {
|
||||||
m, dir, g, c := newSlice4Manager(t)
|
m, dir, g, c := newSlice4Manager(t)
|
||||||
// A PROVEN, FRESH copy whose unit cannot be opened — the realistic half-copied mirror. Proven and
|
g.points, g.cannotBackUp = nil, true
|
||||||
// fresh on purpose: a fixture that is also unproven would be refused by the proven check alone,
|
|
||||||
// and a red-proof that drops the restorable check would then pass inertly (observed on the first
|
|
||||||
// run of red-proof 2, 2026-09-13).
|
|
||||||
g.rp = UpdateRestorePoint{Restorable: false, Proven: true, ProvenAt: slice4T0.Add(-time.Hour)}
|
|
||||||
err := m.StartGuardedUpdate("nextcloud")
|
err := m.StartGuardedUpdate("nextcloud")
|
||||||
var ref *UpdateRefusal
|
var ref *UpdateRefusal
|
||||||
if !errors.As(err, &ref) || ref.Reason != "no_backup" {
|
if !errors.As(err, &ref) || ref.Reason != "no_backup" {
|
||||||
t.Fatalf("an app with no restorable copy must be REFUSED (no_backup), got %v", err)
|
t.Fatalf("no copy anywhere and no way to back up must be REFUSED (no_backup), got %v", err)
|
||||||
}
|
}
|
||||||
if want := fmt.Sprintf(MsgUpdateNoBackupFmt, "nextcloud"); ref.Message != want {
|
if want := fmt.Sprintf(MsgUpdateNoBackupFmt, "nextcloud"); ref.Message != want {
|
||||||
t.Errorf("message = %q", ref.Message)
|
t.Errorf("message = %q", ref.Message)
|
||||||
@@ -304,10 +313,10 @@ func TestSlice4_C_NoRestorableCopyRefusesBeforeAnythingMoves(t *testing.T) {
|
|||||||
if pinOf(t, dir) != "nextcloud:31.0.14-apache" || len(c.list()) != 0 {
|
if pinOf(t, dir) != "nextcloud:31.0.14-apache" || len(c.list()) != 0 {
|
||||||
t.Error("a refused update must move nothing")
|
t.Error("a refused update must move nothing")
|
||||||
}
|
}
|
||||||
// A copy that exists but was never PROVEN is not a copy (R-101).
|
for _, call := range g.callList() {
|
||||||
g.rp = UpdateRestorePoint{Restorable: true, Proven: false}
|
if call == "BackupNow" {
|
||||||
if ref := m.UpdatePreflight("nextcloud"); ref == nil || ref.Reason != "no_backup" {
|
t.Error("a refused update must not try to back up")
|
||||||
t.Errorf("an unproven copy must refuse too, got %v", ref)
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -412,8 +421,8 @@ func TestSlice4_F_HealthFailureHoldsTheAppAndKeepsTheNewPin(t *testing.T) {
|
|||||||
if !held {
|
if !held {
|
||||||
t.Fatal("an app that did not come up must be HELD")
|
t.Fatal("an app that did not come up must be HELD")
|
||||||
}
|
}
|
||||||
if !g.holdProvenAt.Equal(g.rp.ProvenAt) {
|
if !g.holdRP.ProvenAt.Equal(g.points[0].ProvenAt) || g.holdRP.Tier != UpdateTierSecondDrive {
|
||||||
t.Errorf("the hold must name the PROVEN copy date %s, got %s", g.rp.ProvenAt, g.holdProvenAt)
|
t.Errorf("the hold must name the PROVEN copy it leans on (%+v), got %+v", g.points[0], g.holdRP)
|
||||||
}
|
}
|
||||||
if st.UpdateError != "HELD-SENTENCE" {
|
if st.UpdateError != "HELD-SENTENCE" {
|
||||||
t.Errorf("the page must carry the hold's own sentence, got %q", st.UpdateError)
|
t.Errorf("the page must carry the hold's own sentence, got %q", st.UpdateError)
|
||||||
@@ -466,6 +475,7 @@ func simulateAdvanced(t *testing.T, m *Manager, dir string) updateJournalEntry {
|
|||||||
StartedAt: slice4T0, PrevPin: map[string]string{"web": "nextcloud:31.0.14-apache"},
|
StartedAt: slice4T0, PrevPin: map[string]string{"web": "nextcloud:31.0.14-apache"},
|
||||||
PrevCompose: filepath.Join(dir, preUpdateComposeFile), PrevApplied: filepath.Join(dir, preUpdateAppliedFile),
|
PrevCompose: filepath.Join(dir, preUpdateComposeFile), PrevApplied: filepath.Join(dir, preUpdateAppliedFile),
|
||||||
ProvenCopyAt: slice4T0.Add(-time.Hour).Format(time.RFC3339),
|
ProvenCopyAt: slice4T0.Add(-time.Hour).Format(time.RFC3339),
|
||||||
|
ProvenTier: UpdateTierLocal,
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -530,8 +540,8 @@ func TestSlice4_G_InterruptedAfterUpResumesTheHealthWait(t *testing.T) {
|
|||||||
if got := strings.Join(c.list(), " | "); got != "up -d --remove-orphans | down" {
|
if got := strings.Join(c.list(), " | "); got != "up -d --remove-orphans | down" {
|
||||||
t.Errorf("resumption re-runs `up` then stops the failed app; compose calls = %q", got)
|
t.Errorf("resumption re-runs `up` then stops the failed app; compose calls = %q", got)
|
||||||
}
|
}
|
||||||
if !g.holdProvenAt.Equal(slice4T0.Add(-time.Hour)) {
|
if !g.holdRP.ProvenAt.Equal(slice4T0.Add(-time.Hour)) || g.holdRP.Tier != UpdateTierLocal {
|
||||||
t.Errorf("the resumed hold must name the journaled proven copy date, got %s", g.holdProvenAt)
|
t.Errorf("the resumed hold must name the journaled copy (tier %d at %s), got %+v", UpdateTierLocal, slice4T0.Add(-time.Hour), g.holdRP)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,175 @@
|
|||||||
|
package stacks
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"errors"
|
||||||
|
"fmt"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
)
|
||||||
|
|
||||||
|
// R-475 — any backup tier lets an app update (operator ruling 2026-09-13, controller v0.239.0).
|
||||||
|
// Tier order and the Tier-3 timeout are the backup side's (internal/backup/update_tiers_test.go);
|
||||||
|
// here is the update job's half: which copy it leans on, when it backs up first, when it refuses.
|
||||||
|
// Every age here is read against slice4T0, the same clock the job reads (R-457).
|
||||||
|
|
||||||
|
func hasGuardCall(g *fakeGuards, name string) bool {
|
||||||
|
for _, c := range g.callList() {
|
||||||
|
if c == name {
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
|
||||||
|
// failedUpdateHold runs an update whose new version never becomes healthy, so the hold records the
|
||||||
|
// copy the job chose. That is the observable consequence of the choice: the customer is told to
|
||||||
|
// restore from exactly that copy.
|
||||||
|
func failedUpdateHold(t *testing.T, points, afterBackup []UpdateRestorePoint) (*fakeGuards, *Stack) {
|
||||||
|
t.Helper()
|
||||||
|
m, _, g, _ := newSlice4Manager(t)
|
||||||
|
g.points, g.pointsAfterBackup = points, afterBackup
|
||||||
|
m.updateHealthFn = func(context.Context, string, time.Duration) (bool, string) { return false, "crash loop" }
|
||||||
|
if err := m.StartGuardedUpdate("nextcloud"); err != nil {
|
||||||
|
t.Fatalf("start: %v", err)
|
||||||
|
}
|
||||||
|
return g, waitUpdateDone(t, m, "nextcloud")
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestR475_G_TheSecondDriveIsChosenWhenPresent(t *testing.T) {
|
||||||
|
fresh := slice4T0.Add(-time.Hour)
|
||||||
|
g, st := failedUpdateHold(t, []UpdateRestorePoint{
|
||||||
|
{Tier: UpdateTierSecondDrive, ProvenAt: fresh},
|
||||||
|
{Tier: UpdateTierLocal, ProvenAt: fresh.Add(30 * time.Minute)},
|
||||||
|
{Tier: UpdateTierOffsite, ProvenAt: fresh},
|
||||||
|
}, nil)
|
||||||
|
if st.UpdatePhase != UpdatePhaseFailed || g.holdRP.Tier != UpdateTierSecondDrive || !g.holdRP.ProvenAt.Equal(fresh) {
|
||||||
|
t.Errorf("with a fresh second-drive copy that copy is chosen; hold = %+v phase=%q", g.holdRP, st.UpdatePhase)
|
||||||
|
}
|
||||||
|
if hasGuardCall(g, "BackupNow") {
|
||||||
|
t.Error("a fresh copy needs no backup first")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestR475_H_TheOwnUnitAloneCarriesTheUpdate(t *testing.T) {
|
||||||
|
fresh := slice4T0.Add(-2 * time.Hour)
|
||||||
|
g, _ := failedUpdateHold(t, []UpdateRestorePoint{{Tier: UpdateTierLocal, ProvenAt: fresh}}, nil)
|
||||||
|
if g.holdRP.Tier != UpdateTierLocal || !g.holdRP.ProvenAt.Equal(fresh) {
|
||||||
|
t.Errorf("an app with only its own recovery unit must update against it; hold = %+v", g.holdRP)
|
||||||
|
}
|
||||||
|
if hasGuardCall(g, "BackupNow") {
|
||||||
|
t.Error("a fresh own unit needs no backup first — Tier 2 must not be required anywhere")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestR475_I_TheOffsiteCopyAloneCarriesTheUpdate(t *testing.T) {
|
||||||
|
fresh := slice4T0.Add(-3 * time.Hour)
|
||||||
|
g, _ := failedUpdateHold(t, []UpdateRestorePoint{{Tier: UpdateTierOffsite, ProvenAt: fresh}}, nil)
|
||||||
|
if g.holdRP.Tier != UpdateTierOffsite || !g.holdRP.ProvenAt.Equal(fresh) || hasGuardCall(g, "BackupNow") {
|
||||||
|
t.Errorf("an app with only a fresh off-site copy must update against it; hold = %+v calls=%v", g.holdRP, g.callList())
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestR475_K_NoCopyAnywhereIsBackedUpFirstAndTheUpdateProceeds(t *testing.T) {
|
||||||
|
m, dir, g, _ := newSlice4Manager(t)
|
||||||
|
g.points = nil
|
||||||
|
g.pointsAfterBackup = []UpdateRestorePoint{{Tier: UpdateTierLocal, ProvenAt: slice4T0.Add(-time.Minute)}}
|
||||||
|
if ref := m.UpdatePreflight("nextcloud"); ref != nil {
|
||||||
|
t.Fatalf("an app with no copy but a working backup must NOT be refused, got %q (%s)", ref.Message, ref.Reason)
|
||||||
|
}
|
||||||
|
if err := m.StartGuardedUpdate("nextcloud"); err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
st := waitUpdateDone(t, m, "nextcloud")
|
||||||
|
if st.UpdatePhase != UpdatePhaseDone {
|
||||||
|
t.Fatalf("the backup made a Tier-1 copy, so the update completes; phase=%q err=%q", st.UpdatePhase, st.UpdateError)
|
||||||
|
}
|
||||||
|
if !hasGuardCall(g, "BackupNow") {
|
||||||
|
t.Error("with no copy anywhere the job must back up FIRST")
|
||||||
|
}
|
||||||
|
if got := pinOf(t, dir); got != "nextcloud:34.0.1-apache" {
|
||||||
|
t.Errorf("pin = %q", got)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestR475_K_NoCopyAndTheBackupFailsMovesNothing(t *testing.T) {
|
||||||
|
m, dir, g, c := newSlice4Manager(t)
|
||||||
|
g.points = nil
|
||||||
|
g.backupErr = errors.New("nincs elég szabad hely")
|
||||||
|
if err := m.StartGuardedUpdate("nextcloud"); err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
st := waitUpdateDone(t, m, "nextcloud")
|
||||||
|
if want := fmt.Sprintf(MsgUpdateBackupFailFmt, g.backupErr); st.UpdateError != want {
|
||||||
|
t.Errorf("UpdateError = %q, want %q", st.UpdateError, want)
|
||||||
|
}
|
||||||
|
if pinOf(t, dir) != "nextcloud:31.0.14-apache" || len(c.list()) != 0 {
|
||||||
|
t.Error("nothing may move when the backup-first failed")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// L — nothing anywhere and no way to back up is refused (TestSlice4_C_… pins the refusal itself). The
|
||||||
|
// other half pins that the refusal is not wider than that: a copy that EXISTS still carries it.
|
||||||
|
func TestR475_L_AnExistingCopyStillCarriesAnAppThatCannotBeBackedUp(t *testing.T) {
|
||||||
|
m, _, g, _ := newSlice4Manager(t)
|
||||||
|
g.cannotBackUp = true
|
||||||
|
g.points = []UpdateRestorePoint{{Tier: UpdateTierOffsite, ProvenAt: slice4T0.Add(-time.Hour)}}
|
||||||
|
if ref := m.UpdatePreflight("nextcloud"); ref != nil {
|
||||||
|
t.Fatalf("a fresh off-site copy must carry the update even when no backup can be taken now, got %q", ref.Message)
|
||||||
|
}
|
||||||
|
g.points = nil
|
||||||
|
if ref := m.UpdatePreflight("nextcloud"); ref == nil || ref.Reason != "no_backup" {
|
||||||
|
t.Fatalf("positive control: with no copy the same app must be refused, got %v", ref)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// M — a stale copy on ANY tier is stale. Each case would pass under a Tier-2-only age check except the
|
||||||
|
// first, which is the control proving the fixture's clock and limit are real.
|
||||||
|
func TestR475_M_TheAgeRuleAppliesToTheChosenTier(t *testing.T) {
|
||||||
|
stale, fresh, justNow := slice4T0.Add(-30*time.Hour), slice4T0.Add(-time.Hour), slice4T0.Add(-time.Minute)
|
||||||
|
cases := []struct {
|
||||||
|
name string
|
||||||
|
points []UpdateRestorePoint
|
||||||
|
afterBackup []UpdateRestorePoint
|
||||||
|
wantBackup bool
|
||||||
|
wantTier int
|
||||||
|
}{
|
||||||
|
{"control: a stale second-drive copy is backed up first", []UpdateRestorePoint{{Tier: UpdateTierSecondDrive, ProvenAt: stale}},
|
||||||
|
[]UpdateRestorePoint{{Tier: UpdateTierSecondDrive, ProvenAt: justNow}}, true, UpdateTierSecondDrive},
|
||||||
|
{"a stale own unit is backed up first", []UpdateRestorePoint{{Tier: UpdateTierLocal, ProvenAt: stale}},
|
||||||
|
[]UpdateRestorePoint{{Tier: UpdateTierLocal, ProvenAt: justNow}}, true, UpdateTierLocal},
|
||||||
|
{"a stale off-site copy is backed up first", []UpdateRestorePoint{{Tier: UpdateTierOffsite, ProvenAt: stale}},
|
||||||
|
[]UpdateRestorePoint{{Tier: UpdateTierLocal, ProvenAt: justNow}, {Tier: UpdateTierOffsite, ProvenAt: stale}}, true, UpdateTierLocal},
|
||||||
|
{"a stale second drive does not block a fresh own unit", []UpdateRestorePoint{{Tier: UpdateTierSecondDrive, ProvenAt: stale}, {Tier: UpdateTierLocal, ProvenAt: fresh}},
|
||||||
|
nil, false, UpdateTierLocal},
|
||||||
|
}
|
||||||
|
for _, tc := range cases {
|
||||||
|
t.Run(tc.name, func(t *testing.T) {
|
||||||
|
g, st := failedUpdateHold(t, tc.points, tc.afterBackup)
|
||||||
|
if got := hasGuardCall(g, "BackupNow"); got != tc.wantBackup {
|
||||||
|
t.Errorf("backed up first = %v, want %v (calls %v)", got, tc.wantBackup, g.callList())
|
||||||
|
}
|
||||||
|
if g.holdRP.Tier != tc.wantTier {
|
||||||
|
t.Errorf("the chosen copy is tier %d, want %d (hold %+v, err %q)", g.holdRP.Tier, tc.wantTier, g.holdRP, st.UpdateError)
|
||||||
|
}
|
||||||
|
if age := slice4T0.Sub(g.holdRP.ProvenAt); age > 24*time.Hour {
|
||||||
|
t.Errorf("the chosen copy is %s old — past the age limit", age)
|
||||||
|
}
|
||||||
|
})
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestR475_M_TheLimitIsInclusiveOnOneClock(t *testing.T) {
|
||||||
|
fresh := freshRestorePoint(slice4T0, 24*time.Hour)
|
||||||
|
for _, tier := range []int{UpdateTierSecondDrive, UpdateTierLocal, UpdateTierOffsite} {
|
||||||
|
if !fresh(UpdateRestorePoint{Tier: tier, ProvenAt: slice4T0.Add(-24 * time.Hour)}) {
|
||||||
|
t.Errorf("tier %d: exactly the limit old is still fresh", tier)
|
||||||
|
}
|
||||||
|
if fresh(UpdateRestorePoint{Tier: tier, ProvenAt: slice4T0.Add(-24*time.Hour - time.Second)}) {
|
||||||
|
t.Errorf("tier %d: a second past the limit is stale", tier)
|
||||||
|
}
|
||||||
|
if fresh(UpdateRestorePoint{Tier: tier}) {
|
||||||
|
t.Errorf("tier %d: a copy with no proven time is never fresh", tier)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -62,7 +62,7 @@ func TestSlice4_BackupRowUnitFieldsAreUnchangedByTheExtraction(t *testing.T) {
|
|||||||
// manager here on purpose, so the unguarded StartStack call panics and this test fails.
|
// manager here on purpose, so the unguarded StartStack call panics and this test fails.
|
||||||
func TestSlice4_DriveReturnGateSkipsAHeldApp(t *testing.T) {
|
func TestSlice4_DriveReturnGateSkipsAHeldApp(t *testing.T) {
|
||||||
s, _, m := newOffboxWebServer(t)
|
s, _, m := newOffboxWebServer(t)
|
||||||
if err := m.HoldAfterFailedUpdate("held", time.Now(), time.Now()); err != nil {
|
if err := m.HoldAfterFailedUpdate("held", time.Now(), time.Now(), backup.UpdateTierLocal); err != nil {
|
||||||
t.Fatal(err)
|
t.Fatal(err)
|
||||||
}
|
}
|
||||||
defer func() {
|
defer func() {
|
||||||
|
|||||||
Reference in New Issue
Block a user