controller v0.239.0: any backup tier lets an app update (R-475)
gates / gates (push) Successful in 14s

Operator ruling 2026-09-13. The update precondition walks Tier 2, Tier 1
(own recovery unit, "helyi") and Tier 3 (off-site, 15 s bound; unreachable
counts as absent with a WARN) and leans on the first FRESH copy; the
backup_max_age rule applies to whichever tier is chosen. No copy anywhere:
back up first. Refused only when nothing exists and no backup can be taken.
RunAppBackupNow tolerates a Tier-2 failure (WARN) and marks the captured
unit proven current. The hold names the tier (második meghajtó / saját
meghajtó / távoli mentés) and the date; pre-v0.239.0 holds keep their text.
A successful off-site restore now lifts an update hold. The backups page
still uses Tier2UnitRestorePoint unchanged.

Scenarios G-M tested; red-proofs M, L, the tail and the off-site clear in
felhom.eu documentation/audits/rulings-r472-r475-2026-09-13/.
This commit is contained in:
2026-09-13 17:16:24 +02:00
parent f946b0d0ca
commit b93c1543da
16 changed files with 1028 additions and 114 deletions
+51
View File
@@ -1,3 +1,54 @@
## v0.239.0 — any backup tier lets an app update (2026-09-13, R-475)
**MinAgent: 0.129.0** (unchanged)
**Operator ruling 2026-09-13: every backup counts.** Until now the update's precondition was the
Tier-2 unit predicate alone, so an app with no second-drive copy could never be updated. On demo-hp
`gokapi` and `nextcloud` were in exactly that state, each with a fresh recovery unit on its own drive.
**WHAT CHANGED.**
- **`backup.Manager.UpdateRestorePoints`**, beside `Tier2UnitRestorePoint`. It walks the tiers in the
ruling's order: Tier 2 (second drive), Tier 1 (the app's own unit, `ListRestorePoints`, „helyi"),
Tier 3 (`OffsiteInventoryList`, bounded by 15 s). It returns the first copy the caller accepts and
stops there, so a fresh Tier-2 copy never reaches the network. An unreachable off-site repository
counts as ABSENT, with a WARN. A box with no off-site target is plainly absent, with no WARN.
- **The age rule is one rule.** stacks passes `freshRestorePoint` (`update.backup_max_age`) as the
acceptance test, so the limit applies to whichever tier is chosen. The first FRESH copy wins, so a
stale Tier-2 mirror never forces a backup while the app's own unit is minutes old.
- **No copy anywhere → back up first**, then re-read every tier. The preflight refuses `no_backup`
only when there is no copy on any tier AND `CanBackUpApp` says no backup can be taken now. The
Hungarian sentence changed to say that; it no longer tells the customer to switch on the 2nd backup.
- **`RunAppBackupNow` tolerates no Tier-2 target.** A Tier-2 failure after the unit capture is a WARN.
The captured unit's manifest is marked proven current, because the capture's checksum skip leaves it
untouched on a quiet app, and Tier 1 is aged by the newest artifact's mtime. Without that mark an
app with no database and no volume would be refused forever — the ProvenCopyTime trap, one tier down.
- **The hold names the tier.** `settings.RestoreHold.CopyTier`; the sentence now ends *„Visszaállítható
a Mentések oldalon ebből a biztonsági mentésből: <második meghajtó | saját meghajtó | távoli mentés>,
<dátum>."* A hold written by v0.237.0–v0.238.1 has no tier and keeps its original sentence
(`UpdateHoldLegacyFmt`). The update journal records `proven_tier`, so a resumed update names the
right copy.
- **A successful off-site restore lifts an update hold.** That path never went through
`RestoreFromRecoveryUnitAt`, so a hold naming „távoli mentés" could otherwise never be cleared.
- **Unchanged:** the backups page still calls `Tier2UnitRestorePoint` for „Teljes visszaállítás".
**TESTS.** `internal/stacks/update_tiers_test.go` — G (Tier 2 chosen when present), H (own unit alone),
I (off-site alone), K (nothing anywhere: backed up first and the update completes; a failing backup
moves nothing), L (an existing copy still carries an app that cannot be backed up; control: without it
the app is refused), M (a stale copy on each tier is backed up first; a stale Tier 2 does not block a
fresh Tier 1; the limit is inclusive on one clock). `internal/backup/update_tiers_test.go` — the tier
order and early stop, H/I at the source, J (unreachable → absent + WARN; no target → silent; a hanging
repository ends at the bound), the hold text per tier and the legacy text, the tolerant pre-backup
tail, `RunAppBackupNow` really calling it, `CanBackUpApp`, and the off-site restore clearing the hold.
`cmd/controller/r475_wiring_test.go` — the tier numbers agree across packages; the adapter reads every
tier and never `Tier2UnitRestorePoint`; the hold is told the tier. Slice 4's tests were moved onto the
new seam (`TestSlice4_C` is now the L refusal).
**Red-proofs** (felhom.eu `documentation/audits/rulings-r472-r475-2026-09-13/`): **M** — age checked
only for Tier 2: three M cases fail (a 30-hour-old own unit and off-site copy carry the update with no
backup). **L** — refuse even with a copy: the L test fails. **Tail** — the proven-current mark misses:
the tail test fails on the mtime. **Off-site clear** removed: the hold stays and the test fails.
## v0.238.1 — the nightly backup leaves an app alone WHILE it is being updated, not only once it is held (2026-09-13, slice 4 follow-up)
**MinAgent: 0.129.0** (unchanged)
+3 -2
View File
@@ -122,8 +122,9 @@
| `Manager.recordInstalledImages` (v0.233.0) | controller/internal/stacks/installed.go | `(name, stackDir string, env []string)` | writing down what each compose SERVICE is ACTUALLY running, into `app.yaml.installed_images` | Called after a successful compose up from `StartStack`/`RestartStack`/`UpdateStack`/`runComposeDeploy`. **Reads the CONTAINER, never `docker-compose.yml`** — that file is the value the syncer has already moved (spike §3: 25 minutes of disagreement). **A failed write NEVER refuses the action** — the deliberate OPPOSITE of `SetDesiredState`: intent refused, observation logged at ERROR. **NOT from `StartStackServices`** (the R-47 DB-only window would overwrite a complete record with a partial one). Skips the write when ref+digest are unchanged, and carries `at` forward so it means "running since". Its OWN seam (`installedExecFn`) with a **context + 30 s timeout** — the two existing exec helpers have neither |
| `Manager.SetPin` / `AdoptPins` / `RenderPlanFor` / `AppliedComposePath` (v0.235.0) | controller/internal/stacks/pin.go | `SetPin(name, stackDir, pin, composeSrc) error` | THE version freeze — `app.yaml.pinned_images` + the stored `applied-compose.yml` | **`PinnedImages` is INTENT, `InstalledImages` is an OBSERVATION — never feed one from the other** (the R-166 category error, one field over). Four writers only: deploy, the guarded update (via `advancePinToCatalog`, which advances the pin and re-renders BEFORE the pull, and REFUSES the update if the pin cannot be written; a failed pull puts the pin BACK via `SetPin`), the restore adapter (this is what closes R-441), and `AdoptPins`. `AdoptPins` reuses `observationCoversTemplate` — do NOT write a second completeness rule — and skips loudly rather than inventing a pin. Absent pin = pre-v0.235.0 behaviour |
| `Manager.UpdatePreflight` / `StartGuardedUpdate` / `RecoverUpdates` / `ResumeInterruptedUpdates` (v0.237.0) | controller/internal/stacks/update.go | `UpdatePreflight(name) *UpdateRefusal`; `StartGuardedUpdate(name) error` | THE update — refusals, then a 202 job with phases on `Stack.Updating/UpdatePhase/UpdateError` | **Never report an update complete before health is known (R-443).** Every cheap refusal runs BEFORE the intent write. Safety dump BEFORE the pin moves; pin BEFORE pull; pull failure → pin back; health failure → stop + HOLD, pin stays. Journal-before-mutate (`update-journal.json`); `RecoverUpdates` MUST run before the boot sweep and `ResumeInterruptedUpdates` AFTER `SetUpdateGuards`. Seams: `updateComposeFn`, `updateHealthFn`, `updateMemoryFn`, `updateDiskFreeFn`, `updateNowFn` (R-457: the age check and the test read ONE clock). Unwired guards ⇒ every update refused |
| `stacks.UpdateGuards` + `updateGuardsAdapter` (v0.237.0) | controller/internal/stacks/update.go, controller/cmd/controller/main.go | `HoldFor`, `Busy`, `RestorePoint`, `BackupNow`, `SafetyDump`, `HoldAfterFailedUpdate` | the ONLY bridge from the update job to the backup side (stacks cannot import backup) | Wired by `stackMgr.SetUpdateGuards` — pinned by `TestSlice4_UpdateGuardsAreWiredAtStartup`. Add a guard HERE, never by importing backup into stacks |
| `backup.Manager.Tier2UnitRestorePoint` + `Tier2RestorePoint.ProvenCopyTime` (v0.237.0) | controller/internal/backup/update_guard.go | `(stack) (Tier2RestorePoint, error)` | "can this app be restored from Tier 2, and from when" — the predicate that gates BOTH the „Teljes visszaállítás" action and an update | **ONE predicate, two callers** (extracted from `buildAppBackupRows`, not copied). `CopyDate` is what the page NAMES (the package date, R-403); `ProvenCopyTime` is how OLD the data is — the last successful copy, because the manifest's `created_at` moves only when the DEFINITION changes (measured: a fresh dump under a 22-h-older manifest). Do not age a copy by `CopyDate` |
| `stacks.UpdateGuards` + `updateGuardsAdapter` (v0.237.0; tiers v0.239.0) | controller/internal/stacks/update.go, controller/cmd/controller/main.go | `HoldFor`, `Busy`, `RestorePoints`, `CanBackUp`, `BackupNow`, `SafetyDump`, `HoldAfterFailedUpdate` | the ONLY bridge from the update job to the backup side (stacks cannot import backup) | Wired by `stackMgr.SetUpdateGuards` — pinned by `TestSlice4_UpdateGuardsAreWiredAtStartup`. Add a guard HERE, never by importing backup into stacks |
| `backup.Manager.Tier2UnitRestorePoint` + `Tier2RestorePoint.ProvenCopyTime` (v0.237.0) | controller/internal/backup/update_guard.go | `(stack) (Tier2RestorePoint, error)` | "can this app be restored from Tier 2, and from when" — the predicate that gates BOTH the „Teljes visszaállítás" action and an update | **ONE predicate, two callers** (extracted from `buildAppBackupRows`, not copied). `CopyDate` is what the page NAMES (the package date, R-403); `ProvenCopyTime` is how OLD the data is — the last successful copy, because the manifest's `created_at` moves only when the DEFINITION changes (measured: a fresh dump under a 22-h-older manifest). Do not age a copy by `CopyDate`. **Since v0.239.0 the UPDATE no longer calls it directly** — it goes through `UpdateRestorePoints` (next row); the page still does |
| `backup.Manager.UpdateRestorePoints` + `CanBackUpApp` (v0.239.0, R-475) | controller/internal/backup/update_guard.go | `(ctx, stack, accept func(UpdateTierPoint) bool) (UpdateTierPoint, bool, []UpdateTierPoint)` | "which backup can this update lean on" — walks Tier 2, 1, 3 and returns the first copy `accept` admits | **The age rule lives in the caller's `accept`** (stacks' `freshRestorePoint`), so it is ONE rule for every tier. Stops at the first accepted copy, so Tier 3 (restic, 15 s bound, unreachable = absent + WARN) is reached only when needed. Seams: `updateTier2PointFn`, `updateTier1PointsFn`, `updateOffsiteInvFn`. Tier numbers are pinned equal across stacks/backup by `TestR475_TierConstantsAgree` |
| `backup.Manager.HoldAfterFailedUpdate` / `RunAppBackupNow` / `WriteUpdateSafetyDump` / `UpdateBusy` (v0.237.0) | controller/internal/backup/update_guard.go | see file | the update's hold, per-app backup-now, safety dump, busy check | The hold is `settings.RestoreHold` with `Reason: update_failed` — SAME store and gate as R-379, never a second map. `RunAppBackupNow` composes the nightly legs for ONE app (admission, DB dump, volume dump, capture, Tier-2) — do not write a second backup orchestration. A successful unit restore lifts an UPDATE hold only. **`isHeld` is ALSO true while a guarded update is moving the app (`SetUpdatingCheck`, v0.238.1)** — found live: the periodic capture overwrote a primary unit during a health wait |
| `stacks.Manager.memoryVerdict` (v0.237.0) | controller/internal/stacks/deploy.go | `(newReq, newLimit, releasedReq, releasedLimit int) (refusal, warning string)` | the deploy's memory check, shared with the update | An update RELEASES the app's current request first. Deploy passes `0, 0` and is byte-identical in wording and log line |
| `Syncer.SetRenderPlanFn` + `renderSource` (v0.235.0) | controller/internal/sync/sync.go | `func(appName string) stacks.RenderPlan` | the catalog render table | **NIL-SAFE: no seam = copy verbatim = the old product.** Catalog images == pin → verbatim (fixes flow + self-healing, both deliberately kept); differ → the WHOLE stored definition, **never a ref substitution into a newer template** (`wger 2.6`). `.felhom.yml` always verbatim (R-458). The syncer must NEVER read app.yaml. Re-reads the applied file before writing it — a test caught it writing an empty compose over a live app |
+16 -7
View File
@@ -570,22 +570,31 @@ job, and answers **202**. The page polls `GET /api/stacks/{name}`.
**Refused with 409 before anything moves:** held; a backup, restore, app-data op or quiesce holding
the app; a migration; already updating; deploying; not enough memory for the NEW template's request
(the deploy's own `memoryVerdict`, releasing the app's current request); less than **2 GB** free on
the Docker data root (a fixed floor — image sizes are not known without a registry query); and **no
restorable backup** (no openable Tier-2 recovery unit with a proven copy).
the Docker data root (a fixed floor — image sizes are not known without a registry query); and — since
v0.239.0 — **no copy on any backup tier AND no way to take one now** (drive unresolvable, disconnected,
or a migration running).
**The sequence.** A proven copy older than `update.backup_max_age` is refreshed first
(`RunAppBackupNow`: this app's DB dump, volume dump, unit capture, Tier-2 copy). Then a database
**Any backup tier counts (v0.239.0, R-475).** The update leans on the first FRESH copy in the order
second drive (Tier 2), the app's own recovery unit (Tier 1, „helyi"), off-site (Tier 3, looked up with
a 15 s bound — unreachable counts as absent, with a WARN). `update.backup_max_age` applies to whichever
tier is chosen. An app with no copy anywhere is backed up first. Tier 2 is required nowhere in the
update path; the backups page's „Teljes visszaállítás" still reads the Tier-2 predicate alone.
**The sequence.** With no fresh copy on any tier the app is backed up first (`RunAppBackupNow`: this
app's DB dump, volume dump, unit capture, then a Tier-2 copy whose failure is only a WARN — the unit
just captured is marked proven current, so a quiet app's own unit counts as fresh). Then a database
safety dump, then the pin moves, then pull, `up`, and the health wait (`.felhom.yml` check, or 60 s of
every container running for an app with none; bounded by `update.health_timeout`). **A failed pull
puts the pin back. An app that does not become healthy is stopped and HELD** — the pin stays on the
new version, and the hold sentence names the backup it can be restored from. A successful unit
restore lifts an update hold.
new version, and the hold sentence names the tier and the date of the copy it can be restored from
(„második meghajtó" / „saját meghajtó" / „távoli mentés"). A successful unit restore — and, since
v0.239.0, a successful off-site restore — lifts an update hold.
**Config (`controller.yaml`):**
```yaml
update:
backup_max_age: 24h # a proven Tier-2 copy older than this is refreshed before the update
backup_max_age: 24h # the chosen copy (any tier) must be younger than this, or the app is backed up first
health_timeout: 5m # how long the new version has to become healthy before the app is held
```
+25 -9
View File
@@ -3378,16 +3378,32 @@ func (a *updateGuardsAdapter) Busy(name string) (bool, string) {
return a.b.UpdateBusy(name)
}
func (a *updateGuardsAdapter) RestorePoint(name string) (stacks.UpdateRestorePoint, error) {
// RestorePoints (R-475) reads EVERY tier through backup.UpdateRestorePoints. It must never go back to
// Tier2UnitRestorePoint alone — pinned by TestR475_AdapterReadsEveryTier.
func (a *updateGuardsAdapter) RestorePoints(ctx context.Context, name string, accept func(stacks.UpdateRestorePoint) bool) (stacks.UpdateRestorePoint, bool, []stacks.UpdateRestorePoint) {
if a.b == nil {
return stacks.UpdateRestorePoint{}, fmt.Errorf("backup is not enabled on this box")
return stacks.UpdateRestorePoint{}, false, nil
}
rp, err := a.b.Tier2UnitRestorePoint(name)
if err != nil {
return stacks.UpdateRestorePoint{}, err
conv := func(p backup.UpdateTierPoint) stacks.UpdateRestorePoint {
return stacks.UpdateRestorePoint{Tier: p.Tier, ProvenAt: p.At}
}
at, proven := rp.ProvenCopyTime()
return stacks.UpdateRestorePoint{Restorable: rp.Restorable, Proven: proven, ProvenAt: at}, nil
var acc func(backup.UpdateTierPoint) bool
if accept != nil {
acc = func(p backup.UpdateTierPoint) bool { return accept(conv(p)) }
}
chosen, ok, seen := a.b.UpdateRestorePoints(ctx, name, acc)
out := make([]stacks.UpdateRestorePoint, 0, len(seen))
for _, p := range seen {
out = append(out, conv(p))
}
return conv(chosen), ok, out
}
func (a *updateGuardsAdapter) CanBackUp(name string) (bool, string) {
if a.b == nil {
return false, "backup is not enabled on this box"
}
return a.b.CanBackUpApp(name)
}
func (a *updateGuardsAdapter) BackupNow(ctx context.Context, name string) error {
@@ -3404,9 +3420,9 @@ func (a *updateGuardsAdapter) SafetyDump(ctx context.Context, name string) ([]st
return a.b.WriteUpdateSafetyDump(ctx, name)
}
func (a *updateGuardsAdapter) HoldAfterFailedUpdate(name string, at, provenCopyAt time.Time) error {
func (a *updateGuardsAdapter) HoldAfterFailedUpdate(name string, at time.Time, rp stacks.UpdateRestorePoint) error {
if a.b == nil {
return fmt.Errorf("backup is not enabled on this box — the hold cannot be recorded")
}
return a.b.HoldAfterFailedUpdate(name, at, provenCopyAt)
return a.b.HoldAfterFailedUpdate(name, at, rp.ProvenAt, rp.Tier)
}
@@ -0,0 +1,71 @@
package main
import (
"go/ast"
"go/parser"
"go/token"
"strings"
"testing"
"gitea.dooplex.hu/admin/felhom-controller/internal/backup"
"gitea.dooplex.hu/admin/felhom-controller/internal/stacks"
)
// R-475 — the update reads EVERY backup tier, through the one adapter.
func TestR475_TierConstantsAgree(t *testing.T) {
if stacks.UpdateTierLocal != backup.UpdateTierLocal ||
stacks.UpdateTierSecondDrive != backup.UpdateTierSecondDrive ||
stacks.UpdateTierOffsite != backup.UpdateTierOffsite {
t.Fatal("stacks and backup number the tiers differently — a hold would name the wrong copy")
}
}
// adapterMethodSelectors returns the selector names used in updateGuardsAdapter.<name>'s body.
func adapterMethodSelectors(t *testing.T, name string) string {
t.Helper()
fset := token.NewFileSet()
f, err := parser.ParseFile(fset, "main.go", nil, 0)
if err != nil {
t.Fatal(err)
}
for _, d := range f.Decls {
fn, ok := d.(*ast.FuncDecl)
if !ok || fn.Recv == nil || fn.Name.Name != name || fn.Body == nil {
continue
}
star, ok := fn.Recv.List[0].Type.(*ast.StarExpr)
if !ok {
continue
}
if id, ok := star.X.(*ast.Ident); !ok || id.Name != "updateGuardsAdapter" {
continue
}
var names []string
ast.Inspect(fn.Body, func(n ast.Node) bool {
if sel, ok := n.(*ast.SelectorExpr); ok {
names = append(names, sel.Sel.Name)
}
return true
})
return " " + strings.Join(names, " ") + " "
}
t.Fatalf("updateGuardsAdapter.%s not found in main.go", name)
return ""
}
func TestR475_AdapterReadsEveryTier(t *testing.T) {
rp := adapterMethodSelectors(t, "RestorePoints")
if !strings.Contains(rp, " UpdateRestorePoints ") {
t.Errorf("the adapter must read every tier through backup.UpdateRestorePoints; selectors:%s", rp)
}
if strings.Contains(rp, " Tier2UnitRestorePoint ") {
t.Error("the Tier-2-only predicate must not come back into the update path (R-475)")
}
if cb := adapterMethodSelectors(t, "CanBackUp"); !strings.Contains(cb, " CanBackUpApp ") {
t.Errorf("CanBackUp must ask backup.CanBackUpApp; selectors:%s", cb)
}
if h := adapterMethodSelectors(t, "HoldAfterFailedUpdate"); !strings.Contains(h, " Tier ") {
t.Errorf("the hold must be told the chosen TIER, or it cannot name it; selectors:%s", h)
}
}
+15 -8
View File
@@ -24,8 +24,9 @@ import (
// every path here refuses, or fails at the pin (the app's catalog template is absent on purpose).
type apiFakeGuards struct {
b *backup.Manager
rp stacks.UpdateRestorePoint
b *backup.Manager
points []stacks.UpdateRestorePoint
cannotBackUp bool
// blindToHolds makes the manager-side preflight NOT see holds, so a test can prove the ROUTER's
// own hold check refuses — the two layers are each pinned separately (the preflight's by
// TestSlice4_D_CheapRefusals/held). Without it, removing either layer passes inertly, because the
@@ -40,14 +41,20 @@ func (g *apiFakeGuards) HoldFor(n string) (bool, string) {
return g.b.RestoreHoldFor(n)
}
func (g *apiFakeGuards) Busy(string) (bool, string) { return false, "" }
func (g *apiFakeGuards) RestorePoint(string) (stacks.UpdateRestorePoint, error) {
return g.rp, nil
func (g *apiFakeGuards) RestorePoints(_ context.Context, _ string, accept func(stacks.UpdateRestorePoint) bool) (stacks.UpdateRestorePoint, bool, []stacks.UpdateRestorePoint) {
for _, p := range g.points {
if accept == nil || accept(p) {
return p, true, g.points
}
}
return stacks.UpdateRestorePoint{}, false, g.points
}
func (g *apiFakeGuards) CanBackUp(string) (bool, string) { return !g.cannotBackUp, "fake: no drive" }
func (g *apiFakeGuards) BackupNow(context.Context, string) error { return nil }
func (g *apiFakeGuards) SafetyDump(context.Context, string) ([]string, error) {
return nil, nil
}
func (g *apiFakeGuards) HoldAfterFailedUpdate(string, time.Time, time.Time) error { return nil }
func (g *apiFakeGuards) HoldAfterFailedUpdate(string, time.Time, stacks.UpdateRestorePoint) error { return nil }
const slice4AppYAML = "deployed: true\nenv: {}\npinned_images:\n app: nginx:1.27\n"
@@ -82,7 +89,7 @@ func newSlice4Router(t *testing.T) (*Router, *settings.Settings, *apiFakeGuards,
t.Fatal(err)
}
b := backup.NewManager(cfg, sett, lg)
g := &apiFakeGuards{b: b, rp: stacks.UpdateRestorePoint{Restorable: true, Proven: true, ProvenAt: time.Now().Add(-time.Hour)}}
g := &apiFakeGuards{b: b, points: []stacks.UpdateRestorePoint{{Tier: stacks.UpdateTierSecondDrive, ProvenAt: time.Now().Add(-time.Hour)}}}
m.SetUpdateGuards(g)
return &Router{cfg: cfg, stackMgr: m, backupMgr: b, logger: lg}, sett, g, dir
}
@@ -105,7 +112,7 @@ func postUpdate(t *testing.T, r *Router) (int, apiResponse) {
// message — which is what proves the router line is the one doing it.
func TestR439_UpdateOfAHeldAppIsRefused(t *testing.T) {
r, sett, g, dir := newSlice4Router(t)
g.rp = stacks.UpdateRestorePoint{} // the preflight's own refusal would say "no backup" — not the hold
g.points, g.cannotBackUp = nil, true // the preflight's own refusal would say "no backup" — not the hold
g.blindToHolds = true // only the router's line can produce the hold's sentence
if err := sett.SetRestoreHold(settings.RestoreHold{Stack: "app", At: "2026-09-13T08:00:00Z", Reason: settings.HoldReasonUpdateFailed, CopyDate: "2026-09-13T01:30:00Z"}); err != nil {
t.Fatal(err)
@@ -124,7 +131,7 @@ func TestR439_UpdateOfAHeldAppIsRefused(t *testing.T) {
func TestSlice4_Router_NoBackupIs409AndRecordsNothing(t *testing.T) {
r, _, g, dir := newSlice4Router(t)
g.rp = stacks.UpdateRestorePoint{Restorable: false}
g.points, g.cannotBackUp = nil, true
before, _ := os.ReadFile(filepath.Join(dir, "app.yaml"))
code, resp := postUpdate(t, r)
if code != http.StatusConflict || resp.Error != fmt.Sprintf(stacks.MsgUpdateNoBackupFmt, "app") {
+7
View File
@@ -173,6 +173,13 @@ type Manager struct {
// updatingCheck (slice 4) — nil-safe; see isHeld / SetUpdatingCheck.
updatingCheck func(stackName string) bool
// R-475 update-precondition seams, one per tier. Nil → the real Tier2UnitRestorePoint /
// ListRestorePoints / OffsiteInventoryList. They let a test reach Tier 1 and Tier 3 without a drive
// or a restic repository; see UpdateRestorePoints.
updateTier2PointFn func(stackName string) (Tier2RestorePoint, error)
updateTier1PointsFn func(stackName string) ([]RestorePoint, bool)
updateOffsiteInvFn func(ctx context.Context) (OffsiteInventory, error)
// R-354 volume-REPLAY seam — the mirror of the F17 DB seams above, so the off-site path's new
// volume leg is unit-testable without Docker. Nil → the real restoreDockerVolumesFrom.
volumeReplayFrom func(stackName, dumpDir string) (int, error)
@@ -344,7 +344,11 @@ func (m *Manager) RestoreHoldFor(stack string) (bool, string) {
if h.CopyDate != "" {
copyDate = fmtHoldTime(h.CopyDate)
}
return true, fmt.Sprintf(UpdateHoldFmt, stack, fmtHoldTime(h.At), copyDate)
// R-475: name the tier when the hold recorded one; an older hold keeps its own sentence.
if label := UpdateTierLabel(h.CopyTier); label != "" && h.CopyDate != "" {
return true, fmt.Sprintf(UpdateHoldFmt, stack, fmtHoldTime(h.At), label, copyDate)
}
return true, fmt.Sprintf(UpdateHoldLegacyFmt, stack, fmtHoldTime(h.At), copyDate)
}
when := h.At
if t, err := time.Parse(time.RFC3339, h.At); err == nil {
@@ -824,6 +828,11 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string, ack
if err := restartStack(); err != nil {
return res, fmt.Errorf("a(z) %s újraindítása sikertelen a fájlok visszaállítása után: %w", stack, err)
}
// R-475: an update hold may now name the OFF-SITE copy, and this is the route back it names. The
// unit restore clears it in RestoreFromRecoveryUnitAt; this path never went through that function,
// so without this line a successful off-site restore would leave the app refusing its next start.
// Pinned by TestR475_OffsiteRestoreClearsAnUpdateHold.
m.clearUpdateHoldAfterRestore(stack)
if err := m.waitForHealthy(stack, 90*time.Second); err != nil {
m.logger.Printf("[WARN] [offbox] %s reconstituted but health check failed: %v", stack, err)
}
@@ -68,7 +68,7 @@ func TestSlice4_UpdateHoldTextNamesTheTimeAndTheCopy(t *testing.T) {
m := &Manager{logger: log.New(io.Discard, "", 0), settings: sett}
at := time.Date(2026, 9, 13, 8, 0, 0, 0, time.UTC)
copyAt := time.Date(2026, 9, 13, 1, 30, 0, 0, time.UTC)
if err := m.HoldAfterFailedUpdate("bookstack", at, copyAt); err != nil {
if err := m.HoldAfterFailedUpdate("bookstack", at, copyAt, UpdateTierSecondDrive); err != nil {
t.Fatal(err)
}
h, ok := sett.GetRestoreHold("bookstack")
@@ -77,7 +77,7 @@ func TestSlice4_UpdateHoldTextNamesTheTimeAndTheCopy(t *testing.T) {
}
held, why := m.RestoreHoldFor("bookstack")
// Budapest is UTC+2 in September: 08:00Z → 10:00, 01:30Z → 03:30.
want := fmt.Sprintf(UpdateHoldFmt, "bookstack", "2026-09-13 10:00", "2026-09-13 03:30")
want := fmt.Sprintf(UpdateHoldFmt, "bookstack", "2026-09-13 10:00", "második meghajtó", "2026-09-13 03:30")
if !held || why != want {
t.Errorf("hold text =\n%q\nwant\n%q", why, want)
}
@@ -98,7 +98,7 @@ func TestSlice4_RestoreHoldTextIsUnchanged(t *testing.T) {
func TestSlice4_ASuccessfulRestoreClearsOnlyAnUpdateHold(t *testing.T) {
sett := slice4Settings(t)
m := &Manager{logger: log.New(io.Discard, "", 0), settings: sett}
_ = m.HoldAfterFailedUpdate("upd", time.Now(), time.Now())
_ = m.HoldAfterFailedUpdate("upd", time.Now(), time.Now(), UpdateTierLocal)
_ = sett.SetRestoreHold(settings.RestoreHold{Stack: "rst", At: "2026-08-22T14:00:00Z"})
m.clearUpdateHoldAfterRestore("upd")
m.clearUpdateHoldAfterRestore("rst")
@@ -136,7 +136,7 @@ func TestSlice4_UpdateBusy(t *testing.T) {
func TestSlice4_NightlyLegsLeaveAHeldAppAlone(t *testing.T) {
h := newAdmissionHarness(t, "held", "free")
h.m.settings = slice4Settings(t)
if err := h.m.HoldAfterFailedUpdate("held", time.Now(), time.Now()); err != nil {
if err := h.m.HoldAfterFailedUpdate("held", time.Now(), time.Now(), UpdateTierSecondDrive); err != nil {
t.Fatal(err)
}
h.m.runVolumeDumps()
+183 -9
View File
@@ -4,6 +4,7 @@ import (
"context"
"errors"
"fmt"
"os"
"path/filepath"
"time"
@@ -93,6 +94,149 @@ func (p Tier2RestorePoint) ProvenCopyTime() (time.Time, bool) {
return t, true
}
// ── R-475: any backup tier lets an app update (operator ruling 2026-09-13, controller v0.239.0) ────
//
// Until v0.239.0 the update's precondition was Tier2UnitRestorePoint alone, so an app with no second
// drive could never be updated — even with a fresh recovery unit on its own drive and an off-site
// snapshot from last night. The ruling: every backup counts. Tier2UnitRestorePoint itself is NOT
// changed; the backups page still calls it for the „Teljes visszaállítás" action, which really does
// restore from the second drive only.
// Backup tiers, as the update precondition and the hold sentence name them.
const (
UpdateTierLocal = 1 // the app's own recovery unit on its drive — „helyi" on the restore page
UpdateTierSecondDrive = 2 // the Tier-2 mirror on another drive
UpdateTierOffsite = 3 // the off-site restic repository
)
// updateTierOrder is the preference order the ruling set: the second drive, then the app's own unit,
// then off-site. The first tier holding a copy the caller ACCEPTS is chosen.
var updateTierOrder = []int{UpdateTierSecondDrive, UpdateTierLocal, UpdateTierOffsite}
// UpdateTierLabel is a tier's name in the customer's hold sentence. "" for an unknown tier.
func UpdateTierLabel(tier int) string {
switch tier {
case UpdateTierSecondDrive:
return "második meghajtó"
case UpdateTierLocal:
return "saját meghajtó"
case UpdateTierOffsite:
return "távoli mentés"
}
return ""
}
// updateOffsiteCheckTimeout bounds the off-site lookup. An update must not stall on an unreachable
// Storage Box: past this the off-site copy counts as ABSENT (with a WARN), and the update carries on
// with backing up first. A var only so a test can shorten it.
var updateOffsiteCheckTimeout = 15 * time.Second
// UpdateTierPoint is one proven, restorable copy of an app on one tier.
type UpdateTierPoint struct {
Tier int
// At is when the data in that copy was last proven written: Tier 2 ProvenCopyTime, Tier 1 the
// newest artifact of the unit (ListRestorePoints), Tier 3 the newest snapshot for the app.
At time.Time
}
// UpdateRestorePoints walks the tiers in preference order (2, 1, 3) and returns the FIRST copy that
// accept admits (nil accepts any), whether one was found, and every copy it looked at on the way.
//
// It stops at the first accepted copy, so a box with a fresh second-drive copy never touches the
// network. The AGE rule is the caller's (stacks applies backup_max_age through accept) — that is what
// makes "the age rule applies to whichever tier is chosen" one rule, not three (R-475 Scenario M).
func (m *Manager) UpdateRestorePoints(ctx context.Context, stackName string, accept func(UpdateTierPoint) bool) (UpdateTierPoint, bool, []UpdateTierPoint) {
var seen []UpdateTierPoint
for _, tier := range updateTierOrder {
p, ok := m.updateTierPoint(ctx, stackName, tier)
if !ok {
continue
}
seen = append(seen, p)
if accept == nil || accept(p) {
return p, true, seen
}
}
return UpdateTierPoint{}, false, seen
}
func (m *Manager) updateTierPoint(ctx context.Context, stackName string, tier int) (UpdateTierPoint, bool) {
switch tier {
case UpdateTierSecondDrive:
get := m.updateTier2PointFn
if get == nil {
get = m.Tier2UnitRestorePoint
}
rp, err := get(stackName)
if err != nil {
if m.isDebug() {
m.logger.Printf("[DEBUG] [backup] update precondition for %s: no Tier-2 copy (%v)", stackName, err)
}
return UpdateTierPoint{}, false
}
at, ok := rp.ProvenCopyTime()
return UpdateTierPoint{Tier: tier, At: at}, ok
case UpdateTierLocal:
list := m.updateTier1PointsFn
if list == nil {
list = m.ListRestorePoints
}
pts, _ := list(stackName)
for _, rp := range pts {
if at, err := time.Parse(time.RFC3339, rp.Time); err == nil {
return UpdateTierPoint{Tier: tier, At: at}, true
}
}
return UpdateTierPoint{}, false
case UpdateTierOffsite:
inv := m.updateOffsiteInvFn
if inv == nil {
if m.settings == nil || !m.OffboxConfigured() {
return UpdateTierPoint{}, false
}
inv = m.OffsiteInventoryList
}
cctx, cancel := context.WithTimeout(ctx, updateOffsiteCheckTimeout)
defer cancel()
got, err := inv(cctx)
if err != nil {
if !errors.Is(err, errNoOffsiteTarget) {
m.logger.Printf("[WARN] [backup] update precondition for %s: the off-site copy could not be checked within %s (%v) — counted as ABSENT", stackName, updateOffsiteCheckTimeout, err)
}
return UpdateTierPoint{}, false
}
for _, a := range got.Apps {
if a.App == stackName && !a.LatestAt.IsZero() {
return UpdateTierPoint{Tier: tier, At: a.LatestAt}, true
}
}
}
return UpdateTierPoint{}, false
}
// CanBackUpApp reports whether "back up first" can run for this app at all right now — the second
// half of R-475 Scenario L: an app with no copy anywhere is refused only when this is false too.
// Cheap and read-only; RunAppBackupNow re-checks everything when it actually runs.
func (m *Manager) CanBackUpApp(stackName string) (bool, string) {
if m == nil {
return false, "backup is not enabled on this box"
}
if m.stackProvider == nil {
return false, "stack provider not configured"
}
if m.migrationActive() {
return false, "a data migration is running"
}
drivePath := m.GetAppDrivePath(stackName)
if drivePath == "" || !filepath.IsAbs(drivePath) {
return false, "the app's drive cannot be resolved"
}
if m.settings != nil && (m.settings.IsDisconnected(drivePath) || m.settings.IsDecommissioned(drivePath)) {
return false, fmt.Sprintf("the app's drive %s is not available", drivePath)
}
return true, ""
}
// UpdateBusy reports whether something else is ALREADY touching this app's data, which refuses an
// update before anything moves (slice 4 Scenario D). The reason is operator-English; the customer
// sentence is chosen by the caller.
@@ -142,6 +286,7 @@ func (m *Manager) RunAppBackupNow(ctx context.Context, stackName string) error {
}
m.logger.Printf("[INFO] [backup] update pre-backup for %s: starting (DB dump → volume dump → unit capture → Tier 2)", stackName)
start := time.Now()
var nsRoot string
legErr := func() error {
defer m.releaseRunning()
defer m.beginAdmissionRun()()
@@ -156,7 +301,7 @@ func (m *Manager) RunAppBackupNow(ctx context.Context, stackName string) error {
if !m.admitApp(stackName) {
return fmt.Errorf("nincs elég szabad hely a mentéshez a(z) %s meghajtón", drivePath)
}
nsRoot := m.namespaceRoot(drivePath)
nsRoot = m.namespaceRoot(drivePath)
discover := m.discoverDBs
if discover == nil {
@@ -203,16 +348,35 @@ func (m *Manager) RunAppBackupNow(ctx context.Context, stackName string) error {
return legErr
}
m.updatePreBackupTail(stackName, nsRoot, time.Now())
m.logger.Printf("[INFO] [backup] update pre-backup for %s: complete in %s", stackName, time.Since(start).Round(time.Millisecond))
return nil
}
// updatePreBackupTail is what "back up first" does after the capture succeeded (R-475).
//
// 1. It marks the app's OWN unit as proven current NOW. CaptureRecoveryUnit leaves the manifest alone
// when nothing changed (the checksum skip), and Tier 1's age is the newest artifact's mtime — so on
// an app with no database and no named volume a fresh "back up first" would leave Tier 1 as old as
// its last definition change, and the update would be refused forever. That is the trap
// ProvenCopyTime documents for Tier 2, one tier down. The capture has just compared the unit with
// the live definition, so "current as of now" is exactly what it established.
// 2. It runs the Tier-2 copy, and a Tier-2 failure is a WARN, not a failure: the update may lean on
// any tier, and the Tier-1 unit it can lean on was just written. Pinned by
// TestR475_PreBackupTail_Tier2FailureIsAWarnAndTheOwnUnitIsFresh.
func (m *Manager) updatePreBackupTail(stackName, nsRoot string, now time.Time) {
if nsRoot != "" {
if err := os.Chtimes(RecoveryUnitManifestPath(nsRoot, stackName), now, now); err != nil {
m.logger.Printf("[WARN] [backup] update pre-backup for %s: could not mark the recovery unit as proven current (%v) — its own-unit copy may read older than it is", stackName, err)
}
}
runOne := m.perAppTier2
if runOne == nil {
runOne = m.RunTier2
}
if err := runOne(stackName); err != nil {
m.logger.Printf("[ERROR] [backup] update pre-backup for %s: Tier 2 copy FAILED: %v", stackName, err)
return fmt.Errorf("a másodlagos másolat elkészítése sikertelen: %w", err)
m.logger.Printf("[WARN] [backup] update pre-backup for %s: Tier 2 copy FAILED: %v — not fatal: the app's own recovery unit was just captured, and an update may lean on any tier (R-475)", stackName, err)
}
m.logger.Printf("[INFO] [backup] update pre-backup for %s: complete in %s", stackName, time.Since(start).Round(time.Millisecond))
return nil
}
// WriteUpdateSafetyDump takes the last-minute database copy an update makes just before it moves the
@@ -244,9 +408,18 @@ func (m *Manager) WriteUpdateSafetyDump(ctx context.Context, stackName string) (
}
// UpdateHoldFmt is the customer sentence for an app held after a failed update. Arguments: the app,
// the time of the failure, and the PROVEN date of the copy it can be restored from. One named string so
// a test asserts it verbatim instead of retyping Hungarian (R-364).
// the time of the failure, the TIER of the copy it can be restored from (UpdateTierLabel), and that
// copy's PROVEN date. One named string so a test asserts it verbatim instead of retyping Hungarian
// (R-364). Since v0.239.0 (R-475) it names the tier: the copy may be on any of three, and each is
// restored from a different place on the Mentések page.
const UpdateHoldFmt = "A(z) %s frissítése %s-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. " +
"Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. " +
"Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: %s, %s."
// UpdateHoldLegacyFmt is the v0.237.0–v0.238.1 sentence, kept for a hold written before the tier was
// recorded (CopyTier 0) — every such hold named a Tier-2 copy, but it did not SAY so, and rewriting
// it now would state a fact the record does not hold.
const UpdateHoldLegacyFmt = "A(z) %s frissítése %s-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. " +
"Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. " +
"Visszaállítható a(z) %s-i biztonsági mentésből a Mentések oldalon."
@@ -276,7 +449,7 @@ func fmtHoldTime(rfc3339 string) string {
// has just stopped the app on the strength of this record, and an unrecorded hold is a stopped app
// that the next restart button will quietly start again. The caller logs it at ERROR and keeps the
// failure on the page.
func (m *Manager) HoldAfterFailedUpdate(stackName string, at time.Time, copyDate time.Time) error {
func (m *Manager) HoldAfterFailedUpdate(stackName string, at time.Time, copyDate time.Time, copyTier int) error {
if m == nil || m.settings == nil {
return fmt.Errorf("no settings wired — the update hold for %s cannot be persisted", stackName)
}
@@ -287,11 +460,12 @@ func (m *Manager) HoldAfterFailedUpdate(stackName string, at time.Time, copyDate
}
if !copyDate.IsZero() {
h.CopyDate = copyDate.UTC().Format(time.RFC3339)
h.CopyTier = copyTier
}
if err := m.settings.SetRestoreHold(h); err != nil {
return fmt.Errorf("persisting the update hold for %s: %w", stackName, err)
}
m.logger.Printf("[WARN] [backup] %s is HELD STOPPED after a failed update (restore point: %s)", stackName, h.CopyDate)
m.logger.Printf("[WARN] [backup] %s is HELD STOPPED after a failed update (restore point: tier %d %q, %s)", stackName, h.CopyTier, UpdateTierLabel(h.CopyTier), h.CopyDate)
return nil
}
@@ -0,0 +1,317 @@
package backup
import (
"bytes"
"context"
"errors"
"fmt"
"go/ast"
"go/parser"
"go/token"
"io"
"log"
"os"
"path/filepath"
"strings"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-controller/internal/settings"
)
// R-475 — the backup side of "any backup tier lets an app update" (controller v0.239.0).
var r475T0 = time.Date(2026, 9, 13, 12, 0, 0, 0, time.UTC)
func r475Manager() (*Manager, *bytes.Buffer) {
var buf bytes.Buffer
return &Manager{logger: log.New(&buf, "", 0)}, &buf
}
func tier2At(at time.Time) func(string) (Tier2RestorePoint, error) {
return func(string) (Tier2RestorePoint, error) {
ts := at.UTC().Format(time.RFC3339)
return Tier2RestorePoint{Restorable: true, CopyDateProven: true, CopyLastSuccess: ts, CopyDate: ts}, nil
}
}
func noTier2(string) (Tier2RestorePoint, error) {
return Tier2RestorePoint{}, errors.New("no Tier-2 copy recorded for this app")
}
func tier1At(at time.Time) func(string) ([]RestorePoint, bool) {
return func(string) ([]RestorePoint, bool) {
return []RestorePoint{{Time: at.UTC().Format(time.RFC3339), ShortID: "helyi", Tier: 1}}, true
}
}
func noTier1(string) ([]RestorePoint, bool) { return []RestorePoint{}, true }
func offsiteWith(app string, at time.Time) func(context.Context) (OffsiteInventory, error) {
return func(context.Context) (OffsiteInventory, error) {
return OffsiteInventory{Apps: []OffsiteInventoryApp{{App: app, LatestAt: at}}}, nil
}
}
func noOffsiteTarget(context.Context) (OffsiteInventory, error) {
return OffsiteInventory{}, errNoOffsiteTarget
}
func tiersOf(ps []UpdateTierPoint) []int {
out := make([]int, 0, len(ps))
for _, p := range ps {
out = append(out, p.Tier)
}
return out
}
// G — the order is 2, 1, 3, and the walk STOPS at the first accepted copy (so a box with a fresh
// second-drive copy never reaches the network).
func TestR475_G_TierOrderIsSecondDriveThenOwnUnitThenOffsite(t *testing.T) {
m, _ := r475Manager()
offsiteCalls := 0
m.updateTier2PointFn = tier2At(r475T0.Add(-1 * time.Hour))
m.updateTier1PointsFn = tier1At(r475T0.Add(-2 * time.Hour))
m.updateOffsiteInvFn = func(ctx context.Context) (OffsiteInventory, error) {
offsiteCalls++
return offsiteWith("gokapi", r475T0.Add(-3*time.Hour))(ctx)
}
ctx := context.Background()
p, ok, seen := m.UpdateRestorePoints(ctx, "gokapi", nil)
if !ok || p.Tier != UpdateTierSecondDrive || fmt.Sprint(tiersOf(seen)) != "[2]" || offsiteCalls != 0 {
t.Errorf("all three present: want tier 2 and a stop; got %+v ok=%v seen=%v offsite calls=%d", p, ok, tiersOf(seen), offsiteCalls)
}
p, ok, seen = m.UpdateRestorePoints(ctx, "gokapi", func(p UpdateTierPoint) bool { return p.Tier != UpdateTierSecondDrive })
if !ok || p.Tier != UpdateTierLocal || fmt.Sprint(tiersOf(seen)) != "[2 1]" {
t.Errorf("tier 2 refused: want tier 1; got %+v seen=%v", p, tiersOf(seen))
}
p, ok, seen = m.UpdateRestorePoints(ctx, "gokapi", func(p UpdateTierPoint) bool { return p.Tier == UpdateTierOffsite })
if !ok || p.Tier != UpdateTierOffsite || !p.At.Equal(r475T0.Add(-3*time.Hour)) || fmt.Sprint(tiersOf(seen)) != "[2 1 3]" {
t.Errorf("tiers 2 and 1 refused: want tier 3; got %+v seen=%v", p, tiersOf(seen))
}
if _, ok, seen = m.UpdateRestorePoints(ctx, "gokapi", func(UpdateTierPoint) bool { return false }); ok || len(seen) != 3 {
t.Errorf("nothing accepted: want not found with all three seen; ok=%v seen=%v", ok, tiersOf(seen))
}
}
func TestR475_H_OwnUnitOnly(t *testing.T) {
m, _ := r475Manager()
m.updateTier2PointFn = noTier2
m.updateTier1PointsFn = tier1At(r475T0.Add(-2 * time.Hour))
m.updateOffsiteInvFn = noOffsiteTarget
p, ok, _ := m.UpdateRestorePoints(context.Background(), "gokapi", nil)
if !ok || p.Tier != UpdateTierLocal || !p.At.Equal(r475T0.Add(-2*time.Hour)) {
t.Fatalf("got %+v ok=%v", p, ok)
}
// A Tier-2 record that was only ATTEMPTED is not a copy (R-101) — the own unit still wins.
m.updateTier2PointFn = func(string) (Tier2RestorePoint, error) {
return Tier2RestorePoint{Restorable: true, CopyDateProven: false}, nil
}
if p, ok, _ = m.UpdateRestorePoints(context.Background(), "gokapi", nil); !ok || p.Tier != UpdateTierLocal {
t.Errorf("an unproven Tier-2 record must be skipped; got %+v", p)
}
// No unit on disk (ListRestorePoints' empty list) is no copy.
m.updateTier1PointsFn = noTier1
if _, ok, _ = m.UpdateRestorePoints(context.Background(), "gokapi", nil); ok {
t.Error("no copy on any tier must be not found")
}
}
func TestR475_I_OffsiteOnly(t *testing.T) {
m, _ := r475Manager()
m.updateTier2PointFn, m.updateTier1PointsFn = noTier2, noTier1
m.updateOffsiteInvFn = offsiteWith("gokapi", r475T0.Add(-5*time.Hour))
if p, ok, _ := m.UpdateRestorePoints(context.Background(), "gokapi", nil); !ok || p.Tier != UpdateTierOffsite {
t.Fatalf("got %+v ok=%v", p, ok)
}
// Another app's snapshot is not this app's copy.
if _, ok, _ := m.UpdateRestorePoints(context.Background(), "nextcloud", nil); ok {
t.Error("a snapshot tagged for a different app must not count")
}
}
// J — an unreachable off-site repository is ABSENT with a WARN, and it is bounded in time.
func TestR475_J_OffsiteUnreachableIsAbsentWithAWarn(t *testing.T) {
m, buf := r475Manager()
m.updateTier2PointFn, m.updateTier1PointsFn = noTier2, noTier1
m.updateOffsiteInvFn = func(context.Context) (OffsiteInventory, error) {
return OffsiteInventory{}, errors.New("ssh: connect to host: connection timed out")
}
if _, ok, _ := m.UpdateRestorePoints(context.Background(), "gokapi", nil); ok {
t.Error("an unreachable off-site copy must count as absent")
}
if !strings.Contains(buf.String(), "[WARN]") || !strings.Contains(buf.String(), "counted as ABSENT") {
t.Errorf("an unreachable off-site copy must WARN; log = %q", buf.String())
}
// Control: a box with NO off-site target is plainly absent — that is not a fault, so no WARN.
buf.Reset()
m.updateOffsiteInvFn = noOffsiteTarget
if _, ok, _ := m.UpdateRestorePoints(context.Background(), "gokapi", nil); ok || strings.Contains(buf.String(), "WARN") {
t.Errorf("no off-site target: want absent and silent; ok=%v log=%q", ok, buf.String())
}
// The bound: a repository that never answers is given up on at the timeout.
old := updateOffsiteCheckTimeout
updateOffsiteCheckTimeout = 50 * time.Millisecond
defer func() { updateOffsiteCheckTimeout = old }()
buf.Reset()
m.updateOffsiteInvFn = func(ctx context.Context) (OffsiteInventory, error) {
<-ctx.Done()
return OffsiteInventory{}, ctx.Err()
}
start := time.Now()
_, ok, _ := m.UpdateRestorePoints(context.Background(), "gokapi", nil)
if took := time.Since(start); ok || took > 5*time.Second || !strings.Contains(buf.String(), "counted as ABSENT") {
t.Errorf("a hanging off-site check must end at its bound as absent with a WARN; ok=%v took=%s log=%q", ok, took, buf.String())
}
}
// The hold names the tier. The labels are the ruling's exact words.
func TestR475_HoldTextNamesTheTier(t *testing.T) {
at := time.Date(2026, 9, 13, 8, 0, 0, 0, time.UTC)
copyAt := time.Date(2026, 9, 13, 1, 30, 0, 0, time.UTC)
for tier, label := range map[int]string{
UpdateTierSecondDrive: "második meghajtó",
UpdateTierLocal: "saját meghajtó",
UpdateTierOffsite: "távoli mentés",
} {
sett := slice4Settings(t)
m := &Manager{logger: log.New(io.Discard, "", 0), settings: sett}
if err := m.HoldAfterFailedUpdate("gokapi", at, copyAt, tier); err != nil {
t.Fatal(err)
}
held, why := m.RestoreHoldFor("gokapi")
want := fmt.Sprintf(UpdateHoldFmt, "gokapi", "2026-09-13 10:00", label, "2026-09-13 03:30")
if !held || why != want {
t.Errorf("tier %d: hold text =\n%q\nwant\n%q", tier, why, want)
}
if !strings.HasSuffix(why, "ebből a biztonsági mentésből: "+label+", 2026-09-13 03:30.") {
t.Errorf("tier %d: the sentence must END naming the tier and the date, got %q", tier, why)
}
}
}
func TestR475_AHoldWrittenBeforeTheTierKeepsItsSentence(t *testing.T) {
sett := slice4Settings(t)
m := &Manager{logger: log.New(io.Discard, "", 0), settings: sett}
if err := sett.SetRestoreHold(settings.RestoreHold{Stack: "uptime-kuma", At: "2026-09-13T10:18:02Z",
Reason: settings.HoldReasonUpdateFailed, CopyDate: "2026-09-13T10:09:51Z"}); err != nil {
t.Fatal(err)
}
_, why := m.RestoreHoldFor("uptime-kuma")
// The exact sentence quoted in audits/slice4-2026-09-13 (13-H-refusals.txt), live on v0.238.1.
if want := fmt.Sprintf(UpdateHoldLegacyFmt, "uptime-kuma", "2026-09-13 12:18", "2026-09-13 12:09"); why != want {
t.Errorf("a v0.238.1 hold must keep its sentence:\n%q\nwant\n%q", why, want)
}
}
// RunAppBackupNow's tail: a Tier-2 failure no longer fails "back up first", and the own unit it
// captured reads as fresh even when the capture found nothing to rewrite.
//
// COMPANION RED-PROOF (REPORT.md): aim the Chtimes at a path that does not exist — this test fails on
// the mtime: "back up first" would then leave a quiet app's own unit as old as its last definition change.
func TestR475_PreBackupTail_Tier2FailureIsAWarnAndTheOwnUnitIsFresh(t *testing.T) {
nsRoot := t.TempDir()
mp := RecoveryUnitManifestPath(nsRoot, "gokapi")
if err := os.MkdirAll(filepath.Dir(mp), 0o755); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(mp, []byte(`{"app_name":"gokapi"}`), 0o644); err != nil {
t.Fatal(err)
}
old := time.Now().Add(-30 * time.Hour)
if err := os.Chtimes(mp, old, old); err != nil {
t.Fatal(err)
}
m, buf := r475Manager()
tier2Ran := false
m.perAppTier2 = func(string) error { tier2Ran = true; return errors.New("no second drive with room") }
now := time.Now()
m.updatePreBackupTail("gokapi", nsRoot, now)
if !tier2Ran {
t.Fatal("the Tier-2 copy must still be attempted")
}
if !strings.Contains(buf.String(), "[WARN]") || !strings.Contains(buf.String(), "Tier 2 copy FAILED") {
t.Errorf("a Tier-2 failure must be logged as a WARN; log = %q", buf.String())
}
fi, err := os.Stat(mp)
if err != nil {
t.Fatal(err)
}
if d := fi.ModTime().Sub(now); d < -time.Second || d > time.Second {
t.Errorf("the own unit must read as proven NOW (mtime %s, now %s)", fi.ModTime(), now)
}
// Control: the same unit through ListRestorePoints' own rule reads fresh — the newest artifact.
buf.Reset()
m.perAppTier2 = func(string) error { return nil }
m.updatePreBackupTail("gokapi", nsRoot, now)
if strings.Contains(buf.String(), "WARN") {
t.Errorf("a successful Tier-2 copy must not WARN; log = %q", buf.String())
}
}
// The tail is really what RunAppBackupNow runs — a helper tested and never called is the "seam built
// but never wired" shape this project has shipped repeatedly.
func TestR475_RunAppBackupNowUsesTheTolerantTail(t *testing.T) {
fset := token.NewFileSet()
f, err := parser.ParseFile(fset, "update_guard.go", nil, 0)
if err != nil {
t.Fatal(err)
}
var calls []string
found := false
for _, d := range f.Decls {
fn, ok := d.(*ast.FuncDecl)
if !ok || fn.Name.Name != "RunAppBackupNow" {
continue
}
found = true
ast.Inspect(fn.Body, func(n ast.Node) bool {
if sel, ok := n.(*ast.SelectorExpr); ok {
calls = append(calls, sel.Sel.Name)
}
return true
})
}
joined := " " + strings.Join(calls, " ") + " "
if !found || !strings.Contains(joined, " updatePreBackupTail ") {
t.Fatalf("RunAppBackupNow must end in updatePreBackupTail; selectors = %s", joined)
}
if strings.Contains(joined, " RunTier2 ") || strings.Contains(joined, " perAppTier2 ") {
t.Error("RunAppBackupNow must not run the Tier-2 copy itself any more — only through the tolerant tail")
}
}
func TestR475_CanBackUpApp(t *testing.T) {
var nilM *Manager
if ok, why := nilM.CanBackUpApp("gokapi"); ok || why == "" {
t.Error("no backup manager: cannot back up, with a reason")
}
m, _ := r475Manager()
if ok, why := m.CanBackUpApp("gokapi"); ok || !strings.Contains(why, "stack provider") {
t.Errorf("no stack provider: cannot back up; got ok=%v why=%q", ok, why)
}
}
// Tier 3's route back is the off-site restore, and it never went through RestoreFromRecoveryUnitAt —
// so it must lift an update hold itself, or a hold naming „távoli mentés" could never be cleared.
//
// COMPANION RED-PROOF (REPORT.md): remove the clearUpdateHoldAfterRestore call from
// ReconstituteFromOffsite — this test fails with the hold still in place.
func TestR475_OffsiteRestoreClearsAnUpdateHold(t *testing.T) {
m, prov, _ := reconFixture(t, "run1", "2026-07-19T06:00:00Z", "")
prov.composePath = writeLiveCompose(t, noDBCompose)
m.discoverDBs = func(context.Context) ([]DiscoveredDB, error) { return nil, nil }
if err := m.HoldAfterFailedUpdate("immich", r475T0, r475T0.Add(-time.Hour), UpdateTierOffsite); err != nil {
t.Fatal(err)
}
if held, why := m.RestoreHoldFor("immich"); !held || !strings.Contains(why, "távoli mentés") {
t.Fatalf("fixture: the app must be held naming the off-site copy, got held=%v %q", held, why)
}
if _, err := m.ReconstituteFromOffsite(context.Background(), "immich", false); err != nil {
t.Fatalf("reconstitute: %v", err)
}
if held, _ := m.RestoreHoldFor("immich"); held {
t.Error("a successful off-site restore is the route back the hold names — the hold must be lifted")
}
}
+4
View File
@@ -1598,6 +1598,10 @@ type RestoreHold struct {
// CopyDate is the RFC3339 time of the proven backup the customer is told they can restore from.
// Only set for HoldReasonUpdateFailed.
CopyDate string `json:"copy_date,omitempty"`
// CopyTier (R-475, v0.239.0) is WHICH backup tier CopyDate belongs to: 1 the app's own recovery
// unit, 2 the second drive, 3 off-site. 0 means a hold written before the field existed; it keeps
// its original sentence (backup.UpdateHoldLegacyFmt). Only set for HoldReasonUpdateFailed.
CopyTier int `json:"copy_tier,omitempty"`
}
// Hold reasons. See RestoreHold.Reason.
+98 -35
View File
@@ -6,6 +6,7 @@ import (
"fmt"
"os"
"path/filepath"
"strings"
"time"
"gitea.dooplex.hu/admin/felhom-controller/internal/system"
@@ -82,7 +83,7 @@ const (
MsgUpdateAlreadyFmt = "A(z) %s frissítése már folyamatban van."
MsgUpdateBusy = "A frissítés most nem indítható: mentés/visszaállítás folyamatban. Próbáld újra, ha befejeződött."
MsgUpdateMigrating = "A frissítés most nem indítható: adatáthelyezés folyamatban."
MsgUpdateNoBackupFmt = "A(z) %s nem frissíthető, mert nincs olyan biztonsági mentése, amelyből vissza lehetne állítani. Kapcsold be a 2. mentést az alkalmazás mentési beállításainál a Mentések oldalon, és várd meg az első sikeres másolatot — utána a frissítés elindítható."
MsgUpdateNoBackupFmt = "A(z) %s nem frissíthető, mert nincs olyan biztonsági mentése, amelyből vissza lehetne állítani, és most új mentés sem készíthető róla. Ellenőrizd a Mentések oldalon, hogy az alkalmazás meghajtója elérhető-e — utána a frissítés elindítható."
MsgUpdateDiskFmt = "Nincs elég szabad hely a frissítéshez: %.1f GB szabad, az új verzió letöltéséhez legalább %.0f GB szükséges."
MsgUpdateBackupFailFmt = "A frissítés nem indult el, mert a frissítés előtti biztonsági mentés nem sikerült: %v. Az alkalmazás változatlanul fut tovább."
MsgUpdateBackupNoUnit = "A frissítés nem indult el: a frissítés előtti mentés lefutott, de nem jött létre friss, visszaállítható másolat. Az alkalmazás változatlanul fut tovább."
@@ -107,11 +108,55 @@ const updateSettleWindow = 60 * time.Second
// updatePollEvery is how often the health wait re-reads the stack.
const updatePollEvery = 5 * time.Second
// UpdateRestorePoint is the precondition answer, reduced to what the update needs.
// Backup tiers (R-475), mirroring backup.UpdateTier* — stacks cannot import backup, so
// TestR475_TierConstantsAgree (cmd/controller) pins the two sets equal.
const (
UpdateTierLocal = 1 // the app's own recovery unit
UpdateTierSecondDrive = 2 // the Tier-2 mirror on another drive
UpdateTierOffsite = 3 // off-site
)
// UpdateRestorePoint is one proven, restorable copy the update may lean on. The backup side returns
// only proven, restorable copies (never an attempt clock, never an unopenable unit), so there is no
// "maybe" field here to forget to check.
type UpdateRestorePoint struct {
Restorable bool // an openable recovery unit exists in the Tier-2 copy
Proven bool // a copy actually succeeded (never an attempt clock)
ProvenAt time.Time // when the data in that copy was last proven copied
Tier int // UpdateTierSecondDrive / UpdateTierLocal / UpdateTierOffsite
ProvenAt time.Time // when the data in that copy was last proven written
}
func updateTierName(tier int) string {
switch tier {
case UpdateTierSecondDrive:
return "Tier 2 (second drive)"
case UpdateTierLocal:
return "Tier 1 (own recovery unit)"
case UpdateTierOffsite:
return "Tier 3 (off-site)"
}
return fmt.Sprintf("tier %d", tier)
}
// freshRestorePoint is THE age rule (backup_max_age), applied to whichever tier is being considered
// — R-475 Scenario M: a stale copy on ANY tier is stale.
//
// COMPANION RED-PROOF M (REPORT.md): check the age only for Tier 2 (let any other tier through
// whatever its age). TestR475_M_TheAgeRuleAppliesToTheChosenTier then fails: a 30-hour-old copy of
// the app's own unit carries the update with no backup first.
func freshRestorePoint(now time.Time, maxAge time.Duration) func(UpdateRestorePoint) bool {
return func(p UpdateRestorePoint) bool {
return !p.ProvenAt.IsZero() && now.Sub(p.ProvenAt) <= maxAge
}
}
func describeRestorePoints(now time.Time, pts []UpdateRestorePoint) string {
if len(pts) == 0 {
return "none"
}
parts := make([]string, 0, len(pts))
for _, p := range pts {
parts = append(parts, fmt.Sprintf("%s at %s (%s old)", updateTierName(p.Tier), p.ProvenAt.UTC().Format(time.RFC3339), now.Sub(p.ProvenAt).Round(time.Minute)))
}
return strings.Join(parts, "; ")
}
// UpdateGuards is everything the update needs from the backup side. The stacks package cannot import
@@ -119,10 +164,14 @@ type UpdateRestorePoint struct {
type UpdateGuards interface {
HoldFor(name string) (bool, string)
Busy(name string) (bool, string)
RestorePoint(name string) (UpdateRestorePoint, error)
// RestorePoints walks the tiers in preference order (2, 1, 3) and returns the first copy accept
// admits (nil = any), whether one was found, and every copy looked at (R-475).
RestorePoints(ctx context.Context, name string, accept func(UpdateRestorePoint) bool) (UpdateRestorePoint, bool, []UpdateRestorePoint)
// CanBackUp reports whether "back up first" can run for this app now (R-475 Scenario L).
CanBackUp(name string) (bool, string)
BackupNow(ctx context.Context, name string) error
SafetyDump(ctx context.Context, name string) ([]string, error)
HoldAfterFailedUpdate(name string, at, provenCopyAt time.Time) error
HoldAfterFailedUpdate(name string, at time.Time, rp UpdateRestorePoint) error
}
// SetUpdateGuards wires the backup side. INIT-ONLY. Unwired, every update is refused (fail closed):
@@ -199,10 +248,19 @@ func (m *Manager) UpdatePreflight(name string) *UpdateRefusal {
if m.IsMigrating() {
return m.refuseUpdate(name, "migrating", MsgUpdateMigrating, "a data migration is running")
}
rp, err := g.RestorePoint(name)
if err != nil || !rp.Restorable || !rp.Proven {
return m.refuseUpdate(name, "no_backup", fmt.Sprintf(MsgUpdateNoBackupFmt, name),
fmt.Sprintf("no restorable proven Tier-2 unit (restorable=%v proven=%v err=%v)", rp.Restorable, rp.Proven, err))
// R-475: any tier counts, and an app with no copy at all is backed up first by the job. So the only
// refusal left here is Scenario L — no copy on any tier AND no way to make one now. (With a copy
// but no way to back up, the job still applies the age rule and refuses then if the copy is stale.)
// Tier 3 is looked at only on this branch, so an ordinary update never waits on the network here.
//
// COMPANION RED-PROOF (REPORT.md): drop the `!found` condition. TestR475_L then fails — an app
// with a copy but no way to back up is refused too.
if canBackUp, why := g.CanBackUp(name); !canBackUp {
if _, found, seen := g.RestorePoints(context.Background(), name, nil); !found {
return m.refuseUpdate(name, "no_backup", fmt.Sprintf(MsgUpdateNoBackupFmt, name),
fmt.Sprintf("no copy on any tier (found: %s) and no backup can be taken now: %s", describeRestorePoints(m.now(), seen), why))
}
m.logger.Printf("[WARN] [stacks] update %s: no backup can be taken now (%s) — an existing copy must carry the update", name, why)
}
if ref := m.updateMemoryRefusal(name, st); ref != nil {
return ref
@@ -383,15 +441,15 @@ func (m *Manager) runGuardedUpdate(ctx context.Context, name string) {
fail(MsgUpdateNoGuards, "no UpdateGuards wired")
return
}
rp, err := g.RestorePoint(name)
if err != nil || !rp.Restorable || !rp.Proven {
fail(fmt.Sprintf(MsgUpdateNoBackupFmt, name), fmt.Sprintf("precondition vanished: restorable=%v proven=%v err=%v", rp.Restorable, rp.Proven, err))
return
}
// R-475: the precondition is a copy on ANY tier, chosen in the order 2, 1, 3, and the age rule
// applies to whichever tier is chosen. The first FRESH copy wins — not merely the first copy — so a
// stale second-drive mirror never forces a backup while the app's own unit is minutes old.
maxAge := m.backupMaxAge()
if age := start.Sub(rp.ProvenAt); age > maxAge {
m.logger.Printf("[INFO] [stacks] update %s: the proven copy is %s old (limit %s) — backing up first", name, age.Round(time.Minute), maxAge)
rp, ok, seen := g.RestorePoints(ctx, name, freshRestorePoint(start, maxAge))
if ok {
m.logger.Printf("[INFO] [stacks] update %s: precondition met — %s copy from %s (%s old, limit %s)", name, updateTierName(rp.Tier), rp.ProvenAt.UTC().Format(time.RFC3339), start.Sub(rp.ProvenAt).Round(time.Minute), maxAge)
} else {
m.logger.Printf("[INFO] [stacks] update %s: no copy younger than %s on any tier (found: %s) — backing up first", name, maxAge, describeRestorePoints(start, seen))
if !m.enterUpdatePhase(name, &entry, UpdatePhaseBackingUp) {
fail(MsgUpdateJournalFailed, "journal write failed")
return
@@ -400,15 +458,16 @@ func (m *Manager) runGuardedUpdate(ctx context.Context, name string) {
fail(fmt.Sprintf(MsgUpdateBackupFailFmt, err), "pre-update backup: "+err.Error())
return
}
rp, err = g.RestorePoint(name)
if err != nil || !rp.Restorable || !rp.Proven || m.now().Sub(rp.ProvenAt) > maxAge {
fail(MsgUpdateBackupNoUnit, fmt.Sprintf("after the backup: restorable=%v proven=%v at=%s err=%v", rp.Restorable, rp.Proven, rp.ProvenAt.Format(time.RFC3339), err))
now := m.now()
rp, ok, seen = g.RestorePoints(ctx, name, freshRestorePoint(now, maxAge))
if !ok {
fail(MsgUpdateBackupNoUnit, fmt.Sprintf("after the backup there is still no copy younger than %s on any tier (found: %s)", maxAge, describeRestorePoints(now, seen)))
return
}
} else {
m.logger.Printf("[INFO] [stacks] update %s: precondition met — proven copy from %s (%s old, limit %s)", name, rp.ProvenAt.UTC().Format(time.RFC3339), age.Round(time.Minute), maxAge)
m.logger.Printf("[INFO] [stacks] update %s: precondition met after the backup — %s copy from %s", name, updateTierName(rp.Tier), rp.ProvenAt.UTC().Format(time.RFC3339))
}
entry.ProvenCopyAt = rp.ProvenAt.UTC().Format(time.RFC3339)
entry.ProvenTier = rp.Tier
// SAFETY DUMP BEFORE THE PIN MOVES — "a minute ago", before any migration can have run.
if !m.enterUpdatePhase(name, &entry, UpdatePhaseSafetyDump) {
@@ -472,20 +531,20 @@ func (m *Manager) runGuardedUpdate(ctx context.Context, name string) {
}
if !m.enterUpdatePhase(name, &entry, UpdatePhaseStarting) {
m.failAndHold(ctx, name, dir, env, rp.ProvenAt, "journal write failed before up")
m.failAndHold(ctx, name, dir, env, rp, "journal write failed before up")
return
}
if _, err := m.updateCompose(dir, env, "up", "-d", "--remove-orphans"); err != nil {
// Containers may already have been recreated on the new image — something may have run.
m.failAndHold(ctx, name, dir, env, rp.ProvenAt, "compose up failed: "+err.Error())
m.failAndHold(ctx, name, dir, env, rp, "compose up failed: "+err.Error())
return
}
m.verifyAndConclude(ctx, name, dir, env, rp.ProvenAt, start, &entry)
m.verifyAndConclude(ctx, name, dir, env, rp, start, &entry)
}
// verifyAndConclude is the TRUTH half (R-443): success is declared only after the app's health is
// known, and a failure holds the app.
func (m *Manager) verifyAndConclude(ctx context.Context, name, dir string, env []string, provenAt, start time.Time, entry *updateJournalEntry) {
func (m *Manager) verifyAndConclude(ctx context.Context, name, dir string, env []string, rp UpdateRestorePoint, start time.Time, entry *updateJournalEntry) {
if !m.enterUpdatePhase(name, entry, UpdatePhaseVerifying) {
m.logger.Printf("[ERROR] [stacks] update %s: could not journal the verifying phase — verifying anyway", name)
}
@@ -493,7 +552,7 @@ func (m *Manager) verifyAndConclude(ctx context.Context, name, dir string, env [
waitStart := m.now()
healthy, detail := m.updateHealth(ctx, name, timeout)
if !healthy {
m.failAndHold(ctx, name, dir, env, provenAt, "not healthy: "+detail)
m.failAndHold(ctx, name, dir, env, rp, "not healthy: "+detail)
return
}
m.logger.Printf("[INFO] [stacks] update %s: healthy after %s (%s)", name, m.now().Sub(waitStart).Round(time.Second), detail)
@@ -506,7 +565,7 @@ func (m *Manager) verifyAndConclude(ctx context.Context, name, dir string, env [
}
// failAndHold is Scenario F: stop the app, record the hold, tell the customer the route back.
func (m *Manager) failAndHold(ctx context.Context, name, dir string, env []string, provenAt time.Time, why string) {
func (m *Manager) failAndHold(ctx context.Context, name, dir string, env []string, rp UpdateRestorePoint, why string) {
m.logger.Printf("[ERROR] [stacks] update %s FAILED after the new version was started: %s — stopping and HOLDING the app; the pin stays on the new version (its migration may have run)", name, why)
if _, err := m.updateCompose(dir, env, "down"); err != nil {
m.logger.Printf("[ERROR] [stacks] update %s: stopping the failed app also failed: %v", name, err)
@@ -514,7 +573,7 @@ func (m *Manager) failAndHold(ctx context.Context, name, dir string, env []strin
msg := MsgUpdateHoldUnsaved
if g := m.guards(); g == nil {
m.logger.Printf("[ERROR] [stacks] update %s: no UpdateGuards — the hold CANNOT be recorded", name)
} else if err := g.HoldAfterFailedUpdate(name, m.now(), provenAt); err != nil {
} else if err := g.HoldAfterFailedUpdate(name, m.now(), rp); err != nil {
m.logger.Printf("[ERROR] [stacks] update %s: %v", name, err)
} else if _, why := g.HoldFor(name); why != "" {
msg = why
@@ -619,6 +678,9 @@ type updateJournalEntry struct {
PrevCompose string `json:"prev_compose,omitempty"`
PrevApplied string `json:"prev_applied,omitempty"`
ProvenCopyAt string `json:"proven_copy_at,omitempty"`
// ProvenTier (R-475) — which tier ProvenCopyAt belongs to, so a resumed update that fails names
// the right copy. 0 in a journal written by v0.238.1 or older.
ProvenTier int `json:"proven_tier,omitempty"`
}
type updateJournal struct {
@@ -787,16 +849,17 @@ func (m *Manager) ResumeInterruptedUpdates(ctx context.Context) int {
continue
}
provenAt, _ := time.Parse(time.RFC3339, e.ProvenCopyAt)
rp := UpdateRestorePoint{Tier: e.ProvenTier, ProvenAt: provenAt}
dir := filepath.Dir(st.ComposePath)
go func(name, dir string, e updateJournalEntry, provenAt time.Time) {
go func(name, dir string, e updateJournalEntry, rp UpdateRestorePoint) {
env := m.stackEnv(dir)
m.logger.Printf("[INFO] [stacks] update %s: resuming after a controller restart — `up -d` then the health wait", name)
if _, err := m.updateCompose(dir, env, "up", "-d", "--remove-orphans"); err != nil {
m.failAndHold(ctx, name, dir, env, provenAt, "resumed compose up failed: "+err.Error())
m.failAndHold(ctx, name, dir, env, rp, "resumed compose up failed: "+err.Error())
return
}
m.verifyAndConclude(ctx, name, dir, env, provenAt, e.StartedAt, &e)
}(name, dir, e, provenAt)
m.verifyAndConclude(ctx, name, dir, env, rp, e.StartedAt, &e)
}(name, dir, e, rp)
}
return len(names)
}
+48 -38
View File
@@ -25,13 +25,15 @@ type fakeGuards struct {
held bool
holdWhy string
busy bool
rp UpdateRestorePoint
rpErr error
rpAfterBackup *UpdateRestorePoint
// points are the copies the backup side holds, in tier order (R-475); pointsAfterBackup replaces
// them when BackupNow succeeds (nil = the backup changed nothing).
points []UpdateRestorePoint
pointsAfterBackup []UpdateRestorePoint
cannotBackUp bool
backupErr error
dumpErr error
holdErr error
holdProvenAt time.Time
holdRP UpdateRestorePoint
pinAtDump string
stackDir string
}
@@ -48,18 +50,29 @@ func (f *fakeGuards) HoldFor(string) (bool, string) {
return f.held, f.holdWhy
}
func (f *fakeGuards) Busy(string) (bool, string) { return f.busy, "fake busy" }
func (f *fakeGuards) RestorePoint(string) (UpdateRestorePoint, error) {
f.note("RestorePoint")
func (f *fakeGuards) RestorePoints(_ context.Context, _ string, accept func(UpdateRestorePoint) bool) (UpdateRestorePoint, bool, []UpdateRestorePoint) {
f.note("RestorePoints")
f.mu.Lock()
defer f.mu.Unlock()
return f.rp, f.rpErr
var seen []UpdateRestorePoint
for _, p := range f.points {
seen = append(seen, p)
if accept == nil || accept(p) {
return p, true, seen
}
}
return UpdateRestorePoint{}, false, seen
}
func (f *fakeGuards) CanBackUp(string) (bool, string) {
f.note("CanBackUp")
return !f.cannotBackUp, "fake: the drive is gone"
}
func (f *fakeGuards) BackupNow(context.Context, string) error {
f.note("BackupNow")
f.mu.Lock()
defer f.mu.Unlock()
if f.backupErr == nil && f.rpAfterBackup != nil {
f.rp = *f.rpAfterBackup
if f.backupErr == nil && f.pointsAfterBackup != nil {
f.points = f.pointsAfterBackup
}
return f.backupErr
}
@@ -72,14 +85,14 @@ func (f *fakeGuards) SafetyDump(context.Context, string) ([]string, error) {
}
return []string{"/fake/pre-restore-x.sql"}, f.dumpErr
}
func (f *fakeGuards) HoldAfterFailedUpdate(_ string, _ time.Time, provenAt time.Time) error {
func (f *fakeGuards) HoldAfterFailedUpdate(_ string, _ time.Time, rp UpdateRestorePoint) error {
f.note("HoldAfterFailedUpdate")
f.mu.Lock()
defer f.mu.Unlock()
if f.holdErr != nil {
return f.holdErr
}
f.held, f.holdWhy, f.holdProvenAt = true, "HELD-SENTENCE", provenAt
f.held, f.holdWhy, f.holdRP = true, "HELD-SENTENCE", rp
return nil
}
@@ -111,7 +124,7 @@ func newSlice4Manager(t *testing.T) (*Manager, string, *fakeGuards, *composeRec)
m, dir := newPinManager(t, pinTplOld, pinTplNew,
"deployed: true\nenv: {}\npinned_images:\n web: nextcloud:31.0.14-apache\n")
mustWrite(t, AppliedComposePath(dir), pinTplOld)
g := &fakeGuards{rp: UpdateRestorePoint{Restorable: true, Proven: true, ProvenAt: slice4T0.Add(-1 * time.Hour)}, stackDir: dir}
g := &fakeGuards{points: []UpdateRestorePoint{{Tier: UpdateTierSecondDrive, ProvenAt: slice4T0.Add(-1 * time.Hour)}}, stackDir: dir}
c := &composeRec{fail: map[string]error{}}
m.updateGuards = g
m.updateComposeFn = c.fn
@@ -212,9 +225,8 @@ func TestSlice4_A_SuccessIsDeclaredOnlyAfterHealth(t *testing.T) {
if got, want := strings.Join(c.list(), " | "), "pull | up -d --remove-orphans"; got != want {
t.Errorf("compose calls = %q, want %q", got, want)
}
// RestorePoint twice by design: once in the preflight (the refusal), once inside the job (the
// precondition must still hold when the job actually starts).
if got := strings.Join(g.callList(), ","); got != "RestorePoint,RestorePoint,SafetyDump" {
// R-475: the preflight asks only whether a backup could be taken; the job reads the copies once.
if got := strings.Join(g.callList(), ","); got != "CanBackUp,RestorePoints,SafetyDump" {
t.Errorf("a fresh copy needs no backup-first; guard calls = %s", got)
}
}
@@ -223,8 +235,8 @@ func TestSlice4_A_SuccessIsDeclaredOnlyAfterHealth(t *testing.T) {
func TestSlice4_B_StaleCopyIsRefreshedFirst(t *testing.T) {
m, dir, g, _ := newSlice4Manager(t)
g.rp.ProvenAt = slice4T0.Add(-30 * time.Hour) // > 24 h default
g.rpAfterBackup = &UpdateRestorePoint{Restorable: true, Proven: true, ProvenAt: slice4T0.Add(-1 * time.Minute)}
g.points[0].ProvenAt = slice4T0.Add(-30 * time.Hour) // > 24 h default
g.pointsAfterBackup = []UpdateRestorePoint{{Tier: UpdateTierSecondDrive, ProvenAt: slice4T0.Add(-1 * time.Minute)}}
if err := m.StartGuardedUpdate("nextcloud"); err != nil {
t.Fatal(err)
}
@@ -233,7 +245,7 @@ func TestSlice4_B_StaleCopyIsRefreshedFirst(t *testing.T) {
t.Fatalf("with a successful backup-first the update completes, got phase=%q err=%q", st.UpdatePhase, st.UpdateError)
}
calls := strings.Join(g.callList(), ",")
if !strings.HasPrefix(calls, "RestorePoint,RestorePoint,BackupNow,RestorePoint,SafetyDump") {
if !strings.HasPrefix(calls, "CanBackUp,RestorePoints,BackupNow,RestorePoints,SafetyDump") {
t.Errorf("a stale copy must be backed up FIRST and the precondition re-read; calls = %s", calls)
}
if got := pinOf(t, dir); got != "nextcloud:34.0.1-apache" {
@@ -243,7 +255,7 @@ func TestSlice4_B_StaleCopyIsRefreshedFirst(t *testing.T) {
func TestSlice4_B_BackupFailureMovesNothing(t *testing.T) {
m, dir, g, c := newSlice4Manager(t)
g.rp.ProvenAt = slice4T0.Add(-30 * time.Hour)
g.points[0].ProvenAt = slice4T0.Add(-30 * time.Hour)
g.backupErr = errors.New("disk full")
if err := m.StartGuardedUpdate("nextcloud"); err != nil {
t.Fatal(err)
@@ -265,7 +277,7 @@ func TestSlice4_B_BackupFailureMovesNothing(t *testing.T) {
func TestSlice4_B_BackupThatYieldsNoFreshUnitRefuses(t *testing.T) {
m, dir, g, c := newSlice4Manager(t)
g.rp.ProvenAt = slice4T0.Add(-30 * time.Hour) // stays stale: rpAfterBackup nil
g.points[0].ProvenAt = slice4T0.Add(-30 * time.Hour) // stays stale: pointsAfterBackup nil
if err := m.StartGuardedUpdate("nextcloud"); err != nil {
t.Fatal(err)
}
@@ -278,21 +290,18 @@ func TestSlice4_B_BackupThatYieldsNoFreshUnitRefuses(t *testing.T) {
}
}
// ── C: no backup exists that could restore this app ───────────────────────────────────────────────
// ── C / R-475 L: no copy on any tier, and no way to make one ─────────────────────────────────────
// COMPANION RED-PROOF 2 (REPORT.md): make the precondition in UpdatePreflight proceed when
// !rp.Restorable. This test then fails with the update started.
func TestSlice4_C_NoRestorableCopyRefusesBeforeAnythingMoves(t *testing.T) {
// Until v0.239.0 this refused any app without a restorable Tier-2 unit. R-475: every tier counts and
// an app with nothing is backed up first, so the refusal is now only "nothing anywhere AND no backup
// can be taken". The K half (nothing, but a backup CAN be taken) is TestR475_K in update_tiers_test.go.
func TestSlice4_C_NoCopyAndNoWayToBackUpRefusesBeforeAnythingMoves(t *testing.T) {
m, dir, g, c := newSlice4Manager(t)
// A PROVEN, FRESH copy whose unit cannot be opened — the realistic half-copied mirror. Proven and
// fresh on purpose: a fixture that is also unproven would be refused by the proven check alone,
// and a red-proof that drops the restorable check would then pass inertly (observed on the first
// run of red-proof 2, 2026-09-13).
g.rp = UpdateRestorePoint{Restorable: false, Proven: true, ProvenAt: slice4T0.Add(-time.Hour)}
g.points, g.cannotBackUp = nil, true
err := m.StartGuardedUpdate("nextcloud")
var ref *UpdateRefusal
if !errors.As(err, &ref) || ref.Reason != "no_backup" {
t.Fatalf("an app with no restorable copy must be REFUSED (no_backup), got %v", err)
t.Fatalf("no copy anywhere and no way to back up must be REFUSED (no_backup), got %v", err)
}
if want := fmt.Sprintf(MsgUpdateNoBackupFmt, "nextcloud"); ref.Message != want {
t.Errorf("message = %q", ref.Message)
@@ -304,10 +313,10 @@ func TestSlice4_C_NoRestorableCopyRefusesBeforeAnythingMoves(t *testing.T) {
if pinOf(t, dir) != "nextcloud:31.0.14-apache" || len(c.list()) != 0 {
t.Error("a refused update must move nothing")
}
// A copy that exists but was never PROVEN is not a copy (R-101).
g.rp = UpdateRestorePoint{Restorable: true, Proven: false}
if ref := m.UpdatePreflight("nextcloud"); ref == nil || ref.Reason != "no_backup" {
t.Errorf("an unproven copy must refuse too, got %v", ref)
for _, call := range g.callList() {
if call == "BackupNow" {
t.Error("a refused update must not try to back up")
}
}
}
@@ -412,8 +421,8 @@ func TestSlice4_F_HealthFailureHoldsTheAppAndKeepsTheNewPin(t *testing.T) {
if !held {
t.Fatal("an app that did not come up must be HELD")
}
if !g.holdProvenAt.Equal(g.rp.ProvenAt) {
t.Errorf("the hold must name the PROVEN copy date %s, got %s", g.rp.ProvenAt, g.holdProvenAt)
if !g.holdRP.ProvenAt.Equal(g.points[0].ProvenAt) || g.holdRP.Tier != UpdateTierSecondDrive {
t.Errorf("the hold must name the PROVEN copy it leans on (%+v), got %+v", g.points[0], g.holdRP)
}
if st.UpdateError != "HELD-SENTENCE" {
t.Errorf("the page must carry the hold's own sentence, got %q", st.UpdateError)
@@ -466,6 +475,7 @@ func simulateAdvanced(t *testing.T, m *Manager, dir string) updateJournalEntry {
StartedAt: slice4T0, PrevPin: map[string]string{"web": "nextcloud:31.0.14-apache"},
PrevCompose: filepath.Join(dir, preUpdateComposeFile), PrevApplied: filepath.Join(dir, preUpdateAppliedFile),
ProvenCopyAt: slice4T0.Add(-time.Hour).Format(time.RFC3339),
ProvenTier: UpdateTierLocal,
}
}
@@ -530,8 +540,8 @@ func TestSlice4_G_InterruptedAfterUpResumesTheHealthWait(t *testing.T) {
if got := strings.Join(c.list(), " | "); got != "up -d --remove-orphans | down" {
t.Errorf("resumption re-runs `up` then stops the failed app; compose calls = %q", got)
}
if !g.holdProvenAt.Equal(slice4T0.Add(-time.Hour)) {
t.Errorf("the resumed hold must name the journaled proven copy date, got %s", g.holdProvenAt)
if !g.holdRP.ProvenAt.Equal(slice4T0.Add(-time.Hour)) || g.holdRP.Tier != UpdateTierLocal {
t.Errorf("the resumed hold must name the journaled copy (tier %d at %s), got %+v", UpdateTierLocal, slice4T0.Add(-time.Hour), g.holdRP)
}
}
@@ -0,0 +1,175 @@
package stacks
import (
"context"
"errors"
"fmt"
"testing"
"time"
)
// R-475 — any backup tier lets an app update (operator ruling 2026-09-13, controller v0.239.0).
// Tier order and the Tier-3 timeout are the backup side's (internal/backup/update_tiers_test.go);
// here is the update job's half: which copy it leans on, when it backs up first, when it refuses.
// Every age here is read against slice4T0, the same clock the job reads (R-457).
func hasGuardCall(g *fakeGuards, name string) bool {
for _, c := range g.callList() {
if c == name {
return true
}
}
return false
}
// failedUpdateHold runs an update whose new version never becomes healthy, so the hold records the
// copy the job chose. That is the observable consequence of the choice: the customer is told to
// restore from exactly that copy.
func failedUpdateHold(t *testing.T, points, afterBackup []UpdateRestorePoint) (*fakeGuards, *Stack) {
t.Helper()
m, _, g, _ := newSlice4Manager(t)
g.points, g.pointsAfterBackup = points, afterBackup
m.updateHealthFn = func(context.Context, string, time.Duration) (bool, string) { return false, "crash loop" }
if err := m.StartGuardedUpdate("nextcloud"); err != nil {
t.Fatalf("start: %v", err)
}
return g, waitUpdateDone(t, m, "nextcloud")
}
func TestR475_G_TheSecondDriveIsChosenWhenPresent(t *testing.T) {
fresh := slice4T0.Add(-time.Hour)
g, st := failedUpdateHold(t, []UpdateRestorePoint{
{Tier: UpdateTierSecondDrive, ProvenAt: fresh},
{Tier: UpdateTierLocal, ProvenAt: fresh.Add(30 * time.Minute)},
{Tier: UpdateTierOffsite, ProvenAt: fresh},
}, nil)
if st.UpdatePhase != UpdatePhaseFailed || g.holdRP.Tier != UpdateTierSecondDrive || !g.holdRP.ProvenAt.Equal(fresh) {
t.Errorf("with a fresh second-drive copy that copy is chosen; hold = %+v phase=%q", g.holdRP, st.UpdatePhase)
}
if hasGuardCall(g, "BackupNow") {
t.Error("a fresh copy needs no backup first")
}
}
func TestR475_H_TheOwnUnitAloneCarriesTheUpdate(t *testing.T) {
fresh := slice4T0.Add(-2 * time.Hour)
g, _ := failedUpdateHold(t, []UpdateRestorePoint{{Tier: UpdateTierLocal, ProvenAt: fresh}}, nil)
if g.holdRP.Tier != UpdateTierLocal || !g.holdRP.ProvenAt.Equal(fresh) {
t.Errorf("an app with only its own recovery unit must update against it; hold = %+v", g.holdRP)
}
if hasGuardCall(g, "BackupNow") {
t.Error("a fresh own unit needs no backup first — Tier 2 must not be required anywhere")
}
}
func TestR475_I_TheOffsiteCopyAloneCarriesTheUpdate(t *testing.T) {
fresh := slice4T0.Add(-3 * time.Hour)
g, _ := failedUpdateHold(t, []UpdateRestorePoint{{Tier: UpdateTierOffsite, ProvenAt: fresh}}, nil)
if g.holdRP.Tier != UpdateTierOffsite || !g.holdRP.ProvenAt.Equal(fresh) || hasGuardCall(g, "BackupNow") {
t.Errorf("an app with only a fresh off-site copy must update against it; hold = %+v calls=%v", g.holdRP, g.callList())
}
}
func TestR475_K_NoCopyAnywhereIsBackedUpFirstAndTheUpdateProceeds(t *testing.T) {
m, dir, g, _ := newSlice4Manager(t)
g.points = nil
g.pointsAfterBackup = []UpdateRestorePoint{{Tier: UpdateTierLocal, ProvenAt: slice4T0.Add(-time.Minute)}}
if ref := m.UpdatePreflight("nextcloud"); ref != nil {
t.Fatalf("an app with no copy but a working backup must NOT be refused, got %q (%s)", ref.Message, ref.Reason)
}
if err := m.StartGuardedUpdate("nextcloud"); err != nil {
t.Fatal(err)
}
st := waitUpdateDone(t, m, "nextcloud")
if st.UpdatePhase != UpdatePhaseDone {
t.Fatalf("the backup made a Tier-1 copy, so the update completes; phase=%q err=%q", st.UpdatePhase, st.UpdateError)
}
if !hasGuardCall(g, "BackupNow") {
t.Error("with no copy anywhere the job must back up FIRST")
}
if got := pinOf(t, dir); got != "nextcloud:34.0.1-apache" {
t.Errorf("pin = %q", got)
}
}
func TestR475_K_NoCopyAndTheBackupFailsMovesNothing(t *testing.T) {
m, dir, g, c := newSlice4Manager(t)
g.points = nil
g.backupErr = errors.New("nincs elég szabad hely")
if err := m.StartGuardedUpdate("nextcloud"); err != nil {
t.Fatal(err)
}
st := waitUpdateDone(t, m, "nextcloud")
if want := fmt.Sprintf(MsgUpdateBackupFailFmt, g.backupErr); st.UpdateError != want {
t.Errorf("UpdateError = %q, want %q", st.UpdateError, want)
}
if pinOf(t, dir) != "nextcloud:31.0.14-apache" || len(c.list()) != 0 {
t.Error("nothing may move when the backup-first failed")
}
}
// L — nothing anywhere and no way to back up is refused (TestSlice4_C_… pins the refusal itself). The
// other half pins that the refusal is not wider than that: a copy that EXISTS still carries it.
func TestR475_L_AnExistingCopyStillCarriesAnAppThatCannotBeBackedUp(t *testing.T) {
m, _, g, _ := newSlice4Manager(t)
g.cannotBackUp = true
g.points = []UpdateRestorePoint{{Tier: UpdateTierOffsite, ProvenAt: slice4T0.Add(-time.Hour)}}
if ref := m.UpdatePreflight("nextcloud"); ref != nil {
t.Fatalf("a fresh off-site copy must carry the update even when no backup can be taken now, got %q", ref.Message)
}
g.points = nil
if ref := m.UpdatePreflight("nextcloud"); ref == nil || ref.Reason != "no_backup" {
t.Fatalf("positive control: with no copy the same app must be refused, got %v", ref)
}
}
// M — a stale copy on ANY tier is stale. Each case would pass under a Tier-2-only age check except the
// first, which is the control proving the fixture's clock and limit are real.
func TestR475_M_TheAgeRuleAppliesToTheChosenTier(t *testing.T) {
stale, fresh, justNow := slice4T0.Add(-30*time.Hour), slice4T0.Add(-time.Hour), slice4T0.Add(-time.Minute)
cases := []struct {
name string
points []UpdateRestorePoint
afterBackup []UpdateRestorePoint
wantBackup bool
wantTier int
}{
{"control: a stale second-drive copy is backed up first", []UpdateRestorePoint{{Tier: UpdateTierSecondDrive, ProvenAt: stale}},
[]UpdateRestorePoint{{Tier: UpdateTierSecondDrive, ProvenAt: justNow}}, true, UpdateTierSecondDrive},
{"a stale own unit is backed up first", []UpdateRestorePoint{{Tier: UpdateTierLocal, ProvenAt: stale}},
[]UpdateRestorePoint{{Tier: UpdateTierLocal, ProvenAt: justNow}}, true, UpdateTierLocal},
{"a stale off-site copy is backed up first", []UpdateRestorePoint{{Tier: UpdateTierOffsite, ProvenAt: stale}},
[]UpdateRestorePoint{{Tier: UpdateTierLocal, ProvenAt: justNow}, {Tier: UpdateTierOffsite, ProvenAt: stale}}, true, UpdateTierLocal},
{"a stale second drive does not block a fresh own unit", []UpdateRestorePoint{{Tier: UpdateTierSecondDrive, ProvenAt: stale}, {Tier: UpdateTierLocal, ProvenAt: fresh}},
nil, false, UpdateTierLocal},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
g, st := failedUpdateHold(t, tc.points, tc.afterBackup)
if got := hasGuardCall(g, "BackupNow"); got != tc.wantBackup {
t.Errorf("backed up first = %v, want %v (calls %v)", got, tc.wantBackup, g.callList())
}
if g.holdRP.Tier != tc.wantTier {
t.Errorf("the chosen copy is tier %d, want %d (hold %+v, err %q)", g.holdRP.Tier, tc.wantTier, g.holdRP, st.UpdateError)
}
if age := slice4T0.Sub(g.holdRP.ProvenAt); age > 24*time.Hour {
t.Errorf("the chosen copy is %s old — past the age limit", age)
}
})
}
}
func TestR475_M_TheLimitIsInclusiveOnOneClock(t *testing.T) {
fresh := freshRestorePoint(slice4T0, 24*time.Hour)
for _, tier := range []int{UpdateTierSecondDrive, UpdateTierLocal, UpdateTierOffsite} {
if !fresh(UpdateRestorePoint{Tier: tier, ProvenAt: slice4T0.Add(-24 * time.Hour)}) {
t.Errorf("tier %d: exactly the limit old is still fresh", tier)
}
if fresh(UpdateRestorePoint{Tier: tier, ProvenAt: slice4T0.Add(-24*time.Hour - time.Second)}) {
t.Errorf("tier %d: a second past the limit is stale", tier)
}
if fresh(UpdateRestorePoint{Tier: tier}) {
t.Errorf("tier %d: a copy with no proven time is never fresh", tier)
}
}
}
@@ -62,7 +62,7 @@ func TestSlice4_BackupRowUnitFieldsAreUnchangedByTheExtraction(t *testing.T) {
// manager here on purpose, so the unguarded StartStack call panics and this test fails.
func TestSlice4_DriveReturnGateSkipsAHeldApp(t *testing.T) {
s, _, m := newOffboxWebServer(t)
if err := m.HoldAfterFailedUpdate("held", time.Now(), time.Now()); err != nil {
if err := m.HoldAfterFailedUpdate("held", time.Now(), time.Now(), backup.UpdateTierLocal); err != nil {
t.Fatal(err)
}
defer func() {