From cbcca030618b2327233fe5819ded69775cacc788 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Sun, 13 Sep 2026 12:25:09 +0200 Subject: [PATCH] v0.238.1: the nightly backup leaves an app alone WHILE it is being updated, not only once it is held (slice 4 follow-up) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Found live in v0.238.0 Scenario F on demo-hp: during an update's 5-minute health wait the app is not yet held, and the periodic recovery-unit capture at 10:17:09 wrote the never-started definition (alpine:3.20) into its PRIMARY unit, 53 s before the hold landed. The Tier-2 mirror the hold names survived only because Tier 2 runs daily; a nightly Tier 2 inside a verify window would have mirrored the broken definition over the copy the customer is told to restore from. backup.Manager.isHeld — consulted by the capture sweep, the Tier-2 run and the volume dump — is now also true while a guarded update is moving the app, via SetUpdatingCheck wired in main.go to stacks.Manager.IsUpdating. Test with positive control + red-proof; wiring pinned. Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- CHANGELOG.md | 26 ++++++++++++++++++ REUSE.md | 2 +- controller/README.md | 2 +- controller/cmd/controller/main.go | 5 ++++ .../cmd/controller/slice4_wiring_test.go | 9 +++++++ controller/internal/backup/backup.go | 3 +++ .../backup/slice4_update_guard_test.go | 27 +++++++++++++++++++ controller/internal/backup/update_guard.go | 23 ++++++++++++++-- 8 files changed, 93 insertions(+), 4 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index b5a1444..eeffb58 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,3 +1,29 @@ +## v0.238.1 — the nightly backup leaves an app alone WHILE it is being updated, not only once it is held (2026-09-13, slice 4 follow-up) + +**MinAgent: 0.129.0** (unchanged) + +**Found live, not in review.** v0.238.0 Scenario F on demo-hp: an update to `alpine:3.20` (an image +that exits at once) waited its 5-minute health timeout and was held at 10:18:02 — correctly. But the +periodic recovery-unit capture ran at **10:17:09**, inside that wait, when the app was updating and +not yet held, and wrote the never-started definition into the app's **primary** unit: + +``` +primary-unit: image: alpine:3.20 manifest "created_at": "2026-09-13T10:17:09Z" +tier2-mirror: image: louislam/uptime-kuma:2.4.0 manifest "created_at": "2026-09-13T10:09:51Z" +``` + +The Tier-2 mirror the hold text names survived only because the Tier-2 run is daily. A nightly Tier-2 +falling inside a verify window would have mirrored the broken definition over the very copy the +customer is told to restore from. + +**Fix.** `backup.Manager.isHeld` — the predicate the capture sweep, the Tier-2 run and the volume dump +already consult since v0.237.0 — is also true while a guarded update is moving the app, through a new +`SetUpdatingCheck` seam wired in `main.go` to `stacks.Manager.IsUpdating`. + +**Tests.** `TestSlice4_NightlyLegsLeaveAnAppMidUpdateAlone` (all three legs skip the updating app; the +other app is still mirrored — positive control), red-proofed by removing the clause (the app is then +dumped, captured and mirrored); `TestSlice4_UpdatingCheckIsWiredAtStartup`. + ## v0.238.0 — the page follows the update, and a held app offers no way to start it (2026-09-13, update arc slice 4 Part 4) **MinAgent: 0.129.0** (unchanged) diff --git a/REUSE.md b/REUSE.md index b184632..4e62258 100644 --- a/REUSE.md +++ b/REUSE.md @@ -124,7 +124,7 @@ | `Manager.UpdatePreflight` / `StartGuardedUpdate` / `RecoverUpdates` / `ResumeInterruptedUpdates` (v0.237.0) | controller/internal/stacks/update.go | `UpdatePreflight(name) *UpdateRefusal`; `StartGuardedUpdate(name) error` | THE update — refusals, then a 202 job with phases on `Stack.Updating/UpdatePhase/UpdateError` | **Never report an update complete before health is known (R-443).** Every cheap refusal runs BEFORE the intent write. Safety dump BEFORE the pin moves; pin BEFORE pull; pull failure → pin back; health failure → stop + HOLD, pin stays. Journal-before-mutate (`update-journal.json`); `RecoverUpdates` MUST run before the boot sweep and `ResumeInterruptedUpdates` AFTER `SetUpdateGuards`. Seams: `updateComposeFn`, `updateHealthFn`, `updateMemoryFn`, `updateDiskFreeFn`, `updateNowFn` (R-457: the age check and the test read ONE clock). Unwired guards ⇒ every update refused | | `stacks.UpdateGuards` + `updateGuardsAdapter` (v0.237.0) | controller/internal/stacks/update.go, controller/cmd/controller/main.go | `HoldFor`, `Busy`, `RestorePoint`, `BackupNow`, `SafetyDump`, `HoldAfterFailedUpdate` | the ONLY bridge from the update job to the backup side (stacks cannot import backup) | Wired by `stackMgr.SetUpdateGuards` — pinned by `TestSlice4_UpdateGuardsAreWiredAtStartup`. Add a guard HERE, never by importing backup into stacks | | `backup.Manager.Tier2UnitRestorePoint` + `Tier2RestorePoint.ProvenCopyTime` (v0.237.0) | controller/internal/backup/update_guard.go | `(stack) (Tier2RestorePoint, error)` | "can this app be restored from Tier 2, and from when" — the predicate that gates BOTH the „Teljes visszaállítás" action and an update | **ONE predicate, two callers** (extracted from `buildAppBackupRows`, not copied). `CopyDate` is what the page NAMES (the package date, R-403); `ProvenCopyTime` is how OLD the data is — the last successful copy, because the manifest's `created_at` moves only when the DEFINITION changes (measured: a fresh dump under a 22-h-older manifest). Do not age a copy by `CopyDate` | -| `backup.Manager.HoldAfterFailedUpdate` / `RunAppBackupNow` / `WriteUpdateSafetyDump` / `UpdateBusy` (v0.237.0) | controller/internal/backup/update_guard.go | see file | the update's hold, per-app backup-now, safety dump, busy check | The hold is `settings.RestoreHold` with `Reason: update_failed` — SAME store and gate as R-379, never a second map. `RunAppBackupNow` composes the nightly legs for ONE app (admission, DB dump, volume dump, capture, Tier-2) — do not write a second backup orchestration. A successful unit restore lifts an UPDATE hold only | +| `backup.Manager.HoldAfterFailedUpdate` / `RunAppBackupNow` / `WriteUpdateSafetyDump` / `UpdateBusy` (v0.237.0) | controller/internal/backup/update_guard.go | see file | the update's hold, per-app backup-now, safety dump, busy check | The hold is `settings.RestoreHold` with `Reason: update_failed` — SAME store and gate as R-379, never a second map. `RunAppBackupNow` composes the nightly legs for ONE app (admission, DB dump, volume dump, capture, Tier-2) — do not write a second backup orchestration. A successful unit restore lifts an UPDATE hold only. **`isHeld` is ALSO true while a guarded update is moving the app (`SetUpdatingCheck`, v0.238.1)** — found live: the periodic capture overwrote a primary unit during a health wait | | `stacks.Manager.memoryVerdict` (v0.237.0) | controller/internal/stacks/deploy.go | `(newReq, newLimit, releasedReq, releasedLimit int) (refusal, warning string)` | the deploy's memory check, shared with the update | An update RELEASES the app's current request first. Deploy passes `0, 0` and is byte-identical in wording and log line | | `Syncer.SetRenderPlanFn` + `renderSource` (v0.235.0) | controller/internal/sync/sync.go | `func(appName string) stacks.RenderPlan` | the catalog render table | **NIL-SAFE: no seam = copy verbatim = the old product.** Catalog images == pin → verbatim (fixes flow + self-healing, both deliberately kept); differ → the WHOLE stored definition, **never a ref substitution into a newer template** (`wger 2.6`). `.felhom.yml` always verbatim (R-458). The syncer must NEVER read app.yaml. Re-reads the applied file before writing it — a test caught it writing an empty compose over a live app | | `Stack.CatalogImages` vs `Stack.TemplateImages` (v0.235.0) | controller/internal/stacks/manager.go | both `map[string]string` | badge input vs "what the next `up -d` gives this app" | **THE TRAP: same type, same shape, opposite meaning after the freeze.** `TemplateImages` reads the LIVE (possibly frozen) compose file; `CatalogImages` reads the syncer's clone. `web.compareInstalledToTemplate` MUST use `CatalogImages` or it answers „Naprakész" on exactly the apps that are behind, with every test green. Red-proved | diff --git a/controller/README.md b/controller/README.md index 41abef3..187f115 100644 --- a/controller/README.md +++ b/controller/README.md @@ -594,7 +594,7 @@ the boot sweep) puts a pin back or marks an interrupted update for `ResumeInterr **Every unattended start path honours a hold:** the boot sweep and the app-stop guard (as before), and since v0.237.0 the drive-return gate and the nightly volume dump. The nightly capture and Tier-2 run -skip a held app so its restore point is not overwritten. +skip a held app so its restore point is not overwritten — and, since v0.238.1, an app whose update is still in progress (the periodic capture overwrote a primary unit during a health wait, found live). **The page (v0.238.0).** `Frissítés` follows the job — the button shows the phase label and the page reloads when the update ends. An updating card offers no lifecycle button; a held card shows the hold diff --git a/controller/cmd/controller/main.go b/controller/cmd/controller/main.go index d2cfd22..00b6850 100644 --- a/controller/cmd/controller/main.go +++ b/controller/cmd/controller/main.go @@ -554,6 +554,11 @@ func main() { // An UNWIRED manager also refuses (fail closed); TestSlice4_UpdateGuardsAreWiredAtStartup walks // this file for the call, because a seam built and never wired has shipped here seven times. stackMgr.SetUpdateGuards(&updateGuardsAdapter{b: backupMgr, q: quiesceLoop}) + // v0.238.1: the nightly legs (capture, Tier 2, volume dump) leave an app alone WHILE it is being + // updated, not only once it is held — found live in Scenario F, see backup.Manager.isHeld. + if backupMgr != nil { + backupMgr.SetUpdatingCheck(stackMgr.IsUpdating) + } if n := stackMgr.ResumeInterruptedUpdates(ctx); n > 0 { logger.Printf("[WARN] [update] resumed %d interrupted update(s)", n) } diff --git a/controller/cmd/controller/slice4_wiring_test.go b/controller/cmd/controller/slice4_wiring_test.go index f61192b..eb48d29 100644 --- a/controller/cmd/controller/slice4_wiring_test.go +++ b/controller/cmd/controller/slice4_wiring_test.go @@ -102,3 +102,12 @@ func TestSlice4_DriveStartGate_NamesAnUpdateHold(t *testing.T) { t.Errorf("ok=%v why=%q", ok, why) } } + +// v0.238.1: the backup manager must be told which apps are mid-update, or the nightly legs overwrite a +// restore point during an update's health wait (found live, Scenario F). +func TestSlice4_UpdatingCheckIsWiredAtStartup(t *testing.T) { + lines, _, _ := slice4CallLines(t) + if len(lines["SetUpdatingCheck"]) == 0 { + t.Fatal("backupMgr.SetUpdatingCheck is never called — the nightly legs cannot see an update in progress") + } +} diff --git a/controller/internal/backup/backup.go b/controller/internal/backup/backup.go index 2ca529d..8d54106 100644 --- a/controller/internal/backup/backup.go +++ b/controller/internal/backup/backup.go @@ -170,6 +170,9 @@ type Manager struct { // disconnected) can be unit-tested without Docker. Nil → the real DumpAppVolumesSafe. dumpVolumesSafe func(stackName string) error + // updatingCheck (slice 4) — nil-safe; see isHeld / SetUpdatingCheck. + updatingCheck func(stackName string) bool + // R-354 volume-REPLAY seam — the mirror of the F17 DB seams above, so the off-site path's new // volume leg is unit-testable without Docker. Nil → the real restoreDockerVolumesFrom. volumeReplayFrom func(stackName, dumpDir string) (int, error) diff --git a/controller/internal/backup/slice4_update_guard_test.go b/controller/internal/backup/slice4_update_guard_test.go index 4d586d3..7e765ec 100644 --- a/controller/internal/backup/slice4_update_guard_test.go +++ b/controller/internal/backup/slice4_update_guard_test.go @@ -166,3 +166,30 @@ func TestSlice4_NightlyLegsLeaveAHeldAppAlone(t *testing.T) { t.Errorf("positive control: the unheld app must still be mirrored, got %v", mirrored) } } + +// v0.238.1 — the gap Scenario F found live: during the update's health wait the app is not yet held, +// and the periodic capture wrote the never-started new definition into its primary unit. An app a +// guarded update is moving must be left alone by all three nightly legs, exactly like a held one. +// +// COMPANION RED-PROOF (REPORT.md): delete the updatingCheck clause from isHeld — the updating app is +// then dumped, captured and mirrored, and this test fails. +func TestSlice4_NightlyLegsLeaveAnAppMidUpdateAlone(t *testing.T) { + h := newAdmissionHarness(t, "updating", "free") + h.m.settings = slice4Settings(t) + h.m.SetUpdatingCheck(func(name string) bool { return name == "updating" }) + h.m.runVolumeDumps() + h.m.captureAllRecoveryUnits() + var mirrored []string + h.m.perAppTier2 = func(name string) error { mirrored = append(mirrored, name); return nil } + h.m.RunAllTier2() + for _, list := range [][]string{h.volDumped, h.prov.stopped, h.prov.infoHits, mirrored} { + for _, n := range list { + if n == "updating" { + t.Fatalf("a nightly leg touched an app MID-UPDATE (dumped=%v stopped=%v captured=%v mirrored=%v)", h.volDumped, h.prov.stopped, h.prov.infoHits, mirrored) + } + } + } + if len(mirrored) != 1 || mirrored[0] != "free" { + t.Errorf("positive control: the app not being updated must still be mirrored, got %v", mirrored) + } +} diff --git a/controller/internal/backup/update_guard.go b/controller/internal/backup/update_guard.go index 468703b..d1e7085 100644 --- a/controller/internal/backup/update_guard.go +++ b/controller/internal/backup/update_guard.go @@ -295,10 +295,29 @@ func (m *Manager) HoldAfterFailedUpdate(stackName string, at time.Time, copyDate return nil } -// isHeld reports whether an app carries ANY hold. Used by the nightly legs to leave a held app alone. +// isHeld reports whether the nightly legs must leave an app alone: it carries ANY hold, OR a guarded +// update is moving it right now. +// +// THE SECOND HALF WAS FOUND LIVE, v0.238.0 Scenario F on demo-hp 2026-09-13. During the update's +// 5-minute health wait the app is not yet held, and the periodic capture ran at 10:17:09 and wrote +// the NEW definition (alpine:3.20, which never started) into the app's PRIMARY unit, 53 s before the +// hold landed at 10:18:02. The Tier-2 mirror the hold names was intact only because the Tier-2 run is +// daily — a nightly Tier-2 falling inside a verify window would have mirrored the broken definition +// over the very copy the customer is told to restore from. An app mid-update has a restore point that +// must not move, exactly like a held one. func (m *Manager) isHeld(stackName string) bool { held, _ := m.RestoreHoldFor(stackName) - return held + if held { + return true + } + return m.updatingCheck != nil && m.updatingCheck(stackName) +} + +// SetUpdatingCheck wires the "is a guarded update moving this app" question (stacks.Manager.IsUpdating). +// INIT-ONLY, in main.go — pinned by TestSlice4_UpdatingCheckIsWiredAtStartup. The backup package cannot +// import stacks, which is why it is a seam. +func (m *Manager) SetUpdatingCheck(fn func(stackName string) bool) { + m.updatingCheck = fn } // clearUpdateHoldAfterRestore lifts an UPDATE hold once a person has restored the app successfully.