v0.238.1: the nightly backup leaves an app alone WHILE it is being updated, not only once it is held (slice 4 follow-up)
gates / gates (push) Successful in 13s
gates / gates (push) Successful in 13s
Found live in v0.238.0 Scenario F on demo-hp: during an update's 5-minute health wait the app is not yet held, and the periodic recovery-unit capture at 10:17:09 wrote the never-started definition (alpine:3.20) into its PRIMARY unit, 53 s before the hold landed. The Tier-2 mirror the hold names survived only because Tier 2 runs daily; a nightly Tier 2 inside a verify window would have mirrored the broken definition over the copy the customer is told to restore from. backup.Manager.isHeld — consulted by the capture sweep, the Tier-2 run and the volume dump — is now also true while a guarded update is moving the app, via SetUpdatingCheck wired in main.go to stacks.Manager.IsUpdating. Test with positive control + red-proof; wiring pinned. Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -1,3 +1,29 @@
|
||||
## v0.238.1 — the nightly backup leaves an app alone WHILE it is being updated, not only once it is held (2026-09-13, slice 4 follow-up)
|
||||
|
||||
**MinAgent: 0.129.0** (unchanged)
|
||||
|
||||
**Found live, not in review.** v0.238.0 Scenario F on demo-hp: an update to `alpine:3.20` (an image
|
||||
that exits at once) waited its 5-minute health timeout and was held at 10:18:02 — correctly. But the
|
||||
periodic recovery-unit capture ran at **10:17:09**, inside that wait, when the app was updating and
|
||||
not yet held, and wrote the never-started definition into the app's **primary** unit:
|
||||
|
||||
```
|
||||
primary-unit: image: alpine:3.20 manifest "created_at": "2026-09-13T10:17:09Z"
|
||||
tier2-mirror: image: louislam/uptime-kuma:2.4.0 manifest "created_at": "2026-09-13T10:09:51Z"
|
||||
```
|
||||
|
||||
The Tier-2 mirror the hold text names survived only because the Tier-2 run is daily. A nightly Tier-2
|
||||
falling inside a verify window would have mirrored the broken definition over the very copy the
|
||||
customer is told to restore from.
|
||||
|
||||
**Fix.** `backup.Manager.isHeld` — the predicate the capture sweep, the Tier-2 run and the volume dump
|
||||
already consult since v0.237.0 — is also true while a guarded update is moving the app, through a new
|
||||
`SetUpdatingCheck` seam wired in `main.go` to `stacks.Manager.IsUpdating`.
|
||||
|
||||
**Tests.** `TestSlice4_NightlyLegsLeaveAnAppMidUpdateAlone` (all three legs skip the updating app; the
|
||||
other app is still mirrored — positive control), red-proofed by removing the clause (the app is then
|
||||
dumped, captured and mirrored); `TestSlice4_UpdatingCheckIsWiredAtStartup`.
|
||||
|
||||
## v0.238.0 — the page follows the update, and a held app offers no way to start it (2026-09-13, update arc slice 4 Part 4)
|
||||
|
||||
**MinAgent: 0.129.0** (unchanged)
|
||||
|
||||
@@ -124,7 +124,7 @@
|
||||
| `Manager.UpdatePreflight` / `StartGuardedUpdate` / `RecoverUpdates` / `ResumeInterruptedUpdates` (v0.237.0) | controller/internal/stacks/update.go | `UpdatePreflight(name) *UpdateRefusal`; `StartGuardedUpdate(name) error` | THE update — refusals, then a 202 job with phases on `Stack.Updating/UpdatePhase/UpdateError` | **Never report an update complete before health is known (R-443).** Every cheap refusal runs BEFORE the intent write. Safety dump BEFORE the pin moves; pin BEFORE pull; pull failure → pin back; health failure → stop + HOLD, pin stays. Journal-before-mutate (`update-journal.json`); `RecoverUpdates` MUST run before the boot sweep and `ResumeInterruptedUpdates` AFTER `SetUpdateGuards`. Seams: `updateComposeFn`, `updateHealthFn`, `updateMemoryFn`, `updateDiskFreeFn`, `updateNowFn` (R-457: the age check and the test read ONE clock). Unwired guards ⇒ every update refused |
|
||||
| `stacks.UpdateGuards` + `updateGuardsAdapter` (v0.237.0) | controller/internal/stacks/update.go, controller/cmd/controller/main.go | `HoldFor`, `Busy`, `RestorePoint`, `BackupNow`, `SafetyDump`, `HoldAfterFailedUpdate` | the ONLY bridge from the update job to the backup side (stacks cannot import backup) | Wired by `stackMgr.SetUpdateGuards` — pinned by `TestSlice4_UpdateGuardsAreWiredAtStartup`. Add a guard HERE, never by importing backup into stacks |
|
||||
| `backup.Manager.Tier2UnitRestorePoint` + `Tier2RestorePoint.ProvenCopyTime` (v0.237.0) | controller/internal/backup/update_guard.go | `(stack) (Tier2RestorePoint, error)` | "can this app be restored from Tier 2, and from when" — the predicate that gates BOTH the „Teljes visszaállítás" action and an update | **ONE predicate, two callers** (extracted from `buildAppBackupRows`, not copied). `CopyDate` is what the page NAMES (the package date, R-403); `ProvenCopyTime` is how OLD the data is — the last successful copy, because the manifest's `created_at` moves only when the DEFINITION changes (measured: a fresh dump under a 22-h-older manifest). Do not age a copy by `CopyDate` |
|
||||
| `backup.Manager.HoldAfterFailedUpdate` / `RunAppBackupNow` / `WriteUpdateSafetyDump` / `UpdateBusy` (v0.237.0) | controller/internal/backup/update_guard.go | see file | the update's hold, per-app backup-now, safety dump, busy check | The hold is `settings.RestoreHold` with `Reason: update_failed` — SAME store and gate as R-379, never a second map. `RunAppBackupNow` composes the nightly legs for ONE app (admission, DB dump, volume dump, capture, Tier-2) — do not write a second backup orchestration. A successful unit restore lifts an UPDATE hold only |
|
||||
| `backup.Manager.HoldAfterFailedUpdate` / `RunAppBackupNow` / `WriteUpdateSafetyDump` / `UpdateBusy` (v0.237.0) | controller/internal/backup/update_guard.go | see file | the update's hold, per-app backup-now, safety dump, busy check | The hold is `settings.RestoreHold` with `Reason: update_failed` — SAME store and gate as R-379, never a second map. `RunAppBackupNow` composes the nightly legs for ONE app (admission, DB dump, volume dump, capture, Tier-2) — do not write a second backup orchestration. A successful unit restore lifts an UPDATE hold only. **`isHeld` is ALSO true while a guarded update is moving the app (`SetUpdatingCheck`, v0.238.1)** — found live: the periodic capture overwrote a primary unit during a health wait |
|
||||
| `stacks.Manager.memoryVerdict` (v0.237.0) | controller/internal/stacks/deploy.go | `(newReq, newLimit, releasedReq, releasedLimit int) (refusal, warning string)` | the deploy's memory check, shared with the update | An update RELEASES the app's current request first. Deploy passes `0, 0` and is byte-identical in wording and log line |
|
||||
| `Syncer.SetRenderPlanFn` + `renderSource` (v0.235.0) | controller/internal/sync/sync.go | `func(appName string) stacks.RenderPlan` | the catalog render table | **NIL-SAFE: no seam = copy verbatim = the old product.** Catalog images == pin → verbatim (fixes flow + self-healing, both deliberately kept); differ → the WHOLE stored definition, **never a ref substitution into a newer template** (`wger 2.6`). `.felhom.yml` always verbatim (R-458). The syncer must NEVER read app.yaml. Re-reads the applied file before writing it — a test caught it writing an empty compose over a live app |
|
||||
| `Stack.CatalogImages` vs `Stack.TemplateImages` (v0.235.0) | controller/internal/stacks/manager.go | both `map[string]string` | badge input vs "what the next `up -d` gives this app" | **THE TRAP: same type, same shape, opposite meaning after the freeze.** `TemplateImages` reads the LIVE (possibly frozen) compose file; `CatalogImages` reads the syncer's clone. `web.compareInstalledToTemplate` MUST use `CatalogImages` or it answers „Naprakész" on exactly the apps that are behind, with every test green. Red-proved |
|
||||
|
||||
@@ -594,7 +594,7 @@ the boot sweep) puts a pin back or marks an interrupted update for `ResumeInterr
|
||||
|
||||
**Every unattended start path honours a hold:** the boot sweep and the app-stop guard (as before), and
|
||||
since v0.237.0 the drive-return gate and the nightly volume dump. The nightly capture and Tier-2 run
|
||||
skip a held app so its restore point is not overwritten.
|
||||
skip a held app so its restore point is not overwritten — and, since v0.238.1, an app whose update is still in progress (the periodic capture overwrote a primary unit during a health wait, found live).
|
||||
|
||||
**The page (v0.238.0).** `Frissítés` follows the job — the button shows the phase label and the page
|
||||
reloads when the update ends. An updating card offers no lifecycle button; a held card shows the hold
|
||||
|
||||
@@ -554,6 +554,11 @@ func main() {
|
||||
// An UNWIRED manager also refuses (fail closed); TestSlice4_UpdateGuardsAreWiredAtStartup walks
|
||||
// this file for the call, because a seam built and never wired has shipped here seven times.
|
||||
stackMgr.SetUpdateGuards(&updateGuardsAdapter{b: backupMgr, q: quiesceLoop})
|
||||
// v0.238.1: the nightly legs (capture, Tier 2, volume dump) leave an app alone WHILE it is being
|
||||
// updated, not only once it is held — found live in Scenario F, see backup.Manager.isHeld.
|
||||
if backupMgr != nil {
|
||||
backupMgr.SetUpdatingCheck(stackMgr.IsUpdating)
|
||||
}
|
||||
if n := stackMgr.ResumeInterruptedUpdates(ctx); n > 0 {
|
||||
logger.Printf("[WARN] [update] resumed %d interrupted update(s)", n)
|
||||
}
|
||||
|
||||
@@ -102,3 +102,12 @@ func TestSlice4_DriveStartGate_NamesAnUpdateHold(t *testing.T) {
|
||||
t.Errorf("ok=%v why=%q", ok, why)
|
||||
}
|
||||
}
|
||||
|
||||
// v0.238.1: the backup manager must be told which apps are mid-update, or the nightly legs overwrite a
|
||||
// restore point during an update's health wait (found live, Scenario F).
|
||||
func TestSlice4_UpdatingCheckIsWiredAtStartup(t *testing.T) {
|
||||
lines, _, _ := slice4CallLines(t)
|
||||
if len(lines["SetUpdatingCheck"]) == 0 {
|
||||
t.Fatal("backupMgr.SetUpdatingCheck is never called — the nightly legs cannot see an update in progress")
|
||||
}
|
||||
}
|
||||
|
||||
@@ -170,6 +170,9 @@ type Manager struct {
|
||||
// disconnected) can be unit-tested without Docker. Nil → the real DumpAppVolumesSafe.
|
||||
dumpVolumesSafe func(stackName string) error
|
||||
|
||||
// updatingCheck (slice 4) — nil-safe; see isHeld / SetUpdatingCheck.
|
||||
updatingCheck func(stackName string) bool
|
||||
|
||||
// R-354 volume-REPLAY seam — the mirror of the F17 DB seams above, so the off-site path's new
|
||||
// volume leg is unit-testable without Docker. Nil → the real restoreDockerVolumesFrom.
|
||||
volumeReplayFrom func(stackName, dumpDir string) (int, error)
|
||||
|
||||
@@ -166,3 +166,30 @@ func TestSlice4_NightlyLegsLeaveAHeldAppAlone(t *testing.T) {
|
||||
t.Errorf("positive control: the unheld app must still be mirrored, got %v", mirrored)
|
||||
}
|
||||
}
|
||||
|
||||
// v0.238.1 — the gap Scenario F found live: during the update's health wait the app is not yet held,
|
||||
// and the periodic capture wrote the never-started new definition into its primary unit. An app a
|
||||
// guarded update is moving must be left alone by all three nightly legs, exactly like a held one.
|
||||
//
|
||||
// COMPANION RED-PROOF (REPORT.md): delete the updatingCheck clause from isHeld — the updating app is
|
||||
// then dumped, captured and mirrored, and this test fails.
|
||||
func TestSlice4_NightlyLegsLeaveAnAppMidUpdateAlone(t *testing.T) {
|
||||
h := newAdmissionHarness(t, "updating", "free")
|
||||
h.m.settings = slice4Settings(t)
|
||||
h.m.SetUpdatingCheck(func(name string) bool { return name == "updating" })
|
||||
h.m.runVolumeDumps()
|
||||
h.m.captureAllRecoveryUnits()
|
||||
var mirrored []string
|
||||
h.m.perAppTier2 = func(name string) error { mirrored = append(mirrored, name); return nil }
|
||||
h.m.RunAllTier2()
|
||||
for _, list := range [][]string{h.volDumped, h.prov.stopped, h.prov.infoHits, mirrored} {
|
||||
for _, n := range list {
|
||||
if n == "updating" {
|
||||
t.Fatalf("a nightly leg touched an app MID-UPDATE (dumped=%v stopped=%v captured=%v mirrored=%v)", h.volDumped, h.prov.stopped, h.prov.infoHits, mirrored)
|
||||
}
|
||||
}
|
||||
}
|
||||
if len(mirrored) != 1 || mirrored[0] != "free" {
|
||||
t.Errorf("positive control: the app not being updated must still be mirrored, got %v", mirrored)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -295,10 +295,29 @@ func (m *Manager) HoldAfterFailedUpdate(stackName string, at time.Time, copyDate
|
||||
return nil
|
||||
}
|
||||
|
||||
// isHeld reports whether an app carries ANY hold. Used by the nightly legs to leave a held app alone.
|
||||
// isHeld reports whether the nightly legs must leave an app alone: it carries ANY hold, OR a guarded
|
||||
// update is moving it right now.
|
||||
//
|
||||
// THE SECOND HALF WAS FOUND LIVE, v0.238.0 Scenario F on demo-hp 2026-09-13. During the update's
|
||||
// 5-minute health wait the app is not yet held, and the periodic capture ran at 10:17:09 and wrote
|
||||
// the NEW definition (alpine:3.20, which never started) into the app's PRIMARY unit, 53 s before the
|
||||
// hold landed at 10:18:02. The Tier-2 mirror the hold names was intact only because the Tier-2 run is
|
||||
// daily — a nightly Tier-2 falling inside a verify window would have mirrored the broken definition
|
||||
// over the very copy the customer is told to restore from. An app mid-update has a restore point that
|
||||
// must not move, exactly like a held one.
|
||||
func (m *Manager) isHeld(stackName string) bool {
|
||||
held, _ := m.RestoreHoldFor(stackName)
|
||||
return held
|
||||
if held {
|
||||
return true
|
||||
}
|
||||
return m.updatingCheck != nil && m.updatingCheck(stackName)
|
||||
}
|
||||
|
||||
// SetUpdatingCheck wires the "is a guarded update moving this app" question (stacks.Manager.IsUpdating).
|
||||
// INIT-ONLY, in main.go — pinned by TestSlice4_UpdatingCheckIsWiredAtStartup. The backup package cannot
|
||||
// import stacks, which is why it is a seam.
|
||||
func (m *Manager) SetUpdatingCheck(fn func(stackName string) bool) {
|
||||
m.updatingCheck = fn
|
||||
}
|
||||
|
||||
// clearUpdateHoldAfterRestore lifts an UPDATE hold once a person has restored the app successfully.
|
||||
|
||||
Reference in New Issue
Block a user