v0.262.0: six defects two drill nights found in the update, remove and hold paths
gates / gates (push) Successful in 26s
gates / gates (push) Successful in 26s
R-630 (P1): waitUpdateHealthy kept the probe inside `if hc != nil && len(hc.Checks) > 0`, and when findProbeContainer returned "" its else set last="no probe container" and LOOPED - the settle path sat in the outer else, unreachable. So verifying could only time out and failAndHold then stopped a working app. Measured on paperless-ngx: three containers healthy, failed at +313.0s, front door 404 after. It now falls through to the same settle path with a WARN naming the candidates. The probe target is decidable now: HealthCheckConfig.Container plus findProbeContainerMeta resolve by exact stack name -> explicit container -> a UNIQUE prefix -> nothing with the candidates returned. The old rule took the FIRST prefix match. A skipped stack records why instead of silence. R-634 (half): RemoveStack refused on the !Deployed FLAG while the machine had containers, a compose file and an app.yaml. It now asks whether anything EXISTS. The mechanism producing the bad record is still not diagnosed and R-634 stays open for it. R-633/R-626: RemoveStack consults UpdateGuards.Busy and IsUpdating and refuses with the app's own sentence - the product already refused this clash for update and for restore. And because `down` returning 0 is a request not a result, the project is watched for 25s afterwards, anything carrying its label is removed by name with its labels logged, and the answer carries `verified`. R-621: failAndHold writes compose logs --tail 400 into <stackdir>/hold-logs/<ts>/ BEFORE the down that destroys them. Two existing tests pin the compose sequence and correctly caught the new step; their expectations are updated with the reason that the ORDER is the assertion. R-614: RemoveStack calls ClearUpdateState. NOT in this release: R-625 (a held app still renders an Update button). Named, not half-done. Three new sentences, each born as a key in both bundles. Four red-proofs seen failing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -433,8 +433,11 @@ func TestSlice4_F_HealthFailureHoldsTheAppAndKeepsTheNewPin(t *testing.T) {
|
||||
if got := pinOf(t, dir); got != "nextcloud:34.0.1-apache" {
|
||||
t.Errorf("the pin must STAY on the new version (its migration may have run), got %q", got)
|
||||
}
|
||||
if got := strings.Join(c.list(), " | "); got != "pull | up -d --remove-orphans | down" {
|
||||
t.Errorf("the failed app must be stopped; compose calls = %q", got)
|
||||
// The `logs` call between `up` and `down` is R-621: the hold keeps the app's own log BEFORE the
|
||||
// `down` destroys it. The order is the assertion — a capture after the `down` would read empty,
|
||||
// which is exactly how two drill nights lost the only evidence of why an update failed.
|
||||
if got := strings.Join(c.list(), " | "); got != "pull | up -d --remove-orphans | logs --no-color --tail 400 | down" {
|
||||
t.Errorf("the failed app must be stopped, and its log kept FIRST; compose calls = %q", got)
|
||||
}
|
||||
if journalExists(m) {
|
||||
t.Error("the journal is cleared once the hold (the durable record) is written")
|
||||
@@ -537,8 +540,9 @@ func TestSlice4_G_InterruptedAfterUpResumesTheHealthWait(t *testing.T) {
|
||||
if held, _ := g.HoldFor("nextcloud"); !held || st.UpdatePhase != UpdatePhaseFailed {
|
||||
t.Errorf("a resumed update that is still unhealthy must end HELD; held=%v phase=%q", held, st.UpdatePhase)
|
||||
}
|
||||
if got := strings.Join(c.list(), " | "); got != "up -d --remove-orphans | down" {
|
||||
t.Errorf("resumption re-runs `up` then stops the failed app; compose calls = %q", got)
|
||||
// Same R-621 capture on the resumed path — a hold reached by resumption keeps its evidence too.
|
||||
if got := strings.Join(c.list(), " | "); got != "up -d --remove-orphans | logs --no-color --tail 400 | down" {
|
||||
t.Errorf("resumption re-runs `up`, keeps the log, then stops the failed app; compose calls = %q", got)
|
||||
}
|
||||
if !g.holdRP.ProvenAt.Equal(slice4T0.Add(-time.Hour)) || g.holdRP.Tier != UpdateTierLocal {
|
||||
t.Errorf("the resumed hold must name the journaled copy (tier %d at %s), got %+v", UpdateTierLocal, slice4T0.Add(-time.Hour), g.holdRP)
|
||||
|
||||
Reference in New Issue
Block a user