v0.262.0: six defects two drill nights found in the update, remove and hold paths
gates / gates (push) Successful in 26s

R-630 (P1): waitUpdateHealthy kept the probe inside `if hc != nil && len(hc.Checks) > 0`, and when
findProbeContainer returned "" its else set last="no probe container" and LOOPED - the settle path
sat in the outer else, unreachable. So verifying could only time out and failAndHold then stopped a
working app. Measured on paperless-ngx: three containers healthy, failed at +313.0s, front door 404
after. It now falls through to the same settle path with a WARN naming the candidates.

The probe target is decidable now: HealthCheckConfig.Container plus findProbeContainerMeta resolve
by exact stack name -> explicit container -> a UNIQUE prefix -> nothing with the candidates
returned. The old rule took the FIRST prefix match. A skipped stack records why instead of silence.

R-634 (half): RemoveStack refused on the !Deployed FLAG while the machine had containers, a compose
file and an app.yaml. It now asks whether anything EXISTS. The mechanism producing the bad record is
still not diagnosed and R-634 stays open for it.

R-633/R-626: RemoveStack consults UpdateGuards.Busy and IsUpdating and refuses with the app's own
sentence - the product already refused this clash for update and for restore. And because `down`
returning 0 is a request not a result, the project is watched for 25s afterwards, anything carrying
its label is removed by name with its labels logged, and the answer carries `verified`.

R-621: failAndHold writes compose logs --tail 400 into <stackdir>/hold-logs/<ts>/ BEFORE the down
that destroys them. Two existing tests pin the compose sequence and correctly caught the new step;
their expectations are updated with the reason that the ORDER is the assertion.

R-614: RemoveStack calls ClearUpdateState.

NOT in this release: R-625 (a held app still renders an Update button). Named, not half-done.

Three new sentences, each born as a key in both bundles. Four red-proofs seen failing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-22 20:45:19 +02:00
parent 93cee16843
commit b793638484
12 changed files with 4946 additions and 4404 deletions
+68
View File
@@ -1,3 +1,71 @@
## v0.262.0 — six defects two drill nights found in the update, remove and hold paths (2026-09-22, R-630/R-634/R-633/R-626/R-621/R-614)
**MinAgent: 0.131.0** (unchanged). **Two** new sentences, each BORN AS A KEY in both bundles and
registered in `i18n_go_keys.json`: `health.no_probe_container` and
`err.stacks.az_alkalmazason_mentes_vagy_visszaallitas_fut`. A third, `stacks.held_restore_needed`,
was written and then **removed** — it belongs to R-625, which is not in this release, and an unused
bundle key is a promise that was not kept. The go-parity gate caught both new keys as UNLISTED
before the push and was right to.
**The updates themselves were never the problem.** Two nights walked all 53 apps; what broke was the
machinery around them.
### R-630 — an app with NO health probe was STOPPED by a successful update (P1)
`waitUpdateHealthy` kept the probe inside `if hc != nil && len(hc.Checks) > 0`, and when
`findProbeContainer` returned `""` its `else` set `last = "no probe container"` and **looped** — the
settle path that judges an app with no declared check sat in the outer `else`, unreachable. So
`verifying` could only ever time out, and `failAndHold` then stopped a working app. Measured on
paperless-ngx: all three containers `healthy`, `failed` at **+313.0 s**, front door 404 afterwards,
and the controller's own words — `not healthy within 5m0s (last: no probe container)`.
It now falls through to the same settle path with a WARN naming the candidates, and the journal says
`no probe container — settled on container state` so it is distinguishable from `no health check
declared` without reading the log. **A stack with no probe is not healthy and not failing — it is
settled on container state (`09` §3), and never a reason to stop a running app.**
**The target is decidable now.** `HealthCheckConfig.Container` (precedent
`InitialCredentials.Container`) and `findProbeContainerMeta` resolve by **exact stack name →
explicit `container` → a UNIQUE prefix → nothing, with the candidates returned**. The old rule took
the FIRST prefix match: for `immich` — four `immich-*` containers, no exact match — that is whichever
the container list yields, and it was seen live on `outline` probing `outline-postgres:3000` during
startup. A skipped stack now also records WHY instead of going silent.
### R-634 — an app running and serving while recorded as not deployed, and then unremovable (P1, half)
`RemoveStack` refused on `!stack.Deployed` — a **flag** — while the machine had containers, a compose
file and an `app.yaml`. Three apps were measured in that state, one of them answering its own
`/_health` with 200 on three containers, and both remove calls said `stack "x" is not deployed`. The
only exit was a shell, and a household has none. The refusal now asks whether anything **exists**.
**The mechanism that produces the bad record is still not diagnosed** and R-634 stays open for it.
### R-633 / R-626 — remove during a restore reported success and left a ghost
A remove sent while a restore was in flight tore down what existed; the restore's own `compose up`
re-created it seventeen seconds later. Both calls returned success, the record read `deployed:
false`, and a container restarted for hours with a live public route. `RemoveStack` now consults
`UpdateGuards.Busy` and `IsUpdating` and refuses with the app's own sentence — the product already
refused this clash for `update` and for `restore`, and `remove` was the one door without a lock. And
because `down` returning 0 is a request rather than a result, the project is watched for 25 s
afterwards, anything carrying its label is removed by name with its labels logged, and the answer
carries `verified`.
### R-621 — a hold destroyed the evidence of why
`failAndHold` now writes `compose logs --no-color --tail 400` into `<stackdir>/hold-logs/<ts>/`
**before** the `down`. Best-effort by design: a hold must never fail because its evidence could not
be written.
### R-614 — a removed app left its update phase behind
`RemoveStack` calls `ClearUpdateState`. The name is the only thing a new install shares with the old
one, so the record goes when the app does.
### Not in this release
**R-625** (a held app still renders an Update button that then refuses) is **not fixed here** — named
in the report rather than half-done.
## v0.261.0 — the controller no longer swaps itself out from under an app update (2026-09-21, R-608/R-609)
**MinAgent: 0.131.0** (unchanged). Hungarian pages byte-identical; the two new sentences are BORN AS