v0.262.0: six defects two drill nights found in the update, remove and hold paths
gates / gates (push) Successful in 26s
gates / gates (push) Successful in 26s
R-630 (P1): waitUpdateHealthy kept the probe inside `if hc != nil && len(hc.Checks) > 0`, and when findProbeContainer returned "" its else set last="no probe container" and LOOPED - the settle path sat in the outer else, unreachable. So verifying could only time out and failAndHold then stopped a working app. Measured on paperless-ngx: three containers healthy, failed at +313.0s, front door 404 after. It now falls through to the same settle path with a WARN naming the candidates. The probe target is decidable now: HealthCheckConfig.Container plus findProbeContainerMeta resolve by exact stack name -> explicit container -> a UNIQUE prefix -> nothing with the candidates returned. The old rule took the FIRST prefix match. A skipped stack records why instead of silence. R-634 (half): RemoveStack refused on the !Deployed FLAG while the machine had containers, a compose file and an app.yaml. It now asks whether anything EXISTS. The mechanism producing the bad record is still not diagnosed and R-634 stays open for it. R-633/R-626: RemoveStack consults UpdateGuards.Busy and IsUpdating and refuses with the app's own sentence - the product already refused this clash for update and for restore. And because `down` returning 0 is a request not a result, the project is watched for 25s afterwards, anything carrying its label is removed by name with its labels logged, and the answer carries `verified`. R-621: failAndHold writes compose logs --tail 400 into <stackdir>/hold-logs/<ts>/ BEFORE the down that destroys them. Two existing tests pin the compose sequence and correctly caught the new step; their expectations are updated with the reason that the ORDER is the assertion. R-614: RemoveStack calls ClearUpdateState. NOT in this release: R-625 (a held app still renders an Update button). Named, not half-done. Three new sentences, each born as a key in both bundles. Four red-proofs seen failing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -1,3 +1,71 @@
|
||||
## v0.262.0 — six defects two drill nights found in the update, remove and hold paths (2026-09-22, R-630/R-634/R-633/R-626/R-621/R-614)
|
||||
|
||||
**MinAgent: 0.131.0** (unchanged). **Two** new sentences, each BORN AS A KEY in both bundles and
|
||||
registered in `i18n_go_keys.json`: `health.no_probe_container` and
|
||||
`err.stacks.az_alkalmazason_mentes_vagy_visszaallitas_fut`. A third, `stacks.held_restore_needed`,
|
||||
was written and then **removed** — it belongs to R-625, which is not in this release, and an unused
|
||||
bundle key is a promise that was not kept. The go-parity gate caught both new keys as UNLISTED
|
||||
before the push and was right to.
|
||||
|
||||
**The updates themselves were never the problem.** Two nights walked all 53 apps; what broke was the
|
||||
machinery around them.
|
||||
|
||||
### R-630 — an app with NO health probe was STOPPED by a successful update (P1)
|
||||
|
||||
`waitUpdateHealthy` kept the probe inside `if hc != nil && len(hc.Checks) > 0`, and when
|
||||
`findProbeContainer` returned `""` its `else` set `last = "no probe container"` and **looped** — the
|
||||
settle path that judges an app with no declared check sat in the outer `else`, unreachable. So
|
||||
`verifying` could only ever time out, and `failAndHold` then stopped a working app. Measured on
|
||||
paperless-ngx: all three containers `healthy`, `failed` at **+313.0 s**, front door 404 afterwards,
|
||||
and the controller's own words — `not healthy within 5m0s (last: no probe container)`.
|
||||
|
||||
It now falls through to the same settle path with a WARN naming the candidates, and the journal says
|
||||
`no probe container — settled on container state` so it is distinguishable from `no health check
|
||||
declared` without reading the log. **A stack with no probe is not healthy and not failing — it is
|
||||
settled on container state (`09` §3), and never a reason to stop a running app.**
|
||||
|
||||
**The target is decidable now.** `HealthCheckConfig.Container` (precedent
|
||||
`InitialCredentials.Container`) and `findProbeContainerMeta` resolve by **exact stack name →
|
||||
explicit `container` → a UNIQUE prefix → nothing, with the candidates returned**. The old rule took
|
||||
the FIRST prefix match: for `immich` — four `immich-*` containers, no exact match — that is whichever
|
||||
the container list yields, and it was seen live on `outline` probing `outline-postgres:3000` during
|
||||
startup. A skipped stack now also records WHY instead of going silent.
|
||||
|
||||
### R-634 — an app running and serving while recorded as not deployed, and then unremovable (P1, half)
|
||||
|
||||
`RemoveStack` refused on `!stack.Deployed` — a **flag** — while the machine had containers, a compose
|
||||
file and an `app.yaml`. Three apps were measured in that state, one of them answering its own
|
||||
`/_health` with 200 on three containers, and both remove calls said `stack "x" is not deployed`. The
|
||||
only exit was a shell, and a household has none. The refusal now asks whether anything **exists**.
|
||||
**The mechanism that produces the bad record is still not diagnosed** and R-634 stays open for it.
|
||||
|
||||
### R-633 / R-626 — remove during a restore reported success and left a ghost
|
||||
|
||||
A remove sent while a restore was in flight tore down what existed; the restore's own `compose up`
|
||||
re-created it seventeen seconds later. Both calls returned success, the record read `deployed:
|
||||
false`, and a container restarted for hours with a live public route. `RemoveStack` now consults
|
||||
`UpdateGuards.Busy` and `IsUpdating` and refuses with the app's own sentence — the product already
|
||||
refused this clash for `update` and for `restore`, and `remove` was the one door without a lock. And
|
||||
because `down` returning 0 is a request rather than a result, the project is watched for 25 s
|
||||
afterwards, anything carrying its label is removed by name with its labels logged, and the answer
|
||||
carries `verified`.
|
||||
|
||||
### R-621 — a hold destroyed the evidence of why
|
||||
|
||||
`failAndHold` now writes `compose logs --no-color --tail 400` into `<stackdir>/hold-logs/<ts>/`
|
||||
**before** the `down`. Best-effort by design: a hold must never fail because its evidence could not
|
||||
be written.
|
||||
|
||||
### R-614 — a removed app left its update phase behind
|
||||
|
||||
`RemoveStack` calls `ClearUpdateState`. The name is the only thing a new install shares with the old
|
||||
one, so the record goes when the app does.
|
||||
|
||||
### Not in this release
|
||||
|
||||
**R-625** (a held app still renders an Update button that then refuses) is **not fixed here** — named
|
||||
in the report rather than half-done.
|
||||
|
||||
## v0.261.0 — the controller no longer swaps itself out from under an app update (2026-09-21, R-608/R-609)
|
||||
|
||||
**MinAgent: 0.131.0** (unchanged). Hungarian pages byte-identical; the two new sentences are BORN AS
|
||||
|
||||
Reference in New Issue
Block a user