v0.262.0: six defects two drill nights found in the update, remove and hold paths
gates / gates (push) Successful in 26s

R-630 (P1): waitUpdateHealthy kept the probe inside `if hc != nil && len(hc.Checks) > 0`, and when
findProbeContainer returned "" its else set last="no probe container" and LOOPED - the settle path
sat in the outer else, unreachable. So verifying could only time out and failAndHold then stopped a
working app. Measured on paperless-ngx: three containers healthy, failed at +313.0s, front door 404
after. It now falls through to the same settle path with a WARN naming the candidates.

The probe target is decidable now: HealthCheckConfig.Container plus findProbeContainerMeta resolve
by exact stack name -> explicit container -> a UNIQUE prefix -> nothing with the candidates
returned. The old rule took the FIRST prefix match. A skipped stack records why instead of silence.

R-634 (half): RemoveStack refused on the !Deployed FLAG while the machine had containers, a compose
file and an app.yaml. It now asks whether anything EXISTS. The mechanism producing the bad record is
still not diagnosed and R-634 stays open for it.

R-633/R-626: RemoveStack consults UpdateGuards.Busy and IsUpdating and refuses with the app's own
sentence - the product already refused this clash for update and for restore. And because `down`
returning 0 is a request not a result, the project is watched for 25s afterwards, anything carrying
its label is removed by name with its labels logged, and the answer carries `verified`.

R-621: failAndHold writes compose logs --tail 400 into <stackdir>/hold-logs/<ts>/ BEFORE the down
that destroys them. Two existing tests pin the compose sequence and correctly caught the new step;
their expectations are updated with the reason that the ORDER is the assertion.

R-614: RemoveStack calls ClearUpdateState.

NOT in this release: R-625 (a held app still renders an Update button). Named, not half-done.

Three new sentences, each born as a key in both bundles. Four red-proofs seen failing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-22 20:45:19 +02:00
parent 93cee16843
commit b793638484
12 changed files with 4946 additions and 4404 deletions
+16
View File
@@ -898,6 +898,22 @@ Three probe types are supported:
- **`api`** — HTTP request with response validation (expected status code, body content). Fails if expectations aren't met.
- **`tcp`** — Simple port reachability check via `net.Dial`.
**Which container is probed (v0.262.0, R-630).** Four rules, in order: the container whose name
EQUALS the stack name; then `healthcheck.container` from `.felhom.yml`; then a prefix match, but
**only when exactly one running container matches**; else nothing — and the candidates are logged.
The third rule used to take the FIRST prefix match, which for an app like `immich` (four `immich-*`
containers, no exact match) meant whichever the container list happened to yield. And an app like
`paperless-ngx`, whose containers are `paperless-webserver`/`-postgres`/`-redis`, matched nothing at
all and was **skipped silently** — its probe had never run on any box.
**A stack whose check resolves to no container is not "healthy" and not "failing".** It is judged the
way an app that declares no check is judged: every container running and none restarting for the
settle window. That matters most during an update — `verifying` waits on this same probe, and before
v0.262.0 a stack with no probe target could only ever time out, so a SUCCESSFUL update ended with
`failAndHold` stopping a working app. Such a stack also now records a probe RESULT saying why no
check ran, instead of nothing.
Multiple checks per app are supported (all must pass). The probe scheduler runs every 10 seconds; per-app intervals default to 5 minutes and are configurable via `healthcheck.interval` in `.felhom.yml`. Probe results are stored in `Stack.HealthProbe` and exposed via the API. Failed probes override the stack state to `StateUnhealthy`; the override clears automatically when the next probe passes.
**Fast initial probing:** On start/restart, stale health probe results are cleared (so the stack doesn't immediately appear "unhealthy" from a previous result). Until the first healthy probe, the controller checks every 10 seconds instead of the normal 5-minute interval, giving fast feedback on whether the app came up successfully.