v0.262.0: six defects two drill nights found in the update, remove and hold paths
gates / gates (push) Successful in 26s
gates / gates (push) Successful in 26s
R-630 (P1): waitUpdateHealthy kept the probe inside `if hc != nil && len(hc.Checks) > 0`, and when findProbeContainer returned "" its else set last="no probe container" and LOOPED - the settle path sat in the outer else, unreachable. So verifying could only time out and failAndHold then stopped a working app. Measured on paperless-ngx: three containers healthy, failed at +313.0s, front door 404 after. It now falls through to the same settle path with a WARN naming the candidates. The probe target is decidable now: HealthCheckConfig.Container plus findProbeContainerMeta resolve by exact stack name -> explicit container -> a UNIQUE prefix -> nothing with the candidates returned. The old rule took the FIRST prefix match. A skipped stack records why instead of silence. R-634 (half): RemoveStack refused on the !Deployed FLAG while the machine had containers, a compose file and an app.yaml. It now asks whether anything EXISTS. The mechanism producing the bad record is still not diagnosed and R-634 stays open for it. R-633/R-626: RemoveStack consults UpdateGuards.Busy and IsUpdating and refuses with the app's own sentence - the product already refused this clash for update and for restore. And because `down` returning 0 is a request not a result, the project is watched for 25s afterwards, anything carrying its label is removed by name with its labels logged, and the answer carries `verified`. R-621: failAndHold writes compose logs --tail 400 into <stackdir>/hold-logs/<ts>/ BEFORE the down that destroys them. Two existing tests pin the compose sequence and correctly caught the new step; their expectations are updated with the reason that the ORDER is the assertion. R-614: RemoveStack calls ClearUpdateState. NOT in this release: R-625 (a held app still renders an Update button). Named, not half-done. Three new sentences, each born as a key in both bundles. Four red-proofs seen failing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -898,6 +898,22 @@ Three probe types are supported:
|
||||
- **`api`** — HTTP request with response validation (expected status code, body content). Fails if expectations aren't met.
|
||||
- **`tcp`** — Simple port reachability check via `net.Dial`.
|
||||
|
||||
**Which container is probed (v0.262.0, R-630).** Four rules, in order: the container whose name
|
||||
EQUALS the stack name; then `healthcheck.container` from `.felhom.yml`; then a prefix match, but
|
||||
**only when exactly one running container matches**; else nothing — and the candidates are logged.
|
||||
|
||||
The third rule used to take the FIRST prefix match, which for an app like `immich` (four `immich-*`
|
||||
containers, no exact match) meant whichever the container list happened to yield. And an app like
|
||||
`paperless-ngx`, whose containers are `paperless-webserver`/`-postgres`/`-redis`, matched nothing at
|
||||
all and was **skipped silently** — its probe had never run on any box.
|
||||
|
||||
**A stack whose check resolves to no container is not "healthy" and not "failing".** It is judged the
|
||||
way an app that declares no check is judged: every container running and none restarting for the
|
||||
settle window. That matters most during an update — `verifying` waits on this same probe, and before
|
||||
v0.262.0 a stack with no probe target could only ever time out, so a SUCCESSFUL update ended with
|
||||
`failAndHold` stopping a working app. Such a stack also now records a probe RESULT saying why no
|
||||
check ran, instead of nothing.
|
||||
|
||||
Multiple checks per app are supported (all must pass). The probe scheduler runs every 10 seconds; per-app intervals default to 5 minutes and are configurable via `healthcheck.interval` in `.felhom.yml`. Probe results are stored in `Stack.HealthProbe` and exposed via the API. Failed probes override the stack state to `StateUnhealthy`; the override clears automatically when the next probe passes.
|
||||
|
||||
**Fast initial probing:** On start/restart, stale health probe results are cleared (so the stack doesn't immediately appear "unhealthy" from a previous result). Until the first healthy probe, the controller checks every 10 seconds instead of the normal 5-minute interval, giving fast feedback on whether the app came up successfully.
|
||||
|
||||
Reference in New Issue
Block a user