Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Pg8ANF97SEeKYSN5Jxw3qJ
3.2 KiB
REPORT — F1: controller-swap verify hardening, v0.47.0
Task: close F1 from the no-mercy testrun — a controller image with no HEALTHCHECK that
crash-loops could land a single "Running" inspect poll → the swap marked it healthy → no
rollback (alpine tagged as the controller passed in ~4 s, then Restarting (0)).
Baseline: agent main @ 9b0d6c2 (live 0.46.0) → v0.47.0 @ 3844df7, sha 2d5a3ce4….
The fix (internal/localapi/controllerswap.go)
controllerHealthynow reads a 4thdocker inspect -ffield{{.RestartCount}}:running && RestartCount>0→ not-ok (a process that already crash-restarted isn't stably up, regardless of healthcheck). It also returns aneedsDwellsignal for the no-healthcheck (none) case. It stays a pure point-in-time predicate.verifyadds a stability dwell: a no-healthcheck image must report ok onverifyDwell(=3) consecutive polls before acceptance; a realhealthyresult is trusted immediately (Docker already gated it). Any not-ok resets the dwell. Timeout → the existing rollback path runs.- No change to
writeImage, the sudoers grants, or the state-file/rollback orchestration.
§3 sudoers — no change (confirmed live)
The grant is pct exec [0-9]* -- docker inspect -f *; * spans the extended template. Live
sudo -n -l of the new …|{{.RestartCount}} template → exit 0 (matches). No new grant.
Tests (green: go build/vet/test ./...)
- F1 red-proof (
TestF1_Verify_CrashLoopRestartCountBlocks):RestartCount>0→verifyfalse +controllerHealthyreturns(false, starting). Companion: the same image withrc=0+ dwell=1 does verify — proving the RestartCount check is what blocks the crasher (the oldrunning + none → okshape had no guard). - Dwell (
…DwellSingleOkThenCrash): a single ok then a crash →verifyfalse; companion dwell=1 accepts the single ok. - Real healthcheck (
…RealHealthcheckPassesPromptly):healthy→verifytrue promptly, no false rollback. ExistingRollbackOnUnhealthy/HealthyWithNoHealthcheckstay green.
Live re-test (felhom-pve guest 9201) — the previously-found defect is now FIXED
Re-ran the exact F1 scenario: tagged alpine as felhom-controller:9.9.9 (no healthcheck, crash-loops),
swapped to it via the real agent /controller/swap:
swap state: failed error: "new controller did not become healthy within timeout"
controller-swap: rolled back to previous, controller healthy previous=…:0.91.0
running controller: …:0.91.0 Up (healthy)
Before the fix this image was marked done in ~4 s with no rollback; now the verify (polling
the …|{{.RestartCount}} template) never accepts it → 90 s timeout → rollback to 0.91.0. The
rollback's verify of the real 0.91.0 (which has a healthcheck) passed promptly — the real-controller
path is unaffected. Bad tag removed; final state: agent 0.47.0, controller 0.91.0 healthy, capabilities
45/45, channel up.
NOT changed
writeImage, the FELHOM_CONTROLLERSWAP grants, the state-file/rollback orchestration, the
capability/drive-gate/channel logic. Verify predicate + dwell only.