Files
felhom-agent/REPORT.md
T

3.2 KiB

REPORT — F1: controller-swap verify hardening, v0.47.0

Task: close F1 from the no-mercy testrun — a controller image with no HEALTHCHECK that crash-loops could land a single "Running" inspect poll → the swap marked it healthy → no rollback (alpine tagged as the controller passed in ~4 s, then Restarting (0)).

Baseline: agent main @ 9b0d6c2 (live 0.46.0) → v0.47.0 @ 3844df7, sha 2d5a3ce4….

The fix (internal/localapi/controllerswap.go)

  • controllerHealthy now reads a 4th docker inspect -f field {{.RestartCount}}: running && RestartCount>0 → not-ok (a process that already crash-restarted isn't stably up, regardless of healthcheck). It also returns a needsDwell signal for the no-healthcheck (none) case. It stays a pure point-in-time predicate.
  • verify adds a stability dwell: a no-healthcheck image must report ok on verifyDwell (=3) consecutive polls before acceptance; a real healthy result is trusted immediately (Docker already gated it). Any not-ok resets the dwell. Timeout → the existing rollback path runs.
  • No change to writeImage, the sudoers grants, or the state-file/rollback orchestration.

§3 sudoers — no change (confirmed live)

The grant is pct exec [0-9]* -- docker inspect -f *; * spans the extended template. Live sudo -n -l of the new …|{{.RestartCount}} template → exit 0 (matches). No new grant.

Tests (green: go build/vet/test ./...)

  • F1 red-proof (TestF1_Verify_CrashLoopRestartCountBlocks): RestartCount>0verify false + controllerHealthy returns (false, starting). Companion: the same image with rc=0 + dwell=1 does verify — proving the RestartCount check is what blocks the crasher (the old running + none → ok shape had no guard).
  • Dwell (…DwellSingleOkThenCrash): a single ok then a crash → verify false; companion dwell=1 accepts the single ok.
  • Real healthcheck (…RealHealthcheckPassesPromptly): healthyverify true promptly, no false rollback. Existing RollbackOnUnhealthy / HealthyWithNoHealthcheck stay green.

Live re-test (felhom-pve guest 9201) — the previously-found defect is now FIXED

Re-ran the exact F1 scenario: tagged alpine as felhom-controller:9.9.9 (no healthcheck, crash-loops), swapped to it via the real agent /controller/swap:

swap state: failed   error: "new controller did not become healthy within timeout"
controller-swap: rolled back to previous, controller healthy   previous=…:0.91.0
running controller: …:0.91.0 Up (healthy)

Before the fix this image was marked done in ~4 s with no rollback; now the verify (polling the …|{{.RestartCount}} template) never accepts it → 90 s timeout → rollback to 0.91.0. The rollback's verify of the real 0.91.0 (which has a healthcheck) passed promptly — the real-controller path is unaffected. Bad tag removed; final state: agent 0.47.0, controller 0.91.0 healthy, capabilities 45/45, channel up.

NOT changed

writeImage, the FELHOM_CONTROLLERSWAP grants, the state-file/rollback orchestration, the capability/drive-gate/channel logic. Verify predicate + dwell only.