Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Pg8ANF97SEeKYSN5Jxw3qJ
2.8 KiB
REPORT — F2: alert on born/persistent-down channel, v0.91.0
Task: close F2 (the transition-only alerting gap the full-stack testrun surfaced) — a channel failure present at startup/reseed was dashboard-only, no operator email, forever. Part of the F2 (Part A) + prevention (Part B) slice; this repo is Part A (controller).
Baseline: controller main @ 77bccf1 (live 0.90.0) → v0.91.0 @ b2ad871 (image
…/felhom-controller:0.91.0). Companions: hub v0.21.0 (F2 hub checkers), agent v0.46.0 (B.1).
The gap + fix (internal/channelhealth/checker.go)
Pre-fix Check(): a confirmed (past-debounce) down with prev=="" hit // first observation seeds, no alert, then prev==newState every cycle → never alerted. A controller that boots into a
broken channel (e.g. right after a leaf regen + controller restart) was dashboard-only.
Fix — an alerted flag (have we emitted for the CURRENT down-spell?):
- UP:
alerted=false(re-arm); ifprevwasdown:*→NotifyRecovered. - DOWN, transient +
consecutiveDown < N: unchanged debounce (suppress; unseeded → assume up). - DOWN, confirmed: dashboard down. If
prevwas up/unseeded or the reason changed → new spell →alerted=false. Thenif !alerted → NotifyDown; alerted=true. Removed theprev==""silent-seed-for-down branch. - Net: non-transient born-down alerts on cycle 1; transient born-down still needs N≥2 (cold-boot agent-not-yet-up race); healthy first-obs still silent; reason-change re-alerts; recovery re-arms.
Tests (all green: go build/vet/test ./...)
F2 born-down red-proof (TestF2_BornDownNonTransient_AlertsOnce: one alert on cycle 1) + companion
(…OldSeedSilentLogicWouldNotAlert: the old prev=="" path would not have alerted); born-down
transient still debounced; recovery re-arms the spell; healthy first-obs silent. The existing
debounce/transition/classify tests still pass.
Live validation (felhom-pve guest 9201) — PASS
Induced a born-down non-transient: regenerated the agent leaf (channel breaks) then restarted the
controller so its FIRST channel observation was pin_mismatch:
[channel] agent channel DOWN (unseeded->down:pin_mismatch): … TLS pin mismatch …
Event pushed: agent_channel_pin_mismatch (error) — … (HTTP 200)
unseeded->down = the born-down case, and it ALERTED (~60 s, one cycle) — vs the pre-fix
silent seed. Recovered via R1 leaf restore: [channel] agent channel recovered +
agent_channel_recovered, leaf==pin MATCH, dashboard banner cleared. This is the capstone's missing
half — detection at startup, not only on a live transition.
NOT changed
The swap, capability, or drive-gate logic; the debounce threshold; the agentapi pin/transport. Alerting behaviour only.