R-271: the channel recovery after a controller restart is reported

The agent_channel_unauthorized alert's remedy is a re-bootstrap (a restart), and an unseeded->up
first observation was silent, so following the instruction guaranteed no recovery event. A down
alert now leaves a marker in the data dir; the first UP after a restart sends the recovery and
clears it. A restart with no alert outstanding stays silent.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-05 21:52:05 +02:00
parent 4867ec1804
commit f885100d29
3 changed files with 117 additions and 2 deletions
+2 -1
View File
@@ -2002,7 +2002,8 @@ func main() {
// launches the job immediately since the scheduler is already started.
if cfg.LocalAPI.Endpoint != "" {
chSink := channelSink{notifier: notifier, alertMgr: alertMgr}
chChecker := channelhealth.New(webServer.ProbeAgentChannel, chSink, logger)
chChecker := channelhealth.New(webServer.ProbeAgentChannel, chSink, logger).
WithAlertMarker(filepath.Join(cfg.Paths.DataDir, "channelhealth-down-alerted")) // R-271
sched.Every("agent-channel-health", 60*time.Second, chChecker.Check)
}