R-271: the channel recovery after a controller restart is reported
The agent_channel_unauthorized alert's remedy is a re-bootstrap (a restart), and an unseeded->up first observation was silent, so following the instruction guaranteed no recovery event. A down alert now leaves a marker in the data dir; the first UP after a restart sends the recovery and clears it. A restart with no alert outstanding stays silent. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -2002,7 +2002,8 @@ func main() {
|
||||
// launches the job immediately since the scheduler is already started.
|
||||
if cfg.LocalAPI.Endpoint != "" {
|
||||
chSink := channelSink{notifier: notifier, alertMgr: alertMgr}
|
||||
chChecker := channelhealth.New(webServer.ProbeAgentChannel, chSink, logger)
|
||||
chChecker := channelhealth.New(webServer.ProbeAgentChannel, chSink, logger).
|
||||
WithAlertMarker(filepath.Join(cfg.Paths.DataDir, "channelhealth-down-alerted")) // R-271
|
||||
sched.Every("agent-channel-health", 60*time.Second, chChecker.Check)
|
||||
}
|
||||
|
||||
|
||||
Reference in New Issue
Block a user