v0.132.0: the slow crash-loop counter (R-539, operator ruling 3 of 2026-09-16)
gates / gates (push) Successful in 12s
gates / gates (push) Successful in 12s
Beside the unchanged 3-in-15-minutes brake, a second counter: restarts the
supervisor performed in the last 24 hours. At the fifth the heartbeat stanza
sets slow_crashloop_since (moving at most once per 24 h), slow_crashloop and
restarts_24h; hub v0.117.0 mints controller_slow_crashloop (warning,
operator-only) when the timestamp moves. It never stops restarting.
Persisted per guest (tmp+rename, 0600) so an agent restart or reboot does not
reset it - unlike the fast record, whose reason for staying in memory (a
persisted give-up outliving the fix) does not apply to a counter that only
warns. Deliberate kills count. The startup line prints the new limits.
Red-proofs seen failing: no counter; the once-per-24h guard removed ('the
operator would be mailed per restart'); the save removed ('Restarts24h:1'
after an agent restart). Negative control: restarts 7 h apart never raise it.
go build/vet/test ./... green, 30 packages.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -1,3 +1,29 @@
|
||||
## Unreleased (to become v0.132.0) — a controller that dies slowly is reported, not just restarted (2026-09-17, R-539)
|
||||
|
||||
**MinAgent impact:** none required by any controller. Hub **v0.117.0** turns the new fields into
|
||||
`controller_slow_crashloop`; an older hub ignores them.
|
||||
|
||||
- **R-539 (operator ruling 3 of 2026-09-16) — the slow crash-loop counter.** The 3-restarts-in-15-minutes
|
||||
brake cannot see a controller that dies every 20 minutes (measured 2026-09-16, R-531: four restarts,
|
||||
none accumulating, only an `info` event that mails nobody). Beside it, unchanged, a second counter:
|
||||
restarts the supervisor performed in the last **24 hours**; at the **fifth**, the heartbeat's
|
||||
`controller_supervisor` stanza sets `slow_crashloop_since` (and `slow_crashloop: true`,
|
||||
`restarts_24h`). The hub mails on that timestamp MOVING; it moves **at most once per 24 hours**. It
|
||||
does **not** stop restarting — the fast brake remains the only brake.
|
||||
- **Persisted per guest** at `/var/lib/felhom-agent/guests/<vmid>/controller-restarts-24h.json`
|
||||
(tmp + rename, 0600), so an agent restart or a host reboot does not reset it. The fast record stays
|
||||
in memory; the reason it does (a persisted "give up" could outlive the fix) does not apply to a
|
||||
counter that only warns. Unreadable or corrupt → WARN and a clean start, never a blocked supervisor.
|
||||
- **Deliberate kills count.** The supervisor cannot tell an operator's `docker kill` from a crash
|
||||
(measured 2026-09-15); a controller killed five times a day is worth a line either way.
|
||||
- The supervisor's startup line now prints `slow_crashloop_max=5 slow_crashloop_window=24h0m0s`.
|
||||
|
||||
**Red-proofs, each seen failing:** five restarts 20 minutes apart raise it
|
||||
(`TestControllerSupervisor_SlowCrashloop` — fails with no counter; fails again with the once-per-24-hours
|
||||
guard removed, "the operator would be mailed per restart"); the counter survives an agent restart
|
||||
(`…SlowCounterSurvivesAgentRestart` — fails with the save removed, `Restarts24h:1`); restarts seven hours
|
||||
apart never raise it (the negative control). Wire shape extended with `restarts_24h` and `slow_crashloop`.
|
||||
|
||||
## v0.131.0 — a dead controller comes back by itself; the backup status speaks per tier (2026-09-15, R-523 / R-517 / R-518)
|
||||
|
||||
> **RELEASED 2026-09-15** by `scripts/release-agent.sh` — tag `v0.131.0`, sha256 `1118b552f7e775fbde9544c7764ede7e6046e0a7db16ae8494d07a18e3c2ac9c`. Delivered to boxes by the controller v0.243.0 floor (declared MinAgent), not by hand.
|
||||
|
||||
Reference in New Issue
Block a user