v0.132.0: the slow crash-loop counter (R-539, operator ruling 3 of 2026-09-16)
gates / gates (push) Successful in 12s

Beside the unchanged 3-in-15-minutes brake, a second counter: restarts the
supervisor performed in the last 24 hours. At the fifth the heartbeat stanza
sets slow_crashloop_since (moving at most once per 24 h), slow_crashloop and
restarts_24h; hub v0.117.0 mints controller_slow_crashloop (warning,
operator-only) when the timestamp moves. It never stops restarting.

Persisted per guest (tmp+rename, 0600) so an agent restart or reboot does not
reset it - unlike the fast record, whose reason for staying in memory (a
persisted give-up outliving the fix) does not apply to a counter that only
warns. Deliberate kills count. The startup line prints the new limits.

Red-proofs seen failing: no counter; the once-per-24h guard removed ('the
operator would be mailed per restart'); the save removed ('Restarts24h:1'
after an agent restart). Negative control: restarts 7 h apart never raise it.
go build/vet/test ./... green, 30 packages.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-17 10:26:42 +02:00
parent e98b857684
commit 18d03bd437
4 changed files with 229 additions and 3 deletions
+26
View File
@@ -1,3 +1,29 @@
## Unreleased (to become v0.132.0) — a controller that dies slowly is reported, not just restarted (2026-09-17, R-539)
**MinAgent impact:** none required by any controller. Hub **v0.117.0** turns the new fields into
`controller_slow_crashloop`; an older hub ignores them.
- **R-539 (operator ruling 3 of 2026-09-16) — the slow crash-loop counter.** The 3-restarts-in-15-minutes
brake cannot see a controller that dies every 20 minutes (measured 2026-09-16, R-531: four restarts,
none accumulating, only an `info` event that mails nobody). Beside it, unchanged, a second counter:
restarts the supervisor performed in the last **24 hours**; at the **fifth**, the heartbeat's
`controller_supervisor` stanza sets `slow_crashloop_since` (and `slow_crashloop: true`,
`restarts_24h`). The hub mails on that timestamp MOVING; it moves **at most once per 24 hours**. It
does **not** stop restarting — the fast brake remains the only brake.
- **Persisted per guest** at `/var/lib/felhom-agent/guests/<vmid>/controller-restarts-24h.json`
(tmp + rename, 0600), so an agent restart or a host reboot does not reset it. The fast record stays
in memory; the reason it does (a persisted "give up" could outlive the fix) does not apply to a
counter that only warns. Unreadable or corrupt → WARN and a clean start, never a blocked supervisor.
- **Deliberate kills count.** The supervisor cannot tell an operator's `docker kill` from a crash
(measured 2026-09-15); a controller killed five times a day is worth a line either way.
- The supervisor's startup line now prints `slow_crashloop_max=5 slow_crashloop_window=24h0m0s`.
**Red-proofs, each seen failing:** five restarts 20 minutes apart raise it
(`TestControllerSupervisor_SlowCrashloop` — fails with no counter; fails again with the once-per-24-hours
guard removed, "the operator would be mailed per restart"); the counter survives an agent restart
(`…SlowCounterSurvivesAgentRestart` — fails with the save removed, `Restarts24h:1`); restarts seven hours
apart never raise it (the negative control). Wire shape extended with `restarts_24h` and `slow_crashloop`.
## v0.131.0 — a dead controller comes back by itself; the backup status speaks per tier (2026-09-15, R-523 / R-517 / R-518)
> **RELEASED 2026-09-15** by `scripts/release-agent.sh` — tag `v0.131.0`, sha256 `1118b552f7e775fbde9544c7764ede7e6046e0a7db16ae8494d07a18e3c2ac9c`. Delivered to boxes by the controller v0.243.0 floor (declared MinAgent), not by hand.