Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
3.9 KiB
REPORT — agent v0.132.0: a controller that dies slowly is reported, not just restarted
2026-09-17. R-539, operator ruling 3 of 2026-09-16. Architecture: felhom.eu/documentation/architecture/03-host-agent.md (the supervisor and its budget).
Claims in the brief that turned out wrong — named first
- „Emit
controller_slow_crashloopfrom the agent." The agent has no event channel. It carries timestamps in its heartbeat stanza and the hub mints the event when one moves. So the build was three wire fields here plus a hub checker change (hub v0.117.0), not an event registration alone. - „Five kills 20 min apart is too long — make the interval configurable for the test." Not needed, and not done: the proof restart from B.4(a) was already persisted, so four more real kills ~8 minutes apart reached five in 24 hours inside the session, with the production window. 8 minutes keeps any 15-minute window at two restarts, so the fast brake never interferes.
What shipped
internal/localapi/controllersupervisor.go: beside the unchanged 3-in-15 brake, restarts the supervisor
performed in the last 24 hours. At the fifth, slow_crashloop_since moves (at most once per 24 h);
slow_crashloop and restarts_24h ride the stanza (internal/hub/report.go). Persisted per guest in
/var/lib/felhom-agent/guests/<vmid>/controller-restarts-24h.json (tmp + rename, 0600); unreadable or
corrupt → WARN and a clean start. Never stops restarting. Deliberate kills count. The startup line prints
the new limits.
Red-proofs — each seen failing, then passing
| test | break | failure seen |
|---|---|---|
TestControllerSupervisor_SlowCrashloop |
no counter | five restarts 20 minutes apart did not raise slow_crashloop — this is R-539 |
| same | once-per-24h guard removed | the raise moved again on the 6th restart … mailed per restart |
…SlowCounterSurvivesAgentRestart |
save removed | the agent restart reset the slow counter … Restarts24h:1 |
Negative control: restarts 7 h apart never raise it. Wire shape extended (restarts_24h, slow_crashloop).
Gates and release
go build ./... && go vet ./... && go test ./... — green, 30 packages. agent_gates.py --fast — all OK.
Released by scripts/release-agent.sh: tag v0.132.0, sha256
4afe815749a41b327ebe4a98a4557ad2acffb1fa71740a3a8b8003473b835321, verified by independent download.
Not vouched for Day-0 installs (the operator's act). Code 18d03bd, release record CHANGELOG commit after it.
Delivery — operator-signed, per box (ruling 1 of 2026-09-16)
| box | signed | authorised → completed | committed | startup line |
|---|---|---|---|---|
demo-hp (demo-hp-bb76ea) |
08:27:36Z | 08:29:25Z | 08:30:29Z | slow_crashloop_max=5 slow_crashloop_window=24h0m0s |
N100 (demo-felhom-8363b5) |
08:27:36Z | 08:34:48Z | 08:35:52Z | same |
Peti's box: not touched. Evidence: felhom.eu/documentation/audits/evidence-chaos-fixes-2026-09-17/partC-agent-update.txt.
Live validation — PASS, production window
On demo-hp guest 9201: five controller kills (08:51:53Z from B.4(a), then 08:57, 09:05, 09:13, 09:21),
each restarted by the agent in 42–56 s, never by hand; the fast brake never armed. At #5:
SLOW CRASH-LOOP — … restarts_24h=5 window=24h0m0s threshold=5; persisted file holds the five times and the
raise. The hub minted controller_slow_crashloop at 09:29:40Z and exactly one operator mail arrived.
Evidence: …/partC-live-slow-crashloop.txt.
Teardown
Nothing provisioned. Stated effect on the demo box: 9201's counter holds five restarts and stays raised until 09:22Z on 2026-09-18; it cannot mail again inside that window.
Observations
- The persisted file writes the zero raise time as
"0001-01-01T00:00:00Z"(Goomitemptydoes not omit a zerotime.Time). NOT-A-FINDING: the loader reads it back as zero (IsZero), which the live file and the restart test both exercised; cosmetic only.