Files
felhom-agent/REPORT.md
T

3.9 KiB
Raw Blame History

REPORT — agent v0.132.0: a controller that dies slowly is reported, not just restarted

2026-09-17. R-539, operator ruling 3 of 2026-09-16. Architecture: felhom.eu/documentation/architecture/03-host-agent.md (the supervisor and its budget).

Claims in the brief that turned out wrong — named first

  1. „Emit controller_slow_crashloop from the agent." The agent has no event channel. It carries timestamps in its heartbeat stanza and the hub mints the event when one moves. So the build was three wire fields here plus a hub checker change (hub v0.117.0), not an event registration alone.
  2. „Five kills 20 min apart is too long — make the interval configurable for the test." Not needed, and not done: the proof restart from B.4(a) was already persisted, so four more real kills ~8 minutes apart reached five in 24 hours inside the session, with the production window. 8 minutes keeps any 15-minute window at two restarts, so the fast brake never interferes.

What shipped

internal/localapi/controllersupervisor.go: beside the unchanged 3-in-15 brake, restarts the supervisor performed in the last 24 hours. At the fifth, slow_crashloop_since moves (at most once per 24 h); slow_crashloop and restarts_24h ride the stanza (internal/hub/report.go). Persisted per guest in /var/lib/felhom-agent/guests/<vmid>/controller-restarts-24h.json (tmp + rename, 0600); unreadable or corrupt → WARN and a clean start. Never stops restarting. Deliberate kills count. The startup line prints the new limits.

Red-proofs — each seen failing, then passing

test break failure seen
TestControllerSupervisor_SlowCrashloop no counter five restarts 20 minutes apart did not raise slow_crashloop — this is R-539
same once-per-24h guard removed the raise moved again on the 6th restart … mailed per restart
…SlowCounterSurvivesAgentRestart save removed the agent restart reset the slow counter … Restarts24h:1

Negative control: restarts 7 h apart never raise it. Wire shape extended (restarts_24h, slow_crashloop).

Gates and release

go build ./... && go vet ./... && go test ./... — green, 30 packages. agent_gates.py --fast — all OK. Released by scripts/release-agent.sh: tag v0.132.0, sha256 4afe815749a41b327ebe4a98a4557ad2acffb1fa71740a3a8b8003473b835321, verified by independent download. Not vouched for Day-0 installs (the operator's act). Code 18d03bd, release record CHANGELOG commit after it.

Delivery — operator-signed, per box (ruling 1 of 2026-09-16)

box signed authorised → completed committed startup line
demo-hp (demo-hp-bb76ea) 08:27:36Z 08:29:25Z 08:30:29Z slow_crashloop_max=5 slow_crashloop_window=24h0m0s
N100 (demo-felhom-8363b5) 08:27:36Z 08:34:48Z 08:35:52Z same

Peti's box: not touched. Evidence: felhom.eu/documentation/audits/evidence-chaos-fixes-2026-09-17/partC-agent-update.txt.

Live validation — PASS, production window

On demo-hp guest 9201: five controller kills (08:51:53Z from B.4(a), then 08:57, 09:05, 09:13, 09:21), each restarted by the agent in 42–56 s, never by hand; the fast brake never armed. At #5: SLOW CRASH-LOOP — … restarts_24h=5 window=24h0m0s threshold=5; persisted file holds the five times and the raise. The hub minted controller_slow_crashloop at 09:29:40Z and exactly one operator mail arrived. Evidence: …/partC-live-slow-crashloop.txt.

Teardown

Nothing provisioned. Stated effect on the demo box: 9201's counter holds five restarts and stays raised until 09:22Z on 2026-09-18; it cannot mail again inside that window.

Observations

  1. The persisted file writes the zero raise time as "0001-01-01T00:00:00Z" (Go omitempty does not omit a zero time.Time). NOT-A-FINDING: the loader reads it back as zero (IsZero), which the live file and the restart test both exercised; cosmetic only.