From 77cd70f7c0b6822cb40419e794066c14272807cd Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Thu, 17 Sep 2026 11:30:45 +0200 Subject: [PATCH] REPORT: agent v0.132.0 - slow crash loop, signed delivery, live proof with the production window Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- REPORT.md | 92 +++++++++++++++++++++++++++++++++---------------------- 1 file changed, 55 insertions(+), 37 deletions(-) diff --git a/REPORT.md b/REPORT.md index 6a11fe1..dd3e6c4 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,47 +1,65 @@ -# REPORT — agent v0.131.0: the controller comes back by itself; backup status per tier (2026-09-15) +# REPORT — agent v0.132.0: a controller that dies slowly is reported, not just restarted -Task: *before the volunteer — the big night's P1 fixes*, Parts A and C (agent half). Architecture: `03-host-agent.md` -§4 (the "healing a crashed controller" sentence), `07-backup-architecture.md` §6. +**2026-09-17.** R-539, operator ruling 3 of 2026-09-16. Architecture: `felhom.eu/documentation/architecture/03-host-agent.md` (the supervisor and its budget). -## Measured first (A.1) -Docker 29.8.0, throwaway containers on scratch 9202: after `docker kill`, **both** `--restart unless-stopped` and -`--restart always` stayed `exited (137)` 60 s later. The task's claim was right; a policy change is not a fix. +## Claims in the brief that turned out wrong — named first + +1. **„Emit `controller_slow_crashloop` from the agent."** The agent has no event channel. It carries + timestamps in its heartbeat stanza and the **hub** mints the event when one moves. So the build was + three wire fields here **plus** a hub checker change (hub v0.117.0), not an event registration alone. +2. **„Five kills 20 min apart is too long — make the interval configurable for the test."** Not needed, + and not done: the proof restart from B.4(a) was already persisted, so four more real kills ~8 minutes + apart reached five in 24 hours inside the session, with the **production** window. 8 minutes keeps any + 15-minute window at two restarts, so the fast brake never interferes. ## What shipped -- **Controller supervisor** (`internal/localapi/controllersupervisor.go`): every 30 s, for provisioned felhom-pool guests - that are running, restart `felhom-controller-bootstrap.service` on the second not-running observation. Guards: swap - in flight, host-side park marker, locked / vzdump-busy / stopped guest, unknown docker answer, unprovisioned guest, - 3 restarts in 15 min → 30 min pause. Record rides the host report as `controller_supervisor`; hub v0.114.0 mints the - events. Same GuestExecutor and sudoers grants as the swap — no new privilege. -- **Per-tier backup status**: `GET /backup/status` (untargeted) gains `tiers[]` — newest success (record, or storage - after a restart), last attempt kept apart, storage presence; `GET /backup/tiers` gains `storage`. -- **Golden script**: `--restart always`. No golden baked (R-468); existing boxes keep `unless-stopped`. -## Red-proofs (each seen failing, then restored) -- remove the restart call → `the killed controller was NOT restarted — this is R-523 (restarts=0)` -- remove the backoff block → `crash-looping controller restarted 10 times in 10 minutes — want exactly 3` -- `last_success` from the newest attempt → `pbs tier reports a failed attempt as its last success` +`internal/localapi/controllersupervisor.go`: beside the unchanged 3-in-15 brake, restarts the supervisor +performed in the last 24 hours. At the fifth, `slow_crashloop_since` moves (at most once per 24 h); +`slow_crashloop` and `restarts_24h` ride the stanza (`internal/hub/report.go`). Persisted per guest in +`/var/lib/felhom-agent/guests//controller-restarts-24h.json` (tmp + rename, 0600); unreadable or +corrupt → WARN and a clean start. Never stops restarting. Deliberate kills count. The startup line prints +the new limits. -## Release and delivery -`scripts/release-agent.sh 0.131.0`: tag `v0.131.0`, sha256 `1118b552f7e775fbde9544c7764ede7e6046e0a7db16ae8494d07a18e3c2ac9c`, -verified by download. **The task's "the floor delivers the agent" was wrong** (R-530): the hub holds a floor above the -box's agent; agents update only by an operator-signed job. On the operator's keys: `felhom-opsign -op agent_update` for -`demo-hp-bb76ea` only → authorized 08:44:16Z, committed 08:45:21Z, `controller-supervisor: started`. demo-felhom and -Peti's box stay on 0.130.0. +## Red-proofs — each seen failing, then passing -## Live validation on demo-hp guest 9201 (agent 0.131.0, controller 0.243.0) -| moment | result | -|---|---| -| idle kill 08:53:27Z | restarted 08:54:22Z; dashboard 200 **59 s** after the kill | -| parked + kill | stayed dead 100 s, `the guest is PARKED — leaving it` every sweep; unpark → 200 in **25 s** | -| kill 10 s into a swap | `during a controller SWAP — the swap owns it` ×3; the swap rolled back itself, healthy 09:02:06Z | -| kill 5 s into a deploy | **not measured**: the three test restarts had filled the budget, so the guard paused (as designed) and the hub mailed `controller_crashloop` | -| resume after the pause | pause held to 09:35:52Z; restarted 09:36:24Z; dashboard 200 at 09:36:30Z | -Two invalid attempts, both marked in the evidence: a swap POST sent over plain HTTP (400, no swap), and a "health 200" -line during the pause that the container state contradicts. +| test | break | failure seen | +|---|---|---| +| `TestControllerSupervisor_SlowCrashloop` | no counter | `five restarts 20 minutes apart did not raise slow_crashloop — this is R-539` | +| same | once-per-24h guard removed | `the raise moved again on the 6th restart … mailed per restart` | +| `…SlowCounterSurvivesAgentRestart` | save removed | `the agent restart reset the slow counter … Restarts24h:1` | + +Negative control: restarts 7 h apart never raise it. Wire shape extended (`restarts_24h`, `slow_crashloop`). + +## Gates and release + +`go build ./... && go vet ./... && go test ./...` — green, 30 packages. `agent_gates.py --fast` — all OK. +Released by `scripts/release-agent.sh`: tag `v0.132.0`, sha256 +`4afe815749a41b327ebe4a98a4557ad2acffb1fa71740a3a8b8003473b835321`, verified by independent download. +**Not vouched** for Day-0 installs (the operator's act). Code `18d03bd`, release record `CHANGELOG` commit after it. + +## Delivery — operator-signed, per box (ruling 1 of 2026-09-16) + +| box | signed | authorised → completed | committed | startup line | +|---|---|---|---|---| +| demo-hp (`demo-hp-bb76ea`) | 08:27:36Z | 08:29:25Z | 08:30:29Z | `slow_crashloop_max=5 slow_crashloop_window=24h0m0s` | +| N100 (`demo-felhom-8363b5`) | 08:27:36Z | 08:34:48Z | 08:35:52Z | same | + +Peti's box: not touched. Evidence: `felhom.eu/documentation/audits/evidence-chaos-fixes-2026-09-17/partC-agent-update.txt`. + +## Live validation — PASS, production window + +On demo-hp guest 9201: five controller kills (08:51:53Z from B.4(a), then 08:57, 09:05, 09:13, 09:21), +each restarted by the agent in 42–56 s, never by hand; the fast brake never armed. At #5: +`SLOW CRASH-LOOP — … restarts_24h=5 window=24h0m0s threshold=5`; persisted file holds the five times and the +raise. The hub minted `controller_slow_crashloop` at 09:29:40Z and **exactly one** operator mail arrived. +Evidence: `…/partC-live-slow-crashloop.txt`. ## Teardown -Machine: 9201 back on its controller (see resume); the throwaway homebox deploy left no container. Host: park marker -removed; nothing else changed. Hub: demo-hp floor override 0.243.0 kept (it delivers this release). -Evidence: `felhom.eu/documentation/audits/evidence-p1fixes-2026-09-15/A*`. +Nothing provisioned. Stated effect on the demo box: 9201's counter holds five restarts and stays raised +until 09:22Z on 2026-09-18; it cannot mail again inside that window. + +## Observations + +1. The persisted file writes the zero raise time as `"0001-01-01T00:00:00Z"` (Go `omitempty` does not omit a zero `time.Time`). **NOT-A-FINDING: the loader reads it back as zero (`IsZero`), which the live file and the restart test both exercised; cosmetic only.**