Files
felhom-agent/REPORT.md
T
2026-09-15 11:37:07 +02:00

3.5 KiB
Raw Blame History

REPORT — agent v0.131.0: the controller comes back by itself; backup status per tier (2026-09-15)

Task: before the volunteer — the big night's P1 fixes, Parts A and C (agent half). Architecture: 03-host-agent.md §4 (the "healing a crashed controller" sentence), 07-backup-architecture.md §6.

Measured first (A.1)

Docker 29.8.0, throwaway containers on scratch 9202: after docker kill, both --restart unless-stopped and --restart always stayed exited (137) 60 s later. The task's claim was right; a policy change is not a fix.

What shipped

  • Controller supervisor (internal/localapi/controllersupervisor.go): every 30 s, for provisioned felhom-pool guests that are running, restart felhom-controller-bootstrap.service on the second not-running observation. Guards: swap in flight, host-side park marker, locked / vzdump-busy / stopped guest, unknown docker answer, unprovisioned guest, 3 restarts in 15 min → 30 min pause. Record rides the host report as controller_supervisor; hub v0.114.0 mints the events. Same GuestExecutor and sudoers grants as the swap — no new privilege.
  • Per-tier backup status: GET /backup/status (untargeted) gains tiers[] — newest success (record, or storage after a restart), last attempt kept apart, storage presence; GET /backup/tiers gains storage.
  • Golden script: --restart always. No golden baked (R-468); existing boxes keep unless-stopped.

Red-proofs (each seen failing, then restored)

  • remove the restart call → the killed controller was NOT restarted — this is R-523 (restarts=0)
  • remove the backoff block → crash-looping controller restarted 10 times in 10 minutes — want exactly 3
  • last_success from the newest attempt → pbs tier reports a failed attempt as its last success

Release and delivery

scripts/release-agent.sh 0.131.0: tag v0.131.0, sha256 1118b552f7e775fbde9544c7764ede7e6046e0a7db16ae8494d07a18e3c2ac9c, verified by download. The task's "the floor delivers the agent" was wrong (R-530): the hub holds a floor above the box's agent; agents update only by an operator-signed job. On the operator's keys: felhom-opsign -op agent_update for demo-hp-bb76ea only → authorized 08:44:16Z, committed 08:45:21Z, controller-supervisor: started. demo-felhom and Peti's box stay on 0.130.0.

Live validation on demo-hp guest 9201 (agent 0.131.0, controller 0.243.0)

moment result
idle kill 08:53:27Z restarted 08:54:22Z; dashboard 200 59 s after the kill
parked + kill stayed dead 100 s, the guest is PARKED — leaving it every sweep; unpark → 200 in 25 s
kill 10 s into a swap during a controller SWAP — the swap owns it ×3; the swap rolled back itself, healthy 09:02:06Z
kill 5 s into a deploy not measured: the three test restarts had filled the budget, so the guard paused (as designed) and the hub mailed controller_crashloop
resume after the pause pause held to 09:35:52Z; restarted 09:36:24Z; dashboard 200 at 09:36:30Z
Two invalid attempts, both marked in the evidence: a swap POST sent over plain HTTP (400, no swap), and a "health 200"
line during the pause that the container state contradicts.

Teardown

Machine: 9201 back on its controller (see resume); the throwaway homebox deploy left no container. Host: park marker removed; nothing else changed. Hub: demo-hp floor override 0.243.0 kept (it delivers this release).

Evidence: felhom.eu/documentation/audits/evidence-p1fixes-2026-09-15/A*.