Files
felhom-agent/REPORT.md
T
2026-09-15 11:37:07 +02:00

48 lines
3.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# REPORT — agent v0.131.0: the controller comes back by itself; backup status per tier (2026-09-15)
Task: *before the volunteer — the big night's P1 fixes*, Parts A and C (agent half). Architecture: `03-host-agent.md`
§4 (the "healing a crashed controller" sentence), `07-backup-architecture.md` §6.
## Measured first (A.1)
Docker 29.8.0, throwaway containers on scratch 9202: after `docker kill`, **both** `--restart unless-stopped` and
`--restart always` stayed `exited (137)` 60 s later. The task's claim was right; a policy change is not a fix.
## What shipped
- **Controller supervisor** (`internal/localapi/controllersupervisor.go`): every 30 s, for provisioned felhom-pool guests
that are running, restart `felhom-controller-bootstrap.service` on the second not-running observation. Guards: swap
in flight, host-side park marker, locked / vzdump-busy / stopped guest, unknown docker answer, unprovisioned guest,
3 restarts in 15 min → 30 min pause. Record rides the host report as `controller_supervisor`; hub v0.114.0 mints the
events. Same GuestExecutor and sudoers grants as the swap — no new privilege.
- **Per-tier backup status**: `GET /backup/status` (untargeted) gains `tiers[]` — newest success (record, or storage
after a restart), last attempt kept apart, storage presence; `GET /backup/tiers` gains `storage`.
- **Golden script**: `--restart always`. No golden baked (R-468); existing boxes keep `unless-stopped`.
## Red-proofs (each seen failing, then restored)
- remove the restart call → `the killed controller was NOT restarted — this is R-523 (restarts=0)`
- remove the backoff block → `crash-looping controller restarted 10 times in 10 minutes — want exactly 3`
- `last_success` from the newest attempt → `pbs tier reports a failed attempt as its last success`
## Release and delivery
`scripts/release-agent.sh 0.131.0`: tag `v0.131.0`, sha256 `1118b552f7e775fbde9544c7764ede7e6046e0a7db16ae8494d07a18e3c2ac9c`,
verified by download. **The task's "the floor delivers the agent" was wrong** (R-530): the hub holds a floor above the
box's agent; agents update only by an operator-signed job. On the operator's keys: `felhom-opsign -op agent_update` for
`demo-hp-bb76ea` only → authorized 08:44:16Z, committed 08:45:21Z, `controller-supervisor: started`. demo-felhom and
Peti's box stay on 0.130.0.
## Live validation on demo-hp guest 9201 (agent 0.131.0, controller 0.243.0)
| moment | result |
|---|---|
| idle kill 08:53:27Z | restarted 08:54:22Z; dashboard 200 **59 s** after the kill |
| parked + kill | stayed dead 100 s, `the guest is PARKED — leaving it` every sweep; unpark → 200 in **25 s** |
| kill 10 s into a swap | `during a controller SWAP — the swap owns it` ×3; the swap rolled back itself, healthy 09:02:06Z |
| kill 5 s into a deploy | **not measured**: the three test restarts had filled the budget, so the guard paused (as designed) and the hub mailed `controller_crashloop` |
| resume after the pause | pause held to 09:35:52Z; restarted 09:36:24Z; dashboard 200 at 09:36:30Z |
Two invalid attempts, both marked in the evidence: a swap POST sent over plain HTTP (400, no swap), and a "health 200"
line during the pause that the container state contradicts.
## Teardown
Machine: 9201 back on its controller (see resume); the throwaway homebox deploy left no container. Host: park marker
removed; nothing else changed. Hub: demo-hp floor override 0.243.0 kept (it delivers this release).
Evidence: `felhom.eu/documentation/audits/evidence-p1fixes-2026-09-15/A*`.