e98b857684
gates / gates (push) Successful in 13s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
48 lines
3.5 KiB
Markdown
48 lines
3.5 KiB
Markdown
# REPORT — agent v0.131.0: the controller comes back by itself; backup status per tier (2026-09-15)
|
||
|
||
Task: *before the volunteer — the big night's P1 fixes*, Parts A and C (agent half). Architecture: `03-host-agent.md`
|
||
§4 (the "healing a crashed controller" sentence), `07-backup-architecture.md` §6.
|
||
|
||
## Measured first (A.1)
|
||
Docker 29.8.0, throwaway containers on scratch 9202: after `docker kill`, **both** `--restart unless-stopped` and
|
||
`--restart always` stayed `exited (137)` 60 s later. The task's claim was right; a policy change is not a fix.
|
||
|
||
## What shipped
|
||
- **Controller supervisor** (`internal/localapi/controllersupervisor.go`): every 30 s, for provisioned felhom-pool guests
|
||
that are running, restart `felhom-controller-bootstrap.service` on the second not-running observation. Guards: swap
|
||
in flight, host-side park marker, locked / vzdump-busy / stopped guest, unknown docker answer, unprovisioned guest,
|
||
3 restarts in 15 min → 30 min pause. Record rides the host report as `controller_supervisor`; hub v0.114.0 mints the
|
||
events. Same GuestExecutor and sudoers grants as the swap — no new privilege.
|
||
- **Per-tier backup status**: `GET /backup/status` (untargeted) gains `tiers[]` — newest success (record, or storage
|
||
after a restart), last attempt kept apart, storage presence; `GET /backup/tiers` gains `storage`.
|
||
- **Golden script**: `--restart always`. No golden baked (R-468); existing boxes keep `unless-stopped`.
|
||
|
||
## Red-proofs (each seen failing, then restored)
|
||
- remove the restart call → `the killed controller was NOT restarted — this is R-523 (restarts=0)`
|
||
- remove the backoff block → `crash-looping controller restarted 10 times in 10 minutes — want exactly 3`
|
||
- `last_success` from the newest attempt → `pbs tier reports a failed attempt as its last success`
|
||
|
||
## Release and delivery
|
||
`scripts/release-agent.sh 0.131.0`: tag `v0.131.0`, sha256 `1118b552f7e775fbde9544c7764ede7e6046e0a7db16ae8494d07a18e3c2ac9c`,
|
||
verified by download. **The task's "the floor delivers the agent" was wrong** (R-530): the hub holds a floor above the
|
||
box's agent; agents update only by an operator-signed job. On the operator's keys: `felhom-opsign -op agent_update` for
|
||
`demo-hp-bb76ea` only → authorized 08:44:16Z, committed 08:45:21Z, `controller-supervisor: started`. demo-felhom and
|
||
Peti's box stay on 0.130.0.
|
||
|
||
## Live validation on demo-hp guest 9201 (agent 0.131.0, controller 0.243.0)
|
||
| moment | result |
|
||
|---|---|
|
||
| idle kill 08:53:27Z | restarted 08:54:22Z; dashboard 200 **59 s** after the kill |
|
||
| parked + kill | stayed dead 100 s, `the guest is PARKED — leaving it` every sweep; unpark → 200 in **25 s** |
|
||
| kill 10 s into a swap | `during a controller SWAP — the swap owns it` ×3; the swap rolled back itself, healthy 09:02:06Z |
|
||
| kill 5 s into a deploy | **not measured**: the three test restarts had filled the budget, so the guard paused (as designed) and the hub mailed `controller_crashloop` |
|
||
| resume after the pause | pause held to 09:35:52Z; restarted 09:36:24Z; dashboard 200 at 09:36:30Z |
|
||
Two invalid attempts, both marked in the evidence: a swap POST sent over plain HTTP (400, no swap), and a "health 200"
|
||
line during the pause that the container state contradicts.
|
||
|
||
## Teardown
|
||
Machine: 9201 back on its controller (see resume); the throwaway homebox deploy left no container. Host: park marker
|
||
removed; nothing else changed. Hub: demo-hp floor override 0.243.0 kept (it delivers this release).
|
||
|
||
Evidence: `felhom.eu/documentation/audits/evidence-p1fixes-2026-09-15/A*`.
|