From e98b857684f41db47283188a44105ae2b73a4125 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Tue, 15 Sep 2026 11:37:07 +0200 Subject: [PATCH] REPORT: v0.131.0 supervisor + per-tier status, delivery and live validation Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- REPORT.md | 81 +++++++++++++++++++++++++++---------------------------- 1 file changed, 39 insertions(+), 42 deletions(-) diff --git a/REPORT.md b/REPORT.md index 220eabe..6a11fe1 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,50 +1,47 @@ -# REPORT — agent v0.129.0: a correct code for an earlier package (R-311, 2026-08-12) +# REPORT — agent v0.131.0: the controller comes back by itself; backup status per tier (2026-09-15) -## What changed and why +Task: *before the volunteer — the big night's P1 fixes*, Parts A and C (agent half). Architecture: `03-host-agent.md` +§4 (the "healing a crashed controller" sentence), `07-backup-architecture.md` §6. -Yesterday's drill proved a retained escrow package **works** — unsealed with the old recovery code, it -opened a set-aside store and restored planted files byte-identical — while this agent answered that -same correct code with *"the recovery code did not open the sealed bundle"*. Nothing had ever tried -the retained packages, so a correct-but-earlier code and a mistype were genuinely indistinguishable. +## Measured first (A.1) +Docker 29.8.0, throwaway containers on scratch 9202: after `docker kill`, **both** `--restart unless-stopped` and +`--restart always` stayed `exited (137)` 60 s later. The task's claim was right; a policy change is not a fix. -- `internal/hub/client.go` — `FetchRetainedIdentityEscrow` → `GET /api/v1/hosts//escrow/retained` - (hub ≥ v0.103.0). **A 404 is a clean "none"**, not a fault: an older hub must not turn into a failed - recovery. -- `internal/escrow/recover.go` — optional `FetchRetained`, `ErrCodeOpensRetained` + - `RetainedOpenedError{SupersededAt, KeyFingerprint, Index, HasResticPassword}`. Consulted **only** - after the current package refuses. -- `internal/localapi/escrow_recover.go` — a **fifth** case on the R-224 switch: **422**, with - `opens_retained`, `superseded_at`, `retained_has_restic_pw`. Added to the switch, not a restructure. -- `cmd/felhom-agent/main.go` — the retained fetcher wired on the same self-scoped hub client. +## What shipped +- **Controller supervisor** (`internal/localapi/controllersupervisor.go`): every 30 s, for provisioned felhom-pool guests + that are running, restart `felhom-controller-bootstrap.service` on the second not-running observation. Guards: swap + in flight, host-side park marker, locked / vzdump-busy / stopped guest, unknown docker answer, unprovisioned guest, + 3 restarts in 15 min → 30 min pause. Record rides the host report as `controller_supervisor`; hub v0.114.0 mints the + events. Same GuestExecutor and sudoers grants as the swap — no new privilege. +- **Per-tier backup status**: `GET /backup/status` (untargeted) gains `tiers[]` — newest success (record, or storage + after a restart), last attempt kept apart, storage presence; `GET /backup/tiers` gains `storage`. +- **Golden script**: `--restart always`. No golden baked (R-468); existing boxes keep `unless-stopped`. -## Fail-safe, in every direction +## Red-proofs (each seen failing, then restored) +- remove the restart call → `the killed controller was NOT restarted — this is R-523 (restarts=0)` +- remove the backoff block → `crash-looping controller restarted 10 times in 10 minutes — want exactly 3` +- `last_success` from the newest attempt → `pbs tier reports a failed attempt as its last success` -nil fetcher · hub without the route (404) · transport failure · malformed package → **the original -refusal stands, unchanged**. The worst outcome of this feature breaking is the behaviour before it. -Attempts bounded (`MaxRetainedTried`, default 6) — each unwrap is ~1 s of scrypt, so an unbounded loop -would turn one wrong code into a minutes-long hang. +## Release and delivery +`scripts/release-agent.sh 0.131.0`: tag `v0.131.0`, sha256 `1118b552f7e775fbde9544c7764ede7e6046e0a7db16ae8494d07a18e3c2ac9c`, +verified by download. **The task's "the floor delivers the agent" was wrong** (R-530): the hub holds a floor above the +box's agent; agents update only by an operator-signed job. On the operator's keys: `felhom-opsign -op agent_update` for +`demo-hp-bb76ea` only → authorized 08:44:16Z, committed 08:45:21Z, `controller-supervisor: started`. demo-felhom and +Peti's box stay on 0.130.0. -## Tests — 7, with REAL age crypto +## Live validation on demo-hp guest 9201 (agent 0.131.0, controller 0.243.0) +| moment | result | +|---|---| +| idle kill 08:53:27Z | restarted 08:54:22Z; dashboard 200 **59 s** after the kill | +| parked + kill | stayed dead 100 s, `the guest is PARKED — leaving it` every sweep; unpark → 200 in **25 s** | +| kill 10 s into a swap | `during a controller SWAP — the swap owns it` ×3; the swap rolled back itself, healthy 09:02:06Z | +| kill 5 s into a deploy | **not measured**: the three test restarts had filled the budget, so the guard paused (as designed) and the hub mailed `controller_crashloop` | +| resume after the pause | pause held to 09:35:52Z; restarted 09:36:24Z; dashboard 200 at 09:36:30Z | +Two invalid attempts, both marked in the evidence: a swap POST sent over plain HTTP (400, no swap), and a "health 200" +line during the pause that the container state contradicts. -Real crypto because the two situations are indistinguishable **at the unwrap**; a faked unwrap would -prove nothing about what was broken. Full suite green (`go build`/`vet`/`test ./...`), agent gates OK. +## Teardown +Machine: 9201 back on its controller (see resume); the throwaway homebox deploy left no container. Host: park marker +removed; nothing else changed. Hub: demo-hp floor override 0.243.0 kept (it delivers this release). -**Red-proof, mutation asserted applied before the run:** remove the `tryRetained` block from -`RecoverOffsiteRepoPassword` → -`err = escrow: the recovery code did not unwrap the identity escrow (wrong recovery code…)` → -`TestRecover_CodeOpensRetainedPackage_IsNotAWrongCode` FAILS. **The lie returns, in those words.** -That is the layer the lie actually lives in: removing the *controller's* case yields the neutral -message instead, because R-224's safe default catches it. - -## Released and deployed - -`release-agent.sh 0.129.0` — tagged `v0.129.0`, published, **verified by independent download**, -sha256 `53a54f0620afbd6d…`. Installed on `felhom-pve`, `felhom-agent --version` = 0.129.0, unit active, -journal clean (normal PBS verify cycle). **NOT VOUCHED** — that stays the operator's act. - -## Bypass, stated as required - -`git push --no-verify` was used **once** for the code push. The `release-complete` gate refuses a -CHANGELOG entry whose tag and package do not exist, and `release-agent.sh` refuses a tree that is not -pushed — circular by construction. The bypass was immediately followed by the real release; gates were -re-run afterwards and are **green**, and the tag+package now exist. +Evidence: `felhom.eu/documentation/audits/evidence-p1fixes-2026-09-15/A*`.