# REPORT — controller v0.246.0: an interrupted restore is told; the recovery-code reminder waits **2026-09-17.** R-550 (operator ruling „fix") and R-546. Architecture read first: `felhom.eu/documentation/architecture/07-backup-architecture.md` §3 (Lane 1) and `08-alarm-ladder.md` §6.2. ## Claims in the brief that turned out wrong — named first 1. **„Readiness = `escrow.pbs_storage_id` set."** Readiness is the agent's own preflight `ok` over **five** blocking items (`pbs_storage_id`, `dr_tier`, `age_binary`, `hub_upload`, `sudo_grant`; `felhom-agent/internal/localapi/escrow_ceremony.go` `handleEscrowPreflight`). `pbs_storage_id` was the one item R-546's box happened to show. The controller now reads the combined `ok`, never a copy. 2. **„The failure shows a raw error"** — not in a browser. The escrow page already listed the checklist and showed its start form only when the preflight was `ok`. The raw `-storage` stderr came from the chaos-night harness calling `POST /api/escrow/start` directly. What a browser user DID meet: red crosses, the agent's English detail in muted brackets, no start button and no word about waiting. Both paths are fixed; the direct one is now refused server-side. 3. **„The backup tiers persist their records atomically"** (read from `07`, not re-checked) — **checked and TRUE.** The tier run records live in `settings.json`, and `Settings.save()` (`internal/settings/settings.go:727-752`) writes a `.tmp` and renames it, with `.bak` recovery on a corrupt primary. The restore record follows the same shape through the backup package's own `atomicWrite` (tmp + rename; like `save()`, no fsync). ## What shipped **B.1 — the restore record survives a restart (R-550).** A design REVERSED by ruling, recorded in the file header, in `07` §3 and in `CONTEXT.md`: `restore-status.json` in `DataDir`, written at both ends of an op (`internal/backup/restore_record.go`). A record still marked running at startup becomes a failed, `Interrupted` result — „A visszaállítás megszakadt (a doboz újraindult) — indítsd el újra." — kept per app until that app's next restore; `restore_interrupted` (warning, household) pushed once. Wired from `main()` (`cmd/controller/restore_record_wiring.go`). Shown as a card on `/backups/restore` and in the off-site wizard's outcome card. Notification cooldowns stay in memory. **B.2 — the reminder waits (R-546).** `internal/web/escrow_readiness.go`: the agent's preflight `ok`, cached 60 s, probed only while paused. The bar is held back while not ready; `/backup/escrow` shows a waiting card that polls and reloads; `POST /api/escrow/start` refuses 409 before staging. Unknown readiness keeps the bar (fail loud). Transitions logged at INFO. **B.3 — the guide.** `VOLUNTEER-first-hour.md`: the recovery code is now step 7, after the first apps, and begins „Amikor a sárga sáv megjelenik (a beállítás után néhány perccel)". **MinAgent: 0.131.0, unchanged** — the preflight has been served since agent 0.88.0. ## Red-proofs — each seen failing, then passing | test | break | failure seen | |---|---|---| | `TestRestoreRecord_InterruptedRestoreSurvivesRestart` | no persistence | `StartedAt:0001-01-01 … Last:` — the chaos-night value | | `TestRestoreRecordAtStartup_RaisesInterruptedOnce` | helper skips loading | `raised no restore_interrupted: []` | | `TestMainWiresRestoreRecord` | main() not wired | `does not call both … inert seam` | | `TestR550_RestorePageShowsInterruptedRestore` | handler line removed | `the restore page does not show the interrupted restore` | | `TestR546_NotReadyBoxHoldsTheBar` | readiness not consulted | `the bar urges a ceremony the box cannot run yet` | | `TestR546_EscrowPageWaitsWhenNotReady` | flag never set | `does not tell the household the box is still preparing` | | `TestR546_StartRefusedWhenNotReady` | gate removed | `= 200 … job_id escrow-1` | Controls: finished restore stays finished; the next restore of that app clears the notice (another app's does not); a ready box shows the bar; unknown readiness shows the bar; no interruption → no card. Two of my own test mistakes, corrected: a `-storage` needle that matched the menu's `nav-group-storage` id (narrowed to `requires -storage`); the three start-order tests updated to `preflight,…` — their load-bearing assertion (stage BEFORE start) unchanged. ## Gates `go build ./... && go vet ./... && go test ./...` — green, 28 packages. `controller_gates.py --fast` — all OK (`golden-notice` advisory; golden stays 0.245.0 under the waiver to 2026-09-27). ## Live validation (endpoint-level; no browser on DooPlex) **B.4(a) — PASS, on demo-hp guest 9201 with a throwaway `homebox`.** Restore accepted 08:51:51Z; 2 s in the status and the file on disk both said `running … homebox`; `docker kill felhom-controller`; the agent's supervisor restarted it in 42 s (not by hand). After restart the status read `ok:false … "A visszaállítás megszakadt …" interrupted:true`; `/backups/restore` showed the card (negative control 0); the controller logged `restore_interrupted pushed OK (HTTP 200)`; the hub stores the event under demo-hp. A second restore of homebox completed `ok:true` and the card was gone (0). Evidence: `felhom.eu/documentation/audits/evidence-chaos-fixes-2026-09-17/partB4a-*.txt`. **B.4(b) — NOT live-validated, and why (R-551).** No Tier-0 box is both paused and agent-connected: 9201 is escrowed (bar off by design; making it paused would be hand-set state on the standing demo box, and a real ceremony would supersede its escrow — the hub keeps one `host_escrow` row per host); 9202 has no local-API token. Proven by the three tests above through `ServeHTTP` with a fake agent; chaos night measured live that the agent preflight is red ~17 min after a bind and turns green by itself. **The ceremony was deliberately NOT run on demo-hp.** ## Teardown - **Machine:** homebox removed with data and backups (0 containers, 0 volumes); 10 of 10 standing apps up. My first removal was refused 409 („still running — stop it first") — the product was right; stopped, then removed. Password and scripts shredded in the guest (0 left). - **Host:** nothing provisioned. 9201 stays on controller 0.246.0 (the release under validation). - **Hub:** the `restore_interrupted` and `controller_restarted_by_agent` events stay as history. ## Observations 1. An interrupted-restore notice for an app that is then removed stays for ever. **FILED: R-552** 2. No Tier-0 box can exercise the paused + agent-connected escrow state. **FILED: R-551** 3. `PushEvent` is best-effort (3 attempts, 3 s apart); `restore_interrupted` inherits that. **NOT-A-FINDING: the page notice comes from the persisted file and does not depend on the event; the event-drop path is the existing, recorded behaviour of every controller event.**