Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
6.8 KiB
REPORT — controller v0.246.0: an interrupted restore is told; the recovery-code reminder waits
2026-09-17. R-550 (operator ruling „fix") and R-546. Architecture read first:
felhom.eu/documentation/architecture/07-backup-architecture.md §3 (Lane 1) and
08-alarm-ladder.md §6.2.
Claims in the brief that turned out wrong — named first
- „Readiness =
escrow.pbs_storage_idset." Readiness is the agent's own preflightokover five blocking items (pbs_storage_id,dr_tier,age_binary,hub_upload,sudo_grant;felhom-agent/internal/localapi/escrow_ceremony.gohandleEscrowPreflight).pbs_storage_idwas the one item R-546's box happened to show. The controller now reads the combinedok, never a copy. - „The failure shows a raw error" — not in a browser. The escrow page already listed the checklist
and showed its start form only when the preflight was
ok. The raw-storagestderr came from the chaos-night harness callingPOST /api/escrow/startdirectly. What a browser user DID meet: red crosses, the agent's English detail in muted brackets, no start button and no word about waiting. Both paths are fixed; the direct one is now refused server-side. - „The backup tiers persist their records atomically" (read from
07, not re-checked) — checked and TRUE. The tier run records live insettings.json, andSettings.save()(internal/settings/settings.go:727-752) writes a.tmpand renames it, with.bakrecovery on a corrupt primary. The restore record follows the same shape through the backup package's ownatomicWrite(tmp + rename; likesave(), no fsync).
What shipped
B.1 — the restore record survives a restart (R-550). A design REVERSED by ruling, recorded in the
file header, in 07 §3 and in CONTEXT.md: restore-status.json in DataDir, written at both ends of an
op (internal/backup/restore_record.go). A record still marked running at startup becomes a failed,
Interrupted result — „A visszaállítás megszakadt (a doboz újraindult) — indítsd el újra." — kept per
app until that app's next restore; restore_interrupted (warning, household) pushed once. Wired from
main() (cmd/controller/restore_record_wiring.go). Shown as a card on /backups/restore and in the
off-site wizard's outcome card. Notification cooldowns stay in memory.
B.2 — the reminder waits (R-546). internal/web/escrow_readiness.go: the agent's preflight ok,
cached 60 s, probed only while paused. The bar is held back while not ready; /backup/escrow shows a
waiting card that polls and reloads; POST /api/escrow/start refuses 409 before staging. Unknown
readiness keeps the bar (fail loud). Transitions logged at INFO.
B.3 — the guide. VOLUNTEER-first-hour.md: the recovery code is now step 7, after the first apps,
and begins „Amikor a sárga sáv megjelenik (a beállítás után néhány perccel)".
MinAgent: 0.131.0, unchanged — the preflight has been served since agent 0.88.0.
Red-proofs — each seen failing, then passing
| test | break | failure seen |
|---|---|---|
TestRestoreRecord_InterruptedRestoreSurvivesRestart |
no persistence | StartedAt:0001-01-01 … Last:<nil> — the chaos-night value |
TestRestoreRecordAtStartup_RaisesInterruptedOnce |
helper skips loading | raised no restore_interrupted: [] |
TestMainWiresRestoreRecord |
main() not wired | does not call both … inert seam |
TestR550_RestorePageShowsInterruptedRestore |
handler line removed | the restore page does not show the interrupted restore |
TestR546_NotReadyBoxHoldsTheBar |
readiness not consulted | the bar urges a ceremony the box cannot run yet |
TestR546_EscrowPageWaitsWhenNotReady |
flag never set | does not tell the household the box is still preparing |
TestR546_StartRefusedWhenNotReady |
gate removed | = 200 … job_id escrow-1 |
Controls: finished restore stays finished; the next restore of that app clears the notice (another app's
does not); a ready box shows the bar; unknown readiness shows the bar; no interruption → no card.
Two of my own test mistakes, corrected: a -storage needle that matched the menu's nav-group-storage
id (narrowed to requires -storage); the three start-order tests updated to preflight,… — their
load-bearing assertion (stage BEFORE start) unchanged.
Gates
go build ./... && go vet ./... && go test ./... — green, 28 packages. controller_gates.py --fast —
all OK (golden-notice advisory; golden stays 0.245.0 under the waiver to 2026-09-27).
Live validation (endpoint-level; no browser on DooPlex)
B.4(a) — PASS, on demo-hp guest 9201 with a throwaway homebox. Restore accepted 08:51:51Z; 2 s in
the status and the file on disk both said running … homebox; docker kill felhom-controller; the
agent's supervisor restarted it in 42 s (not by hand). After restart the status read
ok:false … "A visszaállítás megszakadt …" interrupted:true; /backups/restore showed the card
(negative control 0); the controller logged restore_interrupted pushed OK (HTTP 200); the hub stores the
event under demo-hp. A second restore of homebox completed ok:true and the card was gone (0).
Evidence: felhom.eu/documentation/audits/evidence-chaos-fixes-2026-09-17/partB4a-*.txt.
B.4(b) — NOT live-validated, and why (R-551). No Tier-0 box is both paused and agent-connected:
9201 is escrowed (bar off by design; making it paused would be hand-set state on the standing demo box,
and a real ceremony would supersede its escrow — the hub keeps one host_escrow row per host); 9202 has
no local-API token. Proven by the three tests above through ServeHTTP with a fake agent; chaos night
measured live that the agent preflight is red ~17 min after a bind and turns green by itself.
The ceremony was deliberately NOT run on demo-hp.
Teardown
- Machine: homebox removed with data and backups (0 containers, 0 volumes); 10 of 10 standing apps up. My first removal was refused 409 („still running — stop it first") — the product was right; stopped, then removed. Password and scripts shredded in the guest (0 left).
- Host: nothing provisioned. 9201 stays on controller 0.246.0 (the release under validation).
- Hub: the
restore_interruptedandcontroller_restarted_by_agentevents stay as history.
Observations
- An interrupted-restore notice for an app that is then removed stays for ever. FILED: R-552
- No Tier-0 box can exercise the paused + agent-connected escrow state. FILED: R-551
PushEventis best-effort (3 attempts, 3 s apart);restore_interruptedinherits that. NOT-A-FINDING: the page notice comes from the persisted file and does not depend on the event; the event-drop path is the existing, recorded behaviour of every controller event.