Files
felhom-controller/REPORT.md
T
2026-09-17 10:59:21 +02:00

6.8 KiB

REPORT — controller v0.246.0: an interrupted restore is told; the recovery-code reminder waits

2026-09-17. R-550 (operator ruling „fix") and R-546. Architecture read first: felhom.eu/documentation/architecture/07-backup-architecture.md §3 (Lane 1) and 08-alarm-ladder.md §6.2.

Claims in the brief that turned out wrong — named first

  1. „Readiness = escrow.pbs_storage_id set." Readiness is the agent's own preflight ok over five blocking items (pbs_storage_id, dr_tier, age_binary, hub_upload, sudo_grant; felhom-agent/internal/localapi/escrow_ceremony.go handleEscrowPreflight). pbs_storage_id was the one item R-546's box happened to show. The controller now reads the combined ok, never a copy.
  2. „The failure shows a raw error" — not in a browser. The escrow page already listed the checklist and showed its start form only when the preflight was ok. The raw -storage stderr came from the chaos-night harness calling POST /api/escrow/start directly. What a browser user DID meet: red crosses, the agent's English detail in muted brackets, no start button and no word about waiting. Both paths are fixed; the direct one is now refused server-side.
  3. „The backup tiers persist their records atomically" (read from 07, not re-checked) — checked and TRUE. The tier run records live in settings.json, and Settings.save() (internal/settings/settings.go:727-752) writes a .tmp and renames it, with .bak recovery on a corrupt primary. The restore record follows the same shape through the backup package's own atomicWrite (tmp + rename; like save(), no fsync).

What shipped

B.1 — the restore record survives a restart (R-550). A design REVERSED by ruling, recorded in the file header, in 07 §3 and in CONTEXT.md: restore-status.json in DataDir, written at both ends of an op (internal/backup/restore_record.go). A record still marked running at startup becomes a failed, Interrupted result — „A visszaállítás megszakadt (a doboz újraindult) — indítsd el újra." — kept per app until that app's next restore; restore_interrupted (warning, household) pushed once. Wired from main() (cmd/controller/restore_record_wiring.go). Shown as a card on /backups/restore and in the off-site wizard's outcome card. Notification cooldowns stay in memory.

B.2 — the reminder waits (R-546). internal/web/escrow_readiness.go: the agent's preflight ok, cached 60 s, probed only while paused. The bar is held back while not ready; /backup/escrow shows a waiting card that polls and reloads; POST /api/escrow/start refuses 409 before staging. Unknown readiness keeps the bar (fail loud). Transitions logged at INFO.

B.3 — the guide. VOLUNTEER-first-hour.md: the recovery code is now step 7, after the first apps, and begins „Amikor a sárga sáv megjelenik (a beállítás után néhány perccel)".

MinAgent: 0.131.0, unchanged — the preflight has been served since agent 0.88.0.

Red-proofs — each seen failing, then passing

test break failure seen
TestRestoreRecord_InterruptedRestoreSurvivesRestart no persistence StartedAt:0001-01-01 … Last:<nil> — the chaos-night value
TestRestoreRecordAtStartup_RaisesInterruptedOnce helper skips loading raised no restore_interrupted: []
TestMainWiresRestoreRecord main() not wired does not call both … inert seam
TestR550_RestorePageShowsInterruptedRestore handler line removed the restore page does not show the interrupted restore
TestR546_NotReadyBoxHoldsTheBar readiness not consulted the bar urges a ceremony the box cannot run yet
TestR546_EscrowPageWaitsWhenNotReady flag never set does not tell the household the box is still preparing
TestR546_StartRefusedWhenNotReady gate removed = 200 … job_id escrow-1

Controls: finished restore stays finished; the next restore of that app clears the notice (another app's does not); a ready box shows the bar; unknown readiness shows the bar; no interruption → no card. Two of my own test mistakes, corrected: a -storage needle that matched the menu's nav-group-storage id (narrowed to requires -storage); the three start-order tests updated to preflight,… — their load-bearing assertion (stage BEFORE start) unchanged.

Gates

go build ./... && go vet ./... && go test ./... — green, 28 packages. controller_gates.py --fast — all OK (golden-notice advisory; golden stays 0.245.0 under the waiver to 2026-09-27).

Live validation (endpoint-level; no browser on DooPlex)

B.4(a) — PASS, on demo-hp guest 9201 with a throwaway homebox. Restore accepted 08:51:51Z; 2 s in the status and the file on disk both said running … homebox; docker kill felhom-controller; the agent's supervisor restarted it in 42 s (not by hand). After restart the status read ok:false … "A visszaállítás megszakadt …" interrupted:true; /backups/restore showed the card (negative control 0); the controller logged restore_interrupted pushed OK (HTTP 200); the hub stores the event under demo-hp. A second restore of homebox completed ok:true and the card was gone (0). Evidence: felhom.eu/documentation/audits/evidence-chaos-fixes-2026-09-17/partB4a-*.txt.

B.4(b) — NOT live-validated, and why (R-551). No Tier-0 box is both paused and agent-connected: 9201 is escrowed (bar off by design; making it paused would be hand-set state on the standing demo box, and a real ceremony would supersede its escrow — the hub keeps one host_escrow row per host); 9202 has no local-API token. Proven by the three tests above through ServeHTTP with a fake agent; chaos night measured live that the agent preflight is red ~17 min after a bind and turns green by itself. The ceremony was deliberately NOT run on demo-hp.

Teardown

  • Machine: homebox removed with data and backups (0 containers, 0 volumes); 10 of 10 standing apps up. My first removal was refused 409 („still running — stop it first") — the product was right; stopped, then removed. Password and scripts shredded in the guest (0 left).
  • Host: nothing provisioned. 9201 stays on controller 0.246.0 (the release under validation).
  • Hub: the restore_interrupted and controller_restarted_by_agent events stay as history.

Observations

  1. An interrupted-restore notice for an app that is then removed stays for ever. FILED: R-552
  2. No Tier-0 box can exercise the paused + agent-connected escrow state. FILED: R-551
  3. PushEvent is best-effort (3 attempts, 3 s apart); restore_interrupted inherits that. NOT-A-FINDING: the page notice comes from the persisted file and does not depend on the event; the event-drop path is the existing, recorded behaviour of every controller event.