diff --git a/CHANGELOG.md b/CHANGELOG.md index 125d1bb..83fbc94 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,3 +1,36 @@ +## v0.246.0 — an interrupted restore is told, and the recovery-code reminder waits until the box can take it (2026-09-17, R-550 / R-546) + +**MinAgent: 0.131.0** (unchanged — the readiness check reads the escrow preflight, served since agent +0.88.0; nothing here needs agent 0.132.0). **Requires hub v0.117.0** for `restore_interrupted` to be +accepted; an older hub 400s that one event and nothing else changes. + +- **R-550 (operator ruling "fix", 2026-09-17) — the restore record survives a restart.** A DESIGN + REVERSED: `opstatus.go` was in memory by choice ("same precedent as notification cooldowns"). Chaos + night round 10 measured the cost — a restore accepted, the box hard-reset four seconds later, and + `/api/backup/restore-status` answering the Go zero value; nobody could learn whether it finished. + Now `restore-status.json` in `DataDir`, written atomically at both ends of an op + (`internal/backup/restore_record.go`). At startup a record still marked running becomes a failed, + `Interrupted` result — „A visszaállítás megszakadt (a doboz újraindult) — indítsd el újra." — kept + per app until that app's next restore, shown as a „Megszakadt visszaállítás" card on + `/backups/restore` and in the off-site wizard's outcome card, and raised ONCE as + `restore_interrupted` (warning, for the household). Cooldowns stay in memory. +- **R-546 — the recovery-code reminder waits until the box can run the ceremony.** The R-543 bar now + consults the agent's OWN preflight `ok` (every blocking item — not a copy of `pbs_storage_id`), cached + 60 s and probed only while paused; held back while not ready. `/backup/escrow` shows a waiting card + that polls and reloads itself instead of red crosses and the agent's English diagnostics. + `POST /api/escrow/start` refuses 409 with the same sentence BEFORE staging or starting — the direct + path chaos night used, which ran the ceremony and returned the raw `-storage` stderr. Unknown + readiness keeps the bar (fail loud). The transition is logged at INFO. + +**Red-proofs, each seen failing then passing:** the restore record across a restart (blank status, +`StartedAt:0001-01-01` — the chaos-night symptom exactly); a finished restore stays finished; the next +restore of that app clears its notice; `main()` calls both startup functions (AST); the startup helper +with loading skipped pushes nothing; the restore page card with the handler line removed; the bar held +back (readiness check removed → bar shown); the waiting card (never set → checklist); the start refusal +(gate removed → 200 and `stage,start` ran). Controls: a ready box shows the bar; unknown readiness shows +the bar; a box with no interruption shows no card. The three existing start-order tests now expect +`preflight` first — their load-bearing assertion (stage BEFORE start) is unchanged. + ## v0.245.0 — the household is asked for the recovery code, and the page says „szünetel" until then (2026-09-16, R-543) **MinAgent: 0.131.0** (unchanged — nothing here needs a newer agent) diff --git a/controller/README.md b/controller/README.md index a1ad72d..a8339b4 100644 --- a/controller/README.md +++ b/controller/README.md @@ -1333,6 +1333,24 @@ truncated. > Those figures do not extrapolate: the structure check's cost tracks the index, read-data's tracks > the data. +### The restore record survives a restart (v0.246.0, R-550) + +The restore op-status (`internal/backup/opstatus.go`, served at `GET /api/backup/restore-status`) was +in memory only. Chaos night round 10 hard-reset a box four seconds into a restore; afterwards the +status was the Go zero value and nothing told the household whether the restore finished. **Operator +ruling 2026-09-17 reversed the in-memory choice for the restore record only** (notification cooldowns +stay in memory). + +- `restore-status.json` in `DataDir` (beside `settings.json`), written atomically at **both** ends of an op. +- At startup (`loadRestoreRecordAtStartup`, before any page is served) a record still marked running + becomes a failed result — „A visszaállítás megszakadt (a doboz újraindult) — indítsd el újra." — + and a per-app notice that stays until **that app's** next restore. +- `restore_interrupted` (warning, for the household; hub v0.117.0) is pushed **once** per + interruption, after the notifier exists (`reportInterruptedRestore`). Best-effort like every + `PushEvent` (3 attempts, 3 s apart); the page notice does not depend on it. +- `/backups/restore` shows a „Megszakadt visszaállítás" card per interrupted app; the off-site + wizard's „Eredmény" card falls back to the interrupted record („Észlelve: …"). + ### Restore refusals (v0.226.0) Three guards added on the off-site restore surface, all server-side: @@ -1585,6 +1603,15 @@ dismissal, back at the next visit, gone for good when the state is `escrowed`. I pages bypass that function, and a session check keeps it off the public guest share page. The pause itself is UNCHANGED — it is the zero-knowledge escrow design, not a defect. +**…but only once the box can do it (v0.246.0, R-546).** For the first minutes after a bind the +agent's escrow preflight is not `ok` (typically no PBS storage yet — ~17 min measured). The bar now +consults the agent's own `ok` (`escrow_readiness.go`, cached 60 s, probed only while paused) and is +held back while the agent says not ready; `/backup/escrow` shows a waiting card („A doboz még készül +— … pár perc múlva …") that polls the preflight and reloads itself, instead of a red checklist; and +`POST /api/escrow/start` refuses 409 with the same sentence **before** staging or starting. Readiness +**unknown** (agent unreachable) keeps the bar — the fail-loud rule of R-543. A re-ceremony on an +escrowed box keeps the checklist, where a red row is a real fault. + #### Restore (`internal/backup/restore.go`) Both **Tier 1** (restic) and **Tier 2** (rsync) restores are supported. All deployed apps diff --git a/controller/cmd/controller/main.go b/controller/cmd/controller/main.go index cdc752a..b991dc6 100644 --- a/controller/cmd/controller/main.go +++ b/controller/cmd/controller/main.go @@ -527,6 +527,8 @@ func main() { // --- Initialize backup manager (app-data only: DB dumps + Docker-volume tars) --- var backupMgr *backup.Manager + // R-550: the restore the last stop interrupted, found when the record loads; raised once the notifier exists. + var interruptedRestore *backup.RestoreOpResult stackProv := &stackAdapter{ mgr: stackMgr, getStoragePaths: func() []settings.StoragePath { return sett.GetStoragePaths() }, @@ -534,6 +536,8 @@ func main() { } if cfg.Backup.Enabled { backupMgr = backup.NewManager(cfg, sett, logger) + // R-550: load the persisted restore record BEFORE any page can ask for it. + interruptedRestore = loadRestoreRecordAtStartup(backupMgr, cfg.Paths.DataDir, logger) // R-166: use the guard that already ran Recover at startup, not a second one over the same // file (see SetAppStopGuard — one file, one owner). backupMgr.SetAppStopGuard(appStopGuard) @@ -580,6 +584,8 @@ func main() { // --- Initialize notifier --- notifier := notify.New(cfg.Hub.URL, cfg.Hub.APIKey, cfg.Customer.ID, sett, logger, cfg.Logging.Level == "debug") + // R-550: an interrupted restore is raised once, now that the notifier exists. + reportInterruptedRestore(notifier, interruptedRestore) // R-97a: wire the quiesce loop's hub-event seam. It MUST happen here and not in // startQuiesceLoop, because the notifier is constructed after it — and it must happen at all, diff --git a/controller/cmd/controller/restore_record_wiring.go b/controller/cmd/controller/restore_record_wiring.go new file mode 100644 index 0000000..b196e85 --- /dev/null +++ b/controller/cmd/controller/restore_record_wiring.go @@ -0,0 +1,48 @@ +package main + +import ( + "fmt" + "log" + "path/filepath" + + "gitea.dooplex.hu/admin/felhom-controller/internal/backup" +) + +// R-550 — the restore record survives a restart, and an interruption is raised ONCE. +// +// Two functions so main() makes two plain calls a test can see (TestMainWiresRestoreRecord): the +// record must be loaded BEFORE the web server serves a restore page, and the event can only be pushed +// AFTER the notifier exists, which main() constructs later. + +// restoreRecordFileName lives in DataDir, beside settings.json — the SSD state dir that survives a +// container recreation. +const restoreRecordFileName = "restore-status.json" + +// restoreEventPusher is the notifier method main() hands in (an interface so a test can record it). +type restoreEventPusher interface { + PushEvent(eventType, severity, message string, details interface{}) +} + +// loadRestoreRecordAtStartup wires persistence and returns the restore the last stop interrupted. +func loadRestoreRecordAtStartup(mgr *backup.Manager, dataDir string, logger *log.Logger) *backup.RestoreOpResult { + if mgr == nil { + return nil + } + path := filepath.Join(dataDir, restoreRecordFileName) + mgr.SetRestoreRecordPath(path) + res := mgr.LoadRestoreRecord() + if logger != nil { + logger.Printf("[INFO] [backup] restore record wired: %s (interrupted at startup: %t)", path, res != nil) + } + return res +} + +// reportInterruptedRestore raises restore_interrupted (warning; for the household — hub v0.117.0). +func reportInterruptedRestore(p restoreEventPusher, res *backup.RestoreOpResult) { + if p == nil || res == nil { + return + } + p.PushEvent("restore_interrupted", "warning", + fmt.Sprintf("A(z) %s visszaállítása megszakadt, mert a doboz újraindult — indítsd el újra.", res.Stack), + map[string]any{"stack": res.Stack, "op": res.Op}) +} diff --git a/controller/cmd/controller/restore_record_wiring_test.go b/controller/cmd/controller/restore_record_wiring_test.go new file mode 100644 index 0000000..e791701 --- /dev/null +++ b/controller/cmd/controller/restore_record_wiring_test.go @@ -0,0 +1,82 @@ +package main + +import ( + "go/ast" + "go/parser" + "go/token" + "io" + "log" + "testing" + + "gitea.dooplex.hu/admin/felhom-controller/internal/backup" +) + +type recordedPush struct{ eventType, severity, message string } + +type pushRecorder struct{ got []recordedPush } + +func (r *pushRecorder) PushEvent(et, sev, msg string, _ interface{}) { + r.got = append(r.got, recordedPush{et, sev, msg}) +} + +// R-550 — the functions main() calls: a restore left in flight is found at startup and raised as +// restore_interrupted (warning), exactly once. +// +// RED-PROOF: make loadRestoreRecordAtStartup skip LoadRestoreRecord → nothing is pushed → +// "an interrupted restore raised no restore_interrupted". +func TestRestoreRecordAtStartup_RaisesInterruptedOnce(t *testing.T) { + dir := t.TempDir() + lg := log.New(io.Discard, "", 0) + + before := new(backup.Manager) + before.SetRestoreRecordPath(dir + "/" + restoreRecordFileName) + before.BeginRestoreOp("restore", "gokapi") // the box stops here + + rec := &pushRecorder{} + reportInterruptedRestore(rec, loadRestoreRecordAtStartup(new(backup.Manager), dir, lg)) + if len(rec.got) != 1 || rec.got[0].eventType != "restore_interrupted" || rec.got[0].severity != "warning" { + t.Fatalf("an interrupted restore raised no restore_interrupted (warning): %+v", rec.got) + } + + again := &pushRecorder{} + reportInterruptedRestore(again, loadRestoreRecordAtStartup(new(backup.Manager), dir, lg)) + if len(again.got) != 0 { + t.Fatalf("the same interruption was raised again on the next start: %+v", again.got) + } +} + +// The seam must be CALLED from main(): both functions, in that order. +func TestMainWiresRestoreRecord(t *testing.T) { + fset := token.NewFileSet() + f, err := parser.ParseFile(fset, "main.go", nil, 0) + if err != nil { + t.Fatalf("parse main.go: %v", err) + } + pos := map[string]token.Pos{} + for _, d := range f.Decls { + fn, ok := d.(*ast.FuncDecl) + if !ok || fn.Name.Name != "main" { + continue + } + ast.Inspect(fn.Body, func(n ast.Node) bool { + if c, ok := n.(*ast.CallExpr); ok { + if id, ok := c.Fun.(*ast.Ident); ok { + if id.Name == "loadRestoreRecordAtStartup" || id.Name == "reportInterruptedRestore" { + if _, seen := pos[id.Name]; !seen { + pos[id.Name] = c.Pos() + } + } + } + } + return true + }) + } + load, okL := pos["loadRestoreRecordAtStartup"] + report, okR := pos["reportInterruptedRestore"] + if !okL || !okR { + t.Fatalf("main() does not call both restore-record functions (load=%t report=%t) — the persistence is an inert seam", okL, okR) + } + if load > report { + t.Fatalf("main() reports before it loads — nothing would ever be reported") + } +} diff --git a/controller/internal/backup/backup.go b/controller/internal/backup/backup.go index 7e9707f..442a672 100644 --- a/controller/internal/backup/backup.go +++ b/controller/internal/backup/backup.go @@ -264,6 +264,9 @@ type Manager struct { opStack string opStartedAt time.Time opLast *RestoreOpResult + // R-550: persistence of the above (restore_record.go) and the per-app interrupted notices. + opRecordPath string + opInterrupted map[string]RestoreOpResult // Cached status for page rendering (refreshed periodically) cachedStatus *FullBackupStatus diff --git a/controller/internal/backup/opstatus.go b/controller/internal/backup/opstatus.go index c2a5ce3..8667a4a 100644 --- a/controller/internal/backup/opstatus.go +++ b/controller/internal/backup/opstatus.go @@ -6,8 +6,8 @@ import "time" // backups page can show a progress banner (running → success/failure) instead of blocking the HTTP // request until the restore completes. It is display-only and mutex-guarded on the Manager's `mu`; // it does NOT gate concurrency (that stays the restore functions' internal single-flight acquire). -// In-memory only — lost on a controller restart (same precedent as notification cooldowns); a page -// load mid-op after a restart simply shows no banner. +// PERSISTED since v0.246.0 (R-550, operator ruling 2026-09-17 — a reversal of the original in-memory +// choice, for the restore record only; notification cooldowns stay in memory). See restore_record.go. // RestoreOpResult is the terminal record of the most recent restore op. type RestoreOpResult struct { @@ -16,6 +16,9 @@ type RestoreOpResult struct { OK bool `json:"ok"` Message string `json:"message"` FinishedAt time.Time `json:"finished_at"` + // Interrupted marks a restore that was still in flight when the controller stopped (R-550): found + // at the next start, recorded as a failure with RestoreInterruptedMessage. + Interrupted bool `json:"interrupted,omitempty"` } // RestoreResultWindow bounds how long a finished restore still counts as "what just happened". @@ -54,6 +57,9 @@ func (m *Manager) BeginRestoreOp(op, stack string) { m.opName = op m.opStack = stack m.opStartedAt = time.Now() + // A new restore of this app supersedes its interrupted notice (R-550). + delete(m.opInterrupted, stack) + m.persistRestoreRecordLocked() } // EndRestoreOp records the terminal result (called from the goroutine on completion, success or @@ -69,6 +75,7 @@ func (m *Manager) EndRestoreOp(ok bool, message string) { FinishedAt: time.Now(), } m.opRunning = false + m.persistRestoreRecordLocked() } // RestoreStatus returns a deep copy of the current restore op-status for the page/API. diff --git a/controller/internal/backup/restore_record.go b/controller/internal/backup/restore_record.go new file mode 100644 index 0000000..7bd72e8 --- /dev/null +++ b/controller/internal/backup/restore_record.go @@ -0,0 +1,133 @@ +package backup + +import ( + "encoding/json" + "os" + "sort" + "time" +) + +// R-550 (operator ruling "fix", 2026-09-17) — the restore record survives a restart. +// +// A DESIGN REVERSED, AND RECORDED AS SUCH. opstatus.go was deliberately in-memory ("same precedent as +// notification cooldowns"). Chaos night round 10 measured the cost: a restore accepted at 23:26:08Z, +// the box hard-reset four seconds later, and afterwards the status answered the Go zero value — the +// household had pressed a button, been told it started, and could never learn whether it finished. +// The operator reversed the choice for the RESTORE RECORD ONLY; notification cooldowns stay in memory. +// +// Shape: one JSON file in the controller's state directory (DataDir, beside settings.json), written +// atomically (atomicWrite: tmp + rename) at BOTH ends of an op. At startup a record still marked +// running is, by construction, a restore nothing is running any more: it becomes a terminal failure +// (Interrupted, RestoreInterruptedMessage) and a per-app notice that stays until that app's next +// restore. LoadRestoreRecord reports the conversion ONCE, so the caller raises restore_interrupted once. +// +// No path set (tests that build a bare Manager, a box with backup disabled) = the old in-memory +// behaviour, silently: persistence is a property of the wired controller, not of every Manager. + +// RestoreInterruptedMessage is what the household reads when the box stopped mid-restore. +const RestoreInterruptedMessage = "A visszaállítás megszakadt (a doboz újraindult) — indítsd el újra." + +// restoreRecordFile is the on-disk shape. +type restoreRecordFile struct { + Running bool `json:"running"` + Op string `json:"op,omitempty"` + Stack string `json:"stack,omitempty"` + StartedAt time.Time `json:"started_at,omitempty"` + Last *RestoreOpResult `json:"last,omitempty"` + Interrupted map[string]RestoreOpResult `json:"interrupted,omitempty"` +} + +// SetRestoreRecordPath wires persistence. Call before LoadRestoreRecord and before serving requests. +func (m *Manager) SetRestoreRecordPath(path string) { + m.mu.Lock() + defer m.mu.Unlock() + m.opRecordPath = path +} + +// LoadRestoreRecord restores the record at startup. It returns the restore that was interrupted by the +// stop — non-nil exactly once per interruption — so the caller can log it and raise restore_interrupted. +func (m *Manager) LoadRestoreRecord() *RestoreOpResult { + m.mu.Lock() + defer m.mu.Unlock() + if m.opRecordPath == "" { + return nil + } + b, err := os.ReadFile(m.opRecordPath) + if err != nil { + if !os.IsNotExist(err) && m.logger != nil { + m.logger.Printf("[WARN] [backup] restore record unreadable (%v) — starting with no restore history", err) + } + return nil + } + var rec restoreRecordFile + if err := json.Unmarshal(b, &rec); err != nil { + if m.logger != nil { + m.logger.Printf("[WARN] [backup] restore record corrupt (%v) — starting with no restore history", err) + } + return nil + } + m.opLast = rec.Last + m.opInterrupted = rec.Interrupted + if !rec.Running || rec.Stack == "" { + if m.logger != nil && len(m.opInterrupted) > 0 { + m.logger.Printf("[INFO] [backup] restore record loaded: %d interrupted restore notice(s) still shown", len(m.opInterrupted)) + } + return nil + } + res := RestoreOpResult{ + Op: rec.Op, Stack: rec.Stack, OK: false, + Message: RestoreInterruptedMessage, FinishedAt: time.Now(), Interrupted: true, + } + m.opLast = &res + if m.opInterrupted == nil { + m.opInterrupted = map[string]RestoreOpResult{} + } + m.opInterrupted[rec.Stack] = res + m.opRunning = false + if m.logger != nil { + m.logger.Printf("[WARN] [backup] restore of %s (%s, started %s) was INTERRUPTED by a controller stop — recorded as failed; the household is told to run it again", + rec.Stack, rec.Op, rec.StartedAt.Format(time.RFC3339)) + } + m.persistRestoreRecordLocked() + out := res + return &out +} + +// InterruptedRestore reports the standing interrupted-restore notice for one app, if any. +func (m *Manager) InterruptedRestore(stack string) (RestoreOpResult, bool) { + m.mu.Lock() + defer m.mu.Unlock() + r, ok := m.opInterrupted[stack] + return r, ok +} + +// InterruptedRestores lists every standing notice, sorted by app, for the restore page. +func (m *Manager) InterruptedRestores() []RestoreOpResult { + m.mu.Lock() + defer m.mu.Unlock() + out := make([]RestoreOpResult, 0, len(m.opInterrupted)) + for _, r := range m.opInterrupted { + out = append(out, r) + } + sort.Slice(out, func(i, j int) bool { return out[i].Stack < out[j].Stack }) + return out +} + +// persistRestoreRecordLocked writes the current op-status. Caller holds m.mu. A failed write is logged +// and the in-memory status carries on — the page still works for this process's lifetime. +func (m *Manager) persistRestoreRecordLocked() { + if m.opRecordPath == "" { + return + } + rec := restoreRecordFile{ + Running: m.opRunning, Op: m.opName, Stack: m.opStack, StartedAt: m.opStartedAt, + Last: m.opLast, Interrupted: m.opInterrupted, + } + b, err := json.Marshal(rec) + if err == nil { + err = atomicWrite(m.opRecordPath, b, 0o600) + } + if err != nil && m.logger != nil { + m.logger.Printf("[WARN] [backup] could not persist the restore record to %s: %v (kept in memory)", m.opRecordPath, err) + } +} diff --git a/controller/internal/backup/restore_record_test.go b/controller/internal/backup/restore_record_test.go new file mode 100644 index 0000000..d9dcd5a --- /dev/null +++ b/controller/internal/backup/restore_record_test.go @@ -0,0 +1,96 @@ +package backup + +import ( + "io" + "log" + "path/filepath" + "strings" + "testing" +) + +func recordManager(path string) *Manager { + m := &Manager{logger: log.New(io.Discard, "", 0)} + m.SetRestoreRecordPath(path) + return m +} + +// R-550 (operator ruling "fix", 2026-09-17). Chaos night round 10: a restore was accepted and the box +// was hard-reset four seconds later. Afterwards /api/backup/restore-status answered the Go zero value +// and nothing told the household whether the restore had finished — the op-status was in-memory only. +// +// The consequence asserted: a restore that was IN FLIGHT when the controller stopped is, after the +// next start, a FAILED result for that app saying it was interrupted — and loading it reports it once, +// so the caller can raise restore_interrupted. +// +// RED-PROOF: keep the op-status in memory (no record written / read) → the second Manager's status is +// blank → "after a restart the interrupted restore left no trace". +func TestRestoreRecord_InterruptedRestoreSurvivesRestart(t *testing.T) { + path := filepath.Join(t.TempDir(), "restore-status.json") + m1 := recordManager(path) + m1.BeginRestoreOp("restore", "gokapi") + // the box stops here — no EndRestoreOp + + m2 := recordManager(path) + got := m2.LoadRestoreRecord() + st := m2.RestoreStatus() + if got == nil || st.Last == nil { + t.Fatalf("after a restart the interrupted restore left no trace: loaded=%v status=%+v", got, st) + } + if st.Running { + t.Fatalf("a restore from before the restart is still shown as RUNNING — nothing is running it: %+v", st) + } + if st.Last.OK || !st.Last.Interrupted || st.Last.Stack != "gokapi" || st.Last.Op != "restore" { + t.Fatalf("interrupted record = %+v, want a failed, interrupted restore of gokapi", st.Last) + } + // ASCII fragment of the Hungarian sentence, with a negative control. + if !strings.Contains(st.Last.Message, "megszakadt") || strings.Contains(st.Last.Message, "zzzz-not-present") { + t.Fatalf("message %q does not say the restore was interrupted", st.Last.Message) + } + if rec, ok := m2.InterruptedRestore("gokapi"); !ok || !rec.Interrupted { + t.Fatalf("InterruptedRestore(gokapi) = %+v, %v — the restore page would not show it", rec, ok) + } + // Reported ONCE: a third start must not raise the event again for the same interruption. + m3 := recordManager(path) + if again := m3.LoadRestoreRecord(); again != nil { + t.Fatalf("the same interruption was reported again on the next start: %+v", again) + } + if _, ok := m3.InterruptedRestore("gokapi"); !ok { + t.Fatalf("the interrupted notice vanished on the next start — it must stay until gokapi is restored again") + } +} + +// Control: a restore that FINISHED is still that result after a restart — never re-labelled interrupted. +func TestRestoreRecord_FinishedRestoreStaysFinished(t *testing.T) { + path := filepath.Join(t.TempDir(), "restore-status.json") + m1 := recordManager(path) + m1.BeginRestoreOp("restore", "mealie") + m1.EndRestoreOp(true, "kész") + + m2 := recordManager(path) + if got := m2.LoadRestoreRecord(); got != nil { + t.Fatalf("a finished restore was reported as interrupted: %+v", got) + } + st := m2.RestoreStatus() + if st.Last == nil || !st.Last.OK || st.Last.Interrupted || st.Last.Message != "kész" { + t.Fatalf("finished record after restart = %+v, want ok 'kész'", st.Last) + } +} + +// The notice lasts until that app's NEXT restore — and only that app's. +func TestRestoreRecord_NextRestoreOfThatAppClearsTheNotice(t *testing.T) { + path := filepath.Join(t.TempDir(), "restore-status.json") + m1 := recordManager(path) + m1.BeginRestoreOp("restore", "gokapi") + m2 := recordManager(path) + m2.LoadRestoreRecord() + + m2.BeginRestoreOp("restore", "mealie") // a different app + m2.EndRestoreOp(true, "kész") + if _, ok := m2.InterruptedRestore("gokapi"); !ok { + t.Fatalf("restoring ANOTHER app cleared gokapi's interrupted notice") + } + m2.BeginRestoreOp("restore", "gokapi") // the same app, again + if _, ok := m2.InterruptedRestore("gokapi"); ok { + t.Fatalf("starting gokapi's restore again did not clear its interrupted notice") + } +} diff --git a/controller/internal/web/escrow_banner.go b/controller/internal/web/escrow_banner.go index 780a5c3..559951f 100644 --- a/controller/internal/web/escrow_banner.go +++ b/controller/internal/web/escrow_banner.go @@ -53,6 +53,11 @@ func (s *Server) escrowBannerVisible(r *http.Request) bool { if r == nil || !s.hasAdminSession(r) || !s.escrowPaused() { return false } + // R-546: while the agent says the ceremony cannot run yet, do not urge the household into it. + // Unknown readiness keeps the bar (fail loud). + if ready, known := s.escrowReadiness(r.Context(), false); known && !ready { + return false + } if c, err := r.Cookie(escrowBannerCookie); err == nil && c.Value == "1" { return false } diff --git a/controller/internal/web/escrow_handlers.go b/controller/internal/web/escrow_handlers.go index 2200147..a65fb5b 100644 --- a/controller/internal/web/escrow_handlers.go +++ b/controller/internal/web/escrow_handlers.go @@ -127,6 +127,15 @@ func (s *Server) escrowWizardPageHandler(w http.ResponseWriter, r *http.Request) "OffboxConfigured": s.backupMgr != nil && s.backupMgr.OffboxConfigured(), "AgentSupported": escrowAgentSupported(agentVer), } + // R-546: a paused box whose agent says the ceremony cannot run yet gets the waiting card, not a red + // checklist and a start form that would refuse. Only for the FIRST ceremony — an escrowed box's + // re-ceremony keeps the checklist, where a red row is a real fault to show. + if !escrowed { + if ready, known := s.escrowReadiness(r.Context(), true); known && !ready { + data["EscrowNotReady"] = true + data["EscrowNotReadyMessage"] = escrowNotReadyMessage + } + } s.executeTemplate(w, r, "backups_escrow", data) } @@ -188,6 +197,16 @@ func (s *Server) escrowStartAPIHandler(w http.ResponseWriter, r *http.Request) { } s.clearAuthFailures(ip) + // (1b) R-546: readiness. The page hides the start form while the agent's preflight is red, but a + // direct POST (the path chaos night used) ran the ceremony and returned the agent's raw stderr. + // Refuse here, BEFORE anything is staged or started. Unknown readiness does not refuse: the + // agent-conn and version gates below answer that case, as before. + if ready, known := s.escrowReadiness(r.Context(), true); known && !ready { + s.logger.Printf("[INFO] [web] escrow start refused: the box is not ready yet (agent preflight not ok)") + escrowJSON(w, http.StatusConflict, nil, escrowNotReadyMessage) + return + } + // (2) Re-stage-first (only when offsite is configured — Scenario C boxes skip it). if s.backupMgr != nil && s.backupMgr.OffboxConfigured() { if err := s.escrowStage(r.Context()); err != nil { diff --git a/controller/internal/web/escrow_readiness.go b/controller/internal/web/escrow_readiness.go new file mode 100644 index 0000000..8f9d388 --- /dev/null +++ b/controller/internal/web/escrow_readiness.go @@ -0,0 +1,104 @@ +package web + +import ( + "context" + "sort" + "strings" + "sync" + "time" +) + +// R-546 — is the box READY to run the escrow ceremony? +// +// Measured 2026-09-16/17 (chaos night Phase 0): after a fresh bind the box has no PBS storage for ~17 +// minutes. The agent's preflight is red, the ceremony refuses, and the R-543 reminder bar (v0.245.0) +// was on every page urging the household into it. The escrow page itself already hid its start form +// behind a red checklist — but it said nothing about WAITING, and a direct POST /api/escrow/start +// (the path chaos night used) ran the ceremony and returned the agent's raw "-storage" stderr. +// +// THE DEFINITION IS THE AGENT'S OWN `ok`: every blocking preflight item (pbs_storage_id, dr_tier, +// age_binary, hub_upload, sudo_grant; staged_secret is informational). Not a controller-side copy of +// one of those items — the escrow.pbs_storage_id reading of R-546 was the SYMPTOM that box showed, and +// a copy of one item is a second definition that drifts when the agent grows a sixth. +// +// Three answers, and UNKNOWN keeps today's behaviour (the bar shows): an unreachable agent must not +// silence a reminder that R-543 made loud on purpose. +// +// Cost: the bar hangs off every page render, so the answer is cached for escrowReadyTTL, and a probe +// is only ever made while the box is paused (escrowPaused gates it) — a finished box never asks. + +const ( + escrowReadyTTL = 60 * time.Second + escrowReadyTimeout = 3 * time.Second +) + +type escrowReadinessCache struct { + mu sync.Mutex + ready bool + known bool + checkedAt time.Time + lastState string // for transition-only INFO logging +} + +// escrowReadiness returns (ready, known). fresh=true skips the cache (the escrow page and the start +// API, where one request justifies one probe); fresh=false is the bar's cached read. +func (s *Server) escrowReadiness(ctx context.Context, fresh bool) (ready, known bool) { + c := &s.escrowReady + c.mu.Lock() + defer c.mu.Unlock() + if !fresh && !c.checkedAt.IsZero() && time.Since(c.checkedAt) < escrowReadyTTL { + return c.ready, c.known + } + pctx, cancel := context.WithTimeout(ctx, escrowReadyTimeout) + defer cancel() + ready, known, why := s.probeEscrowReadiness(pctx) + c.ready, c.known, c.checkedAt = ready, known, time.Now() + + state := "unknown" + if known && ready { + state = "ready" + } else if known { + state = "not-ready" + } + if state != c.lastState { + switch state { + case "ready": + s.logger.Printf("[INFO] [web] escrow readiness: READY (agent preflight ok) — the recovery-code reminder is shown") + case "not-ready": + s.logger.Printf("[INFO] [web] escrow readiness: NOT READY (%s) — reminder held back; the escrow page says the box is still preparing", why) + default: + s.logger.Printf("[INFO] [web] escrow readiness: UNKNOWN (%s) — reminder shown (fail loud)", why) + } + c.lastState = state + } else if s.isDebug() { + s.logger.Printf("[DEBUG] [web] escrow readiness: %s (%s)", state, why) + } + return ready, known +} + +// probeEscrowReadiness asks the agent. Never an error to the caller: failure is "unknown". +func (s *Server) probeEscrowReadiness(ctx context.Context) (ready, known bool, why string) { + agent, err := s.escrowAgentConn() + if err != nil { + return false, false, "agent unavailable: " + err.Error() + } + pf, err := agent.EscrowPreflight(ctx) + if err != nil { + return false, false, "preflight failed: " + err.Error() + } + if pf.OK { + return true, true, "preflight ok" + } + var failing []string + for _, it := range pf.Items { + if !it.OK && it.ID != "staged_secret" { + failing = append(failing, it.ID) + } + } + sort.Strings(failing) + return false, true, "failing: " + strings.Join(failing, ", ") +} + +// escrowNotReadyMessage is the one Hungarian sentence for "wait" — the page card and the API refusal +// say the same thing. +const escrowNotReadyMessage = "A doboz még készül — a távoli mentés kulcsát pár perc múlva tudod létrehozni. Ez az oldal magától frissül." diff --git a/controller/internal/web/escrow_wizard_test.go b/controller/internal/web/escrow_wizard_test.go index 2e755d8..d7f6486 100644 --- a/controller/internal/web/escrow_wizard_test.go +++ b/controller/internal/web/escrow_wizard_test.go @@ -87,6 +87,9 @@ func newEscrowWizardHarness(t *testing.T) *escrowWizardHarness { version: "0.88.0", startResp: agentapi.EscrowCeremonyStartResponse{JobID: "escrow-1", Phase: "running"}, startStatus: http.StatusAccepted, + // R-546 (v0.246.0): the start API now asks the agent's preflight first; this harness is a + // READY box. A not-ready box is r546_escrow_readiness_test.go. + pf: agentapi.EscrowPreflightResponse{OK: true}, } h.s = &Server{cfg: cfg, backupMgr: h.m, settings: sett, logger: lg} h.s.escrowAgentFn = func() (escrowAgent, error) { return h.agent, nil } @@ -131,8 +134,8 @@ func TestEscrowStart_StagesBeforeTrigger(t *testing.T) { if w.Code != http.StatusOK { t.Fatalf("start: got %d (%s)", w.Code, w.Body.String()) } - if got := strings.Join(h.order, ","); got != "stage,start" { - t.Fatalf("call order = %q, want stage BEFORE start (a ceremony without the staged secret mints a hash-less blob)", got) + if got := strings.Join(h.order, ","); got != "preflight,stage,start" { // R-546: readiness first + t.Fatalf("call order = %q, want preflight, then stage BEFORE start (a ceremony without the staged secret mints a hash-less blob)", got) } if !strings.Contains(w.Body.String(), "escrow-1") { t.Fatalf("response lacks the job id: %s", w.Body.String()) @@ -147,8 +150,8 @@ func TestEscrowStart_ReceremonyStagesToo(t *testing.T) { if w := postStart(t, h.s, wizardPassword); w.Code != http.StatusOK { t.Fatalf("re-ceremony start: got %d", w.Code) } - if got := strings.Join(h.order, ","); got != "stage,start" { - t.Fatalf("re-ceremony call order = %q, want stage,start", got) + if got := strings.Join(h.order, ","); got != "preflight,stage,start" { // R-546: readiness first + t.Fatalf("re-ceremony call order = %q, want preflight,stage,start", got) } } @@ -159,8 +162,8 @@ func TestEscrowStart_NoOffboxSkipsStaging(t *testing.T) { if w := postStart(t, h.s, wizardPassword); w.Code != http.StatusOK { t.Fatalf("start: got %d", w.Code) } - if got := strings.Join(h.order, ","); got != "start" { - t.Fatalf("call order = %q, want start only (no staging without offbox)", got) + if got := strings.Join(h.order, ","); got != "preflight,start" { // R-546: readiness first + t.Fatalf("call order = %q, want preflight then start only (no staging without offbox)", got) } } diff --git a/controller/internal/web/handlers.go b/controller/internal/web/handlers.go index 153821d..1f05d60 100644 --- a/controller/internal/web/handlers.go +++ b/controller/internal/web/handlers.go @@ -1067,6 +1067,10 @@ func (s *Server) backupsAppsHandler(w http.ResponseWriter, r *http.Request) { // restore-to-verify list and the .fab export/import loop. func (s *Server) backupsRestoreHandler(w http.ResponseWriter, r *http.Request) { data := s.backupsCommonData("backups-restore", "Biztonsági mentés — Visszaállítás", r) + // R-550: restores the box was running when it stopped — shown until that app is restored again. + if s.backupMgr != nil { + data["InterruptedRestores"] = s.backupMgr.InterruptedRestores() + } s.backupsOffboxData(data) // restore-to-verify lists the offbox-toggled apps // Full-restore two-step reveal (§7.2): after the size+headroom prepare step, offboxRestoreHandler // redirects here with the app + human size so the confirm section can show the size BEFORE starting. diff --git a/controller/internal/web/r546_escrow_readiness_test.go b/controller/internal/web/r546_escrow_readiness_test.go new file mode 100644 index 0000000..ed28934 --- /dev/null +++ b/controller/internal/web/r546_escrow_readiness_test.go @@ -0,0 +1,113 @@ +package web + +import ( + "errors" + "net/http" + "strings" + "testing" + + "gitea.dooplex.hu/admin/felhom-controller/internal/agentapi" +) + +// ── R-546 — do not send the household to a ceremony the box cannot run yet ──────────────────────── +// +// Measured 2026-09-16/17 (chaos night Phase 0): for ~17 minutes after the bind the box has no PBS +// storage, the agent's preflight is red, and the ceremony refuses. Meanwhile the R-543 bar is on every +// page urging the household there. Readiness is the AGENT'S OWN `ok` — every blocking preflight item +// (pbs_storage_id, dr_tier, age_binary, hub_upload, sudo_grant), never a copy of one of them. + +const escrowNotReadyMarker = `id="escrow-not-ready"` + +func readinessAgent(ok bool) *fakeEscrowAgent { + order := []string{} + pf := agentapi.EscrowPreflightResponse{OK: ok} + if !ok { + pf.Items = []agentapi.EscrowPreflightItem{{ID: "pbs_storage_id", OK: false, Detail: "escrow.pbs_storage_id not configured"}} + } + return &fakeEscrowAgent{order: &order, version: "0.132.0", pf: pf} +} + +// RED-PROOF: remove the readiness check from escrowBannerVisible → the bar renders on a box that +// cannot run the ceremony → "the bar urges a ceremony the box cannot run yet". +func TestR546_NotReadyBoxHoldsTheBar(t *testing.T) { + s := escrowServer(t, "pending") + a := readinessAgent(false) + s.escrowAgentFn = func() (escrowAgent, error) { return a, nil } + + body := getPage(t, s, "/dashboard").Body.String() + if strings.Contains(body, escrowBarSentence) { + t.Fatal("R-546: the bar urges a ceremony the box cannot run yet (agent preflight is red)") + } +} + +// Control: once the agent says ready, the R-543 bar is back. +func TestR546_ReadyBoxShowsTheBar(t *testing.T) { + s := escrowServer(t, "pending") + a := readinessAgent(true) + s.escrowAgentFn = func() (escrowAgent, error) { return a, nil } + if body := getPage(t, s, "/dashboard").Body.String(); !strings.Contains(body, escrowBarSentence) { + t.Fatal("a READY, paused box no longer shows the R-543 bar") + } +} + +// Fail loud: readiness UNKNOWN (agent unreachable) keeps the reminder — silence is not a safe default. +func TestR546_UnknownReadinessStillAsks(t *testing.T) { + s := escrowServer(t, "pending") + s.escrowAgentFn = func() (escrowAgent, error) { return nil, errors.New("agent unreachable") } + if body := getPage(t, s, "/dashboard").Body.String(); !strings.Contains(body, escrowBarSentence) { + t.Fatal("an unknown readiness hid the reminder — an unreachable agent must not silence R-543") + } +} + +// The page itself: a calm waiting card, no start form, and none of the agent's raw diagnostics. +// RED-PROOF: never set EscrowNotReady → the checklist + start form render → "start form offered". +func TestR546_EscrowPageWaitsWhenNotReady(t *testing.T) { + s := escrowServer(t, "pending") + a := readinessAgent(false) + s.escrowAgentFn = func() (escrowAgent, error) { return a, nil } + + rec := getPage(t, s, "/backup/escrow") + if rec.Code != 200 { + t.Fatalf("GET /backup/escrow = %d: %s", rec.Code, rec.Body.String()) + } + body := rec.Body.String() + if !strings.Contains(body, escrowNotReadyMarker) || !strings.Contains(body, "A doboz m") { + t.Fatal("R-546: the escrow page does not tell the household the box is still preparing") + } + if strings.Contains(body, `id="start-form"`) { + t.Fatal("R-546: the start form is offered on a box whose ceremony would refuse") + } + // The exact texts chaos night met: the preflight detail, and the ceremony stderr ("…escrow-create + // requires -storage …"). NOT a bare "-storage": the menu carries id="nav-group-storage". + for _, raw := range []string{"pbs_storage_id not configured", "requires -storage"} { + if strings.Contains(body, raw) { + t.Fatalf("R-546: the agent's raw diagnostic %q reached the household's page", raw) + } + } + if strings.Contains(body, "zzzz-not-present") { + t.Fatal("negative control matched") + } +} + +// The direct path — the one that actually produced the stderr in chaos night: POST /api/escrow/start +// on a not-ready box is refused BEFORE anything is staged or started, in Hungarian. +// RED-PROOF: remove the readiness gate from escrowStartAPIHandler → 200 and "stage,start" run. +func TestR546_StartRefusedWhenNotReady(t *testing.T) { + h := newEscrowWizardHarness(t) + h.configureOffbox(t, "pending") + h.agent.pf = agentapi.EscrowPreflightResponse{OK: false, + Items: []agentapi.EscrowPreflightItem{{ID: "pbs_storage_id", OK: false}}} + + w := postStart(t, h.s, wizardPassword) + if w.Code != http.StatusConflict { + t.Fatalf("start on a not-ready box = %d (%s), want 409", w.Code, w.Body.String()) + } + for _, op := range h.order { + if op == "stage" || op == "start" { + t.Fatalf("a not-ready box still ran %q (order %v) — the refusal must come first", op, h.order) + } + } + if !strings.Contains(w.Body.String(), "A doboz m") { + t.Fatalf("the refusal is not the Hungarian waiting sentence: %s", w.Body.String()) + } +} diff --git a/controller/internal/web/r550_restore_interrupted_test.go b/controller/internal/web/r550_restore_interrupted_test.go new file mode 100644 index 0000000..b8f88b0 --- /dev/null +++ b/controller/internal/web/r550_restore_interrupted_test.go @@ -0,0 +1,57 @@ +package web + +import ( + "path/filepath" + "strings" + "testing" + "time" +) + +// ── R-550 — an interrupted restore is SHOWN on the restore page ───────────────────────────────── +// +// Chaos night round 10: a restore interrupted by a hard reset left the page blank. The record now +// survives (internal/backup/restore_record.go); these drive the REAL page through ServeHTTP, so they +// bite on the wiring — a record that exists but that no page renders is the inert-seam shape. + +const interruptedCardFragment = "Megszakadt vissza" // ASCII fragment of the card heading + +// RED-PROOF: drop `data["InterruptedRestores"]` from backupsRestoreHandler → the card never renders → +// "the restore page does not show the interrupted restore". +func TestR550_RestorePageShowsInterruptedRestore(t *testing.T) { + s := newDashboardServer(t, time.Time{}) + if s.backupMgr == nil { + t.Fatal("fixture invalid: no backup manager, so nothing here measures the restore page") + } + s.backupMgr.SetRestoreRecordPath(filepath.Join(t.TempDir(), "restore-status.json")) + s.backupMgr.BeginRestoreOp("restore", "gokapi") // the box stops here … + if s.backupMgr.LoadRestoreRecord() == nil { // … and this is the next start + t.Fatal("fixture invalid: the record did not convert to an interrupted restore") + } + + rec := getPage(t, s, "/backups/restore") + if rec.Code != 200 { + t.Fatalf("GET /backups/restore = %d: %s", rec.Code, rec.Body.String()) + } + body := rec.Body.String() + if !strings.Contains(body, interruptedCardFragment) || !strings.Contains(body, "megszakadt") { + t.Fatalf("the restore page does not show the interrupted restore — the household is left not knowing, which is R-550") + } + if !strings.Contains(body, "gokapi") { + t.Fatalf("the interrupted card does not name the app") + } + if strings.Contains(body, "zzzz-not-present") { + t.Fatal("negative control matched — the fragment search is broken") + } +} + +// The other branch: a box with no interrupted restore shows no such card. +func TestR550_NoInterruptedRestoreNoCard(t *testing.T) { + s := newDashboardServer(t, time.Time{}) + rec := getPage(t, s, "/backups/restore") + if rec.Code != 200 { + t.Fatalf("GET /backups/restore = %d", rec.Code) + } + if strings.Contains(rec.Body.String(), interruptedCardFragment) { + t.Fatalf("a box with no interrupted restore shows the interrupted card") + } +} diff --git a/controller/internal/web/restore_wizard.go b/controller/internal/web/restore_wizard.go index 524349e..f049ba9 100644 --- a/controller/internal/web/restore_wizard.go +++ b/controller/internal/web/restore_wizard.go @@ -289,6 +289,10 @@ func (s *Server) backupsRestoreWizardHandler(w http.ResponseWriter, r *http.Requ // survives a reload, which the flash does not. if in.HasRecentResult { data["LastResult"] = st.Last + } else if ir, ok := s.backupMgr.InterruptedRestore(app); ok { + // R-550: a restore of THIS app the box was running when it stopped — shown past the recency + // window, until the app is restored again. + data["LastResult"] = &ir } s.executeTemplate(w, r, "backups_restore_wizard", data) diff --git a/controller/internal/web/server.go b/controller/internal/web/server.go index fc1aba8..e6e48c0 100644 --- a/controller/internal/web/server.go +++ b/controller/internal/web/server.go @@ -101,6 +101,8 @@ type Server struct { escrowAgentFn func() (escrowAgent, error) escrowStageFn func(ctx context.Context) error escrowStaleFn func() bool + // R-546: the cached agent-preflight readiness behind the reminder bar (escrow_readiness.go). + escrowReady escrowReadinessCache // escrowSealedAtFn (v0.200.0, R-193) reports WHEN the hub's sealed recovery package was created — // the one non-secret fact the recovery screen may state before a code is entered. Wired via diff --git a/controller/internal/web/templates/backups_escrow.html b/controller/internal/web/templates/backups_escrow.html index bf28168..cdc059d 100644 --- a/controller/internal/web/templates/backups_escrow.html +++ b/controller/internal/web/templates/backups_escrow.html @@ -13,7 +13,27 @@ {{.Domain}} -{{if not .AgentSupported}} +{{if .EscrowNotReady}} + +
Ehhez nem kell semmit tenned. Ha fél óra múlva is ezt látod, szólj az üzemeltetőnek.
+A doboz újraindult, miközben a visszaállítás futott, ezért nem tudjuk, hogy befejeződött-e. Indítsd el újra lent, ugyanabból a mentésből.
+Befejezve: {{fmtTime .FinishedAt}}. Ha szeretnéd, alább újra indíthatsz egy visszaállítást.
+{{if .Interrupted}}Észlelve{{else}}Befejezve{{end}}: {{fmtTime .FinishedAt}}. Ha szeretnéd, alább újra indíthatsz egy visszaállítást.