v0.246.0: an interrupted restore is told; the recovery-code reminder waits until the box can take it
gates / gates (push) Successful in 15s
gates / gates (push) Successful in 15s
MinAgent: 0.131.0 (unchanged). Requires hub v0.117.0 for restore_interrupted. R-550 (operator ruling: fix). A design reversed and recorded: the restore op-status was in memory by choice. Now restore-status.json in DataDir, written atomically at both ends of an op. At startup a record still marked running becomes a failed, interrupted result kept per app until that app's next restore, shown on /backups/restore and the off-site wizard, and raised once as restore_interrupted. Cooldowns stay in memory. R-546. The R-543 reminder bar consults the agent's own preflight ok (every blocking item, not a copy of pbs_storage_id), cached 60 s, probed only while paused. /backup/escrow shows a waiting card that polls and reloads instead of red crosses and English diagnostics. POST /api/escrow/start refuses 409 before staging or starting - the direct path chaos night used. Unknown readiness keeps the bar. Red-proofs (each seen failing): restore record across restart; main() calls both startup functions; startup helper with loading skipped; restore page card; bar held back; waiting card; start refusal. go build/vet/test ./... green, 28 packages; controller_gates --fast all OK. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -1,3 +1,36 @@
|
|||||||
|
## v0.246.0 — an interrupted restore is told, and the recovery-code reminder waits until the box can take it (2026-09-17, R-550 / R-546)
|
||||||
|
|
||||||
|
**MinAgent: 0.131.0** (unchanged — the readiness check reads the escrow preflight, served since agent
|
||||||
|
0.88.0; nothing here needs agent 0.132.0). **Requires hub v0.117.0** for `restore_interrupted` to be
|
||||||
|
accepted; an older hub 400s that one event and nothing else changes.
|
||||||
|
|
||||||
|
- **R-550 (operator ruling "fix", 2026-09-17) — the restore record survives a restart.** A DESIGN
|
||||||
|
REVERSED: `opstatus.go` was in memory by choice ("same precedent as notification cooldowns"). Chaos
|
||||||
|
night round 10 measured the cost — a restore accepted, the box hard-reset four seconds later, and
|
||||||
|
`/api/backup/restore-status` answering the Go zero value; nobody could learn whether it finished.
|
||||||
|
Now `restore-status.json` in `DataDir`, written atomically at both ends of an op
|
||||||
|
(`internal/backup/restore_record.go`). At startup a record still marked running becomes a failed,
|
||||||
|
`Interrupted` result — „A visszaállítás megszakadt (a doboz újraindult) — indítsd el újra." — kept
|
||||||
|
per app until that app's next restore, shown as a „Megszakadt visszaállítás" card on
|
||||||
|
`/backups/restore` and in the off-site wizard's outcome card, and raised ONCE as
|
||||||
|
`restore_interrupted` (warning, for the household). Cooldowns stay in memory.
|
||||||
|
- **R-546 — the recovery-code reminder waits until the box can run the ceremony.** The R-543 bar now
|
||||||
|
consults the agent's OWN preflight `ok` (every blocking item — not a copy of `pbs_storage_id`), cached
|
||||||
|
60 s and probed only while paused; held back while not ready. `/backup/escrow` shows a waiting card
|
||||||
|
that polls and reloads itself instead of red crosses and the agent's English diagnostics.
|
||||||
|
`POST /api/escrow/start` refuses 409 with the same sentence BEFORE staging or starting — the direct
|
||||||
|
path chaos night used, which ran the ceremony and returned the raw `-storage` stderr. Unknown
|
||||||
|
readiness keeps the bar (fail loud). The transition is logged at INFO.
|
||||||
|
|
||||||
|
**Red-proofs, each seen failing then passing:** the restore record across a restart (blank status,
|
||||||
|
`StartedAt:0001-01-01` — the chaos-night symptom exactly); a finished restore stays finished; the next
|
||||||
|
restore of that app clears its notice; `main()` calls both startup functions (AST); the startup helper
|
||||||
|
with loading skipped pushes nothing; the restore page card with the handler line removed; the bar held
|
||||||
|
back (readiness check removed → bar shown); the waiting card (never set → checklist); the start refusal
|
||||||
|
(gate removed → 200 and `stage,start` ran). Controls: a ready box shows the bar; unknown readiness shows
|
||||||
|
the bar; a box with no interruption shows no card. The three existing start-order tests now expect
|
||||||
|
`preflight` first — their load-bearing assertion (stage BEFORE start) is unchanged.
|
||||||
|
|
||||||
## v0.245.0 — the household is asked for the recovery code, and the page says „szünetel" until then (2026-09-16, R-543)
|
## v0.245.0 — the household is asked for the recovery code, and the page says „szünetel" until then (2026-09-16, R-543)
|
||||||
|
|
||||||
**MinAgent: 0.131.0** (unchanged — nothing here needs a newer agent)
|
**MinAgent: 0.131.0** (unchanged — nothing here needs a newer agent)
|
||||||
|
|||||||
@@ -1333,6 +1333,24 @@ truncated.
|
|||||||
> Those figures do not extrapolate: the structure check's cost tracks the index, read-data's tracks
|
> Those figures do not extrapolate: the structure check's cost tracks the index, read-data's tracks
|
||||||
> the data.
|
> the data.
|
||||||
|
|
||||||
|
### The restore record survives a restart (v0.246.0, R-550)
|
||||||
|
|
||||||
|
The restore op-status (`internal/backup/opstatus.go`, served at `GET /api/backup/restore-status`) was
|
||||||
|
in memory only. Chaos night round 10 hard-reset a box four seconds into a restore; afterwards the
|
||||||
|
status was the Go zero value and nothing told the household whether the restore finished. **Operator
|
||||||
|
ruling 2026-09-17 reversed the in-memory choice for the restore record only** (notification cooldowns
|
||||||
|
stay in memory).
|
||||||
|
|
||||||
|
- `restore-status.json` in `DataDir` (beside `settings.json`), written atomically at **both** ends of an op.
|
||||||
|
- At startup (`loadRestoreRecordAtStartup`, before any page is served) a record still marked running
|
||||||
|
becomes a failed result — „A visszaállítás megszakadt (a doboz újraindult) — indítsd el újra." —
|
||||||
|
and a per-app notice that stays until **that app's** next restore.
|
||||||
|
- `restore_interrupted` (warning, for the household; hub v0.117.0) is pushed **once** per
|
||||||
|
interruption, after the notifier exists (`reportInterruptedRestore`). Best-effort like every
|
||||||
|
`PushEvent` (3 attempts, 3 s apart); the page notice does not depend on it.
|
||||||
|
- `/backups/restore` shows a „Megszakadt visszaállítás" card per interrupted app; the off-site
|
||||||
|
wizard's „Eredmény" card falls back to the interrupted record („Észlelve: …").
|
||||||
|
|
||||||
### Restore refusals (v0.226.0)
|
### Restore refusals (v0.226.0)
|
||||||
|
|
||||||
Three guards added on the off-site restore surface, all server-side:
|
Three guards added on the off-site restore surface, all server-side:
|
||||||
@@ -1585,6 +1603,15 @@ dismissal, back at the next visit, gone for good when the state is `escrowed`. I
|
|||||||
pages bypass that function, and a session check keeps it off the public guest share page. The pause
|
pages bypass that function, and a session check keeps it off the public guest share page. The pause
|
||||||
itself is UNCHANGED — it is the zero-knowledge escrow design, not a defect.
|
itself is UNCHANGED — it is the zero-knowledge escrow design, not a defect.
|
||||||
|
|
||||||
|
**…but only once the box can do it (v0.246.0, R-546).** For the first minutes after a bind the
|
||||||
|
agent's escrow preflight is not `ok` (typically no PBS storage yet — ~17 min measured). The bar now
|
||||||
|
consults the agent's own `ok` (`escrow_readiness.go`, cached 60 s, probed only while paused) and is
|
||||||
|
held back while the agent says not ready; `/backup/escrow` shows a waiting card („A doboz még készül
|
||||||
|
— … pár perc múlva …") that polls the preflight and reloads itself, instead of a red checklist; and
|
||||||
|
`POST /api/escrow/start` refuses 409 with the same sentence **before** staging or starting. Readiness
|
||||||
|
**unknown** (agent unreachable) keeps the bar — the fail-loud rule of R-543. A re-ceremony on an
|
||||||
|
escrowed box keeps the checklist, where a red row is a real fault.
|
||||||
|
|
||||||
#### Restore (`internal/backup/restore.go`)
|
#### Restore (`internal/backup/restore.go`)
|
||||||
|
|
||||||
Both **Tier 1** (restic) and **Tier 2** (rsync) restores are supported. All deployed apps
|
Both **Tier 1** (restic) and **Tier 2** (rsync) restores are supported. All deployed apps
|
||||||
|
|||||||
@@ -527,6 +527,8 @@ func main() {
|
|||||||
|
|
||||||
// --- Initialize backup manager (app-data only: DB dumps + Docker-volume tars) ---
|
// --- Initialize backup manager (app-data only: DB dumps + Docker-volume tars) ---
|
||||||
var backupMgr *backup.Manager
|
var backupMgr *backup.Manager
|
||||||
|
// R-550: the restore the last stop interrupted, found when the record loads; raised once the notifier exists.
|
||||||
|
var interruptedRestore *backup.RestoreOpResult
|
||||||
stackProv := &stackAdapter{
|
stackProv := &stackAdapter{
|
||||||
mgr: stackMgr,
|
mgr: stackMgr,
|
||||||
getStoragePaths: func() []settings.StoragePath { return sett.GetStoragePaths() },
|
getStoragePaths: func() []settings.StoragePath { return sett.GetStoragePaths() },
|
||||||
@@ -534,6 +536,8 @@ func main() {
|
|||||||
}
|
}
|
||||||
if cfg.Backup.Enabled {
|
if cfg.Backup.Enabled {
|
||||||
backupMgr = backup.NewManager(cfg, sett, logger)
|
backupMgr = backup.NewManager(cfg, sett, logger)
|
||||||
|
// R-550: load the persisted restore record BEFORE any page can ask for it.
|
||||||
|
interruptedRestore = loadRestoreRecordAtStartup(backupMgr, cfg.Paths.DataDir, logger)
|
||||||
// R-166: use the guard that already ran Recover at startup, not a second one over the same
|
// R-166: use the guard that already ran Recover at startup, not a second one over the same
|
||||||
// file (see SetAppStopGuard — one file, one owner).
|
// file (see SetAppStopGuard — one file, one owner).
|
||||||
backupMgr.SetAppStopGuard(appStopGuard)
|
backupMgr.SetAppStopGuard(appStopGuard)
|
||||||
@@ -580,6 +584,8 @@ func main() {
|
|||||||
|
|
||||||
// --- Initialize notifier ---
|
// --- Initialize notifier ---
|
||||||
notifier := notify.New(cfg.Hub.URL, cfg.Hub.APIKey, cfg.Customer.ID, sett, logger, cfg.Logging.Level == "debug")
|
notifier := notify.New(cfg.Hub.URL, cfg.Hub.APIKey, cfg.Customer.ID, sett, logger, cfg.Logging.Level == "debug")
|
||||||
|
// R-550: an interrupted restore is raised once, now that the notifier exists.
|
||||||
|
reportInterruptedRestore(notifier, interruptedRestore)
|
||||||
|
|
||||||
// R-97a: wire the quiesce loop's hub-event seam. It MUST happen here and not in
|
// R-97a: wire the quiesce loop's hub-event seam. It MUST happen here and not in
|
||||||
// startQuiesceLoop, because the notifier is constructed after it — and it must happen at all,
|
// startQuiesceLoop, because the notifier is constructed after it — and it must happen at all,
|
||||||
|
|||||||
@@ -0,0 +1,48 @@
|
|||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"fmt"
|
||||||
|
"log"
|
||||||
|
"path/filepath"
|
||||||
|
|
||||||
|
"gitea.dooplex.hu/admin/felhom-controller/internal/backup"
|
||||||
|
)
|
||||||
|
|
||||||
|
// R-550 — the restore record survives a restart, and an interruption is raised ONCE.
|
||||||
|
//
|
||||||
|
// Two functions so main() makes two plain calls a test can see (TestMainWiresRestoreRecord): the
|
||||||
|
// record must be loaded BEFORE the web server serves a restore page, and the event can only be pushed
|
||||||
|
// AFTER the notifier exists, which main() constructs later.
|
||||||
|
|
||||||
|
// restoreRecordFileName lives in DataDir, beside settings.json — the SSD state dir that survives a
|
||||||
|
// container recreation.
|
||||||
|
const restoreRecordFileName = "restore-status.json"
|
||||||
|
|
||||||
|
// restoreEventPusher is the notifier method main() hands in (an interface so a test can record it).
|
||||||
|
type restoreEventPusher interface {
|
||||||
|
PushEvent(eventType, severity, message string, details interface{})
|
||||||
|
}
|
||||||
|
|
||||||
|
// loadRestoreRecordAtStartup wires persistence and returns the restore the last stop interrupted.
|
||||||
|
func loadRestoreRecordAtStartup(mgr *backup.Manager, dataDir string, logger *log.Logger) *backup.RestoreOpResult {
|
||||||
|
if mgr == nil {
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
path := filepath.Join(dataDir, restoreRecordFileName)
|
||||||
|
mgr.SetRestoreRecordPath(path)
|
||||||
|
res := mgr.LoadRestoreRecord()
|
||||||
|
if logger != nil {
|
||||||
|
logger.Printf("[INFO] [backup] restore record wired: %s (interrupted at startup: %t)", path, res != nil)
|
||||||
|
}
|
||||||
|
return res
|
||||||
|
}
|
||||||
|
|
||||||
|
// reportInterruptedRestore raises restore_interrupted (warning; for the household — hub v0.117.0).
|
||||||
|
func reportInterruptedRestore(p restoreEventPusher, res *backup.RestoreOpResult) {
|
||||||
|
if p == nil || res == nil {
|
||||||
|
return
|
||||||
|
}
|
||||||
|
p.PushEvent("restore_interrupted", "warning",
|
||||||
|
fmt.Sprintf("A(z) %s visszaállítása megszakadt, mert a doboz újraindult — indítsd el újra.", res.Stack),
|
||||||
|
map[string]any{"stack": res.Stack, "op": res.Op})
|
||||||
|
}
|
||||||
@@ -0,0 +1,82 @@
|
|||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"go/ast"
|
||||||
|
"go/parser"
|
||||||
|
"go/token"
|
||||||
|
"io"
|
||||||
|
"log"
|
||||||
|
"testing"
|
||||||
|
|
||||||
|
"gitea.dooplex.hu/admin/felhom-controller/internal/backup"
|
||||||
|
)
|
||||||
|
|
||||||
|
type recordedPush struct{ eventType, severity, message string }
|
||||||
|
|
||||||
|
type pushRecorder struct{ got []recordedPush }
|
||||||
|
|
||||||
|
func (r *pushRecorder) PushEvent(et, sev, msg string, _ interface{}) {
|
||||||
|
r.got = append(r.got, recordedPush{et, sev, msg})
|
||||||
|
}
|
||||||
|
|
||||||
|
// R-550 — the functions main() calls: a restore left in flight is found at startup and raised as
|
||||||
|
// restore_interrupted (warning), exactly once.
|
||||||
|
//
|
||||||
|
// RED-PROOF: make loadRestoreRecordAtStartup skip LoadRestoreRecord → nothing is pushed →
|
||||||
|
// "an interrupted restore raised no restore_interrupted".
|
||||||
|
func TestRestoreRecordAtStartup_RaisesInterruptedOnce(t *testing.T) {
|
||||||
|
dir := t.TempDir()
|
||||||
|
lg := log.New(io.Discard, "", 0)
|
||||||
|
|
||||||
|
before := new(backup.Manager)
|
||||||
|
before.SetRestoreRecordPath(dir + "/" + restoreRecordFileName)
|
||||||
|
before.BeginRestoreOp("restore", "gokapi") // the box stops here
|
||||||
|
|
||||||
|
rec := &pushRecorder{}
|
||||||
|
reportInterruptedRestore(rec, loadRestoreRecordAtStartup(new(backup.Manager), dir, lg))
|
||||||
|
if len(rec.got) != 1 || rec.got[0].eventType != "restore_interrupted" || rec.got[0].severity != "warning" {
|
||||||
|
t.Fatalf("an interrupted restore raised no restore_interrupted (warning): %+v", rec.got)
|
||||||
|
}
|
||||||
|
|
||||||
|
again := &pushRecorder{}
|
||||||
|
reportInterruptedRestore(again, loadRestoreRecordAtStartup(new(backup.Manager), dir, lg))
|
||||||
|
if len(again.got) != 0 {
|
||||||
|
t.Fatalf("the same interruption was raised again on the next start: %+v", again.got)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The seam must be CALLED from main(): both functions, in that order.
|
||||||
|
func TestMainWiresRestoreRecord(t *testing.T) {
|
||||||
|
fset := token.NewFileSet()
|
||||||
|
f, err := parser.ParseFile(fset, "main.go", nil, 0)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("parse main.go: %v", err)
|
||||||
|
}
|
||||||
|
pos := map[string]token.Pos{}
|
||||||
|
for _, d := range f.Decls {
|
||||||
|
fn, ok := d.(*ast.FuncDecl)
|
||||||
|
if !ok || fn.Name.Name != "main" {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
ast.Inspect(fn.Body, func(n ast.Node) bool {
|
||||||
|
if c, ok := n.(*ast.CallExpr); ok {
|
||||||
|
if id, ok := c.Fun.(*ast.Ident); ok {
|
||||||
|
if id.Name == "loadRestoreRecordAtStartup" || id.Name == "reportInterruptedRestore" {
|
||||||
|
if _, seen := pos[id.Name]; !seen {
|
||||||
|
pos[id.Name] = c.Pos()
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return true
|
||||||
|
})
|
||||||
|
}
|
||||||
|
load, okL := pos["loadRestoreRecordAtStartup"]
|
||||||
|
report, okR := pos["reportInterruptedRestore"]
|
||||||
|
if !okL || !okR {
|
||||||
|
t.Fatalf("main() does not call both restore-record functions (load=%t report=%t) — the persistence is an inert seam", okL, okR)
|
||||||
|
}
|
||||||
|
if load > report {
|
||||||
|
t.Fatalf("main() reports before it loads — nothing would ever be reported")
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -264,6 +264,9 @@ type Manager struct {
|
|||||||
opStack string
|
opStack string
|
||||||
opStartedAt time.Time
|
opStartedAt time.Time
|
||||||
opLast *RestoreOpResult
|
opLast *RestoreOpResult
|
||||||
|
// R-550: persistence of the above (restore_record.go) and the per-app interrupted notices.
|
||||||
|
opRecordPath string
|
||||||
|
opInterrupted map[string]RestoreOpResult
|
||||||
|
|
||||||
// Cached status for page rendering (refreshed periodically)
|
// Cached status for page rendering (refreshed periodically)
|
||||||
cachedStatus *FullBackupStatus
|
cachedStatus *FullBackupStatus
|
||||||
|
|||||||
@@ -6,8 +6,8 @@ import "time"
|
|||||||
// backups page can show a progress banner (running → success/failure) instead of blocking the HTTP
|
// backups page can show a progress banner (running → success/failure) instead of blocking the HTTP
|
||||||
// request until the restore completes. It is display-only and mutex-guarded on the Manager's `mu`;
|
// request until the restore completes. It is display-only and mutex-guarded on the Manager's `mu`;
|
||||||
// it does NOT gate concurrency (that stays the restore functions' internal single-flight acquire).
|
// it does NOT gate concurrency (that stays the restore functions' internal single-flight acquire).
|
||||||
// In-memory only — lost on a controller restart (same precedent as notification cooldowns); a page
|
// PERSISTED since v0.246.0 (R-550, operator ruling 2026-09-17 — a reversal of the original in-memory
|
||||||
// load mid-op after a restart simply shows no banner.
|
// choice, for the restore record only; notification cooldowns stay in memory). See restore_record.go.
|
||||||
|
|
||||||
// RestoreOpResult is the terminal record of the most recent restore op.
|
// RestoreOpResult is the terminal record of the most recent restore op.
|
||||||
type RestoreOpResult struct {
|
type RestoreOpResult struct {
|
||||||
@@ -16,6 +16,9 @@ type RestoreOpResult struct {
|
|||||||
OK bool `json:"ok"`
|
OK bool `json:"ok"`
|
||||||
Message string `json:"message"`
|
Message string `json:"message"`
|
||||||
FinishedAt time.Time `json:"finished_at"`
|
FinishedAt time.Time `json:"finished_at"`
|
||||||
|
// Interrupted marks a restore that was still in flight when the controller stopped (R-550): found
|
||||||
|
// at the next start, recorded as a failure with RestoreInterruptedMessage.
|
||||||
|
Interrupted bool `json:"interrupted,omitempty"`
|
||||||
}
|
}
|
||||||
|
|
||||||
// RestoreResultWindow bounds how long a finished restore still counts as "what just happened".
|
// RestoreResultWindow bounds how long a finished restore still counts as "what just happened".
|
||||||
@@ -54,6 +57,9 @@ func (m *Manager) BeginRestoreOp(op, stack string) {
|
|||||||
m.opName = op
|
m.opName = op
|
||||||
m.opStack = stack
|
m.opStack = stack
|
||||||
m.opStartedAt = time.Now()
|
m.opStartedAt = time.Now()
|
||||||
|
// A new restore of this app supersedes its interrupted notice (R-550).
|
||||||
|
delete(m.opInterrupted, stack)
|
||||||
|
m.persistRestoreRecordLocked()
|
||||||
}
|
}
|
||||||
|
|
||||||
// EndRestoreOp records the terminal result (called from the goroutine on completion, success or
|
// EndRestoreOp records the terminal result (called from the goroutine on completion, success or
|
||||||
@@ -69,6 +75,7 @@ func (m *Manager) EndRestoreOp(ok bool, message string) {
|
|||||||
FinishedAt: time.Now(),
|
FinishedAt: time.Now(),
|
||||||
}
|
}
|
||||||
m.opRunning = false
|
m.opRunning = false
|
||||||
|
m.persistRestoreRecordLocked()
|
||||||
}
|
}
|
||||||
|
|
||||||
// RestoreStatus returns a deep copy of the current restore op-status for the page/API.
|
// RestoreStatus returns a deep copy of the current restore op-status for the page/API.
|
||||||
|
|||||||
@@ -0,0 +1,133 @@
|
|||||||
|
package backup
|
||||||
|
|
||||||
|
import (
|
||||||
|
"encoding/json"
|
||||||
|
"os"
|
||||||
|
"sort"
|
||||||
|
"time"
|
||||||
|
)
|
||||||
|
|
||||||
|
// R-550 (operator ruling "fix", 2026-09-17) — the restore record survives a restart.
|
||||||
|
//
|
||||||
|
// A DESIGN REVERSED, AND RECORDED AS SUCH. opstatus.go was deliberately in-memory ("same precedent as
|
||||||
|
// notification cooldowns"). Chaos night round 10 measured the cost: a restore accepted at 23:26:08Z,
|
||||||
|
// the box hard-reset four seconds later, and afterwards the status answered the Go zero value — the
|
||||||
|
// household had pressed a button, been told it started, and could never learn whether it finished.
|
||||||
|
// The operator reversed the choice for the RESTORE RECORD ONLY; notification cooldowns stay in memory.
|
||||||
|
//
|
||||||
|
// Shape: one JSON file in the controller's state directory (DataDir, beside settings.json), written
|
||||||
|
// atomically (atomicWrite: tmp + rename) at BOTH ends of an op. At startup a record still marked
|
||||||
|
// running is, by construction, a restore nothing is running any more: it becomes a terminal failure
|
||||||
|
// (Interrupted, RestoreInterruptedMessage) and a per-app notice that stays until that app's next
|
||||||
|
// restore. LoadRestoreRecord reports the conversion ONCE, so the caller raises restore_interrupted once.
|
||||||
|
//
|
||||||
|
// No path set (tests that build a bare Manager, a box with backup disabled) = the old in-memory
|
||||||
|
// behaviour, silently: persistence is a property of the wired controller, not of every Manager.
|
||||||
|
|
||||||
|
// RestoreInterruptedMessage is what the household reads when the box stopped mid-restore.
|
||||||
|
const RestoreInterruptedMessage = "A visszaállítás megszakadt (a doboz újraindult) — indítsd el újra."
|
||||||
|
|
||||||
|
// restoreRecordFile is the on-disk shape.
|
||||||
|
type restoreRecordFile struct {
|
||||||
|
Running bool `json:"running"`
|
||||||
|
Op string `json:"op,omitempty"`
|
||||||
|
Stack string `json:"stack,omitempty"`
|
||||||
|
StartedAt time.Time `json:"started_at,omitempty"`
|
||||||
|
Last *RestoreOpResult `json:"last,omitempty"`
|
||||||
|
Interrupted map[string]RestoreOpResult `json:"interrupted,omitempty"`
|
||||||
|
}
|
||||||
|
|
||||||
|
// SetRestoreRecordPath wires persistence. Call before LoadRestoreRecord and before serving requests.
|
||||||
|
func (m *Manager) SetRestoreRecordPath(path string) {
|
||||||
|
m.mu.Lock()
|
||||||
|
defer m.mu.Unlock()
|
||||||
|
m.opRecordPath = path
|
||||||
|
}
|
||||||
|
|
||||||
|
// LoadRestoreRecord restores the record at startup. It returns the restore that was interrupted by the
|
||||||
|
// stop — non-nil exactly once per interruption — so the caller can log it and raise restore_interrupted.
|
||||||
|
func (m *Manager) LoadRestoreRecord() *RestoreOpResult {
|
||||||
|
m.mu.Lock()
|
||||||
|
defer m.mu.Unlock()
|
||||||
|
if m.opRecordPath == "" {
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
b, err := os.ReadFile(m.opRecordPath)
|
||||||
|
if err != nil {
|
||||||
|
if !os.IsNotExist(err) && m.logger != nil {
|
||||||
|
m.logger.Printf("[WARN] [backup] restore record unreadable (%v) — starting with no restore history", err)
|
||||||
|
}
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
var rec restoreRecordFile
|
||||||
|
if err := json.Unmarshal(b, &rec); err != nil {
|
||||||
|
if m.logger != nil {
|
||||||
|
m.logger.Printf("[WARN] [backup] restore record corrupt (%v) — starting with no restore history", err)
|
||||||
|
}
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
m.opLast = rec.Last
|
||||||
|
m.opInterrupted = rec.Interrupted
|
||||||
|
if !rec.Running || rec.Stack == "" {
|
||||||
|
if m.logger != nil && len(m.opInterrupted) > 0 {
|
||||||
|
m.logger.Printf("[INFO] [backup] restore record loaded: %d interrupted restore notice(s) still shown", len(m.opInterrupted))
|
||||||
|
}
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
res := RestoreOpResult{
|
||||||
|
Op: rec.Op, Stack: rec.Stack, OK: false,
|
||||||
|
Message: RestoreInterruptedMessage, FinishedAt: time.Now(), Interrupted: true,
|
||||||
|
}
|
||||||
|
m.opLast = &res
|
||||||
|
if m.opInterrupted == nil {
|
||||||
|
m.opInterrupted = map[string]RestoreOpResult{}
|
||||||
|
}
|
||||||
|
m.opInterrupted[rec.Stack] = res
|
||||||
|
m.opRunning = false
|
||||||
|
if m.logger != nil {
|
||||||
|
m.logger.Printf("[WARN] [backup] restore of %s (%s, started %s) was INTERRUPTED by a controller stop — recorded as failed; the household is told to run it again",
|
||||||
|
rec.Stack, rec.Op, rec.StartedAt.Format(time.RFC3339))
|
||||||
|
}
|
||||||
|
m.persistRestoreRecordLocked()
|
||||||
|
out := res
|
||||||
|
return &out
|
||||||
|
}
|
||||||
|
|
||||||
|
// InterruptedRestore reports the standing interrupted-restore notice for one app, if any.
|
||||||
|
func (m *Manager) InterruptedRestore(stack string) (RestoreOpResult, bool) {
|
||||||
|
m.mu.Lock()
|
||||||
|
defer m.mu.Unlock()
|
||||||
|
r, ok := m.opInterrupted[stack]
|
||||||
|
return r, ok
|
||||||
|
}
|
||||||
|
|
||||||
|
// InterruptedRestores lists every standing notice, sorted by app, for the restore page.
|
||||||
|
func (m *Manager) InterruptedRestores() []RestoreOpResult {
|
||||||
|
m.mu.Lock()
|
||||||
|
defer m.mu.Unlock()
|
||||||
|
out := make([]RestoreOpResult, 0, len(m.opInterrupted))
|
||||||
|
for _, r := range m.opInterrupted {
|
||||||
|
out = append(out, r)
|
||||||
|
}
|
||||||
|
sort.Slice(out, func(i, j int) bool { return out[i].Stack < out[j].Stack })
|
||||||
|
return out
|
||||||
|
}
|
||||||
|
|
||||||
|
// persistRestoreRecordLocked writes the current op-status. Caller holds m.mu. A failed write is logged
|
||||||
|
// and the in-memory status carries on — the page still works for this process's lifetime.
|
||||||
|
func (m *Manager) persistRestoreRecordLocked() {
|
||||||
|
if m.opRecordPath == "" {
|
||||||
|
return
|
||||||
|
}
|
||||||
|
rec := restoreRecordFile{
|
||||||
|
Running: m.opRunning, Op: m.opName, Stack: m.opStack, StartedAt: m.opStartedAt,
|
||||||
|
Last: m.opLast, Interrupted: m.opInterrupted,
|
||||||
|
}
|
||||||
|
b, err := json.Marshal(rec)
|
||||||
|
if err == nil {
|
||||||
|
err = atomicWrite(m.opRecordPath, b, 0o600)
|
||||||
|
}
|
||||||
|
if err != nil && m.logger != nil {
|
||||||
|
m.logger.Printf("[WARN] [backup] could not persist the restore record to %s: %v (kept in memory)", m.opRecordPath, err)
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,96 @@
|
|||||||
|
package backup
|
||||||
|
|
||||||
|
import (
|
||||||
|
"io"
|
||||||
|
"log"
|
||||||
|
"path/filepath"
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
)
|
||||||
|
|
||||||
|
func recordManager(path string) *Manager {
|
||||||
|
m := &Manager{logger: log.New(io.Discard, "", 0)}
|
||||||
|
m.SetRestoreRecordPath(path)
|
||||||
|
return m
|
||||||
|
}
|
||||||
|
|
||||||
|
// R-550 (operator ruling "fix", 2026-09-17). Chaos night round 10: a restore was accepted and the box
|
||||||
|
// was hard-reset four seconds later. Afterwards /api/backup/restore-status answered the Go zero value
|
||||||
|
// and nothing told the household whether the restore had finished — the op-status was in-memory only.
|
||||||
|
//
|
||||||
|
// The consequence asserted: a restore that was IN FLIGHT when the controller stopped is, after the
|
||||||
|
// next start, a FAILED result for that app saying it was interrupted — and loading it reports it once,
|
||||||
|
// so the caller can raise restore_interrupted.
|
||||||
|
//
|
||||||
|
// RED-PROOF: keep the op-status in memory (no record written / read) → the second Manager's status is
|
||||||
|
// blank → "after a restart the interrupted restore left no trace".
|
||||||
|
func TestRestoreRecord_InterruptedRestoreSurvivesRestart(t *testing.T) {
|
||||||
|
path := filepath.Join(t.TempDir(), "restore-status.json")
|
||||||
|
m1 := recordManager(path)
|
||||||
|
m1.BeginRestoreOp("restore", "gokapi")
|
||||||
|
// the box stops here — no EndRestoreOp
|
||||||
|
|
||||||
|
m2 := recordManager(path)
|
||||||
|
got := m2.LoadRestoreRecord()
|
||||||
|
st := m2.RestoreStatus()
|
||||||
|
if got == nil || st.Last == nil {
|
||||||
|
t.Fatalf("after a restart the interrupted restore left no trace: loaded=%v status=%+v", got, st)
|
||||||
|
}
|
||||||
|
if st.Running {
|
||||||
|
t.Fatalf("a restore from before the restart is still shown as RUNNING — nothing is running it: %+v", st)
|
||||||
|
}
|
||||||
|
if st.Last.OK || !st.Last.Interrupted || st.Last.Stack != "gokapi" || st.Last.Op != "restore" {
|
||||||
|
t.Fatalf("interrupted record = %+v, want a failed, interrupted restore of gokapi", st.Last)
|
||||||
|
}
|
||||||
|
// ASCII fragment of the Hungarian sentence, with a negative control.
|
||||||
|
if !strings.Contains(st.Last.Message, "megszakadt") || strings.Contains(st.Last.Message, "zzzz-not-present") {
|
||||||
|
t.Fatalf("message %q does not say the restore was interrupted", st.Last.Message)
|
||||||
|
}
|
||||||
|
if rec, ok := m2.InterruptedRestore("gokapi"); !ok || !rec.Interrupted {
|
||||||
|
t.Fatalf("InterruptedRestore(gokapi) = %+v, %v — the restore page would not show it", rec, ok)
|
||||||
|
}
|
||||||
|
// Reported ONCE: a third start must not raise the event again for the same interruption.
|
||||||
|
m3 := recordManager(path)
|
||||||
|
if again := m3.LoadRestoreRecord(); again != nil {
|
||||||
|
t.Fatalf("the same interruption was reported again on the next start: %+v", again)
|
||||||
|
}
|
||||||
|
if _, ok := m3.InterruptedRestore("gokapi"); !ok {
|
||||||
|
t.Fatalf("the interrupted notice vanished on the next start — it must stay until gokapi is restored again")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Control: a restore that FINISHED is still that result after a restart — never re-labelled interrupted.
|
||||||
|
func TestRestoreRecord_FinishedRestoreStaysFinished(t *testing.T) {
|
||||||
|
path := filepath.Join(t.TempDir(), "restore-status.json")
|
||||||
|
m1 := recordManager(path)
|
||||||
|
m1.BeginRestoreOp("restore", "mealie")
|
||||||
|
m1.EndRestoreOp(true, "kész")
|
||||||
|
|
||||||
|
m2 := recordManager(path)
|
||||||
|
if got := m2.LoadRestoreRecord(); got != nil {
|
||||||
|
t.Fatalf("a finished restore was reported as interrupted: %+v", got)
|
||||||
|
}
|
||||||
|
st := m2.RestoreStatus()
|
||||||
|
if st.Last == nil || !st.Last.OK || st.Last.Interrupted || st.Last.Message != "kész" {
|
||||||
|
t.Fatalf("finished record after restart = %+v, want ok 'kész'", st.Last)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The notice lasts until that app's NEXT restore — and only that app's.
|
||||||
|
func TestRestoreRecord_NextRestoreOfThatAppClearsTheNotice(t *testing.T) {
|
||||||
|
path := filepath.Join(t.TempDir(), "restore-status.json")
|
||||||
|
m1 := recordManager(path)
|
||||||
|
m1.BeginRestoreOp("restore", "gokapi")
|
||||||
|
m2 := recordManager(path)
|
||||||
|
m2.LoadRestoreRecord()
|
||||||
|
|
||||||
|
m2.BeginRestoreOp("restore", "mealie") // a different app
|
||||||
|
m2.EndRestoreOp(true, "kész")
|
||||||
|
if _, ok := m2.InterruptedRestore("gokapi"); !ok {
|
||||||
|
t.Fatalf("restoring ANOTHER app cleared gokapi's interrupted notice")
|
||||||
|
}
|
||||||
|
m2.BeginRestoreOp("restore", "gokapi") // the same app, again
|
||||||
|
if _, ok := m2.InterruptedRestore("gokapi"); ok {
|
||||||
|
t.Fatalf("starting gokapi's restore again did not clear its interrupted notice")
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -53,6 +53,11 @@ func (s *Server) escrowBannerVisible(r *http.Request) bool {
|
|||||||
if r == nil || !s.hasAdminSession(r) || !s.escrowPaused() {
|
if r == nil || !s.hasAdminSession(r) || !s.escrowPaused() {
|
||||||
return false
|
return false
|
||||||
}
|
}
|
||||||
|
// R-546: while the agent says the ceremony cannot run yet, do not urge the household into it.
|
||||||
|
// Unknown readiness keeps the bar (fail loud).
|
||||||
|
if ready, known := s.escrowReadiness(r.Context(), false); known && !ready {
|
||||||
|
return false
|
||||||
|
}
|
||||||
if c, err := r.Cookie(escrowBannerCookie); err == nil && c.Value == "1" {
|
if c, err := r.Cookie(escrowBannerCookie); err == nil && c.Value == "1" {
|
||||||
return false
|
return false
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -127,6 +127,15 @@ func (s *Server) escrowWizardPageHandler(w http.ResponseWriter, r *http.Request)
|
|||||||
"OffboxConfigured": s.backupMgr != nil && s.backupMgr.OffboxConfigured(),
|
"OffboxConfigured": s.backupMgr != nil && s.backupMgr.OffboxConfigured(),
|
||||||
"AgentSupported": escrowAgentSupported(agentVer),
|
"AgentSupported": escrowAgentSupported(agentVer),
|
||||||
}
|
}
|
||||||
|
// R-546: a paused box whose agent says the ceremony cannot run yet gets the waiting card, not a red
|
||||||
|
// checklist and a start form that would refuse. Only for the FIRST ceremony — an escrowed box's
|
||||||
|
// re-ceremony keeps the checklist, where a red row is a real fault to show.
|
||||||
|
if !escrowed {
|
||||||
|
if ready, known := s.escrowReadiness(r.Context(), true); known && !ready {
|
||||||
|
data["EscrowNotReady"] = true
|
||||||
|
data["EscrowNotReadyMessage"] = escrowNotReadyMessage
|
||||||
|
}
|
||||||
|
}
|
||||||
s.executeTemplate(w, r, "backups_escrow", data)
|
s.executeTemplate(w, r, "backups_escrow", data)
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -188,6 +197,16 @@ func (s *Server) escrowStartAPIHandler(w http.ResponseWriter, r *http.Request) {
|
|||||||
}
|
}
|
||||||
s.clearAuthFailures(ip)
|
s.clearAuthFailures(ip)
|
||||||
|
|
||||||
|
// (1b) R-546: readiness. The page hides the start form while the agent's preflight is red, but a
|
||||||
|
// direct POST (the path chaos night used) ran the ceremony and returned the agent's raw stderr.
|
||||||
|
// Refuse here, BEFORE anything is staged or started. Unknown readiness does not refuse: the
|
||||||
|
// agent-conn and version gates below answer that case, as before.
|
||||||
|
if ready, known := s.escrowReadiness(r.Context(), true); known && !ready {
|
||||||
|
s.logger.Printf("[INFO] [web] escrow start refused: the box is not ready yet (agent preflight not ok)")
|
||||||
|
escrowJSON(w, http.StatusConflict, nil, escrowNotReadyMessage)
|
||||||
|
return
|
||||||
|
}
|
||||||
|
|
||||||
// (2) Re-stage-first (only when offsite is configured — Scenario C boxes skip it).
|
// (2) Re-stage-first (only when offsite is configured — Scenario C boxes skip it).
|
||||||
if s.backupMgr != nil && s.backupMgr.OffboxConfigured() {
|
if s.backupMgr != nil && s.backupMgr.OffboxConfigured() {
|
||||||
if err := s.escrowStage(r.Context()); err != nil {
|
if err := s.escrowStage(r.Context()); err != nil {
|
||||||
|
|||||||
@@ -0,0 +1,104 @@
|
|||||||
|
package web
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"sort"
|
||||||
|
"strings"
|
||||||
|
"sync"
|
||||||
|
"time"
|
||||||
|
)
|
||||||
|
|
||||||
|
// R-546 — is the box READY to run the escrow ceremony?
|
||||||
|
//
|
||||||
|
// Measured 2026-09-16/17 (chaos night Phase 0): after a fresh bind the box has no PBS storage for ~17
|
||||||
|
// minutes. The agent's preflight is red, the ceremony refuses, and the R-543 reminder bar (v0.245.0)
|
||||||
|
// was on every page urging the household into it. The escrow page itself already hid its start form
|
||||||
|
// behind a red checklist — but it said nothing about WAITING, and a direct POST /api/escrow/start
|
||||||
|
// (the path chaos night used) ran the ceremony and returned the agent's raw "-storage" stderr.
|
||||||
|
//
|
||||||
|
// THE DEFINITION IS THE AGENT'S OWN `ok`: every blocking preflight item (pbs_storage_id, dr_tier,
|
||||||
|
// age_binary, hub_upload, sudo_grant; staged_secret is informational). Not a controller-side copy of
|
||||||
|
// one of those items — the escrow.pbs_storage_id reading of R-546 was the SYMPTOM that box showed, and
|
||||||
|
// a copy of one item is a second definition that drifts when the agent grows a sixth.
|
||||||
|
//
|
||||||
|
// Three answers, and UNKNOWN keeps today's behaviour (the bar shows): an unreachable agent must not
|
||||||
|
// silence a reminder that R-543 made loud on purpose.
|
||||||
|
//
|
||||||
|
// Cost: the bar hangs off every page render, so the answer is cached for escrowReadyTTL, and a probe
|
||||||
|
// is only ever made while the box is paused (escrowPaused gates it) — a finished box never asks.
|
||||||
|
|
||||||
|
const (
|
||||||
|
escrowReadyTTL = 60 * time.Second
|
||||||
|
escrowReadyTimeout = 3 * time.Second
|
||||||
|
)
|
||||||
|
|
||||||
|
type escrowReadinessCache struct {
|
||||||
|
mu sync.Mutex
|
||||||
|
ready bool
|
||||||
|
known bool
|
||||||
|
checkedAt time.Time
|
||||||
|
lastState string // for transition-only INFO logging
|
||||||
|
}
|
||||||
|
|
||||||
|
// escrowReadiness returns (ready, known). fresh=true skips the cache (the escrow page and the start
|
||||||
|
// API, where one request justifies one probe); fresh=false is the bar's cached read.
|
||||||
|
func (s *Server) escrowReadiness(ctx context.Context, fresh bool) (ready, known bool) {
|
||||||
|
c := &s.escrowReady
|
||||||
|
c.mu.Lock()
|
||||||
|
defer c.mu.Unlock()
|
||||||
|
if !fresh && !c.checkedAt.IsZero() && time.Since(c.checkedAt) < escrowReadyTTL {
|
||||||
|
return c.ready, c.known
|
||||||
|
}
|
||||||
|
pctx, cancel := context.WithTimeout(ctx, escrowReadyTimeout)
|
||||||
|
defer cancel()
|
||||||
|
ready, known, why := s.probeEscrowReadiness(pctx)
|
||||||
|
c.ready, c.known, c.checkedAt = ready, known, time.Now()
|
||||||
|
|
||||||
|
state := "unknown"
|
||||||
|
if known && ready {
|
||||||
|
state = "ready"
|
||||||
|
} else if known {
|
||||||
|
state = "not-ready"
|
||||||
|
}
|
||||||
|
if state != c.lastState {
|
||||||
|
switch state {
|
||||||
|
case "ready":
|
||||||
|
s.logger.Printf("[INFO] [web] escrow readiness: READY (agent preflight ok) — the recovery-code reminder is shown")
|
||||||
|
case "not-ready":
|
||||||
|
s.logger.Printf("[INFO] [web] escrow readiness: NOT READY (%s) — reminder held back; the escrow page says the box is still preparing", why)
|
||||||
|
default:
|
||||||
|
s.logger.Printf("[INFO] [web] escrow readiness: UNKNOWN (%s) — reminder shown (fail loud)", why)
|
||||||
|
}
|
||||||
|
c.lastState = state
|
||||||
|
} else if s.isDebug() {
|
||||||
|
s.logger.Printf("[DEBUG] [web] escrow readiness: %s (%s)", state, why)
|
||||||
|
}
|
||||||
|
return ready, known
|
||||||
|
}
|
||||||
|
|
||||||
|
// probeEscrowReadiness asks the agent. Never an error to the caller: failure is "unknown".
|
||||||
|
func (s *Server) probeEscrowReadiness(ctx context.Context) (ready, known bool, why string) {
|
||||||
|
agent, err := s.escrowAgentConn()
|
||||||
|
if err != nil {
|
||||||
|
return false, false, "agent unavailable: " + err.Error()
|
||||||
|
}
|
||||||
|
pf, err := agent.EscrowPreflight(ctx)
|
||||||
|
if err != nil {
|
||||||
|
return false, false, "preflight failed: " + err.Error()
|
||||||
|
}
|
||||||
|
if pf.OK {
|
||||||
|
return true, true, "preflight ok"
|
||||||
|
}
|
||||||
|
var failing []string
|
||||||
|
for _, it := range pf.Items {
|
||||||
|
if !it.OK && it.ID != "staged_secret" {
|
||||||
|
failing = append(failing, it.ID)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
sort.Strings(failing)
|
||||||
|
return false, true, "failing: " + strings.Join(failing, ", ")
|
||||||
|
}
|
||||||
|
|
||||||
|
// escrowNotReadyMessage is the one Hungarian sentence for "wait" — the page card and the API refusal
|
||||||
|
// say the same thing.
|
||||||
|
const escrowNotReadyMessage = "A doboz még készül — a távoli mentés kulcsát pár perc múlva tudod létrehozni. Ez az oldal magától frissül."
|
||||||
@@ -87,6 +87,9 @@ func newEscrowWizardHarness(t *testing.T) *escrowWizardHarness {
|
|||||||
version: "0.88.0",
|
version: "0.88.0",
|
||||||
startResp: agentapi.EscrowCeremonyStartResponse{JobID: "escrow-1", Phase: "running"},
|
startResp: agentapi.EscrowCeremonyStartResponse{JobID: "escrow-1", Phase: "running"},
|
||||||
startStatus: http.StatusAccepted,
|
startStatus: http.StatusAccepted,
|
||||||
|
// R-546 (v0.246.0): the start API now asks the agent's preflight first; this harness is a
|
||||||
|
// READY box. A not-ready box is r546_escrow_readiness_test.go.
|
||||||
|
pf: agentapi.EscrowPreflightResponse{OK: true},
|
||||||
}
|
}
|
||||||
h.s = &Server{cfg: cfg, backupMgr: h.m, settings: sett, logger: lg}
|
h.s = &Server{cfg: cfg, backupMgr: h.m, settings: sett, logger: lg}
|
||||||
h.s.escrowAgentFn = func() (escrowAgent, error) { return h.agent, nil }
|
h.s.escrowAgentFn = func() (escrowAgent, error) { return h.agent, nil }
|
||||||
@@ -131,8 +134,8 @@ func TestEscrowStart_StagesBeforeTrigger(t *testing.T) {
|
|||||||
if w.Code != http.StatusOK {
|
if w.Code != http.StatusOK {
|
||||||
t.Fatalf("start: got %d (%s)", w.Code, w.Body.String())
|
t.Fatalf("start: got %d (%s)", w.Code, w.Body.String())
|
||||||
}
|
}
|
||||||
if got := strings.Join(h.order, ","); got != "stage,start" {
|
if got := strings.Join(h.order, ","); got != "preflight,stage,start" { // R-546: readiness first
|
||||||
t.Fatalf("call order = %q, want stage BEFORE start (a ceremony without the staged secret mints a hash-less blob)", got)
|
t.Fatalf("call order = %q, want preflight, then stage BEFORE start (a ceremony without the staged secret mints a hash-less blob)", got)
|
||||||
}
|
}
|
||||||
if !strings.Contains(w.Body.String(), "escrow-1") {
|
if !strings.Contains(w.Body.String(), "escrow-1") {
|
||||||
t.Fatalf("response lacks the job id: %s", w.Body.String())
|
t.Fatalf("response lacks the job id: %s", w.Body.String())
|
||||||
@@ -147,8 +150,8 @@ func TestEscrowStart_ReceremonyStagesToo(t *testing.T) {
|
|||||||
if w := postStart(t, h.s, wizardPassword); w.Code != http.StatusOK {
|
if w := postStart(t, h.s, wizardPassword); w.Code != http.StatusOK {
|
||||||
t.Fatalf("re-ceremony start: got %d", w.Code)
|
t.Fatalf("re-ceremony start: got %d", w.Code)
|
||||||
}
|
}
|
||||||
if got := strings.Join(h.order, ","); got != "stage,start" {
|
if got := strings.Join(h.order, ","); got != "preflight,stage,start" { // R-546: readiness first
|
||||||
t.Fatalf("re-ceremony call order = %q, want stage,start", got)
|
t.Fatalf("re-ceremony call order = %q, want preflight,stage,start", got)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -159,8 +162,8 @@ func TestEscrowStart_NoOffboxSkipsStaging(t *testing.T) {
|
|||||||
if w := postStart(t, h.s, wizardPassword); w.Code != http.StatusOK {
|
if w := postStart(t, h.s, wizardPassword); w.Code != http.StatusOK {
|
||||||
t.Fatalf("start: got %d", w.Code)
|
t.Fatalf("start: got %d", w.Code)
|
||||||
}
|
}
|
||||||
if got := strings.Join(h.order, ","); got != "start" {
|
if got := strings.Join(h.order, ","); got != "preflight,start" { // R-546: readiness first
|
||||||
t.Fatalf("call order = %q, want start only (no staging without offbox)", got)
|
t.Fatalf("call order = %q, want preflight then start only (no staging without offbox)", got)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -1067,6 +1067,10 @@ func (s *Server) backupsAppsHandler(w http.ResponseWriter, r *http.Request) {
|
|||||||
// restore-to-verify list and the .fab export/import loop.
|
// restore-to-verify list and the .fab export/import loop.
|
||||||
func (s *Server) backupsRestoreHandler(w http.ResponseWriter, r *http.Request) {
|
func (s *Server) backupsRestoreHandler(w http.ResponseWriter, r *http.Request) {
|
||||||
data := s.backupsCommonData("backups-restore", "Biztonsági mentés — Visszaállítás", r)
|
data := s.backupsCommonData("backups-restore", "Biztonsági mentés — Visszaállítás", r)
|
||||||
|
// R-550: restores the box was running when it stopped — shown until that app is restored again.
|
||||||
|
if s.backupMgr != nil {
|
||||||
|
data["InterruptedRestores"] = s.backupMgr.InterruptedRestores()
|
||||||
|
}
|
||||||
s.backupsOffboxData(data) // restore-to-verify lists the offbox-toggled apps
|
s.backupsOffboxData(data) // restore-to-verify lists the offbox-toggled apps
|
||||||
// Full-restore two-step reveal (§7.2): after the size+headroom prepare step, offboxRestoreHandler
|
// Full-restore two-step reveal (§7.2): after the size+headroom prepare step, offboxRestoreHandler
|
||||||
// redirects here with the app + human size so the confirm section can show the size BEFORE starting.
|
// redirects here with the app + human size so the confirm section can show the size BEFORE starting.
|
||||||
|
|||||||
@@ -0,0 +1,113 @@
|
|||||||
|
package web
|
||||||
|
|
||||||
|
import (
|
||||||
|
"errors"
|
||||||
|
"net/http"
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
|
||||||
|
"gitea.dooplex.hu/admin/felhom-controller/internal/agentapi"
|
||||||
|
)
|
||||||
|
|
||||||
|
// ── R-546 — do not send the household to a ceremony the box cannot run yet ────────────────────────
|
||||||
|
//
|
||||||
|
// Measured 2026-09-16/17 (chaos night Phase 0): for ~17 minutes after the bind the box has no PBS
|
||||||
|
// storage, the agent's preflight is red, and the ceremony refuses. Meanwhile the R-543 bar is on every
|
||||||
|
// page urging the household there. Readiness is the AGENT'S OWN `ok` — every blocking preflight item
|
||||||
|
// (pbs_storage_id, dr_tier, age_binary, hub_upload, sudo_grant), never a copy of one of them.
|
||||||
|
|
||||||
|
const escrowNotReadyMarker = `id="escrow-not-ready"`
|
||||||
|
|
||||||
|
func readinessAgent(ok bool) *fakeEscrowAgent {
|
||||||
|
order := []string{}
|
||||||
|
pf := agentapi.EscrowPreflightResponse{OK: ok}
|
||||||
|
if !ok {
|
||||||
|
pf.Items = []agentapi.EscrowPreflightItem{{ID: "pbs_storage_id", OK: false, Detail: "escrow.pbs_storage_id not configured"}}
|
||||||
|
}
|
||||||
|
return &fakeEscrowAgent{order: &order, version: "0.132.0", pf: pf}
|
||||||
|
}
|
||||||
|
|
||||||
|
// RED-PROOF: remove the readiness check from escrowBannerVisible → the bar renders on a box that
|
||||||
|
// cannot run the ceremony → "the bar urges a ceremony the box cannot run yet".
|
||||||
|
func TestR546_NotReadyBoxHoldsTheBar(t *testing.T) {
|
||||||
|
s := escrowServer(t, "pending")
|
||||||
|
a := readinessAgent(false)
|
||||||
|
s.escrowAgentFn = func() (escrowAgent, error) { return a, nil }
|
||||||
|
|
||||||
|
body := getPage(t, s, "/dashboard").Body.String()
|
||||||
|
if strings.Contains(body, escrowBarSentence) {
|
||||||
|
t.Fatal("R-546: the bar urges a ceremony the box cannot run yet (agent preflight is red)")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Control: once the agent says ready, the R-543 bar is back.
|
||||||
|
func TestR546_ReadyBoxShowsTheBar(t *testing.T) {
|
||||||
|
s := escrowServer(t, "pending")
|
||||||
|
a := readinessAgent(true)
|
||||||
|
s.escrowAgentFn = func() (escrowAgent, error) { return a, nil }
|
||||||
|
if body := getPage(t, s, "/dashboard").Body.String(); !strings.Contains(body, escrowBarSentence) {
|
||||||
|
t.Fatal("a READY, paused box no longer shows the R-543 bar")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Fail loud: readiness UNKNOWN (agent unreachable) keeps the reminder — silence is not a safe default.
|
||||||
|
func TestR546_UnknownReadinessStillAsks(t *testing.T) {
|
||||||
|
s := escrowServer(t, "pending")
|
||||||
|
s.escrowAgentFn = func() (escrowAgent, error) { return nil, errors.New("agent unreachable") }
|
||||||
|
if body := getPage(t, s, "/dashboard").Body.String(); !strings.Contains(body, escrowBarSentence) {
|
||||||
|
t.Fatal("an unknown readiness hid the reminder — an unreachable agent must not silence R-543")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The page itself: a calm waiting card, no start form, and none of the agent's raw diagnostics.
|
||||||
|
// RED-PROOF: never set EscrowNotReady → the checklist + start form render → "start form offered".
|
||||||
|
func TestR546_EscrowPageWaitsWhenNotReady(t *testing.T) {
|
||||||
|
s := escrowServer(t, "pending")
|
||||||
|
a := readinessAgent(false)
|
||||||
|
s.escrowAgentFn = func() (escrowAgent, error) { return a, nil }
|
||||||
|
|
||||||
|
rec := getPage(t, s, "/backup/escrow")
|
||||||
|
if rec.Code != 200 {
|
||||||
|
t.Fatalf("GET /backup/escrow = %d: %s", rec.Code, rec.Body.String())
|
||||||
|
}
|
||||||
|
body := rec.Body.String()
|
||||||
|
if !strings.Contains(body, escrowNotReadyMarker) || !strings.Contains(body, "A doboz m") {
|
||||||
|
t.Fatal("R-546: the escrow page does not tell the household the box is still preparing")
|
||||||
|
}
|
||||||
|
if strings.Contains(body, `id="start-form"`) {
|
||||||
|
t.Fatal("R-546: the start form is offered on a box whose ceremony would refuse")
|
||||||
|
}
|
||||||
|
// The exact texts chaos night met: the preflight detail, and the ceremony stderr ("…escrow-create
|
||||||
|
// requires -storage …"). NOT a bare "-storage": the menu carries id="nav-group-storage".
|
||||||
|
for _, raw := range []string{"pbs_storage_id not configured", "requires -storage"} {
|
||||||
|
if strings.Contains(body, raw) {
|
||||||
|
t.Fatalf("R-546: the agent's raw diagnostic %q reached the household's page", raw)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if strings.Contains(body, "zzzz-not-present") {
|
||||||
|
t.Fatal("negative control matched")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The direct path — the one that actually produced the stderr in chaos night: POST /api/escrow/start
|
||||||
|
// on a not-ready box is refused BEFORE anything is staged or started, in Hungarian.
|
||||||
|
// RED-PROOF: remove the readiness gate from escrowStartAPIHandler → 200 and "stage,start" run.
|
||||||
|
func TestR546_StartRefusedWhenNotReady(t *testing.T) {
|
||||||
|
h := newEscrowWizardHarness(t)
|
||||||
|
h.configureOffbox(t, "pending")
|
||||||
|
h.agent.pf = agentapi.EscrowPreflightResponse{OK: false,
|
||||||
|
Items: []agentapi.EscrowPreflightItem{{ID: "pbs_storage_id", OK: false}}}
|
||||||
|
|
||||||
|
w := postStart(t, h.s, wizardPassword)
|
||||||
|
if w.Code != http.StatusConflict {
|
||||||
|
t.Fatalf("start on a not-ready box = %d (%s), want 409", w.Code, w.Body.String())
|
||||||
|
}
|
||||||
|
for _, op := range h.order {
|
||||||
|
if op == "stage" || op == "start" {
|
||||||
|
t.Fatalf("a not-ready box still ran %q (order %v) — the refusal must come first", op, h.order)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if !strings.Contains(w.Body.String(), "A doboz m") {
|
||||||
|
t.Fatalf("the refusal is not the Hungarian waiting sentence: %s", w.Body.String())
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,57 @@
|
|||||||
|
package web
|
||||||
|
|
||||||
|
import (
|
||||||
|
"path/filepath"
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
)
|
||||||
|
|
||||||
|
// ── R-550 — an interrupted restore is SHOWN on the restore page ─────────────────────────────────
|
||||||
|
//
|
||||||
|
// Chaos night round 10: a restore interrupted by a hard reset left the page blank. The record now
|
||||||
|
// survives (internal/backup/restore_record.go); these drive the REAL page through ServeHTTP, so they
|
||||||
|
// bite on the wiring — a record that exists but that no page renders is the inert-seam shape.
|
||||||
|
|
||||||
|
const interruptedCardFragment = "Megszakadt vissza" // ASCII fragment of the card heading
|
||||||
|
|
||||||
|
// RED-PROOF: drop `data["InterruptedRestores"]` from backupsRestoreHandler → the card never renders →
|
||||||
|
// "the restore page does not show the interrupted restore".
|
||||||
|
func TestR550_RestorePageShowsInterruptedRestore(t *testing.T) {
|
||||||
|
s := newDashboardServer(t, time.Time{})
|
||||||
|
if s.backupMgr == nil {
|
||||||
|
t.Fatal("fixture invalid: no backup manager, so nothing here measures the restore page")
|
||||||
|
}
|
||||||
|
s.backupMgr.SetRestoreRecordPath(filepath.Join(t.TempDir(), "restore-status.json"))
|
||||||
|
s.backupMgr.BeginRestoreOp("restore", "gokapi") // the box stops here …
|
||||||
|
if s.backupMgr.LoadRestoreRecord() == nil { // … and this is the next start
|
||||||
|
t.Fatal("fixture invalid: the record did not convert to an interrupted restore")
|
||||||
|
}
|
||||||
|
|
||||||
|
rec := getPage(t, s, "/backups/restore")
|
||||||
|
if rec.Code != 200 {
|
||||||
|
t.Fatalf("GET /backups/restore = %d: %s", rec.Code, rec.Body.String())
|
||||||
|
}
|
||||||
|
body := rec.Body.String()
|
||||||
|
if !strings.Contains(body, interruptedCardFragment) || !strings.Contains(body, "megszakadt") {
|
||||||
|
t.Fatalf("the restore page does not show the interrupted restore — the household is left not knowing, which is R-550")
|
||||||
|
}
|
||||||
|
if !strings.Contains(body, "gokapi") {
|
||||||
|
t.Fatalf("the interrupted card does not name the app")
|
||||||
|
}
|
||||||
|
if strings.Contains(body, "zzzz-not-present") {
|
||||||
|
t.Fatal("negative control matched — the fragment search is broken")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The other branch: a box with no interrupted restore shows no such card.
|
||||||
|
func TestR550_NoInterruptedRestoreNoCard(t *testing.T) {
|
||||||
|
s := newDashboardServer(t, time.Time{})
|
||||||
|
rec := getPage(t, s, "/backups/restore")
|
||||||
|
if rec.Code != 200 {
|
||||||
|
t.Fatalf("GET /backups/restore = %d", rec.Code)
|
||||||
|
}
|
||||||
|
if strings.Contains(rec.Body.String(), interruptedCardFragment) {
|
||||||
|
t.Fatalf("a box with no interrupted restore shows the interrupted card")
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -289,6 +289,10 @@ func (s *Server) backupsRestoreWizardHandler(w http.ResponseWriter, r *http.Requ
|
|||||||
// survives a reload, which the flash does not.
|
// survives a reload, which the flash does not.
|
||||||
if in.HasRecentResult {
|
if in.HasRecentResult {
|
||||||
data["LastResult"] = st.Last
|
data["LastResult"] = st.Last
|
||||||
|
} else if ir, ok := s.backupMgr.InterruptedRestore(app); ok {
|
||||||
|
// R-550: a restore of THIS app the box was running when it stopped — shown past the recency
|
||||||
|
// window, until the app is restored again.
|
||||||
|
data["LastResult"] = &ir
|
||||||
}
|
}
|
||||||
|
|
||||||
s.executeTemplate(w, r, "backups_restore_wizard", data)
|
s.executeTemplate(w, r, "backups_restore_wizard", data)
|
||||||
|
|||||||
@@ -101,6 +101,8 @@ type Server struct {
|
|||||||
escrowAgentFn func() (escrowAgent, error)
|
escrowAgentFn func() (escrowAgent, error)
|
||||||
escrowStageFn func(ctx context.Context) error
|
escrowStageFn func(ctx context.Context) error
|
||||||
escrowStaleFn func() bool
|
escrowStaleFn func() bool
|
||||||
|
// R-546: the cached agent-preflight readiness behind the reminder bar (escrow_readiness.go).
|
||||||
|
escrowReady escrowReadinessCache
|
||||||
|
|
||||||
// escrowSealedAtFn (v0.200.0, R-193) reports WHEN the hub's sealed recovery package was created —
|
// escrowSealedAtFn (v0.200.0, R-193) reports WHEN the hub's sealed recovery package was created —
|
||||||
// the one non-secret fact the recovery screen may state before a code is entered. Wired via
|
// the one non-secret fact the recovery screen may state before a code is entered. Wired via
|
||||||
|
|||||||
@@ -13,7 +13,27 @@
|
|||||||
<span class="domain-badge">{{.Domain}}</span>
|
<span class="domain-badge">{{.Domain}}</span>
|
||||||
</div>
|
</div>
|
||||||
|
|
||||||
{{if not .AgentSupported}}
|
{{if .EscrowNotReady}}
|
||||||
|
<!-- R-546: the box cannot run the ceremony yet (the agent's own preflight is not ok — typically the
|
||||||
|
first minutes after a bind, before the PBS storage is applied). No checklist of red crosses, no
|
||||||
|
start form that would refuse, none of the agent's raw diagnostics: one sentence, and the page
|
||||||
|
polls and reloads itself when the agent says ready. -->
|
||||||
|
<div class="settings-card" id="escrow-not-ready">
|
||||||
|
<div class="alert alert-info">{{.EscrowNotReadyMessage}}</div>
|
||||||
|
<p class="form-hint">Ehhez nem kell semmit tenned. Ha fél óra múlva is ezt látod, szólj az üzemeltetőnek.</p>
|
||||||
|
</div>
|
||||||
|
<script>
|
||||||
|
(function(){
|
||||||
|
function check(){
|
||||||
|
fetch('/api/escrow/preflight',{headers:{'Accept':'application/json'}})
|
||||||
|
.then(function(r){ return r.json(); })
|
||||||
|
.then(function(j){ if(j && j.data && j.data.ok){ window.location.reload(); } })
|
||||||
|
.catch(function(){});
|
||||||
|
}
|
||||||
|
setInterval(check, 20000);
|
||||||
|
})();
|
||||||
|
</script>
|
||||||
|
{{else if not .AgentSupported}}
|
||||||
<div class="alert alert-info">A funkcióhoz az ügynök frissítése szükséges — a frissítés automatikusan megérkezik. Próbálja újra később.</div>
|
<div class="alert alert-info">A funkcióhoz az ügynök frissítése szükséges — a frissítés automatikusan megérkezik. Próbálja újra később.</div>
|
||||||
{{else}}
|
{{else}}
|
||||||
|
|
||||||
|
|||||||
@@ -7,6 +7,18 @@
|
|||||||
</div>
|
</div>
|
||||||
|
|
||||||
{{template "backups_flash" .}}
|
{{template "backups_flash" .}}
|
||||||
|
|
||||||
|
{{with .InterruptedRestores}}
|
||||||
|
<!-- R-550: a restore the box was running when it stopped. Persisted, so it survives the restart that
|
||||||
|
interrupted it; shown until that app's next restore. -->
|
||||||
|
<div class="settings-card">
|
||||||
|
<h3>Megszakadt visszaállítás</h3>
|
||||||
|
{{range .}}
|
||||||
|
<div class="alert alert-error"><strong>{{.Stack}}</strong>: {{.Message}}</div>
|
||||||
|
{{end}}
|
||||||
|
<p class="form-hint">A doboz újraindult, miközben a visszaállítás futott, ezért nem tudjuk, hogy befejeződött-e. Indítsd el újra lent, ugyanabból a mentésből.</p>
|
||||||
|
</div>
|
||||||
|
{{end}}
|
||||||
{{template "restore_banner" .}}
|
{{template "restore_banner" .}}
|
||||||
|
|
||||||
{{if not .Backup}}
|
{{if not .Backup}}
|
||||||
|
|||||||
@@ -34,7 +34,7 @@
|
|||||||
<div class="settings-card">
|
<div class="settings-card">
|
||||||
<h3>Eredmény</h3>
|
<h3>Eredmény</h3>
|
||||||
<div class="alert {{if .OK}}alert-info{{else}}alert-error{{end}}">{{.Message}}</div>
|
<div class="alert {{if .OK}}alert-info{{else}}alert-error{{end}}">{{.Message}}</div>
|
||||||
<p class="form-hint">Befejezve: {{fmtTime .FinishedAt}}. Ha szeretnéd, alább újra indíthatsz egy visszaállítást.</p>
|
<p class="form-hint">{{if .Interrupted}}Észlelve{{else}}Befejezve{{end}}: {{fmtTime .FinishedAt}}. Ha szeretnéd, alább újra indíthatsz egy visszaállítást.</p>
|
||||||
</div>
|
</div>
|
||||||
{{end}}
|
{{end}}
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user