v0.246.0: an interrupted restore is told; the recovery-code reminder waits until the box can take it
gates / gates (push) Successful in 15s

MinAgent: 0.131.0 (unchanged). Requires hub v0.117.0 for restore_interrupted.

R-550 (operator ruling: fix). A design reversed and recorded: the restore
op-status was in memory by choice. Now restore-status.json in DataDir, written
atomically at both ends of an op. At startup a record still marked running
becomes a failed, interrupted result kept per app until that app's next
restore, shown on /backups/restore and the off-site wizard, and raised once as
restore_interrupted. Cooldowns stay in memory.

R-546. The R-543 reminder bar consults the agent's own preflight ok (every
blocking item, not a copy of pbs_storage_id), cached 60 s, probed only while
paused. /backup/escrow shows a waiting card that polls and reloads instead of
red crosses and English diagnostics. POST /api/escrow/start refuses 409 before
staging or starting - the direct path chaos night used. Unknown readiness keeps
the bar.

Red-proofs (each seen failing): restore record across restart; main() calls
both startup functions; startup helper with loading skipped; restore page card;
bar held back; waiting card; start refusal. go build/vet/test ./... green, 28
packages; controller_gates --fast all OK.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-17 10:45:29 +02:00
parent 714d5bce09
commit 0fe315b759
21 changed files with 788 additions and 10 deletions
+9 -2
View File
@@ -6,8 +6,8 @@ import "time"
// backups page can show a progress banner (running → success/failure) instead of blocking the HTTP
// request until the restore completes. It is display-only and mutex-guarded on the Manager's `mu`;
// it does NOT gate concurrency (that stays the restore functions' internal single-flight acquire).
// In-memory only — lost on a controller restart (same precedent as notification cooldowns); a page
// load mid-op after a restart simply shows no banner.
// PERSISTED since v0.246.0 (R-550, operator ruling 2026-09-17 — a reversal of the original in-memory
// choice, for the restore record only; notification cooldowns stay in memory). See restore_record.go.
// RestoreOpResult is the terminal record of the most recent restore op.
type RestoreOpResult struct {
@@ -16,6 +16,9 @@ type RestoreOpResult struct {
OK bool `json:"ok"`
Message string `json:"message"`
FinishedAt time.Time `json:"finished_at"`
// Interrupted marks a restore that was still in flight when the controller stopped (R-550): found
// at the next start, recorded as a failure with RestoreInterruptedMessage.
Interrupted bool `json:"interrupted,omitempty"`
}
// RestoreResultWindow bounds how long a finished restore still counts as "what just happened".
@@ -54,6 +57,9 @@ func (m *Manager) BeginRestoreOp(op, stack string) {
m.opName = op
m.opStack = stack
m.opStartedAt = time.Now()
// A new restore of this app supersedes its interrupted notice (R-550).
delete(m.opInterrupted, stack)
m.persistRestoreRecordLocked()
}
// EndRestoreOp records the terminal result (called from the goroutine on completion, success or
@@ -69,6 +75,7 @@ func (m *Manager) EndRestoreOp(ok bool, message string) {
FinishedAt: time.Now(),
}
m.opRunning = false
m.persistRestoreRecordLocked()
}
// RestoreStatus returns a deep copy of the current restore op-status for the page/API.