v0.246.0: an interrupted restore is told; the recovery-code reminder waits until the box can take it
gates / gates (push) Successful in 15s

MinAgent: 0.131.0 (unchanged). Requires hub v0.117.0 for restore_interrupted.

R-550 (operator ruling: fix). A design reversed and recorded: the restore
op-status was in memory by choice. Now restore-status.json in DataDir, written
atomically at both ends of an op. At startup a record still marked running
becomes a failed, interrupted result kept per app until that app's next
restore, shown on /backups/restore and the off-site wizard, and raised once as
restore_interrupted. Cooldowns stay in memory.

R-546. The R-543 reminder bar consults the agent's own preflight ok (every
blocking item, not a copy of pbs_storage_id), cached 60 s, probed only while
paused. /backup/escrow shows a waiting card that polls and reloads instead of
red crosses and English diagnostics. POST /api/escrow/start refuses 409 before
staging or starting - the direct path chaos night used. Unknown readiness keeps
the bar.

Red-proofs (each seen failing): restore record across restart; main() calls
both startup functions; startup helper with loading skipped; restore page card;
bar held back; waiting card; start refusal. go build/vet/test ./... green, 28
packages; controller_gates --fast all OK.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-17 10:45:29 +02:00
parent 714d5bce09
commit 0fe315b759
21 changed files with 788 additions and 10 deletions
+27
View File
@@ -1333,6 +1333,24 @@ truncated.
> Those figures do not extrapolate: the structure check's cost tracks the index, read-data's tracks
> the data.
### The restore record survives a restart (v0.246.0, R-550)
The restore op-status (`internal/backup/opstatus.go`, served at `GET /api/backup/restore-status`) was
in memory only. Chaos night round 10 hard-reset a box four seconds into a restore; afterwards the
status was the Go zero value and nothing told the household whether the restore finished. **Operator
ruling 2026-09-17 reversed the in-memory choice for the restore record only** (notification cooldowns
stay in memory).
- `restore-status.json` in `DataDir` (beside `settings.json`), written atomically at **both** ends of an op.
- At startup (`loadRestoreRecordAtStartup`, before any page is served) a record still marked running
becomes a failed result — „A visszaállítás megszakadt (a doboz újraindult) — indítsd el újra." —
and a per-app notice that stays until **that app's** next restore.
- `restore_interrupted` (warning, for the household; hub v0.117.0) is pushed **once** per
interruption, after the notifier exists (`reportInterruptedRestore`). Best-effort like every
`PushEvent` (3 attempts, 3 s apart); the page notice does not depend on it.
- `/backups/restore` shows a „Megszakadt visszaállítás" card per interrupted app; the off-site
wizard's „Eredmény" card falls back to the interrupted record („Észlelve: …").
### Restore refusals (v0.226.0)
Three guards added on the off-site restore surface, all server-side:
@@ -1585,6 +1603,15 @@ dismissal, back at the next visit, gone for good when the state is `escrowed`. I
pages bypass that function, and a session check keeps it off the public guest share page. The pause
itself is UNCHANGED — it is the zero-knowledge escrow design, not a defect.
**…but only once the box can do it (v0.246.0, R-546).** For the first minutes after a bind the
agent's escrow preflight is not `ok` (typically no PBS storage yet — ~17 min measured). The bar now
consults the agent's own `ok` (`escrow_readiness.go`, cached 60 s, probed only while paused) and is
held back while the agent says not ready; `/backup/escrow` shows a waiting card („A doboz még készül
— … pár perc múlva …") that polls the preflight and reloads itself, instead of a red checklist; and
`POST /api/escrow/start` refuses 409 with the same sentence **before** staging or starting. Readiness
**unknown** (agent unreachable) keeps the bar — the fail-loud rule of R-543. A re-ceremony on an
escrowed box keeps the checklist, where a red row is a real fault.
#### Restore (`internal/backup/restore.go`)
Both **Tier 1** (restic) and **Tier 2** (rsync) restores are supported. All deployed apps