Commit Graph

4 Commits

Author SHA1 Message Date
admin 0fe315b759 v0.246.0: an interrupted restore is told; the recovery-code reminder waits until the box can take it
gates / gates (push) Successful in 15s
MinAgent: 0.131.0 (unchanged). Requires hub v0.117.0 for restore_interrupted.

R-550 (operator ruling: fix). A design reversed and recorded: the restore
op-status was in memory by choice. Now restore-status.json in DataDir, written
atomically at both ends of an op. At startup a record still marked running
becomes a failed, interrupted result kept per app until that app's next
restore, shown on /backups/restore and the off-site wizard, and raised once as
restore_interrupted. Cooldowns stay in memory.

R-546. The R-543 reminder bar consults the agent's own preflight ok (every
blocking item, not a copy of pbs_storage_id), cached 60 s, probed only while
paused. /backup/escrow shows a waiting card that polls and reloads instead of
red crosses and English diagnostics. POST /api/escrow/start refuses 409 before
staging or starting - the direct path chaos night used. Unknown readiness keeps
the bar.

Red-proofs (each seen failing): restore record across restart; main() calls
both startup functions; startup helper with loading skipped; restore page card;
bar held back; waiting card; start refusal. go build/vet/test ./... green, 28
packages; controller_gates --fast all OK.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 10:45:29 +02:00
admin 985388c6e9 R-351: the restore compares where the backup says the data lived; second press cannot start a second run
gates / gates (push) Successful in 10s
Part 3 (not droppable) and the engine half of Part 2. No version bump yet - one bump and
one bake at the end of the session.

PART 3a - a second press really did start a second run. Established with a test BEFORE any
change: both offboxReconstituteHandler and offboxPlaceHandler answered "...elindult" and
overwrote the first restore's op/stack. Cause: every restore handler gated on
backupMgr.IsRunning() - the CONCURRENCY flag, which the restore goroutine acquires AFTER the
handler returns (offbox_reconstitute.go:180, offbox_restore.go:393). Seven sites. The wizard
had read the correct flag since v0.154.0 and said so in a comment; the handlers never moved.
New Server.restoreOpBlocked() reads BOTH flags - the display flag covers the whole off-box
restore, the concurrency flag is the only one the nightly backup holds - and the refusal now
names the running app and a route.

PART 3b - the page DOES refresh; the defect was the RESULT. backups_shared.html gated the
terminal result on a page-local sawRunning flag, so a restore that finished before the page
was opened, or inside one 3s poll, was shown to nobody. The 2026-08-21 OpenGist restore took
8.666s and no screen ever said it completed - the answer existed only in docker logs.
RestoreOpStatus.LastRecent now carries the server's verdict. The 10-minute window moved to
internal/backup as RestoreResultWindow and internal/web's constant is an alias: one
expression, two surfaces. Also removed the wizard's self-contradiction, which said the state
refreshes automatically AND that you must refresh the page.

PART 2 (engine) - every recovery unit manifest has carried drive and namespace_root since
schema 1, and NO non-test code read either back. The reconstitution opened the manifest and
took only the coherence stamp, then resolved its destination from the live app. A restore
into a different destination succeeded silently under a green message. New
backup/offbox_placement.go: CheckPlacement (pure, total), PlacementMismatchMessage,
recordedPlacementFromScratch. Compared before the safety dump and before the first byte.
A mismatch is NAMED and refused; ackPlacementChange lets the customer proceed deliberately -
a separate field from confirm=1, because one click must not carry two decisions. An UNKNOWN
recording is never a mismatch: refusing on an absence would strand every pre-field unit.
The not-installed refusal (R-253) now names the drive the backup recorded.

RED-PROOFS, each mutation asserted applied and reverted to 0:
  B  both guards removed (count asserted 2) -> the restore WAS seen starting with no drive
     attached: no error, full 3.00s run, wrote into /tmp/mutant-destination
  C  Mismatch forced false -> the silent divergent restore returned
  E  Known() forced true  -> the fabricated empty prefill appeared
  D  Mismatch forced true -> 8 ordinary reconstitute tests broke, proving reachability both ways
Note on D: the existing fixtures write a schema-1 manifest with NO drive, so they are
scenario-E shaped. The matching case is covered in the scenario table, not by them.

Gates 11/11 OK. Suite 28 packages ok. Hungarian verified as hex, no BOM, no mojibake sentinels.

NOT in this commit, still open: Part 2's scenario-A prefill UI, Part 1's deploy-page
visibility line, Part 1's specification document, Part 4's measurement.
2026-08-21 21:04:16 +02:00
admin 68b3a3932e controller: F7 atomic volume dumps + F6 no-single-copy + F5 stale-primary sweep (WIP, pre-build)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017CDMFpFx84pfviCTVuGGhf
2026-07-12 09:00:16 +02:00
admin c529a455af feat(backup): async restore family — no proxy-timeout error page on a succeeding restore (v0.102.0)
Re-adjudicates F4: /backup/restore, /backup/tier2/restore, /backup/offbox/restore
blocked the HTTP request until completion, so through cloudflared's 100s cap a
customer got an error page while the restore succeeded (offbox worse — bounded
on r.Context(), canceling the SFTP restore mid-flight). Convert all three to the
offboxRun async shape: fast-path IsRunning refuse, background goroutine
(offbox ctx off r.Context() -> Background+30m), instant redirect. Add mutex-
guarded op-status (opstatus.go) + GET /api/backup/restore-status + a 3s-polling
backups.html banner (neutral running, red on failure). Restore single-flight
unchanged. Tests + red-proof (sync handler blocks indefinitely vs <500ms async).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PSK5g6qYLknKj8u3QAFEr6
2026-07-06 20:23:49 +02:00