Files
felhom-controller/controller/internal/backup/opstatus.go
T
admin 985388c6e9
gates / gates (push) Successful in 10s
R-351: the restore compares where the backup says the data lived; second press cannot start a second run
Part 3 (not droppable) and the engine half of Part 2. No version bump yet - one bump and
one bake at the end of the session.

PART 3a - a second press really did start a second run. Established with a test BEFORE any
change: both offboxReconstituteHandler and offboxPlaceHandler answered "...elindult" and
overwrote the first restore's op/stack. Cause: every restore handler gated on
backupMgr.IsRunning() - the CONCURRENCY flag, which the restore goroutine acquires AFTER the
handler returns (offbox_reconstitute.go:180, offbox_restore.go:393). Seven sites. The wizard
had read the correct flag since v0.154.0 and said so in a comment; the handlers never moved.
New Server.restoreOpBlocked() reads BOTH flags - the display flag covers the whole off-box
restore, the concurrency flag is the only one the nightly backup holds - and the refusal now
names the running app and a route.

PART 3b - the page DOES refresh; the defect was the RESULT. backups_shared.html gated the
terminal result on a page-local sawRunning flag, so a restore that finished before the page
was opened, or inside one 3s poll, was shown to nobody. The 2026-08-21 OpenGist restore took
8.666s and no screen ever said it completed - the answer existed only in docker logs.
RestoreOpStatus.LastRecent now carries the server's verdict. The 10-minute window moved to
internal/backup as RestoreResultWindow and internal/web's constant is an alias: one
expression, two surfaces. Also removed the wizard's self-contradiction, which said the state
refreshes automatically AND that you must refresh the page.

PART 2 (engine) - every recovery unit manifest has carried drive and namespace_root since
schema 1, and NO non-test code read either back. The reconstitution opened the manifest and
took only the coherence stamp, then resolved its destination from the live app. A restore
into a different destination succeeded silently under a green message. New
backup/offbox_placement.go: CheckPlacement (pure, total), PlacementMismatchMessage,
recordedPlacementFromScratch. Compared before the safety dump and before the first byte.
A mismatch is NAMED and refused; ackPlacementChange lets the customer proceed deliberately -
a separate field from confirm=1, because one click must not carry two decisions. An UNKNOWN
recording is never a mismatch: refusing on an absence would strand every pre-field unit.
The not-installed refusal (R-253) now names the drive the backup recorded.

RED-PROOFS, each mutation asserted applied and reverted to 0:
  B  both guards removed (count asserted 2) -> the restore WAS seen starting with no drive
     attached: no error, full 3.00s run, wrote into /tmp/mutant-destination
  C  Mismatch forced false -> the silent divergent restore returned
  E  Known() forced true  -> the fabricated empty prefill appeared
  D  Mismatch forced true -> 8 ordinary reconstitute tests broke, proving reachability both ways
Note on D: the existing fixtures write a schema-1 manifest with NO drive, so they are
scenario-E shaped. The matching case is covered in the scenario table, not by them.

Gates 11/11 OK. Suite 28 packages ok. Hungarian verified as hex, no BOM, no mojibake sentinels.

NOT in this commit, still open: Part 2's scenario-A prefill UI, Part 1's deploy-page
visibility line, Part 1's specification document, Part 4's measurement.
2026-08-21 21:04:16 +02:00

98 lines
4.0 KiB
Go

package backup
import "time"
// Restore op-status (Part B): a lightweight, in-memory surface for the ASYNC restore family so the
// backups page can show a progress banner (running → success/failure) instead of blocking the HTTP
// request until the restore completes. It is display-only and mutex-guarded on the Manager's `mu`;
// it does NOT gate concurrency (that stays the restore functions' internal single-flight acquire).
// In-memory only — lost on a controller restart (same precedent as notification cooldowns); a page
// load mid-op after a restart simply shows no banner.
// RestoreOpResult is the terminal record of the most recent restore op.
type RestoreOpResult struct {
Op string `json:"op"` // "restore" | "tier2-restore" | "offbox-restore"
Stack string `json:"stack"`
OK bool `json:"ok"`
Message string `json:"message"`
FinishedAt time.Time `json:"finished_at"`
}
// RestoreResultWindow bounds how long a finished restore still counts as "what just happened".
//
// ONE EXPRESSION, TWO SURFACES (R-351). The wizard's phase strip and the list page's banner both
// need this bound, and until now only the wizard had one — it lived in internal/web as an
// unexported constant. A second copy in the JS would be exactly the "two copies that already
// differed" shape this repo keeps paying for, so the window is defined HERE, beside the status it
// bounds, and both surfaces read it from the payload.
const RestoreResultWindow = 10 * time.Minute
// RestoreOpStatus is the shape served at GET /api/backup/restore-status.
type RestoreOpStatus struct {
Running bool `json:"running"`
Op string `json:"op,omitempty"`
Stack string `json:"stack,omitempty"`
StartedAt time.Time `json:"started_at,omitempty"`
Last *RestoreOpResult `json:"last,omitempty"`
// LastRecent reports whether Last finished recently enough to still be worth showing to someone
// who was NOT watching when it happened.
//
// R-351, the measured defect: the banner's JS gated the terminal result on a page-local
// `sawRunning` flag, so a restore that finished before the page was opened — or in under one
// poll interval — was shown to nobody. The 2026-08-21 OpenGist restore finished in 8.7s and the
// operator could not tell from any screen whether it had completed; the answer existed only in
// a container log. A result nobody can see is the same defect class as no result at all.
LastRecent bool `json:"last_recent,omitempty"`
}
// BeginRestoreOp marks a restore op in flight (called by the handler just before launching the
// background goroutine). Idempotent enough for display; concurrency is enforced elsewhere.
func (m *Manager) BeginRestoreOp(op, stack string) {
m.mu.Lock()
defer m.mu.Unlock()
m.opRunning = true
m.opName = op
m.opStack = stack
m.opStartedAt = time.Now()
}
// EndRestoreOp records the terminal result (called from the goroutine on completion, success or
// failure). Message carries the error (failure) or a human note like the scratch path (offbox).
func (m *Manager) EndRestoreOp(ok bool, message string) {
m.mu.Lock()
defer m.mu.Unlock()
m.opLast = &RestoreOpResult{
Op: m.opName,
Stack: m.opStack,
OK: ok,
Message: message,
FinishedAt: time.Now(),
}
m.opRunning = false
}
// RestoreStatus returns a deep copy of the current restore op-status for the page/API.
func (m *Manager) RestoreStatus() RestoreOpStatus {
m.mu.Lock()
defer m.mu.Unlock()
st := RestoreOpStatus{
Running: m.opRunning,
Op: m.opName,
Stack: m.opStack,
StartedAt: m.opStartedAt,
}
if m.opLast != nil {
cp := *m.opLast
st.Last = &cp
// Recency is decided here, on the clock the result was stamped with, so neither surface has
// to hold its own copy of the window. A zero FinishedAt is never recent — an unstamped
// result must not be drawn as "just now".
if !cp.FinishedAt.IsZero() {
if d := time.Since(cp.FinishedAt); d >= 0 && d < RestoreResultWindow {
st.LastRecent = true
}
}
}
return st
}