R-351: the restore compares where the backup says the data lived; second press cannot start a second run
gates / gates (push) Successful in 10s

Part 3 (not droppable) and the engine half of Part 2. No version bump yet - one bump and
one bake at the end of the session.

PART 3a - a second press really did start a second run. Established with a test BEFORE any
change: both offboxReconstituteHandler and offboxPlaceHandler answered "...elindult" and
overwrote the first restore's op/stack. Cause: every restore handler gated on
backupMgr.IsRunning() - the CONCURRENCY flag, which the restore goroutine acquires AFTER the
handler returns (offbox_reconstitute.go:180, offbox_restore.go:393). Seven sites. The wizard
had read the correct flag since v0.154.0 and said so in a comment; the handlers never moved.
New Server.restoreOpBlocked() reads BOTH flags - the display flag covers the whole off-box
restore, the concurrency flag is the only one the nightly backup holds - and the refusal now
names the running app and a route.

PART 3b - the page DOES refresh; the defect was the RESULT. backups_shared.html gated the
terminal result on a page-local sawRunning flag, so a restore that finished before the page
was opened, or inside one 3s poll, was shown to nobody. The 2026-08-21 OpenGist restore took
8.666s and no screen ever said it completed - the answer existed only in docker logs.
RestoreOpStatus.LastRecent now carries the server's verdict. The 10-minute window moved to
internal/backup as RestoreResultWindow and internal/web's constant is an alias: one
expression, two surfaces. Also removed the wizard's self-contradiction, which said the state
refreshes automatically AND that you must refresh the page.

PART 2 (engine) - every recovery unit manifest has carried drive and namespace_root since
schema 1, and NO non-test code read either back. The reconstitution opened the manifest and
took only the coherence stamp, then resolved its destination from the live app. A restore
into a different destination succeeded silently under a green message. New
backup/offbox_placement.go: CheckPlacement (pure, total), PlacementMismatchMessage,
recordedPlacementFromScratch. Compared before the safety dump and before the first byte.
A mismatch is NAMED and refused; ackPlacementChange lets the customer proceed deliberately -
a separate field from confirm=1, because one click must not carry two decisions. An UNKNOWN
recording is never a mismatch: refusing on an absence would strand every pre-field unit.
The not-installed refusal (R-253) now names the drive the backup recorded.

RED-PROOFS, each mutation asserted applied and reverted to 0:
  B  both guards removed (count asserted 2) -> the restore WAS seen starting with no drive
     attached: no error, full 3.00s run, wrote into /tmp/mutant-destination
  C  Mismatch forced false -> the silent divergent restore returned
  E  Known() forced true  -> the fabricated empty prefill appeared
  D  Mismatch forced true -> 8 ordinary reconstitute tests broke, proving reachability both ways
Note on D: the existing fixtures write a schema-1 manifest with NO drive, so they are
scenario-E shaped. The matching case is covered in the scenario table, not by them.

Gates 11/11 OK. Suite 28 packages ok. Hungarian verified as hex, no BOM, no mojibake sentinels.

NOT in this commit, still open: Part 2's scenario-A prefill UI, Part 1's deploy-page
visibility line, Part 1's specification document, Part 4's measurement.
This commit is contained in:
2026-08-21 21:04:16 +02:00
parent 2fa1efc5e5
commit 985388c6e9
14 changed files with 796 additions and 30 deletions
@@ -0,0 +1,146 @@
package backup
import (
"fmt"
"io/fs"
"os"
"path/filepath"
"strings"
)
// R-351 — THE RESTORE ALREADY KNOWS WHERE THE APP LIVED. IT JUST NEVER LOOKED.
//
// Every recovery unit's manifest.json carries `drive` and `namespace_root` (recovery_unit.go:48-49),
// written at capture time from the app's own live placement. Measured on demo-hp 2026-08-21:
//
// opengist drive=/mnt/sys_drive nsroot=/mnt/sys_drive/felhom-data
// calibre-web drive=/mnt/felhom-drives/hdd_1 nsroot=/mnt/felhom-drives/hdd_1
//
// Before this file, NO non-test code in the repository read either field back. `grep -rE
// '\.Drive\b|\.NamespaceRoot\b' --include=*.go` returned only the appbackup.NamespaceRoot FUNCTION
// and Tier2Target's unrelated field. The reconstitution opened the manifest
// (offbox_reconstitute.go:235) and took only the coherence stamp from it, then resolved its
// destination from the LIVE app instead. One side wrote the fact; the other never received it —
// the same shape as several defects closed this month.
//
// THE CONSEQUENCE THAT MADE THIS WORTH FIXING: a restore into a destination that differs from the
// one the backup recorded succeeded SILENTLY, under a green message. Nothing compared them.
//
// WHAT THIS IS NOT. It is not a lock. A domain can legitimately change and hardware can move, so a
// difference is NAMED and the customer decides — it is their own previous answer being shown back to
// them, not a rule imposed on them. What must never happen is the difference passing unremarked.
// RecordedPlacement is what a backup says about where an app's data lived when it was captured.
// Empty fields mean the manifest did not record them — an older unit, or one written before the
// field existed. That is an UNKNOWN and is never rendered as a value.
type RecordedPlacement struct {
Drive string // manifest.Drive — the in-guest mount (HDD_PATH), or the system data path
NamespaceRoot string // manifest.NamespaceRoot — the resolved felhom-data namespace root
}
// Known reports whether the backup recorded a destination at all. A unit whose manifest predates the
// field, or could not be read, is NOT known — and "we cannot tell" is a different answer from "they
// match", which is why this is a method rather than a `!= ""` scattered over the callers.
func (p RecordedPlacement) Known() bool { return strings.TrimSpace(p.Drive) != "" }
// PlacementCheck is the verdict of comparing what the backup recorded against where the restore is
// actually about to write.
type PlacementCheck struct {
// Recorded is what the backup said. Zero value when the manifest carried nothing.
Recorded RecordedPlacement
// LiveDrive / LiveNamespaceRoot are where this restore will write, resolved the way the
// reconstitution resolves it today.
LiveDrive string
LiveNamespaceRoot string
// Known mirrors Recorded.Known(), carried on the verdict so a template never has to re-derive it.
Known bool
// Mismatch is true ONLY when the backup recorded a destination AND it differs from the live one.
// An unknown recording is never a mismatch: refusing on an absence would block every pre-field
// unit, and asserting a match we cannot see would be worse.
Mismatch bool
}
// CheckPlacement compares a manifest's recorded placement against the live destination.
//
// Deliberately PURE and total: a nil manifest, an empty manifest and a manifest whose drive is blank
// all produce the same honest "not known, no mismatch" verdict. Paths are compared Cleaned, because
// `/mnt/x` and `/mnt/x/` are the same destination and a trailing slash must not manufacture a
// mismatch the customer then has to dismiss. Comparison is on the DRIVE, not the namespace root: the
// root is derived from the drive (namespaceRoot appends felhom-data only on the system-data
// fallback), so comparing both would report one difference twice.
func CheckPlacement(man *RecoveryManifest, liveDrive, liveNamespaceRoot string) PlacementCheck {
c := PlacementCheck{
LiveDrive: strings.TrimSpace(liveDrive),
LiveNamespaceRoot: strings.TrimSpace(liveNamespaceRoot),
}
if man != nil {
c.Recorded = RecordedPlacement{
Drive: strings.TrimSpace(man.Drive),
NamespaceRoot: strings.TrimSpace(man.NamespaceRoot),
}
}
c.Known = c.Recorded.Known()
if !c.Known || c.LiveDrive == "" {
return c
}
c.Mismatch = filepath.Clean(c.Recorded.Drive) != filepath.Clean(c.LiveDrive)
return c
}
// scratchManifestMaxDepth bounds the walk below. The unit sits a handful of levels under the scratch
// root (the restore mirrors the snapshot's absolute path), and an unbounded walk over a scratch that
// also holds a full userdata tree would stat a customer's entire library to find one small file.
const scratchManifestMaxDepth = 8
// recordedPlacementFromScratch reads the placement out of the unit manifest inside a PREPARED
// restore scratch, without restoring anything.
//
// It exists for the not-installed refusal (scenario B), which fires BEFORE the live namespace root
// can be resolved — so it cannot use the manifest the reconstitution opens later. The scratch is
// already on local disk by then; this is a bounded walk and a file read, never a network call.
//
// Returns the zero value on ANY failure — unreadable, absent, malformed. A refusal that cannot name
// the recorded place must fall back to the plain sentence rather than print an empty path, which
// would read as "the backup says it lived nowhere".
func (m *Manager) recordedPlacementFromScratch(scratch string) RecordedPlacement {
var found RecordedPlacement
root := filepath.Clean(scratch)
rootDepth := strings.Count(root, string(os.PathSeparator))
_ = filepath.WalkDir(root, func(path string, d fs.DirEntry, err error) error {
if err != nil {
return nil // an unreadable subtree is not fatal: keep looking elsewhere
}
if d.IsDir() {
if strings.Count(filepath.Clean(path), string(os.PathSeparator))-rootDepth >= scratchManifestMaxDepth {
return fs.SkipDir
}
return nil
}
if d.Name() != "manifest.json" {
return nil
}
if man := readManifest(path); man != nil && strings.TrimSpace(man.Drive) != "" {
found = RecordedPlacement{
Drive: strings.TrimSpace(man.Drive),
NamespaceRoot: strings.TrimSpace(man.NamespaceRoot),
}
return fs.SkipAll
}
return nil
})
return found
}
// PlacementMismatchMessage is the Hungarian refusal shown when the destination differs from the one
// the backup recorded. It NAMES BOTH VALUES — which is the whole point: "a destination differs" that
// does not say from what leaves the customer with a decision they cannot make.
//
// It is a refusal with a route, not a dead end: the caller re-offers the action with the
// acknowledgement field set, so the customer can proceed deliberately (scenario C).
func PlacementMismatchMessage(stack string, c PlacementCheck) string {
return fmt.Sprintf(
"A(z) %s mentése szerint az adatok korábban itt voltak: %s. Most viszont ide állna vissza: %s. "+
"Ez lehet szándékos — például ha meghajtót cseréltél —, de magától nem folytatjuk. "+
"Ha így jó, erősítsd meg alább, és a visszaállítás az új helyre fut.",
stack, c.Recorded.Drive, c.LiveDrive)
}
@@ -0,0 +1,167 @@
package backup
import (
"context"
"io/fs"
"os"
"path/filepath"
"strings"
"testing"
)
// scratchManifestPath finds the unit manifest the fixture wrote into the prepared scratch, so a test
// can rewrite it. Fails loudly rather than returning "": a silently-missing manifest would make the
// tests below pass for the wrong reason.
func scratchManifestPath(t *testing.T, m *Manager, stack string) string {
t.Helper()
scratch, _, err := m.offboxRestoreScratchDir(stack)
if err != nil {
t.Fatalf("fixture: scratch dir: %v", err)
}
var found string
_ = filepath.WalkDir(scratch, func(p string, d fs.DirEntry, err error) error {
if err == nil && !d.IsDir() && d.Name() == "manifest.json" {
found = p
return fs.SkipAll
}
return nil
})
if found == "" {
t.Fatal("fixture: no unit manifest in the prepared scratch — the test would prove nothing")
}
return found
}
// R-351 SCENARIO B — THE APP IS NOT INSTALLED AND THE BACKUP NAMES WHERE IT LIVED.
//
// The refusal already existed (R-253) and correctly stopped the restore. What it did NOT do was say
// where the data belonged, so the person on 2026-08-21 had to remember the drive and the address
// themselves — both of which the backup was holding the whole time.
//
// WRONG OUTCOME THIS PINS: "it starts, and writes somewhere else." The assertions below are that the
// restore is refused AND that the app was never touched — not merely that an error came back.
func TestReconstitute_NotInstalled_RefusalNamesTheRecordedDrive(t *testing.T) {
const recordedDrive = "/mnt/felhom-drives/hdd_1"
m, prov, _ := reconFixture(t, "20260719T060000Z", "2026-07-19T06:00:00Z", pgDump(1))
// The backup records a drive — the one this box no longer has attached.
manPath := scratchManifestPath(t, m, "immich")
if err := writeManifest(manPath, &RecoveryManifest{
SchemaVersion: 2, AppName: "immich",
Drive: recordedDrive, NamespaceRoot: recordedDrive,
}); err != nil {
t.Fatal(err)
}
// The app is not installed: no live HDD path. This is the rebuilt-machine shape.
prov.hdd = map[string]string{}
if got := prov.GetStackHDDPath("immich"); got != "" {
t.Fatalf("fixture: the app must look uninstalled, got hdd=%q", got)
}
callsBefore := len(prov.calls)
_, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err == nil {
t.Fatal("a restore with no destination must be REFUSED, not started")
}
if !strings.Contains(err.Error(), recordedDrive) {
t.Errorf("the refusal must name the drive the backup recorded (%s); got: %v", recordedDrive, err)
}
// The route, not just the reason. „Alkalmaz" is the ASCII stem of „Alkalmazások" — matching the
// stem keeps accented bytes out of the comparison (strict rule 7).
if !strings.Contains(err.Error(), "Alkalmaz") {
t.Errorf("the refusal must name a route the person can take; got: %v", err)
}
// THE OBSERVABLE THAT MATTERS: nothing was stopped, placed or started. A refusal that still
// touched the stack would be the defect wearing an error message.
if len(prov.calls) != callsBefore {
t.Errorf("a refused restore must not touch the app; provider calls: %v", prov.calls[callsBefore:])
}
}
// R-351 SCENARIO C — THE DESTINATION DIFFERS FROM THE ONE THE BACKUP RECORDED.
//
// Not blocked outright (a drive can legitimately change) and not silently accepted (which is how a
// restore lands in the wrong place under a green message). Named, and then the customer's own
// deliberate acknowledgement carries it.
func TestReconstitute_MismatchedDestination_RefusesThenProceedsOnAcknowledgement(t *testing.T) {
const recordedDrive = "/mnt/felhom-drives/hdd_1"
m, prov, _ := reconFixture(t, "20260719T060000Z", "2026-07-19T06:00:00Z", pgDump(1))
manPath := scratchManifestPath(t, m, "immich")
if err := writeManifest(manPath, &RecoveryManifest{
SchemaVersion: 2, AppName: "immich",
Drive: recordedDrive, NamespaceRoot: recordedDrive,
}); err != nil {
t.Fatal(err)
}
// The app IS installed — on a different drive from the recorded one (the fixture's temp dir).
live := prov.GetStackHDDPath("immich")
if live == "" || live == recordedDrive {
t.Fatalf("fixture: the live drive must exist and differ from the recorded one; got %q", live)
}
callsBefore := len(prov.calls)
_, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err == nil {
t.Fatal("a divergent destination must be NAMED and refused, never silently accepted")
}
for _, must := range []string{recordedDrive, live} {
if !strings.Contains(err.Error(), must) {
t.Errorf("the refusal must name %q so the customer can decide; got: %v", must, err)
}
}
if len(prov.calls) != callsBefore {
t.Errorf("the refusal must not touch the app; provider calls: %v", prov.calls[callsBefore:])
}
// ...and the customer may go ahead deliberately. Blocking outright is the other wrong outcome.
res, err := m.ReconstituteFromOffsite(context.Background(), "immich", true)
if err != nil {
t.Fatalf("an acknowledged placement change must proceed: %v", err)
}
if !res.Placement.Mismatch {
t.Error("the outcome must still record that the destination differed — a success that forgets the difference reads as a match")
}
if res.Placement.Recorded.Drive != recordedDrive {
t.Errorf("the outcome must carry the recorded drive, got %q", res.Placement.Recorded.Drive)
}
}
// recordedPlacementFromScratch must be TOTAL: an absent scratch, an unreadable one and a manifest
// with no drive all yield the zero value, never a fabricated empty path.
func TestRecordedPlacementFromScratch_UnknownStaysUnknown(t *testing.T) {
m := &Manager{}
if got := m.recordedPlacementFromScratch(filepath.Join(t.TempDir(), "does-not-exist")); got.Known() {
t.Errorf("an absent scratch must not yield a recorded placement, got %+v", got)
}
empty := t.TempDir()
if got := m.recordedPlacementFromScratch(empty); got.Known() {
t.Errorf("an empty scratch must not yield a recorded placement, got %+v", got)
}
noDrive := t.TempDir()
if err := writeManifest(filepath.Join(noDrive, "manifest.json"),
&RecoveryManifest{SchemaVersion: 1, AppName: "immich"}); err != nil {
t.Fatal(err)
}
if got := m.recordedPlacementFromScratch(noDrive); got.Known() {
t.Errorf("a manifest with no drive is an UNKNOWN, not a value; got %+v", got)
}
withDrive := t.TempDir()
deep := filepath.Join(withDrive, "a", "b", "c", "backups", "primary", "immich")
if err := os.MkdirAll(deep, 0o755); err != nil {
t.Fatal(err)
}
if err := writeManifest(filepath.Join(deep, "manifest.json"),
&RecoveryManifest{SchemaVersion: 2, AppName: "immich", Drive: "/mnt/sys_drive"}); err != nil {
t.Fatal(err)
}
got := m.recordedPlacementFromScratch(withDrive)
if !got.Known() || got.Drive != "/mnt/sys_drive" {
t.Errorf("a nested unit manifest must be found, got %+v", got)
}
}
@@ -0,0 +1,135 @@
package backup
import (
"strings"
"testing"
)
// R-351 — THE SCENARIO TABLE from the task, as a truth table.
//
// Each row states the WRONG outcome it exists to prevent, because a row whose expectation is only
// "want X" tells the next reader nothing about why the value matters.
//
// The live values in these fixtures are the ones measured on demo-hp 2026-08-21, not invented ones:
//
// calibre-web (declares a data path) drive=/mnt/felhom-drives/hdd_1
// opengist (declares NONE, 40-of-53) drive=/mnt/sys_drive
func TestCheckPlacement_ScenarioTable(t *testing.T) {
const dataDrive = "/mnt/felhom-drives/hdd_1"
const sysDrive = "/mnt/sys_drive"
for _, tc := range []struct {
name string
man *RecoveryManifest
liveDrive string
wantKnown bool
wantMismatch bool
wrongOutcome string
}{
{
name: "A/D - recorded drive is where we are restoring",
man: &RecoveryManifest{Drive: dataDrive, NamespaceRoot: dataDrive},
liveDrive: dataDrive,
wantKnown: true,
wantMismatch: false,
wrongOutcome: "a new question or obstacle on the ordinary path",
},
{
name: "C - the destination differs from what the backup recorded",
man: &RecoveryManifest{Drive: dataDrive, NamespaceRoot: dataDrive},
liveDrive: "/mnt/felhom-drives/hdd_2",
wantKnown: true,
wantMismatch: true,
wrongOutcome: "silently accepted - a restore into the wrong place under a green message",
},
{
name: "E - the backup records no drive (older unit, missing field)",
man: &RecoveryManifest{Drive: "", NamespaceRoot: ""},
liveDrive: dataDrive,
wantKnown: false,
wantMismatch: false,
wrongOutcome: "an empty prefill presented as if it were the recorded value",
},
{
name: "E - the manifest could not be read at all",
man: nil,
liveDrive: dataDrive,
wantKnown: false,
wantMismatch: false,
wrongOutcome: "a nil manifest treated as a match, or as a refusal that strands every old backup",
},
{
name: "no declared data path - the app has no field, but the record is still the truth",
man: &RecoveryManifest{Drive: sysDrive, NamespaceRoot: sysDrive + "/felhom-data"},
liveDrive: sysDrive,
wantKnown: true,
wantMismatch: false,
wrongOutcome: "falling into the unknown case just because the app has no storage field",
},
{
name: "no declared data path - and it moved to a real drive",
man: &RecoveryManifest{Drive: sysDrive, NamespaceRoot: sysDrive + "/felhom-data"},
liveDrive: dataDrive,
wantKnown: true,
wantMismatch: true,
wrongOutcome: "the 40-of-53 class silently exempted from the mismatch check",
},
{
name: "a trailing slash is the same destination",
man: &RecoveryManifest{Drive: dataDrive + "/", NamespaceRoot: dataDrive},
liveDrive: dataDrive,
wantKnown: true,
wantMismatch: false,
wrongOutcome: "a manufactured mismatch the customer has to dismiss for no reason",
},
{
name: "live destination unresolvable",
man: &RecoveryManifest{Drive: dataDrive, NamespaceRoot: dataDrive},
liveDrive: "",
wantKnown: true,
wantMismatch: false,
wrongOutcome: "comparing against nothing and calling it a difference; the not-installed refusal owns this case",
},
} {
t.Run(tc.name, func(t *testing.T) {
c := CheckPlacement(tc.man, tc.liveDrive, tc.liveDrive)
if c.Known != tc.wantKnown {
t.Errorf("Known = %v, want %v — wrong outcome guarded: %s", c.Known, tc.wantKnown, tc.wrongOutcome)
}
if c.Mismatch != tc.wantMismatch {
t.Errorf("Mismatch = %v, want %v — wrong outcome guarded: %s", c.Mismatch, tc.wantMismatch, tc.wrongOutcome)
}
})
}
}
// The refusal must NAME BOTH VALUES. "The destination differs" without saying from what leaves the
// customer holding a decision they have no way to make — which is the same defect in a politer form.
func TestPlacementMismatchMessage_NamesBothPlaces(t *testing.T) {
c := CheckPlacement(
&RecoveryManifest{Drive: "/mnt/felhom-drives/hdd_1"},
"/mnt/sys_drive", "/mnt/sys_drive/felhom-data")
if !c.Mismatch {
t.Fatal("fixture: these differ, or the message under test is never reached")
}
msg := PlacementMismatchMessage("opengist", c)
for _, must := range []string{"opengist", "/mnt/felhom-drives/hdd_1", "/mnt/sys_drive"} {
if !strings.Contains(msg, must) {
t.Errorf("the refusal must name %q; got: %s", must, msg)
}
}
}
// An UNKNOWN recording must never be rendered as a value. Scenario E's wrong outcome is precisely an
// empty string shown where a recorded path belongs, which reads as "the backup says it lived
// nowhere" — a fabricated fact.
func TestRecordedPlacement_UnknownIsNotAValue(t *testing.T) {
for _, drive := range []string{"", " ", "\t"} {
if (RecordedPlacement{Drive: drive}).Known() {
t.Errorf("a blank drive (%q) must not count as a recorded value", drive)
}
}
if !(RecordedPlacement{Drive: "/mnt/sys_drive"}).Known() {
t.Error("a real recorded drive must count as known, or the whole check is inert")
}
}
@@ -70,6 +70,12 @@ type OffsiteReconstituteResult struct {
OffsiteRunID string // "" for a pre-v0.148 snapshot — an unverified pair OffsiteRunID string // "" for a pre-v0.148 snapshot — an unverified pair
Skewed bool // the snapshot carries no coherence stamp: files and DB may differ in age Skewed bool // the snapshot carries no coherence stamp: files and DB may differ in age
LooksEmpty bool // R-44 sniff on the dump about to be replayed LooksEmpty bool // R-44 sniff on the dump about to be replayed
// Placement (R-351) is what the backup recorded about where this app's data lived, compared
// against where this restore actually wrote. Carried on the RESULT and not only on the refusal,
// so a restore that proceeded into a different destination says so in its own outcome rather
// than reporting a bare success — a warning beside a success is read as a success, so the
// difference has to survive into the message.
Placement PlacementCheck
} }
// fullPlaceCopier returns the FULL-restore file copier (nil seam → rsyncRestoreOverwrite). // fullPlaceCopier returns the FULL-restore file copier (nil seam → rsyncRestoreOverwrite).
@@ -166,7 +172,11 @@ func (m *Manager) dumpForSafety(ctx context.Context, db DiscoveredDB, dumpDir st
// overwritten to the snapshot's version (extras survive, nothing deleted), then the snapshot's own // overwritten to the snapshot's version (extras survive, nothing deleted), then the snapshot's own
// DB dump replayed, with a safety dump of the current database taken first. Requires a completed // DB dump replayed, with a safety dump of the current database taken first. Requires a completed
// FULL scratch restore (RestoreOffboxScratch with full=true). Single-flight. // FULL scratch restore (RestoreOffboxScratch with full=true). Single-flight.
func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string) (OffsiteReconstituteResult, error) { // ackPlacementChange (R-351) is the customer's DELIBERATE acknowledgement that the destination
// differs from the one the backup recorded. It is a separate act from the restore's own confirm:
// folding it into `confirm=1` would mean one click carried two decisions, which is precisely what
// R-48 exists to prevent.
func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string, ackPlacementChange bool) (OffsiteReconstituteResult, error) {
var res OffsiteReconstituteResult var res OffsiteReconstituteResult
if !m.OffboxConfigured() { if !m.OffboxConfigured() {
return res, fmt.Errorf("off-box backup not configured") return res, fmt.Errorf("off-box backup not configured")
@@ -203,6 +213,16 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string) (Of
// destination is the app's own HDD path, which is a drive the CUSTOMER chooses at deploy // destination is the app's own HDD path, which is a drive the CUSTOMER chooses at deploy
// time, and picking it for them is the decision this whole recovery path exists to leave // time, and picking it for them is the decision this whole recovery path exists to leave
// with them. // with them.
// R-351: the refusal now NAMES the place the backup recorded, when it can read it. The
// prepared scratch already contains the unit, so this is a local file read — no network call,
// nothing restored, and it happens on a path that was going to refuse anyway. Telling
// somebody to reinstall without telling them where the data belongs is what forced the
// 2026-08-21 operator to remember two values the backup already held.
if rec := m.recordedPlacementFromScratch(scratch); rec.Known() {
return res, fmt.Errorf("a(z) %s nincs telepítve, ezért nincs hová visszaállítani az adatait. "+
"A mentése szerint az adatai itt voltak: %s. Telepítsd újra az alkalmazást (Alkalmazások) "+
"ugyanerre a helyre, utána ez a visszaállítás működni fog", stack, rec.Drive)
}
return res, fmt.Errorf("a(z) %s nincs telepítve, ezért nincs hová visszaállítani az adatait — "+ return res, fmt.Errorf("a(z) %s nincs telepítve, ezért nincs hová visszaállítani az adatait — "+
"telepítsd újra az alkalmazást (Alkalmazások), utána ez a visszaállítás működni fog", stack) "telepítsd újra az alkalmazást (Alkalmazások), utána ez a visszaállítás működni fog", stack)
} }
@@ -232,7 +252,8 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string) (Of
return res, fmt.Errorf("a pillanatképben nincs mentési egység — a visszaállítás nem indítható") return res, fmt.Errorf("a pillanatképben nincs mentési egység — a visszaállítás nem indítható")
} }
scratchDumpDir := filepath.Join(scratchUnit, "db-dumps") scratchDumpDir := filepath.Join(scratchUnit, "db-dumps")
if man := readManifest(filepath.Join(scratchUnit, "manifest.json")); man != nil { man := readManifest(filepath.Join(scratchUnit, "manifest.json"))
if man != nil {
res.OffsiteRunID = man.OffsiteRunID res.OffsiteRunID = man.OffsiteRunID
if man.DumpsAt != "" { if man.DumpsAt != "" {
if t, pErr := time.Parse(time.RFC3339, man.DumpsAt); pErr == nil { if t, pErr := time.Parse(time.RFC3339, man.DumpsAt); pErr == nil {
@@ -240,6 +261,23 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string) (Of
} }
} }
} }
// --- WHERE THE BACKUP SAYS THIS DATA LIVED (R-351) ------------------------------------------
// The manifest we just opened has carried `drive` and `namespace_root` since schema 1, and until
// now nothing read them back. Compared HERE, before the safety dump and before the first byte is
// placed, so the refusal costs nothing and leaves the app completely untouched.
//
// An UNKNOWN recording (a pre-field unit, or one we could not read) is not a mismatch and does
// not refuse: blocking on an absence would strand every older backup, and CheckPlacement returns
// that case explicitly rather than letting it fall through as "they match".
res.Placement = CheckPlacement(man, hdd, liveNs)
if res.Placement.Mismatch && !ackPlacementChange {
return res, fmt.Errorf("%s", PlacementMismatchMessage(stack, res.Placement))
}
if res.Placement.Mismatch {
m.logger.Printf("[WARN] [offbox] %s: restoring into %s, but the backup recorded %s — the customer acknowledged the change",
stack, res.Placement.LiveDrive, res.Placement.Recorded.Drive)
}
// A pre-v0.148 snapshot carries no stamp: its dump was whatever the 02:30 local run left behind, // A pre-v0.148 snapshot carries no stamp: its dump was whatever the 02:30 local run left behind,
// so the pair's two halves may be hours or days apart. Surfaced, never blocked — the confirm // so the pair's two halves may be hours or days apart. Surfaced, never blocked — the confirm
// dialog says so and the safety dump makes it reversible. // dialog says so and the safety dump makes it reversible.
@@ -174,7 +174,7 @@ func reconFixture(t *testing.T, runID, dumpsAt string, dumpBody string) (*Manage
func TestReconstituteReplaysDBAndOrdersOperations(t *testing.T) { func TestReconstituteReplaysDBAndOrdersOperations(t *testing.T) {
m, prov, imported := reconFixture(t, "20260719T060000Z", "2026-07-19T06:00:00Z", pgDump(1)) m, prov, imported := reconFixture(t, "20260719T060000Z", "2026-07-19T06:00:00Z", pgDump(1))
res, err := m.ReconstituteFromOffsite(context.Background(), "immich") res, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err != nil { if err != nil {
t.Fatalf("reconstitute: %v", err) t.Fatalf("reconstitute: %v", err)
} }
@@ -220,7 +220,7 @@ func TestReconstituteRefusesWhenSafetyDumpFails(t *testing.T) {
var copied bool var copied bool
m.SetOffboxFullPlaceCopier(func(_, _ string) (int, error) { copied = true; return 1, nil }) m.SetOffboxFullPlaceCopier(func(_, _ string) (int, error) { copied = true; return 1, nil })
_, err := m.ReconstituteFromOffsite(context.Background(), "immich") _, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err == nil { if err == nil {
t.Fatal("expected a refusal when the safety dump cannot be taken") t.Fatal("expected a refusal when the safety dump cannot be taken")
} }
@@ -246,7 +246,7 @@ func TestReconstituteNoDBAppMakesNoDumpOrImportCalls(t *testing.T) {
return DumpResult{DB: d} return DumpResult{DB: d}
}) })
res, err := m.ReconstituteFromOffsite(context.Background(), "immich") res, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err != nil { if err != nil {
t.Fatalf("reconstitute: %v", err) t.Fatalf("reconstitute: %v", err)
} }
@@ -267,7 +267,7 @@ func TestReconstituteNoDBAppMakesNoDumpOrImportCalls(t *testing.T) {
func TestReconstituteSurfacesLegacySkewedPair(t *testing.T) { func TestReconstituteSurfacesLegacySkewedPair(t *testing.T) {
m, _, imported := reconFixture(t, "", "", pgDump(1)) m, _, imported := reconFixture(t, "", "", pgDump(1))
res, err := m.ReconstituteFromOffsite(context.Background(), "immich") res, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err != nil { if err != nil {
t.Fatalf("a legacy pair must still be restorable, got refusal: %v", err) t.Fatalf("a legacy pair must still be restorable, got refusal: %v", err)
} }
@@ -286,7 +286,7 @@ func TestReconstituteFlagsCustomerEmptyDump(t *testing.T) {
// A valid postgres dump whose accounts table has NO rows — the 2026-07-19 shape exactly. // A valid postgres dump whose accounts table has NO rows — the 2026-07-19 shape exactly.
m, _, _ := reconFixture(t, "run1", "2026-07-19T06:00:00Z", pgDump(0)) m, _, _ := reconFixture(t, "run1", "2026-07-19T06:00:00Z", pgDump(0))
res, err := m.ReconstituteFromOffsite(context.Background(), "immich") res, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err != nil { if err != nil {
t.Fatalf("the sniff must never block a restore: %v", err) t.Fatalf("the sniff must never block a restore: %v", err)
} }
+26
View File
@@ -18,6 +18,15 @@ type RestoreOpResult struct {
FinishedAt time.Time `json:"finished_at"` FinishedAt time.Time `json:"finished_at"`
} }
// RestoreResultWindow bounds how long a finished restore still counts as "what just happened".
//
// ONE EXPRESSION, TWO SURFACES (R-351). The wizard's phase strip and the list page's banner both
// need this bound, and until now only the wizard had one — it lived in internal/web as an
// unexported constant. A second copy in the JS would be exactly the "two copies that already
// differed" shape this repo keeps paying for, so the window is defined HERE, beside the status it
// bounds, and both surfaces read it from the payload.
const RestoreResultWindow = 10 * time.Minute
// RestoreOpStatus is the shape served at GET /api/backup/restore-status. // RestoreOpStatus is the shape served at GET /api/backup/restore-status.
type RestoreOpStatus struct { type RestoreOpStatus struct {
Running bool `json:"running"` Running bool `json:"running"`
@@ -25,6 +34,15 @@ type RestoreOpStatus struct {
Stack string `json:"stack,omitempty"` Stack string `json:"stack,omitempty"`
StartedAt time.Time `json:"started_at,omitempty"` StartedAt time.Time `json:"started_at,omitempty"`
Last *RestoreOpResult `json:"last,omitempty"` Last *RestoreOpResult `json:"last,omitempty"`
// LastRecent reports whether Last finished recently enough to still be worth showing to someone
// who was NOT watching when it happened.
//
// R-351, the measured defect: the banner's JS gated the terminal result on a page-local
// `sawRunning` flag, so a restore that finished before the page was opened — or in under one
// poll interval — was shown to nobody. The 2026-08-21 OpenGist restore finished in 8.7s and the
// operator could not tell from any screen whether it had completed; the answer existed only in
// a container log. A result nobody can see is the same defect class as no result at all.
LastRecent bool `json:"last_recent,omitempty"`
} }
// BeginRestoreOp marks a restore op in flight (called by the handler just before launching the // BeginRestoreOp marks a restore op in flight (called by the handler just before launching the
@@ -66,6 +84,14 @@ func (m *Manager) RestoreStatus() RestoreOpStatus {
if m.opLast != nil { if m.opLast != nil {
cp := *m.opLast cp := *m.opLast
st.Last = &cp st.Last = &cp
// Recency is decided here, on the clock the result was stamped with, so neither surface has
// to hold its own copy of the window. A zero FinishedAt is never recent — an unstamped
// result must not be drawn as "just now".
if !cp.FinishedAt.IsZero() {
if d := time.Since(cp.FinishedAt); d >= 0 && d < RestoreResultWindow {
st.LastRecent = true
}
}
} }
return st return st
} }
@@ -0,0 +1,84 @@
package backup
import (
"testing"
"time"
)
// R-351 — A FINISHED RESTORE MUST BE VISIBLE TO SOMEONE WHO WAS NOT WATCHING.
//
// THE MEASURED CASE (demo-hp, 2026-08-21). An off-box reconstitution refused at 16:37:14, the person
// reinstalled and ran the local unit restore, and it completed at 16:39:25 — in 8.666s. No screen
// ever said so. The banner's JS gated its terminal result on a page-local `sawRunning` flag, so the
// result was rendered only for a browser that happened to be open and polling across the transition.
// Land on the page a moment later and the banner stayed hidden: "completed" and "never ran" looked
// identical. The answer existed only in `docker logs felhom-controller`, which a customer cannot
// reach.
//
// This pins the CONSEQUENCE — the status carries a recency verdict a late arrival can act on — not
// the mechanism. Mutating LastRecent to stay false must fail this test.
func TestRestoreStatus_LastRecent(t *testing.T) {
for _, tc := range []struct {
name string
finishedAt time.Time
wantRecent bool
why string
}{
{
name: "just finished",
finishedAt: time.Now().Add(-9 * time.Second),
wantRecent: true,
why: "the 8.7s OpenGist restore: finished before a poll could see it running",
},
{
name: "inside the window",
finishedAt: time.Now().Add(-RestoreResultWindow + time.Minute),
wantRecent: true,
why: "still answers \"what just happened\"",
},
{
name: "outside the window",
finishedAt: time.Now().Add(-RestoreResultWindow - time.Minute),
wantRecent: false,
why: "landing here later must not claim a restore just finished",
},
{
name: "unstamped result",
finishedAt: time.Time{},
wantRecent: false,
why: "a zero timestamp is an UNKNOWN and must never be drawn as \"just now\" (presence is not success)",
},
} {
t.Run(tc.name, func(t *testing.T) {
m := &Manager{}
m.opLast = &RestoreOpResult{
Op: "restore", Stack: "opengist", OK: true,
Message: "Restore completed", FinishedAt: tc.finishedAt,
}
st := m.RestoreStatus()
if st.Last == nil {
t.Fatal("fixture: the status must carry the terminal result")
}
if st.LastRecent != tc.wantRecent {
t.Errorf("LastRecent = %v, want %v — %s", st.LastRecent, tc.wantRecent, tc.why)
}
})
}
}
// A RUNNING op must not be reported as a recent RESULT — the banner shows one or the other, and
// conflating them would put a success message over an operation that is still writing.
func TestRestoreStatus_RunningIsNotAResult(t *testing.T) {
m := &Manager{}
m.BeginRestoreOp("offbox-reconstitute", "opengist")
st := m.RestoreStatus()
if !st.Running {
t.Fatal("fixture: the op must be reported running")
}
if st.Last != nil {
t.Errorf("a fresh op has no terminal result yet, got %+v", st.Last)
}
if st.LastRecent {
t.Error("a running op must never carry a recent-result verdict")
}
}
@@ -71,7 +71,7 @@ func TestReconstituteReplaysWithOnlyTheDBServiceUp(t *testing.T) {
return nil return nil
} }
res, err := m.ReconstituteFromOffsite(context.Background(), "immich") res, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err != nil { if err != nil {
t.Fatalf("reconstitute: %v", err) t.Fatalf("reconstitute: %v", err)
} }
@@ -109,7 +109,7 @@ func TestReconstituteNoDBAppNeverStartsServicesOnly(t *testing.T) {
prov.composePath = writeLiveCompose(t, noDBCompose) prov.composePath = writeLiveCompose(t, noDBCompose)
m.discoverDBs = func(context.Context) ([]DiscoveredDB, error) { return nil, nil } m.discoverDBs = func(context.Context) ([]DiscoveredDB, error) { return nil, nil }
res, err := m.ReconstituteFromOffsite(context.Background(), "immich") res, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err != nil { if err != nil {
t.Fatalf("a no-DB app must restore unchanged, got: %v", err) t.Fatalf("a no-DB app must restore unchanged, got: %v", err)
} }
@@ -141,7 +141,7 @@ func TestReconstituteRefusesWhenNoDBServiceIdentifiable(t *testing.T) {
var copied bool var copied bool
m.SetOffboxFullPlaceCopier(func(_, _ string) (int, error) { copied = true; return 1, nil }) m.SetOffboxFullPlaceCopier(func(_, _ string) (int, error) { copied = true; return 1, nil })
_, err := m.ReconstituteFromOffsite(context.Background(), "immich") _, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err == nil { if err == nil {
t.Fatal("expected a refusal: a dump exists but no database service can be started for it") t.Fatal("expected a refusal: a dump exists but no database service can be started for it")
} }
@@ -270,7 +270,7 @@ func TestReconstituteReplayFailureStillBringsTheStackUp(t *testing.T) {
return context.DeadlineExceeded return context.DeadlineExceeded
} }
res, err := m.ReconstituteFromOffsite(context.Background(), "immich") res, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err == nil { if err == nil {
t.Fatal("a failed replay must be surfaced, not swallowed") t.Fatal("a failed replay must be surfaced, not swallowed")
} }
@@ -292,7 +292,7 @@ func TestReconstituteDBOnlyStartFailureStillBringsTheStackUp(t *testing.T) {
m, prov, imported := reconFixture(t, "run1", "2026-07-19T06:00:00Z", pgDump(1)) m, prov, imported := reconFixture(t, "run1", "2026-07-19T06:00:00Z", pgDump(1))
prov.startSvcErr = context.DeadlineExceeded prov.startSvcErr = context.DeadlineExceeded
if _, err := m.ReconstituteFromOffsite(context.Background(), "immich"); err == nil { if _, err := m.ReconstituteFromOffsite(context.Background(), "immich", false); err == nil {
t.Fatal("a failed DB-only start must be surfaced") t.Fatal("a failed DB-only start must be surfaced")
} }
if !prov.fullStarted { if !prov.fullStarted {
+4 -4
View File
@@ -1401,8 +1401,8 @@ func (s *Server) backupRestoreHandler(w http.ResponseWriter, r *http.Request) {
// Part B: restore is a long SYNCHRONOUS op (F4 — through cloudflared's hard 100s cap the customer // Part B: restore is a long SYNCHRONOUS op (F4 — through cloudflared's hard 100s cap the customer
// got an error page while it silently succeeded). Fast-path refuse a concurrent op, then run it in // got an error page while it silently succeeded). Fast-path refuse a concurrent op, then run it in
// a BACKGROUND goroutine (survives the request; the poll banner shows progress → result). // a BACKGROUND goroutine (survives the request; the poll banner shows progress → result).
if s.backupMgr.IsRunning() { if msg, blocked := s.restoreOpBlocked(); blocked {
http.Redirect(w, r, "/backups/restore?flash_error="+url.QueryEscape("Egy mentési/visszaállítási művelet már fut."), http.StatusFound) http.Redirect(w, r, "/backups/restore?flash_error="+url.QueryEscape(msg), http.StatusFound)
return return
} }
s.logger.Printf("[WARN] [web] Restore requested (async): stack=%s, snapshot=%s from %s", stackName, snapshotID, r.RemoteAddr) s.logger.Printf("[WARN] [web] Restore requested (async): stack=%s, snapshot=%s from %s", stackName, snapshotID, r.RemoteAddr)
@@ -1460,8 +1460,8 @@ func (s *Server) backupTier2RestoreHandler(w http.ResponseWriter, r *http.Reques
return return
} }
// Part B (same async shape as backupRestoreHandler): fast-path refuse, then background goroutine. // Part B (same async shape as backupRestoreHandler): fast-path refuse, then background goroutine.
if s.backupMgr.IsRunning() { if msg, blocked := s.restoreOpBlocked(); blocked {
http.Redirect(w, r, "/backups/apps?flash_error="+url.QueryEscape("Egy mentési/visszaállítási művelet már fut."), http.StatusFound) http.Redirect(w, r, "/backups/apps?flash_error="+url.QueryEscape(msg), http.StatusFound)
return return
} }
+15 -11
View File
@@ -352,10 +352,10 @@ func (s *Server) offboxRestoreHandler(w http.ResponseWriter, r *http.Request) {
} }
// Fast-path refuse a concurrent op, then run async on a BACKGROUND context (a proxy read-timeout on // Fast-path refuse a concurrent op, then run async on a BACKGROUND context (a proxy read-timeout on
// r.Context() would CANCEL the SFTP restore mid-flight — the F4 lesson). // r.Context() would CANCEL the SFTP restore mid-flight — the F4 lesson).
if s.backupMgr.IsRunning() { if msg, blocked := s.restoreOpBlocked(); blocked {
// Same silence class as the size gate: a refusal that starts nothing must still be findable. // Same silence class as the size gate: a refusal that starts nothing must still be findable.
s.logger.Printf("[WARN] [web] off-box restore refused for %s (mode=%s): another backup/restore op is already running", app, mode) s.logger.Printf("[WARN] [web] off-box restore refused for %s (mode=%s): another backup/restore op is already running", app, mode)
offboxRedirectTo(w, r, restoreWizardPath(app), "Egy mentési/visszaállítási művelet már fut.", true) offboxRedirectTo(w, r, restoreWizardPath(app), msg, true)
return return
} }
full := mode == "full" full := mode == "full"
@@ -433,15 +433,19 @@ func (s *Server) offboxReconstituteHandler(w http.ResponseWriter, r *http.Reques
offboxRedirectTo(w, r, restoreWizardPath(app), "A teljes visszaállítás megerősítés nélkül nem hajtható végre.", true) offboxRedirectTo(w, r, restoreWizardPath(app), "A teljes visszaállítás megerősítés nélkül nem hajtható végre.", true)
return return
} }
if s.backupMgr.IsRunning() { if msg, blocked := s.restoreOpBlocked(); blocked {
offboxRedirectTo(w, r, restoreWizardPath(app), "Egy mentési/visszaállítási művelet már fut.", true) offboxRedirectTo(w, r, restoreWizardPath(app), msg, true)
return return
} }
// R-351: a SEPARATE field from `confirm`. The restore's own confirm answers "overwrite my live
// data"; this one answers "yes, into a different place than the backup recorded". One checkbox
// carrying both would be the two-decisions-one-button shape R-48 removed from this surface.
ackPlacement := r.FormValue("ack_placement") == "1"
s.backupMgr.BeginRestoreOp("offbox-reconstitute", app) s.backupMgr.BeginRestoreOp("offbox-reconstitute", app)
go func() { go func() {
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Minute) ctx, cancel := context.WithTimeout(context.Background(), 60*time.Minute)
defer cancel() defer cancel()
res, err := s.backupMgr.ReconstituteFromOffsite(ctx, app) res, err := s.backupMgr.ReconstituteFromOffsite(ctx, app, ackPlacement)
if err != nil { if err != nil {
s.logger.Printf("[ERROR] [web] off-box reconstitute %s (async): %v", app, err) s.logger.Printf("[ERROR] [web] off-box reconstitute %s (async): %v", app, err)
s.backupMgr.EndRestoreOp(false, "A teljes visszaállítás sikertelen: "+err.Error()) s.backupMgr.EndRestoreOp(false, "A teljes visszaállítás sikertelen: "+err.Error())
@@ -521,8 +525,8 @@ func (s *Server) offboxPlaceHandler(w http.ResponseWriter, r *http.Request) {
offboxRedirectTo(w, r, "/backups/restore", "Hiányzó alkalmazás.", true) offboxRedirectTo(w, r, "/backups/restore", "Hiányzó alkalmazás.", true)
return return
} }
if s.backupMgr.IsRunning() { if msg, blocked := s.restoreOpBlocked(); blocked {
offboxRedirectTo(w, r, restoreWizardPath(app), "Egy mentési/visszaállítási művelet már fut.", true) offboxRedirectTo(w, r, restoreWizardPath(app), msg, true)
return return
} }
s.backupMgr.BeginRestoreOp("offbox-place", app) s.backupMgr.BeginRestoreOp("offbox-place", app)
@@ -554,8 +558,8 @@ func (s *Server) sharesRestoreHandler(w http.ResponseWriter, r *http.Request) {
offboxRedirectTo(w, r, "/backups/restore", "A távoli mentési cél nincs beállítva.", true) offboxRedirectTo(w, r, "/backups/restore", "A távoli mentési cél nincs beállítva.", true)
return return
} }
if s.backupMgr.IsRunning() { if msg, blocked := s.restoreOpBlocked(); blocked {
offboxRedirectTo(w, r, "/backups/restore", "Egy mentési/visszaállítási művelet már fut.", true) offboxRedirectTo(w, r, "/backups/restore", msg, true)
return return
} }
s.backupMgr.BeginRestoreOp("shares-restore", backup.SharesDisplayName) s.backupMgr.BeginRestoreOp("shares-restore", backup.SharesDisplayName)
@@ -580,8 +584,8 @@ func (s *Server) sharesPlaceHandler(w http.ResponseWriter, r *http.Request) {
offboxRedirectTo(w, r, "/backups/restore", "A távoli mentési cél nincs beállítva.", true) offboxRedirectTo(w, r, "/backups/restore", "A távoli mentési cél nincs beállítva.", true)
return return
} }
if s.backupMgr.IsRunning() { if msg, blocked := s.restoreOpBlocked(); blocked {
offboxRedirectTo(w, r, "/backups/restore", "Egy mentési/visszaállítási művelet már fut.", true) offboxRedirectTo(w, r, "/backups/restore", msg, true)
return return
} }
s.backupMgr.BeginRestoreOp("shares-place", backup.SharesDisplayName) s.backupMgr.BeginRestoreOp("shares-place", backup.SharesDisplayName)
@@ -0,0 +1,116 @@
package web
import (
"net/http/httptest"
"net/url"
"strings"
"testing"
"gitea.dooplex.hu/admin/felhom-controller/internal/settings"
)
// R-351 — THE SECOND PRESS. Part 3 asked whether a second press really starts a second run or is
// refused somewhere deeper. It is NOT refused: it starts a second run, and the customer is told so.
//
// WHY THE EXISTING GUARD DOES NOT CATCH IT. Every restore handler gates on `backupMgr.IsRunning()`
// (offbox_handlers.go:355/436/524/557/583, handlers.go:1404/1463). That reads `m.running`, the
// CONCURRENCY single-flight, which is acquired INSIDE the restore function on the background
// goroutine (offbox_reconstitute.go:180, offbox_restore.go:393) — not by the handler. So between the
// handler's check and the goroutine's acquire there is a window in which `IsRunning()` is false while
// a restore is unmistakably in flight. `restoreOpInFlight` and the whole wizard already read the
// other flag (`RestoreStatus().Running`, set synchronously by BeginRestoreOp) for exactly this
// reason — see the long note on restoreOpInFlight. The handlers were never moved over.
//
// THE FIXTURE IS THE LIVE STATE, NOT AN INVENTED ONE: display flag set, concurrency flag NOT held.
// That is precisely what the box looks like for the entire duration of an off-box restore.
//
// This test pins the CONSEQUENCE (does a second run start?), not the mechanism (which flag is read),
// per CLAUDE.md's preference. Mutating the new guard back to `IsRunning()` must make it fail.
func TestRestoreHandlers_SecondPressDoesNotStartASecondRun(t *testing.T) {
const firstApp = "alpha"
const secondApp = "beta"
s, sett, m := newOffboxWebServer(t)
if err := sett.SetOffboxTarget(&settings.OffboxTarget{
Enabled: true, Host: "nas.local", Port: 22, User: "felhom", RepoPath: "/srv/repo",
Schedule: "daily", EscrowState: "escrowed",
}); err != nil {
t.Fatal(err)
}
if err := m.WriteOffboxSecrets("PRIVATE-KEY-MATERIAL", "nas.local ssh-ed25519 AAAAhostkey"); err != nil {
t.Fatal(err)
}
if !m.OffboxConfigured() {
t.Fatal("fixture: the target must be configured, or the handler exits earlier and proves nothing")
}
// A restore is in flight, exactly as the live box has it.
m.BeginRestoreOp("offbox-restore", firstApp)
// The fixture must reproduce the GAP, or this test is vacuous: the display flag says running,
// the concurrency flag — the one every handler reads — says it is not.
if !m.RestoreStatus().Running {
t.Fatal("fixture: the display flag must report a running op")
}
if m.IsRunning() {
t.Fatal("fixture: the concurrency flag must NOT be held — that gap IS the defect under test")
}
post := func(path string, form url.Values) *httptest.ResponseRecorder {
r := httptest.NewRequest("POST", path, strings.NewReader(form.Encode()))
r.Header.Set("Content-Type", "application/x-www-form-urlencoded")
w := httptest.NewRecorder()
switch path {
case "/backup/offbox/reconstitute":
s.offboxReconstituteHandler(w, r)
case "/backup/offbox/place":
s.offboxPlaceHandler(w, r)
default:
t.Fatalf("unrouted path %q", path)
}
return w
}
for _, tc := range []struct {
name string
path string
form url.Values
}{
{"reconstitute", "/backup/offbox/reconstitute", url.Values{"app": {secondApp}, "confirm": {"1"}}},
{"place", "/backup/offbox/place", url.Values{"app": {secondApp}}},
} {
t.Run(tc.name, func(t *testing.T) {
w := post(tc.path, tc.form)
if w.Code != 302 {
t.Fatalf("the handler redirects; got %d", w.Code)
}
loc := w.Header().Get("Location")
// CONSEQUENCE 1 — the customer must not be told a second restore started.
// "elind" is the ASCII stem shared by every „elindult" flash; matching it avoids putting
// accented bytes through a comparison (strict rule 7).
if strings.Contains(loc, "elind") {
t.Errorf("a second press while a restore runs must NOT report a started restore. Location: %q", loc)
}
// CONSEQUENCE 2 — the in-flight op must still be the FIRST one. If the handler ran,
// BeginRestoreOp overwrote the op name and stack, so the first restore's identity is
// gone from the status the banner reads.
st := m.RestoreStatus()
if st.Op != "offbox-restore" || st.Stack != firstApp {
t.Errorf("the first restore's identity was overwritten by the second press: op=%q stack=%q (want offbox-restore/%s)",
st.Op, st.Stack, firstApp)
}
// CONSEQUENCE 3 — a refusal must name a route the person can act on, not just a reason.
// The wizard path is the route; it is where the live status is shown.
if !strings.Contains(loc, "/backups/restore") {
t.Errorf("the refusal must route somewhere actionable; got %q", loc)
}
// Restore the fixture for the next subtest — a handler that (today) ran will have
// clobbered it.
m.BeginRestoreOp("offbox-restore", firstApp)
})
}
}
+43 -1
View File
@@ -113,7 +113,12 @@ const (
// Without a bound the last result would light that phase forever — landing on the page a week later // Without a bound the last result would light that phase forever — landing on the page a week later
// would claim you had just finished a restore. Same reasoning as escrowCeremonyGraceWindow; shorter, // would claim you had just finished a restore. Same reasoning as escrowCeremonyGraceWindow; shorter,
// because this answers "what just happened", not "are we still waiting". // because this answers "what just happened", not "are we still waiting".
const restoreResultWindow = 10 * time.Minute //
// R-351: this is now an ALIAS, not a second value. The list page's banner needs the same bound, and
// the payload carries the verdict (RestoreOpStatus.LastRecent), so the window is defined once in
// internal/backup beside the status it bounds. Keeping a separate literal here is how the two
// surfaces would drift.
const restoreResultWindow = backup.RestoreResultWindow
// hasRecentRestoreResult reports whether THIS app has a just-finished restore to show. Pure (the // hasRecentRestoreResult reports whether THIS app has a just-finished restore to show. Pure (the
// clock is a parameter) so the boundary and the wrong-app case are table-testable. // clock is a parameter) so the boundary and the wrong-app case are table-testable.
@@ -182,6 +187,43 @@ func restoreOpInFlight(st backup.RestoreOpStatus) bool {
return st.Running return st.Running
} }
// restoreOpBlocked reports whether a NEW restore must be refused right now, and returns the
// Hungarian refusal to show. It reads BOTH flags, deliberately:
//
// - `RestoreStatus().Running` — the DISPLAY flag, set synchronously by `BeginRestoreOp` in the
// handler. It is the only one that is true for the WHOLE duration of an off-box restore, which
// is what makes it the right flag to refuse on.
// - `IsRunning()` — the CONCURRENCY flag. The nightly backup run holds this one and never calls
// `BeginRestoreOp`, so dropping it would open a hole the old guard did close. Kept, not replaced.
//
// R-351, the measured defect: every restore handler read ONLY `IsRunning()`, which the restore
// goroutine acquires AFTER the handler has already returned (offbox_reconstitute.go:180,
// offbox_restore.go:393). A second press inside that window started a second run and was told
// „…elindult". Pinned by TestRestoreHandlers_SecondPressDoesNotStartASecondRun, which asserts the
// CONSEQUENCE — that the first restore's identity survives the second press — rather than which
// flag was read.
//
// The refusal names a reason AND a route: the page it redirects to is the wizard, which carries the
// live status banner, so „ezen az oldalon" is a true instruction and not a gesture.
func (s *Server) restoreOpBlocked() (string, bool) {
if s.backupMgr == nil {
return "", false
}
if st := s.backupMgr.RestoreStatus(); restoreOpInFlight(st) {
subject := "Egy visszaállítási művelet"
if st.Stack != "" {
subject = "Egy visszaállítási művelet (" + st.Stack + ")"
}
return subject + " már fut, ezért most nem indítható újabb. Az állapotát ezen az oldalon " +
"követheted; amint befejeződik, újra indíthatsz visszaállítást.", true
}
if s.backupMgr.IsRunning() {
return "Egy mentési művelet már fut, ezért most nem indítható visszaállítás. Az állapotát " +
"ezen az oldalon követheted; amint befejeződik, újra indíthatsz visszaállítást.", true
}
return "", false
}
// backupsRestoreWizardHandler renders GET /backups/restore/app?name=<app> — the single entry the // backupsRestoreWizardHandler renders GET /backups/restore/app?name=<app> — the single entry the
// list page now offers per app. // list page now offers per app.
// //
@@ -45,7 +45,11 @@
<div class="settings-card"> <div class="settings-card">
<h3>Végrehajtás</h3> <h3>Végrehajtás</h3>
<p>Jelenleg egy mentési vagy visszaállítási művelet fut{{with .RunningStack}} ({{.}}){{end}}. Amíg ez tart, új visszaállítás nem indítható.</p> <p>Jelenleg egy mentési vagy visszaállítási művelet fut{{with .RunningStack}} ({{.}}){{end}}. Amíg ez tart, új visszaállítás nem indítható.</p>
<p class="form-hint">Az állapot fent automatikusan frissül. A művelet befejezése után frissítsd az oldalt.</p> <!-- R-351: this used to say the state refreshes automatically AND that you must refresh the page
when it finishes. Both cannot be true, and the second half was the one people believed. The
banner now carries the terminal result too (RestoreOpStatus.LastRecent), so the sentence can
describe what the page actually does. -->
<p class="form-hint">Az állapot fent automatikusan frissül, és a művelet eredménye is ott jelenik meg, amint elkészült.</p>
<div class="form-actions"> <div class="form-actions">
<a href="/backups/restore/app?name={{.App}}" class="btn btn-sm btn-outline">Állapot frissítése</a> <a href="/backups/restore/app?name={{.App}}" class="btn btn-sm btn-outline">Állapot frissítése</a>
</div> </div>
@@ -40,7 +40,11 @@
banner.textContent = opLabel(st.op) + ' folyamatban' + (st.stack ? ': ' + st.stack : '') + '…'; banner.textContent = opLabel(st.op) + ' folyamatban' + (st.stack ? ': ' + st.stack : '') + '…';
return; return;
} }
if (st.last && sawRunning) { // R-351: `sawRunning` alone showed a terminal result ONLY to a page that watched the op
// happen. A restore that finished before this page was opened — or inside one poll interval
// — was shown to nobody, which is how a completed restore became unknowable from any screen.
// `last_recent` is the server's verdict, using the one window in internal/backup.
if (st.last && (sawRunning || st.last_recent)) {
banner.style.display = 'block'; banner.style.display = 'block';
if (st.last.ok) { if (st.last.ok) {
banner.className = 'flash flash-success'; banner.className = 'flash flash-success';