R-351: the restore compares where the backup says the data lived; second press cannot start a second run
gates / gates (push) Successful in 10s

Part 3 (not droppable) and the engine half of Part 2. No version bump yet - one bump and
one bake at the end of the session.

PART 3a - a second press really did start a second run. Established with a test BEFORE any
change: both offboxReconstituteHandler and offboxPlaceHandler answered "...elindult" and
overwrote the first restore's op/stack. Cause: every restore handler gated on
backupMgr.IsRunning() - the CONCURRENCY flag, which the restore goroutine acquires AFTER the
handler returns (offbox_reconstitute.go:180, offbox_restore.go:393). Seven sites. The wizard
had read the correct flag since v0.154.0 and said so in a comment; the handlers never moved.
New Server.restoreOpBlocked() reads BOTH flags - the display flag covers the whole off-box
restore, the concurrency flag is the only one the nightly backup holds - and the refusal now
names the running app and a route.

PART 3b - the page DOES refresh; the defect was the RESULT. backups_shared.html gated the
terminal result on a page-local sawRunning flag, so a restore that finished before the page
was opened, or inside one 3s poll, was shown to nobody. The 2026-08-21 OpenGist restore took
8.666s and no screen ever said it completed - the answer existed only in docker logs.
RestoreOpStatus.LastRecent now carries the server's verdict. The 10-minute window moved to
internal/backup as RestoreResultWindow and internal/web's constant is an alias: one
expression, two surfaces. Also removed the wizard's self-contradiction, which said the state
refreshes automatically AND that you must refresh the page.

PART 2 (engine) - every recovery unit manifest has carried drive and namespace_root since
schema 1, and NO non-test code read either back. The reconstitution opened the manifest and
took only the coherence stamp, then resolved its destination from the live app. A restore
into a different destination succeeded silently under a green message. New
backup/offbox_placement.go: CheckPlacement (pure, total), PlacementMismatchMessage,
recordedPlacementFromScratch. Compared before the safety dump and before the first byte.
A mismatch is NAMED and refused; ackPlacementChange lets the customer proceed deliberately -
a separate field from confirm=1, because one click must not carry two decisions. An UNKNOWN
recording is never a mismatch: refusing on an absence would strand every pre-field unit.
The not-installed refusal (R-253) now names the drive the backup recorded.

RED-PROOFS, each mutation asserted applied and reverted to 0:
  B  both guards removed (count asserted 2) -> the restore WAS seen starting with no drive
     attached: no error, full 3.00s run, wrote into /tmp/mutant-destination
  C  Mismatch forced false -> the silent divergent restore returned
  E  Known() forced true  -> the fabricated empty prefill appeared
  D  Mismatch forced true -> 8 ordinary reconstitute tests broke, proving reachability both ways
Note on D: the existing fixtures write a schema-1 manifest with NO drive, so they are
scenario-E shaped. The matching case is covered in the scenario table, not by them.

Gates 11/11 OK. Suite 28 packages ok. Hungarian verified as hex, no BOM, no mojibake sentinels.

NOT in this commit, still open: Part 2's scenario-A prefill UI, Part 1's deploy-page
visibility line, Part 1's specification document, Part 4's measurement.
This commit is contained in:
2026-08-21 21:04:16 +02:00
parent 2fa1efc5e5
commit 985388c6e9
14 changed files with 796 additions and 30 deletions
@@ -0,0 +1,146 @@
package backup
import (
"fmt"
"io/fs"
"os"
"path/filepath"
"strings"
)
// R-351 — THE RESTORE ALREADY KNOWS WHERE THE APP LIVED. IT JUST NEVER LOOKED.
//
// Every recovery unit's manifest.json carries `drive` and `namespace_root` (recovery_unit.go:48-49),
// written at capture time from the app's own live placement. Measured on demo-hp 2026-08-21:
//
// opengist drive=/mnt/sys_drive nsroot=/mnt/sys_drive/felhom-data
// calibre-web drive=/mnt/felhom-drives/hdd_1 nsroot=/mnt/felhom-drives/hdd_1
//
// Before this file, NO non-test code in the repository read either field back. `grep -rE
// '\.Drive\b|\.NamespaceRoot\b' --include=*.go` returned only the appbackup.NamespaceRoot FUNCTION
// and Tier2Target's unrelated field. The reconstitution opened the manifest
// (offbox_reconstitute.go:235) and took only the coherence stamp from it, then resolved its
// destination from the LIVE app instead. One side wrote the fact; the other never received it —
// the same shape as several defects closed this month.
//
// THE CONSEQUENCE THAT MADE THIS WORTH FIXING: a restore into a destination that differs from the
// one the backup recorded succeeded SILENTLY, under a green message. Nothing compared them.
//
// WHAT THIS IS NOT. It is not a lock. A domain can legitimately change and hardware can move, so a
// difference is NAMED and the customer decides — it is their own previous answer being shown back to
// them, not a rule imposed on them. What must never happen is the difference passing unremarked.
// RecordedPlacement is what a backup says about where an app's data lived when it was captured.
// Empty fields mean the manifest did not record them — an older unit, or one written before the
// field existed. That is an UNKNOWN and is never rendered as a value.
type RecordedPlacement struct {
Drive string // manifest.Drive — the in-guest mount (HDD_PATH), or the system data path
NamespaceRoot string // manifest.NamespaceRoot — the resolved felhom-data namespace root
}
// Known reports whether the backup recorded a destination at all. A unit whose manifest predates the
// field, or could not be read, is NOT known — and "we cannot tell" is a different answer from "they
// match", which is why this is a method rather than a `!= ""` scattered over the callers.
func (p RecordedPlacement) Known() bool { return strings.TrimSpace(p.Drive) != "" }
// PlacementCheck is the verdict of comparing what the backup recorded against where the restore is
// actually about to write.
type PlacementCheck struct {
// Recorded is what the backup said. Zero value when the manifest carried nothing.
Recorded RecordedPlacement
// LiveDrive / LiveNamespaceRoot are where this restore will write, resolved the way the
// reconstitution resolves it today.
LiveDrive string
LiveNamespaceRoot string
// Known mirrors Recorded.Known(), carried on the verdict so a template never has to re-derive it.
Known bool
// Mismatch is true ONLY when the backup recorded a destination AND it differs from the live one.
// An unknown recording is never a mismatch: refusing on an absence would block every pre-field
// unit, and asserting a match we cannot see would be worse.
Mismatch bool
}
// CheckPlacement compares a manifest's recorded placement against the live destination.
//
// Deliberately PURE and total: a nil manifest, an empty manifest and a manifest whose drive is blank
// all produce the same honest "not known, no mismatch" verdict. Paths are compared Cleaned, because
// `/mnt/x` and `/mnt/x/` are the same destination and a trailing slash must not manufacture a
// mismatch the customer then has to dismiss. Comparison is on the DRIVE, not the namespace root: the
// root is derived from the drive (namespaceRoot appends felhom-data only on the system-data
// fallback), so comparing both would report one difference twice.
func CheckPlacement(man *RecoveryManifest, liveDrive, liveNamespaceRoot string) PlacementCheck {
c := PlacementCheck{
LiveDrive: strings.TrimSpace(liveDrive),
LiveNamespaceRoot: strings.TrimSpace(liveNamespaceRoot),
}
if man != nil {
c.Recorded = RecordedPlacement{
Drive: strings.TrimSpace(man.Drive),
NamespaceRoot: strings.TrimSpace(man.NamespaceRoot),
}
}
c.Known = c.Recorded.Known()
if !c.Known || c.LiveDrive == "" {
return c
}
c.Mismatch = filepath.Clean(c.Recorded.Drive) != filepath.Clean(c.LiveDrive)
return c
}
// scratchManifestMaxDepth bounds the walk below. The unit sits a handful of levels under the scratch
// root (the restore mirrors the snapshot's absolute path), and an unbounded walk over a scratch that
// also holds a full userdata tree would stat a customer's entire library to find one small file.
const scratchManifestMaxDepth = 8
// recordedPlacementFromScratch reads the placement out of the unit manifest inside a PREPARED
// restore scratch, without restoring anything.
//
// It exists for the not-installed refusal (scenario B), which fires BEFORE the live namespace root
// can be resolved — so it cannot use the manifest the reconstitution opens later. The scratch is
// already on local disk by then; this is a bounded walk and a file read, never a network call.
//
// Returns the zero value on ANY failure — unreadable, absent, malformed. A refusal that cannot name
// the recorded place must fall back to the plain sentence rather than print an empty path, which
// would read as "the backup says it lived nowhere".
func (m *Manager) recordedPlacementFromScratch(scratch string) RecordedPlacement {
var found RecordedPlacement
root := filepath.Clean(scratch)
rootDepth := strings.Count(root, string(os.PathSeparator))
_ = filepath.WalkDir(root, func(path string, d fs.DirEntry, err error) error {
if err != nil {
return nil // an unreadable subtree is not fatal: keep looking elsewhere
}
if d.IsDir() {
if strings.Count(filepath.Clean(path), string(os.PathSeparator))-rootDepth >= scratchManifestMaxDepth {
return fs.SkipDir
}
return nil
}
if d.Name() != "manifest.json" {
return nil
}
if man := readManifest(path); man != nil && strings.TrimSpace(man.Drive) != "" {
found = RecordedPlacement{
Drive: strings.TrimSpace(man.Drive),
NamespaceRoot: strings.TrimSpace(man.NamespaceRoot),
}
return fs.SkipAll
}
return nil
})
return found
}
// PlacementMismatchMessage is the Hungarian refusal shown when the destination differs from the one
// the backup recorded. It NAMES BOTH VALUES — which is the whole point: "a destination differs" that
// does not say from what leaves the customer with a decision they cannot make.
//
// It is a refusal with a route, not a dead end: the caller re-offers the action with the
// acknowledgement field set, so the customer can proceed deliberately (scenario C).
func PlacementMismatchMessage(stack string, c PlacementCheck) string {
return fmt.Sprintf(
"A(z) %s mentése szerint az adatok korábban itt voltak: %s. Most viszont ide állna vissza: %s. "+
"Ez lehet szándékos — például ha meghajtót cseréltél —, de magától nem folytatjuk. "+
"Ha így jó, erősítsd meg alább, és a visszaállítás az új helyre fut.",
stack, c.Recorded.Drive, c.LiveDrive)
}
@@ -0,0 +1,167 @@
package backup
import (
"context"
"io/fs"
"os"
"path/filepath"
"strings"
"testing"
)
// scratchManifestPath finds the unit manifest the fixture wrote into the prepared scratch, so a test
// can rewrite it. Fails loudly rather than returning "": a silently-missing manifest would make the
// tests below pass for the wrong reason.
func scratchManifestPath(t *testing.T, m *Manager, stack string) string {
t.Helper()
scratch, _, err := m.offboxRestoreScratchDir(stack)
if err != nil {
t.Fatalf("fixture: scratch dir: %v", err)
}
var found string
_ = filepath.WalkDir(scratch, func(p string, d fs.DirEntry, err error) error {
if err == nil && !d.IsDir() && d.Name() == "manifest.json" {
found = p
return fs.SkipAll
}
return nil
})
if found == "" {
t.Fatal("fixture: no unit manifest in the prepared scratch — the test would prove nothing")
}
return found
}
// R-351 SCENARIO B — THE APP IS NOT INSTALLED AND THE BACKUP NAMES WHERE IT LIVED.
//
// The refusal already existed (R-253) and correctly stopped the restore. What it did NOT do was say
// where the data belonged, so the person on 2026-08-21 had to remember the drive and the address
// themselves — both of which the backup was holding the whole time.
//
// WRONG OUTCOME THIS PINS: "it starts, and writes somewhere else." The assertions below are that the
// restore is refused AND that the app was never touched — not merely that an error came back.
func TestReconstitute_NotInstalled_RefusalNamesTheRecordedDrive(t *testing.T) {
const recordedDrive = "/mnt/felhom-drives/hdd_1"
m, prov, _ := reconFixture(t, "20260719T060000Z", "2026-07-19T06:00:00Z", pgDump(1))
// The backup records a drive — the one this box no longer has attached.
manPath := scratchManifestPath(t, m, "immich")
if err := writeManifest(manPath, &RecoveryManifest{
SchemaVersion: 2, AppName: "immich",
Drive: recordedDrive, NamespaceRoot: recordedDrive,
}); err != nil {
t.Fatal(err)
}
// The app is not installed: no live HDD path. This is the rebuilt-machine shape.
prov.hdd = map[string]string{}
if got := prov.GetStackHDDPath("immich"); got != "" {
t.Fatalf("fixture: the app must look uninstalled, got hdd=%q", got)
}
callsBefore := len(prov.calls)
_, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err == nil {
t.Fatal("a restore with no destination must be REFUSED, not started")
}
if !strings.Contains(err.Error(), recordedDrive) {
t.Errorf("the refusal must name the drive the backup recorded (%s); got: %v", recordedDrive, err)
}
// The route, not just the reason. „Alkalmaz" is the ASCII stem of „Alkalmazások" — matching the
// stem keeps accented bytes out of the comparison (strict rule 7).
if !strings.Contains(err.Error(), "Alkalmaz") {
t.Errorf("the refusal must name a route the person can take; got: %v", err)
}
// THE OBSERVABLE THAT MATTERS: nothing was stopped, placed or started. A refusal that still
// touched the stack would be the defect wearing an error message.
if len(prov.calls) != callsBefore {
t.Errorf("a refused restore must not touch the app; provider calls: %v", prov.calls[callsBefore:])
}
}
// R-351 SCENARIO C — THE DESTINATION DIFFERS FROM THE ONE THE BACKUP RECORDED.
//
// Not blocked outright (a drive can legitimately change) and not silently accepted (which is how a
// restore lands in the wrong place under a green message). Named, and then the customer's own
// deliberate acknowledgement carries it.
func TestReconstitute_MismatchedDestination_RefusesThenProceedsOnAcknowledgement(t *testing.T) {
const recordedDrive = "/mnt/felhom-drives/hdd_1"
m, prov, _ := reconFixture(t, "20260719T060000Z", "2026-07-19T06:00:00Z", pgDump(1))
manPath := scratchManifestPath(t, m, "immich")
if err := writeManifest(manPath, &RecoveryManifest{
SchemaVersion: 2, AppName: "immich",
Drive: recordedDrive, NamespaceRoot: recordedDrive,
}); err != nil {
t.Fatal(err)
}
// The app IS installed — on a different drive from the recorded one (the fixture's temp dir).
live := prov.GetStackHDDPath("immich")
if live == "" || live == recordedDrive {
t.Fatalf("fixture: the live drive must exist and differ from the recorded one; got %q", live)
}
callsBefore := len(prov.calls)
_, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err == nil {
t.Fatal("a divergent destination must be NAMED and refused, never silently accepted")
}
for _, must := range []string{recordedDrive, live} {
if !strings.Contains(err.Error(), must) {
t.Errorf("the refusal must name %q so the customer can decide; got: %v", must, err)
}
}
if len(prov.calls) != callsBefore {
t.Errorf("the refusal must not touch the app; provider calls: %v", prov.calls[callsBefore:])
}
// ...and the customer may go ahead deliberately. Blocking outright is the other wrong outcome.
res, err := m.ReconstituteFromOffsite(context.Background(), "immich", true)
if err != nil {
t.Fatalf("an acknowledged placement change must proceed: %v", err)
}
if !res.Placement.Mismatch {
t.Error("the outcome must still record that the destination differed — a success that forgets the difference reads as a match")
}
if res.Placement.Recorded.Drive != recordedDrive {
t.Errorf("the outcome must carry the recorded drive, got %q", res.Placement.Recorded.Drive)
}
}
// recordedPlacementFromScratch must be TOTAL: an absent scratch, an unreadable one and a manifest
// with no drive all yield the zero value, never a fabricated empty path.
func TestRecordedPlacementFromScratch_UnknownStaysUnknown(t *testing.T) {
m := &Manager{}
if got := m.recordedPlacementFromScratch(filepath.Join(t.TempDir(), "does-not-exist")); got.Known() {
t.Errorf("an absent scratch must not yield a recorded placement, got %+v", got)
}
empty := t.TempDir()
if got := m.recordedPlacementFromScratch(empty); got.Known() {
t.Errorf("an empty scratch must not yield a recorded placement, got %+v", got)
}
noDrive := t.TempDir()
if err := writeManifest(filepath.Join(noDrive, "manifest.json"),
&RecoveryManifest{SchemaVersion: 1, AppName: "immich"}); err != nil {
t.Fatal(err)
}
if got := m.recordedPlacementFromScratch(noDrive); got.Known() {
t.Errorf("a manifest with no drive is an UNKNOWN, not a value; got %+v", got)
}
withDrive := t.TempDir()
deep := filepath.Join(withDrive, "a", "b", "c", "backups", "primary", "immich")
if err := os.MkdirAll(deep, 0o755); err != nil {
t.Fatal(err)
}
if err := writeManifest(filepath.Join(deep, "manifest.json"),
&RecoveryManifest{SchemaVersion: 2, AppName: "immich", Drive: "/mnt/sys_drive"}); err != nil {
t.Fatal(err)
}
got := m.recordedPlacementFromScratch(withDrive)
if !got.Known() || got.Drive != "/mnt/sys_drive" {
t.Errorf("a nested unit manifest must be found, got %+v", got)
}
}
@@ -0,0 +1,135 @@
package backup
import (
"strings"
"testing"
)
// R-351 — THE SCENARIO TABLE from the task, as a truth table.
//
// Each row states the WRONG outcome it exists to prevent, because a row whose expectation is only
// "want X" tells the next reader nothing about why the value matters.
//
// The live values in these fixtures are the ones measured on demo-hp 2026-08-21, not invented ones:
//
// calibre-web (declares a data path) drive=/mnt/felhom-drives/hdd_1
// opengist (declares NONE, 40-of-53) drive=/mnt/sys_drive
func TestCheckPlacement_ScenarioTable(t *testing.T) {
const dataDrive = "/mnt/felhom-drives/hdd_1"
const sysDrive = "/mnt/sys_drive"
for _, tc := range []struct {
name string
man *RecoveryManifest
liveDrive string
wantKnown bool
wantMismatch bool
wrongOutcome string
}{
{
name: "A/D - recorded drive is where we are restoring",
man: &RecoveryManifest{Drive: dataDrive, NamespaceRoot: dataDrive},
liveDrive: dataDrive,
wantKnown: true,
wantMismatch: false,
wrongOutcome: "a new question or obstacle on the ordinary path",
},
{
name: "C - the destination differs from what the backup recorded",
man: &RecoveryManifest{Drive: dataDrive, NamespaceRoot: dataDrive},
liveDrive: "/mnt/felhom-drives/hdd_2",
wantKnown: true,
wantMismatch: true,
wrongOutcome: "silently accepted - a restore into the wrong place under a green message",
},
{
name: "E - the backup records no drive (older unit, missing field)",
man: &RecoveryManifest{Drive: "", NamespaceRoot: ""},
liveDrive: dataDrive,
wantKnown: false,
wantMismatch: false,
wrongOutcome: "an empty prefill presented as if it were the recorded value",
},
{
name: "E - the manifest could not be read at all",
man: nil,
liveDrive: dataDrive,
wantKnown: false,
wantMismatch: false,
wrongOutcome: "a nil manifest treated as a match, or as a refusal that strands every old backup",
},
{
name: "no declared data path - the app has no field, but the record is still the truth",
man: &RecoveryManifest{Drive: sysDrive, NamespaceRoot: sysDrive + "/felhom-data"},
liveDrive: sysDrive,
wantKnown: true,
wantMismatch: false,
wrongOutcome: "falling into the unknown case just because the app has no storage field",
},
{
name: "no declared data path - and it moved to a real drive",
man: &RecoveryManifest{Drive: sysDrive, NamespaceRoot: sysDrive + "/felhom-data"},
liveDrive: dataDrive,
wantKnown: true,
wantMismatch: true,
wrongOutcome: "the 40-of-53 class silently exempted from the mismatch check",
},
{
name: "a trailing slash is the same destination",
man: &RecoveryManifest{Drive: dataDrive + "/", NamespaceRoot: dataDrive},
liveDrive: dataDrive,
wantKnown: true,
wantMismatch: false,
wrongOutcome: "a manufactured mismatch the customer has to dismiss for no reason",
},
{
name: "live destination unresolvable",
man: &RecoveryManifest{Drive: dataDrive, NamespaceRoot: dataDrive},
liveDrive: "",
wantKnown: true,
wantMismatch: false,
wrongOutcome: "comparing against nothing and calling it a difference; the not-installed refusal owns this case",
},
} {
t.Run(tc.name, func(t *testing.T) {
c := CheckPlacement(tc.man, tc.liveDrive, tc.liveDrive)
if c.Known != tc.wantKnown {
t.Errorf("Known = %v, want %v — wrong outcome guarded: %s", c.Known, tc.wantKnown, tc.wrongOutcome)
}
if c.Mismatch != tc.wantMismatch {
t.Errorf("Mismatch = %v, want %v — wrong outcome guarded: %s", c.Mismatch, tc.wantMismatch, tc.wrongOutcome)
}
})
}
}
// The refusal must NAME BOTH VALUES. "The destination differs" without saying from what leaves the
// customer holding a decision they have no way to make — which is the same defect in a politer form.
func TestPlacementMismatchMessage_NamesBothPlaces(t *testing.T) {
c := CheckPlacement(
&RecoveryManifest{Drive: "/mnt/felhom-drives/hdd_1"},
"/mnt/sys_drive", "/mnt/sys_drive/felhom-data")
if !c.Mismatch {
t.Fatal("fixture: these differ, or the message under test is never reached")
}
msg := PlacementMismatchMessage("opengist", c)
for _, must := range []string{"opengist", "/mnt/felhom-drives/hdd_1", "/mnt/sys_drive"} {
if !strings.Contains(msg, must) {
t.Errorf("the refusal must name %q; got: %s", must, msg)
}
}
}
// An UNKNOWN recording must never be rendered as a value. Scenario E's wrong outcome is precisely an
// empty string shown where a recorded path belongs, which reads as "the backup says it lived
// nowhere" — a fabricated fact.
func TestRecordedPlacement_UnknownIsNotAValue(t *testing.T) {
for _, drive := range []string{"", " ", "\t"} {
if (RecordedPlacement{Drive: drive}).Known() {
t.Errorf("a blank drive (%q) must not count as a recorded value", drive)
}
}
if !(RecordedPlacement{Drive: "/mnt/sys_drive"}).Known() {
t.Error("a real recorded drive must count as known, or the whole check is inert")
}
}
@@ -70,6 +70,12 @@ type OffsiteReconstituteResult struct {
OffsiteRunID string // "" for a pre-v0.148 snapshot — an unverified pair
Skewed bool // the snapshot carries no coherence stamp: files and DB may differ in age
LooksEmpty bool // R-44 sniff on the dump about to be replayed
// Placement (R-351) is what the backup recorded about where this app's data lived, compared
// against where this restore actually wrote. Carried on the RESULT and not only on the refusal,
// so a restore that proceeded into a different destination says so in its own outcome rather
// than reporting a bare success — a warning beside a success is read as a success, so the
// difference has to survive into the message.
Placement PlacementCheck
}
// fullPlaceCopier returns the FULL-restore file copier (nil seam → rsyncRestoreOverwrite).
@@ -166,7 +172,11 @@ func (m *Manager) dumpForSafety(ctx context.Context, db DiscoveredDB, dumpDir st
// overwritten to the snapshot's version (extras survive, nothing deleted), then the snapshot's own
// DB dump replayed, with a safety dump of the current database taken first. Requires a completed
// FULL scratch restore (RestoreOffboxScratch with full=true). Single-flight.
func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string) (OffsiteReconstituteResult, error) {
// ackPlacementChange (R-351) is the customer's DELIBERATE acknowledgement that the destination
// differs from the one the backup recorded. It is a separate act from the restore's own confirm:
// folding it into `confirm=1` would mean one click carried two decisions, which is precisely what
// R-48 exists to prevent.
func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string, ackPlacementChange bool) (OffsiteReconstituteResult, error) {
var res OffsiteReconstituteResult
if !m.OffboxConfigured() {
return res, fmt.Errorf("off-box backup not configured")
@@ -203,6 +213,16 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string) (Of
// destination is the app's own HDD path, which is a drive the CUSTOMER chooses at deploy
// time, and picking it for them is the decision this whole recovery path exists to leave
// with them.
// R-351: the refusal now NAMES the place the backup recorded, when it can read it. The
// prepared scratch already contains the unit, so this is a local file read — no network call,
// nothing restored, and it happens on a path that was going to refuse anyway. Telling
// somebody to reinstall without telling them where the data belongs is what forced the
// 2026-08-21 operator to remember two values the backup already held.
if rec := m.recordedPlacementFromScratch(scratch); rec.Known() {
return res, fmt.Errorf("a(z) %s nincs telepítve, ezért nincs hová visszaállítani az adatait. "+
"A mentése szerint az adatai itt voltak: %s. Telepítsd újra az alkalmazást (Alkalmazások) "+
"ugyanerre a helyre, utána ez a visszaállítás működni fog", stack, rec.Drive)
}
return res, fmt.Errorf("a(z) %s nincs telepítve, ezért nincs hová visszaállítani az adatait — "+
"telepítsd újra az alkalmazást (Alkalmazások), utána ez a visszaállítás működni fog", stack)
}
@@ -232,7 +252,8 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string) (Of
return res, fmt.Errorf("a pillanatképben nincs mentési egység — a visszaállítás nem indítható")
}
scratchDumpDir := filepath.Join(scratchUnit, "db-dumps")
if man := readManifest(filepath.Join(scratchUnit, "manifest.json")); man != nil {
man := readManifest(filepath.Join(scratchUnit, "manifest.json"))
if man != nil {
res.OffsiteRunID = man.OffsiteRunID
if man.DumpsAt != "" {
if t, pErr := time.Parse(time.RFC3339, man.DumpsAt); pErr == nil {
@@ -240,6 +261,23 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string) (Of
}
}
}
// --- WHERE THE BACKUP SAYS THIS DATA LIVED (R-351) ------------------------------------------
// The manifest we just opened has carried `drive` and `namespace_root` since schema 1, and until
// now nothing read them back. Compared HERE, before the safety dump and before the first byte is
// placed, so the refusal costs nothing and leaves the app completely untouched.
//
// An UNKNOWN recording (a pre-field unit, or one we could not read) is not a mismatch and does
// not refuse: blocking on an absence would strand every older backup, and CheckPlacement returns
// that case explicitly rather than letting it fall through as "they match".
res.Placement = CheckPlacement(man, hdd, liveNs)
if res.Placement.Mismatch && !ackPlacementChange {
return res, fmt.Errorf("%s", PlacementMismatchMessage(stack, res.Placement))
}
if res.Placement.Mismatch {
m.logger.Printf("[WARN] [offbox] %s: restoring into %s, but the backup recorded %s — the customer acknowledged the change",
stack, res.Placement.LiveDrive, res.Placement.Recorded.Drive)
}
// A pre-v0.148 snapshot carries no stamp: its dump was whatever the 02:30 local run left behind,
// so the pair's two halves may be hours or days apart. Surfaced, never blocked — the confirm
// dialog says so and the safety dump makes it reversible.
@@ -174,7 +174,7 @@ func reconFixture(t *testing.T, runID, dumpsAt string, dumpBody string) (*Manage
func TestReconstituteReplaysDBAndOrdersOperations(t *testing.T) {
m, prov, imported := reconFixture(t, "20260719T060000Z", "2026-07-19T06:00:00Z", pgDump(1))
res, err := m.ReconstituteFromOffsite(context.Background(), "immich")
res, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err != nil {
t.Fatalf("reconstitute: %v", err)
}
@@ -220,7 +220,7 @@ func TestReconstituteRefusesWhenSafetyDumpFails(t *testing.T) {
var copied bool
m.SetOffboxFullPlaceCopier(func(_, _ string) (int, error) { copied = true; return 1, nil })
_, err := m.ReconstituteFromOffsite(context.Background(), "immich")
_, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err == nil {
t.Fatal("expected a refusal when the safety dump cannot be taken")
}
@@ -246,7 +246,7 @@ func TestReconstituteNoDBAppMakesNoDumpOrImportCalls(t *testing.T) {
return DumpResult{DB: d}
})
res, err := m.ReconstituteFromOffsite(context.Background(), "immich")
res, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err != nil {
t.Fatalf("reconstitute: %v", err)
}
@@ -267,7 +267,7 @@ func TestReconstituteNoDBAppMakesNoDumpOrImportCalls(t *testing.T) {
func TestReconstituteSurfacesLegacySkewedPair(t *testing.T) {
m, _, imported := reconFixture(t, "", "", pgDump(1))
res, err := m.ReconstituteFromOffsite(context.Background(), "immich")
res, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err != nil {
t.Fatalf("a legacy pair must still be restorable, got refusal: %v", err)
}
@@ -286,7 +286,7 @@ func TestReconstituteFlagsCustomerEmptyDump(t *testing.T) {
// A valid postgres dump whose accounts table has NO rows — the 2026-07-19 shape exactly.
m, _, _ := reconFixture(t, "run1", "2026-07-19T06:00:00Z", pgDump(0))
res, err := m.ReconstituteFromOffsite(context.Background(), "immich")
res, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err != nil {
t.Fatalf("the sniff must never block a restore: %v", err)
}
+26
View File
@@ -18,6 +18,15 @@ type RestoreOpResult struct {
FinishedAt time.Time `json:"finished_at"`
}
// RestoreResultWindow bounds how long a finished restore still counts as "what just happened".
//
// ONE EXPRESSION, TWO SURFACES (R-351). The wizard's phase strip and the list page's banner both
// need this bound, and until now only the wizard had one — it lived in internal/web as an
// unexported constant. A second copy in the JS would be exactly the "two copies that already
// differed" shape this repo keeps paying for, so the window is defined HERE, beside the status it
// bounds, and both surfaces read it from the payload.
const RestoreResultWindow = 10 * time.Minute
// RestoreOpStatus is the shape served at GET /api/backup/restore-status.
type RestoreOpStatus struct {
Running bool `json:"running"`
@@ -25,6 +34,15 @@ type RestoreOpStatus struct {
Stack string `json:"stack,omitempty"`
StartedAt time.Time `json:"started_at,omitempty"`
Last *RestoreOpResult `json:"last,omitempty"`
// LastRecent reports whether Last finished recently enough to still be worth showing to someone
// who was NOT watching when it happened.
//
// R-351, the measured defect: the banner's JS gated the terminal result on a page-local
// `sawRunning` flag, so a restore that finished before the page was opened — or in under one
// poll interval — was shown to nobody. The 2026-08-21 OpenGist restore finished in 8.7s and the
// operator could not tell from any screen whether it had completed; the answer existed only in
// a container log. A result nobody can see is the same defect class as no result at all.
LastRecent bool `json:"last_recent,omitempty"`
}
// BeginRestoreOp marks a restore op in flight (called by the handler just before launching the
@@ -66,6 +84,14 @@ func (m *Manager) RestoreStatus() RestoreOpStatus {
if m.opLast != nil {
cp := *m.opLast
st.Last = &cp
// Recency is decided here, on the clock the result was stamped with, so neither surface has
// to hold its own copy of the window. A zero FinishedAt is never recent — an unstamped
// result must not be drawn as "just now".
if !cp.FinishedAt.IsZero() {
if d := time.Since(cp.FinishedAt); d >= 0 && d < RestoreResultWindow {
st.LastRecent = true
}
}
}
return st
}
@@ -0,0 +1,84 @@
package backup
import (
"testing"
"time"
)
// R-351 — A FINISHED RESTORE MUST BE VISIBLE TO SOMEONE WHO WAS NOT WATCHING.
//
// THE MEASURED CASE (demo-hp, 2026-08-21). An off-box reconstitution refused at 16:37:14, the person
// reinstalled and ran the local unit restore, and it completed at 16:39:25 — in 8.666s. No screen
// ever said so. The banner's JS gated its terminal result on a page-local `sawRunning` flag, so the
// result was rendered only for a browser that happened to be open and polling across the transition.
// Land on the page a moment later and the banner stayed hidden: "completed" and "never ran" looked
// identical. The answer existed only in `docker logs felhom-controller`, which a customer cannot
// reach.
//
// This pins the CONSEQUENCE — the status carries a recency verdict a late arrival can act on — not
// the mechanism. Mutating LastRecent to stay false must fail this test.
func TestRestoreStatus_LastRecent(t *testing.T) {
for _, tc := range []struct {
name string
finishedAt time.Time
wantRecent bool
why string
}{
{
name: "just finished",
finishedAt: time.Now().Add(-9 * time.Second),
wantRecent: true,
why: "the 8.7s OpenGist restore: finished before a poll could see it running",
},
{
name: "inside the window",
finishedAt: time.Now().Add(-RestoreResultWindow + time.Minute),
wantRecent: true,
why: "still answers \"what just happened\"",
},
{
name: "outside the window",
finishedAt: time.Now().Add(-RestoreResultWindow - time.Minute),
wantRecent: false,
why: "landing here later must not claim a restore just finished",
},
{
name: "unstamped result",
finishedAt: time.Time{},
wantRecent: false,
why: "a zero timestamp is an UNKNOWN and must never be drawn as \"just now\" (presence is not success)",
},
} {
t.Run(tc.name, func(t *testing.T) {
m := &Manager{}
m.opLast = &RestoreOpResult{
Op: "restore", Stack: "opengist", OK: true,
Message: "Restore completed", FinishedAt: tc.finishedAt,
}
st := m.RestoreStatus()
if st.Last == nil {
t.Fatal("fixture: the status must carry the terminal result")
}
if st.LastRecent != tc.wantRecent {
t.Errorf("LastRecent = %v, want %v — %s", st.LastRecent, tc.wantRecent, tc.why)
}
})
}
}
// A RUNNING op must not be reported as a recent RESULT — the banner shows one or the other, and
// conflating them would put a success message over an operation that is still writing.
func TestRestoreStatus_RunningIsNotAResult(t *testing.T) {
m := &Manager{}
m.BeginRestoreOp("offbox-reconstitute", "opengist")
st := m.RestoreStatus()
if !st.Running {
t.Fatal("fixture: the op must be reported running")
}
if st.Last != nil {
t.Errorf("a fresh op has no terminal result yet, got %+v", st.Last)
}
if st.LastRecent {
t.Error("a running op must never carry a recent-result verdict")
}
}
@@ -71,7 +71,7 @@ func TestReconstituteReplaysWithOnlyTheDBServiceUp(t *testing.T) {
return nil
}
res, err := m.ReconstituteFromOffsite(context.Background(), "immich")
res, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err != nil {
t.Fatalf("reconstitute: %v", err)
}
@@ -109,7 +109,7 @@ func TestReconstituteNoDBAppNeverStartsServicesOnly(t *testing.T) {
prov.composePath = writeLiveCompose(t, noDBCompose)
m.discoverDBs = func(context.Context) ([]DiscoveredDB, error) { return nil, nil }
res, err := m.ReconstituteFromOffsite(context.Background(), "immich")
res, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err != nil {
t.Fatalf("a no-DB app must restore unchanged, got: %v", err)
}
@@ -141,7 +141,7 @@ func TestReconstituteRefusesWhenNoDBServiceIdentifiable(t *testing.T) {
var copied bool
m.SetOffboxFullPlaceCopier(func(_, _ string) (int, error) { copied = true; return 1, nil })
_, err := m.ReconstituteFromOffsite(context.Background(), "immich")
_, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err == nil {
t.Fatal("expected a refusal: a dump exists but no database service can be started for it")
}
@@ -270,7 +270,7 @@ func TestReconstituteReplayFailureStillBringsTheStackUp(t *testing.T) {
return context.DeadlineExceeded
}
res, err := m.ReconstituteFromOffsite(context.Background(), "immich")
res, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
if err == nil {
t.Fatal("a failed replay must be surfaced, not swallowed")
}
@@ -292,7 +292,7 @@ func TestReconstituteDBOnlyStartFailureStillBringsTheStackUp(t *testing.T) {
m, prov, imported := reconFixture(t, "run1", "2026-07-19T06:00:00Z", pgDump(1))
prov.startSvcErr = context.DeadlineExceeded
if _, err := m.ReconstituteFromOffsite(context.Background(), "immich"); err == nil {
if _, err := m.ReconstituteFromOffsite(context.Background(), "immich", false); err == nil {
t.Fatal("a failed DB-only start must be surfaced")
}
if !prov.fullStarted {