R-351: the restore compares where the backup says the data lived; second press cannot start a second run
gates / gates (push) Successful in 10s
gates / gates (push) Successful in 10s
Part 3 (not droppable) and the engine half of Part 2. No version bump yet - one bump and
one bake at the end of the session.
PART 3a - a second press really did start a second run. Established with a test BEFORE any
change: both offboxReconstituteHandler and offboxPlaceHandler answered "...elindult" and
overwrote the first restore's op/stack. Cause: every restore handler gated on
backupMgr.IsRunning() - the CONCURRENCY flag, which the restore goroutine acquires AFTER the
handler returns (offbox_reconstitute.go:180, offbox_restore.go:393). Seven sites. The wizard
had read the correct flag since v0.154.0 and said so in a comment; the handlers never moved.
New Server.restoreOpBlocked() reads BOTH flags - the display flag covers the whole off-box
restore, the concurrency flag is the only one the nightly backup holds - and the refusal now
names the running app and a route.
PART 3b - the page DOES refresh; the defect was the RESULT. backups_shared.html gated the
terminal result on a page-local sawRunning flag, so a restore that finished before the page
was opened, or inside one 3s poll, was shown to nobody. The 2026-08-21 OpenGist restore took
8.666s and no screen ever said it completed - the answer existed only in docker logs.
RestoreOpStatus.LastRecent now carries the server's verdict. The 10-minute window moved to
internal/backup as RestoreResultWindow and internal/web's constant is an alias: one
expression, two surfaces. Also removed the wizard's self-contradiction, which said the state
refreshes automatically AND that you must refresh the page.
PART 2 (engine) - every recovery unit manifest has carried drive and namespace_root since
schema 1, and NO non-test code read either back. The reconstitution opened the manifest and
took only the coherence stamp, then resolved its destination from the live app. A restore
into a different destination succeeded silently under a green message. New
backup/offbox_placement.go: CheckPlacement (pure, total), PlacementMismatchMessage,
recordedPlacementFromScratch. Compared before the safety dump and before the first byte.
A mismatch is NAMED and refused; ackPlacementChange lets the customer proceed deliberately -
a separate field from confirm=1, because one click must not carry two decisions. An UNKNOWN
recording is never a mismatch: refusing on an absence would strand every pre-field unit.
The not-installed refusal (R-253) now names the drive the backup recorded.
RED-PROOFS, each mutation asserted applied and reverted to 0:
B both guards removed (count asserted 2) -> the restore WAS seen starting with no drive
attached: no error, full 3.00s run, wrote into /tmp/mutant-destination
C Mismatch forced false -> the silent divergent restore returned
E Known() forced true -> the fabricated empty prefill appeared
D Mismatch forced true -> 8 ordinary reconstitute tests broke, proving reachability both ways
Note on D: the existing fixtures write a schema-1 manifest with NO drive, so they are
scenario-E shaped. The matching case is covered in the scenario table, not by them.
Gates 11/11 OK. Suite 28 packages ok. Hungarian verified as hex, no BOM, no mojibake sentinels.
NOT in this commit, still open: Part 2's scenario-A prefill UI, Part 1's deploy-page
visibility line, Part 1's specification document, Part 4's measurement.
This commit is contained in:
@@ -70,6 +70,12 @@ type OffsiteReconstituteResult struct {
|
||||
OffsiteRunID string // "" for a pre-v0.148 snapshot — an unverified pair
|
||||
Skewed bool // the snapshot carries no coherence stamp: files and DB may differ in age
|
||||
LooksEmpty bool // R-44 sniff on the dump about to be replayed
|
||||
// Placement (R-351) is what the backup recorded about where this app's data lived, compared
|
||||
// against where this restore actually wrote. Carried on the RESULT and not only on the refusal,
|
||||
// so a restore that proceeded into a different destination says so in its own outcome rather
|
||||
// than reporting a bare success — a warning beside a success is read as a success, so the
|
||||
// difference has to survive into the message.
|
||||
Placement PlacementCheck
|
||||
}
|
||||
|
||||
// fullPlaceCopier returns the FULL-restore file copier (nil seam → rsyncRestoreOverwrite).
|
||||
@@ -166,7 +172,11 @@ func (m *Manager) dumpForSafety(ctx context.Context, db DiscoveredDB, dumpDir st
|
||||
// overwritten to the snapshot's version (extras survive, nothing deleted), then the snapshot's own
|
||||
// DB dump replayed, with a safety dump of the current database taken first. Requires a completed
|
||||
// FULL scratch restore (RestoreOffboxScratch with full=true). Single-flight.
|
||||
func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string) (OffsiteReconstituteResult, error) {
|
||||
// ackPlacementChange (R-351) is the customer's DELIBERATE acknowledgement that the destination
|
||||
// differs from the one the backup recorded. It is a separate act from the restore's own confirm:
|
||||
// folding it into `confirm=1` would mean one click carried two decisions, which is precisely what
|
||||
// R-48 exists to prevent.
|
||||
func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string, ackPlacementChange bool) (OffsiteReconstituteResult, error) {
|
||||
var res OffsiteReconstituteResult
|
||||
if !m.OffboxConfigured() {
|
||||
return res, fmt.Errorf("off-box backup not configured")
|
||||
@@ -203,6 +213,16 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string) (Of
|
||||
// destination is the app's own HDD path, which is a drive the CUSTOMER chooses at deploy
|
||||
// time, and picking it for them is the decision this whole recovery path exists to leave
|
||||
// with them.
|
||||
// R-351: the refusal now NAMES the place the backup recorded, when it can read it. The
|
||||
// prepared scratch already contains the unit, so this is a local file read — no network call,
|
||||
// nothing restored, and it happens on a path that was going to refuse anyway. Telling
|
||||
// somebody to reinstall without telling them where the data belongs is what forced the
|
||||
// 2026-08-21 operator to remember two values the backup already held.
|
||||
if rec := m.recordedPlacementFromScratch(scratch); rec.Known() {
|
||||
return res, fmt.Errorf("a(z) %s nincs telepítve, ezért nincs hová visszaállítani az adatait. "+
|
||||
"A mentése szerint az adatai itt voltak: %s. Telepítsd újra az alkalmazást (Alkalmazások) "+
|
||||
"ugyanerre a helyre, utána ez a visszaállítás működni fog", stack, rec.Drive)
|
||||
}
|
||||
return res, fmt.Errorf("a(z) %s nincs telepítve, ezért nincs hová visszaállítani az adatait — "+
|
||||
"telepítsd újra az alkalmazást (Alkalmazások), utána ez a visszaállítás működni fog", stack)
|
||||
}
|
||||
@@ -232,7 +252,8 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string) (Of
|
||||
return res, fmt.Errorf("a pillanatképben nincs mentési egység — a visszaállítás nem indítható")
|
||||
}
|
||||
scratchDumpDir := filepath.Join(scratchUnit, "db-dumps")
|
||||
if man := readManifest(filepath.Join(scratchUnit, "manifest.json")); man != nil {
|
||||
man := readManifest(filepath.Join(scratchUnit, "manifest.json"))
|
||||
if man != nil {
|
||||
res.OffsiteRunID = man.OffsiteRunID
|
||||
if man.DumpsAt != "" {
|
||||
if t, pErr := time.Parse(time.RFC3339, man.DumpsAt); pErr == nil {
|
||||
@@ -240,6 +261,23 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string) (Of
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// --- WHERE THE BACKUP SAYS THIS DATA LIVED (R-351) ------------------------------------------
|
||||
// The manifest we just opened has carried `drive` and `namespace_root` since schema 1, and until
|
||||
// now nothing read them back. Compared HERE, before the safety dump and before the first byte is
|
||||
// placed, so the refusal costs nothing and leaves the app completely untouched.
|
||||
//
|
||||
// An UNKNOWN recording (a pre-field unit, or one we could not read) is not a mismatch and does
|
||||
// not refuse: blocking on an absence would strand every older backup, and CheckPlacement returns
|
||||
// that case explicitly rather than letting it fall through as "they match".
|
||||
res.Placement = CheckPlacement(man, hdd, liveNs)
|
||||
if res.Placement.Mismatch && !ackPlacementChange {
|
||||
return res, fmt.Errorf("%s", PlacementMismatchMessage(stack, res.Placement))
|
||||
}
|
||||
if res.Placement.Mismatch {
|
||||
m.logger.Printf("[WARN] [offbox] %s: restoring into %s, but the backup recorded %s — the customer acknowledged the change",
|
||||
stack, res.Placement.LiveDrive, res.Placement.Recorded.Drive)
|
||||
}
|
||||
// A pre-v0.148 snapshot carries no stamp: its dump was whatever the 02:30 local run left behind,
|
||||
// so the pair's two halves may be hours or days apart. Surfaced, never blocked — the confirm
|
||||
// dialog says so and the safety dump makes it reversible.
|
||||
|
||||
Reference in New Issue
Block a user