R-351: the restore compares where the backup says the data lived; second press cannot start a second run
gates / gates (push) Successful in 10s

Part 3 (not droppable) and the engine half of Part 2. No version bump yet - one bump and
one bake at the end of the session.

PART 3a - a second press really did start a second run. Established with a test BEFORE any
change: both offboxReconstituteHandler and offboxPlaceHandler answered "...elindult" and
overwrote the first restore's op/stack. Cause: every restore handler gated on
backupMgr.IsRunning() - the CONCURRENCY flag, which the restore goroutine acquires AFTER the
handler returns (offbox_reconstitute.go:180, offbox_restore.go:393). Seven sites. The wizard
had read the correct flag since v0.154.0 and said so in a comment; the handlers never moved.
New Server.restoreOpBlocked() reads BOTH flags - the display flag covers the whole off-box
restore, the concurrency flag is the only one the nightly backup holds - and the refusal now
names the running app and a route.

PART 3b - the page DOES refresh; the defect was the RESULT. backups_shared.html gated the
terminal result on a page-local sawRunning flag, so a restore that finished before the page
was opened, or inside one 3s poll, was shown to nobody. The 2026-08-21 OpenGist restore took
8.666s and no screen ever said it completed - the answer existed only in docker logs.
RestoreOpStatus.LastRecent now carries the server's verdict. The 10-minute window moved to
internal/backup as RestoreResultWindow and internal/web's constant is an alias: one
expression, two surfaces. Also removed the wizard's self-contradiction, which said the state
refreshes automatically AND that you must refresh the page.

PART 2 (engine) - every recovery unit manifest has carried drive and namespace_root since
schema 1, and NO non-test code read either back. The reconstitution opened the manifest and
took only the coherence stamp, then resolved its destination from the live app. A restore
into a different destination succeeded silently under a green message. New
backup/offbox_placement.go: CheckPlacement (pure, total), PlacementMismatchMessage,
recordedPlacementFromScratch. Compared before the safety dump and before the first byte.
A mismatch is NAMED and refused; ackPlacementChange lets the customer proceed deliberately -
a separate field from confirm=1, because one click must not carry two decisions. An UNKNOWN
recording is never a mismatch: refusing on an absence would strand every pre-field unit.
The not-installed refusal (R-253) now names the drive the backup recorded.

RED-PROOFS, each mutation asserted applied and reverted to 0:
  B  both guards removed (count asserted 2) -> the restore WAS seen starting with no drive
     attached: no error, full 3.00s run, wrote into /tmp/mutant-destination
  C  Mismatch forced false -> the silent divergent restore returned
  E  Known() forced true  -> the fabricated empty prefill appeared
  D  Mismatch forced true -> 8 ordinary reconstitute tests broke, proving reachability both ways
Note on D: the existing fixtures write a schema-1 manifest with NO drive, so they are
scenario-E shaped. The matching case is covered in the scenario table, not by them.

Gates 11/11 OK. Suite 28 packages ok. Hungarian verified as hex, no BOM, no mojibake sentinels.

NOT in this commit, still open: Part 2's scenario-A prefill UI, Part 1's deploy-page
visibility line, Part 1's specification document, Part 4's measurement.
This commit is contained in:
2026-08-21 21:04:16 +02:00
parent 2fa1efc5e5
commit 985388c6e9
14 changed files with 796 additions and 30 deletions
@@ -0,0 +1,146 @@
package backup
import (
"fmt"
"io/fs"
"os"
"path/filepath"
"strings"
)
// R-351 — THE RESTORE ALREADY KNOWS WHERE THE APP LIVED. IT JUST NEVER LOOKED.
//
// Every recovery unit's manifest.json carries `drive` and `namespace_root` (recovery_unit.go:48-49),
// written at capture time from the app's own live placement. Measured on demo-hp 2026-08-21:
//
// opengist drive=/mnt/sys_drive nsroot=/mnt/sys_drive/felhom-data
// calibre-web drive=/mnt/felhom-drives/hdd_1 nsroot=/mnt/felhom-drives/hdd_1
//
// Before this file, NO non-test code in the repository read either field back. `grep -rE
// '\.Drive\b|\.NamespaceRoot\b' --include=*.go` returned only the appbackup.NamespaceRoot FUNCTION
// and Tier2Target's unrelated field. The reconstitution opened the manifest
// (offbox_reconstitute.go:235) and took only the coherence stamp from it, then resolved its
// destination from the LIVE app instead. One side wrote the fact; the other never received it —
// the same shape as several defects closed this month.
//
// THE CONSEQUENCE THAT MADE THIS WORTH FIXING: a restore into a destination that differs from the
// one the backup recorded succeeded SILENTLY, under a green message. Nothing compared them.
//
// WHAT THIS IS NOT. It is not a lock. A domain can legitimately change and hardware can move, so a
// difference is NAMED and the customer decides — it is their own previous answer being shown back to
// them, not a rule imposed on them. What must never happen is the difference passing unremarked.
// RecordedPlacement is what a backup says about where an app's data lived when it was captured.
// Empty fields mean the manifest did not record them — an older unit, or one written before the
// field existed. That is an UNKNOWN and is never rendered as a value.
type RecordedPlacement struct {
Drive string // manifest.Drive — the in-guest mount (HDD_PATH), or the system data path
NamespaceRoot string // manifest.NamespaceRoot — the resolved felhom-data namespace root
}
// Known reports whether the backup recorded a destination at all. A unit whose manifest predates the
// field, or could not be read, is NOT known — and "we cannot tell" is a different answer from "they
// match", which is why this is a method rather than a `!= ""` scattered over the callers.
func (p RecordedPlacement) Known() bool { return strings.TrimSpace(p.Drive) != "" }
// PlacementCheck is the verdict of comparing what the backup recorded against where the restore is
// actually about to write.
type PlacementCheck struct {
// Recorded is what the backup said. Zero value when the manifest carried nothing.
Recorded RecordedPlacement
// LiveDrive / LiveNamespaceRoot are where this restore will write, resolved the way the
// reconstitution resolves it today.
LiveDrive string
LiveNamespaceRoot string
// Known mirrors Recorded.Known(), carried on the verdict so a template never has to re-derive it.
Known bool
// Mismatch is true ONLY when the backup recorded a destination AND it differs from the live one.
// An unknown recording is never a mismatch: refusing on an absence would block every pre-field
// unit, and asserting a match we cannot see would be worse.
Mismatch bool
}
// CheckPlacement compares a manifest's recorded placement against the live destination.
//
// Deliberately PURE and total: a nil manifest, an empty manifest and a manifest whose drive is blank
// all produce the same honest "not known, no mismatch" verdict. Paths are compared Cleaned, because
// `/mnt/x` and `/mnt/x/` are the same destination and a trailing slash must not manufacture a
// mismatch the customer then has to dismiss. Comparison is on the DRIVE, not the namespace root: the
// root is derived from the drive (namespaceRoot appends felhom-data only on the system-data
// fallback), so comparing both would report one difference twice.
func CheckPlacement(man *RecoveryManifest, liveDrive, liveNamespaceRoot string) PlacementCheck {
c := PlacementCheck{
LiveDrive: strings.TrimSpace(liveDrive),
LiveNamespaceRoot: strings.TrimSpace(liveNamespaceRoot),
}
if man != nil {
c.Recorded = RecordedPlacement{
Drive: strings.TrimSpace(man.Drive),
NamespaceRoot: strings.TrimSpace(man.NamespaceRoot),
}
}
c.Known = c.Recorded.Known()
if !c.Known || c.LiveDrive == "" {
return c
}
c.Mismatch = filepath.Clean(c.Recorded.Drive) != filepath.Clean(c.LiveDrive)
return c
}
// scratchManifestMaxDepth bounds the walk below. The unit sits a handful of levels under the scratch
// root (the restore mirrors the snapshot's absolute path), and an unbounded walk over a scratch that
// also holds a full userdata tree would stat a customer's entire library to find one small file.
const scratchManifestMaxDepth = 8
// recordedPlacementFromScratch reads the placement out of the unit manifest inside a PREPARED
// restore scratch, without restoring anything.
//
// It exists for the not-installed refusal (scenario B), which fires BEFORE the live namespace root
// can be resolved — so it cannot use the manifest the reconstitution opens later. The scratch is
// already on local disk by then; this is a bounded walk and a file read, never a network call.
//
// Returns the zero value on ANY failure — unreadable, absent, malformed. A refusal that cannot name
// the recorded place must fall back to the plain sentence rather than print an empty path, which
// would read as "the backup says it lived nowhere".
func (m *Manager) recordedPlacementFromScratch(scratch string) RecordedPlacement {
var found RecordedPlacement
root := filepath.Clean(scratch)
rootDepth := strings.Count(root, string(os.PathSeparator))
_ = filepath.WalkDir(root, func(path string, d fs.DirEntry, err error) error {
if err != nil {
return nil // an unreadable subtree is not fatal: keep looking elsewhere
}
if d.IsDir() {
if strings.Count(filepath.Clean(path), string(os.PathSeparator))-rootDepth >= scratchManifestMaxDepth {
return fs.SkipDir
}
return nil
}
if d.Name() != "manifest.json" {
return nil
}
if man := readManifest(path); man != nil && strings.TrimSpace(man.Drive) != "" {
found = RecordedPlacement{
Drive: strings.TrimSpace(man.Drive),
NamespaceRoot: strings.TrimSpace(man.NamespaceRoot),
}
return fs.SkipAll
}
return nil
})
return found
}
// PlacementMismatchMessage is the Hungarian refusal shown when the destination differs from the one
// the backup recorded. It NAMES BOTH VALUES — which is the whole point: "a destination differs" that
// does not say from what leaves the customer with a decision they cannot make.
//
// It is a refusal with a route, not a dead end: the caller re-offers the action with the
// acknowledgement field set, so the customer can proceed deliberately (scenario C).
func PlacementMismatchMessage(stack string, c PlacementCheck) string {
return fmt.Sprintf(
"A(z) %s mentése szerint az adatok korábban itt voltak: %s. Most viszont ide állna vissza: %s. "+
"Ez lehet szándékos — például ha meghajtót cseréltél —, de magától nem folytatjuk. "+
"Ha így jó, erősítsd meg alább, és a visszaállítás az új helyre fut.",
stack, c.Recorded.Drive, c.LiveDrive)
}