R-353/R-357/R-358/R-360: the restore tells the truth (v0.226.0)
gates / gates (push) Successful in 11s
gates / gates (push) Successful in 11s
Four defects on the restore surface, all proven on demo-hp during the 2026-08-21 backup-truth drill, all still in shipped code. They share one acceptance idea: a restore surface must state what it actually did, and must refuse what it cannot do. VERSION NOTE. The task specifying this targeted v0.224.0 against baselinef8c9390. Both were consumed earlier the same day by R-330 (0.224.0) and R-331 (0.225.0). Drift re-confirmed against live Gitea before the first edit, operator authorised proceeding, every symbol the spec named re-verified present at the real baselinee5eee50. R-353 -- a restore that gave back nothing still said it worked. RestoreFromRecoveryUnit returned only error, so the surface printed "<app> visszaallitva (<snapshot>)." -- equally true of a run that returned an entire dataset and one that returned nothing. The count already existed and was discarded one line deep: restoreDockerVolumesFrom always returned it, the wrapper threw it away. Now (UnitRestoreResult, error), carrying replayed counts AND what the manifest LISTED, because zero-replayed has two causes that are opposite news. Three cases, three sentences, and EVERY one is a claim about the BACKUP, never about the app -- this path has no SafetyDump discriminator, and 07-backup-architecture 6.3 records that an absent dump says nothing about the app (R-361 destroyed canonical .sql files for four months). R-357 -- the destructive restore had no free-space gate. offbox_reconstitute.go contained ZERO references to offboxFree; all three existing gates guard non-destructive paths. The gate now sits before mapOffsiteRestorePaths, writeSafetyDump and StopStack, so a refusal costs nothing. Position IS the fix, which is why the test asserts StopStack was never called. No headroom multiplier (matches PlaceOffsiteRestore; the x1.1 elsewhere predicts a download). Fail closed on either probe <= 0 -- otherwise `free < need` with need==0 is FALSE and an unmeasurable scratch sails through: a gate present and inert. R-358 -- a failed download was offered as a good one. The gate answered "the directory exists and is non-empty", which is exactly what a part-way restic run leaves. Now a completion marker written 0600 atomically AFTER restic returns nil, with any stale one cleared BEFORE it starts; both orders pinned by an AST test because resticStep is not a seam. Both handlers refuse server-side: the wizard flags control a button, and a hidden button is not a guard. SCENARIO F ANSWERED, and worse than the question assumed: a unit-only scratch IS reachable through the real flow, by the most ordinary route. "Ellenorzo visszaallitas" (mode=unit, advertised non-destructive) writes the SAME directory -- offboxRestoreScratchDir ignores `full` and --include limits what restic extracts, never where -- so a customer who ran the SAFE restore was then offered the destructive one over a unit-only copy. Filed R-396; the marker closes it. R-360 -- the delete refused only while a BACKUP ran. IsRunning() is FALSE for the whole of a verification restore; the five sibling handlers all use restoreOpBlocked(). Its doc comment claimed it already did this, which is why nobody looked -- corrected in place. Red-proofs, each printing the pre-fix behaviour, in CHANGELOG and REPORT. The first R-357 red-proof exposed a hollow test OF MY OWN and it is recorded rather than quietly fixed: the fixture refused earlier at the placement stat pre-pass, so `stops == 0` passed against the pre-fix code. Fixture corrected, assertions reordered so a removed gate reports the outage rather than "no error returned". Green gate clean: 28 packages, rc 0. All 12 controller gates OK.
This commit is contained in:
@@ -144,6 +144,16 @@ type Manager struct {
|
||||
// Windows `go test` host has no `df`). Nil → the real diskFreeBytes (df --output=avail).
|
||||
offboxFreeFn func(path string) int64
|
||||
|
||||
// offboxLatestSnapFn (R-357) overrides the restic snapshot lookup, and it exists for one reason:
|
||||
// without it, ReconstituteFromOffsite's new headroom gate cannot be tested at the level that
|
||||
// matters. Scenario D's claim is not "the error string is right" — it is "the app was NEVER
|
||||
// STOPPED", and reaching the gate at all requires getting past offboxLatestSnapshot, which shells
|
||||
// out to restic. A test that shells to restic is not a unit test, and a gate proven only by
|
||||
// reading the code is exactly the class of assurance this project has been burned by.
|
||||
//
|
||||
// INIT/TEST ONLY. Nil in production, where offboxLatestSnapshot runs unchanged.
|
||||
offboxLatestSnapFn func(ctx context.Context, stack string) (string, []string, error)
|
||||
|
||||
// F17 restore seams — overridable in tests so the .sql re-import orchestration can be unit-tested
|
||||
// without Docker. Default to the real DiscoverDatabases / ImportDump (lazy-init in reimportDBDumps).
|
||||
discoverDBs func(ctx context.Context) ([]DiscoveredDB, error)
|
||||
|
||||
@@ -547,6 +547,44 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string, ack
|
||||
}
|
||||
liveNs := m.namespaceRoot(hdd)
|
||||
|
||||
// --- R-357: FREE SPACE, BEFORE ANYTHING IS TOUCHED ------------------------------------------
|
||||
//
|
||||
// This file contained ZERO references to offboxFree until now. The three headroom gates that
|
||||
// existed all guarded NON-destructive paths (offbox_restore.go: the download sizer, the prepare
|
||||
// gate, and PlaceOffsiteRestore's missing-only merge). The one path that stops the customer's app
|
||||
// and overwrites their live data had none.
|
||||
//
|
||||
// Measured on demo-hp 2026-08-21: it stopped the app, ran out of disk part-way, left 2 of 5 planted
|
||||
// items in place and restarted the app — a half-restored dataset presented as a completed restore.
|
||||
//
|
||||
// POSITION IS THE WHOLE FIX. This sits before mapOffsiteRestorePaths, before writeSafetyDump and
|
||||
// well before StopStack, so a refusal costs the customer nothing at all — the app never goes down.
|
||||
// A gate after StopStack would turn a refusal into an outage, which is the shape it exists to
|
||||
// prevent. Scenario D asserts the non-effect (StopStack call count 0), not the error string.
|
||||
//
|
||||
// NO HEADROOM MULTIPLIER, deliberately, and stated so the next reader does not "fix" it:
|
||||
// OffboxRestorePrepareFull uses ×1.1 because it is sizing a DOWNLOAD whose final size it is
|
||||
// predicting. This is a local copy of a tree that already exists on disk, so its size is known
|
||||
// exactly — the same reasoning PlaceOffsiteRestore's gate uses, and this matches it.
|
||||
free, need := m.offboxFree()(liveNs), m.offboxSize()(scratch)
|
||||
switch {
|
||||
case need <= 0:
|
||||
// FAIL CLOSED. Without this the comparison below is `free < 0`, which is false, and an
|
||||
// unmeasurable scratch would sail straight through into the destructive phase — the gate
|
||||
// present and inert, which is worse than no gate because it reads as protection.
|
||||
m.logger.Printf("[ERROR] [offbox] %s: REFUSING the destructive restore — the scratch size could not be measured (scratch=%s)", stack, scratch)
|
||||
return res, fmt.Errorf(offsiteSizeUnknownMsg)
|
||||
case free <= 0:
|
||||
// Same direction for the other probe. The customer sentence is shared with the case above
|
||||
// (the operator asked for one wording); the LOG line above and below is what distinguishes
|
||||
// which probe failed.
|
||||
m.logger.Printf("[ERROR] [offbox] %s: REFUSING the destructive restore — free space on the live namespace could not be measured (liveNs=%s)", stack, liveNs)
|
||||
return res, fmt.Errorf(offsiteSizeUnknownMsg)
|
||||
case free < need:
|
||||
m.logger.Printf("[WARN] [offbox] %s: REFUSING the destructive restore — need %d B, free %d B on %s; the app was NOT stopped", stack, need, free, liveNs)
|
||||
return res, fmt.Errorf(offsiteNoSpaceMsgFmt, humanizeBytes(need), humanizeBytes(free))
|
||||
}
|
||||
|
||||
placements, err := mapOffsiteRestorePaths(paths, stack, scratch, liveNs)
|
||||
if err != nil {
|
||||
return res, err // whole-placement refusal (no partial writes)
|
||||
|
||||
@@ -30,6 +30,20 @@ const (
|
||||
// SetOffboxFreeFn overrides the restore free-space probe (tests; the Windows go-test host has no df).
|
||||
func (m *Manager) SetOffboxFreeFn(fn func(path string) int64) { m.offboxFreeFn = fn }
|
||||
|
||||
// WriteScratchMarkerForTest exposes the marker writer to the web package's flow test. Test-only by
|
||||
// name so a production caller reads as obviously wrong: only RestoreOffboxScratch may certify a
|
||||
// scratch, because only it knows whether the download finished.
|
||||
func (m *Manager) WriteScratchMarkerForTest(scratch, snapshotID string, full bool) error {
|
||||
return m.writeScratchMarker(scratch, snapshotID, full)
|
||||
}
|
||||
|
||||
// SetOffboxLatestSnapshotFn overrides the restic snapshot lookup (tests; no restic needed). See the
|
||||
// field comment on Manager.offboxLatestSnapFn for why this seam exists rather than a code-reading
|
||||
// argument that the R-357 gate sits early enough.
|
||||
func (m *Manager) SetOffboxLatestSnapshotFn(fn func(ctx context.Context, stack string) (string, []string, error)) {
|
||||
m.offboxLatestSnapFn = fn
|
||||
}
|
||||
|
||||
// SetOffboxFullPlaceCopier overrides the FULL-restore overwrite copier (tests; no rsync needed).
|
||||
func (m *Manager) SetOffboxFullPlaceCopier(fn func(src, dst string) (int, error)) {
|
||||
m.offboxFullPlaceCopier = fn
|
||||
@@ -88,6 +102,9 @@ func offboxUnitPathOf(paths []string, stack string) string {
|
||||
// `snapshots latest --tag <stack> --json`. When the tag spans more than one group (old unit-only shape
|
||||
// + new enlarged shape), it returns the newest by time.
|
||||
func (m *Manager) offboxLatestSnapshot(ctx context.Context, stack string) (id string, paths []string, err error) {
|
||||
if m.offboxLatestSnapFn != nil {
|
||||
return m.offboxLatestSnapFn(ctx, stack)
|
||||
}
|
||||
t := m.settings.GetOffboxTarget()
|
||||
base, env := m.offboxBaseArgs(t)
|
||||
sctx, cancel := context.WithTimeout(ctx, offboxProbeTimeout)
|
||||
@@ -239,14 +256,14 @@ func (m *Manager) RestoreOffboxScratch(ctx context.Context, stack string, full b
|
||||
size, serr := m.offboxSnapshotSize(ctx, id)
|
||||
if serr != nil {
|
||||
// SizeUnknown never renders as fits — fail closed.
|
||||
return fmt.Errorf("A mentés mérete nem állapítható meg — a teljes visszaállítás biztonsági okból nem indítható.")
|
||||
return fmt.Errorf(offsiteSizeUnknownMsg)
|
||||
}
|
||||
need := size + size/10 // ×1.1
|
||||
if free < need {
|
||||
return fmt.Errorf("Nincs elég szabad hely a visszaállításhoz (%s szükséges, %s szabad).", humanizeBytes(need), humanizeBytes(free))
|
||||
return fmt.Errorf(offsiteNoSpaceMsgFmt, humanizeBytes(need), humanizeBytes(free))
|
||||
}
|
||||
} else if free < offboxUnitOnlyFreeFloor {
|
||||
return fmt.Errorf("Nincs elég szabad hely a visszaállításhoz (%s szükséges, %s szabad).", humanizeBytes(offboxUnitOnlyFreeFloor), humanizeBytes(free))
|
||||
return fmt.Errorf(offsiteNoSpaceMsgFmt, humanizeBytes(offboxUnitOnlyFreeFloor), humanizeBytes(free))
|
||||
}
|
||||
// F-A1 hygiene: drop the legacy rootfs scratch (DataDir/offbox-restore/<app>) best-effort.
|
||||
legacy := filepath.Join(m.cfg.Paths.DataDir, "offbox-restore", stack)
|
||||
@@ -260,6 +277,11 @@ func (m *Manager) RestoreOffboxScratch(ctx context.Context, stack string, full b
|
||||
if err := os.MkdirAll(scratch, 0o755); err != nil {
|
||||
return fmt.Errorf("restore dir: %w", err)
|
||||
}
|
||||
// R-358: a marker from a PREVIOUS run must never certify this one. Cleared here, before restic
|
||||
// touches anything, so the window in which a stale certificate could vouch for a part-copy does not
|
||||
// exist. If this run fails, the scratch is left with files and NO marker — which is precisely the
|
||||
// state OffboxFullScratchReady must read as "not ready".
|
||||
m.clearScratchMarker(scratch)
|
||||
t := m.settings.GetOffboxTarget()
|
||||
base, env := m.offboxBaseArgs(t)
|
||||
rctx, cancel := context.WithTimeout(ctx, offboxBackupTimeout)
|
||||
@@ -274,9 +296,94 @@ func (m *Manager) RestoreOffboxScratch(ctx context.Context, stack string, full b
|
||||
return fmt.Errorf("offbox restore %s: %w: %s", stack, rerr, truncate(out))
|
||||
}
|
||||
m.logger.Printf("[INFO] [offbox] restored %s (%s, full=%v) → %s", stack, id, full, scratch)
|
||||
// R-358: the completion certificate, written ONLY now — after restic returned nil. Writing it
|
||||
// earlier would certify a download that has not happened, which is the defect with an extra step.
|
||||
// Written for full=false runs too: the `full` field inside it, not its presence, is what
|
||||
// distinguishes a unit-only scratch from a complete one.
|
||||
if err := m.writeScratchMarker(scratch, id, full); err != nil {
|
||||
// The restore itself succeeded, so this is not an error to fail the operation on — but it is
|
||||
// NOT silent, and the consequence is stated: without the marker the scratch reads as not-ready,
|
||||
// which is the fail-closed direction. Better a re-run than a placement over an uncertified copy.
|
||||
m.logger.Printf("[ERROR] [offbox] %s: restore succeeded but the completion marker could not be written: %v — the scratch will read as NOT ready and the download must be re-run", stack, err)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// --- R-358: the scratch completion marker ------------------------------------------------------
|
||||
//
|
||||
// THE DEFECT. `OffboxFullScratchReady` used to answer "the directory exists and is non-empty". A restic
|
||||
// download that failed part-way leaves exactly that: a directory with files in it. So the product
|
||||
// offered „Teljes visszaállítás indítása" over a part-copy, and pressing it reported success —
|
||||
// observed on demo-hp 2026-08-21. A non-empty directory is evidence that something was written, never
|
||||
// that everything was.
|
||||
//
|
||||
// The marker is the missing fact: not "are there files" but "did the run that wrote them FINISH, and
|
||||
// was it the full one". Only the run itself can know that, so only the run writes it.
|
||||
//
|
||||
// It lives at the scratch ROOT, which is safe from placement for a reason worth stating rather than
|
||||
// assuming: `mapOffsiteRestorePaths` builds placements from the SNAPSHOT's own path list, not from a
|
||||
// directory walk, so a file that exists only locally is invisible to it. That is pinned by
|
||||
// TestR358_MarkerIsNeverPlaced rather than left as a comment.
|
||||
const scratchMarkerName = ".felhom-restore-complete.json"
|
||||
|
||||
// scratchMarker is the on-disk completion certificate. `Schema` is carried so a future format change
|
||||
// is a refusal rather than a misreading — an unrecognised schema fails closed like every other
|
||||
// unreadable marker.
|
||||
type scratchMarker struct {
|
||||
Schema int `json:"schema"`
|
||||
SnapshotID string `json:"snapshot_id"`
|
||||
Full bool `json:"full"`
|
||||
FinishedAt string `json:"finished_at"`
|
||||
}
|
||||
|
||||
const scratchMarkerSchema = 1
|
||||
|
||||
// clearScratchMarker removes any existing marker, best-effort. A failure to remove is logged and NOT
|
||||
// returned: the caller is about to overwrite the scratch anyway, and refusing a restore because a stale
|
||||
// certificate would not delete trades a real capability for a bookkeeping problem.
|
||||
func (m *Manager) clearScratchMarker(scratch string) {
|
||||
if err := os.Remove(filepath.Join(scratch, scratchMarkerName)); err != nil && !os.IsNotExist(err) {
|
||||
m.logger.Printf("[WARN] [offbox] could not clear the stale scratch marker in %s: %v", scratch, err)
|
||||
}
|
||||
}
|
||||
|
||||
// writeScratchMarker writes the certificate atomically (tmp + fsync + rename) at mode 0600. Atomic
|
||||
// because a torn marker read as valid is the one failure this whole mechanism cannot tolerate — it
|
||||
// would certify a part-copy, which is the original defect wearing a new hat.
|
||||
func (m *Manager) writeScratchMarker(scratch, snapshotID string, full bool) error {
|
||||
data, err := json.Marshal(scratchMarker{
|
||||
Schema: scratchMarkerSchema,
|
||||
SnapshotID: snapshotID,
|
||||
Full: full,
|
||||
FinishedAt: time.Now().UTC().Format(time.RFC3339),
|
||||
})
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
final := filepath.Join(scratch, scratchMarkerName)
|
||||
tmp := final + ".tmp"
|
||||
f, err := os.OpenFile(tmp, os.O_WRONLY|os.O_CREATE|os.O_TRUNC, 0o600)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
if _, err := f.Write(data); err != nil {
|
||||
f.Close()
|
||||
os.Remove(tmp)
|
||||
return err
|
||||
}
|
||||
if err := f.Sync(); err != nil {
|
||||
f.Close()
|
||||
os.Remove(tmp)
|
||||
return err
|
||||
}
|
||||
if err := f.Close(); err != nil {
|
||||
os.Remove(tmp)
|
||||
return err
|
||||
}
|
||||
return os.Rename(tmp, final)
|
||||
}
|
||||
|
||||
|
||||
// OffboxRestorePrepareFull resolves the latest snapshot's restore-size and verifies scratch headroom
|
||||
// for a FULL restore WITHOUT starting it (the two-step size-first gate). Returns the human size on
|
||||
// success, or a Hungarian error to flash on refusal (size unknown / no headroom — fail-closed).
|
||||
@@ -293,7 +400,7 @@ func (m *Manager) OffboxRestorePrepareFull(ctx context.Context, stack string) (s
|
||||
}
|
||||
size, serr := m.offboxSnapshotSize(ctx, id)
|
||||
if serr != nil {
|
||||
return "", fmt.Errorf("A mentés mérete nem állapítható meg — a teljes visszaállítás biztonsági okból nem indítható.")
|
||||
return "", fmt.Errorf(offsiteSizeUnknownMsg)
|
||||
}
|
||||
_, nsRoot, derr := m.offboxRestoreScratchDir(stack)
|
||||
if derr != nil {
|
||||
@@ -301,13 +408,37 @@ func (m *Manager) OffboxRestorePrepareFull(ctx context.Context, stack string) (s
|
||||
}
|
||||
need := size + size/10
|
||||
if free := m.offboxFree()(nsRoot); free < need {
|
||||
return "", fmt.Errorf("Nincs elég szabad hely a visszaállításhoz (%s szükséges, %s szabad).", humanizeBytes(need), humanizeBytes(free))
|
||||
return "", fmt.Errorf(offsiteNoSpaceMsgFmt, humanizeBytes(need), humanizeBytes(free))
|
||||
}
|
||||
return humanizeBytes(size), nil
|
||||
}
|
||||
|
||||
// OffboxFullScratchReady reports whether a (non-empty) full-restore scratch exists for stack — the gate
|
||||
// for showing the place-to-live action. PlaceOffsiteRestore re-validates per-path completeness.
|
||||
// R-357 customer-facing refusal strings, shared by every headroom gate on the off-site restore
|
||||
// surface. Named constants because a test asserts them verbatim and because the destructive gate added
|
||||
// in v0.226.0 MUST read identically to the two non-destructive ones that predate it — a customer who
|
||||
// meets this refusal on one path and a differently-worded one on another has to work out whether they
|
||||
// are the same problem.
|
||||
const (
|
||||
offsiteNoSpaceMsgFmt = "Nincs elég szabad hely a visszaállításhoz (%s szükséges, %s szabad)."
|
||||
offsiteSizeUnknownMsg = "A mentés mérete nem állapítható meg — a teljes visszaállítás biztonsági okból nem indítható."
|
||||
)
|
||||
|
||||
// OffboxFullScratchReady reports whether a COMPLETED FULL restore scratch exists for stack — the gate
|
||||
// for the place-to-live and reconstitute actions.
|
||||
//
|
||||
// R-358 — WHAT THIS USED TO ANSWER, AND WHY IT WAS THE WRONG QUESTION. It used to be "the directory
|
||||
// exists and is non-empty", and its doc comment reassured the reader that
|
||||
// `PlaceOffsiteRestore re-validates per-path completeness`. That sentence is what made the weak gate
|
||||
// look adequate, and it is not true in the way it reads: PlaceOffsiteRestore stats the top-level
|
||||
// PLACEMENTS, not the files inside them, so a placement directory that exists but was only half
|
||||
// downloaded passes it. A restic run that died part-way leaves a non-empty directory, so the product
|
||||
// offered „Teljes visszaállítás indítása" over a part-copy and reported success on it (demo-hp,
|
||||
// 2026-08-21).
|
||||
//
|
||||
// It now asks the only question that distinguishes them: did the run that wrote this scratch FINISH,
|
||||
// and was it the full one. Every other answer — no marker, unreadable marker, wrong schema, full=false
|
||||
// — is FALSE, and says at WARN which one it was. **Fail closed: an unreadable marker is not a
|
||||
// completion certificate.**
|
||||
func (m *Manager) OffboxFullScratchReady(stack string) bool {
|
||||
if !isSafeStackName(stack) {
|
||||
return false
|
||||
@@ -319,8 +450,27 @@ func (m *Manager) OffboxFullScratchReady(stack string) bool {
|
||||
if fi, sErr := os.Stat(scratch); sErr != nil || !fi.IsDir() {
|
||||
return false
|
||||
}
|
||||
entries, _ := os.ReadDir(scratch)
|
||||
return len(entries) > 0
|
||||
data, rErr := os.ReadFile(filepath.Join(scratch, scratchMarkerName))
|
||||
if rErr != nil {
|
||||
if !os.IsNotExist(rErr) {
|
||||
m.logger.Printf("[WARN] [offbox] %s: scratch completion marker unreadable (%v) — treating the copy as INCOMPLETE", stack, rErr)
|
||||
}
|
||||
return false
|
||||
}
|
||||
var mk scratchMarker
|
||||
if uErr := json.Unmarshal(data, &mk); uErr != nil {
|
||||
m.logger.Printf("[WARN] [offbox] %s: scratch completion marker does not parse (%v) — treating the copy as INCOMPLETE", stack, uErr)
|
||||
return false
|
||||
}
|
||||
if mk.Schema != scratchMarkerSchema {
|
||||
m.logger.Printf("[WARN] [offbox] %s: scratch completion marker has schema %d, expected %d — treating the copy as INCOMPLETE", stack, mk.Schema, scratchMarkerSchema)
|
||||
return false
|
||||
}
|
||||
if !mk.Full {
|
||||
m.logger.Printf("[INFO] [offbox] %s: scratch holds a UNIT-ONLY restore (snapshot %s) — not a full copy, so place-to-live stays closed", stack, mk.SnapshotID)
|
||||
return false
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
// placement is one source→dest pair for place-to-live: src is the reconstructed absolute path under the
|
||||
@@ -432,7 +582,7 @@ func (m *Manager) PlaceOffsiteRestore(ctx context.Context, stack string) error {
|
||||
// F-3a-1b: headroom gate — a missing-only merge copies at most the scratch size; refuse before any
|
||||
// copy if the live drive lacks that (conservative — scratch and live often share a drive).
|
||||
if free, need := m.offboxFree()(liveNs), m.offboxSize()(scratch); free < need {
|
||||
return fmt.Errorf("Nincs elég szabad hely a visszaállításhoz (%s szükséges, %s szabad).", humanizeBytes(need), humanizeBytes(free))
|
||||
return fmt.Errorf(offsiteNoSpaceMsgFmt, humanizeBytes(need), humanizeBytes(free))
|
||||
}
|
||||
placements, err := mapOffsiteRestorePaths(paths, stack, scratch, liveNs)
|
||||
if err != nil {
|
||||
|
||||
@@ -0,0 +1,176 @@
|
||||
package backup
|
||||
|
||||
import (
|
||||
"context"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"sync/atomic"
|
||||
"testing"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-controller/internal/settings"
|
||||
)
|
||||
|
||||
// ── R-357 — the ONE restore that deletes and replaces data had no free-space gate ────────────────
|
||||
//
|
||||
// The two gentler paths both check (offbox_restore.go: the prepare gate and PlaceOffsiteRestore's
|
||||
// missing-only merge). `offbox_reconstitute.go` contained ZERO references to offboxFree.
|
||||
//
|
||||
// Measured on demo-hp 2026-08-21: the destructive restore stopped the app, ran out of disk part-way,
|
||||
// left 2 of 5 planted items in place and restarted the app — a half-restored dataset presented as a
|
||||
// completed restore.
|
||||
//
|
||||
// WHAT THESE ASSERT IS THE NON-EFFECT, NOT THE ERROR STRING. The value of the fix is that the app is
|
||||
// never stopped: a gate placed after StopStack would turn a refusal into an outage and would still
|
||||
// return the right sentence. So `stops == 0` is the assertion that can fail on a wrong-but-plausible
|
||||
// implementation, and the message is checked second.
|
||||
|
||||
// r357Provider counts StopStack so the non-effect is measurable, and reports the app as deployed so
|
||||
// the reconstitute reaches the gate rather than refusing earlier for an unrelated reason.
|
||||
type r357Provider struct {
|
||||
hdd string
|
||||
stops int32
|
||||
}
|
||||
|
||||
func (p *r357Provider) GetStackComposePath(string) (string, bool) { return "", false }
|
||||
func (p *r357Provider) ListDeployedStacks() []StackSummary {
|
||||
return []StackSummary{{Name: "paperless-ngx"}}
|
||||
}
|
||||
func (p *r357Provider) GetStackHDDMounts(string) []string { return nil }
|
||||
func (p *r357Provider) GetStackHDDPath(string) string { return p.hdd }
|
||||
func (p *r357Provider) GetImportRoot() string { return "" }
|
||||
func (p *r357Provider) GetDockerVolumes(string) []string { return nil }
|
||||
func (p *r357Provider) StopStack(string) error { atomic.AddInt32(&p.stops, 1); return nil }
|
||||
func (p *r357Provider) StartStack(string) error { return nil }
|
||||
func (p *r357Provider) RefreshAndIsRunning(string) bool { return true }
|
||||
func (p *r357Provider) GetStackRecoveryInfo(string) (RecoveryInfo, bool) {
|
||||
return RecoveryInfo{}, false
|
||||
}
|
||||
func (p *r357Provider) RecoverStackSecrets(string, []string) map[string]string { return nil }
|
||||
func (p *r357Provider) RecreateStackDefinitionFromUnit(string, string, map[string]string) error {
|
||||
return nil
|
||||
}
|
||||
func (p *r357Provider) StartStackServices(string, []string) error { return nil }
|
||||
func (p *r357Provider) GetStackClassifiedBinds(string) ([]ClassifiedBind, bool) {
|
||||
return nil, false
|
||||
}
|
||||
|
||||
// newR357Manager builds a manager whose reconstitute reaches the headroom gate: a real scratch on
|
||||
// disk, a stubbed snapshot lookup (no restic), and the app reported deployed.
|
||||
func newR357Manager(t *testing.T) (*Manager, *r357Provider) {
|
||||
t.Helper()
|
||||
m, sett := newOffboxManager(t)
|
||||
drive := t.TempDir()
|
||||
if err := sett.AddStoragePath(settings.StoragePath{Path: drive, Label: "drive", Schedulable: true}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
prov := &r357Provider{hdd: drive}
|
||||
m.SetStackProvider(prov)
|
||||
|
||||
scratch, _, err := m.offboxRestoreScratchDir("paperless-ngx")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
// THE FIXTURE MUST REACH StopStack WHEN THE GATE IS REMOVED, or `stops == 0` proves nothing.
|
||||
// The first draft of this test did not: without the gate the run refused earlier, at the stat
|
||||
// pre-pass over the placements, so the assertion passed against the pre-fix code. That is a hollow
|
||||
// test, and the red-proof is what exposed it — recorded here because the near-miss is the lesson.
|
||||
//
|
||||
// So the scratch is populated the way a real completed download leaves it: `mapOffsiteRestorePaths`
|
||||
// builds each src as filepath.Join(scratch, <full snapshot path>), so the snapshot's own absolute
|
||||
// path is mirrored underneath the scratch.
|
||||
const oldNs = "/mnt/old"
|
||||
snapPaths := []string{oldNs + "/backups/primary/paperless-ngx", oldNs + "/appdata/paperless-ngx"}
|
||||
for _, sp := range snapPaths {
|
||||
if err := os.MkdirAll(filepath.Join(scratch, sp), 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
m.SetOffboxLatestSnapshotFn(func(context.Context, string) (string, []string, error) {
|
||||
return "snap-1", snapPaths, nil
|
||||
})
|
||||
// No Docker in a unit test: the undo copy is seamed out. It runs BEFORE StopStack, so leaving it
|
||||
// real would make the fixture fail for a reason that has nothing to do with the gate.
|
||||
m.SetSafetyDumpFn(func(context.Context, DiscoveredDB, string) DumpResult { return DumpResult{} })
|
||||
return m, prov
|
||||
}
|
||||
|
||||
func TestR357_DestructiveRestoreRefusesWithoutHeadroom(t *testing.T) {
|
||||
m, prov := newR357Manager(t)
|
||||
m.SetOffboxSizer(func(string) int64 { return 1024 * 1024 }) // the scratch is 1 MB
|
||||
m.SetOffboxFreeFn(func(string) int64 { return 300 * 1024 }) // 300 KB free
|
||||
|
||||
_, err := m.ReconstituteFromOffsite(context.Background(), "paperless-ngx", false)
|
||||
|
||||
// THE ASSERTION THAT MATTERS, AND IT IS CHECKED FIRST ON PURPOSE. On 2026-08-21 the app went down
|
||||
// and came back over a half-written dataset. A gate placed after StopStack would return the right
|
||||
// sentence and still take the outage, so the error text cannot be the primary assertion — and if
|
||||
// this is checked second, a removed gate reports "no error returned" instead of naming the outage.
|
||||
if n := atomic.LoadInt32(&prov.stops); n != 0 {
|
||||
t.Fatalf("THE APP WAS STOPPED (%d call(s)) for a restore with 300 KB free for a 1 MB copy — "+
|
||||
"the whole point of this gate is that the customer's app never goes down for a restore "+
|
||||
"that cannot run (err=%v)", n, err)
|
||||
}
|
||||
if err == nil {
|
||||
t.Fatal("the destructive restore proceeded with 300 KB free for a 1 MB copy")
|
||||
}
|
||||
for _, want := range []string{"Nincs elég szabad hely", "szükséges", "szabad"} {
|
||||
if !strings.Contains(err.Error(), want) {
|
||||
t.Errorf("the refusal must name need and free like its two siblings; missing %q in %q", want, err.Error())
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestR357_UnknownSizeFailsClosed(t *testing.T) {
|
||||
// The fail-open hole this closes is subtle: with need = 0 the comparison `free < need` is FALSE,
|
||||
// so an unmeasurable scratch sailed straight into the destructive phase. A gate that is present
|
||||
// and inert is worse than no gate, because it reads as protection.
|
||||
m, prov := newR357Manager(t)
|
||||
m.SetOffboxSizer(func(string) int64 { return 0 })
|
||||
m.SetOffboxFreeFn(func(string) int64 { return 100 << 30 }) // plenty free — irrelevant
|
||||
|
||||
_, err := m.ReconstituteFromOffsite(context.Background(), "paperless-ngx", false)
|
||||
if n := atomic.LoadInt32(&prov.stops); n != 0 {
|
||||
t.Fatalf("THE APP WAS STOPPED (%d call(s)) on an unknown size — fail-closed means refuse "+
|
||||
"BEFORE the stop, not report an error after it (err=%v)", n, err)
|
||||
}
|
||||
if err == nil {
|
||||
t.Fatal("an unmeasurable scratch was allowed into the destructive restore")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "nem állapítható meg") {
|
||||
t.Errorf("want the size-unknown wording already used by OffboxRestorePrepareFull; got %q", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestR357_UnknownFreeSpaceFailsClosed(t *testing.T) {
|
||||
// The mirror hole: free = 0 and need = 0 also compares false. Both probes fail closed.
|
||||
m, prov := newR357Manager(t)
|
||||
m.SetOffboxSizer(func(string) int64 { return 1024 })
|
||||
m.SetOffboxFreeFn(func(string) int64 { return 0 })
|
||||
|
||||
_, err := m.ReconstituteFromOffsite(context.Background(), "paperless-ngx", false)
|
||||
if n := atomic.LoadInt32(&prov.stops); n != 0 {
|
||||
t.Fatalf("THE APP WAS STOPPED (%d call(s)) on an unknown free reading (err=%v)", n, err)
|
||||
}
|
||||
if err == nil {
|
||||
t.Fatal("an unmeasurable free-space reading was allowed into the destructive restore")
|
||||
}
|
||||
}
|
||||
|
||||
func TestR357_AmpleSpaceIsUnchanged(t *testing.T) {
|
||||
// The gate must not become a new way to fail an ordinary restore. With room to spare it does not
|
||||
// fire, and the run proceeds past it — which here means it fails LATER, for its own unrelated
|
||||
// reasons, never with a headroom sentence.
|
||||
m, _ := newR357Manager(t)
|
||||
m.SetOffboxSizer(func(string) int64 { return 1024 })
|
||||
m.SetOffboxFreeFn(func(string) int64 { return 100 << 30 })
|
||||
|
||||
_, err := m.ReconstituteFromOffsite(context.Background(), "paperless-ngx", false)
|
||||
if err != nil && strings.Contains(err.Error(), "Nincs elég szabad hely") {
|
||||
t.Fatalf("the headroom gate fired with 100 GiB free for a 1 KiB copy: %v", err)
|
||||
}
|
||||
if err != nil && strings.Contains(err.Error(), "nem állapítható meg") {
|
||||
t.Fatalf("the size-unknown branch fired on a measurable scratch: %v", err)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,248 @@
|
||||
package backup
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"go/ast"
|
||||
"go/parser"
|
||||
"go/token"
|
||||
"log"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-controller/internal/settings"
|
||||
)
|
||||
|
||||
// ── R-358 — a failed download was offered as a good one ──────────────────────────────────────────
|
||||
//
|
||||
// `OffboxFullScratchReady` used to answer "the directory exists and is non-empty". A restic run that
|
||||
// dies part-way leaves exactly that. So the product showed „Teljes visszaállítás indítása" over a
|
||||
// part-copy and pressing it reported success — observed on demo-hp 2026-08-21.
|
||||
//
|
||||
// A non-empty directory is evidence that SOMETHING was written, never that everything was. The marker
|
||||
// carries the only fact that distinguishes them: did the run FINISH, and was it the full one.
|
||||
|
||||
// newR358Manager gives a manager whose scratch resolves into a t.TempDir().
|
||||
func newR358Manager(t *testing.T) (*Manager, string) {
|
||||
t.Helper()
|
||||
m, sett := newOffboxManager(t)
|
||||
drive := t.TempDir()
|
||||
if err := sett.AddStoragePath(settings.StoragePath{Path: drive, Label: "drive", Schedulable: true}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
m.SetStackProvider(&offbox3aProvider{
|
||||
hdd: map[string]string{"kimai": drive}, binds: map[string][]ClassifiedBind{}, has: map[string]bool{},
|
||||
})
|
||||
scratch, _, err := m.offboxRestoreScratchDir("kimai")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := os.MkdirAll(scratch, 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return m, scratch
|
||||
}
|
||||
|
||||
// writeScratchPayload plants the files a part-way restic run leaves behind.
|
||||
func writeScratchPayload(t *testing.T, scratch string) {
|
||||
t.Helper()
|
||||
if err := os.MkdirAll(filepath.Join(scratch, "backups", "primary", "kimai"), 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := os.WriteFile(filepath.Join(scratch, "backups", "primary", "kimai", "half.tar"), []byte("partial"), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestR358_FailedRestoreLeavesNoUsableScratch(t *testing.T) {
|
||||
m, scratch := newR358Manager(t)
|
||||
writeScratchPayload(t, scratch) // files present, run never finished → no marker
|
||||
|
||||
if m.OffboxFullScratchReady("kimai") {
|
||||
t.Fatal("a part-copy was reported READY — this is the defect: the customer is offered " +
|
||||
"„Teljes visszaállítás indítása" + " over a download that never finished")
|
||||
}
|
||||
}
|
||||
|
||||
func TestR358_UnitOnlyScratchIsNotFullReady(t *testing.T) {
|
||||
m, scratch := newR358Manager(t)
|
||||
writeScratchPayload(t, scratch)
|
||||
if err := m.writeScratchMarker(scratch, "snap-1", false); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if m.OffboxFullScratchReady("kimai") {
|
||||
t.Fatal("a UNIT-ONLY scratch was reported ready for a full restore — both modes write the " +
|
||||
"same directory, so only the marker's `full` field separates them")
|
||||
}
|
||||
}
|
||||
|
||||
func TestR358_StaleMarkerIsClearedBeforeTheRun(t *testing.T) {
|
||||
m, scratch := newR358Manager(t)
|
||||
if err := m.writeScratchMarker(scratch, "snap-OLD", true); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if !m.OffboxFullScratchReady("kimai") {
|
||||
t.Fatal("fixture wrong: a valid full marker should read ready")
|
||||
}
|
||||
// What the next run does before restic touches anything.
|
||||
m.clearScratchMarker(scratch)
|
||||
writeScratchPayload(t, scratch) // ...and then that run dies part-way
|
||||
|
||||
if m.OffboxFullScratchReady("kimai") {
|
||||
t.Fatal("a marker from a PREVIOUS run certified a later part-copy — the stale certificate " +
|
||||
"is the whole reason the clear happens before restic, not after")
|
||||
}
|
||||
}
|
||||
|
||||
func TestR358_UnreadableMarkerFailsClosed(t *testing.T) {
|
||||
var buf bytes.Buffer
|
||||
m, scratch := newR358Manager(t)
|
||||
m.logger = log.New(&buf, "", 0)
|
||||
writeScratchPayload(t, scratch)
|
||||
if err := os.WriteFile(filepath.Join(scratch, scratchMarkerName), []byte("{not json"), 0o600); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
if m.OffboxFullScratchReady("kimai") {
|
||||
t.Fatal("an unparseable marker was treated as a completion certificate")
|
||||
}
|
||||
if !strings.Contains(buf.String(), "WARN") {
|
||||
t.Errorf("a scratch refused for an unreadable marker must say so — silence makes a refusal "+
|
||||
"indistinguishable from a missing download; log was %q", buf.String())
|
||||
}
|
||||
}
|
||||
|
||||
func TestR358_WrongSchemaFailsClosed(t *testing.T) {
|
||||
m, scratch := newR358Manager(t)
|
||||
writeScratchPayload(t, scratch)
|
||||
if err := os.WriteFile(filepath.Join(scratch, scratchMarkerName),
|
||||
[]byte(`{"schema":99,"snapshot_id":"s","full":true,"finished_at":"2026-08-30T00:00:00Z"}`), 0o600); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if m.OffboxFullScratchReady("kimai") {
|
||||
t.Fatal("a marker with an unrecognised schema was accepted — a format we cannot read is not a certificate")
|
||||
}
|
||||
}
|
||||
|
||||
func TestR358_CompletedFullScratchStillReady(t *testing.T) {
|
||||
// The happy path is unchanged: a finished full download is still offered.
|
||||
m, scratch := newR358Manager(t)
|
||||
writeScratchPayload(t, scratch)
|
||||
if err := m.writeScratchMarker(scratch, "snap-1", true); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if !m.OffboxFullScratchReady("kimai") {
|
||||
t.Fatal("a COMPLETED full restore is no longer offered — the fix broke the thing it protects")
|
||||
}
|
||||
}
|
||||
|
||||
func TestR358_MarkerIsWrittenAt0600AndAtomically(t *testing.T) {
|
||||
m, scratch := newR358Manager(t)
|
||||
if err := m.writeScratchMarker(scratch, "snap-1", true); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
fi, err := os.Stat(filepath.Join(scratch, scratchMarkerName))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if fi.Mode().Perm() != 0o600 {
|
||||
t.Errorf("marker mode = %v, want 0600", fi.Mode().Perm())
|
||||
}
|
||||
// The tmp file must not survive: a leftover .tmp beside the marker is a torn write that a later
|
||||
// reader could mistake for the real thing.
|
||||
if _, err := os.Stat(filepath.Join(scratch, scratchMarkerName+".tmp")); !os.IsNotExist(err) {
|
||||
t.Error("the temporary marker file was left behind")
|
||||
}
|
||||
}
|
||||
|
||||
// TestR358_MarkerIsNeverPlaced pins the assumption the whole design rests on: placement is driven by
|
||||
// the SNAPSHOT's own path list, not by a directory walk, so a file that exists only locally cannot be
|
||||
// copied into the customer's live data. Stated as a test rather than trusted as a comment — the spec
|
||||
// asked for exactly this, and "a comment asserting an invariant needs a test pinning it" is a standing
|
||||
// rule earned nine times over in this project.
|
||||
func TestR358_MarkerIsNeverPlaced(t *testing.T) {
|
||||
const stack = "kimai"
|
||||
oldNs := "/mnt/old"
|
||||
scratch := t.TempDir()
|
||||
liveNs := t.TempDir()
|
||||
snapPaths := []string{
|
||||
oldNs + "/backups/primary/" + stack,
|
||||
oldNs + "/appdata/" + stack,
|
||||
}
|
||||
|
||||
placements, err := mapOffsiteRestorePaths(snapPaths, stack, scratch, liveNs)
|
||||
if err != nil {
|
||||
t.Fatalf("mapOffsiteRestorePaths: %v", err)
|
||||
}
|
||||
if len(placements) == 0 {
|
||||
t.Fatal("fixture produced no placements — the test would prove nothing")
|
||||
}
|
||||
for _, pl := range placements {
|
||||
if strings.Contains(pl.src, scratchMarkerName) || strings.Contains(pl.dst, scratchMarkerName) {
|
||||
t.Fatalf("the completion marker entered a placement (src=%q dst=%q) — it would be copied "+
|
||||
"into the customer's live data", pl.src, pl.dst)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestR358_MarkerIsClearedBeforeResticAndWrittenAfter walks the AST of RestoreOffboxScratch.
|
||||
//
|
||||
// It exists because `resticStep` is not a seam — a test cannot run the real download without restic,
|
||||
// so the ORDER of the three calls cannot be proven by execution here. Order is the entire safety
|
||||
// property: a marker written before restic certifies a download that has not happened, and a clear
|
||||
// that runs after it leaves a stale certificate covering a fresh part-copy. A substring search would
|
||||
// not do: a commented-out call satisfies strings.Contains, which a sibling test in this project
|
||||
// records paying for.
|
||||
func TestR358_MarkerIsClearedBeforeResticAndWrittenAfter(t *testing.T) {
|
||||
fset := token.NewFileSet()
|
||||
f, err := parser.ParseFile(fset, "offbox_restore.go", nil, 0)
|
||||
if err != nil {
|
||||
t.Fatalf("parse offbox_restore.go: %v", err)
|
||||
}
|
||||
var body *ast.BlockStmt
|
||||
for _, d := range f.Decls {
|
||||
if fn, ok := d.(*ast.FuncDecl); ok && fn.Name.Name == "RestoreOffboxScratch" && fn.Body != nil {
|
||||
body = fn.Body
|
||||
}
|
||||
}
|
||||
if body == nil {
|
||||
t.Fatal("RestoreOffboxScratch not found")
|
||||
}
|
||||
|
||||
var order []string
|
||||
ast.Inspect(body, func(n ast.Node) bool {
|
||||
call, ok := n.(*ast.CallExpr)
|
||||
if !ok {
|
||||
return true
|
||||
}
|
||||
if sel, ok := call.Fun.(*ast.SelectorExpr); ok {
|
||||
switch sel.Sel.Name {
|
||||
case "clearScratchMarker", "resticStep", "writeScratchMarker":
|
||||
order = append(order, sel.Sel.Name)
|
||||
}
|
||||
}
|
||||
return true
|
||||
})
|
||||
|
||||
idx := func(name string) int {
|
||||
for i, n := range order {
|
||||
if n == name {
|
||||
return i
|
||||
}
|
||||
}
|
||||
return -1
|
||||
}
|
||||
clear, restic, write := idx("clearScratchMarker"), idx("resticStep"), idx("writeScratchMarker")
|
||||
if clear < 0 || restic < 0 || write < 0 {
|
||||
t.Fatalf("RestoreOffboxScratch does not call all three (order seen: %v) — the marker is not wired", order)
|
||||
}
|
||||
if !(clear < restic) {
|
||||
t.Errorf("the stale marker is cleared AFTER restic runs (order %v) — a previous run's "+
|
||||
"certificate would cover this run's part-copy", order)
|
||||
}
|
||||
if !(restic < write) {
|
||||
t.Errorf("the marker is written BEFORE restic returns (order %v) — that certifies a download "+
|
||||
"that has not happened, which is the defect with an extra step", order)
|
||||
}
|
||||
}
|
||||
@@ -164,7 +164,7 @@ func TestReconstituteRefusesWhenNoDBServiceIdentifiable(t *testing.T) {
|
||||
func TestRestoreFromUnitRefusesWhenNoDBServiceIdentifiable(t *testing.T) {
|
||||
m, prov, _ := r47UnitFixture(t, noDBCompose, true)
|
||||
|
||||
err := m.RestoreFromRecoveryUnit("app")
|
||||
_, err := m.RestoreFromRecoveryUnit("app")
|
||||
if err == nil {
|
||||
t.Fatal("expected a refusal: the unit carries a dump but names no startable database service")
|
||||
}
|
||||
@@ -202,7 +202,7 @@ func TestRestoreFromUnitReplaysWithOnlyTheDBServiceUp(t *testing.T) {
|
||||
return nil
|
||||
}
|
||||
|
||||
if err := m.RestoreFromRecoveryUnit("app"); err != nil {
|
||||
if _, err := m.RestoreFromRecoveryUnit("app"); err != nil {
|
||||
t.Fatalf("restore-from-unit: %v", err)
|
||||
}
|
||||
if len(*imported) != 1 {
|
||||
@@ -227,7 +227,7 @@ func TestRestoreFromUnitReplaysWithOnlyTheDBServiceUp(t *testing.T) {
|
||||
func TestRestoreFromUnitNoDumpsTakesOneFullStart(t *testing.T) {
|
||||
m, prov, imported := r47UnitFixture(t, noDBCompose, false)
|
||||
|
||||
if err := m.RestoreFromRecoveryUnit("app"); err != nil {
|
||||
if _, err := m.RestoreFromRecoveryUnit("app"); err != nil {
|
||||
t.Fatalf("restore-from-unit: %v", err)
|
||||
}
|
||||
if len(prov.gotServices) != 0 {
|
||||
@@ -251,7 +251,7 @@ func TestRestoreFromUnitIgnoresSafetyDumpsWhenDecidingToReplay(t *testing.T) {
|
||||
mustWrite(t, filepath.Join(AppDBDumpPath(prov.hdd, "app"),
|
||||
preRestoreDumpPrefix+"20260720T101010Z-app-postgres.sql"), pgDump(1))
|
||||
|
||||
if err := m.RestoreFromRecoveryUnit("app"); err != nil {
|
||||
if _, err := m.RestoreFromRecoveryUnit("app"); err != nil {
|
||||
t.Fatalf("a lone safety dump must not turn into a refusal: %v", err)
|
||||
}
|
||||
if len(prov.gotServices) != 0 {
|
||||
@@ -327,7 +327,7 @@ func TestRestoreFromUnitReplayFailureStillBringsTheStackUp(t *testing.T) {
|
||||
return context.DeadlineExceeded
|
||||
}
|
||||
|
||||
err := m.RestoreFromRecoveryUnit("app")
|
||||
_, err := m.RestoreFromRecoveryUnit("app")
|
||||
if err == nil {
|
||||
t.Fatal("a failed replay must be surfaced, not swallowed")
|
||||
}
|
||||
|
||||
@@ -64,7 +64,7 @@ func (m *Manager) RestoreApp(stackName, snapshotID string) error {
|
||||
if m.isDebug() {
|
||||
m.logger.Printf("[DEBUG] RestoreApp: step 2/3 — restoring Docker volumes for %s", stackName)
|
||||
}
|
||||
if err := m.restoreDockerVolumes(stackName, drivePath); err != nil {
|
||||
if _, err := m.restoreDockerVolumes(stackName, drivePath); err != nil {
|
||||
m.logger.Printf("[ERROR] RESTORE volume restore failed for %s: %v", stackName, err)
|
||||
dataErr = err
|
||||
}
|
||||
@@ -98,10 +98,18 @@ func (m *Manager) RestoreApp(stackName, snapshotID string) error {
|
||||
return nil
|
||||
}
|
||||
|
||||
// restoreDockerVolumes populates Docker volumes from the tars in the app's LIVE recovery unit.
|
||||
func (m *Manager) restoreDockerVolumes(stackName, drivePath string) error {
|
||||
_, err := m.restoreDockerVolumesFrom(stackName, AppVolumeDumpPath(m.namespaceRoot(drivePath), stackName))
|
||||
return err
|
||||
// restoreDockerVolumes populates Docker volumes from the tars in the app's LIVE recovery unit, and
|
||||
// returns HOW MANY it replayed.
|
||||
//
|
||||
// R-353: the count used to be discarded here. `restoreDockerVolumesFrom` has always returned it, so
|
||||
// the fact existed one call deep and was thrown away one line later — which left the unit-restore
|
||||
// path structurally unable to tell a customer whether any data came back. On 2026-08-21 an opengist
|
||||
// restore reported completion over a unit holding manifest.json and compose/ and nothing else, and no
|
||||
// screen could have said otherwise. Discarding a fact the caller needs is cheaper to fix than to
|
||||
// re-derive: the caller cannot count volumes afterwards without re-reading the directory the restore
|
||||
// has already consumed.
|
||||
func (m *Manager) restoreDockerVolumes(stackName, drivePath string) (int, error) {
|
||||
return m.restoreDockerVolumesFrom(stackName, AppVolumeDumpPath(m.namespaceRoot(drivePath), stackName))
|
||||
}
|
||||
|
||||
// restoreDockerVolumesFrom is restoreDockerVolumes with an EXPLICIT dump directory, and it returns how
|
||||
|
||||
@@ -47,7 +47,7 @@ func TestRestoreGeneratesMissingResettableSecret(t *testing.T) {
|
||||
return genValue, true
|
||||
}
|
||||
|
||||
if err := m.RestoreFromRecoveryUnit("app"); err != nil {
|
||||
if _, err := m.RestoreFromRecoveryUnit("app"); err != nil {
|
||||
t.Fatalf("restore must proceed for a missing RESETTABLE secret: %v", err)
|
||||
}
|
||||
if fake.gotEnv == nil {
|
||||
@@ -90,7 +90,7 @@ func TestRestoreProceedsWhenNoGenerator(t *testing.T) {
|
||||
m.generateSecret = func(string, string) (string, bool) { return "", false } // no spec (Scenario G)
|
||||
}
|
||||
|
||||
if err := m.RestoreFromRecoveryUnit("app"); err != nil {
|
||||
if _, err := m.RestoreFromRecoveryUnit("app"); err != nil {
|
||||
t.Fatalf("restore must still proceed: %v", err)
|
||||
}
|
||||
if _, present := fake.gotEnv["DB_PASSWORD"]; present {
|
||||
@@ -122,7 +122,7 @@ func TestRestoreGenerationNeverReachesDataKeys(t *testing.T) {
|
||||
return "eager-value", true
|
||||
}
|
||||
|
||||
if err := m.RestoreFromRecoveryUnit("app"); err == nil {
|
||||
if _, err := m.RestoreFromRecoveryUnit("app"); err == nil {
|
||||
t.Fatal("missing data-key must still refuse fail-closed")
|
||||
}
|
||||
if len(genCalls) != 0 {
|
||||
|
||||
@@ -129,6 +129,34 @@ func hasReplayableDump(dumpDir string) bool {
|
||||
return false
|
||||
}
|
||||
|
||||
// UnitRestoreResult is what a local recovery-unit restore actually did, so the surface can STATE it
|
||||
// rather than report a bare completion.
|
||||
//
|
||||
// It exists for the same reason OffsiteReconstituteResult does, and it is the same lesson arriving on
|
||||
// the other path: on 2026-08-21 an opengist restore reported "Restore-from-unit completed" over a unit
|
||||
// holding manifest.json and compose/ and nothing else, and no screen could have told the customer that
|
||||
// no data had been returned (R-353).
|
||||
//
|
||||
// The Manifest* counts are carried BECAUSE zero-replayed has two causes and they are not the same
|
||||
// fact. A unit that lists no dumps means THE BACKUP held no data. A unit that lists dumps none of which
|
||||
// replayed means something is wrong and the customer's live data was left untouched. R-355 is the
|
||||
// standing rule this obeys: a claim about the APP must never be inferred from a counter — and here it
|
||||
// is not merely unproven but unprovable, because 07-backup-architecture §6.3 records that an app's
|
||||
// canonical .sql could be absent from the unit for reasons that have nothing to do with whether the app
|
||||
// has a database (R-361 destroyed exactly that file for four months).
|
||||
type UnitRestoreResult struct {
|
||||
// VolumesReplayed is how many named-volume tars were unpacked into live Docker volumes.
|
||||
VolumesReplayed int
|
||||
// DBsReplayed is how many .sql dumps were imported. Never inferred from the presence of a database
|
||||
// service — only a completed import increments it.
|
||||
DBsReplayed int
|
||||
// ManifestVolumes is len(manifest.VolumeDumps): what the unit CLAIMS it captured. The gap between
|
||||
// this and VolumesReplayed is the whole of Scenario C.
|
||||
ManifestVolumes int
|
||||
// ManifestDBs is len(manifest.DBDumps): the same claim for the database leg.
|
||||
ManifestDBs int
|
||||
}
|
||||
|
||||
// RestoreFromRecoveryUnit recreates an app from its on-drive recovery unit.
|
||||
//
|
||||
// It reads the unit manifest, takes the portable secrets from the UNIT and the rest from the guest's
|
||||
@@ -140,15 +168,16 @@ func hasReplayableDump(dumpDir string) bool {
|
||||
// D5: this no longer needs the guest. A restore with the guest's app.yaml absent succeeds, which is
|
||||
// pinned by TestRestoreFromRecoveryUnitWithGuestAbsent — the withheld class is regenerated (O4) and
|
||||
// only a data key missing from BOTH sources still refuses.
|
||||
func (m *Manager) RestoreFromRecoveryUnit(stackName string) error {
|
||||
func (m *Manager) RestoreFromRecoveryUnit(stackName string) (UnitRestoreResult, error) {
|
||||
var res UnitRestoreResult
|
||||
if m.stackProvider == nil {
|
||||
return fmt.Errorf("stack provider not configured")
|
||||
return res, fmt.Errorf("stack provider not configured")
|
||||
}
|
||||
|
||||
m.mu.Lock()
|
||||
if m.running {
|
||||
m.mu.Unlock()
|
||||
return fmt.Errorf("backup or restore already in progress")
|
||||
return res, fmt.Errorf("backup or restore already in progress")
|
||||
}
|
||||
m.running = true
|
||||
m.mu.Unlock()
|
||||
@@ -160,7 +189,7 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) error {
|
||||
|
||||
drivePath := m.GetAppDrivePath(stackName)
|
||||
if drivePath == "" || !filepath.IsAbs(drivePath) {
|
||||
return fmt.Errorf("cannot determine drive path for %s", stackName)
|
||||
return res, fmt.Errorf("cannot determine drive path for %s", stackName)
|
||||
}
|
||||
nsRoot := m.namespaceRoot(drivePath)
|
||||
|
||||
@@ -170,9 +199,18 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) error {
|
||||
m.mu.Lock()
|
||||
m.running = false // RestoreApp re-acquires the running flag
|
||||
m.mu.Unlock()
|
||||
return m.RestoreApp(stackName, "")
|
||||
// The fallback path has no unit and therefore no manifest to count against: a ZERO result is
|
||||
// the honest answer, not a missing one. The surface must be able to tell "nothing came back"
|
||||
// from "we never looked", and it can — ManifestVolumes/ManifestDBs are zero too, which is
|
||||
// Scenario B's shape and reads as "the backup held no data", which is exactly true of a box
|
||||
// with no recovery unit.
|
||||
return res, m.RestoreApp(stackName, "")
|
||||
}
|
||||
|
||||
// R-353: what the unit CLAIMS it holds, recorded before any mutation. Read from the manifest that
|
||||
// was just parsed above, so the claim and the outcome are counted from the same document.
|
||||
res.ManifestVolumes, res.ManifestDBs = len(manifest.VolumeDumps), len(manifest.DBDumps)
|
||||
|
||||
composeDir := RecoveryUnitComposePath(nsRoot, stackName)
|
||||
nonSecretEnv, unitSecrets := readUnitEnv(filepath.Join(composeDir, "app.yaml"), manifest.PortableSecretEnvVars)
|
||||
|
||||
@@ -185,7 +223,7 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) error {
|
||||
fullEnv, missing, err := reconcileRestoreSecrets(nonSecretEnv, unitSecrets, guestSecrets, manifest.SecretEnvVars, manifest.DataKeyEnvVars)
|
||||
if err != nil {
|
||||
m.logger.Printf("[ERROR] [backup] Restore REFUSED for %s: %v", stackName, err)
|
||||
return err
|
||||
return res, err
|
||||
}
|
||||
// O4: a missing RESETTABLE secret used to redeploy blank (compose "Defaulting to a blank
|
||||
// string" → exit 1). Generate a replacement via the deploy flow's generator instead —
|
||||
@@ -242,7 +280,7 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) error {
|
||||
hasDumps := hasReplayableDump(AppDBDumpPath(nsRoot, stackName))
|
||||
if hasDumps && len(dbServices) == 0 {
|
||||
m.logger.Printf("[ERROR] [backup] Restore REFUSED for %s: a .sql dump exists but no database service is identifiable in the unit's compose", stackName)
|
||||
return fmt.Errorf("Az adatbázis-szolgáltatás nem azonosítható a(z) %s alkalmazásban — a visszaállítás biztonsági okból nem indult el.", stackName)
|
||||
return res, fmt.Errorf("Az adatbázis-szolgáltatás nem azonosítható a(z) %s alkalmazásban — a visszaállítás biztonsági okból nem indult el.", stackName)
|
||||
}
|
||||
|
||||
// Stop, restore named-volume data, recreate the definition, replay the DB with ONLY the database
|
||||
@@ -252,12 +290,17 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) error {
|
||||
if err := m.stackProvider.StopStack(stackName); err != nil {
|
||||
m.logger.Printf("[WARN] [backup] could not stop %s before restore: %v (continuing)", stackName, err)
|
||||
}
|
||||
if err := m.restoreDockerVolumes(stackName, drivePath); err != nil {
|
||||
m.logger.Printf("[ERROR] [backup] volume restore for %s: %v", stackName, err)
|
||||
dataErr = err
|
||||
// R-353: the count is captured even when the replay errors — a partial replay is a fact the
|
||||
// customer's sentence has to be built from, and discarding it on the error path is how Scenario C
|
||||
// would end up wearing Scenario B's wording.
|
||||
replayed, volErr := m.restoreDockerVolumes(stackName, drivePath)
|
||||
res.VolumesReplayed = replayed
|
||||
if volErr != nil {
|
||||
m.logger.Printf("[ERROR] [backup] volume restore for %s: %v", stackName, volErr)
|
||||
dataErr = volErr
|
||||
}
|
||||
if err := m.stackProvider.RecreateStackDefinitionFromUnit(stackName, composeDir, fullEnv); err != nil {
|
||||
return fmt.Errorf("recreating %s from unit: %w", stackName, err)
|
||||
return res, fmt.Errorf("recreating %s from unit: %w", stackName, err)
|
||||
}
|
||||
// F17: the captured .sql dump is the authoritative logical DB state — replay it AFTER the volume
|
||||
// restore, so the dump WINS over any volume-tar copy of the database.
|
||||
@@ -270,23 +313,27 @@ func (m *Manager) RestoreFromRecoveryUnit(stackName string) error {
|
||||
if dataErr == nil {
|
||||
dataErr = err
|
||||
}
|
||||
} else if _, err := m.reimportDBDumpsCtx(stackName, nsRoot); err != nil {
|
||||
} else if n, err := m.reimportDBDumpsCtx(stackName, nsRoot); err != nil {
|
||||
res.DBsReplayed = n // partial credit: whatever imported before the failure really did import
|
||||
m.logger.Printf("[ERROR] [backup] DB re-import for %s: %v", stackName, err)
|
||||
if dataErr == nil {
|
||||
dataErr = err
|
||||
}
|
||||
} else {
|
||||
res.DBsReplayed = n
|
||||
}
|
||||
}
|
||||
if err := m.stackProvider.StartStack(stackName); err != nil {
|
||||
return fmt.Errorf("starting %s after restore from unit: %w", stackName, err)
|
||||
return res, fmt.Errorf("starting %s after restore from unit: %w", stackName, err)
|
||||
}
|
||||
if err := m.waitForHealthy(stackName, 90*time.Second); err != nil {
|
||||
m.logger.Printf("[WARN] [backup] %s restored but health check failed: %v", stackName, err)
|
||||
}
|
||||
|
||||
if dataErr != nil {
|
||||
return fmt.Errorf("restore of %s from unit completed with data errors: %w", stackName, dataErr)
|
||||
return res, fmt.Errorf("restore of %s from unit completed with data errors: %w", stackName, dataErr)
|
||||
}
|
||||
m.logger.Printf("[INFO] [backup] Restore-from-unit completed: %s", stackName)
|
||||
return nil
|
||||
m.logger.Printf("[INFO] [backup] Restore-from-unit completed: %s — %d volume(s) of %d listed, %d database(s) of %d listed",
|
||||
stackName, res.VolumesReplayed, res.ManifestVolumes, res.DBsReplayed, res.ManifestDBs)
|
||||
return res, nil
|
||||
}
|
||||
|
||||
@@ -78,7 +78,7 @@ func TestRestoreFromRecoveryUnitWithGuestAbsent(t *testing.T) {
|
||||
m := &Manager{logger: log.New(io.Discard, "", 0),
|
||||
systemDataPath: filepath.Join(drive, "..", "sys"), stackProvider: fake}
|
||||
|
||||
if err := m.RestoreFromRecoveryUnit("app"); err != nil {
|
||||
if _, err := m.RestoreFromRecoveryUnit("app"); err != nil {
|
||||
t.Fatalf("restore must succeed from the drive alone, got: %v", err)
|
||||
}
|
||||
if fake.gotEnv == nil {
|
||||
@@ -105,7 +105,7 @@ func TestRestoreFromRecoveryUnitGuestAbsentStillFailsClosed(t *testing.T) {
|
||||
m := &Manager{logger: log.New(io.Discard, "", 0),
|
||||
systemDataPath: filepath.Join(drive, "..", "sys"), stackProvider: fake}
|
||||
|
||||
err := m.RestoreFromRecoveryUnit("app")
|
||||
_, err := m.RestoreFromRecoveryUnit("app")
|
||||
if err == nil {
|
||||
t.Fatal("expected fail-closed refusal when the data key is in neither source")
|
||||
}
|
||||
@@ -148,7 +148,7 @@ func TestRestoreFromRecoveryUnitOrchestration(t *testing.T) {
|
||||
}
|
||||
m := &Manager{logger: log.New(io.Discard, "", 0), systemDataPath: filepath.Join(drive, "..", "sys"),
|
||||
stackProvider: fake}
|
||||
if err := m.RestoreFromRecoveryUnit("app"); err != nil {
|
||||
if _, err := m.RestoreFromRecoveryUnit("app"); err != nil {
|
||||
t.Fatalf("restore: %v", err)
|
||||
}
|
||||
if fake.gotEnv == nil {
|
||||
@@ -169,7 +169,7 @@ func TestRestoreFromRecoveryUnitOrchestration(t *testing.T) {
|
||||
secrets: map[string]string{"DB_PASSWORD": "pw", "SECRET_KEY": "deadbeef"}}
|
||||
m := &Manager{logger: log.New(io.Discard, "", 0),
|
||||
systemDataPath: filepath.Join(drive, "..", "sys"), stackProvider: fake}
|
||||
if err := m.RestoreFromRecoveryUnit("app"); err != nil {
|
||||
if _, err := m.RestoreFromRecoveryUnit("app"); err != nil {
|
||||
t.Fatalf("a pre-D5 unit must still restore: %v", err)
|
||||
}
|
||||
if fake.gotEnv["SECRET_KEY"] != "deadbeef" {
|
||||
@@ -185,7 +185,7 @@ func TestRestoreFromRecoveryUnitOrchestration(t *testing.T) {
|
||||
}
|
||||
m := &Manager{logger: log.New(io.Discard, "", 0), systemDataPath: filepath.Join(drive, "..", "sys"),
|
||||
stackProvider: fake}
|
||||
err := m.RestoreFromRecoveryUnit("app")
|
||||
_, err := m.RestoreFromRecoveryUnit("app")
|
||||
if err == nil {
|
||||
t.Fatal("expected fail-closed refusal, got nil")
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user