v0.218.0: attribute a DB container by its compose project, and replay volumes on the off-site restore
gates / gates (push) Successful in 11s
gates / gates (push) Successful in 11s
R-355 (first, because it is the only one where data can be lost for good). paperless-ngx's PostgreSQL was dumped into backups/primary/paperless/db-dumps/ — a directory for a stack that does not exist, on the system drive — while the app's own unit recorded db_dumps: null. The same misattribution reached writeSafetyDump, so a destructive restore of that app took NO undo copy and the fail-closed refusal was never reached. Fixed by reading the compose project label, which is the stack name by construction (compose runs with cmd.Dir set to the stack dir and no -p). The old derivation stays as the fallback and an unresolvable attribution is now loud. Catalogue sweep, proven able to convict: one affected app of 53. The fix is in the controller, not the catalogue. R-354. ReconstituteFromOffsite skipped every unit placement and the volume archives live inside the unit, so the off-site restore had no volume leg at all — proven live with planted files: calibre-web's 1,422,848-byte config archive was in the unit, the snapshot and the checking folder, and the restore reported success without it. For the 40 of 53 apps that declare no data drive that archive is the whole dataset. restoreDockerVolumesFrom is the local path's own replay with an explicit directory: ONE implementation, two callers. Volumes replay before the database and inside the stopped window. VolumesReplayed reaches the message. The comment beside the skip was half false and is corrected; the half that still holds — the live unit is the local path's source — is named, and scenario D fingerprints the whole live unit across the operation. Seven red-proofs, each asserted applied and reverted. Two found defects in the tests, not the code: scenario D passed with the unit guard removed because the fingerprint had been narrowed and was blind to the unit root.
This commit is contained in:
@@ -141,6 +141,10 @@ type Manager struct {
|
||||
// disconnected) can be unit-tested without Docker. Nil → the real DumpAppVolumesSafe.
|
||||
dumpVolumesSafe func(stackName string) error
|
||||
|
||||
// R-354 volume-REPLAY seam — the mirror of the F17 DB seams above, so the off-site path's new
|
||||
// volume leg is unit-testable without Docker. Nil → the real restoreDockerVolumesFrom.
|
||||
volumeReplayFrom func(stackName, dumpDir string) (int, error)
|
||||
|
||||
// F7 tar seam — the ONE docker exec inside DumpAppVolumes, overridable so the atomic-write
|
||||
// behaviour (tmp+fsync+rename; the last good `.tar` survives a mid-write failure) is unit-testable
|
||||
// without Docker. It must write the tar to `<dumpDir>/<volName>.tar.tmp` and return the combined
|
||||
|
||||
@@ -62,14 +62,20 @@ const preRestoreDumpPrefix = "pre-restore-"
|
||||
// OffsiteReconstituteResult reports what a reconstitution actually did, so the flash can state an
|
||||
// OUTCOME instead of a mechanism. Every field here exists because the v0.147 flash could not say it.
|
||||
type OffsiteReconstituteResult struct {
|
||||
SnapshotID string
|
||||
FilesPlaced int
|
||||
DBsReplayed int
|
||||
SafetyDump string // path of the pre-restore dump (the undo), "" when the app has no DB
|
||||
DumpsAt time.Time // when the snapshot's DB half was taken (zero = unknown/legacy unit)
|
||||
OffsiteRunID string // "" for a pre-v0.148 snapshot — an unverified pair
|
||||
Skewed bool // the snapshot carries no coherence stamp: files and DB may differ in age
|
||||
LooksEmpty bool // R-44 sniff on the dump about to be replayed
|
||||
SnapshotID string
|
||||
FilesPlaced int
|
||||
DBsReplayed int
|
||||
// VolumesReplayed (R-354) is how many named-volume archives came back from the snapshot. It is on
|
||||
// the result for the same reason every other field here is: so the OUTCOME can state what
|
||||
// happened rather than a mechanism. Without it the message said "5 fájl visszaállítva" over a
|
||||
// restore that had silently dropped a 1.4 MB volume archive — a true sentence leaving a false
|
||||
// impression, which is the shape this surface keeps having removed from it.
|
||||
VolumesReplayed int
|
||||
SafetyDump string // path of the pre-restore dump (the undo), "" when the app has no DB
|
||||
DumpsAt time.Time // when the snapshot's DB half was taken (zero = unknown/legacy unit)
|
||||
OffsiteRunID string // "" for a pre-v0.148 snapshot — an unverified pair
|
||||
Skewed bool // the snapshot carries no coherence stamp: files and DB may differ in age
|
||||
LooksEmpty bool // R-44 sniff on the dump about to be replayed
|
||||
// Placement (R-351) is what the backup recorded about where this app's data lived, compared
|
||||
// against where this restore actually wrote. Carried on the RESULT and not only on the refusal,
|
||||
// so a restore that proceeded into a different destination says so in its own outcome rather
|
||||
@@ -344,8 +350,19 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string, ack
|
||||
for _, pl := range placements {
|
||||
if pl.isUnit {
|
||||
// The live recovery unit is still never overwritten — it is the LOCAL restore path's
|
||||
// source and clobbering it would trade one recovery route for another. The snapshot's
|
||||
// dump is replayed from the scratch unit instead, so nothing is lost by skipping it.
|
||||
// source and clobbering it would trade one recovery route for another. THAT reason is
|
||||
// sound and still holds; it is why this skip stays.
|
||||
//
|
||||
// R-354 — THE SECOND HALF OF THIS COMMENT USED TO BE FALSE AND IS CORRECTED HERE. It said
|
||||
// "the snapshot's dump is replayed from the scratch unit instead, so nothing is lost by
|
||||
// skipping it". That was true of the DATABASE dump and false of the VOLUME archives, which
|
||||
// live in the same unit and were replayed by nothing at all. Skipping the placement is
|
||||
// correct; treating the skip as harmless was not. Measured live 2026-08-21: calibre-web's
|
||||
// 1 422 848-byte `calibre_web_config.tar` was in the unit, in the snapshot and in the
|
||||
// verification folder, and the restore reported "5 fájl visszaállítva" without it.
|
||||
//
|
||||
// Both legs are now replayed FROM THE SCRATCH UNIT below — volumes first, then the DB, so
|
||||
// the logical dump still wins over any volume-tar copy of the same database.
|
||||
continue
|
||||
}
|
||||
n, cErr := copier(pl.src, pl.dst)
|
||||
@@ -360,6 +377,30 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string, ack
|
||||
res.FilesPlaced += n
|
||||
}
|
||||
|
||||
// --- NAMED VOLUMES (R-354) ------------------------------------------------------------------
|
||||
// Replayed from the SCRATCH unit, exactly as the database dump is, and for the same reason: the
|
||||
// live unit is never overwritten by a placement, so the snapshot's copy exists only under the
|
||||
// scratch. Same helper as the local restore path — one implementation, two callers.
|
||||
//
|
||||
// ORDER IS LOAD-BEARING and mirrors RestoreFromRecoveryUnit: volumes FIRST, database after, so a
|
||||
// logical .sql dump still wins over whatever copy of the same database a volume tar happens to
|
||||
// contain. It also has to happen inside the stopped window, because replacing a named volume means
|
||||
// removing it, and Docker refuses that while a container holds it.
|
||||
volReplay := m.volumeReplayFrom
|
||||
if volReplay == nil {
|
||||
volReplay = m.restoreDockerVolumesFrom
|
||||
}
|
||||
nVols, vErr := volReplay(stack, filepath.Join(scratchUnit, "volume-dumps"))
|
||||
res.VolumesReplayed = nVols
|
||||
if vErr != nil {
|
||||
// A partial replay must never read as a completion. Bring the app back up rather than leaving
|
||||
// an outage, then surface it — the same shape the file leg above uses.
|
||||
if sErr := restartStack(); sErr != nil {
|
||||
m.logger.Printf("[WARN] [offbox] %s: restart after failed volume replay also failed: %v", stack, sErr)
|
||||
}
|
||||
return res, fmt.Errorf("a(z) %s adatkötetének visszaállítása sikertelen: %w", stack, vErr)
|
||||
}
|
||||
|
||||
// --- DATABASE -------------------------------------------------------------------------------
|
||||
// The DB container must be UP for the replay (ImportDump talks to it with its own discovered
|
||||
// credentials), but NOTHING ELSE may be — R-47. Until v0.153.0 this was a full StartStack, which
|
||||
|
||||
@@ -0,0 +1,216 @@
|
||||
package backup
|
||||
|
||||
import (
|
||||
"context"
|
||||
"crypto/sha256"
|
||||
"encoding/hex"
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"sort"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// R-354. The off-site reconstitution replayed the snapshot's DATABASE and never its named VOLUMES,
|
||||
// because the volume archives live inside the recovery unit and the unit placement is (correctly)
|
||||
// skipped. Measured live 2026-08-21: calibre-web's 1 422 848-byte `calibre_web_config.tar` was in the
|
||||
// unit, in the off-site snapshot and in the verification folder, and the restore returned five files,
|
||||
// reported success, and did not replay it. For the 40 of 53 catalogue apps that declare no data drive,
|
||||
// that archive is the entire dataset.
|
||||
|
||||
// seedScratchVolumes writes volume tars into the scratch unit the reconstitution will read from, and
|
||||
// returns that directory.
|
||||
func seedScratchVolumes(t *testing.T, m *Manager, stack string, names ...string) string {
|
||||
t.Helper()
|
||||
scratch, _, err := m.offboxRestoreScratchDir(stack)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
unit := findScratchUnitDir(scratch, stack)
|
||||
if unit == "" {
|
||||
t.Fatalf("no scratch unit dir for %s under %s", stack, scratch)
|
||||
}
|
||||
dir := filepath.Join(unit, "volume-dumps")
|
||||
if err := os.MkdirAll(dir, 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
for _, n := range names {
|
||||
if err := os.WriteFile(filepath.Join(dir, n), []byte("tar:"+n), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
return dir
|
||||
}
|
||||
|
||||
func fingerprintTree(t *testing.T, root string) string {
|
||||
t.Helper()
|
||||
var lines []string
|
||||
_ = filepath.Walk(root, func(p string, fi os.FileInfo, err error) error {
|
||||
if err != nil || fi.IsDir() {
|
||||
return nil
|
||||
}
|
||||
b, rErr := os.ReadFile(p)
|
||||
if rErr != nil {
|
||||
return nil
|
||||
}
|
||||
// The pre-restore undo copies are the ONE documented write into the live unit; everything
|
||||
// else in the tree must be byte-identical across the operation.
|
||||
if strings.HasPrefix(filepath.Base(p), preRestoreDumpPrefix) {
|
||||
return nil
|
||||
}
|
||||
sum := sha256.Sum256(b)
|
||||
rel, _ := filepath.Rel(root, p)
|
||||
lines = append(lines, rel+":"+hex.EncodeToString(sum[:]))
|
||||
return nil
|
||||
})
|
||||
sort.Strings(lines)
|
||||
return strings.Join(lines, "\n")
|
||||
}
|
||||
|
||||
// TestR354_ScenarioA_VolumeOnlyAppGetsItsVolumeBack is the case that matters: an app whose data is
|
||||
// entirely in a named volume. Before the fix this returned nothing and said it had succeeded.
|
||||
func TestR354_ScenarioA_VolumeOnlyAppGetsItsVolumeBack(t *testing.T) {
|
||||
m, _, _ := reconFixture(t, "20260719T060000Z", "2026-07-19T06:00:00Z", "")
|
||||
// A volume-only app places no files at all — the shape that reported "0 fájl" and success.
|
||||
m.SetOffboxFullPlaceCopier(func(_, _ string) (int, error) { return 0, nil })
|
||||
m.discoverDBs = func(context.Context) ([]DiscoveredDB, error) { return nil, nil }
|
||||
wantDir := seedScratchVolumes(t, m, "immich", "immich_immich_data.tar")
|
||||
|
||||
var gotDir string
|
||||
var gotStack string
|
||||
m.volumeReplayFrom = func(stack, dumpDir string) (int, error) {
|
||||
gotStack, gotDir = stack, dumpDir
|
||||
return 1, nil
|
||||
}
|
||||
|
||||
res, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
|
||||
if err != nil {
|
||||
t.Fatalf("ReconstituteFromOffsite: %v", err)
|
||||
}
|
||||
if res.VolumesReplayed != 1 {
|
||||
t.Errorf("VolumesReplayed = %d, want 1 — the snapshot's volume did not come back", res.VolumesReplayed)
|
||||
}
|
||||
if gotStack != "immich" {
|
||||
t.Errorf("replayed for stack %q, want %q", gotStack, "immich")
|
||||
}
|
||||
// It must read the SCRATCH unit, never the live one.
|
||||
if gotDir != wantDir {
|
||||
t.Errorf("volume replay read %q, want the scratch unit's %q", gotDir, wantDir)
|
||||
}
|
||||
}
|
||||
|
||||
// TestR354_ScenarioB_BothLegsReturnAndAreCounted — declared files AND a volume.
|
||||
func TestR354_ScenarioB_BothLegsReturn(t *testing.T) {
|
||||
m, _, _ := reconFixture(t, "20260719T060000Z", "2026-07-19T06:00:00Z", pgDump(1))
|
||||
seedScratchVolumes(t, m, "immich", "immich_a.tar", "immich_b.tar")
|
||||
m.volumeReplayFrom = func(_, _ string) (int, error) { return 2, nil }
|
||||
|
||||
res, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
|
||||
if err != nil {
|
||||
t.Fatalf("ReconstituteFromOffsite: %v", err)
|
||||
}
|
||||
if res.FilesPlaced != 3 {
|
||||
t.Errorf("FilesPlaced = %d, want 3", res.FilesPlaced)
|
||||
}
|
||||
if res.VolumesReplayed != 2 {
|
||||
t.Errorf("VolumesReplayed = %d, want 2", res.VolumesReplayed)
|
||||
}
|
||||
if res.DBsReplayed != 1 {
|
||||
t.Errorf("DBsReplayed = %d, want 1", res.DBsReplayed)
|
||||
}
|
||||
}
|
||||
|
||||
// TestR354_ScenarioC_NoVolumesIsUnchanged — a snapshot with no volume archives must behave exactly as
|
||||
// before. The real helper runs here (no seam), so the absent-directory path is the one under test.
|
||||
func TestR354_ScenarioC_NoVolumeArchivesIsANoOp(t *testing.T) {
|
||||
m, _, _ := reconFixture(t, "20260719T060000Z", "2026-07-19T06:00:00Z", pgDump(1))
|
||||
// deliberately NO seedScratchVolumes and NO seam
|
||||
|
||||
res, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
|
||||
if err != nil {
|
||||
t.Fatalf("a snapshot without volume archives must not fail the restore: %v", err)
|
||||
}
|
||||
if res.VolumesReplayed != 0 {
|
||||
t.Errorf("VolumesReplayed = %d, want 0", res.VolumesReplayed)
|
||||
}
|
||||
}
|
||||
|
||||
// TestR354_ScenarioD_LiveRecoveryUnitIsNeverWritten. The skip that caused R-354 also protects the
|
||||
// local restore path's own source, and that reason still holds. Fingerprint the live unit across the
|
||||
// whole operation and compare — the doctrine's own answer to the R-181 class, where every test
|
||||
// asserted a mechanism inside one function and the tree still moved.
|
||||
func TestR354_ScenarioD_LiveRecoveryUnitIsNeverWritten(t *testing.T) {
|
||||
m, _, _ := reconFixture(t, "20260719T060000Z", "2026-07-19T06:00:00Z", pgDump(1))
|
||||
_, liveNs, err := m.offboxRestoreScratchDir("immich")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
liveUnit := RecoveryUnitPath(liveNs, "immich")
|
||||
liveVols := filepath.Join(liveUnit, "volume-dumps")
|
||||
if err := os.MkdirAll(liveVols, 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
// The live unit's own copy — the local restore path's source. It must survive untouched.
|
||||
if err := os.WriteFile(filepath.Join(liveVols, "immich_immich_data.tar"), []byte("THE LIVE UNIT COPY"), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := os.WriteFile(filepath.Join(liveUnit, "manifest.json"), []byte(`{"app_name":"immich"}`), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
// Fingerprint the WHOLE live unit. Narrowing this to the volume archives made the test blind to a
|
||||
// placement writing into the unit ROOT — caught by red-proof 6, which passed against the narrowed
|
||||
// version. The excluded set is exactly the pre-restore undo copies, and those are asserted below.
|
||||
before := fingerprintTree(t, liveUnit)
|
||||
|
||||
seedScratchVolumes(t, m, "immich", "immich_immich_data.tar")
|
||||
m.volumeReplayFrom = func(_, _ string) (int, error) { return 1, nil }
|
||||
// A copier that really writes, so "the unit is never written to" is observable rather than assumed.
|
||||
m.SetOffboxFullPlaceCopier(func(src, dst string) (int, error) {
|
||||
_ = os.MkdirAll(dst, 0o755)
|
||||
return 1, os.WriteFile(filepath.Join(dst, "PLACED"), []byte("from the copier"), 0o644)
|
||||
})
|
||||
|
||||
if _, err := m.ReconstituteFromOffsite(context.Background(), "immich", false); err != nil {
|
||||
t.Fatalf("ReconstituteFromOffsite: %v", err)
|
||||
}
|
||||
|
||||
if after := fingerprintTree(t, liveUnit); after != before {
|
||||
t.Errorf("the LIVE recovery unit's backup content changed across the restore — the local path's source was clobbered\nbefore:\n%s\nafter:\n%s", before, after)
|
||||
}
|
||||
|
||||
// The ONE write into the live unit that IS expected: the pre-restore safety dump. It lives in the
|
||||
// app's own db-dumps dir deliberately (see preRestoreDumpPrefix) — it is the undo, and an undo the
|
||||
// customer cannot see is not much of one. Asserted here rather than merely excluded, so "the unit
|
||||
// is untouched" cannot quietly come to mean "the undo stopped being written".
|
||||
undo, _ := filepath.Glob(filepath.Join(liveUnit, "db-dumps", "pre-restore-*.sql"))
|
||||
if len(undo) != 1 {
|
||||
t.Errorf("expected exactly one pre-restore undo copy in the live unit, found %d", len(undo))
|
||||
}
|
||||
}
|
||||
|
||||
// TestR354_ScenarioE_PartialReplayIsAFailure — a volume replay that fails must never read as a
|
||||
// completion, and must name what failed.
|
||||
func TestR354_ScenarioE_PartialReplayIsReportedAsFailure(t *testing.T) {
|
||||
m, prov, _ := reconFixture(t, "20260719T060000Z", "2026-07-19T06:00:00Z", pgDump(1))
|
||||
seedScratchVolumes(t, m, "immich", "immich_a.tar", "immich_b.tar")
|
||||
m.volumeReplayFrom = func(_, _ string) (int, error) {
|
||||
return 1, fmt.Errorf("failed to restore 1 volume(s): [immich_b]")
|
||||
}
|
||||
|
||||
res, err := m.ReconstituteFromOffsite(context.Background(), "immich", false)
|
||||
if err == nil {
|
||||
t.Fatal("a partial volume replay must be reported as a failure, not a completion")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "immich_b") {
|
||||
t.Errorf("the failure must name the volume that did not come back; got %q", err.Error())
|
||||
}
|
||||
// Best-effort bring-up: a failed restore must not also be an outage.
|
||||
if !prov.fullStarted {
|
||||
t.Error("the app was left stopped after a failed volume replay")
|
||||
}
|
||||
// The count of what DID come back is still carried, so the report can say "1 of 2".
|
||||
if res.VolumesReplayed != 1 {
|
||||
t.Errorf("VolumesReplayed = %d, want the partial count 1", res.VolumesReplayed)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,125 @@
|
||||
package backup
|
||||
|
||||
import (
|
||||
"context"
|
||||
"io"
|
||||
"log"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// R-355's second half. The naming defect does not stop at the backup: writeSafetyDump filters the
|
||||
// discovered databases with `db.StackName == stackName`, so an app whose database is attributed to the
|
||||
// wrong stack has NO database as far as the destructive restore is concerned. It therefore takes no
|
||||
// undo copy, and the fail-closed refusal that protects every other app cannot fire — the guard is not
|
||||
// bypassed, it is never reached.
|
||||
//
|
||||
// Measured live on 2026-08-21: a destructive restore of `paperless-ngx` ran to completion over a live
|
||||
// 72-table PostgreSQL with `find /mnt -name "pre-restore-*"` empty both before and after.
|
||||
|
||||
func newSafetyTestManager() *Manager {
|
||||
return &Manager{logger: log.New(io.Discard, "", 0)}
|
||||
}
|
||||
|
||||
// TestR355_SafetyDumpIsTakenForTheCorrectlyAttributedApp is scenario B: the undo copy exists.
|
||||
func TestR355_SafetyDumpIsTakenForTheCorrectlyAttributedApp(t *testing.T) {
|
||||
nsRoot := t.TempDir()
|
||||
m := newSafetyTestManager()
|
||||
|
||||
// Post-fix attribution: the compose project label resolves this container to `paperless-ngx`.
|
||||
m.discoverDBs = func(ctx context.Context) ([]DiscoveredDB, error) {
|
||||
return []DiscoveredDB{
|
||||
{StackName: "paperless-ngx", DBType: DBTypePostgres, ContainerName: "paperless-postgres", ContainerID: "cid"},
|
||||
}, nil
|
||||
}
|
||||
m.safetyDumpFn = func(ctx context.Context, db DiscoveredDB, dumpDir string) DumpResult {
|
||||
p := filepath.Join(dumpDir, string(db.StackName)+"-"+string(db.DBType)+".sql")
|
||||
if err := os.MkdirAll(dumpDir, 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := os.WriteFile(p, []byte("-- 72 tables\n"), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return DumpResult{DB: db, FilePath: p, Size: 13}
|
||||
}
|
||||
|
||||
safety, err := m.writeSafetyDump(context.Background(), "paperless-ngx", nsRoot)
|
||||
if err != nil {
|
||||
t.Fatalf("writeSafetyDump: %v", err)
|
||||
}
|
||||
if safety == "" {
|
||||
t.Fatal("no safety dump was taken for an app that HAS a database — the restore would proceed with no undo")
|
||||
}
|
||||
if _, err := os.Stat(safety); err != nil {
|
||||
t.Fatalf("the safety dump path %q is not on disk: %v", safety, err)
|
||||
}
|
||||
if !strings.Contains(filepath.Base(safety), "pre-restore-") {
|
||||
t.Errorf("the undo copy must carry the pre-restore prefix so it can never be replayed as a source; got %q", filepath.Base(safety))
|
||||
}
|
||||
// The consequence that matters: it lives inside THIS app's unit, not a phantom's.
|
||||
if got, want := filepath.Dir(safety), AppDBDumpPath(nsRoot, "paperless-ngx"); got != want {
|
||||
t.Errorf("undo copy written to %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
// TestR355_MisattributedAppGetsNoUndoCopy demonstrates the WRONG OUTCOME — the state the fix removes.
|
||||
// It models the pre-fix attribution (`paperless`) against a restore of `paperless-ngx` and asserts the
|
||||
// undo silently does not happen. This is the shape that made a destructive restore unrecoverable.
|
||||
func TestR355_MisattributedAppGetsNoUndoCopy(t *testing.T) {
|
||||
nsRoot := t.TempDir()
|
||||
m := newSafetyTestManager()
|
||||
|
||||
// PRE-FIX attribution: deriveStackName gave `paperless` for container `paperless-postgres`.
|
||||
m.discoverDBs = func(ctx context.Context) ([]DiscoveredDB, error) {
|
||||
return []DiscoveredDB{
|
||||
{StackName: "paperless", DBType: DBTypePostgres, ContainerName: "paperless-postgres", ContainerID: "cid"},
|
||||
}, nil
|
||||
}
|
||||
called := false
|
||||
m.safetyDumpFn = func(ctx context.Context, db DiscoveredDB, dumpDir string) DumpResult {
|
||||
called = true
|
||||
return DumpResult{DB: db}
|
||||
}
|
||||
|
||||
safety, err := m.writeSafetyDump(context.Background(), "paperless-ngx", nsRoot)
|
||||
if err != nil {
|
||||
t.Fatalf("writeSafetyDump: %v", err)
|
||||
}
|
||||
if safety != "" || called {
|
||||
t.Fatalf("precondition lost: the misattributed shape now takes an undo copy (safety=%q called=%v) — "+
|
||||
"this test documents the defect and must keep failing to find one", safety, called)
|
||||
}
|
||||
// And this is precisely why it was invisible: no error, no dump, and the caller reads
|
||||
// `hasDB == false` — indistinguishable from an app that genuinely has no database.
|
||||
}
|
||||
|
||||
// TestR355_RestoreRefusesWhenTheUndoCannotBeTaken is scenario C, the fail-closed direction. The
|
||||
// invariant already existed and was proven working live on 2026-08-21 for `romm`; this pins it for the
|
||||
// app that could not reach it before, so the two cannot drift apart.
|
||||
func TestR355_RestoreRefusesWhenTheUndoCannotBeTaken(t *testing.T) {
|
||||
nsRoot := t.TempDir()
|
||||
m := newSafetyTestManager()
|
||||
|
||||
m.discoverDBs = func(ctx context.Context) ([]DiscoveredDB, error) {
|
||||
return []DiscoveredDB{
|
||||
{StackName: "paperless-ngx", DBType: DBTypePostgres, ContainerName: "paperless-postgres", ContainerID: "cid"},
|
||||
}, nil
|
||||
}
|
||||
m.safetyDumpFn = func(ctx context.Context, db DiscoveredDB, dumpDir string) DumpResult {
|
||||
return DumpResult{DB: db, Error: os.ErrPermission}
|
||||
}
|
||||
|
||||
safety, err := m.writeSafetyDump(context.Background(), "paperless-ngx", nsRoot)
|
||||
if err == nil {
|
||||
t.Fatal("a database that cannot be dumped must be a hard error — the undo would not exist")
|
||||
}
|
||||
if safety != "" {
|
||||
t.Errorf("a failed undo must return no path, got %q", safety)
|
||||
}
|
||||
// The message must say the restore did not start, because that is the customer's only signal.
|
||||
if !strings.Contains(err.Error(), "nem indult el") {
|
||||
t.Errorf("the refusal must state that the restore did not start; got %q", err.Error())
|
||||
}
|
||||
}
|
||||
@@ -98,15 +98,35 @@ func (m *Manager) RestoreApp(stackName, snapshotID string) error {
|
||||
return nil
|
||||
}
|
||||
|
||||
// restoreDockerVolumes populates Docker volumes from tar files in the volume dump directory.
|
||||
// restoreDockerVolumes populates Docker volumes from the tars in the app's LIVE recovery unit.
|
||||
func (m *Manager) restoreDockerVolumes(stackName, drivePath string) error {
|
||||
dumpDir := AppVolumeDumpPath(m.namespaceRoot(drivePath), stackName)
|
||||
_, err := m.restoreDockerVolumesFrom(stackName, AppVolumeDumpPath(m.namespaceRoot(drivePath), stackName))
|
||||
return err
|
||||
}
|
||||
|
||||
// restoreDockerVolumesFrom is restoreDockerVolumes with an EXPLICIT dump directory, and it returns how
|
||||
// many volumes it replayed.
|
||||
//
|
||||
// R-354. The off-site reconstitution needs exactly this, for the same reason reimportDBDumpsFrom
|
||||
// exists beside reimportDBDumps: the snapshot's archives live under the restored SCRATCH unit, because
|
||||
// the live unit is deliberately never overwritten by a placement. Until now no such variant existed,
|
||||
// so the off-site path had no way to replay a volume and simply did not — the tar sat in the unit, in
|
||||
// the snapshot and in the verification folder, and the restore reported success without it. For an app
|
||||
// whose data is entirely in a named volume — 40 of the 53 in the catalogue — that is everything the
|
||||
// customer owns.
|
||||
//
|
||||
// ONE implementation, two callers. A second copy of this loop is what produced the divergence in the
|
||||
// first place: the local path replayed volumes and the off-site path did not, and nothing compared the
|
||||
// two.
|
||||
//
|
||||
// It only ever READS dumpDir; the recovery unit is never written to here, on either path.
|
||||
func (m *Manager) restoreDockerVolumesFrom(stackName, dumpDir string) (int, error) {
|
||||
entries, err := os.ReadDir(dumpDir)
|
||||
if err != nil {
|
||||
if os.IsNotExist(err) {
|
||||
return nil // No volume dumps to restore
|
||||
return 0, nil // No volume dumps to restore
|
||||
}
|
||||
return fmt.Errorf("reading volume dump dir: %w", err)
|
||||
return 0, fmt.Errorf("reading volume dump dir: %w", err)
|
||||
}
|
||||
|
||||
var restored int
|
||||
@@ -154,11 +174,12 @@ func (m *Manager) restoreDockerVolumes(stackName, drivePath string) error {
|
||||
m.logger.Printf("[INFO] [backup] Restored %d Docker volume(s) for %s", restored, stackName)
|
||||
}
|
||||
// F17: a per-volume failure used to be a swallowed WARN; surface it so the restore is reported as
|
||||
// failed rather than silently partial.
|
||||
// failed rather than silently partial. The count is returned ALONGSIDE the error, not instead of
|
||||
// it: a caller that replayed three of four volumes needs both numbers to say what happened.
|
||||
if len(failed) > 0 {
|
||||
return fmt.Errorf("failed to restore %d volume(s): %v", len(failed), failed)
|
||||
return restored, fmt.Errorf("failed to restore %d volume(s): %v", len(failed), failed)
|
||||
}
|
||||
return nil
|
||||
return restored, nil
|
||||
}
|
||||
|
||||
// waitForHealthy waits for a stack to reach running state after restore.
|
||||
|
||||
Reference in New Issue
Block a user