v0.218.0: attribute a DB container by its compose project, and replay volumes on the off-site restore
gates / gates (push) Successful in 11s

R-355 (first, because it is the only one where data can be lost for good). paperless-ngx's
PostgreSQL was dumped into backups/primary/paperless/db-dumps/ — a directory for a stack that
does not exist, on the system drive — while the app's own unit recorded db_dumps: null. The
same misattribution reached writeSafetyDump, so a destructive restore of that app took NO undo
copy and the fail-closed refusal was never reached. Fixed by reading the compose project label,
which is the stack name by construction (compose runs with cmd.Dir set to the stack dir and no
-p). The old derivation stays as the fallback and an unresolvable attribution is now loud.
Catalogue sweep, proven able to convict: one affected app of 53. The fix is in the controller,
not the catalogue.

R-354. ReconstituteFromOffsite skipped every unit placement and the volume archives live inside
the unit, so the off-site restore had no volume leg at all — proven live with planted files:
calibre-web's 1,422,848-byte config archive was in the unit, the snapshot and the checking
folder, and the restore reported success without it. For the 40 of 53 apps that declare no data
drive that archive is the whole dataset. restoreDockerVolumesFrom is the local path's own replay
with an explicit directory: ONE implementation, two callers. Volumes replay before the database
and inside the stopped window. VolumesReplayed reaches the message.

The comment beside the skip was half false and is corrected; the half that still holds — the
live unit is the local path's source — is named, and scenario D fingerprints the whole live unit
across the operation.

Seven red-proofs, each asserted applied and reverted. Two found defects in the tests, not the
code: scenario D passed with the unit guard removed because the fingerprint had been narrowed
and was blind to the unit root.
This commit is contained in:
2026-08-22 09:43:22 +02:00
parent f94543ee5c
commit 5ce3a44645
10 changed files with 871 additions and 24 deletions
@@ -62,14 +62,20 @@ const preRestoreDumpPrefix = "pre-restore-"
// OffsiteReconstituteResult reports what a reconstitution actually did, so the flash can state an
// OUTCOME instead of a mechanism. Every field here exists because the v0.147 flash could not say it.
type OffsiteReconstituteResult struct {
SnapshotID string
FilesPlaced int
DBsReplayed int
SafetyDump string // path of the pre-restore dump (the undo), "" when the app has no DB
DumpsAt time.Time // when the snapshot's DB half was taken (zero = unknown/legacy unit)
OffsiteRunID string // "" for a pre-v0.148 snapshot — an unverified pair
Skewed bool // the snapshot carries no coherence stamp: files and DB may differ in age
LooksEmpty bool // R-44 sniff on the dump about to be replayed
SnapshotID string
FilesPlaced int
DBsReplayed int
// VolumesReplayed (R-354) is how many named-volume archives came back from the snapshot. It is on
// the result for the same reason every other field here is: so the OUTCOME can state what
// happened rather than a mechanism. Without it the message said "5 fájl visszaállítva" over a
// restore that had silently dropped a 1.4 MB volume archive — a true sentence leaving a false
// impression, which is the shape this surface keeps having removed from it.
VolumesReplayed int
SafetyDump string // path of the pre-restore dump (the undo), "" when the app has no DB
DumpsAt time.Time // when the snapshot's DB half was taken (zero = unknown/legacy unit)
OffsiteRunID string // "" for a pre-v0.148 snapshot — an unverified pair
Skewed bool // the snapshot carries no coherence stamp: files and DB may differ in age
LooksEmpty bool // R-44 sniff on the dump about to be replayed
// Placement (R-351) is what the backup recorded about where this app's data lived, compared
// against where this restore actually wrote. Carried on the RESULT and not only on the refusal,
// so a restore that proceeded into a different destination says so in its own outcome rather
@@ -344,8 +350,19 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string, ack
for _, pl := range placements {
if pl.isUnit {
// The live recovery unit is still never overwritten — it is the LOCAL restore path's
// source and clobbering it would trade one recovery route for another. The snapshot's
// dump is replayed from the scratch unit instead, so nothing is lost by skipping it.
// source and clobbering it would trade one recovery route for another. THAT reason is
// sound and still holds; it is why this skip stays.
//
// R-354 — THE SECOND HALF OF THIS COMMENT USED TO BE FALSE AND IS CORRECTED HERE. It said
// "the snapshot's dump is replayed from the scratch unit instead, so nothing is lost by
// skipping it". That was true of the DATABASE dump and false of the VOLUME archives, which
// live in the same unit and were replayed by nothing at all. Skipping the placement is
// correct; treating the skip as harmless was not. Measured live 2026-08-21: calibre-web's
// 1 422 848-byte `calibre_web_config.tar` was in the unit, in the snapshot and in the
// verification folder, and the restore reported "5 fájl visszaállítva" without it.
//
// Both legs are now replayed FROM THE SCRATCH UNIT below — volumes first, then the DB, so
// the logical dump still wins over any volume-tar copy of the same database.
continue
}
n, cErr := copier(pl.src, pl.dst)
@@ -360,6 +377,30 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string, ack
res.FilesPlaced += n
}
// --- NAMED VOLUMES (R-354) ------------------------------------------------------------------
// Replayed from the SCRATCH unit, exactly as the database dump is, and for the same reason: the
// live unit is never overwritten by a placement, so the snapshot's copy exists only under the
// scratch. Same helper as the local restore path — one implementation, two callers.
//
// ORDER IS LOAD-BEARING and mirrors RestoreFromRecoveryUnit: volumes FIRST, database after, so a
// logical .sql dump still wins over whatever copy of the same database a volume tar happens to
// contain. It also has to happen inside the stopped window, because replacing a named volume means
// removing it, and Docker refuses that while a container holds it.
volReplay := m.volumeReplayFrom
if volReplay == nil {
volReplay = m.restoreDockerVolumesFrom
}
nVols, vErr := volReplay(stack, filepath.Join(scratchUnit, "volume-dumps"))
res.VolumesReplayed = nVols
if vErr != nil {
// A partial replay must never read as a completion. Bring the app back up rather than leaving
// an outage, then surface it — the same shape the file leg above uses.
if sErr := restartStack(); sErr != nil {
m.logger.Printf("[WARN] [offbox] %s: restart after failed volume replay also failed: %v", stack, sErr)
}
return res, fmt.Errorf("a(z) %s adatkötetének visszaállítása sikertelen: %w", stack, vErr)
}
// --- DATABASE -------------------------------------------------------------------------------
// The DB container must be UP for the replay (ImportDump talks to it with its own discovered
// credentials), but NOTHING ELSE may be — R-47. Until v0.153.0 this was a full StartStack, which