v0.218.0: attribute a DB container by its compose project, and replay volumes on the off-site restore
gates / gates (push) Successful in 11s
gates / gates (push) Successful in 11s
R-355 (first, because it is the only one where data can be lost for good). paperless-ngx's PostgreSQL was dumped into backups/primary/paperless/db-dumps/ — a directory for a stack that does not exist, on the system drive — while the app's own unit recorded db_dumps: null. The same misattribution reached writeSafetyDump, so a destructive restore of that app took NO undo copy and the fail-closed refusal was never reached. Fixed by reading the compose project label, which is the stack name by construction (compose runs with cmd.Dir set to the stack dir and no -p). The old derivation stays as the fallback and an unresolvable attribution is now loud. Catalogue sweep, proven able to convict: one affected app of 53. The fix is in the controller, not the catalogue. R-354. ReconstituteFromOffsite skipped every unit placement and the volume archives live inside the unit, so the off-site restore had no volume leg at all — proven live with planted files: calibre-web's 1,422,848-byte config archive was in the unit, the snapshot and the checking folder, and the restore reported success without it. For the 40 of 53 apps that declare no data drive that archive is the whole dataset. restoreDockerVolumesFrom is the local path's own replay with an explicit directory: ONE implementation, two callers. Volumes replay before the database and inside the stopped window. VolumesReplayed reaches the message. The comment beside the skip was half false and is corrected; the half that still holds — the live unit is the local path's source — is named, and scenario D fingerprints the whole live unit across the operation. Seven red-proofs, each asserted applied and reverted. Two found defects in the tests, not the code: scenario D passed with the unit guard removed because the fingerprint had been narrowed and was blind to the unit root.
This commit is contained in:
@@ -62,14 +62,20 @@ const preRestoreDumpPrefix = "pre-restore-"
|
||||
// OffsiteReconstituteResult reports what a reconstitution actually did, so the flash can state an
|
||||
// OUTCOME instead of a mechanism. Every field here exists because the v0.147 flash could not say it.
|
||||
type OffsiteReconstituteResult struct {
|
||||
SnapshotID string
|
||||
FilesPlaced int
|
||||
DBsReplayed int
|
||||
SafetyDump string // path of the pre-restore dump (the undo), "" when the app has no DB
|
||||
DumpsAt time.Time // when the snapshot's DB half was taken (zero = unknown/legacy unit)
|
||||
OffsiteRunID string // "" for a pre-v0.148 snapshot — an unverified pair
|
||||
Skewed bool // the snapshot carries no coherence stamp: files and DB may differ in age
|
||||
LooksEmpty bool // R-44 sniff on the dump about to be replayed
|
||||
SnapshotID string
|
||||
FilesPlaced int
|
||||
DBsReplayed int
|
||||
// VolumesReplayed (R-354) is how many named-volume archives came back from the snapshot. It is on
|
||||
// the result for the same reason every other field here is: so the OUTCOME can state what
|
||||
// happened rather than a mechanism. Without it the message said "5 fájl visszaállítva" over a
|
||||
// restore that had silently dropped a 1.4 MB volume archive — a true sentence leaving a false
|
||||
// impression, which is the shape this surface keeps having removed from it.
|
||||
VolumesReplayed int
|
||||
SafetyDump string // path of the pre-restore dump (the undo), "" when the app has no DB
|
||||
DumpsAt time.Time // when the snapshot's DB half was taken (zero = unknown/legacy unit)
|
||||
OffsiteRunID string // "" for a pre-v0.148 snapshot — an unverified pair
|
||||
Skewed bool // the snapshot carries no coherence stamp: files and DB may differ in age
|
||||
LooksEmpty bool // R-44 sniff on the dump about to be replayed
|
||||
// Placement (R-351) is what the backup recorded about where this app's data lived, compared
|
||||
// against where this restore actually wrote. Carried on the RESULT and not only on the refusal,
|
||||
// so a restore that proceeded into a different destination says so in its own outcome rather
|
||||
@@ -344,8 +350,19 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string, ack
|
||||
for _, pl := range placements {
|
||||
if pl.isUnit {
|
||||
// The live recovery unit is still never overwritten — it is the LOCAL restore path's
|
||||
// source and clobbering it would trade one recovery route for another. The snapshot's
|
||||
// dump is replayed from the scratch unit instead, so nothing is lost by skipping it.
|
||||
// source and clobbering it would trade one recovery route for another. THAT reason is
|
||||
// sound and still holds; it is why this skip stays.
|
||||
//
|
||||
// R-354 — THE SECOND HALF OF THIS COMMENT USED TO BE FALSE AND IS CORRECTED HERE. It said
|
||||
// "the snapshot's dump is replayed from the scratch unit instead, so nothing is lost by
|
||||
// skipping it". That was true of the DATABASE dump and false of the VOLUME archives, which
|
||||
// live in the same unit and were replayed by nothing at all. Skipping the placement is
|
||||
// correct; treating the skip as harmless was not. Measured live 2026-08-21: calibre-web's
|
||||
// 1 422 848-byte `calibre_web_config.tar` was in the unit, in the snapshot and in the
|
||||
// verification folder, and the restore reported "5 fájl visszaállítva" without it.
|
||||
//
|
||||
// Both legs are now replayed FROM THE SCRATCH UNIT below — volumes first, then the DB, so
|
||||
// the logical dump still wins over any volume-tar copy of the same database.
|
||||
continue
|
||||
}
|
||||
n, cErr := copier(pl.src, pl.dst)
|
||||
@@ -360,6 +377,30 @@ func (m *Manager) ReconstituteFromOffsite(ctx context.Context, stack string, ack
|
||||
res.FilesPlaced += n
|
||||
}
|
||||
|
||||
// --- NAMED VOLUMES (R-354) ------------------------------------------------------------------
|
||||
// Replayed from the SCRATCH unit, exactly as the database dump is, and for the same reason: the
|
||||
// live unit is never overwritten by a placement, so the snapshot's copy exists only under the
|
||||
// scratch. Same helper as the local restore path — one implementation, two callers.
|
||||
//
|
||||
// ORDER IS LOAD-BEARING and mirrors RestoreFromRecoveryUnit: volumes FIRST, database after, so a
|
||||
// logical .sql dump still wins over whatever copy of the same database a volume tar happens to
|
||||
// contain. It also has to happen inside the stopped window, because replacing a named volume means
|
||||
// removing it, and Docker refuses that while a container holds it.
|
||||
volReplay := m.volumeReplayFrom
|
||||
if volReplay == nil {
|
||||
volReplay = m.restoreDockerVolumesFrom
|
||||
}
|
||||
nVols, vErr := volReplay(stack, filepath.Join(scratchUnit, "volume-dumps"))
|
||||
res.VolumesReplayed = nVols
|
||||
if vErr != nil {
|
||||
// A partial replay must never read as a completion. Bring the app back up rather than leaving
|
||||
// an outage, then surface it — the same shape the file leg above uses.
|
||||
if sErr := restartStack(); sErr != nil {
|
||||
m.logger.Printf("[WARN] [offbox] %s: restart after failed volume replay also failed: %v", stack, sErr)
|
||||
}
|
||||
return res, fmt.Errorf("a(z) %s adatkötetének visszaállítása sikertelen: %w", stack, vErr)
|
||||
}
|
||||
|
||||
// --- DATABASE -------------------------------------------------------------------------------
|
||||
// The DB container must be UP for the replay (ImportDump talks to it with its own discovered
|
||||
// credentials), but NOTHING ELSE may be — R-47. Until v0.153.0 this was a full StartStack, which
|
||||
|
||||
Reference in New Issue
Block a user