v0.218.0: attribute a DB container by its compose project, and replay volumes on the off-site restore
gates / gates (push) Successful in 11s

R-355 (first, because it is the only one where data can be lost for good). paperless-ngx's
PostgreSQL was dumped into backups/primary/paperless/db-dumps/ — a directory for a stack that
does not exist, on the system drive — while the app's own unit recorded db_dumps: null. The
same misattribution reached writeSafetyDump, so a destructive restore of that app took NO undo
copy and the fail-closed refusal was never reached. Fixed by reading the compose project label,
which is the stack name by construction (compose runs with cmd.Dir set to the stack dir and no
-p). The old derivation stays as the fallback and an unresolvable attribution is now loud.
Catalogue sweep, proven able to convict: one affected app of 53. The fix is in the controller,
not the catalogue.

R-354. ReconstituteFromOffsite skipped every unit placement and the volume archives live inside
the unit, so the off-site restore had no volume leg at all — proven live with planted files:
calibre-web's 1,422,848-byte config archive was in the unit, the snapshot and the checking
folder, and the restore reported success without it. For the 40 of 53 apps that declare no data
drive that archive is the whole dataset. restoreDockerVolumesFrom is the local path's own replay
with an explicit directory: ONE implementation, two callers. Volumes replay before the database
and inside the stopped window. VolumesReplayed reaches the message.

The comment beside the skip was half false and is corrected; the half that still holds — the
live unit is the local path's source — is named, and scenario D fingerprints the whole live unit
across the operation.

Seven red-proofs, each asserted applied and reverted. Two found defects in the tests, not the
code: scenario D passed with the unit guard removed because the fingerprint had been narrowed
and was blind to the unit root.
This commit is contained in:
2026-08-22 09:43:22 +02:00
parent f94543ee5c
commit 5ce3a44645
10 changed files with 871 additions and 24 deletions
+28 -7
View File
@@ -98,15 +98,35 @@ func (m *Manager) RestoreApp(stackName, snapshotID string) error {
return nil
}
// restoreDockerVolumes populates Docker volumes from tar files in the volume dump directory.
// restoreDockerVolumes populates Docker volumes from the tars in the app's LIVE recovery unit.
func (m *Manager) restoreDockerVolumes(stackName, drivePath string) error {
dumpDir := AppVolumeDumpPath(m.namespaceRoot(drivePath), stackName)
_, err := m.restoreDockerVolumesFrom(stackName, AppVolumeDumpPath(m.namespaceRoot(drivePath), stackName))
return err
}
// restoreDockerVolumesFrom is restoreDockerVolumes with an EXPLICIT dump directory, and it returns how
// many volumes it replayed.
//
// R-354. The off-site reconstitution needs exactly this, for the same reason reimportDBDumpsFrom
// exists beside reimportDBDumps: the snapshot's archives live under the restored SCRATCH unit, because
// the live unit is deliberately never overwritten by a placement. Until now no such variant existed,
// so the off-site path had no way to replay a volume and simply did not — the tar sat in the unit, in
// the snapshot and in the verification folder, and the restore reported success without it. For an app
// whose data is entirely in a named volume — 40 of the 53 in the catalogue — that is everything the
// customer owns.
//
// ONE implementation, two callers. A second copy of this loop is what produced the divergence in the
// first place: the local path replayed volumes and the off-site path did not, and nothing compared the
// two.
//
// It only ever READS dumpDir; the recovery unit is never written to here, on either path.
func (m *Manager) restoreDockerVolumesFrom(stackName, dumpDir string) (int, error) {
entries, err := os.ReadDir(dumpDir)
if err != nil {
if os.IsNotExist(err) {
return nil // No volume dumps to restore
return 0, nil // No volume dumps to restore
}
return fmt.Errorf("reading volume dump dir: %w", err)
return 0, fmt.Errorf("reading volume dump dir: %w", err)
}
var restored int
@@ -154,11 +174,12 @@ func (m *Manager) restoreDockerVolumes(stackName, drivePath string) error {
m.logger.Printf("[INFO] [backup] Restored %d Docker volume(s) for %s", restored, stackName)
}
// F17: a per-volume failure used to be a swallowed WARN; surface it so the restore is reported as
// failed rather than silently partial.
// failed rather than silently partial. The count is returned ALONGSIDE the error, not instead of
// it: a caller that replayed three of four volumes needs both numbers to say what happened.
if len(failed) > 0 {
return fmt.Errorf("failed to restore %d volume(s): %v", len(failed), failed)
return restored, fmt.Errorf("failed to restore %d volume(s): %v", len(failed), failed)
}
return nil
return restored, nil
}
// waitForHealthy waits for a stack to reach running state after restore.