R-638 option A: a replay meets the copy's own schema on the two side paths (09 §3 decision 154)

Slice 1 — the no-manifest fallback RestoreApp no longer starts the WHOLE stack
at the current definition before the replay (a newer app could migrate the
restored data underneath it). New order: resolve DB services from the live
compose -> stop -> volumes -> DB-only start (StartStackServices, the same
helper the unit restore uses, R-47) -> replay -> full start -> health wait.
A dump with no identifiable DB service is refused before any mutation (same
gate and message as the unit and off-site paths). A failed volume leg skips
the replay. restoreDockerVolumes now goes through the existing
volumeReplayFrom seam (nil in production) so the order is testable without
Docker.

Slice 2 — a unit restore whose volume leg failed no longer calls the
importer. Everything else on that failure path is unchanged: dataErr is
returned as "completed with data errors", the unit's definition is written
and the app is fully started.

No loader change, no new delete step. Red-proofs in felhom.eu
documentation/audits/design-build-2026-10-06/B/ (red-slice1-fallback-order.txt,
red-slice2-no-replay-after-volume-failure.txt).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-06 15:11:00 +02:00
parent 9d2d62831a
commit 9c945688c0
3 changed files with 342 additions and 18 deletions
+10 -1
View File
@@ -449,7 +449,16 @@ func (m *Manager) RestoreFromRecoveryUnitAtWith(stackName, unitDir string, opt U
// R-47: the replay happens with ONLY the database service up. This used to run after
// RecreateStackFromUnit had already brought the WHOLE stack up, letting the application rebuild
// schema objects underneath the replay (H4, DIAG-immich-restore-round2-2026-07-19).
if hasDumps {
//
// R-638 (option A, 09 §3 decision 154): NOT when the volume leg failed. The replay overlays the copy
// and removes only what the copy knows about, so it is safe only over the copy's OWN database files
// — which a failed volume leg does not guarantee (the volume may still be the live, newer one, or
// half-filled). The failure is already in dataErr and returned below; the definition is still the
// unit's and the app is still started, exactly as before. Pinned by
// TestR638_UnitVolumeFailureNeverCallsTheImporter.
if hasDumps && volErr != nil {
m.logger.Printf("[ERROR] [backup] %s: NOT replaying the database copy — the volume restore failed, so the database is not known to be the copy's own", stackName)
} else if hasDumps {
if err := m.stackProvider.StartStackServices(stackName, dbServices); err != nil {
m.logger.Printf("[ERROR] [backup] DB-only start for %s: %v", stackName, err)
if dataErr == nil {