v0.148.0 — coherent snapshot pairs + an offsite restore that actually restores (R-43 + R-44)
Closes the two findings from DIAG-immich-restore-2026-07-19. Viktor deleted 11
immich photos to test offsite restore; both runs flashed success and the photos
stayed gone. Two independent defects.
R-43 — no offsite path could restore a database. All three buttons were
file-only: the two "visszaállítás" actions staged to a scratch folder and never
touched postgres, and place-to-live merged only MISSING files. For a DB-indexed
app the bytes returned and the app still could not see them. The dump was
carried INTO every snapshot and could never be replayed OUT of one.
New ReconstituteFromOffsite (/backup/offbox/reconstitute): safety dump → stop →
files overwritten to the snapshot version → start → the snapshot's own dump
replayed → health wait. Two invariants:
- nothing is ever deleted (-a, no --ignore-existing, no --delete): a file
created after the snapshot survives as an extra;
- the undo exists before the act — the pre-restore- dump is verified ON DISK
before anything is stopped, overwritten or replayed; if it cannot be taken
the operation refuses with zero changes.
The replay reads the SCRATCH unit: the live unit is never overwritten, so
replaying from it would replay the current DB over itself and restore nothing.
R-44 — a manual push shipped an unrefreshed dump (up to ~24h old). That day's
predated the customer's account by four hours and probed to asset:0/user:0/
album:0 inside 52MB whose bulk was immich's shipped geodata. Every run, manual
AND nightly, now refreshes dumps + units BEFORE capturing. Order is the
mechanism: the gap can only ADD files the DB does not reference yet, never
remove one it does. Manifests carry offsite_run_id + dumps_at, so coherence is
verifiable at restore time rather than assumed; the periodic refresh carries a
prior stamp forward and never invents one.
Honesty surfaces, all warn-level and none a gate: unstamped (pre-v0.148) pairs
report their skew, ValidateDump gained an EXACT-match accounts-table sniff for
customer-empty dumps, the completion flash states an outcome instead of a
mechanism, and the missing-only button now says what it does NOT do.
11 tests; 5 red-proofs run and reverted. Two of those found real test weaknesses
rather than confirming strength — the first undo mutation was caught by a second
guard, and the first table-matching test did not discriminate between the two
matchers at all. Both tests were rewritten to the cases that separate them.
NOT in scope: R-41's catalog invariant check, nightly cadence, retention, quota
math, tier-2, and v0.147.x progress semantics beyond one added phase line.
Live acceptance (§9) has NOT run: no capability-map flip, customer-restore row
stays MISSING, R-3 stays DRAFT.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01P9Nn14TWGzKoqAJAiVwC2s
This commit is contained in:
@@ -59,6 +59,16 @@ type Manager struct {
|
||||
// offboxPlaceCopier (3a) — the place-to-live missing-only merge seam (nil → rsyncRestoreMissing,
|
||||
// the `-a --ignore-existing` additive copy). Never rsyncMirror (--delete trap).
|
||||
offboxPlaceCopier func(src, dst string) (int, error)
|
||||
// offboxFullPlaceCopier (R-43, v0.148.0) — the FULL-restore overwrite seam (nil →
|
||||
// rsyncRestoreOverwrite: `-a` with NO --ignore-existing and NO --delete). Distinct from
|
||||
// offboxPlaceCopier on purpose: the two have opposite semantics for an existing file.
|
||||
offboxFullPlaceCopier func(src, dst string) (int, error)
|
||||
// safetyDumpFn (R-43) — the pre-restore safety-dump seam (nil → the real DumpOne), so the
|
||||
// "never replay without an undo on disk" refusal is unit-testable without Docker.
|
||||
safetyDumpFn func(ctx context.Context, db DiscoveredDB, dumpDir string) DumpResult
|
||||
// offsitePreDumpFn (R-44) — the offsite dump pre-phase seam (nil → runDBDumpsInternal), so the
|
||||
// dumps-strictly-before-capture ordering is observable in a test without Docker or restic.
|
||||
offsitePreDumpFn func(ctx context.Context) error
|
||||
// offboxFreeFn (3a) — the free-space probe for the restore headroom gate, overridable in tests (the
|
||||
// Windows `go test` host has no `df`). Nil → the real diskFreeBytes (df --output=avail).
|
||||
offboxFreeFn func(path string) int64
|
||||
@@ -126,6 +136,13 @@ type Manager struct {
|
||||
lastDBDump *DBDumpStatus
|
||||
running bool
|
||||
|
||||
// R-43/R-44 (v0.148.0) — the coherence stamp of the offsite run in flight, read by
|
||||
// CaptureRecoveryUnit so each unit records WHICH run took the dumps sitting beside its files.
|
||||
// Set for the duration of the dump pre-phase + capture, cleared after; "" means "no offsite run
|
||||
// is establishing coherence right now" (the periodic refresh and the local 02:30 dump leg).
|
||||
offsiteRunID string
|
||||
offsiteRunDumpAt string
|
||||
|
||||
// Restore op-status (Part B, opstatus.go) — display-only async-restore progress, under `mu`.
|
||||
opRunning bool
|
||||
opName string
|
||||
@@ -310,6 +327,28 @@ func (m *Manager) RunDBDumps(ctx context.Context) error {
|
||||
return m.runDBDumpsInternal(ctx)
|
||||
}
|
||||
|
||||
// offsiteRunStamp returns the in-flight offsite run's coherence stamp ("" when none).
|
||||
func (m *Manager) offsiteRunStamp() (runID, dumpsAt string) {
|
||||
m.mu.Lock()
|
||||
defer m.mu.Unlock()
|
||||
return m.offsiteRunID, m.offsiteRunDumpAt
|
||||
}
|
||||
|
||||
// beginOffsiteRunStamp marks the start of an offsite run's coherence window and returns the cleanup.
|
||||
// The stamp is what CaptureRecoveryUnit writes into each unit manifest, so it must be live across
|
||||
// BOTH the dump leg and the unit capture that follows it — those two together are the pair.
|
||||
func (m *Manager) beginOffsiteRunStamp(runID string) func() {
|
||||
m.mu.Lock()
|
||||
m.offsiteRunID = runID
|
||||
m.offsiteRunDumpAt = time.Now().UTC().Format(time.RFC3339)
|
||||
m.mu.Unlock()
|
||||
return func() {
|
||||
m.mu.Lock()
|
||||
m.offsiteRunID, m.offsiteRunDumpAt = "", ""
|
||||
m.mu.Unlock()
|
||||
}
|
||||
}
|
||||
|
||||
// runDBDumpsInternal is the implementation of RunDBDumps. Caller must hold the running flag.
|
||||
func (m *Manager) runDBDumpsInternal(ctx context.Context) error {
|
||||
start := time.Now()
|
||||
|
||||
Reference in New Issue
Block a user