docs: DIAGNOSE immich offsite restore — findings + ROADMAP R-43/R-44

Viktor deleted 11 immich photos to test offsite restore; both place runs
merged 0 files and flashed success. Two-layered root cause:

- A UI delete means TRASH, not deletion — no file left the disk, so there
  was genuinely nothing missing to restore. The test proved nothing.
- A valid test would also have failed: no offsite path loads a database.
  All three buttons are file-only (offbox_restore.go:171-232, :346-421).
  Files return; the library stays empty. The only DB replay is
  RestoreFromRecoveryUnit, reachable solely from the LOCAL restore panel.

Compounding: the manual push never refreshes dumps (offbox_handlers.go:
203-227) — today's unit dump is from 02:30 and probes to asset:0 / user:0 /
album:0, predating both the admin user and the photos. Had a full restore
loaded it, it would have written an empty DB over the live one, destroying
the trashed rows that were the only surviving recovery path.

Answer to "can a customer trust same-day offsite?" for a DB-indexed app:
no — files come back, content does not.

No product code changed. Photos left in trash (recovery not wanted).
Capability map: customer-restore row stays MISSING by finding, not by
absence of attempt; offsite-restore row flagged scope-contested pending
a ruling on whether CAMPAIGN-6D exercised the DB half.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01P9Nn14TWGzKoqAJAiVwC2s
This commit is contained in:
2026-07-19 08:47:34 +02:00
parent 317037f8eb
commit f3d1ba3062
4 changed files with 269 additions and 2 deletions
@@ -0,0 +1,201 @@
# DIAGNOSE — immich offsite restore "succeeds", photos do not reappear (2026-07-19)
> **Class:** diagnosis. No product code was changed. Findings + evidence only.
> **Box:** demo-felhom (felhom-pve), guest 9201, controller 0.146.0, immich v3.0.3.
> **Trigger:** Viktor deliberately deleted photos in the immich UI to **test the offsite restore
> feature**. The photos were never wanted back — the exercise was the point.
---
## TL;DR
The restore did not fail. **It was never invoked on missing data, and even if it had been, it could
not have brought the photos back.**
1. Deleting in the immich UI moves assets to **Lomtár (trash)**. No file leaves the disk. Both
„csak a hiányzó fájlok" runs therefore merged **0 files** — correctly, there was nothing missing.
The success flash was truthful about mechanism and empty as a test result.
2. A *valid* test (trash emptied, files genuinely gone) would also have failed to restore the
photos: **no offsite button loads a database.** The files would return; the library would still
be empty, because immich's truth lives in postgres and nothing replays the dump.
3. The dump that ships is additionally **stale by design** — it comes from the 02:30 local run and
is never refreshed before a manual push. Today's dump contains **0 users, 0 assets**.
**One-line answer to "can a customer trust same-day offsite?"** — For a DB-indexed app: **no.
Files come back; content does not.** The backup half is honest; the restore half cannot
reconstitute the app.
---
## Timeline (guest/DB clock; host is CEST, guest is UTC — CEST shown)
| CEST | Event | Evidence |
|---|---|---|
| 07-18 17:10 | old pre-v3 user tree `dccc13fe…` last written | dir mtime |
| 07-19 02:30 | local `db-dump` run writes `immich-postgres.sql` | mtime `00:30:00 UTC` |
| 07-19 07:56:25 | **admin user `4dcfb81e…` created** (fresh v3 deploy) | `user.createdAt` |
| 07-19 07:57:0609 | Viktor uploads 11 assets | `asset.createdAt` |
| 07-19 07:58:15 | **manual offsite backup** — captures the 11 files | `[offbox] backed up immich (…1 mandatory path(s))` |
| 07-19 08:00:08 / 08:00:15 | Viktor deletes assets in UI → **trash**, 2 batches | `asset.deletedAt` |
| 07-19 08:01:14 | full restore → **scratch folder**, not live | `[offbox] restored immich (360304f8, full=true) → …/backups/offsite-restore/immich` |
| 07-19 08:27:54 | place (missing-only) → **0 files merged** | `[offbox] placed immich from offsite scratch: 0 file(s) merged (missing-only)` |
| 07-19 08:29:09 | full restore → scratch again | same as above |
| 07-19 08:29:57 | place (missing-only) → **0 files merged** | same as above |
The backup at 07:58:15 ran **after** the upload and **before** the delete — so the snapshot's
*file* half is correct and current. Only the DB half is not.
---
## Evidence
### DB state (live, at 08:34 CEST)
```
total | trashed | live
11 | 11 | 0 -- all 11 assets trashed, none live
users: 1 -- 4dcfb81e-073a-49fb-938e-7fffb51dbac4 / admin@felhom.eu
```
**`updatedAt` == `deletedAt` exactly on all 11 rows** — nothing has modified these rows since the
08:00 deletion. Conclusive: **none of today's restore operations touched the database.**
All 11 `originalPath` values resolve to files that exist on disk (11 OK / 0 missing, after mapping
the container path `/usr/src/app/upload``/mnt/felhom-drives/hdd_1/appdata/immich`). The photos
are fully recoverable by an ordinary in-app un-trash; **left in trash at Viktor's instruction.**
### The decisive dump probe
The recovery unit that actually ships:
`/mnt/felhom-drives/hdd_1/backups/primary/immich/db-dumps/immich-postgres.sql`
— 51 954 452 B, mtime **02:30 CEST**.
```
asset: 0 rows
user: 0 rows
album: 0 rows
```
The 52 MB is entirely `geodata_places` / `naturalearth_countries` — immich's shipped reference
data. **A dump that looks substantial by size contains zero customer content.** It was taken at
02:30, four hours before the admin user existed (07:56:25) and five and a half hours before the
photos were uploaded (07:57).
The same is true of the unit's volume dumps: `immich_immich_postgres_data.tar` (323 088 384 B) is
also from 02:30 CEST.
> **Method note / correction.** The first probe in this session read
> `appdata/immich/backups/immich-db-backup-20260719T020000-…sql.gz` — that is *immich's own*
> self-backup, not the controller's. It is independently also empty (`asset: 0`, `user: 0`), but
> the figure that matters for restore is the controller unit dump above. Both were probed; both
> are empty; the conclusion is unchanged.
### Safety consequence
Had a full restore actually loaded that dump into postgres, it would have written an **empty
database over the live one — destroying the trashed rows that are the only remaining recovery
path.** It did not, because (a) full restore only stages to a scratch folder and never touches a
DB, and (b) the place path's `live recovery unit present — not overwriting` guard skipped the unit
entirely. The photos survived by the combination of two behaviours, neither of which is a
deliberate safety feature for this case.
---
## Mechanism (source, controller 0.146.0)
| Question | Answer | Cite |
|---|---|---|
| Does a manual offsite run refresh DB dumps before pushing? | **No.** `offboxRunHandler``RunOffboxBackup` directly; never calls `RunDBDumps` / `captureAllRecoveryUnits`. It ships whatever `.sql` already exists. | `web/offbox_handlers.go:203-227`, `backup/offbox.go:574-759` |
| Where do dumps come from, then? | The separate local `db-dump` daily at **02:30**; offsite runs at **04:15**. A *manual* offsite run at any hour ships dumps up to ~24 h old. | `cmd/controller/main.go:542,611`, `config/config.go:317` |
| What does the recovery unit contain? | compose/, app.yaml, manifest.json — and it **enumerates** existing dump filenames, never creates them (`listFileNames`). | `backup/recovery_unit.go:68-154` (105-106) |
| „csak a hiányzó fájlok" — what does it do? | `rsync -a --ignore-existing`. Files only. **Never touches a DB. Never stops/starts the stack.** | `backup/offbox_restore.go:346-421`, `tier2_restore.go:148-162` |
| „Teljes visszaállítás" — restores to where? | A **scratch/verify folder** (`backups/offsite-restore/<app>`), never live data, and **never loads a dump into postgres**. | `backup/offbox_restore.go:171-232`, `:130-164` |
| Does *any* offsite path replay a DB dump? | **No.** The only code that replays into a live DB is `RestoreFromRecoveryUnit`, reachable solely from the **local** restore panel. | `backup/restore_unit.go:74``web/handlers.go:1108` |
| Is there any dump-freshness / RPO check? | **None anywhere.** Zero `RPO` hits. `offboxUnitTime` is only a two-drive tiebreak. The `DBValidationCache` exists but no offsite/restore path reads it. | `backup/offbox.go:827-837`, `backup.go:364-370` |
### The place-success message
```
"A(z) " + app + " hiányzó fájljai helyreállítva az élő adatok közé." // offbox_handlers.go:354
```
Fixed string, built from the app name only. `PlaceOffsiteRestore` returns `error` alone, so the
merged-file count is **structurally unavailable** to the flash — it exists only in the English log
(`offbox_restore.go:419`). **With 0 files merged the customer still reads an unconditional claim
that files were restored.**
Tier-2 already solves this: `"Nincs hiányzó fájl — minden fájl megvan a helyén."`
(`web/handlers.go:1154`). The offsite path has no such branch — a straight parity gap.
---
## Findings
### 1. [HIGH → R-43] No offsite restore path can restore a database
All three offsite buttons are file-only. The DB dump is faithfully carried *into* the snapshot and
can never be replayed *out of* it by any offsite action. For every DB-indexed app — which is most
of the catalog — offsite restore cannot reconstitute the application. The customer is left with
correct bytes on disk and an app that cannot see them.
This outranks the staleness finding: even a perfectly fresh dump would not have restored the
photos today.
### 2. [HIGH → R-44] Manual offsite push ships an unrefreshed dump
"Backed up now" is false for the DB half. Today: files from 07:57, dump from 02:30 — containing
neither the user nor the assets. A same-day restore silently loses same-day state. No UI surface
discloses the DB's actual age.
### 3. [MED] Success semantics are mechanism-shaped, not outcome-shaped
Every flash today told the truth about what the code did while the customer watched nothing happen.
The zero-file case is the sharpest form: an unconditional "helyreállítva" when nothing was
restored. Feeds the v0.147.0 feedback arc — restore completions should state the **outcome**
(„N fájl visszaállítva; az adatbázis NEM — az alkalmazás tartalma nem változott").
### 4. [MED, new] The "live recovery unit present" guard silently skips the dump
`offbox_restore.go:400-405` skips the **entire** unit — compose, manifest **and db-dumps**
whenever a local `backups/primary/<app>` exists, which is true on any box whose local backup has
run. So on a normal box the offsite dump is never even staged to disk. Correct as
"don't clobber the local unit"; invisible as "your DB was not part of this restore".
### 5. [LOW, new] Orphaned pre-v3 user tree
`upload|thumbs|encoded-video/dccc13fe-d629-4980-98e5-f2ea88cdb790` — 10 files / 36 MB + 3.9 MB +
15 MB, mtime 07-18 17:10, stranded by the immich 2→3 redeploy. The live DB has never known about
it. Dead weight, not today's issue; worth a sweep policy for major redeploys.
### 6. [LOW, unresolved] Storage figure discrepancy
The brief cites **704.6 MiB** used on the library storage. Measured: **126 MB** total
(`upload` 72 M, `encoded-video` 30 M, `backups` 18 M, `thumbs` 8 M). Not reconciled in this
session. If 704.6 MiB came off a controller Storage page, that gap is its own defect and needs a
separate look.
### 7. [Method] A UI delete cannot test restore
Deleting in an app's own UI usually means *trash*, not deletion. Any "restore drill" scripted that
way will merge 0 files, flash success, and prove nothing. A real drill must empty the trash (or
otherwise remove the files) **and** verify the app's content afterwards, not just the file count.
This belongs in R-3's "first restore by the customer" step.
---
## Capability map
The **customer-restore row STAYS MISSING** and gains this diagnosis as a note. The C6 evidence
attempt found a product gap — that is the rehearsal system working, not failing. **R-3 stays
DRAFT.**
---
## What was NOT done
- No product fixes — no changes to restore/backup code, labels, or flashes.
- No `restic prune`/`forget`, no snapshot deletion, no escrow or check-folder changes.
- The 11 assets were **left in trash** (Viktor: recovery not wanted). immich's default 30-day
auto-purge applies; they are recoverable by un-trash until then.
- The genuine destructive test (empty trash → attempt real restore) was **offered and not run**
it needs explicit go-ahead, and the dump probe already settles the question without it.