diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 9f32461..f9f8d7c 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -58,13 +58,13 @@ | Nightly DB dumps (postgres/mariadb autodiscovery), atomic writes | controller v0.118 | **PROVEN-LIVE** | `CAMPAIGN-2` T-BAK-FULL (pg+mariadb autodiscovered); atomicity `CAMPAIGN-6B` P4 + `CAMPAIGN-6E` B1/B2 (SIGKILL mid-write → only `.tar.tmp` touched, last-good byte-unchanged); DB restore `CAMPAIGN-6D` P-FAB | (Cited `CAMPAIGN-3` F7 is the *finding* of non-atomic writes, and T-RST-DB was auth-hollow — corrected to the 6B/6E fix-proofs.) T-6E-1 dir-fsync asymmetry (LOW) → R-10 | | Tier-2 secondary-drive copy: class-driven legs, v2 relpath layout, NAS-target exclusion, safe-remove boundary | controller v0.135 | **PROVEN-LIVE** | `CAMPAIGN-6E-2026-07-15` (P-TIER2 deep-4 PASS), `CAMPAIGN-6C` | | | Offsite (restic → Hetzner Storage Box): mandatory class only, raw-data quota, enlargement gate, retention regrouping | controller v0.134, agent, hub | **PROVEN-LIVE** | `CAMPAIGN-6D-2026-07-15` (mandatory-only P-IMMICH; enlargement gate fired at real 50GiB quota P3-DELIVERY); `VALIDATION-offbox-storagebox-2026-07-09` (byte-perfect round-trip) | Raw-data quota (SP-1) + retention regrouping (SP-2) are `SPIKE-restic-snapshot-shape` **dry-run** verdicts — mechanism validated, not fired in a live product run; only the enlargement gate is live-fired. **Reinstall-continuity (controller v0.142.0, 2026-07-17):** a recreated data volume that orphaned the repo (new passphrase can't open the old keys) is now CLASSIFIED (`wrong password or no key found`) → explicit ORPHANED card + event (not nightly-spam) + a move-aside (never-delete) reset (unclaimed auto / claimed confirm), instead of a raw nightly restic error. Fake-based scenarios + red-proofs. **The live leg FIRED on its own during the 2026-07-18 rehearsal** (`tests/VALIDATION-n100-rehearsal-2026-07-18.md`, S7): after a RESET + re-enable, the first offsite run hit the previous lifecycle's ciphertext and the guard **classified it, pushed `offbox_repo_orphaned`, skipped the run and showed the card (16:58:14)** rather than nightly-spamming a raw restic error; the operator-confirmed reset then **moved the repo aside (never deleted) to `.orphaned-20260718` and re-initialised (16:59:26→16:59:32)**, and the next run produced 2 snapshots / 48.717 MiB. The guard behaved exactly as designed — **the finding is that it had to fire at all** (R-32: RESET destroys custody, so the ciphertext it leaves behind is dead by design and should be purged, while the move-aside guard stays correct for reinstall-WITHOUT-RESET). `DIAGNOSE-offbox-repo-orphaned-2026-07-17` | -| Offsite restore: local-preferred scratch, unit-only default, full two-step, missing-only place-to-live | controller v0.134/134.1/135 | **PROVEN-LIVE** — *scope contested, ruling needed* | `CAMPAIGN-6D` accept legs (immich end-to-end from offsite alone) | **2026-07-19:** `audits/DIAG-immich-restore-2026-07-19.md` finds **no offsite path loads a DB dump** — all three buttons are file-only (R-43). The mechanics in this row's title are each proven; the phrase "**immich end-to-end from offsite alone**" is what is contested, since a DB-indexed app cannot be reconstituted by any offsite action. Status left as-is pending Viktor's read of whether 6D's accept leg actually exercised the DB half or only the file half. **2026-07-19, controller v0.148.0:** the DB half now EXISTS (R-43 reconstitution + R-44 coherent pairs) — but it is **PARTIAL, not PROVEN-LIVE**: shipped and deployed to demo 9201, with no live acceptance run behind it. The flip needs the §9 evidence (upload → push → empty the trash for real → one button → photos back in the timeline) | +| Offsite restore: local-preferred scratch, unit-only default, full two-step, missing-only place-to-live | controller v0.134/134.1/135 | **PARTIAL** — *scope corrected 2026-07-19* | `CAMPAIGN-6D` accept legs (immich end-to-end from offsite alone) | **2026-07-19:** `audits/DIAG-immich-restore-2026-07-19.md` finds **no offsite path loads a DB dump** — all three buttons are file-only (R-43). The mechanics in this row's title are each proven; the phrase "**immich end-to-end from offsite alone**" is what is contested, since a DB-indexed app cannot be reconstituted by any offsite action. **RULED 2026-07-19 (Viktor): 6D's destruction hit the FILE TREE ONLY — the database survived in its named volume** (`immich_postgres_data` is a named volume in both the v2 and v3 template eras), so "immich end-to-end from offsite alone" **overclaimed scope**: the file half was proven, the DB half was never destroyed and therefore never restored. Row downgraded PROVEN-LIVE → **PARTIAL**, scope-corrected. Evidence: `audits/DIAG-immich-restore-2026-07-19.md` (no offsite path could replay a DB at all) + the P-FAB destructive re-import (the proven-replay evidence, on the LOCAL path). **2026-07-19, controller v0.148.0:** the DB half now exists in code (R-43 + R-44) and its replay reached a live box — but **round 2 found it aborts against a running app** (`audits/DIAG-immich-restore-round2-2026-07-19.md`, H4: the replay races immich's own schema repair; `clip_index` recreated by the app 2 s before the dump's CREATE INDEX). Stays PARTIAL until that closes in v0.149 and a clean §9 run exists | | Manual `.fab` export/import: class-scoped capture, browser up/download, tunnel-proof chunking | controller v0.125/128/130/136 | **PROVEN-LIVE** | `CAMPAIGN-6D` P-FAB / Accept #1 (1.7 GB full circle, byte-identical, app boots); chunking `CAMPAIGN-6B` P2 (100 MiB via real CF edge, 120 MiB→413) | Chunking proven at the real CF edge via `curl --resolve`; the **rendered browser file-picker** upload leg is still Viktor's open full-circle test (6C ran it NOT-RUN). C6B-F1 was the 6B *finding*; fix verified in 6D | | Guest-loss DR: PBS restore with full-fidelity layout from archive, restore-test verification | agent v0.75/0.76, PBS | **PROVEN-LIVE** | `CAMPAIGN-2` T-P9-DESTROY-RESTORE (whole-guest `pct restore` of 9201 → running+healthy) + T-PBS-VERIFY (`verify_state: ok`, 13 snapshots); `DRILL-GL6-2026-07-08` Phase 0d (restore-test `mount_parity: ok`) | (Cited `VALIDATION-newbox-restore` is offbox **restic** file-restore, wrong tier — corrected.) Real **offsite** guest-loss round-trip still R1-blocked → S5 DR drill | | PBS-DR secret self-heal on reused-peer re-provision | hub v0.56 | **IMPLEMENTED** | hub v0.56.0 (`pbsdrheal/reconciler.go`, `RestageHostPBSSecret`, all §10 red-proofs); `SPIKE-pbsdr-selfheal-2026-07-15` (root cause) | Reconciler is **scoped to one host** (`PBSDRHEAL_ONLY_HOST`), not fleet-wide; already fired live hands-free on drill qm300 (07-15) — real-customer firing + fleet-wide widening pending | | Crash/power-loss mid-backup/mid-migration → self-heal on next run | controller, agent | **PROVEN-LIVE** | `CAMPAIGN-6D` P5-REST (SIGKILL mid-offbox → auto-restart ~15s, run marked failed not false-success, no stale lock); `CAMPAIGN-6E` B1-B3 | (Cited `CAMPAIGN-2` T-RBT-* legs were empty / auth-hollow — corrected.) Live mid-**migration** crash→self-heal is the weakest sub-claim (P5-REST is mid-backup) | | Soft-quota: usage bar, pre-push enlargement block, customer notification | controller v0.109/134, hub v0.41/55 | **PROVEN-LIVE** | 6D/6E; hub OffsiteChecker | | -| **A customer (not the operator) performs a restore via UI alone** | all | **MISSING** (as evidence) | — | Alpha will produce this; script it into R-3. **2026-07-19:** the C6 evidence attempt ran and found a **product gap instead of evidence** — `audits/DIAG-immich-restore-2026-07-19.md`. A customer-driven UI restore of a DB-indexed app cannot currently succeed (R-43 file-only restore, R-44 stale dump), so this row cannot flip until those close. Row stays MISSING **by finding, not by absence of attempt** — the rehearsal system working, not failing. **2026-07-19: the blocking product gaps are CLOSED in controller v0.148.0** (R-43 + R-44 shipped), so this row is now blocked only on the evidence run itself, not on missing capability. It flips the moment the §9 acceptance produces screenshots + the outcome flash + a snapshot ID. Method note for R-3's script: deleting in an app's own UI usually means *trash*, not deletion, so a drill written that way merges 0 files, flashes success and proves nothing — a real drill must empty the trash **and** verify the app's *content*, not the file count | +| **A customer (not the operator) performs a restore via UI alone** | all | **MISSING** (as evidence) | — | Alpha will produce this; script it into R-3. **2026-07-19:** the C6 evidence attempt ran and found a **product gap instead of evidence** — `audits/DIAG-immich-restore-2026-07-19.md`. A customer-driven UI restore of a DB-indexed app cannot currently succeed (R-43 file-only restore, R-44 stale dump), so this row cannot flip until those close. Row stays MISSING **by finding, not by absence of attempt** — the rehearsal system working, not failing. **2026-07-19: the blocking product gaps are CLOSED in controller v0.148.0** (R-43 + R-44 shipped), so this row is now blocked only on the evidence run itself, not on missing capability. It flips the moment the §9 acceptance produces screenshots + the outcome flash + a snapshot ID. **2026-07-19 round 2 — PARTIAL EVIDENCE ONLY, row NOT flipped** (`audits/DIAG-immich-restore-round2-2026-07-19.md`): a deliberate run from snapshot `49e7cb46` did recover all 11 assets (`status=active`, files resolve), but the operation **reported failure** and left immich reporting schema drift, because the replay aborted against the running app (H4). Photos back ≠ clean acceptance. Blocked on H4 closing in v0.149, then one clean run. Method note for R-3's script: deleting in an app's own UI usually means *trash*, not deletion, so a drill written that way merges 0 files, flashes success and proves nothing — a real drill must empty the trash **and** verify the app's *content*, not the file count | ## D. Storage & devices diff --git a/documentation/audits/DIAG-immich-restore-2026-07-19.md b/documentation/audits/DIAG-immich-restore-2026-07-19.md index d516a63..2466084 100644 --- a/documentation/audits/DIAG-immich-restore-2026-07-19.md +++ b/documentation/audits/DIAG-immich-restore-2026-07-19.md @@ -174,12 +174,16 @@ run. So on a normal box the offsite dump is never even staged to disk. Correct a 15 MB, mtime 07-18 17:10, stranded by the immich 2→3 redeploy. The live DB has never known about it. Dead weight, not today's issue; worth a sweep policy for major redeploys. -### 6. [LOW, unresolved] Storage figure discrepancy +### 6. [CLOSED — not a defect] Storage figure discrepancy The brief cites **704.6 MiB** used on the library storage. Measured: **126 MB** total -(`upload` 72 M, `encoded-video` 30 M, `backups` 18 M, `thumbs` 8 M). Not reconciled in this -session. If 704.6 MiB came off a controller Storage page, that gap is its own defect and needs a -separate look. +(`upload` 72 M, `encoded-video` 30 M, `backups` 18 M, `thumbs` 8 M). + +**RULED 2026-07-19 (Viktor): resolved, no controller defect.** The 704.6 MiB came from **immich's +own Tárhely widget** — immich-internal accounting — not from any controller Storage page. The two +numbers measure different things and were never expected to agree. The same explanation covers the +„650 MiB → 1.4 GiB" figure quoted during round 2, which likewise does not match the controller's +measurement of the tree. Closed; no follow-up item. ### 7. [Method] A UI delete cannot test restore diff --git a/documentation/audits/DIAG-immich-restore-round2-2026-07-19.md b/documentation/audits/DIAG-immich-restore-round2-2026-07-19.md new file mode 100644 index 0000000..f6e3798 --- /dev/null +++ b/documentation/audits/DIAG-immich-restore-round2-2026-07-19.md @@ -0,0 +1,231 @@ +# DIAGNOSE — v0.148.0 full-restore acceptance: files back, timeline empty (round 2, 2026-07-19) + +> **Class:** diagnosis. No product code changed. **Access mode: SSH** (the box was reachable from +> the workstation; every step below is direct evidence, none inferred). +> **Companion:** `DIAG-immich-restore-2026-07-19.md` (round 1, which produced R-43 + R-44). + +--- + +## TL;DR + +**H1 confirmed: the reconstitution never ran.** The operator clicked the old missing-only button; +`/backup/offbox/reconstitute` was **never hit** (`reconstituted lines: 0`, `safety dump lines: 0`, +`replay lines: 0`). Files came back (34 merged — hence the disk growth), the database did not. + +**Then Phase-3 recovery found a NEW defect (H4), which is the more important result.** Running the +real sequence deliberately, the v0.148.0 path executed correctly — safety dump → stop → start → +replay — and the **replay aborted**: `ERROR: relation "clip_index" already exists`. + +Root cause, proven to the second: **the replay races the application's own schema repair.** The +reconstitution starts the stack *before* replaying (ImportDump needs a live container), which gives +immich-server a window to recreate schema objects the dump is about to create. + +``` +10:58:25 controller: replaying DB dump into immich-postgres +10:58:33 immich-server: "Reindexing clip_index" → "Reindexed clip_index" ← app recreates it +10:58:35 controller: ERROR relation "clip_index" already exists — exit status 3 +``` + +**The photos ARE back** (11 assets, `status=active`, `deletedAt` null, all 11 files resolve, owned +by the live user) — because `pg_dump` emits COPY data *before* CREATE INDEX, so the abort landed +after the data. **That success is accidental.** A collision earlier in the script would abort before +the data and leave a genuinely half-restored database, reported identically. + +**This is NOT clean §9 acceptance evidence.** The run reported failure, and immich now reports +schema drift (indexes the aborted script never created). Recovery is PARTIAL. + +--- + +## Phase 1 — facts + +### 1. Endpoint timeline (controller log, guest UTC) + +| UTC | Event | +|---|---| +| 10:23:13 | controller 0.148.0 starts | +| 10:28:51 | manual offsite run starts — **v0.148 dump pre-phase runs** | +| 10:29:56 | `pre-push dump leg completed in 1m5.04s — snapshot pair is coherent` | +| 10:32:52 | `backed up immich (1 mandatory path(s))` | +| 10:33:41 | `backup OK: 3 app(s), 6 snapshot(s), 4m45s` | +| 10:38:32 | `restored immich (49e7cb46, full=true) → …/backups/offsite-restore/immich` (staging) | +| **10:39:55** | **`placed immich from offsite scratch: 34 file(s) merged (missing-only)`** | + +Counted over the same window: `reconstituted: 0 · safety dump: 0 · replaying DB dump: 0 · +missing-only: 1`. The reconstitute endpoint appears nowhere in the log. **H1.** + +### 2. Which snapshot / which dump + +Only two snapshots exist for immich: + +- `6df12205` — 2026-07-18T17:13:46Z +- `49e7cb46` — 2026-07-19T10:30:01Z ← staged, correct + +**No snapshot was pushed after the deletion**, so there is no post-delete empty-state snapshot to +restore by accident. `49e7cb46`'s unit carries `offsite_run_id: 20260719T102851Z` / +`dumps_at: 2026-07-19T10:28:51Z`, and its dump (`immich-postgres.sql`, mtime 10:28:52, 52 393 708 B) +probes to **`asset: 11` / `user: 1`**. + +> **R-44 is working exactly as designed.** Round 1's dump was `asset: 0 / user: 0`. This one holds +> the customer's data and is stamped as coherent with the files beside it. The pair mechanism is +> not implicated in this failure. + +The `pre-restore-` safety dump exists: `pre-restore-20260719T105811Z-immich-postgres.sql`, +51 964 807 B, 10:58:11 — written before anything was stopped or overwritten, as designed. + +### 3. DB truth + +Before Phase 3: `total 0 / trashed 0 / live 0`, 1 user — the trash was genuinely emptied, this was +real data loss (unlike round 1, where the rows survived as trashed). + +After Phase 3: **`asset: 11, all status=active, deletedAt null, isOffline false, all 11 owned by +the live user`**; all 11 `originalPath` values resolve to files on disk. Every timeline-visibility +condition is satisfied. + +### 4. File truth — what the 1.2 GB actually is + +The operator's suspicion (immich's own DB-backup feature accumulating) is a **minor** contributor: +one 18 MB file. The dominant item is something else entirely. + +| Item | Bytes | Share | +|---|---:|---:| +| `volume-dumps/immich_immich_ml_cache.tar` | 823 660 032 | **~60%** | +| `volume-dumps/immich_immich_postgres_data.tar` | 308 251 136 | ~23% | +| `db-dumps/immich-postgres.sql` | 52 393 708 | ~4% | +| `volume-dumps/immich_immich_redis_data.tar` | 6 358 016 | <1% | +| `appdata/immich` tree (below) | ~127 MB | ~10% | +| — of which `upload/` (originals, 2 user trees) | 72 MB | | +| — of which `encoded-video/` | 30 MB | | +| — of which `backups/` (immich's OWN nightly dump) | 18 MB | | +| — of which `thumbs/` | 8 MB | | + +**The single biggest thing in the customer's offsite backup is immich's machine-learning model +cache — 786 MiB of re-downloadable model weights.** Second is a raw tar of the postgres data +directory, which duplicates the logical `.sql` dump captured beside it (both are in every +snapshot). Actual irreplaceable customer content — the originals — is 72 MB, and even that includes +the ~36 MB stranded pre-v3 user tree (`dccc13fe…`) from round 1. + +So roughly **1.1 GB of a 1.2 GB "photo backup" is cache and duplication**, at the customer's offsite +quota and transfer cost. + +### 5. Phase-3 run — the new defect + +`pg_dump` is invoked `--clean --if-exists`, and `psql` with `ON_ERROR_STOP=1` (so an error aborts +rather than half-applying — correct). Plain-format pg_dump order is DROP → CREATE TABLE → COPY data +→ CREATE INDEX/constraints. The abort landed in the final stage: + +- 10:58:25 replay begins; the DROP + CREATE TABLE + COPY phases succeed (11 assets land) +- 10:58:33 **immich-server** logs `Reindexing clip_index` → `Reindexed clip_index` +- 10:58:35 the dump's own `CREATE INDEX clip_index` fails: *already exists* → exit 3 + +immich-server then logs `Detected schema drift` and lists indexes that are missing and need to be +created (`user_updatedAt_id_idx`, `library_ownerId_idx`, `stack_primaryAssetId_idx`, +`asset_id_timeline_notDeleted_idx`, …) — the ones the aborted script never reached. + +--- + +## Phase 2 — verdict + +| Hypothesis | Verdict | Decisive evidence | +|---|---|---| +| **H1** — mangled row caused a mis-click; missing-only ran, replay never did | **CONFIRMED** | `placed immich … 34 file(s) merged (missing-only)` at 10:39:55; `reconstituted / safety dump / replaying DB dump` counts all **0**; the reconstitute endpoint appears nowhere in the log | +| **H2** — full restore ran but replayed a wrong/stale unit | **REJECTED** | The full path never executed at all. Independently, the staging used `49e7cb46` — the correct pre-delete snapshot, stamped coherent, dump holding 11 assets | +| **H3** — replay ran correctly, immich-side hides the assets | **REJECTED** | No replay occurred; the DB was genuinely `asset: 0`. After the Phase-3 replay the assets are visible-by-every-criterion, so nothing immich-side hides them | +| **H4 (new)** — the replay races the app's own schema repair | **CONFIRMED** | `clip_index` recreated by immich-server at 10:58:33, dump's CREATE INDEX fails at 10:58:35; resulting schema drift reported by immich itself | + +--- + +## Phase 3 — recovery: PARTIAL, stopped as instructed + +One deliberate run, via the exact endpoints the buttons post to (the endpoint route was used rather +than a click **because the mangled 4-button row makes the correct target genuinely ambiguous** — +finding 1 below): + +1. `POST /backup/offbox/restore` `app=immich mode=full confirm=1` → staged `49e7cb46`, 1.3 G scratch. +2. `POST /backup/offbox/reconstitute` `app=immich confirm=1` → safety dump → stop → start → replay → + **abort** (H4). + +Result: **photos recovered and visible-by-every-DB-criterion; the operation reported failure; the +schema is incomplete.** Per the runbook — state is partially mutated, so **no second attempt was +made**. The safety dump from 10:58:11 is intact and untouched. + +> **Screenshot still needed from Viktor** (cannot be captured from SSH): the immich timeline showing +> the 11 photos. Requested; not yet held. + +--- + +## Phase 4 — findings + +### 1. [HIGH] Restore UX — adjacent buttons whose difference is "data returns vs data cannot return" + +Independent of the H-verdict, and the direct cause of H1: + +- **The second step is hidden.** „Teljes visszaállítás indítása" only appears after + „…előkészítése" has been pressed, with zero signposting that a second step exists or that the + first one did nothing to the live data. +- **The row overflows.** Four buttons plus hint text overlap into an unreadable row, so the honest + labels — the ones that distinguish a file-only merge from a real restore — are exactly what gets + lost. +- **The proven consequence class:** *two adjacent controls whose difference is "your data comes + back" versus "your data cannot come back" must not be separable only by layout.* This is now + demonstrated, not theorised: a competent operator who had read the code pressed the wrong one. + +Direction (already ruled in principle, spec rides v0.149): a single „Visszaállítás…" guided dialog — +one intent, visible phases, the wizard precedent. + +### 2. [HIGH, new] The DB replay races the application's own schema repair + +The reconstitution starts the stack before replaying, because `ImportDump` needs a live container. +That hands the application a window in which it recreates schema objects the dump is about to +create. Here it cost only the index phase; the data had already landed because pg_dump orders COPY +before CREATE INDEX. **That ordering is luck, not design** — a collision earlier in the script +aborts before the data and produces a half-restored database reported identically. + +Note the same start-then-replay shape exists on the LOCAL restore path +(`RestoreFromRecoveryUnit` → `RecreateStackFromUnit` → `reimportDBDumpsCtx`), so this is a class +defect, not an offsite-only one. Direction for v0.149 (not decided here): bring up the DB container +alone for the replay and start the app only afterwards, or quiesce the app's schema management for +the duration. + +### 3. [MED, new] A partially-successful restore reports as a flat failure + +The customer was told „az adatbázis visszaállítása sikertelen" while their photos were, in fact, +restored. This is the mirror image of round 1's defect (success reported over a no-op) and the same +class: **the message describes the mechanism's exit status, not the outcome.** Worse here, because +the true state — data restored, schema incomplete — is neither "success" nor "failure" and the UI +has no way to say it. + +### 4. [MED] Template classification — 1.1 GB of a 1.2 GB backup is cache and duplication + +Candidates to exclude from the offsite capture set, **recorded not changed**: + +- `immich_ml_cache` (786 MiB) — re-downloadable model weights, pure cache. Strong candidate. +- `immich_postgres_data` tar (294 MiB) — duplicates the logical `.sql` dump captured beside it. + Keeping both doubles the DB's footprint in every snapshot; the dump is already the authoritative + copy on the restore path. +- `upload/backups/` (18 MiB, immich's own nightly dump) — a backup inside a backup; grows daily. +- The stranded pre-v3 `dccc13fe…` tree (~36 MiB) from round 1, still unreferenced by any DB. + +### 5. Standing rulings, carried into the repo + +- **`00-capability-map.md:61`** — ruled: CAMPAIGN-6D's destruction hit the **file tree only**; the + database survived in its named volume (`immich_postgres_data` is a named volume in both eras), so + "immich end-to-end from offsite alone" **overclaimed scope**. Row → PARTIAL, scope-corrected. +- **The 704.6 MiB "discrepancy"** — ruled **resolved, not a defect**: the figure was immich's own + Tárhely widget (immich-internal accounting), never a controller page. Closed in the round-1 DIAG. + The same explanation covers today's „650 MiB → 1.4 GiB" observation, which likewise does not match + the controller's own measurement of the tree (127 MB). + +### 6. Capability map — customer-restore row + +**NOT flipped.** The Phase-3 run recovered the photos but reported failure and left schema drift, so +it is not the clean acceptance evidence §9 asks for. Recorded as +*"partial evidence captured 2026-07-19 — blocked on H4"* rather than flipped. + +--- + +## What was NOT done + +- No code, label or layout fixes — all of the above rides v0.149. +- No `restic forget`/`prune`, no snapshot deletion. +- The safety dumps were read, never modified or removed. +- **No second restore attempt** after the partially-mutating first run. diff --git a/documentation/backlog/ROADMAP.md b/documentation/backlog/ROADMAP.md index 1e8633b..f46e63a 100644 --- a/documentation/backlog/ROADMAP.md +++ b/documentation/backlog/ROADMAP.md @@ -86,6 +86,9 @@ | R-44 | **[P2-HIGH] A manual offsite push ships an unrefreshed DB dump — "backed up now" is false for the DB half.** `offboxRunHandler` → `RunOffboxBackup` goes straight to the restic push and never calls `RunDBDumps` / `captureAllRecoveryUnits` (`controller/internal/web/offbox_handlers.go:203-227`, `backup/offbox.go:574-759`); the recovery unit merely **enumerates** existing dump filenames via `listFileNames`, never creates them (`backup/recovery_unit.go:105-106`). Dumps come only from the separate local `db-dump` daily at **02:30** (`cmd/controller/main.go:542`), with the scheduled offsite at 04:15 — so a *manual* run at any other hour ships a dump up to ~24 h old. **There is no freshness check and no RPO surface anywhere:** zero `RPO` hits across `controller/`; `offboxUnitTime` is only a two-drive tiebreak (`offbox.go:827-837`); the `DBValidationCache` exists (`backup.go:364-370`) but no offsite or restore path reads it. | S–M | **SHIPPED controller v0.148.0 (2026-07-19)** | **SHIPPED:** every offsite run — **manual AND nightly** — now refreshes the DB/volume dumps and recovery units (`offsitePreDump` → `runDBDumpsInternal`) BEFORE the restic capture, so each snapshot is an internally coherent `{DB@T, files@T}` bundle and retention becomes a history of restorable points. Order is the mechanism and is red-proofed (moving the capture first yields `[capture dump]`): the gap can only ADD files the DB does not reference yet, never remove one it does. This also makes the nightly ordering **structural** rather than a coincidence of two scheduler entries at 02:30 and 04:15. Each unit manifest carries `offsite_run_id` + `dumps_at`, so a pair's coherence is verifiable at restore time instead of assumed; the periodic refresh carries a prior stamp forward and never invents one. A dump-leg failure is a loud WARN that does NOT abort the push (data-first: a degraded backup beats none). Honesty surfaces, all warn-level and none a gate: an unstamped pre-v0.148 pair reports its skew in the confirm, and `ValidateDump` gained an **exact-match** accounts-table sniff for customer-empty dumps (a substring match on "user" would flag `user_metadata`/`album_user`/`user_audit` on every healthy single-user box — red-proofed). — **Evidence: `audits/DIAG-immich-restore-2026-07-19.md`.** Today's unit dump `immich-postgres.sql` (51 954 452 B, mtime **02:30 CEST**) probed to **`asset: 0 rows`, `user: 0 rows`, `album: 0 rows`** — the 52 MB is entirely immich's shipped `geodata_places`/`naturalearth_countries` reference data. It predates both the admin user (created 07:56:25) and the photos (07:57). Same for the unit's `immich_immich_postgres_data.tar` (323 MB, also 02:30). **A dump that looks substantial by size can contain zero customer content** — size is not a health signal, and nothing in the product says otherwise. **Latent hazard:** had a full restore actually loaded that dump it would have written an empty DB over the live one, destroying the trashed rows that were the only surviving recovery path. Direction: dump-before-push on manual runs (the honest fix), **or** an explicit RPO line in the UI („adatbázis-állapot: ") so the operator/customer can see what they are actually shipping. Cheap interim: surface dump mtime + row-count sanity from the existing `DBValidationCache` on `/backups/restore` | | R-45 | **[P2] Unified async-job feedback.** Every long operation invents its own progress surface, or none. Tonight produced three more one-off cards (v0.147.x: samba bring-up, offsite progress, restore result) on top of two existing patterns (deploy 3-step panel; storage-init/netstorage status poll). They agree on nothing: some use `{ok,data}` envelopes and some raw JSON, some poll 1 s / 1.5 s / 3 s, some are in-memory-only and lie after a restart, and each re-implements single-flight + snapshot + phase→Hungarian mapping. | M | idea | Origin: 2026-07-19 feedback slice 1 (controller v0.147.0). The cases to generalise from are all in-tree: `web/storage_init_job.go` (the best shape — acquire/release/set/snapshot), `web/netstorage_job.go`, `web/samba_ensure_job.go`, `backup/opstatus.go`, `backup/offbox_progress.go`. Shape: one job registry + one poll endpoint + one client-side renderer, phases declared per job. **Two lessons tonight that any framework must encode:** (1) a terminal state must be **probed, not inferred** — `compose up -d` exits 0 on a crash-loop; (2) a progress source that reports nothing is normal, not broken — restic reports 0 bytes for a whole incremental run, and a bar that sits at 0% is worse than no bar. Also fixes the restart hole: in-memory job state currently vanishes and the card silently disagrees with reality | | R-46 | **[P2] Verification copies need a customer-visible browse surface and an expiry.** v0.147.0 made them *visible* (listed with path/size/date, individually deletable) — but the customer still cannot LOOK INSIDE a verification restore to confirm the file they wanted is really there, which is the entire point of a verification restore, and nothing ever removes them. | S–M | idea | Origin: 2026-07-19 feedback slice 4a, registered as the explicit follow-up to it. Two gaps, deliberately designed together because they are the same object: (a) **the invisible-result gap** — a read-only browse of `backups/offsite-restore/` (the FileBrowser infra stack already exists and already serves scoped roots, so this may be a mount rather than new code); (b) **the disk-lifecycle gap** — auto-expiry after N days with the count/size surfaced before it fires, so a drive is never quietly filled by verification restores nobody remembers taking. Pairs with R-43: a browse surface is also how a customer would discover that a DB-indexed app's files came back but the app still cannot see them | +| R-47 | **[P2-HIGH] The DB replay races the application's own schema repair — a restore can abort half-applied.** The reconstitution starts the stack BEFORE replaying (ImportDump needs a live container), which hands the app a window to recreate schema objects the dump is about to create. Proven to the second on 2026-07-19: controller began the replay 10:58:25, **immich-server logged `Reindexing clip_index` → `Reindexed clip_index` at 10:58:33**, and the dump's own `CREATE INDEX clip_index` failed at 10:58:35 with `already exists` (exit 3, `ON_ERROR_STOP=1`). immich then reported **schema drift** — the indexes the aborted script never reached. | M | idea | **Evidence: `audits/DIAG-immich-restore-round2-2026-07-19.md` (H4, live on demo-felhom).** The photos survived only because `pg_dump` emits COPY data BEFORE CREATE INDEX, so the abort landed after the rows — **that ordering is luck, not design**: a collision earlier in the script aborts before the data and leaves a genuinely half-restored database, reported identically. **Class defect, not offsite-only:** the same start-then-replay shape is on the LOCAL path (`RestoreFromRecoveryUnit` → `RecreateStackFromUnit` → `reimportDBDumpsCtx`), so the local restore carries the same race. Direction (NOT decided — needs a spec): bring up the DB container alone for the replay and start the app only afterwards, or quiesce the app's schema management for the duration. Note `--clean --if-exists` + `ON_ERROR_STOP=1` are both CORRECT and should stay — the bug is the window, not the flags. Blocks the §9 acceptance and therefore the customer-restore map row | +| R-48 | **[P2-HIGH] Restore controls are separable only by layout — and the difference between them is whether the data comes back.** The offsite restore row renders four buttons plus hint text into an overlapping, unreadable line, and the decisive second step („Teljes visszaállítás indítása") appears ONLY after „…előkészítése" was pressed, with no signposting that a second step exists or that the first one did nothing to live data. | M | idea | **Evidence: `audits/DIAG-immich-restore-round2-2026-07-19.md` (finding 1) — this is not theoretical: it is the CAUSE of the round-2 incident.** An operator who had read the code pressed the missing-only button instead of the full restore; the controller log shows `/backup/offbox/reconstitute` was never hit at all. The rule this establishes, worth stating once and applying beyond this page: **two adjacent controls whose difference is "your data comes back" vs "your data cannot come back" must not be distinguishable only by layout.** Direction (ruled in principle, spec rides v0.149): collapse to a single „Visszaállítás…" guided dialog — one intent, visible phases, the escrow-wizard precedent. Pairs with R-45 (the phases are exactly the async-feedback surface) and R-46 | +| R-49 | **[P2] The offsite capture set is ~90% cache and duplication — 1.1 GB of a 1.2 GB immich "photo backup".** Measured 2026-07-19: `immich_ml_cache.tar` **823 660 032 B (~60%)** — re-downloadable ML model weights; `immich_postgres_data.tar` **308 251 136 B (~23%)** — a raw tar of the postgres data dir that DUPLICATES the logical `.sql` dump captured beside it; `upload/backups/` **18 MB** — immich's own nightly dump, a backup inside the backup, growing daily; plus the stranded pre-v3 `dccc13fe…` tree (~36 MB) no DB has ever referenced. Actual irreplaceable content: **72 MB of originals**. | S–M | idea | **Evidence: `audits/DIAG-immich-restore-round2-2026-07-19.md` §4 (full byte breakdown).** This is the customer's offsite quota and transfer cost, and it lands on the Hetzner sub-account they are billed for. **Recorded, deliberately not changed** — a capture-set exclusion is a data-loss-shaped decision and gets its own ruling, not a drive-by edit. Candidates in priority order: (a) `immich_ml_cache` — pure cache, strongest case; (b) the `postgres_data` volume tar where a logical dump of the same DB is already captured (the dump is what the restore path actually replays); (c) `upload/backups/`. Likely generalises past immich into a template-classification rule about cache volumes and self-backup directories, so it should be specified against the catalog, not one app | ## Pre-invite checklist — what stands between here and the first remote tester