07 §6.6 which version a restore brings back; version-travel evidence (A1, A5 red-proofs + live, A7, D1-D4); R-698 filed
gates / gates (push) Successful in 27s
gates / gates (push) Successful in 27s
This commit is contained in:
@@ -816,6 +816,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-695** | **[P3-LOW] Two kept-data Deletes in the same second can leave a self-perpetuating empty kept folder.** Seen 2026-09-25 on 9202 (v0.274.0, teardown): two Deletes at 13:35:16 each started `SyncFileBrowserMounts` in a goroutine; one restarted the file browser with the OLD bind list, Docker recreated the deleted folder `kept/nextcloud/2026-09-25_141014` EMPTY (root, 13:35:17), and the next sync LISTED that empty folder as a kept item and kept binding it — so the bind recreated the folder and the folder kept the bind. Broken by removing the folder and one controller restart. Harm: an empty 0-byte row and a stale empty source in the view; no data. **Fix direction:** single-flight the file-browser sync (the second waits and re-reads), and never list or bind an EMPTY dated kept folder. `audits/night-2026-09-26/G/T1-teardown-9202.txt` | **OPEN — P3; owner: CC** |
|
||||
| **R-696** | **[P2-MEDIUM] Tier 1's "proven at" is the unit MANIFEST's refresh time, not its data's — so the kept pre-conversion copy was released on a backup taken BEFORE the conversion.** Seen 2026-09-26 on demo-hp (v0.274.0, the automatic leg): docmost converted 16 → 18 at 02:18:52Z; at 02:20:16Z the recovery unit was re-captured (its pins changed) — `manifest.json` `created_at` 02:20:16Z — while its `db-dumps/docmost-postgres.sql` is from **02:15:01Z, on PostgreSQL 16**. `UpdateRestorePoints` reads Tier 1's time from `ListRestorePoints` (the unit's time), so at 02:30:16Z `ReleaseConversionCopies` logged "a backup proven on 18: Tier 1 at 02:20:16Z" and removed the 16 datadir copy. **Harm tonight: low** — the 02:15 logical dump holds the rows the conversion's check proved equal, and the next night's dump is on 18. **The class is wider:** the update's own precondition reads the same time, and a unit is re-captured on every controller release or pin change (`CaptureRecoveryUnit`), so a stale dump can read as minutes old (9202 2026-09-25 11:06: "Tier 1 copy 2m0s old" from the 11:04 restart, dumps older). On 9202 the release was right by luck (dump 11:57 > conversion 11:13). **Fix direction:** Tier 1's time = the newest DATA in the unit (its db/volume dumps' own times), never the manifest's; the release additionally requires a dump of the converted service written after the conversion. Presence-is-not-success, `CLAUDE.md`. `audits/night-2026-09-26/G/G4-demo-hp-docmost-after.txt` **-- A1 SPIKE 2026-09-26 (9202, controller 0.274.0, `audits/version-travel-2026-09-26/A1/README.md`) — the class is WIDER than the time: the refresh re-captures the unit's DEFINITION too, so the unit pairs new pins with old data.** Measured both shapes: an app step (docmost 0.95.0 → 0.96.0) — the refresh 2 min after the update wrote 0.96.0 into `compose/` over 0.95.0's dump; a restore in that window came back whole only because docmost migrated the old data forward at its first start (four migrations, no guard, no undo). An engine step (PostgreSQL 16 → 18) — the refresh wrote the 18 definition over the 16 dump and 16 datadir tar; a restore in that window poured the 16 datadir back, `postgres:18` REFUSED it, the replay timed out and **the app was left DOWN**. The second-drive mirror, written before the update, brought it back whole at 16. From source: the OFF-SITE restore never writes the definition at all (it uses the live compose), so after any update it mixes versions until the next night's off-site run. Being fixed as Part A of the version-travel brief (controller v0.275.0). | **OPEN — P2; owner: CC** |
|
||||
| **R-697** | **[P3-LOW] A restore drops the `conversion_copy` record but not the kept pre-conversion volume — the copy is then never released.** Seen 2026-09-26 on 9202 (A1, controller 0.274.0): after docmost's 16 → 18 conversion `app.yaml` held `conversion_copy` (`docmost_docmost_postgres_data.pre-update-20260926T073003Z`); a restore from the own unit, then one from the second drive, both left `conversion_copy=None` while the volume stayed. `PersistUnitRedeployConfig` builds a fresh `AppConfig` (Deployed, DeployedAt, Env, LockedFields), so every record kept beside the env is dropped by a restore — the hourly release reads the record, so the volume is orphaned. **Harm: disk only** (the size of the old datadir), never data. **Fix direction:** carry `conversion_copy` across the restore's app.yaml rewrite (the release then removes it when a backup on the new major is proven), and say in the log when a restore supersedes one. `audits/version-travel-2026-09-26/A1/S4-restore-window-step2.txt`, `S5-restore-tier2-window.txt` | **OPEN — P3; owner: CC** |
|
||||
| **R-698** | **[P3-LOW] A backup stores the image's NAME, not the image — a restore of a version its maker has deleted cannot start.** `RecoveryManifest.image_pins` ("image NOT stored — re-pulled on restore"); since controller v0.275.0 each data file also records its running `ref@digest`, and a restore brings the data back AT ITS OWN VERSION (`07` §6.6) — so a restore asks for exactly the old image. **Measured 2026-09-26** (`audits/version-travel-2026-09-26/A7/`, registry HEADs, no pulls): the catalog's 42 ladder `ref@digest` pairs all resolve (200); an invented digest answers 404 on Docker Hub and ghcr.io (negative control). Not measured: the digests recorded on boxes (older than any ladder entry), how often makers delete versions, the catalog's 66 digest-less compose lines. **Options (decide nothing yet):** (a) keep — a restore of a deleted version fails at the pull and the household uses the next copy or a newer version; (b) mirror every INSTALLED image into the DooPlex registry, restore falls back to it — storage + bandwidth on DooPlex, a new part on the recovery path; (c) mirror only ladder-named versions — bounded, misses pre-ladder boxes; (d) `docker save` into the unit — hundreds of MB per app per copy on every tier. | **OPEN — P3; owner: operator (a decision), CC measures** |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
|
||||
Reference in New Issue
Block a user