Files
felhom.eu/documentation/audits/day-2026-10-08/design-R-893.md
T

5.6 KiB
Raw Blame History

R-893 — after a failed off-site database replay, the rollback mixes new and old: a one-page design (2026-10-08)

Baselines read: felhom-controller a0370b4, felhom.eu b2dce901. Architecture: 07-backup-architecture.md §6.3 („[FACT] 2026-10-06 — restoring over a newer schema", the R-893 known limit) and „[DESIGN] 2026-08-22 — the failure ladder: replay → rollback → hold"; 09 §3 decision 154 (R-638 option A). Status: design only, nothing built. Not measured yet — the 9202 measurement the row asks for is slice 0 below.

1. The problem (read in source today, controller/internal/backup/offbox_reconstitute.go)

The off-site restore of one app does, in order:

  1. :758 — the undo copy: a logical dump of the LIVE database (writeSafetyDump, :214). Nothing else is saved.
  2. :804-811 — when the snapshot was written at another app version, the snapshot's older definition is written.
  3. :812-842 — the snapshot's files are copied over the live ones (placements).
  4. :856 — the named volumes are replaced with the snapshot's tars (this includes the database volume).
  5. :873-917 — DB-only start, replay of the snapshot's dump. If the replay fails, rollbackSafetyDump (:456) loads the step-1 dump (NEWER state) into the database that now sits on the step-4 (OLDER) volume, then the app starts.

What the household gets after a failed replay:

  • the database: the older volume with the newer dump laid over it — the loader drops only what the dump knows, so tables only the older version had stay (07 §6.3);
  • files and other volumes: the snapshot's state (steps 3–4 are never undone);
  • the definition: the older one when step 2 ran — the rollback branch never writes the live one back.

And the sentence shown (hu.json:1261, en.json:1266, key err.backup.db_restore_failed_rolled_back) says „your data is back as it was before the restore, and the app is still running". That is true only of the database rows, and not of the files, the other volumes or the version.

2. Options

What Costs Risk
A Pre-restore copy of the app's volumes. After step 1, while stopped, tar the live named volumes to an undo folder beside the undo dump (the recovery-unit volume-dump helper; one implementation). On a failed replay: put back those volumes, the live definition and the live files' copy → the exact state before the restore. A fit check refuses to start when the copy does not fit (R-685's lesson). Disk = one copy of the app's volumes for the restore's duration; time ≈ one volume dump. ~1 session + a 9202 drill. Small: the helpers exist; the new part is the undo of steps 2–4. Files placed by step 3 need the same treatment, or step 3 must write to a staging folder first.
B R-638 option B: a loader that rebuilds (drop and re-create the database before loading). The newer dump over the older volume then gives exactly the newer database. Changes the loader under EVERY restore path, Postgres and MariaDB. ~2 sessions. High blast radius; still leaves files, other volumes and the definition at the snapshot's state.
C Hold instead of a mixed start. On a failed replay when step 2 ran (version changed) or step 4 replaced any volume: roll the database back as now, write the live definition back, and hold the app stopped (the ladder's third step, holdAppAfterFailedRollback) with an honest sentence and an operator event. Small: one branch, one sentence pair (hu/en), tests. ~½ session. The app is down until the operator acts; no household data is changed beyond what happens today.

3. The pick — C now, then A

C removes the worst outcome (an older app running on mixed data while the screen says „back as it was") at no disk cost and without touching the loader. A is the real fix — it makes „back as it was" true — and is built after the 9202 measurement shows how big the volume copy is for the standing apps. B stays a note: it fixes only the database and widens the risk to every restore.

4. First slices and their red tests

  • Slice 0 — measure (9202 only, throwaway app): off-site snapshot at version N, Update to N+1 (migrating), then an off-site restore with the replay forced to fail (the rollbackImport/reimportDBDumpsFrom seam is for tests; on the box, a snapshot dump truncated by hand in the scratch copy). Record: tables, files, app version, the sentence. Evidence copied off before teardown.
  • Slice 1 (C): red test beside r379_rollback_test.go — ReconstituteFromOffsite with defineFromSnapshot true and a failing replay must (a) call the definition write with the LIVE definition, (b) NOT start the app, (c) leave a hold, (d) return a sentence that does not contain „visszakerültek" / „back as it was". Today (a)–(d) all fail — TestR379_ScenarioA pins the opposite for its fixture, which has no version change, so it stays green.
  • Slice 2 (A): red test — after a failed replay, the volume-replay seam is called a second time with the undo folder, and the app starts; a too-small free space refuses before step 1 with nothing touched.

5. One question for the operator

When an off-site restore of one app fails half-way, which do you want: the app stopped and waiting for us (no extra disk), or the app put back exactly as it was (needs free space for one more copy of that app's data during the restore)? My pick: stopped now, „exactly as it was" next. If you do nothing: after such a failure the app keeps starting on a mix of the older backup's files and the newer database, and the screen tells the household everything is back as it was.