Files
felhom.eu/documentation/audits/night-burndown-2026-10-05/design-R-638.md
T
2026-10-05 23:32:12 +02:00

5.0 KiB

R-638 — restoring over a newer schema — design proposal (burn-down night 2026-10-05, no code)

Baselines read: felhom-controller ef199c5, felhom.eu b37902ce. Architecture: 07-backup-architecture.md §6 ("replay → rollback → hold", line 596).

1. The problem

The database loader replays a copy on top of the live database, so it only removes what the copy knows about. Measured 2026-09-23 on 9202: after docmost 0.95 → 0.96 migrated, replaying the older copy FAILED on PostgreSQL (rc 3, a new table's foreign key blocked the drop); on MariaDB (romm 5.0 → 5.3) it "succeeded" and left 12 newer tables behind (audits/update-rulings-2026-09-23/README.md Part 1, docmost-45, romm-44).

2. What the code does today (read in source)

  • Loader: psql -v ON_ERROR_STOP=1 --single-transaction / plain mariadb over the live DB (appbackup/dbdump.go:740-786). Copies are made with --clean --if-exists (dbdump.go:312).
  • Unit restore (the restore the hold sentence names): stop → volume tars REPLACE the named volumes (volume rm -f + create + untar, backup/restore.go:154-178) → definition from the unit, i.e. the data's own version (restore_unit.go:362-383, 444) → DB-only start → replay (restore_unit.go:438-458).
  • Off-site restore: same order — volumes from the scratch unit, then replay (offbox_reconstitute.go:856, 881); a failed volume leg restarts and stops there (:858-862).
  • Catalog: all 17 templates with a PostgreSQL or MariaDB data dir keep it in a NAMED volume (read in app-catalog-felhom.eu/templates/*/docker-compose.yml).
  • Inferred from the three above: on the two main paths the replay meets the copy's OWN schema, so R-638 is probably moot there. Not measured — that is the measurement the row owes.
  • Remaining exposures (inferred from source):
    1. No-manifest fallback RestoreApp: volumes back, then the WHOLE stack starts at the CURRENT definition (restore.go:67, 76) — a newer app can migrate the old data — then the replay runs (restore.go:82). This is the R-638 shape exactly.
    2. Unit restore with a failed volume leg still replays (restore_unit.go:439-442 sets the error and continues to :458).
    3. Rollback after a failed off-site replay pours the NEWER pre-restore copy over the OLDER volume just put back (offbox_reconstitute.go:898, :456-463) — the reverse direction; tables the migration removed would stay.

3. Options

A. Measure, then close the three gaps by order, not by loader. Fallback: start only DB services at the restored volume, replay, then start the app. Unit restore: do not replay when the DB's volume leg failed.

  • Cost: small, in two files. Risk: low; no change to what the loader does.
  • Measure first: the named unit restore after a real migration (docmost, romm) on 9202.

B. Make the loader rebuild instead of overlay. PostgreSQL: DROP SCHEMA public CASCADE; CREATE SCHEMA public; + the copy, in ONE transaction (measured working: rc 0, 1.38 s). MariaDB: drop every table first, then load.

  • Cost: medium. Fixes every caller at once, the rollback included.
  • Can go wrong: on MariaDB a failed load now leaves an EMPTY database (DDL is not transactional, 07 line 619) — the rollback must catch it. On PostgreSQL ≥ 15 the app user may not own public (speculative; must be measured per app). Objects in other schemas stay (immich-style extensions — unmeasured).

C. Refuse a replay when the copy's version is older than the live one. Uses the unit's data version record (restore_unit.go:362). Cost: small. Leaves the household with no restore at all in that case.

4. The pick — PROPOSAL for the operator, not a decision

Option A. Read in source, both shipped restore paths already put the copy's own database files back before they load the copy. So the loader is not the weak point; the order on three side paths is. A keeps the loader the drills have proven, and it adds no new delete step on customer data. B is the fallback if the measurement shows the main paths still fail.

5. First slice and its proof

  • Slice 0, measurement only (9202, throwaway docmost + romm): copy at the old version → update and migrate → unit restore. Positive observable: the replay log line Imported DB dump AND a read-back of a row written before the copy. Control from a different channel: \dt / SHOW TABLES counted against the copy's own table list (the 6 and 12 newer tables must be GONE). Evidence off the box before teardown.
  • Slice 1 (fallback order): red test first — a fake stack provider records call order; assert no full StartStack happens before the replay in RestoreApp. Fails today (restore.go:76 before :82).
  • Slice 2: red test — a unit restore whose volume leg errors must NOT call the importer. Fails today.

6. Open questions for the operator

  1. If the measurement shows the main paths are safe, may the row close on A alone, with B kept as a note? If you do nothing: the three side paths stay as they are.
  2. The rollback in the reverse direction (newer copy over an older volume): fix it now, or record it as a known limit?