d549082be9
gates / gates (push) Successful in 2m7s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
41 lines
5.0 KiB
Markdown
41 lines
5.0 KiB
Markdown
# R-638 — restoring over a newer schema — design proposal (burn-down night 2026-10-05, no code)
|
|
|
|
Baselines read: felhom-controller `ef199c5`, felhom.eu `b37902ce`. Architecture: `07-backup-architecture.md` §6 ("replay → rollback → hold", line 596).
|
|
|
|
## 1. The problem
|
|
The database loader replays a copy on top of the live database, so it only removes what the copy knows about. Measured 2026-09-23 on 9202: after docmost 0.95 → 0.96 migrated, replaying the older copy FAILED on PostgreSQL (rc 3, a new table's foreign key blocked the drop); on MariaDB (romm 5.0 → 5.3) it "succeeded" and left 12 newer tables behind (`audits/update-rulings-2026-09-23/README.md` Part 1, `docmost-45`, `romm-44`).
|
|
|
|
## 2. What the code does today (read in source)
|
|
- Loader: `psql -v ON_ERROR_STOP=1 --single-transaction` / plain `mariadb` over the live DB (`appbackup/dbdump.go:740-786`). Copies are made with `--clean --if-exists` (`dbdump.go:312`).
|
|
- **Unit restore** (the restore the hold sentence names): stop → volume tars REPLACE the named volumes (`volume rm -f` + create + untar, `backup/restore.go:154-178`) → definition from the unit, i.e. the data's own version (`restore_unit.go:362-383, 444`) → DB-only start → replay (`restore_unit.go:438-458`).
|
|
- **Off-site restore**: same order — volumes from the scratch unit, then replay (`offbox_reconstitute.go:856, 881`); a failed volume leg restarts and stops there (`:858-862`).
|
|
- Catalog: all 17 templates with a PostgreSQL or MariaDB data dir keep it in a NAMED volume (read in `app-catalog-felhom.eu/templates/*/docker-compose.yml`).
|
|
- Inferred from the three above: on the two main paths the replay meets the copy's OWN schema, so R-638 is probably moot there. **Not measured** — that is the measurement the row owes.
|
|
- Remaining exposures (inferred from source):
|
|
1. **No-manifest fallback** `RestoreApp`: volumes back, then the WHOLE stack starts at the CURRENT definition (`restore.go:67, 76`) — a newer app can migrate the old data — then the replay runs (`restore.go:82`). This is the R-638 shape exactly.
|
|
2. **Unit restore with a failed volume leg** still replays (`restore_unit.go:439-442` sets the error and continues to `:458`).
|
|
3. **Rollback after a failed off-site replay** pours the NEWER pre-restore copy over the OLDER volume just put back (`offbox_reconstitute.go:898`, `:456-463`) — the reverse direction; tables the migration removed would stay.
|
|
|
|
## 3. Options
|
|
**A. Measure, then close the three gaps by order, not by loader.** Fallback: start only DB services at the restored volume, replay, then start the app. Unit restore: do not replay when the DB's volume leg failed.
|
|
- Cost: small, in two files. Risk: low; no change to what the loader does.
|
|
- Measure first: the named unit restore after a real migration (docmost, romm) on 9202.
|
|
|
|
**B. Make the loader rebuild instead of overlay.** PostgreSQL: `DROP SCHEMA public CASCADE; CREATE SCHEMA public;` + the copy, in ONE transaction (measured working: rc 0, 1.38 s). MariaDB: drop every table first, then load.
|
|
- Cost: medium. Fixes every caller at once, the rollback included.
|
|
- Can go wrong: on MariaDB a failed load now leaves an EMPTY database (DDL is not transactional, `07` line 619) — the rollback must catch it. On PostgreSQL ≥ 15 the app user may not own `public` (speculative; must be measured per app). Objects in other schemas stay (immich-style extensions — unmeasured).
|
|
|
|
**C. Refuse a replay when the copy's version is older than the live one.** Uses the unit's data version record (`restore_unit.go:362`). Cost: small. Leaves the household with no restore at all in that case.
|
|
|
|
## 4. The pick — PROPOSAL for the operator, not a decision
|
|
Option A. Read in source, both shipped restore paths already put the copy's own database files back before they load the copy. So the loader is not the weak point; the order on three side paths is. A keeps the loader the drills have proven, and it adds no new delete step on customer data. B is the fallback if the measurement shows the main paths still fail.
|
|
|
|
## 5. First slice and its proof
|
|
- Slice 0, measurement only (9202, throwaway docmost + romm): copy at the old version → update and migrate → unit restore. Positive observable: the replay log line `Imported DB dump` AND a read-back of a row written before the copy. Control from a different channel: `\dt` / `SHOW TABLES` counted against the copy's own table list (the 6 and 12 newer tables must be GONE). Evidence off the box before teardown.
|
|
- Slice 1 (fallback order): red test first — a fake stack provider records call order; assert no full `StartStack` happens before the replay in `RestoreApp`. Fails today (`restore.go:76` before `:82`).
|
|
- Slice 2: red test — a unit restore whose volume leg errors must NOT call the importer. Fails today.
|
|
|
|
## 6. Open questions for the operator
|
|
1. If the measurement shows the main paths are safe, may the row close on A alone, with B kept as a note? If you do nothing: the three side paths stay as they are.
|
|
2. The rollback in the reverse direction (newer copy over an older volume): fix it now, or record it as a known limit?
|