5a349d9884
gates / gates (push) Successful in 26s
- 09 §3: decision 19 (the copy method is chosen by a bake-off) and 20 (the full-system backup waits for the update leg, inside its window; built later). - Bake-off on 9202, docmost / romm / vikunja: both methods pass every case; the folder copy wins because an app with no database server gets no dump, so dump-and-load would need the folder copy anyway. 1-5 s extra downtime, ~420 MB/s, disk = the volumes. - R-645 filed: lifting an update hold by hand lets the recovery unit be re-captured with the failed definition within seconds. Documents and evidence only; product code follows in the controller. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
60 lines
5.0 KiB
Markdown
60 lines
5.0 KiB
Markdown
# The undo bake-off, 2026-09-23 — copy the folder (F) against dump and load (D)
|
||
|
||
`09` §3 decision 19. Venue: scratch guest **9202**, controller v0.262.1, drill catalog (reset to live
|
||
`cfcfe5278428` first; `a837c3a7d8bd` FROM states), `update.health_timeout: 90s`. Same three apps and
|
||
edges as the morning spike (`../update-rulings-2026-09-23/`), each a real migrating edge made to fail
|
||
a deliberately wrong probe in the DRILL template only. Seed A before the backup, seed B after it,
|
||
both through the app's own front door. Driver: `bakeoff.py` (stages per app: prep → fcopy → break →
|
||
update → undoF → update → undoD → cutoff). Time-boxed at 2 h; took ~40 min.
|
||
|
||
**Method, stated so it is not over-read.** Both undos were done BY HAND in the order the product
|
||
would take them. The hold was lifted with the operator CLI + a controller restart (the only exit that
|
||
exists) and the product's own Start supplied the app's secrets.
|
||
|
||
## The table
|
||
|
||
| criterion | limit | **F** docmost / romm / vikunja | **D** docmost / romm / vikunja |
|
||
|---|---|---|---|
|
||
| seed A and seed B read back after the undo | both, every app | **yes·yes / yes·yes / yes·yes** | yes·yes / yes·yes / — (no safety dump exists: R-641) |
|
||
| tables / ledger equal to before the update | equal | **42 · 48 = / alembic 0095 = ¹ / 36 · 117 =** | 42 · 48 = / **24 base tables** · 0095 = / — |
|
||
| a cut-off copy detected before anything is swapped or loaded | detected | **docmost 4 of 4 cuts** (container exit 137, 42.9–68.4 of 69.2 MB copied, no finished-marker); **romm** (105.8 of 160.2 MB, no marker); vikunja **could not be cut** — 2.9 MB completes before a kill lands ² | docmost and romm: the half file lacks `-- PostgreSQL database dump complete` / `-- Dump completed` → refused; vikunja — |
|
||
| extra downtime before the update | ≤ 30 s | stop + copy: **≈ 2.4 s / ≈ 5.3 s / ≈ 1 s** (stop 0.52 / 3.79 / 0.18 s; copies 0.4–1.0 s per volume incl. the helper's start) | 0 — the dump runs against the running database |
|
||
| rate | stated | cp -a **185 MB (Nextcloud's MariaDB) in 0.81 s; 2 GiB in 4.87 s ≈ 420 MB/s**; tar 0.53 s / 4.55 s. Page cache could NOT be dropped in the guest, so these are warm-ish. **A 5 GB database ≈ 12 s at this rate**; on a cold cache or a spinning system disk expect 25–50 s | `mariadb-dump` of the same Nextcloud: 0.44 s, 1.06 MB |
|
||
| disk needed | stated; refuse near the 2 GB floor | the app's named volumes: **51 MB / 175 MB / 2.5 MB** (Nextcloud 185 MB) | the dump: 136 KB / 63 KB / — (Nextcloud 1 MB) |
|
||
| product code the build needs | estimated | **≈ 300 lines**, one mechanism for every app class | ≈ 450: the marker, two engine-specific empty-then-load paths, **and F's own machinery anyway** for apps with no database server |
|
||
|
||
¹ romm's base-table count was not captured on the F run (a quoting bug in the evidence script, fixed
|
||
before the D run, which read 24). F puts the whole datadir back byte-for-byte (copy bytes = source
|
||
bytes, file counts equal, `romm-2-fcopy.txt`), and the migration ledger read back equal.
|
||
² The marker logic is the same one proven on the other two; a copy that finishes cannot be cut.
|
||
|
||
## The choice — F, copy the folder (decision 19)
|
||
|
||
**Both methods pass every row.** Decision 19 then says the folder copy wins on simplicity unless its
|
||
downtime or disk cost fails the limits — and neither does: ≤ 5.3 s for these three, and the disk
|
||
cost is refused near the 2 GB floor by construction. **The decisive fact is the third column:** an
|
||
app with no database server gets no safety dump, so D would have to carry F's volume copy anyway.
|
||
F is one mechanism; D is F plus two engine loaders.
|
||
|
||
**Copy tool: `cp -a` into a sibling volume** (`<volume>.pre-update-<stamp>`), not `tar`: the two ran at
|
||
the same speed, and a sibling volume puts the undo back with the same `cp -a` and needs no space on
|
||
the app's own drive.
|
||
|
||
## Three things the bake-off found that the build must respect
|
||
|
||
1. **The undo must keep its own copy of the old definition.** Lifting the hold made the box re-capture
|
||
docmost's recovery unit ten seconds later — with the NEW definition — and a pin-back that read the
|
||
unit started the new version on the restored data, which migrated again (`docmost-40-undoF.txt`,
|
||
invalid run, kept as evidence; redo `docmost-41`). R-639, seen live.
|
||
2. **Killing `docker run` does not stop the copy** — the container runs on and finishes
|
||
(`docmost-60-cutoff.txt`, first block). The product must judge a copy by the helper container's own
|
||
exit status and the finished-marker it writes last, never by the client.
|
||
3. **Where the copy is taken decides the downtime.** Taken after the pull and just before `up` — where
|
||
the update stops the app anyway to recreate it — the extra downtime is the copy alone.
|
||
|
||
## Evidence
|
||
|
||
`docmost-*`, `romm-*`, `vikunja-*` (per stage), `70`–`73` (the rate test on Nextcloud, removed after),
|
||
`01`/`02` (repoint + three controls). Nextcloud and every `.pre-undo`/`.cut`/`.rate` volume were
|
||
removed by name at the end of the phase (`73-rate-teardown.txt`).
|