The undo bake-off, 2026-09-23 — copy the folder (F) against dump and load (D)
09 §3 decision 19. Venue: scratch guest 9202, controller v0.262.1, drill catalog (reset to live
cfcfe5278428 first; a837c3a7d8bd FROM states), update.health_timeout: 90s. Same three apps and
edges as the morning spike (../update-rulings-2026-09-23/), each a real migrating edge made to fail
a deliberately wrong probe in the DRILL template only. Seed A before the backup, seed B after it,
both through the app's own front door. Driver: bakeoff.py (stages per app: prep → fcopy → break →
update → undoF → update → undoD → cutoff). Time-boxed at 2 h; took ~40 min.
Method, stated so it is not over-read. Both undos were done BY HAND in the order the product would take them. The hold was lifted with the operator CLI + a controller restart (the only exit that exists) and the product's own Start supplied the app's secrets.
The table
| criterion | limit | F docmost / romm / vikunja | D docmost / romm / vikunja |
|---|---|---|---|
| seed A and seed B read back after the undo | both, every app | yes·yes / yes·yes / yes·yes | yes·yes / yes·yes / — (no safety dump exists: R-641) |
| tables / ledger equal to before the update | equal | 42 · 48 = / alembic 0095 = ¹ / 36 · 117 = | 42 · 48 = / 24 base tables · 0095 = / — |
| a cut-off copy detected before anything is swapped or loaded | detected | docmost 4 of 4 cuts (container exit 137, 42.9–68.4 of 69.2 MB copied, no finished-marker); romm (105.8 of 160.2 MB, no marker); vikunja could not be cut — 2.9 MB completes before a kill lands ² | docmost and romm: the half file lacks -- PostgreSQL database dump complete / -- Dump completed → refused; vikunja — |
| extra downtime before the update | ≤ 30 s | stop + copy: ≈ 2.4 s / ≈ 5.3 s / ≈ 1 s (stop 0.52 / 3.79 / 0.18 s; copies 0.4–1.0 s per volume incl. the helper's start) | 0 — the dump runs against the running database |
| rate | stated | cp -a 185 MB (Nextcloud's MariaDB) in 0.81 s; 2 GiB in 4.87 s ≈ 420 MB/s; tar 0.53 s / 4.55 s. Page cache could NOT be dropped in the guest, so these are warm-ish. A 5 GB database ≈ 12 s at this rate; on a cold cache or a spinning system disk expect 25–50 s | mariadb-dump of the same Nextcloud: 0.44 s, 1.06 MB |
| disk needed | stated; refuse near the 2 GB floor | the app's named volumes: 51 MB / 175 MB / 2.5 MB (Nextcloud 185 MB) | the dump: 136 KB / 63 KB / — (Nextcloud 1 MB) |
| product code the build needs | estimated | ≈ 300 lines, one mechanism for every app class | ≈ 450: the marker, two engine-specific empty-then-load paths, and F's own machinery anyway for apps with no database server |
¹ romm's base-table count was not captured on the F run (a quoting bug in the evidence script, fixed
before the D run, which read 24). F puts the whole datadir back byte-for-byte (copy bytes = source
bytes, file counts equal, romm-2-fcopy.txt), and the migration ledger read back equal.
² The marker logic is the same one proven on the other two; a copy that finishes cannot be cut.
The choice — F, copy the folder (decision 19)
Both methods pass every row. Decision 19 then says the folder copy wins on simplicity unless its downtime or disk cost fails the limits — and neither does: ≤ 5.3 s for these three, and the disk cost is refused near the 2 GB floor by construction. The decisive fact is the third column: an app with no database server gets no safety dump, so D would have to carry F's volume copy anyway. F is one mechanism; D is F plus two engine loaders.
Copy tool: cp -a into a sibling volume (<volume>.pre-update-<stamp>), not tar: the two ran at
the same speed, and a sibling volume puts the undo back with the same cp -a and needs no space on
the app's own drive.
Three things the bake-off found that the build must respect
- The undo must keep its own copy of the old definition. Lifting the hold made the box re-capture
docmost's recovery unit ten seconds later — with the NEW definition — and a pin-back that read the
unit started the new version on the restored data, which migrated again (
docmost-40-undoF.txt, invalid run, kept as evidence; redodocmost-41). R-639, seen live. - Killing
docker rundoes not stop the copy — the container runs on and finishes (docmost-60-cutoff.txt, first block). The product must judge a copy by the helper container's own exit status and the finished-marker it writes last, never by the client. - Where the copy is taken decides the downtime. Taken after the pull and just before
up— where the update stops the app anyway to recreate it — the extra downtime is the copy alone.
Evidence
docmost-*, romm-*, vikunja-* (per stage), 70–73 (the rate test on Nextcloud, removed after),
01/02 (repoint + three controls). Nextcloud and every .pre-undo/.cut/.rate volume were
removed by name at the end of the phase (73-rate-teardown.txt).