Files
felhom.eu/documentation/audits/undo-bakeoff-2026-09-23/README.md
T
admin 5a349d9884
gates / gates (push) Successful in 26s
Undo bake-off: copy the folder wins (09 §3 decisions 19-20, §6.1a)
- 09 §3: decision 19 (the copy method is chosen by a bake-off) and 20 (the
  full-system backup waits for the update leg, inside its window; built later).
- Bake-off on 9202, docmost / romm / vikunja: both methods pass every case;
  the folder copy wins because an app with no database server gets no dump,
  so dump-and-load would need the folder copy anyway. 1-5 s extra downtime,
  ~420 MB/s, disk = the volumes.
- R-645 filed: lifting an update hold by hand lets the recovery unit be
  re-captured with the failed definition within seconds.

Documents and evidence only; product code follows in the controller.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 10:28:56 +02:00

5.0 KiB
Raw Blame History

The undo bake-off, 2026-09-23 — copy the folder (F) against dump and load (D)

09 §3 decision 19. Venue: scratch guest 9202, controller v0.262.1, drill catalog (reset to live cfcfe5278428 first; a837c3a7d8bd FROM states), update.health_timeout: 90s. Same three apps and edges as the morning spike (../update-rulings-2026-09-23/), each a real migrating edge made to fail a deliberately wrong probe in the DRILL template only. Seed A before the backup, seed B after it, both through the app's own front door. Driver: bakeoff.py (stages per app: prep → fcopy → break → update → undoF → update → undoD → cutoff). Time-boxed at 2 h; took ~40 min.

Method, stated so it is not over-read. Both undos were done BY HAND in the order the product would take them. The hold was lifted with the operator CLI + a controller restart (the only exit that exists) and the product's own Start supplied the app's secrets.

The table

criterion limit F docmost / romm / vikunja D docmost / romm / vikunja
seed A and seed B read back after the undo both, every app yes·yes / yes·yes / yes·yes yes·yes / yes·yes / — (no safety dump exists: R-641)
tables / ledger equal to before the update equal 42 · 48 = / alembic 0095 = ¹ / 36 · 117 = 42 · 48 = / 24 base tables · 0095 = / —
a cut-off copy detected before anything is swapped or loaded detected docmost 4 of 4 cuts (container exit 137, 42.9–68.4 of 69.2 MB copied, no finished-marker); romm (105.8 of 160.2 MB, no marker); vikunja could not be cut — 2.9 MB completes before a kill lands ² docmost and romm: the half file lacks -- PostgreSQL database dump complete / -- Dump completed → refused; vikunja —
extra downtime before the update ≤ 30 s stop + copy: ≈ 2.4 s / ≈ 5.3 s / ≈ 1 s (stop 0.52 / 3.79 / 0.18 s; copies 0.4–1.0 s per volume incl. the helper's start) 0 — the dump runs against the running database
rate stated cp -a 185 MB (Nextcloud's MariaDB) in 0.81 s; 2 GiB in 4.87 s ≈ 420 MB/s; tar 0.53 s / 4.55 s. Page cache could NOT be dropped in the guest, so these are warm-ish. A 5 GB database ≈ 12 s at this rate; on a cold cache or a spinning system disk expect 25–50 s mariadb-dump of the same Nextcloud: 0.44 s, 1.06 MB
disk needed stated; refuse near the 2 GB floor the app's named volumes: 51 MB / 175 MB / 2.5 MB (Nextcloud 185 MB) the dump: 136 KB / 63 KB / — (Nextcloud 1 MB)
product code the build needs estimated ≈ 300 lines, one mechanism for every app class ≈ 450: the marker, two engine-specific empty-then-load paths, and F's own machinery anyway for apps with no database server

¹ romm's base-table count was not captured on the F run (a quoting bug in the evidence script, fixed before the D run, which read 24). F puts the whole datadir back byte-for-byte (copy bytes = source bytes, file counts equal, romm-2-fcopy.txt), and the migration ledger read back equal. ² The marker logic is the same one proven on the other two; a copy that finishes cannot be cut.

The choice — F, copy the folder (decision 19)

Both methods pass every row. Decision 19 then says the folder copy wins on simplicity unless its downtime or disk cost fails the limits — and neither does: ≤ 5.3 s for these three, and the disk cost is refused near the 2 GB floor by construction. The decisive fact is the third column: an app with no database server gets no safety dump, so D would have to carry F's volume copy anyway. F is one mechanism; D is F plus two engine loaders.

Copy tool: cp -a into a sibling volume (<volume>.pre-update-<stamp>), not tar: the two ran at the same speed, and a sibling volume puts the undo back with the same cp -a and needs no space on the app's own drive.

Three things the bake-off found that the build must respect

  1. The undo must keep its own copy of the old definition. Lifting the hold made the box re-capture docmost's recovery unit ten seconds later — with the NEW definition — and a pin-back that read the unit started the new version on the restored data, which migrated again (docmost-40-undoF.txt, invalid run, kept as evidence; redo docmost-41). R-639, seen live.
  2. Killing docker run does not stop the copy — the container runs on and finishes (docmost-60-cutoff.txt, first block). The product must judge a copy by the helper container's own exit status and the finished-marker it writes last, never by the client.
  3. Where the copy is taken decides the downtime. Taken after the pull and just before up — where the update stops the app anyway to recreate it — the extra downtime is the copy alone.

Evidence

docmost-*, romm-*, vikunja-* (per stage), 70–73 (the rate test on Nextcloud, removed after), 01/02 (repoint + three controls). Nextcloud and every .pre-undo/.cut/.rate volume were removed by name at the end of the phase (73-rate-teardown.txt).