Files
felhom.eu/documentation/audits/undo-bakeoff-2026-09-23/README.md
T
admin 5a349d9884
gates / gates (push) Successful in 26s
Undo bake-off: copy the folder wins (09 §3 decisions 19-20, §6.1a)
- 09 §3: decision 19 (the copy method is chosen by a bake-off) and 20 (the
  full-system backup waits for the update leg, inside its window; built later).
- Bake-off on 9202, docmost / romm / vikunja: both methods pass every case;
  the folder copy wins because an app with no database server gets no dump,
  so dump-and-load would need the folder copy anyway. 1-5 s extra downtime,
  ~420 MB/s, disk = the volumes.
- R-645 filed: lifting an update hold by hand lets the recovery unit be
  re-captured with the failed definition within seconds.

Documents and evidence only; product code follows in the controller.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 10:28:56 +02:00

60 lines
5.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# The undo bake-off, 2026-09-23 — copy the folder (F) against dump and load (D)
`09` §3 decision 19. Venue: scratch guest **9202**, controller v0.262.1, drill catalog (reset to live
`cfcfe5278428` first; `a837c3a7d8bd` FROM states), `update.health_timeout: 90s`. Same three apps and
edges as the morning spike (`../update-rulings-2026-09-23/`), each a real migrating edge made to fail
a deliberately wrong probe in the DRILL template only. Seed A before the backup, seed B after it,
both through the app's own front door. Driver: `bakeoff.py` (stages per app: prep → fcopy → break →
update → undoF → update → undoD → cutoff). Time-boxed at 2 h; took ~40 min.
**Method, stated so it is not over-read.** Both undos were done BY HAND in the order the product
would take them. The hold was lifted with the operator CLI + a controller restart (the only exit that
exists) and the product's own Start supplied the app's secrets.
## The table
| criterion | limit | **F** docmost / romm / vikunja | **D** docmost / romm / vikunja |
|---|---|---|---|
| seed A and seed B read back after the undo | both, every app | **yes·yes / yes·yes / yes·yes** | yes·yes / yes·yes / — (no safety dump exists: R-641) |
| tables / ledger equal to before the update | equal | **42 · 48 = / alembic 0095 = ¹ / 36 · 117 =** | 42 · 48 = / **24 base tables** · 0095 = / — |
| a cut-off copy detected before anything is swapped or loaded | detected | **docmost 4 of 4 cuts** (container exit 137, 42.9–68.4 of 69.2 MB copied, no finished-marker); **romm** (105.8 of 160.2 MB, no marker); vikunja **could not be cut** — 2.9 MB completes before a kill lands ² | docmost and romm: the half file lacks `-- PostgreSQL database dump complete` / `-- Dump completed` → refused; vikunja — |
| extra downtime before the update | ≤ 30 s | stop + copy: **≈ 2.4 s / ≈ 5.3 s / ≈ 1 s** (stop 0.52 / 3.79 / 0.18 s; copies 0.4–1.0 s per volume incl. the helper's start) | 0 — the dump runs against the running database |
| rate | stated | cp -a **185 MB (Nextcloud's MariaDB) in 0.81 s; 2 GiB in 4.87 s ≈ 420 MB/s**; tar 0.53 s / 4.55 s. Page cache could NOT be dropped in the guest, so these are warm-ish. **A 5 GB database ≈ 12 s at this rate**; on a cold cache or a spinning system disk expect 25–50 s | `mariadb-dump` of the same Nextcloud: 0.44 s, 1.06 MB |
| disk needed | stated; refuse near the 2 GB floor | the app's named volumes: **51 MB / 175 MB / 2.5 MB** (Nextcloud 185 MB) | the dump: 136 KB / 63 KB / — (Nextcloud 1 MB) |
| product code the build needs | estimated | **≈ 300 lines**, one mechanism for every app class | ≈ 450: the marker, two engine-specific empty-then-load paths, **and F's own machinery anyway** for apps with no database server |
¹ romm's base-table count was not captured on the F run (a quoting bug in the evidence script, fixed
before the D run, which read 24). F puts the whole datadir back byte-for-byte (copy bytes = source
bytes, file counts equal, `romm-2-fcopy.txt`), and the migration ledger read back equal.
² The marker logic is the same one proven on the other two; a copy that finishes cannot be cut.
## The choice — F, copy the folder (decision 19)
**Both methods pass every row.** Decision 19 then says the folder copy wins on simplicity unless its
downtime or disk cost fails the limits — and neither does: ≤ 5.3 s for these three, and the disk
cost is refused near the 2 GB floor by construction. **The decisive fact is the third column:** an
app with no database server gets no safety dump, so D would have to carry F's volume copy anyway.
F is one mechanism; D is F plus two engine loaders.
**Copy tool: `cp -a` into a sibling volume** (`<volume>.pre-update-<stamp>`), not `tar`: the two ran at
the same speed, and a sibling volume puts the undo back with the same `cp -a` and needs no space on
the app's own drive.
## Three things the bake-off found that the build must respect
1. **The undo must keep its own copy of the old definition.** Lifting the hold made the box re-capture
docmost's recovery unit ten seconds later — with the NEW definition — and a pin-back that read the
unit started the new version on the restored data, which migrated again (`docmost-40-undoF.txt`,
invalid run, kept as evidence; redo `docmost-41`). R-639, seen live.
2. **Killing `docker run` does not stop the copy** — the container runs on and finishes
(`docmost-60-cutoff.txt`, first block). The product must judge a copy by the helper container's own
exit status and the finished-marker it writes last, never by the client.
3. **Where the copy is taken decides the downtime.** Taken after the pull and just before `up` — where
the update stops the app anyway to recreate it — the extra downtime is the copy alone.
## Evidence
`docmost-*`, `romm-*`, `vikunja-*` (per stage), `70`–`73` (the rate test on Nextcloud, removed after),
`01`/`02` (repoint + three controls). Nextcloud and every `.pre-undo`/`.cut`/`.rate` volume were
removed by name at the end of the phase (`73-rate-teardown.txt`).