# Update rulings 2026-09-23 — the undo spike, the ladder spike, the memory watch Venue: scratch guest **9202** on demo-hp, controller v0.262.1, pointed at the **drill catalog** (`admin/app-catalog-drill`, reset to live `02844ae0a579` first) with `update.health_timeout: 90s`. The live catalog carried no test reference at any point (control 3 in `02-three-controls.txt`). Architecture document for the area: `architecture/09-update-architecture.md` (§3 decisions 11–18, §6.1a, §6.4). **Method honesty.** The product has no undo path, so the undo was performed BY HAND, in the order the product would take it, each step timed. Where a product path exists it was used: the Update, the Start, the Remove, the backup. Lifting the hold used the operator CLI `--clear-restore-hold` plus a controller restart (the only exit that exists today). **An attempt to decrypt the app's secrets so compose could be run by hand was refused by the session's safety guard; that route was dropped, not worked around** — the product's own Start supplies the secrets. --- ## Part 1 — the automatic undo, by hand ### Does a product path already load a safety dump back? **Yes, one — and the update never calls it.** `backup.Manager.rollbackSafetyDump` (`internal/backup/offbox_reconstitute.go:413`) re-applies a `pre-restore-*` undo copy through `ImportDump` — but only inside `ReconstituteFromOffsite`. `runGuardedUpdate` → `failAndHold` (`internal/stacks/update.go:766`) writes the safety dump in phase 3 and never reads it again. **And `failAndHold` deletes the journal's pre-update definition copies** (`removePreUpdateCopies`), so after a hold the old definition survives only in the recovery unit's `compose/` directory. It was there in all three cases because a backup preceded each update; that is not guaranteed. ### The four cases — each a REAL migration that then failed a deliberately wrong probe | case | edge (drill) | what migrated | the safety dump | old version on the migrated data, nothing loaded | the load | undo → healthy | seed A (before backup) | seed B (after backup) | |---|---|---|---|---|---|---|---|---| | **PostgreSQL** docmost | 0.95.0 → 0.96.0, held 95.3 s | 4 migrations, 42 → 48 tables (`docmost-31`) | 135 816 B, DB only, **holds B** (tier copy does not) | **REFUSES** — *corrupted migrations: previously executed migration 20260824T211732-page-title-trgm-index is missing* | product semantics **FAILS** (below); fixed load **1.38 s** | **14.8 s** | yes | **yes** | | **MariaDB** romm | 5.0.0 → 5.3.0, held 102.6 s | alembic 0095 → 0128, 27 → 39 tables | 62 943 B, DB only, **holds B** | **REFUSES** — *Can't locate revision identified by '0128_hltb_main_story_column'* | product semantics **1.25 s**, rc 0 — **12 new tables left behind** | **36.5 s** | yes | **yes** | | **Volume data, no DB server** vikunja | 2.3.0 → 2.6.0, held 93.3 s | *Ran all migrations successfully* (SQLite in a volume) | **none** — *the app has no database — nothing to copy (no-op)* | **STARTS AND SERVES** in 0.7 s, A, B and B's attachment read back | — | 0.7 s | yes | **yes** | | vikunja, **if it had refused** | the only other copy: the tier unit's volume tar | — | — | — | volume put back from the tier copy **0.86 s** | — | yes | **NO — lost** | | **Files on disk** | vikunja's attachment (a file in the files volume, written after the backup); romm's two drive folders | none touched the files | the undo copy never holds files | the attachment read back after the update AND after the undo | — | — | — | attachment yes | Three different apps, three mechanisms of refusal now measured (Nextcloud's version check §4, docmost's migration ledger, RomM's alembic revision) — and in each refusing case **the undo made the old version start**, because it met the data it knew. **romm's drive folders were empty before, after the update and after the undo** (tree hash `e3b0c442…` throughout) — so they measured nothing, and are reported as unmeasured, not as "untouched". ### The finding that shapes decision 15's build: the product's loader cannot undo a migration `ImportDump` replays a `pg_dump --clean --if-exists` file over the live database. **Over a database the new version migrated, that FAILS**: the new version created six tables (`oauth_*`, `public_spaces`, `siem_destinations`) whose foreign keys point at old tables, and the dump's own `DROP … workspaces_pkey` is refused — *cannot drop constraint workspaces_pkey on table public.workspaces because other objects depend on it* — rc 3 in 0.40 s, database unchanged (`docmost-45`). The same loader on MariaDB "succeeds" (`FOREIGN_KEY_CHECKS=0`) and leaves the new version's **12 tables** behind; RomM 5.0.0 happens to ignore them. **What worked:** empty the schema and load the copy **in ONE transaction** — `DROP SCHEMA public CASCADE; CREATE SCHEMA public;` + the dump, `psql --single-transaction ON_ERROR_STOP=1`: rc 0 in 1.38 s, 42 tables, 48 ledger rows, `pg_trgm` and `unaccent` back. ### The wrong case — a truncated undo copy | engine / loader | exit | what it left | is the outcome honest? | |---|---|---|---| | PostgreSQL, product semantics | rc 3 | unchanged (it failed on the foreign key before reaching the cut) | yes — nothing moved; the hold's named copy is intact | | **PostgreSQL, the fixed atomic load** | **rc 0 (!)** | **42 tables and 48 ledger rows — and 0 users, 0 spaces, 0 constraints, 0 indexes** (`docmost-47`, measured in a scratch database beside the real one) | **NO.** psql treats end-of-file inside a `COPY` as end of data and commits. The undo would report success, the old version would start on an EMPTY database, and a health check would pass on it. No hold. | | MariaDB, product semantics | rc 1 in 0.94 s | **half-replaced**: alembic back to 0095, later tables still the new version's — neither state | partly — the load is not transactional; the hold names the tier copy, which can still bring the app back | **The check that separates them already exists in the files:** the whole PostgreSQL copy ends with `-- PostgreSQL database dump complete` (and a `\unrestrict` line), the MariaDB copy with `-- Dump completed`; the truncated copies have neither. **`ValidateDump` does not look for them** — it checks the header and one `CREATE TABLE` (`internal/appbackup/dbdump.go:415`, read from source), so it would accept both truncated copies. ### Seconds | | find the copy | old definition + pin back | load | start → healthy (old probe) | undo total | |---|---|---|---|---|---| | docmost | 0.04 s | 0.03 s | 1.38 s | 14.8 s | **≈ 16 s** | | romm | 0.04 s | 0.04 s | 1.25 s | 36.5 s | **≈ 38 s** | | vikunja | — | 0.02 s | none | 0.7 s | **≈ 1 s** | Plus the failing health wait that precedes any undo (`update.health_timeout`, 90 s here, 5 min by default). Lifting the hold by hand cost 15.8 s (CLI + restart) and is NOT part of a product undo, which would never hold in the first place. ### What the product must add — the list decision 15's build starts from 1. **Keep the pre-update copies until the undo is over** — compose, applied definition, pin **and the old `.felhom.yml`**. `failAndHold` deletes the first three today; the fourth was never kept. 2. **The undo step inside `failAndHold`:** `pinBack` (exists) → DB service up alone → validated load → full start → **health check with the OLD `.felhom.yml` probe** (the new one may name a port the old version does not answer — it did in this spike by construction) → `undone`, or HOLD. 3. **The loader must empty the database first, atomically.** PostgreSQL: drop and recreate exactly the schemas the dump creates, in the same transaction as the load. MariaDB: drop every table first (`FOREIGN_KEY_CHECKS=0`); its DDL is not transactional, so a failed load is a HOLD with a sentence saying the database is in neither state. 4. **Validate the copy's completion marker before loading**, both engines. A truncated PostgreSQL copy is otherwise a silent, successful load of an empty database. 5. **Apps with no database server need their own last-second copy**: a tar of the data volumes at safety-dump time (vikunja: 2.2 MB, put back in 0.86 s). Without it, an app whose old version refuses loses everything written since the last backup. Vikunja's old version happens to start — **a per-app fact the test record should carry**, not a rule. 6. **Files on disk:** nothing measured was touched by a migration, and the undo does nothing to files. A step that does rewrite files must say so — decision 13's *files may change* mark is that place. 7. **The Start button says `start completed` (200) while the app crash-loops** (docmost and romm, both negative controls). The undo's success must be the probe, never the start's return. 8. **No retry:** after an undo the badge reads „Frissítés elérhető" again, because the catalog is still ahead. The caller must remember the failed step (decision 15), or it presses it every night. --- ## Part 2 — the ladder ### Today's behaviour, measured (`51-part2-ladder-today.txt`) vikunja installed at **2.3.0**; the drill catalog then got step **B = 2.4.0** (`610ed1a66dee`) and step **C = 2.5.0** (`c71807f5d47e`); one sync, one press. **Result: `done` in 9.5 s, pin, installed record, live compose and container all `vikunja/vikunja:2.5.0`; 2.4.0 never ran.** The box jumps. ### Can the box see the steps? No (`50-part2-clone-depth.txt`) Both 9202 (drill) and 9201 (live) hold a **shallow clone of depth 1**: `is-shallow-repository true`, `rev-list --count HEAD` = **1**, one commit visible per template. Source agrees: `sync.go:283` clones `--depth 1`, `sync.go:300` fetches `--depth 1`. **The brief's premise that "the box keeps a git clone" with history is wrong** — it keeps a clone of the newest commit only. ### The format — two options, one recommended **Size is not the cost.** The whole catalog history is 529 KiB packed, 286 commits. **The cost is that history does not contain the tested steps.** romm's compose has **16 commits, 3 of which move an image**; the step that works today is *15f9ebf's images with f4eb94f's template* (two workers, 768M) — **the commit that moved the image is the one that OOM-looped on demo-hp**. A box walking history would apply the definition that is known to be broken. | option | what it is | cost | |---|---|---| | **A — `update_ladder:` in `.felhom.yml` (recommended)** | per app, one entry per step: `from` and `to` refs per service, the digest per ref (decision 17), the test record (verdict, date, harness version, memory peak), the marks (`files_may_change`, `needs_person: ""`), and — for every step that is NOT the last — the step's own complete definition in `templates//steps/.yml`. The last step is the current `docker-compose.yml`. | every image move adds an entry (the catalog gate enforces it: no entry, no move — decision 13); an intermediate definition that needs a fix is fixed in its step file too, or the step is retired; entries older than the support window are pruned. `.felhom.yml` already flows to frozen apps (§5.4), so the box sees the ladder while frozen; step files are read from the clone. | | B — the box walks the catalog's git history | deepen the clone, treat each image-moving commit as a step | every box carries full history; a history rewrite breaks every box; a commit is not a tested step and cannot be made one retroactively; **and it applies the broken intermediate definition measured above** | The test record and the marks of decision 13 live in the same entry, so one gate reads one place. --- ## Part 3 — the memory watch in the harness (`app-catalog-felhom.eu/scripts/upgrade-test.py` v2) After an edge reads back, the harness runs the new version for `--soak` seconds (default 600) under four light callers and samples every 15 s, per container: memory against the compose limit, the **kernel's own `oom_kill` counter read host-side** (`/sys/fs/cgroup/system.slice/docker-.scope/ memory.events` — works on images with no shell, and does not depend on Docker's OOMKilled flag, which has read false for real kills on this kind of guest), the peak, and restarts. A kill or a restart → `failed`; a peak above 80 % → mark `memory_tight`. Run on 9202 in `/opt/upg` with raw compose (no controller in the path), after Part 1's apps were removed so the fixed `container_name`s could not collide. | edge | template | verdict | memory | |---|---|---|---| | **M1old** — the red-proof | as promoted, catalog `15f9ebf`: 512M, four workers | **failed** | seeded, migrated (2 lines), read back — then **first kernel OOM kill at +76 s**, peak 512 MiB = **100 %**, restarts **0** (the container kept running, which is how it hid on demo-hp), Docker's OOMKilled read true here; abort `refuses` | | **M1** — the positive control | current: 768M, two workers | **proven** + mark `memory_tight` | 608.5 s under 11 429 requests (5 712 × 200, 5 717 × 401): **0 kernel OOM kills, 0 restarts**, peak 621 MiB = **81 %** of 768 MiB — the watch passes the fix and still flags the thin headroom R-635 left open | **Is ten minutes long enough?** For this failure, under load, yes by a wide margin: +76 s. demo-hp's first kill came two hours in only because nothing was loading the app. What ten minutes cannot see is a leak that grows over days; that is a monitoring question (R-636), not a test-bench one. Evidence: `harness/M1old.log`, `harness/evidence/M1old/` (verdict, memory samples, logs), and the same for M1. --- ## Controls and teardown - `02-three-controls.txt` — 9202 followed the drill catalog; 9201 stayed on the live one; the live catalog's `main` was `02844ae0a579` before and after. - `60-part1-teardown.txt` — docmost, romm, vikunja removed through the product; romm's drive data was kept by the product (R-442's fail-closed refusal on 9202, as in every earlier drill) and removed by name at the end.