Update rulings 2026-09-23 — the undo spike, the ladder spike, the memory watch
Venue: scratch guest 9202 on demo-hp, controller v0.262.1, pointed at the drill catalog
(admin/app-catalog-drill, reset to live 02844ae0a579 first) with update.health_timeout: 90s.
The live catalog carried no test reference at any point (control 3 in 02-three-controls.txt).
Architecture document for the area: architecture/09-update-architecture.md (§3 decisions 11–18,
§6.1a, §6.4).
Method honesty. The product has no undo path, so the undo was performed BY HAND, in the order
the product would take it, each step timed. Where a product path exists it was used: the Update, the
Start, the Remove, the backup. Lifting the hold used the operator CLI --clear-restore-hold plus a
controller restart (the only exit that exists today). An attempt to decrypt the app's secrets so
compose could be run by hand was refused by the session's safety guard; that route was dropped, not
worked around — the product's own Start supplies the secrets.
Part 1 — the automatic undo, by hand
Does a product path already load a safety dump back?
Yes, one — and the update never calls it. backup.Manager.rollbackSafetyDump
(internal/backup/offbox_reconstitute.go:413) re-applies a pre-restore-* undo copy through
ImportDump — but only inside ReconstituteFromOffsite. runGuardedUpdate → failAndHold
(internal/stacks/update.go:766) writes the safety dump in phase 3 and never reads it again.
And failAndHold deletes the journal's pre-update definition copies (removePreUpdateCopies),
so after a hold the old definition survives only in the recovery unit's compose/ directory. It was
there in all three cases because a backup preceded each update; that is not guaranteed.
The four cases — each a REAL migration that then failed a deliberately wrong probe
| case | edge (drill) | what migrated | the safety dump | old version on the migrated data, nothing loaded | the load | undo → healthy | seed A (before backup) | seed B (after backup) |
|---|---|---|---|---|---|---|---|---|
| PostgreSQL docmost | 0.95.0 → 0.96.0, held 95.3 s | 4 migrations, 42 → 48 tables (docmost-31) |
135 816 B, DB only, holds B (tier copy does not) | REFUSES — corrupted migrations: previously executed migration 20260824T211732-page-title-trgm-index is missing | product semantics FAILS (below); fixed load 1.38 s | 14.8 s | yes | yes |
| MariaDB romm | 5.0.0 → 5.3.0, held 102.6 s | alembic 0095 → 0128, 27 → 39 tables | 62 943 B, DB only, holds B | REFUSES — Can't locate revision identified by '0128_hltb_main_story_column' | product semantics 1.25 s, rc 0 — 12 new tables left behind | 36.5 s | yes | yes |
| Volume data, no DB server vikunja | 2.3.0 → 2.6.0, held 93.3 s | Ran all migrations successfully (SQLite in a volume) | none — the app has no database — nothing to copy (no-op) | STARTS AND SERVES in 0.7 s, A, B and B's attachment read back | — | 0.7 s | yes | yes |
| vikunja, if it had refused | the only other copy: the tier unit's volume tar | — | — | — | volume put back from the tier copy 0.86 s | — | yes | NO — lost |
| Files on disk | vikunja's attachment (a file in the files volume, written after the backup); romm's two drive folders | none touched the files | the undo copy never holds files | the attachment read back after the update AND after the undo | — | — | — | attachment yes |
Three different apps, three mechanisms of refusal now measured (Nextcloud's version check §4, docmost's migration ledger, RomM's alembic revision) — and in each refusing case the undo made the old version start, because it met the data it knew.
romm's drive folders were empty before, after the update and after the undo (tree hash
e3b0c442… throughout) — so they measured nothing, and are reported as unmeasured, not as "untouched".
The finding that shapes decision 15's build: the product's loader cannot undo a migration
ImportDump replays a pg_dump --clean --if-exists file over the live database. Over a database
the new version migrated, that FAILS: the new version created six tables (oauth_*,
public_spaces, siem_destinations) whose foreign keys point at old tables, and the dump's own
DROP … workspaces_pkey is refused — cannot drop constraint workspaces_pkey on table
public.workspaces because other objects depend on it — rc 3 in 0.40 s, database unchanged
(docmost-45). The same loader on MariaDB "succeeds" (FOREIGN_KEY_CHECKS=0) and leaves the new
version's 12 tables behind; RomM 5.0.0 happens to ignore them.
What worked: empty the schema and load the copy in ONE transaction —
DROP SCHEMA public CASCADE; CREATE SCHEMA public; + the dump, psql --single-transaction ON_ERROR_STOP=1: rc 0 in 1.38 s, 42 tables, 48 ledger rows, pg_trgm and unaccent back.
The wrong case — a truncated undo copy
| engine / loader | exit | what it left | is the outcome honest? |
|---|---|---|---|
| PostgreSQL, product semantics | rc 3 | unchanged (it failed on the foreign key before reaching the cut) | yes — nothing moved; the hold's named copy is intact |
| PostgreSQL, the fixed atomic load | rc 0 (!) | 42 tables and 48 ledger rows — and 0 users, 0 spaces, 0 constraints, 0 indexes (docmost-47, measured in a scratch database beside the real one) |
NO. psql treats end-of-file inside a COPY as end of data and commits. The undo would report success, the old version would start on an EMPTY database, and a health check would pass on it. No hold. |
| MariaDB, product semantics | rc 1 in 0.94 s | half-replaced: alembic back to 0095, later tables still the new version's — neither state | partly — the load is not transactional; the hold names the tier copy, which can still bring the app back |
The check that separates them already exists in the files: the whole PostgreSQL copy ends with
-- PostgreSQL database dump complete (and a \unrestrict line), the MariaDB copy with
-- Dump completed; the truncated copies have neither. ValidateDump does not look for them — it
checks the header and one CREATE TABLE (internal/appbackup/dbdump.go:415, read from source), so it
would accept both truncated copies.
Seconds
| find the copy | old definition + pin back | load | start → healthy (old probe) | undo total | |
|---|---|---|---|---|---|
| docmost | 0.04 s | 0.03 s | 1.38 s | 14.8 s | ≈ 16 s |
| romm | 0.04 s | 0.04 s | 1.25 s | 36.5 s | ≈ 38 s |
| vikunja | — | 0.02 s | none | 0.7 s | ≈ 1 s |
Plus the failing health wait that precedes any undo (update.health_timeout, 90 s here, 5 min by
default). Lifting the hold by hand cost 15.8 s (CLI + restart) and is NOT part of a product undo,
which would never hold in the first place.
What the product must add — the list decision 15's build starts from
- Keep the pre-update copies until the undo is over — compose, applied definition, pin and the
old
.felhom.yml.failAndHolddeletes the first three today; the fourth was never kept. - The undo step inside
failAndHold:pinBack(exists) → DB service up alone → validated load → full start → health check with the OLD.felhom.ymlprobe (the new one may name a port the old version does not answer — it did in this spike by construction) →undone, or HOLD. - The loader must empty the database first, atomically. PostgreSQL: drop and recreate exactly
the schemas the dump creates, in the same transaction as the load. MariaDB: drop every table first
(
FOREIGN_KEY_CHECKS=0); its DDL is not transactional, so a failed load is a HOLD with a sentence saying the database is in neither state. - Validate the copy's completion marker before loading, both engines. A truncated PostgreSQL copy is otherwise a silent, successful load of an empty database.
- Apps with no database server need their own last-second copy: a tar of the data volumes at safety-dump time (vikunja: 2.2 MB, put back in 0.86 s). Without it, an app whose old version refuses loses everything written since the last backup. Vikunja's old version happens to start — a per-app fact the test record should carry, not a rule.
- Files on disk: nothing measured was touched by a migration, and the undo does nothing to files. A step that does rewrite files must say so — decision 13's files may change mark is that place.
- The Start button says
start completed(200) while the app crash-loops (docmost and romm, both negative controls). The undo's success must be the probe, never the start's return. - No retry: after an undo the badge reads „Frissítés elérhető" again, because the catalog is still ahead. The caller must remember the failed step (decision 15), or it presses it every night.
Part 2 — the ladder
Today's behaviour, measured (51-part2-ladder-today.txt)
vikunja installed at 2.3.0; the drill catalog then got step B = 2.4.0 (610ed1a66dee) and
step C = 2.5.0 (c71807f5d47e); one sync, one press. Result: done in 9.5 s, pin, installed
record, live compose and container all vikunja/vikunja:2.5.0; 2.4.0 never ran. The box jumps.
Can the box see the steps? No (50-part2-clone-depth.txt)
Both 9202 (drill) and 9201 (live) hold a shallow clone of depth 1: is-shallow-repository true,
rev-list --count HEAD = 1, one commit visible per template. Source agrees:
sync.go:283 clones --depth 1, sync.go:300 fetches --depth 1. The brief's premise that "the
box keeps a git clone" with history is wrong — it keeps a clone of the newest commit only.
The format — two options, one recommended
Size is not the cost. The whole catalog history is 529 KiB packed, 286 commits.
The cost is that history does not contain the tested steps. romm's compose has 16 commits, 3 of which move an image; the step that works today is 15f9ebf's images with f4eb94f's template (two workers, 768M) — the commit that moved the image is the one that OOM-looped on demo-hp. A box walking history would apply the definition that is known to be broken.
| option | what it is | cost |
|---|---|---|
A — update_ladder: in .felhom.yml (recommended) |
per app, one entry per step: from and to refs per service, the digest per ref (decision 17), the test record (verdict, date, harness version, memory peak), the marks (files_may_change, needs_person: "<why>"), and — for every step that is NOT the last — the step's own complete definition in templates/<app>/steps/<to>.yml. The last step is the current docker-compose.yml. |
every image move adds an entry (the catalog gate enforces it: no entry, no move — decision 13); an intermediate definition that needs a fix is fixed in its step file too, or the step is retired; entries older than the support window are pruned. .felhom.yml already flows to frozen apps (§5.4), so the box sees the ladder while frozen; step files are read from the clone. |
| B — the box walks the catalog's git history | deepen the clone, treat each image-moving commit as a step | every box carries full history; a history rewrite breaks every box; a commit is not a tested step and cannot be made one retroactively; and it applies the broken intermediate definition measured above |
The test record and the marks of decision 13 live in the same entry, so one gate reads one place.
Part 3 — the memory watch in the harness (app-catalog-felhom.eu/scripts/upgrade-test.py v2)
After an edge reads back, the harness runs the new version for --soak seconds (default 600) under
four light callers and samples every 15 s, per container: memory against the compose limit, the
kernel's own oom_kill counter read host-side (/sys/fs/cgroup/system.slice/docker-<id>.scope/ memory.events — works on images with no shell, and does not depend on Docker's OOMKilled flag, which
has read false for real kills on this kind of guest), the peak, and restarts. A kill or a restart →
failed; a peak above 80 % → mark memory_tight.
Run on 9202 in /opt/upg with raw compose (no controller in the path), after Part 1's apps were
removed so the fixed container_names could not collide.
| edge | template | verdict | memory |
|---|---|---|---|
| M1old — the red-proof | as promoted, catalog 15f9ebf: 512M, four workers |
failed | seeded, migrated (2 lines), read back — then first kernel OOM kill at +76 s, peak 512 MiB = 100 %, restarts 0 (the container kept running, which is how it hid on demo-hp), Docker's OOMKilled read true here; abort refuses |
| M1 — the positive control | current: 768M, two workers | proven + mark memory_tight |
608.5 s under 11 429 requests (5 712 × 200, 5 717 × 401): 0 kernel OOM kills, 0 restarts, peak 621 MiB = 81 % of 768 MiB — the watch passes the fix and still flags the thin headroom R-635 left open |
Is ten minutes long enough? For this failure, under load, yes by a wide margin: +76 s. demo-hp's first kill came two hours in only because nothing was loading the app. What ten minutes cannot see is a leak that grows over days; that is a monitoring question (R-636), not a test-bench one.
Evidence: harness/M1old.log, harness/evidence/M1old/ (verdict, memory samples, logs), and the same
for M1.
Controls and teardown
02-three-controls.txt— 9202 followed the drill catalog; 9201 stayed on the live one; the live catalog'smainwas02844ae0a579before and after.60-part1-teardown.txt— docmost, romm, vikunja removed through the product; romm's drive data was kept by the product (R-442's fail-closed refusal on 9202, as in every earlier drill) and removed by name at the end.