Files
felhom.eu/documentation/audits/update-rulings-2026-09-23
admin 4c92beab8f
gates / gates (push) Successful in 26s
Update arc: the undo and ladder spiked, the build plan for the 2026-09-23 rulings
- Part 1 (09 §6.1a, audit): the undo performed by hand on 9202 for docmost
  (PostgreSQL), romm (MariaDB) and vikunja (SQLite volume) - all three came
  back with data written before AND after the backup. The product's loader
  cannot do it: over a migrated PG database it fails on the new tables'
  foreign keys; over MariaDB it leaves them behind. A truncated PG copy loads
  rc 0 into an empty database. No-DB apps have no last-second copy.
- Part 2: one press jumps A -> C; the box's catalog clone is depth 1.
  Ladder format recommended: update_ladder in .felhom.yml, not git history.
- Part 3: memory watch red-proof results (harness change in the catalog repo).
- Part 4 (09 §6.4): ten parts, ~22 evenings; one open point (R-643).
- Rows R-637..R-644 opened; R-446/450/451/462/463 updated. STATUS, CONTEXT.

No product code. Live catalog untouched; 9202 back on it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 08:54:41 +02:00
..
…
…
…

Update rulings 2026-09-23 — the undo spike, the ladder spike, the memory watch

Venue: scratch guest 9202 on demo-hp, controller v0.262.1, pointed at the drill catalog (admin/app-catalog-drill, reset to live 02844ae0a579 first) with update.health_timeout: 90s. The live catalog carried no test reference at any point (control 3 in 02-three-controls.txt). Architecture document for the area: architecture/09-update-architecture.md (§3 decisions 11–18, §6.1a, §6.4).

Method honesty. The product has no undo path, so the undo was performed BY HAND, in the order the product would take it, each step timed. Where a product path exists it was used: the Update, the Start, the Remove, the backup. Lifting the hold used the operator CLI --clear-restore-hold plus a controller restart (the only exit that exists today). An attempt to decrypt the app's secrets so compose could be run by hand was refused by the session's safety guard; that route was dropped, not worked around — the product's own Start supplies the secrets.


Part 1 — the automatic undo, by hand

Does a product path already load a safety dump back?

Yes, one — and the update never calls it. backup.Manager.rollbackSafetyDump (internal/backup/offbox_reconstitute.go:413) re-applies a pre-restore-* undo copy through ImportDump — but only inside ReconstituteFromOffsite. runGuardedUpdate → failAndHold (internal/stacks/update.go:766) writes the safety dump in phase 3 and never reads it again. And failAndHold deletes the journal's pre-update definition copies (removePreUpdateCopies), so after a hold the old definition survives only in the recovery unit's compose/ directory. It was there in all three cases because a backup preceded each update; that is not guaranteed.

The four cases — each a REAL migration that then failed a deliberately wrong probe

case edge (drill) what migrated the safety dump old version on the migrated data, nothing loaded the load undo → healthy seed A (before backup) seed B (after backup)
PostgreSQL docmost 0.95.0 → 0.96.0, held 95.3 s 4 migrations, 42 → 48 tables (docmost-31) 135 816 B, DB only, holds B (tier copy does not) REFUSES — corrupted migrations: previously executed migration 20260824T211732-page-title-trgm-index is missing product semantics FAILS (below); fixed load 1.38 s 14.8 s yes yes
MariaDB romm 5.0.0 → 5.3.0, held 102.6 s alembic 0095 → 0128, 27 → 39 tables 62 943 B, DB only, holds B REFUSES — Can't locate revision identified by '0128_hltb_main_story_column' product semantics 1.25 s, rc 0 — 12 new tables left behind 36.5 s yes yes
Volume data, no DB server vikunja 2.3.0 → 2.6.0, held 93.3 s Ran all migrations successfully (SQLite in a volume) none — the app has no database — nothing to copy (no-op) STARTS AND SERVES in 0.7 s, A, B and B's attachment read back — 0.7 s yes yes
vikunja, if it had refused the only other copy: the tier unit's volume tar — — — volume put back from the tier copy 0.86 s — yes NO — lost
Files on disk vikunja's attachment (a file in the files volume, written after the backup); romm's two drive folders none touched the files the undo copy never holds files the attachment read back after the update AND after the undo — — — attachment yes

Three different apps, three mechanisms of refusal now measured (Nextcloud's version check §4, docmost's migration ledger, RomM's alembic revision) — and in each refusing case the undo made the old version start, because it met the data it knew.

romm's drive folders were empty before, after the update and after the undo (tree hash e3b0c442… throughout) — so they measured nothing, and are reported as unmeasured, not as "untouched".

The finding that shapes decision 15's build: the product's loader cannot undo a migration

ImportDump replays a pg_dump --clean --if-exists file over the live database. Over a database the new version migrated, that FAILS: the new version created six tables (oauth_*, public_spaces, siem_destinations) whose foreign keys point at old tables, and the dump's own DROP … workspaces_pkey is refused — cannot drop constraint workspaces_pkey on table public.workspaces because other objects depend on it — rc 3 in 0.40 s, database unchanged (docmost-45). The same loader on MariaDB "succeeds" (FOREIGN_KEY_CHECKS=0) and leaves the new version's 12 tables behind; RomM 5.0.0 happens to ignore them.

What worked: empty the schema and load the copy in ONE transaction — DROP SCHEMA public CASCADE; CREATE SCHEMA public; + the dump, psql --single-transaction ON_ERROR_STOP=1: rc 0 in 1.38 s, 42 tables, 48 ledger rows, pg_trgm and unaccent back.

The wrong case — a truncated undo copy

engine / loader exit what it left is the outcome honest?
PostgreSQL, product semantics rc 3 unchanged (it failed on the foreign key before reaching the cut) yes — nothing moved; the hold's named copy is intact
PostgreSQL, the fixed atomic load rc 0 (!) 42 tables and 48 ledger rows — and 0 users, 0 spaces, 0 constraints, 0 indexes (docmost-47, measured in a scratch database beside the real one) NO. psql treats end-of-file inside a COPY as end of data and commits. The undo would report success, the old version would start on an EMPTY database, and a health check would pass on it. No hold.
MariaDB, product semantics rc 1 in 0.94 s half-replaced: alembic back to 0095, later tables still the new version's — neither state partly — the load is not transactional; the hold names the tier copy, which can still bring the app back

The check that separates them already exists in the files: the whole PostgreSQL copy ends with -- PostgreSQL database dump complete (and a \unrestrict line), the MariaDB copy with -- Dump completed; the truncated copies have neither. ValidateDump does not look for them — it checks the header and one CREATE TABLE (internal/appbackup/dbdump.go:415, read from source), so it would accept both truncated copies.

Seconds

find the copy old definition + pin back load start → healthy (old probe) undo total
docmost 0.04 s 0.03 s 1.38 s 14.8 s ≈ 16 s
romm 0.04 s 0.04 s 1.25 s 36.5 s ≈ 38 s
vikunja — 0.02 s none 0.7 s ≈ 1 s

Plus the failing health wait that precedes any undo (update.health_timeout, 90 s here, 5 min by default). Lifting the hold by hand cost 15.8 s (CLI + restart) and is NOT part of a product undo, which would never hold in the first place.

What the product must add — the list decision 15's build starts from

  1. Keep the pre-update copies until the undo is over — compose, applied definition, pin and the old .felhom.yml. failAndHold deletes the first three today; the fourth was never kept.
  2. The undo step inside failAndHold: pinBack (exists) → DB service up alone → validated load → full start → health check with the OLD .felhom.yml probe (the new one may name a port the old version does not answer — it did in this spike by construction) → undone, or HOLD.
  3. The loader must empty the database first, atomically. PostgreSQL: drop and recreate exactly the schemas the dump creates, in the same transaction as the load. MariaDB: drop every table first (FOREIGN_KEY_CHECKS=0); its DDL is not transactional, so a failed load is a HOLD with a sentence saying the database is in neither state.
  4. Validate the copy's completion marker before loading, both engines. A truncated PostgreSQL copy is otherwise a silent, successful load of an empty database.
  5. Apps with no database server need their own last-second copy: a tar of the data volumes at safety-dump time (vikunja: 2.2 MB, put back in 0.86 s). Without it, an app whose old version refuses loses everything written since the last backup. Vikunja's old version happens to start — a per-app fact the test record should carry, not a rule.
  6. Files on disk: nothing measured was touched by a migration, and the undo does nothing to files. A step that does rewrite files must say so — decision 13's files may change mark is that place.
  7. The Start button says start completed (200) while the app crash-loops (docmost and romm, both negative controls). The undo's success must be the probe, never the start's return.
  8. No retry: after an undo the badge reads „Frissítés elérhető" again, because the catalog is still ahead. The caller must remember the failed step (decision 15), or it presses it every night.

Part 2 — the ladder

Today's behaviour, measured (51-part2-ladder-today.txt)

vikunja installed at 2.3.0; the drill catalog then got step B = 2.4.0 (610ed1a66dee) and step C = 2.5.0 (c71807f5d47e); one sync, one press. Result: done in 9.5 s, pin, installed record, live compose and container all vikunja/vikunja:2.5.0; 2.4.0 never ran. The box jumps.

Can the box see the steps? No (50-part2-clone-depth.txt)

Both 9202 (drill) and 9201 (live) hold a shallow clone of depth 1: is-shallow-repository true, rev-list --count HEAD = 1, one commit visible per template. Source agrees: sync.go:283 clones --depth 1, sync.go:300 fetches --depth 1. The brief's premise that "the box keeps a git clone" with history is wrong — it keeps a clone of the newest commit only.

Size is not the cost. The whole catalog history is 529 KiB packed, 286 commits.

The cost is that history does not contain the tested steps. romm's compose has 16 commits, 3 of which move an image; the step that works today is 15f9ebf's images with f4eb94f's template (two workers, 768M) — the commit that moved the image is the one that OOM-looped on demo-hp. A box walking history would apply the definition that is known to be broken.

option what it is cost
A — update_ladder: in .felhom.yml (recommended) per app, one entry per step: from and to refs per service, the digest per ref (decision 17), the test record (verdict, date, harness version, memory peak), the marks (files_may_change, needs_person: "<why>"), and — for every step that is NOT the last — the step's own complete definition in templates/<app>/steps/<to>.yml. The last step is the current docker-compose.yml. every image move adds an entry (the catalog gate enforces it: no entry, no move — decision 13); an intermediate definition that needs a fix is fixed in its step file too, or the step is retired; entries older than the support window are pruned. .felhom.yml already flows to frozen apps (§5.4), so the box sees the ladder while frozen; step files are read from the clone.
B — the box walks the catalog's git history deepen the clone, treat each image-moving commit as a step every box carries full history; a history rewrite breaks every box; a commit is not a tested step and cannot be made one retroactively; and it applies the broken intermediate definition measured above

The test record and the marks of decision 13 live in the same entry, so one gate reads one place.


Part 3 — the memory watch in the harness (app-catalog-felhom.eu/scripts/upgrade-test.py v2)

After an edge reads back, the harness runs the new version for --soak seconds (default 600) under four light callers and samples every 15 s, per container: memory against the compose limit, the kernel's own oom_kill counter read host-side (/sys/fs/cgroup/system.slice/docker-<id>.scope/ memory.events — works on images with no shell, and does not depend on Docker's OOMKilled flag, which has read false for real kills on this kind of guest), the peak, and restarts. A kill or a restart → failed; a peak above 80 % → mark memory_tight.

Run on 9202 in /opt/upg with raw compose (no controller in the path), after Part 1's apps were removed so the fixed container_names could not collide.

edge template verdict memory
M1old — the red-proof as promoted, catalog 15f9ebf: 512M, four workers failed seeded, migrated (2 lines), read back — then first kernel OOM kill at +76 s, peak 512 MiB = 100 %, restarts 0 (the container kept running, which is how it hid on demo-hp), Docker's OOMKilled read true here; abort refuses
M1 — the positive control current: 768M, two workers proven + mark memory_tight 608.5 s under 11 429 requests (5 712 × 200, 5 717 × 401): 0 kernel OOM kills, 0 restarts, peak 621 MiB = 81 % of 768 MiB — the watch passes the fix and still flags the thin headroom R-635 left open

Is ten minutes long enough? For this failure, under load, yes by a wide margin: +76 s. demo-hp's first kill came two hours in only because nothing was loading the app. What ten minutes cannot see is a leak that grows over days; that is a monitoring question (R-636), not a test-bench one.

Evidence: harness/M1old.log, harness/evidence/M1old/ (verdict, memory samples, logs), and the same for M1.


Controls and teardown

  • 02-three-controls.txt — 9202 followed the drill catalog; 9201 stayed on the live one; the live catalog's main was 02844ae0a579 before and after.
  • 60-part1-teardown.txt — docmost, romm, vikunja removed through the product; romm's drive data was kept by the product (R-442's fail-closed refusal on 9202, as in every earlier drill) and removed by name at the end.