4c92beab8f
gates / gates (push) Successful in 26s
- Part 1 (09 §6.1a, audit): the undo performed by hand on 9202 for docmost (PostgreSQL), romm (MariaDB) and vikunja (SQLite volume) - all three came back with data written before AND after the backup. The product's loader cannot do it: over a migrated PG database it fails on the new tables' foreign keys; over MariaDB it leaves them behind. A truncated PG copy loads rc 0 into an empty database. No-DB apps have no last-second copy. - Part 2: one press jumps A -> C; the box's catalog clone is depth 1. Ladder format recommended: update_ladder in .felhom.yml, not git history. - Part 3: memory watch red-proof results (harness change in the catalog repo). - Part 4 (09 §6.4): ten parts, ~22 evenings; one open point (R-643). - Rows R-637..R-644 opened; R-446/450/451/462/463 updated. STATUS, CONTEXT. No product code. Live catalog untouched; 9202 back on it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
179 lines
14 KiB
Markdown
179 lines
14 KiB
Markdown
# Update rulings 2026-09-23 — the undo spike, the ladder spike, the memory watch
|
||
|
||
Venue: scratch guest **9202** on demo-hp, controller v0.262.1, pointed at the **drill catalog**
|
||
(`admin/app-catalog-drill`, reset to live `02844ae0a579` first) with `update.health_timeout: 90s`.
|
||
The live catalog carried no test reference at any point (control 3 in `02-three-controls.txt`).
|
||
Architecture document for the area: `architecture/09-update-architecture.md` (§3 decisions 11–18,
|
||
§6.1a, §6.4).
|
||
|
||
**Method honesty.** The product has no undo path, so the undo was performed BY HAND, in the order
|
||
the product would take it, each step timed. Where a product path exists it was used: the Update, the
|
||
Start, the Remove, the backup. Lifting the hold used the operator CLI `--clear-restore-hold` plus a
|
||
controller restart (the only exit that exists today). **An attempt to decrypt the app's secrets so
|
||
compose could be run by hand was refused by the session's safety guard; that route was dropped, not
|
||
worked around** — the product's own Start supplies the secrets.
|
||
|
||
---
|
||
|
||
## Part 1 — the automatic undo, by hand
|
||
|
||
### Does a product path already load a safety dump back?
|
||
|
||
**Yes, one — and the update never calls it.** `backup.Manager.rollbackSafetyDump`
|
||
(`internal/backup/offbox_reconstitute.go:413`) re-applies a `pre-restore-*` undo copy through
|
||
`ImportDump` — but only inside `ReconstituteFromOffsite`. `runGuardedUpdate` → `failAndHold`
|
||
(`internal/stacks/update.go:766`) writes the safety dump in phase 3 and never reads it again.
|
||
**And `failAndHold` deletes the journal's pre-update definition copies** (`removePreUpdateCopies`),
|
||
so after a hold the old definition survives only in the recovery unit's `compose/` directory. It was
|
||
there in all three cases because a backup preceded each update; that is not guaranteed.
|
||
|
||
### The four cases — each a REAL migration that then failed a deliberately wrong probe
|
||
|
||
| case | edge (drill) | what migrated | the safety dump | old version on the migrated data, nothing loaded | the load | undo → healthy | seed A (before backup) | seed B (after backup) |
|
||
|---|---|---|---|---|---|---|---|---|
|
||
| **PostgreSQL** docmost | 0.95.0 → 0.96.0, held 95.3 s | 4 migrations, 42 → 48 tables (`docmost-31`) | 135 816 B, DB only, **holds B** (tier copy does not) | **REFUSES** — *corrupted migrations: previously executed migration 20260824T211732-page-title-trgm-index is missing* | product semantics **FAILS** (below); fixed load **1.38 s** | **14.8 s** | yes | **yes** |
|
||
| **MariaDB** romm | 5.0.0 → 5.3.0, held 102.6 s | alembic 0095 → 0128, 27 → 39 tables | 62 943 B, DB only, **holds B** | **REFUSES** — *Can't locate revision identified by '0128_hltb_main_story_column'* | product semantics **1.25 s**, rc 0 — **12 new tables left behind** | **36.5 s** | yes | **yes** |
|
||
| **Volume data, no DB server** vikunja | 2.3.0 → 2.6.0, held 93.3 s | *Ran all migrations successfully* (SQLite in a volume) | **none** — *the app has no database — nothing to copy (no-op)* | **STARTS AND SERVES** in 0.7 s, A, B and B's attachment read back | — | 0.7 s | yes | **yes** |
|
||
| vikunja, **if it had refused** | the only other copy: the tier unit's volume tar | — | — | — | volume put back from the tier copy **0.86 s** | — | yes | **NO — lost** |
|
||
| **Files on disk** | vikunja's attachment (a file in the files volume, written after the backup); romm's two drive folders | none touched the files | the undo copy never holds files | the attachment read back after the update AND after the undo | — | — | — | attachment yes |
|
||
|
||
Three different apps, three mechanisms of refusal now measured (Nextcloud's version check §4, docmost's
|
||
migration ledger, RomM's alembic revision) — and in each refusing case **the undo made the old version
|
||
start**, because it met the data it knew.
|
||
|
||
**romm's drive folders were empty before, after the update and after the undo** (tree hash
|
||
`e3b0c442…` throughout) — so they measured nothing, and are reported as unmeasured, not as "untouched".
|
||
|
||
### The finding that shapes decision 15's build: the product's loader cannot undo a migration
|
||
|
||
`ImportDump` replays a `pg_dump --clean --if-exists` file over the live database. **Over a database
|
||
the new version migrated, that FAILS**: the new version created six tables (`oauth_*`,
|
||
`public_spaces`, `siem_destinations`) whose foreign keys point at old tables, and the dump's own
|
||
`DROP … workspaces_pkey` is refused — *cannot drop constraint workspaces_pkey on table
|
||
public.workspaces because other objects depend on it* — rc 3 in 0.40 s, database unchanged
|
||
(`docmost-45`). The same loader on MariaDB "succeeds" (`FOREIGN_KEY_CHECKS=0`) and leaves the new
|
||
version's **12 tables** behind; RomM 5.0.0 happens to ignore them.
|
||
|
||
**What worked:** empty the schema and load the copy **in ONE transaction** —
|
||
`DROP SCHEMA public CASCADE; CREATE SCHEMA public;` + the dump, `psql --single-transaction
|
||
ON_ERROR_STOP=1`: rc 0 in 1.38 s, 42 tables, 48 ledger rows, `pg_trgm` and `unaccent` back.
|
||
|
||
### The wrong case — a truncated undo copy
|
||
|
||
| engine / loader | exit | what it left | is the outcome honest? |
|
||
|---|---|---|---|
|
||
| PostgreSQL, product semantics | rc 3 | unchanged (it failed on the foreign key before reaching the cut) | yes — nothing moved; the hold's named copy is intact |
|
||
| **PostgreSQL, the fixed atomic load** | **rc 0 (!)** | **42 tables and 48 ledger rows — and 0 users, 0 spaces, 0 constraints, 0 indexes** (`docmost-47`, measured in a scratch database beside the real one) | **NO.** psql treats end-of-file inside a `COPY` as end of data and commits. The undo would report success, the old version would start on an EMPTY database, and a health check would pass on it. No hold. |
|
||
| MariaDB, product semantics | rc 1 in 0.94 s | **half-replaced**: alembic back to 0095, later tables still the new version's — neither state | partly — the load is not transactional; the hold names the tier copy, which can still bring the app back |
|
||
|
||
**The check that separates them already exists in the files:** the whole PostgreSQL copy ends with
|
||
`-- PostgreSQL database dump complete` (and a `\unrestrict` line), the MariaDB copy with
|
||
`-- Dump completed`; the truncated copies have neither. **`ValidateDump` does not look for them** — it
|
||
checks the header and one `CREATE TABLE` (`internal/appbackup/dbdump.go:415`, read from source), so it
|
||
would accept both truncated copies.
|
||
|
||
### Seconds
|
||
|
||
| | find the copy | old definition + pin back | load | start → healthy (old probe) | undo total |
|
||
|---|---|---|---|---|---|
|
||
| docmost | 0.04 s | 0.03 s | 1.38 s | 14.8 s | **≈ 16 s** |
|
||
| romm | 0.04 s | 0.04 s | 1.25 s | 36.5 s | **≈ 38 s** |
|
||
| vikunja | — | 0.02 s | none | 0.7 s | **≈ 1 s** |
|
||
|
||
Plus the failing health wait that precedes any undo (`update.health_timeout`, 90 s here, 5 min by
|
||
default). Lifting the hold by hand cost 15.8 s (CLI + restart) and is NOT part of a product undo,
|
||
which would never hold in the first place.
|
||
|
||
### What the product must add — the list decision 15's build starts from
|
||
|
||
1. **Keep the pre-update copies until the undo is over** — compose, applied definition, pin **and the
|
||
old `.felhom.yml`**. `failAndHold` deletes the first three today; the fourth was never kept.
|
||
2. **The undo step inside `failAndHold`:** `pinBack` (exists) → DB service up alone → validated load →
|
||
full start → **health check with the OLD `.felhom.yml` probe** (the new one may name a port the old
|
||
version does not answer — it did in this spike by construction) → `undone`, or HOLD.
|
||
3. **The loader must empty the database first, atomically.** PostgreSQL: drop and recreate exactly
|
||
the schemas the dump creates, in the same transaction as the load. MariaDB: drop every table first
|
||
(`FOREIGN_KEY_CHECKS=0`); its DDL is not transactional, so a failed load is a HOLD with a sentence
|
||
saying the database is in neither state.
|
||
4. **Validate the copy's completion marker before loading**, both engines. A truncated PostgreSQL copy
|
||
is otherwise a silent, successful load of an empty database.
|
||
5. **Apps with no database server need their own last-second copy**: a tar of the data volumes at
|
||
safety-dump time (vikunja: 2.2 MB, put back in 0.86 s). Without it, an app whose old version
|
||
refuses loses everything written since the last backup. Vikunja's old version happens to start —
|
||
**a per-app fact the test record should carry**, not a rule.
|
||
6. **Files on disk:** nothing measured was touched by a migration, and the undo does nothing to files.
|
||
A step that does rewrite files must say so — decision 13's *files may change* mark is that place.
|
||
7. **The Start button says `start completed` (200) while the app crash-loops** (docmost and romm, both
|
||
negative controls). The undo's success must be the probe, never the start's return.
|
||
8. **No retry:** after an undo the badge reads „Frissítés elérhető" again, because the catalog is still
|
||
ahead. The caller must remember the failed step (decision 15), or it presses it every night.
|
||
|
||
---
|
||
|
||
## Part 2 — the ladder
|
||
|
||
### Today's behaviour, measured (`51-part2-ladder-today.txt`)
|
||
|
||
vikunja installed at **2.3.0**; the drill catalog then got step **B = 2.4.0** (`610ed1a66dee`) and
|
||
step **C = 2.5.0** (`c71807f5d47e`); one sync, one press. **Result: `done` in 9.5 s, pin, installed
|
||
record, live compose and container all `vikunja/vikunja:2.5.0`; 2.4.0 never ran.** The box jumps.
|
||
|
||
### Can the box see the steps? No (`50-part2-clone-depth.txt`)
|
||
|
||
Both 9202 (drill) and 9201 (live) hold a **shallow clone of depth 1**: `is-shallow-repository true`,
|
||
`rev-list --count HEAD` = **1**, one commit visible per template. Source agrees:
|
||
`sync.go:283` clones `--depth 1`, `sync.go:300` fetches `--depth 1`. **The brief's premise that "the
|
||
box keeps a git clone" with history is wrong** — it keeps a clone of the newest commit only.
|
||
|
||
### The format — two options, one recommended
|
||
|
||
**Size is not the cost.** The whole catalog history is 529 KiB packed, 286 commits.
|
||
|
||
**The cost is that history does not contain the tested steps.** romm's compose has **16 commits, 3
|
||
of which move an image**; the step that works today is *15f9ebf's images with f4eb94f's template*
|
||
(two workers, 768M) — **the commit that moved the image is the one that OOM-looped on demo-hp**. A box
|
||
walking history would apply the definition that is known to be broken.
|
||
|
||
| option | what it is | cost |
|
||
|---|---|---|
|
||
| **A — `update_ladder:` in `.felhom.yml` (recommended)** | per app, one entry per step: `from` and `to` refs per service, the digest per ref (decision 17), the test record (verdict, date, harness version, memory peak), the marks (`files_may_change`, `needs_person: "<why>"`), and — for every step that is NOT the last — the step's own complete definition in `templates/<app>/steps/<to>.yml`. The last step is the current `docker-compose.yml`. | every image move adds an entry (the catalog gate enforces it: no entry, no move — decision 13); an intermediate definition that needs a fix is fixed in its step file too, or the step is retired; entries older than the support window are pruned. `.felhom.yml` already flows to frozen apps (§5.4), so the box sees the ladder while frozen; step files are read from the clone. |
|
||
| B — the box walks the catalog's git history | deepen the clone, treat each image-moving commit as a step | every box carries full history; a history rewrite breaks every box; a commit is not a tested step and cannot be made one retroactively; **and it applies the broken intermediate definition measured above** |
|
||
|
||
The test record and the marks of decision 13 live in the same entry, so one gate reads one place.
|
||
|
||
---
|
||
|
||
## Part 3 — the memory watch in the harness (`app-catalog-felhom.eu/scripts/upgrade-test.py` v2)
|
||
|
||
After an edge reads back, the harness runs the new version for `--soak` seconds (default 600) under
|
||
four light callers and samples every 15 s, per container: memory against the compose limit, the
|
||
**kernel's own `oom_kill` counter read host-side** (`/sys/fs/cgroup/system.slice/docker-<id>.scope/
|
||
memory.events` — works on images with no shell, and does not depend on Docker's OOMKilled flag, which
|
||
has read false for real kills on this kind of guest), the peak, and restarts. A kill or a restart →
|
||
`failed`; a peak above 80 % → mark `memory_tight`.
|
||
|
||
Run on 9202 in `/opt/upg` with raw compose (no controller in the path), after Part 1's apps were
|
||
removed so the fixed `container_name`s could not collide.
|
||
|
||
| edge | template | verdict | memory |
|
||
|---|---|---|---|
|
||
| **M1old** — the red-proof | as promoted, catalog `15f9ebf`: 512M, four workers | **failed** | seeded, migrated (2 lines), read back — then **first kernel OOM kill at +76 s**, peak 512 MiB = **100 %**, restarts **0** (the container kept running, which is how it hid on demo-hp), Docker's OOMKilled read true here; abort `refuses` |
|
||
| **M1** — the positive control | current: 768M, two workers | **proven** + mark `memory_tight` | 608.5 s under 11 429 requests (5 712 × 200, 5 717 × 401): **0 kernel OOM kills, 0 restarts**, peak 621 MiB = **81 %** of 768 MiB — the watch passes the fix and still flags the thin headroom R-635 left open |
|
||
|
||
**Is ten minutes long enough?** For this failure, under load, yes by a wide margin: +76 s. demo-hp's
|
||
first kill came two hours in only because nothing was loading the app. What ten minutes cannot see is
|
||
a leak that grows over days; that is a monitoring question (R-636), not a test-bench one.
|
||
|
||
Evidence: `harness/M1old.log`, `harness/evidence/M1old/` (verdict, memory samples, logs), and the same
|
||
for M1.
|
||
|
||
---
|
||
|
||
## Controls and teardown
|
||
|
||
- `02-three-controls.txt` — 9202 followed the drill catalog; 9201 stayed on the live one; the live
|
||
catalog's `main` was `02844ae0a579` before and after.
|
||
- `60-part1-teardown.txt` — docmost, romm, vikunja removed through the product; romm's drive data was
|
||
kept by the product (R-442's fail-closed refusal on 9202, as in every earlier drill) and removed by
|
||
name at the end.
|