- Part 1 (09 §6.1a, audit): the undo performed by hand on 9202 for docmost (PostgreSQL), romm (MariaDB) and vikunja (SQLite volume) - all three came back with data written before AND after the backup. The product's loader cannot do it: over a migrated PG database it fails on the new tables' foreign keys; over MariaDB it leaves them behind. A truncated PG copy loads rc 0 into an empty database. No-DB apps have no last-second copy. - Part 2: one press jumps A -> C; the box's catalog clone is depth 1. Ladder format recommended: update_ladder in .felhom.yml, not git history. - Part 3: memory watch red-proof results (harness change in the catalog repo). - Part 4 (09 §6.4): ten parts, ~22 evenings; one open point (R-643). - Rows R-637..R-644 opened; R-446/450/451/462/463 updated. STATUS, CONTEXT. No product code. Live catalog untouched; 9202 back on it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
8.9 KiB
REPORT — update arc: the operator's rulings recorded, the undo and ladder spiked, the memory watch, the build plan
2026-09-23. Repos touched: felhom.eu (docs, register, STATUS, CONTEXT, evidence),
app-catalog-felhom.eu (scripts/ only — the memory watch), admin/app-catalog-drill (drill
commits, reset to live main at the end). felhom-controller, felhom-agent, hub: read only.
Architecture read first and named: documentation/architecture/09-update-architecture.md (all of it),
07-backup-architecture.md §6.
Baselines verified live before starting: controller b9deec19077b, agent d9864a94bf62, felhom.eu
267dcad01bcf, catalog 02844ae0a579 — all equal to the brief.
1. Not done, or changed from the brief — first
| item | state |
|---|---|
Part 0 — rulings into 09 |
done, commit 805ad1e (documents only) |
| Part 1 — the undo, four cases + the wrong case | done, with three changes. (a) The fourth case, "files on disk", was measured on vikunja's attachment (a file in a volume), not on a bind-mounted drive folder: romm's two drive folders stayed EMPTY throughout — they measured nothing and are reported as unmeasured. (b) Starting the app by hand needed its decrypted secrets; the session's safety guard refused that, and it was not worked around. The product's own Start was used instead, after lifting the hold with the operator CLI + a controller restart. (c) The wrong case was run on BOTH engines, and on PostgreSQL with both the product's loader and the fixed one — the fixed one produced the session's most important finding (a truncated copy loads rc 0). |
| Part 2 — the ladder | done |
| Part 3 — the memory watch + red-proof | done; harness v2. C3, the harness's standing negative control, was NOT run: its template's container_name: privatebin collides with the privatebin the controller runs on 9202. The memory watch has its own pair instead — M1old must fail, M1 must pass. |
| Part 4 — the build plan | done, 09 §6.4 — with one open point for the operator the brief did not expect (§6) |
Claims in the brief that turned out wrong:
- "A safety dump exists for every app class" — false. An app with no database server gets none
(
update safety dump for vikunja: the app has no database — nothing to copy (no-op)), measured. - "The box keeps a git clone of the catalog" with history — false. Depth 1 on both demo
guests,
rev-list --count HEAD= 1;sync.go:283/:300clone and fetch--depth 1. - "
stacks.update_windowis unread" — true, and the grep was widened fromconfig.go+setup/handlers.goto the whole controller repo: the only other hits areconfigs/controller.yaml.exampleand the i18n base file. No Go code reads it. - The register held 329 row lines by
grep -c '^| \*\*R-', not 326; the highest id was R-636 as stated.
2. Part 0 — the rulings
09 §3 gains decisions 11–18 in the existing shape; decision 3's second half and §6.1's abort
paragraph are marked REPLACED with pointers; §4 says why the undo is not a rollback; §3b is kept,
headed ANSWERED, each question pointing at its decision; §6.2 rewritten to the ruled shape; the slices
table updated. Register: R-450, R-451, R-446, R-463 cite the decisions.
3. Part 1 — the undo, by hand
Full evidence and tables: documentation/audits/update-rulings-2026-09-23/README.md.
| docmost (PostgreSQL) | romm (MariaDB) | vikunja (volume, no DB server) | |
|---|---|---|---|
| held after | 95.3 s | 102.6 s | 93.3 s |
| old version on migrated data | refuses | refuses | starts |
| safety dump | 135 816 B, DB only | 62 943 B, DB only | none |
product loader (ImportDump semantics) |
FAILS, rc 3, foreign keys | rc 0, 12 tables left behind | — |
| fixed load (empty schema + copy, one transaction) | rc 0, 1.38 s | (product loader sufficed) | — |
| undo → healthy on the OLD probe | ≈ 16 s | ≈ 38 s | ≈ 1 s |
| data before / after the backup | yes / yes | yes / yes | yes / yes (attachment too) |
Does a product path load a safety dump back? Yes — rollbackSafetyDump, but only the off-site
restore calls it; the update never reads its own dump, and failAndHold deletes the pre-update
definition copies.
The wrong case: PostgreSQL + product loader → rc 3, nothing changed (honest hold). PostgreSQL +
the fixed atomic loader + a half-length copy → rc 0 and an EMPTY database (0 users, 0 constraints,
0 indexes) that still shows 42 tables and 48 ledger rows — no hold, a dishonest success. MariaDB +
half copy → rc 1, half the tables already replaced (not atomic). The whole copies carry an end marker
the truncated ones lack; ValidateDump does not check it.
4. Part 2 — the ladder
One press on vikunja two steps behind: 2.3.0 → 2.5.0 in 9.5 s; 2.4.0 never ran. The box cannot see
2.4.0 (depth-1 clone). Recommended format: update_ladder: in .felhom.yml, intermediate steps with
their own definition — not git history, because romm's image-moving commit is the definition that
OOM-looped on demo-hp. Full comparison in the audit.
5. Part 3 — the memory watch
upgrade-test.py v2: after a successful readback, --soak seconds (default 600) of light load
(4 callers), sampling every 15 s the kernel's oom_kill counter read host-side from the container's
cgroup, the peak, the limit, the restarts and Docker's OOMKilled flag. Kill or restart → failed;
peak > 80 % → mark memory_tight. New Romm fixture; edges M1 (current template) and M1old
(the template as promoted, 15f9ebf).
- M1old (red-proof):
failed— first OOM kill at +76 s, peak 512 MiB = 100 % of the limit, restarts 0 (the container kept running — the shape that hid it on demo-hp), abortrefuses. Ten minutes is ample for this failure under load. - M1 (positive control): proven + mark
memory_tight— 608.5 s under 11 429 requests (5 712 × 200, 5 717 × 401): 0 kernel OOM kills, 0 restarts, peak 621 MiB = 81 % of 768 MiB — the watch passes the fix and still flags the thin headroom R-635 left open.
6. Part 4 — the build plan
09 §6.4: ten parts, ≈ 22 evenings, recommended order undo → sentences in the household's language
- notifier honesty → test record + gate + memory check → ladder → digests → the automatic leg → R-636 → R-625 → PostgreSQL conversion; fleet view deferred by ruling. One open point needs the operator (R-643): as ruled, the update leg sits between the off-site leg (W+105m) and the full-system gate (W+2h) — at most 15 minutes a night. Recommendation: the full-system backup waits for the leg inside its own four-hour window.
7. Rows
Opened (8): R-637 (build the undo), R-638 (the loader cannot replay over a newer schema — and the restore the hold names is UNMEASURED after a real schema migration), R-639 (pre-update copies deleted on hold), R-640 (a truncated PostgreSQL copy loads rc 0 into an empty database), R-641 (no-DB apps have no last-second copy), R-642 (Start returns 200 over a crash loop), R-643 (the ≤15-minute leg, operator), R-644 (gokapi crash-looping on 9202 at session start, not caused here). Updated: R-446, R-450, R-451, R-462, R-463. Closed: none. 329 → 337.
8. Teardown — three layers
| layer | state |
|---|---|
| machine — guest 9202 | controller.yaml restored from the saved copy and read back identical (live catalog, no update: block); catalog cache re-cloned from the live repo (02844ae); docmost, romm, vikunja removed through the product (no containers, no volumes); romm's drive folder (kept by the product, R-442) removed by name; /opt/upg and every temp file in /root removed; the eight test images removed by name — vikunja:2.4.0 was absent, i.e. never pulled, which is the jump seen from a second side; no prune. Containers afterwards: the same three apps as at the start (gokapi still crash-looping — R-644, pre-existing). 82-teardown-guest.txt |
| host — demo-hp | nothing provisioned; only transient /tmp files, removed |
| hub | nothing touched — no hub call was made |
| drill repo | reset to live main 02844ae0a579; has_actions: false; image lines identical to live. The local clone's push URL to the LIVE catalog was disabled at the start. 81-teardown-drill-repo.txt |
One instrumentation slip, recorded: a background watcher and the first teardown both wrote the same temporary script file on demo-hp at the same moment, so the first teardown never ran (its output file held the watcher's lines). Caught by reading the file, re-run after the watcher ended; the second run is the one recorded.
Fences: DooPlex, Peti's box, ep0, the demo guests' apps, drill-r50, tester-1 and the hub were
not touched. The live catalog's main was 02844ae0a579 before and after.
unproven.py --summary: walked 20 / partial 17 / built 14 / missing 4 — not walked 35 of 55, unchanged.