Files
felhom.eu/REPORT.md
T
admin 4c92beab8f
gates / gates (push) Successful in 26s
Update arc: the undo and ladder spiked, the build plan for the 2026-09-23 rulings
- Part 1 (09 §6.1a, audit): the undo performed by hand on 9202 for docmost
  (PostgreSQL), romm (MariaDB) and vikunja (SQLite volume) - all three came
  back with data written before AND after the backup. The product's loader
  cannot do it: over a migrated PG database it fails on the new tables'
  foreign keys; over MariaDB it leaves them behind. A truncated PG copy loads
  rc 0 into an empty database. No-DB apps have no last-second copy.
- Part 2: one press jumps A -> C; the box's catalog clone is depth 1.
  Ladder format recommended: update_ladder in .felhom.yml, not git history.
- Part 3: memory watch red-proof results (harness change in the catalog repo).
- Part 4 (09 §6.4): ten parts, ~22 evenings; one open point (R-643).
- Rows R-637..R-644 opened; R-446/450/451/462/463 updated. STATUS, CONTEXT.

No product code. Live catalog untouched; 9202 back on it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 08:54:41 +02:00

8.9 KiB
Raw Blame History

REPORT — update arc: the operator's rulings recorded, the undo and ladder spiked, the memory watch, the build plan

2026-09-23. Repos touched: felhom.eu (docs, register, STATUS, CONTEXT, evidence), app-catalog-felhom.eu (scripts/ only — the memory watch), admin/app-catalog-drill (drill commits, reset to live main at the end). felhom-controller, felhom-agent, hub: read only. Architecture read first and named: documentation/architecture/09-update-architecture.md (all of it), 07-backup-architecture.md §6. Baselines verified live before starting: controller b9deec19077b, agent d9864a94bf62, felhom.eu 267dcad01bcf, catalog 02844ae0a579 — all equal to the brief.


1. Not done, or changed from the brief — first

item state
Part 0 — rulings into 09 done, commit 805ad1e (documents only)
Part 1 — the undo, four cases + the wrong case done, with three changes. (a) The fourth case, "files on disk", was measured on vikunja's attachment (a file in a volume), not on a bind-mounted drive folder: romm's two drive folders stayed EMPTY throughout — they measured nothing and are reported as unmeasured. (b) Starting the app by hand needed its decrypted secrets; the session's safety guard refused that, and it was not worked around. The product's own Start was used instead, after lifting the hold with the operator CLI + a controller restart. (c) The wrong case was run on BOTH engines, and on PostgreSQL with both the product's loader and the fixed one — the fixed one produced the session's most important finding (a truncated copy loads rc 0).
Part 2 — the ladder done
Part 3 — the memory watch + red-proof done; harness v2. C3, the harness's standing negative control, was NOT run: its template's container_name: privatebin collides with the privatebin the controller runs on 9202. The memory watch has its own pair instead — M1old must fail, M1 must pass.
Part 4 — the build plan done, 09 §6.4 — with one open point for the operator the brief did not expect (§6)

Claims in the brief that turned out wrong:

  1. "A safety dump exists for every app class" — false. An app with no database server gets none (update safety dump for vikunja: the app has no database — nothing to copy (no-op)), measured.
  2. "The box keeps a git clone of the catalog" with history — false. Depth 1 on both demo guests, rev-list --count HEAD = 1; sync.go:283/:300 clone and fetch --depth 1.
  3. "stacks.update_window is unread" — true, and the grep was widened from config.go + setup/handlers.go to the whole controller repo: the only other hits are configs/controller.yaml.example and the i18n base file. No Go code reads it.
  4. The register held 329 row lines by grep -c '^| \*\*R-', not 326; the highest id was R-636 as stated.

2. Part 0 — the rulings

09 §3 gains decisions 11–18 in the existing shape; decision 3's second half and §6.1's abort paragraph are marked REPLACED with pointers; §4 says why the undo is not a rollback; §3b is kept, headed ANSWERED, each question pointing at its decision; §6.2 rewritten to the ruled shape; the slices table updated. Register: R-450, R-451, R-446, R-463 cite the decisions.

3. Part 1 — the undo, by hand

Full evidence and tables: documentation/audits/update-rulings-2026-09-23/README.md.

docmost (PostgreSQL) romm (MariaDB) vikunja (volume, no DB server)
held after 95.3 s 102.6 s 93.3 s
old version on migrated data refuses refuses starts
safety dump 135 816 B, DB only 62 943 B, DB only none
product loader (ImportDump semantics) FAILS, rc 3, foreign keys rc 0, 12 tables left behind —
fixed load (empty schema + copy, one transaction) rc 0, 1.38 s (product loader sufficed) —
undo → healthy on the OLD probe ≈ 16 s ≈ 38 s ≈ 1 s
data before / after the backup yes / yes yes / yes yes / yes (attachment too)

Does a product path load a safety dump back? Yes — rollbackSafetyDump, but only the off-site restore calls it; the update never reads its own dump, and failAndHold deletes the pre-update definition copies.

The wrong case: PostgreSQL + product loader → rc 3, nothing changed (honest hold). PostgreSQL + the fixed atomic loader + a half-length copy → rc 0 and an EMPTY database (0 users, 0 constraints, 0 indexes) that still shows 42 tables and 48 ledger rows — no hold, a dishonest success. MariaDB + half copy → rc 1, half the tables already replaced (not atomic). The whole copies carry an end marker the truncated ones lack; ValidateDump does not check it.

4. Part 2 — the ladder

One press on vikunja two steps behind: 2.3.0 → 2.5.0 in 9.5 s; 2.4.0 never ran. The box cannot see 2.4.0 (depth-1 clone). Recommended format: update_ladder: in .felhom.yml, intermediate steps with their own definition — not git history, because romm's image-moving commit is the definition that OOM-looped on demo-hp. Full comparison in the audit.

5. Part 3 — the memory watch

upgrade-test.py v2: after a successful readback, --soak seconds (default 600) of light load (4 callers), sampling every 15 s the kernel's oom_kill counter read host-side from the container's cgroup, the peak, the limit, the restarts and Docker's OOMKilled flag. Kill or restart → failed; peak > 80 % → mark memory_tight. New Romm fixture; edges M1 (current template) and M1old (the template as promoted, 15f9ebf).

  • M1old (red-proof): failed — first OOM kill at +76 s, peak 512 MiB = 100 % of the limit, restarts 0 (the container kept running — the shape that hid it on demo-hp), abort refuses. Ten minutes is ample for this failure under load.
  • M1 (positive control): proven + mark memory_tight — 608.5 s under 11 429 requests (5 712 × 200, 5 717 × 401): 0 kernel OOM kills, 0 restarts, peak 621 MiB = 81 % of 768 MiB — the watch passes the fix and still flags the thin headroom R-635 left open.

6. Part 4 — the build plan

09 §6.4: ten parts, ≈ 22 evenings, recommended order undo → sentences in the household's language

  • notifier honesty → test record + gate + memory check → ladder → digests → the automatic leg → R-636 → R-625 → PostgreSQL conversion; fleet view deferred by ruling. One open point needs the operator (R-643): as ruled, the update leg sits between the off-site leg (W+105m) and the full-system gate (W+2h) — at most 15 minutes a night. Recommendation: the full-system backup waits for the leg inside its own four-hour window.

7. Rows

Opened (8): R-637 (build the undo), R-638 (the loader cannot replay over a newer schema — and the restore the hold names is UNMEASURED after a real schema migration), R-639 (pre-update copies deleted on hold), R-640 (a truncated PostgreSQL copy loads rc 0 into an empty database), R-641 (no-DB apps have no last-second copy), R-642 (Start returns 200 over a crash loop), R-643 (the ≤15-minute leg, operator), R-644 (gokapi crash-looping on 9202 at session start, not caused here). Updated: R-446, R-450, R-451, R-462, R-463. Closed: none. 329 → 337.

8. Teardown — three layers

layer state
machine — guest 9202 controller.yaml restored from the saved copy and read back identical (live catalog, no update: block); catalog cache re-cloned from the live repo (02844ae); docmost, romm, vikunja removed through the product (no containers, no volumes); romm's drive folder (kept by the product, R-442) removed by name; /opt/upg and every temp file in /root removed; the eight test images removed by name — vikunja:2.4.0 was absent, i.e. never pulled, which is the jump seen from a second side; no prune. Containers afterwards: the same three apps as at the start (gokapi still crash-looping — R-644, pre-existing). 82-teardown-guest.txt
host — demo-hp nothing provisioned; only transient /tmp files, removed
hub nothing touched — no hub call was made
drill repo reset to live main 02844ae0a579; has_actions: false; image lines identical to live. The local clone's push URL to the LIVE catalog was disabled at the start. 81-teardown-drill-repo.txt

One instrumentation slip, recorded: a background watcher and the first teardown both wrote the same temporary script file on demo-hp at the same moment, so the first teardown never ran (its output file held the watcher's lines). Caught by reading the file, re-run after the watcher ended; the second run is the one recorded.

Fences: DooPlex, Peti's box, ep0, the demo guests' apps, drill-r50, tester-1 and the hub were not touched. The live catalog's main was 02844ae0a579 before and after.

unproven.py --summary: walked 20 / partial 17 / built 14 / missing 4 — not walked 35 of 55, unchanged.