"Upgrading nextcloud from 34.0.1.2 ..." then the seeded user read back
through occ (a never-created uid reads absent on the same call). The
bench's drive tree changed under the app (its appdata), so the entry is
marked files_may_change: an automatic update needs a fresh copy of the
files (09 decision 13). Memory: nextcloud's own 7.9 % (cgroup 100 % =
file cache), db 24.2 %; 0 kills. Abort refuses ("the version of the data
is higher") — the undo, not the old image, is the way back. Box (9202):
done, read back, R-626 clean. 09 decision 21.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-22's move took immich-server to v3.2.2 and left machine-learning at
v3.0.3; this step aligns them. Bench: the admin + an album (its own API)
read back; memory watch 12 053 requests all 200: server 51.4 %, ML 10.2 %,
postgres 35.1 % own memory (its cgroup peak 100 % is file cache, decision
22); 0 kills; abort starts-and-serves. Box (9202): done, read back, R-626
clean. 09 decision 21.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The engine alone moves (R-450: an engine change gets its own edge). Asked of
the engine, not the log: "already upgraded to 11.8.9-MariaDB" after
MARIADB_AUTO_UPGRADE's mariadb-upgrade ran (Phase 1/8 on the box). Bench:
the seeded user read back; memory: kimai 6 % own (cgroup 100 %, cache),
kimai-db 37.9 %; 0 kills; abort starts-and-serves. Box (9202): done, read
back, R-626 clean. 09 decision 21.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Bench at 512M: proven but memory_tight (its own memory 94.6 %); per the
bar the limit is raised in the same commit and the watch re-run once:
at 768M 60 %, 0 kills, 11 940 requests all 200; abort starts-and-serves.
Box (9202, guarded Update, at 512M): done, read back, R-626 clean.
The komga fixture was fixed tonight (users/me lives at /api/v2).
09 decision 21.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
1.15 (v1.15.2 today; its digest is in the entry) moved every page under /-/
(/login now 404, /-/login serves) and marks its session cookie Secure —
the account read back through its own page, a never-created user 404.
Memory: app 77.8 % of 128M, cgroup peak 100 % (file cache), 0 kills.
Box (9202): done, read back, R-626 clean. R-654 for the /login bookmark.
09 decision 21.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- memory watch: load no longer follows redirects to the unresolvable test
domain (every request had been 'err' on box-fixture apps), and sends the
app's own Host; the app's own memory (anon) is sampled beside the cgroup
peak, and memory_tight reads anon where measured (09 decision 22, CC).
- ladder writer: memory_peak_pct = anon (else cgroup peak), with
memory_basis and memory_cgroup_peak_pct beside it.
- fixtures: opengist 1.15 serves under /-/ and marks its cookie Secure
(readback = the account's own page + a never-created user 404); komga's
user endpoint is /api/v2/users/me; a wishlist fixture (form actions).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
update_ladder: in .felhom.yml, one JSON entry per line (spiked live on
controller v0.266.0 and v0.267.0 first). Two gates: check-test-record.py
(static, CI too) and check-test-record-move.py (history + registry for
moved refs only). 16 decoys, 3 red-proofs. The ONLY writer is
upgrade-test.py --write-ladder (bench AND box proven, digests resolved).
Harness v3: box fixtures on the bench, files_may_change.
Backfill: the 21 moves of 2026-09-22, 21 proven from their records.
No image: line moved.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
After an edge reads back, --soak seconds (default 600) of light load while
the kernel's own oom_kill counter is read host-side from the container's
cgroup. A kill or restart turns proven into failed; a peak over 80% of the
limit adds the memory_tight mark. New Romm fixture; edges M1 / M1old.
Red-proof on scratch 9202: M1old (template as promoted, 512M, 4 workers)
OOM-killed at +76 s -> failed. M1 (current, 768M, 2 workers) proven, 0
kills in 608.5 s, peak 81% -> memory_tight.
Test code only; no template changed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Suite 56 -> 64. Both faults re-introduced one at a time and refused; an explicit container no
service declares refused; a unique prefix still resolves; the name moving in a comment convicts
nothing; and the no-PyYAML mode CI runs is covered both ways.
The unique-prefix case first used paperless-ngx and was WRONG - paperless-webserver does not begin
with paperless-ngx, so the gate was right to convict. Rewritten with immich.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
No image: line moved.
paperless-ngx has no container named after its stack, so its probe had NEVER run on any box.
immich has four immich-* containers and no exact match, so the old first-prefix rule picked
whichever came first - possibly the database.
The gate now resolves the target by the same four rules as findProbeContainerMeta: exact name,
explicit container, a UNIQUE prefix, else refuse - and refusing is right, because verifying waits
on this probe and a successful update of such an app gets stopped.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Walked through the product's own guarded Update on scratch guest 9202 during the
twenty-eight drill, seeded and read back through the app's own front door.
Evidence: felhom.eu/documentation/audits/the-28-2026-09-22/apps/termix/verdict.json
catalog_since -> 2026-09-22 (R-452).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Walked through the product's own guarded Update on scratch guest 9202 during the
twenty-eight drill, seeded and read back through the app's own front door.
Evidence: felhom.eu/documentation/audits/the-28-2026-09-22/apps/sonarr/verdict.json
catalog_since -> 2026-09-22 (R-452).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Walked through the product's own guarded Update on scratch guest 9202 during the
twenty-eight drill, seeded and read back through the app's own front door.
Evidence: felhom.eu/documentation/audits/the-28-2026-09-22/apps/radarr/verdict.json
catalog_since -> 2026-09-22 (R-452).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Walked through the product's own guarded Update on scratch guest 9202 during the
twenty-eight drill, seeded and read back through the app's own front door.
Evidence: felhom.eu/documentation/audits/the-28-2026-09-22/apps/immich/verdict.json
catalog_since -> 2026-09-22 (R-452).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Walked through the product's own guarded Update on scratch guest 9202 during the
twenty-eight drill, seeded and read back through the app's own front door.
Evidence: felhom.eu/documentation/audits/the-28-2026-09-22/apps/ghost/verdict.json
catalog_since -> 2026-09-22 (R-452).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Walked through the product's own guarded Update on scratch guest 9202 during the
twenty-eight drill, seeded and read back through the app's own front door.
Evidence: felhom.eu/documentation/audits/the-28-2026-09-22/apps/emby/verdict.json
catalog_since -> 2026-09-22 (R-452).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
WEB_SERVER_CONCURRENCY=2, limit left at 768M. No image: line moved.
768M was still a guess and it was wrong: kills slowed from ~12/min to ~7/min and stopped nothing.
Measured instead - each warm uvicorn worker holds ~216 MiB, so four plus the master reach ~882 MiB,
and the cgroup's memory.peak read exactly 768 MiB.
/init:143 runs --workers "${WEB_SERVER_CONCURRENCY:-4}". Four is a server default; this is one
household on one small box. Two measure ~450 MiB, fit 768M with headroom, and halve the CPU churn.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
No image: line moved, so no catalog_since moved.
Measured on demo-hp after this morning's promotion: OOMKilled true, 4530 gunicorn worker SIGKILLs
in six hours, ~500% CPU in a permanent restart storm, host load 5.2 while otherwise idle. It ran
clean for two hours first, which is why the walk on the scratch guest did not catch it.
The update reported `done` and the app read `running` the whole time - nginx answers 200 while the
workers behind it die. Nothing alarmed; the operator heard the fans.
768M is a measured first step, not a final answer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
CI job 877 on 15d7c2b failed: the runner has no PyYAML and the gate answered INCONCLUSIVE, which
the runner rightly refuses to call a pass. A gate red on every push is bypassed within a week.
Falls back to a line reader and announces mode: DEGRADED. Over all 53 apps it returns exactly what
the full reader returns. Five more decoy cases with PyYAML shadowed out prove it still convicts the
three real faults; suite now 56 cases.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The ONE engine move in this series, and it is deliberate. Walked through the product's own
guarded Update on scratch guest 9202 during the update night: 217.4 s end to end, seed read back
through nextcloud's own front door before AND after, healthy after.
Permitted since R-469 (Slice 4 shipped): MARIADB_AUTO_UPGRADE=1 is already in this template and
converts the datadir. check-engine-major.py was run against this exact commit and ALLOWS it by
name, rather than being assumed to.
catalog_since -> 2026-09-22 (R-452).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Walked through the product's own guarded Update on scratch guest 9202 during the
update night, seeded and read back through the app's own front door.
Evidence: felhom.eu/documentation/audits/update-night-2026-09-21/apps/vikunja/verdict.json
catalog_since -> 2026-09-22 (R-452).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Walked through the product's own guarded Update on scratch guest 9202 during the
update night, seeded and read back through the app's own front door.
Evidence: felhom.eu/documentation/audits/update-night-2026-09-21/apps/tandoor/verdict.json
catalog_since -> 2026-09-22 (R-452).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS