- R-732: the first-start geodata import runs up to 9 concurrent 5000-row INSERTs; the database needs ~400 MB anon + ~170 MB touched shared_buffers (the image's FIXED 512MB, not host-RAM sizing). 512M fits only with swap (bench swap 0: 61-104 kills; 9202 swap 512 MiB: survived by swapping). Controls: swap alone, limit alone flip it; shared_buffers 128MB alone does not. Catalog 56c4888: v3.2.4 + 768M, proven with swap off on both venues. audits/immich-first-start-2026-09-30/A-cause.md. - R-730: scripts/iso/build-felhom-iso.sh refuses an uncommitted/untracked/unpushed tree (no bypass), records repo-commit from the gate and iso-v<version>; test iso/test/clean-tree.sh, red-proof run (status check removed -> 2 of 4 cases fail -> restored). - R-731 narrowed (gitea 28.0.0 GA; mariadb 13.0 a short-term Rolling line). R-676 note. - New rows R-733 (bench has no swap, boxes 512 MiB), R-734 (immich .immich markers -> files_may_change). - STATUS: the golden line corrected (no bake is due; 0.283.1 is the newest release). Register 364 -> 366. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
5.3 KiB
Why immich's database was killed on its first start (R-732) — the cause, measured
2026-09-30, by day. Bench: LXC 9401 on demo-hp (6 cores, 10 GiB, swap 0), rebuilt for this. Box: scratch guest 9202 on
demo-hp (7 cores, 25 GiB, swap 512 MiB). Both disks are raw images on the same nvme-scratch store; both run cgroup v2.
Image: ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0 at the ladder's digest 1a078b23… (PostgreSQL 16.10).
Sampler: tools/pgsample.sh — the database container's memory.current, memory.stat (anon, file, shmem, file_dirty,
file_writeback) and memory.events (oom, oom_kill) every 2 s from the container's birth. Every run is a fresh install.
The mechanism, in one paragraph
On its very first start (and again whenever an update ships a new geodata file), immich imports ~228 000 places into
geodata_places in 5 000-row INSERTs, up to 9 at once (server/src/repositories/map.repository.ts v3.2.2 L284–321:
if (futures.length >= 9) await Promise.all(futures)). Each INSERT is its own database process. Together they need
~400 MB of their own memory (anon), on top of ~170 MB of shared_buffers the import touches (the image's own
/etc/postgresql/postgresql.conf fixes shared_buffers = 512MB, work_mem = 16MB — S1-image-conf.txt, S1-settings.txt).
About 575 MB does not fit a 512 MB limit. With swap the kernel pushes ~70–110 MB of it out and the import survives at the
limit (that is why 9202 passed this morning and at 15:27). Without swap the kernel kills a database process — always
the one running insert into "geodata_places" (S1-postgres.log) — PostgreSQL restarts every backend, the server's worker
dies, and the cycle repeats (61–104 kills in 7–15 min).
The claims the brief listed, each tested
| candidate | test | result |
|---|---|---|
| 1. the image sizes memory from the HOST's RAM | pg_settings.source + the image's config source |
false — shared_buffers 512MB and work_mem 16MB come from the image's fixed postgresql.conf (configuration file /etc/postgresql/postgresql.conf); the rest are PostgreSQL defaults. The container sees the host's RAM (30.7 GB on the bench, 25.9 GB on the box) but nothing reads it |
| 2. the import needs more than 512M at the image's defaults, on any host | limit changed alone | true without swap: 1024M → 0 kills (peak anon 403 MB); 768M → 0 kills twice (412/394 MB) |
3a. shared_buffers |
512MB → 128MB alone (command: flag), limit 512M, no swap |
false — still 104 kills; the import finished only after 14 min of retries (S3) |
| 3b. dirty page cache (slow writeback) | file_dirty sampled |
false — 2–13 MB dirty at the kills (S4) |
3c. /dev/shm |
read | 64M, 1.1M used — not involved |
| bench vs box = host RAM | configs | false — same host; the difference is swap (box 512 MiB, bench 0) |
| swap | the bench given 512 MiB swap alone, limit 512M | true — 0 kills, swap.peak 74.7 MB (S6); on the box swap.peak 119 MB, pswpout 27 669 pages (box/S2-box-positive.txt) |
The two curves (peaks; the full 2-second series are in the files)
| run | limit | swap | peak current | peak anon | oom_kill | outcome |
|---|---|---|---|---|---|---|
| bench S1 — positive control, R-732 repeated | 512M | 0 | 512M | 420M | 84 (in 631 s) | never healthy |
| bench S4 — positive again, with dirty pages | 512M | 0 | 511M | 421M | 24 (in 186 s) | never healthy |
| box S2 — same template, fresh | 512M | 512 MiB | 511M | 280M (+~108 MB swapped) | 0 | import done in 20 s, at 99 % of the limit |
| bench S3 — shared_buffers 128MB only | 512M | 0 | 512M | 434M | 104 | import after ~14 min |
| bench S6 — swap only | 512M | 512 MiB | 511M | 320M | 0 | healthy, import 8.2 s |
| bench S5 — limit only | 1024M | 0 | 956M | 403M | 0 | healthy, import 8.3 s |
| bench S7, S8 — the fix | 768M | 0 | 765M / 759M | 412M / 394M (54 % / 51 %) | 0 / 0 | healthy, import 7.7 / 7.9 s |
Control pair: S1/S4 (fail) vs S6 (only swap changed → pass) and vs S5/S7/S8 (only the limit changed → pass); S3 (only
shared_buffers changed → still fails). The cause is named on the variables that flip the result.
Why settings cannot fix it
The 400 MB is nine concurrent INSERT backends; their memory is the statement's own (5 000 rows × 11 columns), not
work_mem or shared_buffers. immich exposes no knob for the import's concurrency. Lowering shared_buffers saves at most
the ~40–60 MB actually touched and was measured not to be enough (S3). So the limit rises: 768M is the smallest step with
room (anon 51–54 %, under the 80 % memory_tight line; shmem ~170 MB and file cache fill the rest, reclaimable).
Not measured, stated
- Whether a real customer guest has swap: both demo guests have 512 MiB; the golden's guest config is not recorded here. The fix does not depend on it (proven with swap off on both venues).
- The PostgreSQL 14 line upstream runs uses the SAME
postgresql.ssd.conf(base-images copies one file for every major) — the same 512MBshared_buffers; its import behaviour would be the same. Note only. - The chaos-night restarts of 2026-09-17 (R-676) fit this mechanism (a DB connection dropped during the first-start import on a 6 GB guest) — consistent, not re-measured.