Files
felhom.eu/documentation/audits/immich-first-start-2026-09-30/A-cause.md
T
admin 2866f6a318
gates / gates (push) Successful in 27s
immich's first start: cause measured, fixed in the catalog (R-732 closed); ISO clean-tree gate (R-730 closed)
- R-732: the first-start geodata import runs up to 9 concurrent 5000-row INSERTs; the database needs
  ~400 MB anon + ~170 MB touched shared_buffers (the image's FIXED 512MB, not host-RAM sizing). 512M fits
  only with swap (bench swap 0: 61-104 kills; 9202 swap 512 MiB: survived by swapping). Controls: swap
  alone, limit alone flip it; shared_buffers 128MB alone does not. Catalog 56c4888: v3.2.4 + 768M,
  proven with swap off on both venues. audits/immich-first-start-2026-09-30/A-cause.md.
- R-730: scripts/iso/build-felhom-iso.sh refuses an uncommitted/untracked/unpushed tree (no bypass),
  records repo-commit from the gate and iso-v<version>; test iso/test/clean-tree.sh, red-proof run
  (status check removed -> 2 of 4 cases fail -> restored).
- R-731 narrowed (gitea 28.0.0 GA; mariadb 13.0 a short-term Rolling line). R-676 note.
- New rows R-733 (bench has no swap, boxes 512 MiB), R-734 (immich .immich markers -> files_may_change).
- STATUS: the golden line corrected (no bake is due; 0.283.1 is the newest release). Register 364 -> 366.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-30 16:18:17 +02:00

5.3 KiB
Raw Blame History

Why immich's database was killed on its first start (R-732) — the cause, measured

2026-09-30, by day. Bench: LXC 9401 on demo-hp (6 cores, 10 GiB, swap 0), rebuilt for this. Box: scratch guest 9202 on demo-hp (7 cores, 25 GiB, swap 512 MiB). Both disks are raw images on the same nvme-scratch store; both run cgroup v2. Image: ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0 at the ladder's digest 1a078b23… (PostgreSQL 16.10). Sampler: tools/pgsample.sh — the database container's memory.current, memory.stat (anon, file, shmem, file_dirty, file_writeback) and memory.events (oom, oom_kill) every 2 s from the container's birth. Every run is a fresh install.

The mechanism, in one paragraph

On its very first start (and again whenever an update ships a new geodata file), immich imports ~228 000 places into geodata_places in 5 000-row INSERTs, up to 9 at once (server/src/repositories/map.repository.ts v3.2.2 L284–321: if (futures.length >= 9) await Promise.all(futures)). Each INSERT is its own database process. Together they need ~400 MB of their own memory (anon), on top of ~170 MB of shared_buffers the import touches (the image's own /etc/postgresql/postgresql.conf fixes shared_buffers = 512MB, work_mem = 16MB — S1-image-conf.txt, S1-settings.txt). About 575 MB does not fit a 512 MB limit. With swap the kernel pushes ~70–110 MB of it out and the import survives at the limit (that is why 9202 passed this morning and at 15:27). Without swap the kernel kills a database process — always the one running insert into "geodata_places" (S1-postgres.log) — PostgreSQL restarts every backend, the server's worker dies, and the cycle repeats (61–104 kills in 7–15 min).

The claims the brief listed, each tested

candidate test result
1. the image sizes memory from the HOST's RAM pg_settings.source + the image's config source false — shared_buffers 512MB and work_mem 16MB come from the image's fixed postgresql.conf (configuration file /etc/postgresql/postgresql.conf); the rest are PostgreSQL defaults. The container sees the host's RAM (30.7 GB on the bench, 25.9 GB on the box) but nothing reads it
2. the import needs more than 512M at the image's defaults, on any host limit changed alone true without swap: 1024M → 0 kills (peak anon 403 MB); 768M → 0 kills twice (412/394 MB)
3a. shared_buffers 512MB → 128MB alone (command: flag), limit 512M, no swap false — still 104 kills; the import finished only after 14 min of retries (S3)
3b. dirty page cache (slow writeback) file_dirty sampled false — 2–13 MB dirty at the kills (S4)
3c. /dev/shm read 64M, 1.1M used — not involved
bench vs box = host RAM configs false — same host; the difference is swap (box 512 MiB, bench 0)
swap the bench given 512 MiB swap alone, limit 512M true — 0 kills, swap.peak 74.7 MB (S6); on the box swap.peak 119 MB, pswpout 27 669 pages (box/S2-box-positive.txt)

The two curves (peaks; the full 2-second series are in the files)

run limit swap peak current peak anon oom_kill outcome
bench S1 — positive control, R-732 repeated 512M 0 512M 420M 84 (in 631 s) never healthy
bench S4 — positive again, with dirty pages 512M 0 511M 421M 24 (in 186 s) never healthy
box S2 — same template, fresh 512M 512 MiB 511M 280M (+~108 MB swapped) 0 import done in 20 s, at 99 % of the limit
bench S3 — shared_buffers 128MB only 512M 0 512M 434M 104 import after ~14 min
bench S6 — swap only 512M 512 MiB 511M 320M 0 healthy, import 8.2 s
bench S5 — limit only 1024M 0 956M 403M 0 healthy, import 8.3 s
bench S7, S8 — the fix 768M 0 765M / 759M 412M / 394M (54 % / 51 %) 0 / 0 healthy, import 7.7 / 7.9 s

Control pair: S1/S4 (fail) vs S6 (only swap changed → pass) and vs S5/S7/S8 (only the limit changed → pass); S3 (only shared_buffers changed → still fails). The cause is named on the variables that flip the result.

Why settings cannot fix it

The 400 MB is nine concurrent INSERT backends; their memory is the statement's own (5 000 rows × 11 columns), not work_mem or shared_buffers. immich exposes no knob for the import's concurrency. Lowering shared_buffers saves at most the ~40–60 MB actually touched and was measured not to be enough (S3). So the limit rises: 768M is the smallest step with room (anon 51–54 %, under the 80 % memory_tight line; shmem ~170 MB and file cache fill the rest, reclaimable).

Not measured, stated

  • Whether a real customer guest has swap: both demo guests have 512 MiB; the golden's guest config is not recorded here. The fix does not depend on it (proven with swap off on both venues).
  • The PostgreSQL 14 line upstream runs uses the SAME postgresql.ssd.conf (base-images copies one file for every major) — the same 512MB shared_buffers; its import behaviour would be the same. Note only.
  • The chaos-night restarts of 2026-09-17 (R-676) fit this mechanism (a DB connection dropped during the first-start import on a 6 GB guest) — consistent, not re-measured.