immich's first start: cause measured, fixed in the catalog (R-732 closed); ISO clean-tree gate (R-730 closed)
gates / gates (push) Successful in 27s
gates / gates (push) Successful in 27s
- R-732: the first-start geodata import runs up to 9 concurrent 5000-row INSERTs; the database needs ~400 MB anon + ~170 MB touched shared_buffers (the image's FIXED 512MB, not host-RAM sizing). 512M fits only with swap (bench swap 0: 61-104 kills; 9202 swap 512 MiB: survived by swapping). Controls: swap alone, limit alone flip it; shared_buffers 128MB alone does not. Catalog 56c4888: v3.2.4 + 768M, proven with swap off on both venues. audits/immich-first-start-2026-09-30/A-cause.md. - R-730: scripts/iso/build-felhom-iso.sh refuses an uncommitted/untracked/unpushed tree (no bypass), records repo-commit from the gate and iso-v<version>; test iso/test/clean-tree.sh, red-proof run (status check removed -> 2 of 4 cases fail -> restored). - R-731 narrowed (gitea 28.0.0 GA; mariadb 13.0 a short-term Rolling line). R-676 note. - New rows R-733 (bench has no swap, boxes 512 MiB), R-734 (immich .immich markers -> files_may_change). - STATUS: the golden line corrected (no bake is due; 0.283.1 is the newest release). Register 364 -> 366. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -803,7 +803,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-652** | **[P3-LOW] The memory watch counted the kernel's file cache as the app's memory.** MEASURED 2026-09-23 night on the bench: nextcloud 34.0.4 read **100 %** of its 1 GiB and immich's PostgreSQL **100 %**, each with **0** kernel `oom_kill`s — `memory.peak` includes page cache, which the kernel drops before it kills anything. Under the watch as built (R-635 follow-up) both would be marked `memory_tight`, and the gate would demand a raised `mem_limit` — a customer-box capacity figure — for cache. **Done the same night (09 §3 decision 22, CC — operator may reverse):** the watch samples the app's own memory (`anon` of `memory.stat`) every 15 s; the mark and the ladder's `memory_peak_pct` read it where measured; the cgroup peak stays beside it (`memory_cgroup_peak_pct`). **Open:** romm's backfilled entry carries M1's 80.9 % cgroup peak (measured before the anon sample existed) — re-measure it on its next step; and decide whether an app whose anon is low but whose cgroup stays pinned at its limit (cache thrash) should be marked at all. Evidence: `audits/night-2026-09-23/apps/nextcloud/bench-1024M/`, `apps/immich/bench-noanon/`. | **READY — P3; owner: CC (catalog harness)** |
|
||||
| **R-654** | **[P3-LOW] opengist 1.15 moved every page under `/-/` — a household's `/login` bookmark answers 404 after the update.** MEASURED 2026-09-23 night: 1.13 serves `/login`, `/register`, `/all`; 1.15.2 answers **404** on all three and serves `/-/login`, `/-/register`, `/-/all`; `/` redirects to `/-/all`. The app, its data and its probe (`/healthcheck`) are fine, and a household arriving at the root lands correctly — only a deep link breaks. 1.15 also marks its session cookie `Secure`. **Needs:** a line in opengist's `app_info` if the operator wants households told; nothing in the product. Evidence: `apps/opengist-oldfixture/`, `apps/opengist/`. | **READY — P3; owner: operator (copy decision) / CC (writes it)** |
|
||||
| **R-675** | **[P3-LOW] The unit-only restore's refusal for a file app still points to „Fájlok visszaállítása" instead of the second drive's whole restore.** `missingFileLegsRefusal` predates decision 26 (v0.269.0); when a whole copy exists on the second drive the sentence should name it. | **READY — P3; owner: CC (controller)** |
|
||||
| **R-676** | **[P3-LOW] Watch: immich's first start restarted 12 times — decision 28's crash-loop stop (6 in 10 min) would stop it.** From the 2026-09-17 chaos night (DB connection dropped during the first-start geocoding import on a 6 GB guest; it did not recover that night). No healthy app in any drill evidence restarts on a first start (1831 samples, 40 live containers), so the threshold stands; this row exists so the first immich install under v0.269.x is watched. `audits/night-2026-09-24/A3/40-first-start-restarts.txt` **2026-09-25 night (read from source, v0.271.0): a DEPLOY's first start is NOT covered by decision 28's suppression** — `Deploying` clears when `compose up -d` returns (`deploy.go` "Clear deploying flag"), and `ObserveUnhealthy` then samples the app; an automatic update's step, verify and undo ARE covered (`Updating`, pinned by `TestD28_NoCrashLoopStopDuringAnAutomaticStep`). So a first start that restarts ≥ 6 times in 10 min is stopped — which R-676 already accepts for a broken first start; a healthy slow first start would be stopped too. | **OPEN — P3; owner: CC (watch)** |
|
||||
| **R-676** | **[P3-LOW] Watch: immich's first start restarted 12 times — decision 28's crash-loop stop (6 in 10 min) would stop it.** From the 2026-09-17 chaos night (DB connection dropped during the first-start geocoding import on a 6 GB guest; it did not recover that night). No healthy app in any drill evidence restarts on a first start (1831 samples, 40 live containers), so the threshold stands; this row exists so the first immich install under v0.269.x is watched. `audits/night-2026-09-24/A3/40-first-start-restarts.txt` **2026-09-25 night (read from source, v0.271.0): a DEPLOY's first start is NOT covered by decision 28's suppression** — `Deploying` clears when `compose up -d` returns (`deploy.go` "Clear deploying flag"), and `ObserveUnhealthy` then samples the app; an automatic update's step, verify and undo ARE covered (`Updating`, pinned by `TestD28_NoCrashLoopStopDuringAnAutomaticStep`). So a first start that restarts ≥ 6 times in 10 min is stopped — which R-676 already accepts for a broken first start; a healthy slow first start would be stopped too. **-- 2026-09-30: the first-start restarts are explained.** immich's first-start geodata import OOM-kills its database at 512M on a guest with no swap (R-732, measured: 61–104 kills); the 2026-09-17 chaos-night case (DB connection dropped during the import on a 6 GB guest) fits it. Fixed in the catalog (`56c4888`, 768M). The watch itself (decision 28 on a DEPLOY's first start) is unchanged. |**OPEN — P3; owner: CC (watch)** |
|
||||
| **R-682** | **[P3-LOW] A Remove interrupted by a controller kill leaves the app half-removed: containers gone, the app still listed as installed (and held).** MEASURED 2026-09-24 on 9202 (chaos round 9): the kill 2 s after the Remove press answered the household `502 Bad Gateway`; after the restart `chaoscrash` read deployed, stopped, `unhealthy_stop`, with NO container left. Pressing Remove again completed it cleanly (200, only the catalog template left). Recoverable by the household's own second press; nothing tells them to press it. **Fix direction:** the remove journals its intent and finishes (or says it was interrupted) at boot, as the update does. `audits/night-2026-09-24/E/round-09*.json`, `E/round-09b-remove-again.txt` | **READY — P3; owner: CC (controller)** |
|
||||
| **R-683** | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** |
|
||||
| **R-687** | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent), and **did not occur on the demo boxes' real night either** (2026-09-25: both legs ended by 04:19, before the gate opened at 04:30, and no whole-box backup was due on either box) — unit + red-proof only (`TestD20_GateWaitsForTheLeg`). **Also cosmetic:** a leg with no steps reports `"steps": null` to the hub, not `[]`. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` **-- NARROWED 2026-09-25 (controller v0.273.0):** the cosmetic `"steps": null` → `[]` and the taken `files_may_change` step's missing log line are FIXED (red-proofed, `audits/night-2026-09-26/F/`). Items (1)–(4) stay; (4) did not occur on 2026-09-25 either (demo-felhom's whole-box backup ran at 07:29, three hours after its leg; demo-hp had none due). **-- 2026-09-28 (night 27/28):** (4) did not occur again — on demo-hp the leg ended 04:23:54 and the whole-guest backup began 04:37:06, after the gate opened at 04:30; demo-felhom's backup ran at 07:36 (`audits/evidence-golden-0276-2026-09-28/phaseD2-night-read.txt`). **-- 2026-09-30 (by day, demo-hp 9201): item (4) PROVEN LIVE.** The night chain pressed by hand, the window moved to W = now − 2h05m the moment the leg started, `quiesce.poll_interval` 1m: `[quiesce] full-system backup due and inside its window, but the automatic update leg is running … deferring` at 11:35:11 and 11:36:11 UTC while bookstack (55.1 s) and kimai (75.1 s) stepped; the leg's end line at 11:36:29; the backup quiesced at 11:37:11 (the first poll after), job done 11:47:19, the agent's `backup: completed` 9.98 GB. Config and window put back and read back (`audits/pg-last-six-2026-09-30/C/`). **Found, cosmetic, manual chain only:** the deferral names the moved window's W+5h (16:29) while the manual leg's own deadline was its start + the leg length (16:49). | **OPEN — P3, gaps (1)–(3) + the manual-chain deferral text; item (4) proven live 2026-09-30; owner: CC** |
|
||||
@@ -841,9 +841,11 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-727** | **[P2-MEDIUM] The whole-guest restore test picks a PREVIOUS box's archive, fails on its key, and the household sees a bare ✗ labelled with the wrong tier.** MEASURED 2026-09-30 01:52 UTC on the drill box: the agent's restore test chose `felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z` („newest settled archive … has not been proven") — written by a drill box of 2026-09-16, still in the customer's ep0 namespace next to tonight's own `2026-09-29T19:37:07Z` — and failed `wrong key - unable to verify signature since manifest's key 6b:ca:5f:3f… does not match provided key de:51:7a:18…`. One operator mail (`restore_test_failed`); the hub then re-logged the stored failure at every 15-minute report for 5 hours. The household's backups page read „✗ Visszaállítás ellenőrizve 2026-09-30 03:52 — Helyi tároló (local)" — the local tier had not failed; the pbs tier had, and the page says neither which nor why. Root: host delete leaves the old box's archives in the namespace (R-526's shape). **Fix direction:** the restore test skips (and reports as foreign) archives whose key fingerprint is not this host's; the page names the tier that failed. **FIXED 2026-09-30 (agent v0.138.0 + decision 51):** measured — a PBS archive carries its key FINGERPRINT (PVE content `encrypted`), not a host id; the storage carries its own (`encryption-key`). The restore test skips an archive written with another key (logged by name). The ✗ card names the tier (controller v0.283.0) — **and the 2026-09-30 claim "the page blames the local tier" was my misreading**: the „Helyi tároló (local)" after the ✗ was the next section's heading. ep0: the three drill archives in `tester-1` removed, other namespaces byte-identical. Delivered by signed jobs to both demo boxes (340 s); their due-check reads normally on 0.138.0. Evidence: `audits/evidence-fixes-first-tester-2026-09-30/``partC/`. | **CLOSED 2026-09-30 — agent v0.138.0** |
|
||||
| **R-728** | **[P3-LOW] A customer created with one press was created TWICE, and the first of its two connect mails holds a dead link.** MEASURED 2026-09-30 on `Tester-2`: the hub logged `Customer config created: Tester-2` twice in the same second and two self-bind mints (hashes `c40df008…`, `6a1cbef4…`); a mint replaces the previous link (single-active), so one of the two identical mails the tester received answers „expired". Cause not established (a double form submit, or the handler run twice). **Fix direction:** make the create idempotent within a few seconds (or disable the button on submit), and pin it. The workaround for the tester is in STATUS. | **READY — rank P3-LOW; owner: CC (hub)** |
|
||||
| **R-729** | **[P3-LOW] An off-site target, once saved on the page, cannot be removed through the product.** MEASURED 2026-09-30 on 9202: `/backup/offbox/config` refuses an empty address and no route clears the target; the session removed its throwaway target from `settings.json` by hand, with the controller stopped (harness teardown on a scratch guest). A household that tries its own NAS and gives up keeps a disabled target forever. **Fix direction:** a „Távoli mentési cél törlése" press that clears the target (never the repository). | **READY — rank P3-LOW; owner: CC (controller)** |
|
||||
| **R-730** | **[P3-LOW] The published ISO 1.29.0 was built from a tree that git cannot name, so it cannot be tagged.** MEASURED 2026-09-30 from the build record: the manifest says `repo-commit 8d539f97` (2026-09-18 17:48), built 20:41:45; the menu-width fix the published image carries (`grub-release.cfg.tmpl`, „… / text") was committed 35 minutes AFTER the build, in `31eeb36` (21:16). `scripts/iso/build-felhom-iso.sh:498` records `git rev-parse HEAD` with no dirty-tree check, so the manifest names a commit that is not the image. **And the brief's tag name would have been wrong:** the `installer-v*` tags are the host-install SCRIPT's versions (`SCRIPT_VERSION` 1.28.0 on main; the website serves `/scripts/` from `installer-v1.28.0`), a different number line from the ISO — an `installer-v1.29.0` tag would block the next script release. **Needs:** the ISO build refuses a dirty tree (the workspace clean-tree gate, applied to this script) and records `iso-v<version>`; optionally prove 1.29.0's source by content (the ISO's `grub.cfg` against `31eeb36`) and tag that. `audits/pg-last-six-2026-09-30/E/E4-installer-tag.txt` | **READY — rank P3-LOW; owner: CC (scripts/iso)** |
|
||||
| **R-731** | **[P3-LOW] The catalog-currency comparison cannot see an upstream that changes its TAG SHAPE, and two newest tags need a release check.** MEASURED 2026-09-30 (`audits/catalog-currency-2026-09-30.md` §2 item 5): the same-shape rule read three apps as up to date that were not — gramps-web (`v25.6.0` → upstream dropped the `v`, at `26.9.1`), jellyfin (`10.11.11` → two-part `12.1`), kimai (`apache-2.57.0` → plain `2.67.0`; the plain tag's digest equals `apache`'s, so kimai moved on 2026-09-30). A control pass over every shape caught them. Not checked: whether `mariadb:13.0` and `gitea/gitea:28.0.0` (pushed 2026-09-30 00:15 UTC) are general releases. **Needs:** the currency script's shape-switch control made standing (it is in the audit's tools today), and the two release checks before either is walked. | **READY — rank P3-LOW; owner: CC (catalog audit tooling)** |
|
||||
| **R-732** | **[P2-MEDIUM] immich's FIRST start at today's catalog pin could not finish on the bench: its database container was OOM-killed at 512 MiB.** MEASURED 2026-09-30, twice (the second run alone on the bench): at FROM (v3.2.2, `…/postgres:16-vectorchord0.4.3-pgvectors0.2.0` at the ladder's own digest `1a078b23…`), the server's first-start reverse-geocoding import dropped its DB connection („terminating connection because of crash of another server process"), and the harness's 420 s settle never saw it healthy; a live read of the postgres cgroup showed `oom 413, oom_kill 36`, `OOMKilled true`. **The same template on 9202 the same hour** (a fresh install, kept data moved aside) seeded in 35 s and its v3.2.4 step ended `done` — but that first container was replaced by the update, so its kill counter is gone: whether a household's fresh install on a smaller box hits this is NOT measured. R-676's watch (a first start restarting ≥ 6 times in 10 min is stopped by decision 28) makes the consequence a stopped app. So immich's v3.2.4 step stays unpublished (box proven, bench inconclusive twice — Part F's stop rule). **Needs:** a sampled first-start memory watch of immich-postgres on a fresh install (bench and box, `anon` + `oom_kill` from the container's birth), then a measured `mem_limit`. `audits/pg-last-six-2026-09-30/F/immich-first-start-oom.txt`, `F/benchq-q6.txt`, `F/benchq-q10.txt` | **READY — rank P2-MEDIUM; owner: CC (catalog)** |
|
||||
| **R-730** | **[P3-LOW] The published ISO 1.29.0 was built from a tree that git cannot name, so it cannot be tagged.** MEASURED 2026-09-30 from the build record: the manifest says `repo-commit 8d539f97` (2026-09-18 17:48), built 20:41:45; the menu-width fix the published image carries (`grub-release.cfg.tmpl`, „… / text") was committed 35 minutes AFTER the build, in `31eeb36` (21:16). `scripts/iso/build-felhom-iso.sh:498` records `git rev-parse HEAD` with no dirty-tree check, so the manifest names a commit that is not the image. **And the brief's tag name would have been wrong:** the `installer-v*` tags are the host-install SCRIPT's versions (`SCRIPT_VERSION` 1.28.0 on main; the website serves `/scripts/` from `installer-v1.28.0`), a different number line from the ISO — an `installer-v1.29.0` tag would block the next script release. **Needs:** the ISO build refuses a dirty tree (the workspace clean-tree gate, applied to this script) and records `iso-v<version>`; optionally prove 1.29.0's source by content (the ISO's `grub.cfg` against `31eeb36`) and tag that. `audits/pg-last-six-2026-09-30/E/E4-installer-tag.txt` **-- 2026-09-30: BUILT.** `scripts/iso/build-felhom-iso.sh` now runs a clean-tree gate right after argument parsing: any uncommitted or untracked change, or HEAD ≠ `origin/main`, refuses the build (no bypass flag); the manifest records `repo-commit` from the gate and `iso-version-tag : iso-v<version>`. Test `scripts/iso/test/clean-tree.sh` (4 cases, throwaway repo via the `FELHOM_ISO_REPO` seam); **red-proof run:** the status check deleted → cases 2 and 4 FAIL → restored → all pass. `test/rootpw-emission.sh` now builds from the same kind of throwaway clean repo (it cannot run on DooPlex either way: the assistant image is absent, 9 FAILs before and after). ISO 1.29.0 stays untagged: its source cannot be proven. | **CLOSED 2026-09-30 — the gate + test; 1.29.0 untagged by decision (its build tree is not a commit)** |
|
||||
| **R-731** | **[P3-LOW] The catalog-currency comparison cannot see an upstream that changes its TAG SHAPE, and two newest tags need a release check.** MEASURED 2026-09-30 (`audits/catalog-currency-2026-09-30.md` §2 item 5): the same-shape rule read three apps as up to date that were not — gramps-web (`v25.6.0` → upstream dropped the `v`, at `26.9.1`), jellyfin (`10.11.11` → two-part `12.1`), kimai (`apache-2.57.0` → plain `2.67.0`; the plain tag's digest equals `apache`'s, so kimai moved on 2026-09-30). A control pass over every shape caught them. Not checked: whether `mariadb:13.0` and `gitea/gitea:28.0.0` (pushed 2026-09-30 00:15 UTC) are general releases. **Needs:** the currency script's shape-switch control made standing (it is in the audit's tools today), and the two release checks before either is walked. **-- 2026-09-30: the two release checks answered** (`audits/immich-first-start-2026-09-30/D/D2-release-checks.txt`): **gitea v28.0.0** is a general release (GitHub: not prerelease, not draft, 2026-09-29) — upstream renumbered 1.27.x → 28; **mariadb 13.0** is `Stable` but a short-term `Rolling` line (no EOL date), while 12.3 and 11.8 are the LTS lines — so a move of any of the four MariaDB apps to 13.0 would leave LTS. Not done: the shape-switch control made standing. | **NARROWED 2026-09-30 — only the standing shape-switch control is left; owner: CC** |
|
||||
| **R-732** | **[P2-MEDIUM] immich's FIRST start at today's catalog pin could not finish on the bench: its database container was OOM-killed at 512 MiB.** MEASURED 2026-09-30, twice (the second run alone on the bench): at FROM (v3.2.2, `…/postgres:16-vectorchord0.4.3-pgvectors0.2.0` at the ladder's own digest `1a078b23…`), the server's first-start reverse-geocoding import dropped its DB connection („terminating connection because of crash of another server process"), and the harness's 420 s settle never saw it healthy; a live read of the postgres cgroup showed `oom 413, oom_kill 36`, `OOMKilled true`. **The same template on 9202 the same hour** (a fresh install, kept data moved aside) seeded in 35 s and its v3.2.4 step ended `done` — but that first container was replaced by the update, so its kill counter is gone: whether a household's fresh install on a smaller box hits this is NOT measured. R-676's watch (a first start restarting ≥ 6 times in 10 min is stopped by decision 28) makes the consequence a stopped app. So immich's v3.2.4 step stays unpublished (box proven, bench inconclusive twice — Part F's stop rule). **Needs:** a sampled first-start memory watch of immich-postgres on a fresh install (bench and box, `anon` + `oom_kill` from the container's birth), then a measured `mem_limit`. `audits/pg-last-six-2026-09-30/F/immich-first-start-oom.txt`, `F/benchq-q6.txt`, `F/benchq-q10.txt` **-- 2026-09-30 (afternoon): CAUSE MEASURED and FIXED.** immich's first-start geodata import runs up to 9 concurrent 5000-row INSERTs; the database then needs ~400 MB anon + ~170 MB touched shared_buffers (the image's own `postgresql.conf` fixes 512MB — NOT sized from host RAM). At 512M without swap it is OOM-killed (bench: 61–104 kills); with 512 MiB swap it survives by swapping ~70–110 MB (why 9202 passed). Controls, one variable each: swap alone → 0 kills; limit 1024M alone → 0; `shared_buffers` 128MB alone → still 104 kills. **An update that ships a new geodata file re-runs the import** (v3.2.2 → v3.2.4: 228 294 → 228 571 places), so the fix had to ride the step. Catalog `56c4888`: immich v3.2.4 + `immich-postgres` 768M (`mem_limit` 4096M → 4480M; +256 MB per immich box). Proven with swap OFF: fresh install at the new definition on the bench ×2 (anon 409/412 MB = 53 %, 0 kills) and on 9202 (368 MB = 48 %, 0 kills); the step on both venues (bench proven, 0 kills in the 10-min watch; box done 58.5 s, album read back, running limit 768M). `audits/immich-first-start-2026-09-30/A-cause.md`. | **CLOSED 2026-09-30 — catalog `56c4888` (768M + v3.2.4); an installed immich gets it with its next guarded Update** |
|
||||
| **R-733** | **[P3-LOW] The test bench has NO swap and the boxes have 512 MiB — so a box proof can pass on swap where the bench fails, and nobody records whether a customer guest has swap.** MEASURED 2026-09-30 (R-732): immich's first start was OOM-killed 61–104 times on the bench (swap 0) and passed on 9202 by swapping ~108 MB; the bench given 512 MiB swap passed too. demo-hp 9201, 9202 and demo-felhom 9201 all read `swap: 512`; the golden's guest config is not recorded in its bake evidence, so a customer guest's swap is NOT measured. The harness's memory watch judges `anon` against the limit and never reads `memory.swap.current`. **Needs:** the golden's `swap` read and recorded; the box walk and the harness report `memory.swap.peak` beside `anon`; a decision whether proofs run with swap off (the stricter venue, as R-732's fix was proven). | **READY — rank P3-LOW; owner: CC (harness + golden evidence)** |
|
||||
| **R-734** | **[P3-LOW] The harness marks immich `files_may_change` because immich rewrites six 13-byte `.immich` folder markers at every start.** MEASURED 2026-09-30 on the bench (v3.2.2 → v3.2.4): the bind-tree hash of `appdata/immich` changed; the only changed files were `{encoded-video,library,backups,profile,thumbs,upload}/.immich`, rewritten at each start — no household file. The mark is honest by the harness's rule and the ladder writer copies it (never edited by hand), so immich's v3.2.4 night step needs a fresh WHOLE copy (decision 13); on a box without one the night leg skips it and a person presses. The 2026-09-23 immich entry did not carry it (`files_changed []`). **Needs:** a decision whether app-owned marker files are excluded from the file hash (a per-template ignore list, or a size/name rule), or the mark stays. | **READY — rank P3-LOW; owner: CC (harness); the rule change needs a word** |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
|
||||
Reference in New Issue
Block a user