09 decision 43: calcom -> 18; R-704 crash-loop hold survives reinstall
gates / gates (push) Successful in 25s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-28 11:49:50 +02:00
parent e129476750
commit 2af6896348
3 changed files with 8 additions and 5 deletions
+2 -2
View File
@@ -15,8 +15,8 @@
> would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`**
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
> **2026-09-28 — decided by CC unattended (operator may reverse): `09` §3 decision 43 — claper → PostgreSQL 17** (its
> upstream runs 15; decision 42's rule). Also that day: golden 0.276.0 baked + vouched with agent 0.137.0; controller
> **2026-09-28 — decided by CC unattended (operator may reverse): `09` §3 decision 43 — claper → PostgreSQL 17, calcom → 18** (by
> each upstream; decision 42's rule; calcom's memory 768M → 1536M first, R-703 closed). Also that day: golden 0.276.0 baked + vouched with agent 0.137.0; controller
> v0.277.0 (kept data loads from the off-site copy, R-691 (2)); floor 0.277.0; R-702 (claper default admin, P1) and
> R-703 (calcom OOM at 768 MB) filed.
@@ -477,14 +477,16 @@ R-636's louder repeated alarm.
two WHEN the upstream runs the newer major — paperless's does (18, Django ~5.2.5, psycopg 3); tandoor's does not,
and 17 is the newest major its Django documents, with no mount change. Read 2026-09-27 from each upstream's
compose and requirements (`audits/version-travel-2026-09-26/B/`). Reversible: the catalog pins the major.
43. **claper → PostgreSQL 17; calcom not moved** — *decided by CC unattended 2026-09-28 — operator may reverse.*
43. **claper → PostgreSQL 17, calcom → 18** — *decided by CC unattended 2026-09-28 — operator may reverse.*
*One sentence:* which major does claper's conversion target? **Options:** (a) 17 — no mount change; (b) 18 —
the mount moves to `/var/lib/postgresql`. **Costs:** (a) a later second conversion if claper's upstream moves to
18; (b) claper would run a major its own upstream compose (`postgres:15`, v2.5.0) has not run. **Why (a):**
decision 42's rule — the newer major only when the app's own upstream runs it; claper's does not. Proven on both
venues (bench: 32 tables equal, peaks 75.8 %/78.3 %; box 9202: converted through the guarded Update in 7.2 s),
catalog `4a249b9`, `audits/pg-calcom-claper-2026-09-28/`. **calcom** is not moved: its FROM version does not
start at the catalog's memory limit (R-703), so no conversion can be proven. Reversible: the catalog pins the major.
catalog `4a249b9`, `audits/pg-calcom-claper-2026-09-28/`. **calcom → 18** by the same rule: its upstream compose
runs untagged `postgres` at `/var/lib/postgresql` (Prisma 6.16.1 in v6.2.0). First its memory limit had to be fixed
(768M → 1536M, R-703); then both venues proved it (bench: 122 tables equal, `anon` 61.5 %; box: converted through the
guarded Update in 68 s), catalog `037f956`. Reversible: the catalog pins the major.
---
+1
View File
@@ -815,6 +815,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-701** | **[P3-LOW] demo-hp's whole-guest restore test can never run: every 6 h it picks the right archive and the space preflight refuses it.** Read 2026-09-28 (agent 0.137.0, `audits/version-travel-2026-09-26/D1/D1-cycle-demo-hp.txt`): 20:13 and 02:13 CEST both skipped the golden file and the deleted guest 9100's archive (R-689 working), chose `felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z` (21.6 GiB), and were refused — "needs 31.0 GiB free, has 21.1 GiB" on `local-lvm` — logged `ERROR scheduled restore-test FAILED`. Same refusal first seen 2026-09-24 (R-672's delivery). So demo-hp's whole-guest tier is never proven, and the refusal is SAFE (nothing created). Not measured: whether each refusal reaches the hub or the operator as a failure. **Options (decide nothing yet):** (a) reclaim thin-pool space (`pct fstrim`, R-444) and see if 31 GiB frees; (b) restore-test into `nvme-scratch` instead of `local-lvm` — a config change on the host; (c) accept: demo-hp is a small box, record the tier as not testable there. **-- 2026-09-28 option (a) MEASURED — NOT ENOUGH:** `pct fstrim` 9201 + 9202 from the host (rc 0): `local-lvm` 62.32 % → **50.60 %**, free 20.3 → **26.6 GiB** — still under the 31 GiB the preflight needs. The pool is 53.9 GiB and guest 9201 itself holds ~26 GiB, so no trim can reach 31 GiB. The agent's own verdict after the trim: `tier=felhom-pbs due=true` (archive 2026-09-24T20:06:25Z, not proven). `nvme-scratch` has ~820 GiB free. **Left to the operator: (b) or (c).** `audits/evidence-golden-0276-2026-09-28/phaseD1-reclaim.txt` | **OPEN — P3; owner: operator ((b) or (c)); (a) measured, not enough** |
| **R-702** | **[P1-HIGH] Every claper install creates an admin `admin@claper.co` with the public password `claper`, and the app is published on the household's domain.** Measured 2026-09-28 on scratch guest 9202 (catalog `claper` template, `ghcr.io/claperco/claper:2.5` = 2.5.1): the image's own start command runs `Claper.Release.seeds`, which logs `Created default admin user: Email: admin@claper.co`; asked through claper's own CLI (`bin/claper rpc`), `get_user_by_email_and_password("admin@claper.co", "claper")` answered **true**, an unknown e-mail answered false (control). The template routes `<sub>.<domain>` through traefik and the tunnel, so any claper a household installs can be logged into by anyone who knows claper's README. Not measured: whether any box runs claper today (R-632 lists it as never deployed), whether upstream reads an env var for the seed admin. **Needs:** a decision on the fix shape — (a) the catalog passes a generated admin password (if upstream supports it), (b) the controller changes the seeded admin's password after the first start, (c) pull claper from the catalog until (a)/(b). Evidence: `audits/pg-calcom-claper-2026-09-28/box/C0-claper-default-admin.txt`. | **OPEN — P1; owner: operator (fix shape), CC implements** |
| **R-703** | **[P2] calcom v6.2.0 cannot start at its catalog memory limit — a fresh install crash-loops and the box stops it.** Measured 2026-09-28 on 9202: install from the live catalog → `crash_loop — 6 in 10m0s; STOPPING it (decision 28)`; one Start later, the container's own cgroup counted `oom_kill 1` per start at `memory.max` 805306368 (768 MiB) while `anon` reached ~700 MB during `turbo run start` (`signal: 'SIGKILL'`); Docker reported `OOMKilled=false` (R-528's shape). So calcom in the live catalog cannot run on any box. Not measured: the limit it needs. **Needs:** a measured limit (a bench watch at 1.5–2 GiB), then the catalog change. Until then calcom's PostgreSQL move is `inconclusive — the FROM version does not run`. Evidence: `audits/pg-calcom-claper-2026-09-28/box/C0-calcom-crash.txt`, `C0-calcom-memory.txt`. **-- 2026-09-28 later: FIXED in the catalog (`9555e73`, alone in its commit): memory 768M → 1536M.** Measured on 9202 through the drill catalog: at 2048M a 12-minute watch read `anon` steady ~780 MiB, 0 kills; at 1536M, sampled every 2 s from the container's birth, `anon` peaked at **817 MiB (53 %)** during start, `memory.peak` 1075 MiB, 0 kills, healthy. 1024M would sit at the 80 % `memory_tight` line. The seed route works at the new limit (`Calcom` fixture, catalog `b35fc7f`). `…/box/R703-0*.txt` | **CLOSED — catalog `9555e73`, 2026-09-28** |
| **R-704** | **[P3-LOW] The box's crash-loop stop (decision 28) outlives the app: after remove and reinstall, the new install is still held.** Measured 2026-09-28 on 9202: calcom crash-looped at 08:22 and 08:28 (`unhealthy_stop`, `crash_loop`, trip 2, recorded 08:28:59Z); it was then REMOVED through the product twice and installed fresh twice (09:14:51Z the last). At 09:45 the new, healthy install's Update was refused `409 held` with the crash-loop sentence („…újra és újra összeomlott…"), and `GET /api/stacks/calcom` carried the old `hold_reason` while `state=running`. Start lifted it (`the unhealthy stop is LIFTED by Start`). So a household that removes a crash-looping app and installs it again (the obvious fix) finds its updates refused for a crash of a previous install. Not measured: whether the nightly update leg also skips it; whether other holds (restore hold) behave the same. **Fix direction:** the remove clears the app's box-set holds, as `DeleteAppBackupPrefs` clears its backup preferences (R-474). Evidence: `audits/pg-calcom-claper-2026-09-28/box/calcom/hold.txt`, `…/box/calcom/move.txt`. | **OPEN — P3; owner: CC** |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.