Evidence (rulings 2026-10-01): Parts A-D so far; register: R-743, R-744, R-745, R-746, R-749, R-751 closed
gates / gates (push) Successful in 25s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-01 08:40:04 +02:00
parent b2f6c1626c
commit daacf84e30
23 changed files with 1113 additions and 6 deletions
+6 -6
View File
@@ -854,15 +854,15 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-740** | **[P2-MEDIUM] A security fix that upstream ships under the SAME tag (`postgres:18-alpine`, `redis:7-alpine`, `mariadb:12.3` …) reaches NO box — not at night, and not by the household's Update button either, because the catalog never records a re-test of a tag at a new digest.** MEASURED 2026-09-30 (more-night-apps brief, Part C). **(1) The night leg, from source** (`felhom-controller` `d48da6c`): an app at the head tag whose running digest is older than the ladder's tested one reads Behind (`stacks/updateorder.go:86–111` `digestBehind`), but `legCandidate` finds no entry whose `from` is its pin and skips it — `unattended.go:433–435`, `return LegSkipNoTestRecord // at the head, behind only by something no step records`. A unit walk on a scratch copy of the controller confirms it: order Behind → the leg presses nothing, `skipped — no_test_record`; the household's Update button in the same state ends `done` and runs the tested digest (`audits/more-night-apps-2026-09-30/C/C2-*`). **(2) But the catalog never produces that state for a floating tag:** the only writer refuses a step that moves no image (`upgrade-test.py:919`, „--move: nothing moves"), and no ladder in the catalog has an entry with `from == to` or two entries with the same `to` (`C/C4-same-tag-retest.txt`). So the tested digest of `postgres:18-alpine` stays the one of the day the step was tested, a box installed since runs that digest, and the badge reads „Naprakész" while upstream has moved. **(3) How often it matters:** of the 11 distinct floating lines in the catalog today (26 services), the 8 on Docker Hub were pushed 9, 9, 9, 9, 12, 12, 30 and 105 days ago — **7 of 8 within 30 days** (`C/C3-floating-repush.txt`; a push is not proof that the image content changed; ghcr gives no date). Today every tested digest still equals what the registry serves (the tests came after the pushes). Only a box installed BEFORE a step's test can read Behind by digest; the demo boxes have none (`C/C1-digests-demo-hp-9201.txt`, 18 of 18 equal). **(4) Decision 30's cost line says the older image runs „until someone (or the automatic leg) presses Update"** — the automatic leg never does, and the manual press moves the digest only in the legacy case above. **(5) What the box would do if the catalog wrote a same-tag re-test** (unit walk): the leg PRESSES it with no controller change — the newest entry whose `from` is the pin is taken, the update renders the new tested digest, the next night reads Current. So option (a)'s cost is in the catalog. **Needs: a decision (STATUS, 2026-09-30).** **-- BUILT 2026-09-30 late (decision 52, option A):** catalog `6a3ead9` — a re-test entry (`from` == `to`, `digest_from`, `box_evidence`), refused by the gates with no new digest, a digest the registry stopped serving, no box proof, or a `digest_from` that is not the previous digest (8 decoys, seen red); `upgrade-test.py --retest`; the ONE command `scripts/retest-floating.py` (`--dry-run`, `--engines-only`). **Proven end to end on 9202:** docmost at an older `redis:7-alpine` digest, the re-test, „run tonight's chain now”, the leg pressed it (`step ended done after 95.0 s`), the new digest runs, data read back, badge current. **Run for real:** no database/redis line differs today (two exact tags do — R-743). Not a cron job (runbook `runbooks/monthly-floating-retest.md` says why); who presses it monthly is the operator's word (STATUS). `audits/night-rulings-2026-09-30/` | **CLOSED 2026-09-30 — built (catalog `6a3ead9`); the monthly run is a standing step** |
| **R-741** | **[P3-LOW] For a few seconds after a fresh install, an `after_install` app answers its PUBLIC default login through the household's front door — the box changes the password only after the app already serves.** MEASURED 2026-09-30 on scratch guest 9202 (calibre-web, controller 0.283.1): polling `GET /opds` with `admin:admin123` (the image's README default) through traefik once a second from the deploy press: 404 until the route existed, **200 at 16:42:06**, 401 from 16:42:07 on (`after_install` done). The first install the same day logged the app started at 15:02:36 and `after_install … done` at 15:02:54 — a fixture probe at 15:02:52 still saw the new password refused, so that window was up to ~18 s. Decision 45 („no app is published with a login a stranger knows") holds after the window, not during it. Who can reach it: anyone who can resolve the app's name in those seconds; calibre-web is not behind the setup gate. Applies to every `after_install` app (calibre-web, mealie, wger, bookstack, claper, …) — not measured for the others. **Needs:** the route published only after `after_install` succeeds (the gate's own route hold, or traefik labels applied after), or the app behind the setup gate until then. `audits/more-night-apps-2026-09-30/A/A3-calibre-default-login-window.txt` **-- FIXED 2026-09-30 late: controller v0.284.x** — an `after_install` app is installed HELD behind the setup gate's door until the login is replaced (or the household says it changed it). **Live on 9202, as a stranger:** calibre-web 0 of 192 default-login tries got in (hold before the first start 20:36:32; opened by after_install 20:37:18, 17 s after the app was up; then 401); mealie 0 of 97 (then mealie's own lock — R-747). The positive control with the generated password was not obtained (calibre-web: a backup stopped the app at that moment; mealie: the lock). Red-proofs RP-IH1..4. `audits/night-rulings-2026-09-30/` | **CLOSED 2026-09-30 — controller v0.284.2** |
| **R-742** | **[P3-LOW] zipline 4.8.0 cannot be reached from 4.6.1 in one step: it refuses to start until the database has run the release before it.** MEASURED 2026-09-30 on both venues, the same way: 4.6.1 → 4.8.0 — the new container restarts (exit 1) with `Error: cannot safely migrate from prisma to drizzle: expected migration 20260508022000_add_file_folder_created_at_index was not applied. To resolve this, repair the database with the previous (latest before this) Zipline release before upgrading` (bench `to-full.log`, 15×). **On box 9202 the product's guarded Update saw it unhealthy after 1 m 30 s and UNDID it by itself in 20 s — the previous version back on the data from before the update** (`box/zipline/step.txt`): the undo worked on a real upstream failure, not a drill. Bench verdict `failed`, box `undone`; nothing written. **Also found:** the zipline fixture had two faults, both fixed (`/` answers 301 → wait on `/api/healthcheck`; a wrong-password control first trips its login limit → 429 on the right one; the right password now goes first). **Needs:** the ladder climbs 4.6.1 → 4.7.x → 4.8.0 (two tested steps — the ladder exists for exactly this). `audits/more-night-apps-2026-09-30/bench/apps/zipline/`, `box/zipline/` **-- DONE 2026-09-30 (evening):** the two steps, each proven on both venues — 4.6.1 → 4.7.0 (catalog `a9700e2`; box 103.6 s, bench 0 kills) and 4.7.0 → 4.8.0 (`fb87030`; box 32.8 s, bench 0 kills). | **CLOSED 2026-09-30 — catalog `a9700e2` + `fb87030`** |
| **R-743** | **[P3-LOW] Exact version tags are re-pushed under the same name too — not only floating lines.** MEASURED 2026-09-30 by `retest-floating.py --dry-run` on the live catalog (`6a3ead9`): `nextcloud:34.0.4-apache` (tested `a5ace30c…`, registry `37b10988…`) and `lscr.io/linuxserver/sonarr:4.0.20` (tested `a5c1a5fe…`, registry `f247545d…`) — the registry now serves another build under a tag the ladder tested. No database or redis line differed. Decision 52 starts with the engine lines, so these were NOT re-tested (`--engines-only`). linuxserver rebuilds its images weekly under the same tags, so every linuxserver app will show up here. **Needs:** a word on whether the monthly run covers every app (the command does; `--engines-only` is the switch) or only engines. `audits/night-rulings-2026-09-30/A/` | **WAITING-ON-OPERATOR — rank P3-LOW; owner: VIKTOR decides, CC runs** |
| **R-744** | **[P3-LOW] outline's fixture cannot seed outline 1.10.1: no `csrfToken` cookie after `installation.create`.** MEASURED 2026-09-30 on both venues (bench and 9202, the end-to-end re-test's first attempt): `POST /api/installation.create` answered 302 with the session, and `GET /home` set no `csrfToken` cookie (1.9.1 did; the 1.9.1 → 1.10.1 step read back that afternoon because its read-back uses the API key made at 1.9.1). So the NEXT outline step, and any re-test of outline, stops at C1 with „no csrfToken cookie from GET /home". **Needs:** the fixture reads outline 1.10.1's CSRF the way its page does (measure first). `audits/night-rulings-2026-09-30/A/e2e/*outline-attempt*` | **READY — rank P3-LOW; owner: CC (catalog harness)** |
| **R-745** | **[P3-LOW] Old CONTROLLER images are never deleted: ~50 versions on each demo box.** MEASURED 2026-09-30 after v0.284.2's one-time sweep: every image no container uses on demo-hp and on the N100 is a `felhom-controller` tag (0.201.0 … 0.283.1); the N100 still reads 3.6 GB reclaimable. Decision 53 covers APP images, and the sweep leaves the controller's own images alone on purpose (the self-update may need the previous one). A release is ~150 MB and there are several a day. **Needs:** the same rule for the controller — keep the running and the previous tag — as a decision (the self-update's rollback reads which image?). `audits/night-rulings-2026-09-30/E/` | **READY — rank P3-LOW; owner: CC (controller); the rule needs a word** |
| **R-746** | **[P3-LOW] `image_digest.resolve` ignores a `@digest` in its argument — it answers the TAG's current digest.** MEASURED 2026-09-30: `resolve('redis:7-alpine@sha256:000…0')` returned `sha256:858f…` (the tag's), while a manifest request for that digest answers 404. Every gate today passes it a plain tag, so no gate is wrong; a caller that passes `ref@digest` to ask "is THIS digest still served" gets a false yes. **Needs:** refuse a digest-carrying ref, or ask for the manifest by digest (with a test). `scripts/image_digest.py` | **READY — rank P3-LOW; owner: CC (catalog)** |
| **R-743** | **[P3-LOW] Exact version tags are re-pushed under the same name too — not only floating lines.** MEASURED 2026-09-30 by `retest-floating.py --dry-run` on the live catalog (`6a3ead9`): `nextcloud:34.0.4-apache` (tested `a5ace30c…`, registry `37b10988…`) and `lscr.io/linuxserver/sonarr:4.0.20` (tested `a5c1a5fe…`, registry `f247545d…`) — the registry now serves another build under a tag the ladder tested. No database or redis line differed. Decision 52 starts with the engine lines, so these were NOT re-tested (`--engines-only`). linuxserver rebuilds its images weekly under the same tags, so every linuxserver app will show up here. **Needs:** a word on whether the monthly run covers every app (the command does; `--engines-only` is the switch) or only engines. `audits/night-rulings-2026-09-30/A/` **-- 2026-10-01:** Operator ruling `09` §3 decisions 54 (a CC session the operator starts with the standing brief runs it monthly) and 55 (every app with a proven ladder; `--engines-only` a switch). The first full run: nextcloud and sonarr re-tested on both venues and written (catalog `3b59dfb` nextcloud, `1a37032` sonarr), ~17 min per app (bench ~13 incl. the 10-minute watch, box ~4). `audits/retest-2026-10/`. | **CLOSED 2026-10-01 — decisions 54, 55; both tags re-tested** |
| **R-744** | **[P3-LOW] outline's fixture cannot seed outline 1.10.1: no `csrfToken` cookie after `installation.create`.** MEASURED 2026-09-30 on both venues (bench and 9202, the end-to-end re-test's first attempt): `POST /api/installation.create` answered 302 with the session, and `GET /home` set no `csrfToken` cookie (1.9.1 did; the 1.9.1 → 1.10.1 step read back that afternoon because its read-back uses the API key made at 1.9.1). So the NEXT outline step, and any re-test of outline, stops at C1 with „no csrfToken cookie from GET /home". **Needs:** the fixture reads outline 1.10.1's CSRF the way its page does (measure first). `audits/night-rulings-2026-09-30/A/e2e/*outline-attempt*` **-- 2026-10-01:** Outline 1.10 names the cookie `__Host-csrfToken` on a secure request (`server/utils/csrf.ts`); the fixture accepts both names (catalog `9fc7052`). Proven on 9202 at the live pin 1.10.1: seed, read back, unknown id and wrong key refused. Outline has no newer release than 1.10.1, so no step. `audits/rulings-2026-10-01/D/D2-outline-fixture-9202.txt` | **CLOSED 2026-10-01 — fixture fixed, proven on 9202** |
| **R-745** | **[P3-LOW] Old CONTROLLER images are never deleted: ~50 versions on each demo box.** MEASURED 2026-09-30 after v0.284.2's one-time sweep: every image no container uses on demo-hp and on the N100 is a `felhom-controller` tag (0.201.0 … 0.283.1); the N100 still reads 3.6 GB reclaimable. Decision 53 covers APP images, and the sweep leaves the controller's own images alone on purpose (the self-update may need the previous one). A release is ~150 MB and there are several a day. **Needs:** the same rule for the controller — keep the running and the previous tag — as a decision (the self-update's rollback reads which image?). `audits/night-rulings-2026-09-30/E/` **-- 2026-10-01:** Decision 56 built: controller v0.285.0 keeps the running image + the previous (the swap record; version order as fallback), deletes older/untagged ones. **Measured first:** the agent rolls back to the RUNNING image (what `/etc/felhom-controller-image` named), never to a previous one the controller hands it. After the floor: demo-hp 84 → 2 controller images (Docker images 15.2 → 11.65 GB), the N100 76 → 2 (6.07 → 1.11 GB), 9202 5 → 2. `audits/rulings-2026-10-01/B/` | **CLOSED 2026-10-01 — controller v0.285.0, floor 0.285.0** |
| **R-746** | **[P3-LOW] `image_digest.resolve` ignores a `@digest` in its argument — it answers the TAG's current digest.** MEASURED 2026-09-30: `resolve('redis:7-alpine@sha256:000…0')` returned `sha256:858f…` (the tag's), while a manifest request for that digest answers 404. Every gate today passes it a plain tag, so no gate is wrong; a caller that passes `ref@digest` to ask "is THIS digest still served" gets a false yes. **Needs:** refuse a digest-carrying ref, or ask for the manifest by digest (with a test). `scripts/image_digest.py` **-- 2026-10-01:** `resolve()` asks the registry for the manifest BY DIGEST when the ref carries one; a malformed digest is refused without a request (catalog `804884a`, `scripts/test_image_digest.py`; red: the old resolver fails 3 of 4). Live: `redis:7-alpine@sha256:000…0` → HTTP 404 (was the tag's digest). `audits/rulings-2026-10-01/D/D1-r746-digest.txt` | **CLOSED 2026-10-01 — catalog 804884a** |
| **R-747** | **[P3-LOW] A stranger can lock the household out of mealie with five wrong logins.** MEASURED 2026-09-30 on 9202 (the R-741 proof): after the install hold opened, a stranger's default-login tries were refused (401) and after five of them mealie answered 423 (locked) to every login — the generated, correct password included. mealie's own brute-force guard, on an app published on the internet; the setup gate and the install hold do not cover an app after its first setup. Not measured: how long the lock lasts. **Needs:** measure the lock's length; decide whether the page tells the household what to do. `audits/night-rulings-2026-09-30/C/C3-mealie-poll.txt` | **READY — rank P3-LOW; owner: CC** |
| **R-748** | **[P3-LOW] The register-shape gate skipped every row whose id has a letter suffix — so R-88a, R-88b and R-209a were never shape-checked, and its count read 3 short.** FOUND 2026-09-30 (late) while counting the register: `register_shape_gate.py` matched `R-\d+` only; the brief's „the reviewer's regex undercounted by 3” is the same three rows. Fixed the same session: `R-\d+[a-z]?`; decoy `suffix-row-eaten-state` (a suffixed row with its state cell eaten) seen passing with the old pattern and convicted with the new. The register is **382** rows by either count now. | **CLOSED 2026-09-30 — `scripts/register_shape_gate.py`** |
| **R-749** | **[P3-LOW] `retest-floating.py` could never start on a fresh bench: it checked for `/opt/upg/upgrade-test.py` on the bench BEFORE the step that copies it there.** FOUND 2026-10-01 at the first full monthly run (decision 55): bench 9401 freshly created by the runbook, the run answered „CANNOT START — missing: the bench LXC 9401 on demo-hp with /opt/upg” in one minute. The runbook says the command syncs the bench itself — it does, but only after the check. On 2026-09-30 the bench had been synced by hand earlier, so nobody saw it. **Needs:** the check asks for what the bench must bring (docker, python3), the sync then provides `/opt/upg`. `audits/rulings-2026-10-01/A/` | **OPEN — rank P3-LOW; owner: CC (catalog harness)** |
| **R-749** | **[P3-LOW] `retest-floating.py` could never start on a fresh bench: it checked for `/opt/upg/upgrade-test.py` on the bench BEFORE the step that copies it there.** FOUND 2026-10-01 at the first full monthly run (decision 55): bench 9401 freshly created by the runbook, the run answered „CANNOT START — missing: the bench LXC 9401 on demo-hp with /opt/upg” in one minute. The runbook says the command syncs the bench itself — it does, but only after the check. On 2026-09-30 the bench had been synced by hand earlier, so nobody saw it. **Needs:** the check asks for what the bench must bring (docker, python3), the sync then provides `/opt/upg`. `audits/rulings-2026-10-01/A/` **-- 2026-10-01:** Fixed the same session (catalog `9e53205`): the check asks for docker + python3; after the sync `/opt/upg/upgrade-test.py` is required. The re-run started at once and finished both apps. | **CLOSED 2026-10-01 — catalog 9e53205** |
| **R-750** | **[P3-LOW] The registry no longer holds controller releases older than 0.213.0 (2026-08-12) — something removed them, and nothing records what.** MEASURED 2026-10-01 (anonymous registry API, `audits/rulings-2026-10-01/B/`): `felhom-controller` has 91 tags, the oldest release 0.213.0; `0.201.0` answers 404; Gitea's package list starts 2026-08-12. No runbook, row or memory names a clean-up. Today nothing needs those versions: a box runs a newer one, a whole-guest restore brings the guest's own Docker store back (mp0 `backup=1`), and decision 56 deletes only on the box. **But** a box or a backup that names a removed version cannot pull it again (R-698's shape, for the controller). **Needs:** find what removed them (a Gitea clean-up rule?), and record the rule — or say it was a one-time act. Read-only on DooPlex. | **OPEN — rank P3-LOW; owner: operator (Gitea settings), CC measures** |
| **R-751** | **[P2] The image clean-up after an app update could crash the whole controller: it re-read the app after a rescan and dereferenced a nil stack when the app was gone.** FOUND 2026-10-01 by the full test suite (controller v0.284.2): `RetainImagesAfterUpdate` runs in a goroutine; `TestR705_TheManualLegRunsByDay` removed its temp dir under it → `panic: invalid memory address` at `image_retention.go:291`. In a box the same happens when an app is removed (or its compose vanishes) between an update's end and the clean-up — a panic in a goroutine ends the process (the agent's supervisor restarts it). Fixed in v0.285.0 the same session: it returns when the app is gone; `TestRetainImagesAfterUpdate_AppGoneDoesNotPanic` seen panicking on the old code; both retention seams are no-ops in the stacks tests (`TestMain`), so no test leaves the goroutine running. `audits/rulings-2026-10-01/B/B1-red-proofs.txt` | **OPEN — fixed in controller v0.285.0, closes when the floor delivers it; owner: CC** |
| **R-751** | **[P2] The image clean-up after an app update could crash the whole controller: it re-read the app after a rescan and dereferenced a nil stack when the app was gone.** FOUND 2026-10-01 by the full test suite (controller v0.284.2): `RetainImagesAfterUpdate` runs in a goroutine; `TestR705_TheManualLegRunsByDay` removed its temp dir under it → `panic: invalid memory address` at `image_retention.go:291`. In a box the same happens when an app is removed (or its compose vanishes) between an update's end and the clean-up — a panic in a goroutine ends the process (the agent's supervisor restarts it). Fixed in v0.285.0 the same session: it returns when the app is gone; `TestRetainImagesAfterUpdate_AppGoneDoesNotPanic` seen panicking on the old code; both retention seams are no-ops in the stacks tests (`TestMain`), so no test leaves the goroutine running. `audits/rulings-2026-10-01/B/B1-red-proofs.txt` **-- 2026-10-01:** Delivered: floor 0.285.0 reached both demo boxes in ~6 s (hub `managed floor SERVED … from declared`). | **CLOSED 2026-10-01 — controller v0.285.0, floor 0.285.0** |
| **R-752** | **[P3-LOW] Four more catalog apps let a stranger lock the household out with wrong passwords for a known login name — like mealie (R-747).** READ 2026-10-01 in each app's source at its pinned tag (not measured live): **calibre-web-automated v4.0.8** — Flask-Limiter on the login keyed on the lowercased USERNAME, 3/minute and 40/day, checked before the password; the default login is `admin` → up to a day; no env switch (a database setting). **wger 2.7** — django-axes keyed on IP, 10 failures, 30 min, each failure restarts it; behind traefik every client has traefik's IP → everyone is locked out (`AXES_*` env vars exist; `AXES_IPWARE_PROXY_COUNT` 0). **Grafana 13.2.3** — per-account, 5 failures in a sliding 5 minutes; a slow trickle keeps it closed (`GF_SECURITY_*`). **BookStack 26.09.1** — key `email|ip`, 5 tries, 60 s, hard-coded; `APP_PROXIES` empty, so the key is the e-mail alone. gokapi (3 s delay, no lock) and claper (per-IP 10/min, no account lock) cannot. **Needs:** per app, the smallest fix that keeps a guessing guard (calibre-web-automated and wger first — longest and broadest), each proven on 9202 as R-747's was. | **OPEN — rank P3-LOW; owner: CC** |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.