Lockouts (R-752): decisions 58-60 (decided by CC unattended), the one address behind the tunnel (R-753), the registry answered (R-750); STATUS, report, evidence
gates / gates (push) Successful in 27s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-01 12:49:20 +02:00
parent 92a60c62bd
commit 8dab40c7a7
14 changed files with 287 additions and 41 deletions
+2 -2
View File
@@ -861,9 +861,9 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-747** | **[P3-LOW] A stranger can lock the household out of mealie with five wrong logins.** MEASURED 2026-09-30 on 9202 (the R-741 proof): after the install hold opened, a stranger's default-login tries were refused (401) and after five of them mealie answered 423 (locked) to every login — the generated, correct password included. mealie's own brute-force guard, on an app published on the internet; the setup gate and the install hold do not cover an app after its first setup. Not measured: how long the lock lasts. **Needs:** measure the lock's length; decide whether the page tells the household what to do. `audits/night-rulings-2026-09-30/C/C3-mealie-poll.txt` **-- 2026-10-01:** Measured at v3.28.0 (source + 9202): 5 wrong logins lock the ACCOUNT (not the IP) for `SECURITY_USER_LOCKOUT_TIME` hours (default 24); the lock is lifted by an hourly job; an admin can unlock others via `POST /api/admin/users/unlock`, but the household's only admin is the locked account. Both login names are public (`admin`, `changeme@example.com`). **Fixed** (`09` §3 decision 57, decided by CC unattended — operator may reverse): `SECURITY_USER_LOCKOUT_TIME=1`, catalog `a4597cd`; on 9202 the right password answered 423 for 120 min, then 200; a wrong one still 401 (`audits/rulings-2026-10-01/C/`). **Left:** a stranger can renew the lock every hour (per account, public name — option (d) in decision 57 would end that); the page does not tell the household why it is locked; an INSTALLED mealie takes the new setting only when its compose is rendered again (not measured which act does that). | **NARROWED — the lock is 1–2 h; renewal and the page remain; owner: CC** |
| **R-748** | **[P3-LOW] The register-shape gate skipped every row whose id has a letter suffix — so R-88a, R-88b and R-209a were never shape-checked, and its count read 3 short.** FOUND 2026-09-30 (late) while counting the register: `register_shape_gate.py` matched `R-\d+` only; the brief's „the reviewer's regex undercounted by 3” is the same three rows. Fixed the same session: `R-\d+[a-z]?`; decoy `suffix-row-eaten-state` (a suffixed row with its state cell eaten) seen passing with the old pattern and convicted with the new. The register is **382** rows by either count now. | **CLOSED 2026-09-30 — `scripts/register_shape_gate.py`** |
| **R-749** | **[P3-LOW] `retest-floating.py` could never start on a fresh bench: it checked for `/opt/upg/upgrade-test.py` on the bench BEFORE the step that copies it there.** FOUND 2026-10-01 at the first full monthly run (decision 55): bench 9401 freshly created by the runbook, the run answered „CANNOT START — missing: the bench LXC 9401 on demo-hp with /opt/upg” in one minute. The runbook says the command syncs the bench itself — it does, but only after the check. On 2026-09-30 the bench had been synced by hand earlier, so nobody saw it. **Needs:** the check asks for what the bench must bring (docker, python3), the sync then provides `/opt/upg`. `audits/rulings-2026-10-01/A/` **-- 2026-10-01:** Fixed the same session (catalog `9e53205`): the check asks for docker + python3; after the sync `/opt/upg/upgrade-test.py` is required. The re-run started at once and finished both apps. | **CLOSED 2026-10-01 — catalog 9e53205** |
| **R-750** | **[P3-LOW] The registry no longer holds controller releases older than 0.213.0 (2026-08-12) — something removed them, and nothing records what.** MEASURED 2026-10-01 (anonymous registry API, `audits/rulings-2026-10-01/B/`): `felhom-controller` has 91 tags, the oldest release 0.213.0; `0.201.0` answers 404; Gitea's package list starts 2026-08-12. No runbook, row or memory names a clean-up. Today nothing needs those versions: a box runs a newer one, a whole-guest restore brings the guest's own Docker store back (mp0 `backup=1`), and decision 56 deletes only on the box. **But** a box or a backup that names a removed version cannot pull it again (R-698's shape, for the controller). **Needs:** find what removed them (a Gitea clean-up rule?), and record the rule — or say it was a one-time act. Read-only on DooPlex. | **OPEN — rank P3-LOW; owner: operator (Gitea settings), CC measures** |
| **R-750** | **[P3-LOW] The registry no longer holds controller releases older than 0.213.0 (2026-08-12) — something removed them, and nothing records what.** MEASURED 2026-10-01 (anonymous registry API, `audits/rulings-2026-10-01/B/`): `felhom-controller` has 91 tags, the oldest release 0.213.0; `0.201.0` answers 404; Gitea's package list starts 2026-08-12. No runbook, row or memory names a clean-up. Today nothing needs those versions: a box runs a newer one, a whole-guest restore brings the guest's own Docker store back (mp0 `backup=1`), and decision 56 deletes only on the box. **But** a box or a backup that names a removed version cannot pull it again (R-698's shape, for the controller). **Needs:** find what removed them (a Gitea clean-up rule?), and record the rule — or say it was a one-time act. Read-only on DooPlex. **-- 2026-10-01 (afternoon):** **Answered (read only, `audits/lockouts-2026-10-01/C/C1-registry-read.txt`).** Not a Gitea rule: `package_cleanup_rule` is EMPTY (READ ONLY query on the shared CNPG); `[cron.cleanup_packages]` only runs rules and Gitea's own expired-data clean-up. The cause is a MANUAL run of DooPlex's `~/git/misc-scripts/gitea-image-prune.sh --all --keep 7 --apply --reclaim` on the night of 2026-08-22/23, recorded in homelab-manifests HM-024 (the Gitea volume was full: `felhom-golden` 22 → 3 versions, /data 14.8 → 4.4 G); `--keep 7` per container package explains controller 0.213.0 (2026-08-12) as the oldest left. Nothing schedules it (crontabs, timers, cluster CronJobs read). If run again as its usage text says, it keeps 7 controller releases (about a day) and `--type generic --keep 3` would cut the agent to 3; it protects no vouched version. **Needs:** the operator's word on a written rule (STATUS). | **ANSWERED — WAITING-ON-OPERATOR: keep the manual prune or write a rule; owner: operator** |
| **R-751** | **[P2] The image clean-up after an app update could crash the whole controller: it re-read the app after a rescan and dereferenced a nil stack when the app was gone.** FOUND 2026-10-01 by the full test suite (controller v0.284.2): `RetainImagesAfterUpdate` runs in a goroutine; `TestR705_TheManualLegRunsByDay` removed its temp dir under it → `panic: invalid memory address` at `image_retention.go:291`. In a box the same happens when an app is removed (or its compose vanishes) between an update's end and the clean-up — a panic in a goroutine ends the process (the agent's supervisor restarts it). Fixed in v0.285.0 the same session: it returns when the app is gone; `TestRetainImagesAfterUpdate_AppGoneDoesNotPanic` seen panicking on the old code; both retention seams are no-ops in the stacks tests (`TestMain`), so no test leaves the goroutine running. `audits/rulings-2026-10-01/B/B1-red-proofs.txt` **-- 2026-10-01:** Delivered: floor 0.285.0 reached both demo boxes in ~6 s (hub `managed floor SERVED … from declared`). | **CLOSED 2026-10-01 — controller v0.285.0, floor 0.285.0** |
| **R-752** | **[P3-LOW] Four more catalog apps let a stranger lock the household out with wrong passwords for a known login name — like mealie (R-747).** READ 2026-10-01 in each app's source at its pinned tag (not measured live): **calibre-web-automated v4.0.8** — Flask-Limiter on the login keyed on the lowercased USERNAME, 3/minute and 40/day, checked before the password; the default login is `admin` → up to a day; no env switch (a database setting). **wger 2.7** — django-axes keyed on IP, 10 failures, 30 min, each failure restarts it; behind traefik every client has traefik's IP → everyone is locked out (`AXES_*` env vars exist; `AXES_IPWARE_PROXY_COUNT` 0). **Grafana 13.2.3** — per-account, 5 failures in a sliding 5 minutes; a slow trickle keeps it closed (`GF_SECURITY_*`). **BookStack 26.09.1** — key `email|ip`, 5 tries, 60 s, hard-coded; `APP_PROXIES` empty, so the key is the e-mail alone. gokapi (3 s delay, no lock) and claper (per-IP 10/min, no account lock) cannot. **Needs:** per app, the smallest fix that keeps a guessing guard (calibre-web-automated and wger first — longest and broadest), each proven on 9202 as R-747's was. | **OPEN — rank P3-LOW; owner: CC** |
| **R-752** | **[P3-LOW] Four more catalog apps let a stranger lock the household out with wrong passwords for a known login name — like mealie (R-747).** READ 2026-10-01 in each app's source at its pinned tag (not measured live): **calibre-web-automated v4.0.8** — Flask-Limiter on the login keyed on the lowercased USERNAME, 3/minute and 40/day, checked before the password; the default login is `admin` → up to a day; no env switch (a database setting). **wger 2.7** — django-axes keyed on IP, 10 failures, 30 min, each failure restarts it; behind traefik every client has traefik's IP → everyone is locked out (`AXES_*` env vars exist; `AXES_IPWARE_PROXY_COUNT` 0). **Grafana 13.2.3** — per-account, 5 failures in a sliding 5 minutes; a slow trickle keeps it closed (`GF_SECURITY_*`). **BookStack 26.09.1** — key `email|ip`, 5 tries, 60 s, hard-coded; `APP_PROXIES` empty, so the key is the e-mail alone. gokapi (3 s delay, no lock) and claper (per-IP 10/min, no account lock) cannot. **Needs:** per app, the smallest fix that keeps a guessing guard (calibre-web-automated and wger first — longest and broadest), each proven on 9202 as R-747's was. **-- 2026-10-01 (afternoon):** **Measured on 9202, each through traefik as a stranger with the public name** (`audits/lockouts-2026-10-01/B/`): **wger** — control: 10 wrong on `admin` locked the second member too; FIXED (decision 58, catalog `82fff32`): username, 5 min, database handler — the second member unaffected, admin in again at 7.5 min (each try during a lock restarts it — measured: 8-minute retries kept a 15-minute lock closed 40+ min). **BookStack** — 1.0 min, kept (decision 59). **Grafana** — 5.0 min, kept (decision 60); a trickle did not hold the household out once the burst aged. **calibre-web-automated** — the form locks 3/min (1.2 min measured) and **40/day per name: after 40 wrong tries in 14 min the right password was refused 2 min later still; only an app restart cleared it** (in-memory store); OPDS has its own 3/min per name (`cps/main.py:75`), no daily limit. No knob for the daily length; both fixes have a household cost — operator decision in STATUS. Installed apps: a settings-only change reaches the stack file at the next sync (images equal, ≤15 min) and the running app at the next `compose up -d` — Restart/Start (measured: the env changed only at Restart), an Update, or a backup's restart (`backup.go:972`, read). | **NARROWED — wger fixed, BookStack/Grafana kept; calibre-web waits for the operator; owner: operator (calibre-web), CC** |
| **R-753** | **[P3-LOW] Behind the tunnel every visitor reaches an app with the SAME address — the tunnel container's — so every per-address guard is an "everyone" guard and every app's log is blind.** MEASURED 2026-10-01 (`audits/lockouts-2026-10-01/A/A1-client-address.txt`): on demo-hp through its real tunnel, a request from DooPlex's public address reached traefik as `172.18.0.5` (cloudflared, in the guest on `traefik-public`) and BookStack as `172.18.0.3` (traefik); on 9202 an echo container showed `X-Forwarded-For`/`X-Real-Ip` = the sending container for the tunnel's hop (traefik DROPS the incoming chain — good: a client cannot forge it) and the real address from the LAN; `CF-Connecting-IP` passes untouched and is FORGEABLE from the LAN. **No box-wide fix taken:** trusting cloudflared in traefik passes Cloudflare's appended chain, whose LEFTMOST entry the client writes — every app reading the leftmost address would believe it; cloudflared's address is docker-assigned; a single-address rewrite needs a traefik plugin (a new dependency). Per-app fixes trust no header (R-752). **Needs (operator):** whether to build a safe version (cloudflared on a fixed-address network + traefik trusting only it + per-app proxy counts), or keep "one address" and fix per app. Only ONE outside address was available (DooPlex has no IPv6); a second was not measured. | **OPEN — rank P3-LOW; owner: operator (direction), CC measures** |
| **R-754** | **[P3-LOW] `01-topology-and-trust.md` §7 says cloudflared runs on the Proxmox HOST as an agent-managed service; on every box it runs INSIDE the guest as a container the controller renders.** READ 2026-10-01: `felhom-controller` `internal/infra/templates/cloudflared-compose.yml.tmpl` (`container_name: cloudflared`, network `traefik-public`); demo-hp's guest 9201 runs `cloudflared` (ingress `*.enkisfelhom.hu -> https://traefik`); R-505 saw the same in VM 331. A design decision that the build does not follow — the document or the build is wrong, and only the operator decides which (R-370: a design decision is not a defect). | **OPEN — rank P3-LOW; owner: operator (which is right)** |
| **R-755** | **[P3-LOW] wger runs Django's DEVELOPMENT server in production: `manage.py runserver`, because the template does not set `WGER_USE_GUNICORN=True`.** MEASURED 2026-10-01 on 9202 (`ps` in the wger container: `python3 manage.py runserver 0.0.0.0:8000`); wger 2.7's `extras/docker/production/entrypoint.sh:81-87` runs gunicorn only with that switch. Django's own documentation says runserver is not for production (one process, not hardened). Not changed this session (a different change from R-752's; needs its own bench + box proof, memory watch included). | **OPEN — rank P3-LOW; owner: CC (catalog)** |