Golden 0.286.1 baked + vouched (record); 01 §5 visitor addresses + §7 cloudflared in the guest (R-754); register: R-753/754/772/773 closed, R-775 narrowed, R-776..R-782 opened (410 → 417); live evidence (demo-hp real tunnel, 9202)
gates / gates (push) Successful in 31s
gates / gates (push) Successful in 31s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -864,8 +864,8 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-750** | **[P3-LOW] The registry no longer holds controller releases older than 0.213.0 (2026-08-12) — something removed them, and nothing records what.** MEASURED 2026-10-01 (anonymous registry API, `audits/rulings-2026-10-01/B/`): `felhom-controller` has 91 tags, the oldest release 0.213.0; `0.201.0` answers 404; Gitea's package list starts 2026-08-12. No runbook, row or memory names a clean-up. Today nothing needs those versions: a box runs a newer one, a whole-guest restore brings the guest's own Docker store back (mp0 `backup=1`), and decision 56 deletes only on the box. **But** a box or a backup that names a removed version cannot pull it again (R-698's shape, for the controller). **Needs:** find what removed them (a Gitea clean-up rule?), and record the rule — or say it was a one-time act. Read-only on DooPlex. **-- 2026-10-01 (afternoon):** **Answered (read only, `audits/lockouts-2026-10-01/C/C1-registry-read.txt`).** Not a Gitea rule: `package_cleanup_rule` is EMPTY (READ ONLY query on the shared CNPG); `[cron.cleanup_packages]` only runs rules and Gitea's own expired-data clean-up. The cause is a MANUAL run of DooPlex's `~/git/misc-scripts/gitea-image-prune.sh --all --keep 7 --apply --reclaim` on the night of 2026-08-22/23, recorded in homelab-manifests HM-024 (the Gitea volume was full: `felhom-golden` 22 → 3 versions, /data 14.8 → 4.4 G); `--keep 7` per container package explains controller 0.213.0 (2026-08-12) as the oldest left. Nothing schedules it (crontabs, timers, cluster CronJobs read). If run again as its usage text says, it keeps 7 controller releases (about a day) and `--type generic --keep 3` would cut the agent to 3; it protects no vouched version. **Needs:** the operator's word on a written rule (STATUS). **-- 2026-10-01 (late afternoon):** The rule built (`09` §3 decision 62, operator): `admin/misc-scripts` `c9d5ed5` — `gitea-image-prune.sh` keeps the newest 20 + every version in use (floor, vouched golden, vouched agent, `min_agent`, the golden's baked images, the running hub), refuses (exit 3) when that list is unreadable; `tests/test-prune-plan.sh` red-proofed. Live dry-run: protected controller 0.285.0, golden 0.285.0, agent 0.138.0 + 0.131.0, hub 0.126.0, felhom-samba 1.1.0; would delete 70 controller and 8 hub tags — nothing deleted (no `--apply`). `audits/calibre-name-and-prune-2026-10-01/B/` | **CLOSED 2026-10-01 — decision 62, misc-scripts c9d5ed5** |
|
||||
| **R-751** | **[P2] The image clean-up after an app update could crash the whole controller: it re-read the app after a rescan and dereferenced a nil stack when the app was gone.** FOUND 2026-10-01 by the full test suite (controller v0.284.2): `RetainImagesAfterUpdate` runs in a goroutine; `TestR705_TheManualLegRunsByDay` removed its temp dir under it → `panic: invalid memory address` at `image_retention.go:291`. In a box the same happens when an app is removed (or its compose vanishes) between an update's end and the clean-up — a panic in a goroutine ends the process (the agent's supervisor restarts it). Fixed in v0.285.0 the same session: it returns when the app is gone; `TestRetainImagesAfterUpdate_AppGoneDoesNotPanic` seen panicking on the old code; both retention seams are no-ops in the stacks tests (`TestMain`), so no test leaves the goroutine running. `audits/rulings-2026-10-01/B/B1-red-proofs.txt` **-- 2026-10-01:** Delivered: floor 0.285.0 reached both demo boxes in ~6 s (hub `managed floor SERVED … from declared`). | **CLOSED 2026-10-01 — controller v0.285.0, floor 0.285.0** |
|
||||
| **R-752** | **[P3-LOW] Four more catalog apps let a stranger lock the household out with wrong passwords for a known login name — like mealie (R-747).** READ 2026-10-01 in each app's source at its pinned tag (not measured live): **calibre-web-automated v4.0.8** — Flask-Limiter on the login keyed on the lowercased USERNAME, 3/minute and 40/day, checked before the password; the default login is `admin` → up to a day; no env switch (a database setting). **wger 2.7** — django-axes keyed on IP, 10 failures, 30 min, each failure restarts it; behind traefik every client has traefik's IP → everyone is locked out (`AXES_*` env vars exist; `AXES_IPWARE_PROXY_COUNT` 0). **Grafana 13.2.3** — per-account, 5 failures in a sliding 5 minutes; a slow trickle keeps it closed (`GF_SECURITY_*`). **BookStack 26.09.1** — key `email|ip`, 5 tries, 60 s, hard-coded; `APP_PROXIES` empty, so the key is the e-mail alone. gokapi (3 s delay, no lock) and claper (per-IP 10/min, no account lock) cannot. **Needs:** per app, the smallest fix that keeps a guessing guard (calibre-web-automated and wger first — longest and broadest), each proven on 9202 as R-747's was. **-- 2026-10-01 (afternoon):** **Measured on 9202, each through traefik as a stranger with the public name** (`audits/lockouts-2026-10-01/B/`): **wger** — control: 10 wrong on `admin` locked the second member too; FIXED (decision 58, catalog `82fff32`): username, 5 min, database handler — the second member unaffected, admin in again at 7.5 min (each try during a lock restarts it — measured: 8-minute retries kept a 15-minute lock closed 40+ min). **BookStack** — 1.0 min, kept (decision 59). **Grafana** — 5.0 min, kept (decision 60); a trickle did not hold the household out once the burst aged. **calibre-web-automated** — the form locks 3/min (1.2 min measured) and **40/day per name: after 40 wrong tries in 14 min the right password was refused 2 min later still; only an app restart cleared it** (in-memory store); OPDS has its own 3/min per name (`cps/main.py:75`), no daily limit. No knob for the daily length; both fixes have a household cost — operator decision in STATUS. Installed apps: a settings-only change reaches the stack file at the next sync (images equal, ≤15 min) and the running app at the next `compose up -d` — Restart/Start (measured: the env changed only at Restart), an Update, or a backup's restart (`backup.go:972`, read). **-- 2026-10-01 (late afternoon):** calibre-web done by operator ruling `09` §3 decision 61 (catalog `e9f50b5`): a generated `ADMIN_USER` (`secret`, `hex:5`); `after_install` renames `admin` to it and proves it. 9202: 40 wrong tries on `admin`, the household in at once with its own name (form and OPDS). demo-hp renamed by hand; its name is in the operator's credentials file. All four apps answered (wger 58, BookStack 59, Grafana 60, calibre-web 61). An installed calibre-web is given a made-up name by the box — R-757. | **CLOSED 2026-10-01 — decisions 58–61** |
|
||||
| **R-753** | **[P3-LOW] Behind the tunnel every visitor reaches an app with the SAME address — the tunnel container's — so every per-address guard is an "everyone" guard and every app's log is blind.** MEASURED 2026-10-01 (`audits/lockouts-2026-10-01/A/A1-client-address.txt`): on demo-hp through its real tunnel, a request from DooPlex's public address reached traefik as `172.18.0.5` (cloudflared, in the guest on `traefik-public`) and BookStack as `172.18.0.3` (traefik); on 9202 an echo container showed `X-Forwarded-For`/`X-Real-Ip` = the sending container for the tunnel's hop (traefik DROPS the incoming chain — good: a client cannot forge it) and the real address from the LAN; `CF-Connecting-IP` passes untouched and is FORGEABLE from the LAN. **No box-wide fix taken:** trusting cloudflared in traefik passes Cloudflare's appended chain, whose LEFTMOST entry the client writes — every app reading the leftmost address would believe it; cloudflared's address is docker-assigned; a single-address rewrite needs a traefik plugin (a new dependency). Per-app fixes trust no header (R-752). **Needs (operator):** whether to build a safe version (cloudflared on a fixed-address network + traefik trusting only it + per-app proxy counts), or keep "one address" and fix per app. Only ONE outside address was available (DooPlex has no IPv6); a second was not measured. | **OPEN — rank P3-LOW; owner: operator (direction), CC measures** |
|
||||
| **R-754** | **[P3-LOW] `01-topology-and-trust.md` §7 says cloudflared runs on the Proxmox HOST as an agent-managed service; on every box it runs INSIDE the guest as a container the controller renders.** READ 2026-10-01: `felhom-controller` `internal/infra/templates/cloudflared-compose.yml.tmpl` (`container_name: cloudflared`, network `traefik-public`); demo-hp's guest 9201 runs `cloudflared` (ingress `*.enkisfelhom.hu -> https://traefik`); R-505 saw the same in VM 331. A design decision that the build does not follow — the document or the build is wrong, and only the operator decides which (R-370: a design decision is not a defect). | **OPEN — rank P3-LOW; owner: operator (which is right)** |
|
||||
| **R-753** | **[P3-LOW] Behind the tunnel every visitor reaches an app with the SAME address — the tunnel container's — so every per-address guard is an "everyone" guard and every app's log is blind.** MEASURED 2026-10-01 (`audits/lockouts-2026-10-01/A/A1-client-address.txt`): on demo-hp through its real tunnel, a request from DooPlex's public address reached traefik as `172.18.0.5` (cloudflared, in the guest on `traefik-public`) and BookStack as `172.18.0.3` (traefik); on 9202 an echo container showed `X-Forwarded-For`/`X-Real-Ip` = the sending container for the tunnel's hop (traefik DROPS the incoming chain — good: a client cannot forge it) and the real address from the LAN; `CF-Connecting-IP` passes untouched and is FORGEABLE from the LAN. **No box-wide fix taken:** trusting cloudflared in traefik passes Cloudflare's appended chain, whose LEFTMOST entry the client writes — every app reading the leftmost address would believe it; cloudflared's address is docker-assigned; a single-address rewrite needs a traefik plugin (a new dependency). Per-app fixes trust no header (R-752). **Needs (operator):** whether to build a safe version (cloudflared on a fixed-address network + traefik trusting only it + per-app proxy counts), or keep "one address" and fix per app. Only ONE outside address was available (DooPlex has no IPv6); a second was not measured. **-- 2026-10-01 (evening): operator ruling `09` §3 decision 63 (extend the box's own gate; first the box tells visitors apart). BUILT, controller v0.286.1 (v0.286.0 never floored):** cloudflared alone on `felhom-tunnel` at the fixed `172.16.253.2`, traefik trusts forwarded headers from it only, an entrypoint middleware removes client-writable host/path/address headers (measured: Cloudflare passes a client's `X-Forwarded-Host`/`-Port` and APPENDS to its `X-Forwarded-For`); the controller reads the hop traefik saw (`clientaddr.go`). 19 catalog apps that read the LEFTMOST entry have the chain removed on their router (catalog `04e9516`..`50e4fb4`). **Live:** demo-hp's REAL tunnel — the app sees `6.6.6.6,37.191.56.193, 172.16.253.2`; the stranger's 5 wrong dashboard logins (rotating forged XFF) locked only `37.191.56.193`; docmost (reset) still throttles a rotating forger at try 11; every demo-hp app answers through the tunnel; 9202 simulated tunnel: the household from another address in at once, LAN/impostor forgeries counted as themselves. Both demo boxes reconciled traefik + cloudflared by themselves. Design + sweep of all 56: `audits/visitors-2026-10-01/A/`. Left: settings for right-walking apps (R-776), Emby/Jellyfin LAN rights (R-777), the roll-back window (R-778), a second outside address on the real tunnel (R-779). | **CLOSED 2026-10-01 — controller v0.286.1 (rows R-776..R-779 carry what is left)** |
|
||||
| **R-754** | **[P3-LOW] `01-topology-and-trust.md` §7 says cloudflared runs on the Proxmox HOST as an agent-managed service; on every box it runs INSIDE the guest as a container the controller renders.** READ 2026-10-01: `felhom-controller` `internal/infra/templates/cloudflared-compose.yml.tmpl` (`container_name: cloudflared`, network `traefik-public`); demo-hp's guest 9201 runs `cloudflared` (ingress `*.enkisfelhom.hu -> https://traefik`); R-505 saw the same in VM 331. A design decision that the build does not follow — the document or the build is wrong, and only the operator decides which (R-370: a design decision is not a defect). **-- 2026-10-01 (evening):** the operator's brief: the build is right, correct the document. `01-topology-and-trust.md` §7 now says cloudflared runs INSIDE the guest as a controller-rendered container (and, since v0.286.0, at a fixed address on `felhom-tunnel`), and states the consequence: the data path is not independent of the guest. | **CLOSED 2026-10-01 — document corrected** |
|
||||
| **R-755** | **[P3-LOW] wger runs Django's DEVELOPMENT server in production: `manage.py runserver`, because the template does not set `WGER_USE_GUNICORN=True`.** MEASURED 2026-10-01 on 9202 (`ps` in the wger container: `python3 manage.py runserver 0.0.0.0:8000`); wger 2.7's `extras/docker/production/entrypoint.sh:81-87` runs gunicorn only with that switch. Django's own documentation says runserver is not for production (one process, not hardened). Not changed this session (a different change from R-752's; needs its own bench + box proof, memory watch included). **-- 2026-10-01 (checklist pilot):** re-measured on 9202 at the live template: still `manage.py runserver` (`audits/new-app-checklist-2026-10-01/C/C2-static-reads.txt`). The same server question now carries R-762 — runserver with DEBUG off serves no `/static` and no `/media`, so the fix for both is one decision (upstream's nginx front, or gunicorn + a static server). | **OPEN — rank P3-LOW; owner: CC (catalog)** |
|
||||
| **R-756** | **[P3-LOW] On 9202, "remove with drive data" refuses calibre-web with 409 „…/scratch_hdd/userdata/calibre-web tárhely jelenleg nem elérhető", while the controller container lists that folder.** MEASURED twice on 2026-10-01 (`audits/lockouts-2026-10-01/B/B1…`, `audits/calibre-name-and-prune-2026-10-01/A/A1…`): `POST /api/stacks/calibre-web/remove` with `remove_hdd_data` → 409; `docker exec felhom-controller ls -ld /mnt/felhom-drives/scratch_hdd/userdata/calibre-web` → the directory (dated 2026-09-22). The walk then removed the app keeping the data (R-442's fail-closed answer). Either the drive is not a registered drive on this scratch box (a test-venue artefact) or the resolver reads another path than the one it names. Not measured which. **-- 2026-10-01 (night, new apps):** the same 409 for Grimmory (`remove_hdd_data`), and a drive app's per-app backup is NOT offered for restore after a remove — the controller logged `drive … not mounted — skipping ensure (held by drive gate)`; the unit lives on that drive. So on 9202 checklist row 2.5 cannot be shown for any drive app until its scratch drive is a registered drive (`audits/new-apps-2026-10-01/box/grimmory/restore-why.txt`). | **OPEN — rank P3-LOW; owner: CC** |
|
||||
| **R-757** | **[P3-LOW] A template that gains a generated `secret` field makes the box INVENT that value for apps already installed — for calibre-web a login name the app never got.** MEASURED 2026-10-01 on demo-hp: 9 minutes after catalog `e9f50b5` (decision 61) synced, the controller logged `InjectMissingFields … injected missing fields: ADMIN_USER` (deploy.go:1337) and the app page's reveal returned a 10-character name that was NOT calibre-web's login (the app had kept its own name; `after_install` runs only after a fresh install). The page did not list the field, but the reveal answers it, and the password field's text now says "the user name above". demo-hp was fixed by renaming the app's user to the box's recorded name (credentials file updated). Any other installed calibre-web gets the same made-up name at its next sync while its login stays `admin` (no other box has one today: N100 none, Tester-2 not registered). **Needs:** InjectMissingFields must not invent a value an app has to have been GIVEN (a field consumed only by `after_install`), or such a field needs an "installed apps: ask" path. `audits/calibre-name-and-prune-2026-10-01/A/A2…, A3…` | **OPEN — rank P3-LOW; owner: CC (controller)** |
|
||||
@@ -883,10 +883,17 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-769** | **[P3-LOW] Pinchflat is not built: upstream is paused (no release in 2026, last push 2025-12-16) and its last release has no image tag.** READ 2026-10-01: ghcr's newest version tag `v2025.6.6`, `latest` amd64 only; runs as root by default; an unanswered 30 GB yt-dlp memory report (#866); third parties call it unmaintained (community-scripts #15968). Forks with images exist (Pinchflat-NGX, MorganKryze). **Needs:** the operator's word on a fork (a new upstream), after the YouTube sentence (R-767). `audits/new-apps-2026-10-01/FIT.md` | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator** |
|
||||
| **R-770** | **[P3-LOW] Invidious — fit check only; the recommendation is not to build it.** READ 2026-10-01: playback needs `invidious-companion` (rolling `latest`, no version tags); PostgreSQL 14 (EOL 2026-11); `registration_enabled: true` by default; upstream: a bot check means „your IP is blocked from YouTube”, a 429 can last 24 h, triggered by „someone on your network” — on our boxes that IP is the household's. One bad period in 2026 (March, ~2 weeks). No report found of a family's other devices being bot-checked (inference). **Needs:** the operator's go / no-go. `audits/new-apps-2026-10-01/FIT.md` | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator** |
|
||||
| **R-771** | **[P3-LOW] moonlight-web — fit check only; not buildable through an HTTP-only tunnel at usable latency.** READ 2026-10-01: two unrelated projects (MrCreativ3001/moonlight-web-stream, the original; linckosz/moonlight-web); both need Sunshine/Apollo/Wolf on a gaming PC on the LAN and WebRTC over UDP (40000-40100/udp; linckosz recommends host networking and sends telemetry by default); both have a WebSocket fallback (high latency, all video through the tunnel); a logged-in user controls the PC's desktop. **Needs:** the operator's go / no-go (LAN-only use would need a different publishing model). `audits/new-apps-2026-10-01/FIT.md` | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator** |
|
||||
| **R-772** | **[P3-LOW] A health probe that finds NO container to probe records `healthy: true` — and the next check is 5 minutes away.** MEASURED 2026-10-01 on 9202 (controller 0.285.0, Karakeep): with the app container stopped, the probe's record read `{type: none, target: karakeep, healthy: true, message_key: health.no_probe_container}` and stayed so 4+ minutes (`RunHealthProbes: skipping karakeep — last check …, effective interval 5m0s, healthy=true`). The app's STATE did read `degraded` within 10 s, so the household saw it; the probe record alone says healthy about a check that did not run (the presence-is-not-success shape). Not checked: which readers use `health_probe.healthy` (the guarded Update's `verifying`? the hub report?). **Needs:** a not-run probe recorded as not-healthy (or `unknown`), with a test that pins it. `audits/new-apps-2026-10-01/box/karakeep-768M/neg-and-crawl.txt` | **READY — rank P3-LOW; owner: CC (controller)** |
|
||||
| **R-773** | **[P3-LOW] After a remove + restore, an app's sign-up route block (decision 47) is gone; only the app's own switch still refuses.** MEASURED 2026-10-01 on 9202 (Karakeep): before — `/signup` 403 and tRPC `users.create` 403; after the household removed the app (keeping backups) and pressed restore — `/signup` answers 200, `users.create` still 403 (the restored env keeps `DISABLE_SIGNUPS=true`). For an app with NO own switch (opengist, wishlist: block only) a remove + restore would reopen sign-up entirely — not measured. **Needs:** the restore re-applies the lock record's block (or the gate) for an app whose template has `signup_block`, with a test; then measured on a block-only app. `audits/new-apps-2026-10-01/box/karakeep/restore.txt` | **READY — rank P3-LOW; owner: CC (controller)** |
|
||||
| **R-772** | **[P3-LOW] A health probe that finds NO container to probe records `healthy: true` — and the next check is 5 minutes away.** MEASURED 2026-10-01 on 9202 (controller 0.285.0, Karakeep): with the app container stopped, the probe's record read `{type: none, target: karakeep, healthy: true, message_key: health.no_probe_container}` and stayed so 4+ minutes (`RunHealthProbes: skipping karakeep — last check …, effective interval 5m0s, healthy=true`). The app's STATE did read `degraded` within 10 s, so the household saw it; the probe record alone says healthy about a check that did not run (the presence-is-not-success shape). Not checked: which readers use `health_probe.healthy` (the guarded Update's `verifying`? the hub report?). **Needs:** a not-run probe recorded as not-healthy (or `unknown`), with a test that pins it. `audits/new-apps-2026-10-01/box/karakeep-768M/neg-and-crawl.txt` **-- FIXED controller v0.286.1:** a not-run probe records `healthy: false, not_checked: true`, is looked at on the next tick (0.286.0 still waited out the last healthy record's 5-minute interval — found live, fixed in .1), the state stays the containers' (R-630 holds), the app page says so. Readers checked: `health_probe` feeds only the state override and the API/page; the hub gets container states, never the probe record. **Live on 9202 (0.286.1):** paperless-webserver stopped 19:48:33 → `healthy false, not_checked true` at 19:48:47; started 19:49:05 → healthy at 19:49:47. Red-proofs RP-D1, RP-D1b. `audits/visitors-2026-10-01/D/`. | **CLOSED 2026-10-01 — controller v0.286.1** |
|
||||
| **R-773** | **[P3-LOW] After a remove + restore, an app's sign-up route block (decision 47) is gone; only the app's own switch still refuses.** MEASURED 2026-10-01 on 9202 (Karakeep): before — `/signup` 403 and tRPC `users.create` 403; after the household removed the app (keeping backups) and pressed restore — `/signup` answers 200, `users.create` still 403 (the restored env keeps `DISABLE_SIGNUPS=true`). For an app with NO own switch (opengist, wishlist: block only) a remove + restore would reopen sign-up entirely — not measured. **Needs:** the restore re-applies the lock record's block (or the gate) for an app whose template has `signup_block`, with a test; then measured on a block-only app. `audits/new-apps-2026-10-01/box/karakeep/restore.txt` **-- FIXED controller v0.286.0:** a REMOVED app restored from its backup gets the lock record (`opened_by: restore`) and the block written before anything starts; the loop sets the app's own switch; an installed app the household never closed keeps what it had (decision 49). **Live on 9202:** Karakeep — before: `/signup` 403, `users.create` 403; removed keeping backups; restored → record `opened_by: restore`, `native_lock: applied`, block file present, a stranger's `/signup` 403 and `users.create` 403, the bookmark read back. Red-proof RP-D2. Block-only apps (opengist, wishlist) share the code path, not measured. `audits/visitors-2026-10-01/D/r773-live.txt`. | **CLOSED 2026-10-01 — controller v0.286.0** |
|
||||
| **R-774** | **[P3-LOW] Two things the new apps' pages do not show yet: Karakeep's mail-ON path is unproven, and its official phone app reports crashes to its makers.** READ/MEASURED 2026-10-01: Karakeep's `smtp_mapping` (plaintext :2526, as Cal.com) was proven only with mail OFF (a fresh install boots) — 9202 has no hub, so the relay cannot be exercised there; the official mobile app ships Sentry crash reporting with a hard-coded DSN (`apps/mobile/app/_layout.tsx`, FIT.md). **Needs:** one password-reset mail from Karakeep on a hub-enabled box (demo-hp 9201, a throwaway install); and a sentence on the page about the phone app (copy, freeze). `audits/new-apps-2026-10-01/bench/karakeep-mail-off-boot.txt` | **READY — rank P3-LOW; owner: CC (catalog)** |
|
||||
| **R-775** | **[P2-MEDIUM] Grimmory: a stranger's 5 wrong sign-ins lock EVERY visitor out of the web login for 15 minutes — so Grimmory was not published.** MEASURED 2026-10-01 on 9202 (drill catalog, v3.4.1, through traefik): after 5 wrong tries for `admin` every further sign-in answered 429 — the household's right password AND a different name — and stayed 429 for 10+ minutes of retries. Read in the jar: `AuthRateLimitService` — Caffeine `expireAfterWrite(ofMinutes(15))`, `MAX_ATTEMPTS 5`, keys `login:ip:` and `login:user:`; Spring `forward-headers-strategy: native` takes the address from X-Forwarded-For, and behind the tunnel every visitor is the tunnel container's address (R-753) — the wger shape (R-752), with no setting to change it. Everything else in the checklist passed (bench + box step v3.4.1 → v3.5.0, gate by its own probe, OPDS through traefik); two smaller findings for the publishing session: on a reinstall over the first install's kept books, a new upload was saved to the drive but not added to the library (`box/grimmory/reinstall-c1.txt`, not investigated); and the remove + restore round trip (2.5) cannot be shown on 9202 for a drive app — its backup lives on the scratch drive, which is not a registered drive (R-756). The template waits in `audits/new-apps-2026-10-01/wip/grimmory/`. **Needs (operator):** (A) publish with a sentence on the page that wrong guesses by others can lock the login for 15 minutes (MEASURED: during the lock an e-reader's OPDS feed still answered 200 with its own login, wrong 401 — `box/grimmory/opds-under-lock.txt`), or (B) wait until the box passes each visitor's real address (R-753). `audits/new-apps-2026-10-01/box/grimmory/throttle.txt` | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (A/B), CC publishes** |
|
||||
| **R-775** | **[P2-MEDIUM] Grimmory: a stranger's 5 wrong sign-ins lock EVERY visitor out of the web login for 15 minutes — so Grimmory was not published.** MEASURED 2026-10-01 on 9202 (drill catalog, v3.4.1, through traefik): after 5 wrong tries for `admin` every further sign-in answered 429 — the household's right password AND a different name — and stayed 429 for 10+ minutes of retries. Read in the jar: `AuthRateLimitService` — Caffeine `expireAfterWrite(ofMinutes(15))`, `MAX_ATTEMPTS 5`, keys `login:ip:` and `login:user:`; Spring `forward-headers-strategy: native` takes the address from X-Forwarded-For, and behind the tunnel every visitor is the tunnel container's address (R-753) — the wger shape (R-752), with no setting to change it. Everything else in the checklist passed (bench + box step v3.4.1 → v3.5.0, gate by its own probe, OPDS through traefik); two smaller findings for the publishing session: on a reinstall over the first install's kept books, a new upload was saved to the drive but not added to the library (`box/grimmory/reinstall-c1.txt`, not investigated); and the remove + restore round trip (2.5) cannot be shown on 9202 for a drive app — its backup lives on the scratch drive, which is not a registered drive (R-756). The template waits in `audits/new-apps-2026-10-01/wip/grimmory/`. **Needs (operator):** (A) publish with a sentence on the page that wrong guesses by others can lock the login for 15 minutes (MEASURED: during the lock an e-reader's OPDS feed still answered 200 with its own login, wrong 401 — `box/grimmory/opds-under-lock.txt`), or (B) wait until the box passes each visitor's real address (R-753). `audits/new-apps-2026-10-01/box/grimmory/throttle.txt` **-- 2026-10-01 (evening):** option B's precondition SHIPPED (controller v0.286.1, R-753): Grimmory's Tomcat RemoteIpValve walks from the right and counts `172.16.0.0/12` as a proxy (READ in source, Spring Boot 4.1.1 — not yet measured with Grimmory's own lock), so `login:ip:` becomes per visitor; `login:user:` still lets a stranger lock the public name `admin` 15 min. A third route was spiked and passed: Grimmory behind the permanent family gate with its e-reader paths excepted (R-780). Recommendation: publish behind the family gate if R-780 is built; otherwise B with a measured 3.6. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (A / B / behind R-780), CC publishes** |
|
||||
| **R-776** | **[P3-LOW] Right-walking catalog apps need ONE setting to see each visitor since v0.286 (R-753); without it they keep the tunnel's one address (a stranger can still trip their per-address limits for everyone).** READ in source (`audits/visitors-2026-10-01/A/sweep/`): kimai `TRUSTED_PROXIES=127.0.0.1,172.16.0.0/12`; zipline `CORE_TRUST_PROXY=true` + `CORE_TRUSTED_PROXIES=172.16.0.0/12`; vikunja `VIKUNJA_SERVICE_IPEXTRACTIONMETHOD=xff`; nextcloud `TRUSTED_PROXIES=172.16.0.0/12`; n8n `N8N_PROXY_HOPS=2` (optional). BookStack `APP_PROXIES` measured this session (see its catalog commit). **Needs:** per app, the setting + checklist 3.6 re-measured on 9202 through the simulated tunnel (the BookStack shape, `audits/visitors-2026-10-01/tools/bookstack_36.py`). Count readers (calibre-web, tandoor, wger) stay as they are — a fixed count is wrong for one of the two paths. | **READY — rank P3-LOW; owner: CC (catalog)** |
|
||||
| **R-777** | **[P2-MEDIUM] Emby and Jellyfin treat every internet visitor as being on the LAN — users with "remote access" off can sign in from the internet, IP filters and remote limits are skipped.** READ in source (`audits/visitors-2026-10-01/A/sweep/sweep-1.md`, `sweep-2.md`), not measured live: Jellyfin with `KnownProxies` empty uses the TCP peer (traefik, private) → "LAN"; Emby reads the leftmost XFF (its chain is now removed by the R-753 reset, so it sees cloudflared's private address → "LAN", as before). True before R-753 too; R-753 neither caused nor fixed it. **Needs:** measure on 9202 (a user with remote access off, through the simulated tunnel); Jellyfin: `KnownProxies` `172.16.0.0/12` in `network.xml` (no env — an `after_install` or a seed file); Emby: no setting fixes it (its `LocalNetworkSubnets` still counts private ranges) — a page sentence or a decision. | **READY — rank P2-MEDIUM; owner: CC (measure), operator (Emby route)** |
|
||||
| **R-778** | **[P3-LOW] A box that rolls back to a controller ≤ 0.285 after v0.286 keeps the new traefik (an old controller never rewrites a running traefik) — and the old `clientIP` believes the LEFTMOST X-Forwarded-For, which a stranger then writes: the dashboard's login counter becomes dodgeable until the box moves forward again.** Reasoned from the code (old `claim.go` clientIP + v0.286 traefik trust), not measured. The floor never moves back; the window is the self-update's crash roll-back. **Needs:** decide whether that window matters (it closes at the next floor); if it does, a 0.285.x patch that reads the rightmost hop, or the self-update refusing to roll back across v0.286. | **OPEN — rank P3-LOW; owner: CC** |
|
||||
| **R-779** | **[P3-LOW] Part A's "two outside addresses seen as two" is proven through the simulated tunnel only; on the REAL tunnel the second outside address (ep0, one request allowed) was refused by Cloudflare's edge with 403 and never reached the box.** Measured 2026-10-01 19:51 UTC (`audits/visitors-2026-10-01/A/L2-demo-hp-real-tunnel.txt`): no log line on demo-hp; demo-hp's box has no geo restriction in its settings, so a Cloudflare ZONE rule (country or bot, not read) refused a German datacenter address. DooPlex's own address on the real tunnel was seen as itself. **Needs:** one sign-in from a second Hungarian address (the operator's phone off wifi) while DooPlex is locked out — 2 minutes; and say which Cloudflare rule refused ep0. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (a phone), CC reads the logs** |
|
||||
| **R-780** | **[P2-MEDIUM] A permanent household gate with family accounts — the spike PASSED; the build waits for the operator's go (`09` §3 decision 63).** Measured on 9202 2026-10-01 (`audits/permanent-gate-2026-10-01/VERDICT.md`, exit test committed before): a stranger reached nothing of Grimmory or MeTube (36 requests incl. websockets); family members with their own logins got in for 30 days, logout worked; a stranger's guesses locked only the stranger; Grimmory's OPDS/Kobo/KOReader worked through anchored path exceptions with the app's own login; the family login never opened the dashboard; +0.4 ms per request; the gate down → apps refuse (500). **Build requirement F1:** exceptions must be anchored (`PathPrefix(/api/v1/opds)` also matched `/api/v1/opdsx`). **Cost:** 2 sessions (controller release; catalog — Grimmory and MeTube published behind it). If nothing is decided: Grimmory stays held (R-775) and MeTube unpublished (R-767). | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (go/no-go), CC builds** |
|
||||
| **R-781** | **[P3-LOW] The catalog's `scripts/test_gate_decoys.py` fails 4 of its own "genuine" onboarding cases — on the untouched tree.** Measured 2026-10-01 (same 4 FAIL lines before and after this session's checklist edit): the harness copies the catalog into a temp tree with a fake sibling `felhom.eu`, and the REAL published records (radicale, karakeep, dawarich) name evidence that the fake sibling lacks, so a complete record reads incomplete. The gate itself (`onboarding`) is green on the real tree. **Needs:** the genuine cases build their own records, or the fake sibling mirrors the cited evidence. | **READY — rank P3-LOW; owner: CC (catalog)** |
|
||||
| **R-782** | **[P3-LOW] Two side observations of the R-753 sweep, inferred, not measured:** glance's seeded `glance.yml` has no `auth:` block (the dashboard is public to anyone with the address), and homepage's `/api/*` refuses a Host not in `HOMEPAGE_ALLOWED_HOSTS`, which the template does not set (widgets may 400). **Needs:** measure both on 9202; glance: decide whether a public link dashboard is intended (the setup gate does not cover it after setup). | **READY — rank P3-LOW; owner: CC (catalog)** |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
|
||||
Reference in New Issue
Block a user