Files
felhom.eu/REPORT-lockouts-2026-10-01.md
T

6.3 KiB
Raw Permalink Blame History

REPORT — strangers and lockouts (R-752), the one address behind the tunnel (R-753), the registry (R-750) — 2026-10-01 afternoon

Evidence: documentation/audits/lockouts-2026-10-01/ (A, B, C, T, tools). Architecture read: 01-topology-and-trust.md §5, §7; 09 §3 decisions 45–47, 57; 06 (the tunnel is not described there). Baselines (live Gitea ~10:55 CEST): controller c1b123c64955, agent d766666ff8cf, felhom.eu a6a9f0b2458e, catalog 83636352ea10 — all matched. Register 387 rows; highest R-752; last decision 57.

The Part table

Part done / not done / changed why
Operator note (decision 57 kept) done — 09 §3 + CONTEXT first
A — the client address done — measured; no box-wide fix (R-753) trusting cloudflared would pass a client-written leftmost address; a single-address rewrite needs a plugin
A1 two outside addresses changed — one (DooPlex 37.191.56.193; no IPv6 here) the "same address for everyone" result does not depend on a second one
A1 demo-hp done, read only — two GETs of a 404 path, then the logs —
B1 calibre-web measured; not fixed — operator decision (STATUS) no knob for the daily lock; both fixes cost the household
B2 wger done — decision 58, catalog 82fff32; control + two fix runs on 9202 the first fix (15 min) proved every try during a lock restarts it; changed to 5 min
B3 Grafana done — decision 60: no change (5.0 min measured; trickle measured) already short
B4 BookStack done — decision 59: no change (1.0 min measured) already short; APP_PROXIES would not help through the tunnel
B installed apps done — measured on 9202 see below
C — the registry done, read only — cause found (R-750 answered) —
D — release / golden not done — not needed Part A built nothing

Claims in the brief that turned out wrong (or right)

  1. "Apps see traefik's address for every client" — half right. The app's TCP peer is traefik, but X-Forwarded-For carries cloudflared's container address through the tunnel (the same for everyone) and the REAL address from the LAN. Apps that read it (calibre-web's ProxyFix, TRUSTED_PROXY_COUNT 1) still see one address for every tunnel visitor.
  2. "calibre-web has no env switch" — right for the limiter (a database setting, config_ratelimiter); it has an env TRUSTED_PROXY_COUNT, irrelevant here (the login limit is keyed on the user name).
  3. "BookStack's 60 s is hard-coded" — right (ThrottlesLogins.php:82 5 tries, :90 1 minute).
  4. "A Gitea cleanup rule removed the old versions" — wrong. No rule exists; a manual prune script did (HM-024).
  5. R-752's own claims: calibre-web "up to a day" — right (measured: still locked 2 min after the minute window; only a restart cleared it). My own earlier guess that calibre-web's OPDS door had no limit — wrong: 3/minute per name. Grafana "a slow trickle keeps it closed indefinitely" — not as measured: the household got in once the burst aged out, and a success resets the count. wger "everyone at once" — right (measured).
  6. 01 §7 "cloudflared runs on the host" — the build differs: it runs in the guest (R-754).

Part A — the answer

path the app's TCP peer X-Forwarded-For / X-Real-Ip the real client is in forgeable?
tunnel traefik cloudflared's container — same for every visitor CF-Connecting-IP only XFF no (traefik drops it); CF-Connecting-IP not through the tunnel, yes from the LAN
LAN traefik the real LAN address XFF / X-Real-Ip no

Part B — per app (9202, the public name, a stranger through traefik)

app setting (pinned tag) measured before fix after
wger 2.7 settings/main.py:268-272 (AXES_* env), settings_global.py:485 reset-on-failure True 10 wrong → the second member locked too username, 5 min, DB handler (decision 58) other member fine; admin in at 7.5 min with one retry; wrong still refused
BookStack 26.09.1 ThrottlesLogins.php:66,82,90 locked 1.0 min none (59) —
Grafana 13.2.3 login_attempt.go:14,65-85, defaults.ini:498-507 locked 5.0 min; trickle: in after the burst aged none (60) —
calibre-web-automated v4.0.8 cps/web.py:2218-2219 (3/min, 40/day per name), cps/main.py:75 (OPDS 3/min) form 1.2 min; 40 wrong in 14 min → refused 2+ min later; restart cleared operator —

What an installed app gets, and when (measured with wger): a settings-only change reaches the app's stack file at the next catalog sync (when its images equal the catalog's; ≤ 15 min); the RUNNING app keeps the old value until the next compose up -d — the app page's Restart or Start (measured: the env changed exactly at Restart), an Update, or a backup's restart of the app (backup.go:972, read, not measured). An app pinned to an older version than the catalog gets nothing until its Update (the frozen render, 09 §5.4).

Part C — the registry (read only)

package_cleanup_rule empty; Gitea logs only to the console and the pod started 2026-08-23, so August logs are gone. The cause is recorded in homelab-manifests HM-024: gitea-image-prune.sh --all --keep 7 --apply --reclaim the night of 2026-08-22/23 (the Gitea volume was full). Nothing schedules it. STATUS carries the decision (keep / a written rule, pick: a rule).

Rows

387 → 390. Opened R-753 (one address behind the tunnel), R-754 (01 §7 vs the build), R-755 (wger on runserver). Narrowed R-752. Answered R-750 (waiting on the operator). Closed none.

Teardown

  • Machine: 9202 back on the live catalog (repo_url read back), the same six containers as at the start; the apps this session installed (bookstack, grafana, calibre-web, wger ×5) removed through the product — calibre-web's drive data kept because the remove refused the drive path (R-442's fail-closed rule; its folder predates today); the echo container and both probe images removed. Drill catalog reset to live (82fff32).
  • Host: demo-hp untouched except two read-only GETs through its tunnel and log reads; pct list unchanged.
  • Hub: nothing. Gitea / DooPlex: read only (one READ ONLY database transaction, config and log reads).