From 8dab40c7a7867d76b06b06449e7bb12d83c78482 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Thu, 1 Oct 2026 12:49:20 +0200 Subject: [PATCH] Lockouts (R-752): decisions 58-60 (decided by CC unattended), the one address behind the tunnel (R-753), the registry answered (R-750); STATUS, report, evidence Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- CONTEXT.md | 8 ++ REPORT-lockouts-2026-10-01.md | 79 +++++++++++++++++++ STATUS.md | 68 ++++++++-------- .../architecture/09-update-architecture.md | 25 ++++++ .../audits/lockouts-2026-10-01/B/B3-drill.txt | 1 + .../B/B4-wger-fix-drill.txt | 15 ++++ .../B/B5-wger-fix-5min.txt | 37 +++++++++ .../B/B6-installed-app-gets-setting.txt | 14 ++++ .../audits/lockouts-2026-10-01/README.md | 15 ++++ .../T/T1-9202-repoint-live.txt | 16 ++++ .../lockouts-2026-10-01/T/T2-drill-reset.txt | 6 ++ .../lockouts-2026-10-01/tools/b_installed.py | 33 ++++++++ .../lockouts-2026-10-01/tools/b_wger.py | 7 +- documentation/backlog/OPEN-ITEMS.md | 4 +- 14 files changed, 287 insertions(+), 41 deletions(-) create mode 100644 REPORT-lockouts-2026-10-01.md create mode 100644 documentation/audits/lockouts-2026-10-01/B/B5-wger-fix-5min.txt create mode 100644 documentation/audits/lockouts-2026-10-01/B/B6-installed-app-gets-setting.txt create mode 100644 documentation/audits/lockouts-2026-10-01/README.md create mode 100644 documentation/audits/lockouts-2026-10-01/T/T1-9202-repoint-live.txt create mode 100644 documentation/audits/lockouts-2026-10-01/T/T2-drill-reset.txt create mode 100644 documentation/audits/lockouts-2026-10-01/tools/b_installed.py diff --git a/CONTEXT.md b/CONTEXT.md index 71b0f75c..e5b0da11 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -16,6 +16,14 @@ > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +> **2026-10-01 (afternoon) — strangers and lockouts (R-752), the one address behind the tunnel (R-753), the registry (R-750).** +> Decisions 58–60 (decided by CC unattended — operator may reverse): wger `AXES_LOCKOUT_PARAMETERS=username`, +> `AXES_COOLOFF_TIME=5`, database handler (catalog `82fff32`); BookStack and Grafana unchanged (1 and 5 min, measured). +> calibre-web's day-long lock waits for the operator. Behind the tunnel every visitor reaches an app as cloudflared's +> container address; traefik drops forged `X-Forwarded-For`; `CF-Connecting-IP` is forgeable from the LAN — no box-wide +> fix (no controller release). Registry: no Gitea cleanup rule; a manual `gitea-image-prune.sh --keep 7` run (homelab +> HM-024, 2026-08-22/23) removed the old versions; nothing schedules it. Report: `REPORT-lockouts-2026-10-01.md`. + > **Operator note 2026-10-01 (afternoon):** decision 57 kept — mealie stays on the 1-hour lock; no secret login name. > **2026-10-01 — the three rulings built; decision 57 decided by CC unattended (operator may reverse).** 54/55: the first diff --git a/REPORT-lockouts-2026-10-01.md b/REPORT-lockouts-2026-10-01.md new file mode 100644 index 00000000..28526c47 --- /dev/null +++ b/REPORT-lockouts-2026-10-01.md @@ -0,0 +1,79 @@ +# REPORT — strangers and lockouts (R-752), the one address behind the tunnel (R-753), the registry (R-750) — 2026-10-01 afternoon + +Evidence: `documentation/audits/lockouts-2026-10-01/` (A, B, C, T, tools). +Architecture read: `01-topology-and-trust.md` §5, §7; `09` §3 decisions 45–47, 57; `06` (the tunnel is not described +there). Baselines (live Gitea ~10:55 CEST): controller `c1b123c64955`, agent `d766666ff8cf`, felhom.eu `a6a9f0b2458e`, +catalog `83636352ea10` — all matched. Register 387 rows; highest R-752; last decision 57. + +## The Part table + +| Part | done / not done / changed | why | +|---|---|---| +| Operator note (decision 57 kept) | **done** — `09` §3 + CONTEXT | first | +| **A — the client address** | **done — measured; no box-wide fix** (R-753) | trusting cloudflared would pass a client-written leftmost address; a single-address rewrite needs a plugin | +| A1 two outside addresses | **changed — one** (DooPlex 37.191.56.193; no IPv6 here) | the "same address for everyone" result does not depend on a second one | +| A1 demo-hp | **done, read only** — two GETs of a 404 path, then the logs | — | +| **B1 calibre-web** | **measured; not fixed — operator decision** (STATUS) | no knob for the daily lock; both fixes cost the household | +| **B2 wger** | **done — decision 58**, catalog `82fff32`; control + two fix runs on 9202 | the first fix (15 min) proved every try during a lock restarts it; changed to 5 min | +| **B3 Grafana** | **done — decision 60: no change** (5.0 min measured; trickle measured) | already short | +| **B4 BookStack** | **done — decision 59: no change** (1.0 min measured) | already short; `APP_PROXIES` would not help through the tunnel | +| B installed apps | **done** — measured on 9202 | see below | +| **C — the registry** | **done, read only** — cause found (R-750 answered) | — | +| **D — release / golden** | **not done — not needed** | Part A built nothing | + +## Claims in the brief that turned out wrong (or right) + +1. **"Apps see traefik's address for every client"** — half right. The app's TCP peer is traefik, but `X-Forwarded-For` + carries cloudflared's container address through the tunnel (the same for everyone) and the REAL address from the LAN. + Apps that read it (calibre-web's ProxyFix, `TRUSTED_PROXY_COUNT` 1) still see one address for every tunnel visitor. +2. **"calibre-web has no env switch"** — right for the limiter (a database setting, `config_ratelimiter`); it has an env + `TRUSTED_PROXY_COUNT`, irrelevant here (the login limit is keyed on the user name). +3. **"BookStack's 60 s is hard-coded"** — right (`ThrottlesLogins.php:82` 5 tries, `:90` 1 minute). +4. **"A Gitea cleanup rule removed the old versions"** — wrong. No rule exists; a manual prune script did (HM-024). +5. R-752's own claims: calibre-web "up to a day" — **right** (measured: still locked 2 min after the minute window; only a + restart cleared it). My own earlier guess that calibre-web's OPDS door had no limit — **wrong**: 3/minute per name. + Grafana "a slow trickle keeps it closed indefinitely" — **not as measured**: the household got in once the burst aged + out, and a success resets the count. wger "everyone at once" — **right** (measured). +6. `01` §7 "cloudflared runs on the host" — **the build differs**: it runs in the guest (R-754). + +## Part A — the answer + +| path | the app's TCP peer | X-Forwarded-For / X-Real-Ip | the real client is in | forgeable? | +|---|---|---|---|---| +| tunnel | traefik | cloudflared's container — same for every visitor | `CF-Connecting-IP` only | XFF no (traefik drops it); `CF-Connecting-IP` not through the tunnel, **yes from the LAN** | +| LAN | traefik | the real LAN address | XFF / X-Real-Ip | no | + +## Part B — per app (9202, the public name, a stranger through traefik) + +| app | setting (pinned tag) | measured before | fix | after | +|---|---|---|---|---| +| wger 2.7 | `settings/main.py:268-272` (`AXES_*` env), `settings_global.py:485` reset-on-failure True | 10 wrong → the second member locked too | username, 5 min, DB handler (decision 58) | other member fine; admin in at 7.5 min with one retry; wrong still refused | +| BookStack 26.09.1 | `ThrottlesLogins.php:66,82,90` | locked 1.0 min | none (59) | — | +| Grafana 13.2.3 | `login_attempt.go:14,65-85`, `defaults.ini:498-507` | locked 5.0 min; trickle: in after the burst aged | none (60) | — | +| calibre-web-automated v4.0.8 | `cps/web.py:2218-2219` (3/min, 40/day per name), `cps/main.py:75` (OPDS 3/min) | form 1.2 min; 40 wrong in 14 min → refused 2+ min later; restart cleared | **operator** | — | + +**What an installed app gets, and when (measured with wger):** a settings-only change reaches the app's stack file at the +next catalog sync (when its images equal the catalog's; ≤ 15 min); the RUNNING app keeps the old value until the next +`compose up -d` — the app page's Restart or Start (measured: the env changed exactly at Restart), an Update, or a +backup's restart of the app (`backup.go:972`, read, not measured). An app pinned to an older version than the catalog +gets nothing until its Update (the frozen render, `09` §5.4). + +## Part C — the registry (read only) + +`package_cleanup_rule` empty; Gitea logs only to the console and the pod started 2026-08-23, so August logs are gone. The +cause is recorded in homelab-manifests HM-024: `gitea-image-prune.sh --all --keep 7 --apply --reclaim` the night of +2026-08-22/23 (the Gitea volume was full). Nothing schedules it. STATUS carries the decision (keep / a written rule, pick: a rule). + +## Rows + +**387 → 390.** Opened R-753 (one address behind the tunnel), R-754 (`01` §7 vs the build), R-755 (wger on runserver). +Narrowed R-752. Answered R-750 (waiting on the operator). Closed none. + +## Teardown + +- **Machine:** 9202 back on the live catalog (`repo_url` read back), the same six containers as at the start; the apps + this session installed (bookstack, grafana, calibre-web, wger ×5) removed through the product — calibre-web's drive + data kept because the remove refused the drive path (R-442's fail-closed rule; its folder predates today); the echo + container and both probe images removed. Drill catalog reset to live (`82fff32`). +- **Host:** demo-hp untouched except two read-only GETs through its tunnel and log reads; `pct list` unchanged. +- **Hub:** nothing. **Gitea / DooPlex:** read only (one READ ONLY database transaction, config and log reads). diff --git a/STATUS.md b/STATUS.md index cf6de322..d9240a6c 100644 --- a/STATUS.md +++ b/STATUS.md @@ -2,54 +2,50 @@ **Ready for the first real tester (Tester-2): yes. You confirmed the tunnel route and the connect mails (2026-09-30).** -**Updated 2026-10-01. Both demo boxes run controller 0.285.0 and host agent 0.138.0. Hub 0.126.0. New installs get golden 0.285.0 with agent 0.138.0.** +**Updated 2026-10-01 (afternoon). Both demo boxes run controller 0.285.0 and host agent 0.138.0. Hub 0.126.0. New installs get golden 0.285.0 with agent 0.138.0.** **Tester-2 — read only, from the hub.** The customer record exists. Tester-2's box has not registered yet. -## A decision I took myself (you may reverse it) +## Decisions I took myself (you may reverse them) -- **mealie: a stranger can now lock the household out for 1 to 2 hours, not a whole day.** mealie locks the account after - five wrong passwords, and its login name is public. I set the lock to mealie's shortest time (1 hour; its hourly - clean-up makes it up to 2). Tested on the scratch box: locked for 2 hours, then the right password worked. The guard - against password guessing stays. **What it does not fix:** a stranger can lock it again every hour. The other way would - be a secret login name, but that changes what the household types — so I did not do it. +- **wger: a stranger now locks only the name they try, for 5 minutes.** Before, ten wrong tries locked out every + household member for 30 minutes. Tested: the other member could still log in; the targeted one was back in after + 7.5 minutes. Short on purpose: in wger, every try during a lock restarts it, also the household's own. +- **BookStack and Grafana: no change.** Their locks are already short: 1 minute and 5 minutes, measured. +- **mealie stays on the 1-hour lock**, as you said. -## Your three decisions of today — built +## What I measured -- **The monthly security re-test covers every tested app (your 2A).** First full run today: **nextcloud and sonarr** had - a new build under the same name. Both passed on the test bench and on the scratch box, and are now in the catalog. Boxes - take them at night. It took about 17 minutes per app. -- **You start the monthly run (your 1A).** Give the standing brief `MONTHLY-security-retest.md` to CC once a month. The - runbook now says so. -- **A box keeps only two controller versions (your 3A).** The new controller 0.285.0 deletes the older ones. The HP demo - box went from 84 versions to 2 (Docker images 15.2 → 11.7 GB). The N100 went from 76 to 2 (6.1 → 1.1 GB). All apps kept - running. I checked first: if an update fails, the box goes back to the version it was running, and that one is always kept. +- **Behind the tunnel, every visitor looks like the same address** to an app. So an app that locks "by address" locks + out everyone at once (that was wger). Apps cannot see who is who. I did not change this box-wide: the easy way would + let a stranger fake an address. Written down. +- **Old controller versions in the registry:** no automatic rule removed them. Someone ran a clean-up script by hand + on 22–23 August, when the registry disk was full. Nothing runs it on a schedule. +- **An installed app takes a new setting** at its next restart, update or night backup — not by itself. -## What else I did, and it worked +## What else I found -- **A crash fixed.** The image clean-up after an app update could crash the controller if the app was removed right - after. The test suite found it. Fixed in 0.285.0. -- **Two test-tool fixes:** the digest check now refuses an image build that is gone, and the outline test works with - outline's newest version. -- **New controller 0.285.0** is on both demo boxes, and **a new golden 0.285.0** is baked and vouched. The golden check is - green. +- wger runs a development web server, not a production one. Written down. +- The design document says the tunnel runs on the host. It runs inside the box. One of the two is wrong. Written down. -## What broke, and what I did - -- **The monthly re-test could not start on a fresh test bench.** It checked for its own files before it copied them. Fixed. -- **Four more apps can be locked by strangers like mealie:** calibre-web (up to a day), wger (30 minutes, and everyone at - once), Grafana (5 minutes), BookStack (1 minute). I read this in their code; not tested on a box yet. Written down. -- **Very old controller versions are gone from our own registry** (everything before mid-August). Nothing needs them - today. I do not know what removed them. Written down. - -**Rows.** Today: 6 closed, 1 narrowed, 4 opened. The list went from 383 to 387 rows. +**Rows.** Today (afternoon): 0 closed, 2 narrowed or answered, 3 opened. The list went from 387 to 390 rows. ## What needs you -1. **Nothing urgent.** If you want the household to be safe from the hourly re-lock in mealie, say so: then I give it a - secret login name, and the app page shows it. If you do nothing, a stranger who keeps trying can keep mealie locked. -2. **Do you know what removed the old controller versions from the registry?** If you do nothing, I look at the Gitea - settings next time (read only). +1. **calibre-web: a stranger can lock the household out for a day** (40 wrong tries in 14 minutes). There is no setting + for a shorter lock. Two ways: + - **A — a secret login name** (recommended): the box makes one at install and shows it on the app page next to + the password. Strangers cannot aim at it. Cost: the household types a strange name. + - **B — switch the login limit off:** no lock at all. Cost: no limit on password guessing in the login form. + If you do nothing: a stranger who knows the name `admin` can keep the household out for a day. +2. **The registry clean-up script:** today it keeps only the newest 7 versions, about one day of controller releases, + and it does not protect the versions in use. Two ways: + - **A — a written rule** (recommended): keep the newest 20 releases, plus every version a golden, the floor or the + vouched agent names. Change the script to match. + - **B — keep it as it is:** run by hand when the disk fills. + If you do nothing: the next manual run can delete a version a box or a backup still needs. +3. **Smaller:** should apps see each visitor's real address (a bigger build), or do we keep fixing per app? And which is + right for the tunnel: the design document (on the host) or the build (inside the box)? Nothing breaks if you wait. ## Standing steps diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index f705452c..e283df29 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -637,6 +637,31 @@ R-636's louder repeated alarm. every hour. Catalog `a4597cd`. **Operator, 2026-10-01 (afternoon): kept** — mealie stays on the 1-hour lock; no secret login name (option d not taken). +### 2026-10-01 (afternoon) — decided by CC unattended, operator may reverse (R-752) + +The question for all three: how long may a stranger's wrong passwords, aimed at a PUBLIC login name, keep the household +out? Behind the tunnel every visitor has one address (R-753), so no fix may lean on the visitor's address or on any +header a client can send. Each measured on 9202 (`audits/lockouts-2026-10-01/B/`). + +58. **wger: lock only the targeted name, for 5 minutes, counted in the database** — `AXES_LOCKOUT_PARAMETERS=username`, + `AXES_COOLOFF_TIME=5`, `AXES_HANDLER=axes.handlers.database.AxesDatabaseHandler` (catalog `82fff32`). **Options:** (a) + keep ip_address — 10 wrong tries lock EVERY member for 30 min (measured: the second member locked too); (b) username, + 30 min; (c) username, 5 min; (d) axes off — no guard. **Why (c):** wger 2.7 hard-codes + `AXES_RESET_COOL_OFF_ON_FAILURE_DURING_LOCKOUT = True`, so every try during a lock restarts it — measured: retrying + every 8 minutes kept a 15-minute lock closed for 40+ minutes; at 5 minutes one retry still let the household in at + 7.5 min. The guard stays: 10 tries per 5 minutes per name. The database handler answers axes' own warning W001 (the + default cache is per process) and keeps the count over a restart. +59. **BookStack: no change** — its throttle is 5 tries then 60 s, hard-coded (`ThrottlesLogins.php:82,90`), keyed + `email|ip` where ip is traefik's (`APP_PROXIES` empty), so in effect per name. Measured: locked 1.0 min, then the + right password works. `APP_PROXIES` would only move the key to the tunnel's one address (R-753) — no gain. +60. **Grafana: no change** — per name, 5 failures in a sliding 5 minutes (`loginattemptimpl/login_attempt.go:14`, + `defaults.ini:498-507`); a blocked try is not counted and a successful login resets the count. Measured: locked 5.0 + min; under a one-try-a-minute trickle the household got in after the burst aged out. The lock shows as "wrong + password" to the household (Grafana hides it). Raising the attempt limit would not stop a script and weakens the guard. + +calibre-web-automated is NOT decided here: its only long lock (40 tries a day per name, `cps/web.py:2218`) has no knob for +its length, and both fixes cost something the household would notice — operator decision in STATUS. + ### 2026-09-30 (day) — operator notes, recorded before the work - **The day brief runs by day.** Every backup and automatic-update test is started by hand — the night chain's debug diff --git a/documentation/audits/lockouts-2026-10-01/B/B3-drill.txt b/documentation/audits/lockouts-2026-10-01/B/B3-drill.txt index 73dc4326..5f046e00 100644 --- a/documentation/audits/lockouts-2026-10-01/B/B3-drill.txt +++ b/documentation/audits/lockouts-2026-10-01/B/B3-drill.txt @@ -1,3 +1,4 @@ # drill 2026-10-01T09:47:56Z drill a4597cd live 8363635 2e9514d DRILL wger: axes by username, 15 min, database handler (R-752 proof on 9202) +2026-10-01T10:31:39Z drill: 6b0520e DRILL wger: cool-off 5 min (R-752 proof) diff --git a/documentation/audits/lockouts-2026-10-01/B/B4-wger-fix-drill.txt b/documentation/audits/lockouts-2026-10-01/B/B4-wger-fix-drill.txt index 3aa97109..27ee3e46 100644 --- a/documentation/audits/lockouts-2026-10-01/B/B4-wger-fix-drill.txt +++ b/documentation/audits/lockouts-2026-10-01/B/B4-wger-fix-drill.txt @@ -23,3 +23,18 @@ python3 manage.py runserver 0.0.0.0:8000 09:50:26 household — admin RIGHT password: locked | second member RIGHT password: ok 09:58:26 +8.0 min admin RIGHT password DURING the lock (does it restart the 15 min?): locked 09:58:26 second member, same moment: ok +10:06:26 +16.0 min admin RIGHT password: locked +10:14:27 +24.0 min admin RIGHT password: locked +10:22:27 +32.0 min admin RIGHT password: locked +10:30:27 +40.0 min admin RIGHT password: locked +10:30:27 a wrong one after: locked | right again: locked +10:30:31 level=WARNING ts=2026-10-01 12:30:27,488 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=223 message=AXES: Repeated login failure by {username: "********************", i +level=WARNING ts=2026-10-01 12:30:27,498 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=253 message=AXES: Locking out {username: "********************", ip_address: "** +level=INFO ts=2026-10-01 12:30:27,662 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=130 message=AXES: User login failed, running database handler for failure. +level=INFO ts=2026-10-01 12:30:27,662 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=425 message=AXES: Cleaned up 0 expired access attempts from database that were olde +level=WARNING ts=2026-10-01 12:30:27,665 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=223 message=AXES: Repeated login failure by {username: "********************", i +level=WARNING ts=2026-10-01 12:30:27,673 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=253 message=AXES: Locking out {username: "********************", ip_address: "** +12:30:41 [X] stop -> 200 {'ok': True, 'message': 'Stack wger stop completed'} +12:31:13 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'wger', 'volumes_removed': ['wger_wger_data', 'wger_wger_media'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note +12:31:21 [X] after remove: deployed=False leftovers='/opt/docker/stacks/wger' +10:31:21 removed; deployed = False diff --git a/documentation/audits/lockouts-2026-10-01/B/B5-wger-fix-5min.txt b/documentation/audits/lockouts-2026-10-01/B/B5-wger-fix-5min.txt new file mode 100644 index 00000000..8e5c7a65 --- /dev/null +++ b/documentation/audits/lockouts-2026-10-01/B/B5-wger-fix-5min.txt @@ -0,0 +1,37 @@ +10:31:42 ##### wger fix — controller gitea.dooplex.hu/admin/felhom-controller:0.285.0 +10:31:53 the box's catalog copy: - AXES_LOCKOUT_PARAMETERS=username +12:31:53 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['ADMIN_PASSWORD'] +12:31:53 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +12:33:28 [1] deployed, controller state=running, pinned={'wger': 'wger/server:2.7'} +10:33:28 deploy: True +10:33:41 axes settings inside the app: ['username'] 10 0:05:00 True axes.handlers.database.AxesDatabaseHandler django.core.cache.backends.locmem.LocMemCache +python3 manage.py runserver 0.0.0.0:8000 +/usr/bin/python3 manage.py runserver 0.0.0.0:8000 +10:33:46 second member: made second +10:33:46 positive controls — admin: ok | second: ok +10:33:47 stranger wrong try 1 (admin): wrong +10:33:47 stranger wrong try 2 (admin): wrong +10:33:47 stranger wrong try 3 (admin): wrong +10:33:48 stranger wrong try 4 (admin): wrong +10:33:48 stranger wrong try 5 (admin): wrong +10:33:48 stranger wrong try 6 (admin): wrong +10:33:49 stranger wrong try 7 (admin): wrong +10:33:49 stranger wrong try 8 (admin): wrong +10:33:49 stranger wrong try 9 (admin): wrong +10:33:50 stranger wrong try 10 (admin): locked +10:33:50 stranger wrong try 11 (admin): locked +10:33:50 household — admin RIGHT password: locked | second member RIGHT password: ok +10:35:50 +2.0 min admin RIGHT password DURING the lock (does it restart the cool-off?): locked +10:35:51 second member, same moment: ok +10:41:21 +7.5 min admin RIGHT password: ok +10:41:22 a wrong one after: wrong | right again: ok +10:41:25 level=INFO ts=2026-10-01 12:41:21,471 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=425 message=AXES: Cleaned up 1 expired access attempts from database that were olde +level=INFO ts=2026-10-01 12:41:21,819 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=130 message=AXES: User login failed, running database handler for failure. +level=INFO ts=2026-10-01 12:41:21,820 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=425 message=AXES: Cleaned up 0 expired access attempts from database that were olde +level=WARNING ts=2026-10-01 12:41:21,822 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=201 message=AXES: New login failure by {username: "********************", ip_add +level=INFO ts=2026-10-01 12:41:22,053 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=295 message=AXES: Successful login by {username: "********************", ip_address +level=INFO ts=2026-10-01 12:41:22,054 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=425 message=AXES: Cleaned up 0 expired access attempts from database that were olde +12:41:35 [X] stop -> 200 {'ok': True, 'message': 'Stack wger stop completed'} +12:42:07 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'wger', 'volumes_removed': ['wger_wger_data', 'wger_wger_media'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note +12:42:15 [X] after remove: deployed=False leftovers='/opt/docker/stacks/wger' +10:42:15 removed; deployed = False diff --git a/documentation/audits/lockouts-2026-10-01/B/B6-installed-app-gets-setting.txt b/documentation/audits/lockouts-2026-10-01/B/B6-installed-app-gets-setting.txt new file mode 100644 index 00000000..44e4e4ec --- /dev/null +++ b/documentation/audits/lockouts-2026-10-01/B/B6-installed-app-gets-setting.txt @@ -0,0 +1,14 @@ +12:42:24 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['ADMIN_PASSWORD'] +12:42:24 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +12:44:00 [1] deployed, controller state=running, pinned={'wger': 'wger/server:2.7'} +10:44:00 deploy: True +10:44:23 running env: AXES_COOLOFF_TIME=5 | container created 2026-10-01T10:42:53.627288276Z +10:44:24 drill: ea94447 DRILL wger: cool-off 6 (installed-app probe) +10:44:31 stack-dir compose after 7 s of sync rounds: - AXES_COOLOFF_TIME=6 +10:44:34 running env (no action yet): AXES_COOLOFF_TIME=5 | container created 2026-10-01T10:42:53.627288276Z +10:44:45 press Restart: ('200', {'ok': True, 'data': {'state': 'starting'}, 'message': 'Stack wger restart requested — state now: starting'}) +10:45:18 running env after Restart: AXES_COOLOFF_TIME=6 | container created 2026-10-01T10:44:34.755369792Z +12:45:29 [X] stop -> 200 {'ok': True, 'message': 'Stack wger stop completed'} +12:46:01 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'wger', 'volumes_removed': ['wger_wger_data', 'wger_wger_media'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note +12:46:09 [X] after remove: deployed=False leftovers='/opt/docker/stacks/wger' +10:46:09 removed; deployed = False diff --git a/documentation/audits/lockouts-2026-10-01/README.md b/documentation/audits/lockouts-2026-10-01/README.md new file mode 100644 index 00000000..3bea299d --- /dev/null +++ b/documentation/audits/lockouts-2026-10-01/README.md @@ -0,0 +1,15 @@ +# Strangers and lockouts (R-752), the one address behind the tunnel (R-753), the registry (R-750) — 2026-10-01 afternoon + +Report and Part table: `felhom.eu/REPORT-lockouts-2026-10-01.md`. + +| folder | what | +|---|---| +| `A/A1-client-address.txt` | demo-hp through its real tunnel (traefik + BookStack logs), 9202 echo container (tunnel hop simulated, LAN with forged headers), the answer table, why no box-wide fix | +| `B/B1` | BookStack, Grafana (+ trickle), calibre-web (+ OPDS, + the daily limit, + restart) with the LIVE templates | +| `B/B2a`, `B/B2` | wger control (B2a: a wrong login URL — 404 — kept as a record; B2: the real control) | +| `B/B3` | the drill commits and 9202 pointed at the drill | +| `B/B4`, `B/B5` | wger fix at 15 min (retries kept it closed 40+ min) and at 5 min (in at 7.5 min) | +| `B/B6` | what an installed app gets: the stack file at sync, the running env only at Restart | +| `C/C1-registry-read.txt` | Gitea read only: no cleanup rule, the database's first versions, the manual prune (HM-024), nothing scheduled | +| `T/` | 9202 back on live, the drill reset | +| `tools/` | `lk.py` (login probes), `b_three.py`, `b_wger.py`, `b_installed.py`, `echo.sh`, `walk.py`, `repoint.py` | diff --git a/documentation/audits/lockouts-2026-10-01/T/T1-9202-repoint-live.txt b/documentation/audits/lockouts-2026-10-01/T/T1-9202-repoint-live.txt new file mode 100644 index 00000000..db67da0f --- /dev/null +++ b/documentation/audits/lockouts-2026-10-01/T/T1-9202-repoint-live.txt @@ -0,0 +1,16 @@ +git: + branch: main + repo_url: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git + sync_interval: 15m + token: + username: "" +hub: +0 + +felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.285.0 +filebrowser gtstef/filebrowser:1.3.3-stable +paperless-postgres postgres:18-alpine +paperless-redis redis:7-alpine +paperless-webserver ghcr.io/paperless-ngx/paperless-ngx:2.20.15 +traefik traefik:v3.6.7 +0 diff --git a/documentation/audits/lockouts-2026-10-01/T/T2-drill-reset.txt b/documentation/audits/lockouts-2026-10-01/T/T2-drill-reset.txt new file mode 100644 index 00000000..ec5c9e02 --- /dev/null +++ b/documentation/audits/lockouts-2026-10-01/T/T2-drill-reset.txt @@ -0,0 +1,6 @@ +# drill reset 2026-10-01T10:47:28Z +before: ea94447 live: 82fff32 +ea94447 DRILL wger: cool-off 6 (installed-app probe) +6b0520e DRILL wger: cool-off 5 min (R-752 proof) +2e9514d DRILL wger: axes by username, 15 min, database handler (R-752 proof on 9202) +after: 82fff32 diff --git a/documentation/audits/lockouts-2026-10-01/tools/b_installed.py b/documentation/audits/lockouts-2026-10-01/tools/b_installed.py new file mode 100644 index 00000000..647124b1 --- /dev/null +++ b/documentation/audits/lockouts-2026-10-01/tools/b_installed.py @@ -0,0 +1,33 @@ +#!/usr/bin/env python3 +"""What an INSTALLED app gets when the catalog changes only a SETTING (R-752): wger installed from the drill +(AXES_COOLOFF_TIME=5); then a drill-only commit sets 6; sync + rescan; read the stack-dir compose and the RUNNING +container's env; press the app page's Restart (POST /api/stacks/wger/restart); read the env again. Removed after.""" +import subprocess, time +import walk as w +import lk + +D = "/mnt/5_hdd/felhom.eu/drill/app-catalog-drill" +ENV = "docker inspect wger --format '{{range .Config.Env}}{{println .}}{{end}}' | grep AXES_COOLOFF; docker inspect wger --format 'container created {{.Created}}'" +FILE = "grep -h AXES_COOLOFF /opt/docker/stacks/wger/docker-compose.yml" +w.login() +lk.p("deploy:", w.deploy("wger", "fitness")) +time.sleep(20) +lk.p("running env:", w.guest(ENV).strip().replace("\n", " | ")) +r = subprocess.run(["bash", "-c", f"cd {D} && sed -i 's/^ - AXES_COOLOFF_TIME=5$/ - AXES_COOLOFF_TIME=6/' templates/wger/docker-compose.yml && " + "git commit -qam 'DRILL wger: cool-off 6 (installed-app probe)' && git push -q origin main 2>&1 | grep -v ^remote; git log --oneline -1"], + capture_output=True, text=True) +lk.p("drill:", r.stdout.strip(), r.stderr.strip()[-200:]) +t0 = time.time() +for _ in range(20): + w.sync_rescan() + f = w.guest(FILE).strip() + if "=6" in f: + break + time.sleep(15) +lk.p(f"stack-dir compose after {(time.time() - t0):.0f} s of sync rounds:", f) +lk.p("running env (no action yet):", w.guest(ENV).strip().replace("\n", " | ")) +lk.p("press Restart:", w.ctl("POST", "/api/stacks/wger/restart")) +time.sleep(30) +lk.p("running env after Restart:", w.guest(ENV).strip().replace("\n", " | ")) +w.remove("wger") +lk.p("removed; deployed =", w.stack("wger").get("deployed")) diff --git a/documentation/audits/lockouts-2026-10-01/tools/b_wger.py b/documentation/audits/lockouts-2026-10-01/tools/b_wger.py index f512fb2b..17f2edbe 100644 --- a/documentation/audits/lockouts-2026-10-01/tools/b_wger.py +++ b/documentation/audits/lockouts-2026-10-01/tools/b_wger.py @@ -18,7 +18,7 @@ if MODE == "fix": for _ in range(12): w.sync_rescan() cat = w.guest("grep -rh AXES_LOCKOUT_PARAMETERS /var/lib/docker/volumes/felhom-controller-data/_data/ --include=docker-compose.yml 2>/dev/null | sort -u").strip() - if "username" in cat: + if "COOLOFF_TIME=5" in w.guest("grep -rh AXES_COOLOFF_TIME /var/lib/docker/volumes/felhom-controller-data/_data/ --include=docker-compose.yml 2>/dev/null"): break time.sleep(15) lk.p("the box's catalog copy:", cat) @@ -55,9 +55,10 @@ if MODE == "fix": r = lk.wger(who, p or pw) lk.p(f"+{(time.time() - t0) / 60:.1f} min {label}:", r) return r - at(8, "admin RIGHT password DURING the lock (does it restart the 15 min?)") + sched = [float(x) for x in (sys.argv[2].split(",") if len(sys.argv) > 2 else ["8", "16", "24", "32", "40"])] + at(sched[0], "admin RIGHT password DURING the lock (does it restart the cool-off?)") lk.p("second member, same moment:", lk.wger("second", pw2)) - for m in (16, 24, 32, 40): + for m in sched[1:]: if at(m, "admin RIGHT password") == "ok": break lk.p("a wrong one after:", lk.wger("admin", "wrong-after"), "| right again:", lk.wger("admin", pw)) diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 00a488ad..1b79ebdb 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -861,9 +861,9 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-747** | **[P3-LOW] A stranger can lock the household out of mealie with five wrong logins.** MEASURED 2026-09-30 on 9202 (the R-741 proof): after the install hold opened, a stranger's default-login tries were refused (401) and after five of them mealie answered 423 (locked) to every login — the generated, correct password included. mealie's own brute-force guard, on an app published on the internet; the setup gate and the install hold do not cover an app after its first setup. Not measured: how long the lock lasts. **Needs:** measure the lock's length; decide whether the page tells the household what to do. `audits/night-rulings-2026-09-30/C/C3-mealie-poll.txt` **-- 2026-10-01:** Measured at v3.28.0 (source + 9202): 5 wrong logins lock the ACCOUNT (not the IP) for `SECURITY_USER_LOCKOUT_TIME` hours (default 24); the lock is lifted by an hourly job; an admin can unlock others via `POST /api/admin/users/unlock`, but the household's only admin is the locked account. Both login names are public (`admin`, `changeme@example.com`). **Fixed** (`09` §3 decision 57, decided by CC unattended — operator may reverse): `SECURITY_USER_LOCKOUT_TIME=1`, catalog `a4597cd`; on 9202 the right password answered 423 for 120 min, then 200; a wrong one still 401 (`audits/rulings-2026-10-01/C/`). **Left:** a stranger can renew the lock every hour (per account, public name — option (d) in decision 57 would end that); the page does not tell the household why it is locked; an INSTALLED mealie takes the new setting only when its compose is rendered again (not measured which act does that). | **NARROWED — the lock is 1–2 h; renewal and the page remain; owner: CC** | | **R-748** | **[P3-LOW] The register-shape gate skipped every row whose id has a letter suffix — so R-88a, R-88b and R-209a were never shape-checked, and its count read 3 short.** FOUND 2026-09-30 (late) while counting the register: `register_shape_gate.py` matched `R-\d+` only; the brief's „the reviewer's regex undercounted by 3” is the same three rows. Fixed the same session: `R-\d+[a-z]?`; decoy `suffix-row-eaten-state` (a suffixed row with its state cell eaten) seen passing with the old pattern and convicted with the new. The register is **382** rows by either count now. | **CLOSED 2026-09-30 — `scripts/register_shape_gate.py`** | | **R-749** | **[P3-LOW] `retest-floating.py` could never start on a fresh bench: it checked for `/opt/upg/upgrade-test.py` on the bench BEFORE the step that copies it there.** FOUND 2026-10-01 at the first full monthly run (decision 55): bench 9401 freshly created by the runbook, the run answered „CANNOT START — missing: the bench LXC 9401 on demo-hp with /opt/upg” in one minute. The runbook says the command syncs the bench itself — it does, but only after the check. On 2026-09-30 the bench had been synced by hand earlier, so nobody saw it. **Needs:** the check asks for what the bench must bring (docker, python3), the sync then provides `/opt/upg`. `audits/rulings-2026-10-01/A/` **-- 2026-10-01:** Fixed the same session (catalog `9e53205`): the check asks for docker + python3; after the sync `/opt/upg/upgrade-test.py` is required. The re-run started at once and finished both apps. | **CLOSED 2026-10-01 — catalog 9e53205** | -| **R-750** | **[P3-LOW] The registry no longer holds controller releases older than 0.213.0 (2026-08-12) — something removed them, and nothing records what.** MEASURED 2026-10-01 (anonymous registry API, `audits/rulings-2026-10-01/B/`): `felhom-controller` has 91 tags, the oldest release 0.213.0; `0.201.0` answers 404; Gitea's package list starts 2026-08-12. No runbook, row or memory names a clean-up. Today nothing needs those versions: a box runs a newer one, a whole-guest restore brings the guest's own Docker store back (mp0 `backup=1`), and decision 56 deletes only on the box. **But** a box or a backup that names a removed version cannot pull it again (R-698's shape, for the controller). **Needs:** find what removed them (a Gitea clean-up rule?), and record the rule — or say it was a one-time act. Read-only on DooPlex. | **OPEN — rank P3-LOW; owner: operator (Gitea settings), CC measures** | +| **R-750** | **[P3-LOW] The registry no longer holds controller releases older than 0.213.0 (2026-08-12) — something removed them, and nothing records what.** MEASURED 2026-10-01 (anonymous registry API, `audits/rulings-2026-10-01/B/`): `felhom-controller` has 91 tags, the oldest release 0.213.0; `0.201.0` answers 404; Gitea's package list starts 2026-08-12. No runbook, row or memory names a clean-up. Today nothing needs those versions: a box runs a newer one, a whole-guest restore brings the guest's own Docker store back (mp0 `backup=1`), and decision 56 deletes only on the box. **But** a box or a backup that names a removed version cannot pull it again (R-698's shape, for the controller). **Needs:** find what removed them (a Gitea clean-up rule?), and record the rule — or say it was a one-time act. Read-only on DooPlex. **-- 2026-10-01 (afternoon):** **Answered (read only, `audits/lockouts-2026-10-01/C/C1-registry-read.txt`).** Not a Gitea rule: `package_cleanup_rule` is EMPTY (READ ONLY query on the shared CNPG); `[cron.cleanup_packages]` only runs rules and Gitea's own expired-data clean-up. The cause is a MANUAL run of DooPlex's `~/git/misc-scripts/gitea-image-prune.sh --all --keep 7 --apply --reclaim` on the night of 2026-08-22/23, recorded in homelab-manifests HM-024 (the Gitea volume was full: `felhom-golden` 22 → 3 versions, /data 14.8 → 4.4 G); `--keep 7` per container package explains controller 0.213.0 (2026-08-12) as the oldest left. Nothing schedules it (crontabs, timers, cluster CronJobs read). If run again as its usage text says, it keeps 7 controller releases (about a day) and `--type generic --keep 3` would cut the agent to 3; it protects no vouched version. **Needs:** the operator's word on a written rule (STATUS). | **ANSWERED — WAITING-ON-OPERATOR: keep the manual prune or write a rule; owner: operator** | | **R-751** | **[P2] The image clean-up after an app update could crash the whole controller: it re-read the app after a rescan and dereferenced a nil stack when the app was gone.** FOUND 2026-10-01 by the full test suite (controller v0.284.2): `RetainImagesAfterUpdate` runs in a goroutine; `TestR705_TheManualLegRunsByDay` removed its temp dir under it → `panic: invalid memory address` at `image_retention.go:291`. In a box the same happens when an app is removed (or its compose vanishes) between an update's end and the clean-up — a panic in a goroutine ends the process (the agent's supervisor restarts it). Fixed in v0.285.0 the same session: it returns when the app is gone; `TestRetainImagesAfterUpdate_AppGoneDoesNotPanic` seen panicking on the old code; both retention seams are no-ops in the stacks tests (`TestMain`), so no test leaves the goroutine running. `audits/rulings-2026-10-01/B/B1-red-proofs.txt` **-- 2026-10-01:** Delivered: floor 0.285.0 reached both demo boxes in ~6 s (hub `managed floor SERVED … from declared`). | **CLOSED 2026-10-01 — controller v0.285.0, floor 0.285.0** | -| **R-752** | **[P3-LOW] Four more catalog apps let a stranger lock the household out with wrong passwords for a known login name — like mealie (R-747).** READ 2026-10-01 in each app's source at its pinned tag (not measured live): **calibre-web-automated v4.0.8** — Flask-Limiter on the login keyed on the lowercased USERNAME, 3/minute and 40/day, checked before the password; the default login is `admin` → up to a day; no env switch (a database setting). **wger 2.7** — django-axes keyed on IP, 10 failures, 30 min, each failure restarts it; behind traefik every client has traefik's IP → everyone is locked out (`AXES_*` env vars exist; `AXES_IPWARE_PROXY_COUNT` 0). **Grafana 13.2.3** — per-account, 5 failures in a sliding 5 minutes; a slow trickle keeps it closed (`GF_SECURITY_*`). **BookStack 26.09.1** — key `email|ip`, 5 tries, 60 s, hard-coded; `APP_PROXIES` empty, so the key is the e-mail alone. gokapi (3 s delay, no lock) and claper (per-IP 10/min, no account lock) cannot. **Needs:** per app, the smallest fix that keeps a guessing guard (calibre-web-automated and wger first — longest and broadest), each proven on 9202 as R-747's was. | **OPEN — rank P3-LOW; owner: CC** | +| **R-752** | **[P3-LOW] Four more catalog apps let a stranger lock the household out with wrong passwords for a known login name — like mealie (R-747).** READ 2026-10-01 in each app's source at its pinned tag (not measured live): **calibre-web-automated v4.0.8** — Flask-Limiter on the login keyed on the lowercased USERNAME, 3/minute and 40/day, checked before the password; the default login is `admin` → up to a day; no env switch (a database setting). **wger 2.7** — django-axes keyed on IP, 10 failures, 30 min, each failure restarts it; behind traefik every client has traefik's IP → everyone is locked out (`AXES_*` env vars exist; `AXES_IPWARE_PROXY_COUNT` 0). **Grafana 13.2.3** — per-account, 5 failures in a sliding 5 minutes; a slow trickle keeps it closed (`GF_SECURITY_*`). **BookStack 26.09.1** — key `email|ip`, 5 tries, 60 s, hard-coded; `APP_PROXIES` empty, so the key is the e-mail alone. gokapi (3 s delay, no lock) and claper (per-IP 10/min, no account lock) cannot. **Needs:** per app, the smallest fix that keeps a guessing guard (calibre-web-automated and wger first — longest and broadest), each proven on 9202 as R-747's was. **-- 2026-10-01 (afternoon):** **Measured on 9202, each through traefik as a stranger with the public name** (`audits/lockouts-2026-10-01/B/`): **wger** — control: 10 wrong on `admin` locked the second member too; FIXED (decision 58, catalog `82fff32`): username, 5 min, database handler — the second member unaffected, admin in again at 7.5 min (each try during a lock restarts it — measured: 8-minute retries kept a 15-minute lock closed 40+ min). **BookStack** — 1.0 min, kept (decision 59). **Grafana** — 5.0 min, kept (decision 60); a trickle did not hold the household out once the burst aged. **calibre-web-automated** — the form locks 3/min (1.2 min measured) and **40/day per name: after 40 wrong tries in 14 min the right password was refused 2 min later still; only an app restart cleared it** (in-memory store); OPDS has its own 3/min per name (`cps/main.py:75`), no daily limit. No knob for the daily length; both fixes have a household cost — operator decision in STATUS. Installed apps: a settings-only change reaches the stack file at the next sync (images equal, ≤15 min) and the running app at the next `compose up -d` — Restart/Start (measured: the env changed only at Restart), an Update, or a backup's restart (`backup.go:972`, read). | **NARROWED — wger fixed, BookStack/Grafana kept; calibre-web waits for the operator; owner: operator (calibre-web), CC** | | **R-753** | **[P3-LOW] Behind the tunnel every visitor reaches an app with the SAME address — the tunnel container's — so every per-address guard is an "everyone" guard and every app's log is blind.** MEASURED 2026-10-01 (`audits/lockouts-2026-10-01/A/A1-client-address.txt`): on demo-hp through its real tunnel, a request from DooPlex's public address reached traefik as `172.18.0.5` (cloudflared, in the guest on `traefik-public`) and BookStack as `172.18.0.3` (traefik); on 9202 an echo container showed `X-Forwarded-For`/`X-Real-Ip` = the sending container for the tunnel's hop (traefik DROPS the incoming chain — good: a client cannot forge it) and the real address from the LAN; `CF-Connecting-IP` passes untouched and is FORGEABLE from the LAN. **No box-wide fix taken:** trusting cloudflared in traefik passes Cloudflare's appended chain, whose LEFTMOST entry the client writes — every app reading the leftmost address would believe it; cloudflared's address is docker-assigned; a single-address rewrite needs a traefik plugin (a new dependency). Per-app fixes trust no header (R-752). **Needs (operator):** whether to build a safe version (cloudflared on a fixed-address network + traefik trusting only it + per-app proxy counts), or keep "one address" and fix per app. Only ONE outside address was available (DooPlex has no IPv6); a second was not measured. | **OPEN — rank P3-LOW; owner: operator (direction), CC measures** | | **R-754** | **[P3-LOW] `01-topology-and-trust.md` §7 says cloudflared runs on the Proxmox HOST as an agent-managed service; on every box it runs INSIDE the guest as a container the controller renders.** READ 2026-10-01: `felhom-controller` `internal/infra/templates/cloudflared-compose.yml.tmpl` (`container_name: cloudflared`, network `traefik-public`); demo-hp's guest 9201 runs `cloudflared` (ingress `*.enkisfelhom.hu -> https://traefik`); R-505 saw the same in VM 331. A design decision that the build does not follow — the document or the build is wrong, and only the operator decides which (R-370: a design decision is not a defect). | **OPEN — rank P3-LOW; owner: operator (which is right)** | | **R-755** | **[P3-LOW] wger runs Django's DEVELOPMENT server in production: `manage.py runserver`, because the template does not set `WGER_USE_GUNICORN=True`.** MEASURED 2026-10-01 on 9202 (`ps` in the wger container: `python3 manage.py runserver 0.0.0.0:8000`); wger 2.7's `extras/docker/production/entrypoint.sh:81-87` runs gunicorn only with that switch. Django's own documentation says runserver is not for production (one process, not hardened). Not changed this session (a different change from R-752's; needs its own bench + box proof, memory watch included). | **OPEN — rank P3-LOW; owner: CC (catalog)** |