Lockouts (R-752): decisions 58-60 (decided by CC unattended), the one address behind the tunnel (R-753), the registry answered (R-750); STATUS, report, evidence
gates / gates (push) Successful in 27s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-01 12:49:20 +02:00
parent 92a60c62bd
commit 8dab40c7a7
14 changed files with 287 additions and 41 deletions
+8
View File
@@ -16,6 +16,14 @@
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
> **2026-10-01 (afternoon) — strangers and lockouts (R-752), the one address behind the tunnel (R-753), the registry (R-750).**
> Decisions 58–60 (decided by CC unattended — operator may reverse): wger `AXES_LOCKOUT_PARAMETERS=username`,
> `AXES_COOLOFF_TIME=5`, database handler (catalog `82fff32`); BookStack and Grafana unchanged (1 and 5 min, measured).
> calibre-web's day-long lock waits for the operator. Behind the tunnel every visitor reaches an app as cloudflared's
> container address; traefik drops forged `X-Forwarded-For`; `CF-Connecting-IP` is forgeable from the LAN — no box-wide
> fix (no controller release). Registry: no Gitea cleanup rule; a manual `gitea-image-prune.sh --keep 7` run (homelab
> HM-024, 2026-08-22/23) removed the old versions; nothing schedules it. Report: `REPORT-lockouts-2026-10-01.md`.
> **Operator note 2026-10-01 (afternoon):** decision 57 kept — mealie stays on the 1-hour lock; no secret login name.
> **2026-10-01 — the three rulings built; decision 57 decided by CC unattended (operator may reverse).** 54/55: the first
+79
View File
@@ -0,0 +1,79 @@
# REPORT — strangers and lockouts (R-752), the one address behind the tunnel (R-753), the registry (R-750) — 2026-10-01 afternoon
Evidence: `documentation/audits/lockouts-2026-10-01/` (A, B, C, T, tools).
Architecture read: `01-topology-and-trust.md` §5, §7; `09` §3 decisions 45–47, 57; `06` (the tunnel is not described
there). Baselines (live Gitea ~10:55 CEST): controller `c1b123c64955`, agent `d766666ff8cf`, felhom.eu `a6a9f0b2458e`,
catalog `83636352ea10` — all matched. Register 387 rows; highest R-752; last decision 57.
## The Part table
| Part | done / not done / changed | why |
|---|---|---|
| Operator note (decision 57 kept) | **done** — `09` §3 + CONTEXT | first |
| **A — the client address** | **done — measured; no box-wide fix** (R-753) | trusting cloudflared would pass a client-written leftmost address; a single-address rewrite needs a plugin |
| A1 two outside addresses | **changed — one** (DooPlex 37.191.56.193; no IPv6 here) | the "same address for everyone" result does not depend on a second one |
| A1 demo-hp | **done, read only** — two GETs of a 404 path, then the logs | — |
| **B1 calibre-web** | **measured; not fixed — operator decision** (STATUS) | no knob for the daily lock; both fixes cost the household |
| **B2 wger** | **done — decision 58**, catalog `82fff32`; control + two fix runs on 9202 | the first fix (15 min) proved every try during a lock restarts it; changed to 5 min |
| **B3 Grafana** | **done — decision 60: no change** (5.0 min measured; trickle measured) | already short |
| **B4 BookStack** | **done — decision 59: no change** (1.0 min measured) | already short; `APP_PROXIES` would not help through the tunnel |
| B installed apps | **done** — measured on 9202 | see below |
| **C — the registry** | **done, read only** — cause found (R-750 answered) | — |
| **D — release / golden** | **not done — not needed** | Part A built nothing |
## Claims in the brief that turned out wrong (or right)
1. **"Apps see traefik's address for every client"** — half right. The app's TCP peer is traefik, but `X-Forwarded-For`
carries cloudflared's container address through the tunnel (the same for everyone) and the REAL address from the LAN.
Apps that read it (calibre-web's ProxyFix, `TRUSTED_PROXY_COUNT` 1) still see one address for every tunnel visitor.
2. **"calibre-web has no env switch"** — right for the limiter (a database setting, `config_ratelimiter`); it has an env
`TRUSTED_PROXY_COUNT`, irrelevant here (the login limit is keyed on the user name).
3. **"BookStack's 60 s is hard-coded"** — right (`ThrottlesLogins.php:82` 5 tries, `:90` 1 minute).
4. **"A Gitea cleanup rule removed the old versions"** — wrong. No rule exists; a manual prune script did (HM-024).
5. R-752's own claims: calibre-web "up to a day" — **right** (measured: still locked 2 min after the minute window; only a
restart cleared it). My own earlier guess that calibre-web's OPDS door had no limit — **wrong**: 3/minute per name.
Grafana "a slow trickle keeps it closed indefinitely" — **not as measured**: the household got in once the burst aged
out, and a success resets the count. wger "everyone at once" — **right** (measured).
6. `01` §7 "cloudflared runs on the host" — **the build differs**: it runs in the guest (R-754).
## Part A — the answer
| path | the app's TCP peer | X-Forwarded-For / X-Real-Ip | the real client is in | forgeable? |
|---|---|---|---|---|
| tunnel | traefik | cloudflared's container — same for every visitor | `CF-Connecting-IP` only | XFF no (traefik drops it); `CF-Connecting-IP` not through the tunnel, **yes from the LAN** |
| LAN | traefik | the real LAN address | XFF / X-Real-Ip | no |
## Part B — per app (9202, the public name, a stranger through traefik)
| app | setting (pinned tag) | measured before | fix | after |
|---|---|---|---|---|
| wger 2.7 | `settings/main.py:268-272` (`AXES_*` env), `settings_global.py:485` reset-on-failure True | 10 wrong → the second member locked too | username, 5 min, DB handler (decision 58) | other member fine; admin in at 7.5 min with one retry; wrong still refused |
| BookStack 26.09.1 | `ThrottlesLogins.php:66,82,90` | locked 1.0 min | none (59) | — |
| Grafana 13.2.3 | `login_attempt.go:14,65-85`, `defaults.ini:498-507` | locked 5.0 min; trickle: in after the burst aged | none (60) | — |
| calibre-web-automated v4.0.8 | `cps/web.py:2218-2219` (3/min, 40/day per name), `cps/main.py:75` (OPDS 3/min) | form 1.2 min; 40 wrong in 14 min → refused 2+ min later; restart cleared | **operator** | — |
**What an installed app gets, and when (measured with wger):** a settings-only change reaches the app's stack file at the
next catalog sync (when its images equal the catalog's; ≤ 15 min); the RUNNING app keeps the old value until the next
`compose up -d` — the app page's Restart or Start (measured: the env changed exactly at Restart), an Update, or a
backup's restart of the app (`backup.go:972`, read, not measured). An app pinned to an older version than the catalog
gets nothing until its Update (the frozen render, `09` §5.4).
## Part C — the registry (read only)
`package_cleanup_rule` empty; Gitea logs only to the console and the pod started 2026-08-23, so August logs are gone. The
cause is recorded in homelab-manifests HM-024: `gitea-image-prune.sh --all --keep 7 --apply --reclaim` the night of
2026-08-22/23 (the Gitea volume was full). Nothing schedules it. STATUS carries the decision (keep / a written rule, pick: a rule).
## Rows
**387 → 390.** Opened R-753 (one address behind the tunnel), R-754 (`01` §7 vs the build), R-755 (wger on runserver).
Narrowed R-752. Answered R-750 (waiting on the operator). Closed none.
## Teardown
- **Machine:** 9202 back on the live catalog (`repo_url` read back), the same six containers as at the start; the apps
this session installed (bookstack, grafana, calibre-web, wger ×5) removed through the product — calibre-web's drive
data kept because the remove refused the drive path (R-442's fail-closed rule; its folder predates today); the echo
container and both probe images removed. Drill catalog reset to live (`82fff32`).
- **Host:** demo-hp untouched except two read-only GETs through its tunnel and log reads; `pct list` unchanged.
- **Hub:** nothing. **Gitea / DooPlex:** read only (one READ ONLY database transaction, config and log reads).
+32 -36
View File
@@ -2,54 +2,50 @@
**Ready for the first real tester (Tester-2): yes. You confirmed the tunnel route and the connect mails (2026-09-30).**
**Updated 2026-10-01. Both demo boxes run controller 0.285.0 and host agent 0.138.0. Hub 0.126.0. New installs get golden 0.285.0 with agent 0.138.0.**
**Updated 2026-10-01 (afternoon). Both demo boxes run controller 0.285.0 and host agent 0.138.0. Hub 0.126.0. New installs get golden 0.285.0 with agent 0.138.0.**
**Tester-2 — read only, from the hub.** The customer record exists. Tester-2's box has not registered yet.
## A decision I took myself (you may reverse it)
## Decisions I took myself (you may reverse them)
- **mealie: a stranger can now lock the household out for 1 to 2 hours, not a whole day.** mealie locks the account after
five wrong passwords, and its login name is public. I set the lock to mealie's shortest time (1 hour; its hourly
clean-up makes it up to 2). Tested on the scratch box: locked for 2 hours, then the right password worked. The guard
against password guessing stays. **What it does not fix:** a stranger can lock it again every hour. The other way would
be a secret login name, but that changes what the household types — so I did not do it.
- **wger: a stranger now locks only the name they try, for 5 minutes.** Before, ten wrong tries locked out every
household member for 30 minutes. Tested: the other member could still log in; the targeted one was back in after
7.5 minutes. Short on purpose: in wger, every try during a lock restarts it, also the household's own.
- **BookStack and Grafana: no change.** Their locks are already short: 1 minute and 5 minutes, measured.
- **mealie stays on the 1-hour lock**, as you said.
## Your three decisions of today — built
## What I measured
- **The monthly security re-test covers every tested app (your 2A).** First full run today: **nextcloud and sonarr** had
a new build under the same name. Both passed on the test bench and on the scratch box, and are now in the catalog. Boxes
take them at night. It took about 17 minutes per app.
- **You start the monthly run (your 1A).** Give the standing brief `MONTHLY-security-retest.md` to CC once a month. The
runbook now says so.
- **A box keeps only two controller versions (your 3A).** The new controller 0.285.0 deletes the older ones. The HP demo
box went from 84 versions to 2 (Docker images 15.2 → 11.7 GB). The N100 went from 76 to 2 (6.1 → 1.1 GB). All apps kept
running. I checked first: if an update fails, the box goes back to the version it was running, and that one is always kept.
- **Behind the tunnel, every visitor looks like the same address** to an app. So an app that locks "by address" locks
out everyone at once (that was wger). Apps cannot see who is who. I did not change this box-wide: the easy way would
let a stranger fake an address. Written down.
- **Old controller versions in the registry:** no automatic rule removed them. Someone ran a clean-up script by hand
on 22–23 August, when the registry disk was full. Nothing runs it on a schedule.
- **An installed app takes a new setting** at its next restart, update or night backup — not by itself.
## What else I did, and it worked
## What else I found
- **A crash fixed.** The image clean-up after an app update could crash the controller if the app was removed right
after. The test suite found it. Fixed in 0.285.0.
- **Two test-tool fixes:** the digest check now refuses an image build that is gone, and the outline test works with
outline's newest version.
- **New controller 0.285.0** is on both demo boxes, and **a new golden 0.285.0** is baked and vouched. The golden check is
green.
- wger runs a development web server, not a production one. Written down.
- The design document says the tunnel runs on the host. It runs inside the box. One of the two is wrong. Written down.
## What broke, and what I did
- **The monthly re-test could not start on a fresh test bench.** It checked for its own files before it copied them. Fixed.
- **Four more apps can be locked by strangers like mealie:** calibre-web (up to a day), wger (30 minutes, and everyone at
once), Grafana (5 minutes), BookStack (1 minute). I read this in their code; not tested on a box yet. Written down.
- **Very old controller versions are gone from our own registry** (everything before mid-August). Nothing needs them
today. I do not know what removed them. Written down.
**Rows.** Today: 6 closed, 1 narrowed, 4 opened. The list went from 383 to 387 rows.
**Rows.** Today (afternoon): 0 closed, 2 narrowed or answered, 3 opened. The list went from 387 to 390 rows.
## What needs you
1. **Nothing urgent.** If you want the household to be safe from the hourly re-lock in mealie, say so: then I give it a
secret login name, and the app page shows it. If you do nothing, a stranger who keeps trying can keep mealie locked.
2. **Do you know what removed the old controller versions from the registry?** If you do nothing, I look at the Gitea
settings next time (read only).
1. **calibre-web: a stranger can lock the household out for a day** (40 wrong tries in 14 minutes). There is no setting
for a shorter lock. Two ways:
- **A — a secret login name** (recommended): the box makes one at install and shows it on the app page next to
the password. Strangers cannot aim at it. Cost: the household types a strange name.
- **B — switch the login limit off:** no lock at all. Cost: no limit on password guessing in the login form.
If you do nothing: a stranger who knows the name `admin` can keep the household out for a day.
2. **The registry clean-up script:** today it keeps only the newest 7 versions, about one day of controller releases,
and it does not protect the versions in use. Two ways:
- **A — a written rule** (recommended): keep the newest 20 releases, plus every version a golden, the floor or the
vouched agent names. Change the script to match.
- **B — keep it as it is:** run by hand when the disk fills.
If you do nothing: the next manual run can delete a version a box or a backup still needs.
3. **Smaller:** should apps see each visitor's real address (a bigger build), or do we keep fixing per app? And which is
right for the tunnel: the design document (on the host) or the build (inside the box)? Nothing breaks if you wait.
## Standing steps
@@ -637,6 +637,31 @@ R-636's louder repeated alarm.
every hour. Catalog `a4597cd`.
**Operator, 2026-10-01 (afternoon): kept** — mealie stays on the 1-hour lock; no secret login name (option d not taken).
### 2026-10-01 (afternoon) — decided by CC unattended, operator may reverse (R-752)
The question for all three: how long may a stranger's wrong passwords, aimed at a PUBLIC login name, keep the household
out? Behind the tunnel every visitor has one address (R-753), so no fix may lean on the visitor's address or on any
header a client can send. Each measured on 9202 (`audits/lockouts-2026-10-01/B/`).
58. **wger: lock only the targeted name, for 5 minutes, counted in the database** — `AXES_LOCKOUT_PARAMETERS=username`,
`AXES_COOLOFF_TIME=5`, `AXES_HANDLER=axes.handlers.database.AxesDatabaseHandler` (catalog `82fff32`). **Options:** (a)
keep ip_address — 10 wrong tries lock EVERY member for 30 min (measured: the second member locked too); (b) username,
30 min; (c) username, 5 min; (d) axes off — no guard. **Why (c):** wger 2.7 hard-codes
`AXES_RESET_COOL_OFF_ON_FAILURE_DURING_LOCKOUT = True`, so every try during a lock restarts it — measured: retrying
every 8 minutes kept a 15-minute lock closed for 40+ minutes; at 5 minutes one retry still let the household in at
7.5 min. The guard stays: 10 tries per 5 minutes per name. The database handler answers axes' own warning W001 (the
default cache is per process) and keeps the count over a restart.
59. **BookStack: no change** — its throttle is 5 tries then 60 s, hard-coded (`ThrottlesLogins.php:82,90`), keyed
`email|ip` where ip is traefik's (`APP_PROXIES` empty), so in effect per name. Measured: locked 1.0 min, then the
right password works. `APP_PROXIES` would only move the key to the tunnel's one address (R-753) — no gain.
60. **Grafana: no change** — per name, 5 failures in a sliding 5 minutes (`loginattemptimpl/login_attempt.go:14`,
`defaults.ini:498-507`); a blocked try is not counted and a successful login resets the count. Measured: locked 5.0
min; under a one-try-a-minute trickle the household got in after the burst aged out. The lock shows as "wrong
password" to the household (Grafana hides it). Raising the attempt limit would not stop a script and weakens the guard.
calibre-web-automated is NOT decided here: its only long lock (40 tries a day per name, `cps/web.py:2218`) has no knob for
its length, and both fixes cost something the household would notice — operator decision in STATUS.
### 2026-09-30 (day) — operator notes, recorded before the work
- **The day brief runs by day.** Every backup and automatic-update test is started by hand — the night chain's debug
@@ -1,3 +1,4 @@
# drill 2026-10-01T09:47:56Z
drill a4597cd live 8363635
2e9514d DRILL wger: axes by username, 15 min, database handler (R-752 proof on 9202)
2026-10-01T10:31:39Z drill: 6b0520e DRILL wger: cool-off 5 min (R-752 proof)
@@ -23,3 +23,18 @@ python3 manage.py runserver 0.0.0.0:8000
09:50:26 household — admin RIGHT password: locked | second member RIGHT password: ok
09:58:26 +8.0 min admin RIGHT password DURING the lock (does it restart the 15 min?): locked
09:58:26 second member, same moment: ok
10:06:26 +16.0 min admin RIGHT password: locked
10:14:27 +24.0 min admin RIGHT password: locked
10:22:27 +32.0 min admin RIGHT password: locked
10:30:27 +40.0 min admin RIGHT password: locked
10:30:27 a wrong one after: locked | right again: locked
10:30:31 level=WARNING ts=2026-10-01 12:30:27,488 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=223 message=AXES: Repeated login failure by {username: "********************", i
level=WARNING ts=2026-10-01 12:30:27,498 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=253 message=AXES: Locking out {username: "********************", ip_address: "**
level=INFO ts=2026-10-01 12:30:27,662 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=130 message=AXES: User login failed, running database handler for failure.
level=INFO ts=2026-10-01 12:30:27,662 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=425 message=AXES: Cleaned up 0 expired access attempts from database that were olde
level=WARNING ts=2026-10-01 12:30:27,665 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=223 message=AXES: Repeated login failure by {username: "********************", i
level=WARNING ts=2026-10-01 12:30:27,673 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=253 message=AXES: Locking out {username: "********************", ip_address: "**
12:30:41 [X] stop -> 200 {'ok': True, 'message': 'Stack wger stop completed'}
12:31:13 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'wger', 'volumes_removed': ['wger_wger_data', 'wger_wger_media'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note
12:31:21 [X] after remove: deployed=False leftovers='/opt/docker/stacks/wger'
10:31:21 removed; deployed = False
@@ -0,0 +1,37 @@
10:31:42 ##### wger fix — controller gitea.dooplex.hu/admin/felhom-controller:0.285.0
10:31:53 the box's catalog copy: - AXES_LOCKOUT_PARAMETERS=username
12:31:53 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['ADMIN_PASSWORD']
12:31:53 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
12:33:28 [1] deployed, controller state=running, pinned={'wger': 'wger/server:2.7'}
10:33:28 deploy: True
10:33:41 axes settings inside the app: ['username'] 10 0:05:00 True axes.handlers.database.AxesDatabaseHandler django.core.cache.backends.locmem.LocMemCache
python3 manage.py runserver 0.0.0.0:8000
/usr/bin/python3 manage.py runserver 0.0.0.0:8000
10:33:46 second member: made second
10:33:46 positive controls — admin: ok | second: ok
10:33:47 stranger wrong try 1 (admin): wrong
10:33:47 stranger wrong try 2 (admin): wrong
10:33:47 stranger wrong try 3 (admin): wrong
10:33:48 stranger wrong try 4 (admin): wrong
10:33:48 stranger wrong try 5 (admin): wrong
10:33:48 stranger wrong try 6 (admin): wrong
10:33:49 stranger wrong try 7 (admin): wrong
10:33:49 stranger wrong try 8 (admin): wrong
10:33:49 stranger wrong try 9 (admin): wrong
10:33:50 stranger wrong try 10 (admin): locked
10:33:50 stranger wrong try 11 (admin): locked
10:33:50 household — admin RIGHT password: locked | second member RIGHT password: ok
10:35:50 +2.0 min admin RIGHT password DURING the lock (does it restart the cool-off?): locked
10:35:51 second member, same moment: ok
10:41:21 +7.5 min admin RIGHT password: ok
10:41:22 a wrong one after: wrong | right again: ok
10:41:25 level=INFO ts=2026-10-01 12:41:21,471 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=425 message=AXES: Cleaned up 1 expired access attempts from database that were olde
level=INFO ts=2026-10-01 12:41:21,819 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=130 message=AXES: User login failed, running database handler for failure.
level=INFO ts=2026-10-01 12:41:21,820 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=425 message=AXES: Cleaned up 0 expired access attempts from database that were olde
level=WARNING ts=2026-10-01 12:41:21,822 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=201 message=AXES: New login failure by {username: "********************", ip_add
level=INFO ts=2026-10-01 12:41:22,053 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=295 message=AXES: Successful login by {username: "********************", ip_address
level=INFO ts=2026-10-01 12:41:22,054 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=425 message=AXES: Cleaned up 0 expired access attempts from database that were olde
12:41:35 [X] stop -> 200 {'ok': True, 'message': 'Stack wger stop completed'}
12:42:07 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'wger', 'volumes_removed': ['wger_wger_data', 'wger_wger_media'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note
12:42:15 [X] after remove: deployed=False leftovers='/opt/docker/stacks/wger'
10:42:15 removed; deployed = False
@@ -0,0 +1,14 @@
12:42:24 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['ADMIN_PASSWORD']
12:42:24 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
12:44:00 [1] deployed, controller state=running, pinned={'wger': 'wger/server:2.7'}
10:44:00 deploy: True
10:44:23 running env: AXES_COOLOFF_TIME=5 | container created 2026-10-01T10:42:53.627288276Z
10:44:24 drill: ea94447 DRILL wger: cool-off 6 (installed-app probe)
10:44:31 stack-dir compose after 7 s of sync rounds: - AXES_COOLOFF_TIME=6
10:44:34 running env (no action yet): AXES_COOLOFF_TIME=5 | container created 2026-10-01T10:42:53.627288276Z
10:44:45 press Restart: ('200', {'ok': True, 'data': {'state': 'starting'}, 'message': 'Stack wger restart requested — state now: starting'})
10:45:18 running env after Restart: AXES_COOLOFF_TIME=6 | container created 2026-10-01T10:44:34.755369792Z
12:45:29 [X] stop -> 200 {'ok': True, 'message': 'Stack wger stop completed'}
12:46:01 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'wger', 'volumes_removed': ['wger_wger_data', 'wger_wger_media'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note
12:46:09 [X] after remove: deployed=False leftovers='/opt/docker/stacks/wger'
10:46:09 removed; deployed = False
@@ -0,0 +1,15 @@
# Strangers and lockouts (R-752), the one address behind the tunnel (R-753), the registry (R-750) — 2026-10-01 afternoon
Report and Part table: `felhom.eu/REPORT-lockouts-2026-10-01.md`.
| folder | what |
|---|---|
| `A/A1-client-address.txt` | demo-hp through its real tunnel (traefik + BookStack logs), 9202 echo container (tunnel hop simulated, LAN with forged headers), the answer table, why no box-wide fix |
| `B/B1` | BookStack, Grafana (+ trickle), calibre-web (+ OPDS, + the daily limit, + restart) with the LIVE templates |
| `B/B2a`, `B/B2` | wger control (B2a: a wrong login URL — 404 — kept as a record; B2: the real control) |
| `B/B3` | the drill commits and 9202 pointed at the drill |
| `B/B4`, `B/B5` | wger fix at 15 min (retries kept it closed 40+ min) and at 5 min (in at 7.5 min) |
| `B/B6` | what an installed app gets: the stack file at sync, the running env only at Restart |
| `C/C1-registry-read.txt` | Gitea read only: no cleanup rule, the database's first versions, the manual prune (HM-024), nothing scheduled |
| `T/` | 9202 back on live, the drill reset |
| `tools/` | `lk.py` (login probes), `b_three.py`, `b_wger.py`, `b_installed.py`, `echo.sh`, `walk.py`, `repoint.py` |
@@ -0,0 +1,16 @@
git:
branch: main
repo_url: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git
sync_interval: 15m
token: <redacted>
username: ""
hub:
0
felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.285.0
filebrowser gtstef/filebrowser:1.3.3-stable
paperless-postgres postgres:18-alpine
paperless-redis redis:7-alpine
paperless-webserver ghcr.io/paperless-ngx/paperless-ngx:2.20.15
traefik traefik:v3.6.7
0
@@ -0,0 +1,6 @@
# drill reset 2026-10-01T10:47:28Z
before: ea94447 live: 82fff32
ea94447 DRILL wger: cool-off 6 (installed-app probe)
6b0520e DRILL wger: cool-off 5 min (R-752 proof)
2e9514d DRILL wger: axes by username, 15 min, database handler (R-752 proof on 9202)
after: 82fff32
@@ -0,0 +1,33 @@
#!/usr/bin/env python3
"""What an INSTALLED app gets when the catalog changes only a SETTING (R-752): wger installed from the drill
(AXES_COOLOFF_TIME=5); then a drill-only commit sets 6; sync + rescan; read the stack-dir compose and the RUNNING
container's env; press the app page's Restart (POST /api/stacks/wger/restart); read the env again. Removed after."""
import subprocess, time
import walk as w
import lk
D = "/mnt/5_hdd/felhom.eu/drill/app-catalog-drill"
ENV = "docker inspect wger --format '{{range .Config.Env}}{{println .}}{{end}}' | grep AXES_COOLOFF; docker inspect wger --format 'container created {{.Created}}'"
FILE = "grep -h AXES_COOLOFF /opt/docker/stacks/wger/docker-compose.yml"
w.login()
lk.p("deploy:", w.deploy("wger", "fitness"))
time.sleep(20)
lk.p("running env:", w.guest(ENV).strip().replace("\n", " | "))
r = subprocess.run(["bash", "-c", f"cd {D} && sed -i 's/^ - AXES_COOLOFF_TIME=5$/ - AXES_COOLOFF_TIME=6/' templates/wger/docker-compose.yml && "
"git commit -qam 'DRILL wger: cool-off 6 (installed-app probe)' && git push -q origin main 2>&1 | grep -v ^remote; git log --oneline -1"],
capture_output=True, text=True)
lk.p("drill:", r.stdout.strip(), r.stderr.strip()[-200:])
t0 = time.time()
for _ in range(20):
w.sync_rescan()
f = w.guest(FILE).strip()
if "=6" in f:
break
time.sleep(15)
lk.p(f"stack-dir compose after {(time.time() - t0):.0f} s of sync rounds:", f)
lk.p("running env (no action yet):", w.guest(ENV).strip().replace("\n", " | "))
lk.p("press Restart:", w.ctl("POST", "/api/stacks/wger/restart"))
time.sleep(30)
lk.p("running env after Restart:", w.guest(ENV).strip().replace("\n", " | "))
w.remove("wger")
lk.p("removed; deployed =", w.stack("wger").get("deployed"))
@@ -18,7 +18,7 @@ if MODE == "fix":
for _ in range(12):
w.sync_rescan()
cat = w.guest("grep -rh AXES_LOCKOUT_PARAMETERS /var/lib/docker/volumes/felhom-controller-data/_data/ --include=docker-compose.yml 2>/dev/null | sort -u").strip()
if "username" in cat:
if "COOLOFF_TIME=5" in w.guest("grep -rh AXES_COOLOFF_TIME /var/lib/docker/volumes/felhom-controller-data/_data/ --include=docker-compose.yml 2>/dev/null"):
break
time.sleep(15)
lk.p("the box's catalog copy:", cat)
@@ -55,9 +55,10 @@ if MODE == "fix":
r = lk.wger(who, p or pw)
lk.p(f"+{(time.time() - t0) / 60:.1f} min {label}:", r)
return r
at(8, "admin RIGHT password DURING the lock (does it restart the 15 min?)")
sched = [float(x) for x in (sys.argv[2].split(",") if len(sys.argv) > 2 else ["8", "16", "24", "32", "40"])]
at(sched[0], "admin RIGHT password DURING the lock (does it restart the cool-off?)")
lk.p("second member, same moment:", lk.wger("second", pw2))
for m in (16, 24, 32, 40):
for m in sched[1:]:
if at(m, "admin RIGHT password") == "ok":
break
lk.p("a wrong one after:", lk.wger("admin", "wrong-after"), "| right again:", lk.wger("admin", pw))
+2 -2
View File
@@ -861,9 +861,9 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-747** | **[P3-LOW] A stranger can lock the household out of mealie with five wrong logins.** MEASURED 2026-09-30 on 9202 (the R-741 proof): after the install hold opened, a stranger's default-login tries were refused (401) and after five of them mealie answered 423 (locked) to every login — the generated, correct password included. mealie's own brute-force guard, on an app published on the internet; the setup gate and the install hold do not cover an app after its first setup. Not measured: how long the lock lasts. **Needs:** measure the lock's length; decide whether the page tells the household what to do. `audits/night-rulings-2026-09-30/C/C3-mealie-poll.txt` **-- 2026-10-01:** Measured at v3.28.0 (source + 9202): 5 wrong logins lock the ACCOUNT (not the IP) for `SECURITY_USER_LOCKOUT_TIME` hours (default 24); the lock is lifted by an hourly job; an admin can unlock others via `POST /api/admin/users/unlock`, but the household's only admin is the locked account. Both login names are public (`admin`, `changeme@example.com`). **Fixed** (`09` §3 decision 57, decided by CC unattended — operator may reverse): `SECURITY_USER_LOCKOUT_TIME=1`, catalog `a4597cd`; on 9202 the right password answered 423 for 120 min, then 200; a wrong one still 401 (`audits/rulings-2026-10-01/C/`). **Left:** a stranger can renew the lock every hour (per account, public name — option (d) in decision 57 would end that); the page does not tell the household why it is locked; an INSTALLED mealie takes the new setting only when its compose is rendered again (not measured which act does that). | **NARROWED — the lock is 1–2 h; renewal and the page remain; owner: CC** |
| **R-748** | **[P3-LOW] The register-shape gate skipped every row whose id has a letter suffix — so R-88a, R-88b and R-209a were never shape-checked, and its count read 3 short.** FOUND 2026-09-30 (late) while counting the register: `register_shape_gate.py` matched `R-\d+` only; the brief's „the reviewer's regex undercounted by 3” is the same three rows. Fixed the same session: `R-\d+[a-z]?`; decoy `suffix-row-eaten-state` (a suffixed row with its state cell eaten) seen passing with the old pattern and convicted with the new. The register is **382** rows by either count now. | **CLOSED 2026-09-30 — `scripts/register_shape_gate.py`** |
| **R-749** | **[P3-LOW] `retest-floating.py` could never start on a fresh bench: it checked for `/opt/upg/upgrade-test.py` on the bench BEFORE the step that copies it there.** FOUND 2026-10-01 at the first full monthly run (decision 55): bench 9401 freshly created by the runbook, the run answered „CANNOT START — missing: the bench LXC 9401 on demo-hp with /opt/upg” in one minute. The runbook says the command syncs the bench itself — it does, but only after the check. On 2026-09-30 the bench had been synced by hand earlier, so nobody saw it. **Needs:** the check asks for what the bench must bring (docker, python3), the sync then provides `/opt/upg`. `audits/rulings-2026-10-01/A/` **-- 2026-10-01:** Fixed the same session (catalog `9e53205`): the check asks for docker + python3; after the sync `/opt/upg/upgrade-test.py` is required. The re-run started at once and finished both apps. | **CLOSED 2026-10-01 — catalog 9e53205** |
| **R-750** | **[P3-LOW] The registry no longer holds controller releases older than 0.213.0 (2026-08-12) — something removed them, and nothing records what.** MEASURED 2026-10-01 (anonymous registry API, `audits/rulings-2026-10-01/B/`): `felhom-controller` has 91 tags, the oldest release 0.213.0; `0.201.0` answers 404; Gitea's package list starts 2026-08-12. No runbook, row or memory names a clean-up. Today nothing needs those versions: a box runs a newer one, a whole-guest restore brings the guest's own Docker store back (mp0 `backup=1`), and decision 56 deletes only on the box. **But** a box or a backup that names a removed version cannot pull it again (R-698's shape, for the controller). **Needs:** find what removed them (a Gitea clean-up rule?), and record the rule — or say it was a one-time act. Read-only on DooPlex. | **OPEN — rank P3-LOW; owner: operator (Gitea settings), CC measures** |
| **R-750** | **[P3-LOW] The registry no longer holds controller releases older than 0.213.0 (2026-08-12) — something removed them, and nothing records what.** MEASURED 2026-10-01 (anonymous registry API, `audits/rulings-2026-10-01/B/`): `felhom-controller` has 91 tags, the oldest release 0.213.0; `0.201.0` answers 404; Gitea's package list starts 2026-08-12. No runbook, row or memory names a clean-up. Today nothing needs those versions: a box runs a newer one, a whole-guest restore brings the guest's own Docker store back (mp0 `backup=1`), and decision 56 deletes only on the box. **But** a box or a backup that names a removed version cannot pull it again (R-698's shape, for the controller). **Needs:** find what removed them (a Gitea clean-up rule?), and record the rule — or say it was a one-time act. Read-only on DooPlex. **-- 2026-10-01 (afternoon):** **Answered (read only, `audits/lockouts-2026-10-01/C/C1-registry-read.txt`).** Not a Gitea rule: `package_cleanup_rule` is EMPTY (READ ONLY query on the shared CNPG); `[cron.cleanup_packages]` only runs rules and Gitea's own expired-data clean-up. The cause is a MANUAL run of DooPlex's `~/git/misc-scripts/gitea-image-prune.sh --all --keep 7 --apply --reclaim` on the night of 2026-08-22/23, recorded in homelab-manifests HM-024 (the Gitea volume was full: `felhom-golden` 22 → 3 versions, /data 14.8 → 4.4 G); `--keep 7` per container package explains controller 0.213.0 (2026-08-12) as the oldest left. Nothing schedules it (crontabs, timers, cluster CronJobs read). If run again as its usage text says, it keeps 7 controller releases (about a day) and `--type generic --keep 3` would cut the agent to 3; it protects no vouched version. **Needs:** the operator's word on a written rule (STATUS). | **ANSWERED — WAITING-ON-OPERATOR: keep the manual prune or write a rule; owner: operator** |
| **R-751** | **[P2] The image clean-up after an app update could crash the whole controller: it re-read the app after a rescan and dereferenced a nil stack when the app was gone.** FOUND 2026-10-01 by the full test suite (controller v0.284.2): `RetainImagesAfterUpdate` runs in a goroutine; `TestR705_TheManualLegRunsByDay` removed its temp dir under it → `panic: invalid memory address` at `image_retention.go:291`. In a box the same happens when an app is removed (or its compose vanishes) between an update's end and the clean-up — a panic in a goroutine ends the process (the agent's supervisor restarts it). Fixed in v0.285.0 the same session: it returns when the app is gone; `TestRetainImagesAfterUpdate_AppGoneDoesNotPanic` seen panicking on the old code; both retention seams are no-ops in the stacks tests (`TestMain`), so no test leaves the goroutine running. `audits/rulings-2026-10-01/B/B1-red-proofs.txt` **-- 2026-10-01:** Delivered: floor 0.285.0 reached both demo boxes in ~6 s (hub `managed floor SERVED … from declared`). | **CLOSED 2026-10-01 — controller v0.285.0, floor 0.285.0** |
| **R-752** | **[P3-LOW] Four more catalog apps let a stranger lock the household out with wrong passwords for a known login name — like mealie (R-747).** READ 2026-10-01 in each app's source at its pinned tag (not measured live): **calibre-web-automated v4.0.8** — Flask-Limiter on the login keyed on the lowercased USERNAME, 3/minute and 40/day, checked before the password; the default login is `admin` → up to a day; no env switch (a database setting). **wger 2.7** — django-axes keyed on IP, 10 failures, 30 min, each failure restarts it; behind traefik every client has traefik's IP → everyone is locked out (`AXES_*` env vars exist; `AXES_IPWARE_PROXY_COUNT` 0). **Grafana 13.2.3** — per-account, 5 failures in a sliding 5 minutes; a slow trickle keeps it closed (`GF_SECURITY_*`). **BookStack 26.09.1** — key `email|ip`, 5 tries, 60 s, hard-coded; `APP_PROXIES` empty, so the key is the e-mail alone. gokapi (3 s delay, no lock) and claper (per-IP 10/min, no account lock) cannot. **Needs:** per app, the smallest fix that keeps a guessing guard (calibre-web-automated and wger first — longest and broadest), each proven on 9202 as R-747's was. | **OPEN — rank P3-LOW; owner: CC** |
| **R-752** | **[P3-LOW] Four more catalog apps let a stranger lock the household out with wrong passwords for a known login name — like mealie (R-747).** READ 2026-10-01 in each app's source at its pinned tag (not measured live): **calibre-web-automated v4.0.8** — Flask-Limiter on the login keyed on the lowercased USERNAME, 3/minute and 40/day, checked before the password; the default login is `admin` → up to a day; no env switch (a database setting). **wger 2.7** — django-axes keyed on IP, 10 failures, 30 min, each failure restarts it; behind traefik every client has traefik's IP → everyone is locked out (`AXES_*` env vars exist; `AXES_IPWARE_PROXY_COUNT` 0). **Grafana 13.2.3** — per-account, 5 failures in a sliding 5 minutes; a slow trickle keeps it closed (`GF_SECURITY_*`). **BookStack 26.09.1** — key `email|ip`, 5 tries, 60 s, hard-coded; `APP_PROXIES` empty, so the key is the e-mail alone. gokapi (3 s delay, no lock) and claper (per-IP 10/min, no account lock) cannot. **Needs:** per app, the smallest fix that keeps a guessing guard (calibre-web-automated and wger first — longest and broadest), each proven on 9202 as R-747's was. **-- 2026-10-01 (afternoon):** **Measured on 9202, each through traefik as a stranger with the public name** (`audits/lockouts-2026-10-01/B/`): **wger** — control: 10 wrong on `admin` locked the second member too; FIXED (decision 58, catalog `82fff32`): username, 5 min, database handler — the second member unaffected, admin in again at 7.5 min (each try during a lock restarts it — measured: 8-minute retries kept a 15-minute lock closed 40+ min). **BookStack** — 1.0 min, kept (decision 59). **Grafana** — 5.0 min, kept (decision 60); a trickle did not hold the household out once the burst aged. **calibre-web-automated** — the form locks 3/min (1.2 min measured) and **40/day per name: after 40 wrong tries in 14 min the right password was refused 2 min later still; only an app restart cleared it** (in-memory store); OPDS has its own 3/min per name (`cps/main.py:75`), no daily limit. No knob for the daily length; both fixes have a household cost — operator decision in STATUS. Installed apps: a settings-only change reaches the stack file at the next sync (images equal, ≤15 min) and the running app at the next `compose up -d` — Restart/Start (measured: the env changed only at Restart), an Update, or a backup's restart (`backup.go:972`, read). | **NARROWED — wger fixed, BookStack/Grafana kept; calibre-web waits for the operator; owner: operator (calibre-web), CC** |
| **R-753** | **[P3-LOW] Behind the tunnel every visitor reaches an app with the SAME address — the tunnel container's — so every per-address guard is an "everyone" guard and every app's log is blind.** MEASURED 2026-10-01 (`audits/lockouts-2026-10-01/A/A1-client-address.txt`): on demo-hp through its real tunnel, a request from DooPlex's public address reached traefik as `172.18.0.5` (cloudflared, in the guest on `traefik-public`) and BookStack as `172.18.0.3` (traefik); on 9202 an echo container showed `X-Forwarded-For`/`X-Real-Ip` = the sending container for the tunnel's hop (traefik DROPS the incoming chain — good: a client cannot forge it) and the real address from the LAN; `CF-Connecting-IP` passes untouched and is FORGEABLE from the LAN. **No box-wide fix taken:** trusting cloudflared in traefik passes Cloudflare's appended chain, whose LEFTMOST entry the client writes — every app reading the leftmost address would believe it; cloudflared's address is docker-assigned; a single-address rewrite needs a traefik plugin (a new dependency). Per-app fixes trust no header (R-752). **Needs (operator):** whether to build a safe version (cloudflared on a fixed-address network + traefik trusting only it + per-app proxy counts), or keep "one address" and fix per app. Only ONE outside address was available (DooPlex has no IPv6); a second was not measured. | **OPEN — rank P3-LOW; owner: operator (direction), CC measures** |
| **R-754** | **[P3-LOW] `01-topology-and-trust.md` §7 says cloudflared runs on the Proxmox HOST as an agent-managed service; on every box it runs INSIDE the guest as a container the controller renders.** READ 2026-10-01: `felhom-controller` `internal/infra/templates/cloudflared-compose.yml.tmpl` (`container_name: cloudflared`, network `traefik-public`); demo-hp's guest 9201 runs `cloudflared` (ingress `*.enkisfelhom.hu -> https://traefik`); R-505 saw the same in VM 331. A design decision that the build does not follow — the document or the build is wrong, and only the operator decides which (R-370: a design decision is not a defect). | **OPEN — rank P3-LOW; owner: operator (which is right)** |
| **R-755** | **[P3-LOW] wger runs Django's DEVELOPMENT server in production: `manage.py runserver`, because the template does not set `WGER_USE_GUNICORN=True`.** MEASURED 2026-10-01 on 9202 (`ps` in the wger container: `python3 manage.py runserver 0.0.0.0:8000`); wger 2.7's `extras/docker/production/entrypoint.sh:81-87` runs gunicorn only with that switch. Django's own documentation says runserver is not for production (one process, not hardened). Not changed this session (a different change from R-752's; needs its own bench + box proof, memory watch included). | **OPEN — rank P3-LOW; owner: CC (catalog)** |