Register: R-753 (one address behind the tunnel), R-754 (01 §7 says cloudflared on the host), R-755 (wger on runserver); evidence so far
gates / gates (push) Successful in 27s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-01 11:58:57 +02:00
parent 939553c82f
commit 92a60c62bd
15 changed files with 1276 additions and 0 deletions
@@ -0,0 +1,44 @@
# Part A — what client address an app sees (2026-10-01 ~09:00 UTC)
## demo-hp 9201, through its REAL tunnel (read only: two GETs of a 404 path, then the logs)
From DooPlex, public IPv4 37.191.56.193 (no IPv6 here, so only ONE outside address was available):
curl https://wiki.enkisfelhom.hu/probe-1790844959-a (once plain, once with X-Forwarded-For: 6.6.6.6)
cloudflared remote config (its own log): ingress *.enkisfelhom.hu -> https://traefik (noTLSVerify)
traefik access log (ClientHost = the peer it believes):
172.18.0.5 - - [01/Oct/2026:08:56:09 +0000] "GET /probe-1790844959-a HTTP/1.1" 404 ... "bookstack@docker" "http://172.18.0.7:80"
172.18.0.5 - - [01/Oct/2026:08:56:09 +0000] "GET /probe-1790844959-a HTTP/1.1" 404 ... "bookstack@docker" "http://172.18.0.7:80"
(172.18.0.5 = the cloudflared CONTAINER, 172.18.0.3 = traefik, docker network traefik-public 172.18.0.0/16)
BookStack's own nginx log ($remote_addr):
172.18.0.3 - - [01/Oct/2026:10:56:09 +0200] "GET /probe-1790844959-a HTTP/1.1" 404 ... "curl/8.14.1"
172.18.0.3 - - [01/Oct/2026:10:56:10 +0200] "GET /probe-1790844959-a HTTP/1.1" 404 ... "curl/8.14.1"
## 9202 (scratch, no tunnel): an echo container (traefik/whoami v1.11) behind traefik with the catalog's usual labels (tools/echo.sh)
traefik 172.18.0.6, echo 172.18.0.5
T — the tunnel's hop SIMULATED: a container on traefik-public (where cloudflared sits) sends what cloudflared sends
(X-Forwarded-For: 6.6.6.6, 203.0.113.9 · CF-Connecting-IP: 203.0.113.9):
RemoteAddr: 172.18.0.6:42874 (traefik)
X-Forwarded-For: 172.18.0.7 (the sending container — the forwarded chain was DROPPED by traefik)
X-Real-Ip: 172.18.0.7
Cf-Connecting-Ip: 203.0.113.9 (passed through untouched — traefik does not know this header)
L — from the LAN (DooPlex 192.168.0.180) straight to 9202:443, forging X-Forwarded-For 6.6.6.6, CF-Connecting-IP 7.7.7.7,
X-Real-IP 8.8.4.4:
RemoteAddr: 172.18.0.6:42874
X-Forwarded-For: 192.168.0.180 (forgery dropped; the real LAN address — docker's DNAT keeps the source)
X-Real-Ip: 192.168.0.180
Cf-Connecting-Ip: 7.7.7.7 (THE FORGERY ARRIVES — any device that reaches :443 directly can set it)
Echo container and both test images removed afterwards.
## The answer
| path | the app's TCP peer | X-Forwarded-For / X-Real-Ip | the real client is in | forgeable by the client? |
|--------|--------------------|-----------------------------|------------------------------|--------------------------|
| tunnel | traefik | cloudflared's container IP — THE SAME FOR EVERY VISITOR | CF-Connecting-IP only (set by Cloudflare's edge) | XFF: no (traefik drops it). CF-Connecting-IP: not through the tunnel (the edge overwrites it), YES from the LAN |
| LAN | traefik | the real LAN address | X-Forwarded-For / X-Real-Ip | no (traefik drops a forged chain) |
## Why no box-wide fix (Part A4 stops here)
traefik `forwardedHeaders.trustedIPs: [cloudflared]` would keep the tunnel's chain. Cloudflare APPENDS to a client's own
X-Forwarded-For (Cloudflare docs, not measured here), so an app would then receive "<anything the client wrote>, <real
client>, <cloudflared>". Every app that reads the LEFTMOST address (a common default) would then believe a client-written
value — a stranger could dodge or aim any per-address guard. Today that header is useless but honest. Trusting
CF-Connecting-IP instead is forgeable from the LAN (L above). cloudflared's address is also not fixed (172.18.0.5 here,
docker-assigned). A safe version needs a fixed-address network for cloudflared plus a rewrite to ONE address — not
available in traefik without a plugin (a new external dependency). So: per app, trusting no header.
@@ -0,0 +1,124 @@
09:03:07 ##### bookstack — controller gitea.dooplex.hu/admin/felhom-controller:0.285.0
11:03:07 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['ADMIN_PASSWORD']
11:03:07 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
11:04:07 [1] deployed, controller state=running, pinned={'bookstack': 'lscr.io/linuxserver/bookstack:26.09.1', 'bookstack-db': 'mariadb:12.3'}
09:04:07 deploy: True
09:04:16 generated password length: 30
09:04:16 positive control — right password: ok
09:04:17 stranger wrong try 1: wrong
09:04:17 stranger wrong try 2: wrong
09:04:18 stranger wrong try 3: wrong
09:04:18 stranger wrong try 4: wrong
09:04:19 stranger wrong try 5: wrong
09:04:19 stranger wrong try 6: locked
09:04:19 household, RIGHT password, right after: locked
09:04:34 +0.3 min right password: locked
09:04:50 +0.5 min right password: locked
09:05:05 +0.8 min right password: locked
09:05:20 +1.0 min right password: ok
09:05:20 RESULT bookstack: the right password worked again after 1.0 min
09:05:21 a wrong one after: wrong | right again: ok
09:05:24
11:05:29 [X] stop -> 200 {'ok': True, 'message': 'Stack bookstack stop completed'}
11:06:01 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'bookstack', 'volumes_removed': ['bookstack_bookstack_config', 'bookstack_bookstack_db_data'], 'hdd_paths_removed': [], 'hdd_pa
11:06:09 [X] after remove: deployed=False leftovers='/opt/docker/stacks/bookstack'
09:06:09 removed; deployed = False
09:06:12 ##### grafana — controller gitea.dooplex.hu/admin/felhom-controller:0.285.0
11:06:12 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['GF_SECURITY_ADMIN_PASSWORD']
11:06:12 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
11:06:53 [1] deployed, controller state=running, pinned={'grafana': 'grafana/grafana:13.2.3'}
09:06:53 deploy: True
09:07:01 generated password length: 30
09:07:01 positive control — right password: ok
09:07:01 stranger wrong try 1: wrong
09:07:01 stranger wrong try 2: wrong
09:07:01 stranger wrong try 3: wrong
09:07:01 stranger wrong try 4: wrong
09:07:01 stranger wrong try 5: wrong
09:07:01 stranger wrong try 6: wrong
09:07:01 household, RIGHT password, right after: wrong
09:07:17 +0.3 min right password: wrong
09:07:32 +0.5 min right password: wrong
09:07:47 +0.8 min right password: wrong
09:08:02 +1.0 min right password: wrong
09:08:17 +1.3 min right password: wrong
09:08:32 +1.5 min right password: wrong
09:08:47 +1.8 min right password: wrong
09:09:02 +2.0 min right password: wrong
09:09:17 +2.3 min right password: wrong
09:09:32 +2.5 min right password: wrong
09:09:47 +2.8 min right password: wrong
09:10:02 +3.0 min right password: wrong
09:10:17 +3.3 min right password: wrong
09:10:32 +3.5 min right password: wrong
09:10:47 +3.8 min right password: wrong
09:11:02 +4.0 min right password: wrong
09:11:17 +4.3 min right password: wrong
09:11:32 +4.5 min right password: wrong
09:11:47 +4.8 min right password: wrong
09:12:02 +5.0 min right password: ok
09:12:02 RESULT grafana: the right password worked again after 5.0 min
09:12:02 a wrong one after: wrong | right again: ok
09:12:02 ## grafana: the TRICKLE — 5 wrong, then one wrong try per minute for 8 minutes; the right password each minute
09:12:02 burst wrong 1: wrong
09:12:02 burst wrong 2: wrong
09:12:02 burst wrong 3: wrong
09:12:02 burst wrong 4: wrong
09:12:02 burst wrong 5: wrong
09:13:02 minute 1: stranger wrong -> wrong, household right -> wrong
09:14:02 minute 2: stranger wrong -> wrong, household right -> wrong
09:15:02 minute 3: stranger wrong -> wrong, household right -> wrong
09:16:03 minute 4: stranger wrong -> wrong, household right -> wrong
09:17:03 minute 5: stranger wrong -> wrong, household right -> ok
09:18:03 minute 6: stranger wrong -> wrong, household right -> ok
09:19:03 minute 7: stranger wrong -> wrong, household right -> ok
09:20:03 minute 8: stranger wrong -> wrong, household right -> ok
09:20:03 stop the trickle; wait for the right password:
09:20:18 +0.3 min right password: ok
09:20:18 RESULT grafana trickle: open again 0.3 min after the last wrong try
09:20:21 logger=authn.service t=2026-10-01T11:16:02.977306539+02:00 level=info msg="Failed to authenticate request" client=auth.client.form error="[password-auth.failed] too many consecutive incorrect login at
logger=context userId=0 orgId=0 uname= t=2026-10-01T11:16:02.978088015+02:00 level=info msg="Request Completed" method=POST path=/login status=401 remote_addr=192.168.0.180 time_ms=1 duration=1.228039
logger=authn.service t=2026-10-01T11:16:03.000927142+02:00 level=info msg="Failed to authenticate request" client=auth.client.form error="[password-auth.failed] too many consecutive incorrect login at
logger=context userId=0 orgId=0 uname= t=2026-10-01T11:16:03.001553164+02:00 level=info msg="Request Completed" method=POST path=/login status=401 remote_addr=192.168.0.180 time_ms=1 duration=1.066954
11:20:22 [X] stop -> 200 {'ok': True, 'message': 'Stack grafana stop completed'}
11:20:54 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'grafana', 'volumes_removed': ['grafana_grafana_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note': 'Az alk
11:21:02 [X] after remove: deployed=False leftovers='/opt/docker/stacks/grafana'
09:21:02 removed; deployed = False
09:21:05 ##### calibre-web — controller gitea.dooplex.hu/admin/felhom-controller:0.285.0
11:21:08 [1] made the drive paths this app requires: ['/mnt/felhom-drives/scratch_hdd/userdata/calibre-web']
11:21:08 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['HDD_PATH', 'ADMIN_PASSWORD']
11:21:08 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
11:21:48 [1] deployed, controller state=running, pinned={'calibre-web': 'crocodilestick/calibre-web-automated:v4.0.8'}
09:21:48 deploy: True
09:21:57 generated password length: 30
09:21:57 positive control — right password: ok
09:21:57 stranger wrong try 1: wrong
09:21:57 stranger wrong try 2: wrong
09:21:58 stranger wrong try 3: wrong
09:21:58 stranger wrong try 4: locked
09:21:58 household, RIGHT password, right after: locked
09:23:08 +1.2 min right password: ok
09:23:08 RESULT calibre-web: the right password worked again after 1.2 min
09:23:08 a wrong one after: wrong | right again: ok
09:23:08 ## calibre-web: OPDS basic auth (the book-feed door) — 60 wrong tries as a stranger
09:23:13 OPDS wrong tries: {'locked': 57, 'wrong': 3}
09:23:13 OPDS right password after: locked | the login form, right password: ok
09:23:13 ## calibre-web: the DAILY limit — 40 wrong tries at the 3/minute pace (~14 min), then the right password
09:26:18 wrong try 10: wrong
09:29:23 wrong try 20: wrong
09:32:28 wrong try 30: wrong
09:36:34 wrong try 40: wrong
09:37:36 household, RIGHT password after 40: locked
09:39:46 household, RIGHT password 2 min later (past the minute window): locked
09:39:54 restart the app (the limiter's store is in memory): restarted
09:39:59 household, RIGHT password after the restart: ok
09:40:02 [2026-10-01 11:39:46,068] INFO {flask-limiter:1102} ratelimit 40 per 1 day (admin) exceeded at endpoint: web.login_post
[cwa-init] Checking for leftover lock files from previous instance...
[cwa-init] No leftover lock files to remove. Ending service...
[2026-10-01 11:39:56,532] WARN {py.warnings:110} /lsiopy/lib/python3.13/site-packages/flask_limiter/extension.py:324: UserWarning: Using the in-memory storage for tracking rate limits as no storage w
11:40:13 [X] stop -> 200 {'ok': True, 'message': 'Stack calibre-web stop completed'}
11:40:18 [X] remove (with drive data) -> 409 {'ok': False, 'error': 'A(z) /mnt/felhom-drives/scratch_hdd/userdata/calibre-web tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghaj
11:40:18 [X] refused because the drive path cannot be resolved (R-442, fail-closed and right) — removing the app and KEEPING the drive data instead
11:40:45 [X] remove (keeping drive data) -> 200 {'ok': True, 'data': {'removed': 'calibre-web', 'volumes_removed': ['calibre-web_calibre_web_config'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'back
11:40:53 [X] after remove: deployed=False leftovers='/opt/docker/stacks/calibre-web'
09:40:53 removed; deployed = False
@@ -0,0 +1,31 @@
09:44:40 ##### wger control — controller gitea.dooplex.hu/admin/felhom-controller:0.285.0
11:44:40 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['ADMIN_PASSWORD']
11:44:40 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
11:46:11 [1] deployed, controller state=running, pinned={'wger': 'wger/server:2.7'}
09:46:11 deploy: True
09:46:22 gunicorn workers: 0
09:46:27 axes settings inside the app:
09:46:31 second member: made second
09:46:32 positive controls — admin: ok | second: ok
09:46:32 stranger wrong try 1 (admin): wrong
09:46:32 stranger wrong try 2 (admin): wrong
09:46:33 stranger wrong try 3 (admin): wrong
09:46:33 stranger wrong try 4 (admin): wrong
09:46:34 stranger wrong try 5 (admin): wrong
09:46:34 stranger wrong try 6 (admin): wrong
09:46:34 stranger wrong try 7 (admin): wrong
09:46:35 stranger wrong try 8 (admin): wrong
09:46:35 stranger wrong try 9 (admin): wrong
09:46:35 stranger wrong try 10 (admin): locked
09:46:35 stranger wrong try 11 (admin): locked
09:46:36 household — admin RIGHT password: locked | second member RIGHT password: locked
09:46:39 level=WARNING ts=2026-10-01 11:46:35,782 module=cache path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/cache.py line=130 message=AXES: Repeated login failure by {username: "********************", ip_addr
level=WARNING ts=2026-10-01 11:46:35,783 module=cache path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/cache.py line=146 message=AXES: Locking out {username: "********************", ip_address: "********
level=WARNING ts=2026-10-01 11:46:35,946 module=cache path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/cache.py line=130 message=AXES: Repeated login failure by {username: "********************", ip_addr
level=WARNING ts=2026-10-01 11:46:35,947 module=cache path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/cache.py line=146 message=AXES: Locking out {username: "********************", ip_address: "********
level=WARNING ts=2026-10-01 11:46:36,120 module=cache path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/cache.py line=130 message=AXES: Repeated login failure by {username: "********************", ip_addr
level=WARNING ts=2026-10-01 11:46:36,120 module=cache path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/cache.py line=146 message=AXES: Locking out {username: "********************", ip_address: "********
11:46:49 [X] stop -> 200 {'ok': True, 'message': 'Stack wger stop completed'}
11:47:21 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'wger', 'volumes_removed': ['wger_wger_data', 'wger_wger_media'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note
11:47:29 [X] after remove: deployed=False leftovers='/opt/docker/stacks/wger'
09:47:29 removed; deployed = False
@@ -0,0 +1,75 @@
09:41:19 ##### wger control — controller gitea.dooplex.hu/admin/felhom-controller:0.285.0
11:41:19 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['ADMIN_PASSWORD']
11:41:20 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
11:42:55 [1] deployed, controller state=running, pinned={'wger': 'wger/server:2.7'}
09:42:55 deploy: True
09:43:08 axes settings inside the app: ['ip_address'] 10 0:30:00 True axes.handlers.cache.AxesCacheHandler
09:43:13 second member: made second
09:43:13 positive controls — admin: other:404 HTTP/2 404
content-language: en
content-type: text/html; charset=utf-8
date: Th | second: other:404 HTTP/2 404
content-language: en
content-type: text/html; charset=utf-8
date: Th
09:43:13 stranger wrong try 1 (admin): other:404 HTTP/2 404
content-language: en
content-type: text/html; charset=utf-8
date: Th
09:43:13 stranger wrong try 2 (admin): other:404 HTTP/2 404
content-language: en
content-type: text/html; charset=utf-8
date: Th
09:43:13 stranger wrong try 3 (admin): other:404 HTTP/2 404
content-language: en
content-type: text/html; charset=utf-8
date: Th
09:43:13 stranger wrong try 4 (admin): other:404 HTTP/2 404
content-language: en
content-type: text/html; charset=utf-8
date: Th
09:43:13 stranger wrong try 5 (admin): other:404 HTTP/2 404
content-language: en
content-type: text/html; charset=utf-8
date: Th
09:43:13 stranger wrong try 6 (admin): other:404 HTTP/2 404
content-language: en
content-type: text/html; charset=utf-8
date: Th
09:43:13 stranger wrong try 7 (admin): other:404 HTTP/2 404
content-language: en
content-type: text/html; charset=utf-8
date: Th
09:43:13 stranger wrong try 8 (admin): other:404 HTTP/2 404
content-language: en
content-type: text/html; charset=utf-8
date: Th
09:43:13 stranger wrong try 9 (admin): other:404 HTTP/2 404
content-language: en
content-type: text/html; charset=utf-8
date: Th
09:43:13 stranger wrong try 10 (admin): other:404 HTTP/2 404
content-language: en
content-type: text/html; charset=utf-8
date: Th
09:43:13 stranger wrong try 11 (admin): other:404 HTTP/2 404
content-language: en
content-type: text/html; charset=utf-8
date: Th
09:43:13 household — admin RIGHT password: other:404 HTTP/2 404
content-language: en
content-type: text/html; charset=utf-8
date: Th | second member RIGHT password: other:404 HTTP/2 404
content-language: en
content-type: text/html; charset=utf-8
date: Th
09:43:16 Apply all migrations: account, actstream, allauth_idp_oidc, auth, authtoken, axes, config, contenttypes, core, easy_thumbnails, exercises, gallery, gym, mailer, manager, measurements, mfa, nutrition, sessions, sites, s
level=INFO ts=2026-10-01 11:42:41,281 module=apps path=/home/wger/.local/lib/python3.12/site-packages/axes/apps.py line=53 message=AXES: BEGIN version 8.3.1, blocking by ip_address
?: (axes.W001) You are using the django-axes cache handler for login attempt tracking. Your cache configuration is however invalid and will not work correctly with django-axes. This can leave security holes in your login
level=INFO ts=2026-10-01 11:42:43,527 module=apps path=/home/wger/.local/lib/python3.12/site-packages/axes/apps.py line=53 message=AXES: BEGIN version 8.3.1, blocking by ip_address
level=INFO ts=2026-10-01 11:42:44,714 module=apps path=/home/wger/.local/lib/python3.12/site-packages/axes/apps.py line=53 message=AXES: BEGIN version 8.3.1, blocking by ip_address
?: (axes.W001) You are using the django-axes cache handler for login attempt tracking. Your cache configuration is however invalid and will not work correctly with django-axes. This can leave security holes in your login
11:43:27 [X] stop -> 200 {'ok': True, 'message': 'Stack wger stop completed'}
11:43:58 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'wger', 'volumes_removed': ['wger_wger_data', 'wger_wger_media'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note
11:44:07 [X] after remove: deployed=False leftovers='/opt/docker/stacks/wger'
09:44:07 removed; deployed = False
@@ -0,0 +1,10 @@
git:
branch: main
repo_url: https://gitea.dooplex.hu/admin/app-catalog-drill.git
sync_interval: 15m
token: <redacted>
username: "admin"
hub:
update:
health_timeout: 90s
@@ -0,0 +1,3 @@
# drill 2026-10-01T09:47:56Z
drill a4597cd live 8363635
2e9514d DRILL wger: axes by username, 15 min, database handler (R-752 proof on 9202)
@@ -0,0 +1,25 @@
09:48:26 ##### wger fix — controller gitea.dooplex.hu/admin/felhom-controller:0.285.0
09:48:33 the box's catalog copy: - AXES_LOCKOUT_PARAMETERS=username
11:48:33 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['ADMIN_PASSWORD']
11:48:33 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
11:50:04 [1] deployed, controller state=running, pinned={'wger': 'wger/server:2.7'}
09:50:04 deploy: True
09:50:17 axes settings inside the app: ['username'] 10 0:15:00 True axes.handlers.database.AxesDatabaseHandler django.core.cache.backends.locmem.LocMemCache
python3 manage.py runserver 0.0.0.0:8000
/usr/bin/python3 manage.py runserver 0.0.0.0:8000
09:50:21 second member: made second
09:50:22 positive controls — admin: ok | second: ok
09:50:22 stranger wrong try 1 (admin): wrong
09:50:22 stranger wrong try 2 (admin): wrong
09:50:23 stranger wrong try 3 (admin): wrong
09:50:23 stranger wrong try 4 (admin): wrong
09:50:23 stranger wrong try 5 (admin): wrong
09:50:24 stranger wrong try 6 (admin): wrong
09:50:24 stranger wrong try 7 (admin): wrong
09:50:24 stranger wrong try 8 (admin): wrong
09:50:25 stranger wrong try 9 (admin): wrong
09:50:25 stranger wrong try 10 (admin): locked
09:50:25 stranger wrong try 11 (admin): locked
09:50:26 household — admin RIGHT password: locked | second member RIGHT password: ok
09:58:26 +8.0 min admin RIGHT password DURING the lock (does it restart the 15 min?): locked
09:58:26 second member, same moment: ok
@@ -0,0 +1,41 @@
# Part C — what removed the old controller versions from the registry (R-750). READ ONLY. 2026-10-01T09:01:40Z
## 1. Gitea's package cleanup rules — NONE
Gitea 1.26.2 (API /version). Database gitea on the shared CNPG (postgresql-2), one READ ONLY transaction:
SELECT ... FROM package_cleanup_rule; -> (0 rows)
app.ini (the pod's, regenerated at every pod start; secrets filtered out of this read):
[packages] ENABLED = true
[cron.cleanup_packages] ENABLED = true, RUN_AT_START = true
— that cron runs the cleanup RULES (none) and Gitea's own expired-data clean-up; it deletes no tagged version by itself.
The admin cron API answered 403 (the available token has no write:admin) — same as HM-024 recorded.
## 2. What the database shows (READ ONLY)
package first remaining version versions
felhom-act-runner 2026-08-02 3
felhom-agent 2026-08-03 20
felhom-controller 2026-08-12 288 (tags + digest manifests)
felhom-golden 2026-08-22 21
felhom-hub 2026-08-08 90
felhom-samba 2026-07-18 6
package_version ids run 17..4201; the felhom-controller package itself is id 2 (it existed long before 08-12).
## 3. The cause — a MANUAL run of a DooPlex script, recorded in homelab-manifests (HM-024, CHANGELOG 2026-08-23)
"gitea-image-prune.sh --all --keep 7 --apply --reclaim reported success overnight and freed nothing visible"
— then the generic fix: felhom-golden 22 -> 3 versions; /data 14.8 G -> 4.4 G; Longhorn actualSize 46.5 -> 7.0 GB.
~/git/misc-scripts/gitea-image-prune.sh history:
7739c83 2026-08-23 13:47:58 +0200 gitea-image-prune.sh: add --type, so generic packages are prunable
5b4d8ec 2026-06-17 09:28:55 +0200 gitea-image-prune.sh: auto-discover credentials from git
761dc38 2026-06-17 09:16:07 +0200 Add gitea-image-prune.sh: inspect/prune Gitea container images + reclaim disk
Its own usage text: "./gitea-image-prune.sh --all --keep 7 --apply --yes --reclaim # containers"
"./gitea-image-prune.sh --type generic --all --keep 3 --apply --yes --reclaim"
--keep 7 per CONTAINER package on 2026-08-22/23 leaves felhom-controller's oldest at 0.213.0 (2026-08-12) — matches.
The record names WHY: "build-felhom-{hub,controller}.sh push :<version> AND :latest on every build ... the Gitea
Longhorn PVC keeps filling."
## 4. Will it remove more? Only when someone runs it again — nothing schedules it
kisfenyo crontab: one unrelated line (jarrs.eu backup sync); root crontab: empty; systemd timers: none matching;
/etc/cron.d, cron.daily: none; cluster CronJobs: authentik-shm-cleaner, crashloop-restarter, renovate,
postgresql-backup, longhorn backup-daily/weekly — none touches packages.
If run again as its usage text says: felhom-controller would keep 7 releases (about one day of releases today), and
"--type generic --all --keep 3" would cut felhom-agent to 3 versions — HM-024 itself refused that sweep ("felhom-agent's
retention was not a decision anyone had made"). Nothing in the script protects the VOUCHED agent or golden.
@@ -0,0 +1,72 @@
#!/usr/bin/env python3
"""R-752 on 9202 with the LIVE catalog templates (no change yet): BookStack, Grafana, calibre-web-automated.
Per app: install through the product, wait for the install hold to open, positive control, a stranger's wrong tries
with the PUBLIC login name, the household's right password right after, then how long until it works again, then
a wrong one still refused. One app at a time; each removed at the end."""
import sys, time
import walk as w
import lk
APPS = sys.argv[1:] or ["bookstack", "grafana", "calibre-web"]
SUB = {"bookstack": "wiki", "grafana": "grafana", "calibre-web": "books"}
KEY = {"bookstack": "ADMIN_PASSWORD", "grafana": "GF_SECURITY_ADMIN_PASSWORD", "calibre-web": "ADMIN_PASSWORD"}
NAME = {"bookstack": "admin@admin.com", "grafana": "admin", "calibre-web": "admin"}
PROBE = {"bookstack": lk.bookstack, "grafana": lk.grafana, "calibre-web": lk.calibre}
w.login()
for app in APPS:
sub, name, probe = SUB[app], NAME[app], PROBE[app]
lk.p(f"##### {app} — controller {w.guest('docker inspect felhom-controller --format {{.Config.Image}}').strip()}")
lk.p("deploy:", w.deploy(app, sub))
for _ in range(120):
time.sleep(5)
logs = w.guest(f"docker logs --since 20m felhom-controller 2>&1 | grep -E '{app}: install hold OPENED|{app}.*deployed successfully'")
if "hold OPENED" in logs or (app == "grafana" and "deployed successfully" in logs):
break
w.wait_app(sub, "/login", want=("200",), tries=60)
pw = (w.GENERATED.get(app) or {}).get(KEY[app], "")
lk.p("generated password length:", len(pw))
lk.p("positive control — right password:", probe(name, pw))
tries = 6 if app != "calibre-web" else 4
for i in range(1, tries + 1):
lk.p(f"stranger wrong try {i}:", probe(name, f"wrong-{i}-{time.time()}"))
lk.p("household, RIGHT password, right after:", probe(name, pw))
# calibre-web counts EVERY try (the household's too) against 3/minute and 40/day — poll it slowly
every, cap = (70, 6 * 60) if app == "calibre-web" else (15, 20 * 60)
waited = lk.wait_until(lambda: probe(name, pw), "ok", every, cap, "right password")
lk.p(f"RESULT {app}: the right password worked again after {waited if waited is None else round(waited, 1)} min")
lk.p("a wrong one after:", probe(name, "wrong-after"), "| right again:", probe(name, pw))
if app == "calibre-web":
lk.p("## calibre-web: OPDS basic auth (the book-feed door) — 60 wrong tries as a stranger")
res = [lk.calibre_opds(name, f"opds-wrong-{i}") for i in range(60)]
lk.p("OPDS wrong tries:", {r: res.count(r) for r in set(res)})
lk.p("OPDS right password after:", lk.calibre_opds(name, pw), "| the login form, right password:", lk.calibre(name, pw))
lk.p("## calibre-web: the DAILY limit — 40 wrong tries at the 3/minute pace (~14 min), then the right password")
n = 0
while n < 40:
for _ in range(3):
if n < 40:
n += 1
r = lk.calibre(name, f"daily-{n}")
if n % 10 == 0 or r not in ("wrong",):
lk.p(f" wrong try {n}: {r}")
time.sleep(61)
lk.p("household, RIGHT password after 40:", lk.calibre(name, pw))
time.sleep(130)
lk.p("household, RIGHT password 2 min later (past the minute window):", lk.calibre(name, pw))
lk.p("restart the app (the limiter's store is in memory):", w.guest("docker restart calibre-web >/dev/null && echo restarted").strip())
w.wait_app(sub, "/login", want=("200",), tries=60)
lk.p("household, RIGHT password after the restart:", lk.calibre(name, pw))
if app == "grafana":
lk.p("## grafana: the TRICKLE — 5 wrong, then one wrong try per minute for 8 minutes; the right password each minute")
for i in range(5):
lk.p(f"burst wrong {i + 1}:", probe(name, f"burst-{i}-{time.time()}"))
for m in range(8):
time.sleep(60)
lk.p(f" minute {m + 1}: stranger wrong -> {probe(name, f'trickle-{m}')}, household right -> {probe(name, pw)}")
lk.p("stop the trickle; wait for the right password:")
waited = lk.wait_until(lambda: probe(name, pw), "ok", 15, 10 * 60, "right password")
lk.p(f"RESULT grafana trickle: open again {waited if waited is None else round(waited, 1)} min after the last wrong try")
lk.p(w.guest(f"docker logs --since 30m {app} 2>&1 | grep -i -E 'lock|block|throttl|limit' | tail -4 | cut -c1-200").strip())
w.remove(app)
lk.p("removed; deployed =", w.stack(app).get("deployed"))
@@ -0,0 +1,66 @@
#!/usr/bin/env python3
"""R-752 wger on 9202. argv[1] = control | fix.
control: the LIVE template (axes keyed on ip_address, 10 failures, 30 min). fix: the drill template
(AXES_LOCKOUT_PARAMETERS=username, AXES_COOLOFF_TIME=15).
Both: install, a SECOND household member made with wger's own Django (manage.py shell — seeding, not a login),
positive controls, 10 wrong REST logins for the public name `admin` as a stranger, then admin's and the second
member's RIGHT passwords. fix only: one right-password try at +8 min (a try DURING the lock — does it extend it?),
then single tries at +16, +24, +32 min until it works; a wrong one after."""
import sys, time
import walk as w
import lk
MODE = sys.argv[1]
SUB = "fitness"
w.login()
lk.p(f"##### wger {MODE} — controller {w.guest('docker inspect felhom-controller --format {{.Config.Image}}').strip()}")
if MODE == "fix":
for _ in range(12):
w.sync_rescan()
cat = w.guest("grep -rh AXES_LOCKOUT_PARAMETERS /var/lib/docker/volumes/felhom-controller-data/_data/ --include=docker-compose.yml 2>/dev/null | sort -u").strip()
if "username" in cat:
break
time.sleep(15)
lk.p("the box's catalog copy:", cat)
lk.p("deploy:", w.deploy("wger", SUB))
for _ in range(120):
time.sleep(5)
if "hold OPENED" in w.guest("docker logs --since 20m felhom-controller 2>&1 | grep 'wger: install hold OPENED'"):
break
w.wait_app(SUB, "/en/user/login", want=("200",), tries=72)
lk.p("axes settings inside the app:", w.guest("""cat > /tmp/axs.py <<'PY'
import os, sys
sys.path.insert(0, '/home/wger/src'); os.chdir('/home/wger/src')
os.environ.setdefault('DJANGO_SETTINGS_MODULE', 'settings.main')
import django; django.setup()
from django.conf import settings as s
print(s.AXES_LOCKOUT_PARAMETERS, s.AXES_FAILURE_LIMIT, s.AXES_COOLOFF_TIME, s.AXES_RESET_COOL_OFF_ON_FAILURE_DURING_LOCKOUT, s.AXES_HANDLER, s.CACHES['default']['BACKEND'])
PY
docker cp /tmp/axs.py wger:/tmp/axs.py && docker exec wger python3 /tmp/axs.py 2>&1 | tail -1; docker exec wger sh -c 'ps -eo args' | grep -E '[g]unicorn|[r]unserver|[u]vicorn' | head -3""").strip())
pw = (w.GENERATED.get("wger") or {}).get("ADMIN_PASSWORD", "")
pw2 = "Second-" + str(int(time.time()))
lk.p("second member:", w.guest(
"docker exec wger python3 -c \"import os,sys; sys.path.insert(0,'/home/wger/src'); os.chdir('/home/wger/src'); "
"os.environ.setdefault('DJANGO_SETTINGS_MODULE','settings.main'); import django; django.setup(); "
"from django.contrib.auth.models import User; u=User.objects.create_user('second', 'second@gate.invalid', sys.argv[1]); print('made', u.username)\" " + pw2).strip())
lk.p("positive controls — admin:", lk.wger("admin", pw), "| second:", lk.wger("second", pw2))
for i in range(1, 12):
lk.p(f"stranger wrong try {i} (admin):", lk.wger("admin", f"wrong-{i}-{time.time()}"))
t0 = time.time()
lk.p("household — admin RIGHT password:", lk.wger("admin", pw), "| second member RIGHT password:", lk.wger("second", pw2))
if MODE == "fix":
def at(minutes, label, who="admin", p=None):
while time.time() - t0 < minutes * 60:
time.sleep(10)
r = lk.wger(who, p or pw)
lk.p(f"+{(time.time() - t0) / 60:.1f} min {label}:", r)
return r
at(8, "admin RIGHT password DURING the lock (does it restart the 15 min?)")
lk.p("second member, same moment:", lk.wger("second", pw2))
for m in (16, 24, 32, 40):
if at(m, "admin RIGHT password") == "ok":
break
lk.p("a wrong one after:", lk.wger("admin", "wrong-after"), "| right again:", lk.wger("admin", pw))
lk.p(w.guest("docker logs --since 60m wger 2>&1 | grep -i -E 'axes|lock' | tail -6 | cut -c1-220").strip())
w.remove("wger")
lk.p("removed; deployed =", w.stack("wger").get("deployed"))
@@ -0,0 +1,11 @@
set -e
docker rm -f echo-probe cf-sim >/dev/null 2>&1 || true
docker run -d --name echo-probe --network traefik-public \
-l traefik.enable=true -l 'traefik.http.routers.echo-probe.rule=Host(`echo-probe.enkisfelhom.hu`)' \
-l traefik.http.routers.echo-probe.entrypoints=websecure -l traefik.http.routers.echo-probe.tls=true \
-l traefik.http.services.echo-probe.loadbalancer.server.port=80 traefik/whoami:v1.11 >/dev/null
sleep 5
echo "traefik ip: $(docker inspect traefik --format '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}') echo ip: $(docker inspect echo-probe --format '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}')"
echo "## T: the tunnel's hop simulated — a container on traefik-public (where cloudflared sits) sending what cloudflared sends"
docker run --rm --name cf-sim --network traefik-public curlimages/curl:8.11.1 -sk https://traefik/ -H 'Host: echo-probe.enkisfelhom.hu' \
-H 'X-Forwarded-For: 6.6.6.6, 203.0.113.9' -H 'CF-Connecting-IP: 203.0.113.9' -H 'X-Forwarded-Proto: https' | grep -iE '^(RemoteAddr|X-Forwarded-For|X-Real-Ip|Cf-Connecting-Ip|IP: 172)'
@@ -0,0 +1,151 @@
#!/usr/bin/env python3
"""Login probes for R-752 on 9202 — each one a fresh, cookie-less visitor through traefik (a STRANGER unless it holds
the right password). Every probe returns one word: ok / wrong / locked / other:<detail>. EVIDENCE, NOT PRODUCT."""
import json, os, re, subprocess, tempfile, time
from datetime import datetime, timezone
import walk as w
def now():
return datetime.now(timezone.utc).strftime("%H:%M:%S")
def p(*a):
print(now(), *a, flush=True)
def curl(sub, path, *extra, data=None, jar=None, timeout=20):
args = ["curl", "-sk", "--max-time", str(timeout), "-D", "-", "-H", f"Host: {sub}.{w.DOMAIN}"]
if jar:
args += ["-b", jar, "-c", jar]
args += list(extra)
if data is not None:
args += ["--data-raw", data]
args.append(f"{w.BASE}{path}")
r = subprocess.run(args, capture_output=True, text=True)
return r.stdout or ""
def status(out):
m = re.findall(r"(?m)^HTTP/\S+ (\d{3})", out)
return m[-1] if m else "000"
def location(out):
m = re.findall(r"(?im)^location:\s*(\S+)", out)
return m[-1] if m else ""
def form_token(html, name):
m = re.search(r'name="%s"\s+value="([^"]+)"' % name, html) or re.search(r'value="([^"]+)"\s+name="%s"' % name, html)
return m.group(1) if m else ""
def enc(**kw):
from urllib.parse import urlencode
return urlencode(kw)
# ---- BookStack: the web form (Laravel CSRF _token), key email|ip, 5 tries / 60 s ----------------------------------
def bookstack(email, pw, sub="wiki"):
jar = tempfile.mktemp()
try:
page = curl(sub, "/login", jar=jar)
tok = form_token(page, "_token")
out = curl(sub, "/login", "-H", "Content-Type: application/x-www-form-urlencoded", jar=jar,
data=enc(_token=tok, email=email, password=pw))
loc = location(out)
if status(out) == "302" and not loc.rstrip("/").endswith("/login"):
return "ok"
back = curl(sub, "/login", jar=jar)
if re.search(r"(?i)too many|throttle|t[uú]l sok", back):
return "locked"
if status(out) in ("302", "422", "200"):
return "wrong"
return "other:" + status(out)
finally:
try:
os.unlink(jar)
except OSError:
pass
# ---- Grafana: POST /login JSON ----------------------------------------------------------------------------------
def grafana(user, pw, sub="grafana"):
out = curl(sub, "/login", "-H", "Content-Type: application/json", data=json.dumps({"user": user, "password": pw}))
body = out.split("\r\n\r\n")[-1]
if status(out) == "200":
return "ok"
if re.search(r"(?i)temporarily blocked|too many", body):
return "locked"
if status(out) in ("401", "400"):
return "wrong"
return "other:" + status(out) + " " + body[:80]
# ---- calibre-web-automated: the web form (Flask-WTF csrf_token), and OPDS basic auth -----------------------------
def calibre(user, pw, sub="books"):
jar = tempfile.mktemp()
try:
page = curl(sub, "/login", jar=jar)
tok = form_token(page, "csrf_token")
out = curl(sub, "/login", "-H", "Content-Type: application/x-www-form-urlencoded", jar=jar,
data=enc(csrf_token=tok, username=user, password=pw, next="/"))
if status(out) == "302" and "/login" not in location(out):
return "ok"
if re.search(r"(?i)wait one minute|v[aá]rj", out):
return "locked"
if re.search(r"(?i)wrong username|hib[aá]s", out) or status(out) == "200":
return "wrong"
return "other:" + status(out)
finally:
try:
os.unlink(jar)
except OSError:
pass
def calibre_opds(user, pw, sub="books"):
out = curl(sub, "/opds", "-u", f"{user}:{pw}")
c = status(out)
return {"200": "ok", "401": "wrong", "429": "locked"}.get(c, "other:" + c)
# ---- wger: the web login form (Django CSRF), as the catalog's own fixture logs in; django-axes on authenticate() ----
def wger(user, pw, sub="fitness"):
jar = tempfile.mktemp()
try:
o = f"https://{sub}.{w.DOMAIN}"
hdr = ["-H", f"Origin: {o}", "-H", f"Referer: {o}/en/user/login"]
page = curl(sub, "/en/user/login", *hdr, jar=jar)
m = re.search(r'name="csrfmiddlewaretoken" value="([^"]+)"', page)
if not m:
return "other:no-csrf " + status(page)
out = curl(sub, "/en/user/login", *hdr, "-H", "Content-Type: application/x-www-form-urlencoded", jar=jar,
data=enc(csrfmiddlewaretoken=m.group(1), login=user, password=pw))
c = status(out)
jt = open(jar).read() if os.path.exists(jar) else ""
if c == "302" and "sessionid" in jt:
return "ok"
if c in ("429", "403") or re.search(r"(?i)account locked|too many", out):
return "locked"
if c == "200":
return "wrong"
return "other:" + c
finally:
try:
os.unlink(jar)
except OSError:
pass
def wait_until(fn, want, every, cap_s, label):
"""Poll fn() every `every` s until it answers `want`; returns minutes waited (None if never)."""
t0 = time.time()
while time.time() - t0 < cap_s:
time.sleep(every)
r = fn()
p(f" +{(time.time() - t0) / 60:.1f} min {label}: {r}")
if r == want:
return (time.time() - t0) / 60
return None
@@ -0,0 +1,60 @@
#!/usr/bin/env python3
"""Point guest 9202 at the drill catalog (and a 90 s health timeout), or restore the saved config.
`09` §6.5: `git.repo_url` alone is INERT (R-615) — the cache dir must go too. The saved copy is
`controller.yaml.pre-lockouts1001` (NOT the older `.pre-28`, which a restore must never pick up).
"""
import re, sys, io
sys.path.insert(0, '.')
import walk as w
VOL = "/var/lib/docker/volumes/felhom-controller-data/_data"
DRILL_REPO = "https://gitea.dooplex.hu/admin/app-catalog-drill.git"
def creds():
for l in io.open("/home/kisfenyo/.git-credentials").read().strip().split("\n"):
m = re.match(r'https://(admin):([^@]+)@gitea\.dooplex\.hu', l)
if m:
return m.group(1), m.group(2)
raise SystemExit("no admin credential")
def to_drill():
u, t = creds()
print(w.guest(f"""
set -e
test -f {VOL}/controller.yaml.pre-lockouts1001 || cp -p {VOL}/controller.yaml {VOL}/controller.yaml.pre-lockouts1001
python3 - <<'PY'
import re
p = "{VOL}/controller.yaml"
s = open(p).read()
s = re.sub(r'(^\\s+repo_url: ).*$', r'\\g<1>{DRILL_REPO}', s, count=1, flags=re.M)
s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+token: ).*$', r'\\g<1>"{t}"', s, count=1, flags=re.M)
s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+username: ).*$', r'\\g<1>"{u}"', s, count=1, flags=re.M)
if not re.search(r'^update:', s, re.M):
s += "update:\\n health_timeout: 90s\\n"
open(p, "w").write(s)
PY
rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache
docker restart felhom-controller >/dev/null
sleep 15
grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: <redacted>/'
grep -A2 '^update:' {VOL}/controller.yaml
"""))
def restore():
print(w.guest(f"""
set -e
cp -p {VOL}/controller.yaml.pre-lockouts1001 {VOL}/controller.yaml
rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache
docker restart felhom-controller >/dev/null
sleep 15
grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: <redacted>/'
grep -c '^update:' {VOL}/controller.yaml || true
"""))
if __name__ == "__main__":
to_drill() if sys.argv[1] == "drill" else restore()
@@ -0,0 +1,560 @@
#!/usr/bin/env python3
"""walk.py — ONE app's full update walk on guest 9202, through the product's own endpoints.
EVIDENCE, NOT PRODUCT. It presses exactly the buttons a person presses:
POST /api/stacks/<n>/deploy · POST /api/sync · POST /api/stacks/rescan
POST /api/stacks/<n>/update · POST /api/stacks/<n>/remove
and reads GET /api/stacks/<n>. No controller code exists for it.
The walk, per `09` §6.4 and the update-night brief §4:
1 deploy from the DRILL catalog at the LIVE pin
2 seed through the app's OWN front door (R-156: never a volume, never SQL)
3 read the seed back <- control C1; a fixture that cannot prove itself proves nothing
4 „Mentés most"
5 commit the real one-step bump to the DRILL repo, sync, rescan, read the badge in BOTH languages
6 press the guarded Update, record every phase with timestamps
7 read the seed back through the front door
8 the four version observables side by side
9 write the verdict record in `09`'s JSON shape
`inconclusive` is a first-class verdict and is NEVER collapsed into `failed`.
"""
import argparse, json, os, re, subprocess, sys, time
from datetime import datetime, timezone
SC = os.environ.get('SC', '/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/b4d68b9b-a6cf-4d21-8220-ece956837fc3/scratchpad')
EV = os.environ.get('EV', '/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/lockouts-2026-10-01')
DRILL = "/mnt/5_hdd/felhom.eu/drill/app-catalog-drill"
# GUEST=9201 selects demo-hp's hub-enabled guest (the mail proof); default 9202, the scratch guest.
GUEST = os.environ.get("GUEST", "9202")
BASE = os.environ.get("BASE") or {"9202": "https://192.168.0.114", "9201": "https://192.168.0.155"}[GUEST]
DOMAIN = os.environ.get("DOMAIN", "enkisfelhom.hu")
HOSTHDR = f"Host: felhom.{DOMAIN}"
HP = "demo-hp"
LOG = []
def say(*a):
line = " ".join(str(x) for x in a)
ts = datetime.now().strftime("%H:%M:%S")
print(f"{ts} {line}", flush=True)
LOG.append(f"{ts} {line}")
def sh(args, timeout=300, inp=None):
try:
return subprocess.run(args, capture_output=True, text=True, timeout=timeout, input=inp)
except (subprocess.TimeoutExpired, OSError) as e:
return subprocess.CompletedProcess(args, 124, "", f"{e}")
def guest(script, timeout=600):
"""Run a bash script inside guest 9202. Piped as a file — never as an argument (quoting)."""
# ONE TEMP FILE PER CALL (night 2026-09-23): the shared /tmp/w<guest>.sh swapped scripts under
# two concurrent walks (memory: guest-helper-shares-one-tmp-file).
import secrets as _s
t = f"/tmp/w{GUEST}-{os.getpid()}-{_s.token_hex(4)}.sh"
r = sh(["ssh", "-o", "ConnectTimeout=20", "-o", "StrictHostKeyChecking=accept-new", HP,
f"export LC_ALL=C; cat > {t}; pct push {GUEST} {t} {t} >/dev/null 2>&1; "
f"pct exec {GUEST} -- bash {t}; pct exec {GUEST} -- rm -f {t}; rm -f {t}"],
timeout=timeout, inp=script)
return r.stdout or ""
def login():
pw = open(f"{SC}/.ctlpw").read().strip()
sh(["curl", "-sk", "-D", f"{SC}/hdr{os.getpid()}.txt", "-o", "/dev/null", "-H", HOSTHDR,
"-X", "POST", "--data-urlencode", f"password={pw}", f"{BASE}/login"])
h = open(f"{SC}/hdr{os.getpid()}.txt").read()
m = re.search(r"felhom_session=[A-Za-z0-9._-]+", h, re.I)
if not m:
sys.exit("login failed: no session cookie")
open(f"{SC}/sess{os.getpid()}.txt", "w").write(m.group(0))
r = sh(["curl", "-sk", "-L", "-H", HOSTHDR, "-H", f"Cookie: {m.group(0)}", f"{BASE}/"])
c = re.search(r'<meta name="csrf-token" content="([^"]+)"', r.stdout or "")
if not c:
sys.exit("login failed: no csrf token")
open(f"{SC}/csrf{os.getpid()}.txt", "w").write(c.group(1))
def ctl(method, path, data=None, raw=False, tries=2):
"""One controller API call. Re-logs in once on a 302/401 — the controller's session store is
in memory, so any controller restart during the night invalidates it silently."""
for attempt in range(tries):
sess = open(f"{SC}/sess{os.getpid()}.txt").read().strip()
csrf = open(f"{SC}/csrf{os.getpid()}.txt").read().strip()
args = ["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", "-w", "\n%{http_code}"]
if method != "GET":
args += ["-H", f"X-CSRF-Token: {csrf}", "-H", "Content-Type: application/json",
"-X", method]
if data is not None:
args += ["--data", json.dumps(data)]
args.append(f"{BASE}{path}")
r = sh(args)
body, _, code = (r.stdout or "").rpartition("\n")
if code.strip() in ("302", "401") and attempt + 1 < tries:
login()
continue
if raw:
return code.strip(), body
try:
return code.strip(), json.loads(body)
except Exception:
return code.strip(), {"_raw": body[:600]}
return code.strip(), {"_raw": body[:600]}
def page(path):
sess = open(f"{SC}/sess{os.getpid()}.txt").read().strip()
r = sh(["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", f"{BASE}{path}"])
return r.stdout or ""
GATE = {} # sub -> the felhom_gate cookie the household's browser would hold (setup gate, v0.280.0)
def _loc(out):
m = re.search(r"(?im)^location:\s*(\S+)", out or "")
return m.group(1) if m else ""
def gate_cookie(sub):
"""Pass the setup gate (`09` decision 46) the way the HOUSEHOLD does: the app host redirects to the
dashboard's /__gate/start, which (with the dashboard session) redirects back to the app host's
/__felhom_gate/cb, which sets `felhom_gate`. Never printed."""
import urllib.parse
sess = open(f"{SC}/sess{os.getpid()}.txt").read().strip()
r = sh(["curl", "-sk", "-D", "-", "-o", "/dev/null", "-H", "Accept: text/html", "-H", f"Host: {sub}.{DOMAIN}", f"{BASE}/"])
loc = _loc(r.stdout)
if "/__gate/start" not in loc:
GATE[sub] = ""
return "" # not gated (open, or no gate for this app)
u = urllib.parse.urlsplit(loc)
r = sh(["curl", "-sk", "-D", "-", "-o", "/dev/null", "-H", "Accept: text/html", "-H", f"Host: {u.hostname}", "-H", f"Cookie: {sess}",
f"{BASE}{u.path}?{u.query}"])
loc = _loc(r.stdout)
u = urllib.parse.urlsplit(loc)
if "/__felhom_gate/cb" not in u.path:
say(f" gate: the dashboard did not hand back a callback for {sub} ({loc[:80]})")
return ""
r = sh(["curl", "-sk", "-D", "-", "-o", "/dev/null", "-H", "Accept: text/html", "-H", f"Host: {u.hostname}", f"{BASE}{u.path}?{u.query}"])
m = re.search(r"(?im)^set-cookie:\s*(felhom_gate=[^;\r\n]+)", r.stdout or "")
GATE[sub] = m.group(1) if m else ""
say(f" gate: {sub} is gated — passed as the household (cookie {'set' if GATE[sub] else 'NOT set'})")
return GATE[sub]
def app_curl(sub, path, *extra, method=None, data=None, timeout=45, _retry=True):
"""A call to the APP's own front door on 9202 — the household's route, not ours. Carries the setup
gate's cookie when the app is gated, merged into a fixture's own Cookie header (never a second one)."""
raw = list(extra)
gc = GATE[sub] if sub in GATE else gate_cookie(sub)
ext = list(raw)
if gc:
merged = False
for i, a in enumerate(ext):
if isinstance(a, str) and a.lower().startswith("cookie:") and i > 0 and ext[i - 1] == "-H":
ext[i] = a + "; " + gc
merged = True
if not merged:
ext = ["-H", f"Cookie: {gc}"] + ext
args = ["curl", "-sSk", "--max-time", str(timeout), "-H", f"Host: {sub}.{DOMAIN}",
"-w", "\n%{http_code} %{redirect_url}"]
if method:
args += ["-X", method]
if data is not None:
args += ["--data-binary", "@-"]
args += ext + [f"{BASE}{path}"]
r = sh(args, timeout=timeout + 30, inp=data)
body, _, tail = (r.stdout or "").rpartition("\n")
code, _, redir = tail.strip().partition(" ")
if _retry and "/__gate/start" in redir:
GATE.pop(sub, None) # the gate cookie expired or was never taken — log in as the household again
return app_curl(sub, path, *raw, method=method, data=data, timeout=timeout, _retry=False)
return r.returncode, code.strip(), body
def stack(name):
_, d = ctl("GET", f"/api/stacks/{name}")
return (d.get("data") or {}) if isinstance(d, dict) else {}
def wait_app(sub, path="/", want=("200", "302", "303", "401", "403"), tries=60, delay=5):
"""Settling says the container runs; this says the APP answers. Not the same thing."""
last = None
for _ in range(tries):
rc, code, _ = app_curl(sub, path, timeout=15)
last = (rc, code)
if rc == 0 and code in want:
return True
time.sleep(delay)
say(f" app never answered on {sub}{path} (last rc={last[0]} code={last[1]})")
return False
# ------------------------------------------------------------------ the walk
DRIVE = "/mnt/felhom-drives/scratch_hdd/userdata"
# What THIS run generated for a deploy, per app. Deploy secrets are ENCRYPTED AT REST in
# `app.yaml` (`ENC:…`), which is right and which means a fixture cannot read an app's admin
# password back off the box — the household sees it once. So the value the harness itself
# generated is kept here for the life of the run, and nowhere else.
GENERATED = {}
def deploy_values(name, sub):
"""Fill EVERY required deploy field the way the wizard would, by asking the box what this app
asks for — `GET /api/stacks/<n>/deploy-fields` — instead of assuming DOMAIN+SUBDOMAIN.
Measured 2026-09-21: three apps in one batch refused at the deploy with a correct 400 because
a required field was absent — `HDD_PATH` (navidrome, audiobookshelf) and an admin password
(grafana). The refusals happen BEFORE anything is created (`deploy.go:324`), which is the only
reason this was safe to discover by running it (live-probes rule).
A `path` field must name a directory that ALREADY EXISTS (`deploy.go:330`), so one is made on
the scratch drive first — the same act the drive browser performs for a household.
"""
code, d = ctl("GET", f"/api/stacks/{name}/deploy-fields")
fields = (((d.get("data") or {}).get("metadata") or {}).get("deploy_fields")) or []
values = {"DOMAIN": DOMAIN, "SUBDOMAIN": sub}
made = []
for f in fields:
ev, ty = f.get("env_var"), f.get("type")
if ev in values:
continue
# `type: password` is MANDATORY whatever `required` says — `deploy.go:305-312` refuses
# when the caller sends none, deliberately ("the user needs to know their password"),
# while `.felhom.yml` declares `required: false` and the API serves that verbatim. A
# caller that trusts the contract gets a 400. Measured tonight on grafana; filed.
if not f.get("required") and ty != "password":
continue # the controller generates the optional secrets itself
if ty == "path":
p = f"{DRIVE}/{name}"
values[ev] = p
made.append(p)
elif ty in ("secret", "password"):
import secrets as _s
values[ev] = "Drill-" + _s.token_hex(12)
GENERATED.setdefault(name, {})[ev] = values[ev]
elif f.get("default"):
values[ev] = f["default"]
else:
values[ev] = f"drill-{name}"
if made:
guest("mkdir -p " + " ".join(made) + "; ls -ld " + " ".join(made))
say(f" [1] made the drive paths this app requires: {made}")
extra = [k for k in values if k not in ("DOMAIN", "SUBDOMAIN")]
if extra:
say(f" [1] required fields filled beyond DOMAIN/SUBDOMAIN: {extra}")
return values
def deploy(name, sub, extra_values=None):
st = stack(name)
if st.get("deployed"):
say(f" [1] {name} already deployed — reusing")
return True
values = deploy_values(name, sub)
if extra_values:
values.update(extra_values)
body = {"values": values}
if os.environ.get("KEPT"): # decision 36: the household's answer when the drive holds old data ("fresh" moves it aside, deletes nothing)
body["kept_data"] = os.environ["KEPT"]
code, d = ctl("POST", f"/api/stacks/{name}/deploy", body)
say(f" [1] deploy -> {code} {str(d)[:120]}")
if code != "202":
return False
# WAIT FOR `deployed`, NOT FOR `running`. Measured 2026-09-21 on tandoor: docker reported the
# container `healthy` while the controller's own state read `unhealthy` — a gate on `running`
# alone therefore times out on an app that is up. The state is RECORDED rather than required;
# the real gate is the fixture's own `wait_app`, which asks whether the APP answers.
seen = None
for _ in range(90):
time.sleep(5)
st = stack(name)
seen = st.get("state")
# `deployed` alone is NOT enough and `state` alone is NOT right. Measured 2026-09-21:
# tandoor reads `unhealthy` while serving (R-618), so gating on "running" hangs; and romm
# read `deployed=True, state=degraded, pinned_images=None` twenty seconds in, i.e. the
# deploy had not finished writing app.yaml. The PIN is the deploy's own completion mark
# (`runComposeDeploy` writes it), so that is what to wait for.
pins = (st.get("app_config") or {}).get("pinned_images")
if st.get("deployed") and pins and seen in ("running", "unhealthy", "degraded"):
say(f" [1] deployed, controller state={seen}, "
f"pinned={(st.get('app_config') or {}).get('pinned_images')}")
if seen != "running":
say(f" [1] NOTE: the controller's own state is {seen!r}, not 'running' — recorded, "
f"not treated as a failure; the fixture's front-door wait is the real gate")
return True
say(f" [1] never became deployed (last controller state={seen!r})")
return False
def backup_now(name):
"""R-648 (2026-09-23): NO whole-box „Mentés most" from a drill, ever.
`POST /api/backup/run` is the only backup endpoint and it is WHOLE-BOX: on 9201 it stopped and
restarted 9 of 10 standing apps twice, and on 9202 it broke a deploy in flight (R-634). The product
has NO per-app backup endpoint (router.go: /backup/run, /backup/tier2 only); the per-app backup
exists only inside the guarded update, whose `backing-up` phase calls RunAppBackupNow for the one
app. So this presses nothing: the update takes the throwaway app's own backup, and says so in its
phase list. A seed written "after the backup" is therefore written before the update's own backup
— the undo's last-second copy is still the one that must bring it back."""
say(f" [4] backup press SKIPPED for {name} (R-648: whole-box only; the update's backing-up phase backs up {name} alone)")
return None
def drill_bump(app, frm, to, service_hint=None):
"""Serialised across concurrent walks: one git working tree, one lock."""
import fcntl
with open(f"{SC}/drill.lock", "w") as lk:
fcntl.flock(lk, fcntl.LOCK_EX)
sh(["git", "-C", DRILL, "pull", "-q", "--rebase", "origin", "main"], timeout=120)
return _drill_bump(app, frm, to, service_hint)
def _drill_bump(app, frm, to, service_hint=None):
"""Commit the edge to the DRILL repo. catalog_since set by hand (the drill repo has no gates).
`frm`/`to` may be comma-separated lists of the SAME length: an app whose own version lives in
two images (adventurelog's backend and frontend) moves both in one edge, while its engine
sidecar stays where it is — `09` §3b Q3's rule is per SERVICE, and an app-half edge must move
every service that carries the app's own version and no others.
"""
comp = f"{DRILL}/templates/{app}/docker-compose.yml"
fy = f"{DRILL}/templates/{app}/.felhom.yml"
s = open(comp).read()
froms = [x.strip() for x in frm.split(",") if x.strip()]
tos = [x.strip() for x in to.split(",") if x.strip()]
if len(froms) != len(tos):
say(f" [5] from/to lists differ in length: {froms} vs {tos}")
return None
for f1, t1 in zip(froms, tos):
if f"image: {f1}" not in s:
say(f" [5] FROM ref not found in compose: {f1}")
return None
s = s.replace(f"image: {f1}", f"image: {t1}")
open(comp, "w").write(s)
f = open(fy).read()
today = datetime.now().strftime("%Y-%m-%d")
f = re.sub(r'^catalog_since:.*$', f'catalog_since: "{today}"', f, count=1, flags=re.M)
open(fy, "w").write(f)
sh(["git", "-C", DRILL, "add", "-A"])
sh(["git", "-C", DRILL, "commit", "-q", "-m", f"DRILL {app}: {frm} -> {to}"])
r = sh(["git", "-C", DRILL, "push", "-q", "origin", "main"], timeout=120)
h = sh(["git", "-C", DRILL, "rev-parse", "--short=12", "HEAD"]).stdout.strip()
say(f" [5] drill commit {h}: {app} {frm} -> {to} (push rc={r.returncode})")
return h
def sync_rescan(expect_app=None, expect_ref=None, tries=12, delay=5):
"""Sync, rescan, and — when told what to expect — WAIT FOR THE BADGE TO CATCH UP.
R-607: `POST /api/sync` answers "nincs valtozas" while the catalog HAS moved, and
`catalog_images` stays stale until a separate rescan. Tonight showed the rescan alone is not
enough either: mealie's badge read "Naprakesz" seconds after its bump was pushed, and the
Update that followed moved nothing and still reported "Frissitve". So when the caller knows
which reference should appear, this polls for it and SAYS HOW LONG IT TOOK — which is the
NUMBER R-607 asks for and has never had.
"""
t0 = time.time()
ctl("POST", "/api/sync")
time.sleep(2)
ctl("POST", "/api/stacks/rescan")
time.sleep(2)
if not expect_app or not expect_ref:
return None
for i in range(tries):
cat = stack(expect_app).get("catalog_images") or {}
if expect_ref in cat.values():
waited = round(time.time() - t0, 1)
if i:
say(f" [sync] the badge needed {waited}s and {i+1} sync+rescan rounds to catch up "
f"to {expect_ref} — R-607's window, measured")
return waited
time.sleep(delay)
ctl("POST", "/api/sync")
time.sleep(1)
ctl("POST", "/api/stacks/rescan")
say(f" [sync] the badge NEVER caught up to {expect_ref} in {round(time.time()-t0,1)}s — "
f"catalog_images = {stack(expect_app).get('catalog_images')}")
return None
def badges(name):
out = {}
for lang, suffix in (("hu", ""), ("en", "?lang=en")):
h = page(f"/apps/{name}{suffix}")
m = re.findall(r'<span class="tag tag-[^"]*"[^>]*title="([^"]*)"[^>]*>([^<]*)<', h)
out[lang] = [{"title": a.strip(), "text": b.strip()} for a, b in m][:3]
return out
def press_update(name, poll=1.0, cap_s=1800):
code, d = ctl("POST", f"/api/stacks/{name}/update")
say(f" [6] Update -> {code} {str(d)[:220]}")
if code not in ("202", "200"):
return {"accepted": False, "http": code, "refusal": d, "phases": [], "duration_s": 0}
phases, seen, t0 = [], None, time.time()
while time.time() - t0 < cap_s:
st = stack(name)
ph = st.get("update_phase")
if ph != seen:
seen = ph
rec = {"t": round(time.time() - t0, 1), "phase": ph,
"label": st.get("update_phase_label"), "updating": st.get("updating"),
"error": st.get("update_error"), "hold": st.get("hold_reason")}
phases.append(rec)
say(f" +{rec['t']:>6.1f}s phase={ph} label={rec['label']} "
f"err={rec['error']} hold={rec['hold']}")
if not st.get("updating") and ph in ("done", "failed", "undone", None) and time.time() - t0 > 3:
break
time.sleep(poll)
st = stack(name)
return {"accepted": True, "http": code, "phases": phases,
"duration_s": round(time.time() - t0, 1),
"final_phase": st.get("update_phase"), "update_error": st.get("update_error"),
"hold_reason": st.get("hold_reason"), "state": st.get("state")}
def observables(name):
st = stack(name)
ac = st.get("app_config") or {}
live = guest(f"""
grep -E '^\\s+image:' /opt/docker/stacks/{name}/docker-compose.yml 2>/dev/null | sed 's/^ *//'
echo '---inspect---'
for c in $(docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'); do
echo -n "$c "; docker inspect "$c" --format '{{{{.Config.Image}}}} running={{{{.State.Running}}}} restarts={{{{.RestartCount}}}}'
done
""")
a, _, b = live.partition("---inspect---")
return {
"pinned_images": ac.get("pinned_images"),
"installed_images": {k: (v.get("ref") if isinstance(v, dict) else v)
for k, v in (ac.get("installed_images") or {}).items()},
"catalog_images": st.get("catalog_images"),
"live_compose_image_lines": [x for x in a.strip().splitlines() if x.strip()],
"docker_inspect": [x for x in b.strip().splitlines() if x.strip()],
}
def app_logs(name, lines=400):
"""The app's own container log, DECODED. The endpoint answers a JSON envelope whose `logs` is
one string with escaped newlines — a scan over the envelope sees a single enormous line and
finds nothing, which reads exactly like "the app printed no migration line" and is not. R-96
rule 3 in a new place: an absent line is not evidence when the instrument cannot see lines."""
code, d = ctl("GET", f"/api/stacks/{name}/logs?lines={lines}")
if isinstance(d, dict):
data = d.get("data")
if isinstance(data, dict) and isinstance(data.get("logs"), str):
return data["logs"]
if isinstance(d.get("_raw"), str):
return d["_raw"]
return str(d)
def write_verdict(rec, appdir):
os.makedirs(appdir, exist_ok=True)
p = os.path.join(appdir, "verdict.json")
json.dump(rec, open(p, "w"), indent=2, ensure_ascii=False)
say(f" [9] verdict {rec['verdict']} -> {p}")
def remove(name):
"""Remove through the PRODUCT, never `docker rm` (live-probes rule). The remove endpoint
refuses a running stack — `409 still running` — so the stop is part of the act, not a tidy-up."""
c1, d1 = ctl("POST", f"/api/stacks/{name}/stop")
say(f" [X] stop -> {c1} {str(d1)[:100]}")
for _ in range(24):
time.sleep(5)
if stack(name).get("state") != "running":
break
code, d = ctl("POST", f"/api/stacks/{name}/remove",
{"remove_hdd_data": True, "remove_backups": True})
say(f" [X] remove (with drive data) -> {code} {str(d)[:160]}")
if code == "409":
# R-442's fail-closed guard: when the storage subsystem cannot RESOLVE the app's drive
# path, the removal is REFUSED and the app is kept rather than half-deleted. On guest 9202
# `/api/disks` answers `agent not configured`, so every app deployed with an HDD_PATH hits
# this. The household's other choice — remove the app, KEEP the data — is accepted, and the
# harness takes it, then tidies its own directory by name at teardown.
say(" [X] refused because the drive path cannot be resolved (R-442, fail-closed and right)"
" — removing the app and KEEPING the drive data instead")
code, d = ctl("POST", f"/api/stacks/{name}/remove",
{"remove_hdd_data": False, "remove_backups": True})
say(f" [X] remove (keeping drive data) -> {code} {str(d)[:160]}")
time.sleep(5)
st = stack(name)
left = guest(f"ls -d /opt/docker/stacks/{name} 2>/dev/null; "
f"docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'")
say(f" [X] after remove: deployed={st.get('deployed')} leftovers={left.strip()!r}")
return code
def app_env(name, key):
"""Read one deploy value the CUSTOMER was given (e.g. the generated admin password) from the
app's own `app.yaml`. This is not seeding — it is how the household logs in; the controller
shows them the same value. Data still goes in through the app's own front door."""
out = guest(f"grep -E '^\\s*{key}:' /opt/docker/stacks/{name}/app.yaml 2>/dev/null | head -1")
if ":" in out:
return out.split(":", 1)[1].strip().strip('"').strip("'")
return ""
def snapshots(name):
"""The restorable copies the backups page offers for this app."""
code, d = ctl("GET", f"/api/backup/snapshots?stack={name}")
data = d.get("data") if isinstance(d, dict) else None
if isinstance(data, dict):
for k in ("snapshots", "items", "restore_points"):
if isinstance(data.get(k), list):
return data[k]
return data if isinstance(data, list) else []
def restore(name, snapshot_id=None, wait_s=1200):
"""The household's own way out: the „Visszaállítás a mentésből" button on the backups page.
A FORM post, not an API call — `POST /backup/restore` with `_csrf`, `stack_name`,
`snapshot_id` — because that is the button the sentence tells them to press.
"""
snaps = snapshots(name)
if snapshot_id is None:
if not snaps:
say(f" [R] no restorable copy offered for {name}")
return {"ok": False, "why": "no snapshot offered", "snapshots": snaps}
first = snaps[0]
snapshot_id = first.get("id") or first.get("snapshot_id") or first.get("short_id")
say(f" [R] restoring {name} from snapshot {snapshot_id!r} (of {len(snaps)} offered)")
sess = open(f"{SC}/sess{os.getpid()}.txt").read().strip()
csrf = open(f"{SC}/csrf{os.getpid()}.txt").read().strip()
r = sh(["curl", "-sk", "-D", "-", "-o", "/dev/null", "-H", HOSTHDR, "-H", f"Cookie: {sess}",
"-X", "POST",
"--data-urlencode", f"_csrf={csrf}",
"--data-urlencode", f"stack_name={name}",
"--data-urlencode", f"snapshot_id={snapshot_id}",
f"{BASE}/backup/restore"], timeout=180)
head = (r.stdout or "").split("\n")[0].strip()
loc = [l for l in (r.stdout or "").split("\n") if l.lower().startswith("location:")]
say(f" [R] POST /backup/restore -> {head} {loc[:1]}")
t0 = time.time()
last = None
while time.time() - t0 < wait_s:
code, d = ctl("GET", "/api/backup/restore-status")
dd = d.get("data") or {}
cur = (dd.get("running"), dd.get("phase") or dd.get("state"), dd.get("message"))
if cur != last:
say(f" +{round(time.time()-t0,1):>6.1f}s restore {cur}")
last = cur
if not dd.get("running", False) and time.time() - t0 > 5:
break
time.sleep(2)
st = stack(name)
say(f" [R] after restore: state={st.get('state')} hold={st.get('hold_reason')!r} "
f"phase={st.get('update_phase')}")
return {"ok": True, "snapshot_id": snapshot_id, "snapshots": snaps,
"http": head, "location": loc[:1], "seconds": round(time.time() - t0, 1),
"state_after": st.get("state"), "hold_after": st.get("hold_reason"),
"observables_after": observables(name)}
+3
View File
@@ -864,6 +864,9 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-750** | **[P3-LOW] The registry no longer holds controller releases older than 0.213.0 (2026-08-12) — something removed them, and nothing records what.** MEASURED 2026-10-01 (anonymous registry API, `audits/rulings-2026-10-01/B/`): `felhom-controller` has 91 tags, the oldest release 0.213.0; `0.201.0` answers 404; Gitea's package list starts 2026-08-12. No runbook, row or memory names a clean-up. Today nothing needs those versions: a box runs a newer one, a whole-guest restore brings the guest's own Docker store back (mp0 `backup=1`), and decision 56 deletes only on the box. **But** a box or a backup that names a removed version cannot pull it again (R-698's shape, for the controller). **Needs:** find what removed them (a Gitea clean-up rule?), and record the rule — or say it was a one-time act. Read-only on DooPlex. | **OPEN — rank P3-LOW; owner: operator (Gitea settings), CC measures** |
| **R-751** | **[P2] The image clean-up after an app update could crash the whole controller: it re-read the app after a rescan and dereferenced a nil stack when the app was gone.** FOUND 2026-10-01 by the full test suite (controller v0.284.2): `RetainImagesAfterUpdate` runs in a goroutine; `TestR705_TheManualLegRunsByDay` removed its temp dir under it → `panic: invalid memory address` at `image_retention.go:291`. In a box the same happens when an app is removed (or its compose vanishes) between an update's end and the clean-up — a panic in a goroutine ends the process (the agent's supervisor restarts it). Fixed in v0.285.0 the same session: it returns when the app is gone; `TestRetainImagesAfterUpdate_AppGoneDoesNotPanic` seen panicking on the old code; both retention seams are no-ops in the stacks tests (`TestMain`), so no test leaves the goroutine running. `audits/rulings-2026-10-01/B/B1-red-proofs.txt` **-- 2026-10-01:** Delivered: floor 0.285.0 reached both demo boxes in ~6 s (hub `managed floor SERVED … from declared`). | **CLOSED 2026-10-01 — controller v0.285.0, floor 0.285.0** |
| **R-752** | **[P3-LOW] Four more catalog apps let a stranger lock the household out with wrong passwords for a known login name — like mealie (R-747).** READ 2026-10-01 in each app's source at its pinned tag (not measured live): **calibre-web-automated v4.0.8** — Flask-Limiter on the login keyed on the lowercased USERNAME, 3/minute and 40/day, checked before the password; the default login is `admin` → up to a day; no env switch (a database setting). **wger 2.7** — django-axes keyed on IP, 10 failures, 30 min, each failure restarts it; behind traefik every client has traefik's IP → everyone is locked out (`AXES_*` env vars exist; `AXES_IPWARE_PROXY_COUNT` 0). **Grafana 13.2.3** — per-account, 5 failures in a sliding 5 minutes; a slow trickle keeps it closed (`GF_SECURITY_*`). **BookStack 26.09.1** — key `email|ip`, 5 tries, 60 s, hard-coded; `APP_PROXIES` empty, so the key is the e-mail alone. gokapi (3 s delay, no lock) and claper (per-IP 10/min, no account lock) cannot. **Needs:** per app, the smallest fix that keeps a guessing guard (calibre-web-automated and wger first — longest and broadest), each proven on 9202 as R-747's was. | **OPEN — rank P3-LOW; owner: CC** |
| **R-753** | **[P3-LOW] Behind the tunnel every visitor reaches an app with the SAME address — the tunnel container's — so every per-address guard is an "everyone" guard and every app's log is blind.** MEASURED 2026-10-01 (`audits/lockouts-2026-10-01/A/A1-client-address.txt`): on demo-hp through its real tunnel, a request from DooPlex's public address reached traefik as `172.18.0.5` (cloudflared, in the guest on `traefik-public`) and BookStack as `172.18.0.3` (traefik); on 9202 an echo container showed `X-Forwarded-For`/`X-Real-Ip` = the sending container for the tunnel's hop (traefik DROPS the incoming chain — good: a client cannot forge it) and the real address from the LAN; `CF-Connecting-IP` passes untouched and is FORGEABLE from the LAN. **No box-wide fix taken:** trusting cloudflared in traefik passes Cloudflare's appended chain, whose LEFTMOST entry the client writes — every app reading the leftmost address would believe it; cloudflared's address is docker-assigned; a single-address rewrite needs a traefik plugin (a new dependency). Per-app fixes trust no header (R-752). **Needs (operator):** whether to build a safe version (cloudflared on a fixed-address network + traefik trusting only it + per-app proxy counts), or keep "one address" and fix per app. Only ONE outside address was available (DooPlex has no IPv6); a second was not measured. | **OPEN — rank P3-LOW; owner: operator (direction), CC measures** |
| **R-754** | **[P3-LOW] `01-topology-and-trust.md` §7 says cloudflared runs on the Proxmox HOST as an agent-managed service; on every box it runs INSIDE the guest as a container the controller renders.** READ 2026-10-01: `felhom-controller` `internal/infra/templates/cloudflared-compose.yml.tmpl` (`container_name: cloudflared`, network `traefik-public`); demo-hp's guest 9201 runs `cloudflared` (ingress `*.enkisfelhom.hu -> https://traefik`); R-505 saw the same in VM 331. A design decision that the build does not follow — the document or the build is wrong, and only the operator decides which (R-370: a design decision is not a defect). | **OPEN — rank P3-LOW; owner: operator (which is right)** |
| **R-755** | **[P3-LOW] wger runs Django's DEVELOPMENT server in production: `manage.py runserver`, because the template does not set `WGER_USE_GUNICORN=True`.** MEASURED 2026-10-01 on 9202 (`ps` in the wger container: `python3 manage.py runserver 0.0.0.0:8000`); wger 2.7's `extras/docker/production/entrypoint.sh:81-87` runs gunicorn only with that switch. Django's own documentation says runserver is not for production (one process, not hardened). Not changed this session (a different change from R-752's; needs its own bench + box proof, memory watch included). | **OPEN — rank P3-LOW; owner: CC (catalog)** |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.