diff --git a/STATUS.md b/STATUS.md index a521784b..d32a5211 100644 --- a/STATUS.md +++ b/STATUS.md @@ -3,7 +3,7 @@ **Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop) is off; nothing was sent to it.** **Updated 2026-10-08 (afternoon): hub 0.143.1; demo-hp, demo-felhom and Tester 1 run agent 0.153.0 and controller -0.303.0 (nothing delivered today — tonight is the second kernel night). The open-items list is at 130. Reports: +0.303.0 (nothing delivered today — tonight is the second kernel night). The open-items list is at 129. Reports: `REPORT-day-2026-10-08.md` (morning), `REPORT-day2-2026-10-08.md` (afternoon).** ## Afternoon (2026-10-08): your four answers built, and one sheet of decisions @@ -29,6 +29,7 @@ | D7 | Should the hub mail you an error when even one off-site snapshot disappears outside a clean-up window it opened? | Yes | ~½ session, hub | A deletion through any other key stays silent unless it removes over half of a household's history | | D8 | After a failed off-site restore of one app: keep the app stopped for support now, and „put back exactly as it was" next? | Yes, both, in that order | ~½ session now; ~1 session + a test + disk space for one copy later | The app restarts on a mix of old files and a newer database, and the screen says all is back as it was | | D9 | May CC install the DooPlex job that removes a deleted customer's copy (dry run first, then daily)? | Yes, after tomorrow's releases are read back | ~30 min; one new local token on DooPlex | A deleted customer's copy stays on DooPlex, and the privacy-notice line „within 30 days" is not true | +| D10 | wger's install check still assumes 100 MB; measured 250 MB with its proper server. Raise it to 256 MB? | Yes — only boxes with room can install it | One line in the catalog; a box short of memory can no longer install wger | The install check lets wger onto a box that cannot hold it (wger is hidden today, so no one meets this yet) | Full designs: `documentation/audits/day-2026-10-08/`. @@ -57,6 +58,7 @@ Full designs: `documentation/audits/day-2026-10-08/`. in the menu goes to the same page in the other language. - **A missing-page (404) page exists now.** - **The language link is now a globe icon**, like the dashboard's: a click shows „Magyar" and „English". +- **The contact form's lost program code was searched for everywhere I can reach: not found.** The running program is now saved in git, so it cannot be lost. A plan for a replacement is written. One question for you: is the February folder on your Windows computer? **Needs you:** two choices in `REPORT-website-refresh.md`. If nothing: SparkyFitness stays listed; the contact mailer's source stays lost. diff --git a/documentation/audits/day-2026-10-08/r762/9202/README.md b/documentation/audits/day-2026-10-08/r762/9202/README.md new file mode 100644 index 00000000..d8691196 --- /dev/null +++ b/documentation/audits/day-2026-10-08/r762/9202/README.md @@ -0,0 +1,33 @@ +# R-762 on scratch 9202 (2026-10-08, 09:50–10:09 CEST) + +wger was installed fresh from the drill catalog (commit `bebbac8`: wger un-hidden plus the gunicorn definition, DRILL +only; reverted as `95876a3`). The install went through the product's deploy endpoint (`tools/box9202.py install`). +`tools/repoint.py` moved 9202's catalog setting to the drill catalog and back. + +| check | result | +|---|---| +| workers | 2 × „Booting worker" (`wger.log`); env `WGER_USE_GUNICORN=True`, `WEB_CONCURRENCY=2` | +| login page | 200 | +| CSS (the page's 3 links, through wger-files) | 200 / 200 / 200 (2481 B, 277042 B, 1006 B, `text/css`) | +| photo | POST `/api/v2/gallery/` 201; read back 200, 179 B, `image/png` | +| control | an unknown `/static/` file 404 | +| seed | a weight entry 201, read back 200 | +| 600 s watch (10,944 requests: 8,208 × 200, 2,736 × 302) | anon peak 258,273,280 B = 246.3 MiB = 64.1 % of 384M (memory.stat `anon`, every 2 s); oom 0, oom_kill 0; restarts 0, oomkilled false, healthy (wger and wger-files) | +| after the watch (`plateau-check.txt`) | anon 246 MiB for 1,200 more login-page loads; oom_kill 0; restarts 0 | + +memory.peak reached the limit (402,657,280 B) because it counts the page cache (`after-watch-memory.txt`: file 22 MiB, +kernel 49 MiB). The anon figure is what counts. + +The bench figure was 170 MiB (44 %). The bench load was 302 redirects only. Here the load also rendered the login page +and served a photo upload. The cause of the difference was not measured. + +**Removal through the product** (`run.txt`): stop 200, remove (with data) 200, volumes `wger_wger_data`, +`wger_wger_media` and `wger_wger_static` removed. Afterwards: no wger container, no wger volume, `deployed=False`. +`/opt/docker/stacks/wger` holds only the catalog's synced `.felhom.yml` and `docker-compose.yml`. It held the same two +files before the test, because every catalog app has this directory. + +**Put back:** `repo_url` = `https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git`. `controller.yaml` sha256 +`7739ad7b…` is the same before and after. The save copy was deleted. Deployed apps are paperless-ngx and privatebin, +as before. + +The generated admin password was never written. The image's default password in `wger.log` is redacted. diff --git a/documentation/audits/day-2026-10-08/r762/9202/after-watch-memory.txt b/documentation/audits/day-2026-10-08/r762/9202/after-watch-memory.txt new file mode 100644 index 00000000..e78f385c --- /dev/null +++ b/documentation/audits/day-2026-10-08/r762/9202/after-watch-memory.txt @@ -0,0 +1,14 @@ +anon 258273280 +file 23224320 +kernel 51847168 +shmem 0 +peak=402657280 +wger-files 7.48MiB / 32MiB +wger 295.8MiB / 384MiB + PID RSS COMMAND + 1 3804 /bin/bash /home/wger/entrypoint.sh + 58 132060 /usr/bin/python3 /home/wger/.local/bin/gunicorn wger.wsgi:application --preload --bind 0.0.0.0:8000 + 59 127124 /usr/bin/python3 /home/wger/.local/bin/gunicorn wger.wsgi:application --preload --bind 0.0.0.0:8000 + 60 121560 /usr/bin/python3 /home/wger/.local/bin/gunicorn wger.wsgi:application --preload --bind 0.0.0.0:8000 + 464 4316 ps -o pid,rss,args + diff --git a/documentation/audits/day-2026-10-08/r762/9202/box-verdict-wger.json b/documentation/audits/day-2026-10-08/r762/9202/box-verdict-wger.json new file mode 100644 index 00000000..d0dc0a1e --- /dev/null +++ b/documentation/audits/day-2026-10-08/r762/9202/box-verdict-wger.json @@ -0,0 +1,32 @@ +{ + "app": "wger", + "venue": "scratch guest 9202 on demo-hp, drill catalog, the product's deploy endpoint", + "drill_commit": "bebbac8 DRILL wger: un-hidden + the R-762 gunicorn definition (DRILL only, 2026-10-08)", + "pinned": { + "wger": "wger/server:2.7", + "wger-files": "nginx:1.30.5-alpine" + }, + "seed_read": true, + "login_page": "200", + "css": { + "/static/css/workout-manager.7007d84ce531.css": "200 2481 text/css", + "/static/bootstrap-compiled.80a6279921f8.css": "200 277042 text/css", + "/static/css/bootstrap-custom.400ad578123c.css": "200 1006 text/css" + }, + "photo_post": "201", + "photo_bytes_posted": 179, + "photo_path": "/media/gallery/1/f3391508-f895-4fc4-9688-7a42dc158932.png", + "photo_read": "200 179 image/png", + "control_unknown_static": "404 153 text/html", + "container": "cgroup=/sys/fs/cgroup/system.slice/docker-42589be792575b34221f0ff8ab0141f576e21db37a1ce1a4214e7e3385e1da46.scope\nWGER_USE_GUNICORN=True\nWEB_CONCURRENCY=2\nUsing gunicorn on port 8000...\n[2026-10-08 09:54:38 +0200] [58] [INFO] Starting gunicorn 26.1.0\n[2026-10-08 09:54:38 +0200] [58] [INFO] Listening at: http://0.0.0.0:8000 (58)\n[2026-10-08 09:54:38 +0200] [59] [INFO] Booting worker with pid: 59\n[2026-10-08 09:54:38 +0200] [60] [INFO] Booting worker with pid: 60\n402653184\n", + "watch_s": 600.1, + "watch_requests": 10944, + "watch_codes": { + "200": 8208, + "302": 2736 + }, + "watch": "anon_max=258273280\nmemory.max=402653184\noom 0\noom_kill 0\nrestarts=0 oomkilled=false status=running health=healthy started=2026-10-08T07:53:22.480653193Z\nrestarts=0 oomkilled=false status=running health=healthy started=2026-10-08T07:53:22.678387772Z\n2\n", + "anon_max_bytes": 258273280, + "anon_max_mib": 246.3, + "anon_pct_of_384M": 64.1 +} \ No newline at end of file diff --git a/documentation/audits/day-2026-10-08/r762/9202/plateau-check.txt b/documentation/audits/day-2026-10-08/r762/9202/plateau-check.txt new file mode 100644 index 00000000..7b07b351 --- /dev/null +++ b/documentation/audits/day-2026-10-08/r762/9202/plateau-check.txt @@ -0,0 +1,15 @@ +08:06:00 round=1 gets=100 anon_MiB=246 +08:06:05 round=2 gets=100 anon_MiB=246 +08:06:10 round=3 gets=100 anon_MiB=246 +08:06:14 round=4 gets=100 anon_MiB=246 +08:06:19 round=5 gets=100 anon_MiB=246 +08:06:23 round=6 gets=100 anon_MiB=246 +08:06:28 round=7 gets=100 anon_MiB=246 +08:06:32 round=8 gets=100 anon_MiB=246 +08:06:37 round=9 gets=100 anon_MiB=246 +08:06:41 round=10 gets=100 anon_MiB=246 +08:06:46 round=11 gets=100 anon_MiB=246 +08:06:50 round=12 gets=100 anon_MiB=246 +oom_kill 0 +restarts=0 + diff --git a/documentation/audits/day-2026-10-08/r762/9202/run.txt b/documentation/audits/day-2026-10-08/r762/9202/run.txt new file mode 100644 index 00000000..c0775320 --- /dev/null +++ b/documentation/audits/day-2026-10-08/r762/9202/run.txt @@ -0,0 +1,39 @@ +09:50:14 drill: bebbac8 DRILL wger: un-hidden + the R-762 gunicorn definition (DRILL only, 2026-10-08) +09:52:49 synced template on the box: 2 18:lifecycle: available +09:54:56 wger: POST /api/v2/weightentry/ http=201 +09:54:56 wger: readback of the seeded weight entry http=200 found=True +09:54:56 login page -> 200; CSS {'/static/css/workout-manager.7007d84ce531.css': '200 2481 text/css', '/static/bootstrap-compiled.80a6279921f8.css': '200 277042 text/css', '/static/css/bootstrap-custom.400ad578123c.css': '200 1006 text/css'} +09:54:57 POST /api/v2/gallery/ (179 B PNG) -> 201; read back /media/gallery/1/f3391508-f895-4fc4-9688-7a42dc158932.png -> 200 179 image/png; control -> 404 153 text/html +09:55:00 container: +cgroup=/sys/fs/cgroup/system.slice/docker-42589be792575b34221f0ff8ab0141f576e21db37a1ce1a4214e7e3385e1da46.scope +WGER_USE_GUNICORN=True +WEB_CONCURRENCY=2 +Using gunicorn on port 8000... +[2026-10-08 09:54:38 +0200] [58] [INFO] Starting gunicorn 26.1.0 +[2026-10-08 09:54:38 +0200] [58] [INFO] Listening at: http://0.0.0.0:8000 (58) +[2026-10-08 09:54:38 +0200] [59] [INFO] Booting worker with pid: 59 +[2026-10-08 09:54:38 +0200] [60] [INFO] Booting worker with pid: 60 +402653184 + +10:05:10 watch: +anon_max=258273280 +memory.max=402653184 +oom 0 +oom_kill 0 +restarts=0 oomkilled=false status=running health=healthy started=2026-10-08T07:53:22.480653193Z +restarts=0 oomkilled=false status=running health=healthy started=2026-10-08T07:53:22.678387772Z +2 + +10:05:13 RECORD {"app": "wger", "venue": "scratch guest 9202 on demo-hp, drill catalog, the product's deploy endpoint", "drill_commit": "bebbac8 DRILL wger: un-hidden + the R-762 gunicorn definition (DRILL only, 2026-10-08)", "pinned": {"wger": "wger/server:2.7", "wger-files": "nginx:1.30.5-alpine"}, "seed_read": true, "login_page": "200", "css": {"/static/css/workout-manager.7007d84ce531.css": "200 2481 text/css", "/static/bootstrap-compiled.80a6279921f8.css": "200 277042 text/css", "/static/css/bootstrap-custom.400ad578123c.css": "200 1006 text/css"}, "photo_post": "201", "photo_bytes_posted": 179, "photo_path": "/media/gallery/1/f3391508-f895-4fc4-9688-7a42dc158932.png", "photo_read": "200 179 image/png", "control_unknown_static": "404 153 text/html", "watch_s": 600.1, "watch_requests": 10944, "watch_codes": {"200": 8208, "302": 2736}, "anon_max_bytes": 258273280, "anon_max_mib": 246.3, "anon_pct_of_384M": 64.1} +10:07:50 remove -> 200 +10:07:53 after remove: +vols: +stackdir: +total 28 +drwxr-xr-x 2 root root 4096 Oct 8 08:07 . +drwxr-xr-x 65 root root 4096 Oct 2 05:49 .. +-rw-r--r-- 1 root root 7619 Oct 8 07:50 .felhom.yml +-rw-r--r-- 1 root root 8672 Oct 8 07:50 docker-compose.yml +end + +10:07:53 stack: deployed=False state=not_deployed diff --git a/documentation/audits/day-2026-10-08/r762/9202/stdout.txt b/documentation/audits/day-2026-10-08/r762/9202/stdout.txt new file mode 100644 index 00000000..fcdbba14 --- /dev/null +++ b/documentation/audits/day-2026-10-08/r762/9202/stdout.txt @@ -0,0 +1,30 @@ +09:50:14 drill: bebbac8 DRILL wger: un-hidden + the R-762 gunicorn definition (DRILL only, 2026-10-08) +09:52:49 synced template on the box: 2 18:lifecycle: available +09:52:49 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['ADMIN_PASSWORD'] +09:52:50 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +09:54:55 [1] deployed, controller state=running, pinned={'wger': 'wger/server:2.7', 'wger-files': 'nginx:1.30.5-alpine'} +09:54:56 wger: POST /api/v2/weightentry/ http=201 +09:54:56 wger: readback of the seeded weight entry http=200 found=True +09:54:56 login page -> 200; CSS {'/static/css/workout-manager.7007d84ce531.css': '200 2481 text/css', '/static/bootstrap-compiled.80a6279921f8.css': '200 277042 text/css', '/static/css/bootstrap-custom.400ad578123c.css': '200 1006 text/css'} +09:54:57 POST /api/v2/gallery/ (179 B PNG) -> 201; read back /media/gallery/1/f3391508-f895-4fc4-9688-7a42dc158932.png -> 200 179 image/png; control -> 404 153 text/html +09:55:00 container: +cgroup=/sys/fs/cgroup/system.slice/docker-42589be792575b34221f0ff8ab0141f576e21db37a1ce1a4214e7e3385e1da46.scope +WGER_USE_GUNICORN=True +WEB_CONCURRENCY=2 +Using gunicorn on port 8000... +[2026-10-08 09:54:38 +0200] [58] [INFO] Starting gunicorn 26.1.0 +[2026-10-08 09:54:38 +0200] [58] [INFO] Listening at: http://0.0.0.0:8000 (58) +[2026-10-08 09:54:38 +0200] [59] [INFO] Booting worker with pid: 59 +[2026-10-08 09:54:38 +0200] [60] [INFO] Booting worker with pid: 60 +402653184 + +10:05:10 watch: +anon_max=258273280 +memory.max=402653184 +oom 0 +oom_kill 0 +restarts=0 oomkilled=false status=running health=healthy started=2026-10-08T07:53:22.480653193Z +restarts=0 oomkilled=false status=running health=healthy started=2026-10-08T07:53:22.678387772Z +2 + +10:05:13 RECORD {"app": "wger", "venue": "scratch guest 9202 on demo-hp, drill catalog, the product's deploy endpoint", "drill_commit": "bebbac8 DRILL wger: un-hidden + the R-762 gunicorn definition (DRILL only, 2026-10-08)", "pinned": {"wger": "wger/server:2.7", "wger-files": "nginx:1.30.5-alpine"}, "seed_read": true, "login_page": "200", "css": {"/static/css/workout-manager.7007d84ce531.css": "200 2481 text/css", "/static/bootstrap-compiled.80a6279921f8.css": "200 277042 text/css", "/static/css/bootstrap-custom.400ad578123c.css": "200 1006 text/css"}, "photo_post": "201", "photo_bytes_posted": 179, "photo_path": "/media/gallery/1/f3391508-f895-4fc4-9688-7a42dc158932.png", "photo_read": "200 179 image/png", "control_unknown_static": "404 153 text/html", "watch_s": 600.1, "watch_requests": 10944, "watch_codes": {"200": 8208, "302": 2736}, "anon_max_bytes": 258273280, "anon_max_mib": 246.3, "anon_pct_of_384M": 64.1} diff --git a/documentation/audits/day-2026-10-08/r762/9202/tools/box9202.py b/documentation/audits/day-2026-10-08/r762/9202/tools/box9202.py new file mode 100644 index 00000000..e63e58df --- /dev/null +++ b/documentation/audits/day-2026-10-08/r762/9202/tools/box9202.py @@ -0,0 +1,152 @@ +"""box9202.py — R-762 9202 half (2026-10-08 afternoon). ONE process, through the product. +phase install: drill commit (wger un-hidden + the gunicorn definition, DRILL only), sync, deploy, seed, login page, +CSS, photo, workers, env; then a ~10-minute watch (anon from memory.stat, oom_kill, restarts) with light load. +The removal is a separate phase (remove) so a failure leaves something to look at. +Evidence -> $EVD (no secrets: the generated admin password is never written).""" +import json, os, re, struct, subprocess, sys, tempfile, time, zlib, secrets +sys.path.insert(0, "/mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts") +import box_walk as w +import upgrade_fixtures_box as fixtures + +APP, SUB = "wger", "fitness" +EVD = os.environ["EVD"]; os.makedirs(EVD, exist_ok=True) +log = open(f"{EVD}/run.txt", "a", buffering=1) +D = "/mnt/5_hdd/felhom.eu/drill/app-catalog-drill" +NEWC = "/mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/templates/wger/docker-compose.yml" +REC = {"app": APP, "venue": "scratch guest 9202 on demo-hp, drill catalog, the product's deploy endpoint"} + + +def say(*a): + w.say(*a); log.write(time.strftime("%H:%M:%S ") + " ".join(map(str, a)) + "\n") + + +def git(*a): + return subprocess.run(["git", "-C", D, *a], check=True, capture_output=True, text=True).stdout + + +def png(path, n=64): + rgb = secrets.token_bytes(3) + raw = b"".join(b"\x00" + rgb * n for _ in range(n)) + def chunk(t, d): + return struct.pack(">I", len(d)) + t + d + struct.pack(">I", zlib.crc32(t + d) & 0xffffffff) + data = (b"\x89PNG\r\n\x1a\n" + chunk(b"IHDR", struct.pack(">IIBBBBB", n, n, 8, 2, 0, 0, 0)) + + chunk(b"IDAT", zlib.compress(raw)) + chunk(b"IEND", b"")) + open(path, "wb").write(data) + return len(data) + + +def fetch(path): + """code, bytes, content-type of one GET through the household's route (traefik, gate cookie). Body discarded.""" + gc = w.GATE.get(SUB) or w.gate_cookie(SUB) + args = ["curl", "-sSk", "--max-time", "45", "-H", f"Host: {SUB}.{w.DOMAIN}", "-o", "/dev/null", + "-w", "%{http_code} %{size_download} %{content_type}"] + if gc: + args += ["-H", f"Cookie: {gc}"] + r = w.sh(args + [f"{w.BASE}{path}"], timeout=90) + return (r.stdout or "").strip() + + +def css_links(): + rc, code, page = w.app_curl(SUB, "/en/user/login") + links = re.findall(r'(/static/[^"]+\.css)', page or "")[:3] + return code, {l: fetch(l) for l in links} + + +CG = 'ID=$(docker inspect -f "{{.Id}}" wger); CG=$(find /sys/fs/cgroup -maxdepth 6 -type d -name "*$ID*" | head -1)' + +phase = sys.argv[1] +w.login() +if phase == "install": + if w.stack(APP).get("deployed"): + sys.exit(say("STOP: wger already deployed on 9202 — not this run's") or 1) + git("pull", "-q", "--rebase", "origin", "main") + fy = f"{D}/templates/{APP}/.felhom.yml" + s = open(fy).read() + open(fy, "w").write(s.replace("lifecycle: hidden", "lifecycle: available", 1)) + open(f"{D}/templates/{APP}/docker-compose.yml", "w").write(open(NEWC).read()) + git("add", f"templates/{APP}/.felhom.yml", f"templates/{APP}/docker-compose.yml") + git("commit", "-q", "-m", "DRILL wger: un-hidden + the R-762 gunicorn definition (DRILL only, 2026-10-08)") + git("push", "-q", "origin", "main") + REC["drill_commit"] = git("log", "--oneline", "-1").strip() + say("drill:", REC["drill_commit"]) + for _ in range(12): + w.sync_rescan(); time.sleep(5) + stk = w.guest(f"grep -c WGER_USE_GUNICORN /opt/docker/stacks/{APP}/docker-compose.yml; grep -n lifecycle /opt/docker/stacks/{APP}/.felhom.yml").split() + if stk and stk[0] == "1" and "available" in " ".join(stk): + break + say("synced template on the box:", " ".join(stk)) + if not w.deploy(APP, SUB): + sys.exit(say("RESULT the install did not complete") or 1) + REC["pinned"] = (w.stack(APP).get("app_config") or {}).get("pinned_images") + fx = fixtures.FIXTURES[APP] + tok = fx.seed(w, SUB, say) + if tok is None: + sys.exit(say(f"RESULT seed failed: {getattr(fx, 'tried', '')}") or 1) + REC["seed_read"] = fx.verify(w, SUB, tok, say) + REC["login_page"], REC["css"] = css_links() + say(f"login page -> {REC['login_page']}; CSS {REC['css']}") + jar = tempfile.mktemp(prefix="wger-jar-") + hdr, why = fx._login(w, SUB, tok["pw"], jar) + pic = tempfile.mktemp(prefix="wger-photo-", suffix=".png") + size = png(pic) + rc, code, out = w.app_curl(SUB, "/api/v2/gallery/", *hdr, "-F", f"image=@{pic};type=image/png", "-F", + "date=2026-10-08", "-F", "description=r762", method="POST") + img = None + try: + img = json.loads(out).get("image") + except Exception: + say(f" gallery body {out[:200]}") + path = re.sub(r"^https?://[^/]+", "", img or "") + REC["photo_post"] = code; REC["photo_bytes_posted"] = size; REC["photo_path"] = path + REC["photo_read"] = fetch(path) if path else None + REC["control_unknown_static"] = fetch("/static/r762-nonexistent.css") + say(f"POST /api/v2/gallery/ ({size} B PNG) -> {code}; read back {path} -> {REC['photo_read']}; control -> {REC['control_unknown_static']}") + for f in (jar, pic): + if os.path.exists(f): + os.unlink(f) + g = w.guest(CG + ''' +echo "cgroup=$CG" +docker exec wger sh -c 'env | grep -E "^(WGER_USE_GUNICORN|WEB_CONCURRENCY)="' +docker logs wger 2>&1 | grep -E "Using gunicorn|Using django|Booting worker|Listening at|Starting gunicorn|Starting development server" +cat $CG/memory.max''') + say("container:\n" + g) + REC["container"] = g + # the watch: ~10 min, anon sampled every 2 s in the guest; light load from here (login page + CSS + photo) + w.guest(CG + ''' +rm -f /root/r762w.*; touch /root/r762w.on +nohup sh -c 'max=0; while [ -f /root/r762w.on ]; do a=$(awk "\\$1==\\"anon\\"{print \\$2}" '"$CG"'/memory.stat 2>/dev/null); [ -n "$a" ] && [ "$a" -gt "$max" ] && max=$a && echo $max > /root/r762w.max; sleep 2; done' >/dev/null 2>&1 & +echo started''') + t0 = time.time(); n = 0; codes = {} + while time.time() - t0 < 600: + for p in ("/en/user/login", list(REC["css"])[0] if REC["css"] else "/en/user/login", path or "/en/user/login"): + c = fetch(p).split(" ")[0]; codes[c] = codes.get(c, 0) + 1; n += 1 + rc, c, _ = w.app_curl(SUB, "/en/dashboard"); codes[c] = codes.get(c, 0) + 1; n += 1 + REC["watch_s"] = round(time.time() - t0, 1); REC["watch_requests"] = n; REC["watch_codes"] = codes + g = w.guest(CG + ''' +rm -f /root/r762w.on; sleep 3 +echo "anon_max=$(cat /root/r762w.max)" +echo "memory.max=$(cat $CG/memory.max)" +grep -E "^(oom|oom_kill) " $CG/memory.events +docker inspect -f "restarts={{.RestartCount}} oomkilled={{.State.OOMKilled}} status={{.State.Status}} health={{.State.Health.Status}} started={{.State.StartedAt}}" wger wger-files +docker logs wger 2>&1 | grep -c "Booting worker" +rm -f /root/r762w.*''') + say("watch:\n" + g) + REC["watch"] = g + m = re.search(r"anon_max=(\d+)", g) + if m: + a = int(m.group(1)); REC["anon_max_bytes"] = a + REC["anon_max_mib"] = round(a / 1048576, 1); REC["anon_pct_of_384M"] = round(100 * a / (384 * 1048576), 1) + logs = w.guest("docker logs wger 2>&1 | tail -n 300") + pw = tok.get("pw") or "" + if pw: + logs = logs.replace(pw, "") + open(f"{EVD}/wger.log", "w").write(logs) + json.dump(REC, open(f"{EVD}/box-verdict-wger.json", "w"), indent=2, ensure_ascii=False) + say("RECORD", json.dumps({k: REC[k] for k in REC if k not in ("container", "watch")}, ensure_ascii=False)) +elif phase == "remove": + say(f"remove -> {w.remove(APP)}") + g = w.guest("docker ps -a --format '{{.Names}}' | grep -i wger; echo vols:; docker volume ls --format '{{.Name}}' | grep -i wger; " + "echo stackdir:; ls -la /opt/docker/stacks/wger; echo end") + say("after remove:\n" + g) + st = w.stack(APP) + say(f"stack: deployed={st.get('deployed')} state={st.get('state')}") diff --git a/documentation/audits/day-2026-10-08/r762/9202/tools/repoint.py b/documentation/audits/day-2026-10-08/r762/9202/tools/repoint.py new file mode 100644 index 00000000..353da0e1 --- /dev/null +++ b/documentation/audits/day-2026-10-08/r762/9202/tools/repoint.py @@ -0,0 +1,37 @@ +#!/usr/bin/env python3 +"""Point 9202 at the drill catalog, or put the saved controller.yaml back. `09` §6.5.""" +import io, os, re, sys +sys.path.insert(0, "/mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts") +import box_walk as w +VOL = "/var/lib/docker/volumes/felhom-controller-data/_data" +SAVE = f"{VOL}/controller.yaml.pre-r762-1008" +DRILL = "https://gitea.dooplex.hu/admin/app-catalog-drill.git" +def creds(): + for l in io.open(os.path.expanduser("~/.git-credentials")).read().split("\n"): + m = re.match(r"https://(admin):([^@]+)@gitea\.dooplex\.hu", l) + if m: return m.group(1), m.group(2) + sys.exit("no admin credential") +if sys.argv[1] == "drill": + u, t = creds() + out = w.guest(f"""set -e +cp -p {VOL}/controller.yaml {SAVE} +python3 - <<'PY' +import re +p = "{VOL}/controller.yaml"; s = open(p).read() +s = re.sub(r'(^\\s+repo_url: ).*$', r'\\g<1>{DRILL}', s, count=1, flags=re.M) +s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+token: ).*$', r'\\g<1>"{t}"', s, count=1, flags=re.M) +s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+username: ).*$', r'\\g<1>"{u}"', s, count=1, flags=re.M) +open(p, "w").write(s) +PY +rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache +docker restart felhom-controller >/dev/null +grep -n 'repo_url' {VOL}/controller.yaml +""") + print(out.replace(t, "")) +elif sys.argv[1] == "restore": + print(w.guest(f"""set -e +cp -p {SAVE} {VOL}/controller.yaml +rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache +docker restart felhom-controller >/dev/null +grep -n 'repo_url' {VOL}/controller.yaml; grep -c 'token: ""' {VOL}/controller.yaml || true +cmp {SAVE} {VOL}/controller.yaml && echo RESTORED-IDENTICAL""")) diff --git a/documentation/audits/day-2026-10-08/r762/9202/wger.log b/documentation/audits/day-2026-10-08/r762/9202/wger.log new file mode 100644 index 00000000..2eebe6da --- /dev/null +++ b/documentation/audits/day-2026-10-08/r762/9202/wger.log @@ -0,0 +1,300 @@ + Applying allauth_idp_oidc.0001_initial... OK + Applying allauth_idp_oidc.0002_client_default_scopes... OK + Applying allauth_idp_oidc.0003_client_allow_uri_wildcards... OK + Applying contenttypes.0002_remove_content_type_name... OK + Applying auth.0002_alter_permission_name_max_length... OK + Applying auth.0003_alter_user_email_max_length... OK + Applying auth.0004_alter_user_username_opts... OK + Applying auth.0005_alter_user_last_login_null... OK + Applying auth.0006_require_contenttypes_0002... OK + Applying auth.0007_alter_validators_add_error_messages... OK + Applying auth.0008_alter_user_username_max_length... OK + Applying auth.0009_alter_user_last_name_max_length... OK + Applying auth.0010_alter_group_name_max_length... OK + Applying auth.0011_update_proxy_permissions... OK + Applying auth.0012_alter_user_first_name_max_length... OK + Applying authtoken.0001_initial... OK + Applying authtoken.0002_auto_20160226_1747... OK + Applying authtoken.0003_tokenproxy... OK + Applying authtoken.0004_alter_tokenproxy_options... OK + Applying axes.0001_initial... OK + Applying axes.0002_auto_20151217_2044... OK + Applying axes.0003_auto_20160322_0929... OK + Applying axes.0004_auto_20181024_1538... OK + Applying axes.0005_remove_accessattempt_trusted... OK + Applying axes.0006_remove_accesslog_trusted... OK + Applying axes.0007_alter_accessattempt_unique_together... OK + Applying axes.0008_accessfailurelog... OK + Applying axes.0009_add_session_hash... OK + Applying axes.0010_accessattemptexpiration... OK + Applying gym.0001_initial... OK + Applying core.0001_initial... OK + Applying config.0001_initial... OK + Applying config.0002_auto_20190618_1617... OK + Applying config.0003_delete_languageconfig... OK + Applying sessions.0001_initial... OK + Applying weight.0001_initial... OK + Applying weight.0002_auto_20150604_2139... OK + Applying weight.0003_auto_20160416_1030... OK + Applying weight.0004_multiple_weight_entries_per_day... OK + Applying weight.0005_add_uuid... OK + Applying nutrition.0001_initial... OK + Applying nutrition.0002_auto_20170101_1538... OK + Applying nutrition.0003_auto_20170118_2308... OK + Applying nutrition.0004_auto_20200819_2310... OK + Applying nutrition.0005_logitem... OK + Applying nutrition.0006_auto_20201201_0653... OK + Applying nutrition.0007_auto_20201214_0013... OK + Applying nutrition.0008_auto_20210102_1446... OK + Applying nutrition.0009_meal_name... OK + Applying nutrition.0010_logitem_meal... OK + Applying nutrition.0011_alter_logitem_datetime... OK + Applying nutrition.0012_alter_ingredient_license_author... OK + Applying gym.0002_auto_20151003_1944... OK + Applying gym.0003_auto_20151003_2008... OK + Applying gym.0004_auto_20151003_2357... OK + Applying gym.0005_auto_20151023_1522... OK + Applying gym.0006_auto_20160214_1013... OK + Applying gym.0007_auto_20170123_0920... OK + Applying gym.0008_auto_20190618_1617... OK + Applying exercises.0001_initial... OK + Applying manager.0001_initial... OK + Applying manager.0002_auto_20150202_2040... OK + Applying manager.0004_auto_20150609_1603... OK + Applying core.0002_auto_20141225_1512... OK + Applying core.0003_auto_20150217_1554... OK + Applying core.0004_auto_20150217_1914... OK + Applying core.0005_auto_20151025_2236... OK + Applying core.0006_auto_20151025_2237... OK + Applying core.0007_repetitionunit... OK + Applying core.0008_weightunit... OK + Applying core.0009_auto_20160303_2340... OK + Applying core.0010_auto_20170403_0144... OK + Applying core.0011_auto_20201201_0653... OK + Applying core.0012_auto_20210210_1228... OK + Applying core.0013_auto_20210726_1729... OK + Applying nutrition.0013_ingredient_image... OK + Applying nutrition.0014_license_information... OK + Applying nutrition.0015_alter_ingredient_creation_date_and_more... OK + Applying nutrition.0016_alter_logitem_options_and_more... OK + Applying nutrition.0017_remove_nutritionplan_language_alter_logitem_meal... OK + Applying nutrition.0018_nutritionplan_goal_carbs_nutritionplan_goal_energy_and_more... OK + Applying nutrition.0019_alter_image_license_author_and_more... OK + Applying nutrition.0020_full_text_search... OK + Applying nutrition.0021_add_fibers_field... OK + Applying nutrition.0022_add_remote_id_increase_author_field_length... OK + Applying nutrition.0023_fiber_spelling... OK + Applying nutrition.0024_remove_ingredient_status... OK + Applying nutrition.0025_add_last_image_check... OK + Applying nutrition.0026_add_start_and_end_fields... OK + Applying nutrition.0027_prefill_end_date... OK + Applying nutrition.0028_ingredient_dietary_properties... OK + Applying nutrition.0029_ingredient_nutriscore... OK + Applying manager.0005_auto_20160303_2008... OK + Applying manager.0006_auto_20160303_2138... OK + Applying manager.0007_auto_20160311_2258... OK + Applying manager.0008_auto_20190618_1617... OK + Applying manager.0009_auto_20201202_1559... OK + Applying manager.0010_auto_20210102_1446... OK + Applying manager.0011_remove_set_exercises... OK + Applying manager.0012_auto_20210430_1449... OK + Applying manager.0013_set_comment... OK + Applying manager.0014_auto_20210717_1858... OK + Applying manager.0015_auto_20211028_1113... OK + Applying exercises.0002_auto_20150307_1841... OK + Applying exercises.0003_auto_20160921_2000... OK + Applying exercises.0004_auto_20170404_0114... OK + Applying exercises.0005_auto_20190618_1617... OK + Applying exercises.0006_auto_20201203_0203... OK + Applying exercises.0007_auto_20201203_1042... OK + Applying exercises.0008_exercisebase... OK + Applying exercises.0009_auto_20201211_0139... OK + Applying exercises.0010_auto_20201211_0205... OK + Applying exercises.0011_auto_20201214_0033... OK + Applying exercises.0012_auto_20210327_1219... OK + Applying exercises.0013_auto_20210503_1232... OK + Applying exercises.0014_exerciseimage_style... OK + Applying exercises.0015_exercise_videos... OK + Applying exercises.0016_exercisealias... OK + Applying exercises.0017_muscle_name_en... OK + Applying core.0013_userprofile_email_verified... OK + Applying core.0014_merge_20210818_1735... OK + Applying exercises.0018_delete_pending_exercises... OK + Applying exercises.0019_exercise_crowdsourcing_changes... OK + Applying manager.0016_move_to_exercise_base... OK + Applying manager.0017_alter_workoutlog_exercise_base... OK + Applying exercises.0020_historicalexerciseimage_historicalexercisevideo... OK + Applying exercises.0021_deletionlog... OK + Applying exercises.0022_alter_exercise_license_author_and_more... OK + Applying exercises.0023_make_uuid_unique... OK + Applying exercises.0024_license_information... OK + Applying exercises.0025_rename_update_date_exercise_last_update_and_more... OK + Applying exercises.0026_deletionlog_replaced_by... OK + Applying exercises.0027_alter_deletionlog_replaced_by_and_more... OK + Applying exercises.0028_add_uuid_alias_and_comments... OK + Applying exercises.0029_full_text_search... OK + Applying exercises.0030_increase_author_field_length... OK + Applying exercises.0032_rename_exercise... OK + Applying core.0015_alter_language_short_name... OK + Applying core.0016_alter_language_short_name... OK + Applying manager.0018_flexible_routines... OK + Applying manager.0019_flexible_routines_migration... OK + Applying manager.0021_flexible_routines_cleanup... OK + Applying core.0017_language_full_name_en... OK + Applying core.0018_rounding... OK + Applying core.0019_delete_daysofweek... OK + Applying core.0020_add_trophies_enabled_to_userprofile... OK + Applying core.0021_add_unit_type_to_repetitionunit... OK + Applying nutrition.0030_add_indices... OK + Applying nutrition.0031_start_weight_unit_merge... OK + Applying nutrition.0032_continue_weight_unit_merge... OK + Applying nutrition.0033_finalize_weight_unit_merge... OK + Applying nutrition.0034_ingredient_trigram_gin_index... OK + Applying nutrition.0035_add_uuids... OK + Applying nutrition.0036_alter_image_license_author_and_more... OK + Applying nutrition.0037_powersync_synced_ingredient_tables... OK + Applying measurements.0001_initial... OK + Applying measurements.0002_auto_20210722_1042... OK + Applying measurements.0003_alter_measurement_unique_together_and_more... OK + Applying measurements.0004_add_uuids... OK + Applying measurements.0005_alter_measurement_date... OK + Applying trophies.0001_initial... OK + Applying trophies.0002_load_initial_trophies... OK + Applying manager.0022_alter_rir_type... OK + Applying manager.0023_change_validators... OK + Applying manager.0024_log_and_session_uuid... OK + Applying manager.0025_change_pk_to_uuid... OK + Applying trophies.0003_migrate_context_data_uuids... OK + Applying manager.0026_change_pk_to_uuid_swap... OK + Applying core.0022_move_email_verified_to_emailaddress... OK + Applying core.0023_create_publication... OK + Applying manager.0027_cleanup_fields... OK + Applying manager.0028_backfill_session_day... OK + Applying gallery.0001_initial... OK + Applying exercises.0033_uniqueness_constraint_translations... OK + Applying exercises.0034_add_exercise_image_dimensions... OK + Applying exercises.0035_add_is_ai_generated... OK + Applying exercises.0036_add_markdown_description_field... OK + Applying exercises.0037_replace_variation_with_uuid_field... OK + Applying exercises.0038_sync_model_changes... OK + Applying exercises.0039_translation_alias_trigram_gin_index... OK + Applying exercises.0040_alter_exercise_license_author_and_more... OK + Applying core.0024_backfill_emailaddress... OK + Applying core.0025_remove_unused_fields_in_userprofile... OK + Applying core.0026_alter_userprofile_birthdate_alter_userprofile_height... OK + Applying core.0027_powersync_publication... OK + Applying core.0028_longlivedsession... OK + Applying core.0029_userprofile_timezone... OK + Applying easy_thumbnails.0001_initial... OK + Applying easy_thumbnails.0002_thumbnaildimensions... OK + Applying mailer.0001_initial... OK + Applying mailer.0002_auto_20190618_1617... OK + Applying mailer.0003_auto_20201201_0653... OK + Applying manager.0029_alter_workoutsession_options_and_more... OK + Applying measurements.0006_health_sync... OK + Applying measurements.0007_migrate_weight... OK + Applying measurements.0008_dynamic_type... OK + Applying mfa.0001_initial... OK + Applying mfa.0002_authenticator_timestamps... OK + Applying mfa.0003_authenticator_type_uniq... OK + Applying sites.0001_initial... OK + Applying sites.0002_alter_domain_unique... OK + Applying socialaccount.0001_initial... OK + Applying socialaccount.0002_token_max_lengths... OK + Applying socialaccount.0003_extra_data_default_dict... OK + Applying socialaccount.0004_app_provider_id_settings... OK + Applying socialaccount.0005_socialtoken_nullable_app... OK + Applying socialaccount.0006_alter_socialaccount_extra_data... OK + Applying token_blacklist.0001_initial... OK + Applying token_blacklist.0002_outstandingtoken_jti_hex... OK + Applying token_blacklist.0003_auto_20171017_2007... OK + Applying token_blacklist.0004_auto_20171017_2013... OK + Applying token_blacklist.0005_remove_outstandingtoken_jti... OK + Applying token_blacklist.0006_auto_20171017_2113... OK + Applying token_blacklist.0007_auto_20171017_2214... OK + Applying token_blacklist.0008_migrate_to_bigautofield... OK + Applying token_blacklist.0010_fix_migrate_to_bigautofield... OK + Applying token_blacklist.0011_linearizes_history... OK + Applying token_blacklist.0012_alter_outstandingtoken_user... OK + Applying token_blacklist.0013_alter_blacklistedtoken_options_and_more... OK + Applying weight.0006_delete_weightentry... OK +*** Using settings from env: settings.main +Installed 1 object(s) from 1 fixture(s) +Installed 33 object(s) from 1 fixture(s) +Installed 7 object(s) from 1 fixture(s) +Installed 3 object(s) from 1 fixture(s) +Installed 5 object(s) from 1 fixture(s) +Installed 8 object(s) from 1 fixture(s) +Installed 6 object(s) from 1 fixture(s) +Installed 1 object(s) from 1 fixture(s) +Installed 12 object(s) from 1 fixture(s) +Installed 16 object(s) from 1 fixture(s) +Installed 8 object(s) from 1 fixture(s) +Installed 872 object(s) from 1 fixture(s) +Installed 2429 object(s) from 1 fixture(s) +Installed 1 object(s) from 1 fixture(s) +Installed 1 object(s) from 1 fixture(s) +Installed 1 object(s) from 1 fixture(s) +*** Using settings from env: settings.main +*** Password for user admin was reset to '' +Installed 3 object(s) from 1 fixture(s) +Running in production mode, running collectstatic now +level=INFO ts=2026-10-08 09:54:12,682 module=apps path=/home/wger/.local/lib/python3.12/site-packages/axes/apps.py line=53 message=AXES: BEGIN version 8.3.1, blocking by username + +11362 static files copied to '/home/wger/static', 11362 post-processed. +Performing database migrations +level=INFO ts=2026-10-08 09:54:32,571 module=apps path=/home/wger/.local/lib/python3.12/site-packages/axes/apps.py line=53 message=AXES: BEGIN version 8.3.1, blocking by username +System check identified some issues: + +WARNINGS: +?: (axes.W006) AXES_LOCKOUT_PARAMETERS does not contain 'ip_address'. This configuration allows attackers to bypass rate limits by rotating User-Agents or Cookies. + HINT: Add 'ip_address' to AXES_LOCKOUT_PARAMETERS. +Operations to perform: + Apply all migrations: account, actstream, allauth_idp_oidc, auth, authtoken, axes, config, contenttypes, core, easy_thumbnails, exercises, gallery, gym, mailer, manager, measurements, mfa, nutrition, sessions, sites, socialaccount, token_blacklist, trophies, weight +Running migrations: + No migrations to apply. + Your models in app(s): 'exercises', 'gallery' have changes that are not yet reflected in a migration, and so won't be applied. + Run 'manage.py makemigrations' to make new migrations, and then re-run 'manage.py migrate' to apply them. +level=INFO ts=2026-10-08 09:54:35,564 module=apps path=/home/wger/.local/lib/python3.12/site-packages/axes/apps.py line=53 message=AXES: BEGIN version 8.3.1, blocking by username +System check identified some issues: + +WARNINGS: +?: (axes.W006) AXES_LOCKOUT_PARAMETERS does not contain 'ip_address'. This configuration allows attackers to bypass rate limits by rotating User-Agents or Cookies. + HINT: Add 'ip_address' to AXES_LOCKOUT_PARAMETERS. +Set site URL to fitness.enkisfelhom.hu +Using gunicorn on port 8000... +level=INFO ts=2026-10-08 09:54:38,089 module=apps path=/home/wger/.local/lib/python3.12/site-packages/axes/apps.py line=53 message=AXES: BEGIN version 8.3.1, blocking by username +[2026-10-08 09:54:38 +0200] [58] [INFO] Starting gunicorn 26.1.0 +[2026-10-08 09:54:38 +0200] [58] [INFO] Listening at: http://0.0.0.0:8000 (58) +[2026-10-08 09:54:38 +0200] [58] [INFO] Using worker: sync +[2026-10-08 09:54:38 +0200] [59] [INFO] Booting worker with pid: 59 +[2026-10-08 09:54:38 +0200] [60] [INFO] Booting worker with pid: 60 +[2026-10-08 09:54:38 +0200] [58] [INFO] Control socket listening at /home/wger/.gunicorn/gunicorn.ctl +level=WARNING ts=2026-10-08 09:54:53,158 module=wsgi path=/home/wger/.local/lib/python3.12/site-packages/gunicorn/http/wsgi.py line=450 message=WSGI app sent body bytes on a no-body response (method=HEAD status=200); dropping per RFC 9110. +level=INFO ts=2026-10-08 09:54:56,132 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=295 message=AXES: Successful login by {username: "********************", ip_address: "********************", user_agent: "curl/8.14.1", path_info: "/en/user/login"}. +level=INFO ts=2026-10-08 09:54:56,134 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=425 message=AXES: Cleaned up 0 expired access attempts from database that were older than 2026-10-08 07:49:56.005561+00:00 +level=WARNING ts=2026-10-08 09:54:56,295 module=log path=/home/wger/.local/lib/python3.12/site-packages/django/utils/log.py line=249 message=Forbidden: /api/v2/weightentry/ +level=INFO ts=2026-10-08 09:54:56,491 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=295 message=AXES: Successful login by {username: "********************", ip_address: "********************", user_agent: "curl/8.14.1", path_info: "/en/user/login"}. +level=INFO ts=2026-10-08 09:54:56,492 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=425 message=AXES: Cleaned up 0 expired access attempts from database that were older than 2026-10-08 07:49:56.371620+00:00 +level=INFO ts=2026-10-08 09:54:57,028 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=295 message=AXES: Successful login by {username: "********************", ip_address: "********************", user_agent: "curl/8.14.1", path_info: "/en/user/login"}. +level=INFO ts=2026-10-08 09:54:57,029 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=425 message=AXES: Cleaned up 0 expired access attempts from database that were older than 2026-10-08 07:49:56.907310+00:00 +level=WARNING ts=2026-10-08 09:55:23,254 module=wsgi path=/home/wger/.local/lib/python3.12/site-packages/gunicorn/http/wsgi.py line=450 message=WSGI app sent body bytes on a no-body response (method=HEAD status=200); dropping per RFC 9110. +level=WARNING ts=2026-10-08 09:55:53,725 module=wsgi path=/home/wger/.local/lib/python3.12/site-packages/gunicorn/http/wsgi.py line=450 message=WSGI app sent body bytes on a no-body response (method=HEAD status=200); dropping per RFC 9110. +level=WARNING ts=2026-10-08 09:56:23,813 module=wsgi path=/home/wger/.local/lib/python3.12/site-packages/gunicorn/http/wsgi.py line=450 message=WSGI app sent body bytes on a no-body response (method=HEAD status=200); dropping per RFC 9110. +level=WARNING ts=2026-10-08 09:56:53,908 module=wsgi path=/home/wger/.local/lib/python3.12/site-packages/gunicorn/http/wsgi.py line=450 message=WSGI app sent body bytes on a no-body response (method=HEAD status=200); dropping per RFC 9110. +level=WARNING ts=2026-10-08 09:57:24,005 module=wsgi path=/home/wger/.local/lib/python3.12/site-packages/gunicorn/http/wsgi.py line=450 message=WSGI app sent body bytes on a no-body response (method=HEAD status=200); dropping per RFC 9110. +level=WARNING ts=2026-10-08 09:57:54,097 module=wsgi path=/home/wger/.local/lib/python3.12/site-packages/gunicorn/http/wsgi.py line=450 message=WSGI app sent body bytes on a no-body response (method=HEAD status=200); dropping per RFC 9110. +level=WARNING ts=2026-10-08 09:58:24,191 module=wsgi path=/home/wger/.local/lib/python3.12/site-packages/gunicorn/http/wsgi.py line=450 message=WSGI app sent body bytes on a no-body response (method=HEAD status=200); dropping per RFC 9110. +level=WARNING ts=2026-10-08 09:58:54,279 module=wsgi path=/home/wger/.local/lib/python3.12/site-packages/gunicorn/http/wsgi.py line=450 message=WSGI app sent body bytes on a no-body response (method=HEAD status=200); dropping per RFC 9110. +level=WARNING ts=2026-10-08 09:59:24,378 module=wsgi path=/home/wger/.local/lib/python3.12/site-packages/gunicorn/http/wsgi.py line=450 message=WSGI app sent body bytes on a no-body response (method=HEAD status=200); dropping per RFC 9110. +level=WARNING ts=2026-10-08 09:59:54,478 module=wsgi path=/home/wger/.local/lib/python3.12/site-packages/gunicorn/http/wsgi.py line=450 message=WSGI app sent body bytes on a no-body response (method=HEAD status=200); dropping per RFC 9110. +level=WARNING ts=2026-10-08 10:00:24,571 module=wsgi path=/home/wger/.local/lib/python3.12/site-packages/gunicorn/http/wsgi.py line=450 message=WSGI app sent body bytes on a no-body response (method=HEAD status=200); dropping per RFC 9110. +level=WARNING ts=2026-10-08 10:00:54,667 module=wsgi path=/home/wger/.local/lib/python3.12/site-packages/gunicorn/http/wsgi.py line=450 message=WSGI app sent body bytes on a no-body response (method=HEAD status=200); dropping per RFC 9110. +level=WARNING ts=2026-10-08 10:01:24,758 module=wsgi path=/home/wger/.local/lib/python3.12/site-packages/gunicorn/http/wsgi.py line=450 message=WSGI app sent body bytes on a no-body response (method=HEAD status=200); dropping per RFC 9110. +level=WARNING ts=2026-10-08 10:01:54,852 module=wsgi path=/home/wger/.local/lib/python3.12/site-packages/gunicorn/http/wsgi.py line=450 message=WSGI app sent body bytes on a no-body response (method=HEAD status=200); dropping per RFC 9110. +level=WARNING ts=2026-10-08 10:02:24,938 module=wsgi path=/home/wger/.local/lib/python3.12/site-packages/gunicorn/http/wsgi.py line=450 message=WSGI app sent body bytes on a no-body response (method=HEAD status=200); dropping per RFC 9110. +level=WARNING ts=2026-10-08 10:02:55,018 module=wsgi path=/home/wger/.local/lib/python3.12/site-packages/gunicorn/http/wsgi.py line=450 message=WSGI app sent body bytes on a no-body response (method=HEAD status=200); dropping per RFC 9110. +level=WARNING ts=2026-10-08 10:03:25,106 module=wsgi path=/home/wger/.local/lib/python3.12/site-packages/gunicorn/http/wsgi.py line=450 message=WSGI app sent body bytes on a no-body response (method=HEAD status=200); dropping per RFC 9110. +level=WARNING ts=2026-10-08 10:03:55,203 module=wsgi path=/home/wger/.local/lib/python3.12/site-packages/gunicorn/http/wsgi.py line=450 message=WSGI app sent body bytes on a no-body response (method=HEAD status=200); dropping per RFC 9110. +level=WARNING ts=2026-10-08 10:04:25,294 module=wsgi path=/home/wger/.local/lib/python3.12/site-packages/gunicorn/http/wsgi.py line=450 message=WSGI app sent body bytes on a no-body response (method=HEAD status=200); dropping per RFC 9110. +level=WARNING ts=2026-10-08 10:04:55,381 module=wsgi path=/home/wger/.local/lib/python3.12/site-packages/gunicorn/http/wsgi.py line=450 message=WSGI app sent body bytes on a no-body response (method=HEAD status=200); dropping per RFC 9110. diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index e998bb22..a5ca0a75 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -32,6 +32,8 @@ The full text of every row below: `git show b2dce901b2:documentation/backlog/OPE | Row | What | Closed | Evidence | |---|---|---|---| +| **R-415** | **The hub enforced uniqueness on `customer_id` only: two customers could be given the same or a nested domain.** (P3) | CLOSED 2026-10-08 — FIXED on hub main (ships with the next hub release): create and edit refuse a domain equal to, containing or under another customer's, or under felhom.eu | `hub/internal/store/domain_conflict.go`, `TestR415_CreateAndEditRefuseConflictingDomain` (red-proved); re-filed the same day after the 2026-10-03 triage lost it; full text `git show 78121aa475:documentation/backlog/OPEN-ITEMS.md`. | +| **R-762** | **wger served no CSS, no JavaScript and no uploaded photo; and ran Django's development server.** (P3) | CLOSED 2026-10-08 — both halves proven and pushed: the files (catalog `cf1ed43`, 2026-10-06) and gunicorn with 2 workers (catalog `eec9a0d`, CI job 1546 success; operator ruling 3, decision 182, for the drill-catalog write) | Bench 9401: anon 170 MiB = 44 % of 384M, 0 kills, 0 restarts over 606 s; 9202 (drill catalog, fresh install through the product): 2 × „Booting worker", anon 246 MiB = 64 %, 0 kills, 0 restarts over 600 s / 10,944 requests, login 200, 3 CSS 200, photo 201 → 200, unknown static 404 (control). No ladder step applies (images unchanged; `ladder.check_entry` refuses a same-digest re-test; `09` §5.4 renders the catalog template). Evidence `audits/day-2026-10-08/r762/`. Follow-ups: R-905 (static files in backups); `mem_request` decision D10. wger stays `lifecycle: hidden`. | | **R-177** | **There was no operator-triggerable „run the fill check now" path.** (P4) | CLOSED 2026-10-08 — NO LONGER TRUE: since controller v0.297.0 (R-363) the fill check also runs every 10 minutes, so no door is needed for it | `felhom-controller/controller/cmd/controller/main.go:1579` (`sched.Every("fill-watch-interval", …)`), `:2478` (`fillWatchInterval = 10 * time.Minute`), pinned by `cmd/controller/r363_fillwatch_interval_test.go`; each run logs „checked N filesystem(s)" (readable through the hub's controller-log pull). Design `audits/day-2026-10-08/design-R-314-279-177.md`. | | **R-298** | **The `/storage` page's unregistered list filtered on `role==='user-data'`, so a drive that is also the backup target could never be registered.** (P3) | CLOSED 2026-10-08 — NOT REPRODUCIBLE / NOT TRUE: the agent's role comes from the storage type and backing disk and never from being the backup target, so a backup-target drive on a non-system disk IS user-data | `felhom-agent/internal/storage/role.go:172-186` (unchanged since 2026-07-13); `internal/localapi/disks.go:212-233` and `:1239-1249`; August topology `audits/evidence-rehearsal-2026-08-09/GATE0-demo-hp-before.txt:47-84` → user-data; read-only today: demo-felhom `felhom-backup` on `sdb` (root on `sda`). Side fix on controller main `a40729a`: no eject/format button on a backup-target drive (the agent refuses the eject, 403). | | **R-900** | **The website's FAQ said the household's data is not with a third party, while copies and traffic go to processors.** (P2) | CLOSED 2026-10-08 — PUBLISHED: the GDPR answer on `gyik.html` and `en/faq.html` (visible text and structured data) now names the box, the encrypted copies at Hetzner in the EU, and Cloudflare's view of remote-access traffic; text approved by the operator in chat (`09` §3 decision 180) | website commit `b2dce901`; live read-back 2026-10-08: the new sentence 2× on each page (visible + JSON-LD), the old „nem harmadik félnél" 0×; rest of the site searched (ASCII fragments, positive control) — no other page makes the claim. Left for R-813: the contact form's consent line „harmadik félnek nem adjuk ki", which the consent draft already replaces. | diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 4b3656cf..66ed8030 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -128,13 +128,13 @@ stopping line that lies. |---|---|---|---|---|---|---|---| | **R-562** | Apps & catalog | P3 | **[P3-LOW] Dates and sizes are not formatted for any locale — and the Hungarian pages disagree with themselves.** FOUND 2026-09-17 by the i18n inventory §2.8: the two template date layouts differ (`2006. 01. 02. 15:04` Hungarian vs `2006-01-02 15:04` ISO); 10 layout literals in `internal/web` Go and 25 elsewhere pick formats ad hoc; sizes print a decimal POINT (`%.1f GB`, 4 helpers) where Hungarian uses a comma; `timeAgo`/`nextRunLabel`/`pruneLabel` produce Hungarian words outside the three converted pages. Not changed by v0.247.0 (Hungarian bytes are frozen by the parity rule). **Fix shape:** one date and one size formatter per language in `internal/i18n`, the Hungarian output deliberately changed in ONE reviewed release with the parity fixtures re-captured for that release only and the change named in its CHANGELOG. Needs an operator word on the Hungarian format (comma, date style). | **READY - rank P3-LOW; owner: CC** **2026-10-06 night: not started — the row needs the operator's word on the Hungarian format (decimal comma, date style) before any code.** | — | — | CC | | **R-676** | Apps & catalog | P3 | **[P3-LOW] Watch: immich's first start restarted 12 times — decision 28's crash-loop stop (6 in 10 min) would stop it.** From the 2026-09-17 chaos night (DB connection dropped during the first-start geocoding import on a 6 GB guest; it did not recover that night). No healthy app in any drill evidence restarts on a first start (1831 samples, 40 live containers), so the threshold stands; this row exists so the first immich install under v0.269.x is watched. `audits/night-2026-09-24/A3/40-first-start-restarts.txt` **2026-09-25 night (read from source, v0.271.0): a DEPLOY's first start is NOT covered by decision 28's suppression** — `Deploying` clears when `compose up -d` returns (`deploy.go` "Clear deploying flag"), and `ObserveUnhealthy` then samples the app; an automatic update's step, verify and undo ARE covered (`Updating`, pinned by `TestD28_NoCrashLoopStopDuringAnAutomaticStep`). So a first start that restarts ≥ 6 times in 10 min is stopped — which R-676 already accepts for a broken first start; a healthy slow first start would be stopped too. **-- 2026-09-30: the first-start restarts are explained.** immich's first-start geodata import OOM-kills its database at 512M on a guest with no swap (R-732, measured: 61–104 kills); the 2026-09-17 chaos-night case (DB connection dropped during the import on a 6 GB guest) fits it. Fixed in the catalog (`56c4888`, 768M). The watch itself (decision 28 on a DEPLOY's first start) is unchanged. | **OPEN — P3; owner: CC (watch)** | — | — | CC | -| **R-762** | Apps & catalog | P3 | **[P2-MEDIUM] wger serves no CSS or JavaScript and no uploaded photo: every static file and every `/media/` file answers 404.** MEASURED 2026-10-01 on 9202 (drill catalog, the live template `82fff32`, wger 2.7), found by checklist rows 1.7 and 2.8: the login page links `/static/css/workout-manager.css`, `/static/bootstrap-compiled.css` — both 404 through traefik; the static root inside the container is empty (4 KB). A progress photo posted to `/api/v2/gallery/` answered 201 and the file is on the media volume, but `GET /media/gallery/…png` answers 404 signed in, without a session, and straight at the app inside the container. **Cause, read in the image:** the entrypoint runs `collectstatic` only when `DJANGO_DEBUG == "False"` and the template sets no `DJANGO_DEBUG`; and wger serves `/media/` only in development (`urls.py:393` „served like this during development only”) — upstream's production setup puts nginx in front for `/static` and `/media`. So the household gets an unstyled app and photos that never show. The same lines stand since the template was written (the 2026-09-29 template too). Not checked: whether any box runs wger (on 2026-09-30 none reported to the hub). **Needs:** `DJANGO_DEBUG=False` (collectstatic) and something that serves `/static` + `/media` (upstream's nginx sidecar, or the gunicorn switch of R-755 plus a static server), proven on the bench and on 9202 with a page that loads its CSS and a photo read back. Owner decides together with R-755 (same server question). `audits/new-app-checklist-2026-10-01/C/C8-signup-guest-media-static.txt`, `C/C4-seed-photo-size.txt` **-- 2026-10-01 (operator):** wger is `lifecycle: hidden` until this and its twin are fixed (catalog `55b8c8a`; read back on 9202: not on the app list, mealie control present). **Merged 2026-10-05 from R-755 (duplicate):** `templates/wger/docker-compose.yml` still sets no `WGER_USE_GUNICORN` — the gunicorn switch is the same server question. **-- 2026-10-06 (afternoon), measured again on 9202 (live template, wger 2.7):** `/static/css/workout-manager.css` 404 straight at the app, `/home/wger/static` 4 KB, settings `DEBUG False`; the image has gunicorn but no whitenoise and runs Django's `runserver` (no `WGER_USE_GUNICORN`, R-755). So `DJANGO_DEBUG=False` alone would collect the files and still serve none: the fix needs a server for `/static` + `/media` (a second container) — medium, not taken. `audits/r890-instructions-2026-10-06/C/wger.txt`. **-- 2026-10-08 (day):** the gunicorn half PROVEN ON THE BENCH (9401, `wger/server:2.7`, limit 384M): `WGER_USE_GUNICORN=True` + `WEB_CONCURRENCY=2` (the image runs gunicorn only on that switch and passes no `-w`; the container log shows 2 × „Booting worker"); anon peak 166–170 MiB = 43–44 %, oom_kill 0, restarts 0 over a 606 s watch (11,584 requests; harness verdict `proven`); login 200, all three CSS links 200 via `wger-files`, a photo uploaded 201 and read back 200; an unknown `/static/` file 404 (control). **Not done: 9202** — wger is hidden, the live catalog refuses a hidden deploy, and the drill-catalog write that un-hides it there was refused by the permission check (stop, per the brief). **Nothing pushed**; the tested definition is in the evidence. **No ladder step applies:** the change moves no image (from == to), the ladder writer refuses a same-digest re-test, and `09` §5.4 renders the catalog template when the images equal the pinned ones, so the compose change reaches an installed wger at its next `up -d`. Still true: the collected style files (283 MB) go into every wger backup. `audits/day-2026-10-08/r762/`. | **READY — narrowed 2026-10-08: the bench proof is done; the 9202 proof is ALLOWED since the operator's ruling 2026-10-08 09:04 (`09` §3 decision 182: CC may write the shared test catalog).** wger stays `lifecycle: hidden`. | — | Operator allows the drill-catalog write (fast-forward the drill to live main, un-hide wger there, point 9202 at it) → prove on 9202 (login + CSS + photo 200, memory, 10-min watch) → push the two env lines (no ladder entry) → close | CC | | **R-76** | Apps & catalog | P4 | **FileBrowser-created folders break the setgid chain, and a drop-zone's mode is not stable** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-26, size S, roadmap state `idea (surfaced by the R-75 spike, 2026-07-26)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-06 night: not a catalog fix** — the setgid chain is set by the controller's FileBrowser setup; a controller row. | — | Two related findings from `audits/SPIKE-catalog-data-paths-2026-07-26.md` P3/P5, both **pre-existing** and deliberately left alone by that spike. **(a)** FileBrowser Quantum 1.3.3 creates files `0644` and folders `0755` and does **not** propagate the setgid bit — even though the entrypoint wrapper's `umask 002` really is in effect (`/proc/1/status` `Umask: 0002`). Group inheritance itself works (a file uploaded into a 2775 group-100 dir landed group 100, not the process gid 1000), so the convention's *group* half holds and only its *mode* half is lost. The consequence is proven with a control: inside a UI-created `0755` folder a gid-1000 process's file landed group **1000**, while the identical write into the 2775 parent landed group **100**. So **any folder a customer creates through FileBrowser breaks the shared-group chain one level down.** Latent today — every userdata-touching catalog app that declares an identity declares uid/gid **1000**, the same uid FileBrowser runs as, so owner permissions mask it; it bites the day a content app runs as a different non-root uid with gid 1000. The comment at `infra/infra.go:156` is right that the image ignores `-e UMASK` but does not say t | CC | | **R-577** | Apps & catalog | P4 | **[P3-LOW] A guest SHARE visitor has no way to pick a language, and the household's setting is the wrong default for them.** FOUND 2026-09-18 by localisation slice 2 release C (R-557, controller v0.254.0): every other page a person can reach now carries a language globe — the dashboard (the household's setting), and the sign-in and claim pages (the visitor's own cookie). The two guest share pages (`launcher_shared`, `launcher_share_password`) deliberately do NOT, and `TestGuestSharePagesHaveNoGlobe` pins that so it stays a decision rather than an oversight. **Why it is the operator's and not CC's:** a share visitor is a stranger the household sent a link to, and what language they are shown is a promise the SHARE FEATURE makes, not an implementation detail. The `felhom_lang` cookie already built would fit them exactly (display-only, their own browser, never the household's setting). **Fix shape, if the operator says yes:** add `{{template "lang_globe" .}}` to both shells with the anonymous form, and one render case per page per language. | **READY - rank P3-LOW; owner: operator (the decision), CC (the change)** **Re-ranked 2026-10-03: P3->P4: feature decision for the operator; Hungarian default works today.** | — | — | operator | | **R-707** | Apps & catalog | P4 | **[P2] 37 apps still start with a login a stranger can take (`09` §3 decision 45).** Audit of all 53 apps: `app-catalog-felhom.eu/FIRST-ADMIN.md` (class, fix route, status, measured or read). Open: **3 hard-coded defaults** — calibre-web (`admin / admin123`, measured working on demo-hp and 9202; its own `cps.py -s` route needs a generated password WITH a special character — our generator is letters+digits, a controller change), mealie (`changeme@example.com / MyPassword`), wger (`admin / adminadmin`); **34 open first-run screens** (the first visitor creates the admin: actualbudget, adventurelog, audiobookshelf, calcom, docmost, emby, ghost, gitea, gramps-web, home-assistant, homebox, immich, jellyfin, komga, n8n, navidrome, opengist, outline, papra, plant-it, radarr, rallly, recipe-importer, romm, seerr, sonarr, sparkyfitness, tandoor, termix, uptime-kuma, vikunja, wanderer, wishlist, zipline). **Stale notes:** romm's `default_creds` `admin / admin` answers 401 on demo-hp (like a wrong password) — the page now warns with a login that does not exist; zipline's looks stale too. **Measured on demo-hp 2026-09-28 (read-only):** bookstack's default still logs in on the INSTALLED app (the fix is for new installs; the page now warns). Each fix: route (a) env or (b) the app's own CLI/API via `after_install:`, proven on 9202 with the default failing and the generated password working; route (c) a page sentence. Several sessions (operator, 2026-09-28). **2026-09-29 (controller v0.280.0, catalog `d0e7e2e`):** every class-3 app fixed — mealie, wger, calibre-web by `after_install` (calibre-web with the new `password:24:special`), proven on 9202 fresh installs (`audits/login-gate-2026-09-29/D/`); the setup gate (decision 46, spike PASSED) built and live on immich, n8n, audiobookshelf (probes measured) and uptime-kuma (button) (`…/C/`); romm's and zipline's stale notes removed. **Left: 30 class-4 apps** — gate each (probe measured on 9202 where one exists — 11 upstream candidates listed in `…/B/B-VERDICT.md` §3; the button otherwise). **2026-09-29 afternoon (controller v0.281.0, catalog `6faf432`):** 28 more class-4 apps gated — 32 of 34 — each proven on 9202 (`audits/gate-rollout-2026-09-29/`B): stranger → gate page / 401, household reached the first-setup screen, the gate opened (9 by a measured probe, the rest by the press), the app answered after. seerr, outline, rallly: gated, their opening needs a media server / e-mail (not proven). **Left:** wanderer (R-714); plant-it is not installable. | **NARROWED** (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — seerr, outline and rallly are gated, but the gate OPENING is not proven (needs a media server / e-mail)) — **CLOSED — 2026-09-29 (the rest → R-714)** | — | — | CC | | **R-769** | Apps & catalog | P4 | **[P3-LOW] Pinchflat is not built: upstream is paused (no release in 2026, last push 2025-12-16) and its last release has no image tag.** READ 2026-10-01: ghcr's newest version tag `v2025.6.6`, `latest` amd64 only; runs as root by default; an unanswered 30 GB yt-dlp memory report (#866); third parties call it unmaintained (community-scripts #15968). Forks with images exist (Pinchflat-NGX, MorganKryze). **Needs:** the operator's word on a fork (a new upstream), after the YouTube sentence (R-767). `audits/new-apps-2026-10-01/FIT.md` | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator** **Re-ranked 2026-10-03: P3→P4: a new-app idea waiting on a decision.** | — | — | operator | | **R-770** | Apps & catalog | P4 | **[P3-LOW] Invidious — fit check only; the recommendation is not to build it.** READ 2026-10-01: playback needs `invidious-companion` (rolling `latest`, no version tags); PostgreSQL 14 (EOL 2026-11); `registration_enabled: true` by default; upstream: a bot check means „your IP is blocked from YouTube”, a 429 can last 24 h, triggered by „someone on your network” — on our boxes that IP is the household's. One bad period in 2026 (March, ~2 weeks). No report found of a family's other devices being bot-checked (inference). **Needs:** the operator's go / no-go. `audits/new-apps-2026-10-01/FIT.md` | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator** **Re-ranked 2026-10-03: P3→P4: a new-app idea waiting on a decision.** | — | — | operator | | **R-771** | Apps & catalog | P4 | **[P3-LOW] moonlight-web — fit check only; not buildable through an HTTP-only tunnel at usable latency.** READ 2026-10-01: two unrelated projects (MrCreativ3001/moonlight-web-stream, the original; linckosz/moonlight-web); both need Sunshine/Apollo/Wolf on a gaming PC on the LAN and WebRTC over UDP (40000-40100/udp; linckosz recommends host networking and sends telemetry by default); both have a WebSocket fallback (high latency, all video through the tunnel); a logged-in user controls the PC's desktop. **Needs:** the operator's go / no-go (LAN-only use would need a different publishing model). `audits/new-apps-2026-10-01/FIT.md` | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator** **Re-ranked 2026-10-03: P3→P4: a new-app idea waiting on a decision.** | — | — | operator | +| **R-905** | Apps & catalog | P4 | **wger's collected style files (283 MB) go into every wger backup, though the app rebuilds them at every start.** Since R-762's fix (`DJANGO_DEBUG=False` + `collectstatic`, catalog `cf1ed43`) the static files sit in a named volume that the backup legs copy like data (measured 283 MB, `audits/design-build-2026-10-06/`); `collectstatic` regenerates them, so they are rebuildable bytes in every Tier-1 unit, Tier-2 copy and off-site snapshot. wger is `lifecycle: hidden` and no box runs it (hub read 2026-10-08). Options to weigh: a non-backed-up volume class for regenerable data, or an anonymous volume. Found by the R-762 helpers 2026-10-08. | **OPEN — needs a design before wger is shown** | — | Design the „regenerable volume" exclusion (catalog + controller backup legs) | CC | ## App updates — 4 rows (P3 4) @@ -197,7 +197,7 @@ stopping line that lies. | **R-861** | Security & access | P2 | **The agent's sudoers lets the agent user reach root without the operator key, so "root-minimized" (`03` §3) overstates it and the root-owned trust files (decision 93, the bundle's R17) are defence in depth, not a boundary.** READ 2026-10-04 from `felhom-agent/configs/felhom-agent.sudoers` (not exploited): `FELHOM_GUESTHOOK` installs `/tmp/felhom-guest-hook-*.sh` as a hookscript Proxmox runs as root at guest start, and `pct reboot` is granted; `FELHOM_INTERMEDIARY` installs a script + a systemd unit that run as root at boot; `FELHOM_ESCROW` runs `/usr/local/bin/felhom-agent` as root, and `FELHOM_SELFUPDATE apply` accepts a sha the agent itself passes. A compromised agent PROCESS is therefore root on its host. Fix direction: each of the four becomes a root-owned wrapper that checks its own input (fixed content or a signature), like `felhom-os-apply`; delivered by the config bundle. `11` §5.4.2, `03` §11. | **NARROWED 2026-10-05 — FIXED agent v0.146.1 for every root path found (nine, not four), delivered to demo-hp, demo-felhom and Tester 1 by a step bundle (R-880); live on both demo boxes: `sudo -l` 93/93 (64 commands allowed, 29 attacks refused — 23 of them allowed before), capability probe 67/67, a staged unit over /etc/sudoers.d refused. Design `03` §3.1, decision 122. LEFT, each named there: (a) the controller-swap image ref is guest-scoped (a compromised agent can run a chosen pinned-registry image in the guest); (b) the felhom-op SSH key is hub-delivered, not signed (felhom-op's sudo is scoped, not root); (c) the escrow ceremony hands the agent R by design (the box's PBS key). Tester 2: not delivered (offline).** **2026-10-06 night: design written** (`audits/night-burndown-2026-10-06/design-R-861.md`). Correction: (a) is not "pinned-registry" — the `tee` content is unchecked by sudo, so any image from any registry runs in the guest with the docker socket (`03` §3.1 corrected). Pick: (a) close before the first paying customer (a `felhom-priv-apply controller-image` verb; ~1–2 h, rides the bundle); (b) and (c) accept for the first customers. Waits for the operator. **2026-10-07 07:58: `09` §3 decision 165 — (a) A1 yes before the first paying customer; (b) B3 accept + B2 hygiene in the same bundle; (c) C2 accept.** **2026-10-07: (a) A1 and (b) B2 DELIVERED (agent 0.151.0 + bundle on demo-hp, demo-felhom, Tester 1; probe 68/68).** Live on demo-hp: no `tee` grant left in `sudo -l -U felhom-agent`; the verb `felhom-priv-apply ^controller-image [0-9]+$` is the route; a hand-fed `docker.io/library/alpine:latest` → `REFUSED [I1]` rc 3, the guest's image file unchanged; the old `pct exec … tee` asks for a password; felhom-op's pct lines anchored (`audits/day-2026-10-07/C/C-live-demo-hp.txt`). **LEFT:** one managed controller swap seen through the verb — no newer controller existed today; the next controller release shows it. (c) accepted (decision 165). | — | the operator decides whether (a)–(c) are accepted or need work before the first paying customer | CC | | **R-132** | Security & access | P3 | **`curl -w '%{redirect_url}'` reconstructs the request URL WITH its basic-auth credential** — so a `-u ":$HUB_PW"` call that never put the password in a URL still printed it **Merged 2026-10-05 from R-350 (duplicate):** (1) 2026-08-20 occurrence: POST /configuration/artifacts answers 303; leak lives only in the CC transcript under ~/.claude/projects/ on DooPlex, not in git/evidence (checked then). (2) `-v` and `--libcurl` also re-render the credential, not only %{redirect_url}; confirm redirects with %{http_code} + follow-up GET. (3) Rotation path: hub /configuration form (current_password/new_password/confirm_password); DB override wins over ConfigMap (break-glass); CC can rotate file-to-file without printing (operator-present-one-time-secrets) if asked. | **WAITING-ON-OPERATOR** — **ACTION: rotate `HUB_PW`** **Folded R-580 2026-10-03** (the same `curl -w %{redirect_url}` credential echo, seen again 2026-09-18). | — | Happened on 2026-07-31 while red-proofing the R-120 gate: the hub operator password was written to the session transcript by the write-out format, not by the request. `-u` is safe; the *reporting* was not. Rule: read the redirect from `-D -` and grep `^Location:`, never `%{redirect_url}`, on any authenticated call. Rotate the hub password (`/configuration` → Login password; ConfigMap `auth.password_hash` is the reset path) and update `~/.config/credentials` | Viktor | | **R-137** | Security & access | P3 | **Cloudflare geo-WAF rules are zone-scoped and non-namespaced — four cross-tenant faults.** `globalRuleDesc = "[felhom-geo] Global"` (`waf.go:18`) is one literal description per ZONE; `appRuleDescPrefix` keys by app name with no customer (`waf.go:21`); `BuildGlobalExpression` has no positive hostname scoping (`waf.go:241`); `applyDiff` deletes every `[felhom-geo]` rule not in THIS box's desired set (`geosync.go:320`) | READY (M) — **blocks shared-zone onboarding** | — | With two customers in one zone: they overwrite each other's Global rule forever; one customer's country policy applies zone-wide; per-app rules collide by name; and disabling the feature for one (or the hub's `RemoveGeoRules`) wipes them all. Interim mitigation, no code: keep geo-restriction OFF for every shared-zone customer. Fix = namespace descriptions by `customer_id` + add `http.host ends_with ""` to both expressions — a TWO-REPO change (controller + hub `RemoveGeoRules`). Same audit §5.1 | CC | -| **R-138** | Security & access | P3 | **A shared-zone `cf_api_token` is a zone-wide DNS-write capability on a customer's box** — written 0600 to `/opt/docker/stacks/traefik/.env` (`controller/internal/infra/infra.go:123`) **-- 2026-10-08 design:** `audits/day-2026-10-08/design-R-138.md` — the shared-zone premise is gone (own domain per customer, `01` §7), but nothing enforces it and an account-wide token has the same reach; option B (hub refuses a duplicate or nested domain) needs no decision and also covers R-415. Question D6 on STATUS's decision sheet. | READY (S) **2026-10-05 (burn-down night): NEEDS A DESIGN.** No notion of a „shared zone” exists anywhere; the token is typed into the hub form, so the guard belongs on the hub side with that notion defined. | — | Today each box holds a token for a zone nobody else uses, so the blast radius is one customer. Under a shared customer zone, one compromised tester box could repoint every other tester's DNS. The ACME path is already switchable — an empty token selects HTTP-01 (`traefik.yml.tmpl`) — so the fix is policy plus a guard that refuses to hand a shared-zone customer a zone-scoped token. Same audit §5.2 | CC | +| **R-138** | Security & access | P3 | **A shared-zone `cf_api_token` is a zone-wide DNS-write capability on a customer's box** — written 0600 to `/opt/docker/stacks/traefik/.env` (`controller/internal/infra/infra.go:123`) **-- 2026-10-08 design:** `audits/day-2026-10-08/design-R-138.md` — the shared-zone premise is gone (own domain per customer, `01` §7), but nothing enforces it and an account-wide token has the same reach; option B (hub refuses a duplicate or nested domain) needs no decision and also covers R-415. Question D6 on STATUS's decision sheet. **-- 2026-10-08 (afternoon):** option B BUILT on hub main (the duplicate/nested/felhom.eu domain guard, R-415 closed with it). Left: option C (check the key's reach with Cloudflare) — question D6. | READY (S) **2026-10-05 (burn-down night): NEEDS A DESIGN.** No notion of a „shared zone” exists anywhere; the token is typed into the hub form, so the guard belongs on the hub side with that notion defined. | — | Today each box holds a token for a zone nobody else uses, so the blast radius is one customer. Under a shared customer zone, one compromised tester box could repoint every other tester's DNS. The ACME path is already switchable — an empty token selects HTTP-01 (`traefik.yml.tmpl`) — so the fix is policy plus a guard that refuses to hand a shared-zone customer a zone-scoped token. Same audit §5.2 | CC | | **R-255** | Security & access | P3 | **The check that would catch a fourth secret-in-the-body covers 4 of 27 pages, and the cheap gate that covers all 36 templates is blind to the shape that actually shipped.** Filed 2026-08-08 while closing R-254, **because a partial guard reported as complete is worse than no guard — it stops the next person looking.** **Two nets, both measured.** **(1) `scripts/secret_in_markup_gate.py`** reads all 36 templates and convicts any `{{ … }}` naming a secret unless allowlisted with a reason. It catches `{{.RetrievalPassword}}` and `{{.InitialCreds.Password}}`, **and it catches a launder through a local variable** because the assignment itself names the secret (`{{$v := .InitialCreds.Password}}` is convicted — verified). **It is blind to a secret arriving under a NEUTRAL PAGE-DATA KEY** — `data["Tagline"] = creds.Password` then `{{.AppInfo.Tagline}}` passes it cleanly, also verified. **That is exactly the shape of R-254 site two** (`value="{{$val}}"` inside an `{{if eq .Type "secret"}}` branch), so the gate **would not have caught one of the three instances it was written for.** **(2) The runtime body assertion** — render the page and grep the response for a sentinel — catches every shape, including that one (demonstrated on the same planted leak the gate missed). But it needs each page's data to be constructible in a test, and **only 4 of 27 page templates have that today**: `settings_security`, `app_info`, `deploy`, `backups_restore` — the four that were touched by R-249/R-252/R-253/R-254 and therefore got their own tests. **The other 23 pages have no runtime coverage at all.** **What closing this needs, so the cost is not re-estimated:** a per-page data fixture for the remaining 23 (most need a wired `Server` — `stackMgr`, `backupMgr`, agent seams), then one table-driven test that renders each with a sentinel substituted for every string in its data and asserts the sentinel is absent. **That is real scaffolding, which is why it was NOT built inside R-254's session** rather than half-built and declared done. | **READY** — owner Viktor | — | — | operator | | **R-338** | Security & access | P3 | **`demo-hp` is not on the R-50 island at all, and `operations/nodes.md` states that it is.** The page records both fleet boxes as island-migrated 2026-07-25. True of `felhom-pve`; **false of `demo-hp`**, whose `agent.json` has `listen_addr: 192.168.0.87:8443` — the customer LAN address — and **no `island_bridge`/`island_guest_addr` keys at all**, whose guest 9201 has `net0` only (no `eth1`), and whose `vmbr9` exists with **zero members**. The controller's `controller.yaml` points at the LAN address, so the box works; this is inventory drift, not breakage. **Two costs.** A session trusting the page addresses the wrong endpoint — that happened on 2026-08-18 and the resulting timeout was briefly read as a fault. And the agent's local API is **bound to the customer LAN on this box** rather than to a point-to-point island, which is the exposure R-50 was built to remove — so a documented security property is claimed for a box that does not have it **Checked from source 2026-10-05 (burn-down round 2):** nodes.md:86-88 still claims demo-hp is on the R-50 island (`local_api` on 169.254.253.1:8443/vmbr9, guest eth1). git blame: that claim dates from e6b5fa1e (2026-07-30); the 2026-09-21 edit bcdd5b20 re-read addresses but only reworded the lan_resolver clause -- the island claim was NOT re-verified after the reprovision. Agent config path /etc/felhom-agent/agent.json (felhom-agent cmd/felhom-agent/main.go:171), island keys island_bridge (internal/config/config.go:246). | **READY (S) — NEW 2026-08-18** | — | Decide which is true: migrate `demo-hp` to the island, or correct `nodes.md`. Leaving both is the one option that keeps the doc lying | Viktor decides; CC executes | | **R-616** | Security & access | P3 | **[P3-LOW] The catalog credentials are stored in PLAINTEXT in the box's catalog clone and are printed by an ordinary `git remote -v`.** FOUND 2026-09-21 on guest 9202 while pointing it at a private drill catalog. `Syncer.buildRepoURL` injects `username:token` into the HTTPS URL, and `git clone` persists that URL as the clone's `origin`, so `/catalog-cache/.git/config` holds the token in the clear and **any** diagnostic that prints the remote leaks it — which is what happened in this session's own transcript, and is the same shape as R-580 (`curl -w '%{redirect_url}'`). `maskRepoURL` exists and is used for the LOG lines, so the masking intent is already there; the stored remote is the half that was missed. **INERT ON THE FLEET TODAY** — the live catalog is public and `git.token` is empty on every real box — which is exactly why it should be fixed before it is not: the day the catalog goes private, every box carries a readable credential and every support session that runs `git remote -v` prints it. **Fix shape:** store the remote WITHOUT credentials and supply them per-fetch (a credential helper, `http.extraHeader`, or `GIT_ASKPASS`), and a test asserting the clone's stored `origin` contains no `@`. **Operator action from tonight, unrelated to the fix:** the Gitea `admin` token used for the drill repo was printed by that command and must be rotated. Evidence: `audits/update-night-2026-09-21/05-9202-follows-drill.txt` (redacted). | **READY — rank P3-LOW; owner: CC (controller); one operator action (rotate the Gitea admin token)** **2026-10-05 (burn-down night): FIXED on controller `main`** (`28a5203`; the catalog clone stores no credentials; the token is supplied per fetch; R-615's repo comparison ignores credentials on both sides, so a token never re-clones (`TestR616_TokenSetSameRepoNoRecloneOriginClean`) and a credentialed origin is cleaned at the next pull. The operator's Gitea admin token rotation (the row's second half) is still owed). Ships with the next controller release; close after delivery. **2026-10-06: DELIVERED** in controller v0.298.0 (the clone stores no credentials). Left: the operator's Gitea admin token rotation. | — | — | CC + operator | @@ -257,7 +257,6 @@ stopping line that lies. | **R-719** | Hub & operator | P4 | **[P2-MEDIUM] A customer who already exists never gets a fresh self-bind link when their new box registers: the last link expires in 7 days and nothing re-sends it.** MEASURED 2026-09-29 (new-household drill, `tester-1`): the previous link went out 2026-09-17 07:25 UTC at a host delete and expired 2026-09-24; the box registered at 19:11:30 UTC and its console told the volunteer to open the link from their e-mail — there was none that worked. Hub source: the link is sent at customer creation, RESET, e-mail set on a box-less customer and host delete (`selfbind_mint.go` callers `hosts.go:908`, `configs.go:850`, `customer_reset.go:162`) — never on appliance registration. The volunteer guide says the operator needs to press nothing. The operator pressed „Send self-bind link" (the mail arrived in 1 s) — an operator step the volunteer depends on, recorded, not an intervention. **Fix direction:** send the link when an unclaimed appliance registers while a box-less customer waits with no live link (R-509's first fix shape), or the guide's operator part says: press it the day the volunteer installs. Evidence: `audits/evidence-drill-new-household-2026-09-30/` `phase0/operator-steps.txt`. **CHANGED AND BUILT 2026-09-30 (hub v0.126.0) — the brief's shape was not buildable:** a box registers UNCLAIMED (uuid, MACs, host keys, hardware — nothing of a customer), so "send the link when their box registers" would mail every waiting customer. Built instead: the expired AND used link pages offer „Új linket kérek" → a fresh link to the address registered for that link's customer, only when it has no box, ≤1/h per customer, identical answer for any token (no oracle). Live: the button on the real hub, the same page for a made-up token, no mail; the mint+send path unit-proven (RP42). Evidence: `audits/evidence-fixes-first-tester-2026-09-30/``partD/`. | **WAITING-ON-OPERATOR** (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the operator has not reviewed the changed page shape ("operator may prefer another"), and mint+send is unit-proven only) — **CLOSED 2026-09-30 — hub v0.126.0 (changed shape; operator may prefer another)** | — | — | operator | | **R-814** | Hub & operator | P4 | `PBS-storage-1` (u629193, box 611421) still `status=active`, 19.9 MB | **VERIFY** (2026-10-03 triage: a July watch row with no id; given R-814. WAITING-ON-OPERATOR — no record found that the box was deleted.) — WAITING-ON-OPERATOR | operator console | Delete the box | operator | | **R-844** | Hub & operator | P4 | **The household's OS-update line exists only on the hub's customer timeline.** 2026-10-04: the box itself has no event surface for agent results (the controller UI shows no timeline), so `os_update_applied` is a hub customer event (info: recorded, never mailed). Its stored text is the hub's English sentence; the hu/en bundle text (`mail.event.os_update_applied`) is used only if it is ever mailed. Fix direction: a controller-side line (the controller already polls the agent's local API) when the box gets a household timeline. `audits/os-guest-lane-2026-10-04/partG/hub-customer-timeline-demo-hp.txt` | **READY — owner: CC** **2026-10-05 (burn-down night): NEEDS A DESIGN** — a household timeline on the box does not exist yet. | — | — | CC | -| **R-415** | Hub & operator | P3 | **The hub enforces uniqueness on `customer_id` only: two customers can be given the same domain, or one inside another, silently.** `domain` has no UNIQUE/CHECK and the create path rejects only a duplicate id (row text at `git show 71b8c8c6^:documentation/backlog/OPEN-ITEMS.md`). **RE-FILED 2026-10-08:** the 2026-10-03 triage commit `71b8c8c6` removed it from this register and it never reached CLOSED-ITEMS — lost, not closed (found by the R-138 design). Since the operator's 2026-09-14 ruling (`01` §7: every customer has their own domain) a duplicate or nested domain is a ruled-out state that nothing refuses. `audits/day-2026-10-08/design-R-138.md` option B covers it. | **READY (XS) — no decision needed: build R-138 option B (refuse a domain equal to, containing or under another customer's, or under `felhom.eu`, at save)** | — | Build the hub guard with its red test; ships with a hub release | CC | ## Business & legal — 6 rows (P2 4, P4 2) @@ -277,7 +276,7 @@ stopping line that lies. |---|---|---|---|---|---|---|---| | **R-733** | Process & tooling | P3 | **[P3-LOW] The test bench has NO swap and the boxes have 512 MiB — so a box proof can pass on swap where the bench fails, and nobody records whether a customer guest has swap.** MEASURED 2026-09-30 (R-732): immich's first start was OOM-killed 61–104 times on the bench (swap 0) and passed on 9202 by swapping ~108 MB; the bench given 512 MiB swap passed too. demo-hp 9201, 9202 and demo-felhom 9201 all read `swap: 512`; the golden's guest config is not recorded in its bake evidence, so a customer guest's swap is NOT measured. The harness's memory watch judges `anon` against the limit and never reads `memory.swap.current`. **Needs:** the golden's `swap` read and recorded; the box walk and the harness report `memory.swap.peak` beside `anon`; a decision whether proofs run with swap off (the stricter venue, as R-732's fix was proven). | **READY — rank P3-LOW; owner: CC (harness + golden evidence)** **2026-10-06 night: NARROWED** — the harness records each container's `swap_peak` and the venue's swap (catalog branch `night-held-2026-10-06`, `3f4611c`; SwapRecorded red-proved; bench SwapTotal 0 kB, 9202 524288 kB). LEFT: the golden's swap in its bake evidence, the box walk's record, and whether proofs run with swap off. | — | — | CC | | **R-887** | Process & tooling | P3 | **Some CI jobs are never run, and Gitea fails them ~10–13 minutes later with no log.** Seen 2026-10-05: felhom.eu job 1361 (commit `1122b5c`) and felhom-controller job 1357 (`114ff27`): every step reads `failure`, including the first fetch, the log API answers `file does not exist`, and the runner pod's log has no `task` line for them (its task ids are job id + 1). A re-run through the API ran the controller job normally (success in 32 s) but the felhom.eu job was again never picked up and failed after ~12 min. The runner pod (`gitea-system/act-runner`, image `felhom-act-runner:0.1.0`) had restarted 5 times ~142 min earlier, around the Longhorn instance-manager restart (R-882). Suspected, NOT measured: a stale runner registration claims jobs it never runs — the session's Gitea token cannot list runners (`read:admin` scope). Consequence: a red CI verdict that is not about the code, and **no failure mail** (the alarm step never runs either), so only the pull check sees it. **CORRECTED 2026-10-05 18:21 (operator's screenshot of Gitea → Site Administration → Runners): ONE runner only — ID 2, `felhom-gates-runner`, v0.6.1, label `felhom-gates`, Idle, last online „now". There is no old registration; the stale-registration guess (this row's first text and the reviewer's) was WRONG.** **RE-DIAGNOSED 2026-10-05 (round 2), from the logs that survive:** (1) **„lost in a runner restart" does NOT fit** — the runner pod last restarted 13:24:42Z (`restartCount 5`, all around the 13:20Z Longhorn restart), the lost attempts started 1.5–2.5 h later. (2) **FOUR attempts were lost, not two:** controller job 1357 (start 15:05:41Z → failed 15:18:38Z), felhom.eu job 1359 (15:15:37 → 15:28:38 — the previous session blamed that one on the BusyBox fault; the runner never ran it), job 1361 (15:33:21 → 15:43:38) and its API re-run (15:46:48 → 15:58:38). None has a `task` line in the runner log; every one was failed at a :38-second mark on a 5-minute step, 10–13 min after it was handed out — **the shape of Gitea's periodic „zombie task" stop** (a task assigned to a runner that never reports is failed after ~10 min; no log exists because none was written). (3) The runner's task ids are NOT job id + 1 (the controller re-run was task 1363). (4) **Gitea's own log for the window is gone** — the pod log starts 16:16:32Z (rotated), so the assignment side cannot be read. **Likely mechanism, NOT proven:** the runner's fetch-task request timed out on its side after Gitea had already assigned the task, so the task was orphaned. In that same hour this session polled Gitea's jobs API hard (15 pages every 15 s per wait loop) and Gitea logged „slow" requests — a plausible load cause, and the session's own. Mitigation taken: the session's CI waiter now polls once a minute. **Nothing changed on DooPlex.** **MECHANISM SEEN 2026-10-05 17:15–17:28Z, with Gitea's own log (round 2):** catalog run 1368 (`4828dc7`) — 17:15:14 the job is marked started; 17:15:16 `router: slow POST /api/actions/runner.v1.RunnerService/FetchTask for 10.42.0.42 (the runner), elapsed 3192ms`, then `UpdateRepoRunsNumbers … context canceled` and `GetActionWorkflow: EOF` — **the runner abandoned its fetch after Gitea had assigned the task**; the runner log has no line for task 1371; 17:28:39 `actions/clear_tasks.go:174 stopTasks() [W] Cannot transfer logs of task 1371` — Gitea's zombie-task stop. **The load at that minute:** an outside crawler (216.73.216.78) walking commit pages and `archive/*.tar.gz`, and THIS session's CI waiter, whose 15-page job listings took 13–31 s each. An API re-run passed in 7 s. **Done in-session:** the waiter now asks `GET …/actions/runs?head_sha=` once a minute (1 s). **Not done (DooPlex, the operator's):** the runner's fetch timeout and Gitea's exposure to the crawler. | **OPEN** **DATED CHECK 2026-10-12 (DUE-CHECKS):** if no job was lost since 2026-10-05 16:00Z (no completed job whose runner log has no `task` line / whose log API answers `file does not exist`), close. **NIGHT WATCH 2026-10-05/06 (burn-down night): 2 jobs lost of ~30 runs** — felhom.eu run 1384 (`4aa4d837`, 21:23→21:33Z, no log) and felhom-controller run 1401 (`c67b26be`, 00:45→00:58Z, no log); each re-run once through the API and each passed (2 m 05 s, 57 s). The night's waiter made one filtered call a minute. So the 2026-10-12 close condition („no job lost since 2026-10-05 16:00Z") is already NOT met. `audits/night-burndown-2026-10-05/r887-lost-jobs.txt`. **NIGHT WATCH 2026-10-06/07 (second burn-down night): 0 jobs lost** — every push checked by its commit (felhom.eu, controller, agent, catalog main and the held branch), each completed `success`, one filtered call per check. | — | Operator: decide whether to raise the act-runner fetch timeout and/or rate-limit the public Gitea pages the crawler walks; meanwhile re-run a lost job via `POST /repos/admin//actions/runs//rerun`. Keep the 2026-10-12 check | operator | -| **R-902** | Process & tooling | P3 | **The contact form's mailer has no source in any repository: the running image cannot be rebuilt or changed.** FOUND 2026-10-08 (website refresh, brief Part C5). `manifests/contact-mailer.yaml` names `contact-mailer:latest` with `imagePullPolicy: Never` (imported into k3s by hand); the binary in the pod is `/app/contact-mailer` dated 2026-02-05, Go module `github.com/felhom/contact-mailer` (devel). Searched: `git log --all` of felhom.eu (only the manifest), every repo under `/mnt/5_hdd/felhom.eu/git`, `find /mnt/5_hdd/felhom.eu ~ -maxdepth 5 -iname '*contact-mailer*'`. READ from the binary (strings, read-only `kubectl exec … cat`): ONE send per form (`sendViaResend`) to `TO_EMAIL`, `reply_to` = the visitor; nothing is sent back to the visitor; the operator's notification is Hungarian (`buildEmailHTML`, „Csatolmány … fájl"). The English form needs no mailer change (its subject labels carry „(EN)"). `audits/website-refresh-2026-10-08/mailer.md` | **OPEN — owner: operator** (only the operator can say where the Feb 2026 source went) | — | Find the source and commit it (or rebuild a reviewed equivalent); until then a node rebuild loses the image | operator | +| **R-902** | Process & tooling | P3 | **The contact form's mailer has no source in any repository: the running image cannot be rebuilt or changed.** FOUND 2026-10-08 (website refresh, brief Part C5). `manifests/contact-mailer.yaml` names `contact-mailer:latest` with `imagePullPolicy: Never` (imported into k3s by hand); the binary in the pod is `/app/contact-mailer` dated 2026-02-05, Go module `github.com/felhom/contact-mailer` (devel). Searched: `git log --all` of felhom.eu (only the manifest), every repo under `/mnt/5_hdd/felhom.eu/git`, `find /mnt/5_hdd/felhom.eu ~ -maxdepth 5 -iname '*contact-mailer*'`. READ from the binary (strings, read-only `kubectl exec … cat`): ONE send per form (`sendViaResend`) to `TO_EMAIL`, `reply_to` = the visitor; nothing is sent back to the visitor; the operator's notification is Hungarian (`buildEmailHTML`, „Csatolmány … fájl"). The English form needs no mailer change (its subject labels carry „(EN)"). `audits/website-refresh-2026-10-08/mailer.md` | **OPEN — NARROWED 2026-10-08 (search done, NOT FOUND): owner: operator.** A read-only search found no source on DooPlex (55,548 `*.go`/`go.mod`/`Dockerfile*` files, every disk, media/Longhorn/containerd/PBS excluded by name), in the web editor's storage and file history, in all 10 Gitea repositories' history, or in Docker's build cache (oldest record 2026-09-28). The binary's own clues: one file `/build/main.go`, go1.23.12, no `vcs.*` stamp; image built 2026-02-05 10:40 by BuildKit. The binary is now KEPT in git (`audits/mailer-source-2026-10-08/contact-mailer.bin`), so a DooPlex rebuild no longer loses the program. A replacement plan is written (`audits/mailer-source-2026-10-08/PLAN-replacement.md`); one question left: is the February folder on the operator's Windows workstation? | — | Operator: look on the workstation; if absent, build the replacement from the plan (attended deploy) | operator | | **R-206** | Process & tooling | P4 | **The build-cache cap and the weekly prune exist only as a hand-edited `/etc/docker/daemon.json` on DooPlex — not in Ansible, so a rebuild loses them.** The `node_housekeeping` role must also carry the prune, which today it is forbidden to run | **READY (M) — NEW 2026-08-05** | — | **The spike validated the recipe; this row builds it.** Three parts. **(a) Template `/etc/docker/daemon.json`** with the **`policy` array** form — **the flat form (`{"gc":{"reservedSpace":…}}`) is SILENTLY IGNORED**, measured: the daemon starts, logs nothing, and `docker buildx inspect` still reports the built-in defaults. **The oracle is `docker buildx inspect`, never `dockerd --validate`** — the validator returned `configuration OK` for a bogus key AND for the config that then **crashed the daemon** (`filter` takes one value per policy entry, not an array; `error initializing buildkit: filters expect only one value`). **(b) Narrow the role's Docker ban** (`node-housekeeping.sh.j2:10-14`) to permit exactly `docker builder prune -af` and nothing else — the ban's stated premise ("Docker here runs only unrelated jarr-* dev containers") is obsolete: the growth is Felhom Go build cache. **The measured prune is SYNCHRONOUS** (150.35 GB back at t+0, two consecutive polls <1 MB apart within 60 s) — **unlike containerd's image GC, so it needs no `settle_imagefs` equivalent**, but it MUST measure the filesystem rather than trust the command: `prune` claimed **156.9 GB** and the filesystem returned **150.35 GB**, the 6.5 GB gap being layers still shared with images. **(c) A restart-safety note in the role:** a bad `daemon.json` takes the daemon down AND leaves the `unless-stopped` dev containers stopped — they needed a manual `docker start` — so the role must restart-and-verify, not validate-and-assume. Recipe + every measurement: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` | CC | | **R-209a** | Process & tooling | P4 | **The SSD2 move has NOT survived a reboot, so by this project's own standard it is not fully validated** | **WATCHING — NEW 2026-08-05** | the next DooPlex reboot | **Operator ruled explicitly: do NOT reboot DooPlex.** Uptime verified unbroken (7 weeks 6 days, since 2026-06-10). **The distinction is stated rather than glossed: the MECHANISM is proven** — the guard is wired into both units and containerd refuses to start when a required mount's device is absent — **but the CONSEQUENCE is not**: that a real boot mounts `/mnt/ssd_2` before containerd starts, in this host's actual ordering. Mount-ordering reasoning is precisely the class this project has been burned by (`RequiresMountsFor` RE-MOUNTS rather than refusing — the ep0 lesson), and `CLAUDE.md` prefers a consequence assertion over a mechanism one. **Two deliberate consequences: (1)** the rollback copy `/var/lib/containerd.pre-move-2026-08-05` (**34.3 GB on `/`**) **STAYS** until a reboot validates — which is why `/` sits at 54% and not lower; deleting it now would trade a cheap 34 GB for the only cheap way back. **(2)** validation is **automatic and needs no one to remember it**: `felhom-store-postboot-check.service` (oneshot, enabled, dry-run PASS at install) runs at **every** boot and writes `RESULT: PASS`/`FAIL` to `/var/log/felhom-store-postboot-check.log`, asserting positively that `/mnt/ssd_2` is mounted, that containerd's root is on it, that **`/var/lib/containerd` does NOT exist** (the empty-store trap), that ≥100 images are visible and that both dev containers run. **Next action: after the next reboot — planned or not — read that file; on PASS, `rm -rf /var/lib/containerd.pre-move-2026-08-05` returns ~34 GB to `/`** **P3's prune already removed the urgency: `/` went 86% → 53% used and SSD1's Longhorn disk went `Schedulable=False (DiskPressure)` → `Schedulable=True` (18.85% → 50.32% available).** The move was ruled "cap then move"; the cap is in and the pressure is gone, so this is now a deliberate choice rather than a rescue. **The numbers, measured (`Crucial-SSD-240G`, `/mnt/ssd_2/data/longhorn`, `storageMaximum` 235,148,750,848):** available today **214,958,080,000 (91.41%)**; 25% floor **58,787,187,712**. Moving the whole containerd tree at steady state (~31.5 GB images + ≤30 GB cache ≈ 65 GB) leaves **63.77%, i.e. +38.8 pp above the floor — comfortably safe as measured.** **But `storageScheduled` on SSD2 is 139,586,437,120 while `df` says only 20,094,939,136 is actually used** — Longhorn has overcommitted 6.9× — and if those volumes ever inflate to their scheduled size, the same disk lands at **13.00%, i.e. 12 pp BELOW the floor → `Schedulable=False`**, which is exactly the failure that just took SSD1 out. **Recommendation: do the move only together with setting `storageReserved` on SSD2 to cover the containerd tree (~80 GB); SSD2 reserving zero while HDD2 and HDD4 each reserve 500 GB is an anomaly in its own right.** Mechanism, if it goes ahead: **containerd's `root` in `/etc/containerd/config.toml`** (the key is present but commented out) — **not** Docker's `data-root`, which would move only 0.62 GB. Guard: `RequiresMountsFor=/mnt/ssd_2` on `containerd.service` **and** `docker.service`, remembering that **`RequiresMountsFor` RE-MOUNTS rather than refusing** ([[ep0-datastore-volume-move-2026-07-27]]) — so it must be tested with a genuinely absent device, and the move is not validated until it has survived a **reboot**. Full pre-analysis: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` §P6 | operator + CC | | **R-230** | Process & tooling | P4 | **Three instruction/memory follow-ups deliberately left by the part-2 session (2026-08-06), each needing a decision rather than an implementation.** (a) **A ruling is owed on auto-written staleness.** The hand-written `CLAUDE.md` files are now clean of version literals and expired blocks — the gate enforces it — but `MEMORY.md`, which Claude writes and which is the LARGER half of what loads (8.4k tokens vs the root file's 6.6k), carries **21 lines with component version literals**, **5 with bare host addresses**, and an entry still reading *"demo boxes REMOTE till ~08-02"* — the same expired-TEMPORARY class the gate was built to kill, now surviving in the one file the gate's content rules do not cover. **Partly actioned 2026-08-06 (close-out), and the ruling is STILL OWED:** the **three statements that were actively false** were corrected — `R-193 decision open` (closed 2026-08-05), `demo boxes REMOTE till ~08-02` (the box answers on the home LAN), `OPEN R-25b` (shipped 2026-07-21) — and gate check 6 now **WARNs** on version literals, host addresses, expired statements and stale-open citations in the index. WARN, never FAIL: Claude writes that file between sessions, so a hard failure would refuse a human's push over a line no human typed, and the warning is read by the model that will next edit it. **The remaining 32 version literals and 4 host addresses were deliberately left** for that loop. What is still owed is the bulk-correction ruling. **Correcting the premise:** the earlier report's "three expired statements" were all FALSE POSITIVES — each matched an ISO date inside a markdown link target, i.e. a filename — while the one real expired claim carried no ISO date at all. (b) **CLOSED 2026-08-06 (close-out)** — the workspace-root `CLAUDE.md` **is now a relative symlink** to the versioned copy, so the divergence class is gone rather than policed. Check 5 learned two shapes: for a link it asserts the target resolves to a real file (**a dangling link is worse than a diverged copy — the instructions load NOTHING and there is no content left to notice is wrong**), for two files byte-identity as before, so a clone elsewhere is unaffected. **Proven, not assumed:** three fresh sessions logged `session_start` for the link path, and a fourth **with no tools at all** quoted standing rule 1 verbatim — the content reaches the model, not just the path. (c) **The spec-as-failing-test pilot**, approved in principle and not started (was R-229(d)). | **READY** — owner Viktor | — | — | operator | diff --git a/hub/CHANGELOG.md b/hub/CHANGELOG.md index 644dd43b..777704c2 100644 --- a/hub/CHANGELOG.md +++ b/hub/CHANGELOG.md @@ -1,9 +1,16 @@ -## Unreleased (2026-10-08) — an alarm when a box never backs up off-site because its escrow is pending (R-243; `09` §3 decision 179); a deleted customer's audit rows go after 1 year (R-901; decision 181); the operator's older-recovery-package mail (R-304; decision 183) — ships with tomorrow's hub release +## Unreleased (2026-10-08) — an alarm when a box never backs up off-site because its escrow is pending (R-243; `09` §3 decision 179); a deleted customer's audit rows go after 1 year (R-901; decision 181); the operator's older-recovery-package mail (R-304; decision 183); no duplicate or nested customer domain (R-415, R-138 option B) — ships with tomorrow's hub release **Operator action on deploy: none.** Expect ONE `offsite_escrow_pending` mail for **Tester 2** on the first sweep after the deploy: its latest report (2026-10-04) says off-site ON, escrow `pending`, no successful run ever — the state the operator believes it is in (decision 170). +- **R-415 / R-138 option B (operator ruling 2026-09-14, `01` §7 — every customer has their own domain):** the create and + edit paths refuse a domain that equals, contains or lies under another customer's, or is under `felhom.eu`, BEFORE + anything is generated or provisioned; the form re-renders with the submitted values and one sentence; a store error + refuses too. Label-boundary matching, case-insensitive, trailing dot ignored. `store.DomainConflict`; test + `TestR415_CreateAndEditRefuseConflictingDomain` (red-proved: without the create guard „b1 with example.hu was + created"). Live data read first (read-only DB copy, deleted): the four customers' domains are distinct and none is + under felhom.eu. The form's example and one old test fixture said `kovacs.felhom.eu` — corrected to `kovacs.hu`. - **R-304 option C (decision 183):** new event type `recovery_older_package` (sent by the controller of the same day when a household's code opens, or may open, an older sealed escrow package): in `allowedEventTypes` and `notify.operatorOnlyEvents`, no household text. Test `TestR304_RecoveryOlderPackageIsAllowlistedAndOperatorOnly` diff --git a/hub/internal/store/domain_conflict.go b/hub/internal/store/domain_conflict.go new file mode 100644 index 00000000..69d1f533 --- /dev/null +++ b/hub/internal/store/domain_conflict.go @@ -0,0 +1,47 @@ +package store + +import "strings" + +// R-415 / R-138 option B (2026-10-08; operator ruling 2026-09-14, `01` §7: every customer has their OWN domain, never a +// name under felhom.eu). The hub enforced uniqueness on customer_id only, so two customers could be given the same +// domain — or one inside the other, a shared zone in all but name — silently. DomainConflict answers, for a domain about +// to be saved for customerID, which rule it breaks ("" = none): +// +// - "felhom.eu" — the domain is felhom.eu or a name under it; +// - "" — another customer's domain equals it, contains it, or lies under it. +// +// Matching is case-insensitive, ignores a trailing dot, and is on LABEL boundaries („notexample.hu" is not under +// „example.hu"). An empty domain conflicts with nothing. Pinned by TestR415_* (domain_conflict_test.go) and the +// handler test TestR415_CreateAndEditRefuseConflictingDomain. +func (s *Store) DomainConflict(customerID, domain string) (string, error) { + d := normDomain(domain) + if d == "" { + return "", nil + } + if d == "felhom.eu" || strings.HasSuffix(d, ".felhom.eu") { + return "felhom.eu", nil + } + rows, err := s.db.Query(`SELECT customer_id, domain FROM customer_configs WHERE customer_id != ? AND domain != ''`, customerID) + if err != nil { + return "", err + } + defer rows.Close() + for rows.Next() { + var id, other string + if err := rows.Scan(&id, &other); err != nil { + return "", err + } + o := normDomain(other) + if o == "" { + continue + } + if d == o || strings.HasSuffix(d, "."+o) || strings.HasSuffix(o, "."+d) { + return id, nil + } + } + return "", rows.Err() +} + +func normDomain(d string) string { + return strings.TrimSuffix(strings.ToLower(strings.TrimSpace(d)), ".") +} diff --git a/hub/internal/web/configs.go b/hub/internal/web/configs.go index 62a2f063..297eb544 100644 --- a/hub/internal/web/configs.go +++ b/hub/internal/web/configs.go @@ -741,6 +741,18 @@ func (s *Server) handleConfigCreate(w http.ResponseWriter, r *http.Request) { return } + // R-415 / R-138 option B: a domain equal to, inside or containing another customer's, or under felhom.eu, is refused + // BEFORE anything is generated or provisioned. + if msg := s.domainConflictMessage(customerID, r.FormValue("domain")); msg != "" { + s.renderConfigForm(w, r, true, &store.CustomerConfig{ + CustomerID: customerID, + CustomerName: r.FormValue("customer_name"), + Domain: r.FormValue("domain"), + Email: r.FormValue("email"), + }, nil, msg) + return + } + // Generate credentials. // // R-597: the Owner passphrase follows the language the operator is choosing ON THIS FORM, not @@ -861,6 +873,13 @@ func (s *Server) handleConfigUpdate(w http.ResponseWriter, r *http.Request, cust s.renderConfigForm(w, r, false, cfg, submitted, "Display Name and Domain are required.") return } + // R-415 / R-138 option B — the same guard on edit (a customer's own current domain never conflicts with itself). + if msg := s.domainConflictMessage(customerID, cfg.Domain); msg != "" { + var submitted map[string]interface{} + _ = json.Unmarshal([]byte(buildConfigJSON(r)), &submitted) + s.renderConfigForm(w, r, false, cfg, submitted, msg) + return + } cfg.ConfigJSON = buildConfigJSON(r) @@ -1769,3 +1788,22 @@ func (s *Server) handleGeoDisable(w http.ResponseWriter, r *http.Request, custom w.Header().Set("Content-Type", "application/json") json.NewEncoder(w).Encode(map[string]interface{}{"ok": true, "message": "Geo-restriction removed from Cloudflare."}) } + +// domainConflictMessage is the operator's sentence for a refused domain ("" = allowed). A store error refuses too +// (fail closed: the save can be retried; a duplicate domain cannot be undone once boxes use it). +func (s *Server) domainConflictMessage(customerID, domain string) string { + other, err := s.store.DomainConflict(customerID, domain) + if err != nil { + s.logger.Printf("[ERROR] domain check for %s failed: %v — save refused", customerID, err) + return "The domain could not be checked against the other customers — nothing was saved. Try again." + } + switch other { + case "": + return "" + case "felhom.eu": + return fmt.Sprintf("Domain %q is under felhom.eu — every customer has their own domain (01 §7). Nothing was saved.", strings.TrimSpace(domain)) + default: + s.logger.Printf("[WARN] domain %q for %s refused: it overlaps customer %s's domain (R-415)", strings.TrimSpace(domain), customerID, other) + return fmt.Sprintf("Domain %q equals, contains or lies under the domain of customer %q — every customer has their own domain (01 §7). Nothing was saved.", strings.TrimSpace(domain), other) + } +} diff --git a/hub/internal/web/configs_debug_test.go b/hub/internal/web/configs_debug_test.go index db330341..cf18fe4a 100644 --- a/hub/internal/web/configs_debug_test.go +++ b/hub/internal/web/configs_debug_test.go @@ -64,13 +64,13 @@ func TestConfigUpdate_DebugSurvivesRebuild_OffsiteUntouched(t *testing.T) { const id = "cust-dbg" if err := st.SaveCustomerConfig(&store.CustomerConfig{ - CustomerID: id, CustomerName: "Kovács", Domain: "kovacs.felhom.eu", ConfigJSON: "{}", + CustomerID: id, CustomerName: "Kovács", Domain: "kovacs.hu", ConfigJSON: "{}", }); err != nil { t.Fatalf("seed customer: %v", err) } // 1) First save WITH offsite enabled → provisions + merges the descriptor. No debug yet. - const offsiteForm = "customer_name=Kov%C3%A1cs&domain=kovacs.felhom.eu&dr_tier=on&offsite_enabled=on&offsite_type=shared&offsite_quota_gb=50" + const offsiteForm = "customer_name=Kov%C3%A1cs&domain=kovacs.hu&dr_tier=on&offsite_enabled=on&offsite_type=shared&offsite_quota_gb=50" w := httptest.NewRecorder() s.handleConfigUpdate(w, postForm("/configs/"+id+"/edit", offsiteForm), id) if w.Code != http.StatusSeeOther { diff --git a/hub/internal/web/r415_domain_conflict_test.go b/hub/internal/web/r415_domain_conflict_test.go new file mode 100644 index 00000000..4f995abc --- /dev/null +++ b/hub/internal/web/r415_domain_conflict_test.go @@ -0,0 +1,62 @@ +package web + +import ( + "net/http" + "net/http/httptest" + "net/url" + "strings" + "testing" +) + +// R-415 / R-138 option B: every customer has their own domain (`01` §7). The CONSEQUENCE asserted: a refused create +// stores nothing; an allowed one is stored; an edit keeping its own domain is allowed. +// RED-PROOF: delete the domainConflictMessage block in handleConfigCreate → „b" with „example.hu" is created → FAILS. +func TestR415_CreateAndEditRefuseConflictingDomain(t *testing.T) { + s, st := newTestServer(t) + post := func(id, domain string) *httptest.ResponseRecorder { + form := url.Values{"customer_id": {id}, "customer_name": {"T " + id}, "email": {"t@example.org"}, "domain": {domain}} + req := httptest.NewRequest(http.MethodPost, "/configs/new", strings.NewReader(form.Encode())) + req.Header.Set("Content-Type", "application/x-www-form-urlencoded") + rr := httptest.NewRecorder() + s.handleConfigCreate(rr, req) + return rr + } + if rr := post("a", "example.hu"); rr.Code != http.StatusSeeOther { + t.Fatalf("create a: %d %s", rr.Code, rr.Body.String()) + } + for _, c := range []struct{ id, domain string }{{"b1", "example.hu"}, {"b2", "x.EXAMPLE.hu."}, {"b3", "hu"}, {"b4", "t1.felhom.eu"}, {"b5", "felhom.eu"}} { + rr := post(c.id, c.domain) + if rr.Code == http.StatusSeeOther { + t.Errorf("%s with %q was created — it must be refused", c.id, c.domain) + } + if !strings.Contains(rr.Body.String(), "Nothing was saved") { + t.Errorf("%s with %q: the refusal must say nothing was saved; got %d", c.id, c.domain, rr.Code) + } + if got, _ := st.GetCustomerConfig(c.id); got != nil { + t.Errorf("%s with %q: a config was stored", c.id, c.domain) + } + } + for _, c := range []struct{ id, domain string }{{"c1", "example2.hu"}, {"c2", "notexample.hu"}} { + if rr := post(c.id, c.domain); rr.Code != http.StatusSeeOther { + t.Errorf("%s with %q must be allowed; got %d %s", c.id, c.domain, rr.Code, rr.Body.String()) + } + } + // Edit: „a" keeping its own domain is allowed; „c1" moving under „a"'s is refused and keeps its old domain. + edit := func(id, domain string) *httptest.ResponseRecorder { + form := url.Values{"customer_name": {"T " + id}, "email": {"t@example.org"}, "domain": {domain}} + req := httptest.NewRequest(http.MethodPost, "/configs/"+id, strings.NewReader(form.Encode())) + req.Header.Set("Content-Type", "application/x-www-form-urlencoded") + rr := httptest.NewRecorder() + s.handleConfigUpdate(rr, req, id) + return rr + } + if rr := edit("a", "example.hu"); strings.Contains(rr.Body.String(), "Nothing was saved") { + t.Errorf("editing a with its own domain must not conflict with itself: %s", rr.Body.String()) + } + if rr := edit("c1", "shop.example.hu"); !strings.Contains(rr.Body.String(), "Nothing was saved") { + t.Errorf("moving c1 under a's domain must be refused; got %d", rr.Code) + } + if got, _ := st.GetCustomerConfig("c1"); got == nil || got.Domain != "example2.hu" { + t.Errorf("c1's domain must stay example2.hu after a refused edit; got %+v", got) + } +} diff --git a/hub/internal/web/templates/config_form_body.html b/hub/internal/web/templates/config_form_body.html index 84f08151..bf4b3c0f 100644 --- a/hub/internal/web/templates/config_form_body.html +++ b/hub/internal/web/templates/config_form_body.html @@ -25,7 +25,7 @@