From a406efc7d52807398eacba39a3bb1e94355a9da7 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Sat, 10 Oct 2026 10:02:33 +0200 Subject: [PATCH] release 2026-10-10: hub 0.145.0 deployed (security check PASS live), controller 0.305.0 delivered, kernel button offers -22; R-921/R-922 released; R-925 consequences; build skill: sync only the hub when other resources drift Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- REPORT-release-2026-10-10.md | 57 +++++++++++++++++++ STATUS.md | 17 +++++- .../audits/release-2026-10-10/delivered.txt | 25 ++++++++ .../release-2026-10-10/floors-0.305.0.txt | 7 +++ .../release-2026-10-10/hub-0.145.0-checks.txt | 23 ++++++++ .../release-2026-10-10/hub-0.145.0-deploy.txt | 8 +++ .../release-2026-10-10/system-page-kernel.txt | 5 ++ documentation/backlog/OPEN-ITEMS.md | 6 +- skills/felhom-build-deploy/SKILL.md | 7 ++- 9 files changed, 149 insertions(+), 6 deletions(-) create mode 100644 REPORT-release-2026-10-10.md create mode 100644 documentation/audits/release-2026-10-10/delivered.txt create mode 100644 documentation/audits/release-2026-10-10/floors-0.305.0.txt create mode 100644 documentation/audits/release-2026-10-10/hub-0.145.0-checks.txt create mode 100644 documentation/audits/release-2026-10-10/hub-0.145.0-deploy.txt create mode 100644 documentation/audits/release-2026-10-10/system-page-kernel.txt diff --git a/REPORT-release-2026-10-10.md b/REPORT-release-2026-10-10.md new file mode 100644 index 00000000..cdd30087 --- /dev/null +++ b/REPORT-release-2026-10-10.md @@ -0,0 +1,57 @@ +# REPORT — 2026-10-10: the waiting fixes released (a security hole closed), and the kernel approval rule + +| Part | What | Result | +|---|---|---| +| A1 | Hub 0.145.0 (operator present) | Deployed 07:51Z, `Deployment/hub` only; `/healthz` 200, `/system` 200; **live check of the hole: PASS** | +| A2 | Controller 0.305.0; agent | Controller on demo-hp, demo-felhom, Tester 1 (floors 07:58:50Z, all three healthy by 07:59:08Z). Agent: no change since 0.154.0 — no release | +| A3 | MAIL-HOLD after the deploy | Marker absent, no banner, 0 MAIL-HOLD log lines | +| B | Kernel approval rule (`09` §3 decision 195) | Built, red-proved, deployed; the System page offers **7.0.14-22** — not clicked | + +**Register: before 137 · after 137 · opened 0 · closed 0.** R-921 and R-922 marked released; R-925 got two measured +consequences. + +## A1 — hub 0.145.0 +- Contents: the SECURITY fix (`/preferences`, `/notify` refuse another household's key), R-922 (`email_cleared`), + MAIL-HOLD, two log lines without the address, the kernel rule. `go test ./...` rc 0; CI 1624 (code), 1625 (manifest). +- **Image push:** DooPlex's saved registry login is stale since the R-925 rotation (token endpoint 401 for it, 200 for + the current one) — the push used a one-off `DOCKER_CONFIG` login in the scratchpad, password file→stdin, shredded after. +- **Deploy:** the `felhom` app also showed 3 Secrets + `Deployment/umami` OutOfSync (R-925's de-gitting); a whole-app sync + would have pushed git's view over the rotated Secrets, so only `Deployment/hub` was synced. Pod image 0.145.0, log + `felhom-hub 0.145.0 starting`. +- **Live check (two channels):** Tester 1's key (read file→file from its controller.yaml, 64 chars, never printed): + own household, own settings → **200**; demo-felhom's household, demo-felhom's own settings → **403 „customer_id does + not match the key"**. Second channel: both households' stored rows hashed from a hub DB copy (with -wal/-shm) before + the deploy and after. +- **My mistake, corrected:** the control request stored Tester 1's event list as `["null"]` — my baseline script read + the stored JSON `null` as a list holding the word null. No real event was enabled by it. Restored with the same + own-household request carrying `enabled_events: null`; both rows then **equal the pre-deploy baseline** (hashes). + +## A2 — controller 0.305.0 +R-921 pre-check + R-922 `email_cleared`. CI 1627; image in the registry (anonymous 200). Floors 0.305.0 with MinAgent +0.131.0 for demo-hp, demo-felhom, tester-1; global floor and Tester 2 untouched. Read-back per box: sudo +`felhom-priv-apply controller-image 9201` → `WROTE … 0.305.0` → agent „new controller healthy" (host journal) and +`docker ps` 0.305.0 healthy (guest). + +## B — the kernel rule +`KernelStatus` now offers the newest kernel in every ring-0 box's set of kernels booted healthily after a night stage; +per box and kernel only the newest ended step counts. 6 tests (`kernel_approval_test.go`); against the old function two +fail (today's case: „booted different kernels" where -22 was due; and the fell-back case). Behaviour change to note: a +box whose NEWEST step fell back on -23 no longer blocks the button — -22 (healthy earlier) is offered instead. +Live: the System page shows „Approve kernel set" (2 packages, first seen 07:52Z); the hub DB's candidate is +`proxmox-kernel-7.0` + `proxmox-kernel-7.0.14-22-pve-signed`. `11` §5.11 and `09` §3 updated; poster facts: no fact +changed (it says only „the operator approves on the System page"). + +## Instruction-file edit (rule 5) +`skills/felhom-build-deploy/SKILL.md`, hub step 4: added — read what is OutOfSync before syncing; if more than +`Deployment/hub`, sync only it (command given); a 401 push from DooPlex is the stale login (R-925). Why: both measured today. + +## Not done +R-922 live (a real household clear) and R-921 live (a two-tier night) — not exercised today. No release of the agent. + +## Decisions for the operator +1. **Refresh DooPlex's registry login** (`docker login gitea.dooplex.hu` with the new admin password, once, at your + keyboard; or a scoped push token). Pick: do it. **If you do nothing:** every build session must use a one-off login, + and a session that does not know this stops at the push. +2. **„Approve kernel set" for 7.0.14-22** — yours to click. Pick: click it (both demo boxes booted -22 healthily; ring 1 + still takes it only by a signed job). **If you do nothing:** ring 1 gets no kernel; tonight demo-hp moves to -23, and + once it boots -23 healthily the button will offer -23 instead. diff --git a/STATUS.md b/STATUS.md index 22c434f9..53722544 100644 --- a/STATUS.md +++ b/STATUS.md @@ -2,8 +2,21 @@ **Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop) is off; nothing was sent to it.** -**Updated 2026-10-09 (afternoon): hub 0.144.0; demo-hp, demo-felhom and Tester 1 run agent 0.154.0 and controller -0.304.0. The open-items list is at 138. Reports: `REPORT-break-the-circle-2026-10-09.md`, `REPORT-dooplex-survival-2026-10-09.md`, `REPORT-day4-2026-10-09.md`.** +**Updated 2026-10-10: hub 0.145.0; demo-hp, demo-felhom and Tester 1 run agent 0.154.0 and controller +0.305.0. The open-items list is at 137. Reports: `REPORT-release-2026-10-10.md`, `REPORT-break-the-circle-2026-10-09.md`, `REPORT-dooplex-survival-2026-10-09.md`, `REPORT-day4-2026-10-09.md`.** + +## Saturday 2026-10-10: released — the security hole is closed, and the kernel button shows + +- **Hub 0.145.0 is live** (you were present). One box's key could change another household's mail settings. Now it + cannot: checked live with Tester 1's key against demo-felhom — refused (403); the same key for its own household — + accepted. Nothing changed in either household. +- **Controller 0.305.0 runs on demo-hp, demo-felhom and Tester 1.** The agent did not change (still 0.154.0). +- **Also live:** a household's cleared mail address is deleted; a restored hub can start quiet; a backup no longer + stops the apps while the box's other backup is still running. +- **„Approve kernel set" shows now, for 7.0.14-22** (your new rule: the newest kernel every demo box started + healthily). Not clicked — that is yours. +- **Found on the way:** DooPlex's saved login for the image registry still has the old password (it changed + yesterday), so builds cannot push until it is refreshed. I used a one-off login and destroyed it. ## Evening (2026-10-09): the system poster is in the repository, and staying true is now a rule diff --git a/documentation/audits/release-2026-10-10/delivered.txt b/documentation/audits/release-2026-10-10/delivered.txt new file mode 100644 index 00000000..a2b1d342 --- /dev/null +++ b/documentation/audits/release-2026-10-10/delivered.txt @@ -0,0 +1,25 @@ +== 2026-10-10T07:59:12Z read-back +--- demo-hp +agent: felhom-agent 0.154.0 +gitea.dooplex.hu/admin/felhom-controller:0.305.0 +gitea.dooplex.hu/admin/felhom-controller:0.305.0 Up 11 seconds (healthy) +StartedAt=2026-10-10T07:59:02.745744507Z +2026-10-10T09:59:00+02:00 demo-hp sudo[894563]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201 +2026-10-10T09:59:01+02:00 demo-hp felhom-priv-apply[894595]: felhom-priv-apply: WROTE controller-image 9201 gitea.dooplex.hu/admin/felhom-controller:0.305.0 +2026-10-10T09:59:08+02:00 demo-hp felhom-agent[1262]: time=2026-10-10T09:59:08.372+02:00 level=INFO msg="controller-swap: new controller healthy" vmid=9201 target=gitea.dooplex.hu/admin/felhom-control +--- felhom-pve +agent: felhom-agent 0.154.0 +gitea.dooplex.hu/admin/felhom-controller:0.305.0 +gitea.dooplex.hu/admin/felhom-controller:0.305.0 Up 14 seconds (healthy) +StartedAt=2026-10-10T07:58:58.628209837Z +2026-10-10T09:58:56+02:00 demo-felhom sudo[284694]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201 +2026-10-10T09:58:57+02:00 demo-felhom felhom-priv-apply[284701]: felhom-priv-apply: WROTE controller-image 9201 gitea.dooplex.hu/admin/felhom-controller:0.305.0 +2026-10-10T09:59:07+02:00 demo-felhom felhom-agent[1205]: time=2026-10-10T09:59:07.187+02:00 level=INFO msg="controller-swap: new controller healthy" vmid=9201 target=gitea.dooplex.hu/admin/felhom-con +--- root@192.168.0.154 +agent: felhom-agent 0.154.0 +gitea.dooplex.hu/admin/felhom-controller:0.305.0 +gitea.dooplex.hu/admin/felhom-controller:0.305.0 Up 13 seconds (healthy) +StartedAt=2026-10-10T07:59:02.717345534Z +2026-10-10T09:59:00+02:00 felhom sudo[418770]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-priv-apply controller-image 9201 +2026-10-10T09:59:01+02:00 felhom felhom-priv-apply[418798]: felhom-priv-apply: WROTE controller-image 9201 gitea.dooplex.hu/admin/felhom-controller:0.305.0 +2026-10-10T09:59:08+02:00 felhom felhom-agent[1161]: time=2026-10-10T09:59:08.339+02:00 level=INFO msg="controller-swap: new controller healthy" vmid=9201 target=gitea.dooplex.hu/admin/felhom-controll diff --git a/documentation/audits/release-2026-10-10/floors-0.305.0.txt b/documentation/audits/release-2026-10-10/floors-0.305.0.txt new file mode 100644 index 00000000..b75f6340 --- /dev/null +++ b/documentation/audits/release-2026-10-10/floors-0.305.0.txt @@ -0,0 +1,7 @@ +== floors 2026-10-10T07:58:50Z: 0.305.0 with min_agent 0.131.0; global floor not touched; Tester 2 not touched +demo-hp: 303 Location: /customers/demo-hp?flash=floor_set +demo-felhom: 303 Location: /customers/demo-felhom?flash=floor_set +tester-1: 303 Location: /customers/tester-1?flash=floor_set +2026/10/10 09:58:50 [INFO] Customer demo-hp controller-version floor override set to "0.305.0" (declared MinAgent "0.131.0") +2026/10/10 09:58:50 [INFO] Customer demo-felhom controller-version floor override set to "0.305.0" (declared MinAgent "0.131.0") +2026/10/10 09:58:51 [INFO] Customer tester-1 controller-version floor override set to "0.305.0" (declared MinAgent "0.131.0") diff --git a/documentation/audits/release-2026-10-10/hub-0.145.0-checks.txt b/documentation/audits/release-2026-10-10/hub-0.145.0-checks.txt new file mode 100644 index 00000000..f7f48928 --- /dev/null +++ b/documentation/audits/release-2026-10-10/hub-0.145.0-checks.txt @@ -0,0 +1,23 @@ +## 2026-10-10T07:52:30Z after the deploy +GET /health: 302 +GET /system: 200 26979 bytes; MAIL-HOLD banner on the page: False +MAIL-HOLD marker in /data: absent +hub log MAIL-HOLD lines since start: 0 +GET /healthz: 200 +## live check of the hole (each request sends that household's CURRENT settings unchanged; the key is never printed) +control (own household): POST /api/v1/preferences customer_id=tester-1 with Tester 1's key -> 200 {"status":"ok"} +TEST (another household): POST /api/v1/preferences customer_id=demo-felhom with Tester 1's key -> 403 Forbidden: customer_id does not match the key +RESULT: PASS — another household's settings refused, own accepted +## the stored rows after (hub DB copy with -wal/-shm, read-only, shredded) — compared with the baseline taken before the deploy +demo-felhom row present sha256: 5ef8fc1ff15b +tester-1 row present sha256: 38de0bad97e1 +demo-felhom UNCHANGED +tester-1 CHANGED +hub log for the two requests: +2026/10/10 09:52:48 [INFO] Notification preferences updated for tester-1: address set=true, events=[null] +## 2026-10-10T07:53:23Z CORRECTION: the control request stored tester-1's events as ["null"] (the baseline script misread the stored JSON null as a list holding the word null; no real event became enabled). Restored with the same own-household request carrying enabled_events: null +restore POST (own household) -> 200 +demo-felhom row present sha256: 5ef8fc1ff15b +tester-1 row present sha256: 5583ffee1d37 +demo-felhom UNCHANGED vs the pre-deploy baseline +tester-1 UNCHANGED vs the pre-deploy baseline diff --git a/documentation/audits/release-2026-10-10/hub-0.145.0-deploy.txt b/documentation/audits/release-2026-10-10/hub-0.145.0-deploy.txt new file mode 100644 index 00000000..20b91a34 --- /dev/null +++ b/documentation/audits/release-2026-10-10/hub-0.145.0-deploy.txt @@ -0,0 +1,8 @@ +2026-10-10T07:51:41Z +operation: Succeeded +Deployment/hub Synced deployment.apps/hub configured +Waiting for deployment "hub" rollout to finish: 0 of 1 updated replicas are available... +deployment "hub" successfully rolled out +hub-d9b69dc69-cthcb gitea.dooplex.hu/admin/felhom-hub:0.145.0 ready=true +2026/10/10 09:51:47 [INFO] felhom-hub 0.145.0 starting +2026/10/10 09:52:15 [INFO] Listening on :8080 diff --git a/documentation/audits/release-2026-10-10/system-page-kernel.txt b/documentation/audits/release-2026-10-10/system-page-kernel.txt new file mode 100644 index 00000000..0ada9890 --- /dev/null +++ b/documentation/audits/release-2026-10-10/system-page-kernel.txt @@ -0,0 +1,5 @@ +kernel candidate block: kernel : 2 packages, first seen 2026-10-10 07:52 UTC Approve kernel set +approve-kernel form on the page: 1 +offered candidate (first seen today): {'fingerprint': '3aaea2861453e986', 'first_seen': '2026-10-10 02:40:08', 'packages_json': '[{"name":"containerd.io","version":"2.4.1-2~debian.13~trixie","origin":"Docker"},{"name":"docker-buildx-plugin","version":"0.38.0-1~debian.13~trixie","origin":"Docker"},{"name":"docker-ce","version":"5:29.9.0-1~debian.13~trixie","origin":"Docker"},{"name":"docker-ce-cli","version":"5:29.9.0-1~debian.13~trixie","origin":"Docker"},{"name":"docker-ce-rootless-extras","version":"5:29.9.0-1~debian.13~trixie","origin":"Docker"},{"name":"docker-compose-plugin","version":"5.6.0-1~debian.13~trixie","origin":"Docker"}]'} +offered candidate (first seen today): {'fingerprint': '7e9f2c2b988d3af7', 'first_seen': '2026-10-10 02:42:08', 'packages_json': '[{"name":"apparmor","version":"4.1.1-pmx1","origin":"Proxmox"},{"name":"ceph-common","version":"19.2.6-pve4","origin":"Proxmox"},{"name":"ceph-fuse","version":"19.2.6-pve4","origin":"Proxmox"},{"name":"chrony","version":"4.8-4~bpo13+2","origin":"Proxmox"},{"name":"corosync","version":"3.1.10-pve3","origin":"Proxmox"},{"name":"dmeventd","version":"2:1.02.205-2+pmx1","origin":"Proxmox"},{"name":"dmsetup","version":"2:1.02.205-2+pmx1","origin":"Proxmox"},{"name":"fonts-font-logos","version":"1.0.1-3","origin":"Proxmox"},{"name":"frr","version":"10.6.1-1+pve3","origin":"Proxmox"},{"name":"frr-pythontools","version":"10.6.1-1+pve3","origin":"Proxmox"},{"name":"ifupdown2","version":"3.3.0-1+pmx12","origin":"Proxmox"},{"name":"ksm-control-daemon","version":"1.5-1","origin":"Proxmox"},{"name":"libapparmor1","version":"4.1.1-pmx1","origin":"Proxmox"},{"name":"libcephfs2","version":"19.2.6-pve4","origin":"Proxmox"},{"name":"libcfg7","version":"3.1.10-pve3","origin":"Proxmox"},{"name":"libcmap4","version":"3.1.10-pve3","origin":"Proxmox"},{"name":"libcorosync-common4","version":"3.1.10-pve3","origin":"Proxmox"},{"name":"libcpg4","version":"3.1.10-pve3","origin":"Proxmox"},{"name":"libcrypt-openssl-rsa-perl","version":"0.35-1.1","origin":"Proxmox"},{"name":"libdevmapper-event1.02.1","version":"2:1.02.205-2+pmx1","origin":"Proxmox"},{"name":"libdevmapper1.02.1","version":"2:1.02.205-2+pmx1","origin":"Proxmox"},{"name":"libjs-extjs","version":"7.0.0-7","origin":"Proxmox"},{"name":"libjs-qrcodejs","version":"1.20230525-pve1","origin":"Proxmox"},{"name":"libknet1t64","version":"1.35-pve2","origin":"Proxmox"},{"name":"liblvm2cmd2.03","version":"2.03.31-2+pmx1","origin":"Proxmox"},{"name":"libnozzle1t64","version":"1.35-pve2","origin":"Proxmox"},{"name":"libnss-systemd","version":"257.13-1~deb13u1","origin":"Proxmox"},{"name":"libnvpair3linux","version":"2.4.4-pve1","origin":"Proxmox"},{"name":"libpam-systemd","version":"257.13-1~deb13u1","origin":"Proxmox"},{"name":"libproxmox-acme-perl","version":"1.7.2","origin":"Proxmox"},{"name":"libproxmox-acme-plugins","version":"1.7.2","origin":"Proxmox"},{"name":"libproxmox-backup-qemu0","version":"2.0.3","origin":"Proxmox"},{"name":"libproxmox-rs-perl","version":"0.4.1","origin":"Proxmox"},{"name":"libpve-access-control","version":"9.1.2","origin":"Proxmox"},{"name":"libpve-apiclient-perl","version":"3.4.3","origin":"Proxmox"},{"name":"libpve-cluster-api-perl","version":"9.1.6","origin":"Proxmox"},{"name":"libpve-cluster-perl","version":"9.1.6","origin":"Proxmox"},{"name":"libpve-common-perl","version":"9.2.3","origin":"Proxmox"},{"name":"libpve-guest-common-perl","version":"6.0.5","origin":"Proxmox"},{"name":"libpve-http-server-perl","version":"6.0.5","origin":"Proxmox"},{"name":"libpve-network-api-perl","version":"1.6.7","origin":"Proxmox"},{"name":"libpve-network-perl","version":"1.6.7","origin":"Proxmox"},{"name":"libpve-notify-perl","version":"9.1.6","origin":"Proxmox"},{"name":"libpve-rs-perl","version":"0.15.3","origin":"Proxmox"},{"name":"libpve-storage-perl","version":"9.1.12","origin":"Proxmox"},{"name":"libquorum5","version":"3.1.10-pve3","origin":"Proxmox"},{"name":"librados2","version":"19.2.6-pve4","origin":"Proxmox"},{"name":"librados2-perl","version":"1.5.1","origin":"Proxmox"},{"name":"libradosstriper1","version":"19.2.6-pve4","origin":"Proxmox"},{"name":"librbd1","version":"19.2.6-pve4","origin":"Proxmox"},{"name":"librgw2","version":"19.2.6-pve4","origin":"Proxmox"},{"name":"librrd8t64","version":"1.7.2-4.2+pve4","origin":"Proxmox"},{"name":"librrds-perl","version":"1.7.2-4.2+pve4","origin":"Proxmox"},{"name":"libsystemd-shared","version":"257.13-1~deb13u1","origin":"Proxmox"},{"name":"libsystemd0","version":"257.13-1~deb13u1","origin":"Proxmox"},{"name":"libtpms0","version":"0.9.7+pve2","origin":"Proxmox"},{"name":"libudev1","version":"257.13-1~deb13u1","origin":"Proxmox"},{"name":"libuutil3linux","version":"2.4.4-pve1","origin":"Proxmox"},{"name":"libvotequorum8","version":"3.1.10-pve3","origin":"Proxmox"},{"name":"libzfs7linux","version":"2.4.4-pve1","origin":"Proxmox"},{"name":"libzpool7linux","version":"2.4.4-pve1","origin":"Proxmox"},{"name":"lvm2","version":"2.03.31-2+pmx1","origin":"Proxmox"},{"name":"lxc-pve","version":"7.0.0-2","origin":"Proxmox"},{"name":"lxcfs","version":"7.0.0-pve1","origin":"Proxmox"},{"name":"novnc-pve","version":"1.7.0-2","origin":"Proxmox"},{"name":"proxmox-archive-keyring","version":"4.0","origin":"Proxmox"},{"name":"proxmox-backup-client","version":"4.2.8-1","origin":"Proxmox"},{"name":"proxmox-backup-file-restore","version":"4.2.8-1","origin":"Proxmox"},{"name":"proxmox-backup-restore-image","version":"1.0.0","origin":"Proxmox"},{"name":"proxmox-enterprise-support-keyring","version":"1.1","origin":"Proxmox"},{"name":"proxmox-firewall","version":"1.2.3","origin":"Proxmox"},{"name":"proxmox-firewall-data","version":"0.1","origin":"Proxmox"},{"name":"proxmox-first-boot","version":"9.2.8","origin":"Proxmox"},{"name":"proxmox-grub","version":"2.12-9+pmx2","origin":"Proxmox"},{"name":"proxmox-mail-forward","version":"1.0.3","origin":"Proxmox"},{"name":"proxmox-mini-journalreader","version":"1.7","origin":"Proxmox"},{"name":"proxmox-offline-mirror-docs","version":"0.7.4","origin":"Proxmox"},{"name":"proxmox-offline-mirror-helper","version":"0.7.4","origin":"Proxmox"},{"name":"proxmox-secure-boot-support","version":"2.0.6","origin":"Proxmox"},{"name":"proxmox-termproxy","version":"2.1.0","origin":"Proxmox"},{"name":"proxmox-ve","version":"9.2.0","origin":"Proxmox"},{"name":"proxmox-websocket-tunnel","version":"1.0.0","origin":"Proxmox"},{"name":"proxmox-widget-toolkit","version":"5.2.10","origin":"Proxmox"},{"name":"pve-cluster","version":"9.1.6","origin":"Proxmox"},{"name":"pve-container","version":"6.1.14","origin":"Proxmox"},{"name":"pve-docs","version":"9.2.13","origin":"Proxmox"},{"name":"pve-edk2-firmware","version":"4.2026.08-1","origin":"Proxmox"},{"name":"pve-edk2-firmware-aarch64","version":"4.2026.08-1","origin":"Proxmox"},{"name":"pve-edk2-firmware-legacy","version":"4.2026.08-1","origin":"Proxmox"},{"name":"pve-edk2-firmware-ovmf","version":"4.2026.08-1","origin":"Proxmox"},{"name":"pve-esxi-import-tools","version":"1.0.1","origin":"Proxmox"},{"name":"pve-firewall","version":"6.0.6","origin":"Proxmox"},{"name":"pve-ha-manager","version":"5.2.5","origin":"Proxmox"},{"name":"pve-i18n","version":"3.10.0","origin":"Proxmox"},{"name":"pve-lxc-syscalld","version":"2.0.2","origin":"Proxmox"},{"name":"pve-manager","version":"9.2.21","origin":"Proxmox"},{"name":"pve-nvidia-vgpu-helper","version":"0.3.1","origin":"Proxmox"},{"name":"pve-qemu-kvm","version":"11.0.3-4","origin":"Proxmox"},{"name":"pve-xtermjs","version":"6.0.0-2","origin":"Proxmox"},{"name":"pve-yew-mobile-gui","version":"0.8.0","origin":"Proxmox"},{"name":"pve-yew-mobile-i18n","version":"3.10.0","origin":"Proxmox"},{"name":"python3-ceph-argparse","version":"19.2.6-pve4","origin":"Proxmox"},{"name":"python3-ceph-common","version":"19.2.6-pve4","origin":"Proxmox"},{"name":"python3-cephfs","version":"19.2.6-pve4","origin":"Proxmox"},{"name":"python3-rados","version":"19.2.6-pve4","origin":"Proxmox"},{"name":"python3-rbd","version":"19.2.6-pve4","origin":"Proxmox"},{"name":"python3-rgw","version":"19.2.6-pve4","origin":"Proxmox"},{"name":"qemu-server","version":"9.2.10","origin":"Proxmox"},{"name":"rrdcached","version":"1.7.2-4.2+pve4","origin":"Proxmox"},{"name":"smartmontools","version":"7.5-pve2","origin":"Proxmox"},{"name":"spiceterm","version":"3.4.2","origin":"Proxmox"},{"name":"swtpm","version":"0.8.0+pve3","origin":"Proxmox"},{"name":"swtpm-libs","version":"0.8.0+pve3","origin":"Proxmox"},{"name":"swtpm-tools","version":"0.8.0+pve3","origin":"Proxmox"},{"name":"systemd","version":"257.13-1~deb13u1","origin":"Proxmox"},{"name":"systemd-sysv","version":"257.13-1~deb13u1","origin":"Proxmox"},{"name":"udev","version":"257.13-1~deb13u1","origin":"Proxmox"},{"name":"vncterm","version":"1.9.2","origin":"Proxmox"},{"name":"zfs-initramfs","version":"2.4.4-pve1","origin":"Proxmox"},{"name":"zfs-zed","version":"2.4.4-pve1","origin":"Proxmox"},{"name":"zfsutils-linux","version":"2.4.4-pve1","origin":"Proxmox"}]'} +offered candidate (first seen today): {'fingerprint': 'd6f354ee7bc6d5f9', 'first_seen': '2026-10-10 07:52:31', 'packages_json': '[{"name":"proxmox-kernel-7.0","version":"7.0.14-22","origin":"Proxmox Debian Repository"},{"name":"proxmox-kernel-7.0.14-22-pve-signed","version":"7.0.14-22","origin":"Proxmox Debian Repository"}]'} diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index d6dcd123..74d64f0f 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -166,7 +166,7 @@ stopping line that lies. | **R-540** | Backup & restore | P3 | **[P3-LOW] The hub knows exactly ONE off-site pool box, so there is no rule for what happens when it fills.** Read from source 2026-09-16 while making off-site the default: `HETZNER_POOL_BOX_ID` is a single value, and every shared customer becomes a sub-account on that box. With off-site now ON for every new customer (hub v0.116.0) the box fills faster, and the fill warning (80%/90% of the box, `monitor/offsite.go`) tells the operator it is filling but nothing says which box a new customer should land on. **Needs a selection rule** (least-full, or explicit per-customer), not a bigger box. No customer is at risk today: the pool box read 0.3% full (2.7 GB of 1 TB), Σ shared quota 150 GB, oversub 0.15x. | **READY — rank P3-LOW; owner: CC (hub)** **2026-10-05 (burn-down night): NEEDS A DESIGN (and money).** A selection rule needs a multi-box configuration and a second box. Next: an operator pick of the rule and its trigger (e.g. add box 2 at 70 %). | — | — | CC | | **R-548** | Backup & restore | P3 | **[P3-LOW] The whole-guest backup’s LOCAL tier cannot fit on a small-system-disk box, and will retry on that tier for ever.** MEASURED 2026-09-17 (chaos night) on `tester-1-022354`: a whole-guest backup wrote a **~29 GB source** (`mp0` = `local-lvm:vm-9201-disk-1`, 70 G provisioned, 40.58 % used, `backup=1`) into `pve-root`, which on a 32 GB system disk is **14 GB total with ~4.9 GB free**. Two samples thirty seconds apart showed the archive growing ~497 MB while free space fell ~475 MB — **~16 MB/s, i.e. under four minutes to a full `/`** on the nested PVE. **The product’s behaviour is correct and legible throughout:** it failed the tier and said which one — `whole_guest_backup_failed` (error, operator-only): „Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s)” — its status surface agreed (`target_id:"local"`, `success:false`, `size_bytes:0`), and the **off-site tier then ran from the same snapshot and succeeded in ~8½ minutes, encrypted to ep0, consuming no local disk at all**. So the data still left the house. What is filed is the loop: on a box shaped like this the local tier can **never** succeed, and it keeps retrying on a backoff for ever, burning I/O and risking `/` each time. **Honest caveat:** the 32 GB system disk is this drill’s fixture choice, so the row is conditional on disk size — but nothing in the product checks whether the local target could ever hold the source before trying. **Fix shape:** compare the source size against the target’s free space before starting the local tier and skip it with a clear reason, rather than discovering it at ~16 MB/s. Evidence: `audits/evidence-chaos-night-2026-09-17/round-6.txt`. **-- 2026-09-30, the class on a real box: demo-hp's whole-guest LOCAL tier was refused by the space preflight (R-685) 10 times since 2026-09-27** („local has 12.1 GiB free; the last archive of guest 9201 was 8.9 GiB, so a new one needs about 12.1 GiB") — its newest local archive was 2026-09-29 04:42, 30 h old at the day's read (the weekly PBS tier, last 2026-09-24, is its whole copy meanwhile). The refusal is correct and named; what is missing is that nothing gives the box the room back (the host's `local` holds the 8.9 GiB archive of the only guest plus templates on a 39 GB root). `audits/pg-last-six-2026-09-30/C/`. | **READY — rank P3-LOW; owner: CC** | — | — | CC | | **R-698** | Backup & restore | P3 | **[P3-LOW] A backup stores the image's NAME, not the image — a restore of a version its maker has deleted cannot start.** `RecoveryManifest.image_pins` ("image NOT stored — re-pulled on restore"); since controller v0.275.0 each data file also records its running `ref@digest`, and a restore brings the data back AT ITS OWN VERSION (`07` §6.6) — so a restore asks for exactly the old image. **Measured 2026-09-26** (`audits/version-travel-2026-09-26/A7/`, registry HEADs, no pulls): the catalog's 42 ladder `ref@digest` pairs all resolve (200); an invented digest answers 404 on Docker Hub and ghcr.io (negative control). Not measured: the digests recorded on boxes (older than any ladder entry), how often makers delete versions, the catalog's 66 digest-less compose lines. **Options (decide nothing yet):** (a) keep — a restore of a deleted version fails at the pull and the household uses the next copy or a newer version; (b) mirror every INSTALLED image into the DooPlex registry, restore falls back to it — storage + bandwidth on DooPlex, a new part on the recovery path; (c) mirror only ladder-named versions — bounded, misses pre-ladder boxes; (d) `docker save` into the unit — hundreds of MB per app per copy on every tier. **-- 2026-09-30 late (decision 53):** a box now keeps only an app's running and previous image; a restore to an older version re-pulls it — as every restore already did. The limit above is unchanged. | **OPEN — P3; owner: operator (a decision), CC measures** **Operator ruling 2026-10-05 18:23: kept OPEN as a known risk to a household's restore; owner the operator; not worked on in the burn-down.** | — | — | operator | -| **R-921** | Backup & restore | P3 | **When the off-site tier answers BUSY, the household still gets a short stop of every app — with no copy.** MEASURED 2026-10-08 night on demo-hp (read-back 2026-10-09, `audits/dooplex-survival-2026-10-09/partE/R-518.txt`): local tier 20:20–20:25Z (own stop, < 2 min), then at 20:26:14Z the agent refused the off-site request (`backup refused — a heavy operation is already in flight`, `busy=backup:local` — the night OS step ran 20:27–20:29Z), yet the controller's metrics show all 15 app containers down at 20:26:19Z and back at 20:27:19Z; the retry ran the off-site tier at 20:34Z with its own short stop. So a two-tier night costs THREE stops, one of them for nothing. This was R-518's noted "unmeasured" risk. Fix shape: ask the agent whether it is free (or reserve the slot) BEFORE quiescing, and resume at once on BUSY. | **NARROWED 2026-10-09 — built, not released (controller):** before stopping apps the controller asks the agent which tiers have a job in flight (`GET /backup/status`); another tier's job in flight → no stop, the tier stays due (red-proved: 3 apps stopped for a refused tier before); a refusal after the stop still resumes the apps at once (pinned). Also fixes a stop after a long upload or a restart mid-upload. **LEFT: the measured case itself** — the agent's host-wide busy lock (the night OS step, restore test, fstrim) is served on NO endpoint, so the controller cannot see it before stopping; needs one agent field (e.g. `busy` on `GET /backup/status`) + a controller pre-check behind a feature probe. `07` §6.4. | next controller release; an agent change | Release the controller; build the agent `busy` field + its controller check; read back a two-tier night | CC | +| **R-921** | Backup & restore | P3 | **When the off-site tier answers BUSY, the household still gets a short stop of every app — with no copy.** MEASURED 2026-10-08 night on demo-hp (read-back 2026-10-09, `audits/dooplex-survival-2026-10-09/partE/R-518.txt`): local tier 20:20–20:25Z (own stop, < 2 min), then at 20:26:14Z the agent refused the off-site request (`backup refused — a heavy operation is already in flight`, `busy=backup:local` — the night OS step ran 20:27–20:29Z), yet the controller's metrics show all 15 app containers down at 20:26:19Z and back at 20:27:19Z; the retry ran the off-site tier at 20:34Z with its own short stop. So a two-tier night costs THREE stops, one of them for nothing. This was R-518's noted "unmeasured" risk. Fix shape: ask the agent whether it is free (or reserve the slot) BEFORE quiescing, and resume at once on BUSY. | **NARROWED 2026-10-09 — built; RELEASED 2026-10-10 (controller 0.305.0 on demo-hp, demo-felhom, Tester 1):** before stopping apps the controller asks the agent which tiers have a job in flight (`GET /backup/status`); another tier's job in flight → no stop, the tier stays due (red-proved: 3 apps stopped for a refused tier before); a refusal after the stop still resumes the apps at once (pinned). Also fixes a stop after a long upload or a restart mid-upload. **LEFT: the measured case itself** — the agent's host-wide busy lock (the night OS step, restore test, fstrim) is served on NO endpoint, so the controller cannot see it before stopping; needs one agent field (e.g. `busy` on `GET /backup/status`) + a controller pre-check behind a feature probe. `07` §6.4. | an agent change | Build the agent `busy` field + its controller check; read back a two-tier night | CC | | **R-91** | Backup & restore | P4 | Old 13 GB datastore copy at `/srv/pbs-felhom` on ep0's root disk **Checked from source 2026-10-05 (burn-down round 2):** Gate is long past (row waits on demo-felhom's first post-migration PBS backup, migration 2026-07-27). Last positive record of the copy: audits/CAMPAIGN-9-restore-proof-2026-07-28.md:759 'ep0 : /srv/pbs-felhom rollback copy intact (13G)'; CONTEXT.md:3666 still says it is 13 G of dead weight awaiting R-91. No later record of deletion found (grep srv/pbs-felhom across felhom.eu). Deleting is on ep0 (protected) and needs an operator word. | WATCHING | demo-felhom's first **post-migration** PBS backup | Delete once it lands; fix `CONTEXT.md:1018` same commit | CC | | **R-164** | Backup & restore | P4 | **C2's chain: the DB volume tar cannot be dropped until a SOUND dump predicate exists.** The unit carries both a volume tar and a SQL dump; the restore uses **both** — the dump is authoritative and replayed *after* the tar so it WINS (F17), with only the DB service up (R-47) — `internal/backup/restore_unit.go:262-266`. Dropping the DB container's tar would halve DB-app units **and** close the R-127(b) initdb-skip password trap (restored PGDATA ⇒ `POSTGRES_PASSWORD` ignored). | **BLOCKED** — on the predicate | a dump-validity predicate that is not `accounts has rows` | **The obvious gate is DEAD, measured:** `ValidateDump` warns when the `accounts` table is empty, and that warning was **correct** — the live DB genuinely had 0 accounts, and seeding one stopped the warning and put the row in the dump. But **a fresh appliance legitimately has zero accounts**, so promoting that predicate to a gate would **block every new customer's first backup**. Order: (1) a sound predicate — dump vs **live** per-table counts, not an absolute expectation; (2) warn→gate; (3) tar-drop. **Until (1), the tar is load-bearing** — not because dumps are bad, but because nothing can yet prove one is good. Pairs with **R-127** | CC | | **R-213** | Backup & restore | P4 | **Putting files back in place — the half the recovery screen deliberately does not do.** The screen (R-193, controller v0.200.0) unlocks the repository and LISTS what is in it; restoring is per-app and lives in the backups area, and the operator ruled the two separate on 2026-08-05: *a screen that unlocks and then offers to overwrite is two decisions wearing one button*. What is missing is the step after the listing — a customer who can now SEE their files still has to work out, per app, which restore to choose. **The operator named its requirement: a live-versus-backup comparison** — the customer must be able to see what would change before anything is overwritten | **OPEN — not started, deliberately** | the comparison design (nothing exists for it yet) | Design the live-vs-backup comparison, then the put-back flow on top of it. Do NOT fold it into the recovery screen | Operator + CC | @@ -208,7 +208,7 @@ stopping line that lies. | **R-870** | Security & access | P3 | **Tester 1's two Cloudflare credentials — the zone API token (`infrastructure.cf_api_token`) and the tunnel token (`infrastructure.cf_tunnel_token`) of the hub's `customer_configs` row `tester-1` — were printed into the 2026-10-04 night session's transcript** (not into any file): a read-only query selected `substr(config_json,1,400)`, and both values sit in the first 400 characters. Tester 1 is CC's disposable test customer (`enkicsifelhom.hu`). **Not rotated, by the operator's ruling of 2026-10-05 06:49 (option B).** **Rotation, whenever chosen (3 steps):** in the Cloudflare dashboard create a new API token for the `enkicsifelhom.hu` zone with the same permissions and refresh the Tester 1 tunnel's token (Zero Trust → Networks → Tunnels → the tunnel → refresh token) → hub → Configs → `tester-1` → Edit → the two Cloudflare fields → Save, then confirm on the box that cloudflared reconnected (`docker ps` health `healthy`) → delete the old API token. Rule for sessions (as R-831): never select a whole config row — name the fields, and never `config_json` without `json_extract` of a non-secret field. | **WAITING-ON-OPERATOR — rotation is his call (ruled: not now)** **Not rotated by the operator's rulings (2026-10-04 „keep using the current one"; 2026-10-05 option B) — restated 2026-10-05 18:23; the steps stay here.** | — | rotate when chosen | operator | | **R-908** | Security & access | P3 | **An old Resend API key is still in the history of the `homelab-manifests` repository; it was replaced on 2026-06-29, but whether it is also revoked at Resend is unknown.** FOUND 2026-10-08 by the contact-mailer source search (`audits/mailer-source-2026-10-08/SEARCH.md`, „A side finding"): commit c9648cd (2026-02-05, „added mailer pod") put a key literal in the old `felhom-system/contact-mailer.yaml` comment; the file was deleted in ee93b50 but the history keeps it. Compared by hash only, never printed: it is NOT today's key (`Secret/resend-api`, rotated 2026-06-29, felhom.eu feea0606). If the old key still works at Resend, anyone with read access to that repository can send mail as felhom.eu. | **OPEN — owner: operator** | — | In the Resend console, check that the old key is revoked (revoke it if not); rewriting the repository's history is NOT needed once it is revoked | operator | | **R-525** | Security & access | P4 | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** **Re-ranked 2026-10-03: P3->P4: comfort feature needing a new unmeasured mechanism; the default password hole is closed.** | — | — | CC | -| **R-925** | Security & access | P1 | **`manifests/felhom.secret.yaml` carries REAL secret values, and the Gitea repo that holds it is readable by anyone on the internet with no login.** FOUND 2026-10-09 by a background security review of the legal-pages commit, which flagged the served `` comments; chasing what those comments pointed at found this instead. **MEASURED, in this order:** (1) `scripts/manifest_bearer_gate.py` itself prints `manifests/felhom.secret.yaml:39 KNOWN-BACKLOG committed secret 65cee3c4…7a86`; (2) the file holds **live-shaped values, not placeholders** — `SECRET_KEY` (69 chars), `SUPERUSER_EMAIL`/`SUPERUSER_PASSWORD` (18), Umami `APP_SECRET` (66) and `POSTGRES_PASSWORD` (34), and a `username`/`password` pair (18); values were never printed, only measured by length; (3) `gitea.dooplex.hu` resolves to **37.191.56.193**, the same public address as the website; (4) an **anonymous** `curl` returns HTTP 200 and 1686 bytes for that exact file, and 200 for `hub/internal/store/store.go` and `manifests/hub.yaml`; (5) **the off-network control**, which is what makes (4) mean anything — fetched from outside the operator's network entirely: a real path returns the file's first line, a nonsense path returns 404, so the 200 is genuine anonymous read from the internet and not a LAN-only ACL. **This contradicts the project's own written model**: `documentation/runbooks/secrets.md:3-4` says *“Secret values are never committed to git. The manifests in `manifests/` carry only placeholders + comments that point here.”* That promise is false today, which is the P1 wording in this register's own scale — see the ranking note below. The same runbook already states the rule that governs the fix: *“The git-history copy stays alive until the value is ROTATED — de-git alone kills nothing.”* **Known neighbours, none of which cover this:** R-887 records that an outside crawler walks the public Gitea pages and leaves *“Gitea's exposure to the crawler”* as the operator's undone item; R-580 still owes a Gitea admin token rotation. Neither says a secret file is world-readable. **Why P2 and not P1, stated so the operator can overrule it:** the exposed credentials guard **analytics and an undeployed healthchecks instance, not household data** — measured: `umami-db` is a **ClusterIP** service on 5432 with no external IP, so the Postgres password is not reachable from the internet; the realistic harm is session forgery against the public `stats.felhom.eu` via `APP_SECRET`, and reuse of those passwords anywhere else. No customer box, hub token or escrow key is in this file. **CC did NOT change anything**: making the repo private could break the public day-0 path (the installer is fetched from a public tag, R-110), and rotation plus repository visibility are operator decisions on production infrastructure. **PARTLY FIXED 2026-10-09, and RE-RANKED P2 -> P1 on what it turned out to be.** The exposed value was not an analytics password: it was the **Gitea `admin` account password** (`is_admin: true`; `/api/v1/admin/users` answered 200), published in a repo `gitea.dooplex.hu` serves anonymously to the internet. That is push access to every repo — including the one whose `website/` is git-synced live and whose `scripts/` is published by tag to **every new box installer** (R-110), i.e. a supply-chain path onto customer hardware. My first ranking said P2 because I had only measured the analytics blast radius; the credential test is what corrected it. **DONE (CC, operator-instructed):** (1) `umami-config` `APP_SECRET` + `POSTGRES_PASSWORD` rotated — the password changed **inside Postgres** too (`ALTER USER`), because the env var is only read at first init; verified by a real beacon returning 200 with a bogus-site-id control still returning 400. (2) `healthchecks-config` `SECRET_KEY` + `SUPERUSER_PASSWORD` rotated (nothing consumes them — there is no healthchecks Deployment). (3) `gitea-creds` no longer holds the admin password at all: it holds a **scoped token** (`read:package` + `read:repository`). **Both scopes are load-bearing and the second was found by breaking it** — a package-only token made the hub log `Template fetch: unexpected status 403`, because the template fetcher reads a raw file out of the `felhom-controller` repo. (4) The Gitea **admin password** was changed and `gitea-system/gitea-admin` updated; **verified: new password 200, old published password 401 — the leak is dead.** (5) The file is removed from git and `.gitignore`'s `*secret*` now applies to it; `manifest_bearer_gate.py`'s `KNOWN_BACKLOG` carve-out is gone, red-proofed with a decoy (exit 1 with, 0 without). **AN INCIDENT CAUSED BY THE FIX, recorded because it is the useful part:** the `rollout restart` needed to pick up the new umami secret put umami into **CrashLoopBackOff and took `stats.felhom.eu` down (503)**. Cause was not the rotation — at `memory: 512Mi` the pod runs for months but **cannot restart**, because startup (Prisma + Next.js) peaks over the limit and is OOMKilled (exit 137). Raised to 1Gi, **in the manifest, not just live** (the `.claude/rules/manifests.md` rule: a bare `kubectl set` is reverted by the next sync and the fix silently disappears). Service restored and verified. New values are in a 0600 file on DooPlex (`~/rotated-secrets-2026-10-09.txt`), never echoed; the operator moves them to the password manager and deletes it. **WHAT REMAINS, and it is bigger than what was fixed:** that one password is still the admin password in about **12 other namespaces** — `nextcloud`, `paperless`, `bookstack` (a **database root** password), `tandoor`, `calibre`, `adventurelog`, `gokapi`, `qbittorrent` (x2), `servarr`, `homepage` (x2). Rotating Gitea does not touch them: each is its own login and each is still the string that was published. **Also owed:** a Gitea token for user `kisfenyo` sits in plaintext in the `origin` URL of the local `homelab-manifests` clone (`.git/config`); CC printed it to a session transcript while investigating, so it should be rotated regardless — this is R-580's shape (store the remote without credentials). **Not a finding:** `homelab-manifests` itself is private (404 anonymously) and does not contain the password; ArgoCD's repo credential is a separate token, so the rotation did not touch it. | **NARROWED 2026-10-09 — the felhom.eu half is DONE and verified (Gitea admin password rotated, old one now 401; umami + healthchecks rotated; file de-gitted; gate carve-out removed). What remains is the ~12 reused logins elsewhere in the homelab and the repo-visibility decision; owner: operator** | — | **OPERATOR RULED 2026-10-09: (a) he rotates the remaining services himself — CC's scope stopped at the Felhom boundary; (b) the repo STAYS anonymously readable for now**, because making it private breaks the website git-sync and the installer tag fetch, which both clone with no credentials (R-110). The standing consequence of (b): no secret may ever enter this repo again, which `manifest_bearer_gate.py` now enforces with no exemption; and if (b) is reversed, the 32 `` comments on /adatkezeles must be stripped in the same change. **The exact checklist of the 12 remaining secrets** (namespace / secret / key, re-measured after the rotation, with the two traps that bite — a DB password is not changed by editing the Secret, and a pod that has run for months may not restart) is in `documentation/runbooks/secrets.md`. Operator: **(1)** work that checklist, `bookstack-db/root-password` first because it is a database root password and every one of these hosts answers on the public internet; (`bookstack` first — it is a database root password); **(2)** rotate the `kisfenyo` Gitea token embedded in the local `homelab-manifests` remote URL, and store the remote without credentials; **(3)** decide whether `gitea.dooplex.hu` should answer anonymously at all — remembering the website git-sync clones it with no credentials and the installer is fetched from a public tag, so making it private breaks both unless they get credentials first, and the 32 `` comments on /adatkezeles become a map of it; **(4)** move `~/rotated-secrets-2026-10-09.txt` into the password manager and delete it. If nothing is done: the published password keeps opening a dozen services, even though Gitea itself is now safe | operator | +| **R-925** | Security & access | P1 | **`manifests/felhom.secret.yaml` carries REAL secret values, and the Gitea repo that holds it is readable by anyone on the internet with no login.** FOUND 2026-10-09 by a background security review of the legal-pages commit, which flagged the served `` comments; chasing what those comments pointed at found this instead. **MEASURED, in this order:** (1) `scripts/manifest_bearer_gate.py` itself prints `manifests/felhom.secret.yaml:39 KNOWN-BACKLOG committed secret 65cee3c4…7a86`; (2) the file holds **live-shaped values, not placeholders** — `SECRET_KEY` (69 chars), `SUPERUSER_EMAIL`/`SUPERUSER_PASSWORD` (18), Umami `APP_SECRET` (66) and `POSTGRES_PASSWORD` (34), and a `username`/`password` pair (18); values were never printed, only measured by length; (3) `gitea.dooplex.hu` resolves to **37.191.56.193**, the same public address as the website; (4) an **anonymous** `curl` returns HTTP 200 and 1686 bytes for that exact file, and 200 for `hub/internal/store/store.go` and `manifests/hub.yaml`; (5) **the off-network control**, which is what makes (4) mean anything — fetched from outside the operator's network entirely: a real path returns the file's first line, a nonsense path returns 404, so the 200 is genuine anonymous read from the internet and not a LAN-only ACL. **This contradicts the project's own written model**: `documentation/runbooks/secrets.md:3-4` says *“Secret values are never committed to git. The manifests in `manifests/` carry only placeholders + comments that point here.”* That promise is false today, which is the P1 wording in this register's own scale — see the ranking note below. The same runbook already states the rule that governs the fix: *“The git-history copy stays alive until the value is ROTATED — de-git alone kills nothing.”* **Known neighbours, none of which cover this:** R-887 records that an outside crawler walks the public Gitea pages and leaves *“Gitea's exposure to the crawler”* as the operator's undone item; R-580 still owes a Gitea admin token rotation. Neither says a secret file is world-readable. **Why P2 and not P1, stated so the operator can overrule it:** the exposed credentials guard **analytics and an undeployed healthchecks instance, not household data** — measured: `umami-db` is a **ClusterIP** service on 5432 with no external IP, so the Postgres password is not reachable from the internet; the realistic harm is session forgery against the public `stats.felhom.eu` via `APP_SECRET`, and reuse of those passwords anywhere else. No customer box, hub token or escrow key is in this file. **CC did NOT change anything**: making the repo private could break the public day-0 path (the installer is fetched from a public tag, R-110), and rotation plus repository visibility are operator decisions on production infrastructure. **PARTLY FIXED 2026-10-09, and RE-RANKED P2 -> P1 on what it turned out to be.** The exposed value was not an analytics password: it was the **Gitea `admin` account password** (`is_admin: true`; `/api/v1/admin/users` answered 200), published in a repo `gitea.dooplex.hu` serves anonymously to the internet. That is push access to every repo — including the one whose `website/` is git-synced live and whose `scripts/` is published by tag to **every new box installer** (R-110), i.e. a supply-chain path onto customer hardware. My first ranking said P2 because I had only measured the analytics blast radius; the credential test is what corrected it. **DONE (CC, operator-instructed):** (1) `umami-config` `APP_SECRET` + `POSTGRES_PASSWORD` rotated — the password changed **inside Postgres** too (`ALTER USER`), because the env var is only read at first init; verified by a real beacon returning 200 with a bogus-site-id control still returning 400. (2) `healthchecks-config` `SECRET_KEY` + `SUPERUSER_PASSWORD` rotated (nothing consumes them — there is no healthchecks Deployment). (3) `gitea-creds` no longer holds the admin password at all: it holds a **scoped token** (`read:package` + `read:repository`). **Both scopes are load-bearing and the second was found by breaking it** — a package-only token made the hub log `Template fetch: unexpected status 403`, because the template fetcher reads a raw file out of the `felhom-controller` repo. (4) The Gitea **admin password** was changed and `gitea-system/gitea-admin` updated; **verified: new password 200, old published password 401 — the leak is dead.** (5) The file is removed from git and `.gitignore`'s `*secret*` now applies to it; `manifest_bearer_gate.py`'s `KNOWN_BACKLOG` carve-out is gone, red-proofed with a decoy (exit 1 with, 0 without). **AN INCIDENT CAUSED BY THE FIX, recorded because it is the useful part:** the `rollout restart` needed to pick up the new umami secret put umami into **CrashLoopBackOff and took `stats.felhom.eu` down (503)**. Cause was not the rotation — at `memory: 512Mi` the pod runs for months but **cannot restart**, because startup (Prisma + Next.js) peaks over the limit and is OOMKilled (exit 137). Raised to 1Gi, **in the manifest, not just live** (the `.claude/rules/manifests.md` rule: a bare `kubectl set` is reverted by the next sync and the fix silently disappears). Service restored and verified. New values are in a 0600 file on DooPlex (`~/rotated-secrets-2026-10-09.txt`), never echoed; the operator moves them to the password manager and deletes it. **WHAT REMAINS, and it is bigger than what was fixed:** that one password is still the admin password in about **12 other namespaces** — `nextcloud`, `paperless`, `bookstack` (a **database root** password), `tandoor`, `calibre`, `adventurelog`, `gokapi`, `qbittorrent` (x2), `servarr`, `homepage` (x2). Rotating Gitea does not touch them: each is its own login and each is still the string that was published. **Also owed:** a Gitea token for user `kisfenyo` sits in plaintext in the `origin` URL of the local `homelab-manifests` clone (`.git/config`); CC printed it to a session transcript while investigating, so it should be rotated regardless — this is R-580's shape (store the remote without credentials). **Not a finding:** `homelab-manifests` itself is private (404 anonymously) and does not contain the password; ArgoCD's repo credential is a separate token, so the rotation did not touch it. | **NARROWED 2026-10-09 — the felhom.eu half is DONE and verified (Gitea admin password rotated, old one now 401; umami + healthchecks rotated; file de-gitted; gate carve-out removed). What remains is the ~12 reused logins elsewhere in the homelab and the repo-visibility decision; owner: operator** **2026-10-10 (release session), two consequences measured:** (a) DooPlex's own Docker login for `gitea.dooplex.hu` (`~/.docker/config.json`) still holds the OLD admin password — the registry token endpoint answers 401 for it and 200 for the current one — so every image push from DooPlex fails `unauthorized` until it is refreshed; the 0.145.0/0.305.0 pushes used a one-off login in a scratch `DOCKER_CONFIG`, shredded after. (b) The `felhom` ArgoCD app reads OutOfSync for `Secret/gitea-creds`, `Secret/healthchecks-config`, `Secret/umami-config` and `Deployment/umami`: a whole-app sync would push git's (de-gitted) view over the rotated live Secrets — the hub deploy synced ONLY `Deployment/hub` (`audits/release-2026-10-10/hub-0.145.0-deploy.txt`). Refreshing the DooPlex login is a DooPlex change (operator's word). | — | **OPERATOR RULED 2026-10-09: (a) he rotates the remaining services himself — CC's scope stopped at the Felhom boundary; (b) the repo STAYS anonymously readable for now**, because making it private breaks the website git-sync and the installer tag fetch, which both clone with no credentials (R-110). The standing consequence of (b): no secret may ever enter this repo again, which `manifest_bearer_gate.py` now enforces with no exemption; and if (b) is reversed, the 32 `` comments on /adatkezeles must be stripped in the same change. **The exact checklist of the 12 remaining secrets** (namespace / secret / key, re-measured after the rotation, with the two traps that bite — a DB password is not changed by editing the Secret, and a pod that has run for months may not restart) is in `documentation/runbooks/secrets.md`. Operator: **(1)** work that checklist, `bookstack-db/root-password` first because it is a database root password and every one of these hosts answers on the public internet; (`bookstack` first — it is a database root password); **(2)** rotate the `kisfenyo` Gitea token embedded in the local `homelab-manifests` remote URL, and store the remote without credentials; **(3)** decide whether `gitea.dooplex.hu` should answer anonymously at all — remembering the website git-sync clones it with no credentials and the installer is fetched from a public tag, so making it private breaks both unless they get credentials first, and the 32 `` comments on /adatkezeles become a map of it; **(4)** move `~/rotated-secrets-2026-10-09.txt` into the password manager and delete it. If nothing is done: the published password keeps opening a dozen services, even though Gitea itself is now safe | operator | | **R-779** | Security & access | P4 | **[P3-LOW] Part A's "two outside addresses seen as two" is proven through the simulated tunnel only; on the REAL tunnel the second outside address (ep0, one request allowed) was refused by Cloudflare's edge with 403 and never reached the box.** Measured 2026-10-01 19:51 UTC (`audits/visitors-2026-10-01/A/L2-demo-hp-real-tunnel.txt`): no log line on demo-hp; demo-hp's box has no geo restriction in its settings, so a Cloudflare ZONE rule (country or bot, not read) refused a German datacenter address. DooPlex's own address on the real tunnel was seen as itself. **Needs:** one sign-in from a second Hungarian address (the operator's phone off wifi) while DooPlex is locked out — 2 minutes; and say which Cloudflare rule refused ep0. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (a phone), CC reads the logs** **Re-ranked 2026-10-03: P3→P4: a proof gap on the real tunnel; operator-only follow-up.** | — | — | CC + operator | | **R-904** | Security & access | P4 | **Cloudflare can read every household's app traffic; replacing it with our own relay is a later item.** Facts (reviewer discussion 2026-10-08; `01`): app traffic and the dashboard reach the box through the Cloudflare Tunnel (`01` §5 trust table, rows end-user ↔ apps and customer ↔ controller UI; §7), and the tunnel's public end is Cloudflare's edge, where TLS ends — so Cloudflare can technically read that traffic (the FAQ says so since 2026-10-08, R-900; „TLS ends at the edge" is not written in `01` — add it there). What Cloudflare gives today, free: inbound reach with no router setup, the CGNAT answer (`01` §4, §7); certificates (`01` §7, the free tier covers one level below a zone); the geo-WAF the hub enforces (`01` §5 last row, §7); flood protection (not in `01`). The alternative named: our own EU relay over WireGuard with TLS passthrough by SNI, certificates on the box, the geo-block on the relay. Its costs: one more machine the operator keeps up, and a single point of reach for every box; weaker flood protection; about a week of work after a spike. Operator ruling 2026-10-08 09:07 (`09` §3 decision 184): a later item. | **DEFERRED — after the first customers (operator ruling 2026-10-08 09:07)** | — | A spike after the first customers (the relay's reach, cost and flood behaviour, measured) | operator | | **R-913** | Security & access | P4 | **The Cloudflare token check (R-138, decision 190) reads what a token can SEE, not what it can WRITE.** FOUND 2026-10-08 by the security review of the R-138 build: `GET /zones` lists zones the token can read; a hand-built token with Zone:Read on the customer's zone and DNS:Edit on ALL zones would pass the check and could still change every household's DNS. The token wizard's single „Specific zone" scope does not build such a token, and a customer token cannot read its own policies (that needs „API Tokens Read"). Written as a limit in `01` §7. **Options:** (a) the operator mints every customer token from one recipe and the hub checks nothing more; (b) the hub mints the token itself with the operator's account token (one zone, DNS:Edit) — a new privileged credential on the hub; (c) a negative probe: with the pasted token, try to READ the DNS records of another customer's zone by its id (the hub knows the ids) and refuse on success — catches the read side only. | **OPEN — needs the operator** (a is free today; b is a design) | operator: which of a/b/c | Pick a/b/c; if nothing: the check stands as built and the limit stays written in `01` §7 | operator | @@ -248,7 +248,7 @@ stopping line that lies. | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| -| **R-922** | Hub & operator | P2 | **A household that clears its mail address on the dashboard does not get it cleared on the hub — the hub keeps the old address.** SEEN 2026-10-09 on Tester 1 (`audits/release-2026-10-09/proofs/D3/d3.txt`): the push with an empty address answered 200 and the hub row kept the old address (with no events, so no mail is sent). Cause: the empty-email no-clobber guard (`hub/internal/api/handler.go`, v0.71.0, audit F12) protects against an unconfigured box wiping a seeded address, and cannot tell that from a household's deliberate clear. Personal data the household removed stays on our side with no stated end — the same family as R-901's deletion rules. | **NARROWED 2026-10-09 — option A BUILT, not released** (operator ruling 11:19): the controller sends `email_cleared: true` while a household-cleared address stays empty; the hub deletes its stored notification address on that flag and keeps the F12 guard for every other empty push (red-proved both sides; the wire-contract gate now checks this push). Privacy-notice draft: one retention row. Ships with the next hub + controller release (hub first). **LEFT:** the operator-registered address `customer_configs.email` is a separate copy and stays — and the kernel notice, claim codes and self-bind mails read THAT one (`hub/internal/notify/dispatcher.go` `SendKernelNotice`), so a household that cleared its address still gets kernel notices; which address counts is a rule for the operator. | operator rule for the second copy | Release; then the operator: does a household clear also stop mails to the registered address? | CC (release), operator (rule) | +| **R-922** | Hub & operator | P2 | **A household that clears its mail address on the dashboard does not get it cleared on the hub — the hub keeps the old address.** SEEN 2026-10-09 on Tester 1 (`audits/release-2026-10-09/proofs/D3/d3.txt`): the push with an empty address answered 200 and the hub row kept the old address (with no events, so no mail is sent). Cause: the empty-email no-clobber guard (`hub/internal/api/handler.go`, v0.71.0, audit F12) protects against an unconfigured box wiping a seeded address, and cannot tell that from a household's deliberate clear. Personal data the household removed stays on our side with no stated end — the same family as R-901's deletion rules. | **NARROWED 2026-10-09 — option A BUILT; RELEASED 2026-10-10** (hub 0.145.0 deployed, controller 0.305.0 on demo-hp, demo-felhom, Tester 1; a household's clear not yet exercised live) (operator ruling 11:19): the controller sends `email_cleared: true` while a household-cleared address stays empty; the hub deletes its stored notification address on that flag and keeps the F12 guard for every other empty push (red-proved both sides; the wire-contract gate now checks this push). Privacy-notice draft: one retention row. Ships with the next hub + controller release (hub first). **LEFT:** the operator-registered address `customer_configs.email` is a separate copy and stays — and the kernel notice, claim codes and self-bind mails read THAT one (`hub/internal/notify/dispatcher.go` `SendKernelNotice`), so a household that cleared its address still gets kernel notices; which address counts is a rule for the operator. | operator rule for the second copy | The operator: does a household clear also stop mails to the registered address? CC: one live clear on a demo box | operator (rule), CC (live proof) | | **R-31** | Hub & operator | P3 | **[P2-HIGH] Offsite provisioning is synchronous with no status affordance.** Save runs the Hetzner sync in-request, so the request can hit the nginx 504 **while succeeding server-side**: the operator cannot tell failed from slow, and a retry races the first attempt. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: operator-only; a known workaround (click once, wait, verify) exists.** **2026-10-06 night: the race half fixed on felhom.eu main (hub, unreleased):** a second Save while the first still provisions is refused with 409 („already running — wait about a minute, then reload; do not save again"); nothing saved, nothing created; per customer, in memory. `TestProvision_R31_*`, red-proof `audits/night-burndown-2026-10-06/hub/R-31-red.txt`. LEFT: the async save with a status card (the escrow-card idiom). | — | Direction: make it async + a status card, reusing the proven **awaiting-card/poll idiom** (v0.138.0 escrow card). **Interim mitigation belongs in R-3 as an operator note: click once, wait, verify — do not re-click.** | CC | | **R-244** | Hub & operator | P3 | **The customer DELETE cascade leaves `app_log_issues` behind, and it is systematic across every venue ever torn down.** Found **2026-08-07** while verifying the `finalwalk` teardown with a **full census** (every table, every column) rather than a per-table query. After a cascade that logged `COMPLETE … full teardown`, **61 rows still matched `finalwalk`**. Four of the five sources are **deliberate and correct** — the cascade's own header states *"Provenance/events are NEVER wiped — audit outlives every tier"*: `events` 16, `notification_log` 14, `host_deletions` 1, `customer_resets` 1. **The fifth is a gap:** `app_log_issues` 29 rows, which the residue purge does not touch (its logged leg covers `reports`/`app_telemetry`/`app_log_tails`/`log_tail_requests`/`notif_prefs`/`selfbind_tokens`/`appliance_registrations` — not this table). **It is not a `finalwalk` quirk:** rows still reference **`c11` 40, `rewalk` 20, `part4` 24** — all three torn down 2026-08-06, whose ledger recorded *"0 occurrences"*. **That prior claim was measured with a narrower query and does not survive a full census; the correction is recorded rather than the measurement quietly redone.** **Why it was probably never written, established rather than assumed:** the table is a **fleet-wide aggregate** keyed on `app_name`+`fingerprint` with an `affected_customers` JSON list — of the 29 `finalwalk` rows, **12 reference only `finalwalk`** (orphans, safely deletable) and **17 are shared with LIVE customers** (`demo-felhom`, `peti-felhom`, …) and **must not be deleted, only de-referenced.** A naive `DELETE … WHERE customer LIKE` would destroy a live customer's issue history — which is very likely why the leg does not exist, and is the reason this is not a one-line fix. **Severity is LOW and stated plainly: no secret material is involved** — app name, fingerprint, message text, counts, timestamps. What survives is a deleted customer's *identifier* inside an aggregate row. **Proposed shape:** a residue leg that (a) removes the customer id from `affected_customers`/`context_customer`, and (b) deletes rows whose `affected_customers` becomes empty; plus a one-off sweep for the four already-torn-down venues. **The general lesson is the reusable part:** *a per-table absence query is not a census.* The teardown verification is now a full-schema sweep, and that is what found this. **Not fixed** — a cascade change needs its own red-proof and this session was scoped as a spike plus two operations. Evidence: `tests/teardown-finalwalk-2026-08-07.md`. **⚠ STILL OWED, AND NOW MEASURED RATHER THAN ESTIMATED (2026-08-08 census, read-only, no truncation).** `app_log_issues` holds **1309 rows**; **71 reference a torn-down venue** (`finalwalk`, `c11`, `rewalk`, `part4`); of those **44 are ORPHANS** — they name only torn-down customers and are safely deletable — and **27 are SHARED with a live customer** (`demo-felhom`, `peti-felhom`, …) and **must be de-referenced, never deleted**. 1238 rows are untouched. **The 27 are exactly why the leg was never written**, and why a `DELETE … WHERE customer LIKE` would destroy a live customer's issue history. **What it needs, precisely:** a cascade leg that (a) removes the customer id from `affected_customers` / `context_customer`, and (b) deletes only rows whose `affected_customers` becomes empty; plus a one-off sweep for the four venues already gone. **Why it was NOT done on 2026-08-08:** the fix is hub code, and that session's scope forbade a hub version bump; a hand-run SQL mutation over 71 rows — 27 of them needing surgical de-referencing — with no tested code path and no red-proof is precisely the shape that goes wrong on a live database. **It accumulates one venue at a time, so the next walk adds to it**; the numbers above mean the next session starts from data rather than a guess. **⚠ IT GREW AGAIN, AS PREDICTED — walk5 teardown, 2026-08-08.** The fifth walk's venue was torn down with a full-schema census taken **before and after**: **168 rows → 67**. Of the 67, **37 are by design** (`events` 21, `notification_log` 14, `host_deletions` 1, `customer_resets` 1) and **30 are `app_log_issues`** — this row's gap, and the count was **predicted in the pre-run enumeration rather than discovered afterwards**, which is the difference from the ledger that once recorded *"0 occurrences"* from a narrower query. **The running total across torn-down venues therefore rises from 71 to ~101 rows** (`finalwalk`, `c11`, `rewalk`, `part4`, now `walk5`) — the shared-with-a-live-customer subset must still be **de-referenced, never deleted**. **It accumulates one venue at a time and it did so again.** Evidence: `tests/walk5-r201-2026-08-07/teardown-walk5-2026-08-08.md`. **2026-09-25:** `peti-felhom` is no longer a live customer (deleted through the cascade, journal #20); 8 `app_log_issues` rows still name it — the same gap. | **READY** — owner Viktor | — | — | operator | | **R-882** | Hub & operator | P3 | **Longhorn on DooPlex could not grow a volume online: its `instance-manager` (116 days up) called a host process that no longer existed** — `nsenter: cannot open /host/proc/196610/ns/mnt` on every expansion retry, and an offline growth was blocked by the expansion's own attachment ticket (found 2026-10-05 growing `hub-data` to 2 Gi). A restart of the instance-manager (operator-approved) fixed it: 77/77 volumes back `attached/healthy` in 110 s. **Why the cached PID went stale was not established** (likely a containerd/k3s or iscsid restart after the instance-manager started), so it will recur after the next such restart and stay invisible until a volume needs to grow. `audits/hub-db-offsite-2026-10-05/partA/step1-*.txt` | **OPEN** | — | Find which host process the PID was and whether Longhorn 1.10.x re-resolves it; until then, before growing any volume, check the instance-manager's age against the last k3s/containerd restart | operator | diff --git a/skills/felhom-build-deploy/SKILL.md b/skills/felhom-build-deploy/SKILL.md index aae89c77..1195a91e 100644 --- a/skills/felhom-build-deploy/SKILL.md +++ b/skills/felhom-build-deploy/SKILL.md @@ -161,7 +161,12 @@ is never placed on the system by an interactive install, so day-0 rides a `.deb` # 1. commit+push code 2. build+push image (LOCAL) cd $FELHOM_ROOT/build/felhom-hub && ./build.sh --push # 3. bump manifests/hub.yaml image tag → , commit, push -# 4. hard-refresh + sync (argocd CLI is not logged in — drive the Application CR) +# 4. hard-refresh, then READ what is OutOfSync before syncing. If anything besides Deployment/hub is OutOfSync +# (since 2026-10-09: three Secrets de-gitted by R-925 + Deployment/umami), sync ONLY the hub — a whole-app sync +# pushes git's view over rotated live Secrets: +# sudo kubectl -n argocd patch application felhom --type merge -p '{"operation":{"initiatedBy":{"username":"cc"},"sync":{"revision":"","resources":[{"group":"apps","kind":"Deployment","name":"hub","namespace":"felhom-system"}]}}}' +# Image push 401 "unauthorized" from DooPlex = its saved Docker login predates the R-925 rotation (see R-925). +# (argocd CLI is not logged in — drive the Application CR) sudo kubectl -n argocd annotate application felhom argocd.argoproj.io/refresh=hard --overwrite; sleep 8; sudo kubectl -n argocd get application felhom -o jsonpath='{.status.sync.status} {.status.sync.revision}{"\n"}' sudo kubectl -n argocd patch application felhom --type merge -p '{"operation":{"initiatedBy":{"username":"cc"},"sync":{"syncStrategy":{"apply":{}}}}}' # 5. verify: Synced/Healthy + rollout + image tag + startup log