From e467785fd3efe717451db037b3eb4f82cd9a5dac Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Fri, 9 Oct 2026 13:21:10 +0200 Subject: [PATCH] security: rotate the published secrets and de-git felhom.secret.yaml (R-925, P1) WHY .gitignore "was not working": it was working. git never consults .gitignore for a file it ALREADY TRACKS. The rule `*secret*` matched fine -- proved by dropping an untracked copy in and watching check-ignore name `.gitignore:3:*secret*`. The file had been tracked since feea0606, which is ironically the commit that de-gitted the Resend key. WHAT THE EXPOSED VALUE ACTUALLY WAS. Not an analytics password: the GITEA ADMIN ACCOUNT PASSWORD (is_admin true; /api/v1/admin/users answered 200), in a repo gitea.dooplex.hu serves anonymously to the internet. That is push access to every repo -- including the one whose website/ is git-synced live and whose scripts/ is published by tag to every new box installer (R-110). Re-ranked P2 -> P1 on that measurement; my first ranking had only measured the analytics blast radius. Every committed value was still live. Nothing had ever been rotated. ROTATED (values never echoed; written to a 0600 file on DooPlex): umami-config APP_SECRET + POSTGRES_PASSWORD. The password was changed INSIDE postgres (ALTER USER) as well as in the Secret -- the env var is only read at first init, so patching the Secret alone would have changed nothing. healthchecks-config SECRET_KEY + SUPERUSER_PASSWORD (nothing consumes them, there is no healthchecks Deployment). gitea-creds no longer holds the admin password at all: a SCOPED token (read:package + read:repository). gitea admin new random password; gitea-system/gitea-admin updated. VERIFIED, not assumed: - new admin password -> 200, OLD PUBLISHED PASSWORD -> 401 (the leak is dead) - umami: a real beacon returns 200 (so the app authenticates to postgres and writes) while a bogus site id still returns 400 (so the 200 means something) - hub: "Registry version check: latest = 0.304.0" AND "Template fetched (5881 bytes)", no auth failures - BOTH token scopes are load-bearing, and the second was found by breaking it: a package-only token made the hub log "Template fetch: unexpected status 403", because the template fetcher reads a raw file out of the felhom-controller repo, not the registry. AN INCIDENT CAUSED BY THE FIX, recorded because it is the useful part: the rollout restart needed to pick up the new umami secret put umami into CrashLoopBackOff and took stats.felhom.eu down (503) for ~4 minutes. Not the rotation -- at memory 512Mi that pod runs for months but CANNOT RESTART: startup (Prisma + Next.js) peaks over the limit and is OOMKilled (exit 137). Raised to 1Gi IN THE MANIFEST, not just live, per .claude/rules/manifests.md ("never bare kubectl set -- the next sync reverts it and the fix silently disappears"). THE GATE: KNOWN_BACKLOG is removed from manifest_bearer_gate.py, as its own comment instructed. Red-proofed with a decoy: exit 1 with it, exit 0 without. An exemption kept this visible for three months and changed nothing. WHAT REMAINS (operator, and it is bigger than what was fixed): the same password is still the admin password in ~12 other namespaces -- nextcloud, paperless, bookstack (a DATABASE ROOT password), tandoor, calibre, adventurelog, gokapi, qbittorrent, servarr, homepage. Rotating Gitea does not touch them. Also owed: a kisfenyo Gitea token sits in plaintext in the local homelab-manifests remote URL and was printed to a session transcript during this investigation, so it should be replaced regardless (R-580's shape). NOT a finding: homelab-manifests is private (404 anonymously) and does not contain the password; ArgoCD's repo credential is a separate token and was untouched by the rotation. --- REUSE.md | 4 +- STATUS.md | 23 +++++++++- documentation/backlog/OPEN-ITEMS.md | 4 +- documentation/runbooks/secrets.md | 70 ++++++++++++++++++++++++++--- manifests/felhom.secret.yaml | 51 --------------------- manifests/umami.yaml | 9 +++- scripts/manifest_bearer_gate.py | 12 ++--- 7 files changed, 105 insertions(+), 68 deletions(-) delete mode 100644 manifests/felhom.secret.yaml diff --git a/REUSE.md b/REUSE.md index f4b62247..1b9e0dff 100644 --- a/REUSE.md +++ b/REUSE.md @@ -176,7 +176,7 @@ | Install-profile gate (shell) | scripts/felhom-host-install.sh `--mode appliance\|byo` (GL-2, v1.10.0) | Mandatory-flag profile (no default), refusals at argv time BEFORE any prompt/step, risky step gated at its CALL SITE (one auditable place — never a branch inside the step), mode persisted to state.json + resume-mismatch refusal, `FELHOM_INSTALL_STATE_DIR` override for harness isolation. Harness: scripts/hostinstall-mode-harness.sh (static refusal matrix + grep-invariants + PVE dry-transcript tier; red-proofs run against a mutated scratch copy). | | Disclosure↔uninstall parity (shell) | scripts/felhom-host-install.sh `_uninstall_statement` + harness GL4-D (v1.11.0) | Every host artifact the byo disclosure names must be removed OR explicitly listed KEPT by `run_uninstall`; the harness greps the parity (token list). New install-time artifact ⇒ add its removal + disclosure line + parity token in the SAME commit. Drive data rule: plain `umount` only, never `-l`/`-f`, never any format op under /mnt/felhom-drives. | | Website deploy (manifest) | manifests/webpage.yaml | git-sync sidecar (sparse-checkout `/website/` + `/scripts/`, `--link=current`) + init container waits for first sync; nginx serves `current/website`; push to main = deployed, no image build. | -| Secret handling (manifest) | manifests/hub.yaml (env, ~L142) | Secrets via `secretKeyRef` to OUT-OF-BAND secrets created per documentation/runbooks/secrets.md — never inline stringData (see §3). `report-api` (the operator bearer, v0.53.0) is deliberately NOT `optional:` — a missing Secret fails Ready instead of booting an unauthenticatable hub. `scripts/manifest_bearer_gate.py` (run after ANY manifests/ change) blocks bearer-shaped (64-hex) literals. ERRATA (2026-07-03): `gitea-creds` is COMMITTED in manifests/felhom.secret.yaml AND live-consumed by hub.yaml — rotation + de-git is a pending operator task (spike SPIKE-a1 appendix). | +| Secret handling (manifest) | manifests/hub.yaml (env, ~L142) | Secrets via `secretKeyRef` to OUT-OF-BAND secrets created per documentation/runbooks/secrets.md — never inline stringData (see §3). `report-api` (the operator bearer, v0.53.0) is deliberately NOT `optional:` — a missing Secret fails Ready instead of booting an unauthenticatable hub. `scripts/manifest_bearer_gate.py` (run after ANY manifests/ change) blocks bearer-shaped (64-hex) literals. RESOLVED 2026-10-09 (R-925), was ERRATA 2026-07-03: `gitea-creds` was COMMITTED in the since-deleted `felhom.secret.yaml` manifest while being live-consumed by hub.yaml — and that value was the **Gitea `admin` account password**, in a repo gitea.dooplex.hu serves anonymously to the internet. It is rotated (old password now 401) and the file is removed from git; `gitea-creds` now holds a **scoped token** (`read:package` + `read:repository`), never the admin password. BOTH scopes are load-bearing: the registry version checker needs the first and the template fetcher needs the second, and a package-only token makes the hub log `Template fetch: unexpected status 403`. `KNOWN_BACKLOG` is gone from `manifest_bearer_gate.py`, so a reintroduction now FAILS instead of printing a visible-but-harmless note. Recreate procedure: documentation/runbooks/secrets.md. | | Hub deploy (GitOps) | manifests/hub.yaml `image:` (~L129) | Pinned explicit tag, bumped in git, deliberate ArgoCD sync (auto-sync OFF). Code push alone deploys nothing. | ## 3. Dangerous lookalikes — do NOT reuse @@ -187,7 +187,7 @@ | `(*Handler).handleNotify` + `formatNotificationEmail` + `sendResendEmail` (hub/internal/api/handler.go ~L1289/1624/1589) | Legacy pre-dispatcher notification trio: no cooldowns, no operator channel, no allowedEventTypes gate, duplicate Hungarian formatter. Controller path is FROZEN until slice-10 cutover. | `POST /api/v1/event` → `Dispatcher.ProcessEvent` + `notify.Format*Email` | | Severity `"critical"` POSTed to a PRE-v0.31.0 hub | Fixed in hub v0.31.0 (`handleEvent` now accepts critical). Older hubs coerce `critical` → `"info"`, which never notifies — silent alert loss. Case-variants (`"Critical"`) still coerce on every version. | Against an old hub send `warning`/`error`; otherwise lowercase `critical` is safe | | `compareVersions` for anything security-ish (hub/internal/web/server.go ~L571) | Returns 0 (equal) on unparseable input — a garbage version passes a floor check. `gitea.compareSemver` behaves differently (lexical fallback). | Validate input with `normalizeFloorInput` first; then compareVersions is safe | -| Inline `stringData` secrets à la manifests/felhom.secret.yaml | Commits real credentials to git (healthchecks superuser pw, umami APP_SECRET/POSTGRES_PASSWORD, gitea-creds admin password still live there). | Out-of-band `kubectl create secret` + `secretKeyRef` (hub.yaml resend-api pattern; runbook documentation/runbooks/secrets.md) | +| Inline `stringData` secrets, as the deleted `felhom.secret.yaml` manifest did | Commits real credentials to git. This was not hypothetical: that file held the healthchecks superuser password, the umami `APP_SECRET`/`POSTGRES_PASSWORD` and — worst — the **Gitea `admin` account password**, in a repo served anonymously to the internet, and every value was still live three months after being flagged. All rotated and the file de-gitted 2026-10-09 (R-925). A `KNOWN_BACKLOG` exemption kept it *visible* in the gate and changed nothing for three months, which is the lesson: an exemption is not a plan. | Out-of-band `kubectl create secret` + `secretKeyRef` (hub.yaml resend-api pattern; runbook documentation/runbooks/secrets.md) | | `kubectl apply` / `kubectl set image` on manifests/ | ArgoCD app `felhom` reverts drift on next sync; live state lies about git. | Edit manifest in git → push → ArgoCD sync (CLAUDE.md steps 3–5) | | `:latest` image tag in manifests | Re-push doesn't change the manifest → no redeploy; Synced/Rollback misreport. | Pinned version tag, bumped per deploy | | The report's `backup` object for snapshots / repo size / integrity (`snapshot_count`, `repo_size_mb`, `integrity_ok`) | **No producer since slice 8C** — the controller's `buildBackupReport` leaves all four zero deliberately and says so. Rendering them gave every customer `Snapshots 0 · Repo Size 0 MB · Integrity Unknown` indefinitely (R-331, measured on demo-hp over a repository holding 67 snapshots). `integrity_ok` is worse than stale: the controller runs no integrity check at all, so it can only ever be "Unknown" or a lie. | The report's `offsite` object (`backup.OffboxReportStatus`) via `reportBackupCard` — and check `stats_known` before trusting a zero | diff --git a/STATUS.md b/STATUS.md index eba88a08..82c756b8 100644 --- a/STATUS.md +++ b/STATUS.md @@ -47,7 +47,28 @@ - **Still to prove:** the lost off-site copy alarm (Tester 1's next clean-up window, about 12 October), and the failed-restore hold (needs a scratch off-site store). -## Needs you soon (2026-10-09): passwords are sitting in a repository anyone on the internet can read +## Done and still owed (2026-10-09): the published password is dead for Gitea — but it still opens a dozen other things + +Following up the finding below, I measured what that password actually was. It was **your Gitea admin login**, +published on the internet. That is push access to every repository — including the one the website is served from +and the one new boxes download their installer from. I rotated it, on your instruction. + +- **Dead now:** the old password returns "unauthorized" at Gitea. The visitor counter and the health-check tool + got fresh random passwords too. The hub no longer holds your admin password at all — it holds a **limited token** + that can only read the things it needs. +- **I broke something briefly and fixed it, and you should know:** restarting the visitor counter to pick up its + new password took **stats.felhom.eu down for about four minutes**. Not the password — that service had a memory + limit it could run under for months but could never *restart* under. It is raised now, in the file as well as on + the server, so the next restart works. +- **Still open, and this is the bigger half:** the same password is the admin password for about **a dozen other + services** — Nextcloud, Paperless, Bookstack (that one is a *database* password), Tandoor, Calibre, qBittorrent + and others. Rotating Gitea did nothing for those. Each needs its own new password. +- **Also:** a Gitea token of yours sits in plain text inside one local repository's settings, and I printed it to + my own session while investigating. Worth replacing. +- **Your passwords are in a protected file on DooPlex** (`rotated-secrets-2026-10-09.txt`). Move them into your + password manager and delete it. + +## The finding itself (2026-10-09): passwords were sitting in a repository anyone on the internet can read I found this while publishing the legal pages, not because I went looking — a security check flagged something small in the new page, and following it led here. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 6892e730..d9db7cc5 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -192,7 +192,7 @@ stopping line that lies. | **R-331** | Storage & devices | P4 | **Disk health Phase 3 — growth-rate detection, and retiring the static 64.** The v0.215.0 count backstop (64 unreadable sectors → Hiba) is **a judgement from ONE drive**: the observed benign excursion peaked at 16 and cleared inside an hour, and the terminal run passed 64 at 13 Aug 11:28 and never came back. It is deliberately a backstop BEHIND the sustain rule, not the primary signal, but it is still a magic number tuned on a single sample and it will be wrong for some drive. With Phase 2's history the box can ask the question that actually matters — *is this count climbing, and how fast* — which distinguishes a drive with eight stable aging sectors from one adding forty a day, something no static threshold can do. Revisit 64 when that exists | **READY (M) — NEW 2026-08-14** | R-330 | Growth-rate rule over persisted samples; re-derive or delete the static 64 | CC | | **R-352** | Storage & devices | P4 | **Four screens state something untrue about where an app's data goes, and the configured default is consulted by nothing that places data.** Measured on `demo-hp` 2026-08-21. **(1)** **40 of 53** catalogue templates declare no data path (`grep -rl 'env_var: HDD_PATH' --include='.felhom.yml'` → 13; total 53); those apps get **no storage field and no default** — their data lands in a named Docker volume on the system drive. **(2)** `GetDefaultStoragePath()` has exactly **three** non-test callers — the metrics collector (`cmd/controller/main.go:410`), the dashboard SystemInfo panel (`web/server.go:733`) and `.fab` import landing (`handler_export_upload.go:154`). **The deploy route never reads it.** Its field comment `// new apps use this by default` (`internal/settings/settings.go:453`) has never been true — an invariant with no test pinning it. **(3)** The first-tier backup follows the data onto the same disk (`backup/backup.go:324-334` → `systemDataPath`), so data and nearest copy share one device for a customer doing nothing wrong — the posture Tier 2 refuses outright at `tier2.go:329`. **(4)** „1 alkalmazás használja" on the Drives page counts only `Env["HDD_PATH"] == path` (`web/handlers.go:2118`), so it can never include the 40-class; it truthfully means *„1 of the apps that CAN use a drive does"*. | **NARROWED** — **PARTLY CLOSED 2026-08-21** — visibility shipped; **placement OPEN** | — | **Shipped tonight (visibility only, no placement change, nothing migrated):** the deploy page now states where the app's data will live before the button is pressed, naming the system drive for the 40-class and the selected drive for the 13. **⚠ RE-FRAMED 2026-08-22 — THE FOUR MEASUREMENTS STAND; TWO OF THE CONCLUSIONS DRAWN FROM THEM DO NOT.** (1) is a measurement and is correct, but "those apps get no storage field and no default" is not a deprivation: the architecture places **hot** data (DB/config/cache) on fast storage inside the guest and states that placement is **ENFORCED** (`documentation/architecture/01-topology-and-trust.md:150-152`). The 40 are all-hot apps; the 13 are the ones with **bulk** content, which belongs on an attached drive. There is no choice being denied. (3) **overstated one risk and understated a distinction.** Since R-165 the guest carries a small OS rootfs plus **ONE** data volume at `/var/lib/felhom`; `/var/lib/docker` and `/mnt/sys_drive` are two **binds of that same volume** (`felhom-agent/configs/build-golden.sh:29-40, 99`) — the `mp0`/`mp1` split assumed here was retired 2026-08-03. **Real risk:** a physical-disk failure loses the data and its first-tier copy together — which is what the off-site and whole-machine tiers exist for, and which is equally true of a drive-resident app whose unit sits beside its data by design. **Overstated risk:** a full data volume stopping the operating system — the OS rootfs is a separate volume and the capture floor refuses per app before exhaustion (`00-capability-map.md:94`), watched working 2026-08-21 with the volume at 99% and all 15 containers healthy. **The comparison to Tier 2's same-disk refusal (`tier2.go:329`) is withdrawn:** Tier 2 refuses a SECOND copy on the same disk; Tier 1's unit is meant to sit beside the data. **(2) and (4) are untouched and remain correct** — (2) is now filed on its own as **R-368** with its scope measured, and (4) needs no ruling: the count is honest and only easy to misread. **The specification for the rest is filed at `documentation/backlog/SPEC-app-data-placement-2026-08-21.md`** (corrected 2026-08-22, framing marked inline, measurements kept) and lists the five points a ruling must settle (compose-template vs controller, existing deployments, when the SSD is legitimately right, `IsDefault` must become true or go away *with a test*, and the Drives-page count). **An earlier recommendation to refuse deployment until a drive is registered was WITHDRAWN** — it assumed the customer had failed to choose; they had no choice to make. | **Viktor rules**, CC executes | -## Security & access — 16 rows (P2 2, P3 10, P4 4) +## Security & access — 16 rows (P1 1, P2 1, P3 10, P4 4) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -208,7 +208,7 @@ stopping line that lies. | **R-870** | Security & access | P3 | **Tester 1's two Cloudflare credentials — the zone API token (`infrastructure.cf_api_token`) and the tunnel token (`infrastructure.cf_tunnel_token`) of the hub's `customer_configs` row `tester-1` — were printed into the 2026-10-04 night session's transcript** (not into any file): a read-only query selected `substr(config_json,1,400)`, and both values sit in the first 400 characters. Tester 1 is CC's disposable test customer (`enkicsifelhom.hu`). **Not rotated, by the operator's ruling of 2026-10-05 06:49 (option B).** **Rotation, whenever chosen (3 steps):** in the Cloudflare dashboard create a new API token for the `enkicsifelhom.hu` zone with the same permissions and refresh the Tester 1 tunnel's token (Zero Trust → Networks → Tunnels → the tunnel → refresh token) → hub → Configs → `tester-1` → Edit → the two Cloudflare fields → Save, then confirm on the box that cloudflared reconnected (`docker ps` health `healthy`) → delete the old API token. Rule for sessions (as R-831): never select a whole config row — name the fields, and never `config_json` without `json_extract` of a non-secret field. | **WAITING-ON-OPERATOR — rotation is his call (ruled: not now)** **Not rotated by the operator's rulings (2026-10-04 „keep using the current one"; 2026-10-05 option B) — restated 2026-10-05 18:23; the steps stay here.** | — | rotate when chosen | operator | | **R-908** | Security & access | P3 | **An old Resend API key is still in the history of the `homelab-manifests` repository; it was replaced on 2026-06-29, but whether it is also revoked at Resend is unknown.** FOUND 2026-10-08 by the contact-mailer source search (`audits/mailer-source-2026-10-08/SEARCH.md`, „A side finding"): commit c9648cd (2026-02-05, „added mailer pod") put a key literal in the old `felhom-system/contact-mailer.yaml` comment; the file was deleted in ee93b50 but the history keeps it. Compared by hash only, never printed: it is NOT today's key (`Secret/resend-api`, rotated 2026-06-29, felhom.eu feea0606). If the old key still works at Resend, anyone with read access to that repository can send mail as felhom.eu. | **OPEN — owner: operator** | — | In the Resend console, check that the old key is revoked (revoke it if not); rewriting the repository's history is NOT needed once it is revoked | operator | | **R-525** | Security & access | P4 | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** **Re-ranked 2026-10-03: P3->P4: comfort feature needing a new unmeasured mechanism; the default password hole is closed.** | — | — | CC | -| **R-925** | Security & access | P2 | **`manifests/felhom.secret.yaml` carries REAL secret values, and the Gitea repo that holds it is readable by anyone on the internet with no login.** FOUND 2026-10-09 by a background security review of the legal-pages commit, which flagged the served `` comments; chasing what those comments pointed at found this instead. **MEASURED, in this order:** (1) `scripts/manifest_bearer_gate.py` itself prints `manifests/felhom.secret.yaml:39 KNOWN-BACKLOG committed secret 65cee3c4…7a86`; (2) the file holds **live-shaped values, not placeholders** — `SECRET_KEY` (69 chars), `SUPERUSER_EMAIL`/`SUPERUSER_PASSWORD` (18), Umami `APP_SECRET` (66) and `POSTGRES_PASSWORD` (34), and a `username`/`password` pair (18); values were never printed, only measured by length; (3) `gitea.dooplex.hu` resolves to **37.191.56.193**, the same public address as the website; (4) an **anonymous** `curl` returns HTTP 200 and 1686 bytes for that exact file, and 200 for `hub/internal/store/store.go` and `manifests/hub.yaml`; (5) **the off-network control**, which is what makes (4) mean anything — fetched from outside the operator's network entirely: a real path returns the file's first line, a nonsense path returns 404, so the 200 is genuine anonymous read from the internet and not a LAN-only ACL. **This contradicts the project's own written model**: `documentation/runbooks/secrets.md:3-4` says *“Secret values are never committed to git. The manifests in `manifests/` carry only placeholders + comments that point here.”* That promise is false today, which is the P1 wording in this register's own scale — see the ranking note below. The same runbook already states the rule that governs the fix: *“The git-history copy stays alive until the value is ROTATED — de-git alone kills nothing.”* **Known neighbours, none of which cover this:** R-887 records that an outside crawler walks the public Gitea pages and leaves *“Gitea's exposure to the crawler”* as the operator's undone item; R-580 still owes a Gitea admin token rotation. Neither says a secret file is world-readable. **Why P2 and not P1, stated so the operator can overrule it:** the exposed credentials guard **analytics and an undeployed healthchecks instance, not household data** — measured: `umami-db` is a **ClusterIP** service on 5432 with no external IP, so the Postgres password is not reachable from the internet; the realistic harm is session forgery against the public `stats.felhom.eu` via `APP_SECRET`, and reuse of those passwords anywhere else. No customer box, hub token or escrow key is in this file. **CC did NOT change anything**: making the repo private could break the public day-0 path (the installer is fetched from a public tag, R-110), and rotation plus repository visibility are operator decisions on production infrastructure. | **OPEN — owner: operator (rotate + decide repo visibility); CC can do the de-git and the gate once told** | — | Operator, in this order: **(1) ROTATE** the Umami `APP_SECRET` and `POSTGRES_PASSWORD` and the healthchecks `SECRET_KEY`/`SUPERUSER_PASSWORD` — de-gitting first changes nothing, the history keeps the old values; **(2) decide** whether `gitea.dooplex.hu` should be readable anonymously at all, remembering the public installer tag depends on some public path (R-110); **COUPLED, and easy to miss:** `website/adatkezeles.html` serves 32 `` comments naming `hub/internal/...`, `manifests/...` and architecture paths — harmless while the repo is public (they point at files anyone can already read), but **the moment the repo is made private those comments become a map of it**, so step (2) and stripping them are ONE change, not two. The traceability they carry is already kept in `documentation/legal/*-1.0.md`, so stripping costs nothing. **(3) then** tell CC to move the values out of `manifests/` into out-of-band Secrets per `secrets.md` and extend `manifest_bearer_gate.py` so the KNOWN-BACKLOG exemption becomes a failure. If nothing is done: live credentials stay downloadable by anyone who finds the repo, and the secrets runbook keeps stating a rule the repo breaks | operator | +| **R-925** | Security & access | P1 | **`manifests/felhom.secret.yaml` carries REAL secret values, and the Gitea repo that holds it is readable by anyone on the internet with no login.** FOUND 2026-10-09 by a background security review of the legal-pages commit, which flagged the served `` comments; chasing what those comments pointed at found this instead. **MEASURED, in this order:** (1) `scripts/manifest_bearer_gate.py` itself prints `manifests/felhom.secret.yaml:39 KNOWN-BACKLOG committed secret 65cee3c4…7a86`; (2) the file holds **live-shaped values, not placeholders** — `SECRET_KEY` (69 chars), `SUPERUSER_EMAIL`/`SUPERUSER_PASSWORD` (18), Umami `APP_SECRET` (66) and `POSTGRES_PASSWORD` (34), and a `username`/`password` pair (18); values were never printed, only measured by length; (3) `gitea.dooplex.hu` resolves to **37.191.56.193**, the same public address as the website; (4) an **anonymous** `curl` returns HTTP 200 and 1686 bytes for that exact file, and 200 for `hub/internal/store/store.go` and `manifests/hub.yaml`; (5) **the off-network control**, which is what makes (4) mean anything — fetched from outside the operator's network entirely: a real path returns the file's first line, a nonsense path returns 404, so the 200 is genuine anonymous read from the internet and not a LAN-only ACL. **This contradicts the project's own written model**: `documentation/runbooks/secrets.md:3-4` says *“Secret values are never committed to git. The manifests in `manifests/` carry only placeholders + comments that point here.”* That promise is false today, which is the P1 wording in this register's own scale — see the ranking note below. The same runbook already states the rule that governs the fix: *“The git-history copy stays alive until the value is ROTATED — de-git alone kills nothing.”* **Known neighbours, none of which cover this:** R-887 records that an outside crawler walks the public Gitea pages and leaves *“Gitea's exposure to the crawler”* as the operator's undone item; R-580 still owes a Gitea admin token rotation. Neither says a secret file is world-readable. **Why P2 and not P1, stated so the operator can overrule it:** the exposed credentials guard **analytics and an undeployed healthchecks instance, not household data** — measured: `umami-db` is a **ClusterIP** service on 5432 with no external IP, so the Postgres password is not reachable from the internet; the realistic harm is session forgery against the public `stats.felhom.eu` via `APP_SECRET`, and reuse of those passwords anywhere else. No customer box, hub token or escrow key is in this file. **CC did NOT change anything**: making the repo private could break the public day-0 path (the installer is fetched from a public tag, R-110), and rotation plus repository visibility are operator decisions on production infrastructure. **PARTLY FIXED 2026-10-09, and RE-RANKED P2 -> P1 on what it turned out to be.** The exposed value was not an analytics password: it was the **Gitea `admin` account password** (`is_admin: true`; `/api/v1/admin/users` answered 200), published in a repo `gitea.dooplex.hu` serves anonymously to the internet. That is push access to every repo — including the one whose `website/` is git-synced live and whose `scripts/` is published by tag to **every new box installer** (R-110), i.e. a supply-chain path onto customer hardware. My first ranking said P2 because I had only measured the analytics blast radius; the credential test is what corrected it. **DONE (CC, operator-instructed):** (1) `umami-config` `APP_SECRET` + `POSTGRES_PASSWORD` rotated — the password changed **inside Postgres** too (`ALTER USER`), because the env var is only read at first init; verified by a real beacon returning 200 with a bogus-site-id control still returning 400. (2) `healthchecks-config` `SECRET_KEY` + `SUPERUSER_PASSWORD` rotated (nothing consumes them — there is no healthchecks Deployment). (3) `gitea-creds` no longer holds the admin password at all: it holds a **scoped token** (`read:package` + `read:repository`). **Both scopes are load-bearing and the second was found by breaking it** — a package-only token made the hub log `Template fetch: unexpected status 403`, because the template fetcher reads a raw file out of the `felhom-controller` repo. (4) The Gitea **admin password** was changed and `gitea-system/gitea-admin` updated; **verified: new password 200, old published password 401 — the leak is dead.** (5) The file is removed from git and `.gitignore`'s `*secret*` now applies to it; `manifest_bearer_gate.py`'s `KNOWN_BACKLOG` carve-out is gone, red-proofed with a decoy (exit 1 with, 0 without). **AN INCIDENT CAUSED BY THE FIX, recorded because it is the useful part:** the `rollout restart` needed to pick up the new umami secret put umami into **CrashLoopBackOff and took `stats.felhom.eu` down (503)**. Cause was not the rotation — at `memory: 512Mi` the pod runs for months but **cannot restart**, because startup (Prisma + Next.js) peaks over the limit and is OOMKilled (exit 137). Raised to 1Gi, **in the manifest, not just live** (the `.claude/rules/manifests.md` rule: a bare `kubectl set` is reverted by the next sync and the fix silently disappears). Service restored and verified. New values are in a 0600 file on DooPlex (`~/rotated-secrets-2026-10-09.txt`), never echoed; the operator moves them to the password manager and deletes it. **WHAT REMAINS, and it is bigger than what was fixed:** that one password is still the admin password in about **12 other namespaces** — `nextcloud`, `paperless`, `bookstack` (a **database root** password), `tandoor`, `calibre`, `adventurelog`, `gokapi`, `qbittorrent` (x2), `servarr`, `homepage` (x2). Rotating Gitea does not touch them: each is its own login and each is still the string that was published. **Also owed:** a Gitea token for user `kisfenyo` sits in plaintext in the `origin` URL of the local `homelab-manifests` clone (`.git/config`); CC printed it to a session transcript while investigating, so it should be rotated regardless — this is R-580's shape (store the remote without credentials). **Not a finding:** `homelab-manifests` itself is private (404 anonymously) and does not contain the password; ArgoCD's repo credential is a separate token, so the rotation did not touch it. | **NARROWED 2026-10-09 — the felhom.eu half is DONE and verified (Gitea admin password rotated, old one now 401; umami + healthchecks rotated; file de-gitted; gate carve-out removed). What remains is the ~12 reused logins elsewhere in the homelab and the repo-visibility decision; owner: operator** | — | Operator: **(1)** change the admin password of the ~12 other services that still use the published string (`bookstack` first — it is a database root password); **(2)** rotate the `kisfenyo` Gitea token embedded in the local `homelab-manifests` remote URL, and store the remote without credentials; **(3)** decide whether `gitea.dooplex.hu` should answer anonymously at all — remembering the website git-sync clones it with no credentials and the installer is fetched from a public tag, so making it private breaks both unless they get credentials first, and the 32 `` comments on /adatkezeles become a map of it; **(4)** move `~/rotated-secrets-2026-10-09.txt` into the password manager and delete it. If nothing is done: the published password keeps opening a dozen services, even though Gitea itself is now safe | operator | | **R-779** | Security & access | P4 | **[P3-LOW] Part A's "two outside addresses seen as two" is proven through the simulated tunnel only; on the REAL tunnel the second outside address (ep0, one request allowed) was refused by Cloudflare's edge with 403 and never reached the box.** Measured 2026-10-01 19:51 UTC (`audits/visitors-2026-10-01/A/L2-demo-hp-real-tunnel.txt`): no log line on demo-hp; demo-hp's box has no geo restriction in its settings, so a Cloudflare ZONE rule (country or bot, not read) refused a German datacenter address. DooPlex's own address on the real tunnel was seen as itself. **Needs:** one sign-in from a second Hungarian address (the operator's phone off wifi) while DooPlex is locked out — 2 minutes; and say which Cloudflare rule refused ep0. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (a phone), CC reads the logs** **Re-ranked 2026-10-03: P3→P4: a proof gap on the real tunnel; operator-only follow-up.** | — | — | CC + operator | | **R-904** | Security & access | P4 | **Cloudflare can read every household's app traffic; replacing it with our own relay is a later item.** Facts (reviewer discussion 2026-10-08; `01`): app traffic and the dashboard reach the box through the Cloudflare Tunnel (`01` §5 trust table, rows end-user ↔ apps and customer ↔ controller UI; §7), and the tunnel's public end is Cloudflare's edge, where TLS ends — so Cloudflare can technically read that traffic (the FAQ says so since 2026-10-08, R-900; „TLS ends at the edge" is not written in `01` — add it there). What Cloudflare gives today, free: inbound reach with no router setup, the CGNAT answer (`01` §4, §7); certificates (`01` §7, the free tier covers one level below a zone); the geo-WAF the hub enforces (`01` §5 last row, §7); flood protection (not in `01`). The alternative named: our own EU relay over WireGuard with TLS passthrough by SNI, certificates on the box, the geo-block on the relay. Its costs: one more machine the operator keeps up, and a single point of reach for every box; weaker flood protection; about a week of work after a spike. Operator ruling 2026-10-08 09:07 (`09` §3 decision 184): a later item. | **DEFERRED — after the first customers (operator ruling 2026-10-08 09:07)** | — | A spike after the first customers (the relay's reach, cost and flood behaviour, measured) | operator | | **R-913** | Security & access | P4 | **The Cloudflare token check (R-138, decision 190) reads what a token can SEE, not what it can WRITE.** FOUND 2026-10-08 by the security review of the R-138 build: `GET /zones` lists zones the token can read; a hand-built token with Zone:Read on the customer's zone and DNS:Edit on ALL zones would pass the check and could still change every household's DNS. The token wizard's single „Specific zone" scope does not build such a token, and a customer token cannot read its own policies (that needs „API Tokens Read"). Written as a limit in `01` §7. **Options:** (a) the operator mints every customer token from one recipe and the hub checks nothing more; (b) the hub mints the token itself with the operator's account token (one zone, DNS:Edit) — a new privileged credential on the hub; (c) a negative probe: with the pasted token, try to READ the DNS records of another customer's zone by its id (the hub knows the ids) and refuse on success — catches the read side only. | **OPEN — needs the operator** (a is free today; b is a design) | operator: which of a/b/c | Pick a/b/c; if nothing: the check stands as built and the limit stays written in `01` §7 | operator | diff --git a/documentation/runbooks/secrets.md b/documentation/runbooks/secrets.md index 5fec9d3f..6e428763 100644 --- a/documentation/runbooks/secrets.md +++ b/documentation/runbooks/secrets.md @@ -160,9 +160,69 @@ silently. Rotate the ep0 token with `proxmox-backup-manager user generate-token` --- -## Other committed secrets (tracked, NOT yet de-gitted — backlog) +## `felhom.secret.yaml` — ROTATED AND DE-GITTED 2026-10-09 (R-925) -`manifests/felhom.secret.yaml` still commits other plaintext secrets (`healthchecks-config` `SECRET_KEY` -/ `SUPERUSER_PASSWORD`, `umami-config`, `gitea-creds`). These are **out of scope** for the Resend -rotation but are the same hygiene problem; de-git them the same way (out-of-band `Secret/...` + -placeholder) when touched. Tracked here so the gap is visible. +**What was wrong.** `manifests/felhom.secret.yaml` committed three Secrets in plaintext, and +`gitea.dooplex.hu` serves this repo to **anyone on the internet with no login** (measured from +off-network: a real path returns the file, a nonsense path returns 404). Every committed value was +still live — nothing had ever been rotated. Worst of all, `gitea-creds` was the **Gitea `admin` +account password** (`is_admin: true`, `/api/v1/admin/users` answered 200), and the *same string* was +reused as the admin password in **15 cluster Secrets** across the homelab. + +**What was done, in this order** (the order matters — de-gitting alone kills nothing, the history +keeps the value): + +1. **`umami-config`** — new random `APP_SECRET` and `POSTGRES_PASSWORD`. The password was changed + **inside Postgres** (`ALTER USER umami WITH PASSWORD`) as well as in the Secret; the env var alone + would not have changed it, because it is only read when the database is first initialised. +2. **`healthchecks-config`** — new random `SECRET_KEY` and `SUPERUSER_PASSWORD`. Nothing consumes + them (there is no healthchecks Deployment), so this was free. +3. **`gitea-creds`** — the hub no longer holds the admin password at all. It now holds a **scoped + Gitea access token** (`read:package` + `read:repository`). Both scopes are required and were each + verified against their real endpoint *before* the swap: + `GET /v2/admin/felhom-controller/tags/list` (the registry version checker) and + `GET /admin/felhom-controller/raw/branch/main/controller/configs/controller.yaml.example` + (the template fetcher — a `read:package`-only token gets **403** here, which is how the missing + scope was found). +4. **The Gitea `admin` password** was changed to a new random value and + `gitea-system/gitea-admin` updated. Verified: the new password returns **200**, the old published + one returns **401**. + +**The values** were written to a 0600 file on DooPlex (`~/rotated-secrets-2026-10-09.txt`) and are to +be moved into the password manager and the file deleted. They were never echoed to a terminal. + +**The file is gone from git** and `.gitignore`'s `*secret*` rule now applies to it (it never did +before: **`.gitignore` is not consulted for a file git already tracks**, which is why the rule looked +broken). `scripts/manifest_bearer_gate.py`'s `KNOWN_BACKLOG` carve-out was removed in the same +commit, so a reintroduction now fails the gate. + +### Recreating these Secrets (they are no longer in `manifests/`) + +ArgoCD does not manage them any more, exactly like `Secret/resend-api` above. The `felhom` app has +`automated.enabled: false` and no prune, so removing the file does **not** delete them. To recreate +one from the password manager, on 192.168.0.180 as `kisfenyo`, writing the value to a 0600 temp file +first so it is never echoed: + +```bash +# umami-config: APP_SECRET + POSTGRES_PASSWORD. +# If POSTGRES_PASSWORD changes, ALTER USER in the database too, or umami cannot connect: +# kubectl exec -n felhom-system deploy/umami-db -- psql -U umami -d umami \ +# -c "ALTER USER umami WITH PASSWORD '';" +sudo kubectl create secret generic umami-config -n felhom-system \ + --from-file=APP_SECRET=/path/app_secret --from-file=POSTGRES_PASSWORD=/path/pg_pw \ + --dry-run=client -o yaml | sudo kubectl apply -f - + +# gitea-creds: username=admin, password = a SCOPED TOKEN (read:package + read:repository), +# never the admin account password. +sudo kubectl create secret generic gitea-creds -n felhom-system \ + --from-literal=username=admin --from-file=password=/path/token \ + --dry-run=client -o yaml | sudo kubectl apply -f - + +# healthchecks-config: only needed if healthchecks is ever deployed; wire EMAIL_HOST_PASSWORD to +# secretKeyRef resend-api/RESEND_API_KEY rather than inlining it (see the Resend section above). +``` + +**Still owed, and NOT done here** (they are outside this repo): the same password is still the admin +password in ~12 other namespaces — `nextcloud`, `paperless`, `bookstack` (a **database root** +password), `tandoor`, `calibre`, `adventurelog`, `gokapi`, `qbittorrent`, `servarr`, `homepage`. +Rotating Gitea does not touch those; each is its own login and each is still the published string. diff --git a/manifests/felhom.secret.yaml b/manifests/felhom.secret.yaml deleted file mode 100644 index 2746fe7c..00000000 --- a/manifests/felhom.secret.yaml +++ /dev/null @@ -1,51 +0,0 @@ -apiVersion: v1 -kind: Secret -metadata: - name: healthchecks-config - namespace: felhom-system -type: Opaque -stringData: - # === REQUIRED: Generate a random key === - # python3 -c "import secrets; print(secrets.token_urlsafe(50))" - SECRET_KEY: "jumZn0XOcO1oDs77siMCgfkg0S2JGLHUiKBAxGleUF0KodBg6SHj-mcLNPxt29Wb6pk" - - # === REQUIRED: Superuser for first login === - SUPERUSER_EMAIL: "admin@felhom.eu" - SUPERUSER_PASSWORD: "doodooP4ssWD001!" - - # === REQUIRED: SMTP via Resend.com === - EMAIL_HOST: "smtp.resend.com" - EMAIL_PORT: "587" - EMAIL_HOST_USER: "resend" - # Resend API key — NOT committed. Healthchecks is not currently deployed; when it is, wire its - # EMAIL_HOST_PASSWORD to the out-of-band Secret/resend-api (key RESEND_API_KEY) via secretKeyRef - # instead of inlining a value here. See documentation/runbooks/secrets.md. - EMAIL_HOST_PASSWORD: "" - EMAIL_USE_TLS: "True" - EMAIL_USE_VERIFICATION: "False" - DEFAULT_FROM_EMAIL: "monitoring@felhom.eu" ---- -# NOTE: the Resend API key (formerly Secret/contact-mailer-config RESEND_API_KEY) is no longer -# committed. It lives in the out-of-band Secret/resend-api (key RESEND_API_KEY), created imperatively -# per documentation/runbooks/secrets.md. Both the hub and contact-mailer now read from resend-api. ---- -apiVersion: v1 -kind: Secret -metadata: - name: umami-config - namespace: felhom-system -type: Opaque -stringData: - APP_SECRET: "65cee3c4826b2478bcd304d81d3f2193983544d159a7b89dbf2d528479927a86" - POSTGRES_PASSWORD: "a764dc950f1f065d8b6f5e402d6420dd" ---- -apiVersion: v1 -kind: Secret -metadata: - name: gitea-creds - namespace: felhom-system -type: Opaque -stringData: - username: "admin" - password: "doodooP4ssWD001!" ---- \ No newline at end of file diff --git a/manifests/umami.yaml b/manifests/umami.yaml index 30ad07e3..953de128 100644 --- a/manifests/umami.yaml +++ b/manifests/umami.yaml @@ -208,12 +208,17 @@ spec: value: "1" - name: TZ value: "Europe/Budapest" + # 1Gi, NOT 512Mi (raised 2026-10-09). MEASURED: at 512Mi this pod runs fine for months but + # CANNOT RESTART — startup (Prisma migrate + Next.js) peaks above the limit, so the container + # reaches "Ready", is OOMKilled (exit 137) and enters CrashLoopBackOff. Found the hard way + # during the R-925 secret rotation: a `rollout restart` took stats.felhom.eu down (503) until + # the limit was raised. Steady state is well under this; the headroom is for startup only. resources: requests: - memory: "128Mi" + memory: "256Mi" cpu: "50m" limits: - memory: "512Mi" + memory: "1Gi" cpu: "500m" livenessProbe: httpGet: diff --git a/scripts/manifest_bearer_gate.py b/scripts/manifest_bearer_gate.py index 04f0e339..a29ccca6 100644 --- a/scripts/manifest_bearer_gate.py +++ b/scripts/manifest_bearer_gate.py @@ -21,11 +21,13 @@ ROOT = "manifests" # longer strings still match at 64+, but ordinary short ids never do). BEARER = re.compile(r"(?