diff --git a/STATUS.md b/STATUS.md index fc31f66e..5e7e0afa 100644 --- a/STATUS.md +++ b/STATUS.md @@ -35,6 +35,27 @@ - **Still to prove:** the lost off-site copy alarm (Tester 1's next clean-up window, about 12 October), and the failed-restore hold (needs a scratch off-site store). +## Needs you soon (2026-10-09): passwords are sitting in a repository anyone on the internet can read + +I found this while publishing the legal pages, not because I went looking — a security check flagged something +small in the new page, and following it led here. + +- **What is true:** `gitea.dooplex.hu` answers to anyone, with no login, from anywhere on the internet. I proved + that from outside your network, not just from DooPlex. One of the files it hands out is a manifest that + contains **real passwords**, not placeholders: the visitor-counter's database password and app key, and the + login for the health-check tool. I did not print or copy any of the values. +- **How bad:** not as bad as it sounds, and I want to be exact. These guard the **statistics and monitoring + bits, not any household's data.** The database itself is not reachable from the internet — I checked. No + customer box, no hub key, no backup key is in that file. But they are live passwords, publicly downloadable. +- **Your own runbook already says this must not happen** — "secret values are never committed to git" — so the + rule is right and the repository is breaking it. +- **What to do, in this order:** (1) **change those passwords** — deleting them from the files changes nothing, + the history keeps them; (2) decide whether that repository should be readable by strangers at all, bearing in + mind the installer people download depends on something public; (3) then I move the values out properly and + make the check refuse them in future. +- **I changed nothing.** Making the repository private could break the installer that new boxes fetch, and + changing passwords on live services is yours to time. It is written up as a register item with the evidence. + ## Legal pages (2026-10-09): the website finally says what it does with people's data - **Two new pages are live:** (what we do with data) and diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 39f381dc..f45af33a 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -192,7 +192,7 @@ stopping line that lies. | **R-331** | Storage & devices | P4 | **Disk health Phase 3 — growth-rate detection, and retiring the static 64.** The v0.215.0 count backstop (64 unreadable sectors → Hiba) is **a judgement from ONE drive**: the observed benign excursion peaked at 16 and cleared inside an hour, and the terminal run passed 64 at 13 Aug 11:28 and never came back. It is deliberately a backstop BEHIND the sustain rule, not the primary signal, but it is still a magic number tuned on a single sample and it will be wrong for some drive. With Phase 2's history the box can ask the question that actually matters — *is this count climbing, and how fast* — which distinguishes a drive with eight stable aging sectors from one adding forty a day, something no static threshold can do. Revisit 64 when that exists | **READY (M) — NEW 2026-08-14** | R-330 | Growth-rate rule over persisted samples; re-derive or delete the static 64 | CC | | **R-352** | Storage & devices | P4 | **Four screens state something untrue about where an app's data goes, and the configured default is consulted by nothing that places data.** Measured on `demo-hp` 2026-08-21. **(1)** **40 of 53** catalogue templates declare no data path (`grep -rl 'env_var: HDD_PATH' --include='.felhom.yml'` → 13; total 53); those apps get **no storage field and no default** — their data lands in a named Docker volume on the system drive. **(2)** `GetDefaultStoragePath()` has exactly **three** non-test callers — the metrics collector (`cmd/controller/main.go:410`), the dashboard SystemInfo panel (`web/server.go:733`) and `.fab` import landing (`handler_export_upload.go:154`). **The deploy route never reads it.** Its field comment `// new apps use this by default` (`internal/settings/settings.go:453`) has never been true — an invariant with no test pinning it. **(3)** The first-tier backup follows the data onto the same disk (`backup/backup.go:324-334` → `systemDataPath`), so data and nearest copy share one device for a customer doing nothing wrong — the posture Tier 2 refuses outright at `tier2.go:329`. **(4)** „1 alkalmazás használja" on the Drives page counts only `Env["HDD_PATH"] == path` (`web/handlers.go:2118`), so it can never include the 40-class; it truthfully means *„1 of the apps that CAN use a drive does"*. | **NARROWED** — **PARTLY CLOSED 2026-08-21** — visibility shipped; **placement OPEN** | — | **Shipped tonight (visibility only, no placement change, nothing migrated):** the deploy page now states where the app's data will live before the button is pressed, naming the system drive for the 40-class and the selected drive for the 13. **⚠ RE-FRAMED 2026-08-22 — THE FOUR MEASUREMENTS STAND; TWO OF THE CONCLUSIONS DRAWN FROM THEM DO NOT.** (1) is a measurement and is correct, but "those apps get no storage field and no default" is not a deprivation: the architecture places **hot** data (DB/config/cache) on fast storage inside the guest and states that placement is **ENFORCED** (`documentation/architecture/01-topology-and-trust.md:150-152`). The 40 are all-hot apps; the 13 are the ones with **bulk** content, which belongs on an attached drive. There is no choice being denied. (3) **overstated one risk and understated a distinction.** Since R-165 the guest carries a small OS rootfs plus **ONE** data volume at `/var/lib/felhom`; `/var/lib/docker` and `/mnt/sys_drive` are two **binds of that same volume** (`felhom-agent/configs/build-golden.sh:29-40, 99`) — the `mp0`/`mp1` split assumed here was retired 2026-08-03. **Real risk:** a physical-disk failure loses the data and its first-tier copy together — which is what the off-site and whole-machine tiers exist for, and which is equally true of a drive-resident app whose unit sits beside its data by design. **Overstated risk:** a full data volume stopping the operating system — the OS rootfs is a separate volume and the capture floor refuses per app before exhaustion (`00-capability-map.md:94`), watched working 2026-08-21 with the volume at 99% and all 15 containers healthy. **The comparison to Tier 2's same-disk refusal (`tier2.go:329`) is withdrawn:** Tier 2 refuses a SECOND copy on the same disk; Tier 1's unit is meant to sit beside the data. **(2) and (4) are untouched and remain correct** — (2) is now filed on its own as **R-368** with its scope measured, and (4) needs no ruling: the count is honest and only easy to misread. **The specification for the rest is filed at `documentation/backlog/SPEC-app-data-placement-2026-08-21.md`** (corrected 2026-08-22, framing marked inline, measurements kept) and lists the five points a ruling must settle (compose-template vs controller, existing deployments, when the SSD is legitimately right, `IsDefault` must become true or go away *with a test*, and the Drives-page count). **An earlier recommendation to refuse deployment until a drive is registered was WITHDRAWN** — it assumed the customer had failed to choose; they had no choice to make. | **Viktor rules**, CC executes | -## Security & access — 15 rows (P2 1, P3 10, P4 4) +## Security & access — 16 rows (P2 2, P3 10, P4 4) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -208,6 +208,7 @@ stopping line that lies. | **R-870** | Security & access | P3 | **Tester 1's two Cloudflare credentials — the zone API token (`infrastructure.cf_api_token`) and the tunnel token (`infrastructure.cf_tunnel_token`) of the hub's `customer_configs` row `tester-1` — were printed into the 2026-10-04 night session's transcript** (not into any file): a read-only query selected `substr(config_json,1,400)`, and both values sit in the first 400 characters. Tester 1 is CC's disposable test customer (`enkicsifelhom.hu`). **Not rotated, by the operator's ruling of 2026-10-05 06:49 (option B).** **Rotation, whenever chosen (3 steps):** in the Cloudflare dashboard create a new API token for the `enkicsifelhom.hu` zone with the same permissions and refresh the Tester 1 tunnel's token (Zero Trust → Networks → Tunnels → the tunnel → refresh token) → hub → Configs → `tester-1` → Edit → the two Cloudflare fields → Save, then confirm on the box that cloudflared reconnected (`docker ps` health `healthy`) → delete the old API token. Rule for sessions (as R-831): never select a whole config row — name the fields, and never `config_json` without `json_extract` of a non-secret field. | **WAITING-ON-OPERATOR — rotation is his call (ruled: not now)** **Not rotated by the operator's rulings (2026-10-04 „keep using the current one"; 2026-10-05 option B) — restated 2026-10-05 18:23; the steps stay here.** | — | rotate when chosen | operator | | **R-908** | Security & access | P3 | **An old Resend API key is still in the history of the `homelab-manifests` repository; it was replaced on 2026-06-29, but whether it is also revoked at Resend is unknown.** FOUND 2026-10-08 by the contact-mailer source search (`audits/mailer-source-2026-10-08/SEARCH.md`, „A side finding"): commit c9648cd (2026-02-05, „added mailer pod") put a key literal in the old `felhom-system/contact-mailer.yaml` comment; the file was deleted in ee93b50 but the history keeps it. Compared by hash only, never printed: it is NOT today's key (`Secret/resend-api`, rotated 2026-06-29, felhom.eu feea0606). If the old key still works at Resend, anyone with read access to that repository can send mail as felhom.eu. | **OPEN — owner: operator** | — | In the Resend console, check that the old key is revoked (revoke it if not); rewriting the repository's history is NOT needed once it is revoked | operator | | **R-525** | Security & access | P4 | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** **Re-ranked 2026-10-03: P3->P4: comfort feature needing a new unmeasured mechanism; the default password hole is closed.** | — | — | CC | +| **R-925** | Security & access | P2 | **`manifests/felhom.secret.yaml` carries REAL secret values, and the Gitea repo that holds it is readable by anyone on the internet with no login.** FOUND 2026-10-09 by a background security review of the legal-pages commit, which flagged the served `` comments; chasing what those comments pointed at found this instead. **MEASURED, in this order:** (1) `scripts/manifest_bearer_gate.py` itself prints `manifests/felhom.secret.yaml:39 KNOWN-BACKLOG committed secret 65cee3c4…7a86`; (2) the file holds **live-shaped values, not placeholders** — `SECRET_KEY` (69 chars), `SUPERUSER_EMAIL`/`SUPERUSER_PASSWORD` (18), Umami `APP_SECRET` (66) and `POSTGRES_PASSWORD` (34), and a `username`/`password` pair (18); values were never printed, only measured by length; (3) `gitea.dooplex.hu` resolves to **37.191.56.193**, the same public address as the website; (4) an **anonymous** `curl` returns HTTP 200 and 1686 bytes for that exact file, and 200 for `hub/internal/store/store.go` and `manifests/hub.yaml`; (5) **the off-network control**, which is what makes (4) mean anything — fetched from outside the operator's network entirely: a real path returns the file's first line, a nonsense path returns 404, so the 200 is genuine anonymous read from the internet and not a LAN-only ACL. **This contradicts the project's own written model**: `documentation/runbooks/secrets.md:3-4` says *“Secret values are never committed to git. The manifests in `manifests/` carry only placeholders + comments that point here.”* That promise is false today, which is the P1 wording in this register's own scale — see the ranking note below. The same runbook already states the rule that governs the fix: *“The git-history copy stays alive until the value is ROTATED — de-git alone kills nothing.”* **Known neighbours, none of which cover this:** R-887 records that an outside crawler walks the public Gitea pages and leaves *“Gitea's exposure to the crawler”* as the operator's undone item; R-580 still owes a Gitea admin token rotation. Neither says a secret file is world-readable. **Why P2 and not P1, stated so the operator can overrule it:** the exposed credentials guard **analytics and an undeployed healthchecks instance, not household data** — measured: `umami-db` is a **ClusterIP** service on 5432 with no external IP, so the Postgres password is not reachable from the internet; the realistic harm is session forgery against the public `stats.felhom.eu` via `APP_SECRET`, and reuse of those passwords anywhere else. No customer box, hub token or escrow key is in this file. **CC did NOT change anything**: making the repo private could break the public day-0 path (the installer is fetched from a public tag, R-110), and rotation plus repository visibility are operator decisions on production infrastructure. | **OPEN — owner: operator (rotate + decide repo visibility); CC can do the de-git and the gate once told** | — | Operator, in this order: **(1) ROTATE** the Umami `APP_SECRET` and `POSTGRES_PASSWORD` and the healthchecks `SECRET_KEY`/`SUPERUSER_PASSWORD` — de-gitting first changes nothing, the history keeps the old values; **(2) decide** whether `gitea.dooplex.hu` should be readable anonymously at all, remembering the public installer tag depends on some public path (R-110); **COUPLED, and easy to miss:** `website/adatkezeles.html` serves 32 `` comments naming `hub/internal/...`, `manifests/...` and architecture paths — harmless while the repo is public (they point at files anyone can already read), but **the moment the repo is made private those comments become a map of it**, so step (2) and stripping them are ONE change, not two. The traceability they carry is already kept in `documentation/legal/*-1.0.md`, so stripping costs nothing. **(3) then** tell CC to move the values out of `manifests/` into out-of-band Secrets per `secrets.md` and extend `manifest_bearer_gate.py` so the KNOWN-BACKLOG exemption becomes a failure. If nothing is done: live credentials stay downloadable by anyone who finds the repo, and the secrets runbook keeps stating a rule the repo breaks | operator | | **R-779** | Security & access | P4 | **[P3-LOW] Part A's "two outside addresses seen as two" is proven through the simulated tunnel only; on the REAL tunnel the second outside address (ep0, one request allowed) was refused by Cloudflare's edge with 403 and never reached the box.** Measured 2026-10-01 19:51 UTC (`audits/visitors-2026-10-01/A/L2-demo-hp-real-tunnel.txt`): no log line on demo-hp; demo-hp's box has no geo restriction in its settings, so a Cloudflare ZONE rule (country or bot, not read) refused a German datacenter address. DooPlex's own address on the real tunnel was seen as itself. **Needs:** one sign-in from a second Hungarian address (the operator's phone off wifi) while DooPlex is locked out — 2 minutes; and say which Cloudflare rule refused ep0. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (a phone), CC reads the logs** **Re-ranked 2026-10-03: P3→P4: a proof gap on the real tunnel; operator-only follow-up.** | — | — | CC + operator | | **R-904** | Security & access | P4 | **Cloudflare can read every household's app traffic; replacing it with our own relay is a later item.** Facts (reviewer discussion 2026-10-08; `01`): app traffic and the dashboard reach the box through the Cloudflare Tunnel (`01` §5 trust table, rows end-user ↔ apps and customer ↔ controller UI; §7), and the tunnel's public end is Cloudflare's edge, where TLS ends — so Cloudflare can technically read that traffic (the FAQ says so since 2026-10-08, R-900; „TLS ends at the edge" is not written in `01` — add it there). What Cloudflare gives today, free: inbound reach with no router setup, the CGNAT answer (`01` §4, §7); certificates (`01` §7, the free tier covers one level below a zone); the geo-WAF the hub enforces (`01` §5 last row, §7); flood protection (not in `01`). The alternative named: our own EU relay over WireGuard with TLS passthrough by SNI, certificates on the box, the geo-block on the relay. Its costs: one more machine the operator keeps up, and a single point of reach for every box; weaker flood protection; about a week of work after a spike. Operator ruling 2026-10-08 09:07 (`09` §3 decision 184): a later item. | **DEFERRED — after the first customers (operator ruling 2026-10-08 09:07)** | — | A spike after the first customers (the relay's reach, cost and flood behaviour, measured) | operator | | **R-913** | Security & access | P4 | **The Cloudflare token check (R-138, decision 190) reads what a token can SEE, not what it can WRITE.** FOUND 2026-10-08 by the security review of the R-138 build: `GET /zones` lists zones the token can read; a hand-built token with Zone:Read on the customer's zone and DNS:Edit on ALL zones would pass the check and could still change every household's DNS. The token wizard's single „Specific zone" scope does not build such a token, and a customer token cannot read its own policies (that needs „API Tokens Read"). Written as a limit in `01` §7. **Options:** (a) the operator mints every customer token from one recipe and the hub checks nothing more; (b) the hub mints the token itself with the operator's account token (one zone, DNS:Edit) — a new privileged credential on the hub; (c) a negative probe: with the pasted token, try to READ the DNS records of another customer's zone by its id (the hub knows the ids) and refuse on success — catches the read side only. | **OPEN — needs the operator** (a is free today; b is a design) | operator: which of a/b/c | Pick a/b/c; if nothing: the check stands as built and the limit stays written in `01` §7 | operator |