website: one picture viewer (assets/gallery.js) for app and dashboard pictures — declared data-gallery sets, a click on the big picture closes, JS-off still opens the file; site gate 20 + 4 decoys; demo-hp language read back (hu); R-908 (old Resend key in homelab-manifests history); 131 -> 132
gates / gates (push) Successful in 4m25s

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qRErfCoiTkvDK9N5XHbzb
This commit is contained in:
2026-10-08 13:14:24 +02:00
parent 460135a15a
commit 79c53b5f1a
27 changed files with 711 additions and 734 deletions
+2 -1
View File
@@ -191,7 +191,7 @@ stopping line that lies.
| **R-331** | Storage & devices | P4 | **Disk health Phase 3 — growth-rate detection, and retiring the static 64.** The v0.215.0 count backstop (64 unreadable sectors → Hiba) is **a judgement from ONE drive**: the observed benign excursion peaked at 16 and cleared inside an hour, and the terminal run passed 64 at 13 Aug 11:28 and never came back. It is deliberately a backstop BEHIND the sustain rule, not the primary signal, but it is still a magic number tuned on a single sample and it will be wrong for some drive. With Phase 2's history the box can ask the question that actually matters — *is this count climbing, and how fast* — which distinguishes a drive with eight stable aging sectors from one adding forty a day, something no static threshold can do. Revisit 64 when that exists | **READY (M) — NEW 2026-08-14** | R-330 | Growth-rate rule over persisted samples; re-derive or delete the static 64 | CC |
| **R-352** | Storage & devices | P4 | **Four screens state something untrue about where an app's data goes, and the configured default is consulted by nothing that places data.** Measured on `demo-hp` 2026-08-21. **(1)** **40 of 53** catalogue templates declare no data path (`grep -rl 'env_var: HDD_PATH' --include='.felhom.yml'` → 13; total 53); those apps get **no storage field and no default** — their data lands in a named Docker volume on the system drive. **(2)** `GetDefaultStoragePath()` has exactly **three** non-test callers — the metrics collector (`cmd/controller/main.go:410`), the dashboard SystemInfo panel (`web/server.go:733`) and `.fab` import landing (`handler_export_upload.go:154`). **The deploy route never reads it.** Its field comment `// new apps use this by default` (`internal/settings/settings.go:453`) has never been true — an invariant with no test pinning it. **(3)** The first-tier backup follows the data onto the same disk (`backup/backup.go:324-334` → `systemDataPath`), so data and nearest copy share one device for a customer doing nothing wrong — the posture Tier 2 refuses outright at `tier2.go:329`. **(4)** „1 alkalmazás használja" on the Drives page counts only `Env["HDD_PATH"] == path` (`web/handlers.go:2118`), so it can never include the 40-class; it truthfully means *„1 of the apps that CAN use a drive does"*. | **NARROWED** — **PARTLY CLOSED 2026-08-21** — visibility shipped; **placement OPEN** | — | **Shipped tonight (visibility only, no placement change, nothing migrated):** the deploy page now states where the app's data will live before the button is pressed, naming the system drive for the 40-class and the selected drive for the 13. **⚠ RE-FRAMED 2026-08-22 — THE FOUR MEASUREMENTS STAND; TWO OF THE CONCLUSIONS DRAWN FROM THEM DO NOT.** (1) is a measurement and is correct, but "those apps get no storage field and no default" is not a deprivation: the architecture places **hot** data (DB/config/cache) on fast storage inside the guest and states that placement is **ENFORCED** (`documentation/architecture/01-topology-and-trust.md:150-152`). The 40 are all-hot apps; the 13 are the ones with **bulk** content, which belongs on an attached drive. There is no choice being denied. (3) **overstated one risk and understated a distinction.** Since R-165 the guest carries a small OS rootfs plus **ONE** data volume at `/var/lib/felhom`; `/var/lib/docker` and `/mnt/sys_drive` are two **binds of that same volume** (`felhom-agent/configs/build-golden.sh:29-40, 99`) — the `mp0`/`mp1` split assumed here was retired 2026-08-03. **Real risk:** a physical-disk failure loses the data and its first-tier copy together — which is what the off-site and whole-machine tiers exist for, and which is equally true of a drive-resident app whose unit sits beside its data by design. **Overstated risk:** a full data volume stopping the operating system — the OS rootfs is a separate volume and the capture floor refuses per app before exhaustion (`00-capability-map.md:94`), watched working 2026-08-21 with the volume at 99% and all 15 containers healthy. **The comparison to Tier 2's same-disk refusal (`tier2.go:329`) is withdrawn:** Tier 2 refuses a SECOND copy on the same disk; Tier 1's unit is meant to sit beside the data. **(2) and (4) are untouched and remain correct** — (2) is now filed on its own as **R-368** with its scope measured, and (4) needs no ruling: the count is honest and only easy to misread. **The specification for the rest is filed at `documentation/backlog/SPEC-app-data-placement-2026-08-21.md`** (corrected 2026-08-22, framing marked inline, measurements kept) and lists the five points a ruling must settle (compose-template vs controller, existing deployments, when the SSD is legitimately right, `IsDefault` must become true or go away *with a test*, and the Drives-page count). **An earlier recommendation to refuse deployment until a drive is registered was WITHDRAWN** — it assumed the customer had failed to choose; they had no choice to make. | **Viktor rules**, CC executes |
## Security & access — 14 rows (P2 1, P3 11, P4 2)
## Security & access — 15 rows (P2 1, P3 12, P4 2)
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|---|---|---|---|---|---|---|---|
@@ -207,6 +207,7 @@ stopping line that lies.
| **R-782** | Security & access | P3 | **[P3-LOW] Two side observations of the R-753 sweep, inferred, not measured:** glance's seeded `glance.yml` has no `auth:` block (the dashboard is public to anyone with the address), and homepage's `/api/*` refuses a Host not in `HOMEPAGE_ALLOWED_HOSTS`, which the template does not set (widgets may 400). **Needs:** measure both on 9202; glance: decide whether a public link dashboard is intended (the setup gate does not cover it after setup). | **READY — rank P3-LOW; owner: CC (catalog)** **2026-10-06 night: homepage fixed on catalog branch `night-held-2026-10-06`** (`093e5ed`): `HOMEPAGE_ALLOWED_HOSTS=${SUBDOMAIN}.${DOMAIN}`; bench: `/api/services` 400 → 200. Glance measured: public (`GET /` 200 with no login, no auth block, `/login` 404) — whether it gets a login is the operator's question; the row stays open for it (`cat/R-782-red.txt`). | — | — | CC |
| **R-831** | Security & access | P3 | **The Hetzner storage API token (`HETZNER_TOKEN`, the storage project's token in `Secret/storagebox`) was printed into the 2026-10-03 session transcript** — CC read the gitignored `manifests/storagebox.secret.yaml` and its redaction pattern missed the quoted value. It can create, reset and delete Storage Box sub-accounts. Not rotated by the operator's choice (decision 73). **Rotation, whenever chosen (3 steps):** create a new token in the storage project in the Hetzner console → patch `Secret/storagebox` key `HETZNER_TOKEN` in `felhom-system` and `kubectl rollout restart deployment/hub` → delete the old token in the console. Rule for sessions: never print a file that holds secrets — read the one field needed. | **WAITING-ON-OPERATOR — rotation is his call** **Not rotated by the operator's rulings (2026-10-04 „keep using the current one"; 2026-10-05 option B) — restated 2026-10-05 18:23; the steps stay here.** | — | rotate when chosen | operator |
| **R-870** | Security & access | P3 | **Tester 1's two Cloudflare credentials — the zone API token (`infrastructure.cf_api_token`) and the tunnel token (`infrastructure.cf_tunnel_token`) of the hub's `customer_configs` row `tester-1` — were printed into the 2026-10-04 night session's transcript** (not into any file): a read-only query selected `substr(config_json,1,400)`, and both values sit in the first 400 characters. Tester 1 is CC's disposable test customer (`enkicsifelhom.hu`). **Not rotated, by the operator's ruling of 2026-10-05 06:49 (option B).** **Rotation, whenever chosen (3 steps):** in the Cloudflare dashboard create a new API token for the `enkicsifelhom.hu` zone with the same permissions and refresh the Tester 1 tunnel's token (Zero Trust → Networks → Tunnels → the tunnel → refresh token) → hub → Configs → `tester-1` → Edit → the two Cloudflare fields → Save, then confirm on the box that cloudflared reconnected (`docker ps` health `healthy`) → delete the old API token. Rule for sessions (as R-831): never select a whole config row — name the fields, and never `config_json` without `json_extract` of a non-secret field. | **WAITING-ON-OPERATOR — rotation is his call (ruled: not now)** **Not rotated by the operator's rulings (2026-10-04 „keep using the current one"; 2026-10-05 option B) — restated 2026-10-05 18:23; the steps stay here.** | — | rotate when chosen | operator |
| **R-908** | Security & access | P3 | **An old Resend API key is still in the history of the `homelab-manifests` repository; it was replaced on 2026-06-29, but whether it is also revoked at Resend is unknown.** FOUND 2026-10-08 by the contact-mailer source search (`audits/mailer-source-2026-10-08/SEARCH.md`, „A side finding"): commit c9648cd (2026-02-05, „added mailer pod") put a key literal in the old `felhom-system/contact-mailer.yaml` comment; the file was deleted in ee93b50 but the history keeps it. Compared by hash only, never printed: it is NOT today's key (`Secret/resend-api`, rotated 2026-06-29, felhom.eu feea0606). If the old key still works at Resend, anyone with read access to that repository can send mail as felhom.eu. | **OPEN — owner: operator** | — | In the Resend console, check that the old key is revoked (revoke it if not); rewriting the repository's history is NOT needed once it is revoked | operator |
| **R-525** | Security & access | P4 | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** **Re-ranked 2026-10-03: P3->P4: comfort feature needing a new unmeasured mechanism; the default password hole is closed.** | — | — | CC |
| **R-779** | Security & access | P4 | **[P3-LOW] Part A's "two outside addresses seen as two" is proven through the simulated tunnel only; on the REAL tunnel the second outside address (ep0, one request allowed) was refused by Cloudflare's edge with 403 and never reached the box.** Measured 2026-10-01 19:51 UTC (`audits/visitors-2026-10-01/A/L2-demo-hp-real-tunnel.txt`): no log line on demo-hp; demo-hp's box has no geo restriction in its settings, so a Cloudflare ZONE rule (country or bot, not read) refused a German datacenter address. DooPlex's own address on the real tunnel was seen as itself. **Needs:** one sign-in from a second Hungarian address (the operator's phone off wifi) while DooPlex is locked out — 2 minutes; and say which Cloudflare rule refused ep0. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (a phone), CC reads the logs** **Re-ranked 2026-10-03: P3→P4: a proof gap on the real tunnel; operator-only follow-up.** | — | — | CC + operator |
| **R-904** | Security & access | P4 | **Cloudflare can read every household's app traffic; replacing it with our own relay is a later item.** Facts (reviewer discussion 2026-10-08; `01`): app traffic and the dashboard reach the box through the Cloudflare Tunnel (`01` §5 trust table, rows end-user ↔ apps and customer ↔ controller UI; §7), and the tunnel's public end is Cloudflare's edge, where TLS ends — so Cloudflare can technically read that traffic (the FAQ says so since 2026-10-08, R-900; „TLS ends at the edge" is not written in `01` — add it there). What Cloudflare gives today, free: inbound reach with no router setup, the CGNAT answer (`01` §4, §7); certificates (`01` §7, the free tier covers one level below a zone); the geo-WAF the hub enforces (`01` §5 last row, §7); flood protection (not in `01`). The alternative named: our own EU relay over WireGuard with TLS passthrough by SNI, certificates on the box, the geo-block on the relay. Its costs: one more machine the operator keeps up, and a single point of reach for every box; weaker flood protection; about a week of work after a spike. Operator ruling 2026-10-08 09:07 (`09` §3 decision 184): a later item. | **DEFERRED — after the first customers (operator ruling 2026-10-08 09:07)** | — | A spike after the first customers (the relay's reach, cost and flood behaviour, measured) | operator |