# Felhom Controller Architecture — Part 1: Topology & Trust > **How to read this document.** Two kinds of statement appear, and where this document marks them it > marks them like this — the same wording as `07-backup-architecture.md:11-17`, carried here on > 2026-08-22 (R-376) so a reader meets one convention and not eight: > > - **[DESIGN]** — a decision taken. Not derived from code; the code may not implement it yet. > - **[FACT]** — an observed property, carrying a `file:line`, a live command output or a citation. > > **Statements in this document are NOT yet all marked.** Marking them wholesale is a large judgement > exercise and a wrong mark is worse than none, so only what a session touches is marked (R-376). > **An unmarked statement therefore means "not yet classified", never "observed".** That ambiguity is > exactly what cost this project three sessions in August 2026: the hot/bulk placement decision sat > unmarked beside a marked `[FACT]`, and was read as an observation and reported as a defect. **Status:** draft (decisions from the topology/trust design sessions). **Platform facts** referenced here live in `docs/proxmox-platform.md`; this document records *Felhom's decisions*, not Proxmox behaviour. --- ## 1. Model at a glance Three components. **Control is always box-initiated** — the hub never connects *into* a customer box. ``` operator side customer box (per Proxmox host) ┌───────────────────┐ ┌───────────────────────────────────────────┐ │ HUB │ │ Proxmox host │ │ (dooplex.hu, k3s) │ │ ┌──────────────┐ │ │ - report sink │◀──poll──┤ │ HOST AGENT │ operator-tier │ │ - signed jobs │ signed │ │ (Proxmox │ • all Proxmox ops │ │ - dashboard │ jobs │ │ token) │ • provision / restore │ │ - customer record│ │ └──────┬───────┘ • storage mgmt │ │ - PBS namespace │ │ │ local constrained API │ └─────────▲─────────┘ │ ┌──────▼───────────────────────────────┐ │ │ │ │ customer LXC (one per customer) │ │ │ direct, app- │ │ ┌──────────────┐ Docker: │ │ └───────────────────┼───┤ │ IN-GUEST │ [app] [app] ... │ │ domain reports │ │ │ CONTROLLER │ (Docker containers)│ │ │ │ (Docker-only)│ │ │ │ │ └──────────────┘ │ │ │ └───────────────────────────────────────┘ │ └───────────────────────────────────────────┘ PBS (offsite) ◀── outbound, client-side-encrypted backups ── customer box end-users / customer ◀── Cloudflare Tunnel ── apps + controller UI ``` --- ## 2. The customer node - One **Proxmox host** per box (PVE 9.2, Debian 13, LVM-thin). - **Default workload topology:** one **customer LXC**, Docker inside it, each app a Docker container/stack. Apps are isolated at the Docker layer (separate containers, networks, volumes, cgroup limits); they share one LXC/kernel/Docker daemon. - **Escape hatch:** promote an individual app to its own guest (LXC or VM) only for a specific reason — a non-Linux/Windows app, a genuinely untrusted or exposed app needing hard isolation, or a resource hog needing guarantees. - **Multi-tenant:** one customer per host is the home default; multiple customer LXCs on one host (a company environment) is **not precluded** — the agent manages a *set* of guests. The only multi-tenant-specific work deferred to "if it becomes real" is resource fairness (per-guest disk/RAM/CPU quotas). - **A scratch guest is a second guest of the SAME customer, unenrolled** (operator ruling 2026-09-13, R-481; built as LXC 9202 on demo-hp, persists). The hub ties one host to one customer, so a second enrolled customer on a box is not a thing the product does. The scratch guest runs with the hub, the tunnel, the agent link, off-site and self-update all off, and its controller image is set by hand — the one place that is allowed. Two of its properties were *decided by CC unattended — operator may reverse*: **it never binds the customer's real data drive** (a throwaway must not be able to reach real data) and **it never starts cloudflared** (a second connector would serve the public domain from a scratch box). --- ## 3. Components & responsibilities | | **Hub** | **Host agent** | **In-guest controller** | |---|---|---|---| | Runs on | dooplex.hu (k3s) | the Proxmox host | the customer LXC | | Tier | operator backend | operator (high-privilege) | customer-facing (app) | | Holds | customer records, signed-job source, PBS namespaces, escrowed keys | the **only** Proxmox API token; per-host operator identity | **no Proxmox creds**; its own hub API key + a local-API token to the agent | | Does | reporting sink, dashboard, job queue, source of durable truth | all Proxmox ops (provision, restore, snapshot, backup, storage mgmt, LXC lifecycle); polls hub for signed jobs; exposes a constrained local API to the controller; **per-guest authorization gate** | Docker/app lifecycle, catalog deploy, customer UI, app-level (data-layer) backup; reports app-domain to the hub directly | | Never does | initiate a connection *into* a box | — | touch the Proxmox API directly | **Key separation:** the controller manages Docker; the agent manages Proxmox. The controller's only path to guest-level operations (snapshot-before-deploy, "grow my RAM") is a constrained **local API call to the agent**, which the agent authorizes (scoped to that controller's own guest) and executes with its operator-tier token. This consolidates all Proxmox access and all per-guest authorization in one auditable place and leaves the guest with zero Proxmox credentials. --- ## 4. Control plane — box-initiated - CGNAT does **not** force this: the Cloudflare Tunnel already makes a box reachable through Cloudflare's edge. We *choose* box-initiated control for the smallest attack surface — the box exposes no control endpoint at all. - The agent and the controller **poll** the hub; the hub never initiates inbound. - Operator actions are delivered as **signed jobs**: the agent verifies an operator signature before executing, so a compromised hub database alone cannot forge commands. - All operator-initiated actions are recorded in a **customer-visible audit log**. --- ## 5. Trust boundaries | Boundary | What crosses | Mechanism | Blast radius if breached | |---|---|---|---| | end-user ↔ apps | app traffic | Cloudflare Tunnel → Traefik (Host routing) | that app | | customer ↔ controller UI | management UI | Cloudflare Tunnel; UI auth (bcrypt) | the customer's own box | | controller ↔ agent | snapshot/resize/backup requests | local constrained RPC; agent authorizes per-guest | the controller's own guest only | | agent ↔ hub | reports + signed jobs | outbound poll; signed jobs | one box; signed jobs limit forgery | | controller ↔ hub | app-domain reports/jobs (incl. geo desired-state) | outbound, own API key | app-domain of one customer | | box ↔ PBS | encrypted backups | outbound; per-customer namespace; client-side encryption | ciphertext only (operator can't read) | | guest ↔ Proxmox host | **(none direct)** | the guest holds no Proxmox creds; all via the agent | — | | hub ↔ Cloudflare API | geo-restriction WAF (enforcement) | the **hub** holds the CF API token; reconciles geo desired-state → WAF | the customer's zone/WAF | **Every app is on the internet from its first minute** (`*.domain` through the tunnel), so an app's FIRST admin login is a trust boundary too: a default password, or a "first visitor creates the admin" screen, is open to a stranger until the household acts. Rule and per-app status: `09` §3 decision 45 and `app-catalog-felhom.eu/FIRST-ADMIN.md` (the audit of all 53 apps, 2026-09-28). **Who is the visitor — the address (recorded 2026-10-01, R-753, `09` §3 decision 63, controller ≥ 0.286.0).** Rule: *never believe an address a client can write.* Through the tunnel every visitor used to reach traefik as cloudflared's one docker-assigned address, so every per-address guard (the dashboard's login counter, an app's lock) was an "everyone" guard a stranger could aim at the household. Now cloudflared has a fixed address and traefik trusts forwarded headers from that one address only: an app receives `X-Forwarded-For: , , 172.16.253.2` (Cloudflare APPENDS to a client's own chain — measured), a LAN visitor arrives as itself, and every other peer's chain is dropped. traefik's entrypoint middleware `felhom-forwarded` removes every header a client could write a host, a path or an address into (`X-Forwarded-Host`, `Forwarded`, `True-Client-Ip`, …) and fixes `X-Forwarded-Port: 443`. **Readers take the visitor from the RIGHT.** The controller (`clientaddr.go`) believes the chain only from traefik, takes the hop traefik saw, and for the tunnel's hop reads `CF-Connecting-IP` (the edge refuses a client-sent one). Catalog apps that read the LEFTMOST entry have the chain removed on their router. Design and measurements: `audits/visitors-2026-10-01/A/DESIGN.md`. A permanent household gate with family accounts in front of an app (Grimmory, MeTube) was SPIKED on this base and passed (`audits/permanent-gate-2026-10-01/VERDICT.md`) and is BUILT since controller 0.287.0 (`09` §3 decision 64) — the family gate, at the end of this section. **Who may reach an app, and through what (recorded 2026-09-29 — no document said it before; spike finding F1).** Every app is reached only through the box's traefik (no catalog app publishes a host port except crafty-controller's game ports; none uses host networking — read from the catalog 2026-09-29). traefik routes by host name: the tunnel's `*.domain` and the LAN both land there. The dashboard (`felhom.`) has its own password; its session cookie is **host-only** and never reaches an app host. An app answers anyone who reaches its host, with the app's own login — **except while its setup gate is closed** (`09` §3 decision 46, controller ≥ 0.280.0): then traefik asks the controller first (`forwardAuth`), and only a browser holding a gate cookie for that one host gets through. The cookie is minted after a valid dashboard session vouched for the browser (a 60-second, one-use token bound to the host, on the dashboard's own `/__gate/start`). The controller is in an app's request path ONLY while its gate is closed; once open, the gate's traefik router is removed and the app is reached exactly as without it. The gate decides who creates the first admin. **Who may sign up afterwards** is `09` §3 decision 47 (operator ruling 2026-09-29): once the first admin exists, open sign-up is closed; only the admin adds people, from the app's own user page. An app that cannot close it says so on its page. Mechanism (controller ≥ 0.281.0): the box keeps a small traefik router on the app's own sign-up address once the gate opens, answered "sign-up is closed" by the controller; the household opens it for 15 minutes from the app page to let a family member in. The controller is in THAT address's path only. Since controller 0.282.0 there are two locks where the app has its own switch: the box also sets the app's own "no sign-up" setting (`after_setup`), and the block matches any letter case and extra slashes. An app installed before decision 47 gets both only when the household presses "Close sign-up now" (decision 49); wanderer, which is not gated, the same way. **The family gate (`09` §3 decisions 63/64, controller ≥ 0.287.0).** An app whose template says `family_gate: true` is PERMANENTLY behind a forwardAuth door, the setup gate's mechanism with three differences: it never opens; its priority is below the install hold, the setup gate and the sign-up block; and the cookie it accepts is a FAMILY cookie. Each family member has their own name and password (the dashboard's Biztonság → Család card adds, resets and removes them; the password is shown once, stored as a bcrypt hash in the controller's data). A member signs in on the dashboard host's `/__family/login`; that sets a family session cookie on `Path=/__family` only — **never accepted by the dashboard's own RequireAuth** — and each app host gets a host-only cookie through a 60-second one-use token, as the setup gate does. The session store is asked on EVERY request, so a reset, a removal or a logout ends access at the next request. The family sign-in locks per real visitor (5 a minute) and per name (10 in 10 minutes) — never "everyone" (it reads the visitor from the right, as above). `family_gate_except:` lists LITERAL path prefixes a phone or e-reader app calls (Grimmory's OPDS, Kobo, KOReader, Komga API); each becomes a router without the door, anchored at a segment boundary (`^/prefix(/|$)`), so the app's OWN login decides there and a look-alike (`/api/v1/opdsx`) stays behind the gate. The gate is recorded at install (and at a removed app's restore), so a catalog change never gates or un-gates an installed app. Controller down → the gated app answers an error, never the app. Measured on 9202: `audits/family-gate-2026-10-02/A/items.txt`. --- ## 6. Enrollment & identity - **Physical presence at provisioning** (on-site install, or pre-imaged-and-delivered). This removes any zero-touch remote-enrollment problem. - A **one-time retrieval code** mints durable identity. Single-use (burned on the successful config fetch) plus a short *pre-use* TTL; one-click regenerate for the only real failure case (fetch fails before anything is persisted). After the fetch, the code is irrelevant — everything downstream runs on durable credentials, so retries don't need it. - **Order:** the agent enrolls first (and, running as root at setup, mints its own scoped operator-tier Proxmox token), then provisions the customer LXC from the golden template and deploys the controller into it — injecting the controller's hub API key and its local-API token. The controller is the agent's product, never the other way around. - The **hub customer record is the durable source of truth**, and it survives box loss: identity, domain, **Cloudflare tunnel token**, **PBS namespace**, **storage manifest**, a **mirrored app inventory** (bottom-up reality, not operator-declared intent — apps themselves restore from the PBS guest snapshot, never re-deployed from this record; see `05` §1/§9), and the **escrowed (zero-knowledge) backup key**. This is what makes hardware replacement possible. --- ## 7. Networking - **Cloudflare Tunnel** provides inbound access to apps and the controller UI (the CGNAT solution). Tunnel token lives in the hub record → **reused on new hardware during DR**, so DNS/routing stay intact through an outage. - **Outbound only** for control/report/backup (poll to hub, push to PBS). No inbound control endpoint exists in the chosen model. - **Every customer has their OWN domain — never a name under `felhom.eu`** (operator ruling 2026-09-14). The free Cloudflare tier's certificate covers one level below a zone, so nested names such as `felhom..felhom.eu` are not covered; the customer's own domain is a few thousand forints a year and is **included in the customer's price**. The dashboard is `felhom.`, each app `.`. **The tunnel is created by the operator** per `runbooks/day0-install.md` A.1 today; the hub creating it at customer creation (for a domain already on Cloudflare) is register row R-494, P3, not blocking. Measured reason this was ruled now: the 2026-09-14 first-hour drill used a `*.felhom.eu` customer domain with no tunnel, and the dashboard link in the setup-code mail did not resolve. - **Tunnel placement: INSIDE the guest** (corrected 2026-10-01, R-754 — the operator's brief of that evening: the build is right, correct the document). `cloudflared` is a container the CONTROLLER renders and keeps up (`internal/infra`, `EnsureBaseStack`, a protected stack), with the tunnel token from `controller.yaml`. *This page used to say it ran on the Proxmox host as an agent-managed systemd service; no box has ever been built that way* (read: the controller's template; seen running in demo-hp's guest 9201 and in R-505's VM 331). Consequence, stated: the data path is NOT independent of the guest — a dead guest is an unreachable box, and the controller (not the agent) restarts the tunnel. Since controller v0.286.0 it sits ALONE on the `felhom-tunnel` network at the fixed address `172.16.253.2` (traefik at `.3`), so traefik can believe forwarded headers from it and from nothing else (§5). Geo-restriction WAF is **hub-enforced** (the hub holds the CF API token; the controller only reports geo desired-state). --- ## 8. Storage & backup > **2026-09-30 (operator, `09` §3 decisions 50–51):** a new box's apps are in the off-site copy by default (07 §6); > the whole-guest restore test takes only the box's OWN archives — an earlier box's archive left in the customer's > ep0 namespace is never this box's proof, and the three drill archives found there are removed. **Tiers** (escalating failure scope): | Layer | Mechanism | Survives | Note | |---|---|---|---| | Snapshot | LVM-thin snapshot (transient) | *logical* loss only | whole-LXC rollback; **not a backup** | | Local — second storage | vzdump to `dir`/`nfs`/`cifs` | primary-disk failure (USB) / box death (NAS) | first *real* backup tier | | Offsite — PBS | dedup'd, incremental, encrypted | site loss | the DR substrate; paid tier | - **Storage manifest** (hub-held, agent-reconciled): per target → type, durable identity (UUID / `server:/export` / repo+fingerprint), **class** (fast/slow + rough IOPS, set once at attach), role, encrypted credentials, schedule/retention. The agent creates the Proxmox storages, continuously checks presence/reachability, and reports per-target status (a disconnected target → actionable notification). - **[DESIGN] App data placement is per-volume, not per-app:** `.felhom.yml` classifies each volume **hot** (DB/config/cache → fast storage, enforced) vs **bulk** (media/files → may be slow). A photo app's DB stays on SSD while its blobs go to the USB. > **Marked [DESIGN] on 2026-08-22 (R-376), and the pointer is honest about what it can point at.** > **This decision was never recorded as a decision anywhere** — it was established by reading, not > by citation: it exists as this bullet and nowhere else, with no dated entry in the decision log > and no `R-` row. A log entry was written on 2026-08-22 (`CONTEXT.md`, "App data placement is a > DECISION") **to give it a home, not to claim it was decided then**; the choice is older than the > entry and its original date is not on record. > > **What being unmarked cost.** The consequence of this bullet — that 40 of 53 catalogue templates > declare no configurable path because they are all-hot — is stated as **[FACT]** at > `07-backup-architecture.md:296-299`. A reader met a marked observation beside an unmarked choice > and reasonably asked whether it *should* be so. **Between 19 and 22 August that reader called this > decision a defect in four places** (R-370), and one session's work went into correcting the record > rather than into the product. - **Backup scoping:** hot data (LXC rootfs) rides the guest `vzdump` → tiers + PBS. Bulk data on external mount points is **excluded** from the guest vzdump (per-mount `backup` flag) and gets its own per-volume policy (file-level to a tier, slower cadence — or explicitly *not* backed up for re-downloadable content, with the customer informed). - **Tiers double as the DR restore-source priority:** restore from the fastest *surviving* source (local if still attachable, PBS on true site loss). - **Key custody (zero-knowledge default):** three tiers the customer chooses — *customer-only* / *zero-knowledge escrow (default)* / *operator-managed*. Default escrows the **PBS passphrase-protected keyfile** in the hub, wrapped under a **customer recovery code** the operator can't open; DR needs the customer's code. Access-notification is an audit signal, never the primary guard. (Don't build bespoke crypto — use PBS's native keyfile passphrase.) --- ## 9. Disaster recovery - **Guest-loss (host + agent alive):** the agent restores the guest from the fastest surviving tier, **resets identity** (MAC/hostname — see `proxmox-platform.md`), boots it, controller returns. Validated mechanics: Phase 2. - **Host / hardware-loss (agent gone):** re-provision (§6) in **restore mode** — the hub, knowing the customer has PBS backups, hands the freshly-enrolled agent the existing identity + PBS namespace + a restore directive instead of a clean-provision directive. The agent restores from PBS; the controller returns on the same domain (tunnel reused from the hub record). DR = provisioning + a restore mode, not a separate mechanism. - **Snapshot-before-deploy:** controller asks the agent to snapshot, deploys, runs its post-deploy health check, asks the agent to roll back on failure. (Transient snapshot, §8.) --- ## 10. How this embodies the product values - **Zero-knowledge offsite** — the operator holds the offsite backup but cannot read it. - **Box-initiated control + signed jobs** — no standing operator backdoor; a hub compromise alone can't forge commands. - **Customer-visible audit log** — every operator action is visible to the customer. - **Never hold data hostage** — subscriptions cover ongoing labour (monitoring, offsite, support, new deployments); the customer's data and deployed apps remain recoverable by the customer (recovery code), with nothing locked behind the operator. --- ## 11. Open sub-decisions (carried into later parts) - **RTO/RPO targets** → drive the backup + offsite-replication schedule (§8). - Offboarding / decommission (scenario 6) — not yet designed; must honour "never hold data hostage" in credential revocation + data hand-off. - Multi-tenant resource fairness — deferred until multi-tenant is real (§2). --- ## Appendix — relationship to the spike - **Phase 0** → §2: LXC-default for the workload; overhead numbers. - **Phase 1** → §3/§5: validated the privilege boundary (create/allocate is operator-tier). The guest-side scoped-backup-token it proved possible is **not** used — we chose the agent-mediated path — but it confirmed restore = operator-tier, which shapes the agent. - **Phase 2** → §8/§9: backup→restore round-trip; identity reset on restore. --- ## Changelog — design-review + Phase-3 fold-in (2026-06-08) - §5 trust boundaries: **added `hub ↔ Cloudflare API`** row (hub holds the CF token, enforces geo→WAF); controller↔hub row notes it carries geo desired-state (S4). - §7 networking: **tunnel placement resolved → host** (agent-managed systemd service); geo is hub-enforced (S4/S5). - §11 open items: removed the now-resolved **tunnel placement** and **self-update flow** entries (S5; self-update designed in 03 §11). - §6 durable record: **"declarative app inventory" → "mirrored app inventory"** — aligns the wording with the locked two-driver model (`05` §1: apps are bottom-up mirror, never operator-declared; `05` §9: apps restore from the PBS guest snapshot, not re-deployed from this record).