Question A: YES, no code change. customer.domain is a trimmed string with no UNIQUE, no CHECK, no format rule (store.go:114, configs.go:673), copied verbatim into controller.yaml (configgen.go:48), and every one of its 30 consumers on the box interpolates it without parsing. Zero hits for registrable/eTLD/publicsuffix across both repos. Nothing creates DNS records (zero hits for dns_records) — the two Cloudflare clients are WAF-only. And the zone-ownership assumption is a SWITCH, not a requirement: traefik.yml.tmpl selects DNS-01 when cf_api_token is set and HTTP-01 when it is empty. The real blocker is Cloudflare, proven live: the edge certificate covers exactly one wildcard level (SAN = demo-felhom.eu, *.demo-felhom.eu), so a two-label hostname — which a per-tester subdomain forces — gets "tls alert handshake failure" and no peer certificate at all. That makes Advanced Certificate Manager a prerequisite of the separate-domain plan, not an optional extra. Whether ACM is available on the account could not be established read-only: the only Cloudflare tokens in reach are the Zone:DNS:Edit tokens on the demo boxes, which the fence forbids using. Question C, measured rather than reasoned: r.Cookie returns the FIRST match and never tries the others (BOGUS+real = 302, real+BOGUS = 200), so a tossed cookie wins outright — DoS and confusion, not takeover, since it fails closed on mutations. CSRF is a single choke point (server.go:256) and the token carries the whole load against a same-registrable-domain attacker. But it is SKIPPED entirely when no session cookie is present, which with browser-cached Basic auth is cross-origin CSRF on every mutating route (R-135). Agreeing with the separate-domain recommendation, with the caveat the brief asked for: it is necessary but not sufficient. It does not solve Question D, because that is a shared-zone problem and the new domain is a shared zone. Filed R-133..R-138: duplicate domains accepted; hub/controller zone-resolvers disagree on depth; CSRF skipped on the no-cookie path; __Host- rename (one line, preconditions verified met); geo-WAF rules zone-scoped and non-namespaced (four cross-tenant faults, blocks shared-zone onboarding); shared-zone cf_api_token is a zone-wide DNS-write capability on a customer's box. Nothing created: no customer, DNS record, tunnel, route or code change.
30 KiB
RECON — can a tester get a Felhom-issued subdomain instead of owning a domain? (2026-07-31)
Read-only. No customer, DNS record, tunnel, route or code change was created. Two hub login sessions were created (in-memory, 7-day expiry) to test cookie precedence; nothing else was mutated.
Headline: the product accepts a Felhom-issued subdomain today with no code change. The
blocker is not Felhom — it is Cloudflare's edge certificate depth, proven live: a two-label
hostname under a zone gets tls alert handshake failure and no certificate at all. Everything
else follows from that.
1. Baselines — read fresh, none carried
| What | Value | How |
|---|---|---|
controller main |
v0.188.0 | HEAD=origin/main=4115e88, clean |
controller on demo-felhom (9201) |
0.187.0 | pct exec 9201 -- cat /etc/felhom-controller-image |
controller on demo-hp (9201) |
0.187.0 | same, on demo-hp |
| golden-baked controller | 0.188.0 (sha 7353d8be…) |
hub /configuration, selected option |
| agent published + vouched | 0.119.0 (sha e37aca82…) |
hub manifest; both boxes run 0.119.0 |
| hub | 0.84.0 | deploy/hub image tag |
| ISO | v1.25.0 (scripts/iso/build-felhom-iso.sh:44); host-install v1.22.0 |
source read |
felhom.eu HEAD |
b4edc08 |
git rev-parse |
| MinAgent | 0.113.0 | hub /configuration |
Both the hub (0.83.0→0.84.0) and the agent (0.118.1→0.119.0) moved since this morning's tester-gate run, which is exactly why the brief says to read rather than carry.
Not mine, left untouched: hub/internal/api/testdata/host-report.golden.json is dirty in the
shared worktree — another session's WIP. Not staged, not reverted.
2. Question A — does the product accept a Felhom-issued subdomain?
2.1 The domain's full path, hop by hop
| # | Hop | Citation | What happens |
|---|---|---|---|
| 1 | Operator form | hub/internal/web/configs.go:673 |
Domain: strings.TrimSpace(r.FormValue("domain")) — trimmed, nothing else |
| 2 | Hub DB | hub/internal/store/store.go:114 |
domain TEXT NOT NULL DEFAULT '' — no UNIQUE, no CHECK, no format |
| 3 | controller.yaml generation |
hub/internal/configgen/configgen.go:48 |
setNested(base, []string{"customer","domain"}, cfg.Domain) — verbatim copy |
| 4 | Box pull | felhom-host-install.sh config-retrieve → /opt/docker/felhom-controller/controller.yaml |
verified present on demo-felhom |
| 5 | Controller config load | controller/internal/config/config.go:363,379 |
env override FELHOM_CUSTOMER_DOMAIN; the only check is non-empty |
| 6 | Consumers | 30 call sites (§2.3) | all plain string interpolation |
2.2 Constraints found — the answer is "almost none"
- No format validation anywhere. Grepped the hub and controller for
valid|regexp|match|checkneardomain: no hits on the customer domain. - No registrable-domain / eTLD logic anywhere.
grep -riE 'registrable|etld|publicsuffix'across both repos: zero hits. Nothing computes "the registrable part" of a customer domain. - No "must not be a subdomain" rule. None exists.
- No uniqueness on domain.
hub/internal/web/configs.go:644rejects a duplicatecustomer_idonly. Two customers may be given the identical domain string with no complaint (→ R-133). - The only permissive-pattern check is the agent's
internal/lanresolver/lanresolver.go:225domainRe = ^[ \t]*domain:[ \t]*"?([A-Za-z0-9.-]+)"?— dots allowed, so a subdomain parses fine.
2.3 What consumes it on the box — all interpolation, no parsing
| Surface | Citation | Shape |
|---|---|---|
| Controller Traefik route | controller/internal/infra/infra.go:214 |
Host(`felhom.%s`) |
| Wildcard TLS SAN | same, :217-223 |
main: "*.<domain>", sans: ["<domain>"] |
| App Traefik rules | catalog compose | Host(`${SUBDOMAIN}.${DOMAIN}`) |
| App env | controller/internal/stacks/manager.go:1187, deploy.go:560 |
DOMAIN=<domain> |
| Controller-host guard | controller/internal/web/server.go:571 |
controllerHost := "felhom." + domain |
| Report / DR recipe | controller/internal/report/builder.go:53 |
https://felhom.<domain> |
| FileBrowser | controller/internal/stacks/infra.go:157 |
domain passed to compose render |
| Share launcher, escrow, export, claim, data-path cards | 9 further sites | template data only |
| SMB | controller/internal/infra/samba.go:65 |
NetBIOS name from smb.EffectiveServerName() — not the domain |
None of these parse the domain. A dotted subdomain flows through every one of them unchanged.
2.4 Does anything assume the customer owns the zone? — Yes, but it is a switch, not a requirement
This is the question the brief flagged, and the answer is better than feared. The ACME challenge is
config-selected at controller/internal/infra/templates/traefik.yml.tmpl:
certificatesResolvers:
letsencrypt:
acme:
email: {{.ACMEEmail}}
{{- if .CFAPIToken}}
dnsChallenge:
provider: cloudflare
{{- else}}
httpChallenge:
entryPoint: web
{{- end}}
- With
cf_api_token→ DNS-01, which does needZone:DNS:Editfor that zone, and the token is written to a 0600.envon the box (controller/internal/infra/infra.go:123). - Without it → HTTP-01, which needs no DNS credential at all.
The wildcard SAN is gated on the same flag: controller/internal/stacks/infra.go:211 —
wildcardTLS := CFAPIToken != "" && Customer.Email != "".
Live on demo-felhom: DNS-01 is what actually runs — traefik.yml shows
dnsChallenge: provider: cloudflare, a 0600 .env exists, cloudflared is up, and the dynamic route
carries main: "*.demo-felhom.eu". So both demo boxes hold a zone-write Cloudflare token because
they own their zones, not because the code demands one.
2.5 Does anything create DNS records automatically? — No
grep -rn 'dns_records' across both repos: zero hits. Neither the hub nor the controller ever
creates, updates or deletes a DNS record. The two Cloudflare clients that exist are WAF-only:
hub/internal/cloudflare/unblock.go— deletes[felhom-geo]WAF rules.controller/internal/cloudflare/{waf,geosync,zone}.go— reconciles[felhom-geo]WAF rules.
The cloudflared template says it outright: "Routes are configured in the Cloudflare dashboard (Zero Trust > Networks > Tunnels > Public Hostname)". All DNS and tunnel routing is manual operator work today.
2.6 Verdict on Question A
YES — the product accepts a Felhom-issued subdomain today, with no code change.
customer.domain = "t1.<felhom-owned-zone>" flows through the hub, controller.yaml, Traefik, the
tunnel, app env and every URL builder without touching a single constraint. Nothing needs DNS API
credentials if cf_api_token is left empty (HTTP-01), and nothing creates DNS records either way.
The blocker is elsewhere — see Question B.4. Two code-level caveats that are not blockers:
- Zone-resolution depth asymmetry (→ R-134). The controller's
GetZoneID(controller/internal/cloudflare/zone.go:18) strips labels progressively (any depth). The hub'sresolveZone(hub/internal/cloudflare/unblock.go:115) tries the exact name and thenparentDomain— which strips exactly one label (:136). Fort1.felhomcust.huboth succeed; for a two-label customer domain the hub silently fails to find the zone while the controller succeeds. Two implementations of one concept, disagreeing. - Geo-restriction must stay OFF for shared-zone customers — Question D, and it is severe.
3. Question B — Cloudflare routing
3.1 How the demo boxes reach the internet
One tunnel per customer, created by hand. cf_tunnel_token is a per-customer field the operator
pastes into the hub (hub/internal/web/configs.go:1338), stored in config_json.infrastructure,
delivered in controller.yaml, and passed straight to cloudflared as TUNNEL_TOKEN
(controller/internal/infra/templates/cloudflared-compose.yml.tmpl). If it is empty, cloudflared is
simply not deployed — controller/internal/stacks/infra.go:67 logs "LAN-only node". The token is
the entire config; there is no config.yml, and public hostnames live only in the dashboard.
3.2 The live DNS picture
felhom.eu A 37.191.56.193 (NOT proxied)
hub.felhom.eu CNAME dooplex.hopto.org → (NOT proxied — no cf-ray, no server: cloudflare)
files.felhom.eu CNAME dooplex.hopto.org
ep0.felhom.eu A 167.233.158.164
demo-felhom.eu (apex: no A)
felhom.demo-felhom.eu 188.114.97.0 / 188.114.96.0 (Cloudflare proxy)
poll.demo-felhom.eu 188.114.97.0 / 188.114.96.0 (Cloudflare proxy)
felhom.enkisfelhom.hu 188.114.96.0 / 188.114.97.0 (Cloudflare proxy)
All four zones (felhom.eu, demo-felhom.eu, enkisfelhom.hu, sajatfelhom.hu) sit on the same
Cloudflare nameservers (martha/tony.ns.cloudflare.com) — one account.
Two facts that matter, both established rather than assumed:
felhom.euis NOT behind Cloudflare.hub.felhom.euandfelhom.eureturn nocf-rayand noserver: cloudflare; they resolve through a DDNS name to a home IP and are served by DooPlex's k3s nginx ingress. The hub is not protected by, or reachable through, the tunnel fabric.- There is no
*.felhom.euwildcard today. A random label underfelhom.euresolves to nothing; the same probe underdemo-felhom.euandenkisfelhom.hureturns Cloudflare proxy IPs. So the cookie exposure in Question C is latent, not live — it becomes live the moment a tester box is put underfelhom.eu.
3.3 Proxied wildcards — verified, not remembered
The brief warns against trusting a remembered plan restriction. Checked empirically: a random
label under a customer zone resolves to Cloudflare proxy anycast IPs and returns a real HTTP
response with server: cloudflare and a cf-ray. So a proxied wildcard DNS record works on this
account today — the historical "wildcard proxying is Enterprise-only" restriction does not bind here.
Two wildcards in the same account already route to two different tunnels (each box runs its own
cloudflared with its own token). Distinct wildcard records at distinct names
(*.t1.zone → tunnel A, *.t2.zone → tunnel B) are ordinary DNS with no conflict.
3.4 The actual blocker — edge certificate depth
Proven live, and this is the finding that decides the shape of the whole plan.
The Cloudflare edge certificate for a zone covers exactly one wildcard level:
$ openssl s_client -connect poll.demo-felhom.eu:443 -servername poll.demo-felhom.eu
subject=CN=demo-felhom.eu
issuer=C=US, O=Google Trust Services, CN=WE1
X509v3 Subject Alternative Name:
DNS:demo-felhom.eu, DNS:*.demo-felhom.eu
A two-label hostname resolves — the wildcard DNS record matches at any depth — but the TLS handshake is refused outright:
$ openssl s_client -connect zz.poll.demo-felhom.eu:443 -servername zz.poll.demo-felhom.eu
CONNECTED(00000003)
ssl3_read_bytes:tls alert handshake failure ... SSL alert number 40
no peer certificate available
$ curl https://zz.poll.demo-felhom.eu/
curl: (35) TLS connect error: ... tls alert handshake failure
$ curl https://poll.demo-felhom.eu/ # one-level control
HTTP 404 # TLS fine; just no tunnel route for a random name
Consequence. A Felhom-issued subdomain means customer.domain = t1.felhomcust.hu, so the
controller lives at felhom.t1.felhomcust.hu and apps at poll.t1.felhomcust.hu — two labels under
the zone. Under Universal SSL those hostnames get no certificate, and the tester's browser gets a
hard TLS error before anything Felhom-related runs.
Options, in order of how much I can vouch for them:
| Option | Cost | Status |
|---|---|---|
Cloudflare Advanced Certificate Manager / Total TLS — issue *.t1.felhomcust.hu at the edge |
paid add-on (~$10/mo, one zone) | Could not verify read-only — needs the dashboard or a token I am fenced from using. This is the single thing to check before committing to the plan. |
| One zone per tester | one domain per tester, ~2–4k HUF/yr each | Works today, no code change — but it is the cost the plan exists to avoid |
Flat naming (poll-t1.felhomcust.hu) |
— | Does not work without code change. The controller hard-derives felhom.<domain> (infra.go:214, web/server.go:571); every tester's controller would want the identical felhom.felhomcust.hu |
| Grey-cloud + DNS-01 wildcard on the box | — | Traefik can issue *.t1.felhomcust.hu itself via DNS-01, and the browser would see that cert — but grey-cloud means no tunnel, so the box needs a public IP and port-forwarding. Loses the reason the tunnel exists |
3.5 What must exist for one tester box — the create-list
Assuming a Felhom-owned zone felhomcust.hu already on Cloudflare, and ACM/Total TLS confirmed:
| # | Where | Action |
|---|---|---|
| 1 | Cloudflare → zone felhomcust.hu → SSL/TLS → Edge Certificates |
Enable Total TLS / ACM so depth-2 hostnames get a cert. Verify first — §3.4. |
| 2 | Cloudflare → Zero Trust → Networks → Tunnels | Create tunnel felhom-t1. Copy its token (this is the cf_tunnel_token) |
| 3 | Same tunnel → Public Hostname | felhom.t1.felhomcust.hu → https://traefik:443 (or per the existing demo tunnels' pattern — I could not read those routes, §8) |
| 4 | Same tunnel → Public Hostname, one per app | e.g. poll.t1.felhomcust.hu → same origin. Cloudflare creates the proxied CNAME to <tunnel>.cfargotunnel.com automatically when you add a public hostname |
| 5 | Cloudflare → DNS (only if you prefer a wildcard over per-app entries) | *.t1.felhomcust.hu CNAME → <tunnel-id>.cfargotunnel.com, proxied. Verified workable in §3.3 |
| 6 | Hub → Customers → + Add Customer | customer_id = t1; domain = t1.felhomcust.hu; email = the tester's real mailbox (see §6); DR tier / offsite per policy |
| 7 | Hub → the customer's Edit tab | Paste cf_tunnel_token from step 2. Leave cf_api_token EMPTY — §2.4 (HTTP-01), §5 (a shared-zone token is a cross-tenant write capability) |
| 8 | — | Do not enable geo-restriction for this customer. §5. |
No DNS record and no tunnel is created by Felhom; steps 1–5 are entirely operator work in the Cloudflare dashboard.
4. Question C — cookies, CSRF, and the domain recommendation
4.1 The cookie, reconfirmed at source and on the wire
hub/internal/web/server.go:656-663 sets hub_session with Path: "/", HttpOnly: true,
SameSite: SameSiteLaxMode, Secure: isSecure, no Domain, no __Host- prefix — as the
brief states. Confirmed on the live production path:
set-cookie: hub_session=<REDACTED>; Path=/; Max-Age=604800; HttpOnly; Secure; SameSite=Lax
isSecure is computed from r.TLS != nil || X-Forwarded-Proto == "https". Over the k3s ingress that
holds; over the plain-HTTP ClusterIP the same login returns the cookie without Secure.
4.2 What the hub does with two same-named cookies — tested, not reasoned
Against the live hub, GET /configuration, varying only the Cookie: header:
| Cookie header | Result |
|---|---|
hub_session=<real> |
HTTP 200 — authenticated |
hub_session=BOGUS; hub_session=<real> |
HTTP 302 → /login — rejected |
hub_session=<real>; hub_session=BOGUS |
HTTP 200 — authenticated |
hub_session=BOGUS |
HTTP 302 → /login |
r.Cookie("hub_session") returns the FIRST match in header order and never tries the others. A
tossed cookie that sorts first therefore wins outright, and the real session is silently ignored.
Exploitability, stated precisely. Per RFC 6265 §5.4 a browser orders cookies longer Path
first. Any page on any *.felhom.eu host can set
hub_session=X; Domain=felhom.eu; Path=/configuration — a longer path than the real Path=/ cookie —
so the browser sends the attacker's value first for every /configuration* request.
What that does and does not buy an attacker:
- It is not session takeover. The attacker cannot mint a value present in the hub's in-memory
sessionsmap, so the tossed cookie never authenticates. - It is a persistent, hard-to-diagnose denial of service against the operator UI, on paths of the attacker's choosing, invisible in the UI and not clearable from the hub.
- It fails closed on mutations.
csrfToken(r)(server.go:698) reads the same first cookie and returns"", andvalidateCSRF(:678) returns false for an unknown session → 403.
4.3 What the CSRF defence actually rests on
A single choke point, and it is the token. hub/internal/web/server.go:256-269 gates every
non-GET/HEAD/OPTIONS web route before the route switch, exempting only /login, isPublicBindPath,
and the no-password-configured case. /api/v1/ is Bearer-authenticated and separately exempt. So the
token is enforced on every mutating web route — the brief's worry that it "would be carrying the whole
load" is correct, and it does carry it.
Measured, using POST /zz-recon-nonexistent so the CSRF gate is isolated with zero mutation
(the middleware at :256 runs before the route switch at :271; a 404 therefore means "past the
gate"):
| Request | Result |
|---|---|
session cookie, no _csrf |
403 — CSRF enforced |
session cookie, valid _csrf |
404 — past the gate |
Basic auth, no cookie, no _csrf |
404 — CSRF SKIPPED |
Basic auth + bogus cookie, no _csrf |
403 — CSRF enforced |
The third row is validateCSRF's documented early return: "No session cookie — likely Basic Auth or
programmatic access; skip CSRF" (server.go:679-683).
Why that is a real hole and not just an API convenience (→ R-135). Browsers cache HTTP Basic
credentials per origin and resend them automatically, including on cross-origin requests, and
SameSite does not govern the Authorization header. So if the operator has ever authenticated to
hub.felhom.eu with Basic auth in their browser, an attacker page anywhere — not merely same-site —
can POST to any mutating hub route: the credentials ride along, auth passes, and CSRF is skipped
because there is no cookie. The condition ("operator used Basic auth in a browser") is what makes this
latent rather than live, and it is a condition, not a guarantee.
Note on SameSite=Lax: against a genuinely cross-site attacker it already blocks the cookie, so the
cookie path is safe there. Against a same-registrable-domain attacker (tester.felhom.eu →
hub.felhom.eu) Lax provides no boundary at all, and the CSRF token is the only thing left.
4.4 Should customer boxes live under felhom.eu? — No. And that is necessary, not sufficient.
I agree with the recommendation, for the stated reason, and it is cheap. A different registrable
domain makes tossing structurally impossible, restores SameSite as a real boundary, and permanently
answers "the tester owns no domain" for every future customer at one domain's annual cost.
But my findings say it does not finish the job, and the brief asked to be told if so:
- Question D is not solved by moving zones — it is a shared-zone problem, and the new domain is a shared zone. The geo-WAF collisions in §5 apply identically there.
- The nesting requirement collides with edge-certificate depth (§3.4). The separate domain forces two-label hostnames, which is precisely what Universal SSL will not serve. The separate domain makes ACM a prerequisite, not an optional extra. That is a real, ongoing cost the recommendation should carry openly.
So: adopt the separate registrable domain, and treat Cloudflare ACM/Total TLS on that zone as part of its price. If ACM turns out to be unavailable or unexpectedly expensive, the honest fallback is one cheap zone per tester — which keeps every property except the "one domain forever" economics.
Independently, and regardless: rename the cookie to __Host-hub_session. Verified on the live
production response that all three prefix preconditions already hold (Secure, Path=/, no Domain),
so it genuinely is a one-line change — with one caveat worth writing into the ticket: a browser
rejects a __Host- cookie that lacks Secure, and isSecure is conditional, so any plain-HTTP
browser access to the hub would stop working. Non-browser access uses Basic auth and is unaffected.
Filed as R-136; not implemented here.
5. Question D — what breaks with many customers in one zone
What I checked, including what came back clean:
| Surface | Checked | Result |
|---|---|---|
Hub uniqueness on domain |
store.go:114, configs.go:644 |
No constraint — duplicate domains accepted silently (R-133) |
| App subdomain collisions | catalog Host(${SUBDOMAIN}.${DOMAIN}) |
Clean if each tester has their own label — poll.t1.z vs poll.t2.z |
| Controller hostname | infra.go:214, web/server.go:571 |
Clean with nesting, fatal with flat naming — every box wants felhom.<domain> |
| SMB / NetBIOS names | controller/internal/infra/samba.go:63-65 |
Clean — workgroup = WORKGROUP fixed, NetBIOS from settings, not the domain. LAN-scoped anyway |
storage.cfg / PVE storage names |
agent internal/storage/ |
Clean — no domain-derived names found |
| Host identity | hub log sess-g-e687c3 |
Clean — <customer_id>-<random>; the agent never reads customer.domain except in lanresolver |
| Tunnel naming | dashboard-only, token is the whole config | Clean — tunnels are independent objects; no Felhom-side naming |
| LE certificate rate limits | — | Not a factor while the edge terminates TLS; becomes one only in the grey-cloud/DNS-01 topology |
| LAN split-horizon DNS | agent internal/lanresolver/lanresolver.go:6-13 |
Scoped correctly with nesting (local=/t1.felhomcust.hu/); would hijack the entire zone on that LAN if a customer's domain were the bare zone — another reason flat naming is wrong |
| Cloudflare geo-WAF rules | controller/internal/cloudflare/{waf,geosync}.go |
BROKEN — four distinct cross-tenant faults (R-137) |
5.1 The geo-WAF collision — the one real breakage
The rules are written at zone scope with non-namespaced identities:
waf.go:18—globalRuleDesc = "[felhom-geo] Global"— one literal description for the whole zone.waf.go:21—appRuleDescPrefix = "[felhom-geo] app:"+ the app name, with no customer in the key.waf.go:241—BuildGlobalExpressionreturns(not ip.src.country in {...})plushttp.host neexclusions only. There is no positive hostname scoping: the rule matches every hostname in the zone.geosync.go:320—applyDiffdeletes every existing[felhom-geo]rule not in this box's desired set.
With two customers in one zone and geo-restriction enabled by either:
- They overwrite each other forever. Both want a rule described
[felhom-geo] Globalwith different expressions → each sync sees "expression changed" →UpdateRule→ flip-flop on every tick. - One customer's country policy applies to everyone in the zone, because the expression is not hostname-scoped.
- Per-app rules collide by name — two testers running
ralllyboth produce[felhom-geo] app:rallly. - Disabling the feature for one customer wipes them all —
deleteAllRules(geosync.go:167), and the hub's operator "unblock" (hub/internal/cloudflare/unblock.go:22RemoveGeoRules) deletes all[felhom-geo]rules in the resolved zone.
Mitigation for tester #1 (no code): leave geo-restriction off for every shared-zone customer.
Fix (code, R-137): namespace the descriptions by customer_id and add a positive
http.host ends_with "<customer domain>" term to both expressions. That is ~4 functions across
waf.go/geosync.go plus the hub's RemoveGeoRules — a two-repo change, i.e. exactly the class
that dr-recipe-three-producers-hub-allowlist warns about. Small-to-medium, not trivial.
5.2 The shared-zone credential problem
If a shared-zone customer is ever given a cf_api_token (for DNS-01 or geo), that token is
Zone:DNS:Edit for the whole shared zone, written to a 0600 .env on a box in a stranger's home
(controller/internal/infra/infra.go:123). One compromised tester box could repoint every other
tester's DNS. Today each box holds a token only for a zone the operator owns and no one shares, so
this is new exposure created by the shared-zone plan. Keep cf_api_token empty for shared-zone
customers — which §2.4 shows is supported (HTTP-01). Recorded as part of R-138.
6. The operator's step list for tester #1
Ordered, with the prerequisites that bite if taken out of order.
- Choose and register the customer domain. One registrable domain for all customers, e.g.
felhomcust.hu. Not a subdomain offelhom.eu— §4.4. - Add the zone to Cloudflare (same account) and point the registrar's nameservers at it.
- Verify ACM / Total TLS on that zone and enable it. §3.4 — do this before promising a tester a date; without it every hostname the tester touches fails TLS outright. If ACM is not viable, fall back to one zone per tester and stop here.
- Create the tunnel
felhom-t1in Zero Trust → Networks → Tunnels. Keep the token; it iscf_tunnel_token. - Add the public hostnames on that tunnel:
felhom.t1.felhomcust.huplus one per app the tester will run — or a single proxied wildcard*.t1.felhomcust.hu→<tunnel-id>.cfargotunnel.com. - Create the customer in the hub —
customer_idt1,domaint1.felhomcust.hu, andemail= the tester's own working mailbox.R-119 prerequisite — this is the step that strands a tester. The dashboard claim code is generated hub-side, bcrypt-hashed, and emailed only; there is no operator-side route to the plaintext. It is sent when the customer record is created and again on
POST /configs/<id>/claim-resend. The tester's mailbox must be on the record before their box boots, or the box installs perfectly and nobody can claim it. Confirmed live in this morning's tester-gate run, where the operator had to relay the code by hand. - Paste
cf_tunnel_tokeninto the customer's Edit tab. Leavecf_api_tokenempty (§2.4, §5.2). - Do not enable geo-restriction for this customer (§5.1).
- Boot the box from the v1.25.0 generic ISO. It self-registers as an unclaimed appliance; bind it
to
t1on the hub's Hosts page; the credentials are delivered once and day-0 completes unattended. - The tester claims the dashboard with the emailed code and sets their own password.
- Sanity-check TLS first, before the tester does: open
https://felhom.t1.felhomcust.hu. If it fails with a handshake error rather than a Felhom page, step 3 did not take.
7. R-n rows opened
Registered in documentation/backlog/OPEN-ITEMS.md; none implemented.
| ID | What | Size |
|---|---|---|
| R-133 | Hub enforces uniqueness on customer_id only — two customers can be given the same domain with no complaint |
XS |
| R-134 | Zone-resolution depth asymmetry: controller strips labels progressively, hub strips exactly one (unblock.go:136) |
XS |
| R-135 | validateCSRF returns true when no session cookie is present — with browser-cached Basic auth this is cross-origin CSRF on every mutating hub route |
S, security |
| R-136 | Rename hub_session → __Host-hub_session (all preconditions verified met on the production path) |
XS, one line |
| R-137 | Cloudflare geo-WAF rules are zone-scoped and non-namespaced — four cross-tenant faults; blocks geo-restriction for any shared zone | M, two-repo |
| R-138 | A shared-zone cf_api_token is a zone-wide DNS-write capability sitting on a customer's box — policy + a guard |
S |
8. What could not be established read-only
- Whether Cloudflare ACM / Total TLS is available or enabled on this account. This is the single
decision-critical unknown (§3.4). It needs the Cloudflare dashboard or an API token. The only
Cloudflare tokens in reach are the
Zone:DNS:Edittokens in/opt/docker/stacks/traefik/.envon the demo boxes — those can WRITE, so per the fence I did not use them. The hub's own namespace holds no Cloudflare credential at all (checked: 14 secrets, none Cloudflare-related). Operator action: read it off the dashboard. - The existing tunnels' public-hostname route tables — same reason. So step 5 of §3.5 names the shape but not the exact origin URL the demo tunnels use; copy it from an existing tunnel.
- Whether Cloudflare passes HTTP-01
/.well-known/acme-challenge/through a proxied hostname to the tunnel origin. Relevant only ifcf_api_tokenis left empty, which §3.5 step 7 recommends. Would need a live issuance attempt on a real hostname — i.e. tester #1's first boot. If it fails, the fallback is to setcf_api_tokenand accept §5.2's exposure, or rely on the edge cert alone. - Browser Basic-auth resend behaviour underlying R-135 is stated from the specification, not
measured here; a browser is needed and
claude-in-chromeis unavailable on DooPlex.
9. What was not done
No customer, DNS record, tunnel, route or code change was created. No demo box, drill-r50, guest 9201
or Peti's cluster was touched — the only writes anywhere were two hub login sessions (in-memory, 7-day
expiry, tokens shredded) and this document plus the six register rows. curl -w '%{redirect_url}' was
not used (R-132).