Empirical spike (demo arc, docs-only, no product code). Measures the two hub->box 'sync now' transports. Verdict: (b) long-poll = primary for user config (hub-only, reaches controller, no ep0/WG/agent coupling; binding ceiling = twin 60s nginx proxy_read_timeout + hub WriteTimeout, both config lifts); (a) ep0-relayed poke = agent-plane nudge (~0.42s, zero ep0/box infra change, EKEYREJECTED-confined). Viktor Q2 answered: hub.felhom.eu is DNS-only, no Cloudflare. All live mutations reverted; arc re-verified green.
18 KiB
SPIKE — immediate-sync transport (hub→box "sync now" signal) — 2026-07-16
STATUS: COMPLETE — verdict below. Empirically measures the two candidate transports for the hub→box immediacy signal BEFORE any production spec. No production code shipped; no agent/controller/hub image; every live mutation inventoried and reverted (§CLEANUP, verified). The question under test (Viktor's ruling): a user-triggered config change must round-trip in seconds, not the 15-min report cycle — so the hub needs a "sync now" wake-up and the box then pulls config via the NORMAL channel (report ACK / desired-state fetch; the hub never delivers config by connecting inbound).
Class: SPIKE (empirical; no product code). Repos: felhom.eu (this doc only). All others
read-only for grounding: hub/internal/wgsync/{client,reconciler}.go, hub/cmd/hub/main.go,
manifests/hub.yaml, felhom-agent/internal/desired/syncer.go.
Baselines re-confirmed at session start (git HEAD): felhom.eu a50822af v0.57.0 · felhom-agent
c040c180 v0.88.0 · felhom-controller 8f3564c1 v0.138.0. Live hub image :0.57.0.
Probe ends (demo arc ONLY; Peti's arc untouched): box = felhom-pve/demo-felhom
(192.168.0.162, LAN SSH = safety line) · endpoint = felhom-hetzner/ep0 (167.233.158.164,
root SSH, wg0 10.77.0.1) · hub pod on DooPlex k3s (192.168.0.180, kubectl -n felhom-system).
Verdict (one line)
Both transports are GO and they are COMPLEMENTARY, not competing. (b) long-poll delivers to
the controller (the config puller) with a hub-only change and zero ep0 coupling, and is
the lighter fit for Viktor's user-config round-trip — but only after two 60 s ceilings are raised
(nginx proxy_read_timeout measured at 60 s, and the hub's in-code WriteTimeout 60 s);
today a naive long-poll is capped at ~55 s → ~1 reconnect/box/min. (a) ep0-relayed poke delivers
to the agent (WG terminates host-side) in ~0.42 s per fresh SSH session (≈27 ms+18 ms if the
session is pooled), needs zero infra change on ep0 or the box (origination from wg0's own IP
never touches the forward chain — proven), and is cryptographically confined to registered peers —
but it needs an agent train (a new UDP listener→heartbeat trigger) and a hub forced-command key
on ep0. Recommend (b) as the primary immediacy path for customer/app config (hub+controller
only), with (a) reserved as the agent-plane desired-state nudge already aligned with the OOB /
mutual-repair arc. Interval shortening stays REJECTED.
P0 — ground truth (no changes) — GATE PASSED
All three ends reachable → proceed. Redacted reads only; wg show <if> dump never run (S1 key-leak
ban); private keys never printed.
DNS / fronting — Viktor's Q2 answered: DNS-only, NO Cloudflare
- Public:
hub.felhom.eu→ CNAMEdooplex.hopto.org→ 37.191.56.193 (the operator's One-Hungary fixed-cable line — the SAME public IPv4 the demo box sits behind). Nocf-ray, noserver: cloudflare. - On-LAN split-horizon (resolver
192.168.0.250):hub.felhom.eu→ 192.168.0.192 (thenginx-internalingress LB,nginx-system/ingress-nginx-internal-controller). curl -sIon-LAN/healthz→HTTP/1.1 200, HSTS, no CDN headers. From ep0 via the public IP (--resolve hub:443:37.191.56.193) →HTTP/2 403= the HU geo-gate (ep0 is in DE; the gate is a location-levelconfiguration-snippeton the hub/path). The path reaches nginx over HTTP/2 and returns immediately.- Real ingress chain:
[client] → internet → 37.191.56.193 (operator home-router NAT) → k3s nginx-internal → hub pod:8080. Self-hosted, no CDN in front. (Controller→hub and agent→hub both ride this same DNS-only nginx path — the controller's Cloudflare tunnel is its inbound edge only.)
Box (felhom-pve / demo-felhom)
wg-felhom: Address 10.77.0.2/32, MTU 1280; peer f3d1ZI7…(ep0) Endpoint 167.233.158.164:443,
AllowedIPs 10.77.0.1/32, 10.77.0.250/32, PersistentKeepalive 25. registered.json: assigned
10.77.0.2/32, generation 11, pubkey aqMpucg…. Handshake fresh (~1 min); wg-quick@wg-felhom
felhom-agentboth active; agent 0.88.0. Box nft empty (no host belt; pve-firewall disabled) — matches the 07-05 baseline.
- Note vs 07-05: the box's WG identity rotated (new pubkey) on the R-21 ISO reinstall but the
assigned IP was preserved (
10.77.0.2). The boxAllowedIPsnow already carries10.77.0.250/32— i.e. the 07-05 spike's finding-1 gap ("render the operator /32 into the agent conf") has since been closed; the OOB operator peer is durably rendered.
Endpoint (felhom-hetzner, wg0 :443)
Three peers: box aqMpucg… → 10.77.0.2/32 (fresh); a second real box zxy… → 10.77.0.3/32
(handshake ~3h49m — not touched); operator uDPlMa0… → 10.77.0.250/32 (no endpoint — dormant
OOB peer). ip_forward = 1 (ipv6 fwd 0). nft inet filter: input policy DROP with allows
(established, lo, tcp/22, udp/443, tcp/8007 iif wg0, icmp); forward policy DROP and empty (no
accept rules). authorized_keys: root 2 lines, felhom-peersync 2 lines.
- Note vs 07-05:
ip_forwardwas 0 at the 07-05 baseline and is 1 now — the OOB operator-peer feature has shipped in the interim. The live baseline (ip_forward=1, forward chain empty) is therefore the cleanup target — NOT the stale "must be 0" assumption in the plan.
Hub pod
hub-7b97cfb48d-qrbpb Running, image :0.57.0. exec works (Alpine, root). Reconciler ALIVE:
[INFO] wgsync: pushed 3 peers to 167.233.158.164:22 every 5 min — so hub-pod → ep0:22 already
works in production (the forced-command peersync channel a poke-relay would clone).
P1 — ep0→box UDP origination (the poke's last hop) — GO, zero endpoint change
Listener on the box bound to its WG IP (10.77.0.2:19999) + tcpdump -ni wg-felhom; ep0 originates
from wg0.
- Delivery, packet-proven: ep0
python3 sendto("10.77.0.2",19999)from10.77.0.1. Box:RECV b'sync-poke-p1' from ('10.77.0.1', 47280). tcpdump captured 2 pkts10.77.0.1.* > 10.77.0.2.19999(python + ancvariant). - The load-bearing claim — ZERO mutation needed:
ip_forward1 → 1 unchanged;forwardchain still empty policy-drop; no nft/peer/route change. Endpoint-originated traffic leaves via wg0's output path and never traverses theforwardchain — so the poke needs noip_forward, no forward rule, no nft edit. (This is exactly why forwarding being on or off is irrelevant to the relay: the poke is origination, not transit.) - Latency over the tunnel: ping avg 29.0 ms (min 22.6 / max 36.9 / mdev 4.8); HTTP
round-trip connect ~23–31 ms, total ~55–66 ms (5/5
http=200). - Negative — cryptographically confined to registered peers: poking an unregistered
in-subnet IP (
10.77.0.199; the route exists,dev wg0 src 10.77.0.1) fails at the kernel:OSError: [Errno 126] Required key not available(EKEYREJECTED) — WireGuard refuses to encrypt to an IP no peer owns. wg0 peer count unchanged (3). The future forced command can therefore only ever reach a registered peer's/32; there is no "spray the subnet" failure mode.
P2 — hub-pod → ep0 → box poke round-trip — GO, ~0.42 s per poke (target was ≤2–3 s)
Throwaway keypair generated on the DooPlex host (/tmp/spike-poke-key, 0600, private key never
printed); public key installed on ep0 root authorized_keys as
command="/root/spike-poke.sh",restrict … (root ak 2 → 3, 1 spike line). The forced command is
one UDP sendto to the box.
- Pod capability: the hub image is Alpine — it has
ncbut nossh, nobash, nopython3. So a production poke MUST reuse the in-processgolang.org/x/crypto/sshclient (thewgsyncpattern), not a shelled-out binary. Pod→ep0:22 reachability from the pod netns confirmed:nc -z -w4 167.233.158.164 22= OK (and the reconciler proves it continuously). - End-to-end, 5 timed pokes (DooPlex host
--ssh--> ep0 forced-cmd --> UDP --> box): all 5 delivered (boxRECV #1–5 b'hub-poke' from 10.77.0.1), remote returnedpoke-fired.ssh_wall: 0.636 s (cold), 0.423, 0.439, 0.414, 0.412 → warm median ~0.42 s. - Cost decomposition: forced-command body (python startup +
sendto) = 18–20 ms (bare python3 = 11–12 ms); DooPlex→ep0 net RTT = 27 ms. So ~0.40 s of the 0.42 s is SSH TCP+KEX+auth+shell-spawn — the session setup dominates, the poke itself is negligible. Task-2 lever: a pooled/persistent SSH session (or the in-process client holding one channel) collapses per-poke to ~1 RTT (27 ms) + body; even a fresh session per poke is already sub-second. - Failure behavior: with no box listener, the poke still fires on ep0 and the datagram is
silently discarded at the box — no error path back to the hub (a lost poke is harmless by
construction; the next scheduled report reconciles). Stale handshake: NOT force-tested —
forcing it means stopping keepalive/the tunnel, which risks the live offsite/PBS path (STOP
rule). From WG semantics + observation:
PersistentKeepalive 25kept the handshake < 125 s fresh throughout; a poke arriving on a cold handshake triggers WG's auto-rehandshake on first tx (~1 extra RTT), datagram buffered then delivered — not dropped.
P3 — long-poll hold-time through the REAL ingress chain
Method (why no DNS record was created): a throwaway sleeper (a python:3.14-bookworm pod
that sleeps ?hold=N before a 200) + Service + Ingress on host probe.felhom.eu, class
nginx-internal, in felhom-system. Reached with curl --resolve probe.felhom.eu:443:<ip> -k
from each vantage — so no DNS record was needed (the plan's fallback "ask Viktor for an A
record" was avoided), the live hub Ingress was never touched, and the sleeper's own server block
carries no geo snippet (so the remote ep0 vantage is not geo-blocked). Default cert → curl -k.
The nginx-internal ConfigMap sets no custom proxy timeouts → nginx built-in defaults.
Pass 1 — DEFAULT ingress, on-LAN box (direct nginx, no NAT)
| hold N | result |
|---|---|
| 30 s | 200 @ 30.0 s |
| 55 s | 200 @ 55.0 s |
| 70 s | KILLED |
| 100 s | KILLED |
nginx error log (authoritative): upstream timed out (110: Operation timed out) while reading response header from upstream … GET /?hold=70 HTTP/2.0. Killing layer = nginx
proxy_read_timeout = 60 s (built-in default). Each request appears twice ~60 s apart =
ingress-nginx's default single upstream retry on timeout, so the client-visible 504 lands at
~120 s (2 × 60 s); the client's own cut at 110 s showed http=000 just before it. The sleeper's
BrokenPipeError confirms nginx abandoned the read at 60 s.
- H(default) = ~55 s usable / hard kill at 60 s → naive per-box reconnect rate ≈ 1/min.
- In-code twin ceiling: the hub's own
http.Server{WriteTimeout: 60s}(hub/cmd/hub/main.go:339–341) caps a held response at ~60 s regardless of ingress — coincident with nginx's 60 s. Both must be raised for any long-poll (nginx via annotation; the hub per-endpoint viahttp.ResponseController).
Pass 2 — TUNED sleeper ingress (proxy-read-timeout/proxy-send-timeout=3600, proxy-next-upstream=off)
- on-LAN
hold=120→200 @ 120.0 s(was killed under default → nginx WAS the layer, and raising it lifts the ceiling cleanly; on-LAN there is no further layer below the pure-nginx value). - on-LAN
hold=600→200 @ 600.0 s(pure tuned-nginx holds a full 10 min cleanly — no further on-LAN layer below 600 s). - ep0 remote (through the operator home-router NAT
37.191.56.193):hold=2→200 @ 2.1 s(routes, not geo-blocked);hold=200→200 @ 200.09 s;hold=600→200 @ 600.09 s— the operator home NAT holds an idle held connection for the full 10 min without eviction. So on this arc the NAT is not the binding constraint (≤ 600 s); the only ceiling today is the twin 60 s nginx+hub limits, both config-fixable. Once raised, a long-poll can hold many minutes even through the operator NAT → reconnect rate drops from ~1/min toward ~1/10-min.
Vantage caveat (honest ledger)
The demo box is on-LAN, co-located behind the SAME operator NAT as the hub, so it does not
traverse a real customer's home-NAT → internet → operator-NAT double-NAT path. The ep0 vantage
represents the ingress-side (operator) NAT only; a customer-side home-router idle timeout is
unrepresented by this arc. The load-bearing measurement — that nginx's 60 s default is the
binding ceiling today, and it is a config lift — holds regardless of vantage.
P4 — verdict: (a) poke-relay vs (b) long-poll
| Dimension | (a) ep0-relayed poke | (b) long-poll |
|---|---|---|
| Measured latency | ~0.42 s per fresh SSH session (≈27 ms + 18 ms if pooled); last hop ep0→box ~29 ms | delivers on the next held-GET completion; today capped at ~55 s by the nginx+hub 60 s ceilings (both fixable) |
| Reaches | the agent (WG terminates host-side) → desired-state nudge | the controller (the config puller) directly — also usable on the agent channel |
| Hub change | new forced-command poke key + in-process SSH poke call | new hanging-GET endpoint + raise nginx & hub WriteTimeout |
| ep0 change | none (P1: origination needs no forward/nft/sysctl) | none |
| Box/agent change | agent train — new UDP listener → "heartbeat now" trigger | none if the controller holds the GET (it already pulls config); or the agent |
| ep0 coupling | yes (rides the peersync SSH pattern) | no (direct hub↔box outbound over the existing ingress) |
| WG down | poke fails (WG is the transport) — but agent self-heals WG in ~15 s (07-05) and a lost poke is harmless; next report catches up | unaffected (uses the ingress path, not WG) |
| Hub restart | forced-command is stateless; resumes on next trigger | held GETs drop → box must reconnect (reconnect logic needed) |
| NAT rebind | WG keepalive re-handshakes; poke buffered | held GET dies silently on rebind → client hold must be < NAT idle timeout |
Recommended Task-2 shape — complementary, not either/or:
- (b) long-poll is the primary immediacy path for Viktor's actual ask (user-triggered customer/app config). It reaches the controller — the component that already pulls config — with a hub-only change (a hanging-GET endpoint) plus raising the two 60 s ceilings, no agent train, no ep0 coupling, no WG dependency. The box holds a GET with a hold window comfortably under the ingress/NAT ceiling (e.g. 45–55 s today, or longer once the ceilings are raised) and reconnects; the hub completes the GET the instant operator intent changes.
- (a) poke-relay is the natural agent-plane nudge for host-domain desired-state (guests, wireguard) — it reaches the agent, needs zero ep0/box infra change (P1), rides the OOB / mutual-repair arc already in place, and is sub-second. Its cost is an agent train (UDP listener → heartbeat trigger). Reserve it for the agent's desired-state, where the long-poll's "reaches the controller" advantage doesn't apply.
Interval shortening stays REJECTED. Shrinking the 900 s hub poll / 15-min report cycle fleet-wide multiplies steady-state load on every box and the hub continuously, for an event that is rare and user-triggered, and still does not deliver true immediacy (latency is bounded by the shortened interval, not seconds). A wake signal delivers seconds-latency at ~zero idle cost. Rejected.
CLEANUP — inventory & assertion
Every live mutation, reverted (revised to the LIVE P0 baseline, not the stale 07-05 assumptions):
- ep0: spike forced-command line removed from
/root/.ssh/authorized_keys(grep -c spike-immediate-sync= 0, back to 2 lines);/root/spike-poke.shshred -u'd;ip_forwardstill 1 (never touched — the poke needs none);forwardchain still empty; wg0 peer list unchanged (3, incl. the untouched second box10.77.0.3and dormant operator10.77.0.250). - DooPlex / k8s: sleeper
Deployment+Service+Ingress+ConfigMap(all labelledapp: spike-immediate-sync) deleted;/tmp/spike-sleeper.yamland/tmp/spike-poke-key*removed. No change to the hub Ingress, the hub Deployment, or any felhom-system object. - Box: all spike listeners / tcpdumps / http.servers gone (timeout-bounded);
/tmp/spike-*removed; nft still empty;wg-felhomunchanged. - Final re-verify (all green): box
wg-quick@wg-felhom+felhom-agentactive, tunnel live (pvesm status felhom-offsite= active; box↔ep0 last-handshake epoch matches on both ends), guest 9201 running +felhom-controller Up (healthy); ep0 wg0 3 peers unchanged (second box10.77.0.3+ dormant operator10.77.0.250untouched); hub pod 1/1 Running, reconciler still pushing 3 peers. ep0ip_forward1 (never changed),forwardchain empty, rootauthorized_keysback to 2 lines (0 spike lines). - No repo/hub/agent/manifest production change — this commit is docs-only.
Method bar honored
wg show <if> dump never run (S1 key-leak ban); only latest-handshakes/redacted conf reads.
Throwaway keys generated on the box that owns them, 0600, private keys never printed. Every negative
carries kernel/nft/packet-level evidence (EKEYREJECTED errno; nginx upstream-timeout log line;
tcpdump counts), not bare timeouts. The live baseline was re-measured rather than assumed — the two
places the plan's assumptions had gone stale (ip_forward now 1; the operator /32 now rendered)
are recorded as findings. No probe touched Peti's arc, the hub DB, the live peer list, or the
box's rendered WG conf.