Files
felhom.eu/documentation/audits/SPIKE-immediate-sync-transport-2026-07-16.md
admin 10e07f5747 docs(spike): immediate-sync transport — poke-relay vs long-poll, both GO/complementary
Empirical spike (demo arc, docs-only, no product code). Measures the two hub->box
'sync now' transports. Verdict: (b) long-poll = primary for user config (hub-only,
reaches controller, no ep0/WG/agent coupling; binding ceiling = twin 60s nginx
proxy_read_timeout + hub WriteTimeout, both config lifts); (a) ep0-relayed poke
= agent-plane nudge (~0.42s, zero ep0/box infra change, EKEYREJECTED-confined).
Viktor Q2 answered: hub.felhom.eu is DNS-only, no Cloudflare. All live mutations
reverted; arc re-verified green.
2026-07-16 20:00:01 +02:00

18 KiB
Raw Permalink Blame History

SPIKE — immediate-sync transport (hub→box "sync now" signal) — 2026-07-16

STATUS: COMPLETE — verdict below. Empirically measures the two candidate transports for the hub→box immediacy signal BEFORE any production spec. No production code shipped; no agent/controller/hub image; every live mutation inventoried and reverted (§CLEANUP, verified). The question under test (Viktor's ruling): a user-triggered config change must round-trip in seconds, not the 15-min report cycle — so the hub needs a "sync now" wake-up and the box then pulls config via the NORMAL channel (report ACK / desired-state fetch; the hub never delivers config by connecting inbound).

Class: SPIKE (empirical; no product code). Repos: felhom.eu (this doc only). All others read-only for grounding: hub/internal/wgsync/{client,reconciler}.go, hub/cmd/hub/main.go, manifests/hub.yaml, felhom-agent/internal/desired/syncer.go.

Baselines re-confirmed at session start (git HEAD): felhom.eu a50822af v0.57.0 · felhom-agent c040c180 v0.88.0 · felhom-controller 8f3564c1 v0.138.0. Live hub image :0.57.0.

Probe ends (demo arc ONLY; Peti's arc untouched): box = felhom-pve/demo-felhom (192.168.0.162, LAN SSH = safety line) · endpoint = felhom-hetzner/ep0 (167.233.158.164, root SSH, wg0 10.77.0.1) · hub pod on DooPlex k3s (192.168.0.180, kubectl -n felhom-system).


Verdict (one line)

Both transports are GO and they are COMPLEMENTARY, not competing. (b) long-poll delivers to the controller (the config puller) with a hub-only change and zero ep0 coupling, and is the lighter fit for Viktor's user-config round-trip — but only after two 60 s ceilings are raised (nginx proxy_read_timeout measured at 60 s, and the hub's in-code WriteTimeout 60 s); today a naive long-poll is capped at ~55 s → ~1 reconnect/box/min. (a) ep0-relayed poke delivers to the agent (WG terminates host-side) in ~0.42 s per fresh SSH session (≈27 ms+18 ms if the session is pooled), needs zero infra change on ep0 or the box (origination from wg0's own IP never touches the forward chain — proven), and is cryptographically confined to registered peers — but it needs an agent train (a new UDP listener→heartbeat trigger) and a hub forced-command key on ep0. Recommend (b) as the primary immediacy path for customer/app config (hub+controller only), with (a) reserved as the agent-plane desired-state nudge already aligned with the OOB / mutual-repair arc. Interval shortening stays REJECTED.


P0 — ground truth (no changes) — GATE PASSED

All three ends reachable → proceed. Redacted reads only; wg show <if> dump never run (S1 key-leak ban); private keys never printed.

DNS / fronting — Viktor's Q2 answered: DNS-only, NO Cloudflare

  • Public: hub.felhom.eu → CNAME dooplex.hopto.org37.191.56.193 (the operator's One-Hungary fixed-cable line — the SAME public IPv4 the demo box sits behind). No cf-ray, no server: cloudflare.
  • On-LAN split-horizon (resolver 192.168.0.250): hub.felhom.eu192.168.0.192 (the nginx-internal ingress LB, nginx-system/ingress-nginx-internal-controller).
  • curl -sI on-LAN /healthzHTTP/1.1 200, HSTS, no CDN headers. From ep0 via the public IP (--resolve hub:443:37.191.56.193) → HTTP/2 403 = the HU geo-gate (ep0 is in DE; the gate is a location-level configuration-snippet on the hub / path). The path reaches nginx over HTTP/2 and returns immediately.
  • Real ingress chain: [client] → internet → 37.191.56.193 (operator home-router NAT) → k3s nginx-internal → hub pod:8080. Self-hosted, no CDN in front. (Controller→hub and agent→hub both ride this same DNS-only nginx path — the controller's Cloudflare tunnel is its inbound edge only.)

Box (felhom-pve / demo-felhom)

wg-felhom: Address 10.77.0.2/32, MTU 1280; peer f3d1ZI7…(ep0) Endpoint 167.233.158.164:443, AllowedIPs 10.77.0.1/32, 10.77.0.250/32, PersistentKeepalive 25. registered.json: assigned 10.77.0.2/32, generation 11, pubkey aqMpucg…. Handshake fresh (~1 min); wg-quick@wg-felhom

  • felhom-agent both active; agent 0.88.0. Box nft empty (no host belt; pve-firewall disabled) — matches the 07-05 baseline.
  • Note vs 07-05: the box's WG identity rotated (new pubkey) on the R-21 ISO reinstall but the assigned IP was preserved (10.77.0.2). The box AllowedIPs now already carries 10.77.0.250/32 — i.e. the 07-05 spike's finding-1 gap ("render the operator /32 into the agent conf") has since been closed; the OOB operator peer is durably rendered.

Endpoint (felhom-hetzner, wg0 :443)

Three peers: box aqMpucg… → 10.77.0.2/32 (fresh); a second real box zxy… → 10.77.0.3/32 (handshake ~3h49m — not touched); operator uDPlMa0… → 10.77.0.250/32 (no endpoint — dormant OOB peer). ip_forward = 1 (ipv6 fwd 0). nft inet filter: input policy DROP with allows (established, lo, tcp/22, udp/443, tcp/8007 iif wg0, icmp); forward policy DROP and empty (no accept rules). authorized_keys: root 2 lines, felhom-peersync 2 lines.

  • Note vs 07-05: ip_forward was 0 at the 07-05 baseline and is 1 now — the OOB operator-peer feature has shipped in the interim. The live baseline (ip_forward=1, forward chain empty) is therefore the cleanup target — NOT the stale "must be 0" assumption in the plan.

Hub pod

hub-7b97cfb48d-qrbpb Running, image :0.57.0. exec works (Alpine, root). Reconciler ALIVE: [INFO] wgsync: pushed 3 peers to 167.233.158.164:22 every 5 min — so hub-pod → ep0:22 already works in production (the forced-command peersync channel a poke-relay would clone).


P1 — ep0→box UDP origination (the poke's last hop) — GO, zero endpoint change

Listener on the box bound to its WG IP (10.77.0.2:19999) + tcpdump -ni wg-felhom; ep0 originates from wg0.

  1. Delivery, packet-proven: ep0 python3 sendto("10.77.0.2",19999) from 10.77.0.1. Box: RECV b'sync-poke-p1' from ('10.77.0.1', 47280). tcpdump captured 2 pkts 10.77.0.1.* > 10.77.0.2.19999 (python + a nc variant).
  2. The load-bearing claim — ZERO mutation needed: ip_forward 1 → 1 unchanged; forward chain still empty policy-drop; no nft/peer/route change. Endpoint-originated traffic leaves via wg0's output path and never traverses the forward chain — so the poke needs no ip_forward, no forward rule, no nft edit. (This is exactly why forwarding being on or off is irrelevant to the relay: the poke is origination, not transit.)
  3. Latency over the tunnel: ping avg 29.0 ms (min 22.6 / max 36.9 / mdev 4.8); HTTP round-trip connect ~2331 ms, total ~5566 ms (5/5 http=200).
  4. Negative — cryptographically confined to registered peers: poking an unregistered in-subnet IP (10.77.0.199; the route exists, dev wg0 src 10.77.0.1) fails at the kernel: OSError: [Errno 126] Required key not available (EKEYREJECTED) — WireGuard refuses to encrypt to an IP no peer owns. wg0 peer count unchanged (3). The future forced command can therefore only ever reach a registered peer's /32; there is no "spray the subnet" failure mode.

P2 — hub-pod → ep0 → box poke round-trip — GO, ~0.42 s per poke (target was ≤23 s)

Throwaway keypair generated on the DooPlex host (/tmp/spike-poke-key, 0600, private key never printed); public key installed on ep0 root authorized_keys as command="/root/spike-poke.sh",restrict … (root ak 2 → 3, 1 spike line). The forced command is one UDP sendto to the box.

  1. Pod capability: the hub image is Alpine — it has nc but no ssh, no bash, no python3. So a production poke MUST reuse the in-process golang.org/x/crypto/ssh client (the wgsync pattern), not a shelled-out binary. Pod→ep0:22 reachability from the pod netns confirmed: nc -z -w4 167.233.158.164 22 = OK (and the reconciler proves it continuously).
  2. End-to-end, 5 timed pokes (DooPlex host --ssh--> ep0 forced-cmd --> UDP --> box): all 5 delivered (box RECV #15 b'hub-poke' from 10.77.0.1), remote returned poke-fired. ssh_wall: 0.636 s (cold), 0.423, 0.439, 0.414, 0.412 → warm median ~0.42 s.
  3. Cost decomposition: forced-command body (python startup + sendto) = 1820 ms (bare python3 = 1112 ms); DooPlex→ep0 net RTT = 27 ms. So ~0.40 s of the 0.42 s is SSH TCP+KEX+auth+shell-spawn — the session setup dominates, the poke itself is negligible. Task-2 lever: a pooled/persistent SSH session (or the in-process client holding one channel) collapses per-poke to ~1 RTT (27 ms) + body; even a fresh session per poke is already sub-second.
  4. Failure behavior: with no box listener, the poke still fires on ep0 and the datagram is silently discarded at the box — no error path back to the hub (a lost poke is harmless by construction; the next scheduled report reconciles). Stale handshake: NOT force-tested — forcing it means stopping keepalive/the tunnel, which risks the live offsite/PBS path (STOP rule). From WG semantics + observation: PersistentKeepalive 25 kept the handshake < 125 s fresh throughout; a poke arriving on a cold handshake triggers WG's auto-rehandshake on first tx (~1 extra RTT), datagram buffered then delivered — not dropped.

P3 — long-poll hold-time through the REAL ingress chain

Method (why no DNS record was created): a throwaway sleeper (a python:3.14-bookworm pod that sleeps ?hold=N before a 200) + Service + Ingress on host probe.felhom.eu, class nginx-internal, in felhom-system. Reached with curl --resolve probe.felhom.eu:443:<ip> -k from each vantage — so no DNS record was needed (the plan's fallback "ask Viktor for an A record" was avoided), the live hub Ingress was never touched, and the sleeper's own server block carries no geo snippet (so the remote ep0 vantage is not geo-blocked). Default cert → curl -k. The nginx-internal ConfigMap sets no custom proxy timeouts → nginx built-in defaults.

Pass 1 — DEFAULT ingress, on-LAN box (direct nginx, no NAT)

hold N result
30 s 200 @ 30.0 s
55 s 200 @ 55.0 s
70 s KILLED
100 s KILLED

nginx error log (authoritative): upstream timed out (110: Operation timed out) while reading response header from upstream … GET /?hold=70 HTTP/2.0. Killing layer = nginx proxy_read_timeout = 60 s (built-in default). Each request appears twice ~60 s apart = ingress-nginx's default single upstream retry on timeout, so the client-visible 504 lands at ~120 s (2 × 60 s); the client's own cut at 110 s showed http=000 just before it. The sleeper's BrokenPipeError confirms nginx abandoned the read at 60 s.

  • H(default) = ~55 s usable / hard kill at 60 s → naive per-box reconnect rate ≈ 1/min.
  • In-code twin ceiling: the hub's own http.Server{WriteTimeout: 60s} (hub/cmd/hub/main.go:339341) caps a held response at ~60 s regardless of ingress — coincident with nginx's 60 s. Both must be raised for any long-poll (nginx via annotation; the hub per-endpoint via http.ResponseController).

Pass 2 — TUNED sleeper ingress (proxy-read-timeout/proxy-send-timeout=3600, proxy-next-upstream=off)

  • on-LAN hold=120200 @ 120.0 s (was killed under default → nginx WAS the layer, and raising it lifts the ceiling cleanly; on-LAN there is no further layer below the pure-nginx value).
  • on-LAN hold=600200 @ 600.0 s (pure tuned-nginx holds a full 10 min cleanly — no further on-LAN layer below 600 s).
  • ep0 remote (through the operator home-router NAT 37.191.56.193): hold=2200 @ 2.1 s (routes, not geo-blocked); hold=200200 @ 200.09 s; hold=600200 @ 600.09 s — the operator home NAT holds an idle held connection for the full 10 min without eviction. So on this arc the NAT is not the binding constraint (≤ 600 s); the only ceiling today is the twin 60 s nginx+hub limits, both config-fixable. Once raised, a long-poll can hold many minutes even through the operator NAT → reconnect rate drops from ~1/min toward ~1/10-min.

Vantage caveat (honest ledger)

The demo box is on-LAN, co-located behind the SAME operator NAT as the hub, so it does not traverse a real customer's home-NAT → internet → operator-NAT double-NAT path. The ep0 vantage represents the ingress-side (operator) NAT only; a customer-side home-router idle timeout is unrepresented by this arc. The load-bearing measurement — that nginx's 60 s default is the binding ceiling today, and it is a config lift — holds regardless of vantage.


P4 — verdict: (a) poke-relay vs (b) long-poll

Dimension (a) ep0-relayed poke (b) long-poll
Measured latency ~0.42 s per fresh SSH session (≈27 ms + 18 ms if pooled); last hop ep0→box ~29 ms delivers on the next held-GET completion; today capped at ~55 s by the nginx+hub 60 s ceilings (both fixable)
Reaches the agent (WG terminates host-side) → desired-state nudge the controller (the config puller) directly — also usable on the agent channel
Hub change new forced-command poke key + in-process SSH poke call new hanging-GET endpoint + raise nginx & hub WriteTimeout
ep0 change none (P1: origination needs no forward/nft/sysctl) none
Box/agent change agent train — new UDP listener → "heartbeat now" trigger none if the controller holds the GET (it already pulls config); or the agent
ep0 coupling yes (rides the peersync SSH pattern) no (direct hub↔box outbound over the existing ingress)
WG down poke fails (WG is the transport) — but agent self-heals WG in ~15 s (07-05) and a lost poke is harmless; next report catches up unaffected (uses the ingress path, not WG)
Hub restart forced-command is stateless; resumes on next trigger held GETs drop → box must reconnect (reconnect logic needed)
NAT rebind WG keepalive re-handshakes; poke buffered held GET dies silently on rebind → client hold must be < NAT idle timeout

Recommended Task-2 shape — complementary, not either/or:

  • (b) long-poll is the primary immediacy path for Viktor's actual ask (user-triggered customer/app config). It reaches the controller — the component that already pulls config — with a hub-only change (a hanging-GET endpoint) plus raising the two 60 s ceilings, no agent train, no ep0 coupling, no WG dependency. The box holds a GET with a hold window comfortably under the ingress/NAT ceiling (e.g. 4555 s today, or longer once the ceilings are raised) and reconnects; the hub completes the GET the instant operator intent changes.
  • (a) poke-relay is the natural agent-plane nudge for host-domain desired-state (guests, wireguard) — it reaches the agent, needs zero ep0/box infra change (P1), rides the OOB / mutual-repair arc already in place, and is sub-second. Its cost is an agent train (UDP listener → heartbeat trigger). Reserve it for the agent's desired-state, where the long-poll's "reaches the controller" advantage doesn't apply.

Interval shortening stays REJECTED. Shrinking the 900 s hub poll / 15-min report cycle fleet-wide multiplies steady-state load on every box and the hub continuously, for an event that is rare and user-triggered, and still does not deliver true immediacy (latency is bounded by the shortened interval, not seconds). A wake signal delivers seconds-latency at ~zero idle cost. Rejected.


CLEANUP — inventory & assertion

Every live mutation, reverted (revised to the LIVE P0 baseline, not the stale 07-05 assumptions):

  • ep0: spike forced-command line removed from /root/.ssh/authorized_keys (grep -c spike-immediate-sync = 0, back to 2 lines); /root/spike-poke.sh shred -u'd; ip_forward still 1 (never touched — the poke needs none); forward chain still empty; wg0 peer list unchanged (3, incl. the untouched second box 10.77.0.3 and dormant operator 10.77.0.250).
  • DooPlex / k8s: sleeper Deployment + Service + Ingress + ConfigMap (all labelled app: spike-immediate-sync) deleted; /tmp/spike-sleeper.yaml and /tmp/spike-poke-key* removed. No change to the hub Ingress, the hub Deployment, or any felhom-system object.
  • Box: all spike listeners / tcpdumps / http.servers gone (timeout-bounded); /tmp/spike-* removed; nft still empty; wg-felhom unchanged.
  • Final re-verify (all green): box wg-quick@wg-felhom + felhom-agent active, tunnel live (pvesm status felhom-offsite = active; box↔ep0 last-handshake epoch matches on both ends), guest 9201 running + felhom-controller Up (healthy); ep0 wg0 3 peers unchanged (second box 10.77.0.3 + dormant operator 10.77.0.250 untouched); hub pod 1/1 Running, reconciler still pushing 3 peers. ep0 ip_forward 1 (never changed), forward chain empty, root authorized_keys back to 2 lines (0 spike lines).
  • No repo/hub/agent/manifest production change — this commit is docs-only.

Method bar honored

wg show <if> dump never run (S1 key-leak ban); only latest-handshakes/redacted conf reads. Throwaway keys generated on the box that owns them, 0600, private keys never printed. Every negative carries kernel/nft/packet-level evidence (EKEYREJECTED errno; nginx upstream-timeout log line; tcpdump counts), not bare timeouts. The live baseline was re-measured rather than assumed — the two places the plan's assumptions had gone stale (ip_forward now 1; the operator /32 now rendered) are recorded as findings. No probe touched Peti's arc, the hub DB, the live peer list, or the box's rendered WG conf.