Files
felhom.eu/documentation/audits/SPIKE-immediate-sync-transport-2026-07-16.md
T
admin 10e07f5747 docs(spike): immediate-sync transport — poke-relay vs long-poll, both GO/complementary
Empirical spike (demo arc, docs-only, no product code). Measures the two hub->box
'sync now' transports. Verdict: (b) long-poll = primary for user config (hub-only,
reaches controller, no ep0/WG/agent coupling; binding ceiling = twin 60s nginx
proxy_read_timeout + hub WriteTimeout, both config lifts); (a) ep0-relayed poke
= agent-plane nudge (~0.42s, zero ep0/box infra change, EKEYREJECTED-confined).
Viktor Q2 answered: hub.felhom.eu is DNS-only, no Cloudflare. All live mutations
reverted; arc re-verified green.
2026-07-16 20:00:01 +02:00

258 lines
18 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# SPIKE — immediate-sync transport (hub→box "sync now" signal) — 2026-07-16
> **STATUS: COMPLETE — verdict below.** Empirically measures the two candidate transports for the
> hub→box immediacy signal BEFORE any production spec. No production code shipped; no
> agent/controller/hub image; every live mutation inventoried and reverted (§CLEANUP, verified).
> The question under test (Viktor's ruling): a user-triggered config change must round-trip in
> **seconds**, not the 15-min report cycle — so the hub needs a "sync now" wake-up and the box then
> pulls config via the NORMAL channel (report ACK / desired-state fetch; the hub never delivers
> config by connecting inbound).
**Class:** SPIKE (empirical; no product code). **Repos:** `felhom.eu` (this doc only). All others
read-only for grounding: `hub/internal/wgsync/{client,reconciler}.go`, `hub/cmd/hub/main.go`,
`manifests/hub.yaml`, `felhom-agent/internal/desired/syncer.go`.
**Baselines re-confirmed at session start (git HEAD):** felhom.eu `a50822af` v0.57.0 · felhom-agent
`c040c180` v0.88.0 · felhom-controller `8f3564c1` v0.138.0. Live hub image `:0.57.0`.
**Probe ends (demo arc ONLY; Peti's arc untouched):** box = `felhom-pve`/demo-felhom
(`192.168.0.162`, LAN SSH = safety line) · endpoint = `felhom-hetzner`/ep0 (`167.233.158.164`,
root SSH, wg0 `10.77.0.1`) · hub pod on DooPlex k3s (`192.168.0.180`, `kubectl -n felhom-system`).
---
## Verdict (one line)
**Both transports are GO and they are COMPLEMENTARY, not competing.** (b) **long-poll** delivers to
the **controller** (the config puller) with a **hub-only** change and **zero ep0 coupling**, and is
the lighter fit for Viktor's user-config round-trip — but only after two 60 s ceilings are raised
(nginx `proxy_read_timeout` **measured at 60 s**, and the hub's in-code `WriteTimeout` **60 s**);
today a naive long-poll is capped at ~55 s → ~1 reconnect/box/min. (a) **ep0-relayed poke** delivers
to the **agent** (WG terminates host-side) in **~0.42 s per fresh SSH session** (≈27 ms+18 ms if the
session is pooled), needs **zero infra change on ep0 or the box** (origination from wg0's own IP
never touches the forward chain — proven), and is cryptographically confined to registered peers —
but it needs an **agent train** (a new UDP listener→heartbeat trigger) and a hub forced-command key
on ep0. Recommend **(b) as the primary immediacy path for customer/app config** (hub+controller
only), with **(a) reserved as the agent-plane desired-state nudge** already aligned with the OOB /
mutual-repair arc. **Interval shortening stays REJECTED.**
---
## P0 — ground truth (no changes) — GATE PASSED
All three ends reachable → proceed. Redacted reads only; `wg show <if> dump` never run (S1 key-leak
ban); private keys never printed.
### DNS / fronting — **Viktor's Q2 answered: DNS-only, NO Cloudflare**
- **Public:** `hub.felhom.eu` → CNAME `dooplex.hopto.org`**37.191.56.193** (the operator's
One-Hungary fixed-cable line — the SAME public IPv4 the demo box sits behind). No `cf-ray`, no
`server: cloudflare`.
- **On-LAN split-horizon** (resolver `192.168.0.250`): `hub.felhom.eu`**192.168.0.192** (the
`nginx-internal` ingress LB, `nginx-system/ingress-nginx-internal-controller`).
- `curl -sI` on-LAN `/healthz``HTTP/1.1 200`, HSTS, no CDN headers. From ep0 via the public IP
(`--resolve hub:443:37.191.56.193`) → `HTTP/2 403` = the **HU geo-gate** (ep0 is in DE; the gate
is a location-level `configuration-snippet` on the hub `/` path). The path reaches nginx over
HTTP/2 and returns immediately.
- **Real ingress chain:** `[client] → internet → 37.191.56.193 (operator home-router NAT) → k3s
nginx-internal → hub pod:8080`. Self-hosted, no CDN in front. (Controller→hub and agent→hub both
ride this same DNS-only nginx path — the controller's Cloudflare tunnel is its *inbound* edge only.)
### Box (`felhom-pve` / demo-felhom)
`wg-felhom`: `Address 10.77.0.2/32`, MTU 1280; peer `f3d1ZI7…`(ep0) `Endpoint 167.233.158.164:443`,
`AllowedIPs 10.77.0.1/32, 10.77.0.250/32`, `PersistentKeepalive 25`. `registered.json`: assigned
`10.77.0.2/32`, **generation 11**, pubkey `aqMpucg…`. Handshake fresh (~1 min); `wg-quick@wg-felhom`
+ `felhom-agent` both active; **agent 0.88.0**. Box **nft empty** (no host belt; pve-firewall
disabled) — matches the 07-05 baseline.
- *Note vs 07-05:* the box's WG **identity rotated** (new pubkey) on the R-21 ISO reinstall but the
assigned IP was preserved (`10.77.0.2`). The box `AllowedIPs` now **already carries
`10.77.0.250/32`** — i.e. the 07-05 spike's finding-1 gap ("render the operator /32 into the agent
conf") has since been **closed**; the OOB operator peer is durably rendered.
### Endpoint (`felhom-hetzner`, wg0 :443)
Three peers: box `aqMpucg… → 10.77.0.2/32` (fresh); a **second real box** `zxy… → 10.77.0.3/32`
(handshake ~3h49m — **not touched**); operator `uDPlMa0… → 10.77.0.250/32` (no endpoint — dormant
OOB peer). **`ip_forward = 1`** (ipv6 fwd 0). `nft inet filter`: `input` policy DROP with allows
(established, lo, tcp/22, udp/443, tcp/8007 iif wg0, icmp); `forward` policy DROP and **empty** (no
accept rules). `authorized_keys`: root **2** lines, `felhom-peersync` **2** lines.
- *Note vs 07-05:* `ip_forward` was **0** at the 07-05 baseline and is **1** now — the OOB
operator-peer feature has shipped in the interim. **The live baseline (`ip_forward=1`, forward
chain empty) is therefore the cleanup target — NOT the stale "must be 0" assumption in the plan.**
### Hub pod
`hub-7b97cfb48d-qrbpb` Running, image `:0.57.0`. `exec` works (Alpine, root). **Reconciler ALIVE:**
`[INFO] wgsync: pushed 3 peers to 167.233.158.164:22` every 5 min — so **hub-pod → ep0:22 already
works in production** (the forced-command peersync channel a poke-relay would clone).
---
## P1 — ep0→box UDP origination (the poke's last hop) — **GO, zero endpoint change**
Listener on the box bound to its WG IP (`10.77.0.2:19999`) + `tcpdump -ni wg-felhom`; ep0 originates
from wg0.
1. **Delivery, packet-proven:** ep0 `python3 sendto("10.77.0.2",19999)` from `10.77.0.1`. Box:
`RECV b'sync-poke-p1' from ('10.77.0.1', 47280)`. tcpdump captured **2** pkts
`10.77.0.1.* > 10.77.0.2.19999` (python + a `nc` variant).
2. **The load-bearing claim — ZERO mutation needed:** `ip_forward` **1 → 1 unchanged**; `forward`
chain **still empty policy-drop**; no nft/peer/route change. Endpoint-**originated** traffic
leaves via wg0's output path and **never traverses the `forward` chain** — so the poke needs no
`ip_forward`, no forward rule, no nft edit. (This is exactly why forwarding being on or off is
irrelevant to the relay: the poke is origination, not transit.)
3. **Latency over the tunnel:** ping **avg 29.0 ms** (min 22.6 / max 36.9 / mdev 4.8); HTTP
round-trip connect ~2331 ms, total ~5566 ms (5/5 `http=200`).
4. **Negative — cryptographically confined to registered peers:** poking an **unregistered**
in-subnet IP (`10.77.0.199`; the route exists, `dev wg0 src 10.77.0.1`) fails at the kernel:
`OSError: [Errno 126] Required key not available` (EKEYREJECTED) — WireGuard refuses to encrypt
to an IP no peer owns. wg0 peer count unchanged (3). The future forced command can therefore
**only** ever reach a registered peer's `/32`; there is no "spray the subnet" failure mode.
---
## P2 — hub-pod → ep0 → box poke round-trip — **GO, ~0.42 s per poke (target was ≤23 s)**
Throwaway keypair generated **on the DooPlex host** (`/tmp/spike-poke-key`, 0600, private key never
printed); **public** key installed on ep0 root `authorized_keys` as
`command="/root/spike-poke.sh",restrict …` (root ak **2 → 3**, 1 spike line). The forced command is
one UDP `sendto` to the box.
1. **Pod capability:** the hub image is **Alpine** — it has `nc` but **no `ssh`, no `bash`, no
`python3`**. So a production poke MUST reuse the **in-process `golang.org/x/crypto/ssh`** client
(the `wgsync` pattern), not a shelled-out binary. Pod→ep0:22 reachability from the **pod netns**
confirmed: `nc -z -w4 167.233.158.164 22` = OK (and the reconciler proves it continuously).
2. **End-to-end, 5 timed pokes** (DooPlex host `--ssh--> ep0 forced-cmd --> UDP --> box`): all 5
delivered (box `RECV #15 b'hub-poke' from 10.77.0.1`), remote returned `poke-fired`.
`ssh_wall`: **0.636 s (cold), 0.423, 0.439, 0.414, 0.412** → warm median **~0.42 s**.
3. **Cost decomposition:** forced-command body (python startup + `sendto`) = **1820 ms** (bare
python3 = 1112 ms); DooPlex→ep0 net RTT = **27 ms**. So **~0.40 s of the 0.42 s is SSH
TCP+KEX+auth+shell-spawn** — the session setup dominates, the poke itself is negligible.
**Task-2 lever:** a **pooled/persistent** SSH session (or the in-process client holding one
channel) collapses per-poke to ~**1 RTT (27 ms) + body**; even a fresh session per poke is already
sub-second.
4. **Failure behavior:** with **no box listener**, the poke still fires on ep0 and the datagram is
silently discarded at the box — **no error path back to the hub** (a lost poke is harmless by
construction; the next scheduled report reconciles). **Stale handshake:** NOT force-tested —
forcing it means stopping keepalive/the tunnel, which risks the live offsite/PBS path (STOP
rule). From WG semantics + observation: `PersistentKeepalive 25` kept the handshake < 125 s fresh
throughout; a poke arriving on a cold handshake triggers WG's auto-rehandshake on first tx (~1
extra RTT), datagram buffered then delivered — **not dropped**.
---
## P3 — long-poll hold-time through the REAL ingress chain
**Method (why no DNS record was created):** a throwaway **sleeper** (a `python:3.14-bookworm` pod
that `sleep`s `?hold=N` before a 200) + Service + **Ingress on host `probe.felhom.eu`, class
`nginx-internal`**, in `felhom-system`. Reached with `curl --resolve probe.felhom.eu:443:<ip> -k`
from each vantage — so **no DNS record was needed** (the plan's fallback "ask Viktor for an A
record" was avoided), **the live hub Ingress was never touched**, and the sleeper's own server block
carries **no geo snippet** (so the remote ep0 vantage is not geo-blocked). Default cert → `curl -k`.
The `nginx-internal` ConfigMap sets **no** custom proxy timeouts → nginx built-in defaults.
### Pass 1 — DEFAULT ingress, on-LAN box (direct nginx, no NAT)
| hold N | result |
|---|---|
| 30 s | `200 @ 30.0 s` |
| 55 s | `200 @ 55.0 s` |
| 70 s | **KILLED** |
| 100 s | **KILLED** |
nginx error log (authoritative): `upstream timed out (110: Operation timed out) while reading
response header from upstream … GET /?hold=70 HTTP/2.0`. **Killing layer = nginx
`proxy_read_timeout` = 60 s** (built-in default). Each request appears **twice** ~60 s apart =
ingress-nginx's default single upstream **retry on timeout**, so the client-visible 504 lands at
**~120 s** (2 × 60 s); the client's own cut at 110 s showed `http=000` just before it. The sleeper's
`BrokenPipeError` confirms nginx abandoned the read at 60 s.
- **H(default) = ~55 s usable / hard kill at 60 s** → naive per-box reconnect rate ≈ **1/min**.
- **In-code twin ceiling:** the hub's own `http.Server{WriteTimeout: 60s}` (`hub/cmd/hub/main.go:339341`)
caps a held response at ~60 s regardless of ingress — coincident with nginx's 60 s. **Both must be
raised** for any long-poll (nginx via annotation; the hub per-endpoint via `http.ResponseController`).
### Pass 2 — TUNED sleeper ingress (`proxy-read-timeout`/`proxy-send-timeout`=3600, `proxy-next-upstream`=off)
- **on-LAN** `hold=120` → `200 @ 120.0 s` (was killed under default → **nginx WAS the layer**, and
raising it lifts the ceiling cleanly; on-LAN there is no further layer below the pure-nginx value).
- **on-LAN** `hold=600` → `200 @ 600.0 s` (pure tuned-nginx holds a full **10 min** cleanly — no
further on-LAN layer below 600 s).
- **ep0 remote** (through the operator home-router NAT `37.191.56.193`): `hold=2` → `200 @ 2.1 s`
(routes, not geo-blocked); `hold=200` → `200 @ 200.09 s`; **`hold=600` → `200 @ 600.09 s`** — the
**operator home NAT holds an idle held connection for the full 10 min** without eviction. So on
this arc the NAT is **not** the binding constraint (≤ 600 s); the only ceiling today is the twin
60 s nginx+hub limits, **both config-fixable**. Once raised, a long-poll can hold **many minutes**
even through the operator NAT → reconnect rate drops from ~1/min toward ~1/10-min.
### Vantage caveat (honest ledger)
The demo box is **on-LAN, co-located behind the SAME operator NAT** as the hub, so it does **not**
traverse a real customer's `home-NAT → internet → operator-NAT` double-NAT path. The ep0 vantage
represents the **ingress-side (operator) NAT** only; a **customer-side home-router idle timeout** is
**unrepresented** by this arc. The load-bearing measurement — that **nginx's 60 s default is the
binding ceiling today, and it is a config lift** — holds regardless of vantage.
---
## P4 — verdict: (a) poke-relay vs (b) long-poll
| Dimension | **(a) ep0-relayed poke** | **(b) long-poll** |
|---|---|---|
| Measured latency | **~0.42 s** per fresh SSH session (≈**27 ms + 18 ms** if pooled); last hop ep0→box **~29 ms** | delivers on the next held-GET completion; today **capped at ~55 s** by the nginx+hub 60 s ceilings (both fixable) |
| Reaches | the **agent** (WG terminates host-side) → desired-state nudge | the **controller** (the config puller) directly — also usable on the agent channel |
| Hub change | new forced-command poke key + in-process SSH poke call | new hanging-GET endpoint + **raise nginx & hub WriteTimeout** |
| ep0 change | **none** (P1: origination needs no forward/nft/sysctl) | **none** |
| Box/agent change | **agent train** — new UDP listener → "heartbeat now" trigger | **none** if the controller holds the GET (it already pulls config); or the agent |
| ep0 coupling | **yes** (rides the peersync SSH pattern) | **no** (direct hub↔box outbound over the existing ingress) |
| WG down | poke fails (WG is the transport) — but agent self-heals WG in ~15 s (07-05) and a lost poke is harmless; next report catches up | **unaffected** (uses the ingress path, not WG) |
| Hub restart | forced-command is stateless; resumes on next trigger | held GETs drop → box must reconnect (reconnect logic needed) |
| NAT rebind | WG keepalive re-handshakes; poke buffered | held GET dies silently on rebind → client hold must be < NAT idle timeout |
**Recommended Task-2 shape — complementary, not either/or:**
- **(b) long-poll is the primary immediacy path for Viktor's actual ask** (user-triggered
customer/app **config**). It reaches the **controller — the component that already pulls config** —
with a **hub-only** change (a hanging-GET endpoint) plus raising the two 60 s ceilings, **no agent
train, no ep0 coupling, no WG dependency**. The box holds a GET with a hold window comfortably
under the ingress/NAT ceiling (e.g. 4555 s today, or longer once the ceilings are raised) and
reconnects; the hub completes the GET the instant operator intent changes.
- **(a) poke-relay is the natural agent-plane nudge** for host-domain desired-state (guests,
wireguard) — it reaches the **agent**, needs **zero ep0/box infra change** (P1), rides the OOB /
mutual-repair arc already in place, and is **sub-second**. Its cost is an agent train (UDP
listener → heartbeat trigger). Reserve it for the agent's desired-state, where the long-poll's
"reaches the controller" advantage doesn't apply.
**Interval shortening stays REJECTED.** Shrinking the 900 s hub poll / 15-min report cycle
fleet-wide multiplies steady-state load on **every** box and the hub **continuously**, for an event
that is **rare and user-triggered**, and still does not deliver true immediacy (latency is bounded by
the shortened interval, not seconds). A wake signal delivers seconds-latency at ~zero idle cost.
Rejected.
---
## CLEANUP — inventory & assertion
Every live mutation, reverted (revised to the LIVE P0 baseline, not the stale 07-05 assumptions):
- **ep0:** spike forced-command line removed from `/root/.ssh/authorized_keys` (`grep -c
spike-immediate-sync` = **0**, back to **2** lines); `/root/spike-poke.sh` `shred -u`'d;
**`ip_forward` still 1** (never touched — the poke needs none); `forward` chain still empty; wg0
peer list unchanged (3, incl. the untouched second box `10.77.0.3` and dormant operator
`10.77.0.250`).
- **DooPlex / k8s:** sleeper `Deployment` + `Service` + `Ingress` + `ConfigMap` (all labelled
`app: spike-immediate-sync`) deleted; `/tmp/spike-sleeper.yaml` and `/tmp/spike-poke-key*` removed.
**No change to the hub Ingress, the hub Deployment, or any felhom-system object.**
- **Box:** all spike listeners / tcpdumps / http.servers gone (timeout-bounded); `/tmp/spike-*`
removed; **nft still empty**; `wg-felhom` unchanged.
- **Final re-verify (all green):** box `wg-quick@wg-felhom` + `felhom-agent` **active**, tunnel live
(`pvesm status felhom-offsite` = **active**; box↔ep0 last-handshake epoch matches on both ends),
guest 9201 **running** + `felhom-controller Up (healthy)`; ep0 wg0 **3 peers unchanged**
(second box `10.77.0.3` + dormant operator `10.77.0.250` untouched); hub pod **1/1 Running**,
reconciler still pushing 3 peers. ep0 `ip_forward` **1** (never changed), `forward` chain empty,
root `authorized_keys` back to **2** lines (0 spike lines).
- **No repo/hub/agent/manifest production change** — this commit is docs-only.
---
## Method bar honored
`wg show <if> dump` never run (S1 key-leak ban); only `latest-handshakes`/redacted conf reads.
Throwaway keys generated on the box that owns them, 0600, private keys never printed. Every negative
carries kernel/nft/packet-level evidence (EKEYREJECTED errno; nginx upstream-timeout log line;
tcpdump counts), not bare timeouts. The live baseline was re-measured rather than assumed — the two
places the plan's assumptions had gone stale (`ip_forward` now 1; the operator `/32` now rendered)
are recorded as findings. No probe touched Peti's arc, the hub DB, the live peer list, or the
box's rendered WG conf.