Files
felhom.eu/documentation/audits/SPIKE-oob-wg-operator-peer-2026-07-05.md
T
admin 2db92c8837 docs(audit): OOB-over-WG operator-peer spike (2026-07-05) — GO
Validates operator-inbound access over the existing offsite WG arc (doc 06):
operator peer forwarded operator->box only, box sshd gated to the operator /32,
§4.5 box<->box isolation intact (both negatives counter-proven), mutual repair
real (agent self-healed a stopped tunnel in ~15s unaided). One TASK-shaping gap:
the operator /32 must be a RENDERED conf field — a runtime `wg set` is wiped by
the agent's own self-heal. All live-arc changes reverted + baseline re-verified.

Docs-only; no code/hub/agent/manifest change.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PSK5g6qYLknKj8u3QAFEr6
2026-07-05 17:01:35 +02:00

23 KiB

SPIKE — OOB management over the existing WG arc (operator peer, isolation, mutual repair) — 2026-07-05

STATUS: COMPLETE — verdict below. Empirically validates the unproven mechanisms of the OOB (out-of-band operator access) design BEFORE the production spec. No production code shipped; no agent binary change; every config change on the live demo arc reverted at cleanup (§13, verified). The design under test: an operator peer on the existing hub-and-spoke WG arc (doc 06), endpoint-forwarded operator→box only, box sshd reachable solely from the operator tunnel IP, while §4.5 box↔box isolation stays intact — and the hub desired-state channel as the tunnel's repair path.

Class: SPIKE (empirical; no product code). Repos: felhom.eu (this doc only); felhom-agent read-only for grounding (internal/wgtunnel/{manager,loop}.go, internal/desired/syncer.go, configs/felhom-agent.sudoers); felhom-hub read-only (internal/wgsync/reconciler.go, internal/store/wg.go).

Probe ends (the LIVE demo arc, ground-truthed in P0):

  • Endpoint = felhom-hetzner (Hetzner CX23, Debian 13, 167.233.158.164, DNS ep0.felhom.eu), the dev offsite endpoint: WG server wg0 on 443/UDP, tunnel 10.77.0.0/24, endpoint 10.77.0.1, PBS felhom-offsite. Peer list is hub-driven (internal/wgsync SSH push → felhom-peersync). CC-SSH-reachable as root ✓ (the P0 access gate passed).
  • Box = felhom-pve / node demo-felhom (192.168.0.162), agent v0.70.0, wg-felhom client 10.77.0.2/32, tunnel to ep0:443. CC's LAN SSH to it was the safety line throughout.

Verdict (one line): GO — an operator peer on the existing arc reaches a box's sshd through the endpoint with a single added /32, at ~24 ms interactive latency, WITHOUT breaking §4.5 isolation (both negatives held with packet-level evidence); mutual repair is REAL and already built (the agent self-healed a stopped tunnel in ~15 s with zero hub/LAN involvement, and OOB works with the agent dead); the one load-bearing gap the TASK must close is that the operator /32 must become a RENDERED conf field — a runtime-only wg set is wiped by the agent's own self-heal/restart.


0. What "production ep0" this demo could and could not represent

The demo endpoint (felhom-hetzner) is a real public dual-stack cloud VM with a public IPv4, 443/UDP WG, hub-driven peersync, and a real PBS — architecturally identical to the intended production ep0. So forwarding, isolation, and the SSH-through-endpoint path are represented faithfully. What it does not represent: (a) a CGNAT customer line — the box here is felhom-pve on the operator's One-Hungary fixed-cable line (public IPv4 37.191.56.193, single-NAT), same caveat as the prior spike §5; the OOB path rides the same outbound-mapping mechanism, so this is low-risk but unproven on true 100.64/10. (b) A second real customer box — the "dummy customer" was a throwaway netns peer at the endpoint, which is sufficient for the topology-isolation negative but is not a full second agent. (c) The production ep0 AAAA is still mis-set per the standing OPERATOR follow-up (memory: "OPERATOR fix ep0 AAAA"); this spike used v4 throughout, consistent with the agent's v4-pin (doc 06 §4.2).


1. P0 — ground-truth discovery (no changes) — GATE PASSED

Endpoint wg0 (redacted; wg show prints only the public key — dump is BANNED, S1 rule):

interface: wg0   listening port: 443
peer: yV5hFFZF548r6kq2Nl/CiSyOtFb0xYnvEAjw1ofrTCg=   (= the box)
  endpoint: 37.191.56.193:54900   allowed ips: 10.77.0.2/32
  latest handshake: ~1.5 min ago   transfer: 5.51 GiB rx / 16.43 GiB tx
wg0.conf: [Interface] Address 10.77.0.1/24, ListenPort 443, MTU 1420 ; [Peer] box only
sysctl: net.ipv4.ip_forward = 0   net.ipv6.conf.all.forwarding = 0
nft: table inet filter { input: drop-policy, allows 22, 443/udp, 8007 iif wg0 only, icmp ;
                         forward: policy DROP (empty) }
PBS: proxmox-backup + proxmox-backup-proxy active

Box wg-felhom (redacted):

wg-felhom.conf (agent-rendered, "DO NOT EDIT"): Address 10.77.0.2/32, MTU 1280,
  Peer f3d1ZI7…=  Endpoint 167.233.158.164:443  AllowedIPs 10.77.0.1/32  PersistentKeepalive 25
unit wg-quick@wg-felhom: enabled, active (exited)   handshake fresh
agent marker /var/lib/felhom-agent/wg/registered.json:
  {"pubkey":"yV5hFFZF548r6kq2Nl/CiSyOtFb0xYnvEAjw1ofrTCg=","assigned_ip":"10.77.0.2/32","generation":5}
sshd: port 22, ListenAddress 0.0.0.0:22 + [::]:22 (NOT interface-scoped today)
pve-firewall: disabled ;  nft list tables: EMPTY (no host nft tables at baseline)

Reconcile model (source + journal):

  • Box→tunnel reconcile is CONTINUOUS, not generation-gated. wgtunnel.Loop.Run (internal/wgtunnel/loop.go) calls Manager.Apply on every tickwg_tunnel.interval_seconds = 60 (agent.json, confirmed) — AND on each desired-state nudge. The desired-state fetch is generation-gated (desired/syncer.go OnEnvelope, hub poll poll_seconds = 900), but the wireguard block is cached and re-applied every 60 s regardless. Consequence for probes: any hand-edit to wg-felhom.conf or a runtime wg set that diverges from the agent's rendered conf is at risk of being reconciled within ≤60 s. This is why the AllowedIPs widening (P2) used runtime-only wg set (which the agent does NOT touch — it only rewrites the conf and restarts on a conf-hash change) and why the self-heal in P5b wiped it (§P5).
  • Endpoint peer-list reconcile (hub/internal/wgsync/reconciler.go): declarative FULL-list push every 5 min (or on a mutation Trigger); drift (a manually-added/removed peer) is erased on the next push. Proven live at cleanup (§13).

Gate: endpoint is CC-SSH-accessible as root ✓ → proceed.


2. P1 — operator peer, endpoint-side — GO

Throwaway operator + dummy-customer keypairs generated on the endpoint (wg genkey, 0600, private keys never printed). Registered BOTH into the hub wg_peers registry (not hand-added to wg0) so they survive the 5-min reconciler — via sqlite3 /data/hub.db in the hub pod (stdin-piped SQL to avoid the base64 =/+ quoting trap):

operator  cdmN4U+fjR18zBk+SKoceJQyz9HgA9+hN8/FiKF1u0o=  10.77.0.250
dummy     tQ6eJC9y8pSKtam44DaSrrtYBZYUvZG2z1stS6PsaD4=  10.77.0.251

The hub reconciler pushed both to wg0 in ~150 s (next 5-min tick fired early on a warm loop).

Endpoint forwarding, enabled NARROWLYip_forward=1 (was 0) + a named, one-flush nft chain felhom_spike_oob hung off forward, ALL rules scoped iifname wg0 oifname wg0:

table inet filter { chain forward { … ; jump felhom_spike_oob } }
chain felhom_spike_oob {
  iifname "wg0" oifname "wg0" ct state established,related accept          # conntrack replies
  iifname "wg0" oifname "wg0" ip saddr 10.77.0.250 ip daddr 10.77.0.2 counter accept   # operator→box
  iifname "wg0" oifname "wg0" ip daddr 10.77.0.250 counter drop            # box→operator NEW: drop
  iifname "wg0" oifname "wg0" counter drop                                 # box↔box (any other): drop
}

The operator "laptop" is a netns client on the endpoint (spike-op netns, veth to a 192.168.99.0/30 link, wg-op interface, MTU 1280, AllowedIPs 10.77.0.0/24) — an acceptable stand-in; the crypto+routing path through wg0 is identical to a remote laptop. GO: operator peer handshakes with wg0 and pings the endpoint 10.77.0.1 at 0.2 ms (on-box netns).


3. P2 — box-side AllowedIPs widening + SSH-over-tunnel + PBS — GO

The minimal box change, runtime-only (nothing persists past a wg-quick restart — see P5):

wg set wg-felhom peer f3d1ZI7…= allowed-ips 10.77.0.1/32,10.77.0.250/32
ip route add 10.77.0.250/32 dev wg-felhom

Exactly one added /32 (+ the return route). wg-felhom.conf mtime UNCHANGED (2026-07-04) — the agent's rendered conf was not touched.

From the operator netns, through the endpoint's forwarding:

ping 10.77.0.2  → 5/5, avg 24.7 ms, TTL 63  (TTL 63 = exactly one forwarded hop, i.e. via ep0)
ssh root@10.77.0.2 "hostname; pct list; uptime"  → interactive session in 1.205 s:
    demo-felhom ;  9201 running demo-felhom (+9001/9100/9999) ;  load avg 0.43

Interactive SSH over the tunnel works end-to-end.

PBS-path check (mandatory, immediately after the box change):

wg show wg-felhom latest-handshakes → age 73 s (fresh, advancing)
pvesm status --storage felhom-offsite → active   (PVE reaches PBS at 10.77.0.1:8007 over the tunnel)
openssl s_client -connect 10.77.0.1:8007 → CONNECTED, Certificate chain present

PBS unaffected = GO condition met. Adding the operator reach did not perturb the backup path.


4. P3 — isolation negative probes (§4.5 must survive) — GO (both negatives, counter-proven)

Negative 1 — dummy customer → box: DROPPED. The dummy peer handshakes with the endpoint (0.2 ms to 10.77.0.1 — it IS a registered peer) but:

spike-cust → ping 10.77.0.2  → 3/3 packets LOST (100%)
spike-cust → tcp/22 10.77.0.2 → REFUSED/TIMEOUT
felhom_spike_oob box↔box drop counter: 0 → 8 packets / 552 bytes   (the drops are real, not a route miss)

Negative 2 — box → operator: NEW connections DROPPED, replies flow. With an operator→box SSH session OPEN (so ESTABLISHED replies were flowing), from the box:

felhom-pve → ping 10.77.0.250        → 3/3 LOST
felhom-pve → tcp/22 & tcp/80 10.77.0.250 → REFUSED/TIMEOUT
felhom_spike_oob box→operator drop counter: 0 → 13 packets / 852 bytes
  (while the operator→box accept counter incremented normally — conntrack replies unaffected)

Both hold with packet-level evidence (nft counters), not merely "command timed out". §4.5 isolation survives the operator-peer addition: a box still cannot reach another box, and a box cannot initiate to the operator.


5. P4 — sshd scoping on the box (belt rule) — GO (all three)

A named, one-flush nft table on felhom-pve (baseline had NO host nft tables), independent hook so LAN and other interfaces are untouched:

table inet felhom_spike_oob {
  chain input { type filter hook input priority -10; policy accept;
    iifname "wg-felhom" tcp dport 22 ip saddr 10.77.0.250 counter accept   # operator only
    iifname "wg-felhom" tcp dport 22 counter drop                          # any other tunnel src
  } }

Verified:

  1. LAN SSH from CC — unaffected (the safety line): the rule matches iifname wg-felhom only; LAN commands kept running throughout. ✓
  2. Operator SSH over the tunnel — still works (accept counter 27 pkts). ✓
  3. Defense-in-depth — dummy refused at the HOST even if the endpoint belt failed: I temporarily simulated an endpoint-belt failure (added an endpoint forward-accept for dummy→box AND widened the box AllowedIPs to include 10.77.0.251/32). Result: dummy→box ping now passes (endpoint forwards it, 27 ms) but dummy→box SSH is dropped at the host (host belt drop counter 5) — Connection timed out. The two layers are independent; the host belt holds when the endpoint layer is bypassed. The temporary widening was reverted immediately after. ✓

6. P5 — mutual repair (the crown probe)

6a. Agent dead, wg alive — GO

systemctl stop felhom-agent (via LAN, simulating no other way in). Then from the operator peer over the tunnel: SSH into the box (agent DOWN → OOB does not depend on the agent), and systemctl reset-failed felhom-agent; systemctl start felhom-agent from that session → active, felhom-agent 0.70.0. OOB works with the agent down AND repairs it.

6b. wg dead, agent alive — GO, ~15 s, and this is THE central TASK finding

systemctl stop wg-quick@wg-felhom (interface gone). Operator SSH over the tunnel now Connection timed out (OOB down, confirmed). Then, watching the endpoint's view of the box peer handshake with zero LAN intervention:

box re-handshake observed at +15 s   (agent log: "wgtunnel: service inactive with current conf —
                                       re-enabling (self-heal)")

The agent's reconcile loop already covers a stopped unitManager.ensureTunnelLocked's self-heal branch (internal/wgtunnel/manager.go:394-402): conf hash current + isActive() false → systemctl enable --now. No wg_repair-shaped desired action is needed; the "wg dead, agent alive" repair path exists today and fires within one 60 s tick (here 15 s). This closes O-B (mutual repair via the desired-state channel) — with a caveat that becomes the TASK's main work:

THE GAP. The self-heal rebuilds wg-felhom purely from the agent-rendered conf, whose AllowedIPs is derived solely from pbs_tunnel_ip/32 (renderConf, manager.go:239: AllowedIPs = <pbsIP>/32). After the self-heal, the box's live AllowedIPs was back to only 10.77.0.1/32 — the operator /32 I added at runtime was wiped, and OOB SSH failed until I re-added it. So OOB survival across the agent's own self-heal/restart/reboot REQUIRES the operator peer's /32 to be part of what the agent RENDERS into the conf — a runtime wg set is not durable. (Re-adding 10.77.0.250/32 at runtime immediately restored OOB, confirming the mechanism; it is only the durability that needs the rendered-conf change.)

6c. Failure-domain sweep with an open OOB session — GO

Over a single open operator→box tunnel session: attempted cloudflared stop (it is NOT a guest systemd unit — it runs as a Docker container in the controller stack, so the unit stop no-op'd; noted), then pct stop 9201 (whole guest down). The session stayed alive (ALIVE_2, handshake fresh) — the tunnel terminates host-side on felhom-pve, so guest death cannot touch it — and pct start 9201 from the SAME session brought the guest back. Post-check: felhom-controller Up (healthy), cloudflared/nextcloud containers Up (healthy). OOB survives the failure domains it exists for (guest down; and by construction cloudflared/controller down, since OOB is host-side and independent of the guest's own edge).


7. P6 — usability + environment (record-only)

  • Interactive feel over MTU 1280: top -bn1, pct exec 9201 -- … (into the guest), and a 20 MB dd | wc -c bulk pull all ran cleanly over the tunnel at ~24 ms RTT — no black-holing, no stalls (MTU 1280 is the doc 06 §4.3 floor; interactive OOB is well within it).
  • Endpoint load during the probes: load average 0.10 — WG forwarding + a shell session is negligible on a CX23.
  • Active port: 443/UDP (wg show wg0 listen-port → 443). Relevant to CGNAT: this is the friendly-port path already; the box's NAT source was 37.191.56.193:47922 (single-NAT public v4, not CGNAT — the standing caveat).
  • Keepalive: PersistentKeepalive 25 on the box; handshakes stayed fresh (<125 s) throughout.

8. Findings that must shape the TASK

  1. Render the operator /32 into the agent conf — the one required agent change (§6b). OOB reach must survive the agent's self-heal/restart/reboot, so the operator peer's tunnel /32 must be a field the agent puts in wg-felhom.conf AllowedIPs (today hard-derived from pbs_tunnel_ip alone). Shape options for the TASK to pick: (a) extend the hub wireguard desired-state block with an optional oob_peer_ip (or a small extra_allowed_ips list) and have renderConf append it; (b) a separate [Peer]-less allowed-ips widening is NOT enough (WG ties AllowedIPs to the server peer here, and all tunnel traffic transits the one endpoint peer, so the operator /32 just joins the existing peer's AllowedIPs). Keep the derivation deterministic and validated (the existing netip.ParsePrefix//32 checks). Nothing else in the agent needs to change for O-A/O-B — reconcile continuity + self-heal already do the repair.

  2. Mutual repair (O-B) is already built — do NOT add a wg_repair action. The 60 s reconcile tick + ensureTunnelLocked self-heal restored a stopped tunnel in 15 s unaided (§6b). The TASK should rely on this, and at most add the tunnel-health→alert wiring already slated for S6 (doc 06 §4.6 stanza exists). Endpoint-side drift repair likewise already works (5-min declarative full-list push erased the spike peers from wg0.conf at cleanup, §13).

  3. The exact nft rulesets that worked (copy-paste ready). Endpoint forward chain and the host belt are in §2 and §5 verbatim. Both are named, one-flush tables. Production shape:

    • Endpoint: ip_forward=1 + a felhom_oob forward chain: ct established,related accept; saddr <operator/32> daddr <box/32> accept per (operator,box) pair; a catch-all iifname wg0 oifname wg0 … drop to preserve §4.5 for every non-OOB pair. Note this makes the endpoint's forward posture per-pair allow-list instead of doc 06 §4.5's blanket "forwarding OFF" — the TASK must state that the OOB feature deliberately turns forwarding ON but gates it to operator→box pairs only, and that the box↔box drop is now an explicit rule, not an absent capability. Peersync (felhom-peersync.sh) would need to learn these forward rules, OR they live in the endpoint's static nftables and only the peer list stays hub-driven (simpler; recommended — the operator set is small and stable).
    • Box (belt, defense-in-depth): a felhom_oob inet table, input hook, iifname wg-felhom tcp dport 22 ip saddr <operator/32> accept then iifname wg-felhom tcp dport 22 drop. Scope strictly to wg-felhom; never touch LAN sshd (the safety line). This belongs in the agent's managed surface if the box is to enforce it (a new narrow grant), OR ship as host-install static config; the spike proved the rule works, the TASK picks the ownership.
  4. AllowedIPs method: runtime wg set is NOT durable (finding 1); the durable path is the rendered conf. For the endpoint side, the operator peer is a normal hub wg_peers row (survives the reconciler) — that part needs no new mechanism, only that the operator peer is not a customer host (host_id empty, as the spike used) and is excluded from any customer-facing accounting.

  5. sshd scoping shape. Today box sshd listens on 0.0.0.0:22 + [::]:22 (P0). The spike scoped access with an nft belt on iifname wg-felhom, NOT by changing ListenAddress — deliberate: an interface-bound ListenAddress would fight the LAN safety line and the tunnel's late bring-up. The TASK should keep sshd listening broadly and gate at the packet layer (belt rule), exactly as §5. The endpoint forward rule is the primary gate; the box belt is defense-in-depth (both proven independent in §5).

  6. Measured latencies (production budget input). Operator→box forwarded RTT ~24 ms (One Hungary ↔ Hetzner-DE); interactive SSH incl. pct list 1.2 s wall; self-heal repair ~15 s; endpoint peer-push convergence ~150 s (≤5 min bound). All comfortably interactive.

  7. Reconcile continuity constraint (from P0, load-bearing for any future OOB probe/op). The box applies its conf every 60 s. Any OOB mechanism that mutates box wg state MUST go through the agent's rendered conf (finding 1), never a side-channel wg set/hand-edit — those are silently reverted within a tick. This is a feature (drift-repair) to build with, not around.

  8. Endpoint-access + production prerequisites discovered. (a) The demo endpoint is a real public VM, so O-A is faithfully represented; production ep0 needs the AAAA fixed (standing OPERATOR item) though the v4-pin means OOB rides v4 regardless. (b) cloudflared in the guest is a Docker container, not a systemd unit — any runbook step that "stops cloudflared" must docker stop it, not systemctl. (c) The hub wg_peers insert path for a NON-host operator peer exists in the store (AddWGPeer with empty host_id — the S1 unbound-admin-row shape); a production OOB feature can reuse it or add an explicit operator-peer admin endpoint.


9. What this spike did NOT prove (honest ledger)

  • True CGNAT OOB traversal — the box is single-NAT public-v4 (same caveat as the prior spike §5). OOB rides the same outbound WG mapping as the backup path, so low-risk, but unproven on 100.64/10.
  • A second real agent box — the dummy customer was a netns peer (sufficient for the topology isolation negative; not a full agent). The box↔box drop is proven by topology + counters, which is the load-bearing property.
  • The rendered-conf operator-/32 change itself — finding 1 is the diagnosis (runtime wg set wiped by self-heal, re-add restored OOB); the agent code change to render it is TASK work, not shipped here.
  • Endpoint peersync owning the forward rules — the spike put forward rules in ad-hoc nft; the TASK decides static-vs-hub-driven (finding 3 recommends static forward rules + hub-driven peer list).
  • Long-horizon hold of an idle OOB session (minutes-scale only here; the backup-path 32-min soak is in doc 06 §7).

10. Cleanup assertion (§13 of the plan) — VERIFIED

Hub: both spike wg_peers rows deleted (WHERE note LIKE 'spike-oob-%'); registry back to the single real peer demo-felhom-01 / 10.77.0.2. Endpoint felhom-hetzner: operator+dummy peers removed from live wg0; the 5-min reconciler then rewrote wg0.conf to the box-only peer (drift-repair confirmed live — the persisted conf self-healed); felhom_spike_oob forward chain + jump rule deleted; ip_forward restored to 0 (the P0 value); spike-op/spike-cust netns + veths deleted; throwaway WG + SSH keys shred -u'd, /root/spike-oob removed. Final: nft forward chain is the empty drop-policy chain (P0), wg0 lists only the box, ip_forward=0. Box felhom-pve: runtime AllowedIPs back to 10.77.0.1/32 (operator /32 removed) + return route deleted; felhom_spike_oob host table deleted (nft list tables empty, as P0); the spike-oob-operator line removed from /root/.ssh/authorized_keys (grep count 0); /tmp/wgdown_t0 removed. Final baseline re-verified: felhom-agent active, wg-quick@wg-felhom active + handshake fresh (age 92 s), pvesm status felhom-offsite active, guest 9201 running + controller healthy, no host nft tables, LAN SSH fine. No repo/hub/agent/manifest production change — this commit is docs-only.


11. Method bar honored

wg show <if> dump NEVER run (private-key leak ban, S1); only latest-handshakes + redacted conf reads. Every mutation was a named one-flush nft table or a runtime-only wg set, inventoried as made. PBS-path re-checked after the box wg change (P2). Negative probes carry nft-counter evidence, not just timeouts. Surprises recorded as findings (the self-heal wipe of the runtime /32 — finding 1 — is the headline example).