docs(audit): OOB-over-WG operator-peer spike (2026-07-05) — GO
Validates operator-inbound access over the existing offsite WG arc (doc 06): operator peer forwarded operator->box only, box sshd gated to the operator /32, §4.5 box<->box isolation intact (both negatives counter-proven), mutual repair real (agent self-healed a stopped tunnel in ~15s unaided). One TASK-shaping gap: the operator /32 must be a RENDERED conf field — a runtime `wg set` is wiped by the agent's own self-heal. All live-arc changes reverted + baseline re-verified. Docs-only; no code/hub/agent/manifest change. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PSK5g6qYLknKj8u3QAFEr6
This commit is contained in:
@@ -0,0 +1,366 @@
|
||||
# SPIKE — OOB management over the existing WG arc (operator peer, isolation, mutual repair) — 2026-07-05
|
||||
|
||||
> **STATUS: COMPLETE — verdict below.** Empirically validates the unproven mechanisms of the OOB
|
||||
> (out-of-band operator access) design BEFORE the production spec. No production code shipped; no
|
||||
> agent binary change; every config change on the live demo arc reverted at cleanup (§13, verified).
|
||||
> The design under test: an **operator peer on the existing hub-and-spoke WG arc** (doc 06),
|
||||
> endpoint-forwarded **operator→box only**, box sshd reachable solely from the operator tunnel IP,
|
||||
> while §4.5 box↔box isolation stays intact — and the hub desired-state channel as the tunnel's
|
||||
> **repair path**.
|
||||
|
||||
**Class:** SPIKE (empirical; no product code). **Repos:** felhom.eu (this doc only); felhom-agent
|
||||
read-only for grounding (`internal/wgtunnel/{manager,loop}.go`, `internal/desired/syncer.go`,
|
||||
`configs/felhom-agent.sudoers`); felhom-hub read-only (`internal/wgsync/reconciler.go`,
|
||||
`internal/store/wg.go`).
|
||||
|
||||
**Probe ends (the LIVE demo arc, ground-truthed in P0):**
|
||||
- **Endpoint** = `felhom-hetzner` (Hetzner CX23, Debian 13, `167.233.158.164`, DNS `ep0.felhom.eu`),
|
||||
the **dev** offsite endpoint: WG server `wg0` on **443/UDP**, tunnel `10.77.0.0/24`, endpoint
|
||||
`10.77.0.1`, PBS `felhom-offsite`. Peer list is hub-driven (`internal/wgsync` SSH push →
|
||||
`felhom-peersync`). CC-SSH-reachable as root ✓ (the P0 access gate passed).
|
||||
- **Box** = `felhom-pve` / node `demo-felhom` (`192.168.0.162`), agent **v0.70.0**, `wg-felhom`
|
||||
client `10.77.0.2/32`, tunnel to `ep0:443`. CC's LAN SSH to it was the safety line throughout.
|
||||
|
||||
**Verdict (one line):** **GO — an operator peer on the existing arc reaches a box's sshd through the
|
||||
endpoint with a single added `/32`, at ~24 ms interactive latency, WITHOUT breaking §4.5 isolation
|
||||
(both negatives held with packet-level evidence); mutual repair is REAL and already built (the
|
||||
agent self-healed a stopped tunnel in ~15 s with zero hub/LAN involvement, and OOB works with the
|
||||
agent dead); the one load-bearing gap the TASK must close is that the operator `/32` must become a
|
||||
RENDERED conf field — a runtime-only `wg set` is wiped by the agent's own self-heal/restart.**
|
||||
|
||||
---
|
||||
|
||||
## 0. What "production ep0" this demo could and could not represent
|
||||
|
||||
The demo endpoint (`felhom-hetzner`) **is** a real public dual-stack cloud VM with a public IPv4,
|
||||
443/UDP WG, hub-driven peersync, and a real PBS — architecturally identical to the intended
|
||||
production `ep0`. So forwarding, isolation, and the SSH-through-endpoint path are represented
|
||||
faithfully. What it does **not** represent: (a) a **CGNAT** customer line — the box here is
|
||||
`felhom-pve` on the operator's One-Hungary fixed-cable line (public IPv4 `37.191.56.193`,
|
||||
single-NAT), same caveat as the prior spike §5; the OOB path rides the same outbound-mapping
|
||||
mechanism, so this is low-risk but unproven on true `100.64/10`. (b) A **second real customer box**
|
||||
— the "dummy customer" was a throwaway netns peer at the endpoint, which is sufficient for the
|
||||
topology-isolation negative but is not a full second agent. (c) The production `ep0` **AAAA** is
|
||||
still mis-set per the standing OPERATOR follow-up (memory: "OPERATOR fix ep0 AAAA"); this spike
|
||||
used v4 throughout, consistent with the agent's v4-pin (doc 06 §4.2).
|
||||
|
||||
---
|
||||
|
||||
## 1. P0 — ground-truth discovery (no changes) — **GATE PASSED**
|
||||
|
||||
**Endpoint `wg0` (redacted; `wg show` prints only the public key — `dump` is BANNED, S1 rule):**
|
||||
```
|
||||
interface: wg0 listening port: 443
|
||||
peer: yV5hFFZF548r6kq2Nl/CiSyOtFb0xYnvEAjw1ofrTCg= (= the box)
|
||||
endpoint: 37.191.56.193:54900 allowed ips: 10.77.0.2/32
|
||||
latest handshake: ~1.5 min ago transfer: 5.51 GiB rx / 16.43 GiB tx
|
||||
wg0.conf: [Interface] Address 10.77.0.1/24, ListenPort 443, MTU 1420 ; [Peer] box only
|
||||
sysctl: net.ipv4.ip_forward = 0 net.ipv6.conf.all.forwarding = 0
|
||||
nft: table inet filter { input: drop-policy, allows 22, 443/udp, 8007 iif wg0 only, icmp ;
|
||||
forward: policy DROP (empty) }
|
||||
PBS: proxmox-backup + proxmox-backup-proxy active
|
||||
```
|
||||
**Box `wg-felhom` (redacted):**
|
||||
```
|
||||
wg-felhom.conf (agent-rendered, "DO NOT EDIT"): Address 10.77.0.2/32, MTU 1280,
|
||||
Peer f3d1ZI7…= Endpoint 167.233.158.164:443 AllowedIPs 10.77.0.1/32 PersistentKeepalive 25
|
||||
unit wg-quick@wg-felhom: enabled, active (exited) handshake fresh
|
||||
agent marker /var/lib/felhom-agent/wg/registered.json:
|
||||
{"pubkey":"yV5hFFZF548r6kq2Nl/CiSyOtFb0xYnvEAjw1ofrTCg=","assigned_ip":"10.77.0.2/32","generation":5}
|
||||
sshd: port 22, ListenAddress 0.0.0.0:22 + [::]:22 (NOT interface-scoped today)
|
||||
pve-firewall: disabled ; nft list tables: EMPTY (no host nft tables at baseline)
|
||||
```
|
||||
**Reconcile model (source + journal):**
|
||||
- **Box→tunnel reconcile is CONTINUOUS, not generation-gated.** `wgtunnel.Loop.Run`
|
||||
(`internal/wgtunnel/loop.go`) calls `Manager.Apply` **on every tick** — `wg_tunnel.interval_seconds
|
||||
= 60` (agent.json, confirmed) — AND on each desired-state nudge. The desired-state *fetch* is
|
||||
generation-gated (`desired/syncer.go` `OnEnvelope`, hub poll `poll_seconds = 900`), but the
|
||||
wireguard block is cached and re-applied every 60 s regardless. **Consequence for probes:** any
|
||||
hand-edit to `wg-felhom.conf` or a runtime `wg set` that diverges from the agent's rendered conf
|
||||
is at risk of being reconciled within ≤60 s. This is why the AllowedIPs widening (P2) used
|
||||
runtime-only `wg set` (which the agent does NOT touch — it only rewrites the *conf* and restarts
|
||||
on a conf-hash change) and why the self-heal in P5b wiped it (§P5).
|
||||
- **Endpoint peer-list reconcile** (`hub/internal/wgsync/reconciler.go`): declarative FULL-list push
|
||||
every **5 min** (or on a mutation Trigger); drift (a manually-added/removed peer) is erased on the
|
||||
next push. Proven live at cleanup (§13).
|
||||
|
||||
**Gate:** endpoint is CC-SSH-accessible as root ✓ → proceed.
|
||||
|
||||
---
|
||||
|
||||
## 2. P1 — operator peer, endpoint-side — **GO**
|
||||
|
||||
Throwaway operator + dummy-customer keypairs generated on the endpoint (`wg genkey`, 0600,
|
||||
private keys never printed). Registered BOTH into the **hub** `wg_peers` registry (not hand-added
|
||||
to `wg0`) so they survive the 5-min reconciler — via `sqlite3 /data/hub.db` in the hub pod
|
||||
(stdin-piped SQL to avoid the base64 `=`/`+` quoting trap):
|
||||
```
|
||||
operator cdmN4U+fjR18zBk+SKoceJQyz9HgA9+hN8/FiKF1u0o= 10.77.0.250
|
||||
dummy tQ6eJC9y8pSKtam44DaSrrtYBZYUvZG2z1stS6PsaD4= 10.77.0.251
|
||||
```
|
||||
The hub reconciler pushed both to `wg0` in **~150 s** (next 5-min tick fired early on a warm loop).
|
||||
|
||||
**Endpoint forwarding, enabled NARROWLY** — `ip_forward=1` (was 0) + a named, one-flush nft chain
|
||||
`felhom_spike_oob` hung off `forward`, ALL rules scoped `iifname wg0 oifname wg0`:
|
||||
```
|
||||
table inet filter { chain forward { … ; jump felhom_spike_oob } }
|
||||
chain felhom_spike_oob {
|
||||
iifname "wg0" oifname "wg0" ct state established,related accept # conntrack replies
|
||||
iifname "wg0" oifname "wg0" ip saddr 10.77.0.250 ip daddr 10.77.0.2 counter accept # operator→box
|
||||
iifname "wg0" oifname "wg0" ip daddr 10.77.0.250 counter drop # box→operator NEW: drop
|
||||
iifname "wg0" oifname "wg0" counter drop # box↔box (any other): drop
|
||||
}
|
||||
```
|
||||
The operator "laptop" is a **netns client** on the endpoint (`spike-op` netns, veth to a
|
||||
`192.168.99.0/30` link, `wg-op` interface, MTU 1280, `AllowedIPs 10.77.0.0/24`) — an acceptable
|
||||
stand-in; the crypto+routing path through `wg0` is identical to a remote laptop. **GO:** operator
|
||||
peer handshakes with `wg0` and pings the endpoint `10.77.0.1` at **0.2 ms** (on-box netns).
|
||||
|
||||
---
|
||||
|
||||
## 3. P2 — box-side AllowedIPs widening + SSH-over-tunnel + PBS — **GO**
|
||||
|
||||
**The minimal box change, runtime-only** (nothing persists past a wg-quick restart — see P5):
|
||||
```
|
||||
wg set wg-felhom peer f3d1ZI7…= allowed-ips 10.77.0.1/32,10.77.0.250/32
|
||||
ip route add 10.77.0.250/32 dev wg-felhom
|
||||
```
|
||||
Exactly one added `/32` (+ the return route). `wg-felhom.conf` mtime UNCHANGED (2026-07-04) — the
|
||||
agent's rendered conf was not touched.
|
||||
|
||||
From the operator netns, through the endpoint's forwarding:
|
||||
```
|
||||
ping 10.77.0.2 → 5/5, avg 24.7 ms, TTL 63 (TTL 63 = exactly one forwarded hop, i.e. via ep0)
|
||||
ssh root@10.77.0.2 "hostname; pct list; uptime" → interactive session in 1.205 s:
|
||||
demo-felhom ; 9201 running demo-felhom (+9001/9100/9999) ; load avg 0.43
|
||||
```
|
||||
Interactive SSH over the tunnel works end-to-end.
|
||||
|
||||
**PBS-path check (mandatory, immediately after the box change):**
|
||||
```
|
||||
wg show wg-felhom latest-handshakes → age 73 s (fresh, advancing)
|
||||
pvesm status --storage felhom-offsite → active (PVE reaches PBS at 10.77.0.1:8007 over the tunnel)
|
||||
openssl s_client -connect 10.77.0.1:8007 → CONNECTED, Certificate chain present
|
||||
```
|
||||
**PBS unaffected = GO condition met.** Adding the operator reach did not perturb the backup path.
|
||||
|
||||
---
|
||||
|
||||
## 4. P3 — isolation negative probes (§4.5 must survive) — **GO (both negatives, counter-proven)**
|
||||
|
||||
**Negative 1 — dummy customer → box: DROPPED.** The dummy peer handshakes with the endpoint (0.2 ms
|
||||
to `10.77.0.1` — it IS a registered peer) but:
|
||||
```
|
||||
spike-cust → ping 10.77.0.2 → 3/3 packets LOST (100%)
|
||||
spike-cust → tcp/22 10.77.0.2 → REFUSED/TIMEOUT
|
||||
felhom_spike_oob box↔box drop counter: 0 → 8 packets / 552 bytes (the drops are real, not a route miss)
|
||||
```
|
||||
**Negative 2 — box → operator: NEW connections DROPPED, replies flow.** With an operator→box SSH
|
||||
session OPEN (so ESTABLISHED replies were flowing), from the box:
|
||||
```
|
||||
felhom-pve → ping 10.77.0.250 → 3/3 LOST
|
||||
felhom-pve → tcp/22 & tcp/80 10.77.0.250 → REFUSED/TIMEOUT
|
||||
felhom_spike_oob box→operator drop counter: 0 → 13 packets / 852 bytes
|
||||
(while the operator→box accept counter incremented normally — conntrack replies unaffected)
|
||||
```
|
||||
Both hold with **packet-level evidence** (nft counters), not merely "command timed out". §4.5
|
||||
isolation survives the operator-peer addition: a box still cannot reach another box, and a box
|
||||
cannot initiate to the operator.
|
||||
|
||||
---
|
||||
|
||||
## 5. P4 — sshd scoping on the box (belt rule) — **GO (all three)**
|
||||
|
||||
A named, one-flush nft table on `felhom-pve` (baseline had NO host nft tables), independent hook
|
||||
so LAN and other interfaces are untouched:
|
||||
```
|
||||
table inet felhom_spike_oob {
|
||||
chain input { type filter hook input priority -10; policy accept;
|
||||
iifname "wg-felhom" tcp dport 22 ip saddr 10.77.0.250 counter accept # operator only
|
||||
iifname "wg-felhom" tcp dport 22 counter drop # any other tunnel src
|
||||
} }
|
||||
```
|
||||
Verified:
|
||||
1. **LAN SSH from CC — unaffected** (the safety line): the rule matches `iifname wg-felhom` only;
|
||||
LAN commands kept running throughout. ✓
|
||||
2. **Operator SSH over the tunnel — still works** (accept counter 27 pkts). ✓
|
||||
3. **Defense-in-depth — dummy refused at the HOST even if the endpoint belt failed:** I temporarily
|
||||
simulated an endpoint-belt failure (added an endpoint forward-accept for dummy→box AND widened
|
||||
the box AllowedIPs to include `10.77.0.251/32`). Result: dummy→box **ping now passes** (endpoint
|
||||
forwards it, 27 ms) but dummy→box **SSH is dropped at the host** (host belt drop counter 5) —
|
||||
`Connection timed out`. The two layers are independent; the host belt holds when the endpoint
|
||||
layer is bypassed. The temporary widening was reverted immediately after. ✓
|
||||
|
||||
---
|
||||
|
||||
## 6. P5 — mutual repair (the crown probe)
|
||||
|
||||
### 6a. Agent dead, wg alive — **GO**
|
||||
`systemctl stop felhom-agent` (via LAN, simulating no other way in). Then from the operator peer
|
||||
**over the tunnel**: SSH into the box (agent DOWN → OOB does not depend on the agent), and
|
||||
`systemctl reset-failed felhom-agent; systemctl start felhom-agent` from that session →
|
||||
`active`, `felhom-agent 0.70.0`. **OOB works with the agent down AND repairs it.**
|
||||
|
||||
### 6b. wg dead, agent alive — **GO, ~15 s, and this is THE central TASK finding**
|
||||
`systemctl stop wg-quick@wg-felhom` (interface gone). Operator SSH over the tunnel now
|
||||
`Connection timed out` (OOB down, confirmed). Then, watching the endpoint's view of the box peer
|
||||
handshake with **zero LAN intervention**:
|
||||
```
|
||||
box re-handshake observed at +15 s (agent log: "wgtunnel: service inactive with current conf —
|
||||
re-enabling (self-heal)")
|
||||
```
|
||||
The agent's reconcile loop **already covers a stopped unit** — `Manager.ensureTunnelLocked`'s
|
||||
self-heal branch (`internal/wgtunnel/manager.go:394-402`): conf hash current + `isActive()` false →
|
||||
`systemctl enable --now`. No `wg_repair`-shaped desired action is needed; the "wg dead, agent alive"
|
||||
repair path exists today and fires within one 60 s tick (here 15 s). **This closes O-B (mutual
|
||||
repair via the desired-state channel) — with a caveat that becomes the TASK's main work:**
|
||||
|
||||
> **THE GAP.** The self-heal rebuilds `wg-felhom` **purely from the agent-rendered conf**, whose
|
||||
> `AllowedIPs` is derived solely from `pbs_tunnel_ip/32` (`renderConf`, `manager.go:239`:
|
||||
> `AllowedIPs = <pbsIP>/32`). After the self-heal, the box's live AllowedIPs was back to **only
|
||||
> `10.77.0.1/32`** — the operator `/32` I added at runtime was **wiped**, and OOB SSH failed until I
|
||||
> re-added it. So **OOB survival across the agent's own self-heal/restart/reboot REQUIRES the
|
||||
> operator peer's `/32` to be part of what the agent RENDERS into the conf** — a runtime `wg set` is
|
||||
> not durable. (Re-adding `10.77.0.250/32` at runtime immediately restored OOB, confirming the
|
||||
> mechanism; it is only the *durability* that needs the rendered-conf change.)
|
||||
|
||||
### 6c. Failure-domain sweep with an open OOB session — **GO**
|
||||
Over a single open operator→box tunnel session: attempted `cloudflared` stop (it is NOT a guest
|
||||
systemd unit — it runs as a **Docker container** in the controller stack, so the unit stop no-op'd;
|
||||
noted), then `pct stop 9201` (whole guest down). **The session stayed alive** (`ALIVE_2`, handshake
|
||||
fresh) — the tunnel terminates **host-side** on `felhom-pve`, so guest death cannot touch it — and
|
||||
`pct start 9201` from the SAME session brought the guest back. Post-check: `felhom-controller Up
|
||||
(healthy)`, `cloudflared`/`nextcloud` containers `Up (healthy)`. OOB survives the failure domains it
|
||||
exists for (guest down; and by construction cloudflared/controller down, since OOB is host-side and
|
||||
independent of the guest's own edge).
|
||||
|
||||
---
|
||||
|
||||
## 7. P6 — usability + environment (record-only)
|
||||
|
||||
- **Interactive feel over MTU 1280:** `top -bn1`, `pct exec 9201 -- …` (into the guest), and a
|
||||
20 MB `dd | wc -c` bulk pull all ran cleanly over the tunnel at ~24 ms RTT — no black-holing, no
|
||||
stalls (MTU 1280 is the doc 06 §4.3 floor; interactive OOB is well within it).
|
||||
- **Endpoint load** during the probes: `load average 0.10` — WG forwarding + a shell session is
|
||||
negligible on a CX23.
|
||||
- **Active port: 443/UDP** (`wg show wg0 listen-port → 443`). Relevant to CGNAT: this is the
|
||||
friendly-port path already; the box's NAT source was `37.191.56.193:47922` (single-NAT public v4,
|
||||
not CGNAT — the standing caveat).
|
||||
- **Keepalive:** `PersistentKeepalive 25` on the box; handshakes stayed fresh (<125 s) throughout.
|
||||
|
||||
---
|
||||
|
||||
## 8. Findings that must shape the TASK
|
||||
|
||||
1. **Render the operator `/32` into the agent conf — the one required agent change (§6b).** OOB
|
||||
reach must survive the agent's self-heal/restart/reboot, so the operator peer's tunnel `/32`
|
||||
must be a field the agent puts in `wg-felhom.conf` `AllowedIPs` (today hard-derived from
|
||||
`pbs_tunnel_ip` alone). Shape options for the TASK to pick: (a) extend the hub `wireguard`
|
||||
desired-state block with an optional `oob_peer_ip` (or a small `extra_allowed_ips` list) and have
|
||||
`renderConf` append it; (b) a separate `[Peer]`-less allowed-ips widening is NOT enough (WG ties
|
||||
AllowedIPs to the *server* peer here, and all tunnel traffic transits the one endpoint peer, so
|
||||
the operator `/32` just joins the existing peer's AllowedIPs). Keep the derivation deterministic
|
||||
and validated (the existing `netip.ParsePrefix`/`/32` checks). **Nothing else in the agent needs
|
||||
to change for O-A/O-B** — reconcile continuity + self-heal already do the repair.
|
||||
|
||||
2. **Mutual repair (O-B) is already built — do NOT add a `wg_repair` action.** The 60 s reconcile
|
||||
tick + `ensureTunnelLocked` self-heal restored a stopped tunnel in 15 s unaided (§6b). The TASK
|
||||
should *rely on* this, and at most add the tunnel-health→alert wiring already slated for S6 (doc
|
||||
06 §4.6 stanza exists). Endpoint-side drift repair likewise already works (5-min declarative
|
||||
full-list push erased the spike peers from `wg0.conf` at cleanup, §13).
|
||||
|
||||
3. **The exact nft rulesets that worked (copy-paste ready).** Endpoint forward chain and the
|
||||
host belt are in §2 and §5 verbatim. Both are named, one-flush tables. Production shape:
|
||||
- **Endpoint:** `ip_forward=1` + a `felhom_oob` forward chain: `ct established,related accept`;
|
||||
`saddr <operator/32> daddr <box/32> accept` **per (operator,box) pair**; a catch-all
|
||||
`iifname wg0 oifname wg0 … drop` to preserve §4.5 for every non-OOB pair. Note this makes the
|
||||
endpoint's forward posture **per-pair allow-list** instead of doc 06 §4.5's blanket
|
||||
"forwarding OFF" — the TASK must state that the OOB feature deliberately turns forwarding ON
|
||||
but gates it to operator→box pairs only, and that the box↔box drop is now an explicit rule, not
|
||||
an absent capability. Peersync (`felhom-peersync.sh`) would need to learn these forward rules,
|
||||
OR they live in the endpoint's static nftables and only the peer *list* stays hub-driven
|
||||
(simpler; recommended — the operator set is small and stable).
|
||||
- **Box (belt, defense-in-depth):** a `felhom_oob` inet table, `input` hook, `iifname wg-felhom
|
||||
tcp dport 22 ip saddr <operator/32> accept` then `iifname wg-felhom tcp dport 22 drop`. Scope
|
||||
strictly to `wg-felhom`; never touch LAN sshd (the safety line). This belongs in the agent's
|
||||
managed surface if the box is to enforce it (a new narrow grant), OR ship as host-install
|
||||
static config; the spike proved the *rule* works, the TASK picks the ownership.
|
||||
|
||||
4. **AllowedIPs method: runtime `wg set` is NOT durable (finding 1); the durable path is the
|
||||
rendered conf.** For the endpoint side, the operator peer is a normal hub `wg_peers` row
|
||||
(survives the reconciler) — that part needs no new mechanism, only that the operator peer is
|
||||
*not* a customer host (host_id empty, as the spike used) and is excluded from any customer-facing
|
||||
accounting.
|
||||
|
||||
5. **sshd scoping shape.** Today box sshd listens on `0.0.0.0:22 + [::]:22` (P0). The spike scoped
|
||||
access with an **nft belt on `iifname wg-felhom`**, NOT by changing `ListenAddress` — deliberate:
|
||||
an interface-bound `ListenAddress` would fight the LAN safety line and the tunnel's late bring-up.
|
||||
The TASK should keep sshd listening broadly and gate at the packet layer (belt rule), exactly as
|
||||
§5. The endpoint forward rule is the primary gate; the box belt is defense-in-depth (both proven
|
||||
independent in §5).
|
||||
|
||||
6. **Measured latencies (production budget input).** Operator→box forwarded RTT **~24 ms** (One
|
||||
Hungary ↔ Hetzner-DE); interactive SSH incl. `pct list` **1.2 s** wall; self-heal repair **~15 s**;
|
||||
endpoint peer-push convergence **~150 s** (≤5 min bound). All comfortably interactive.
|
||||
|
||||
7. **Reconcile continuity constraint (from P0, load-bearing for any future OOB probe/op).** The box
|
||||
applies its conf every 60 s. Any OOB mechanism that mutates box wg state MUST go through the
|
||||
agent's rendered conf (finding 1), never a side-channel `wg set`/hand-edit — those are silently
|
||||
reverted within a tick. This is a *feature* (drift-repair) to build with, not around.
|
||||
|
||||
8. **Endpoint-access + production prerequisites discovered.** (a) The demo endpoint is a real public
|
||||
VM, so O-A is faithfully represented; production `ep0` needs the **AAAA fixed** (standing
|
||||
OPERATOR item) though the v4-pin means OOB rides v4 regardless. (b) `cloudflared` in the guest is
|
||||
a **Docker container, not a systemd unit** — any runbook step that "stops cloudflared" must
|
||||
`docker stop` it, not `systemctl`. (c) The hub `wg_peers` insert path for a NON-host operator
|
||||
peer exists in the store (`AddWGPeer` with empty host_id — the S1 unbound-admin-row shape); a
|
||||
production OOB feature can reuse it or add an explicit operator-peer admin endpoint.
|
||||
|
||||
---
|
||||
|
||||
## 9. What this spike did NOT prove (honest ledger)
|
||||
|
||||
- **True CGNAT** OOB traversal — the box is single-NAT public-v4 (same caveat as the prior spike
|
||||
§5). OOB rides the same outbound WG mapping as the backup path, so low-risk, but unproven on
|
||||
`100.64/10`.
|
||||
- **A second real agent box** — the dummy customer was a netns peer (sufficient for the topology
|
||||
isolation negative; not a full agent). The box↔box drop is proven by topology + counters, which
|
||||
is the load-bearing property.
|
||||
- **The rendered-conf operator-`/32` change itself** — finding 1 is the *diagnosis* (runtime `wg
|
||||
set` wiped by self-heal, re-add restored OOB); the agent code change to render it is TASK work,
|
||||
not shipped here.
|
||||
- **Endpoint peersync owning the forward rules** — the spike put forward rules in ad-hoc nft; the
|
||||
TASK decides static-vs-hub-driven (finding 3 recommends static forward rules + hub-driven peer
|
||||
list).
|
||||
- **Long-horizon hold** of an idle OOB session (minutes-scale only here; the backup-path 32-min
|
||||
soak is in doc 06 §7).
|
||||
|
||||
---
|
||||
|
||||
## 10. Cleanup assertion (§13 of the plan) — VERIFIED
|
||||
|
||||
**Hub:** both spike `wg_peers` rows deleted (`WHERE note LIKE 'spike-oob-%'`); registry back to the
|
||||
single real peer `demo-felhom-01 / 10.77.0.2`.
|
||||
**Endpoint `felhom-hetzner`:** operator+dummy peers removed from live `wg0`; the 5-min reconciler
|
||||
then rewrote `wg0.conf` to the box-only peer (drift-repair confirmed live — the persisted conf
|
||||
self-healed); `felhom_spike_oob` forward chain + jump rule deleted; `ip_forward` restored to **0**
|
||||
(the P0 value); `spike-op`/`spike-cust` netns + veths deleted; throwaway WG + SSH keys `shred -u`'d,
|
||||
`/root/spike-oob` removed. Final: `nft` forward chain is the empty drop-policy chain (P0), `wg0`
|
||||
lists only the box, `ip_forward=0`.
|
||||
**Box `felhom-pve`:** runtime AllowedIPs back to `10.77.0.1/32` (operator `/32` removed) + return
|
||||
route deleted; `felhom_spike_oob` host table deleted (`nft list tables` empty, as P0); the
|
||||
`spike-oob-operator` line removed from `/root/.ssh/authorized_keys` (grep count 0); `/tmp/wgdown_t0`
|
||||
removed. Final baseline re-verified: `felhom-agent` active, `wg-quick@wg-felhom` active + handshake
|
||||
fresh (age 92 s), `pvesm status felhom-offsite` active, guest 9201 running + controller healthy, no
|
||||
host nft tables, LAN SSH fine.
|
||||
**No repo/hub/agent/manifest production change** — this commit is docs-only.
|
||||
|
||||
---
|
||||
|
||||
## 11. Method bar honored
|
||||
|
||||
`wg show <if> dump` NEVER run (private-key leak ban, S1); only `latest-handshakes` + redacted conf
|
||||
reads. Every mutation was a named one-flush nft table or a runtime-only `wg set`, inventoried as
|
||||
made. PBS-path re-checked after the box wg change (P2). Negative probes carry nft-counter evidence,
|
||||
not just timeouts. Surprises recorded as findings (the self-heal wipe of the runtime `/32` — finding
|
||||
1 — is the headline example).
|
||||
Reference in New Issue
Block a user