Files
felhom.eu/documentation/runbooks/offsite-endpoint.md
T
admin d2cc02344c runbook: offsite-endpoint.md — executed + validated live on felhom-hetzner (S1)
All 8 steps run on the dev endpoint; live-run corrections folded in: enterprise-
repo removal after PBS install, proxmox-backup-client is a separate package,
throwaway-token namespace proof, wg-show-dump-leaks-the-private-key warning
(incident: first server key leaked to session log -> rotated on the spot),
AAAA-must-be-::1 + DNS-propagation-lag notes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PSK5g6qYLknKj8u3QAFEr6
2026-07-04 00:12:07 +02:00

10 KiB
Raw Blame History

RUNBOOK — offsite endpoint provisioning (WG server + PBS + peersync surface)

What this creates: the Felhom offsite endpoint (doc 06 §5) on a blank Debian 13 VM — the WG server on 443/UDP, the offsite PBS with a /srv datastore, a drop-everything-else firewall, and the hub-driven felhom-peersync reconcile surface. Re-running it on a fresh VM re-creates the endpoint from nothing (that is the DR story, step 8).

Validated: 2026-07-03 on the dev/test endpoint felhom-hetzner (Hetzner CX23, Debian 13, 167.233.158.164 / 2a01:4f8:1c16:7aa1::1) with hub v0.32.0. The production endpoint is a later re-run of this runbook on a production VM.

Companion piece: the hub side (env + Secret) is manifests/hub.yaml + step 6/7 below. The peersync script's source of truth is scripts/felhom-peersync.sh (v1.0.0).

Parameters used throughout (adjust for a new endpoint):

Param Dev value
VM felhom-hetzner, 167.233.158.164 (NOT the live jarrs.eu box — different server)
DNS name ep0.felhom.eu (A + AAAA, DNS-only / grey-cloud — proxying breaks WG/UDP + PBS)
WG port 443/udp (spike P4: handshakes + carries identically to 51820)
Tunnel subnet 10.77.0.0/24; endpoint in-tunnel IP 10.77.0.1
PBS datastore felhom-offsite at /srv/pbs-felhom (never under /root — the backup user cannot traverse it; spike gotcha)

0. Preconditions (operator)

  • The VM is powered on and root-SSH-reachable.
  • DNS ep0.felhom.eu → A 167.233.158.164 + AAAA 2a01:4f8:1c16:7aa1::1, grey-cloud. Verify: nslookup ep0.felhom.eu 1.1.1.1 returns the A record. The AAAA must be the box's full ::1 address — a bare ...:7aa1:: (the subnet zero address) resolves but points at nothing (live-run finding). Home-resolver propagation can lag public DNS by minutes — a client-side wg-quick up that fails to resolve right after record creation just needs a retry.
  • Sanity: ssh root@167.233.158.164 hostnamefelhom-hetzner (the throwaway CX23), not any production box.

1. Base (on the box, as root)

apt update && DEBIAN_FRONTEND=noninteractive apt full-upgrade -y
printf 'PasswordAuthentication no\n' > /etc/ssh/sshd_config.d/90-felhom.conf && systemctl reload ssh
sshd -T | grep -i '^passwordauthentication'   # must print: passwordauthentication no

(The explicit drop-in beats auditing defaults — Hetzner images have no explicit line at all.)

2. WG server

apt install -y wireguard-tools          # module ships in the stock Debian 13 kernel
umask 077
wg genkey > /etc/wireguard/wg0.key      # server private key — NEVER displayed/copied anywhere
wg pubkey < /etc/wireguard/wg0.key      # RECORD this pubkey → hub registration (step 7)
cat > /etc/wireguard/wg0.conf.head <<EOF
[Interface]
Address = 10.77.0.1/24
ListenPort = 443
MTU = 1420
PrivateKey = $(cat /etc/wireguard/wg0.key)
EOF
chmod 600 /etc/wireguard/wg0.conf.head
cp -p /etc/wireguard/wg0.conf.head /etc/wireguard/wg0.conf   # zero peers initially
systemctl enable --now wg-quick@wg0
wg show wg0

Verify: wg show wg0listening port: 443, no peers. No SaveConfig anywhere — the peersync script owns persistence (head-file + generated peers; wg-quick save would rewrite the conf nondeterministically).

NEVER run wg show <if> dump in a logged/shared session — field 1 of its interface line is the private key. Plain wg show prints only the public key. (Live-run lesson: a dump | cut leaked the first server key into a session log → the key was rotated on the spot. If it happens to you: regenerate the keypair, rebuild wg0.conf.head + wg0.conf, restart wg-quick@wg0, re-PUT the new pubkey to the hub — 2 minutes, zero peer downtime beyond the restart.)

3. Firewall (nftables, v4+v6)

cat > /etc/nftables.conf <<'EOF'
#!/usr/sbin/nft -f
flush ruleset
table inet filter {
    chain input {
        type filter hook input priority 0; policy drop;
        ct state established,related accept
        iif "lo" accept
        tcp dport 22 accept
        udp dport 443 accept
        tcp dport 8007 iifname "wg0" accept
        ip protocol icmp accept
        meta l4proto ipv6-icmp accept
    }
    chain forward {
        type filter hook forward priority 0; policy drop;
    }
}
EOF
systemctl enable --now nftables
nft list ruleset | head -20
sysctl net.ipv4.ip_forward

Verify (doc 06 §4.5 — no inter-peer routing by topology): forward chain policy drop, net.ipv4.ip_forward = 0. From an outside vantage: ssh (22) connects; 8007 publicly CLOSEDcurl -sk --max-time 8 https://<public-ip>:8007/ must time out (the spike-P2b firewall proof, repeated). 443/udp is proven by the first WG handshake in step 7's validation.

4. PBS

curl -fsSL https://enterprise.proxmox.com/debian/proxmox-release-trixie.gpg \
  -o /usr/share/keyrings/proxmox-archive-keyring.gpg
echo 'deb [signed-by=/usr/share/keyrings/proxmox-archive-keyring.gpg] http://download.proxmox.com/debian/pbs trixie pbs-no-subscription' \
  > /etc/apt/sources.list.d/pbs.list
apt update && DEBIAN_FRONTEND=noninteractive apt install -y --no-install-recommends \
  proxmox-backup-server proxmox-backup-client   # client is a SEPARATE package (namespace ops need it)
rm -f /etc/apt/sources.list.d/pbs-enterprise.sources   # the install adds the ENTERPRISE repo → 401s every apt update
apt update
proxmox-backup-manager datastore create felhom-offsite /srv/pbs-felhom
# prove namespace ops (per-customer tenancy, doc 06 D5) with a THROWAWAY token, then remove it
# (root@pam has no password on a cloud image; S4 wires the real per-customer tokens):
umask 077
proxmox-backup-manager user generate-token root@pam nsproof > /root/.nsproof
proxmox-backup-manager acl update /datastore/felhom-offsite DatastoreAdmin --auth-id 'root@pam!nsproof'
export PBS_PASSWORD="$(tr -d '\n' < /root/.nsproof | sed -n 's/.*"value": "\([^"]*\)".*/\1/p')"
export PBS_FINGERPRINT="$(proxmox-backup-manager cert info | grep -i fingerprint | awk '{print $3}')"
proxmox-backup-client namespace create test-ns --repository 'root@pam!nsproof@localhost:felhom-offsite'
proxmox-backup-client namespace delete test-ns --repository 'root@pam!nsproof@localhost:felhom-offsite'
proxmox-backup-manager user delete-token root@pam nsproof && rm -f /root/.nsproof

Verify: systemctl is-active proxmox-backup proxmox-backup-proxy → both active; the datastore parent is /srv (world-traversable — a datastore under /root fails with Permission denied for the backup user). GUI: reachable ONLY via SSH port-forward (ssh -L 8007:127.0.0.1:8007 root@<box>) or through the tunnel — by design, no public 8007.

5. Peersync surface

useradd -m -s /bin/sh felhom-peersync   # real shell required for the forced command;
                                        # access is bounded by the key options + sudoers, not the shell
apt install -y jq
# install the script (from the repo; strip CRLF if staging via a Windows checkout):
tr -d '\r' < scripts/felhom-peersync.sh > /usr/local/bin/felhom-peersync
chown root:root /usr/local/bin/felhom-peersync && chmod 0755 /usr/local/bin/felhom-peersync
cat > /etc/sudoers.d/felhom-peersync <<'EOF'
felhom-peersync ALL=(root) NOPASSWD: /usr/local/bin/felhom-peersync
EOF
visudo -c

Verify: visudo -c → "parsed OK". The authorized_keys entry is written in step 6 (it needs the hub's pubkey):

restrict,command="sudo /usr/local/bin/felhom-peersync" ssh-ed25519 AAAA... hub@felhom

(restrict kills pty/forwarding/X11/agent in one word; the forced command overrides whatever the client asks to run.)

6. Hub credential (build server + the box + k3s)

On the build server (192.168.0.180):

ssh-keygen -t ed25519 -f wg-endpoint-ssh -N "" -C hub@felhom
  • The pubkey (wg-endpoint-ssh.pub) goes into step 5's authorized_keys line on the box: printf 'restrict,command="sudo /usr/local/bin/felhom-peersync" %s\n' "$(cat wg-endpoint-ssh.pub)" > /home/felhom-peersync/.ssh/authorized_keys (as root on the box; .ssh dir 0700, file 0600, owner felhom-peersync).
  • Read the endpoint's host key on the box itself (never trust ssh-keyscan alone): cat /etc/ssh/ssh_host_ed25519_key.pub
  • Create the k8s Secret, then shred the local private key — from this point it exists ONLY in the Secret:
sudo kubectl -n felhom-system create secret generic wg-endpoint-ssh \
  --from-file=key=wg-endpoint-ssh \
  --from-literal=hostkey='ssh-ed25519 AAAA...the-host-key-line'
shred -u wg-endpoint-ssh wg-endpoint-ssh.pub

Then apply the manifest + roll the hub (manifests/hub.yaml already mounts the Secret as optional): expect [INFO] WG peer-sync enabled (endpoint 167.233.158.164:22, ...) in the hub log (before the Secret existed it says disabled (endpoint not configured)).

7. Register the endpoint in the hub

curl -s -X PUT https://hub.felhom.eu/api/v1/admin/wg/endpoint \
  -H "Authorization: Bearer <GLOBAL-KEY out-of-band>" \
  -d '{"dns_name":"ep0.felhom.eu","wg_port":443,"server_pubkey":"<step-2 pubkey>",
       "tunnel_subnet":"10.77.0.0/24","pbs_tunnel_ip":"10.77.0.1"}'

Validation loop (the S1 done-criterion): POST /api/v1/admin/wg/peers {"pubkey":...} with a throwaway pubkey → response sync:"ok" → on the box wg show wg0 peers lists it → DELETE /api/v1/admin/wg/peers {"pubkey":...} → gone from wg show. A live handshake + curl -sk https://10.77.0.1:8007/ through a throwaway client tunnel proves 443/udp + the wg0-only 8007 rule end-to-end.

8. Re-provision from nothing (the DR story)

The rebuild is steps 17 on a fresh VM. What is lost vs regenerable:

  • WG server keypair — regenerable, but every existing peer's config then points at a dead server pubkey. Until S2/S3 deliver endpoint coords via desired-state, re-keying is a manual client-side update (dev-acceptable; note per doc 06 §7). Register the new pubkey via step 7; peers re-converge as they're updated.
  • The hub's peer registry survives (it lives in the hub DB): a rebuilt endpoint converges on the first reconciler tick (≤5 min) or the next mutation — no re-registration needed.
  • The PBS datastore is the real loss — customer ciphertext. PBS-side redundancy is a deferred economics call (doc 06 §7); until then, endpoint loss = offsite history loss (customers still hold local backups + a re-seedable offsite).
  • The hub's SSH credential + pinned host key must be rotated on rebuild (new box = new host key): repeat step 6 (kubectl delete secret wg-endpoint-ssh first), roll the hub.