docs(H1): doc06 §4.5/§4.6 amendment + endpoint runbook §9 + scripts CHANGELOG + REPORT + CONTEXT

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PSK5g6qYLknKj8u3QAFEr6
This commit is contained in:
2026-07-05 23:03:33 +02:00
parent b70f2d0763
commit 61f4898d30
5 changed files with 113 additions and 49 deletions
@@ -224,16 +224,30 @@ production endpoint exists.
11.4-minute fully-idle window (P2) at ~150 B/s of overhead traffic; **further proven through a
live mobile-carrier NAT for a 32-minute fully-idle soak, zero stalls** (2026-07-04 CGNAT smoke
test, §7).
- **4.5 Isolation.** Per-peer `/32` `AllowedIPs`; IP forwarding stays **off** on the endpoint; its
firewall admits, from the WG interface, only the PBS port — so a box can reach the PBS API and
nothing else, and boxes cannot see each other **by topology** (spike P6). Public surface: SSH
(operator) + the WG UDP port, nothing more. PBS tenancy on top: namespace + per-customer token +
per-customer key (D5).
- **4.6 Tunnel health → hub.** The tunnel is a storage dependency, so it reports like one — the
storage-manifest model (01 §8: agent "continuously checks presence/reachability, and reports
per-target status; a disconnected target → actionable notification") gains a tunnel-health
input: no handshake within ~3 keepalive periods → the offsite target reports unreachable → the
existing alerting path carries it. No new alarm channel.
- **4.5 Isolation.** Per-peer `/32` `AllowedIPs`; boxes cannot see each other (spike P6). PBS tenancy
on top: namespace + per-customer token + per-customer key (D5).
**AMENDED 2026-07-05 (TASK H1 — OOB operator access).** Forwarding is no longer blanket-**off**; it
is **ON but per-pair allow-listed**. The endpoint runs `net.ipv4.ip_forward=1` (sysctl.d) and a
static forward posture: `ct established,related accept`; **per (operator, box) pair** `ip saddr
<operator/32> ip daddr <box/32> accept`; and **box↔box `iifname wg0 oifname wg0` DROP is now an
EXPLICIT rule** (previously implicit under the absent capability), backed by the base-chain `policy
drop`. These rules live in the endpoint's static nftables (NOT in `felhom-peersync`, which still
manages only the peer *list*). Net effect: the operator peer reaches a box's `felhom-sshd`; boxes
still cannot reach each other or the operator (only conntrack replies flow). The box side adds a
second layer independent of the endpoint: a dedicated `felhom-sshd` on a claimed non-22 port, gated
by the host-local `inet felhom_oob` belt (reachable only from the operator `/32` over `wg-felhom`;
the customer's stock sshd on :22 is never touched). Live-proven: operator→box SSH works; a dummy
tunnel peer is dropped box↔box (counter); the operator `/32` is **rendered** into the box's
`wg-felhom` AllowedIPs so it survives the agent's self-heal ([OF-1]); PBS unaffected.
- **4.6 Tunnel health → hub.** The tunnel is a storage dependency, so it reports like one: no handshake
within ~3 keepalive periods → the offsite target reports unreachable → the existing alerting path
carries it. No new alarm channel.
**EXTENDED 2026-07-05 (TASK H1).** The agent's heartbeat now also carries an **`oob` stanza**
(`felhom_sshd_active`, `felhom_sshd_port`, `reachable`, `config_invalid`, `operator_peer_configured`,
`operator_key_configured`, `wg_handshake_age_s`) — the operator's "can I get into this box right
now, and if not, why" signal. It reaches the hub over HTTPS even when `felhom-sshd` or the tunnel is
DOWN (channel independence). The hub raises a transition-based **`oob_degraded`/`oob_recovered`**
warning (felhom-sshd down while the operator peer is configured, OR config invalid).
---
@@ -311,3 +311,38 @@ The rebuild is **steps 17 on a fresh VM**. What is lost vs regenerable:
(customers still hold local backups + a re-seedable offsite).
- The hub's SSH credential + pinned host key must be **rotated on rebuild** (new box = new
host key): repeat step 6 (`kubectl delete secret wg-endpoint-ssh` first), roll the hub.
## 9. OOB operator forwarding (TASK H1 — 2026-07-05)
The endpoint gains a **per-pair operator→box forward** posture so the operator peer can reach each
box's `felhom-sshd` (doc 06 §4.5 amended: forwarding ON but per-pair allow-listed; box↔box drop is
now explicit). Peersync is UNCHANGED — it still manages only the peer *list*; these forward rules are
STATIC endpoint config.
```sh
# 1. permanent forwarding
echo 'net.ipv4.ip_forward = 1' > /etc/sysctl.d/99-felhom-oob.conf
sysctl -w net.ipv4.ip_forward=1
# 2. forward posture in the STATIC nftables filter forward chain (add to /etc/nftables.conf's
# `chain forward` — which keeps `policy drop`). ONE accept rule per (operator, box) pair; the
# box↔box drop is explicit. Reload path so replies + the PBS path are unaffected (INPUT hook).
# <operator/32> = GET /api/v1/admin/wg/operator-peer ; <box/32> = each host's assigned_ip.
ct state established,related accept
iifname "wg0" oifname "wg0" ip saddr <operator/32> ip daddr <box/32> counter accept
iifname "wg0" oifname "wg0" counter drop
```
**Register the operator peer** (hub, global key) — it becomes an UNBOUND `wg_peers` row (peersync
pushes it to wg0) and its `/32` flows to every box as `oob_peer_ip`:
```sh
curl -s -X PUT https://hub.felhom.eu/api/v1/admin/wg/operator-peer \
-H "Authorization: Bearer <GLOBAL-KEY>" \
-d '{"pubkey":"<operator WG pubkey>","assigned_ip":"10.77.0.250","ssh_pubkey":"ssh-ed25519 AAAA… operator"}'
```
**Box side:** `felhom-host-install … --enable-oob` (static felhom-sshd + belt) + `oob.enabled=true`
in `agent.json`. The agent renders the config, claims a port, writes `felhom-op`'s authorized_keys
from `oob_operator_ssh_key`, and fills the belt sets. Verify: from the operator peer,
`ssh -p <claimed-port> felhom-op@<box tunnel IP>`; the host belt drops any non-operator tunnel source.
Re-verify PBS (`pvesm status --storage felhom-offsite` on the box) after the forward change.