docs(audit): OOB-over-WG operator-peer spike (2026-07-05) — GO

Validates operator-inbound access over the existing offsite WG arc (doc 06):
operator peer forwarded operator->box only, box sshd gated to the operator /32,
§4.5 box<->box isolation intact (both negatives counter-proven), mutual repair
real (agent self-healed a stopped tunnel in ~15s unaided). One TASK-shaping gap:
the operator /32 must be a RENDERED conf field — a runtime `wg set` is wiped by
the agent's own self-heal. All live-arc changes reverted + baseline re-verified.

Docs-only; no code/hub/agent/manifest change.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PSK5g6qYLknKj8u3QAFEr6
This commit is contained in:
2026-07-05 17:01:35 +02:00
parent b6bad953b5
commit 2db92c8837
@@ -0,0 +1,366 @@
# SPIKE — OOB management over the existing WG arc (operator peer, isolation, mutual repair) — 2026-07-05
> **STATUS: COMPLETE — verdict below.** Empirically validates the unproven mechanisms of the OOB
> (out-of-band operator access) design BEFORE the production spec. No production code shipped; no
> agent binary change; every config change on the live demo arc reverted at cleanup (§13, verified).
> The design under test: an **operator peer on the existing hub-and-spoke WG arc** (doc 06),
> endpoint-forwarded **operator→box only**, box sshd reachable solely from the operator tunnel IP,
> while §4.5 box↔box isolation stays intact — and the hub desired-state channel as the tunnel's
> **repair path**.
**Class:** SPIKE (empirical; no product code). **Repos:** felhom.eu (this doc only); felhom-agent
read-only for grounding (`internal/wgtunnel/{manager,loop}.go`, `internal/desired/syncer.go`,
`configs/felhom-agent.sudoers`); felhom-hub read-only (`internal/wgsync/reconciler.go`,
`internal/store/wg.go`).
**Probe ends (the LIVE demo arc, ground-truthed in P0):**
- **Endpoint** = `felhom-hetzner` (Hetzner CX23, Debian 13, `167.233.158.164`, DNS `ep0.felhom.eu`),
the **dev** offsite endpoint: WG server `wg0` on **443/UDP**, tunnel `10.77.0.0/24`, endpoint
`10.77.0.1`, PBS `felhom-offsite`. Peer list is hub-driven (`internal/wgsync` SSH push →
`felhom-peersync`). CC-SSH-reachable as root ✓ (the P0 access gate passed).
- **Box** = `felhom-pve` / node `demo-felhom` (`192.168.0.162`), agent **v0.70.0**, `wg-felhom`
client `10.77.0.2/32`, tunnel to `ep0:443`. CC's LAN SSH to it was the safety line throughout.
**Verdict (one line):** **GO — an operator peer on the existing arc reaches a box's sshd through the
endpoint with a single added `/32`, at ~24 ms interactive latency, WITHOUT breaking §4.5 isolation
(both negatives held with packet-level evidence); mutual repair is REAL and already built (the
agent self-healed a stopped tunnel in ~15 s with zero hub/LAN involvement, and OOB works with the
agent dead); the one load-bearing gap the TASK must close is that the operator `/32` must become a
RENDERED conf field — a runtime-only `wg set` is wiped by the agent's own self-heal/restart.**
---
## 0. What "production ep0" this demo could and could not represent
The demo endpoint (`felhom-hetzner`) **is** a real public dual-stack cloud VM with a public IPv4,
443/UDP WG, hub-driven peersync, and a real PBS — architecturally identical to the intended
production `ep0`. So forwarding, isolation, and the SSH-through-endpoint path are represented
faithfully. What it does **not** represent: (a) a **CGNAT** customer line — the box here is
`felhom-pve` on the operator's One-Hungary fixed-cable line (public IPv4 `37.191.56.193`,
single-NAT), same caveat as the prior spike §5; the OOB path rides the same outbound-mapping
mechanism, so this is low-risk but unproven on true `100.64/10`. (b) A **second real customer box**
— the "dummy customer" was a throwaway netns peer at the endpoint, which is sufficient for the
topology-isolation negative but is not a full second agent. (c) The production `ep0` **AAAA** is
still mis-set per the standing OPERATOR follow-up (memory: "OPERATOR fix ep0 AAAA"); this spike
used v4 throughout, consistent with the agent's v4-pin (doc 06 §4.2).
---
## 1. P0 — ground-truth discovery (no changes) — **GATE PASSED**
**Endpoint `wg0` (redacted; `wg show` prints only the public key — `dump` is BANNED, S1 rule):**
```
interface: wg0 listening port: 443
peer: yV5hFFZF548r6kq2Nl/CiSyOtFb0xYnvEAjw1ofrTCg= (= the box)
endpoint: 37.191.56.193:54900 allowed ips: 10.77.0.2/32
latest handshake: ~1.5 min ago transfer: 5.51 GiB rx / 16.43 GiB tx
wg0.conf: [Interface] Address 10.77.0.1/24, ListenPort 443, MTU 1420 ; [Peer] box only
sysctl: net.ipv4.ip_forward = 0 net.ipv6.conf.all.forwarding = 0
nft: table inet filter { input: drop-policy, allows 22, 443/udp, 8007 iif wg0 only, icmp ;
forward: policy DROP (empty) }
PBS: proxmox-backup + proxmox-backup-proxy active
```
**Box `wg-felhom` (redacted):**
```
wg-felhom.conf (agent-rendered, "DO NOT EDIT"): Address 10.77.0.2/32, MTU 1280,
Peer f3d1ZI7…= Endpoint 167.233.158.164:443 AllowedIPs 10.77.0.1/32 PersistentKeepalive 25
unit wg-quick@wg-felhom: enabled, active (exited) handshake fresh
agent marker /var/lib/felhom-agent/wg/registered.json:
{"pubkey":"yV5hFFZF548r6kq2Nl/CiSyOtFb0xYnvEAjw1ofrTCg=","assigned_ip":"10.77.0.2/32","generation":5}
sshd: port 22, ListenAddress 0.0.0.0:22 + [::]:22 (NOT interface-scoped today)
pve-firewall: disabled ; nft list tables: EMPTY (no host nft tables at baseline)
```
**Reconcile model (source + journal):**
- **Box→tunnel reconcile is CONTINUOUS, not generation-gated.** `wgtunnel.Loop.Run`
(`internal/wgtunnel/loop.go`) calls `Manager.Apply` **on every tick**`wg_tunnel.interval_seconds
= 60` (agent.json, confirmed) — AND on each desired-state nudge. The desired-state *fetch* is
generation-gated (`desired/syncer.go` `OnEnvelope`, hub poll `poll_seconds = 900`), but the
wireguard block is cached and re-applied every 60 s regardless. **Consequence for probes:** any
hand-edit to `wg-felhom.conf` or a runtime `wg set` that diverges from the agent's rendered conf
is at risk of being reconciled within ≤60 s. This is why the AllowedIPs widening (P2) used
runtime-only `wg set` (which the agent does NOT touch — it only rewrites the *conf* and restarts
on a conf-hash change) and why the self-heal in P5b wiped it (§P5).
- **Endpoint peer-list reconcile** (`hub/internal/wgsync/reconciler.go`): declarative FULL-list push
every **5 min** (or on a mutation Trigger); drift (a manually-added/removed peer) is erased on the
next push. Proven live at cleanup (§13).
**Gate:** endpoint is CC-SSH-accessible as root ✓ → proceed.
---
## 2. P1 — operator peer, endpoint-side — **GO**
Throwaway operator + dummy-customer keypairs generated on the endpoint (`wg genkey`, 0600,
private keys never printed). Registered BOTH into the **hub** `wg_peers` registry (not hand-added
to `wg0`) so they survive the 5-min reconciler — via `sqlite3 /data/hub.db` in the hub pod
(stdin-piped SQL to avoid the base64 `=`/`+` quoting trap):
```
operator cdmN4U+fjR18zBk+SKoceJQyz9HgA9+hN8/FiKF1u0o= 10.77.0.250
dummy tQ6eJC9y8pSKtam44DaSrrtYBZYUvZG2z1stS6PsaD4= 10.77.0.251
```
The hub reconciler pushed both to `wg0` in **~150 s** (next 5-min tick fired early on a warm loop).
**Endpoint forwarding, enabled NARROWLY**`ip_forward=1` (was 0) + a named, one-flush nft chain
`felhom_spike_oob` hung off `forward`, ALL rules scoped `iifname wg0 oifname wg0`:
```
table inet filter { chain forward { … ; jump felhom_spike_oob } }
chain felhom_spike_oob {
iifname "wg0" oifname "wg0" ct state established,related accept # conntrack replies
iifname "wg0" oifname "wg0" ip saddr 10.77.0.250 ip daddr 10.77.0.2 counter accept # operator→box
iifname "wg0" oifname "wg0" ip daddr 10.77.0.250 counter drop # box→operator NEW: drop
iifname "wg0" oifname "wg0" counter drop # box↔box (any other): drop
}
```
The operator "laptop" is a **netns client** on the endpoint (`spike-op` netns, veth to a
`192.168.99.0/30` link, `wg-op` interface, MTU 1280, `AllowedIPs 10.77.0.0/24`) — an acceptable
stand-in; the crypto+routing path through `wg0` is identical to a remote laptop. **GO:** operator
peer handshakes with `wg0` and pings the endpoint `10.77.0.1` at **0.2 ms** (on-box netns).
---
## 3. P2 — box-side AllowedIPs widening + SSH-over-tunnel + PBS — **GO**
**The minimal box change, runtime-only** (nothing persists past a wg-quick restart — see P5):
```
wg set wg-felhom peer f3d1ZI7…= allowed-ips 10.77.0.1/32,10.77.0.250/32
ip route add 10.77.0.250/32 dev wg-felhom
```
Exactly one added `/32` (+ the return route). `wg-felhom.conf` mtime UNCHANGED (2026-07-04) — the
agent's rendered conf was not touched.
From the operator netns, through the endpoint's forwarding:
```
ping 10.77.0.2 → 5/5, avg 24.7 ms, TTL 63 (TTL 63 = exactly one forwarded hop, i.e. via ep0)
ssh root@10.77.0.2 "hostname; pct list; uptime" → interactive session in 1.205 s:
demo-felhom ; 9201 running demo-felhom (+9001/9100/9999) ; load avg 0.43
```
Interactive SSH over the tunnel works end-to-end.
**PBS-path check (mandatory, immediately after the box change):**
```
wg show wg-felhom latest-handshakes → age 73 s (fresh, advancing)
pvesm status --storage felhom-offsite → active (PVE reaches PBS at 10.77.0.1:8007 over the tunnel)
openssl s_client -connect 10.77.0.1:8007 → CONNECTED, Certificate chain present
```
**PBS unaffected = GO condition met.** Adding the operator reach did not perturb the backup path.
---
## 4. P3 — isolation negative probes (§4.5 must survive) — **GO (both negatives, counter-proven)**
**Negative 1 — dummy customer → box: DROPPED.** The dummy peer handshakes with the endpoint (0.2 ms
to `10.77.0.1` — it IS a registered peer) but:
```
spike-cust → ping 10.77.0.2 → 3/3 packets LOST (100%)
spike-cust → tcp/22 10.77.0.2 → REFUSED/TIMEOUT
felhom_spike_oob box↔box drop counter: 0 → 8 packets / 552 bytes (the drops are real, not a route miss)
```
**Negative 2 — box → operator: NEW connections DROPPED, replies flow.** With an operator→box SSH
session OPEN (so ESTABLISHED replies were flowing), from the box:
```
felhom-pve → ping 10.77.0.250 → 3/3 LOST
felhom-pve → tcp/22 & tcp/80 10.77.0.250 → REFUSED/TIMEOUT
felhom_spike_oob box→operator drop counter: 0 → 13 packets / 852 bytes
(while the operator→box accept counter incremented normally — conntrack replies unaffected)
```
Both hold with **packet-level evidence** (nft counters), not merely "command timed out". §4.5
isolation survives the operator-peer addition: a box still cannot reach another box, and a box
cannot initiate to the operator.
---
## 5. P4 — sshd scoping on the box (belt rule) — **GO (all three)**
A named, one-flush nft table on `felhom-pve` (baseline had NO host nft tables), independent hook
so LAN and other interfaces are untouched:
```
table inet felhom_spike_oob {
chain input { type filter hook input priority -10; policy accept;
iifname "wg-felhom" tcp dport 22 ip saddr 10.77.0.250 counter accept # operator only
iifname "wg-felhom" tcp dport 22 counter drop # any other tunnel src
} }
```
Verified:
1. **LAN SSH from CC — unaffected** (the safety line): the rule matches `iifname wg-felhom` only;
LAN commands kept running throughout. ✓
2. **Operator SSH over the tunnel — still works** (accept counter 27 pkts). ✓
3. **Defense-in-depth — dummy refused at the HOST even if the endpoint belt failed:** I temporarily
simulated an endpoint-belt failure (added an endpoint forward-accept for dummy→box AND widened
the box AllowedIPs to include `10.77.0.251/32`). Result: dummy→box **ping now passes** (endpoint
forwards it, 27 ms) but dummy→box **SSH is dropped at the host** (host belt drop counter 5) —
`Connection timed out`. The two layers are independent; the host belt holds when the endpoint
layer is bypassed. The temporary widening was reverted immediately after. ✓
---
## 6. P5 — mutual repair (the crown probe)
### 6a. Agent dead, wg alive — **GO**
`systemctl stop felhom-agent` (via LAN, simulating no other way in). Then from the operator peer
**over the tunnel**: SSH into the box (agent DOWN → OOB does not depend on the agent), and
`systemctl reset-failed felhom-agent; systemctl start felhom-agent` from that session →
`active`, `felhom-agent 0.70.0`. **OOB works with the agent down AND repairs it.**
### 6b. wg dead, agent alive — **GO, ~15 s, and this is THE central TASK finding**
`systemctl stop wg-quick@wg-felhom` (interface gone). Operator SSH over the tunnel now
`Connection timed out` (OOB down, confirmed). Then, watching the endpoint's view of the box peer
handshake with **zero LAN intervention**:
```
box re-handshake observed at +15 s (agent log: "wgtunnel: service inactive with current conf —
re-enabling (self-heal)")
```
The agent's reconcile loop **already covers a stopped unit**`Manager.ensureTunnelLocked`'s
self-heal branch (`internal/wgtunnel/manager.go:394-402`): conf hash current + `isActive()` false →
`systemctl enable --now`. No `wg_repair`-shaped desired action is needed; the "wg dead, agent alive"
repair path exists today and fires within one 60 s tick (here 15 s). **This closes O-B (mutual
repair via the desired-state channel) — with a caveat that becomes the TASK's main work:**
> **THE GAP.** The self-heal rebuilds `wg-felhom` **purely from the agent-rendered conf**, whose
> `AllowedIPs` is derived solely from `pbs_tunnel_ip/32` (`renderConf`, `manager.go:239`:
> `AllowedIPs = <pbsIP>/32`). After the self-heal, the box's live AllowedIPs was back to **only
> `10.77.0.1/32`** — the operator `/32` I added at runtime was **wiped**, and OOB SSH failed until I
> re-added it. So **OOB survival across the agent's own self-heal/restart/reboot REQUIRES the
> operator peer's `/32` to be part of what the agent RENDERS into the conf** — a runtime `wg set` is
> not durable. (Re-adding `10.77.0.250/32` at runtime immediately restored OOB, confirming the
> mechanism; it is only the *durability* that needs the rendered-conf change.)
### 6c. Failure-domain sweep with an open OOB session — **GO**
Over a single open operator→box tunnel session: attempted `cloudflared` stop (it is NOT a guest
systemd unit — it runs as a **Docker container** in the controller stack, so the unit stop no-op'd;
noted), then `pct stop 9201` (whole guest down). **The session stayed alive** (`ALIVE_2`, handshake
fresh) — the tunnel terminates **host-side** on `felhom-pve`, so guest death cannot touch it — and
`pct start 9201` from the SAME session brought the guest back. Post-check: `felhom-controller Up
(healthy)`, `cloudflared`/`nextcloud` containers `Up (healthy)`. OOB survives the failure domains it
exists for (guest down; and by construction cloudflared/controller down, since OOB is host-side and
independent of the guest's own edge).
---
## 7. P6 — usability + environment (record-only)
- **Interactive feel over MTU 1280:** `top -bn1`, `pct exec 9201 -- …` (into the guest), and a
20 MB `dd | wc -c` bulk pull all ran cleanly over the tunnel at ~24 ms RTT — no black-holing, no
stalls (MTU 1280 is the doc 06 §4.3 floor; interactive OOB is well within it).
- **Endpoint load** during the probes: `load average 0.10` — WG forwarding + a shell session is
negligible on a CX23.
- **Active port: 443/UDP** (`wg show wg0 listen-port → 443`). Relevant to CGNAT: this is the
friendly-port path already; the box's NAT source was `37.191.56.193:47922` (single-NAT public v4,
not CGNAT — the standing caveat).
- **Keepalive:** `PersistentKeepalive 25` on the box; handshakes stayed fresh (<125 s) throughout.
---
## 8. Findings that must shape the TASK
1. **Render the operator `/32` into the agent conf — the one required agent change (§6b).** OOB
reach must survive the agent's self-heal/restart/reboot, so the operator peer's tunnel `/32`
must be a field the agent puts in `wg-felhom.conf` `AllowedIPs` (today hard-derived from
`pbs_tunnel_ip` alone). Shape options for the TASK to pick: (a) extend the hub `wireguard`
desired-state block with an optional `oob_peer_ip` (or a small `extra_allowed_ips` list) and have
`renderConf` append it; (b) a separate `[Peer]`-less allowed-ips widening is NOT enough (WG ties
AllowedIPs to the *server* peer here, and all tunnel traffic transits the one endpoint peer, so
the operator `/32` just joins the existing peer's AllowedIPs). Keep the derivation deterministic
and validated (the existing `netip.ParsePrefix`/`/32` checks). **Nothing else in the agent needs
to change for O-A/O-B** — reconcile continuity + self-heal already do the repair.
2. **Mutual repair (O-B) is already built — do NOT add a `wg_repair` action.** The 60 s reconcile
tick + `ensureTunnelLocked` self-heal restored a stopped tunnel in 15 s unaided (§6b). The TASK
should *rely on* this, and at most add the tunnel-health→alert wiring already slated for S6 (doc
06 §4.6 stanza exists). Endpoint-side drift repair likewise already works (5-min declarative
full-list push erased the spike peers from `wg0.conf` at cleanup, §13).
3. **The exact nft rulesets that worked (copy-paste ready).** Endpoint forward chain and the
host belt are in §2 and §5 verbatim. Both are named, one-flush tables. Production shape:
- **Endpoint:** `ip_forward=1` + a `felhom_oob` forward chain: `ct established,related accept`;
`saddr <operator/32> daddr <box/32> accept` **per (operator,box) pair**; a catch-all
`iifname wg0 oifname wg0 … drop` to preserve §4.5 for every non-OOB pair. Note this makes the
endpoint's forward posture **per-pair allow-list** instead of doc 06 §4.5's blanket
"forwarding OFF" — the TASK must state that the OOB feature deliberately turns forwarding ON
but gates it to operator→box pairs only, and that the box↔box drop is now an explicit rule, not
an absent capability. Peersync (`felhom-peersync.sh`) would need to learn these forward rules,
OR they live in the endpoint's static nftables and only the peer *list* stays hub-driven
(simpler; recommended — the operator set is small and stable).
- **Box (belt, defense-in-depth):** a `felhom_oob` inet table, `input` hook, `iifname wg-felhom
tcp dport 22 ip saddr <operator/32> accept` then `iifname wg-felhom tcp dport 22 drop`. Scope
strictly to `wg-felhom`; never touch LAN sshd (the safety line). This belongs in the agent's
managed surface if the box is to enforce it (a new narrow grant), OR ship as host-install
static config; the spike proved the *rule* works, the TASK picks the ownership.
4. **AllowedIPs method: runtime `wg set` is NOT durable (finding 1); the durable path is the
rendered conf.** For the endpoint side, the operator peer is a normal hub `wg_peers` row
(survives the reconciler) — that part needs no new mechanism, only that the operator peer is
*not* a customer host (host_id empty, as the spike used) and is excluded from any customer-facing
accounting.
5. **sshd scoping shape.** Today box sshd listens on `0.0.0.0:22 + [::]:22` (P0). The spike scoped
access with an **nft belt on `iifname wg-felhom`**, NOT by changing `ListenAddress` — deliberate:
an interface-bound `ListenAddress` would fight the LAN safety line and the tunnel's late bring-up.
The TASK should keep sshd listening broadly and gate at the packet layer (belt rule), exactly as
§5. The endpoint forward rule is the primary gate; the box belt is defense-in-depth (both proven
independent in §5).
6. **Measured latencies (production budget input).** Operator→box forwarded RTT **~24 ms** (One
Hungary ↔ Hetzner-DE); interactive SSH incl. `pct list` **1.2 s** wall; self-heal repair **~15 s**;
endpoint peer-push convergence **~150 s** (≤5 min bound). All comfortably interactive.
7. **Reconcile continuity constraint (from P0, load-bearing for any future OOB probe/op).** The box
applies its conf every 60 s. Any OOB mechanism that mutates box wg state MUST go through the
agent's rendered conf (finding 1), never a side-channel `wg set`/hand-edit — those are silently
reverted within a tick. This is a *feature* (drift-repair) to build with, not around.
8. **Endpoint-access + production prerequisites discovered.** (a) The demo endpoint is a real public
VM, so O-A is faithfully represented; production `ep0` needs the **AAAA fixed** (standing
OPERATOR item) though the v4-pin means OOB rides v4 regardless. (b) `cloudflared` in the guest is
a **Docker container, not a systemd unit** — any runbook step that "stops cloudflared" must
`docker stop` it, not `systemctl`. (c) The hub `wg_peers` insert path for a NON-host operator
peer exists in the store (`AddWGPeer` with empty host_id — the S1 unbound-admin-row shape); a
production OOB feature can reuse it or add an explicit operator-peer admin endpoint.
---
## 9. What this spike did NOT prove (honest ledger)
- **True CGNAT** OOB traversal — the box is single-NAT public-v4 (same caveat as the prior spike
§5). OOB rides the same outbound WG mapping as the backup path, so low-risk, but unproven on
`100.64/10`.
- **A second real agent box** — the dummy customer was a netns peer (sufficient for the topology
isolation negative; not a full agent). The box↔box drop is proven by topology + counters, which
is the load-bearing property.
- **The rendered-conf operator-`/32` change itself** — finding 1 is the *diagnosis* (runtime `wg
set` wiped by self-heal, re-add restored OOB); the agent code change to render it is TASK work,
not shipped here.
- **Endpoint peersync owning the forward rules** — the spike put forward rules in ad-hoc nft; the
TASK decides static-vs-hub-driven (finding 3 recommends static forward rules + hub-driven peer
list).
- **Long-horizon hold** of an idle OOB session (minutes-scale only here; the backup-path 32-min
soak is in doc 06 §7).
---
## 10. Cleanup assertion (§13 of the plan) — VERIFIED
**Hub:** both spike `wg_peers` rows deleted (`WHERE note LIKE 'spike-oob-%'`); registry back to the
single real peer `demo-felhom-01 / 10.77.0.2`.
**Endpoint `felhom-hetzner`:** operator+dummy peers removed from live `wg0`; the 5-min reconciler
then rewrote `wg0.conf` to the box-only peer (drift-repair confirmed live — the persisted conf
self-healed); `felhom_spike_oob` forward chain + jump rule deleted; `ip_forward` restored to **0**
(the P0 value); `spike-op`/`spike-cust` netns + veths deleted; throwaway WG + SSH keys `shred -u`'d,
`/root/spike-oob` removed. Final: `nft` forward chain is the empty drop-policy chain (P0), `wg0`
lists only the box, `ip_forward=0`.
**Box `felhom-pve`:** runtime AllowedIPs back to `10.77.0.1/32` (operator `/32` removed) + return
route deleted; `felhom_spike_oob` host table deleted (`nft list tables` empty, as P0); the
`spike-oob-operator` line removed from `/root/.ssh/authorized_keys` (grep count 0); `/tmp/wgdown_t0`
removed. Final baseline re-verified: `felhom-agent` active, `wg-quick@wg-felhom` active + handshake
fresh (age 92 s), `pvesm status felhom-offsite` active, guest 9201 running + controller healthy, no
host nft tables, LAN SSH fine.
**No repo/hub/agent/manifest production change** — this commit is docs-only.
---
## 11. Method bar honored
`wg show <if> dump` NEVER run (private-key leak ban, S1); only `latest-handshakes` + redacted conf
reads. Every mutation was a named one-flush nft table or a runtime-only `wg set`, inventoried as
made. PBS-path re-checked after the box wg change (P2). Negative probes carry nft-counter evidence,
not just timeouts. Surprises recorded as findings (the self-heal wipe of the runtime `/32` — finding
1 — is the headline example).