R-50 island-bridge: SPIKED -> GO (probes P1-P8 pass live on t740 drill)

Provisioned nested-PVE drill 'drill-r50' (qm300 on demo-hp) via the v1.25.0
nested-vm ISO through the real day-0, then ran the R-50 empirical spike:
- vmbr9 portless island bridge + guest island NIC hot-add (LAN undisturbed)
- F1 replay money shot: LAN move survives on the island; LAN-literal bind
  reproduces the 2026-07-20 daemon-exit bug verbatim
- dnsmasq trap confirmed live + lan_resolver.host_ip fix proven
- pin address-independent (leaf SHA-256 unchanged, HTTP 200 over island)
- survival matrix: agent/guest/host-cold-reboot all return on the island

Docs: SPIKE verdict BLOCKED->GO, ROADMAP R-50 SPIKED->GO, nodes.md drill VM,
REPORT overwrite.
This commit is contained in:
2026-07-25 12:19:09 +02:00
parent 5d56f93755
commit 515e0c3cc7
4 changed files with 117 additions and 61 deletions
@@ -8,11 +8,60 @@ fixed private addresses, so the control plane survives any router/lease/site mov
**validate the whole chain end-to-end on the DRILL environment (qm300 / `demo-vm-felhom-2f4b00`)**
before any production spec exists.
## VERDICT UP FRONT: SPIKE STILL BLOCKED (empirically) — no probeable drill environment exists.
## VERDICT UP FRONT: **GO** — validated end-to-end on a live drill (2026-07-25, evening).
**Two empirical attempts, both blocked (2026-07-25). Design half below stands; GO/NO-GO PENDING.**
**A nested-PVE drill appliance was provisioned on the t740 (`demo-hp`) through the REAL day-0 pipeline,
and probes P1P8 ALL PASS.** The island-bridge control plane works end-to-end, survives the exact F1
failure that motivated the row, the pin is address-independent as predicted, the dnsmasq trap is
confirmed live AND its fix proven, and the whole topology survives a cold host reboot untouched.
**Recommendation: GO — write the production implementation spec (Phase A/B/C skeleton at the bottom).**
See **“## Empirical validation (2026-07-25 PM) — probes P1P8”** below. The two earlier BLOCKED attempts
(no drill existed) are retained further down for provenance.
### Empirical attempt 2 (2026-07-25, PM) — the t740 was checked; there is no drill VM there.
### Method caveat (honest scope)
The probes drove the mechanism via **manual, surgical config edits** on the drill (append a `vmbr9`
stanza + `ifreload`, `pct set -net1`, hand-edit `agent.json` `listen_addr`/`lan_resolver.host_ip` and
the host-side `bootstrap.json` endpoint, restart). They prove the **runtime chain** is sound and
reboot-durable. They do **not** exercise the *provisioning* path that a real rollout needs
(`felhom-host-install.sh` writing these at install/golden-bake time) — that is Phase A of the impl spec,
still a future task. What is now de-risked: every assumption the spec rests on is empirically true.
## Empirical validation (2026-07-25 PM) — probes P1P8, ALL PASS
**Drill:** nested-PVE appliance `drill-r50` (QEMU VM **300** on `demo-hp`/t740), installed from the
current release ISO `felhom-pve-9.2-1-v1.25.0-nested-vm-generic-mkimage.iso` through the **real day-0**
(self-register → operator bind → deliver → guest provision). Reached the exact R-50 starting condition:
agent `local_api.listen_addr = 192.168.0.176:8443` (the LAN literal) **and** the guest's
`bootstrap.json local_api.endpoint = 192.168.0.176:8443` — the F1 literal baked in both places. Nested
guest **9201** runs the controller (agent 0.93.0, golden 0.161.0). Access to the drill-PVE:
break-glass root via hub `host_recovery/drill-r50-0a4f9a`, reached through `demo-hp`.
(Day-0 tail: **claim** left PENDING — customer-email-gated, no inbox on a scratch customer; **escrow**
N/A — no DR/PBS tier on the minimal drill. Neither gates the control-plane probes.)
| Probe | What it validates | Result |
|---|---|---|
| **P1** snapshot | rollback point before mutation | **PASS**`qm snapshot 300 r50pre` (retained) |
| **P2 (3)** `vmbr9` create | portless host bridge, non-disruptive | **PASS**`vmbr9 169.254.253.1/30` up via `ifreload -a`; **vmbr0 untouched** (`.176/24`), LAN gw still reachable, guest kept running |
| **P3 (4)** guest NIC hot-add | live hotplug, LAN leg undisturbed | **PASS**`pct set 9201 -net1 …ip=169.254.253.2/30`; `eth1` up on the island, **`eth0`/LAN undisturbed** (`192.168.0.15/24`), host↔guest island ping both ways |
| **P4 (7)** dnsmasq trap **+ fix** | the Finding-1 coupling, live | **PASS (both)** — with `host_ip` **unset**, moving `listen_addr` to the island rebound dnsmasq to **`169.254.253.1:53`** and **LAN DNS `192.168.0.176:53` died**; setting `lan_resolver.host_ip=192.168.0.176` put dnsmasq **back on `192.168.0.176:53`** while the API bind stayed on the island |
| **P5 (6)** pin | pin survives the address move | **PASS** — served leaf SHA-256 over the island **identical** (`4ef1d953fe…f219bab`); authenticated `GET https://169.254.253.1:8443/storage` from the guest → **HTTP 200**. No cert re-issue. Finding 6 confirmed live |
| **P6 (5)** **F1 replay** | control plane survives a LAN move | **PASS — the money shot.** LAN moved `192.168.0.176 → .200`, `listen_addr` left on the island → agent **stays `active`**, still bound `169.254.253.1:8443`, control plane **HTTP 200**. **Contrast (original bug reproduced):** set `listen_addr` back to the now-absent `.176``level=ERROR "daemon: exited with error" err="localapi: bind 192.168.0.176:8443: listen…"`, systemd `status=1/FAILURE`, nothing bound — the 2026-07-20 incident, verbatim |
| **P7 (8)** survival matrix | topology is reboot-durable | **PASS****agent restart**: island bind + 200; **guest reboot**: `eth1` island NIC + controller + 200 returned; **host (qm300) COLD reboot**: `vmbr9`, agent island bind, dnsmasq on the LAN IP, guest autostart (`onboot=1`) with `eth1`, controller healthy, control plane **HTTP 200****all with zero intervention** |
**All source-grounded findings are now empirically confirmed:** the F1 double-bake (agent bind + guest
dial) moves atomically and works; the pin is DER-based and address-transparent (Finding 6); the dnsmasq
trap is real and its `lan_resolver.host_ip = LAN IP` fix works (Finding 1); link-local `169.254.253.x/30`
is a clean host-internal fabric that no LAN move perturbs. **The drill was left in the working island
configuration** (not rolled back); snapshot `r50pre` preserves the clean LAN-literal day-0 if a re-run
is wanted.
### GO/NO-GO: **GO.** Proceed to the implementation spec (Phase A/B/C below). No blocker remains; the one
residual is that the *provisioning* path (host-install writing the island config + golden-bake carrying
the new bootstrap template) is unbuilt — that IS the impl task, now de-risked.
---
### Empirical attempt 2 (2026-07-25, PM) — the t740 was checked; there is no drill VM there. — **SUPERSEDED by the validation above (a drill was then provisioned and probed).**
The operator ruled (2026-07-25) that **drill + build VMs are hosted on the HP t740 (`demo-hp`) from
now on**, so this half was re-attempted there. Discovery (read-only, no LAN scan): the t740 = `demo-hp`
(Tailscale `100.76.96.79` / LAN `192.168.0.87`), accessed via the hub-vaulted G1 break-glass root
@@ -56,7 +105,7 @@ next drill run) start from a real base rather than zero.
---
## Probes 38 — NOT RUN (drill env absent)
## Probes 38 — NOT RUN (drill env absent) — **SUPERSEDED: all now PASS, see “Empirical validation (2026-07-25 PM)” above.**
| Probe | Would validate | Status |
|---|---|---|
File diff suppressed because one or more lines are too long
+14 -8
View File
@@ -65,14 +65,20 @@ No trace of `192.168.100.2` remains. `wg-felhom` `10.77.0.3/32` up to the hub. G
them off DooPlex (the production k3s node). **This is a VM-HOSTING ruling only; the build-PIPELINE
relocation to the t740 is NOT ruled or implemented here.**
**Current state (verified 2026-07-25):** the ruling is **forward-looking and not yet realized.** The
t740 hosts **no drill/build VM yet**`qm list` is empty, `/etc/pve/qemu-server/` is empty, it is a
standalone PVE node running only its own LXC guest 9201. The historical drill appliance is still
`drill.qcow2` on **DooPlex** (`/mnt/5_hdd/felhom.eu/drill/`, ~18G, powered off — a golden-bake VM whose
nested guest is purged after each bake). felhom-pve likewise has no drill VM. **A nested-PVE drill VM
(agent + a nested guest) must be provisioned on the t740** before the R-50 island-bridge empirical spike
(`audits/SPIKE-island-bridge-2026-07-25.md`) can run its P2P7 probes — that spike is blocked on exactly
this. The t740's ~30 GB RAM / ~49 GB free local-lvm suit it as the drill host.
**Current state (updated 2026-07-25 PM):** the ruling is **realized** — the t740 now hosts the first
drill appliance. **QEMU VM `300` = `drill-r50`**, a nested PVE-in-a-VM (8 GiB RAM, 4 vCPU cpu=host, 32 GiB
local-lvm disk, OVMF/SB-off, one NIC on `vmbr0` DHCP). Installed from
`felhom-pve-9.2-1-v1.25.0-nested-vm-generic-mkimage.iso` through the **real day-0** (self-register →
operator bind to scratch customer `drill-r50` → deliver → nested guest **9201** provisioned, controller
0.161.0 / agent 0.93.0 healthy). Its own break-glass root is vaulted in the hub
`host_recovery/drill-r50-0a4f9a`; reach it as `root@192.168.0.176` **through `demo-hp`** (it has no key
and no tailnet — it is a peer on demo-hp's LAN). It has snapshot **`r50pre`** (clean LAN-literal day-0).
This VM was provisioned to unblock and run the **R-50 island-bridge empirical spike**
(`audits/SPIKE-island-bridge-2026-07-25.md`) — **probes P1P8 PASSED 2026-07-25, verdict GO**; the drill
was left in its working island configuration. It is a throwaway: destroy with `qm stop 300 && qm destroy
300 --purge 1` (and delete the `drill-r50` customer + appliance #9 hub-side) when no longer needed. The
historical golden-bake `drill.qcow2` still lives on **DooPlex** (`/mnt/5_hdd/felhom.eu/drill/`, ~18G,
powered off) and is unrelated. The build-PIPELINE relocation to the t740 remains unbuilt.
### Access — there is no baked SSH key