Files
felhom.eu/documentation/audits/SPIKE-firstboot-nic-sweep-2026-07-22.md
T
admin c043ba9abe docs(spike): R-59/R-60 Phase-0 — the NIC sweep mechanics are GO, measured on a nested wrong-NIC install
Nested VM 9310 on felhom-pve (throwaway vmbr9 rig): re-point
bridge-ports -> ifreload (0.36s, raises the port itself) -> dhclient
lease 3.3s -> hub HTTPS 302 in 160ms; a dead candidate costs exactly
its timeout bound (rc=124 at 20s). Bonus finding: the auto-installer
picks the NIC that leases and bakes the lease as STATIC — so wrong-NIC
boxes come in two variants (192.168.100.2 fallback bake vs
truth-changed-after-install), and the gate must trigger on
hub-unreachable, never on the fallback signature. Eight design inputs
(F-P1..F-P8) recorded for the Part-3 gate. STOP-0 discharged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UuFPHmHNrCJj1VhY6QdDMU
2026-07-22 10:50:43 +02:00

105 lines
7.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# SPIKE — first-boot NIC sweep mechanics (R-59/R-60 Phase-0), 2026-07-22
**Question:** on a Proxmox box that installed against the wrong NIC (the demo-hp shape: no DHCP on
any cable-bearing port at install time → the installer bakes its **192.168.100.2 static fallback**
into `vmbr0` and completes), does this sequence work, and is it bounded:
*re-point `vmbr0`'s `bridge-ports` to a candidate NIC → reload networking → bounded DHCP on vmbr0 →
hub HTTPS probe*?
**Verdict up front: GO.** The whole success leg is ~4 s; a dead candidate costs exactly its DHCP
bound (20 s in the probe). Details and the design inputs below.
**Vehicle:** PVE-in-a-VM on felhom-pve (nested VM 9310, throwaway), installed from the v1.24.0
pairing ISO (uncommitted `nested-probe` profile = `nested-vm` + the ops SSH key for dead-network
access). Guest: PVE 9.2, `ifupdown2 3.3.0-1+pmx12`, `isc-dhcp-client 4.4.3-P1-8`. NICs: `ens18`
(net0) / `ens19` (net1).
**Access-path deviation (recorded):** the probe needs a shell on a box whose network is by
definition dead. The plan was to use the installer's fallback address as the way in (host carries
`192.168.100.1/24` on the port-less throwaway `vmbr9`, `ssh root@192.168.100.2` over the dead
bridge, root key = the probe profile's baked ops key). The run then took a different door (below),
but the vmbr9 address was mounted and stays part of the Part-4 drill plan.
## Reproduction — and an unplanned second trap variant
The plan was both NICs dark from `qm start`. In practice the tap re-point to the dead bridge landed
~12 min after start, and **the installer had already leased on `ens19` (192.168.0.141) in that
window** — after which the flip cut it off. What that accident bought:
- **The auto-installer picks the NIC that leases, not the first by name.** `bridge-ports ens19`
(net1, the second NIC) was baked, with the lease written as a **STATIC** stanza
(`address 192.168.0.141/24`, `gateway 192.168.0.1`). So the t740's "picked wrong" is not
name-order bad luck: at that site the 4-port card was the only cabled thing, no NIC leased, and
the installer fell back. Two distinct wrong-NIC install variants therefore exist:
1. **no lease anywhere → 192.168.100.2 fallback baked static** (demo-hp, live-proven 2026-07-21;
README t740 trap) — R-59's detection signature;
2. **a lease existed at install time but the truth changed after** (cable moved / switch port
died) → a plausible-looking static config that is nonetheless dead — no 192.168.100.2 to key
on. **The gate must therefore key on "hub unreachable", not on the fallback signature**; the
signature is a diagnosis detail for the screen/log, not the trigger.
- **The silent surface is real and was watched live for 30 minutes:** `felhom-bootstrap` in the
pairing loop logged `registration did not go through (no network yet?) — retrying` every ~42 s
(≈12 s curl DNS failure + 30 s sleep), unit `activating`, console showing only the login prompt.
Nothing on the physical screen names the problem — exactly R-59's complaint.
## Probe log — the sweep walked by hand (detached; `/root/probe-sweep.log`)
Shape at probe time (the "cable move"): `vmbr0` configured for `ens19` = dead (tap on vmbr9),
`ens18` = live (tap on vmbr0, site DHCP + internet available). Script scheduled via
`systemd-run --on-active=45s`, both taps flipped, box left alone.
| Step | Result | Wall time |
|---|---|---|
| Leg 1: `timeout 20 dhclient -1 -v vmbr0` on the DEAD shape | rc=124 (killed by `timeout`; dhclient never gives up on its own) | **20.00 s** (the full bound, exactly) |
| Rewrite `bridge-ports ens19``ens18` (tmp + `mv`, original kept) | one-line diff | ~10 ms |
| `ifreload -a` | rc=0; **raised the admin-DOWN candidate itself**, port `forwarding` immediately (STP off, fd 0) | **0.36 s** |
| `timeout 30 dhclient -1 -v vmbr0` | DISCOVER→OFFER→REQUEST→ACK, bound 192.168.0.245 | **3.32 s** |
| `curl -fsS --max-time 15 https://hub.felhom.eu/` | HTTP **302**, rc=0 | **0.16 s** |
Post state: box reachable again at its baked static address over the new port (the static stanza
was deliberately not touched — the probe changed only `bridge-ports`); dhclient's lease sits as a
second address on vmbr0; default route unchanged.
## Findings (design inputs for Part 3)
- **F-P1 — the sequence works and is tightly bounded.** Success ≈ 4 s; failed candidate = its DHCP
bound + ε. With a 20 s bound: worst case ≈ `NICs × ~21 s` + one hub-probe timeout each. GO.
- **F-P2 — unused physical NICs are admin-DOWN and their carrier is unreadable while down**
(`/sys/class/net/*/carrier` reads empty). A carrier-based candidate filter must
`ip link set <nic> up` + settle (~2 s) before reading. `ifreload -a` brings the *configured*
bridge-port up on its own, so the sweep proper needs no manual link handling — only the
*diagnosis table* (NIC/MAC/cable) needs the raise-first step.
- **F-P3 — carrier is necessary but not sufficient** (virtio on a dead bridge shows carrier=1; a
cabled port on a router with a dead uplink would too). Carrier selects/orders candidates; only
DHCP-lease + hub-probe decides. Fine — that is the R-60 design anyway.
- **F-P4 — `dhclient -1` must be externally bounded** (`timeout N dhclient -1`); on a dead segment
it retries past any useful window (rc=124 at whatever bound is set). After a killed attempt,
`pkill` dhclient before the next candidate — the second run then starts clean (proven: leg 2's
dhclient ran without pidfile/lease complaints after leg 1's kill + pkill).
- **F-P5 — treat ANY HTTP response as "hub reachable".** The hub answers `/` with a 302; `-f`
only fails ≥ 400, and even a 4xx would prove TLS+HTTP reachability. The gate's probe should key
on "got an HTTP status" (curl rc 0, or rc 22 with a status code), not on 200.
- **F-P6 — the trigger is "hub unreachable", never the 192.168.100.2 signature** (see variant 2
above). The signature is named in the log/screen when present, as the likely cause.
- **F-P7 — persisting only `bridge-ports` is enough to restore reachability in variant 2** (the
static address stays valid when the cable moved ports). In variant 1 (fallback bake) the address
itself is garbage — after a candidate wins, the persisted config must also carry the working
addressing. Simplest shape that covers both: persist `bridge-ports <winner>` + `iface vmbr0 inet
dhcp` (the winner is by definition a leasing NIC — DHCP is what proved it). Note for the R-59
screen loop: a box that self-healed onto DHCP addressing keeps working across router reboots;
the operator can still pin it later.
- **F-P8 (observation, out of scope):** the pairing loop's register-failure line prints every
cycle (~42 s) — the v1.21.0 quieting covered the 204 wait but not the no-network register path.
The Part-3 gate makes this moot on the no-network path (the gate holds before register), but it
is worth knowing the line exists.
## Cleanup state after the probe
Nested VM 9310 **stopped** (destroyed in Part 4 after the drill re-uses the rig); `vmbr9` +
host-side `192.168.100.1/24` still mounted for Part 4; felhom-pve `/etc/network/interfaces`
untouched throughout (imperative `ip link` bridge only; pre-task copy sha
`c7ace087183cb72e031ad0cd1974b9239d539747528dd19523d993f8e9c4ec7a` — final diff happens at Part-4
cleanup). The probe box self-registered at the production hub as an unclaimed appliance while its
network was up (expected; the record can be discarded, or will simply be superseded by the Part-4
drill box's registration).