Files
felhom.eu/documentation/audits/SPIKE-firstboot-nic-sweep-2026-07-22.md
admin a12c6f9730 docs(v1.24.0): R-59/R-60/R-61 SHIPPED — CHANGELOG, README, ROADMAP (+R-62), runbook, capability map, drill evidence, REPORT
Virgin-ISO nested drill closed the train: dead-NIC install baked the
fallback (incl. the dead default gateway), the R-59 screen painted
(capture committed beside the spike doc), the cable move healed +
registered at the hub in 23s unaided, and the build's rootpw file
matched the installed box's shadow hash. R-59 SHIPPED with the recorded
deviation (first-boot gate; installer-initrd abort out of scope by
operator ack). R-60 SHIPPED (spike + drill cited; F-P9 route-flush fix
included). R-61 slice 1 SHIPPED. New R-62 row (hub delete-dialog
cosmetics, XS). Capability map: new PROVEN-LIVE row (nested != metal,
said so). Cleanup verified: felhom-pve interfaces byte-identical,
bridge/VMs/ISO removed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UuFPHmHNrCJj1VhY6QdDMU
2026-07-22 11:45:55 +02:00

134 lines
10 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# SPIKE — first-boot NIC sweep mechanics (R-59/R-60 Phase-0), 2026-07-22
**Question:** on a Proxmox box that installed against the wrong NIC (the demo-hp shape: no DHCP on
any cable-bearing port at install time → the installer bakes its **192.168.100.2 static fallback**
into `vmbr0` and completes), does this sequence work, and is it bounded:
*re-point `vmbr0`'s `bridge-ports` to a candidate NIC → reload networking → bounded DHCP on vmbr0 →
hub HTTPS probe*?
**Verdict up front: GO.** The whole success leg is ~4 s; a dead candidate costs exactly its DHCP
bound (20 s in the probe). Details and the design inputs below.
**Vehicle:** PVE-in-a-VM on felhom-pve (nested VM 9310, throwaway), installed from the v1.24.0
pairing ISO (uncommitted `nested-probe` profile = `nested-vm` + the ops SSH key for dead-network
access). Guest: PVE 9.2, `ifupdown2 3.3.0-1+pmx12`, `isc-dhcp-client 4.4.3-P1-8`. NICs: `ens18`
(net0) / `ens19` (net1).
**Access-path deviation (recorded):** the probe needs a shell on a box whose network is by
definition dead. The plan was to use the installer's fallback address as the way in (host carries
`192.168.100.1/24` on the port-less throwaway `vmbr9`, `ssh root@192.168.100.2` over the dead
bridge, root key = the probe profile's baked ops key). The run then took a different door (below),
but the vmbr9 address was mounted and stays part of the Part-4 drill plan.
## Reproduction — and an unplanned second trap variant
The plan was both NICs dark from `qm start`. In practice the tap re-point to the dead bridge landed
~12 min after start, and **the installer had already leased on `ens19` (192.168.0.141) in that
window** — after which the flip cut it off. What that accident bought:
- **The auto-installer picks the NIC that leases, not the first by name.** `bridge-ports ens19`
(net1, the second NIC) was baked, with the lease written as a **STATIC** stanza
(`address 192.168.0.141/24`, `gateway 192.168.0.1`). So the t740's "picked wrong" is not
name-order bad luck: at that site the 4-port card was the only cabled thing, no NIC leased, and
the installer fell back. Two distinct wrong-NIC install variants therefore exist:
1. **no lease anywhere → 192.168.100.2 fallback baked static** (demo-hp, live-proven 2026-07-21;
README t740 trap) — R-59's detection signature;
2. **a lease existed at install time but the truth changed after** (cable moved / switch port
died) → a plausible-looking static config that is nonetheless dead — no 192.168.100.2 to key
on. **The gate must therefore key on "hub unreachable", not on the fallback signature**; the
signature is a diagnosis detail for the screen/log, not the trigger.
- **The silent surface is real and was watched live for 30 minutes:** `felhom-bootstrap` in the
pairing loop logged `registration did not go through (no network yet?) — retrying` every ~42 s
(≈12 s curl DNS failure + 30 s sleep), unit `activating`, console showing only the login prompt.
Nothing on the physical screen names the problem — exactly R-59's complaint.
## Probe log — the sweep walked by hand (detached; `/root/probe-sweep.log`)
Shape at probe time (the "cable move"): `vmbr0` configured for `ens19` = dead (tap on vmbr9),
`ens18` = live (tap on vmbr0, site DHCP + internet available). Script scheduled via
`systemd-run --on-active=45s`, both taps flipped, box left alone.
| Step | Result | Wall time |
|---|---|---|
| Leg 1: `timeout 20 dhclient -1 -v vmbr0` on the DEAD shape | rc=124 (killed by `timeout`; dhclient never gives up on its own) | **20.00 s** (the full bound, exactly) |
| Rewrite `bridge-ports ens19``ens18` (tmp + `mv`, original kept) | one-line diff | ~10 ms |
| `ifreload -a` | rc=0; **raised the admin-DOWN candidate itself**, port `forwarding` immediately (STP off, fd 0) | **0.36 s** |
| `timeout 30 dhclient -1 -v vmbr0` | DISCOVER→OFFER→REQUEST→ACK, bound 192.168.0.245 | **3.32 s** |
| `curl -fsS --max-time 15 https://hub.felhom.eu/` | HTTP **302**, rc=0 | **0.16 s** |
Post state: box reachable again at its baked static address over the new port (the static stanza
was deliberately not touched — the probe changed only `bridge-ports`); dhclient's lease sits as a
second address on vmbr0; default route unchanged.
## Findings (design inputs for Part 3)
- **F-P1 — the sequence works and is tightly bounded.** Success ≈ 4 s; failed candidate = its DHCP
bound + ε. With a 20 s bound: worst case ≈ `NICs × ~21 s` + one hub-probe timeout each. GO.
- **F-P2 — unused physical NICs are admin-DOWN and their carrier is unreadable while down**
(`/sys/class/net/*/carrier` reads empty). A carrier-based candidate filter must
`ip link set <nic> up` + settle (~2 s) before reading. `ifreload -a` brings the *configured*
bridge-port up on its own, so the sweep proper needs no manual link handling — only the
*diagnosis table* (NIC/MAC/cable) needs the raise-first step.
- **F-P3 — carrier is necessary but not sufficient** (virtio on a dead bridge shows carrier=1; a
cabled port on a router with a dead uplink would too). Carrier selects/orders candidates; only
DHCP-lease + hub-probe decides. Fine — that is the R-60 design anyway.
- **F-P4 — `dhclient -1` must be externally bounded** (`timeout N dhclient -1`); on a dead segment
it retries past any useful window (rc=124 at whatever bound is set). After a killed attempt,
`pkill` dhclient before the next candidate — the second run then starts clean (proven: leg 2's
dhclient ran without pidfile/lease complaints after leg 1's kill + pkill).
- **F-P5 — treat ANY HTTP response as "hub reachable".** The hub answers `/` with a 302; `-f`
only fails ≥ 400, and even a 4xx would prove TLS+HTTP reachability. The gate's probe should key
on "got an HTTP status" (curl rc 0, or rc 22 with a status code), not on 200.
- **F-P6 — the trigger is "hub unreachable", never the 192.168.100.2 signature** (see variant 2
above). The signature is named in the log/screen when present, as the likely cause.
- **F-P7 — persisting only `bridge-ports` is enough to restore reachability in variant 2** (the
static address stays valid when the cable moved ports). In variant 1 (fallback bake) the address
itself is garbage — after a candidate wins, the persisted config must also carry the working
addressing. Simplest shape that covers both: persist `bridge-ports <winner>` + `iface vmbr0 inet
dhcp` (the winner is by definition a leasing NIC — DHCP is what proved it). Note for the R-59
screen loop: a box that self-healed onto DHCP addressing keeps working across router reboots;
the operator can still pin it later.
- **F-P8 (observation, out of scope):** the pairing loop's register-failure line prints every
cycle (~42 s) — the v1.21.0 quieting covered the 204 wait but not the no-network register path.
The Part-3 gate makes this moot on the no-network path (the gate holds before register), but it
is worth knowing the line exists.
## Part-4 drill — the shipped v1.24.0 ISO, end to end (same day)
Two drill rounds on the same rig (both NICs on the dead `vmbr9` from `qm create`, so no lease
exists at ANY point of the install — the strict demo-hp shape):
**Round 1 (VM 9311, first v1.24.0 build) — reproduced the trap AND caught a real gate defect.**
The installer baked the fallback (`address 192.168.100.2/24` + **`gateway 192.168.100.1`** +
first NIC) and completed; the gate diagnosed, swept, painted the R-59 screen and retried — but
after the cable move the heal FAILED and the gate oscillated: the sweep's candidate leased fine,
yet the hub probe rode the **baked stale default route** (dhclient-script never replaces an
existing default route) and failed, so the sweep dutifully restored byte-identically every cycle.
**F-P9 (fix, shipped in `felhom-bootstrap.sh` before the final build): flush vmbr0's
addresses+routes after the re-point, before the bounded dhclient — a candidate is judged on the
lease's OWN addressing/routing.** Fix validated by hot-patching 9311 live: heal + hub
registration 33 s after unit restart (pairing banner with code on the console).
**Round 2 (VM 9312, virgin final ISO `c25cdbbf…`, zero patching) — the full claim:**
| Leg | Evidence |
|---|---|
| (a) dead-NIC install reproduces | installer env on 192.168.100.2 at 11:32; installed box: static fallback + dead gateway + `bridge-ports ens18` baked |
| R-59 screen renders | console screendump committed beside this doc (`SPIKE-firstboot-nic-sweep-2026-07-22-screen.png`): header, NIC table (név/MAC/kábel/sebesség), fallback-signature explanation naming 192.168.100.2, remedy line |
| failed-sweep restore is safe | after full sweep cycles (with F-P9 flushes) the runtime had `192.168.100.2/24` back on vmbr0 — `ifreload` re-applies the flushed pristine config; interfaces file byte-shape intact |
| (b) sweep to the moved cable, unaided | net1 → vmbr0 at 11:35:27; `network self-heal: vmbr0 -> ens19` + **hub registration at 11:35:50 — 23 s, no human action**; persisted `iface vmbr0 inet dhcp` + `bridge-ports ens19`, original in `interfaces.felhom-bak`; runtime 192.168.0.88 with a correct default route |
| (c) unclaimed appliance at the hub | `registered — appliance token stored (0600)`, pairing code delivered + banner on console; unit `activating` in the normal pairing wait |
| R-61 lockout closed | the build's `.rootpw.txt` password recomputed against the installed box's `/etc/shadow` root hash (`openssl passwd -6 -salt <shadow-salt>`) → **exact match**; comparison done on the build host, the plaintext never left it |
## Cleanup (verified at the end of the run)
Nested VMs 9310/9311/9312 destroyed (`--purge`); `vmbr9` deleted (its temporary
`192.168.100.1/24` went with it); the transferred ISO and screendump temp files removed from
felhom-pve; `brctl show` shows vmbr0 with exactly its original ports (`enp1s0`, `veth9201i0`);
**felhom-pve `/etc/network/interfaces` diffed byte-identical against the pre-task copy** (sha
`c7ace087183cb72e031ad0cd1974b9239d539747528dd19523d993f8e9c4ec7a`) — the file was never touched
(the throwaway bridge lived only as an imperative `ip link` object). Residue by design: **three
unclaimed-appliance registrations at the production hub** (the probe box, the hot-patched drill
box, the virgin drill box — all three machines destroyed); the operator can discard them from the
Unclaimed list at leisure.