Files
felhom.eu/REPORT.md
T
admin 7fffde3e13 R-50 runbook: Part C (agent 0.95.0 -> demo-hp) + qm300 forensic
- deployed agent 0.95.0 to demo-hp via break-glass (0.93.0->0.95.0), binary
  sha-verified, caps 68/68 ok, guest 9201 untouched, hub confirms 0.95.0
- forensic: qm300 had no qmdestroy; it died with the mid-July N100 reprovision
- nodes.md: fleet agents 0.95.0; demo-hp deploy note
2026-07-25 12:25:04 +02:00

5.2 KiB
Raw Blame History

REPORT — R-50 island-bridge drill + empirical spike (RUNBOOK, 2026-07-25 PM)

What ran

Provisioned the first nested-PVE drill appliance on the t740 (demo-hp) and ran the R-50 island-bridge empirical spike end-to-end. Verdict: GO.

Part A — drill provisioned through the REAL day-0

  • VM 300 drill-r50 on demo-hp: 8 GiB / 4 vCPU (cpu=host) / 32 GiB local-lvm / OVMF (SB off) / one NIC on vmbr0 DHCP. Nested-virt already enabled on the host (A2 no-op; no QEMU VM was running at A1).
  • Installed from felhom-pve-9.2-1-v1.25.0-nested-vm-generic-mkimage.iso (current release; the RUNBOOK's stated v1.22.0 was both stale and the wrong profile — v1.22.0-nested-canary is a deliberate match-nothing installer that aborts touching no disk; corrected to the v1.25.0 nested-vm profile. Friction finding, not a day-0 defect).
  • Day-0 drove itself: install → first-boot → self-register (appliance #9, pairing 9MB-4QX) → operator bind to a scratch customer drill-r50 (minimal config, no DNS/offsite/PBS side-effects) → deliver → nested guest 9201 provisioned, controller stack healthy (agent 0.93.0, golden 0.161.0).
  • R-50 starting condition reached & confirmed: agent listen_addr = 192.168.0.176:8443 and guest bootstrap.json local_api.endpoint = 192.168.0.176:8443 — the F1 LAN literal baked in both places.
  • Snapshot r50pre taken (clean day-0). Break-glass vaulted host_recovery/drill-r50-0a4f9a.
  • Day-0 tail not driven: claim is customer-email-gated (scratch customer has no inbox; forcing it needs a real email + a production Resend send) — left PENDING; escrow is N/A (no DR/PBS tier on the minimal drill). Neither gates the control-plane probes. Operator decision if the full customer-side tail should be exercised (needs an email address + DR-tier choice).

Part B — probes P1P8, ALL PASS (verdict GO)

  • P2 vmbr9 portless island bridge 169.254.253.1/30 — vmbr0 untouched, LAN gw reachable.
  • P3 guest island NIC eth1 169.254.253.2/30 hot-added — eth0/LAN undisturbed, bidirectional island ping.
  • P4 (dnsmasq trap + fix) — with host_ip unset, moving listen_addr to the island rebound dnsmasq to 169.254.253.1:53 and killed LAN DNS; lan_resolver.host_ip=192.168.0.176 restored it while the API bind stayed on the island. Finding 1 confirmed + fixed, live.
  • P5 (pin) — served leaf SHA-256 over the island identical (4ef1d953…), authenticated GET /storage from the guest over the island → HTTP 200. No cert re-issue (Finding 6 live).
  • P6 (F1 replay — money shot) — LAN 192.168.0.176→.200, listen_addr on the island → agent stays active + serves (200). LAN-literal contrast reproduced the 2026-07-20 bug verbatim (localapi: bind 192.168.0.176:8443 → daemon exit 1).
  • P7 (survival) — agent restart / guest reboot / host cold reboot all return the control plane on the island with zero intervention (vmbr9, island bind, dnsmasq on LAN IP, guest autostart, controller).
  • Method caveat: probes drove the runtime chain via manual config edits; the provisioning path (host-install + golden-bake bootstrap template) is the implementation task — now fully de-risked.

Docs updated

  • documentation/audits/SPIKE-island-bridge-2026-07-25.md — verdict flipped BLOCKED → GO, empirical P1P8 results added, old BLOCKED/NOT-RUN sections marked superseded.
  • documentation/backlog/ROADMAP.md — R-50 → SPIKED → GO; next = Phase A/B/C production spec.
  • documentation/operations/nodes.md — the drill VM 300 now documented on the t740 (access, snapshot, teardown).

State left behind

Drill VM 300 left running in the working island configuration (snapshot r50pre preserves clean day-0). Scratch hub records drill-r50 (customer + appliance #9) remain — throwaway; teardown noted in nodes.md. demo-hp host networking, guest 9201, and felhom-pve were untouched.

Part C — agent 0.95.0 deployed to demo-hp (DONE)

0.93.0 → 0.95.0 via the sanctioned method (green gate build/vet/test pass; configs/ unchanged since 0.93.0 so binary-only). scp via break-glass → backup .bak-0.93.0install -m0755systemctl restart. Verified: deployed sha256 == local build (b5fa7c5d…), version 0.95.0, service active, capabilities 68/68 ok, 0 degraded, local-api listening, all watchdogs up, no ERROR/panic; guest 9201 untouched/running; hub confirms demo-hp reporting 0.95.0. (Note: the hub's vouched Day-0 artifact still reads 0.93.0 — the manifest vouch is a separate operator step, not needed for a running host; flag for follow-up if new installs should get 0.95.0.) Both fleet nodes now run 0.95.0.

Forensic — where did qm300 go? (DONE, read-only)

On felhom-pve (demo-felhom, PVE 9.2.2): no discrete qmdestroy. No 300.conf, no vm-300 LVs, no :300: in any task log (incl. archives), qm list empty; task history only reaches back to ~2026-07-16. qm300 (the old drill VM) vanished when the N100 was fully reprovisioned/reborn mid-July — the whole node (task DB + LVM) was recreated from scratch, taking qm300 with it. felhom-pve left untouched.