Files
felhom.eu/documentation/tests/VALIDATION-n100-baremetal-2026-07-16.md
T
admin 3172df1927 scripts v1.18.0 — R-21 slice B: firmware loader option --loader shim|mkimage (F1)
Closes N100 F1 (HIGH): cheap AMI (AN3PLUS 0.01-class) UEFI firmware can't
relocate the ISO's stock signed GRUB from USB (relocation 0x0). The run's live
grub-mkimage workaround is now a first-class pipeline mode.

- build-felhom-iso.sh: --loader shim|mkimage (default shim, byte-for-byte
  unchanged; profile-settable FELHOM_LOADER; --loader wins). Loud banner +
  manifest loader:/grub-mkimage: fields + -mkimage filename suffix.
- mkimage-surgery.sh (new): post-prepare-iso, in the assistant container. Builds
  a monolithic grub-mkimage loader from the ISO's own GRUB (module set from its
  grub.cfg; embedded search --fs-uuid -> configfile the real menu). Swaps it into
  the ISO9660 tree (real lowercase path) + the efi.img ESP; xorriso re-master
  preserves BIOS-hybrid + UEFI + GPT-ESP, drops only Apple HFS+/APM. Recipe from
  the N100 run evidence, not re-derived.
- Dockerfile.assistant: grub-common + grub-efi-amd64-bin + mtools + dosfstools.
  profiles/n100.profile (new, mkimage + SB-off note).
- Validated on nested VM 311 (RUNBOOK-B legs): leg1 shim boots+installs under
  OVMF SB-enforcing + SeaBIOS; leg2 mkimage boots+installs under SB-off; leg3
  (red-proof) mkimage under SB-enforcing FAILS Access Denied (unsigned -> SB must
  be OFF); leg4 surgery byte-identical payload. bash -n + shellcheck clean.
  Physical N100 closure folds into the rehearsal (n100-safety match-nothing ISO
  built + sha-recorded, unbooted). PXE stays a deferred R-21 note.
2026-07-17 14:28:10 +02:00

14 KiB

VALIDATION — N100 bare-metal reinstall via Felhom ISO (R-21 physical run + onboarding rehearsal), 2026-07-16

Supervised run: Viktor at the box + keyboard, CC driving Phase-0 prep and the SSH-side legs. The demo host (felhom-pve / the N100 serving demo-felhom.eu) was reinstalled clean-slate from a pipeline-built Felhom ISO, then walked through the friend-alpha onboarding path with Viktor as customer zero. No production code changed — this document + the ISO build artifacts (kept on DooPlex ~/felhom-iso/, ~/n100-baremetal/) are the outputs. Findings will be tackled separately.

Headline — read this first. The R-21 pipeline itself was flawless: the ISO was correct, and the full unattended chain (install → enroll → provision → escrow) reached rc-0 on the first attempt (zero retries) — the exact terminal success that was operator-gated in slice A, now proven on real, customer-shaped hardware. The one hard obstacle was not our software: this cheap N100's early AMI firmware cannot UEFI-boot GRUB from USB (relocation 0x0 is not implemented yet) — the single class of thing nested virt could never surface. We worked around it live by rebuilding the stick's loader from the box's own known-good GRUB. Everything downstream worked; the onboarding surfaced a crop of reused-customer / clean-slate edges (the Peti-pattern data the run was designed to produce) plus a couple of UX bugs. Net: core objectives green, ~7 findings to work.


Run header

Hardware Intel N100 (AlderLake-N), 16 GB; AMI Aptio BIOS AN3PLUS 0.01 (build 2024-06-11), UEFI 2.8, no CSM/Legacy. DMI all "Default string" (mfr/product/serial/baseboard/chassis); system UUID 001E91E9-… real
Disks internal SATA M.2 SSD AirDisk 512 GB serial QDF922W009654S30EX (target); external USB HDD Toshiba MQ04ABF100 (ADATA HD710 PRO enclosure) serial 65NOP3HDT, 1 TB — held the Felhom secondary backups; + a 64 GB SD card + the 119 GB Samsung boot stick
Customer demo-felhom (Demo Ügyfél), domain demo-felhom.eu, email admin@felhom.eu, config MANAGED, created 145 d ago (previously onboarded)
ISO felhom-pve-9.2-1-v1.16.0-n100-demo.iso (pipeline scripts/iso/, secret-bearing), profile: filter.ID_SERIAL_SHORT="QDF922W009654S30EX", appliance mode, my ed25519 key baked (durable post-wipe access)
Outcome SUCCESS — clean-slate reinstall, host enrolled, guest 9201 provisioned, BookStack deployed, demo-felhom.eu live. Install wall-clock ≈ 3 min; first-boot → rc-0 ≈ a few min

What succeeded (the core objectives)

  • Pipeline correctness on metal. The ISO built by build-felhom-iso.sh installed and ran the first-boot chain exactly as designed. felhom-bootstrap fetched felhom-host-install.sh from the public felhom.eu/scripts/ channel and ran it to rc-0 on the first try (NRestarts=0): done-flag written, unit disabled, /etc/felhom/bootstrap.env shredded. This closes slice A's operator-gated boundary on real hardware.
  • Serial-filter safety — proven on real hardware. The install targeted only the SSD; the external USB HDD's canary file was byte-identical before/after (sha256 de6e00cb…, host-verified). The HDD even moved device node (sddsdb) across the reinstall — irrelevant, because we filter by serial, which is exactly why it was safe. (Match-nothing / wrong-disk fail-safe was also proven through the pipeline earlier in slice A, Scenario D.)
  • Auto-installer over a complete prior install. It did not abort on the pre-existing LVM — it wiped and installed. The spike's "prior-LVM abort" (SPIKE-baremetal-iso §S2b) was an artifact of a half-wiped disk; a clean, complete prior PVE install is simply overwritten. (So the runbook's manual blkdiscard/STOP-2 step was not needed here.)
  • PBS-DR reconciler self-healed on the reused peer. Hub showed all four steps done: host enrolled (demo-felhom-01), WG tunnel peer registered, descriptor provisioned (namespace demo-felhom, token felhom@pbs!demo-felhom), key escrow present / ceremony done — the S8 headline, first real firing on physical customer-shaped hardware.
  • UEFI + Secure Boot was disabled during troubleshooting; the installed system boots fine from the SSD. (The SB-enforcing path was proven in the spike on reference firmware.)
  • DMI verdict for slice C — settled on real cheap hardware: every DMI identity string is "Default string"; only the system UUID and NIC MAC are usable. Slice-C keying must be MAC + UUID, never DMI serials.

The one hard obstacle — GRUB won't UEFI-boot from USB (firmware)

Booting the stick showed only:

relocation 0x0 is not implemented yet
Aborted. Press any key to exit.

and fell through to the internal Proxmox. Diagnosis ladder (all still safe — pre-wipe):

  1. Secure Boot disabled → same error (so it is not an SB/shim-verification issue). BIOS confirmed SB Disabled / Not Active.
  2. Bypassed shim (swapped the stick's \EFI\BOOT\BOOTX64.EFI from shim → the ISO's grubx64.efi) → same error (so it is not shim; it is GRUB failing — the message is GRUB's).
  3. Our ISO is correct — the identical prepared image boots in reference UEFI (OVMF), and the box's installed GRUB 2.12-9+pmx2 boots fine from the SSD. So: this early AMI firmware cannot relocate the ISO's (signed/shim-oriented) GRUB build when loaded from USB.

Workaround that fixed it (done live over SSH from the still-bootable old system): rebuilt the stick's BOOTX64.EFI with grub-mkimage from the box's own working 2.12-9+pmx2 GRUB, embedding every module the ISO's grub.cfg needs (so nothing loads from the ISO's problem build) and an embedded config that search --fs-uuides the ISO and configfiles its real menu. That booted straight into the installer. The installed system's SSD GRUB then booted normally (SSD path always worked here).

This is the highest-value slice-B input: some cheap boards need a firmware-compatible boot loader. Options to fold into the pipeline — (a) build BOOTX64.EFI with grub-mkimage at ISO-prep time instead of shipping the shim-oriented grub, or (b) ship a documented field-fix (the exact grub-mkimage recipe used here), or (c) offer a PXE/network-boot path (the assistant's --pxe artifacts) for boards where USB-grub is broken.

Findings

# Sev Finding Root cause Disposition / fix
F1 HIGH ISO won't UEFI-boot GRUB from USB on this AMI AN3PLUS 0.01 firmware (relocation 0x0…) firmware can't relocate the ISO's signed GRUB from USB; SB-off and shim-bypass don't help worked around live (self-built grub-mkimage loader from the box's own GRUB). PIPELINE-FIXED in scripts v1.18.0 (2026-07-17): build-felhom-iso.sh --loader mkimage bakes the monolithic grub-mkimage loader in as a first-class mode (recipe reproduced from this run's evidence, not re-derived); profiles/n100.profile uses it. Validated on nested VM 311 — mkimage boots + auto-installs under OVMF Secure Boot OFF (leg 2), and under SB enforcing FAILS with firmware Access Denied (leg 3, red-proof) → the loader is unsigned, so the target board's Secure Boot must be OFF (documented). Physical closure on the real AMI board still pending — it folds into the supervised N100 rehearsal (an optional zero-risk n100-safety match-nothing ISO is built + sha-recorded for a pre-flight).
F2 MEDIUM No claim-code email on reinstall of an existing customer claim state (claim_code_generation:2, hash, issued 2026-07-13) is hub/customer-level and is delivered to the fresh box; an existing code ⇒ no re-issue/re-email. Fresh box has no password set need a "re-issue claim code" operator action (bump generation + email). ⚠ also verify whether the reinstalled unclaimed box is properly gated or accidentally open (F-4 class) ERRATUM 2026-07-16 (Viktor): the ⚠ is RETRACTED — the claim gate WAS presented at felhom.demo-felhom.eu; the customer self-served a new code, claimed, and set a password. F2 is a continuity/UX gap, not a gating hole. SHIPPED hub v0.57.0claim.ReissueForReenroll auto-issues a reset code on clean-slate re-enroll (host-enroll mint path).
F3 MEDIUM Offsite target missing on the fresh controller → escrow blocked offsite transient password is "delivered to the controller once" — it went to the old box; the fresh controller never got it. Hub showed provisioned + escrow-done → hub/controller desync "Re-issue offsite credentials" in the hub restaged it (done during the run; controller picks up next config refresh). Codify: reinstall must re-issue offsite. SHIPPED hub v0.57.0 — the re-enroll mint path calls the same machinery (ReissueOffsiteForCustomer) automatically.
F4 MEDIUM PBS-DR read 403s every tick agent token felhom-agent@pve!agent has FelhomAgentStore on /storage/**felhom-pbs** only, but the customer's PVE STORAGE ID is felhom-offsite (non-default, the demo's adopted manual entry) → GET /storage/felhom-offsite -> 403 (missing Datastore.Allocate) install ACL must grant on the config's storage id ERRATUM/DISPOSITION 2026-07-16: an installer fix is not feasible — the DR storage id lives in the agent-domain pbs_dr descriptor (web/pbsdr.go StorageID), provisioned after WG registration, so step_agent_config() cannot know it at ACL-grant time. The real block is a bootstrap circularity: the agent's reconcile tick does a token-auth GET /storage/<id> pre-check that 403s and aborts before its own root-run felhom-pbs-apply grant sets the ACL. Root fix is agent-side (proceed to the root-run apply despite the pre-check 403, or run the pre-check as root) — logged as a ROADMAP agent-train item; the demo was unblocked live with a one-shot pveum grant on /storage/felhom-offsite. Every default-storage-id (all new/Peti installs) already works — F4 only bites non-default ids.
F5 MEDIUM Guest RAM = 2 GB (too low on a 16 GB host) golden default; no --memory passed (appliance), no per-customer/host-aware sizing make guest RAM/cores configurable (hub config or auto-size at provision). Currently only the install --memory flag exists and isn't surfaced
F6 MEDIUM Drive "initialize" formats but doesn't mount/attach; UI shows nothing on confirm, the format client disconnects ("mkfs continues detached; poll GET /disks/format/status") — mkfs completes (storage: formatted device /dev/sdb ext4) but the post-mkfs mount+register is aborted and the UI never polls the status users reasonably expect initialize to also mount+register. Fix the confirm→poll flow so the wizard finishes the mount+attach and reports progress/completion. (Viktor recovered by re-running the attach flow manually → hdd_1 at /mnt/hdd_1)
F7 LOW "Vissza" (Back) on /storage/init and /storage/attach routes to /settings frontend route bug point Back → /storage

Non-findings / notes: the HDD initially not showing under "attach" was CC's leftover read-only canary mount (fixed, not a product bug). BIOS State After G3 = S5 — the box stays off after a power cut; for a headless server this should be "Power On / Last State" (set before final sign-off). Hub "Containers 0/0 → 2/2" was report lag, not a bug.

Current state (end of run)

  • felhom-pve = PVE 9.2.2 on the SSD; agent 0.88.0; up ~30 min; boots normally from the SSD.
  • Guest 9201 running (demo-felhom): felhom-controller (0.136.0), traefik, cloudflared, filebrowser, bookstack + bookstack-db — demo-felhom.eu live.
  • Storage: local + local-lvm (system, protected); hdd_1 = the external HDD, reformatted ext4 at /mnt/hdd_1 (old secondary backups gone, per the reformat choice). SD card + Samsung stick still attached.
  • Offsite: re-issued (staged, picking up). DR tier blocked on F4 (403) until the ACL is granted.
  • Onboarding not fully completed: claim/password not set (F2), escrow ceremony not run (F3/F4).
  • The Samsung boot stick still carries the ISO with the CC-built GRUB workaround (secret-bearing — wipe at teardown per the runbook).

Evidence

DooPlex ~/n100-baremetal/pre/ (G0.2 harvest: dmidecode, udev per-disk, efibootmgr, canary-before, summary with the rulings + identity), ~/n100-baremetal/evidence/, ~/felhom-iso/out/ (the ISO + sha/manifest, secret-bearing, 0600). BIOS/console photos + hub screenshots captured in the session. Console error, GRUB-fix steps, and the rc-0 chain journal are quoted inline above.

Recommendations (to tackle next)

  1. F1 (slice B): decide the boot-loader strategy — pipeline-built grub-mkimage BOOTX64.EFI, documented field recipe, and/or a PXE path — so R-21 works on cheap boards, not just reference UEFI.
  2. F4: grant the DR ACL on the config's storage id (quick unblock) and fix the installer to derive the ACL storage set from the customer's PVE STORAGE ID.
  3. F2 + F3 (Peti-pattern, R-1): the reinstall-of-an-existing-customer path needs first-class support — re-issue claim code (+ verify gating), and re-issue offsite/PBS on reprovision. This is directly on Peti's convergence path.
  4. F6: finish the drive-initialize wizard (mount+attach + status polling).
  5. F5 / F7 / G3-power-on: guest RAM configurability, the Back-route fix, and set the BIOS power-on policy.