3172df1927
Closes N100 F1 (HIGH): cheap AMI (AN3PLUS 0.01-class) UEFI firmware can't relocate the ISO's stock signed GRUB from USB (relocation 0x0). The run's live grub-mkimage workaround is now a first-class pipeline mode. - build-felhom-iso.sh: --loader shim|mkimage (default shim, byte-for-byte unchanged; profile-settable FELHOM_LOADER; --loader wins). Loud banner + manifest loader:/grub-mkimage: fields + -mkimage filename suffix. - mkimage-surgery.sh (new): post-prepare-iso, in the assistant container. Builds a monolithic grub-mkimage loader from the ISO's own GRUB (module set from its grub.cfg; embedded search --fs-uuid -> configfile the real menu). Swaps it into the ISO9660 tree (real lowercase path) + the efi.img ESP; xorriso re-master preserves BIOS-hybrid + UEFI + GPT-ESP, drops only Apple HFS+/APM. Recipe from the N100 run evidence, not re-derived. - Dockerfile.assistant: grub-common + grub-efi-amd64-bin + mtools + dosfstools. profiles/n100.profile (new, mkimage + SB-off note). - Validated on nested VM 311 (RUNBOOK-B legs): leg1 shim boots+installs under OVMF SB-enforcing + SeaBIOS; leg2 mkimage boots+installs under SB-off; leg3 (red-proof) mkimage under SB-enforcing FAILS Access Denied (unsigned -> SB must be OFF); leg4 surgery byte-identical payload. bash -n + shellcheck clean. Physical N100 closure folds into the rehearsal (n100-safety match-nothing ISO built + sha-recorded, unbooted). PXE stays a deferred R-21 note.
135 lines
14 KiB
Markdown
135 lines
14 KiB
Markdown
# VALIDATION — N100 bare-metal reinstall via Felhom ISO (R-21 physical run + onboarding rehearsal), 2026-07-16
|
|
|
|
> Supervised run: Viktor at the box + keyboard, CC driving Phase-0 prep and the SSH-side legs. The
|
|
> demo host (felhom-pve / the N100 serving demo-felhom.eu) was reinstalled clean-slate from a
|
|
> pipeline-built Felhom ISO, then walked through the friend-alpha onboarding path with Viktor as
|
|
> customer zero. **No production code changed** — this document + the ISO build artifacts (kept on
|
|
> DooPlex `~/felhom-iso/`, `~/n100-baremetal/`) are the outputs. Findings will be tackled separately.
|
|
>
|
|
> **Headline — read this first.** The R-21 pipeline itself was **flawless**: the ISO was correct, and
|
|
> the full unattended chain (install → enroll → provision → escrow) reached **rc-0 on the first
|
|
> attempt (zero retries)** — the exact terminal success that was operator-gated in slice A, now
|
|
> **proven on real, customer-shaped hardware**. The *one hard obstacle* was **not** our software: this
|
|
> cheap N100's early AMI firmware **cannot UEFI-boot GRUB from USB** (`relocation 0x0 is not
|
|
> implemented yet`) — the single class of thing nested virt could never surface. We worked around it
|
|
> live by rebuilding the stick's loader from the box's own known-good GRUB. Everything downstream
|
|
> worked; the onboarding surfaced a crop of **reused-customer / clean-slate edges** (the Peti-pattern
|
|
> data the run was designed to produce) plus a couple of UX bugs. Net: **core objectives green,
|
|
> ~7 findings to work.**
|
|
|
|
---
|
|
|
|
## Run header
|
|
|
|
| | |
|
|
|---|---|
|
|
| Hardware | Intel **N100** (AlderLake-N), 16 GB; AMI **Aptio** BIOS `AN3PLUS 0.01` (build 2024-06-11), UEFI 2.8, no CSM/Legacy. **DMI all "Default string"** (mfr/product/serial/baseboard/chassis); system UUID `001E91E9-…` real |
|
|
| Disks | internal SATA M.2 SSD **AirDisk 512 GB** serial `QDF922W009654S30EX` (target); external USB HDD **Toshiba MQ04ABF100** (ADATA HD710 PRO enclosure) serial `65NOP3HDT`, 1 TB — held the Felhom **secondary backups**; + a 64 GB SD card + the 119 GB Samsung boot stick |
|
|
| Customer | `demo-felhom` (Demo Ügyfél), domain **demo-felhom.eu**, email admin@felhom.eu, config MANAGED, created 145 d ago (previously onboarded) |
|
|
| ISO | `felhom-pve-9.2-1-v1.16.0-n100-demo.iso` (pipeline `scripts/iso/`, secret-bearing), profile: `filter.ID_SERIAL_SHORT="QDF922W009654S30EX"`, appliance mode, my ed25519 key baked (durable post-wipe access) |
|
|
| Outcome | **SUCCESS** — clean-slate reinstall, host enrolled, guest 9201 provisioned, BookStack deployed, demo-felhom.eu live. Install wall-clock ≈ 3 min; first-boot → rc-0 ≈ a few min |
|
|
|
|
## What succeeded (the core objectives)
|
|
|
|
- **Pipeline correctness on metal.** The ISO built by `build-felhom-iso.sh` installed and ran the
|
|
first-boot chain exactly as designed. `felhom-bootstrap` fetched `felhom-host-install.sh` from the
|
|
public `felhom.eu/scripts/` channel and ran it to **rc-0 on the first try (`NRestarts=0`)**: done-flag
|
|
written, unit disabled, `/etc/felhom/bootstrap.env` **shredded**. This **closes slice A's
|
|
operator-gated boundary on real hardware.**
|
|
- **Serial-filter safety — proven on real hardware.** The install targeted only the SSD; the external
|
|
USB HDD's canary file was **byte-identical** before/after (`sha256 de6e00cb…`, host-verified). The
|
|
HDD even moved device node (`sdd`→`sdb`) across the reinstall — irrelevant, because we filter by
|
|
**serial**, which is exactly why it was safe. (Match-nothing / wrong-disk fail-safe was also proven
|
|
through the pipeline earlier in slice A, Scenario D.)
|
|
- **Auto-installer over a complete prior install.** It did **not** abort on the pre-existing LVM — it
|
|
wiped and installed. The spike's "prior-LVM abort" (SPIKE-baremetal-iso §S2b) was an artifact of a
|
|
*half*-wiped disk; a **clean, complete prior PVE install is simply overwritten.** (So the runbook's
|
|
manual `blkdiscard`/STOP-2 step was not needed here.)
|
|
- **PBS-DR reconciler self-healed on the reused peer.** Hub showed all four steps `done`: host
|
|
enrolled (`demo-felhom-01`), WG tunnel peer registered, descriptor provisioned (namespace
|
|
`demo-felhom`, token `felhom@pbs!demo-felhom`), **key escrow present / ceremony done** — the S8
|
|
headline, first real firing on physical customer-shaped hardware.
|
|
- **UEFI + Secure Boot** was disabled during troubleshooting; the installed system boots fine from the
|
|
SSD. (The SB-enforcing path was proven in the spike on reference firmware.)
|
|
- **DMI verdict for slice C — settled on real cheap hardware:** every DMI identity string is
|
|
**"Default string"**; only the **system UUID** and **NIC MAC** are usable. Slice-C keying must be
|
|
**MAC + UUID**, never DMI serials.
|
|
|
|
## The one hard obstacle — GRUB won't UEFI-boot from USB (firmware)
|
|
|
|
Booting the stick showed only:
|
|
```
|
|
relocation 0x0 is not implemented yet
|
|
Aborted. Press any key to exit.
|
|
```
|
|
and fell through to the internal Proxmox. Diagnosis ladder (all still safe — pre-wipe):
|
|
1. **Secure Boot disabled** → same error (so it is *not* an SB/shim-verification issue). BIOS confirmed
|
|
SB `Disabled / Not Active`.
|
|
2. **Bypassed shim** (swapped the stick's `\EFI\BOOT\BOOTX64.EFI` from shim → the ISO's `grubx64.efi`)
|
|
→ **same error** (so it is *not* shim; it is **GRUB** failing — the message is GRUB's).
|
|
3. Our ISO is **correct** — the identical prepared image boots in reference UEFI (OVMF), and the box's
|
|
*installed* GRUB `2.12-9+pmx2` boots fine **from the SSD**. So: this early AMI firmware cannot
|
|
relocate the **ISO's (signed/shim-oriented) GRUB build when loaded from USB**.
|
|
|
|
**Workaround that fixed it (done live over SSH from the still-bootable old system):** rebuilt the
|
|
stick's `BOOTX64.EFI` with `grub-mkimage` **from the box's own working `2.12-9+pmx2` GRUB**, embedding
|
|
every module the ISO's `grub.cfg` needs (so nothing loads from the ISO's problem build) and an embedded
|
|
config that `search --fs-uuid`es the ISO and `configfile`s its real menu. **That booted straight into
|
|
the installer.** The installed system's SSD GRUB then booted normally (SSD path always worked here).
|
|
|
|
This is the highest-value **slice-B input**: some cheap boards need a firmware-compatible boot loader.
|
|
Options to fold into the pipeline — (a) build `BOOTX64.EFI` with `grub-mkimage` at ISO-prep time
|
|
instead of shipping the shim-oriented grub, or (b) ship a documented field-fix (the exact `grub-mkimage`
|
|
recipe used here), or (c) offer a **PXE/network-boot** path (the assistant's `--pxe` artifacts) for
|
|
boards where USB-grub is broken.
|
|
|
|
## Findings
|
|
|
|
| # | Sev | Finding | Root cause | Disposition / fix |
|
|
|---|-----|---------|-----------|-------------------|
|
|
| **F1** | **HIGH** | ISO won't UEFI-boot GRUB from USB on this AMI `AN3PLUS 0.01` firmware (`relocation 0x0…`) | firmware can't relocate the ISO's signed GRUB from USB; SB-off and shim-bypass don't help | **worked around live** (self-built `grub-mkimage` loader from the box's own GRUB). **PIPELINE-FIXED in scripts v1.18.0 (2026-07-17):** `build-felhom-iso.sh --loader mkimage` bakes the monolithic grub-mkimage loader in as a first-class mode (recipe reproduced from this run's evidence, not re-derived); `profiles/n100.profile` uses it. Validated on nested VM 311 — mkimage boots + auto-installs under OVMF **Secure Boot OFF** (leg 2), and under **SB enforcing FAILS** with firmware `Access Denied` (leg 3, red-proof) → **the loader is unsigned, so the target board's Secure Boot must be OFF** (documented). **Physical closure on the real AMI board still pending** — it folds into the supervised N100 rehearsal (an optional zero-risk `n100-safety` match-nothing ISO is built + sha-recorded for a pre-flight). |
|
|
| **F2** | MEDIUM | No claim-code email on reinstall of an existing customer | claim state (`claim_code_generation:2`, hash, issued 2026-07-13) is **hub/customer-level** and is delivered to the fresh box; an existing code ⇒ no re-issue/re-email. Fresh box has **no password set** | need a **"re-issue claim code"** operator action (bump generation + email). ~~**⚠ also verify** whether the reinstalled *unclaimed* box is properly gated or accidentally **open** (F-4 class)~~ **ERRATUM 2026-07-16 (Viktor):** the ⚠ is RETRACTED — the claim gate WAS presented at felhom.demo-felhom.eu; the customer self-served a new code, claimed, and set a password. F2 is a continuity/UX gap, not a gating hole. **SHIPPED hub v0.57.0** — `claim.ReissueForReenroll` auto-issues a reset code on clean-slate re-enroll (host-enroll mint path). |
|
|
| **F3** | MEDIUM | Offsite target missing on the fresh controller → escrow blocked | offsite transient password is "delivered to the controller **once**" — it went to the *old* box; the fresh controller never got it. Hub showed provisioned + escrow-done → **hub/controller desync** | **"Re-issue offsite credentials"** in the hub restaged it (done during the run; controller picks up next config refresh). Codify: reinstall must re-issue offsite. **SHIPPED hub v0.57.0** — the re-enroll mint path calls the same machinery (`ReissueOffsiteForCustomer`) automatically. |
|
|
| **F4** | MEDIUM | PBS-DR read 403s every tick | agent token `felhom-agent@pve!agent` has `FelhomAgentStore` on `/storage/**felhom-pbs**` only, but the customer's **PVE STORAGE ID is `felhom-offsite`** (non-default, the demo's adopted manual entry) → `GET /storage/felhom-offsite -> 403 (missing Datastore.Allocate)` | ~~install ACL must grant on the **config's storage id**~~ **ERRATUM/DISPOSITION 2026-07-16:** an installer fix is **not feasible** — the DR storage id lives in the agent-domain **pbs_dr descriptor** (`web/pbsdr.go` `StorageID`), provisioned *after* WG registration, so `step_agent_config()` cannot know it at ACL-grant time. The real block is a bootstrap circularity: the agent's reconcile tick does a **token-auth** `GET /storage/<id>` pre-check that 403s and aborts **before** its own root-run `felhom-pbs-apply grant` sets the ACL. Root fix is **agent-side** (proceed to the root-run apply despite the pre-check 403, or run the pre-check as root) — logged as a ROADMAP agent-train item; the demo was unblocked live with a one-shot `pveum` grant on `/storage/felhom-offsite`. Every default-storage-id (all new/Peti installs) already works — F4 only bites non-default ids. |
|
|
| **F5** | MEDIUM | Guest RAM = 2 GB (too low on a 16 GB host) | **golden default**; no `--memory` passed (appliance), no per-customer/host-aware sizing | make guest RAM/cores configurable (hub config or auto-size at provision). Currently only the install `--memory` flag exists and isn't surfaced | **FIXED — host-install v1.17.0 (2026-07-17):** appliance mode auto-sizes RAM=clamp(host-4096, min 4096, max host-2048, ceil host-1024) + cores=host-1 (min 2) when no explicit cap; explicit `--memory`/`--cores` win. Harness red-proof (8/16/32 GB). Live at the next from-scratch rehearsal. |
|
|
| **F6** | MEDIUM | Drive "initialize" formats but doesn't mount/attach; UI shows nothing | on confirm, the **format client disconnects** ("mkfs continues detached; poll GET /disks/format/status") — mkfs completes (`storage: formatted device /dev/sdb ext4`) but the **post-mkfs mount+register is aborted** and the UI never polls the status | users reasonably expect *initialize* to also mount+register. Fix the confirm→poll flow so the wizard finishes the mount+attach and reports progress/completion. (Viktor recovered by re-running the *attach* flow manually → `hdd_1` at `/mnt/hdd_1`) | **FIXED — controller v0.141.0 (2026-07-17):** init runs as a DETACHED job (survives disconnect) the wizard polls (`GET /api/storage/init/status`, 3-step progress); a slow mkfs is followed via the agent's `/disks/format/status`; register is marker-last. Live-validated on a 64 GB scratch USB → mounted+registered at `/mnt/felhom-drives/scratch1`. |
|
|
| **F7** | LOW | "Vissza" (Back) on `/storage/init` and `/storage/attach` routes to `/settings` | frontend route bug | point Back → `/storage` | **FIXED — controller v0.141.0 (2026-07-17):** both Vissza anchors → `/storage` (test-guarded). |
|
|
|
|
**Non-findings / notes:** the HDD initially not showing under "attach" was **CC's leftover read-only
|
|
canary mount** (fixed, not a product bug). BIOS **State After G3 = S5** — the box stays *off* after a
|
|
power cut; for a headless server this should be "Power On / Last State" (set before final sign-off).
|
|
Hub "Containers 0/0 → 2/2" was report lag, not a bug.
|
|
|
|
## Current state (end of run)
|
|
|
|
- felhom-pve = **PVE 9.2.2** on the SSD; **agent 0.88.0**; up ~30 min; boots normally from the SSD.
|
|
- Guest **9201 running** (demo-felhom): `felhom-controller` (0.136.0), `traefik`, `cloudflared`,
|
|
`filebrowser`, **`bookstack` + `bookstack-db`** — demo-felhom.eu live.
|
|
- Storage: `local` + `local-lvm` (system, protected); **`hdd_1`** = the external HDD, reformatted ext4
|
|
at `/mnt/hdd_1` (old secondary backups **gone**, per the reformat choice). SD card + Samsung stick
|
|
still attached.
|
|
- Offsite: re-issued (staged, picking up). **DR tier blocked on F4 (403)** until the ACL is granted.
|
|
- Onboarding not fully completed: **claim/password not set** (F2), **escrow ceremony not run** (F3/F4).
|
|
- The Samsung boot stick still carries the ISO with the CC-built GRUB workaround (secret-bearing —
|
|
wipe at teardown per the runbook).
|
|
|
|
## Evidence
|
|
|
|
DooPlex `~/n100-baremetal/pre/` (G0.2 harvest: dmidecode, udev per-disk, efibootmgr, canary-before,
|
|
summary with the rulings + identity), `~/n100-baremetal/evidence/`, `~/felhom-iso/out/` (the ISO +
|
|
sha/manifest, secret-bearing, 0600). BIOS/console photos + hub screenshots captured in the session.
|
|
Console error, GRUB-fix steps, and the rc-0 chain journal are quoted inline above.
|
|
|
|
## Recommendations (to tackle next)
|
|
|
|
1. **F1 (slice B):** decide the boot-loader strategy — pipeline-built `grub-mkimage` `BOOTX64.EFI`,
|
|
documented field recipe, and/or a PXE path — so R-21 works on cheap boards, not just reference UEFI.
|
|
2. **F4:** grant the DR ACL on the config's storage id (quick unblock) and fix the installer to derive
|
|
the ACL storage set from the customer's PVE STORAGE ID.
|
|
3. **F2 + F3 (Peti-pattern, R-1):** the **reinstall-of-an-existing-customer** path needs first-class
|
|
support — re-issue claim code (+ verify gating), and re-issue offsite/PBS on reprovision. This is
|
|
directly on Peti's convergence path.
|
|
4. **F6:** finish the drive-initialize wizard (mount+attach + status polling).
|
|
5. **F5 / F7 / G3-power-on:** guest RAM configurability, the Back-route fix, and set the BIOS power-on
|
|
policy.
|