docs: R-59/R-60/R-61 from the HP install + second-hardware pairing proof
R-59 no-DHCP install must hard-abort (it baked 192.168.100.2 static and completed - a box that can never call home). R-60 first-boot NIC sweep self-heal. R-61 the baked root password must be knowable; a fixed well-known password is explicitly rejected. Positive evidence same-session: R-21 slice C PROVEN on a SECOND, virgin board (HP t740) - and the shim loader booted with Secure Boot ENABLED, retiring the assumption that Felhom installs need SB off. Fresh-box floor lift 0.153.0 -> 0.156.0 during day-0 cited on the publish-train row. ISO README gains the t740 five-NIC trap: the 4-port igb card gets no lease, the onboard r8169 port does.
This commit is contained in:
@@ -31,7 +31,7 @@
|
||||
|---|---|---|---|---|
|
||||
| Appliance day-0 install: golden image → first boot → auto-confirm (zero clicks) → claimable box | installer, agent, hub, golden | **PROVEN-LIVE** (nested VM) | `DRILL-day0-vm-2026-07-12`, `DRILL-day0-take2-2026-07-12` | First firing on real customer hardware pending → R-1 |
|
||||
| BYO install: `--mode byo`, mandatory caps, host-mutation disclosure, coexistence guards | installer v1.15+, agent | **PARTIAL** | `DRILL-GL6-2026-07-08` (demo box); GL-8 coexistence fixes | Peti clean-slate reinstall on proxmox2 is the first real BYO run of the current path → R-1 |
|
||||
| Bare-metal Felhom ISO (blank hardware → zero-touch auto-install → first-boot `host-install`); selectable UEFI loader; **universal secret-free / operator-bind** mode | scripts v1.19.0 (`scripts/iso/`) + hub v0.62.0 + assistant container | **PROVEN-LIVE** (physical N100, one pass, 2026-07-18) | **`tests/VALIDATION-n100-rehearsal-2026-07-18.md` — the full chain on real metal in a single pass:** the generic reusable pairing ISO (v1.20.0, `--loader mkimage`, SB off) booted the cheap AMI board that F1 had blocked, installed unattended, and the box **self-registered as an unclaimed appliance at 16:17:14 — the same second it first booted** (`appliance_registrations` id=3), then bound → credential-delivered → day-0 SUCCESS 16:32:32 → floor-lifted to current. **F1 is closed on physical hardware.** Prior nested legs: slice A `SPIKE-baremetal-iso-2026-07-16` (build gate, disk-filter fail-safe, stub→host-install fetch); slice B RUNBOOK-B (shim boots+installs OVMF SB-enforcing + SeaBIOS; `--loader mkimage` boots+installs SB-off; mkimage SB-enforcing **FAILS** `Access Denied`; surgery byte-identical); **slice C (2026-07-17): the GENERIC secret-free ISO** — box self-registers as an unclaimed appliance (`POST /api/v1/appliance/register`, one-shot poll delivery, 404-no-oracle — all live-verified through the public ingress), operator binds on the Hosts page, hub delivers credentials once; bootstrap harness proves direct(zero-appliance-calls)/pairing/delivery; artifact proven secret-free (baked env = hub URL only) | **F1 loader caveat:** `--loader mkimage` fixes cheap AMI firmware that can't USB-boot the stock GRUB — UNSIGNED → **Secure Boot must be OFF**; default `shim` keeps SB. **Slice C bind is operator-password-gated** (CC stages, Viktor binds) → the live boot→register→bind→day-0 composition + physical N100 boot fold into the supervised rehearsal (R-1). Customer-facing **self-bind page = R-27 slice 1 SHIPPED (hub v0.66.0, 2026-07-17)** — see the dedicated self-bind row |
|
||||
| Bare-metal Felhom ISO (blank hardware → zero-touch auto-install → first-boot `host-install`); selectable UEFI loader; **universal secret-free / operator-bind** mode | scripts v1.19.0 (`scripts/iso/`) + hub v0.62.0 + assistant container | **PROVEN-LIVE on TWO different boards** (N100 2026-07-18; HP t740 2026-07-21) | **`tests/VALIDATION-n100-rehearsal-2026-07-18.md` — the full chain on real metal in a single pass:** the generic reusable pairing ISO (v1.20.0, `--loader mkimage`, SB off) booted the cheap AMI board that F1 had blocked, installed unattended, and the box **self-registered as an unclaimed appliance at 16:17:14 — the same second it first booted** (`appliance_registrations` id=3), then bound → credential-delivered → day-0 SUCCESS 16:32:32 → floor-lifted to current. **F1 is closed on physical hardware.** Prior nested legs: slice A `SPIKE-baremetal-iso-2026-07-16` (build gate, disk-filter fail-safe, stub→host-install fetch); slice B RUNBOOK-B (shim boots+installs OVMF SB-enforcing + SeaBIOS; `--loader mkimage` boots+installs SB-off; mkimage SB-enforcing **FAILS** `Access Denied`; surgery byte-identical); **slice C (2026-07-17): the GENERIC secret-free ISO** — box self-registers as an unclaimed appliance (`POST /api/v1/appliance/register`, one-shot poll delivery, 404-no-oracle — all live-verified through the public ingress), operator binds on the Hosts page, hub delivers credentials once; bootstrap harness proves direct(zero-appliance-calls)/pairing/delivery; artifact proven secret-free (baked env = hub URL only) | **F1 loader caveat:** `--loader mkimage` fixes cheap AMI firmware that can't USB-boot the stock GRUB — UNSIGNED → **Secure Boot must be OFF**; default `shim` keeps SB. **Slice C bind is operator-password-gated** (CC stages, Viktor binds) → the live boot→register→bind→day-0 composition + physical N100 boot fold into the supervised rehearsal (R-1). Customer-facing **self-bind page = R-27 slice 1 SHIPPED (hub v0.66.0, 2026-07-17)** — see the dedicated self-bind row | **Second board, 2026-07-21 (demo-hp, HP t740 / Ryzen V1756B / AMI M42):** the whole chain ran on virgin hardware in one pass — armed install → self-registration as an unclaimed appliance → operator bind → day-0 → running guest 9201 + agent 0.92.1 as `demo-hp-bb76ea`. **The shim loader booted with Secure Boot ENABLED**, which retires the assumption that Felhom installs need SB off — that was an N100-firmware workaround. The exact-serial disk filter took the system SSD and left the box's 1TB NVMe untouched/unenrolled on hardware it had never seen. Two failures filed rather than smoothed over: **R-59** (no DHCP → the installer baked a static fallback instead of aborting) and **R-61** (baked root password unknowable → no console access).
|
||||
| Customer claim: one-time emailed code → customer sets own password (bcrypt, operator never sees it) | controller v0.122, hub v0.50 | **PROVEN-LIVE** (drill VM) | `DRILL-day0-vm-2026-07-12` §10/F-4 (gate ON via real edge; claimed, code consumed) | Never executed by a non-Viktor human → R-3. **Deliverability (R-4), gmail half DONE 2026-07-18:** the rehearsal's claim email was the first sent under the tightened DMARC `p=quarantine` and **landed in the gmail Inbox, not spam** (`tests/VALIDATION-n100-rehearsal-2026-07-18.md`). **freemail.hu remains Viktor's open half.** (Dropped mis-cited `CAMPAIGN-4` F-C — that is the escrow-claim 502, not password claim) |
|
||||
| Customer binds their own appliance (self-service): operator-sent 7-day tokenized capability link → public two-factor `/bind/<token>` (console pairing code + retrieval passphrase) → hub stages the bind, no operator | hub v0.66.0 + ISO scripts v1.20.0 | **PROVEN-LIVE** (real customer-zero bind on metal, 2026-07-18) | **`tests/VALIDATION-n100-rehearsal-2026-07-18.md`:** operator minted + emailed the link 16:28:55 (7-day TTL, expiry 2026-07-25 recorded); **the customer bound their own box at 16:29:55 with `attempts=0`, `locked=0`** — `appliance_bound` carries source **`customer_selfbind`**, and the credential was delivered **26 s later** with no operator action. Hub-side lifecycle in `hub-state.txt` (`selfbind_tokens` mint→email→consume). Prior unit evidence: hub v0.66.0 (`web/selfbind.go`, `store/selfbind.go`; Scenarios A–F + F1/F2; 4 red-proofs verified red — THE TRAP `/bind/` exemption, no-oracle, lockout, single-active); GC verdict §3 (no appliance GC → TTL stands alone) | R-27 **slice 1**. No appliance list ever rendered; wrong code == wrong passphrase (one generic failure); 5-attempt lockout → call support; expiry falls back to operator-bind. **Live first-run DONE 2026-07-18** (rehearsal; the console banner rendered on the real ISO). **R-27b** (controller second-box dismissable prompt) deferred; **multi-box-per-link** = repeated operator sends |
|
||||
| Escrow ceremony: customer-facing wizard, one-shot R claim, operator zero-knowledge | controller v0.127, agent v0.88/0.89 | **PROVEN-LIVE** (drill VM, endpoint-exact) | agent v0.88.0 REPORT (ceremony ~4s, one-shot claim 200→410, R absent from every payload); `SPIKE-controller-escrow-2026-07-13` | **Customer-facing browser wizard FIRST LIVE FIRING 2026-07-18** (`tests/VALIDATION-n100-rehearsal-2026-07-18.md`, S6): customer zero drove the wizard on the reborn box — ceremony started 16:56:29, recovery code claimed one-shot 16:56:39 (absent from logs by design), hub-verified and `EscrowState` auto-confirmed 16:56:41, **offsite runs enabled 12 s after the ceremony began**; the v0.138.0 „megerősítésre vár, legfeljebb 15 perc" awaiting card rendered and flipped on the ACK (operator screenshots: Viktor's set). Honest caveat: at a 12-second confirm the awaiting window is so short that catching *both* states on screen is luck, not procedure. Prior: endpoints driven on the drill VM. **agent v0.89.0:** `/escrow/preflight` `pbs_storage_id` row now live-reloads (reads current agent.json) — a pbsdr convergence that seeds the id flips it green with NO service restart. **hub v0.60.0 (data-first retention):** a re-escrow with a DIFFERENT sealed passphrase no longer destroys the old blob — the hub RETAINS it (`host_escrow_superseded`), so a previous passphrase stays recoverable with its recovery code (turns the reinstall-orphan incident from "history destroyed" into "history recoverable"). Guided-recovery flow = R-26. Red-proof `TestSaveHostEscrow_RetainsSuperseded`. **hub v0.60.1 — custody survives the host lifecycle:** host deletion (with the escrow ack) DEMOTES the current blob to retained custody (moved into `host_escrow_superseded`, never destroyed; existing superseded rows spared); the customer Danger-zone Delete is the one true purge point (cascades both escrow tables incl. already-deleted hosts). No operator path through host lifecycle can lose a blob. Red-proofs `TestDeleteHost_DemotesEscrowNeverDestroys` + `TestDeleteCustomer_PurgesEscrowCustody`. **agent v0.93.0 (2026-07-21) — recovery codes can no longer contain a hyphenated word.** The EFF large list holds exactly four entries containing the hyphen the words are joined with (`drop-down`, `felt-tip`, `t-shirt`, `yo-yo`); drawing one produced a code that reads as 11 words instead of 10 — ambiguous to transcribe in exactly the situation R exists for. They are now excluded **from GENERATION only**: the draw space goes 7776 → 7772 and a 10-word code 129.248 → 129.241 bits, still well clear of the 128-bit floor. **Every code already issued remains valid** — R is verified as a whole passphrase by the PBS scrypt KDF and is never re-split, so no customer needs to re-run a ceremony. This also retired the long-standing ~1/5 `TestGenerateRecoveryCode_EntropyAndFormat` flake, which was this defect and not a flaky test |
|
||||
@@ -114,7 +114,7 @@
|
||||
| Customer/host management: 8-tab detail, scoped auto-refresh, safe stale-host deletion, capability chips | hub v0.47–0.53 | **PROVEN-LIVE** | hub v0.53.0 dead-host roll-up live on the Peti cluster (proxmox1 down 23h); `CAMPAIGN-4-2026-07-13` (operator UI driven live); `DRILL-day0-take2` F-16 (offsite/freeze buttons live) | 8-tab render + capability chips are **render-test-validated** (hub UI is password-gated; CC cannot log in). (Cited "daily operator use" was a no-doc citation; `AUDIT-hub-gui-2026-06-30` predates these features at hub v0.25) |
|
||||
| Config/state change round-trips in **seconds** (hub↔box immediacy; 15-min cycle stays the backbone): box→hub out-of-cycle report (Dir 1) + hub→box `GET /api/v1/wait` long-poll wake (Dir 2) | controller v0.139/140, hub v0.58/0.63 | **PROVEN-LIVE** (2026-07-21) | Transport proven live through the real DNS-only ingress: `SPIKE-immediate-sync-transport-2026-07-16` + hub v0.58.0 / controller v0.140.0 REPORTs — 240 s no-annotation hold (25 s heartbeat defeats nginx's 60 s `proxy_read_timeout`, no ingress change), 0.047 s wake-on-change, hub `rollout restart` = 1 WARN + 0-storm reconnect; Dir-1 2 s box→hub round-trip live in controller v0.139.0 | The operator-UI **save→apply** round-trip is not fired end-to-end live (hub UI password-gated; CC can't log in) → R-23; the wake transport and the ACK→config_version→`ConfigRefresher` delivery chain are each proven, only the UI-triggered bump leg is unexercised. Agent-plane (host-domain desired-state) poke **first slice PROVEN-LIVE** (Direction-2a, agent v0.89.0 + hub v0.59.0, 2026-07-17): contentless ep0-relayed UDP poke → agent immediate desired-state cycle, per `SPIKE-immediate-sync-transport-2026-07-16` P4. Full path live-proven: a real operator manifest save fired `poke: sync-poke delivered to 10.77.0.2`; the box (0.89.0) received it and logged `poke received → triggering an immediate desired-state cycle` → `out-of-band report triggered` — **~31 ms ep0→box, sub-ms to the report cycle** (WG-confined, from 10.77.0.1 to the 10.77.0.2-bound socket); save→tick ≈ ~0.45 s (SSH-dominated), well under ≤2–3 s. R-13 first slice (listener+sender only; the rest of the mutual-repair arc stays open). **System-initiated immediacy wired (hub v0.63.0, this REPORT):** the mutation sites that only OPERATOR actions used to notify now fire the correct plane's notifier when the hub itself mints state — agent-plane pokes at `PBSDRAutoProvision` (the observed slice-C lag), `ReissuePBSDR` (also the pbsdrheal escalation), `handlePBSDRReissue`, and the two admin desired-state api writers; controller-plane bump at `reissueOnReenroll`. Unit-tested + red-proofed, not yet fired on a real system event (folds into the rehearsal bind sequence). Still PARTIAL: the R-23 operator-UI save→apply leg and the agent **fast-tick-until-first-convergence** SECONDARY (the WG-registration leg a poke can't reach pre-tunnel) remain unfired live. **Fast-tick SHIPPED (agent v0.90.0, R-28):** while any desired-state item is unapplied — incl. the pre-tunnel window a poke can't reach — the agent pulses the out-of-band trigger every 30 s and self-disarms on convergence (state-based; four cached sources; LOUD states excluded). LIVE on both demo agents (the `fast-tick armed: 30s …` startup line verified). **REAL-ONBOARDING PROOF DONE — `tests/VALIDATION-n100-rehearsal-2026-07-18.md` (ledger 8, S5):** on a genuine first onboarding on metal, **every post-bind leg landed seconds apart with no ~15-minute stall anywhere** — bind 16:29:55 → credential delivered 16:30:21 (26 s) → agent 0.90.0 up 16:30:49 → **WG registered + tunnel applied 16:30:51 (~2 s)** → poke listener 16:30:54 → controller 16:32:28 → floor-lifted and running current 16:32:39. **Bind → running-current = 2 min 44 s.** The PBS-DR descriptor auto-provisioned on the same cadence (agent `converged state=applied` 16:45:53) — though see the DR-tier row: the descriptor converged while the credential behind it was already stale (R-39). The pre-tunnel fast-tick window is therefore proven in its real setting; the remaining PARTIAL is the R-23 operator-UI save→apply leg alone **2026-07-21 — THE LAST PARTIAL LEG IS CLOSED (R-23(a) restart leg).** The operator saved the global floor to a version the box did NOT run (0.153.0 → **v0.154.0**) and the managed self-update fired **exactly once**: `06:57:13Z` UpdateState pending (`initiated_by=auto-floor`) → `06:57:17Z` agent `controller swap requested` → `06:57:21Z` container restarted → `06:57:29Z` `new controller healthy`. **Save → healthy on the new version = 16 s.** Over a 39-minute window: swap requests **1**, agent-driven bootstrap restarts **1**, rollbacks **0**, container `RestartCount` **0**; `VerifyStartup` confirmed on the next boot and the following periodic check logged `Current version 0.154.0 is up to date` (the at/above-floor branch doing nothing, as designed). The 2026-07-20 attempt could not prove this because it targeted an already-running version. Evidence: `felhom-controller/REPORT.md` §6. |
|
||||
| Customer right-sizes guest RAM from the controller (agent-enforced bounds, live cgroup apply, no reboot) | agent v0.90.0 + controller v0.143.0 (R-24) | **PROVEN-LIVE** (grow **and** shrink on metal, 2026-07-18) | **`tests/VALIDATION-n100-rehearsal-2026-07-18.md` (ledger 9) — the apply is now proven in both directions on a normal-sized box:** customer zero **shrank 11675 → 8192 MB at 16:50:22** and **grew 8192 → 12288 MB at 17:02:17**, each a live cgroup apply with **no reboot** (`local-api: guest-memory resized` in the agent journal, `[web] memory resized` in the controller log), and the new total rippled into the deploy page's memory math at 17:05:15 (`total=12288MB`). F5 auto-sizing had landed the guest at 11675 MB. Controller-direct (R-24's hub-desired-state framing SUPERSEDED, Viktor 2026-07-17). Agent `GET`/`POST /guest/memory` enforces every bound FRESH (min 2048 / max host_total−2048 / shrink floor max(2048, usage+512)) + verify-after-apply; PVE `SetConfig` hot-applies (**Phase-0 PROVEN** on the nested box: maxmem moves with the guest running, /proc/meminfo ripples via lxcfs, no reboot). Controller "Szerver memória (RAM)" card + code→Hungarian map, gated on `FeatureGuestMemoryResize` (MinAgent 0.90.0). LIVE-validated end-to-end through the real endpoint on the demo (above_max + below_min refusals render the Hungarian, agent English never leaks; SupportYes via the version header) | Row complete as of the 2026-07-18 rehearsal — the refusals were proven on the nested demo, the **applies** on the N100. Cores stay observation |
|
||||
| Publish train: MinAgent floors, gated auto-Reissue, version channels, floor-field-LAST rules | hub v0.45/0.53, agent | **PARTIAL** | `runbooks/publish-train-rules.md`; demo-fleet updates proven | **Box-side floor lift PROVEN-LIVE on a fresh install** (`tests/VALIDATION-n100-rehearsal-2026-07-18.md`): the day-0 golden deployed controller **0.143.0** at 16:32:28 and the managed floor lifted it to **0.145.0 by 16:32:34 — a 5-second, fully unattended update inside the first minute of controller life**, `update-state.json` recording `initiated_by: auto-floor` with `controller_updated` pushed to the hub. So the *mechanism* is no longer nested-only. **Still never proven on a real REMOTE customer** — parked trains `RUNBOOK-publish-0.79/0.81/0.85-*` await Peti → R-1. **Action before first invite: rebuild the golden to 0.145.x** now that this evidence is banked, so fresh boxes don't sit two versions stale |
|
||||
| Publish train: MinAgent floors, gated auto-Reissue, version channels, floor-field-LAST rules | hub v0.45/0.53, agent | **PARTIAL** | `runbooks/publish-train-rules.md`; demo-fleet updates proven | **Box-side floor lift PROVEN-LIVE on a fresh install** (`tests/VALIDATION-n100-rehearsal-2026-07-18.md`): the day-0 golden deployed controller **0.143.0** at 16:32:28 and the managed floor lifted it to **0.145.0 by 16:32:34 — a 5-second, fully unattended update inside the first minute of controller life**, `update-state.json` recording `initiated_by: auto-floor` with `controller_updated` pushed to the hub. So the *mechanism* is no longer nested-only. **Still never proven on a real REMOTE customer** — parked trains `RUNBOOK-publish-0.79/0.81/0.85-*` await Peti → R-1. **Action before first invite: rebuild the golden to 0.145.x** now that this evidence is banked, so fresh boxes don't sit two versions stale | **Fresh-box floor lift proven 2026-07-21 (demo-hp):** a brand-new box came up on golden **0.153.0** and self-updated to the fleet floor **0.156.0** during day-0, unattended — the floor mechanism works on first contact, not only on boxes that have been in the fleet a while. That is the R-23(a) single-swap behaviour observed on a machine with no history at all.
|
||||
| Agent self-update: A/B slots, crash-loop auto-rollback, operator-signed | agent v0.70+ | **PROVEN-LIVE** (demo) | `SPIKE-agent-selfupdate-2026-07-05` | Remote-customer proof pending → R-1 |
|
||||
| Controller self-update: anonymous registry, no credentials in guest | controller v0.112 | **PROVEN-LIVE** (demo) | 07-10 arc | |
|
||||
| Offsite provisioning: Hetzner API, sub-account per customer, host-key pinning, credential re-issue | hub v0.37–0.39 | **PROVEN-LIVE** | `VALIDATION-offsite-provisioning-e2e-2026-07-09`, `SPIKE-hetzner-api-provisioning-2026-07-09` | |
|
||||
|
||||
File diff suppressed because one or more lines are too long
+19
-2
@@ -144,8 +144,10 @@ producer steps re-run each pass).
|
||||
default stays at the stock signed `shim`**, so Secure Boot keeps working. `mkimage` exists only to
|
||||
work around the N100's AMI firmware GRUB defect and is unsigned.
|
||||
|
||||
**PROVEN on this board (safety boot, operator-photographed 2026-07-21):** the shim loader booted, the
|
||||
answer file was fetched and parsed, and the match-nothing filter refused with **zero disk writes** —
|
||||
**PROVEN on this board (2026-07-21):** the safety boot loaded shim, fetched and parsed the answer file,
|
||||
and the match-nothing filter refused with **zero disk writes**; the armed ISO then installed end to end.
|
||||
**Secure Boot stayed ENABLED throughout** (`mokutil --sb-state` on the installed box → `SecureBoot
|
||||
enabled`) — so SB-off is an N100-firmware workaround, not a Felhom requirement.
|
||||
so the loader, Secure Boot setting, and network/answer path are all confirmed before anything
|
||||
destructive existed on a stick. `enp1s0f0` auto-detect verified on-board.
|
||||
|
||||
@@ -158,6 +160,21 @@ from the written stick — the profile is what you intended, the ISO is what you
|
||||
osirrox -indev <iso> -extract /answer.toml /tmp/a.toml && sed -n '/^\[disk-setup\]/,$p' /tmp/a.toml
|
||||
```
|
||||
|
||||
**t740 NIC TRAP — the 4-port card gets no DHCP; use the onboard port.** This board presents FIVE
|
||||
wired NICs and the install picks wrong:
|
||||
|
||||
| interface | driver | what it is | at this site |
|
||||
|---|---|---|---|
|
||||
| `enp1s0f0`–`f3` | `igb` | the 4-port expansion card | **no DHCP lease** (no link) |
|
||||
| `enp2s0f0` | `r8169` | **the onboard port** | leases fine, 1000 Mb |
|
||||
| `wlo1` | `iwlwifi` | wifi | unused |
|
||||
|
||||
Plug the cable into the **onboard** port. If the install already happened on the wrong port, the
|
||||
symptom is nasty: the installer does not abort, it bakes its **192.168.100.2 fallback as a STATIC
|
||||
`vmbr0` address** and completes, so the box looks installed and can never reach the hub (R-59).
|
||||
Repair on the console: point `bridge-ports` at `enp2s0f0` in `/etc/network/interfaces`, set the
|
||||
correct address (or `dhcp`), `ifreload -a`. R-60 is the self-heal that would make this unnecessary.
|
||||
|
||||
**Confirm the serial is the system disk and not a data drive.** On this board the SanDisk X600 128GB
|
||||
(`sda`) is the system disk; the 1TB NVMe is the future data drive and must stay OUTSIDE the filter —
|
||||
it joins later through the normal Tárhely flow, not the installer.
|
||||
|
||||
Reference in New Issue
Block a user