# The kernel lane — what the spike measured, and a design (R-836, R-812 option B; `09` §3 decisions 164, 171) > **Ruled 2026-10-07 14:43 (decision 172): A with C. BUILT the same day — `11-os-updates.md` §5.11, proofs in > `audits/kernel-lane-2026-10-07/`.** Spike 2026-10-07, operator present. Boxes: the Tester 1 box (VM 341 on the HP box, OVMF, Secure Boot off), demo-felhom (Intel N100, Secure Boot off), demo-hp (AMD, Secure Boot ON). Kernels: 7.0.2-6 and 7.0.14-20. 24 reboots, 2 operator power cycles. Evidence per boot in this folder (`//before.txt`, `after.txt`, `result.txt`; screenshots for the VM). Tools: `tools/`. ## 1. The problem, measured Installing a kernel makes it the GRUB default at once (R-836, seen again today on Tester 1). A new kernel that fails before userspace then stays the default on every power cycle. `--next-boot` and `grub-reboot` are not one-shots on these hosts: `/boot` is ext4 on LVM, and GRUB cannot clear `next_entry` there (R-836, 2026-10-04). ## 2. What was measured today | Candidate | Tester 1 (VM) | demo-felhom | demo-hp (Secure Boot on) | |---|---|---|---| | **2 — a one-shot flag in a GRUB env block on the ESP (vfat)** | one-shot → new kernel; next boot → old; a panic → back to old by itself | same | same | | **1 — UEFI `BootNext` to a copy of the signed loader whose `grub.cfg` names the target** | same | same | same | | **3 — a hardware watchdog armed by systemd during the reboot** | FAIL: frozen until `qm reset` | FAIL: frozen until the operator's power cycle | FAIL: frozen until the operator's power cycle | - A **panic** (no init, `panic=10`) is recovered by 1 and 2 with no person: the crash boot shows in the screenshots (Tester 1) and as a longer gap between boots (51–117 s against 15–34 s). - A **freeze** (`panic=0`) is recovered by nothing: the machine reset during the reboot clears the watchdog timer, on the emulated i6300ESB, Intel's TCO and AMD's SP5100 alike. Proxmox also blacklists hardware watchdog drivers by default. - My first hang simulation (`init=` only) did NOT hang: initramfs-tools falls back to `/sbin/init`. `rdinit=` and `init=` both missing forces the panic. Worth knowing for every future test of this kind. ## 3. Options for the lane **A. Candidate 2 — the ESP flag.** A GRUB snippet reads `felhom_next` from an env file on the ESP, boots it once and clears it. Writes nothing to firmware. Works with Secure Boot (it is plain `grub.cfg`). Costs: one `/etc/grub.d` file the config bundle owns; the flag is written by the root wrapper. Can go wrong: an ESP that is not vfat or a GRUB without `fat`/`save_env` — both checked once per box before the lane runs. **B. Candidate 1 — `BootNext`.** The firmware clears it. Costs: a second loader copy on the ESP that must follow every shim/GRUB update, and an NVRAM write per kernel step — some boards wear or reorder NVRAM (demo-hp already holds 41 boot entries). Can go wrong: a firmware that ignores `BootNext`. **C. Turn freezes into panics.** Add `softlockup_panic=1 hardlockup_panic=1 hung_task_panic=1 panic=10` to the one-shot entry, so a lockup the kernel can detect becomes a panic that A or B recovers. NOT measured (no clean way to simulate a lockup was tried). A true dead freeze stays a person's power cycle. ## 4. The pick — a proposal **A, with C on the one-shot entry.** The new kernel boots once from the ESP flag. If the box comes back healthy (the host health rule of `11` §8.2, the guest running), a userspace "boot good" step makes it the default. Otherwise the next reboot returns to the old kernel by itself. A freeze still needs a person — the design says so to the household rather than promising more. ## 5. The night slot The kernel step ends the night: after the host step (`11` §8.2) and only on a night whose whole-guest backup succeeded. The reboot then costs ~1–3 min of the household's apps (measured today: 43–181 s back, the guest starts by itself). Ring 0 (the demo boxes) first; ring 1 by signed job after the operator's approval, as the Docker and Proxmox sets. ## 6. First slice and its proof Build A in the wrapper (`kernel` layer, the ESP snippet in the bundle, the "boot good" step), red tests first (a plan without the ESP flag refused; a kernel name outside the approved set refused; the snippet clears the flag). Live proof on the Tester 1 box: one step to a new kernel; a forced panic falls back with no person (screenshots + boot gap). ## 7. Questions for the operator 1. **May the box restart at night for a kernel update, and do we tell the household?** That is a promise to users (the design does not decide it). If you do nothing: kernels stay manual and no box restarts by itself. 2. **A frozen new kernel needs a person to switch the box off and on.** Accept that for the first customers (with the household told what to do), or block the kernel lane until C is measured? If you do nothing: the lane is not built. ## 8. State left on the boxes Every test entry, the ESP flag, the `BootNext` loader and its firmware entry, and the watchdog config are removed (`/G9-cleanup.txt`). Tester 1 and demo-felhom: default 7.0.2-6 (running), 7.0.14-20 installed and booted healthily in the spike. demo-hp: default 7.0.14-20 (running). VM 341's test watchdog device is removed (applies at its next start). Also seen: after a reboot, demo-felhom answered over Tailscale only after more than 6 minutes (the LAN at once).