Files
felhom.eu/documentation/audits/kernel-spike-2026-10-07/DESIGN-kernel-lane.md
T
2026-10-07 18:19:03 +02:00

5.4 KiB
Raw Blame History

The kernel lane — what the spike measured, and a design (R-836, R-812 option B; 09 §3 decisions 164, 171)

Ruled 2026-10-07 14:43 (decision 172): A with C. BUILT the same day — 11-os-updates.md §5.11, proofs in audits/kernel-lane-2026-10-07/.

Spike 2026-10-07, operator present. Boxes: the Tester 1 box (VM 341 on the HP box, OVMF, Secure Boot off), demo-felhom (Intel N100, Secure Boot off), demo-hp (AMD, Secure Boot ON). Kernels: 7.0.2-6 and 7.0.14-20. 24 reboots, 2 operator power cycles. Evidence per boot in this folder (<box>/<step>/before.txt, after.txt, result.txt; screenshots for the VM). Tools: tools/.

1. The problem, measured

Installing a kernel makes it the GRUB default at once (R-836, seen again today on Tester 1). A new kernel that fails before userspace then stays the default on every power cycle. --next-boot and grub-reboot are not one-shots on these hosts: /boot is ext4 on LVM, and GRUB cannot clear next_entry there (R-836, 2026-10-04).

2. What was measured today

Candidate Tester 1 (VM) demo-felhom demo-hp (Secure Boot on)
2 — a one-shot flag in a GRUB env block on the ESP (vfat) one-shot → new kernel; next boot → old; a panic → back to old by itself same same
1 — UEFI BootNext to a copy of the signed loader whose grub.cfg names the target same same same
3 — a hardware watchdog armed by systemd during the reboot FAIL: frozen until qm reset FAIL: frozen until the operator's power cycle FAIL: frozen until the operator's power cycle
  • A panic (no init, panic=10) is recovered by 1 and 2 with no person: the crash boot shows in the screenshots (Tester 1) and as a longer gap between boots (51–117 s against 15–34 s).
  • A freeze (panic=0) is recovered by nothing: the machine reset during the reboot clears the watchdog timer, on the emulated i6300ESB, Intel's TCO and AMD's SP5100 alike. Proxmox also blacklists hardware watchdog drivers by default.
  • My first hang simulation (init= only) did NOT hang: initramfs-tools falls back to /sbin/init. rdinit= and init= both missing forces the panic. Worth knowing for every future test of this kind.

3. Options for the lane

A. Candidate 2 — the ESP flag. A GRUB snippet reads felhom_next from an env file on the ESP, boots it once and clears it. Writes nothing to firmware. Works with Secure Boot (it is plain grub.cfg). Costs: one /etc/grub.d file the config bundle owns; the flag is written by the root wrapper. Can go wrong: an ESP that is not vfat or a GRUB without fat/save_env — both checked once per box before the lane runs.

B. Candidate 1 — BootNext. The firmware clears it. Costs: a second loader copy on the ESP that must follow every shim/GRUB update, and an NVRAM write per kernel step — some boards wear or reorder NVRAM (demo-hp already holds 41 boot entries). Can go wrong: a firmware that ignores BootNext.

C. Turn freezes into panics. Add softlockup_panic=1 hardlockup_panic=1 hung_task_panic=1 panic=10 to the one-shot entry, so a lockup the kernel can detect becomes a panic that A or B recovers. NOT measured (no clean way to simulate a lockup was tried). A true dead freeze stays a person's power cycle.

4. The pick — a proposal

A, with C on the one-shot entry. The new kernel boots once from the ESP flag. If the box comes back healthy (the host health rule of 11 §8.2, the guest running), a userspace "boot good" step makes it the default. Otherwise the next reboot returns to the old kernel by itself. A freeze still needs a person — the design says so to the household rather than promising more.

5. The night slot

The kernel step ends the night: after the host step (11 §8.2) and only on a night whose whole-guest backup succeeded. The reboot then costs ~1–3 min of the household's apps (measured today: 43–181 s back, the guest starts by itself). Ring 0 (the demo boxes) first; ring 1 by signed job after the operator's approval, as the Docker and Proxmox sets.

6. First slice and its proof

Build A in the wrapper (kernel layer, the ESP snippet in the bundle, the "boot good" step), red tests first (a plan without the ESP flag refused; a kernel name outside the approved set refused; the snippet clears the flag). Live proof on the Tester 1 box: one step to a new kernel; a forced panic falls back with no person (screenshots + boot gap).

7. Questions for the operator

  1. May the box restart at night for a kernel update, and do we tell the household? That is a promise to users (the design does not decide it). If you do nothing: kernels stay manual and no box restarts by itself.
  2. A frozen new kernel needs a person to switch the box off and on. Accept that for the first customers (with the household told what to do), or block the kernel lane until C is measured? If you do nothing: the lane is not built.

8. State left on the boxes

Every test entry, the ESP flag, the BootNext loader and its firmware entry, and the watchdog config are removed (<box>/G9-cleanup.txt). Tester 1 and demo-felhom: default 7.0.2-6 (running), 7.0.14-20 installed and booted healthily in the spike. demo-hp: default 7.0.14-20 (running). VM 341's test watchdog device is removed (applies at its next start). Also seen: after a reboot, demo-felhom answered over Tailscale only after more than 6 minutes (the LAN at once).