Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
5.4 KiB
The kernel lane — what the spike measured, and a design (R-836, R-812 option B; 09 §3 decisions 164, 171)
Ruled 2026-10-07 14:43 (decision 172): A with C. BUILT the same day —
11-os-updates.md§5.11, proofs inaudits/kernel-lane-2026-10-07/.
Spike 2026-10-07, operator present. Boxes: the Tester 1 box (VM 341 on the HP box, OVMF, Secure Boot off), demo-felhom
(Intel N100, Secure Boot off), demo-hp (AMD, Secure Boot ON). Kernels: 7.0.2-6 and 7.0.14-20. 24 reboots, 2 operator
power cycles. Evidence per boot in this folder (<box>/<step>/before.txt, after.txt, result.txt; screenshots for
the VM). Tools: tools/.
1. The problem, measured
Installing a kernel makes it the GRUB default at once (R-836, seen again today on Tester 1). A new kernel that fails
before userspace then stays the default on every power cycle. --next-boot and grub-reboot are not one-shots on these
hosts: /boot is ext4 on LVM, and GRUB cannot clear next_entry there (R-836, 2026-10-04).
2. What was measured today
| Candidate | Tester 1 (VM) | demo-felhom | demo-hp (Secure Boot on) |
|---|---|---|---|
| 2 — a one-shot flag in a GRUB env block on the ESP (vfat) | one-shot → new kernel; next boot → old; a panic → back to old by itself | same | same |
1 — UEFI BootNext to a copy of the signed loader whose grub.cfg names the target |
same | same | same |
| 3 — a hardware watchdog armed by systemd during the reboot | FAIL: frozen until qm reset |
FAIL: frozen until the operator's power cycle | FAIL: frozen until the operator's power cycle |
- A panic (no init,
panic=10) is recovered by 1 and 2 with no person: the crash boot shows in the screenshots (Tester 1) and as a longer gap between boots (51–117 s against 15–34 s). - A freeze (
panic=0) is recovered by nothing: the machine reset during the reboot clears the watchdog timer, on the emulated i6300ESB, Intel's TCO and AMD's SP5100 alike. Proxmox also blacklists hardware watchdog drivers by default. - My first hang simulation (
init=only) did NOT hang: initramfs-tools falls back to/sbin/init.rdinit=andinit=both missing forces the panic. Worth knowing for every future test of this kind.
3. Options for the lane
A. Candidate 2 — the ESP flag. A GRUB snippet reads felhom_next from an env file on the ESP, boots it once and
clears it. Writes nothing to firmware. Works with Secure Boot (it is plain grub.cfg). Costs: one /etc/grub.d file the
config bundle owns; the flag is written by the root wrapper. Can go wrong: an ESP that is not vfat or a GRUB without
fat/save_env — both checked once per box before the lane runs.
B. Candidate 1 — BootNext. The firmware clears it. Costs: a second loader copy on the ESP that must follow every
shim/GRUB update, and an NVRAM write per kernel step — some boards wear or reorder NVRAM (demo-hp already holds 41 boot
entries). Can go wrong: a firmware that ignores BootNext.
C. Turn freezes into panics. Add softlockup_panic=1 hardlockup_panic=1 hung_task_panic=1 panic=10 to the one-shot
entry, so a lockup the kernel can detect becomes a panic that A or B recovers. NOT measured (no clean way to simulate a
lockup was tried). A true dead freeze stays a person's power cycle.
4. The pick — a proposal
A, with C on the one-shot entry. The new kernel boots once from the ESP flag. If the box comes back healthy (the host
health rule of 11 §8.2, the guest running), a userspace "boot good" step makes it the default. Otherwise the next
reboot returns to the old kernel by itself. A freeze still needs a person — the design says so to the household rather
than promising more.
5. The night slot
The kernel step ends the night: after the host step (11 §8.2) and only on a night whose whole-guest backup succeeded.
The reboot then costs ~1–3 min of the household's apps (measured today: 43–181 s back, the guest starts by itself).
Ring 0 (the demo boxes) first; ring 1 by signed job after the operator's approval, as the Docker and Proxmox sets.
6. First slice and its proof
Build A in the wrapper (kernel layer, the ESP snippet in the bundle, the "boot good" step), red tests first (a plan
without the ESP flag refused; a kernel name outside the approved set refused; the snippet clears the flag). Live proof on
the Tester 1 box: one step to a new kernel; a forced panic falls back with no person (screenshots + boot gap).
7. Questions for the operator
- May the box restart at night for a kernel update, and do we tell the household? That is a promise to users (the design does not decide it). If you do nothing: kernels stay manual and no box restarts by itself.
- A frozen new kernel needs a person to switch the box off and on. Accept that for the first customers (with the household told what to do), or block the kernel lane until C is measured? If you do nothing: the lane is not built.
8. State left on the boxes
Every test entry, the ESP flag, the BootNext loader and its firmware entry, and the watchdog config are removed
(<box>/G9-cleanup.txt). Tester 1 and demo-felhom: default 7.0.2-6 (running), 7.0.14-20 installed and booted healthily
in the spike. demo-hp: default 7.0.14-20 (running). VM 341's test watchdog device is removed (applies at its next start).
Also seen: after a reboot, demo-felhom answered over Tailscale only after more than 6 minutes (the LAN at once).