Kernel spike done (R-836): ESP one-shot and BootNext pass on all three boxes; watchdog fails; design with two operator questions
gates / gates (push) Successful in 3m3s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-07 13:08:03 +02:00
parent 28da46a764
commit c79e475511
39 changed files with 201 additions and 1 deletions
@@ -0,0 +1,75 @@
# The kernel lane — what the spike measured, and a design (R-836, R-812 option B; `09` §3 decisions 164, 171)
Spike 2026-10-07, operator present. Boxes: the Tester 1 box (VM 341 on the HP box, OVMF, Secure Boot off), demo-felhom
(Intel N100, Secure Boot off), demo-hp (AMD, Secure Boot ON). Kernels: 7.0.2-6 and 7.0.14-20. 24 reboots, 2 operator
power cycles. Evidence per boot in this folder (`<box>/<step>/before.txt`, `after.txt`, `result.txt`; screenshots for
the VM). Tools: `tools/`.
## 1. The problem, measured
Installing a kernel makes it the GRUB default at once (R-836, seen again today on Tester 1). A new kernel that fails
before userspace then stays the default on every power cycle. `--next-boot` and `grub-reboot` are not one-shots on these
hosts: `/boot` is ext4 on LVM, and GRUB cannot clear `next_entry` there (R-836, 2026-10-04).
## 2. What was measured today
| Candidate | Tester 1 (VM) | demo-felhom | demo-hp (Secure Boot on) |
|---|---|---|---|
| **2 — a one-shot flag in a GRUB env block on the ESP (vfat)** | one-shot → new kernel; next boot → old; a panic → back to old by itself | same | same |
| **1 — UEFI `BootNext` to a copy of the signed loader whose `grub.cfg` names the target** | same | same | same |
| **3 — a hardware watchdog armed by systemd during the reboot** | FAIL: frozen until `qm reset` | FAIL: frozen until the operator's power cycle | FAIL: frozen until the operator's power cycle |
- A **panic** (no init, `panic=10`) is recovered by 1 and 2 with no person: the crash boot shows in the screenshots
(Tester 1) and as a longer gap between boots (51–117 s against 15–34 s).
- A **freeze** (`panic=0`) is recovered by nothing: the machine reset during the reboot clears the watchdog timer, on
the emulated i6300ESB, Intel's TCO and AMD's SP5100 alike. Proxmox also blacklists hardware watchdog drivers by default.
- My first hang simulation (`init=` only) did NOT hang: initramfs-tools falls back to `/sbin/init`. `rdinit=` and `init=`
both missing forces the panic. Worth knowing for every future test of this kind.
## 3. Options for the lane
**A. Candidate 2 — the ESP flag.** A GRUB snippet reads `felhom_next` from an env file on the ESP, boots it once and
clears it. Writes nothing to firmware. Works with Secure Boot (it is plain `grub.cfg`). Costs: one `/etc/grub.d` file the
config bundle owns; the flag is written by the root wrapper. Can go wrong: an ESP that is not vfat or a GRUB without
`fat`/`save_env` — both checked once per box before the lane runs.
**B. Candidate 1 — `BootNext`.** The firmware clears it. Costs: a second loader copy on the ESP that must follow every
shim/GRUB update, and an NVRAM write per kernel step — some boards wear or reorder NVRAM (demo-hp already holds 41 boot
entries). Can go wrong: a firmware that ignores `BootNext`.
**C. Turn freezes into panics.** Add `softlockup_panic=1 hardlockup_panic=1 hung_task_panic=1 panic=10` to the one-shot
entry, so a lockup the kernel can detect becomes a panic that A or B recovers. NOT measured (no clean way to simulate a
lockup was tried). A true dead freeze stays a person's power cycle.
## 4. The pick — a proposal
**A, with C on the one-shot entry.** The new kernel boots once from the ESP flag. If the box comes back healthy (the host
health rule of `11` §8.2, the guest running), a userspace "boot good" step makes it the default. Otherwise the next
reboot returns to the old kernel by itself. A freeze still needs a person — the design says so to the household rather
than promising more.
## 5. The night slot
The kernel step ends the night: after the host step (`11` §8.2) and only on a night whose whole-guest backup succeeded.
The reboot then costs ~1–3 min of the household's apps (measured today: 43–181 s back, the guest starts by itself).
Ring 0 (the demo boxes) first; ring 1 by signed job after the operator's approval, as the Docker and Proxmox sets.
## 6. First slice and its proof
Build A in the wrapper (`kernel` layer, the ESP snippet in the bundle, the "boot good" step), red tests first (a plan
without the ESP flag refused; a kernel name outside the approved set refused; the snippet clears the flag). Live proof on
the Tester 1 box: one step to a new kernel; a forced panic falls back with no person (screenshots + boot gap).
## 7. Questions for the operator
1. **May the box restart at night for a kernel update, and do we tell the household?** That is a promise to users (the
design does not decide it). If you do nothing: kernels stay manual and no box restarts by itself.
2. **A frozen new kernel needs a person to switch the box off and on.** Accept that for the first customers (with the
household told what to do), or block the kernel lane until C is measured? If you do nothing: the lane is not built.
## 8. State left on the boxes
Every test entry, the ESP flag, the `BootNext` loader and its firmware entry, and the watchdog config are removed
(`<box>/G9-cleanup.txt`). Tester 1 and demo-felhom: default 7.0.2-6 (running), 7.0.14-20 installed and booted healthily
in the spike. demo-hp: default 7.0.14-20 (running). VM 341's test watchdog device is removed (applies at its next start).
Also seen: after a reboot, demo-felhom answered over Tailscale only after more than 6 minutes (the LAN at once).