Kernel spike done (R-836): ESP one-shot and BootNext pass on all three boxes; watchdog fails; design with two operator questions
gates / gates (push) Successful in 3m3s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-07 13:08:03 +02:00
parent 28da46a764
commit c79e475511
39 changed files with 201 additions and 1 deletions
@@ -0,0 +1,75 @@
# The kernel lane — what the spike measured, and a design (R-836, R-812 option B; `09` §3 decisions 164, 171)
Spike 2026-10-07, operator present. Boxes: the Tester 1 box (VM 341 on the HP box, OVMF, Secure Boot off), demo-felhom
(Intel N100, Secure Boot off), demo-hp (AMD, Secure Boot ON). Kernels: 7.0.2-6 and 7.0.14-20. 24 reboots, 2 operator
power cycles. Evidence per boot in this folder (`<box>/<step>/before.txt`, `after.txt`, `result.txt`; screenshots for
the VM). Tools: `tools/`.
## 1. The problem, measured
Installing a kernel makes it the GRUB default at once (R-836, seen again today on Tester 1). A new kernel that fails
before userspace then stays the default on every power cycle. `--next-boot` and `grub-reboot` are not one-shots on these
hosts: `/boot` is ext4 on LVM, and GRUB cannot clear `next_entry` there (R-836, 2026-10-04).
## 2. What was measured today
| Candidate | Tester 1 (VM) | demo-felhom | demo-hp (Secure Boot on) |
|---|---|---|---|
| **2 — a one-shot flag in a GRUB env block on the ESP (vfat)** | one-shot → new kernel; next boot → old; a panic → back to old by itself | same | same |
| **1 — UEFI `BootNext` to a copy of the signed loader whose `grub.cfg` names the target** | same | same | same |
| **3 — a hardware watchdog armed by systemd during the reboot** | FAIL: frozen until `qm reset` | FAIL: frozen until the operator's power cycle | FAIL: frozen until the operator's power cycle |
- A **panic** (no init, `panic=10`) is recovered by 1 and 2 with no person: the crash boot shows in the screenshots
(Tester 1) and as a longer gap between boots (51–117 s against 15–34 s).
- A **freeze** (`panic=0`) is recovered by nothing: the machine reset during the reboot clears the watchdog timer, on
the emulated i6300ESB, Intel's TCO and AMD's SP5100 alike. Proxmox also blacklists hardware watchdog drivers by default.
- My first hang simulation (`init=` only) did NOT hang: initramfs-tools falls back to `/sbin/init`. `rdinit=` and `init=`
both missing forces the panic. Worth knowing for every future test of this kind.
## 3. Options for the lane
**A. Candidate 2 — the ESP flag.** A GRUB snippet reads `felhom_next` from an env file on the ESP, boots it once and
clears it. Writes nothing to firmware. Works with Secure Boot (it is plain `grub.cfg`). Costs: one `/etc/grub.d` file the
config bundle owns; the flag is written by the root wrapper. Can go wrong: an ESP that is not vfat or a GRUB without
`fat`/`save_env` — both checked once per box before the lane runs.
**B. Candidate 1 — `BootNext`.** The firmware clears it. Costs: a second loader copy on the ESP that must follow every
shim/GRUB update, and an NVRAM write per kernel step — some boards wear or reorder NVRAM (demo-hp already holds 41 boot
entries). Can go wrong: a firmware that ignores `BootNext`.
**C. Turn freezes into panics.** Add `softlockup_panic=1 hardlockup_panic=1 hung_task_panic=1 panic=10` to the one-shot
entry, so a lockup the kernel can detect becomes a panic that A or B recovers. NOT measured (no clean way to simulate a
lockup was tried). A true dead freeze stays a person's power cycle.
## 4. The pick — a proposal
**A, with C on the one-shot entry.** The new kernel boots once from the ESP flag. If the box comes back healthy (the host
health rule of `11` §8.2, the guest running), a userspace "boot good" step makes it the default. Otherwise the next
reboot returns to the old kernel by itself. A freeze still needs a person — the design says so to the household rather
than promising more.
## 5. The night slot
The kernel step ends the night: after the host step (`11` §8.2) and only on a night whose whole-guest backup succeeded.
The reboot then costs ~1–3 min of the household's apps (measured today: 43–181 s back, the guest starts by itself).
Ring 0 (the demo boxes) first; ring 1 by signed job after the operator's approval, as the Docker and Proxmox sets.
## 6. First slice and its proof
Build A in the wrapper (`kernel` layer, the ESP snippet in the bundle, the "boot good" step), red tests first (a plan
without the ESP flag refused; a kernel name outside the approved set refused; the snippet clears the flag). Live proof on
the Tester 1 box: one step to a new kernel; a forced panic falls back with no person (screenshots + boot gap).
## 7. Questions for the operator
1. **May the box restart at night for a kernel update, and do we tell the household?** That is a promise to users (the
design does not decide it). If you do nothing: kernels stay manual and no box restarts by itself.
2. **A frozen new kernel needs a person to switch the box off and on.** Accept that for the first customers (with the
household told what to do), or block the kernel lane until C is measured? If you do nothing: the lane is not built.
## 8. State left on the boxes
Every test entry, the ESP flag, the `BootNext` loader and its firmware entry, and the watchdog config are removed
(`<box>/G9-cleanup.txt`). Tester 1 and demo-felhom: default 7.0.2-6 (running), 7.0.14-20 installed and booted healthily
in the spike. demo-hp: default 7.0.14-20 (running). VM 341's test watchdog device is removed (applies at its next start).
Also seen: after a reboot, demo-felhom answered over Tailscale only after more than 6 minutes (the LAN at once).
@@ -0,0 +1,7 @@
done
felhom lines in grub.cfg: 0
GRUB_DEFAULT="gnulinux-advanced-1af1fcc6-639c-416b-a7e5-c4470d41a502>gnulinux-7.0.2-6-pve-advanced-1af1fcc6-639c-416b-a7e5-c4470d41a502"
7.0.2-6-pve
BootOrder: 0001,0003,0002,0000
BOOT
proxmox
@@ -0,0 +1,3 @@
lrwxrwxrwx 1 root root 9 Oct 7 12:55 /dev/felhom-hwwd -> watchdog1
/sys/class/watchdog/watchdog0 Software Watchdog timeout=10 state=active
/sys/class/watchdog/watchdog1 SP5100 TCO timer timeout=30 state=active
@@ -0,0 +1,7 @@
done
felhom lines in grub.cfg: 0
GRUB_DEFAULT="gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.14-20-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43"
7.0.14-20-pve
BootOrder: 0003,0029,0001,0002,0006,0007,0019,001A,001B,001C,001D,001E,0008,0009
BOOT
proxmox
@@ -0,0 +1,7 @@
7.0.14-20-pve
BOOT_IMAGE=/boot/vmlinuz-7.0.14-20-pve root=/dev/mapper/pve-root ro quiet
env:
BootCurrent: 0003
BootOrder: 0003,0029,0001,0002,0006,0007,0019,001A,001B,001C,001D,001E,0008,0009,0016,0017,001F,0020,0021,0022,0023,0024,0025,0026,0027,0028,000A,0004,0005,000B
-1 37b2e108fa0b44f985cfa026b908807e Mon 2026-10-05 09:56:46 CEST Wed 2026-10-07 12:47:43 CEST
0 cd3ae3a17ea142119259132a38995458 Wed 2026-10-07 12:48:16 CEST Wed 2026-10-07 12:48:29 CEST
@@ -0,0 +1 @@
r1-baseline: PASS kernel=7.0.14-20-pve back=181s gap-between-boots=33s
@@ -0,0 +1,7 @@
7.0.2-6-pve
BOOT_IMAGE=/boot/vmlinuz-7.0.2-6-pve root=/dev/mapper/pve-root ro quiet panic=10
env: felhom_next=
BootCurrent: 0003
BootOrder: 0003,0029,0001,0002,0006,0007,0019,001A,001B,001C,001D,001E,0008,0009,0016,0017,001F,0020,0021,0022,0023,0024,0025,0026,0027,0028,000A,0004,0005,000B
-1 cd3ae3a17ea142119259132a38995458 Wed 2026-10-07 12:48:16 CEST Wed 2026-10-07 12:48:41 CEST
0 05f158b202134ce6a285a295d5a40f31 Wed 2026-10-07 12:49:15 CEST Wed 2026-10-07 12:49:26 CEST
@@ -0,0 +1,4 @@
7.0.14-20-pve
felhom_next=felhom-spike-good
BootOrder: 0003,0029,0001,0002,0006,0007,0019,001A,001B,001C,001D,001E,0008,0009,0016,0017,001F,0020,0021,0022,0023,0024,0025,0026,0027,0028,000A,0004,0005,000B
2026-10-07T10:48:31Z
@@ -0,0 +1 @@
cd3ae3a1-7ea1-4211-9259-132a38995458
@@ -0,0 +1 @@
r2-c2-good: PASS kernel=7.0.2-6-pve back=53s gap-between-boots=34s
@@ -0,0 +1,7 @@
7.0.14-20-pve
BOOT_IMAGE=/boot/vmlinuz-7.0.14-20-pve root=/dev/mapper/pve-root ro quiet
env: felhom_next=
BootCurrent: 0003
BootOrder: 0003,0029,0001,0002,0006,0007,0019,001A,001B,001C,001D,001E,0008,0009,0016,0017,001F,0020,0021,0022,0023,0024,0025,0026,0027,0028,000A,0004,0005,000B
-1 05f158b202134ce6a285a295d5a40f31 Wed 2026-10-07 12:49:15 CEST Wed 2026-10-07 12:49:39 CEST
0 321e492ea4e64ea9943db9a544e17700 Wed 2026-10-07 12:50:13 CEST Wed 2026-10-07 12:50:24 CEST
@@ -0,0 +1,4 @@
7.0.2-6-pve
felhom_next=
BootOrder: 0003,0029,0001,0002,0006,0007,0019,001A,001B,001C,001D,001E,0008,0009,0016,0017,001F,0020,0021,0022,0023,0024,0025,0026,0027,0028,000A,0004,0005,000B
2026-10-07T10:49:27Z
@@ -0,0 +1 @@
05f158b2-0213-4ce6-a285-a295d5a40f31
@@ -0,0 +1 @@
r3-c2-plain: PASS kernel=7.0.14-20-pve back=55s gap-between-boots=34s
@@ -0,0 +1,7 @@
7.0.14-20-pve
BOOT_IMAGE=/boot/vmlinuz-7.0.14-20-pve root=/dev/mapper/pve-root ro quiet
env: felhom_next=
BootCurrent: 0003
BootOrder: 0003,0029,0001,0002,0006,0007,0019,001A,001B,001C,001D,001E,0008,0009,0016,0017,001F,0020,0021,0022,0023,0024,0025,0026,0027,0028,000A,0004,0005,000B
-1 321e492ea4e64ea9943db9a544e17700 Wed 2026-10-07 12:50:13 CEST Wed 2026-10-07 12:50:33 CEST
0 cf9353298ce94a438c68cce1cdf5de26 Wed 2026-10-07 12:51:47 CEST Wed 2026-10-07 12:51:59 CEST
@@ -0,0 +1,4 @@
7.0.14-20-pve
felhom_next=felhom-spike-panic
BootOrder: 0003,0029,0001,0002,0006,0007,0019,001A,001B,001C,001D,001E,0008,0009,0016,0017,001F,0020,0021,0022,0023,0024,0025,0026,0027,0028,000A,0004,0005,000B
2026-10-07T10:50:25Z
@@ -0,0 +1 @@
321e492e-a4e6-4ea9-943d-b9a544e17700
@@ -0,0 +1 @@
r4-c2-panic: PASS kernel=7.0.14-20-pve back=91s gap-between-boots=74s
@@ -0,0 +1,7 @@
7.0.2-6-pve
BOOT_IMAGE=/boot/vmlinuz-7.0.2-6-pve root=/dev/mapper/pve-root ro quiet panic=10
env: felhom_next=
BootCurrent: 0000
BootOrder: 0003,0029,0001,0002,0006,0007,0019,001A,001B,001C,001D,001E,0008,0009,0016,0017,001F,0020,0021,0022,0023,0024,0025,0026,0027,0028,0000,000C,000A
-1 cf9353298ce94a438c68cce1cdf5de26 Wed 2026-10-07 12:51:47 CEST Wed 2026-10-07 12:52:11 CEST
0 18039e256a30463485e70b98c856d620 Wed 2026-10-07 12:53:06 CEST Wed 2026-10-07 12:53:20 CEST
@@ -0,0 +1,5 @@
7.0.14-20-pve
felhom_next=
BootNext: 0000
BootOrder: 0003,0029,0001,0002,0006,0007,0019,001A,001B,001C,001D,001E,0008,0009,0016,0017,001F,0020,0021,0022,0023,0024,0025,0026,0027,0028,000A,0004,0005,000B
2026-10-07T10:52:01Z
@@ -0,0 +1 @@
cf935329-8ce9-4a43-8c68-cce1cdf5de26
@@ -0,0 +1 @@
r5-c1-good: PASS kernel=7.0.2-6-pve back=75s gap-between-boots=55s
@@ -0,0 +1,7 @@
7.0.14-20-pve
BOOT_IMAGE=/boot/vmlinuz-7.0.14-20-pve root=/dev/mapper/pve-root ro quiet
env: felhom_next=
BootCurrent: 0003
BootOrder: 0003,0029,0001,0002,0006,0007,0019,001A,001B,001C,001D,001E,0008,0009,0016,0017,001F,0020,0021,0022,0023,0024,0025,0026,0027,0028,0004,000A
-1 18039e256a30463485e70b98c856d620 Wed 2026-10-07 12:53:06 CEST Wed 2026-10-07 12:53:29 CEST
0 982d4e0ad5ef47a8a7d0b5f110e9ad13 Wed 2026-10-07 12:55:26 CEST Wed 2026-10-07 12:55:37 CEST
@@ -0,0 +1,5 @@
7.0.2-6-pve
felhom_next=
BootNext: 0000
BootOrder: 0003,0029,0001,0002,0006,0007,0019,001A,001B,001C,001D,001E,0008,0009,0016,0017,001F,0020,0021,0022,0023,0024,0025,0026,0027,0028,0000,000C,000A
2026-10-07T10:53:21Z
@@ -0,0 +1 @@
18039e25-6a30-4634-85e7-0b98c856d620
@@ -0,0 +1 @@
r6-c1-panic: PASS kernel=7.0.14-20-pve back=134s gap-between-boots=117s
@@ -0,0 +1,10 @@
7.0.14-20-pve
env: felhom_next=
-1 982d4e0ad5ef47a8a7d0b5f110e9ad13 Wed 2026-10-07 12:55:26 CEST Wed 2026-10-07 12:57:53 CEST
0 8ea2cdffd1664af69937156f4ee659a5 Wed 2026-10-07 13:05:28 CEST Wed 2026-10-07 13:06:04 CEST
/sys/class/watchdog/watchdog0 Software Watchdog state=active
341 night1004-tester1 running 8192 200.00 1816
VMID Status Lock Name
9201 running demo-hp
9202 running demo-hp-scratch
9401 stopped upgrade-harness
@@ -0,0 +1,4 @@
7.0.14-20-pve
felhom_next=felhom-spike-freeze
BootOrder: 0003,0029,0001,0002,0006,0007,0019,001A,001B,001C,001D,001E,0008,0009,0016,0017,001F,0020,0021,0022,0023,0024,0025,0026,0027,0028,0004,000A
2026-10-07T10:56:00Z
@@ -0,0 +1 @@
982d4e0a-d5ef-47a8-a7d0-b5f110e9ad13
@@ -0,0 +1 @@
r7-c3-freeze: FAIL — sp5100_tco armed by systemd (RebootWatchdogSec=90s) did NOT reset the frozen kernel; dead (no ping) 6 min after the reboot until the operator's power cycle; then the default 7.0.14-20
@@ -0,0 +1,7 @@
done
felhom lines in grub.cfg: 0
GRUB_DEFAULT="gnulinux-advanced-e7c978ee-70aa-454b-a5bb-53dc1d19f16e>gnulinux-7.0.2-6-pve-advanced-e7c978ee-70aa-454b-a5bb-53dc1d19f16e"
7.0.2-6-pve
BootOrder: 0004,0003,0000,0001
BOOT
proxmox
@@ -0,0 +1,10 @@
#!/bin/bash
# End of the spike: remove the test entries, the ESP one-shot and the BootNext loader + firmware entry. KEEP the explicit
# GRUB_DEFAULT the setup wrote (the kernel this box runs and booted healthily); the original file stays as *.felhom-spike-orig.
set -uo pipefail
rm -f /etc/grub.d/01_felhom_oneshot /etc/grub.d/41_felhom_spike /boot/efi/EFI/proxmox/felhom-oneshot.env
bash /root/spike-bootnext-undo.sh >/dev/null 2>&1
update-grub 2>&1 | tail -1
echo "felhom lines in grub.cfg: $(grep -c felhom /boot/grub/grub.cfg)"
grep -m1 '^GRUB_DEFAULT' /etc/default/grub; uname -r
efibootmgr | grep -E 'BootOrder|felhom|BootNext' | cut -c1-80; ls /boot/efi/EFI
+1 -1
View File
@@ -217,7 +217,7 @@ stopping line that lies.
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|---|---|---|---|---|---|---|---|
| **R-812** | Box system & updates | P2 | **[P2] A box never receives operating-system security updates — not the Proxmox host, not the guest's Debian, not its Docker engine.** SEARCHED 2026-10-03 (read-only): `felhom-controller`, `felhom-agent`, `app-catalog-felhom.eu` and `felhom.eu` hold no `apt-get upgrade`, `apt full-upgrade`, `unattended-upgrades`, `pveupgrade` or `needrestart` that runs on a box. The installer aligns the host's Proxmox repositories to no-subscription *"so the box can pull security updates"* and then says plainly *"No upgrades are run"* (`scripts/felhom-host-install.sh:2133-2136`). The guest's Docker engine is installed when the golden is BAKED (`felhom-agent/configs/build-golden.sh:124-125`), so a fresh install gets that week's engine and an installed box keeps it forever. The only `apt full-upgrade` in the project is a by-hand step for the off-site endpoint ep0 (`documentation/runbooks/offsite-endpoint.md:41`), not a box. App images ARE updated (the update arc); the layer under them is not. The intention, with its scope, is **R-808** in `ROADMAP.md`. | **NARROWED 2026-10-04 (evening) — the guest's DOCKER engine slow lane is BUILT and proven live (agent v0.142.0, hub v0.132.0; `11` §5.8): live-restore on everywhere, ring 0 steps under a root-owned mark, ring 1 and undo only by a signed job the wrapper re-verifies; the operator approves each engine set on the System page. LEFT: the kernel lane (R-836); existing boxes (R-840).** Earlier: NARROWED 2026-10-04 (afternoon) — the HOST's Debian fast lane is BUILT and proven live too (agent v0.141.1, hub v0.131.1; `11` §8.2, `audits/os-host-lane-2026-10-04/`): appliances only, never kernel/boot/firmware, after a healthy guest step; fleet view and four alarms (§8.3). LEFT: the Docker and kernel slow lanes (R-836); existing boxes (R-840).** Earlier: NARROWED AGAIN 2026-10-04 (day) — the GUEST's Debian fast lane is BUILT and proven live (agent v0.140.0, hub v0.130.0, installer 1.29.0; `11` §8.1, `audits/os-guest-lane-2026-10-04/`). LEFT: the host, Docker and kernel lanes; the undo (R-842); existing boxes (R-840).** Earlier: NARROWED 2026-10-04 — the SPIKE is done (`11` §7.1, corrections C1–C12, `audits/os-updates-spike-2026-10-04/`); no product code yet. LEFT: the build steps of `11` §8, each with the operator's go; the §5.3 snapshot question is in STATUS; preconditions R-835, R-836, R-837.** **2026-10-06 night: design written** (`audits/night-burndown-2026-10-06/design-R-812.md`) — what is left is `11` §8 step 6: the host's Proxmox-origin packages and its kernel are never updated (read tonight: demo-hp 77 and demo-felhom 78 Proxmox packages pending on the 2026-10-05 lists, `pve-manager` 9.2.2 → 9.2.21; demo-felhom still runs its install-time kernel 7.0.2-6). Pick: a `pve` slow lane for the Proxmox packages without the kernel, approved per set like the Docker engine; the kernel stays an operator-run step until R-836's fallback is measured. Proposed split: close R-812 when that lane ships; R-836 carries the kernel. Two operator questions in the design. **2026-10-07 07:58: `09` §3 decisions 163–164 — option A (the Proxmox package lane, per-set approval) YES; option B (the kernel lane) YES before the first paying customer, starting with a reboot spike carried by R-836.** | — | — | CC + operator |
| **R-836** | Box system & updates | P2 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin <new> --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot <new>`: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** **2026-10-07 07:58: raised P3 → P2 (`09` §3 decision 164 — the kernel lane before the first paying customer); the spike runs with the operator present.** | — | — | CC |
| **R-836** | Box system & updates | P2 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin <new> --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot <new>`: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** **2026-10-07 07:58: raised P3 → P2 (`09` §3 decision 164 — the kernel lane before the first paying customer); the spike runs with the operator present.** **2026-10-07 SPIKE DONE (operator present; 24 reboots, 2 power cycles; `audits/kernel-spike-2026-10-07/DESIGN-kernel-lane.md`).** On Tester 1 (VM), demo-felhom and demo-hp (Secure Boot on): **a one-shot flag in a GRUB env block on the ESP (vfat) WORKS** — new kernel once, then the old one; a panicking new kernel (`panic=10`) falls back with no person. **UEFI `BootNext` WORKS too.** **A hardware watchdog armed by systemd during the reboot FAILS on all three** (i6300ESB, Intel TCO, AMD SP5100): the reset clears the timer, so a FROZEN new kernel needs a power cycle. Pick: the ESP flag + lockup-to-panic kernel options (unmeasured). Two operator questions in the design (night restarts and the household; a freeze needs a person). Boxes left clean on a healthily booted default. | — | — | CC |
| **R-862** | Box system & updates | P3 | **Tester 2 cannot take the config bundle until one by-hand bootstrap is done: its `felhom-os-apply` (agent 0.142.0) predates the bundle mode, and no signed job can write a root file on it.** FOUND 2026-10-04 (R-840 build, Part C): the route reaches every box installed from installer 1.31.0 on, and the demo boxes (bootstrapped by CC); Tester 2 has every root file of agent 0.142.0 (installer 1.30.0) and lacks only the R-858 wrapper fix, which matters only for a Docker step it gets solely from a signed job. CC has no route to Tester 2 (its door admits only the operator's WireGuard peer; `felhom-op` cannot become root). The steps: `runbooks/config-bundle.md` "Tester 2". Then CC signs the bundle and reads it back. | **WAITING-ON-OPERATOR** — the operator said (2026-10-04 ~19:05) he will try through his tunnel **2026-10-07 07:58 (`09` §3 decision 170): the operator believes Tester 2 never set up the off-site backup (no escrow); the laptop may stay off for days; the tester plans to add an HDD.** **2026-10-07 (Part H, read only):** in Tester 2's last 10 notifications (2026-10-05 09:16 → 2026-10-07 06:58) the household got 0 mails, the operator 10 (host_down ×9, expected_dbdump_missed ×1); control: the household channel shows on the other three customers. So the household is not mailed daily while the box is off; no row (`audits/day-2026-10-07/H/`). | the operator's WireGuard tunnel | the operator runs the three bootstrap commands; CC sends the bundle | operator |
| **R-35** | Box system & updates | P3 | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-05 (burn-down night): NEEDS A DESIGN.** Hot-apply needs a per-setting ruling; persisting sessions puts login tokens on disk (and into the whole-guest archive, `07` §5). Next: the operator's pick between them. | — | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` | CC |
| **R-78** | Box system & updates | P3 | **`local_api` authority ruling — auto-reconcile vs detect-only** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-25, size M, roadmap state `idea (deferred OUT of R-77 on purpose)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **An OWED OPERATOR DECISION, not a defect — moved because an owed ruling hidden among feature ideas is the shape this session exists to remove.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | R-77 ships detection because the fix is genuinely undecided, and **both directions can lose customer-visible function**. **Direction 1 (today):** `controller.yaml` wins and drift is silent → the 2026-07-25 island migration blinded the whole fleet's control plane for 17.5 h (drive gate, guest-reboot recovery, quiesce/backup all degrade). R-77 makes that loud but does not stop it recurring. **Direction 2 (`bootstrap.json` wins, auto-reconcile on boot):** a guest whose `controller.yaml` is CORRECT and whose `bootstrap.json` is stale — a half-completed re-provision, a hand-repaired guest, a setup-wizard box — gets a **working channel clobbered on the next restart**, fleet-wide and silently, during a routine deploy. That is not obviously better than the bug. Needs a spike: which writer is authoritative per field (endpoint vs fingerprint vs token — `mergeLocalAPI` replaces the whole block, so they cannot be reconciled independently today), whether the agent side should stamp a generation/mtime so 'newer wins' is even expressible, and whether reconcile should require an operator ack. Until then the drift alert plus a manual edit is the supported path. | CC |