hub v0.143.0 (code): the kernel lane — the day-before household mail, the night instruction, the operator's kernel set (R-836, decision 172)
gates / gates (push) Successful in 2m47s

KernelDue / KernelNotify (09-20 h Budapest, one per 20 h, max 3, registered
address, only an accepted mail counts) / os_update.kernel {kver, tonight}
(no mail, no step) / layer kernel ingest + operator events / Approve kernel
set after every ring-0 box booted it healthily after a night stage / two
System page cells. 11 §5.11 written; §5.10 status corrected (proven).
Installer uninstall knows the two GRUB generators (unreleased).
Evidence: audits/kernel-lane-2026-10-07/ (red-proofs, boot timing).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-07 15:34:14 +02:00
parent 0c5dbfbfc1
commit ab3b7ea2f4
29 changed files with 1453 additions and 11 deletions
+76 -3
View File
@@ -469,7 +469,7 @@ Decision 88 (R-851): *"Yes, but maybe not indefinitely."* Evidence `audits/os-do
---
### 5.10 The Proxmox package lane — BUILT 2026-10-07, unreleased, not yet proven live (R-812 option A, `09` §3 decision 163) `[FACT]`
### 5.10 The Proxmox package lane — BUILT 2026-10-07 (agent v0.151.0, hub v0.142.0), proven on demo-felhom (R-812 option A, `09` §3 decision 163) `[FACT]`
**What it updates:** the host's Proxmox USERSPACE packages (pve-manager, qemu-server, pve-container, lxc-pve,
libpve-*, proxmox-widget-toolkit, the ceph client libraries …) — the 77–78 packages behind on the demo boxes on
@@ -498,9 +498,81 @@ libpve-*, proxmox-widget-toolkit, the ceph client libraries …) — the 77–78
when every ring-0 box ran the set in `DockerNightsEffective` (2) healthy night steps since it was first seen and none
was unhealthy — „healthy" is the box's own verdict above (no memory-kill check; that is the Docker engine's, R-528).
An approval nudges no box: ring 1 takes the set only by a signed `os_pve_step` (signed per box, as `os_docker_step`).
- **Not yet:** the live proof on demo-felhom (ring 0); the undo runbook (by hand: `apt-get install <name>=<old>` —
- **Proven live 2026-10-07** on demo-felhom (ring 0, a signed `os_pve_step`): 65 packages in 70 s, pve-manager
9.2.2 → 9.2.21, healthy, every container kept its id, the hub logged the report (`audits/day-2026-10-07/B/RESULT.md`).
- **Not yet:** the undo runbook (by hand: `apt-get install <name>=<old>` —
Proxmox keeps 30–66 old versions, C2).
### 5.11 The kernel lane — BUILT 2026-10-07 (agent v0.152.0, hub v0.143.0; R-836, `09` §3 decisions 164, 171, 172) `[FACT]`
**The ruling (decision 172):** design option A (a one-shot flag on the ESP) with option C (lockup → panic) on the one-shot
entry. A box may restart at night for a kernel update; **the household is mailed the day before** with what to do if
the box is not back by morning; each kernel set is approved by the operator after ring 0 ran it; **a frozen new kernel
that needs a person's power cycle is accepted** for the first customers. Why A: measured in the spike on the Tester 1
VM, demo-felhom and demo-hp (Secure Boot on) — the ESP flag and UEFI `BootNext` both boot a kernel once and fall back
after a panic; a hardware watchdog armed during the reboot never fired on any of the three
(`audits/kernel-spike-2026-10-07/DESIGN-kernel-lane.md`). A writes nothing to firmware.
**The boot side** (config bundle, two GRUB generators): `/etc/grub.d/01_felhom_oneshot` reads `felhom_next` from a GRUB
env block on the ESP (`EFI/felhom/oneshot.env`), clears and saves it BEFORE the menu, and — only if it names an
installed kernel — boots that kernel's one-shot entry. `/etc/grub.d/42_felhom_oneshot` gives each installed kernel an
entry `felhom-oneshot-<ver>`: the normal entry plus `softlockup_panic=1 hardlockup_panic=1 hung_task_panic=1 panic=10`
(option C; sorted after `10_linux`, so never entry 0 and never the default). A panic on the one-shot restarts the box in
10 s into the default — the old kernel. Without a vfat ESP both print nothing.
**The default** is the kernel the box RUNS, pinned in `/etc/default/grub.d/zz-felhom-kernel-default.cfg` before the
install and proved from grub.cfg after it (an install would otherwise make the newest kernel the default — R-836's
defect). Only a healthy one-shot boot moves it.
**The root wrapper** (`felhom-os-apply`, layer `kernel`, lane slow, an appliance, authority = a signed `os_kernel_step`
or the root-owned ring-0 mark): `apply` STAGES (installs the set — the series meta-package, at most one new kernel
image, the boot helper and firmware — pins the default, writes the flag; never reboots); `kernel-reboot` reboots a
staged step; `kernel-boot` says what became of it after a boot; `kernel-good` moves the default; `kernel-revert`
reboots ONCE into the old kernel; `kernel-cancel`; `kernel-status`. Refusals R20 (the box cannot do a one-shot: not
UEFI, no vfat ESP, a separate /boot, GRUB without fat/loadenv, the generators missing, a hand pin), R21 (the crash
guard tripped or saw an unclean boot within its window), R22 (the step's phase does not allow it; never two steps
within 20 h), R23 (not exactly one newer kernel, or not the one signed / told). State `/var/lib/felhom-kernel/state.json`.
**The night slot** (agent): the kernel step ENDS the night leg — after the host step (and the Proxmox step when one ran)
ended healthy, after the whole-guest backup (the leg runs only after it), under the same heavy-op gate, trigger `night`
only, at most one per night. Ring 0 stages the pending kernel and reboots; ring 1 reboots only a kernel a signed
`os_kernel_step` staged by day (the job never reboots). The hub hears `staged` before the reboot.
**"Boot good" and the self-revert** (agent, at every start): on the new kernel the agent judges the boot for **20
minutes** — the host health rule (§8.2: the Proxmox daemons and the agent active, the customer guest running and its
own rule passing, the tunnel running) AND the box reached the hub (the `judging` report itself). Healthy → the new
kernel becomes the default (`applied`). Not healthy by the deadline → `health_failed`, then ONE self-revert into the old
kernel (`self_reverted` after it boots; a revert that comes back on the new kernel is `revert_failed` and never
retried). The wait, measured 2026-10-07 by a plain reboot (`audits/kernel-lane-2026-10-07/B/`): every container healthy
68 s after the reboot on demo-felhom (the hub reached at 63 s; Tailscale back at 53 s — the spike's > 6 min did not
recur) and 272 s on demo-hp (the hub at 189 s). A box that never comes back raises the hub's existing `host_stale`
(45 min, `alerting.stale_threshold`) and `host_down` (90 min) to the operator.
**The crash guard** (§5.9): the step's reboot and the self-revert are orderly (the clean-stop marker), so a step adds at
most ONE unclean boot (a panic before userspace adds none: the planned reboot's marker is still there); with R21 the
step starts only with none in the window — the box cannot reach the 3rd unclean boot that leaves it off. Pinned by
`KernelStepCannotLeaveTheBoxOff`.
**The household mail** (hub): a box is DUE when ring 0 has a pending kernel (its host report's `proxmox-kernel-X.Y`
upgrade) or a ring-1 box runs a kernel a signed job staged, and no step for that kernel has ended (`applied`,
`fell_back`, `health_failed`, `self_reverted`, `revert_failed`, or a refusal R20/R23/R3/R12 — the operator decides; never
retried by itself). Its household gets ONE mail (`mail.kernel.*`, its language, informal; to the registered address)
between 09:00 and 20:00 Budapest time, at most once per 20 h and three times per kernel. The box's `os_update.kernel`
block says `tonight` only within 24 h of a mail the mail service ACCEPTED — **no mail, no step**. Operator events
`os_kernel_step` and `os_kernel_notice`; the household's timeline gets the usual "security fixes installed" line after
`applied`.
**Approval** (hub): "Approve kernel set" on the System page appears when every ring-0 box's newest ended kernel step is
`applied` for the same kernel, after a NIGHT stage. The approved set is the series meta-package and the signed image. An
approval nudges no box: ring 1 takes it only through a signed `os_kernel_step` (stage), then its own told night.
**The System page** shows per box: the running kernel, the next boot, the default (amber while a one-shot flag names
another kernel), and the kernel step (the newest result, the kernel due, whether the household was told for tonight).
**Not measured:** option C itself — the Proxmox kernel ships no lockup test module (`CONFIG_TEST_LOCKUP` is not set on
7.0.14-20), so a soft lockup turning into a panic is recorded as unmeasured (no tool was built). A true dead freeze
needs a person (accepted, decision 172). Evidence and proofs: `audits/kernel-lane-2026-10-07/`.
## 6. Risks and edge cases
| # | What can go wrong | What the design does |
@@ -580,7 +652,8 @@ Each step returns to the operator for go or no-go.
after enrolment (16:24/16:25 UTC: 49 guest + 106 host packages, 110 s + 44 s, healthy). An installer pass would cost
~45 s and run BEFORE any whole-guest backup exists (no undo) — not built.
6. **Slow lane: host kernel and Proxmox packages, with the reboot.** **Split 2026-10-07 (decisions 163–164):** the Proxmox
USERSPACE packages are §5.10 (BUILT, unreleased, no reboot); the kernel lane is R-836 (a spike with reboots first).
USERSPACE packages are §5.10 (BUILT, proven on demo-felhom); the kernel lane is §5.11 (**BUILT 2026-10-07**, agent
v0.152.0 + hub v0.143.0, decision 172; the spike with reboots came first, `audits/kernel-spike-2026-10-07/`).
7. **Later:** the Proxmox major upgrade (PVE 9 → 10), drilled on ring 0 first.
---