Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
9.6 KiB
R-812 — what is left of OS updates: the Proxmox packages and the kernel — design proposal (burn-down night 2026-10-06, no code)
Baselines read: felhom-agent 74b5eae, felhom.eu 8e2dc204. Architecture: 11-os-updates.md (§1, §3, §5.2, §5.5, §5.6,
§5.9, §7.1, §8). Related rows: R-836 (the kernel's fallback), R-808 (the intention, ROADMAP.md).
1. The problem
11 §8 steps 1–5 are BUILT: the guest's Debian, the host's Debian, the fleet view, the Docker engine, the root files
(R-840, closed). Step 6 is not: the host's Proxmox packages and its kernel are never updated on any box. The host
fast lane takes only Debian-origin packages, so everything from the Proxmox repository waits for ever — including
packages with ordinary names (zfsutils-linux, shim-signed, corosync, ceph-common, amd64-microcode, 11 C3).
Measured tonight (read only; apt list --upgradable, pveversion, dpkg -l, ls /boot; no apt update — the package
lists are from 2026-10-05, the last daily refresh):
- demo-hp:
pve-manager9.2.2 → 9.2.21 pending; 77 packages pending, all from the Proxmox repository (the Debian lane has taken the rest):qemu-server9.1.15 → 9.2.10,pve-container6.1.10 → 6.1.14,proxmox-backup-client4.2.0 → 4.2.7,zfsutils-linux2.4.2 → 2.4.4,shim-signed1.48 → 1.51,pve-firmware,corosync,ceph-*,frr,amd64-microcode,libpve-*. Kernel: runs 7.0.14-20 (installed BY HAND in the 2026-10-04 spike), 7.0.2-6 kept. - demo-felhom:
pve-manager9.2.2; 78 pending; runs kernel 7.0.2-6, the install-time kernel — 7.0.14-20 is in the repository and not installed. - So a box keeps its install-time Proxmox and kernel. Proxmox security fixes (for example in
pve-manager's web API, inqemu-server, in the Secure Boot loadershim-signed) never arrive. The spike counted 79 Proxmox packages "not covered" on the host (11§7.1).
2. What the code does today (read in source)
- The fast-lane origin rule:
felhom-agent/configs/felhom-os-apply:63FAST_ORIGINS = ("Debian", "Debian-Security"); a plan package of any other origin is refused R2 (:491-492, and in the simulation:1164origin_ok). - The host's name rule on top:
:91-92HOST_SLOW_RE(linux-*,proxmox-kernel*,pve-kernel*,pve-firmware,firmware-*,grub*,shim*,systemd-boot*,*-microcode,efibootmgr) → refusal R14 (:493-494,:1234-1235); the pending selection skips them (:1136). - Layers and lanes:
:437-442—guest/hostare lanefastonly;dockeris the onlyslowlayer. There is no host slow layer. - Tests pinning this:
configs/test_felhom_os_apply.py:417test_R2_non_debian_origin_in_the_plan,:422test_R2_non_debian_origin_in_the_simulation,:584/:590test_R14_kernel_package_*,:639test_pending_fast_skips_proxmox_docker_and_kernel(a pendingpve-manager9.2.21 from „Proxmox Debian Repository" is NOT taken). - What a Proxmox step restarts (spike
partH/H5-proxmox-simulate.txt, read from the maintainer scripts, not run):pve-manager→ pvedaemon, pveproxy, pvescheduler, pvestatd, spiceproxy;pve-cluster→ pmxcfs (/etc/pvegone for seconds — the agent's API calls fail meanwhile);qemu-server,pve-firewall,pve-ha-*,corosync,chrony,zfs-zed. No guest is stopped by these scripts (read, not measured). One new package (proxmox-firewall-data). - The kernel (
11§5.6, R-836, measured on demo-hp with the operator's word): installing a kernel makes it the GRUB default at once;--next-bootandgrub-rebootare NOT one-shots here (/bootis ext4 on LVM; GRUB cannot write its env block); a kernel that hangs before userspace stays the default on every power cycle.sp5100_tcoexists on demo-hp, blacklisted, never armed. demo-felhom (Intel N100) has no watchdog measured.
3. Options
A. A host slow lane for the Proxmox USERSPACE packages, no kernel, no reboot.
- What: a new layer
pve(laneslow) in the wrapper: origin „Proxmox Debian Repository" only, still never a name inHOST_SLOW_RE(kernel, boot, firmware, microcode stay out), no removals, no new package except a named allow-list (theproxmox-firewall-datakind). Approval like the Docker engine set (11§5.8): the operator approves one SET per box generation on the System page after ring 0 ran it; ring 1 only by signed job. Runs in the night leg after the host Debian step, under the same heavy-op gate. Health = the host rule (HostHealthVerdict) +pveversionreports the new version. - Cost: wrapper layer + refusals + tests; hub candidate/approval/System page (copies the Docker set's code); ~1.5 days.
- Can go wrong: pmxcfs restart while the agent writes
/etc/pve(the gate already excludes backups/restore-tests; the agent's own reconcile must wait too); a Proxmox point release that needs a newer kernel (proxmox-vedepends onproxmox-default-kernel— the plan must not pull the kernel in: the simulation refusal R14 already catches it); undo is only „install the previous version" (Proxmox keeps 30–66 versions,11C2 — good). - Brings the 77–78 pending packages, minus the boot-chain ones, to every box. Leaves
shim-signed,pve-firmware, microcode and the kernel out.
B. The kernel lane, with a fallback that works on these hosts — measured first.
- What: install the kernel (plus
shim-signed,pve-firmware, microcode, which also only act at boot), keep the old one as the permanent default, boot the new one ONCE, and make it the default only from userspace after a healthy boot. The one-shot needs one of: (1) UEFIBootNextwith a second boot entry whose GRUB config defaults to the new kernel — the firmware clears BootNext itself on use, so a hang falls back on the next power cycle; (2) a GRUB env block on the ESP (vfat, writable by GRUB); (3) arming a hardware watchdog early (sp5100_tcoon demo-hp) so a hang power-cycles into the old default. None is measured. - Cost: a spike with ≥4 reboots per box on both demo hosts (Secure Boot ON on demo-hp, OFF on demo-felhom; AMD vs Intel), then the lane: ~3–5 days. Every customer box restarts all its apps once per kernel (≈1–2 min at night).
- Can go wrong: firmware that ignores
BootNext(seen on consumer boards); Secure Boot refusing a second entry; a hang with no watchdog still needs a person — it is a dead box at a household. Telling households that the box may restart at night is a promise to users (11§5.7) — the operator's.
C. The kernel only by an operator-present maintenance step.
- What: no automatic kernel lane. The System page shows „kernel behind / reboot needed"; the operator runs the step per box by a signed job at a time they choose, ready to power-cycle (or ask the household to).
- Cost: small (a signed job and a runbook), but a person per box per kernel; does not scale past a handful of boxes.
- Can go wrong: kernels are never installed because nobody schedules them — the state today with a button.
4. The pick — PROPOSAL for the operator, not a decision
A now, C for the kernel until B is measured. A closes the biggest gap (78 Proxmox packages, the web API and qemu/lxc tooling) with no reboot and the same approval shape the operator already uses for Docker. The kernel stays a separate, operator-scheduled act (C) until R-836's fallback is measured on both demo hosts; then B replaces C.
R-812 should split. Close R-812 (P2) when A ships to ring 0 and ring 1, with R-836 (kernel fallback, P3) carrying the kernel. The kernel's urgency then becomes its own question: a P2 row „kernel security fixes reach a box" only if the operator wants it before the first paying customer (question 2).
5. First slice and its proof
- Build: wrapper layer
pvelaneslow— origin „Proxmox Debian Repository" only; R14 still applies; R2 for any other origin; refusal for a removal or an unlisted new package; appliance only (R12's root-owned install record). - Red test first (must FAIL on today's code): a host plan with
pve-manager 9.2.21origin „Proxmox Debian Repository", layerpve, laneslow→ today refused R12 („layer is not guest, host or docker"); after: installed. Companion tests that must stay refused: the same plan carryingproxmox-kernel-7.0(R14) orshim-signed(R14); a Debian package in apveplan (R2);test_pending_fast_skips_proxmox_docker_and_kernelunchanged. - Live proof on demo-felhom (ring 0, Tier 0, its Proxmox is the install-time one): one approved set of the pending
Proxmox packages, no kernel. Positive observable:
pveversionreads 9.2.21; the host health rule passes; the guest kept running (containerStartedAtunchanged). Control from a different channel: the hub's System page row for the box, and an app on the guest answering 200 every 5 s through the step. No reboot needed, so no operator word for a reboot; the operator's word is the approval click, as for Docker. Evidence off the box at the end.
6. Open questions for the operator
- May Proxmox's own packages (no kernel) update by night, after you approve each set on the System page, as the Docker engine does? Pick: yes (option A). If you do nothing: every box keeps its install-time Proxmox; 78 packages behind on the demo boxes today, and growing.
- Must kernel security fixes reach customer boxes before the first paying customer? If yes, B is a ~1-week arc with reboots on both demo boxes (your word before each). Pick: no — the kernel by your hand (C) until B is measured. If you do nothing: kernels stay at the install-time version; a kernel hole stays open until you run the step by hand.