Files
felhom.eu/documentation/audits/night-burndown-2026-10-06/design-R-812.md
T

9.6 KiB
Raw Blame History

R-812 — what is left of OS updates: the Proxmox packages and the kernel — design proposal (burn-down night 2026-10-06, no code)

Baselines read: felhom-agent 74b5eae, felhom.eu 8e2dc204. Architecture: 11-os-updates.md (§1, §3, §5.2, §5.5, §5.6, §5.9, §7.1, §8). Related rows: R-836 (the kernel's fallback), R-808 (the intention, ROADMAP.md).

1. The problem

11 §8 steps 1–5 are BUILT: the guest's Debian, the host's Debian, the fleet view, the Docker engine, the root files (R-840, closed). Step 6 is not: the host's Proxmox packages and its kernel are never updated on any box. The host fast lane takes only Debian-origin packages, so everything from the Proxmox repository waits for ever — including packages with ordinary names (zfsutils-linux, shim-signed, corosync, ceph-common, amd64-microcode, 11 C3).

Measured tonight (read only; apt list --upgradable, pveversion, dpkg -l, ls /boot; no apt update — the package lists are from 2026-10-05, the last daily refresh):

  • demo-hp: pve-manager 9.2.2 → 9.2.21 pending; 77 packages pending, all from the Proxmox repository (the Debian lane has taken the rest): qemu-server 9.1.15 → 9.2.10, pve-container 6.1.10 → 6.1.14, proxmox-backup-client 4.2.0 → 4.2.7, zfsutils-linux 2.4.2 → 2.4.4, shim-signed 1.48 → 1.51, pve-firmware, corosync, ceph-*, frr, amd64-microcode, libpve-*. Kernel: runs 7.0.14-20 (installed BY HAND in the 2026-10-04 spike), 7.0.2-6 kept.
  • demo-felhom: pve-manager 9.2.2; 78 pending; runs kernel 7.0.2-6, the install-time kernel — 7.0.14-20 is in the repository and not installed.
  • So a box keeps its install-time Proxmox and kernel. Proxmox security fixes (for example in pve-manager's web API, in qemu-server, in the Secure Boot loader shim-signed) never arrive. The spike counted 79 Proxmox packages "not covered" on the host (11 §7.1).

2. What the code does today (read in source)

  • The fast-lane origin rule: felhom-agent/configs/felhom-os-apply:63 FAST_ORIGINS = ("Debian", "Debian-Security"); a plan package of any other origin is refused R2 (:491-492, and in the simulation :1164 origin_ok).
  • The host's name rule on top: :91-92 HOST_SLOW_RE (linux-*, proxmox-kernel*, pve-kernel*, pve-firmware, firmware-*, grub*, shim*, systemd-boot*, *-microcode, efibootmgr) → refusal R14 (:493-494, :1234-1235); the pending selection skips them (:1136).
  • Layers and lanes: :437-442 — guest/host are lane fast only; docker is the only slow layer. There is no host slow layer.
  • Tests pinning this: configs/test_felhom_os_apply.py:417 test_R2_non_debian_origin_in_the_plan, :422 test_R2_non_debian_origin_in_the_simulation, :584/:590 test_R14_kernel_package_*, :639 test_pending_fast_skips_proxmox_docker_and_kernel (a pending pve-manager 9.2.21 from „Proxmox Debian Repository" is NOT taken).
  • What a Proxmox step restarts (spike partH/H5-proxmox-simulate.txt, read from the maintainer scripts, not run): pve-manager → pvedaemon, pveproxy, pvescheduler, pvestatd, spiceproxy; pve-cluster → pmxcfs (/etc/pve gone for seconds — the agent's API calls fail meanwhile); qemu-server, pve-firewall, pve-ha-*, corosync, chrony, zfs-zed. No guest is stopped by these scripts (read, not measured). One new package (proxmox-firewall-data).
  • The kernel (11 §5.6, R-836, measured on demo-hp with the operator's word): installing a kernel makes it the GRUB default at once; --next-boot and grub-reboot are NOT one-shots here (/boot is ext4 on LVM; GRUB cannot write its env block); a kernel that hangs before userspace stays the default on every power cycle. sp5100_tco exists on demo-hp, blacklisted, never armed. demo-felhom (Intel N100) has no watchdog measured.

3. Options

A. A host slow lane for the Proxmox USERSPACE packages, no kernel, no reboot.

  • What: a new layer pve (lane slow) in the wrapper: origin „Proxmox Debian Repository" only, still never a name in HOST_SLOW_RE (kernel, boot, firmware, microcode stay out), no removals, no new package except a named allow-list (the proxmox-firewall-data kind). Approval like the Docker engine set (11 §5.8): the operator approves one SET per box generation on the System page after ring 0 ran it; ring 1 only by signed job. Runs in the night leg after the host Debian step, under the same heavy-op gate. Health = the host rule (HostHealthVerdict) + pveversion reports the new version.
  • Cost: wrapper layer + refusals + tests; hub candidate/approval/System page (copies the Docker set's code); ~1.5 days.
  • Can go wrong: pmxcfs restart while the agent writes /etc/pve (the gate already excludes backups/restore-tests; the agent's own reconcile must wait too); a Proxmox point release that needs a newer kernel (proxmox-ve depends on proxmox-default-kernel — the plan must not pull the kernel in: the simulation refusal R14 already catches it); undo is only „install the previous version" (Proxmox keeps 30–66 versions, 11 C2 — good).
  • Brings the 77–78 pending packages, minus the boot-chain ones, to every box. Leaves shim-signed, pve-firmware, microcode and the kernel out.

B. The kernel lane, with a fallback that works on these hosts — measured first.

  • What: install the kernel (plus shim-signed, pve-firmware, microcode, which also only act at boot), keep the old one as the permanent default, boot the new one ONCE, and make it the default only from userspace after a healthy boot. The one-shot needs one of: (1) UEFI BootNext with a second boot entry whose GRUB config defaults to the new kernel — the firmware clears BootNext itself on use, so a hang falls back on the next power cycle; (2) a GRUB env block on the ESP (vfat, writable by GRUB); (3) arming a hardware watchdog early (sp5100_tco on demo-hp) so a hang power-cycles into the old default. None is measured.
  • Cost: a spike with ≥4 reboots per box on both demo hosts (Secure Boot ON on demo-hp, OFF on demo-felhom; AMD vs Intel), then the lane: ~3–5 days. Every customer box restarts all its apps once per kernel (≈1–2 min at night).
  • Can go wrong: firmware that ignores BootNext (seen on consumer boards); Secure Boot refusing a second entry; a hang with no watchdog still needs a person — it is a dead box at a household. Telling households that the box may restart at night is a promise to users (11 §5.7) — the operator's.

C. The kernel only by an operator-present maintenance step.

  • What: no automatic kernel lane. The System page shows „kernel behind / reboot needed"; the operator runs the step per box by a signed job at a time they choose, ready to power-cycle (or ask the household to).
  • Cost: small (a signed job and a runbook), but a person per box per kernel; does not scale past a handful of boxes.
  • Can go wrong: kernels are never installed because nobody schedules them — the state today with a button.

4. The pick — PROPOSAL for the operator, not a decision

A now, C for the kernel until B is measured. A closes the biggest gap (78 Proxmox packages, the web API and qemu/lxc tooling) with no reboot and the same approval shape the operator already uses for Docker. The kernel stays a separate, operator-scheduled act (C) until R-836's fallback is measured on both demo hosts; then B replaces C.

R-812 should split. Close R-812 (P2) when A ships to ring 0 and ring 1, with R-836 (kernel fallback, P3) carrying the kernel. The kernel's urgency then becomes its own question: a P2 row „kernel security fixes reach a box" only if the operator wants it before the first paying customer (question 2).

5. First slice and its proof

  • Build: wrapper layer pve lane slow — origin „Proxmox Debian Repository" only; R14 still applies; R2 for any other origin; refusal for a removal or an unlisted new package; appliance only (R12's root-owned install record).
  • Red test first (must FAIL on today's code): a host plan with pve-manager 9.2.21 origin „Proxmox Debian Repository", layer pve, lane slow → today refused R12 („layer is not guest, host or docker"); after: installed. Companion tests that must stay refused: the same plan carrying proxmox-kernel-7.0 (R14) or shim-signed (R14); a Debian package in a pve plan (R2); test_pending_fast_skips_proxmox_docker_and_kernel unchanged.
  • Live proof on demo-felhom (ring 0, Tier 0, its Proxmox is the install-time one): one approved set of the pending Proxmox packages, no kernel. Positive observable: pveversion reads 9.2.21; the host health rule passes; the guest kept running (container StartedAt unchanged). Control from a different channel: the hub's System page row for the box, and an app on the guest answering 200 every 5 s through the step. No reboot needed, so no operator word for a reboot; the operator's word is the approval click, as for Docker. Evidence off the box at the end.

6. Open questions for the operator

  1. May Proxmox's own packages (no kernel) update by night, after you approve each set on the System page, as the Docker engine does? Pick: yes (option A). If you do nothing: every box keeps its install-time Proxmox; 78 packages behind on the demo boxes today, and growing.
  2. Must kernel security fixes reach customer boxes before the first paying customer? If yes, B is a ~1-week arc with reboots on both demo boxes (your word before each). Pick: no — the kernel by your hand (C) until B is measured. If you do nothing: kernels stay at the install-time version; a kernel hole stays open until you run the step by hand.