Files
felhom.eu/documentation/audits/night-burndown-2026-10-06/design-R-812.md
T

109 lines
9.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# R-812 — what is left of OS updates: the Proxmox packages and the kernel — design proposal (burn-down night 2026-10-06, no code)
Baselines read: felhom-agent `74b5eae`, felhom.eu `8e2dc204`. Architecture: `11-os-updates.md` (§1, §3, §5.2, §5.5, §5.6,
§5.9, §7.1, §8). Related rows: R-836 (the kernel's fallback), R-808 (the intention, `ROADMAP.md`).
## 1. The problem
`11` §8 steps 1–5 are BUILT: the guest's Debian, the host's Debian, the fleet view, the Docker engine, the root files
(R-840, closed). Step 6 is not: **the host's Proxmox packages and its kernel are never updated on any box.** The host
fast lane takes only Debian-origin packages, so everything from the Proxmox repository waits for ever — including
packages with ordinary names (`zfsutils-linux`, `shim-signed`, `corosync`, `ceph-common`, `amd64-microcode`, `11` C3).
Measured tonight (read only; `apt list --upgradable`, `pveversion`, `dpkg -l`, `ls /boot`; no `apt update` — the package
lists are from 2026-10-05, the last daily refresh):
- **demo-hp:** `pve-manager` 9.2.2 → 9.2.21 pending; **77 packages pending, all from the Proxmox repository** (the Debian
lane has taken the rest): `qemu-server` 9.1.15 → 9.2.10, `pve-container` 6.1.10 → 6.1.14, `proxmox-backup-client`
4.2.0 → 4.2.7, `zfsutils-linux` 2.4.2 → 2.4.4, `shim-signed` 1.48 → 1.51, `pve-firmware`, `corosync`, `ceph-*`, `frr`,
`amd64-microcode`, `libpve-*`. Kernel: runs 7.0.14-20 (installed BY HAND in the 2026-10-04 spike), 7.0.2-6 kept.
- **demo-felhom:** `pve-manager` 9.2.2; **78 pending**; runs kernel **7.0.2-6, the install-time kernel** — 7.0.14-20 is in
the repository and not installed.
- So a box keeps its install-time Proxmox and kernel. Proxmox security fixes (for example in `pve-manager`'s web API, in
`qemu-server`, in the Secure Boot loader `shim-signed`) never arrive. The spike counted 79 Proxmox packages "not
covered" on the host (`11` §7.1).
## 2. What the code does today (read in source)
- The fast-lane origin rule: `felhom-agent/configs/felhom-os-apply:63` `FAST_ORIGINS = ("Debian", "Debian-Security")`;
a plan package of any other origin is refused **R2** (`:491-492`, and in the simulation `:1164` `origin_ok`).
- The host's name rule on top: `:91-92` `HOST_SLOW_RE` (`linux-*`, `proxmox-kernel*`, `pve-kernel*`, `pve-firmware`,
`firmware-*`, `grub*`, `shim*`, `systemd-boot*`, `*-microcode`, `efibootmgr`) → refusal **R14** (`:493-494`, `:1234-1235`);
the pending selection skips them (`:1136`).
- Layers and lanes: `:437-442` — `guest`/`host` are lane `fast` only; `docker` is the only `slow` layer. There is no
host slow layer.
- Tests pinning this: `configs/test_felhom_os_apply.py:417` `test_R2_non_debian_origin_in_the_plan`, `:422`
`test_R2_non_debian_origin_in_the_simulation`, `:584`/`:590` `test_R14_kernel_package_*`, `:639`
`test_pending_fast_skips_proxmox_docker_and_kernel` (a pending `pve-manager` 9.2.21 from „Proxmox Debian Repository"
is NOT taken).
- What a Proxmox step restarts (spike `partH/H5-proxmox-simulate.txt`, read from the maintainer scripts, not run):
`pve-manager` → pvedaemon, pveproxy, pvescheduler, pvestatd, spiceproxy; `pve-cluster` → pmxcfs (`/etc/pve` gone for
seconds — the agent's API calls fail meanwhile); `qemu-server`, `pve-firewall`, `pve-ha-*`, `corosync`, `chrony`,
`zfs-zed`. **No guest is stopped by these scripts** (read, not measured). One new package (`proxmox-firewall-data`).
- The kernel (`11` §5.6, R-836, measured on demo-hp with the operator's word): installing a kernel makes it the GRUB
default at once; `--next-boot` and `grub-reboot` are NOT one-shots here (`/boot` is ext4 on LVM; GRUB cannot write its
env block); a kernel that hangs before userspace stays the default on every power cycle. `sp5100_tco` exists on
demo-hp, blacklisted, never armed. demo-felhom (Intel N100) has no watchdog measured.
## 3. Options
**A. A host slow lane for the Proxmox USERSPACE packages, no kernel, no reboot.**
- What: a new layer `pve` (lane `slow`) in the wrapper: origin „Proxmox Debian Repository" only, still never a name in
`HOST_SLOW_RE` (kernel, boot, firmware, microcode stay out), no removals, no new package except a named allow-list (the
`proxmox-firewall-data` kind). Approval like the Docker engine set (`11` §5.8): the operator approves one SET per box
generation on the System page after ring 0 ran it; ring 1 only by signed job. Runs in the night leg after the host
Debian step, under the same heavy-op gate. Health = the host rule (`HostHealthVerdict`) + `pveversion` reports the
new version.
- Cost: wrapper layer + refusals + tests; hub candidate/approval/System page (copies the Docker set's code); ~1.5 days.
- Can go wrong: pmxcfs restart while the agent writes `/etc/pve` (the gate already excludes backups/restore-tests; the
agent's own reconcile must wait too); a Proxmox point release that needs a newer kernel (`proxmox-ve` depends on
`proxmox-default-kernel` — the plan must not pull the kernel in: the simulation refusal R14 already catches it);
undo is only „install the previous version" (Proxmox keeps 30–66 versions, `11` C2 — good).
- Brings the 77–78 pending packages, minus the boot-chain ones, to every box. Leaves `shim-signed`, `pve-firmware`,
microcode and the kernel out.
**B. The kernel lane, with a fallback that works on these hosts — measured first.**
- What: install the kernel (plus `shim-signed`, `pve-firmware`, microcode, which also only act at boot), keep the old
one as the permanent default, boot the new one ONCE, and make it the default only from userspace after a healthy boot.
The one-shot needs one of: (1) **UEFI `BootNext`** with a second boot entry whose GRUB config defaults to the new kernel
— the firmware clears BootNext itself on use, so a hang falls back on the next power cycle; (2) a GRUB env block on
the ESP (vfat, writable by GRUB); (3) arming a hardware watchdog early (`sp5100_tco` on demo-hp) so a hang power-cycles
into the old default. None is measured.
- Cost: a spike with ≥4 reboots per box on both demo hosts (Secure Boot ON on demo-hp, OFF on demo-felhom; AMD vs Intel),
then the lane: ~3–5 days. Every customer box restarts all its apps once per kernel (≈1–2 min at night).
- Can go wrong: firmware that ignores `BootNext` (seen on consumer boards); Secure Boot refusing a second entry; a
hang with no watchdog still needs a person — it is a dead box at a household. Telling households that the box may
restart at night is a **promise to users** (`11` §5.7) — the operator's.
**C. The kernel only by an operator-present maintenance step.**
- What: no automatic kernel lane. The System page shows „kernel behind / reboot needed"; the operator runs the step per
box by a signed job at a time they choose, ready to power-cycle (or ask the household to).
- Cost: small (a signed job and a runbook), but a person per box per kernel; does not scale past a handful of boxes.
- Can go wrong: kernels are never installed because nobody schedules them — the state today with a button.
## 4. The pick — PROPOSAL for the operator, not a decision
**A now, C for the kernel until B is measured.** A closes the biggest gap (78 Proxmox packages, the web API and
qemu/lxc tooling) with no reboot and the same approval shape the operator already uses for Docker. The kernel stays a
separate, operator-scheduled act (C) until R-836's fallback is measured on both demo hosts; then B replaces C.
**R-812 should split.** Close R-812 (P2) when A ships to ring 0 and ring 1, with R-836 (kernel fallback, P3) carrying
the kernel. The kernel's urgency then becomes its own question: a P2 row „kernel security fixes reach a box" only if the
operator wants it before the first paying customer (question 2).
## 5. First slice and its proof
- Build: wrapper layer `pve` lane `slow` — origin „Proxmox Debian Repository" only; R14 still applies; R2 for any other
origin; refusal for a removal or an unlisted new package; appliance only (R12's root-owned install record).
- Red test first (must FAIL on today's code): a host plan with `pve-manager 9.2.21` origin „Proxmox Debian Repository",
layer `pve`, lane `slow` → today refused R12 („layer is not guest, host or docker"); after: installed. Companion
tests that must stay refused: the same plan carrying `proxmox-kernel-7.0` (R14) or `shim-signed` (R14); a Debian
package in a `pve` plan (R2); `test_pending_fast_skips_proxmox_docker_and_kernel` unchanged.
- Live proof on **demo-felhom** (ring 0, Tier 0, its Proxmox is the install-time one): one approved set of the pending
Proxmox packages, no kernel. Positive observable: `pveversion` reads 9.2.21; the host health rule passes; the guest
kept running (container `StartedAt` unchanged). Control from a different channel: the hub's System page row for the
box, and an app on the guest answering 200 every 5 s through the step. No reboot needed, so no operator word for a
reboot; the operator's word is the approval click, as for Docker. Evidence off the box at the end.
## 6. Open questions for the operator
1. **May Proxmox's own packages (no kernel) update by night, after you approve each set on the System page, as the
Docker engine does?** Pick: yes (option A). If you do nothing: every box keeps its install-time Proxmox; 78 packages
behind on the demo boxes today, and growing.
2. **Must kernel security fixes reach customer boxes before the first paying customer?** If yes, B is a ~1-week arc with
reboots on both demo boxes (your word before each). Pick: no — the kernel by your hand (C) until B is measured. If you
do nothing: kernels stay at the install-time version; a kernel hole stays open until you run the step by hand.