Files
felhom.eu/documentation/architecture/11-os-updates.md
T

37 KiB
Raw Blame History

11 — Operating-system updates: the host, the guest and the Docker engine

Status NOT RATIFIED — a PROPOSAL with operator rulings, corrected by the 2026-10-04 spike (§7.1, C1–C12); §8 step 2 BUILT 2026-10-04 (§8.1). Ratification is Viktor's review, not an editor's.
Written 2026-10-04, by the reviewer (project Claude), before any spike.
Verified against felhom.eu d07a1a9 · felhom-controller 99a1497 (v0.290.0) · felhom-agent d766666 (v0.138.0) · hub v0.128.0
Freshness CURRENT as of 2026-10-04. The spike TASK-backup-close-and-os-updates-spike-2026-10-04 adds measurements here as [FACT] and corrects every claim it disproves. Mark this file STALE when it falls behind what the product does.

How to read this document. Each statement has a label:

  • [FACT]: an observed property, with a file:line, a register row or an audit path.
  • [RULED]: an operator decision, with its date.
  • [PROPOSAL]: the reviewer's design. It is not decided and not built.
  • OPEN: a question that nobody has answered yet. The spike measures it. Nobody guesses it.

Why this file exists. No architecture document covered operating-system updates. 00 §G marks the capability MISSING (added 2026-10-03). The finding is R-812. The roadmap intention is R-808. The register carries the work. This file carries the reasoning. The source is the truth.

Spike corrections, 2026-10-04 (audits/os-updates-spike-2026-10-04/). Each is made in place: the reviewer's text is struck (like this) and the measured text follows, labelled [FACT], with its id.

Id Where What the spike changed
C1 §2 The agent may run apt-get install for TWO packages, not one (dnsmasq and wireguard-tools).
C2 §5.3 Debian keeps two versions, not one; for box packages no fix was replaced within 14 days in 3 months; snapshot.debian.org works from a box in seconds.
C3 §5.2 The lanes must follow the package's ORIGIN, not its name: 40 Proxmox-repository packages have ordinary names (ZFS, the Secure Boot shim, Ceph, corosync, chrony, CPU microcode).
C4 §5.6 --next-boot is NOT a one-shot on these GRUB hosts. The fallback works only after a boot that reaches userspace. Installing a kernel alone makes it the default. Only a software watchdog runs.
C5 §5.6, §6 row 4 Docker's live-restore keeps every container running across an engine update (measured). Turning it OFF again stops every container and starts none.
C6 §5.6 A guest snapshot works on LVM-thin (customer guests) but not on dir storage; the snapshot rollback itself is unmeasured; a backup-restore undo took 73 s.
C7 §6 row 2 A killed apt run does not recover by itself; the repair took ~5 s.
C8 §1, §6 row 14 cloudflared is not a host package: it is a container in the guest, pinned by the controller since June. It belongs with the controller's infrastructure pins, not this file's lanes.
C9 §5.3 The approved list must record what ring 0 RUNS healthy, not only what it installed that night.
C10 §5.5 Restore-tests and agent updates are not windowed; the host's apt timers install nothing today.
C11 §5.2 A Debian (fast-lane) update leaves PID 1, lxc-start, the Proxmox daemons and dockerd on the old library: its full effect needs a restart the fast lane does not do.
C12 §6 row 9 The two demo hosts differ by 7 packages, including the CPU microcode (AMD vs Intel) and Secure Boot (on vs off).

0. In plain language

A box runs three layers that we install and never update: the Proxmox host, the small Debian system inside the guest, and the Docker engine. App images are updated (see 09). The layer under them is not. A box lives in a home for years, so this is a security gap.

The proposal has two lanes. The fast lane applies Debian's security fixes automatically, but only the exact versions that already ran well on the demo boxes for 1–2 days. The slow lane covers the kernel, Proxmox and Docker. These cause restarts or big changes, so a person approves them one version at a time, and they run only at night. The tested versions are recorded automatically from what the demo boxes installed. Nobody keeps a hand-written list.


1. Scope

In scope.

  • The Proxmox host on an appliance install: the Debian 13 base, the Proxmox packages, and the kernel.
  • The guest (the customer LXC): its Debian 13 packages.
  • The Docker engine inside the guest (docker-ce, docker-ce-cli, containerd.io).
  • How the household and the operator are told, and how a failed update is undone.

Out of scope.

  • App images. 09 covers them (the ladder, the monthly same-tag re-test).
  • The controller and agent binaries. Their own self-update covers them (03, the self-update section; 09 R-608).
  • DooPlex and ep0. The operator updates them by hand (runbooks/offsite-endpoint.md).
  • A BYO host (the owner brought their own Proxmox). [FACT] The installer leaves its repositories alone: "apt repo alignment skipped (byo — the owner manages repos)" (scripts/felhom-host-install.sh, align_apt_repos). [PROPOSAL] On a BYO box we update the guest and Docker only, never the host.
  • The Proxmox MAJOR upgrade (PVE 9 → 10). It is a later step of its own, drilled first (R-808 item 5).
  • cloudflared and the other infrastructure images (traefik, filebrowser). [FACT] (C8) They are containers in the guest, pinned in the controller (internal/infra/infra.go:26: cloudflare/cloudflared:2026.6.0, since 2026-06-11) and baked into the golden; upstream was 2026.9.3 on 2026-10-04. They move only by a controller release — the app-image question (09), not an OS package. The gap is R-838.

2. What a box runs, as measured

Layer What it is Source of packages Evidence
Host Proxmox VE 9.2 on Debian 13, LVM-thin Debian mirrors + pve-no-subscription (enterprise repo switched off on appliance installs) [FACT] 01 §2; felhom-host-install.sh align_apt_repos (~L2130–2136)
Host kernel The only kernel on the box. The guest has none. pve-no-subscription (proxmox-kernel-*) [FACT] LXC shares the host kernel (01 §2: "one LXC/kernel/Docker daemon")
Guest Unprivileged LXC, nesting=1,keyctl=1, from debian-13-standard_13.1-2 Debian mirrors [FACT] felhom-agent/configs/build-golden.sh:67,101-103
Docker engine docker-ce, docker-ce-cli, containerd.io, installed at golden BAKE time download.docker.com/linux/debian trixie stable [FACT] build-golden.sh:108-125
Docker settings containerd-snapshotter: false, json-file log caps. No live-restore. baked daemon.json [FACT] build-golden.sh:138-144
Controller A container in the guest. A Docker engine restart restarts it too. Felhom registry [FACT] 03 §1

[FACT] Nothing updates any of these layers today (R-812, searched 2026-10-03). The installer says "No upgrades are run — repo alignment only" (felhom-host-install.sh ~L2135). A fresh install gets the Docker engine that was current when its golden was baked, and keeps it.

[FACT] The agent may not run apt today, with one exception. Its sudoers allowlist (felhom-agent/configs/felhom-agent.sudoers) holds one apt line: apt-get install -y -q dnsmasq. [FACT] (C1) The agent may run apt-get install for two named packages: felhom-agent.sudoers:58 (apt-get install -y -q dnsmasq, used by internal/lanresolver/lanresolver.go:107) and :194 (apt-get install -y -q wireguard-tools, used by internal/wgtunnel/manager.go:628). Nothing else. apt as root runs package scripts as root, so a broad apt grant is a full root grant. §5.4 proposes how to avoid that.


3. Operator rulings

2026-10-03. R-808 is on the roadmap at P2: "every box receives operating-system security patches on a schedule, and a failed update is undone."

[RULED] 2026-10-04: the fast lane follows an approved list, with a 1–2 day wait. Every update runs on the demo boxes first. The other boxes install only the exact versions that the demo boxes ran without trouble. The exact wait is set when the feature is built. Rejected: Debian's own unattended-upgrades with no wait. It is simpler and common, but a bad update would reach every customer at the same time.

[RULED] 2026-10-04: the off-site backup topic is closed first. The same brief carries both topics, and the backup part runs first.

The two-lane split (§5.2) is the reviewer's proposal. The operator's ruling above assumes it, but he has not ruled on it as such.


4. The constraints that shape the design

  1. Unattended. Nobody is at the box. Nobody answers a question that apt asks.
  2. No screen. If the box does not boot, the household sees only that nothing works.
  3. One kernel for everything. A host kernel update needs a host reboot. A reboot stops every app.
  4. The controller lives in Docker. A Docker engine update restarts the controller in the middle of its own work. So the controller cannot drive a Docker update. The agent, on the host, must.
  5. The host has no whole-system backup. The guest has three tiers (07). The host has none. A host update can only be undone by installing the previous version again.
  6. Every update must already have run on a box we own. This is the lesson of the update arc (09 §3 decision 13: "the test decides").
  7. Few packages. Each extra package is one more thing to update and break. The guest is the Debian standard template plus Docker. Keep it that way.

5. The proposed shape [PROPOSAL]

5.1 Two rings

  • Ring 0: demo-felhom (N100) and demo-hp. They are disposable (runbooks/target-selection.md), and they have different hardware. They take every update first.
  • Ring 1: every other box. It takes only what ring 0 approved.

5.2 Two lanes

Fast lane Slow lane
What Debian packages on host and guest, from trixie-security and the stable point releases, except the slow-lane list Kernel (proxmox-kernel-*), Proxmox (pve-*, proxmox-*, lxc-pve, qemu-server, …), Docker (docker-ce*, containerd.io)
Restart A service restart at most. No reboot. Kernel: a host reboot. Docker: every container restarts. Proxmox: its services restart.
Approval Automatic. A version is approved when ring 0 ran it and stayed healthy for the wait (1–2 days). A person approves one version at a time, like a controller floor.
Urgent fix The operator can approve a version on the same day once ring 0 has run it. The same.
When The night window (§5.5) The night window. A kernel reboot only on a night the operator scheduled.

Which packages count as slow lane is a list the spike checks (OPEN Q7). A Debian package that restarts something big (for example systemd, libc6, openssh-server) may belong in the slow lane too.

[FACT] (C3) The lane must be decided by ORIGIN, not by name. On demo-hp, 40 pending packages come from the Proxmox repository under ordinary names: zfsutils-linux, zfs-zed, libzfs7linux…, shim-signed and friends (the Secure Boot loader), ceph-common/librados2…, corosync, chrony, frr, amd64-microcode (partH/H2-proxmox-origin-debian-names.txt). A name rule (pve-*, proxmox-*) would have put them in the fast lane. The fast lane is: origin Debian or Debian-Security, and nothing else. On demo-hp that selection was 108 packages and pulled in zero Proxmox packages.

[FACT] (C11) What a fast-lane run restarts, and what it does not. Guest (9202, 49 packages incl. libc6): the packages' own scripts restarted postfix, journald, networkd; dockerd, containerd, sshd, dbus, logind, cron kept the old libc; no container stopped (13 samples). Host (demo-hp, 108 packages): dnsmasq, postfix, journald restarted; systemd (PID 1), lxc-start, pveproxy, pvedaemon, pvestatd, pvescheduler, watchdog-mux, sshd, zed, chronyd kept the old libc; both guests and the agent stayed up. So the fast lane is safe to run unattended, but a libc fix is only fully in force after a reboot (host) or a Docker restart (guest) — which are slow-lane acts. [PROPOSAL] the box reports "restart needed" (processes on deleted libraries) and the slow lane's next reboot picks it up.

5.3 The approved list (the "tested versions" record)

Nobody writes the list by hand. It fills itself:

  1. A ring-0 box updates. Afterwards it reports to the hub the exact package=version it installed, per layer (host, guest), with each package's origin (Debian-Security, Debian, Proxmox, Docker).
  2. The box then reports health for the wait period: the agent, the controller, every app's health, and the guest's network.
  3. When the wait passes with ring 0 healthy, the hub marks that set approved. It is one record: an OS release, with an id and a date. The hub stores it. The register does not.
  4. A ring-1 box asks the hub for the newest approved OS release. For each package it has installed, if the approved version is newer, it installs that exact version. It never installs a version newer than the approved one.
  5. Each box reports which OS release it runs, how many updates are waiting, whether it needs a reboot, and which packages it has that no approved list covers (see §6, edge case 9).

OPEN Q1 is the weak point. The design works only if a box can still download the approved version a week later. Debian's main and security archives keep only the newest version of each package. If Debian publishes a newer fix between approval and install, the approved version is gone. Options: the box waits for the next approval; or the box uses snapshot.debian.org (Debian's own dated archive, still signed by Debian); or Felhom runs a package cache. The spike measures how often this happens.

[FACT] (C2) Q1 measured. The live Debian archives keep two versions: the point-release one in trixie (main) and the newest in trixie-security; intermediate versions are gone (openssl: installed u1, main u2, security u3). Over 2026-07-04..10-04, 167 trixie security advisories; for the 517 source packages installed on a box, no package got a second advisory within 2, 7 or 14 days (the 3 within 2 days were chromium and webkit2gtk, not on a box). snapshot.debian.org answers a box: a dated index in 2.3–3.0 s, a gone exact version (openssl 3.5.6-1~deb13u1) downloaded in 2.0 s, Debian-signed. Proxmox and Docker keep many old versions (pve-manager 66, docker-ce 46). So with a 1–2 day wait the approved version is almost always still live; the rare miss is fetched from the snapshot taken at approval time. [PROPOSAL] each OS release records its approval timestamp; a box installs from its own sources, and only for a Debian package that is no longer there, from snapshot.debian.org/archive/<debian|debian-security>/<timestamp>. This is the operator decision in STATUS.

[FACT] (C9) Approve what ring 0 RUNS, not what it installed. In the simulation on demo-felhom (partI/demo-felhom-simulation.txt), every one of the 108 host and 49 guest approved versions was installable and downloadable — but curl, libcurl* and libssh2 were "not covered" in the guest, only because scratch guest 9202 ALREADY ran the newer version and so installed nothing. [PROPOSAL] step 1 reports the full installed package=version set after the run, and approval covers every version ring 0 runs healthy.

5.4 Who runs it, and with what permission

  • The agent runs every OS update, for the host and for the guest (pct exec). The controller does not, because of constraint 4.
  • The agent gets no general apt permission. Like felhom-selfupdate-guarded (03, the self-update section), a small root-owned wrapper does the work. It accepts only an approved list and refuses everything else:
    • no package removal;
    • no downgrade, except the undo of §5.6, which an operator job signs;
    • no package that the box does not already have, unless the approved list records it as a dependency that the same update pulled in on ring 0 (a kernel update installs a NEW package name each time, for example proxmox-kernel-6.x.y-z-pve-signed, pulled by proxmox-default-kernel);
    • no package source other than the ones the installer set up;
    • non-interactive, and it always keeps the existing config file (--force-confold) and reports the conflict.
  • Package signatures stay the publishers' (Debian, Proxmox, Docker). The hub sends only names and versions. A broken-into hub can choose an older version or no version. It cannot make a box install a package that the publisher did not sign.

5.4.1 The root wrapper's interface — DRAFT (2026-10-04, design only, nothing installed) [PROPOSAL]

felhom-os-apply — root-owned (0755 root:root), installed by the installer beside felhom-selfupdate-guarded; the agent's sudoers gets exactly /usr/local/sbin/felhom-os-apply --plan /var/lib/felhom-agent/os/plan-*.json and … --repair-only. For the guest it runs on the host and enters the guest with pct exec <vmid> -- itself, so the agent needs no pct exec … apt line. Built from what Parts G and H measured.

Input — one JSON plan file (written by the agent, from the hub's approved OS release):

{
  "release_id": "os-2026-10-04-1", "approved_at": "2026-10-04T08:00:00Z",
  "layer": "host",                       // "host" | "guest"
  "vmid": 9201,                          // guest only
  "snapshot": "20261004T080000Z",        // the snapshot.debian.org timestamp of approval (C2)
  "packages": [ {"name": "libc6", "version": "2.41-12+deb13u4", "origin": "Debian"} ],
  "allow_new": ["proxmox-kernel-7.0.14-20-pve-signed"],   // slow lane only, signed operator job
  "lane": "fast"                         // "fast" | "slow"
}

Order of work: (1) refuse checks below; (2) repair first — dpkg --configure -a then apt-get -f install, logging what it repaired (C7); (3) apt-get -s install of exactly name=version for every package that is installed AND older; (4) refuse if the simulation would remove, downgrade, add an unlisted package, or touch a package whose candidate origin is not the plan's; (5) download — from the box's own sources, or for a Debian version no longer there, from snapshot.debian.org/archive/<archive>/<snapshot> with a temporary sources list it deletes after (C2); (6) install with DEBIAN_FRONTEND=noninteractive APT_LISTCHANGES_FRONTEND=none -o Dpkg::Options::=--force-confold -o Dpkg::Options::=--force-confdef; (7) apt-get clean; (8) report.

Refusals (each exits non-zero with one line os-apply: REFUSED: <reason> and changes nothing):

# Refuses when
R1 the plan file is not under /var/lib/felhom-agent/os/, not owned by the agent, or not valid JSON
R2 lane is fast and any package's origin is not Debian / Debian-Security (C3)
R3 lane is slow and the plan is not carried by a verified signed operator job (R-530's mechanism)
R4 the simulation removes any package
R5 the simulation downgrades any package (the operator undo is a separate signed op, §5.6)
R6 the simulation installs a package that is neither installed nor in allow_new
R7 a listed version is not downloadable from the sources the installer set up or the named snapshot
R8 free space on / (or the guest's rootfs) is below 3× the download size, minimum 500 MB (edge case 8)
R9 another apt/dpkg holds the lock, or the per-guest lane lock is held (a backup, a restore-test, C10)
R10 layer is guest and the vmid is not the box's own customer guest
R11 the plan names a package twice, or a version that is not a Debian version string

Log lines (to the journal, tag felhom-os-apply, and echoed for the agent to forward to the hub):

os-apply: START release=<id> layer=<host|guest:vmid> lane=<fast|slow> packages=<n>
os-apply: REPAIR configured=<n> fixed=<n>          (always printed; 0 0 when nothing was half-done)
os-apply: PLAN upgrade=<n> already=<n> not-installed=<n> from-snapshot=<n> download=<bytes>
os-apply: REFUSED: <R-number> <reason>
os-apply: CONFFILE kept <path> (new version saved as <path>.dpkg-dist)
os-apply: DONE rc=0 seconds=<s> upgraded=<n> restarted=<unit,…> restart-needed=<process,…> reboot-needed=<yes|no>
os-apply: FAILED rc=<n> step=<download|install> — dpkg state: <dpkg --audit first line>

restart-needed lists processes still mapping deleted libraries (C11); reboot-needed is yes when that list holds PID 1 or lxc-start, or a kernel was installed.

5.5 When

Inside the household's night window, after the backups:

W        DB dump
W+60m    Tier 2
W+105m   off-site → app updates (until W+5h at most)
[W+2h, W+6h)  whole-guest backup (agent)
after it      OS updates — guest first, then host

The OS leg starts after the whole-guest backup has finished, so the guest's newest full copy is minutes old. This is the opposite order to app updates, which run before the whole-guest backup (07 §6.1, 09 decision 11). The reason: the whole-guest backup is the guest's undo.

[FACT] (C10) Q8 measured (guest UTC; demo-hp W = 02:30): db-dump 02:30, tier-2 03:30, off-site ~04:15, whole-guest gate [04:30, 08:30), controller self-update 04:30, offsite-integrity 06:00. Host: apt-daily and apt-daily-upgrade run daily but install nothing (no unattended-upgrades, no APT::Periodic); pve-daily-update refreshes the lists daily. Restore-tests are NOT windowed — they run on a cadence at any hour (demo-felhom 10:38 daily, demo-hp 16:43 and 22:46); agent updates arrive by signed job at any hour. [PROPOSAL] the OS leg takes the same per-guest lane lock the restore-test and the whole-guest backup take, rather than a clock slot.

OPEN Q8: what else runs then. The controller's self-update (default 04:30, and after any hub report when a floor is above it, R-608), the agent's self-update, the restore-tests, and PBS jobs. The OS leg must never overlap a backup, a restore-test or a self-update.

5.6 How a failed update is undone

"Rollback" is not used (09 §4). The shapes:

Layer Undo Limit
Guest packages Restore the guest snapshot taken just before the update It also undoes app data written after the snapshot. Use it only inside the health window, before apps have written much. After that, install the previous version (needs OPEN Q1). [FACT] (C6) Customer guests are on LVM-thin and can snapshot; scratch 9202 (dir storage) cannot (snapshot feature is not available), so the measured undo was a whole-guest backup + restore: 73 s down, all apps healthy, libc back. The snapshot rollback itself is unmeasured (R-837).
Docker engine Install the previous version (Docker's repository keeps old versions — OPEN Q2) Every container restarts again. [FACT] (C5) Docker keeps 46 docker-ce versions. A step (either way) without live-restore: 6 of 6 containers restart, the app silent 26.5–30 s, healthy at +42–45 s. With live-restore on: 0 restarts, no gap, also across a containerd step. But a restart that turns live-restore OFF stops every container and starts NONE (unless-stopped ignored) — on 9202 they stayed down until restarted by hand; systemctl reload turns it on but not off.
Host packages Install the previous version Only if the source still has it (OPEN Q1/Q2). [FACT] rsync back to its pre-update version: not found; libpng16-16t64 back to the point-release version: worked. A Debian undo needs the snapshot archive (C2).
Host kernel Boot the previous kernel. Proxmox can boot a new kernel once (proxmox-boot-tool kernel pin <ver> --next-boot). If that boot fails, the next boot uses the old kernel again. Make the new kernel permanent only after a healthy boot. [FACT] (C4) Both demo hosts boot UEFI + GRUB (no proxmox-boot-tool ESPs). Installing a kernel makes it the GRUB default at once. --next-boot on GRUB writes an ordinary GRUB_DEFAULT + update-grub; proxmox-boot-cleanup.service clears it only once a boot reaches userspace. Measured on demo-hp (old kernel pinned permanently FIRST, new pinned for next boot): boot 1 → 7.0.14-20-pve, healthy, 60 s; boot 2 → back on 7.0.2-6-pve, healthy, 76 s. So the fallback works after a boot that succeeds; read from the code, a kernel that hangs before userspace stays the default on every power cycle. [PROPOSAL] use GRUB's own one-shot (GRUB_DEFAULT=saved + grub-reboot) with the old kernel as the saved default — unmeasured (R-836). If the new kernel hangs, someone must switch the box off and on. The spike checks whether a hardware watchdog can do that (OPEN Q4). [FACT] Only softdog runs (loaded by watchdog-mux); it cannot rescue a kernel that never boots. demo-hp has an AMD FCH whose sp5100_tco driver ships but is not loaded; untested. A hang still needs a person.

5.7 Telling people

  • Operator: a hub event for each update and each failure; a fleet view showing each box's OS release, how far behind it is, and whether it needs a reboot. A box that is more than N days behind the newest approved release raises an alarm (08).
  • Household: one line on the timeline in both languages, informal voice: what was updated and whether the box restarted. Telling households in advance that the box may restart at night is a promise to users. That is the operator's decision when the slow lane is built.

6. Risks and edge cases

# What can go wrong What the design does
1 A new kernel does not boot Boot it once (--next-boot); the old kernel stays the default. A hard hang needs a power cycle (OPEN Q4).
2 Power cut during an update: dpkg is half done The wrapper runs dpkg --configure -a and apt-get -f install first, every time, and reports what it repaired (spike measures, Q5). [FACT] (C7) Killed after 15 unpacks: 5 packages iU, 4 triggers pending; the next ordinary apt-get install REFUSES (Unmet dependencies) — nothing repairs it by itself. The two commands repaired it in 1.4 s + 3.9 s; apps stayed up.
3 An update asks a question (changed config file, service restart prompt) Non-interactive, keep the old config, report the conflict.
4 A Docker update stops every app, and the controller Stop the apps cleanly first, like before a backup. Measure whether live-restore keeps containers running (Q3). The agent drives it. [FACT] (C5) live-restore keeps them running (0 restarts). The controller keeps running too; its docker calls fail for the seconds dockerd is down (logged errors, no app event). Turning live-restore on is a golden + fleet change — and turning it off later must not be a plain restart (R-835).
5 The approved version is no longer downloadable OPEN Q1. Until it is answered, the box waits for the next approval and reports it.
6 An urgent security hole The operator approves the same day once ring 0 has run it.
7 A box was off for months It catches up through the newest approved release. It does not install every release in between. Packages are not like app data; one apt step is enough. Kernel and Docker still go one approved step at a time.
8 The system disk is full Check free space before downloading; clean the package cache after; refuse and report.
9 A customer box has a package that ring 0 does not have (other hardware: firmware, CPU microcode, NIC drivers) It is never approved, so it is never updated. The box reports it as "not covered", and the hub raises it. Fix: add matching hardware to ring 0, or approve it by hand. [FACT] (C12) The two demo hosts differ by 7 packages: amd64-microcode, proxmox-secure-boot-support, felhom-bootstrap (demo-hp) vs intel-microcode, proxmox-first-boot, tailscale, tailscale-archive-keyring (demo-felhom); Secure Boot is ON on demo-hp, OFF on demo-felhom. Ring 0 covers both CPU vendors today.
10 The package source is down, or its signing key changes (Docker has done this) The update fails cleanly and the box reports it. A key change is a slow-lane act for a person.
11 A BYO host Only the guest and Docker are updated (§1).
12 Two boxes on ring 0 is a small sample Accepted for now. When there are customers, the first tester boxes can become a second ring.
13 A broken-into hub sends a harmful list The wrapper refuses removals, downgrades and new packages, and the publishers' signatures still apply (§5.4).
14 cloudflared on the host Not covered by this file yet. How it is installed and updated is OPEN Q9. It faces the internet, so it matters. [FACT] (C8) Not on the host: a pinned container in the guest (§1). Four months behind upstream on 2026-10-04 (R-838).
15 The no-subscription Proxmox repository gets less testing than the enterprise one Ring 0 is our test. The enterprise repository costs a yearly fee per box. That is a money decision for the operator, later (Q10).

7. Open questions the spike must answer

Q Question How to answer
Q1 Can a box install an exact Debian version one week after approval? How often is it already superseded? apt-cache madison on host and guest for security packages. Debian's security announcement history. Whether snapshot.debian.org is reachable and fast enough.
Q2 Do the Proxmox and Docker sources keep older versions? apt-cache madison for pve-manager, proxmox-kernel-*, docker-ce, containerd.io.
Q3 What does a Docker engine update do to running containers, with and without live-restore? On scratch guest 9202: step from one pinned version to the next; time the downtime.
Q4 Does proxmox-boot-tool kernel pin --next-boot work on both demo hosts' boot setups (UEFI or legacy, GRUB or systemd-boot)? Does either box have a hardware watchdog? On demo-hp, with the operator's go before each reboot.
Q5 What does an interrupted apt run leave, and does the repair recover it? Kill a run on 9202, after a snapshot.
Q6 How far behind are the demo boxes today, per layer and per source? How long does catching up take, and what restarts? apt-get -s upgrade (read only) first; then a real run on 9202 and on one demo host.
Q7 Which Debian packages restart something big? needrestart in list mode after a run.
Q8 What else runs in the night window, and where does the OS leg fit? Read the timers on host and guest.
Q9 How is cloudflared installed and updated on the host? Read the installer and the agent.
Q10 What does the Proxmox enterprise repository cost per box per year, and what does it add? The publisher's price page. Record only. The operator decides later.

7.1 Answers — the 2026-10-04 spike [FACT]

All evidence: audits/os-updates-spike-2026-10-04/ (its README.md carries every number).

Q Answer in one line Detail
Q1 Usually yes: Debian keeps 2 versions; no box package was re-fixed within 14 days in 3 months; the snapshot archive serves a gone version in 2 s. C2
Q2 Yes for Proxmox (30–66 versions) and Docker (18–46); Debian only the point-release version. C2, §5.6
Q3 Without live-restore: every container restarts, ~30 s of silence. With it: none. Switching it off is a trap. C5
Q4 --next-boot falls back only after a boot that succeeds (measured); a hang keeps the new kernel (code). Software watchdog only. C4
Q5 Not by itself; dpkg --configure -a + apt-get -f install repair it in ~5 s. C7
Q6 Hosts 188 pending each, guests 54–59; guest Debian 24 s, host Debian 60 s, kernel 47 s. Three guests, three Docker versions. audit README
Q7 Few restarts by script; libc leaves PID 1, lxc-start, Proxmox daemons and dockerd on the old library. Proxmox packages restart their own daemons. C11
Q8 The backup legs are windowed; restore-tests and agent updates are not; host apt timers are inert. C10
Q9 Not a host package: a pinned guest container, 4 months behind. C8
Q10 €120 (Community) to €1,100 (Premium) per CPU socket per year, net; every tier includes the Enterprise Repository. Money — the operator's. audit README

Sample approved list built from what ring 0 installed (partI/sample-approved-list.tsv: 108 host + 49 guest packages, with origin) and simulated read-only on demo-felhom: all 157 would install, exact version, downloadable; not covered: 79 Proxmox + 1 Tailscale on the host (slow lane / not ours), 6 Docker + 4 Debian in the guest (C9).

8. Build order [PROPOSAL]

Each step returns to the operator for go or no-go.

  1. Spike (measure Q1–Q10; no product code).
  2. Guest Debian, fast lane. Lowest risk: a snapshot undo exists. BUILT 2026-10-04 — agent v0.140.0, hub v0.130.0, installer 1.29.0; §8.1. There is no snapshot undo (R-837, measured).
  3. Host Debian, fast lane (no kernel, no Proxmox packages).
  4. Fleet view and alarms (§5.7).
  5. Slow lane: Docker engine.
  6. Slow lane: host kernel and Proxmox packages, with the reboot.
  7. Later: the Proxmox major upgrade (PVE 9 → 10), drilled on ring 0 first.

8.1 Step 2 as BUILT (2026-10-04) [FACT]

Evidence: audits/os-guest-lane-2026-10-04/ (parts A–G). Brief: guest fast lane, decisions 78–80 (09 §3).

  • The undo (R-837, measured first): none automatic. PVE refuses ANY snapshot of a customer guest — PVE/AbstractConfig.pm:755-757 skips non-snapshot mounts only for a snapshot named vzdump, and every customer guest carries the host-path binds mp8/mp9. As the agent's token and as root: snapshot feature is not available. The token's role HAS VM.Snapshot / VM.Snapshot.Rollback. So §5.6's first row does not exist: a failed health check stops, reports health_failed, and the hub mails the operator; the whole-guest backup taken minutes earlier is the undo, by hand. The choice of a real undo is in STATUS (R-842).
  • The wrapper felhom-os-apply (agent repo configs/), Python 3 stdlib, per §5.4.1 with refusals R1–R13 (R12: the host layer; R13: dpkg still broken after the repair). Changed from the draft (decided by CC unattended — operator may reverse): one sudoers entry (--plan <file>); the plan's mode field (inventory / apply / health) replaces a separate --repair-only (the repair runs first on every apply); Python, because a JSON plan cannot be parsed safely in sh. The box's own customer guest is "the guest that binds /mnt/felhom-drives" (R10).
  • The leg (agent internal/osupdate): after a SUCCESSFUL primary whole-guest backup, inside the backup's goroutine before the host-wide heavy-op gate is released (so never beside a backup or a restore-test, C10), 90 s after the backup, at most once per 20 h. Decided by CC unattended — operator may reverse: the 90 s settle and the 20 h gap; the controller's own self-update (04:30) is not detected — the 5-minute health wait absorbs a restart.
  • The health rule (HealthVerdict, pinned): docker answers, the guest resolves deb.debian.org, the controller's health check is healthy, and every container running at the START of the leg runs again (healthy if it was). The baseline merges the inventory's reading with the apply's own — found live: an app stopped between them escaped the first rule. Wait 5 min, poll 15 s.
  • Rings and the switch (hub, per box): ring 0 = demo-hp + demo-felhom, everything else ring 1; switch ON by default; OFF → the box reports, installs nothing. A box with no os_update block (older hub) = ring 1, ON, no release.
  • The approval rule (the ruled "1–2 day wait"): every Debian / Debian-Security package=version that ALL ring-0 boxes having it agree on; approved when, since that set was first seen, 24 h passed with every ring-0 report healthy and every ring-0 box completed 1 post-backup night run (OS_APPROVE_AFTER, OS_APPROVE_NIGHTS; an override is logged as a TEST configuration). The approval time is the snapshot.debian.org timestamp (decision 79). "Approve now" is an operator event. The 24 h / 1 night numbers: decided by CC unattended within the ruled 1–2 days.
  • The household's line: hub event os_update_applied (info: on the household's hub timeline, never mailed; hu/en in the bundle). There is no surface on the box itself (R-844).
  • Measured live: ring 0 — 53 Debian packages on each demo box (18.7–31.7 s inside the wrapper; 174–226 s for the whole leg incl. the inventory), healthy, 6 Docker updates "not covered"; approval — with a 2-minute TEST wait, a 272-package release approved automatically, then the ruled values restored; ring 1 — demo-felhom installed exactly the 3 approved versions it lacked (269 already current) and left the newer Docker packages alone; a deliberately stopped app → health_failed after 5 min, the operator mailed, the household line recorded.
  • Not delivered to existing boxes by the product: the wrapper and the sudoers line reach a box only through the installer; the signed agent update replaces the binary only (R-840). The demo boxes got them BY HAND.

9. Where the rest lives

  • The finding: R-812 (backlog/OPEN-ITEMS.md). The intention: R-808 (backlog/ROADMAP.md).
  • Related: R-604 (a per-customer floor hides a box from global raises; the same risk applies to an OS-release floor), R-530 (agents update only by a signed job per box).
  • App updates: 09. Backups and the night chain: 07 §6.1. The agent's permissions: 03 §3.