31bdb4b549
gates / gates (push) Successful in 31s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
529 lines
44 KiB
Markdown
529 lines
44 KiB
Markdown
# 11 — Operating-system updates: the host, the guest and the Docker engine
|
||
|
||
> | | |
|
||
> |---|---|
|
||
> | **Status** | **NOT RATIFIED — a PROPOSAL with operator rulings, corrected by the 2026-10-04 spike (§7.1, C1–C12); §8 step 2 BUILT 2026-10-04 (§8.1).** Ratification is Viktor's review, not an editor's. |
|
||
> | **Written** | 2026-10-04, by the reviewer (project Claude), before any spike. |
|
||
> | **Verified against** | felhom.eu `d07a1a9` · felhom-controller `99a1497` (v0.290.0) · felhom-agent `d766666` (v0.138.0) · hub v0.128.0 |
|
||
> | **Freshness** | **CURRENT** as of 2026-10-04. The spike `TASK-backup-close-and-os-updates-spike-2026-10-04` adds measurements here as `[FACT]` and corrects every claim it disproves. Mark this file STALE when it falls behind what the product does. |
|
||
>
|
||
> **How to read this document.** Each statement has a label:
|
||
>
|
||
> - **[FACT]**: an observed property, with a `file:line`, a register row or an audit path.
|
||
> - **[RULED]**: an operator decision, with its date.
|
||
> - **[PROPOSAL]**: the reviewer's design. It is not decided and not built.
|
||
> - **OPEN**: a question that nobody has answered yet. The spike measures it. Nobody guesses it.
|
||
>
|
||
> **Why this file exists.** No architecture document covered operating-system updates. `00` §G marks the
|
||
> capability MISSING (added 2026-10-03). The finding is **R-812**. The roadmap intention is **R-808**.
|
||
> **The register carries the work. This file carries the reasoning. The source is the truth.**
|
||
>
|
||
> **Spike corrections, 2026-10-04** (`audits/os-updates-spike-2026-10-04/`). Each is made in place: the reviewer's
|
||
> text is struck (~~like this~~) and the measured text follows, labelled `[FACT]`, with its id.
|
||
>
|
||
> | Id | Where | What the spike changed |
|
||
> |---|---|---|
|
||
> | C1 | §2 | The agent may run `apt-get install` for TWO packages, not one (dnsmasq and wireguard-tools). |
|
||
> | C2 | §5.3 | Debian keeps two versions, not one; for box packages no fix was replaced within 14 days in 3 months; `snapshot.debian.org` works from a box in seconds. |
|
||
> | C3 | §5.2 | The lanes must follow the package's ORIGIN, not its name: 40 Proxmox-repository packages have ordinary names (ZFS, the Secure Boot shim, Ceph, corosync, chrony, CPU microcode). |
|
||
> | C4 | §5.6 | `--next-boot` is NOT a one-shot on these GRUB hosts. The fallback works only after a boot that reaches userspace. Installing a kernel alone makes it the default. Only a software watchdog runs. |
|
||
> | C5 | §5.6, §6 row 4 | Docker's `live-restore` keeps every container running across an engine update (measured). Turning it OFF again stops every container and starts none. |
|
||
> | C6 | §5.6 | A guest snapshot works on LVM-thin (customer guests) but not on `dir` storage; the snapshot rollback itself is unmeasured; a backup-restore undo took 73 s. |
|
||
> | C7 | §6 row 2 | A killed `apt` run does not recover by itself; the repair took ~5 s. |
|
||
> | C8 | §1, §6 row 14 | `cloudflared` is not a host package: it is a container in the guest, pinned by the controller since June. It belongs with the controller's infrastructure pins, not this file's lanes. |
|
||
> | C9 | §5.3 | The approved list must record what ring 0 RUNS healthy, not only what it installed that night. |
|
||
> | C10 | §5.5 | Restore-tests and agent updates are not windowed; the host's `apt` timers install nothing today. |
|
||
> | C11 | §5.2 | A Debian (fast-lane) update leaves PID 1, `lxc-start`, the Proxmox daemons and dockerd on the old library: its full effect needs a restart the fast lane does not do. |
|
||
> | C12 | §6 row 9 | The two demo hosts differ by 7 packages, including the CPU microcode (AMD vs Intel) and Secure Boot (on vs off). |
|
||
|
||
---
|
||
|
||
## 0. In plain language
|
||
|
||
A box runs three layers that we install and never update: the Proxmox host, the small Debian system
|
||
inside the guest, and the Docker engine. App images are updated (see `09`). The layer under them is not.
|
||
A box lives in a home for years, so this is a security gap.
|
||
|
||
The proposal has two lanes. **The fast lane** applies Debian's security fixes automatically, but only
|
||
the exact versions that already ran well on the demo boxes for 1–2 days. **The slow lane** covers the
|
||
kernel, Proxmox and Docker. These cause restarts or big changes, so a person approves them one version at
|
||
a time, and they run only at night. The tested versions are recorded automatically from what the demo
|
||
boxes installed. Nobody keeps a hand-written list.
|
||
|
||
---
|
||
|
||
## 1. Scope
|
||
|
||
**In scope.**
|
||
- The Proxmox host on an **appliance** install: the Debian 13 base, the Proxmox packages, and the kernel.
|
||
- The guest (the customer LXC): its Debian 13 packages.
|
||
- The Docker engine inside the guest (`docker-ce`, `docker-ce-cli`, `containerd.io`).
|
||
- How the household and the operator are told, and how a failed update is undone.
|
||
|
||
**Out of scope.**
|
||
- **App images.** `09` covers them (the ladder, the monthly same-tag re-test).
|
||
- **The controller and agent binaries.** Their own self-update covers them (`03`, the self-update section; `09` R-608).
|
||
- **DooPlex and ep0.** The operator updates them by hand (`runbooks/offsite-endpoint.md`).
|
||
- **A BYO host** (the owner brought their own Proxmox). `[FACT]` The installer leaves its repositories
|
||
alone: *"apt repo alignment skipped (byo — the owner manages repos)"* (`scripts/felhom-host-install.sh`,
|
||
`align_apt_repos`). `[PROPOSAL]` On a BYO box we update the guest and Docker only, never the host.
|
||
- **The Proxmox MAJOR upgrade** (PVE 9 → 10). It is a later step of its own, drilled first (R-808 item 5).
|
||
- **`cloudflared` and the other infrastructure images** (traefik, filebrowser). **[FACT] (C8)** They are containers in
|
||
the guest, pinned in the controller (`internal/infra/infra.go:26`: `cloudflare/cloudflared:2026.6.0`, since
|
||
2026-06-11) and baked into the golden; upstream was `2026.9.3` on 2026-10-04. They move only by a controller release
|
||
— the app-image question (`09`), not an OS package. The gap is R-838.
|
||
|
||
---
|
||
|
||
## 2. What a box runs, as measured
|
||
|
||
| Layer | What it is | Source of packages | Evidence |
|
||
|---|---|---|---|
|
||
| Host | Proxmox VE 9.2 on Debian 13, LVM-thin | Debian mirrors + `pve-no-subscription` (enterprise repo switched off on appliance installs) | `[FACT]` `01` §2; `felhom-host-install.sh` `align_apt_repos` (~L2130–2136) |
|
||
| Host kernel | The only kernel on the box. The guest has none. | `pve-no-subscription` (`proxmox-kernel-*`) | `[FACT]` LXC shares the host kernel (`01` §2: "one LXC/kernel/Docker daemon") |
|
||
| Guest | Unprivileged LXC, `nesting=1,keyctl=1`, from `debian-13-standard_13.1-2` | Debian mirrors | `[FACT]` `felhom-agent/configs/build-golden.sh:67,101-103` |
|
||
| Docker engine | `docker-ce`, `docker-ce-cli`, `containerd.io`, installed at golden BAKE time | `download.docker.com/linux/debian trixie stable` | `[FACT]` `build-golden.sh:108-125` |
|
||
| Docker settings | `containerd-snapshotter: false`, json-file log caps. **No `live-restore`.** | baked `daemon.json` | `[FACT]` `build-golden.sh:138-144` |
|
||
| Controller | A container in the guest. A Docker engine restart restarts it too. | Felhom registry | `[FACT]` `03` §1 |
|
||
|
||
**[FACT] Nothing updates any of these layers today** (R-812, searched 2026-10-03). The installer
|
||
says *"No upgrades are run — repo alignment only"* (`felhom-host-install.sh` ~L2135). A fresh install
|
||
gets the Docker engine that was current when its golden was baked, and keeps it.
|
||
|
||
~~**[FACT] The agent may not run `apt` today, with one exception.** Its sudoers allowlist
|
||
(`felhom-agent/configs/felhom-agent.sudoers`) holds one apt line: `apt-get install -y -q dnsmasq`.~~
|
||
**[FACT] (C1) The agent may run `apt-get install` for two named packages:** `felhom-agent.sudoers:58`
|
||
(`apt-get install -y -q dnsmasq`, used by `internal/lanresolver/lanresolver.go:107`) and `:194`
|
||
(`apt-get install -y -q wireguard-tools`, used by `internal/wgtunnel/manager.go:628`). Nothing else.
|
||
`apt` as root runs package scripts as root, so a broad `apt` grant is a full root grant. §5.4
|
||
proposes how to avoid that.
|
||
|
||
---
|
||
|
||
## 3. Operator rulings
|
||
|
||
**2026-10-03.** R-808 is on the roadmap at P2: *"every box receives operating-system security
|
||
patches on a schedule, and a failed update is undone."*
|
||
|
||
**[RULED] 2026-10-04: the fast lane follows an approved list, with a 1–2 day wait.** Every update
|
||
runs on the demo boxes first. The other boxes install only the exact versions that the demo boxes ran
|
||
without trouble. The exact wait is set when the feature is built. **Rejected:** Debian's own
|
||
`unattended-upgrades` with no wait. It is simpler and common, but a bad update would reach every
|
||
customer at the same time.
|
||
|
||
**[RULED] 2026-10-04: the off-site backup topic is closed first.** The same brief carries both
|
||
topics, and the backup part runs first.
|
||
|
||
The two-lane split (§5.2) is the reviewer's proposal. The operator's ruling above assumes it, but he
|
||
has not ruled on it as such.
|
||
|
||
---
|
||
|
||
## 4. The constraints that shape the design
|
||
|
||
1. **Unattended.** Nobody is at the box. Nobody answers a question that `apt` asks.
|
||
2. **No screen.** If the box does not boot, the household sees only that nothing works.
|
||
3. **One kernel for everything.** A host kernel update needs a host reboot. A reboot stops every app.
|
||
4. **The controller lives in Docker.** A Docker engine update restarts the controller in the middle of
|
||
its own work. So the controller cannot drive a Docker update. The agent, on the host, must.
|
||
5. **The host has no whole-system backup.** The guest has three tiers (`07`). The host has none.
|
||
A host update can only be undone by installing the previous version again.
|
||
6. **Every update must already have run on a box we own.** This is the lesson of the update arc
|
||
(`09` §3 decision 13: "the test decides").
|
||
7. **Few packages.** Each extra package is one more thing to update and break. The guest is the
|
||
Debian standard template plus Docker. Keep it that way.
|
||
|
||
---
|
||
|
||
## 5. The proposed shape `[PROPOSAL]`
|
||
|
||
### 5.1 Two rings
|
||
|
||
- **Ring 0:** demo-felhom (N100) and demo-hp. They are disposable (`runbooks/target-selection.md`), and
|
||
they have different hardware. They take every update first.
|
||
- **Ring 1:** every other box. It takes only what ring 0 approved.
|
||
|
||
### 5.2 Two lanes
|
||
|
||
| | Fast lane | Slow lane |
|
||
|---|---|---|
|
||
| What | Debian packages on host and guest, from `trixie-security` and the stable point releases, **except** the slow-lane list | Kernel (`proxmox-kernel-*`), Proxmox (`pve-*`, `proxmox-*`, `lxc-pve`, `qemu-server`, …), Docker (`docker-ce*`, `containerd.io`) |
|
||
| Restart | A service restart at most. No reboot. | Kernel: a host reboot. Docker: every container restarts. Proxmox: its services restart. |
|
||
| Approval | Automatic. A version is approved when ring 0 ran it and stayed healthy for the wait (1–2 days). | A person approves one version at a time, like a controller floor. |
|
||
| Urgent fix | The operator can approve a version on the same day once ring 0 has run it. | The same. |
|
||
| When | The night window (§5.5) | The night window. A kernel reboot only on a night the operator scheduled. |
|
||
|
||
~~Which packages count as slow lane is a list the spike checks (OPEN Q7). A Debian package that restarts
|
||
something big (for example `systemd`, `libc6`, `openssh-server`) may belong in the slow lane too.~~
|
||
|
||
**[FACT] (C3) The lane must be decided by ORIGIN, not by name.** On demo-hp, 40 pending packages come from the
|
||
Proxmox repository under ordinary names: `zfsutils-linux`, `zfs-zed`, `libzfs7linux`…, **`shim-signed` and friends (the
|
||
Secure Boot loader)**, `ceph-common`/`librados2`…, `corosync`, `chrony`, `frr`, `amd64-microcode`
|
||
(`partH/H2-proxmox-origin-debian-names.txt`). A name rule (`pve-*`, `proxmox-*`) would have put them in the fast lane.
|
||
The fast lane is: origin `Debian` or `Debian-Security`, and nothing else. On demo-hp that selection was 108 packages
|
||
and pulled in **zero** Proxmox packages.
|
||
|
||
**[FACT] (C11) What a fast-lane run restarts, and what it does not.** Guest (9202, 49 packages incl. libc6): the
|
||
packages' own scripts restarted postfix, journald, networkd; **dockerd, containerd, sshd, dbus, logind, cron** kept the
|
||
old libc; no container stopped (13 samples). Host (demo-hp, 108 packages): dnsmasq, postfix, journald restarted;
|
||
**systemd (PID 1), `lxc-start`, pveproxy, pvedaemon, pvestatd, pvescheduler, watchdog-mux, sshd, zed, chronyd** kept the
|
||
old libc; both guests and the agent stayed up. So the fast lane is safe to run unattended, but a libc fix is only
|
||
fully in force after a reboot (host) or a Docker restart (guest) — which are slow-lane acts. `[PROPOSAL]` the box
|
||
reports "restart needed" (processes on deleted libraries) and the slow lane's next reboot picks it up.
|
||
|
||
### 5.3 The approved list (the "tested versions" record)
|
||
|
||
Nobody writes the list by hand. It fills itself:
|
||
|
||
1. A ring-0 box updates. Afterwards it reports to the hub the exact `package=version` it installed,
|
||
per layer (host, guest), with each package's origin (`Debian-Security`, `Debian`, `Proxmox`,
|
||
`Docker`).
|
||
2. The box then reports health for the wait period: the agent, the controller, every app's health,
|
||
and the guest's network.
|
||
3. When the wait passes with ring 0 healthy, the hub marks that set **approved**. It is one record:
|
||
an **OS release**, with an id and a date. The hub stores it. The register does not.
|
||
4. A ring-1 box asks the hub for the newest approved OS release. For each package it has installed,
|
||
if the approved version is newer, it installs that **exact** version. It never installs a version
|
||
newer than the approved one.
|
||
5. Each box reports which OS release it runs, how many updates are waiting, whether it needs a reboot,
|
||
and which packages it has that **no approved list covers** (see §6, edge case 9).
|
||
|
||
~~**OPEN Q1 is the weak point.** The design works only if a box can still download the approved version
|
||
a week later. Debian's main and security archives keep only the newest version of each package. If
|
||
Debian publishes a newer fix between approval and install, the approved version is gone. Options:
|
||
the box waits for the next approval; or the box uses `snapshot.debian.org` (Debian's own dated archive,
|
||
still signed by Debian); or Felhom runs a package cache. The spike measures how often this happens.~~
|
||
|
||
**[FACT] (C2) Q1 measured.** The live Debian archives keep **two** versions: the point-release one in `trixie`
|
||
(main) and the newest in `trixie-security`; intermediate versions are gone (openssl: installed `u1`, main `u2`,
|
||
security `u3`). Over 2026-07-04..10-04, 167 trixie security advisories; for the 517 source packages installed on a box,
|
||
**no package got a second advisory within 2, 7 or 14 days** (the 3 within 2 days were chromium and webkit2gtk, not on a
|
||
box). `snapshot.debian.org` answers a box: a dated index in **2.3–3.0 s**, a gone exact version
|
||
(`openssl 3.5.6-1~deb13u1`) downloaded in **2.0 s**, Debian-signed. Proxmox and Docker keep many old versions
|
||
(pve-manager 66, docker-ce 46). So with a 1–2 day wait the approved version is almost always still live; the rare
|
||
miss is fetched from the snapshot taken at approval time. `[PROPOSAL]` each OS release records its approval
|
||
timestamp; a box installs from its own sources, and only for a Debian package that is no longer there, from
|
||
`snapshot.debian.org/archive/<debian|debian-security>/<timestamp>`. This is the operator decision in STATUS.
|
||
|
||
**[FACT] (C9) Approve what ring 0 RUNS, not what it installed.** In the simulation on demo-felhom
|
||
(`partI/demo-felhom-simulation.txt`), every one of the 108 host and 49 guest approved versions was installable and
|
||
downloadable — but `curl`, `libcurl*` and `libssh2` were "not covered" in the guest, only because scratch guest 9202
|
||
ALREADY ran the newer version and so installed nothing. `[PROPOSAL]` step 1 reports the full installed
|
||
`package=version` set after the run, and approval covers every version ring 0 runs healthy.
|
||
|
||
### 5.4 Who runs it, and with what permission
|
||
|
||
- **The agent runs every OS update**, for the host and for the guest (`pct exec`). The controller does
|
||
not, because of constraint 4.
|
||
- **The agent gets no general `apt` permission.** Like `felhom-selfupdate-guarded` (`03`, the self-update section),
|
||
a small root-owned wrapper does the work. It accepts only an approved list and refuses everything else:
|
||
- no package removal;
|
||
- no downgrade, except the undo of §5.6, which an operator job signs;
|
||
- no package that the box does not already have, unless the approved list records it as a dependency
|
||
that the same update pulled in on ring 0 (a kernel update installs a NEW package name each time,
|
||
for example `proxmox-kernel-6.x.y-z-pve-signed`, pulled by `proxmox-default-kernel`);
|
||
- no package source other than the ones the installer set up;
|
||
- non-interactive, and it always keeps the existing config file (`--force-confold`) and reports the
|
||
conflict.
|
||
- **Package signatures stay the publishers'** (Debian, Proxmox, Docker). The hub sends only names and
|
||
versions. A broken-into hub can choose an older version or no version. It cannot make a box install a
|
||
package that the publisher did not sign.
|
||
|
||
### 5.4.1 The root wrapper's interface — DRAFT (2026-10-04, design only, nothing installed) `[PROPOSAL]`
|
||
|
||
`felhom-os-apply` — root-owned (`0755 root:root`), installed by the installer beside `felhom-selfupdate-guarded`;
|
||
the agent's sudoers gets exactly `/usr/local/sbin/felhom-os-apply --plan /var/lib/felhom-agent/os/plan-*.json` and
|
||
`… --repair-only`. For the guest it runs on the host and enters the guest with `pct exec <vmid> --` itself, so the
|
||
agent needs no `pct exec … apt` line. Built from what Parts G and H measured.
|
||
|
||
**Input — one JSON plan file** (written by the agent, from the hub's approved OS release):
|
||
|
||
```json
|
||
{
|
||
"release_id": "os-2026-10-04-1", "approved_at": "2026-10-04T08:00:00Z",
|
||
"layer": "host", // "host" | "guest"
|
||
"vmid": 9201, // guest only
|
||
"snapshot": "20261004T080000Z", // the snapshot.debian.org timestamp of approval (C2)
|
||
"packages": [ {"name": "libc6", "version": "2.41-12+deb13u4", "origin": "Debian"} ],
|
||
"allow_new": ["proxmox-kernel-7.0.14-20-pve-signed"], // slow lane only, signed operator job
|
||
"lane": "fast" // "fast" | "slow"
|
||
}
|
||
```
|
||
|
||
**Order of work:** (1) refuse checks below; (2) **repair first** — `dpkg --configure -a` then `apt-get -f install`,
|
||
logging what it repaired (C7); (3) `apt-get -s install` of exactly `name=version` for every package that is
|
||
installed AND older; (4) refuse if the simulation would remove, downgrade, add an unlisted package, or touch a package
|
||
whose candidate origin is not the plan's; (5) download — from the box's own sources, or for a Debian version no
|
||
longer there, from `snapshot.debian.org/archive/<archive>/<snapshot>` with a temporary sources list it deletes after
|
||
(C2); (6) install with `DEBIAN_FRONTEND=noninteractive APT_LISTCHANGES_FRONTEND=none -o Dpkg::Options::=--force-confold
|
||
-o Dpkg::Options::=--force-confdef`; (7) `apt-get clean`; (8) report.
|
||
|
||
**Refusals** (each exits non-zero with one line `os-apply: REFUSED: <reason>` and changes nothing):
|
||
|
||
| # | Refuses when |
|
||
|---|---|
|
||
| R1 | the plan file is not under `/var/lib/felhom-agent/os/`, not owned by the agent, or not valid JSON |
|
||
| R2 | `lane` is `fast` and any package's origin is not `Debian` / `Debian-Security` (C3) |
|
||
| R3 | `lane` is `slow` and the plan is not carried by a verified signed operator job (R-530's mechanism) |
|
||
| R4 | the simulation removes any package |
|
||
| R5 | the simulation downgrades any package (the operator undo is a separate signed op, §5.6) |
|
||
| R6 | the simulation installs a package that is neither installed nor in `allow_new` |
|
||
| R7 | a listed version is not downloadable from the sources the installer set up or the named snapshot |
|
||
| R8 | free space on `/` (or the guest's rootfs) is below 3× the download size, minimum 500 MB (edge case 8) |
|
||
| R9 | another apt/dpkg holds the lock, or the per-guest lane lock is held (a backup, a restore-test, C10) |
|
||
| R10 | `layer` is `guest` and the vmid is not the box's own customer guest |
|
||
| R11 | the plan names a package twice, or a version that is not a Debian version string |
|
||
|
||
**Log lines** (to the journal, tag `felhom-os-apply`, and echoed for the agent to forward to the hub):
|
||
|
||
```
|
||
os-apply: START release=<id> layer=<host|guest:vmid> lane=<fast|slow> packages=<n>
|
||
os-apply: REPAIR configured=<n> fixed=<n> (always printed; 0 0 when nothing was half-done)
|
||
os-apply: PLAN upgrade=<n> already=<n> not-installed=<n> from-snapshot=<n> download=<bytes>
|
||
os-apply: REFUSED: <R-number> <reason>
|
||
os-apply: CONFFILE kept <path> (new version saved as <path>.dpkg-dist)
|
||
os-apply: DONE rc=0 seconds=<s> upgraded=<n> restarted=<unit,…> restart-needed=<process,…> reboot-needed=<yes|no>
|
||
os-apply: FAILED rc=<n> step=<download|install> — dpkg state: <dpkg --audit first line>
|
||
```
|
||
|
||
`restart-needed` lists processes still mapping deleted libraries (C11); `reboot-needed` is yes when that list holds
|
||
PID 1 or `lxc-start`, or a kernel was installed.
|
||
|
||
### 5.5 When
|
||
|
||
Inside the household's night window, after the backups:
|
||
|
||
```
|
||
W DB dump
|
||
W+60m Tier 2
|
||
W+105m off-site → app updates (until W+5h at most)
|
||
[W+2h, W+6h) whole-guest backup (agent)
|
||
after it OS updates — guest first, then host
|
||
```
|
||
|
||
The OS leg starts **after the whole-guest backup has finished**, so the guest's newest full copy is
|
||
minutes old. This is the opposite order to app updates, which run before the whole-guest backup
|
||
(`07` §6.1, `09` decision 11). The reason: the whole-guest backup is the guest's undo.
|
||
|
||
**[FACT] (C10) Q8 measured** (guest UTC; demo-hp W = 02:30): db-dump 02:30, tier-2 03:30, off-site ~04:15,
|
||
whole-guest gate [04:30, 08:30), controller self-update 04:30, offsite-integrity 06:00. Host: `apt-daily` and
|
||
`apt-daily-upgrade` run daily but install nothing (no `unattended-upgrades`, no `APT::Periodic`); `pve-daily-update`
|
||
refreshes the lists daily. **Restore-tests are NOT windowed** — they run on a cadence at any hour (demo-felhom 10:38
|
||
daily, demo-hp 16:43 and 22:46); agent updates arrive by signed job at any hour. `[PROPOSAL]` the OS leg takes the same
|
||
per-guest lane lock the restore-test and the whole-guest backup take, rather than a clock slot.
|
||
|
||
~~**OPEN Q8:** what else runs then. The controller's self-update (default 04:30, and after any hub report
|
||
when a floor is above it, R-608), the agent's self-update, the restore-tests, and PBS jobs. The OS leg
|
||
must never overlap a backup, a restore-test or a self-update.~~
|
||
|
||
### 5.6 How a failed update is undone
|
||
|
||
"Rollback" is not used (`09` §4). The shapes:
|
||
|
||
| Layer | Undo | Limit |
|
||
|---|---|---|
|
||
| Guest packages | Restore the guest snapshot taken just before the update | It also undoes app data written after the snapshot. Use it only inside the health window, before apps have written much. After that, install the previous version (needs OPEN Q1). **[FACT] (C6)** Customer guests are on LVM-thin and can snapshot; scratch 9202 (`dir` storage) cannot (`snapshot feature is not available`), so the measured undo was a whole-guest backup + restore: **73 s down**, all apps healthy, libc back. The snapshot rollback itself is unmeasured (R-837). |
|
||
| Docker engine | Install the previous version (Docker's repository keeps old versions — OPEN Q2) | Every container restarts again. **[FACT] (C5)** Docker keeps 46 `docker-ce` versions. A step (either way) without `live-restore`: 6 of 6 containers restart, the app silent **26.5–30 s**, healthy at +42–45 s. With `live-restore` on: **0 restarts, no gap**, also across a containerd step. But a restart that turns `live-restore` OFF stops every container and starts NONE (`unless-stopped` ignored) — on 9202 they stayed down until restarted by hand; `systemctl reload` turns it on but not off. |
|
||
| Host packages | Install the previous version | Only if the source still has it (OPEN Q1/Q2). **[FACT]** `rsync` back to its pre-update version: `not found`; `libpng16-16t64` back to the point-release version: worked. A Debian undo needs the snapshot archive (C2). **[FACT, 2026-10-04]** Proved by hand: `runbooks/os-updates-host-undo.md` (demo-hp, `tzdata` back one version from `snapshot.debian.org`, held, released). Use the NEW version's `first_seen` as the timestamp when there is no previous host release; a package that pins its siblings (`eject` → `libmount1 =`) goes back only with them. |
|
||
| Host kernel | ~~Boot the previous kernel. Proxmox can boot a new kernel **once** (`proxmox-boot-tool kernel pin <ver> --next-boot`). If that boot fails, the next boot uses the old kernel again. Make the new kernel permanent only after a healthy boot.~~ **[FACT] (C4)** Both demo hosts boot UEFI + GRUB (no proxmox-boot-tool ESPs). **Installing a kernel makes it the GRUB default at once.** `--next-boot` on GRUB writes an ordinary `GRUB_DEFAULT` + `update-grub`; `proxmox-boot-cleanup.service` clears it only once a boot reaches userspace. Measured on demo-hp (old kernel pinned permanently FIRST, new pinned for next boot): boot 1 → `7.0.14-20-pve`, healthy, 60 s; boot 2 → back on `7.0.2-6-pve`, healthy, 76 s. **So the fallback works after a boot that succeeds; read from the code, a kernel that hangs before userspace stays the default on every power cycle.** ~~`[PROPOSAL]` use GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel as the saved default — unmeasured (R-836).~~ **[FACT, 2026-10-04, demo-hp, operator's word before each reboot] GRUB's one-shot is NOT a one-shot here either.** With `GRUB_DEFAULT=saved` (old 7.0.2-6 saved) and `grub-reboot` 7.0.14-20: boot 1 → 7.0.14-20, **Secure Boot ON, booted fine** (signed kernel, shim → GRUB); but `/boot` is ext4 on LVM, GRUB cannot write its environment block there (`grub-reboot` warns so itself), `next_entry` was never cleared, and boot 2 with no command → **7.0.14-20 again**. A kernel lane needs a writable env block (the ESP) or a userspace "boot good" step — R-836. demo-hp left on 7.0.14-20, saved default 7.0.14-20, both kernels installed (`audits/os-host-lane-2026-10-04/partE/`). | ~~If the new kernel hangs, someone must switch the box off and on. The spike checks whether a hardware watchdog can do that (OPEN Q4).~~ **[FACT]** Only `softdog` runs (loaded by `watchdog-mux`); it cannot rescue a kernel that never boots. demo-hp has an AMD FCH whose `sp5100_tco` driver ships but is not loaded; untested. A hang still needs a person. **[FACT, 2026-10-04]** `sp5100_tco` is blacklisted by the Proxmox kernel package; loaded by hand it answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0 — read from sysfs, never opened, so never armed; unloaded). `kernel.panic = 0`: a panic leaves the host stopped (R-851). |
|
||
|
||
### 5.7 Telling people
|
||
|
||
- **Operator:** a hub event for each update and each failure; a fleet view showing each box's OS release,
|
||
how far behind it is, and whether it needs a reboot. A box that is more than N days behind the newest
|
||
approved release raises an alarm (`08`).
|
||
- **Household:** one line on the timeline in both languages, informal voice: what was updated and
|
||
whether the box restarted. Telling households in advance that the box may restart at night is a
|
||
**promise to users**. That is the operator's decision when the slow lane is built.
|
||
|
||
### 5.8 The Docker engine slow lane — DESIGN (2026-10-04, nothing built) `[PROPOSAL]`
|
||
|
||
Built from C5 (measured) and R-835. **The one decision it needs is in STATUS: `live-restore` on, fleet-wide.**
|
||
|
||
- **What moves:** `docker-ce`, `docker-ce-cli`, `containerd.io`, `docker-buildx-plugin`, `docker-compose-plugin`,
|
||
`docker-ce-rootless-extras` in the customer guest — today the six "not covered" packages on every box. One
|
||
approved **engine set** at a time, like a controller floor: the operator approves it (slow lane, §5.2), after ring 0
|
||
has run it for at least 2 nights healthy. Never two steps in one night.
|
||
- **Precondition: `live-restore` ON.** Without it an engine step restarts every container — 26.5–30 s of silence,
|
||
healthy at +42–45 s (C5). With it: 0 restarts, no gap, also across a containerd step. Turning it ON is safe
|
||
(`systemctl reload docker` applies it without a restart, C5); turning it OFF later by a plain restart stops every
|
||
container and starts none (R-835) — so it is turned on once, by the golden and by a one-time fleet step, and never
|
||
turned off by the lane.
|
||
- **Who and when:** the agent, through the same wrapper (`lane: slow`, refusal R3: only inside a verified signed
|
||
operator job, R-530's mechanism), in the guest, after the guest and host fast-lane steps, under the same heavy-op
|
||
gate, on a night the operator scheduled. Debian origin rule replaced by "origin `Docker CE`, exactly these names".
|
||
- **Health:** the guest rule (§8.1) plus `docker version` reports the approved engine, and every container running
|
||
at the start is running with the SAME container id (proof that `live-restore` held). A changed id is
|
||
`health_failed` even if the app is healthy — it means the households' apps restarted when they should not have.
|
||
- **Undo:** install the previous engine set (Docker's repository keeps 46 versions, C2) — by an operator job, with
|
||
`live-restore` still on, so the undo is also restart-free.
|
||
- **Not covered here:** the golden's own engine (baked weekly; a new golden carries the approved set), and BYO hosts
|
||
(the guest is ours on both, so the lane applies there too).
|
||
|
||
---
|
||
|
||
## 6. Risks and edge cases
|
||
|
||
| # | What can go wrong | What the design does |
|
||
|---|---|---|
|
||
| 1 | A new kernel does not boot | Boot it once (`--next-boot`); the old kernel stays the default. A hard hang needs a power cycle (OPEN Q4). |
|
||
| 2 | Power cut during an update: `dpkg` is half done | The wrapper runs `dpkg --configure -a` and `apt-get -f install` first, every time, and reports what it repaired (spike measures, Q5). **[FACT] (C7)** Killed after 15 unpacks: 5 packages `iU`, 4 triggers pending; the next ordinary `apt-get install` REFUSES (`Unmet dependencies`) — nothing repairs it by itself. The two commands repaired it in 1.4 s + 3.9 s; apps stayed up. |
|
||
| 3 | An update asks a question (changed config file, service restart prompt) | Non-interactive, keep the old config, report the conflict. |
|
||
| 4 | A Docker update stops every app, and the controller | Stop the apps cleanly first, like before a backup. Measure whether `live-restore` keeps containers running (Q3). The agent drives it. **[FACT] (C5)** `live-restore` keeps them running (0 restarts). The controller keeps running too; its `docker` calls fail for the seconds dockerd is down (logged errors, no app event). Turning `live-restore` on is a golden + fleet change — and turning it off later must not be a plain restart (R-835). |
|
||
| 5 | The approved version is no longer downloadable | OPEN Q1. Until it is answered, the box waits for the next approval and reports it. |
|
||
| 6 | An urgent security hole | The operator approves the same day once ring 0 has run it. |
|
||
| 7 | A box was off for months | It catches up through the newest approved release. It does not install every release in between. Packages are not like app data; one `apt` step is enough. Kernel and Docker still go one approved step at a time. |
|
||
| 8 | The system disk is full | Check free space before downloading; clean the package cache after; refuse and report. |
|
||
| 9 | A customer box has a package that ring 0 does not have (other hardware: firmware, CPU microcode, NIC drivers) | It is never approved, so it is never updated. The box reports it as "not covered", and the hub raises it. Fix: add matching hardware to ring 0, or approve it by hand. **[FACT] (C12)** The two demo hosts differ by 7 packages: `amd64-microcode`, `proxmox-secure-boot-support`, `felhom-bootstrap` (demo-hp) vs `intel-microcode`, `proxmox-first-boot`, `tailscale`, `tailscale-archive-keyring` (demo-felhom); Secure Boot is ON on demo-hp, OFF on demo-felhom. Ring 0 covers both CPU vendors today. |
|
||
| 10 | The package source is down, or its signing key changes (Docker has done this) | The update fails cleanly and the box reports it. A key change is a slow-lane act for a person. |
|
||
| 11 | A BYO host | Only the guest and Docker are updated (§1). |
|
||
| 12 | Two boxes on ring 0 is a small sample | Accepted for now. When there are customers, the first tester boxes can become a second ring. |
|
||
| 13 | A broken-into hub sends a harmful list | The wrapper refuses removals, downgrades and new packages, and the publishers' signatures still apply (§5.4). |
|
||
| 14 | `cloudflared` on the host | ~~Not covered by this file yet. How it is installed and updated is OPEN Q9. It faces the internet, so it matters.~~ **[FACT] (C8)** Not on the host: a pinned container in the guest (§1). Four months behind upstream on 2026-10-04 (R-838). |
|
||
| 15 | The no-subscription Proxmox repository gets less testing than the enterprise one | Ring 0 is our test. The enterprise repository costs a yearly fee per box. That is a **money** decision for the operator, later (Q10). |
|
||
|
||
---
|
||
|
||
## 7. Open questions the spike must answer
|
||
|
||
| Q | Question | How to answer |
|
||
|---|---|---|
|
||
| Q1 | Can a box install an exact Debian version one week after approval? How often is it already superseded? | `apt-cache madison` on host and guest for security packages. Debian's security announcement history. Whether `snapshot.debian.org` is reachable and fast enough. |
|
||
| Q2 | Do the Proxmox and Docker sources keep older versions? | `apt-cache madison` for `pve-manager`, `proxmox-kernel-*`, `docker-ce`, `containerd.io`. |
|
||
| Q3 | What does a Docker engine update do to running containers, with and without `live-restore`? | On scratch guest 9202: step from one pinned version to the next; time the downtime. |
|
||
| Q4 | Does `proxmox-boot-tool kernel pin --next-boot` work on both demo hosts' boot setups (UEFI or legacy, GRUB or systemd-boot)? Does either box have a hardware watchdog? | On demo-hp, with the operator's go before each reboot. |
|
||
| Q5 | What does an interrupted `apt` run leave, and does the repair recover it? | Kill a run on 9202, after a snapshot. |
|
||
| Q6 | How far behind are the demo boxes today, per layer and per source? How long does catching up take, and what restarts? | `apt-get -s upgrade` (read only) first; then a real run on 9202 and on one demo host. |
|
||
| Q7 | Which Debian packages restart something big? | `needrestart` in list mode after a run. |
|
||
| Q8 | What else runs in the night window, and where does the OS leg fit? | Read the timers on host and guest. |
|
||
| Q9 | How is `cloudflared` installed and updated on the host? | Read the installer and the agent. |
|
||
| Q10 | What does the Proxmox enterprise repository cost per box per year, and what does it add? | The publisher's price page. Record only. The operator decides later. |
|
||
|
||
---
|
||
|
||
### 7.1 Answers — the 2026-10-04 spike `[FACT]`
|
||
|
||
All evidence: `audits/os-updates-spike-2026-10-04/` (its `README.md` carries every number).
|
||
|
||
| Q | Answer in one line | Detail |
|
||
|---|---|---|
|
||
| Q1 | Usually yes: Debian keeps 2 versions; no box package was re-fixed within 14 days in 3 months; the snapshot archive serves a gone version in 2 s. | C2 |
|
||
| Q2 | Yes for Proxmox (30–66 versions) and Docker (18–46); Debian only the point-release version. | C2, §5.6 |
|
||
| Q3 | Without `live-restore`: every container restarts, ~30 s of silence. With it: none. Switching it off is a trap. | C5 |
|
||
| Q4 | `--next-boot` falls back only after a boot that succeeds (measured); a hang keeps the new kernel (code). Software watchdog only. | C4 |
|
||
| Q5 | Not by itself; `dpkg --configure -a` + `apt-get -f install` repair it in ~5 s. | C7 |
|
||
| Q6 | Hosts 188 pending each, guests 54–59; guest Debian 24 s, host Debian 60 s, kernel 47 s. Three guests, three Docker versions. | audit README |
|
||
| Q7 | Few restarts by script; libc leaves PID 1, `lxc-start`, Proxmox daemons and dockerd on the old library. Proxmox packages restart their own daemons. | C11 |
|
||
| Q8 | The backup legs are windowed; restore-tests and agent updates are not; host apt timers are inert. | C10 |
|
||
| Q9 | Not a host package: a pinned guest container, 4 months behind. | C8 |
|
||
| Q10 | €120 (Community) to €1,100 (Premium) per CPU socket per year, net; every tier includes the Enterprise Repository. Money — the operator's. | audit README |
|
||
|
||
**Sample approved list** built from what ring 0 installed (`partI/sample-approved-list.tsv`: 108 host + 49 guest
|
||
packages, with origin) and simulated read-only on demo-felhom: **all 157 would install, exact version, downloadable**;
|
||
not covered: 79 Proxmox + 1 Tailscale on the host (slow lane / not ours), 6 Docker + 4 Debian in the guest (C9).
|
||
|
||
## 8. Build order `[PROPOSAL]`
|
||
|
||
Each step returns to the operator for go or no-go.
|
||
|
||
1. **Spike** (measure Q1–Q10; no product code).
|
||
2. **Guest Debian, fast lane.** ~~Lowest risk: a snapshot undo exists.~~ **BUILT 2026-10-04** — agent v0.140.0, hub
|
||
v0.130.0, installer 1.29.0; §8.1. **There is no snapshot undo** (R-837, measured).
|
||
3. **Host Debian, fast lane** (no kernel, no Proxmox packages). **BUILT 2026-10-04** — agent v0.141.1, hub v0.131.1;
|
||
§8.2.
|
||
4. **Fleet view and alarms** (§5.7). **BUILT 2026-10-04** — hub v0.131.0/v0.131.1; §8.3.
|
||
5. **Slow lane: Docker engine.**
|
||
6. **Slow lane: host kernel and Proxmox packages, with the reboot.**
|
||
7. **Later:** the Proxmox major upgrade (PVE 9 → 10), drilled on ring 0 first.
|
||
|
||
---
|
||
|
||
### 8.1 Step 2 as BUILT (2026-10-04) `[FACT]`
|
||
|
||
Evidence: `audits/os-guest-lane-2026-10-04/` (parts A–G). Brief: guest fast lane, decisions 78–80 (`09` §3).
|
||
|
||
- **The undo (R-837, measured first): none automatic.** PVE refuses ANY snapshot of a customer guest —
|
||
`PVE/AbstractConfig.pm:755-757` skips non-snapshot mounts only for a snapshot named `vzdump`, and every customer
|
||
guest carries the host-path binds mp8/mp9. As the agent's token and as root: `snapshot feature is not available`.
|
||
The token's role HAS `VM.Snapshot` / `VM.Snapshot.Rollback`. So §5.6's first row does not exist: a failed health
|
||
check stops, reports `health_failed`, and the hub mails the operator; the whole-guest backup taken minutes earlier
|
||
is the undo, by hand. The choice of a real undo is in STATUS (R-842).
|
||
- **The wrapper** `felhom-os-apply` (agent repo `configs/`), Python 3 stdlib, per §5.4.1 with refusals R1–R13 (R12:
|
||
the host layer; R13: dpkg still broken after the repair). **Changed from the draft** *(decided by CC unattended —
|
||
operator may reverse)*: one sudoers entry (`--plan <file>`); the plan's `mode` field (`inventory` / `apply` /
|
||
`health`) replaces a separate `--repair-only` (the repair runs first on every apply); Python, because a JSON plan
|
||
cannot be parsed safely in sh. The box's own customer guest is "the guest that binds `/mnt/felhom-drives`" (R10).
|
||
- **The leg** (agent `internal/osupdate`): after a SUCCESSFUL primary whole-guest backup, inside the backup's
|
||
goroutine before the host-wide heavy-op gate is released (so never beside a backup or a restore-test, C10), 90 s
|
||
after the backup, at most once per 20 h. *Decided by CC unattended — operator may reverse:* the 90 s settle and the
|
||
20 h gap; the controller's own self-update (04:30) is not detected — the 5-minute health wait absorbs a restart.
|
||
- **The health rule** (`HealthVerdict`, pinned): docker answers, the guest resolves `deb.debian.org`, the controller's
|
||
health check is `healthy`, and every container running at the START of the leg runs again (healthy if it was). The
|
||
baseline merges the inventory's reading with the apply's own — found live: an app stopped between them escaped the
|
||
first rule. Wait 5 min, poll 15 s.
|
||
- **Rings and the switch** (hub, per box): ring 0 = demo-hp + demo-felhom, everything else ring 1; switch ON by
|
||
default; OFF → the box reports, installs nothing. A box with no `os_update` block (older hub) = ring 1, ON, no
|
||
release.
|
||
- **The approval rule (the ruled "1–2 day wait")**: every Debian / Debian-Security package=version that ALL ring-0
|
||
boxes having it agree on; approved when, since that set was first seen, **24 h** passed with every ring-0 report
|
||
healthy and every ring-0 box completed **1** post-backup night run (`OS_APPROVE_AFTER`, `OS_APPROVE_NIGHTS`; an
|
||
override is logged as a TEST configuration). The approval time is the snapshot.debian.org timestamp (decision 79).
|
||
"Approve now" is an operator event. *The 24 h / 1 night numbers: decided by CC unattended within the ruled 1–2 days.*
|
||
- **The household's line**: hub event `os_update_applied` (info: on the household's hub timeline, never mailed;
|
||
hu/en in the bundle). There is no surface on the box itself (R-844).
|
||
- **Measured live**: ring 0 — 53 Debian packages on each demo box (18.7–31.7 s inside the wrapper; 174–226 s for the
|
||
whole leg incl. the inventory), healthy, 6 Docker updates "not covered"; approval — with a 2-minute TEST wait, a
|
||
272-package release approved automatically, then the ruled values restored; ring 1 — demo-felhom installed exactly
|
||
the 3 approved versions it lacked (269 already current) and left the newer Docker packages alone; a deliberately
|
||
stopped app → `health_failed` after 5 min, the operator mailed, the household line recorded.
|
||
- **Not delivered to existing boxes by the product**: the wrapper and the sudoers line reach a box only through the
|
||
installer; the signed agent update replaces the binary only (R-840). The demo boxes got them BY HAND.
|
||
|
||
### 8.2 Step 3 as BUILT (2026-10-04) `[FACT]`
|
||
|
||
Evidence: `audits/os-host-lane-2026-10-04/` (parts A–G).
|
||
|
||
- **Where it runs:** appliances only. The proof is the ROOT-owned install record `/var/lib/felhom-install/state.json`
|
||
`mode: appliance` (written by the installer as root); the agent-writable `agent.json` `deployment_mode` is not
|
||
trusted for this. A BYO host gets no host step (wrapper refusal **R12**, now lifted only for lane fast / layer host
|
||
on an appliance). *Decided by CC unattended — operator may reverse.*
|
||
- **What:** origin `Debian` / `Debian-Security` only, and never a kernel, boot or firmware package (name pattern
|
||
`HOST_SLOW_RE` — `linux-*`, `proxmox-kernel*`, `pve-kernel*`, `pve-firmware`, `firmware-*`, `grub*`, `shim*`,
|
||
`systemd-boot*`, `*-microcode`, `efibootmgr`; refusal **R14**). The hub leaves the same names out of the host
|
||
candidate.
|
||
- **When:** in the same leg, after the guest step, under the same heavy-op gate. A failed or unhealthy guest step
|
||
skips the host step.
|
||
- **The host health rule** (`HostHealthVerdict`, pinned in `internal/osupdate`): `felhom-agent`, `pveproxy`,
|
||
`pvedaemon`, `pvestatd` and `pve-cluster` are `active`; the customer guest runs; the guest health rule (§8.1)
|
||
passes; the tunnel is `running` — an `unknown` tunnel does not fail it, `not_running` does. Same 5-minute wait.
|
||
- **Separate approved sets:** host and guest releases are separate (`os-host-…`, `os-guest-…`), each by the same rule
|
||
(24 h, 1 night of THAT layer, every ring-0 box). Ring 1 receives `host_release` beside `release`.
|
||
- **Reboot needed:** reported when PID 1 or `lxc-start` maps a replaced file, with the date of the first scanned
|
||
report that said so. **Never reboots.** The host is scanned on every pass, so a reboot clears it (v0.141.1 — v0.141.0
|
||
hid `lxc-start` and never cleared, R-846).
|
||
- **Undo:** by hand, `runbooks/os-updates-host-undo.md` (proved). No automatic undo.
|
||
- **Speed (R-845):** one call per layer; both steps with nothing to install 23–32 s; a 108-package host pass 70 s.
|
||
- **Measured live:** demo-felhom ring 0 installed 108 Debian host packages, all Debian origin (checked against apt),
|
||
healthy; a 605-package host release approved (TEST wait 2 min / 0 nights, logged, reverted to 24 h + 1 night);
|
||
demo-felhom as ring 1 installed exactly the one version it lacked (604 already current).
|
||
|
||
### 8.3 Step 4 as BUILT (2026-10-04) `[FACT]`
|
||
|
||
- **The fleet view** (`GET /os/fleet`, operator): one line per box — ring, switch, the tunnel, and per layer: the
|
||
release, the last outcome, the last successful leg, pending, not covered, restart needed, reboot needed since, the
|
||
wrapper's own seconds.
|
||
- **The four alarms** (`08` §6.3), operator-only, hourly, at most weekly while true: no successful OS leg for
|
||
**7 days** while the switch is ON (naming the likely reason); reboot needed for **14 days**; ring 0 approved nothing
|
||
for **7 days** while it has pending fast-lane updates; not-covered fast-lane packages for **14 days**. The four
|
||
numbers are configuration (`OS_ALARM_*`). *Decided by CC unattended — operator may reverse:* 7 days = a week of
|
||
missed nights is past any normal hiccup (a box off for a weekend does not alarm); 14 days for reboot and coverage =
|
||
two weekly golden cycles, both need a person anyway.
|
||
- **The tunnel** (R-841): `running` / `not_running` / `unknown`; `tunnel_down` after two `not_running` reports.
|
||
|
||
## 9. Where the rest lives
|
||
|
||
- The finding: **R-812** (`backlog/OPEN-ITEMS.md`). The intention: **R-808** (`backlog/ROADMAP.md`).
|
||
- Related: **R-604** (a per-customer floor hides a box from global raises; the same risk applies to an
|
||
OS-release floor), **R-530** (agents update only by a signed job per box).
|
||
- App updates: `09`. Backups and the night chain: `07` §6.1. The agent's permissions: `03` §3.
|