Files

670 lines
57 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 11 — Operating-system updates: the host, the guest and the Docker engine
> **How to read this document.** Where a statement is marked, it is marked like this — the same wording as
> `07-backup-architecture.md:11-17`, carried here on 2026-10-05 (R-376, the three documents written after the
> 2026-08-22 pass):
>
> - **[DESIGN]** — a decision taken. Not derived from code; the code may not implement it yet.
> - **[FACT]** — an observed property, carrying a `file:line`, a live command output or a citation.
>
> **An unmarked statement means "not yet classified", never "observed"** (R-376).
> | | |
> |---|---|
> | **Status** | **NOT RATIFIED — a PROPOSAL with operator rulings, corrected by the 2026-10-04 spike (§7.1, C1–C12); §8 step 2 BUILT 2026-10-04 (§8.1).** Ratification is Viktor's review, not an editor's. |
> | **Written** | 2026-10-04, by the reviewer (project Claude), before any spike. |
> | **Verified against** | felhom.eu `d07a1a9` · felhom-controller `99a1497` (v0.290.0) · felhom-agent `d766666` (v0.138.0) · hub v0.128.0 |
> | **Freshness** | **CURRENT** as of 2026-10-04. The spike `TASK-backup-close-and-os-updates-spike-2026-10-04` adds measurements here as `[FACT]` and corrects every claim it disproves. Mark this file STALE when it falls behind what the product does. |
>
> **How to read this document.** Each statement has a label:
>
> - **[FACT]**: an observed property, with a `file:line`, a register row or an audit path.
> - **[RULED]**: an operator decision, with its date.
> - **[PROPOSAL]**: the reviewer's design. It is not decided and not built.
> - **OPEN**: a question that nobody has answered yet. The spike measures it. Nobody guesses it.
>
> **Why this file exists.** No architecture document covered operating-system updates. `00` §G marks the
> capability MISSING (added 2026-10-03). The finding is **R-812**. The roadmap intention is **R-808**.
> **The register carries the work. This file carries the reasoning. The source is the truth.**
>
> **Spike corrections, 2026-10-04** (`audits/os-updates-spike-2026-10-04/`). Each is made in place: the reviewer's
> text is struck (~~like this~~) and the measured text follows, labelled `[FACT]`, with its id.
>
> | Id | Where | What the spike changed |
> |---|---|---|
> | C1 | §2 | The agent may run `apt-get install` for TWO packages, not one (dnsmasq and wireguard-tools). |
> | C2 | §5.3 | Debian keeps two versions, not one; for box packages no fix was replaced within 14 days in 3 months; `snapshot.debian.org` works from a box in seconds. |
> | C3 | §5.2 | The lanes must follow the package's ORIGIN, not its name: 40 Proxmox-repository packages have ordinary names (ZFS, the Secure Boot shim, Ceph, corosync, chrony, CPU microcode). |
> | C4 | §5.6 | `--next-boot` is NOT a one-shot on these GRUB hosts. The fallback works only after a boot that reaches userspace. Installing a kernel alone makes it the default. Only a software watchdog runs. |
> | C5 | §5.6, §6 row 4 | Docker's `live-restore` keeps every container running across an engine update (measured). Turning it OFF again stops every container and starts none. |
> | C6 | §5.6 | A guest snapshot works on LVM-thin (customer guests) but not on `dir` storage; the snapshot rollback itself is unmeasured; a backup-restore undo took 73 s. |
> | C7 | §6 row 2 | A killed `apt` run does not recover by itself; the repair took ~5 s. |
> | C8 | §1, §6 row 14 | `cloudflared` is not a host package: it is a container in the guest, pinned by the controller since June. It belongs with the controller's infrastructure pins, not this file's lanes. |
> | C9 | §5.3 | The approved list must record what ring 0 RUNS healthy, not only what it installed that night. |
> | C10 | §5.5 | Restore-tests and agent updates are not windowed; the host's `apt` timers install nothing today. |
> | C11 | §5.2 | A Debian (fast-lane) update leaves PID 1, `lxc-start`, the Proxmox daemons and dockerd on the old library: its full effect needs a restart the fast lane does not do. |
> | C12 | §6 row 9 | The two demo hosts differ by 7 packages, including the CPU microcode (AMD vs Intel) and Secure Boot (on vs off). |
---
## 0. In plain language
A box runs three layers that we install and never update: the Proxmox host, the small Debian system
inside the guest, and the Docker engine. App images are updated (see `09`). The layer under them is not.
A box lives in a home for years, so this is a security gap.
The proposal has two lanes. **The fast lane** applies Debian's security fixes automatically, but only
the exact versions that already ran well on the demo boxes for 1–2 days. **The slow lane** covers the
kernel, Proxmox and Docker. These cause restarts or big changes, so a person approves them one version at
a time, and they run only at night. The tested versions are recorded automatically from what the demo
boxes installed. Nobody keeps a hand-written list.
---
## 1. Scope
**In scope.**
- The Proxmox host on an **appliance** install: the Debian 13 base, the Proxmox packages, and the kernel.
- The guest (the customer LXC): its Debian 13 packages.
- The Docker engine inside the guest (`docker-ce`, `docker-ce-cli`, `containerd.io`).
- How the household and the operator are told, and how a failed update is undone.
**Out of scope.**
- **App images.** `09` covers them (the ladder, the monthly same-tag re-test).
- **The controller and agent binaries.** Their own self-update covers them (`03`, the self-update section; `09` R-608).
- **DooPlex and ep0.** The operator updates them by hand (`runbooks/offsite-endpoint.md`).
- **A BYO host** (the owner brought their own Proxmox). `[FACT]` The installer leaves its repositories
alone: *"apt repo alignment skipped (byo — the owner manages repos)"* (`scripts/felhom-host-install.sh`,
`align_apt_repos`). `[PROPOSAL]` On a BYO box we update the guest and Docker only, never the host.
- **The Proxmox MAJOR upgrade** (PVE 9 → 10). It is a later step of its own, drilled first (R-808 item 5).
- **`cloudflared` and the other infrastructure images** (traefik, filebrowser). **[FACT] (C8)** They are containers in
the guest, pinned in the controller (`internal/infra/infra.go:26`: `cloudflare/cloudflared:2026.6.0`, since
2026-06-11) and baked into the golden; upstream was `2026.9.3` on 2026-10-04. They move only by a controller release
— the app-image question (`09`), not an OS package. The gap is R-838.
---
## 2. What a box runs, as measured
| Layer | What it is | Source of packages | Evidence |
|---|---|---|---|
| Host | Proxmox VE 9.2 on Debian 13, LVM-thin | Debian mirrors + `pve-no-subscription` (enterprise repo switched off on appliance installs) | `[FACT]` `01` §2; `felhom-host-install.sh` `align_apt_repos` (~L2130–2136) |
| Host kernel | The only kernel on the box. The guest has none. | `pve-no-subscription` (`proxmox-kernel-*`) | `[FACT]` LXC shares the host kernel (`01` §2: "one LXC/kernel/Docker daemon") |
| Guest | Unprivileged LXC, `nesting=1,keyctl=1`, from `debian-13-standard_13.1-2` | Debian mirrors | `[FACT]` `felhom-agent/configs/build-golden.sh:67,101-103` |
| Docker engine | `docker-ce`, `docker-ce-cli`, `containerd.io`, installed at golden BAKE time | `download.docker.com/linux/debian trixie stable` | `[FACT]` `build-golden.sh:108-125` |
| Docker settings | `containerd-snapshotter: false`, json-file log caps. **No `live-restore`.** | baked `daemon.json` | `[FACT]` `build-golden.sh:138-144` |
| Controller | A container in the guest. A Docker engine restart restarts it too. | Felhom registry | `[FACT]` `03` §1 |
**[FACT] Nothing updates any of these layers today** (R-812, searched 2026-10-03). The installer
says *"No upgrades are run — repo alignment only"* (`felhom-host-install.sh` ~L2135). A fresh install
gets the Docker engine that was current when its golden was baked, and keeps it.
~~**[FACT] The agent may not run `apt` today, with one exception.** Its sudoers allowlist
(`felhom-agent/configs/felhom-agent.sudoers`) holds one apt line: `apt-get install -y -q dnsmasq`.~~
**[FACT] (C1) The agent may run `apt-get install` for two named packages:** `felhom-agent.sudoers:58`
(`apt-get install -y -q dnsmasq`, used by `internal/lanresolver/lanresolver.go:107`) and `:194`
(`apt-get install -y -q wireguard-tools`, used by `internal/wgtunnel/manager.go:628`). Nothing else.
`apt` as root runs package scripts as root, so a broad `apt` grant is a full root grant. §5.4
proposes how to avoid that.
---
## 3. Operator rulings
**2026-10-03.** R-808 is on the roadmap at P2: *"every box receives operating-system security
patches on a schedule, and a failed update is undone."*
**[RULED] 2026-10-04: the fast lane follows an approved list, with a 1–2 day wait.** Every update
runs on the demo boxes first. The other boxes install only the exact versions that the demo boxes ran
without trouble. The exact wait is set when the feature is built. **Rejected:** Debian's own
`unattended-upgrades` with no wait. It is simpler and common, but a bad update would reach every
customer at the same time.
**[RULED] 2026-10-04: the off-site backup topic is closed first.** The same brief carries both
topics, and the backup part runs first.
The two-lane split (§5.2) is the reviewer's proposal. The operator's ruling above assumes it, but he
has not ruled on it as such.
---
## 4. The constraints that shape the design
1. **Unattended.** Nobody is at the box. Nobody answers a question that `apt` asks.
2. **No screen.** If the box does not boot, the household sees only that nothing works.
3. **One kernel for everything.** A host kernel update needs a host reboot. A reboot stops every app.
4. **The controller lives in Docker.** A Docker engine update restarts the controller in the middle of
its own work. So the controller cannot drive a Docker update. The agent, on the host, must.
5. **The host has no whole-system backup.** The guest has three tiers (`07`). The host has none.
A host update can only be undone by installing the previous version again.
6. **Every update must already have run on a box we own.** This is the lesson of the update arc
(`09` §3 decision 13: "the test decides").
7. **Few packages.** Each extra package is one more thing to update and break. The guest is the
Debian standard template plus Docker. Keep it that way.
---
## 5. The proposed shape `[PROPOSAL]`
### 5.1 Two rings
- **Ring 0:** demo-felhom (N100) and demo-hp. They are disposable (`runbooks/target-selection.md`), and
they have different hardware. They take every update first.
- **Ring 1:** every other box. It takes only what ring 0 approved.
### 5.2 Two lanes
| | Fast lane | Slow lane |
|---|---|---|
| What | Debian packages on host and guest, from `trixie-security` and the stable point releases, **except** the slow-lane list | Kernel (`proxmox-kernel-*`), Proxmox (`pve-*`, `proxmox-*`, `lxc-pve`, `qemu-server`, …), Docker (`docker-ce*`, `containerd.io`) |
| Restart | A service restart at most. No reboot. | Kernel: a host reboot. Docker: every container restarts. Proxmox: its services restart. |
| Approval | Automatic. A version is approved when ring 0 ran it and stayed healthy for the wait (1–2 days). | A person approves one version at a time, like a controller floor. |
| Urgent fix | The operator can approve a version on the same day once ring 0 has run it. | The same. |
| When | The night window (§5.5) | The night window. A kernel reboot only on a night the operator scheduled. |
~~Which packages count as slow lane is a list the spike checks (OPEN Q7). A Debian package that restarts
something big (for example `systemd`, `libc6`, `openssh-server`) may belong in the slow lane too.~~
**[FACT] (C3) The lane must be decided by ORIGIN, not by name.** On demo-hp, 40 pending packages come from the
Proxmox repository under ordinary names: `zfsutils-linux`, `zfs-zed`, `libzfs7linux`…, **`shim-signed` and friends (the
Secure Boot loader)**, `ceph-common`/`librados2`…, `corosync`, `chrony`, `frr`, `amd64-microcode`
(`partH/H2-proxmox-origin-debian-names.txt`). A name rule (`pve-*`, `proxmox-*`) would have put them in the fast lane.
The fast lane is: origin `Debian` or `Debian-Security`, and nothing else. On demo-hp that selection was 108 packages
and pulled in **zero** Proxmox packages.
**[FACT] (C11) What a fast-lane run restarts, and what it does not.** Guest (9202, 49 packages incl. libc6): the
packages' own scripts restarted postfix, journald, networkd; **dockerd, containerd, sshd, dbus, logind, cron** kept the
old libc; no container stopped (13 samples). Host (demo-hp, 108 packages): dnsmasq, postfix, journald restarted;
**systemd (PID 1), `lxc-start`, pveproxy, pvedaemon, pvestatd, pvescheduler, watchdog-mux, sshd, zed, chronyd** kept the
old libc; both guests and the agent stayed up. So the fast lane is safe to run unattended, but a libc fix is only
fully in force after a reboot (host) or a Docker restart (guest) — which are slow-lane acts. `[PROPOSAL]` the box
reports "restart needed" (processes on deleted libraries) and the slow lane's next reboot picks it up.
### 5.3 The approved list (the "tested versions" record)
Nobody writes the list by hand. It fills itself:
1. A ring-0 box updates. Afterwards it reports to the hub the exact `package=version` it installed,
per layer (host, guest), with each package's origin (`Debian-Security`, `Debian`, `Proxmox`,
`Docker`).
2. The box then reports health for the wait period: the agent, the controller, every app's health,
and the guest's network.
3. When the wait passes with ring 0 healthy, the hub marks that set **approved**. It is one record:
an **OS release**, with an id and a date. The hub stores it. The register does not.
4. A ring-1 box asks the hub for the newest approved OS release. For each package it has installed,
if the approved version is newer, it installs that **exact** version. It never installs a version
newer than the approved one.
5. Each box reports which OS release it runs, how many updates are waiting, whether it needs a reboot,
and which packages it has that **no approved list covers** (see §6, edge case 9).
~~**OPEN Q1 is the weak point.** The design works only if a box can still download the approved version
a week later. Debian's main and security archives keep only the newest version of each package. If
Debian publishes a newer fix between approval and install, the approved version is gone. Options:
the box waits for the next approval; or the box uses `snapshot.debian.org` (Debian's own dated archive,
still signed by Debian); or Felhom runs a package cache. The spike measures how often this happens.~~
**[FACT] (C2) Q1 measured.** The live Debian archives keep **two** versions: the point-release one in `trixie`
(main) and the newest in `trixie-security`; intermediate versions are gone (openssl: installed `u1`, main `u2`,
security `u3`). Over 2026-07-04..10-04, 167 trixie security advisories; for the 517 source packages installed on a box,
**no package got a second advisory within 2, 7 or 14 days** (the 3 within 2 days were chromium and webkit2gtk, not on a
box). `snapshot.debian.org` answers a box: a dated index in **2.3–3.0 s**, a gone exact version
(`openssl 3.5.6-1~deb13u1`) downloaded in **2.0 s**, Debian-signed. Proxmox and Docker keep many old versions
(pve-manager 66, docker-ce 46). So with a 1–2 day wait the approved version is almost always still live; the rare
miss is fetched from the snapshot taken at approval time. `[PROPOSAL]` each OS release records its approval
timestamp; a box installs from its own sources, and only for a Debian package that is no longer there, from
`snapshot.debian.org/archive/<debian|debian-security>/<timestamp>`. This is the operator decision in STATUS.
**[FACT] (C9) Approve what ring 0 RUNS, not what it installed.** In the simulation on demo-felhom
(`partI/demo-felhom-simulation.txt`), every one of the 108 host and 49 guest approved versions was installable and
downloadable — but `curl`, `libcurl*` and `libssh2` were "not covered" in the guest, only because scratch guest 9202
ALREADY ran the newer version and so installed nothing. `[PROPOSAL]` step 1 reports the full installed
`package=version` set after the run, and approval covers every version ring 0 runs healthy.
### 5.3.1 Test approvals end with the test — BUILT 2026-10-04 (hub v0.133.0, R-859) `[FACT]`
An approval made while a TEST override (`OS_APPROVE_AFTER`, `OS_APPROVE_NIGHTS`, `OS_DOCKER_APPROVE_NIGHTS`) is active
carries a `test` mark (amber on the System page). At every hub start WITHOUT an override, every test approval that no
real approval has superseded is cancelled: never served again, ring-1 boxes bumped, one operator event
`os_release_cancelled` each; what boxes installed stays; the ruled wait approves the same set again as a real release.
A one-time backfill marked the AUTOMATIC approvals made under 24 h after first seen; the 2026-10-04 guest and host test
approvals (which Tester 2 installed on its first night) were cancelled at 18:20 UTC. The operator's Docker button
approval of that day stays in force (decided by CC unattended — operator may reverse). Runbook:
`runbooks/os-updates-test-waits.md`. Evidence: `audits/r840-config-bundle-2026-10-04/partD/`.
### 5.4 Who runs it, and with what permission
- **The agent runs every OS update**, for the host and for the guest (`pct exec`). The controller does
not, because of constraint 4.
- **The agent gets no general `apt` permission.** Like `felhom-selfupdate-guarded` (`03`, the self-update section),
a small root-owned wrapper does the work. It accepts only an approved list and refuses everything else:
- no package removal;
- no downgrade, except the undo of §5.6, which an operator job signs;
- no package that the box does not already have, unless the approved list records it as a dependency
that the same update pulled in on ring 0 (a kernel update installs a NEW package name each time,
for example `proxmox-kernel-6.x.y-z-pve-signed`, pulled by `proxmox-default-kernel`);
- no package source other than the ones the installer set up;
- non-interactive, and it always keeps the existing config file (`--force-confold`) and reports the
conflict.
- **Package signatures stay the publishers'** (Debian, Proxmox, Docker). The hub sends only names and
versions. A broken-into hub can choose an older version or no version. It cannot make a box install a
package that the publisher did not sign.
### 5.4.1 The root wrapper's interface — DRAFT (2026-10-04, design only, nothing installed) `[PROPOSAL]`
`felhom-os-apply` — root-owned (`0755 root:root`), installed by the installer beside `felhom-selfupdate-guarded`;
the agent's sudoers gets exactly `/usr/local/sbin/felhom-os-apply --plan /var/lib/felhom-agent/os/plan-*.json` and
`… --repair-only`. For the guest it runs on the host and enters the guest with `pct exec <vmid> --` itself, so the
agent needs no `pct exec … apt` line. Built from what Parts G and H measured.
**Input — one JSON plan file** (written by the agent, from the hub's approved OS release):
```json
{
"release_id": "os-2026-10-04-1", "approved_at": "2026-10-04T08:00:00Z",
"layer": "host", // "host" | "guest"
"vmid": 9201, // guest only
"snapshot": "20261004T080000Z", // the snapshot.debian.org timestamp of approval (C2)
"packages": [ {"name": "libc6", "version": "2.41-12+deb13u4", "origin": "Debian"} ],
"allow_new": ["proxmox-kernel-7.0.14-20-pve-signed"], // slow lane only, signed operator job
"lane": "fast" // "fast" | "slow"
}
```
**Order of work:** (1) refuse checks below; (2) **repair first** — `dpkg --configure -a` then `apt-get -f install`,
logging what it repaired (C7); (3) `apt-get -s install` of exactly `name=version` for every package that is
installed AND older; (4) refuse if the simulation would remove, downgrade, add an unlisted package, or touch a package
whose candidate origin is not the plan's; (5) download — from the box's own sources, or for a Debian version no
longer there, from `snapshot.debian.org/archive/<archive>/<snapshot>` with a temporary sources list it deletes after
(C2); (6) install with `DEBIAN_FRONTEND=noninteractive APT_LISTCHANGES_FRONTEND=none -o Dpkg::Options::=--force-confold
-o Dpkg::Options::=--force-confdef`; (7) `apt-get clean`; (8) report.
**Refusals** (each exits non-zero with one line `os-apply: REFUSED: <reason>` and changes nothing):
| # | Refuses when |
|---|---|
| R1 | the plan file is not under `/var/lib/felhom-agent/os/`, not owned by the agent, or not valid JSON |
| R2 | `lane` is `fast` and any package's origin is not `Debian` / `Debian-Security` (C3) |
| R3 | `lane` is `slow` and the plan is not carried by a verified signed operator job (R-530's mechanism) |
| R4 | the simulation removes any package |
| R5 | the simulation downgrades any package (the operator undo is a separate signed op, §5.6) |
| R6 | the simulation installs a package that is neither installed nor in `allow_new` |
| R7 | a listed version is not downloadable from the sources the installer set up or the named snapshot |
| R8 | free space on `/` (or the guest's rootfs) is below 3× the download size, minimum 500 MB (edge case 8). The download size is the `--print-uris` total WITHOUT `-s` (R-865, agent v0.144.0 — with `-s` it read 0 B) |
| R9 | another apt/dpkg holds the lock, or the per-guest lane lock is held (a backup, a restore-test, C10) |
| R10 | `layer` is `guest` and the vmid is not the box's own customer guest |
| R11 | the plan names a package twice, or a version that is not a Debian version string |
**Log lines** (to the journal, tag `felhom-os-apply`, and echoed for the agent to forward to the hub):
```
os-apply: START release=<id> layer=<host|guest:vmid> lane=<fast|slow> packages=<n>
os-apply: REPAIR configured=<n> fixed=<n> (always printed; 0 0 when nothing was half-done)
os-apply: PLAN upgrade=<n> already=<n> not-installed=<n> from-snapshot=<n> download=<bytes>
os-apply: REFUSED: <R-number> <reason>
os-apply: CONFFILE kept <path> (new version saved as <path>.dpkg-dist)
os-apply: DONE rc=0 seconds=<s> upgraded=<n> restarted=<unit,…> restart-needed=<process,…> reboot-needed=<yes|no>
os-apply: FAILED rc=<n> step=<download|install> — dpkg state: <dpkg --audit first line>
```
`restart-needed` lists processes still mapping deleted libraries (C11); `reboot-needed` is yes when that list holds
PID 1 or `lxc-start`, or a kernel was installed.
### 5.4.2 The config bundle: a box's root-owned files by a signed job — BUILT 2026-10-04 (agent v0.143.0, hub v0.133.0, installer 1.31.0, R-840) `[FACT]`
Decision 96. Before it, the signed `agent_update` replaced only the binary; wrappers, units and sudoers lines reached
an installed box by reinstall or by hand.
- **One source of truth.** `BUNDLE_FILES` in `felhom-os-apply` is the ONE table of root-owned paths (22: sudoers ×2, the
five wrappers, the crash guard and its units + config, the agent and rollback units, the start-limit drop-in, the mgmt
watchdog and its tmpfiles/units, the OOB belt's four files). `scripts/build-config-bundle.py` builds the bundle from it
reproducibly; `release-agent.sh` publishes it beside the binary; the hub vouches its sha with the agent (exact-name
lookup); the installer (1.31.0) installs it through the same code (`--install-bundle`, root only, refused through
sudo). A test fails on a root path the installer names that the bundle lacks.
- **The route.** Signed `agent_config_update` {agent_version, bundle_sha256}. The agent is a courier; the root wrapper
re-verifies the signature against the root-owned signers file, the host binding (`os-trust.json`), the window and its
own nonce, then the sha, every path (R16) and every content check before the first write: `visudo -cf`, `sh/bash -n`,
Python compile, unit sections, no `RuntimeDirectory=` (G1), `User=felhom-agent`, `nft -c`, and that the route
survives (the sudoers keeps the `--plan` line; the new wrapper keeps the bundle mode). Policies: replace / if-absent
(`crash-guard.conf`, an operator setting) / oob (only on a box with the belt). Atomic per file, sudoers last,
previous copies kept (last 3). Self-check: `visudo -c`, `sudo -l -U felhom-agent` lists the route, the new wrapper's
`--self-check`, the self-update wrapper's usage, the crash guard's status equals `kernel.panic`; any failure puts
every previous copy back. A newly installed crash guard is started (`enable --now`): `kernel.panic` for this boot, no
reboot. Record `/etc/felhom/config-bundle.json`; the agent reports it (`system.config_bundle`, "none" when absent);
the facts mode adds drift (files changed by hand).
- **The trust root is not changed by a bundle** (R17, tested). A missing signers file is created only with the
installer's pinned key, only after a job that key signed (pinned equal to the installer by a test). Signer rotation is
a later, separate act (`04` §3).
- **Bootstrap (corrects the brief).** The self-update wrapper cannot install a bundle (fixed sh, binary-only), and no
signed job can write a root file on a box whose `felhom-os-apply` predates 0.143.0. Such a box needs ONE by-hand step
(`scripts/felhom-bundle-bootstrap.sh`: only the new `felhom-os-apply`). Done on both demo boxes; Tester 2 needs the
operator (`runbooks/config-bundle.md`).
- **Visibility.** System page "Root files" column; alarm `os_config_bundle_behind` after 7 days
(`OS_ALARM_BUNDLE_BEHIND_AFTER`; decided by CC unattended — operator may reverse).
- **Measured live** (`audits/r840-config-bundle-2026-10-04/partB/`): both demo boxes had every file equal to the release
except `felhom-os-apply`; after the bootstrap the signed bundle wrote 0 of 22 (21 same, 1 setting kept), self-check
ok, capability probe 71/71. demo-hp: a wrong-sha job refused (nothing changed); a bundle with one deliberate change
wrote exactly that file; the 0.143.0 bundle undid it; a replayed job was rejected. The installer path on demo-felhom:
0 written, services active. 22 of 22 wrapper rules red-proved.
- **Not a boundary yet — R-861.** The agent's sudoers already lets the agent user reach root without the operator key
(a hookscript, a boot unit, the escrow self-test run as root, the self-update of its own binary). The trust-root rule
is defence in depth until R-861 is closed.
### 5.5 When
Inside the household's night window, after the backups:
```
W DB dump
W+60m Tier 2
W+105m off-site → app updates (until W+5h at most)
[W+2h, W+6h) whole-guest backup (agent)
after it OS updates — guest first, then host
```
The OS leg starts **after the whole-guest backup has finished**, so the guest's newest full copy is
minutes old. This is the opposite order to app updates, which run before the whole-guest backup
(`07` §6.1, `09` decision 11). The reason: the whole-guest backup is the guest's undo.
**[FACT] (C10) Q8 measured** (guest UTC; demo-hp W = 02:30): db-dump 02:30, tier-2 03:30, off-site ~04:15,
whole-guest gate [04:30, 08:30), controller self-update 04:30, offsite-integrity 06:00. Host: `apt-daily` and
`apt-daily-upgrade` run daily but install nothing (no `unattended-upgrades`, no `APT::Periodic`); `pve-daily-update`
refreshes the lists daily. **Restore-tests are NOT windowed** — they run on a cadence at any hour (demo-felhom 10:38
daily, demo-hp 16:43 and 22:46); agent updates arrive by signed job at any hour. `[PROPOSAL]` the OS leg takes the same
per-guest lane lock the restore-test and the whole-guest backup take, rather than a clock slot.
~~**OPEN Q8:** what else runs then. The controller's self-update (default 04:30, and after any hub report
when a floor is above it, R-608), the agent's self-update, the restore-tests, and PBS jobs. The OS leg
must never overlap a backup, a restore-test or a self-update.~~
### 5.6 How a failed update is undone
"Rollback" is not used (`09` §4). The shapes:
| Layer | Undo | Limit |
|---|---|---|
| Guest packages | Restore the guest snapshot taken just before the update | It also undoes app data written after the snapshot. Use it only inside the health window, before apps have written much. After that, install the previous version (needs OPEN Q1). **[FACT] (C6)** Customer guests are on LVM-thin and can snapshot; scratch 9202 (`dir` storage) cannot (`snapshot feature is not available`), so the measured undo was a whole-guest backup + restore: **73 s down**, all apps healthy, libc back. The snapshot rollback itself is unmeasured (R-837). |
| Docker engine | Install the previous version (Docker's repository keeps old versions — OPEN Q2) | Every container restarts again. **[FACT] (C5)** Docker keeps 46 `docker-ce` versions. A step (either way) without `live-restore`: 6 of 6 containers restart, the app silent **26.5–30 s**, healthy at +42–45 s. With `live-restore` on: **0 restarts, no gap**, also across a containerd step. But a restart that turns `live-restore` OFF stops every container and starts NONE (`unless-stopped` ignored) — on 9202 they stayed down until restarted by hand; `systemctl reload` turns it on but not off. |
| Host packages | Install the previous version | Only if the source still has it (OPEN Q1/Q2). **[FACT]** `rsync` back to its pre-update version: `not found`; `libpng16-16t64` back to the point-release version: worked. A Debian undo needs the snapshot archive (C2). **[FACT, 2026-10-04]** Proved by hand: `runbooks/os-updates-host-undo.md` (demo-hp, `tzdata` back one version from `snapshot.debian.org`, held, released). Use the NEW version's `first_seen` as the timestamp when there is no previous host release; a package that pins its siblings (`eject` → `libmount1 =`) goes back only with them. |
| Host kernel | ~~Boot the previous kernel. Proxmox can boot a new kernel **once** (`proxmox-boot-tool kernel pin <ver> --next-boot`). If that boot fails, the next boot uses the old kernel again. Make the new kernel permanent only after a healthy boot.~~ **[FACT] (C4)** Both demo hosts boot UEFI + GRUB (no proxmox-boot-tool ESPs). **Installing a kernel makes it the GRUB default at once.** `--next-boot` on GRUB writes an ordinary `GRUB_DEFAULT` + `update-grub`; `proxmox-boot-cleanup.service` clears it only once a boot reaches userspace. Measured on demo-hp (old kernel pinned permanently FIRST, new pinned for next boot): boot 1 → `7.0.14-20-pve`, healthy, 60 s; boot 2 → back on `7.0.2-6-pve`, healthy, 76 s. **So the fallback works after a boot that succeeds; read from the code, a kernel that hangs before userspace stays the default on every power cycle.** ~~`[PROPOSAL]` use GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel as the saved default — unmeasured (R-836).~~ **[FACT, 2026-10-04, demo-hp, operator's word before each reboot] GRUB's one-shot is NOT a one-shot here either.** With `GRUB_DEFAULT=saved` (old 7.0.2-6 saved) and `grub-reboot` 7.0.14-20: boot 1 → 7.0.14-20, **Secure Boot ON, booted fine** (signed kernel, shim → GRUB); but `/boot` is ext4 on LVM, GRUB cannot write its environment block there (`grub-reboot` warns so itself), `next_entry` was never cleared, and boot 2 with no command → **7.0.14-20 again**. A kernel lane needs a writable env block (the ESP) or a userspace "boot good" step — R-836. demo-hp left on 7.0.14-20, saved default 7.0.14-20, both kernels installed (`audits/os-host-lane-2026-10-04/partE/`). | ~~If the new kernel hangs, someone must switch the box off and on. The spike checks whether a hardware watchdog can do that (OPEN Q4).~~ **[FACT]** Only `softdog` runs (loaded by `watchdog-mux`); it cannot rescue a kernel that never boots. demo-hp has an AMD FCH whose `sp5100_tco` driver ships but is not loaded; untested. A hang still needs a person. **[FACT, 2026-10-04]** `sp5100_tco` is blacklisted by the Proxmox kernel package; loaded by hand it answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0 — read from sysfs, never opened, so never armed; unloaded). `kernel.panic = 0`: a panic leaves the host stopped (R-851). |
### 5.7 Telling people
- **Operator:** a hub event for each update and each failure; a fleet view showing each box's OS release,
how far behind it is, and whether it needs a reboot. A box that is more than N days behind the newest
approved release raises an alarm (`08`).
- **[FACT, hub v0.132.0] The System page** (`/system`, R-852, decision 89): one row per box — ring and switch with
buttons, the tunnel, host Proxmox / running and next-boot kernel / Debian / release / pending / not covered / held /
reboot needed / `kernel.panic` / oops / crash restarts / the guard, guest Debian / release / pending / restart needed,
Docker engine / containerd / live-restore / release, the last leg — and above it the releases, what ring 0 runs, "Approve
now" and "Approve Docker set". Colours are the alarm thresholds (decision 94). Hosts shows Proxmox / kernel too.
- **Household:** one line on the timeline in both languages, informal voice: what was updated and
whether the box restarted. Telling households in advance that the box may restart at night is a
**promise to users**. That is the operator's decision when the slow lane is built.
### 5.8 The Docker engine slow lane — BUILT 2026-10-04 (agent v0.142.0, hub v0.132.0) `[FACT]`
**As built** (evidence `audits/os-docker-crash-2026-10-04/partB/`; decisions 87, 93):
- **live-restore ON**: the golden bakes it (`build-golden.sh` 3.1.0, fail-closed assertion); an installed box gets it once
by the wrapper's `live-restore-on` (merge into daemon.json + `systemctl reload docker`). Measured: the same container ids
after (9202 6/6 by hand — R10 refuses a scratch guest by design —, demo-hp 24/24, demo-felhom 5/5).
- **The step** runs in the night leg after a healthy guest and host step, **ring 0 only**, `select pending-docker`,
allowed by the wrapper only with the root-owned `ring0_slow_lane` mark. **Ring 1 and every undo** only through a signed
`os_docker_step` the wrapper re-verifies itself (decision 93). Measured: 29.7.x → 29.8.2 on both demo boxes, every id
kept; a signed undo to 29.7.2 on demo-hp (41 s, ids kept) and back; demo-felhom as ring 1 by a signed job; a replayed
job refused by the agent (nonce).
- **Approval**: the hub never approves a Docker set automatically; the System page's button works after every ring-0 box
ran the set in 2 healthy night Docker steps (`OS_DOCKER_APPROVE_NIGHTS` TEST override, logged). An approval nudges no
box. Undo: `runbooks/os-updates-docker-undo.md`.
**The design as written before the build:**
Built from C5 (measured) and R-835. **The one decision it needs is in STATUS: `live-restore` on, fleet-wide.**
- **What moves:** `docker-ce`, `docker-ce-cli`, `containerd.io`, `docker-buildx-plugin`, `docker-compose-plugin`,
`docker-ce-rootless-extras` in the customer guest — today the six "not covered" packages on every box. One
approved **engine set** at a time, like a controller floor: the operator approves it (slow lane, §5.2), after ring 0
has run it for at least 2 nights healthy. Never two steps in one night.
- **Precondition: `live-restore` ON.** Without it an engine step restarts every container — 26.5–30 s of silence,
healthy at +42–45 s (C5). With it: 0 restarts, no gap, also across a containerd step. Turning it ON is safe
(`systemctl reload docker` applies it without a restart, C5); turning it OFF later by a plain restart stops every
container and starts none (R-835) — so it is turned on once, by the golden and by a one-time fleet step, and never
turned off by the lane.
- **Who and when:** the agent, through the same wrapper (`lane: slow`, refusal R3: only inside a verified signed
operator job, R-530's mechanism), in the guest, after the guest and host fast-lane steps, under the same heavy-op
gate, on a night the operator scheduled. Debian origin rule replaced by "origin `Docker CE`, exactly these names".
- **Health:** the guest rule (§8.1) plus `docker version` reports the approved engine, and every container running
at the start is running with the SAME container id (proof that `live-restore` held). A changed id is
`health_failed` even if the app is healthy — it means the households' apps restarted when they should not have.
- **Undo:** install the previous engine set (Docker's repository keeps 46 versions, C2) — by an operator job, with
`live-restore` still on, so the undo is also restart-free.
- **Not covered here:** the golden's own engine (baked weekly; a new golden carries the approved set), and BYO hosts
(the guest is ours on both, so the lane applies there too).
### 5.9 A crashed host restarts, with a limit — BUILT 2026-10-04 (agent v0.142.0, installer 1.30.0) `[FACT]`
Decision 88 (R-851): *"Yes, but maybe not indefinitely."* Evidence `audits/os-docker-crash-2026-10-04/partC/`.
- **Measured first (demo-hp, the operator's word before each crash):** `kernel.panic = 10` + `echo c >
/proc/sysrq-trigger` → the box restarted by itself in 54 s, on the same kernel (the saved default), the agent up 18 s
after the boot. `kernel.panic` set by `sysctl -w` is gone after the restart (back to 0) — it must be set at every boot.
- **The crash signal** (decision 90): `efi_pstore` is on, yet a real panic saved NOTHING; the journal and `last` show
only "no shutdown". The guard uses a **clean-stop marker** (the unit's ExecStop at every orderly shutdown); a boot
without it followed a crash, a power cut or a hard reset — counted alike.
- **The guard** (`felhom-crash-guard`, early boot unit + hourly re-arm timer): armed → `kernel.panic = 10`; the 2nd
unclean boot within 60 minutes TRIPS it (`kernel.panic = 0`), so **the 3rd crash within the hour leaves the box off**
(decision 92 — the operator's words); re-arms after 24 h of normal running or `felhom-crash-guard rearm`. A crash before
the unit runs (very early boot) leaves the box off: the safe side. `panic_on_oops` stays 0; an oops is reported
(decision 91).
- **Telling people:** each unclean boot = an operator mail (`host_crash_restart`) + the household's line; the trip = an
alarm (`host_crash_guard_tripped`); re-arm and oops announced once (`08` §6.3). The System page shows the guard.
- **Measured live:** crash 1 → back in 54 s; crash 2 → back in 53 s, guard tripped; crash 3 → **stayed off** until the
operator switched it on; the hub mailed the trip (15 min after the boot — R-853); re-armed by hand.
---
## 6. Risks and edge cases
| # | What can go wrong | What the design does |
|---|---|---|
| 1 | A new kernel does not boot | Boot it once (`--next-boot`); the old kernel stays the default. A hard hang needs a power cycle (OPEN Q4). |
| 2 | Power cut during an update: `dpkg` is half done | The wrapper runs `dpkg --configure -a` and `apt-get -f install` first, every time, and reports what it repaired (spike measures, Q5). **[FACT] (C7)** Killed after 15 unpacks: 5 packages `iU`, 4 triggers pending; the next ordinary `apt-get install` REFUSES (`Unmet dependencies`) — nothing repairs it by itself. The two commands repaired it in 1.4 s + 3.9 s; apps stayed up. |
| 3 | An update asks a question (changed config file, service restart prompt) | Non-interactive, keep the old config, report the conflict. |
| 4 | A Docker update stops every app, and the controller | Stop the apps cleanly first, like before a backup. Measure whether `live-restore` keeps containers running (Q3). The agent drives it. **[FACT] (C5)** `live-restore` keeps them running (0 restarts). The controller keeps running too; its `docker` calls fail for the seconds dockerd is down (logged errors, no app event). Turning `live-restore` on is a golden + fleet change — and turning it off later must not be a plain restart (R-835). |
| 5 | The approved version is no longer downloadable | OPEN Q1. Until it is answered, the box waits for the next approval and reports it. |
| 6 | An urgent security hole | The operator approves the same day once ring 0 has run it. |
| 7 | A box was off for months | It catches up through the newest approved release. It does not install every release in between. Packages are not like app data; one `apt` step is enough. Kernel and Docker still go one approved step at a time. |
| 8 | The system disk is full | Check free space before downloading; clean the package cache after; refuse and report. |
| 9 | A customer box has a package that ring 0 does not have (other hardware: firmware, CPU microcode, NIC drivers) | It is never approved, so it is never updated. The box reports it as "not covered", and the hub raises it. Fix: add matching hardware to ring 0, or approve it by hand. **[FACT] (C12)** The two demo hosts differ by 7 packages: `amd64-microcode`, `proxmox-secure-boot-support`, `felhom-bootstrap` (demo-hp) vs `intel-microcode`, `proxmox-first-boot`, `tailscale`, `tailscale-archive-keyring` (demo-felhom); Secure Boot is ON on demo-hp, OFF on demo-felhom. Ring 0 covers both CPU vendors today. |
| 10 | The package source is down, or its signing key changes (Docker has done this) | The update fails cleanly and the box reports it. A key change is a slow-lane act for a person. |
| 11 | A BYO host | Only the guest and Docker are updated (§1). |
| 12 | Two boxes on ring 0 is a small sample | Accepted for now. When there are customers, the first tester boxes can become a second ring. |
| 13 | A broken-into hub sends a harmful list | The wrapper refuses removals, downgrades and new packages, and the publishers' signatures still apply (§5.4). |
| 14 | `cloudflared` on the host | ~~Not covered by this file yet. How it is installed and updated is OPEN Q9. It faces the internet, so it matters.~~ **[FACT] (C8)** Not on the host: a pinned container in the guest (§1). Four months behind upstream on 2026-10-04 (R-838). |
| 15 | The no-subscription Proxmox repository gets less testing than the enterprise one | Ring 0 is our test. The enterprise repository costs a yearly fee per box. That is a **money** decision for the operator, later (Q10). |
---
## 7. Open questions the spike must answer
| Q | Question | How to answer |
|---|---|---|
| Q1 | Can a box install an exact Debian version one week after approval? How often is it already superseded? | `apt-cache madison` on host and guest for security packages. Debian's security announcement history. Whether `snapshot.debian.org` is reachable and fast enough. |
| Q2 | Do the Proxmox and Docker sources keep older versions? | `apt-cache madison` for `pve-manager`, `proxmox-kernel-*`, `docker-ce`, `containerd.io`. |
| Q3 | What does a Docker engine update do to running containers, with and without `live-restore`? | On scratch guest 9202: step from one pinned version to the next; time the downtime. |
| Q4 | Does `proxmox-boot-tool kernel pin --next-boot` work on both demo hosts' boot setups (UEFI or legacy, GRUB or systemd-boot)? Does either box have a hardware watchdog? | On demo-hp, with the operator's go before each reboot. |
| Q5 | What does an interrupted `apt` run leave, and does the repair recover it? | Kill a run on 9202, after a snapshot. |
| Q6 | How far behind are the demo boxes today, per layer and per source? How long does catching up take, and what restarts? | `apt-get -s upgrade` (read only) first; then a real run on 9202 and on one demo host. |
| Q7 | Which Debian packages restart something big? | `needrestart` in list mode after a run. |
| Q8 | What else runs in the night window, and where does the OS leg fit? | Read the timers on host and guest. |
| Q9 | How is `cloudflared` installed and updated on the host? | Read the installer and the agent. |
| Q10 | What does the Proxmox enterprise repository cost per box per year, and what does it add? | The publisher's price page. Record only. The operator decides later. |
---
### 7.1 Answers — the 2026-10-04 spike `[FACT]`
All evidence: `audits/os-updates-spike-2026-10-04/` (its `README.md` carries every number).
| Q | Answer in one line | Detail |
|---|---|---|
| Q1 | Usually yes: Debian keeps 2 versions; no box package was re-fixed within 14 days in 3 months; the snapshot archive serves a gone version in 2 s. | C2 |
| Q2 | Yes for Proxmox (30–66 versions) and Docker (18–46); Debian only the point-release version. | C2, §5.6 |
| Q3 | Without `live-restore`: every container restarts, ~30 s of silence. With it: none. Switching it off is a trap. | C5 |
| Q4 | `--next-boot` falls back only after a boot that succeeds (measured); a hang keeps the new kernel (code). Software watchdog only. | C4 |
| Q5 | Not by itself; `dpkg --configure -a` + `apt-get -f install` repair it in ~5 s. | C7 |
| Q6 | Hosts 188 pending each, guests 54–59; guest Debian 24 s, host Debian 60 s, kernel 47 s. Three guests, three Docker versions. | audit README |
| Q7 | Few restarts by script; libc leaves PID 1, `lxc-start`, Proxmox daemons and dockerd on the old library. Proxmox packages restart their own daemons. | C11 |
| Q8 | The backup legs are windowed; restore-tests and agent updates are not; host apt timers are inert. | C10 |
| Q9 | Not a host package: a pinned guest container, 4 months behind. | C8 |
| Q10 | €120 (Community) to €1,100 (Premium) per CPU socket per year, net; every tier includes the Enterprise Repository. Money — the operator's. | audit README |
**Sample approved list** built from what ring 0 installed (`partI/sample-approved-list.tsv`: 108 host + 49 guest
packages, with origin) and simulated read-only on demo-felhom: **all 157 would install, exact version, downloadable**;
not covered: 79 Proxmox + 1 Tailscale on the host (slow lane / not ours), 6 Docker + 4 Debian in the guest (C9).
## 8. Build order `[PROPOSAL]`
Each step returns to the operator for go or no-go.
1. **Spike** (measure Q1–Q10; no product code).
2. **Guest Debian, fast lane.** ~~Lowest risk: a snapshot undo exists.~~ **BUILT 2026-10-04** — agent v0.140.0, hub
v0.130.0, installer 1.29.0; §8.1. **There is no snapshot undo** (R-837, measured).
3. **Host Debian, fast lane** (no kernel, no Proxmox packages). **BUILT 2026-10-04** — agent v0.141.1, hub v0.131.1;
§8.2.
4. **Fleet view and alarms** (§5.7). **BUILT 2026-10-04** — hub v0.131.0/v0.131.1; §8.3.
5. **Slow lane: Docker engine.** **BUILT 2026-10-04** — agent v0.142.0, hub v0.132.0; §5.8.
**Root files to installed boxes (R-840): BUILT 2026-10-04** — agent v0.143.0, hub v0.133.0, installer 1.31.0; §5.4.2.
**Test approvals end with the test (R-859): BUILT** — hub v0.133.0; §5.3.1. **The golden carries the approved guest
release** (`build-golden.sh` 3.2.0 `GOLDEN_GUEST_PKGS`; golden 0.293.0 baked with none in force — 49 Debian updates
pending for the next real approval; the host stays with its first night's OS leg). **A host pass at install time is
not needed:** measured on Tester 2, the agent's first OS leg ran right after the box's FIRST whole-guest backup, 17 min
after enrolment (16:24/16:25 UTC: 49 guest + 106 host packages, 110 s + 44 s, healthy). An installer pass would cost
~45 s and run BEFORE any whole-guest backup exists (no undo) — not built.
6. **Slow lane: host kernel and Proxmox packages, with the reboot.**
7. **Later:** the Proxmox major upgrade (PVE 9 → 10), drilled on ring 0 first.
---
### 8.1 Step 2 as BUILT (2026-10-04) `[FACT]`
Evidence: `audits/os-guest-lane-2026-10-04/` (parts A–G). Brief: guest fast lane, decisions 78–80 (`09` §3).
- **The undo (R-837, measured first): none automatic.** PVE refuses ANY snapshot of a customer guest —
`PVE/AbstractConfig.pm:755-757` skips non-snapshot mounts only for a snapshot named `vzdump`, and every customer
guest carries the host-path binds mp8/mp9. As the agent's token and as root: `snapshot feature is not available`.
The token's role HAS `VM.Snapshot` / `VM.Snapshot.Rollback`. So §5.6's first row does not exist: a failed health
check stops, reports `health_failed`, and the hub mails the operator; the whole-guest backup taken minutes earlier
is the undo, by hand. The choice of a real undo is in STATUS (R-842).
- **The wrapper** `felhom-os-apply` (agent repo `configs/`), Python 3 stdlib, per §5.4.1 with refusals R1–R13 (R12:
the host layer; R13: dpkg still broken after the repair). **Changed from the draft** *(decided by CC unattended —
operator may reverse)*: one sudoers entry (`--plan <file>`); the plan's `mode` field (`inventory` / `apply` /
`health`) replaces a separate `--repair-only` (the repair runs first on every apply); Python, because a JSON plan
cannot be parsed safely in sh. The box's own customer guest is "the guest that binds `/mnt/felhom-drives`" (R10).
- **The leg** (agent `internal/osupdate`): after a SUCCESSFUL primary whole-guest backup, inside the backup's
goroutine before the host-wide heavy-op gate is released (so never beside a backup or a restore-test, C10), 90 s
after the backup, at most once per 20 h. *Decided by CC unattended — operator may reverse:* the 90 s settle and the
20 h gap; the controller's own self-update (04:30) is not detected — the 5-minute health wait absorbs a restart.
- **The health rule** (`HealthVerdict`, pinned): docker answers, the guest resolves `deb.debian.org`, the controller's
health check is `healthy`, and every container running at the START of the leg runs again (healthy if it was). The
baseline merges the inventory's reading with the apply's own — found live: an app stopped between them escaped the
first rule. Wait 5 min, poll 15 s.
- **Rings and the switch** (hub, per box): ring 0 = demo-hp + demo-felhom, everything else ring 1; switch ON by
default; OFF → the box reports, installs nothing. A box with no `os_update` block (older hub) = ring 1, ON, no
release.
- **The approval rule (the ruled "1–2 day wait")**: every Debian / Debian-Security package=version that ALL ring-0
boxes having it agree on; approved when, since that set was first seen, **24 h** passed with every ring-0 report
healthy and every ring-0 box completed **1** post-backup night run (`OS_APPROVE_AFTER`, `OS_APPROVE_NIGHTS`; an
override is logged as a TEST configuration). The approval time is the snapshot.debian.org timestamp (decision 79).
"Approve now" is an operator event. *The 24 h / 1 night numbers: decided by CC unattended within the ruled 1–2 days.*
- **The household's line**: hub event `os_update_applied` (info: on the household's hub timeline, never mailed;
hu/en in the bundle). There is no surface on the box itself (R-844).
- **Measured live**: ring 0 — 53 Debian packages on each demo box (18.7–31.7 s inside the wrapper; 174–226 s for the
whole leg incl. the inventory), healthy, 6 Docker updates "not covered"; approval — with a 2-minute TEST wait, a
272-package release approved automatically, then the ruled values restored; ring 1 — demo-felhom installed exactly
the 3 approved versions it lacked (269 already current) and left the newer Docker packages alone; a deliberately
stopped app → `health_failed` after 5 min, the operator mailed, the household line recorded.
- **Not delivered to existing boxes by the product**: the wrapper and the sudoers line reach a box only through the
installer; the signed agent update replaces the binary only (R-840). The demo boxes got them BY HAND.
### 8.2 Step 3 as BUILT (2026-10-04) `[FACT]`
Evidence: `audits/os-host-lane-2026-10-04/` (parts A–G).
- **Where it runs:** appliances only. The proof is the ROOT-owned install record `/var/lib/felhom-install/state.json`
`mode: appliance` (written by the installer as root); the agent-writable `agent.json` `deployment_mode` is not
trusted for this. A BYO host gets no host step (wrapper refusal **R12**, now lifted only for lane fast / layer host
on an appliance). *Decided by CC unattended — operator may reverse.*
- **What:** origin `Debian` / `Debian-Security` only, and never a kernel, boot or firmware package (name pattern
`HOST_SLOW_RE` — `linux-*`, `proxmox-kernel*`, `pve-kernel*`, `pve-firmware`, `firmware-*`, `grub*`, `shim*`,
`systemd-boot*`, `*-microcode`, `efibootmgr`; refusal **R14**). The hub leaves the same names out of the host
candidate.
- **When:** in the same leg, after the guest step, under the same heavy-op gate. A failed or unhealthy guest step
skips the host step.
- **The host health rule** (`HostHealthVerdict`, pinned in `internal/osupdate`): `felhom-agent`, `pveproxy`,
`pvedaemon`, `pvestatd` and `pve-cluster` are `active`; the customer guest runs; the guest health rule (§8.1)
passes; the tunnel is `running` — an `unknown` tunnel does not fail it, `not_running` does. Same 5-minute wait.
- **Separate approved sets:** host and guest releases are separate (`os-host-…`, `os-guest-…`), each by the same rule
(24 h, 1 night of THAT layer, every ring-0 box). Ring 1 receives `host_release` beside `release`.
- **Reboot needed:** reported when PID 1 or `lxc-start` maps a replaced file, with the date of the first scanned
report that said so. **Never reboots.** The host is scanned on every pass, so a reboot clears it (v0.141.1 — v0.141.0
hid `lxc-start` and never cleared, R-846).
- **Undo:** by hand, `runbooks/os-updates-host-undo.md` (proved). No automatic undo.
- **Speed (R-845):** one call per layer; both steps with nothing to install 23–32 s; a 108-package host pass 70 s.
- **Measured live:** demo-felhom ring 0 installed 108 Debian host packages, all Debian origin (checked against apt),
healthy; a 605-package host release approved (TEST wait 2 min / 0 nights, logged, reverted to 24 h + 1 night);
demo-felhom as ring 1 installed exactly the one version it lacked (604 already current).
### 8.3 Step 4 as BUILT (2026-10-04) `[FACT]`
- **The fleet view** (`GET /os/fleet`, operator): one line per box — ring, switch, the tunnel, and per layer: the
release, the last outcome, the last successful leg, pending, not covered, restart needed, reboot needed since, the
wrapper's own seconds.
- **The four alarms** (`08` §6.3), operator-only, hourly, at most weekly while true: no successful OS leg for
**7 days** while the switch is ON (naming the likely reason); reboot needed for **14 days**; ring 0 approved nothing
for **7 days** while it has pending fast-lane updates; not-covered fast-lane packages for **14 days**. The four
numbers are configuration (`OS_ALARM_*`). *Decided by CC unattended — operator may reverse:* 7 days = a week of
missed nights is past any normal hiccup (a box off for a weekend does not alarm); 14 days for reboot and coverage =
two weekly golden cycles, both need a person anyway.
- **The tunnel** (R-841): `running` / `not_running` / `unknown`; `tunnel_down` after two `not_running` reports.
### 8.4 The night's fixes as BUILT (2026-10-05) `[FACT]`
Agent v0.144.0 + v0.144.1 (wrapper and agent; `09` decisions 106–108). Evidence `audits/night-fixes-2026-10-05/`.
- **R8 measures the real download** (R-865): `--print-uris` without `-s`; the installed wrapper read 12 802 456 B for 13
pending guest upgrades on demo-hp (it read 0 B before). `--print-uris` alone downloads nothing (measured on 9202).
- **A killed pass still reports** (R-868): the wrapper writes `report-<run>-<layer>-apply.json` beside the plan before it
prints the report, and it survives a dead reader (a killed agent closes the stderr pipe — v0.144.0 died there,
measured); the agent sends a kept copy at start and every 5 minutes, then deletes it. Live: the A5 shape on demo-hp →
ONE `applied` report (13 packages) reached the hub 5 minutes later. A pass lock (flock) keeps the sender off a running
pass, also across the daemon and a selftest.
- **The debug pass with the hub away** (R-866): the daemon saves the hub's block (`os-update-block.json`); the selftest
uses it when the hub is unreachable and says `block=SAVED(…)` in its header. The daemon itself does not load the saved
block at start (decision 107).
- **A power cut in the middle of an update (night A1, run by day on demo-hp, decision 101):** crash at 06:13:55 UTC while
dpkg ran; the crash guard brought the host back in 37 s (1 unclean boot, still armed); every app healthy again within
4 minutes; dpkg `--audit` clean, the 13 packages split 1 new / 12 old, all `ii`. **The next pass FAILED** (`E: dpkg was
interrupted, you must manually run 'dpkg --configure -a'`): the repair step runs only on a non-clean `--audit`, and a
crash can leave only dpkg's update journal behind — **R-876 (P2, open)**. By hand: `dpkg --configure -a` in the guest
(runbook `crash-guard.md`), then the next pass installed the 12. Operator mail `os_update_failed` (true); no
household mail; the household's timeline showed "Controller elindult".
### 8.5 The self-repair after a power cut as BUILT (2026-10-05, agent v0.145.0) `[FACT]`
- **R-876 fixed.** The wrapper reads `dpkg --audit` AND dpkg's update journal (`/var/lib/dpkg/updates/`) in ONE call
and repairs when either shows something; as a belt, when apt itself says "dpkg was interrupted", it repairs and
retries once (`09` decision 118). A clean pass still costs one call (R-845's speed, pinned by a test).
- **Proven live** (operator's go, demo-hp, `audits/catchup-2026-10-05/partD/`): 13 guest packages rolled back, crash
at 07:56:03 UTC during dpkg's unpack, back by itself (new boot 07:56:41, guard armed, 1 unclean boot in the window);
at boot `--audit` clean, journal 1 file, the same shape as §8.4. **The next pass, with nobody touching the box:**
`REPAIR configured=0 journal=1`, then `DONE rc=0 upgraded=12`, healthy; the guest's package list equals the one
before the rollback. No mail (the pass did not fail); the household's timeline: "Controller elindult".
- **R-874.** The agent's restore-test due-check runs 30 minutes after start (then every interval), so a box with short
power-on sessions is restore-tested (`09` decision 117).
## 9. Where the rest lives
- The finding: **R-812** (`backlog/OPEN-ITEMS.md`). The intention: **R-808** (`backlog/ROADMAP.md`).
- Related: **R-604** (a per-customer floor hides a box from global raises; the same risk applies to an
OS-release floor), **R-530** (agents update only by a signed job per box).
- App updates: `09`. Backups and the night chain: `07` §6.1. The agent's permissions: `03` §3.