diff --git a/CONTEXT.md b/CONTEXT.md index 5d0064d4..623c6893 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -16,6 +16,11 @@ > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +> **Rulings 2026-10-04 (day) — recorded before the work (off-site close + OS update spike).** `09` §3 decisions **75** +> (the off-site topic closes first, one brief), **76** (the OS fast lane follows an approved list with a 1–2 day wait; +> `unattended-upgrades` with no wait rejected) and **77** (`architecture/11-os-updates.md` added unchanged, NOT RATIFIED, +> then corrected by measurement). + > **Rulings 2026-10-04 (morning) — recorded before the work.** `09` §3 decisions **71** (ep0's DooPlex copy keeps 8 weekly copies; a third place later, R-832), **72** (tester-1's unpinned keys removed via the registrar), **73** (the transcript-exposed Hetzner storage token is not rotated now; R-831), and **74** (the set-aside deletion is the hub's, after a 7-day wait — from the brief's Part E). > **Rulings 2026-10-03 (afternoon) — recorded before the work (off-site lock build).** `09` §3 decisions **68** (the box prunes only inside a weekly hub-opened window; option 2 rejected; no box-side retention in the interim), **69** (the hub is the key registrar; the box never receives the sub-account password; the hub stores it encrypted with a key outside the database; daily `authorized_keys` check) and **70** (nightly PBS pull-sync of ep0's `felhom-offsite` to DooPlex; the fence opens for that brief's Part F acts only). diff --git a/documentation/README.md b/documentation/README.md index 954f4d79..78af5fb1 100644 --- a/documentation/README.md +++ b/documentation/README.md @@ -30,6 +30,7 @@ The operator-tier agent and the Proxmox platform. - [`architecture/04-control-plane-authorization.md`](architecture/04-control-plane-authorization.md) — signing, escrow, authz - [`architecture/02-controller-module-map.md`](architecture/02-controller-module-map.md) — **historical** v0.33 planning map; the live map is [`controller/module-map.md`](controller/module-map.md) - [`proxmox-platform.md`](proxmox-platform.md) — Proxmox platform reference +- [`architecture/11-os-updates.md`](architecture/11-os-updates.md) — operating-system updates: host, guest, Docker engine (**NOT RATIFIED**, 2026-10-04) ### Hub (operator backend) — `architecture/05` - [`architecture/05-hub-architecture.md`](architecture/05-hub-architecture.md) — hub architecture (v0.11.0) diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 2cb9a6bc..22468f26 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -232,7 +232,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis | **The hub reports LOSS OF VISIBILITY into either off-site store (not just how full it is)** | hub **v0.106.0** (R-339) | **IMPLEMENTED — deliberately NOT proven-live** | Both box checkers count consecutive failed fetch windows and emit `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default 3 windows (≈30–45 min), each with a paired `*_recovered` all-clear routed via `recoveredPairedDownTypes` — required because the recoveries are severity `info`, which `severityNotifies` drops. Scopes stay customer-less (`pbsdr-box` / `pool-box`) → operator channel only. Fill logic untouched: a degraded read still drives no band transition. Evidence: `internal/monitor/box_reachability_test.go` + the cross-package wiring test in `internal/notify/`, which asserts an actual operator mail rather than a map entry. **Filed BECAUSE of a measured gap**, not a hypothesis: the 2026-08-18 ep0 outage ran 9 h 37 m with the hub silent | **The gap that remains is R-340**, and it is not small: the ep0 read is the `usage` op, which rides the LOCAL API daemon — the daemon that incident explicitly cleared — so this check would have shown GREEN for that entire outage. It closes "ep0 is unreachable as a host"; it does not close what actually happened. **No live or constructed outage has exercised the emit path**, and one cannot be manufactured against ep0 (Tier 2, protected) | | Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | | | Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | | -| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | — | **MISSING** | `felhom-host-install.sh:2133-2136` ("No upgrades are run") | Nothing runs them after install → finding R-812, intention R-808 (added 2026-10-03) | +| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | — | **MISSING** | `felhom-host-install.sh:2133-2136` ("No upgrades are run") | Nothing runs them after install → finding R-812, intention R-808 (added 2026-10-03). Design: `architecture/11-os-updates.md` (NOT RATIFIED, 2026-10-04) | | **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** | | **A deletion of a customer's off-site history is NOTICED within a day** | hub **v0.111.0** (R-431) | **IMPLEMENTED — not yet PROVEN-LIVE** | 09-01 | `hub/internal/monitor/offsite.go` — third signal beside FILL and STALENESS. **On the hub deliberately:** a detector on the box is one the deletion can silence. Alarms when the reported count falls by more than HALF and by at least 5, guarded by `StatsKnown` (R-331), the declared `State` (R-204) and run success (R-100). **Threshold reasoned, not invented:** over 12 898 reports every decrease lands on ZERO and predates `stats_known`; in the 380-report `stats_known` window there are none. **ACCEPTANCE: 9 009 real points replayed → ZERO alarms** (`offsite_r431_test.go`, fixture committed). **What PROVEN-LIVE would need and this does NOT have:** a real drop observed on a live box producing a real mail — the live firing done at ship time was driven through the hub's own path with synthetic counts, which is an end-to-end delivery proof, not a proof that a genuine deletion is caught. | diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index 45d91c17..f492f25d 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -742,6 +742,17 @@ its length, and both fixes cost something the household would notice — operato time; the hub deletes only `.orphaned-<…>`, never the live repository; every request, cancel and deletion is an operator event. *Operator brief 2026-10-04 (Part E).* Built hub v0.128.0 + controller v0.290.0. +### 2026-10-04 (day) — three operator rulings (recorded before the work) + +75. **The off-site topic is closed before the operating-system update work starts.** One brief carries both; the + backup part runs first. *Operator ruling 2026-10-04 (morning).* +76. **The OS update fast lane follows an approved list, with a 1–2 day wait** (`11` §3). Every update runs on the demo + boxes first; other boxes install only the exact versions that ran well there. **Rejected:** Debian's + `unattended-upgrades` with no wait — a bad update would reach every customer at the same time. *Operator ruling + 2026-10-04 (morning) (R-808, R-812).* +77. **`architecture/11-os-updates.md` is added as the reviewer wrote it, NOT RATIFIED**, committed unchanged first and + then corrected by the spike's measurements, each correction named. *Operator ruling 2026-10-04 (morning).* + ### 2026-09-30 (day) — operator notes, recorded before the work - **The day brief runs by day.** Every backup and automatic-update test is started by hand — the night chain's debug diff --git a/documentation/architecture/11-os-updates.md b/documentation/architecture/11-os-updates.md new file mode 100644 index 00000000..ec017b23 --- /dev/null +++ b/documentation/architecture/11-os-updates.md @@ -0,0 +1,275 @@ +# 11 — Operating-system updates: the host, the guest and the Docker engine + +> | | | +> |---|---| +> | **Status** | **NOT RATIFIED — a PROPOSAL with one operator ruling.** Ratification is Viktor's review, not an editor's. | +> | **Written** | 2026-10-04, by the reviewer (project Claude), before any spike. | +> | **Verified against** | felhom.eu `d07a1a9` · felhom-controller `99a1497` (v0.290.0) · felhom-agent `d766666` (v0.138.0) · hub v0.128.0 | +> | **Freshness** | **CURRENT** as of 2026-10-04. The spike `TASK-backup-close-and-os-updates-spike-2026-10-04` adds measurements here as `[FACT]` and corrects every claim it disproves. Mark this file STALE when it falls behind what the product does. | +> +> **How to read this document.** Each statement has a label: +> +> - **[FACT]**: an observed property, with a `file:line`, a register row or an audit path. +> - **[RULED]**: an operator decision, with its date. +> - **[PROPOSAL]**: the reviewer's design. It is not decided and not built. +> - **OPEN**: a question that nobody has answered yet. The spike measures it. Nobody guesses it. +> +> **Why this file exists.** No architecture document covered operating-system updates. `00` §G marks the +> capability MISSING (added 2026-10-03). The finding is **R-812**. The roadmap intention is **R-808**. +> **The register carries the work. This file carries the reasoning. The source is the truth.** + +--- + +## 0. In plain language + +A box runs three layers that we install and never update: the Proxmox host, the small Debian system +inside the guest, and the Docker engine. App images are updated (see `09`). The layer under them is not. +A box lives in a home for years, so this is a security gap. + +The proposal has two lanes. **The fast lane** applies Debian's security fixes automatically, but only +the exact versions that already ran well on the demo boxes for 1–2 days. **The slow lane** covers the +kernel, Proxmox and Docker. These cause restarts or big changes, so a person approves them one version at +a time, and they run only at night. The tested versions are recorded automatically from what the demo +boxes installed. Nobody keeps a hand-written list. + +--- + +## 1. Scope + +**In scope.** +- The Proxmox host on an **appliance** install: the Debian 13 base, the Proxmox packages, and the kernel. +- The guest (the customer LXC): its Debian 13 packages. +- The Docker engine inside the guest (`docker-ce`, `docker-ce-cli`, `containerd.io`). +- How the household and the operator are told, and how a failed update is undone. + +**Out of scope.** +- **App images.** `09` covers them (the ladder, the monthly same-tag re-test). +- **The controller and agent binaries.** Their own self-update covers them (`03`, the self-update section; `09` R-608). +- **DooPlex and ep0.** The operator updates them by hand (`runbooks/offsite-endpoint.md`). +- **A BYO host** (the owner brought their own Proxmox). `[FACT]` The installer leaves its repositories + alone: *"apt repo alignment skipped (byo — the owner manages repos)"* (`scripts/felhom-host-install.sh`, + `align_apt_repos`). `[PROPOSAL]` On a BYO box we update the guest and Docker only, never the host. +- **The Proxmox MAJOR upgrade** (PVE 9 → 10). It is a later step of its own, drilled first (R-808 item 5). + +--- + +## 2. What a box runs, as measured + +| Layer | What it is | Source of packages | Evidence | +|---|---|---|---| +| Host | Proxmox VE 9.2 on Debian 13, LVM-thin | Debian mirrors + `pve-no-subscription` (enterprise repo switched off on appliance installs) | `[FACT]` `01` §2; `felhom-host-install.sh` `align_apt_repos` (~L2130–2136) | +| Host kernel | The only kernel on the box. The guest has none. | `pve-no-subscription` (`proxmox-kernel-*`) | `[FACT]` LXC shares the host kernel (`01` §2: "one LXC/kernel/Docker daemon") | +| Guest | Unprivileged LXC, `nesting=1,keyctl=1`, from `debian-13-standard_13.1-2` | Debian mirrors | `[FACT]` `felhom-agent/configs/build-golden.sh:67,101-103` | +| Docker engine | `docker-ce`, `docker-ce-cli`, `containerd.io`, installed at golden BAKE time | `download.docker.com/linux/debian trixie stable` | `[FACT]` `build-golden.sh:108-125` | +| Docker settings | `containerd-snapshotter: false`, json-file log caps. **No `live-restore`.** | baked `daemon.json` | `[FACT]` `build-golden.sh:138-144` | +| Controller | A container in the guest. A Docker engine restart restarts it too. | Felhom registry | `[FACT]` `03` §1 | + +**[FACT] Nothing updates any of these layers today** (R-812, searched 2026-10-03). The installer +says *"No upgrades are run — repo alignment only"* (`felhom-host-install.sh` ~L2135). A fresh install +gets the Docker engine that was current when its golden was baked, and keeps it. + +**[FACT] The agent may not run `apt` today, with one exception.** Its sudoers allowlist +(`felhom-agent/configs/felhom-agent.sudoers`) holds one apt line: `apt-get install -y -q dnsmasq`. +`apt` as root runs package scripts as root, so a broad `apt` grant is a full root grant. §5.4 +proposes how to avoid that. + +--- + +## 3. Operator rulings + +**2026-10-03.** R-808 is on the roadmap at P2: *"every box receives operating-system security +patches on a schedule, and a failed update is undone."* + +**[RULED] 2026-10-04: the fast lane follows an approved list, with a 1–2 day wait.** Every update +runs on the demo boxes first. The other boxes install only the exact versions that the demo boxes ran +without trouble. The exact wait is set when the feature is built. **Rejected:** Debian's own +`unattended-upgrades` with no wait. It is simpler and common, but a bad update would reach every +customer at the same time. + +**[RULED] 2026-10-04: the off-site backup topic is closed first.** The same brief carries both +topics, and the backup part runs first. + +The two-lane split (§5.2) is the reviewer's proposal. The operator's ruling above assumes it, but he +has not ruled on it as such. + +--- + +## 4. The constraints that shape the design + +1. **Unattended.** Nobody is at the box. Nobody answers a question that `apt` asks. +2. **No screen.** If the box does not boot, the household sees only that nothing works. +3. **One kernel for everything.** A host kernel update needs a host reboot. A reboot stops every app. +4. **The controller lives in Docker.** A Docker engine update restarts the controller in the middle of + its own work. So the controller cannot drive a Docker update. The agent, on the host, must. +5. **The host has no whole-system backup.** The guest has three tiers (`07`). The host has none. + A host update can only be undone by installing the previous version again. +6. **Every update must already have run on a box we own.** This is the lesson of the update arc + (`09` §3 decision 13: "the test decides"). +7. **Few packages.** Each extra package is one more thing to update and break. The guest is the + Debian standard template plus Docker. Keep it that way. + +--- + +## 5. The proposed shape `[PROPOSAL]` + +### 5.1 Two rings + +- **Ring 0:** demo-felhom (N100) and demo-hp. They are disposable (`runbooks/target-selection.md`), and + they have different hardware. They take every update first. +- **Ring 1:** every other box. It takes only what ring 0 approved. + +### 5.2 Two lanes + +| | Fast lane | Slow lane | +|---|---|---| +| What | Debian packages on host and guest, from `trixie-security` and the stable point releases, **except** the slow-lane list | Kernel (`proxmox-kernel-*`), Proxmox (`pve-*`, `proxmox-*`, `lxc-pve`, `qemu-server`, …), Docker (`docker-ce*`, `containerd.io`) | +| Restart | A service restart at most. No reboot. | Kernel: a host reboot. Docker: every container restarts. Proxmox: its services restart. | +| Approval | Automatic. A version is approved when ring 0 ran it and stayed healthy for the wait (1–2 days). | A person approves one version at a time, like a controller floor. | +| Urgent fix | The operator can approve a version on the same day once ring 0 has run it. | The same. | +| When | The night window (§5.5) | The night window. A kernel reboot only on a night the operator scheduled. | + +Which packages count as slow lane is a list the spike checks (OPEN Q7). A Debian package that restarts +something big (for example `systemd`, `libc6`, `openssh-server`) may belong in the slow lane too. + +### 5.3 The approved list (the "tested versions" record) + +Nobody writes the list by hand. It fills itself: + +1. A ring-0 box updates. Afterwards it reports to the hub the exact `package=version` it installed, + per layer (host, guest), with each package's origin (`Debian-Security`, `Debian`, `Proxmox`, + `Docker`). +2. The box then reports health for the wait period: the agent, the controller, every app's health, + and the guest's network. +3. When the wait passes with ring 0 healthy, the hub marks that set **approved**. It is one record: + an **OS release**, with an id and a date. The hub stores it. The register does not. +4. A ring-1 box asks the hub for the newest approved OS release. For each package it has installed, + if the approved version is newer, it installs that **exact** version. It never installs a version + newer than the approved one. +5. Each box reports which OS release it runs, how many updates are waiting, whether it needs a reboot, + and which packages it has that **no approved list covers** (see §6, edge case 9). + +**OPEN Q1 is the weak point.** The design works only if a box can still download the approved version +a week later. Debian's main and security archives keep only the newest version of each package. If +Debian publishes a newer fix between approval and install, the approved version is gone. Options: +the box waits for the next approval; or the box uses `snapshot.debian.org` (Debian's own dated archive, +still signed by Debian); or Felhom runs a package cache. The spike measures how often this happens. + +### 5.4 Who runs it, and with what permission + +- **The agent runs every OS update**, for the host and for the guest (`pct exec`). The controller does + not, because of constraint 4. +- **The agent gets no general `apt` permission.** Like `felhom-selfupdate-guarded` (`03`, the self-update section), + a small root-owned wrapper does the work. It accepts only an approved list and refuses everything else: + - no package removal; + - no downgrade, except the undo of §5.6, which an operator job signs; + - no package that the box does not already have, unless the approved list records it as a dependency + that the same update pulled in on ring 0 (a kernel update installs a NEW package name each time, + for example `proxmox-kernel-6.x.y-z-pve-signed`, pulled by `proxmox-default-kernel`); + - no package source other than the ones the installer set up; + - non-interactive, and it always keeps the existing config file (`--force-confold`) and reports the + conflict. +- **Package signatures stay the publishers'** (Debian, Proxmox, Docker). The hub sends only names and + versions. A broken-into hub can choose an older version or no version. It cannot make a box install a + package that the publisher did not sign. + +### 5.5 When + +Inside the household's night window, after the backups: + +``` +W DB dump +W+60m Tier 2 +W+105m off-site → app updates (until W+5h at most) +[W+2h, W+6h) whole-guest backup (agent) +after it OS updates — guest first, then host +``` + +The OS leg starts **after the whole-guest backup has finished**, so the guest's newest full copy is +minutes old. This is the opposite order to app updates, which run before the whole-guest backup +(`07` §6.1, `09` decision 11). The reason: the whole-guest backup is the guest's undo. + +**OPEN Q8:** what else runs then. The controller's self-update (default 04:30, and after any hub report +when a floor is above it, R-608), the agent's self-update, the restore-tests, and PBS jobs. The OS leg +must never overlap a backup, a restore-test or a self-update. + +### 5.6 How a failed update is undone + +"Rollback" is not used (`09` §4). The shapes: + +| Layer | Undo | Limit | +|---|---|---| +| Guest packages | Restore the guest snapshot taken just before the update | It also undoes app data written after the snapshot. Use it only inside the health window, before apps have written much. After that, install the previous version (needs OPEN Q1). | +| Docker engine | Install the previous version (Docker's repository keeps old versions — OPEN Q2) | Every container restarts again | +| Host packages | Install the previous version | Only if the source still has it (OPEN Q1/Q2) | +| Host kernel | Boot the previous kernel. Proxmox can boot a new kernel **once** (`proxmox-boot-tool kernel pin --next-boot`). If that boot fails, the next boot uses the old kernel again. Make the new kernel permanent only after a healthy boot. | If the new kernel hangs, someone must switch the box off and on. The spike checks whether a hardware watchdog can do that (OPEN Q4). | + +### 5.7 Telling people + +- **Operator:** a hub event for each update and each failure; a fleet view showing each box's OS release, + how far behind it is, and whether it needs a reboot. A box that is more than N days behind the newest + approved release raises an alarm (`08`). +- **Household:** one line on the timeline in both languages, informal voice: what was updated and + whether the box restarted. Telling households in advance that the box may restart at night is a + **promise to users**. That is the operator's decision when the slow lane is built. + +--- + +## 6. Risks and edge cases + +| # | What can go wrong | What the design does | +|---|---|---| +| 1 | A new kernel does not boot | Boot it once (`--next-boot`); the old kernel stays the default. A hard hang needs a power cycle (OPEN Q4). | +| 2 | Power cut during an update: `dpkg` is half done | The wrapper runs `dpkg --configure -a` and `apt-get -f install` first, every time, and reports what it repaired (spike measures, Q5). | +| 3 | An update asks a question (changed config file, service restart prompt) | Non-interactive, keep the old config, report the conflict. | +| 4 | A Docker update stops every app, and the controller | Stop the apps cleanly first, like before a backup. Measure whether `live-restore` keeps containers running (Q3). The agent drives it. | +| 5 | The approved version is no longer downloadable | OPEN Q1. Until it is answered, the box waits for the next approval and reports it. | +| 6 | An urgent security hole | The operator approves the same day once ring 0 has run it. | +| 7 | A box was off for months | It catches up through the newest approved release. It does not install every release in between. Packages are not like app data; one `apt` step is enough. Kernel and Docker still go one approved step at a time. | +| 8 | The system disk is full | Check free space before downloading; clean the package cache after; refuse and report. | +| 9 | A customer box has a package that ring 0 does not have (other hardware: firmware, CPU microcode, NIC drivers) | It is never approved, so it is never updated. The box reports it as "not covered", and the hub raises it. Fix: add matching hardware to ring 0, or approve it by hand. | +| 10 | The package source is down, or its signing key changes (Docker has done this) | The update fails cleanly and the box reports it. A key change is a slow-lane act for a person. | +| 11 | A BYO host | Only the guest and Docker are updated (§1). | +| 12 | Two boxes on ring 0 is a small sample | Accepted for now. When there are customers, the first tester boxes can become a second ring. | +| 13 | A broken-into hub sends a harmful list | The wrapper refuses removals, downgrades and new packages, and the publishers' signatures still apply (§5.4). | +| 14 | `cloudflared` on the host | Not covered by this file yet. How it is installed and updated is OPEN Q9. It faces the internet, so it matters. | +| 15 | The no-subscription Proxmox repository gets less testing than the enterprise one | Ring 0 is our test. The enterprise repository costs a yearly fee per box. That is a **money** decision for the operator, later (Q10). | + +--- + +## 7. Open questions the spike must answer + +| Q | Question | How to answer | +|---|---|---| +| Q1 | Can a box install an exact Debian version one week after approval? How often is it already superseded? | `apt-cache madison` on host and guest for security packages. Debian's security announcement history. Whether `snapshot.debian.org` is reachable and fast enough. | +| Q2 | Do the Proxmox and Docker sources keep older versions? | `apt-cache madison` for `pve-manager`, `proxmox-kernel-*`, `docker-ce`, `containerd.io`. | +| Q3 | What does a Docker engine update do to running containers, with and without `live-restore`? | On scratch guest 9202: step from one pinned version to the next; time the downtime. | +| Q4 | Does `proxmox-boot-tool kernel pin --next-boot` work on both demo hosts' boot setups (UEFI or legacy, GRUB or systemd-boot)? Does either box have a hardware watchdog? | On demo-hp, with the operator's go before each reboot. | +| Q5 | What does an interrupted `apt` run leave, and does the repair recover it? | Kill a run on 9202, after a snapshot. | +| Q6 | How far behind are the demo boxes today, per layer and per source? How long does catching up take, and what restarts? | `apt-get -s upgrade` (read only) first; then a real run on 9202 and on one demo host. | +| Q7 | Which Debian packages restart something big? | `needrestart` in list mode after a run. | +| Q8 | What else runs in the night window, and where does the OS leg fit? | Read the timers on host and guest. | +| Q9 | How is `cloudflared` installed and updated on the host? | Read the installer and the agent. | +| Q10 | What does the Proxmox enterprise repository cost per box per year, and what does it add? | The publisher's price page. Record only. The operator decides later. | + +--- + +## 8. Build order `[PROPOSAL]` + +Each step returns to the operator for go or no-go. + +1. **Spike** (measure Q1–Q10; no product code). +2. **Guest Debian, fast lane.** Lowest risk: a snapshot undo exists. +3. **Host Debian, fast lane** (no kernel, no Proxmox packages). +4. **Fleet view and alarms** (§5.7). +5. **Slow lane: Docker engine.** +6. **Slow lane: host kernel and Proxmox packages, with the reboot.** +7. **Later:** the Proxmox major upgrade (PVE 9 → 10), drilled on ring 0 first. + +--- + +## 9. Where the rest lives + +- The finding: **R-812** (`backlog/OPEN-ITEMS.md`). The intention: **R-808** (`backlog/ROADMAP.md`). +- Related: **R-604** (a per-customer floor hides a box from global raises; the same risk applies to an + OS-release floor), **R-530** (agents update only by a signed job per box). +- App updates: `09`. Backups and the night chain: `07` §6.1. The agent's permissions: `03` §3.