docs: rulings 2026-10-04 (decisions 75-77); add 11-os-updates.md verbatim (NOT RATIFIED)
gates / gates (push) Successful in 29s
gates / gates (push) Successful in 29s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -232,7 +232,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
||||
| **The hub reports LOSS OF VISIBILITY into either off-site store (not just how full it is)** | hub **v0.106.0** (R-339) | **IMPLEMENTED — deliberately NOT proven-live** | Both box checkers count consecutive failed fetch windows and emit `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default 3 windows (≈30–45 min), each with a paired `*_recovered` all-clear routed via `recoveredPairedDownTypes` — required because the recoveries are severity `info`, which `severityNotifies` drops. Scopes stay customer-less (`pbsdr-box` / `pool-box`) → operator channel only. Fill logic untouched: a degraded read still drives no band transition. Evidence: `internal/monitor/box_reachability_test.go` + the cross-package wiring test in `internal/notify/`, which asserts an actual operator mail rather than a map entry. **Filed BECAUSE of a measured gap**, not a hypothesis: the 2026-08-18 ep0 outage ran 9 h 37 m with the hub silent | **The gap that remains is R-340**, and it is not small: the ep0 read is the `usage` op, which rides the LOCAL API daemon — the daemon that incident explicitly cleared — so this check would have shown GREEN for that entire outage. It closes "ep0 is unreachable as a host"; it does not close what actually happened. **No live or constructed outage has exercised the emit path**, and one cannot be manufactured against ep0 (Tier 2, protected) |
|
||||
| Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | |
|
||||
| Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | |
|
||||
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | — | **MISSING** | `felhom-host-install.sh:2133-2136` ("No upgrades are run") | Nothing runs them after install → finding R-812, intention R-808 (added 2026-10-03) |
|
||||
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | — | **MISSING** | `felhom-host-install.sh:2133-2136` ("No upgrades are run") | Nothing runs them after install → finding R-812, intention R-808 (added 2026-10-03). Design: `architecture/11-os-updates.md` (NOT RATIFIED, 2026-10-04) |
|
||||
| **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** |
|
||||
| **A deletion of a customer's off-site history is NOTICED within a day** | hub **v0.111.0** (R-431) | **IMPLEMENTED — not yet PROVEN-LIVE** | 09-01 | `hub/internal/monitor/offsite.go` — third signal beside FILL and STALENESS. **On the hub deliberately:** a detector on the box is one the deletion can silence. Alarms when the reported count falls by more than HALF and by at least 5, guarded by `StatsKnown` (R-331), the declared `State` (R-204) and run success (R-100). **Threshold reasoned, not invented:** over 12 898 reports every decrease lands on ZERO and predates `stats_known`; in the 380-report `stats_known` window there are none. **ACCEPTANCE: 9 009 real points replayed → ZERO alarms** (`offsite_r431_test.go`, fixture committed). **What PROVEN-LIVE would need and this does NOT have:** a real drop observed on a live box producing a real mail — the live firing done at ship time was driven through the hub's own path with synthetic counts, which is an end-to-end delivery proof, not a proof that a genuine deletion is caught. |
|
||||
|
||||
|
||||
@@ -742,6 +742,17 @@ its length, and both fixes cost something the household would notice — operato
|
||||
time; the hub deletes only `<repo>.orphaned-<…>`, never the live repository; every request, cancel and deletion is
|
||||
an operator event. *Operator brief 2026-10-04 (Part E).* Built hub v0.128.0 + controller v0.290.0.
|
||||
|
||||
### 2026-10-04 (day) — three operator rulings (recorded before the work)
|
||||
|
||||
75. **The off-site topic is closed before the operating-system update work starts.** One brief carries both; the
|
||||
backup part runs first. *Operator ruling 2026-10-04 (morning).*
|
||||
76. **The OS update fast lane follows an approved list, with a 1–2 day wait** (`11` §3). Every update runs on the demo
|
||||
boxes first; other boxes install only the exact versions that ran well there. **Rejected:** Debian's
|
||||
`unattended-upgrades` with no wait — a bad update would reach every customer at the same time. *Operator ruling
|
||||
2026-10-04 (morning) (R-808, R-812).*
|
||||
77. **`architecture/11-os-updates.md` is added as the reviewer wrote it, NOT RATIFIED**, committed unchanged first and
|
||||
then corrected by the spike's measurements, each correction named. *Operator ruling 2026-10-04 (morning).*
|
||||
|
||||
### 2026-09-30 (day) — operator notes, recorded before the work
|
||||
|
||||
- **The day brief runs by day.** Every backup and automatic-update test is started by hand — the night chain's debug
|
||||
|
||||
@@ -0,0 +1,275 @@
|
||||
# 11 — Operating-system updates: the host, the guest and the Docker engine
|
||||
|
||||
> | | |
|
||||
> |---|---|
|
||||
> | **Status** | **NOT RATIFIED — a PROPOSAL with one operator ruling.** Ratification is Viktor's review, not an editor's. |
|
||||
> | **Written** | 2026-10-04, by the reviewer (project Claude), before any spike. |
|
||||
> | **Verified against** | felhom.eu `d07a1a9` · felhom-controller `99a1497` (v0.290.0) · felhom-agent `d766666` (v0.138.0) · hub v0.128.0 |
|
||||
> | **Freshness** | **CURRENT** as of 2026-10-04. The spike `TASK-backup-close-and-os-updates-spike-2026-10-04` adds measurements here as `[FACT]` and corrects every claim it disproves. Mark this file STALE when it falls behind what the product does. |
|
||||
>
|
||||
> **How to read this document.** Each statement has a label:
|
||||
>
|
||||
> - **[FACT]**: an observed property, with a `file:line`, a register row or an audit path.
|
||||
> - **[RULED]**: an operator decision, with its date.
|
||||
> - **[PROPOSAL]**: the reviewer's design. It is not decided and not built.
|
||||
> - **OPEN**: a question that nobody has answered yet. The spike measures it. Nobody guesses it.
|
||||
>
|
||||
> **Why this file exists.** No architecture document covered operating-system updates. `00` §G marks the
|
||||
> capability MISSING (added 2026-10-03). The finding is **R-812**. The roadmap intention is **R-808**.
|
||||
> **The register carries the work. This file carries the reasoning. The source is the truth.**
|
||||
|
||||
---
|
||||
|
||||
## 0. In plain language
|
||||
|
||||
A box runs three layers that we install and never update: the Proxmox host, the small Debian system
|
||||
inside the guest, and the Docker engine. App images are updated (see `09`). The layer under them is not.
|
||||
A box lives in a home for years, so this is a security gap.
|
||||
|
||||
The proposal has two lanes. **The fast lane** applies Debian's security fixes automatically, but only
|
||||
the exact versions that already ran well on the demo boxes for 1–2 days. **The slow lane** covers the
|
||||
kernel, Proxmox and Docker. These cause restarts or big changes, so a person approves them one version at
|
||||
a time, and they run only at night. The tested versions are recorded automatically from what the demo
|
||||
boxes installed. Nobody keeps a hand-written list.
|
||||
|
||||
---
|
||||
|
||||
## 1. Scope
|
||||
|
||||
**In scope.**
|
||||
- The Proxmox host on an **appliance** install: the Debian 13 base, the Proxmox packages, and the kernel.
|
||||
- The guest (the customer LXC): its Debian 13 packages.
|
||||
- The Docker engine inside the guest (`docker-ce`, `docker-ce-cli`, `containerd.io`).
|
||||
- How the household and the operator are told, and how a failed update is undone.
|
||||
|
||||
**Out of scope.**
|
||||
- **App images.** `09` covers them (the ladder, the monthly same-tag re-test).
|
||||
- **The controller and agent binaries.** Their own self-update covers them (`03`, the self-update section; `09` R-608).
|
||||
- **DooPlex and ep0.** The operator updates them by hand (`runbooks/offsite-endpoint.md`).
|
||||
- **A BYO host** (the owner brought their own Proxmox). `[FACT]` The installer leaves its repositories
|
||||
alone: *"apt repo alignment skipped (byo — the owner manages repos)"* (`scripts/felhom-host-install.sh`,
|
||||
`align_apt_repos`). `[PROPOSAL]` On a BYO box we update the guest and Docker only, never the host.
|
||||
- **The Proxmox MAJOR upgrade** (PVE 9 → 10). It is a later step of its own, drilled first (R-808 item 5).
|
||||
|
||||
---
|
||||
|
||||
## 2. What a box runs, as measured
|
||||
|
||||
| Layer | What it is | Source of packages | Evidence |
|
||||
|---|---|---|---|
|
||||
| Host | Proxmox VE 9.2 on Debian 13, LVM-thin | Debian mirrors + `pve-no-subscription` (enterprise repo switched off on appliance installs) | `[FACT]` `01` §2; `felhom-host-install.sh` `align_apt_repos` (~L2130–2136) |
|
||||
| Host kernel | The only kernel on the box. The guest has none. | `pve-no-subscription` (`proxmox-kernel-*`) | `[FACT]` LXC shares the host kernel (`01` §2: "one LXC/kernel/Docker daemon") |
|
||||
| Guest | Unprivileged LXC, `nesting=1,keyctl=1`, from `debian-13-standard_13.1-2` | Debian mirrors | `[FACT]` `felhom-agent/configs/build-golden.sh:67,101-103` |
|
||||
| Docker engine | `docker-ce`, `docker-ce-cli`, `containerd.io`, installed at golden BAKE time | `download.docker.com/linux/debian trixie stable` | `[FACT]` `build-golden.sh:108-125` |
|
||||
| Docker settings | `containerd-snapshotter: false`, json-file log caps. **No `live-restore`.** | baked `daemon.json` | `[FACT]` `build-golden.sh:138-144` |
|
||||
| Controller | A container in the guest. A Docker engine restart restarts it too. | Felhom registry | `[FACT]` `03` §1 |
|
||||
|
||||
**[FACT] Nothing updates any of these layers today** (R-812, searched 2026-10-03). The installer
|
||||
says *"No upgrades are run — repo alignment only"* (`felhom-host-install.sh` ~L2135). A fresh install
|
||||
gets the Docker engine that was current when its golden was baked, and keeps it.
|
||||
|
||||
**[FACT] The agent may not run `apt` today, with one exception.** Its sudoers allowlist
|
||||
(`felhom-agent/configs/felhom-agent.sudoers`) holds one apt line: `apt-get install -y -q dnsmasq`.
|
||||
`apt` as root runs package scripts as root, so a broad `apt` grant is a full root grant. §5.4
|
||||
proposes how to avoid that.
|
||||
|
||||
---
|
||||
|
||||
## 3. Operator rulings
|
||||
|
||||
**2026-10-03.** R-808 is on the roadmap at P2: *"every box receives operating-system security
|
||||
patches on a schedule, and a failed update is undone."*
|
||||
|
||||
**[RULED] 2026-10-04: the fast lane follows an approved list, with a 1–2 day wait.** Every update
|
||||
runs on the demo boxes first. The other boxes install only the exact versions that the demo boxes ran
|
||||
without trouble. The exact wait is set when the feature is built. **Rejected:** Debian's own
|
||||
`unattended-upgrades` with no wait. It is simpler and common, but a bad update would reach every
|
||||
customer at the same time.
|
||||
|
||||
**[RULED] 2026-10-04: the off-site backup topic is closed first.** The same brief carries both
|
||||
topics, and the backup part runs first.
|
||||
|
||||
The two-lane split (§5.2) is the reviewer's proposal. The operator's ruling above assumes it, but he
|
||||
has not ruled on it as such.
|
||||
|
||||
---
|
||||
|
||||
## 4. The constraints that shape the design
|
||||
|
||||
1. **Unattended.** Nobody is at the box. Nobody answers a question that `apt` asks.
|
||||
2. **No screen.** If the box does not boot, the household sees only that nothing works.
|
||||
3. **One kernel for everything.** A host kernel update needs a host reboot. A reboot stops every app.
|
||||
4. **The controller lives in Docker.** A Docker engine update restarts the controller in the middle of
|
||||
its own work. So the controller cannot drive a Docker update. The agent, on the host, must.
|
||||
5. **The host has no whole-system backup.** The guest has three tiers (`07`). The host has none.
|
||||
A host update can only be undone by installing the previous version again.
|
||||
6. **Every update must already have run on a box we own.** This is the lesson of the update arc
|
||||
(`09` §3 decision 13: "the test decides").
|
||||
7. **Few packages.** Each extra package is one more thing to update and break. The guest is the
|
||||
Debian standard template plus Docker. Keep it that way.
|
||||
|
||||
---
|
||||
|
||||
## 5. The proposed shape `[PROPOSAL]`
|
||||
|
||||
### 5.1 Two rings
|
||||
|
||||
- **Ring 0:** demo-felhom (N100) and demo-hp. They are disposable (`runbooks/target-selection.md`), and
|
||||
they have different hardware. They take every update first.
|
||||
- **Ring 1:** every other box. It takes only what ring 0 approved.
|
||||
|
||||
### 5.2 Two lanes
|
||||
|
||||
| | Fast lane | Slow lane |
|
||||
|---|---|---|
|
||||
| What | Debian packages on host and guest, from `trixie-security` and the stable point releases, **except** the slow-lane list | Kernel (`proxmox-kernel-*`), Proxmox (`pve-*`, `proxmox-*`, `lxc-pve`, `qemu-server`, …), Docker (`docker-ce*`, `containerd.io`) |
|
||||
| Restart | A service restart at most. No reboot. | Kernel: a host reboot. Docker: every container restarts. Proxmox: its services restart. |
|
||||
| Approval | Automatic. A version is approved when ring 0 ran it and stayed healthy for the wait (1–2 days). | A person approves one version at a time, like a controller floor. |
|
||||
| Urgent fix | The operator can approve a version on the same day once ring 0 has run it. | The same. |
|
||||
| When | The night window (§5.5) | The night window. A kernel reboot only on a night the operator scheduled. |
|
||||
|
||||
Which packages count as slow lane is a list the spike checks (OPEN Q7). A Debian package that restarts
|
||||
something big (for example `systemd`, `libc6`, `openssh-server`) may belong in the slow lane too.
|
||||
|
||||
### 5.3 The approved list (the "tested versions" record)
|
||||
|
||||
Nobody writes the list by hand. It fills itself:
|
||||
|
||||
1. A ring-0 box updates. Afterwards it reports to the hub the exact `package=version` it installed,
|
||||
per layer (host, guest), with each package's origin (`Debian-Security`, `Debian`, `Proxmox`,
|
||||
`Docker`).
|
||||
2. The box then reports health for the wait period: the agent, the controller, every app's health,
|
||||
and the guest's network.
|
||||
3. When the wait passes with ring 0 healthy, the hub marks that set **approved**. It is one record:
|
||||
an **OS release**, with an id and a date. The hub stores it. The register does not.
|
||||
4. A ring-1 box asks the hub for the newest approved OS release. For each package it has installed,
|
||||
if the approved version is newer, it installs that **exact** version. It never installs a version
|
||||
newer than the approved one.
|
||||
5. Each box reports which OS release it runs, how many updates are waiting, whether it needs a reboot,
|
||||
and which packages it has that **no approved list covers** (see §6, edge case 9).
|
||||
|
||||
**OPEN Q1 is the weak point.** The design works only if a box can still download the approved version
|
||||
a week later. Debian's main and security archives keep only the newest version of each package. If
|
||||
Debian publishes a newer fix between approval and install, the approved version is gone. Options:
|
||||
the box waits for the next approval; or the box uses `snapshot.debian.org` (Debian's own dated archive,
|
||||
still signed by Debian); or Felhom runs a package cache. The spike measures how often this happens.
|
||||
|
||||
### 5.4 Who runs it, and with what permission
|
||||
|
||||
- **The agent runs every OS update**, for the host and for the guest (`pct exec`). The controller does
|
||||
not, because of constraint 4.
|
||||
- **The agent gets no general `apt` permission.** Like `felhom-selfupdate-guarded` (`03`, the self-update section),
|
||||
a small root-owned wrapper does the work. It accepts only an approved list and refuses everything else:
|
||||
- no package removal;
|
||||
- no downgrade, except the undo of §5.6, which an operator job signs;
|
||||
- no package that the box does not already have, unless the approved list records it as a dependency
|
||||
that the same update pulled in on ring 0 (a kernel update installs a NEW package name each time,
|
||||
for example `proxmox-kernel-6.x.y-z-pve-signed`, pulled by `proxmox-default-kernel`);
|
||||
- no package source other than the ones the installer set up;
|
||||
- non-interactive, and it always keeps the existing config file (`--force-confold`) and reports the
|
||||
conflict.
|
||||
- **Package signatures stay the publishers'** (Debian, Proxmox, Docker). The hub sends only names and
|
||||
versions. A broken-into hub can choose an older version or no version. It cannot make a box install a
|
||||
package that the publisher did not sign.
|
||||
|
||||
### 5.5 When
|
||||
|
||||
Inside the household's night window, after the backups:
|
||||
|
||||
```
|
||||
W DB dump
|
||||
W+60m Tier 2
|
||||
W+105m off-site → app updates (until W+5h at most)
|
||||
[W+2h, W+6h) whole-guest backup (agent)
|
||||
after it OS updates — guest first, then host
|
||||
```
|
||||
|
||||
The OS leg starts **after the whole-guest backup has finished**, so the guest's newest full copy is
|
||||
minutes old. This is the opposite order to app updates, which run before the whole-guest backup
|
||||
(`07` §6.1, `09` decision 11). The reason: the whole-guest backup is the guest's undo.
|
||||
|
||||
**OPEN Q8:** what else runs then. The controller's self-update (default 04:30, and after any hub report
|
||||
when a floor is above it, R-608), the agent's self-update, the restore-tests, and PBS jobs. The OS leg
|
||||
must never overlap a backup, a restore-test or a self-update.
|
||||
|
||||
### 5.6 How a failed update is undone
|
||||
|
||||
"Rollback" is not used (`09` §4). The shapes:
|
||||
|
||||
| Layer | Undo | Limit |
|
||||
|---|---|---|
|
||||
| Guest packages | Restore the guest snapshot taken just before the update | It also undoes app data written after the snapshot. Use it only inside the health window, before apps have written much. After that, install the previous version (needs OPEN Q1). |
|
||||
| Docker engine | Install the previous version (Docker's repository keeps old versions — OPEN Q2) | Every container restarts again |
|
||||
| Host packages | Install the previous version | Only if the source still has it (OPEN Q1/Q2) |
|
||||
| Host kernel | Boot the previous kernel. Proxmox can boot a new kernel **once** (`proxmox-boot-tool kernel pin <ver> --next-boot`). If that boot fails, the next boot uses the old kernel again. Make the new kernel permanent only after a healthy boot. | If the new kernel hangs, someone must switch the box off and on. The spike checks whether a hardware watchdog can do that (OPEN Q4). |
|
||||
|
||||
### 5.7 Telling people
|
||||
|
||||
- **Operator:** a hub event for each update and each failure; a fleet view showing each box's OS release,
|
||||
how far behind it is, and whether it needs a reboot. A box that is more than N days behind the newest
|
||||
approved release raises an alarm (`08`).
|
||||
- **Household:** one line on the timeline in both languages, informal voice: what was updated and
|
||||
whether the box restarted. Telling households in advance that the box may restart at night is a
|
||||
**promise to users**. That is the operator's decision when the slow lane is built.
|
||||
|
||||
---
|
||||
|
||||
## 6. Risks and edge cases
|
||||
|
||||
| # | What can go wrong | What the design does |
|
||||
|---|---|---|
|
||||
| 1 | A new kernel does not boot | Boot it once (`--next-boot`); the old kernel stays the default. A hard hang needs a power cycle (OPEN Q4). |
|
||||
| 2 | Power cut during an update: `dpkg` is half done | The wrapper runs `dpkg --configure -a` and `apt-get -f install` first, every time, and reports what it repaired (spike measures, Q5). |
|
||||
| 3 | An update asks a question (changed config file, service restart prompt) | Non-interactive, keep the old config, report the conflict. |
|
||||
| 4 | A Docker update stops every app, and the controller | Stop the apps cleanly first, like before a backup. Measure whether `live-restore` keeps containers running (Q3). The agent drives it. |
|
||||
| 5 | The approved version is no longer downloadable | OPEN Q1. Until it is answered, the box waits for the next approval and reports it. |
|
||||
| 6 | An urgent security hole | The operator approves the same day once ring 0 has run it. |
|
||||
| 7 | A box was off for months | It catches up through the newest approved release. It does not install every release in between. Packages are not like app data; one `apt` step is enough. Kernel and Docker still go one approved step at a time. |
|
||||
| 8 | The system disk is full | Check free space before downloading; clean the package cache after; refuse and report. |
|
||||
| 9 | A customer box has a package that ring 0 does not have (other hardware: firmware, CPU microcode, NIC drivers) | It is never approved, so it is never updated. The box reports it as "not covered", and the hub raises it. Fix: add matching hardware to ring 0, or approve it by hand. |
|
||||
| 10 | The package source is down, or its signing key changes (Docker has done this) | The update fails cleanly and the box reports it. A key change is a slow-lane act for a person. |
|
||||
| 11 | A BYO host | Only the guest and Docker are updated (§1). |
|
||||
| 12 | Two boxes on ring 0 is a small sample | Accepted for now. When there are customers, the first tester boxes can become a second ring. |
|
||||
| 13 | A broken-into hub sends a harmful list | The wrapper refuses removals, downgrades and new packages, and the publishers' signatures still apply (§5.4). |
|
||||
| 14 | `cloudflared` on the host | Not covered by this file yet. How it is installed and updated is OPEN Q9. It faces the internet, so it matters. |
|
||||
| 15 | The no-subscription Proxmox repository gets less testing than the enterprise one | Ring 0 is our test. The enterprise repository costs a yearly fee per box. That is a **money** decision for the operator, later (Q10). |
|
||||
|
||||
---
|
||||
|
||||
## 7. Open questions the spike must answer
|
||||
|
||||
| Q | Question | How to answer |
|
||||
|---|---|---|
|
||||
| Q1 | Can a box install an exact Debian version one week after approval? How often is it already superseded? | `apt-cache madison` on host and guest for security packages. Debian's security announcement history. Whether `snapshot.debian.org` is reachable and fast enough. |
|
||||
| Q2 | Do the Proxmox and Docker sources keep older versions? | `apt-cache madison` for `pve-manager`, `proxmox-kernel-*`, `docker-ce`, `containerd.io`. |
|
||||
| Q3 | What does a Docker engine update do to running containers, with and without `live-restore`? | On scratch guest 9202: step from one pinned version to the next; time the downtime. |
|
||||
| Q4 | Does `proxmox-boot-tool kernel pin --next-boot` work on both demo hosts' boot setups (UEFI or legacy, GRUB or systemd-boot)? Does either box have a hardware watchdog? | On demo-hp, with the operator's go before each reboot. |
|
||||
| Q5 | What does an interrupted `apt` run leave, and does the repair recover it? | Kill a run on 9202, after a snapshot. |
|
||||
| Q6 | How far behind are the demo boxes today, per layer and per source? How long does catching up take, and what restarts? | `apt-get -s upgrade` (read only) first; then a real run on 9202 and on one demo host. |
|
||||
| Q7 | Which Debian packages restart something big? | `needrestart` in list mode after a run. |
|
||||
| Q8 | What else runs in the night window, and where does the OS leg fit? | Read the timers on host and guest. |
|
||||
| Q9 | How is `cloudflared` installed and updated on the host? | Read the installer and the agent. |
|
||||
| Q10 | What does the Proxmox enterprise repository cost per box per year, and what does it add? | The publisher's price page. Record only. The operator decides later. |
|
||||
|
||||
---
|
||||
|
||||
## 8. Build order `[PROPOSAL]`
|
||||
|
||||
Each step returns to the operator for go or no-go.
|
||||
|
||||
1. **Spike** (measure Q1–Q10; no product code).
|
||||
2. **Guest Debian, fast lane.** Lowest risk: a snapshot undo exists.
|
||||
3. **Host Debian, fast lane** (no kernel, no Proxmox packages).
|
||||
4. **Fleet view and alarms** (§5.7).
|
||||
5. **Slow lane: Docker engine.**
|
||||
6. **Slow lane: host kernel and Proxmox packages, with the reboot.**
|
||||
7. **Later:** the Proxmox major upgrade (PVE 9 → 10), drilled on ring 0 first.
|
||||
|
||||
---
|
||||
|
||||
## 9. Where the rest lives
|
||||
|
||||
- The finding: **R-812** (`backlog/OPEN-ITEMS.md`). The intention: **R-808** (`backlog/ROADMAP.md`).
|
||||
- Related: **R-604** (a per-customer floor hides a box from global raises; the same risk applies to an
|
||||
OS-release floor), **R-530** (agents update only by a signed job per box).
|
||||
- App updates: `09`. Backups and the night chain: `07` §6.1. The agent's permissions: `03` §3.
|
||||
Reference in New Issue
Block a user