> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
> **Rulings 2026-10-04 (day) — recorded before the work (off-site close + OS update spike).**`09` §3 decisions **75**
> (the off-site topic closes first, one brief), **76** (the OS fast lane follows an approved list with a 1–2 day wait;
> `unattended-upgrades` with no wait rejected) and **77** (`architecture/11-os-updates.md` added unchanged, NOT RATIFIED,
> then corrected by measurement).
> **Rulings 2026-10-04 (morning) — recorded before the work.**`09` §3 decisions **71** (ep0's DooPlex copy keeps 8 weekly copies; a third place later, R-832), **72** (tester-1's unpinned keys removed via the registrar), **73** (the transcript-exposed Hetzner storage token is not rotated now; R-831), and **74** (the set-aside deletion is the hub's, after a 7-day wait — from the brief's Part E).
> **Rulings 2026-10-03 (afternoon) — recorded before the work (off-site lock build).**`09` §3 decisions **68** (the box prunes only inside a weekly hub-opened window; option 2 rejected; no box-side retention in the interim), **69** (the hub is the key registrar; the box never receives the sub-account password; the hub stores it encrypted with a key outside the database; daily `authorized_keys` check) and **70** (nightly PBS pull-sync of ep0's `felhom-offsite` to DooPlex; the fence opens for that brief's Part F acts only).
- [`architecture/02-controller-module-map.md`](architecture/02-controller-module-map.md) — **historical** v0.33 planning map; the live map is [`controller/module-map.md`](controller/module-map.md)
| **The hub reports LOSS OF VISIBILITY into either off-site store (not just how full it is)** | hub **v0.106.0** (R-339) | **IMPLEMENTED — deliberately NOT proven-live** | Both box checkers count consecutive failed fetch windows and emit `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default 3 windows (≈30–45 min), each with a paired `*_recovered` all-clear routed via `recoveredPairedDownTypes` — required because the recoveries are severity `info`, which `severityNotifies` drops. Scopes stay customer-less (`pbsdr-box` / `pool-box`) → operator channel only. Fill logic untouched: a degraded read still drives no band transition. Evidence: `internal/monitor/box_reachability_test.go` + the cross-package wiring test in `internal/notify/`, which asserts an actual operator mail rather than a map entry. **Filed BECAUSE of a measured gap**, not a hypothesis: the 2026-08-18 ep0 outage ran 9 h 37 m with the hub silent | **The gap that remains is R-340**, and it is not small: the ep0 read is the `usage` op, which rides the LOCAL API daemon — the daemon that incident explicitly cleared — so this check would have shown GREEN for that entire outage. It closes "ep0 is unreachable as a host"; it does not close what actually happened. **No live or constructed outage has exercised the emit path**, and one cannot be manufactured against ep0 (Tier 2, protected) |
| Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | |
| **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** |
| **A deletion of a customer's off-site history is NOTICED within a day** | hub **v0.111.0** (R-431) | **IMPLEMENTED — not yet PROVEN-LIVE** | 09-01 | `hub/internal/monitor/offsite.go` — third signal beside FILL and STALENESS. **On the hub deliberately:** a detector on the box is one the deletion can silence. Alarms when the reported count falls by more than HALF and by at least 5, guarded by `StatsKnown` (R-331), the declared `State` (R-204) and run success (R-100). **Threshold reasoned, not invented:** over 12 898 reports every decrease lands on ZERO and predates `stats_known`; in the 380-report `stats_known` window there are none. **ACCEPTANCE: 9 009 real points replayed → ZERO alarms** (`offsite_r431_test.go`, fixture committed). **What PROVEN-LIVE would need and this does NOT have:** a real drop observed on a live box producing a real mail — the live firing done at ship time was driven through the hub's own path with synthetic counts, which is an end-to-end delivery proof, not a proof that a genuine deletion is caught. |
> | **Freshness** | **CURRENT** as of 2026-10-04. The spike `TASK-backup-close-and-os-updates-spike-2026-10-04` adds measurements here as `[FACT]` and corrects every claim it disproves. Mark this file STALE when it falls behind what the product does. |
>
> **How to read this document.** Each statement has a label:
>
> - **[FACT]**: an observed property, with a `file:line`, a register row or an audit path.
> - **[RULED]**: an operator decision, with its date.
> - **[PROPOSAL]**: the reviewer's design. It is not decided and not built.
> - **OPEN**: a question that nobody has answered yet. The spike measures it. Nobody guesses it.
>
> **Why this file exists.** No architecture document covered operating-system updates. `00` §G marks the
> capability MISSING (added 2026-10-03). The finding is **R-812**. The roadmap intention is **R-808**.
> **The register carries the work. This file carries the reasoning. The source is the truth.**
---
## 0. In plain language
A box runs three layers that we install and never update: the Proxmox host, the small Debian system
inside the guest, and the Docker engine. App images are updated (see `09`). The layer under them is not.
A box lives in a home for years, so this is a security gap.
The proposal has two lanes. **The fast lane** applies Debian's security fixes automatically, but only
the exact versions that already ran well on the demo boxes for 1–2 days. **The slow lane** covers the
kernel, Proxmox and Docker. These cause restarts or big changes, so a person approves them one version at
a time, and they run only at night. The tested versions are recorded automatically from what the demo
boxes installed. Nobody keeps a hand-written list.
---
## 1. Scope
**In scope.**
- The Proxmox host on an **appliance** install: the Debian 13 base, the Proxmox packages, and the kernel.
- The guest (the customer LXC): its Debian 13 packages.
- The Docker engine inside the guest (`docker-ce`, `docker-ce-cli`, `containerd.io`).
- How the household and the operator are told, and how a failed update is undone.
**Out of scope.**
- **App images.**`09` covers them (the ladder, the monthly same-tag re-test).
- **The controller and agent binaries.** Their own self-update covers them (`03`, the self-update section; `09` R-608).
- **DooPlex and ep0.** The operator updates them by hand (`runbooks/offsite-endpoint.md`).
- **A BYO host** (the owner brought their own Proxmox). `[FACT]` The installer leaves its repositories
`align_apt_repos`). `[PROPOSAL]` On a BYO box we update the guest and Docker only, never the host.
- **The Proxmox MAJOR upgrade** (PVE 9 → 10). It is a later step of its own, drilled first (R-808 item 5).
---
## 2. What a box runs, as measured
| Layer | What it is | Source of packages | Evidence |
|---|---|---|---|
| Host | Proxmox VE 9.2 on Debian 13, LVM-thin | Debian mirrors + `pve-no-subscription` (enterprise repo switched off on appliance installs) | `[FACT]``01` §2; `felhom-host-install.sh``align_apt_repos` (~L2130–2136) |
| Host kernel | The only kernel on the box. The guest has none. | `pve-no-subscription` (`proxmox-kernel-*`) | `[FACT]` LXC shares the host kernel (`01` §2: "one LXC/kernel/Docker daemon") |
| Controller | A container in the guest. A Docker engine restart restarts it too. | Felhom registry | `[FACT]``03` §1 |
**[FACT] Nothing updates any of these layers today** (R-812, searched 2026-10-03). The installer
says *"No upgrades are run — repo alignment only"* (`felhom-host-install.sh` ~L2135). A fresh install
gets the Docker engine that was current when its golden was baked, and keeps it.
**[FACT] The agent may not run `apt` today, with one exception.** Its sudoers allowlist
(`felhom-agent/configs/felhom-agent.sudoers`) holds one apt line: `apt-get install -y -q dnsmasq`.
`apt` as root runs package scripts as root, so a broad `apt` grant is a full root grant. §5.4
proposes how to avoid that.
---
## 3. Operator rulings
**2026-10-03.** R-808 is on the roadmap at P2: *"every box receives operating-system security
patches on a schedule, and a failed update is undone."*
**[RULED] 2026-10-04: the fast lane follows an approved list, with a 1–2 day wait.** Every update
runs on the demo boxes first. The other boxes install only the exact versions that the demo boxes ran
without trouble. The exact wait is set when the feature is built. **Rejected:** Debian's own
`unattended-upgrades` with no wait. It is simpler and common, but a bad update would reach every
customer at the same time.
**[RULED] 2026-10-04: the off-site backup topic is closed first.** The same brief carries both
topics, and the backup part runs first.
The two-lane split (§5.2) is the reviewer's proposal. The operator's ruling above assumes it, but he
has not ruled on it as such.
---
## 4. The constraints that shape the design
1. **Unattended.** Nobody is at the box. Nobody answers a question that `apt` asks.
2. **No screen.** If the box does not boot, the household sees only that nothing works.
3. **One kernel for everything.** A host kernel update needs a host reboot. A reboot stops every app.
4. **The controller lives in Docker.** A Docker engine update restarts the controller in the middle of
its own work. So the controller cannot drive a Docker update. The agent, on the host, must.
5. **The host has no whole-system backup.** The guest has three tiers (`07`). The host has none.
A host update can only be undone by installing the previous version again.
6. **Every update must already have run on a box we own.** This is the lesson of the update arc
(`09` §3 decision 13: "the test decides").
7. **Few packages.** Each extra package is one more thing to update and break. The guest is the
Debian standard template plus Docker. Keep it that way.
---
## 5. The proposed shape `[PROPOSAL]`
### 5.1 Two rings
- **Ring 0:** demo-felhom (N100) and demo-hp. They are disposable (`runbooks/target-selection.md`), and
they have different hardware. They take every update first.
- **Ring 1:** every other box. It takes only what ring 0 approved.
### 5.2 Two lanes
| | Fast lane | Slow lane |
|---|---|---|
| What | Debian packages on host and guest, from `trixie-security` and the stable point releases, **except** the slow-lane list | Kernel (`proxmox-kernel-*`), Proxmox (`pve-*`, `proxmox-*`, `lxc-pve`, `qemu-server`, …), Docker (`docker-ce*`, `containerd.io`) |
| Restart | A service restart at most. No reboot. | Kernel: a host reboot. Docker: every container restarts. Proxmox: its services restart. |
| Approval | Automatic. A version is approved when ring 0 ran it and stayed healthy for the wait (1–2 days). | A person approves one version at a time, like a controller floor. |
| Urgent fix | The operator can approve a version on the same day once ring 0 has run it. | The same. |
| When | The night window (§5.5) | The night window. A kernel reboot only on a night the operator scheduled. |
Which packages count as slow lane is a list the spike checks (OPEN Q7). A Debian package that restarts
something big (for example `systemd`, `libc6`, `openssh-server`) may belong in the slow lane too.
### 5.3 The approved list (the "tested versions" record)
Nobody writes the list by hand. It fills itself:
1. A ring-0 box updates. Afterwards it reports to the hub the exact `package=version` it installed,
per layer (host, guest), with each package's origin (`Debian-Security`, `Debian`, `Proxmox`,
`Docker`).
2. The box then reports health for the wait period: the agent, the controller, every app's health,
and the guest's network.
3. When the wait passes with ring 0 healthy, the hub marks that set **approved**. It is one record:
an **OS release**, with an id and a date. The hub stores it. The register does not.
4. A ring-1 box asks the hub for the newest approved OS release. For each package it has installed,
if the approved version is newer, it installs that **exact** version. It never installs a version
newer than the approved one.
5. Each box reports which OS release it runs, how many updates are waiting, whether it needs a reboot,
and which packages it has that **no approved list covers** (see §6, edge case 9).
**OPEN Q1 is the weak point.** The design works only if a box can still download the approved version
a week later. Debian's main and security archives keep only the newest version of each package. If
Debian publishes a newer fix between approval and install, the approved version is gone. Options:
the box waits for the next approval; or the box uses `snapshot.debian.org` (Debian's own dated archive,
still signed by Debian); or Felhom runs a package cache. The spike measures how often this happens.
### 5.4 Who runs it, and with what permission
- **The agent runs every OS update**, for the host and for the guest (`pct exec`). The controller does
not, because of constraint 4.
- **The agent gets no general `apt` permission.** Like `felhom-selfupdate-guarded` (`03`, the self-update section),
a small root-owned wrapper does the work. It accepts only an approved list and refuses everything else:
- no package removal;
- no downgrade, except the undo of §5.6, which an operator job signs;
- no package that the box does not already have, unless the approved list records it as a dependency
that the same update pulled in on ring 0 (a kernel update installs a NEW package name each time,
for example `proxmox-kernel-6.x.y-z-pve-signed`, pulled by `proxmox-default-kernel`);
- no package source other than the ones the installer set up;
- non-interactive, and it always keeps the existing config file (`--force-confold`) and reports the
conflict.
- **Package signatures stay the publishers'** (Debian, Proxmox, Docker). The hub sends only names and
versions. A broken-into hub can choose an older version or no version. It cannot make a box install a
package that the publisher did not sign.
### 5.5 When
Inside the household's night window, after the backups:
```
W DB dump
W+60m Tier 2
W+105m off-site → app updates (until W+5h at most)
[W+2h, W+6h) whole-guest backup (agent)
after it OS updates — guest first, then host
```
The OS leg starts **after the whole-guest backup has finished**, so the guest's newest full copy is
minutes old. This is the opposite order to app updates, which run before the whole-guest backup
(`07` §6.1, `09` decision 11). The reason: the whole-guest backup is the guest's undo.
**OPEN Q8:** what else runs then. The controller's self-update (default 04:30, and after any hub report
when a floor is above it, R-608), the agent's self-update, the restore-tests, and PBS jobs. The OS leg
must never overlap a backup, a restore-test or a self-update.
### 5.6 How a failed update is undone
"Rollback" is not used (`09` §4). The shapes:
| Layer | Undo | Limit |
|---|---|---|
| Guest packages | Restore the guest snapshot taken just before the update | It also undoes app data written after the snapshot. Use it only inside the health window, before apps have written much. After that, install the previous version (needs OPEN Q1). |
| Docker engine | Install the previous version (Docker's repository keeps old versions — OPEN Q2) | Every container restarts again |
| Host packages | Install the previous version | Only if the source still has it (OPEN Q1/Q2) |
| Host kernel | Boot the previous kernel. Proxmox can boot a new kernel **once** (`proxmox-boot-tool kernel pin <ver> --next-boot`). If that boot fails, the next boot uses the old kernel again. Make the new kernel permanent only after a healthy boot. | If the new kernel hangs, someone must switch the box off and on. The spike checks whether a hardware watchdog can do that (OPEN Q4). |
### 5.7 Telling people
- **Operator:** a hub event for each update and each failure; a fleet view showing each box's OS release,
how far behind it is, and whether it needs a reboot. A box that is more than N days behind the newest
approved release raises an alarm (`08`).
- **Household:** one line on the timeline in both languages, informal voice: what was updated and
whether the box restarted. Telling households in advance that the box may restart at night is a
**promise to users**. That is the operator's decision when the slow lane is built.
---
## 6. Risks and edge cases
| # | What can go wrong | What the design does |
|---|---|---|
| 1 | A new kernel does not boot | Boot it once (`--next-boot`); the old kernel stays the default. A hard hang needs a power cycle (OPEN Q4). |
| 2 | Power cut during an update: `dpkg` is half done | The wrapper runs `dpkg --configure -a` and `apt-get -f install` first, every time, and reports what it repaired (spike measures, Q5). |
| 3 | An update asks a question (changed config file, service restart prompt) | Non-interactive, keep the old config, report the conflict. |
| 4 | A Docker update stops every app, and the controller | Stop the apps cleanly first, like before a backup. Measure whether `live-restore` keeps containers running (Q3). The agent drives it. |
| 5 | The approved version is no longer downloadable | OPEN Q1. Until it is answered, the box waits for the next approval and reports it. |
| 6 | An urgent security hole | The operator approves the same day once ring 0 has run it. |
| 7 | A box was off for months | It catches up through the newest approved release. It does not install every release in between. Packages are not like app data; one `apt` step is enough. Kernel and Docker still go one approved step at a time. |
| 8 | The system disk is full | Check free space before downloading; clean the package cache after; refuse and report. |
| 9 | A customer box has a package that ring 0 does not have (other hardware: firmware, CPU microcode, NIC drivers) | It is never approved, so it is never updated. The box reports it as "not covered", and the hub raises it. Fix: add matching hardware to ring 0, or approve it by hand. |
| 10 | The package source is down, or its signing key changes (Docker has done this) | The update fails cleanly and the box reports it. A key change is a slow-lane act for a person. |
| 11 | A BYO host | Only the guest and Docker are updated (§1). |
| 12 | Two boxes on ring 0 is a small sample | Accepted for now. When there are customers, the first tester boxes can become a second ring. |
| 13 | A broken-into hub sends a harmful list | The wrapper refuses removals, downgrades and new packages, and the publishers' signatures still apply (§5.4). |
| 14 | `cloudflared` on the host | Not covered by this file yet. How it is installed and updated is OPEN Q9. It faces the internet, so it matters. |
| 15 | The no-subscription Proxmox repository gets less testing than the enterprise one | Ring 0 is our test. The enterprise repository costs a yearly fee per box. That is a **money** decision for the operator, later (Q10). |
---
## 7. Open questions the spike must answer
| Q | Question | How to answer |
|---|---|---|
| Q1 | Can a box install an exact Debian version one week after approval? How often is it already superseded? | `apt-cache madison` on host and guest for security packages. Debian's security announcement history. Whether `snapshot.debian.org` is reachable and fast enough. |
| Q2 | Do the Proxmox and Docker sources keep older versions? | `apt-cache madison` for `pve-manager`, `proxmox-kernel-*`, `docker-ce`, `containerd.io`. |
| Q3 | What does a Docker engine update do to running containers, with and without `live-restore`? | On scratch guest 9202: step from one pinned version to the next; time the downtime. |
| Q4 | Does `proxmox-boot-tool kernel pin --next-boot` work on both demo hosts' boot setups (UEFI or legacy, GRUB or systemd-boot)? Does either box have a hardware watchdog? | On demo-hp, with the operator's go before each reboot. |
| Q5 | What does an interrupted `apt` run leave, and does the repair recover it? | Kill a run on 9202, after a snapshot. |
| Q6 | How far behind are the demo boxes today, per layer and per source? How long does catching up take, and what restarts? | `apt-get -s upgrade` (read only) first; then a real run on 9202 and on one demo host. |
| Q7 | Which Debian packages restart something big? | `needrestart` in list mode after a run. |
| Q8 | What else runs in the night window, and where does the OS leg fit? | Read the timers on host and guest. |
| Q9 | How is `cloudflared` installed and updated on the host? | Read the installer and the agent. |
| Q10 | What does the Proxmox enterprise repository cost per box per year, and what does it add? | The publisher's price page. Record only. The operator decides later. |
---
## 8. Build order `[PROPOSAL]`
Each step returns to the operator for go or no-go.
1. **Spike** (measure Q1–Q10; no product code).
2. **Guest Debian, fast lane.** Lowest risk: a snapshot undo exists.
3. **Host Debian, fast lane** (no kernel, no Proxmox packages).
4. **Fleet view and alarms** (§5.7).
5. **Slow lane: Docker engine.**
6. **Slow lane: host kernel and Proxmox packages, with the reboot.**
7. **Later:** the Proxmox major upgrade (PVE 9 → 10), drilled on ring 0 first.
---
## 9. Where the rest lives
- The finding: **R-812** (`backlog/OPEN-ITEMS.md`). The intention: **R-808** (`backlog/ROADMAP.md`).
- Related: **R-604** (a per-customer floor hides a box from global raises; the same risk applies to an
OS-release floor), **R-530** (agents update only by a signed job per box).
- App updates: `09`. Backups and the night chain: `07` §6.1. The agent's permissions: `03` §3.
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.