docs: OS updates steps 3+4 BUILT (11 §8.2/§8.3, §5.6 kernel facts, §5.8 Docker slow lane design), 00/03/07/08 updated, decisions 84-86 (CC unattended), host undo runbook (proved), register: R-841 R-845 R-846 R-850 closed, R-848 R-849 R-851 opened, R-836 R-812 narrowed (332 -> 333); STATUS; live evidence
gates / gates (push) Successful in 31s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-04 14:47:21 +02:00
parent 0e970ba384
commit 31bdb4b549
36 changed files with 1679 additions and 16 deletions
+10
View File
@@ -16,6 +16,16 @@
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
> **2026-10-04 (late afternoon) — OS updates: host fast lane + fleet view + alarms BUILT (`11` §8.2–§8.3); the tunnel
> status is true (R-841).** Agent v0.141.0 → v0.141.1, hub v0.131.0 → v0.131.1, controller v0.292.0 (cloudflared
> readiness health check). **Decided by CC unattended — operator may reverse:** `09` §3 decisions **84** (appliance proof
> = the root-owned install record), **85** (alarm numbers 7/14/7/14 days, configuration), **86** (a same-session patch
> release of agent and hub for the host "reboot needed" defect, R-846). Live: `tunnel_down`/`tunnel_recovered` on
> demo-hp; ring 0 host pass on demo-felhom (108 Debian packages); a 605-package host release; ring 1 exact install;
> by-hand host undo proved; leg 23–32 s with nothing to install. Kernel spike (R-836, narrowed): GRUB's one-shot is not
> a one-shot on LVM `/boot`; Secure Boot fine; `kernel.panic = 0` (R-851); `sp5100_tco` answers. demo-hp now runs and
> defaults to kernel 7.0.14-20. Docker slow lane designed (`11` §5.8). `REPORT-os-host-lane-2026-10-04.md`.
> **Rulings 2026-10-04 (~12:20) — recorded before the work (host fast lane brief).** `09` §3 decisions **81** (R-842 A:
> the undo is the whole-guest backup by hand; R-842 closed), **82** (R-840 not built now — "There are no older boxes";
> row kept open with the reviewer's note) and **83** (next: R-841, the host fast lane, OS-update improvements).
+40 -2
View File
@@ -2,8 +2,46 @@
**Ready for the first real tester (Tester-2): yes. You confirmed the tunnel route and the connect mails (2026-09-30).**
**Updated 2026-10-04 (afternoon): guest system updates are automatic on the demo boxes. Both demo boxes run
controller 0.291.0 and host agent 0.140.0. Hub 0.130.0. New installs get golden 0.291.0.**
**Updated 2026-10-04 (evening): the HOST's security fixes install themselves too, the hub has a fleet view and four
OS alarms, and the tunnel status is true. Both demo boxes run controller 0.292.0 and host agent 0.141.1. Hub 0.131.1.
New installs: see the golden line in the section below.**
## Today (2026-10-04, evening): host fixes, the fleet view, a true tunnel status
**Decisions I took myself (you may reverse each):**
- Only a box the installer itself recorded as "appliance" gets host updates (a record the agent cannot change).
- The four alarm times: no update run for 7 days; "reboot needed" for 14 days; the demo boxes approve nothing for 7
days; a box has fixes nobody approved for 14 days. All four are settings.
- I released the host agent and the hub a second time today. The first release said "reboot needed" wrongly; the new
14-day alarm would then have mailed you about boxes you had already rebooted.
**One decision for you (a safe default if you say nothing) — the Docker engine updates (designed, not built):**
1. **Turn on Docker's "live-restore" on every box.** With it, a Docker engine update restarts no app (measured: 0
restarts). Without it, every app stops for about 30 seconds per engine update.
- **A (my pick):** turn it on — in the new-install image and once on existing boxes. It goes on without restarting
anything. Cost: it must never be turned off by a plain restart again (that stops every app and starts none).
- **B:** leave it off. Every Docker update then means ~30 seconds of every app being down, at night.
- **If you say nothing:** nothing changes; Docker updates stay unbuilt.
**What I did:**
- **The host's Debian fixes now install themselves**, after the guest's, on the same night run, only on appliances,
never a kernel or boot package, never a reboot. Proven: demo-felhom installed 108 host fixes and stayed healthy; all
108 came from Debian. A box made "ring 1" for a test installed exactly the one approved host version it lacked.
- **A way to put one host package back by hand** is written and proven on demo-hp (and the test taught it two fixes).
- **The fleet view** in the hub: one line per box with its updates, "reboot needed since", and the tunnel.
- **The tunnel status is now true:** running, not running, or unknown. I blocked demo-hp's tunnel: after two reports
(about 30 minutes) you got a "tunnel down" mail; when I unblocked it, "recovered". A simply stopped tunnel heals
itself within 5 minutes, before the hub can even see it.
- **The night run is fast:** 23–32 seconds when there is nothing to install (target was under 60).
- **The kernel test on demo-hp (your two reboots):** Secure Boot works with it, but GRUB's "boot once" does not work
on our boxes: the second plain reboot came up on the NEW kernel again. So a new kernel that hangs would stay. That
must be solved before kernel updates. demo-hp now runs the newer kernel, healthy. A watchdog chip exists on demo-hp.
- **Found and fixed during the live test:** "reboot needed" was wrong in two ways (fixed in the second release).
- **Rows:** 2 closed, 2 opened-and-closed the same day, 3 opened. The list went from 332 to 333.
**Needs you later (nothing breaks if you wait):**
- **A host that crashes does not restart by itself** (Linux's "panic" setting is off). Changing it changes how every
box behaves, so it is your call, together with kernel updates.
## Today (2026-10-04, afternoon): the guest's security fixes install themselves
@@ -232,7 +232,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
| **The hub reports LOSS OF VISIBILITY into either off-site store (not just how full it is)** | hub **v0.106.0** (R-339) | **IMPLEMENTED — deliberately NOT proven-live** | Both box checkers count consecutive failed fetch windows and emit `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default 3 windows (≈30–45 min), each with a paired `*_recovered` all-clear routed via `recoveredPairedDownTypes` — required because the recoveries are severity `info`, which `severityNotifies` drops. Scopes stay customer-less (`pbsdr-box` / `pool-box`) → operator channel only. Fill logic untouched: a degraded read still drives no band transition. Evidence: `internal/monitor/box_reachability_test.go` + the cross-package wiring test in `internal/notify/`, which asserts an actual operator mail rather than a map entry. **Filed BECAUSE of a measured gap**, not a hypothesis: the 2026-08-18 ep0 outage ran 9 h 37 m with the hub silent | **The gap that remains is R-340**, and it is not small: the ep0 read is the `usage` op, which rides the LOCAL API daemon — the daemon that incident explicitly cleared — so this check would have shown GREEN for that entire outage. It closes "ep0 is unreachable as a host"; it does not close what actually happened. **No live or constructed outage has exercised the emit path**, and one cannot be manufactured against ep0 (Tier 2, protected) |
| Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | |
| Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | |
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.140.0, hub v0.130.0 | **PARTIAL — the GUEST's Debian fast lane is PROVEN-LIVE (2026-10-04); the host, Docker and the kernel are MISSING** | `audits/os-guest-lane-2026-10-04/` — ring 0 (both demo boxes) installed 53 Debian fixes each, healthy; the hub approved a 272-package release (TEST wait 2 min, then the ruled 24 h + 1 night restored); demo-felhom as ring 1 installed exactly the 3 approved versions; a stopped app → `health_failed` + operator mail. Design `architecture/11-os-updates.md` §8.1 | **No automatic undo** (a customer guest cannot be snapshotted, R-837 → R-842); existing boxes need the wrapper + sudoers by hand (R-840); host / Docker / kernel lanes not built (R-812, R-835, R-836); `felhom-host-install.sh:2133-2136` still runs no host upgrades |
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.141.1, hub v0.131.1 | **PARTIAL — the GUEST and HOST Debian fast lanes are PROVEN-LIVE (2026-10-04), with the fleet view and four operator alarms; Docker and the kernel are MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/` — ring 0 on both demo hosts (demo-felhom 108 Debian host packages, healthy, 0 Proxmox-origin); a 605-package host release approved (TEST wait, then the ruled 24 h + 1 night restored); demo-felhom as ring 1 installed exactly the one host version it lacked; the by-hand host undo proved (`runbooks/os-updates-host-undo.md`); `tunnel_down` / `tunnel_recovered` live. Design `architecture/11-os-updates.md` §8.1–§8.3 | **No automatic undo** (guest: last night's backup, decision 81; host: the by-hand runbook); existing boxes need the wrapper + sudoers by hand (R-840); Docker and kernel lanes not built (R-812, R-835, R-836); a host panic is not restarted (R-851) |
| **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** |
| **A deletion of a customer's off-site history is NOTICED within a day** | hub **v0.111.0** (R-431) | **IMPLEMENTED — not yet PROVEN-LIVE** | 09-01 | `hub/internal/monitor/offsite.go` — third signal beside FILL and STALENESS. **On the hub deliberately:** a detector on the box is one the deletion can silence. Alarms when the reported count falls by more than HALF and by at least 5, guarded by `StatsKnown` (R-331), the declared `State` (R-204) and run success (R-100). **Threshold reasoned, not invented:** over 12 898 reports every decrease lands on ZERO and predates `stats_known`; in the 380-report `stats_known` window there are none. **ACCEPTANCE: 9 009 real points replayed → ZERO alarms** (`offsite_r431_test.go`, fixture committed). **What PROVEN-LIVE would need and this does NOT have:** a real drop observed on a live box producing a real mail — the live firing done at ship time was driven through the hub's own path with synthetic counts, which is an end-to-end delivery proof, not a proof that a genuine deletion is caught. |
+3 -3
View File
@@ -47,7 +47,7 @@ Owns:
1. **Proxmox lifecycle** — create/start/stop/destroy guests, snapshots, storage allocation. Via a scoped Proxmox API token (the **`FelhomAgent` operator role** — `proxmox-platform.md` §3.6, validated Phase 3 B3) for everything the API covers; raw host ops only where unavoidable.
2. **Storage management** — attach/classify targets, reconcile the storage manifest, mount USB-by-UUID, present mounts into guests.
3. **Backup/restore orchestration** — vzdump to the tiers, PBS, snapshot management, and the **self-restore-test**.
4. **Host & tunnel monitoring** — host metrics, guest up/down, storage-target status, and `cloudflared` health; reports the host domain to the hub. **[FACT, 2026-10-04] The "cloudflared health" leg reads a host unit that does not exist:** `internal/hub/cloudflared.go` runs `systemctl is-active cloudflared` on the HOST, but cloudflared is a container in the guest (below), so every box reports `inactive` (`Unit cloudflared.service could not be found`, demo-hp). R-841.
4. **Host & tunnel monitoring** — host metrics, guest up/down, storage-target status, and `cloudflared` health; reports the host domain to the hub. **[FACT, 2026-10-04] The "cloudflared health" leg reads a host unit that does not exist:** `internal/hub/cloudflared.go` runs `systemctl is-active cloudflared` on the HOST, but cloudflared is a container in the guest (below), so every box reports `inactive` (`Unit cloudflared.service could not be found`, demo-hp). R-841. **[FACT, FIXED agent v0.141.0 / controller v0.292.0, R-841 CLOSED]** The agent now reads the guest's `cloudflared` container through the existing `pct exec [0-9]* -- docker inspect -f *` sudoers line: state, exit code and the Docker health status of a check the controller adds (`cloudflared tunnel --metrics localhost:20241 ready` → cloudflared's own `/ready`, 200 only with a connection). Three states: `running` (healthy), `not_running` (stopped, absent, or running but NOT connected), `unknown` (could not ask, or the check is still starting) — `unknown` never alarms. No new sudoers line.
5. **Provisioning** — provision a guest **by restoring the golden base image** (§9), deploy the controller into it, hand it its bootstrap config; also **build and refresh the golden base image** itself.
6. **Hub control loop** — poll for desired state + signed jobs, reconcile, execute, report, heartbeat.
7. **Local API** — the per-guest authorization gate the controller calls.
@@ -64,7 +64,7 @@ Explicitly does **not**:
- **Native Go binary, systemd service** on the host: boot-start, `Restart=always`, systemd watchdog (kill+restart on hang), journald logging, resource limits.
- **Root-minimized (boundary settled — Phase 3 B3).** The agent runs as a **non-root** service user with the scoped `FelhomAgent` token for all API-covered work + a **narrow `sudoers` allowlist** for true host ops. Per Phase 3 (B3) the boundary is settled: the entire per-customer guest lifecycle — provision (by restore, §9), config, start/stop, snapshot, backup, **restore**, destroy — is token-covered. Genuine OS-root is confined to: (1) building/refreshing the **golden base image** (`keyctl` create is `root@pam`-only — one-time at enrollment + a maintenance cadence, §9); (2) **host mounts** (USB mount-by-UUID, systemd mount units / fstab); (3) **SMART / hardware sensors**. Root therefore never sits on the per-customer path. See `proxmox-platform.md` §3.6 for the role + boundary table.
- ~~**`cloudflared` is a separate systemd service**, not embedded in the agent. … The agent **manages and health-watches** it (see §5) but the tunnel does not live or die with the agent process.~~ **[FACT, corrected 2026-10-04 — `11-os-updates.md` C8, R-838]** `cloudflared` is a **container in the customer guest** (`cloudflare/cloudflared:<pin>`), rendered and kept up by the in-guest controller (`felhom-controller/controller/internal/infra/infra.go`, `internal/stacks/infra.go`) and baked into the golden; there is no host systemd unit. It is still NOT embedded in the agent, so the data path survives the agent's death — and the agent neither manages it nor (see R-841) sees its health. Its version moves only by a controller release (the pin), on the monthly re-test (`runbooks/monthly-floating-retest.md` "Infrastructure pins").
- ~~**`cloudflared` is a separate systemd service**, not embedded in the agent. … The agent **manages and health-watches** it (see §5) but the tunnel does not live or die with the agent process.~~ **[FACT, corrected 2026-10-04 — `11-os-updates.md` C8, R-838]** `cloudflared` is a **container in the customer guest** (`cloudflare/cloudflared:<pin>`), rendered and kept up by the in-guest controller (`felhom-controller/controller/internal/infra/infra.go`, `internal/stacks/infra.go`) and baked into the golden; there is no host systemd unit. It is still NOT embedded in the agent, so the data path survives the agent's death — and the agent does not manage it; it READS its health (R-841, agent v0.141.0). Its version moves only by a controller release (the pin), on the monthly re-test (`runbooks/monthly-floating-retest.md` "Infrastructure pins").
## 4. Control model — reconcile + signed destructive ops
@@ -143,7 +143,7 @@ notification) is the control.
**Box-initiated poll.** The hub never connects inbound. Each poll cycle exchanges:
- **Up:** heartbeat + a host-domain state report — host CPU/RAM/disk, per-guest up/down + spec, storage-target status (USB connected? NFS/CIFS reachable? PBS reachable?), last backup per target, last restore-test result, `cloudflared` health (a host-unit probe that always reads `inactive` — R-841), agent + controller versions, audit-log tail.
- **Up:** heartbeat + a host-domain state report — host CPU/RAM/disk, per-guest up/down + spec, storage-target status (USB connected? NFS/CIFS reachable? PBS reachable?), last backup per target, last restore-test result, `cloudflared` health (`running` / `not_running` / `unknown` from the guest container's readiness check — R-841), agent + controller versions, audit-log tail.
- **Down:** the current desired state, any pending signed one-shot jobs, and config (poll interval, update window, policy changes).
**Dead-man's-switch (essential, not optional).** In a box-initiated model the heartbeat
@@ -366,7 +366,10 @@ never moved by an update or an undo, and the unit restore accepts it.
whole-guest backup (the controller drives it, inside [W+2h, W+6h)) ends SUCCESSFULLY on the primary tier, the agent
waits 90 s and runs the guest's Debian fast lane — still holding the host-wide heavy-op gate, so it never overlaps a
backup or a restore-test; at most once per 20 h. The backup minutes old is the guest's undo (no snapshot is possible,
R-837). A failed or missed backup → no OS leg that night.
R-837). A failed or missed backup → no OS leg that night. **[FACT, 2026-10-04 — agent v0.141.1, `11` §8.2]** On an
appliance, the HOST step follows under the same gate: after a healthy guest step only (a failed or unhealthy guest
step skips it), Debian-origin fixes only, never a kernel, boot or firmware package, never a reboot. Measured: both
steps with nothing to install, 23–32 s; a 108-package host pass, 70 s.
**[FACT]** The three nightly legs derive from **one** customer-settable window start W at fixed
offsets — db-dump at W, Tier-2 at W+60m, offsite at W+105m — so they can never be misordered
@@ -310,6 +310,32 @@ message, not a wider cooldown.
---
## 6.3 Box alarms outside the app ladder: the tunnel and OS updates [DESIGN, hub v0.131.0, 2026-10-04]
These are **operator-only** (the household can act on none of them; `operatorOnlyEvents`, pinned by
`TestOSUpdateEvents_OperatorOnlyExceptApplied`). They go through the same dispatcher and severity contract (§6.1):
`info` is recorded and never mailed; `warning` and `error` are mailed.
| Event | Severity | Raised when | Cleared | Pinned by |
|---|---|---|---|---|
| `tunnel_down` | error | the box's two newest host reports say the tunnel is `not_running` (more than one report cycle, 15 min) and the one before did not | the first `running` after it → `tunnel_recovered` (info) | `api/tunnel_test.go` |
| `os_update_stale` | warning | no successful OS leg for **7 days** while the switch is ON (agents that can run the leg only); the mail names the likely reason (box not reporting / the last leg's failure / no good night backup) | a successful leg | `TestAlarm_StaleLeg`, `TestAlarm_StaleNamesTheReason` |
| `os_reboot_needed` | warning | the host has needed a reboot for **14 days** (from the FIRST scanned report that said so) | a scanned pass that finds nothing (agent ≥ 0.141.1 scans the host every pass) | `TestAlarm_RebootNeeded`, `TestRebootNeeded_ClearedByAScannedPass` |
| `os_ring0_stalled` | error | ring 0 approved nothing for **7 days** in a layer while it has pending FAST-lane updates (a pending kernel does not count) | a new release | `TestAlarm_Ring0Stalled` |
| `os_not_covered` | warning | a ring-1 box has had fast-lane packages no approved release names for **14 days** | the packages are covered or gone | `TestAlarm_NotCovered` |
- **`unknown` never alarms** (R-96 rule 3): a probe that could not ask is neither up nor down. An `unknown` report
breaks a `not_running` run.
- **A stopped cloudflared heals itself before the hub can see it** (measured 2026-10-04): the controller's
protected-container check recreates it within 5 minutes, and the host reports every 15. So `tunnel_down` catches
what the box cannot heal — a running container with no connection (wrong token, blocked network).
- The OS alarms are checked **hourly**, re-sent at most **once a week** while true, and forgotten when false, so the
next occurrence is announced again. The four numbers are configuration (`OS_ALARM_STALE_AFTER`,
`OS_ALARM_REBOOT_AFTER`, `OS_ALARM_RING0_STALL_AFTER`, `OS_ALARM_NOT_COVERED_AFTER`) — *decided by CC unattended,
operator may reverse* (`11` §8.3).
---
## 7. The intent test [DESIGN, R-386 — CLOSED controller v0.223.0]
**"The customer stopped this" is asked of the FIELD THAT RECORDS IT, never inferred from the state.**
@@ -778,6 +778,22 @@ its length, and both fixes cost something the household would notice — operato
83. **Next: the tunnel status (R-841), the host fast lane (`11` §8 step 3), and other OS-update improvements.**
*Operator ruling 2026-10-04 ~12:20.*
### 2026-10-04 (afternoon) — decided by CC unattended — operator may reverse (host fast lane brief)
84. **Which record proves a box is an appliance, for the host fast lane?** Options: (a) `agent.json`
`deployment_mode` — the agent can write it, so a compromised agent could claim "appliance" and unlock host
updates; (b) the installer's ROOT-owned `/var/lib/felhom-install/state.json` `mode` — written once, as root, at
install. **Chosen (b):** the wrapper is the root fence and must not trust a file the agent can change. Cost: a box
whose install record is missing gets no host step (fails closed). `11` §8.2.
85. **The four OS alarm thresholds.** Options: shorter (3/7 days — noisy: a weekend away alarms) or longer (14/30 —
a stopped box goes unseen for weeks). **Chosen:** no OS leg 7 days, reboot needed 14, ring 0 stalled 7, not covered
14; all four are configuration. Cost: a real stop is seen after a week, not a night. `11` §8.3, `08` §6.3.
86. **A second release of the agent (v0.141.1) and the hub (v0.131.1) in the same session**, against "one release per
repo". Options: (a) keep v0.141.0 / v0.131.0 and file the defect — the host "reboot needed" stays wrong (it hid
`lxc-start`, and a reboot never cleared it), so the new 14-day alarm would fire on rebooted hosts; (b) patch now.
**Chosen (b):** a known-false operator alarm is worse than an extra release; both patches are small and red-proved.
R-846.
### 2026-09-30 (day) — operator notes, recorded before the work
- **The day brief runs by day.** Every backup and automatic-update test is started by hand — the night chain's debug
+70 -4
View File
@@ -324,8 +324,8 @@ must never overlap a backup, a restore-test or a self-update.~~
|---|---|---|
| Guest packages | Restore the guest snapshot taken just before the update | It also undoes app data written after the snapshot. Use it only inside the health window, before apps have written much. After that, install the previous version (needs OPEN Q1). **[FACT] (C6)** Customer guests are on LVM-thin and can snapshot; scratch 9202 (`dir` storage) cannot (`snapshot feature is not available`), so the measured undo was a whole-guest backup + restore: **73 s down**, all apps healthy, libc back. The snapshot rollback itself is unmeasured (R-837). |
| Docker engine | Install the previous version (Docker's repository keeps old versions — OPEN Q2) | Every container restarts again. **[FACT] (C5)** Docker keeps 46 `docker-ce` versions. A step (either way) without `live-restore`: 6 of 6 containers restart, the app silent **26.5–30 s**, healthy at +42–45 s. With `live-restore` on: **0 restarts, no gap**, also across a containerd step. But a restart that turns `live-restore` OFF stops every container and starts NONE (`unless-stopped` ignored) — on 9202 they stayed down until restarted by hand; `systemctl reload` turns it on but not off. |
| Host packages | Install the previous version | Only if the source still has it (OPEN Q1/Q2). **[FACT]** `rsync` back to its pre-update version: `not found`; `libpng16-16t64` back to the point-release version: worked. A Debian undo needs the snapshot archive (C2). |
| Host kernel | ~~Boot the previous kernel. Proxmox can boot a new kernel **once** (`proxmox-boot-tool kernel pin <ver> --next-boot`). If that boot fails, the next boot uses the old kernel again. Make the new kernel permanent only after a healthy boot.~~ **[FACT] (C4)** Both demo hosts boot UEFI + GRUB (no proxmox-boot-tool ESPs). **Installing a kernel makes it the GRUB default at once.** `--next-boot` on GRUB writes an ordinary `GRUB_DEFAULT` + `update-grub`; `proxmox-boot-cleanup.service` clears it only once a boot reaches userspace. Measured on demo-hp (old kernel pinned permanently FIRST, new pinned for next boot): boot 1 → `7.0.14-20-pve`, healthy, 60 s; boot 2 → back on `7.0.2-6-pve`, healthy, 76 s. **So the fallback works after a boot that succeeds; read from the code, a kernel that hangs before userspace stays the default on every power cycle.** `[PROPOSAL]` use GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel as the saved default — unmeasured (R-836). | ~~If the new kernel hangs, someone must switch the box off and on. The spike checks whether a hardware watchdog can do that (OPEN Q4).~~ **[FACT]** Only `softdog` runs (loaded by `watchdog-mux`); it cannot rescue a kernel that never boots. demo-hp has an AMD FCH whose `sp5100_tco` driver ships but is not loaded; untested. A hang still needs a person. |
| Host packages | Install the previous version | Only if the source still has it (OPEN Q1/Q2). **[FACT]** `rsync` back to its pre-update version: `not found`; `libpng16-16t64` back to the point-release version: worked. A Debian undo needs the snapshot archive (C2). **[FACT, 2026-10-04]** Proved by hand: `runbooks/os-updates-host-undo.md` (demo-hp, `tzdata` back one version from `snapshot.debian.org`, held, released). Use the NEW version's `first_seen` as the timestamp when there is no previous host release; a package that pins its siblings (`eject` → `libmount1 =`) goes back only with them. |
| Host kernel | ~~Boot the previous kernel. Proxmox can boot a new kernel **once** (`proxmox-boot-tool kernel pin <ver> --next-boot`). If that boot fails, the next boot uses the old kernel again. Make the new kernel permanent only after a healthy boot.~~ **[FACT] (C4)** Both demo hosts boot UEFI + GRUB (no proxmox-boot-tool ESPs). **Installing a kernel makes it the GRUB default at once.** `--next-boot` on GRUB writes an ordinary `GRUB_DEFAULT` + `update-grub`; `proxmox-boot-cleanup.service` clears it only once a boot reaches userspace. Measured on demo-hp (old kernel pinned permanently FIRST, new pinned for next boot): boot 1 → `7.0.14-20-pve`, healthy, 60 s; boot 2 → back on `7.0.2-6-pve`, healthy, 76 s. **So the fallback works after a boot that succeeds; read from the code, a kernel that hangs before userspace stays the default on every power cycle.** ~~`[PROPOSAL]` use GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel as the saved default — unmeasured (R-836).~~ **[FACT, 2026-10-04, demo-hp, operator's word before each reboot] GRUB's one-shot is NOT a one-shot here either.** With `GRUB_DEFAULT=saved` (old 7.0.2-6 saved) and `grub-reboot` 7.0.14-20: boot 1 → 7.0.14-20, **Secure Boot ON, booted fine** (signed kernel, shim → GRUB); but `/boot` is ext4 on LVM, GRUB cannot write its environment block there (`grub-reboot` warns so itself), `next_entry` was never cleared, and boot 2 with no command → **7.0.14-20 again**. A kernel lane needs a writable env block (the ESP) or a userspace "boot good" step — R-836. demo-hp left on 7.0.14-20, saved default 7.0.14-20, both kernels installed (`audits/os-host-lane-2026-10-04/partE/`). | ~~If the new kernel hangs, someone must switch the box off and on. The spike checks whether a hardware watchdog can do that (OPEN Q4).~~ **[FACT]** Only `softdog` runs (loaded by `watchdog-mux`); it cannot rescue a kernel that never boots. demo-hp has an AMD FCH whose `sp5100_tco` driver ships but is not loaded; untested. A hang still needs a person. **[FACT, 2026-10-04]** `sp5100_tco` is blacklisted by the Proxmox kernel package; loaded by hand it answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0 — read from sysfs, never opened, so never armed; unloaded). `kernel.panic = 0`: a panic leaves the host stopped (R-851). |
### 5.7 Telling people
@@ -336,6 +336,30 @@ must never overlap a backup, a restore-test or a self-update.~~
whether the box restarted. Telling households in advance that the box may restart at night is a
**promise to users**. That is the operator's decision when the slow lane is built.
### 5.8 The Docker engine slow lane — DESIGN (2026-10-04, nothing built) `[PROPOSAL]`
Built from C5 (measured) and R-835. **The one decision it needs is in STATUS: `live-restore` on, fleet-wide.**
- **What moves:** `docker-ce`, `docker-ce-cli`, `containerd.io`, `docker-buildx-plugin`, `docker-compose-plugin`,
`docker-ce-rootless-extras` in the customer guest — today the six "not covered" packages on every box. One
approved **engine set** at a time, like a controller floor: the operator approves it (slow lane, §5.2), after ring 0
has run it for at least 2 nights healthy. Never two steps in one night.
- **Precondition: `live-restore` ON.** Without it an engine step restarts every container — 26.5–30 s of silence,
healthy at +42–45 s (C5). With it: 0 restarts, no gap, also across a containerd step. Turning it ON is safe
(`systemctl reload docker` applies it without a restart, C5); turning it OFF later by a plain restart stops every
container and starts none (R-835) — so it is turned on once, by the golden and by a one-time fleet step, and never
turned off by the lane.
- **Who and when:** the agent, through the same wrapper (`lane: slow`, refusal R3: only inside a verified signed
operator job, R-530's mechanism), in the guest, after the guest and host fast-lane steps, under the same heavy-op
gate, on a night the operator scheduled. Debian origin rule replaced by "origin `Docker CE`, exactly these names".
- **Health:** the guest rule (§8.1) plus `docker version` reports the approved engine, and every container running
at the start is running with the SAME container id (proof that `live-restore` held). A changed id is
`health_failed` even if the app is healthy — it means the households' apps restarted when they should not have.
- **Undo:** install the previous engine set (Docker's repository keeps 46 versions, C2) — by an operator job, with
`live-restore` still on, so the undo is also restart-free.
- **Not covered here:** the golden's own engine (baked weekly; a new golden carries the approved set), and BYO hosts
(the guest is ours on both, so the lane applies there too).
---
## 6. Risks and edge cases
@@ -405,8 +429,9 @@ Each step returns to the operator for go or no-go.
1. **Spike** (measure Q1–Q10; no product code).
2. **Guest Debian, fast lane.** ~~Lowest risk: a snapshot undo exists.~~ **BUILT 2026-10-04** — agent v0.140.0, hub
v0.130.0, installer 1.29.0; §8.1. **There is no snapshot undo** (R-837, measured).
3. **Host Debian, fast lane** (no kernel, no Proxmox packages).
4. **Fleet view and alarms** (§5.7).
3. **Host Debian, fast lane** (no kernel, no Proxmox packages). **BUILT 2026-10-04** — agent v0.141.1, hub v0.131.1;
§8.2.
4. **Fleet view and alarms** (§5.7). **BUILT 2026-10-04** — hub v0.131.0/v0.131.1; §8.3.
5. **Slow lane: Docker engine.**
6. **Slow lane: host kernel and Proxmox packages, with the reboot.**
7. **Later:** the Proxmox major upgrade (PVE 9 → 10), drilled on ring 0 first.
@@ -454,6 +479,47 @@ Evidence: `audits/os-guest-lane-2026-10-04/` (parts A–G). Brief: guest fast la
- **Not delivered to existing boxes by the product**: the wrapper and the sudoers line reach a box only through the
installer; the signed agent update replaces the binary only (R-840). The demo boxes got them BY HAND.
### 8.2 Step 3 as BUILT (2026-10-04) `[FACT]`
Evidence: `audits/os-host-lane-2026-10-04/` (parts A–G).
- **Where it runs:** appliances only. The proof is the ROOT-owned install record `/var/lib/felhom-install/state.json`
`mode: appliance` (written by the installer as root); the agent-writable `agent.json` `deployment_mode` is not
trusted for this. A BYO host gets no host step (wrapper refusal **R12**, now lifted only for lane fast / layer host
on an appliance). *Decided by CC unattended — operator may reverse.*
- **What:** origin `Debian` / `Debian-Security` only, and never a kernel, boot or firmware package (name pattern
`HOST_SLOW_RE` — `linux-*`, `proxmox-kernel*`, `pve-kernel*`, `pve-firmware`, `firmware-*`, `grub*`, `shim*`,
`systemd-boot*`, `*-microcode`, `efibootmgr`; refusal **R14**). The hub leaves the same names out of the host
candidate.
- **When:** in the same leg, after the guest step, under the same heavy-op gate. A failed or unhealthy guest step
skips the host step.
- **The host health rule** (`HostHealthVerdict`, pinned in `internal/osupdate`): `felhom-agent`, `pveproxy`,
`pvedaemon`, `pvestatd` and `pve-cluster` are `active`; the customer guest runs; the guest health rule (§8.1)
passes; the tunnel is `running` — an `unknown` tunnel does not fail it, `not_running` does. Same 5-minute wait.
- **Separate approved sets:** host and guest releases are separate (`os-host-…`, `os-guest-…`), each by the same rule
(24 h, 1 night of THAT layer, every ring-0 box). Ring 1 receives `host_release` beside `release`.
- **Reboot needed:** reported when PID 1 or `lxc-start` maps a replaced file, with the date of the first scanned
report that said so. **Never reboots.** The host is scanned on every pass, so a reboot clears it (v0.141.1 — v0.141.0
hid `lxc-start` and never cleared, R-846).
- **Undo:** by hand, `runbooks/os-updates-host-undo.md` (proved). No automatic undo.
- **Speed (R-845):** one call per layer; both steps with nothing to install 23–32 s; a 108-package host pass 70 s.
- **Measured live:** demo-felhom ring 0 installed 108 Debian host packages, all Debian origin (checked against apt),
healthy; a 605-package host release approved (TEST wait 2 min / 0 nights, logged, reverted to 24 h + 1 night);
demo-felhom as ring 1 installed exactly the one version it lacked (604 already current).
### 8.3 Step 4 as BUILT (2026-10-04) `[FACT]`
- **The fleet view** (`GET /os/fleet`, operator): one line per box — ring, switch, the tunnel, and per layer: the
release, the last outcome, the last successful leg, pending, not covered, restart needed, reboot needed since, the
wrapper's own seconds.
- **The four alarms** (`08` §6.3), operator-only, hourly, at most weekly while true: no successful OS leg for
**7 days** while the switch is ON (naming the likely reason); reboot needed for **14 days**; ring 0 approved nothing
for **7 days** while it has pending fast-lane updates; not-covered fast-lane packages for **14 days**. The four
numbers are configuration (`OS_ALARM_*`). *Decided by CC unattended — operator may reverse:* 7 days = a week of
missed nights is past any normal hiccup (a box off for a weekend does not alarm); 14 days for reboot and coverage =
two weekly golden cycles, both need a person anyway.
- **The tunnel** (R-841): `running` / `not_running` / `unknown`; `tunnel_down` after two `not_running` reports.
## 9. Where the rest lives
- The finding: **R-812** (`backlog/OPEN-ITEMS.md`). The intention: **R-808** (`backlog/ROADMAP.md`).
@@ -3,3 +3,17 @@
11:36:47 container=running/unhealthy hub=running alarm=none
11:37:49 container=running/unhealthy hub=not_running alarm=none
11:38:51 container=running/unhealthy hub=not_running alarm=none
11:39:53 container=running/unhealthy hub=not_running alarm=none
11:40:55 container=running/unhealthy hub=not_running alarm=none
11:41:57 container=running/unhealthy hub=not_running alarm=none
11:42:59 container=running/unhealthy hub=not_running alarm=none
11:44:01 container=running/unhealthy hub=not_running alarm=none
11:45:03 container=running/unhealthy hub=not_running alarm=none
11:46:05 container=running/unhealthy hub=not_running alarm=none
11:47:07 container=running/unhealthy hub=not_running alarm=none
11:48:09 container=running/unhealthy hub=not_running alarm=none
11:49:10 container=running/unhealthy hub=not_running alarm=none
11:50:12 container=running/unhealthy hub=not_running alarm=none
11:51:14 container=running/unhealthy hub=not_running alarm=none
11:52:16 container=running/unhealthy hub=not_running alarm=none
11:53:18 container=running/unhealthy hub=not_running alarm=2026/10/04 13:52:29 [WARN] host demo-hp-bb76ea tunnel: tunnel_down (container running but the tunnel is NOT connected (cloudflared /ready fails))
@@ -3,3 +3,6 @@ Chain DOCKER-USER (1 references)
target prot opt source destination
DROP tcp -- 0.0.0.0/0 0.0.0.0/0 tcp dpt:7844
DROP udp -- 0.0.0.0/0 0.0.0.0/0 udp dpt:7844
unblock at 2026-10-04T11:53:29Z
Chain DOCKER-USER (1 references)
target prot opt source destination
@@ -0,0 +1,14 @@
11:53:39 container=unhealthy hub=not_running
11:54:41 container=healthy hub=not_running
11:55:43 container=healthy hub=not_running
11:56:45 container=healthy hub=not_running
11:57:47 container=healthy hub=not_running
11:58:48 container=healthy hub=not_running
11:59:50 container=healthy hub=not_running
12:00:52 container=healthy hub=not_running
12:01:54 container=healthy hub=not_running
12:02:56 container=healthy hub=not_running
12:03:58 container=healthy hub=not_running
12:05:00 container=healthy hub=not_running
12:06:02 container=healthy hub=not_running
12:07:04 container=healthy hub=not_running
@@ -0,0 +1,5 @@
2026/10/04 13:52:29 [WARN] host demo-hp-bb76ea tunnel: tunnel_down (container running but the tunnel is NOT connected (cloudflared /ready fails))
2026/10/04 13:52:30 [INFO] Operator email sent for demo-hp/tunnel_down
2026/10/04 13:52:29 [WARN] host demo-hp-bb76ea tunnel: tunnel_down (container running but the tunnel is NOT connected (cloudflared /ready fails))
2026/10/04 13:52:30 [INFO] Operator email sent for demo-hp/tunnel_down
2026/10/04 14:07:29 [WARN] host demo-hp-bb76ea tunnel: tunnel_recovered (connected)
@@ -0,0 +1,135 @@
=== felhom-agent 0.141.1 selftest=os-update vmid=9201 ring=0 enabled=true guest-release=false host-release=false appliance=true ===
time=2026-10-04T14:22:52.476+02:00 level=INFO msg="osupdate: START" run=20261004T122252Z layer=guest vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T122252Z
time=2026-10-04T14:23:08.079+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T122252Z layer=guest:9201 lane=fast mode=apply select=pending-fast packages=0"
time=2026-10-04T14:23:08.079+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
time=2026-10-04T14:23:08.079+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0"
time=2026-10-04T14:23:08.079+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)"
time=2026-10-04T14:23:08.080+02:00 level=INFO msg="osupdate: DONE" run=20261004T122252Z layer=guest vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=6 not_covered=6 restart_needed=0 reboot_needed=false wrapper_seconds=15.5
time=2026-10-04T14:23:08.089+02:00 level=INFO msg="osupdate: START" run=20261004T122252Z layer=host vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T122252Z
time=2026-10-04T14:23:22.788+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T122252Z layer=host lane=fast mode=apply select=pending-fast packages=0"
time=2026-10-04T14:23:22.788+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
time=2026-10-04T14:23:22.788+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0"
time=2026-10-04T14:23:22.788+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)"
time=2026-10-04T14:23:23.852+02:00 level=INFO msg="osupdate: DONE" run=20261004T122252Z layer=host vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=78 not_covered=78 restart_needed=0 reboot_needed=false wrapper_seconds=14.6
--- os-update report (guest) ---
{
"health_reason": "",
"healthy": true,
"mode": "apply",
"not_covered": [
"docker-ce-cli",
"containerd.io",
"docker-ce",
"docker-buildx-plugin",
"docker-ce-rootless-extras",
"docker-compose-plugin"
],
"outcome": "nothing",
"pending": 6,
"reboot_needed": false,
"refused": null,
"release_id": "ring0-20261004T122252Z",
"restart_needed": null,
"ring": 0,
"run_id": "20261004T122252Z",
"upgraded": [],
"wrapper_seconds": 15.5
}
--- os-update report (host) ---
{
"health_reason": "",
"healthy": true,
"mode": "apply",
"not_covered": [
"frr",
"shim-signed-common",
"proxmox-secure-boot-support",
"shim-unsigned",
"shim-helpers-amd64-signed",
"shim-signed",
"amd64-microcode",
"libradosstriper1",
"librgw2",
"ceph-common",
"librbd1",
"librados2",
"python3-cephfs",
"libcephfs2",
"python3-rgw",
"python3-rados",
"python3-ceph-argparse",
"python3-ceph-common",
"python3-rbd",
"ceph-fuse",
"chrony",
"libcorosync-common4",
"libcfg7",
"libcmap4",
"libcpg4",
"libknet1t64",
"libnozzle1t64",
"libquorum5",
"libvotequorum8",
"corosync",
"frr-pythontools",
"libjs-extjs",
"libnvpair3linux",
"libproxmox-acme-plugins",
"libproxmox-backup-qemu0",
"pve-qemu-kvm",
"libpve-notify-perl",
"libpve-cluster-api-perl",
"libpve-cluster-perl",
"pve-cluster",
"libpve-access-control",
"libpve-apiclient-perl",
"librados2-perl",
"proxmox-backup-client",
"proxmox-backup-file-restore",
"pve-manager",
"libproxmox-acme-perl",
"libpve-common-perl",
"libpve-guest-common-perl",
"qemu-server",
"libpve-storage-perl",
"pve-edk2-firmware-legacy",
"pve-edk2-firmware-ovmf",
"libpve-network-api-perl",
"libpve-network-perl",
"proxmox-firewall-data",
"pve-firewall",
"pve-container",
"pve-ha-manager",
"novnc-pve",
"proxmox-enterprise-support-keyring",
"proxmox-mini-journalreader",
"proxmox-widget-toolkit",
"pve-docs",
"pve-i18n",
"pve-xtermjs",
"pve-yew-mobile-i18n",
"pve-yew-mobile-gui",
"libuutil3linux",
"libzfs7linux",
"libzpool7linux",
"proxmox-kernel-helper",
"pve-edk2-firmware-aarch64",
"pve-edk2-firmware",
"pve-firmware",
"zfs-initramfs",
"zfsutils-linux",
"zfs-zed"
],
"outcome": "nothing",
"pending": 78,
"reboot_needed": false,
"refused": null,
"release_id": "ring0-20261004T122252Z",
"restart_needed": [],
"ring": 0,
"run_id": "20261004T122252Z",
"upgraded": [],
"wrapper_seconds": 14.6
}
pass took 31.4s
WALL_SECONDS=31.520457543
@@ -0,0 +1,2 @@
{"ok":true}
200
@@ -0,0 +1,43 @@
+ PKG=tzdata
++ grep -oE 'tzdata:amd64 \([^,]+' /var/log/apt/history.log
++ tail -1
++ sed 's/.*(//'
+ OLD=2026b-0+deb13u1
++ dpkg-query -W '-f=${Version}' tzdata
+ NEW=2026c-0+deb13u1
+ echo OLD=2026b-0+deb13u1 NEW=2026c-0+deb13u1
OLD=2026b-0+deb13u1 NEW=2026c-0+deb13u1
++ curl -s 'https://snapshot.debian.org/mr/binary/tzdata/2026c-0+deb13u1/binfiles?fileinfo=1'
++ python3 -c '
import json,sys; d=json.load(sys.stdin)
print(sorted(f["first_seen"] for v in d["fileinfo"].values() for f in v)[0])'
+ TS=20260831T204404Z
+ . /etc/os-release
++ PRETTY_NAME='Debian GNU/Linux 13 (trixie)'
++ NAME='Debian GNU/Linux'
++ VERSION_ID=13
++ VERSION='13 (trixie)'
++ VERSION_CODENAME=trixie
++ DEBIAN_VERSION_FULL=13.7
++ ID=debian
++ HOME_URL=https://www.debian.org/
++ SUPPORT_URL=https://www.debian.org/support
++ BUG_REPORT_URL=https://bugs.debian.org/
+ printf 'deb [check-valid-until=no] http://snapshot.debian.org/archive/debian/%s %s main\ndeb [check-valid-until=no] http://snapshot.debian.org/archive/debian-security/%s %s-security main\n' 20260831T204404Z trixie 20260831T204404Z trixie
+ apt-get -q update
+ apt-get -s install --allow-downgrades tzdata=2026b-0+deb13u1
+ grep -E '^(Inst|Remv)|downgraded'
0 upgraded, 0 newly installed, 1 downgraded, 0 to remove and 78 not upgraded.
Inst tzdata [2026c-0+deb13u1] (2026b-0+deb13u1 Debian:13.6/stable [all])
+ DEBIAN_FRONTEND=noninteractive
+ apt-get -y -q install --allow-downgrades tzdata=2026b-0+deb13u1
+ rm -f /etc/apt/sources.list.d/felhom-undo-snapshot.list
+ apt-get -q update
+ dpkg-query -W tzdata
tzdata 2026b-0+deb13u1
+ ls /etc/apt/sources.list.d/
ceph.sources
debian.sources
pve-enterprise.sources
pve-no-subscription.sources
tailscale.list
@@ -0,0 +1,178 @@
=== felhom-agent 0.141.1 selftest=os-update vmid=9201 ring=1 enabled=true guest-release=true host-release=true appliance=true ===
time=2026-10-04T14:40:46.795+02:00 level=INFO msg="osupdate: START" run=20261004T124046Z layer=guest vmid=9201 ring=1 trigger=debug enabled=true release=os-guest-20261004-123933
time=2026-10-04T14:40:56.895+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=os-guest-20261004-123933 layer=guest:9201 lane=fast mode=apply select=listed packages=272"
time=2026-10-04T14:40:56.895+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
time=2026-10-04T14:40:56.895+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=272 not-installed=0 from-snapshot=0"
time=2026-10-04T14:40:56.895+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)"
time=2026-10-04T14:40:56.896+02:00 level=INFO msg="osupdate: DONE" run=20261004T124046Z layer=guest vmid=9201 ring=1 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=6 not_covered=6 restart_needed=0 reboot_needed=false wrapper_seconds=10.1
time=2026-10-04T14:40:56.907+02:00 level=INFO msg="osupdate: START" run=20261004T124046Z layer=host vmid=9201 ring=1 trigger=debug enabled=true release=os-host-20261004-124034
time=2026-10-04T14:41:10.034+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=os-host-20261004-124034 layer=host lane=fast mode=apply select=listed packages=605"
time=2026-10-04T14:41:10.034+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
time=2026-10-04T14:41:10.034+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=1 already=604 not-installed=0 from-snapshot=0"
time=2026-10-04T14:41:10.034+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=1.4 upgraded=1 restart-needed=agetty,blkmapd,chronyd,cron,dbus-daemon,dmeventd,ksmtuned,lxc-monitord,lxc-start,lxcfs,pmxcfs,proxmox-firewal,pve-firewall,pve-ha-crm,pve-ha-lrm,pve-lxc-syscall,pvedaemon,pvedaemon worke,pvefw-logger,pveproxy,pveproxy worker,pvescheduler,pvestatd,qmeventd,rpcbind,rrdcached,smartd,spiceproxy,spiceproxy work,sshd,systemd-logind,systemd-udevd,watchdog-mux,zed reboot-needed=yes"
time=2026-10-04T14:41:10.815+02:00 level=INFO msg="osupdate: DONE" run=20261004T124046Z layer=host vmid=9201 ring=1 trigger=debug outcome=applied healthy=true reason="" upgraded=1 pending=80 not_covered=80 restart_needed=34 reboot_needed=true wrapper_seconds=13.1
--- os-update report (guest) ---
{
"health_reason": "",
"healthy": true,
"mode": "apply",
"not_covered": [
"docker-ce-cli",
"containerd.io",
"docker-ce",
"docker-buildx-plugin",
"docker-ce-rootless-extras",
"docker-compose-plugin"
],
"outcome": "nothing",
"pending": 6,
"reboot_needed": false,
"refused": null,
"release_id": "os-guest-20261004-123933",
"restart_needed": null,
"ring": 1,
"run_id": "20261004T124046Z",
"upgraded": [],
"wrapper_seconds": 10.1
}
--- os-update report (host) ---
{
"health_reason": "",
"healthy": true,
"mode": "apply",
"not_covered": [
"frr",
"shim-signed-common",
"shim-unsigned",
"shim-helpers-amd64-signed",
"shim-signed",
"libradosstriper1",
"librgw2",
"ceph-common",
"librbd1",
"librados2",
"python3-cephfs",
"libcephfs2",
"python3-rgw",
"python3-rados",
"python3-ceph-argparse",
"python3-ceph-common",
"python3-rbd",
"ceph-fuse",
"chrony",
"libcorosync-common4",
"libcfg7",
"libcmap4",
"libcpg4",
"libknet1t64",
"libnozzle1t64",
"libquorum5",
"libvotequorum8",
"corosync",
"frr-pythontools",
"libjs-extjs",
"libnvpair3linux",
"libproxmox-acme-plugins",
"libproxmox-backup-qemu0",
"pve-qemu-kvm",
"libpve-notify-perl",
"libpve-cluster-api-perl",
"libpve-cluster-perl",
"pve-cluster",
"libpve-access-control",
"libpve-apiclient-perl",
"librados2-perl",
"proxmox-backup-client",
"proxmox-backup-file-restore",
"pve-manager",
"libproxmox-acme-perl",
"libpve-common-perl",
"libpve-guest-common-perl",
"qemu-server",
"libpve-storage-perl",
"pve-edk2-firmware-legacy",
"pve-edk2-firmware-ovmf",
"libpve-network-api-perl",
"libpve-network-perl",
"proxmox-firewall-data",
"pve-firewall",
"pve-container",
"pve-ha-manager",
"novnc-pve",
"proxmox-enterprise-support-keyring",
"proxmox-mini-journalreader",
"proxmox-widget-toolkit",
"pve-docs",
"pve-i18n",
"pve-xtermjs",
"pve-yew-mobile-i18n",
"pve-yew-mobile-gui",
"libuutil3linux",
"libzfs7linux",
"libzpool7linux",
"proxmox-first-boot",
"pve-firmware",
"proxmox-kernel-7.0.14-20-pve-signed",
"proxmox-kernel-7.0",
"proxmox-kernel-helper",
"pve-edk2-firmware-aarch64",
"pve-edk2-firmware",
"zfs-initramfs",
"zfsutils-linux",
"zfs-zed",
"tailscale"
],
"outcome": "applied",
"pending": 80,
"reboot_needed": true,
"refused": null,
"release_id": "os-host-20261004-124034",
"restart_needed": [
"agetty",
"blkmapd",
"chronyd",
"cron",
"dbus-daemon",
"dmeventd",
"ksmtuned",
"lxc-monitord",
"lxc-start",
"lxcfs",
"pmxcfs",
"proxmox-firewal",
"pve-firewall",
"pve-ha-crm",
"pve-ha-lrm",
"pve-lxc-syscall",
"pvedaemon",
"pvedaemon worke",
"pvefw-logger",
"pveproxy",
"pveproxy worker",
"pvescheduler",
"pvestatd",
"qmeventd",
"rpcbind",
"rrdcached",
"smartd",
"spiceproxy",
"spiceproxy work",
"sshd",
"systemd-logind",
"systemd-udevd",
"watchdog-mux",
"zed"
],
"ring": 1,
"run_id": "20261004T124046Z",
"upgraded": [
{
"name": "tzdata",
"version": "2026c-0+deb13u1",
"origin": ""
}
],
"wrapper_seconds": 13.1
}
pass took 24.1s
tzdata 2026c-0+deb13u1
@@ -0,0 +1,11 @@
+ PKG=eject
+ OLD=2.41-5
+ dpkg -s eject
+ grep -E '^(Status|Version)'
Status: install ok installed
Version: 2.41.5-0+deb13u1
+ curl -s 'https://snapshot.debian.org/mr/binary/eject/2.41-5/binfiles?fileinfo=1'
+ python3 -c '
import json,sys; d=json.load(sys.stdin)
print(sorted(f["first_seen"] for v in d["fileinfo"].values() for f in v)[0])'
20250510T015204Z
@@ -0,0 +1,28 @@
+ PKG=eject
+ OLD=2.41-5
+ TS=20250510T015204Z
+ . /etc/os-release
++ PRETTY_NAME='Debian GNU/Linux 13 (trixie)'
++ NAME='Debian GNU/Linux'
++ VERSION_ID=13
++ VERSION='13 (trixie)'
++ VERSION_CODENAME=trixie
++ DEBIAN_VERSION_FULL=13.7
++ ID=debian
++ HOME_URL=https://www.debian.org/
++ SUPPORT_URL=https://www.debian.org/support
++ BUG_REPORT_URL=https://bugs.debian.org/
+ cat
++ date +%s
+ S=1791116624
+ apt-get -q update
+ tail -3
Get:7 http://snapshot.debian.org/archive/debian/20250510T015204Z trixie/main amd64 Packages [9681 kB]
Fetched 9904 kB in 5s (2165 kB/s)
Reading package lists...
++ date +%s
update_seconds=6
+ echo update_seconds=6
+ apt-get -s install --allow-downgrades eject=2.41-5
+ grep -E '^(Inst|Remv|Conf)|downgraded|newly|Err|E:'
E: Version '2.41-5' for 'eject' was not found
@@ -0,0 +1,35 @@
+ PKG=eject
+ OLD=2.41-5
++ dpkg-query -W '-f=${Version}' eject
+ NEW=2.41.5-0+deb13u1
++ curl -s 'https://snapshot.debian.org/mr/binary/eject/2.41.5-0+deb13u1/binfiles?fileinfo=1'
++ python3 -c '
import json,sys; d=json.load(sys.stdin)
print(sorted(f["first_seen"] for v in d["fileinfo"].values() for f in v)[0])'
TS=20260814T165831Z
+ TS=20260814T165831Z
+ echo TS=20260814T165831Z
+ . /etc/os-release
++ PRETTY_NAME='Debian GNU/Linux 13 (trixie)'
++ NAME='Debian GNU/Linux'
++ VERSION_ID=13
++ VERSION='13 (trixie)'
++ VERSION_CODENAME=trixie
++ DEBIAN_VERSION_FULL=13.7
++ ID=debian
++ HOME_URL=https://www.debian.org/
++ SUPPORT_URL=https://www.debian.org/support
++ BUG_REPORT_URL=https://bugs.debian.org/
+ cat
+ apt-get -q update
+ tail -1
Reading package lists...
+ apt-cache madison eject
eject | 2.41.5-0+deb13u1 | http://deb.debian.org/debian trixie/main amd64 Packages
eject | 2.41.5-0+deb13u1 | http://security.debian.org/debian-security trixie-security/main amd64 Packages
eject | 2.41.5-0+deb13u1 | http://snapshot.debian.org/archive/debian-security/20260814T165831Z trixie-security/main amd64 Packages
eject | 2.41-5 | http://snapshot.debian.org/archive/debian/20260814T165831Z trixie/main amd64 Packages
+ apt-get -s install --allow-downgrades eject=2.41-5
+ grep -E '^(Inst|Remv|Conf)|downgraded|newly|Err|E:'
E: Unable to correct problems, you have held broken packages.
E: The following information from --solver 3.0 may provide additional context:
@@ -0,0 +1,37 @@
+ rm -f /etc/apt/sources.list.d/felhom-undo-snapshot.list
+ PKG=tzdata
+ OLD=2026b-0+deb13u1
++ dpkg-query -W '-f=${Version}' tzdata
+ NEW=2026c-0+deb13u1
+ echo NEW=2026c-0+deb13u1
NEW=2026c-0+deb13u1
++ curl -s 'https://snapshot.debian.org/mr/binary/tzdata/2026c-0+deb13u1/binfiles?fileinfo=1'
++ python3 -c '
import json,sys; d=json.load(sys.stdin)
print(sorted(f["first_seen"] for v in d["fileinfo"].values() for f in v)[0])'
TS=20260831T204404Z
+ TS=20260831T204404Z
+ echo TS=20260831T204404Z
+ . /etc/os-release
++ PRETTY_NAME='Debian GNU/Linux 13 (trixie)'
++ NAME='Debian GNU/Linux'
++ VERSION_ID=13
++ VERSION='13 (trixie)'
++ VERSION_CODENAME=trixie
++ DEBIAN_VERSION_FULL=13.7
++ ID=debian
++ HOME_URL=https://www.debian.org/
++ SUPPORT_URL=https://www.debian.org/support
++ BUG_REPORT_URL=https://bugs.debian.org/
+ cat
+ apt-get -q update
+ tail -1
Reading package lists...
+ apt-cache madison tzdata
tzdata | 2026c-0+deb13u1 | http://deb.debian.org/debian trixie/main amd64 Packages
tzdata | 2026b-0+deb13u1 | http://snapshot.debian.org/archive/debian/20260831T204404Z trixie/main amd64 Packages
+ apt-get -s install --allow-downgrades tzdata=2026b-0+deb13u1
+ grep -E '^(Inst|Remv|Conf)|downgraded|newly|Err|E:'
0 upgraded, 0 newly installed, 1 downgraded, 0 to remove and 77 not upgraded.
Inst tzdata [2026c-0+deb13u1] (2026b-0+deb13u1 Debian:13.6/stable [all])
Conf tzdata (2026b-0+deb13u1 Debian:13.6/stable [all])
@@ -0,0 +1,38 @@
+ PKG=tzdata
+ OLD=2026b-0+deb13u1
++ date +%s
+ S=1791116670
+ DEBIAN_FRONTEND=noninteractive
+ apt-get -y -q install --allow-downgrades tzdata=2026b-0+deb13u1
+ grep -E '^(Unpacking|Setting up)|downgraded'
0 upgraded, 0 newly installed, 1 downgraded, 0 to remove and 77 not upgraded.
Unpacking tzdata (2026b-0+deb13u1) over (2026c-0+deb13u1) ...
Setting up tzdata (2026b-0+deb13u1) ...
+ apt-mark hold tzdata
tzdata set on hold.
+ rm -f /etc/apt/sources.list.d/felhom-undo-snapshot.list
+ apt-get -q update
+ tail -1
Reading package lists...
++ date +%s
undo_seconds=4
+ echo undo_seconds=4
+ dpkg -s tzdata
+ grep -E '^(Status|Version)'
Status: hold ok installed
Version: 2026b-0+deb13u1
+ apt-mark showhold
tzdata
+ systemctl is-active pveproxy pvedaemon pvestatd pve-cluster felhom-agent
active
active
active
active
active
+ pct status 9201
status: running
+ ls /etc/apt/sources.list.d/
ceph.sources
debian.sources
pve-enterprise.sources
pve-no-subscription.sources
@@ -0,0 +1,135 @@
=== felhom-agent 0.141.1 selftest=os-update vmid=9201 ring=0 enabled=true guest-release=false host-release=false appliance=true ===
time=2026-10-04T14:24:41.910+02:00 level=INFO msg="osupdate: START" run=20261004T122441Z layer=guest vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T122441Z
time=2026-10-04T14:24:57.242+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T122441Z layer=guest:9201 lane=fast mode=apply select=pending-fast packages=0"
time=2026-10-04T14:24:57.242+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
time=2026-10-04T14:24:57.242+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0"
time=2026-10-04T14:24:57.242+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)"
time=2026-10-04T14:24:57.243+02:00 level=INFO msg="osupdate: DONE" run=20261004T122441Z layer=guest vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=6 not_covered=6 restart_needed=0 reboot_needed=false wrapper_seconds=15.3
time=2026-10-04T14:24:57.251+02:00 level=INFO msg="osupdate: START" run=20261004T122441Z layer=host vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T122441Z
time=2026-10-04T14:25:12.127+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T122441Z layer=host lane=fast mode=apply select=pending-fast packages=0"
time=2026-10-04T14:25:12.127+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
time=2026-10-04T14:25:12.127+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0"
time=2026-10-04T14:25:12.127+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)"
time=2026-10-04T14:25:13.136+02:00 level=INFO msg="osupdate: DONE" run=20261004T122441Z layer=host vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=78 not_covered=78 restart_needed=0 reboot_needed=false wrapper_seconds=14.8
--- os-update report (guest) ---
{
"health_reason": "",
"healthy": true,
"mode": "apply",
"not_covered": [
"docker-ce-cli",
"containerd.io",
"docker-ce",
"docker-buildx-plugin",
"docker-ce-rootless-extras",
"docker-compose-plugin"
],
"outcome": "nothing",
"pending": 6,
"reboot_needed": false,
"refused": null,
"release_id": "ring0-20261004T122441Z",
"restart_needed": null,
"ring": 0,
"run_id": "20261004T122441Z",
"upgraded": [],
"wrapper_seconds": 15.3
}
--- os-update report (host) ---
{
"health_reason": "",
"healthy": true,
"mode": "apply",
"not_covered": [
"frr",
"shim-signed-common",
"proxmox-secure-boot-support",
"shim-unsigned",
"shim-helpers-amd64-signed",
"shim-signed",
"amd64-microcode",
"libradosstriper1",
"librgw2",
"ceph-common",
"librbd1",
"librados2",
"python3-cephfs",
"libcephfs2",
"python3-rgw",
"python3-rados",
"python3-ceph-argparse",
"python3-ceph-common",
"python3-rbd",
"ceph-fuse",
"chrony",
"libcorosync-common4",
"libcfg7",
"libcmap4",
"libcpg4",
"libknet1t64",
"libnozzle1t64",
"libquorum5",
"libvotequorum8",
"corosync",
"frr-pythontools",
"libjs-extjs",
"libnvpair3linux",
"libproxmox-acme-plugins",
"libproxmox-backup-qemu0",
"pve-qemu-kvm",
"libpve-notify-perl",
"libpve-cluster-api-perl",
"libpve-cluster-perl",
"pve-cluster",
"libpve-access-control",
"libpve-apiclient-perl",
"librados2-perl",
"proxmox-backup-client",
"proxmox-backup-file-restore",
"pve-manager",
"libproxmox-acme-perl",
"libpve-common-perl",
"libpve-guest-common-perl",
"qemu-server",
"libpve-storage-perl",
"pve-edk2-firmware-legacy",
"pve-edk2-firmware-ovmf",
"libpve-network-api-perl",
"libpve-network-perl",
"proxmox-firewall-data",
"pve-firewall",
"pve-container",
"pve-ha-manager",
"novnc-pve",
"proxmox-enterprise-support-keyring",
"proxmox-mini-journalreader",
"proxmox-widget-toolkit",
"pve-docs",
"pve-i18n",
"pve-xtermjs",
"pve-yew-mobile-i18n",
"pve-yew-mobile-gui",
"libuutil3linux",
"libzfs7linux",
"libzpool7linux",
"proxmox-kernel-helper",
"pve-edk2-firmware-aarch64",
"pve-edk2-firmware",
"pve-firmware",
"zfs-initramfs",
"zfsutils-linux",
"zfs-zed"
],
"outcome": "nothing",
"pending": 78,
"reboot_needed": false,
"refused": null,
"release_id": "ring0-20261004T122441Z",
"restart_needed": [],
"ring": 0,
"run_id": "20261004T122441Z",
"upgraded": [],
"wrapper_seconds": 14.8
}
pass took 31.2s
tzdata 2026b-0+deb13u1
@@ -0,0 +1,143 @@
Canceled hold on tzdata.
=== felhom-agent 0.141.1 selftest=os-update vmid=9201 ring=0 enabled=true guest-release=false host-release=false appliance=true ===
time=2026-10-04T14:25:22.024+02:00 level=INFO msg="osupdate: START" run=20261004T122522Z layer=guest vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T122522Z
time=2026-10-04T14:25:37.515+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T122522Z layer=guest:9201 lane=fast mode=apply select=pending-fast packages=0"
time=2026-10-04T14:25:37.515+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
time=2026-10-04T14:25:37.515+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0"
time=2026-10-04T14:25:37.515+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)"
time=2026-10-04T14:25:37.516+02:00 level=INFO msg="osupdate: DONE" run=20261004T122522Z layer=guest vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=6 not_covered=6 restart_needed=0 reboot_needed=false wrapper_seconds=15.4
time=2026-10-04T14:25:37.526+02:00 level=INFO msg="osupdate: START" run=20261004T122522Z layer=host vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T122522Z
time=2026-10-04T14:25:57.778+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T122522Z layer=host lane=fast mode=apply select=pending-fast packages=0"
time=2026-10-04T14:25:57.778+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
time=2026-10-04T14:25:57.778+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=1 already=0 not-installed=0 from-snapshot=0"
time=2026-10-04T14:25:57.778+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=2.0 upgraded=1 restart-needed=- reboot-needed=no"
time=2026-10-04T14:25:58.809+02:00 level=INFO msg="osupdate: DONE" run=20261004T122522Z layer=host vmid=9201 ring=0 trigger=debug outcome=applied healthy=true reason="" upgraded=1 pending=78 not_covered=78 restart_needed=0 reboot_needed=false wrapper_seconds=20.2
--- os-update report (guest) ---
{
"health_reason": "",
"healthy": true,
"mode": "apply",
"not_covered": [
"docker-ce-cli",
"containerd.io",
"docker-ce",
"docker-buildx-plugin",
"docker-ce-rootless-extras",
"docker-compose-plugin"
],
"outcome": "nothing",
"pending": 6,
"reboot_needed": false,
"refused": null,
"release_id": "ring0-20261004T122522Z",
"restart_needed": null,
"ring": 0,
"run_id": "20261004T122522Z",
"upgraded": [],
"wrapper_seconds": 15.4
}
--- os-update report (host) ---
{
"health_reason": "",
"healthy": true,
"mode": "apply",
"not_covered": [
"frr",
"shim-signed-common",
"proxmox-secure-boot-support",
"shim-unsigned",
"shim-helpers-amd64-signed",
"shim-signed",
"amd64-microcode",
"libradosstriper1",
"librgw2",
"ceph-common",
"librbd1",
"librados2",
"python3-cephfs",
"libcephfs2",
"python3-rgw",
"python3-rados",
"python3-ceph-argparse",
"python3-ceph-common",
"python3-rbd",
"ceph-fuse",
"chrony",
"libcorosync-common4",
"libcfg7",
"libcmap4",
"libcpg4",
"libknet1t64",
"libnozzle1t64",
"libquorum5",
"libvotequorum8",
"corosync",
"frr-pythontools",
"libjs-extjs",
"libnvpair3linux",
"libproxmox-acme-plugins",
"libproxmox-backup-qemu0",
"pve-qemu-kvm",
"libpve-notify-perl",
"libpve-cluster-api-perl",
"libpve-cluster-perl",
"pve-cluster",
"libpve-access-control",
"libpve-apiclient-perl",
"librados2-perl",
"proxmox-backup-client",
"proxmox-backup-file-restore",
"pve-manager",
"libproxmox-acme-perl",
"libpve-common-perl",
"libpve-guest-common-perl",
"qemu-server",
"libpve-storage-perl",
"pve-edk2-firmware-legacy",
"pve-edk2-firmware-ovmf",
"libpve-network-api-perl",
"libpve-network-perl",
"proxmox-firewall-data",
"pve-firewall",
"pve-container",
"pve-ha-manager",
"novnc-pve",
"proxmox-enterprise-support-keyring",
"proxmox-mini-journalreader",
"proxmox-widget-toolkit",
"pve-docs",
"pve-i18n",
"pve-xtermjs",
"pve-yew-mobile-i18n",
"pve-yew-mobile-gui",
"libuutil3linux",
"libzfs7linux",
"libzpool7linux",
"proxmox-kernel-helper",
"pve-edk2-firmware-aarch64",
"pve-edk2-firmware",
"pve-firmware",
"zfs-initramfs",
"zfsutils-linux",
"zfs-zed"
],
"outcome": "applied",
"pending": 78,
"reboot_needed": false,
"refused": null,
"release_id": "ring0-20261004T122522Z",
"restart_needed": [],
"ring": 0,
"run_id": "20261004T122522Z",
"upgraded": [
{
"name": "tzdata",
"version": "2026c-0+deb13u1",
"origin": ""
}
],
"wrapper_seconds": 20.2
}
pass took 36.9s
tzdata 2026c-0+deb13u1
0
@@ -0,0 +1,104 @@
{
"boxes": [
{
"HostID": "demo-felhom-8363b5",
"Ring": 1,
"Enabled": true,
"Tunnel": "running",
"Guest": {
"ReleaseID": "os-guest-20261004-123933",
"LastOutcome": "nothing",
"LastAt": "2026-10-04T12:40:56Z",
"LastSuccessfulLeg": "2026-10-04T12:40:56Z",
"Pending": 6,
"NotCovered": 6,
"NotCoveredFast": 0,
"RestartNeeded": 0,
"RebootNeededSince": "2026-10-04T09:20:04Z",
"WrapperPassSeconds": 10.1
},
"Host": {
"ReleaseID": "os-host-20261004-124034",
"LastOutcome": "applied",
"LastAt": "2026-10-04T12:41:10Z",
"LastSuccessfulLeg": "2026-10-04T12:41:10Z",
"Pending": 80,
"NotCovered": 80,
"NotCoveredFast": 0,
"RestartNeeded": 34,
"RebootNeededSince": "2026-10-04T11:53:20Z",
"WrapperPassSeconds": 13.1
}
},
{
"HostID": "demo-hp-bb76ea",
"Ring": 0,
"Enabled": true,
"Tunnel": "unknown",
"Guest": {
"ReleaseID": "ring0-20261004T122522Z",
"LastOutcome": "nothing",
"LastAt": "2026-10-04T12:25:37Z",
"LastSuccessfulLeg": "2026-10-04T12:25:37Z",
"Pending": 6,
"NotCovered": 6,
"NotCoveredFast": 0,
"RestartNeeded": 0,
"RebootNeededSince": "2026-10-04T09:21:28Z",
"WrapperPassSeconds": 15.4
},
"Host": {
"ReleaseID": "ring0-20261004T122522Z",
"LastOutcome": "applied",
"LastAt": "2026-10-04T12:25:58Z",
"LastSuccessfulLeg": "2026-10-04T12:25:58Z",
"Pending": 78,
"NotCovered": 78,
"NotCoveredFast": 0,
"RestartNeeded": 0,
"RebootNeededSince": "0001-01-01T00:00:00Z",
"WrapperPassSeconds": 20.2
}
},
{
"HostID": "drill-r50-0a4f9a",
"Ring": 1,
"Enabled": true,
"Tunnel": "inactive",
"Guest": {
"ReleaseID": "",
"LastOutcome": "",
"LastAt": "0001-01-01T00:00:00Z",
"LastSuccessfulLeg": "0001-01-01T00:00:00Z",
"Pending": 0,
"NotCovered": 0,
"NotCoveredFast": 0,
"RestartNeeded": 0,
"RebootNeededSince": "0001-01-01T00:00:00Z",
"WrapperPassSeconds": 0
},
"Host": {
"ReleaseID": "",
"LastOutcome": "",
"LastAt": "0001-01-01T00:00:00Z",
"LastSuccessfulLeg": "0001-01-01T00:00:00Z",
"Pending": 0,
"NotCovered": 0,
"NotCoveredFast": 0,
"RestartNeeded": 0,
"RebootNeededSince": "0001-01-01T00:00:00Z",
"WrapperPassSeconds": 0
}
}
],
"latest_guest_release": {
"approved_at": "2026-10-04T12:39:33Z",
"approved_by": "auto",
"id": "os-guest-20261004-123933"
},
"latest_host_release": {
"approved_at": "2026-10-04T12:40:34Z",
"approved_by": "auto",
"id": "os-host-20261004-124034"
}
}
@@ -0,0 +1,172 @@
=== felhom-agent 0.141.1 selftest=os-update vmid=9201 ring=0 enabled=true guest-release=false host-release=false appliance=true ===
time=2026-10-04T13:52:57.534+02:00 level=INFO msg="osupdate: START" run=20261004T115257Z layer=guest vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T115257Z
time=2026-10-04T13:53:09.479+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T115257Z layer=guest:9201 lane=fast mode=apply select=pending-fast packages=0"
time=2026-10-04T13:53:09.479+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
time=2026-10-04T13:53:09.479+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0"
time=2026-10-04T13:53:09.479+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)"
time=2026-10-04T13:53:09.480+02:00 level=INFO msg="osupdate: DONE" run=20261004T115257Z layer=guest vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=6 not_covered=6 restart_needed=0 reboot_needed=false wrapper_seconds=11.9
time=2026-10-04T13:53:09.491+02:00 level=INFO msg="osupdate: START" run=20261004T115257Z layer=host vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T115257Z
time=2026-10-04T13:53:19.927+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T115257Z layer=host lane=fast mode=apply select=pending-fast packages=0"
time=2026-10-04T13:53:19.927+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
time=2026-10-04T13:53:19.927+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0"
time=2026-10-04T13:53:19.927+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)"
time=2026-10-04T13:53:20.705+02:00 level=INFO msg="osupdate: DONE" run=20261004T115257Z layer=host vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=80 not_covered=80 restart_needed=34 reboot_needed=true wrapper_seconds=10.4
--- os-update report (guest) ---
{
"health_reason": "",
"healthy": true,
"mode": "apply",
"not_covered": [
"docker-ce-cli",
"containerd.io",
"docker-ce",
"docker-buildx-plugin",
"docker-ce-rootless-extras",
"docker-compose-plugin"
],
"outcome": "nothing",
"pending": 6,
"reboot_needed": false,
"refused": null,
"release_id": "ring0-20261004T115257Z",
"restart_needed": null,
"ring": 0,
"run_id": "20261004T115257Z",
"upgraded": [],
"wrapper_seconds": 11.9
}
--- os-update report (host) ---
{
"health_reason": "",
"healthy": true,
"mode": "apply",
"not_covered": [
"frr",
"shim-signed-common",
"shim-unsigned",
"shim-helpers-amd64-signed",
"shim-signed",
"libradosstriper1",
"librgw2",
"ceph-common",
"librbd1",
"librados2",
"python3-cephfs",
"libcephfs2",
"python3-rgw",
"python3-rados",
"python3-ceph-argparse",
"python3-ceph-common",
"python3-rbd",
"ceph-fuse",
"chrony",
"libcorosync-common4",
"libcfg7",
"libcmap4",
"libcpg4",
"libknet1t64",
"libnozzle1t64",
"libquorum5",
"libvotequorum8",
"corosync",
"frr-pythontools",
"libjs-extjs",
"libnvpair3linux",
"libproxmox-acme-plugins",
"libproxmox-backup-qemu0",
"pve-qemu-kvm",
"libpve-notify-perl",
"libpve-cluster-api-perl",
"libpve-cluster-perl",
"pve-cluster",
"libpve-access-control",
"libpve-apiclient-perl",
"librados2-perl",
"proxmox-backup-client",
"proxmox-backup-file-restore",
"pve-manager",
"libproxmox-acme-perl",
"libpve-common-perl",
"libpve-guest-common-perl",
"qemu-server",
"libpve-storage-perl",
"pve-edk2-firmware-legacy",
"pve-edk2-firmware-ovmf",
"libpve-network-api-perl",
"libpve-network-perl",
"proxmox-firewall-data",
"pve-firewall",
"pve-container",
"pve-ha-manager",
"novnc-pve",
"proxmox-enterprise-support-keyring",
"proxmox-mini-journalreader",
"proxmox-widget-toolkit",
"pve-docs",
"pve-i18n",
"pve-xtermjs",
"pve-yew-mobile-i18n",
"pve-yew-mobile-gui",
"libuutil3linux",
"libzfs7linux",
"libzpool7linux",
"proxmox-first-boot",
"pve-firmware",
"proxmox-kernel-7.0.14-20-pve-signed",
"proxmox-kernel-7.0",
"proxmox-kernel-helper",
"pve-edk2-firmware-aarch64",
"pve-edk2-firmware",
"zfs-initramfs",
"zfsutils-linux",
"zfs-zed",
"tailscale"
],
"outcome": "nothing",
"pending": 80,
"reboot_needed": true,
"refused": null,
"release_id": "ring0-20261004T115257Z",
"restart_needed": [
"agetty",
"blkmapd",
"chronyd",
"cron",
"dbus-daemon",
"dmeventd",
"ksmtuned",
"lxc-monitord",
"lxc-start",
"lxcfs",
"pmxcfs",
"proxmox-firewal",
"pve-firewall",
"pve-ha-crm",
"pve-ha-lrm",
"pve-lxc-syscall",
"pvedaemon",
"pvedaemon worke",
"pvefw-logger",
"pveproxy",
"pveproxy worker",
"pvescheduler",
"pvestatd",
"qmeventd",
"rpcbind",
"rrdcached",
"smartd",
"spiceproxy",
"spiceproxy work",
"sshd",
"systemd-logind",
"systemd-udevd",
"watchdog-mux",
"zed"
],
"ring": 0,
"run_id": "20261004T115257Z",
"upgraded": [],
"wrapper_seconds": 10.4
}
pass took 23.2s
WALL_SECONDS=23.268350584
@@ -0,0 +1,125 @@
+ uname -r
7.0.2-6-pve
+ mokutil --sb-state
SecureBoot enabled
+ od -An -tx1 /sys/firmware/efi/efivars/SecureBoot-8be4df61-93ca-11d2-aa0d-00e098032b8c
+ head -1
06 00 00 00 01
+ ls /boot/vmlinuz-7.0.14-20-pve /boot/vmlinuz-7.0.2-6-pve
/boot/vmlinuz-7.0.14-20-pve
/boot/vmlinuz-7.0.2-6-pve
+ dpkg -l 'proxmox-kernel-*'
+ grep '^ii'
+ awk '{print $2,$3}'
proxmox-kernel-7.0 7.0.14-20
proxmox-kernel-7.0.14-20-pve-signed 7.0.14-20
proxmox-kernel-7.0.2-6-pve-signed 7.0.2-6
proxmox-kernel-helper 9.1.0+fde2
+ proxmox-boot-tool status
+ tail -5
Re-executing '/usr/sbin/proxmox-boot-tool' in new private mount namespace..
E: /etc/kernel/proxmox-boot-uuids does not exist.
+ grep -E '^GRUB_DEFAULT|^GRUB_TIMEOUT|^GRUB_SAVEDEFAULT' /etc/default/grub
GRUB_DEFAULT=0
GRUB_TIMEOUT=5
+ grub-editenv list
+ '[' -d /sys/firmware/efi ']'
+ efibootmgr
+ head -6
BootCurrent: 0003
Timeout: 0 seconds
BootOrder: 0003,0029,0001,0002,0006,0007,0019,001A,001B,001C,001D,001E,0008,0009,0016,0017,001F,0020,0021,0022,0023,0024,0025,0026,0027,0028,000A,0004,0005,000B
Boot0001* USB Floppy/CD VenMedia(b6fef66f-1495-4584-a836-3492d1984a8d,0500000001)0000424f
Boot0002* USB Hard Drive VenMedia(b6fef66f-1495-4584-a836-3492d1984a8d,0200000001)0000424f
Boot0003* proxmox HD(2,GPT,175383fb-546d-430e-9b4c-73ec2379169d,0x800,0x200000)/File(\EFI\proxmox\shimx64.efi)
+ sysctl kernel.panic kernel.panic_on_oops
kernel.panic = 0
kernel.panic_on_oops = 0
+ lsmod
+ grep -i -E 'sp5100|wdt|watchdog'
+ ls -l /dev/watchdog /dev/watchdog0
crw------- 1 root root 10, 130 Oct 4 09:46 /dev/watchdog
crw------- 1 root root 243, 0 Oct 4 09:46 /dev/watchdog0
+ wdctl
+ head -12
Device: /dev/watchdog0
Identity: Software Watchdog [version 0]
Timeout: 10 seconds
Timeleft: 9 seconds
Pre-timeout: 0 seconds
Pre-timeout governor: noop
Available pre-timeout governors: noop
FLAG DESCRIPTION STATUS BOOT-STATUS
KEEPALIVEPING Keep alive ping reply 1 0
MAGICCLOSE Supports magic close char 0 0
PRETIMEOUT Pretimeout (in seconds) 0 0
SETTIMEOUT Set timeout (in seconds) 0 0
+ dmesg
+ grep -i -E 'sp5100|watchdog'
+ tail -5
[ 0.353227] NMI watchdog: Enabled. Permanently consumes one hw-PMU counter.
+ apt-cache policy proxmox-default-kernel
+ head -4
proxmox-default-kernel:
Installed: 2.1.0
Candidate: 2.1.0
Version table:
+ apt-cache search --names-only '^proxmox-kernel-[0-9.]+-[0-9]+-pve-signed$'
+ sort -V
+ tail -3
proxmox-kernel-7.0.14-18-pve-signed - Proxmox Kernel Image (signed)
proxmox-kernel-7.0.14-19-pve-signed - Proxmox Kernel Image (signed)
proxmox-kernel-7.0.14-20-pve-signed - Proxmox Kernel Image (signed)
+ cat /proc/cmdline
BOOT_IMAGE=/boot/vmlinuz-7.0.2-6-pve root=/dev/mapper/pve-root ro quiet
+ grep -nE '^menuentry|^submenu|^\s+menuentry' /boot/grub/grub.cfg
+ cut -c1-140
+ head -8
26: menuentry_id_option="--id"
28: menuentry_id_option=""
110:menuentry 'Proxmox VE GNU/Linux' --class proxmox --class gnu-linux --class gnu --class os $menuentry_id_option 'gnulinux-simple-529c0c3d
128:submenu 'Advanced options for Proxmox VE GNU/Linux' $menuentry_id_option 'gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43' {
129: menuentry 'Proxmox VE GNU/Linux, with Linux 7.0.14-20-pve' --class proxmox --class gnu-linux --class gnu --class os $menuentry_id_optio
147: menuentry 'Proxmox VE GNU/Linux, with Linux 7.0.14-20-pve (recovery mode)' --class proxmox --class gnu-linux --class gnu --class os $me
165: menuentry 'Proxmox VE GNU/Linux, with Linux 7.0.2-6-pve' --class proxmox --class gnu-linux --class gnu --class os $menuentry_id_option
183: menuentry 'Proxmox VE GNU/Linux, with Linux 7.0.2-6-pve (recovery mode)' --class proxmox --class gnu-linux --class gnu --class os $menu
+ ls -l --time-style=+%F_%T /boot/grub/grub.cfg
-rw------- 1 root root 13743 2026-10-04_09:47:05 /boot/grub/grub.cfg
+ uptime -s
2026-10-04 09:46:15
+ last -x reboot
+ head -4
reboot system boot 7.0.2-6-pve Sun Oct 4 09:46 - still running
reboot system boot 7.0.14-20-pve Sun Oct 4 09:44 - 09:45 (00:00)
shutdown system down 7.0.14-20-pve Sun Oct 4 09:45 - 09:46 (00:00)
reboot system boot 7.0.2-6-pve Fri Aug 21 17:44 - 09:43 (43+15:58)
+ ls /boot/efi/EFI/proxmox/
BOOTX64.CSV
fbx64.efi
grub.cfg
grubx64.efi
mmx64.efi
shimx64.efi
+ cat /boot/efi/EFI/proxmox/grub.cfg
+ head -5
search.fs_uuid 529c0c3d-b48e-4d01-989d-43fd5d7dbb43 root lvmid/zVqGDa-V6js-HB26-RyZr-d0ft-fUB3-blL75R/iAs9TN-W8rL-ViG8-jptz-1tcY-3djE-MRqamk
set prefix=($root)'/boot/grub'
configfile $prefix/grub.cfg
+ lspci -nn
+ grep -i -E 'smbus|fch'
00:14.0 SMBus [0c05]: Advanced Micro Devices, Inc. [AMD] FCH SMBus Controller [1022:790b] (rev 61)
00:14.3 ISA bridge [0601]: Advanced Micro Devices, Inc. [AMD] FCH LPC Bridge [1022:790e] (rev 51)
06:00.0 SATA controller [0106]: Advanced Micro Devices, Inc. [AMD] FCH SATA Controller [AHCI mode] [1022:7901] (rev 61)
+ modinfo -F filename sp5100_tco
/lib/modules/7.0.2-6-pve/kernel/drivers/watchdog/sp5100_tco.ko
+ grep -rl sp5100 /lib/modprobe.d /etc/modprobe.d
/lib/modprobe.d/blacklist_proxmox-kernel-7.0.14-20-pve.conf
/lib/modprobe.d/blacklist_proxmox-kernel-7.0.2-6-pve.conf
+ grep -h sp5100 /lib/modprobe.d/aliases.conf /lib/modprobe.d/blacklist_proxmox-kernel-7.0.14-20-pve.conf /lib/modprobe.d/blacklist_proxmox-kernel-7.0.2-6-pve.conf /lib/modprobe.d/fbdev-blacklist.conf /lib/modprobe.d/proxmox_prevent_autoload_proxmox-kernel-7.0.14-20-pve.conf /lib/modprobe.d/proxmox_prevent_autoload_proxmox-kernel-7.0.2-6-pve.conf /lib/modprobe.d/systemd.conf /etc/modprobe.d/amd64-microcode-blacklist.conf /etc/modprobe.d/pve-blacklist.conf /etc/modprobe.d/zfs.conf
+ head -3
blacklist sp5100_tco
blacklist sp5100_tco
+ modprobe -n -v sp5100_tco
insmod /lib/modules/7.0.2-6-pve/kernel/drivers/watchdog/sp5100_tco.ko
+ ls /sys/class/watchdog/
watchdog0
@@ -0,0 +1,20 @@
+ findmnt -no SOURCE,FSTYPE --target /boot
/dev/mapper/pve-root ext4
+ ls -l /boot/grub/grubenv
-rw-r--r-- 1 root root 1024 Oct 4 09:46 /boot/grub/grubenv
+ cp -p /etc/default/grub /root/grub.default.bak-2026-10-04
+ sed -i 's/^GRUB_DEFAULT=0$/GRUB_DEFAULT=saved/' /etc/default/grub
+ grep '^GRUB_DEFAULT' /etc/default/grub
GRUB_DEFAULT=saved
+ update-grub
+ tail -3
Found memtest86+ 32bit image: /boot/memtest86+ia32.bin
Adding boot menu entry for UEFI Firmware Settings ...
done
+ grub-set-default 'gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.2-6-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43'
+ grub-editenv list
saved_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.2-6-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43
+ grep -n 'set default' /boot/grub/grub.cfg
+ head -4
17: set default="${next_entry}"
22: set default="${saved_entry}"
@@ -0,0 +1,14 @@
+ grub-reboot 'gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.14-20-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43'
WARNING: Detected GRUB environment block on lvm device
gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.14-20-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43 will remain the default boot entry until manually cleared with:
grub-editenv /boot/grub/grubenv unset next_entry
+ grub-editenv list
saved_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.2-6-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43
next_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.14-20-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43
+ uname -r
7.0.2-6-pve
+ date -u +%FT%TZ
2026-10-04T12:26:10Z
reboot1 issued 2026-10-04T12:26:10Z
@@ -0,0 +1,19 @@
+ uname -r
7.0.14-20-pve
+ uptime -s
2026-10-04 14:27:03
+ mokutil --sb-state
SecureBoot enabled
+ grub-editenv list
saved_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.2-6-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43
next_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.14-20-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43
+ cat /proc/cmdline
BOOT_IMAGE=/boot/vmlinuz-7.0.14-20-pve root=/dev/mapper/pve-root ro quiet
active
active
active
active
active
status: running
starting
starting
@@ -0,0 +1,33 @@
reboot2 issued 2026-10-04T12:35:19Z, no grub command; grubenv as after reboot 1
ssh back after ~40s
System is going down. Unprivileged users are not permitted to log in anymore. For technical details, see pam_nologin(8).
+ uname -r
7.0.14-20-pve
+ uptime -s
2026-10-04 14:27:03
+ grub-editenv list
saved_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.2-6-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43
next_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.14-20-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43
+ cat /proc/cmdline
BOOT_IMAGE=/boot/vmlinuz-7.0.14-20-pve root=/dev/mapper/pve-root ro quiet
--- after the real restart, 2026-10-04T12:36:27Z
+ uname -r
7.0.14-20-pve
+ uptime -s
2026-10-04 14:36:11
+ mokutil --sb-state
SecureBoot enabled
+ grub-editenv list
saved_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.2-6-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43
next_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.14-20-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43
+ cat /proc/cmdline
BOOT_IMAGE=/boot/vmlinuz-7.0.14-20-pve root=/dev/mapper/pve-root ro quiet
+ systemctl is-active pveproxy pvedaemon pvestatd pve-cluster felhom-agent
activating
active
active
active
inactive
+ pct status 9201
status: stopped
@@ -0,0 +1,27 @@
+ grub-editenv /boot/grub/grubenv unset next_entry
+ grub-set-default 'gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.14-20-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43'
+ grub-editenv list
saved_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.14-20-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43
+ grep '^GRUB_DEFAULT' /etc/default/grub
GRUB_DEFAULT=saved
+ dpkg -l 'proxmox-kernel-*-pve-signed'
+ awk '{print $2}'
+ grep '^ii'
proxmox-kernel-7.0.14-20-pve-signed
proxmox-kernel-7.0.2-6-pve-signed
+ uname -r
7.0.14-20-pve
+ systemctl is-active pveproxy pvedaemon pvestatd pve-cluster felhom-agent
active
active
active
active
active
+ pct status 9201
status: running
+ sleep 45
+ pct exec 9201 -- docker inspect -f '{{.Name}} {{.State.Health.Status}}' cloudflared felhom-controller
/cloudflared healthy
/felhom-controller healthy
+ sysctl kernel.panic
kernel.panic = 0
@@ -0,0 +1,74 @@
+ modprobe sp5100_tco
rc=0
+ echo rc=0
+ lsmod
+ grep sp5100
sp5100_tco 20480 0
+ dmesg
+ grep -i sp5100
+ tail -5
[ 85.993589] sp5100_tco: SP5100/SB800 TCO WatchDog Timer Driver
[ 85.993780] sp5100-tco sp5100-tco: Using 0xfeb00000 for watchdog MMIO address
[ 85.993918] sp5100-tco sp5100-tco: initialized. heartbeat=60 sec (nowayout=0)
== /sys/class/watchdog/watchdog0
+ for w in /sys/class/watchdog/watchdog*
+ echo '== /sys/class/watchdog/watchdog0'
+ for f in identity state timeout nowayout bootstatus status
++ cat /sys/class/watchdog/watchdog0/identity
identity=Software Watchdog
+ printf '%s=%s\n' identity 'Software Watchdog'
+ for f in identity state timeout nowayout bootstatus status
++ cat /sys/class/watchdog/watchdog0/state
state=active
+ printf '%s=%s\n' state active
+ for f in identity state timeout nowayout bootstatus status
++ cat /sys/class/watchdog/watchdog0/timeout
timeout=10
+ printf '%s=%s\n' timeout 10
+ for f in identity state timeout nowayout bootstatus status
++ cat /sys/class/watchdog/watchdog0/nowayout
nowayout=0
+ printf '%s=%s\n' nowayout 0
+ for f in identity state timeout nowayout bootstatus status
++ cat /sys/class/watchdog/watchdog0/bootstatus
bootstatus=0
+ printf '%s=%s\n' bootstatus 0
+ for f in identity state timeout nowayout bootstatus status
++ cat /sys/class/watchdog/watchdog0/status
status=0x8000
== /sys/class/watchdog/watchdog1
+ printf '%s=%s\n' status 0x8000
+ for w in /sys/class/watchdog/watchdog*
+ echo '== /sys/class/watchdog/watchdog1'
+ for f in identity state timeout nowayout bootstatus status
++ cat /sys/class/watchdog/watchdog1/identity
identity=SP5100 TCO timer
+ printf '%s=%s\n' identity 'SP5100 TCO timer'
+ for f in identity state timeout nowayout bootstatus status
++ cat /sys/class/watchdog/watchdog1/state
state=inactive
+ printf '%s=%s\n' state inactive
+ for f in identity state timeout nowayout bootstatus status
++ cat /sys/class/watchdog/watchdog1/timeout
timeout=60
+ printf '%s=%s\n' timeout 60
+ for f in identity state timeout nowayout bootstatus status
++ cat /sys/class/watchdog/watchdog1/nowayout
nowayout=0
+ printf '%s=%s\n' nowayout 0
+ for f in identity state timeout nowayout bootstatus status
++ cat /sys/class/watchdog/watchdog1/bootstatus
bootstatus=0
+ printf '%s=%s\n' bootstatus 0
+ for f in identity state timeout nowayout bootstatus status
++ cat /sys/class/watchdog/watchdog1/status
status=0x0
+ printf '%s=%s\n' status 0x0
+ rmmod sp5100_tco
rmmod_rc=0
+ echo rmmod_rc=0
+ lsmod
+ grep -c sp5100
0
+ ls /sys/class/watchdog/
watchdog0
@@ -1,2 +1,4 @@
demo-felhom wrapper 4729769ce32e25e6 755 root; sudoers 02df92d751f1780a
demo-hp wrapper 4729769ce32e25e6 755 root; sudoers 02df92d751f1780a
demo-felhom wrapper 51e100ad81945f67
demo-hp wrapper 51e100ad81945f67
+11
View File
@@ -26,6 +26,17 @@
---
## 2026-10-04 (afternoon) — OS updates, host fast lane + fleet view + alarms (agent v0.141.0/v0.141.1, hub v0.131.0/v0.131.1, controller v0.292.0)
> Evidence: `audits/os-host-lane-2026-10-04/`.
| Row | What | Closed | Evidence |
|---|---|---|---|
| **R-841** | **Every box reported its tunnel `inactive`** (the agent asked a host unit that does not exist). Agent v0.141.0 reads the guest's `cloudflared` container and controller v0.292.0's Docker health check on cloudflared's own `/ready` (200 only with a connection); three states `running` / `not_running` / `unknown`; hub v0.131.0 alarms `tunnel_down` after two `not_running` reports and `tunnel_recovered` on the next `running`. LIVE on demo-hp: port 7844 blocked → `tunnel_down` mailed after the 2nd report; unblocked → `tunnel_recovered`. **Reasoning kept: `unknown` never alarms; a container state alone says "up" for a dead tunnel (wrong token: running, `/ready` 503). A plain `docker stop` is healed by the controller's protected-container check within 5 min — before the 15-min host report sees it.** | CLOSED 2026-10-04 — FIXED | `partA/` |
| **R-845** | **The OS leg was slow.** One `pct exec` per package (~0.9 s each) replaced by one call per layer; restart scan only after an install (host: every pass, v0.141.1); repair only when `dpkg --audit` reports. MEASURED: nothing to install, both layers — 23.3 s (demo-felhom), 31.5 s (demo-hp); before, guest only, nothing to install — 14.0 s; the 174–245 s passes of R-845 were passes WITH an install. A 108-package host install pass: 70 s. | CLOSED 2026-10-04 — FIXED | `partD/`, `partB/live/` |
| **R-846** | **Host "reboot needed" was wrong in agent v0.141.0** (found live on demo-felhom): the restart scan skipped every cgroup containing `lxc`, hiding `lxc-start` (`0::/lxc.monitor/<vmid>`, 20 deleted maps after libc6); and it scanned only after an install, so a reboot never cleared it (the hub's 14-day alarm would fire on a rebooted host). Agent v0.141.1 (`:/lxc/`, host scans every pass, `reboot_scanned`) + hub v0.131.1. Red-proved. | CLOSED 2026-10-04 — FIXED | `partB/live-defects-redproofs.txt` |
| **R-850** | **Hub v0.131.0 put the layer into the release fingerprint**, so the unchanged guest set (272 packages) counted as new and waited its 24 h again. One-time; ring 1 kept the previous release meanwhile; approved again under the TEST wait (`os-guest-20261004-123933`). Nothing to fix. | CLOSED 2026-10-04 — ONE-TIME, NO ACTION | hub log 2026-10-04 |
## 2026-10-04 (~12:20) — operator ruling
| Row | What | Closed | Evidence |
+6 -5
View File
@@ -308,7 +308,7 @@ stopping line that lies.
|---|---|---|---|---|---|---|---|
| **R-530** | Box system & updates | P2 | **[P2-MEDIUM] A floor does not deliver an agent: agents update only by an operator-signed `agent_update` job per box, and nothing records which boxes still run 0.130.0.** MEASURED 2026-09-15: the hub HOLDS a floor whose declared MinAgent is above the box's agent (`api/handler.go` ResolveManagedFloor); the agent's only update path is `signedjobs` + `selfupdate.Executor`. demo-hp reached 0.131.0 by `felhom-opsign -op agent_update` (key `felhom-op-1`) at 08:44:16Z and its controller floor was then SERVED in 3 s. **demo-felhom (N100) and Peti's box still run 0.130.0** — not touched (Peti fenced; N100 not asked). **What it needs:** the operator signs per box, or rules a fleet rollout step. **NARROWED 2026-09-16 (operator ruling 1):** the keys stay on DooPlex owner-only and CC may sign `agent_update` until the first PAYING customer (testers excluded) — recorded in `CONTEXT.md` + `04-control-plane-authorization.md` §3.1. Both demo boxes now run agent 0.131.0 (demo-hp 2026-09-15, demo-felhom 2026-09-16, each by a per-box signed job; Peti's box untouched, still 0.130.0). **What remains:** a fleet rollout step — signing per box does not scale past a handful, and nothing lists which boxes are behind. **2026-09-25:** Peti's box was RETIRED (operator ruling) — it no longer counts as a box behind. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (signing)** | — | — | operator |
| **R-604** | Box system & updates | P2 | **[P2-MEDIUM] A per-customer controller floor silently excludes that box from every global floor raise, and NOTHING says so — demo-hp missed four of them.** FOUND 2026-09-21 while raising the global floor to 0.259.0 at the operator's request. The raise logged `Global controller-version floor set to "0.259.0"` and then `managed floor SERVED for demo-felhom` — **and nothing at all for demo-hp**, which went on reporting every few minutes and stayed on 0.258.0. Cause: `customer_configs.min_controller_version` for demo-hp held **`0.243.0`**, a per-customer override that wins over the global. It is a **leftover from the 2026-09-16 drill**, whose golden was 0.243.0; **R-343 measured on 2026-08-18 that all five rows were EMPTY and recorded that as a safety property** — it stopped being true and nothing surfaced the change. demo-hp had therefore silently missed the raises to 0.253.0, 0.254.0, 0.257.0 and 0.259.0. **Why it is invisible rather than merely quiet:** `managed floor SERVED` fires **once per CHANGE** (`h.floorNotes`, `api/handler.go:600`), deliberately, because a box reports every few minutes — so a box whose override never changes is silent for ever, and its silence is indistinguishable from the silence of a box that already had the line. A session that raises the floor reads one SERVED line and reasonably concludes the fleet took it. **CLEARED** for demo-hp the same session (rollback line: POST `/customers/demo-hp/floor` with `min_controller_version=0.243.0`, `min_agent=0.131.0`); it then self-updated 0.258.0 → 0.259.0 in **under four minutes**, healthy, `settle-gate: GO — at/above floor 0.259.0`, and its claim page answers **"Wrong or expired code"** in English — the floor delivered the FIX, not a version string, to a box nobody hand-deployed. All five overrides are now empty. **Fix shape (smallest first):** the floor-raise page shows which customers carry an override and would NOT be moved, before the save; or the raise logs one line per customer naming the ones it skipped and why. A raise that quietly reaches half the fleet is worse than one that refuses. | **READY — rank P2-MEDIUM; owner: CC (hub)** | — | — | CC |
| **R-812** | Box system & updates | P2 | **[P2] A box never receives operating-system security updates — not the Proxmox host, not the guest's Debian, not its Docker engine.** SEARCHED 2026-10-03 (read-only): `felhom-controller`, `felhom-agent`, `app-catalog-felhom.eu` and `felhom.eu` hold no `apt-get upgrade`, `apt full-upgrade`, `unattended-upgrades`, `pveupgrade` or `needrestart` that runs on a box. The installer aligns the host's Proxmox repositories to no-subscription *"so the box can pull security updates"* and then says plainly *"No upgrades are run"* (`scripts/felhom-host-install.sh:2133-2136`). The guest's Docker engine is installed when the golden is BAKED (`felhom-agent/configs/build-golden.sh:124-125`), so a fresh install gets that week's engine and an installed box keeps it forever. The only `apt full-upgrade` in the project is a by-hand step for the off-site endpoint ep0 (`documentation/runbooks/offsite-endpoint.md:41`), not a box. App images ARE updated (the update arc); the layer under them is not. The intention, with its scope, is **R-808** in `ROADMAP.md`. | **NARROWED AGAIN 2026-10-04 (day) — the GUEST's Debian fast lane is BUILT and proven live (agent v0.140.0, hub v0.130.0, installer 1.29.0; `11` §8.1, `audits/os-guest-lane-2026-10-04/`). LEFT: the host, Docker and kernel lanes; the undo (R-842); existing boxes (R-840).** Earlier: NARROWED 2026-10-04 — the SPIKE is done (`11` §7.1, corrections C1–C12, `audits/os-updates-spike-2026-10-04/`); no product code yet. LEFT: the build steps of `11` §8, each with the operator's go; the §5.3 snapshot question is in STATUS; preconditions R-835, R-836, R-837.** | — | — | CC + operator |
| **R-812** | Box system & updates | P2 | **[P2] A box never receives operating-system security updates — not the Proxmox host, not the guest's Debian, not its Docker engine.** SEARCHED 2026-10-03 (read-only): `felhom-controller`, `felhom-agent`, `app-catalog-felhom.eu` and `felhom.eu` hold no `apt-get upgrade`, `apt full-upgrade`, `unattended-upgrades`, `pveupgrade` or `needrestart` that runs on a box. The installer aligns the host's Proxmox repositories to no-subscription *"so the box can pull security updates"* and then says plainly *"No upgrades are run"* (`scripts/felhom-host-install.sh:2133-2136`). The guest's Docker engine is installed when the golden is BAKED (`felhom-agent/configs/build-golden.sh:124-125`), so a fresh install gets that week's engine and an installed box keeps it forever. The only `apt full-upgrade` in the project is a by-hand step for the off-site endpoint ep0 (`documentation/runbooks/offsite-endpoint.md:41`), not a box. App images ARE updated (the update arc); the layer under them is not. The intention, with its scope, is **R-808** in `ROADMAP.md`. | **NARROWED 2026-10-04 (afternoon) — the HOST's Debian fast lane is BUILT and proven live too (agent v0.141.1, hub v0.131.1; `11` §8.2, `audits/os-host-lane-2026-10-04/`): appliances only, never kernel/boot/firmware, after a healthy guest step; fleet view and four alarms (§8.3). LEFT: the Docker and kernel slow lanes (R-836); existing boxes (R-840).** Earlier: NARROWED AGAIN 2026-10-04 (day) — the GUEST's Debian fast lane is BUILT and proven live (agent v0.140.0, hub v0.130.0, installer 1.29.0; `11` §8.1, `audits/os-guest-lane-2026-10-04/`). LEFT: the host, Docker and kernel lanes; the undo (R-842); existing boxes (R-840).** Earlier: NARROWED 2026-10-04 — the SPIKE is done (`11` §7.1, corrections C1–C12, `audits/os-updates-spike-2026-10-04/`); no product code yet. LEFT: the build steps of `11` §8, each with the operator's go; the §5.3 snapshot question is in STATUS; preconditions R-835, R-836, R-837.** | — | — | CC + operator |
| **R-35** | Box system & updates | P3 | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` | CC |
| **R-50b** | Box system & updates | P3 | **[P2] A root-owned privileged host artifact is delivered unversioned from `main` — "which wrapper is on this host?" is unanswerable.** `configs/felhom-pbs-apply` installs to `/usr/local/sbin/felhom-pbs-apply` (0755 root:root) and is the pinned sudoers vector for `create\ |reconcile\|grant` against `/etc/pve/priv/storage`. It is fetched by `felhom-host-install.sh:1914` via `fetch_raw`, which hits `raw/branch/main/<path>` — **no tag, no pin, no checksum, and no record in the Day-0 artifact manifest**, unlike the agent binary (sha256-vouched) and the golden image. Three consequences: (1) two hosts installed a week apart can carry different privileged wrapper code while both reporting the same agent version; (2) a host hotfixed in place (felhom-pve, 2026-07-18) is indistinguishable from one that fetched the same content — the fleet has no inventory of it; (3) an accidental push to `main` reaches the next install of every host with no review gate between commit and root-owned deployment. | **NARROWED** — **(a) SHIPPED 2026-07-21; (b)/(c) open** — moved from `ROADMAP.md` 2026-10-03: it states a checkable fact about the shipped product, so it is a FINDING (the sorting rule). **Re-ranked 2026-10-03: [P2] → P3 — operator-only; leg (a) shipped.** | — | **Surfaced 2026-07-21 while stopping the R-39 v0.90.1 publish** (`felhom-controller/REPORT.md` §5): the publish was cancelled precisely because the version number would have claimed to carry a fix that in fact rides this unversioned channel. Candidate shapes, in increasing cost: (a) record the wrapper's sha256 in the Day-0 artifact manifest beside the agent binary and have the agent report the installed file's hash, so drift is at least *visible*; (b) `fetch_raw` takes a pinned ref (tag or commit) supplied by the manifest rather than `main`; (c) the wrapper becomes a published generic-registry artifact with the same gate ladder as the agent binary. **(a) is the cheap honest first step and would have caught this class already.** Pairs with R-39 (whose remaining fleet half is specced separately) **(a) SHIPPED 2026-07-21 — hub v0.68.0 + agent v0.91.2.** `ArtifactManifest.WrapperSHA256` + an operator field; agents report the installed wrapper's sha256 each cycle and the host page surfaces a mismatch. **An unknown on EITHER side reads as quiet, never as drift** — lighting every host amber on rollout day is how a warning becomes background noise. Live confirmation of exactly the problem: felhom-pve's July-18 in-place hotfix hashed `2888f2ea…`, matching **no commit anyone could name**; it now reports `104db0a4…` against a vouchable manifest value. **(b)/(c) REMAIN OPEN:** the wrapper is still fetched unversioned from `raw/branch/main` — this makes drift *visible*, it does not fix the channel. Also recorded: the 0440 sudoers file is not agent-readable, so its drift stays invisible. **Flips (2026-10-03):** `00` §A "The installer is PUBLISHED, not pushed" — the same discipline for the privileged wrappers. **Re-ranked 2026-10-03:** [P2] → P3: operator-only; (a) shipped 2026-07-21 (the report carries the wrapper sha256), (b)/(c) open. **Finding-shaped** — an R-424 instance; check against today's product before building. | CC |
| **R-78** | Box system & updates | P3 | **`local_api` authority ruling — auto-reconcile vs detect-only** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-25, size M, roadmap state `idea (deferred OUT of R-77 on purpose)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **An OWED OPERATOR DECISION, not a defect — moved because an owed ruling hidden among feature ideas is the shape this session exists to remove.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | R-77 ships detection because the fix is genuinely undecided, and **both directions can lose customer-visible function**. **Direction 1 (today):** `controller.yaml` wins and drift is silent → the 2026-07-25 island migration blinded the whole fleet's control plane for 17.5 h (drive gate, guest-reboot recovery, quiesce/backup all degrade). R-77 makes that loud but does not stop it recurring. **Direction 2 (`bootstrap.json` wins, auto-reconcile on boot):** a guest whose `controller.yaml` is CORRECT and whose `bootstrap.json` is stale — a half-completed re-provision, a hand-repaired guest, a setup-wizard box — gets a **working channel clobbered on the next restart**, fleet-wide and silently, during a routine deploy. That is not obviously better than the bug. Needs a spike: which writer is authoritative per field (endpoint vs fingerprint vs token — `mergeLocalAPI` replaces the whole block, so they cannot be reconciled independently today), whether the agent side should stamp a generation/mtime so 'newer wins' is even expressible, and whether reconcile should require an operator ack. Until then the drift alert plus a manual edit is the supported path. | CC |
@@ -324,10 +324,10 @@ stopping line that lies.
| **R-194** | Box system & updates | P4 | **PVE's permission cache delays every grant-state verdict by an unknown amount, so "the agent can read it" and "the ACL exists" are not the same measurement.** Observed twice while validating R-190's self-repair on demo-felhom 2026-08-04: both ACL rows for `/storage/felhom-backup` were deleted, and `GET /access/permissions` continued to report `Datastore.AllocateSpace` present — for **~40 s** in one run and **~16 minutes** in another. During that window the capability probe reads healthy and the self-repair does not fire | **OPEN** | — | **Why it matters beyond the delay:** it puts a floor under how fast a lost grant can be noticed, it makes any single permission read a lagging indicator, and — the interesting part — **it is a candidate contributor to R-190's own timeline**: a grant removed at an unknown moment could keep working until a cache expiry, which is exactly the shape of *worked at 04:44, refused at 09:24*. That does not explain what removed it, but it may explain when the refusal SURFACED, and the two have been treated as the same instant. **Not a defect in our code** — it is PVE behaviour, and the mitigation already tolerates it (the repair fires on the next probe after the cache clears). **What is worth deciding:** whether the store-grant probe should ALSO consult the storage content listing as a second signal, since that appeared to reflect the loss immediately (`{"data":[]}` while the permission read still said present) — two signals disagreeing is itself information, and today only one of them is read | CC |
| **R-373** | Box system & updates | P4 | **`SysDataGrowGB` is the intended lever for the system-data volume, it works, and nothing sets it.** Written down 2026-08-02 in `audits/SPIKE-recovery-unit-space-2026-08-02.md:230-232`, under an explicit *"### Not filed"* heading: the 20 G / 50 G mismatch was ruled a tier-sizing decision rather than a defect, *"`SysDataGrowGB` is the intended lever and it works; nothing sets it."* A lever with no caller is the same shape as R-368's comment — a setting that names a behaviour nothing invokes. **Age when filed: 20 days.** | **OPEN — LOW** | R-368 (same shape) | Either wire it to something an operator can reach, or remove it and record the sizing decision where a reader will meet it. | CC |
| **R-835** | Box system & updates | P3 | **Turning Docker's `live-restore` OFF with a restart stops every running container and starts none.** MEASURED 2026-10-04 on scratch 9202: `live-restore` on (via `systemctl reload docker`, which does enable it) kept all 6 containers running across two engine steps; `systemctl reload` with the baked `daemon.json` did NOT turn it off; a `systemctl restart docker` did — and the new daemon stopped every container (`Exited (0)`, `Removing stale sandbox … isRestore=false`) and restarted none, though all are `unless-stopped`. Nothing brought them back for 3.5 min. A precondition for the Docker slow lane (`11` C5): if `live-restore` ships, turning it off must be a guarded act (stop apps first), never a plain restart. `audits/os-updates-spike-2026-10-04/partG/` | **READY — design input, owner: CC** | — | — | CC |
| **R-836** | Box system & updates | P3 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin <new> --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **READY — measure before the kernel slow lane; owner: CC + operator (reboots)** | — | — | CC |
| **R-836** | Box system & updates | P3 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin <new> --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot <new>`: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** | — | — | CC |
| **R-851** | Box system & updates | P3 | **A host that panics stays stopped: `kernel.panic = 0`.** READ 2026-10-04 on demo-hp (os-host-lane Part E): `kernel.panic = 0`, `kernel.panic_on_oops = 0`. The brief assumed the host restarts after a panic; it does not — the box stays down until a person power-cycles it, and only `softdog` (dead in a panic) runs. A fix (`kernel.panic = 10` set by the installer, maybe `panic_on_oops`) changes how every box behaves and what a household sees, so it is the operator's call, together with the kernel lane. `audits/os-host-lane-2026-10-04/partE/e0-readonly-demo-hp.txt` | **READY — operator decision (with the kernel lane); owner: operator** | — | — | operator |
| **R-839** | Box system & updates | P3 | **After a Docker restart that stopped every container, the boot sweep HELD an app whose `app.yaml` `HDD_PATH` names its user folder instead of the drive.** MEASURED 2026-10-04 on scratch 9202 (paperless-ngx): `bootrecon` logged `drive /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx is not a live mountpoint — NOT starting it`, although the drive `/mnt/felhom-drives/scratch_hdd` IS a mountpoint; the app stayed down until started by hand. The customer boxes' 9201s showed no hold. **Not diagnosed:** which writer put a per-app path in `HDD_PATH` on 9202, and whether a household box can get it. `audits/os-updates-spike-2026-10-04/partG/SUMMARY.md` | **READY — diagnose; owner: CC** | — | — | CC |
| **R-840** | Box system & updates | P2 | **A new root wrapper or a new sudoers line cannot reach a box that is already installed — there is no product route.** FOUND 2026-10-04 (guest OS lane, Part B 1): the agent's signed `agent_update` replaces ONLY the binary (`configs/felhom-selfupdate-guarded` swaps `/usr/local/bin/felhom-agent`); wrappers (`/usr/local/sbin/felhom-*`) and `/etc/sudoers.d/felhom-agent` are written only by `felhom-host-install.sh` step 5 on a new box. So agent v0.140.0's OS leg is DEAD on every existing box (the capability probe will show `osapply-run` missing) until someone installs `felhom-os-apply` + the FELHOM_OSAPPLY line by hand — which is what the demo boxes got this session. The `wireguard-tools` line (S3) reached the fleet the same way: by reinstall or by hand. **Proposal:** a signed `agent_config_update` op (operator-signed like `agent_update`) whose params pin the agent TAG and the sha256 of a config bundle (sudoers + wrappers); the self-update wrapper installs it as root after `visudo -cf` and a syntax check, keeps the previous copies, and the capability probe confirms. `audits/os-guest-lane-2026-10-04/README.md` | **READY — design + operator go; owner: CC** **RULED 2026-10-04 ~12:20 (`09` §3 decision 82): not built now — "There are no older boxes"; no installed box exists outside the two demo boxes, which take wrapper changes by hand. Kept open, not re-ranked. Reviewer's note (as written): every box installed from now on becomes an "older box" the first time the wrapper or the sudoers line changes again (the 2026-10-04 host-lane brief changes the wrapper); R-840 must be solved before the first box that cannot be reached by hand.** | — | — | CC |
| **R-841** | Box system & updates | P3 | **The agent's `cloudflared` health probe reads a host systemd unit that does not exist — every box reports its tunnel `inactive`.** FOUND 2026-10-04: `felhom-agent/internal/hub/cloudflared.go` runs `systemctl is-active cloudflared` on the HOST; cloudflared is a container in the GUEST (`11` C8), so demo-hp answers `inactive` / `Unit cloudflared.service could not be found`, and the hub stores that in `cloudflared_status` for every box. The field is equally consistent with "tunnel down" and "never checked" (R-96 rule 3). Fix direction: read the guest's container state (the agent already may `pct exec * -- docker inspect -f *`), or drop the field; and say so in `03` (corrected 2026-10-04). | **READY — owner: CC** | — | — | CC |
## Monitoring & notifications — 24 rows (P2 2, P3 16, P4 6)
@@ -358,7 +358,7 @@ stopping line that lies.
| **R-348** | Monitoring & notifications | P4 | **Every agent restart blanks the reported backup list for up to ~18 hours, and the comment that covers it says "unaffected".** Observed 2026-08-20 while deploying R-344: the first host reports after `demo-hp`'s agent restart carry **`0 backups`** (11:15:50 and 11:30:52 CEST, two consecutive), while the box's own `pvesm list` shows archives present on **both** tiers. `internal/backup/store.go`'s `Store` is in-memory and `byTarget` is repopulated only when a backup **runs** — daily for the local tier, weekly for offsite — so the field reads 0 until the next run. `restore_tests` did **not** blank, because that half has a durable on-disk companion (`RestoreTestState`, R-189). **It blinds no alarm, and that was CHECKED rather than assumed.** `hub/internal/monitor/deadline.go` scans back over stored reports with a 7-day `backupEvidenceLookback` whose own comment names this exact case — *"when the LATEST report carries none... and against an agent that stayed restarted for days"* — and `pbs_snapshots` stayed populated at 2 regardless. So this is an observability wart, **not** a safety hole, and it is filed at that severity deliberately. **What is actually wrong is the comment.** The `Store` doc says *"Backups are unaffected — their freshness has a ground truth on the storage (R-84)"*. That is true of the **consequence** and false of the **field**, and it sits three lines below a paragraph explaining that the very same sentence about restore-tests *"used to be here and it is now FALSE"* — so the file already carries one correction of this shape and invites the next reader to trust the surviving half. | **READY (XS) — NEW 2026-08-20** | — | Say what is measured: the field IS lost on restart and repopulates only when a backup runs; the freshness VERDICT is unaffected because the hub looks back 7 days. **Name `backupEvidenceLookback` in the comment** so the cross-repo dependency is visible from the agent side — today the agent's claim of safety rests on a hub constant it does not mention. Per the workspace rule, a comment asserting an invariant needs a test pinning it: the pin belongs on the HUB side, asserting the verdict survives a report carrying `backups: []`. | CC |
| **R-371** | Monitoring & notifications | P4 | **The off-site tier is the only backup tier that announces nothing on success.** Written down 2026-08-05 in `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md:508-513` and explicitly *"recorded, not filed"*: the off-site run emits **no hub event at all**, while both lesser tiers do (`db_dump_completed`, `crossdrive_completed`). Failures are covered by `backup_run_failures` and staleness by the hub's 8-day tier deadline, which is why it was judged a wrinkle. **Still true 2026-08-22** — the 2026-08-21 drill's own event dump shows `db_dump_completed` and six `crossdrive_completed` rows and no off-site success event. **Age when filed: 17 days.** | **OPEN — LOW** | — | Either emit one, or record deliberately that the highest-value tier is silent on success and say why. | CC |
## Hub & operator — 23 rows (P2 1, P3 7, P4 15)
## Hub & operator — 24 rows (P2 1, P3 7, P4 16)
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|---|---|---|---|---|---|---|---|
@@ -384,7 +384,8 @@ stopping line that lies.
| **R-719** | Hub & operator | P4 | **[P2-MEDIUM] A customer who already exists never gets a fresh self-bind link when their new box registers: the last link expires in 7 days and nothing re-sends it.** MEASURED 2026-09-29 (new-household drill, `tester-1`): the previous link went out 2026-09-17 07:25 UTC at a host delete and expired 2026-09-24; the box registered at 19:11:30 UTC and its console told the volunteer to open the link from their e-mail — there was none that worked. Hub source: the link is sent at customer creation, RESET, e-mail set on a box-less customer and host delete (`selfbind_mint.go` callers `hosts.go:908`, `configs.go:850`, `customer_reset.go:162`) — never on appliance registration. The volunteer guide says the operator needs to press nothing. The operator pressed „Send self-bind link" (the mail arrived in 1 s) — an operator step the volunteer depends on, recorded, not an intervention. **Fix direction:** send the link when an unclaimed appliance registers while a box-less customer waits with no live link (R-509's first fix shape), or the guide's operator part says: press it the day the volunteer installs. Evidence: `audits/evidence-drill-new-household-2026-09-30/` `phase0/operator-steps.txt`. **CHANGED AND BUILT 2026-09-30 (hub v0.126.0) — the brief's shape was not buildable:** a box registers UNCLAIMED (uuid, MACs, host keys, hardware — nothing of a customer), so "send the link when their box registers" would mail every waiting customer. Built instead: the expired AND used link pages offer „Új linket kérek" → a fresh link to the address registered for that link's customer, only when it has no box, ≤1/h per customer, identical answer for any token (no oracle). Live: the button on the real hub, the same page for a made-up token, no mail; the mint+send path unit-proven (RP42). Evidence: `audits/evidence-fixes-first-tester-2026-09-30/``partD/`. | **WAITING-ON-OPERATOR** (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the operator has not reviewed the changed page shape ("operator may prefer another"), and mint+send is unit-proven only) — **CLOSED 2026-09-30 — hub v0.126.0 (changed shape; operator may prefer another)** | — | — | operator |
| **R-814** | Hub & operator | P4 | `PBS-storage-1` (u629193, box 611421) still `status=active`, 19.9 MB | **VERIFY** (2026-10-03 triage: a July watch row with no id; given R-814. WAITING-ON-OPERATOR — no record found that the box was deleted.) — WAITING-ON-OPERATOR | operator console | Delete the box | operator |
| **R-844** | Hub & operator | P4 | **The household's OS-update line exists only on the hub's customer timeline.** 2026-10-04: the box itself has no event surface for agent results (the controller UI shows no timeline), so `os_update_applied` is a hub customer event (info: recorded, never mailed). Its stored text is the hub's English sentence; the hu/en bundle text (`mail.event.os_update_applied`) is used only if it is ever mailed. Fix direction: a controller-side line (the controller already polls the agent's local API) when the box gets a household timeline. `audits/os-guest-lane-2026-10-04/partG/hub-customer-timeline-demo-hp.txt` | **READY — owner: CC** | — | — | CC |
| **R-845** | Hub & operator | P4 | **One OS-leg pass takes 3–4 minutes even when it installs 3 packages**: the wrapper's inventory (an `apt-get update`, `apt-cache policy` over every installed package, a `/proc/*/maps` scan for restart-needed, two simulations) dominates; measured 174–245 s per pass on the demo boxes vs 3.8–31.7 s for the install itself. It runs at night under the heavy-op gate, so it delays a restore-test by minutes, nothing worse. Fix direction: one `apt-cache policy` per run and the restart scan only after an install. `audits/os-guest-lane-2026-10-04/partG/` | **READY — owner: CC** | — | — | CC |
| **R-848** | Hub & operator | P4 | **A held host package is invisible to the hub.** MEASURED 2026-10-04 on demo-hp (undo runbook proof): with `tzdata` held after a by-hand undo, the wrapper's `pending` stayed 78 — apt's simulation leaves held packages out, so the fleet view shows nothing and no alarm can see a hold that was forgotten. The hold lives only in the incident's register row (`runbooks/os-updates-host-undo.md`). Fix direction: the wrapper reports `apt-mark showhold` and the fleet line shows it. | **READY — owner: CC** | — | — | CC |
| **R-849** | Hub & operator | P4 | **The fleet view's GUEST "reboot needed since" never clears.** 2026-10-04: the guest is scanned only after an install (R-845, one `pct exec`), so a guest restart is never seen; the guest line keeps the date of the last install that said "needed". No alarm reads the guest line (the reboot alarm is host-only), so it is display only. Fix direction: scan the guest on every pass too (one `pct exec`, ~1 s) or hide the guest date. `audits/os-host-lane-2026-10-04/partC/live/fleet-during-ring1-test.json` | **READY — owner: CC** | — | — | CC |
## Business & legal — 7 rows (P2 4, P4 3)
@@ -0,0 +1,81 @@
# Put one host package back by hand (OS updates, host fast lane)
> **When:** the hub mailed `os_update_health_failed` for the **host** layer (the box's base system), or the operator
> sees a host problem that started with a host OS pass, and ONE package is the suspect.
> **Who:** the operator, or CC with the operator's word, as root on the box's Proxmox host (`ssh <host>`).
> **Owner design:** `architecture/11-os-updates.md` §5.6 (host row), §8 step 3. **Proved** on demo-hp 2026-10-04 with
> `tzdata` (2026c → 2026b, held; then released and re-installed by the next ring-0 pass) and again on demo-felhom for
> the ring-1 test — evidence `audits/os-host-lane-2026-10-04/partB/undo/` and `partB/ring1/`.
There is **no automatic undo** for the host. A host cannot be snapshotted the way the old design assumed, and the
fast lane only ever installs Debian / Debian-Security packages, never a kernel, boot or firmware package (wrapper
refusals R14, R12). So the undo is small: install the previous version of the suspect package from
`snapshot.debian.org`, which keeps every version Debian ever published.
## 1. Find the suspect and its previous version
```bash
# the host pass's own apt run (Requested-By: felhom-agent); each line "name:arch (old, new)"
grep -B2 -A6 "Requested-By: felhom-agent" /var/log/apt/history.log | tail -20
PKG=<name>; OLD=<old version from that line>
```
## 2. Pick the snapshot timestamp
- **Normal case:** the timestamp of the **previous host release** — the hub fleet view (`GET /os/fleet`, the box's
host line names its release; the release's approval time IS its snapshot time, `YYYYMMDDTHHMMSSZ`). At that time
the old version was the current one.
- **No previous host release** (a ring-0 box, or the first host pass): ask snapshot.debian.org when the **NEW**
(installed) version first appeared, and use that time. At that moment the suite still carried the old version.
**Do not use the OLD version's `first_seen`:** that is when it reached Debian *unstable*, and `trixie` may not
have had it yet — measured on demo-hp 2026-10-04: `eject 2.41-5` at its own `first_seen` → `Version not found`.
```bash
NEW=$(dpkg-query -W -f='${Version}' "$PKG")
curl -s "https://snapshot.debian.org/mr/binary/$PKG/$NEW/binfiles?fileinfo=1" | python3 -c '
import json,sys; d=json.load(sys.stdin)
print(sorted(f["first_seen"] for v in d["fileinfo"].values() for f in v)[0])'
# → e.g. 20260831T204404Z
TS=<that timestamp>
```
## 3. Install the old version from the snapshot, then remove the snapshot source
```bash
. /etc/os-release
cat > /etc/apt/sources.list.d/felhom-undo-snapshot.list <<EOF
deb [check-valid-until=no] http://snapshot.debian.org/archive/debian/$TS $VERSION_CODENAME main
deb [check-valid-until=no] http://snapshot.debian.org/archive/debian-security/$TS $VERSION_CODENAME-security main
EOF
apt-get -q update
apt-get -s install --allow-downgrades "$PKG=$OLD" # READ the simulation: only $PKG may change. If apt
# refuses, the package pins its siblings to the SAME
# version (measured: eject needs libmount1 = 2.41-5) —
# list them all in one command, or stop.
apt-get -y install --allow-downgrades "$PKG=$OLD"
apt-mark hold "$PKG" # or the next ring-0 night installs the new one again
rm -f /etc/apt/sources.list.d/felhom-undo-snapshot.list
apt-get -q update
dpkg -s "$PKG" | grep -E "^(Status|Version)"
```
**The hold is the important line.** Without it the next night pass (ring 0) or the next approved release (ring 1)
installs the new version again. Write the hold into the register row that tracks the incident; take it off
(`apt-mark unhold "$PKG"`) when a fixed version is out.
## 4. Check the box
- `systemctl is-active pveproxy pvedaemon pvestatd pve-cluster felhom-agent` → all `active`.
- `pct status <customer vmid>` → `running`.
- The host health rule (`11` §8.2) is what the leg checks; run the debug action to see it pass:
`sudo -u felhom-agent /usr/local/bin/felhom-agent --config /etc/felhom-agent/agent.json --selftest=os-update -vmid <vmid>`
— **note: on ring 0 this also installs every pending fast-lane fix.** A held package is NOT reported as pending
at all (measured on demo-hp: `pending` stayed 78 with `tzdata` held) — **the hub cannot see a hold**, so the hold
lives only in the register row (R-848).
## What this runbook does NOT cover
- A kernel: the fast lane never installs one (R14). The kernel lane is `11` §5.6 (not built).
- A Proxmox-origin package (pve-*, qemu-server, …): never installed by the fast lane (R12/origin rule). Proxmox's
own repository keeps old versions; that undo is not written yet.
- The customer guest: its undo is last night's whole-guest backup (decision 81).