docs: OS updates steps 3+4 BUILT (11 §8.2/§8.3, §5.6 kernel facts, §5.8 Docker slow lane design), 00/03/07/08 updated, decisions 84-86 (CC unattended), host undo runbook (proved), register: R-841 R-845 R-846 R-850 closed, R-848 R-849 R-851 opened, R-836 R-812 narrowed (332 -> 333); STATUS; live evidence
gates / gates (push) Successful in 31s
gates / gates (push) Successful in 31s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
+10
@@ -16,6 +16,16 @@
|
||||
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
|
||||
|
||||
|
||||
> **2026-10-04 (late afternoon) — OS updates: host fast lane + fleet view + alarms BUILT (`11` §8.2–§8.3); the tunnel
|
||||
> status is true (R-841).** Agent v0.141.0 → v0.141.1, hub v0.131.0 → v0.131.1, controller v0.292.0 (cloudflared
|
||||
> readiness health check). **Decided by CC unattended — operator may reverse:** `09` §3 decisions **84** (appliance proof
|
||||
> = the root-owned install record), **85** (alarm numbers 7/14/7/14 days, configuration), **86** (a same-session patch
|
||||
> release of agent and hub for the host "reboot needed" defect, R-846). Live: `tunnel_down`/`tunnel_recovered` on
|
||||
> demo-hp; ring 0 host pass on demo-felhom (108 Debian packages); a 605-package host release; ring 1 exact install;
|
||||
> by-hand host undo proved; leg 23–32 s with nothing to install. Kernel spike (R-836, narrowed): GRUB's one-shot is not
|
||||
> a one-shot on LVM `/boot`; Secure Boot fine; `kernel.panic = 0` (R-851); `sp5100_tco` answers. demo-hp now runs and
|
||||
> defaults to kernel 7.0.14-20. Docker slow lane designed (`11` §5.8). `REPORT-os-host-lane-2026-10-04.md`.
|
||||
|
||||
> **Rulings 2026-10-04 (~12:20) — recorded before the work (host fast lane brief).** `09` §3 decisions **81** (R-842 A:
|
||||
> the undo is the whole-guest backup by hand; R-842 closed), **82** (R-840 not built now — "There are no older boxes";
|
||||
> row kept open with the reviewer's note) and **83** (next: R-841, the host fast lane, OS-update improvements).
|
||||
|
||||
@@ -2,8 +2,46 @@
|
||||
|
||||
**Ready for the first real tester (Tester-2): yes. You confirmed the tunnel route and the connect mails (2026-09-30).**
|
||||
|
||||
**Updated 2026-10-04 (afternoon): guest system updates are automatic on the demo boxes. Both demo boxes run
|
||||
controller 0.291.0 and host agent 0.140.0. Hub 0.130.0. New installs get golden 0.291.0.**
|
||||
**Updated 2026-10-04 (evening): the HOST's security fixes install themselves too, the hub has a fleet view and four
|
||||
OS alarms, and the tunnel status is true. Both demo boxes run controller 0.292.0 and host agent 0.141.1. Hub 0.131.1.
|
||||
New installs: see the golden line in the section below.**
|
||||
|
||||
## Today (2026-10-04, evening): host fixes, the fleet view, a true tunnel status
|
||||
|
||||
**Decisions I took myself (you may reverse each):**
|
||||
- Only a box the installer itself recorded as "appliance" gets host updates (a record the agent cannot change).
|
||||
- The four alarm times: no update run for 7 days; "reboot needed" for 14 days; the demo boxes approve nothing for 7
|
||||
days; a box has fixes nobody approved for 14 days. All four are settings.
|
||||
- I released the host agent and the hub a second time today. The first release said "reboot needed" wrongly; the new
|
||||
14-day alarm would then have mailed you about boxes you had already rebooted.
|
||||
|
||||
**One decision for you (a safe default if you say nothing) — the Docker engine updates (designed, not built):**
|
||||
1. **Turn on Docker's "live-restore" on every box.** With it, a Docker engine update restarts no app (measured: 0
|
||||
restarts). Without it, every app stops for about 30 seconds per engine update.
|
||||
- **A (my pick):** turn it on — in the new-install image and once on existing boxes. It goes on without restarting
|
||||
anything. Cost: it must never be turned off by a plain restart again (that stops every app and starts none).
|
||||
- **B:** leave it off. Every Docker update then means ~30 seconds of every app being down, at night.
|
||||
- **If you say nothing:** nothing changes; Docker updates stay unbuilt.
|
||||
|
||||
**What I did:**
|
||||
- **The host's Debian fixes now install themselves**, after the guest's, on the same night run, only on appliances,
|
||||
never a kernel or boot package, never a reboot. Proven: demo-felhom installed 108 host fixes and stayed healthy; all
|
||||
108 came from Debian. A box made "ring 1" for a test installed exactly the one approved host version it lacked.
|
||||
- **A way to put one host package back by hand** is written and proven on demo-hp (and the test taught it two fixes).
|
||||
- **The fleet view** in the hub: one line per box with its updates, "reboot needed since", and the tunnel.
|
||||
- **The tunnel status is now true:** running, not running, or unknown. I blocked demo-hp's tunnel: after two reports
|
||||
(about 30 minutes) you got a "tunnel down" mail; when I unblocked it, "recovered". A simply stopped tunnel heals
|
||||
itself within 5 minutes, before the hub can even see it.
|
||||
- **The night run is fast:** 23–32 seconds when there is nothing to install (target was under 60).
|
||||
- **The kernel test on demo-hp (your two reboots):** Secure Boot works with it, but GRUB's "boot once" does not work
|
||||
on our boxes: the second plain reboot came up on the NEW kernel again. So a new kernel that hangs would stay. That
|
||||
must be solved before kernel updates. demo-hp now runs the newer kernel, healthy. A watchdog chip exists on demo-hp.
|
||||
- **Found and fixed during the live test:** "reboot needed" was wrong in two ways (fixed in the second release).
|
||||
- **Rows:** 2 closed, 2 opened-and-closed the same day, 3 opened. The list went from 332 to 333.
|
||||
|
||||
**Needs you later (nothing breaks if you wait):**
|
||||
- **A host that crashes does not restart by itself** (Linux's "panic" setting is off). Changing it changes how every
|
||||
box behaves, so it is your call, together with kernel updates.
|
||||
|
||||
## Today (2026-10-04, afternoon): the guest's security fixes install themselves
|
||||
|
||||
|
||||
@@ -232,7 +232,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
||||
| **The hub reports LOSS OF VISIBILITY into either off-site store (not just how full it is)** | hub **v0.106.0** (R-339) | **IMPLEMENTED — deliberately NOT proven-live** | Both box checkers count consecutive failed fetch windows and emit `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default 3 windows (≈30–45 min), each with a paired `*_recovered` all-clear routed via `recoveredPairedDownTypes` — required because the recoveries are severity `info`, which `severityNotifies` drops. Scopes stay customer-less (`pbsdr-box` / `pool-box`) → operator channel only. Fill logic untouched: a degraded read still drives no band transition. Evidence: `internal/monitor/box_reachability_test.go` + the cross-package wiring test in `internal/notify/`, which asserts an actual operator mail rather than a map entry. **Filed BECAUSE of a measured gap**, not a hypothesis: the 2026-08-18 ep0 outage ran 9 h 37 m with the hub silent | **The gap that remains is R-340**, and it is not small: the ep0 read is the `usage` op, which rides the LOCAL API daemon — the daemon that incident explicitly cleared — so this check would have shown GREEN for that entire outage. It closes "ep0 is unreachable as a host"; it does not close what actually happened. **No live or constructed outage has exercised the emit path**, and one cannot be manufactured against ep0 (Tier 2, protected) |
|
||||
| Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | |
|
||||
| Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | |
|
||||
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.140.0, hub v0.130.0 | **PARTIAL — the GUEST's Debian fast lane is PROVEN-LIVE (2026-10-04); the host, Docker and the kernel are MISSING** | `audits/os-guest-lane-2026-10-04/` — ring 0 (both demo boxes) installed 53 Debian fixes each, healthy; the hub approved a 272-package release (TEST wait 2 min, then the ruled 24 h + 1 night restored); demo-felhom as ring 1 installed exactly the 3 approved versions; a stopped app → `health_failed` + operator mail. Design `architecture/11-os-updates.md` §8.1 | **No automatic undo** (a customer guest cannot be snapshotted, R-837 → R-842); existing boxes need the wrapper + sudoers by hand (R-840); host / Docker / kernel lanes not built (R-812, R-835, R-836); `felhom-host-install.sh:2133-2136` still runs no host upgrades |
|
||||
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.141.1, hub v0.131.1 | **PARTIAL — the GUEST and HOST Debian fast lanes are PROVEN-LIVE (2026-10-04), with the fleet view and four operator alarms; Docker and the kernel are MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/` — ring 0 on both demo hosts (demo-felhom 108 Debian host packages, healthy, 0 Proxmox-origin); a 605-package host release approved (TEST wait, then the ruled 24 h + 1 night restored); demo-felhom as ring 1 installed exactly the one host version it lacked; the by-hand host undo proved (`runbooks/os-updates-host-undo.md`); `tunnel_down` / `tunnel_recovered` live. Design `architecture/11-os-updates.md` §8.1–§8.3 | **No automatic undo** (guest: last night's backup, decision 81; host: the by-hand runbook); existing boxes need the wrapper + sudoers by hand (R-840); Docker and kernel lanes not built (R-812, R-835, R-836); a host panic is not restarted (R-851) |
|
||||
| **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** |
|
||||
| **A deletion of a customer's off-site history is NOTICED within a day** | hub **v0.111.0** (R-431) | **IMPLEMENTED — not yet PROVEN-LIVE** | 09-01 | `hub/internal/monitor/offsite.go` — third signal beside FILL and STALENESS. **On the hub deliberately:** a detector on the box is one the deletion can silence. Alarms when the reported count falls by more than HALF and by at least 5, guarded by `StatsKnown` (R-331), the declared `State` (R-204) and run success (R-100). **Threshold reasoned, not invented:** over 12 898 reports every decrease lands on ZERO and predates `stats_known`; in the 380-report `stats_known` window there are none. **ACCEPTANCE: 9 009 real points replayed → ZERO alarms** (`offsite_r431_test.go`, fixture committed). **What PROVEN-LIVE would need and this does NOT have:** a real drop observed on a live box producing a real mail — the live firing done at ship time was driven through the hub's own path with synthetic counts, which is an end-to-end delivery proof, not a proof that a genuine deletion is caught. |
|
||||
|
||||
|
||||
@@ -47,7 +47,7 @@ Owns:
|
||||
1. **Proxmox lifecycle** — create/start/stop/destroy guests, snapshots, storage allocation. Via a scoped Proxmox API token (the **`FelhomAgent` operator role** — `proxmox-platform.md` §3.6, validated Phase 3 B3) for everything the API covers; raw host ops only where unavoidable.
|
||||
2. **Storage management** — attach/classify targets, reconcile the storage manifest, mount USB-by-UUID, present mounts into guests.
|
||||
3. **Backup/restore orchestration** — vzdump to the tiers, PBS, snapshot management, and the **self-restore-test**.
|
||||
4. **Host & tunnel monitoring** — host metrics, guest up/down, storage-target status, and `cloudflared` health; reports the host domain to the hub. **[FACT, 2026-10-04] The "cloudflared health" leg reads a host unit that does not exist:** `internal/hub/cloudflared.go` runs `systemctl is-active cloudflared` on the HOST, but cloudflared is a container in the guest (below), so every box reports `inactive` (`Unit cloudflared.service could not be found`, demo-hp). R-841.
|
||||
4. **Host & tunnel monitoring** — host metrics, guest up/down, storage-target status, and `cloudflared` health; reports the host domain to the hub. **[FACT, 2026-10-04] The "cloudflared health" leg reads a host unit that does not exist:** `internal/hub/cloudflared.go` runs `systemctl is-active cloudflared` on the HOST, but cloudflared is a container in the guest (below), so every box reports `inactive` (`Unit cloudflared.service could not be found`, demo-hp). R-841. **[FACT, FIXED agent v0.141.0 / controller v0.292.0, R-841 CLOSED]** The agent now reads the guest's `cloudflared` container through the existing `pct exec [0-9]* -- docker inspect -f *` sudoers line: state, exit code and the Docker health status of a check the controller adds (`cloudflared tunnel --metrics localhost:20241 ready` → cloudflared's own `/ready`, 200 only with a connection). Three states: `running` (healthy), `not_running` (stopped, absent, or running but NOT connected), `unknown` (could not ask, or the check is still starting) — `unknown` never alarms. No new sudoers line.
|
||||
5. **Provisioning** — provision a guest **by restoring the golden base image** (§9), deploy the controller into it, hand it its bootstrap config; also **build and refresh the golden base image** itself.
|
||||
6. **Hub control loop** — poll for desired state + signed jobs, reconcile, execute, report, heartbeat.
|
||||
7. **Local API** — the per-guest authorization gate the controller calls.
|
||||
@@ -64,7 +64,7 @@ Explicitly does **not**:
|
||||
|
||||
- **Native Go binary, systemd service** on the host: boot-start, `Restart=always`, systemd watchdog (kill+restart on hang), journald logging, resource limits.
|
||||
- **Root-minimized (boundary settled — Phase 3 B3).** The agent runs as a **non-root** service user with the scoped `FelhomAgent` token for all API-covered work + a **narrow `sudoers` allowlist** for true host ops. Per Phase 3 (B3) the boundary is settled: the entire per-customer guest lifecycle — provision (by restore, §9), config, start/stop, snapshot, backup, **restore**, destroy — is token-covered. Genuine OS-root is confined to: (1) building/refreshing the **golden base image** (`keyctl` create is `root@pam`-only — one-time at enrollment + a maintenance cadence, §9); (2) **host mounts** (USB mount-by-UUID, systemd mount units / fstab); (3) **SMART / hardware sensors**. Root therefore never sits on the per-customer path. See `proxmox-platform.md` §3.6 for the role + boundary table.
|
||||
- ~~**`cloudflared` is a separate systemd service**, not embedded in the agent. … The agent **manages and health-watches** it (see §5) but the tunnel does not live or die with the agent process.~~ **[FACT, corrected 2026-10-04 — `11-os-updates.md` C8, R-838]** `cloudflared` is a **container in the customer guest** (`cloudflare/cloudflared:<pin>`), rendered and kept up by the in-guest controller (`felhom-controller/controller/internal/infra/infra.go`, `internal/stacks/infra.go`) and baked into the golden; there is no host systemd unit. It is still NOT embedded in the agent, so the data path survives the agent's death — and the agent neither manages it nor (see R-841) sees its health. Its version moves only by a controller release (the pin), on the monthly re-test (`runbooks/monthly-floating-retest.md` "Infrastructure pins").
|
||||
- ~~**`cloudflared` is a separate systemd service**, not embedded in the agent. … The agent **manages and health-watches** it (see §5) but the tunnel does not live or die with the agent process.~~ **[FACT, corrected 2026-10-04 — `11-os-updates.md` C8, R-838]** `cloudflared` is a **container in the customer guest** (`cloudflare/cloudflared:<pin>`), rendered and kept up by the in-guest controller (`felhom-controller/controller/internal/infra/infra.go`, `internal/stacks/infra.go`) and baked into the golden; there is no host systemd unit. It is still NOT embedded in the agent, so the data path survives the agent's death — and the agent does not manage it; it READS its health (R-841, agent v0.141.0). Its version moves only by a controller release (the pin), on the monthly re-test (`runbooks/monthly-floating-retest.md` "Infrastructure pins").
|
||||
|
||||
## 4. Control model — reconcile + signed destructive ops
|
||||
|
||||
@@ -143,7 +143,7 @@ notification) is the control.
|
||||
|
||||
**Box-initiated poll.** The hub never connects inbound. Each poll cycle exchanges:
|
||||
|
||||
- **Up:** heartbeat + a host-domain state report — host CPU/RAM/disk, per-guest up/down + spec, storage-target status (USB connected? NFS/CIFS reachable? PBS reachable?), last backup per target, last restore-test result, `cloudflared` health (a host-unit probe that always reads `inactive` — R-841), agent + controller versions, audit-log tail.
|
||||
- **Up:** heartbeat + a host-domain state report — host CPU/RAM/disk, per-guest up/down + spec, storage-target status (USB connected? NFS/CIFS reachable? PBS reachable?), last backup per target, last restore-test result, `cloudflared` health (`running` / `not_running` / `unknown` from the guest container's readiness check — R-841), agent + controller versions, audit-log tail.
|
||||
- **Down:** the current desired state, any pending signed one-shot jobs, and config (poll interval, update window, policy changes).
|
||||
|
||||
**Dead-man's-switch (essential, not optional).** In a box-initiated model the heartbeat
|
||||
|
||||
@@ -366,7 +366,10 @@ never moved by an update or an undo, and the unit restore accepts it.
|
||||
whole-guest backup (the controller drives it, inside [W+2h, W+6h)) ends SUCCESSFULLY on the primary tier, the agent
|
||||
waits 90 s and runs the guest's Debian fast lane — still holding the host-wide heavy-op gate, so it never overlaps a
|
||||
backup or a restore-test; at most once per 20 h. The backup minutes old is the guest's undo (no snapshot is possible,
|
||||
R-837). A failed or missed backup → no OS leg that night.
|
||||
R-837). A failed or missed backup → no OS leg that night. **[FACT, 2026-10-04 — agent v0.141.1, `11` §8.2]** On an
|
||||
appliance, the HOST step follows under the same gate: after a healthy guest step only (a failed or unhealthy guest
|
||||
step skips it), Debian-origin fixes only, never a kernel, boot or firmware package, never a reboot. Measured: both
|
||||
steps with nothing to install, 23–32 s; a 108-package host pass, 70 s.
|
||||
|
||||
**[FACT]** The three nightly legs derive from **one** customer-settable window start W at fixed
|
||||
offsets — db-dump at W, Tier-2 at W+60m, offsite at W+105m — so they can never be misordered
|
||||
|
||||
@@ -310,6 +310,32 @@ message, not a wider cooldown.
|
||||
|
||||
---
|
||||
|
||||
## 6.3 Box alarms outside the app ladder: the tunnel and OS updates [DESIGN, hub v0.131.0, 2026-10-04]
|
||||
|
||||
These are **operator-only** (the household can act on none of them; `operatorOnlyEvents`, pinned by
|
||||
`TestOSUpdateEvents_OperatorOnlyExceptApplied`). They go through the same dispatcher and severity contract (§6.1):
|
||||
`info` is recorded and never mailed; `warning` and `error` are mailed.
|
||||
|
||||
| Event | Severity | Raised when | Cleared | Pinned by |
|
||||
|---|---|---|---|---|
|
||||
| `tunnel_down` | error | the box's two newest host reports say the tunnel is `not_running` (more than one report cycle, 15 min) and the one before did not | the first `running` after it → `tunnel_recovered` (info) | `api/tunnel_test.go` |
|
||||
| `os_update_stale` | warning | no successful OS leg for **7 days** while the switch is ON (agents that can run the leg only); the mail names the likely reason (box not reporting / the last leg's failure / no good night backup) | a successful leg | `TestAlarm_StaleLeg`, `TestAlarm_StaleNamesTheReason` |
|
||||
| `os_reboot_needed` | warning | the host has needed a reboot for **14 days** (from the FIRST scanned report that said so) | a scanned pass that finds nothing (agent ≥ 0.141.1 scans the host every pass) | `TestAlarm_RebootNeeded`, `TestRebootNeeded_ClearedByAScannedPass` |
|
||||
| `os_ring0_stalled` | error | ring 0 approved nothing for **7 days** in a layer while it has pending FAST-lane updates (a pending kernel does not count) | a new release | `TestAlarm_Ring0Stalled` |
|
||||
| `os_not_covered` | warning | a ring-1 box has had fast-lane packages no approved release names for **14 days** | the packages are covered or gone | `TestAlarm_NotCovered` |
|
||||
|
||||
- **`unknown` never alarms** (R-96 rule 3): a probe that could not ask is neither up nor down. An `unknown` report
|
||||
breaks a `not_running` run.
|
||||
- **A stopped cloudflared heals itself before the hub can see it** (measured 2026-10-04): the controller's
|
||||
protected-container check recreates it within 5 minutes, and the host reports every 15. So `tunnel_down` catches
|
||||
what the box cannot heal — a running container with no connection (wrong token, blocked network).
|
||||
- The OS alarms are checked **hourly**, re-sent at most **once a week** while true, and forgotten when false, so the
|
||||
next occurrence is announced again. The four numbers are configuration (`OS_ALARM_STALE_AFTER`,
|
||||
`OS_ALARM_REBOOT_AFTER`, `OS_ALARM_RING0_STALL_AFTER`, `OS_ALARM_NOT_COVERED_AFTER`) — *decided by CC unattended,
|
||||
operator may reverse* (`11` §8.3).
|
||||
|
||||
---
|
||||
|
||||
## 7. The intent test [DESIGN, R-386 — CLOSED controller v0.223.0]
|
||||
|
||||
**"The customer stopped this" is asked of the FIELD THAT RECORDS IT, never inferred from the state.**
|
||||
|
||||
@@ -778,6 +778,22 @@ its length, and both fixes cost something the household would notice — operato
|
||||
83. **Next: the tunnel status (R-841), the host fast lane (`11` §8 step 3), and other OS-update improvements.**
|
||||
*Operator ruling 2026-10-04 ~12:20.*
|
||||
|
||||
### 2026-10-04 (afternoon) — decided by CC unattended — operator may reverse (host fast lane brief)
|
||||
|
||||
84. **Which record proves a box is an appliance, for the host fast lane?** Options: (a) `agent.json`
|
||||
`deployment_mode` — the agent can write it, so a compromised agent could claim "appliance" and unlock host
|
||||
updates; (b) the installer's ROOT-owned `/var/lib/felhom-install/state.json` `mode` — written once, as root, at
|
||||
install. **Chosen (b):** the wrapper is the root fence and must not trust a file the agent can change. Cost: a box
|
||||
whose install record is missing gets no host step (fails closed). `11` §8.2.
|
||||
85. **The four OS alarm thresholds.** Options: shorter (3/7 days — noisy: a weekend away alarms) or longer (14/30 —
|
||||
a stopped box goes unseen for weeks). **Chosen:** no OS leg 7 days, reboot needed 14, ring 0 stalled 7, not covered
|
||||
14; all four are configuration. Cost: a real stop is seen after a week, not a night. `11` §8.3, `08` §6.3.
|
||||
86. **A second release of the agent (v0.141.1) and the hub (v0.131.1) in the same session**, against "one release per
|
||||
repo". Options: (a) keep v0.141.0 / v0.131.0 and file the defect — the host "reboot needed" stays wrong (it hid
|
||||
`lxc-start`, and a reboot never cleared it), so the new 14-day alarm would fire on rebooted hosts; (b) patch now.
|
||||
**Chosen (b):** a known-false operator alarm is worse than an extra release; both patches are small and red-proved.
|
||||
R-846.
|
||||
|
||||
### 2026-09-30 (day) — operator notes, recorded before the work
|
||||
|
||||
- **The day brief runs by day.** Every backup and automatic-update test is started by hand — the night chain's debug
|
||||
|
||||
@@ -324,8 +324,8 @@ must never overlap a backup, a restore-test or a self-update.~~
|
||||
|---|---|---|
|
||||
| Guest packages | Restore the guest snapshot taken just before the update | It also undoes app data written after the snapshot. Use it only inside the health window, before apps have written much. After that, install the previous version (needs OPEN Q1). **[FACT] (C6)** Customer guests are on LVM-thin and can snapshot; scratch 9202 (`dir` storage) cannot (`snapshot feature is not available`), so the measured undo was a whole-guest backup + restore: **73 s down**, all apps healthy, libc back. The snapshot rollback itself is unmeasured (R-837). |
|
||||
| Docker engine | Install the previous version (Docker's repository keeps old versions — OPEN Q2) | Every container restarts again. **[FACT] (C5)** Docker keeps 46 `docker-ce` versions. A step (either way) without `live-restore`: 6 of 6 containers restart, the app silent **26.5–30 s**, healthy at +42–45 s. With `live-restore` on: **0 restarts, no gap**, also across a containerd step. But a restart that turns `live-restore` OFF stops every container and starts NONE (`unless-stopped` ignored) — on 9202 they stayed down until restarted by hand; `systemctl reload` turns it on but not off. |
|
||||
| Host packages | Install the previous version | Only if the source still has it (OPEN Q1/Q2). **[FACT]** `rsync` back to its pre-update version: `not found`; `libpng16-16t64` back to the point-release version: worked. A Debian undo needs the snapshot archive (C2). |
|
||||
| Host kernel | ~~Boot the previous kernel. Proxmox can boot a new kernel **once** (`proxmox-boot-tool kernel pin <ver> --next-boot`). If that boot fails, the next boot uses the old kernel again. Make the new kernel permanent only after a healthy boot.~~ **[FACT] (C4)** Both demo hosts boot UEFI + GRUB (no proxmox-boot-tool ESPs). **Installing a kernel makes it the GRUB default at once.** `--next-boot` on GRUB writes an ordinary `GRUB_DEFAULT` + `update-grub`; `proxmox-boot-cleanup.service` clears it only once a boot reaches userspace. Measured on demo-hp (old kernel pinned permanently FIRST, new pinned for next boot): boot 1 → `7.0.14-20-pve`, healthy, 60 s; boot 2 → back on `7.0.2-6-pve`, healthy, 76 s. **So the fallback works after a boot that succeeds; read from the code, a kernel that hangs before userspace stays the default on every power cycle.** `[PROPOSAL]` use GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel as the saved default — unmeasured (R-836). | ~~If the new kernel hangs, someone must switch the box off and on. The spike checks whether a hardware watchdog can do that (OPEN Q4).~~ **[FACT]** Only `softdog` runs (loaded by `watchdog-mux`); it cannot rescue a kernel that never boots. demo-hp has an AMD FCH whose `sp5100_tco` driver ships but is not loaded; untested. A hang still needs a person. |
|
||||
| Host packages | Install the previous version | Only if the source still has it (OPEN Q1/Q2). **[FACT]** `rsync` back to its pre-update version: `not found`; `libpng16-16t64` back to the point-release version: worked. A Debian undo needs the snapshot archive (C2). **[FACT, 2026-10-04]** Proved by hand: `runbooks/os-updates-host-undo.md` (demo-hp, `tzdata` back one version from `snapshot.debian.org`, held, released). Use the NEW version's `first_seen` as the timestamp when there is no previous host release; a package that pins its siblings (`eject` → `libmount1 =`) goes back only with them. |
|
||||
| Host kernel | ~~Boot the previous kernel. Proxmox can boot a new kernel **once** (`proxmox-boot-tool kernel pin <ver> --next-boot`). If that boot fails, the next boot uses the old kernel again. Make the new kernel permanent only after a healthy boot.~~ **[FACT] (C4)** Both demo hosts boot UEFI + GRUB (no proxmox-boot-tool ESPs). **Installing a kernel makes it the GRUB default at once.** `--next-boot` on GRUB writes an ordinary `GRUB_DEFAULT` + `update-grub`; `proxmox-boot-cleanup.service` clears it only once a boot reaches userspace. Measured on demo-hp (old kernel pinned permanently FIRST, new pinned for next boot): boot 1 → `7.0.14-20-pve`, healthy, 60 s; boot 2 → back on `7.0.2-6-pve`, healthy, 76 s. **So the fallback works after a boot that succeeds; read from the code, a kernel that hangs before userspace stays the default on every power cycle.** ~~`[PROPOSAL]` use GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel as the saved default — unmeasured (R-836).~~ **[FACT, 2026-10-04, demo-hp, operator's word before each reboot] GRUB's one-shot is NOT a one-shot here either.** With `GRUB_DEFAULT=saved` (old 7.0.2-6 saved) and `grub-reboot` 7.0.14-20: boot 1 → 7.0.14-20, **Secure Boot ON, booted fine** (signed kernel, shim → GRUB); but `/boot` is ext4 on LVM, GRUB cannot write its environment block there (`grub-reboot` warns so itself), `next_entry` was never cleared, and boot 2 with no command → **7.0.14-20 again**. A kernel lane needs a writable env block (the ESP) or a userspace "boot good" step — R-836. demo-hp left on 7.0.14-20, saved default 7.0.14-20, both kernels installed (`audits/os-host-lane-2026-10-04/partE/`). | ~~If the new kernel hangs, someone must switch the box off and on. The spike checks whether a hardware watchdog can do that (OPEN Q4).~~ **[FACT]** Only `softdog` runs (loaded by `watchdog-mux`); it cannot rescue a kernel that never boots. demo-hp has an AMD FCH whose `sp5100_tco` driver ships but is not loaded; untested. A hang still needs a person. **[FACT, 2026-10-04]** `sp5100_tco` is blacklisted by the Proxmox kernel package; loaded by hand it answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0 — read from sysfs, never opened, so never armed; unloaded). `kernel.panic = 0`: a panic leaves the host stopped (R-851). |
|
||||
|
||||
### 5.7 Telling people
|
||||
|
||||
@@ -336,6 +336,30 @@ must never overlap a backup, a restore-test or a self-update.~~
|
||||
whether the box restarted. Telling households in advance that the box may restart at night is a
|
||||
**promise to users**. That is the operator's decision when the slow lane is built.
|
||||
|
||||
### 5.8 The Docker engine slow lane — DESIGN (2026-10-04, nothing built) `[PROPOSAL]`
|
||||
|
||||
Built from C5 (measured) and R-835. **The one decision it needs is in STATUS: `live-restore` on, fleet-wide.**
|
||||
|
||||
- **What moves:** `docker-ce`, `docker-ce-cli`, `containerd.io`, `docker-buildx-plugin`, `docker-compose-plugin`,
|
||||
`docker-ce-rootless-extras` in the customer guest — today the six "not covered" packages on every box. One
|
||||
approved **engine set** at a time, like a controller floor: the operator approves it (slow lane, §5.2), after ring 0
|
||||
has run it for at least 2 nights healthy. Never two steps in one night.
|
||||
- **Precondition: `live-restore` ON.** Without it an engine step restarts every container — 26.5–30 s of silence,
|
||||
healthy at +42–45 s (C5). With it: 0 restarts, no gap, also across a containerd step. Turning it ON is safe
|
||||
(`systemctl reload docker` applies it without a restart, C5); turning it OFF later by a plain restart stops every
|
||||
container and starts none (R-835) — so it is turned on once, by the golden and by a one-time fleet step, and never
|
||||
turned off by the lane.
|
||||
- **Who and when:** the agent, through the same wrapper (`lane: slow`, refusal R3: only inside a verified signed
|
||||
operator job, R-530's mechanism), in the guest, after the guest and host fast-lane steps, under the same heavy-op
|
||||
gate, on a night the operator scheduled. Debian origin rule replaced by "origin `Docker CE`, exactly these names".
|
||||
- **Health:** the guest rule (§8.1) plus `docker version` reports the approved engine, and every container running
|
||||
at the start is running with the SAME container id (proof that `live-restore` held). A changed id is
|
||||
`health_failed` even if the app is healthy — it means the households' apps restarted when they should not have.
|
||||
- **Undo:** install the previous engine set (Docker's repository keeps 46 versions, C2) — by an operator job, with
|
||||
`live-restore` still on, so the undo is also restart-free.
|
||||
- **Not covered here:** the golden's own engine (baked weekly; a new golden carries the approved set), and BYO hosts
|
||||
(the guest is ours on both, so the lane applies there too).
|
||||
|
||||
---
|
||||
|
||||
## 6. Risks and edge cases
|
||||
@@ -405,8 +429,9 @@ Each step returns to the operator for go or no-go.
|
||||
1. **Spike** (measure Q1–Q10; no product code).
|
||||
2. **Guest Debian, fast lane.** ~~Lowest risk: a snapshot undo exists.~~ **BUILT 2026-10-04** — agent v0.140.0, hub
|
||||
v0.130.0, installer 1.29.0; §8.1. **There is no snapshot undo** (R-837, measured).
|
||||
3. **Host Debian, fast lane** (no kernel, no Proxmox packages).
|
||||
4. **Fleet view and alarms** (§5.7).
|
||||
3. **Host Debian, fast lane** (no kernel, no Proxmox packages). **BUILT 2026-10-04** — agent v0.141.1, hub v0.131.1;
|
||||
§8.2.
|
||||
4. **Fleet view and alarms** (§5.7). **BUILT 2026-10-04** — hub v0.131.0/v0.131.1; §8.3.
|
||||
5. **Slow lane: Docker engine.**
|
||||
6. **Slow lane: host kernel and Proxmox packages, with the reboot.**
|
||||
7. **Later:** the Proxmox major upgrade (PVE 9 → 10), drilled on ring 0 first.
|
||||
@@ -454,6 +479,47 @@ Evidence: `audits/os-guest-lane-2026-10-04/` (parts A–G). Brief: guest fast la
|
||||
- **Not delivered to existing boxes by the product**: the wrapper and the sudoers line reach a box only through the
|
||||
installer; the signed agent update replaces the binary only (R-840). The demo boxes got them BY HAND.
|
||||
|
||||
### 8.2 Step 3 as BUILT (2026-10-04) `[FACT]`
|
||||
|
||||
Evidence: `audits/os-host-lane-2026-10-04/` (parts A–G).
|
||||
|
||||
- **Where it runs:** appliances only. The proof is the ROOT-owned install record `/var/lib/felhom-install/state.json`
|
||||
`mode: appliance` (written by the installer as root); the agent-writable `agent.json` `deployment_mode` is not
|
||||
trusted for this. A BYO host gets no host step (wrapper refusal **R12**, now lifted only for lane fast / layer host
|
||||
on an appliance). *Decided by CC unattended — operator may reverse.*
|
||||
- **What:** origin `Debian` / `Debian-Security` only, and never a kernel, boot or firmware package (name pattern
|
||||
`HOST_SLOW_RE` — `linux-*`, `proxmox-kernel*`, `pve-kernel*`, `pve-firmware`, `firmware-*`, `grub*`, `shim*`,
|
||||
`systemd-boot*`, `*-microcode`, `efibootmgr`; refusal **R14**). The hub leaves the same names out of the host
|
||||
candidate.
|
||||
- **When:** in the same leg, after the guest step, under the same heavy-op gate. A failed or unhealthy guest step
|
||||
skips the host step.
|
||||
- **The host health rule** (`HostHealthVerdict`, pinned in `internal/osupdate`): `felhom-agent`, `pveproxy`,
|
||||
`pvedaemon`, `pvestatd` and `pve-cluster` are `active`; the customer guest runs; the guest health rule (§8.1)
|
||||
passes; the tunnel is `running` — an `unknown` tunnel does not fail it, `not_running` does. Same 5-minute wait.
|
||||
- **Separate approved sets:** host and guest releases are separate (`os-host-…`, `os-guest-…`), each by the same rule
|
||||
(24 h, 1 night of THAT layer, every ring-0 box). Ring 1 receives `host_release` beside `release`.
|
||||
- **Reboot needed:** reported when PID 1 or `lxc-start` maps a replaced file, with the date of the first scanned
|
||||
report that said so. **Never reboots.** The host is scanned on every pass, so a reboot clears it (v0.141.1 — v0.141.0
|
||||
hid `lxc-start` and never cleared, R-846).
|
||||
- **Undo:** by hand, `runbooks/os-updates-host-undo.md` (proved). No automatic undo.
|
||||
- **Speed (R-845):** one call per layer; both steps with nothing to install 23–32 s; a 108-package host pass 70 s.
|
||||
- **Measured live:** demo-felhom ring 0 installed 108 Debian host packages, all Debian origin (checked against apt),
|
||||
healthy; a 605-package host release approved (TEST wait 2 min / 0 nights, logged, reverted to 24 h + 1 night);
|
||||
demo-felhom as ring 1 installed exactly the one version it lacked (604 already current).
|
||||
|
||||
### 8.3 Step 4 as BUILT (2026-10-04) `[FACT]`
|
||||
|
||||
- **The fleet view** (`GET /os/fleet`, operator): one line per box — ring, switch, the tunnel, and per layer: the
|
||||
release, the last outcome, the last successful leg, pending, not covered, restart needed, reboot needed since, the
|
||||
wrapper's own seconds.
|
||||
- **The four alarms** (`08` §6.3), operator-only, hourly, at most weekly while true: no successful OS leg for
|
||||
**7 days** while the switch is ON (naming the likely reason); reboot needed for **14 days**; ring 0 approved nothing
|
||||
for **7 days** while it has pending fast-lane updates; not-covered fast-lane packages for **14 days**. The four
|
||||
numbers are configuration (`OS_ALARM_*`). *Decided by CC unattended — operator may reverse:* 7 days = a week of
|
||||
missed nights is past any normal hiccup (a box off for a weekend does not alarm); 14 days for reboot and coverage =
|
||||
two weekly golden cycles, both need a person anyway.
|
||||
- **The tunnel** (R-841): `running` / `not_running` / `unknown`; `tunnel_down` after two `not_running` reports.
|
||||
|
||||
## 9. Where the rest lives
|
||||
|
||||
- The finding: **R-812** (`backlog/OPEN-ITEMS.md`). The intention: **R-808** (`backlog/ROADMAP.md`).
|
||||
|
||||
@@ -3,3 +3,17 @@
|
||||
11:36:47 container=running/unhealthy hub=running alarm=none
|
||||
11:37:49 container=running/unhealthy hub=not_running alarm=none
|
||||
11:38:51 container=running/unhealthy hub=not_running alarm=none
|
||||
11:39:53 container=running/unhealthy hub=not_running alarm=none
|
||||
11:40:55 container=running/unhealthy hub=not_running alarm=none
|
||||
11:41:57 container=running/unhealthy hub=not_running alarm=none
|
||||
11:42:59 container=running/unhealthy hub=not_running alarm=none
|
||||
11:44:01 container=running/unhealthy hub=not_running alarm=none
|
||||
11:45:03 container=running/unhealthy hub=not_running alarm=none
|
||||
11:46:05 container=running/unhealthy hub=not_running alarm=none
|
||||
11:47:07 container=running/unhealthy hub=not_running alarm=none
|
||||
11:48:09 container=running/unhealthy hub=not_running alarm=none
|
||||
11:49:10 container=running/unhealthy hub=not_running alarm=none
|
||||
11:50:12 container=running/unhealthy hub=not_running alarm=none
|
||||
11:51:14 container=running/unhealthy hub=not_running alarm=none
|
||||
11:52:16 container=running/unhealthy hub=not_running alarm=none
|
||||
11:53:18 container=running/unhealthy hub=not_running alarm=2026/10/04 13:52:29 [WARN] host demo-hp-bb76ea tunnel: tunnel_down (container running but the tunnel is NOT connected (cloudflared /ready fails))
|
||||
|
||||
@@ -3,3 +3,6 @@ Chain DOCKER-USER (1 references)
|
||||
target prot opt source destination
|
||||
DROP tcp -- 0.0.0.0/0 0.0.0.0/0 tcp dpt:7844
|
||||
DROP udp -- 0.0.0.0/0 0.0.0.0/0 udp dpt:7844
|
||||
unblock at 2026-10-04T11:53:29Z
|
||||
Chain DOCKER-USER (1 references)
|
||||
target prot opt source destination
|
||||
|
||||
@@ -0,0 +1,14 @@
|
||||
11:53:39 container=unhealthy hub=not_running
|
||||
11:54:41 container=healthy hub=not_running
|
||||
11:55:43 container=healthy hub=not_running
|
||||
11:56:45 container=healthy hub=not_running
|
||||
11:57:47 container=healthy hub=not_running
|
||||
11:58:48 container=healthy hub=not_running
|
||||
11:59:50 container=healthy hub=not_running
|
||||
12:00:52 container=healthy hub=not_running
|
||||
12:01:54 container=healthy hub=not_running
|
||||
12:02:56 container=healthy hub=not_running
|
||||
12:03:58 container=healthy hub=not_running
|
||||
12:05:00 container=healthy hub=not_running
|
||||
12:06:02 container=healthy hub=not_running
|
||||
12:07:04 container=healthy hub=not_running
|
||||
@@ -0,0 +1,5 @@
|
||||
2026/10/04 13:52:29 [WARN] host demo-hp-bb76ea tunnel: tunnel_down (container running but the tunnel is NOT connected (cloudflared /ready fails))
|
||||
2026/10/04 13:52:30 [INFO] Operator email sent for demo-hp/tunnel_down
|
||||
2026/10/04 13:52:29 [WARN] host demo-hp-bb76ea tunnel: tunnel_down (container running but the tunnel is NOT connected (cloudflared /ready fails))
|
||||
2026/10/04 13:52:30 [INFO] Operator email sent for demo-hp/tunnel_down
|
||||
2026/10/04 14:07:29 [WARN] host demo-hp-bb76ea tunnel: tunnel_recovered (connected)
|
||||
@@ -0,0 +1,135 @@
|
||||
=== felhom-agent 0.141.1 selftest=os-update vmid=9201 ring=0 enabled=true guest-release=false host-release=false appliance=true ===
|
||||
time=2026-10-04T14:22:52.476+02:00 level=INFO msg="osupdate: START" run=20261004T122252Z layer=guest vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T122252Z
|
||||
time=2026-10-04T14:23:08.079+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T122252Z layer=guest:9201 lane=fast mode=apply select=pending-fast packages=0"
|
||||
time=2026-10-04T14:23:08.079+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
|
||||
time=2026-10-04T14:23:08.079+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0"
|
||||
time=2026-10-04T14:23:08.079+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)"
|
||||
time=2026-10-04T14:23:08.080+02:00 level=INFO msg="osupdate: DONE" run=20261004T122252Z layer=guest vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=6 not_covered=6 restart_needed=0 reboot_needed=false wrapper_seconds=15.5
|
||||
time=2026-10-04T14:23:08.089+02:00 level=INFO msg="osupdate: START" run=20261004T122252Z layer=host vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T122252Z
|
||||
time=2026-10-04T14:23:22.788+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T122252Z layer=host lane=fast mode=apply select=pending-fast packages=0"
|
||||
time=2026-10-04T14:23:22.788+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
|
||||
time=2026-10-04T14:23:22.788+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0"
|
||||
time=2026-10-04T14:23:22.788+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)"
|
||||
time=2026-10-04T14:23:23.852+02:00 level=INFO msg="osupdate: DONE" run=20261004T122252Z layer=host vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=78 not_covered=78 restart_needed=0 reboot_needed=false wrapper_seconds=14.6
|
||||
--- os-update report (guest) ---
|
||||
{
|
||||
"health_reason": "",
|
||||
"healthy": true,
|
||||
"mode": "apply",
|
||||
"not_covered": [
|
||||
"docker-ce-cli",
|
||||
"containerd.io",
|
||||
"docker-ce",
|
||||
"docker-buildx-plugin",
|
||||
"docker-ce-rootless-extras",
|
||||
"docker-compose-plugin"
|
||||
],
|
||||
"outcome": "nothing",
|
||||
"pending": 6,
|
||||
"reboot_needed": false,
|
||||
"refused": null,
|
||||
"release_id": "ring0-20261004T122252Z",
|
||||
"restart_needed": null,
|
||||
"ring": 0,
|
||||
"run_id": "20261004T122252Z",
|
||||
"upgraded": [],
|
||||
"wrapper_seconds": 15.5
|
||||
}
|
||||
--- os-update report (host) ---
|
||||
{
|
||||
"health_reason": "",
|
||||
"healthy": true,
|
||||
"mode": "apply",
|
||||
"not_covered": [
|
||||
"frr",
|
||||
"shim-signed-common",
|
||||
"proxmox-secure-boot-support",
|
||||
"shim-unsigned",
|
||||
"shim-helpers-amd64-signed",
|
||||
"shim-signed",
|
||||
"amd64-microcode",
|
||||
"libradosstriper1",
|
||||
"librgw2",
|
||||
"ceph-common",
|
||||
"librbd1",
|
||||
"librados2",
|
||||
"python3-cephfs",
|
||||
"libcephfs2",
|
||||
"python3-rgw",
|
||||
"python3-rados",
|
||||
"python3-ceph-argparse",
|
||||
"python3-ceph-common",
|
||||
"python3-rbd",
|
||||
"ceph-fuse",
|
||||
"chrony",
|
||||
"libcorosync-common4",
|
||||
"libcfg7",
|
||||
"libcmap4",
|
||||
"libcpg4",
|
||||
"libknet1t64",
|
||||
"libnozzle1t64",
|
||||
"libquorum5",
|
||||
"libvotequorum8",
|
||||
"corosync",
|
||||
"frr-pythontools",
|
||||
"libjs-extjs",
|
||||
"libnvpair3linux",
|
||||
"libproxmox-acme-plugins",
|
||||
"libproxmox-backup-qemu0",
|
||||
"pve-qemu-kvm",
|
||||
"libpve-notify-perl",
|
||||
"libpve-cluster-api-perl",
|
||||
"libpve-cluster-perl",
|
||||
"pve-cluster",
|
||||
"libpve-access-control",
|
||||
"libpve-apiclient-perl",
|
||||
"librados2-perl",
|
||||
"proxmox-backup-client",
|
||||
"proxmox-backup-file-restore",
|
||||
"pve-manager",
|
||||
"libproxmox-acme-perl",
|
||||
"libpve-common-perl",
|
||||
"libpve-guest-common-perl",
|
||||
"qemu-server",
|
||||
"libpve-storage-perl",
|
||||
"pve-edk2-firmware-legacy",
|
||||
"pve-edk2-firmware-ovmf",
|
||||
"libpve-network-api-perl",
|
||||
"libpve-network-perl",
|
||||
"proxmox-firewall-data",
|
||||
"pve-firewall",
|
||||
"pve-container",
|
||||
"pve-ha-manager",
|
||||
"novnc-pve",
|
||||
"proxmox-enterprise-support-keyring",
|
||||
"proxmox-mini-journalreader",
|
||||
"proxmox-widget-toolkit",
|
||||
"pve-docs",
|
||||
"pve-i18n",
|
||||
"pve-xtermjs",
|
||||
"pve-yew-mobile-i18n",
|
||||
"pve-yew-mobile-gui",
|
||||
"libuutil3linux",
|
||||
"libzfs7linux",
|
||||
"libzpool7linux",
|
||||
"proxmox-kernel-helper",
|
||||
"pve-edk2-firmware-aarch64",
|
||||
"pve-edk2-firmware",
|
||||
"pve-firmware",
|
||||
"zfs-initramfs",
|
||||
"zfsutils-linux",
|
||||
"zfs-zed"
|
||||
],
|
||||
"outcome": "nothing",
|
||||
"pending": 78,
|
||||
"reboot_needed": false,
|
||||
"refused": null,
|
||||
"release_id": "ring0-20261004T122252Z",
|
||||
"restart_needed": [],
|
||||
"ring": 0,
|
||||
"run_id": "20261004T122252Z",
|
||||
"upgraded": [],
|
||||
"wrapper_seconds": 14.6
|
||||
}
|
||||
pass took 31.4s
|
||||
WALL_SECONDS=31.520457543
|
||||
@@ -0,0 +1,2 @@
|
||||
{"ok":true}
|
||||
200
|
||||
+43
@@ -0,0 +1,43 @@
|
||||
+ PKG=tzdata
|
||||
++ grep -oE 'tzdata:amd64 \([^,]+' /var/log/apt/history.log
|
||||
++ tail -1
|
||||
++ sed 's/.*(//'
|
||||
+ OLD=2026b-0+deb13u1
|
||||
++ dpkg-query -W '-f=${Version}' tzdata
|
||||
+ NEW=2026c-0+deb13u1
|
||||
+ echo OLD=2026b-0+deb13u1 NEW=2026c-0+deb13u1
|
||||
OLD=2026b-0+deb13u1 NEW=2026c-0+deb13u1
|
||||
++ curl -s 'https://snapshot.debian.org/mr/binary/tzdata/2026c-0+deb13u1/binfiles?fileinfo=1'
|
||||
++ python3 -c '
|
||||
import json,sys; d=json.load(sys.stdin)
|
||||
print(sorted(f["first_seen"] for v in d["fileinfo"].values() for f in v)[0])'
|
||||
+ TS=20260831T204404Z
|
||||
+ . /etc/os-release
|
||||
++ PRETTY_NAME='Debian GNU/Linux 13 (trixie)'
|
||||
++ NAME='Debian GNU/Linux'
|
||||
++ VERSION_ID=13
|
||||
++ VERSION='13 (trixie)'
|
||||
++ VERSION_CODENAME=trixie
|
||||
++ DEBIAN_VERSION_FULL=13.7
|
||||
++ ID=debian
|
||||
++ HOME_URL=https://www.debian.org/
|
||||
++ SUPPORT_URL=https://www.debian.org/support
|
||||
++ BUG_REPORT_URL=https://bugs.debian.org/
|
||||
+ printf 'deb [check-valid-until=no] http://snapshot.debian.org/archive/debian/%s %s main\ndeb [check-valid-until=no] http://snapshot.debian.org/archive/debian-security/%s %s-security main\n' 20260831T204404Z trixie 20260831T204404Z trixie
|
||||
+ apt-get -q update
|
||||
+ apt-get -s install --allow-downgrades tzdata=2026b-0+deb13u1
|
||||
+ grep -E '^(Inst|Remv)|downgraded'
|
||||
0 upgraded, 0 newly installed, 1 downgraded, 0 to remove and 78 not upgraded.
|
||||
Inst tzdata [2026c-0+deb13u1] (2026b-0+deb13u1 Debian:13.6/stable [all])
|
||||
+ DEBIAN_FRONTEND=noninteractive
|
||||
+ apt-get -y -q install --allow-downgrades tzdata=2026b-0+deb13u1
|
||||
+ rm -f /etc/apt/sources.list.d/felhom-undo-snapshot.list
|
||||
+ apt-get -q update
|
||||
+ dpkg-query -W tzdata
|
||||
tzdata 2026b-0+deb13u1
|
||||
+ ls /etc/apt/sources.list.d/
|
||||
ceph.sources
|
||||
debian.sources
|
||||
pve-enterprise.sources
|
||||
pve-no-subscription.sources
|
||||
tailscale.list
|
||||
+178
@@ -0,0 +1,178 @@
|
||||
=== felhom-agent 0.141.1 selftest=os-update vmid=9201 ring=1 enabled=true guest-release=true host-release=true appliance=true ===
|
||||
time=2026-10-04T14:40:46.795+02:00 level=INFO msg="osupdate: START" run=20261004T124046Z layer=guest vmid=9201 ring=1 trigger=debug enabled=true release=os-guest-20261004-123933
|
||||
time=2026-10-04T14:40:56.895+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=os-guest-20261004-123933 layer=guest:9201 lane=fast mode=apply select=listed packages=272"
|
||||
time=2026-10-04T14:40:56.895+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
|
||||
time=2026-10-04T14:40:56.895+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=272 not-installed=0 from-snapshot=0"
|
||||
time=2026-10-04T14:40:56.895+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)"
|
||||
time=2026-10-04T14:40:56.896+02:00 level=INFO msg="osupdate: DONE" run=20261004T124046Z layer=guest vmid=9201 ring=1 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=6 not_covered=6 restart_needed=0 reboot_needed=false wrapper_seconds=10.1
|
||||
time=2026-10-04T14:40:56.907+02:00 level=INFO msg="osupdate: START" run=20261004T124046Z layer=host vmid=9201 ring=1 trigger=debug enabled=true release=os-host-20261004-124034
|
||||
time=2026-10-04T14:41:10.034+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=os-host-20261004-124034 layer=host lane=fast mode=apply select=listed packages=605"
|
||||
time=2026-10-04T14:41:10.034+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
|
||||
time=2026-10-04T14:41:10.034+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=1 already=604 not-installed=0 from-snapshot=0"
|
||||
time=2026-10-04T14:41:10.034+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=1.4 upgraded=1 restart-needed=agetty,blkmapd,chronyd,cron,dbus-daemon,dmeventd,ksmtuned,lxc-monitord,lxc-start,lxcfs,pmxcfs,proxmox-firewal,pve-firewall,pve-ha-crm,pve-ha-lrm,pve-lxc-syscall,pvedaemon,pvedaemon worke,pvefw-logger,pveproxy,pveproxy worker,pvescheduler,pvestatd,qmeventd,rpcbind,rrdcached,smartd,spiceproxy,spiceproxy work,sshd,systemd-logind,systemd-udevd,watchdog-mux,zed reboot-needed=yes"
|
||||
time=2026-10-04T14:41:10.815+02:00 level=INFO msg="osupdate: DONE" run=20261004T124046Z layer=host vmid=9201 ring=1 trigger=debug outcome=applied healthy=true reason="" upgraded=1 pending=80 not_covered=80 restart_needed=34 reboot_needed=true wrapper_seconds=13.1
|
||||
--- os-update report (guest) ---
|
||||
{
|
||||
"health_reason": "",
|
||||
"healthy": true,
|
||||
"mode": "apply",
|
||||
"not_covered": [
|
||||
"docker-ce-cli",
|
||||
"containerd.io",
|
||||
"docker-ce",
|
||||
"docker-buildx-plugin",
|
||||
"docker-ce-rootless-extras",
|
||||
"docker-compose-plugin"
|
||||
],
|
||||
"outcome": "nothing",
|
||||
"pending": 6,
|
||||
"reboot_needed": false,
|
||||
"refused": null,
|
||||
"release_id": "os-guest-20261004-123933",
|
||||
"restart_needed": null,
|
||||
"ring": 1,
|
||||
"run_id": "20261004T124046Z",
|
||||
"upgraded": [],
|
||||
"wrapper_seconds": 10.1
|
||||
}
|
||||
--- os-update report (host) ---
|
||||
{
|
||||
"health_reason": "",
|
||||
"healthy": true,
|
||||
"mode": "apply",
|
||||
"not_covered": [
|
||||
"frr",
|
||||
"shim-signed-common",
|
||||
"shim-unsigned",
|
||||
"shim-helpers-amd64-signed",
|
||||
"shim-signed",
|
||||
"libradosstriper1",
|
||||
"librgw2",
|
||||
"ceph-common",
|
||||
"librbd1",
|
||||
"librados2",
|
||||
"python3-cephfs",
|
||||
"libcephfs2",
|
||||
"python3-rgw",
|
||||
"python3-rados",
|
||||
"python3-ceph-argparse",
|
||||
"python3-ceph-common",
|
||||
"python3-rbd",
|
||||
"ceph-fuse",
|
||||
"chrony",
|
||||
"libcorosync-common4",
|
||||
"libcfg7",
|
||||
"libcmap4",
|
||||
"libcpg4",
|
||||
"libknet1t64",
|
||||
"libnozzle1t64",
|
||||
"libquorum5",
|
||||
"libvotequorum8",
|
||||
"corosync",
|
||||
"frr-pythontools",
|
||||
"libjs-extjs",
|
||||
"libnvpair3linux",
|
||||
"libproxmox-acme-plugins",
|
||||
"libproxmox-backup-qemu0",
|
||||
"pve-qemu-kvm",
|
||||
"libpve-notify-perl",
|
||||
"libpve-cluster-api-perl",
|
||||
"libpve-cluster-perl",
|
||||
"pve-cluster",
|
||||
"libpve-access-control",
|
||||
"libpve-apiclient-perl",
|
||||
"librados2-perl",
|
||||
"proxmox-backup-client",
|
||||
"proxmox-backup-file-restore",
|
||||
"pve-manager",
|
||||
"libproxmox-acme-perl",
|
||||
"libpve-common-perl",
|
||||
"libpve-guest-common-perl",
|
||||
"qemu-server",
|
||||
"libpve-storage-perl",
|
||||
"pve-edk2-firmware-legacy",
|
||||
"pve-edk2-firmware-ovmf",
|
||||
"libpve-network-api-perl",
|
||||
"libpve-network-perl",
|
||||
"proxmox-firewall-data",
|
||||
"pve-firewall",
|
||||
"pve-container",
|
||||
"pve-ha-manager",
|
||||
"novnc-pve",
|
||||
"proxmox-enterprise-support-keyring",
|
||||
"proxmox-mini-journalreader",
|
||||
"proxmox-widget-toolkit",
|
||||
"pve-docs",
|
||||
"pve-i18n",
|
||||
"pve-xtermjs",
|
||||
"pve-yew-mobile-i18n",
|
||||
"pve-yew-mobile-gui",
|
||||
"libuutil3linux",
|
||||
"libzfs7linux",
|
||||
"libzpool7linux",
|
||||
"proxmox-first-boot",
|
||||
"pve-firmware",
|
||||
"proxmox-kernel-7.0.14-20-pve-signed",
|
||||
"proxmox-kernel-7.0",
|
||||
"proxmox-kernel-helper",
|
||||
"pve-edk2-firmware-aarch64",
|
||||
"pve-edk2-firmware",
|
||||
"zfs-initramfs",
|
||||
"zfsutils-linux",
|
||||
"zfs-zed",
|
||||
"tailscale"
|
||||
],
|
||||
"outcome": "applied",
|
||||
"pending": 80,
|
||||
"reboot_needed": true,
|
||||
"refused": null,
|
||||
"release_id": "os-host-20261004-124034",
|
||||
"restart_needed": [
|
||||
"agetty",
|
||||
"blkmapd",
|
||||
"chronyd",
|
||||
"cron",
|
||||
"dbus-daemon",
|
||||
"dmeventd",
|
||||
"ksmtuned",
|
||||
"lxc-monitord",
|
||||
"lxc-start",
|
||||
"lxcfs",
|
||||
"pmxcfs",
|
||||
"proxmox-firewal",
|
||||
"pve-firewall",
|
||||
"pve-ha-crm",
|
||||
"pve-ha-lrm",
|
||||
"pve-lxc-syscall",
|
||||
"pvedaemon",
|
||||
"pvedaemon worke",
|
||||
"pvefw-logger",
|
||||
"pveproxy",
|
||||
"pveproxy worker",
|
||||
"pvescheduler",
|
||||
"pvestatd",
|
||||
"qmeventd",
|
||||
"rpcbind",
|
||||
"rrdcached",
|
||||
"smartd",
|
||||
"spiceproxy",
|
||||
"spiceproxy work",
|
||||
"sshd",
|
||||
"systemd-logind",
|
||||
"systemd-udevd",
|
||||
"watchdog-mux",
|
||||
"zed"
|
||||
],
|
||||
"ring": 1,
|
||||
"run_id": "20261004T124046Z",
|
||||
"upgraded": [
|
||||
{
|
||||
"name": "tzdata",
|
||||
"version": "2026c-0+deb13u1",
|
||||
"origin": ""
|
||||
}
|
||||
],
|
||||
"wrapper_seconds": 13.1
|
||||
}
|
||||
pass took 24.1s
|
||||
tzdata 2026c-0+deb13u1
|
||||
@@ -0,0 +1,11 @@
|
||||
+ PKG=eject
|
||||
+ OLD=2.41-5
|
||||
+ dpkg -s eject
|
||||
+ grep -E '^(Status|Version)'
|
||||
Status: install ok installed
|
||||
Version: 2.41.5-0+deb13u1
|
||||
+ curl -s 'https://snapshot.debian.org/mr/binary/eject/2.41-5/binfiles?fileinfo=1'
|
||||
+ python3 -c '
|
||||
import json,sys; d=json.load(sys.stdin)
|
||||
print(sorted(f["first_seen"] for v in d["fileinfo"].values() for f in v)[0])'
|
||||
20250510T015204Z
|
||||
@@ -0,0 +1,28 @@
|
||||
+ PKG=eject
|
||||
+ OLD=2.41-5
|
||||
+ TS=20250510T015204Z
|
||||
+ . /etc/os-release
|
||||
++ PRETTY_NAME='Debian GNU/Linux 13 (trixie)'
|
||||
++ NAME='Debian GNU/Linux'
|
||||
++ VERSION_ID=13
|
||||
++ VERSION='13 (trixie)'
|
||||
++ VERSION_CODENAME=trixie
|
||||
++ DEBIAN_VERSION_FULL=13.7
|
||||
++ ID=debian
|
||||
++ HOME_URL=https://www.debian.org/
|
||||
++ SUPPORT_URL=https://www.debian.org/support
|
||||
++ BUG_REPORT_URL=https://bugs.debian.org/
|
||||
+ cat
|
||||
++ date +%s
|
||||
+ S=1791116624
|
||||
+ apt-get -q update
|
||||
+ tail -3
|
||||
Get:7 http://snapshot.debian.org/archive/debian/20250510T015204Z trixie/main amd64 Packages [9681 kB]
|
||||
Fetched 9904 kB in 5s (2165 kB/s)
|
||||
Reading package lists...
|
||||
++ date +%s
|
||||
update_seconds=6
|
||||
+ echo update_seconds=6
|
||||
+ apt-get -s install --allow-downgrades eject=2.41-5
|
||||
+ grep -E '^(Inst|Remv|Conf)|downgraded|newly|Err|E:'
|
||||
E: Version '2.41-5' for 'eject' was not found
|
||||
@@ -0,0 +1,35 @@
|
||||
+ PKG=eject
|
||||
+ OLD=2.41-5
|
||||
++ dpkg-query -W '-f=${Version}' eject
|
||||
+ NEW=2.41.5-0+deb13u1
|
||||
++ curl -s 'https://snapshot.debian.org/mr/binary/eject/2.41.5-0+deb13u1/binfiles?fileinfo=1'
|
||||
++ python3 -c '
|
||||
import json,sys; d=json.load(sys.stdin)
|
||||
print(sorted(f["first_seen"] for v in d["fileinfo"].values() for f in v)[0])'
|
||||
TS=20260814T165831Z
|
||||
+ TS=20260814T165831Z
|
||||
+ echo TS=20260814T165831Z
|
||||
+ . /etc/os-release
|
||||
++ PRETTY_NAME='Debian GNU/Linux 13 (trixie)'
|
||||
++ NAME='Debian GNU/Linux'
|
||||
++ VERSION_ID=13
|
||||
++ VERSION='13 (trixie)'
|
||||
++ VERSION_CODENAME=trixie
|
||||
++ DEBIAN_VERSION_FULL=13.7
|
||||
++ ID=debian
|
||||
++ HOME_URL=https://www.debian.org/
|
||||
++ SUPPORT_URL=https://www.debian.org/support
|
||||
++ BUG_REPORT_URL=https://bugs.debian.org/
|
||||
+ cat
|
||||
+ apt-get -q update
|
||||
+ tail -1
|
||||
Reading package lists...
|
||||
+ apt-cache madison eject
|
||||
eject | 2.41.5-0+deb13u1 | http://deb.debian.org/debian trixie/main amd64 Packages
|
||||
eject | 2.41.5-0+deb13u1 | http://security.debian.org/debian-security trixie-security/main amd64 Packages
|
||||
eject | 2.41.5-0+deb13u1 | http://snapshot.debian.org/archive/debian-security/20260814T165831Z trixie-security/main amd64 Packages
|
||||
eject | 2.41-5 | http://snapshot.debian.org/archive/debian/20260814T165831Z trixie/main amd64 Packages
|
||||
+ apt-get -s install --allow-downgrades eject=2.41-5
|
||||
+ grep -E '^(Inst|Remv|Conf)|downgraded|newly|Err|E:'
|
||||
E: Unable to correct problems, you have held broken packages.
|
||||
E: The following information from --solver 3.0 may provide additional context:
|
||||
@@ -0,0 +1,37 @@
|
||||
+ rm -f /etc/apt/sources.list.d/felhom-undo-snapshot.list
|
||||
+ PKG=tzdata
|
||||
+ OLD=2026b-0+deb13u1
|
||||
++ dpkg-query -W '-f=${Version}' tzdata
|
||||
+ NEW=2026c-0+deb13u1
|
||||
+ echo NEW=2026c-0+deb13u1
|
||||
NEW=2026c-0+deb13u1
|
||||
++ curl -s 'https://snapshot.debian.org/mr/binary/tzdata/2026c-0+deb13u1/binfiles?fileinfo=1'
|
||||
++ python3 -c '
|
||||
import json,sys; d=json.load(sys.stdin)
|
||||
print(sorted(f["first_seen"] for v in d["fileinfo"].values() for f in v)[0])'
|
||||
TS=20260831T204404Z
|
||||
+ TS=20260831T204404Z
|
||||
+ echo TS=20260831T204404Z
|
||||
+ . /etc/os-release
|
||||
++ PRETTY_NAME='Debian GNU/Linux 13 (trixie)'
|
||||
++ NAME='Debian GNU/Linux'
|
||||
++ VERSION_ID=13
|
||||
++ VERSION='13 (trixie)'
|
||||
++ VERSION_CODENAME=trixie
|
||||
++ DEBIAN_VERSION_FULL=13.7
|
||||
++ ID=debian
|
||||
++ HOME_URL=https://www.debian.org/
|
||||
++ SUPPORT_URL=https://www.debian.org/support
|
||||
++ BUG_REPORT_URL=https://bugs.debian.org/
|
||||
+ cat
|
||||
+ apt-get -q update
|
||||
+ tail -1
|
||||
Reading package lists...
|
||||
+ apt-cache madison tzdata
|
||||
tzdata | 2026c-0+deb13u1 | http://deb.debian.org/debian trixie/main amd64 Packages
|
||||
tzdata | 2026b-0+deb13u1 | http://snapshot.debian.org/archive/debian/20260831T204404Z trixie/main amd64 Packages
|
||||
+ apt-get -s install --allow-downgrades tzdata=2026b-0+deb13u1
|
||||
+ grep -E '^(Inst|Remv|Conf)|downgraded|newly|Err|E:'
|
||||
0 upgraded, 0 newly installed, 1 downgraded, 0 to remove and 77 not upgraded.
|
||||
Inst tzdata [2026c-0+deb13u1] (2026b-0+deb13u1 Debian:13.6/stable [all])
|
||||
Conf tzdata (2026b-0+deb13u1 Debian:13.6/stable [all])
|
||||
@@ -0,0 +1,38 @@
|
||||
+ PKG=tzdata
|
||||
+ OLD=2026b-0+deb13u1
|
||||
++ date +%s
|
||||
+ S=1791116670
|
||||
+ DEBIAN_FRONTEND=noninteractive
|
||||
+ apt-get -y -q install --allow-downgrades tzdata=2026b-0+deb13u1
|
||||
+ grep -E '^(Unpacking|Setting up)|downgraded'
|
||||
0 upgraded, 0 newly installed, 1 downgraded, 0 to remove and 77 not upgraded.
|
||||
Unpacking tzdata (2026b-0+deb13u1) over (2026c-0+deb13u1) ...
|
||||
Setting up tzdata (2026b-0+deb13u1) ...
|
||||
+ apt-mark hold tzdata
|
||||
tzdata set on hold.
|
||||
+ rm -f /etc/apt/sources.list.d/felhom-undo-snapshot.list
|
||||
+ apt-get -q update
|
||||
+ tail -1
|
||||
Reading package lists...
|
||||
++ date +%s
|
||||
undo_seconds=4
|
||||
+ echo undo_seconds=4
|
||||
+ dpkg -s tzdata
|
||||
+ grep -E '^(Status|Version)'
|
||||
Status: hold ok installed
|
||||
Version: 2026b-0+deb13u1
|
||||
+ apt-mark showhold
|
||||
tzdata
|
||||
+ systemctl is-active pveproxy pvedaemon pvestatd pve-cluster felhom-agent
|
||||
active
|
||||
active
|
||||
active
|
||||
active
|
||||
active
|
||||
+ pct status 9201
|
||||
status: running
|
||||
+ ls /etc/apt/sources.list.d/
|
||||
ceph.sources
|
||||
debian.sources
|
||||
pve-enterprise.sources
|
||||
pve-no-subscription.sources
|
||||
@@ -0,0 +1,135 @@
|
||||
=== felhom-agent 0.141.1 selftest=os-update vmid=9201 ring=0 enabled=true guest-release=false host-release=false appliance=true ===
|
||||
time=2026-10-04T14:24:41.910+02:00 level=INFO msg="osupdate: START" run=20261004T122441Z layer=guest vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T122441Z
|
||||
time=2026-10-04T14:24:57.242+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T122441Z layer=guest:9201 lane=fast mode=apply select=pending-fast packages=0"
|
||||
time=2026-10-04T14:24:57.242+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
|
||||
time=2026-10-04T14:24:57.242+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0"
|
||||
time=2026-10-04T14:24:57.242+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)"
|
||||
time=2026-10-04T14:24:57.243+02:00 level=INFO msg="osupdate: DONE" run=20261004T122441Z layer=guest vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=6 not_covered=6 restart_needed=0 reboot_needed=false wrapper_seconds=15.3
|
||||
time=2026-10-04T14:24:57.251+02:00 level=INFO msg="osupdate: START" run=20261004T122441Z layer=host vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T122441Z
|
||||
time=2026-10-04T14:25:12.127+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T122441Z layer=host lane=fast mode=apply select=pending-fast packages=0"
|
||||
time=2026-10-04T14:25:12.127+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
|
||||
time=2026-10-04T14:25:12.127+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0"
|
||||
time=2026-10-04T14:25:12.127+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)"
|
||||
time=2026-10-04T14:25:13.136+02:00 level=INFO msg="osupdate: DONE" run=20261004T122441Z layer=host vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=78 not_covered=78 restart_needed=0 reboot_needed=false wrapper_seconds=14.8
|
||||
--- os-update report (guest) ---
|
||||
{
|
||||
"health_reason": "",
|
||||
"healthy": true,
|
||||
"mode": "apply",
|
||||
"not_covered": [
|
||||
"docker-ce-cli",
|
||||
"containerd.io",
|
||||
"docker-ce",
|
||||
"docker-buildx-plugin",
|
||||
"docker-ce-rootless-extras",
|
||||
"docker-compose-plugin"
|
||||
],
|
||||
"outcome": "nothing",
|
||||
"pending": 6,
|
||||
"reboot_needed": false,
|
||||
"refused": null,
|
||||
"release_id": "ring0-20261004T122441Z",
|
||||
"restart_needed": null,
|
||||
"ring": 0,
|
||||
"run_id": "20261004T122441Z",
|
||||
"upgraded": [],
|
||||
"wrapper_seconds": 15.3
|
||||
}
|
||||
--- os-update report (host) ---
|
||||
{
|
||||
"health_reason": "",
|
||||
"healthy": true,
|
||||
"mode": "apply",
|
||||
"not_covered": [
|
||||
"frr",
|
||||
"shim-signed-common",
|
||||
"proxmox-secure-boot-support",
|
||||
"shim-unsigned",
|
||||
"shim-helpers-amd64-signed",
|
||||
"shim-signed",
|
||||
"amd64-microcode",
|
||||
"libradosstriper1",
|
||||
"librgw2",
|
||||
"ceph-common",
|
||||
"librbd1",
|
||||
"librados2",
|
||||
"python3-cephfs",
|
||||
"libcephfs2",
|
||||
"python3-rgw",
|
||||
"python3-rados",
|
||||
"python3-ceph-argparse",
|
||||
"python3-ceph-common",
|
||||
"python3-rbd",
|
||||
"ceph-fuse",
|
||||
"chrony",
|
||||
"libcorosync-common4",
|
||||
"libcfg7",
|
||||
"libcmap4",
|
||||
"libcpg4",
|
||||
"libknet1t64",
|
||||
"libnozzle1t64",
|
||||
"libquorum5",
|
||||
"libvotequorum8",
|
||||
"corosync",
|
||||
"frr-pythontools",
|
||||
"libjs-extjs",
|
||||
"libnvpair3linux",
|
||||
"libproxmox-acme-plugins",
|
||||
"libproxmox-backup-qemu0",
|
||||
"pve-qemu-kvm",
|
||||
"libpve-notify-perl",
|
||||
"libpve-cluster-api-perl",
|
||||
"libpve-cluster-perl",
|
||||
"pve-cluster",
|
||||
"libpve-access-control",
|
||||
"libpve-apiclient-perl",
|
||||
"librados2-perl",
|
||||
"proxmox-backup-client",
|
||||
"proxmox-backup-file-restore",
|
||||
"pve-manager",
|
||||
"libproxmox-acme-perl",
|
||||
"libpve-common-perl",
|
||||
"libpve-guest-common-perl",
|
||||
"qemu-server",
|
||||
"libpve-storage-perl",
|
||||
"pve-edk2-firmware-legacy",
|
||||
"pve-edk2-firmware-ovmf",
|
||||
"libpve-network-api-perl",
|
||||
"libpve-network-perl",
|
||||
"proxmox-firewall-data",
|
||||
"pve-firewall",
|
||||
"pve-container",
|
||||
"pve-ha-manager",
|
||||
"novnc-pve",
|
||||
"proxmox-enterprise-support-keyring",
|
||||
"proxmox-mini-journalreader",
|
||||
"proxmox-widget-toolkit",
|
||||
"pve-docs",
|
||||
"pve-i18n",
|
||||
"pve-xtermjs",
|
||||
"pve-yew-mobile-i18n",
|
||||
"pve-yew-mobile-gui",
|
||||
"libuutil3linux",
|
||||
"libzfs7linux",
|
||||
"libzpool7linux",
|
||||
"proxmox-kernel-helper",
|
||||
"pve-edk2-firmware-aarch64",
|
||||
"pve-edk2-firmware",
|
||||
"pve-firmware",
|
||||
"zfs-initramfs",
|
||||
"zfsutils-linux",
|
||||
"zfs-zed"
|
||||
],
|
||||
"outcome": "nothing",
|
||||
"pending": 78,
|
||||
"reboot_needed": false,
|
||||
"refused": null,
|
||||
"release_id": "ring0-20261004T122441Z",
|
||||
"restart_needed": [],
|
||||
"ring": 0,
|
||||
"run_id": "20261004T122441Z",
|
||||
"upgraded": [],
|
||||
"wrapper_seconds": 14.8
|
||||
}
|
||||
pass took 31.2s
|
||||
tzdata 2026b-0+deb13u1
|
||||
@@ -0,0 +1,143 @@
|
||||
Canceled hold on tzdata.
|
||||
=== felhom-agent 0.141.1 selftest=os-update vmid=9201 ring=0 enabled=true guest-release=false host-release=false appliance=true ===
|
||||
time=2026-10-04T14:25:22.024+02:00 level=INFO msg="osupdate: START" run=20261004T122522Z layer=guest vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T122522Z
|
||||
time=2026-10-04T14:25:37.515+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T122522Z layer=guest:9201 lane=fast mode=apply select=pending-fast packages=0"
|
||||
time=2026-10-04T14:25:37.515+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
|
||||
time=2026-10-04T14:25:37.515+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0"
|
||||
time=2026-10-04T14:25:37.515+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)"
|
||||
time=2026-10-04T14:25:37.516+02:00 level=INFO msg="osupdate: DONE" run=20261004T122522Z layer=guest vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=6 not_covered=6 restart_needed=0 reboot_needed=false wrapper_seconds=15.4
|
||||
time=2026-10-04T14:25:37.526+02:00 level=INFO msg="osupdate: START" run=20261004T122522Z layer=host vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T122522Z
|
||||
time=2026-10-04T14:25:57.778+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T122522Z layer=host lane=fast mode=apply select=pending-fast packages=0"
|
||||
time=2026-10-04T14:25:57.778+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
|
||||
time=2026-10-04T14:25:57.778+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=1 already=0 not-installed=0 from-snapshot=0"
|
||||
time=2026-10-04T14:25:57.778+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=2.0 upgraded=1 restart-needed=- reboot-needed=no"
|
||||
time=2026-10-04T14:25:58.809+02:00 level=INFO msg="osupdate: DONE" run=20261004T122522Z layer=host vmid=9201 ring=0 trigger=debug outcome=applied healthy=true reason="" upgraded=1 pending=78 not_covered=78 restart_needed=0 reboot_needed=false wrapper_seconds=20.2
|
||||
--- os-update report (guest) ---
|
||||
{
|
||||
"health_reason": "",
|
||||
"healthy": true,
|
||||
"mode": "apply",
|
||||
"not_covered": [
|
||||
"docker-ce-cli",
|
||||
"containerd.io",
|
||||
"docker-ce",
|
||||
"docker-buildx-plugin",
|
||||
"docker-ce-rootless-extras",
|
||||
"docker-compose-plugin"
|
||||
],
|
||||
"outcome": "nothing",
|
||||
"pending": 6,
|
||||
"reboot_needed": false,
|
||||
"refused": null,
|
||||
"release_id": "ring0-20261004T122522Z",
|
||||
"restart_needed": null,
|
||||
"ring": 0,
|
||||
"run_id": "20261004T122522Z",
|
||||
"upgraded": [],
|
||||
"wrapper_seconds": 15.4
|
||||
}
|
||||
--- os-update report (host) ---
|
||||
{
|
||||
"health_reason": "",
|
||||
"healthy": true,
|
||||
"mode": "apply",
|
||||
"not_covered": [
|
||||
"frr",
|
||||
"shim-signed-common",
|
||||
"proxmox-secure-boot-support",
|
||||
"shim-unsigned",
|
||||
"shim-helpers-amd64-signed",
|
||||
"shim-signed",
|
||||
"amd64-microcode",
|
||||
"libradosstriper1",
|
||||
"librgw2",
|
||||
"ceph-common",
|
||||
"librbd1",
|
||||
"librados2",
|
||||
"python3-cephfs",
|
||||
"libcephfs2",
|
||||
"python3-rgw",
|
||||
"python3-rados",
|
||||
"python3-ceph-argparse",
|
||||
"python3-ceph-common",
|
||||
"python3-rbd",
|
||||
"ceph-fuse",
|
||||
"chrony",
|
||||
"libcorosync-common4",
|
||||
"libcfg7",
|
||||
"libcmap4",
|
||||
"libcpg4",
|
||||
"libknet1t64",
|
||||
"libnozzle1t64",
|
||||
"libquorum5",
|
||||
"libvotequorum8",
|
||||
"corosync",
|
||||
"frr-pythontools",
|
||||
"libjs-extjs",
|
||||
"libnvpair3linux",
|
||||
"libproxmox-acme-plugins",
|
||||
"libproxmox-backup-qemu0",
|
||||
"pve-qemu-kvm",
|
||||
"libpve-notify-perl",
|
||||
"libpve-cluster-api-perl",
|
||||
"libpve-cluster-perl",
|
||||
"pve-cluster",
|
||||
"libpve-access-control",
|
||||
"libpve-apiclient-perl",
|
||||
"librados2-perl",
|
||||
"proxmox-backup-client",
|
||||
"proxmox-backup-file-restore",
|
||||
"pve-manager",
|
||||
"libproxmox-acme-perl",
|
||||
"libpve-common-perl",
|
||||
"libpve-guest-common-perl",
|
||||
"qemu-server",
|
||||
"libpve-storage-perl",
|
||||
"pve-edk2-firmware-legacy",
|
||||
"pve-edk2-firmware-ovmf",
|
||||
"libpve-network-api-perl",
|
||||
"libpve-network-perl",
|
||||
"proxmox-firewall-data",
|
||||
"pve-firewall",
|
||||
"pve-container",
|
||||
"pve-ha-manager",
|
||||
"novnc-pve",
|
||||
"proxmox-enterprise-support-keyring",
|
||||
"proxmox-mini-journalreader",
|
||||
"proxmox-widget-toolkit",
|
||||
"pve-docs",
|
||||
"pve-i18n",
|
||||
"pve-xtermjs",
|
||||
"pve-yew-mobile-i18n",
|
||||
"pve-yew-mobile-gui",
|
||||
"libuutil3linux",
|
||||
"libzfs7linux",
|
||||
"libzpool7linux",
|
||||
"proxmox-kernel-helper",
|
||||
"pve-edk2-firmware-aarch64",
|
||||
"pve-edk2-firmware",
|
||||
"pve-firmware",
|
||||
"zfs-initramfs",
|
||||
"zfsutils-linux",
|
||||
"zfs-zed"
|
||||
],
|
||||
"outcome": "applied",
|
||||
"pending": 78,
|
||||
"reboot_needed": false,
|
||||
"refused": null,
|
||||
"release_id": "ring0-20261004T122522Z",
|
||||
"restart_needed": [],
|
||||
"ring": 0,
|
||||
"run_id": "20261004T122522Z",
|
||||
"upgraded": [
|
||||
{
|
||||
"name": "tzdata",
|
||||
"version": "2026c-0+deb13u1",
|
||||
"origin": ""
|
||||
}
|
||||
],
|
||||
"wrapper_seconds": 20.2
|
||||
}
|
||||
pass took 36.9s
|
||||
tzdata 2026c-0+deb13u1
|
||||
0
|
||||
@@ -0,0 +1,104 @@
|
||||
{
|
||||
"boxes": [
|
||||
{
|
||||
"HostID": "demo-felhom-8363b5",
|
||||
"Ring": 1,
|
||||
"Enabled": true,
|
||||
"Tunnel": "running",
|
||||
"Guest": {
|
||||
"ReleaseID": "os-guest-20261004-123933",
|
||||
"LastOutcome": "nothing",
|
||||
"LastAt": "2026-10-04T12:40:56Z",
|
||||
"LastSuccessfulLeg": "2026-10-04T12:40:56Z",
|
||||
"Pending": 6,
|
||||
"NotCovered": 6,
|
||||
"NotCoveredFast": 0,
|
||||
"RestartNeeded": 0,
|
||||
"RebootNeededSince": "2026-10-04T09:20:04Z",
|
||||
"WrapperPassSeconds": 10.1
|
||||
},
|
||||
"Host": {
|
||||
"ReleaseID": "os-host-20261004-124034",
|
||||
"LastOutcome": "applied",
|
||||
"LastAt": "2026-10-04T12:41:10Z",
|
||||
"LastSuccessfulLeg": "2026-10-04T12:41:10Z",
|
||||
"Pending": 80,
|
||||
"NotCovered": 80,
|
||||
"NotCoveredFast": 0,
|
||||
"RestartNeeded": 34,
|
||||
"RebootNeededSince": "2026-10-04T11:53:20Z",
|
||||
"WrapperPassSeconds": 13.1
|
||||
}
|
||||
},
|
||||
{
|
||||
"HostID": "demo-hp-bb76ea",
|
||||
"Ring": 0,
|
||||
"Enabled": true,
|
||||
"Tunnel": "unknown",
|
||||
"Guest": {
|
||||
"ReleaseID": "ring0-20261004T122522Z",
|
||||
"LastOutcome": "nothing",
|
||||
"LastAt": "2026-10-04T12:25:37Z",
|
||||
"LastSuccessfulLeg": "2026-10-04T12:25:37Z",
|
||||
"Pending": 6,
|
||||
"NotCovered": 6,
|
||||
"NotCoveredFast": 0,
|
||||
"RestartNeeded": 0,
|
||||
"RebootNeededSince": "2026-10-04T09:21:28Z",
|
||||
"WrapperPassSeconds": 15.4
|
||||
},
|
||||
"Host": {
|
||||
"ReleaseID": "ring0-20261004T122522Z",
|
||||
"LastOutcome": "applied",
|
||||
"LastAt": "2026-10-04T12:25:58Z",
|
||||
"LastSuccessfulLeg": "2026-10-04T12:25:58Z",
|
||||
"Pending": 78,
|
||||
"NotCovered": 78,
|
||||
"NotCoveredFast": 0,
|
||||
"RestartNeeded": 0,
|
||||
"RebootNeededSince": "0001-01-01T00:00:00Z",
|
||||
"WrapperPassSeconds": 20.2
|
||||
}
|
||||
},
|
||||
{
|
||||
"HostID": "drill-r50-0a4f9a",
|
||||
"Ring": 1,
|
||||
"Enabled": true,
|
||||
"Tunnel": "inactive",
|
||||
"Guest": {
|
||||
"ReleaseID": "",
|
||||
"LastOutcome": "",
|
||||
"LastAt": "0001-01-01T00:00:00Z",
|
||||
"LastSuccessfulLeg": "0001-01-01T00:00:00Z",
|
||||
"Pending": 0,
|
||||
"NotCovered": 0,
|
||||
"NotCoveredFast": 0,
|
||||
"RestartNeeded": 0,
|
||||
"RebootNeededSince": "0001-01-01T00:00:00Z",
|
||||
"WrapperPassSeconds": 0
|
||||
},
|
||||
"Host": {
|
||||
"ReleaseID": "",
|
||||
"LastOutcome": "",
|
||||
"LastAt": "0001-01-01T00:00:00Z",
|
||||
"LastSuccessfulLeg": "0001-01-01T00:00:00Z",
|
||||
"Pending": 0,
|
||||
"NotCovered": 0,
|
||||
"NotCoveredFast": 0,
|
||||
"RestartNeeded": 0,
|
||||
"RebootNeededSince": "0001-01-01T00:00:00Z",
|
||||
"WrapperPassSeconds": 0
|
||||
}
|
||||
}
|
||||
],
|
||||
"latest_guest_release": {
|
||||
"approved_at": "2026-10-04T12:39:33Z",
|
||||
"approved_by": "auto",
|
||||
"id": "os-guest-20261004-123933"
|
||||
},
|
||||
"latest_host_release": {
|
||||
"approved_at": "2026-10-04T12:40:34Z",
|
||||
"approved_by": "auto",
|
||||
"id": "os-host-20261004-124034"
|
||||
}
|
||||
}
|
||||
+172
@@ -0,0 +1,172 @@
|
||||
=== felhom-agent 0.141.1 selftest=os-update vmid=9201 ring=0 enabled=true guest-release=false host-release=false appliance=true ===
|
||||
time=2026-10-04T13:52:57.534+02:00 level=INFO msg="osupdate: START" run=20261004T115257Z layer=guest vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T115257Z
|
||||
time=2026-10-04T13:53:09.479+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T115257Z layer=guest:9201 lane=fast mode=apply select=pending-fast packages=0"
|
||||
time=2026-10-04T13:53:09.479+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
|
||||
time=2026-10-04T13:53:09.479+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0"
|
||||
time=2026-10-04T13:53:09.479+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)"
|
||||
time=2026-10-04T13:53:09.480+02:00 level=INFO msg="osupdate: DONE" run=20261004T115257Z layer=guest vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=6 not_covered=6 restart_needed=0 reboot_needed=false wrapper_seconds=11.9
|
||||
time=2026-10-04T13:53:09.491+02:00 level=INFO msg="osupdate: START" run=20261004T115257Z layer=host vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T115257Z
|
||||
time=2026-10-04T13:53:19.927+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T115257Z layer=host lane=fast mode=apply select=pending-fast packages=0"
|
||||
time=2026-10-04T13:53:19.927+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
|
||||
time=2026-10-04T13:53:19.927+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0"
|
||||
time=2026-10-04T13:53:19.927+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)"
|
||||
time=2026-10-04T13:53:20.705+02:00 level=INFO msg="osupdate: DONE" run=20261004T115257Z layer=host vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=80 not_covered=80 restart_needed=34 reboot_needed=true wrapper_seconds=10.4
|
||||
--- os-update report (guest) ---
|
||||
{
|
||||
"health_reason": "",
|
||||
"healthy": true,
|
||||
"mode": "apply",
|
||||
"not_covered": [
|
||||
"docker-ce-cli",
|
||||
"containerd.io",
|
||||
"docker-ce",
|
||||
"docker-buildx-plugin",
|
||||
"docker-ce-rootless-extras",
|
||||
"docker-compose-plugin"
|
||||
],
|
||||
"outcome": "nothing",
|
||||
"pending": 6,
|
||||
"reboot_needed": false,
|
||||
"refused": null,
|
||||
"release_id": "ring0-20261004T115257Z",
|
||||
"restart_needed": null,
|
||||
"ring": 0,
|
||||
"run_id": "20261004T115257Z",
|
||||
"upgraded": [],
|
||||
"wrapper_seconds": 11.9
|
||||
}
|
||||
--- os-update report (host) ---
|
||||
{
|
||||
"health_reason": "",
|
||||
"healthy": true,
|
||||
"mode": "apply",
|
||||
"not_covered": [
|
||||
"frr",
|
||||
"shim-signed-common",
|
||||
"shim-unsigned",
|
||||
"shim-helpers-amd64-signed",
|
||||
"shim-signed",
|
||||
"libradosstriper1",
|
||||
"librgw2",
|
||||
"ceph-common",
|
||||
"librbd1",
|
||||
"librados2",
|
||||
"python3-cephfs",
|
||||
"libcephfs2",
|
||||
"python3-rgw",
|
||||
"python3-rados",
|
||||
"python3-ceph-argparse",
|
||||
"python3-ceph-common",
|
||||
"python3-rbd",
|
||||
"ceph-fuse",
|
||||
"chrony",
|
||||
"libcorosync-common4",
|
||||
"libcfg7",
|
||||
"libcmap4",
|
||||
"libcpg4",
|
||||
"libknet1t64",
|
||||
"libnozzle1t64",
|
||||
"libquorum5",
|
||||
"libvotequorum8",
|
||||
"corosync",
|
||||
"frr-pythontools",
|
||||
"libjs-extjs",
|
||||
"libnvpair3linux",
|
||||
"libproxmox-acme-plugins",
|
||||
"libproxmox-backup-qemu0",
|
||||
"pve-qemu-kvm",
|
||||
"libpve-notify-perl",
|
||||
"libpve-cluster-api-perl",
|
||||
"libpve-cluster-perl",
|
||||
"pve-cluster",
|
||||
"libpve-access-control",
|
||||
"libpve-apiclient-perl",
|
||||
"librados2-perl",
|
||||
"proxmox-backup-client",
|
||||
"proxmox-backup-file-restore",
|
||||
"pve-manager",
|
||||
"libproxmox-acme-perl",
|
||||
"libpve-common-perl",
|
||||
"libpve-guest-common-perl",
|
||||
"qemu-server",
|
||||
"libpve-storage-perl",
|
||||
"pve-edk2-firmware-legacy",
|
||||
"pve-edk2-firmware-ovmf",
|
||||
"libpve-network-api-perl",
|
||||
"libpve-network-perl",
|
||||
"proxmox-firewall-data",
|
||||
"pve-firewall",
|
||||
"pve-container",
|
||||
"pve-ha-manager",
|
||||
"novnc-pve",
|
||||
"proxmox-enterprise-support-keyring",
|
||||
"proxmox-mini-journalreader",
|
||||
"proxmox-widget-toolkit",
|
||||
"pve-docs",
|
||||
"pve-i18n",
|
||||
"pve-xtermjs",
|
||||
"pve-yew-mobile-i18n",
|
||||
"pve-yew-mobile-gui",
|
||||
"libuutil3linux",
|
||||
"libzfs7linux",
|
||||
"libzpool7linux",
|
||||
"proxmox-first-boot",
|
||||
"pve-firmware",
|
||||
"proxmox-kernel-7.0.14-20-pve-signed",
|
||||
"proxmox-kernel-7.0",
|
||||
"proxmox-kernel-helper",
|
||||
"pve-edk2-firmware-aarch64",
|
||||
"pve-edk2-firmware",
|
||||
"zfs-initramfs",
|
||||
"zfsutils-linux",
|
||||
"zfs-zed",
|
||||
"tailscale"
|
||||
],
|
||||
"outcome": "nothing",
|
||||
"pending": 80,
|
||||
"reboot_needed": true,
|
||||
"refused": null,
|
||||
"release_id": "ring0-20261004T115257Z",
|
||||
"restart_needed": [
|
||||
"agetty",
|
||||
"blkmapd",
|
||||
"chronyd",
|
||||
"cron",
|
||||
"dbus-daemon",
|
||||
"dmeventd",
|
||||
"ksmtuned",
|
||||
"lxc-monitord",
|
||||
"lxc-start",
|
||||
"lxcfs",
|
||||
"pmxcfs",
|
||||
"proxmox-firewal",
|
||||
"pve-firewall",
|
||||
"pve-ha-crm",
|
||||
"pve-ha-lrm",
|
||||
"pve-lxc-syscall",
|
||||
"pvedaemon",
|
||||
"pvedaemon worke",
|
||||
"pvefw-logger",
|
||||
"pveproxy",
|
||||
"pveproxy worker",
|
||||
"pvescheduler",
|
||||
"pvestatd",
|
||||
"qmeventd",
|
||||
"rpcbind",
|
||||
"rrdcached",
|
||||
"smartd",
|
||||
"spiceproxy",
|
||||
"spiceproxy work",
|
||||
"sshd",
|
||||
"systemd-logind",
|
||||
"systemd-udevd",
|
||||
"watchdog-mux",
|
||||
"zed"
|
||||
],
|
||||
"ring": 0,
|
||||
"run_id": "20261004T115257Z",
|
||||
"upgraded": [],
|
||||
"wrapper_seconds": 10.4
|
||||
}
|
||||
pass took 23.2s
|
||||
WALL_SECONDS=23.268350584
|
||||
@@ -0,0 +1,125 @@
|
||||
+ uname -r
|
||||
7.0.2-6-pve
|
||||
+ mokutil --sb-state
|
||||
SecureBoot enabled
|
||||
+ od -An -tx1 /sys/firmware/efi/efivars/SecureBoot-8be4df61-93ca-11d2-aa0d-00e098032b8c
|
||||
+ head -1
|
||||
06 00 00 00 01
|
||||
+ ls /boot/vmlinuz-7.0.14-20-pve /boot/vmlinuz-7.0.2-6-pve
|
||||
/boot/vmlinuz-7.0.14-20-pve
|
||||
/boot/vmlinuz-7.0.2-6-pve
|
||||
+ dpkg -l 'proxmox-kernel-*'
|
||||
+ grep '^ii'
|
||||
+ awk '{print $2,$3}'
|
||||
proxmox-kernel-7.0 7.0.14-20
|
||||
proxmox-kernel-7.0.14-20-pve-signed 7.0.14-20
|
||||
proxmox-kernel-7.0.2-6-pve-signed 7.0.2-6
|
||||
proxmox-kernel-helper 9.1.0+fde2
|
||||
+ proxmox-boot-tool status
|
||||
+ tail -5
|
||||
Re-executing '/usr/sbin/proxmox-boot-tool' in new private mount namespace..
|
||||
E: /etc/kernel/proxmox-boot-uuids does not exist.
|
||||
+ grep -E '^GRUB_DEFAULT|^GRUB_TIMEOUT|^GRUB_SAVEDEFAULT' /etc/default/grub
|
||||
GRUB_DEFAULT=0
|
||||
GRUB_TIMEOUT=5
|
||||
+ grub-editenv list
|
||||
+ '[' -d /sys/firmware/efi ']'
|
||||
+ efibootmgr
|
||||
+ head -6
|
||||
BootCurrent: 0003
|
||||
Timeout: 0 seconds
|
||||
BootOrder: 0003,0029,0001,0002,0006,0007,0019,001A,001B,001C,001D,001E,0008,0009,0016,0017,001F,0020,0021,0022,0023,0024,0025,0026,0027,0028,000A,0004,0005,000B
|
||||
Boot0001* USB Floppy/CD VenMedia(b6fef66f-1495-4584-a836-3492d1984a8d,0500000001)0000424f
|
||||
Boot0002* USB Hard Drive VenMedia(b6fef66f-1495-4584-a836-3492d1984a8d,0200000001)0000424f
|
||||
Boot0003* proxmox HD(2,GPT,175383fb-546d-430e-9b4c-73ec2379169d,0x800,0x200000)/File(\EFI\proxmox\shimx64.efi)
|
||||
+ sysctl kernel.panic kernel.panic_on_oops
|
||||
kernel.panic = 0
|
||||
kernel.panic_on_oops = 0
|
||||
+ lsmod
|
||||
+ grep -i -E 'sp5100|wdt|watchdog'
|
||||
+ ls -l /dev/watchdog /dev/watchdog0
|
||||
crw------- 1 root root 10, 130 Oct 4 09:46 /dev/watchdog
|
||||
crw------- 1 root root 243, 0 Oct 4 09:46 /dev/watchdog0
|
||||
+ wdctl
|
||||
+ head -12
|
||||
Device: /dev/watchdog0
|
||||
Identity: Software Watchdog [version 0]
|
||||
Timeout: 10 seconds
|
||||
Timeleft: 9 seconds
|
||||
Pre-timeout: 0 seconds
|
||||
Pre-timeout governor: noop
|
||||
Available pre-timeout governors: noop
|
||||
FLAG DESCRIPTION STATUS BOOT-STATUS
|
||||
KEEPALIVEPING Keep alive ping reply 1 0
|
||||
MAGICCLOSE Supports magic close char 0 0
|
||||
PRETIMEOUT Pretimeout (in seconds) 0 0
|
||||
SETTIMEOUT Set timeout (in seconds) 0 0
|
||||
+ dmesg
|
||||
+ grep -i -E 'sp5100|watchdog'
|
||||
+ tail -5
|
||||
[ 0.353227] NMI watchdog: Enabled. Permanently consumes one hw-PMU counter.
|
||||
+ apt-cache policy proxmox-default-kernel
|
||||
+ head -4
|
||||
proxmox-default-kernel:
|
||||
Installed: 2.1.0
|
||||
Candidate: 2.1.0
|
||||
Version table:
|
||||
+ apt-cache search --names-only '^proxmox-kernel-[0-9.]+-[0-9]+-pve-signed$'
|
||||
+ sort -V
|
||||
+ tail -3
|
||||
proxmox-kernel-7.0.14-18-pve-signed - Proxmox Kernel Image (signed)
|
||||
proxmox-kernel-7.0.14-19-pve-signed - Proxmox Kernel Image (signed)
|
||||
proxmox-kernel-7.0.14-20-pve-signed - Proxmox Kernel Image (signed)
|
||||
+ cat /proc/cmdline
|
||||
BOOT_IMAGE=/boot/vmlinuz-7.0.2-6-pve root=/dev/mapper/pve-root ro quiet
|
||||
+ grep -nE '^menuentry|^submenu|^\s+menuentry' /boot/grub/grub.cfg
|
||||
+ cut -c1-140
|
||||
+ head -8
|
||||
26: menuentry_id_option="--id"
|
||||
28: menuentry_id_option=""
|
||||
110:menuentry 'Proxmox VE GNU/Linux' --class proxmox --class gnu-linux --class gnu --class os $menuentry_id_option 'gnulinux-simple-529c0c3d
|
||||
128:submenu 'Advanced options for Proxmox VE GNU/Linux' $menuentry_id_option 'gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43' {
|
||||
129: menuentry 'Proxmox VE GNU/Linux, with Linux 7.0.14-20-pve' --class proxmox --class gnu-linux --class gnu --class os $menuentry_id_optio
|
||||
147: menuentry 'Proxmox VE GNU/Linux, with Linux 7.0.14-20-pve (recovery mode)' --class proxmox --class gnu-linux --class gnu --class os $me
|
||||
165: menuentry 'Proxmox VE GNU/Linux, with Linux 7.0.2-6-pve' --class proxmox --class gnu-linux --class gnu --class os $menuentry_id_option
|
||||
183: menuentry 'Proxmox VE GNU/Linux, with Linux 7.0.2-6-pve (recovery mode)' --class proxmox --class gnu-linux --class gnu --class os $menu
|
||||
+ ls -l --time-style=+%F_%T /boot/grub/grub.cfg
|
||||
-rw------- 1 root root 13743 2026-10-04_09:47:05 /boot/grub/grub.cfg
|
||||
+ uptime -s
|
||||
2026-10-04 09:46:15
|
||||
+ last -x reboot
|
||||
+ head -4
|
||||
reboot system boot 7.0.2-6-pve Sun Oct 4 09:46 - still running
|
||||
reboot system boot 7.0.14-20-pve Sun Oct 4 09:44 - 09:45 (00:00)
|
||||
shutdown system down 7.0.14-20-pve Sun Oct 4 09:45 - 09:46 (00:00)
|
||||
reboot system boot 7.0.2-6-pve Fri Aug 21 17:44 - 09:43 (43+15:58)
|
||||
+ ls /boot/efi/EFI/proxmox/
|
||||
BOOTX64.CSV
|
||||
fbx64.efi
|
||||
grub.cfg
|
||||
grubx64.efi
|
||||
mmx64.efi
|
||||
shimx64.efi
|
||||
+ cat /boot/efi/EFI/proxmox/grub.cfg
|
||||
+ head -5
|
||||
search.fs_uuid 529c0c3d-b48e-4d01-989d-43fd5d7dbb43 root lvmid/zVqGDa-V6js-HB26-RyZr-d0ft-fUB3-blL75R/iAs9TN-W8rL-ViG8-jptz-1tcY-3djE-MRqamk
|
||||
set prefix=($root)'/boot/grub'
|
||||
configfile $prefix/grub.cfg
|
||||
+ lspci -nn
|
||||
+ grep -i -E 'smbus|fch'
|
||||
00:14.0 SMBus [0c05]: Advanced Micro Devices, Inc. [AMD] FCH SMBus Controller [1022:790b] (rev 61)
|
||||
00:14.3 ISA bridge [0601]: Advanced Micro Devices, Inc. [AMD] FCH LPC Bridge [1022:790e] (rev 51)
|
||||
06:00.0 SATA controller [0106]: Advanced Micro Devices, Inc. [AMD] FCH SATA Controller [AHCI mode] [1022:7901] (rev 61)
|
||||
+ modinfo -F filename sp5100_tco
|
||||
/lib/modules/7.0.2-6-pve/kernel/drivers/watchdog/sp5100_tco.ko
|
||||
+ grep -rl sp5100 /lib/modprobe.d /etc/modprobe.d
|
||||
/lib/modprobe.d/blacklist_proxmox-kernel-7.0.14-20-pve.conf
|
||||
/lib/modprobe.d/blacklist_proxmox-kernel-7.0.2-6-pve.conf
|
||||
+ grep -h sp5100 /lib/modprobe.d/aliases.conf /lib/modprobe.d/blacklist_proxmox-kernel-7.0.14-20-pve.conf /lib/modprobe.d/blacklist_proxmox-kernel-7.0.2-6-pve.conf /lib/modprobe.d/fbdev-blacklist.conf /lib/modprobe.d/proxmox_prevent_autoload_proxmox-kernel-7.0.14-20-pve.conf /lib/modprobe.d/proxmox_prevent_autoload_proxmox-kernel-7.0.2-6-pve.conf /lib/modprobe.d/systemd.conf /etc/modprobe.d/amd64-microcode-blacklist.conf /etc/modprobe.d/pve-blacklist.conf /etc/modprobe.d/zfs.conf
|
||||
+ head -3
|
||||
blacklist sp5100_tco
|
||||
blacklist sp5100_tco
|
||||
+ modprobe -n -v sp5100_tco
|
||||
insmod /lib/modules/7.0.2-6-pve/kernel/drivers/watchdog/sp5100_tco.ko
|
||||
+ ls /sys/class/watchdog/
|
||||
watchdog0
|
||||
@@ -0,0 +1,20 @@
|
||||
+ findmnt -no SOURCE,FSTYPE --target /boot
|
||||
/dev/mapper/pve-root ext4
|
||||
+ ls -l /boot/grub/grubenv
|
||||
-rw-r--r-- 1 root root 1024 Oct 4 09:46 /boot/grub/grubenv
|
||||
+ cp -p /etc/default/grub /root/grub.default.bak-2026-10-04
|
||||
+ sed -i 's/^GRUB_DEFAULT=0$/GRUB_DEFAULT=saved/' /etc/default/grub
|
||||
+ grep '^GRUB_DEFAULT' /etc/default/grub
|
||||
GRUB_DEFAULT=saved
|
||||
+ update-grub
|
||||
+ tail -3
|
||||
Found memtest86+ 32bit image: /boot/memtest86+ia32.bin
|
||||
Adding boot menu entry for UEFI Firmware Settings ...
|
||||
done
|
||||
+ grub-set-default 'gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.2-6-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43'
|
||||
+ grub-editenv list
|
||||
saved_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.2-6-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43
|
||||
+ grep -n 'set default' /boot/grub/grub.cfg
|
||||
+ head -4
|
||||
17: set default="${next_entry}"
|
||||
22: set default="${saved_entry}"
|
||||
@@ -0,0 +1,14 @@
|
||||
+ grub-reboot 'gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.14-20-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43'
|
||||
|
||||
WARNING: Detected GRUB environment block on lvm device
|
||||
gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.14-20-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43 will remain the default boot entry until manually cleared with:
|
||||
grub-editenv /boot/grub/grubenv unset next_entry
|
||||
|
||||
+ grub-editenv list
|
||||
saved_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.2-6-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43
|
||||
next_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.14-20-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43
|
||||
+ uname -r
|
||||
7.0.2-6-pve
|
||||
+ date -u +%FT%TZ
|
||||
2026-10-04T12:26:10Z
|
||||
reboot1 issued 2026-10-04T12:26:10Z
|
||||
@@ -0,0 +1,19 @@
|
||||
+ uname -r
|
||||
7.0.14-20-pve
|
||||
+ uptime -s
|
||||
2026-10-04 14:27:03
|
||||
+ mokutil --sb-state
|
||||
SecureBoot enabled
|
||||
+ grub-editenv list
|
||||
saved_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.2-6-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43
|
||||
next_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.14-20-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43
|
||||
+ cat /proc/cmdline
|
||||
BOOT_IMAGE=/boot/vmlinuz-7.0.14-20-pve root=/dev/mapper/pve-root ro quiet
|
||||
active
|
||||
active
|
||||
active
|
||||
active
|
||||
active
|
||||
status: running
|
||||
starting
|
||||
starting
|
||||
@@ -0,0 +1,33 @@
|
||||
reboot2 issued 2026-10-04T12:35:19Z, no grub command; grubenv as after reboot 1
|
||||
ssh back after ~40s
|
||||
System is going down. Unprivileged users are not permitted to log in anymore. For technical details, see pam_nologin(8).
|
||||
|
||||
+ uname -r
|
||||
7.0.14-20-pve
|
||||
+ uptime -s
|
||||
2026-10-04 14:27:03
|
||||
+ grub-editenv list
|
||||
saved_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.2-6-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43
|
||||
next_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.14-20-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43
|
||||
+ cat /proc/cmdline
|
||||
BOOT_IMAGE=/boot/vmlinuz-7.0.14-20-pve root=/dev/mapper/pve-root ro quiet
|
||||
--- after the real restart, 2026-10-04T12:36:27Z
|
||||
+ uname -r
|
||||
7.0.14-20-pve
|
||||
+ uptime -s
|
||||
2026-10-04 14:36:11
|
||||
+ mokutil --sb-state
|
||||
SecureBoot enabled
|
||||
+ grub-editenv list
|
||||
saved_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.2-6-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43
|
||||
next_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.14-20-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43
|
||||
+ cat /proc/cmdline
|
||||
BOOT_IMAGE=/boot/vmlinuz-7.0.14-20-pve root=/dev/mapper/pve-root ro quiet
|
||||
+ systemctl is-active pveproxy pvedaemon pvestatd pve-cluster felhom-agent
|
||||
activating
|
||||
active
|
||||
active
|
||||
active
|
||||
inactive
|
||||
+ pct status 9201
|
||||
status: stopped
|
||||
@@ -0,0 +1,27 @@
|
||||
+ grub-editenv /boot/grub/grubenv unset next_entry
|
||||
+ grub-set-default 'gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.14-20-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43'
|
||||
+ grub-editenv list
|
||||
saved_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.14-20-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43
|
||||
+ grep '^GRUB_DEFAULT' /etc/default/grub
|
||||
GRUB_DEFAULT=saved
|
||||
+ dpkg -l 'proxmox-kernel-*-pve-signed'
|
||||
+ awk '{print $2}'
|
||||
+ grep '^ii'
|
||||
proxmox-kernel-7.0.14-20-pve-signed
|
||||
proxmox-kernel-7.0.2-6-pve-signed
|
||||
+ uname -r
|
||||
7.0.14-20-pve
|
||||
+ systemctl is-active pveproxy pvedaemon pvestatd pve-cluster felhom-agent
|
||||
active
|
||||
active
|
||||
active
|
||||
active
|
||||
active
|
||||
+ pct status 9201
|
||||
status: running
|
||||
+ sleep 45
|
||||
+ pct exec 9201 -- docker inspect -f '{{.Name}} {{.State.Health.Status}}' cloudflared felhom-controller
|
||||
/cloudflared healthy
|
||||
/felhom-controller healthy
|
||||
+ sysctl kernel.panic
|
||||
kernel.panic = 0
|
||||
@@ -0,0 +1,74 @@
|
||||
+ modprobe sp5100_tco
|
||||
rc=0
|
||||
+ echo rc=0
|
||||
+ lsmod
|
||||
+ grep sp5100
|
||||
sp5100_tco 20480 0
|
||||
+ dmesg
|
||||
+ grep -i sp5100
|
||||
+ tail -5
|
||||
[ 85.993589] sp5100_tco: SP5100/SB800 TCO WatchDog Timer Driver
|
||||
[ 85.993780] sp5100-tco sp5100-tco: Using 0xfeb00000 for watchdog MMIO address
|
||||
[ 85.993918] sp5100-tco sp5100-tco: initialized. heartbeat=60 sec (nowayout=0)
|
||||
== /sys/class/watchdog/watchdog0
|
||||
+ for w in /sys/class/watchdog/watchdog*
|
||||
+ echo '== /sys/class/watchdog/watchdog0'
|
||||
+ for f in identity state timeout nowayout bootstatus status
|
||||
++ cat /sys/class/watchdog/watchdog0/identity
|
||||
identity=Software Watchdog
|
||||
+ printf '%s=%s\n' identity 'Software Watchdog'
|
||||
+ for f in identity state timeout nowayout bootstatus status
|
||||
++ cat /sys/class/watchdog/watchdog0/state
|
||||
state=active
|
||||
+ printf '%s=%s\n' state active
|
||||
+ for f in identity state timeout nowayout bootstatus status
|
||||
++ cat /sys/class/watchdog/watchdog0/timeout
|
||||
timeout=10
|
||||
+ printf '%s=%s\n' timeout 10
|
||||
+ for f in identity state timeout nowayout bootstatus status
|
||||
++ cat /sys/class/watchdog/watchdog0/nowayout
|
||||
nowayout=0
|
||||
+ printf '%s=%s\n' nowayout 0
|
||||
+ for f in identity state timeout nowayout bootstatus status
|
||||
++ cat /sys/class/watchdog/watchdog0/bootstatus
|
||||
bootstatus=0
|
||||
+ printf '%s=%s\n' bootstatus 0
|
||||
+ for f in identity state timeout nowayout bootstatus status
|
||||
++ cat /sys/class/watchdog/watchdog0/status
|
||||
status=0x8000
|
||||
== /sys/class/watchdog/watchdog1
|
||||
+ printf '%s=%s\n' status 0x8000
|
||||
+ for w in /sys/class/watchdog/watchdog*
|
||||
+ echo '== /sys/class/watchdog/watchdog1'
|
||||
+ for f in identity state timeout nowayout bootstatus status
|
||||
++ cat /sys/class/watchdog/watchdog1/identity
|
||||
identity=SP5100 TCO timer
|
||||
+ printf '%s=%s\n' identity 'SP5100 TCO timer'
|
||||
+ for f in identity state timeout nowayout bootstatus status
|
||||
++ cat /sys/class/watchdog/watchdog1/state
|
||||
state=inactive
|
||||
+ printf '%s=%s\n' state inactive
|
||||
+ for f in identity state timeout nowayout bootstatus status
|
||||
++ cat /sys/class/watchdog/watchdog1/timeout
|
||||
timeout=60
|
||||
+ printf '%s=%s\n' timeout 60
|
||||
+ for f in identity state timeout nowayout bootstatus status
|
||||
++ cat /sys/class/watchdog/watchdog1/nowayout
|
||||
nowayout=0
|
||||
+ printf '%s=%s\n' nowayout 0
|
||||
+ for f in identity state timeout nowayout bootstatus status
|
||||
++ cat /sys/class/watchdog/watchdog1/bootstatus
|
||||
bootstatus=0
|
||||
+ printf '%s=%s\n' bootstatus 0
|
||||
+ for f in identity state timeout nowayout bootstatus status
|
||||
++ cat /sys/class/watchdog/watchdog1/status
|
||||
status=0x0
|
||||
+ printf '%s=%s\n' status 0x0
|
||||
+ rmmod sp5100_tco
|
||||
rmmod_rc=0
|
||||
+ echo rmmod_rc=0
|
||||
+ lsmod
|
||||
+ grep -c sp5100
|
||||
0
|
||||
+ ls /sys/class/watchdog/
|
||||
watchdog0
|
||||
@@ -1,2 +1,4 @@
|
||||
demo-felhom wrapper 4729769ce32e25e6 755 root; sudoers 02df92d751f1780a
|
||||
demo-hp wrapper 4729769ce32e25e6 755 root; sudoers 02df92d751f1780a
|
||||
demo-felhom wrapper 51e100ad81945f67
|
||||
demo-hp wrapper 51e100ad81945f67
|
||||
|
||||
@@ -26,6 +26,17 @@
|
||||
|
||||
---
|
||||
|
||||
## 2026-10-04 (afternoon) — OS updates, host fast lane + fleet view + alarms (agent v0.141.0/v0.141.1, hub v0.131.0/v0.131.1, controller v0.292.0)
|
||||
|
||||
> Evidence: `audits/os-host-lane-2026-10-04/`.
|
||||
|
||||
| Row | What | Closed | Evidence |
|
||||
|---|---|---|---|
|
||||
| **R-841** | **Every box reported its tunnel `inactive`** (the agent asked a host unit that does not exist). Agent v0.141.0 reads the guest's `cloudflared` container and controller v0.292.0's Docker health check on cloudflared's own `/ready` (200 only with a connection); three states `running` / `not_running` / `unknown`; hub v0.131.0 alarms `tunnel_down` after two `not_running` reports and `tunnel_recovered` on the next `running`. LIVE on demo-hp: port 7844 blocked → `tunnel_down` mailed after the 2nd report; unblocked → `tunnel_recovered`. **Reasoning kept: `unknown` never alarms; a container state alone says "up" for a dead tunnel (wrong token: running, `/ready` 503). A plain `docker stop` is healed by the controller's protected-container check within 5 min — before the 15-min host report sees it.** | CLOSED 2026-10-04 — FIXED | `partA/` |
|
||||
| **R-845** | **The OS leg was slow.** One `pct exec` per package (~0.9 s each) replaced by one call per layer; restart scan only after an install (host: every pass, v0.141.1); repair only when `dpkg --audit` reports. MEASURED: nothing to install, both layers — 23.3 s (demo-felhom), 31.5 s (demo-hp); before, guest only, nothing to install — 14.0 s; the 174–245 s passes of R-845 were passes WITH an install. A 108-package host install pass: 70 s. | CLOSED 2026-10-04 — FIXED | `partD/`, `partB/live/` |
|
||||
| **R-846** | **Host "reboot needed" was wrong in agent v0.141.0** (found live on demo-felhom): the restart scan skipped every cgroup containing `lxc`, hiding `lxc-start` (`0::/lxc.monitor/<vmid>`, 20 deleted maps after libc6); and it scanned only after an install, so a reboot never cleared it (the hub's 14-day alarm would fire on a rebooted host). Agent v0.141.1 (`:/lxc/`, host scans every pass, `reboot_scanned`) + hub v0.131.1. Red-proved. | CLOSED 2026-10-04 — FIXED | `partB/live-defects-redproofs.txt` |
|
||||
| **R-850** | **Hub v0.131.0 put the layer into the release fingerprint**, so the unchanged guest set (272 packages) counted as new and waited its 24 h again. One-time; ring 1 kept the previous release meanwhile; approved again under the TEST wait (`os-guest-20261004-123933`). Nothing to fix. | CLOSED 2026-10-04 — ONE-TIME, NO ACTION | hub log 2026-10-04 |
|
||||
|
||||
## 2026-10-04 (~12:20) — operator ruling
|
||||
|
||||
| Row | What | Closed | Evidence |
|
||||
|
||||
@@ -308,7 +308,7 @@ stopping line that lies.
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| **R-530** | Box system & updates | P2 | **[P2-MEDIUM] A floor does not deliver an agent: agents update only by an operator-signed `agent_update` job per box, and nothing records which boxes still run 0.130.0.** MEASURED 2026-09-15: the hub HOLDS a floor whose declared MinAgent is above the box's agent (`api/handler.go` ResolveManagedFloor); the agent's only update path is `signedjobs` + `selfupdate.Executor`. demo-hp reached 0.131.0 by `felhom-opsign -op agent_update` (key `felhom-op-1`) at 08:44:16Z and its controller floor was then SERVED in 3 s. **demo-felhom (N100) and Peti's box still run 0.130.0** — not touched (Peti fenced; N100 not asked). **What it needs:** the operator signs per box, or rules a fleet rollout step. **NARROWED 2026-09-16 (operator ruling 1):** the keys stay on DooPlex owner-only and CC may sign `agent_update` until the first PAYING customer (testers excluded) — recorded in `CONTEXT.md` + `04-control-plane-authorization.md` §3.1. Both demo boxes now run agent 0.131.0 (demo-hp 2026-09-15, demo-felhom 2026-09-16, each by a per-box signed job; Peti's box untouched, still 0.130.0). **What remains:** a fleet rollout step — signing per box does not scale past a handful, and nothing lists which boxes are behind. **2026-09-25:** Peti's box was RETIRED (operator ruling) — it no longer counts as a box behind. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (signing)** | — | — | operator |
|
||||
| **R-604** | Box system & updates | P2 | **[P2-MEDIUM] A per-customer controller floor silently excludes that box from every global floor raise, and NOTHING says so — demo-hp missed four of them.** FOUND 2026-09-21 while raising the global floor to 0.259.0 at the operator's request. The raise logged `Global controller-version floor set to "0.259.0"` and then `managed floor SERVED for demo-felhom` — **and nothing at all for demo-hp**, which went on reporting every few minutes and stayed on 0.258.0. Cause: `customer_configs.min_controller_version` for demo-hp held **`0.243.0`**, a per-customer override that wins over the global. It is a **leftover from the 2026-09-16 drill**, whose golden was 0.243.0; **R-343 measured on 2026-08-18 that all five rows were EMPTY and recorded that as a safety property** — it stopped being true and nothing surfaced the change. demo-hp had therefore silently missed the raises to 0.253.0, 0.254.0, 0.257.0 and 0.259.0. **Why it is invisible rather than merely quiet:** `managed floor SERVED` fires **once per CHANGE** (`h.floorNotes`, `api/handler.go:600`), deliberately, because a box reports every few minutes — so a box whose override never changes is silent for ever, and its silence is indistinguishable from the silence of a box that already had the line. A session that raises the floor reads one SERVED line and reasonably concludes the fleet took it. **CLEARED** for demo-hp the same session (rollback line: POST `/customers/demo-hp/floor` with `min_controller_version=0.243.0`, `min_agent=0.131.0`); it then self-updated 0.258.0 → 0.259.0 in **under four minutes**, healthy, `settle-gate: GO — at/above floor 0.259.0`, and its claim page answers **"Wrong or expired code"** in English — the floor delivered the FIX, not a version string, to a box nobody hand-deployed. All five overrides are now empty. **Fix shape (smallest first):** the floor-raise page shows which customers carry an override and would NOT be moved, before the save; or the raise logs one line per customer naming the ones it skipped and why. A raise that quietly reaches half the fleet is worse than one that refuses. | **READY — rank P2-MEDIUM; owner: CC (hub)** | — | — | CC |
|
||||
| **R-812** | Box system & updates | P2 | **[P2] A box never receives operating-system security updates — not the Proxmox host, not the guest's Debian, not its Docker engine.** SEARCHED 2026-10-03 (read-only): `felhom-controller`, `felhom-agent`, `app-catalog-felhom.eu` and `felhom.eu` hold no `apt-get upgrade`, `apt full-upgrade`, `unattended-upgrades`, `pveupgrade` or `needrestart` that runs on a box. The installer aligns the host's Proxmox repositories to no-subscription *"so the box can pull security updates"* and then says plainly *"No upgrades are run"* (`scripts/felhom-host-install.sh:2133-2136`). The guest's Docker engine is installed when the golden is BAKED (`felhom-agent/configs/build-golden.sh:124-125`), so a fresh install gets that week's engine and an installed box keeps it forever. The only `apt full-upgrade` in the project is a by-hand step for the off-site endpoint ep0 (`documentation/runbooks/offsite-endpoint.md:41`), not a box. App images ARE updated (the update arc); the layer under them is not. The intention, with its scope, is **R-808** in `ROADMAP.md`. | **NARROWED AGAIN 2026-10-04 (day) — the GUEST's Debian fast lane is BUILT and proven live (agent v0.140.0, hub v0.130.0, installer 1.29.0; `11` §8.1, `audits/os-guest-lane-2026-10-04/`). LEFT: the host, Docker and kernel lanes; the undo (R-842); existing boxes (R-840).** Earlier: NARROWED 2026-10-04 — the SPIKE is done (`11` §7.1, corrections C1–C12, `audits/os-updates-spike-2026-10-04/`); no product code yet. LEFT: the build steps of `11` §8, each with the operator's go; the §5.3 snapshot question is in STATUS; preconditions R-835, R-836, R-837.** | — | — | CC + operator |
|
||||
| **R-812** | Box system & updates | P2 | **[P2] A box never receives operating-system security updates — not the Proxmox host, not the guest's Debian, not its Docker engine.** SEARCHED 2026-10-03 (read-only): `felhom-controller`, `felhom-agent`, `app-catalog-felhom.eu` and `felhom.eu` hold no `apt-get upgrade`, `apt full-upgrade`, `unattended-upgrades`, `pveupgrade` or `needrestart` that runs on a box. The installer aligns the host's Proxmox repositories to no-subscription *"so the box can pull security updates"* and then says plainly *"No upgrades are run"* (`scripts/felhom-host-install.sh:2133-2136`). The guest's Docker engine is installed when the golden is BAKED (`felhom-agent/configs/build-golden.sh:124-125`), so a fresh install gets that week's engine and an installed box keeps it forever. The only `apt full-upgrade` in the project is a by-hand step for the off-site endpoint ep0 (`documentation/runbooks/offsite-endpoint.md:41`), not a box. App images ARE updated (the update arc); the layer under them is not. The intention, with its scope, is **R-808** in `ROADMAP.md`. | **NARROWED 2026-10-04 (afternoon) — the HOST's Debian fast lane is BUILT and proven live too (agent v0.141.1, hub v0.131.1; `11` §8.2, `audits/os-host-lane-2026-10-04/`): appliances only, never kernel/boot/firmware, after a healthy guest step; fleet view and four alarms (§8.3). LEFT: the Docker and kernel slow lanes (R-836); existing boxes (R-840).** Earlier: NARROWED AGAIN 2026-10-04 (day) — the GUEST's Debian fast lane is BUILT and proven live (agent v0.140.0, hub v0.130.0, installer 1.29.0; `11` §8.1, `audits/os-guest-lane-2026-10-04/`). LEFT: the host, Docker and kernel lanes; the undo (R-842); existing boxes (R-840).** Earlier: NARROWED 2026-10-04 — the SPIKE is done (`11` §7.1, corrections C1–C12, `audits/os-updates-spike-2026-10-04/`); no product code yet. LEFT: the build steps of `11` §8, each with the operator's go; the §5.3 snapshot question is in STATUS; preconditions R-835, R-836, R-837.** | — | — | CC + operator |
|
||||
| **R-35** | Box system & updates | P3 | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` | CC |
|
||||
| **R-50b** | Box system & updates | P3 | **[P2] A root-owned privileged host artifact is delivered unversioned from `main` — "which wrapper is on this host?" is unanswerable.** `configs/felhom-pbs-apply` installs to `/usr/local/sbin/felhom-pbs-apply` (0755 root:root) and is the pinned sudoers vector for `create\ |reconcile\|grant` against `/etc/pve/priv/storage`. It is fetched by `felhom-host-install.sh:1914` via `fetch_raw`, which hits `raw/branch/main/<path>` — **no tag, no pin, no checksum, and no record in the Day-0 artifact manifest**, unlike the agent binary (sha256-vouched) and the golden image. Three consequences: (1) two hosts installed a week apart can carry different privileged wrapper code while both reporting the same agent version; (2) a host hotfixed in place (felhom-pve, 2026-07-18) is indistinguishable from one that fetched the same content — the fleet has no inventory of it; (3) an accidental push to `main` reaches the next install of every host with no review gate between commit and root-owned deployment. | **NARROWED** — **(a) SHIPPED 2026-07-21; (b)/(c) open** — moved from `ROADMAP.md` 2026-10-03: it states a checkable fact about the shipped product, so it is a FINDING (the sorting rule). **Re-ranked 2026-10-03: [P2] → P3 — operator-only; leg (a) shipped.** | — | **Surfaced 2026-07-21 while stopping the R-39 v0.90.1 publish** (`felhom-controller/REPORT.md` §5): the publish was cancelled precisely because the version number would have claimed to carry a fix that in fact rides this unversioned channel. Candidate shapes, in increasing cost: (a) record the wrapper's sha256 in the Day-0 artifact manifest beside the agent binary and have the agent report the installed file's hash, so drift is at least *visible*; (b) `fetch_raw` takes a pinned ref (tag or commit) supplied by the manifest rather than `main`; (c) the wrapper becomes a published generic-registry artifact with the same gate ladder as the agent binary. **(a) is the cheap honest first step and would have caught this class already.** Pairs with R-39 (whose remaining fleet half is specced separately) **(a) SHIPPED 2026-07-21 — hub v0.68.0 + agent v0.91.2.** `ArtifactManifest.WrapperSHA256` + an operator field; agents report the installed wrapper's sha256 each cycle and the host page surfaces a mismatch. **An unknown on EITHER side reads as quiet, never as drift** — lighting every host amber on rollout day is how a warning becomes background noise. Live confirmation of exactly the problem: felhom-pve's July-18 in-place hotfix hashed `2888f2ea…`, matching **no commit anyone could name**; it now reports `104db0a4…` against a vouchable manifest value. **(b)/(c) REMAIN OPEN:** the wrapper is still fetched unversioned from `raw/branch/main` — this makes drift *visible*, it does not fix the channel. Also recorded: the 0440 sudoers file is not agent-readable, so its drift stays invisible. **Flips (2026-10-03):** `00` §A "The installer is PUBLISHED, not pushed" — the same discipline for the privileged wrappers. **Re-ranked 2026-10-03:** [P2] → P3: operator-only; (a) shipped 2026-07-21 (the report carries the wrapper sha256), (b)/(c) open. **Finding-shaped** — an R-424 instance; check against today's product before building. | CC |
|
||||
| **R-78** | Box system & updates | P3 | **`local_api` authority ruling — auto-reconcile vs detect-only** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-25, size M, roadmap state `idea (deferred OUT of R-77 on purpose)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **An OWED OPERATOR DECISION, not a defect — moved because an owed ruling hidden among feature ideas is the shape this session exists to remove.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | R-77 ships detection because the fix is genuinely undecided, and **both directions can lose customer-visible function**. **Direction 1 (today):** `controller.yaml` wins and drift is silent → the 2026-07-25 island migration blinded the whole fleet's control plane for 17.5 h (drive gate, guest-reboot recovery, quiesce/backup all degrade). R-77 makes that loud but does not stop it recurring. **Direction 2 (`bootstrap.json` wins, auto-reconcile on boot):** a guest whose `controller.yaml` is CORRECT and whose `bootstrap.json` is stale — a half-completed re-provision, a hand-repaired guest, a setup-wizard box — gets a **working channel clobbered on the next restart**, fleet-wide and silently, during a routine deploy. That is not obviously better than the bug. Needs a spike: which writer is authoritative per field (endpoint vs fingerprint vs token — `mergeLocalAPI` replaces the whole block, so they cannot be reconciled independently today), whether the agent side should stamp a generation/mtime so 'newer wins' is even expressible, and whether reconcile should require an operator ack. Until then the drift alert plus a manual edit is the supported path. | CC |
|
||||
@@ -324,10 +324,10 @@ stopping line that lies.
|
||||
| **R-194** | Box system & updates | P4 | **PVE's permission cache delays every grant-state verdict by an unknown amount, so "the agent can read it" and "the ACL exists" are not the same measurement.** Observed twice while validating R-190's self-repair on demo-felhom 2026-08-04: both ACL rows for `/storage/felhom-backup` were deleted, and `GET /access/permissions` continued to report `Datastore.AllocateSpace` present — for **~40 s** in one run and **~16 minutes** in another. During that window the capability probe reads healthy and the self-repair does not fire | **OPEN** | — | **Why it matters beyond the delay:** it puts a floor under how fast a lost grant can be noticed, it makes any single permission read a lagging indicator, and — the interesting part — **it is a candidate contributor to R-190's own timeline**: a grant removed at an unknown moment could keep working until a cache expiry, which is exactly the shape of *worked at 04:44, refused at 09:24*. That does not explain what removed it, but it may explain when the refusal SURFACED, and the two have been treated as the same instant. **Not a defect in our code** — it is PVE behaviour, and the mitigation already tolerates it (the repair fires on the next probe after the cache clears). **What is worth deciding:** whether the store-grant probe should ALSO consult the storage content listing as a second signal, since that appeared to reflect the loss immediately (`{"data":[]}` while the permission read still said present) — two signals disagreeing is itself information, and today only one of them is read | CC |
|
||||
| **R-373** | Box system & updates | P4 | **`SysDataGrowGB` is the intended lever for the system-data volume, it works, and nothing sets it.** Written down 2026-08-02 in `audits/SPIKE-recovery-unit-space-2026-08-02.md:230-232`, under an explicit *"### Not filed"* heading: the 20 G / 50 G mismatch was ruled a tier-sizing decision rather than a defect, *"`SysDataGrowGB` is the intended lever and it works; nothing sets it."* A lever with no caller is the same shape as R-368's comment — a setting that names a behaviour nothing invokes. **Age when filed: 20 days.** | **OPEN — LOW** | R-368 (same shape) | Either wire it to something an operator can reach, or remove it and record the sizing decision where a reader will meet it. | CC |
|
||||
| **R-835** | Box system & updates | P3 | **Turning Docker's `live-restore` OFF with a restart stops every running container and starts none.** MEASURED 2026-10-04 on scratch 9202: `live-restore` on (via `systemctl reload docker`, which does enable it) kept all 6 containers running across two engine steps; `systemctl reload` with the baked `daemon.json` did NOT turn it off; a `systemctl restart docker` did — and the new daemon stopped every container (`Exited (0)`, `Removing stale sandbox … isRestore=false`) and restarted none, though all are `unless-stopped`. Nothing brought them back for 3.5 min. A precondition for the Docker slow lane (`11` C5): if `live-restore` ships, turning it off must be a guarded act (stop apps first), never a plain restart. `audits/os-updates-spike-2026-10-04/partG/` | **READY — design input, owner: CC** | — | — | CC |
|
||||
| **R-836** | Box system & updates | P3 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin <new> --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **READY — measure before the kernel slow lane; owner: CC + operator (reboots)** | — | — | CC |
|
||||
| **R-836** | Box system & updates | P3 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin <new> --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot <new>`: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** | — | — | CC |
|
||||
| **R-851** | Box system & updates | P3 | **A host that panics stays stopped: `kernel.panic = 0`.** READ 2026-10-04 on demo-hp (os-host-lane Part E): `kernel.panic = 0`, `kernel.panic_on_oops = 0`. The brief assumed the host restarts after a panic; it does not — the box stays down until a person power-cycles it, and only `softdog` (dead in a panic) runs. A fix (`kernel.panic = 10` set by the installer, maybe `panic_on_oops`) changes how every box behaves and what a household sees, so it is the operator's call, together with the kernel lane. `audits/os-host-lane-2026-10-04/partE/e0-readonly-demo-hp.txt` | **READY — operator decision (with the kernel lane); owner: operator** | — | — | operator |
|
||||
| **R-839** | Box system & updates | P3 | **After a Docker restart that stopped every container, the boot sweep HELD an app whose `app.yaml` `HDD_PATH` names its user folder instead of the drive.** MEASURED 2026-10-04 on scratch 9202 (paperless-ngx): `bootrecon` logged `drive /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx is not a live mountpoint — NOT starting it`, although the drive `/mnt/felhom-drives/scratch_hdd` IS a mountpoint; the app stayed down until started by hand. The customer boxes' 9201s showed no hold. **Not diagnosed:** which writer put a per-app path in `HDD_PATH` on 9202, and whether a household box can get it. `audits/os-updates-spike-2026-10-04/partG/SUMMARY.md` | **READY — diagnose; owner: CC** | — | — | CC |
|
||||
| **R-840** | Box system & updates | P2 | **A new root wrapper or a new sudoers line cannot reach a box that is already installed — there is no product route.** FOUND 2026-10-04 (guest OS lane, Part B 1): the agent's signed `agent_update` replaces ONLY the binary (`configs/felhom-selfupdate-guarded` swaps `/usr/local/bin/felhom-agent`); wrappers (`/usr/local/sbin/felhom-*`) and `/etc/sudoers.d/felhom-agent` are written only by `felhom-host-install.sh` step 5 on a new box. So agent v0.140.0's OS leg is DEAD on every existing box (the capability probe will show `osapply-run` missing) until someone installs `felhom-os-apply` + the FELHOM_OSAPPLY line by hand — which is what the demo boxes got this session. The `wireguard-tools` line (S3) reached the fleet the same way: by reinstall or by hand. **Proposal:** a signed `agent_config_update` op (operator-signed like `agent_update`) whose params pin the agent TAG and the sha256 of a config bundle (sudoers + wrappers); the self-update wrapper installs it as root after `visudo -cf` and a syntax check, keeps the previous copies, and the capability probe confirms. `audits/os-guest-lane-2026-10-04/README.md` | **READY — design + operator go; owner: CC** **RULED 2026-10-04 ~12:20 (`09` §3 decision 82): not built now — "There are no older boxes"; no installed box exists outside the two demo boxes, which take wrapper changes by hand. Kept open, not re-ranked. Reviewer's note (as written): every box installed from now on becomes an "older box" the first time the wrapper or the sudoers line changes again (the 2026-10-04 host-lane brief changes the wrapper); R-840 must be solved before the first box that cannot be reached by hand.** | — | — | CC |
|
||||
| **R-841** | Box system & updates | P3 | **The agent's `cloudflared` health probe reads a host systemd unit that does not exist — every box reports its tunnel `inactive`.** FOUND 2026-10-04: `felhom-agent/internal/hub/cloudflared.go` runs `systemctl is-active cloudflared` on the HOST; cloudflared is a container in the GUEST (`11` C8), so demo-hp answers `inactive` / `Unit cloudflared.service could not be found`, and the hub stores that in `cloudflared_status` for every box. The field is equally consistent with "tunnel down" and "never checked" (R-96 rule 3). Fix direction: read the guest's container state (the agent already may `pct exec * -- docker inspect -f *`), or drop the field; and say so in `03` (corrected 2026-10-04). | **READY — owner: CC** | — | — | CC |
|
||||
|
||||
## Monitoring & notifications — 24 rows (P2 2, P3 16, P4 6)
|
||||
|
||||
@@ -358,7 +358,7 @@ stopping line that lies.
|
||||
| **R-348** | Monitoring & notifications | P4 | **Every agent restart blanks the reported backup list for up to ~18 hours, and the comment that covers it says "unaffected".** Observed 2026-08-20 while deploying R-344: the first host reports after `demo-hp`'s agent restart carry **`0 backups`** (11:15:50 and 11:30:52 CEST, two consecutive), while the box's own `pvesm list` shows archives present on **both** tiers. `internal/backup/store.go`'s `Store` is in-memory and `byTarget` is repopulated only when a backup **runs** — daily for the local tier, weekly for offsite — so the field reads 0 until the next run. `restore_tests` did **not** blank, because that half has a durable on-disk companion (`RestoreTestState`, R-189). **It blinds no alarm, and that was CHECKED rather than assumed.** `hub/internal/monitor/deadline.go` scans back over stored reports with a 7-day `backupEvidenceLookback` whose own comment names this exact case — *"when the LATEST report carries none... and against an agent that stayed restarted for days"* — and `pbs_snapshots` stayed populated at 2 regardless. So this is an observability wart, **not** a safety hole, and it is filed at that severity deliberately. **What is actually wrong is the comment.** The `Store` doc says *"Backups are unaffected — their freshness has a ground truth on the storage (R-84)"*. That is true of the **consequence** and false of the **field**, and it sits three lines below a paragraph explaining that the very same sentence about restore-tests *"used to be here and it is now FALSE"* — so the file already carries one correction of this shape and invites the next reader to trust the surviving half. | **READY (XS) — NEW 2026-08-20** | — | Say what is measured: the field IS lost on restart and repopulates only when a backup runs; the freshness VERDICT is unaffected because the hub looks back 7 days. **Name `backupEvidenceLookback` in the comment** so the cross-repo dependency is visible from the agent side — today the agent's claim of safety rests on a hub constant it does not mention. Per the workspace rule, a comment asserting an invariant needs a test pinning it: the pin belongs on the HUB side, asserting the verdict survives a report carrying `backups: []`. | CC |
|
||||
| **R-371** | Monitoring & notifications | P4 | **The off-site tier is the only backup tier that announces nothing on success.** Written down 2026-08-05 in `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md:508-513` and explicitly *"recorded, not filed"*: the off-site run emits **no hub event at all**, while both lesser tiers do (`db_dump_completed`, `crossdrive_completed`). Failures are covered by `backup_run_failures` and staleness by the hub's 8-day tier deadline, which is why it was judged a wrinkle. **Still true 2026-08-22** — the 2026-08-21 drill's own event dump shows `db_dump_completed` and six `crossdrive_completed` rows and no off-site success event. **Age when filed: 17 days.** | **OPEN — LOW** | — | Either emit one, or record deliberately that the highest-value tier is silent on success and say why. | CC |
|
||||
|
||||
## Hub & operator — 23 rows (P2 1, P3 7, P4 15)
|
||||
## Hub & operator — 24 rows (P2 1, P3 7, P4 16)
|
||||
|
||||
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
@@ -384,7 +384,8 @@ stopping line that lies.
|
||||
| **R-719** | Hub & operator | P4 | **[P2-MEDIUM] A customer who already exists never gets a fresh self-bind link when their new box registers: the last link expires in 7 days and nothing re-sends it.** MEASURED 2026-09-29 (new-household drill, `tester-1`): the previous link went out 2026-09-17 07:25 UTC at a host delete and expired 2026-09-24; the box registered at 19:11:30 UTC and its console told the volunteer to open the link from their e-mail — there was none that worked. Hub source: the link is sent at customer creation, RESET, e-mail set on a box-less customer and host delete (`selfbind_mint.go` callers `hosts.go:908`, `configs.go:850`, `customer_reset.go:162`) — never on appliance registration. The volunteer guide says the operator needs to press nothing. The operator pressed „Send self-bind link" (the mail arrived in 1 s) — an operator step the volunteer depends on, recorded, not an intervention. **Fix direction:** send the link when an unclaimed appliance registers while a box-less customer waits with no live link (R-509's first fix shape), or the guide's operator part says: press it the day the volunteer installs. Evidence: `audits/evidence-drill-new-household-2026-09-30/` `phase0/operator-steps.txt`. **CHANGED AND BUILT 2026-09-30 (hub v0.126.0) — the brief's shape was not buildable:** a box registers UNCLAIMED (uuid, MACs, host keys, hardware — nothing of a customer), so "send the link when their box registers" would mail every waiting customer. Built instead: the expired AND used link pages offer „Új linket kérek" → a fresh link to the address registered for that link's customer, only when it has no box, ≤1/h per customer, identical answer for any token (no oracle). Live: the button on the real hub, the same page for a made-up token, no mail; the mint+send path unit-proven (RP42). Evidence: `audits/evidence-fixes-first-tester-2026-09-30/``partD/`. | **WAITING-ON-OPERATOR** (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the operator has not reviewed the changed page shape ("operator may prefer another"), and mint+send is unit-proven only) — **CLOSED 2026-09-30 — hub v0.126.0 (changed shape; operator may prefer another)** | — | — | operator |
|
||||
| **R-814** | Hub & operator | P4 | `PBS-storage-1` (u629193, box 611421) still `status=active`, 19.9 MB | **VERIFY** (2026-10-03 triage: a July watch row with no id; given R-814. WAITING-ON-OPERATOR — no record found that the box was deleted.) — WAITING-ON-OPERATOR | operator console | Delete the box | operator |
|
||||
| **R-844** | Hub & operator | P4 | **The household's OS-update line exists only on the hub's customer timeline.** 2026-10-04: the box itself has no event surface for agent results (the controller UI shows no timeline), so `os_update_applied` is a hub customer event (info: recorded, never mailed). Its stored text is the hub's English sentence; the hu/en bundle text (`mail.event.os_update_applied`) is used only if it is ever mailed. Fix direction: a controller-side line (the controller already polls the agent's local API) when the box gets a household timeline. `audits/os-guest-lane-2026-10-04/partG/hub-customer-timeline-demo-hp.txt` | **READY — owner: CC** | — | — | CC |
|
||||
| **R-845** | Hub & operator | P4 | **One OS-leg pass takes 3–4 minutes even when it installs 3 packages**: the wrapper's inventory (an `apt-get update`, `apt-cache policy` over every installed package, a `/proc/*/maps` scan for restart-needed, two simulations) dominates; measured 174–245 s per pass on the demo boxes vs 3.8–31.7 s for the install itself. It runs at night under the heavy-op gate, so it delays a restore-test by minutes, nothing worse. Fix direction: one `apt-cache policy` per run and the restart scan only after an install. `audits/os-guest-lane-2026-10-04/partG/` | **READY — owner: CC** | — | — | CC |
|
||||
| **R-848** | Hub & operator | P4 | **A held host package is invisible to the hub.** MEASURED 2026-10-04 on demo-hp (undo runbook proof): with `tzdata` held after a by-hand undo, the wrapper's `pending` stayed 78 — apt's simulation leaves held packages out, so the fleet view shows nothing and no alarm can see a hold that was forgotten. The hold lives only in the incident's register row (`runbooks/os-updates-host-undo.md`). Fix direction: the wrapper reports `apt-mark showhold` and the fleet line shows it. | **READY — owner: CC** | — | — | CC |
|
||||
| **R-849** | Hub & operator | P4 | **The fleet view's GUEST "reboot needed since" never clears.** 2026-10-04: the guest is scanned only after an install (R-845, one `pct exec`), so a guest restart is never seen; the guest line keeps the date of the last install that said "needed". No alarm reads the guest line (the reboot alarm is host-only), so it is display only. Fix direction: scan the guest on every pass too (one `pct exec`, ~1 s) or hide the guest date. `audits/os-host-lane-2026-10-04/partC/live/fleet-during-ring1-test.json` | **READY — owner: CC** | — | — | CC |
|
||||
|
||||
## Business & legal — 7 rows (P2 4, P4 3)
|
||||
|
||||
|
||||
@@ -0,0 +1,81 @@
|
||||
# Put one host package back by hand (OS updates, host fast lane)
|
||||
|
||||
> **When:** the hub mailed `os_update_health_failed` for the **host** layer (the box's base system), or the operator
|
||||
> sees a host problem that started with a host OS pass, and ONE package is the suspect.
|
||||
> **Who:** the operator, or CC with the operator's word, as root on the box's Proxmox host (`ssh <host>`).
|
||||
> **Owner design:** `architecture/11-os-updates.md` §5.6 (host row), §8 step 3. **Proved** on demo-hp 2026-10-04 with
|
||||
> `tzdata` (2026c → 2026b, held; then released and re-installed by the next ring-0 pass) and again on demo-felhom for
|
||||
> the ring-1 test — evidence `audits/os-host-lane-2026-10-04/partB/undo/` and `partB/ring1/`.
|
||||
|
||||
There is **no automatic undo** for the host. A host cannot be snapshotted the way the old design assumed, and the
|
||||
fast lane only ever installs Debian / Debian-Security packages, never a kernel, boot or firmware package (wrapper
|
||||
refusals R14, R12). So the undo is small: install the previous version of the suspect package from
|
||||
`snapshot.debian.org`, which keeps every version Debian ever published.
|
||||
|
||||
## 1. Find the suspect and its previous version
|
||||
|
||||
```bash
|
||||
# the host pass's own apt run (Requested-By: felhom-agent); each line "name:arch (old, new)"
|
||||
grep -B2 -A6 "Requested-By: felhom-agent" /var/log/apt/history.log | tail -20
|
||||
PKG=<name>; OLD=<old version from that line>
|
||||
```
|
||||
|
||||
## 2. Pick the snapshot timestamp
|
||||
|
||||
- **Normal case:** the timestamp of the **previous host release** — the hub fleet view (`GET /os/fleet`, the box's
|
||||
host line names its release; the release's approval time IS its snapshot time, `YYYYMMDDTHHMMSSZ`). At that time
|
||||
the old version was the current one.
|
||||
- **No previous host release** (a ring-0 box, or the first host pass): ask snapshot.debian.org when the **NEW**
|
||||
(installed) version first appeared, and use that time. At that moment the suite still carried the old version.
|
||||
**Do not use the OLD version's `first_seen`:** that is when it reached Debian *unstable*, and `trixie` may not
|
||||
have had it yet — measured on demo-hp 2026-10-04: `eject 2.41-5` at its own `first_seen` → `Version not found`.
|
||||
|
||||
```bash
|
||||
NEW=$(dpkg-query -W -f='${Version}' "$PKG")
|
||||
curl -s "https://snapshot.debian.org/mr/binary/$PKG/$NEW/binfiles?fileinfo=1" | python3 -c '
|
||||
import json,sys; d=json.load(sys.stdin)
|
||||
print(sorted(f["first_seen"] for v in d["fileinfo"].values() for f in v)[0])'
|
||||
# → e.g. 20260831T204404Z
|
||||
TS=<that timestamp>
|
||||
```
|
||||
|
||||
## 3. Install the old version from the snapshot, then remove the snapshot source
|
||||
|
||||
```bash
|
||||
. /etc/os-release
|
||||
cat > /etc/apt/sources.list.d/felhom-undo-snapshot.list <<EOF
|
||||
deb [check-valid-until=no] http://snapshot.debian.org/archive/debian/$TS $VERSION_CODENAME main
|
||||
deb [check-valid-until=no] http://snapshot.debian.org/archive/debian-security/$TS $VERSION_CODENAME-security main
|
||||
EOF
|
||||
apt-get -q update
|
||||
apt-get -s install --allow-downgrades "$PKG=$OLD" # READ the simulation: only $PKG may change. If apt
|
||||
# refuses, the package pins its siblings to the SAME
|
||||
# version (measured: eject needs libmount1 = 2.41-5) —
|
||||
# list them all in one command, or stop.
|
||||
apt-get -y install --allow-downgrades "$PKG=$OLD"
|
||||
apt-mark hold "$PKG" # or the next ring-0 night installs the new one again
|
||||
rm -f /etc/apt/sources.list.d/felhom-undo-snapshot.list
|
||||
apt-get -q update
|
||||
dpkg -s "$PKG" | grep -E "^(Status|Version)"
|
||||
```
|
||||
|
||||
**The hold is the important line.** Without it the next night pass (ring 0) or the next approved release (ring 1)
|
||||
installs the new version again. Write the hold into the register row that tracks the incident; take it off
|
||||
(`apt-mark unhold "$PKG"`) when a fixed version is out.
|
||||
|
||||
## 4. Check the box
|
||||
|
||||
- `systemctl is-active pveproxy pvedaemon pvestatd pve-cluster felhom-agent` → all `active`.
|
||||
- `pct status <customer vmid>` → `running`.
|
||||
- The host health rule (`11` §8.2) is what the leg checks; run the debug action to see it pass:
|
||||
`sudo -u felhom-agent /usr/local/bin/felhom-agent --config /etc/felhom-agent/agent.json --selftest=os-update -vmid <vmid>`
|
||||
— **note: on ring 0 this also installs every pending fast-lane fix.** A held package is NOT reported as pending
|
||||
at all (measured on demo-hp: `pending` stayed 78 with `tzdata` held) — **the hub cannot see a hold**, so the hold
|
||||
lives only in the register row (R-848).
|
||||
|
||||
## What this runbook does NOT cover
|
||||
|
||||
- A kernel: the fast lane never installs one (R14). The kernel lane is `11` §5.6 (not built).
|
||||
- A Proxmox-origin package (pve-*, qemu-server, …): never installed by the fast lane (R12/origin rule). Proxmox's
|
||||
own repository keeps old versions; that undo is not written yet.
|
||||
- The customer guest: its undo is last night's whole-guest backup (decision 81).
|
||||
Reference in New Issue
Block a user