diff --git a/CONTEXT.md b/CONTEXT.md index bc0795ee..a25c3087 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -16,6 +16,16 @@ > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +> **2026-10-04 (late afternoon) — OS updates: host fast lane + fleet view + alarms BUILT (`11` §8.2–§8.3); the tunnel +> status is true (R-841).** Agent v0.141.0 → v0.141.1, hub v0.131.0 → v0.131.1, controller v0.292.0 (cloudflared +> readiness health check). **Decided by CC unattended — operator may reverse:** `09` §3 decisions **84** (appliance proof +> = the root-owned install record), **85** (alarm numbers 7/14/7/14 days, configuration), **86** (a same-session patch +> release of agent and hub for the host "reboot needed" defect, R-846). Live: `tunnel_down`/`tunnel_recovered` on +> demo-hp; ring 0 host pass on demo-felhom (108 Debian packages); a 605-package host release; ring 1 exact install; +> by-hand host undo proved; leg 23–32 s with nothing to install. Kernel spike (R-836, narrowed): GRUB's one-shot is not +> a one-shot on LVM `/boot`; Secure Boot fine; `kernel.panic = 0` (R-851); `sp5100_tco` answers. demo-hp now runs and +> defaults to kernel 7.0.14-20. Docker slow lane designed (`11` §5.8). `REPORT-os-host-lane-2026-10-04.md`. + > **Rulings 2026-10-04 (~12:20) — recorded before the work (host fast lane brief).** `09` §3 decisions **81** (R-842 A: > the undo is the whole-guest backup by hand; R-842 closed), **82** (R-840 not built now — "There are no older boxes"; > row kept open with the reviewer's note) and **83** (next: R-841, the host fast lane, OS-update improvements). diff --git a/STATUS.md b/STATUS.md index 2e5871d1..7da59f97 100644 --- a/STATUS.md +++ b/STATUS.md @@ -2,8 +2,46 @@ **Ready for the first real tester (Tester-2): yes. You confirmed the tunnel route and the connect mails (2026-09-30).** -**Updated 2026-10-04 (afternoon): guest system updates are automatic on the demo boxes. Both demo boxes run -controller 0.291.0 and host agent 0.140.0. Hub 0.130.0. New installs get golden 0.291.0.** +**Updated 2026-10-04 (evening): the HOST's security fixes install themselves too, the hub has a fleet view and four +OS alarms, and the tunnel status is true. Both demo boxes run controller 0.292.0 and host agent 0.141.1. Hub 0.131.1. +New installs: see the golden line in the section below.** + +## Today (2026-10-04, evening): host fixes, the fleet view, a true tunnel status + +**Decisions I took myself (you may reverse each):** +- Only a box the installer itself recorded as "appliance" gets host updates (a record the agent cannot change). +- The four alarm times: no update run for 7 days; "reboot needed" for 14 days; the demo boxes approve nothing for 7 + days; a box has fixes nobody approved for 14 days. All four are settings. +- I released the host agent and the hub a second time today. The first release said "reboot needed" wrongly; the new + 14-day alarm would then have mailed you about boxes you had already rebooted. + +**One decision for you (a safe default if you say nothing) — the Docker engine updates (designed, not built):** +1. **Turn on Docker's "live-restore" on every box.** With it, a Docker engine update restarts no app (measured: 0 + restarts). Without it, every app stops for about 30 seconds per engine update. + - **A (my pick):** turn it on — in the new-install image and once on existing boxes. It goes on without restarting + anything. Cost: it must never be turned off by a plain restart again (that stops every app and starts none). + - **B:** leave it off. Every Docker update then means ~30 seconds of every app being down, at night. + - **If you say nothing:** nothing changes; Docker updates stay unbuilt. + +**What I did:** +- **The host's Debian fixes now install themselves**, after the guest's, on the same night run, only on appliances, + never a kernel or boot package, never a reboot. Proven: demo-felhom installed 108 host fixes and stayed healthy; all + 108 came from Debian. A box made "ring 1" for a test installed exactly the one approved host version it lacked. +- **A way to put one host package back by hand** is written and proven on demo-hp (and the test taught it two fixes). +- **The fleet view** in the hub: one line per box with its updates, "reboot needed since", and the tunnel. +- **The tunnel status is now true:** running, not running, or unknown. I blocked demo-hp's tunnel: after two reports + (about 30 minutes) you got a "tunnel down" mail; when I unblocked it, "recovered". A simply stopped tunnel heals + itself within 5 minutes, before the hub can even see it. +- **The night run is fast:** 23–32 seconds when there is nothing to install (target was under 60). +- **The kernel test on demo-hp (your two reboots):** Secure Boot works with it, but GRUB's "boot once" does not work + on our boxes: the second plain reboot came up on the NEW kernel again. So a new kernel that hangs would stay. That + must be solved before kernel updates. demo-hp now runs the newer kernel, healthy. A watchdog chip exists on demo-hp. +- **Found and fixed during the live test:** "reboot needed" was wrong in two ways (fixed in the second release). +- **Rows:** 2 closed, 2 opened-and-closed the same day, 3 opened. The list went from 332 to 333. + +**Needs you later (nothing breaks if you wait):** +- **A host that crashes does not restart by itself** (Linux's "panic" setting is off). Changing it changes how every + box behaves, so it is your call, together with kernel updates. ## Today (2026-10-04, afternoon): the guest's security fixes install themselves diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index dbf2edf3..7a4001fb 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -232,7 +232,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis | **The hub reports LOSS OF VISIBILITY into either off-site store (not just how full it is)** | hub **v0.106.0** (R-339) | **IMPLEMENTED — deliberately NOT proven-live** | Both box checkers count consecutive failed fetch windows and emit `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default 3 windows (≈30–45 min), each with a paired `*_recovered` all-clear routed via `recoveredPairedDownTypes` — required because the recoveries are severity `info`, which `severityNotifies` drops. Scopes stay customer-less (`pbsdr-box` / `pool-box`) → operator channel only. Fill logic untouched: a degraded read still drives no band transition. Evidence: `internal/monitor/box_reachability_test.go` + the cross-package wiring test in `internal/notify/`, which asserts an actual operator mail rather than a map entry. **Filed BECAUSE of a measured gap**, not a hypothesis: the 2026-08-18 ep0 outage ran 9 h 37 m with the hub silent | **The gap that remains is R-340**, and it is not small: the ep0 read is the `usage` op, which rides the LOCAL API daemon — the daemon that incident explicitly cleared — so this check would have shown GREEN for that entire outage. It closes "ep0 is unreachable as a host"; it does not close what actually happened. **No live or constructed outage has exercised the emit path**, and one cannot be manufactured against ep0 (Tier 2, protected) | | Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | | | Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | | -| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.140.0, hub v0.130.0 | **PARTIAL — the GUEST's Debian fast lane is PROVEN-LIVE (2026-10-04); the host, Docker and the kernel are MISSING** | `audits/os-guest-lane-2026-10-04/` — ring 0 (both demo boxes) installed 53 Debian fixes each, healthy; the hub approved a 272-package release (TEST wait 2 min, then the ruled 24 h + 1 night restored); demo-felhom as ring 1 installed exactly the 3 approved versions; a stopped app → `health_failed` + operator mail. Design `architecture/11-os-updates.md` §8.1 | **No automatic undo** (a customer guest cannot be snapshotted, R-837 → R-842); existing boxes need the wrapper + sudoers by hand (R-840); host / Docker / kernel lanes not built (R-812, R-835, R-836); `felhom-host-install.sh:2133-2136` still runs no host upgrades | +| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.141.1, hub v0.131.1 | **PARTIAL — the GUEST and HOST Debian fast lanes are PROVEN-LIVE (2026-10-04), with the fleet view and four operator alarms; Docker and the kernel are MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/` — ring 0 on both demo hosts (demo-felhom 108 Debian host packages, healthy, 0 Proxmox-origin); a 605-package host release approved (TEST wait, then the ruled 24 h + 1 night restored); demo-felhom as ring 1 installed exactly the one host version it lacked; the by-hand host undo proved (`runbooks/os-updates-host-undo.md`); `tunnel_down` / `tunnel_recovered` live. Design `architecture/11-os-updates.md` §8.1–§8.3 | **No automatic undo** (guest: last night's backup, decision 81; host: the by-hand runbook); existing boxes need the wrapper + sudoers by hand (R-840); Docker and kernel lanes not built (R-812, R-835, R-836); a host panic is not restarted (R-851) | | **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** | | **A deletion of a customer's off-site history is NOTICED within a day** | hub **v0.111.0** (R-431) | **IMPLEMENTED — not yet PROVEN-LIVE** | 09-01 | `hub/internal/monitor/offsite.go` — third signal beside FILL and STALENESS. **On the hub deliberately:** a detector on the box is one the deletion can silence. Alarms when the reported count falls by more than HALF and by at least 5, guarded by `StatsKnown` (R-331), the declared `State` (R-204) and run success (R-100). **Threshold reasoned, not invented:** over 12 898 reports every decrease lands on ZERO and predates `stats_known`; in the 380-report `stats_known` window there are none. **ACCEPTANCE: 9 009 real points replayed → ZERO alarms** (`offsite_r431_test.go`, fixture committed). **What PROVEN-LIVE would need and this does NOT have:** a real drop observed on a live box producing a real mail — the live firing done at ship time was driven through the hub's own path with synthetic counts, which is an end-to-end delivery proof, not a proof that a genuine deletion is caught. | diff --git a/documentation/architecture/03-host-agent.md b/documentation/architecture/03-host-agent.md index 5d78497a..bdaa0eb3 100644 --- a/documentation/architecture/03-host-agent.md +++ b/documentation/architecture/03-host-agent.md @@ -47,7 +47,7 @@ Owns: 1. **Proxmox lifecycle** — create/start/stop/destroy guests, snapshots, storage allocation. Via a scoped Proxmox API token (the **`FelhomAgent` operator role** — `proxmox-platform.md` §3.6, validated Phase 3 B3) for everything the API covers; raw host ops only where unavoidable. 2. **Storage management** — attach/classify targets, reconcile the storage manifest, mount USB-by-UUID, present mounts into guests. 3. **Backup/restore orchestration** — vzdump to the tiers, PBS, snapshot management, and the **self-restore-test**. -4. **Host & tunnel monitoring** — host metrics, guest up/down, storage-target status, and `cloudflared` health; reports the host domain to the hub. **[FACT, 2026-10-04] The "cloudflared health" leg reads a host unit that does not exist:** `internal/hub/cloudflared.go` runs `systemctl is-active cloudflared` on the HOST, but cloudflared is a container in the guest (below), so every box reports `inactive` (`Unit cloudflared.service could not be found`, demo-hp). R-841. +4. **Host & tunnel monitoring** — host metrics, guest up/down, storage-target status, and `cloudflared` health; reports the host domain to the hub. **[FACT, 2026-10-04] The "cloudflared health" leg reads a host unit that does not exist:** `internal/hub/cloudflared.go` runs `systemctl is-active cloudflared` on the HOST, but cloudflared is a container in the guest (below), so every box reports `inactive` (`Unit cloudflared.service could not be found`, demo-hp). R-841. **[FACT, FIXED agent v0.141.0 / controller v0.292.0, R-841 CLOSED]** The agent now reads the guest's `cloudflared` container through the existing `pct exec [0-9]* -- docker inspect -f *` sudoers line: state, exit code and the Docker health status of a check the controller adds (`cloudflared tunnel --metrics localhost:20241 ready` → cloudflared's own `/ready`, 200 only with a connection). Three states: `running` (healthy), `not_running` (stopped, absent, or running but NOT connected), `unknown` (could not ask, or the check is still starting) — `unknown` never alarms. No new sudoers line. 5. **Provisioning** — provision a guest **by restoring the golden base image** (§9), deploy the controller into it, hand it its bootstrap config; also **build and refresh the golden base image** itself. 6. **Hub control loop** — poll for desired state + signed jobs, reconcile, execute, report, heartbeat. 7. **Local API** — the per-guest authorization gate the controller calls. @@ -64,7 +64,7 @@ Explicitly does **not**: - **Native Go binary, systemd service** on the host: boot-start, `Restart=always`, systemd watchdog (kill+restart on hang), journald logging, resource limits. - **Root-minimized (boundary settled — Phase 3 B3).** The agent runs as a **non-root** service user with the scoped `FelhomAgent` token for all API-covered work + a **narrow `sudoers` allowlist** for true host ops. Per Phase 3 (B3) the boundary is settled: the entire per-customer guest lifecycle — provision (by restore, §9), config, start/stop, snapshot, backup, **restore**, destroy — is token-covered. Genuine OS-root is confined to: (1) building/refreshing the **golden base image** (`keyctl` create is `root@pam`-only — one-time at enrollment + a maintenance cadence, §9); (2) **host mounts** (USB mount-by-UUID, systemd mount units / fstab); (3) **SMART / hardware sensors**. Root therefore never sits on the per-customer path. See `proxmox-platform.md` §3.6 for the role + boundary table. -- ~~**`cloudflared` is a separate systemd service**, not embedded in the agent. … The agent **manages and health-watches** it (see §5) but the tunnel does not live or die with the agent process.~~ **[FACT, corrected 2026-10-04 — `11-os-updates.md` C8, R-838]** `cloudflared` is a **container in the customer guest** (`cloudflare/cloudflared:`), rendered and kept up by the in-guest controller (`felhom-controller/controller/internal/infra/infra.go`, `internal/stacks/infra.go`) and baked into the golden; there is no host systemd unit. It is still NOT embedded in the agent, so the data path survives the agent's death — and the agent neither manages it nor (see R-841) sees its health. Its version moves only by a controller release (the pin), on the monthly re-test (`runbooks/monthly-floating-retest.md` "Infrastructure pins"). +- ~~**`cloudflared` is a separate systemd service**, not embedded in the agent. … The agent **manages and health-watches** it (see §5) but the tunnel does not live or die with the agent process.~~ **[FACT, corrected 2026-10-04 — `11-os-updates.md` C8, R-838]** `cloudflared` is a **container in the customer guest** (`cloudflare/cloudflared:`), rendered and kept up by the in-guest controller (`felhom-controller/controller/internal/infra/infra.go`, `internal/stacks/infra.go`) and baked into the golden; there is no host systemd unit. It is still NOT embedded in the agent, so the data path survives the agent's death — and the agent does not manage it; it READS its health (R-841, agent v0.141.0). Its version moves only by a controller release (the pin), on the monthly re-test (`runbooks/monthly-floating-retest.md` "Infrastructure pins"). ## 4. Control model — reconcile + signed destructive ops @@ -143,7 +143,7 @@ notification) is the control. **Box-initiated poll.** The hub never connects inbound. Each poll cycle exchanges: -- **Up:** heartbeat + a host-domain state report — host CPU/RAM/disk, per-guest up/down + spec, storage-target status (USB connected? NFS/CIFS reachable? PBS reachable?), last backup per target, last restore-test result, `cloudflared` health (a host-unit probe that always reads `inactive` — R-841), agent + controller versions, audit-log tail. +- **Up:** heartbeat + a host-domain state report — host CPU/RAM/disk, per-guest up/down + spec, storage-target status (USB connected? NFS/CIFS reachable? PBS reachable?), last backup per target, last restore-test result, `cloudflared` health (`running` / `not_running` / `unknown` from the guest container's readiness check — R-841), agent + controller versions, audit-log tail. - **Down:** the current desired state, any pending signed one-shot jobs, and config (poll interval, update window, policy changes). **Dead-man's-switch (essential, not optional).** In a box-initiated model the heartbeat diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index f7238f00..4b026079 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -366,7 +366,10 @@ never moved by an update or an undo, and the unit restore accepts it. whole-guest backup (the controller drives it, inside [W+2h, W+6h)) ends SUCCESSFULLY on the primary tier, the agent waits 90 s and runs the guest's Debian fast lane — still holding the host-wide heavy-op gate, so it never overlaps a backup or a restore-test; at most once per 20 h. The backup minutes old is the guest's undo (no snapshot is possible, -R-837). A failed or missed backup → no OS leg that night. +R-837). A failed or missed backup → no OS leg that night. **[FACT, 2026-10-04 — agent v0.141.1, `11` §8.2]** On an +appliance, the HOST step follows under the same gate: after a healthy guest step only (a failed or unhealthy guest +step skips it), Debian-origin fixes only, never a kernel, boot or firmware package, never a reboot. Measured: both +steps with nothing to install, 23–32 s; a 108-package host pass, 70 s. **[FACT]** The three nightly legs derive from **one** customer-settable window start W at fixed offsets — db-dump at W, Tier-2 at W+60m, offsite at W+105m — so they can never be misordered diff --git a/documentation/architecture/08-alarm-ladder.md b/documentation/architecture/08-alarm-ladder.md index 5d72e1a7..160c8dfe 100644 --- a/documentation/architecture/08-alarm-ladder.md +++ b/documentation/architecture/08-alarm-ladder.md @@ -310,6 +310,32 @@ message, not a wider cooldown. --- +## 6.3 Box alarms outside the app ladder: the tunnel and OS updates [DESIGN, hub v0.131.0, 2026-10-04] + +These are **operator-only** (the household can act on none of them; `operatorOnlyEvents`, pinned by +`TestOSUpdateEvents_OperatorOnlyExceptApplied`). They go through the same dispatcher and severity contract (§6.1): +`info` is recorded and never mailed; `warning` and `error` are mailed. + +| Event | Severity | Raised when | Cleared | Pinned by | +|---|---|---|---|---| +| `tunnel_down` | error | the box's two newest host reports say the tunnel is `not_running` (more than one report cycle, 15 min) and the one before did not | the first `running` after it → `tunnel_recovered` (info) | `api/tunnel_test.go` | +| `os_update_stale` | warning | no successful OS leg for **7 days** while the switch is ON (agents that can run the leg only); the mail names the likely reason (box not reporting / the last leg's failure / no good night backup) | a successful leg | `TestAlarm_StaleLeg`, `TestAlarm_StaleNamesTheReason` | +| `os_reboot_needed` | warning | the host has needed a reboot for **14 days** (from the FIRST scanned report that said so) | a scanned pass that finds nothing (agent ≥ 0.141.1 scans the host every pass) | `TestAlarm_RebootNeeded`, `TestRebootNeeded_ClearedByAScannedPass` | +| `os_ring0_stalled` | error | ring 0 approved nothing for **7 days** in a layer while it has pending FAST-lane updates (a pending kernel does not count) | a new release | `TestAlarm_Ring0Stalled` | +| `os_not_covered` | warning | a ring-1 box has had fast-lane packages no approved release names for **14 days** | the packages are covered or gone | `TestAlarm_NotCovered` | + +- **`unknown` never alarms** (R-96 rule 3): a probe that could not ask is neither up nor down. An `unknown` report + breaks a `not_running` run. +- **A stopped cloudflared heals itself before the hub can see it** (measured 2026-10-04): the controller's + protected-container check recreates it within 5 minutes, and the host reports every 15. So `tunnel_down` catches + what the box cannot heal — a running container with no connection (wrong token, blocked network). +- The OS alarms are checked **hourly**, re-sent at most **once a week** while true, and forgotten when false, so the + next occurrence is announced again. The four numbers are configuration (`OS_ALARM_STALE_AFTER`, + `OS_ALARM_REBOOT_AFTER`, `OS_ALARM_RING0_STALL_AFTER`, `OS_ALARM_NOT_COVERED_AFTER`) — *decided by CC unattended, + operator may reverse* (`11` §8.3). + +--- + ## 7. The intent test [DESIGN, R-386 — CLOSED controller v0.223.0] **"The customer stopped this" is asked of the FIELD THAT RECORDS IT, never inferred from the state.** diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index e81d5c6c..f10c53c5 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -778,6 +778,22 @@ its length, and both fixes cost something the household would notice — operato 83. **Next: the tunnel status (R-841), the host fast lane (`11` §8 step 3), and other OS-update improvements.** *Operator ruling 2026-10-04 ~12:20.* +### 2026-10-04 (afternoon) — decided by CC unattended — operator may reverse (host fast lane brief) + +84. **Which record proves a box is an appliance, for the host fast lane?** Options: (a) `agent.json` + `deployment_mode` — the agent can write it, so a compromised agent could claim "appliance" and unlock host + updates; (b) the installer's ROOT-owned `/var/lib/felhom-install/state.json` `mode` — written once, as root, at + install. **Chosen (b):** the wrapper is the root fence and must not trust a file the agent can change. Cost: a box + whose install record is missing gets no host step (fails closed). `11` §8.2. +85. **The four OS alarm thresholds.** Options: shorter (3/7 days — noisy: a weekend away alarms) or longer (14/30 — + a stopped box goes unseen for weeks). **Chosen:** no OS leg 7 days, reboot needed 14, ring 0 stalled 7, not covered + 14; all four are configuration. Cost: a real stop is seen after a week, not a night. `11` §8.3, `08` §6.3. +86. **A second release of the agent (v0.141.1) and the hub (v0.131.1) in the same session**, against "one release per + repo". Options: (a) keep v0.141.0 / v0.131.0 and file the defect — the host "reboot needed" stays wrong (it hid + `lxc-start`, and a reboot never cleared it), so the new 14-day alarm would fire on rebooted hosts; (b) patch now. + **Chosen (b):** a known-false operator alarm is worse than an extra release; both patches are small and red-proved. + R-846. + ### 2026-09-30 (day) — operator notes, recorded before the work - **The day brief runs by day.** Every backup and automatic-update test is started by hand — the night chain's debug diff --git a/documentation/architecture/11-os-updates.md b/documentation/architecture/11-os-updates.md index 399faa88..0f723141 100644 --- a/documentation/architecture/11-os-updates.md +++ b/documentation/architecture/11-os-updates.md @@ -324,8 +324,8 @@ must never overlap a backup, a restore-test or a self-update.~~ |---|---|---| | Guest packages | Restore the guest snapshot taken just before the update | It also undoes app data written after the snapshot. Use it only inside the health window, before apps have written much. After that, install the previous version (needs OPEN Q1). **[FACT] (C6)** Customer guests are on LVM-thin and can snapshot; scratch 9202 (`dir` storage) cannot (`snapshot feature is not available`), so the measured undo was a whole-guest backup + restore: **73 s down**, all apps healthy, libc back. The snapshot rollback itself is unmeasured (R-837). | | Docker engine | Install the previous version (Docker's repository keeps old versions — OPEN Q2) | Every container restarts again. **[FACT] (C5)** Docker keeps 46 `docker-ce` versions. A step (either way) without `live-restore`: 6 of 6 containers restart, the app silent **26.5–30 s**, healthy at +42–45 s. With `live-restore` on: **0 restarts, no gap**, also across a containerd step. But a restart that turns `live-restore` OFF stops every container and starts NONE (`unless-stopped` ignored) — on 9202 they stayed down until restarted by hand; `systemctl reload` turns it on but not off. | -| Host packages | Install the previous version | Only if the source still has it (OPEN Q1/Q2). **[FACT]** `rsync` back to its pre-update version: `not found`; `libpng16-16t64` back to the point-release version: worked. A Debian undo needs the snapshot archive (C2). | -| Host kernel | ~~Boot the previous kernel. Proxmox can boot a new kernel **once** (`proxmox-boot-tool kernel pin --next-boot`). If that boot fails, the next boot uses the old kernel again. Make the new kernel permanent only after a healthy boot.~~ **[FACT] (C4)** Both demo hosts boot UEFI + GRUB (no proxmox-boot-tool ESPs). **Installing a kernel makes it the GRUB default at once.** `--next-boot` on GRUB writes an ordinary `GRUB_DEFAULT` + `update-grub`; `proxmox-boot-cleanup.service` clears it only once a boot reaches userspace. Measured on demo-hp (old kernel pinned permanently FIRST, new pinned for next boot): boot 1 → `7.0.14-20-pve`, healthy, 60 s; boot 2 → back on `7.0.2-6-pve`, healthy, 76 s. **So the fallback works after a boot that succeeds; read from the code, a kernel that hangs before userspace stays the default on every power cycle.** `[PROPOSAL]` use GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel as the saved default — unmeasured (R-836). | ~~If the new kernel hangs, someone must switch the box off and on. The spike checks whether a hardware watchdog can do that (OPEN Q4).~~ **[FACT]** Only `softdog` runs (loaded by `watchdog-mux`); it cannot rescue a kernel that never boots. demo-hp has an AMD FCH whose `sp5100_tco` driver ships but is not loaded; untested. A hang still needs a person. | +| Host packages | Install the previous version | Only if the source still has it (OPEN Q1/Q2). **[FACT]** `rsync` back to its pre-update version: `not found`; `libpng16-16t64` back to the point-release version: worked. A Debian undo needs the snapshot archive (C2). **[FACT, 2026-10-04]** Proved by hand: `runbooks/os-updates-host-undo.md` (demo-hp, `tzdata` back one version from `snapshot.debian.org`, held, released). Use the NEW version's `first_seen` as the timestamp when there is no previous host release; a package that pins its siblings (`eject` → `libmount1 =`) goes back only with them. | +| Host kernel | ~~Boot the previous kernel. Proxmox can boot a new kernel **once** (`proxmox-boot-tool kernel pin --next-boot`). If that boot fails, the next boot uses the old kernel again. Make the new kernel permanent only after a healthy boot.~~ **[FACT] (C4)** Both demo hosts boot UEFI + GRUB (no proxmox-boot-tool ESPs). **Installing a kernel makes it the GRUB default at once.** `--next-boot` on GRUB writes an ordinary `GRUB_DEFAULT` + `update-grub`; `proxmox-boot-cleanup.service` clears it only once a boot reaches userspace. Measured on demo-hp (old kernel pinned permanently FIRST, new pinned for next boot): boot 1 → `7.0.14-20-pve`, healthy, 60 s; boot 2 → back on `7.0.2-6-pve`, healthy, 76 s. **So the fallback works after a boot that succeeds; read from the code, a kernel that hangs before userspace stays the default on every power cycle.** ~~`[PROPOSAL]` use GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel as the saved default — unmeasured (R-836).~~ **[FACT, 2026-10-04, demo-hp, operator's word before each reboot] GRUB's one-shot is NOT a one-shot here either.** With `GRUB_DEFAULT=saved` (old 7.0.2-6 saved) and `grub-reboot` 7.0.14-20: boot 1 → 7.0.14-20, **Secure Boot ON, booted fine** (signed kernel, shim → GRUB); but `/boot` is ext4 on LVM, GRUB cannot write its environment block there (`grub-reboot` warns so itself), `next_entry` was never cleared, and boot 2 with no command → **7.0.14-20 again**. A kernel lane needs a writable env block (the ESP) or a userspace "boot good" step — R-836. demo-hp left on 7.0.14-20, saved default 7.0.14-20, both kernels installed (`audits/os-host-lane-2026-10-04/partE/`). | ~~If the new kernel hangs, someone must switch the box off and on. The spike checks whether a hardware watchdog can do that (OPEN Q4).~~ **[FACT]** Only `softdog` runs (loaded by `watchdog-mux`); it cannot rescue a kernel that never boots. demo-hp has an AMD FCH whose `sp5100_tco` driver ships but is not loaded; untested. A hang still needs a person. **[FACT, 2026-10-04]** `sp5100_tco` is blacklisted by the Proxmox kernel package; loaded by hand it answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0 — read from sysfs, never opened, so never armed; unloaded). `kernel.panic = 0`: a panic leaves the host stopped (R-851). | ### 5.7 Telling people @@ -336,6 +336,30 @@ must never overlap a backup, a restore-test or a self-update.~~ whether the box restarted. Telling households in advance that the box may restart at night is a **promise to users**. That is the operator's decision when the slow lane is built. +### 5.8 The Docker engine slow lane — DESIGN (2026-10-04, nothing built) `[PROPOSAL]` + +Built from C5 (measured) and R-835. **The one decision it needs is in STATUS: `live-restore` on, fleet-wide.** + +- **What moves:** `docker-ce`, `docker-ce-cli`, `containerd.io`, `docker-buildx-plugin`, `docker-compose-plugin`, + `docker-ce-rootless-extras` in the customer guest — today the six "not covered" packages on every box. One + approved **engine set** at a time, like a controller floor: the operator approves it (slow lane, §5.2), after ring 0 + has run it for at least 2 nights healthy. Never two steps in one night. +- **Precondition: `live-restore` ON.** Without it an engine step restarts every container — 26.5–30 s of silence, + healthy at +42–45 s (C5). With it: 0 restarts, no gap, also across a containerd step. Turning it ON is safe + (`systemctl reload docker` applies it without a restart, C5); turning it OFF later by a plain restart stops every + container and starts none (R-835) — so it is turned on once, by the golden and by a one-time fleet step, and never + turned off by the lane. +- **Who and when:** the agent, through the same wrapper (`lane: slow`, refusal R3: only inside a verified signed + operator job, R-530's mechanism), in the guest, after the guest and host fast-lane steps, under the same heavy-op + gate, on a night the operator scheduled. Debian origin rule replaced by "origin `Docker CE`, exactly these names". +- **Health:** the guest rule (§8.1) plus `docker version` reports the approved engine, and every container running + at the start is running with the SAME container id (proof that `live-restore` held). A changed id is + `health_failed` even if the app is healthy — it means the households' apps restarted when they should not have. +- **Undo:** install the previous engine set (Docker's repository keeps 46 versions, C2) — by an operator job, with + `live-restore` still on, so the undo is also restart-free. +- **Not covered here:** the golden's own engine (baked weekly; a new golden carries the approved set), and BYO hosts + (the guest is ours on both, so the lane applies there too). + --- ## 6. Risks and edge cases @@ -405,8 +429,9 @@ Each step returns to the operator for go or no-go. 1. **Spike** (measure Q1–Q10; no product code). 2. **Guest Debian, fast lane.** ~~Lowest risk: a snapshot undo exists.~~ **BUILT 2026-10-04** — agent v0.140.0, hub v0.130.0, installer 1.29.0; §8.1. **There is no snapshot undo** (R-837, measured). -3. **Host Debian, fast lane** (no kernel, no Proxmox packages). -4. **Fleet view and alarms** (§5.7). +3. **Host Debian, fast lane** (no kernel, no Proxmox packages). **BUILT 2026-10-04** — agent v0.141.1, hub v0.131.1; + §8.2. +4. **Fleet view and alarms** (§5.7). **BUILT 2026-10-04** — hub v0.131.0/v0.131.1; §8.3. 5. **Slow lane: Docker engine.** 6. **Slow lane: host kernel and Proxmox packages, with the reboot.** 7. **Later:** the Proxmox major upgrade (PVE 9 → 10), drilled on ring 0 first. @@ -454,6 +479,47 @@ Evidence: `audits/os-guest-lane-2026-10-04/` (parts A–G). Brief: guest fast la - **Not delivered to existing boxes by the product**: the wrapper and the sudoers line reach a box only through the installer; the signed agent update replaces the binary only (R-840). The demo boxes got them BY HAND. +### 8.2 Step 3 as BUILT (2026-10-04) `[FACT]` + +Evidence: `audits/os-host-lane-2026-10-04/` (parts A–G). + +- **Where it runs:** appliances only. The proof is the ROOT-owned install record `/var/lib/felhom-install/state.json` + `mode: appliance` (written by the installer as root); the agent-writable `agent.json` `deployment_mode` is not + trusted for this. A BYO host gets no host step (wrapper refusal **R12**, now lifted only for lane fast / layer host + on an appliance). *Decided by CC unattended — operator may reverse.* +- **What:** origin `Debian` / `Debian-Security` only, and never a kernel, boot or firmware package (name pattern + `HOST_SLOW_RE` — `linux-*`, `proxmox-kernel*`, `pve-kernel*`, `pve-firmware`, `firmware-*`, `grub*`, `shim*`, + `systemd-boot*`, `*-microcode`, `efibootmgr`; refusal **R14**). The hub leaves the same names out of the host + candidate. +- **When:** in the same leg, after the guest step, under the same heavy-op gate. A failed or unhealthy guest step + skips the host step. +- **The host health rule** (`HostHealthVerdict`, pinned in `internal/osupdate`): `felhom-agent`, `pveproxy`, + `pvedaemon`, `pvestatd` and `pve-cluster` are `active`; the customer guest runs; the guest health rule (§8.1) + passes; the tunnel is `running` — an `unknown` tunnel does not fail it, `not_running` does. Same 5-minute wait. +- **Separate approved sets:** host and guest releases are separate (`os-host-…`, `os-guest-…`), each by the same rule + (24 h, 1 night of THAT layer, every ring-0 box). Ring 1 receives `host_release` beside `release`. +- **Reboot needed:** reported when PID 1 or `lxc-start` maps a replaced file, with the date of the first scanned + report that said so. **Never reboots.** The host is scanned on every pass, so a reboot clears it (v0.141.1 — v0.141.0 + hid `lxc-start` and never cleared, R-846). +- **Undo:** by hand, `runbooks/os-updates-host-undo.md` (proved). No automatic undo. +- **Speed (R-845):** one call per layer; both steps with nothing to install 23–32 s; a 108-package host pass 70 s. +- **Measured live:** demo-felhom ring 0 installed 108 Debian host packages, all Debian origin (checked against apt), + healthy; a 605-package host release approved (TEST wait 2 min / 0 nights, logged, reverted to 24 h + 1 night); + demo-felhom as ring 1 installed exactly the one version it lacked (604 already current). + +### 8.3 Step 4 as BUILT (2026-10-04) `[FACT]` + +- **The fleet view** (`GET /os/fleet`, operator): one line per box — ring, switch, the tunnel, and per layer: the + release, the last outcome, the last successful leg, pending, not covered, restart needed, reboot needed since, the + wrapper's own seconds. +- **The four alarms** (`08` §6.3), operator-only, hourly, at most weekly while true: no successful OS leg for + **7 days** while the switch is ON (naming the likely reason); reboot needed for **14 days**; ring 0 approved nothing + for **7 days** while it has pending fast-lane updates; not-covered fast-lane packages for **14 days**. The four + numbers are configuration (`OS_ALARM_*`). *Decided by CC unattended — operator may reverse:* 7 days = a week of + missed nights is past any normal hiccup (a box off for a weekend does not alarm); 14 days for reboot and coverage = + two weekly golden cycles, both need a person anyway. +- **The tunnel** (R-841): `running` / `not_running` / `unknown`; `tunnel_down` after two `not_running` reports. + ## 9. Where the rest lives - The finding: **R-812** (`backlog/OPEN-ITEMS.md`). The intention: **R-808** (`backlog/ROADMAP.md`). diff --git a/documentation/audits/os-host-lane-2026-10-04/partA/live/demo-hp-block-states.txt b/documentation/audits/os-host-lane-2026-10-04/partA/live/demo-hp-block-states.txt index b61b1ebf..aa4494ab 100644 --- a/documentation/audits/os-host-lane-2026-10-04/partA/live/demo-hp-block-states.txt +++ b/documentation/audits/os-host-lane-2026-10-04/partA/live/demo-hp-block-states.txt @@ -3,3 +3,17 @@ 11:36:47 container=running/unhealthy hub=running alarm=none 11:37:49 container=running/unhealthy hub=not_running alarm=none 11:38:51 container=running/unhealthy hub=not_running alarm=none +11:39:53 container=running/unhealthy hub=not_running alarm=none +11:40:55 container=running/unhealthy hub=not_running alarm=none +11:41:57 container=running/unhealthy hub=not_running alarm=none +11:42:59 container=running/unhealthy hub=not_running alarm=none +11:44:01 container=running/unhealthy hub=not_running alarm=none +11:45:03 container=running/unhealthy hub=not_running alarm=none +11:46:05 container=running/unhealthy hub=not_running alarm=none +11:47:07 container=running/unhealthy hub=not_running alarm=none +11:48:09 container=running/unhealthy hub=not_running alarm=none +11:49:10 container=running/unhealthy hub=not_running alarm=none +11:50:12 container=running/unhealthy hub=not_running alarm=none +11:51:14 container=running/unhealthy hub=not_running alarm=none +11:52:16 container=running/unhealthy hub=not_running alarm=none +11:53:18 container=running/unhealthy hub=not_running alarm=2026/10/04 13:52:29 [WARN] host demo-hp-bb76ea tunnel: tunnel_down (container running but the tunnel is NOT connected (cloudflared /ready fails)) diff --git a/documentation/audits/os-host-lane-2026-10-04/partA/live/demo-hp-block.txt b/documentation/audits/os-host-lane-2026-10-04/partA/live/demo-hp-block.txt index 746f7f0a..d5d7ab8a 100644 --- a/documentation/audits/os-host-lane-2026-10-04/partA/live/demo-hp-block.txt +++ b/documentation/audits/os-host-lane-2026-10-04/partA/live/demo-hp-block.txt @@ -3,3 +3,6 @@ Chain DOCKER-USER (1 references) target prot opt source destination DROP tcp -- 0.0.0.0/0 0.0.0.0/0 tcp dpt:7844 DROP udp -- 0.0.0.0/0 0.0.0.0/0 udp dpt:7844 +unblock at 2026-10-04T11:53:29Z +Chain DOCKER-USER (1 references) +target prot opt source destination diff --git a/documentation/audits/os-host-lane-2026-10-04/partA/live/demo-hp-recover-states.txt b/documentation/audits/os-host-lane-2026-10-04/partA/live/demo-hp-recover-states.txt new file mode 100644 index 00000000..6a6aec0a --- /dev/null +++ b/documentation/audits/os-host-lane-2026-10-04/partA/live/demo-hp-recover-states.txt @@ -0,0 +1,14 @@ +11:53:39 container=unhealthy hub=not_running +11:54:41 container=healthy hub=not_running +11:55:43 container=healthy hub=not_running +11:56:45 container=healthy hub=not_running +11:57:47 container=healthy hub=not_running +11:58:48 container=healthy hub=not_running +11:59:50 container=healthy hub=not_running +12:00:52 container=healthy hub=not_running +12:01:54 container=healthy hub=not_running +12:02:56 container=healthy hub=not_running +12:03:58 container=healthy hub=not_running +12:05:00 container=healthy hub=not_running +12:06:02 container=healthy hub=not_running +12:07:04 container=healthy hub=not_running diff --git a/documentation/audits/os-host-lane-2026-10-04/partA/live/hub-log-tunnel_down.txt b/documentation/audits/os-host-lane-2026-10-04/partA/live/hub-log-tunnel_down.txt new file mode 100644 index 00000000..0c7302d0 --- /dev/null +++ b/documentation/audits/os-host-lane-2026-10-04/partA/live/hub-log-tunnel_down.txt @@ -0,0 +1,5 @@ +2026/10/04 13:52:29 [WARN] host demo-hp-bb76ea tunnel: tunnel_down (container running but the tunnel is NOT connected (cloudflared /ready fails)) +2026/10/04 13:52:30 [INFO] Operator email sent for demo-hp/tunnel_down +2026/10/04 13:52:29 [WARN] host demo-hp-bb76ea tunnel: tunnel_down (container running but the tunnel is NOT connected (cloudflared /ready fails)) +2026/10/04 13:52:30 [INFO] Operator email sent for demo-hp/tunnel_down +2026/10/04 14:07:29 [WARN] host demo-hp-bb76ea tunnel: tunnel_recovered (connected) diff --git a/documentation/audits/os-host-lane-2026-10-04/partB/live/b2-demo-hp-ring0.log b/documentation/audits/os-host-lane-2026-10-04/partB/live/b2-demo-hp-ring0.log new file mode 100644 index 00000000..f5958a1d --- /dev/null +++ b/documentation/audits/os-host-lane-2026-10-04/partB/live/b2-demo-hp-ring0.log @@ -0,0 +1,135 @@ +=== felhom-agent 0.141.1 selftest=os-update vmid=9201 ring=0 enabled=true guest-release=false host-release=false appliance=true === +time=2026-10-04T14:22:52.476+02:00 level=INFO msg="osupdate: START" run=20261004T122252Z layer=guest vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T122252Z +time=2026-10-04T14:23:08.079+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T122252Z layer=guest:9201 lane=fast mode=apply select=pending-fast packages=0" +time=2026-10-04T14:23:08.079+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0" +time=2026-10-04T14:23:08.079+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0" +time=2026-10-04T14:23:08.079+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)" +time=2026-10-04T14:23:08.080+02:00 level=INFO msg="osupdate: DONE" run=20261004T122252Z layer=guest vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=6 not_covered=6 restart_needed=0 reboot_needed=false wrapper_seconds=15.5 +time=2026-10-04T14:23:08.089+02:00 level=INFO msg="osupdate: START" run=20261004T122252Z layer=host vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T122252Z +time=2026-10-04T14:23:22.788+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T122252Z layer=host lane=fast mode=apply select=pending-fast packages=0" +time=2026-10-04T14:23:22.788+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0" +time=2026-10-04T14:23:22.788+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0" +time=2026-10-04T14:23:22.788+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)" +time=2026-10-04T14:23:23.852+02:00 level=INFO msg="osupdate: DONE" run=20261004T122252Z layer=host vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=78 not_covered=78 restart_needed=0 reboot_needed=false wrapper_seconds=14.6 + --- os-update report (guest) --- + { + "health_reason": "", + "healthy": true, + "mode": "apply", + "not_covered": [ + "docker-ce-cli", + "containerd.io", + "docker-ce", + "docker-buildx-plugin", + "docker-ce-rootless-extras", + "docker-compose-plugin" + ], + "outcome": "nothing", + "pending": 6, + "reboot_needed": false, + "refused": null, + "release_id": "ring0-20261004T122252Z", + "restart_needed": null, + "ring": 0, + "run_id": "20261004T122252Z", + "upgraded": [], + "wrapper_seconds": 15.5 + } + --- os-update report (host) --- + { + "health_reason": "", + "healthy": true, + "mode": "apply", + "not_covered": [ + "frr", + "shim-signed-common", + "proxmox-secure-boot-support", + "shim-unsigned", + "shim-helpers-amd64-signed", + "shim-signed", + "amd64-microcode", + "libradosstriper1", + "librgw2", + "ceph-common", + "librbd1", + "librados2", + "python3-cephfs", + "libcephfs2", + "python3-rgw", + "python3-rados", + "python3-ceph-argparse", + "python3-ceph-common", + "python3-rbd", + "ceph-fuse", + "chrony", + "libcorosync-common4", + "libcfg7", + "libcmap4", + "libcpg4", + "libknet1t64", + "libnozzle1t64", + "libquorum5", + "libvotequorum8", + "corosync", + "frr-pythontools", + "libjs-extjs", + "libnvpair3linux", + "libproxmox-acme-plugins", + "libproxmox-backup-qemu0", + "pve-qemu-kvm", + "libpve-notify-perl", + "libpve-cluster-api-perl", + "libpve-cluster-perl", + "pve-cluster", + "libpve-access-control", + "libpve-apiclient-perl", + "librados2-perl", + "proxmox-backup-client", + "proxmox-backup-file-restore", + "pve-manager", + "libproxmox-acme-perl", + "libpve-common-perl", + "libpve-guest-common-perl", + "qemu-server", + "libpve-storage-perl", + "pve-edk2-firmware-legacy", + "pve-edk2-firmware-ovmf", + "libpve-network-api-perl", + "libpve-network-perl", + "proxmox-firewall-data", + "pve-firewall", + "pve-container", + "pve-ha-manager", + "novnc-pve", + "proxmox-enterprise-support-keyring", + "proxmox-mini-journalreader", + "proxmox-widget-toolkit", + "pve-docs", + "pve-i18n", + "pve-xtermjs", + "pve-yew-mobile-i18n", + "pve-yew-mobile-gui", + "libuutil3linux", + "libzfs7linux", + "libzpool7linux", + "proxmox-kernel-helper", + "pve-edk2-firmware-aarch64", + "pve-edk2-firmware", + "pve-firmware", + "zfs-initramfs", + "zfsutils-linux", + "zfs-zed" + ], + "outcome": "nothing", + "pending": 78, + "reboot_needed": false, + "refused": null, + "release_id": "ring0-20261004T122252Z", + "restart_needed": [], + "ring": 0, + "run_id": "20261004T122252Z", + "upgraded": [], + "wrapper_seconds": 14.6 + } + pass took 31.4s +WALL_SECONDS=31.520457543 diff --git a/documentation/audits/os-host-lane-2026-10-04/partB/ring1/r0-set-ring1.txt b/documentation/audits/os-host-lane-2026-10-04/partB/ring1/r0-set-ring1.txt new file mode 100644 index 00000000..254cd352 --- /dev/null +++ b/documentation/audits/os-host-lane-2026-10-04/partB/ring1/r0-set-ring1.txt @@ -0,0 +1,2 @@ +{"ok":true} + 200 diff --git a/documentation/audits/os-host-lane-2026-10-04/partB/ring1/r1-demo-felhom-tzdata-back.txt b/documentation/audits/os-host-lane-2026-10-04/partB/ring1/r1-demo-felhom-tzdata-back.txt new file mode 100644 index 00000000..e121160b --- /dev/null +++ b/documentation/audits/os-host-lane-2026-10-04/partB/ring1/r1-demo-felhom-tzdata-back.txt @@ -0,0 +1,43 @@ ++ PKG=tzdata +++ grep -oE 'tzdata:amd64 \([^,]+' /var/log/apt/history.log +++ tail -1 +++ sed 's/.*(//' ++ OLD=2026b-0+deb13u1 +++ dpkg-query -W '-f=${Version}' tzdata ++ NEW=2026c-0+deb13u1 ++ echo OLD=2026b-0+deb13u1 NEW=2026c-0+deb13u1 +OLD=2026b-0+deb13u1 NEW=2026c-0+deb13u1 +++ curl -s 'https://snapshot.debian.org/mr/binary/tzdata/2026c-0+deb13u1/binfiles?fileinfo=1' +++ python3 -c ' +import json,sys; d=json.load(sys.stdin) +print(sorted(f["first_seen"] for v in d["fileinfo"].values() for f in v)[0])' ++ TS=20260831T204404Z ++ . /etc/os-release +++ PRETTY_NAME='Debian GNU/Linux 13 (trixie)' +++ NAME='Debian GNU/Linux' +++ VERSION_ID=13 +++ VERSION='13 (trixie)' +++ VERSION_CODENAME=trixie +++ DEBIAN_VERSION_FULL=13.7 +++ ID=debian +++ HOME_URL=https://www.debian.org/ +++ SUPPORT_URL=https://www.debian.org/support +++ BUG_REPORT_URL=https://bugs.debian.org/ ++ printf 'deb [check-valid-until=no] http://snapshot.debian.org/archive/debian/%s %s main\ndeb [check-valid-until=no] http://snapshot.debian.org/archive/debian-security/%s %s-security main\n' 20260831T204404Z trixie 20260831T204404Z trixie ++ apt-get -q update ++ apt-get -s install --allow-downgrades tzdata=2026b-0+deb13u1 ++ grep -E '^(Inst|Remv)|downgraded' +0 upgraded, 0 newly installed, 1 downgraded, 0 to remove and 78 not upgraded. +Inst tzdata [2026c-0+deb13u1] (2026b-0+deb13u1 Debian:13.6/stable [all]) ++ DEBIAN_FRONTEND=noninteractive ++ apt-get -y -q install --allow-downgrades tzdata=2026b-0+deb13u1 ++ rm -f /etc/apt/sources.list.d/felhom-undo-snapshot.list ++ apt-get -q update ++ dpkg-query -W tzdata +tzdata 2026b-0+deb13u1 ++ ls /etc/apt/sources.list.d/ +ceph.sources +debian.sources +pve-enterprise.sources +pve-no-subscription.sources +tailscale.list diff --git a/documentation/audits/os-host-lane-2026-10-04/partB/ring1/r2-demo-felhom-ring1-leg.log b/documentation/audits/os-host-lane-2026-10-04/partB/ring1/r2-demo-felhom-ring1-leg.log new file mode 100644 index 00000000..09da7c55 --- /dev/null +++ b/documentation/audits/os-host-lane-2026-10-04/partB/ring1/r2-demo-felhom-ring1-leg.log @@ -0,0 +1,178 @@ +=== felhom-agent 0.141.1 selftest=os-update vmid=9201 ring=1 enabled=true guest-release=true host-release=true appliance=true === +time=2026-10-04T14:40:46.795+02:00 level=INFO msg="osupdate: START" run=20261004T124046Z layer=guest vmid=9201 ring=1 trigger=debug enabled=true release=os-guest-20261004-123933 +time=2026-10-04T14:40:56.895+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=os-guest-20261004-123933 layer=guest:9201 lane=fast mode=apply select=listed packages=272" +time=2026-10-04T14:40:56.895+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0" +time=2026-10-04T14:40:56.895+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=272 not-installed=0 from-snapshot=0" +time=2026-10-04T14:40:56.895+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)" +time=2026-10-04T14:40:56.896+02:00 level=INFO msg="osupdate: DONE" run=20261004T124046Z layer=guest vmid=9201 ring=1 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=6 not_covered=6 restart_needed=0 reboot_needed=false wrapper_seconds=10.1 +time=2026-10-04T14:40:56.907+02:00 level=INFO msg="osupdate: START" run=20261004T124046Z layer=host vmid=9201 ring=1 trigger=debug enabled=true release=os-host-20261004-124034 +time=2026-10-04T14:41:10.034+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=os-host-20261004-124034 layer=host lane=fast mode=apply select=listed packages=605" +time=2026-10-04T14:41:10.034+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0" +time=2026-10-04T14:41:10.034+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=1 already=604 not-installed=0 from-snapshot=0" +time=2026-10-04T14:41:10.034+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=1.4 upgraded=1 restart-needed=agetty,blkmapd,chronyd,cron,dbus-daemon,dmeventd,ksmtuned,lxc-monitord,lxc-start,lxcfs,pmxcfs,proxmox-firewal,pve-firewall,pve-ha-crm,pve-ha-lrm,pve-lxc-syscall,pvedaemon,pvedaemon worke,pvefw-logger,pveproxy,pveproxy worker,pvescheduler,pvestatd,qmeventd,rpcbind,rrdcached,smartd,spiceproxy,spiceproxy work,sshd,systemd-logind,systemd-udevd,watchdog-mux,zed reboot-needed=yes" +time=2026-10-04T14:41:10.815+02:00 level=INFO msg="osupdate: DONE" run=20261004T124046Z layer=host vmid=9201 ring=1 trigger=debug outcome=applied healthy=true reason="" upgraded=1 pending=80 not_covered=80 restart_needed=34 reboot_needed=true wrapper_seconds=13.1 + --- os-update report (guest) --- + { + "health_reason": "", + "healthy": true, + "mode": "apply", + "not_covered": [ + "docker-ce-cli", + "containerd.io", + "docker-ce", + "docker-buildx-plugin", + "docker-ce-rootless-extras", + "docker-compose-plugin" + ], + "outcome": "nothing", + "pending": 6, + "reboot_needed": false, + "refused": null, + "release_id": "os-guest-20261004-123933", + "restart_needed": null, + "ring": 1, + "run_id": "20261004T124046Z", + "upgraded": [], + "wrapper_seconds": 10.1 + } + --- os-update report (host) --- + { + "health_reason": "", + "healthy": true, + "mode": "apply", + "not_covered": [ + "frr", + "shim-signed-common", + "shim-unsigned", + "shim-helpers-amd64-signed", + "shim-signed", + "libradosstriper1", + "librgw2", + "ceph-common", + "librbd1", + "librados2", + "python3-cephfs", + "libcephfs2", + "python3-rgw", + "python3-rados", + "python3-ceph-argparse", + "python3-ceph-common", + "python3-rbd", + "ceph-fuse", + "chrony", + "libcorosync-common4", + "libcfg7", + "libcmap4", + "libcpg4", + "libknet1t64", + "libnozzle1t64", + "libquorum5", + "libvotequorum8", + "corosync", + "frr-pythontools", + "libjs-extjs", + "libnvpair3linux", + "libproxmox-acme-plugins", + "libproxmox-backup-qemu0", + "pve-qemu-kvm", + "libpve-notify-perl", + "libpve-cluster-api-perl", + "libpve-cluster-perl", + "pve-cluster", + "libpve-access-control", + "libpve-apiclient-perl", + "librados2-perl", + "proxmox-backup-client", + "proxmox-backup-file-restore", + "pve-manager", + "libproxmox-acme-perl", + "libpve-common-perl", + "libpve-guest-common-perl", + "qemu-server", + "libpve-storage-perl", + "pve-edk2-firmware-legacy", + "pve-edk2-firmware-ovmf", + "libpve-network-api-perl", + "libpve-network-perl", + "proxmox-firewall-data", + "pve-firewall", + "pve-container", + "pve-ha-manager", + "novnc-pve", + "proxmox-enterprise-support-keyring", + "proxmox-mini-journalreader", + "proxmox-widget-toolkit", + "pve-docs", + "pve-i18n", + "pve-xtermjs", + "pve-yew-mobile-i18n", + "pve-yew-mobile-gui", + "libuutil3linux", + "libzfs7linux", + "libzpool7linux", + "proxmox-first-boot", + "pve-firmware", + "proxmox-kernel-7.0.14-20-pve-signed", + "proxmox-kernel-7.0", + "proxmox-kernel-helper", + "pve-edk2-firmware-aarch64", + "pve-edk2-firmware", + "zfs-initramfs", + "zfsutils-linux", + "zfs-zed", + "tailscale" + ], + "outcome": "applied", + "pending": 80, + "reboot_needed": true, + "refused": null, + "release_id": "os-host-20261004-124034", + "restart_needed": [ + "agetty", + "blkmapd", + "chronyd", + "cron", + "dbus-daemon", + "dmeventd", + "ksmtuned", + "lxc-monitord", + "lxc-start", + "lxcfs", + "pmxcfs", + "proxmox-firewal", + "pve-firewall", + "pve-ha-crm", + "pve-ha-lrm", + "pve-lxc-syscall", + "pvedaemon", + "pvedaemon worke", + "pvefw-logger", + "pveproxy", + "pveproxy worker", + "pvescheduler", + "pvestatd", + "qmeventd", + "rpcbind", + "rrdcached", + "smartd", + "spiceproxy", + "spiceproxy work", + "sshd", + "systemd-logind", + "systemd-udevd", + "watchdog-mux", + "zed" + ], + "ring": 1, + "run_id": "20261004T124046Z", + "upgraded": [ + { + "name": "tzdata", + "version": "2026c-0+deb13u1", + "origin": "" + } + ], + "wrapper_seconds": 13.1 + } + pass took 24.1s +tzdata 2026c-0+deb13u1 diff --git a/documentation/audits/os-host-lane-2026-10-04/partB/undo/u1-find.txt b/documentation/audits/os-host-lane-2026-10-04/partB/undo/u1-find.txt new file mode 100644 index 00000000..535aad05 --- /dev/null +++ b/documentation/audits/os-host-lane-2026-10-04/partB/undo/u1-find.txt @@ -0,0 +1,11 @@ ++ PKG=eject ++ OLD=2.41-5 ++ dpkg -s eject ++ grep -E '^(Status|Version)' +Status: install ok installed +Version: 2.41.5-0+deb13u1 ++ curl -s 'https://snapshot.debian.org/mr/binary/eject/2.41-5/binfiles?fileinfo=1' ++ python3 -c ' +import json,sys; d=json.load(sys.stdin) +print(sorted(f["first_seen"] for v in d["fileinfo"].values() for f in v)[0])' +20250510T015204Z diff --git a/documentation/audits/os-host-lane-2026-10-04/partB/undo/u2-simulate.txt b/documentation/audits/os-host-lane-2026-10-04/partB/undo/u2-simulate.txt new file mode 100644 index 00000000..712493d0 --- /dev/null +++ b/documentation/audits/os-host-lane-2026-10-04/partB/undo/u2-simulate.txt @@ -0,0 +1,28 @@ ++ PKG=eject ++ OLD=2.41-5 ++ TS=20250510T015204Z ++ . /etc/os-release +++ PRETTY_NAME='Debian GNU/Linux 13 (trixie)' +++ NAME='Debian GNU/Linux' +++ VERSION_ID=13 +++ VERSION='13 (trixie)' +++ VERSION_CODENAME=trixie +++ DEBIAN_VERSION_FULL=13.7 +++ ID=debian +++ HOME_URL=https://www.debian.org/ +++ SUPPORT_URL=https://www.debian.org/support +++ BUG_REPORT_URL=https://bugs.debian.org/ ++ cat +++ date +%s ++ S=1791116624 ++ apt-get -q update ++ tail -3 +Get:7 http://snapshot.debian.org/archive/debian/20250510T015204Z trixie/main amd64 Packages [9681 kB] +Fetched 9904 kB in 5s (2165 kB/s) +Reading package lists... +++ date +%s +update_seconds=6 ++ echo update_seconds=6 ++ apt-get -s install --allow-downgrades eject=2.41-5 ++ grep -E '^(Inst|Remv|Conf)|downgraded|newly|Err|E:' +E: Version '2.41-5' for 'eject' was not found diff --git a/documentation/audits/os-host-lane-2026-10-04/partB/undo/u3-simulate-newts.txt b/documentation/audits/os-host-lane-2026-10-04/partB/undo/u3-simulate-newts.txt new file mode 100644 index 00000000..b73c4290 --- /dev/null +++ b/documentation/audits/os-host-lane-2026-10-04/partB/undo/u3-simulate-newts.txt @@ -0,0 +1,35 @@ ++ PKG=eject ++ OLD=2.41-5 +++ dpkg-query -W '-f=${Version}' eject ++ NEW=2.41.5-0+deb13u1 +++ curl -s 'https://snapshot.debian.org/mr/binary/eject/2.41.5-0+deb13u1/binfiles?fileinfo=1' +++ python3 -c ' +import json,sys; d=json.load(sys.stdin) +print(sorted(f["first_seen"] for v in d["fileinfo"].values() for f in v)[0])' +TS=20260814T165831Z ++ TS=20260814T165831Z ++ echo TS=20260814T165831Z ++ . /etc/os-release +++ PRETTY_NAME='Debian GNU/Linux 13 (trixie)' +++ NAME='Debian GNU/Linux' +++ VERSION_ID=13 +++ VERSION='13 (trixie)' +++ VERSION_CODENAME=trixie +++ DEBIAN_VERSION_FULL=13.7 +++ ID=debian +++ HOME_URL=https://www.debian.org/ +++ SUPPORT_URL=https://www.debian.org/support +++ BUG_REPORT_URL=https://bugs.debian.org/ ++ cat ++ apt-get -q update ++ tail -1 +Reading package lists... ++ apt-cache madison eject + eject | 2.41.5-0+deb13u1 | http://deb.debian.org/debian trixie/main amd64 Packages + eject | 2.41.5-0+deb13u1 | http://security.debian.org/debian-security trixie-security/main amd64 Packages + eject | 2.41.5-0+deb13u1 | http://snapshot.debian.org/archive/debian-security/20260814T165831Z trixie-security/main amd64 Packages + eject | 2.41-5 | http://snapshot.debian.org/archive/debian/20260814T165831Z trixie/main amd64 Packages ++ apt-get -s install --allow-downgrades eject=2.41-5 ++ grep -E '^(Inst|Remv|Conf)|downgraded|newly|Err|E:' +E: Unable to correct problems, you have held broken packages. +E: The following information from --solver 3.0 may provide additional context: diff --git a/documentation/audits/os-host-lane-2026-10-04/partB/undo/u4-tzdata-simulate.txt b/documentation/audits/os-host-lane-2026-10-04/partB/undo/u4-tzdata-simulate.txt new file mode 100644 index 00000000..0bdcca86 --- /dev/null +++ b/documentation/audits/os-host-lane-2026-10-04/partB/undo/u4-tzdata-simulate.txt @@ -0,0 +1,37 @@ ++ rm -f /etc/apt/sources.list.d/felhom-undo-snapshot.list ++ PKG=tzdata ++ OLD=2026b-0+deb13u1 +++ dpkg-query -W '-f=${Version}' tzdata ++ NEW=2026c-0+deb13u1 ++ echo NEW=2026c-0+deb13u1 +NEW=2026c-0+deb13u1 +++ curl -s 'https://snapshot.debian.org/mr/binary/tzdata/2026c-0+deb13u1/binfiles?fileinfo=1' +++ python3 -c ' +import json,sys; d=json.load(sys.stdin) +print(sorted(f["first_seen"] for v in d["fileinfo"].values() for f in v)[0])' +TS=20260831T204404Z ++ TS=20260831T204404Z ++ echo TS=20260831T204404Z ++ . /etc/os-release +++ PRETTY_NAME='Debian GNU/Linux 13 (trixie)' +++ NAME='Debian GNU/Linux' +++ VERSION_ID=13 +++ VERSION='13 (trixie)' +++ VERSION_CODENAME=trixie +++ DEBIAN_VERSION_FULL=13.7 +++ ID=debian +++ HOME_URL=https://www.debian.org/ +++ SUPPORT_URL=https://www.debian.org/support +++ BUG_REPORT_URL=https://bugs.debian.org/ ++ cat ++ apt-get -q update ++ tail -1 +Reading package lists... ++ apt-cache madison tzdata + tzdata | 2026c-0+deb13u1 | http://deb.debian.org/debian trixie/main amd64 Packages + tzdata | 2026b-0+deb13u1 | http://snapshot.debian.org/archive/debian/20260831T204404Z trixie/main amd64 Packages ++ apt-get -s install --allow-downgrades tzdata=2026b-0+deb13u1 ++ grep -E '^(Inst|Remv|Conf)|downgraded|newly|Err|E:' +0 upgraded, 0 newly installed, 1 downgraded, 0 to remove and 77 not upgraded. +Inst tzdata [2026c-0+deb13u1] (2026b-0+deb13u1 Debian:13.6/stable [all]) +Conf tzdata (2026b-0+deb13u1 Debian:13.6/stable [all]) diff --git a/documentation/audits/os-host-lane-2026-10-04/partB/undo/u5-tzdata-undo.txt b/documentation/audits/os-host-lane-2026-10-04/partB/undo/u5-tzdata-undo.txt new file mode 100644 index 00000000..f826433c --- /dev/null +++ b/documentation/audits/os-host-lane-2026-10-04/partB/undo/u5-tzdata-undo.txt @@ -0,0 +1,38 @@ ++ PKG=tzdata ++ OLD=2026b-0+deb13u1 +++ date +%s ++ S=1791116670 ++ DEBIAN_FRONTEND=noninteractive ++ apt-get -y -q install --allow-downgrades tzdata=2026b-0+deb13u1 ++ grep -E '^(Unpacking|Setting up)|downgraded' +0 upgraded, 0 newly installed, 1 downgraded, 0 to remove and 77 not upgraded. +Unpacking tzdata (2026b-0+deb13u1) over (2026c-0+deb13u1) ... +Setting up tzdata (2026b-0+deb13u1) ... ++ apt-mark hold tzdata +tzdata set on hold. ++ rm -f /etc/apt/sources.list.d/felhom-undo-snapshot.list ++ apt-get -q update ++ tail -1 +Reading package lists... +++ date +%s +undo_seconds=4 ++ echo undo_seconds=4 ++ dpkg -s tzdata ++ grep -E '^(Status|Version)' +Status: hold ok installed +Version: 2026b-0+deb13u1 ++ apt-mark showhold +tzdata ++ systemctl is-active pveproxy pvedaemon pvestatd pve-cluster felhom-agent +active +active +active +active +active ++ pct status 9201 +status: running ++ ls /etc/apt/sources.list.d/ +ceph.sources +debian.sources +pve-enterprise.sources +pve-no-subscription.sources diff --git a/documentation/audits/os-host-lane-2026-10-04/partB/undo/u6-leg-with-hold.log b/documentation/audits/os-host-lane-2026-10-04/partB/undo/u6-leg-with-hold.log new file mode 100644 index 00000000..692cd04a --- /dev/null +++ b/documentation/audits/os-host-lane-2026-10-04/partB/undo/u6-leg-with-hold.log @@ -0,0 +1,135 @@ +=== felhom-agent 0.141.1 selftest=os-update vmid=9201 ring=0 enabled=true guest-release=false host-release=false appliance=true === +time=2026-10-04T14:24:41.910+02:00 level=INFO msg="osupdate: START" run=20261004T122441Z layer=guest vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T122441Z +time=2026-10-04T14:24:57.242+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T122441Z layer=guest:9201 lane=fast mode=apply select=pending-fast packages=0" +time=2026-10-04T14:24:57.242+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0" +time=2026-10-04T14:24:57.242+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0" +time=2026-10-04T14:24:57.242+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)" +time=2026-10-04T14:24:57.243+02:00 level=INFO msg="osupdate: DONE" run=20261004T122441Z layer=guest vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=6 not_covered=6 restart_needed=0 reboot_needed=false wrapper_seconds=15.3 +time=2026-10-04T14:24:57.251+02:00 level=INFO msg="osupdate: START" run=20261004T122441Z layer=host vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T122441Z +time=2026-10-04T14:25:12.127+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T122441Z layer=host lane=fast mode=apply select=pending-fast packages=0" +time=2026-10-04T14:25:12.127+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0" +time=2026-10-04T14:25:12.127+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0" +time=2026-10-04T14:25:12.127+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)" +time=2026-10-04T14:25:13.136+02:00 level=INFO msg="osupdate: DONE" run=20261004T122441Z layer=host vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=78 not_covered=78 restart_needed=0 reboot_needed=false wrapper_seconds=14.8 + --- os-update report (guest) --- + { + "health_reason": "", + "healthy": true, + "mode": "apply", + "not_covered": [ + "docker-ce-cli", + "containerd.io", + "docker-ce", + "docker-buildx-plugin", + "docker-ce-rootless-extras", + "docker-compose-plugin" + ], + "outcome": "nothing", + "pending": 6, + "reboot_needed": false, + "refused": null, + "release_id": "ring0-20261004T122441Z", + "restart_needed": null, + "ring": 0, + "run_id": "20261004T122441Z", + "upgraded": [], + "wrapper_seconds": 15.3 + } + --- os-update report (host) --- + { + "health_reason": "", + "healthy": true, + "mode": "apply", + "not_covered": [ + "frr", + "shim-signed-common", + "proxmox-secure-boot-support", + "shim-unsigned", + "shim-helpers-amd64-signed", + "shim-signed", + "amd64-microcode", + "libradosstriper1", + "librgw2", + "ceph-common", + "librbd1", + "librados2", + "python3-cephfs", + "libcephfs2", + "python3-rgw", + "python3-rados", + "python3-ceph-argparse", + "python3-ceph-common", + "python3-rbd", + "ceph-fuse", + "chrony", + "libcorosync-common4", + "libcfg7", + "libcmap4", + "libcpg4", + "libknet1t64", + "libnozzle1t64", + "libquorum5", + "libvotequorum8", + "corosync", + "frr-pythontools", + "libjs-extjs", + "libnvpair3linux", + "libproxmox-acme-plugins", + "libproxmox-backup-qemu0", + "pve-qemu-kvm", + "libpve-notify-perl", + "libpve-cluster-api-perl", + "libpve-cluster-perl", + "pve-cluster", + "libpve-access-control", + "libpve-apiclient-perl", + "librados2-perl", + "proxmox-backup-client", + "proxmox-backup-file-restore", + "pve-manager", + "libproxmox-acme-perl", + "libpve-common-perl", + "libpve-guest-common-perl", + "qemu-server", + "libpve-storage-perl", + "pve-edk2-firmware-legacy", + "pve-edk2-firmware-ovmf", + "libpve-network-api-perl", + "libpve-network-perl", + "proxmox-firewall-data", + "pve-firewall", + "pve-container", + "pve-ha-manager", + "novnc-pve", + "proxmox-enterprise-support-keyring", + "proxmox-mini-journalreader", + "proxmox-widget-toolkit", + "pve-docs", + "pve-i18n", + "pve-xtermjs", + "pve-yew-mobile-i18n", + "pve-yew-mobile-gui", + "libuutil3linux", + "libzfs7linux", + "libzpool7linux", + "proxmox-kernel-helper", + "pve-edk2-firmware-aarch64", + "pve-edk2-firmware", + "pve-firmware", + "zfs-initramfs", + "zfsutils-linux", + "zfs-zed" + ], + "outcome": "nothing", + "pending": 78, + "reboot_needed": false, + "refused": null, + "release_id": "ring0-20261004T122441Z", + "restart_needed": [], + "ring": 0, + "run_id": "20261004T122441Z", + "upgraded": [], + "wrapper_seconds": 14.8 + } + pass took 31.2s +tzdata 2026b-0+deb13u1 diff --git a/documentation/audits/os-host-lane-2026-10-04/partB/undo/u7-unhold-leg.log b/documentation/audits/os-host-lane-2026-10-04/partB/undo/u7-unhold-leg.log new file mode 100644 index 00000000..8d945f85 --- /dev/null +++ b/documentation/audits/os-host-lane-2026-10-04/partB/undo/u7-unhold-leg.log @@ -0,0 +1,143 @@ +Canceled hold on tzdata. +=== felhom-agent 0.141.1 selftest=os-update vmid=9201 ring=0 enabled=true guest-release=false host-release=false appliance=true === +time=2026-10-04T14:25:22.024+02:00 level=INFO msg="osupdate: START" run=20261004T122522Z layer=guest vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T122522Z +time=2026-10-04T14:25:37.515+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T122522Z layer=guest:9201 lane=fast mode=apply select=pending-fast packages=0" +time=2026-10-04T14:25:37.515+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0" +time=2026-10-04T14:25:37.515+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0" +time=2026-10-04T14:25:37.515+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)" +time=2026-10-04T14:25:37.516+02:00 level=INFO msg="osupdate: DONE" run=20261004T122522Z layer=guest vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=6 not_covered=6 restart_needed=0 reboot_needed=false wrapper_seconds=15.4 +time=2026-10-04T14:25:37.526+02:00 level=INFO msg="osupdate: START" run=20261004T122522Z layer=host vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T122522Z +time=2026-10-04T14:25:57.778+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T122522Z layer=host lane=fast mode=apply select=pending-fast packages=0" +time=2026-10-04T14:25:57.778+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0" +time=2026-10-04T14:25:57.778+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=1 already=0 not-installed=0 from-snapshot=0" +time=2026-10-04T14:25:57.778+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=2.0 upgraded=1 restart-needed=- reboot-needed=no" +time=2026-10-04T14:25:58.809+02:00 level=INFO msg="osupdate: DONE" run=20261004T122522Z layer=host vmid=9201 ring=0 trigger=debug outcome=applied healthy=true reason="" upgraded=1 pending=78 not_covered=78 restart_needed=0 reboot_needed=false wrapper_seconds=20.2 + --- os-update report (guest) --- + { + "health_reason": "", + "healthy": true, + "mode": "apply", + "not_covered": [ + "docker-ce-cli", + "containerd.io", + "docker-ce", + "docker-buildx-plugin", + "docker-ce-rootless-extras", + "docker-compose-plugin" + ], + "outcome": "nothing", + "pending": 6, + "reboot_needed": false, + "refused": null, + "release_id": "ring0-20261004T122522Z", + "restart_needed": null, + "ring": 0, + "run_id": "20261004T122522Z", + "upgraded": [], + "wrapper_seconds": 15.4 + } + --- os-update report (host) --- + { + "health_reason": "", + "healthy": true, + "mode": "apply", + "not_covered": [ + "frr", + "shim-signed-common", + "proxmox-secure-boot-support", + "shim-unsigned", + "shim-helpers-amd64-signed", + "shim-signed", + "amd64-microcode", + "libradosstriper1", + "librgw2", + "ceph-common", + "librbd1", + "librados2", + "python3-cephfs", + "libcephfs2", + "python3-rgw", + "python3-rados", + "python3-ceph-argparse", + "python3-ceph-common", + "python3-rbd", + "ceph-fuse", + "chrony", + "libcorosync-common4", + "libcfg7", + "libcmap4", + "libcpg4", + "libknet1t64", + "libnozzle1t64", + "libquorum5", + "libvotequorum8", + "corosync", + "frr-pythontools", + "libjs-extjs", + "libnvpair3linux", + "libproxmox-acme-plugins", + "libproxmox-backup-qemu0", + "pve-qemu-kvm", + "libpve-notify-perl", + "libpve-cluster-api-perl", + "libpve-cluster-perl", + "pve-cluster", + "libpve-access-control", + "libpve-apiclient-perl", + "librados2-perl", + "proxmox-backup-client", + "proxmox-backup-file-restore", + "pve-manager", + "libproxmox-acme-perl", + "libpve-common-perl", + "libpve-guest-common-perl", + "qemu-server", + "libpve-storage-perl", + "pve-edk2-firmware-legacy", + "pve-edk2-firmware-ovmf", + "libpve-network-api-perl", + "libpve-network-perl", + "proxmox-firewall-data", + "pve-firewall", + "pve-container", + "pve-ha-manager", + "novnc-pve", + "proxmox-enterprise-support-keyring", + "proxmox-mini-journalreader", + "proxmox-widget-toolkit", + "pve-docs", + "pve-i18n", + "pve-xtermjs", + "pve-yew-mobile-i18n", + "pve-yew-mobile-gui", + "libuutil3linux", + "libzfs7linux", + "libzpool7linux", + "proxmox-kernel-helper", + "pve-edk2-firmware-aarch64", + "pve-edk2-firmware", + "pve-firmware", + "zfs-initramfs", + "zfsutils-linux", + "zfs-zed" + ], + "outcome": "applied", + "pending": 78, + "reboot_needed": false, + "refused": null, + "release_id": "ring0-20261004T122522Z", + "restart_needed": [], + "ring": 0, + "run_id": "20261004T122522Z", + "upgraded": [ + { + "name": "tzdata", + "version": "2026c-0+deb13u1", + "origin": "" + } + ], + "wrapper_seconds": 20.2 + } + pass took 36.9s +tzdata 2026c-0+deb13u1 +0 diff --git a/documentation/audits/os-host-lane-2026-10-04/partC/live/fleet-during-ring1-test.json b/documentation/audits/os-host-lane-2026-10-04/partC/live/fleet-during-ring1-test.json new file mode 100644 index 00000000..720f635a --- /dev/null +++ b/documentation/audits/os-host-lane-2026-10-04/partC/live/fleet-during-ring1-test.json @@ -0,0 +1,104 @@ +{ + "boxes": [ + { + "HostID": "demo-felhom-8363b5", + "Ring": 1, + "Enabled": true, + "Tunnel": "running", + "Guest": { + "ReleaseID": "os-guest-20261004-123933", + "LastOutcome": "nothing", + "LastAt": "2026-10-04T12:40:56Z", + "LastSuccessfulLeg": "2026-10-04T12:40:56Z", + "Pending": 6, + "NotCovered": 6, + "NotCoveredFast": 0, + "RestartNeeded": 0, + "RebootNeededSince": "2026-10-04T09:20:04Z", + "WrapperPassSeconds": 10.1 + }, + "Host": { + "ReleaseID": "os-host-20261004-124034", + "LastOutcome": "applied", + "LastAt": "2026-10-04T12:41:10Z", + "LastSuccessfulLeg": "2026-10-04T12:41:10Z", + "Pending": 80, + "NotCovered": 80, + "NotCoveredFast": 0, + "RestartNeeded": 34, + "RebootNeededSince": "2026-10-04T11:53:20Z", + "WrapperPassSeconds": 13.1 + } + }, + { + "HostID": "demo-hp-bb76ea", + "Ring": 0, + "Enabled": true, + "Tunnel": "unknown", + "Guest": { + "ReleaseID": "ring0-20261004T122522Z", + "LastOutcome": "nothing", + "LastAt": "2026-10-04T12:25:37Z", + "LastSuccessfulLeg": "2026-10-04T12:25:37Z", + "Pending": 6, + "NotCovered": 6, + "NotCoveredFast": 0, + "RestartNeeded": 0, + "RebootNeededSince": "2026-10-04T09:21:28Z", + "WrapperPassSeconds": 15.4 + }, + "Host": { + "ReleaseID": "ring0-20261004T122522Z", + "LastOutcome": "applied", + "LastAt": "2026-10-04T12:25:58Z", + "LastSuccessfulLeg": "2026-10-04T12:25:58Z", + "Pending": 78, + "NotCovered": 78, + "NotCoveredFast": 0, + "RestartNeeded": 0, + "RebootNeededSince": "0001-01-01T00:00:00Z", + "WrapperPassSeconds": 20.2 + } + }, + { + "HostID": "drill-r50-0a4f9a", + "Ring": 1, + "Enabled": true, + "Tunnel": "inactive", + "Guest": { + "ReleaseID": "", + "LastOutcome": "", + "LastAt": "0001-01-01T00:00:00Z", + "LastSuccessfulLeg": "0001-01-01T00:00:00Z", + "Pending": 0, + "NotCovered": 0, + "NotCoveredFast": 0, + "RestartNeeded": 0, + "RebootNeededSince": "0001-01-01T00:00:00Z", + "WrapperPassSeconds": 0 + }, + "Host": { + "ReleaseID": "", + "LastOutcome": "", + "LastAt": "0001-01-01T00:00:00Z", + "LastSuccessfulLeg": "0001-01-01T00:00:00Z", + "Pending": 0, + "NotCovered": 0, + "NotCoveredFast": 0, + "RestartNeeded": 0, + "RebootNeededSince": "0001-01-01T00:00:00Z", + "WrapperPassSeconds": 0 + } + } + ], + "latest_guest_release": { + "approved_at": "2026-10-04T12:39:33Z", + "approved_by": "auto", + "id": "os-guest-20261004-123933" + }, + "latest_host_release": { + "approved_at": "2026-10-04T12:40:34Z", + "approved_by": "auto", + "id": "os-host-20261004-124034" + } +} diff --git a/documentation/audits/os-host-lane-2026-10-04/partD/after-demo-felhom-agent0.141.1.log b/documentation/audits/os-host-lane-2026-10-04/partD/after-demo-felhom-agent0.141.1.log new file mode 100644 index 00000000..974bcae7 --- /dev/null +++ b/documentation/audits/os-host-lane-2026-10-04/partD/after-demo-felhom-agent0.141.1.log @@ -0,0 +1,172 @@ +=== felhom-agent 0.141.1 selftest=os-update vmid=9201 ring=0 enabled=true guest-release=false host-release=false appliance=true === +time=2026-10-04T13:52:57.534+02:00 level=INFO msg="osupdate: START" run=20261004T115257Z layer=guest vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T115257Z +time=2026-10-04T13:53:09.479+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T115257Z layer=guest:9201 lane=fast mode=apply select=pending-fast packages=0" +time=2026-10-04T13:53:09.479+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0" +time=2026-10-04T13:53:09.479+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0" +time=2026-10-04T13:53:09.479+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)" +time=2026-10-04T13:53:09.480+02:00 level=INFO msg="osupdate: DONE" run=20261004T115257Z layer=guest vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=6 not_covered=6 restart_needed=0 reboot_needed=false wrapper_seconds=11.9 +time=2026-10-04T13:53:09.491+02:00 level=INFO msg="osupdate: START" run=20261004T115257Z layer=host vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T115257Z +time=2026-10-04T13:53:19.927+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T115257Z layer=host lane=fast mode=apply select=pending-fast packages=0" +time=2026-10-04T13:53:19.927+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0" +time=2026-10-04T13:53:19.927+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0" +time=2026-10-04T13:53:19.927+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)" +time=2026-10-04T13:53:20.705+02:00 level=INFO msg="osupdate: DONE" run=20261004T115257Z layer=host vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=80 not_covered=80 restart_needed=34 reboot_needed=true wrapper_seconds=10.4 + --- os-update report (guest) --- + { + "health_reason": "", + "healthy": true, + "mode": "apply", + "not_covered": [ + "docker-ce-cli", + "containerd.io", + "docker-ce", + "docker-buildx-plugin", + "docker-ce-rootless-extras", + "docker-compose-plugin" + ], + "outcome": "nothing", + "pending": 6, + "reboot_needed": false, + "refused": null, + "release_id": "ring0-20261004T115257Z", + "restart_needed": null, + "ring": 0, + "run_id": "20261004T115257Z", + "upgraded": [], + "wrapper_seconds": 11.9 + } + --- os-update report (host) --- + { + "health_reason": "", + "healthy": true, + "mode": "apply", + "not_covered": [ + "frr", + "shim-signed-common", + "shim-unsigned", + "shim-helpers-amd64-signed", + "shim-signed", + "libradosstriper1", + "librgw2", + "ceph-common", + "librbd1", + "librados2", + "python3-cephfs", + "libcephfs2", + "python3-rgw", + "python3-rados", + "python3-ceph-argparse", + "python3-ceph-common", + "python3-rbd", + "ceph-fuse", + "chrony", + "libcorosync-common4", + "libcfg7", + "libcmap4", + "libcpg4", + "libknet1t64", + "libnozzle1t64", + "libquorum5", + "libvotequorum8", + "corosync", + "frr-pythontools", + "libjs-extjs", + "libnvpair3linux", + "libproxmox-acme-plugins", + "libproxmox-backup-qemu0", + "pve-qemu-kvm", + "libpve-notify-perl", + "libpve-cluster-api-perl", + "libpve-cluster-perl", + "pve-cluster", + "libpve-access-control", + "libpve-apiclient-perl", + "librados2-perl", + "proxmox-backup-client", + "proxmox-backup-file-restore", + "pve-manager", + "libproxmox-acme-perl", + "libpve-common-perl", + "libpve-guest-common-perl", + "qemu-server", + "libpve-storage-perl", + "pve-edk2-firmware-legacy", + "pve-edk2-firmware-ovmf", + "libpve-network-api-perl", + "libpve-network-perl", + "proxmox-firewall-data", + "pve-firewall", + "pve-container", + "pve-ha-manager", + "novnc-pve", + "proxmox-enterprise-support-keyring", + "proxmox-mini-journalreader", + "proxmox-widget-toolkit", + "pve-docs", + "pve-i18n", + "pve-xtermjs", + "pve-yew-mobile-i18n", + "pve-yew-mobile-gui", + "libuutil3linux", + "libzfs7linux", + "libzpool7linux", + "proxmox-first-boot", + "pve-firmware", + "proxmox-kernel-7.0.14-20-pve-signed", + "proxmox-kernel-7.0", + "proxmox-kernel-helper", + "pve-edk2-firmware-aarch64", + "pve-edk2-firmware", + "zfs-initramfs", + "zfsutils-linux", + "zfs-zed", + "tailscale" + ], + "outcome": "nothing", + "pending": 80, + "reboot_needed": true, + "refused": null, + "release_id": "ring0-20261004T115257Z", + "restart_needed": [ + "agetty", + "blkmapd", + "chronyd", + "cron", + "dbus-daemon", + "dmeventd", + "ksmtuned", + "lxc-monitord", + "lxc-start", + "lxcfs", + "pmxcfs", + "proxmox-firewal", + "pve-firewall", + "pve-ha-crm", + "pve-ha-lrm", + "pve-lxc-syscall", + "pvedaemon", + "pvedaemon worke", + "pvefw-logger", + "pveproxy", + "pveproxy worker", + "pvescheduler", + "pvestatd", + "qmeventd", + "rpcbind", + "rrdcached", + "smartd", + "spiceproxy", + "spiceproxy work", + "sshd", + "systemd-logind", + "systemd-udevd", + "watchdog-mux", + "zed" + ], + "ring": 0, + "run_id": "20261004T115257Z", + "upgraded": [], + "wrapper_seconds": 10.4 + } + pass took 23.2s +WALL_SECONDS=23.268350584 diff --git a/documentation/audits/os-host-lane-2026-10-04/partE/e0-readonly-demo-hp.txt b/documentation/audits/os-host-lane-2026-10-04/partE/e0-readonly-demo-hp.txt new file mode 100644 index 00000000..4bc88faf --- /dev/null +++ b/documentation/audits/os-host-lane-2026-10-04/partE/e0-readonly-demo-hp.txt @@ -0,0 +1,125 @@ ++ uname -r +7.0.2-6-pve ++ mokutil --sb-state +SecureBoot enabled ++ od -An -tx1 /sys/firmware/efi/efivars/SecureBoot-8be4df61-93ca-11d2-aa0d-00e098032b8c ++ head -1 + 06 00 00 00 01 ++ ls /boot/vmlinuz-7.0.14-20-pve /boot/vmlinuz-7.0.2-6-pve +/boot/vmlinuz-7.0.14-20-pve +/boot/vmlinuz-7.0.2-6-pve ++ dpkg -l 'proxmox-kernel-*' ++ grep '^ii' ++ awk '{print $2,$3}' +proxmox-kernel-7.0 7.0.14-20 +proxmox-kernel-7.0.14-20-pve-signed 7.0.14-20 +proxmox-kernel-7.0.2-6-pve-signed 7.0.2-6 +proxmox-kernel-helper 9.1.0+fde2 ++ proxmox-boot-tool status ++ tail -5 +Re-executing '/usr/sbin/proxmox-boot-tool' in new private mount namespace.. +E: /etc/kernel/proxmox-boot-uuids does not exist. ++ grep -E '^GRUB_DEFAULT|^GRUB_TIMEOUT|^GRUB_SAVEDEFAULT' /etc/default/grub +GRUB_DEFAULT=0 +GRUB_TIMEOUT=5 ++ grub-editenv list ++ '[' -d /sys/firmware/efi ']' ++ efibootmgr ++ head -6 +BootCurrent: 0003 +Timeout: 0 seconds +BootOrder: 0003,0029,0001,0002,0006,0007,0019,001A,001B,001C,001D,001E,0008,0009,0016,0017,001F,0020,0021,0022,0023,0024,0025,0026,0027,0028,000A,0004,0005,000B +Boot0001* USB Floppy/CD VenMedia(b6fef66f-1495-4584-a836-3492d1984a8d,0500000001)0000424f +Boot0002* USB Hard Drive VenMedia(b6fef66f-1495-4584-a836-3492d1984a8d,0200000001)0000424f +Boot0003* proxmox HD(2,GPT,175383fb-546d-430e-9b4c-73ec2379169d,0x800,0x200000)/File(\EFI\proxmox\shimx64.efi) ++ sysctl kernel.panic kernel.panic_on_oops +kernel.panic = 0 +kernel.panic_on_oops = 0 ++ lsmod ++ grep -i -E 'sp5100|wdt|watchdog' ++ ls -l /dev/watchdog /dev/watchdog0 +crw------- 1 root root 10, 130 Oct 4 09:46 /dev/watchdog +crw------- 1 root root 243, 0 Oct 4 09:46 /dev/watchdog0 ++ wdctl ++ head -12 +Device: /dev/watchdog0 +Identity: Software Watchdog [version 0] +Timeout: 10 seconds +Timeleft: 9 seconds +Pre-timeout: 0 seconds +Pre-timeout governor: noop +Available pre-timeout governors: noop +FLAG DESCRIPTION STATUS BOOT-STATUS +KEEPALIVEPING Keep alive ping reply 1 0 +MAGICCLOSE Supports magic close char 0 0 +PRETIMEOUT Pretimeout (in seconds) 0 0 +SETTIMEOUT Set timeout (in seconds) 0 0 ++ dmesg ++ grep -i -E 'sp5100|watchdog' ++ tail -5 +[ 0.353227] NMI watchdog: Enabled. Permanently consumes one hw-PMU counter. ++ apt-cache policy proxmox-default-kernel ++ head -4 +proxmox-default-kernel: + Installed: 2.1.0 + Candidate: 2.1.0 + Version table: ++ apt-cache search --names-only '^proxmox-kernel-[0-9.]+-[0-9]+-pve-signed$' ++ sort -V ++ tail -3 +proxmox-kernel-7.0.14-18-pve-signed - Proxmox Kernel Image (signed) +proxmox-kernel-7.0.14-19-pve-signed - Proxmox Kernel Image (signed) +proxmox-kernel-7.0.14-20-pve-signed - Proxmox Kernel Image (signed) ++ cat /proc/cmdline +BOOT_IMAGE=/boot/vmlinuz-7.0.2-6-pve root=/dev/mapper/pve-root ro quiet ++ grep -nE '^menuentry|^submenu|^\s+menuentry' /boot/grub/grub.cfg ++ cut -c1-140 ++ head -8 +26: menuentry_id_option="--id" +28: menuentry_id_option="" +110:menuentry 'Proxmox VE GNU/Linux' --class proxmox --class gnu-linux --class gnu --class os $menuentry_id_option 'gnulinux-simple-529c0c3d +128:submenu 'Advanced options for Proxmox VE GNU/Linux' $menuentry_id_option 'gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43' { +129: menuentry 'Proxmox VE GNU/Linux, with Linux 7.0.14-20-pve' --class proxmox --class gnu-linux --class gnu --class os $menuentry_id_optio +147: menuentry 'Proxmox VE GNU/Linux, with Linux 7.0.14-20-pve (recovery mode)' --class proxmox --class gnu-linux --class gnu --class os $me +165: menuentry 'Proxmox VE GNU/Linux, with Linux 7.0.2-6-pve' --class proxmox --class gnu-linux --class gnu --class os $menuentry_id_option +183: menuentry 'Proxmox VE GNU/Linux, with Linux 7.0.2-6-pve (recovery mode)' --class proxmox --class gnu-linux --class gnu --class os $menu ++ ls -l --time-style=+%F_%T /boot/grub/grub.cfg +-rw------- 1 root root 13743 2026-10-04_09:47:05 /boot/grub/grub.cfg ++ uptime -s +2026-10-04 09:46:15 ++ last -x reboot ++ head -4 +reboot system boot 7.0.2-6-pve Sun Oct 4 09:46 - still running +reboot system boot 7.0.14-20-pve Sun Oct 4 09:44 - 09:45 (00:00) +shutdown system down 7.0.14-20-pve Sun Oct 4 09:45 - 09:46 (00:00) +reboot system boot 7.0.2-6-pve Fri Aug 21 17:44 - 09:43 (43+15:58) ++ ls /boot/efi/EFI/proxmox/ +BOOTX64.CSV +fbx64.efi +grub.cfg +grubx64.efi +mmx64.efi +shimx64.efi ++ cat /boot/efi/EFI/proxmox/grub.cfg ++ head -5 +search.fs_uuid 529c0c3d-b48e-4d01-989d-43fd5d7dbb43 root lvmid/zVqGDa-V6js-HB26-RyZr-d0ft-fUB3-blL75R/iAs9TN-W8rL-ViG8-jptz-1tcY-3djE-MRqamk +set prefix=($root)'/boot/grub' +configfile $prefix/grub.cfg ++ lspci -nn ++ grep -i -E 'smbus|fch' +00:14.0 SMBus [0c05]: Advanced Micro Devices, Inc. [AMD] FCH SMBus Controller [1022:790b] (rev 61) +00:14.3 ISA bridge [0601]: Advanced Micro Devices, Inc. [AMD] FCH LPC Bridge [1022:790e] (rev 51) +06:00.0 SATA controller [0106]: Advanced Micro Devices, Inc. [AMD] FCH SATA Controller [AHCI mode] [1022:7901] (rev 61) ++ modinfo -F filename sp5100_tco +/lib/modules/7.0.2-6-pve/kernel/drivers/watchdog/sp5100_tco.ko ++ grep -rl sp5100 /lib/modprobe.d /etc/modprobe.d +/lib/modprobe.d/blacklist_proxmox-kernel-7.0.14-20-pve.conf +/lib/modprobe.d/blacklist_proxmox-kernel-7.0.2-6-pve.conf ++ grep -h sp5100 /lib/modprobe.d/aliases.conf /lib/modprobe.d/blacklist_proxmox-kernel-7.0.14-20-pve.conf /lib/modprobe.d/blacklist_proxmox-kernel-7.0.2-6-pve.conf /lib/modprobe.d/fbdev-blacklist.conf /lib/modprobe.d/proxmox_prevent_autoload_proxmox-kernel-7.0.14-20-pve.conf /lib/modprobe.d/proxmox_prevent_autoload_proxmox-kernel-7.0.2-6-pve.conf /lib/modprobe.d/systemd.conf /etc/modprobe.d/amd64-microcode-blacklist.conf /etc/modprobe.d/pve-blacklist.conf /etc/modprobe.d/zfs.conf ++ head -3 +blacklist sp5100_tco +blacklist sp5100_tco ++ modprobe -n -v sp5100_tco +insmod /lib/modules/7.0.2-6-pve/kernel/drivers/watchdog/sp5100_tco.ko ++ ls /sys/class/watchdog/ +watchdog0 diff --git a/documentation/audits/os-host-lane-2026-10-04/partE/e1-saved-default.txt b/documentation/audits/os-host-lane-2026-10-04/partE/e1-saved-default.txt new file mode 100644 index 00000000..47ba51e8 --- /dev/null +++ b/documentation/audits/os-host-lane-2026-10-04/partE/e1-saved-default.txt @@ -0,0 +1,20 @@ ++ findmnt -no SOURCE,FSTYPE --target /boot +/dev/mapper/pve-root ext4 ++ ls -l /boot/grub/grubenv +-rw-r--r-- 1 root root 1024 Oct 4 09:46 /boot/grub/grubenv ++ cp -p /etc/default/grub /root/grub.default.bak-2026-10-04 ++ sed -i 's/^GRUB_DEFAULT=0$/GRUB_DEFAULT=saved/' /etc/default/grub ++ grep '^GRUB_DEFAULT' /etc/default/grub +GRUB_DEFAULT=saved ++ update-grub ++ tail -3 +Found memtest86+ 32bit image: /boot/memtest86+ia32.bin +Adding boot menu entry for UEFI Firmware Settings ... +done ++ grub-set-default 'gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.2-6-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43' ++ grub-editenv list +saved_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.2-6-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43 ++ grep -n 'set default' /boot/grub/grub.cfg ++ head -4 +17: set default="${next_entry}" +22: set default="${saved_entry}" diff --git a/documentation/audits/os-host-lane-2026-10-04/partE/e2-before-reboot1.txt b/documentation/audits/os-host-lane-2026-10-04/partE/e2-before-reboot1.txt new file mode 100644 index 00000000..2f9c9494 --- /dev/null +++ b/documentation/audits/os-host-lane-2026-10-04/partE/e2-before-reboot1.txt @@ -0,0 +1,14 @@ ++ grub-reboot 'gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.14-20-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43' + +WARNING: Detected GRUB environment block on lvm device +gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.14-20-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43 will remain the default boot entry until manually cleared with: + grub-editenv /boot/grub/grubenv unset next_entry + ++ grub-editenv list +saved_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.2-6-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43 +next_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.14-20-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43 ++ uname -r +7.0.2-6-pve ++ date -u +%FT%TZ +2026-10-04T12:26:10Z +reboot1 issued 2026-10-04T12:26:10Z diff --git a/documentation/audits/os-host-lane-2026-10-04/partE/e3-after-reboot1.txt b/documentation/audits/os-host-lane-2026-10-04/partE/e3-after-reboot1.txt new file mode 100644 index 00000000..d551b4db --- /dev/null +++ b/documentation/audits/os-host-lane-2026-10-04/partE/e3-after-reboot1.txt @@ -0,0 +1,19 @@ ++ uname -r +7.0.14-20-pve ++ uptime -s +2026-10-04 14:27:03 ++ mokutil --sb-state +SecureBoot enabled ++ grub-editenv list +saved_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.2-6-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43 +next_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.14-20-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43 ++ cat /proc/cmdline +BOOT_IMAGE=/boot/vmlinuz-7.0.14-20-pve root=/dev/mapper/pve-root ro quiet +active +active +active +active +active +status: running +starting +starting diff --git a/documentation/audits/os-host-lane-2026-10-04/partE/e4-reboot2.txt b/documentation/audits/os-host-lane-2026-10-04/partE/e4-reboot2.txt new file mode 100644 index 00000000..8055e6bc --- /dev/null +++ b/documentation/audits/os-host-lane-2026-10-04/partE/e4-reboot2.txt @@ -0,0 +1,33 @@ +reboot2 issued 2026-10-04T12:35:19Z, no grub command; grubenv as after reboot 1 +ssh back after ~40s +System is going down. Unprivileged users are not permitted to log in anymore. For technical details, see pam_nologin(8). + ++ uname -r +7.0.14-20-pve ++ uptime -s +2026-10-04 14:27:03 ++ grub-editenv list +saved_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.2-6-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43 +next_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.14-20-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43 ++ cat /proc/cmdline +BOOT_IMAGE=/boot/vmlinuz-7.0.14-20-pve root=/dev/mapper/pve-root ro quiet +--- after the real restart, 2026-10-04T12:36:27Z ++ uname -r +7.0.14-20-pve ++ uptime -s +2026-10-04 14:36:11 ++ mokutil --sb-state +SecureBoot enabled ++ grub-editenv list +saved_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.2-6-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43 +next_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.14-20-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43 ++ cat /proc/cmdline +BOOT_IMAGE=/boot/vmlinuz-7.0.14-20-pve root=/dev/mapper/pve-root ro quiet ++ systemctl is-active pveproxy pvedaemon pvestatd pve-cluster felhom-agent +activating +active +active +active +inactive ++ pct status 9201 +status: stopped diff --git a/documentation/audits/os-host-lane-2026-10-04/partE/e5-final-state.txt b/documentation/audits/os-host-lane-2026-10-04/partE/e5-final-state.txt new file mode 100644 index 00000000..7f300e6f --- /dev/null +++ b/documentation/audits/os-host-lane-2026-10-04/partE/e5-final-state.txt @@ -0,0 +1,27 @@ ++ grub-editenv /boot/grub/grubenv unset next_entry ++ grub-set-default 'gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.14-20-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43' ++ grub-editenv list +saved_entry=gnulinux-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43>gnulinux-7.0.14-20-pve-advanced-529c0c3d-b48e-4d01-989d-43fd5d7dbb43 ++ grep '^GRUB_DEFAULT' /etc/default/grub +GRUB_DEFAULT=saved ++ dpkg -l 'proxmox-kernel-*-pve-signed' ++ awk '{print $2}' ++ grep '^ii' +proxmox-kernel-7.0.14-20-pve-signed +proxmox-kernel-7.0.2-6-pve-signed ++ uname -r +7.0.14-20-pve ++ systemctl is-active pveproxy pvedaemon pvestatd pve-cluster felhom-agent +active +active +active +active +active ++ pct status 9201 +status: running ++ sleep 45 ++ pct exec 9201 -- docker inspect -f '{{.Name}} {{.State.Health.Status}}' cloudflared felhom-controller +/cloudflared healthy +/felhom-controller healthy ++ sysctl kernel.panic +kernel.panic = 0 diff --git a/documentation/audits/os-host-lane-2026-10-04/partE/e6-sp5100.txt b/documentation/audits/os-host-lane-2026-10-04/partE/e6-sp5100.txt new file mode 100644 index 00000000..0f690693 --- /dev/null +++ b/documentation/audits/os-host-lane-2026-10-04/partE/e6-sp5100.txt @@ -0,0 +1,74 @@ ++ modprobe sp5100_tco +rc=0 ++ echo rc=0 ++ lsmod ++ grep sp5100 +sp5100_tco 20480 0 ++ dmesg ++ grep -i sp5100 ++ tail -5 +[ 85.993589] sp5100_tco: SP5100/SB800 TCO WatchDog Timer Driver +[ 85.993780] sp5100-tco sp5100-tco: Using 0xfeb00000 for watchdog MMIO address +[ 85.993918] sp5100-tco sp5100-tco: initialized. heartbeat=60 sec (nowayout=0) +== /sys/class/watchdog/watchdog0 ++ for w in /sys/class/watchdog/watchdog* ++ echo '== /sys/class/watchdog/watchdog0' ++ for f in identity state timeout nowayout bootstatus status +++ cat /sys/class/watchdog/watchdog0/identity +identity=Software Watchdog ++ printf '%s=%s\n' identity 'Software Watchdog' ++ for f in identity state timeout nowayout bootstatus status +++ cat /sys/class/watchdog/watchdog0/state +state=active ++ printf '%s=%s\n' state active ++ for f in identity state timeout nowayout bootstatus status +++ cat /sys/class/watchdog/watchdog0/timeout +timeout=10 ++ printf '%s=%s\n' timeout 10 ++ for f in identity state timeout nowayout bootstatus status +++ cat /sys/class/watchdog/watchdog0/nowayout +nowayout=0 ++ printf '%s=%s\n' nowayout 0 ++ for f in identity state timeout nowayout bootstatus status +++ cat /sys/class/watchdog/watchdog0/bootstatus +bootstatus=0 ++ printf '%s=%s\n' bootstatus 0 ++ for f in identity state timeout nowayout bootstatus status +++ cat /sys/class/watchdog/watchdog0/status +status=0x8000 +== /sys/class/watchdog/watchdog1 ++ printf '%s=%s\n' status 0x8000 ++ for w in /sys/class/watchdog/watchdog* ++ echo '== /sys/class/watchdog/watchdog1' ++ for f in identity state timeout nowayout bootstatus status +++ cat /sys/class/watchdog/watchdog1/identity +identity=SP5100 TCO timer ++ printf '%s=%s\n' identity 'SP5100 TCO timer' ++ for f in identity state timeout nowayout bootstatus status +++ cat /sys/class/watchdog/watchdog1/state +state=inactive ++ printf '%s=%s\n' state inactive ++ for f in identity state timeout nowayout bootstatus status +++ cat /sys/class/watchdog/watchdog1/timeout +timeout=60 ++ printf '%s=%s\n' timeout 60 ++ for f in identity state timeout nowayout bootstatus status +++ cat /sys/class/watchdog/watchdog1/nowayout +nowayout=0 ++ printf '%s=%s\n' nowayout 0 ++ for f in identity state timeout nowayout bootstatus status +++ cat /sys/class/watchdog/watchdog1/bootstatus +bootstatus=0 ++ printf '%s=%s\n' bootstatus 0 ++ for f in identity state timeout nowayout bootstatus status +++ cat /sys/class/watchdog/watchdog1/status +status=0x0 ++ printf '%s=%s\n' status 0x0 ++ rmmod sp5100_tco +rmmod_rc=0 ++ echo rmmod_rc=0 ++ lsmod ++ grep -c sp5100 +0 ++ ls /sys/class/watchdog/ +watchdog0 diff --git a/documentation/audits/os-host-lane-2026-10-04/partG/wrapper-copied-by-hand.txt b/documentation/audits/os-host-lane-2026-10-04/partG/wrapper-copied-by-hand.txt index d2588081..517d89e5 100644 --- a/documentation/audits/os-host-lane-2026-10-04/partG/wrapper-copied-by-hand.txt +++ b/documentation/audits/os-host-lane-2026-10-04/partG/wrapper-copied-by-hand.txt @@ -1,2 +1,4 @@ demo-felhom wrapper 4729769ce32e25e6 755 root; sudoers 02df92d751f1780a demo-hp wrapper 4729769ce32e25e6 755 root; sudoers 02df92d751f1780a +demo-felhom wrapper 51e100ad81945f67 +demo-hp wrapper 51e100ad81945f67 diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 86c54abf..2f1b399a 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -26,6 +26,17 @@ --- +## 2026-10-04 (afternoon) — OS updates, host fast lane + fleet view + alarms (agent v0.141.0/v0.141.1, hub v0.131.0/v0.131.1, controller v0.292.0) + +> Evidence: `audits/os-host-lane-2026-10-04/`. + +| Row | What | Closed | Evidence | +|---|---|---|---| +| **R-841** | **Every box reported its tunnel `inactive`** (the agent asked a host unit that does not exist). Agent v0.141.0 reads the guest's `cloudflared` container and controller v0.292.0's Docker health check on cloudflared's own `/ready` (200 only with a connection); three states `running` / `not_running` / `unknown`; hub v0.131.0 alarms `tunnel_down` after two `not_running` reports and `tunnel_recovered` on the next `running`. LIVE on demo-hp: port 7844 blocked → `tunnel_down` mailed after the 2nd report; unblocked → `tunnel_recovered`. **Reasoning kept: `unknown` never alarms; a container state alone says "up" for a dead tunnel (wrong token: running, `/ready` 503). A plain `docker stop` is healed by the controller's protected-container check within 5 min — before the 15-min host report sees it.** | CLOSED 2026-10-04 — FIXED | `partA/` | +| **R-845** | **The OS leg was slow.** One `pct exec` per package (~0.9 s each) replaced by one call per layer; restart scan only after an install (host: every pass, v0.141.1); repair only when `dpkg --audit` reports. MEASURED: nothing to install, both layers — 23.3 s (demo-felhom), 31.5 s (demo-hp); before, guest only, nothing to install — 14.0 s; the 174–245 s passes of R-845 were passes WITH an install. A 108-package host install pass: 70 s. | CLOSED 2026-10-04 — FIXED | `partD/`, `partB/live/` | +| **R-846** | **Host "reboot needed" was wrong in agent v0.141.0** (found live on demo-felhom): the restart scan skipped every cgroup containing `lxc`, hiding `lxc-start` (`0::/lxc.monitor/`, 20 deleted maps after libc6); and it scanned only after an install, so a reboot never cleared it (the hub's 14-day alarm would fire on a rebooted host). Agent v0.141.1 (`:/lxc/`, host scans every pass, `reboot_scanned`) + hub v0.131.1. Red-proved. | CLOSED 2026-10-04 — FIXED | `partB/live-defects-redproofs.txt` | +| **R-850** | **Hub v0.131.0 put the layer into the release fingerprint**, so the unchanged guest set (272 packages) counted as new and waited its 24 h again. One-time; ring 1 kept the previous release meanwhile; approved again under the TEST wait (`os-guest-20261004-123933`). Nothing to fix. | CLOSED 2026-10-04 — ONE-TIME, NO ACTION | hub log 2026-10-04 | + ## 2026-10-04 (~12:20) — operator ruling | Row | What | Closed | Evidence | diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index da81f9b3..a2cf2832 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -308,7 +308,7 @@ stopping line that lies. |---|---|---|---|---|---|---|---| | **R-530** | Box system & updates | P2 | **[P2-MEDIUM] A floor does not deliver an agent: agents update only by an operator-signed `agent_update` job per box, and nothing records which boxes still run 0.130.0.** MEASURED 2026-09-15: the hub HOLDS a floor whose declared MinAgent is above the box's agent (`api/handler.go` ResolveManagedFloor); the agent's only update path is `signedjobs` + `selfupdate.Executor`. demo-hp reached 0.131.0 by `felhom-opsign -op agent_update` (key `felhom-op-1`) at 08:44:16Z and its controller floor was then SERVED in 3 s. **demo-felhom (N100) and Peti's box still run 0.130.0** — not touched (Peti fenced; N100 not asked). **What it needs:** the operator signs per box, or rules a fleet rollout step. **NARROWED 2026-09-16 (operator ruling 1):** the keys stay on DooPlex owner-only and CC may sign `agent_update` until the first PAYING customer (testers excluded) — recorded in `CONTEXT.md` + `04-control-plane-authorization.md` §3.1. Both demo boxes now run agent 0.131.0 (demo-hp 2026-09-15, demo-felhom 2026-09-16, each by a per-box signed job; Peti's box untouched, still 0.130.0). **What remains:** a fleet rollout step — signing per box does not scale past a handful, and nothing lists which boxes are behind. **2026-09-25:** Peti's box was RETIRED (operator ruling) — it no longer counts as a box behind. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (signing)** | — | — | operator | | **R-604** | Box system & updates | P2 | **[P2-MEDIUM] A per-customer controller floor silently excludes that box from every global floor raise, and NOTHING says so — demo-hp missed four of them.** FOUND 2026-09-21 while raising the global floor to 0.259.0 at the operator's request. The raise logged `Global controller-version floor set to "0.259.0"` and then `managed floor SERVED for demo-felhom` — **and nothing at all for demo-hp**, which went on reporting every few minutes and stayed on 0.258.0. Cause: `customer_configs.min_controller_version` for demo-hp held **`0.243.0`**, a per-customer override that wins over the global. It is a **leftover from the 2026-09-16 drill**, whose golden was 0.243.0; **R-343 measured on 2026-08-18 that all five rows were EMPTY and recorded that as a safety property** — it stopped being true and nothing surfaced the change. demo-hp had therefore silently missed the raises to 0.253.0, 0.254.0, 0.257.0 and 0.259.0. **Why it is invisible rather than merely quiet:** `managed floor SERVED` fires **once per CHANGE** (`h.floorNotes`, `api/handler.go:600`), deliberately, because a box reports every few minutes — so a box whose override never changes is silent for ever, and its silence is indistinguishable from the silence of a box that already had the line. A session that raises the floor reads one SERVED line and reasonably concludes the fleet took it. **CLEARED** for demo-hp the same session (rollback line: POST `/customers/demo-hp/floor` with `min_controller_version=0.243.0`, `min_agent=0.131.0`); it then self-updated 0.258.0 → 0.259.0 in **under four minutes**, healthy, `settle-gate: GO — at/above floor 0.259.0`, and its claim page answers **"Wrong or expired code"** in English — the floor delivered the FIX, not a version string, to a box nobody hand-deployed. All five overrides are now empty. **Fix shape (smallest first):** the floor-raise page shows which customers carry an override and would NOT be moved, before the save; or the raise logs one line per customer naming the ones it skipped and why. A raise that quietly reaches half the fleet is worse than one that refuses. | **READY — rank P2-MEDIUM; owner: CC (hub)** | — | — | CC | -| **R-812** | Box system & updates | P2 | **[P2] A box never receives operating-system security updates — not the Proxmox host, not the guest's Debian, not its Docker engine.** SEARCHED 2026-10-03 (read-only): `felhom-controller`, `felhom-agent`, `app-catalog-felhom.eu` and `felhom.eu` hold no `apt-get upgrade`, `apt full-upgrade`, `unattended-upgrades`, `pveupgrade` or `needrestart` that runs on a box. The installer aligns the host's Proxmox repositories to no-subscription *"so the box can pull security updates"* and then says plainly *"No upgrades are run"* (`scripts/felhom-host-install.sh:2133-2136`). The guest's Docker engine is installed when the golden is BAKED (`felhom-agent/configs/build-golden.sh:124-125`), so a fresh install gets that week's engine and an installed box keeps it forever. The only `apt full-upgrade` in the project is a by-hand step for the off-site endpoint ep0 (`documentation/runbooks/offsite-endpoint.md:41`), not a box. App images ARE updated (the update arc); the layer under them is not. The intention, with its scope, is **R-808** in `ROADMAP.md`. | **NARROWED AGAIN 2026-10-04 (day) — the GUEST's Debian fast lane is BUILT and proven live (agent v0.140.0, hub v0.130.0, installer 1.29.0; `11` §8.1, `audits/os-guest-lane-2026-10-04/`). LEFT: the host, Docker and kernel lanes; the undo (R-842); existing boxes (R-840).** Earlier: NARROWED 2026-10-04 — the SPIKE is done (`11` §7.1, corrections C1–C12, `audits/os-updates-spike-2026-10-04/`); no product code yet. LEFT: the build steps of `11` §8, each with the operator's go; the §5.3 snapshot question is in STATUS; preconditions R-835, R-836, R-837.** | — | — | CC + operator | +| **R-812** | Box system & updates | P2 | **[P2] A box never receives operating-system security updates — not the Proxmox host, not the guest's Debian, not its Docker engine.** SEARCHED 2026-10-03 (read-only): `felhom-controller`, `felhom-agent`, `app-catalog-felhom.eu` and `felhom.eu` hold no `apt-get upgrade`, `apt full-upgrade`, `unattended-upgrades`, `pveupgrade` or `needrestart` that runs on a box. The installer aligns the host's Proxmox repositories to no-subscription *"so the box can pull security updates"* and then says plainly *"No upgrades are run"* (`scripts/felhom-host-install.sh:2133-2136`). The guest's Docker engine is installed when the golden is BAKED (`felhom-agent/configs/build-golden.sh:124-125`), so a fresh install gets that week's engine and an installed box keeps it forever. The only `apt full-upgrade` in the project is a by-hand step for the off-site endpoint ep0 (`documentation/runbooks/offsite-endpoint.md:41`), not a box. App images ARE updated (the update arc); the layer under them is not. The intention, with its scope, is **R-808** in `ROADMAP.md`. | **NARROWED 2026-10-04 (afternoon) — the HOST's Debian fast lane is BUILT and proven live too (agent v0.141.1, hub v0.131.1; `11` §8.2, `audits/os-host-lane-2026-10-04/`): appliances only, never kernel/boot/firmware, after a healthy guest step; fleet view and four alarms (§8.3). LEFT: the Docker and kernel slow lanes (R-836); existing boxes (R-840).** Earlier: NARROWED AGAIN 2026-10-04 (day) — the GUEST's Debian fast lane is BUILT and proven live (agent v0.140.0, hub v0.130.0, installer 1.29.0; `11` §8.1, `audits/os-guest-lane-2026-10-04/`). LEFT: the host, Docker and kernel lanes; the undo (R-842); existing boxes (R-840).** Earlier: NARROWED 2026-10-04 — the SPIKE is done (`11` §7.1, corrections C1–C12, `audits/os-updates-spike-2026-10-04/`); no product code yet. LEFT: the build steps of `11` §8, each with the operator's go; the §5.3 snapshot question is in STATUS; preconditions R-835, R-836, R-837.** | — | — | CC + operator | | **R-35** | Box system & updates | P3 | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` | CC | | **R-50b** | Box system & updates | P3 | **[P2] A root-owned privileged host artifact is delivered unversioned from `main` — "which wrapper is on this host?" is unanswerable.** `configs/felhom-pbs-apply` installs to `/usr/local/sbin/felhom-pbs-apply` (0755 root:root) and is the pinned sudoers vector for `create\ |reconcile\|grant` against `/etc/pve/priv/storage`. It is fetched by `felhom-host-install.sh:1914` via `fetch_raw`, which hits `raw/branch/main/` — **no tag, no pin, no checksum, and no record in the Day-0 artifact manifest**, unlike the agent binary (sha256-vouched) and the golden image. Three consequences: (1) two hosts installed a week apart can carry different privileged wrapper code while both reporting the same agent version; (2) a host hotfixed in place (felhom-pve, 2026-07-18) is indistinguishable from one that fetched the same content — the fleet has no inventory of it; (3) an accidental push to `main` reaches the next install of every host with no review gate between commit and root-owned deployment. | **NARROWED** — **(a) SHIPPED 2026-07-21; (b)/(c) open** — moved from `ROADMAP.md` 2026-10-03: it states a checkable fact about the shipped product, so it is a FINDING (the sorting rule). **Re-ranked 2026-10-03: [P2] → P3 — operator-only; leg (a) shipped.** | — | **Surfaced 2026-07-21 while stopping the R-39 v0.90.1 publish** (`felhom-controller/REPORT.md` §5): the publish was cancelled precisely because the version number would have claimed to carry a fix that in fact rides this unversioned channel. Candidate shapes, in increasing cost: (a) record the wrapper's sha256 in the Day-0 artifact manifest beside the agent binary and have the agent report the installed file's hash, so drift is at least *visible*; (b) `fetch_raw` takes a pinned ref (tag or commit) supplied by the manifest rather than `main`; (c) the wrapper becomes a published generic-registry artifact with the same gate ladder as the agent binary. **(a) is the cheap honest first step and would have caught this class already.** Pairs with R-39 (whose remaining fleet half is specced separately) **(a) SHIPPED 2026-07-21 — hub v0.68.0 + agent v0.91.2.** `ArtifactManifest.WrapperSHA256` + an operator field; agents report the installed wrapper's sha256 each cycle and the host page surfaces a mismatch. **An unknown on EITHER side reads as quiet, never as drift** — lighting every host amber on rollout day is how a warning becomes background noise. Live confirmation of exactly the problem: felhom-pve's July-18 in-place hotfix hashed `2888f2ea…`, matching **no commit anyone could name**; it now reports `104db0a4…` against a vouchable manifest value. **(b)/(c) REMAIN OPEN:** the wrapper is still fetched unversioned from `raw/branch/main` — this makes drift *visible*, it does not fix the channel. Also recorded: the 0440 sudoers file is not agent-readable, so its drift stays invisible. **Flips (2026-10-03):** `00` §A "The installer is PUBLISHED, not pushed" — the same discipline for the privileged wrappers. **Re-ranked 2026-10-03:** [P2] → P3: operator-only; (a) shipped 2026-07-21 (the report carries the wrapper sha256), (b)/(c) open. **Finding-shaped** — an R-424 instance; check against today's product before building. | CC | | **R-78** | Box system & updates | P3 | **`local_api` authority ruling — auto-reconcile vs detect-only** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-25, size M, roadmap state `idea (deferred OUT of R-77 on purpose)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **An OWED OPERATOR DECISION, not a defect — moved because an owed ruling hidden among feature ideas is the shape this session exists to remove.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | R-77 ships detection because the fix is genuinely undecided, and **both directions can lose customer-visible function**. **Direction 1 (today):** `controller.yaml` wins and drift is silent → the 2026-07-25 island migration blinded the whole fleet's control plane for 17.5 h (drive gate, guest-reboot recovery, quiesce/backup all degrade). R-77 makes that loud but does not stop it recurring. **Direction 2 (`bootstrap.json` wins, auto-reconcile on boot):** a guest whose `controller.yaml` is CORRECT and whose `bootstrap.json` is stale — a half-completed re-provision, a hand-repaired guest, a setup-wizard box — gets a **working channel clobbered on the next restart**, fleet-wide and silently, during a routine deploy. That is not obviously better than the bug. Needs a spike: which writer is authoritative per field (endpoint vs fingerprint vs token — `mergeLocalAPI` replaces the whole block, so they cannot be reconciled independently today), whether the agent side should stamp a generation/mtime so 'newer wins' is even expressible, and whether reconcile should require an operator ack. Until then the drift alert plus a manual edit is the supported path. | CC | @@ -324,10 +324,10 @@ stopping line that lies. | **R-194** | Box system & updates | P4 | **PVE's permission cache delays every grant-state verdict by an unknown amount, so "the agent can read it" and "the ACL exists" are not the same measurement.** Observed twice while validating R-190's self-repair on demo-felhom 2026-08-04: both ACL rows for `/storage/felhom-backup` were deleted, and `GET /access/permissions` continued to report `Datastore.AllocateSpace` present — for **~40 s** in one run and **~16 minutes** in another. During that window the capability probe reads healthy and the self-repair does not fire | **OPEN** | — | **Why it matters beyond the delay:** it puts a floor under how fast a lost grant can be noticed, it makes any single permission read a lagging indicator, and — the interesting part — **it is a candidate contributor to R-190's own timeline**: a grant removed at an unknown moment could keep working until a cache expiry, which is exactly the shape of *worked at 04:44, refused at 09:24*. That does not explain what removed it, but it may explain when the refusal SURFACED, and the two have been treated as the same instant. **Not a defect in our code** — it is PVE behaviour, and the mitigation already tolerates it (the repair fires on the next probe after the cache clears). **What is worth deciding:** whether the store-grant probe should ALSO consult the storage content listing as a second signal, since that appeared to reflect the loss immediately (`{"data":[]}` while the permission read still said present) — two signals disagreeing is itself information, and today only one of them is read | CC | | **R-373** | Box system & updates | P4 | **`SysDataGrowGB` is the intended lever for the system-data volume, it works, and nothing sets it.** Written down 2026-08-02 in `audits/SPIKE-recovery-unit-space-2026-08-02.md:230-232`, under an explicit *"### Not filed"* heading: the 20 G / 50 G mismatch was ruled a tier-sizing decision rather than a defect, *"`SysDataGrowGB` is the intended lever and it works; nothing sets it."* A lever with no caller is the same shape as R-368's comment — a setting that names a behaviour nothing invokes. **Age when filed: 20 days.** | **OPEN — LOW** | R-368 (same shape) | Either wire it to something an operator can reach, or remove it and record the sizing decision where a reader will meet it. | CC | | **R-835** | Box system & updates | P3 | **Turning Docker's `live-restore` OFF with a restart stops every running container and starts none.** MEASURED 2026-10-04 on scratch 9202: `live-restore` on (via `systemctl reload docker`, which does enable it) kept all 6 containers running across two engine steps; `systemctl reload` with the baked `daemon.json` did NOT turn it off; a `systemctl restart docker` did — and the new daemon stopped every container (`Exited (0)`, `Removing stale sandbox … isRestore=false`) and restarted none, though all are `unless-stopped`. Nothing brought them back for 3.5 min. A precondition for the Docker slow lane (`11` C5): if `live-restore` ships, turning it off must be a guarded act (stop apps first), never a plain restart. `audits/os-updates-spike-2026-10-04/partG/` | **READY — design input, owner: CC** | — | — | CC | -| **R-836** | Box system & updates | P3 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **READY — measure before the kernel slow lane; owner: CC + operator (reboots)** | — | — | CC | +| **R-836** | Box system & updates | P3 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot `: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** | — | — | CC | +| **R-851** | Box system & updates | P3 | **A host that panics stays stopped: `kernel.panic = 0`.** READ 2026-10-04 on demo-hp (os-host-lane Part E): `kernel.panic = 0`, `kernel.panic_on_oops = 0`. The brief assumed the host restarts after a panic; it does not — the box stays down until a person power-cycles it, and only `softdog` (dead in a panic) runs. A fix (`kernel.panic = 10` set by the installer, maybe `panic_on_oops`) changes how every box behaves and what a household sees, so it is the operator's call, together with the kernel lane. `audits/os-host-lane-2026-10-04/partE/e0-readonly-demo-hp.txt` | **READY — operator decision (with the kernel lane); owner: operator** | — | — | operator | | **R-839** | Box system & updates | P3 | **After a Docker restart that stopped every container, the boot sweep HELD an app whose `app.yaml` `HDD_PATH` names its user folder instead of the drive.** MEASURED 2026-10-04 on scratch 9202 (paperless-ngx): `bootrecon` logged `drive /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx is not a live mountpoint — NOT starting it`, although the drive `/mnt/felhom-drives/scratch_hdd` IS a mountpoint; the app stayed down until started by hand. The customer boxes' 9201s showed no hold. **Not diagnosed:** which writer put a per-app path in `HDD_PATH` on 9202, and whether a household box can get it. `audits/os-updates-spike-2026-10-04/partG/SUMMARY.md` | **READY — diagnose; owner: CC** | — | — | CC | | **R-840** | Box system & updates | P2 | **A new root wrapper or a new sudoers line cannot reach a box that is already installed — there is no product route.** FOUND 2026-10-04 (guest OS lane, Part B 1): the agent's signed `agent_update` replaces ONLY the binary (`configs/felhom-selfupdate-guarded` swaps `/usr/local/bin/felhom-agent`); wrappers (`/usr/local/sbin/felhom-*`) and `/etc/sudoers.d/felhom-agent` are written only by `felhom-host-install.sh` step 5 on a new box. So agent v0.140.0's OS leg is DEAD on every existing box (the capability probe will show `osapply-run` missing) until someone installs `felhom-os-apply` + the FELHOM_OSAPPLY line by hand — which is what the demo boxes got this session. The `wireguard-tools` line (S3) reached the fleet the same way: by reinstall or by hand. **Proposal:** a signed `agent_config_update` op (operator-signed like `agent_update`) whose params pin the agent TAG and the sha256 of a config bundle (sudoers + wrappers); the self-update wrapper installs it as root after `visudo -cf` and a syntax check, keeps the previous copies, and the capability probe confirms. `audits/os-guest-lane-2026-10-04/README.md` | **READY — design + operator go; owner: CC** **RULED 2026-10-04 ~12:20 (`09` §3 decision 82): not built now — "There are no older boxes"; no installed box exists outside the two demo boxes, which take wrapper changes by hand. Kept open, not re-ranked. Reviewer's note (as written): every box installed from now on becomes an "older box" the first time the wrapper or the sudoers line changes again (the 2026-10-04 host-lane brief changes the wrapper); R-840 must be solved before the first box that cannot be reached by hand.** | — | — | CC | -| **R-841** | Box system & updates | P3 | **The agent's `cloudflared` health probe reads a host systemd unit that does not exist — every box reports its tunnel `inactive`.** FOUND 2026-10-04: `felhom-agent/internal/hub/cloudflared.go` runs `systemctl is-active cloudflared` on the HOST; cloudflared is a container in the GUEST (`11` C8), so demo-hp answers `inactive` / `Unit cloudflared.service could not be found`, and the hub stores that in `cloudflared_status` for every box. The field is equally consistent with "tunnel down" and "never checked" (R-96 rule 3). Fix direction: read the guest's container state (the agent already may `pct exec * -- docker inspect -f *`), or drop the field; and say so in `03` (corrected 2026-10-04). | **READY — owner: CC** | — | — | CC | ## Monitoring & notifications — 24 rows (P2 2, P3 16, P4 6) @@ -358,7 +358,7 @@ stopping line that lies. | **R-348** | Monitoring & notifications | P4 | **Every agent restart blanks the reported backup list for up to ~18 hours, and the comment that covers it says "unaffected".** Observed 2026-08-20 while deploying R-344: the first host reports after `demo-hp`'s agent restart carry **`0 backups`** (11:15:50 and 11:30:52 CEST, two consecutive), while the box's own `pvesm list` shows archives present on **both** tiers. `internal/backup/store.go`'s `Store` is in-memory and `byTarget` is repopulated only when a backup **runs** — daily for the local tier, weekly for offsite — so the field reads 0 until the next run. `restore_tests` did **not** blank, because that half has a durable on-disk companion (`RestoreTestState`, R-189). **It blinds no alarm, and that was CHECKED rather than assumed.** `hub/internal/monitor/deadline.go` scans back over stored reports with a 7-day `backupEvidenceLookback` whose own comment names this exact case — *"when the LATEST report carries none... and against an agent that stayed restarted for days"* — and `pbs_snapshots` stayed populated at 2 regardless. So this is an observability wart, **not** a safety hole, and it is filed at that severity deliberately. **What is actually wrong is the comment.** The `Store` doc says *"Backups are unaffected — their freshness has a ground truth on the storage (R-84)"*. That is true of the **consequence** and false of the **field**, and it sits three lines below a paragraph explaining that the very same sentence about restore-tests *"used to be here and it is now FALSE"* — so the file already carries one correction of this shape and invites the next reader to trust the surviving half. | **READY (XS) — NEW 2026-08-20** | — | Say what is measured: the field IS lost on restart and repopulates only when a backup runs; the freshness VERDICT is unaffected because the hub looks back 7 days. **Name `backupEvidenceLookback` in the comment** so the cross-repo dependency is visible from the agent side — today the agent's claim of safety rests on a hub constant it does not mention. Per the workspace rule, a comment asserting an invariant needs a test pinning it: the pin belongs on the HUB side, asserting the verdict survives a report carrying `backups: []`. | CC | | **R-371** | Monitoring & notifications | P4 | **The off-site tier is the only backup tier that announces nothing on success.** Written down 2026-08-05 in `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md:508-513` and explicitly *"recorded, not filed"*: the off-site run emits **no hub event at all**, while both lesser tiers do (`db_dump_completed`, `crossdrive_completed`). Failures are covered by `backup_run_failures` and staleness by the hub's 8-day tier deadline, which is why it was judged a wrinkle. **Still true 2026-08-22** — the 2026-08-21 drill's own event dump shows `db_dump_completed` and six `crossdrive_completed` rows and no off-site success event. **Age when filed: 17 days.** | **OPEN — LOW** | — | Either emit one, or record deliberately that the highest-value tier is silent on success and say why. | CC | -## Hub & operator — 23 rows (P2 1, P3 7, P4 15) +## Hub & operator — 24 rows (P2 1, P3 7, P4 16) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -384,7 +384,8 @@ stopping line that lies. | **R-719** | Hub & operator | P4 | **[P2-MEDIUM] A customer who already exists never gets a fresh self-bind link when their new box registers: the last link expires in 7 days and nothing re-sends it.** MEASURED 2026-09-29 (new-household drill, `tester-1`): the previous link went out 2026-09-17 07:25 UTC at a host delete and expired 2026-09-24; the box registered at 19:11:30 UTC and its console told the volunteer to open the link from their e-mail — there was none that worked. Hub source: the link is sent at customer creation, RESET, e-mail set on a box-less customer and host delete (`selfbind_mint.go` callers `hosts.go:908`, `configs.go:850`, `customer_reset.go:162`) — never on appliance registration. The volunteer guide says the operator needs to press nothing. The operator pressed „Send self-bind link" (the mail arrived in 1 s) — an operator step the volunteer depends on, recorded, not an intervention. **Fix direction:** send the link when an unclaimed appliance registers while a box-less customer waits with no live link (R-509's first fix shape), or the guide's operator part says: press it the day the volunteer installs. Evidence: `audits/evidence-drill-new-household-2026-09-30/` `phase0/operator-steps.txt`. **CHANGED AND BUILT 2026-09-30 (hub v0.126.0) — the brief's shape was not buildable:** a box registers UNCLAIMED (uuid, MACs, host keys, hardware — nothing of a customer), so "send the link when their box registers" would mail every waiting customer. Built instead: the expired AND used link pages offer „Új linket kérek" → a fresh link to the address registered for that link's customer, only when it has no box, ≤1/h per customer, identical answer for any token (no oracle). Live: the button on the real hub, the same page for a made-up token, no mail; the mint+send path unit-proven (RP42). Evidence: `audits/evidence-fixes-first-tester-2026-09-30/``partD/`. | **WAITING-ON-OPERATOR** (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the operator has not reviewed the changed page shape ("operator may prefer another"), and mint+send is unit-proven only) — **CLOSED 2026-09-30 — hub v0.126.0 (changed shape; operator may prefer another)** | — | — | operator | | **R-814** | Hub & operator | P4 | `PBS-storage-1` (u629193, box 611421) still `status=active`, 19.9 MB | **VERIFY** (2026-10-03 triage: a July watch row with no id; given R-814. WAITING-ON-OPERATOR — no record found that the box was deleted.) — WAITING-ON-OPERATOR | operator console | Delete the box | operator | | **R-844** | Hub & operator | P4 | **The household's OS-update line exists only on the hub's customer timeline.** 2026-10-04: the box itself has no event surface for agent results (the controller UI shows no timeline), so `os_update_applied` is a hub customer event (info: recorded, never mailed). Its stored text is the hub's English sentence; the hu/en bundle text (`mail.event.os_update_applied`) is used only if it is ever mailed. Fix direction: a controller-side line (the controller already polls the agent's local API) when the box gets a household timeline. `audits/os-guest-lane-2026-10-04/partG/hub-customer-timeline-demo-hp.txt` | **READY — owner: CC** | — | — | CC | -| **R-845** | Hub & operator | P4 | **One OS-leg pass takes 3–4 minutes even when it installs 3 packages**: the wrapper's inventory (an `apt-get update`, `apt-cache policy` over every installed package, a `/proc/*/maps` scan for restart-needed, two simulations) dominates; measured 174–245 s per pass on the demo boxes vs 3.8–31.7 s for the install itself. It runs at night under the heavy-op gate, so it delays a restore-test by minutes, nothing worse. Fix direction: one `apt-cache policy` per run and the restart scan only after an install. `audits/os-guest-lane-2026-10-04/partG/` | **READY — owner: CC** | — | — | CC | +| **R-848** | Hub & operator | P4 | **A held host package is invisible to the hub.** MEASURED 2026-10-04 on demo-hp (undo runbook proof): with `tzdata` held after a by-hand undo, the wrapper's `pending` stayed 78 — apt's simulation leaves held packages out, so the fleet view shows nothing and no alarm can see a hold that was forgotten. The hold lives only in the incident's register row (`runbooks/os-updates-host-undo.md`). Fix direction: the wrapper reports `apt-mark showhold` and the fleet line shows it. | **READY — owner: CC** | — | — | CC | +| **R-849** | Hub & operator | P4 | **The fleet view's GUEST "reboot needed since" never clears.** 2026-10-04: the guest is scanned only after an install (R-845, one `pct exec`), so a guest restart is never seen; the guest line keeps the date of the last install that said "needed". No alarm reads the guest line (the reboot alarm is host-only), so it is display only. Fix direction: scan the guest on every pass too (one `pct exec`, ~1 s) or hide the guest date. `audits/os-host-lane-2026-10-04/partC/live/fleet-during-ring1-test.json` | **READY — owner: CC** | — | — | CC | ## Business & legal — 7 rows (P2 4, P4 3) diff --git a/documentation/runbooks/os-updates-host-undo.md b/documentation/runbooks/os-updates-host-undo.md new file mode 100644 index 00000000..10350af6 --- /dev/null +++ b/documentation/runbooks/os-updates-host-undo.md @@ -0,0 +1,81 @@ +# Put one host package back by hand (OS updates, host fast lane) + +> **When:** the hub mailed `os_update_health_failed` for the **host** layer (the box's base system), or the operator +> sees a host problem that started with a host OS pass, and ONE package is the suspect. +> **Who:** the operator, or CC with the operator's word, as root on the box's Proxmox host (`ssh `). +> **Owner design:** `architecture/11-os-updates.md` §5.6 (host row), §8 step 3. **Proved** on demo-hp 2026-10-04 with +> `tzdata` (2026c → 2026b, held; then released and re-installed by the next ring-0 pass) and again on demo-felhom for +> the ring-1 test — evidence `audits/os-host-lane-2026-10-04/partB/undo/` and `partB/ring1/`. + +There is **no automatic undo** for the host. A host cannot be snapshotted the way the old design assumed, and the +fast lane only ever installs Debian / Debian-Security packages, never a kernel, boot or firmware package (wrapper +refusals R14, R12). So the undo is small: install the previous version of the suspect package from +`snapshot.debian.org`, which keeps every version Debian ever published. + +## 1. Find the suspect and its previous version + +```bash +# the host pass's own apt run (Requested-By: felhom-agent); each line "name:arch (old, new)" +grep -B2 -A6 "Requested-By: felhom-agent" /var/log/apt/history.log | tail -20 +PKG=; OLD= +``` + +## 2. Pick the snapshot timestamp + +- **Normal case:** the timestamp of the **previous host release** — the hub fleet view (`GET /os/fleet`, the box's + host line names its release; the release's approval time IS its snapshot time, `YYYYMMDDTHHMMSSZ`). At that time + the old version was the current one. +- **No previous host release** (a ring-0 box, or the first host pass): ask snapshot.debian.org when the **NEW** + (installed) version first appeared, and use that time. At that moment the suite still carried the old version. + **Do not use the OLD version's `first_seen`:** that is when it reached Debian *unstable*, and `trixie` may not + have had it yet — measured on demo-hp 2026-10-04: `eject 2.41-5` at its own `first_seen` → `Version not found`. + +```bash +NEW=$(dpkg-query -W -f='${Version}' "$PKG") +curl -s "https://snapshot.debian.org/mr/binary/$PKG/$NEW/binfiles?fileinfo=1" | python3 -c ' +import json,sys; d=json.load(sys.stdin) +print(sorted(f["first_seen"] for v in d["fileinfo"].values() for f in v)[0])' +# → e.g. 20260831T204404Z +TS= +``` + +## 3. Install the old version from the snapshot, then remove the snapshot source + +```bash +. /etc/os-release +cat > /etc/apt/sources.list.d/felhom-undo-snapshot.list <` → `running`. +- The host health rule (`11` §8.2) is what the leg checks; run the debug action to see it pass: + `sudo -u felhom-agent /usr/local/bin/felhom-agent --config /etc/felhom-agent/agent.json --selftest=os-update -vmid ` + — **note: on ring 0 this also installs every pending fast-lane fix.** A held package is NOT reported as pending + at all (measured on demo-hp: `pending` stayed 78 with `tzdata` held) — **the hub cannot see a hold**, so the hold + lives only in the register row (R-848). + +## What this runbook does NOT cover + +- A kernel: the fast lane never installs one (R14). The kernel lane is `11` §5.6 (not built). +- A Proxmox-origin package (pve-*, qemu-server, …): never installed by the fast lane (R12/origin rule). Proxmox's + own repository keeps old versions; that undo is not written yet. +- The customer guest: its undo is last night's whole-guest backup (decision 81).