From a55eedcf2cc182101ebd8de7adbf4283470eda38 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Sun, 4 Oct 2026 11:30:14 +0200 Subject: [PATCH] CHANGELOG + REPORT: v0.140.0 released (OS updates, guest fast lane) Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- CHANGELOG.md | 29 +++++++++++++++++++++++++++++ REPORT.md | 17 +++++++++-------- 2 files changed, 38 insertions(+), 8 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 186923b..4620689 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,3 +1,32 @@ +## v0.140.0 — OS updates, guest fast lane (`11-os-updates.md` §8 step 2; `09` §3 decisions 76, 79, 80) + +> **RELEASED 2026-10-04** by `scripts/release-agent.sh` — tag `v0.140.0` (`9cac346`), sha256 +> `ae2d60b794869c51d6b063c8e31e2da98ecdbd36c6b75febd64e2fb4266c1250`, verified by download. Not vouched at release time. + +**MinAgent impact:** none required by any controller. **Needs hub v0.130.0** (`os-report`, the `os_update` block); an +older hub serves no block and the leg then reports and installs nothing (ring 1, no release). + +- **`configs/felhom-os-apply`** — the root wrapper (Python 3, stdlib). One sudoers entry, `FELHOM_OSAPPLY`: + `felhom-os-apply --plan /var/lib/felhom-agent/os/plan-*.json`. Modes `inventory` / `apply` / `health`. Refuses (exit + 2, nothing changed) on R1–R13: plan path/owner/JSON, a non-Debian origin, the slow lane, any removal, any downgrade, + a new or unlisted package, a version not downloadable even from the snapshot, low space, a lock (apt or a guest lock + such as a backup), a vmid that is not the box's own customer guest (it must bind `/mnt/felhom-drives`), malformed + names/versions, the host layer, dpkg still broken after the repair. Repairs first (`dpkg --configure -a`, + `apt-get -f install`). A version Debian already replaced comes from `snapshot.debian.org` at the approval time + (decision 79). Reports the full installed set with origins, pending, restart-needed (outside containers), health. + `configs/test_felhom_os_apply.py`: 35 tests; every refusal red-proved. +- **`internal/osupdate`** — the leg: after a SUCCESSFUL primary whole-guest backup, still holding the heavy-op gate + (never beside another backup or a restore-test), once per night, 90 s after the backup. Ring 0 installs every pending + Debian / Debian-Security fix; ring 1 exactly the hub's newest approved release; switched OFF → reports only. The + health rule: docker answers, the network resolves, the controller is healthy, every container running at the START + of the leg runs (and is healthy if it was) — a 5-minute wait. **No automatic undo:** a customer guest cannot be + snapshotted (R-837). A failure is `health_failed` → the hub mails the operator. +- **`--selftest=os-update -vmid N`** — the debug action (trigger `debug`, never throttled, not a night run). +- **The `--selftest` flag also accepts `wgtunnel`** — it was dispatched but refused since S3 (found by the new + `TestSelftestFlag_AcceptsEveryDispatchedMode`, which also caught `os-update` live). +- Proven live 2026-10-04 on both demo boxes (ring 0: 53 packages each; ring 1: exactly 3 approved versions; a failed + health check → `health_failed`, operator mailed): `felhom.eu/documentation/audits/os-guest-lane-2026-10-04/`. + ## v0.139.0 — a DR restore never lands beside a live original (2026-10-04, R-834) > **RELEASED 2026-10-04** by `scripts/release-agent.sh` — tag `v0.139.0` (`475bdce`), sha256 diff --git a/REPORT.md b/REPORT.md index 10c3c3c..c4d24ad 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,10 +1,11 @@ -# REPORT — 2026-10-04: v0.139.0 (R-834) +# REPORT — 2026-10-04: v0.140.0, OS updates (guest fast lane) -Full session report: `felhom.eu/REPORT-backup-close-os-spike-2026-10-04.md`. +Full session report: `felhom.eu/REPORT-os-guest-lane-2026-10-04.md`. -- **Measured** on demo-hp: the scheduled restore-test's scratch guest has `onboot: 0` and throwaway stand-ins for - mp8/mp9 on every config read until teardown — it was already safe. Evidence `felhom.eu/documentation/audits/backup-close-2026-10-04/partA/`. -- **Fixed:** the DR bring-up refuses beside a live original (source guest present, a drives bind on another guest, - or an unreadable config). On a replaced host it is unchanged. -- Tests `TestRunBringUp_DRRefusesBesideALiveOriginal`, `TestRestoreTest_NoHostPathBindBesideTheOriginal` and two - more; both rules red-proved. No sudoers change (said why in the CHANGELOG). +- `felhom-os-apply` wrapper (R1–R13, repair first, snapshot.debian.org fallback), `FELHOM_OSAPPLY` sudoers, the OS leg + after the primary backup, `--selftest=os-update`. Released `9cac346`, sha256 `ae2d60b7…1250`, verified by download. +- Live on both demo boxes: ring 0 installed 53 packages each, healthy; ring 1 installed exactly the 3 approved versions; + a deliberately failed health check reported `health_failed` and mailed the operator. +- Found and fixed live: the `--selftest` flag refused `os-update` (and `wgtunnel`, since S3); the conffile log line + called an updated file "kept"; an app stopped between the inventory and the apply escaped the health check. +- No automatic undo (R-837: PVE refuses a snapshot of a guest with host-path binds).