## Unreleased — part of v0.153.0: ring 0 stages exactly the told kernel (R-898; `09` §3 decision 176) (2026-10-07) **Delivery: the agent binary only** — no root file changed (the wrapper is unchanged; its tests gained two cases). - `internal/osupdate/kernel.go`: ring 0's night kernel step stages EXACTLY the kernel the household was told about (select `listed`, `KernelSet(kver)` = the series meta-package and the signed image at the kernel's own version) instead of "whatever is pending tonight". Seen 2026-10-07: demo-felhom was told about 7.0.14-20 while its sources offered 7.0.14-22 by night — the old code staged `pending-kernel` and the wrapper refused it (R23), losing the night. A told version that is no longer installable is refused by the wrapper before any change (R7) and the hub tells the household again for the newer kernel (hub v0.143.1). Tests `TestKernel_Ring0ToldNightStagesThenReboots` (red-proved against the old select), `TestKernelSet`; wrapper `test_ring0_listed_installs_the_told_kernel_not_the_newest`, `test_ring0_told_kernel_gone_is_refused_before_any_change`. ## v0.152.0 — the kernel lane (R-836; `09` §3 decisions 164, 172; `11` §5.11) (2026-10-07) Released by `scripts/release-agent.sh`: binary sha256 `95ff42208e36ba49b6e2b09a97e81a6fa11562ecc8042f1ed18d378b8b1f88b1`, config bundle sha256 `f0c2cec374b711b3c131c955012b33fd0ad495337d049b7a743c3eef9e85c20b` (tag `v0.152.0` = `d03ab7f`). Step bundle `0.152.0-step1` sha256 `0b71d32b054cf3b7ade0234ffcbb0df159901f542cde540adaee411db466f48e` (the 0.151.0 bundle with only `felhom-os-apply` replaced; `scripts/build-step-bundle.py`), published as package version `0.152.0-step1`. **Delivery: agent binary, then the STEP bundle `0.152.0-step1`, then the bundle `0.152.0`** — the bundle ADDS two paths (the GRUB generators), and an installed `felhom-os-apply` refuses a path its own table lacks (R16, R-880). A new kernel boots ONCE; if it crashes the box comes back on the old kernel by itself; it becomes the default only after a healthy boot; a booted-but-unhealthy kernel is reverted ONCE by the agent with no person. Built on the spike's candidate 2 (`audits/kernel-spike-2026-10-07/`), with option C on the one-shot entry. - `configs/felhom-grub-oneshot.sh` → `/etc/grub.d/01_felhom_oneshot` (bundle): reads `felhom_next` from a GRUB env block on the ESP (`EFI/felhom/oneshot.env`), clears and saves it BEFORE the menu, and sets the default to that kernel's one-shot entry only when the name is an installed kernel. No vfat ESP → prints nothing. - `configs/felhom-grub-oneshot-entries.sh` → `/etc/grub.d/42_felhom_oneshot` (bundle): one entry per installed kernel, id `felhom-oneshot-`, the normal entry plus `softlockup_panic=1 hardlockup_panic=1 hung_task_panic=1 panic=10` (option C). Sorted after `10_linux`: never entry 0, never the default. - `configs/felhom-os-apply`: layer `kernel` (lane slow; an appliance; authority = a signed `os_kernel_step` or the root-owned ring-0 mark). Modes: `apply` STAGES (select `pending-kernel` or a signed `listed` set; `expect_kver` = the kernel the household was told about): pins the GRUB default to the RUNNING kernel in `/etc/default/grub.d/zz-felhom-kernel-default.cfg` and proves it from grub.cfg, installs, proves the default did not move and the one-shot entry exists, writes the flag; never reboots. `kernel-reboot` (a staged step only), `kernel-boot` (judging | fell_back | self_reverted | revert_failed), `kernel-good` (the new kernel becomes the default, proved), `kernel-revert` (ONE per step; refused when the default is not the old kernel), `kernel-cancel`, `kernel-status`. New refusals: R20 (the box cannot do a one-shot: not UEFI, no vfat ESP, a separate /boot, GRUB without fat/loadenv, the generators missing, a hand pin), R21 (the crash guard tripped or an unclean boot in its window), R22 (the phase does not allow the mode; never two steps within 20 h), R23 (not exactly one newer kernel, or not the one signed / told). State `/var/lib/felhom-kernel/state.json`. Facts carry `kernel_lane`; the next-boot kernel reads the flag and the grub.cfg default. Tests: `KernelLane` (27), red-proof `felhom.eu/documentation/audits/kernel-lane-2026-10-07/A/redproof.txt`. - `configs/felhom-crash-guard` unchanged; `KernelStepCannotLeaveTheBoxOff` (3 tests) pins that a step's planned reboot, one crash and one self-revert add ONE unclean boot (a panic before userspace adds none), so the box cannot stay off. - `internal/osupdate/kernel.go`: the night leg ends with the kernel step — after a healthy host step (and a healthy Proxmox step when one ran), trigger `night` only, on a night the hub's `os_update.kernel` block marks `tonight` (the household was mailed the day before — no mail, no step). Ring 0 stages + reboots; ring 1 reboots only a kernel a signed `os_kernel_step` staged (`KernelStepExecutor`: stage only, under the heavy-op gate). The hub hears `staged` BEFORE the reboot. At every start `KernelAfterBoot`: on the new kernel it JUDGES the boot — `KernelVerdict` = the host health rule (`11` §8.2) AND the box reached the hub (the `judging` report itself) — for 20 minutes (measured: everything healthy 68 s after the reboot on demo-felhom, 272 s on demo-hp; under the hub's 45-minute `host_stale` (`alerting.stale_threshold`)). Healthy → `kernel-good`, outcome `applied`; not healthy → outcome `health_failed`, then ONE `kernel-revert`. Tests: `TestKernel*` (13); red-proofs in the same file. - `internal/hub`: `WireOSUpdate.Kernel` {kver, tonight, notified_at}. `internal/reconcile`: `os_kernel_step` is destructive-class. `cmd/felhom-opsign`: the op is listed. ## v0.151.0 — the agent can no longer hand the guest any image; the Proxmox package lane; the other-key archives reported; the DR directive retired (R-861, R-812 A, R-366, R-105; `09` §3 163, 165, 168, 169) (2026-10-07) Released by `scripts/release-agent.sh`: binary sha256 `0464354f2cdf452a7c5d2a74d9191fe91415fcfa244154480d26d5b30e10b194` config bundle sha256 `bacd1d175a9392bc1755341d01a106abbd72aba48ab575e7a32dd819d8f5da4c` (tag `v0.151.0` = `dd7cdc0`). **Deliver the binary FIRST, then the bundle:** the bundle's sudoers removes the `tee` grant the 0.150.0 binary still uses. No path added (26 → 26), so no step bundle. ### Part of v0.151.0 — the agent can no longer hand the guest any image; felhom-op's pct lines are exact (R-861 (a) A1, (b) B2; `09` §3 decision 165) **Delivery order: agent binary FIRST, then the config bundle.** The new sudoers drops the agent's in-guest `tee` grant; an older binary still calls `tee`, so a bundle that lands before the binary would stop managed controller updates (and the old binary's capability probe would read `controllerswap-write` degraded). No bundle path is added (`felhom-priv-apply` and `/etc/sudoers.d/felhom-op` are already bundle files), so no step bundle. - `configs/felhom-priv-apply`: new verb `controller-image ` — reads the ref on stdin (≤ 256 bytes, ASCII, one optional trailing newline), requires `^gitea\.dooplex\.hu/admin/felhom-controller:[0-9]+\.[0-9]+\.[0-9]+$` (the agent's own `controllerImageRe`), then runs `pct exec -- tee /etc/felhom-controller-image` AS ROOT; refusal rule `I1` (rc 3), a bad vmid `A1` (rc 2); listed in `--self-check`. - `configs/felhom-agent.sudoers` `FELHOM_CONTROLLERSWAP`: `pct ^exec [0-9]+ -- tee /etc/felhom-controller-image$` REMOVED; `/usr/local/sbin/felhom-priv-apply ^controller-image [0-9]+$` added. - `internal/localapi`: `GuestExecutor.GuestExecStdin` replaced by `WriteControllerImage`; `GuestBinder.WriteControllerImage` pipes `ref\n` to the verb through the fenced runner; the swap's `writeImage` calls it. Capability `controllerswap-write` now probes the verb. - `configs/felhom-op.sudoers` (B2, hygiene): `pct start|stop|unlock [0-9]*` → `pct ^start [0-9]+$` etc. (the glob's `*` matched spaces: `pct stop 9201 --skiplock 1` passed). - Tests: `ControllerImage` (5, `configs/test_felhom_priv_apply.py`), `TestSudoersRefusesTheR861Injections` (+3 lines), `TestSudoersAllowsTheControllerImageVerb`, `TestFelhomOpSudoersPctIsExact`, `TestR861_WriteControllerImageUsesTheRootVerb`, `TestControllerSwap_WriteViaRootVerb_NoShell`. Red-proofs: `felhom.eu/documentation/audits/day-2026-10-07/C/`. - `README.md`: the controller-swap paragraph described the removed `tee` path — corrected. ### Part of v0.151.0 — the Proxmox package lane (R-812 option A, `09` §3 decision 163) **MinAgent impact: none** (a new layer; an older hub ignores the pve report). **The bundle carries the new `felhom-os-apply` — deliver it with the binary** (signed `agent_update`, then signed `agent_config_update`). - `configs/felhom-os-apply`: new layer `pve`, lane `slow` only — the host's Proxmox USERSPACE packages: origin `Proxmox Debian Repository` only (R2), never a kernel / boot / firmware / microcode name (R14, `HOST_SLOW_RE` — the kernel is R-836's lane), no removal (R4), no undo (R5), a new package only from `PVE_NEW_ALLOW` (`proxmox-firewall-data`, measured on demo-felhom; R6 otherwise), an appliance only (R12), authority = a signed `os_pve_step` or the root-owned ring-0 mark (R3). Select `pending-pve` (ring 0): installed Proxmox-origin packages with a pending upgrade. The report carries `pve_manager` (pveversion after the step). - `internal/pvegate` (new): the agent's own writes to /etc/pve wait while a pve step runs (pmxcfs restarts); the step waits for writes in flight (bounded, 2 min — then it fails and does not run). Wired at `proxmox.Client.doBody` (every non-GET) and `ExecRunner.RunStdin` (`WritesEtcPVE`: pct config verbs, pvesm, pveum, felhom-pbs-apply create/reconcile). - `internal/osupdate`: `LayerPVE`; the night leg runs the pve step in ring 0 after a healthy host step (an appliance; ring 1 never in the night leg); `PVEHealthVerdict` = the host rule + every running container keeps its id + pveversion reads the installed pve-manager; the pve report carries Proxmox userspace only (the hub's candidate set). `PVEStepExecutor` (signed `os_pve_step`, ring 1, under the heavy-op gate and the /etc/pve gate); `reconcile.ClassOSPVEStep` (destructive-class); `felhom-opsign -op os_pve_step` (params by `-params`). - Tests: wrapper `PVELane` (17; red first — the `pve-manager` plan was refused R12 on the old code), `pvegate` (5), `TestPVEGate_*` + `TestWritesEtcPVE`, `TestPVE_*`, `TestPVEHealthVerdict`, `TestPVEStepExecutor_*`. Red-proofs: `felhom.eu/documentation/audits/day-2026-10-07/B/`. ### Part of v0.151.0 - R-366 slice 2 (`09` §3 decision 168): the restore-test pick records, per tier, the archives it skipped as written with another key (count, oldest, newest — no key material) in a `ForeignKeyLedger`; the host report carries it as `foreign_key_archives.tiers` (the stanza absent until a tier was evaluated since start, `tiers: []` when none — no null on the wire, the report contract forbids it). The hub turns a change into one operator line. Tests `TestR366_PickRecordsArchivesWrittenWithAnotherKey`, `TestR366_EvaluatedWithNoneIsAnEmptyList` (red-proved, `felhom.eu/documentation/audits/day-2026-10-07/E/`). - R-105 option A (`09` §3 decision 169): the `--selftest=escrow-create -directive ` flag and the escrow upload's `directive` field are removed — nothing read the directive; the DR path reads the recipe, tenantsync and the escrow blob. The hub ignores a `directive` from an older agent. ## v0.150.0 — the Docker step proves the engine reports a memory kill; after a restart the agent remembers the last backup per tier; three more SMART counters on the wire (R-528, R-894, R-330; `09` §3 decisions 157, 161) (2026-10-07) Released by `scripts/release-agent.sh`: binary sha256 `a23d1c9085bc7fd4fc48fe0327f6504aa83e6331510fb4a3d23e042dddb26f9c` config bundle sha256 `88456b386d9b1027bd22861cac8c23df004bf9fd9f67644d6595bfca8c94498e` (tag `v0.150.0` = `3a72a48`). **The bundle carries the new `felhom-os-apply` (the memory-kill check) — deliver it with the binary:** signed `agent_update`, then signed `agent_config_update`. No path added (26 → 26), so no step bundle. ### Part of v0.150.0 (2026-10-06 night, later) — after a restart the agent remembers the last backup per tier (R-894); three more SMART counters on the wire (R-330) Ships with the memory-kill check below as v0.150.0, AFTER the 2026-10-07 night read-back. Nothing delivered tonight. - **The defect (measured 2026-10-05 on demo-hp):** the agent restarted at 04:57; at 06:25 the off-site storage answered *Can't connect*; the per-tier backup record is in memory only, so the due-check fell back to an EMPTY record and the 7-day tier (last copy 4 days old) read DUE; the controller asked and vzdump failed. - New `internal/backup/backup_state.go` `BackupSuccessState`: the newest SUCCESSFUL backup per tier and guest, on disk (`/backup-success-state.json`, atomic tmp+rename, 0600). Only successes are written; a corrupt file reads as nothing known. - `internal/localapi` `handleBackupDue`: when the tier's storage CANNOT be read, the saved copy stands in for the in-memory record. A fresh copy → not due („… (storage unreadable — age from the last success saved on disk)"); a copy older than the cadence → DUE; no copy → the old answer (DUE, age unknown). A storage that answers stays the ground truth: an archive absent there is due even when the file remembers one. - Wired in `buildLocalAPIServer` (`LastKnownBackups`); the local API's backup job saves each success. - Tests: `TestBackupDue_R894_*` (restart = a new server and a new state from the same file; fresh / old / none / storage answers / failed backup not saved), `TestBackupSuccessState_*`, `TestR894_LastKnownBackupsIsWiredIntoTheDaemon` (AST). Four red-proofs observed (`felhom.eu/documentation/audits/night-burndown-2026-10-06/s4/`). - **R-330 (disk health Phase 2, the wire only):** the SMART summary carries three more SATA raw counters — `reported_uncorrect` (187), `command_timeout` (188, carried as the vendor reports it; some pack several counters), `udma_crc_errors` (199). Pointer + omitempty: an attribute the drive does not report is OMITTED (unknown), never 0. No verdict reads them yet. Tests `TestParseSMART_R330_*` (two red-proofs, `felhom.eu/documentation/audits/night-burndown-2026-10-06/r330/`). ### Part of v0.150.0 (2026-10-06 night) — the Docker step proves the engine reports a memory kill (`09` §3 decision 157, R-528) To be released as v0.150.0 with its config bundle AFTER the 2026-10-07 night read-back (the night of 2026-10-06 runs v0.149.0 on purpose). - `configs/felhom-os-apply`: after a docker-layer APPLY (after `health_after`) the wrapper runs `oom_check()`: a throwaway container from the image the running controller uses (`--pull never`, `--network none`, no volume, label `felhom.oomcheck=1`, 64 MB cap) asks for one 200 MB block; „pass" only when `OOMKilled=true` AND the `oom` event; it waits 2 s and reads the events window to the guest's epoch + 1 (measured: a window closed in the same second missed the event); the container is always removed. Reported as `oom_check`; it never changes the step's outcome or health. A wrapper-only mode `oom-check` runs the check alone (no apt, no engine change), by hand as root. - `internal/osupdate`: `WrapperReport` and `Report` carry `oom_check` verbatim, on the normal pass and on the kept-copy path (R-868). - `configs/test_felhom_os_apply.py`: its `unittest.main()` sat in the middle of the file, so 11 tests (UnsentReport, SaveReportOnDisk, AgentDiesMidPass, CrashLeftTheJournal) never ran — moved to the end; all pass. - Tests: the OOMCheck class (pass, OOMKilled=false, no event, unreadable image, removal on an inspect error, not on other layers or in health mode, the events window after the settle wait, mode oom-check alone and its refusals); TestDocker_OOMCheckReachesTheHubUnchanged, TestR868_KeptCopyCarriesTheOOMCheck. 12 red-proofs in `felhom.eu/documentation/audits/readback-2026-10-07/F/`. ### Part of v0.150.0 (2026-10-06 evening) — the shared rule file (`09` §3 decision 152); no code change - `.claude/rules/unprompted-work.md` added, byte-identical to the copies in felhom.eu, felhom-controller, app-catalog-felhom.eu and the workspace root (checked with `diff` against the controller's copy and one md5 across all five). Its copies line names five copies. ### Part of v0.150.0 (2026-10-06 afternoon) — instruction files kept true (`09` §3 decision 150); no code change - `CLAUDE.md` „Gates — ONE entry point": the runner runs every gate in its `GATES` table (five: three shared, `published`, `release-complete`); `--fast` skips `published` (network). It said two gates and „all of them". - `CLAUDE.md`: the decoy gate and its audit are named with their `felhom.eu/` prefix (they do not exist in this repo). - `.claude/rules/health-checks.md` (comment): the health-check rule's copies live in felhom.eu `hub.md` and the controller's `gates.md`; it named felhom.eu `CLAUDE.md` „Code quality rules", which holds no such rule. ## v0.149.0 — a weekly disk trim of each customer guest, the crash-boot fact for the controller, the phantom WARN names its runbook (R-444, R-856, R-99; operator rulings `09` §3 139, 143, 140) (2026-10-06) Released by `scripts/release-agent.sh`: binary sha256 `6bcae9c2eb5d97e8285316583870059835793893299e291891a53a4ce505585f` config bundle sha256 `e182c82dcf4a67faa3bcb74dbe4ffa7b06e0b27dc8451cb7574d6339ce91ad66` (tag `v0.149.0` = `f277e61`). **The bundle carries the new sudoers rule for the trim (`FELHOM_FSTRIM`) — deliver it with the binary:** signed `agent_update`, then signed `agent_config_update`. - R-856 (`09` §3 decision 143): new local-API route `GET /host/crash-guard` — passes the host crash guard's last-boot record (present, last_boot_at, last_boot_unclean, tripped) from /var/lib/felhom-crash-guard/state.json to the controller, which waits ~15 min with app mails after a crash boot. Read-only, no Proxmox call, guest-token authed; a missing/unreadable/garbled file answers 200 present:false (never an error page). An older agent answers 404, which the controller reads as unknown (normal 90 s grace) — no controller MinAgent raise needed. - R-444 (`09` §3 decision 139): weekly guest disk trim. New sudoers alias FELHOM_FSTRIM with ONE exact rule `/usr/sbin/pct ^fstrim [0-9]+$` (rides the signed config bundle; decoys pinned by TestSudoersFstrimRuleIsExact) and capability guest-fstrim (non-critical). New internal/fstrim job: each owned RUNNING guest gets `pct fstrim ` once a week - due Wednesday from 10:00 host-local, starts only 10:00-20:59 (never the 01:00-06:59 night), holds the one-heavy-op gate so it never runs beside a backup or restore-test (busy -> deferred to the next hourly tick; a box that was off catches up at its next daytime hour); a failed trim WARNs and is retried at most 3 times that week; bytes parsed from `pct fstrim`'s "(N bytes) trimmed" lines; positive log `fstrim: guest N trimmed X GiB in Ys`; last result per guest persisted in /guest-disk-trim.json and reported as the new omitempty host-report stanza `guest_disk_trim`. Opt-out: agent.json "disk_trim": {"disable": true}. - R-99 (`09` §3 decision 140): the agent's WARN for a PBS archive below the 1 MiB plausibility floor now ends with the pointer to the sanctioned cleanup (`documentation/runbooks/pbs-phantom-cleanup.md`); detection only — nothing is deleted automatically. Dir-storage archives keep the old text. ## unreleased - R-426: scripts/test_gate_decoys.py (new) — the published gate judged against a fake Gitea (127.0.0.1, via GITEA_BASE; never the real registry): 11 cases; COVERS published — `felhom-agent/published` leaves the decoy-coverage EXEMPT list. - R-426: release-complete gate — 10 decoy cases (scratch clone + scratch bare origin + fake Gitea); COVERS release-complete — `felhom-agent/release-complete` leaves the decoy-coverage EXEMPT list. - R-426: the shared reuse-refs/instructions/observations gates get agent-side decoys (9 cases on a scratch clone of this repo); COVERS reuse-refs, instructions, observations — three `felhom-agent/*` entries leave the decoy-coverage EXEMPT list. - **Fixed without a row:** `configs/test_felhom_config_bundle.py` read the two ISO first-boot files that installer 1.32.0 now NAMES under KEPT (R-275) as files the installer writes — `go test ./internal/osupdate` was red on DooPlex from 21:25 to 01:55 (felhom.eu `85de3f9b`); they are listed with why. Test only; agent v0.148.0's code is unaffected. ## v0.148.0 — the host report names the running binary's sha; the format answer carries the new filesystem's UUID (burn-down night: R-349, R-25 agent halves) (2026-10-06) Released by `scripts/release-agent.sh`: binary sha256 `3e68a0870e0e2ce262cb4819294edddeb0a73e8c558a31611a20329a6d9ee283` config bundle sha256 `a6fa4f589d184b58c9911303bd087e300be1e75b3647e4302c9594df6989c4de` (tag `v0.148.0` = `861d32a`). Delivery order as for v0.147.0: signed `agent_update`, then signed `agent_config_update`. - **R-349:** the host report carries `agent_sha256`, the sha256 of the running agent binary (read once from `/proc/self/exe`; empty = unknown), so a hand-built binary under the vouched version name becomes visible. The hub comparison is a separate hub change. Test `TestCollect_AgentSHA256IsTheRunningBinary`; red-proved. - **R-25 (agent half):** `POST /disks/format` and `GET /disks/format/status` return `fs_uuid`, the new filesystem's UUID read back after mkfs only when the bound durable id still resolves to the formatted device and the superblock is the requested type (empty = not verified); `DeviceProbe` gains `FSUUID` from blkid. Tests `TestFormat_*FSUUID*`; three red-proofs. The controller half (mount that UUID) is a controller change. ## v0.147.0 — the recovery recipe spells the root namespace the way PBS does; a removed drive no longer shows the root disk's size; a rotated-out token stops at once; the dnsmasq check looks at the right package (burn-down round 2: R-124, R-118, R-269, R-317) (2026-10-05) Released by `scripts/release-agent.sh`: binary sha256 `642c4d196c48671c14ff653118303c5903af1670b7abb550edeaf73e701cd5b8`, config bundle sha256 `326527d0993c9a62df2f790c7700ca645cedbf0673dcfb6dc1768d8610b8007d` (tag `v0.147.0` = `f1b9b41`). Delivery order as for v0.146.1: signed `agent_update`, then signed `agent_config_update`. MinAgent impact: none (the controller needs nothing new from this agent). Config bundle content unchanged from v0.146.1. - **R-124 (operator ruling 2026-10-05: fix it):** the DR recipe's `pbs.namespace` for a box in PBS's ROOT namespace is now `""` — PBS's own spelling — beside `namespace_state: resolved`; it used to be the word `root`, which no namespace is named, so `--ns root` failed in a recovery. `hub.PBSRootNamespace`; `TestR124_RootNamespaceOnTheWireIsPBSSpelling` (red-proof: back to "root" → FAIL). Runbook: `felhom.eu runbooks/ep0-datastore-copy.md` step 2 says how to read it (and to treat a recorded `root` from older agents as empty). The hub stores the recipe raw; its fixture follows. - **R-118:** the local API's drive list reads a drive's capacity only while its DEVICE is present — with the device gone the bare mountpoint is a directory on the root filesystem, whose size was reported as the drive's. `TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity` (red-proof convicts). - **R-269:** the token store re-reads its shared file whenever it has grown, BEFORE answering — so a token rotated out by another process stops authorizing on its next use (it used to keep working until an unrelated miss). One `stat` per call. `TestTokenStore_RotatedOutTokenRejectedFirst` (red-proof convicts). - **R-317:** the LAN resolver decides whether to install `dnsmasq` by its service UNIT, not by `/usr/sbin/dnsmasq` (which the `dnsmasq-base` package also ships). Same `apt-get install` command; no sudoers change. `TestEnsureDnsmasq_*` (red-proof convicts). Red-proofs: `felhom.eu/documentation/audits/burndown2-2026-10-05/agent-red-proofs.txt`, `r124-red-proof.txt`. Also in this release (no binary effect; from burn-down round 1): - **R-291:** `scripts/retention-policy.json` names where its 10 comes from — the R-267 newest-10 prune of generic packages, established 2026-08-10 (R-287) — instead of „observed, no located ruling"; the non-existent `registry-retention.md` reader is dropped. `check-published-versions.py` still reads 10 (checked). - **R-348:** `internal/backup/store.go` no longer says backups are „unaffected" by a restart: the reported backup list reads 0 until the next backup runs; only the hub's verdict (7-day look-back) is unaffected. ## v0.146.1 — R-861 review fixes: the signed update flips a root-owned copy; no Wants=/continuations in mount units; the escrow read follows no symlink anywhere (2026-10-05) Released by `scripts/release-agent.sh`: binary sha256 `badd6c9a2e40c8bfe856d2d1a203443b21b7eb92ecc35d6090ab44518d4d082a`, config bundle sha256 `42333e969028867ad8142335e6c1bc4040eec231de0d8d330c2d4b2cf7bc3442`. **Supersedes v0.146.0, which was released but never vouched or delivered to any box.** The same order applies: signed `agent_update` first, then the signed `agent_config_update`. **Delivery needs a STEP bundle (R-880, found while delivering).** An installed `felhom-os-apply` checks an incoming bundle's paths against its OWN table (R16), so every box on the v0.145.0 bundle REFUSES the v0.146.1 bundle (it adds 4 paths). `scripts/build-step-bundle.py` builds the transition: the box's current bundle with ONLY `felhom-os-apply` replaced (same paths — the old wrapper accepts it), published as bundle version `0.146.1-step1`; then the release's own bundle. Order on a box: `agent_update` 0.146.1 → `agent_config_update` 0.146.1-step1 → `agent_config_update` 0.146.1. Tests `StepBundle` (the R16 refusal reproduced; the step accepted; exactly one file changed). Tooling only — not in the binary or the bundle. A background security review of the v0.146.0 commit found three holes in the new code; each is fixed and red-proved (`felhom.eu/documentation/audits/hub-safety-2026-10-05/partF/red-proof.txt`, S1–S3): - **S1 — a race in the signed update.** `felhom-os-apply` hashed the agent's staged file and then let the A/B wrapper copy it BY PATH; the agent owns that directory and could swap the file in between. Now the root step reads the file ONCE (`read_staged_once`: O_NOFOLLOW, fstat, owner, size), hashes those bytes, writes them to a root-owned directory (`/var/lib/felhom-os-apply/agent-update/`) and hands ONLY that copy to `felhom-selfupdate-guarded apply`, which now refuses any other directory, a symlink, or a file not owned by root. Tests: `AgentUpdate` (+1), `SelfupdateWrapperConfinement`. - **S2 — an allowlist escape in `felhom-priv-apply`.** `[Unit]` accepted `Wants=`/`Requires=`/`Before=` naming any unit, so a mount unit could start e.g. `reboot.target`. `[Unit]` now holds only `Description` and `After=local-fs-pre.target` (what the renderers write), and any line ending in a backslash (a systemd continuation this parser would read differently) is refused. Tests `test_U2_wants_starts_another_unit`, `test_U2_continuation_line`. - **S3 — a path traversal in the escrow read.** `O_NOFOLLOW` guards only the last component; a symlinked DIRECTORY in the agent's own state dir still redirected the root read. `readStagedNoFollow` now walks the path from `/` with `openat(O_NOFOLLOW)` per component. Test `TestAttach_RefusesASymlinkedDirectory`. ## v0.146.0 — the agent's root grants narrowed: exact sudo patterns, a root content checker, fixed files from the bundle, the signed update checked as root (R-861) (2026-10-05) Released by `scripts/release-agent.sh`: binary sha256 `b860af465076041e07f35fed1b12d64ae2b2985d8995f0ce167418d39c2b00d5`, config bundle sha256 `161c737e523aa7910cf32ce41b83f989569bee55b8c5938e7211c92aef68548e`. **Order on a box: the signed `agent_update` FIRST (the old bundle still grants the old flip), then the signed `agent_config_update`.** Between the two (minutes) the new agent's checker calls are refused and retried; nothing is lost. After the bundle, an agent BELOW 0.146.0 cannot update itself on that box any more (the unsigned flip grant is gone) — deliver both together. Design: `felhom.eu/documentation/architecture/03-host-agent.md` §3.1 (new). Measured before the change (real sudo 1.9.16, a throwaway container): the v0.145.0 sudoers let **23 of 29** attack command lines through; v0.146.0 lets **0** through and still allows all **64** commands the agent's capability check uses. - **Exact patterns.** A sudoers `*` in the arguments also matches spaces: `pct set [0-9]* -onboot 1` matched `pct set 100 --dev0 /dev/sda -onboot 1` (a raw host disk for a guest), `mount --bind /mnt/*/felhom-data /mnt/felhom-drives/*` matched a `..` path onto `/etc/sudoers.d`, `nft add element … *` took a chained `; flush ruleset`. Every varying argument list is now a sudo regex (`^…$`): one value per slot, a fixed character set, no `..`, no extra argument. `TestSudoersRefusesTheR861Injections` (29 attacks) + `TestManifestCoveredBySudoers` (regex-aware now). - **`felhom-priv-apply`** (new root wrapper, in the bundle). A systemd mount/automount unit, a dnsmasq drop-in, the WireGuard config and the OOB sshd config + felhom-op key reach their root-read places only through it: fixed source, fixed destination, CONTENT checked against what the agent's renderers write (no `[Service]`, `Where=` only `/mnt/` or `/mnt/felhom-drives/` and equal to the unit name, no `bind`/`suid`; a network share must carry `nosuid,nodev`; no `dhcp-script=`; no `PostUp=`; the sshd config only the one template with its Port). Its 30 tests (`configs/test_felhom_priv_apply.py`) + Go contract tests feeding each renderer's real output (`internal/privapplytest`). Pre-flight: every live file on both demo boxes reads OK. - **NFS/SMB options gain `nosuid,nodev`** (a set-uid file on a server outside the box never acts on the host). - **Fixed files from the bundle.** The guest pre-start hook (`/var/lib/vz/snippets/felhom-guest-hook.sh`, run as root at every guest start) and the shared drive parent script + unit are bundle files now (byte-identical to the agent's constants, pinned). The agent no longer installs them from `/tmp`; it checks them (`guesthook.SnippetReady`, `ensureSharedParentBoot`) and only registers / enables. - **The signed update is checked as root.** `felhom-os-apply` mode `agent_update` verifies the operator signature (root-owned signers, this host, the window, the nonce), re-hashes the staged binary against the SIGNED sha, then runs the A/B flip; `felhom-selfupdate-guarded apply` is no longer in the agent's sudoers. 7 tests (`AgentUpdate`). - **The root escrow run reads no path from the agent's config.** As root it pins the PVE secret dir and the WireGuard state dir to their defaults, refuses a storage id that is a path, and reads its two staged files without following a symlink (`readStagedNoFollow`) — before, a symlink in the agent's own directory sealed any root file into the blob. - **Not narrowed here (named in `03` §3.1):** `FELHOM_CONTROLLERSWAP` stays guest-scoped (a compromised agent can run a chosen controller image in the guest — the household's data, not host root); `FELHOM_ESCROW` still hands the agent R by design (the agent relays the ceremony); the mkfs / pbs-apply / backup-target wrappers keep a coarse argument and their own checks. - Red-proofs F1–F9: `felhom.eu/documentation/audits/hub-safety-2026-10-05/partF/red-proof.txt` (F1's first run did NOT convict — the name rule masked it — and the test now uses the pair only the Where rule stops). ## v0.145.0 — the OS update repairs itself after a power cut; a short-session box gets restore-tested; "sent late" (R-876, R-874, R-875) (2026-10-05) Released by `scripts/release-agent.sh`: binary sha256 `894da35c7b9e1ac78885690b78352b634e6831e7d99b321573c75d878db8886e`, config bundle sha256 `78c00adce662d2d966b2ac50ebde46cde1ae225f0107c6a7c02b70ec8ce80c4f`. The wrapper changed: a box needs the signed `agent_update` AND the signed `agent_config_update`. - **R-876.** After a crash during an install, `dpkg --audit` can read clean while dpkg's update journal (`/var/lib/dpkg/updates/`) is not — and apt refuses every install until `dpkg --configure -a` (measured on demo-hp 2026-10-05: every later pass failed until a person typed it). The wrapper now reads `--audit` and the journal in ONE `sh -c` call (`DPKG_STATE_SCRIPT`) — a clean pass still costs one call (R-845's speed, pinned) — and repairs when either shows something; as a belt, when apt itself says "dpkg was interrupted", it repairs and retries the install ONCE. `REPAIR` now logs `journal=N`; a journal still not empty after the repair refuses (R13). Tests `CrashLeftTheJournal` (the measured shape, the speed, the belt); 3 red-proofs. - **R-874.** The restore-test's first due-check runs 30 minutes after the agent starts (`DefaultFirstEval`), then every interval; a box whose power-on sessions are shorter than the 6 h interval never evaluated. A crash-looping agent restarting faster than 30 minutes still never evaluates (the earned restraint, pinned). - **R-875.** A kept report's reason is neutral — "sent late — kept on the box until the hub could take it" — the copy cannot tell a killed agent from an absent hub. - Red-proofs: `felhom.eu/documentation/audits/catchup-2026-10-05/part{C,D}/`. ## v0.144.1 — a killed pass really keeps its report: the wrapper survives a dead reader; the agent looks again every 5 minutes (R-868, measured live) (2026-10-05) Released by `scripts/release-agent.sh`: binary sha256 `6ccd521d47e64999017e8eb5bc613d724543cfdc5ef9b13bcae9e3ea8c53b8f3`, config bundle sha256 `e89a9ddfb767e177f8874d56f3dcd3bfd47409d157ff830d362653333bf815e8`. The wrapper changed again: a box needs the signed `agent_update` AND the signed `agent_config_update`. - **Found live on demo-hp 2026-10-05 05:45 UTC with v0.144.0** (the night's A5 shape: kill -9 of the pass and the daemon while apt-get ran): apt finished all 13 packages, but the wrapper's next log line went to a stderr pipe no process read any more → `BrokenPipeError` → the wrapper died before it saved its report copy (journal: `PLAN upgrade=13`, then nothing; no copy; the hub got nothing). v0.144.0's mechanism was right and never reached. `Runner.log` and the final `OSAPPLY-REPORT` line now survive a dead reader (the journal still gets every line). Test `AgentDiesMidPass` drives the REAL `log()` into a pipe that breaks while apt-get runs. - **Also found live:** the restarted daemon looked for kept copies ~7 s before the orphaned wrapper wrote one. The daemon now looks at start and every 5 minutes (`Leg.SendUnsentLoop`); `TestR868_ACopyWrittenAfterTheStartIsSentByTheLoop`. - Red-proofs: `felhom.eu/documentation/audits/night-fixes-2026-10-05/partD/r868-brokenpipe-red-proof.txt`. - A second agent release in one session, against "one release per repo": recorded as `09` decision 108 (operator may reverse) — the alternative was to ship a fix proven not to work. ## v0.144.0 — R8 measures the real download; an OS pass reports even when its agent was killed; the debug pass runs with the hub away (R-865, R-868, R-866) (2026-10-05) Released by `scripts/release-agent.sh`: binary sha256 `f18093c3466749ec4cd47f83f97a401160a24bad1e704f1183051adad14db928`, config bundle `felhom-config-bundle.json` sha256 `6acf42fe46df5223384d767801cb2bf73238ba4dab82debb7811ca2f790591d8`. The wrapper `felhom-os-apply` changed, so a box needs BOTH the signed `agent_update` and the signed `agent_config_update`. - **R-865.** `download_bytes` runs `apt-get --print-uris` WITHOUT `-s`: with `-s` apt prints the simulation and no URI list, so R8 summed 0 B and only its 500 MB floor ever applied. `--print-uris` alone downloads nothing (measured on 9202: the archive cache and the versions unchanged). The test fake now answers like real apt (with `-s`: no URIs), and `test_R8_counts_the_real_download` / `test_download_bytes_never_simulates` pin it. - **R-868.** The wrapper writes every apply pass's report to `/report---apply.json` before it prints it (root writes into the agent's dir: the dir opened O_NOFOLLOW and checked to be the agent's own, the file created O_EXCL|O_NOFOLLOW, 0600, handed to the agent). The plan now carries `run_id`, `trigger`, `ring`, echoed in the report. The agent deletes the copy once the hub has the report; a copy left on disk (the agent was killed, or the hub was away) is sent at the agent's start and before every pass (`Leg.SendUnsent`), then deleted. A pass lock (flock on `pass.lock`, across the daemon and a selftest) keeps the sender off a pass that is still running. - **R-866.** The daemon saves the hub's newest os_update block (`os-update-block.json`); `--selftest=os-update` uses it when the hub cannot be reached and says so in its header (`block=SAVED(