From de4a8d20bafd0127d7e7bdd3a6a3f60a10c39260 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Thu, 8 Oct 2026 06:41:20 +0200 Subject: [PATCH] kernel night 7->8 read back: demo-felhom passed (7.0.14-20 default, ~1.5 min), demo-hp no step (R-899 filed); 126 -> 127 Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- CONTEXT.md | 5 + REPORT.md | 46 ++++----- STATUS.md | 14 ++- .../readback/demo-felhom-night.txt | 99 +++++++++++++++++++ .../readback/demo-hp-why-no-step.txt | 3 + documentation/backlog/OPEN-ITEMS.md | 3 +- 6 files changed, 142 insertions(+), 28 deletions(-) create mode 100644 documentation/audits/kernel-night-2026-10-07/readback/demo-felhom-night.txt create mode 100644 documentation/audits/kernel-night-2026-10-07/readback/demo-hp-why-no-step.txt diff --git a/CONTEXT.md b/CONTEXT.md index 469a4f79..4c237035 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -16,6 +16,11 @@ > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +> **2026-10-08 (morning) — the first real kernel night read back.** demo-felhom: staged exactly the told 7.0.14-20 at +> 04:39:04, healthy 04:40:38, default now 7.0.14-20, apps ~1.5 min away, no alarm. demo-hp: no whole-guest backup that +> night (yesterday's 08:49 press + 24 h cadence > window end) → no step; R-899 filed. Reply-To proven by the operator's +> reply. Next: both boxes due 7.0.14-22 tonight; read back 2026-10-09. Register 126 → 127. + > **2026-10-07 (late evening) — pre-night fixes, the night moved to 7→8 (`09` §3 174–176).** Agent v0.153.0 (ring 0 stages > exactly the told kernel — R-898 closed), controller v0.303.0 (`backup/driveready.go`: captures wait ≤10 min after start > for a live drive mount — R-897 closed), hub v0.143.1 (`reply_to` = operator on every household mail). Armed: demo-felhom diff --git a/REPORT.md b/REPORT.md index 1ba80027..fe2b1e84 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,34 +1,30 @@ -# REPORT — three fixes before the first real kernel night, 2026-10-07 (evening) +# REPORT — read-back of the first real kernel night (2026-10-07 → 08) -The brief said 2026-10-08; at 18:34 it was 2026-10-07. Asked; the operator answered *"Start today, and do both boxes this -night"* (`09` §3 decision 176). So the first real kernel night is **7→8**, read back on the morning of 2026-10-08. - -| Part | Result | +| Check | Result | |---|---| -| Rulings | Decisions 174 (mail text stays + a reply address), 175 (window 09–20, max 3 mails: keep), 176 (today, both boxes tonight) recorded FIRST (`a0ff737e`). | -| A — the told kernel boots (R-898) | **Done, delivered, closed.** Agent v0.153.0: ring 0 stages EXACTLY the told kernel (select `listed`, the set from the kver). Told 20 / sources offer 22 → 20 installs (wrapper test); 20 gone → R7 before any change (wrapper test); the hub keeps R7 temporary and tells the household again for the newer kernel (hub test). Red-proofs `audits/kernel-night-2026-10-07/A/redproof.txt`. | -| B — first backup after a restart waits (R-897) | **Done, delivered, closed.** The reviewer's pick, taken: controller v0.303.0 — for 10 min after the controller starts, a capture for an app on a drive runs only once that drive is a LIVE mount in the controller's own namespace (`/proc/self/mountinfo`, the signal the startup app gate already uses); the 5-min refresh skips, a data run waits; after 10 min it runs and logs once. The agent's bind order is unchanged. Red test: not-bound → no capture; bound → capture. | -| C — a reply reaches a person | **Built and delivered (hub v0.143.1); the header read-back is NOT done.** Every mail to a household now carries `reply_to` = `admin@felhom.eu` (all household templates invite contact: the kernel notice, the event sign-off "contact your operator", the setup/link mails); a mail to the operator carries none. Test pinned and red-proved. One test mail sent through the mail service with the hub's exact fields to `tester1@felhom.eu` — it arrived in Gmail (16:54 UTC). **Read-back tried:** the Gmail tool returns no headers (METADATA_ONLY, no RAW); the mail service's read API refused (`restricted_api_key` — send-only, correctly). Note: the demo households' own address IS `admin@felhom.eu`, so their mails carry no Reply-To by design. | -| D — deliver, arm the night | **Done.** Hub 0.143.1 deployed 18:54 local (operator attending), agent 0.153.0 on all three boxes 19:03 (binary only — no root file changed), controller 0.303.0 on all three 18:50 (per-customer floors; global untouched). **Armed:** demo-felhom → **7.0.14-20-pve**, household mail **18:12**; demo-hp → **7.0.14-22-pve**, household mail **18:38**. No reboot by hand. | +| demo-felhom — step ran | **PASS.** Whole-guest backup 04:35; guest/host/Proxmox steps 04:37–04:38; kernel **staged 04:39:04 as exactly the told 7.0.14-20** (the sources offered 7.0.14-22 — R-898's fix held); restart 04:39:08; boot on 7.0.14-20; judged healthy **04:40:38** (38 s after the agent started); **7.0.14-20 is the default**, flag empty, step `good`. | +| demo-felhom — apps | Away about **1.5 min** (restart 04:39:08 → every container healthy at the 04:40:38 verdict). | +| demo-felhom — alarms | None: no `host_stale`, no backup failure mail (demo-felhom has no drive apps, so R-897's fix was not exercised there). Crash guard armed, 0 unclean boots. Hub logged staged → judging → applied. | +| demo-hp — step ran | **NO.** No whole-guest backup that night: the last one was yesterday's morning press at 08:49:53; with the 24 h cadence it came due at 08:49 today, after the window closed (08:30). No backup → no night leg → no kernel step. Nothing changed: running and default 7.0.14-20. Filed **R-899**. | +| Reply-To | **Proven by you:** your reply to the test mail (2026-10-07 20:29) went to admin@felhom.eu. | +| Your inbox overnight | Only Tester 2's expected missed-backup mails (the laptop is off). | -**Rows: 128 before → 126 after. Opened 0. Closed 2 (R-897, R-898).** +**Rows: 126 before → 127 after. Opened 1 (R-899). Closed 0.** -**How demo-hp got told tonight:** its apt lists were a day old, so its last report named no new kernel. I switched its OS -updates OFF on the hub for ~1 minute, ran the agent's own report-only pass (`--selftest=os-update`; with the switch off it -installs nothing and runs no Docker/Proxmox/kernel step), and switched it back ON (`kernel-night-2026-10-07/demo-hp-inventory-pass.txt`). -The hub mailed 20 s after the host report arrived (18:37:51 → 18:38:11). +**What happens next, by itself:** +- demo-hp is still due 7.0.14-22. A new household mail is allowed after 14:38 today (20 h after the last); its + whole-guest backup is due inside tonight's window → the step should run on the night 8→9. +- demo-felhom now sees 7.0.14-22 pending → due; its household gets a mail after 09:00; it should step to 22 tonight too. +- After that both run 7.0.14-22, and "Approve kernel set" can appear. Read back the morning of 2026-10-09. -**Releases:** agent v0.153.0 (binary `b204ebe6…`, tag at `2d1e5d0`); controller v0.303.0 (`29bebbb`, MinAgent 0.131.0); -hub v0.143.1 (`87765bfa`, deployed `d1457892`, Synced/Healthy). CI green on every push, checked by commit. No change on -ep0, Tester 2 or DooPlex's own system. +**Seen, not fixed:** the after-boot "judging" report carries ring 1 (the agent reads its ring from the hub's block, which +it has not fetched yet in the first second after a boot). Cosmetic: the hub's approval reads its own ring list, not +this field. Not fixed today because it needs an agent release and delivery for a label. -**What the morning read-back must check:** each demo box on its told kernel as the default; the household apps back and -the minutes they were down; no false backup failure after the restart (R-897's fix); the hub's kernel events. The two -boxes run DIFFERENT kernels after tonight, so "Approve kernel set" waits until both boot the same one. +Evidence: `documentation/audits/kernel-night-2026-10-07/readback/`. ## Decisions for you -1. **Check the reply address in one click:** open „TEST — Reply-To check" in Gmail and press Reply — it should go to - admin@felhom.eu. My pick: do it once. If you do nothing: the code and its test say it works; nobody has seen it. -2. **The demo households' address is your own**, so their kernel mails carry no Reply-To. My pick: leave it. If you do - nothing: it stays (a reply to them lands in your catch-all anyway). +1. **R-899 — a daytime "back up now" press skips the next night's backup (and its updates).** My pick: count nights, + not 24 hours (the backup is due when the last one is older than ~20 h at the window's start). If you do nothing: each + daytime press costs one night, and the kernel step waits one more day. diff --git a/STATUS.md b/STATUS.md index ce73654b..2c26e397 100644 --- a/STATUS.md +++ b/STATUS.md @@ -2,8 +2,18 @@ **Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop) is off; nothing was sent to it.** -**Updated 2026-10-07 19:10: hub 0.143.1; demo-hp, demo-felhom and Tester 1 run agent 0.153.0 and controller 0.303.0. -The open-items list is at 126. Report: `REPORT.md`.** +**Updated 2026-10-08 06:50: hub 0.143.1; demo-hp, demo-felhom and Tester 1 run agent 0.153.0 and controller 0.303.0. +The open-items list is at 127. Report: `REPORT.md`.** + +## Morning (2026-10-08): the first real kernel night, read back + +- **demo-felhom passed.** It restarted at 04:39 into the new kernel it was told about, was healthy in 38 seconds, and + kept it. Apps away about 1.5 minutes. No alarm. +- **demo-hp did not restart.** Its nightly full backup did not run, because yesterday's daytime "back up now" press moved + its 24-hour clock past the night. Nothing changed on it. It should go tonight (filed as a small item). +- **The reply address works:** your test reply went to admin@. + +**Needs you:** one choice on the backup clock (see `REPORT.md`). If nothing: a daytime press costs one night. ## Tonight (2026-10-07 → 08): the first real kernel night diff --git a/documentation/audits/kernel-night-2026-10-07/readback/demo-felhom-night.txt b/documentation/audits/kernel-night-2026-10-07/readback/demo-felhom-night.txt new file mode 100644 index 00000000..61537274 --- /dev/null +++ b/documentation/audits/kernel-night-2026-10-07/readback/demo-felhom-night.txt @@ -0,0 +1,99 @@ +2026-10-08T04:38:31+02:00 demo-felhom felhom-agent[206012]: time=2026-10-08T04:38:31.955+02:00 level=INFO msg="osupdate: DONE" run=20261008T023722Z layer=docker vmid=9201 ring=0 trigger=night outcome=nothing healthy=true reason="" upgraded=0 pending= +2026-10-08T04:38:31+02:00 demo-felhom felhom-agent[206012]: time=2026-10-08T04:38:31.968+02:00 level=INFO msg="osupdate: pve step holds the /etc/pve write gate — the agent's own writes wait until it ends" run=20261008T023722Z layer=pve vmid=9201 tr +2026-10-08T04:38:31+02:00 demo-felhom felhom-agent[206012]: time=2026-10-08T04:38:31.968+02:00 level=INFO msg="osupdate: START" run=20261008T023722Z layer=pve vmid=9201 ring=0 trigger=night enabled=true release=ring0-20261008T023722Z +2026-10-08T04:38:32+02:00 demo-felhom felhom-os-apply[725634]: os-apply: START release=ring0-20261008T023722Z layer=pve:9201 lane=slow mode=apply select=pending-pve packages=0 authority=ring0 +2026-10-08T04:38:47+02:00 demo-felhom felhom-os-apply[726282]: os-apply: DONE rc=0 seconds=4.1 upgraded=5 restart-needed=pmxcfs,pve-firewall,pve-ha-crm,pve-ha-lrm,pvedaemon,pvedaemon worke,pveproxy,pveproxy worker,pvescheduler,pvestatd,rrdcached,spic +2026-10-08T04:38:52+02:00 demo-felhom felhom-agent[206012]: time=2026-10-08T04:38:52.894+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261008T023722Z layer=pve:9201 lane=slow mode=apply select=pending-pve packages=0 a +2026-10-08T04:38:52+02:00 demo-felhom felhom-agent[206012]: time=2026-10-08T04:38:52.895+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=4.1 upgraded=5 restart-needed=pmxcfs,pve-firewall,pve-ha-crm,pve-ha-lrm,pvedaemon,pved +2026-10-08T04:38:53+02:00 demo-felhom felhom-agent[206012]: time=2026-10-08T04:38:53.688+02:00 level=INFO msg="osupdate: DONE" run=20261008T023722Z layer=pve vmid=9201 ring=0 trigger=night outcome=applied healthy=true reason="" upgraded=5 pending=0 n +2026-10-08T04:38:53+02:00 demo-felhom felhom-agent[206012]: time=2026-10-08T04:38:53.714+02:00 level=INFO msg="osupdate: pve step released the /etc/pve write gate" run=20261008T023722Z layer=pve vmid=9201 trigger=night +2026-10-08T04:38:54+02:00 demo-felhom felhom-os-apply[726516]: os-apply: START release=ring0-20261008T023722Z layer=kernel lane=slow mode=apply select=listed authority=ring0 running=7.0.2-6-pve +2026-10-08T04:39:02+02:00 demo-felhom felhom-os-apply[727201]: os-apply: KERNEL default = 7.0.2-6-pve (proved from grub.cfg) +2026-10-08T04:39:04+02:00 demo-felhom felhom-os-apply[727259]: os-apply: KERNEL STAGED 7.0.2-6-pve -> 7.0.14-20-pve (default 7.0.2-6-pve, one-shot flag 7.0.14-20-pve); upgraded=1 seconds=1.9 +2026-10-08T04:39:04+02:00 demo-felhom felhom-agent[206012]: time=2026-10-08T04:39:04.532+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261008T023722Z layer=kernel lane=slow mode=apply select=listed authority=ring0 run +2026-10-08T04:39:04+02:00 demo-felhom felhom-agent[206012]: time=2026-10-08T04:39:04.532+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: KERNEL default = 7.0.2-6-pve (proved from grub.cfg)" +2026-10-08T04:39:04+02:00 demo-felhom felhom-agent[206012]: time=2026-10-08T04:39:04.533+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: KERNEL STAGED 7.0.2-6-pve -> 7.0.14-20-pve (default 7.0.2-6-pve, one-shot flag 7.0.14-20-pve); upgraded= +2026-10-08T04:39:04+02:00 demo-felhom felhom-agent[206012]: time=2026-10-08T04:39:04.533+02:00 level=INFO msg="osupdate: DONE" run=20261008T023722Z layer=kernel vmid=9201 trigger=night ring=0 outcome=staged healthy=true reason="" upgraded=1 pending=0 +2026-10-08T04:39:08+02:00 demo-felhom felhom-os-apply[727378]: os-apply: KERNEL REBOOT — one-shot boot of 7.0.14-20-pve (the default stays 7.0.2-6-pve) +2026-10-08T04:39:08+02:00 demo-felhom felhom-agent[206012]: time=2026-10-08T04:39:08.515+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: KERNEL REBOOT — one-shot boot of 7.0.14-20-pve (the default stays 7.0.2-6-pve)" +---BOOT +2026-10-08T04:39:59+02:00 demo-felhom felhom-agent[1268]: time=2026-10-08T04:39:59.718+02:00 level=INFO msg="felhom-agent daemon starting" version=0.153.0 host_id=demo-felhom-8363b5 hub_url=https://hub.felhom.eu interval_s=900 +2026-10-08T04:39:59+02:00 demo-felhom felhom-agent[1268]: time=2026-10-08T04:39:59.852+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: KERNEL AFTER-BOOT oneshot -> judging running=7.0.14-20-pve (7.0.2-6-pve -> 7.0.14-20-pve)" +2026-10-08T04:39:59+02:00 demo-felhom felhom-agent[1268]: time=2026-10-08T04:39:59.853+02:00 level=INFO msg="osupdate: kernel step — judging the one-shot boot" run=boot-20261008T023959Z layer=kernel vmid=0 from=7.0.2-6-pve to=7.0.14-20-pve wait=20m +2026-10-08T04:39:59+02:00 demo-felhom felhom-agent[1268]: time=2026-10-08T04:39:59.942+02:00 level=INFO msg="osupdate: kernel step — the box reached the hub on the new kernel" run=boot-20261008T023959Z layer=kernel vmid=0 after=0s +2026-10-08T04:40:38+02:00 demo-felhom felhom-os-apply[4716]: os-apply: KERNEL default = 7.0.14-20-pve (proved from grub.cfg) +2026-10-08T04:40:38+02:00 demo-felhom felhom-os-apply[4718]: os-apply: KERNEL GOOD 7.0.14-20-pve is the default now (was 7.0.2-6-pve) +2026-10-08T04:40:38+02:00 demo-felhom felhom-agent[1268]: time=2026-10-08T04:40:38.269+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: KERNEL default = 7.0.14-20-pve (proved from grub.cfg)" +2026-10-08T04:40:38+02:00 demo-felhom felhom-agent[1268]: time=2026-10-08T04:40:38.269+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: KERNEL GOOD 7.0.14-20-pve is the default now (was 7.0.2-6-pve)" +2026-10-08T04:40:38+02:00 demo-felhom felhom-agent[1268]: time=2026-10-08T04:40:38.269+02:00 level=INFO msg="osupdate: kernel step — applied" run=boot-20261008T023959Z layer=kernel vmid=0 to=7.0.14-20-pve reason="healthy 38s after the agent started +2026-10-08T04:40:38+02:00 demo-felhom felhom-agent[1268]: time=2026-10-08T04:40:38.269+02:00 level=INFO msg="osupdate: DONE" run=boot-20261008T023959Z layer=kernel vmid=0 outcome=applied healthy=true reason="healthy 38s after the agent started; the n +---GUEST +2026-10-08T04:40:00+02:00 demo-felhom systemd[1]: Starting pve-guests.service - PVE guests... +2026-10-08T04:40:01+02:00 demo-felhom pve-guests[1451]: starting task UPID:demo-felhom:000005C7:000004F9:6AC70281:startall::root@pam: +2026-10-08T04:40:01+02:00 demo-felhom pvesh[1451]: Starting CT 9201 +2026-10-08T04:40:01+02:00 demo-felhom pve-guests[1479]: starting task UPID:demo-felhom:000005C8:000004FA:6AC70281:vzstart:9201:root@pam: +2026-10-08T04:40:01+02:00 demo-felhom pve-guests[1480]: starting CT 9201: UPID:demo-felhom:000005C8:000004FA:6AC70281:vzstart:9201:root@pam: +{ + "authority": "ring0", + "default_before": "7.0.2-6-pve", + "from": "7.0.2-6-pve", + "health_before": { + "guest": { + "containers": { + "cloudflared": { + "health": "healthy", + "id": "bdbd524f10b0e4c37dab5d1faa7df84f820cb3b58ad6dbc8ab2dcb73902fe12d", + "state": "running" + }, + "felhom-controller": { + "health": "healthy", + "id": "dd0ff16d55df84742838fb39ca6e810cde6e20bc70bf36c40d16ac35bc6c2fbf", + "state": "running" + }, + "filebrowser": { + "health": "healthy", + "id": "6bda22106919fbef251af5494395af05b34e095d5e4187b90af85af69470ffa4", + "state": "running" + }, + "opengist": { + "health": "healthy", + "id": "1c0e86675b1353a094361c2519236c41e9ca6502e81be91b2ec53f9ad9de0c5d", + "state": "running" + }, + "traefik": { + "health": "none", + "id": "ddc8abb0ce45574cb0aad245db8c42641139114fe075d06d7f849ffd0441524e", + "state": "running" + } + }, + "controller": "healthy", + "controller_docker_ok": true, + "docker_ok": true, + "network_ok": true + }, + "guest_running": true, + "host_services": { + "felhom-agent": "active", + "pve-cluster": "active", + "pvedaemon": "active", + "pveproxy": "active", + "pvestatd": "active" + } + }, + "judging_since": "2026-10-08T02:39:59Z", + "packages": [ + { + "name": "proxmox-kernel-7.0", + "version": "7.0.14-20" + } + ], + "phase": "good", + "rebooted_at": "2026-10-08T02:39:08Z", + "result_at": "2026-10-08T02:40:38Z", + "self_revert_used": false, + "staged_at": "2026-10-08T02:39:04Z", + "step_id": "ring0-20261008T023722Z", + "to": "7.0.14-20-pve", + "updated_at": "2026-10-08T02:40:38Z", + "vmid": 9201 +} diff --git a/documentation/audits/kernel-night-2026-10-07/readback/demo-hp-why-no-step.txt b/documentation/audits/kernel-night-2026-10-07/readback/demo-hp-why-no-step.txt new file mode 100644 index 00000000..64a3a749 --- /dev/null +++ b/documentation/audits/kernel-night-2026-10-07/readback/demo-hp-why-no-step.txt @@ -0,0 +1,3 @@ +2026-10-07T04:40:48+02:00 demo-hp felhom-agent[297545]: time=2026-10-07T04:40:48.072+02:00 level=INFO msg="backup: completed" vmid=9201 target=local archive=local:backup/vzdump-lxc-9201-2026_10_07-04_35_11.tar.zst size_b +2026-10-07T08:49:53+02:00 demo-hp felhom-agent[297545]: time=2026-10-07T08:49:53.496+02:00 level=INFO msg="backup: completed" vmid=9201 target=local archive=local:backup/vzdump-lxc-9201-2026_10_07-08_45_01.tar.zst size_b +demo-hp last whole-guest backup 2026-10-07 08:49:53 (the morning 'Mentes most' press, R-518 read-back) + 24 h cadence = due 2026-10-08 08:49, after the window [04:30, 08:30) closed -> no backup -> no night leg -> no kernel step. No kernel state written; running and default 7.0.14-20. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 20371906..d350a5f4 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -214,7 +214,8 @@ stopping line that lies. | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| -| **R-836** | Box system & updates | P2 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot `: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** **2026-10-07 07:58: raised P3 → P2 (`09` §3 decision 164 — the kernel lane before the first paying customer); the spike runs with the operator present.** **2026-10-07 SPIKE DONE (operator present; 24 reboots, 2 power cycles; `audits/kernel-spike-2026-10-07/DESIGN-kernel-lane.md`).** On Tester 1 (VM), demo-felhom and demo-hp (Secure Boot on): **a one-shot flag in a GRUB env block on the ESP (vfat) WORKS** — new kernel once, then the old one; a panicking new kernel (`panic=10`) falls back with no person. **UEFI `BootNext` WORKS too.** **A hardware watchdog armed by systemd during the reboot FAILS on all three** (i6300ESB, Intel TCO, AMD SP5100): the reset clears the timer, so a FROZEN new kernel needs a power cycle. Pick: the ESP flag + lockup-to-panic kernel options (unmeasured). Two operator questions in the design (night restarts and the household; a freeze needs a person). Boxes left clean on a healthily booted default. **2026-10-07 14:43 — `09` §3 decision 172: the kernel lane is YES — option A (the ESP one-shot flag) with option C (lockup → panic) on the one-shot entry; a box may restart at night; the household is mailed the day before (what to do if it is not back: unplug, plug back in); each kernel set is operator-approved after ring 0; a frozen new kernel needing a power cycle is ACCEPTED for the first customers. BUILD STARTED 2026-10-07 (the kernel-lane arc).** **2026-10-07 evening: BUILT and proven by hand** — agent v0.152.0 + hub v0.143.0 (`11` §5.11), delivered to demo-hp, demo-felhom and Tester 1. Tester 1 (`audits/kernel-lane-2026-10-07/E/RESULT.md`): a forced panic on the one-shot → back on the old kernel by itself (fell_back, 0 unclean boots counted); a held guest → judged 20 min → ONE self-revert → old kernel; a healthy boot → 7.0.14-22 the default. The household mail went out for demo-felhom 18:12 (Gmail). Option C UNMEASURED (no test_lockup module). **WHERE IT STOPPED (next session):** the ring-0 NIGHT run (Part E3) — tonight demo-felhom refuses the told 7.0.14-20 (R-898, safe); 7.0.14-22 is due on both demo boxes for the night 2026-10-08→09 (mail 2026-10-08 daytime); read back the morning of 2026-10-09; then approve the set on the System page and deliver to the Tester 1 box as ring 1 by a signed os_kernel_step (Part E4 — Tester 1 already runs 7.0.14-22, so the ring-1 proof needs the NEXT kernel). **2026-10-07 evening (decision 176): the night run moved to TONIGHT 7→8.** Pre-night fixes delivered (agent v0.153.0 — ring 0 stages exactly the told kernel, R-898 closed; controller v0.303.0 — the first capture after a boot waits for the drive, R-897 closed; hub v0.143.1 — Reply-To on household mails). Armed at 19:05: demo-felhom due 7.0.14-20 (household told 18:12), demo-hp due 7.0.14-22 (told 18:38; demo-hp's day-old apt lists refreshed by a report-only pass with its updates switched off for ~1 min). **WHERE IT STOPPED: read back the night on 2026-10-08 morning** (each box on its kernel as default, the apps back, minutes down, no false backup failure), then approve the set (the two boxes ran DIFFERENT kernels — the button waits until both boot the same one) and Part E4. | — | — | CC | +| **R-836** | Box system & updates | P2 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot `: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** **2026-10-07 07:58: raised P3 → P2 (`09` §3 decision 164 — the kernel lane before the first paying customer); the spike runs with the operator present.** **2026-10-07 SPIKE DONE (operator present; 24 reboots, 2 power cycles; `audits/kernel-spike-2026-10-07/DESIGN-kernel-lane.md`).** On Tester 1 (VM), demo-felhom and demo-hp (Secure Boot on): **a one-shot flag in a GRUB env block on the ESP (vfat) WORKS** — new kernel once, then the old one; a panicking new kernel (`panic=10`) falls back with no person. **UEFI `BootNext` WORKS too.** **A hardware watchdog armed by systemd during the reboot FAILS on all three** (i6300ESB, Intel TCO, AMD SP5100): the reset clears the timer, so a FROZEN new kernel needs a power cycle. Pick: the ESP flag + lockup-to-panic kernel options (unmeasured). Two operator questions in the design (night restarts and the household; a freeze needs a person). Boxes left clean on a healthily booted default. **2026-10-07 14:43 — `09` §3 decision 172: the kernel lane is YES — option A (the ESP one-shot flag) with option C (lockup → panic) on the one-shot entry; a box may restart at night; the household is mailed the day before (what to do if it is not back: unplug, plug back in); each kernel set is operator-approved after ring 0; a frozen new kernel needing a power cycle is ACCEPTED for the first customers. BUILD STARTED 2026-10-07 (the kernel-lane arc).** **2026-10-07 evening: BUILT and proven by hand** — agent v0.152.0 + hub v0.143.0 (`11` §5.11), delivered to demo-hp, demo-felhom and Tester 1. Tester 1 (`audits/kernel-lane-2026-10-07/E/RESULT.md`): a forced panic on the one-shot → back on the old kernel by itself (fell_back, 0 unclean boots counted); a held guest → judged 20 min → ONE self-revert → old kernel; a healthy boot → 7.0.14-22 the default. The household mail went out for demo-felhom 18:12 (Gmail). Option C UNMEASURED (no test_lockup module). **WHERE IT STOPPED (next session):** the ring-0 NIGHT run (Part E3) — tonight demo-felhom refuses the told 7.0.14-20 (R-898, safe); 7.0.14-22 is due on both demo boxes for the night 2026-10-08→09 (mail 2026-10-08 daytime); read back the morning of 2026-10-09; then approve the set on the System page and deliver to the Tester 1 box as ring 1 by a signed os_kernel_step (Part E4 — Tester 1 already runs 7.0.14-22, so the ring-1 proof needs the NEXT kernel). **2026-10-07 evening (decision 176): the night run moved to TONIGHT 7→8.** Pre-night fixes delivered (agent v0.153.0 — ring 0 stages exactly the told kernel, R-898 closed; controller v0.303.0 — the first capture after a boot waits for the drive, R-897 closed; hub v0.143.1 — Reply-To on household mails). Armed at 19:05: demo-felhom due 7.0.14-20 (household told 18:12), demo-hp due 7.0.14-22 (told 18:38; demo-hp's day-old apt lists refreshed by a report-only pass with its updates switched off for ~1 min). **WHERE IT STOPPED: read back the night on 2026-10-08 morning** (each box on its kernel as default, the apps back, minutes down, no false backup failure), then approve the set (the two boxes ran DIFFERENT kernels — the button waits until both boot the same one) and Part E4. **2026-10-08 read-back (`audits/kernel-night-2026-10-07/readback/`): demo-felhom PASSED its first real kernel night** — whole-guest backup 04:35, OS + Proxmox steps 04:37–04:38, kernel staged 04:39:04 as EXACTLY the told 7.0.14-20 (not the newer 22 — R-898's fix held), restart 04:39:08, new kernel judged healthy 04:40:38 (38 s after the agent started), now the default; apps away ~1.5 min; no operator alarm, no false backup failure (demo-felhom has no drive apps, so R-897's fix was not exercised there); crash guard 0. **demo-hp took NO step** — no whole-guest backup that night (R-899): it stays on 7.0.14-20, still due 7.0.14-22; a new household mail is allowed after 14:38 today, so it should run the night 2026-10-08→09. Reply-To proven: the operator's reply to the test mail went to admin@felhom.eu (Gmail 2026-10-07 18:29 UTC). **WHERE IT STOPPED:** read back demo-hp the morning of 2026-10-09; the approval button waits until both demo boxes booted the SAME kernel (now 20 vs due 22) — likely a step to 22 on demo-felhom later; then Part E4. | — | — | CC | +| **R-899** | Box system & updates | P3 | **A daytime whole-guest backup moves the 24 h cadence past the night window, so the next night has no whole-guest backup — and with it no OS leg and no kernel step.** SEEN 2026-10-08: demo-hp's last whole-guest backup was a morning press at 2026-10-07 08:49:53; due again 2026-10-08 08:49, after the window [04:30, 08:30) closed → no backup, no night leg; the household had been told „tonight" (18:38) and nothing happened (harmless: no change). The kernel lane recovers by itself (a new mail after 20 h, step the next night, inside the 3-mail limit), but every daytime press costs one night. Fix direction (a design, not taken): the cadence counts nights, not 24 h (due when the last whole-guest backup is older than ~20 h at the window's start), or a press does not reset the night's cadence. `audits/kernel-night-2026-10-07/readback/demo-hp-why-no-step.txt` | **OPEN — filed 2026-10-08 (kernel-night read-back).** | — | — | CC | | **R-862** | Box system & updates | P3 | **Tester 2 cannot take the config bundle until one by-hand bootstrap is done: its `felhom-os-apply` (agent 0.142.0) predates the bundle mode, and no signed job can write a root file on it.** FOUND 2026-10-04 (R-840 build, Part C): the route reaches every box installed from installer 1.31.0 on, and the demo boxes (bootstrapped by CC); Tester 2 has every root file of agent 0.142.0 (installer 1.30.0) and lacks only the R-858 wrapper fix, which matters only for a Docker step it gets solely from a signed job. CC has no route to Tester 2 (its door admits only the operator's WireGuard peer; `felhom-op` cannot become root). The steps: `runbooks/config-bundle.md` "Tester 2". Then CC signs the bundle and reads it back. | **WAITING-ON-OPERATOR** — the operator said (2026-10-04 ~19:05) he will try through his tunnel **2026-10-07 07:58 (`09` §3 decision 170): the operator believes Tester 2 never set up the off-site backup (no escrow); the laptop may stay off for days; the tester plans to add an HDD.** **2026-10-07 (Part H, read only):** in Tester 2's last 10 notifications (2026-10-05 09:16 → 2026-10-07 06:58) the household got 0 mails, the operator 10 (host_down ×9, expected_dbdump_missed ×1); control: the household channel shows on the other three customers. So the household is not mailed daily while the box is off; no row (`audits/day-2026-10-07/H/`). | the operator's WireGuard tunnel | the operator runs the three bootstrap commands; CC sends the bundle | operator | | **R-35** | Box system & updates | P3 | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-05 (burn-down night): NEEDS A DESIGN.** Hot-apply needs a per-setting ruling; persisting sessions puts login tokens on disk (and into the whole-guest archive, `07` §5). Next: the operator's pick between them. | — | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` | CC | | **R-78** | Box system & updates | P3 | **`local_api` authority ruling — auto-reconcile vs detect-only** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-25, size M, roadmap state `idea (deferred OUT of R-77 on purpose)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **An OWED OPERATOR DECISION, not a defect — moved because an owed ruling hidden among feature ideas is the shape this session exists to remove.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | R-77 ships detection because the fix is genuinely undecided, and **both directions can lose customer-visible function**. **Direction 1 (today):** `controller.yaml` wins and drift is silent → the 2026-07-25 island migration blinded the whole fleet's control plane for 17.5 h (drive gate, guest-reboot recovery, quiesce/backup all degrade). R-77 makes that loud but does not stop it recurring. **Direction 2 (`bootstrap.json` wins, auto-reconcile on boot):** a guest whose `controller.yaml` is CORRECT and whose `bootstrap.json` is stale — a half-completed re-provision, a hand-repaired guest, a setup-wizard box — gets a **working channel clobbered on the next restart**, fleet-wide and silently, during a routine deploy. That is not obviously better than the bug. Needs a spike: which writer is authoritative per field (endpoint vs fingerprint vs token — `mergeLocalAPI` replaces the whole block, so they cannot be reconciled independently today), whether the agent side should stamp a generation/mtime so 'newer wins' is even expressible, and whether reconcile should require an operator ack. Until then the drift alert plus a manual edit is the supported path. | CC |