kernel lane: report, STATUS, CONTEXT, R-836 where-it-stopped, R-898 filed; 126 -> 128
gates / gates (push) Successful in 3m11s
gates / gates (push) Successful in 3m11s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -6,3 +6,19 @@ uploaded signed op to the hub jobs queue
|
||||
signed: op=agent_update host=demo-felhom-8363b5 guest="" key_id=felhom-op-1 nonce=1064b6dde1dd050dd5f6f7b788990412 expires=2026-10-07T16:04:42Z
|
||||
wrote envelope to /dev/null
|
||||
uploaded signed op to the hub jobs queue
|
||||
== agent_config_update 0.152.0-step1 → demo-hp-bb76ea, 2026-10-07T15:31:38Z
|
||||
signed: op=agent_config_update host=demo-hp-bb76ea guest="" key_id=felhom-op-1 nonce=e87c668823146debf011a1680691118c expires=2026-10-07T16:16:38Z
|
||||
wrote envelope to /dev/null
|
||||
uploaded signed op to the hub jobs queue
|
||||
== agent_config_update 0.152.0-step1 → demo-felhom-8363b5, 2026-10-07T15:31:38Z
|
||||
signed: op=agent_config_update host=demo-felhom-8363b5 guest="" key_id=felhom-op-1 nonce=699531f773854f791e3517174cf382ad expires=2026-10-07T16:16:38Z
|
||||
wrote envelope to /dev/null
|
||||
uploaded signed op to the hub jobs queue
|
||||
== agent_config_update 0.152.0 → demo-hp-bb76ea, 2026-10-07T15:47:04Z
|
||||
signed: op=agent_config_update host=demo-hp-bb76ea guest="" key_id=felhom-op-1 nonce=b126b41931b7123998dcc6f3e182e1dc expires=2026-10-07T16:32:04Z
|
||||
wrote envelope to /dev/null
|
||||
uploaded signed op to the hub jobs queue
|
||||
== agent_config_update 0.152.0 → demo-felhom-8363b5, 2026-10-07T15:47:04Z
|
||||
signed: op=agent_config_update host=demo-felhom-8363b5 guest="" key_id=felhom-op-1 nonce=ac6fde774fe15cfdc4671fbbeddc502f expires=2026-10-07T16:32:04Z
|
||||
wrote envelope to /dev/null
|
||||
uploaded signed op to the hub jobs queue
|
||||
|
||||
@@ -0,0 +1,14 @@
|
||||
2026/10/07 16:17:16 [INFO] osupdates: tester-1-d70be4 reported kernel run 20261007T141703Z: ring=1 mode=apply outcome=staged healthy=true upgraded=1 pending=0 not-covered=0 restart-needed=0 reboot-needed=true wrapper=13.2s
|
||||
2026/10/07 16:20:32 [INFO] osupdates: tester-1-d70be4 reported kernel run boot-20261007T142032Z: ring=1 mode=kernel-boot outcome=fell_back healthy=false upgraded=0 pending=0 not-covered=0 restart-needed=0 reboot-needed=false wrapper=0.0s
|
||||
2026/10/07 16:20:32 [INFO] Operator email sent for tester-1/os_kernel_step
|
||||
2026/10/07 16:35:52 [INFO] osupdates: tester-1-d70be4 reported kernel run 20261007T143539Z: ring=1 mode=apply outcome=staged healthy=true upgraded=1 pending=0 not-covered=0 restart-needed=0 reboot-needed=true wrapper=12.7s
|
||||
2026/10/07 16:38:43 [INFO] osupdates: tester-1-d70be4 reported kernel run boot-20261007T143843Z: ring=1 mode=kernel-boot outcome=judging healthy=false upgraded=0 pending=0 not-covered=0 restart-needed=0 reboot-needed=false wrapper=0.0s
|
||||
2026/10/07 16:58:54 [INFO] osupdates: tester-1-d70be4 reported kernel run boot-20261007T143843Z: ring=1 mode=kernel-revert outcome=health_failed healthy=false upgraded=0 pending=0 not-covered=0 restart-needed=0 reboot-needed=false wrapper=0.0s
|
||||
2026/10/07 16:58:54 [INFO] Operator email sent for tester-1/os_kernel_step
|
||||
2026/10/07 16:59:30 [INFO] osupdates: tester-1-d70be4 reported kernel run boot-20261007T145930Z: ring=1 mode=kernel-boot outcome=self_reverted healthy=false upgraded=0 pending=0 not-covered=0 restart-needed=0 reboot-needed=false wrapper=0.0s
|
||||
2026/10/07 16:59:31 [INFO] Operator email sent for tester-1/os_kernel_step
|
||||
2026/10/07 17:15:18 [INFO] osupdates: tester-1-d70be4 reported kernel run 20261007T151439Z: ring=1 mode=apply outcome=staged healthy=true upgraded=2 pending=0 not-covered=0 restart-needed=0 reboot-needed=true wrapper=39.0s
|
||||
2026/10/07 17:17:53 [INFO] osupdates: tester-1-d70be4 reported kernel run boot-20261007T151753Z: ring=1 mode=kernel-boot outcome=judging healthy=false upgraded=0 pending=0 not-covered=0 restart-needed=0 reboot-needed=false wrapper=0.0s
|
||||
2026/10/07 17:19:10 [INFO] osupdates: tester-1-d70be4 reported kernel run boot-20261007T151753Z: ring=1 mode=kernel-good outcome=applied healthy=true upgraded=0 pending=0 not-covered=0 restart-needed=0 reboot-needed=false wrapper=0.0s
|
||||
2026/10/07 18:12:11 [INFO] kernel notice (7.0.14-20-pve) emailed to the household of demo-felhom (lang=hu)
|
||||
2026/10/07 18:12:11 [INFO] osupdates: kernel 7.0.14-20-pve — the household of demo-felhom-8363b5 was told the box restarts tonight (lang=hu)
|
||||
@@ -1,5 +1,8 @@
|
||||
# The kernel lane — what the spike measured, and a design (R-836, R-812 option B; `09` §3 decisions 164, 171)
|
||||
|
||||
> **Ruled 2026-10-07 14:43 (decision 172): A with C. BUILT the same day — `11-os-updates.md` §5.11, proofs in
|
||||
> `audits/kernel-lane-2026-10-07/`.**
|
||||
|
||||
Spike 2026-10-07, operator present. Boxes: the Tester 1 box (VM 341 on the HP box, OVMF, Secure Boot off), demo-felhom
|
||||
(Intel N100, Secure Boot off), demo-hp (AMD, Secure Boot ON). Kernels: 7.0.2-6 and 7.0.14-20. 24 reboots, 2 operator
|
||||
power cycles. Evidence per boot in this folder (`<box>/<step>/before.txt`, `after.txt`, `result.txt`; screenshots for
|
||||
|
||||
@@ -215,7 +215,8 @@ stopping line that lies.
|
||||
|
||||
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| **R-836** | Box system & updates | P2 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin <new> --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot <new>`: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** **2026-10-07 07:58: raised P3 → P2 (`09` §3 decision 164 — the kernel lane before the first paying customer); the spike runs with the operator present.** **2026-10-07 SPIKE DONE (operator present; 24 reboots, 2 power cycles; `audits/kernel-spike-2026-10-07/DESIGN-kernel-lane.md`).** On Tester 1 (VM), demo-felhom and demo-hp (Secure Boot on): **a one-shot flag in a GRUB env block on the ESP (vfat) WORKS** — new kernel once, then the old one; a panicking new kernel (`panic=10`) falls back with no person. **UEFI `BootNext` WORKS too.** **A hardware watchdog armed by systemd during the reboot FAILS on all three** (i6300ESB, Intel TCO, AMD SP5100): the reset clears the timer, so a FROZEN new kernel needs a power cycle. Pick: the ESP flag + lockup-to-panic kernel options (unmeasured). Two operator questions in the design (night restarts and the household; a freeze needs a person). Boxes left clean on a healthily booted default. **2026-10-07 14:43 — `09` §3 decision 172: the kernel lane is YES — option A (the ESP one-shot flag) with option C (lockup → panic) on the one-shot entry; a box may restart at night; the household is mailed the day before (what to do if it is not back: unplug, plug back in); each kernel set is operator-approved after ring 0; a frozen new kernel needing a power cycle is ACCEPTED for the first customers. BUILD STARTED 2026-10-07 (the kernel-lane arc).** | — | — | CC |
|
||||
| **R-836** | Box system & updates | P2 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin <new> --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot <new>`: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** **2026-10-07 07:58: raised P3 → P2 (`09` §3 decision 164 — the kernel lane before the first paying customer); the spike runs with the operator present.** **2026-10-07 SPIKE DONE (operator present; 24 reboots, 2 power cycles; `audits/kernel-spike-2026-10-07/DESIGN-kernel-lane.md`).** On Tester 1 (VM), demo-felhom and demo-hp (Secure Boot on): **a one-shot flag in a GRUB env block on the ESP (vfat) WORKS** — new kernel once, then the old one; a panicking new kernel (`panic=10`) falls back with no person. **UEFI `BootNext` WORKS too.** **A hardware watchdog armed by systemd during the reboot FAILS on all three** (i6300ESB, Intel TCO, AMD SP5100): the reset clears the timer, so a FROZEN new kernel needs a power cycle. Pick: the ESP flag + lockup-to-panic kernel options (unmeasured). Two operator questions in the design (night restarts and the household; a freeze needs a person). Boxes left clean on a healthily booted default. **2026-10-07 14:43 — `09` §3 decision 172: the kernel lane is YES — option A (the ESP one-shot flag) with option C (lockup → panic) on the one-shot entry; a box may restart at night; the household is mailed the day before (what to do if it is not back: unplug, plug back in); each kernel set is operator-approved after ring 0; a frozen new kernel needing a power cycle is ACCEPTED for the first customers. BUILD STARTED 2026-10-07 (the kernel-lane arc).** **2026-10-07 evening: BUILT and proven by hand** — agent v0.152.0 + hub v0.143.0 (`11` §5.11), delivered to demo-hp, demo-felhom and Tester 1. Tester 1 (`audits/kernel-lane-2026-10-07/E/RESULT.md`): a forced panic on the one-shot → back on the old kernel by itself (fell_back, 0 unclean boots counted); a held guest → judged 20 min → ONE self-revert → old kernel; a healthy boot → 7.0.14-22 the default. The household mail went out for demo-felhom 18:12 (Gmail). Option C UNMEASURED (no test_lockup module). **WHERE IT STOPPED (next session):** the ring-0 NIGHT run (Part E3) — tonight demo-felhom refuses the told 7.0.14-20 (R-898, safe); 7.0.14-22 is due on both demo boxes for the night 2026-10-08→09 (mail 2026-10-08 daytime); read back the morning of 2026-10-09; then approve the set on the System page and deliver to the Tester 1 box as ring 1 by a signed os_kernel_step (Part E4 — Tester 1 already runs 7.0.14-22, so the ring-1 proof needs the NEXT kernel). | — | — | CC |
|
||||
| **R-898** | Box system & updates | P3 | **Ring 0's told kernel can be a day old: the hub names the kernel from the box's LAST host report, the box's sources may offer a newer one by night, and the step then refuses (R23) — one night and one household mail lost.** SEEN 2026-10-07 18:12: demo-felhom's household was mailed for 7.0.14-20-pve (from the night report) while the box's apt already offered 7.0.14-22; the night leg would stage `pending-kernel` → 22 ≠ the told 20 → R23 before any change (safe; read from the code, `felhom-os-apply` Kernel.stage). Fix direction (needs an agent + hub release, so not in this session): ring 0 stages EXACTLY the told kernel (select `listed`, the set derived from the kver — Proxmox keeps old kernel versions), and the hub treats a told-kernel mismatch as transient, not „ended for good". | **OPEN — filed 2026-10-07 (kernel-lane arc); tonight's demo-felhom run shows it.** | — | — | CC |
|
||||
| **R-862** | Box system & updates | P3 | **Tester 2 cannot take the config bundle until one by-hand bootstrap is done: its `felhom-os-apply` (agent 0.142.0) predates the bundle mode, and no signed job can write a root file on it.** FOUND 2026-10-04 (R-840 build, Part C): the route reaches every box installed from installer 1.31.0 on, and the demo boxes (bootstrapped by CC); Tester 2 has every root file of agent 0.142.0 (installer 1.30.0) and lacks only the R-858 wrapper fix, which matters only for a Docker step it gets solely from a signed job. CC has no route to Tester 2 (its door admits only the operator's WireGuard peer; `felhom-op` cannot become root). The steps: `runbooks/config-bundle.md` "Tester 2". Then CC signs the bundle and reads it back. | **WAITING-ON-OPERATOR** — the operator said (2026-10-04 ~19:05) he will try through his tunnel **2026-10-07 07:58 (`09` §3 decision 170): the operator believes Tester 2 never set up the off-site backup (no escrow); the laptop may stay off for days; the tester plans to add an HDD.** **2026-10-07 (Part H, read only):** in Tester 2's last 10 notifications (2026-10-05 09:16 → 2026-10-07 06:58) the household got 0 mails, the operator 10 (host_down ×9, expected_dbdump_missed ×1); control: the household channel shows on the other three customers. So the household is not mailed daily while the box is off; no row (`audits/day-2026-10-07/H/`). | the operator's WireGuard tunnel | the operator runs the three bootstrap commands; CC sends the bundle | operator |
|
||||
| **R-35** | Box system & updates | P3 | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-05 (burn-down night): NEEDS A DESIGN.** Hot-apply needs a per-setting ruling; persisting sessions puts login tokens on disk (and into the whole-guest archive, `07` §5). Next: the operator's pick between them. | — | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` | CC |
|
||||
| **R-78** | Box system & updates | P3 | **`local_api` authority ruling — auto-reconcile vs detect-only** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-25, size M, roadmap state `idea (deferred OUT of R-77 on purpose)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **An OWED OPERATOR DECISION, not a defect — moved because an owed ruling hidden among feature ideas is the shape this session exists to remove.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | R-77 ships detection because the fix is genuinely undecided, and **both directions can lose customer-visible function**. **Direction 1 (today):** `controller.yaml` wins and drift is silent → the 2026-07-25 island migration blinded the whole fleet's control plane for 17.5 h (drive gate, guest-reboot recovery, quiesce/backup all degrade). R-77 makes that loud but does not stop it recurring. **Direction 2 (`bootstrap.json` wins, auto-reconcile on boot):** a guest whose `controller.yaml` is CORRECT and whose `bootstrap.json` is stale — a half-completed re-provision, a hand-repaired guest, a setup-wizard box — gets a **working channel clobbered on the next restart**, fleet-wide and silently, during a routine deploy. That is not obviously better than the bug. Needs a spike: which writer is authoritative per field (endpoint vs fingerprint vs token — `mergeLocalAPI` replaces the whole block, so they cannot be reconciled independently today), whether the agent side should stamp a generation/mtime so 'newer wins' is even expressible, and whether reconcile should require an operator ack. Until then the drift alert plus a manual edit is the supported path. | CC |
|
||||
|
||||
Reference in New Issue
Block a user