Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
3.3 KiB
Part E 1 — the kernel lane by hand on the Tester 1 box (VM 341 on demo-hp), 2026-10-07
Agent v0.152.0 + bundle 0.152.0 (delivered by signed jobs: binary, step bundle 0.152.0-step1, bundle). Hub v0.143.0.
Every stage went through the PRODUCT's ring-1 path: a signed os_kernel_step (felhom-opsign → hub jobs queue → the
agent's KernelStepExecutor → the wrapper, which re-verifies the signature). Each REBOOT was started by hand with the
wrapper's own kernel-reboot mode (the night leg does the same call on a told night). The after-boot judging was the
agent's own, untouched. Order run: (b), (c), (a) — a crash and a revert must start from the old kernel.
| Test | Kernel | What happened (times local) | Verdict |
|---|---|---|---|
| (b) forced panic | 7.0.2-6 → 7.0.14-20 | staged 16:17:16; the one-shot entry broken BY HAND (rdinit= and init= both missing, only in grub.cfg); reboot 16:17:47. Screens: GRUB selected Felhom one-shot: 7.0.14-20-pve (shots/s0115.png), the kernel died, the next GRUB menu selected the default (s0140.png). Back on 7.0.2-6 156 s after the reboot; boot gap 41 s. kernel-boot → fell_back; hub logged it and mailed the operator (os_kernel_step). Flag empty. Crash guard: unclean_boots_in_window 0 (the panic came before userspace — the planned reboot's clean-stop marker still stood, exactly as test_a_panic_before_userspace_is_not_even_counted says). |
PASS — no person |
| (c) booted but unhealthy | 7.0.2-6 → 7.0.14-20 | reset BY HAND first (meta-package back to 7.0.2-6, the step record moved aside — b-panic/5-reset-by-hand.txt). Staged 16:35:52; guest held BY HAND (pct set 9201 --onboot 0); reboot. Up on 7.0.14-20 after 131 s; 16:38:43 judging, the hub reached at once; the guest never ran → at 16:58:54 (20 min 11 s) health_failed to the hub, then ONE kernel-revert; back on 7.0.2-6 at 16:59:22; kernel-boot → self_reverted (reason: vmid 9201 is not running). Guard 0 unclean. Then onboot 1 + pct start by hand. |
PASS — no person |
| (a) healthy | 7.0.2-6 → 7.0.14-22 | the step record moved aside BY HAND (the 20 h one-step-per-night rule). Staged 17:15:18 (2 packages: the meta and the new signed image, 28.3 s); reboot; up on 7.0.14-22 after 131 s; judging 17:17:53; healthy 1 m 17 s after the agent started → kernel-good: the default proved from grub.cfg = 7.0.14-22 (zz-felhom-kernel-default.cfg), outcome applied. Every container healthy. |
PASS |
State left: Tester 1 runs 7.0.14-22-pve as its GRUB default (booted healthily), flag empty, step good, guest running,
onboot 1. Kernels installed: 7.0.2-6, 7.0.14-20, 7.0.14-22.
Option C (E 2): UNMEASURED. The Proxmox kernel ships no lockup test module: CONFIG_TEST_LOCKUP is not set in
/boot/config-7.0.14-20-pve on demo-hp, modinfo test_lockup → not found on both demo boxes. No tool was built (the
brief). The options are on the one-shot entry (grub.cfg line in a-healthy/c-unhealthy setups); whether a soft
lockup becomes a panic there is not shown.
Small things seen: the self-revert's reason is the wrapper's raw refusal JSON (no health reading: {"code": "R10", …}) — readable, not pretty; left as is. The System page's running/next-boot cells read "unknown" for a few minutes after
each boot until the next facts read (10 min cadence) — by design.