2844d0b0e6
gates / gates (push) Successful in 3m35s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
26 lines
3.3 KiB
Markdown
26 lines
3.3 KiB
Markdown
# Part E 1 — the kernel lane by hand on the Tester 1 box (VM 341 on demo-hp), 2026-10-07
|
|
|
|
Agent v0.152.0 + bundle 0.152.0 (delivered by signed jobs: binary, step bundle `0.152.0-step1`, bundle). Hub v0.143.0.
|
|
Every stage went through the PRODUCT's ring-1 path: a signed `os_kernel_step` (felhom-opsign → hub jobs queue → the
|
|
agent's `KernelStepExecutor` → the wrapper, which re-verifies the signature). Each REBOOT was started by hand with the
|
|
wrapper's own `kernel-reboot` mode (the night leg does the same call on a told night). The after-boot judging was the
|
|
agent's own, untouched. Order run: (b), (c), (a) — a crash and a revert must start from the old kernel.
|
|
|
|
| Test | Kernel | What happened (times local) | Verdict |
|
|
|---|---|---|---|
|
|
| (b) forced panic | 7.0.2-6 → 7.0.14-20 | staged 16:17:16; the one-shot entry broken BY HAND (`rdinit=` and `init=` both missing, only in grub.cfg); reboot 16:17:47. Screens: GRUB selected `Felhom one-shot: 7.0.14-20-pve` (`shots/s0115.png`), the kernel died, the next GRUB menu selected the default (`s0140.png`). Back on 7.0.2-6 156 s after the reboot; boot gap 41 s. `kernel-boot` → `fell_back`; hub logged it and mailed the operator (`os_kernel_step`). Flag empty. Crash guard: `unclean_boots_in_window` 0 (the panic came before userspace — the planned reboot's clean-stop marker still stood, exactly as `test_a_panic_before_userspace_is_not_even_counted` says). | **PASS** — no person |
|
|
| (c) booted but unhealthy | 7.0.2-6 → 7.0.14-20 | reset BY HAND first (meta-package back to 7.0.2-6, the step record moved aside — `b-panic/5-reset-by-hand.txt`). Staged 16:35:52; guest held BY HAND (`pct set 9201 --onboot 0`); reboot. Up on 7.0.14-20 after 131 s; 16:38:43 `judging`, the hub reached at once; the guest never ran → at 16:58:54 (20 min 11 s) `health_failed` to the hub, then ONE `kernel-revert`; back on 7.0.2-6 at 16:59:22; `kernel-boot` → `self_reverted` (reason: `vmid 9201 is not running`). Guard 0 unclean. Then `onboot 1` + `pct start` by hand. | **PASS** — no person |
|
|
| (a) healthy | 7.0.2-6 → 7.0.14-22 | the step record moved aside BY HAND (the 20 h one-step-per-night rule). Staged 17:15:18 (2 packages: the meta and the new signed image, 28.3 s); reboot; up on 7.0.14-22 after 131 s; `judging` 17:17:53; healthy 1 m 17 s after the agent started → `kernel-good`: the default proved from grub.cfg = 7.0.14-22 (`zz-felhom-kernel-default.cfg`), outcome `applied`. Every container healthy. | **PASS** |
|
|
|
|
**State left:** Tester 1 runs 7.0.14-22-pve as its GRUB default (booted healthily), flag empty, step `good`, guest running,
|
|
`onboot 1`. Kernels installed: 7.0.2-6, 7.0.14-20, 7.0.14-22.
|
|
|
|
**Option C (E 2): UNMEASURED.** The Proxmox kernel ships no lockup test module: `CONFIG_TEST_LOCKUP is not set` in
|
|
`/boot/config-7.0.14-20-pve` on demo-hp, `modinfo test_lockup` → not found on both demo boxes. No tool was built (the
|
|
brief). The options are on the one-shot entry (`grub.cfg` line in `a-healthy`/`c-unhealthy` setups); whether a soft
|
|
lockup becomes a panic there is not shown.
|
|
|
|
**Small things seen:** the self-revert's reason is the wrapper's raw refusal JSON (`no health reading: {"code": "R10",
|
|
…}`) — readable, not pretty; left as is. The System page's running/next-boot cells read "unknown" for a few minutes after
|
|
each boot until the next facts read (10 min cadence) — by design.
|