kernel lane: report, STATUS, CONTEXT, R-836 where-it-stopped, R-898 filed; 126 -> 128
gates / gates (push) Successful in 3m11s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-07 18:19:03 +02:00
parent 2844d0b0e6
commit 3691ac2307
7 changed files with 117 additions and 32 deletions
+12
View File
@@ -16,6 +16,18 @@
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
> **2026-10-07 (evening) — the kernel lane BUILT (R-836, `09` §3 decisions 172–173; `11` §5.11).** Agent v0.152.0 (+ step
> bundle `0.152.0-step1`, + bundle) on demo-hp, demo-felhom, Tester 1; hub v0.143.0 (deployed attended). One-shot flag on
> the ESP, option C on the one-shot entry, default pinned to the running kernel and moved only by a healthy boot (host
> rule + hub reached, 20 min, measured 68 s / 272 s), ONE self-revert; the hub mails the household 09–20 h the day
> before and sets `os_update.kernel.tonight` only after an accepted mail (no mail, no step). Proven on Tester 1 (panic →
> fell_back; held guest → self_reverted after 20 min; healthy → 7.0.14-22 default). R-32 → P4 (leftovers left).
> **Decided by CC unattended — operator may reverse:** the mail window 09:00–20:00 Budapest, at most once per 20 h and 3
> times per kernel, to the REGISTERED address; the judge wait 20 min; a refusal R20/R23/R3/R12 or any ended step is never
> retried by itself; ring 1 stages only by a signed `os_kernel_step` (no reboot) and reboots on its own told night.
> Open: R-898 (ring 0's told kernel can be a day old — tonight demo-felhom refuses 7.0.14-20 safely; 7.0.14-22 due on
> both demo boxes 2026-10-08), R-897 (post-reboot drive re-bind races the first backup). Register 126 → 128.
> **2026-10-07 (day) — the operator's answers built; the kernel spike (`09` §3 162–171).** Hub 0.142.0 + agent 0.151.0
> (+ bundle) on the three boxes: the Proxmox package lane (R-812 A, proven on demo-felhom), the controller-image verb
> (R-861 a+b), RESET's purge (R-32; the one-time clean-up waits for a main-account login), the other-key line (R-366),
+55 -29
View File
@@ -1,38 +1,64 @@
# REPORT — the operator's answers built; the kernel lane spiked with real reboots (2026-10-07 day)
# REPORT — the kernel lane, 2026-10-07 (evening)
| Part | Result |
|---|---|
| **A** — rulings | `09` §3 decisions 162–170 recorded first, plus 171 (reboots without a per-reboot word, never DooPlex). R-822 closed by ruling. R-836 → P2. Shared rule file: no hub build/deploy without the operator (all five copies identical) |
| **B** — Proxmox package lane (R-812 A) | built (agent `pve` layer + `/etc/pve` write gate; hub candidate/approval/System page), delivered, **proven live on demo-felhom**: one signed step, 65 packages in 70 s, `pve-manager` 9.2.2 → 9.2.21, healthy, guest untouched, app 31/31; hub logged it. R-812 closed |
| **C** — agent image check (R-861 (a) A1, (b) B2) | built, delivered to the three boxes; on demo-hp: no `tee` grant, the verb is the route, an `alpine` ref refused (rc 3, file unchanged), felhom-op's pct lines exact. **Not seen: a full managed swap** (no newer controller today) |
| **D** — RESET purge + clean-up (R-32) | purge through the sub-account's own login built and delivered (hub 0.142.0). **The one-time clean-up STOPPED:** ~2.7 GB belongs to no live customer, reachable only by the pool box's main account — no such login on DooPlex |
| **E** — other-key archives line (R-366 slice 2) | built and delivered; R-366 closed |
| **F** — retire the empty fields (R-105) | built and delivered (columns kept, unread); `05`/`06` corrected; R-105 closed |
| **G** — kernel spike (R-836) | 24 reboots + 2 power cycles on Tester 1 (VM), demo-felhom, demo-hp (Secure Boot on). **ESP one-shot flag: works on all three. UEFI BootNext: works on all three. A watchdog armed by systemd: fails on all three** (a reset clears the timer; a frozen kernel needs a person). Design with two operator questions: `audits/kernel-spike-2026-10-07/DESIGN-kernel-lane.md`. Boxes left clean on a healthily booted default |
| **H** — Tester 2 (read only) | the household got 0 mails while the box was off (operator 10); no row |
| Rulings | Decisions 172 (kernel lane yes: A + C, night restarts, household mailed the day before, a freeze needing a person accepted) and 173 (R-32 leftovers left; P2 → P4, blocked on a safe main-account path) recorded FIRST (`0c5dbfbf`). |
| A — the one-shot boot | **Done.** Wrapper layer `kernel` (agent v0.152.0) + the two GRUB generators in the bundle (option C on the one-shot entry only). Red tests first-class: a box without the ESP flag refused (R20), a kernel outside the set refused (R23), the snippet clears the flag before booting, an install never moves the default. Red-proof `audits/kernel-lane-2026-10-07/A/redproof.txt` (16 mutations, each red). |
| B — boot good, the self-revert | **Done.** Health = the host rule + the hub reached, 20 min (measured: all healthy 68 s after a reboot on demo-felhom, 272 s on demo-hp; the spike's > 6 min Tailscale delay did not recur — no row). A box that never returns raises the hub's existing `host_stale` at 45 min and `host_down` at 90 min. One self-revert per step. Crash guard: a step adds at most ONE unclean boot (a panic before userspace adds none) — pinned by `KernelStepCannotLeaveTheBoxOff`, and seen live (0 counted after the panic). |
| C — night slot, approval, System page | **Done.** The kernel step ends the night leg (after a healthy host step, after the backup, trigger night, once per night); ring 1 stages only by a signed job. Hub v0.143.0 (deployed with you attending): due boxes, "Approve kernel set", two System page cells. `11` §5.11 + §8 step 6 written. |
| D — the household mail | **Done.** One mail the day before, 09:00–20:00 Budapest, its language, registered address; `tonight` only after the mail service accepted it (no mail, no step). Sent for real to demo-felhom at 18:12 (Gmail). Parity green (both locales; tests). |
| E 1 — Tester 1 by hand | **PASS ×3** (`audits/kernel-lane-2026-10-07/E/RESULT.md`): forced panic → back on the old kernel by itself (screens); held guest → judged 20 min → one self-revert; healthy → 7.0.14-22 the default. Tester 1 left on 7.0.14-22 (healthy, default). |
| E 2 — option C | **Unmeasured:** the Proxmox kernel has no lockup test module (`CONFIG_TEST_LOCKUP` not set). No tool built. |
| E 3 — ring-0 night run | **Not yet — next session.** Tonight demo-felhom was told about 7.0.14-20 (from last night's report) but its sources now offer 7.0.14-22: the step will refuse (R23) before any change (R-898). demo-hp's lists were a day old. Both demo boxes are due for 7.0.14-22 on the night 2026-10-08→09. |
| E 4 — Tester 1 as ring 1 | **Not yet** (after E3; Tester 1 already runs 7.0.14-22, so the ring-1 proof needs the next kernel). |
| Small | R-861 (a): no controller release in this arc — not read back. Tailscale delay: not reproduced → no row. |
| Rows before | Rows after | Opened | Closed |
|---|---|---|---|
| **130** | **126** | **0** | **4** (R-822, R-812, R-366, R-105) |
**Rows: 126 before → 128 after. Opened 2 (R-897, R-898). Closed 0.**
**Releases (one per repo changed):** hub **0.142.0** (live, healthz 200 within 50 s); agent **0.151.0** + bundle on demo-hp,
demo-felhom, Tester 1 (probe 68/68 each). **The agent vouch was refused by the hub** (`golden_behind_fleet`: golden 0.301.0
is behind the fleet's controller 0.302.0), so new installs keep agent 0.150.0 until the weekly bake. No controller release
(no controller code changed). CI: one lost job (felhom.eu run 1497, every step failed with no log) re-run once → success
(R-887 count +1).
**Releases:** agent v0.152.0 (tag `d03ab7f`; binary `95ff4220…`, bundle `f0c2cec3…`, step bundle `0.152.0-step1`
`0b71d32b…`), delivered by signed jobs to demo-hp, demo-felhom and Tester 1 (binary → step bundle → bundle). Hub v0.143.0
(`ab3b7ea2`, deployed `0ece2ff7` 15:39, Synced/Healthy). The installer's uninstall now knows the two GRUB files
(unreleased, no tag cut). Nothing on ep0, Tester 2 or DooPlex's own system; no Docker engine step; no delivery after 02:00.
**Helpers:** each prompt carried the brief's fences in full (rule 11). **Security note** on `hub/internal/api/handler.go`
(raised by a background review, no detail): the last three changes log box/customer ids and counts only — nothing fixed.
**Read for this report:** `11-os-updates.md` (§5.3, §5.8–5.10, §8.2 — §5.11 new), `03-host-agent.md`, `07` §6.1,
`audits/kernel-spike-2026-10-07/DESIGN-kernel-lane.md`.
**Teardown:** Tester 1 — spike entries, ESP flag, BootNext loader/entry and watchdog config removed; VM 341's test
watchdog device removed; default 7.0.2-6 (running). demo-felhom — the same; default 7.0.2-6; Proxmox userspace now 9.2.21.
demo-hp — the same; default 7.0.14-20. 7.0.14-20 is installed on Tester 1 and demo-felhom (booted healthily in the
spike) but not their default. Hub — nothing provisioned.
**Also seen:** the hub image was built while one audit text file (outside `hub/`) was uncommitted; the build copies
`hub/` only, and `hub/` equalled the pushed commit (checked by diff). R-897: after demo-hp's timing reboot, one app's
backup failed in the second the agent re-bound the drive.
## Two decisions for you
## The household mail (both languages, as sent)
1. **The kernel lane's promise:** may a box restart at night for a kernel update, and do we tell the household? My pick:
yes, with one line in the household's mail the day before. If you do nothing: kernels stay manual.
2. **The pool box clean-up (~2.7 GB):** give me the pool box's main-account login for one session, or delete the
folders yourself in the Hetzner console. My pick: you delete them in the console (no new credential on DooPlex). If
you do nothing: the leftovers stay; nothing reads them.
**Hungarian** — subject: `[Felhom] Ma éjjel újraindul a Felhom dobozod`
> Szia!
>
> Ma éjjel a Felhom dobozod egy biztonsági frissítés miatt újraindul. Ilyenkor az alkalmazásaid néhány percig nem
> érhetők el, utána maguktól visszajönnek. Neked nem kell semmit tenned.
>
> Ha reggelre a doboz mégsem érhető el: húzd ki a tápkábelét, várj 10 másodpercet, és dugd vissza. A doboz ilyenkor
> magától a korábbi, bevált rendszerrel indul el.
>
> Ha ez sem segít, válaszolj erre a levélre, és segítünk.
>
> Üdv, Felhom.eu
**English** — subject: `[Felhom] Your Felhom box restarts tonight`
> Hi,
>
> Tonight your Felhom box restarts for a security update. Your apps will be away for a few minutes, then come back by
> themselves. You do not need to do anything.
>
> If the box is not back by morning: unplug its power cable, wait 10 seconds, and plug it back in. The box then starts
> its earlier, proven system by itself.
>
> If that does not help, reply to this e-mail and we will help.
>
> Best, Felhom.eu
## Decisions for you
1. **Is the mail text right?** My pick: keep it. If you do nothing: it goes out as written.
2. **Mail time window 09:00–20:00, at most 3 mails per kernel.** My pick: keep (decided by me, you may reverse). If you
do nothing: it stays.
+15 -2
View File
@@ -2,8 +2,21 @@
**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop) is off; nothing was sent to it.**
**Updated 2026-10-07 14:30: hub 0.142.0; demo-hp, demo-felhom and Tester 1 run agent 0.151.0 and controller 0.302.0.
The open-items list is at 126. Report: `REPORT.md`.**
**Updated 2026-10-07 18:40: hub 0.143.0; demo-hp, demo-felhom and Tester 1 run agent 0.152.0 and controller 0.302.0.
The open-items list is at 128. Report: `REPORT.md`.**
## Evening (2026-10-07): the kernel lane is built
- **A new kernel boots once.** If it crashes, the box comes back on the old kernel by itself. If it boots well, it
becomes the default. If it boots but the box is not healthy within 20 minutes, the box restarts once into the old
kernel by itself. All three were proven on the Tester 1 box today.
- **The household is mailed the day before**, between 9:00 and 20:00, in its language. No mail → no restart.
- **You approve each kernel** on the System page after both demo boxes booted it well at night.
- **Tonight:** demo-felhom's household was mailed, but for yesterday's kernel; the box will safely refuse it, because a
newer one appeared. Tomorrow both demo boxes are due for the newest kernel (filed as a small item).
- **Found:** after a restart, the first backup can fail for one app on a drive (the drive is being re-attached). Filed.
**Needs you (none urgent):** read the household mail text in `REPORT.md`. If nothing: it stays as written.
## Day (2026-10-07): your answers built, and the kernel test with real reboots
@@ -6,3 +6,19 @@ uploaded signed op to the hub jobs queue
signed: op=agent_update host=demo-felhom-8363b5 guest="" key_id=felhom-op-1 nonce=1064b6dde1dd050dd5f6f7b788990412 expires=2026-10-07T16:04:42Z
wrote envelope to /dev/null
uploaded signed op to the hub jobs queue
== agent_config_update 0.152.0-step1 → demo-hp-bb76ea, 2026-10-07T15:31:38Z
signed: op=agent_config_update host=demo-hp-bb76ea guest="" key_id=felhom-op-1 nonce=e87c668823146debf011a1680691118c expires=2026-10-07T16:16:38Z
wrote envelope to /dev/null
uploaded signed op to the hub jobs queue
== agent_config_update 0.152.0-step1 → demo-felhom-8363b5, 2026-10-07T15:31:38Z
signed: op=agent_config_update host=demo-felhom-8363b5 guest="" key_id=felhom-op-1 nonce=699531f773854f791e3517174cf382ad expires=2026-10-07T16:16:38Z
wrote envelope to /dev/null
uploaded signed op to the hub jobs queue
== agent_config_update 0.152.0 → demo-hp-bb76ea, 2026-10-07T15:47:04Z
signed: op=agent_config_update host=demo-hp-bb76ea guest="" key_id=felhom-op-1 nonce=b126b41931b7123998dcc6f3e182e1dc expires=2026-10-07T16:32:04Z
wrote envelope to /dev/null
uploaded signed op to the hub jobs queue
== agent_config_update 0.152.0 → demo-felhom-8363b5, 2026-10-07T15:47:04Z
signed: op=agent_config_update host=demo-felhom-8363b5 guest="" key_id=felhom-op-1 nonce=ac6fde774fe15cfdc4671fbbeddc502f expires=2026-10-07T16:32:04Z
wrote envelope to /dev/null
uploaded signed op to the hub jobs queue
@@ -0,0 +1,14 @@
2026/10/07 16:17:16 [INFO] osupdates: tester-1-d70be4 reported kernel run 20261007T141703Z: ring=1 mode=apply outcome=staged healthy=true upgraded=1 pending=0 not-covered=0 restart-needed=0 reboot-needed=true wrapper=13.2s
2026/10/07 16:20:32 [INFO] osupdates: tester-1-d70be4 reported kernel run boot-20261007T142032Z: ring=1 mode=kernel-boot outcome=fell_back healthy=false upgraded=0 pending=0 not-covered=0 restart-needed=0 reboot-needed=false wrapper=0.0s
2026/10/07 16:20:32 [INFO] Operator email sent for tester-1/os_kernel_step
2026/10/07 16:35:52 [INFO] osupdates: tester-1-d70be4 reported kernel run 20261007T143539Z: ring=1 mode=apply outcome=staged healthy=true upgraded=1 pending=0 not-covered=0 restart-needed=0 reboot-needed=true wrapper=12.7s
2026/10/07 16:38:43 [INFO] osupdates: tester-1-d70be4 reported kernel run boot-20261007T143843Z: ring=1 mode=kernel-boot outcome=judging healthy=false upgraded=0 pending=0 not-covered=0 restart-needed=0 reboot-needed=false wrapper=0.0s
2026/10/07 16:58:54 [INFO] osupdates: tester-1-d70be4 reported kernel run boot-20261007T143843Z: ring=1 mode=kernel-revert outcome=health_failed healthy=false upgraded=0 pending=0 not-covered=0 restart-needed=0 reboot-needed=false wrapper=0.0s
2026/10/07 16:58:54 [INFO] Operator email sent for tester-1/os_kernel_step
2026/10/07 16:59:30 [INFO] osupdates: tester-1-d70be4 reported kernel run boot-20261007T145930Z: ring=1 mode=kernel-boot outcome=self_reverted healthy=false upgraded=0 pending=0 not-covered=0 restart-needed=0 reboot-needed=false wrapper=0.0s
2026/10/07 16:59:31 [INFO] Operator email sent for tester-1/os_kernel_step
2026/10/07 17:15:18 [INFO] osupdates: tester-1-d70be4 reported kernel run 20261007T151439Z: ring=1 mode=apply outcome=staged healthy=true upgraded=2 pending=0 not-covered=0 restart-needed=0 reboot-needed=true wrapper=39.0s
2026/10/07 17:17:53 [INFO] osupdates: tester-1-d70be4 reported kernel run boot-20261007T151753Z: ring=1 mode=kernel-boot outcome=judging healthy=false upgraded=0 pending=0 not-covered=0 restart-needed=0 reboot-needed=false wrapper=0.0s
2026/10/07 17:19:10 [INFO] osupdates: tester-1-d70be4 reported kernel run boot-20261007T151753Z: ring=1 mode=kernel-good outcome=applied healthy=true upgraded=0 pending=0 not-covered=0 restart-needed=0 reboot-needed=false wrapper=0.0s
2026/10/07 18:12:11 [INFO] kernel notice (7.0.14-20-pve) emailed to the household of demo-felhom (lang=hu)
2026/10/07 18:12:11 [INFO] osupdates: kernel 7.0.14-20-pve — the household of demo-felhom-8363b5 was told the box restarts tonight (lang=hu)
@@ -1,5 +1,8 @@
# The kernel lane — what the spike measured, and a design (R-836, R-812 option B; `09` §3 decisions 164, 171)
> **Ruled 2026-10-07 14:43 (decision 172): A with C. BUILT the same day — `11-os-updates.md` §5.11, proofs in
> `audits/kernel-lane-2026-10-07/`.**
Spike 2026-10-07, operator present. Boxes: the Tester 1 box (VM 341 on the HP box, OVMF, Secure Boot off), demo-felhom
(Intel N100, Secure Boot off), demo-hp (AMD, Secure Boot ON). Kernels: 7.0.2-6 and 7.0.14-20. 24 reboots, 2 operator
power cycles. Evidence per boot in this folder (`<box>/<step>/before.txt`, `after.txt`, `result.txt`; screenshots for
+2 -1
View File
@@ -215,7 +215,8 @@ stopping line that lies.
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|---|---|---|---|---|---|---|---|
| **R-836** | Box system & updates | P2 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin <new> --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot <new>`: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** **2026-10-07 07:58: raised P3 → P2 (`09` §3 decision 164 — the kernel lane before the first paying customer); the spike runs with the operator present.** **2026-10-07 SPIKE DONE (operator present; 24 reboots, 2 power cycles; `audits/kernel-spike-2026-10-07/DESIGN-kernel-lane.md`).** On Tester 1 (VM), demo-felhom and demo-hp (Secure Boot on): **a one-shot flag in a GRUB env block on the ESP (vfat) WORKS** — new kernel once, then the old one; a panicking new kernel (`panic=10`) falls back with no person. **UEFI `BootNext` WORKS too.** **A hardware watchdog armed by systemd during the reboot FAILS on all three** (i6300ESB, Intel TCO, AMD SP5100): the reset clears the timer, so a FROZEN new kernel needs a power cycle. Pick: the ESP flag + lockup-to-panic kernel options (unmeasured). Two operator questions in the design (night restarts and the household; a freeze needs a person). Boxes left clean on a healthily booted default. **2026-10-07 14:43 — `09` §3 decision 172: the kernel lane is YES — option A (the ESP one-shot flag) with option C (lockup → panic) on the one-shot entry; a box may restart at night; the household is mailed the day before (what to do if it is not back: unplug, plug back in); each kernel set is operator-approved after ring 0; a frozen new kernel needing a power cycle is ACCEPTED for the first customers. BUILD STARTED 2026-10-07 (the kernel-lane arc).** | — | — | CC |
| **R-836** | Box system & updates | P2 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin <new> --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot <new>`: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** **2026-10-07 07:58: raised P3 → P2 (`09` §3 decision 164 — the kernel lane before the first paying customer); the spike runs with the operator present.** **2026-10-07 SPIKE DONE (operator present; 24 reboots, 2 power cycles; `audits/kernel-spike-2026-10-07/DESIGN-kernel-lane.md`).** On Tester 1 (VM), demo-felhom and demo-hp (Secure Boot on): **a one-shot flag in a GRUB env block on the ESP (vfat) WORKS** — new kernel once, then the old one; a panicking new kernel (`panic=10`) falls back with no person. **UEFI `BootNext` WORKS too.** **A hardware watchdog armed by systemd during the reboot FAILS on all three** (i6300ESB, Intel TCO, AMD SP5100): the reset clears the timer, so a FROZEN new kernel needs a power cycle. Pick: the ESP flag + lockup-to-panic kernel options (unmeasured). Two operator questions in the design (night restarts and the household; a freeze needs a person). Boxes left clean on a healthily booted default. **2026-10-07 14:43 — `09` §3 decision 172: the kernel lane is YES — option A (the ESP one-shot flag) with option C (lockup → panic) on the one-shot entry; a box may restart at night; the household is mailed the day before (what to do if it is not back: unplug, plug back in); each kernel set is operator-approved after ring 0; a frozen new kernel needing a power cycle is ACCEPTED for the first customers. BUILD STARTED 2026-10-07 (the kernel-lane arc).** **2026-10-07 evening: BUILT and proven by hand** — agent v0.152.0 + hub v0.143.0 (`11` §5.11), delivered to demo-hp, demo-felhom and Tester 1. Tester 1 (`audits/kernel-lane-2026-10-07/E/RESULT.md`): a forced panic on the one-shot → back on the old kernel by itself (fell_back, 0 unclean boots counted); a held guest → judged 20 min → ONE self-revert → old kernel; a healthy boot → 7.0.14-22 the default. The household mail went out for demo-felhom 18:12 (Gmail). Option C UNMEASURED (no test_lockup module). **WHERE IT STOPPED (next session):** the ring-0 NIGHT run (Part E3) — tonight demo-felhom refuses the told 7.0.14-20 (R-898, safe); 7.0.14-22 is due on both demo boxes for the night 2026-10-08→09 (mail 2026-10-08 daytime); read back the morning of 2026-10-09; then approve the set on the System page and deliver to the Tester 1 box as ring 1 by a signed os_kernel_step (Part E4 — Tester 1 already runs 7.0.14-22, so the ring-1 proof needs the NEXT kernel). | — | — | CC |
| **R-898** | Box system & updates | P3 | **Ring 0's told kernel can be a day old: the hub names the kernel from the box's LAST host report, the box's sources may offer a newer one by night, and the step then refuses (R23) — one night and one household mail lost.** SEEN 2026-10-07 18:12: demo-felhom's household was mailed for 7.0.14-20-pve (from the night report) while the box's apt already offered 7.0.14-22; the night leg would stage `pending-kernel` → 22 ≠ the told 20 → R23 before any change (safe; read from the code, `felhom-os-apply` Kernel.stage). Fix direction (needs an agent + hub release, so not in this session): ring 0 stages EXACTLY the told kernel (select `listed`, the set derived from the kver — Proxmox keeps old kernel versions), and the hub treats a told-kernel mismatch as transient, not „ended for good". | **OPEN — filed 2026-10-07 (kernel-lane arc); tonight's demo-felhom run shows it.** | — | — | CC |
| **R-862** | Box system & updates | P3 | **Tester 2 cannot take the config bundle until one by-hand bootstrap is done: its `felhom-os-apply` (agent 0.142.0) predates the bundle mode, and no signed job can write a root file on it.** FOUND 2026-10-04 (R-840 build, Part C): the route reaches every box installed from installer 1.31.0 on, and the demo boxes (bootstrapped by CC); Tester 2 has every root file of agent 0.142.0 (installer 1.30.0) and lacks only the R-858 wrapper fix, which matters only for a Docker step it gets solely from a signed job. CC has no route to Tester 2 (its door admits only the operator's WireGuard peer; `felhom-op` cannot become root). The steps: `runbooks/config-bundle.md` "Tester 2". Then CC signs the bundle and reads it back. | **WAITING-ON-OPERATOR** — the operator said (2026-10-04 ~19:05) he will try through his tunnel **2026-10-07 07:58 (`09` §3 decision 170): the operator believes Tester 2 never set up the off-site backup (no escrow); the laptop may stay off for days; the tester plans to add an HDD.** **2026-10-07 (Part H, read only):** in Tester 2's last 10 notifications (2026-10-05 09:16 → 2026-10-07 06:58) the household got 0 mails, the operator 10 (host_down ×9, expected_dbdump_missed ×1); control: the household channel shows on the other three customers. So the household is not mailed daily while the box is off; no row (`audits/day-2026-10-07/H/`). | the operator's WireGuard tunnel | the operator runs the three bootstrap commands; CC sends the bundle | operator |
| **R-35** | Box system & updates | P3 | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-05 (burn-down night): NEEDS A DESIGN.** Hot-apply needs a per-setting ruling; persisting sessions puts login tokens on disk (and into the whole-guest archive, `07` §5). Next: the operator's pick between them. | — | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` | CC |
| **R-78** | Box system & updates | P3 | **`local_api` authority ruling — auto-reconcile vs detect-only** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-25, size M, roadmap state `idea (deferred OUT of R-77 on purpose)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **An OWED OPERATOR DECISION, not a defect — moved because an owed ruling hidden among feature ideas is the shape this session exists to remove.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | R-77 ships detection because the fix is genuinely undecided, and **both directions can lose customer-visible function**. **Direction 1 (today):** `controller.yaml` wins and drift is silent → the 2026-07-25 island migration blinded the whole fleet's control plane for 17.5 h (drive gate, guest-reboot recovery, quiesce/backup all degrade). R-77 makes that loud but does not stop it recurring. **Direction 2 (`bootstrap.json` wins, auto-reconcile on boot):** a guest whose `controller.yaml` is CORRECT and whose `bootstrap.json` is stale — a half-completed re-provision, a hand-repaired guest, a setup-wizard box — gets a **working channel clobbered on the next restart**, fleet-wide and silently, during a routine deploy. That is not obviously better than the bug. Needs a spike: which writer is authoritative per field (endpoint vs fingerprint vs token — `mergeLocalAPI` replaces the whole block, so they cannot be reconciled independently today), whether the agent side should stamp a generation/mtime so 'newer wins' is even expressible, and whether reconcile should require an operator ack. Until then the drift alert plus a manual edit is the supported path. | CC |