kernel lane: report, STATUS, CONTEXT, R-836 where-it-stopped, R-898 filed; 126 -> 128
gates / gates (push) Successful in 3m11s
gates / gates (push) Successful in 3m11s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
+12
@@ -16,6 +16,18 @@
|
||||
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
|
||||
|
||||
|
||||
> **2026-10-07 (evening) — the kernel lane BUILT (R-836, `09` §3 decisions 172–173; `11` §5.11).** Agent v0.152.0 (+ step
|
||||
> bundle `0.152.0-step1`, + bundle) on demo-hp, demo-felhom, Tester 1; hub v0.143.0 (deployed attended). One-shot flag on
|
||||
> the ESP, option C on the one-shot entry, default pinned to the running kernel and moved only by a healthy boot (host
|
||||
> rule + hub reached, 20 min, measured 68 s / 272 s), ONE self-revert; the hub mails the household 09–20 h the day
|
||||
> before and sets `os_update.kernel.tonight` only after an accepted mail (no mail, no step). Proven on Tester 1 (panic →
|
||||
> fell_back; held guest → self_reverted after 20 min; healthy → 7.0.14-22 default). R-32 → P4 (leftovers left).
|
||||
> **Decided by CC unattended — operator may reverse:** the mail window 09:00–20:00 Budapest, at most once per 20 h and 3
|
||||
> times per kernel, to the REGISTERED address; the judge wait 20 min; a refusal R20/R23/R3/R12 or any ended step is never
|
||||
> retried by itself; ring 1 stages only by a signed `os_kernel_step` (no reboot) and reboots on its own told night.
|
||||
> Open: R-898 (ring 0's told kernel can be a day old — tonight demo-felhom refuses 7.0.14-20 safely; 7.0.14-22 due on
|
||||
> both demo boxes 2026-10-08), R-897 (post-reboot drive re-bind races the first backup). Register 126 → 128.
|
||||
|
||||
> **2026-10-07 (day) — the operator's answers built; the kernel spike (`09` §3 162–171).** Hub 0.142.0 + agent 0.151.0
|
||||
> (+ bundle) on the three boxes: the Proxmox package lane (R-812 A, proven on demo-felhom), the controller-image verb
|
||||
> (R-861 a+b), RESET's purge (R-32; the one-time clean-up waits for a main-account login), the other-key line (R-366),
|
||||
|
||||
@@ -1,38 +1,64 @@
|
||||
# REPORT — the operator's answers built; the kernel lane spiked with real reboots (2026-10-07 day)
|
||||
# REPORT — the kernel lane, 2026-10-07 (evening)
|
||||
|
||||
| Part | Result |
|
||||
|---|---|
|
||||
| **A** — rulings | `09` §3 decisions 162–170 recorded first, plus 171 (reboots without a per-reboot word, never DooPlex). R-822 closed by ruling. R-836 → P2. Shared rule file: no hub build/deploy without the operator (all five copies identical) |
|
||||
| **B** — Proxmox package lane (R-812 A) | built (agent `pve` layer + `/etc/pve` write gate; hub candidate/approval/System page), delivered, **proven live on demo-felhom**: one signed step, 65 packages in 70 s, `pve-manager` 9.2.2 → 9.2.21, healthy, guest untouched, app 31/31; hub logged it. R-812 closed |
|
||||
| **C** — agent image check (R-861 (a) A1, (b) B2) | built, delivered to the three boxes; on demo-hp: no `tee` grant, the verb is the route, an `alpine` ref refused (rc 3, file unchanged), felhom-op's pct lines exact. **Not seen: a full managed swap** (no newer controller today) |
|
||||
| **D** — RESET purge + clean-up (R-32) | purge through the sub-account's own login built and delivered (hub 0.142.0). **The one-time clean-up STOPPED:** ~2.7 GB belongs to no live customer, reachable only by the pool box's main account — no such login on DooPlex |
|
||||
| **E** — other-key archives line (R-366 slice 2) | built and delivered; R-366 closed |
|
||||
| **F** — retire the empty fields (R-105) | built and delivered (columns kept, unread); `05`/`06` corrected; R-105 closed |
|
||||
| **G** — kernel spike (R-836) | 24 reboots + 2 power cycles on Tester 1 (VM), demo-felhom, demo-hp (Secure Boot on). **ESP one-shot flag: works on all three. UEFI BootNext: works on all three. A watchdog armed by systemd: fails on all three** (a reset clears the timer; a frozen kernel needs a person). Design with two operator questions: `audits/kernel-spike-2026-10-07/DESIGN-kernel-lane.md`. Boxes left clean on a healthily booted default |
|
||||
| **H** — Tester 2 (read only) | the household got 0 mails while the box was off (operator 10); no row |
|
||||
| Rulings | Decisions 172 (kernel lane yes: A + C, night restarts, household mailed the day before, a freeze needing a person accepted) and 173 (R-32 leftovers left; P2 → P4, blocked on a safe main-account path) recorded FIRST (`0c5dbfbf`). |
|
||||
| A — the one-shot boot | **Done.** Wrapper layer `kernel` (agent v0.152.0) + the two GRUB generators in the bundle (option C on the one-shot entry only). Red tests first-class: a box without the ESP flag refused (R20), a kernel outside the set refused (R23), the snippet clears the flag before booting, an install never moves the default. Red-proof `audits/kernel-lane-2026-10-07/A/redproof.txt` (16 mutations, each red). |
|
||||
| B — boot good, the self-revert | **Done.** Health = the host rule + the hub reached, 20 min (measured: all healthy 68 s after a reboot on demo-felhom, 272 s on demo-hp; the spike's > 6 min Tailscale delay did not recur — no row). A box that never returns raises the hub's existing `host_stale` at 45 min and `host_down` at 90 min. One self-revert per step. Crash guard: a step adds at most ONE unclean boot (a panic before userspace adds none) — pinned by `KernelStepCannotLeaveTheBoxOff`, and seen live (0 counted after the panic). |
|
||||
| C — night slot, approval, System page | **Done.** The kernel step ends the night leg (after a healthy host step, after the backup, trigger night, once per night); ring 1 stages only by a signed job. Hub v0.143.0 (deployed with you attending): due boxes, "Approve kernel set", two System page cells. `11` §5.11 + §8 step 6 written. |
|
||||
| D — the household mail | **Done.** One mail the day before, 09:00–20:00 Budapest, its language, registered address; `tonight` only after the mail service accepted it (no mail, no step). Sent for real to demo-felhom at 18:12 (Gmail). Parity green (both locales; tests). |
|
||||
| E 1 — Tester 1 by hand | **PASS ×3** (`audits/kernel-lane-2026-10-07/E/RESULT.md`): forced panic → back on the old kernel by itself (screens); held guest → judged 20 min → one self-revert; healthy → 7.0.14-22 the default. Tester 1 left on 7.0.14-22 (healthy, default). |
|
||||
| E 2 — option C | **Unmeasured:** the Proxmox kernel has no lockup test module (`CONFIG_TEST_LOCKUP` not set). No tool built. |
|
||||
| E 3 — ring-0 night run | **Not yet — next session.** Tonight demo-felhom was told about 7.0.14-20 (from last night's report) but its sources now offer 7.0.14-22: the step will refuse (R23) before any change (R-898). demo-hp's lists were a day old. Both demo boxes are due for 7.0.14-22 on the night 2026-10-08→09. |
|
||||
| E 4 — Tester 1 as ring 1 | **Not yet** (after E3; Tester 1 already runs 7.0.14-22, so the ring-1 proof needs the next kernel). |
|
||||
| Small | R-861 (a): no controller release in this arc — not read back. Tailscale delay: not reproduced → no row. |
|
||||
|
||||
| Rows before | Rows after | Opened | Closed |
|
||||
|---|---|---|---|
|
||||
| **130** | **126** | **0** | **4** (R-822, R-812, R-366, R-105) |
|
||||
**Rows: 126 before → 128 after. Opened 2 (R-897, R-898). Closed 0.**
|
||||
|
||||
**Releases (one per repo changed):** hub **0.142.0** (live, healthz 200 within 50 s); agent **0.151.0** + bundle on demo-hp,
|
||||
demo-felhom, Tester 1 (probe 68/68 each). **The agent vouch was refused by the hub** (`golden_behind_fleet`: golden 0.301.0
|
||||
is behind the fleet's controller 0.302.0), so new installs keep agent 0.150.0 until the weekly bake. No controller release
|
||||
(no controller code changed). CI: one lost job (felhom.eu run 1497, every step failed with no log) re-run once → success
|
||||
(R-887 count +1).
|
||||
**Releases:** agent v0.152.0 (tag `d03ab7f`; binary `95ff4220…`, bundle `f0c2cec3…`, step bundle `0.152.0-step1`
|
||||
`0b71d32b…`), delivered by signed jobs to demo-hp, demo-felhom and Tester 1 (binary → step bundle → bundle). Hub v0.143.0
|
||||
(`ab3b7ea2`, deployed `0ece2ff7` 15:39, Synced/Healthy). The installer's uninstall now knows the two GRUB files
|
||||
(unreleased, no tag cut). Nothing on ep0, Tester 2 or DooPlex's own system; no Docker engine step; no delivery after 02:00.
|
||||
|
||||
**Helpers:** each prompt carried the brief's fences in full (rule 11). **Security note** on `hub/internal/api/handler.go`
|
||||
(raised by a background review, no detail): the last three changes log box/customer ids and counts only — nothing fixed.
|
||||
**Read for this report:** `11-os-updates.md` (§5.3, §5.8–5.10, §8.2 — §5.11 new), `03-host-agent.md`, `07` §6.1,
|
||||
`audits/kernel-spike-2026-10-07/DESIGN-kernel-lane.md`.
|
||||
|
||||
**Teardown:** Tester 1 — spike entries, ESP flag, BootNext loader/entry and watchdog config removed; VM 341's test
|
||||
watchdog device removed; default 7.0.2-6 (running). demo-felhom — the same; default 7.0.2-6; Proxmox userspace now 9.2.21.
|
||||
demo-hp — the same; default 7.0.14-20. 7.0.14-20 is installed on Tester 1 and demo-felhom (booted healthily in the
|
||||
spike) but not their default. Hub — nothing provisioned.
|
||||
**Also seen:** the hub image was built while one audit text file (outside `hub/`) was uncommitted; the build copies
|
||||
`hub/` only, and `hub/` equalled the pushed commit (checked by diff). R-897: after demo-hp's timing reboot, one app's
|
||||
backup failed in the second the agent re-bound the drive.
|
||||
|
||||
## Two decisions for you
|
||||
## The household mail (both languages, as sent)
|
||||
|
||||
1. **The kernel lane's promise:** may a box restart at night for a kernel update, and do we tell the household? My pick:
|
||||
yes, with one line in the household's mail the day before. If you do nothing: kernels stay manual.
|
||||
2. **The pool box clean-up (~2.7 GB):** give me the pool box's main-account login for one session, or delete the
|
||||
folders yourself in the Hetzner console. My pick: you delete them in the console (no new credential on DooPlex). If
|
||||
you do nothing: the leftovers stay; nothing reads them.
|
||||
**Hungarian** — subject: `[Felhom] Ma éjjel újraindul a Felhom dobozod`
|
||||
|
||||
> Szia!
|
||||
>
|
||||
> Ma éjjel a Felhom dobozod egy biztonsági frissítés miatt újraindul. Ilyenkor az alkalmazásaid néhány percig nem
|
||||
> érhetők el, utána maguktól visszajönnek. Neked nem kell semmit tenned.
|
||||
>
|
||||
> Ha reggelre a doboz mégsem érhető el: húzd ki a tápkábelét, várj 10 másodpercet, és dugd vissza. A doboz ilyenkor
|
||||
> magától a korábbi, bevált rendszerrel indul el.
|
||||
>
|
||||
> Ha ez sem segít, válaszolj erre a levélre, és segítünk.
|
||||
>
|
||||
> Üdv, Felhom.eu
|
||||
|
||||
**English** — subject: `[Felhom] Your Felhom box restarts tonight`
|
||||
|
||||
> Hi,
|
||||
>
|
||||
> Tonight your Felhom box restarts for a security update. Your apps will be away for a few minutes, then come back by
|
||||
> themselves. You do not need to do anything.
|
||||
>
|
||||
> If the box is not back by morning: unplug its power cable, wait 10 seconds, and plug it back in. The box then starts
|
||||
> its earlier, proven system by itself.
|
||||
>
|
||||
> If that does not help, reply to this e-mail and we will help.
|
||||
>
|
||||
> Best, Felhom.eu
|
||||
|
||||
## Decisions for you
|
||||
|
||||
1. **Is the mail text right?** My pick: keep it. If you do nothing: it goes out as written.
|
||||
2. **Mail time window 09:00–20:00, at most 3 mails per kernel.** My pick: keep (decided by me, you may reverse). If you
|
||||
do nothing: it stays.
|
||||
|
||||
@@ -2,8 +2,21 @@
|
||||
|
||||
**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop) is off; nothing was sent to it.**
|
||||
|
||||
**Updated 2026-10-07 14:30: hub 0.142.0; demo-hp, demo-felhom and Tester 1 run agent 0.151.0 and controller 0.302.0.
|
||||
The open-items list is at 126. Report: `REPORT.md`.**
|
||||
**Updated 2026-10-07 18:40: hub 0.143.0; demo-hp, demo-felhom and Tester 1 run agent 0.152.0 and controller 0.302.0.
|
||||
The open-items list is at 128. Report: `REPORT.md`.**
|
||||
|
||||
## Evening (2026-10-07): the kernel lane is built
|
||||
|
||||
- **A new kernel boots once.** If it crashes, the box comes back on the old kernel by itself. If it boots well, it
|
||||
becomes the default. If it boots but the box is not healthy within 20 minutes, the box restarts once into the old
|
||||
kernel by itself. All three were proven on the Tester 1 box today.
|
||||
- **The household is mailed the day before**, between 9:00 and 20:00, in its language. No mail → no restart.
|
||||
- **You approve each kernel** on the System page after both demo boxes booted it well at night.
|
||||
- **Tonight:** demo-felhom's household was mailed, but for yesterday's kernel; the box will safely refuse it, because a
|
||||
newer one appeared. Tomorrow both demo boxes are due for the newest kernel (filed as a small item).
|
||||
- **Found:** after a restart, the first backup can fail for one app on a drive (the drive is being re-attached). Filed.
|
||||
|
||||
**Needs you (none urgent):** read the household mail text in `REPORT.md`. If nothing: it stays as written.
|
||||
|
||||
## Day (2026-10-07): your answers built, and the kernel test with real reboots
|
||||
|
||||
|
||||
@@ -6,3 +6,19 @@ uploaded signed op to the hub jobs queue
|
||||
signed: op=agent_update host=demo-felhom-8363b5 guest="" key_id=felhom-op-1 nonce=1064b6dde1dd050dd5f6f7b788990412 expires=2026-10-07T16:04:42Z
|
||||
wrote envelope to /dev/null
|
||||
uploaded signed op to the hub jobs queue
|
||||
== agent_config_update 0.152.0-step1 → demo-hp-bb76ea, 2026-10-07T15:31:38Z
|
||||
signed: op=agent_config_update host=demo-hp-bb76ea guest="" key_id=felhom-op-1 nonce=e87c668823146debf011a1680691118c expires=2026-10-07T16:16:38Z
|
||||
wrote envelope to /dev/null
|
||||
uploaded signed op to the hub jobs queue
|
||||
== agent_config_update 0.152.0-step1 → demo-felhom-8363b5, 2026-10-07T15:31:38Z
|
||||
signed: op=agent_config_update host=demo-felhom-8363b5 guest="" key_id=felhom-op-1 nonce=699531f773854f791e3517174cf382ad expires=2026-10-07T16:16:38Z
|
||||
wrote envelope to /dev/null
|
||||
uploaded signed op to the hub jobs queue
|
||||
== agent_config_update 0.152.0 → demo-hp-bb76ea, 2026-10-07T15:47:04Z
|
||||
signed: op=agent_config_update host=demo-hp-bb76ea guest="" key_id=felhom-op-1 nonce=b126b41931b7123998dcc6f3e182e1dc expires=2026-10-07T16:32:04Z
|
||||
wrote envelope to /dev/null
|
||||
uploaded signed op to the hub jobs queue
|
||||
== agent_config_update 0.152.0 → demo-felhom-8363b5, 2026-10-07T15:47:04Z
|
||||
signed: op=agent_config_update host=demo-felhom-8363b5 guest="" key_id=felhom-op-1 nonce=ac6fde774fe15cfdc4671fbbeddc502f expires=2026-10-07T16:32:04Z
|
||||
wrote envelope to /dev/null
|
||||
uploaded signed op to the hub jobs queue
|
||||
|
||||
@@ -0,0 +1,14 @@
|
||||
2026/10/07 16:17:16 [INFO] osupdates: tester-1-d70be4 reported kernel run 20261007T141703Z: ring=1 mode=apply outcome=staged healthy=true upgraded=1 pending=0 not-covered=0 restart-needed=0 reboot-needed=true wrapper=13.2s
|
||||
2026/10/07 16:20:32 [INFO] osupdates: tester-1-d70be4 reported kernel run boot-20261007T142032Z: ring=1 mode=kernel-boot outcome=fell_back healthy=false upgraded=0 pending=0 not-covered=0 restart-needed=0 reboot-needed=false wrapper=0.0s
|
||||
2026/10/07 16:20:32 [INFO] Operator email sent for tester-1/os_kernel_step
|
||||
2026/10/07 16:35:52 [INFO] osupdates: tester-1-d70be4 reported kernel run 20261007T143539Z: ring=1 mode=apply outcome=staged healthy=true upgraded=1 pending=0 not-covered=0 restart-needed=0 reboot-needed=true wrapper=12.7s
|
||||
2026/10/07 16:38:43 [INFO] osupdates: tester-1-d70be4 reported kernel run boot-20261007T143843Z: ring=1 mode=kernel-boot outcome=judging healthy=false upgraded=0 pending=0 not-covered=0 restart-needed=0 reboot-needed=false wrapper=0.0s
|
||||
2026/10/07 16:58:54 [INFO] osupdates: tester-1-d70be4 reported kernel run boot-20261007T143843Z: ring=1 mode=kernel-revert outcome=health_failed healthy=false upgraded=0 pending=0 not-covered=0 restart-needed=0 reboot-needed=false wrapper=0.0s
|
||||
2026/10/07 16:58:54 [INFO] Operator email sent for tester-1/os_kernel_step
|
||||
2026/10/07 16:59:30 [INFO] osupdates: tester-1-d70be4 reported kernel run boot-20261007T145930Z: ring=1 mode=kernel-boot outcome=self_reverted healthy=false upgraded=0 pending=0 not-covered=0 restart-needed=0 reboot-needed=false wrapper=0.0s
|
||||
2026/10/07 16:59:31 [INFO] Operator email sent for tester-1/os_kernel_step
|
||||
2026/10/07 17:15:18 [INFO] osupdates: tester-1-d70be4 reported kernel run 20261007T151439Z: ring=1 mode=apply outcome=staged healthy=true upgraded=2 pending=0 not-covered=0 restart-needed=0 reboot-needed=true wrapper=39.0s
|
||||
2026/10/07 17:17:53 [INFO] osupdates: tester-1-d70be4 reported kernel run boot-20261007T151753Z: ring=1 mode=kernel-boot outcome=judging healthy=false upgraded=0 pending=0 not-covered=0 restart-needed=0 reboot-needed=false wrapper=0.0s
|
||||
2026/10/07 17:19:10 [INFO] osupdates: tester-1-d70be4 reported kernel run boot-20261007T151753Z: ring=1 mode=kernel-good outcome=applied healthy=true upgraded=0 pending=0 not-covered=0 restart-needed=0 reboot-needed=false wrapper=0.0s
|
||||
2026/10/07 18:12:11 [INFO] kernel notice (7.0.14-20-pve) emailed to the household of demo-felhom (lang=hu)
|
||||
2026/10/07 18:12:11 [INFO] osupdates: kernel 7.0.14-20-pve — the household of demo-felhom-8363b5 was told the box restarts tonight (lang=hu)
|
||||
@@ -1,5 +1,8 @@
|
||||
# The kernel lane — what the spike measured, and a design (R-836, R-812 option B; `09` §3 decisions 164, 171)
|
||||
|
||||
> **Ruled 2026-10-07 14:43 (decision 172): A with C. BUILT the same day — `11-os-updates.md` §5.11, proofs in
|
||||
> `audits/kernel-lane-2026-10-07/`.**
|
||||
|
||||
Spike 2026-10-07, operator present. Boxes: the Tester 1 box (VM 341 on the HP box, OVMF, Secure Boot off), demo-felhom
|
||||
(Intel N100, Secure Boot off), demo-hp (AMD, Secure Boot ON). Kernels: 7.0.2-6 and 7.0.14-20. 24 reboots, 2 operator
|
||||
power cycles. Evidence per boot in this folder (`<box>/<step>/before.txt`, `after.txt`, `result.txt`; screenshots for
|
||||
|
||||
@@ -215,7 +215,8 @@ stopping line that lies.
|
||||
|
||||
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| **R-836** | Box system & updates | P2 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin <new> --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot <new>`: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** **2026-10-07 07:58: raised P3 → P2 (`09` §3 decision 164 — the kernel lane before the first paying customer); the spike runs with the operator present.** **2026-10-07 SPIKE DONE (operator present; 24 reboots, 2 power cycles; `audits/kernel-spike-2026-10-07/DESIGN-kernel-lane.md`).** On Tester 1 (VM), demo-felhom and demo-hp (Secure Boot on): **a one-shot flag in a GRUB env block on the ESP (vfat) WORKS** — new kernel once, then the old one; a panicking new kernel (`panic=10`) falls back with no person. **UEFI `BootNext` WORKS too.** **A hardware watchdog armed by systemd during the reboot FAILS on all three** (i6300ESB, Intel TCO, AMD SP5100): the reset clears the timer, so a FROZEN new kernel needs a power cycle. Pick: the ESP flag + lockup-to-panic kernel options (unmeasured). Two operator questions in the design (night restarts and the household; a freeze needs a person). Boxes left clean on a healthily booted default. **2026-10-07 14:43 — `09` §3 decision 172: the kernel lane is YES — option A (the ESP one-shot flag) with option C (lockup → panic) on the one-shot entry; a box may restart at night; the household is mailed the day before (what to do if it is not back: unplug, plug back in); each kernel set is operator-approved after ring 0; a frozen new kernel needing a power cycle is ACCEPTED for the first customers. BUILD STARTED 2026-10-07 (the kernel-lane arc).** | — | — | CC |
|
||||
| **R-836** | Box system & updates | P2 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin <new> --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot <new>`: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** **2026-10-07 07:58: raised P3 → P2 (`09` §3 decision 164 — the kernel lane before the first paying customer); the spike runs with the operator present.** **2026-10-07 SPIKE DONE (operator present; 24 reboots, 2 power cycles; `audits/kernel-spike-2026-10-07/DESIGN-kernel-lane.md`).** On Tester 1 (VM), demo-felhom and demo-hp (Secure Boot on): **a one-shot flag in a GRUB env block on the ESP (vfat) WORKS** — new kernel once, then the old one; a panicking new kernel (`panic=10`) falls back with no person. **UEFI `BootNext` WORKS too.** **A hardware watchdog armed by systemd during the reboot FAILS on all three** (i6300ESB, Intel TCO, AMD SP5100): the reset clears the timer, so a FROZEN new kernel needs a power cycle. Pick: the ESP flag + lockup-to-panic kernel options (unmeasured). Two operator questions in the design (night restarts and the household; a freeze needs a person). Boxes left clean on a healthily booted default. **2026-10-07 14:43 — `09` §3 decision 172: the kernel lane is YES — option A (the ESP one-shot flag) with option C (lockup → panic) on the one-shot entry; a box may restart at night; the household is mailed the day before (what to do if it is not back: unplug, plug back in); each kernel set is operator-approved after ring 0; a frozen new kernel needing a power cycle is ACCEPTED for the first customers. BUILD STARTED 2026-10-07 (the kernel-lane arc).** **2026-10-07 evening: BUILT and proven by hand** — agent v0.152.0 + hub v0.143.0 (`11` §5.11), delivered to demo-hp, demo-felhom and Tester 1. Tester 1 (`audits/kernel-lane-2026-10-07/E/RESULT.md`): a forced panic on the one-shot → back on the old kernel by itself (fell_back, 0 unclean boots counted); a held guest → judged 20 min → ONE self-revert → old kernel; a healthy boot → 7.0.14-22 the default. The household mail went out for demo-felhom 18:12 (Gmail). Option C UNMEASURED (no test_lockup module). **WHERE IT STOPPED (next session):** the ring-0 NIGHT run (Part E3) — tonight demo-felhom refuses the told 7.0.14-20 (R-898, safe); 7.0.14-22 is due on both demo boxes for the night 2026-10-08→09 (mail 2026-10-08 daytime); read back the morning of 2026-10-09; then approve the set on the System page and deliver to the Tester 1 box as ring 1 by a signed os_kernel_step (Part E4 — Tester 1 already runs 7.0.14-22, so the ring-1 proof needs the NEXT kernel). | — | — | CC |
|
||||
| **R-898** | Box system & updates | P3 | **Ring 0's told kernel can be a day old: the hub names the kernel from the box's LAST host report, the box's sources may offer a newer one by night, and the step then refuses (R23) — one night and one household mail lost.** SEEN 2026-10-07 18:12: demo-felhom's household was mailed for 7.0.14-20-pve (from the night report) while the box's apt already offered 7.0.14-22; the night leg would stage `pending-kernel` → 22 ≠ the told 20 → R23 before any change (safe; read from the code, `felhom-os-apply` Kernel.stage). Fix direction (needs an agent + hub release, so not in this session): ring 0 stages EXACTLY the told kernel (select `listed`, the set derived from the kver — Proxmox keeps old kernel versions), and the hub treats a told-kernel mismatch as transient, not „ended for good". | **OPEN — filed 2026-10-07 (kernel-lane arc); tonight's demo-felhom run shows it.** | — | — | CC |
|
||||
| **R-862** | Box system & updates | P3 | **Tester 2 cannot take the config bundle until one by-hand bootstrap is done: its `felhom-os-apply` (agent 0.142.0) predates the bundle mode, and no signed job can write a root file on it.** FOUND 2026-10-04 (R-840 build, Part C): the route reaches every box installed from installer 1.31.0 on, and the demo boxes (bootstrapped by CC); Tester 2 has every root file of agent 0.142.0 (installer 1.30.0) and lacks only the R-858 wrapper fix, which matters only for a Docker step it gets solely from a signed job. CC has no route to Tester 2 (its door admits only the operator's WireGuard peer; `felhom-op` cannot become root). The steps: `runbooks/config-bundle.md` "Tester 2". Then CC signs the bundle and reads it back. | **WAITING-ON-OPERATOR** — the operator said (2026-10-04 ~19:05) he will try through his tunnel **2026-10-07 07:58 (`09` §3 decision 170): the operator believes Tester 2 never set up the off-site backup (no escrow); the laptop may stay off for days; the tester plans to add an HDD.** **2026-10-07 (Part H, read only):** in Tester 2's last 10 notifications (2026-10-05 09:16 → 2026-10-07 06:58) the household got 0 mails, the operator 10 (host_down ×9, expected_dbdump_missed ×1); control: the household channel shows on the other three customers. So the household is not mailed daily while the box is off; no row (`audits/day-2026-10-07/H/`). | the operator's WireGuard tunnel | the operator runs the three bootstrap commands; CC sends the bundle | operator |
|
||||
| **R-35** | Box system & updates | P3 | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-05 (burn-down night): NEEDS A DESIGN.** Hot-apply needs a per-setting ruling; persisting sessions puts login tokens on disk (and into the whole-guest archive, `07` §5). Next: the operator's pick between them. | — | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` | CC |
|
||||
| **R-78** | Box system & updates | P3 | **`local_api` authority ruling — auto-reconcile vs detect-only** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-25, size M, roadmap state `idea (deferred OUT of R-77 on purpose)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **An OWED OPERATOR DECISION, not a defect — moved because an owed ruling hidden among feature ideas is the shape this session exists to remove.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | R-77 ships detection because the fix is genuinely undecided, and **both directions can lose customer-visible function**. **Direction 1 (today):** `controller.yaml` wins and drift is silent → the 2026-07-25 island migration blinded the whole fleet's control plane for 17.5 h (drive gate, guest-reboot recovery, quiesce/backup all degrade). R-77 makes that loud but does not stop it recurring. **Direction 2 (`bootstrap.json` wins, auto-reconcile on boot):** a guest whose `controller.yaml` is CORRECT and whose `bootstrap.json` is stale — a half-completed re-provision, a hand-repaired guest, a setup-wizard box — gets a **working channel clobbered on the next restart**, fleet-wide and silently, during a routine deploy. That is not obviously better than the bug. Needs a spike: which writer is authoritative per field (endpoint vs fingerprint vs token — `mergeLocalAPI` replaces the whole block, so they cannot be reconciled independently today), whether the agent side should stamp a generation/mtime so 'newer wins' is even expressible, and whether reconcile should require an operator ack. Until then the drift alert plus a manual edit is the supported path. | CC |
|
||||
|
||||
Reference in New Issue
Block a user