Files
felhom.eu/REPORT.md
T

3.8 KiB
Raw Blame History

REPORT — the operator's answers built; the kernel lane spiked with real reboots (2026-10-07 day)

Part Result
A — rulings 09 §3 decisions 162–170 recorded first, plus 171 (reboots without a per-reboot word, never DooPlex). R-822 closed by ruling. R-836 → P2. Shared rule file: no hub build/deploy without the operator (all five copies identical)
B — Proxmox package lane (R-812 A) built (agent pve layer + /etc/pve write gate; hub candidate/approval/System page), delivered, proven live on demo-felhom: one signed step, 65 packages in 70 s, pve-manager 9.2.2 → 9.2.21, healthy, guest untouched, app 31/31; hub logged it. R-812 closed
C — agent image check (R-861 (a) A1, (b) B2) built, delivered to the three boxes; on demo-hp: no tee grant, the verb is the route, an alpine ref refused (rc 3, file unchanged), felhom-op's pct lines exact. Not seen: a full managed swap (no newer controller today)
D — RESET purge + clean-up (R-32) purge through the sub-account's own login built and delivered (hub 0.142.0). The one-time clean-up STOPPED: ~2.7 GB belongs to no live customer, reachable only by the pool box's main account — no such login on DooPlex
E — other-key archives line (R-366 slice 2) built and delivered; R-366 closed
F — retire the empty fields (R-105) built and delivered (columns kept, unread); 05/06 corrected; R-105 closed
G — kernel spike (R-836) 24 reboots + 2 power cycles on Tester 1 (VM), demo-felhom, demo-hp (Secure Boot on). ESP one-shot flag: works on all three. UEFI BootNext: works on all three. A watchdog armed by systemd: fails on all three (a reset clears the timer; a frozen kernel needs a person). Design with two operator questions: audits/kernel-spike-2026-10-07/DESIGN-kernel-lane.md. Boxes left clean on a healthily booted default
H — Tester 2 (read only) the household got 0 mails while the box was off (operator 10); no row
Rows before Rows after Opened Closed
130 126 0 4 (R-822, R-812, R-366, R-105)

Releases (one per repo changed): hub 0.142.0 (live, healthz 200 within 50 s); agent 0.151.0 + bundle on demo-hp, demo-felhom, Tester 1 (probe 68/68 each). The agent vouch was refused by the hub (golden_behind_fleet: golden 0.301.0 is behind the fleet's controller 0.302.0), so new installs keep agent 0.150.0 until the weekly bake. No controller release (no controller code changed). CI: one lost job (felhom.eu run 1497, every step failed with no log) re-run once → success (R-887 count +1).

Helpers: each prompt carried the brief's fences in full (rule 11). Security note on hub/internal/api/handler.go (raised by a background review, no detail): the last three changes log box/customer ids and counts only — nothing fixed.

Teardown: Tester 1 — spike entries, ESP flag, BootNext loader/entry and watchdog config removed; VM 341's test watchdog device removed; default 7.0.2-6 (running). demo-felhom — the same; default 7.0.2-6; Proxmox userspace now 9.2.21. demo-hp — the same; default 7.0.14-20. 7.0.14-20 is installed on Tester 1 and demo-felhom (booted healthily in the spike) but not their default. Hub — nothing provisioned.

Two decisions for you

  1. The kernel lane's promise: may a box restart at night for a kernel update, and do we tell the household? My pick: yes, with one line in the household's mail the day before. If you do nothing: kernels stay manual.
  2. The pool box clean-up (~2.7 GB): give me the pool box's main-account login for one session, or delete the folders yourself in the Hetzner console. My pick: you delete them in the console (no new credential on DooPlex). If you do nothing: the leftovers stay; nothing reads them.