Files
felhom.eu/REPORT-os-host-lane-2026-10-04.md
T

8.4 KiB
Raw Blame History

REPORT — OS updates, build steps 3 + 4 (2026-10-04, afternoon–evening)

Brief: "OS updates, build step 3 + 4" (Parts A–G). Architecture read first: documentation/architecture/11-os-updates.md (owner), 08-alarm-ladder.md, 03-host-agent.md, 07-backup-architecture.md §6.1. Evidence for every claim: documentation/audits/os-host-lane-2026-10-04/ (partA … partG).

Part table

Part What Result Evidence
A The tunnel status is true (R-841) DONE. Three states from the guest container + controller v0.292.0's readiness check; unknown never alarms; tunnel_down after two not_running reports. Live on demo-hp: port 7844 blocked → tunnel_down mailed 11:52 UTC (2nd report); unblocked → tunnel_recovered 12:07 UTC. No new sudoers line. partA/ (red-proofs agent/hub/controller, live/)
B Host fast lane (11 §8 step 3) DONE. R12 lifted for lane fast / layer host on an appliance only (root-owned install record); R14 refuses kernel/boot/firmware; host step after a healthy guest step; host health rule in 11 §8.2; reboot-needed with first date, never reboots; separate per-layer approved sets. Live: demo-felhom ring 0 → 108 Debian host packages (all Debian origin, checked against apt), healthy; demo-hp ring 0 (nothing pending); 605-package host release approved under a TEST wait (2 m / 0 nights, logged, reverted); demo-felhom as ring 1 installed exactly the 1 version it lacked; back to ring 0 and the ruled 24 h + 1 night. Host undo runbook proved on demo-hp (tzdata). partB/ (live/, ring1/, undo/, red-proofs)
C Fleet view + four alarms (11 §8 step 4) DONE. GET /os/fleet: one line per box, ring, switch, tunnel, per layer release/pending/not-covered/last good leg/reboot-since. Alarms os_update_stale, os_reboot_needed, os_ring0_stalled, os_not_covered, hourly, operator-only, each red-proved (9 mutations caught). Numbers in 11 §8.3 / 09 decision 85 as CC-decided. partC/
D The leg is fast (R-845) DONE. Nothing to install, both layers: 23.3 s (demo-felhom), 31.5 s (demo-hp). Before (agent 0.140.0, guest only, nothing to install): 14.0 s. A 108-package host pass: 70 s. partD/, partB/live/
E Kernel one-shot spike on demo-hp (R-836), operator's word before each reboot DONE (measured). Secure Boot ON boots the new kernel fine; grub-reboot is NOT a one-shot here — /boot on LVM, GRUB cannot clear next_entry, reboot 2 (no command) came back on the NEW kernel. kernel.panic = 0. sp5100_tco loads and answers (sysfs only, never armed, unloaded). No kernel installed (7.0.14-20 was already there). Left: runs 7.0.14-20, saved default 7.0.14-20, both kernels installed, GRUB_DEFAULT=saved. partE/
F Docker slow lane design DONE (design only). 11 §5.8. One STATUS decision: live-restore on, fleet-wide. 11 §5.8, STATUS
G Releases, golden, records See below. Agent v0.141.0 + v0.141.1, hub v0.131.0 + v0.131.1 (decision 86 — a second release in each, for a live-found defect), controller v0.292.0. Installer: no new release (see below). partG/

Claims in the brief that turned out wrong (or only half right)

  • cloudflared readiness — it EXISTS (/ready, 200 only with a connection), but nothing exposed it: the metrics port was random and there was no health check. A container state alone lies (wrong token: running, /ready 503). Controller v0.292.0 adds the fixed port + Docker health check; the agent reads it with its existing sudoers line.
  • "Find where the install mode is known" — known in TWO places, and only one is trustworthy: agent.json deployment_mode (the agent can write it) and the installer's root-owned state.json mode. The wrapper trusts only the second (decision 84).
  • grub-reboot with Secure Boot — Secure Boot was not the problem (it booted fine). The one-shot itself fails on these boxes because GRUB cannot write its state on LVM /boot.
  • Panic auto-restart — there is none: kernel.panic = 0; a panicked host stays down (R-851).
  • Under 60 s — met (23–32 s). But the leg was ALREADY under 60 s with nothing to install (14 s, guest only); the 3–4-minute passes of R-845 were passes that installed something.

Releases

  • Agent v0.141.0 (cfba0d0, sha256 6eaad980…) and v0.141.1 (a6bc3f1, sha256 b712f577…) — signed agent_update jobs to both demo boxes (both COMPLETED). v0.141.1 fixes R-846 (host "reboot needed" hid lxc-start, and a reboot never cleared it).
  • Hub v0.131.0 (88b0a2e) and v0.131.1 (fd1f985) — deployed via ArgoCD; image tag verified on the pod.
  • Controller v0.292.0 (09e634d) — floor raised (min_controller_version 0.292.0, min_agent 0.131.0); both demo boxes run it; cloudflared recreated once, healthy.
  • Installer — no new release, deliberately (recommendation not followed, one line why): the installer fetches configs/felhom-os-apply from the VOUCHED agent's tag, the file name did not change and the sudoers line did not change, so vouching agent 0.141.1 delivers the new wrapper to every fresh install with no installer change.
  • Copied by hand to both demo hosts (partG/wrapper-copied-by-hand.txt): /usr/local/sbin/felhom-os-apply only — first from v0.141.0 (sha 4729769c…), then from v0.141.1 (sha 51e100ad…), installed 0755 root:root, the old file kept as /root/felhom-os-apply.bak-0.140.0. /etc/sudoers.d/felhom-agent was NOT copied: it already equals the repo's (02df92d7… on both).

Golden

  • Golden 0.292.0 baked (by a helper agent, RUNBOOK-manual-build §4.0/§4.1, same build-golden.sh bytes as 0.291.0; drill VM 9100 destroyed, drill disk back to virgin): sha256 d6cf8b33…16671, re-hashed by download in the main session — match. Evidence documentation/tests/golden-0.292.0-2026-10-04/ (commit f93738d).
  • Vouched: agent 0.141.1 + golden 0.292.0, min_agent 0.131.0 → artifacts_set (05-vouch.txt). Fresh installs now get agent 0.141.1, its wrapper, and controller 0.292.0.
  • Waiver REMOVED (documentation/tests/golden-waiver.yml deleted): the newest controller (0.292.0) now has its golden, so golden_currency_gate.py passes without it. Golden 0.291.0 alone could NOT have retired it — this session released controller 0.292.0, which put 0.291.0 behind.

Decisions taken by CC unattended (operator may reverse) — 09 §3

  • 84 — the appliance proof is the root-owned install record.
  • 85 — alarm numbers 7 / 14 / 7 / 14 days, all configuration.
  • 86 — a second same-session release of agent and hub (R-846).

Register

Open rows 332 → 333. Closed: R-841, R-845; filed and closed the same day: R-846, R-850. Opened: R-848 (a held host package is invisible to the hub), R-849 (the guest "reboot needed since" never clears), R-851 (a panicked host stays down — operator). Narrowed: R-836 (GRUB one-shot measured: not a one-shot), R-812 (host lane built). unproven.py: unchanged (35 of 55 not walked).

Teardown — three layers

  • Machine: demo-hp guest 9201 — the port-7844 block removed (DOCKER-USER empty, verified); cloudflared healthy. demo-hp host — tzdata back on 2026c, no holds, no snapshot source left; sp5100_tco unloaded; GRUB: GRUB_DEFAULT=saved, saved default 7.0.14-20, next_entry cleared, /etc/default/grub backup at /root/grub.default.bak-2026-10-04. demo-felhom host — tzdata 2026c (re-installed by its ring-1 run), no snapshot source left.
  • Host: both demo hosts run agent 0.141.1 + wrapper 51e100ad…; both ring 0, switch ON.
  • Hub: the TEST approval wait reverted (log: 24 h and 1 night); the TEST releases os-guest-20261004-123933 and os-host-20261004-124034 stay (they are real approvals of what ring 0 runs). The operator mails tunnel_down / tunnel_recovered for demo-hp were sent on purpose (Part A). The hub password copy in the scratchpad was shredded.

CI

Checked by head_sha over every page of the Gitea jobs endpoint (felhom.eu CLAUDE.md recipe): agent cfba0d0 (jobs 1259, 1260), 3bf77c3 (1261), a6bc3f1 (1262, 1263), 2e2e8f5 (1264) — all success; controller 09e634d (1258) success; felhom.eu 88b0a2e (1256), fd1f985 (1265) success. The last docs pushes (31bdb4b and this report's commit): see the session's final message — checked after the push.