Files
felhom.eu/REPORT-os-host-lane-2026-10-04.md

91 lines
8.4 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# REPORT — OS updates, build steps 3 + 4 (2026-10-04, afternoon–evening)
Brief: "OS updates, build step 3 + 4" (Parts A–G). Architecture read first: `documentation/architecture/11-os-updates.md`
(owner), `08-alarm-ladder.md`, `03-host-agent.md`, `07-backup-architecture.md` §6.1. Evidence for every claim:
`documentation/audits/os-host-lane-2026-10-04/` (partA … partG).
## Part table
| Part | What | Result | Evidence |
|---|---|---|---|
| A | The tunnel status is true (R-841) | **DONE.** Three states from the guest container + controller v0.292.0's readiness check; `unknown` never alarms; `tunnel_down` after two `not_running` reports. Live on demo-hp: port 7844 blocked → `tunnel_down` mailed 11:52 UTC (2nd report); unblocked → `tunnel_recovered` 12:07 UTC. No new sudoers line. | `partA/` (red-proofs agent/hub/controller, `live/`) |
| B | Host fast lane (`11` §8 step 3) | **DONE.** R12 lifted for lane fast / layer host on an appliance only (root-owned install record); R14 refuses kernel/boot/firmware; host step after a healthy guest step; host health rule in `11` §8.2; reboot-needed with first date, never reboots; separate per-layer approved sets. Live: demo-felhom ring 0 → 108 Debian host packages (all Debian origin, checked against apt), healthy; demo-hp ring 0 (nothing pending); 605-package host release approved under a TEST wait (2 m / 0 nights, logged, reverted); demo-felhom as ring 1 installed exactly the 1 version it lacked; back to ring 0 and the ruled 24 h + 1 night. Host undo runbook proved on demo-hp (`tzdata`). | `partB/` (`live/`, `ring1/`, `undo/`, red-proofs) |
| C | Fleet view + four alarms (`11` §8 step 4) | **DONE.** `GET /os/fleet`: one line per box, ring, switch, tunnel, per layer release/pending/not-covered/last good leg/reboot-since. Alarms `os_update_stale`, `os_reboot_needed`, `os_ring0_stalled`, `os_not_covered`, hourly, operator-only, each red-proved (9 mutations caught). Numbers in `11` §8.3 / `09` decision 85 as CC-decided. | `partC/` |
| D | The leg is fast (R-845) | **DONE.** Nothing to install, both layers: **23.3 s** (demo-felhom), **31.5 s** (demo-hp). Before (agent 0.140.0, guest only, nothing to install): 14.0 s. A 108-package host pass: 70 s. | `partD/`, `partB/live/` |
| E | Kernel one-shot spike on demo-hp (R-836), operator's word before each reboot | **DONE (measured).** Secure Boot ON boots the new kernel fine; **`grub-reboot` is NOT a one-shot here** — `/boot` on LVM, GRUB cannot clear `next_entry`, reboot 2 (no command) came back on the NEW kernel. `kernel.panic = 0`. `sp5100_tco` loads and answers (sysfs only, never armed, unloaded). No kernel installed (7.0.14-20 was already there). Left: runs 7.0.14-20, saved default 7.0.14-20, both kernels installed, `GRUB_DEFAULT=saved`. | `partE/` |
| F | Docker slow lane design | **DONE (design only).** `11` §5.8. One STATUS decision: `live-restore` on, fleet-wide. | `11` §5.8, STATUS |
| G | Releases, golden, records | See below. Agent **v0.141.0 + v0.141.1**, hub **v0.131.0 + v0.131.1** (decision 86 — a second release in each, for a live-found defect), controller **v0.292.0**. Installer: **no new release** (see below). | `partG/` |
## Claims in the brief that turned out wrong (or only half right)
- **cloudflared readiness** — it EXISTS (`/ready`, 200 only with a connection), but nothing exposed it: the metrics
port was random and there was no health check. A container state alone lies (wrong token: `running`, `/ready` 503).
Controller v0.292.0 adds the fixed port + Docker health check; the agent reads it with its existing sudoers line.
- **"Find where the install mode is known"** — known in TWO places, and only one is trustworthy: `agent.json`
`deployment_mode` (the agent can write it) and the installer's root-owned `state.json` `mode`. The wrapper trusts
only the second (decision 84).
- **`grub-reboot` with Secure Boot** — Secure Boot was not the problem (it booted fine). The one-shot itself fails on
these boxes because GRUB cannot write its state on LVM `/boot`.
- **Panic auto-restart** — there is none: `kernel.panic = 0`; a panicked host stays down (R-851).
- **Under 60 s** — met (23–32 s). But the leg was ALREADY under 60 s with nothing to install (14 s, guest only); the
3–4-minute passes of R-845 were passes that installed something.
## Releases
- **Agent v0.141.0** (`cfba0d0`, sha256 `6eaad980…`) and **v0.141.1** (`a6bc3f1`, sha256 `b712f577…`) — signed
`agent_update` jobs to both demo boxes (both COMPLETED). v0.141.1 fixes R-846 (host "reboot needed" hid `lxc-start`,
and a reboot never cleared it).
- **Hub v0.131.0** (`88b0a2e`) and **v0.131.1** (`fd1f985`) — deployed via ArgoCD; image tag verified on the pod.
- **Controller v0.292.0** (`09e634d`) — floor raised (`min_controller_version` 0.292.0, `min_agent` 0.131.0); both demo
boxes run it; cloudflared recreated once, `healthy`.
- **Installer — no new release, deliberately** (recommendation not followed, one line why): the installer fetches
`configs/felhom-os-apply` from the VOUCHED agent's tag, the file name did not change and the sudoers line did not
change, so vouching agent 0.141.1 delivers the new wrapper to every fresh install with no installer change.
- **Copied by hand to both demo hosts** (`partG/wrapper-copied-by-hand.txt`): `/usr/local/sbin/felhom-os-apply` only —
first from v0.141.0 (sha `4729769c…`), then from v0.141.1 (sha `51e100ad…`), installed `0755 root:root`, the old
file kept as `/root/felhom-os-apply.bak-0.140.0`. `/etc/sudoers.d/felhom-agent` was NOT copied: it already equals the
repo's (`02df92d7…` on both).
## Golden
- **Golden 0.292.0 baked** (by a helper agent, RUNBOOK-manual-build §4.0/§4.1, same `build-golden.sh` bytes as
0.291.0; drill VM 9100 destroyed, drill disk back to `virgin`): sha256 `d6cf8b33…16671`, re-hashed by download in the
main session — match. Evidence `documentation/tests/golden-0.292.0-2026-10-04/` (commit `f93738d`).
- **Vouched:** agent 0.141.1 + golden 0.292.0, `min_agent` 0.131.0 → `artifacts_set` (`05-vouch.txt`). Fresh installs
now get agent 0.141.1, its wrapper, and controller 0.292.0.
- **Waiver REMOVED** (`documentation/tests/golden-waiver.yml` deleted): the newest controller (0.292.0) now has its
golden, so `golden_currency_gate.py` passes without it. Golden 0.291.0 alone could NOT have retired it — this session
released controller 0.292.0, which put 0.291.0 behind.
## Decisions taken by CC unattended (operator may reverse) — `09` §3
- **84** — the appliance proof is the root-owned install record.
- **85** — alarm numbers 7 / 14 / 7 / 14 days, all configuration.
- **86** — a second same-session release of agent and hub (R-846).
## Register
Open rows **332 → 333**. Closed: **R-841**, **R-845**; filed and closed the same day: **R-846**, **R-850**. Opened:
**R-848** (a held host package is invisible to the hub), **R-849** (the guest "reboot needed since" never clears),
**R-851** (a panicked host stays down — operator). Narrowed: **R-836** (GRUB one-shot measured: not a one-shot),
**R-812** (host lane built). `unproven.py`: unchanged (35 of 55 not walked).
## Teardown — three layers
- **Machine:** demo-hp guest 9201 — the port-7844 block removed (`DOCKER-USER` empty, verified); cloudflared
`healthy`. demo-hp host — `tzdata` back on 2026c, no holds, no snapshot source left; `sp5100_tco` unloaded; GRUB:
`GRUB_DEFAULT=saved`, saved default 7.0.14-20, `next_entry` cleared, `/etc/default/grub` backup at
`/root/grub.default.bak-2026-10-04`. demo-felhom host — `tzdata` 2026c (re-installed by its ring-1 run), no snapshot
source left.
- **Host:** both demo hosts run agent 0.141.1 + wrapper `51e100ad…`; both ring 0, switch ON.
- **Hub:** the TEST approval wait reverted (log: 24 h and 1 night); the TEST releases `os-guest-20261004-123933` and
`os-host-20261004-124034` stay (they are real approvals of what ring 0 runs). The operator mails `tunnel_down` /
`tunnel_recovered` for demo-hp were sent on purpose (Part A). The hub password copy in the scratchpad was shredded.
## CI
Checked by `head_sha` over every page of the Gitea `jobs` endpoint (felhom.eu `CLAUDE.md` recipe):
agent `cfba0d0` (jobs 1259, 1260), `3bf77c3` (1261), `a6bc3f1` (1262, 1263), `2e2e8f5` (1264) — all **success**;
controller `09e634d` (1258) **success**; felhom.eu `88b0a2e` (1256), `fd1f985` (1265) **success**. The last docs pushes
(`31bdb4b` and this report's commit): see the session's final message — checked after the push.