golden 0.292.0 vouched with agent 0.141.1; golden waiver deleted (newest controller has its golden); REPORT-os-host-lane-2026-10-04 with the Part table; STATUS/CONTEXT
gates / gates (push) Successful in 33s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-04 14:53:50 +02:00
parent f93738dc49
commit 5efe6daec3
5 changed files with 101 additions and 18 deletions
+2 -1
View File
@@ -24,7 +24,8 @@
> demo-hp; ring 0 host pass on demo-felhom (108 Debian packages); a 605-package host release; ring 1 exact install;
> by-hand host undo proved; leg 23–32 s with nothing to install. Kernel spike (R-836, narrowed): GRUB's one-shot is not
> a one-shot on LVM `/boot`; Secure Boot fine; `kernel.panic = 0` (R-851); `sp5100_tco` answers. demo-hp now runs and
> defaults to kernel 7.0.14-20. Docker slow lane designed (`11` §5.8). `REPORT-os-host-lane-2026-10-04.md`.
> defaults to kernel 7.0.14-20. Docker slow lane designed (`11` §5.8). Golden 0.292.0 baked + vouched with agent 0.141.1 (min_agent
> 0.131.0); the golden waiver deleted (not needed). `REPORT-os-host-lane-2026-10-04.md`.
> **Rulings 2026-10-04 (~12:20) — recorded before the work (host fast lane brief).** `09` §3 decisions **81** (R-842 A:
> the undo is the whole-guest backup by hand; R-842 closed), **82** (R-840 not built now — "There are no older boxes";
+90
View File
@@ -0,0 +1,90 @@
# REPORT — OS updates, build steps 3 + 4 (2026-10-04, afternoon–evening)
Brief: "OS updates, build step 3 + 4" (Parts A–G). Architecture read first: `documentation/architecture/11-os-updates.md`
(owner), `08-alarm-ladder.md`, `03-host-agent.md`, `07-backup-architecture.md` §6.1. Evidence for every claim:
`documentation/audits/os-host-lane-2026-10-04/` (partA … partG).
## Part table
| Part | What | Result | Evidence |
|---|---|---|---|
| A | The tunnel status is true (R-841) | **DONE.** Three states from the guest container + controller v0.292.0's readiness check; `unknown` never alarms; `tunnel_down` after two `not_running` reports. Live on demo-hp: port 7844 blocked → `tunnel_down` mailed 11:52 UTC (2nd report); unblocked → `tunnel_recovered` 12:07 UTC. No new sudoers line. | `partA/` (red-proofs agent/hub/controller, `live/`) |
| B | Host fast lane (`11` §8 step 3) | **DONE.** R12 lifted for lane fast / layer host on an appliance only (root-owned install record); R14 refuses kernel/boot/firmware; host step after a healthy guest step; host health rule in `11` §8.2; reboot-needed with first date, never reboots; separate per-layer approved sets. Live: demo-felhom ring 0 → 108 Debian host packages (all Debian origin, checked against apt), healthy; demo-hp ring 0 (nothing pending); 605-package host release approved under a TEST wait (2 m / 0 nights, logged, reverted); demo-felhom as ring 1 installed exactly the 1 version it lacked; back to ring 0 and the ruled 24 h + 1 night. Host undo runbook proved on demo-hp (`tzdata`). | `partB/` (`live/`, `ring1/`, `undo/`, red-proofs) |
| C | Fleet view + four alarms (`11` §8 step 4) | **DONE.** `GET /os/fleet`: one line per box, ring, switch, tunnel, per layer release/pending/not-covered/last good leg/reboot-since. Alarms `os_update_stale`, `os_reboot_needed`, `os_ring0_stalled`, `os_not_covered`, hourly, operator-only, each red-proved (9 mutations caught). Numbers in `11` §8.3 / `09` decision 85 as CC-decided. | `partC/` |
| D | The leg is fast (R-845) | **DONE.** Nothing to install, both layers: **23.3 s** (demo-felhom), **31.5 s** (demo-hp). Before (agent 0.140.0, guest only, nothing to install): 14.0 s. A 108-package host pass: 70 s. | `partD/`, `partB/live/` |
| E | Kernel one-shot spike on demo-hp (R-836), operator's word before each reboot | **DONE (measured).** Secure Boot ON boots the new kernel fine; **`grub-reboot` is NOT a one-shot here** — `/boot` on LVM, GRUB cannot clear `next_entry`, reboot 2 (no command) came back on the NEW kernel. `kernel.panic = 0`. `sp5100_tco` loads and answers (sysfs only, never armed, unloaded). No kernel installed (7.0.14-20 was already there). Left: runs 7.0.14-20, saved default 7.0.14-20, both kernels installed, `GRUB_DEFAULT=saved`. | `partE/` |
| F | Docker slow lane design | **DONE (design only).** `11` §5.8. One STATUS decision: `live-restore` on, fleet-wide. | `11` §5.8, STATUS |
| G | Releases, golden, records | See below. Agent **v0.141.0 + v0.141.1**, hub **v0.131.0 + v0.131.1** (decision 86 — a second release in each, for a live-found defect), controller **v0.292.0**. Installer: **no new release** (see below). | `partG/` |
## Claims in the brief that turned out wrong (or only half right)
- **cloudflared readiness** — it EXISTS (`/ready`, 200 only with a connection), but nothing exposed it: the metrics
port was random and there was no health check. A container state alone lies (wrong token: `running`, `/ready` 503).
Controller v0.292.0 adds the fixed port + Docker health check; the agent reads it with its existing sudoers line.
- **"Find where the install mode is known"** — known in TWO places, and only one is trustworthy: `agent.json`
`deployment_mode` (the agent can write it) and the installer's root-owned `state.json` `mode`. The wrapper trusts
only the second (decision 84).
- **`grub-reboot` with Secure Boot** — Secure Boot was not the problem (it booted fine). The one-shot itself fails on
these boxes because GRUB cannot write its state on LVM `/boot`.
- **Panic auto-restart** — there is none: `kernel.panic = 0`; a panicked host stays down (R-851).
- **Under 60 s** — met (23–32 s). But the leg was ALREADY under 60 s with nothing to install (14 s, guest only); the
3–4-minute passes of R-845 were passes that installed something.
## Releases
- **Agent v0.141.0** (`cfba0d0`, sha256 `6eaad980…`) and **v0.141.1** (`a6bc3f1`, sha256 `b712f577…`) — signed
`agent_update` jobs to both demo boxes (both COMPLETED). v0.141.1 fixes R-846 (host "reboot needed" hid `lxc-start`,
and a reboot never cleared it).
- **Hub v0.131.0** (`88b0a2e`) and **v0.131.1** (`fd1f985`) — deployed via ArgoCD; image tag verified on the pod.
- **Controller v0.292.0** (`09e634d`) — floor raised (`min_controller_version` 0.292.0, `min_agent` 0.131.0); both demo
boxes run it; cloudflared recreated once, `healthy`.
- **Installer — no new release, deliberately** (recommendation not followed, one line why): the installer fetches
`configs/felhom-os-apply` from the VOUCHED agent's tag, the file name did not change and the sudoers line did not
change, so vouching agent 0.141.1 delivers the new wrapper to every fresh install with no installer change.
- **Copied by hand to both demo hosts** (`partG/wrapper-copied-by-hand.txt`): `/usr/local/sbin/felhom-os-apply` only —
first from v0.141.0 (sha `4729769c…`), then from v0.141.1 (sha `51e100ad…`), installed `0755 root:root`, the old
file kept as `/root/felhom-os-apply.bak-0.140.0`. `/etc/sudoers.d/felhom-agent` was NOT copied: it already equals the
repo's (`02df92d7…` on both).
## Golden
- **Golden 0.292.0 baked** (by a helper agent, RUNBOOK-manual-build §4.0/§4.1, same `build-golden.sh` bytes as
0.291.0; drill VM 9100 destroyed, drill disk back to `virgin`): sha256 `d6cf8b33…16671`, re-hashed by download in the
main session — match. Evidence `documentation/tests/golden-0.292.0-2026-10-04/` (commit `f93738d`).
- **Vouched:** agent 0.141.1 + golden 0.292.0, `min_agent` 0.131.0 → `artifacts_set` (`05-vouch.txt`). Fresh installs
now get agent 0.141.1, its wrapper, and controller 0.292.0.
- **Waiver REMOVED** (`documentation/tests/golden-waiver.yml` deleted): the newest controller (0.292.0) now has its
golden, so `golden_currency_gate.py` passes without it. Golden 0.291.0 alone could NOT have retired it — this session
released controller 0.292.0, which put 0.291.0 behind.
## Decisions taken by CC unattended (operator may reverse) — `09` §3
- **84** — the appliance proof is the root-owned install record.
- **85** — alarm numbers 7 / 14 / 7 / 14 days, all configuration.
- **86** — a second same-session release of agent and hub (R-846).
## Register
Open rows **332 → 333**. Closed: **R-841**, **R-845**; filed and closed the same day: **R-846**, **R-850**. Opened:
**R-848** (a held host package is invisible to the hub), **R-849** (the guest "reboot needed since" never clears),
**R-851** (a panicked host stays down — operator). Narrowed: **R-836** (GRUB one-shot measured: not a one-shot),
**R-812** (host lane built). `unproven.py`: unchanged (35 of 55 not walked).
## Teardown — three layers
- **Machine:** demo-hp guest 9201 — the port-7844 block removed (`DOCKER-USER` empty, verified); cloudflared
`healthy`. demo-hp host — `tzdata` back on 2026c, no holds, no snapshot source left; `sp5100_tco` unloaded; GRUB:
`GRUB_DEFAULT=saved`, saved default 7.0.14-20, `next_entry` cleared, `/etc/default/grub` backup at
`/root/grub.default.bak-2026-10-04`. demo-felhom host — `tzdata` 2026c (re-installed by its ring-1 run), no snapshot
source left.
- **Host:** both demo hosts run agent 0.141.1 + wrapper `51e100ad…`; both ring 0, switch ON.
- **Hub:** the TEST approval wait reverted (log: 24 h and 1 night); the TEST releases `os-guest-20261004-123933` and
`os-host-20261004-124034` stay (they are real approvals of what ring 0 runs). The operator mails `tunnel_down` /
`tunnel_recovered` for demo-hp were sent on purpose (Part A). The hub password copy in the scratchpad was shredded.
## CI
Checked by `head_sha` over every page of the Gitea `jobs` endpoint (felhom.eu `CLAUDE.md` recipe):
agent `cfba0d0` (jobs 1259, 1260), `3bf77c3` (1261), `a6bc3f1` (1262, 1263), `2e2e8f5` (1264) — all **success**;
controller `09e634d` (1258) **success**; felhom.eu `88b0a2e` (1256), `fd1f985` (1265) **success**. The last docs pushes
(`31bdb4b` and this report's commit): see the session's final message — checked after the push.
+2 -1
View File
@@ -4,7 +4,7 @@
**Updated 2026-10-04 (evening): the HOST's security fixes install themselves too, the hub has a fleet view and four
OS alarms, and the tunnel status is true. Both demo boxes run controller 0.292.0 and host agent 0.141.1. Hub 0.131.1.
New installs: see the golden line in the section below.**
New installs get golden 0.292.0 with agent 0.141.1 (vouched).**
## Today (2026-10-04, evening): host fixes, the fleet view, a true tunnel status
@@ -37,6 +37,7 @@ New installs: see the golden line in the section below.**
on our boxes: the second plain reboot came up on the NEW kernel again. So a new kernel that hangs would stay. That
must be solved before kernel updates. demo-hp now runs the newer kernel, healthy. A watchdog chip exists on demo-hp.
- **Found and fixed during the live test:** "reboot needed" was wrong in two ways (fixed in the second release).
- **New-install image baked and vouched** (0.292.0 with host agent 0.141.1). The golden waiver is gone: not needed.
- **Rows:** 2 closed, 2 opened-and-closed the same day, 3 opened. The list went from 332 to 333.
**Needs you later (nothing breaks if you wait):**
@@ -0,0 +1,7 @@
Vouch 2026-10-04 by CC (main session), POST /configuration/artifacts via operator Basic auth on the ClusterIP:
agent_version=0.141.1 agent_sha256=b712f577099fd2d374f648df1e825874821302fe1b46e93c7b4b70044fbe84b5
golden_version=0.292.0 golden_sha256=d6cf8b33ad58e5cfbf9566558044b2e1df692ef29d5353a40e39117794d16671 (re-hashed by download in the main session: match)
min_agent=0.131.0
-> 303 /configuration?flash=artifacts_set
hub log: Artifact manifest set: agent=0.141.1 golden=0.292.0 min_agent="0.131.0" wrapper_sha=false (same as the 0.287.0 and 0.291.0 vouches: no PBS wrapper pin was set before)
Controller floor: already 0.292.0 (min_agent 0.131.0), set earlier this session.
-16
View File
@@ -1,16 +0,0 @@
# documentation/tests/golden-waiver.yml — read by scripts/golden_currency_gate.py
#
# Goldens are on a CADENCE, not per release (operator ruling 2026-09-13, R-468): weekly, and always
# before any drill or fresh install. While this waiver is valid a golden BEHIND the record makes the
# gate ADVISORY instead of red. It never covers a golden that is UNRECORDED (R-385). At most 14 days
# after `issued`, or the gate refuses it. Renew with a new dated commit; delete it at the first
# external install. Issued AFTER the 0.236.0 bake landed — not over a bake that could have been done.
# Renewed 2026-09-27 by CC for 7 days only: the version-travel brief said "no golden" this session, and the
# weekly bake (last: 0.258.0, 2026-09-20) is DUE — STATUS asks the operator for it. Not a bypass: dated, capped.
# Renewed 2026-10-04 by CC for 3 days only: the guest-OS-lane brief releases controller 0.291.0 and agent 0.140.0
# and ENDS with a golden bake + vouch (its Part H); this covers the hours between the release and that bake. Not a
# bypass: dated, capped, and the bake is in the same session.
issued: 2026-10-04
expires: 2026-10-07
reason: pre-customer development; goldens on a weekly cadence (operator ruling 2026-09-13)
register_row: R-468