Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
8.4 KiB
REPORT — OS updates, build steps 3 + 4 (2026-10-04, afternoon–evening)
Brief: "OS updates, build step 3 + 4" (Parts A–G). Architecture read first: documentation/architecture/11-os-updates.md
(owner), 08-alarm-ladder.md, 03-host-agent.md, 07-backup-architecture.md §6.1. Evidence for every claim:
documentation/audits/os-host-lane-2026-10-04/ (partA … partG).
Part table
| Part | What | Result | Evidence |
|---|---|---|---|
| A | The tunnel status is true (R-841) | DONE. Three states from the guest container + controller v0.292.0's readiness check; unknown never alarms; tunnel_down after two not_running reports. Live on demo-hp: port 7844 blocked → tunnel_down mailed 11:52 UTC (2nd report); unblocked → tunnel_recovered 12:07 UTC. No new sudoers line. |
partA/ (red-proofs agent/hub/controller, live/) |
| B | Host fast lane (11 §8 step 3) |
DONE. R12 lifted for lane fast / layer host on an appliance only (root-owned install record); R14 refuses kernel/boot/firmware; host step after a healthy guest step; host health rule in 11 §8.2; reboot-needed with first date, never reboots; separate per-layer approved sets. Live: demo-felhom ring 0 → 108 Debian host packages (all Debian origin, checked against apt), healthy; demo-hp ring 0 (nothing pending); 605-package host release approved under a TEST wait (2 m / 0 nights, logged, reverted); demo-felhom as ring 1 installed exactly the 1 version it lacked; back to ring 0 and the ruled 24 h + 1 night. Host undo runbook proved on demo-hp (tzdata). |
partB/ (live/, ring1/, undo/, red-proofs) |
| C | Fleet view + four alarms (11 §8 step 4) |
DONE. GET /os/fleet: one line per box, ring, switch, tunnel, per layer release/pending/not-covered/last good leg/reboot-since. Alarms os_update_stale, os_reboot_needed, os_ring0_stalled, os_not_covered, hourly, operator-only, each red-proved (9 mutations caught). Numbers in 11 §8.3 / 09 decision 85 as CC-decided. |
partC/ |
| D | The leg is fast (R-845) | DONE. Nothing to install, both layers: 23.3 s (demo-felhom), 31.5 s (demo-hp). Before (agent 0.140.0, guest only, nothing to install): 14.0 s. A 108-package host pass: 70 s. | partD/, partB/live/ |
| E | Kernel one-shot spike on demo-hp (R-836), operator's word before each reboot | DONE (measured). Secure Boot ON boots the new kernel fine; grub-reboot is NOT a one-shot here — /boot on LVM, GRUB cannot clear next_entry, reboot 2 (no command) came back on the NEW kernel. kernel.panic = 0. sp5100_tco loads and answers (sysfs only, never armed, unloaded). No kernel installed (7.0.14-20 was already there). Left: runs 7.0.14-20, saved default 7.0.14-20, both kernels installed, GRUB_DEFAULT=saved. |
partE/ |
| F | Docker slow lane design | DONE (design only). 11 §5.8. One STATUS decision: live-restore on, fleet-wide. |
11 §5.8, STATUS |
| G | Releases, golden, records | See below. Agent v0.141.0 + v0.141.1, hub v0.131.0 + v0.131.1 (decision 86 — a second release in each, for a live-found defect), controller v0.292.0. Installer: no new release (see below). | partG/ |
Claims in the brief that turned out wrong (or only half right)
- cloudflared readiness — it EXISTS (
/ready, 200 only with a connection), but nothing exposed it: the metrics port was random and there was no health check. A container state alone lies (wrong token:running,/ready503). Controller v0.292.0 adds the fixed port + Docker health check; the agent reads it with its existing sudoers line. - "Find where the install mode is known" — known in TWO places, and only one is trustworthy:
agent.jsondeployment_mode(the agent can write it) and the installer's root-ownedstate.jsonmode. The wrapper trusts only the second (decision 84). grub-rebootwith Secure Boot — Secure Boot was not the problem (it booted fine). The one-shot itself fails on these boxes because GRUB cannot write its state on LVM/boot.- Panic auto-restart — there is none:
kernel.panic = 0; a panicked host stays down (R-851). - Under 60 s — met (23–32 s). But the leg was ALREADY under 60 s with nothing to install (14 s, guest only); the 3–4-minute passes of R-845 were passes that installed something.
Releases
- Agent v0.141.0 (
cfba0d0, sha2566eaad980…) and v0.141.1 (a6bc3f1, sha256b712f577…) — signedagent_updatejobs to both demo boxes (both COMPLETED). v0.141.1 fixes R-846 (host "reboot needed" hidlxc-start, and a reboot never cleared it). - Hub v0.131.0 (
88b0a2e) and v0.131.1 (fd1f985) — deployed via ArgoCD; image tag verified on the pod. - Controller v0.292.0 (
09e634d) — floor raised (min_controller_version0.292.0,min_agent0.131.0); both demo boxes run it; cloudflared recreated once,healthy. - Installer — no new release, deliberately (recommendation not followed, one line why): the installer fetches
configs/felhom-os-applyfrom the VOUCHED agent's tag, the file name did not change and the sudoers line did not change, so vouching agent 0.141.1 delivers the new wrapper to every fresh install with no installer change. - Copied by hand to both demo hosts (
partG/wrapper-copied-by-hand.txt):/usr/local/sbin/felhom-os-applyonly — first from v0.141.0 (sha4729769c…), then from v0.141.1 (sha51e100ad…), installed0755 root:root, the old file kept as/root/felhom-os-apply.bak-0.140.0./etc/sudoers.d/felhom-agentwas NOT copied: it already equals the repo's (02df92d7…on both).
Golden
- Golden 0.292.0 baked (by a helper agent, RUNBOOK-manual-build §4.0/§4.1, same
build-golden.shbytes as 0.291.0; drill VM 9100 destroyed, drill disk back tovirgin): sha256d6cf8b33…16671, re-hashed by download in the main session — match. Evidencedocumentation/tests/golden-0.292.0-2026-10-04/(commitf93738d). - Vouched: agent 0.141.1 + golden 0.292.0,
min_agent0.131.0 →artifacts_set(05-vouch.txt). Fresh installs now get agent 0.141.1, its wrapper, and controller 0.292.0. - Waiver REMOVED (
documentation/tests/golden-waiver.ymldeleted): the newest controller (0.292.0) now has its golden, sogolden_currency_gate.pypasses without it. Golden 0.291.0 alone could NOT have retired it — this session released controller 0.292.0, which put 0.291.0 behind.
Decisions taken by CC unattended (operator may reverse) — 09 §3
- 84 — the appliance proof is the root-owned install record.
- 85 — alarm numbers 7 / 14 / 7 / 14 days, all configuration.
- 86 — a second same-session release of agent and hub (R-846).
Register
Open rows 332 → 333. Closed: R-841, R-845; filed and closed the same day: R-846, R-850. Opened:
R-848 (a held host package is invisible to the hub), R-849 (the guest "reboot needed since" never clears),
R-851 (a panicked host stays down — operator). Narrowed: R-836 (GRUB one-shot measured: not a one-shot),
R-812 (host lane built). unproven.py: unchanged (35 of 55 not walked).
Teardown — three layers
- Machine: demo-hp guest 9201 — the port-7844 block removed (
DOCKER-USERempty, verified); cloudflaredhealthy. demo-hp host —tzdataback on 2026c, no holds, no snapshot source left;sp5100_tcounloaded; GRUB:GRUB_DEFAULT=saved, saved default 7.0.14-20,next_entrycleared,/etc/default/grubbackup at/root/grub.default.bak-2026-10-04. demo-felhom host —tzdata2026c (re-installed by its ring-1 run), no snapshot source left. - Host: both demo hosts run agent 0.141.1 + wrapper
51e100ad…; both ring 0, switch ON. - Hub: the TEST approval wait reverted (log: 24 h and 1 night); the TEST releases
os-guest-20261004-123933andos-host-20261004-124034stay (they are real approvals of what ring 0 runs). The operator mailstunnel_down/tunnel_recoveredfor demo-hp were sent on purpose (Part A). The hub password copy in the scratchpad was shredded.
CI
Checked by head_sha over every page of the Gitea jobs endpoint (felhom.eu CLAUDE.md recipe):
agent cfba0d0 (jobs 1259, 1260), 3bf77c3 (1261), a6bc3f1 (1262, 1263), 2e2e8f5 (1264) — all success;
controller 09e634d (1258) success; felhom.eu 88b0a2e (1256), fd1f985 (1265) success. The last docs pushes
(31bdb4b and this report's commit): see the session's final message — checked after the push.