Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
7.9 KiB
REPORT — System page, Docker slow lane, crash restart (2026-10-04, late afternoon–evening)
Brief: "OS updates — a System page …; the Docker engine slow lane built (live-restore ON, decision); a crashed host
restarts by itself, with a limit (decision)". Architecture read first: documentation/architecture/11-os-updates.md
(owner; §5.6, §5.8, §8.1–8.3), 05-hub-architecture.md, 03-host-agent.md, 08-alarm-ladder.md, 07 §6.1. Rulings
recorded before the work: 09 §3 decisions 87–89. Evidence for every claim: documentation/audits/os-docker-crash-2026-10-04/.
Part table
| Part | What | Result | Evidence |
|---|---|---|---|
| A | System page + version report (R-852) | DONE. The box reports Proxmox + kernel (API) and the wrapper's read-only facts (host Debian, next-boot kernel, held packages, taint, kernel.panic, crash guard; guest Debian, Docker, containerd, live-restore); unreadable = unknown. Hub: System tab on every page — per box ring + switch with buttons, tunnel, host, guest, Docker, last leg; releases per layer; what ring 0 runs; "Approve now" (confirm) and "Approve Docker set" (only when allowed). Hosts gets a Proxmox / kernel column. R-849: guest scanned every pass. Rendered with real data from both demo boxes. |
partA/ (live/system-page.html, hosts-page.html, facts) |
| B | Docker engine slow lane (11 §5.8) |
DONE. live-restore on by RELOAD: same ids on 9202 (6/6, by hand — R10 refuses a scratch guest by design), demo-hp (24/24, 7.5 s), demo-felhom (5/5). Ring 0 on both: 29.7.x → 29.8.2, every id kept (122 s / 99 s whole pass). Approval with the page button under a TEST 0-night wait (logged, reverted): os-docker-20261004-142842. Signed undo on demo-hp → 29.7.2 (wrapper authority=signed UNDO, 24/24 ids, 41 s) and back to 29.8.2 by its ring-0 pass. demo-felhom as ring 1: its night pass skips Docker; a signed job with the approved set verified by the wrapper (already current → nothing); the replayed job refused ("nonce already seen"); back to ring 0. |
partB/ |
| C | Crash restart with a limit (R-851) | DONE. Spike (your word before each crash): kernel.panic=10 + echo c → back by itself in 54 s, same kernel; pstore saved nothing; no oops history. Guard built (decision 88, 90–92). Live: crash 1 → 54 s, crash 2 → 53 s and the guard tripped, crash 3 → stayed off (184 s watched) until you switched it on; the hub mailed host_crash_guard_tripped (+3 restart events, 3 household lines); System page red; re-armed by felhom-crash-guard rearm. |
partC/ |
| D | Releases, golden, records | Agent v0.142.0 (signed to both demo boxes), hub v0.132.0, installer 1.30.0 (tagged, public, verified), build-golden.sh 3.1.0. Golden: see below. Records: 11 §5.7/§5.8/§5.9, 00, 03, 07, 08, decisions 90–94, two runbooks. |
partD/, git |
Claims in the brief that turned out wrong (or half right)
- "No box reports its versions" — half right: every box already sent
pveversionandkversioninside the Proxmox API answer the agent reads each report (NodeStatus); nothing stored or showed them. Debian, the next-boot kernel and everything Docker were truly not reported. - "A reload turns live-restore on with no container restart, on the demo boxes too" — TRUE, measured: 24/24 and 5/5 ids kept. But "prove on 9202 first" could not use the product path: 9202 binds a scratch folder, not the drives, so the wrapper refuses it (R10, by design); proved there by hand with the same two steps.
- "A crash boot can be told apart from a clean one" — only from a CLEAN one. Not from a power cut or a hard reset:
efi_pstoreis on, yet a real panic saved nothing on demo-hp. The guard counts every unclean stop (decision 90). - "
echo ccrashes the host andkernel.panicbrings it back" — TRUE (54 s, 53 s). Butsysctl -wdoes not survive the restart (back to 0), so it must be set at every boot — the guard does that. - The guard's count — the brief said both "at 3 … the next crash leaves the box off" (the 4th) and "crashes 3 times within one hour, it stays off" (the 3rd). Built per your words: the 3rd (decision 92).
Decisions taken by CC unattended (operator may reverse) — 09 §3
90 crash signal = clean-stop marker · 91 panic_on_oops stays 0, an oops is mailed · 92 the 3rd unclean stop in 60 min
stays off; 10 s; 24 h · 93 the wrapper verifies Docker authority against root-owned files · 94 page colours = alarm
thresholds.
Releases and what was copied by hand
- Agent v0.142.0 (
b1746c2, sha2567beb3222…de6): signedagent_updateto both demo boxes, both COMPLETED. - Hub v0.132.0 (
175ecfc): image tag verified on the pod; ArgoCD Synced/Healthy. - Installer 1.30.0: tag
installer-v1.30.0, both git-sync refs;https://felhom.eu/scripts/felhom-host-install.shservesSCRIPT_VERSION="1.30.0". - By hand on both demo hosts (R-840;
partD/copied-by-hand-*.txt, hashes equal to tag v0.142.0):/usr/local/sbin/felhom-os-apply,/usr/local/sbin/felhom-crash-guard,/etc/systemd/system/felhom-crash-guard.service,…/felhom-crash-guard-check.service,…/felhom-crash-guard-check.timer,/etc/felhom/crash-guard.conf,/etc/felhom/operator-signers,/etc/felhom/os-trust.json(withring0_slow_lane: true— the demo boxes only); units enabled. Previous wrapper kept as/root/felhom-os-apply.bak-0.141.1. A test binary (felhom-agent-0.142.0-rc1) ran the debug actions before the release and was removed after.
Golden
- Re-baked golden 0.292.0 with
build-golden.sh3.1.0 (by a helper agent, RUNBOOK-manual-build §4.0/§4.1; the documented publish replaces the same version: pre-delete 204, upload 201). The log shows "Docker engine set PINNED", the six approved versions (docker-ce 5:29.8.2, containerd.io 2.3.6, buildx 0.37.1, compose 5.6.0, …) and "live-restore: on". New sha25679a1dce3…d43a, re-hashed by download in the main session: match. Drill VM destroyed, drill disk back tovirgin. Evidencedocumentation/tests/golden-0.292.0-2026-10-04-rebake/(commit208d21d). - Re-vouched: agent 0.142.0 + golden 0.292.0 (new sha),
min_agent0.131.0 →artifacts_set(17:48). Between the re-bake upload and the re-vouch the hub vouched the old sha — a fresh install would have failed closed; none ran (R-857). - The crash guard is NOT in the golden (it is a host program): the installer 1.30.0 installs it.
- The golden waiver stays deleted: the newest controller (0.292.0) has its golden.
Register
Open rows 334 → 333 (333 at the start + R-852 filed first). Closed: R-852, R-835, R-848, R-849, R-851; filed and
closed: R-854. Opened: R-853 (facts reach the hub ~15 min late after a boot), R-855 (cosmetic TEST log line),
R-856 (after a crash the household also gets app mails — your choice later), R-857 (a same-version golden
re-bake: the gate shows the first sha; a window until the re-vouch). Narrowed: R-812 (Docker lane built;
kernel left), R-840 (by hand again). unproven.py: unchanged (35 of 55 not walked).
Teardown — three layers
- Machine: demo-hp and demo-felhom customer guests running, live-restore on, Docker 29.8.2, all apps up. 9202
running, live-restore on (its old daemon.json kept as
/root/daemon.json.bak-2026-10-04in the guest). - Host: both hosts run agent 0.142.0, the crash guard ARMED (
kernel.panic = 10), ring 0, switch ON; the test binary and the id lists removed. demo-hp booted 3 extra times today (the crash test); kernel unchanged (7.0.14-20). - Hub: the TEST Docker wait reverted (log: 2 nights); the approval
os-docker-20261004-142842stays (a real approval of what ring 0 runs). The crash and trip mails for demo-hp were sent on purpose. The hub password copy in the scratchpad shredded at the end. The hub announced the re-arm (host_crash_guard_rearmed, 17:47).
CI
CI_SECTION