# REPORT — System page, Docker slow lane, crash restart (2026-10-04, late afternoon–evening) Brief: "OS updates — a System page …; the Docker engine slow lane built (live-restore ON, decision); a crashed host restarts by itself, with a limit (decision)". Architecture read first: `documentation/architecture/11-os-updates.md` (owner; §5.6, §5.8, §8.1–8.3), `05-hub-architecture.md`, `03-host-agent.md`, `08-alarm-ladder.md`, `07` §6.1. Rulings recorded before the work: `09` §3 decisions 87–89. Evidence for every claim: `documentation/audits/os-docker-crash-2026-10-04/`. ## Part table | Part | What | Result | Evidence | |---|---|---|---| | A | System page + version report (R-852) | **DONE.** The box reports Proxmox + kernel (API) and the wrapper's read-only facts (host Debian, next-boot kernel, held packages, taint, `kernel.panic`, crash guard; guest Debian, Docker, containerd, live-restore); unreadable = `unknown`. Hub: **System** tab on every page — per box ring + switch with buttons, tunnel, host, guest, Docker, last leg; releases per layer; what ring 0 runs; "Approve now" (confirm) and "Approve Docker set" (only when allowed). Hosts gets a Proxmox / kernel column. R-849: guest scanned every pass. Rendered with real data from both demo boxes. | `partA/` (`live/system-page.html`, `hosts-page.html`, facts) | | B | Docker engine slow lane (`11` §5.8) | **DONE.** live-restore on by RELOAD: same ids on 9202 (6/6, by hand — R10 refuses a scratch guest by design), demo-hp (24/24, 7.5 s), demo-felhom (5/5). Ring 0 on both: 29.7.x → **29.8.2**, every id kept (122 s / 99 s whole pass). Approval with the page button under a TEST 0-night wait (logged, reverted): `os-docker-20261004-142842`. **Signed undo** on demo-hp → 29.7.2 (wrapper `authority=signed UNDO`, 24/24 ids, 41 s) and back to 29.8.2 by its ring-0 pass. **demo-felhom as ring 1**: its night pass skips Docker; a signed job with the approved set verified by the wrapper (already current → nothing); the **replayed** job refused ("nonce already seen"); back to ring 0. | `partB/` | | C | Crash restart with a limit (R-851) | **DONE.** Spike (your word before each crash): `kernel.panic=10` + `echo c` → back by itself in 54 s, same kernel; pstore saved nothing; no oops history. Guard built (decision 88, 90–92). Live: crash 1 → 54 s, crash 2 → 53 s and the guard **tripped**, crash 3 → **stayed off** (184 s watched) until you switched it on; the hub mailed `host_crash_guard_tripped` (+3 restart events, 3 household lines); System page red; re-armed by `felhom-crash-guard rearm`. | `partC/` | | D | Releases, golden, records | Agent **v0.142.0** (signed to both demo boxes), hub **v0.132.0**, installer **1.30.0** (tagged, public, verified), `build-golden.sh` 3.1.0. Golden: see below. Records: `11` §5.7/§5.8/§5.9, `00`, `03`, `07`, `08`, decisions 90–94, two runbooks. | `partD/`, git | ## Claims in the brief that turned out wrong (or half right) - **"No box reports its versions"** — half right: every box already sent `pveversion` and `kversion` inside the Proxmox API answer the agent reads each report (`NodeStatus`); nothing stored or showed them. Debian, the next-boot kernel and everything Docker were truly not reported. - **"A reload turns live-restore on with no container restart, on the demo boxes too"** — TRUE, measured: 24/24 and 5/5 ids kept. But "prove on 9202 first" could not use the product path: 9202 binds a scratch folder, not the drives, so the wrapper refuses it (R10, by design); proved there by hand with the same two steps. - **"A crash boot can be told apart from a clean one"** — only from a CLEAN one. Not from a power cut or a hard reset: `efi_pstore` is on, yet a real panic saved nothing on demo-hp. The guard counts every unclean stop (decision 90). - **"`echo c` crashes the host and `kernel.panic` brings it back"** — TRUE (54 s, 53 s). But `sysctl -w` does not survive the restart (back to 0), so it must be set at every boot — the guard does that. - **The guard's count** — the brief said both "at 3 … the next crash leaves the box off" (the 4th) and "crashes 3 times within one hour, it stays off" (the 3rd). Built per your words: the 3rd (decision 92). ## Decisions taken by CC unattended (operator may reverse) — `09` §3 90 crash signal = clean-stop marker · 91 `panic_on_oops` stays 0, an oops is mailed · 92 the 3rd unclean stop in 60 min stays off; 10 s; 24 h · 93 the wrapper verifies Docker authority against root-owned files · 94 page colours = alarm thresholds. ## Releases and what was copied by hand - **Agent v0.142.0** (`b1746c2`, sha256 `7beb3222…de6`): signed `agent_update` to both demo boxes, both COMPLETED. - **Hub v0.132.0** (`175ecfc`): image tag verified on the pod; ArgoCD Synced/Healthy. - **Installer 1.30.0**: tag `installer-v1.30.0`, both git-sync refs; `https://felhom.eu/scripts/felhom-host-install.sh` serves `SCRIPT_VERSION="1.30.0"`. - **By hand on both demo hosts** (R-840; `partD/copied-by-hand-*.txt`, hashes equal to tag v0.142.0): `/usr/local/sbin/felhom-os-apply`, `/usr/local/sbin/felhom-crash-guard`, `/etc/systemd/system/felhom-crash-guard.service`, `…/felhom-crash-guard-check.service`, `…/felhom-crash-guard-check.timer`, `/etc/felhom/crash-guard.conf`, `/etc/felhom/operator-signers`, `/etc/felhom/os-trust.json` (with `ring0_slow_lane: true` — the demo boxes only); units enabled. Previous wrapper kept as `/root/felhom-os-apply.bak-0.141.1`. A test binary (`felhom-agent-0.142.0-rc1`) ran the debug actions before the release and was removed after. ## Golden - **Re-baked golden 0.292.0** with `build-golden.sh` 3.1.0 (by a helper agent, RUNBOOK-manual-build §4.0/§4.1; the documented publish replaces the same version: pre-delete 204, upload 201). The log shows "Docker engine set PINNED", the six approved versions (docker-ce 5:29.8.2, containerd.io 2.3.6, buildx 0.37.1, compose 5.6.0, …) and "live-restore: on". New sha256 `79a1dce3…d43a`, re-hashed by download in the main session: match. Drill VM destroyed, drill disk back to `virgin`. Evidence `documentation/tests/golden-0.292.0-2026-10-04-rebake/` (commit `208d21d`). - **Re-vouched:** agent 0.142.0 + golden 0.292.0 (new sha), `min_agent` 0.131.0 → `artifacts_set` (17:48). Between the re-bake upload and the re-vouch the hub vouched the old sha — a fresh install would have failed closed; none ran (R-857). - The crash guard is NOT in the golden (it is a host program): the installer 1.30.0 installs it. - The golden waiver stays deleted: the newest controller (0.292.0) has its golden. ## Register Open rows **334 → 333** (333 at the start + R-852 filed first). Closed: **R-852, R-835, R-848, R-849, R-851**; filed and closed: **R-854**. Opened: **R-853** (facts reach the hub ~15 min late after a boot), **R-855** (cosmetic TEST log line), **R-856** (after a crash the household also gets app mails — your choice later), **R-857** (a same-version golden re-bake: the gate shows the first sha; a window until the re-vouch). Narrowed: **R-812** (Docker lane built; kernel left), **R-840** (by hand again). `unproven.py`: unchanged (35 of 55 not walked). ## Teardown — three layers - **Machine:** demo-hp and demo-felhom customer guests running, live-restore on, Docker 29.8.2, all apps up. 9202 running, live-restore on (its old daemon.json kept as `/root/daemon.json.bak-2026-10-04` in the guest). - **Host:** both hosts run agent 0.142.0, the crash guard ARMED (`kernel.panic = 10`), ring 0, switch ON; the test binary and the id lists removed. demo-hp booted 3 extra times today (the crash test); kernel unchanged (7.0.14-20). - **Hub:** the TEST Docker wait reverted (log: 2 nights); the approval `os-docker-20261004-142842` stays (a real approval of what ring 0 runs). The crash and trip mails for demo-hp were sent on purpose. The hub password copy in the scratchpad shredded at the end. The hub announced the re-arm (`host_crash_guard_rearmed`, 17:47). ## CI Checked by `head_sha` over every page of the Gitea `jobs` endpoint: felhom.eu — all 10 commits from `bee277f` (rulings) to `21986d0` (re-vouch records) **success** (incl. hub `175ecfc` 1265→, installer, manifests, docs); felhom-agent — `b1746c2` (1273, 1274), `f24dce5` (1275), `42af3ab` (1281) **success**. This report's own commit: checked after the push (see the session's final message). ## Addendum (~18:30) — R-858, found by the operator The N100 showed DOWN from 14:18 UTC. Cause: v0.142.0's Docker step restarted dockerd (14:13); live-restore kept the containers running, but `felhom-controller` and `traefik` bind-mount the socket FILE and kept the deleted inode, so the controller could not reach Docker. My Docker health rule passed it (the controller's own check said healthy) — the rule checked the mechanism, not the consequence. Repaired by restarting the two containers (15:57 UTC). Ruling 95: agent **v0.142.1** (`4950030`, sha256 `003f882a…62bd`) restarts only the socket users after a step and fails health when the controller cannot reach Docker. Proven live on demo-hp before the release (signed undo, then forward): `applied, healthy`, guest / controller / traefik on the same socket inode both times. Ring-0 marks were OFF during the fix, back ON after. Both boxes on 0.142.1; vouched for new installs (golden unchanged). Red-proofs 4/4. Register: R-858 opened and closed (still **333** open). Evidence `partE-incident/`. Not touched: Tester 1 shows DOWN for 4 days on the dashboard — a fenced tester box, outside this brief.