Files
felhom.eu/REPORT-os-docker-crash-2026-10-04.md
T

100 lines
9.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# REPORT — System page, Docker slow lane, crash restart (2026-10-04, late afternoon–evening)
Brief: "OS updates — a System page …; the Docker engine slow lane built (live-restore ON, decision); a crashed host
restarts by itself, with a limit (decision)". Architecture read first: `documentation/architecture/11-os-updates.md`
(owner; §5.6, §5.8, §8.1–8.3), `05-hub-architecture.md`, `03-host-agent.md`, `08-alarm-ladder.md`, `07` §6.1. Rulings
recorded before the work: `09` §3 decisions 87–89. Evidence for every claim: `documentation/audits/os-docker-crash-2026-10-04/`.
## Part table
| Part | What | Result | Evidence |
|---|---|---|---|
| A | System page + version report (R-852) | **DONE.** The box reports Proxmox + kernel (API) and the wrapper's read-only facts (host Debian, next-boot kernel, held packages, taint, `kernel.panic`, crash guard; guest Debian, Docker, containerd, live-restore); unreadable = `unknown`. Hub: **System** tab on every page — per box ring + switch with buttons, tunnel, host, guest, Docker, last leg; releases per layer; what ring 0 runs; "Approve now" (confirm) and "Approve Docker set" (only when allowed). Hosts gets a Proxmox / kernel column. R-849: guest scanned every pass. Rendered with real data from both demo boxes. | `partA/` (`live/system-page.html`, `hosts-page.html`, facts) |
| B | Docker engine slow lane (`11` §5.8) | **DONE.** live-restore on by RELOAD: same ids on 9202 (6/6, by hand — R10 refuses a scratch guest by design), demo-hp (24/24, 7.5 s), demo-felhom (5/5). Ring 0 on both: 29.7.x → **29.8.2**, every id kept (122 s / 99 s whole pass). Approval with the page button under a TEST 0-night wait (logged, reverted): `os-docker-20261004-142842`. **Signed undo** on demo-hp → 29.7.2 (wrapper `authority=signed UNDO`, 24/24 ids, 41 s) and back to 29.8.2 by its ring-0 pass. **demo-felhom as ring 1**: its night pass skips Docker; a signed job with the approved set verified by the wrapper (already current → nothing); the **replayed** job refused ("nonce already seen"); back to ring 0. | `partB/` |
| C | Crash restart with a limit (R-851) | **DONE.** Spike (your word before each crash): `kernel.panic=10` + `echo c` → back by itself in 54 s, same kernel; pstore saved nothing; no oops history. Guard built (decision 88, 90–92). Live: crash 1 → 54 s, crash 2 → 53 s and the guard **tripped**, crash 3 → **stayed off** (184 s watched) until you switched it on; the hub mailed `host_crash_guard_tripped` (+3 restart events, 3 household lines); System page red; re-armed by `felhom-crash-guard rearm`. | `partC/` |
| D | Releases, golden, records | Agent **v0.142.0** (signed to both demo boxes), hub **v0.132.0**, installer **1.30.0** (tagged, public, verified), `build-golden.sh` 3.1.0. Golden: see below. Records: `11` §5.7/§5.8/§5.9, `00`, `03`, `07`, `08`, decisions 90–94, two runbooks. | `partD/`, git |
## Claims in the brief that turned out wrong (or half right)
- **"No box reports its versions"** — half right: every box already sent `pveversion` and `kversion` inside the Proxmox
API answer the agent reads each report (`NodeStatus`); nothing stored or showed them. Debian, the next-boot kernel and
everything Docker were truly not reported.
- **"A reload turns live-restore on with no container restart, on the demo boxes too"** — TRUE, measured: 24/24 and 5/5
ids kept. But "prove on 9202 first" could not use the product path: 9202 binds a scratch folder, not the drives, so the
wrapper refuses it (R10, by design); proved there by hand with the same two steps.
- **"A crash boot can be told apart from a clean one"** — only from a CLEAN one. Not from a power cut or a hard reset:
`efi_pstore` is on, yet a real panic saved nothing on demo-hp. The guard counts every unclean stop (decision 90).
- **"`echo c` crashes the host and `kernel.panic` brings it back"** — TRUE (54 s, 53 s). But `sysctl -w` does not survive
the restart (back to 0), so it must be set at every boot — the guard does that.
- **The guard's count** — the brief said both "at 3 … the next crash leaves the box off" (the 4th) and "crashes 3 times
within one hour, it stays off" (the 3rd). Built per your words: the 3rd (decision 92).
## Decisions taken by CC unattended (operator may reverse) — `09` §3
90 crash signal = clean-stop marker · 91 `panic_on_oops` stays 0, an oops is mailed · 92 the 3rd unclean stop in 60 min
stays off; 10 s; 24 h · 93 the wrapper verifies Docker authority against root-owned files · 94 page colours = alarm
thresholds.
## Releases and what was copied by hand
- **Agent v0.142.0** (`b1746c2`, sha256 `7beb3222…de6`): signed `agent_update` to both demo boxes, both COMPLETED.
- **Hub v0.132.0** (`175ecfc`): image tag verified on the pod; ArgoCD Synced/Healthy.
- **Installer 1.30.0**: tag `installer-v1.30.0`, both git-sync refs; `https://felhom.eu/scripts/felhom-host-install.sh`
serves `SCRIPT_VERSION="1.30.0"`.
- **By hand on both demo hosts** (R-840; `partD/copied-by-hand-*.txt`, hashes equal to tag v0.142.0):
`/usr/local/sbin/felhom-os-apply`, `/usr/local/sbin/felhom-crash-guard`, `/etc/systemd/system/felhom-crash-guard.service`,
`…/felhom-crash-guard-check.service`, `…/felhom-crash-guard-check.timer`, `/etc/felhom/crash-guard.conf`,
`/etc/felhom/operator-signers`, `/etc/felhom/os-trust.json` (with `ring0_slow_lane: true` — the demo boxes only);
units enabled. Previous wrapper kept as `/root/felhom-os-apply.bak-0.141.1`. A test binary (`felhom-agent-0.142.0-rc1`)
ran the debug actions before the release and was removed after.
## Golden
- **Re-baked golden 0.292.0** with `build-golden.sh` 3.1.0 (by a helper agent, RUNBOOK-manual-build §4.0/§4.1; the
documented publish replaces the same version: pre-delete 204, upload 201). The log shows "Docker engine set PINNED",
the six approved versions (docker-ce 5:29.8.2, containerd.io 2.3.6, buildx 0.37.1, compose 5.6.0, …) and
"live-restore: on". New sha256 `79a1dce3…d43a`, re-hashed by download in the main session: match. Drill VM destroyed,
drill disk back to `virgin`. Evidence `documentation/tests/golden-0.292.0-2026-10-04-rebake/` (commit `208d21d`).
- **Re-vouched:** agent 0.142.0 + golden 0.292.0 (new sha), `min_agent` 0.131.0 → `artifacts_set` (17:48). Between the
re-bake upload and the re-vouch the hub vouched the old sha — a fresh install would have failed closed; none ran (R-857).
- The crash guard is NOT in the golden (it is a host program): the installer 1.30.0 installs it.
- The golden waiver stays deleted: the newest controller (0.292.0) has its golden.
## Register
Open rows **334 → 333** (333 at the start + R-852 filed first). Closed: **R-852, R-835, R-848, R-849, R-851**; filed and
closed: **R-854**. Opened: **R-853** (facts reach the hub ~15 min late after a boot), **R-855** (cosmetic TEST log line),
**R-856** (after a crash the household also gets app mails — your choice later), **R-857** (a same-version golden
re-bake: the gate shows the first sha; a window until the re-vouch). Narrowed: **R-812** (Docker lane built;
kernel left), **R-840** (by hand again). `unproven.py`: unchanged (35 of 55 not walked).
## Teardown — three layers
- **Machine:** demo-hp and demo-felhom customer guests running, live-restore on, Docker 29.8.2, all apps up. 9202
running, live-restore on (its old daemon.json kept as `/root/daemon.json.bak-2026-10-04` in the guest).
- **Host:** both hosts run agent 0.142.0, the crash guard ARMED (`kernel.panic = 10`), ring 0, switch ON; the test
binary and the id lists removed. demo-hp booted 3 extra times today (the crash test); kernel unchanged (7.0.14-20).
- **Hub:** the TEST Docker wait reverted (log: 2 nights); the approval `os-docker-20261004-142842` stays (a real
approval of what ring 0 runs). The crash and trip mails for demo-hp were sent on purpose. The hub password copy in
the scratchpad shredded at the end. The hub announced the re-arm (`host_crash_guard_rearmed`, 17:47).
## CI
Checked by `head_sha` over every page of the Gitea `jobs` endpoint: felhom.eu — all 10 commits from `bee277f` (rulings)
to `21986d0` (re-vouch records) **success** (incl. hub `175ecfc` 1265→, installer, manifests, docs); felhom-agent —
`b1746c2` (1273, 1274), `f24dce5` (1275), `42af3ab` (1281) **success**. This report's own commit: checked after the push
(see the session's final message).
## Addendum (~18:30) — R-858, found by the operator
The N100 showed DOWN from 14:18 UTC. Cause: v0.142.0's Docker step restarted dockerd (14:13); live-restore kept the
containers running, but `felhom-controller` and `traefik` bind-mount the socket FILE and kept the deleted inode, so the
controller could not reach Docker. My Docker health rule passed it (the controller's own check said healthy) — the rule
checked the mechanism, not the consequence. Repaired by restarting the two containers (15:57 UTC). Ruling 95: agent
**v0.142.1** (`4950030`, sha256 `003f882a…62bd`) restarts only the socket users after a step and fails health when the
controller cannot reach Docker. Proven live on demo-hp before the release (signed undo, then forward): `applied, healthy`,
guest / controller / traefik on the same socket inode both times. Ring-0 marks were OFF during the fix, back ON after.
Both boxes on 0.142.1; vouched for new installs (golden unchanged). Red-proofs 4/4. Register: R-858 opened and closed
(still **333** open). Evidence `partE-incident/`. Not touched: Tester 1 shows DOWN for 4 days on the dashboard — a
fenced tester box, outside this brief.