Compare commits

...

3 Commits

7 changed files with 154 additions and 13 deletions
+43
View File
@@ -1,3 +1,46 @@
## v0.142.0 — the Docker engine slow lane, the version report, the crash guard (`11` §5.8, §5.9; `09` decisions 87–89)
> **RELEASED 2026-10-04** by `scripts/release-agent.sh` — tag `v0.142.0` (`b1746c2`), sha256
> `7beb32224d6495e9561acfd3ad8a48393a799520f196080011cb27fceb6d1de6`, verified by download. Not vouched at release time.
**MinAgent impact:** none. **Needs hub v0.132.0** for the System page, the Docker approval and the crash events; an older
hub stores the new `system` stanza unread. **Needs the new root files** on an installed box (R-840): the wrapper,
`/etc/felhom/os-trust.json`, `/etc/felhom/operator-signers` and the crash guard — the installer 1.30.0 writes them; the
demo boxes got them by hand.
- **Docker `live-restore` ON** (decision 87). Wrapper mode `live-restore-on`: merge `"live-restore": true` into the
guest's `/etc/docker/daemon.json` and `systemctl reload docker` — never a restart (R-835). An invalid daemon.json is
left alone (R16); a reload that does not enable it puts the old file back. The leg runs it once before a Docker step;
`--selftest=live-restore -vmid N` runs it by hand. Measured: demo-hp 24 containers, demo-felhom 5 — the same ids after.
- **The Docker engine slow lane** (`11` §5.8). Wrapper layer `docker`, lane `slow` only, the six Docker packages only,
origin `Docker CE` only (R2; a Docker package in a fast-lane plan is refused). **R3 — the wrapper checks the authority
itself**, never the agent's config: a ring-1 step or ANY undo needs a signed `os_docker_step` it verifies with
`ssh-keygen -Y verify` against the ROOT-owned `/etc/felhom/operator-signers` (namespace `felhom-op-v1`, the blob's
`key_id`), bound to `/etc/felhom/os-trust.json` `host_id`, inside its time window, never replayed (a root-owned nonce
file), with exactly the signed packages and undo flag; an unsigned ring-0 step needs that file's
`"ring0_slow_lane": true` (the demo boxes only, set by hand). **R15:** live-restore must be on. An undo may downgrade
(`--allow-downgrades`) only inside a signed job. The leg: ring 0 runs the Docker step at night after a healthy guest
and host step (`select pending-docker`); ring 1 never does — only `DockerStepExecutor` (signed job, heavy-op gate).
**Health:** the guest rule + every container running at the start has the SAME id after + the engine reports the
installed version; a changed id is `health_failed`. Measured ring 0: 29.7.x → 29.8.2 on both demo boxes, every id kept.
- **The version report** (R-852, decision 89). Wrapper mode `facts` (read-only): host Debian, running and next-boot
kernel (`next_entry` > saved default > newest installed, by dpkg order), held packages (R-848), kernel taint (oops,
warn), `kernel.panic`, the crash guard state; guest Debian, Docker engine, containerd, live-restore. The host report
gains `system {pve_version, kernel_version, vmid, facts, facts_error}`, read at most every 10 min (~2 s);
`--selftest=os-facts -vmid N`. A value nobody could read is `unknown`.
- **R-849:** the guest is scanned for "restart needed" on every pass too, so a guest restart clears it.
- **The crash guard** (decision 88, R-851): `configs/felhom-crash-guard` + `felhom-crash-guard.service` (early boot;
its ExecStop writes a clean-stop marker) + an hourly re-arm timer + `/etc/felhom/crash-guard.conf`. A boot without
the marker followed an unclean stop (a crash, a power cut or a hard reset — pstore saved nothing for a real panic on
demo-hp, so they cannot be told apart). Armed: `kernel.panic = 10`. After the 2nd unclean boot within 60 min it
TRIPS (`kernel.panic = 0`), so the 3rd crash within the hour leaves the box off; it re-arms after 24 h of normal
running or `felhom-crash-guard rearm`. State in `/var/lib/felhom-crash-guard/state.json` (0644; read by facts).
- **After the tag (main only, not shipped to boxes): `configs/build-golden.sh` 3.1.0** — the golden's `daemon.json`
carries `"live-restore": true` with a fail-closed assertion, and `GOLDEN_DOCKER_PKGS` pins the approved Docker engine
set (all six `name=version`); without it the bake log warns that the set is the newest, not an approved one.
- Tests: wrapper 76 (DockerLane, LiveRestore, Facts, RealSignatureCheck with a throwaway key), crash guard 9, Go leg +
executor; red-proofs `felhom.eu/documentation/audits/os-docker-crash-2026-10-04/partB/agent-redproofs.txt` (17 caught).
## v0.141.1 — "reboot needed" is true on the host (found live on demo-felhom, 2026-10-04)
> **RELEASED 2026-10-04** by `scripts/release-agent.sh` — tag `v0.141.1` (`a6bc3f1`), sha256
+7 -9
View File
@@ -1,11 +1,9 @@
# REPORT — 2026-10-04: v0.141.0 + v0.141.1, the host fast lane, the true tunnel status, the fast leg
# REPORT — 2026-10-04: v0.142.0, Docker slow lane + version report + crash guard
Full session report: `felhom.eu/REPORT-os-host-lane-2026-10-04.md`.
Full session report: `felhom.eu/REPORT-os-docker-crash-2026-10-04.md`.
- Tunnel (R-841): the agent reads the guest's cloudflared container and its health check; three states.
- Host fast lane: the wrapper gains the host layer (R12 appliance proof from the root-owned install record, R14 no
kernel/boot/firmware); the leg runs the host step after a healthy guest step; host health rule.
- Speed (R-845): one call per layer instead of one per package; measured before/after in the session report.
- Tests green; red-proofs in the audit folder.
- v0.141.1 (same day): the host "reboot needed" scan hid `lxc-start` and was never cleared by a reboot — both fixed,
found live on demo-felhom, red-proved.
- Docker live-restore turned on by reload (never a restart); the Docker engine set as a slow lane whose authority the
root wrapper checks itself (signed job against a root-owned key file, or the root-owned ring-0 mark); same-id health.
- The box reports its versions (Proxmox, kernels, Debian, Docker, live-restore, held packages, taint, crash guard).
- The crash guard: a crashed host restarts, the 3rd unclean stop within an hour leaves it off, 24 h re-arm.
- Tests and 17 red-proofs; live on both demo boxes (see the session report).
+21 -3
View File
@@ -61,7 +61,7 @@ set -euo pipefail
# Script provenance — logged into every bake transcript next to the baked controller tag, so an
# archive can always be traced to the script that produced it. Bump on any behavior change.
GOLDEN_SCRIPT_VERSION="3.0.0"
GOLDEN_SCRIPT_VERSION="3.1.0"
VMID="${1:-9100}"
TEMPLATE="${2:-local:vztmpl/debian-13-standard_13.1-2_amd64.tar.zst}"
@@ -112,7 +112,15 @@ for i in $(seq 1 30); do
if pct exec "$VMID" -- getent hosts download.docker.com >/dev/null 2>&1; then break; fi
sleep 1
done
pct exec "$VMID" -- bash -c '
# v3.1.0 (`11` §5.8): GOLDEN_DOCKER_PKGS pins the APPROVED Docker engine set (the hub's newest Docker release, all six
# "name=version"); without it the newest stable set is installed and the bake log says so.
GOLDEN_DOCKER_PKGS="${GOLDEN_DOCKER_PKGS:-}"
if [[ -n "$GOLDEN_DOCKER_PKGS" ]]; then
echo "[golden] Docker engine set PINNED to the approved release: $GOLDEN_DOCKER_PKGS"
else
echo "[golden] WARNING: GOLDEN_DOCKER_PKGS not set — installing the newest stable Docker set, not an approved one"
fi
pct exec "$VMID" -- env GOLDEN_DOCKER_PKGS="$GOLDEN_DOCKER_PKGS" bash -c '
set -e
export DEBIAN_FRONTEND=noninteractive
apt-get update -qq
@@ -122,7 +130,12 @@ pct exec "$VMID" -- bash -c '
echo "deb [signed-by=/etc/apt/keyrings/docker.asc] https://download.docker.com/linux/debian trixie stable" \
> /etc/apt/sources.list.d/docker.list
apt-get update -qq
apt-get install -y -qq docker-ce docker-ce-cli containerd.io >/dev/null
if [ -n "$GOLDEN_DOCKER_PKGS" ]; then
apt-get install -y -qq $GOLDEN_DOCKER_PKGS >/dev/null
else
apt-get install -y -qq docker-ce docker-ce-cli containerd.io >/dev/null
fi
dpkg-query -W containerd.io docker-buildx-plugin docker-ce docker-ce-cli docker-ce-rootless-extras docker-compose-plugin 2>/dev/null | sed "s/^/ installed: /"
'
echo "[golden] baking daemon.json: classic overlay2 driver (containerd-snapshotter OFF) + log rotation …"
# containerd-snapshotter (Docker 28+/29 default) keeps the IMAGE content store under
@@ -135,9 +148,12 @@ echo "[golden] baking daemon.json: classic overlay2 driver (containerd-snapshott
# layout (a container's `df /` reports the single volume, phase-0 spike). Since v3.0.0 /var/lib/docker
# is a BIND of <volume>/docker rather than the mp0 mount itself, wired immediately below; data-root
# still needs no override because the path is unchanged. Log caps kill the most common runaway.
# v3.1.0: "live-restore": true (`09` decision 87) — a Docker engine update then restarts no app (`11` C5). A box made
# from this golden never needs the agent's one-time live-restore-on step. NEVER removed by a plain restart (R-835).
pct exec "$VMID" -- bash -c 'mkdir -p /etc/docker; cat > /etc/docker/daemon.json <<JSON
{
"features": { "containerd-snapshotter": false },
"live-restore": true,
"log-driver": "json-file",
"log-opts": { "max-size": "10m", "max-file": "3" }
}
@@ -175,6 +191,8 @@ pct exec "$VMID" -- bash -c 'systemctl restart docker; sleep 3; docker run --rm
# Guard: the image store MUST be on the data volume now. /var/lib/containerd holding the images would
# mean containerd-snapshotter is still on (the split would leave images on the rootfs).
pct exec "$VMID" -- bash -c 'drv=$(docker info 2>/dev/null | sed -n "s/.*Storage Driver: //p"); [ "$drv" = "overlay2" ] || { echo "[golden] FATAL: storage driver is $drv, expected overlay2 — images would not land on the data volume"; exit 1; }'
# v3.1.0 ASSERTION: live-restore is ON in the running daemon (decision 87), or the bake fails closed.
pct exec "$VMID" -- bash -c 'lr=$(docker info --format "{{.LiveRestoreEnabled}}" 2>/dev/null); [ "$lr" = "true" ] && echo " live-restore: on" || { echo "[golden] FATAL: live-restore is $lr, expected true (decision 87)"; exit 1; }'
# ASSERTION 1 (RETARGETED v3.0.0, not removed). /var/lib/docker must be a real mount — now the V-c
# bind of <volume>/docker rather than the mp0 mount itself. Still fails closed on the same failure:
# if the bind did not take, Docker's data-root silently sits on the OS rootfs and the golden ships
+28 -1
View File
@@ -86,6 +86,11 @@ SIG_NAMESPACE = "felhom-op-v1"
SIGNED_OP = "os_docker_step"
NONCE_FILE = "/var/lib/felhom-os-apply/nonces.json"
DAEMON_JSON = "/etc/docker/daemon.json"
# R-858 (v0.142.1): a Docker engine step restarts dockerd, which RECREATES the socket file. With live-restore the
# containers keep running — and one that bind-mounts the socket FILE keeps the deleted inode: measured 2026-10-04 on
# demo-felhom, the controller and traefik were blind to Docker for 1h44m. After a step that installed something, the
# wrapper restarts exactly the containers that mount one of these paths (never the apps, never the engine).
DOCKER_SOCKETS = ("/var/run/docker.sock", "/run/docker.sock")
CRASH_GUARD_STATE = "/var/lib/felhom-crash-guard/state.json"
@@ -575,7 +580,10 @@ class Apply:
if len(p) >= 4 and p[3]:
cont[p[0]]["id"] = p[3]
nrc, _, _ = self.g(["getent", "hosts", "deb.debian.org"], timeout=30)
return {"docker_ok": rc == 0, "containers": cont,
# R-858: the controller's own health check stayed "healthy" while it could not reach Docker at all — so ask
# the consequence directly: can the controller talk to the engine from inside its container?
crc, _, _ = self.g(["docker", "exec", "felhom-controller", "docker", "version", "--format", "{{.Server.Version}}"], timeout=60)
return {"docker_ok": rc == 0, "containers": cont, "controller_docker_ok": crc == 0,
"controller": cont.get("felhom-controller", {}).get("health", "absent"),
"network_ok": nrc == 0}
@@ -728,6 +736,23 @@ class Apply:
out.append({"name": p["name"], "version": p["to"], "origin": "Debian-Security" if "Debian-Security" in o else "Debian"})
return out
def restart_socket_users(self):
"""R-858: restart ONLY the containers that bind-mount the Docker socket, so they attach to the new one."""
rc, out, _ = self.g(["docker", "ps", "-q", "--no-trunc"], timeout=60)
users = []
for cid in out.split():
irc, iout, _ = self.g(["docker", "inspect", "-f", "{{.Name}}|{{range .Mounts}}{{.Destination}};{{end}}", cid], timeout=60)
if irc != 0 or "|" not in iout:
continue
name, mounts = iout.strip().split("|", 1)
if any(m in DOCKER_SOCKETS for m in mounts.split(";")):
users.append(name.lstrip("/"))
users.sort()
if users:
rrc, _, rerr = self.g(["docker", "restart"] + users, timeout=300)
self.r.log(f"os-apply: SOCKET-USERS restarted={','.join(users)} rc={rrc} (R-858: they held the old docker socket)")
return users
def pending_docker(self):
"""Ring 0 (select pending-docker): the newest pending version of each INSTALLED Docker package, Docker origin."""
rc, pend, remv, _ = self.simulate(["dist-upgrade"])
@@ -837,6 +862,8 @@ class Apply:
return 3, None
self.report["upgraded"] = [{"name": n, "version": v} for n, v in upgrade]
self.report["seconds"] = round(secs, 1)
if self.layer == "docker":
self.report["socket_restarted"] = self.restart_socket_users()
procs, reboot = self.restart_needed()
self.report["restart_needed"] = procs
self.report["docker_restart_needed"] = any(p in ("dockerd", "containerd") for p in procs)
+28
View File
@@ -186,6 +186,14 @@ class Fake:
return 0, self.engine + "\n", ""
if cmd == "docker" and a[1:3] == ["ps", "-q"]:
return 0, "".join(i + "\n" for i in self.ids), ""
if cmd == "docker" and a[1] == "inspect":
mounts = {"aaa111": "/felhom-controller|/var/run/docker.sock;/app/data;", "bbb222": "/app|/data;"}
return 0, mounts.get(a[-1], "/other|;") + "\n", ""
if cmd == "docker" and a[1] == "restart":
self.restarted_containers = a[2:]
return 0, "", ""
if cmd == "docker" and a[1] == "exec":
return (1, "", "Cannot connect to the Docker daemon") if getattr(self, "controller_blind", False) else (0, self.engine + "\n", "")
if cmd == "docker":
ids = self.ids + ["x"] * 2
return 0, f"felhom-controller\trunning\tUp 1 hour (healthy)\t{ids[0]}\napp\trunning\tUp 1 hour (healthy)\t{ids[1]}\n", ""
@@ -804,6 +812,26 @@ class DockerLane(unittest.TestCase):
rc, rep = run(f)
self.assertEqual(rep["plan"]["upgrade"], 0, "an older version on an unsigned step is 'already', never installed")
def test_step_restarts_only_the_socket_users(self):
# R-858: after an engine step, ONLY the container that mounts the docker socket is restarted (here the controller)
f = docker_fake(signed=signed_job())
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertEqual(f.restarted_containers, ["felhom-controller"])
self.assertEqual(rep["socket_restarted"], ["felhom-controller"])
def test_no_install_restarts_nothing(self):
f = docker_fake(signed=signed_job())
f.installed.update({"docker-ce": "5:29.8.2-1~debian.13~trixie", "containerd.io": "2.3.6-1~debian.13~trixie"})
rc, rep = run(f)
self.assertFalse(hasattr(f, "restarted_containers"), "nothing installed -> no container restart")
def test_health_says_when_the_controller_cannot_reach_docker(self):
f = docker_fake(signed=signed_job())
f.controller_blind = True
rc, rep = run(f)
self.assertFalse(rep["health_after"]["controller_docker_ok"])
def test_health_carries_container_ids(self):
f = docker_fake(signed=signed_job())
rc, rep = run(f)
+6
View File
@@ -75,6 +75,9 @@ type Health struct {
NetworkOK bool `json:"network_ok"`
Controller string `json:"controller"`
Containers map[string]Container `json:"containers"`
// ControllerDockerOK: the controller reaches the engine from INSIDE its container (R-858, wrapper ≥ v0.142.1;
// nil from an older wrapper = not checked). Its own health check stayed "healthy" while it was blind.
ControllerDockerOK *bool `json:"controller_docker_ok,omitempty"`
HostServices map[string]string `json:"host_services,omitempty"`
GuestRunning *bool `json:"guest_running,omitempty"`
Guest *Health `json:"guest,omitempty"`
@@ -240,6 +243,9 @@ func HealthVerdict(before, after *Health) (bool, string) {
if after.Controller != "healthy" {
return false, "the controller is " + after.Controller
}
if after.ControllerDockerOK != nil && !*after.ControllerDockerOK {
return false, "the controller cannot reach Docker (it holds an old socket — R-858)"
}
if before == nil {
return true, ""
}
+21
View File
@@ -465,3 +465,24 @@ func TestDocker_ChangedIDIsHealthFailed(t *testing.T) {
t.Fatalf("docker = %+v", p.Docker)
}
}
// R-858: a controller that cannot reach Docker fails the health rule even though its own check says healthy.
// Red-proof: drop the ControllerDockerOK check in HealthVerdict and this fails.
func TestHealthVerdict_ControllerBlindToDockerFails(t *testing.T) {
no, yes2 := false, true
after := guestOK()
after.ControllerDockerOK = &no
if ok, why := HealthVerdict(guestOK(), after); ok || !strings.Contains(why, "R-858") {
t.Fatalf("a blind controller passed: %v %q", ok, why)
}
if ok, _ := DockerHealthVerdict(guestOK(), after, "", ""); ok {
t.Fatal("the Docker rule passed a blind controller")
}
after.ControllerDockerOK = &yes2
if ok, why := HealthVerdict(guestOK(), after); !ok {
t.Fatalf("a seeing controller failed: %s", why)
}
if ok, _ := HealthVerdict(guestOK(), guestOK()); !ok {
t.Fatal("an older wrapper (no field) must not fail")
}
}