Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
3.2 KiB
Put a box's Docker engine back one set (OS updates, Docker slow lane)
When: the hub mailed
os_update_health_failedfor the docker layer, or the System page shows a box whose apps broke after a Docker engine step. Who: the operator, or CC with the operator's word (CC may sign until the first paying customer, R-530 ruling). Owner design:architecture/11-os-updates.md§5.8. Proved on demo-hp 2026-10-04 — evidenceaudits/os-docker-crash-2026-10-04/partB/undo/.
The undo is an ordinary signed os_docker_step with "undo": true that names the previous engine set. The root
wrapper re-verifies the signature itself (against the root-owned /etc/felhom/operator-signers), allows the downgrade
only because the signed job says undo, and refuses unless live-restore is on — so the undo, like the step, restarts
no app. Nothing is done by hand on the box.
1. Find the previous set
- The box's previous Docker report (System page → Docker release column, or the hub's
os_reportsfor the box, layerdocker): itsinstalledlist before the bad step. Or, on the host:pct exec <vmid> -- grep -A3 "Start-Date" /var/log/apt/history.log | tail— each linename:amd64 (old, new). - All six names, each with its old version. Docker's repository keeps old versions (
11C2), so no snapshot is needed.
2. Sign and queue it
cd /mnt/5_hdd/felhom.eu/git/felhom-agent && go build -o /tmp/felhom-opsign ./cmd/felhom-opsign
P='{"release_id":"undo-to-<engine>","undo":true,"packages":[
{"name":"containerd.io","version":"<old>","origin":"Docker CE"},
{"name":"docker-buildx-plugin","version":"<old>","origin":"Docker CE"},
{"name":"docker-ce","version":"<old>","origin":"Docker CE"},
{"name":"docker-ce-cli","version":"<old>","origin":"Docker CE"},
{"name":"docker-ce-rootless-extras","version":"<old>","origin":"Docker CE"},
{"name":"docker-compose-plugin","version":"<old>","origin":"Docker CE"}]}'
KEY=$(sudo kubectl -n felhom-system get secret report-api -o jsonpath='{.data.REPORT_API_KEY}' | base64 -d)
/tmp/felhom-opsign -op os_docker_step -host <host_id> -key-id felhom-op-1 \
-key /mnt/5_hdd/felhom.eu/felhom-op-operational -params "$P" -ttl 45m \
-upload http://<hub ClusterIP>:8080 -hub-key "$KEY"
unset KEY
The box takes the job at its next poll (up to 15 minutes), under the heavy-op gate (never beside a backup).
3. Check
- Agent journal:
signedjobs: AUTHORIZED signed op→os-apply: START … layer=docker:… authority=signed UNDO→signedjobs: signed op COMPLETED. AREFUSED: R3means the signature, host, time window, nonce or package list did not match;R15means live-restore is off. - The hub: a docker-layer report,
outcome applied,healthy,undo: true, the old engine. - The same container ids before and after (the health rule checks it; a changed id is
health_failed).
Afterwards
On a ring-0 box the next night step installs the newest set again — switch the box's updates OFF on the System page first if that set is the bad one, and record it in the incident's register row. A ring-1 box takes a set only by a signed job, so it stays on the undone set until someone signs the next one.