Files
felhom.eu/documentation/runbooks/os-updates-docker-undo.md
T

3.2 KiB

Put a box's Docker engine back one set (OS updates, Docker slow lane)

When: the hub mailed os_update_health_failed for the docker layer, or the System page shows a box whose apps broke after a Docker engine step. Who: the operator, or CC with the operator's word (CC may sign until the first paying customer, R-530 ruling). Owner design: architecture/11-os-updates.md §5.8. Proved on demo-hp 2026-10-04 — evidence audits/os-docker-crash-2026-10-04/partB/undo/.

The undo is an ordinary signed os_docker_step with "undo": true that names the previous engine set. The root wrapper re-verifies the signature itself (against the root-owned /etc/felhom/operator-signers), allows the downgrade only because the signed job says undo, and refuses unless live-restore is on — so the undo, like the step, restarts no app. Nothing is done by hand on the box.

1. Find the previous set

  • The box's previous Docker report (System page → Docker release column, or the hub's os_reports for the box, layer docker): its installed list before the bad step. Or, on the host: pct exec <vmid> -- grep -A3 "Start-Date" /var/log/apt/history.log | tail — each line name:amd64 (old, new).
  • All six names, each with its old version. Docker's repository keeps old versions (11 C2), so no snapshot is needed.

2. Sign and queue it

cd /mnt/5_hdd/felhom.eu/git/felhom-agent && go build -o /tmp/felhom-opsign ./cmd/felhom-opsign
P='{"release_id":"undo-to-<engine>","undo":true,"packages":[
  {"name":"containerd.io","version":"<old>","origin":"Docker CE"},
  {"name":"docker-buildx-plugin","version":"<old>","origin":"Docker CE"},
  {"name":"docker-ce","version":"<old>","origin":"Docker CE"},
  {"name":"docker-ce-cli","version":"<old>","origin":"Docker CE"},
  {"name":"docker-ce-rootless-extras","version":"<old>","origin":"Docker CE"},
  {"name":"docker-compose-plugin","version":"<old>","origin":"Docker CE"}]}'
KEY=$(sudo kubectl -n felhom-system get secret report-api -o jsonpath='{.data.REPORT_API_KEY}' | base64 -d)
/tmp/felhom-opsign -op os_docker_step -host <host_id> -key-id felhom-op-1 \
  -key /mnt/5_hdd/felhom.eu/felhom-op-operational -params "$P" -ttl 45m \
  -upload http://<hub ClusterIP>:8080 -hub-key "$KEY"
unset KEY

The box takes the job at its next poll (up to 15 minutes), under the heavy-op gate (never beside a backup).

3. Check

  • Agent journal: signedjobs: AUTHORIZED signed op → os-apply: START … layer=docker:… authority=signed UNDO → signedjobs: signed op COMPLETED. A REFUSED: R3 means the signature, host, time window, nonce or package list did not match; R15 means live-restore is off.
  • The hub: a docker-layer report, outcome applied, healthy, undo: true, the old engine.
  • The same container ids before and after (the health rule checks it; a changed id is health_failed).

Afterwards

On a ring-0 box the next night step installs the newest set again — switch the box's updates OFF on the System page first if that set is the bad one, and record it in the incident's register row. A ring-1 box takes a set only by a signed job, so it stays on the undone set until someone signs the next one.