Files
felhom.eu/documentation/runbooks/os-updates-guest-undo.md
T
2026-10-04 22:37:25 +02:00

3.6 KiB

Undo a failed OS update on the customer guest — restore the whole-guest backup in place (decision 81, R-842 option A)

When: the hub mailed os_update_health_failed for the guest layer, or the box's apps broke right after a guest OS pass, and installing the previous versions one by one is not enough. Who: the operator, or CC with the operator's word, as root on the box's Proxmox host. Owner design: architecture/11-os-updates.md §5.6 (guest row), 09 §3 decision 81 (option A, by hand; nothing new built). Proved on demo-felhom 2026-10-04 night: 48 s down, every app healthy, the package versions back to the "before" ones — evidence audits/night-2026-10-04/undo/.

What it costs the household: every app is down for under a minute (measured 48 s), and everything written to the guest after the backup is lost — app data in volumes and databases since that backup. Data on the household's drive (/mnt/felhom-drives, a bind mount) is NOT in the whole-guest backup and is NOT touched by the restore. So use it soon after the bad pass — the OS leg runs right after the night backup, so last night's backup is minutes older than the bad pass. After that, prefer putting single packages back (os-updates-host-undo.md, the same snapshot.debian.org method works inside the guest with pct exec).

There is no product route: the agent has no executor for a restore over a live guest (restore_overwrite is classified, not built) and RUNBOOK-manual-guest-restore.md is for a REPLACED host. This page is the route.

1. Pick the archive

The newest whole-guest backup taken BEFORE the bad pass. The pass runs right after the primary backup, so it is normally the newest one on the box's primary backup tier:

pvesm list <primary-tier-storage> | grep -- "-9201-" | tail -3     # e.g. felhom-backup, local
journalctl -u felhom-agent --since -1d | grep -E "backup: completed|osupdate: (START|DONE)"   # the order, with times

Check the archive's config BEFORE restoring (the binds must be this host's — they are, on the same box):

tar -xOf /var/lib/vz/dump/<archive>.tar.zst ./etc/vzdump/pct.conf | grep -E '^(rootfs|mp[0-9]|onboot|hookscript)'

2. Restore in place (this is the downtime)

date -u +%T                                          # down from
pct shutdown 9201 --timeout 60
pct restore 9201 <storage>:backup/<archive>.tar.zst --force --storage local-lvm
pct config 9201 | grep -E '^(rootfs|mp[0-9]|onboot|hookscript|lock)'   # mp8 + mp9 binds, onboot 1, the hook, no lock
pct start 9201
  • --force replaces the running guest's volumes; the new volumes get NEW names (vm-9201-disk-2/3 instead of disk-0/1, measured) — nothing on the box pinned the old names.
  • While the guest is locked by the restore, the agent's guest-power watchdog leaves it alone (measured: guest is stopped but LOCKED — leaving it to the stale-lock path).
  • --storage must be the storage the guest's volumes were on (local-lvm on the appliances; check pct config first).

3. Read back (all must hold)

pct exec 9201 -- docker ps --format "{{.Names}} {{.Status}}"   # every app "(healthy)"; measured 15 s after start
pct exec 9201 -- dpkg-query -W <the packages the bad pass changed>   # the BEFORE versions

The hub: the box's next host report (≤ 15 min) arrives; the household's timeline shows "Controller elindult" — nothing else (measured: no mail).

4. Before the next night

The next night's OS leg would install the same bad versions again. Switch the box's OS updates OFF on the hub's System page (or move it out of ring 0) until a fixed release exists, and say so in the box's event log.