Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
3.6 KiB
Undo a failed OS update on the customer guest — restore the whole-guest backup in place (decision 81, R-842 option A)
When: the hub mailed
os_update_health_failedfor the guest layer, or the box's apps broke right after a guest OS pass, and installing the previous versions one by one is not enough. Who: the operator, or CC with the operator's word, as root on the box's Proxmox host. Owner design:architecture/11-os-updates.md§5.6 (guest row),09§3 decision 81 (option A, by hand; nothing new built). Proved on demo-felhom 2026-10-04 night: 48 s down, every app healthy, the package versions back to the "before" ones — evidenceaudits/night-2026-10-04/undo/.
What it costs the household: every app is down for under a minute (measured 48 s), and everything written to
the guest after the backup is lost — app data in volumes and databases since that backup. Data on the household's
drive (/mnt/felhom-drives, a bind mount) is NOT in the whole-guest backup and is NOT touched by the restore. So use it
soon after the bad pass — the OS leg runs right after the night backup, so last night's backup is minutes older than the
bad pass. After that, prefer putting single packages back (os-updates-host-undo.md, the same snapshot.debian.org
method works inside the guest with pct exec).
There is no product route: the agent has no executor for a restore over a live guest (restore_overwrite is
classified, not built) and RUNBOOK-manual-guest-restore.md is for a REPLACED host. This page is the route.
1. Pick the archive
The newest whole-guest backup taken BEFORE the bad pass. The pass runs right after the primary backup, so it is normally the newest one on the box's primary backup tier:
pvesm list <primary-tier-storage> | grep -- "-9201-" | tail -3 # e.g. felhom-backup, local
journalctl -u felhom-agent --since -1d | grep -E "backup: completed|osupdate: (START|DONE)" # the order, with times
Check the archive's config BEFORE restoring (the binds must be this host's — they are, on the same box):
tar -xOf /var/lib/vz/dump/<archive>.tar.zst ./etc/vzdump/pct.conf | grep -E '^(rootfs|mp[0-9]|onboot|hookscript)'
2. Restore in place (this is the downtime)
date -u +%T # down from
pct shutdown 9201 --timeout 60
pct restore 9201 <storage>:backup/<archive>.tar.zst --force --storage local-lvm
pct config 9201 | grep -E '^(rootfs|mp[0-9]|onboot|hookscript|lock)' # mp8 + mp9 binds, onboot 1, the hook, no lock
pct start 9201
--forcereplaces the running guest's volumes; the new volumes get NEW names (vm-9201-disk-2/3instead ofdisk-0/1, measured) — nothing on the box pinned the old names.- While the guest is locked by the restore, the agent's guest-power watchdog leaves it alone (measured:
guest is stopped but LOCKED — leaving it to the stale-lock path). --storagemust be the storage the guest's volumes were on (local-lvmon the appliances; checkpct configfirst).
3. Read back (all must hold)
pct exec 9201 -- docker ps --format "{{.Names}} {{.Status}}" # every app "(healthy)"; measured 15 s after start
pct exec 9201 -- dpkg-query -W <the packages the bad pass changed> # the BEFORE versions
The hub: the box's next host report (≤ 15 min) arrives; the household's timeline shows "Controller elindult" — nothing else (measured: no mail).
4. Before the next night
The next night's OS leg would install the same bad versions again. Switch the box's OS updates OFF on the hub's System page (or move it out of ring 0) until a fixed release exists, and say so in the box's event log.