# Undo a failed OS update on the customer guest — restore the whole-guest backup in place (decision 81, R-842 option A) > **When:** the hub mailed `os_update_health_failed` for the **guest** layer, or the box's apps broke right after a > guest OS pass, and installing the previous versions one by one is not enough. > **Who:** the operator, or CC with the operator's word, as root on the box's Proxmox host. > **Owner design:** `architecture/11-os-updates.md` §5.6 (guest row), `09` §3 decision 81 (option A, by hand; nothing > new built). **Proved** on demo-felhom 2026-10-04 night: **48 s down**, every app healthy, the package versions back > to the "before" ones — evidence `audits/night-2026-10-04/undo/`. **What it costs the household:** every app is down for under a minute (measured 48 s), and **everything written to the guest after the backup is lost** — app data in volumes and databases since that backup. Data on the household's drive (`/mnt/felhom-drives`, a bind mount) is NOT in the whole-guest backup and is NOT touched by the restore. So use it soon after the bad pass — the OS leg runs right after the night backup, so last night's backup is minutes older than the bad pass. After that, prefer putting single packages back (`os-updates-host-undo.md`, the same `snapshot.debian.org` method works inside the guest with `pct exec`). There is no product route: the agent has no executor for a restore over a live guest (`restore_overwrite` is classified, not built) and `RUNBOOK-manual-guest-restore.md` is for a REPLACED host. This page is the route. ## 1. Pick the archive The newest whole-guest backup taken BEFORE the bad pass. The pass runs right after the primary backup, so it is normally the newest one on the box's primary backup tier: ```bash pvesm list | grep -- "-9201-" | tail -3 # e.g. felhom-backup, local journalctl -u felhom-agent --since -1d | grep -E "backup: completed|osupdate: (START|DONE)" # the order, with times ``` Check the archive's config BEFORE restoring (the binds must be this host's — they are, on the same box): ```bash tar -xOf /var/lib/vz/dump/.tar.zst ./etc/vzdump/pct.conf | grep -E '^(rootfs|mp[0-9]|onboot|hookscript)' ``` ## 2. Restore in place (this is the downtime) ```bash date -u +%T # down from pct shutdown 9201 --timeout 60 pct restore 9201 :backup/.tar.zst --force --storage local-lvm pct config 9201 | grep -E '^(rootfs|mp[0-9]|onboot|hookscript|lock)' # mp8 + mp9 binds, onboot 1, the hook, no lock pct start 9201 ``` - `--force` replaces the running guest's volumes; the new volumes get NEW names (`vm-9201-disk-2/3` instead of `disk-0/1`, measured) — nothing on the box pinned the old names. - While the guest is locked by the restore, the agent's guest-power watchdog leaves it alone (measured: `guest is stopped but LOCKED — leaving it to the stale-lock path`). - `--storage` must be the storage the guest's volumes were on (`local-lvm` on the appliances; check `pct config` first). ## 3. Read back (all must hold) ```bash pct exec 9201 -- docker ps --format "{{.Names}} {{.Status}}" # every app "(healthy)"; measured 15 s after start pct exec 9201 -- dpkg-query -W # the BEFORE versions ``` The hub: the box's next host report (≤ 15 min) arrives; the household's timeline shows "Controller elindult" — nothing else (measured: no mail). ## 4. Before the next night The next night's OS leg would install the same bad versions again. Switch the box's OS updates OFF on the hub's System page (or move it out of ring 0) until a fixed release exists, and say so in the box's event log.