night 2026-10-04: the guest undo runbook (proven 48 s), R-866, A3 in progress
gates / gates (push) Successful in 30s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-04 22:37:25 +02:00
parent a1a32bbd74
commit 76ff0654ea
4 changed files with 73 additions and 1 deletions
@@ -0,0 +1,65 @@
# Undo a failed OS update on the customer guest — restore the whole-guest backup in place (decision 81, R-842 option A)
> **When:** the hub mailed `os_update_health_failed` for the **guest** layer, or the box's apps broke right after a
> guest OS pass, and installing the previous versions one by one is not enough.
> **Who:** the operator, or CC with the operator's word, as root on the box's Proxmox host.
> **Owner design:** `architecture/11-os-updates.md` §5.6 (guest row), `09` §3 decision 81 (option A, by hand; nothing
> new built). **Proved** on demo-felhom 2026-10-04 night: **48 s down**, every app healthy, the package versions back
> to the "before" ones — evidence `audits/night-2026-10-04/undo/`.
**What it costs the household:** every app is down for under a minute (measured 48 s), and **everything written to
the guest after the backup is lost** — app data in volumes and databases since that backup. Data on the household's
drive (`/mnt/felhom-drives`, a bind mount) is NOT in the whole-guest backup and is NOT touched by the restore. So use it
soon after the bad pass — the OS leg runs right after the night backup, so last night's backup is minutes older than the
bad pass. After that, prefer putting single packages back (`os-updates-host-undo.md`, the same `snapshot.debian.org`
method works inside the guest with `pct exec`).
There is no product route: the agent has no executor for a restore over a live guest (`restore_overwrite` is
classified, not built) and `RUNBOOK-manual-guest-restore.md` is for a REPLACED host. This page is the route.
## 1. Pick the archive
The newest whole-guest backup taken BEFORE the bad pass. The pass runs right after the primary backup, so it is
normally the newest one on the box's primary backup tier:
```bash
pvesm list <primary-tier-storage> | grep -- "-9201-" | tail -3 # e.g. felhom-backup, local
journalctl -u felhom-agent --since -1d | grep -E "backup: completed|osupdate: (START|DONE)" # the order, with times
```
Check the archive's config BEFORE restoring (the binds must be this host's — they are, on the same box):
```bash
tar -xOf /var/lib/vz/dump/<archive>.tar.zst ./etc/vzdump/pct.conf | grep -E '^(rootfs|mp[0-9]|onboot|hookscript)'
```
## 2. Restore in place (this is the downtime)
```bash
date -u +%T # down from
pct shutdown 9201 --timeout 60
pct restore 9201 <storage>:backup/<archive>.tar.zst --force --storage local-lvm
pct config 9201 | grep -E '^(rootfs|mp[0-9]|onboot|hookscript|lock)' # mp8 + mp9 binds, onboot 1, the hook, no lock
pct start 9201
```
- `--force` replaces the running guest's volumes; the new volumes get NEW names (`vm-9201-disk-2/3` instead of
`disk-0/1`, measured) — nothing on the box pinned the old names.
- While the guest is locked by the restore, the agent's guest-power watchdog leaves it alone (measured:
`guest is stopped but LOCKED — leaving it to the stale-lock path`).
- `--storage` must be the storage the guest's volumes were on (`local-lvm` on the appliances; check `pct config` first).
## 3. Read back (all must hold)
```bash
pct exec 9201 -- docker ps --format "{{.Names}} {{.Status}}" # every app "(healthy)"; measured 15 s after start
pct exec 9201 -- dpkg-query -W <the packages the bad pass changed> # the BEFORE versions
```
The hub: the box's next host report (≤ 15 min) arrives; the household's timeline shows "Controller elindult" — nothing
else (measured: no mail).
## 4. Before the next night
The next night's OS leg would install the same bad versions again. Switch the box's OS updates OFF on the hub's System
page (or move it out of ring 0) until a fixed release exists, and say so in the box's event log.
@@ -78,4 +78,4 @@ installs the new version again. Write the hold into the register row that tracks
- A kernel: the fast lane never installs one (R14). The kernel lane is `11` §5.6 (not built).
- A Proxmox-origin package (pve-*, qemu-server, …): never installed by the fast lane (R12/origin rule). Proxmox's
own repository keeps old versions; that undo is not written yet.
- The customer guest: its undo is last night's whole-guest backup (decision 81).
- The customer guest: its undo is last night's whole-guest backup (decision 81) — `os-updates-guest-undo.md` (proven 2026-10-04 night, 48 s down).