night 2026-10-04: the guest undo runbook (proven 48 s), R-866, A3 in progress
gates / gates (push) Successful in 30s
gates / gates (push) Successful in 30s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -0,0 +1,65 @@
|
||||
# Undo a failed OS update on the customer guest — restore the whole-guest backup in place (decision 81, R-842 option A)
|
||||
|
||||
> **When:** the hub mailed `os_update_health_failed` for the **guest** layer, or the box's apps broke right after a
|
||||
> guest OS pass, and installing the previous versions one by one is not enough.
|
||||
> **Who:** the operator, or CC with the operator's word, as root on the box's Proxmox host.
|
||||
> **Owner design:** `architecture/11-os-updates.md` §5.6 (guest row), `09` §3 decision 81 (option A, by hand; nothing
|
||||
> new built). **Proved** on demo-felhom 2026-10-04 night: **48 s down**, every app healthy, the package versions back
|
||||
> to the "before" ones — evidence `audits/night-2026-10-04/undo/`.
|
||||
|
||||
**What it costs the household:** every app is down for under a minute (measured 48 s), and **everything written to
|
||||
the guest after the backup is lost** — app data in volumes and databases since that backup. Data on the household's
|
||||
drive (`/mnt/felhom-drives`, a bind mount) is NOT in the whole-guest backup and is NOT touched by the restore. So use it
|
||||
soon after the bad pass — the OS leg runs right after the night backup, so last night's backup is minutes older than the
|
||||
bad pass. After that, prefer putting single packages back (`os-updates-host-undo.md`, the same `snapshot.debian.org`
|
||||
method works inside the guest with `pct exec`).
|
||||
|
||||
There is no product route: the agent has no executor for a restore over a live guest (`restore_overwrite` is
|
||||
classified, not built) and `RUNBOOK-manual-guest-restore.md` is for a REPLACED host. This page is the route.
|
||||
|
||||
## 1. Pick the archive
|
||||
|
||||
The newest whole-guest backup taken BEFORE the bad pass. The pass runs right after the primary backup, so it is
|
||||
normally the newest one on the box's primary backup tier:
|
||||
|
||||
```bash
|
||||
pvesm list <primary-tier-storage> | grep -- "-9201-" | tail -3 # e.g. felhom-backup, local
|
||||
journalctl -u felhom-agent --since -1d | grep -E "backup: completed|osupdate: (START|DONE)" # the order, with times
|
||||
```
|
||||
|
||||
Check the archive's config BEFORE restoring (the binds must be this host's — they are, on the same box):
|
||||
|
||||
```bash
|
||||
tar -xOf /var/lib/vz/dump/<archive>.tar.zst ./etc/vzdump/pct.conf | grep -E '^(rootfs|mp[0-9]|onboot|hookscript)'
|
||||
```
|
||||
|
||||
## 2. Restore in place (this is the downtime)
|
||||
|
||||
```bash
|
||||
date -u +%T # down from
|
||||
pct shutdown 9201 --timeout 60
|
||||
pct restore 9201 <storage>:backup/<archive>.tar.zst --force --storage local-lvm
|
||||
pct config 9201 | grep -E '^(rootfs|mp[0-9]|onboot|hookscript|lock)' # mp8 + mp9 binds, onboot 1, the hook, no lock
|
||||
pct start 9201
|
||||
```
|
||||
|
||||
- `--force` replaces the running guest's volumes; the new volumes get NEW names (`vm-9201-disk-2/3` instead of
|
||||
`disk-0/1`, measured) — nothing on the box pinned the old names.
|
||||
- While the guest is locked by the restore, the agent's guest-power watchdog leaves it alone (measured:
|
||||
`guest is stopped but LOCKED — leaving it to the stale-lock path`).
|
||||
- `--storage` must be the storage the guest's volumes were on (`local-lvm` on the appliances; check `pct config` first).
|
||||
|
||||
## 3. Read back (all must hold)
|
||||
|
||||
```bash
|
||||
pct exec 9201 -- docker ps --format "{{.Names}} {{.Status}}" # every app "(healthy)"; measured 15 s after start
|
||||
pct exec 9201 -- dpkg-query -W <the packages the bad pass changed> # the BEFORE versions
|
||||
```
|
||||
|
||||
The hub: the box's next host report (≤ 15 min) arrives; the household's timeline shows "Controller elindult" — nothing
|
||||
else (measured: no mail).
|
||||
|
||||
## 4. Before the next night
|
||||
|
||||
The next night's OS leg would install the same bad versions again. Switch the box's OS updates OFF on the hub's System
|
||||
page (or move it out of ring 0) until a fixed release exists, and say so in the box's event log.
|
||||
@@ -78,4 +78,4 @@ installs the new version again. Write the hold into the register row that tracks
|
||||
- A kernel: the fast lane never installs one (R14). The kernel lane is `11` §5.6 (not built).
|
||||
- A Proxmox-origin package (pve-*, qemu-server, …): never installed by the fast lane (R12/origin rule). Proxmox's
|
||||
own repository keeps old versions; that undo is not written yet.
|
||||
- The customer guest: its undo is last night's whole-guest backup (decision 81).
|
||||
- The customer guest: its undo is last night's whole-guest backup (decision 81) — `os-updates-guest-undo.md` (proven 2026-10-04 night, 48 s down).
|
||||
|
||||
Reference in New Issue
Block a user