Files
felhom.eu/documentation/audits/os-updates-spike-2026-10-04/partG/SUMMARY.md
T

4.5 KiB
Raw Blame History

Part G — a guest update end to end, on scratch guest 9202 (demo-hp), 2026-10-04 07:06–07:29 UTC

Venue facts: 9202's disks are on nvme-scratch (a dir storage, raw images) — pct snapshot refuses: "snapshot feature is not available". Customer guests (9201 on both demo boxes) are on local-lvm (LVM-thin) and can snapshot. So the undo here was a whole-guest backup (vzdump --mode snapshot fell back to suspend: 98 s, 1.99 GB) and a pct restore --force. A snapshot rollback on LVM-thin was NOT measured (no fenced venue for it this session).

Step Result Evidence
G1 Debian update (49 packages, Docker CE excluded, --force-confold --force-confdef, non-interactive) rc 0, 24.0 s, Fetched 38.1 MB; rootfs +52 MB with the cache, −81 MB after apt-get clean vs before; no config-file prompt or conflict G1.txt, g-evidence/apt-run.log
What restarted maintainer scripts restarted postfix, systemd-journald, systemd-networkd (MainPID changed) G1-after.txt, g-svc-*.txt
What still needs a restart processes mapping deleted libraries after libc6 u3→u4: dockerd, containerd, docker-proxy ×4, sshd, dbus-daemon, systemd-logind, cron, dhclient, agetty (postgres/celery inside containers map deleted files from before — container-internal, unrelated) G1-after.txt
Containers during the run 13 samples, 2 s apart: all 6 containers up the whole time; StartedAt unchanged; controller + every app healthy g-sampler-G1.txt
G2 undo (restore the pre-update backup) stop 5 s, restore 30 s, start 3 s, all healthy 35 s after start; 73 s total downtime; libc6 back to u3 G2.txt, g-restore.log
G3 half-done (killed after 15 Unpacking, 9 Setting up) 5 packages iU (unpacked, not configured), 4 it (trigger pending: debianutils, libc-bin, man-db, systemd). It does NOT recover by itself: the next ordinary apt-get install refuses (E: Unmet dependencies. Try 'apt --fix-broken install'). Containers stayed up (6 running). G3a.txt, g-evidence/apt-kill.log
G3 repair dpkg --configure -a rc 0 1.4 s; apt-get -f install rc 0 3.9 s (5 upgraded); dpkg --audit empty; the remaining 29 applied in 14.1 s; all healthy, nothing restarted G3b.txt, g-evidence/repair*.log
G4 r1 Docker 29.8.0→29.7.2 (setup), no live-restore apt 17.0 s; all 6 containers restarted; paperless front door no answer 30.0 s; controller no answer 3.0 s; all healthy +41.9 s G4-r1.txt, g-evidence/g4-probe-r1-*.txt
G4 r1 step 29.7.2→29.8.0, no live-restore apt 14.6 s; 6 of 6 restarted; paperless no answer 26.5 s; controller 3.0 s; all healthy +44.9 s same
G4 r2 step 29.7.2→29.8.0, live-restore on (set by systemctl reload docker, which DOES enable it) apt 6.6 s; 0 of 6 restarted; no gap at either front door; the controller logged docker ps / docker inspect / docker stats failures for the seconds dockerd was down, and raised no app event G4-r2.txt
G4 r3 step to the golden's 29.8.2 and containerd 2.3.5→2.3.6, live-restore on apt 9.2 s; 0 of 6 restarted; no gap, even across the containerd restart G4-r3.txt
Turning live-restore OFF again systemctl reload docker with the baked daemon.json did not switch it off (docker info still true). A systemctl restart docker did — and the new dockerd (live-restore off) stopped every container still running from the old one (Exited (0)) and restarted none of them, although all are unless-stopped (Removing stale sandbox … isRestore=false). Nothing brought them back for 3.5 min (the host agent supervises 9201 only). felhom-controller-bootstrap.service restarted the controller; its boot sweep started traefik and filebrowser but HELD paperless-ngx: drive /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx is not a live mountpoint — that app's app.yaml HDD_PATH names the app's user folder, not the drive (the drive itself IS a mountpoint). Started by hand. G4-liverestore-off-restart-dockerd.log
End state Docker 29.8.2 / containerd 2.3.6 (= golden 0.290.0), daemon.json byte-identical to the baked one (cmp), live-restore false, Debian fully updated (docker-compose-plugin 5.5.1→5.6.0 left pending), all 6 containers healthy —

Measurement note: the first G4 probe counted any non-200 as down; paperless answers 302 and the controller 404 (no Host header) when up. Re-counted with no answer (000) = down from the raw probe files; the table uses the re-count.