Files
felhom.eu/REPORT-backup-close-os-spike-2026-10-04.md
T

5.9 KiB
Raw Permalink Blame History

REPORT — off-site topic closed; operating-system update spike — 2026-10-04 (day)

Architecture read: 07-backup-architecture.md (Lane 2, §6.1), _recovery-inventory-2026-07-28.md, 03-host-agent.md, 09 §3/§4, 11-os-updates.md. Baselines (re-verified): felhom.eu d07a1a904cbf (hub 0.128.0), controller 99a149756070 (0.290.0), agent d766666ff8cf (0.138.0), catalog 917a779cca67. Register 328, highest R-834. Rulings recorded first as 09 §3 decisions 75–77 and 11 committed verbatim (3885640). Evidence: documentation/audits/backup-close-2026-10-04/ and documentation/audits/os-updates-spike-2026-10-04/.

The Part table

Part Result Notes
A — restored guest safe by default (R-834) done Routes: restore-test (MEASURED safe: onboot 0 and throwaway mp8/mp9 on every poll), DR bring-up (fixed: refuses beside a live original — agent v0.139.0, refused live on demo-hp, nothing created), hand route (new scripts/felhom-restore-beside.sh, proven live on 9298 then destroyed), provisioning (golden, no binds). 4 tests, 2 red-proofs + 1 built-in. Changed: no sudoers line — the restore-test sets onboot 0 through the API, and DR refuses rather than degrades.
B — clean-up cannot wedge (R-833) done Hub v0.129.0 deployed. 4 red-proofs. Lab proof on a real restic 0.14.0 repo: 98 → 13 under a raised cap, default refused before, normal after. Live: 4 bad grants refused (400). Changed: no valid grant placed on a real customer — a demo box would consume it (brief: lab repo only). No controller change needed.
C — returning household (R-726) done Two options in STATUS; pick A. Nothing built.
D — dated check (R-95) done Due 2026-10-12, four checks named in R-95.
E — where we stand done 5 systems surveyed read-only (throwaway apt indexes).
F — exact version later done madison host + guest; DSA history 3 months; snapshot.debian.org from a throwaway container on 9202.
G — guest update on 9202 done, one deviation Changed: 9202 is on dir storage and cannot snapshot, so the undo was a backup + restore (73 s); the snapshot rollback is unmeasured (R-837). G3 interrupted the update straight after the undo (the "apply again" happened as G3's repair) — the same update could not be interrupted once applied.
H — host update on demo-hp done Debian lane 108 packages, 60 s, guests up. One-package undo: rsync gone, libpng worked. Kernel: two reboots on the operator's word — new kernel, then fallback to old. Proxmox simulated only.
I — design record done 11 corrected (C1–C12), wrapper draft §5.4.1, answers §7.1, sample list (157 packages) simulated on demo-felhom, Q10 price recorded. Two STATUS decisions.

Claims that turned out wrong (named)

  1. "Debian's archives keep only the newest version" (11 §5.3) — they keep two: the point-release one and the newest security one; intermediates are gone (C2).
  2. "Proxmox and Docker keep older ones" — true, measured: 30–66 and 18–46 versions.
  3. "--next-boot falls back by itself" (11 §5.6) — only after a boot that reaches userspace; on GRUB it is an ordinary default; a hang keeps the new kernel (code-read). And installing a kernel alone makes it the default (C4).
  4. "The guest has no live-restore" — true. But live-restore is the answer to Q3, and switching it off again is a trap (C5, R-835).
  5. "The agent may not run apt except for dnsmasq" — it may also install wireguard-tools (C1).
  6. "The restore-test guest is safe today" — TRUE, measured (onboot 0, no host bind). The unsafe routes were the DR bring-up and the hand route.
  7. Also wrong in 11: the slow-lane list by name (40 Proxmox packages have plain names, C3); cloudflared "on the host" (it is a guest container, C8); approving what ring 0 installed (C9); a fast-lane run is "a service restart at most" — libc leaves PID 1 and lxc-start on the old library (C11).

Found and handled in-session

  • My own output filter dropped every line containing "perl" — including "paperless". A false "the app vanished" was caught before acting on it; evidence files were saved unfiltered.
  • pkill -f dpkg-deb killed my own shell during G3 (the known trap); the kill itself had landed, and the state was read in a fresh command.
  • The first Docker probe counted 302/404 answers as down; re-counted from the raw probe files with "no answer" as down.
  • A pgrep waiter matched itself and never ended (the known trap); it was harmless and killed by its timeout.

Rows

Closed: R-833, R-834. Opened: R-835 (live-restore off trap), R-836 (kernel hang keeps new kernel), R-837 (snapshot undo unmeasured), R-838 (cloudflared pinned since June, P2), R-839 (boot sweep held an app whose HDD_PATH names its folder). Narrowed: R-812 (spike done), R-95 (dated check), R-726 (waiting on the operator). Register 328 → 331.

Teardown, three layers

  • Machines: scratch VMIDs 990000 (restore-test, torn down by the agent) and 9298 (destroyed); no 9297 was created. 9202: Debian fully updated, Docker 29.8.2 / containerd 2.3.6 (golden 0.290.0's), daemon.json byte-identical to the baked one, all apps healthy; its backup deleted; the debian:trixie probe image removed. demo-hp host: 108 Debian packages + kernel 7.0.14-20-pve installed (110 changes, partH/H-final-host-packages-after.tsv); running 7.0.2-6-pve, next boot 7.0.14-20-pve, no pins; 78 Proxmox packages still pending. demo-felhom: read only (its agent updated to 0.139.0 by signed job). Helper files removed from every host and guest.
  • Host (DooPlex): the lab restic repo removed; scratch copies of the hub password and the DSA list shredded. Agent 0.139.0 released (tag + package, verified by download), not vouched.
  • Hub: v0.129.0 deployed; no grant left pending; weekly windows unchanged (ON).