From 5ab159521dc7b4d6df94ccb3ebd9f7eb67589f82 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Sat, 13 Jun 2026 14:46:20 +0200 Subject: [PATCH] spike: OS/Docker-data storage-split findings (Part B, build-nothing) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Findings report (REPORT-storage-split-spike.md) gating the provisioning spec. Proven on throwaway LXC 9300 (destroyed): data-root on a second local-lvm mountpoint works on an unprivileged LXC (overlayfs/ext4, no idmap, reboot-survives); copy-not-move migration is safe. Key finding: extra CT mountpoints are EXCLUDED from vzdump by default — need backup=1 + a CT restart — so the docker-data mount must be attached with backup=1 (bringup.go:313 omits it today). No agent code changed. Co-Authored-By: Claude Opus 4.8 (1M context) --- CHANGELOG.md | 13 +++++ REPORT-storage-split-spike.md | 90 +++++++++++++++++++++++++++++++++++ 2 files changed, 103 insertions(+) create mode 100644 REPORT-storage-split-spike.md diff --git a/CHANGELOG.md b/CHANGELOG.md index 0db17ff..d121187 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -3,6 +3,19 @@ All notable changes to **felhom-agent** are recorded here. Update on every code change that gets pushed. +## (no version) — storage OS/data-split spike findings (2026-06-13) + +Investigation only — **no code changed**. Findings report: `REPORT-storage-split-spike.md` (gates the +provisioning spec for splitting the controller guest's OS rootfs from its Docker/data onto separate +`local-lvm` volumes). Proven on a throwaway unprivileged LXC (9300, since destroyed): Docker `data-root` +on a second `local-lvm` mountpoint works (overlayfs/ext4, no idmap issue, reboot-survives); the +move-then-verify migration is safe (copy-not-move). **Key finding:** additional LXC mountpoints are +**excluded from vzdump by default** — they need `backup=1` set **and a CT restart** — so the docker-data +mount must be attached with `,backup=1` or named-volume DBs silently fall out of PBS. The exact seam is +`internal/reconcile/bringup.go:313` (`buildConfigParams`), which today builds `mpN` without a `backup=` +flag; `GuestMount` should carry the flag. Per-customer sizes belong in the slice-10 hub storage manifest +(marked at `bringup.go:49-50`); the golden rootfs is hardcoded `8` at `configs/build-golden.sh:40`. + ## v0.28.0 — backup re-target → felhom-pbs (offsite DR) + operator-signed decommission (2026-06-12) **Whole-guest backup now defaults to the offsite PBS tier (real DR).** `BackupConfig.BackupTarget()` diff --git a/REPORT-storage-split-spike.md b/REPORT-storage-split-spike.md new file mode 100644 index 0000000..7e47281 --- /dev/null +++ b/REPORT-storage-split-spike.md @@ -0,0 +1,90 @@ +# Findings — OS/Docker-data storage-split spike (Part B) + +**Date:** 2026-06-13 · **Host:** demo-felhom (PVE 9.2.2) · **Throwaway:** unprivileged LXC 9300 (created, exercised, destroyed) · **Builds nothing** — this report gates the provisioning spec. + +**Goal:** prove the mechanics of splitting the controller guest's **OS rootfs** from its **Docker/data** onto separate `local-lvm` volumes, and the live-migration path, before writing the provisioning spec. Target shape to validate: a small OS rootfs (~32 GB default) + a separate large Docker-data volume (~256 GB), sizes configurable at install. + +**Method:** all risky steps were run on a throwaway unprivileged LXC (9300) replicating guest 9201's config (`unprivileged: 1`, `features nesting=1,keyctl=1`, rootfs on `local-lvm`), never on 9201's live data. Docker 29.5.3 — **identical to production 9201** (verified: both `overlayfs` driver + `io.containerd.snapshotter.v1`). + +--- + +## Verdict summary + +| Item | Verdict | +|---|---| +| B1 — second volume + Docker `data-root` on unprivileged LXC | ✅ **Works.** overlayfs on the ext4 `local-lvm` mountpoint; no idmap/ownership problem; survives reboot. | +| B2 — live move-then-verify migration | ✅ **Works**, copy-not-move is the safety net. Caveat: rsync must exist + its exit code must be checked. | +| B3 — vzdump/PBS coverage of the new mountpoint | ⚠️ **NOT by default — contradicts the spec's premise.** Extra CT mountpoints are **excluded** from vzdump unless `backup=1` is set **and the CT restarted**. | +| B4 — live resize vs fresh-install carving | ✅ Live grow is online/non-disruptive (`pct resize`). Fresh-install carving = installer `hdsize`/`maxroot`/`maxvz` (host-level, documented below). | +| B5 — provisioning seams | ✅ Mapped: golden rootfs size + `BringUpSpec.RootfsGrowGB`/`Mounts`; `bringup.go:313` is the exact spot that today omits `backup=1`. | + +--- + +## B1 — Second volume + Docker `data-root` on an unprivileged LXC ✅ + +- Added a second `local-lvm` mountpoint: `pct set 9300 -mp0 local-lvm:16,mp=/mnt/dockerdata`. PVE created `vm-9300-disk-1`, formatted it **ext4**, mounted it; in-container it is `root:root 0755` — **no idmap/ownership fix needed** (PVE maps the unprivileged offset for a fresh volume automatically). +- Pointed Docker at it via `/etc/docker/daemon.json` `{"data-root":"/mnt/dockerdata"}` → after `systemctl restart docker`, `docker info` reports `Docker Root Dir: /mnt/dockerdata`, `Storage Driver: overlayfs`. +- **overlay works on the mountpoint's filesystem** (ext4): pulled an image + created a named volume + wrote to both the volume and the container's overlay layer. Image/volume bytes landed on `/mnt/dockerdata` (228K → 64M) while the rootfs `/var/lib/docker` stayed flat (232K). +- **Survives reboot:** after `pct reboot`, `data-root` is still `/mnt/dockerdata`, the volume's marker file persists, the image is present. + +**Gotcha (Docker 29 / containerd-snapshotter):** the daemon already auto-starts at boot, so `systemctl start docker` is a no-op and won't pick up a freshly-written `daemon.json` — use `systemctl restart docker`. Also note the driver is the new **`overlayfs`** (containerd-snapshotter), not legacy `overlay2` — relevant to B2. + +## B2 — Live migration sequence (move-then-verify) ✅ + +Proven on the throwaway with seeded **running** apps (postgres with a known row `migrate-survive` in a named volume + nginx). The safe sequence: + +1. `docker stop ` 2. `systemctl stop docker docker.socket` 3. **copy, not move:** `rsync -aHAX --numeric-ids /var/lib/docker/ /mnt/dockerdata/` 4. write `data-root` → `/mnt/dockerdata` in `daemon.json` 5. `systemctl start docker` 6. **verify** images + `docker volume ls` + start apps + query the DB row 7. **only then** reclaim the old location (`rm -rf /var/lib/docker/*`). + +Result: post-switch the `pgdata` volume was present, both apps booted, and the DB row was intact. Reclaiming the old path while apps ran on the new `data-root` did not disturb them — the new `data-root` is fully self-sufficient. + +**Two hard lessons (both nearly caused a silent false-positive):** +- **`rsync` was not installed** on the fresh guest; the first attempt's `rsync … | grep …` swallowed the `127` exit (pipe returns grep's status, `set -e` doesn't fire), so the migration "completed" having copied **nothing**. The migration script MUST install rsync and **check rsync's own exit code**, not a piped one. +- With **containerd-snapshotter** the image content store sits under `data-root/containerd…`; switching `data-root` to an empty volume made images *appear* present (they resolve from the snapshotter's own addressing) while volumes/containers were empty — a misleading partial state. The reliable approach is to **rsync the entire `data-root` with docker stopped** (captures the containerd store + volumes + containers) and verify functional integrity (volume + DB row + boot), or treat images as re-pullable and migrate only the volumes (the irreplaceable data). The copy-not-move discipline meant the source was always intact and recoverable — I reverted `data-root` and the original apps + DB row came straight back. + +## B3 — Backup coverage after the split ⚠️ (the load-bearing finding) + +**The spec's premise is wrong.** It assumed the new mountpoint "should be [captured], as a CT mountpoint unless `backup=0`." The opposite is true on PVE 9.2.2: + +``` +# mp0 with NO backup flag: +INFO: excluding volume mount point mp0 ('/mnt/dockerdata') from backup (disabled) +``` + +Additional LXC mountpoints are **excluded from vzdump by default**. They are included **only** when `backup=1` is set on the mountpoint. And setting it on a running CT is a **pending change** (`[pve:pending]`) — vzdump kept excluding mp0 until the **CT was restarted**, after which: + +``` +# mp0 with backup=1, post-restart: +INFO: including mount point mp0 ('/mnt/dockerdata') in backup +``` + +Confirmed end-to-end: the resulting archive contains `./mnt/dockerdata/volumes/pgdata/_data/…` — the DB data is in PBS/vzdump only once `backup=1` is active. + +**Implication:** if the OS/data split lands without `backup=1` on the docker-data mountpoint, **every named-volume database silently falls out of the PBS whole-guest snapshot** — exactly the class of silent-default trap the spike exists to catch (cf. the PBS `ignore-verified` default). This is the single must-fix for the provisioning spec. + +Honest note (as the spec asked): because **both** volumes stay in PBS once `backup=1` is set, the backup **size does not shrink** — the win of the split is independent sizing/growth and a smaller, faster-to-restore OS rootfs, **not** a smaller backup. (Excluding images from PBS would shrink it but would drop DB coverage — out of scope; not done.) + +## B4 — Resize paths: live vs fresh install ✅ + +- **Live (within the existing 349 GB `local-lvm`), non-disruptive:** `pct resize 9300 rootfs +2G` and `pct resize 9300 mp0 +4G` both grew the volume **and** the in-guest ext4 online (7.8→9.8 GB rootfs, 16→20 GB docker-data) with the CT running and the DB still queryable — no reboot, no downtime. Adding a brand-new mountpoint (`pct set -mpN`) is also non-disruptive to the rootfs, but the **mount itself only appears after a CT restart** on an unprivileged guest (same pending-activation behaviour seen with `backup=1` and with drive binds — consistent with the agent's existing "activate at next boot" note). Grow-only is safe; shrinking LVM-thin volumes is not supported online and was not attempted. The host root (`local`) was **not** repartitioned on the live system. +- **Fresh install (host-level carving — this is where "configurable at install" lives):** the Proxmox ISO installer's *Advanced LVM options* expose `hdsize` (total of the disk PVE uses), `maxroot` (cap on the host `/` = `pve/root`), `maxvz` (cap on the `data` thin pool = `local-lvm`; set to 0 / leave headroom to keep space unallocated), `minfree` (reserved free space in the VG), and `swapsize`. These decide how the SSD is split between a small **host** root and a large **`local-lvm` thin pool**. Note the distinction the spec's target shape blurs: those installer knobs size the **host's** local-vs-thinpool; the **guest's** 32 GB-rootfs + 256 GB-docker-data split is sized at *provision* time via `pct` (rootfs size + an additive mountpoint), carved out of the thin pool — not by the installer. On the demo, the thin pool already has ~350 GB free, so per-guest carving needs no host repartition. + +## B5 — Provisioning integration points ✅ + +Where the guest is sized today, and where the split + configurable sizes slot in: + +- **Golden rootfs size — `configs/build-golden.sh:40`:** `--rootfs "${ROOTFS_STORAGE}:8"` — the OS rootfs is **hardcoded at 8 GB**. The golden also `docker pull`s the controller + app images into `/var/lib/docker` **on the rootfs** (lines ~71–92), so the baked-image set is bounded by this size. `ROOTFS_STORAGE` is the 3rd script arg (default `local-lvm`). The golden runs the same docker-ce as production (→ 29.5.3 overlayfs/containerd-snapshotter). +- **Provision bring-up — `internal/reconcile/bringup.go`:** `BringUpSpec` already carries the right seams: `RestoreStorage` (rootfs target storage), `RootfsGrowGB` (grow-only rootfs resize, applied as its own `ResizeLXC` call at line ~205), and `Mounts []GuestMount{Storage,SizeGB,MountPoint}` (additive `mpN`, attached in `buildConfigParams` at line ~313). So a docker-data volume is a natural `GuestMount`, and the OS rootfs can be grown per-customer via `RootfsGrowGB`. +- **🔴 Exact code spot for the B3 fix — `bringup.go:313`:** today it builds `fmt.Sprintf("%s:%d,mp=%s", m.Storage, m.SizeGB, m.MountPoint)` — **no `backup=` flag**, so any additive mount (including a future docker-data volume) is created `backup=0` = **excluded from PBS**. The docker-data mount MUST be attached with `,backup=1` (and the provision flow must account for the restart-to-activate behaviour). `GuestMount` should gain a `Backup bool` (or always-on for the data mount). +- **Where the size config comes from:** currently caller-provided — the `--selftest=provision`/bring-up path and `cfg.Backup.RestoreStorage`. The in-code marker (`bringup.go:49-50`) states slice 10 wires the **hub storage manifest** into `Mounts`; that is the right home for per-customer sizes, consistent with bootstrap-v2 (the controller already pulls customer-scoped config from the hub). So "configurable at install" = a per-customer field in the hub desired-state → `RootfsGrowGB` + `GuestMount.SizeGB`. + +--- + +## Open questions for the provisioning spec to resolve + +1. **Where does `data-root` get redirected for a *fresh* guest?** The golden bakes images into `/var/lib/docker` on the rootfs. Mounting an empty data volume **over** `/var/lib/docker` hides the baked images. Options: (a) bake `daemon.json` `data-root=/mnt/dockerdata` into the golden and attach the (empty) data mount before first Docker start, re-pulling images on first boot; (b) bake the split into the golden itself (data volume present at golden-build time, images baked onto it); (c) first-boot migration step. Pick one; (b) keeps the baked-image fast-start, (a) is simplest but loses it. +2. **`backup=1` is mandatory on the docker-data mount** (B3). Encode it at `bringup.go:313` and decide whether the OS rootfs alone (no app data) is worth a separate lighter backup cadence. +3. **Restart-to-activate** (B3/B4): attaching a mount or flipping `backup=1` on a running unprivileged guest is pending until reboot. Provisioning attaches mounts *before* first start, so this is a non-issue at provision time, but any *post-hoc* migration of an existing customer guest must schedule a reboot. +4. **Default sizes & source:** confirm 32 GB rootfs / 256 GB docker-data defaults, and wire them through the hub storage manifest (per-customer), sized to the actual SSD via the installer carving facts in B4. +5. **Backup size does not shrink** (B3): set expectations — the split is for independent sizing/growth and a smaller OS-rootfs restore, not a smaller PBS footprint. + +**Nothing was shipped.** Throwaway LXC 9300 and its test archives were destroyed after the spike.