Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
5.4 KiB
REPORT — felhom-agent v0.29.x (OS / Docker-data storage split + lanresolver fix)
Follow-up v0.29.1 (lanresolver): the v0.29.0 re-provision moved 9201's DHCP IP (.151 → .141), but LAN clients kept resolving the old IP. Root cause:
lanresolverupdated the dnsmasq drop-in (address=/<domain>/<ip>) then ransystemctl reload dnsmasq— SIGHUP does not re-read config drop-ins (only/etc/hosts+ cache), so the changedaddress=never took effect. Fixed:reload()→restartDnsmasq()(systemctl restart). Deployed live; host dnsmasq restarted + Pi-hole cache flushed;*.demo-felhom.eunow resolves to .141 and LAN HTTPS returns 200. Future IP moves self-heal on the loop's next tick. (DHCP for the guest is fine — no static IP needed.)
REPORT — felhom-agent v0.29.0 (OS / Docker-data storage split: golden + provision)
Phase 1 of the storage-split slice (Phase 2 = felhom-controller v0.58.0 prevention layer). The
controller guest's OS rootfs and Docker data are carved onto separate local-lvm volumes for
resilience — an isolated OS rootfs stays bootable + agent-recoverable if the Docker volume fills.
Built, the golden re-baked, and live-validated by destroying + re-provisioning guest 9201.
What shipped (v0.29.0)
configs/build-golden.sh— split baked in:--rootfs ${ROOTFS_STORAGE}:${OS_SIZE_GB}(default 32, was hardcoded 8) plus--mp0 ${ROOTFS_STORAGE}:${GOLDEN_DOCKER_GB},mp=/var/lib/docker,backup=1(default 16).daemon.jsonbakes the classic overlay2 driver (features.containerd-snapshotter: false) + log rotation (max-size 10m,max-file 3). Guards: aborts unless/var/lib/dockeris a separate mount, the driver is overlay2, and vzdump includes mp0 (the B3 trap).internal/reconcile/bringup.go—GuestMount.Backupemits,backup=1(closes the spike-B3/B5 silent-DB-loss trap at the mount builder).BringUpSpec.{DataVolGrowGB,DataVolMount}grow the golden-carried Docker-data volume (defaultmp0) online to the per-customer target rather than attaching a fresh empty volume that would shadow the baked images. PlusRootfsGrowGB.- CLI:
--selftest=bring-up|provisiongain-rootfs-grow/-datavol-grow/-datavol-mount. RUNBOOK-provisioning-storage.md— provisioning procedure + fresh-PVE-install thin-pool carving (hdsize/maxroot/maxvz) + the per-customer sizing seam (flags now; hub storage manifest later).- Tests:
buildBringUpConfigbackup=1 emission; bring-up issues rootfs + data-volume resizes.
The critical fix validation caught (overlay2)
Docker 29's default containerd-snapshotter keeps the image store at /var/lib/containerd (on the OS
rootfs), so mounting the data volume at /var/lib/docker only moved named volumes — 1.2 GB of images
stayed on the rootfs (validated live), defeating the split and breaking the controller's statfs("/")
guard. Fix: the classic overlay2 driver stores everything (images + overlay + volumes) under
data-root = the data volume. A start-vs-restart trap (docker-ce auto-starts on install) was also fixed
(daemon.json needs a restart). After the fix, a provisioned guest showed images on the data volume
(/var/lib/docker/overlay2), /var/lib/containerd idle, and a lean OS rootfs.
Live validation
- Re-baked the golden with the split (overlay2): docker OK on overlay2,
/var/lib/dockera separate ext4 mount, vzdump "including mount point mp0 ('/var/lib/docker')", 580 MB archive carrying the images. - Throwaway provision (9301) from the golden: booted in 39 s;
pct configmp0=/var/lib/docker,backup=1,size=256G, rootfs 32G; overlay2; baked images on the data volume; OS isolation — filled the data volume to 100%, the OS rootfs stayed at 4% + writable; log rotation effective. - Destroyed + re-provisioned the live 9201 from the same golden via
--selftest=provision(customerdemo-felhom, retrieval passphrase reused from its old bootstrap;-datavol-grow 240→ 256 GB) + a reboot: bring-up booted in 39 s, the back-half minted a token + attached the bootstrap mp9, and after reboot the controller-bootstrap deployed the baked controller, which pulled its config from the hub. All four infra containers came up from baked images with no pull. Split + overlay2 + lean rootfs confirmed on the real guest; the controller's deploy gate refused (507) on the full data volume. (Phase-2 details in the controller REPORT.)
Provisioning facts confirmed (for the spec/runbook)
- The restore carries all mountpoints (golden mp0 travels in);
freeMountSlotpicks the lowest free slot for a later USB enroll, so mp0 = Docker-data, mp9 = bootstrap, USB = mp1+ — no collision. - Attaching/flag-changing a mountpoint on a running unprivileged guest is pending until reboot — the provision flow's post-provision reboot activates the bootstrap mp9 (and would activate a USB bind).
backup=1on the Docker-data mp is mandatory twice over: keeps named-volume DBs in PBS, AND captures the volume (with baked images) into the golden archive.
Outstanding (demo restoration, not slice validation)
RomM (the HDD app) + USB re-enroll: its data is safe on the host USB (/mnt/felhom-usb/felhom-data,
untouched). Restoring it is the slice-10 enroll flow (assign → guest-attach → reboot → register storage
→ deploy); with the split it binds to mp1+ (mp0 is the Docker-data volume). Documented as the final
restore step.