Evidence-based: pct set won't hot-apply to a running guest; /proc/pid/root inject blocked by mount-locking; nsenter -m loses the host source. Decision: enroll persists (no forced reboot) + user-triggered batched restart button + P3 flags "restart to reconnect". Staging-mp deferred. Remaining build noted in REPORT. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
4.3 KiB
REPORT — slice 10: external user-data drive passthrough (P1 spike + P2)
Agent v0.25.0 (+ controller v0.48.0, golden rebaked). P1 spike PASSED (gate); P2 BUILT + validated live on guest 9201. P3 (self-heal) + P4 (dual-role) are the next phases.
PHASE 1 — SPIKE (GATE PASSED), four proofs on 9201
- 1A host→guest bind:
pct set 9201 -mp0 /mnt/felhom-usb,mp=/mnt/felhom-usb— bind form (host path), neverstorage:size. Propagationshared:49host↔guest automatic. - 1B write — chown, not idmap: idmap not clean (mixed host ownership 1000+0, container-wide,
restart, subuid). Decision (refined with Viktor): chown only a fresh
<drive>/felhom-datanamespace to100000:100000; the customer's existing data is never touched. Guest-root r+w confirmed. - 1C guest→controller-container:
-v /mnt:/mnt:rslave+/mntrshared in the guest → newly-mounted drives propagate into the running container (proven). The de-priv/mntgap, reopened scoped. - 1D app-container: busybox bind read+write; bytes land on
/dev/sdb1.
PHASE 2 — passthrough (Model A: the felhom-data namespace is the in-guest mount)
- P2A agent —
POST /disks/guest-attach(internal/localapi,GuestBinder): self-scoped; creates<drive>/felhom-data, chowns it to the guest base (not -R), andpct set <vmid> -mpN <drive>/felhom-data,mp=/mnt/<name>(RW bind). Idempotent; lowest freempN;wherevalidated. Only Felhom's namespace crosses into the guest — the customer's other on-drive data never does. Tests:TestGuestAttach_*(slot select, idempotency, bad-path, not-configured). - P2B golden —
configs/build-golden.sh: controllerdocker rungains-v /mnt:/mnt:rslave; the bootstrap makes/mntrshared first. Scoped to/mnt(only felhom-data-namespace mounts). - P2C controller (v0.48.0):
agentapi.GuestAttach;runStorageInit/runStorageAttach/handleStorageRegistercallattachIntoGuestafter register (best-effort; P3 heals a miss).
Live validation (9201)
After guest-attach + a guest restart to activate mp0:
- mp0 =
/mnt/felhom-usb SOURCE /dev/sdb1[/felhom-data](Model A; the[/felhom-data]suffix the controller's mount strip already handles). - Controller container mountinfo has
/felhom-data /mnt/felhom-usb … /dev/sdb1. - An app (busybox bind) writes
proof.txt→ present on the host/mnt/felhom-usb/felhom-data/...,dfdevice/dev/sdb1(NOT the rootfs). - Banner cleared:
[PASS] Storage paths: 1 connected, 0 disconnected. go test ./...green (both repos).
LIVE-ACTIVATION — investigated, mechanism chosen (evidence-based)
A drive enrolled into a running unprivileged guest cannot be activated live, and the host-side
bind-inject is blocked (proven on 9201): pct set doesn't hot-apply a mountpoint to a running
guest; /proc/<pid>/root/... bind → mount: bad superblock (unprivileged mount-locking); nsenter -m
into the guest ns loses the host source path. The bind activates at the next guest boot (validated:
reboot → mp0 active /dev/sdb1[/felhom-data], controller sees it, banner clears). Fresh guests from the
rebaked golden are unaffected (mp activates at first boot, before the controller starts).
Decision (Viktor): enroll persists via pct set with NO forced reboot; the UI shows a
"pending activation" state + a user-triggered "Újraindítás most (~30s)" button that batches all
pending drives; P3 self-heal flags "restart to reconnect" for recovered drives. The staging-mp
live-propagation alternative is deferred to its own spike. (Build remaining — see below.)
Not done (next build increments)
- Activation-UX: agent self-reboot endpoint (
POST /guest/reboot→pct reboot <vmid>, token- scoped) the controller calls; controller "pending activation" detection (registered + agent shows the drive present/attached, but not a live mount in the container) + a batched "Újraindítás most (~30s)" button. - P3 self-heal watchdog reconcile (4-state new/enrolled/ejected/decommissioned, durable-id-keyed) + safety rails (flapping backoff, durable-id-only, in-progress-op respect, PVE coordination, re-propagate); flags "restart to reconnect".
- P4 dual-role eligibility + backup-aware wipe warning. Cross-drive backup ENGINE stays out of scope.