Files
felhom.eu/documentation/audits/SPIKE-nas-storage-2026-06-29.md
admin 0719db0d53 spike: NAS network storage (media automount + restic backup target) — findings
Empirical validation on a DooPlex-simulated NAS (SMB+NFS). Verdict READY for the
production network-storage TASK. Key findings: mount must be host-side (unprivileged
LXC blocks NFS/CIFS mount) + bind-propagated into the guest; NFS soft fails-clean +
auto-recovers (SMB hangs-but-contained); UID mapping = container uid + 100000 LXC
offset; ~128MB/s adequate; write integrity intact (atomic-write apps safe on soft);
restic-SFTP backup/restore works + fails-as-job not hang; systemd automount gives
on-demand + idle-unmount. Distinct network-storage class bypasses the drive lifecycle.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-29 11:44:40 +02:00

16 KiB
Raw Permalink Blame History

SPIKE — NAS network storage (media automount + backup target), simulated on DooPlex

Date: 2026-06-29 Class: Spike (empirical validation; output = this doc only). Throwaway probes; nothing merged into agent/controller production packages. Goal: validate the bulk-media network-storage mechanism (storage class 3) — host-side mount of a NAS share → bind into guest 9201 → bind into media container, for both SMB and NFS — before writing the production TASK. Secondary: restic-over-SFTP as a backup target (class 2).

Simulated NAS: real Samba (SMB) + nfs-kernel-server (NFS) already running on DooPlex 192.168.0.180; added an isolated scratch export /srv/nas-sim/media (sample media, ~513 MB incl. a 512 MB file), separate from the live Longhorn-PVC NFS exports and PBS/k3s. Mount probes on the Proxmox host felhom-pve (192.168.0.162) → guest 9201 (live demo guest) → throwaway alpine containers.

Sim-vs-real-NAS caveat (stated up front): this validates the Linux mechanism (kernel mount, propagation, idmap, failure semantics) against a Debian Samba/NFS server. It does not reproduce Synology/QNAP-specific quirks: their SMB dialect ceilings, NFS squash/mapall UIs, per-share ACL models, SMB multichannel, or appliance reboot timing. Those need a confirmation pass on the real appliance before GA.


Verdict: READY to write the production "network storage" TASK (with named caveats)

NFS and SMB both work end-to-end through the host-mount→bind chain for read and write, with a precise UID-mapping recipe; NFS soft mounts achieve clean failure isolation (fail-soft + auto-recover, no guest/df wedge). SMB failure isolation is the one area needing the documented option set (below). Restic-SFTP backup is trivial and failure-isolated. No blockers; see "Caveats / for the production TASK."


Per-Spike-Question findings

Q1 — Mount placement: HOST-side, bind into guest (in-guest mount is kernel-blocked)

Guest 9201 is an unprivileged LXC (unprivileged: 1, features: nesting,keyctl). Mounting NFS or CIFS inside the guest fails:

# pct exec 9201 -- mount -t nfs  ... → mount: permission denied (exit 32)
# pct exec 9201 -- mount -t cifs ... → mount: permission denied (exit 32)
# CapEff: 000001ffffffffff  (full caps *within the userns* — but mounting these fs types needs the INIT userns)

The kernel forbids NFS/CIFS mounts in a user namespace regardless of the in-userns capability set. Therefore the NAS mount must live on the Proxmox host (agent-owned) and be bind-propagated into the guest — exactly mirroring the existing drive-bind architecture.

Propagation works for free via the existing mp8: /mnt/felhom-drives bind: the host path is shared propagation, the guest sees it shared,slave, and existing drive mounts (felhom-flash/felhom-usb) already propagate through. A new host mount under /mnt/felhom-drives/<name> appears live in the guest with no restart — confirmed.

Recommended chain: host: mount NAS → /mnt/felhom-drives/<share> → (auto-propagates) guest: /mnt/felhom-drives/<share>docker -v /mnt/felhom-drives/<share>:/media.

Q2 — SMB and NFS both work through the full chain — YES (read+write)

NFS (mount -t nfs -o vers=4.1,soft,timeo=50,retrans=2,_netdev,noatime): host mount OK → propagated to guest → container read the 512 MB file (md5 verified) and wrote a new file. Effective options included soft,timeo=50,retrans=2.

SMB/CIFS (mount -t cifs -o credentials=…,uid=101000,gid=101000,forceuid,forcegid,file_mode=0664,dir_mode=0775,vers=3.0,_netdev): host mount OK → propagated → container read and wrote over SMB.

Minimal viable option sets:

  • NFS: vers=4.1,soft,timeo=50,retrans=2,_netdev,noatime + the export's all_squash,anonuid=…,anongid=… (see Q5). vers=4.1 avoids the rpcbind/111 + lock-manager surface of v3.
  • SMB: vers=3.0,credentials=<file>,uid=…,gid=…,forceuid,forcegid,file_mode=0664,dir_mode=0775,_netdev (creds file 0600 on the host; never inline). dir_mode/file_mode MUST be plain octal (0775, not 2775).

Q5 — UID / permission mapping: the LXC +100000 offset is the whole story (read+write proven)

The established felhom convention (verified against the live felhom-usb userdata): a dir owned host 100000:101000 appears in the guest as root:1000, mode 2775 (setgid, group-writable); media apps write via gid 1000. So guest/container uid/gid N = host uid/gid N+100000.

A naïve anonuid=1000 export is WRONG — it lands as host-1000, which is outside the guest's 100000165535 idmap and shows as nobody:nogroup (readable because world-readable, but not writable). The fix is to present the share as the host-mapped id:

  • NFS: export …,all_squash,anonuid=101000,anongid=101000 and own the share dir 101000:101000. Then the guest sees 1000:1000, and a uid/gid-1000 container reads existing files and writes new ones; the write lands on the NAS as 101000:101000. (Verified: WRITE_OK, server file -rw-r--r-- 101000 101000.)
  • SMB: mount with uid=101000,gid=101000,forceuid,forcegid (client forces the guest-visible owner to 1000:1000 regardless of server ownership) and the share's force user/force group must resolve to a uid/gid 101000 identity so server-side writes are actually permitted on the 101000-owned dir. (Verified: with force user=nasguest (uid 101000) the container SMB_WRITE_OK; with force user=a uid-1000 identity the write was denied — the mismatch is exactly the trap to document.)

Mapping recipe for the production class: the NAS share for a customer's media must be owned/squashed to container_uid + 100000 (e.g. media app uid 1000 → NAS uses 101000). On a real Synology/QNAP this means an export mapall/anonuid of 101000 (NFS) or an SMB share whose force-user resolves to 101000.

Q3 — Failure isolation (THE risk)

Method: with a media container looping reads, black-holed the NAS from felhom-pve via an iptables DROP on 180 scoped to 192.168.0.162 (DROP, not REJECT → simulates a dead NAS / network black hole; leaves the NFS service up for k3s/Longhorn — the safe way to "disappear" it).

NFS (soft,timeo=50,retrans=2) — clean fail-soft + auto-recover:

  • Uncached stat returned Input/output error after ~16 s (≈ retrans×timeo). It errors, does not hang.
  • df on the mount did not hang (returned rc 0); no reader process in D (uninterruptible) state.
  • The looping reader kept serving from page-cache briefly, then flipped to FAIL once the attribute cache expired and revalidation hit the black-holed server — i.e. fails soft, no wedge.
  • Other apps unaffected: gitea stayed HTTP 200 throughout; the guest stayed healthy (uptime/commands responsive).
  • Recovery on NAS return: after removing the DROP, a fresh read was RECOVERED_READ_OK, and the looping reader auto-recovered on its own (FAIL → PASS) with no remount/intervention. stat worked again.

SMB/CIFS (default options) — hangs, but blast-radius contained:

  • With the NAS black-holed, df on the mount hung the full 40 s timeout (rc 124); an uncached stat hung the full 60 s; the reader's dd went D (uninterruptible) state. CIFS has no soft/hard knob like NFS — a dead server blocks operations until the SMB session/TCP timeout.
  • BUT the blast radius was still contained: gitea stayed HTTP 200, the guest stayed healthy (uptime responsive) — only the process touching the SMB mount wedged.
  • Recovery: on NAS return the wedged dd unblocked on its own (df instant, no D-state, reader PASS again) — no kill/remount needed.

Q3 verdict: prefer NFS soft for the media class — it fails-soft (clean EIO) and never wedges df/the process. SMB hangs the accessing process during the outage (contained, self-healing) and is the fallback only for NAS that speak SMB-only; pair it with automount idle-unmount (below) to minimise the stale-mount window. The WRONG outcome (a hard mount wedging the guest or df box-wide) was not observed in either case — the unprivileged-LXC bind boundary kept the blast radius to the accessing container.

Q4 — Write reliability — integrity perfect; an interrupted in-flight write leaves a partial file

  • Clean write integrity: a container (uid 1000) wrote a 100 MB file through the chain; md5 on the host-mount side == md5 read directly on the NAS server (2541a840…) — byte-intact end to end.
  • Disappear mid-write (write-loop of 1 MB fsync'd files, NAS black-holed at file ~221): every completed file before the cut was byte-intact (exactly 1 MB); the single in-flight file (w221.bin) was left partial/truncated on the NAS; the writer's dd went D-state (writes-in-flight soft-fail slower than reads — soft eventually returns EIO, but the in-flight file is already partially persisted).
  • Safe-option note: soft is correct for read-mostly media, but for write apps a soft-interrupted write can leave a partial file. Mitigations (recommend, in order): (1) use apps with atomic write semantics (temp file + rename) — a partial temp never gets the final name (immich, *arr, qBittorrent all do this); (2) for write-critical non-atomic paths, stage on local SSD then move; (3) hard mounts avoid the partial (the write blocks until the NAS returns then completes) at the cost of the wedge — not recommended as the default.

Q6 — Performance — ~128 MB/s sequential through the chain — ample for streaming

512 MB sequential read through container→bind→guest→host→NFS measured 128.2 MB/s (≈ saturating the gigabit LAN). A 4K stream needs ≈36 MB/s, so this comfortably serves many concurrent HD/4K streams. (Host-side raw read reported 12 GB/s = page-cache served, not representative.) Recommendation: transcode scratch dirs and thumbnail/preview caches stay on the local SSD, not the NAS — they are latency-sensitive random I/O and high-churn; only the bulk media library lives on the NAS.

Q8 (secondary) — restic-over-SFTP backup target — works; failed-job-not-hang on NAS loss

restic init / backup / restore against sftp:user@192.168.0.180:/srv/nas-sim/backup/repo (no kernel mount) succeeded: snapshot … saved, and a restored 20 MB file matched the source md5 (c26d0b59…) exactly. Failure isolation: with SSH/22 black-holed, the backup failed as a userspace job (errored / had to be timed out — set -o sftp.args='-o ConnectTimeout=10' to fail fast instead of retrying the TCP connect). The key property: it is an ordinary killable process with no kernel mountno df/D-state/box wedge, just a failed backup. This is exactly the desired class-2 behaviour and de-risks that feature.

Q7 — Storage-class fit — model it as a distinct "network storage" class, NOT a drive

A NAS share must be a separate registry entry kind from the physical-drive model, because it is network + read-mostly + credentialed and has no device lifecycle:

  • It bypasses the drive enroll → format → eject → decommission → migrate machinery entirely (there is no block device to wipe, no durable-id, no SMART, no data-bearing-wipe gate).
  • It carries instead: kind: network, protocol: nfs|smb, server, export/share, credentials_ref (SMB; out-of-band), mount_opts, mapped_uid/gid (the +100000 recipe), automount: true, idle_unmount.
  • It is selectable as a media app's data path (the HDD_PATH/userdata slot) but must be disqualified as a target for: the DB/app-config class (class 1, SSD-only), and the live-write-critical fast paths. The controller already sees storage as paths; a network-storage path slots in as a new path kind with role-gating "bulk userdata / media only".
  • Health is per-share liveness (is the automount currently mounted + last-access OK), surfaced as "degraded: NAS unreachable" affecting only apps bound to it — never the box's overall health.

The mount chain (mirror the existing drive-bind model):

NAS share ──(host automount, agent-owned)──▶ /mnt/felhom-drives/<share> on felhom-pve
          ──(existing shared→slave propagation via mp8)──▶ same path inside guest 9201
          ──(docker -v /mnt/felhom-drives/<share>:/media)──▶ media container

The mount lives host-side (Q1: unprivileged LXC cannot mount NFS/CIFS); it propagates into the guest for free through the existing /mnt/felhom-drives shared bind — no new guest mountpoint, no guest restart.

Mount via systemd .automount (on-demand + idle-unmount) so an idle NAS reboot is a non-event and the stale-mount window is minimised. autofs is an equivalent alternative; systemd automount needs no extra package on the Proxmox host. Validated: before access the path is an autofs trigger (no real mount); a single ls mounts it on demand (nfs4); after TimeoutIdleSec (tested 20 s) it auto-unmounts; a re-access remounts. So when no app is reading, there is no NAS mount at all — a NAS reboot while idle cannot affect the box, and an in-use outage is governed by the soft semantics above.

Protocol + option sets:

  • NFS (preferred): vers=4.1,soft,timeo=50,retrans=2,noatime,_netdev; export all_squash,anonuid=<U>, anongid=<G> where <U>/<G> = container_uid/gid + 100000. Clean failure isolation.
  • SMB (SMB-only NAS): vers=3.0,credentials=<0600 file>,uid=<U>,gid=<G>,forceuid,forcegid, file_mode=0664,dir_mode=0775,_netdev (octal modes!), share force user/group = the uid/gid-<U>/<G> identity. Hangs-but-contained on NAS loss → always pair with automount idle-unmount.

UID/permission mapping (the +100000 rule): a media app at container uid/gid 1000 = host 101000. The NAS must present/own/squash files as 101000 so the guest sees 1000 and the container reads+writes natively. NFS: anonuid/anongid=101000. SMB: force user/group → uid/gid 101000 + client uid/gid=101000,forceuid,forcegid.

Failure / health model: NFS soft → a down NAS yields EIO to the accessing app only; df/guest/other apps stay healthy; auto-recovers on return. SMB → accessing app blocks (D-state) but contained + self-healing. Model per-share health as "degraded, NAS unreachable"; never fold it into box health. Writes: prefer atomic-write apps or local-staging (a soft-interrupted write leaves a partial in-flight file).

Local-SSD-stays-local: DB + app config/state (class 1) and transcode scratch + thumbnail/preview caches never go on the NAS — SSD only. The NAS holds bulk media library + (class 2) restic-SFTP backup repos.


Caveats / for the production TASK

  • Sim-vs-real-NAS (restated): Synology/QNAP SMB dialect/ACL/mapall specifics, SMB multichannel, and appliance reboot timing are unproven here — confirm on the real appliance before GA.
  • SMB credentials live out-of-band (a 0600 creds file the agent writes); never in git, never in the registry plaintext. The mapping uses a host uid/gid-101000 identity for force user/group.
  • restic-SFTP needs an explicit ConnectTimeout to fail fast; otherwise a down NAS slow-fails the backup job (still no kernel wedge).
  • Atomic-write requirement for write-apps on a soft NFS NAS — document per-app, or stage writes on SSD.

Probe teardown — DONE

Fully torn down and verified:

  • 180: /etc/exports + /etc/samba/smb.conf restored from the pre-spike .bak-nasspike copies (then the baks removed); the [nas-sim] share and /srv/nas-sim export are gone; nassim/nasguest users removed; scratch /srv/nas-sim deleted; the throwaway pubkey removed from authorized_keys; SMB-pw files deleted. Verified intact: both Longhorn-PVC NFS exports still present (count 2); all 6 original samba shares (kisfenyo_home/media/orsi_home/…) preserved; PBS/k3s untouched.
  • felhom-pve: all NAS mounts unmounted (0 remaining), the systemd automount units removed, the throwaway restic binary + SSH keypair removed, /mnt/felhom-drives still holds only the real drives (felhom-flash, felhom-usb). All iptables DROP rules removed throughout.
  • Live guest 9201 unharmed: gitea, rallly, felhom-controller all healthy after the spike.

No secrets in this doc. SMB/restic test credentials were throwaway and stored out-of-band only.