spike: NAS network storage (media automount + restic backup target) — findings

Empirical validation on a DooPlex-simulated NAS (SMB+NFS). Verdict READY for the
production network-storage TASK. Key findings: mount must be host-side (unprivileged
LXC blocks NFS/CIFS mount) + bind-propagated into the guest; NFS soft fails-clean +
auto-recovers (SMB hangs-but-contained); UID mapping = container uid + 100000 LXC
offset; ~128MB/s adequate; write integrity intact (atomic-write apps safe on soft);
restic-SFTP backup/restore works + fails-as-job not hang; systemd automount gives
on-demand + idle-unmount. Distinct network-storage class bypasses the drive lifecycle.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-06-29 11:44:40 +02:00
parent b4ce6c10b5
commit 0719db0d53
@@ -0,0 +1,232 @@
# SPIKE — NAS network storage (media automount + backup target), simulated on DooPlex
**Date:** 2026-06-29
**Class:** Spike (empirical validation; output = this doc only). Throwaway probes; nothing merged into
agent/controller production packages.
**Goal:** validate the **bulk-media network-storage** mechanism (storage class 3) — host-side mount of a NAS
share → bind into guest 9201 → bind into media container, for both **SMB** and **NFS** — before writing the
production TASK. Secondary: restic-over-SFTP as a backup target (class 2).
**Simulated NAS:** real Samba (SMB) + nfs-kernel-server (NFS) already running on **DooPlex 192.168.0.180**;
added an **isolated scratch export** `/srv/nas-sim/media` (sample media, ~513 MB incl. a 512 MB file), separate
from the live Longhorn-PVC NFS exports and PBS/k3s. Mount probes on the Proxmox host **felhom-pve**
(192.168.0.162) → guest **9201** (live demo guest) → throwaway alpine containers.
> **Sim-vs-real-NAS caveat (stated up front):** this validates the Linux *mechanism* (kernel mount, propagation,
> idmap, failure semantics) against a Debian Samba/NFS server. It does **not** reproduce Synology/QNAP-specific
> quirks: their SMB dialect ceilings, NFS `squash`/`mapall` UIs, per-share ACL models, SMB multichannel, or
> appliance reboot timing. Those need a confirmation pass on the real appliance before GA.
---
## Verdict: **READY** to write the production "network storage" TASK (with named caveats)
NFS and SMB both work end-to-end through the host-mount→bind chain for **read and write**, with a precise
UID-mapping recipe; NFS `soft` mounts achieve clean failure isolation (fail-soft + auto-recover, no guest/df
wedge). SMB failure isolation is the one area needing the documented option set (below). Restic-SFTP backup is
trivial and failure-isolated. No blockers; see "Caveats / for the production TASK."
---
## Per-Spike-Question findings
### Q1 — Mount placement: **HOST-side, bind into guest (in-guest mount is kernel-blocked)**
Guest 9201 is an **unprivileged** LXC (`unprivileged: 1`, `features: nesting,keyctl`). Mounting NFS or CIFS
**inside** the guest fails:
```
# pct exec 9201 -- mount -t nfs ... → mount: permission denied (exit 32)
# pct exec 9201 -- mount -t cifs ... → mount: permission denied (exit 32)
# CapEff: 000001ffffffffff (full caps *within the userns* — but mounting these fs types needs the INIT userns)
```
The kernel forbids NFS/CIFS mounts in a user namespace regardless of the in-userns capability set. **Therefore
the NAS mount must live on the Proxmox host (agent-owned)** and be bind-propagated into the guest — exactly
mirroring the existing drive-bind architecture.
**Propagation works for free** via the existing `mp8: /mnt/felhom-drives` bind: the host path is `shared`
propagation, the guest sees it `shared,slave`, and existing drive mounts (felhom-flash/felhom-usb) already
propagate through. A **new host mount under `/mnt/felhom-drives/<name>` appears live in the guest with no
restart** — confirmed.
**Recommended chain:** `host: mount NAS → /mnt/felhom-drives/<share>` → (auto-propagates) `guest:
/mnt/felhom-drives/<share>``docker -v /mnt/felhom-drives/<share>:/media`.
### Q2 — SMB and NFS both work through the full chain — **YES (read+write)**
**NFS** (`mount -t nfs -o vers=4.1,soft,timeo=50,retrans=2,_netdev,noatime`): host mount OK → propagated to
guest → container read the 512 MB file (md5 verified) and **wrote** a new file. Effective options included
`soft,timeo=50,retrans=2`.
**SMB/CIFS** (`mount -t cifs -o credentials=…,uid=101000,gid=101000,forceuid,forcegid,file_mode=0664,dir_mode=0775,vers=3.0,_netdev`):
host mount OK → propagated → container read and **wrote** over SMB.
Minimal viable option sets:
- **NFS:** `vers=4.1,soft,timeo=50,retrans=2,_netdev,noatime` + the export's `all_squash,anonuid=…,anongid=…`
(see Q5). `vers=4.1` avoids the rpcbind/111 + lock-manager surface of v3.
- **SMB:** `vers=3.0,credentials=<file>,uid=…,gid=…,forceuid,forcegid,file_mode=0664,dir_mode=0775,_netdev`
(creds file 0600 on the host; never inline). `dir_mode`/`file_mode` MUST be plain octal (`0775`, not `2775`).
### Q5 — UID / permission mapping: **the LXC +100000 offset is the whole story** (read+write proven)
The established felhom convention (verified against the live `felhom-usb` userdata): a dir owned **host
`100000:101000`** appears in the guest as **`root:1000`**, mode **2775** (setgid, group-writable); media apps
write via **gid 1000**. So **guest/container uid/gid N = host uid/gid N+100000.**
A naïve `anonuid=1000` export is WRONG — it lands as host-1000, which is *outside* the guest's 100000165535
idmap and shows as `nobody:nogroup` (readable because world-readable, but **not writable**). The fix is to
present the share as the **host-mapped** id:
- **NFS:** export `…,all_squash,anonuid=101000,anongid=101000` and own the share dir `101000:101000`. Then the
guest sees `1000:1000`, and a uid/gid-1000 container **reads existing files and writes new ones**; the write
lands on the NAS as `101000:101000`. (Verified: `WRITE_OK`, server file `-rw-r--r-- 101000 101000`.)
- **SMB:** mount with `uid=101000,gid=101000,forceuid,forcegid` (client forces the guest-visible owner to
1000:1000 regardless of server ownership) **and** the share's `force user`/`force group` must resolve to a
**uid/gid 101000** identity so server-side writes are actually permitted on the 101000-owned dir. (Verified:
with `force user=nasguest` (uid 101000) the container `SMB_WRITE_OK`; with `force user`=a uid-1000 identity
the write was denied — the mismatch is exactly the trap to document.)
**Mapping recipe for the production class:** the NAS share for a customer's media must be owned/squashed to
`container_uid + 100000` (e.g. media app uid 1000 → NAS uses **101000**). On a real Synology/QNAP this means an
export `mapall`/`anonuid` of 101000 (NFS) or an SMB share whose force-user resolves to 101000.
### Q3 — Failure isolation (THE risk)
Method: with a media container looping reads, **black-holed** the NAS from felhom-pve via an **iptables DROP on
180 scoped to 192.168.0.162** (DROP, not REJECT → simulates a dead NAS / network black hole; leaves the NFS
service up for k3s/Longhorn — the safe way to "disappear" it).
**NFS (`soft,timeo=50,retrans=2`) — clean fail-soft + auto-recover:**
- Uncached `stat` returned **`Input/output error` after ~16 s** (≈ retrans×timeo). It **errors, does not hang.**
- `df` on the mount did **not** hang (returned rc 0); no reader process in **D** (uninterruptible) state.
- The looping reader kept serving from page-cache briefly, then flipped to **FAIL** once the attribute cache
expired and revalidation hit the black-holed server — i.e. **fails soft, no wedge**.
- **Other apps unaffected:** gitea stayed **HTTP 200** throughout; the guest stayed healthy (`uptime`/commands
responsive).
- **Recovery on NAS return:** after removing the DROP, a fresh read was `RECOVERED_READ_OK`, and **the looping
reader auto-recovered on its own** (`FAIL → PASS`) with **no remount/intervention**. `stat` worked again.
**SMB/CIFS (default options) — hangs, but blast-radius contained:**
- With the NAS black-holed, `df` on the mount **hung the full 40 s timeout** (rc 124); an uncached `stat`
**hung the full 60 s**; the reader's `dd` went **D (uninterruptible)** state. CIFS has **no `soft`/`hard`
knob** like NFS — a dead server blocks operations until the SMB session/TCP timeout.
- BUT the blast radius was still **contained**: gitea stayed **HTTP 200**, the guest stayed healthy
(`uptime` responsive) — only the process touching the SMB mount wedged.
- **Recovery:** on NAS return the wedged `dd` **unblocked on its own** (`df` instant, no D-state, reader
`PASS` again) — no kill/remount needed.
**Q3 verdict:** **prefer NFS `soft` for the media class** — it fails-soft (clean EIO) and never wedges
`df`/the process. **SMB hangs the accessing process** during the outage (contained, self-healing) and is the
fallback only for NAS that speak SMB-only; pair it with **automount idle-unmount** (below) to minimise the
stale-mount window. The WRONG outcome (a hard mount wedging the *guest* or *df* box-wide) was **not** observed
in either case — the unprivileged-LXC bind boundary kept the blast radius to the accessing container.
### Q4 — Write reliability — **integrity perfect; an interrupted in-flight write leaves a partial file**
- **Clean write integrity:** a container (uid 1000) wrote a 100 MB file through the chain; md5 on the
host-mount side == md5 read **directly on the NAS server** (`2541a840…`) — byte-intact end to end.
- **Disappear mid-write** (write-loop of 1 MB fsync'd files, NAS black-holed at file ~221): every **completed**
file before the cut was byte-intact (exactly 1 MB); the **single in-flight file (`w221.bin`) was left
partial/truncated** on the NAS; the writer's `dd` went **D-state** (writes-in-flight soft-fail slower than
reads — soft eventually returns EIO, but the in-flight file is already partially persisted).
- **Safe-option note:** `soft` is correct for read-mostly media, but for **write apps a `soft`-interrupted
write can leave a partial file**. Mitigations (recommend, in order): (1) use apps with **atomic write**
semantics (temp file + `rename`) — a partial temp never gets the final name (immich, *arr, qBittorrent all
do this); (2) for write-critical non-atomic paths, **stage on local SSD then move**; (3) `hard` mounts avoid
the partial (the write blocks until the NAS returns then completes) at the cost of the wedge — **not**
recommended as the default.
### Q6 — Performance — **~128 MB/s sequential through the chain — ample for streaming**
512 MB sequential read through container→bind→guest→host→NFS measured **128.2 MB/s** (≈ saturating the
gigabit LAN). A 4K stream needs ≈36 MB/s, so this comfortably serves many concurrent HD/4K streams. (Host-side
raw read reported 12 GB/s = page-cache served, not representative.) **Recommendation: transcode scratch dirs and
thumbnail/preview caches stay on the local SSD**, not the NAS — they are latency-sensitive random I/O and
high-churn; only the bulk media library lives on the NAS.
### Q8 (secondary) — restic-over-SFTP backup target — **works; failed-job-not-hang on NAS loss**
`restic init` / `backup` / `restore` against `sftp:user@192.168.0.180:/srv/nas-sim/backup/repo` (no kernel
mount) succeeded: `snapshot … saved`, and a restored 20 MB file matched the source md5 (`c26d0b59…`) exactly.
**Failure isolation:** with SSH/22 black-holed, the backup **failed as a userspace job** (errored / had to be
timed out — set `-o sftp.args='-o ConnectTimeout=10'` to fail fast instead of retrying the TCP connect). The
key property: it is an ordinary **killable process with no kernel mount****no `df`/D-state/box wedge**, just
a failed backup. This is exactly the desired class-2 behaviour and de-risks that feature.
### Q7 — Storage-class fit — **model it as a distinct "network storage" class, NOT a drive**
A NAS share must be a **separate registry entry kind** from the physical-drive model, because it is network +
read-mostly + credentialed and has **no device lifecycle**:
- It **bypasses** the drive `enroll → format → eject → decommission → migrate` machinery entirely (there is no
block device to wipe, no durable-id, no SMART, no data-bearing-wipe gate).
- It carries instead: `kind: network`, `protocol: nfs|smb`, `server`, `export/share`, `credentials_ref`
(SMB; out-of-band), `mount_opts`, `mapped_uid/gid` (the +100000 recipe), `automount: true`, `idle_unmount`.
- It is **selectable as a media app's data path** (the `HDD_PATH`/userdata slot) but must be **disqualified as
a target for**: the DB/app-config class (class 1, SSD-only), and the live-write-critical fast paths. The
controller already sees storage as *paths*; a network-storage path slots in as a new path *kind* with
role-gating "bulk userdata / media only".
- Health is **per-share liveness** (is the automount currently mounted + last-access OK), surfaced as
"degraded: NAS unreachable" affecting only apps bound to it — never the box's overall health.
---
## Recommended production architecture
**The mount chain (mirror the existing drive-bind model):**
```
NAS share ──(host automount, agent-owned)──▶ /mnt/felhom-drives/<share> on felhom-pve
──(existing shared→slave propagation via mp8)──▶ same path inside guest 9201
──(docker -v /mnt/felhom-drives/<share>:/media)──▶ media container
```
The mount lives **host-side** (Q1: unprivileged LXC cannot mount NFS/CIFS); it propagates into the guest for
free through the existing `/mnt/felhom-drives` shared bind — **no new guest mountpoint, no guest restart**.
**Mount via systemd `.automount` (on-demand + idle-unmount)** so an idle NAS reboot is a non-event and the
stale-mount window is minimised. autofs is an equivalent alternative; systemd automount needs **no extra
package** on the Proxmox host. **Validated:** before access the path is an `autofs` trigger (no real mount); a
single `ls` mounts it on demand (`nfs4`); after `TimeoutIdleSec` (tested 20 s) it **auto-unmounts**; a
re-access remounts. So when no app is reading, there is **no NAS mount at all** — a NAS reboot while idle
cannot affect the box, and an in-use outage is governed by the `soft` semantics above.
**Protocol + option sets:**
- **NFS (preferred):** `vers=4.1,soft,timeo=50,retrans=2,noatime,_netdev`; export `all_squash,anonuid=<U>,
anongid=<G>` where `<U>/<G> = container_uid/gid + 100000`. Clean failure isolation.
- **SMB (SMB-only NAS):** `vers=3.0,credentials=<0600 file>,uid=<U>,gid=<G>,forceuid,forcegid,
file_mode=0664,dir_mode=0775,_netdev` (octal modes!), share `force user/group` = the uid/gid-`<U>/<G>`
identity. Hangs-but-contained on NAS loss → always pair with automount idle-unmount.
**UID/permission mapping (the +100000 rule):** a media app at container uid/gid 1000 = host 101000. The NAS must
present/own/squash files as **101000** so the guest sees **1000** and the container reads+writes natively. NFS:
`anonuid/anongid=101000`. SMB: `force user/group` → uid/gid 101000 + client `uid/gid=101000,forceuid,forcegid`.
**Failure / health model:** NFS `soft` → a down NAS yields EIO to the accessing app only; `df`/guest/other apps
stay healthy; auto-recovers on return. SMB → accessing app blocks (D-state) but contained + self-healing. Model
per-share health as "degraded, NAS unreachable"; never fold it into box health. **Writes:** prefer atomic-write
apps or local-staging (a `soft`-interrupted write leaves a partial in-flight file).
**Local-SSD-stays-local:** DB + app config/state (class 1) and **transcode scratch + thumbnail/preview caches**
never go on the NAS — SSD only. The NAS holds bulk media library + (class 2) restic-SFTP backup repos.
---
## Caveats / for the production TASK
- **Sim-vs-real-NAS** (restated): Synology/QNAP SMB dialect/ACL/`mapall` specifics, SMB multichannel, and
appliance reboot timing are unproven here — confirm on the real appliance before GA.
- **SMB credentials** live out-of-band (a 0600 creds file the agent writes); never in git, never in the
registry plaintext. The mapping uses a host uid/gid-101000 identity for `force user/group`.
- **restic-SFTP** needs an explicit `ConnectTimeout` to fail fast; otherwise a down NAS slow-fails the backup
job (still no kernel wedge).
- **Atomic-write requirement** for write-apps on a `soft` NFS NAS — document per-app, or stage writes on SSD.
## Probe teardown — DONE
Fully torn down and verified:
- **180:** `/etc/exports` + `/etc/samba/smb.conf` restored from the pre-spike `.bak-nasspike` copies (then the
baks removed); the `[nas-sim]` share and `/srv/nas-sim` export are gone; `nassim`/`nasguest` users removed;
scratch `/srv/nas-sim` deleted; the throwaway pubkey removed from `authorized_keys`; SMB-pw files deleted.
**Verified intact:** both Longhorn-PVC NFS exports still present (count 2); all 6 original samba shares
(`kisfenyo_home`/`media`/`orsi_home`/…) preserved; PBS/k3s untouched.
- **felhom-pve:** all NAS mounts unmounted (0 remaining), the systemd automount units removed, the throwaway
restic binary + SSH keypair removed, `/mnt/felhom-drives` still holds only the real drives
(`felhom-flash`, `felhom-usb`). All iptables DROP rules removed throughout.
- **Live guest 9201 unharmed:** `gitea`, `rallly`, `felhom-controller` all `healthy` after the spike.
No secrets in this doc. SMB/restic test credentials were throwaway and stored out-of-band only.