From 27e2fb05c08ccee05db113359b05caad69a022dd Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Sat, 11 Jul 2026 10:10:16 +0200 Subject: [PATCH] host-install v1.13.0: systemd-journal group for the agent user + NAS feature doc MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The NAS verify pipeline (agent v0.81.0) reads mount-unit journals unprivileged — group membership, NO sudoers grant. Fixes the v1.11.0/1.12.0 header drift. New authoritative feature doc documentation/controller/network-storage-nas.md (verify pipeline, §8 truth table, Q4 error taxonomy, retry=0, Route A recipes incl. the chmod-persists nuance). Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01PSK5g6qYLknKj8u3QAFEr6 --- .../controller/network-storage-nas.md | 126 ++++++++++++++++++ scripts/CHANGELOG.md | 11 ++ scripts/felhom-host-install.sh | 17 ++- 3 files changed, 152 insertions(+), 2 deletions(-) create mode 100644 documentation/controller/network-storage-nas.md diff --git a/documentation/controller/network-storage-nas.md b/documentation/controller/network-storage-nas.md new file mode 100644 index 0000000..1570700 --- /dev/null +++ b/documentation/controller/network-storage-nas.md @@ -0,0 +1,126 @@ +# Network storage (NAS) — bulk-media shares with verify-before-commit + +**Status:** authoritative feature doc, code-verified against felhom-agent v0.81.0 + +felhom-controller v0.113.0 (2026-07-11). +**Evidence base:** `documentation/audits/SPIKE-nas-storage-2026-06-29.md` (the locked mount recipe) +and `documentation/audits/SPIKE-nas-verify-2026-07-11.md` (the verify mechanisms, squash matrix, +error taxonomy — commit b57f6c1). Code: agent `internal/storage/netmount.go` / +`internal/storage/netverify.go` / `internal/localapi/netverifyjob.go`; controller +`internal/web/netstorage_job.go` / `internal/web/netprobe*.go`. + +## What it is + +A customer's NAS serves **bulk media** (film, photo, music) to media apps. The share is mounted +**host-side** by the agent under `/mnt/felhom-drives/` via a systemd `.automount` + `.mount` +pair (on-demand + idle-unmount), propagates into the guest through the existing shared `mp8` bind, +and is selectable as a media app's data path. A NAS is a **distinct storage class from a drive**: +no durable-id, no SMART, no enroll/eject/decommission/wipe/migrate — remove is the only lifecycle +action (`refuseNetworkLifecycle` blocks the drive verbs). + +## The locked mount recipe + +- **NFS (nfs4):** `vers=4.1,soft,timeo=50,retrans=2,noatime,_netdev,retry=0`. `soft` = failure + isolation (clean EIO, never a wedge). **`retry=0`** (SPIKE-nas-verify Q4-vi): without it a + dead-NAS on-demand access wedges the accessing app until systemd's 90 s start cap (measured + 91 s); with it the access fails clean in ~3.8 s (ENODEV) and each autofs re-access is a fresh + attempt. `retry` only governs retrying a FAILED first attempt — the happy path is untouched. + The uid mapping is the EXPORT's job (squash), never the client mount. +- **SMB (cifs):** `vers=3.0,credentials=<0600 file>,uid=,gid=,forceuid, + forcegid,file_mode=0664,dir_mode=0775,_netdev`. No `retry=` (a mount.nfs option; mount.cifs + rejects it). Credentials live in an agent-written 0600 file, never in the registry, never + controller-persisted. +- **The +100000 rule:** container uid/gid N = host N+100000 (unprivileged-LXC idmap). + +## Verify-before-commit (the add pipeline) + +`POST /api/storage/netstorage/add` never registers blind. The controller orchestrates +(detached job, single-flight, ~150 s budget, status on `GET /api/storage/netstorage/add/status`): + +1. **agent add** (`POST /netstorage/add`): role-gate → **sync fast-fail** (spec validation + 2 s + TCP pre-probe; an unreachable server is refused with NOTHING installed) → stage SMB creds → + install unit pair + enable the automount → start the agent's detached verify job. +2. **agent verify** (polled on `GET /netstorage/verify-status` every 2 s): a directory read + through the automount triggers a REAL mount (spike Q1: an in-guest or host-side access wakes + it, ~1 s LAN); success is judged from **/proc/mounts only** (truth table below); a failure is + classified from the mount unit's journal and **auto-rolled-back** agent-side (units + creds). +3. **in-guest uid-1000 write probe:** the controller re-execs itself + (`felhom-controller --netprobe `) with `SysProcAttr.Credential{Uid:1000,Gid:1000}` — + create a `.felhom-proba-` dot-file, write a nonce, read back, compare, delete. This + catches the **squash trap**: an export that mounts fine but denies every uid-1000 write + (spike Q3 row c). A failed delete after a good write is a WARN, not a failure. +4. **register** (`AddStoragePath`, Kind=network, Schedulable) — **the LAST step**. The worst + crash outcome is an agent-side orphan (surfaced in the list as a remove-only "Árva megosztás" + row), never a registered-but-broken path. + +Any failure = **full rollback**: nothing registered, no units left installed, no creds file. +If the agent restarts mid-verify, its in-memory verify slot reports phase `none`; the controller +treats that as verify-lost and rolls back (the orphan row is the belt-and-braces surface). + +### Mount-success truth table (SPIKE-nas-verify §8) + +| Trigger `ReadDir(where)` | /proc/mounts shows nfs4/cifs at where | Verdict | +|---|---|---| +| ok | yes | mount OK → the uid-1000 probe decides writability | +| EACCES / EPERM | yes | mount OK (the agent user just can't read it — a 0700 export is fine) | +| any result | no | mount FAILED → classify from the journal | +| still blocked at 95 s | — | `timeout` (black-holed-but-routed; systemd's 90 s cap resolves it) | + +Readability is NEVER the truth source — /proc/mounts is. + +### Error taxonomy (live-measured; classification is string-based BY DESIGN — every failure is rc=32) + +| Category | Journal substring (verbatim) | Meaning | +|---|---|---| +| `unreachable` | `No route to host` / `Connection refused` / `Connection timed out` (or the 2 s sync pre-probe) | no NAS behind the address | +| `nfs_export` | `reason given by server: No such file or directory` | export missing **OR not permitted — MERGED**: NFSv4 returns the identical string for both (Q4 ii≡iii); the customer message names both possibilities | +| `smb_auth` | `mount error(13)` | wrong SMB username/password | +| `smb_share` | `mount error(2)` | SMB share name not found | +| `timeout` | `Mounting timed out. Terminating` | routed but not answering (systemd 90 s cap) | +| `mount_failed` | (none matched / journal unavailable) | generic; detail carries the raw journal | +| `not_writable` | (probe exit 2) | mounts, but uid-1000 cannot write — the squash trap | +| `probe_io` | (probe exit 3) | write readback failed/differed | + +The journal is read **unprivileged** (`journalctl -u -n 20 -o cat`): the agent user is in +the `systemd-journal` group (host-install ≥ v1.13.0 adds it; existing hosts: +`usermod -aG systemd-journal felhom-agent && systemctl restart felhom-agent`). **No sudoers +grant** — journal access is group-based by rule. Journal unavailable degrades to `mount_failed` +with a hint; the rollback still runs. + +## NAS-side recipes (protocol-honest guidance — what the UI tells the customer) + +- **SMB (Synology, QNAP — the consumer default, listed FIRST in the UI):** a plain user account + with rw on the share is sufficient — **zero server-side uid configuration** (spike Q5). Files + land on the NAS owned by the connecting account; the client mount forces the guest view to + uid/gid 1000. Hardlinks worked on Debian Samba/ext4 — re-verify per appliance. +- **NFS, simple recipe (Route A — spike Q3 row b, ACCEPTED):** enable "map all users / all + squash" on the export, rw, to ANY local user. The guest-uid-1000 app gets full + read/write/rename/**hardlink**/delete. Caveats: + - guest-visible ownership is `nobody:nogroup` (65534) — cosmetic for apps that just read/write; + - `chown` fails with an immediate clean EPERM (no hang); + - **`chmod` SUCCEEDS and persists server-side** (the squashed identity owns every file) — modes + are app-controlled; operators should know a mode change on the NAS side is real. +- **NFS, full-fidelity recipe (TrueNAS / Linux server):** + `rw,all_squash,anonuid=,anongid=` (e.g. 101000 for a uid-1000 app) — + ownership information is exact end-to-end. +- **The pinned WRONG case (what `not_writable` usually means):** `anonuid=` **without** + `all_squash` is a no-op for non-root traffic — the app's wire uid (101000) hits the export + unmapped and fails on others-perms. The consumer recipe REQUIRES the map-ALL-users mode; an + anonymous-uid field alone does nothing. The UI's `not_writable` message and the guidance block + tell this same story. + +## Health model + +Per-share liveness only (`ok | idle | unreachable`): `idle` (automount idle-unmounted) is the +benign steady state; `unreachable` degrades the affected share's apps ONLY — never box health, +never the drive missing→stop cascade (kind-gated). Timing/budget note for integrators: the verify +worst case is systemd's 90 s (black-holed server) — verify traffic therefore rides the detached +job + status poll, never a single long HTTP call (the agentapi client keeps its global 15 s +timeout). + +## Open items + +- Q1c: whether an automount trigger installed BEFORE guest boot propagates at guest start — + needs a restart-window test (the add flow installs while the guest runs, which is proven). +- Synology/QNAP appliance fidelity pass (virtual-dsm) before GA. +- The demo's pre-existing `nas-media` unit (if any) predates `retry=0` and is not rewritten — + re-adding the share re-creates it with the current option string. diff --git a/scripts/CHANGELOG.md b/scripts/CHANGELOG.md index f454ebd..0a77f18 100644 --- a/scripts/CHANGELOG.md +++ b/scripts/CHANGELOG.md @@ -1,5 +1,16 @@ # Felhom scripts — Changelog +## felhom-host-install v1.13.0 — systemd-journal group for the agent user (NAS verify) (2026-07-11) + +- The agent-user setup now adds `felhom-agent` to the **`systemd-journal` group** (idempotent + `usermod -aG`): the NAS verify-before-commit pipeline (agent v0.81.0) classifies mount failures + from an UNPRIVILEGED `journalctl -u ` read — group membership instead of a sudoers grant + (rule: no journalctl via sudo). Existing installs get the one-liner + `usermod -aG systemd-journal felhom-agent && systemctl restart felhom-agent` (applied to the + demo host during this task's live validation; Peti's box only when the NAS feature reaches him). +- Version-header drift fixed: the file header still said v1.11.0 while `SCRIPT_VERSION` was + 1.12.0 (the v1.12.0 bump shipped without a changelog entry or header sync); both now 1.13.0. + ## felhom-host-install v1.11.0 — uninstaller gap-closure + operator-key day-0 fold (TASK GL-4, go-live G6 + G1-follow-up) (2026-07-08) - **Uninstall gap-closure (G6):** diff --git a/scripts/felhom-host-install.sh b/scripts/felhom-host-install.sh index 9e3c989..2144693 100644 --- a/scripts/felhom-host-install.sh +++ b/scripts/felhom-host-install.sh @@ -1,6 +1,6 @@ #!/bin/bash #=============================================================================== -# felhom-host-install.sh v1.11.0 +# felhom-host-install.sh v1.13.0 # Day-0 host-bootstrap for a Felhom Proxmox host (operator-deploy model). # # Run by the operator on a FRESHLY-PVE-INSTALLED box (after a manual PVE install @@ -163,7 +163,7 @@ set -euo pipefail -SCRIPT_VERSION="1.12.0" # keep in sync with the header line at the top of this file +SCRIPT_VERSION="1.13.0" # keep in sync with the header line at the top of this file # Operator signing keys pinned at day-0 (GL-4; doc 04 §3 two-key model). EMPTY by default — the pin # CEREMONY is an operator step: generate the real keypairs OFFLINE, then fill these four constants @@ -1697,6 +1697,19 @@ step_agent_install() { log_success " created service user $AGENT_USER" fi + # systemd-journal group: the NAS verify pipeline (agent v0.81.0) classifies mount failures from + # an UNPRIVILEGED `journalctl -u ` read — group membership instead of a sudoers grant + # (rule: no journalctl via sudo). Idempotent. Existing installs: run the usermod by hand, then + # `systemctl restart felhom-agent`. + if $DRY_RUN; then + log_dry "usermod -aG systemd-journal $AGENT_USER" + elif id -nG "$AGENT_USER" | tr ' ' '\n' | grep -qx systemd-journal; then + log_info " $AGENT_USER already in systemd-journal" + else + usermod -aG systemd-journal "$AGENT_USER" + log_success " added $AGENT_USER to systemd-journal (unprivileged journal read for NAS verify)" + fi + # State dir (the old root deployment may have created it root-owned; StateDirectory= also adjusts # on start, but chown here so the very first start has a writable dir). run mkdir -p "$AGENT_STATE_DIR"