Files
felhom.eu/documentation/audits/DIAG-sharing-2026-07-20.md
T
admin 1d1d60a737 docs: DIAG addendum 2 — S-3 fixed (samba 1.1.0 mDNS + v0.152.0 copy); Finder sidebar stays OPEN
Capture disproved addendum 1's guess: macOS DOES send a correct NBNS query and
IS answered correctly in 140us; it simply never acts on it. Records the proven
per-client matrix and keeps sidebar discovery explicitly unproven.
2026-07-20 13:40:55 +02:00

23 KiB
Raw Blame History

DIAG — Megosztás: SMB unreachable from Mac + /sharing reload loop (2026-07-20)

Class: diagnosis run (RUNBOOK). Read-only. No code changes, no version bumps, no builds, no sanctioned restart was needed (H2 was ruled out at the first probe — the container was already up).

Baselines at run time

Repo main @ Deployed
felhom-controller 8db9232 gitea.dooplex.hu/admin/felhom-controller:0.150.0 on 9201 — Up 23 min (healthy)
felhom-samba gitea.dooplex.hu/admin/felhom-samba:1.0.0 — Up 3 hours

Site: felhom-pve at the remote location, vmbr0 static 192.168.0.162/24, gw 192.168.0.1.


Verdicts

Symptom Verdict Root cause
1. smb://192.168.0.162 fails from the Mac ROOT-CAUSED .162 is the Proxmox host, never an SMB endpoint. smbd lives in guest 9201 and binds 192.168.0.104:445. .162:445 = connection refused, by design.
1b. smb://FELHOM also fails OPEN — narrowed NetBIOS name resolution works on the wire (nmblookupFELHOM<00> = 192.168.0.104), but the stack advertises no mDNS/Bonjour (no avahi, nothing on udp/5353) — the mechanism macOS Finder actually prefers. Cannot be closed without one test from the Mac (§ Mac matrix).
2. /sharing full-page reloads ~1× per second ROOT-CAUSED GET /sharing/status returns phase:"running" on every load (liveness coercion added in b5d78d1, v0.147.0, 2026-07-19); the page's JS treats running as "the job just finished" and fires location.reload() 1.2 s later. Unconditional loop whenever sharing is up.

Hypothesis outcome: H1 confirmed, in a stronger form than stated (not a sticky in-memory job — there was no job at all this process; the handler manufactures the terminal phase). H2 ruled out. H3 confirmed for the .162 half of symptom 1; the \\FELHOM half is a new, separate defect (no Bonjour advertisement) and stays OPEN pending the Mac-side probe.


Phase A — Live state snapshot (guest 9201)

$ docker ps -a
felhom-samba        Up 3 hours    gitea.dooplex.hu/admin/felhom-samba:1.0.0
felhom-controller   Up 23 minutes (healthy)   …/felhom-controller:0.150.0

No Exited, no Restarting, no restart churn → H2 dead on arrival.

$ docker logs felhom-samba --tail 80
[felhom-samba] launching nmbd + wsdd + smbd (server=FELHOM iface=eth0 uid=1000)
[felhom-samba] launching nmbd + wsdd + smbd (server=FELHOM iface=eth0 uid=1000)

(Two lines = the two lifetimes since creation. Clean, no errors.)

$ ip -4 addr
2: eth0@if7 …  inet 192.168.0.104/24 brd 192.168.0.255 scope global dynamic eth0

The guest's LAN IP is 192.168.0.104 (DHCP lease, ~54 min remaining at probe time).

$ ss -tlnp | grep ':445'
LISTEN  192.168.0.104:445   users:(("smbd",pid=2137,fd=28))
LISTEN      127.0.0.1:445   users:(("smbd",pid=2137,fd=29))
LISTEN          [::1]:445   users:(("smbd",pid=2137,fd=30))
$ ss -ulnp | grep -E ':137|:138'
UNCONN  192.168.0.255:137 / 192.168.0.104:137 / 0.0.0.0:137   nmbd pid=2146
UNCONN  192.168.0.255:138 / 192.168.0.104:138 / 0.0.0.0:138   nmbd pid=2146
$ ss -tlnp | grep ':5357'
LISTEN  192.168.0.104:5357  users:(("python3",pid=2147,fd=10))    # wsdd

All three daemons up and bound exactly as infra/samba.go intends. Nothing on udp/5353 — no mDNS responder anywhere in the guest.

Reachability from the Proxmox host (same L2 segment as the Mac):

$ nc -zv -w3 192.168.0.104 445 →  (UNKNOWN) [192.168.0.104] 445 (microsoft-ds) open      rc=0
$ nc -zv -w3 192.168.0.162 445 →  demo-felhom.local [192.168.0.162] 445 : Connection refused  rc=1

SMB is reachable on the LAN at 192.168.0.104:445. .162:445 refused — documents that the host IP was never the SMB endpoint.

No packet filter in the path: /etc/pve/firewall/cluster.fw and /etc/pve/firewall/9201.fw do not exist; pct config 9201 shows net0: name=eth0,bridge=vmbr0,ip=dhcp with no firewall flag.

Service-level proof from inside the container:

$ docker exec felhom-samba smbclient -L 127.0.0.1 -N
Anonymous login successful
$ docker exec felhom-samba nmblookup -B 192.168.0.255 FELHOM
querying FELHOM on 192.168.0.255
192.168.0.104 FELHOM<00>                                    rc=0

NetBIOS name resolution answers correctly on the LAN broadcast.


Phase B — The reload loop

The decisive code

controller/internal/web/sharing_handlers.gosharingStatusHandler (blame b5d78d1, "v0.147.0 — feedback slice 1", 2026-07-19):

phase := sambaPhaseIdle
if job := s.sambaEnsure.snapshot(); job != nil { phase, errMsg = job.Phase, job.Error }
running := s.stackMgr != nil && s.stackMgr.SambaRunning()
// A stale `idle`/`running` job must never contradict reality: liveness wins on a fresh page.
if phase == sambaPhaseIdle && running {
    phase = sambaPhaseRunning        // ← line 246
}

controller/internal/web/templates/sharing.html L313335, the 1.5 s tick:

if (ph === 'running') {
    stop();
    show('alert-success', 'A megosztási szolgáltatás fut.');
    /* Repaint the „Állapot" badge, which was rendered server-side as „áll". */
    setTimeout(function(){ location.reload(); }, 1200);
    return;
}

tick();                       // ← fires IMMEDIATELY on load, before the interval
timer = setInterval(tick, 1500);

The defect. The two halves were written for different worlds and contradict each other:

  • The JS reads running as an edge — "the ensure job I was watching has just succeeded, so the server-rendered badge is stale; reload once to repaint it."
  • Since b5d78d1 the handler reports running as a level — "the container is alive", true on every page load forever after, with no job in sight.

tick() is called synchronously at script end, so the very first poll of every fresh page load returns running → reload scheduled at +1200 ms → new page → new tick() → reload. Period ≈ 1.2 s + load time, matching Viktor's "roughly every second". The stop() is irrelevant: the timer dies, but the page it lived on is destroyed 1.2 s later anyway.

It is unconditional — any customer with sharing enabled and the container up cannot use the Megosztás page at all. setTimeout also survives stop(), so there is no in-page escape.

The live signal (6 polls, 2 s apart, authenticated)

{"data":{"error":"","phase":"running","running":true},"ok":true}      ×6, identical

Constant terminal phase, never idle, never flapping ⇒ H1, and the constancy is structural rather than a leftover job.

Onset correlation

$ docker logs felhom-controller --since 3h | grep -iE 'samba|sharing|megoszt'
(no output)

No ensure job ran in this controller process at alls.sambaEnsure.snapshot() is nil, so phase starts at idle and the reported running comes purely from line 246. This rules out "today's deploy restarts kicked a job that got stuck": there is nothing stuck; the loop needs only a running container and a page load.

Onset is therefore v0.147.0 (2026-07-19), the commit that added line 246 — one day before the report, consistent with "didn't do this before". Today's 0.148→0.149→0.150 deploys are innocent.


Phase C — SMB service correctness

Live smb.conf inside the container matches RenderSambaConfig byte-for-byte in the global block:

[global]
  workgroup = WORKGROUP
  netbios name = FELHOM
  security = user
  map to guest = never
  server min protocol = SMB2
  disable netbios = no
  bind interfaces only = yes
  interfaces = lo eth0
  smb ports = 445
  …
[Share]
  path = /mnt/felhom-drives/hdd_1/shares/Share
  read only = no
  • No baked address literal. interfaces = lo eth0 is interface-scoped, not IP-scoped — it follows the DHCP lease automatically. The .162 mismatch is not config drift; .162 is simply a different machine.
  • One share is exported: Share.
  • Discovery daemons: nmbd (answers NBNS, § Phase A), wsdd (tcp/5357 bound — the Windows path), avahi/mDNS — absent by design of the R-6 spike verdict, which selected smbd + nmbd + wsdd.

The subnet-collision hazard (noticed while mapping the LAN)

$ ip route            (felhom-pve)
default via 192.168.0.1 dev vmbr0
192.168.0.0/24 dev vmbr0 proto kernel scope link src 192.168.0.162
$ ip neigh show dev vmbr0
192.168.0.1    REACHABLE      192.168.0.104  REACHABLE   (the guest)
192.168.0.230  STALE          192.168.0.11   STALE
192.168.0.180  FAILED         ← DooPlex is NOT on this segment

The remote site's LAN is also 192.168.0.0/24 — the same prefix as the DooPlex home LAN, which the Tailscale subnet router advertises. Consequences to keep in mind while the box is away:

  • A successful ping 192.168.0.162 does not prove the Mac is on the remote segment — if the Mac were at home, the same address could resolve to a home device or route over the tailnet.
  • Any 192.168.0.x address is ambiguous across the two sites for the duration of the trip.

This does not change the verdicts above, but it is the reason the Mac-side probe in the next section starts with "which network am I actually on".


Findings

# Sev Finding Evidence Proposed fix GO needed?
S-1FIXED v0.151.0 HIGH /sharing reload-loops forever for every customer with sharing enabled — the page is unusable. sharingStatusHandler reports running as a level; the JS treats it as a one-shot success edge and calls location.reload(). sharing_handlers.go:246 (blame b5d78d1), sharing.html L320326, 6× identical phase:"running" poll XS. Make the JS reload only when it has observed a transition — latch a sawInFlight flag on the first tick and reload on running only if it is set; otherwise just paint the success alert and stop. Server-side alternative (also XS): report the coerced level as a distinct phase (running_steady) the JS renders but never reloads on. Prefer the JS latch — it keeps the honest "liveness wins" contract the handler comment states. Needs a red-proof: a template/JS test or handler test asserting a fresh page load with a live container does not schedule a reload. no — plain bug fix
S-2SHIPPED v0.151.0 MED The UI never tells the customer where to connect. sharing.html L32/L61 show only \\FELHOM; there is no IP anywhere on the page, so a customer who cannot resolve the name has nothing to fall back on and will guess (Viktor guessed the Proxmox host IP). sharing.html L3032, L61; symptom 1 itself S. Render the guest's current LAN address beside the name: \\FELHOM (vagy smb://192.168.0.104), sourced from the same place the dashboard already knows the guest IP. Add a one-line macOS hint (Finder → Ugrás → Csatlakozás a szerverhez). no
S-3FIXED v0.152.0 / samba 1.1.0 MED No mDNS/Bonjour advertisement. smbd + nmbd + wsdd covers Windows (WSD) and legacy NetBIOS, but macOS Finder's preferred path is Bonjour — FELHOM.local does not resolve and the share never appears in the Finder sidebar. Likely the whole of symptom 1b. nothing on udp/5353 (§ Phase A); container log daemon list; infra/samba.go header comment citing the R-6 spike SM. Add avahi-daemon (or wsdd's mDNS sibling) to the felhom-samba image publishing _smb._tcp, host-net so multicast reaches the LAN. Re-runs the R-6 spike's constraint set — treat as a slice, not a hotfix. Confirm with the Mac matrix first (below) so we fix the mechanism that is actually failing. yes — image change + republish
S-4CORE SHIPPED v0.151.0 LOW Success-state semantics: the ensure job's phase never returns to idle after a terminal phase, and SambaRunning() masks that anyway. Harmless today only because S-1's fix removes the consumer that cared. Worth folding into the unified async-job feedback roadmap item the samba_ensure_job.go header already names. samba_ensure_job.go (no reset path); handler line 246 Fold into the roadmap item; no standalone task. no
S-5SHIPPED v0.151.0 INFO The guest holds its LAN address by DHCP (ip=dhcp, lease ~54 min at probe). Any printed IP is not stable across a lease change or a site move, which weakens S-2 unless the page re-reads it live (it should — render it, never persist it). pct config 9201; valid_lft 3251sec Render live, never cache. Consider a DHCP reservation at the customer's router as a documented install step. no

Recommended packaging: S-1 alone as an immediate patch task (it bricks a shipped page). S-2 with it if the guest IP is already available to the sharing handler — same page, same release. S-3 as its own slice with an image republish. S-4 stays on the roadmap.


Mac test matrix (for Viktor — CC cannot drive the Mac)

Run in Terminal, in order. First one is the important one.

# Command Expected if the diagnosis is right What a surprise means
1 ifconfig | grep 'inet 192' an address on 192.168.0.0/24 and route -n get 192.168.0.104 pointing at the Wi-Fi/Ethernet interface, not utun* if it leaves via utun*, you are at home on the tailnet and the whole test is against the wrong site — see the subnet-collision section
2 nc -zv 192.168.0.104 445 succeeded! a timeout with #1 green = something between the Mac and the guest filters 445 (AP client isolation / guest Wi-Fi) — a new finding, tell CC
3 Finder → Ugrás → Csatlakozás a szerverhez → smb://192.168.0.104 login prompt; user felhom, the household password (out-of-band) → the Share folder mounts this is the fix for symptom 1
4 smbutil lookup FELHOM returns 192.168.0.104 empty/error ⇒ macOS is not doing NBNS at all ⇒ S-3 confirmed as the cause of symptom 1b, and the mDNS slice is the real fix
5 dns-sd -B _smb._tcp FELHOM will not appear — predicted by S-3 if it does appear, S-3 is wrong and 1b needs re-opening
6 Remove the stale smb://…162 Finder favourite (Finder → Ugrás → Csatlakozás → select → ), and Kilépés a szerverről any half-mounted entry stops the recurring failure dialog

Also expect the /sharing page to keep reloading until S-1 ships. To read it in the meantime, load it with JavaScript disabled, or read the state from docker ps in the guest.


Actions taken

None. No restart was needed (H2 excluded by the very first probe — the samba container had been up 3 h with clean logs). No files, settings, smb.conf, containers, or guest network state were modified. Every command in this document is a read, except curl -X POST /login to obtain a session for the /sharing/status polls, which mutates nothing.

Observations (noticed, not acted on)

  • nmblookup -U 127.0.0.1 -R FELHOM mis-parses: -R is --name-resolve=ORDER in this Samba build, not --recursion, so the invocation errors out. Diagnostic-tooling trivia only — -B works and was used for the real proof. Worth remembering for future runbooks.
  • smbclient -L 127.0.0.1 -N reports "Anonymous login successful" and then lists nothing — correct and intended under map to guest = never, but it reads like a failure. If a health probe is ever built on this command, it must not treat the empty list as an error.
  • 192.168.0.180 FAILED in the felhom-pve neighbour table is the clean, one-line proof that the box is off the DooPlex segment — a handy check for the rest of the trip.
  • No secrets appear in this document: the household SMB password and the controller login password are stored out-of-band.

Addendum — 2026-07-20, controller v0.151.0 (S-1, S-2, S-4-core, S-5 closed)

Shipped and live on demo guest 9201 the same day this diagnosis was written. Controller commit badf17b, image gitea.dooplex.hu/admin/felhom-controller:0.151.0, Up (healthy).

# Status What changed
S-1 FIXED v0.151.0 The idle->running coercion is gone from sharingStatusHandler. The duty it was written for — "liveness must never be contradicted by a missing job" — was always discharged by the running LEVEL field beside it, and is now pinned by its own named regression test instead of by a phase-channel side effect. sharing.html's <script> block is byte-identical to v0.150.0: the client contract was proven fixed server-side, not papered over.
S-2 SHIPPED v0.151.0 New „Csatlakozás a megosztáshoz" card: the Windows form, the Mac form, and the direct smb://<IP> fallback. Rendered only while sharing is enabled; the address line is absent whenever no address can be derived.
S-4 CORE SHIPPED v0.151.0 consumeIfRunning serves a terminal running exactly once, so a REAL bring-up cannot re-arm the reload on the page it just caused — without it the loop returns after every future image update. failed, needs_password and in-flight phases are never consumed. The unified async-job feedback layer stays the roadmap item (R-45); this is the minimal contract fix, not the framework.
S-5 SHIPPED v0.151.0 The address is derived per render from the samba container's netns and stored nowhere — no settings key, no memo, no package-level cache. Deriving it through the samba container also avoids the trap that made the card necessary: the controller is on a docker bridge, so net.InterfaceAddrs() would answer 172.x.
S-3 OPEN Unchanged — still no mDNS/Bonjour advertisement, and still awaiting the two Mac-side probes from the matrix above (smbutil lookup FELHOM, dns-sd -B _smb._tcp) before an image change is justified. Deliberately out of scope for v0.151.0.

Live validation (endpoint-level — the exact endpoints the UI invokes; no browser on DooPlex)

GET /sharing/status  x3   ->  {"data":{"error":"","phase":"idle","running":true},"ok":true}   (all three)
                              was {"phase":"running","running":true} on every call before the fix
GET /sharing              ->  connect card present; smb://FELHOM and smb://192.168.0.104
                              (192.168.0.104 = the guest's real DHCP lease, derived, not a literal)

The phase:"idle" triple is the server-side proof that the loop's trigger is gone: the client's reload path is reachable only from phase === "running". The human confirmation that the page sits still is Viktor's to give — a browser is the one method not available here.

Tests: 23/23 packages green (twice). Three red-proofs recorded — reinstating the coercion, deleting the serve-once clear, and memoizing the derived address each turned the corresponding test red, then were restored.


Addendum 2 — 2026-07-20, controller v0.152.0 + felhom-samba 1.1.0 (S-3 closed)

The finding that redirected the fix: macOS asks, is answered correctly, and ignores it

Addendum 1 recorded S-3 as "no mDNS, and macOS probably no longer does NetBIOS". The second half was wrong, and a packet capture on the box settled it. On a bare smb://FELHOM from the operator's Mac (192.168.0.11):

11:18:30.222080  192.168.0.11.52844 > 192.168.0.255.137   UDP len 50   (NBNS query)
                 name = EGEFEMEIEPENCACACACACACACACACACA → "FELHOM" + 9 pad + suffix 0x20
                 i.e. FELHOM<20>, the File Server Service suffix — exactly right for SMB
11:18:30.222220  192.168.0.104.137 > 192.168.0.11.52844   UDP len 62   (NBNS response, +140 µs)
                 flags 0x8580 → response, authoritative, RCODE=0 (no error), ANCOUNT=1
                 TTL 259200, NB_FLAGS 0x0000 (unique, B-node), RDATA 192.168.0.104

A textbook-correct positive answer, delivered in 140 microseconds — and macOS never opened a TCP connection. It retried once at 11:18:30.718, was answered again, and gave up. Sixteen seconds later, at 11:18:46, the same Mac queried FELHOM.local over mDNS and went straight to 445.

So NetBIOS on macOS feeds legacy browsing, not smb:// URL resolution. The bare smb://<NÉV> cannot be made to work from a Mac by any change on our side, and nmbd — the thing the R-6/S4b spike was right to insist on — was never the broken part. It is exactly what serves Windows.

Investigated and dismissed: the box answers twice per broadcast query, because nmbd holds 0.0.0.0:137, 192.168.0.104:137 and 192.168.0.255:137, and a subnet broadcast is delivered to two of them. That is standard Samba behaviour, and a duplicated correct answer is still a correct answer that macOS declined to use.

The proven matrix

Client Working form Served by
Windows \\FELHOM nmbd (+ wsdd for the Network view)
macOS smb://FELHOM.local avahi/mDNS — new in samba 1.1.0
Any smb://192.168.0.104 direct

Shipped

  • felhom-samba 1.1.0 — avahi + dbus added. avahi-daemon.conf and the _smb._tcp service file are templated from FELHOM_SERVER_NAME at entrypoint, so renaming the server re-advertises under the new name. Static service file rather than smbd's multicast dns register, because smb.conf is bind-mounted READ-ONLY and owned by the controller's renderer — and it allows publishing _device-info._tcp too. Both daemons are non-fatal on failure: a discovery gap must not turn into a sharing outage.
  • controller v0.152.0 — the connect card's Mac line was smb://<NÉV>, i.e. the one form that provably cannot work. Now smb://<NÉV>.local. The Windows line stays the flat \\<NÉV>. This was a copy defect introduced by v0.151.0 the same day, not merely a missing feature.

Live validation (endpoint-level + on-the-wire from a second machine)

mDNS from felhom-pve (shipped image, not the spike):
  FELHOM.local (A)     -> A 192.168.0.104              (from 192.168.0.104)
  _smb._tcp (browse)   -> PTR FELHOM._smb._tcp.local   (from 192.168.0.104)

daemons in felhom-samba:1.1.0:  smbd · nmbd · wsdd · dbus-daemon · avahi-daemon: running [FELHOM.local]

GET /sharing  ->  smb://FELHOM.local   smb://192.168.0.104   \\FELHOM
GET /sharing/status x2 -> phase:"idle", running:true  (v0.151.0 contract still holding)

Operator confirmation from the Mac: smb://FELHOM.local connects and prompts for credentials. (Username is felhom, the household account — the Mac prefills the local account name, which will be rejected.)

STILL OPEN — Finder-sidebar discovery

Not claimed as shipped. The _smb._tcp record is published and answers browse queries on the wire, but FELHOM did not appear in the operator's Finder sidebar. That Finder window showed no Network/Bonjour section at all, which points at Finder Settings → Sidebar → Locations rather than at the box — but it is unverified either way, so it stays open. Next step is one look at that setting, and dns-sd -B _smb._tcp on the Mac to confirm the advert is seen client-side. Zero-typing discovery was the original ambition of the R-6 spike and is still not demonstrated.