SAN Storage Guide
Flat isometric illustration of a dark server chassis with four pink-lit drive bays, standing on a bright pink platform against a navy dotted background.
san-fundamentals

SAN Zoning and LUN Masking Explained

SAN zoning controls which ports communicate; LUN masking controls volume access. Learn how both fit Fibre Channel fabrics and independent paths.

By SAN Storage Guide Editorial · ·Updated · 8 min read

SAN zoning controls which Fibre Channel ports may communicate. LUN masking controls which logical units the storage array presents to a host. Configure both: being allowed to reach an array port should not imply permission to access every volume behind it.

A storage area network presents block devices to servers across a fabric. This guide explains the zoning, masking and path-design decisions for a homelab or small-business SAN, with Fibre Channel and NVMe-oF as the fabric context.

Block storage is not shared storage

Because the host owns the filesystem, two hosts writing to the same block device without a cluster aware filesystem will corrupt it. That follows from what a SAN storage area network is: the fabric hands the host a raw block device and leaves the filesystem, the locking and the consistency to it. NAS protocols like NFS and SMB arbitrate access on the server side, so concurrent clients are safe by design. A SAN does not do that for you. If more than one host must see the same LUN, you need either a clustered filesystem or a layer above that coordinates ownership.

Choosing a transport

Fibre Channel runs block traffic over a purpose built lossless fabric with its own switches and host bus adapters. It is expensive and it is predictable, which is why it persists in environments that cannot tolerate variance.

iSCSI carries SCSI commands over ordinary TCP/IP. The hardware is cheap and familiar, but the network stops being someone else’s problem. Put storage traffic on isolated VLANs or physically separate switching, and be deliberate about flow control and frame size. Sharing a congested general purpose network with storage is a frequent cause of unexplained latency complaints.

NVMe over Fabrics carries the NVMe command set across a network instead of SCSI. It removes translation overhead that made sense for spinning disks and makes little sense for flash. It can run over Fibre Channel, RDMA capable Ethernet, or plain TCP, and the fabric choice determines how much of the benefit you actually see.

Picking between the transports requires checking host support and recovery behaviour as well as bandwidth. NVMe-oF vs iSCSI vs Fibre Channel explains the command models and fabric choices.

Names you have to get right

Every SAN control in this article keys off a name, so the naming scheme is not clerical detail. It is the identifier that zoning, masking and audit logs all match on, and a typo in it is indistinguishable from a permissions failure.

Fibre Channel identifies endpoints by World Wide Name. Each adapter carries a node name and each of its ports carries a port name, both 64-bit values burned in by the manufacturer and written as eight colon-separated hex bytes. Because they are assigned by the vendor rather than by the site, two ports never collide, but they also carry no meaning: nothing in a WWPN tells you which server it belongs to. Keep an external map from WWPN to host and slot, and update it when an adapter is replaced, because a swapped HBA changes the name that every zone references.

For IQN naming, target discovery, initiator login and CHAP configuration, see What is an iSCSI LUN? Targets, initiators and MPIO on iSCSI Hub. Keep the host identifiers recorded in the fabric plan consistent with the array’s access mappings.

NVMe over Fabrics uses NVMe Qualified Names built on the same reverse-domain pattern, and presents namespaces rather than LUNs. The vocabulary differs but the control model does not: a namespace is still a block device that exactly one host should own unless something above it coordinates ownership.

How a host actually finds a LUN

Discovery is a separate step from access, and conflating the two is behind a lot of “the LUN is not there” tickets.

On a Fibre Channel fabric, check port login, name-server registration and the active zone set before checking array mappings. A working link light alone does not confirm target visibility. Also inspect the default-zone policy: access for devices outside explicit zones depends on that policy.

NVMe over Fabrics provides discovery controllers for learning about subsystems. Treat discovering a subsystem, connecting to it and receiving access to a namespace as separate checks. Record which check failed before changing fabric or array permissions.

SAN zoning vs LUN masking

Zoning happens in the Fibre Channel fabric. It restricts communication between ports according to the active zone set and default-zone policy. Prefer single-initiator zones where the array’s deployment guide calls for them, containing one host port and the target ports it needs.

LUN masking happens on the array. It decides which specific volumes a permitted initiator is allowed to use. Zoning does not select the LUNs assigned to a host, and masking does not replace the fabric’s communication policy.

ControlConfigured onQuestion it answers
SAN zoningFibre Channel switchesMay these host and storage ports communicate?
LUN maskingStorage arrayWhich volumes may this host or host group access?
Host multipathingHost storage stackWhich available path should carry I/O?

The failure that hurts is presenting a LUN already owned by another host. The new host sees an unformatted disk, an administrator formats it, and the original data is gone. Verify ownership before presenting anything.

Zone membership and enforcement are separate choices. WWPN-based membership identifies an endpoint; interface-based membership identifies a switch connection. Soft zoning filters name-server information, while hard zoning enforces traffic restrictions in the forwarding path. Hard zoning is not synonymous with interface-based membership: Cisco MDS supports hardware enforcement for WWN-based zones too. Its zoning guide documents the supported combinations. Record the membership scheme so an HBA replacement or cable move prompts the right review.

Keep the zone database under the same change control as the array. A fabric configuration is a live security control, not switch scratch space, and an exported copy of the active zone set taken before every change is the fastest route back when a change goes wrong.

Multipathing is mandatory, not optional

Every host should reach every LUN through more than one path: two HBAs or NICs, two fabrics, two array controllers. Multipath software on the host collapses those duplicate device nodes into one and handles failover. Skipping it means a single cable, optic or switch reboot takes an application down.

Two things go wrong here. First, multipathing is configured but never tested, so nobody discovers the path failover settings are wrong until an outage. Second, the paths are not genuinely independent, because both run through the same switch or the same physical route. Draw the topology and confirm no single device appears on every path.

A third failure is quieter than either: the paths all work, but the path grouping sends production traffic down the array’s non-optimised route, so everything is permanently slower and nothing ever errors. That check and the rest of the triage order are in SAN latency troubleshooting.

That quiet failure has a name. Most dual-controller arrays are asymmetric: a given volume is owned by one controller at a time, and reaching it through the other controller means the request is forwarded internally before it is served. The SCSI command set standardised by the T10 committee lets the array advertise which paths are the optimised ones, and multipath software is supposed to group paths accordingly and prefer the optimised group. When that advertisement is ignored, either because the host is using a generic device profile or because someone pinned a path group by hand, every read takes the long way through the array’s interconnect. Nothing errors, no alert fires, and the only symptom is latency that has always been slightly worse than the hardware should deliver.

Three settings decide how a path failure feels to the application, and all three are worth setting deliberately rather than inheriting. The path grouping policy decides whether traffic spreads across an entire group or pins to one path until it dies. The retry behaviour decides what happens when every path is gone: fail the I/O immediately and let a cluster manager react, or queue indefinitely and hope the fabric comes back, which keeps the data safe but can hang a host until it does. The failback behaviour decides whether a recovered path is used again automatically, which is usually what you want, or held back until an operator confirms the path is genuinely stable, which is what you want when a marginal optic is flapping. The device-mapper multipath documentation covers the per-array defaults, and a device with no matching entry falls back to generic behaviour that is safe but rarely optimal.

Test the failover rather than the configuration. Pulling one optic at a time while a synthetic write load runs, and confirming both that the paths drop out and that they come back into the right group, is the only check that distinguishes multipathing that works from multipathing that is merely configured.

Capacity planning

Provision for the workload profile, not just the total size. Thin provisioning is useful and it is also a way to oversubscribe an array into an outage, so monitor actual pool consumption rather than allocated capacity. Watch queue depth and latency together, since throughput numbers alone hide the moment a fabric starts queuing.

For a first pass at the numbers, the SAN LUN Capacity Calculator uses raw capacity, representative RAID ratios, FC port speed and initiator count to estimate capacity and aggregate bandwidth. Confirm the array’s actual reserves and supported path layout before purchasing equipment.

The order these controls should be built in

The four concepts above are not independent, and building them out of order is how a SAN acquires the failure modes it keeps for the rest of its life.

Transport first, because it fixes the hardware budget and decides who owns an incident: a Fibre Channel fabric is a separate estate with its own switches and skills, while iSCSI hands storage traffic to the network team whether or not they were told. Naming second, before any zone exists, because renaming a host after fifty zones reference its WWPN is a change window nobody wants. Zoning third, then masking, in that order, so that at no point is a volume reachable by a host that has not yet been given explicit permission to use it. Multipathing last, and tested, because a path layout designed after the fabric is built usually inherits a single shared switch that nobody notices until it reboots.

The single highest-value habit is drawing the topology before touching a switch, then confirming that no device appears on every path from any host to any volume. That drawing catches the shared-switch problem, the shared-optic problem and the single-controller problem in one pass, and it takes minutes. Recovering from any of the three after the array is in production takes a maintenance window and a conversation about downtime.

For the next design decision, compare NVMe-oF vs iSCSI vs Fibre Channel. If the fabric is reachable but slow, follow SAN high latency: where to look first before changing queue or path settings.

Sources

  1. RFC 7143 — Internet Small Computer System Interface (iSCSI) Protocol (Consolidated)
  2. NVM Express Specifications
  3. INCITS T11 — Fibre Channel Interfaces
  4. multipath-tools — device-mapper multipath userspace
  5. INCITS T10 Technical Committee — SCSI Interfaces
  6. Cisco MDS 9000 Fabric Configuration Guide — Configuring and Managing Zones
#san #fibre-channel #zoning#lun-masking#multipathing#nvme-of

Related