SAN Storage Guide
Flat isometric illustration of a hot pink three-unit server stack with blank white label plates and mesh vents on a grey pad with pink link nodes.
performance

SAN High Latency: Where to Look First

Trace SAN high latency through host queues, buffer-credit starvation and ISL oversubscription before checking array cache, rebuilds and thin pools.

By SAN Storage Guide Editorial · ·Updated · 7 min read

“The SAN is slow” is a conclusion, not a symptom, and it is wrong often enough that starting there wastes the first hour. A block I/O crosses an application, a filesystem, a host queue, an adapter, a fabric, an array port, an array cache and finally some media. Latency is measured at the top and blamed at the bottom, and the layer that actually queued the request is usually somewhere in between.

Start a SAN high latency investigation at the host queue, then check multipath selection, Fibre Channel credit starvation and shared inter-switch links. Correlate each symptom with the same time window before moving to the array. The access-control background is in SAN zoning and LUN masking explained.

Turn the complaint into a measurement

Before touching a switch, pin down four things: which host, which volume, what time window, and what number moved. “Slow” that means average latency rose from 0.4 ms to 1.1 ms is a different investigation from “slow” that means a p99 spike to 400 ms twice an hour.

Then check whether the symptom is latency, throughput, or a queue, because they are not interchangeable. The relationship is arithmetic: outstanding I/O equals throughput multiplied by latency. If a host is issuing more concurrent requests than before, latency rises even though nothing broke. A workload that doubled its queue depth and doubled its latency is behaving exactly as the maths says it should, and the array is not at fault.

The single most useful discipline is to measure at both ends over the same interval. If the host reports 8 ms service time and the array reports 0.6 ms for the same volume in the same minute, the missing 7.4 ms is above the array, and the array is where most teams spend the outage.

Queue-depth saturation: host and target port limits

Start where the request starts. On Linux, iostat -x gives per-device average request time and average queue size; compare the multipath device against the individual paths beneath it. Windows exposes the equivalent through logical disk latency counters.

Two host-side conditions produce array-shaped symptoms:

Queue depth exhaustion. Every layer has a queue limit: a per-LUN limit, a per-adapter limit, and a limit on the array’s target port. When the array’s port queue fills, it returns a TASK SET FULL or QUEUE FULL status, and the host responds by throttling. The visible result is latency, and the invisible cause is that too many initiators are pointed at one target port. This is the classic consolidation failure: nothing changed on the array, four more hosts were zoned to the same port, and everyone got slower at once.

A saturated single path. If multipath is configured but all I/O is riding one path, the host sees adapter-level queueing while the fabric looks idle. Check the per-path counters, not just the aggregate.

Multipath, which is wrong more often than it looks

multipath -ll reporting all paths active is not the same as all paths working. Three specific misconfigurations produce steady, unexplained latency:

Path grouping inverted on an ALUA array. Arrays that present asymmetric access mark some paths active/optimised and others active/non-optimised. If the path groups are set so that non-optimised paths carry production I/O, every request takes an internal detour across the array’s controller interconnect. Nothing errors. Everything is slower, permanently. This is the highest-value single check in the entire list.

Path selector unsuited to the topology. Compare the configured selector with the array’s supported host profile. The Linux kernel’s service-time selector estimates service time from in-flight I/O size and configured relative throughput; it does not automatically benchmark every path. A supported selector and correct path groups matter more than choosing a policy by its name.

Paths that are not independent. Two paths that traverse the same switch, the same inter-switch link or the same array controller are two names for one path. Draw the topology and confirm that no single device appears on every route. This matters for availability and it also matters here, because a shared bottleneck makes both paths slow simultaneously and hides itself by looking symmetric.

The fabric: start at the port error counters

On Fibre Channel, go to the port error counters first, and note that they are cumulative: clear the baselines, wait a defined interval, and read them again. Rising CRC errors, encoding errors, or loss-of-sync events on a specific port point at the physical layer — an optic, a patch lead, a dirty connector — and no amount of array tuning will help. A single degrading optic that corrupts a small fraction of frames generates retries and produces exactly the intermittent-latency profile that gets blamed on the array.

Buffer-credit starvation: Fibre Channel slow-drain devices

Fibre Channel transmitters wait when receiver credits are unavailable. A slow-drain device can therefore cause back-pressure beyond its own link. Correlate transmit-wait duration and slow-port events with affected host paths. Cisco’s slow-drain guidance distinguishes how often credits reach zero from how long transmission is blocked; a single instantaneous zero-credit reading is insufficient to diagnose persistent starvation.

Where supported, fabric performance-impact notifications provide additional congestion evidence. Collect the switch’s documented counters over the incident interval and check downstream ports before changing buffer allocations. Buffer capacity, link requirements and slow receivers are different issues.

On NVMe/TCP or iSCSI, inspect interface discards, pause frames and TCP retransmissions over the same interval. Small packets succeeding does not prove the path supports the intended MTU. Verify the network’s supported frame size end to end. For protocol-specific login and path configuration, see What is an iSCSI LUN? Targets, initiators and MPIO on iSCSI Hub.

An inter-switch link (ISL) carries traffic from multiple edge ports. Compare the traffic crossing it with its available bandwidth during the latency window, including the capacity remaining when one member of a port channel is unavailable. Several lightly loaded host ports can still converge on one busy link.

Cisco’s congestion guide separates transmit overutilization from slow-drain behaviour. Check both utilization and credit-wait evidence: a blocked downstream device can cause queuing without sustained line-rate traffic. Map affected hosts to the shared ISL before deciding whether to redistribute traffic, add supported link capacity or investigate a slow receiver.

The array: cache, background jobs and thin pools

Only now is the array a reasonable suspect. Four conditions account for most of it:

Cache saturation. Write cache absorbs bursts and then has to destage to media. When the destage rate cannot keep up, latency rises sharply rather than gradually, because the array switches from cached to write-through behaviour. The signature is a cliff, not a slope.

A background job. RAID rebuilds, capacity rebalancing, snapshot consolidation and replication catch-up all consume backend bandwidth. Correlate the latency window against the array’s job history before anything else, because it is a two-minute check that closes a meaningful share of cases.

Thin pool pressure. A nearly full thin pool degrades before it fails, and allocation-based dashboards do not show it. Monitor actual pool consumption. Watching allocated capacity while the pool fills is how thin provisioning turns into an outage.

Noisy neighbours. One volume’s workload is served from the same pool and the same controllers as everyone else’s. Per-volume latency across the array, sorted, usually identifies the culprit in one look.

When it is not the SAN

If storage metrics remain within the workload’s baseline, inspect application queues and changes in request rate or I/O size. The SAN LUN Capacity Calculator estimates capacity and aggregate FC bandwidth for a planning check. It cannot diagnose latency. For a transport evaluation, compare NVMe-oF vs iSCSI vs Fibre Channel using the same workload and recovery requirements. If the workload wanted an arbitrated share rather than a volume the host owns outright, the SAN vs NAS block storage split is the wider question behind the complaint.

Triage order, condensed

  1. Define the symptom: host, volume, window, metric.
  2. Compare host-reported latency against array-reported latency for the same volume and interval.
  3. Check host queue depth and whether QUEUE FULL is being returned.
  4. Verify multipath path grouping, selector and genuine path independence.
  5. Clear fabric error counters, wait, re-read; look for CRC and sync errors on specific ports.
  6. Look for congestion and slow-drain signatures, or TCP retransmits and MTU mismatch on Ethernet.
  7. Correlate against array background jobs and cache destage behaviour.
  8. Check thin pool consumption as a percentage of real capacity.
  9. If everything agrees, size the fabric for the workload it now has.

Working the list in order is slower for the first ten minutes and faster for everything after that. Starting at step 7 of that list because the array is the thing with a support contract is how outages become long ones.

Sources

  1. dm-service-time — Linux kernel documentation
  2. multipath-tools — device-mapper multipath userspace
  3. INCITS T11 — Fibre Channel Interfaces
  4. Fibre Channel Industry Association
  5. Cisco MDS 9000 Interfaces Configuration Guide — Congestion Management
  6. Cisco — Slow-Drain Device Detection, Troubleshooting, and Automatic Recovery
#san #fibre-channel #multipathing#storage-performance#troubleshooting

Related