When a Fibre Channel (FC) storage area network (SAN) path disappears, the temptation is to bounce a host bus adapter (HBA), reset a switch port or reactivate zoning. Resist it. Those actions may restore service, but they can also erase the best clue: whether the fault lies in the host, physical link, fabric or storage presentation layer. Work the path in order and change one thing at a time.
First establish the scope. Is one logical unit number (LUN) missing from one server, are all LUNs missing from one host port, or is one of several paths degraded across a cluster? A single-path problem usually points to an HBA, optic, cable, switch port or one fabric. Host-wide visibility trouble is more likely to involve zoning, an incorrect World Wide Port Name (WWPN), array-side masking or host multipathing. If unrelated hosts lose access through the same fabric, stop chasing individual servers and investigate the shared switch, inter-switch link or storage target ports.
Safe first checks
- Protect the working path. Confirm multipathing, the host configuration that maintains multiple storage paths, is healthy and at least one independent path remains available before touching a production FC port.
- Record the symptom and time. Capture host events, HBA logs, switch logs, port state and error counters before clearing counters or reseating optics.
- Compare Fabric A with Fabric B. A healthy redundant path is the best control sample available in a live SAN.
- Trace the actual path. Record the host HBA port WWPN, switch port, virtual SAN (VSAN) or fabric, target port WWPN, array controller and affected LUN.
- Check physical state before configuration. A device can’t zone, log in or discover storage over a link that is down or flapping.
- Make one reversible change. Swap or move one known-good component where redundancy permits; don’t alter zoning, firmware and cabling at once.
First, decide which layer has failed
An FC SAN path has distinct stages. The HBA must see link. The attached switch port must come online at a mutually supported speed. Initiator and target must log in to the fabric. Zoning must permit communication within the correct fabric or VSAN. The array must present the correct LUNs to the host initiator, and the operating system must discover and claim those paths.
The sequence matters. A missing LUN isn’t automatically a zoning fault. If the host WWPN is absent from the fabric login database, zoning isn’t the immediate problem. Equally, a healthy link and successful fabric login don’t prove the array has mapped a LUN to that WWPN.
On Cisco MDS-based fabrics, show flogi database is the useful dividing line: it shows whether an initiator or target has completed Fabric Login and received an FCID, the fabric-assigned FC address. Cisco’s validation guidance pairs it with show zoneset active to confirm that the relevant logged-in WWPNs belong to the active zone set, rather than merely existing in an edited but inactive configuration. Cisco’s FC validation example is useful for the troubleshooting logic even where the platform differs.
On Brocade-based fabrics, start with switch and port status, name-server visibility, the active configuration and per-port counters. Command names vary by Fabric OS release and management tooling, but switchshow, portshow, porterrshow, nsshow and cfgshow are the familiar evidence set. Broadcom’s Fabric OS command reference documents portErrShow among the standard diagnostic commands; use syntax for the installed release rather than copying an old runbook blindly. Broadcom’s Fabric OS command reference is a useful starting point.
Physical faults: check link, speed and counters before moving anything
Before disabling a suspect port, establish the expected state: which paths are active, which are standby, how many paths each LUN should have, and whether any host is already running with reduced redundancy.
Physical-layer trouble is often mundane: a badly seated small form-factor pluggable (SFP) transceiver, a contaminated fibre end face, a damaged patch lead, an unsupported optic or a port that has been administratively disabled. It can also be intermittent. A green status light isn’t a clean bill of health.
Check both ends of the connection. Confirm the port is online, identify the negotiated speed and compare it with the corresponding port on the healthy fabric. An unexpected lower speed isn’t necessarily the direct cause of an outage, but it needs an explanation. A new link that comes up only after a forced speed change may be a compatibility or media problem disguised as a configuration fix. Automatic speed negotiation is usually preferable where attached equipment supports it cleanly; forcing a speed to hide a bad optic is poor engineering.
Read error counters as rates, not just totals. Capture the numbers, allow a meaningful period of activity, then capture them again. A counter that rose years ago and remains static is historical debris. A counter rising during the incident is evidence.
- CRC errors: frames are arriving with a failed cyclic redundancy check. Suspect the receive-side media path first: fibre, connectors, transceiver or the transmitting port. A CRC counter doesn’t prove the local switch port is defective; it shows that a bad frame was received, not conclusively where it was corrupted.
- Loss of signal or loss of sync: points more directly to optical signal, cabling, transceiver seating, compatibility or a port repeatedly dropping link.
- Invalid transmission words or encoding errors: can indicate a marginal optical path or port issue. Use the same controlled swap-and-compare approach.
- Link resets and repeated state changes: matter when they correlate with host I/O retries or path failures. One reset during planned maintenance is not a port flapping all day.
Broadcom’s current SANnav port-health material identifies CRC, loss of synchronisation, link failures and invalid transmission words as conditions monitored on FC ports. Its troubleshooting guidance says excessive CRC alerts should make operators suspect a transceiver or cable. That’s a sensible first hypothesis, not a licence to replace components at random. Broadcom’s SANnav troubleshooting guidance supports that distinction.
A controlled isolation test is straightforward in principle: move the suspect host connection to a known-good switch port using a known-good, compatible optic and patch lead, while retaining the other production fabric. If the failure follows the cable or optic, the fabric port is exonerated. If it remains on the original switch port despite known-good media, the port or its configuration becomes more credible. Record exactly what moved. Troubleshooting by memory gets unreliable quickly.
Link up is not the same as fabric access
A port can be physically online yet fail to complete the fabric login sequence. In a switched FC fabric, Fabric Login, usually abbreviated to FLOGI, is how an N_Port, an FC endpoint port, receives a fabric address. No FLOGI entry for the host WWPN means there’s no point debating LUN masking yet: return to the HBA, link, port mode and fabric attachment.
Check that the WWPN in the switch login database matches the physical HBA port you intended to connect. This catches a common operational error: zoning an adapter’s node WWN instead of its port WWPN, or using a virtual WWPN from a blade or virtualisation profile while the host is presenting another identity. Aliases make zoning readable but can conceal the mistake, particularly after an HBA replacement or profile change. During an incident, resolve the alias to the raw WWPN.
On a Cisco fabric, verify the correct VSAN as well. A VSAN is an isolated logical FC fabric carried by the same switching estate. A healthy device can be invisible to its intended peer if it has joined the wrong VSAN. Cisco documentation notes that an E_Port may be isolated because of a port-VSAN mismatch; don’t assume an inter-switch link carries the fabric you expect. Cisco’s VSAN troubleshooting guide covers the failure mode and verification approach.
For sporadic login failures, correlate switch events with the HBA’s driver and firmware log. The host may report link down, remote-port loss, repeated login attempts or a Small Computer System Interface (SCSI) transport reset before storage software reports a failed disk. Broadcom’s Emulex documentation also advises investigating the physical FC connection when repeated link events occur. The Emulex driver troubleshooting reference is dated, but the underlying advice remains sound: treat repeated link events as a physical-path investigation until evidence says otherwise.
When the host and array can see the fabric but not each other
Zoning is an access-control boundary within the FC fabric, not an optional tidying exercise. Check the active configuration, not merely the defined zone database. Confirm that the host initiator port WWPN and intended storage target port WWPN are in the same active zone on the relevant fabric or VSAN.
Single-initiator zoning remains the least troublesome operating model: one host initiator in a zone with the necessary target ports. It makes incidents easier to read, constrains accidental host-to-host visibility and avoids turning the fabric into a broad discovery domain. Storage vendors differ on whether one target per zone is required or multiple target ports are appropriate, so the array’s supported host configuration should win over generic convention. NetApp, for example, recommends zones containing one initiator and one or more target logical interfaces, and dual-fabric zoning to avoid a single component failure becoming an outage. NetApp’s current FC zoning guidance explains that model.
Don’t make zone edits during a live investigation until you have proved the expected WWPNs are logged in and the desired zone set isn’t already active. The usual bad fix is adding a broad catch-all zone, restoring visibility and leaving an unreviewed access-control mess behind. If a temporary diagnostic zone is unavoidable, document it, activate it through the formal change process, test the narrow hypothesis, then remove it.
Visible targets, missing LUNs: move to the array and host
If FLOGI and zoning are correct but the expected disk is absent, the remaining fault domain is normally array presentation or host discovery. On the array, verify that the LUN exists, is online and is mapped to the correct host object or initiator group, and that the group contains the current initiator WWPN. Check controller ownership, target-port health and whether the host personality or operating-system profile meets the vendor’s configuration requirements.
The distinction is fundamental: zoning permits a host to reach a target port; it doesn’t grant access to a LUN. NetApp describes initiator groups as tables of FC host port names that control which initiators can access which LUNs. Other array suppliers use different language—host groups, host objects or mappings—but the diagnostic question is the same. NetApp’s SAN administration documentation describes the separation between FC connectivity and LUN mapping.
Only then should you rescan the host. A rescan is discovery, not repair. It can’t discover a LUN the array hasn’t presented, and repeated scans can complicate an already noisy incident record. On Linux, inspect multipath first with multipath -ll, then check remote-port state under sysfs, the Linux virtual filesystem that exposes device information, if a path is questionable. Red Hat documents that an FC remote port may be Online or Blocked; after link loss exceeds the driver’s configured loss timeout, I/O can fail and devices may be removed. Red Hat’s RHEL 9 Fibre Channel guidance is useful because it separates a temporary transport disruption from a stable storage configuration fault.
On Windows, review Event Viewer, the HBA utility and MultiPath I/O (MPIO) state before trying to bring disks online. Microsoft recommends recording events and recent hardware or zoning changes, and specifically identifies missing paths as a reason to verify zoning, cabling and HBA state on both the server and SAN. Commands such as mpclaim -s -d, alongside the MPIO control panel or PowerShell cmdlets, can show whether Windows sees the storage device but has lost paths. Microsoft’s current MPIO troubleshooting guidance is the right reference for Windows-specific recovery work.
Intermittent paths are usually more dangerous than missing paths
A path that is cleanly down is inconvenient but comparatively easy to isolate. An intermittently failing path can create I/O retries, failovers, latency spikes and application errors while dashboards still show the LUN online through another route. Treat it as a production fault even if multipathing has prevented an obvious outage.
Compare host timestamps with switch counters and array events. If the host reports retries as one FC port increments CRC or loss-of-sync counters, that is a strong physical-path lead. If both fabrics have clean counters but a particular target port disappears from the array’s audit trail, investigate the storage controller and presentation configuration. If failures move with a virtual machine or a particular operating-system build rather than the physical path, look at the HBA driver, multipathing software and supported firmware combinations.
What to do next: build a path worksheet before the next incident. For every production host, record both HBA port WWPNs, switch ports, fabric or VSAN, storage target WWPNs, host group, LUN identifiers and expected multipath count. It isn’t glamorous documentation, but it turns a two-hour outage call into a sequence of checks.
Use redundancy to test safely, not as an excuse to be careless
Properly designed dual-fabric FC lets you isolate one path while service continues on the other. That’s valuable only when the surviving path has been verified, the workload can tolerate a failover and the operating system’s multipathing configuration is understood. A degraded cluster, a boot-from-SAN host or an array with asymmetric path rules may need a maintenance window even for what looks like a small test.
Before disabling a suspect port, establish the expected state: which paths are active, which are standby, how many paths each LUN should have, and whether any host is already running with reduced redundancy. Don’t reset both fabrics to test a theory. Don’t clear counters until they have been captured. And don’t update HBA firmware during an outage merely because it is old; compatibility matrices and change control exist for a reason.
The marketing version of FC resilience is that redundancy makes faults invisible. In practice, redundancy buys time and evidence. Use both to identify the failed component, restore protection and remove the underlying weakness rather than celebrating that applications stayed up.
Sources and further reading
- Cisco: verify FLOGI and active zone sets on an MDS fabric
- Cisco: troubleshooting VSANs, domains and FSPF
- Broadcom: Brocade Fabric OS command reference
- Broadcom: SANnav troubleshooting mode and FC port health indicators
- NetApp: recommended FC and FCoE zoning configurations
- NetApp: SAN management overview and initiator groups
- Red Hat: using Fibre Channel devices in RHEL 9
- Microsoft: Windows Server MPIO troubleshooting guidance
Spot an error?
If something factual looks wrong, outdated or misleading, flag it here. Corrections are reviewed separately from normal article comments and reader questions.


