Messages are displayed which are similar to:Frequent PowerOn Reset Unit Attentions are occurring on path vmhba0:C1:T0:L0. This may indicate a storage problem. Affected device: naa.###################. Affected datastores:vmdatafilescpu0:2659)ScsiCore: 1460: Power-on Reset occurred on vmhbaX:C0:T2:L0cpu7:2055)NMP: nmp_ThrottleLogForDevice:2318: Cmd 0x2a (0x41244038d380, 2056) to dev "naa.###################" on path "vmhbaX:C0:T1:L0" Failed: H:0xb D:0x0 P:0x0 Possible sense data: 0x0 0x0 0x0. Act:NONE
/var/log/vmkernel.log log file reports SCSI warnings against the paths for one or more devices, which indicate permanent device loss.
logs may report H:0x1 SCSI code ("no connection"), or logical unit not supported (H:0x0 D:0x2 P:0x0 Valid sense data: 0x5 0x25 0x0), or logical unit not accessible (H:0x0 D:0x2 P:0x0 Valid sense data: 0x2 0x4 0xa).
esxcli storage san fc list will FcDevice: Adapter: vmhbaX Port ID: 000000 Node Name: 20:##:##:24:##:18:#7:12 Port Name: 21:##:##:24:##:18:#7:12 Speed: 8 Gbps Port Type: NPort Port State: ONLINE Model Description: HPE SN1100Q 16Gb 2p FC HBA Hardware Version:BK3210407-20 F OptionROM Version: 3.68 Firmware Version: 9.15.05 (d0d5) Driver Name: qlnativefcError getting field DriverVersion
YYYY-MM-DDTHH:MM:SS.123Z Wa(180) vmkwarning: cpu52:123456)WARNING: lpfc: vmhba## lpfc_els_rcv_fpin_cgn:7266: 4657 FPIN CONGESTION WARNING Notification type Credit Stall (x2) Event Duration 10000 mSecs.
YYYY-MM-DDTHH:MM:SS.ZZ Wa(180) vmkwarning: cpu1:##)WARNING: NMP: nmpHandleLinkEvent:3998: Marking path vmhba## flaky on link event 2 with timeoutMS = 20000 flakyMarkTC = ####, reEvalFlakyPathTime = 20000
YYYY-MM-DDTHH:MM:SS.ZZ In(14) vobd[#####]: [HardwareCorrelator] ###: [vob.hardware.fpin.fc.congestion.creditstall] FPIN FC credit stall congestion: Host WWPN ##### , target WWPN #####.
YYYY-MM-DDTHH:MM:SS.ZZ In(14) vobd[#####]: [HardwareCorrelator] ###: [esx.problem.hardware.fpin.fc.congestion.creditstall] FPIN FC credit stall congestion: Host WWPN##### , target WWPN #####.
YYYY-MM-DDTHH:MM:SS.ZZ In(14) vobd[#####]: The event ([esx.problem.hardware.fpin.fc.congestion.oversubscription] FPIN FC oversubscription congestion: Host WWPN #####, target WWPN #####.) was sent immediately to hostd;
YYYY-MM-DDTHH:MM:SS.ZZ In(14) vobd[2098149]: [HardwareCorrelator] ###: [vob.hardware.fpin.fc.congestion.oversubscription] FPIN FC oversubscription congestion: Host WWPN #####, target WWPN #####
In addition, one or more of the following errors are observed:
YYYY-MM-DDTHH:MM:SS.ZZ Wa(180) vmkwarning: cpu33:##)WARNING: VMW_SATP_ALUA: satp_alua_getTargetPortInfo:190: Could not get page 83 INQUIRY data for path "vmhba##" - Transient storage condition, suggest retry (195887294)
YYYY-MM-DDTHH:MM:SS.ZZ Wa(180) vmkwarning: cpu38:##)WARNING: ScsiDeviceIO: 1781: Device ########## performance has deteriorated. I/O latency increased from average value of ##### microseconds to #### microseconds.
VMware vSphere ESXi
FPIN (Fabric Performance Impact Notifications) capability was added in ESXi 8.0 U2 to be able to better understand fabric related issues/events. This module will also print to /var/log/vmkernel.log when there are fabric events happening. The events that FPIN tracks and will report on are:
The ESXi Host is a recipient of these notifications, not the source.
These events are triggered by a "Slow Drain" condition in the SAN fabric.
The storage Fabric Switch detects that a Target port or ISL (Inter-Switch Link) is failing to return Buffer-to-Buffer (B2B) credits, creating backpressure.
Common triggers include: