Congestion Oversubscription and Credit Stall events reported in /var/log/vmkernel.log on ESXi Server as well as vCenter Server
search cancel

Congestion Oversubscription and Credit Stall events reported in /var/log/vmkernel.log on ESXi Server as well as vCenter Server

book

Article ID: 390100

calendar_today

Updated On:

Products

VMware vSphere ESXi VMware vSphere ESXi 5.0 VMware vSphere ESXi 5.5 VMware vSphere ESXi 5.x - View VMware vSphere ESXi 6.0 VMware vSphere ESXi 7.0 VMware vSphere ESXi 8.0

Issue/Introduction

Symptoms 

  • Existing VMFS datastores intermittently disappear or become unmounted from specific ESXi hosts within the cluster.
  • In the vSphere Client, the storage device (LUN) is visible under Storage Devices, but the Datastore column shows “Not Consumed.”
  • Running esxcfg-volume -l returns no output, indicating that no mountable volumes are detected.
  • Rescanning storage on an ESXi host fails and may see error "Virtual Reset Issued".
  • Often this results in noticeable Storage Latency in regards to I/O Operations

Messages are displayed which are similar to:

Frequent PowerOn Reset Unit Attentions are occurring on path vmhba0:C1:T0:L0. This may indicate a storage problem. Affected device: naa.###################. Affected datastores:vmdatafiles
cpu0:2659)ScsiCore: 1460: Power-on Reset occurred on vmhbaX:C0:T2:L0
cpu7:2055)NMP: nmp_ThrottleLogForDevice:2318: Cmd 0x2a (0x41244038d380, 2056) to dev "naa.###################" on path "vmhbaX:C0:T1:L0" Failed: H:0xb D:0x0 P:0x0 Possible sense data: 0x0 0x0 0x0. Act:NONE

 

/var/log/vmkernel.log log file  reports SCSI warnings against the paths for one or more devices, which indicate permanent device loss.

logs may report H:0x1 SCSI code ("no connection"), or logical unit not supported (H:0x0 D:0x2 P:0x0 Valid sense data: 0x5 0x25 0x0), or logical unit not accessible (H:0x0 D:0x2 P:0x0 Valid sense data: 0x2 0x4 0xa). 

esxcli storage san fc list will 

FcDevice:
  Adapter: vmhbaX
  Port ID: 000000
  Node Name: 20:##:##:24:##:18:#7:12
  Port Name: 21:##:##:24:##:18:#7:12
  Speed: 8 Gbps
  Port Type: NPort
  Port State: ONLINE
  Model Description: HPE SN1100Q 16Gb 2p FC HBA
  Hardware Version:BK3210407-20  F
  OptionROM Version: 3.68
  Firmware Version: 9.15.05 (d0d5)
  Driver Name: qlnativefc
Error getting field DriverVersion

In the Logs of one or more of the connected ESXi Host(s) one or more of the following errors are observed:
 
  • /var/log/vmkernel.log
YYYY-MM-DDTHH:MM:SS.ZZ In(182) vmkernel: cpu36:##)StorageFPIN: 1276: Report FC FPIN Congestion Oversubscription event (hostWWPN ##### tgtWWPN ##### to vobd. ## events have occurred since last report.
YYYY-MM-DDTHH:MM:SS.ZZ In(182) vmkernel: cpu52:##)StorageFPIN: 1276: Report FC FPIN Congestion Credit Stall event (hostWWPN #####] tgtWWPN #####) to vobd. ## events have occurred since last report.
YYYY-MM-DDTHH:MM:SS.ZZ In(182) vmkernel: cpu52:##)StoragePath: 5394: Calling MPP NMP for link event 2 on adapter vmhba## (hostWWPN=##### targetWWPN=##### targetNum = 12)
 
 
  • /var/log/vmkwarning.log

YYYY-MM-DDTHH:MM:SS.123Z Wa(180) vmkwarning: cpu52:123456)WARNING: lpfc: vmhba## lpfc_els_rcv_fpin_cgn:7266: 4657 FPIN CONGESTION WARNING Notification type Credit Stall (x2) Event Duration 10000 mSecs.

YYYY-MM-DDTHH:MM:SS.ZZ Wa(180) vmkwarning: cpu1:##)WARNING: NMP: nmpHandleLinkEvent:3998: Marking path vmhba## flaky on link event 2 with timeoutMS = 20000 flakyMarkTC = ####, reEvalFlakyPathTime = 20000

 

  • /var/log/vobd.log

YYYY-MM-DDTHH:MM:SS.ZZ In(14) vobd[#####]:  [HardwareCorrelator] ###: [vob.hardware.fpin.fc.congestion.creditstall] FPIN FC credit stall congestion: Host WWPN ##### , target WWPN #####.

YYYY-MM-DDTHH:MM:SS.ZZ In(14) vobd[#####]:  [HardwareCorrelator] ###: [esx.problem.hardware.fpin.fc.congestion.creditstall] FPIN FC credit stall congestion: Host WWPN##### , target WWPN #####.

YYYY-MM-DDTHH:MM:SS.ZZ In(14) vobd[#####]:  The event ([esx.problem.hardware.fpin.fc.congestion.oversubscription] FPIN FC oversubscription congestion: Host WWPN #####, target WWPN #####.) was sent immediately to hostd;
YYYY-MM-DDTHH:MM:SS.ZZ In(14) vobd[2098149]:  [HardwareCorrelator] ###: [vob.hardware.fpin.fc.congestion.oversubscription] FPIN FC oversubscription congestion: Host WWPN #####, target WWPN #####

 

In addition, one or more of the following errors are observed:

  • /var/log/vmkwarning.log

YYYY-MM-DDTHH:MM:SS.ZZ Wa(180) vmkwarning: cpu33:##)WARNING: VMW_SATP_ALUA: satp_alua_getTargetPortInfo:190: Could not get page 83 INQUIRY data for path "vmhba##" - Transient storage condition, suggest retry (195887294)

YYYY-MM-DDTHH:MM:SS.ZZ Wa(180) vmkwarning: cpu38:##)WARNING: ScsiDeviceIO: 1781: Device ########## performance has deteriorated. I/O latency increased from average value of ##### microseconds to #### microseconds.

 

Environment

VMware vSphere ESXi

Cause

FPIN (Fabric Performance Impact Notifications) capability was added in ESXi 8.0 U2 to be able to better understand fabric related issues/events. This module will also print to /var/log/vmkernel.log when there are fabric events happening. The events that FPIN tracks and will report on are:

  • Link Integrity
  • Delivery
  • Congestion
  • Peer Congestion

The ESXi Host is a recipient of these notifications, not the source.

These events are triggered by a "Slow Drain" condition in the SAN fabric.

The storage Fabric Switch detects that a Target port or ISL (Inter-Switch Link) is failing to return Buffer-to-Buffer (B2B) credits, creating backpressure.

Common triggers include:

  • Failing SFP modules or damaged fiber optic cables
  • Oversubscribed storage ports or overloaded storage processors
  • Misconfigured port speeds across the fabric.

Resolution

Perform these steps to identify and mitigate the FC Fabric-Level Congestion:
 
1.) Engage SAN/Fabric Vendor: 
Request that the SAN/Fabric team inspect switch logs for tim_txcrd_z (Time at zero transmit credits) or c3_timeout (Class 3 frame drops) at the time of the event.
 
2.) Verify Physical Health: 
Inspect the SFPs and fiber cables connected to the impacted vmhba ports identified in the logs. Replace any components showing physical degradation.
 
3.) Balance Load:
Use vSphere DRS rules to distribute I/O-intensive virtual machines across more ESXi Hosts to prevent overwhelming individual fabric ports.
 
Ensure storage devices are set to the Round-Robin (VMW_PSP_RR) policy to evenly distribute traffic.
 
5.) Review Fibre Channel Firmware/Drivers:
Ensure HBAs are running the recommended firmware/driver versions as documented in the Broadcom's compatibility guide.
Download Updates: If a defect is confirmed by Broadcom support, ensure you are running the latest version.

 

Additional Information

While Round robin policy is best practice, switching a live system from a Fixed policy to RR can cause momentary I/O blips.
It is recommended to perform this during a Maintenance window if paths are currently unstable.