Mellanox nmlx5_core NICs Experience High Physical CRC Errors From Degraded Switch Ports
search cancel

Mellanox nmlx5_core NICs Experience High Physical CRC Errors From Degraded Switch Ports

book

Article ID: 443237

calendar_today

Updated On:

Products

VMware vSAN

Issue/Introduction

VMware ESXi hosts equipped with Mellanox Technologies ConnectX-6 Lx network adapters utilizing the nmlx5_core driver experience packet drops, throughput degradation, or link flapping. Interrogating the hardware counters reveals an active, rapid accumulation of physical layer frame check sequence errors and lane symbol errors.

Symptoms include:

  • Critical counts for rxCrcErrorsPhy and rxSymbolErrorsPhy across interface lanes.

  • Log snippets or command outputs mirror the following health state:
NIC: vmnic0
   rxCrcErrorsPhy: 290127
   rxSymbolErrorsPhy: 290108
   rxErrLane_0_Phy: 72531
   MAC Address: <REDACTED_MAC_ADDRESSES>
   PCI Device: <REDACTED_SECRETS>

Environment

VMware ESXi 8.0U3

Cause

Defective or degraded upstream physical network switch ports causing physical layer (Layer 1) signal attenuation and data corruption prior to host ingestion.

Resolution

  1. Schedule a maintenance window and place the affected ESXi host into Maintenance Mode via vSphere Client or SDDC Manager.

  2. Physically migrate the fiber optic patch cabling and transceivers from the degraded switch ports to known functional alternative ports on the upstream network switch.

  3. Coordinate with the network engineering team to administratively disable (shutdown) and flag the defective switch ports to prevent future workload allocations.

  4. Clear the host hardware interface statistics.

  5. Monitor the network interface statistics over a 24-hour baseline period to confirm that the error counters remain stationary at zero by running:

       esxcli network nic stats get -n vmnic0

Additional Information

https://knowledge.broadcom.com/external/article/401149/receive-length-errors-detected-on-mellan.html