kernel: sd 4:0:1:0: [sdg] tag#678 task abort on host 4 WARNING: NMP: nmp_ThrottleLogForDevice:3868: H:0x8 D:0x0 P:0x0 . Act:EVAL. qlnativefcEhAbort:2745:SCSI command timeout counter incremented to 4ESXi 8.x
Oracle RAC
The issue is caused by transient storage unresponsiveness or SAN fabric latency.
In this scenario, a critical database I/O (such as a WRITE(10) from the Oracle LGWR process) is issued to the storage array but is not acknowledged. Because the Oracle RAC eviction threshold is typically shorter (~70 seconds) than the default ESXi storage driver timeout for Task Management Aborts (often 120 seconds for qlnativefc), the database cluster initiates a node eviction to protect data integrity before the hypervisor or HBA driver can recover the stuck I/O queue.
The presence of Error Handling (EH) Abort increments in the HBA driver logs without a subsequent LUN or Virtual Reset indicates that the HBA successfully cleared individual commands, but the duration of the stall was sufficient to crash the application layer.
Because the root cause resides in the physical storage layer or fabric, investigate the following areas: