Unable to vMotion, clone, backup, or consolidate VMs on vSAN error "Checksum mismatch"
search cancel

Unable to vMotion, clone, backup, or consolidate VMs on vSAN error "Checksum mismatch"

book

Article ID: 397059

calendar_today

Updated On:

Products

VMware vSAN

Issue/Introduction

Symptoms:

  • Virtual Machine operations such as clones, backups, or consolidations fail.

  • Backup job fails with below error:

Processing VM# Error: Data error (cyclic redundancy check). Asynchronous request operation has failed. [requestsize = 4194304] [offset = 7643157495808] Failed
to upload disk '>' Agent failed to process method {DataTransfer.SyncDisk}.

  • Read errors are reported during the backups and below errors are observed in the vddk logs:

Failed to read 1,048,576 bytes from disk offset 7,643,158,544,384
VixDiskLib_Read failed: error code 1, Unknown error.
Closing VDDK ...
VixDiskLib_Close()
VixDiskLib_EndAccess()
VixDiskLib_Disconnect()
VixDiskLib_Exit()
Closing VDDK ... OK

Copied the VixDiskLib log to "logs/vixDiskLib-14848.log"
VixDiskLib_Read failed: error code 1, Unknown error.
Error: ErrorMessage { msg: "VixDiskLib_Read failed: error code 1, Unknown error." }

  • The vSphere Client displays a "Checksum mismatch" error during these operations.

     "Failed to copy source (/vmfs/volumes/vsan:/namespace/VM_12.vmdk) to destination (/vmfs/volumes/vsan:/namesapce/VM_12.vmdk): Checksum mismatch. "

Validation Steps: 

  • The /var/log/hostd.log file on the ESXi host will record a failed file copy operation resulting from a checksum mismatch, along with a corresponding event indicating an unrecoverable vSAN medium error on the underlying disk group component.

"Failed to copy source (/vmfs/volumes/vsan:/namespace/VM_12.vmdk) to destination (/vmfs/volumes/vsan:/namesapce/VM_12.vmdk): Checksum mismatch."

2025-04-28T06:00:34.299Z: [vSANCorrelator] 4631684299497us: [vob.vsan.dom.unrecoverableerror] vSAN detected an unrecoverable medium or checksum error for component ########-####-####-####-########### on disk group ########-####-####-####-###########.

Failed waiting for data. Error 195887167. Connection closed by remote host, possibly due to timeout. 2025-07-03T19:48:49.447128Z Failed to copy source (/vmfs/volumes/vsan:#####.vmdk) to destination (/vmfs/volumes/vsan:###/####/####.vmdk): Checksum mismatch. Failed to copy one or more disks. A fatal internal error occurred. See the virtual machine's log for more details. 2025-07-03T19:48:51.337343Z vMotion migration [####] failed to read stream keepalive: Connection closed by remote host, possibly due to timeout

  • Further confirmation of the physical drive issue can be found in the /var/log/vobd.log of the host which records the exact CRC mismatch and the corresponding vSAN alerts for the impacted component and disk group:

    2023-01-08T05:32:02.042Z cpu7:2099333)WARNING: LSOM: LSOMReadVerifyChecksum:4397: Throttled: Checksum error detected on component ########-####-####-####-###########, comp offset 210520969216 (computed CRC 0x0 != saved CRC 0x81bf8868 (faked: Y))
    2023-01-05T19:03:27.576Z: [vSANCorrelator] 4820779458341us: [vob.vsan.dom.unrecoverableerror] vSAN detected an unrecoverable medium or checksum error for component ########-####-####-####-########### on disk group ########-####-####-####-###########.
    2022-10-23T03:15:13.877Z: [vSANCorrelator] 22532796272us: [vob.vsan.dom.errorfixed] vSAN detected and fixed a medium or checksum error for component ########-####-####-####-########### on disk group ########-####-####-####-###########.

  • /var/log/vmkernel.log will contain the below message when dealing with a failed vMotion

    vmkernel.log:2026-01-13T04:26:23.389Z Wa(180) vmkwarning: cpu36:44343895)WARNING: Migrate: 257: #########-#####-####-####-############# S: Failed: Checksum mismatch (0xbad003a) @0x420016adc88f
 

Environment

VMware vSAN

Cause

This issue occurs because medium errors (physical bad blocks) are present on the physical storage drive where the vSAN component with checksum errors resides.

When vSAN encounters these bad blocks, it records SCSI sense codes in the host logs indicating the hardware failure.

 

Cause Validation: 

  •  The following entries in the /var/log/vobd.log file can be used to identify the specific component and disk group UUIDs affected by the unrecoverable medium or checksum error

      2025-04-28T06:00:34.299Z: [vSANCorrelator] 4631684299497us: [vob.vsan.dom.unrecoverableerror] vSAN detected an unrecoverable medium or checksum error for component ########-####-####-####-########### on disk group ########-####-####-####-###########.
      2025-04-28T15:15:46.534Z: [vSANCorrelator] 4664996091018us: [vob.vsan.dom.unrecoverableerror] vSAN detected an unrecoverable medium or checksum error for component ########-####-####-####-########### on disk group ########-####-####-####-###########.
      2025-04-28T15:15:46.534Z: [vSANCorrelator] 4665059504699us: [esx.problem.vob.vsan.dom.unrecoverableerror] vSAN detected an unrecoverable medium or checksum error for component ########-####-####-####-########### on disk group ########-####-####-####-###########

  • To confirm the underlying hardware issue, review the /var/log/vmkernel.log file for repeated device command failures reporting valid sense data for medium errors on the physical disk

            2025-05-05T18:02:42.745Z cpu54:2098074)ScsiDeviceIO: 4167: Cmd(0x45dea00a8308) 0x28, CmdSN 0x6098bd81 from world 0 to dev "naa.################" failed H:0x0 D:0x2 P:0x0 Valid sense data: 0x3 0x11 0x0 Medium Error, LBA: 2520853716
      2025-05-05T18:02:43.284Z cpu30:2098080)ScsiDeviceIO: 4167: Cmd(0x45dea66d48c8) 0x28, CmdSN 0x6098bdbd from world 0 to dev "naa.################" failed H:0x0 D:0x2 P:0x0 Valid sense data: 0x3 0x11 0x0 Medium Error, LBA: 2520853716
      2025-05-05T18:02:43.821Z cpu34:2098082)ScsiDeviceIO: 4167: Cmd(0x45de788f2e88) 0x28, CmdSN 0x6098bdf0 from world 0 to dev "naa.################" failed H:0x0 D:0x2 P:0x0 Valid sense data: 0x3 0x11 0x0 Medium Error, LBA: 2520853716

  • Note: Common SCSI medium error codes that may be observed include:
      0x3 0x3 0x0 - PERIPHERAL DEVICE WRITE FAULT
     0x3 0x10 0x0 - ID CRC OR ECC ERROR
     0x3 0x11 0x0 - Unrecovered read error
     0x3 0x31 0x0 - Medium Format corruption

Resolution

Workaround:
 

To allow the backup to proceed , Temporarily disable the object checksum setting in the storage policy.

  1. In the vSphere Client, go to Menu > Policies and Profiles > VM Storage Policies.
  2. Select the existing policy (e.g., vSAN Default Storage Policy) and click Clone.
  3. Name the new policy (e.g., vSAN-No-Checksum).
  4. Under vSAN Rules > Advanced Policy Rules, set Object checksum to Disabled.
  5. Assign the new vSAN-No-Checksum policy to the affected VM disk.
  6. Change the storage policy of the VM back to Original Storage policy or vSAN default storage policy. 
 
For a permanent fix:

 

If medium errors are encountered in the data region of the disk and the disk wasn't failed out by vSAN you can do one of the following:

Using the preferred method, contact the hardware vendor and get the disk replaced.

    1. If Deduplication is not enabled:
      • Perform a pre-check by navigating to vSAN Cluster > Monitor > Data Migration Pre-check. Under Pre-check Data Migration For, select OBJECT, then go to the problematic host and select the problematic disk (as seen in the logs) that needs to be removed from the disk group and re-added.



         
        • Remove the failed disk with medium errors from the disk group and then add it back. Upon adding the disk back to the disk group, the bad blocks are automatically reallocated by the disk for non-use, so they don't get used again.

    2. If Deduplication is enabled:
      • Remove the disk group which contains the disk failing with medium errors from the host and recreate the disk group. Upon recreating the disk group the bad blocks are automatically reallocated by the disk for non-use, so they don't get used again.
        Note: This results in a resync to rebuild data due to the disk/disk group being removed and then recreated/added back.


Reference KB  

VMware vSAN disk encounters medium errors but not failed out by vSAN 

 

Additional Information

Enable alert in vCenter for vSAN checksum errors detected in the host logs

vSAN Disk Or Diskgroup Fails With Medium Errors