Nutanix CVM machine appears as invalid state in vCenter with "corrupt heartbeat detected"
search cancel

Nutanix CVM machine appears as invalid state in vCenter with "corrupt heartbeat detected"

book

Article ID: 444372

calendar_today

Updated On:

Products

VMware vSphere ESX 8.x

Issue/Introduction

  • The CVM is displayed as Invalid in vCenter.
  • Unable to browse the CVM directory at /vmfs/volumes/DataStore_Name/CVM_Name/.
  • The df -h command was also hanging and did not return the expected filesystem information.
  • The "dmesg | less" logs on the ESXi host and identified the following error:
    2026-06-12T05:04:33.681Z cpu49:2098231)FS3: 662: and upload the dump by `voma -m vmfs -f dump -d /vmfs/devices/disks/t10.ATA_____SMC_VD__________________________________e4f57##########000000000:8 -D X`
    2026-06-12T05:04:33.681Z cpu49:2098231)FS3: 665: where X is the dump file name on a DIFFERENT volume
    2026-06-12T05:04:33.682Z cpu49:2098231)FS3: 476: FS3HB 0 0 0 0 00000000-00000000-0000-000000000000
    2026-06-12T05:04:33.682Z cpu49:2098231)FS3: 481: 0 0 0 0 0 00000000-00000000-0000-000000000000
    2026-06-12T05:04:33.682Z cpu49:2098231)FS3: 491: 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
    2026-06-12T05:04:33.682Z cpu49:2098231)FS3: 493: 0 0 0 0 0
    2026-06-12T05:04:34.940Z cpu73:3175332)HBX: 3089: 'NTNX-local-ds-##############-#': HB at offset 3584000 - Waiting for timed out HB:
    2026-06-12T05:04:34.940Z cpu73:3175332)  [HB state abcdef02 offset 3584000 gen 127 stampUS 967329270818 uuid 6a1674dc-#########-####-########c7bc jrnl <FB 1216000> drv 14.81 lockImpl 3 ip 10.172.###.###]
    ESC[7m2026-06-12T05:04:34.941Z cpu49:2098231)WARNING: HBX: 754: 'NTNX-local-ds-##############-#': HB at offset 0 - Volume 6323314c-#########-####-########9c7bc may be damaged on disk. Corrupt heartbeat detected:ESC[0m
    ESC[7m2026-06-12T05:04:34.941Z cpu49:2098231)WARNING:   [HB state 0 offset 0 gen 0 stampUS 0 uuid 00000000-00000000-0000-000000000000 jrnl <FB 0> drv 0.0]ESC[0m
    ESC[7m2026-06-12T05:04:34.941Z cpu49:2098231)WARNING: FS3: 636: VMFS volume NTNX-local-ds-##############-#/6323314c-#########-####-########c7bc on t10.ATA_____SMC_VD__________________________________e4f57##########000000000:8 has been detected corruptedESC[0m
    2026-06-12T05:04:34.941Z cpu49:2098231)FS3: 639: While filing a PR, please report the names of all hosts that attach to this LUN, tests that were running on them,
  • We tried running the smart test from the ESXi host on the M.2 drives where the CVM is present, however, it did not work:

    [root@ESXi-Host:~] esxcli storage vmfs extent list
    Volume Name                                 VMFS UUID                            Extent Number  Device Name                                                                   Partition
    ------------------------------------------  -----------------------------------  -------------  ----------------------------------------------------------------------------  ---------
    NTNX-local-ds-############-#                6323314c-#########-####-########c7bc              0  t10.ATA_____SMC_VD__________________________________e4f57########000000000          8
    OSDATA-#########-####-########c7bc          6a1674dc-#########-####-########c7bc              0  t10.ATA_____SMC_VD__________________________________e4f57########000000000          7
    [root@ESXi-Host:~] esxcli storage core device smart get -d t10.ATA_____SMC_VD__________________________________e4f57########000000000
    Error getting Smart Parameters: Cannot open device
    [root@ESXi-Host:~] esxcli storage core device smart get -d t10.ATA_____SMC_VD__________________________________e4f57########000000000
    Error getting Smart Parameters: Cannot open device
  • Running the voma command, we could see that there were some filesystem errors on the drive:

    [root@ESXi-Host:~] voma -m vmfs -f check -d /vmfs/devices/disks/t10.ATA_____SMC_VD__________________________________e4f57########000000000:8
    Running VMFS Checker version 2.1 in check mode
    Initializing LVM metadata, Basic Checks will be done
     
    Checking for filesystem activity
             Scsi 2 reservation successful                       FS-5 host activity (512 bytes/HB, 2048 HBs).                                                   -
    Phase 1: Checking VMFS header and resource files
       Detected VMFS file system (labeled:'NTNX-local-ds-##############-#') with UUID:6323314c-#########-####-########c7bc, Version 5:81
       Found stale lock [type 10c00003 offset 644133888 v 18, hb offset 3584000
             gen 45, mode 1, owner 6459561b-#########-####-########c7bc mtime 111106
             num 0 gblnum 0 gblgen 0 gblbrk 0]
    Phase 2: Checking VMFS heartbeat region
    ON-DISK ERROR: Invalid HB address <0>
    Phase 3: Checking all file descriptors.
       Found stale lock [type 10c00001 offset 268568576 v 233, hb offset 3584000
             gen 127, mode 1, owner 6a1674dc-#########-####-########c7bc mtime 214
             num 0 gblnum 0 gblgen 0 gblbrk 0]
       Found stale lock [type 10c00001 offset 268572672 v 1744, hb offset 3584000
             gen 121, mode 1, owner 6a0694d4-#########-####-########c7bc mtime 3342
             num 0 gblnum 0 gblgen 0 gblbrk 0]
       Found stale lock [type 10c00001 offset 268574720 v 122, hb offset 3584000
             gen 127, mode 1, owner 6a1674dc-#########-####-########c7bc mtime 758
             num 0 gblnum 0 gblgen 0 gblbrk 0]
    .../....
       Found stale lock [type 10c00001 offset 268910592 v 1714, hb offset 3584000
             gen 113, mode 1, owner 6a0694d4-#########-####-########c7bc mtime 202413
             num 0 gblnum 0 gblgen 0 gblbrk 0]
       Found stale lock [type 10c00001 offset 268955648 v 1773, hb offset 3584000
             gen 127, mode 1, owner 6a1674dc-#########-####-########c7bc mtime 574
             num 0 gblnum 0 gblgen 0 gblbrk 0]
       Found stale lock [type 10c00001 offset 268957696 v 1748, hb offset 3584000
             gen 119, mode 2, owner 00000000-00000000-0000-000000000000 mtime 762
             num 1 gblnum 0 gblgen 0 gblbrk 0]
    Phase 4: Checking pathname and connectivity.
    Phase 5: Checking resource reference counts.
    ON-DISK ERROR: FB inconsistency found: (760,0) allocated in bitmap, but never used
     
    Total Errors Found:           2

Environment

ESXi 8.0

Cause

The issue appears to be caused by VMFS metadata corruption on the underlying storage partition. Although the ESXi host can still recognize the partition and access parts of the datastore, VOMA has identified inconsistencies in the VMFS metadata. These inconsistencies include the heartbeat region, filesystem block allocation information, and multiple stale metadata locks, indicating that portions of the VMFS filesystem have entered an inconsistent or logically corrupted state, thus making the CVM inaccessible.

Resolution

  1. Run the following command to verify the device that was reported to have corrupted data and that you would like to repair.
    voma -m vmfs -f check -d /vmfs/devices/disks/t10.ATA_____SMC_VD__________________________________e4f57########000000000:8
  2. Run the following command to remove the stale entries that were identified during the VOMA check:
    voma -m vmfs -f advfix -p /vmfs/volumes/<datastore for voma output file>/ -d /vmfs/devices/disks/t10.ATA_____SMC_VD___________________________e4f57########000000000:8
  3. When prompted, select 0 for Yes to proceed.
  4. The command will rescan the device and attempt to repair the inconsistencies it detects.
  5. After the stale entries have been corrected, reboot the ESXi host so that the host and Nutanix's storage services can start cleanly and the repaired filesystem state is fully recognized.