After a cluster-wide storage APD (All Paths Down) or PDL (Permanent Device Loss) event recovers — or after a cluster-wide reboot of hosts — vSphere HA (FDM) terminates and attempts to fail over the affected VMs, but a subset of VMs are left in a Powered Off state permanently. HA does not retry or auto-recover these VMs. -Seen on both VMFS (FC/iSCSI) and NFS-backed clusters -Also seen during cluster-wide host reboot testing (e.g. array reboot/infra-failure scenarios), even without a storage APD/PDL trigger. -The VMs themselves are healthy; their configuration/data is intact — they simply never get powered back on by HA.
9.1.x
A race condition in VC HA's (FDM) VM placement engine causes the power-on action for a few VMs to be silently dropped during cluster-wide failover.
Broadcom Engineering are aware of this issue and currently working on a resolution.
Workaround
Manually power on the affected VMs(via VC or API)