VM fails to power on with fail to check pid error due to orphaned snapshot
search cancel

VM fails to power on with fail to check pid error due to orphaned snapshot

book

Article ID: 447062

calendar_today

Updated On:

Products

VMware vCenter Server VMware vSphere ESXi VMware Cloud Director VMware Telco Cloud Infrastructure

Issue/Introduction

A virtual machine running in a VMware Cloud Director environment is unable to power on, despite appearing healthy in the vCD management interface. Attempting to power on or access the VM via the vCenter Server web console results in a "fail to check pid" error. Datastores show adequate physical free space, but a massive (10 TB) virtual machine snapshot is present directly in vCenter Server that is completely invisible or orphaned from the vCD interface.

Environment

ESXi: 7.x

vCenter server: 7.x

VMware cloud director: 10.3

Telco cloud Infrastructure: 2.2

Cause

A non-memory snapshot was left running beyond the recommended timeframe, growing to an exceptionally large size of 9-10 TB. This massive delta disk created a severe storage overhead and file lock conflict. This state prevented the VMX process from successfully initializing the PID check, blocking the power-on sequence.

Resolution

Steps:

  1. Log in to the vCenter Server interface and locate the affected virtual machine.

  2. Identify the hidden/orphaned snapshot that is not visible in the vCD interface.

  3. Initiate a deletion of the orphaned snapshot directly from vCenter Server to trigger consolidation.

  4. Monitor the snapshot deletion/consolidation process to completion.

    Note: The consolidation process may require a significant amount of time depending on the size of the snapshot delta disk (e.g., 10 TB).

  5. Once consolidation is successfully completed, attempt to power on the virtual machine.

Critical Note: The excessive size and age of the orphaned snapshot introduces a high risk of guest OS corruption prior to or during the consolidation process. If the virtual machine fails to boot into the operating system, or if OS-level services remain unresponsive after the snapshot is successfully cleared, the guest OS is considered corrupted. In such cases, the virtual machine must be redeployed from a known good backup or provisioned as a fresh deployment.

Additional Information

FAQ: Delete all Snapshots and Consolidate Snapshots Feature