In the vSphere Replication UI or hms.log, you may see errors similar to:
A replication error has occurred on the vSphere Replication server for replication 'VM_NAME'. Details: 'Error for (datastoreUUID: "UUID"), (diskId: "RDID-UUID"), (pathname: "VM_NAME/disk.vmdk") ObjLib error: Invalid argument; Failed to combine files ... Prune disks could not remove disk instance (instanceKey=####); While completing partial prune'.
vSphere Replication 9.x
This issue is caused by intermittent NFS storage disconnects (All Paths Down / APD events) on the underlying storage layer.
When the storage path fails, the Host-Based Replication (HBR) stream over port 32032 is interrupted. These disconnects prevent the ESXi hosts from performing critical I/O operations on the VMDK files, such as disk consolidation and pruning. This results in stale file locks and a hang in the replication health check (mapping test) as the services cannot reach a consistent state on the impacted datastores.
To restore replication functionality, the underlying storage must be stabilized before refreshing the replication services.
Disk prune errors may also have other causes, see: vSphere Replication for virtual machines are in error state with the consolidation error: "Prune disks could not remove disk instance"