ESXi is running workloads but vCenter shows the host as disconnected or not responding.
01 Start with the ESXi management path, then validate DNS, ports, agents, networking and logs.
11 Verify recovery and the underlying cause after the host reconnects.
Use this runbook when an ESXi host is healthy enough to run workloads but vCenter shows it as disconnected or not responding. Separate workload connectivity from the ESXi management path and troubleshoot the first failing dependency.
Start with the management path
Separate VM workload connectivity from ESXi management connectivity. Check whether the ESXi host can reach its default gateway from DCUI and whether vCenter can reach the ESXi management address.
# From the ESXi DCUI or ESXi Shell, verify the management path
# Confirm the expected management VMkernel adapter and gateway.
# From a management system, test the ESXi management address
Test-NetConnection -Port 443
Validate DNS and IP identity
Confirm forward and reverse DNS resolution from vCenter and the ESXi host. A hostname resolving to an incorrect or multiple IP addresses can prevent vCenter from communicating with the intended management interface.
nslookup
nslookup
ping
Test required ports
Check TCP 443 and the vCenter/ESXi communication path, including port 902 where applicable. A ping result alone does not prove that vSphere management traffic works.
Test-NetConnection -Port 443
Test-NetConnection -Port 902
Check ESXi management agents
Review hostd and vpxa when the host is reachable but vCenter cannot synchronize it. Management services can become unresponsive because of network, storage or resource pressure.
# Review hostd and vpxa status using your approved ESXi management method.
# Capture the original error before restarting services.
# Useful log files
hostd.log
vpxa.log
Check MTU and physical networking
Verify the management VMkernel path end to end. Look for MTU mismatches, LACP issues, VLAN errors, switch-port problems, routing and firewall rules.
# Validate the management VMkernel configuration and path.
# Confirm the VLAN and MTU match the upstream network.
# From a suitable host, test the management endpoint
Test-NetConnection -Port 443
Check storage and host pressure
Datastore latency, APD/PDL conditions, high CPU or memory pressure can affect host responsiveness. Review vmkernel, vobd and vpxa logs before restarting services.
# Review these logs around the first failure
vmkernel.log
vobd.log
vpxa.log
# Correlate timestamps with vCenter tasks and events.
Review logs
Review ESXi and vCenter logs around the time the disconnection began. Search for connection failures, timeouts, SSL, certificates, authentication, handshake, no route and agent unreachable conditions.
# Useful ESXi logs
hostd.log
vpxa.log
vmkernel.log
# Also review
# vCenter tasks and events
# Physical switch logs
# Firewall or ACL logs
Reconnect the host from vCenter
Once network, DNS, time, certificate and management-agent checks are satisfactory, use the Reconnect action in vCenter. After reconnection, validate host alarms, datastores, VM state, networking, HA, DRS, vMotion, backup and monitoring.
# vCenter action:
# Host → Reconnect
# Enter valid ESXi administrator credentials if prompted.
# Confirm the host returns to Connected state.
If reconnect fails, follow the matching path
Use the error message to select the next path. Network errors point to VLAN, switch, routing, gateway, MTU or firewall checks. DNS errors require forward/reverse resolution and FQDN consistency. Service errors require hostd/vpxa and resource checks. SSL errors require time and certificate validation.
# Network
Test-NetConnection -Port 443
# DNS
nslookup
# Then correlate the exact vCenter task error
# with ESXi hostd/vpxa/vmkernel timestamps.
Remove and re-add the host only as a last resort
Before removing the host from vCenter, confirm the reason is understood, HA/DRS impact is known, host configuration is backed up, datastore visibility and VM registration are documented, networking is recorded, and no critical vMotion, backup or replication task is active.
# Do not remove and re-add a production host as the first troubleshooting action.
# Complete the network, DNS, time, certificate, agent and log checks first.
Verify recovery
A successful reconnect does not automatically mean the incident is fully resolved. Confirm the underlying cause and validate related services after the host returns to Connected.
# Post-recovery validation
# Host alarms
# Datastore accessibility
# VM power states
# VM network connectivity
# HA and DRS
# vMotion
# Backup access
# Monitoring status