vSphere ESXTOP Cheatsheet & Troubleshooter
Authoritative real-time performance inspection guide for VMware vSphere. Live interactive terminal emulator,
automated minFree memory threshold calculator, latency alert matrices (%RDY, %CSTP, DAVG, KAVG, N%L), and symptom-to-metric diagnostic playbooks.
Interactive ESXTOP Terminal Simulator & Live Telemetry Console
Authentic live-streaming ESXi 8.0/9.1 shell environment with field selectors, interactive hotkeys, and real-time metric telemetry.
VMware Cloud Foundation 9.1: VCF Operations Real-Time Metrics
In VMware Cloud Foundation 9.1, VMware introduces VCF Operations Real-Time Metrics, bringing granular 2-second streaming telemetry directly into the central VCF Operations console across vCenter, ESXi compute worlds, NSX network fabrics, and vSAN ESA/OSA storage. This architectural leap eliminates the need to open ad-hoc SSH sessions for isolated ESXTOP commands during Sev-1 outages, providing fleet-wide cross-cluster visibility with native root-cause correlation.
DAVG > 25ms or AVGLAT > 25ms) directly to synchronized VMkernel and vSAN
traces with zero context switching.
Architecture Comparison: Classic ESXTOP CLI vs. VCF 9.1 Operations
How VCF 9.1 evolves low-level node diagnostics into an enterprise-wide observability plane.
| Capability Dimension | Classic ESXTOP CLI (Node-Local) | VCF 9.1 Operations Real-Time Metrics |
|---|---|---|
| Telemetry Frequency | Default 2-second screen refresh (customizable via -d delay switch). |
Native 2-second streaming telemetry ingested continuously into time-series store. |
| Access & Security | Requires SSH root shell on individual ESXi hosts; port 22 open; local credential risk. | Web UI single-pane-of-glass with enterprise SSO, role-based access control (RBAC), and full audit logging. |
| Monitoring Scope | Single host in isolation. Impossible to correlate cross-host DRS migrations or vSAN object peers. | Fleet-wide across all Workload Domains (vCenter, ESXi Compute, NSX Fabric, vSAN ESA/OSA). |
| Historical Retention | Ephemeral. Once the terminal session exits, data is lost unless manually run in batch mode
(-b). |
Persistent historical telemetry with live scrubbing, metric playback, and trend regression analysis. |
| Root Cause Correlation | Manual human analysis. Admin must switch between 9 views (c, m,
d, x, etc.).
|
AI-assisted automated anomaly detection correlating VM CPU contention with vSAN disk latency. |
| Log Synchronization | Separate manual grep on /var/log/vmkernel.log and /var/log/vobd.log. |
Natively embedded Operations for Logs with 1-click contextual log drilldown from metric peaks. |
ESXTOP-to-VCF 9.1 Operations Field Translation Matrix
Select an infrastructure domain below to map classic ESXTOP terminal counters to their exact VCF 9.1 Operations widget paths and alert thresholds:
Triage Action: Identifies host pCPU overcommitment or wide-vSMP skew in the Operate dashboard.
Triage Action: VCF Operations highlights oversized vCPU configurations and suggests right-sizing to minimize co-stop latency.
Triage Action: Correlates with host memory state transitions (Soft/Hard) to guide DRS memory rebalancing.
Triage Action: Directly modeled in VCF Operations Capacity Fleet views with automated proactive evacuation.
Triage Action: Highlights HBA queue exhaustion vs. SAN target array bottleneck in live 2s streaming graph.
Triage Action: Isolates client-side DOM latency from underlying storage cache or network resynchronization traffic.
Triage Action: Pinpoints write buffer saturation on specific NVMe/SSD disks in the vSAN ESA or OSA cluster.
Triage Action: Flags physical switch buffer micro-bursts, MTU mismatches, or NIC ring buffer exhaustion.
Verify the Current ESXi Host Memory State & minFree Calculator
The 4 dynamic ESXi host memory states (High, Soft, Hard, Low), real-time SSH ESXTOP verification steps, dynamic minFree formula sizer, and Broadcom KB 404415 batch ESXTOP memory diagnostic automation.
ESXi dynamically transitions across 4 host memory states depending on free physical
memory relative to the host minFree reservation. Select each state below to inspect the
simulated ESXTOP memory screen header and active VMkernel reclamation mechanics.
PMEM /MB: ... free) and compare free RAM against
the min free reservation to determine the active host state:MEM overcommit avg for 1-min, 5-min, and 15-min moving
averages. A value of 0.50 means 50% overcommitted; 5.87 means 587%
overcommitted:Host Memory Savings = shared − sharedcommon. Example: 2575 MB − 355 MB =
2220 MB saved:
MCTLSZ) or M to sort by
MEMSZ. Check if any VM has non-zero SWCUR (swapped) or CACHEUSD
(compressed):
minFree = 899 MB + [ 1% × (Host RAM − 28 GB) ]For the first 28 GB of host RAM, minFree is fixed at 899 MB (or ~900 MB). Every additional GB of RAM contributes 10.24 MB (1%). In pre-vSphere 5, fixed percentages (6%, 4%, 2%, 1%) were used, which caused massive wasted reservations on 1TB+ hosts.
| Host State | Free Memory Condition | Reclamation Technique | Ballooning (MCTLSZ) | Swapping & Compression | Guest VM Impact | Resolution Playbook |
|---|---|---|---|---|---|---|
| 1. HIGH | ≥ 100% minFree (Pre-vSphere 5: > 6%) |
Transparent Page Sharing (TPS) | Inactive (MCTLSZ = 0) |
Inactive (SWCUR = 0, CACHEUSD = 0) |
None (Bare-metal memory access speeds) | Healthy. Normal host operating baseline. |
| 2. SOFT | < 64% minFree (2/3 High; Pre-v5: < 4%) |
Memory Ballooning (vmmemctl) + TPS | Actively Inflates (Target > 0) | Inactive (Compression/Swap avoided) | Low-Moderate. Guest OS pages internal memory to guest swap. | Check cluster DRS, right-size overcommitted VMs, add RAM reservations. |
| 3. HARD | < 32% minFree (1/3 High; Pre-v5: < 2%) |
Swapping (.vswp) + Compression | PAUSED / STOPPED (Prevents OS panic) | Active (ZIP/s > 0,
SWR/s > 0)
|
Severe. High disk swap latency (10-50ms), high %SWPWT. |
URGENT: Force vMotion of large VMs to uncongested hosts immediately. |
| 4. LOW | < 16% minFree (1/6 High; Pre-v5: < 1%) |
Allocation Blocking & Execution Throttling | Halted (Cannot reclaim further) | Continuous aggressive swapping & compression | Critical. VMs blocked from allocating RAM; thread hangs, BSODs, OOM. | EMERGENCY: Power off non-prod VMs, evacuate workloads, open Broadcom P1 case. |
df -h).
# Broadcom KB 404415: 15-Minute Unattended ESXTOP Batch Memory Capture (2-sec intervals)
capture_minutes=15
interval_seconds=2
datastore_path="/vmfs/volumes/<datastore_name>" # Replace with actual VMFS datastore name
# Calculate total samples needed (450 iterations for 15 mins @ 2s)
samples_per_minute=$(expr 60 / ${interval_seconds})
total_samples=$(expr ${capture_minutes} \* ${samples_per_minute})
# Execute batch capture across all performance counters (-ba)
esxtop -ba -d ${interval_seconds} -n ${total_samples} > "${datastore_path}"/$(hostname)_$(date -u +"%Y-%m-%dT%H%M%S")_esxtop_batch_all.csv
# After completion, generate ESXi diagnostic log bundle for Broadcom Support:
vm-support
Memory Tracker, MCTLSZ, SWCUR, ZIP/s, and
MEM Overcommit avg to identify root cause.
Symptom-to-Metric Root Cause Engine
Select the active infrastructure symptom you are troubleshooting to reveal the critical ESXTOP metrics to inspect, root cause mechanics, and step-by-step remediation.
ESXTOP Metric Deep-Dive Library
Full technical reference covering every critical ESXTOP metric, key sequence, healthy vs critical thresholds, and exact resolution steps.
Command-Line & Interactive Hotkey Cheatsheet
Interactive navigation keys, view switching shortcuts, and copyable bash/CLI one-liners for unattended batch-mode ESXTOP performance logging.
Interactive Screen Controls & Safety Reference
Real-time hotkeys for live ESXTOP display manipulation, vCPU group hierarchy expansion, row highlighting, and emergency diagnostic intervention. Commands with stability or operational risks are strictly highlighted below.
SIGKILL signal.
Terminating a virtual machine world causes an immediate ungraceful power-off without guest OS shutdown,
risking file system corruption, uncommitted database transactions, and acute data loss. Terminating
system/kernel worlds can destabilize the ESXi host or trigger a Purple Screen of Death (PSOD).
For VMware Technical Support (GSS) emergency diagnostics only! Never execute in
production without explicit vendor direction.
View Switching & Field Keys
| Key | Target View & Action | Key Fields |
|---|---|---|
| c | Switch to CPU View | D, F |
| m | Switch to Memory View | B, D, J, K, Q |
| n | Switch to Network View | A, B, C, D, E, F, K, L |
| d | Switch to Disk Adapter View | A, B, G, J |
| u | Switch to Disk Device (LUN) View | Default Device Stats |
| v | Switch to Disk VM View | Per-VM Storage Latency |
| x | Switch to vSAN DOM View | DOM Roles & Resync Latency |
| p | Switch to Power Management View | C-State & P-State Freq |
| i | Switch to Interrupts View | Vector & World Interrupts |
| f / o | Add/Remove Fields (f) • Field Order (o) | Toggle columns on/off |
| W | Save Current Configuration | Saves to ~/.esxtop4rc |
| Space | Immediate Refresh & Filter Reset | Forces repaint & unhides rows |
| s + 2 | Set Screen Refresh Delay (e.g. 2s) | Default is 5 seconds |
Production Batch Mode Snippets
Architecture & Troubleshooting FAQs
In-depth technical explanations of VMkernel scheduling, memory overcommitment, storage queue latency, and NUMA node mechanics.
Discussion 0
Author Active`command`for inline code