esxtop CheatSheet Troubleshooting Guide

vSphere ESXTOP Cheatsheet & Troubleshooting Guide (2026) | vNetes Skip to main content
VMWARE CLOUD FOUNDATION 9.1 & VSAN ESA / OSA SPEC
VMware ESXi 6.x • 7.x • 8.x+ Enterprise Performance Suite

vSphere ESXTOP Cheatsheet & Troubleshooter

Authoritative real-time performance inspection guide for VMware vSphere. Live interactive terminal emulator, automated minFree memory threshold calculator, latency alert matrices (%RDY, %CSTP, DAVG, KAVG, N%L), and symptom-to-metric diagnostic playbooks.

Interactive ESXTOP Terminal Simulator & Live Telemetry Console

Authentic live-streaming ESXi 8.0/9.1 shell environment with field selectors, interactive hotkeys, and real-time metric telemetry.

root@esxi-prod-01.vnetes.local:~# esxtop LIVE SSH • ESX 9.1.1.0
14:02:18 up 124 days, 18:42 824 worlds 18 VMs 64 vCPUs CPU Load Avg: 0.18, 0.24, 0.31
Polling: 2.0s
VCF 9.1 Breakthrough Real-Time Telemetry Plane

VMware Cloud Foundation 9.1: VCF Operations Real-Time Metrics

In VMware Cloud Foundation 9.1, VMware introduces VCF Operations Real-Time Metrics, bringing granular 2-second streaming telemetry directly into the central VCF Operations console across vCenter, ESXi compute worlds, NSX network fabrics, and vSAN ESA/OSA storage. This architectural leap eliminates the need to open ad-hoc SSH sessions for isolated ESXTOP commands during Sev-1 outages, providing fleet-wide cross-cluster visibility with native root-cause correlation.

2-Second Streaming
Replaces legacy 5-minute rollups and 20-second vCenter averages. Captures micro-burst latency, instantaneous CPU co-scheduling skew (%CSTP), and flash cache write saturation as they happen in real time.
Management Services Cluster
Retires legacy monolithic virtual appliance VMs. Real-time metric ingestion agents run as containerized microservices on a managed Kubernetes cluster, scaling horizontally across thousands of nodes.
Integrated Operations for Logs
Operations for Logs is now fully unified inside VCF Operations. Correlate a live metric spike (e.g. DAVG > 25ms or AVGLAT > 25ms) directly to synchronized VMkernel and vSAN traces with zero context switching.
4 Unified Workflows
Organized around natural IT operations: Operate (live streaming & alerts), Manage (fleet capacity & host memory states), Protect (security & logs), and Build (APIs & software depot).

Architecture Comparison: Classic ESXTOP CLI vs. VCF 9.1 Operations

How VCF 9.1 evolves low-level node diagnostics into an enterprise-wide observability plane.

Enterprise Multi-Cluster Spec
Capability Dimension Classic ESXTOP CLI (Node-Local) VCF 9.1 Operations Real-Time Metrics
Telemetry Frequency Default 2-second screen refresh (customizable via -d delay switch). Native 2-second streaming telemetry ingested continuously into time-series store.
Access & Security Requires SSH root shell on individual ESXi hosts; port 22 open; local credential risk. Web UI single-pane-of-glass with enterprise SSO, role-based access control (RBAC), and full audit logging.
Monitoring Scope Single host in isolation. Impossible to correlate cross-host DRS migrations or vSAN object peers. Fleet-wide across all Workload Domains (vCenter, ESXi Compute, NSX Fabric, vSAN ESA/OSA).
Historical Retention Ephemeral. Once the terminal session exits, data is lost unless manually run in batch mode (-b). Persistent historical telemetry with live scrubbing, metric playback, and trend regression analysis.
Root Cause Correlation Manual human analysis. Admin must switch between 9 views (c, m, d, x, etc.). AI-assisted automated anomaly detection correlating VM CPU contention with vSAN disk latency.
Log Synchronization Separate manual grep on /var/log/vmkernel.log and /var/log/vobd.log. Natively embedded Operations for Logs with 1-click contextual log drilldown from metric peaks.

ESXTOP-to-VCF 9.1 Operations Field Translation Matrix

Select an infrastructure domain below to map classic ESXTOP terminal counters to their exact VCF 9.1 Operations widget paths and alert thresholds:

%RDY Operate Workflow
CPU Ready Time Contention
Virtual Machine > CPU > Ready (%)
VCF 9.1 Alert Threshold: > 5% Warning, > 10% Critical.
Triage Action: Identifies host pCPU overcommitment or wide-vSMP skew in the Operate dashboard.
Terminal Sync Root Cause Engine
%CSTP Operate Workflow
vSMP Co-Scheduling Skew
Virtual Machine > CPU > Co-Stop (%)
VCF 9.1 Alert Threshold: > 3% Warning.
Triage Action: VCF Operations highlights oversized vCPU configurations and suggests right-sizing to minimize co-stop latency.
Terminal Sync Root Cause Engine
MCTLSZ / SWCUR Manage Workflow
Ballooning & Hypervisor Swap
Host > Memory > Reclaim > Balloon / Swap (MB)
VCF 9.1 Alert Threshold: MCTLSZ > 0 MB (Warning), SWCUR > 0 MB (Critical).
Triage Action: Correlates with host memory state transitions (Soft/Hard) to guide DRS memory rebalancing.
Terminal Sync Root Cause Engine
minFree / State Manage Workflow
Host Memory State (High/Soft/Hard/Low)
Host > Capacity > Free Memory vs minFree Threshold
VCF 9.1 Alert Threshold: Free < 64% minFree (Soft), < 32% (Hard), < 16% (Low).
Triage Action: Directly modeled in VCF Operations Capacity Fleet views with automated proactive evacuation.
Terminal Sync Root Cause Engine
DAVG / KAVG Operate Workflow
Storage Device & Kernel Latency
Storage > Datastore / Adapter > Total Latency (ms)
VCF 9.1 Alert Threshold: DAVG > 20ms (SAN/Array), KAVG > 2ms (Queue Congestion).
Triage Action: Highlights HBA queue exhaustion vs. SAN target array bottleneck in live 2s streaming graph.
Terminal Sync Root Cause Engine
AVGLAT Operate Workflow
vSAN DOM Average Latency
vSAN Cluster > Performance > DOM Average Latency (ms)
VCF 9.1 Alert Threshold: > 10ms Warning, > 25ms Critical.
Triage Action: Isolates client-side DOM latency from underlying storage cache or network resynchronization traffic.
Terminal Sync Root Cause Engine
SDLA Operate Workflow
vSAN Disk Latency Std Deviation
vSAN Disk Group > Disk Latency > Standard Deviation (ms)
VCF 9.1 Alert Threshold: > 5ms Warning, > 20ms Critical.
Triage Action: Pinpoints write buffer saturation on specific NVMe/SSD disks in the vSAN ESA or OSA cluster.
Terminal Sync Root Cause Engine
%DRPTX / %DRPRX Operate Workflow
Network Packet Drops
NSX Edge & Host > Network > Dropped Packets (%)
VCF 9.1 Alert Threshold: > 0% Dropped Packets.
Triage Action: Flags physical switch buffer micro-bursts, MTU mismatches, or NIC ring buffer exhaustion.
Terminal Sync Root Cause Engine

Verify the Current ESXi Host Memory State & minFree Calculator

The 4 dynamic ESXi host memory states (High, Soft, Hard, Low), real-time SSH ESXTOP verification steps, dynamic minFree formula sizer, and Broadcom KB 404415 batch ESXTOP memory diagnostic automation.

Verify the Current ESXi Host Memory State

ESXi dynamically transitions across 4 host memory states depending on free physical memory relative to the host minFree reservation. Select each state below to inspect the simulated ESXTOP memory screen header and active VMkernel reclamation mechanics.

esxtop > Press 'm' 4 Host States
root@esxi-host:~# esxtop (press 'm' for Memory screen)
---------------------------------------------------------------------------------------------------------
PMEM /MB: 102400 total: 2150 vmk, 87850 other, 12400 free
min free: 1619 MB, current state: HIGH
MEM overcommit avg: 0.28, 0.32, 0.35
PSHARE /MB: 12480 shared, 2410 common, 10070 saving
MEMCTL /MB: 0 curr, 0 target, 24576 max
ZIP /MB: 0 zipped, 0 saved, CACHEUSD: 0 MB
SWAP /MB: 0 curr, 0.0 r/s, 0.0 w/s
---------------------------------------------------------------------------------------------------------
Active VMkernel Memory Technique
Transparent Page Sharing (TPS)
State Trigger Threshold
Host Free RAM ≥ 100% of minFree (≥ 1,619 MB for a 100 GB host; Pre-vSphere 5: > 6%).
Hypervisor & Reclamation Behavior
By default, Transparent Page Sharing (TPS) runs periodically in the background. The VMkernel scans guest physical memory, hashes identical 4KB pages, and maps them to a single machine page using Copy-on-Write (COW). No ballooning, no compression, and no hypervisor swapping occur.
Guest VM & Workload Impact
Zero performance impact. Guest operating systems and latency-sensitive workloads execute at bare-metal memory speeds with no page reclamation overhead.
Critical ESXTOP Columns to Verify
PSHARE/MB MEM overcommit avg MCTLSZ = 0 SWCUR = 0 CACHEUSD = 0
Recommended Remediation & Action
Normal operating condition. Host has abundant free machine memory. No administrator intervention needed.
Step-by-Step: How to Verify Host Memory State via SSH
1 Connect via SSH
Open a secure terminal session to the target ESXi host console using root or administrative credentials:
ssh root@<esxi-host-fqdn-or-ip>
2 Switch to Memory (m)
Launch the interactive top monitor and press m to switch from default CPU to the Memory display panel:
esxtop # then press 'm'
3 Read PMEM & State
Inspect line 1 (PMEM /MB: ... free) and compare free RAM against the min free reservation to determine the active host state:
min free: 1619 MB, state: HIGH
4 Check Overcommit Ratio
Inspect MEM overcommit avg for 1-min, 5-min, and 15-min moving averages. A value of 0.50 means 50% overcommitted; 5.87 means 587% overcommitted:
MEM overcommit avg: 1.42, 1.38, 1.25
5 Calculate TPS Savings
Transparent Page Sharing savings formula: Host Memory Savings = shared − sharedcommon. Example: 2575 MB − 355 MB = 2220 MB saved:
PSHARE /MB: 2575 shared, 355 common
6 Isolate Contending VMs
Press B to sort by Balloon size (MCTLSZ) or M to sort by MEMSZ. Check if any VM has non-zero SWCUR (swapped) or CACHEUSD (compressed):
Press 'B' (Sort by MCTLSZ)
Host RAM & Free Memory Simulation
Calculated minFree
1,619 MB
≈ 1.58 GB
Current Memory State
HIGH
Normal TPS page sharing; zero contention
100 GB
Presets:
12.4 GB
VMware ESXi minFree Formula (vSphere 5.x, 6.x, 7.x, 8.x):
minFree = 899 MB + [ 1% × (Host RAM − 28 GB) ]
For the first 28 GB of host RAM, minFree is fixed at 899 MB (or ~900 MB). Every additional GB of RAM contributes 10.24 MB (1%). In pre-vSphere 5, fixed percentages (6%, 4%, 2%, 1%) were used, which caused massive wasted reservations on 1TB+ hosts.
The 4 Different ESXi Host Memory States Breakdown
Active State & Dynamic Reclamation Thresholds
1. HIGH
≥ 100% minFree
Abundant free RAM; default background TPS page sharing.
≥ 1,619 MB
2. SOFT
< 64% minFree (2/3 of High)
Host reclaims memory using Balloon driver (MCTLSZ) + TPS.
< 1,036 MB
3. HARD
< 32% minFree (1/3 of High)
Host invokes Swapping (.vswp) + Compression; ballooning stops!
< 518 MB
4. LOW
< 16% minFree (1/6 of High)
ESXi blocks VMs from allocating additional guest RAM!
< 259 MB
Host State Free Memory Condition Reclamation Technique Ballooning (MCTLSZ) Swapping & Compression Guest VM Impact Resolution Playbook
1. HIGH ≥ 100% minFree
(Pre-vSphere 5: > 6%)
Transparent Page Sharing (TPS) Inactive (MCTLSZ = 0) Inactive (SWCUR = 0, CACHEUSD = 0) None (Bare-metal memory access speeds) Healthy. Normal host operating baseline.
2. SOFT < 64% minFree
(2/3 High; Pre-v5: < 4%)
Memory Ballooning (vmmemctl) + TPS Actively Inflates (Target > 0) Inactive (Compression/Swap avoided) Low-Moderate. Guest OS pages internal memory to guest swap. Check cluster DRS, right-size overcommitted VMs, add RAM reservations.
3. HARD < 32% minFree
(1/3 High; Pre-v5: < 2%)
Swapping (.vswp) + Compression PAUSED / STOPPED (Prevents OS panic) Active (ZIP/s > 0, SWR/s > 0) Severe. High disk swap latency (10-50ms), high %SWPWT. URGENT: Force vMotion of large VMs to uncongested hosts immediately.
4. LOW < 16% minFree
(1/6 High; Pre-v5: < 1%)
Allocation Blocking & Execution Throttling Halted (Cannot reclaim further) Continuous aggressive swapping & compression Critical. VMs blocked from allocating RAM; thread hangs, BSODs, OOM. EMERGENCY: Power off non-prod VMs, evacuate workloads, open Broadcom P1 case.
Broadcom KB 404415: ESXi Host Memory Analysis Using Batch ESXTOP
Production Support Playbook
If your ESXi host experiences sustained high memory usage (>80-90% in vCenter), sudden unexplained consumption jumps (e.g. baseline jumping from 350 GB to 690 GB), or VMs report guest memory starvation, execute this official Broadcom 15-minute batch ESXTOP capture script. Ensure datastore has ≥ 500 MB free space (df -h).
# Broadcom KB 404415: 15-Minute Unattended ESXTOP Batch Memory Capture (2-sec intervals)
capture_minutes=15
interval_seconds=2
datastore_path="/vmfs/volumes/<datastore_name>"   # Replace with actual VMFS datastore name

# Calculate total samples needed (450 iterations for 15 mins @ 2s)
samples_per_minute=$(expr 60 / ${interval_seconds})
total_samples=$(expr ${capture_minutes} \* ${samples_per_minute})

# Execute batch capture across all performance counters (-ba)
esxtop -ba -d ${interval_seconds} -n ${total_samples} > "${datastore_path}"/$(hostname)_$(date -u +"%Y-%m-%dT%H%M%S")_esxtop_batch_all.csv

# After completion, generate ESXi diagnostic log bundle for Broadcom Support:
vm-support
Offline Analysis: Load the generated CSV into VMware VisualEsxtop or Windows Perfmon. Filter for Memory Tracker, MCTLSZ, SWCUR, ZIP/s, and MEM Overcommit avg to identify root cause.

Symptom-to-Metric Root Cause Engine

Select the active infrastructure symptom you are troubleshooting to reveal the critical ESXTOP metrics to inspect, root cause mechanics, and step-by-step remediation.

ESXTOP Metric Deep-Dive Library

Full technical reference covering every critical ESXTOP metric, key sequence, healthy vs critical thresholds, and exact resolution steps.

Command-Line & Interactive Hotkey Cheatsheet

Interactive navigation keys, view switching shortcuts, and copyable bash/CLI one-liners for unattended batch-mode ESXTOP performance logging.

Interactive Screen Controls & Safety Reference

Real-time hotkeys for live ESXTOP display manipulation, vCPU group hierarchy expansion, row highlighting, and emergency diagnostic intervention. Commands with stability or operational risks are strictly highlighted below.

Safe Display & Nav (8 Keys) V, e, l, #, 2, 8, e, 6 • Zero risk in production
Operational Caution (1 Key) Key [4] • Hides active entity from screen
Critical Danger (1 Key) Key [k] • Tech Support Only (SIGKILL)
V
only show virtual machine worlds
Target: Global Display Filter • Hides idle worlds, internal ESXi helper worlds, and background daemons
SAFE DISPLAY
e
Expand/Rollup CPU statistics, show details of all worlds associated with group (GID)
Target: CPU View [c] • Prompts for VM Group ID (GID) to display individual vCPUs and worker threads
SAFE TREE EXPAND
k
kill world, for tech support purposes only!
Target: Process / World ID (WID) • Destructive VMkernel SIGKILL Execution
CRITICAL RISK — TECH SUPPORT ONLY
CRITICAL RISK / UNGRACEFUL TERMINATION HAZARD: Forcefully terminates the targeted ESXi World ID using an unrecoverable SIGKILL signal. Terminating a virtual machine world causes an immediate ungraceful power-off without guest OS shutdown, risking file system corruption, uncommitted database transactions, and acute data loss. Terminating system/kernel worlds can destabilize the ESXi host or trigger a Purple Screen of Death (PSOD). For VMware Technical Support (GSS) emergency diagnostics only! Never execute in production without explicit vendor direction.
l
limit display to a single group (GID), enables you to focus on one VM
Target: CPU / Memory Views • Isolates an individual VM Group ID and suppresses all other background noise
SAFE ISOLATION
#
limiting the number of entitites, for instance the top 5
Target: Global Row Limit • Prompts for max visible row count (e.g. top 5 or top 10) to prevent screen overflow
SAFE DISPLAY
2
highlight a row, moving down
Target: Interactive Table Cursor • Moves the highlighted entity selection bar down one position
SAFE NAVIGATION
8
highlight a row, moving up
Target: Interactive Table Cursor • Moves the highlighted entity selection bar up one position
SAFE NAVIGATION
4
remove selected row from view
Target: Active Screen Display • Diagnostic Suppression (Process Stays Active)
CAUTION — VIEW MUTATION
OPERATIONAL MONITORING CAUTION: Strips the highlighted entity from the current active screen view. The underlying VM or process continues executing in the background! If pressed inadvertently, administrators may mistakenly diagnose that a rogue CPU hog or memory-thrashing VM has stopped. To restore all hidden rows, press Space or switch view tabs.
e
statistics broken down per world
Target: CPU & Scheduling Telemetry • Displays granular performance counters for each individual world within a group
SAFE TELEMETRY
6
statistics broken down per world
Target: Per-World Telemetry • Alternative toggle key to break down performance statistics across individual worlds
SAFE TELEMETRY

View Switching & Field Keys

Key Target View & Action Key Fields
c Switch to CPU View D, F
m Switch to Memory View B, D, J, K, Q
n Switch to Network View A, B, C, D, E, F, K, L
d Switch to Disk Adapter View A, B, G, J
u Switch to Disk Device (LUN) View Default Device Stats
v Switch to Disk VM View Per-VM Storage Latency
x Switch to vSAN DOM View DOM Roles & Resync Latency
p Switch to Power Management View C-State & P-State Freq
i Switch to Interrupts View Vector & World Interrupts
f / o Add/Remove Fields (f) • Field Order (o) Toggle columns on/off
W Save Current Configuration Saves to ~/.esxtop4rc
Space Immediate Refresh & Filter Reset Forces repaint & unhides rows
s + 2 Set Screen Refresh Delay (e.g. 2s) Default is 5 seconds

Production Batch Mode Snippets

Standard 30-Minute Batch Capture (5-sec interval)
esxtop -b -d 5 -n 360 > /tmp/esxtop_capture_$(date +%Y%m%d_%H%M%S).csv
Gzip-Compressed Live Streaming (Prevents Disk Fill)
esxtop -b -d 2 -n 900 | gzip -9c > /vmfs/volumes/shared/esxtop_highres.csv.gz
Replay Mode (Play Back Captured Batch Data)
esxtop -R /tmp/esxtop_capture.csv
Batch Collection with Custom Configuration File
esxtop -b -c ~/.esxtop4rc -d 10 -n 180 > /tmp/esxtop_custom.csv

Architecture & Troubleshooting FAQs

In-depth technical explanations of VMkernel scheduling, memory overcommitment, storage queue latency, and NUMA node mechanics.

Community Discussion & Architecture Q&A

Have questions regarding ESXTOP performance metrics, latency thresholds, or host troubleshooting? Join the engineering discussion below.
vNetes Community
Join the VMware ESXTOP & vSphere Engineering Forum
Share real-world troubleshooting scenarios, discuss %RDY or DAVG/cmd thresholds, or request new cheatsheet columns.
Write a Comment
Snippet copied to clipboard!

Discussion 0

Author Active
Leave a Response
Verified Discussion Policy: Comments are moderated by the engineering team to ensure high-quality technical answers and zero spam.
Join the Discussion
Markdown supported: Use `command` for inline code
Theme Mode