Return
A Multivariate Permutation Cluster Validity Test
DOI:10.1002/asmb.70096.png)
Abstract
En 中文
This paper contributes to the existing literature on cluster analysis presenting a new method to identify the best partition, that is, the optimal number of clusters, when the -means clustering algorithm is adopted. Clustering is an unsupervised learning technique aimed at constructing partitions such that items within the same cluster are similar, while those in different clusters exhibit distinct differences. As an exploratory analysis, the effectiveness of clustering algorithms, particularly non-hierarchical ones, depends heavily on key decisions, such as determining the number of clusters in the final partition. In the literature, various cluster validity indices are available to assist in identifying the optimal partition. However, these indices often produce conflicting recommendations, which may not provide clear guidance to researchers and practitioners. To address this gap, the Multivariate Permutation Cluster Validity (MPCV) test has been introduced to compare partitions across multiple clustering performance metrics. This approach synthesizes information from widely used clustering indices, including the Silhouette, Dunn, Calinski-Harabasz, and Davies-Bouldin indices, aiming to achieve an optimal balance between internal homogeneity and external separability. Our approach is presented and discussed using both simulated and real data, highlighting its main advantages.
Keywords:
cluster analysis
multivariate
permutation
ranking
smart data
Journal
A
IF:
1.5
Papers:
66
Citations:
1.3K

