arrow
Return

DP- k-modes: A self-tuning k-modes clustering algorithm

delete2022-06-01
delete7
PRE
AI
J
Juanying Xie *
王明钊 cover
王明钊 (Mingzhao Wang)
X
Xiaoxiao Lu
X
Xinglin Liu
P
P.W. Grant
DOI:10.1016/j.patrec.2022.04.026delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The k-modes clustering algorithm was proposed by Huang for handling datasets with categorical attributes, however, the dissimilarity measure used limits its applicability. Ng et al. improved on Huang's k-modes algorithm by proposing a new dissimilarity measure between objects. Moreover, both k-modes algorithms require the initial seeds to be randomly chosen and the number of clusters be specified manually. To overcome the limitations of Huang's and Ng's k-modes clustering algorithms, we first extend the clustering algorithm published in Science in 2014 (clustering by fast search and find of density peaks). The optimal initial seeds and the number of clusters of a dataset are determined simultaneously by taking the standard deviation as the self-tuning cutoff distance and the simple match dissimilarity as the distance measurement in the definition of the density of a point. A new dissimilarity measure is proposed to calculate the dissimilarities between objects to improve on that of Ng's k-modes algorithm. The performance of our resulting self-tuning k-modes clustering algorithm was tested on nine datasets (three being relatively large) from the UCI (University of California in Irvine) machine learning repository. The clustering results were compared to those produced by Huang's and Ng's algorithms. Statistical tests of three k-modes algorithms were undertaken to determine whether or not there is significant difference between our self-tuning k-modes algorithm and Huang's and Ng's k-modes algorithms. All these experimental results demonstrate that our proposed k-modes clustering algorithm is superior to Hang's and Ng's k-modes algorithms in terms of clustering accuracy ( ACC ) and the well-known Adjusted Rand Index (ARI) metric. Our self-tuning k-modes algorithm is significantly different from both Huang's and Ng's k modes algorithms, and there is no statistically significant difference between Ng's and Huang's k-modes algorithms.(c) 2022 Elsevier B.V. All rights reserved.
Keywords:
k-Modes clustering
Local density
Density peaks
Standard deviation
Initial seeds
k-Modes clustering
Local density
Density peaks
Standard deviation
Initial seeds

Journal

Pattern Recognition Letters cover
Pattern Recognition Letters
IF:
3.3
Papers:
7.8K
Citations:
1.6W

Organization

S
Shaanxi Normal University
Scholars:
1.6W
Papers: 1.1W
Citations: 1.7W
S
Swansea University
Scholars:
8.3K
Papers: 8.6K
Citations: 1.3W