返回
An Adaptive Clustering Algorithm Based on Local-Density Peaks for Imbalanced Data Without Parameters
DOI:10.1109/TKDE.2021.3138962.png)
摘要
En 中文
Imbalanced data clustering is a challenging problem in machine learning. The main difficulty is caused by the imbalance in both cluster size and data density distribution. To address this problem, we propose a novel clustering algorithm called LDPI based on local-density peaks in this study. First, an initial sub-cluster construction scheme is designed based on a 3-dimensional (3-D) decision graph that can easily detect the initial sub-cluster centers and identify the noise points. Second, a sub-cluster updating strategy is designed, which can automatically identify the false sub-cluster centers and update the initial sub-clusters. Third, a sub-cluster merging scheme is designed, which merges the updated initial sub-clusters into final clusters. Consequently, the proposed algorithm has three advantages: 1) It does not require any input parameters; 2) It can automatically determine the cluster centers and number of clusters; 3) It is suitable for imbalanced datasets and datasets with arbitrary shapes and distributions. The effectiveness of LDPI is demonstrated experimentally and the superiority of LDPI is identified by comparison with 5 state-of-the-art algorithms.
Keyword:
Clustering algorithms
Machine learning algorithms
Machine learning
Computer science
Clustering methods
Task analysis
Shape
Data clustering
density peaks
imbalanced data
multiple centers
期刊
IF:
10.4
论文数:
6.8K
被引数:
3.2W
机构
引用论文
RNN-DBSCAN: A Density-Based Clustering Algorithm Using Reverse Nearest Neighbor Density EstimatesRnn-dbscan: 使用反向最近邻密度估计的基于密度的聚类算法
Study on density peaks clustering based on k-nearest neighbors and principal component analysis基于k近邻和主成分分析的密度峰聚类研究
Integration of self-organizing feature map and K-means algorithm for market segmentation融合自组织特征映射和k-means算法的市场分割

