arrow
Return

A Sampling-Based Density Peaks Clustering Algorithm for Large-Scale Data

delete2023-04-01
delete28
PRE
AI
S
Shifei Ding
C
Chao Li
L
Ling Ding
J
Jian Zhang
L
Lili Guo
DOI:10.1016/j.patcog.2022.109238delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
With the rapid development of information technology, massive amount of data is generated. How to dis-cover useful information to support decision-making has become one of the focuses of scholar's research. Clustering is thought to be one of the main means to deal with large-scale data. Density peaks clustering (DPC) is an effective density-based clustering algorithm which is widely applied in numerous fields be-cause of its satisfactory performance. However, the computational complexity of DPC is O(N2) which is not friendly to large-scale data. To solve this issue, a sampling-based density peaks clustering algorithm for large-scale data (SDPC) is proposed. Firstly, a sampling method is used to reduce the distance cal-culations. Secondly, approximate representatives are identified by an improved TI search strategy which further accelerates the clustering process. Afterwards, the approximate representatives are clustered by DPC. Finally, the remaining points are allocated to the same cluster as its nearest representatives. Exper-imental results on both synthetic datasets and real-world datasets illustrate that SDPC is more efficient than DPC, while its clustering performance maintains the same level as DPC.(c) 2022 Elsevier Ltd. All rights reserved.
Keywords:
Density peaks clustering
Sampling method
TI search strategy
Large-scale data

Journal

Pattern Recognition cover
Pattern Recognition
IF:
7.6
Papers:
1.3W
Citations:
4.5W

Organization

No organization information available