返回
HY-DBSCAN: A hybrid parallel DBSCAN clustering algorithm scalable on distributed-memory computers
DOI:10.1016/j.jpdc.2022.06.005.png)
摘要
En 中文
Dbscan is a density-based clustering algorithm which is well known for its ability to discover clusters of arbitrary shape as well as to distinguish noise. As it is computationally expensive for large datasets, research studies on the parallelization of Dbscan have been received a considerable amount of attention. In this paper we present an exact, efficient and scalable parallel Dbscan algorithm which we call HyDbscan. It employs three major techniques to enable scalable data clustering on distributed-memory computers i) a modified kd-tree for domain decomposition, ii) a spatial indexing approach based on grid and inference, and iii) a cluster merging scheme based on distributed Rem's Union-Find algorithm. Moreover, Hy-Dbscan exploits process level and thread level parallelization. In experiments, we have demonstrated performance and scalability using two scientific datasets on up to 2048 cores of a distributed-memory computer. Through extensive evaluation, we show that Hy-Dbscan significantly outperforms previous state-of-the-art Dbscan implementations.
Keyword:
Dbscan
Density-based clustering
Parallel algorithm
Distributed-memory
期刊
IF:
4
论文数:
3.8K
被引数:
4.8K
机构
暂无机构信息
引用论文
Herpes Simplex Virus 1 (HSV-1) ICP22 Protein Directly Interacts with Cyclin-Dependent Kinase (CDK)9 to Inhibit RNA Polymerase II Transcription Elongation
PLoS ONE
IF0
R&D strategy study of customized furniture with film-laminated wood-based panels based on an analytic hierarchy process/quality function deployment integration approach
BioResources
IF0
没有更多内容

