返回
A new data clustering algorithm based on critical distance methodology
DOI:10.1016/j.eswa.2019.03.051.png)
摘要
En 中文
A variety of algorithms have recently emerged in the field of cluster analysis. Consequently, based on the distribution nature of the data, an appropriate algorithm can be chosen for the purpose of clustering. It is difficult for a user to decide a priori which algorithm would be the most appropriate for a given dataset. Algorithms based on graphs provide good results for this task. However, these algorithms are vulnerable to outliers with limited information about edges contained in the tree to split a dataset. Thus, in several fields, the need for better clustering algorithms increases and for this reason utilizing robust and dynamic algorithms to improve and simplify the whole process of data clustering has become an urgent need. In this paper, we propose a novel distance-based clustering algorithm called the critical distance clustering algorithm. This algorithm depends on the Euclidean distance between data points and some basic mathematical statistics operations. The algorithm is simple, robust, and flexible; it works with quantitative data that are real-valued, not qualitative, and categorical with different dimensions. In this work, 26 experiments are conducted using different types of real and synthetic datasets taken from different fields. The results prove that the new algorithm outperforms some popular clustering algorithms such as MST-based clustering, K-means, and Dbscan. Moreover, the algorithm can precisely produce more reasonable clusters even when the dataset contains outliers and without specifying any parameters in advance. It also provides a number of indicators to evaluate the established clusters and prove the validity of the clustering. (C) 2019 Published by Elsevier Ltd.
Keyword:
Algorithm
Cluster analysis
Euclidean distance
MST
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
7.5
论文数:
3.0W
被引数:
10.2W
机构
引用论文
An efficient hybrid approach based on PSO, ACO and k-means for cluster analysis一种基于PSO、ACO和k-means的高效混合聚类分析方法
Fast density clustering strategies based on the k-means algorithm基于k-means算法的快速密度聚类策略
PATTERN RECOGNITION
IF7.6
Habitat type determines the effects of disturbance on the breeding productivity of the Dartford Warbler Sylvia undata
Ibis
IF0

