返回
DENDIS: A new density-based sampling for clustering algorithm
DOI:10.1016/j.eswa.2016.03.008.png)
摘要
En 中文
To deal with large datasets, sampling can be used as a preprocessing step for clustering. In this paper, an hybrid sampling algorithm is proposed. It is density-based while managing distance concepts to ensure space coverage and fit cluster shapes. At each step a new item is added to the sample: it is chosen as the furthest from the representative in the most important group. A constraint on the hyper volume induced by the samples avoids over sampling in high density areas. The inner structure allows for internal optimization: only a few distances have to be computed. The algorithm behavior is investigated using synthetic and real-world data sets and compared to alternative approaches, at conceptual and empirical levels. The numerical experiments proved it is more parsimonious, faster and more accurate, according to the Rand Index, with both k-means and hierarchical clustering algorithms. (C) 2016 Elsevier Ltd. All rights reserved.
Keyword:
Density
Distance
Space coverage
Clustering
Rand index
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
7.5
论文数:
2.9W
被引数:
10.2W
机构
引用论文
Halogen atom transfer radical cyclization of N-allyl-N-benzyl-2,2-dihaloamides to 2-pyrrolidinones, promoted by Fe0-FeCl3 or CuCl-TMEDA
Tetrahedron
IF0
Comparison of distributed evolutionary k-means clustering algorithms分布式进化k-means聚类算法的比较
NEUROCOMPUTING
IF6.5
A comparative study of efficient initialization methods for the k-means clustering algorithmK-means聚类算法高效初始化方法的比较研究

