返回
Efficient agglomerative hierarchical clustering
DOI:10.1016/j.eswa.2014.09.054.png)
摘要
En 中文
Hierarchical clustering is of great importance in data analytics especially because of the exponential growth of real-world data. Often these data are unlabelled and there is little prior domain knowledge available. One challenge in handling these huge data collections is the computational cost. In this paper, we aim to improve the efficiency by introducing a set of methods of agglomerative hierarchical clustering. Instead of building cluster hierarchies based on raw data points, our approach builds a hierarchy based on a group of centroids. These centroids represent a group of adjacent points in the data space. By this approach, feature extraction or dimensionality reduction is not required. To evaluate our approach, we have conducted a comprehensive experimental study. We tested the approach with different clustering methods (i.e., UPGMA and SLINK), data distributions, (i.e., normal and uniform), and distance measures (i.e., Euclidean and Canberra). The experimental results indicate that, using the centroid based approach, computational cost can be significantly reduced without compromising the clustering performance. The performance of this approach is relatively consistent regardless the variation of the settings, i.e., clustering methods, data distributions, and distance measures. (C) 2014 Elsevier Ltd. All rights reserved.
Keyword:
Clustering analysis
Hybrid clustering
Data mining
Data distribution
Coefficient of correlation
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
7.5
论文数:
2.9W
被引数:
10.2W
机构
引用论文
Impact of Servant Leadership on Project Success Through Mediating Role of Team Motivation and Effectiveness: A Case of Software Industry
Sage Open
IF0
Incremental cluster-based retrieval using compressed cluster-skipping inverted files使用压缩的跳过簇的反向文件进行基于簇的增量检索

