返回
An Efficient Split-Merge Re-Start for the K-Means Algorithm
DOI:10.1109/TKDE.2020.3002926.png)
摘要
En 中文
The K-means algorithm is one of the most popular clustering methods. However, it is a well-known fact that its performance, in terms of quality of the obtained solution and computational load, highly depends upon its initialization phase. For this reason, different initialization techniques have been developed throughout the years to enable its fast convergence to competitive solutions. In this sense, it is common practice to re-start the K-means algorithm several times via one of these techniques and keep the solution with the lowest error. Unfortunately, such a choice is still likely to be a poor approximation of the optimal set of centroids. In this article, we introduce a cheap Split-Merge step that can be used to re-start the K-means algorithm after reaching a fixed point. Under some settings, one can show that this approach reduces the error of the given fixed point without requiring any further iteration of the K-means algorithm. Moreover, experimental results show that this strategy is able to generate approximations with an associated error that is hard to reach for different multi-start methods, such as multi-start Forgy K-means, K-means++ and Hartigan K-means, while also computing a lower amount of distances than the previous algorithms.
Keyword:
K-means
K-means plus
Hartigan K-means
clustering
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
10.4
论文数:
6.8K
被引数:
3.2W
机构
暂无机构信息
引用论文
Loss of Habitat Connectivity Hinders Pair Formation and Juvenile Dispersal of Chucao Tapaculos in Chilean Rainforest
The Condor
IF0
A comparative study of efficient initialization methods for the k-means clustering algorithmK-means聚类算法高效初始化方法的比较研究

