返回
Accelerating the k-Means plus plus Algorithm by Using Geometric Information
DOI:10.1109/ACCESS.2025.3561293.png)
摘要
En 中文
Clustering is a fundamental task in data analysis with applications across a wide range of fields, such as computer vision, pattern recognition, and data mining. Real-world use cases include social network analysis, medical imaging, market segmentation, and anomaly detection, to name a few. In this paper, we propose an acceleration of the exact k-means++ algorithm using geometric information, specifically the Triangle Inequality and additional norm filters, along with a two-step sampling procedure. Our experiments demonstrate that the accelerated version outperforms the standard k-means++ version in terms of the number of visited points and distance calculations, achieving greater speedup as the number of clusters increases. The version utilizing the Triangle Inequality is particularly effective for low-dimensional data, while the additional norm-based filter enhances performance in high-dimensional instances with greater norm variance among points. Additional experiments show the behavior of our algorithms when executed concurrently across multiple jobs and examine how memory performance impacts practical speedup.
Keyword:
Clustering algorithms
Approximation algorithms
Standards
Measurement
Machine learning algorithms
Vectors
Filters
Computational efficiency
Scalability
Prediction algorithms
Clustering
D-2 sampling
k-means plus plus
norm
triangle inequality
期刊
IF:
3.6
论文数:
9.8W
被引数:
29.4W
机构
引用论文
暂无论文信息

