返回
Big data: an optimized approach for cluster initialization
DOI:10.1186/s40537-023-00798-1.png)
摘要
En 中文
The k-means, one of the most widely used clustering algorithm, is not only faster in computation but also produces comparatively better clusters. However, it has two major downsides, first it is sensitive to initialize k value and secondly, especially for larger datasets, the number of iterations could be very large, making it computationally hard. In order to address these issues, we proposed a scalable and cost-effective algorithm, called R-k-means, which provides an optimized solution for better clustering large scale high-dimensional datasets. The algorithm first selects O(R) initial points then reselect O(l) better initial points, using distance probability from dataset. These points are then again clustered into k initial points. An empirical study in a controlled environment was conducted using both simulated and real datasets. Experimental results showed that the proposed approach outperformed as compared to the previous approaches when the size of data increases with increasing number of dimensions.
Keyword:
k-means
Cluster initialization
Large scale data
期刊
IF:
6.4
论文数:
1.5K
被引数:
1.1W
机构
引用论文
A Hybrid MPI/OpenMP Parallelization of K-Means Algorithms Accelerated Using the Triangle Inequality使用三角不等式加速的k-means算法的混合MPI/OpenMP并行化
IEEE ACCESS
IF3.6
Revisiting K-Means and Topic Modeling, a Comparison Study to Cluster Arabic Documents
IEEE ACCESS
IF3.6
k-PbC: an improved cluster center initialization for categorical data clusteringK-pbc: 一种改进的分类数据聚类中心初始化方法
APPLIED INTELLIGENCE
IF3.5

