arrow
返回

A stratified sampling based clustering algorithm for large-scale data

delete2019-01-01
delete42
PRE
AI
X
Xingwang Zhao
J
Jiye Liang *
C
Chuangyin Dang
DOI:10.1016/j.knosys.2018.09.007delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Large-scale data analysis is a challenging and relevant task for present-day research and industry. As a promising data analysis tool, clustering is becoming more important in the era of big data. In large-scale data clustering, sampling is an efficient and most widely used approximation technique. Recently, several sampling-based clustering algorithms have attracted considerable attention in large-scale data analysis owing to their efficiency. However, some of these existing algorithms have low clustering accuracy, whereas others have high computational complexity. To overcome these deficiencies, a stratified sampling based clustering algorithm for large-scale data is proposed in this paper. Its basic steps include: (1) obtaining a number of representative samples from different strata with a stratified sampling scheme, which are formed by locality sensitive hashing technique, (2) partitioning the chosen samples into different clusters using the fuzzy c-means clustering algorithm, (3) assigning the out-of-sample objects into their closest clusters via data labeling technique. The performance of the proposed algorithm is compared with the state-of-the-art sampling-based fuzzy c-means clustering algorithms on several large-scale data sets including synthetic and real ones. The experimental results show that the proposed algorithm outperforms the related algorithms in terms of clustering quality and computational efficiency for large-scale data sets. (C) 2018 Published by Elsevier B.V.
Keyword:
Large-scale data
Fuzzy c-means algorithm
Stratified sampling
Data labeling
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

K
Knowledge-Based Systems
IF:
7.6
论文数:
1.2W
被引数:
4.5W

机构

C
City University of Hong Kong
学者数:
2.3W
论文数: 3.0W
被引数: 6.1W
S
Shanxi University
学者数:
1.3W
论文数: 8.4K
被引数: 1.2W
引用论文

引用论文

A Hybrid Approach to Clustering in Big Data
err2016-10-01
err70
PREAI
errKumar, Dheeraj; Bezdek, James C.; Palaniswami, Marimuthu; Rajasegarar, Sutharshan; Leckie, Christopher; Havens, Timothy Craig
err分享
err收藏
err分享
err收藏
err分享
err收藏
err2002-01-01
err0
errOAAI
errPhilip Boalch
err分享
err收藏
Learning to Hash for Indexing Big Data-A Survey
err2016-01-01
err399
errOAAI
errWang, Jun; Liu, Wei; Kumar, Sanjiv; Chang, Shih-Fu
err分享
err收藏
err分享
err收藏
学者 查看更多内容