Return
Scalable incremental fuzzy consensus clustering algorithm for handling big data
DOI:10.1007/s00500-021-05733-1.png)
Abstract
En 中文
Consensus clustering can produce novel, stable, and robust clustering results. Consensus clustering intends to merge a few existing basic segments into a coordinated one, and this has been broadly perceived as a promising solution for heterogeneous data clustering for big data. Even though many clustering algorithms have been proposed, getting a decent quality segment with high effectiveness is still not yet decided. In this paper, we propose a scalable incremental fuzzy consensus clustering (SIFCC) algorithm for a big data framework. It has been implemented on Apache Spark cluster framework, a distributed data stream environment for handling big data by considering the data as a set of data subsets that are processed incrementally. Sparks work great for iterative algorithms by supporting in-memory calculations, scalability, etc. SIFCC not only facilitates efficient big data clustering, but also improves the quality of clusters, performs storage space optimization, and time complexity during clustering. To establish the comparison, we designed and implemented the scalable model of existing fuzzy consensus clustering (FCC) on Apache Spark cluster, named as a scalable fuzzy consensus clustering (SFCC). Extensive experiments on real-world datasets show that the SIFCC algorithm achieves the better potential for clustering of Big Data in comparison with SFCC.
Keywords:
Fuzzy consensus clustering
Space optimization
Incremental algorithm
Big data
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
2.5
Papers:
1.0W
Citations:
2.1W

