Return
A parallel text clustering method using Spark and hashing
DOI:10.1007/s00607-021-00932-y.png)
Abstract
En 中文
Clustering textual data has become an important task in data analytics since several applications require to automatically organizing large amounts of textual documents into homogeneous topics. The increasing growth of available textual data from web, social networks and open platforms have challenged this task. It becomes important to design scalable clustering method able to effectively organize huge amount of textual data into topics. In this context, we propose a new parallel text clustering method based on Spark framework and hashing. The proposed method deals simultaneously with the issue of clustering huge amount of documents and the issue of high dimensionality of textual data by respectively integrating the divide and conquer approach and implementing a new document hashing strategy. These two facts have shown an important improvement of scalability and a good approximation of clustering quality results. Experiments performed on several large collections of documents have shown the effectiveness of the proposed method compared to existing ones in terms of running time and clustering accuracy.
Keywords:
Text clustering
Parallel computing
Spark framework
Hashing
High-dimensional data
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
C
IF:
2.8
Papers:
2.3K
Citations:
3.5K
Organization
Cited Papers
Optogenetic inhibition of Purkinje cell activity reveals cerebellar control of blood pressure during postural alterations in anesthetized rats
Neuroscience
IF0

