返回
A parallel text clustering method using Spark and hashing
DOI:10.1007/s00607-021-00932-y.png)
摘要
En 中文
Clustering textual data has become an important task in data analytics since several applications require to automatically organizing large amounts of textual documents into homogeneous topics. The increasing growth of available textual data from web, social networks and open platforms have challenged this task. It becomes important to design scalable clustering method able to effectively organize huge amount of textual data into topics. In this context, we propose a new parallel text clustering method based on Spark framework and hashing. The proposed method deals simultaneously with the issue of clustering huge amount of documents and the issue of high dimensionality of textual data by respectively integrating the divide and conquer approach and implementing a new document hashing strategy. These two facts have shown an important improvement of scalability and a good approximation of clustering quality results. Experiments performed on several large collections of documents have shown the effectiveness of the proposed method compared to existing ones in terms of running time and clustering accuracy.
Keyword:
Text clustering
Parallel computing
Spark framework
Hashing
High-dimensional data
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
C
IF:
2.8
论文数:
2.3K
被引数:
3.5K
机构
引用论文
Optogenetic inhibition of Purkinje cell activity reveals cerebellar control of blood pressure during postural alterations in anesthetized rats
Neuroscience
IF0

