arrow
返回

Document clustering

delete2022-06-08
delete3
delete
OA
AI
I
Irene Cozzolino
M
Maria Brigida Ferraro *
DOI:10.1002/wics.1588delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Nowadays, the explosive growth in text data emphasizes the need for developing new and computationally efficient methods and credible theoretical support tailored for analyzing such large-scale data. Given the vast amount of this kind of unstructured data, the majority of it is not classified, hence unsupervised learning techniques show to be useful in this field. Document clustering has proven to be an efficient tool in organizing textual documents and it has been widely applied in different areas from information retrieval to topic modeling. Before introducing the proposals of document clustering algorithms, the principal steps of the whole process, including the mathematical representation of documents and the preprocessing phase, are discussed. Then, the main clustering algorithms used for text data are critically analyzed, considering prototype-based, graph-based, hierarchical, and model-based approaches. This article is categorized under: Statistical Learning and Exploratory Methods of the Data Sciences > Clustering and Classification Statistical Learning and Exploratory Methods of the Data Sciences > Text Mining Data: Types and Structure > Text Data
Keyword:
document clustering
document representation
graph-based methods
hierarchical methods
model-based methods
prototype-based methods
text data

期刊

W
Wiley Interdisciplinary Reviews and Computational Statistics
IF:
5.4
论文数:
201
被引数:
5.1K

机构

S
sapienza university rome
学者数:
6.3W
论文数: 4.7W
被引数: 381
引用论文

引用论文

The crystal structure of GeAsSe
err1981-10-01
err0
PREAI
errF. Hulliger; T. Siegrist
err分享
err收藏
err分享
err收藏
err分享
err收藏
err分享
err收藏
Carbon isotope discrimination in leaf juice of Acacia mangium and its relationship to water-use efficiency
err2009-03-25
err0
PREAI
errLvliu Zou; Guchou Sun; Ping Zhao; Xi’an Cai; Xiaoping Zeng; Xiaojing Liu
err分享
err收藏
学者 查看更多内容