arrow
Return

Textual data summarization using the Self-Organized Co-Clustering model

delete2020-07-01
delete13
delete
OA
AI
M
Margot Selosse *
J
Julien Jacques
C
Christophe Biernacki
DOI:10.1016/j.patcog.2020.107315delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Recently, different studies have demonstrated the use of co-clustering, a data mining technique which simultaneously produces row-clusters of observations and column-clusters of features. The present work introduces a novel co-clustering model to easily summarize textual data in a document-term format. In addition to highlighting homogeneous co-clusters as other existing algorithms do we also distinguish noisy co-clusters from significant co-clusters, which is particularly useful for sparse document-term matrices. Furthermore, our model proposes a structure among the significant co-clusters, thus providing improved interpretability to users. The approach proposed contends with state-of-the-art methods for document and term clustering and offers user-friendly results. The model relies on the Poisson distribution and on a constrained version of the Latent Block Model, which is a probabilistic approach for co-clustering. A Stochastic Expectation-Maximization algorithm is proposed to run the model's inference as well as a model selection criterion to choose the number of co-clusters. Both simulated and real data sets illustrate the efficiency of this model by its ability to easily identify relevant co-clusters. (C) 2020 Elsevier Ltd. All rights reserved.
Keywords:
Co-Clustering
Document-term matrix
Latent block model
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Pattern Recognition cover
Pattern Recognition
IF:
7.6
Papers:
1.3W
Citations:
4.5W

Organization

I
Inria
Scholars:
3.5K
Papers: 2.5K
Citations: 343
U
universite de lille
Scholars:
2.7W
Papers: 2.0W
Citations: 15