返回
Topic modeling combined with classification technique for extractive multi-document text summarization
DOI:10.1007/s00500-020-05207-w.png)
摘要
En 中文
The qualities of human readable summaries available in the datasets are not up to the mark, leading to issues in creating an accurate model for text summarization. Although recent works have been largely built upon this issue and set up a strong platform for further improvements, they still have many limitations. Looking in this direction, the paper proposes a novel methodology for summarizing a corpus of documents to generate a coherent summary using topic modeling and classification technique. The objectives of the propose work are highlighted below: A novel heuristic approach is introduced to find out the actual number of topics that exist in a corpus of documents which handles the stochastic nature of latent dirichlet allocation. A large corpus of documents is handled by minimizing the huge set of sentences into a small set without losing the important one and thus providing a concise and information rich summary at the end. Ensuring that the sentences are arranged as per their importance in the coherent summary. Results of the experiment are compared with the state-of-the-art summary systems. The outcomes of the empirical work show that the proposed model is more promising compared to the well-known text summarization models.
Keyword:
Classification
Extractive
LDA
ROUGE
Silhouette
Summarization
Topic modeling
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
2.5
论文数:
1.0W
被引数:
2.1W
机构
暂无机构信息

