arrow
返回

Extractive text summarization using clustering-based topic modeling

delete2022-10-04
delete7
PRE
AI
R
Ramesh Chandra Belwal *
S
Sawan Rai
A
Atul Gupta
DOI:10.1007/s00500-022-07534-6delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Text summarization is the process of converting the input document into a short form, provided that it preserves the overall meaning associated with it. Primarily, text summarization is achieved in two ways, i.e., abstractive and extractive. Extractive summarizers select a few best sentences out of the input document, while abstractive methods may modify the sentence structure or introduce new sentences. The proposed approach is an extractive text summarization technique, where we have expanded topic modeling specifically to be applied to multiple lower-level specialized entities (i.e., groups) embedded in a single document. Our goal is to overcome the lack of coherence issues found in the summarization techniques. Topic modeling was initially proposed to model text data at the multi-document and word levels without considering sentence modeling. Subsequently, it has been applied at the sentence level and used for the document summarization; however, certain limitations were associated. Topic modeling does not perform as expected when applied to a single document at the sentence level. To address this shortcoming, we have proposed a summarization approach that is incorporated at the individual document and clusters level (instead of the sentence level). We aim to choose the best statement from each group (containing sentences of the same kind) found in the given text. We have tried to select the perfect topic by evaluating the probability distribution of the words and respective topics' at the cluster level. The method is evaluated on two standard datasets and shows significant performance gains over existing text summarization techniques. Compared to other text summarization techniques, the Rouge parameters for automatic evaluation show a considerable improvement in F-measure, precision, and recall of the generated summary. Furthermore, a manual evaluation has demonstrated that the proposed approach outperforms the current state-of-the-art text summarization approaches.
Keyword:
Extractive summarization
Topic modeling
Clustering
Semantic measure

期刊

Soft Computing 封面图
Soft Computing
IF:
2.5
论文数:
1.0W
被引数:
2.1W

机构

引用论文

引用论文

err
IF0
err
err0
PREAI
err
err分享
err收藏
Ionospheric observations of underground nuclear explosions (UNE) using GPS and the Very Large Array
err2013-08-20
err0
errOAAI
errJihye Park; Joseph Helmboldt; Dorota A. Grejner‐Brzezinska; Ralph R. B. von Frese; Thomas L. Wilson
err分享
err收藏
SCALING WITH KNOWN UNCERTAINTY: A SYNTHESIS
err2006-01-01
err0
PREAI
errJIANGUO WU; HARBIN LI; K. BRUCE JONES; ORIE L. LOUCKS
err分享
err收藏
Cross-reactive antibody immunity against SARS-CoV-2 in children and adults
err2021-05-31
err0
errOAAI
errElizabeth Fraley; Cas LeMaster; Dithi Banerjee; Santosh Khanal; Rangaraj Selvarangan; Todd Bradley
err分享
err收藏
Homeobox Gene Involvement in Normal Hematopoiesis and in the Pathogenesis of Childhood Leukemias
err2017-01-01
err0
PREAI
errMaria Adamaki; Maria Goulielmaki; Ioannis Christodoulou; Spiros Vlahopoulos; Vassilis Zoumpourlis
err分享
err收藏
Endoscopic Myotomy for Achalasia贲门失弛缓症的内窥镜肌切开术
err2014-09-01
err0
PREAI
errChristy M. Dunst; Ashwin A. Kurian; Lee L. Swanstrom
err分享
err收藏
学者 查看更多内容