arrow
Return

Improving TextRank Algorithm for Automatic Keyword Extraction with Tolerance Rough Set

delete2021-11-21
delete12
PRE
AI
D
Dong Qiu *
Z
Zheng Qin
DOI:10.1007/s40815-021-01190-ydelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Aiming at the shortcomings of the TextRank method (TM) which only considers the co-occurrence between words and the incipient word importance when extracting keywords, this paper proposes a tolerance rough set (TRS)-based unsupervised keyword extraction method. Generally, how to score the words in a document has a significant influence on the word graph modeling. In this paper, we improve TM in two aspects with TRS theory that is used to mine vocabulary, semantics, grammar and other information in the corpus. First, the degree of words belonging to each document is calculated to form a fuzzy membership matrix, which helps to characterize the incipient word importance. Second, the fuzzy membership of words to each word tolerance class is calculated to form a semantic correlation matrix, which contributes to optimize the transition probability of all graph edges. We apply the proposed methods to the clustering tasks of two datasets, outperforming the strong baselines.
Keywords:
Automatic keyword extraction
Tolerance rough set
Semantic correlation
TextRank

Journal

International Journal of Fuzzy Systems cover
International Journal of Fuzzy Systems
IF:
3.6
Papers:
2.2K
Citations:
4.3K

Organization

C
chongqing university of posts & telecommunications
Scholars:
6.7K
Papers: 5.3K
Citations: 5