Return
Improving TextRank Algorithm for Automatic Keyword Extraction with Tolerance Rough Set
DOI:10.1007/s40815-021-01190-y.png)
Abstract
En 中文
Aiming at the shortcomings of the TextRank method (TM) which only considers the co-occurrence between words and the incipient word importance when extracting keywords, this paper proposes a tolerance rough set (TRS)-based unsupervised keyword extraction method. Generally, how to score the words in a document has a significant influence on the word graph modeling. In this paper, we improve TM in two aspects with TRS theory that is used to mine vocabulary, semantics, grammar and other information in the corpus. First, the degree of words belonging to each document is calculated to form a fuzzy membership matrix, which helps to characterize the incipient word importance. Second, the fuzzy membership of words to each word tolerance class is calculated to form a semantic correlation matrix, which contributes to optimize the transition probability of all graph edges. We apply the proposed methods to the clustering tasks of two datasets, outperforming the strong baselines.
Keywords:
Automatic keyword extraction
Tolerance rough set
Semantic correlation
TextRank
Journal
IF:
3.6
Papers:
2.2K
Citations:
4.3K

