返回
Improving TextRank Algorithm for Automatic Keyword Extraction with Tolerance Rough Set
DOI:10.1007/s40815-021-01190-y.png)
摘要
En 中文
Aiming at the shortcomings of the TextRank method (TM) which only considers the co-occurrence between words and the incipient word importance when extracting keywords, this paper proposes a tolerance rough set (TRS)-based unsupervised keyword extraction method. Generally, how to score the words in a document has a significant influence on the word graph modeling. In this paper, we improve TM in two aspects with TRS theory that is used to mine vocabulary, semantics, grammar and other information in the corpus. First, the degree of words belonging to each document is calculated to form a fuzzy membership matrix, which helps to characterize the incipient word importance. Second, the fuzzy membership of words to each word tolerance class is calculated to form a semantic correlation matrix, which contributes to optimize the transition probability of all graph edges. We apply the proposed methods to the clustering tasks of two datasets, outperforming the strong baselines.
Keyword:
Automatic keyword extraction
Tolerance rough set
Semantic correlation
TextRank
期刊
IF:
3.6
论文数:
2.2K
被引数:
4.3K
机构
引用论文
Dosimetric Impact of Amino Acid Positron Emission Tomography Imaging for Target Delineation in Radiation Treatment Planning for High-Grade Gliomas氨基酸正电子发射断层成像对高分级胶质瘤放射治疗计划中靶区勾画剂量的影响
Automatic Keywords Extraction Based on Co-Occurrence and Semantic Relationships Between Words
IEEE ACCESS
IF3.6

