Return
An improved TextRank algorithm based on complex network features
DOI:10.1142/S0129183125420112.png)
Abstract
En 中文
Keyword extraction has a wide range of applications in the field of natural language processing. Many methods are currently used for keyword extraction, but the accuracy of keyword extraction still needs to be improved. This paper proposes an improved Chinese keyword extraction algorithm C-TextRank, which further improves the accuracy of keyword extraction based on complex networks. Based on TextRank, C-TextRank calculates final weight by fusing the position, the degree and the clustering coefficient of a node. The final weight values are obtained by combining with position coefficient and are applied to the sorting of nodes where Top-K nodes are finally extracted as keywords. This study uses the Southern Weekend, a Chinese news website, as corpus and four keyword extraction algorithms are selected for comparison. Experiments show that C-TextRank performs the best when the weights of node degree and clustering coefficient are set at 0.8 and 0.2, respectively. In terms of indicators including precision, recall and F-score of keyword extraction, our algorithm outperforms other four comparison algorithms. In comparison with TextRank, the average values of precision, recall and F-score are improved by 6.3%, 8.6% and 6.5%, respectively.
Keywords:
Complex network
keyword extraction
TextRank
word co-occurrence network
Journal
I
IF:
1.6
Papers:
143
Citations:
1

