返回
MICRank: Multi-information interconstrained keyphrase extraction
DOI:10.1016/j.eswa.2024.123744.png)
摘要
En 中文
Keyphrase Extraction is an automatic task that involves identifying the key words or phrases that capture the main content of an article. It is useful for various downstream tasks, including text search, text clustering, and text classification. Embedding-based methods for keyphrase extraction have shown promising results by utilizing pre-trained language models to represent candidate phrases and documents separately. These methods then rank the candidate phrases based on the cosine similarity between the document and the candidate phrases embeddings. However, there are mainly two shortcomings in such methods: I) Redundancy errors, when there are partial repetitions of candidate keyphrases, the methods tend to use redundant long phrases as keyphrases; II) Low keyphrase coverage, such as some keyphrases used to describe locally important information are ignored. In this paper, we propose an unsupervised keyphrase extraction method called MICRank, which evaluates the importance of candidate keyphrases from three perspectives: global information, local information, and attribute information, and solved the aforementioned issues. The experimental results on six benchmarks demonstrate that the proposed MICRank method outperforms the state-of-the-art unsupervised keyphrase extraction methods. In addition, this paper improves the judgment criterion of correct keyphrase extraction and introduces a new evaluation metric called S1@M (M is an element of {5,10,15}) to address the issue of synonyms being considered incorrect predictions.
Keyword:
Keyphrase extraction
Local information
Global information
Attribute information
A new evaluation metric
期刊
IF:
7.5
论文数:
3.0W
被引数:
10.2W
机构
引用论文
SIFRank: A New Baseline for Unsupervised Keyphrase Extraction Based on Pre-Trained Language Model
IEEE ACCESS
IF3.6
GCN-based document representation for keyphrase generation enhanced by maximizing mutual information

