arrow
返回

Important citations identification with semi-supervised classification model

delete2022-01-20
delete8
PRE
AI
X
Xin An *
孙
孙新 (Xin Sun)
徐
徐硕 (Shuo Xu)
DOI:10.1007/s11192-021-04212-6delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Given that citations are not equally important, various techniques have been presented to identify important citations on the basis of supervised machine learning models. However, only a small volume of instances have been annotated manually with the labels. To make full use of unlabeled instances and promote the identification performance, the semi-supervised self-training technique is utilized here to identify important citations in this work. After six groups of features are engineered, the SVM and RF models are chosen as the base classifiers for self-training strategy. Then two experiments based on two different types of datasets are conducted. The experiment on the expert-labeled dataset from one single discipline shows that the semi-supervised versions of SVM and RF models significantly improve the performance of the conventional supervised versions when unannotated samples under 75% and 95% confidence level are rejoined to the training set, respectively. The AUC-PR and AUC-ROC of SVM model are 0.8102 and 0.9622, and those of RF model reach 0.9248 and 0.9841, which outperform their counterparts and the benchmark methods in the literature. This demonstrates the effectiveness of our semi-supervised self-training strategy for important citation identification. Another experiment on the author-labeled dataset from multiple disciplines, semi-supervised learning models can perform better than their supervised learning counterparts in term of AUC-PR when the ratio of labeled instances is less than 20%. Compared to our first experiment, insufficient amount of instances from each discipline in our second experiment enables the performance of the models to be unsatisfactory.
Keyword:
Important citation
Semi-supervised learning
Self-training
Expert-labeled dataset
Author-labeled dataset

期刊

Scientometrics 封面图
Scientometrics
IF:
3.5
论文数:
8.1K
被引数:
2.2W

机构

B
Beijing University of Technology
学者数:
2.8W
论文数: 2.1W
被引数: 2.7W
B
beijing forestry university
学者数:
1.9W
论文数: 1.1W
被引数: 3
引用论文

引用论文

Emerging research topics detection with multiple machine learning models
err2019-11-01
err54
PREAI
errXu, Shuo; Hao, Liyuan; An, Xin; Yang, Guancan; Wang, Feifei
err分享
err收藏
Morphological Traits Influence the Uptake Ability of Priority Pollutant Elements by Hypnum cupressiforme and Robinia pseudoacacia Leaves
err2020-01-29
err0
errOAAI
errFiore Capozzi; Anna Di Palma; Maria Cristina Sorrentino; Paola Adamo; Simonetta Giordano; Valeria Spagnuolo
err分享
err收藏
err分享
err收藏
err分享
err收藏
Important citation identification using sentiment analysis of in-text citations
err2021-01-01
err45
PREAI
errAljuaid, Hanan; Iftikhar, Rimsha; Ahmad, Shahbaz; Asif, Muhammad; Afzal, Muhammad Tanvir
err分享
err收藏
Ultraviolet and optical observations of OB associations in M31
err1995-06-01
err0
PREAI
errJ. K. Hill; J. E. Isensee; R. C. Bohlin; K. P. Cheng; P. M. N. Hintzen; R. W. O'Connell; M. S. Roberts; A. M. Smith; Eric P. Smith; T. P. Stecher
err分享
err收藏
学者 查看更多内容