Return
An efficient two-stage computing method for large-scale research interest mining
DOI:10.1016/j.future.2025.108117.png)
Abstract
En 中文
Semantic analysis for academic data is crucial for many scientific services, such as review recommendation, planning research funding directions. Research interest analysis faces challenges in large-scale academic data mining. Traditional methods of representing research interests, such as manual labeling, using statistical or machine learning methods, have limitations. In particular, the computation amount is unacceptable in large-scale multisource information integration. This paper presents an efficient computing method for predicting scholar interests based on the principle of large-scale recommendation systems, consisting of rough and refined sorting. In rough sorting, one-hot encoding, CHI square feature selection, TF-IDF feature extraction, and an SGD-based classifier are used to obtain several top interest labels. In refined sorting, a pre-trained SciBERT model outputs the optimal interest labels. The proposed approach offers two main advantages. Firstly, it improves computational efficiency, as directly using pre-trained models like BERT for large-scale data leads to excessive calculations. Secondly, the algorithm ensures better model performance. Feature selection in the rough sorting stage can avoid the negative impact of irrelevant papers on prediction precision, which is a problem when using pre-trained model directly.
Journal
F
IF:
0
Papers:
642
Citations:
0
Organization
Cited Papers
No cited papers available

