返回
A complex history browsing text categorization method with improved BERT embedding layer
DOI:10.1007/s10489-025-06298-4.png)
摘要
En 中文
For long texts composed of multiple short fragments, the importance of each fragment to the classification task varies. Some fragments have higher discriminative power and positively contribute to the classification, while others lack discriminative power or even mislead it. Existing methods struggle to converge when handling texts with negative examples. This study analyzes user behavior and assigns interest scores to text fragments based on their classification relevance, allowing the model to focus more on important fragments. Building on bidirectional encoder representations from transformers (BERT), we propose an interest encoding layer model for historical browsing texts. By analyzing user behavior and incorporating an improved term frequency-inverse document frequency (TF-IDF) method, the model adds indicators to fragments with higher discriminative power for user behavior analysis, enabling the model to focus more on these during training. Finally, comparative experiments on the BERT model series validate the advantages of the proposed approach.
Keyword:
History browsing text
BERT network
TF-IDF method
Interest embedding
Text categorization
期刊
IF:
3.5
论文数:
7.6K
被引数:
1.7W
机构
引用论文
RETRACTED: Automatic text classification using machine learning and optimization algorithms (Retracted Article)
SOFT COMPUTING
IF2.5
An in-text citation classification predictive model for a scholarly search system学术搜索系统的文本引文分类预测模型
SCIENTOMETRICS
IF3.5

