arrow
Return

A complex history browsing text categorization method with improved BERT embedding layer

delete2025-02-03
delete1
PRE
AI
Y
Yuanhang Wang
周永华 (Yonghua Zhou) *
H
Huiyu Qi
W
Wang, Dingyi
H
Huang, Annan
DOI:10.1007/s10489-025-06298-4delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
For long texts composed of multiple short fragments, the importance of each fragment to the classification task varies. Some fragments have higher discriminative power and positively contribute to the classification, while others lack discriminative power or even mislead it. Existing methods struggle to converge when handling texts with negative examples. This study analyzes user behavior and assigns interest scores to text fragments based on their classification relevance, allowing the model to focus more on important fragments. Building on bidirectional encoder representations from transformers (BERT), we propose an interest encoding layer model for historical browsing texts. By analyzing user behavior and incorporating an improved term frequency-inverse document frequency (TF-IDF) method, the model adds indicators to fragments with higher discriminative power for user behavior analysis, enabling the model to focus more on these during training. Finally, comparative experiments on the BERT model series validate the advantages of the proposed approach.
Keywords:
History browsing text
BERT network
TF-IDF method
Interest embedding
Text categorization

Journal

Applied Intelligence cover
Applied Intelligence
IF:
3.5
Papers:
7.5K
Citations:
1.7W

Organization

B
Beijing Jiaotong University
Scholars:
2.2W
Papers: 1.7W
Citations: 1.2W