arrow
返回

Text classification with improved word embedding and adaptive segmentation

delete2024-03-01
delete5
PRE
AI
G
Guoying Sun
Y
Yanan Cheng
Z
Zhaoxin Zhang *
X
Xiaojun Tong
T
Tingting Chai
DOI:10.1016/j.eswa.2023.121852delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Text classification first needs to convert the text into embedding vectors. Considering that static word embedding models such as Word2vec do not consider the position information of word and the difference of its role in different documents, while dynamic word embedding models such as Bert consume a large amount of time. An improved word embedding model based on pre-trained Word2vec is proposed, which achieves better classification accuracy and much lower classification time than Bert. At first, the concept of Term Document Frequency (TDF) is proposed on the basis of TF-IDF, and the TF-IDF-TDF of each word in different documents is calculated. Then, The positional encoding is added. Finally, in order to reduce the misleading of words with low importance, a filter is designed to set the embedding vector with low importance to zero. Considering that the sequence length that the deep learning model can handle is limited, and the text sequence exceeding the Maximum Sequence Length (MSL) set by the deep learning model will be directly truncated and discarded, an adaptive segmentation model is proposed, which can set different segmentation strategies for different texts according to the length of the text and the MSL. In order to maintain the continuity of adjacent text after segmentation, an adjacent-segment-vector-attended co-attention network is designed. In addition, the multi-channel convolution and the capsule network are designed to further extract deep hidden features. Multiple comparative experiment results show that the proposed model achieves the best Accuracy and Micro-F1 on five long text baseline datasets and six short text baseline datasets. In addition, when the MSL is not set too large compared with the document length in the dataset, the classification results of the proposed model are not affected by it.
Keyword:
Word embedding
Term Document Frequency
Adaptive segmentation model
Maximum Sequence Length
Co-attention network

期刊

Expert Systems with Applications 封面图
Expert Systems with Applications
IF:
7.5
论文数:
3.0W
被引数:
10.2W

机构

H
harbin institute of technology
学者数:
8.0W
论文数: 6.6W
被引数: 66
引用论文

引用论文

Co-attention network with label embedding for text classification
err2022-01-01
err42
PREAI
errLiu, Minqian; Liu, Lizhao; Cao, Junyi; Du, Qing
err分享
err收藏
err分享
err收藏
err分享
err收藏
YAC transgene-mediated olfactory receptor gene choiceYAC转基因介导的嗅觉受体基因选择
err2000-02-01
err0
errOAAI
errFarah A.W. Ebrahimi; James Edmondson; Rodney Rothstein; Andrew Chess
err分享
err收藏
Learning URL Embedding for Malicious Website Detection用于恶意网站检测的URL嵌入学习
err2020-10-01
err45
PREAI
errYan, Xiaodan; Xu, Yang; Cui, Baojiang; Zhang, Shuhan; Guo, Taibiao; Li, Chaoliang
err分享
err收藏
学者 查看更多内容