arrow
Return

Boosting Short Text Classification by Solving the OOV Problem

delete2023-01-01
delete1
PRE
AI
N
Nan Gao *
Y
Yongjian Wang
P
Peng Chen
J
Jijun Tang
DOI:10.1109/TASLP.2023.3316422delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In the field of natural language processing, text classification has received a lot of attention. Compared with long texts, short texts have fewer words and lack contextual semantic information. Existing approaches enrich short text information by linking the external knowledge graph, but they ignore the out-of-vocabulary (OOV) problem during entity linking, especially when dealing with domain-oriented data, which has some rare words or domain-specific nouns. In this article, to alleviate the OOV problem caused by linking the external knowledge graph(KG), we propose a domain knowledge graph and entity complementation strategy to improve the performance of short text classification. Specifically, the external knowledge graph is used to enrich the information of short texts. The self-build domain knowledge graph is used to solve the problem of entities failing to link to the external knowledge graph. Finally, we conduct experiments on various datasets: 1. a labeled Chinese electronic domain dataset; 2. an open-source dataset to test the performance of our algorithm in different data distribution scenarios. The results demonstrate our dual knowledge graph model outperforms the state-of-the-art short text classification methods, especially when the OOV problem is severe.
Keywords:
Dual knowledge graph
knowledge enhancement
out of vocabulary problem
short text classification

Journal

I
IEEE-ACM Transactions on Audio Speech and Language Processing
IF:
5.1
Papers:
2.6K
Citations:
1.1W

Organization

Z
zhejiang university of technology
Scholars:
3.3W
Papers: 2.0W
Citations: 22
U
University of South Carolina System
Scholars:
1.5W
Papers: 1.4W
Citations: 27