arrow
Return

Improving text classification via computing category correlation matrix from text graph

delete2025-01-01
delete0
PRE
AI
张珍 (Zhen Zhang)
M
Mengqiu Liu
X
Xiyuan Jia *
G
Gongxun Miao
X
Xin Wang
H
Hao Ni
G
Guohua Wu
DOI:10.1016/j.csl.2024.101688delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In text classification task, models have shown remarkable accuracy across various datasets. However, confusion often arises when certain categories within the dataset are too similar, causing misclassification of certain samples. This paper proposes an improved method for this problem, through the creation of a three-layer text graph for the corpus, which is used to calculate the Category Correlation Matrix (CCM). Additionally, this paper introduces categoryadaptive contrastive learning for text embedding from the encoder, enhancing the model's ability to distinguish between samples in confusable categories that are easily confused. Soft labels are generated using this matrix to guide the classifier, preventing the model from becoming overconfident with one-hot vectors. The efficacy of this approach was demonstrated through experimental evaluations on three text encoders and six different datasets.
Keywords:
Natural language processing
Text classification
Contrastive learning
Soft label learning

Journal

C
Computer Speech and Language
IF:
3.4
Papers:
1.5K
Citations:
2.6K

Organization

H
Hangzhou Dianzi University
Scholars:
1.3W
Papers: 9.5K
Citations: 7.5K