arrow
Return

SEBGM: Sentence Embedding Based on Generation Model with multi-task learning

delete2024-08-01
delete1
PRE
AI
王茜 cover
王茜 (Qian Wang)
W
Weiqi Zhang
曹裕 (Yu Cao)
D
Dezhong Peng
X
Xu Wang *
DOI:10.1016/j.csl.2024.101647delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Sentence embedding, which aims to learn an effective representation of a sentence, is a significant part for downstream tasks. Recently, using contrastive learning and pre -trained model, most methods of sentence embedding achieve encouraging results. However, on the one hand, these methods utilize discrete data augmentation to obtain positive samples performing contrastive learning, which could distort the original semantic of sentences. On the other hand, most methods directly employ the contrastive frameworks of computer vision to perform contrastive learning, which could confine the contrastive training due to the discrete and sparse text data compared with image data. To solve the issues above, we design a novel contrastive framework based on generation model with multi -task learning by supervised contrastive training on the dataset of natural language inference (NLI) to obtain meaningful sentence embedding (SEBGM). SEBGM makes use of multi -task learning to enhance the usage of wordlevel and sentence -level semantic information of samples. In this way, the positive samples of SEBGM are from NLI rather than data augmentation. Extensive experiments show that our proposed SEBGM can advance the state-of-the-art sentence embedding on the semantic textual similarity (STS) tasks by utilizing multi -task learning.
Keywords:
Sentence embedding
Contrastive learning
Multi-task learning

Journal

C
Computer Speech and Language
IF:
3.4
Papers:
1.5K
Citations:
2.6K

Organization

S
Southwest Petroleum University
Scholars:
1.4W
Papers: 7.8K
Citations: 8.5K
S
sichuan university
Scholars:
11.9W
Papers: 7.7W
Citations: 100