arrow
Return

Self-supervised Contrastive Learning for Content-Centric Speech Representation

delete2026-01-01
delete0
PRE
AI
J
Jinlong Li
L
Ling Dong
W
Wenjun Wang
余正涛 cover
余正涛 (Zhengtao Yu) *
高盛祥 cover
高盛祥 (Shengxiang Gao)
DOI:10.1007/978-981-95-2725-0_1delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Self-supervised learning (SSL) speech models have achieved remarkable performance across various tasks, with the learned representations often exhibiting a high degree of generality and applicability to multiple downstream tasks. However, these representations contain both speech content and some paralinguistic information, which may be redundant for content-focused tasks. Decoupling this redundant information is challenging. To address this issue, we propose a Self-Supervised Contrastive Representation Learning method (SSCRL), which effectively disentangles paralinguistic information from speech content by aligning similar content speech representations in the feature space using self-supervised contrastive learning with pitch perturbation and speaker perturbation features. Experimental results demonstrate that the proposed method, when fine-tuned on the LibriSpeech 100-hour dataset, achieves superior performance across all content-related tasks in the SUPERB Benchmark, generally outperforming prior approaches.
Keywords:
Self-Supervised Fine-Tuning
Feature Disentanglement
Pre-trained Speech Model
Contrastive Learning

Journal

C
CHINESE COMPUTATIONAL LINGUISTICS, CCL 2025
IF:
0
Papers:
27
Citations:
0

Organization

K
kunming university of science & technology
Scholars:
2.4K
Papers: 610
Citations: 0