arrow
Return

Generalized Score Comparison-Based Learning Objective for Deep Speaker Embedding

delete2025-01-01
delete0
delete
OA
AI
M
Min Hyun Han
S
Sung Hwan Mun
N
Nam Soo Kim *
DOI:10.1109/ACCESS.2025.3552790delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In state-of-the-art speaker verification systems, speaker embeddings are trained to be closer to the target speaker prototype, which is either obtained from the other speech samples or constructed with trainable parameters. This can be considered a classification task since the network is trying to learn the features that are most relevant to the corresponding speaker from the input speech. Although classification-based learning demonstrates the ability to extract speaker-related information, it does not guarantee optimal speaker verification performance. In this paper, we propose a score comparison-based learning objective, which guides the training framework to be more consistent with the verification task, enforcing the embedding space to have lower intra-class variance compared to inter-class variance in terms of similarity scores. Furthermore, we propose a generalized loss function for score comparison-based learning, encompassing many conventional training losses and regularization techniques. The proposed technique is compared with the conventional methods using the VoxCeleb, VOiCES, CN-Celeb, and Common Voice datasets. Experimental results demonstrate that the proposed method can boost the performance and make the system more robust to over-fitting in speaker verification tasks.
Keywords:
Training
Measurement
Prototypes
Vectors
Feature extraction
Data mining
Object recognition
Kernel
Focusing
Data models
Speaker verification
deep speaker embedding
metric learning
embedding space

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

S
samsung
Scholars:
8.6K
Papers: 6.4K
Citations: 8
S
seoul national university (snu)
Scholars:
7.2W
Papers: 6.6W
Citations: 86