arrow
Return

Speaker embedding loss for end-to-end speaker diarization without external embedding networks

delete2025-11-21
delete0
delete
OA
AI
J
Jaehee Jung
W
Wooil Kim *
DOI:10.1186/s13636-025-00431-4delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
This paper introduces a novel speaker embedding loss function designed to improve the performance of end-to-end neural diarization (EEND) systems by enhancing speaker discrimination. Unlike previous methods that require additional speaker embedding networks or pre-training on large-scale speaker-labeled datasets, the proposed approach derives speaker-wise embeddings directly from frame-level encoder outputs and ground truth speaker labels. By avoiding external embedding networks, the architecture remains simple and efficient while still facilitating effective learning of speaker characteristics within a unified framework. The proposed loss incorporates cosine similarity to maximize the inter-speaker embedding distance, encouraging the model to learn discriminative embeddings for each speaker, and is integrated with a permutation-free loss to resolve label ambiguity during training. Experimental evaluations were conducted on both simulated and real-world datasets, including LibriSpeech and CALLHOME. The results demonstrate that the proposed speaker embedding loss significantly improves diarization accuracy, achieving a 20.8% relative reduction in Diarization Error Rate (DER) on the two-speaker LibriSpeech dataset (4.99% ->\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\rightarrow$$\end{document} 3.95%). Consistent performance gains were also observed in five-speaker and CALLHOME evaluation settings. These findings underscore the potential of the proposed speaker embedding loss as a lightweight yet effective addition to EEND systems for improved diarization in overlapping and conversational speech scenarios.
Keywords:
Speaker diarization
End-to-end neural diarization
Speaker embedding loss
Speaker label
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

E
EURASIP Journal on Audio Speech and Music Processing
IF:
1.9
Papers:
22
Citations:
0

Organization

I
incheon national university
Scholars:
3.9K
Papers: 4.3K
Citations: 4