Return
Prototype Division for Self-Supervised Speaker Verification
DOI:10.1109/LSP.2024.3377593.png)
Abstract
En 中文
Self-supervised learning has shown promising performance on speaker verification tasks, among which Self-DIstillation with NO labels (DINO) is currently a widely adopted framework. As one of the unsupervised deep clustering methods, the number of valid prototypes in DINO is far less than the speakers in practical applications and remains unchanged throughout the training period, leading to severe speaker confusion and performance degradation. Therefore, a strategy named prototype division (PD) is proposed to iteratively generate fine-grained prototypes in the projection space based on the converged model to separate confused categories, where new prototypes are derived from the neighborhood of the existing valid prototypes by clustering or sampling. The results on Vox1O achieve significant improvements, relatively outperforming the baseline by 31.1% without any auxiliary loss. Further experiments on CN-Celeb also show stable improvement, proving the consistency of the proposed method.
Keywords:
DINO
prototype division
self-supervised learning
speaker verification
Journal
IF:
9.6
Papers:
1.1W
Citations:
1.7W
Organization
Cited Papers
Novel intramolecular blocked isocyanates as stable one-component systems for poly(urea urethane)s
Polymer
IF0
Diaziquone given as a continuous infusion is an active agent for relapsed adult acute nonlymphocytic leukemia
Blood
IF0
Genome-Wide Analysis and Exploration of WRKY Transcription Factor Family Involved in the Regulation of Shoot Branching in Petunia
Genes
IF0

