Return
Sequence-to-Sequence Voice Conversion With Weighted Guided Attention
DOI:10.1109/ACCESS.2025.3647153.png)
Abstract
En 中文
Voice conversion (VC) using non-autoregressive sequence-to-sequence (S2S) models has gained attention in recent years. Compared with conventional framewise VC models, these S2S VC models can control features, such as prosody and speech speed, and enable high-quality and fast conversion. In S2S VC, the accuracy of alignment between source and target speakers strongly affects conversion performance. In this paper, we propose EdenVC and WSAS-VC using attention-based alignment. EdenVC simply introduces a diagonal-guided weighting initially proposed in EdenTTS, while WSAS-VC integrates the attention weights and those based on the alignment path obtained from the monotonic alignment search. We evaluate the conventional MAS-VC, EdenVC, and WSAS-VC using two types of speaker conversion conditions, from native to native speech and from non-native to native speech. Experimental results show that WSAS-VC achieves more stable training than EdenVC and better conversion accuracy than MAS-VC.
Keywords:
Viterbi algorithm
Artificial intelligence
computational and artificial intelligence
learning (artificial intelligence)
learning (artificial intelligence)
speech synthesis
speech synthesis
Viterbi algorithm
speech synthesis
Journal
IF:
3.6
Papers:
9.8W
Citations:
29.4W

