返回
Efficient Parallel Audio Generation Using Group Masked Language Modeling
DOI:10.1109/LSP.2024.3381910.png)
摘要
En 中文
We present a fast and high-quality codec language model for parallel audio generation. While SoundStorm, a state-of-the-art parallel audio generation model, accelerates inference speed compared to autoregressive models, it still suffers from slow inference due to iterative sampling. To resolve this problem, we propose Group-Masked Language Modeling (G-MLM) and Group Iterative Parallel Decoding (G-IPD) for efficient parallel audio generation. Both the training and sampling schemes enable the model to synthesize high-quality audio with a small number of iterations by effectively modeling the group-wise conditional dependencies. In addition, our model employs a cross-attention-based architecture to capture the speaker style of the prompt voice and improves computational efficiency. Experimental results demonstrate that our proposed model outperforms the baselines in prompt-based audio generation.
Keyword:
Parallel audio generation
neural audio codec
期刊
IF:
9.6
论文数:
1.1W
被引数:
1.7W
机构
引用论文
Genome-Wide Analysis and Exploration of WRKY Transcription Factor Family Involved in the Regulation of Shoot Branching in Petunia
Genes
IF0
Research on the Efficiency of Wireless Power Transfer System Based on Multi-Auxiliary Transmitting Coils基于多辅助发射线圈的无线电能传输系统效率研究

