arrow
Return

Efficient Parallel Audio Generation Using Group Masked Language Modeling

delete2024-01-01
delete0
delete
OA
AI
M
Myeonghun Jeong
M
Minchan Kim
J
Joun Yeop Lee
N
Nam Soo Kim *
DOI:10.1109/LSP.2024.3381910delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
We present a fast and high-quality codec language model for parallel audio generation. While SoundStorm, a state-of-the-art parallel audio generation model, accelerates inference speed compared to autoregressive models, it still suffers from slow inference due to iterative sampling. To resolve this problem, we propose Group-Masked Language Modeling (G-MLM) and Group Iterative Parallel Decoding (G-IPD) for efficient parallel audio generation. Both the training and sampling schemes enable the model to synthesize high-quality audio with a small number of iterations by effectively modeling the group-wise conditional dependencies. In addition, our model employs a cross-attention-based architecture to capture the speaker style of the prompt voice and improves computational efficiency. Experimental results demonstrate that our proposed model outperforms the baselines in prompt-based audio generation.
Keywords:
Parallel audio generation
neural audio codec

Journal

IEEE Signal Processing Magazine cover
IEEE Signal Processing Magazine
IF:
9.6
Papers:
1.1W
Citations:
1.7W

Organization

S
samsung
Scholars:
8.6K
Papers: 6.4K
Citations: 8
S
seoul national university (snu)
Scholars:
7.2W
Papers: 6.6W
Citations: 86