1
Return

DualSpec: Text-to-Spatial-Audio Generation via Dual-Spectrogram Guided Diffusion Model

delete2026-03-12
delete0
PRE
AI
L
Lei Zhao
S
Sizhou Chen
L
Linfeng Feng
J
Jichao Zhang
X
Xiao-Lei Zhang
C
Chi Zhang
X
Xuelong Li
DOI:10.1109/tmm.2026.3673480delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Text-to-audio (TTA), which generates audio signals from textual descriptions, has received huge attention in recent years. However, recent works focused on text to monaural audio only. As we know, spatial audio provides more immersive auditory experience than monaural audio, e.g. in virtual reality. To address this issue, we propose a text-to-spatial-audio (TTSA) generation framework named DualSpec. Specifically, it first trains variational autoencoders (VAEs) for extracting the latent acoustic representations from sound event audio. Then, given text that describes sound events and event directions, the proposed method uses the encoder of a pretrained large language model to transform the text into text features. Finally, it trains a diffusion model from the latent acoustic representations and text features for the spatial audio generation. In the inference stage, only the text description is needed to generate spatial audio. Particularly, to improve the synthesis quality and azimuth accuracy of the spatial sound events simultaneously, we propose to use two kinds of acoustic features. One is the Mel spectrograms which is good for improving the synthesis quality, and the other is the short-time Fourier transform spectrograms which is good at improving the azimuth accuracy. We provide a pipeline of constructing spatial audio dataset with text prompts, for the training of the VAEs and diffusion model. We also introduce new spatial-aware evaluation metrics to quantify the azimuth errors of the generated spatial audio recordings. Experimental results demonstrate that the proposed method can generate spatial audio with high directional and event consistency.
Keywords:
Text-to-spatial-audio
audio generation
latent diffusion model

Journal

IEEE Transactions on Multimedia cover
IEEE Transactions on Multimedia
IF:
9.7
Papers:
4.4K
Citations:
2.4W

Organization

T
the university of sydney
Scholars:
2.0K
Papers: 908
Citations: 0
N
northwestern polytechnical university
Scholars:
1.0W
Papers: 3.8K
Citations: 0
C
China Telecom
Scholars:
70
Papers: 61
Citations: 121
Cited Papers

Cited Papers

Citing Papers

Citing Papers