Return
Feature-Enhanced Diffusion Model for Text-Guided Sound Effect Generation
DOI:10.3390/electronics15071358.png)
Abstract
En 中文
This study proposes a feature-enhanced diffusion model based on wavelet transform and Mamba to address the issues of low audio realism, inadequate text relevance, and slow inference speed in text-guided sound effect generation. A wavelet transform-based downsampling module is designed to mitigate the loss of high-frequency feature information during the downsampling process of the diffusion model, thereby enhancing the realism of the generated audio. A multi-scale feature extraction and fusion method is employed to capture both local and global acoustic information, while the channel attention mechanism further strengthens the model’s focus on text-relevant key features. Additionally, an optimization method based on Mamba and adaptive weight adjustment is proposed, which takes advantage of Mamba’s efficient information processing mechanism and learnable parameters to optimize skip connections, improving model training and inference efficiency without adding substantial computational cost. Experiments show that the model achieves FAD and KL scores of 1.608 and 1.609, respectively, reflecting improvements of 33.8% and 26.1% compared to the baseline model.
Keywords:
sound effect generation
text-guided
wavelet transform
Mamba
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
2.6
Papers:
9.6K
Citations:
4.7W

