Return
DiffMusic: Efficient Music Generation From a Single Image Using Diffusion-Based Representations
DOI:10.1109/TASLPRO.2026.3660263.png)
Abstract
En 中文
In this paper, we present DiffMusic, a novel methodology for generating high-quality music from a single image. Existing methodologies have achieved multi-modality by integrating data from various domains or employing high-cost approaches, such as Large Language Models (LLMs), to generate high-quality music. The proposed DiffMusic adopts a diffusion-based approach to generate music descriptions from a single image, aiming to solve the issues encountered in traditional music generation methods. Unlike traditional image captioning, this approach does not simply describe the scene within the image. Instead, it generates descriptions that capture genre-related, melodic, and rhythmic elements essential for music generation. This approach enables the direct conversion of an image into music with a single inference process while allowing seamless integration into various existing music generators. Experimental results demonstrate that our method outperforms existing approaches by 29.9% in terms of Frechet Audio Distance, while remaining more cost-effective.
Keywords:
Visualization
Music
Speech processing
Videos
Training
Diffusion models
Vectors
Instruments
Translation
Transformers
Music generation
diffusion-based representations
large language models
Journal
I
IF:
0
Papers:
151
Citations:
0

