Return
DiffBlender: Composable and versatile multimodal text-to-image diffusion models
DOI:10.1016/j.eswa.2025.129345.png)
Abstract
En 中文
• Introduces DiffBlender to unify multiple input modalities—structure, layout, and attribute—within a single T2I framework. • Utilizes a compact “Blender block” that preserves the pre-trained diffusion parameters, minimizing additional training overhead. • Enables efficient multimodal generation and composability across diverse conditions and user preferences. • Proposes mode-specific guidance for precise control over each modality, ensuring balanced and high-fidelity image synthesis.
Keywords:
DiffBlender
multimodal generation
text-to-image synthesis
diffusion models
modality-specific guidance
Journal
IF:
7.5
Papers:
2.9W
Citations:
10.2W

