Return
Fabric2Fashion: Integrating Multimodal Large Language Models and Diffusion Transformers for Controllable Garment Synthesis
Y
DOI:10.1016/j.patrec.2026.05.004.png)
Abstract
En 中文
• research highlights item 1 Leverages Qwen2.5-VL-72B-Instruct in zero-shot setting to extract fabric attributes and generate seven-element structured prompts (fabric, style, model, background, lighting, accessories, mood), establishing precise semantic guidance for downstream garment synthesis and significantly enhancing user-need alignment. • Research highlights item 2 Introduces a dynamic, iterative channel-dimension fabric reference loop that injects original fabric features at each denoising step via ACE++’s long-context conditioning module. Unlike static ControlNet-based methods, this mechanism progressively reinforces high-frequency details throughout the diffusion trajectory, effectively resolving texture-structure misalignment. • Research highlights item 3 Employs Diffusion Transformer architecture with multi-scale hierarchical denoising to simultaneously enhance high-frequency fabric textures and preserve low-frequency style elements, achieving superior material realism and structural coherence under complex garment configurations.
Keywords:
Fabric attribute extraction
Diffusion Transformer
Multimodal large language models
Controllable garment synthesis
Texture-structure alignment
Journal
IF:
3.3
Papers:
7.8K
Citations:
1.6W
