1
Return

Fabric2Fashion: Integrating Multimodal Large Language Models and Diffusion Transformers for Controllable Garment Synthesis

delete2026-05-07
delete0
PRE
AI
Y
Yinyin Sun *
DOI:10.1016/j.patrec.2026.05.004delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• research highlights item 1 Leverages Qwen2.5-VL-72B-Instruct in zero-shot setting to extract fabric attributes and generate seven-element structured prompts (fabric, style, model, background, lighting, accessories, mood), establishing precise semantic guidance for downstream garment synthesis and significantly enhancing user-need alignment. • Research highlights item 2 Introduces a dynamic, iterative channel-dimension fabric reference loop that injects original fabric features at each denoising step via ACE++’s long-context conditioning module. Unlike static ControlNet-based methods, this mechanism progressively reinforces high-frequency details throughout the diffusion trajectory, effectively resolving texture-structure misalignment. • Research highlights item 3 Employs Diffusion Transformer architecture with multi-scale hierarchical denoising to simultaneously enhance high-frequency fabric textures and preserve low-frequency style elements, achieving superior material realism and structural coherence under complex garment configurations.
Keywords:
Fabric attribute extraction
Diffusion Transformer
Multimodal large language models
Controllable garment synthesis
Texture-structure alignment

Journal

Pattern Recognition Letters cover
Pattern Recognition Letters
IF:
3.3
Papers:
7.8K
Citations:
1.6W

Organization

Q
qingdao university
Scholars:
5.0K
Papers: 1.5K
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers