Return
Texture-preserving multimodal fashion image editing with diffusion models
DOI:10.1016/j.knosys.2025.114769.png)
Abstract
En 中文
Multimodal fashion image editing leverages diverse conditions, including text, human body poses, garment sketches, and fabric textures, to generate personalized clothing designs using deep learning techniques. This process enhances the efficiency of creative design and visual presentation while preserving the model’s identity features. However, current approaches typically employ the CLIP image encoder or a reference network to handle texture-conditional inputs, both of which have notable limitations: (1) the CLIP image encoder struggles to capture fine-grained texture details, resulting in reduced visual fidelity; and (2) the reference network contains a large number of parameters, leading to considerable computational overhead. To address these challenges, we introduce a novel diffusion model, Texture-Preserving Multimodal Garment Designer (TP-MGD). TP-MGD first concatenates the model image with the inpainted region and the texture image along the spatial dimension before feeding them into the diffusion model. The model then utilizes self-attention layers to enable high-fidelity and efficient transfer of texture features. Furthermore, a decoupled cross-attention mechanism is employed to enhance the alignment of texture features. Experimental results demonstrate that TP-MGD achieves a 17.2 % reduction in the FID score while eliminating the need for an image encoder (632.08M parameters) and outperforms existing methods in terms of fidelity and multimodal consistency. The source code is available at: https://github.com/zibingo/TP-MGD .
Journal
K
IF:
7.6
Papers:
1.2W
Citations:
4.5W

