Return
VGen-adapter: A vision generalization adapter for stable diffusion 3
DOI:10.1016/j.neucom.2025.131948.png)
Abstract
En 中文
Recent years have witnessed rapid advancements in large-scale text-to-image generation models, leading to substantial improvements in generated image quality and the continuous enrichment of the associated ecosystem. Concurrently, various control methods tailored for image generation have emerged, enabling diverse control tasks based on input images. The rise of Transformer architectures has led to the progressive replacement of traditional U-Net-based generative models with Diffusion Transformer (DiT) frameworks. However, most existing control methods face significant challenges in directly adapting to DiT architectures, commonly exhibiting limitations such as inadequate control precision, inefficient training processes, excessive parameter scales, and heavy reliance on large training datasets. To address these issues, we introduce a lightweight image control adapter method tailored for DiT-based Stable Diffusion 3 frameworks, termed VGen-adapter (Vision Generalization adapter). The proposed method incorporates an image feature extractor and optimizes the feature fusion strategy, significantly enhancing control precision. Regarding training strategy, VGen-adapter employs a parameter-efficient approach by completely freezing all parameters of the original model while implementing a lightweight design for the adapter module. Through the integration of noise perturbation for data augmentation, our approach enables high-quality image generation using limited text-image datasets. Furthermore, with only 95 M parameters (approximately 5 % of the SD3-Medium model), this architecture achieves substantial computational efficiency. Comparative experiments on the COCO dataset demonstrate that the VGen-adapter achieves excellent control performance while maintaining high efficiency and low parameter count, verifying its feasibility and effectiveness under the DiT architecture.
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W

