返回
LaCon: Late-Constraint Controllable Visual Generation
DOI:10.1109/TIP.2026.3654412.png)
摘要
En 中文
Diffusion models have demonstrated impressive abilities in generating photo-realistic and creative images. To offer more controllability for the generation process of diffusion models, previous studies normally adopt extra modules to integrate condition signals by manipulating the intermediate features of the noise predictors, where they often fail in conditions not seen in the training. Although subsequent studies are motivated to handle multi-condition control, they are mostly resource-consuming to implement, where more generalizable and efficient solutions are expected for controllable visual generation. In this paper, we present a late-constraint controllable visual generation method, namely LaCon, which enables generalization across various modalities and granularities for each single-condition control. LaCon establishes an alignment between the external condition and specific diffusion timesteps, and guides diffusion models to produce conditional results based on this built alignment. Experimental results on prevailing benchmark datasets illustrate the promising performance and generalization capability of LaCon under various conditions and settings. Ablation studies analyze different components in LaCon, illustrating its great potential to offer flexible condition controls for different backbones.
Keyword:
Conditional image animation
controllable visual generation
diffusion models
text-to-image generation
期刊
IF:
13.7
论文数:
1.0W
被引数:
8.4W
机构
引用论文
KT-GAN: Knowledge-Transfer Generative Adversarial Network for Text-to-Image SynthesisKt-gan: 用于文本到图像合成的知识转移生成对抗网络
U2-Net: Going deeper with nested U-structure for salient object detectionU2-Net: 基于嵌套U结构的显著性目标检测
PATTERN RECOGNITION
IF7.6

