arrow
Return

Auto-Locate: A Training-Free Multi-instance Generation for Text-to-Image Diffusion Models

delete2026-01-01
delete0
PRE
AI
T
Tao, Xiangzhi *
K
K. Wang
Z
Zhongyang Hu
N
Naijie Gu
DOI:10.1007/978-981-95-5696-0_18delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Text-to-image diffusion models demonstrate remarkable capabilities in generating high-quality images. However, diffusion models struggle with generating the correct number of instances, leading to issues such as missing or excessive instances, subject fusion, and incorrect attribute binding. These issues can be attributed to incorrect instance numbers, indicating that generating the correct number of instances can effectively avoid them. To address these issues, this paper proposes a training-free method that enables diffusion models to generate images with the correct number of instances. The method automatically identifies non-overlapping instance regions during inference. These regions are then refined via constraint loss functions applied to cross-attention maps. The proposed method, named Auto-Locate, generalizes well to multi-subject and multi-instance scenarios, enabling diffusion models to better handle complex text prompts involving diverse entities. Extensive experiments show that Auto-Locate effectively controls the number of instances in generated images while maintaining high fidelity and diversity, outperforming several baselines on standard benchmarks.
Keywords:
Diffusion models
Image generation
Cross attention
Multiple instances

Journal

P
PATTERN RECOGNITION AND COMPUTER VISION, PRCV 2025, PT III
IF:
0
Papers:
25
Citations:
0

Organization

U
university of science & technology of china, cas
Scholars:
3.2W
Papers: 2.7W
Citations: 74
C
chinese academy of sciences
Scholars:
56.5W
Papers: 44.9W
Citations: 704