arrow
Return

Efficient Multi-Instance Generation With Janus-Pro-Driven Prompt Parsing

delete2026-02-23
delete0
PRE
AI
亓 帆 cover
亓 帆 (Fan Qi)
Y
Yu Duan
M
Mengxian Li
徐常胜 (Changsheng Xu)
DOI:10.1109/TCSVT.2026.3666039delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Recent advances in text-guided diffusion models have revolutionized conditional image generation, yet they struggle to synthesize complex scenes with multiple objects due to imprecise spatial grounding and limited scalability. We address these challenges through two key modules: 1) Janus-Pro-driven Prompt Parsing, a prompt-layout parsing module that bridges text understanding and layout generation via a compact 1B-parameter architecture, and 2) MIGLoRA, a parameter-efficient plug-in integrating Low-Rank Adaptation (LoRA) into UNet (SD1.5) and DiT (SD3) backbones. MIGLoRA is capable of preserving the base model’s parameters and ensuring plug-and-play adaptability, minimizing architectural intrusion while enabling efficient fine-tuning. To support a comprehensive evaluation, we create DescripBox and DescripBox-1024, benchmarks that span diverse scenes and resolutions. The proposed method achieves state-of-the-art performance on COCO and LVIS benchmarks while maintaining parameter efficiency, demonstrating superior layout fidelity and scalability for open-world synthesis. Here is our project page: <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/FanQi-AI/MIGLoRA</uri>
Keywords:
Diffusion models
multi-instance generation
low-rank adaptation
prompt engineering
layout-to-image generation

Journal

IEEE Transactions on Circuits and Systems for Video Technology cover
IEEE Transactions on Circuits and Systems for Video Technology
IF:
11.1
Papers:
612
Citations:
3.1W

Organization

T
tianjin university of technology
Scholars:
1.8K
Papers: 542
Citations: 0
C
chinese academy of sciences
Scholars:
55.9W
Papers: 44.7W
Citations: 704