arrow
Return

Improving Object Detection Models via LLM-Based Training Data Synthesis

delete2025-09-10
delete0
PRE
AI
S
Shi-Ran Ge *
J
Jie Cao *
赫然 cover
赫然 (Ran He)
DOI:10.1007/s11263-025-02560-xdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Despite significant advancements in deep generative models, generating high-quality training data for object detection remains a challenging task, primarily due to the complex requirements of precise annotations and diverse scenes. To address this challenge, we propose a novel framework for generating high-quality training data using Large Language Models (LLMs). First, we introduce the Layout Enhancement and Diverse Imagery Synthesis Framework (LE-DIS), which leverages LLMs to create diverse target scenes and systematically constructs synthetic data. Next, we propose a CLIP-based image-layout quality metric (CILQM) to evaluate the global consistency and category alignment of synthetic data, ensuring high-quality outputs. Finally, we employ a mixup-based strategy (SRMix) that integrates synthetic and real data to produce diverse training samples, enhancing the model’s stability and adaptability. Extensive experiments on the COCO benchmark demonstrate that our approach significantly improves the performance of both Transformer-based and CNN-based object detection models, highlighting the potential of deep generative models in synthesizing high-quality datasets for object detection tasks.
Keywords:
Dataset Synthesis
Object Detection
Image Generation
Large Language Model
Diffusion Model

Journal

International Journal of Computer Vision cover
International Journal of Computer Vision
IF:
9.3
Papers:
3.9K
Citations:
2.8W

Organization

I
Institute of Automation
Scholars:
539
Papers: 287
Citations: 220
Cited Papers

Cited Papers

A Survey on Evaluation of Large Language Models
err2024-03-29
err232
errOAAI
errChang, Yupeng; Wang, Xu; Wang, Jindong; Wu, Yuan; Yang, Linyi; Zhu, Kaijie; Chen, Hao; Yi, Xiaoyuan; Wang, Cunxiang; Wang, Yidong; Ye, Wei; Zhang, Yue; Chang, Yi; Yu, Philip S.; Yang, Qiang; Xie, Xing
errShare
errSave
errShare
errSave
Image Generation: A Review
err2022-03-11
err0
PREAI
errMohamed Elasri; Omar Elharrouss; Somaya Al-Maadeed; Hamid Tairi
errShare
errSave
Layout2image: Image Generation from Layout
err2020-02-24
err11
errOAAI
errZhao, Bo; Yin, Weidong; Meng, Lili; Sigal, Leonid
errShare
errSave