arrow
Return

Learning Compositionality from Multifaceted Synthetic Data for Language-based Object Detection

delete2025-08-12
delete0
PRE
AI
K
Kwanyong Park
S
Sojung An
Y
Yong Jae Lee
D
Donghyun Kim *
DOI:10.1007/s11263-025-02554-9delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Language-based object detection aims to locate target objects from complex language queries. However, current vision-language detectors often struggle to understand complex representations of visual objects (e.g., attributes, shapes, and relationships), especially under complex queries. In this paper, we first conduct a thorough analysis of current language-based detectors to identify their specific weaknesses in compositional understanding. To this end, we propose a novel comprehensive evaluation framework that automatically categorizes test cases by the type and complexity of compositionality, leveraging large language models (LLMs). This reveals that detectors show significant performance drops with increased complexity and consistent failures in specific types, such as spatial and numerical reasoning. To effectively address this, we propose a multifaceted synthetic data consisting of (1) generative model-based synthetic triplets that inherited compositional knowledge from large generative models (e.g., LLMs, diffusion models) in the form of triplets (i.e., image-text-box data); and (2) weakness-targeted synthetic descriptions designed to enhance understanding in vulnerable types like spatial and numeracy concepts. We further introduce a compositional contrastive learning method to better leverage the proposed synthetic data while mitigating the common drawbacks of synthetic data. Consequently, our models trained on proposed multifaceted synthetic data exhibit a significant performance boost in the Omnilabel benchmark by up to +7.1AP and the $$\hbox {D}^{3}$$ benchmark by up to $$+8.4$$ AP upon existing baselines.
Keywords:
Multimodal Synthetic Datasets
Compositional Learning
Transfer Learning
Language-based Object Detection

Journal

International Journal of Computer Vision cover
International Journal of Computer Vision
IF:
9.3
Papers:
3.9K
Citations:
2.8W

Organization

K
Korea University
Scholars:
3.6W
Papers: 3.8W
Citations: 4.4W
U
University of Seoul
Scholars:
3.5K
Papers: 4.5K
Citations: 4.4K
U
university of wisconsin-madison
Scholars:
3.5K
Papers: 1.5K
Citations: 2
researcher View more organizations
Cited Papers

Cited Papers

Structured Domain Randomization: Bridging the Reality Gap by Context-Aware Synthetic Data
err2019-05-01
err0
errOAAI
errAayush Prakash; Shaad Boochoon; Mark Brophy; David Acuna; Eric Cameracci; Gavriel State; Omer Shapira; Stan Birchfield
errShare
errSave
High-Resolution Image Synthesis with Latent Diffusion Models
err2022-06-01
err0
errOAAI
errRobin Rombach; Andreas Blattmann; Dominik Lorenz; Patrick Esser; Bjorn Ommer
errShare
errSave
Sigmoid Loss for Language Image Pre-Training
err2023-10-01
err0
errOAAI
errXiaohua Zhai; Basil Mustafa; Alexander Kolesnikov; Lucas Beyer
errShare
errSave
Scaling Laws of Synthetic Images for Model Training … for Now
err2024-06-16
err0
PREAI
errLijie Fan; Kaifeng Chen; Dilip Krishnan; Dina Katabi; Phillip Isola; Yonglong Tian
errShare
errSave
End-to-End Object Detection with Transformers
err2020-11-03
err0
PREAI
errNicolas Carion; Francisco Massa; Gabriel Synnaeve; Nicolas Usunier; Alexander Kirillov; Sergey Zagoruyko
errShare
errSave
Task2Sim: Towards Effective Pre-training and Transfer from Synthetic Data
err2022-06-01
err0
errOAAI
errSamarth Mishra; Rameswar Panda; Cheng Perng Phoo; Chun-Fu Richard Chen; Leonid Karlinsky; Kate Saenko; Venkatesh Saligrama; Rogerio S. Feris
errShare
errSave
OmniLabel: A Challenging Benchmark for Language-Based Object Detection
err2023-10-01
err0
PREAI
errSchulter,Samuel; G,Vijay Kumar B; Suh,Yumin; Dafnis,Konstantinos M.; Zhang,Zhixing; Zhao,Shiyu; Metaxas,Dimitris
errShare
errSave
MDETR - Modulated Detection for End-to-End Multi-Modal Understanding
err2021-10-01
err0
errOAAI
errAishwarya Kamath; Mannat Singh; Yann LeCun; Gabriel Synnaeve; Ishan Misra; Nicolas Carion
errShare
errSave
Teaching Structured Vision & Language Concepts to Vision & Language Models
err2023-06-01
err0
PREAI
errSivan Doveh; Assaf Arbelle; Sivan Harary; Eli Schwartz; Roei Herzig; Raja Giryes; Rogerio Feris; Rameswar Panda; Shimon Ullman; Leonid Karlinsky
errShare
errSave
researcher View more