Return
An Empirical Study of Validating Synthetic Data for Text-Based Person Retrieval
M
Y
Z
D
J
M
DOI:10.1109/tifs.2026.3715117.png)
Abstract
En 中文
Data plays a pivotal role in Text-Based Person Retrieval (TBPR) research. The mainstream research paradigm heavily relies on real-world person images with manual textual annotations for training models, posing privacy concerns and annotation burdens. Several pioneering efforts have explored synthetic data generation, yet they still depend on real data as a foundation, inheriting the same limitations. The feasibility of training TBPR models purely on synthetic data remains unexplored, and there is currently no systematic study on the effectiveness boundaries of synthetic data across various real-world scenarios. In this work, we present the first comprehensive empirical study of synthetic data for TBPR, focusing on two key dimensions. 1) We propose a unified data synthesis pipeline that can operate entirely without real person data. It combines an inter-class image generation module that produces diverse identity-centric images by means of an automatic prompt construction strategy, and an intra-class augmentation module that enhances identity variation through text-driven image editing. 2) Leveraging this pipeline and automatic textual description generation, we explore the effectiveness of synthetic data in diverse scenarios through extensive experiments, to reveal its practical utility as either a standalone replacement or a complementary augmentation to real data. We release the code, along with the synthetic large-scale dataset generated by our pipeline in <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/Flame-Chasers/SynTBPR</uri>
Keywords:
Text-based person retrieval
data generation
empirical study
Journal
IF:
8
Papers:
5.2K
Citations:
2.3W
