Return
Weakly Supervised Composed Object Re-Identification With Large Models
Y
J
Q
张
DOI:10.1109/tcyb.2026.3681841.png)
Abstract
En 中文
Existing object re-identification (re-ID) and composed image retrieval (CIR) methods capture different aspects of real-world retrieval requirements; re-ID preserves identity but cannot specify desired appearance changes, whereas CIR supports attribute-guided retrieval but does not enforce identity consistency. To bridge this gap, we introduce <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">composed object re-identification</i> (CORI), a new task that requires the retrieved target to simultaneously satisfy identity preservation and text-guided attribute modification. This problem is fundamentally different from existing re-ID and CIR settings and has not been explicitly studied before. To make CORI feasible without costly manual annotation, we propose a weakly supervised framework that leverages large language models (LLMs) and visual question answering (VQA) models to automatically generate reference-to-target descriptions using ID labels alone. We further develop the first baseline model tailored for CORI, which jointly learns multimodal composition and identity-aware matching through shared-weight image encoders, a text encoder, and a compositor module optimized by contrastive, ID, and triplet losses. We also establish four CORI benchmark datasets covering person and vehicle retrieval. Experiments show that the proposed method consistently outperforms representative baselines adapted from existing CIR and re-ID methods for the newly introduced CORI setting, improving Rank@1 by 2.1% and 0.8% on RAP and Celeb-reID-light, and by 9.9% and 9.5% on VeRi-776 and VRIC, respectively.
Keywords:
Composed object re-identification (CORI)
large models
weak supervision
Journal
IF:
10.5
Papers:
1.1W
Citations:
5.0W
