arrow
Return

Language Supervised Multi-Camera Multi-Object Tracking

delete2026-06-08
delete0
PRE
AI
K
Kaige Mao
X
Xiaopeng Hong
范晓鹏 (Xiaopeng Fan)
左旺孟 (Wangmeng Zuo)
DOI:10.1109/TIP.2026.3699088delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Recent multi-camera multi-object tracking (MCMOT) algorithms are primarily trained using per-detection identity annotations, which are complicated to obtain. In contrast, labeling a language description per-object is a more natural and human-friendly way. In this paper, we explore MCMOT in a language-supervised manner (LS-MCMOT) and propose a novel approach LaVST, which performs language-to-vision weakly-supervised learning based on reliable pseudo-labels generated via tracklet-level cross-modality matching. In addition, we design an ID-aware projection self-correction mechanism to correct inaccurate image-to-ground projection in a self-supervised manner. The models trained with our approach exhibit promising performance in LS-MCMOT. Surprisingly, they perform favorably against state-of-the-art identity-supervised methods, especially in cross-dataset evaluation (with an average gain by 20.0% in IDF1), underscoring the potential of language annotations in MCMOT. Codes and language annotations will be available here.
Keywords:
Multi-camera multi-object tracking
language supervision
tracklet-level matching
projection self-correction

Journal

IEEE Transactions on Image Processing cover
IEEE Transactions on Image Processing
IF:
13.7
Papers:
1.0W
Citations:
8.4W

Organization

H
Harbin Institute of Technology
Scholars:
1.3W
Papers: 4.3K
Citations: 8.5W