返回
DETrack: Depth information is predictable for tracking
DOI:10.1016/j.neucom.2024.128906.png)
摘要
En 中文
The purpose of multi-object tracking lies in the estimation of both the bounding boxes of targets and their identities. Nonetheless, occlusion brought by the object interactions often cause identity switches and trajectory loss. Inspired by the human vision of three-dimensional tracking properties, we propose a tracking framework based on depth estimation called DETrack to address this issue. This framework features a Depth Information Module (DIM) under monocular conditions, which can produce depth features as an association cue for multi- object tracking. In addition, to actively retrieves information lost in trajectories, we have also put forward a refind component, which echoes how human vision compensates for objects out of sight. Our framework can seamlessly integrate with most trackers, and introduce introducing an entirely new data dimension to the tracking task. We have tested DETrack using the MOT17 and DanceTrack benchmark datasets and compared it with alternative methods. The test results demonstrate that our technique works effectively with current MOT trackers, and it significantly enhances tracking results based on HOTA, IDF1, and MOTA metrics on both datasets.
Keyword:
Multi-object tracking
Depth estimation
Depth prediction
Occlusion problem
Human vision properties
期刊
IF:
6.5
论文数:
2.5W
被引数:
6.5W

