返回
Pseudo loss active learning for deep visual tracking
DOI:10.1016/j.patcog.2022.108773.png)
摘要
En 中文
In visual tracking tasks, the training data are commonly composed of a large number of video sequences and each frame in the sequences needs to be labeled manually, which is labor-intensive and time-consuming. In addition, considering the similarity among the consecutive frames in the same sequence, there is significant redundancy in the training data. To address these problems, a novel pseudo loss active learning (PLAL) method is developed in this paper. PLAL aims to select the most informative and least redundant data for training to reduce the cost of labeling and maintain competitive tracking results simultaneously. Firstly, the Gaussian distribution based pseudo label is generated for the unlabeled candidates based on the tracking model which is initially trained on a small amount of training data. Then, the pseudo loss based on cross entropy is designed to compute the difference between the pseudo label and the target response map. The pseudo loss measures the uncertainty of the target spatial context which is used as the informativeness criterion of the image frame for selection. Meanwhile, a sampling interval threshold and a temporal penalty are employed for frame selection to avoid drastic variation in target appearance and reduce the redundancy within the consecutive candidate frames. Only the selected frames are labeled by the oracle (human expert) and then added to the training data. Extensive experiments on public benchmarks (OTB2013, OTB2015, VOT2018, UAV123, GOT-10K, TrackingNet, LaSOT, OxUvA and TLP) demonstrate that PLAL method outperforms the baseline and other recent active learning approaches. With only 3% of labeled data from the training dataset, PLAL reaches competitive performance (98-100%) compared to the model trained on the entire training dataset. (C) 2022 Elsevier Ltd. All rights reserved.
Keyword:
Active learning
Visual tracking
Pseudo loss
Pseudo label
期刊
IF:
7.6
论文数:
1.3W
被引数:
4.5W
机构
引用论文
Siamese network for object tracking with multi-granularity appearance representations具有多粒度外观表示的Siamese对象跟踪网络
PATTERN RECOGNITION
IF7.6
MeMu: Metric correlation Siamese network and multi-class negative sampling for visual tracking
PATTERN RECOGNITION
IF7.6
Joint temporal context exploitation and active learning for video segmentation联合时域上下文开发和主动学习的视频分割
PATTERN RECOGNITION
IF7.6
没有更多内容

