Return
Visual tracking with dumbbell selection network
DOI:10.1016/j.neucom.2022.10.031.png)
Abstract
En 中文
The Siamese network-based trackers aim to train the convolutional neural network offline to match tar-get templates and search regions. Recent researches have managed to adopt deep neural networks to extract sufficient semantic information for the Siamese trackers. However, most of the current methods suffer drastic target appearance variations for 1) failing to encode sufficient localization information of targets and 2) neglecting powerful cross-channel interaction information, thus reducing the learned fea-tures' discriminative and representative ability the tracking accuracy. To remedy these issues, in this arti-cle, we propose a Dumbbell Selection Network (DuStNet) by exploring the correlation of the hierarchies of convolutional layers. In concrete, an adaptively Dumbbell Selection mechanism is presented to deal with the targets' appearance deformation by providing rich semantic and localization information. Furthermore, a CSResNet is developed to improve the residual unit in backbones by strengthening the interdependence between the convolution feature channels. We ingeniously employ the Generalized Intersection over Union (GIoU) to supervise the cross-layer feature-map selection, improving tracking accuracy when utilized as a regression loss simultaneously. Our results suggest that the proposed method is robust to significant appearance variations and can generate more accurate bounding boxes in compli-cated scenarios. Extensive experimental results on large-scale benchmark datasets prove our method's effectiveness, which achieves excellent performance on OTB2015, VOT2017, LaSOT, and VOT2019. (C) 2022 Elsevier B.V. All rights reserved.
Keywords:
Visual object tracking
Siamese network
Dumbbell selection
Cropping -Squeeze residual unit
Regression loss
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W

