arrow
返回

Deep Decoupling Classification and Regression for Visual Tracking

delete2023-09-01
delete1
PRE
AI
G
Guang Han *
R
Ruiyu Yang
H
Hua Gao
S
Sam Kwong
DOI:10.1109/TCDS.2022.3202802delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Classification and regression are two tasks that most Siamese-based trackers need to handle. However, most of the existing trackers only learn one feature embedding to handle these two types of task, making it difficult to optimize both simultaneously. To solve this problem, this article tries to deeply decouple classification and regression in the model structure. Specifically, two feature extraction backbone networks are used to divide the model into two branches to extract the heterogeneous features suitable for the two tasks, respectively. Inspired by the core idea of transformer, information interaction and fusion between multiple branches are achieved by the cross-attention mechanism, which can fully exploit the deep information dependence between multiple branches. In addition, the concept of channel-level information interaction is proposed by innovatively changing the generation mode of vector groups in the attention module. The experiments show that double Siamese tracker (DST) designed in this article greatly improves the accuracy of classification and regression. DST runs at 60 frames per second (FPS) on GPU, far above the real-time requirement.
Keyword:
Attention mechanism
Siamese network
transformer
visual tracking

期刊

IEEE Transactions on Cognitive and Developmental Systems 封面图
IEEE Transactions on Cognitive and Developmental Systems
IF:
4.9
论文数:
1.0K
被引数:
3.5K

机构

Z
zhejiang university of technology
学者数:
3.3W
论文数: 2.0W
被引数: 22
C
City University of Hong Kong
学者数:
2.3W
论文数: 3.0W
被引数: 6.1W
引用论文

引用论文

Trust in Virtual Teams: A Multidisciplinary Review and Integration
err2019-01-21
err0
errOAAI
errJanine Viol Hacker; Michael Johnson; Carol Saunders; Amanda L. Thayer
err分享
err收藏
Definitions of primary-progressive multiple sclerosis trajectories by rate of clinical disability progression
err2021-05-01
err0
PREAI
errAnat Achiron; Sapir Dreyer-Alster; Michael Gurevich; Shay Menascu; David Magalashvili; Mark Dolev; Yael Stern; Tomer Ziv-Baran
err分享
err收藏
An efficient self-attention network for skeleton-based action recognition
err2022-03-08
err14
errOAAI
errQin, Xiaofei; Cai, Rui; Yu, Jiabin; He, Changxiang; Zhang, Xuedian
err分享
err收藏
err分享
err收藏
Object tracking: A survey对象跟踪: 一项调查
err2006-12-25
err4.0K
PREAI
errYilmaz, Alper; Javed, Omar; Shah, Mubarak
err分享
err收藏
ImageNet Large Scale Visual Recognition ChallengeImageNet大规模视觉识别挑战
err2015-04-11
err2.7W
PREAI
errRussakovsky, Olga; Deng, Jia; Su, Hao; Krause, Jonathan; Satheesh, Sanjeev; Ma, Sean; Huang, Zhiheng; Karpathy, Andrej; Khosla, Aditya; Bernstein, Michael; Berg, Alexander C.; Fei-Fei, Li
err分享
err收藏
学者 查看更多内容