arrow
Return

A Robust Visual Tracking Method Based on Reconstruction Patch Transformer Tracking

delete2022-08-31
delete4
delete
OA
AI
H
Hui Chen
Z
Zhenhai Wang *
H
Hongyu Tian
L
Lutao Yuan
X
Xing Wang
P
Peng Leng
DOI:10.3390/s22176558delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Recently, the transformer model has progressed from the field of visual classification to target tracking. Its primary method replaces the cross-correlation operation in the Siamese tracker. The backbone of the network is still a convolutional neural network (CNN). However, the existing transformer-based tracker simply deforms the features extracted by the CNN into patches and feeds them into the transformer encoder. Each patch contains a single element of the spatial dimension of the extracted features and inputs into the transformer structure to use cross-attention instead of cross-correlation operations. This paper proposes a reconstruction patch strategy which combines the extracted features with multiple elements of the spatial dimension into a new patch. The reconstruction operation has the following advantages: (1) the correlation between adjacent elements combines well, and the features extracted by the CNN are usable for classification and regression; (2) using the performer operation reduces the amount of network computation and the dimension of the patch sent to the transformer, thereby sharply reducing the network parameters and improving the model-tracking speed.
Keywords:
transformer
cross-attention
CNN
transformer-based tracker
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Sensors cover
Sensors
IF:
3.5
Papers:
7.1W
Citations:
20.9W

Organization

L
linyi university
Scholars:
4.4K
Papers: 3.1K
Citations: 62
Z
zhejiang university
Scholars:
17.4W
Papers: 12.0W
Citations: 152