arrow
返回

Efficient Vision Transformer with Token Sparsification for Event-Based Object Tracking

delete2026-01-20
delete0
PRE
AI
J
Jiqing Zhang
X
Xin Yang
H
Haoming Tang
Y
Yuanchen Wang
H
Huibing Wang *
X
Xianping Fu
DOI:10.1007/s11263-025-02666-2delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The Vision Transformer has gained popularity as a neural network architecture for event-based vision tasks. However, its use on resource-constrained devices is limited due to high computational and memory costs. This paper presents an Efficient Vision Transformer, a novel backbone for object tracking with event cameras. Specifically, we propose two adaptive token sparsification strategies based on the inherent characteristics of event data and tracking tasks, thereby reducing the computation while maintaining comparable performance. Firstly, since events are spatially sparse at pixel locations, we adaptively predict the sparsification ratio based on statistical entropy analysis and subsequently remove search tokens exhibiting dissimilarity to template tokens. Besides, we further eliminate redundant tokens by estimating the importance score of each token given the tracking targets. By hierarchically pruning more than 60% of the input tokens of vanilla Vision Transformer (ViT), our method significantly reduces MAC operations by approximately 25%, while the tracking accuracy drops by no more than 0.5%. Extensive experiments on various event-based tracking datasets demonstrate that our proposed approach outperforms existing state-of-the-art methods in both tracking accuracy and speed.
Keyword:
Event-based camera
Visual object tracking
Transformer

期刊

International Journal of Computer Vision 封面图
International Journal of Computer Vision
IF:
9.3
论文数:
3.9K
被引数:
2.8W

机构

D
dalian maritime university
学者数:
1.4K
论文数: 547
被引数: 0
D
Dalian University of Technology
学者数:
6.0W
论文数: 4.4W
被引数: 5.5W
引用论文

引用论文

AiATrack: Attention in Attention for Transformer Visual Tracking
err2022-10-23
err0
PREAI
errShenyuan Gao; Chunluan Zhou; Chao Ma; Xinggang Wang; Junsong Yuan
err分享
err收藏
Frame-Event Alignment and Fusion Network for High Frame Rate Tracking
err2023-06-01
err0
errOAAI
errJiqing Zhang; Yuanchen Wang; Wenxi Liu; Meng Li; Jinpeng Bai; Baocai Yin; Xin Yang
err分享
err收藏
Adaptive Vision Transformer for Event-Based Human Pose Estimation
err2024-10-28
err0
PREAI
errYu,Nannan; Ma,Tao; Zhang,Jiqing; Zhang,Yuji; Bao,Qirui; Wei,Xiaopeng; Yang,Xin
err分享
err收藏
Autoregressive Visual Tracking
err2023-06-01
err0
PREAI
errXing Wei; Yifan Bai; Yongchao Zheng; Dahu Shi; Yihong Gong
err分享
err收藏
A-ViT: Adaptive Tokens for Efficient Vision Transformer
err2022-06-01
err0
PREAI
errHongxu Yin; Arash Vahdat; Jose M. Alvarez; Arun Mallya; Jan Kautz; Pavlo Molchanov
err分享
err收藏
Flow-Guided Transformer for Video Inpainting
err2022-11-03
err0
errOAAI
errKaidong Zhang; Jingjing Fu; Dong Liu
err分享
err收藏
ODTrack: Online Dense Temporal Token Learning for Visual Tracking
err2024-03-24
err0
errOAAI
errYaozong Zheng; Bineng Zhong; Qihua Liang; Zhiyi Mo; Shengping Zhang; Xianxian Li
err分享
err收藏
学者 查看更多内容