arrow
返回

Convolutional transformer network for fine-grained action recognition

delete2024-02-01
delete7
PRE
AI
Y
Yujun Ma
R
Ruili Wang *
M
Ming Zong
W
Wanting Ji
Y
Yi Wang
B
Baoliu Ye
DOI:10.1016/j.neucom.2023.127027delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Fine-grained action recognition is one of the critical problems in video processing, which aims to recognize similar actions of subtle interactions between humans and objects. Inspired by the remarkable performance of the Transformer in natural language processing, Transformer has been applied to the fine-grained action recognition task. However, Transformer needs abundant training data and extra supervision to achieve comparable results with convolutional neural networks (CNNs). To address these issues, we propose a Convolutional Transformer Network (CTN), which integrates the merits of CNN (e.g., sharing weights, capturing low-level features in videos and locality) and the benefits of Transformer (e.g., dynamic attention and learning long-range dependencies). In this paper, we propose two modifications to the original Transformer: (i) We propose a video-to-tokens module that can extract tokens from extracted spatial-temporal features in videos by 3D convolutions instead of the direct token embedding from raw input video clips; (ii) We completely replace the linear mapping in multi-head self-attention layer with depth-wise convolutional mapping, which applies a depth-wise separable convolution operation on embedded token maps. With these two modifications, our approach can extract effective spatialtemporal features from videos and process the long sequences of tokens encountered in videos. Experimental results demonstrate that our proposed CTN can achieve state-of-the-art accuracy on two fine-grained action recognition datasets (i.e., Epic-Kitchens and Diving 48) with a small computational increase.
Keyword:
Fine-grained action recognition
Transformer
3D convolutions
Spatial-temporal features

期刊

Neurocomputing 封面图
Neurocomputing
IF:
6.5
论文数:
2.5W
被引数:
6.5W

机构

L
liaoning university
学者数:
5.7K
论文数: 3.5K
被引数: 2
S
shanghai institute of technology
学者数:
5.8K
论文数: 3.7K
被引数: 1
N
nanjing university
学者数:
7.8W
论文数: 5.6W
被引数: 87
D
Dalian University of Technology
学者数:
6.0W
论文数: 4.4W
被引数: 5.5W
M
Massey University
学者数:
7.7K
论文数: 7.9K
被引数: 9.6K
学者 查看更多机构
引用论文

引用论文

err分享
err收藏
Contrastive predictive coding with transformer for video representation learning
err2022-04-01
err20
PREAI
errLiu, Yue; Ma, Junqi; Xie, Yufei; Yang, Xuefeng; Tao, Xingzhen; Peng, Lin; Gao, Wei
err分享
err收藏
The mental health of staff working in intensive care during COVID-19
err
IF0
err2020-11-04
err0
errOAAI
errNeil Greenberg; Dale Weston; Charlotte Hall; Tristan Caulfield; Victoria Williamson; Kevin Fong
err分享
err收藏
Spatial-temporal pooling for action recognition in videos
err2021-09-01
err29
PREAI
errWang, Jiaming; Shao, Zhenfeng; Huang, Xiao; Lu, Tao; Zhang, Ruiqian; Lv, Xianwei
err分享
err收藏
The effect of autonomic arousal on attentional focus
err2000-12-01
err0
PREAI
errJ I. Tracy; F Mohamed; S Faro; R Tiver; A Pinus; C Bloomer; A Pyrros; J Harvan
err分享
err收藏
学者 查看更多内容