arrow
Return

Differential motion attention network for efficient action recognition

delete2024-06-13
delete0
PRE
AI
C
Caifeng Liu
F
Fangjie Gu *
DOI:10.1007/s00371-024-03478-0delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Despite the great progresses achieved by commonly-used 3D CNNs and two-stream methods in action recognition, they cause heavy computational burden which are inefficient and even infeasible in real-world scenarios. In this paper, we propose differential motion attention network (DMANet) to specially highlight human dynamics toward efficient action recognition. First, we argue that consecutive frames contain redundant static features and construct a low computational unit for discriminative motion extraction to highlight the human action trajectories across consecutive frames. Second, as not all spatial regions in images play an equal role in depicting human actions, we propose an adaptive protocol to dynamically emphasize informative spatial regions. As an end-to-end lightweight framework, our DMANet outperforms costly 3D CNNs and two-stream methods by 2.3% with only 0.23x\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\times $$\end{document} computations and other efficient methods by 1.6% on Something-Something v1 dataset. Experimental results on two temporal-related datasets and the large-scale scene-related Kinetics-400 dataset prove the efficacy of DMANet. In-depth ablations further give both quantitative and qualitative support on its effects.
Keywords:
Action recognition
Temporal reasoning
Differential motion attention
Efficiency

Journal

Visual Computer cover
Visual Computer
IF:
2.9
Papers:
4.6K
Citations:
6.5K

Organization

D
Dalian University of Technology
Scholars:
5.9W
Papers: 4.4W
Citations: 5.5W