arrow
Return

EAR: Efficient action recognition with local-global temporal aggregation

delete2021-12-01
delete2
PRE
AI
C
Can Zhang
Y
Yuexian Zou *
G
Guang Chen
L
Lei Gan
DOI:10.1016/j.imavis.2021.104329delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Temporal modeling in videos is crucial for action recognition. Traditionally, it involves feature aggregation for both local motion and global semantic. In this paper, we propose an Efficient Action Recognition network (EAR), which includes a Persistence of Appearance (PA) module anda Various-timescale Aggregation (VA) module for local and global temporal aggregations respectively. For local motion aggregation, instead of using the previ-ous time-consuming optical flow, our PA calculates pixel-wise differences in feature space as the motion repre-sentation, which is much more efficient (8196 fps vs. 8 fps in optical flow). Besides, to capture global semantic hints, we propose VA module which adaptively emphasizes expressive features and suppresses less informative ones across various timescales. Empowered by the local-global temporal aggregation, our EAR achieves compet-itive results on six challenging action recognition benchmarks at low FLOPs. (c) 2021 Elsevier B.V. All rights reserved.
Keywords:
Efficient action recognition
Local-global temporal aggregation
Motion representation
Persistence of appearance

Journal

Image and Vision Computing cover
Image and Vision Computing
IF:
4.2
Papers:
4.0K
Citations:
6.7K

Organization

P
peking university
Scholars:
11.8W
Papers: 8.7W
Citations: 146