arrow
Return

Recurrent attention network using spatial-temporal relations for action recognition

delete2018-04-01
delete27
PRE
AI
张明星 (Mingxing Zhang)
杨阳 (Yang Yang) *
Y
Yanli Ji
N
Ning Xie
F
Fumin Shen
DOI:10.1016/j.sigpro.2017.12.008delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Action recognition in videos, which contains many complex and semantic contents, is still a challenging task in computer vision research. In this paper, we propose a novel attention mechanism that leverages the gate system of Long Short Term Memory (LSTM) to compute the attention weights for action recognition. The proposed attention mechanism is embedded in a recurrent attention network that can explore the spatial-temporal relations between different local regions to concentrate important ones. For more accurate attention, we derive a new attention unit from the standard LSTM unit so as how important the local region is only depends on its input gate. Because of exploring spatial-temporal relations and using attention unit, our model can attend more accurately and thus achieve a better action recognition performance. We evaluate our proposed model on three datasets: UCF101, HMDB51 and Hollywood2, and results illustrate that our model outperforms other attention models with significant improvements. (C) 2017 Elsevier B.V. All rights reserved.
Keywords:
Action recognition
Attention mechanism
Spatial-temporal relations
LSTM
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Signal Processing cover
Signal Processing
IF:
3.6
Papers:
9.9K
Citations:
1.7W

Organization

No organization information available