arrow
返回

Spatial and temporal saliency based four-stream network with multi-task learning for action recognition

delete2023-01-01
delete19
PRE
AI
M
Ming Zong
R
Ruili Wang
Y
Yujun Ma *
W
Wanting Ji
DOI:10.1016/j.asoc.2022.109884delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Action recognition is a challenging video understanding task for the following two reasons: (i) the complex video background impairs the recognition of desirable actions, and (ii) the fusion of spatial information and temporal information. In this paper, we proposed a novel spatial and temporal saliency based four-stream network with multi-task learning. The proposed model comprises four streams: an appearance stream (i.e. a spatial stream), a motion stream (i.e. a temporal stream), a novel spatial saliency stream and a novel temporal saliency stream. The spatial stream captures the global spatial information from videos using the sampled RGB video frames as the input. The temporal stream captures the global motion information of each pixel using the sampled stacked optical flow frames as the input. The novel spatial saliency stream is used to acquire spatial saliency information from spatial saliency frames, and the novel temporal saliency stream is used to acquire temporal saliency information from temporal saliency frames. In addition, based on the four streams, multi-task learning based LSTM is adopted, which can share the complementary knowledge between different CNN features extracted from different stacked frames. The multi-task learning based LSTM can capture long-term dependency relationships between the consecutive frames over temporal evolution, which take full advantage of CNNs and LSTMs. We conduct experiments on three popular video action recognition datasets, including the UCF101 action dataset, the HMDB51 action dataset and the large-scale Kinetics action dataset, to verify the effectiveness of the proposed network, and the results demonstrate that the proposed network achieves better performance than the state-of-the-art methods on these action recognition datasets.(c) 2022 Elsevier B.V. All rights reserved.
Keyword:
Action recognition
Spatial saliency
Temporal saliency
Multi-task learning

期刊

Applied Soft Computing 封面图
Applied Soft Computing
IF:
6.6
论文数:
1.4W
被引数:
4.8W

机构

L
liaoning university
学者数:
5.7K
论文数: 3.5K
被引数: 2
P
peking university
学者数:
11.9W
论文数: 8.7W
被引数: 146
M
Massey University
学者数:
7.7K
论文数: 7.9K
被引数: 9.6K
学者 查看更多机构
引用论文

引用论文

Three-Dimensional Geometry of the Heineke–Mikulicz Strictureplasty
err2013-03-01
err0
errOAAI
errLuka Pocivavsek; Efi Efrati; Ke Y.C. Lee; Roger D. Hurst
err分享
err收藏
Climate policy: Steps to China's carbon peak气候政策: 迈向中国碳峰值的步骤
err2015-06-17
err0
errOAAI
errZhu Liu; Dabo Guan; Scott Moore; Henry Lee; Jun Su; Qiang Zhang
err分享
err收藏
err分享
err收藏
学者 查看更多内容