arrow
返回

Multi-Stream Single Network: Efficient Compressed Video Action Recognition With a Single Multi-Input Multi-Output Network

delete2024-01-01
delete0
delete
OA
AI
H
Hayato Terao *
W
Wataru Noguchi
H
Hiroyuki Iizuka
M
Masahito Yamamoto
DOI:10.1109/ACCESS.2024.3363022delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Compressed video action recognition classifies actions using multiple features stored in compressed videos to omit the decoding process for RGB frames and shorten the computation time. Previous methods mostly used multiple networks to process compressed video features and explored the use of lightweight networks without affecting accuracy to reduce the computational complexity further. We have focused on another approach that uses only one network to reduce computational complexity. Our previous study proposed the MussNet model, which consists of independent subnetworks within a single network instead of multiple networks. The subnetworks classify compressed video features independently with a feedforwarding step of a single network and achieved competitive accuracy against previous studies with lower computational complexity. The remaining issue of the MussNet model is how to fuse the independently processed compressed video features. The current MussNet model makes independent predictions from each input and only averages them to fuse the inputs. However, recent studies have shown that intermediate fusion, which fuses features inside the networks, improves accuracy. This study proposes the EFS module that extends the MussNet model into intermediate fusion by disentangling and aggregating the features of the same videos in the hidden vectors while keeping the individual subnetworks. Our experiments show that the EFS module improves the MussNet model's accuracy by 0.4 points for UCF-101 and 1.0 points for HMDB-51, while the additional GFLOPs are only 1% of the MussNet model. These accuracy scores are also competitive against previous studies while keeping one of the lowest computational complexity.
Keyword:
Compressed video action recognition
deep learning, intermediate fusion
multi-input multi-output model
video understanding

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

H
Hokkaido University
学者数:
3.6W
论文数: 2.5W
被引数: 2.6W
引用论文

引用论文

err分享
err收藏
The New Digital Storytelling
err
IF0
err2017-01-01
err0
PREAI
errBryan Alexander
err分享
err收藏
Dense Trajectories and Motion Boundary Descriptors for Action Recognition
err2013-03-06
err1.3K
errOAAI
errWang, Heng; Klaeser, Alexander; Schmid, Cordelia; Liu, Cheng-Lin
err分享
err收藏
The mental health of staff working in intensive care during COVID-19
err
IF0
err2020-11-04
err0
errOAAI
errNeil Greenberg; Dale Weston; Charlotte Hall; Tristan Caulfield; Victoria Williamson; Kevin Fong
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
ImageNet Large Scale Visual Recognition ChallengeImageNet大规模视觉识别挑战
err2015-04-11
err2.7W
PREAI
errRussakovsky, Olga; Deng, Jia; Su, Hao; Krause, Jonathan; Satheesh, Sanjeev; Ma, Sean; Huang, Zhiheng; Karpathy, Andrej; Khosla, Aditya; Bernstein, Michael; Berg, Alexander C.; Fei-Fei, Li
err分享
err收藏
err分享
err收藏
学者 查看更多内容