返回
Weakly Semantic Guided Action Recognition
DOI:10.1109/TMM.2019.2907060.png)
摘要
En 中文
Action recognition plays a fundamental role in computer vision and video analysis. Nevertheless, extracting effective spatial-temporal features remains a challenging task. This paper proposes three simple but effective weakly semantic guided modules (SGMs) for both environment-constrained and cross-domain action recognition. The SGMs are composed of total 3-D convolution and element-wise gated operations; thus, they are efficient and easy to implement. The semantic guidance is obtained in a weakly supervised manner, in which each video clip is labeled with only an action class instead of pixel-level semantics. Benefitting from the semantic guidance, the network [called semantic guided network (SGN)] can focus on the salient parts of the video clips. Consequently, the redundant information can be reduced and the model is more robust to noise. Besides, benefitting from the intrinsic property of SGMs, SGN is totally end-to-end trainable. Quantities of experiments on both environment-constrained (e.g., Penn, HMDB-51, and UCF-101) and cross-domain (e.g., ODAR) action recognition datasets demonstrate its effectiveness. Specifically, SGN gets improvements of 3.7%, 2.1%, and 5.2% for Penn, HMDB-51, and UCF-101 than the baseline ResNet3D, respectively, and SGN ranked third place in the ODAR 2017 challenge.
Keyword:
Semantic guided module
action recognition
cross domain
3D convolution
attention model
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
9.7
论文数:
4.5K
被引数:
2.4W

