arrow
返回

Temporal refinement network: Combining dynamic convolution and multi-scale information for fine-grained action recognition

delete2024-07-01
delete2
PRE
AI
H
Hu, Zhengping *
H
Hehao Zhang
Y
Yulu Wang
Z
Zhe Sun
DOI:10.1016/j.imavis.2024.105058delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Fine-grained action recognition is challenging due to the nearly identical context, limited background information, and less distinct inter-class differences compared to coarse-grained actions.Effectively capturing spatiotemporal information is crucial for fine-grained action recognition models. To address the limitations of coarsegrained models in describing spatio-temporal context, we propose a Temporal Refinement Block (TRB) as an efficient component for fine-grained action recognition. The TRB enables our model to effectively model underlying semantics and global dependencies by generating spatial-temporal kernels of different scales and performing fully connected operations within the temporal dimension. Our experiments demonstrate the effectiveness of TRB in learning latent semantics and global dependencies. To further enhance the framework's performance, we incorporate an enhanced spatio-temporal pyramidal network (TPN) that collects beat information and utilizes dilated convolutions to boost multi-scale features.We refer to the proposed framework as the Temporal Refinement Network, abbreviated as TRN.Our TRN achieves competitive performance on the FineGym and Diving48 benchmarks.
Keyword:
Fine-grained action recognition
Temporal refinement block
Temporal pyramidal network

期刊

Image and Vision Computing 封面图
Image and Vision Computing
IF:
4.2
论文数:
4.1K
被引数:
6.7K

机构

Y
Yanshan University
学者数:
1.7W
论文数: 1.1W
被引数: 1.3W
引用论文

引用论文

GCM: Efficient video recognition with glance and combine module
err2023-01-01
err7
PREAI
errZhou, Yichen; Huang, Ziyuan; Yang, Xulei; Ang, Marcelo; Ng, Teck Khim
err分享
err收藏