1
Return

Action Recognition Method Based on Multi-Scale Dilated Feature Fusion and Decoupled Spatiotemporal Attention Pooling

delete2026-08-13
delete0
delete
OA
AI
H
Hanbo Zhang
J
Jing Huang *
DOI:10.3390/electronics15163581delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Video action recognition requires the joint modeling of spatial appearance information and temporal dynamics. However, existing efficient action recognition methods based on two-dimensional convolution still have limitations in representing multi-scale spatial cues and aggregating key spatiotemporal information. To address these issues, this paper proposes an action recognition network based on Multi-Scale Dilated Feature Fusion and Decoupled Spatiotemporal Attention Pooling, termed MDSTA-Net. The proposed method adopts TSM as the basic temporal modeling framework and ResNet-50 as the backbone network. First, a Multi-Scale Dilated Feature Fusion module (MSDF) is designed to construct continuous multi-scale receptive fields through parallel convolutional branches with different dilation rates. An adaptive branch aggregation mechanism is further introduced to dynamically fuse responses at different scales, thereby enhancing the representation of both local details and broader contextual information. Second, a Decoupled Spatiotemporal Attention Pooling module (DSTAP) is proposed to model key action frames along the temporal dimension and salient discriminative regions along the spatial dimension. A residual pooling path is also incorporated to preserve global semantic information, improving the discriminative capability of video-level action representations. Experimental results on three public datasets, namely Something-Something V2, Kinetics-400, and HMDB51, demonstrate that MDSTA-Net achieves favorable recognition performance compared with several representative methods. Ablation studies further verify the effectiveness of MSDF and DSTAP, indicating that multi-scale spatial feature enhancement and key spatiotemporal information aggregation can effectively improve action recognition performance.
Keywords:
action recognition
multi-scale feature fusion
spatiotemporal attention
residual pooling

Journal

Electronics cover
Electronics
IF:
2.6
Papers:
9.2K
Citations:
4.7W

Organization

Z
Zhejiang Sci-Tech University
Scholars:
1.6W
Papers: 9.9K
Citations: 1.3W
Cited Papers

Cited Papers

Citing Papers

Citing Papers