Return
MLENet: Multi-Level Extraction Network for video action recognition
DOI:10.1016/j.patcog.2024.110614.png)
Abstract
En 中文
Human action recognition is a well -established task in the field of computer vision. However, accurately representing spatio-temporal information remains a challenge due to the complex interplay between human actions, video timing, and scene changes. To address this challenge and improve the efficiency of temporal modeling in videos, we propose MLENet, a novel approach that eliminates contextual data and eliminates the need for laborious optical flow extraction.MLENet incorporates a Temporal Feature Refinement Extraction Module (TFREM) that utilizes Optical Flow Guided Features to enhance attention to local deep detail information. This refinement process significantly enhances the network's capacity for feature learning and expression. Moreover, MLENet is designed to be trained end -to -end, facilitating seamless integration into existing frameworks. Additionally, our model adopts a temporal segmentation structure for sampling, effectively reducing redundant information and improving computational efficiency. Compared to existing video -based action recognition models that require optical flow or other modalities, MLENet achieves substantial performance enhancements while requiring fewer inputs. We validate the effectiveness of our proposed approach on benchmark datasets, including Something -Something V1&V2, UCF-101, and HMDB-51, where MLENet consistently outperforms state-of-the-art models.
Keywords:
Action recognition
Spatio-temporal
Temporal feature refinement extraction module
Motion information
Optical flow guided feature
Journal
IF:
7.6
Papers:
1.3W
Citations:
4.5W

