arrow
返回

Frequency Decoupled Masked Auto-Encoder for Self-Supervised Skeleton-Based Action Recognition

delete2025-01-01
delete0
PRE
AI
Y
Ye Liu *
M
Mingliang Zhai
J
Jun Liu
DOI:10.1109/LSP.2024.3525398delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
In 3D skeleton-based action recognition, the limited availability of supervised data has driven interest in self-supervised learning methods. The reconstruction paradigm using masked auto-encoder (MAE) is an effective and mainstream self-supervised learning approach. However, recent studies indicate that MAE models tend to focus on features within a certain frequency range, which may result in the loss of important information. To address this issue, we propose a frequency decoupled MAE. Specifically, by incorporating a scale-specific frequency feature reconstruction module, we delve into leveraging frequency information as a direct and explicit target for reconstruction, which augments the MAE's capability to discern and accurately reproduce diverse frequency attributes within the data. Moreover, in order to address the issue of unstable gradient updates caused by more complex optimization objectives with frequency reconstruction, we introduce a dual-path network combined with an exponential moving average (EMA) parameter updating strategy to guide the model in stabilizing the training process. We have conducted extensive experiments which have demonstrated the effectiveness of the proposed method.
Keyword:
Training
Skeleton
Image reconstruction
Convolution
Data models
Solid modeling
Frequency-domain analysis
Frequency diversity
Transformers
Target recognition
Skeleton-based action recognition
masked auto-encoder
self-supervised learning
frequency domain

期刊

IEEE Signal Processing Magazine 封面图
IEEE Signal Processing Magazine
IF:
9.6
论文数:
1.1W
被引数:
1.7W

机构

L
Lancaster University
学者数:
9.5K
论文数: 1.1W
被引数: 1.7W