Return
TFC: Time-frequency contrasting network for wearable-based human activity recognition
DOI:10.1016/j.knosys.2025.113373.png)
Abstract
En 中文
Human Activity Recognition (HAR) using sensor data has significantly progressed with various supervised learning architectures, including both traditional CNN/LSTM models and the more recent Transformer-based models. A primary challenge in supervised learning is the requirement for extensive, accurately labeled training data. Self-supervised methods, particularly those employing contrastive learning, offer an innovative solution to this challenge by leveraging unlabeled data. In this study, we introduce a novel self-supervised learning method named Time-Frequency Contrasting (TFC) for limited labeled data in HAR, rooted in the principles of contrastive learning and Bayesian structural time series. This approach interprets sensor data as a combination of time-domain trends, frequency-domain cycles, and error noise. Our objective is to learn a universal representation so as to enhance the performance of human activity recognition in downstream tasks. This is achieved by minimizing the impact of redundant noise and leveraging time-domain prior knowledge to learn time-domain trend features and utilizing frequency-domain prior knowledge to acquire frequency-domain cycle features, respectively. After fine-tuning, TFC achieved Macro F1-scores of 86.39, 95.44, and 80.27 on three publicly available datasets, namely MotionSense, USC-HAD, and UCI-HAR. Additionally, it obtained a F1-score of 97.64 on a custom-built dataset called BARD. Our extensive experiment demonstrate that TFC markedly improves self-supervised activity recognition tasks, especially in scenarios with limited labeled data.
Keywords:
Human activity recognition
Representation learning
Self-supervised learning
Contrastive learning
Deep learning

