Return
Masked signal modeling with self-supervised transformers for stress and affect detection using wearable biosignals
E
S
DOI:10.1016/j.bspc.2026.109607.png)
Abstract
En 中文
Accurate recognition of stress and affective states from wearable physiological signals remains challenging due to limited labeled data, inter-subject variability, and noise inherent in real-world sensing. This study evaluates the performance of a self-supervised MSM-Transformer for stress and affect detection using multimodal physiological data from the WESAD and SWELL-KW datasets. Under a strict Leave-One-Subject-Out Cross-Validation (LOSO-CV) protocol, the proposed model achieves 89.0 % accuracy on the WESAD dataset, outperforming CNN-LSTM, LSTM and a supervised Transformer. On the SWELL-KW dataset, the MSM-Transformer attains 86.9 % accuracy, exceeding CNN-LSTM, LSTM and supervised Transformer. Ablation experiments indicate that MSM pretraining improves accuracy by + 3.0–3.2 % and F1-score by + 3.1 % compared to models trained from scratch. Robustness evaluations show strong stability under severe noise, maintaining 87.4 % accuracy at 5 dB SNR with only a 1.6 % decrease, and under temporal distortions, retaining 83.8 % accuracy with ± 20 % warping. Cross-dataset generalization further supports model scalability, achieving 84.1 % accuracy when pretrained on WESAD and fine-tuned on SWELL-KW. Interpretability assessments demonstrate meaningful physiological alignment, with attention weights strongly correlating with key stress markers, including EDA peaks (r = 0.74), respiration fluctuations (r = 0.68), and BVP amplitude changes (r = 0.61). These results confirm the MSM-Transformer’s superior performance, robustness and physiological relevance in real-time wearable stress monitoring.
Journal
IF:
4.9
Papers:
9.7K
Citations:
2.4W

