Return
Learning discriminative features for micro-expression recognition
DOI:10.1007/s11042-023-15596-3.png)
Abstract
En 中文
Micro-expressions (MEs) have subtle facial muscle movements, and fully extracting the movement features is crucial to recognizing MEs. However, the existing deep models are prone to focus on those regions with relatively prominent facial muscle movements and ignore the ones with less obvious muscle movements, thereby not aggregating the facial movement information comprehensively. To solve this problem, we propose a novel SpatioTemporal Two-Stream Network based on a Class Activation Map (STTSN-CAM) to emphasize the facial regions related to MEs. STTSN-CAM adopts a two-stream structure including the spatial and temporal streams. The spatial stream generates a CAM that represents the regions with high contributions to ME recognition. The temporal stream introduces CAM into the inputs to guide the model to focus on the facial key regions, thereby learning more discriminative features. In addition, a novel data augmentation method is proposed to generate new ME samples with two labels by combining two samples with different ME categories. Furthermore, to accommodate new samples with two labels, a dual-labeling loss function is designed to drive the model to focus on different facial local regions, further strengthening the ability to capture discriminative movement information comprehensively. The experimental results show that the proposed method outperforms most advanced methods, and on CASME II and SMIC datasets, the recognition accuracy of the proposed method is 78.63% and 71.95%, respectively, which is improved by 5.91 and 4.88 percentage points compared with the baseline method.
Keywords:
Micro-Expression recognition
Data augmentation
Dual-Labeling
Dual-Stream networks
Journal
IF:
3
Papers:
1.9W
Citations:
3.2W
Organization
No organization information available

