Return
A micro-expression recognition algorithm fusing visual information with textual semantics
DOI:10.1016/j.eswa.2025.129000.png)
Abstract
En 中文
In micro-expression research, the dependencies among Action Units (AUs) and their associated facial regions have provided valuable cues for feature extraction and have driven substantial progress. However, existing methods primarily rely on visual data and AU labels, which often fail to capture the nuances of facial movements. In contrast, semantic descriptions of AUs offer richer and more interpretable information about facial muscle behavior, yet remain underexplored in micro-expression analysis. In this study, we propose a novel framework that integrates visual information with AU textual semantics to enhance micro-expression analysis. Specifically, we construct detailed AU descriptions and introduce a semantic feature extraction module to capture inter-AU dependencies. To overcome the limitations of long-range inductive bias in local features within CNN models, we develop a region attention mechanism in the visual feature extraction module. Furthermore, a cross-modal attention mechanism is introduced to semantically align and fuse visual features with AU textual semantics at both global and local levels. Our model achieves strong performance, with accuracies of 79.27%, 82.45%, and 80.14% on the SMIC, CASMEII, and SAMM datasets, respectively. On a composite dataset, our method further achieves a UF1 of 0.8461 and a UAR of 0.8510, outperforming or matching state-of-the-art methods. Ablation studies and visualizations further validate the effectiveness and interpretability of our proposed approach.
Keywords:
micro-expression
Action Units
semantic description
cross-modal attention
region attention
Journal
IF:
7.5
Papers:
2.9W
Citations:
10.2W

