1
Return

DPMMN: A dual performer-multi-modal network for emotion recognition

delete2026-01-21
delete0
delete
OA
AI
S
Shivanand S. Gornale
P
Palaiahnakote Shivakumara *
A
Amruta Unki
S
Sunil Vadera
DOI:10.1016/j.image.2025.117464delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Background and Objective: Although emotion recognition systems have been widely advocated, their accuracy can be affected when a person's normal facial features overlap with their expressions when in a particular emotional state. This study, therefore, explores how heatmaps of electroencephalography (EEG) signals can be integrated with facial information to improve the accuracy of emotion recognition systems. Method: The key idea of the proposed work is to fuse EEG signal heatmaps and Facial information for recognizing eight different emotions. For implementing this new idea, we propose a Dual Performer Multi-Modal Network (DPMMN). For each modality, the proposed work integrates modified Vision Transformer (ViT) and Long Short-Term Memory (LSTM). The integration is achieved by concatenating the features extracted from each modality and using them to classify the different emotions. In contrast to a baseline ViT, which uses self-attention layers, the proposed work replaces self-attention layers with the Performer layers through a kernelized attention approach. This results in extracting distinct visual features from EEG signal heatmaps and facial images. Similarly, for capturing temporal features from EEG heatmaps and Facial videos, the proposed LSTM replaces a traditional feed-forward network with a recurrent structure. This step helps to learn sequential dependencies across the patches. Results: A comprehensive evaluation of DPMMN with respect to current state-of-the-art systems shows favorable results, with DPMN achieving 97.02 % in identifying eight distinct emotions on the DEAP benchmark dataset. Conclusion: The proposed work shows that the use of EEG signal heatmap with facial information is better than EEG signal and facial information alone. Similarly, integrating performer layers with ViT and LSTM is better than existing models for extracting distinct features to classify eight emotions.
Keywords:
Brain-computer interface
EEG signals
Vision transformer
Affective computing
Emotions
LSTM
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

S
SIGNAL PROCESSING-IMAGE COMMUNICATION
IF:
2.7
Papers:
18
Citations:
0

Organization

U
University of Salford
Scholars:
2.4K
Papers: 2.5K
Citations: 2.9K
Cited Papers

Cited Papers

Citing Papers

Citing Papers