返回
Deep operational audio-visual emotion recognition
DOI:10.1016/j.neucom.2024.127713.png)
摘要
En 中文
Emotions play a large role in interpersonal communication, marketing, healthcare and the service industry. For this reason, much research has been carried out on emotion classification until today. Audio-visual emotion recognition is a field within artificial intelligence and machine learning that focuses on recognizing and understanding human emotions from both visual and audio cues. It combines computer vision and audio processing techniques to analyze and interpret emotional states expressed by individuals. This paper presents a deep learning model developed over an operational neural network using multiple inputs and aimed at audio-visual emotion recognition. The proposed network utilizes both visual and audio information in an end to end approach. The primary objective of this work is to demonstrate that multi-input models can produce more efficient outcomes compared to single-input models in emotion classification. Another objective is to demonstrate the superior performance of weight calculation methods employed in operational neural networks compared to the conventional weight calculation methods used in convolutional neural networks. Therefore, we want to demonstrate that substituting convolutional neural network approaches with operational neural network methods can yield superior outcomes in emotion categorization models. In the proposed architecture, regular convolutional layers are replaced with operational layers. The experimental results demonstrate that the operational convolutional architecture performs better compared to the classical convolutional neural network architecture.
Keyword:
Audio-visual emotion classification
Operational neural network
Visual geometry group
Multi-input classification
期刊
IF:
6.5
论文数:
2.5W
被引数:
6.5W
机构
引用论文
Thermally Activated Delayed Fluorescence-Based Near-Infrared-II Luminescence and Controlled Size Growth of Silver Nanoclusters热激活延迟荧光基近红外-II区发光及银纳米簇的尺寸可控生长
ACS NANO
IF16
Atmospheric Oxidation Mechanism of 2-Hydroxy-benzothiazole Initiated by Hydroxyl Radicals羟基自由基引发的2-Hydroxy-benzothiazole的大气氧化机制
Speech Emotion Recognition Using Convolution Neural Networks and Multi-Head Convolutional Transformer基于卷积神经网络和多头卷积变换器的语音情感识别
SENSORS
IF3.5
Multimodal Emotion Recognition With Transformer-Based Self Supervised Feature Fusion基于Transformer自监督特征融合的多模态情感识别
IEEE ACCESS
IF3.6

