Return
A dual-modality deep learning approach for fine-grained sentiment analysis using speech and text
DOI:10.1016/j.knosys.2025.115141.png)
Abstract
En 中文
Understanding human emotions from digital data has become a critical component of modern intelligent systems. Traditional unimodal approaches, which rely solely on text or audio, often fail to capture the subtle interplay between semantics and prosody, thereby limiting their ability to detect nuanced sentiments, such as sarcasm or mixed emotions. To address this challenge, we present a dual-modality deep learning framework that integrates textual embeddings from RoBERTa with acoustic representations from Wav2Vec 2.0. The novelty of the research lies in a multi-head cross-modal attention mechanism that dynamically aligns semantic and prosodic cues, together with an interpretable output layer that translates complex emotions into descriptive tokens or emojis for enhanced usability. Extensive evaluations on the CMU-MOSI and CMU-MOSEI benchmarks confirm the effectiveness of the proposed approach, yielding accuracies of 84.99% and 84.33% with F1-scores of 87.01% and 89.43 %, respectively. Compared with the strongest baselines, the framework achieves a +1.9 % gain in F1-score and +0.7% accuracy on CMU-MOSI, along with a +0.33% F1 improvement on CMU-MOSEI, demonstrating both consistency and robustness under class imbalance. These results highlight the superiority of the proposed framework and its potential to advance next-generation emotion-aware systems in practical domains.
Keywords:
Sentiment analysis
Feature fusion
Attention mechanism
Sentiment classification
RoBERTa
Wave2Vec 2.0
Journal
K
IF:
7.6
Papers:
1.2W
Citations:
4.5W

