1
Return

Multimodal text-audio sentiment in clinical aphasia speech using NLP

delete2026-08-11
delete0
delete
OA
AI
S
SB Shamiha Binta Manir
A
AN Anai N. Kothari
W
WL William L. Gross
P
PD Priya Deshpande
DOI:10.3389/fdgth.2026.1740270delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
IntroductionAphasia affects expressive and receptive communication and may influence the affective tone expressed during clinical speech tasks. This study presents an exploratory weakly supervised NLP analysis of positive/negative affective-tone proxies in AphasiaBank transcripts with paired audio.MethodsWe extracted sentence embeddings from DistilBERT (e∈R768) and recording-level acoustic summaries (MFCC13; ZCR; RMS; spectral centroid; and spectral bandwidth; a∈R17). Text and acoustic features were concatenated (x=[e;a]∈R785) and classified using Random Forest models. Sentiment labels were generated using an SST-2-derived weak-supervision pipeline and should be interpreted as pseudo-labels rather than clinical ground truth. To evaluate modality contribution and potential leakage; we compared text-only; audio-only; and fused text-audio models under utterance-level and recording-disjoint splits. A small five-rater evaluation was used to examine human judgment alignment.ResultsUnder the recording-disjoint split; both the text-only and fused text-audio models achieved 97.9% accuracy and 0.791 macro-F1; while the audio-only model achieved 55.4% accuracy and 0.388 macro-F1. These results indicate that classification performance was primarily driven by textual embeddings; while the recording-level acoustic summaries did not improve performance over text-only features. Across the pseudo-labeled corpus; aphasic utterances were more often labeled negative than control utterances. Age-stratified summaries showed subgroup variation in pseudo-label distributions; but these patterns were treated descriptively because labels were model-derived. Human-rater agreement was low for aphasic utterances; indicating that affective-tone interpretation in fragmented clinical speech is ambiguous.DiscussionThese findings should be interpreted as exploratory evidence about weakly supervised affective-tone proxies; not as validated clinical sentiment recognition. The results highlight both the promise of clinical NLP for aphasia discourse analysis and the need for independent human-labeled validation; utterance-aligned acoustic features; and careful control of domain; task; age; and topic bias.
Keywords:
sentiment analysis
acoustic features
aphasia
clinical NLP
weak supervision
affective tone
DistilBERT embeddings
multimodal speech analysis

Journal

F
Frontiers in Digital Health
IF:
3.8
Papers:
2.0K
Citations:
3.1K

Organization

D
Department of Electrical and Computer Engineering
Scholars:
744
Papers: 399
Citations: 6
D
department of surgery
Scholars:
1.8K
Papers: 522
Citations: 0
D
department of anesthesiology
Scholars:
1.2K
Papers: 411
Citations: 5
Cited Papers

Cited Papers

Citing Papers

Citing Papers