1
Return

Automatic analysis of speech representations to assess psychological distress

delete2026-08-13
delete0
delete
OA
AI
S
SF Sara Fernández-Velasco
J
JM Jose Moreno-Mesa
D
DE Daniel Escobar-Grisales
D
DM Diego M. Lopez
J
JR Juan Rafael Orozco-Arroyave
DOI:10.3389/fdgth.2026.1886893delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
BackgroundCurrent mental health diagnostic methods are limited by subjective clinical interpretation. Automatic speech analysis is a promising technology for objective assessment.ObjectiveTo evaluate and compare different speech-based representations (acoustic; phonetic; and time-frequency) and deep learning-based embeddings for discriminating symptoms associated with psychological distress.MethodsA secondary analysis of the Distress Analysis Interview Corpus (DAIC-WOZ) was conducted using recordings from 125 participants (3; 069 responses). Speech representations included phonation; articulation; and prosody features extracted with DisVoice; phonetic features extracted with Phonet; time-frequency representations derived from Mexican hat wavelets; and deep embeddings extracted with the multilingual Wav2Vec 2.0 model XLSR-53. Two classification strategies were addressed at the response and participant levels using a Fully Connected Neural Network (FCNN) and a Support Vector Machine (SVM); respectively.ResultsProsody at the participant level achieved the highest mean performance (F1-score 0.67 ± 0.07; accuracy 0.64 ± 0.10; AUC 0.65 ± 0.12); followed by participant-level phonation (F1-score 0.59 ± 0.16; accuracy 0.61 ± 0.14; AUC 0.65 ± 0.16). Conversely; participant-level aggregation of deep embeddings yielded lower performance (F1-score 0.48 ± 0.19; accuracy 0.55 ± 0.13; AUC 0.52 ± 0.15); failing to surpass traditional features. Response-level performance remained close to chance. Phonet and wavelet representations did not improve performance over prosody or phonation.ConclusionParticipant-level analysis provided more robust and consistent discriminative patterns than response-level approaches. Prosody and phonation achieved the best performance across speech representations; while phonetic; time-frequency; and deep speech representations did not outperform the best acoustic baseline. These findings suggest that; within the evaluated experimental setting; the aggregation strategy appears to have a stronger influence on performance than increasing representational complexity.
Keywords:
digital mental health
depression
speech analysis
psychological distress
post-traumatic stress disorder
vocal biomarkers
DAIC-WOZ
participant-level validation

Journal

F
Frontiers in Digital Health
IF:
3.8
Papers:
2.0K
Citations:
3.1K

Organization

T
telematics department
Scholars:
3
Papers: 1
Citations: 0
G
gita lab
Scholars:
4
Papers: 1
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers