Return
Multimodal Spatiotemporal Representation for Automatic Depression Level Detection
DOI:10.1109/TAFFC.2020.3031345.png)
Abstract
En 中文
Physiological studies have shown that there are some differences in speech and facial activities between depressive and healthy individuals. Based on this fact, we propose a novel spatio-temporal attention (STA) network and a multimodal attention feature fusion (MAFF) strategy to obtain the multimodal representation of depression cues for predicting the individual depression level. Specifically, we first divide the speech amplitude spectrum/video into fixed-length segments and input these segments into the STA network, which not only integrates the spatial and temporal information through attention mechanism, but also emphasizes the audio/video frames related to depression detection. The audio/video segment-level feature is obtained from the output of the last full connection layer of the STA network. Second, this article employs the eigen evolution pooling method to summarize the changes of each dimension of the audio/video segment-level features to aggregate them into the audio/video level feature. Third, the multimodal representation with modal complementary information is generated using the MAFF and inputs into the support vector regression predictor for estimating depression severity. Experimental results on the AVEC2013 and AVEC2014 depression databases illustrate the effectiveness of our method.
Keywords:
Feature extraction
Depression
Two dimensional displays
Spatiotemporal phenomena
Databases
Three-dimensional displays
Image segmentation
Multimodal depression detection
spatio-temporal attention
audio
video segment-level feature
eigen evolution pooling
video level feature
multimodal attention feature fusion
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
9.8
Papers:
1.3K
Citations:
9.1K

