arrow
返回

Semi-supervised classification-aware cross-modal deep adversarial data augmentation

delete2021-12-01
delete8
PRE
AI
S
Shaoqiang Wang
Z
Zhenzhen Wu
G
Gewen He
王淑栋 封面图
王淑栋 (Shudong Wang)
H
Hongwei Sun
F
Fangfang Fan *
DOI:10.1016/j.future.2021.05.029delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Deep neural networks are usually data-starved in real-world applications, while manually annotation can be costly-for example, the audio emotion recognition from the audio. In contrast, the continued research in image-based facial expression recognition grants us a rich source of public available labeled IFER datasets. Using images to support audio emotion recognition with limited labeled data according to their inherent correlations can be a meaningful and challenging task. This paper proposes a system that facilitates knowledge transfer from the labeled visual to the heterogeneous labeled audio domain by learning a joint distribution of examples in different modalities then the system can map an IFER example to a corresponding audio spectrogram. Next, our work reformulates the audio emotion classification into a K+1 class discriminator of GAN-based semi-supervised learning. Good semi-supervised learning requires that the generator does NOT sample from a distribution well matching the true data distribution. Therefore, we demand the generated examples are from the low-density areas of the marginal distribution in the audio spectrogram modality. Concretely, the proposed model translates image samples to audios class-wisely in the form of spectrograms. To harness the decoded samples in a sparsely distributed area and construct a tighter decision boundary, we give a solution to precisely estimate the density on feature space and incorporate low-density pieces with an annealing scheme. Our method requires the network to discriminate against the low-density data points from high-density data points throughout the classification, and we evidence that this technique effectively improves task performance. (C) 2021 Published by Elsevier B.V.
Keyword:
Adversarial network
Data augmentation
Density estimation
Graph representation
Semi supervised learning

期刊

F
Future Generation Computer Systems-The International Journal of eScience
IF:
6.1
论文数:
6.9K
被引数:
2.3W

机构

State University System of Florida 封面图
State University System of Florida
学者数:
12.8W
论文数: 10.9W
被引数: 130
F
Florida State University
学者数:
1.1W
论文数: 8.6K
被引数: 2.0W
W
weifang university of science & technology
学者数:
660
论文数: 592
被引数: 0
C
china university of petroleum
学者数:
4.1W
论文数: 2.7W
被引数: 30
学者 查看更多机构
引用论文

引用论文

err分享
err收藏
err分享
err收藏
The mitochondrial complex I inhibitor rotenone triggers a cerebral tauopathy
err2005-10-10
err0
PREAI
errGünter U. Höglinger; Annie Lannuzel; Myriam Escobar Khondiker; Patrick P. Michel; Charles Duyckaerts; Jean Féger; Pierre Champy; Annick Prigent; Fadia Medja; Anne Lombes; Wolfgang H. Oertel; Merle Ruberg; Etienne C. Hirsch
err分享
err收藏
Modulation of Progressive Leaf Senescence by the Red:Far-Red Ratio of Incident Light
err1989-06-01
err0
PREAI
errJuan J. Guiamet; Jorge G. Willemoes; Edgardo R. Montaldi
err分享
err收藏
A Mouse Model of Uterine Leiomyosarcoma
err2004-01-01
err0
errOAAI
errKaterina Politi; Matthias Szabolcs; Peter Fisher; Ana Kljuic; Thomas Ludwig; Argiris Efstratiadis
err分享
err收藏
Low-Level Laser Therapy at 635 nm for Treatment of Chronic Plantar Fasciitis: A Placebo-Controlled, Randomized Study
err2015-09-01
err0
PREAI
errDavid M. Macias; Michael J. Coughlin; Kerry Zang; Faustin R. Stevens; James R. Jastifer; Jesse F. Doty
err分享
err收藏
err分享
err收藏
Facial expression recognition from near-infrared videos基于近红外视频的人脸表情识别
err2011-08-01
err520
PREAI
errZhao, Guoying; Huang, Xiaohua; Taini, Matti; Li, Stan Z.; Pietikainen, Matti
err分享
err收藏
学者 查看更多内容