arrow
Return

Semi-supervised classification-aware cross-modal deep adversarial data augmentation

delete2021-12-01
delete8
PRE
AI
S
Shaoqiang Wang
Z
Zhenzhen Wu
G
Gewen He
王淑栋 cover
王淑栋 (Shudong Wang)
H
Hongwei Sun
F
Fangfang Fan *
DOI:10.1016/j.future.2021.05.029delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Deep neural networks are usually data-starved in real-world applications, while manually annotation can be costly-for example, the audio emotion recognition from the audio. In contrast, the continued research in image-based facial expression recognition grants us a rich source of public available labeled IFER datasets. Using images to support audio emotion recognition with limited labeled data according to their inherent correlations can be a meaningful and challenging task. This paper proposes a system that facilitates knowledge transfer from the labeled visual to the heterogeneous labeled audio domain by learning a joint distribution of examples in different modalities then the system can map an IFER example to a corresponding audio spectrogram. Next, our work reformulates the audio emotion classification into a K+1 class discriminator of GAN-based semi-supervised learning. Good semi-supervised learning requires that the generator does NOT sample from a distribution well matching the true data distribution. Therefore, we demand the generated examples are from the low-density areas of the marginal distribution in the audio spectrogram modality. Concretely, the proposed model translates image samples to audios class-wisely in the form of spectrograms. To harness the decoded samples in a sparsely distributed area and construct a tighter decision boundary, we give a solution to precisely estimate the density on feature space and incorporate low-density pieces with an annealing scheme. Our method requires the network to discriminate against the low-density data points from high-density data points throughout the classification, and we evidence that this technique effectively improves task performance. (C) 2021 Published by Elsevier B.V.
Keywords:
Adversarial network
Data augmentation
Density estimation
Graph representation
Semi supervised learning

Journal

F
Future Generation Computer Systems-The International Journal of eScience
IF:
6.1
Papers:
6.9K
Citations:
2.3W

Organization

State University System of Florida cover
State University System of Florida
Scholars:
12.8W
Papers: 10.9W
Citations: 130
F
Florida State University
Scholars:
1.1W
Papers: 8.6K
Citations: 2.0W
W
weifang university of science & technology
Scholars:
660
Papers: 592
Citations: 0
C
china university of petroleum
Scholars:
4.1W
Papers: 2.7W
Citations: 30
researcher View more organizations
Cited Papers

Cited Papers

errShare
errSave
errShare
errSave
The mitochondrial complex I inhibitor rotenone triggers a cerebral tauopathy
err2005-10-10
err0
PREAI
errGünter U. Höglinger; Annie Lannuzel; Myriam Escobar Khondiker; Patrick P. Michel; Charles Duyckaerts; Jean Féger; Pierre Champy; Annick Prigent; Fadia Medja; Anne Lombes; Wolfgang H. Oertel; Merle Ruberg; Etienne C. Hirsch
errShare
errSave
Modulation of Progressive Leaf Senescence by the Red:Far-Red Ratio of Incident Light
err1989-06-01
err0
PREAI
errJuan J. Guiamet; Jorge G. Willemoes; Edgardo R. Montaldi
errShare
errSave
A Mouse Model of Uterine Leiomyosarcoma
err2004-01-01
err0
errOAAI
errKaterina Politi; Matthias Szabolcs; Peter Fisher; Ana Kljuic; Thomas Ludwig; Argiris Efstratiadis
errShare
errSave
Low-Level Laser Therapy at 635 nm for Treatment of Chronic Plantar Fasciitis: A Placebo-Controlled, Randomized Study
err2015-09-01
err0
PREAI
errDavid M. Macias; Michael J. Coughlin; Kerry Zang; Faustin R. Stevens; James R. Jastifer; Jesse F. Doty
errShare
errSave
Facial expression recognition from near-infrared videos
err2011-08-01
err520
PREAI
errZhao, Guoying; Huang, Xiaohua; Taini, Matti; Li, Stan Z.; Pietikainen, Matti
errShare
errSave
researcher View more