arrow
返回

Data Augmentation by Guided Deep Interpolation

delete2021-11-01
delete13
PRE
AI
G
Gergely Szlobodnyik *
L
Lóránt Farkas
DOI:10.1016/j.asoc.2021.107680delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
State-of-the-art machine learning algorithms require large amount of high quality data. In practice, however, the sample size is commonly low and data is imbalanced along different class labels. Low sample size and imbalanced class distribution can significantly deteriorate the predictive performance of machine learning models. In order to overcome data quality issues, we propose a novel data augmentation method, Guided Deep Interpolation (GDI). It is based on a convolutional auto-encoder network, which is equipped with an auxiliary linear self-expressive layer. The network is trained by minimizing a composite objective function so that to extract the underlying clustered structure of semantic similarities of data points while high reconstruction quality is also preserved. The trained network is used to define a sampling strategy and a synthetic data generation procedure. Making use of the weights of the self-expressive layer, we introduce a measure of semantic variability to quantify how similar a data point to other data points on average. Based on the proposed measure of semantic variability, a joint distribution is defined. Using the distribution we can draw pairs of similar data points so that one point is semantically underrepresented (isolated) while its pair possesses relatively high semantic variability. A sampled pair is interpolated in the deep feature space of the network so that to increase semantic variability while preserve class label of the semantically underrepresented data point. The trained decoder is used to determine pixel space representations of latent space interpolations. The resulting data augmentation procedure generates synthetic samples by increasing the semantic variability of semantically underrepresented instances in a class label preserving way. Our experimental results show that the proposed method outperforms traditional and generative model-based data augmentation methods on low sample size and imbalanced data sets. (C) 2021 Elsevier B.V. All rights reserved.
Keyword:
Autoencoders
Data augmentation
Imbalanced data
Self-expressiveness
Interpolation
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Applied Soft Computing 封面图
Applied Soft Computing
IF:
6.6
论文数:
1.4W
被引数:
4.8W

机构

N
nokia corporation
学者数:
1.8K
论文数: 1.5K
被引数: 1
引用论文

引用论文

Disease activity in primary progressive multiple sclerosis: a systematic review and meta-analysis
err2023-11-06
err0
errOAAI
errKatelijn M. Blok; Joost van Rosmalen; Nura Tebayna; Joost Smolders; Beatrijs Wokke; Janet de Beukelaar
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
Arsenic removal from aqueous solutions by adsorption using novel MIL-53(Fe) as a highly efficient adsorbent使用新型MIL-53(Fe) 作为高效吸附剂通过吸附从水溶液中去除砷
err2015-01-01
err0
PREAI
errTuan. A. Vu; Giang. H. Le; Canh. D. Dao; Lan. Q. Dang; Kien. T. Nguyen; Quang. K. Nguyen; Phuong. T. Dang; Hoa. T. K. Tran; Quang. T. Duong; Tuyen. V. Nguyen; Gun. D. Lee
err分享
err收藏
High definition IEEE AVS decoder on ARM NEON platform
err2013-09-01
err0
PREAI
errRonggang Wang; Jie Wan; Wenmin Wang; Zhenyu Wang; Shengfu Dong; Wen Gao
err分享
err收藏
Influence of the Thermal Conductivity of Air on the Moisture Homogeneity of a Tray Dryer
err2019-03-31
err0
errOAAI
errEsparza Jessica; Grisales Felipe; Pérez José; Ordóñez Eduardo; Lobatón Fabián
err分享
err收藏
The role of natural killer T cells in a mouse model with spontaneous bile duct inflammation
err2017-02-20
err0
errOAAI
errElisabeth Schrumpf; Xiaojun Jiang; Sebastian Zeissig; Marion J. Pollheimer; Jarl Andreas Anmarkrud; Corey Tan; Mark A. Exley; Tom H. Karlsen; Richard S. Blumberg; Espen Melum
err分享
err收藏
学者 查看更多内容