返回
Training data augmentation using generative models with statistical guarantees for materials informatics
DOI:10.1007/s00500-021-06533-3.png)
摘要
En 中文
In materials double dagger science, the small data size problem occurs frequently when applying machine learning algorithms. This issue is primarily due to, for example, real experimental measurement cost and time-consuming simulations based on materials models, e.g., density functional theory. We address the small data issue by generating training data using generative models. The proposed training data augmentation method can generate data using kernel density estimation models while maintaining statistical guarantees using unbiased estimators of linear regression, which has low computational costs and easy hyper-parameter tuning compared to deep neural networks. In addition, we derive an upper bound for the L-1 loss of the proposed method relative to probability density functions. Experiments were conducted with four benchmark and three materials (i.e., binary compounds, oxide ionic conductivity, and phosphorescent materials) datasets. Regarding the generated data, we examined training and generalization performances using kernel ridge regression compared to those of generative adversarial networks and real-NVPs. For the materials datasets, we analyzed influential factors regarding material properties (conductivity and emission intensity). The experimental results demonstrate that the proposed method can generate data that can be used as new training data.
Keyword:
Linear regression model
Density estimation
Data augmentation
Generative models
Materials informatics
期刊
IF:
2.5
论文数:
1.0W
被引数:
2.1W
机构
引用论文
Chain-growth click copolymerization for the synthesis of branched copolymers with tunable branching densities链增长点击共聚合成具有可调支化密度的支化共聚物
Auto-encoder-based generative models for data augmentation on regression problems
SOFT COMPUTING
IF2.5

