arrow
返回

Reducing training data needs with minimal multilevel machine learning (M3L)

delete2024-06-06
delete6
delete
OA
AI
S
Stefan Heinen
D
Danish Khan
G
Guido Falk von Rudorff
K
Konstantin Karandashev
D
Daniel Jose Arismendi Arrieta
A
Alastair J. A. Price
S
Surajit Nandi
A
Arghya Bhowmik
K
Kersti Hermansson
O
O. Anatole von Lilienfeld *
DOI:10.1088/2632-2153/ad4ae5delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
For many machine learning applications in science, data acquisition, not training, is the bottleneck even when avoiding experiments and relying on computation and simulation. Correspondingly, and in order to reduce cost and carbon footprint, training data efficiency is key. We introduce minimal multilevel machine learning (M3L) which optimizes training data set sizes using a loss function at multiple levels of reference data in order to minimize a combination of prediction error with overall training data acquisition costs (as measured by computational wall-times). Numerical evidence has been obtained for calculated atomization energies and electron affinities of thousands of organic molecules at various levels of theory including HF, MP2, DLPNO-CCSD(T), DFHFCABS, PNOMP2F12, and PNOCCSD(T)F12, and treating them with basis sets TZ, cc-pVTZ, and AVTZ-F12. Our M3L benchmarks for reaching chemical accuracy in distinct chemical compound sub-spaces indicate substantial computational cost reductions by factors of similar to 1.01, 1.1, 3.8, 13.8, and 25.8 when compared to heuristic sub-optimal multilevel machine learning (M2L) for the data sets QM7b, QM9 LCCSD ( T ) , Electrolyte Genome Project, QM9 AE CCSD ( T ) , and QM9 EA CCSD ( T ) , respectively. Furthermore, we use M2L to investigate the performance for 76 density functionals when used within multilevel learning and building on the following levels drawn from the hierarchy of Jacobs Ladder: LDA, GGA, mGGA, and hybrid functionals. Within M2L and the molecules considered, mGGAs do not provide any noticeable advantage over GGAs. Among the functionals considered and in combination with LDA, the three on average top performing GGA and Hybrid levels for atomization energies on QM9 using M3L correspond respectively to PW91, KT2, B97D, and tau-HCTH, B3LYP & lowast; (VWN5), and TPSSH.
Keyword:
quantum machine learning
computational cost
multilevel machine learning
delta learning
quantum chemistry
kernel ridge regression

期刊

M
Machine Learning-Science and Technology
IF:
4.6
论文数:
1.1K
被引数:
3.4K

机构

U
Universitat Kassel
学者数:
4.0K
论文数: 3.5K
被引数: 39
V
Vector Institute for Artificial Intelligence
学者数:
217
论文数: 164
被引数: 4
U
uppsala university
学者数:
3.7W
论文数: 3.4W
被引数: 47
U
university of toronto
学者数:
14.8W
论文数: 12.0W
被引数: 165
学者 查看更多机构
引用论文

引用论文

err分享
err收藏
err分享
err收藏
Cheap Turns Superior: A Linear Regression-Based Correction Method to Reaction Energy from the DFT
err2022-09-16
err2
PREAI
errNandi, Surajit; Busk, Jonas; Jorgensen, Peter Bjorn; Vegge, Tejs; Bhowmik, Arghya
err分享
err收藏
Predicting Materials Properties with Little Data Using Shotgun Transfer Learning使用shot弹枪迁移学习在少量数据的情况下预测材料特性
err2019-09-30
err274
errOAAI
errYamada, Hironao; Liu, Chang; Wu, Stephen; Koyama, Yukinori; Ju, Shenghong; Shiomi, Junichiro; Morikawa, Junko; Yoshida, Ryo
err分享
err收藏
err分享
err收藏
err分享
err收藏
学者 查看更多内容