返回
Active learning of molecular data for task-specific objectives
DOI:10.1063/5.0229834.png)
摘要
En 中文
Active learning (AL) has shown promise to be a particularly data-efficient machine learning approach. Yet, its performance depends on the application, and it is not clear when AL practitioners can expect computational savings. Here, we carry out a systematic AL performance assessment for three diverse molecular datasets and two common scientific tasks: compiling compact, informative datasets and targeted molecular searches. We implemented AL with Gaussian processes (GP) and used the many-body tensor as molecular representation. For the first task, we tested different data acquisition strategies, batch sizes, and GP noise settings. AL was insensitive to the acquisition batch size, and we observed the best AL performance for the acquisition strategy that combines uncertainty reduction with clustering to promote diversity. However, for optimal GP noise settings, AL did not outperform the randomized selection of data points. Conversely, for targeted searches, AL outperformed random sampling and achieved data savings of up to 64%. Our analysis provides insight into this task-specific performance difference in terms of target distributions and data collection strategies. We established that the performance of AL depends on the relative distribution of the target molecules in comparison to the total dataset distribution, with the largest computational savings achieved when their overlap is minimal.
Keyword:
ORBITAL ENERGIES
PREDICTIONS
POTENTIALS
期刊
IF:
3.1
论文数:
7.2W
被引数:
23.2W
机构
引用论文
Efficient and accurate machine-learning interpolation of atomic energies in compositions with many species
PHYSICAL REVIEW B
IF3.7
Comparative study of GeO2/Ge and SiO2/Si structures on anomalous charging of oxide films upon water adsorption revealed by ambient-pressure X-ray photoelectron spectroscopy通过环境压力x射线光电子能谱揭示的GeO2/Ge和SiO2/Si结构对水吸附后氧化膜异常充电的比较研究

