arrow
Return

MolSelector: A Machine Learning Framework for Interpretable Subset Selection from Molecular Science Data Sets

delete
delete0
PRE
AI
C
Caitlin Whitter *
A
Aurora E. Clark
A
Alex Pothen
R
Rajiv Khanna
DOI:10.1021/acs.jcim.6c01332delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文

We describe MolSelector, a machine learning framework for selecting representative subsets of molecular science data sets for efficient and accurate neural network training. As part of MolSelector, we introduce the interpretable Atypicality Score algorithm, which identifies molecules that are typical versus atypical of their data set and selects a representative sample of the data set on that basis. These subsets can be used for subsequent neural network training, leading to fast model training with reduced memory requirements while still maintaining low test set error across several molecular property prediction tasks. Our experiments demonstrated that the Atypicality Score subsets resulted in errors close to the errors obtained when the entire training set is used, while achieving a 3× or greater training time speedup compared to this baseline. Additionally, we analyzed the Atypicality Score algorithm’s typical and atypical molecule assignments to gain insight into the molecular characteristics the algorithm determined most beneficial for subset selection.

Journal

Journal of Chemical Information and Modeling cover
Journal of Chemical Information and Modeling
IF:
5.3
Papers:
9.1K
Citations:
4.0W

Organization

U
University of Utah
Scholars:
3.0W
Papers: 2.2W
Citations: 4.6W
P
Purdue University
Scholars:
2.7W
Papers: 2.1W
Citations: 147