MolSelector: A Machine Learning Framework for Interpretable Subset Selection from Molecular Science Data Sets
Abstract
We describe MolSelector, a machine learning framework for selecting representative subsets of molecular science data sets for efficient and accurate neural network training. As part of MolSelector, we introduce the interpretable Atypicality Score algorithm, which identifies molecules that are typical versus atypical of their data set and selects a representative sample of the data set on that basis. These subsets can be used for subsequent neural network training, leading to fast model training with reduced memory requirements while still maintaining low test set error across several molecular property prediction tasks. Our experiments demonstrated that the Atypicality Score subsets resulted in errors close to the errors obtained when the entire training set is used, while achieving a 3× or greater training time speedup compared to this baseline. Additionally, we analyzed the Atypicality Score algorithm’s typical and atypical molecule assignments to gain insight into the molecular characteristics the algorithm determined most beneficial for subset selection.

