arrow
Return

Informed training set design enables efficient machine learning-assisted directed protein evolution

delete2021-11-01
delete69
delete
OA
AI
B
Bruce J. Wittmann
Y
Yisong Yue
F
Frances H. Arnold *
DOI:10.1016/j.cels.2021.07.008delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Directed evolution of proteins often involves a greedy optimization in which the mutation in the highest fitness variant identified in each round of single-site mutagenesis is fixed. The efficiency of such a singlestep greedy walk depends on the order in which beneficial mutations are identified-the process is path dependent. Here, we investigate and optimize a path-independent machine learning-assisted directed evolution (MLDE) protocol that allows in silico screening of full combinatorial libraries. In particular, we evaluate the importance of different protein encoding strategies, training procedures, models, and training set design strategies on MLDE outcome, finding the most important consideration to be the implementation of strategies that reduce inclusion of minimally informative holes(protein variants with zero or extremely low fitness) in training data. When applied to an epistatic, hole-filled, four-site combinatorial fitness landscape, our optimized protocol achieved the global fitness maximum up to 81-fold more frequently than singlestep greedy optimization. A record of this paper's transparent peer review process is included in the supplemental information.
Keywords:
FITNESS LANDSCAPE
EPISTASIS
DATABASE

Journal

Cell Systems cover
Cell Systems
IF:
7.7
Papers:
1.4K
Citations:
1.0W

Organization

C
California Institute of Technology
Scholars:
2.9W
Papers: 2.5W
Citations: 4.9W