Return
Minimum epistasis interpolation for sequence-function relationships
DOI:10.1038/s41467-020-15512-5.png)
Abstract
En 中文
Massively parallel phenotyping assays have provided unprecedented insight into how multiple mutations combine to determine biological function. While such assays can measure phenotypes for thousands to millions of genotypes in a single experiment, in practice these measurements are not exhaustive, so that there is a need for techniques to impute values for genotypes whose phenotypes have not been directly assayed. Here, we present an imputation method based on inferring the least epistatic possible sequence-function relationship compatible with the data. In particular, we infer the reconstruction where mutational effects change as little as possible across adjacent genetic backgrounds. The resulting models can capture complex higher-order genetic interactions near the data, but approach additivity where data is sparse or absent. We apply the method to high-throughput transcription factor binding assays and use it to explore a fitness landscape for protein G. High-throughput combinatorial mutagenesis assays are useful to screen the function of many different sequences but they are not exhaustive. Here, Zhou and McCandlish develop a method to impute such missing genotype-phenotype data based on inferring the least epistatic sequence-function relationship.
Keywords:
FITNESS LANDSCAPE
PROTEIN EVOLUTION
ORDER EPISTASIS
MODELS
PREDICTABILITY
MUTATIONS
NETWORK
PATHS
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
15.7
Papers:
9.3W
Citations:
91.2W
Organization
Cited Papers
Compact, universal DNA microarrays to comprehensively determine transcription-factor binding site specificities
NATURE BIOTECHNOLOGY
IF41.7
Deep mutational scanning of an RRM domain of the Saccharomyces cerevisiae poly(A)-binding protein
RNA
IF5

