Return
Evaluating deep learning for predicting epigenomic profiles
DOI:10.1038/s42256-022-00570-9.png)
Abstract
En 中文
Deep learning has been successful at predicting epigenomic profiles from DNA sequences. Most approaches frame this task as a binary classification relying on peak callers to define functional activity. Recently, quantitative models have emerged to directly predict the experimental coverage values as a regression. As new models with different architectures and training configurations continue to emerge, a major bottleneck is forming due to the lack of ability to fairly assess the novelty of proposed models and their utility for downstream biological discovery. Here we introduce a unified evaluation framework and use it to compare various binary and quantitative models trained to predict chromatin accessibility data. We highlight various modelling choices that affect generalization performance, including a downstream application of predicting variant effects. In addition, we introduce a robustness metric that can be used to enhance model selection and improve variant effect predictions. Our empirical study largely supports that quantitative modelling of epigenomic profiles leads to better generalizability and interpretability.
Keywords:
TRANSCRIPTION FACTOR-BINDING
Journal
IF:
23.9
Papers:
1.3K
Citations:
1.5W
Organization
Cited Papers
Evaluating the informativeness of deep learning annotations for human complex diseases
NATURE COMMUNICATIONS
IF15.7
Unraveling determinants of transcription factor binding outside the core binding site
GENOME RESEARCH
IF5.5
Genome-wide association study implicates novel loci and reveals candidate effector genes for longitudinal pediatric bone accrual
GENOME BIOLOGY
IF9.4
Sequence-based modeling of three-dimensional genome architecture from kilobase to chromosome scale
NATURE GENETICS
IF31.8
Integration of multiple epigenomic marks improves prediction of variant impact in saturation mutagenesis reporter assay
HUMAN MUTATION
IF3.7
FactorNet: A deep learning framework for predicting cell type specific transcription factor binding from nucleotide-resolution sequential data
METHODS
IF4.3

