arrow
Return

An autoencoder-based deep learning method for genotype imputation

delete2022-11-03
delete6
delete
OA
AI
M
Meng Song
J
Jonathan Greenbaum
J
Joseph Luttrell
W
Weihua Zhou
C
Chong Wu
Z
Zhe Luo
C
Chuan Qiu
L
Lan‐Juan Zhao
K
Kuan‐Jui Su
Q
Qing Tian
H
Hui Shen
H
Huixiao Hong
P
Ping Gong
X
Xinghua Shi
H
Hong‐Wen Deng *
C
Chaoyang Zhang *
DOI:10.3389/frai.2022.1028978delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Genotype imputation has a wide range of applications in genome-wide association study (GWAS), including increasing the statistical power of association tests, discovering trait-associated loci in meta-analyses, and prioritizing causal variants with fine-mapping. In recent years, deep learning (DL) based methods, such as sparse convolutional denoising autoencoder (SCDA), have been developed for genotype imputation. However, it remains a challenging task to optimize the learning process in DL-based methods to achieve high imputation accuracy. To address this challenge, we have developed a convolutional autoencoder (AE) model for genotype imputation and implemented a customized training loop by modifying the training process with a single batch loss rather than the average loss over batches. This modified AE imputation model was evaluated using a yeast dataset, the human leukocyte antigen (HLA) data from the 1,000 Genomes Project (1KGP), and our in-house genotype data from the Louisiana Osteoporosis Study (LOS). Our modified AE imputation model has achieved comparable or better performance than the existing SCDA model in terms of evaluation metrics such as the concordance rate (CR), the Hellinger score, the scaled Euclidean norm (SEN) score, and the imputation quality score (IQS) in all three datasets. Taking the imputation results from the HLA data as an example, the AE model achieved an average CR of 0.9468 and 0.9459, Hellinger score of 0.9765 and 0.9518, SEN score of 0.9977 and 0.9953, and IQS of 0.9515 and 0.9044 at missing ratios of 10% and 20%, respectively. As for the results of LOS data, it achieved an average CR of 0.9005, Hellinger score of 0.9384, SEN score of 0.9940, and IQS of 0.8681 at the missing ratio of 20%. In summary, our proposed method for genotype imputation has a great potential to increase the statistical power of GWAS and improve downstream post-GWAS analyses.
Keywords:
genotype imputation
deep learning
autoencoder
paired sample t-test
GWAS
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

F
Frontiers in Artificial Intelligence
IF:
4.7
Papers:
2.3K
Citations:
4.4K

Organization

U
utmd anderson cancer center
Scholars:
3.0W
Papers: 2.4W
Citations: 27
U
us food & drug administration (fda)
Scholars:
1.4W
Papers: 9.3K
Citations: 3
T
tulane university
Scholars:
1.3W
Papers: 1.0W
Citations: 9
United States Department of Defense cover
United States Department of Defense
Scholars:
2.8W
Papers: 2.3W
Citations: 172
U
university of texas system
Scholars:
18.5W
Papers: 15.6W
Citations: 210
U
University of Southern Mississippi
Scholars:
2.2K
Papers: 2.0K
Citations: 3.6K
M
Michigan Technological University
Scholars:
5.0K
Papers: 4.4K
Citations: 6.4K
researcher View more organizations