Return
Improving polygenic score prediction for underrepresented groups through transfer learning
DOI:10.1038/s41467-026-68696-7.png)
Abstract
En 中文
The advent of large biobanks has substantially increased the accuracy of polygenic scores (PGS). However, most existing PGSs were derived from European-ancestry data and often exhibit reduced predictive performance when applied to individuals of non-European ancestries. Transfer Learning offers a promising strategy to address this limitation by leveraging information learned in one population to improve prediction in another. Here, we introduce GPTL, an R package that implements three Transfer Learning based approaches for developing PGS: (1) gradient descent with early stopping, (2) a penalized regression model that shrinks variant-effect estimates toward prior values, and (3) a Bayesian method with a finite-mixture prior that enables integration of multiple prior sources of information. Using both simulated data and real data from the UK-Biobank and All of Us, we demonstrate that PGS generated with GPTL’s Transfer Learning algorithms consistently outperform single-ancestry PGS and, in many settings, match or exceed the performance of multi-ancestry ensemble-based PGS. Our software can be used with either individual genotype-phenotype data or summary statistics from genome-wide association studies. Polygenic scores often underperform in non‑European ancestries. Here, the authors present GPTL, an R package with three transfer‑learning methods that improve cross‑ancestry PGS using either individual‑level data or GWAS summary statistics.
Keywords:
Polygenic scores
Transfer Learning
Cross-ancestry prediction
Biobanks
Genomic prediction
Journal
IF:
15.7
Papers:
9.2W
Citations:
91.2W

