Return
Pangenome-based genome inference using integer programming
DOI:10.1101/gr.280567.125.png)
Abstract
En 中文
Affordable genotyping methods are essential in genomics. Commonly used genotyping methods primarily support single-nucleotide variants and short indels but neglect structural variants. Additionally, accuracy of read alignments to a reference genome is unreliable in highly polymorphic and repetitive regions, further impacting genotyping performance. Recent works highlight the advantage of pangenome graphs in addressing these challenges. Building on these developments, we propose a rigorous alignment-free genotyping method. Our optimization framework identifies a path through the pangenome graph that maximizes the matches between the path and substrings of sequencing reads (e.g., k-mers) while minimizing recombination events (haplotype switches) along the path. We prove that this problem is NP-hard and develop efficient integer-programming solutions. We benchmark the algorithm using downsampled short-read data sets from homozygous human cell lines with coverage ranging from 0.1x to 10x. Our algorithm accurately estimates complete major histocompatibility complex (MHC) haplotype sequences with small edit distances from the ground-truth sequences, providing a significant advantage over existing methods on low-coverage inputs.
Keywords:
SEQUENCE
IMPUTATION
EFFICIENT
GRAPHS
Journal
IF:
5.5
Papers:
5.6K
Citations:
4.3W
Organization
Cited Papers
Pangenome-based genome inference allows efficient and accurate genotyping across a wide spectrum of variant classes
NATURE GENETICS
IF31.8
The Bovine Pangenome Consortium: democratizing production and accessibility of genome assemblies for global cattle breeds and other bovine species
GENOME BIOLOGY
IF9.4
Efficient phasing and imputation of low-coverage sequencing data using large reference panels
NATURE GENETICS
IF31.8

