Return
Mapping-based genome size estimation
DOI:10.1186/s12864-025-11640-8.png)
Abstract
En 中文
While the size of chromosomes can be measured under a microscope, obtaining the exact size of a genome remains a challenge. Biochemical methods and k-mer distribution-based approaches allow only estimations. An alternative approach to estimate the genome size based on high contiguity assemblies and read mappings is presented here. Analyses of Arabidopsis thaliana and Beta vulgaris data sets are presented to show the impact of different parameters. Oryza sativa, Brachypodium distachyon, Solanum lycopersicum, Vitis vinifera, and Zea mays were also analyzed to demonstrate the broad applicability of this approach. Further, MGSE was also used to analyze Escherichia coli, Saccharomyces cerevisiae, and Caenorhabditis elegans datasets to show its utility beyond plants. Mapping-based Genome Size Estimation (MGSE) and additional scripts are available on GitHub: https://github.com/bpucker/MGSE. MGSE predicts genome sizes based on short reads or long reads requiring a minimal coverage of 5-fold.
Keywords:
Genome size
Short reads
Long reads
Next generation sequencing
Long read sequencing
Read mapping
Nanopore sequencing
Journal
IF:
3.7
Papers:
1.9W
Citations:
5.2W
Organization
Cited Papers
The Structural Features of Thousands of T-DNA Insertion Sites Are Consistent with a Double-Strand Break Repair-Based Insertion Mechanism
MOLECULAR PLANT
IF24.1
Indica rice genome assembly, annotation and mining of blast disease resistance genes
BMC GENOMICS
IF3.7

