arrow
Return

Impact of QTL number, heritability and reference population size on the benefit of machine learning models over GBLUP in genomic prediction

delete2026-07-13
delete0
delete
OA
AI
J
Jifan Yang *
Y
Yvonne C. J. Wientjes
M
M.P.L. Calus
P
Pascal Duenk
DOI:10.1186/s12711-026-01065-6delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
With the successful application of machine learning (ML) in many areas, its potential for genomic prediction (GP) in animal and plant breeding has attracted growing attention. However, despite numerous studies, the benefit of ML models over GBLUP remains controversial. Some studies reported higher accuracy of ML than linear models with small reference populations, but this benefit often disappeared in larger datasets. Additionally, the number of QTL also affects the prediction accuracy of ML models. The aim of this study was to compare the performance of three ML models, random forest (RF), support vector regression (SVR), and multilayer perceptron with residual networks (MLP-ResNet), with that of GBLUP. We simulated livestock populations with different data characteristics, including heritability, reference population size, and the number of QTL underlying the trait. Our results showed that the number of QTL strongly affected the prediction accuracy of RF and MLP-ResNet, but had little effect on GBLUP and SVR. The accuracy of RF dropped markedly with increasing QTL number, whereas the accuracy of MLP-ResNet decreased with increasing QTL number when reference populations exceeded 10,000. Both RF and MLP-ResNet achieved benefits over GBLUP only when the number of QTL was small (i.e. ≤ 100). RF outperformed GBLUP with small reference populations, whereas MLP-ResNet outperformed GBLUP with large populations. SVR showed marginally lower accuracy than GBLUP across scenarios and required more computation time. Among all models, GBLUP showed the lowest dispersion bias. RF and MLP-ResNet achieved higher accuracy than GBLUP only when the number of QTL was small, whereas in all other situations GBLUP outperformed the ML models. The benefit of ML models over GBLUP occurred mainly in scenarios where the infinitesimal model assumption of GBLUP did not hold. Our results suggest that when introducing a new GP model, particularly tree-based or neural network methods, it is essential to evaluate its performance on simulated datasets across different numbers of QTL and reference population sizes. These findings provide insights for applying ML in animal and plant breeding.

Journal

G
Genetics Selection Evolution
IF:
3.1
Papers:
54
Citations:
0

Organization

W
Wageningen University & Research
Scholars:
2.9W
Papers: 2.8W
Citations: 55