Return
Advancing Wheat Single-Nucleotide Polymorphism Data Analysis with Explainable Deep Learning Models
DOI:10.1080/08839514.2025.2565169.png)
Abstract
En 中文
The application of machine learning (ML) and deep learning (DL) is transforming scientific fields, including foodomics. This study advances the application of artificial neural networks (ANNs) for analyzing single-nucleotide polymorphism (SNP) data in foodomics. We introduce underutilized mechanisms in fields, such as data augmentation, dropout, batch normalization, learning rate scheduling, and Bayesian optimization for hyper-parameter optimization (HPO), enabling more robust and generalizable models. Our ANN achieves state-of-the-art performance on a publicly available dataset from prior work, with an average RMSE reduction of 0.017 (10.5%) over the previous ANN model and statistically significant improvements on strong traditional baselines, including Random Forest, XGBoost, LASSO, and Ridge regression. To enhance interpretability, we integrate SHAP (SHapley Additive exPlanations), which highlights the most influential SNP markers contributing to predictions, potentially identifying novel genomic markers. We also emphasize reproducibility, following best practices in code and data sharing. By making both our code and preprocessed dataset publicly available, we aim to support transparency and foster further research. Our results show that ANNs can serve not only as high-performing predictive models but also as explainable tools for SNP analysis in foodomics, contributing to the foundation of explainable artificial intelligence in this emerging field.
Journal
A
IF:
4.3
Papers:
66
Citations:
0

