Return
Interpreting black-box machine learning based species distribution models
DOI:10.1016/j.ecocom.2026.101158.png)
Abstract
En 中文
Species Distribution Models (SDMs) are widely used to analyze the relationship between species occurrence and environmental factors, offering critical ecological and evolutionary insights. However, the complexity of SDMs, coupled with high-dimensional environmental data, can hinder model interpretability, especially machine learning (ML) based SDMs. To address this, we propose and evaluate an interpretable modeling approach that integrates Feature Selection (FS) techniques to enhance both transparency and predictive performance of ML based SDMs. In the present study, we predict the distribution of seven bird species by evaluating the impacts of six univariate filter-based FS methods (each tested with two thresholds), six wrapper methods, two multivariate approaches, and ensemble wrappers on the classification performance and interpretability. We employed four black-box ML classifiers: Extreme Gradient Boosting (XGB), Decision Tree (DT), Random Forest (RF), and Light Gradient Boosting Machine (LGBM) as well three performance criteria: accuracy, Kappa, and F1-score. Moreover, we used four interpretability techniques (Lime, Shapley additive explanations, Accumulated Local Effects, and Global surrogate) to analyze feature importance and understand how selected variables influence predictions. The findings indicate that wrapper methods outperformed both filter methods and their corresponding ensembles. Further-more, the interpretability analysis across the four classifiers indicated that the highly influential features are Temperature seasonality, Maximum Temperature of the Warmest Month, and Minimum Temperature of the Coldest Month, suggesting their critical role in the prediction process and their overall importance in interpretability assessments. Additionally, the findings confirmed the reliability of the interpretability techniques used and highlighted the effectiveness of the Global Surrogate approach in addressing the accuracy-interpretability trade-off across all black-box models.
Keywords:
Species distribution models
Interpretability
Feature selection
Ensemble learning
Machine learning
Global interpretability
Local interpretability
Journal
IF:
3.4
Papers:
951
Citations:
2.3K

