Return
Spatial information explained by prediction models under block cross-validation
DOI:10.1016/j.isprsjprs.2026.09.034.png)
Abstract
En 中文
Spatial information in geographic data arising from the dependence between variable values and their locations fundamentally shapes the performance and evaluation of spatial prediction models. Model evaluation primarily relies on global accuracy metrics such as root mean squared error (RMSE), which quantify prediction error but do not assess spatial structure a model explains in observed data. However, there are still challenges in the spatial information as a measurable quantity in prediction modelling and the assessment of spatial learning performance under spatial block cross-validation. This study develops the spatial information explained (SIE) metric to assess the extent to which spatial prediction models explain spatial information during modelling, through quantifying the relative reduction in spatial information from a response variable to its prediction residuals. SIE integrates block cross-validation residual computation that enforces spatial separation between training and evaluation sets, and mutual information-based quantification of spatial information contents, and softmax-weighted multi-scale aggregation. A spatial simulation study validates that SIE is effective in measuring spatial learning performance. The developed SIE metric is implemented in predicting C4 natural grass area percentage across Australia with multi-source environmental data of photosynthetic advantage, temperature, and precipitation variables, and ten machine learning models evaluated under block and random cross-validation. Results show that SIE effectively distinguishes machine learning models in spatial learning performance, and reveals that model rankings under SIE differ from those under RMSE, and prediction accuracy and SIE represent distinct evaluation dimensions. Specifically, KNN, the neural network, and XGBoost achieve the highest multi-scale SIE of approximately 0.50 together with high prediction accuracy, SIE rankings correlate weakly with accuracy rankings across the ten models, and random cross-validation overstates SIE relative to block cross-validation for all ten models, by 2.1% to 17.0% with a mean of 9.9%. The validation power of SIE arises from the information-theoretic property that mutual information captures all forms of spatially statistical dependence between residuals and spatial location, including nonlinear spatial associations that domain-average accuracy metrics cannot detect. SIE measures spatial learning performance that complements conventional accuracy metrics, with broad applicability to spatial prediction tasks in Earth and social sciences, and advanced geospatial artificial intelligence and foundation modelling, where capturing geographic structure is analytically required.
Keywords:
Spatial information explained
Model validation
Mutual information
Machine learning
Block cross-validation
Interpretability in GeoAI
Journal
IF:
12.2
Papers:
4.4K
Citations:
3.2W
Organization
No organization information available
Cited Papers
No cited papers available

