Return
Gradient Boosting-Based Predictive Uncertainty Estimation for Tabular Data Using Statistical Variance and Proper Scoring Rule
DOI:10.1109/TASE.2026.3671833.png)
Abstract
En 中文
Tabular data is the most widely used data form in real-world applications, and tree-based models are suitable for it due to their model structures. In practice, it is crucial to quantify predictive uncertainty, and several uncertainty estimation methods have been developed for tree-based models. However, some of them compromise point estimation accuracy or incur high computational overhead. To address this limitation, we propose a variance-based method for predictive uncertainty estimation of tabular data using gradient boosting decision tree (VarBoost). VarBoost estimates variance by analytically calculating statistical characteristics based on gradient boosting strategy, ensuring efficient and accurate variance estimation while maintaining point estimation performance. To avoid overfitting or underfitting, hyperparameters are determined via cross-validation using proper scoring rules that balance calibration and sharpness. The effectiveness of VarBoost is validated by comparisons with several state-of-the-art uncertainty estimation methods on a collection of UCI tabular datasets. Experimental results demonstrate that VarBoost achieves superior performance in both point estimation and uncertainty estimation. Note to Practitioners—Quantifying predictive uncertainty for tabular data is crucial in risk-sensitive applications, as it enables reliable decision-making under uncertainty. This article proposes a practical method to quantify predictive uncertainty for tabular data using tree-based models. The implementation consists of two key components: (1) an analytical method for efficiently computing predictive variance from a trained gradient boosting decision tree model, which statistically estimates heteroscedastic uncertainty through gradient boosting; and (2) a hyperparameter selection strategy based on proper scoring rules to optimize the trade-off between calibration and sharpness. Experimental results demonstrate that our method not only provides accurate point predictions but also yields high-quality uncertainty estimates, making it particularly suitable for real-world tabular data scenarios.
Keywords:
Tabular data
uncertainty estimation
gradient boosting
statistical variance
proper scoring rules
Journal
IF:
6.4
Papers:
4.9K
Citations:
1.6W

