Return
Long-term software reliability prediction via nonparametric kernel regression methods
DOI:10.1016/j.infsof.2026.108284.png)
Abstract
En 中文
Context: Software reliability prediction is essential for quality assurance and release decisions in modern software development. Long-horizon prediction remains challenging, especially when fault-count data are limited and testing conditions evolve over time. Objective: This paper aims to develop a unified nonparametric kernel regression approach for long-horizon software reliability prediction, and to clarify the effects of structural flexibility and the sequential updating strategy on predictive performance. Methods: We propose a unified nonparametric kernel regression framework that integrates both single-kernel and multi-kernel estimators with a closed-loop recursive forecasting protocol. The multi-kernel extension represents software reliability growth as a combination of heterogeneous kernels, where the kernel weights and bandwidths are simultaneously optimized by least-squares cross-validation. Experiments are conducted on sixteen real-world software development project datasets, including both classical closed-source and open-source projects on GitHub, under three training cutoffs (20%, 50%, and 80%), corresponding to the early, middle, and late testing phases. Results: The results show that the proposed kernel-based models achieve competitive point-prediction performance compared with representative NHPP-based, wavelet-based, and LSTM-based baselines across the three testing phases. In the early testing phase, MKLL-D achieves the largest number of best-performing cases, while LSTM and Wavelet also perform best on some datasets. In the middle testing phase, Wavelet shows strong dataset-wise consistency, whereas LL-based multi-kernel models remain competitive. In the late testing phase, LSTM and NHPP-based SRMs become more competitive, while LL-based kernel models remain competitive on several datasets. Conclusion: The proposed framework provides a practical nonparametric approach to long-horizon software reliability prediction. The results suggest that multi-kernel local linear models are especially useful in the early testing phase, as shown by the strong performance of MKLL-D. However, this advantage is not consistent across all datasets or phases. Sequential updating is mainly helpful for early-stage local linear estimators; in the middle phase, fixed-parameter configurations become more competitive, and in the late phase the relative performance becomes more dataset-dependent.
Keywords:
Software reliability
Nonparametric prediction
Kernel regression
Multi-kernel learning
Data-driven approach
Journal
IF:
4.3
Papers:
3.8K
Citations:
7.7K
Organization
Cited Papers
No cited papers available

