arrow
Return

Comprehensive predictive analytics for collaborators’ answers, code quality, and dropout: stack overflow case study

delete2025-07-23
delete0
PRE
AI
E
Elijah Zolduoarrati *
S
Sherlock A. Licorish
N
Nigel Stanger
DOI:10.1007/s10664-025-10692-4delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Previous studies that used data from Stack Overflow to develop predictive models often employed limited benchmarks of 3–5 models or adopted arbitrary selection methods. Despite being insightful, their limited scope suggests the need to benchmark more models to avoid overlooking untested algorithms. Our study evaluates 21 algorithms across three tasks: predicting the number of question a user is likely to answer, their code quality violations, and their dropout status. We employed normalisation, standardisation, as well as logarithmic and power transformations paired with Bayesian hyperparameter optimisation and genetic algorithms. CodeBERT, a pre-trained language model for both natural and programming languages, was fine-tuned to classify user dropout given their posts (questions and answers) and code snippets. We found Bagging ensemble models combined with standardisation achieved the highest R2 value (0.821) in predicting users’ answers. The Stochastic Gradient Descent regressor, followed by Bagging and Epsilon Support Vector Machine models, consistently demonstrated superior performance to other benchmarked algorithms in predicting users’ code quality across multiple quality dimensions and languages. Extreme Gradient Boosting paired with log-transformation exhibited the highest F1-score (0.825) in predicting users’ dropout. CodeBERT was able to classify users’ dropout with a final F1-score of 0.809, validating the performance of Extreme Gradient Boosting that was solely based on numerical data. Overall, our benchmarking of 21 algorithms provides multiple insights. Researchers can leverage findings regarding the most suitable models for specific target variables, and practitioners can utilise the identified optimal hyperparameters to reduce the initial search space during their own hyperparameter tuning processes.
Keywords:
Answers
Code Quality
Prediction
Stack Overflow
User Dropout

Journal

Empirical Software Engineering cover
Empirical Software Engineering
IF:
3.6
Papers:
2.0K
Citations:
5.3K

Organization

D
Department of Information Science
Scholars:
36
Papers: 28
Citations: 0
Cited Papers

Cited Papers

Residual geochemical gold grade prediction using extreme gradient boosting
err2022-01-01
err0
errOAAI
errBemah Ibrahim; Fareed Majeed; Anthony Ewusi; Isaac Ahenkorah
errShare
errSave
Transformer Architectures
err2024-01-01
err0
PREAI
errSingh,Pradeep; Raman,Balasubramanian
errShare
errSave
err
IF0
err
err0
PREAI
err
errShare
errSave
Data Cleaning
err2013-01-01
err0
PREAI
errVan den Broeck,Jan; Fadnes,Lars Thore
errShare
errSave
Developing an online hate classifier for multiple social media platforms
err2020-01-02
err118
errOAAI
errSalminen, Joni; Hopf, Maximilian; Chowdhury, Shammur A.; Jung, Soon-gyo; Almerekhi, Hind; Jansen, Bernard J.
errShare
errSave
Aeromagnetic Compensation Algorithm Robust to Outliers of Magnetic Sensor Based on Huber Loss Method
err2019-07-15
err24
PREAI
errGe, Jian; Li, Han; Wang, Hongpeng; Dong, Haobin; Liu, Huan; Wang, Wenjie; Yuan, Zhiwen; Zhu, Jun; Zhang, Haiyang
errShare
errSave
How have views on Software Quality differed over time? Research and practice viewpoints
err2023-01-01
err6
PREAI
errNdukwe, Ifeanyi G.; Licorish, Sherlock A.; Tahir, Amjed; MacDonell, Stephen G.
errShare
errSave
researcher View more