Return
A Lightweight and Accurate Machine Learning Model for Predicting Outliers in Educational Data
Z
S
H
Z
DOI:10.1142/S0218001426510018.png)
Abstract
En 中文
The accumulation of massive amounts of student learning data has provided a solid foundation for predicting academic performance, while also presenting new opportunities for personalized teaching and optimizing teaching resources. It is worth noting that a small number of outlier data points exist within educational data. These data points are difficult to analyze through the mining of latent patterns, and they do not easily capture the dynamic changes in student performance. However, these outlier data points play an indispensable role in the formulation of personalized teaching strategies for students. Traditional machine learning methods have limitations in handling issues related to concentrated performance data distribution and insufficient prediction accuracy for outliers. To address these challenges, this paper proposes a lightweight long short term memory (LSTM) prediction model (BFE-LSTM) enhanced by binning feature engineering. By integrating binning encoding techniques with time-series feature analysis, the model significantly improves prediction accuracy and robustness. Empirical studies demonstrate that, while reducing model complexity, the proposed model optimizes its modeling capability for long tail distributed data, thereby significantly enhancing the prediction accuracy of outliers. Experimental results show that the R2 value of BFE-LSTM model is 0.9471. Compared with the designed baseline LSTM model, the mean squared error (MSE) is reduced to 95.1%, the mean absolute error (MAE) is improved to 94.6%, and the outlier correction error reaches 100%. At the same time, there is no additional burden on computing efficiency. This study provides an efficient and lightweight prediction framework for educational data mining, achieving breakthroughs in feature interpretability, computational efficiency, and model generalization ability, and offering a new methodological support for precise learning situation analysis.
Keywords:
Standardization of multi-source heterogeneous data sets
BFE-LSTM model
outlier prediction
residual model
SHAP
Journal
IF:
1.1
Papers:
161
Citations:
2.0K
