Return
XStack-Net: A stacking-based deep learning framework for robust deepfake detection
DOI:10.1007/s10586-026-06319-y.png)
Abstract
En 中文
The rapid spread of deepfake videos raises significant concerns about digital security, media accuracy, and the reliability of information. To address this issue, this study develops a stacking-based deepfake detection framework called XStack-Net. The proposed framework combines three complementary convolutional neural network architectures—ResNet-50, DenseNet-121, and InceptionV3—and uses XGBoost as the meta-learner. Instead of directly processing full videos, the proposed architectural framework extracts five representative frames from each video and applies Dlib-based face detection and cropping operations to transform video data into a more manageable image dataset. Within this approach, a limited number of frames are selected from each video to extract face regions, creating a balanced image dataset for model training. This approach makes the data preparation and model training process more practical, faster, and computationally more efficient compared to direct video-level processing, while also allowing the model to focus on face regions associated with manipulation. The proposed framework is evaluated using a comprehensive experimental protocol including base model comparisons, meta-learner ablation analysis, analysis of variance, calibration evaluation, and computational cost analysis. In addition to LR, SVM, and ElasticNet, XGBoost was also examined as a final meta-learner, and the results showed that XGBoost’s nonlinear modeling capabilities enabled the most effective combination of basic model outputs. The proposed XGBoost-based stacking framework exhibited stronger classification performance compared to single deep learning models, achieving 98.25% accuracy and 99.81% AUC. Additional experiments conducted on Celeb-DF and OpenForensics datasets revealed strong results under the adopted evaluation protocol. However, performance degradation was observed under the domain shift effect in direct cross-dataset evaluations, indicating that generalization among heterogeneous data distributions remains a challenging problem. Overall, the results demonstrate that XStack-Net offers a practical, transparent, and empirically enhanced framework for frame-based deepfake detection.
Keywords:
Deepfake detection
XGBoost
Stacking framework
Meta-learning
Frame-based deepfake analysis
Cross-dataset generalization
CNN
Journal
C
IF:
4.1
Papers:
5.0K
Citations:
7.5K

