Return
Detecting Deepfake Audio Using Spectrogram-Based Machine Learning Approaches
DOI:10.1109/ACCESS.2025.3602531.png)
Abstract
En 中文
The increasing ability of deep learning models to produce realistic-sounding synthetic speech poses serious problems for privacy, public trust, and digital security. To counter this danger, we offer a methodology for identifying deepfake audio that is based on machine learning. We use three ensemble learning models (Random Forest, Gradient Boosting, and XGBoost) for classification after transforming speech samples into mel-spectrograms to extract time-frequency information. These models, which were trained on the Deep Voice dataset, which included a variety of actual and synthetic samples, were assessed using common metrics such as F1-score, accuracy, precision, and recall. At 99.32% accuracy, XGBoost performed better than the others. These findings show the promise of lightweight, interpretable machine learning methods for identifying fake audio and protecting media authentication, digital forensics, and cybersecurity applications.
Keywords:
Auto encoders
music generation
generative AI
deep learning
Journal
IF:
3.6
Papers:
9.8W
Citations:
29.4W

