arrow
返回

Generating Synthetic Malware Samples Using Generative AI

delete2025-01-01
delete0
delete
OA
AI
T
Tiffany Bao
K
Kylie Trousil
Q
Quang Duy Tran
F
Fabio Di Troia *
Y
Younghee Park *
DOI:10.1109/ACCESS.2025.3556704delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Malware attacks have a significant negative impact on organizations of varied scales in the field of cybersecurity. Recently, malware researchers have increasingly turned to machine learning techniques to combat sophisticated obfuscation methods used in malware. However, collecting a diverse set of malware samples with various obfuscation techniques is challenging and often takes years, especially for newly developed malware. This issue is further compounded by a well-known limitation of machine learning models: their poor performance when training data is scarce. In this paper, we propose a new system for generating synthetic malware samples to augment imbalanced malware dataset. Our approach decomposes malware binary samples into mnemonic opcode sequences, leveraging natural language processing to extract contextual meaning behind malware opcode features to aid the learning of generative AI (GenAI) employed in this paper, Generative Adversarial Networks (GAN), Wasserstein Generative Adversarial Networks with Gradient Penalty (WGAN-GP), and a modified Diffusion model. The experiment results show that augmenting training data with Diffusion-based synthetic data significantly improves classification performance for minor classes by up to 60% on average. This enhancement ultimately leads to an overall malware classification performance of 96%, an 8% improvement. These findings demonstrate the high quality and fidelity of the synthetic data, its robustness, and its potential applications in malware analysis. Specifically, synthetic malware data proves effective in improving the classification of minor malware classes and detection rates, even though the size of known malware data is significantly small.
Keyword:
Computer viruses
Hidden Markov models
Training
Natural language processing
Generative adversarial networks
Feature extraction
Vectors
Data models
Synthetic data
Robustness
Diffusion
GAN
generative AI
malware
natural language processing
machine learning
imbalanced datasets
data augmentation

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

U
University of Wisconsin La Crosse
学者数:
127
论文数: 83
被引数: 325
California State University System 封面图
California State University System
学者数:
2.8W
论文数: 2.4W
被引数: 457
B
boston university
学者数:
3.8W
论文数: 3.2W
被引数: 67
University of Wisconsin System 封面图
University of Wisconsin System
学者数:
6.7W
论文数: 5.8W
被引数: 382
学者 查看更多机构
引用论文

引用论文

暂无论文信息