arrow
Return

Malware detection using pre-trained transformer encoder with byte sequences

delete2025-10-13
delete0
PRE
AI
E
Eun‐Jin Kim
Y
Yun-Kyung Lee
S
Sang-Min Lee
A
Ah Reum Kang
M
M.-K. Kim
Y
Young-Seob Jeong *
DOI:10.1371/journal.pone.0332307delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Ordinary users encounter various documents on the network every day, such as news articles, emails, and messages, and most are vulnerable to malicious attacks. Malicious attack methods continue to evolve, making neural network-based malware detection increasingly appealing to both academia and industry. Recent studies have leveraged byte sequences within files to detect malicious activities, primarily using convolutional neural networks to capture local patterns in the byte sequences. Meanwhile, in natural language processing, Transformer-based language models have demonstrated superior performance across various tasks and have been applied to other domains, such as image analysis and speech recognition. In this paper, we introduce a novel Transformer-based language model for malware detection that processes byte sequences as input. We propose two new pre-training strategies: real-or-fake prediction and same-sequence prediction. Including conventional pre-training strategies such as masked language modeling and next-sentence prediction, we explore all possible combinations of these approaches. By compiling existing byte sequences for malware detection, we construct a benchmark consisting of three file types (PDF, HWP, and MS Office) for pre-training and fine-tuning. Our empirical results demonstrate that our language model outperforms convolutional neural networks in the malware detection task, achieving a macro F1 score improvement of approximately 2.7%p similar to 11.1%p. We believe our language model will serve as a foundation model for malware detection services, and will extend our research to develop a more powerful encoder-based model that can process longer byte sequences.

Journal

PLoS One cover
PLoS One
IF:
2.6
Papers:
2.6W
Citations:
81.6W

Organization

P
Pai Chai University
Scholars:
337
Papers: 339
Citations: 218
C
chungbuk national university
Scholars:
1.7K
Papers: 729
Citations: 0
researcher View more organizations
Cited Papers

Cited Papers

Long Short-Term Memory
err1997-11-01
err0
PREAI
errSepp Hochreiter; Jürgen Schmidhuber
errShare
errSave
A novel graph neural network framework with self-evolutionary mechanism: Application to train-bridge coupled systems
err2024-11-01
err6
PREAI
errZhang, Peng; Zhao, Han; Shao, Zhanjun; Xie, Xiaonan; Hu, Huifang; Zeng, Yingying; Xiang, Ping
errShare
errSave
Empirical study on character level neural network classifier for Chinese text
err2019-04-01
err0
PREAI
errTonglee Chung; Bin Xu; Yongbin Liu; Chunping Ouyang; Siliang Li; Lingyun Luo
errShare
errSave
Enhanced multi-scenario running safety assessment of railway bridges based on graph neural networks with self-evolutionary capability
err2024-11-01
err11
PREAI
errZhang, Peng; Zhao, Han; Shao, Zhanjun; Xie, Xiaonan; Hu, Huifang; Zeng, Yingying; Jiang, Lizhong; Xiang, Ping
errShare
errSave
Static Malware Detection Using Stacked BiLSTM and GPT-2
err2022-01-01
err26
errOAAI
errDemirci, Deniz; Sahin, Nazenin; Sirlancis, Melih; Acarturk, Cengiz
errShare
errSave
High-speed railway seismic response prediction using CNN-LSTM hybrid neural network
err2024-03-11
err43
PREAI
errZhang, Xuebing; Xie, Xiaonan; Tang, Shenghua; Zhao, Han; Shi, Xueji; Wang, Li; Wu, Han; Xiang, Ping
errShare
errSave
researcher View more