arrow
返回

Machine learning-based detection method for malicious PDF files: A temporal classification approach

delete2026-01-21
delete0
PRE
AI
D
Doo-Seop Choi
T
Taeguen Kim
B
BooJoong Kang
E
Eul Gyu Im *
DOI:10.1016/j.asoc.2025.114461delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Cybercriminals increasingly exploit non-executable files that can bypass antivirus software detection and are often opened by users without suspicion. In particular, PDF files have become a primary attack vector for adversaries due to their platform-independent nature and ability to preserve document components across different systems. Malicious PDF files continuously evolve to avoid detection, and traditional detection methods, which rely primar ily on static features from older PDF datasets, show limitations in identifying evolving malicious PDF files. This paper identifies temporal evolution in feature distributions and proposes a novel framework to detect malicious PDF files by introducing temporal classification and addressing the evolved characteristics of recent threats. Through in-depth statistical analysis, we revealed that recent malicious PDF files closely mimic the structural characteristics of legitimate files, exhibiting an 11-fold increase in graphic components and a 21-fold increase in hyperlinks compared to older samples. This finding indicates a significant shift in attack methodologies from traditional script injection to social engineering techniques. To address this challenge, we enhanced the basic fea ture set, comprising 31 structural and metadata-based features initially defined in the CIC-Evasive-PDFMal2022 dataset, by integrating 12 newly identified features, resulting in an enhanced set of 43 features. Experimental results demonstrate that our framework with the enhanced feature set achieves 97.80 % detection accuracy us ing the random forest algorithm, representing a 4.12 % improvement over the basic feature set. The framework maintains balanced performance across all metrics with a recall of 0.96, a precision of 0.98, an F1-score of 0.97, and an AUC of 0.99. Additionally, the framework reduced the false positive rate (FPR) from 2.84 % to 1.12 %, a 1.72 percentage points reduction, which is critical for practical deployment in real-world security environments. The proposed enhanced feature set provides an effective approach for strengthening real-world detection systems, including email attachment scanners and antivirus engines, against evolving PDF-based attacks.
Keyword:
PDF malware
Non-executable malware
Malware detection
Machine learning
Analysis of temporal feature evolution

期刊

Applied Soft Computing 封面图
Applied Soft Computing
IF:
6.6
论文数:
1.4W
被引数:
4.8W

机构

U
University of Southampton
学者数:
735
论文数: 422
被引数: 1
K
Korea University
学者数:
3.6W
论文数: 3.8W
被引数: 4.4W
H
hanyang university
学者数:
2.9W
论文数: 2.7W
被引数: 36
学者 查看更多机构
引用论文

引用论文

Backpropagation Applied to Handwritten Zip Code Recognition反向传播在手写邮政编码识别中的应用
err1989-12-01
err0
PREAI
errY. LeCun; B. Boser; J. S. Denker; D. Henderson; R. E. Howard; W. Hubbard; L. D. Jackel
err分享
err收藏
Glyph: Efficient ML-Based Detection of Heap Spraying Attacks
err2021-01-01
err3
errOAAI
errPierazzi, Fabio; Cristalli, Stefano; Bruschi, Danilo; Colajanni, Michele; Marchetti, Mirco; Lanzi, Andrea
err分享
err收藏
The Whale Optimization Algorithm鲸鱼优化算法
err2016-05-01
err9.5K
PREAI
errMirjalili, Seyedali; Lewis, Andrew
err分享
err收藏
A feature-vector generative adversarial network for evading PDF malware classifiers
err2020-06-01
err0
PREAI
errYuanzhang Li; Yaxiao Wang; Ye Wang; Lishan Ke; Yu-an Tan
err分享
err收藏
Hidden Markov models for malware classification
err2014-05-23
err0
errOAAI
errChinmayee Annachhatre; Thomas H. Austin; Mark Stamp
err分享
err收藏
学者 查看更多内容