Return
Oversampling method using large language model for malicious JavaScript code detection
DOI:10.1016/j.engappai.2026.115821.png)
Abstract
En 中文
Detecting malicious JavaScript code on websites remains a significant issue. Detection methods using machine learning models have been proposed to detect these attacks in real time. In targeted attacks, malware is often customized for a specific organization. To defend against these attacks, machine learning models have to be trained from a small number of past malicious samples. Traditional oversampling techniques applied to JavaScript do not appropriately represent the semantics and syntax of the code. Therefore, the generated samples may not properly represent the features of actual malicious JavaScript code. To address this issue, we propose a method to oversample malicious JavaScript code using CodeLlama, a large language model specialized for code generation. We constructed an imbalanced dataset consisting of 21,744 benign JavaScript code collected by crawling URLs and 8000 publicly available malicious JavaScript code categorized by collection year, and evaluated the accuracy of the proposed method. The recall of our method, which applied oversampling, was about 0.22 higher than existing methods and about 0.21 higher than the baseline without oversampling. Furthermore, we demonstrated that our method improves accuracy by applying it to both classification and feature extraction processes, and that the optimal ratio of original samples to oversampled samples is approximately 1:1 to 1:2.
Keywords:
Oversampling
Large language model
CodeLlama-Instruct
JavaScript
Journal
IF:
8
Papers:
5.7K
Citations:
3.5W
Organization
Cited Papers
No cited papers available

