arrow
Return

Oversampling method using large language model for malicious JavaScript code detection

delete2026-08-01
delete0
PRE
AI
M
Mamoru Mimura *
D
Danjo, Shinjiro
DOI:10.1016/j.engappai.2026.115821delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Detecting malicious JavaScript code on websites remains a significant issue. Detection methods using machine learning models have been proposed to detect these attacks in real time. In targeted attacks, malware is often customized for a specific organization. To defend against these attacks, machine learning models have to be trained from a small number of past malicious samples. Traditional oversampling techniques applied to JavaScript do not appropriately represent the semantics and syntax of the code. Therefore, the generated samples may not properly represent the features of actual malicious JavaScript code. To address this issue, we propose a method to oversample malicious JavaScript code using CodeLlama, a large language model specialized for code generation. We constructed an imbalanced dataset consisting of 21,744 benign JavaScript code collected by crawling URLs and 8000 publicly available malicious JavaScript code categorized by collection year, and evaluated the accuracy of the proposed method. The recall of our method, which applied oversampling, was about 0.22 higher than existing methods and about 0.21 higher than the baseline without oversampling. Furthermore, we demonstrated that our method improves accuracy by applying it to both classification and feature extraction processes, and that the optimal ratio of original samples to oversampled samples is approximately 1:1 to 1:2.
Keywords:
Oversampling
Large language model
CodeLlama-Instruct
JavaScript

Journal

Engineering Applications of Artificial Intelligence cover
Engineering Applications of Artificial Intelligence
IF:
8
Papers:
5.7K
Citations:
3.5W

Organization

N
National Defense Academy
Scholars:
12
Papers: 6
Citations: 0
Cited Papers

Cited Papers

No cited papers available