arrow
返回

Malicious code classification based on opcode sequences and textCNN network

delete2022-06-01
delete20
PRE
AI
Q
Qianhui Wang
Q
Quan Qian *
DOI:10.1016/j.jisa.2022.103151delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
A malicious code classification problem is essential for the network security. Malicious code is the most common means of network attack, which threatens user information and property security. An effective malicious code classification method can improve the efficiency of malicious code detection and the ability to discover new malicious code families. This study proposes a new malicious code classification method to analyze, classify, and detect malicious code. The semantic features of opcode sequences are extracted effectively by introducing the concept of word vectors. Furthermore, the extracted sequence is regarded as a text sentence and then introduced to a text convolutional neural network (textCNN) to identify malicious code families. The experimental results revealed that the model has more than 98% accuracy (with macro-average precision above 98.65% and macro-average recall approximately 98.66%) on the Microsoft Malware Challenge dataset conducted in 2015. Meanwhile, the accuracy of the model on the SOREL-20M dataset is 91.93%. Mostly call instructions are used to call the API, library functions, and other user-defined functions through which the behavior of malicious code is generally realized. Thus, selecting the block that contains call instructions as the key block will reduce the model training speed. After selecting the key block, on average, the number of opcodes on Microsoft Malware Challenge dataset is reduced by 39.07% and has a 98.18% accuracy rate, which is slightly lower than the result obtained by using all opcodes. The number of opcodes on the SOREL20M dataset is reduced by 30.49% on average, and the accuracy is increased to 93.46%. Experimental results show that the proposed algorithm works well and outperforms the results obtained by using byte n-gram representation.
Keyword:
Malicious code classification
Opcode sequences
Word embedding
Text convolutional neural network

期刊

Journal of Information Security and Applications 封面图
Journal of Information Security and Applications
IF:
3.7
论文数:
2.0K
被引数:
4.9K

机构

S
shanghai university
学者数:
3.9W
论文数: 2.7W
被引数: 52
引用论文

引用论文

Effects of titanium and tantalum adhesion layers on the properties of sol-gel derived SrBi2Ta2O9 thin films
err2002-08-01
err0
PREAI
errChing-Chich Leu; Hung-Tao Lin; Chen-Ti Hu; Chao-Hsin Chien; Ming-Jui Yang; Ming-Che Yang; Tiao-Yuan Huang
err分享
err收藏
Preparation and Characterization of Nitrosylruthenium(III) Complexes Containing Diethylenetriamine
err2006-07-19
err0
PREAI
errRisa Asanuma; Hiroshi Tomizawa; Akio Urushiyama; Eiichi Miki; Kunihiko Mizumachi; Tatsujiro Ishimori
err分享
err收藏
SPAM: Signal Processing to Analyze Malware
err2016-03-01
err41
errOAAI
errNataraj, Lakshmanan; Manjunath, B. S.
err分享
err收藏
High magnetic moments on binary yttrium-alkali superatoms
err2013-09-01
err0
PREAI
errHenry González-Ramírez; J. Ulises Reveles; Z. Gómez-Sandoval
err分享
err收藏
没有更多内容