arrow
返回

DeKeDVer: A deep learning-based multi-type software vulnerability classification framework using vulnerability description and source code

delete2023-11-01
delete6
PRE
AI
董玉坤 封面图
董玉坤 (Yukun Dong) *
Y
Yeer Tang
X
Xiaotong Cheng
Y
Yufei Yang
DOI:10.1016/j.infsof.2023.107290delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Context: Software vulnerabilities have confused software developers for a long time. Vulnerability classification is thus crucial, through which we can know the specific type of vulnerability and then conduct targeted repair. Stack of papers have looked into deep learning-based multi-type vulnerability classification, among which most are based on vulnerability descriptions and some are based on source code. While vulnerability descriptions can sometimes mislead vulnerability classification and source code-based approaches have been rarely explored in multi-type vulnerability classification. Objective: We design DeKeDVer (Vulnerability Descriptions and Key Domain based Vulnerability Classifier) with two objectives: (i) to extract more useful information from vulnerability descriptions; (ii) to better utilize the information source code can reflect. Method: In this work, we propose a multi-type vulnerability classifier which combine vulnerability descriptions and source code together. We process vulnerability descriptions and source code of each project separately. For the vulnerability description of a sample, we preprocess it using a specified way we design based on our observations on numerous descriptions and then select text features. After that, Text Recurrent Convolutional Neural Network (TextRCNN) is applied to learn text information. For source code, we leverage its Code Property Graph (CPG) and extract key domain from it which are then embedded. Acquired feature vectors are then fed into Relational Graph Attention Network (RGAT). Result vectors gained from TextRCNN and RGAT are combined together as the feature vector of the current sample. A Multi-Layer Perceptron (MLP) layer is further added to undertake classification. Results: We conduct our experiments on C/C++ projects from NVD. Experimental results show that our work achieves 84.49% in weighted F1-measure which proves our work to be more effective. Conclusion: Our work utilizes information reflected both from vulnerability descriptions and source code to facilitate vulnerability classification and achieves higher weighted F1-measure than existing vulnerability classification tools.
Keyword:
Multi-type vulnerability classification
Vulnerability description
Source code
Text Recurrent Convolutional Neural Network
Relational graph attention network

期刊

Information and Software Technology 封面图
Information and Software Technology
IF:
4.3
论文数:
3.8K
被引数:
7.7K

机构

C
china university of petroleum
学者数:
4.1W
论文数: 2.7W
被引数: 30
引用论文

引用论文

err分享
err收藏
Analyzing bug fix for automatic bug cause classification自动bug原因分类的bug修复分析
err2020-05-01
err33
PREAI
errNi, Zhen; Li, Bin; Sun, Xiaobing; Chen, Tianhao; Tang, Ben; Shi, Xinchen
err分享
err收藏
err分享
err收藏
LSTM: A Search Space OdysseyLSTM: 搜索空间奥德赛
err2017-10-01
err3.7K
errOAAI
errGreff, Klaus; Srivastava, Rupesh K.; Koutnik, Jan; Steunebrink, Bas R.; Schmidhuber, Juergen
err分享
err收藏
Support-vector networks支持向量网络
err1995-09-01
err0
errOAAI
errCorinna Cortes; Vladimir Vapnik
err分享
err收藏
学者 查看更多内容