arrow
返回

Software Vulnerability Discovery via Learning Multi-Domain Knowledge Bases

delete2021-09-01
delete80
PRE
AI
G
Guanjun Lin
J
Jun Zhang *
罗玮 封面图
罗玮 (Wei Luo)
L
Lei Pan
O
Olivier De Vel
P
Paul Montague
向
向阳 (Yang Xiang)
DOI:10.1109/TDSC.2019.2954088delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Machine learning (ML) has great potential in automated code vulnerability discovery. However, automated discovery application driven by off-the-shelf machine learning tools often performs poorly due to the shortage of high-quality training data. The scarceness of vulnerability data is almost always a problem for any developing software project during its early stages, which is referred to as the cold-start problem. This article proposes a framework that utilizes transferable knowledge from pre-existing data sources. In order to improve the detection performance, multiple vulnerability-relevant data sources were selected to form a broader base for learning transferable knowledge. The selected vulnerability-relevant data sources are cross-domain, including historical vulnerability data from different software projects and data from the Software Assurance Reference Database (SARD) consisting of synthetic vulnerability examples and proof-of-concept test cases. To extract the information applicable in vulnerability detection from the cross-domain data sets, we designed a deep-learning-based framework with Long-short Term Memory (LSTM) cells. Our framework combines the heterogeneous data sources to learn unified representations of the patterns of the vulnerable source codes. Empirical studies showed that the unified representations generated by the proposed deep learning networks are feasible and effective, and are transferable for real-world vulnerability detection. Our experiments demonstrated that by leveraging two heterogeneous data sources, the performance of our vulnerability detection outperformed the static vulnerability discovery tool Flawfinder. The findings of this article may stimulate further research in ML-based vulnerability detection using heterogeneous data sources.
Keyword:
Software
Feature extraction
Deep learning
Feeds
Task analysis
Neural networks
Data mining
Vulnerability discovery
representation learning
deep learning
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Dependable and Secure Computing 封面图
IEEE Transactions on Dependable and Secure Computing
IF:
7.5
论文数:
2.5K
被引数:
9.6K

机构

D
defence science & technology
学者数:
896
论文数: 938
被引数: 0
S
Swinburne University of Technology
学者数:
9.3K
论文数: 1.2W
被引数: 2.0W
D
Deakin University
学者数:
2.0W
论文数: 2.1W
被引数: 2.8W
学者 查看更多机构
引用论文

引用论文

Supercurrent in a Double Quantum Dot
err2018-12-20
err0
errOAAI
errJ. C. Estrada Saldaña; A. Vekris; G. Steffensen; R. Žitko; P. Krogstrup; J. Paaske; K. Grove-Rasmussen; J. Nygård
err分享
err收藏
Data-Driven Cybersecurity Incident Prediction: A Survey
err2019-01-01
err210
PREAI
errSun, Nan; Zhang, Jun; Rimba, Paul; Gao, Shang; Zhang, Leo Yu; Xiang, Yang
err分享
err收藏
A Sword with Two Edges: Propagation Studies on Both Positive and Negative Information in Online Social Networks
err2015-03-01
err159
PREAI
errWen, Sheng; Haghighi, Mohammad Sayad; Chen, Chao; Xiang, Yang; Zhou, Wanlei; Jia, Weijia
err分享
err收藏
err分享
err收藏
err分享
err收藏
学者 查看更多内容