arrow
Return

The MalSource Dataset: Quantifying Complexity and Code Reuse in Malware Development

delete2019-12-01
delete40
delete
OA
AI
A
Alejandro Calleja *
J
Juan Tapiador
J
Juan Caballero
DOI:10.1109/TIFS.2018.2885512delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
During the last decades, the problem of malicious and unwanted software (malware) has surged in numbers and sophistication. Malware plays a key role in most of today's cyber-attacks and has consolidated as a commodity in the underground economy. In this paper, we analyze the evolution of malware from 1975 to date from a software engineering perspective. We analyze the source code of 456 samples from 428 unique families and obtain measures of their size, code quality, and estimates of the development casts (effort, time, and number of people). Our results suggest an exponential increment of nearly one order of magnitude per decade in aspects such as size and estimated effort, with code quality metrics similar to those of benign software. We also study the extent to which code reuse is present in our dataset. We detect a significant number of code clones across malware families and report which features and functionalities are more commonly shared. Overall, our results support claims about the increasing complexity of malware and its production progressively becoming an industry.
Keywords:
Computer crime
computer languages
open source software
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Information Forensics and Security cover
IEEE Transactions on Information Forensics and Security
IF:
8
Papers:
5.2K
Citations:
2.3W

Organization

I
imdea software institute
Scholars:
63
Papers: 45
Citations: 0
U
Universidad Carlos III de Madrid
Scholars:
5.5K
Papers: 5.7K
Citations: 4.5K