arrow
返回

Impact of benign sample size on binary classification accuracy

delete2023-01-01
delete6
delete
OA
AI
M
Mamoru Mimura *
DOI:10.1016/j.eswa.2022.118630delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Recently, there has been a significant increase in malware attacks and malicious traffic. Consequently, several machine learning-based detection models have been developed to detect them. However, the detection accuracy of these models is currently evaluated using different methodologies and datasets, with some studies overstating high detection rates. The lack of a common testing approach coupled with the limited datasets used for the experiments make it challenging to compare the performances of these models to identify those that provide superior detection accuracy. A few studies have focused on benign samples and their effects on detection accuracy. The datasets used in the experiments generally consist of benign and malicious samples; hence, binary classification is used in the machine learning models. In the binary classification task, the size of a benign sample affects the classification accuracy of malicious samples, that is, it can either improve or degrade detection accuracy. In this study, we propose a novel metric for evaluating accuracy degradation by increasing benign sample size. We mainly used the FFRI dataset, which consists of 11,243 malware samples and 250,000 benign samples, and evaluated the classification accuracy with extracted strings from the malware. In addition, we obtained other malware samples that we used as supplementary to the main dataset. We increased the number of benign samples for testing by tenfold, while maintaining the malicious sample and benign training sample sizes, which resulted in a decrease of 0.293 in the F1 score. Furthermore, we confirmed that using a sufficiently sized benign training sample set mitigates accuracy degradation. Our metric can be beneficial for evaluating the benign sample size needed in binary classification and comparing accuracy.
Keyword:
Malware
Machine learning
Binary classification
Benign sample
Random forest
Support vector machine
XGBoost
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Expert Systems with Applications 封面图
Expert Systems with Applications
IF:
7.5
论文数:
3.0W
被引数:
10.2W

机构

N
national defense academy - japan
学者数:
595
论文数: 620
被引数: 0
引用论文

引用论文

The arms race: Adversarial search defeats entropy used to detect malware
err2019-03-01
err30
errOAAI
errMenendez, Hector D.; Bhattacharya, Sukriti; Clark, David; Barr, Earl T.
err分享
err收藏
err分享
err收藏
Addressing the class imbalance problem in Twitter spam detection using ensemble learning
err2017-08-01
err89
PREAI
errLiu, Shigang; Wang, Yu; Zhang, Jun; Chen, Chao; Xiang, Yang
err分享
err收藏
Expression of IDO1 and PD-L2 in Patients with Benign Lymphadenopathies and Association with Autoimmune Diseases
err2023-01-27
err0
errOAAI
errMaysaa Abdulla; Christer Sundström; Cecilia Lindskog; Peter Hollander
err分享
err收藏
Allosteric Regulation in Phosphofructokinase from the Extreme Thermophile Thermus thermophilus
err2013-12-27
err0
errOAAI
errMaria S. McGresham; Michelle Lovingshimer; Gregory D. Reinhart
err分享
err收藏
A framework for metamorphic malware analysis and real-time detection变形恶意软件分析和实时检测框架
err2015-02-01
err48
PREAI
errAlam, Shahid; Horspool, R. Nigel; Traore, Issa; Sogukpinar, Ibrahim
err分享
err收藏
学者 查看更多内容