arrow
返回

Interpretable Machine Learning Models for Malicious Domains Detection Using Explainable Artificial Intelligence (XAI)

delete2022-06-16
delete27
delete
OA
AI
N
Nida Aslam *
I
Irfan Ullah Khan
S
Samiha Mirza
A
Alanoud AlOwayed
F
Fatima M. Anis
R
Reef M. Aljuaid
R
Reham Baageel
DOI:10.3390/su14127375delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
With the expansion of the internet, a major threat has emerged involving the spread of malicious domains intended by attackers to perform illegal activities aiming to target governments, violating privacy of organizations, and even manipulating everyday users. Therefore, detecting these harmful domains is necessary to combat the growing network attacks. Machine Learning (ML) models have shown significant outcomes towards the detection of malicious domains. However, the black box nature of the complex ML models obstructs their wide-ranging acceptance in some of the fields. The emergence of Explainable Artificial Intelligence (XAI) has successfully incorporated the interpretability and explicability in the complex models. Furthermore, the post hoc XAI model has enabled the interpretability without affecting the performance of the models. This study aimed to propose an Explainable Artificial Intelligence (XAI) model to detect malicious domains on a recent dataset containing 45,000 samples of malicious and non-malicious domains. In the current study, initially several interpretable ML models, such as Decision Tree (DT) and Naive Bayes (NB), and black box ensemble models, such as Random Forest (RF), Extreme Gradient Boosting (XGB), AdaBoost (AB), and Cat Boost (CB) algorithms, were implemented and found that XGB outperformed the other classifiers. Furthermore, the post hoc XAI global surrogate model (Shapley additive explanations) and local surrogate LIME were used to generate the explanation of the XGB prediction. Two sets of experiments were performed; initially the model was executed using a preprocessed dataset and later with selected features using the Sequential Forward Feature selection algorithm. The results demonstrate that ML algorithms were able to distinguish benign and malicious domains with overall accuracy ranging from 0.8479 to 0.9856. The ensemble classifier XGB achieved the highest result, with an AUC and accuracy of 0.9991 and 0.9856, respectively, before the feature selection algorithm, while there was an AUC of 0.999 and accuracy of 0.9818 after the feature selection algorithm. The proposed model outperformed the benchmark study.
Keyword:
network security
malicious domains
machine learning
ensemble models
explainable artificial intelligence

期刊

Sustainability 封面图
Sustainability
IF:
3.3
论文数:
10.6W
被引数:
28.4W

机构

I
Imam Abdulrahman Bin Faisal University
学者数:
6.3K
论文数: 4.5K
被引数: 5.4K
引用论文

引用论文

O-GlcNAc Modification of the Extracellular Domain of Notch Receptors
err2010-01-01
err0
PREAI
errYuta Sakaidani; Koichi Furukawa; Tetsuya Okajima
err分享
err收藏
Pragmatics
err
IF0
err2007-01-01
err0
PREAI
err
err分享
err收藏
Application of the Social Interaction Model
err1997-04-01
err0
PREAI
errCyndie Koning; Kathy Manyk; Joyce Magill-Evans; Anne Cameron-Sadava
err分享
err收藏
Classification and Explanation for Intrusion Detection System Based on Ensemble Trees and SHAP Method
errSENSORS
IF3.5
err2022-02-03
err75
errOAAI
errLe, Thi-Thu-Huong; Kim, Haeyoung; Kang, Hyoeun; Kim, Howon
err分享
err收藏
学者 查看更多内容