Return
A Filter-Based Feature Selection Framework to Detect Phishing URLs Using Stacking Ensemble Machine Learning
DOI:10.32604/cmes.2025.070311.png)
Abstract
En 中文
Today, phishing is an online attack designed to obtain sensitive information such as credit card and bank account numbers, passwords, and usernames. We can find several anti-phishing solutions, such as heuristic detection, virtual similarity detection, black and white lists, and machine learning (ML). However, phishing attempts remain a problem, and establishing an effective anti-phishing strategy is a work in progress. Furthermore, while most antiphishing solutions achieve the highest levels of accuracy on a given dataset, their methods suffer from an increased number of false positives. These methods are ineffective against zero-hour attacks. Phishing sites with a high False Positive Rate (FPR) are considered genuine because they can cause people to lose a lot of money by visiting them. Feature selection is critical when developing phishing detection strategies. Good feature selection helps improve accuracy; however, duplicate features can also increase noise in the dataset and reduce the accuracy of the algorithm. Therefore, a combination of filter-based feature selection methods is proposed to detect phishing attacks, including constant extraction, and Analysis of Variance (ANOVA) testing. The technique has been tested with different Machine Learning ensemble models, stacking and majority voting to gain A low false positive rate is achieved. Stacked ensemble classifiers (gradient boosting, random forest, support vector machine) achieve 1.31% FPR and 98.17% accuracy on Dataset 1, 2.81% FPR and Dataset 3 shows 2.81% FPR and 97.61% accuracy, while Dataset 2 shows 3.47% FPR and 96.47% accuracy.
Keywords:
Phishing detection
feature selection
stacking ensemble
machine learning
phishing URL
Journal
C
IF:
2.5
Papers:
354
Citations:
0
Organization
Cited Papers
RSTHFS: A Rough Set Theory-Based Hybrid Feature Selection Method for Phishing Website Classification
IEEE Access
IF0
Exploring the Influence of Direct and Indirect Factors on Information Security Policy Compliance: A Systematic Literature Review
IEEE ACCESS
IF3.6
An ensemble classification method based on machine learning models for malicious Uniform Resource Locators (URL)
PLOS ONE
IF0
Enhancing Phishing Detection: A Machine Learning Approach With Feature Selection and Deep Learning Models
IEEE ACCESS
IF3.6

