Design of Intelligent Feature Selection Technique for Phishing Detection

Phishing attacks lead to significant threats to individuals and organizations by gaining unauthorized access. The attackers redirect the users to fake websites and steal their credentials and other confidential data. Various techniques are employed to detect phishing using machine learning algorith...

Full description

Saved in:
Bibliographic Details
Main Authors: Sharvari Sagar Patil, Narendra M. Shekokar, Sridhar Chandramohan Iyer
Format: Article
Language:English
Published: IIUM Press, International Islamic University Malaysia 2025-01-01
Series:International Islamic University Malaysia Engineering Journal
Subjects:
Online Access:https://journals.iium.edu.my/ejournal/index.php/iiumej/article/view/3337
Tags: Add Tag
No Tags, Be the first to tag this record!
Description
Summary:Phishing attacks lead to significant threats to individuals and organizations by gaining unauthorized access. The attackers redirect the users to fake websites and steal their credentials and other confidential data. Various techniques are employed to detect phishing using machine learning algorithms or static detection techniques that use blacklisting of web URLs. The attackers tend to change their approach to launch an attack, making it difficult for traditional phishing detection techniques to safeguard the user. The performance of conventional detection methods relies on exhaustive data and features selected for classification. Features selected for designing detection systems majorly contribute to the performance of the detection system. Phishing detection techniques rely mainly on static features that are selected based on traditional feature selection or ranking techniques. This paper proposes an innovative approach to phishing detection by designing a feature selection technique using reinforcement learning. A novel reinforcement learning agent is designed that uses a dynamic, adaptive, and data-driven approach to improve classifier performance in phishing detection. The technique is designed to select the features using the RL agent dynamically. We have evaluated our technique using the real-world phishing dataset and compared its performance with the existing techniques. Based on the evaluation, our proposed methodology of dynamic feature selection gives the best accuracy of 99.07 % with the random forest classifier model. Our work contributes to advancing phishing detection methodology by developing a dynamic feature selection technique. ABSTRAK: Serangan pancing data membawa ancaman besar kepada individu dan organisasi dengan mendapatkan akses tanpa kebenaran. Penyerang akan mengalihkan pengguna ke laman web palsu dan mencuri maklumat log masuk serta data sulit yang lain. Pelbagai teknik digunakan bagi mengesan pancing data menggunakan algoritma pembelajaran mesin atau teknik pengesanan statik yang menggunakan URL laman web yang disenarai hitam. Penyerang cenderung mengubah pendekatan mereka untuk melancarkan serangan, menjadikan teknik pengesanan pancing data tradisional sukar bagi melindungi pengguna. Prestasi kaedah pengesanan konvensional bergantung kepada data menyeluruh dan ciri-ciri yang dipilih untuk pengelasan. Teknik pengesanan pancing data kebanyakannya bergantung pada ciri-ciri statik yang dipilih berdasarkan kaedah pemilihan atau penarafan ciri tradisional. Kajian ini mencadangkan pendekatan inovatif bagi pengesanan pancing data dengan mereka bentuk teknik pemilihan ciri menggunakan pembelajaran peneguhan. Ejen pembelajaran peneguhan baru, direka menggunakan pendekatan yang dinamik, adaptif, dan berasaskan data bagi memperbaiki prestasi pengelas dalam pengesanan pancing data. Teknik ini direka untuk memilih ciri-ciri secara dinamik menggunakan ejen RL. Teknik ini dinilai menggunakan dataset pancing data sebenar dan dibanding prestasinya dengan teknik sedia ada. Berdasarkan penilaian, metodologi pemilihan ciri dinamik ini memberikan ketepatan terbaik sebanyak 99.07% dengan model pengelasan rawak. Kerja ini merupakan sumbangan kepada kemajuan metodologi pengesanan pancing data dengan membangunkan teknik pemilihan ciri dinamik.
ISSN:1511-788X
2289-7860