00
Days
00
Hrs
00
Min
00
Sec
Submit Your Paper

Hybrid Ensemble Learning For Malicious URL Detection: A Literature Review of Machine Learning, Deep Learning, and Feature-Fusion Approaches

Authors

Ashwini Sable

Independent Researcher (India)

Vijay More

Independent Researcher (India)

Article Information

DOI: 10.51583/IJLTEMAS.2026.150900024

Subject Category: Machine Learning

Volume/Issue: 15/9 | Page No: 289-307

Publication Timeline

Submitted: 2026-09-11

Accepted: 2026-09-16

Published: 2026-10-01

Abstract

The rapid evolution of phishing, malware distribution, and other web-based attacks has made malicious Uniform Resource Locator (URL) detection an important cybersecurity research problem. Traditional blacklist and rule-based mechanisms provide efficient protection against known malicious addresses but may fail when attackers generate previously unseen URLs, employ URL shortening, manipulate lexical structures, or rapidly change domains. Consequently, machine learning and deep learning approaches have increasingly been investigated for detecting malicious URLs using lexical, structural, host-based, contextual, and semantic information. More recently, hybrid and ensemble learning architectures have emerged as promising alternatives to individual classifiers because they combine complementary representations and learning mechanisms. This structured review synthesizes twenty studies published primarily between 2020 and 2026 across IEEE, Elsevier, Springer, and related peer-reviewed venues. The analysis is organized around representation diversity rather than reported accuracy alone and covers handcrafted lexical and structural features, sparse character n-grams, convolutional and recurrent neural representations, transformer-based contextual models, graph neural networks, and heterogeneous ensemble or fusion architectures. The evidence indicates a clear transition from single-representation classifiers toward multi-branch models that combine complementary statistical, sequential, contextual, and relational signals. However, reported benchmark performance is not directly comparable across studies because datasets, class distributions, preprocessing, temporal splits, and operating thresholds vary substantially. Persistent research challenges include cross-dataset generalization, temporal drift, domain-level leakage, class imbalance, false-positive control, adversarial robustness, explainability, and deployment latency. Based on these gaps, the review proposes a four-branch hybrid ensemble architecture and a formal evaluation plan emphasizing ablation, temporal and cross-dataset validation, adversarial testing, low-false-positive operating points, and computational efficiency.

Keywords

malicious URL detection, phishing detection, ensemble learning, hybrid learning, machine learning, deep learning, BERT, Random Forest, boosting, cybersecurity

Downloads

References

1. M. Sameen, K. Han, and S. O. Hwang, “PhishHaven—An efficient real-time AI phishing URLs detection system,” IEEE Access, vol. 8, pp. 83425–83443, 2020, doi: 10.1109/ACCESS.2020.2991403. [Google Scholar] [Crossref]

2. J. Yuan, G. Chen, S. Tian, and X. Pei, “Malicious URL detection based on a parallel neural joint model,” IEEE Access, vol. 9, pp. 9464–9472, 2021, doi: 10.1109/ACCESS.2021.3049625. [Google Scholar] [Crossref]

3. V. K. Nadar, B. Patel, V. Devmane, and U. Bhave, “Detection of phishing websites using machine learning approach,” Proc. 2nd Global Conf. Advancement in Technology (GCAT), 2021, pp. 1–8, doi: 10.1109/GCAT52182.2021.9587682. [Google Scholar] [Crossref]

4. M. Aljabri et al., “Detecting malicious URLs using machine learning techniques: Review and research directions,” IEEE Access, vol. 10, pp. 121395–121417, 2022, doi: 10.1109/ACCESS.2022.3222307. [Google Scholar] [Crossref]

5. A. Saleem Raja, R. Madhubala, N. Rajesh, L. Shaheetha, and N. Arulkumar, “Survey on malicious URL detection techniques,” in Proc. 6th Int. Conf. Trends in Electronics and Informatics (ICOEI), Tirunelveli, India, Apr. 28–30, 2022, pp. 778–781, doi: 10.1109/ICOEI53556.2022.9777221. [Google Scholar] [Crossref]

6. J. A. Kumar, “Hybrid feature-based machine learning method for phishing URL detection,” Proc. 3rd Int. Conf. Secure Cyber Computing and Communication (ICSCCC), 2023, pp. 222–227, doi: 10.1109/ICSCCC58608.2023.10176901. [Google Scholar] [Crossref]

7. R. Ferdaws and N. E. Majd, “Phishing URL detection using machine learning and deep learning,” Proc. IEEE World AI IoT Congress (AIIoT), 2024, pp. 485–490, doi: 10.1109/AIIoT61789.2024.10579005. [Google Scholar] [Crossref]

8. B. Alaladinni et al., “A hybrid approach for malicious URL detection using ML classifiers and graph neural networks,” Proc. 6th Int. Conf. Data Intelligence and Cognitive Informatics (ICDICI), 2025, doi: 10.1109/ICDICI66477.2025.11135390. [Google Scholar] [Crossref]

9. J. Lee and H. Kwon, “Hybrid ensemble learning for malicious URL detection with BERT and boosting models,” IEEE Access, vol. 14, pp. 62045–62058, 2026, doi: 10.1109/ACCESS.2025.3605302. [Google Scholar] [Crossref]

10. Ü. Özmen and E. O. Yildirim, “DistilBERT-based hybrid architecture for phishing URL detection,” IEEE Access, vol. 14, pp. 71720–71737, 2026, doi: 10.1109/ACCESS.2026.3684855. [Google Scholar] [Crossref]

11. N. Reyes-Dorta, P. Caballero-Gil, and C. Rosa-Remedios, “Detection of malicious URLs using machine learning,” Wireless Networks, vol. 30, pp. 7543–7560, 2024, doi: 10.1007/s11276-024-03700-w. [Google Scholar] [Crossref]

12. R. Liu, Y. Wang, Z. Guo, H. Xu, Z. Qin, W. Ma, and F. Zhang, “TransURL: Improving malicious URL detection with multi-layer Transformer encoding and multi-scale pyramid features,” Computer Networks, 2024, doi: 10.1016/j.comnet.2024.110707. [Google Scholar] [Crossref]

13. R. Liu, Y. Wang, H. Xu, Z. Qin, F. Zhang, Y. Liu, and Z. Cao, “PMANet: Malicious URL detection via post-trained language model guided multi-level feature attention network,” Information Fusion, vol. 113, 102638, 2025, doi: 10.1016/j.inffus.2024.102638. [Google Scholar] [Crossref]

14. N. Q. Do, A. Selamat, H. Fujita, and O. Krejcar, “An integrated model based on deep learning classifiers and pre-trained transformer for phishing URL detection,” Future Generation Computer Systems, vol. 161, pp. 269–285, 2024, doi: 10.1016/j.future.2024.06.031. [Google Scholar] [Crossref]

15. N. Q. Do, A. Selamat, O. Krejcar, and H. Fujita, “Detection of malicious URLs using Temporal Convolutional Network and Multi-Head Self-Attention mechanism,” Applied Soft Computing, vol. 169, 112540, 2025, doi: 10.1016/j.asoc.2024.112540. [Google Scholar] [Crossref]

16. K. Barik, S. Misra, and R. Mohan, “Web-based phishing URL detection model using deep learning optimization techniques,” International Journal of Data Science and Analytics, vol. 20, pp. 4449–4471, 2025, doi: 10.1007/s41060-025-00728-9. [Google Scholar] [Crossref]

17. T. Doshi et al., “PhishHunter-XLD: An ensemble approach integrating machine learning and deep learning for phishing URL classification,” Franklin Open, vol. 12, 100349, 2025, doi: 10.1016/j.fraope.2025.100349. [Google Scholar] [Crossref]

18. S. Lokesh R, “Phishing URL detection using machine learning: A comparative study,” in Proc. 6th Int. Conf. Mobile Computing and Sustainable Informatics (ICMCSI), Goathgaun, Nepal, Jan. 7–8, 2025, pp. 1524–1532, doi: 10.1109/ICMCSI64620.2025.10883082. [Google Scholar] [Crossref]

19. J. D. Duarte, P. Chagas Junior, J. P. J. da Costa, E. J. da Costa, L. P. de Melo, R. R. Nunes, C. V. N. G. Soares, and T. E. da Cunha Silva, “Machine learning for early detection of phishing URLs in parked domains: An approach applied to a financial institution,” IEEE Access, vol. 13, pp. 145736–145753, 2025, doi: 10.1109/ACCESS.2025.3599454. [Google Scholar] [Crossref]

20. Y. Tian, Y. Yu, J. Sun, and Y. Wang, “From past to present: A survey of malicious URL detection techniques, datasets and code repositories,” Computer Science Review, vol. 58, Art. no. 100810, 2025, doi: 10.1016/j.cosrev.2025.100810. [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles

© 2026 IJLTEMAS · RSIS International. All rights reserved. ISSN 2278-2540.