Machine learning model for detecting cyberattacks in HTTP traffic in web environments
Main Article Content
Abstract
Introduction. The accelerated growth of web applications has increased HTTP traffic and has made this protocol one of the main targets for those seeking to compromise information security. Objective. To create a technological solution based on a machine learning model capable of detecting cyberattacks in HTTP traffic. Methodology. This was an applied study with a mixed approach. The HTTP DATASET CSIC 2010 was used; eight quantitative features were extracted, and Random Forest, SVM, and XGBoost were evaluated through accuracy, precision, recall, and F1-score, using grid search and five-fold stratified cross-validation. Results. XGBoost achieved the best overall performance, with 91.1% accuracy and an F1-score of 0.888, followed by Random Forest, with 90.7% accuracy and an F1-score of 0.884. SVM achieved a recall of 0.627 and failed to detect many attacks. Conclusion. Random Forest and XGBoost demonstrated the ability to identify anomalous patterns without relying on predefined signatures and constitute safer alternatives than SVM for a future production detection system. The findings support the viability of the proposed machine learning approach for strengthening cybersecurity in web environments. Future validation should use institutional HTTP traffic from the Technical University of Manabí and assess model performance under real operating conditions. General area of study: Engineering and technology. Specific area of study: Cybersecurity and machine learning. Type of study: Original article.
Downloads
Article Details
References
Chen, T., & Guestrin, C. (2016). XGBoost: A Scalable Tree Boosting System. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 785–794. https://doi.org/10.1145/2939672.2939785
Cortes, C., & Vapnik, V. (1995). Support-vector networks. Machine Learning, 20(3), 273–297. https://doi.org/10.1007/bf00994018
Ferriyan, A., Thamrin, A. H., Takeda, K., & Murai, J. (2022). Encrypted malicious traffic detection based on Word2Vec. Electronics, 11(5), 679. https://doi.org/10.3390/electronics11050679
Gniewkowski, M., Maciejewski, H., Surmacz, T. R., & Walentynowicz, W. (2021). HTTP2vec: Embedding of HTTP requests for detection of anomalous traffic. arXiv . https://arxiv.org/abs/2108.01763
Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep Learning. MIT Press. https://books.google.com.ec/books?id=omivDQAAQBAJ
OWASP Foundation. (2021). OWASP Top 10 Web Application Security Risks. OWASP. https://owasp.org/www-project-top-ten/
Python Software Foundation. (2024). Python 3.12 documentation. Python. https://docs.python.org/3.12/
Ring, M., Wunderlich, S., Grudl, D., Landes, D., & Hotho, A. (2019). A survey of network-based intrusion detection data sets. Computer Security, 86, 147-167. https://doi.org/10.1016/j.cose.2019.06.005
Scarfone, K. A., & Mell, P. M. (2007). Guide to intrusion detection and prevention systems (IDPS). National Institute of Standards and Technology. https://csrc.nist.gov/pubs/sp/800/94/final
Scikit-learn Developers. (2024). Ensembles: gradient boosting, random forests, bagging, voting, stacking. User guide – Scikit learn. https://scikit-learn.org/stable/modules/ensemble.html
Sommer, R., & Paxson, V. (2010). Outside the closed world: on using machine learning for network intrusion detection [Proceedings of the IEEE Symposium on Security and Privacy, 305-316]. https://doi.org/10.1109/SP.2010.25
Stallings, W. (2017). Network Security Essentials: applications and Standards (6th ed). Pearson. https://www.pearson.com/en-us/subject-catalog/p/network-security-essentials-applications-and-standards/P200000003333/9780134527338
Tanenbaum, A. S. & Wetherall, D. J. (2012). Redes de computadoras 5e. Pearson Education. https://www.pearsonenespanol.com/mexico/educacion-superior/tanenbaum_index
Tavallaee, M., Bagheri, E., Lu, W., & Ghorbani, A. A. (2009). A detailed analysis of the KDD CUP 99 data set [Proceedings of the IEEE Symposium on Computational Intelligence for Security and Defense Applications, 1-6]. https://doi.org/10.1109/CISDA.2009.5356528
Torrano-Giménez, C., Pérez-Villegas, A., & Álvarez Marañón, G. (2013). HTTP dataset CSIC 2010. Information Security Institute, Spanish National Research Council (CSIC). https://petescully.co.uk/wp-content/uploads/2018/04/http_dataset_csic_2010.pdf
Yu, Y., Yan, H., Guan, H., & Zhou, H. (2018). DeepHTTP: Semantics-structure model with attention for anomalous HTTP traffic detection and pattern mining. arXiv. https://doi.org/10.48550/arXiv.1810.12751