Training Data Matters: Evaluating ML Intrusion Detectors Across Diverse Network Traffic Datasets

Authors

  • Sakunaveeti Vishnu Author
  • S. Ramesh Author

Keywords:

anomaly detection, classification algorithms, cybersecurity, feature selection, random forest

Abstract

The rapid expansion of internet-connected systems has significantly increased the volume and sophistication of cyberattacks, making traditional signature-based Intrusion Detection Systems (IDS) inadequate for identifying novel and zero-day threats. This paper presents a comprehensive study of machine learning approaches for network intrusion detection, examining supervised classification algorithms including Random Forest, Support Vector Machine, Decision Tree, K-Nearest Neighbours, Naïve Bayes, and ensemble-based methods such as XGBoost. The study evaluates these algorithms in the context of widely used benchmark datasets, namely NSL-KDD, CICIDS2017, and UNSW-NB15. IDS development pipeline consisting of data preprocessing, feature selection, model training, and evaluation and highlights practical challenges such as class imbalance, high dimensionality and the evolving nature of attack patterns is discussed. Findings from the earlier studies indicate hybrid models combining feature selection techniques with classifiers such as Random Forest and XGBoost consistently achieve the highest detection accuracy, often exceeding 99% on benchmark datasets. The paper concludes by outlining research directions for building more robust, real-time, and adversarially resilient ML-based IDS suitable for deployment in enterprise, cloud, and IoT network environments.

Published

2026-08-23