MACHINE LEARNING FOR NETWORK INTRUSION DETECTION: A PERFORMANCE COMPARISON OF ACCURACY, FALSE POSITIVES, AND LATENCY
Keywords:
Intrusion Detection; Network Security; Machine Learning; CICIDS2017; Accuracy; False Positive Rate; Latency; Feature Selection; Ensemble MethodsAbstract
The network IDS should provide both a high level of accuracy and low latency to be efficient in practice. Within the framework of this study, we evaluated four machine learning methods on the CICIDS2017 dataset considering their practical performance. They included Random Forest, SVM, KNN, and CNN models. The results show that CNN demonstrated the highest accuracy of 99.2 percent, while providing the FPR of 1.5 percent, but the latency of 115 ms per sample, which is from 2.7 to 1 times more than Random Forest. The accuracy of Random Forest was 98.7 percent, FPR – 1.2 percent and latency – 42 ms per sample. SVM provided 96.4 percent accuracy, 2.8 percent FPR and 67 ms latency. KNN was the worst among the four tested approaches: its accuracy was 94.1 percent, FPR – 4.5 percent and processing time – 89 ms. As a means of improving the models' efficiency, feature selection was conducted, which helped to reduce the number of features by 35 percent and increase the performance of the models in 28 percent without decreasing accuracy. Ensembles also improved zero-day attacks detection in 22 percent comparing to single classifiers. Major limitations of the work are the dataset bias and lack of tests of the methods in real traffic. The results demonstrate that the best balance between performance and cost in the context of enterprise usage is provided by Random Forest with 1.8 to 1 ratio.


