A HYBRID CNN-LSTM DEEP LEARNING MODEL FOR NETWORK INTRUSION DETECTION AND MULTICLASS CYBERATTACK CLASSIFICATION
Keywords:
Network intrusion detection; CNN–LSTM; deep learning; multiclass classification; CICIDS2017; cybersecurity.Abstract
Network intrusion detection has become increasingly dependent on machine learning and deep learning techniques for identifying malicious network activity and discriminating among heterogeneous cyberattack categories. This study develops and evaluates a hybrid convolutional neural network–long short-term memory (CNN–LSTM) architecture for seven-class network intrusion detection using flow-level traffic from the CICIDS2017 dataset. The original dataset comprised 2,830,743 network-flow records with 78 traffic features and 15 labels. To construct a sufficiently represented multiclass classification problem, related attack labels were consolidated into broader attack families, while Infiltration and Heartbleed were excluded because they contained only 36 and 11 observations, respectively. A stratified experimental dataset of 29,146 records was constructed comprising BENIGN, DoS, DDoS, PortScan, Brute Force, Web Attack, and Bot classes. Following data cleaning, the dataset was partitioned into training, validation, and independent testing subsets, with feature standardization parameters estimated exclusively from the training data to prevent preprocessing-related information leakage. The proposed architecture employed one-dimensional convolutional processing to extract local patterns from the feature representation, followed by an LSTM layer to learn dependencies within the transformed feature representation. Class-weighted categorical cross-entropy was employed to mitigate class imbalance during model training. On the held-out test set, the CNN–LSTM achieved an accuracy of 82.20%, macro-precision of 82.25%, macro-recall of 83.90%, and macro-F1 score of 82.74%. Class-level analysis showed the highest F1 score for PortScan (96.17%), followed by Brute Force (94.34%), Web Attack (89.28%), and Bot (87.27%). Under the same experimental framework, the hybrid architecture outperformed the evaluated CNN-only and LSTM-only baselines. These findings indicate that integrating convolutional feature extraction with recurrent representation learning can improve multiclass intrusion classification under the evaluated configuration, although comparatively lower performance for BENIGN and DoS highlights persistent challenges in discriminating overlapping network-traffic patterns.


