BEYOND MODEL COMPLEXITY: A DATA-CENTRIC APPROACH TO RELIABLE AI-DRIVEN CYBERSECURITY PREDICTION
Keywords:
data-centric AI; cybersecurity; intrusion detection; UNSW-NB15; CIC-IDS2017; class imbalance; data leakage; temporal validation; XGBoost; explainability; reliabilityAbstract
Reliable AI-driven cybersecurity prediction depends not only on model architecture but also on how data are constructed, cleaned, balanced, split, and validated. This study develops a data-centric evaluation framework and tests it on two widely used intrusion-detection benchmarks, UNSW-NB15 and CIC-IDS2017. The primary UNSW-NB15 analysis retains the official 175,341-record training and 82,332-record testing partitions; additional uncertainty, intervention, and attack-level experiments use a fixed stratified analysis sample of 80,000 training and 40,000 testing records for computational reproducibility. For CIC-IDS2017, all eight machine-learning CSV files were audited to obtain exact full-file class counts, and a deterministic 103,245-record sample was used for controlled modeling. On UNSW-NB15, class-aware XGBoost increased F1 from 0.898 to 0.926 and reduced the false-positive rate from 25.6% to 14.3% in the controlled sample. Bootstrap analysis estimated an F1 improvement of 0.0278 (95% CI 0.0258-0.0298), while McNemar testing showed a highly significant change in paired classification errors (p < 0.001). Undersampling produced a comparable F1 of 0.926, and training-size analysis showed monotonic gains as usable data increased. Attack-level evaluation revealed that aggregate performance masked weaker recall for Fuzzers (0.743). On CIC-IDS2017, a random split yielded near-saturated F1 (0.996-0.997), whereas a Friday-afternoon temporal holdout reduced class-aware XGBoost F1 to 0.433 and attack recall to 0.277 despite ROC-AUC of 0.918. This contrast demonstrates that apparent model excellence can be dominated by data partitioning and distribution shift. The results support a practical conclusion: trustworthy cybersecurity AI requires data provenance, leakage control, class-aware learning, uncertainty analysis, attack-level auditing, and temporal or cross-dataset validation before additional model complexity is interpreted as progress.


